ML/AIWork
Oracle logo

Senior Core Infrastructure Software Engineer

Oracle · Nashville, US

Job description

The Senior Core Infrastructure Software Engineer is a Senior experienced, hands-on software engineer role responsible for designing, building, and operating next-generation AI systems on Oracle Cloud Infrastructure (OCI). This person will work on building production-grade cloud distributed systems, agentic AI platforms, autonomous workflows, scalable inference infrastructure, and enterprise AI applications used in large-scale, business-critical environments.

This role owns moderately complex components within OCI Cloud platform services or SDKs; leads team-level improvements to integration frameworks and developer tooling. Performs deep debugging across a bounded set of distributed services, driving fixes that protect downstream consumers and upgrade paths. Analyzes usage, performance, and error budgets for specific platform surfaces; implements targeted resilience and capacity optimizations. Authors and curates team-scoped documentation, samples, and adoption guidance. The ideal candidate combines deep distributed systems experience with practical AI-native engineering.

Key Responsibilities include:

  • Implements and contributes to the development for components of distributed systems that support horizontal and vertical scaling including leveraging distributed state management tools.
  • Design, implement, and deliver scalable agentic AI systems capable of reasoning, planning, tool use, workflow execution, multi-step task orchestration, and safe human-in-the-loop escalation.
  • Implement scalability, performance, availability, and durability requirements for owned services and components.
  • Designs and implements functional requirements and testing for assigned features within an existing system
  • Optimize services for high-throughput, low-latency, and large-scale cloud workloads.
  • Implement systems that handle network unreliability, service disruptions, and partition scenarios while meeting SLOs.
  • Establish telemetry, KPIs, dashboards, and alerts to monitor service health, performance, reliability, and customer impact.
  • Drive performance testing, load testing, fault injection, brownout testing, and other validation strategies to ensure correctness and resiliency.
  • Implement secure infrastructure controls for multi-tenant cloud environments, including access controls, encryption, and remediation of security gaps.
  • Develop automation, Infrastructure as Code, and deployment tooling to support safe patching, updates, rollbacks, and operational recovery.
  • Designs and implements automation scripts and tooling used to troubleshoot operational issues.
  • Take ownership of production operations, including troubleshooting, incident response, root cause analysis, and ongoing service improvements.
  • Adheres to change management plans for patching, updating, and rolling back applications.
  • Strengthen operational readiness by improving runbooks, monitoring, change management, deployment safety, and recovery processes.
  • Applies advanced security measures to protect data and applications in multi-tenant environments, including encryption and access controls.
  • Collaborates with the team to ensure cloud infrastructure complies with relevant industry standards and regulations and that documentation is up-to-date

Required Qualifications:

  • Bachelor's, Master's, or Ph.D. in Computer Science, AI/ML, Engineering, or a related field, or equivalent practical experience.
  • 3-7+ years of professional software engineering experience
  • Experience in designing and developing high-scale distributed systems, cloud services, infrastructure platforms, or AI/ML platform services.
  • Strong programming skills in Java or Golang or Python and ability to contribute high-quality production code, reviews, tests, and debugging in complex distributed environments.
  • Strong expertise with Kubernetes, Docker, cloud-native infrastructure, service-to-service communication, scalability, fault tolerance, observability, and performance analysis.
  • Agile Methodologies: Demonstrated ability to use agile methodologies to drive continuous improvement and product delivery
  • Excellent written and verbal communication

Preferred Qualifications:

  • Experience with large scale cloud platforms (e.g., AWS, Azure, Google, Oracle Cloud).
  • Practical experience with orchestration frameworks such as LangGraph, LangChain, CrewAI, AutoGen, LlamaIndex, or similar ecosystems.
  • Understanding of LLM application patterns, including prompt design, structured outputs, function/tool calling, context management, RAG, memory, tool safety, and evaluation.
  • Experience using AI-assisted software development tools such as Codex, Claude Code, Cursor, Copilot, or similar systems in large-scale engineering environments

ML/AI Work links you to the employer's original posting — always verify the details there before applying.

More Generative AI and LLM roles

View all →
Senior Core Infrastructure Software Engineer
Oracle
Apply →