ML/AIWork

AI Inference Engineer

Epicsoft Tech · Remote · Oakland

Job description

About the Role

We are seeking an experienced AI Inference Engineer to help build and optimize high-performance AI model serving infrastructure for production-scale Large Language Models (LLMs). You will work on GPU acceleration, distributed inference systems, Kubernetes deployments, and model optimization to deliver low-latency, scalable AI services.

This position is ideal for engineers with strong systems programming experience who enjoy solving performance challenges across GPUs, distributed computing, and AI infrastructure.

Responsibilities

  • AI Model Serving
  • Build, deploy, and optimize production AI inference services.
  • Work with frameworks such as vLLM, TensorRT-LLM, Triton Inference Server, SGLang, TorchServe, or KServe.
  • Develop scalable inference microservices using C++, Python, and Rust.
  • GPU Performance Optimization
  • Develop and optimize GPU kernels using CUDA or ROCm.
  • Improve GPU utilization, memory efficiency, and inference throughput.
  • Optimize tensor operations and hardware acceleration.
  • Large Language Model Optimization
  • Improve inference performance through:
  • Continuous batching
  • KV Cache optimization
  • Quantization
  • Tensor Parallelism
  • Pipeline Parallelism
  • Speculative Decoding
  • Mixture of Experts (MoE)
  • Distributed Infrastructure
  • Deploy and manage inference services on Kubernetes, OpenShift, or similar container platforms.
  • Support distributed multi-node, multi-GPU environments.
  • Build reliable, fault-tolerant AI serving infrastructure.
  • Performance & Reliability
  • Profile and benchmark production workloads.
  • Troubleshoot latency and throughput bottlenecks.
  • Implement monitoring, telemetry, and observability solutions.
  • Improve system reliability and scalability.

Required Qualifications

  • Bachelor's degree in Computer Science, Software Engineering, or a related field (or equivalent experience).
  • 5+ years of software engineering experience.
  • Hands-on experience building production AI inference systems.
  • Strong programming skills in:
  • C++
  • Python
  • Rust
  • Experience with one or more of the following:
  • vLLM
  • TensorRT-LLM
  • Triton Inference Server
  • SGLang
  • TorchServe
  • KServe
  • Experience with CUDA and/or ROCm GPU programming.
  • Knowledge of LLM inference optimization techniques, including:
  • KV Cache
  • Continuous Batching
  • Quantization
  • Attention Optimization
  • Experience deploying workloads on Kubernetes or OpenShift.
  • Experience working with distributed GPU infrastructure.
  • Strong debugging, benchmarking, and performance tuning skills.

Preferred Qualifications

  • Experience with NVIDIA Dynamo or similar distributed inference platforms.
  • Experience serving:
  • Large Language Models (LLMs)
  • Multimodal Models
  • Mixture of Experts (MoE)
  • Embedding Models
  • Familiarity with OpenAI-compatible APIs.
  • Experience with telemetry, monitoring, and observability tools.

Pay: $65.00 - $75.00 per hour

Experience:

  • CUDA/ROCm: 1 year (Required)
  • production LLM inference: 3 years (Required)
  • vLLM/TensorRT-LLM/Triton production: 1 year (Required)

Ability to Relocate:

  • San Jose, CA 95139: Relocate before starting work (Preferred)

Work Location: Hybrid remote in San Jose, CA 95139

ML/AI Work links you to the employer's original posting — always verify the details there before applying.

More Domain Specializations roles

View all →
$135,200 – $156,000/yr
Epicsoft Tech
Apply →