ML/AIWork

Senior/Staff Applied AI Engineer, Agent Harness

David Joseph & Company · San Francisco, US

Job description

San Francisco, CA · On-site · Full-time

Compensation: $200,000–$300,000 + 1%–2% equity

About the Company

Our client is a seed-stage startup building AI “coworkers” for IT teams — a security-focused product in which an agent operates as a governed identity inside a customer’s environment, requests scoped access on a per-task basis, escalates to a human for approval, and can operate systems it was never handed an API for. The company is well funded for its stage and backed by a strong group of operators and angels from across the IT and security world, with several design partners already live. You’d be one of the earliest engineers alongside a small founding team.

Founded 2026 · 1–10 people · Industry: AI tooling

The Role

You’ll build the layer that turns raw model capability into systems that reliably work for real users. That means developing the core agent harness — the execution loop, tool-use strategies, context construction, and model-facing experimentation — and iterating on agent behavior across real customer workflows and long-horizon tasks. A guiding principle here: the creative work happens once, at authoring time, and what runs afterward is deterministic, compiled, type-checked code rather than stochastic tool-chaining. You’ll own that boundary between what the model decides at runtime and what ships as code.

What you’ll be doing

  • Own the core agent harness: execution loop, tool-use strategies, context construction, and model-facing experimentation
  • Shape agent behavior across real customer workflows and long-horizon tasks
  • Own the line between runtime model decisions and deterministic, compiled execution
  • Build and run evals against replicas of real customer environments — reliable enough to gate a release
  • Diagnose production failures and attribute them to the right layer (model, prompt, tool contract, environment state, retry logic), then harden the system
  • Extend computer-use to systems that have no API, with real production guarantees (isolated per-task sandbox, model-blind credential handling, full action recording, killable sessions)
  • Build the feedback loops and data systems that feed better real-task data into eval and training

Tech stack: Python or TypeScript, modern AI tooling, agent-harness / execution-loop engineering, evals, sandboxed (microVM) execution, computer-use / browser automation.

Requirements

  • Hands-on experience building with agent frameworks or tool-using LLM systems
  • Strong programming ability in Python or TypeScript, and fluency with current AI tooling
  • Practical experience across model evaluation, fine-tuning, and/or prompt design
  • Comfortable owning systems end-to-end and debugging across the full stack
  • Frame problems around systems and real user outcomes, not just model benchmarks
  • Genuinely enjoy chasing down messy real-world failures and turning them into durable fixes
  • Motivated to work at the layer where model capability becomes dependable product
  • Able to work on-site in San Francisco, Monday–Friday, at startup pace

Nice to Haves

  • A standout signal of excellence somewhere — leading a meaningful team at a high-growth company, top competitive-programming results, or elite achievement in a craft, sport, or game
  • A CS or engineering degree from a top program, or equivalent genuinely exceptional experience
  • Time at a strong company on a high-caliber, directly relevant team (agents / AI infrastructure)
  • A track record of shipping agents that actually hold up in production
  • High agency — former founders, broad scope, comfort operating amid real ambiguity
  • Experience building computer-use or browser-automation agents
  • Familiarity with virtualization and sandboxed execution environments, and scaling them
  • Published AI research at top conferences
  • Shipped systems where correctness had to survive partial failure (idempotency, resumability, compensating actions)
  • Integrated enterprise SaaS APIs (identity providers, directory services, ticketing, ERPs)
  • Comfort working with large, messy datasets or production logs
  • Early-engineer experience in a seat where you also talked to customers

Why Join

  • Ground-floor ownership of the core agent harness at a well-backed, security-focused AI startup
  • A meaningful equity stake (1%–2%) at the seed stage
  • Hard, real engineering — production-grade evals, isolated execution, and computer-use on systems with no API, not throwaway demos
  • Small founding team, broad scope, and direct impact on what ships

Details

  • Location — San Francisco, CA
  • Work policy — On-site, Monday–Friday, startup intensity
  • Compensation — $200,000–$300,000 + 1%–2% equity
  • Visa sponsorship — Available (H-1B, O-1, OPT)
  • Employment type — Full-time
  • Benefits — Meals in office, health insurance, unlimited PTO

ML/AI Work links you to the employer's original posting — always verify the details there before applying.

More AI Safety and Evaluation roles

View all →
$200,000 – $300,000/yr
David Joseph & Company
Apply →