Senior/Staff Applied AI Engineer, Agent Harness
David Joseph & Company · San Francisco, US
Job description
San Francisco, CA · On-site · Full-time
Compensation: $200,000–$300,000 + 1%–2% equity
About the Company
Our client is a seed-stage startup building AI “coworkers” for IT teams — a security-focused product in which an agent operates as a governed identity inside a customer’s environment, requests scoped access on a per-task basis, escalates to a human for approval, and can operate systems it was never handed an API for. The company is well funded for its stage and backed by a strong group of operators and angels from across the IT and security world, with several design partners already live. You’d be one of the earliest engineers alongside a small founding team.
Founded 2026 · 1–10 people · Industry: AI tooling
The Role
You’ll build the layer that turns raw model capability into systems that reliably work for real users. That means developing the core agent harness — the execution loop, tool-use strategies, context construction, and model-facing experimentation — and iterating on agent behavior across real customer workflows and long-horizon tasks. A guiding principle here: the creative work happens once, at authoring time, and what runs afterward is deterministic, compiled, type-checked code rather than stochastic tool-chaining. You’ll own that boundary between what the model decides at runtime and what ships as code.
What you’ll be doing
- Own the core agent harness: execution loop, tool-use strategies, context construction, and model-facing experimentation
- Shape agent behavior across real customer workflows and long-horizon tasks
- Own the line between runtime model decisions and deterministic, compiled execution
- Build and run evals against replicas of real customer environments — reliable enough to gate a release
- Diagnose production failures and attribute them to the right layer (model, prompt, tool contract, environment state, retry logic), then harden the system
- Extend computer-use to systems that have no API, with real production guarantees (isolated per-task sandbox, model-blind credential handling, full action recording, killable sessions)
- Build the feedback loops and data systems that feed better real-task data into eval and training
Tech stack: Python or TypeScript, modern AI tooling, agent-harness / execution-loop engineering, evals, sandboxed (microVM) execution, computer-use / browser automation.
Requirements
- Hands-on experience building with agent frameworks or tool-using LLM systems
- Strong programming ability in Python or TypeScript, and fluency with current AI tooling
- Practical experience across model evaluation, fine-tuning, and/or prompt design
- Comfortable owning systems end-to-end and debugging across the full stack
- Frame problems around systems and real user outcomes, not just model benchmarks
- Genuinely enjoy chasing down messy real-world failures and turning them into durable fixes
- Motivated to work at the layer where model capability becomes dependable product
- Able to work on-site in San Francisco, Monday–Friday, at startup pace
Nice to Haves
- A standout signal of excellence somewhere — leading a meaningful team at a high-growth company, top competitive-programming results, or elite achievement in a craft, sport, or game
- A CS or engineering degree from a top program, or equivalent genuinely exceptional experience
- Time at a strong company on a high-caliber, directly relevant team (agents / AI infrastructure)
- A track record of shipping agents that actually hold up in production
- High agency — former founders, broad scope, comfort operating amid real ambiguity
- Experience building computer-use or browser-automation agents
- Familiarity with virtualization and sandboxed execution environments, and scaling them
- Published AI research at top conferences
- Shipped systems where correctness had to survive partial failure (idempotency, resumability, compensating actions)
- Integrated enterprise SaaS APIs (identity providers, directory services, ticketing, ERPs)
- Comfort working with large, messy datasets or production logs
- Early-engineer experience in a seat where you also talked to customers
Why Join
- Ground-floor ownership of the core agent harness at a well-backed, security-focused AI startup
- A meaningful equity stake (1%–2%) at the seed stage
- Hard, real engineering — production-grade evals, isolated execution, and computer-use on systems with no API, not throwaway demos
- Small founding team, broad scope, and direct impact on what ships
Details
- Location — San Francisco, CA
- Work policy — On-site, Monday–Friday, startup intensity
- Compensation — $200,000–$300,000 + 1%–2% equity
- Visa sponsorship — Available (H-1B, O-1, OPT)
- Employment type — Full-time
- Benefits — Meals in office, health insurance, unlimited PTO
ML/AI Work links you to the employer's original posting — always verify the details there before applying.
More AI Safety and Evaluation roles
View all →KI-Trainer / KI-Dozent für Generative KI (m/w/d) – Freelancer
ORYZAR AI GbR · Gelsenkirchen, DE
Senior Manager, AI Solutions
— · Remote · Houston
Agent Engineer
Rifa AI · Remote · San Francisco
Senior Software Engineer, Analytics Data & Applied AI
Unity Technologies · San Francisco, US
AI Product Engineer
NEHO · Remote · Vernier
AI Engineer
Accountor · Espoo, FI