Machine Learning Engineer (3D Human Pose Egocentric)
MaxInsights · San Jose, US
Job description
Job Description:
We are looking for a Machine Learning Engineer to join our core research and development team, focused on recovering accurate 3D human body and hand motion from egocentric (first-person) video.
Human demonstration data is the fuel for robot learning, and the quality of that data is bounded by how well we can reconstruct what the hands and body actually did. In this role, you will own models and pipelines that turn head-mounted and body-mounted camera streams — often wide-FOV, stereo, motion-blurred, and heavily self-occluded — into metrically accurate, temporally stable 3D pose that is directly usable for robot policy training and human-to-robot retargeting.
You will work across the full stack: capture rig and calibration, ground-truth annotation tooling, model training and evaluation, and production deployment at scale. This role suits engineers who are equally comfortable with multi-view geometry and modern deep learning, and who are motivated by hard, measurable accuracy problems on real-world data.
Responsibilities
- Build 3D body and hand pose estimation models for egocentric video, covering 2D/3D keypoints, parametric body and hand models (SMPL/SMPL-X, MANO), and full-sequence motion recovery from monocular and stereo first-person cameras.
- Solve the hard cases specific to the egocentric viewpoint — severe self-occlusion, truncated limbs, extreme perspective foreshortening, hand–object interaction, rapid head motion, and rolling-shutter and motion-blur artifacts.
- Own camera geometry and calibration: fisheye and wide-FOV camera models (Kannala-Brandt, Double Sphere), intrinsic/extrinsic calibration, stereo triangulation, and head-to-body coordinate-frame alignment for metric-scale output.
- Drive temporal consistency and physical plausibility through robust estimation, smoothing and filtering, kinematic and anatomical constraints, contact and penetration reasoning, and multi-view or multi-modal fusion (e.g. IMU, exocentric cameras, marker-based mocap).
- Build the ground-truth and evaluation loop: semi-automatic annotation and keypoint propagation tools, confidence-aware quality gating, and evaluation protocols that separate real accuracy gains from benchmark noise.
- Ship end-to-end systems at scale — large-scale training, high-throughput video inference, and reliable production pipelines over high-bandwidth multi-camera data.
- Translate reconstructed human motion into robot-usable data, collaborating with robotics and product teams on retargeting fidelity for dexterous hands and humanoid end-effectors.
- Contribute to technical design, code quality, and best practices, and help shape the long-term direction of the company’s perception stack.
Minimum Qualifications
- Bachelor’s, Master’s, or PhD in Computer Science, Machine Learning, Computer Vision, Robotics, or a related technical field, or equivalent practical experience.
- 3+ years of experience building and shipping machine learning systems.
- Proven hands-on experience developing and deploying 3D human pose, hand pose, or human motion tracking models from video.
- Working knowledge of multi-view geometry and camera models: projection, calibration, triangulation, rigid-body transforms, and coordinate-frame management.
- Strong proficiency in Python and at least one major deep learning framework (e.g. PyTorch, TensorFlow).
- Solid understanding of modern deep learning concepts, training workflows, model evaluation, and real-world, production-oriented ML pipelines.
- Strong problem-solving skills and the ability to work effectively in a fast-moving, collaborative environment.
Preferred Qualifications
- Direct experience with egocentric or head-mounted perception (AR/VR headsets, smart glasses, chest- or head-mounted capture rigs), including fisheye and stereo pipelines.
- Deep expertise in human kinematics and parametric models — SMPL/SMPL-X, MANO, inverse kinematics, markerless motion capture, and hand–object pose estimation.
- Familiarity with relevant egocentric vision datasets and benchmarks.
- Familiarity with state-of-the-art architectures for video and 3D data (e.g. video transformers, diffusion-based motion priors, 3D CNNs, etc).
- Experience building or operating multi-camera capture systems, time synchronization, and calibration infrastructure.
- Experience with human-to-robot motion retargeting, teleoperation data, or imitation learning pipelines.
- Experience with annotation tooling, active learning, or data quality systems for large-scale video.
- Publications at leading venues (CVPR, ICCV, ECCV, NeurIPS, SIGGRAPH, 3DV), open-source contributions, or demonstrated impact in applied ML or AI systems.
What We Offer
- Competitive salary and options package.
- Comprehensive health, dental, and vision insurance.
- 401(k) plan.
- Paid time off.
- Direct collaboration with leading experts in the field of robotics and AI.
MaxInsights is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
Default Benefits:
- Health insurance
- Vision care
- Dental coverage:
- 401(k)
- Paid holidays
- PTO (Paid Time Off)
- Sick leave
ML/AI Work links you to the employer's original posting — always verify the details there before applying.
More Machine Learning roles
View all →Lead Product Manager - AI Governance & Adoption
Tesco · Remote · Kettering
AI Product Analyst - Integration and Quality
Chubb Insurance · Markham, CA
AI Delivery Coordinator H/F
Lectra · Bordeaux, FR
AI Product Analyst
Chubb Insurance · Markham, CA
AI Enablement Lead
Broadstone Corporate Benefits · Colchester, GB
AI Delivery & Operations Lead
The AA · Remote · Colchester