roboforce

Senior / Staff AI Research Engineer, Data Infrastructure

Apply Now

At a Glance

Location
Milpitas, California, United States
Experience
5+ years
Posted
2026-04-13T14:38:32-04:00

Key Requirements

Required Skills

Deep LearningETLPyTorchPython

Domain Knowledge

  • Robotics

Benefits & Perks

Health Insurance

Health, dental, and vision insurance, 401(k) plan.

Requirements

Strong proficiency in Python and experience building production-grade data pipelines and ETL systems.

Hands-on experience with large-scale dataset management, including versioning, deduplication, quality filtering, and distributed storage (e.g., S3, GCS, HDF5, WebDataset, Zarr).

Experience building or working with post-training infrastructure — SFT pipelines, reward modeling, or RL training loops (e.g., PPO, DPO, rejection sampling).

Familiarity with deep learning frameworks (PyTorch, JAX) and ML training workflows sufficient to collaborate tightly with research teams.

Experience with robotics data collection hardware — teleoperation devices, UMI, GELLO, or similar — and the synchronization and preprocessing challenges they introduce.

Familiarity with robot learning pipelines: imitation learning, behavior cloning, or VLA/VLM fine-tuning workflows.

Compensation & Benefits

Competitive stock options/equity programs.

Health, dental, and vision insurance, 401(k) plan.

Responsibilities

Design and maintain end-to-end data collection pipelines ingesting multimodal demonstration data from teleoperation devices and UMI hardware, including synchronization, versioning, and distributed storage at scale.

Build annotation tooling and data curation workflows — quality filtering, deduplication, episode scoring, and domain reweighting — to produce high-quality training datasets for robot policy learning.

Develop post-SFT reinforcement learning infrastructure: implement reward scoring on demonstrations, mine and categorize failure patterns, and feed curated failure data back into the retraining loop.

Build evaluation and test infrastructure to log policy rollouts on-robot, capture structured results, and surface actionable diagnostics for the research team.

Collaborate with ML researchers to define data schemas, episode formats, and pipeline interfaces that support rapid iteration on VLA and manipulation policy training.

Architect scalable storage and retrieval systems for heterogeneous robot data (vision, proprioception, action, language) across both cloud and on-prem environments.