Summary
✨ AI‑Generated
Build and operate a scalable AI-powered autolabeling pipeline that helps human reviewers process complex visual classification tasks more efficiently. Integrate foundation-model APIs, process structured predictions, connect data workflows, build detailed observability, and run experiments to improve accuracy, latency, cost, and coverage.
Highlights
Engineering role building an AI-powered autolabeling pipeline at fleet scale, with ownership across model integration, data workflows, observability, experimentation, and performance optimization.
Description
Role Overview
Build and operate the autolabeling pipeline that accelerates human annotation throughput for vehicle attribute classification tasks.Develop a pipeline using foundation models such as Gemini, SigLIP, and CLIP to pre-label tasks for human reviewers to verify and correct.Own pipeline engineering, including ingesting queued tasks from the annotator service, calling foundation-model APIs at fleet scale, parsing structured predictions, and writing pre-labels back into the labeling workflow.Partner closely with the team lead, ML engineers, and data infrastructure team to integrate the pipeline with existing Zoox systems.Responsibilities
Build the autolabeling pipeline to ingest queued annotation tasks, dispatch them to foundation-model APIs, parse structured outputs, and write pre-labels back to the labeling workflow.Build the observability layer covering per-task latency, per-model cost, per-attribute coverage, and error-mode dashboards.Set up and execute experiments designed by the team lead and collect outputs in formats suitable for ML engineers to analyze.Integrate the pipeline with existing Zoox systems in partnership with the data infrastructure team.Document the system, write runbooks, and ensure a clean handoff at the end of the engagement.Required Qualifications
3+ years of backend or data pipeline engineering experience.Strong Python skills with comfort in C++.Large-dataset experience using PySpark or equivalent.Understanding of ML fundamentals including model inference, embeddings, structured output, precision, recall, and calibration.Ability to reason about ML data shapes and integration patterns.Experience integrating foundation models such as Gemini, OpenAI, or Anthropic at production scale.Excellent written communication skills for design documents and runbooks.Bonus Qualifications
Databricks experience.End-to-end ML pipeline stewardship from data ingest through inference and monitoring.Experience with annotation tooling or human-in-the-loop ML workflows.Experience with autonomous-systems data pipelines.AWS experience, especially S3, ECS/EKS, and Lambda.Experience working in a shared codebase with ML engineers, including proto schemas and joint deployments.EEO: All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.
Cell phones are not allowed at the client site, shooting of photos and/or videos are strictly prohibited while inside the client facility.