Senior Machine Learning Inference Engineer

Oscar — United States · Posted ~2 hours ago

Senior Full-time Hybrid $250K base + equity

Skills

Python PyTorch GPU infrastructure machine learning inference model serving microservices GPU Triton TensorRT vLLM

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior AI engineering position focused on optimizing inference systems for advanced generative and multimodal models, building scalable serving infrastructure, and improving production performance.

Highlights

High-impact AI infrastructure role with strong compensation, equity opportunities, and ownership over production-scale model performance.

Description

Title: ML Inference Engineer Location: San Francisco, CA Salary: $250k base + equity An AI Unicorn startup is hiring a Senior Machine Learning Inference Engineer for a full-time role. You will be responsible for improving efficiency for AI-native infrastructure powered by generative and multimodal models. The ideal candidate has over 3 years of professional experience and a strong understanding of GPU infrastructure, Python, and PyTorch. This is a highly autonomous role with significant ownership across inference systems and model performance in production. This role is hybrid in San Francisco Bay Area and offers full benefits and equity. Experience: Building AI applications at scale from the ground upStrong understanding of GPU infrastructure including Triton, TensorRT, or vLLM frameworksHands-on experience with Python and PyTorchBuilding model-serving MicroservicesDiffusion and Multimodal model experience is a plus Benefits: Competitive base salaryEquity$401k matchingMedical coverage Oscar Associates Limited (US) is acting as an Employment Agency in relation to this vacancy.