MLOps / Machine Learning Platform Engineer

Remotehunter — United States · Posted ~3 hours ago

Mid Full-time

Skills

MLOps machine learning infrastructure model deployment ML pipelines machine learning

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A machine learning platform engineering role focused on infrastructure for training, evaluation, deployment, and monitoring of AI models. The role combines software engineering and ML operations.

Highlights

Opportunity to build scalable AI infrastructure and improve the reliability and efficiency of machine learning workflows.

Description

About Our Client: This team operates within the machine learning and infrastructure engineering space, addressing the challenges of managing scalable, reliable, and efficient machine learning workflows. They focus on building and maintaining infrastructure that supports high-throughput model training, evaluation, and production deployment, ensuring operational excellence and cost efficiency. About the Opportunity: The MLOPS / ML PLATFORM ENGINEER is responsible for developing and maintaining the infrastructure that enables seamless machine learning workflows from experimentation to production. This role ensures that model training, serving, and monitoring systems operate efficiently and reliably, directly impacting the organization's ability to deploy scalable AI solutions with minimal friction. Responsibilities: Design and manage ML infrastructure for data handling, training, serving, and inference.Build scalable and reproducible training and evaluation pipelines with versioning and scheduling.Optimize compute resources and costs through tuning and cluster management.Operate low-latency model serving APIs with autoscaling and safe deployment strategies.Define and maintain SLOs; monitor latency, cost, drift, and data quality.Manage security aspects, including IAM, secrets, and container security; automate deployments with CI/CD and infrastructure as code.Collaborate with research scientists and AI engineers to transition models to production.Create documentation, templates, and tooling to standardize and accelerate ML workflows. Requirements: Minimum 4 years in ML platform, DevOps, or infrastructure engineering.Proficient with Kubernetes, CI/CD, containers, and cloud platforms (AWS, GCP, or Azure).Experience managing GPU clusters and ML training/inference pipelines.Knowledge of data orchestration and storage formats like Delta, Parquet, Polars, Spark.Demonstrated ability to deploy and operate production ML systems meeting SLOs.Strong Python programming skills and experience with infrastructure automation.Familiarity with observability tools and cost optimization at scale. Pay Range and Compensation Package: The pay range and compensation package for this role will be determined based on the candidate's experience, skills, and other relevant factors. Benefits & Perks: Competitive Salary and Bonus PlanComprehensive health insurance planRetirement savings plan (401k) with company matchRemote working environmentFlexible, unlimited time off policyGenerous paid holiday schedule including 13 holidays Equal Opportunity Statement: Our client is an equal opportunity employer. They celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, or national origin. Note: RemoteHunter is not the Employer of Record (EOR) for this role. Our purpose in this opportunity is to connect exceptional candidates with leading employers. We help job seekers worldwide discover roles that match their goals and guide them to complete their full application directly through the hiring company's career page or ATS.