Summary
✨ AI‑Generated
A senior leadership position for an engineer responsible for designing and operating AI and machine learning infrastructure. The role covers cloud platforms, model lifecycle automation, reliability, and scalable production AI systems.
Highlights
Highly visible leadership role owning AI platform infrastructure, cloud architecture, and machine learning operations with strong technical ownership.
Description
SVP, Lead AI/MLOps Infrastructure Engineer – Full Time – Hybrid
We’re partnering with our client, a fast-growing fintech firm, on a senior-level hire to lead and own the infrastructure behind their AI and machine learning platforms.
This is a highly visible, hands-on leadership role where you’ll own the end-to-end AI/ML platform stack, from training and inference infrastructure to model serving, reliability, and cost, while setting the MLOps roadmap and standards for the team.
The priority here is MLOps and AI infrastructure first, built on deep AWS, Terraform, and production ML / GenAI experience.
What You’ll Be Doing
Own the end-to-end AI/ML platform stack: orchestration, compute (including GPU), storage, and model servingBuild and operate MLOps pipelines across the full model lifecycle: training, validation, versioning, and deploymentProductionize AI/ML and GenAI (LLM) workloads in partnership with ML engineers and data scientistsDesign and manage cloud-native AWS infrastructure using Kubernetes, and own Infrastructure as Code standards (Terraform)Own SLAs/SLOs for model serving and inference; lead monitoring, drift detection, and incident responseBuild internal tooling and standardized environments that boost ML engineer productivity (e.g., MLflow, Kubeflow, Weights & Biases, Ray)Drive data governance, privacy compliance, and cost optimization across training and inference workloadsSet the MLOps roadmap and mentor engineers on infrastructure and MLOps best practices
What They’re Looking For
15+ years of experience in DevOps, SRE, or platform engineering, with AWS as primary cloudProven, hands-on experience building and operating MLOps pipelines in production (key priority)Experience with MLOps tooling, including model registries, experiment tracking, and feature storesExposure to Generative AI / LLM workloads, including AWS BedrockStrong Infrastructure as Code (Terraform) and scripting skills (Python or similar)Solid Linux, systems, and troubleshooting fundamentalsExcellent communicator, comfortable collaborating across teams
Nice to Have
Hands-on Kubernetes, containerized workloads, and cloud networkingExperience in regulated or fintech environmentsBackground optimizing costs for compute-intensive (GPU) workloads
Why This Role
Own and shape the AI platform at a growing fintechHigh-impact leadership role setting MLOps strategy and standardsStrong compensation: $200K–$230K base + bonus + equityComprehensive benefits, including retirement match and unlimited PTOHybrid model: 4 days onsite / 1 day remote (NYC area)
By applying for this job, you agree to receive calls, AI-generated calls, text messages, or emails from Benchmark IT, LLC and its affiliates, and contracted partners.
Frequency varies for text messages.
Message and data rates may apply.
Carriers are not liable for delayed or undelivered messages.
You can reply STOP to cancel and HELP for help.
You can access our privacy policy here: https://bmarkits.com/privacy-policy/