Senior MLOps / ML Platform Engineer

Sigma Software Page — Poland · Posted ~3 hours ago

Senior Full-time

Skills

MLOps ML platform engineering ML orchestration model lifecycle automation observability production engineering scalable infrastructure Machine Learning ML Platforms Observability

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology organization is seeking a Senior MLOps / ML Platform Engineer to build production-grade machine learning infrastructure for a high-scale advertising environment. The role covers ML orchestration, model lifecycle automation, observability, and real-time optimization, with substantial opportunities to influence architecture and solve complex operational challenges.

Highlights

Senior role focused on production-grade ML infrastructure, large-scale workloads, architecture influence, observability, model automation, and collaboration with experienced engineers on a long-term strategic engagement.

Description

Company Description We are looking for a Senior MLOps Engineer to join Sigma Software and help build a production-grade ML platform for a large-scale AdTech ecosystem. You will work on infrastructure powering predictive decision-making systems that process hundreds of millions of auction requests daily. As part of a dedicated engineering team, you will contribute to scalable ML orchestration, model lifecycle automation, observability, and real-time optimization workflows. This role is ideal for engineers with strong production experience who enjoy solving complex platform and operational challenges. We at Sigma Software offer the opportunity to work on cutting-edge ML infrastructure projects, collaborate with experienced engineers, and influence architecture decisions in a long-term strategic engagement. CUSTOMER Our Customer is a technology company operating supply-side infrastructure within the programmatic advertising ecosystem. The company manages a high-load ad exchange platform handling hundreds of millions of auction requests every day and is investing in advanced predictive decision-making capabilities to improve advertiser outcomes and real-time optimization processes. PROJECT Sigma Software is building a predictive modeling and optimization platform integrated with a live ad exchange environment. The solution enables real-time supply scoring and filtering, audience look-alike generation, contextual performance estimation, and multi-objective optimization under operational constraints. The project combines large-scale ML infrastructure, automated model lifecycle management, multi-tenant architecture, and advanced observability practices. The team focuses on delivering reliable, reproducible, and scalable ML systems ready for long-term Customer ownership. Key Technologies: Python, Kubernetes, Docker, GCP, Vertex AI, MLflow, Airflow, Kubeflow, Argo Workflows, Terraform Job Description Build and maintain ML training orchestration pipelines across hourly, daily, and weekly schedulesImplement retries, backfills, and idempotent execution mechanismsDesign and support model registry workflows including versioning, lineage, evaluation gates, and promotion processesDevelop isolated per-advertiser model environments with namespace and configuration separationBuild scalable refresh pipelines and publishing workflows for serving infrastructureImplement shadow mode and champion/challenger deployment strategiesDevelop monitoring and alerting for ML-specific metrics including feature drift, prediction drift, train/serve skew, and calibration decayEnsure reproducibility of ML workflows using containerized environments, pinned dependencies, and data snapshotsMonitor training and scoring costs across tenantsCollaborate with DevOps and SRE engineers on CI/CD and infrastructure automationPrepare operational documentation and platform handover materials Qualifications 5+ years of experience in MLOps, ML platform engineering, or infrastructure engineering supporting production ML systemsStrong Python skills and experience building platform-level tooling and automationHands-on experience with Kubernetes and DockerExperience building CI/CD pipelines for ML workloadsHands-on production experience with MLflow, Kubeflow, Airflow, Argo Workflows, Vertex Pipelines, or similar orchestration and ML lifecycle platformsExperience with ML platforms and model lifecycle tools such as Vertex AI, MLflow, or KubeflowStrong understanding of ML observability including drift detection, train/serve skew monitoring, and incident responseExperience designing or supporting multi-tenant ML systems and isolated model environmentsExperience working with cloud platforms, preferably GCPExperience with infrastructure-as-code tools such as TerraformExperience with Linux environmentsUnderstanding of the ML lifecycle and productionization processesUpper-Intermediate English level or higher WILL BE A PLUS Experience with feature stores and feature consistency managementExperience with large-scale batch scoring systems operating under freshness SLAsFamiliarity with experiment tracking platforms and evaluation gatesExperience with on-premises Kubernetes or bare-metal Linux infrastructureKnowledge of DVC, lakeFS, or other data versioning toolsExperience with Bigtable, Redis, Aerospike, or similar low-latency serving databasesGPU scheduling and training cost optimization experienceFamiliarity with SOC 2, ISO 27001, or GDPR-related compliance requirements Additional Information PERSONAL PROFILE Strong ownership mindset and focus on operational reliabilityAbility to work independently in complex distributed systems environmentsStrong collaboration and communication skillsAnalytical thinking with attention to scalability and maintainabilityComfortable working in fast-paced product-oriented environments