DevOps Engineer – AI and ML Platform

Vallum Associates Limited — United Kingdom · Posted ~11 hours ago

Senior

Skills

Python Kubernetes OpenShift Terraform Backstage Internal Developer Platforms Continuous Delivery Observability Metrics and Telemetry Service Level Objectives Runbook Automation Platform Engineering Google Cloud Platform Containerization KServe Vertex AI Endpoints Model Registry Management Feature Registry Management Model Monitoring Drift Detection GPU Scheduling Canary Deployments Shadow Deployments Troubleshooting Root Cause Analysis GCP Vertex AI Docker CI/CD

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An engineering organization is looking for an experienced DevOps Engineer to build and operate cloud-based AI and machine learning platforms. Responsibilities include developing Python automation, deploying containerized models, managing registries and model monitoring, scheduling GPU workloads, and implementing progressive deployment strategies. Strong Kubernetes or OpenShift engineering experience, infrastructure automation, continuous delivery, and observability expertise are essential. The role emphasizes scalable platform architecture and reliable production operations.

Highlights

Work on advanced AI and machine learning infrastructure, production model serving, GPU scheduling, automated deployments, and scalable cloud platforms. Collaborate across engineering teams while improving platform reliability and developer productivity.

Description

The Role DevOps Engineer will be part of the Engineering Team, who will be responsible for automating workflows, managing cloud infrastructure and streamlining software delivery. Your responsibilities: Working on Python programs – new development and enhancements.Build transformer / Agent architecture on cloud platform (GCP)Packaging and serving models using containers and KServe/Vertex endpoints, as well as Working with feature and model registries, model monitoring and drift detectionGPU SchedulingCanary and Shadow deployments.Collaborate with cross-functional teams to understand requirements and translate them into platform solutions.Troubleshoot platform and deployment issues, perform root-cause analysis and implement permanent fixesYour Profile Essential skills/knowledge/experience: 6+ yrs Experience Kubernetes/OpenShift engineeringTerraform Backstage/IDPContinuous delivery, observability (inc. metrics/telemetry/SLO/run book automation)Platform engineering, control engineeringDesirable skills/knowledge/experience: (As applicable) Good scripting and automation skills using PythonJIRAGood communication and leadership abilities.Learning Attitude