Senior AI Platform Engineer - Consultant

Stealth It Consulting — United Kingdom · Posted ~20 hours ago

Senior Other

Skills

AI platform engineering Cloud infrastructure Hybrid and multi-cloud environments GPU-accelerated computing Kubernetes OpenShift Container platforms Model serving MLOps LLMOps Generative AI infrastructure GPU computing Cloud-native services

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Design, build, and operate the platform infrastructure that supports enterprise AI and generative AI workloads. You will architect GPU-accelerated compute, container platforms, model-serving infrastructure, evaluation and guardrail systems, and MLOps/LLMOps tooling across hybrid and multi-cloud environments. As a senior consultant, you will lead technical delivery and help organizations modernize AI infrastructure and adopt AI safely at scale.

Highlights

Senior consulting opportunity focused on designing and operating infrastructure for enterprise AI and generative AI workloads. The role spans GPU computing, Kubernetes/OpenShift, model serving, MLOps/LLMOps, evaluation and guardrail systems, and hybrid or multi-cloud environments, with bonus and benefits included.

Description

AI Platform Engineer - Senior Consultant - SC Eligibility required - Permanent Locations: London, Manchester, Glasgow + other UK locations Salary: Up to £70,000 + Bonus & Benefits Active SC or SC Eligibility essential As an AI Platform Engineer, you’ll design, build, and operate the infrastructure that enterprise AI and Generative AI workloads run on: the platform layer beneath LLMs, agents, and MLOps pipelines. This spans GPU-accelerated compute and container platforms, model serving and gateway infrastructure, evaluation and guardrail systems, and the MLOps/LLMOps tooling that takes a model from experiment to production. You’ll work across hybrid and multi-cloud environments, helping clients modernize their AI infrastructure and adopt AI safely and at scale. As part of your role, you will: Be a senior or lead engineer on client AI platform engagementsArchitect and deploy AI-ready infrastructure (GPU-accelerated compute, Kubernetes/OpenShift, and cloud-native services) across cloud, on-premises, and hybrid environmentsBuild and operate core AI platform components: model serving and gateway infrastructure, agent orchestration and tool-calling frameworks, evaluation harnesses, and guardrail/governance layersImplement MLOps and LLMOps pipelines (model deployment, monitoring, retraining, and fine-tuning where relevant) using Infrastructure-as-Code, GitOps, and CI/CDEstablish observability, security, and governance frameworks specific to AI systems, including cost attribution and lifecycle managementWork with clients and internal teams to develop new opportunities and shape a strong AI platform engineering cultureLead client workshops, architecture reviews, and technical briefings; provide operational support including monitoring and troubleshootingShare your knowledge and experience with colleagues as you coach and mentor them, while developing your own skills by experimenting with and learning new technologies You’ll bring deep, hands-on experience in most of the areas below, with strong depth in AI/GenAI platform engineering specifically. You don’t need to tick every box. AI & GenAI Platform Engineering Model serving and gateway infrastructure (e.g. vLLM, LiteLLM, managed endpoints), with routing, failover, and per-workload cost attributionAgent orchestration and tool-calling frameworks (e.g. LangGraph or equivalent), including familiarity with the Model Context Protocol (MCP)Evaluation engineering (golden datasets, regression gates in CI, LLM-judge calibration)Guardrail and AI-observability tooling (e.g. NeMo Guardrails, OpenTelemetry GenAI conventions, LangSmith, Braintrust) MLOps & LLMOps Hands-on with MLOps platforms (Azure ML, Databricks, SageMaker) and vector/retrieval databases (Pinecone, Milvus, pgvector)Experience with GPU-accelerated infrastructure and NVIDIA AI Enterprise or equivalent stacksExposure to fine-tuning, RLHF, or SLM distillation is a strong plus Cloud-Native & Infrastructure Deep expertise in Kubernetes and container platforms (OpenShift, AKS, EKS, GKE, or VMware Tanzu)Infrastructure as Code and DevOps practices (Terraform, Bicep, Ansible, GitOps and CI/CD pipelines)5+ years’ experience across Azure, AWS, or GCP; strong DevOps fundamentals