Senior Platform Engineer - Azure Data & AI Platform

Howden Insurance — United Kingdom · Posted ~3 hours ago

Senior Full-time

Skills

Microsoft Azure Terraform Network Architecture MLOps DevSecOps Azure Policy RBAC Entra ID Monitoring and Logging Landing Zones Monitoring Tools

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

We are looking for senior platform engineers to build and operate a cloud-based data and AI platform. This is hands-on infrastructure work involving writing Terraform, designing network architecture, and implementing MLOps pipelines. You will establish governance, security, and monitoring patterns that engineering teams can rely on daily.

Highlights

Build and operate a scalable Azure-based data and AI platform; hands-on engineering role with real code, CLI, and infrastructure ownership; opportunity to implement MLOps and DevSecOps patterns.

Description

Senior Platform Engineer - Azure Data & AI Platform The Role We are looking for senior platform engineers to build and operate our Azure-based data and AI platform. This is hands-on infrastructure and platform work - you'll be writing Terraform, designing network architecture, implementing MLOps (AI model) pipelines, and establishing DevSecOps patterns that engineering teams can use. This isn't a "thought leadership" or "strategy" role. You'll be in the code, in the CLI, and in the infrastructure daily. What You'll Actually Do Azure Platform Engineering (40%) Design and implement Azure landing zones, management groups, and subscription architecture Build and maintain hub-spoke network topologies with proper segmentation and security controls Implement Azure Policy, RBAC, and governance frameworks that balance security with developer productivity Manage identity and access using Entra ID, service principals, managed identities Establish monitoring, logging, and alerting with Azure Monitor, Log Analytics, and Application Insights Cost management and FinOps practices - keeping cloud spend under control without hamstringing teams Databricks & Data Platform (30%) Deploy and configure Azure Databricks workspaces with Unity Catalog for data governance Assisting both data and AI teams with pipeline development Establish Databricks best practices: cluster policies, job scheduling, notebook standards, workspace organization Integrate Databricks with ADLS Gen2, Azure SQL, Synapse, and other data services Set up and maintain CI/CD for Databricks notebooks, jobs, and infrastructure Work with data engineers on performance optimization, cost control, and platform capabilities MLOps & AI Platform (20%) Build ML model deployment pipelines using Azure ML, Databricks MLflow, or both Implement model versioning, experiment tracking, and model registry patterns Establish inference endpoints (batch and real-time) with proper monitoring and governance Create reusable ML pipeline templates and infrastructure-as-code modules Integrate AI services (Azure OpenAI, Cognitive Services) into platform offerings Implement responsible AI guardrails: model monitoring, bias detection, explainability DevSecOps & Platform Enablement (10%) Creation of tools, features and dashboards for the developer platform Build CI/CD pipelines in Azure DevOps or GitHub Actions with proper security scanning Implement shift-left security: SAST/DAST, dependency scanning, infrastructure scanning, secrets management Establish infrastructure-as-code standards with Terraform (or Bicep), including modules and policy enforcement Create self-service tooling and automation for common platform tasks Write documentation that engineers will read and use Participate in on-call rotation for platform incidents What We Need From You Required 5+ years platform/infrastructure engineering - you've built production platforms, not just prototypes Deep Azure knowledge - networking, IAM, storage, compute, PaaS services. You know the difference between service endpoints and private endpoints and when to use each Databricks experience - you've deployed workspaces, configured Unity Catalog, optimized Spark jobs, managed costs Infrastructure as Code - Terraform (preferred) or Bicep. You write modules, understand state management, know how to structure large IaC projects CI/CD pipelines - Azure DevOps or GitHub Actions. You've built multi-stage pipelines with gates, approvals, and security scanning Containerization experience - you know best practices when building and working with both application-based containers and containers holding ML/AI models Security-first mindset - you understand defense in depth, least privilege, network segmentation, and don't treat security as an afterthought MLOps fundamentals - model training vs inference, experiment tracking, model versioning, deployment patterns Python and/or PowerShell - for automation, tooling, and platform utilities Observability stack beyond basic metrics (distributed tracing, log aggregation patterns) Strongly Preferred Experience with Azure landing zones and CAF (Cloud Adoption Framework) Microsoft Purview for data governance and cataloging Experience with Delta Lake, Spark optimization, data quality frameworks Azure networking certifications or equivalent deep knowledge Container orchestration (AKS) API design and management (API Management, App Gateway, Front Door) What Actually Matters Pragmatism over purity - you choose the right tool for the job, not the coolest one Documentation discipline - you document as you build because you know future-you will thank today-you Automation mindset - if you do it twice, you automate it Everything-as-code - if it's not in git, it doesn't exist to you Collaboration skills - you can translate between data scientists, engineers, and business stakeholders Ownership mentality - you build it, you run it, you support it Intellectual honesty - you say "I don't know" when you don't, and then you figure it out What We Offer Actual flexibility: Remote-first with occasional in-office travel for workshops/planning. We care about outcomes, not seat time. Real learning budget: for conferences, training, certifications. We expect you to use it. Tooling: You'll get the equipment and licenses you need to do the job properly. Grown-up engineering culture: PRs are required, branching is mandatory, tests matter Blameless post-mortems when things break Technical decisions driven by evidence and context, not politics or trends We write RFCs for significant changes Add the usual stuff here: How to apply, interview process.