Description
Senior Platform Engineer - Azure Data & AI Platform
The Role
We are looking for senior platform engineers to build and operate our Azure-based data and AI platform.
This is hands-on infrastructure and platform work - you'll be writing Terraform, designing network architecture, implementing MLOps (AI model) pipelines, and establishing DevSecOps patterns that engineering teams can use.
This isn't a "thought leadership" or "strategy" role.
You'll be in the code, in the CLI, and in the infrastructure daily.
What You'll Actually Do
Azure Platform Engineering (40%)
Design and implement Azure landing zones, management groups, and subscription architecture
Build and maintain hub-spoke network topologies with proper segmentation and security controls
Implement Azure Policy, RBAC, and governance frameworks that balance security with developer productivity
Manage identity and access using Entra ID, service principals, managed identities
Establish monitoring, logging, and alerting with Azure Monitor, Log Analytics, and Application Insights
Cost management and FinOps practices - keeping cloud spend under control without hamstringing teams
Databricks & Data Platform (30%)
Deploy and configure Azure Databricks workspaces with Unity Catalog for data governance
Assisting both data and AI teams with pipeline development
Establish Databricks best practices: cluster policies, job scheduling, notebook standards, workspace organization
Integrate Databricks with ADLS Gen2, Azure SQL, Synapse, and other data services
Set up and maintain CI/CD for Databricks notebooks, jobs, and infrastructure
Work with data engineers on performance optimization, cost control, and platform capabilities
MLOps & AI Platform (20%)
Build ML model deployment pipelines using Azure ML, Databricks MLflow, or both
Implement model versioning, experiment tracking, and model registry patterns
Establish inference endpoints (batch and real-time) with proper monitoring and governance
Create reusable ML pipeline templates and infrastructure-as-code modules
Integrate AI services (Azure OpenAI, Cognitive Services) into platform offerings
Implement responsible AI guardrails: model monitoring, bias detection, explainability
DevSecOps & Platform Enablement (10%)
Creation of tools, features and dashboards for the developer platform
Build CI/CD pipelines in Azure DevOps or GitHub Actions with proper security scanning
Implement shift-left security: SAST/DAST, dependency scanning, infrastructure scanning, secrets management
Establish infrastructure-as-code standards with Terraform (or Bicep), including modules and policy enforcement
Create self-service tooling and automation for common platform tasks
Write documentation that engineers will read and use
Participate in on-call rotation for platform incidents
What We Need From You
Required
5+ years platform/infrastructure engineering - you've built production platforms, not just prototypes
Deep Azure knowledge - networking, IAM, storage, compute, PaaS services.
You know the difference between service endpoints and private endpoints and when to use each
Databricks experience - you've deployed workspaces, configured Unity Catalog, optimized Spark jobs, managed costs
Infrastructure as Code - Terraform (preferred) or Bicep.
You write modules, understand state management, know how to structure large IaC projects
CI/CD pipelines - Azure DevOps or GitHub Actions.
You've built multi-stage pipelines with gates, approvals, and security scanning
Containerization experience - you know best practices when building and working with both application-based containers and containers holding ML/AI models
Security-first mindset - you understand defense in depth, least privilege, network segmentation, and don't treat security as an afterthought
MLOps fundamentals - model training vs inference, experiment tracking, model versioning, deployment patterns
Python and/or PowerShell - for automation, tooling, and platform utilities
Observability stack beyond basic metrics (distributed tracing, log aggregation patterns)
Strongly Preferred
Experience with Azure landing zones and CAF (Cloud Adoption Framework)
Microsoft Purview for data governance and cataloging
Experience with Delta Lake, Spark optimization, data quality frameworks
Azure networking certifications or equivalent deep knowledge
Container orchestration (AKS)
API design and management (API Management, App Gateway, Front Door)
What Actually Matters
Pragmatism over purity - you choose the right tool for the job, not the coolest one
Documentation discipline - you document as you build because you know future-you will thank today-you
Automation mindset - if you do it twice, you automate it
Everything-as-code - if it's not in git, it doesn't exist to you
Collaboration skills - you can translate between data scientists, engineers, and business stakeholders
Ownership mentality - you build it, you run it, you support it
Intellectual honesty - you say "I don't know" when you don't, and then you figure it out
What We Offer
Actual flexibility: Remote-first with occasional in-office travel for workshops/planning.
We care about outcomes, not seat time.
Real learning budget: for conferences, training, certifications.
We expect you to use it.
Tooling: You'll get the equipment and licenses you need to do the job properly.
Grown-up engineering culture:
PRs are required, branching is mandatory, tests matter
Blameless post-mortems when things break
Technical decisions driven by evidence and context, not politics or trends
We write RFCs for significant changes
Add the usual stuff here: How to apply, interview process.