Summary
✨ AI‑Generated
A senior DevOps role embedded with an AI-focused engineering team, responsible for deploying applications, building cloud infrastructure, improving reliability, and enabling developers to deliver production-ready solutions faster.
Highlights
Work on cutting-edge AI infrastructure, collaborate with engineering teams, and help scale production systems in a fast-moving environment.
Description
About the role
You'll be the DevOps/SRE engineer embedded with a fast-moving AI application team at a large enterprise.
You'll deploy LLM applications and AI agents to Kubernetes, write the Terraform behind them, and help build a new AI gateway.
The team ships quickly, with a high volume of projects in its backlog, and needs someone who makes AI applications production-ready and makes developers faster.
You'll sit with the application developers day to day and partner with an established Platform Engineering team that runs the enterprise Kubernetes platform.
What you'll do
Help developers containerize, configure and deploy AI applications and agents to Kubernetes.Write and maintain Terraform for databases, storage, networking, identities and AI services.Troubleshoot production workloads: health checks, resource settings, scaling, secrets and failed rollouts.Work with Platform Engineering on the Kong API gateway: routing, authentication and rate limiting.Help build the new AI gateway: centralized access to LLMs (Claude, Gemini, Azure OpenAI), with model routing, fallback, token tracking and cost controls.Host and secure MCP servers, agent runtimes and vector workloads such as pgvector.Manage Harness and GitHub CI/CD: security scanning, environment promotion, approvals and rollback.Build golden-path templates and starter kits so developers can launch new AI solutions quickly.Implement Dynatrace observability, SLOs and incident runbooks, and track AI and cloud spend.Enforce Entra ID SSO, OAuth 2.0 / OIDC, RBAC, secrets management and AI guardrails such as redacting personal data before it reaches a model.
What you bring (required)
5+ years in DevOps, SRE, platform or cloud engineering, including 2+ years supporting AI/ML or LLM-based workloads in production.Hands-on production experience deploying and troubleshooting workloads on Kubernetes.Production Terraform that you personally wrote and maintained.Strong hands-on cloud experience: Azure preferred (including Entra ID); AWS or GCP considered.CI/CD pipelines you built, with environment promotion, approval gates and rollback.Working knowledge of API gateways (Kong preferred) and how applications reach the internet (DNS, WAF/CDN, ingress).The ability to explain end to end how a user's identity controls what an application can access, and which parts you built.Observability experience (Dynatrace preferred), including SLOs, alerting and incident response.Python plus TypeScript/Node.js or Go.Daily use of AI coding tools (Claude Code, Cursor, Copilot or Codex), with a repeatable workflow of rules files, review discipline and guardrails.Clear, concise communication backed by numbers: request volume, latency, uptime, cost.
Nice to have
Kong AI Gateway or LiteLLM · Azure OpenAI / AI Foundry · MCP servers or agent frameworks · LLM observability and evaluation tools · Harness, Dynatrace, CyberArk, Chainguard · Node.js / Next.js / React / Prisma · CKA/CKAD, Terraform Associate, or Azure DevOps Engineer / Azure AI Engineer certification
This role is a strong fit if you…
Are a DevOps engineer or SRE who has already put LLM apps or agents into production, and want to do it full-time.Like being the person developers come to when a deploy breaks.Can say "I wrote that Terraform" and "I owned that pipeline" and back it up.
This role is not the right fit if…
Your main work is building RAG pipelines, prompts or models.
That's valuable work, but this role is about running and securing AI applications, not building them.Your Kubernetes and Terraform experience comes from pipelines and clusters others built.You're looking for a pure infrastructure role with no application-team involvement.
How to stand out
In your application or first conversation, share one real production number (for example, request volume, uptime or cost saved) and one AI application you helped run in production.