AI Platform Engineer

Systems Limited — Jordan · Posted ~2 hours ago

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Description

We are seeking a AI Platform Engineer with approximately 6–12+ years of experience in the field. Builds and operates the shared AI platform infrastructure — the paved road every AI practice builds on top of, so no team reinvents deployment plumbing. Responsibilities: Build and maintain shared AI platform infrastructure — compute provisioning, networking, IAM for AI workloads Own the internal tooling and templates practices use to deploy models/agents consistently Standardize CI/CD pipelines for AI workloads across practices, including shared AI evaluation platforms Manage platform-level cost governance and capacity planning across concurrent engagements Own platform security posture in partnership with AI Security Engineers Partner with MLOps/LLMOps Engineers on the boundary between platform and workload-specific operations Balance competing infrastructure requests from multiple practice leads Document platform capabilities clearly enough that practices can self-serve Forecast and justify platform spend to non-technical leadership Requirements: 6–12+ yrs platform/infrastructure engineering, with 2+ yrs supporting AI/ML workloads specifically Deep cloud infrastructure expertise (IaC, Kubernetes, networking, IAM), including hosting vector/graph databases Experience building internal developer platforms/tooling, not just running infrastructure Experience integrating and operating managed AI/agentic platforms — Microsoft Azure AI Foundry, AWS Bedrock, and Google Vertex AI — alongside self-hosted open-source stacks as a good-to-have Familiarity with multi-tenant capacity planning and cost allocation Experience with platform-level security hardening Cross-practice stakeholder management — balances competing infra requests from multiple practice leads Cost/capacity planning literacy — can forecast and justify platform spend to non-technical leadership Documents platform capabilities clearly enough that practices can self-serve Collaborative — builds shared infrastructure without becoming a bottleneckSuccess metrics: platform uptime/reliability · cost per workload vs. budget · practice self-service adoption rate