Platform Engineer - AI/ML Infrastructure

Bench Talent Cloud — Canada · Posted ~2 hours ago

Full-time

Skills

Kubernetes GitOps Flux CD CI/CD GitHub Actions HashiCorp Vault ExternalSecrets OIDC Microsoft Entra ID Helm Kubernetes manifests Production troubleshooting Artifactory LlamaIndex Cloud KDB.AI

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join a platform engineering team responsible for deploying and operating AI/ML workloads on Kubernetes across development, QA, and production. You will implement GitOps workflows, manage secrets and identity integrations, maintain CI/CD pipelines and Helm configurations, and troubleshoot production systems.

Highlights

Platform engineering role centered on modern AI/ML infrastructure and Kubernetes. Offers hands-on responsibility for multi-environment deployments, GitOps automation, secrets management, identity integration, CI/CD, monitoring, and production reliability.

Description

Platform Engineer - AI/ML Infrastructure About the Role: We're seeking a skilled Platform Engineer to deploy and manage our AI/ML infrastructure, e.g LlamaIndex Cloud and KDB.AI applications on Kubernetes Platform. You'll be responsible for building reliable, scalable, and secure deployment pipelines using modern GitOps practices. What You'll Do: - Deploy and manage LlamaIndex Cloud and KDB.AI applications and similar products supporting AI workloads across Dev/QA/Prod environments - Implement GitOps workflows using Flux CD for automated deployments - Administer Kubernetes clusters across multiple environments - Configure HashiCorp Vault for secrets management and ExternalSecrets integration - Maintain CI/CD pipelines with GitHub Actions and Artifactory - Work with enterprise Identity Management team to configure OIDC authentication with Microsoft Entra ID - Create and maintain Helm charts and Kubernetes manifests - Monitor application performance and troubleshoot production issues - Document procedures, runbooks, and infrastructure patterns Required Skills: Core Technologies: - 3+ years managing production Kubernetes clusters - 2+ years with Flux CD, ArgoCD, or similar GitOps tools - Advanced Helm chart development and management - HashiCorp Vault for secrets management - Artifactory or similar container registries - CI/CD with GitHub Actions, Jenkins, or similar Infrastructure & Database: - PostgreSQL, MongoDB, Redis, RabbitMQ administration - Database HA/failover configurations (PgBouncer, HAProxy) - Linux/Unix systems and shell scripting (Bash, PowerShell) - Kubernetes networking, Ingress, and Gateway API Nice to Have: - Experience with LlamaIndex, LangChain, or AI/ML platforms - Vector databases or KDB.AI knowledge - Temporal.io workflow orchestration - Python/Go for automation