Summary
✨ AI‑Generated
A technology team is seeking an AI/ML Platform Engineer to deploy and support scalable machine-learning workloads across Kubernetes environments. The role covers GitOps deployment, CI/CD, Helm, secrets management, infrastructure automation, and production reliability across development, QA, and production environments.
Highlights
Deploy and operate scalable AI/ML platforms across Kubernetes environments, with strong ownership of GitOps, CI/CD, secrets management, automation, and production reliability.
Description
Position: AI/ML Platform Engineer (Cloud Infrastructure)
Location: 4days/week (Onsite- Toronto, ON)
Job type: Full-Time/Permanent Hiring
We are seeking a skilled AI/ML Platform Engineer to deploy, manage, and support scalable AI/ML infrastructure and applications across Kubernetes environments.
The role will focus on Kubernetes platform administration, GitOps-based deployments, CI/CD, secrets management, infrastructure automation, and production reliability supporting platforms such as LlamaIndex Cloud and KDB.AI.
Roles & Responsibilities
Deploy and manage AI/ML applications and platforms, including LlamaIndex Cloud, KDB.AI, and similar workloads, across Dev, QA, and Production environments.Administer and maintain production Kubernetes clusters across multiple environments.Implement and manage GitOps deployment workflows using Flux CD, Argo CD, or similar tools.Develop and maintain Helm charts and Kubernetes manifests for application deployments.Configure and manage HashiCorp Vault and ExternalSecrets for secure secrets management.Build and maintain CI/CD pipelines using GitHub Actions, Jenkins, Artifactory, or similar technologies.Integrate enterprise identity solutions using OIDC and Microsoft Entra ID.Manage Kubernetes networking, Ingress, and Gateway API configurations.Administer and support PostgreSQL, MongoDB, Redis, and RabbitMQ infrastructure.Configure and maintain database high availability and failover, including technologies such as PgBouncer and HAProxy.Monitor application and infrastructure performance and troubleshoot production issues.Automate operational tasks using Bash, PowerShell, Python, or Go.Create and maintain infrastructure documentation, runbooks, deployment procedures, and platform patterns.Collaborate with development, security, identity, and infrastructure teams to deliver reliable, secure, and scalable AI/ML platforms.Required Skills
3+ years of production Kubernetes administration2+ years of experience with GitOps using Flux CD, Argo CD, or similar toolsAdvanced Helm chart development and Kubernetes manifest managementHashiCorp Vault and secrets-management solutionsArtifactory or similar artifact/container registry platformsCI/CD with GitHub Actions, Jenkins, or similar toolsKubernetes networking, Ingress, and Gateway APIAdministration of PostgreSQL, MongoDB, Redis, and RabbitMQDatabase HA/failover configuration and troubleshootingStrong Linux/Unix administration and shell scriptingBash and/or PowerShell scriptingNice to Have
Experience with LlamaIndex, LangChain, or AI/ML platformsExperience with vector databases or KDB.AIExperience with Temporal.ioProgramming experience with Python or Go for automation