DevOps Platform Engineer | AI & LLMOps Infrastructure

Knk Gt — Canada · Posted ~2 hours ago

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Description

Position: DevOps Platform Engineer | AI & LLMOps Infrastructure Location: Onsite- Toronto, ON Job type: Full-Time/Permanent Hiring As an Agent & DevOps Platform Engineer, you will design, scale, and secure the foundational infrastructure that powers our next-generation developer platforms and autonomous AI/LLM agent frameworks. This is a highly specialized role bridging the gap between advanced cloud-native Platform Engineering and cutting-edge LLMOps. You will not just build standalone AI applications—you will build the highly automated, multi-account AWS environments, robust CI/CD pipelines, and secure IAM patterns that allow autonomous agents and developers to ship code reliably, securely, and at enterprise scale. Key Responsibilities Infrastructure as Code & Automation: Architect, deploy, and manage highly reliable multi-account AWS environments using Terraform, CloudFormation, or AWS CDK.Agent & Tooling Orchestration: Build, deploy, and scale enterprise developer platforms, integrating agentic orchestration frameworks (e.g., LangChain, LlamaIndex) and custom toolchains into production pipelines.Bulletproof Identity & Access: Implement advanced AWS IAM architectures under the principle of least privilege, designing secure cross-account access, service-to-service authentication, and Kubernetes service accounts.Container Platform Engineering: Package and orchestrate distributed container systems using Docker and enterprise container platforms (Amazon EKS, ECS, or Kubernetes), configuring robust networking, ingress, and workload security.LLMOps & Platform Observability: Build the guardrails for production AI workloads, implementing systems for prompt management, tool execution, model evaluation, observability (CloudWatch), and cost-reliability metrics.Scale CI/CD & DevSecOps: Treat security as a core software design constraint. Author high-quality reusable pipeline patterns, build automation workflows, and enable decentralized DevOps practices across multiple business units.System Ownership & Culture: Act as a technical mentor whose coaching is reflected in clean working code, robust pipelines, and reusable modules. Define and solve core platform bottlenecks in ambiguous environments.Required Qualifications Experience: 3–5+ years of hands-on platform engineering, cloud architecture, and CI/CD operations within complex enterprise environments.Programming: Strong proficiency in Python for infrastructure tooling, API integrations, agent services, and advanced production scripting.Technical Ecosystems: Working competence in at least one additional development ecosystem (e.g., Java, .NET, Node.js, or Groovy).Cloud Native & Containers: Deep production experience with AWS cloud-native deployments and container orchestration via Docker and Kubernetes / Amazon EKS / ECS.AWS Security Mastery: Comprehensive knowledge of complex AWS IAM structures, including roles, policies, trust relationships, and secure machine-to-machine authentication.AI/LLMOps Engineering: Direct experience building, deploying, or scaling AI/LLM-based tools or agents in or near production. Solid understanding of prompt management, tool usage, and agent testing.Methodology: Strong grasp of Git-based workflows, release management at scale, and an uncompromising DevSecOps approach to system design.Nice-to-Have Skills Hands-on usage of advanced AWS ecosystem tools: ECR, Lambda, Bedrock, CloudWatch, S3, Secrets Manager, Systems Manager, and VPC networking.Familiarity with Kubernetes internals: Helm, service accounts, ingress controllers, and network policies.Prior experience working with framework tools like LangChain or LlamaIndex.Experience driving agile delivery models using Scrum, Kanban, or SAFe across multiple teams.