AI Infrastructure Engineer

Mission Dev — Canada · Posted ~3 hours ago

Full-time Hybrid

Skills

AI infrastructure inference optimization large-scale models hardware utilization cloud computing decoding processes AI Infrastructure Inference Optimization Cloud Computing Hardware Utilization Large-Scale Models

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

We are an AI infrastructure company specializing in inference optimization for large-scale models. Our platform enhances token throughput and reduces latency by optimizing hardware utilization and decoding processes within a customer's cloud environment. We are looking for an AI Infrastructure Engineer to join our remote-first team. The role requires being available to visit our office approximately two to three times per quarter. This is a full-time, direct employment position with 40 hours per week.

Highlights

Remote-first role with a competitive full-time employment position. Work on cutting-edge AI inference optimization that enhances token throughput and reduces latency for large-scale models. Direct employment with a client focused on improving model serving efficiency without requiring infrastructure or code changes. Opportunity to visit the client's office in Montreal quarterly.

Description

Employment type: Full-time position, 40 hours per week, with direct employment by the client. Location: Remote-first role for candidates based in Canada. Travel requirement: Candidates must be available to visit the client’s office in Montreal approximately two to three times per quarter. Our company description Mission.dev is the next-gen staffing platform for software talent. We help you find, evaluate, and manage top software talent (contractors or direct hires) faster, smarter, and more efficiently. Powered by AI. Backed by real humans. About the client An AI infrastructure company specializing in inference optimization for large-scale models. The platform enhances token throughput and reduces latency by optimizing hardware utilization and decoding processes within a customer's cloud environment. The technology focuses on improving the efficiency of model serving without requiring changes to existing infrastructure or code, ensuring data privacy while reducing operational costs for organizations deploying open-source models. About the Role As a founding Senior AI Infrastructure Engineer, you will report to the Chief Technology Officer to design and operate large-scale infrastructure for AI workloads on public cloud and on-premises environments. You will be an individual contributor with significant influence, combining software engineering with deep systems expertise to build secure and reliable platforms. Your work will focus on enabling efficient model serving at scale, ensuring the infrastructure can support a massive number of concurrent users while maintaining high service quality and performance. What You'll Do Design and manage multi-cloud and on-premises infrastructure for internal and external workloads.Orchestrate large-scale deployments using infrastructure-as-code and container orchestration to support high user concurrency.Benchmark AI inference workloads to identify performance bottlenecks and optimize hardware utilization.Develop and maintain comprehensive monitoring systems to ensure high service quality and reliability.Collaborate with customers to deploy and integrate optimization solutions within their existing infrastructure.Build and maintain an internal laboratory for infrastructure testing and experimentation. What You Bring 5+ years of experience in AI infrastructure or systems engineering.Proficiency in Python and experience building and operating scalable APIs.Extensive experience with Kubernetes, Helm, Ansible, and Terraform.Proven track record of shipping large language model (LLM) serving systems in production environments.Experience with distributed computing frameworks such as Ray, RayClusters, or KubeRay.Demonstrated experience with agentic workflows.Bachelor’s degree in Computer Science or a related technical field. Compensation & Benefits Founding-engineer equity and direct ownership.Remote-friendly work environment. Visit the office 2-3x a quarter in Montreal.Access to extensive compute resources and specialized hardware for experimentation.