Senior Performance Engineer - LLM Inference Optimization

Mission Dev — Canada · Posted ~5 hours ago

Senior Full-time Remote

Skills

LLM inference optimization GPU performance AI infrastructure systems programming performance engineering GPU LLM AI inference ML serving systems runtime optimization

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An early-stage AI infrastructure company is seeking a senior engineer to improve large-scale model inference performance. The role involves GPU optimization, low-level systems work, and building technologies that help AI workloads achieve higher efficiency and throughput.

Highlights

Remote full-time role working on advanced AI infrastructure challenges with opportunities to optimize high-performance computing systems and contribute as an early technical team member.

Description

Agreement type: This is a remote full-time employment role requiring 40 hours per week, with direct hiring by the client. Preferred candidates' locations: Canada Our company description Mission.dev is the next-gen staffing platform for software talent. We help you find, evaluate, and manage top software talent (contractors or direct hires) faster, smarter, and more efficiently. Powered by AI. Backed by real humans. About the client An early-stage AI infrastructure startup specializing in inference optimization. The company develops software to increase compute performance per watt on GPUs, integrating with major serving stacks to provide kernel-level power telemetry. By utilizing a proprietary optimization engine and runtime controller, the platform helps AI teams maintain high throughput under power constraints, dynamically adjusting performance as models and hardware conditions evolve. About the Role As a Senior Performance Engineer and founding team member, you will architect solutions across the entire inference stack. You will have direct ownership over optimizing kernels, shaping the serving engine, and improving orchestration to redefine large-scale model deployment. Your work will directly impact the cost and energy efficiency of AI at scale, pushing GPU utilization toward theoretical limits. This role offers the opportunity to translate complex research into robust, production-ready systems that influence the future of high-performance computing. What You'll Do Design and build high-performance inference systems for large-scale models across multi-GPU and multi-node deployments. Develop custom kernels and runtime paths to maximize GPU utilization and memory efficiency. Extend serving engines by implementing advanced batching, KV-cache management, and scheduling policies. Architect distributed inference systems including autoscaling, load balancing, and failure handling under strict latency SLOs. Profile and analyze the full inference lifecycle to identify and resolve systemic bottlenecks. Implement advanced techniques such as speculative decoding, tensor parallelism, and mixture-of-experts serving. Collaborate with hardware partners to co-design software strategies that improve energy efficiency and reduce operational costs. Contribute to internal tooling and relevant open-source projects within the inference ecosystem. What You Bring Extensive experience building or operating large-scale production inference or training systems. Deep understanding of GPU architectures, including compute, memory bandwidth, and kernel launch characteristics. Proficiency with the modern inference ecosystem and frameworks. Strong foundation in systems and distributed systems, including concurrency, scheduling, and fault tolerance. Expertise in debugging performance issues across layers, from kernel traces to autoscaler dynamics. Ability to own technical problems end-to-end, from initial measurement to long-term operational health. Nice to Haves Contributions to major open-source inference frameworks or libraries. Experience with Kubernetes-based infrastructure and modern observability stacks. Knowledge of energy-aware scheduling or capacity planning for large accelerator fleets. Familiarity with post-training pipelines and their interaction with inference infrastructure. Compensation & Benefits Founding-engineer equity and direct ownership. Access to state-of-the-art compute resources and hardware. Remote-friendly. Montreal preferred for regular in-person work with the founding team.