Site Reliability Engineer

Boson Ai — Canada · Posted ~1 day ago

Mid Remote

Skills

Site reliability engineering AI infrastructure Networking Cluster scheduling Storage GPU systems Infrastructure automation Observability Scalable infrastructure AI training and inference infrastructure GPU clusters High-performance networking Automation

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary

Join a technically ambitious team building and operating dependable infrastructure for large-scale AI training and inference. You will design, automate, and improve systems spanning high-performance networking, GPU clusters, storage, scheduling, and observability. Deep expertise in at least one infrastructure area is valued, along with the curiosity and judgment to collaborate across the stack.

Highlights

Hands-on ownership of large-scale AI infrastructure with opportunities to work across high-performance networking, GPU clusters, storage, scheduling, and reliability. The role offers meaningful technical ownership, cross-functional collaboration, and the chance to improve complex systems for scalability and observability.

Description

About The Role Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. We are looking for a Site Reliability Engineer to help build and operate the infrastructure behind that work. Based in Toronto or remote, you will work across the systems that enable large-scale AI training and serving: high-performance networks, GPU clusters, storage, scheduling, and the operational tooling that keeps them reliable. This is a hands-on role for someone who enjoys taking complex infrastructure from “it works” to dependable, observable, and scalable. You do not need to be an expert in every layer of the stack. We are looking for deep strength in at least one area—networking, cluster scheduling, storage, GPU systems, or AI infrastructure— and the curiosity and judgment to collaborate across the rest. Responsibilities Design, operate, and improve reliable infrastructure for AI training and inference workloadsOwn and automate operational workflows across one or more core areas: networking, compute allocation, storage, GPU/server configuration, or AI platformsBuild monitoring, alerting, runbooks, and incident-response practices that make systems easier to operateDiagnose performance, capacity, and reliability issues across hardware, operating systems, networks, schedulers, and distributed workloadsPartner closely with ML, research, and platform teams to translate workload needs into practical infrastructure improvementsImprove provisioning, configuration management, testing, and deployment automationHelp plan cluster growth, capacity allocation, upgrades, and lifecycle managementContribute to a thoughtful reliability culture through documentation, post-incident learning, and pragmatic engineering standards Minimum Qualifications 4+ years of experience in site reliability engineering, infrastructure engineering, systems engineering, or a related production-operations roleStrong hands-on expertise in at least one of the following:Networking, including firewalls, switching, routing, ASN/BGP configuration, or InfiniBandCluster and systems allocation with Kubernetes, SLURM, MAAS, or similar platformsDistributed storage, particularly CephGPU and server administration, including CUDA drivers, firmware, BIOS, and hardware troubleshootingAI training or model-serving infrastructure Experience operating production systems with a focus on availability, performance, security, and automationStrong Linux administration and scripting skillsA systematic approach to troubleshooting across multiple layers of a complex systemClear written and verbal communication skills, including the ability to work effectively with a distributed team Preferred Qualifications Experience supporting GPU-intensive AI or HPC environments Experience with NVIDIA GPUs, CUDA, NCCL, and high-performance interconnects - Experience with InfiniBand, RDMA, RoCE, or 100Gb+ Ethernet Familiarity with Kubernetes, SLURM, MAAS, Terraform, Ansible, or similar infrastructure tooling Experience operating or tuning Ceph clusters Familiarity with observability tooling such as Prometheus, Grafana, and centralized logging systems Experience with hardware provisioning, firmware management, and bare-metal automation Experience running large-scale distributed training or high-throughput inference workloads Familiarity with cloud and hybrid infrastructure across AWS, GCP, or Azure Boson AI is building AI systems for real-world, business-critical use. If you enjoy solving difficult infrastructure problems and want your work to directly enable the next generation of AI products, we’d love to hear from you. We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.