Senior Infrastructure Engineer

Nexgencloud β€” Canada Β· Posted ~3 hours ago

Senior

Skills

Infrastructure engineering OpenStack Kubernetes Cloud infrastructure Infrastructure operations Infrastructure design

πŸ”“ Log in to save this job, tailor your resume & track your apply process β€” 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A Senior Infrastructure Engineer is sought to take ownership of business-critical cloud infrastructure as global environments scale. The role focuses on designing, operating, and improving large OpenStack and Kubernetes environments, with substantial responsibility for reliability, scalability, and the evolution of AI-oriented infrastructure.

Highlights

High-ownership infrastructure role supporting rapidly scaling global cloud environments. Engineers gain direct responsibility for business-critical infrastructure, with opportunities to work deeply on OpenStack, Kubernetes, platform design, operations, and continuous improvement in an AI-focused environment.

Description

About Nexgen Cloud NexGen Cloud is the company behind Hyperstack, a full-stack AI cloud serving tens of thousands of customers from AI researchers to enterprises running the world's most compute-intensive workloads. We deliver on-demand and private GPU infrastructure to teams who treat performance as a requirement, not a feature. We're a tight-knit, fast-moving team working at the cutting edge of AI cloud infrastructure. We practice what we preach, equipping our people with AI at every level so we can solve harder problems, ship faster, and keep raising the bar for what enterprise GPU infrastructure looks like. THE ROLE: Senior Infrastructure Engineer This role exists because our platform is scaling quickly β€” and complexity comes with it. As we expand our OpenStack and Kubernetes environments globally, we need engineers who can take real ownership of how the platform is designed, operated, and improved. You'll have direct ownership over business-critical infrastructure that impacts performance, reliability, and customer experience. This is not a maintenance role. If you like solving hard problems, owning systems end-to-end, and seeing the impact of your work immediately β€” you'll enjoy this. What You'll Be Doing Rather than a long checklist, here's what success in this role looks like: Own the design, deployment, and operation of OpenStack and Kubernetes environments β€” ensuring platform performance, scalability, and resilience for GPU workloadsBuild and improve infrastructure using infrastructure-as-code and GitOps practices, driving automation across provisioning, deployment, and operational workflowsOptimise GPU workload scheduling using Kubernetes and NVIDIA tooling, and implement monitoring, logging, and alerting to ensure platform stabilityLead incident response and drive continuous improvement of reliability across the platformMaintain strong security controls across infrastructure and container layers β€” RBAC, network policies, and tenant isolationWork closely with Platform, DevOps, AI, Product, and Support teams to align infrastructure capabilities with customer and platform requirements About You We're more interested in how you think and work than in a perfect CV. You'll likely bring a combination of the following: Essential Extensive hands-on Linux systems administration skills and knowledge β€” genuine depth, not surface-level familiarityStrong, proven experience building servers and racks β€” you've physically assembled, cabled, and commissioned hardware, not just specified or overseen itDirect hands-on experience physically working in data centres β€” you've personally stacked and racked hardware on-siteA willingness and ability to travel to Quebec sites as requiredA solid understanding of networking and storage systems Nice to Have Experience installing, racking, and configuring GPU hardware specifically, ideally including NVIDIA platformsProduction experience running OpenStack and/or Kubernetes at scaleExperience with infrastructure automation, CI/CD, and Git-based workflowsBroader exposure to HPC or large-scale compute environmentsContributions to open-source projects What We Offer Competitive salary and annual discretionary bonus schemeEmployee wellbeing benefits25 days of holiday, plus public holidaysFlexible working arrangements (remote or hybrid, depending on role and location)Real ownership and autonomy, with the trust to take initiative and experimentThe opportunity to make a visible, meaningful impact as we scaleClear career progression and growth opportunities in a fast-growing companyA collaborative, international culture built on trust, transparency, and ownershipThe chance to help shape NexGen Cloud's team, culture, and future alongside ambitious, mission-driven colleagues MORE INFORMATION Head over to our NexGen Cloud careers page to view current openings and follow us on LinkedIn and X to learn more about our journey, newest releases and hear exciting news in the neocloud space.