Senior Systems Engineer

Illuminhq — Canada · Posted ~2 hours ago

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Description

Who we need We are hiring a Senior Systems Engineer to join our Infrastructure & Operations team. This is a senior, hands-on infrastructure role focused on Linux systems, virtualization, physical and co-location infrastructure, Kubernetes, storage, automation, and complex production troubleshooting. The successful candidate will act as a senior technical escalation point for systems and infrastructure issues and will work closely with DevOps, Network, DBA, Security, and application teams to deliver reliable, secure, scalable infrastructure in support of illumin's global operations. The ideal candidate has deep expertise in core systems and infrastructure engineering, broad knowledge across adjacent technology domains, and the ability to trace problems across operating systems, compute, storage, networking, virtualization, and container platforms. The role requires strong ownership, sound operational judgment, and a focus on automation and continuous improvement. This is a full-time, hybrid position, working in our Toronto office three days per week. You must be able to travel internationally (US & Europe) without restrictions. How you will make an impact: Systems Administration & Infrastructure Engineering Serve as a senior technical escalation point for Linux systems and infrastructure issues, with strong working knowledge of Windows Server environments.Manage, maintain, and optimize VMware vSphere/ESXi environments and other virtualization platforms, including compute, networking, storage, host lifecycle, and vCenter integration.Operate and support on-premise, co-location, and hybrid infrastructure across multiple sites and regions.Deploy, maintain, and troubleshoot enterprise server hardware, including storage controllers, NICs, firmware, out-of-band management, and hardware lifecycle management.Perform capacity planning, lifecycle management, performance analysis, and reliability improvements across infrastructure platforms. Cloud-Native & Container Platforms Deploy, operate, upgrade, and troubleshoot production Kubernetes clusters at scale.Support Kubernetes node lifecycle, container runtimes, networking, DNS, certificates, persistent storage, observability, and cluster reliability.Work with Containerd, Docker, Helm, and related container and orchestration tooling.Partner with DevOps and application teams to improve platform reliability, deployment practices, and operational standards. Networking & Cross-Platform Troubleshooting Apply strong enterprise networking fundamentals, including TCP/IP, routing, DNS, DHCP, VLANs, load balancing, and BGP, to troubleshoot infrastructure and application connectivity issues.Work collaboratively with the Network Engineering team on Cisco, Juniper, load balancers, routing, and datacenter connectivity; direct ownership of all network platforms is not required.Diagnose complex production issues that span systems, network, storage, virtualization, and application layers using packet captures, logs, metrics, and system-level tooling. Automation, DevOps & Observability Develop and maintain infrastructure automation using Python, Ansible, Bash, APIs, and related tooling; Go experience is an asset.Automate repetitive infrastructure operations, configuration changes, validation, and systems integrations across large-scale environments.Apply DevOps practices to configuration management, CI/CD workflows, infrastructure delivery, and operational automation.Improve monitoring, logging, alerting, capacity management, and observability across infrastructure platforms using tools such as Prometheus, Grafana, Netdata, ELK, or equivalent. Storage & Data Platform Infrastructure Support enterprise and distributed storage technologies, including local storage, SAN/NAS, object storage, and Kubernetes persistent storage.Troubleshoot storage performance, availability, disk, filesystem, and capacity issues across physical, virtual, and containerized environments.Support infrastructure underlying MySQL and other stateful platforms, including compute, storage, networking, high availability, and performance troubleshooting, in collaboration with the DBA team.Operate effectively in high-throughput and low-latency application environments. Operational Excellence Lead and participate in incident response, root-cause analysis, problem management, change management, and disaster recovery activities.Drive proactive monitoring, capacity planning, resiliency, standardization, and infrastructure lifecycle initiatives.Produce clear technical documentation, runbooks, standards, and operational procedures.Mentor junior engineers and share technical knowledge across teams.Collaborate closely with DevOps, Network, DBA, Security, and application teams to deliver infrastructure that meets technical and business requirements. You will bring these required qualifications: Experience: 8-10+ years in systems engineering, infrastructure engineering, systems administration, or a closely related role, with substantial experience supporting production environments. Certifications: Certified Kubernetes Administrator (CKA)ITIL Foundations (v3 or v4) in IT Service Management Technical Expertise: Deep hands-on Linux systems administration and troubleshooting experience, with solid working knowledge of Windows Server.Strong VMware vSphere/ESXi administration and troubleshooting experience; familiarity with other hypervisors is beneficial.Hands-on experience operating production Kubernetes environments, including cluster troubleshooting and lifecycle management.Strong understanding of enterprise networking concepts, including TCP/IP, routing, DNS, DHCP, VLANs, load balancing, and general BGP concepts.Strong automation skills using Python, Ansible, Bash, APIs, or comparable tooling.Experience supporting on-premise and co-location infrastructure, including enterprise server hardware and remote management technologies.Experience with storage technologies and troubleshooting across physical, virtual, and containerized environments.Working knowledge of MySQL architecture, replication/high availability concepts, backup considerations, and common operational troubleshooting.Strong understanding of DevOps practices, configuration management, monitoring, and automation for complex systems.Demonstrated ability to troubleshoot cross-functional production issues independently and collaborate effectively with specialist teams. Your application is even stronger if you have some of these preferred qualifications: Experience with Ceph, MinIO, ZFS, NFS, SAN/NAS, or other distributed and enterprise storage platforms.Experience with Cisco, Juniper, A10, HAProxy, or similar networking and load-balancing platforms.Experience with Prometheus, Grafana, Netdata, ELK, or comparable monitoring and observability platforms.Experience supporting low-latency, high-throughput distributed application environments.Experience with Go or other software-development languages used for infrastructure tooling.Experience with ServiceNow or a comparable ITSM platform.Exposure to GPU, AI/ML, or high-performance compute infrastructure is an asset where relevant to future platform requirements. Compensation At illumin, we believe compensation should be transparent, fair, and reflective of both experience and impact. The salary range for this role is $100,000–$150,000 plus bonus. Compensation is determined based on skills, experience, and the scope and complexity of the role. We value open conversations about compensation and are happy to discuss our approach at any stage of the hiring process. Who we are At illumin, we are transforming the advertising landscape. Our platform offers an integrated space for journey planning, execution, and reporting. It empowers marketers to connect with their audiences in powerful ways through real-time data and easy-to-use visual tools. By seamlessly combining media planning and buying in an intuitive interface, marketers can take complete control of their campaigns, meeting customers wherever they are in the buying journey and maximizing the impact of their ad spend through personalized insights for smarter decision-making. We are at a pivotal moment, evolving into a product-led company with a team of over 100 skilled professionals and new leadership guiding our path forward. By harnessing the power of data, advancing our AI capabilities, and deeply investing in our people, we are preparing for a future that will redefine what’s possible in journey advertising. Our work is guided by two beliefs: that the ability to execute is paramount to success and that we are only as good as our people. As we grow and transform, we are looking for team members (illumineers) who share our bias for speed, delivery over perfection, and an entrepreneurial mindset. Joining us now is a chance to be part of our transformation. What else should you know about us? We are undergoing a transformative shift. We are embracing change and the opportunities that come with it, empowering every illumineer to innovate, experiment, and bring forward new ideas. Whether accessing new technology, restructuring workflows, or expanding your team, you will have full support if you can make the business case. We are a broad and diverse team, but we all share a passion for success, a drive to do more, and a love of creating connections. We hire for talent and commitment and provide the guidelines and guidance to elevate skills, knowledge, and abilities across all areas. This is a place where proven methods meet bold ideas, offering opportunities to grow personally and professionally. To support a healthy work-life balance, we offer a flexible work environment, a meal credit for your in-office days, and a free massage with an RMT in-house every eight weeks. That is in addition to our comprehensive benefits, which include life, AD&D, long-term disability insurance, and coverage for prescriptions, dental, vision, mental health, and professional health services. You will also have access to a workplace advisor, the Vitality Wellness app, and a $300 annual healthcare spending account. Apply now If you want to seize the opportunity to impact a company and influence an industry, and you have 70% of what we are looking for, apply now. We can't promise an interview, but we will consider your whole application. What you can expect from our interview process: A virtual interview with our Talent Acquisition Partner to discuss your interest in the role, your background, and what you are looking for in your next opportunity. The conversation will be recorded. A virtual interview with the Hiring Manager and the VP, Engineer Technology. This will be an opportunity to discuss illumin’s business priorities, growth strategy, and the role you could play in supporting the organization. You will have the opportunity to share how your experience and approach can contribute to maintaining a strong infrastructure while learning more about the broader business. An optional final in-person interview if A) the second interview was virtual and/or B) we feel an additional conversation would help clarify the fit. This interview could be with the same illumin team members, or with our CITO, or a mix thereof as necessary. illumin is firmly committed to diversity within its community and welcomes applications from racialized persons/persons of colour, Indigenous People of North America and the world, persons with disabilities, 2SLGBTQIA+ persons, and those who may contribute to the further diversification of ideas. We are committed to providing equitable opportunities in employment and to providing a workplace which is free from discrimination and harassment. We are equally committed to providing an inclusive and accessible workplace. If you require accommodations at any stage of the interview process, please email us at hr@illumin.com.