Summary
✨ AI‑Generated
A senior engineering role within a globally distributed reliability organization, responsible for operating and improving critical services at significant scale. You will apply an automation-first mindset across private and public cloud environments, using Infrastructure as Code and GitOps to build reusable services, automate operations, and improve reliability, scalability, performance, and observability. The position combines hands-on engineering with architecture, technical leadership, and operational responsibility.
Highlights
Senior-level opportunity within a globally distributed engineering team, focused on large-scale infrastructure, automation, cloud environments, reliability, scalability, performance, and observability. The role offers substantial technical ownership spanning hands-on engineering, architecture, and operational leadership.
Description
This position is listed on behalf of a partner company, who manages all applications and next steps.
Our partner is looking for a Senior Site Reliability / Gitops Engineer based in Canada.
Join a globally distributed Site Reliability Engineering team responsible for operating critical IT services at significant scale.
You will drive an automation-first approach across private and public cloud environments, using Infrastructure as Code and GitOps practices to improve reliability and efficiency.
The role combines hands-on engineering, technical leadership, architecture, and operational responsibility across complex infrastructure.
You will help design reusable services, automate software operations, and strengthen infrastructure practices through modern engineering methods.
Your work will directly influence the reliability, scalability, performance, and observability of production systems used across a global organization.
You will collaborate with engineers, architects, operations teams, and support specialists while contributing feedback and improvements to open-source technologies.
This is an opportunity to take ownership of challenging infrastructure initiatives while mentoring others and shaping next-generation SRE and automation practices.
Accountabilities
Drive the development of automation and GitOps practices within the team while acting as an embedded technical lead.Collaborate closely with the infrastructure architecture function to align technical solutions with broader architecture objectives.Design and architect infrastructure services that can be delivered as reusable products across the organization.Develop and strengthen Infrastructure as Code practices by continuously improving automation, processes, consistency, and reusability.Automate software operations across private and public clouds while accounting for the complexity and operational requirements of distributed systems.Maintain operational responsibility for core services, networks, and infrastructure, ensuring reliability and continuity.Troubleshoot complex infrastructure issues, support capacity planning, investigate performance challenges, and improve system resilience.Implement and maintain observability, monitoring, and alerting solutions using technologies such as Prometheus, Grafana, and Elasticsearch.Collaborate with globally distributed engineering, operations, and support teams to resolve technical challenges and deliver reliable services.Dedicate focused development time to larger engineering projects and the automation of repetitive or manual operational tasks.Share technical knowledge, experience, and best practices through design sessions, mentorship, collaborative implementation, and team development.Take final responsibility for resolving time-critical escalations and ensuring appropriate technical follow-through.
Requirements
Strong understanding of modern hosting architectures and an automation-first approach based on Infrastructure as Code across private and public cloud environments.A product-oriented mindset with an interest in building reusable infrastructure products rather than one-off solutions.Professional experience with Python development, including work on large or complex projects.Hands-on experience with Kubernetes or other container orchestration technologies.Proven experience managing and deploying cloud infrastructure through code and automation.Practical knowledge of Linux networking, routing, firewalls, and related infrastructure concepts.Familiarity with Linux storage technologies, ranging from distributed storage such as Ceph to database-backed systems.Hands-on experience administering enterprise Linux servers.Strong knowledge of cloud computing concepts, architectures, and technologies.Bachelor’s degree or higher, preferably in Computer Science, Engineering, or a related technical discipline.Strong English communication skills across email, chat, video calls, voice communication, and in-person collaboration.Ability to troubleshoot problems across the technology stack, from the Linux kernel through application and web layers, while knowing when to seek input from others.Flexibility, curiosity, and the ability to learn new technologies and approaches quickly.Ability to adapt to fast-changing technical environments and remain focused on delivering reliable outcomes.Experience working effectively within globally distributed teams.Passion for open-source technologies, with familiarity with Ubuntu or Debian considered valuable.
Benefits
Compensation shaped according to geographic location, experience, and performance, with regular compensation reviews.Performance-driven annual bonus or commission in addition to base compensation.Distributed work environment with twice-yearly in-person team sprints.USD 2,000 annual personal learning and development budget.Recognition and performance rewards.Annual holiday leave.Maternity and paternity leave.Team Member Assistance Program and Wellness Platform.Opportunities to travel internationally and meet colleagues in new locations.Priority Pass and travel upgrades for long-haul company events.Opportunity to work on large-scale cloud infrastructure, SRE, GitOps, Infrastructure as Code, automation, observability, and open-source technologies.Dedicated development time for larger technical initiatives and automation projects.Opportunities to provide technical mentorship and influence infrastructure engineering practices.
How Jobgether Works
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements.
Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company.
The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer.
This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR).
You may exercise your rights (access, rectification, erasure, objection) at any time.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information.
These tools assist our recruitment team but do not replace human judgment.
Final hiring decisions are ultimately made by humans.
If you would like more information about how your data is processed, please contact us.