Senior DevOps Engineer

Jobgether — Canada · Posted ~2 hours ago

Senior Full-time Remote

Skills

Google Cloud Infrastructure as Code Kubernetes CI/CD Observability Networking Reliability engineering Cloud architecture Terraform

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior DevOps role within a globally distributed infrastructure team supporting highly available, production-critical financial technology. You will architect and operate cloud infrastructure, build Infrastructure-as-Code and Kubernetes capabilities, improve CI/CD and observability, and create self-service tooling that increases engineering productivity.

Highlights

Join a globally distributed infrastructure team, work asynchronously with significant autonomy, own cloud platform capabilities, automate operations, and directly influence technical decisions.

Description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior DevOps Engineer based in Canada. Join a globally distributed infrastructure team responsible for powering highly available, production-critical financial technology. You will design, automate, and operate cloud infrastructure on Google Cloud while helping engineering teams ship faster and more safely. The role combines cloud architecture, Infrastructure-as-Code, Kubernetes, CI/CD, observability, networking, and reliability engineering. You will take ownership of platform capabilities and build self-service solutions that reduce manual work and improve developer productivity. You will also contribute to incident response, security, capacity planning, and continuous improvements across the infrastructure environment. Working in an async-first international setting, you will have significant autonomy and a direct influence on technical decisions and platform evolution. This is an opportunity to solve complex infrastructure challenges at scale while applying strong Platform-as-a-Product and SRE principles. Accountabilities Design and evolve highly available cloud architecture on Google Cloud, including networking, interconnects, IAM, and resilient infrastructure topologies, using Terraform and GitOps practices.Build and maintain secure CI/CD pipelines for Infrastructure-as-Code, incorporating automated planning and testing, code review, Policy-as-Code guardrails, drift detection, and controlled rollouts.Develop a Platform-as-a-Product approach by creating self-service capabilities, reusable infrastructure patterns, and developer-friendly golden paths that enable teams to provision resources efficiently.Operate and improve production Kubernetes/GKE environments, including workload deployment with Helm, scaling, networking, security, observability, and troubleshooting.Strengthen platform observability across metrics, logs, traces, and alerting using technologies such as Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager.Operate infrastructure services and data platforms, including PostgreSQL and message brokers, while partnering with specialized SRE and database teams on complex operational challenges.Participate in a global Follow-The-Sun on-call rotation, responding to alerts, supporting incidents, leading structured troubleshooting, coordinating escalations, and driving blameless post-mortems and follow-up actions.Apply SRE principles such as SLIs, SLOs, error budgets, and capacity planning to improve reliability and operational maturity.Design and troubleshoot cloud networking components including VPCs, routing, load balancing, DNS, TLS, firewalls, and interconnects.Partner with security and engineering teams to implement secure-by-default infrastructure, least-privilege access, and effective infrastructure security practices.Identify opportunities to eliminate manual toil through automation and continuously improve reliability, security, scalability, cost efficiency, and developer experience.Mentor engineers and lead infrastructure initiatives that raise engineering standards and improve the overall platform. Requirements 5+ years of professional experience in DevOps, SRE, Platform Engineering, Infrastructure Engineering, or a closely related discipline, with experience operating large-scale, highly available production systems.Deep hands-on experience designing and operating cloud architecture on Google Cloud Platform, including landing zones, networking, IAM, and high-availability architectures.Strong Terraform and Infrastructure-as-Code expertise, including structuring large codebases across multiple environments and applying GitOps and least-privilege principles.Proven experience building CI/CD pipelines for Infrastructure-as-Code, including automated plan/apply workflows, code review, Policy-as-Code, drift detection, and safe deployment practices.Significant production experience with Kubernetes, ideally Google Kubernetes Engine (GKE), and Helm-based workload deployment.Strong understanding of cloud and L3/L4-L7 networking fundamentals, including VPCs, routing, load balancing, DNS, TLS, firewalls, and interconnects, with the ability to troubleshoot complex connectivity issues.Practical experience with modern observability platforms covering metrics, logs, traces, and alerting, particularly Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager.Operator-level knowledge of production data stores and messaging systems such as PostgreSQL, RabbitMQ, or similar message brokers.Solid understanding of SRE principles, including SLOs, error budgets, capacity planning, incident management, and root-cause analysis.Strong scripting or programming capabilities in Python, Go, Shell, or equivalent technologies.Demonstrated ability to independently troubleshoot complex production problems and drive issues through resolution.Strong ownership, communication, documentation, and problem-solving skills, with the ability to work effectively in a globally distributed and async-first environment.Willingness to participate in a Follow-The-Sun on-call rotation, including scheduled availability for urgent infrastructure incidents.Experience with Policy-as-Code and IaC quality tools such as OPA/Conftest, Checkov, tflint, or Atlantis is a plus.Experience managing Terraform state, module registries, and versioning at scale is a plus.Experience building internal developer platforms or golden paths using tools such as Backstage or Tilt is a plus.Working knowledge of Go, Linux, Docker/containerd, Ansible, Chef, or Puppet is advantageous.Experience securing containers and Kubernetes/GKE environments is beneficial.Knowledge of advanced Google Cloud security controls, regulated environments, SOC 2, secrets management, or audit logging is advantageous.Familiarity with trading, brokerage, fintech, or other regulated and low-latency environments is a plus.Relevant cloud certifications, particularly Google Cloud Professional certifications, are valued. Benefits Competitive salary and stock options.Health benefits.One-time USD $500 home-office setup allowance for new hires.USD $150 monthly stipend provided through a company expense card.Fully remote work within the eligible location.Opportunity to work with a globally distributed team across multiple regions.High level of autonomy, ownership, and influence over infrastructure and platform strategy.Opportunity to work on highly available, trading-critical systems and complex cloud infrastructure challenges.A strong focus on developer productivity, automation, reliability, and Platform-as-a-Product practices.Exposure to modern cloud, Kubernetes, Infrastructure-as-Code, observability, and SRE technologies.Inclusive environment committed to building a diverse and collaborative workforce. How Jobgether Works We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.