Senior Site Reliability Engineer

Jobgether — Netherlands · Posted ~4 hours ago

Senior Remote

Skills

Kubernetes cloud infrastructure Infrastructure as Code Observability CI/CD SLOs Alerting Incident response Reliability engineering AI workflows Cloud infrastructure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Be part of a remote SRE team that builds and maintains resilient infrastructure across Kubernetes, cloud services, and modern DevOps practices. You will shape observability, automate CI/CD pipelines, and integrate AI workflows to boost reliability and engineering efficiency.

Highlights

We design and operate highly reliable, scalable platforms in a fully remote setting, leveraging modern cloud and Kubernetes practices. The role drives robust observability, automated CI/CD, and AI‑enhanced workflows to improve engineering productivity.

Description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in Netherlands. This is a senior-level SRE opportunity focused on solving complex reliability, infrastructure, and platform challenges in a fully remote environment. You will take ownership of high-impact projects, from solution discovery through delivery, while helping shape platform architecture and long-term reliability strategy. Your work will span Kubernetes, cloud infrastructure, infrastructure as code, observability, CI/CD, and operational excellence. You will also play a key role in establishing strong SLOs, alerting, incident response, and reliability practices across engineering. AI is embedded into the way the team works, with an emphasis on building practical, reusable AI workflows that improve engineering productivity and safety. The role offers significant autonomy, technical influence, and opportunities to mentor engineers while collaborating asynchronously across a global organization. Accountabilities Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust and maintainable technical solutions.Contribute to platform architecture, infrastructure tooling, technical roadmaps, and engineering priorities, advocating for initiatives that improve reliability and developer experience.Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.Use operational and incident metrics to identify systemic reliability issues and influence the team's technical strategy.Resolve cross-team infrastructure and platform requests while turning recurring problems into reusable solutions, automation, documentation, and runbooks.Operate and scale production Kubernetes environments and associated container infrastructure.Build and manage cloud infrastructure using AWS or comparable cloud platforms, with strong emphasis on reliability, scalability, and operational efficiency.Develop and maintain infrastructure as code using Terraform and support automated deployment workflows through modern CI/CD platforms.Use AI natively in infrastructure, operations, and development workflows, creating reusable prompts, skills, tooling, and agentic workflows that improve team-wide productivity and reliability.Design infrastructure and systems with AI-assisted engineering in mind, including clean interfaces, strong observability, secure-by-default patterns, CI protections, and review guardrails.Mentor less-senior engineers through actionable feedback, technical guidance, hiring, onboarding, and RFC discussions.Collaborate with Security on infrastructure hardening, threat mitigation, and defensive security practices.Contribute to infrastructure capacity planning, performance optimization, and cost efficiency.Participate in incident response and on-call rotations, helping maintain high standards of platform availability and reliability. Requirements Solid professional experience in Site Reliability Engineering, DevOps, Platform Engineering, or a closely related discipline.Strong hands-on experience operating and scaling Kubernetes in production, including Docker and the wider container ecosystem.Proven experience designing, building, and managing production cloud infrastructure using AWS or a comparable cloud provider.Strong practical expertise with Terraform and infrastructure-as-code principles.Hands-on experience with reliability engineering frameworks, including SLOs, SLIs, error budgets, alerting, and incident management.Strong observability experience with technologies such as OpenTelemetry, Grafana, Prometheus, or equivalent platforms.Experience designing and operating CI/CD pipelines using GitLab CI, GitHub Actions, or similar technologies.Comfortable working with Golang, Bash, and scripting, with broader programming experience considered an advantage.Demonstrated practical use of AI and agentic workflows in infrastructure, operations, or software engineering, with measurable outcomes beyond simple familiarity with AI tools.Strong understanding of production operations, troubleshooting, automation, and platform reliability.Clear, thoughtful communication skills, particularly in an asynchronous and globally distributed environment.Proactive and curious mindset, with the ability to independently identify problems, take ownership, and drive solutions to completion.Collaborative and respectful approach when working across cultures, time zones, and diverse teams.Experience with a backend programming language such as Elixir, Node.js, Python, or similar is a plus.Experience operating and configuring Linux systems outside cloud environments is beneficial.Security knowledge, including defensive and offensive security concepts, is an advantage. Benefits 100% remote work, with the ability to work from anywhere.Flexible working hours within an async-first working environment.Flexible paid time off to support a healthy balance between work and personal life.16 weeks of paid parental leave.Budget for co-working spaces, learning, and wellness, including gym memberships.Mental health support services.Stock options.Home office budget and IT equipment.Competitive, location-aware compensation designed to support fair and equitable pay across global markets.Annual salary range of $53,300–$119,850 USD, with actual compensation determined by location, experience, skills, training, business needs, and market conditions.Significant autonomy and ownership in a globally distributed engineering organization.Opportunities to influence platform architecture, technical strategy, engineering standards, and reliability practices.Supportive environment focused on professional growth, inclusion, collaboration, and continuous improvement. How Jobgether Works We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.