Senior Site Reliability Engineer

Salt — Netherlands · Posted ~1 day ago

Senior Contract Hybrid €90-€120 per hour

Skills

Site Reliability Engineering Production operations Incident response Root cause analysis Observability Monitoring Alerting Logging Tracing CI/CD Infrastructure as Code Automation Distributed systems Cloud migration Cloud Distributed Systems

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An experienced SRE is sought for a high-performing platform team operating a mission-critical distributed system at enormous scale. You will own production reliability, lead incident response and root-cause analysis, improve observability, build CI/CD and infrastructure automation, and contribute to a major cloud migration. The role is offered on a six-month hybrid contract.

Highlights

High-impact work on large-scale distributed systems processing billions of events, a strong focus on reliability and operational excellence, involvement in a major cloud migration, and a hybrid schedule with a competitive hourly rate.

Description

Senior Site Reliability Engineer (SRE) – Travel – Amsterdam Hourly rate: €90 - €120 Duration: 6-Month Contract Start: ASAP Hybrid: 2 days a week (on-call 1/7) My client is looking for an experienced Site Reliability Engineer to join a high-performing platform team responsible for operating and evolving a mission-critical event processing platform that handles billions of events every day. This is an exciting opportunity to work on large-scale distributed systems, drive operational excellence, and support a major cloud migration initiative while ensuring the reliability, scalability, and performance of a business-critical platform. What You'll Be Doing Own end-to-end reliability of production services.Lead incident response, root cause analysis, post-mortems, and remediation activities.Improve observability through monitoring, alerting, logging, tracing, and dashboards.Build and enhance CI/CD pipelines and Infrastructure-as-Code solutions.Drive automation initiatives to reduce operational toil and improve platform efficiency.Support performance testing, capacity planning, and scalability initiatives.Contribute to the migration of a large-scale event streaming platform to a cloud-native architecture.Collaborate with software engineers and platform teams to improve reliability, resilience, and operational maturity.Participate in a shared on-call rotation and help maintain high service availability. What We're Looking For Proven experience as a Site Reliability Engineer, SRE, Platform Engineer, or DevOps Engineer in high-scale production environments.Strong hands-on experience with:JavaKafkaKubernetesAWSTerraform, Helm, GitOps, or similar Infrastructure-as-Code toolsCI/CD pipelines and deployment automationExperience with modern observability tooling such as Prometheus, Grafana, OpenTelemetry, ELK/EFK, Datadog, or similar.Strong background in incident management, production troubleshooting, and reliability engineering.Experience operating and scaling distributed systems handling high transaction or event volumes.Excellent communication skills and the ability to work effectively across engineering teams. Nice to Have Experience with Confluent Cloud.Experience migrating Kafka workloads from on-premises environments to cloud platforms.Knowledge of large-scale event-driven architectures and stream processing systems.Experience optimising platform performance, scalability, and operational costs.