Senior Site Reliability Engineer

Infinityquest — United Kingdom · Posted ~23 hours ago

Senior

Skills

Site Reliability Engineering Platform Engineering DevOps Splunk Enterprise Splunk Cloud Splunk SPL Elasticsearch ELK Prometheus Grafana Grafana Tempo Distributed Tracing OpenTelemetry Kafka Observability Terraform Infrastructure as Code Python Go Ruby Bash Splunk SPL Kubernetes AWS Azure GCP Ansible Consul CI/CD Service Mesh

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior reliability engineering role focused on building and operating highly observable, resilient platforms. You will work hands-on with log and metric analysis, distributed tracing, telemetry standards, infrastructure as code, cloud environments, event streaming, and automation. Experience with container orchestration, CI/CD, service meshes, and regulated environments is valuable.

Highlights

Work on modern observability and reliability engineering across metrics, logs, traces, distributed tracing, and cloud infrastructure. The role offers exposure to major cloud platforms, infrastructure automation, container orchestration, and regulated environments.

Description

• 7+ years in Site Reliability Engineering, Platform Engineering, or DevOps. • Hands-on experience administering Splunk Enterprise or Splunk Cloud. • Strong knowledge of Splunk SPL. • Experience with Elasticsearch/ELK, Prometheus, Grafana, Grafana Tempo, distributed tracing, OpenTelemetry, and Kafka. • Experience implementing metrics, logs, and traces as part of a modern observability strategy. • Experience with Terraform and Infrastructure as Code. • Programming experience in Python, Go, Ruby, or Bash. Good to have skills – Splunk certification. • Experience with Kubernetes, AWS/Azure/GCP, Ansible, Consul, CI/CD pipelines, and service mesh technologies. • Experience supporting FedRAMP or regulated environments.