AWS SRE / Observability Engineer

Marks Sattin — United Kingdom · Posted ~3 hours ago

Senior Full-time

Skills

AWS Grafana observability monitoring alerting SLI SLO error budgets incident response CloudWatch EC2 ECS EKS Lambda RDS DynamoDB S3 API Gateway Prometheus

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An AWS SRE Engineer with deep Grafana and observability expertise is sought to design and deliver monitoring, alerting, and dashboards. The role covers SLI/SLO management, error budgets, incident response, telemetry integration, and monitoring of a broad range of AWS services.

Highlights

Own production-grade observability across a large AWS environment, design actionable dashboards, define SRE reliability practices, integrate metrics, logs, and traces, and influence platform and application teams.

Description

An established, technology-focused organisation is looking for an AWS SRE Engineer with a deep technical focus on Grafana and observability. You’ll own the design and delivery of monitoring, alerting and dashboards across a large AWS environment, turning metrics, logs and traces into clear, actionable insight for engineering teams and stakeholders. Key Responsibilities Own Grafana: design, build and maintain production-grade dashboards for application/infrastructure health, performance and SLO/SLI reporting.Observability: define metrics (golden signals, RED/USE), integrate logs and traces, and create unified views across services.SRE Practice: define and track SLIs/SLOs, manage error budgets, tune alerting policies and support incident response and post-incident review.AWS Integration: integrate Grafana with CloudWatch and other data sources to monitor EC2, ECS/EKS, Lambda, RDS, DynamoDB, S3, API Gateway, etc.Collaboration: work closely with platform, data and application teams on instrumentation and best practice; provide guidance and training on observability. Required Experience 6+ years in SRE / Observability / Cloud Reliability roles in AWS.Strong, hands-on expertise with Grafana (dashboards, alerts, RBAC, multiple data sources).Solid experience with metrics, logs and traces and operating observability in production.Strong understanding of SLA, SLO, SLI, error budgets and golden signals.Proven track record troubleshooting live production systems using observability tooling. Nice to Have Exposure to Snowflake or Databricks from an observability perspective.Experience with tools such as Prometheus, Loki, OpenTelemetry, OpenSearch.Infrastructure as Code (e.g. Terraform/CloudFormation) for observability and Grafana resources.