Site Reliability Engineer - Observability and Monitoring

Data Nexus Ai — United States · Posted ~4 hours ago

Full-time

Skills

Site Reliability Engineering observability monitoring distributed tracing OpenTelemetry Prometheus Grafana AppDynamics Splunk telemetry pipelines incident response Root Cause Analysis SLI SLO error budgets

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An SRE with strong observability and monitoring expertise is sought to design, build, and scale an end-to-end telemetry framework. The role emphasizes OpenTelemetry, metrics/logs/traces pipelines, distributed tracing, incident response, root cause analysis, and reliability practices based on SLIs, SLOs, and error budgets.

Highlights

Design and scale an end-to-end observability framework, standardize telemetry across logs, metrics, and traces, improve incident detection and response, and use reliability metrics to continuously strengthen systems.

Description

Sire Reliability Engineer (SRE) – Observability & Monitoring About the Role We are seeking a Site Reliability Engineer (SRE) with strong expertise in observability, monitoring, and distributed tracing to join our SRE team. The ideal candidate will help us design, build, and scale an observability framework that provides end-to-end visibility into our systems and applications. A strong focus will be places on OpenTelemetry, as we continue to standardize our telemetry pipeline across logs, metrics, and traces. Responsibilities Design, implement, and maintain observability solutions using OpenTelemetry, Prometheus, Grafana, AppDynamics, and Splunk.Build and manage telemetry pipelines (metrics, logs, traces) ensuring reliable data collection, transformation, and export.Lead initiatives to improve incident detection, response, and post-incident analysis with a strong emphasis on RCA (Root Cause Analysis).Define and maintain SLIs, SLOs, and error budgets to measure and improve system reliability.Partner with development and operations teams to instrument applications and services for better monitoring and tracing coverage.Develop dashboards, alerts, and visualizations to provide actionable insights into system health and performance.Contribute to automation and self-healing practices that improve uptime and reduce operational toil.Stay current with trends in observability and advocate best practices across the engineering organization.Requirements 7+ years of SRE/ Devops/ Cloud/ Infrastructure engineering experience with a focus on monitoring and observability.Strong communication skills with the ability to articulate technical requirements and explain benefits of observability clearly to development teams.Experience working in an area that requires strong business knowledge and the ability to pull together business parters and development teams to perform RCA across distributed systems.Experience implementing an Observability Framework at an organization.Hands-on experience with OpenTelemetry SDKs, collectors, and exporters.Proficiency with observability stacks such as Prometheus, Grafana, Loki, Tempo, Elastic Stack, or Splunk Observability (Splunk/AppDynamics).Strong knowledge on cloud platforms (GCP, or Azure).Hand-on experience on container orchestration using Kubernetes (OCP, GKE, AKS)Familiarity with CI/CD pipelines like (Jenkins and Github actions), infrastructure as code (Terraform/Ansible/ARM/CloudFormation).Experience provisioning infrastructure and capacity planning.Hands-on skills in programming languages like Java, and Python.