Senior Observability Engineer

Dexiansolutions — United States · Posted ~3 hours ago

Senior

Skills

Grafana OpenTelemetry Monitoring Alerting Terraform Infrastructure automation

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology organization is seeking an experienced observability engineer to design and operate enterprise monitoring platforms. The role focuses on metrics, logs, traces, automation, and helping teams adopt consistent reliability practices.

Highlights

Lead modernization of enterprise monitoring systems and build scalable observability platforms using modern infrastructure tools.

Description

We're looking for an experienced Senior Observability Engineer to join our Observability Engineering team. This role focuses on designing, implementing, managing, and automating enterprise‑level observability solutions, with emphasis on Grafana, OpenTelemetry, monitoring, alerting, and telemetry operations. The ideal candidate has deep experience building and running observability platforms at scale, automating infrastructure with Terraform, and helping application and infrastructure teams adopt consistent observability standards. This position will play a major role in modernizing our observability stack and leading the transition from legacy monitoring tools to Grafana‑based solutions. Key Responsibilities Observability Platform Engineering Administer and support Grafana Cloud and on‑prem Grafana environments. Build and deploy observability solutions for metrics, logs, traces, synthetic monitoring, and alerting. Define and maintain observability standards, best practices, and governance. Configure and manage Grafana data sources, alerting, RBAC, folders, teams, and integrations. Ensure the platform is scalable, reliable, resilient, and operationally sound. Automation & Infrastructure as Code Develop and maintain Terraform modules for Grafana infrastructure and configuration. Automate onboarding for applications, dashboards, alerts, and data sources. Build self‑service capabilities to reduce manual work and improve adoption. Integrate observability into CI/CD and infrastructure provisioning workflows. Monitoring, Alerting & Incident Management Design monitoring and alerting strategies aligned to service health and critical workflows. Improve alert quality by reducing noise and strengthening signal accuracy. Support incident response, troubleshooting, RCA, and post‑incident reviews. Continuously enhance operational visibility and platform health. OpenTelemetry & Telemetry Engineering Implement and support OpenTelemetry instrumentation across applications and infrastructure. Define standards for logs, metrics, traces, and telemetry collection. Support telemetry pipelines, agent deployments, and data collection strategies. Guide teams on instrumentation design and observability adoption. Migration & Modernization Lead migration efforts from legacy monitoring tools to Grafana. Assess existing monitoring, logging, alerting, and tracing setups and recommend modernization paths. Build reusable migration patterns, automation, and engineering standards. Partner with application teams to accelerate enterprise observability adoption. Collaboration & Leadership Work closely with development, infrastructure, cloud, and SRE teams. Provide technical leadership and mentorship across engineering groups. Contribute to observability architecture, strategy, and roadmap. Champion observability as a core engineering discipline across the organization. Required Qualifications Bachelor's degree in Computer Science, Engineering, Information Systems, or related field. 5+ years in observability, monitoring, operations, or platform engineering. Hands‑on experience administering Grafana in large enterprise environments. Strong Terraform and IaC experience. Experience with monitoring, alerting, logging, and distributed tracing. Knowledge of OpenTelemetry instrumentation and telemetry pipelines. Strong Linux and cloud administration skills. Scripting experience with Python, PowerShell, Bash, or similar. Understanding of operational excellence, reliability engineering, and incident management. Preferred Qualifications Experience migrating from Splunk, Dynatrace, AppDynamics, New Relic, OpenText OBM, etc. Experience with Grafana Alloy, Tempo, Loki, Mimir, or Prometheus. Experience running observability platforms in AWS. Knowledge of Kubernetes, containers, and cloud‑native observability. Experience designing enterprise observability strategies and governance. Familiarity with CI/CD and DevOps practices. Desired Skills Grafana Administration Terraform OpenTelemetry (OTEL) Monitoring & Alerting Observability Engineering Platform Engineering Linux Administration AWS Cloud Services Automation & Scripting Incident Management Infrastructure as Code Telemetry Pipelines Reliability Engineering Root Cause Analysis Enterprise Monitoring Architecture Dexian stands at the forefront of Talent + Technology solutions with a presence spanning more than 70 locations worldwide and a team exceeding 10,000 professionals. As one of the largest technology and professional staffing companies and one of the largest minority-owned staffing companies in the United States, Dexian combines over 30 years of industry expertise with cutting-edge technologies to deliver comprehensive global services and support.     Dexian connects the right talent and the right technology with the right organizations to deliver trajectory-changing results that help everyone achieve their ambitions and goals. To learn more, please visit https://dexian.com/. Dexian is an Equal Opportunity Employer that recruits and hires qualified candidates without regard to race, religion, sex, sexual orientation, gender identity, age, national origin, ancestry, citizenship, disability, or veteran status.