Senior Monitoring and Observability Engineer

Aptonet Inc — United States · Posted ~5 hours ago

Senior

Skills

monitoring observability Datadog cloud infrastructure Kubernetes automation OpenShift APM cloud

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior reliability engineering position focused on monitoring strategy, cloud and container environments, automation, dashboards, alerting, and operational excellence.

Highlights

Opportunity to design enterprise observability platforms, automate operations, and improve reliability across complex environments.

Description

About the job We are seeking a Senior Monitoring & Observability Engineer to support the client by designing, operating, and continuously improving enterprise monitoring and observability capabilities across hybrid infrastructure, cloud, and containerized environments. This role will focus on monitoring strategy, platform integrations, automation, dashboards, alerting, and service reliability using Datadog and other enterprise observability tools. Responsibilities Design, implement, and maintain enterprise monitoring and observability solutions.Develop and manage dashboards, alerts, logging, metrics, APM, synthetic monitoring, and operational reporting.Assess monitoring coverage across infrastructure and applications, identifying and closing visibility gaps.Support monitoring across Windows, Linux, cloud, Kubernetes/OpenShift, databases, storage, networks, and applications.Deploy, configure, and troubleshoot monitoring agents, integrations, APIs, and platform components.Automate monitoring onboarding, configuration, deployment, and maintenance using scripting, APIs, CI/CD, and Infrastructure-as-Code tools.Integrate observability platforms with ServiceNow, notification systems, and operational workflows.Analyze telemetry and performance data to improve reliability, availability, and incident detection.Support incident response, root cause analysis, and service performance optimization.Partner with operations and engineering teams to improve monitoring standards, alert quality, and operational visibility.Develop reporting and dashboards for technical teams and leadership stakeholders.Maintain monitoring documentation, operational procedures, and platform standards. Required Skills Bachelor's degree and 8+ years of relevant experience, or Master's degree with 6+ years of experience.Strong experience with enterprise monitoring and observability platforms.Hands-on experience with Datadog, or comparable platforms such as Dynatrace, Splunk Observability, ScienceLogic, SolarWinds, New Relic, LogicMonitor, or Prometheus/Grafana.Experience monitoring Windows, Linux, Kubernetes, and/or OpenShift environments.Experience deploying and managing dashboards, alerts, agents, integrations, tagging strategies, and operational reporting.Experience automating monitoring deployments and administration using scripting, APIs, Ansible, CI/CD pipelines, or Infrastructure-as-Code tools.Experience integrating observability platforms with ServiceNow or other ITSM solutions.Strong troubleshooting and dependency analysis skills across infrastructure, applications, and services.Experience analyzing performance, availability, and operational telemetry data.Strong communication skills with the ability to present technical findings to stakeholders.Ability to obtain and maintain SEC Public Trust clearance. Preferred Skills Direct Datadog administration and engineering experience.Experience with Application Performance Monitoring (APM), distributed tracing, and OpenTelemetry.Experience with Datadog APM, Log Management, Synthetic Monitoring, RUM, Network Performance Monitoring, or Database Monitoring.Experience with Red Hat OpenShift, Kubernetes, OpenShift Virtualization, KubeVirt, or related platforms.Experience monitoring Microsoft Azure and AWS environments.Experience with Terraform, monitoring-as-code, and API-driven automation.Experience supporting federal environments governed by FISMA, FedRAMP, or NIST requirements.Relevant certifications in Datadog, Azure, AWS, OpenShift, Terraform, or ITIL.