Description
About the job
We are seeking a Senior Monitoring & Observability Engineer to support the client by designing, operating, and continuously improving enterprise monitoring and observability capabilities across hybrid infrastructure, cloud, and containerized environments.
This role will focus on monitoring strategy, platform integrations, automation, dashboards, alerting, and service reliability using Datadog and other enterprise observability tools.
Responsibilities
Design, implement, and maintain enterprise monitoring and observability solutions.Develop and manage dashboards, alerts, logging, metrics, APM, synthetic monitoring, and operational reporting.Assess monitoring coverage across infrastructure and applications, identifying and closing visibility gaps.Support monitoring across Windows, Linux, cloud, Kubernetes/OpenShift, databases, storage, networks, and applications.Deploy, configure, and troubleshoot monitoring agents, integrations, APIs, and platform components.Automate monitoring onboarding, configuration, deployment, and maintenance using scripting, APIs, CI/CD, and Infrastructure-as-Code tools.Integrate observability platforms with ServiceNow, notification systems, and operational workflows.Analyze telemetry and performance data to improve reliability, availability, and incident detection.Support incident response, root cause analysis, and service performance optimization.Partner with operations and engineering teams to improve monitoring standards, alert quality, and operational visibility.Develop reporting and dashboards for technical teams and leadership stakeholders.Maintain monitoring documentation, operational procedures, and platform standards.
Required Skills
Bachelor's degree and 8+ years of relevant experience, or Master's degree with 6+ years of experience.Strong experience with enterprise monitoring and observability platforms.Hands-on experience with Datadog, or comparable platforms such as Dynatrace, Splunk Observability, ScienceLogic, SolarWinds, New Relic, LogicMonitor, or Prometheus/Grafana.Experience monitoring Windows, Linux, Kubernetes, and/or OpenShift environments.Experience deploying and managing dashboards, alerts, agents, integrations, tagging strategies, and operational reporting.Experience automating monitoring deployments and administration using scripting, APIs, Ansible, CI/CD pipelines, or Infrastructure-as-Code tools.Experience integrating observability platforms with ServiceNow or other ITSM solutions.Strong troubleshooting and dependency analysis skills across infrastructure, applications, and services.Experience analyzing performance, availability, and operational telemetry data.Strong communication skills with the ability to present technical findings to stakeholders.Ability to obtain and maintain SEC Public Trust clearance.
Preferred Skills
Direct Datadog administration and engineering experience.Experience with Application Performance Monitoring (APM), distributed tracing, and OpenTelemetry.Experience with Datadog APM, Log Management, Synthetic Monitoring, RUM, Network Performance Monitoring, or Database Monitoring.Experience with Red Hat OpenShift, Kubernetes, OpenShift Virtualization, KubeVirt, or related platforms.Experience monitoring Microsoft Azure and AWS environments.Experience with Terraform, monitoring-as-code, and API-driven automation.Experience supporting federal environments governed by FISMA, FedRAMP, or NIST requirements.Relevant certifications in Datadog, Azure, AWS, OpenShift, Terraform, or ITIL.