Summary
✨ AI‑Generated
Lead the design and operation of enterprise monitoring and observability solutions. Build telemetry pipelines, instrumentation, dashboards, alerts, distributed tracing, and SLO/SLI monitoring across cloud and containerized environments, while developing automated remediation and self-healing capabilities.
Highlights
Lead enterprise observability initiatives focused on reliability, reduced incident resolution time, proactive monitoring, automated remediation, and embedding reliability practices across engineering teams.
Description
Lead Platform Engineer – Monitoring & ObservabilityJob SummarySeeking a Lead Platform Engineer with strong SRE and Observability experience to design, implement, and support enterprise monitoring solutions.
The role focuses on improving reliability, reducing MTTR, and building proactive monitoring, alerting, dashboards, and automated remediation.Key Responsibilities
Design and manage enterprise monitoring and observability solutions using OpenTelemetry, Elastic, Grafana, OpsRamp, and BigPanda.Build telemetry pipelines, instrumentation, dashboards, alerts, distributed tracing, and SLO/SLI monitoring.Implement monitoring across AWS/Azure, Kubernetes, Docker, and distributed applications.Develop actionable alerts, escalation workflows, and automated remediation/self-healing solutions.Partner with application, platform, and engineering teams to embed observability and reliability into the SDLC.Define and operationalize SLIs, SLOs, error budgets, and reliability standards.Troubleshoot production issues, improve MTTR, and drive continuous reliability improvements.Create documentation, runbooks, and observability standards; mentor junior engineers.Mandatory Skills
5+ years in Software, Systems, SRE, or Reliability Engineering.Strong hands-on experience with Monitoring & Observability in large-scale environments.OpenTelemetry, Elastic Observability (APM/Logs/Metrics/Traces), Grafana, OpsRamp, BigPanda.AWS CloudWatch and/or Azure Monitor.Strong knowledge of metrics, logs, traces, distributed tracing, and SRE practices.Strong scripting/programming skills in Bash, PowerShell, Python, JavaScript, or C-family.Experience with AWS/Azure, Kubernetes, and Docker.Understanding of distributed systems, networking, security, databases, DevSecOps, and performance engineering.Strong troubleshooting, communication, prioritization, and cross-functional collaboration skills.Desired Skills
CI/CD: Jenkins, GitHub, GitLab CI, CircleCI.IaC: Terraform, Ansible.Time-series data and visualization.Microservices and event-driven architectures.REST APIs, JSON, ServiceNow.Additional AWS/Azure monitoring experience.