Senior Observability Engineer

Lancesoft — Canada · Posted ~1 day ago

Senior

Skills

Kubernetes Prometheus Grafana Thanos Loki Metrics monitoring Centralized logging Distributed tracing Alerting SLI/SLO monitoring GitOps Infrastructure as Code Multi-cluster observability Grafana dashboard development RBAC Performance monitoring ELK Splunk Fluent Bit Fluentd Promtail Vector LogQL Lucene

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An experienced observability engineer is needed to lead monitoring across a large enterprise Kubernetes environment. You will manage metrics, logs, traces, alerts, service-level indicators, and objectives, while building dashboards, access controls, and multi-cluster monitoring capabilities. Strong hands-on expertise with cloud-native monitoring tools, GitOps, infrastructure as code, and centralized logging is required. Exposure to predictive monitoring and anomaly detection is beneficial.

Highlights

Own enterprise observability across more than 50 production Kubernetes clusters in a financial services environment. Build advanced monitoring, logging, tracing, and alerting capabilities, develop dashboards and SLI/SLO systems, and work with modern cloud-native infrastructure and automation practices.

Description

We are looking for an experienced Observability Engineer to own enterprise observability across 50+ production Kubernetes clusters in a financial services environment. Strong hands-on experience with Kubernetes, Prometheus, Grafana, Thanos, Loki and modern observability/collection agents.Build and manage metrics, logging, tracing, alerting, ServiceMonitors/PodMonitors, recording rules, and SLI/SLO monitoring.Experience with GitOps, Infrastructure as Code, multi-cluster observability, and cloud object storage.Develop Grafana dashboards, dashboard-as-code, RBAC/multi-tenancy, and performance/business KPI monitoring.Strong centralized logging experience with Loki, ELK, Splunk, Fluent Bit/Fluentd/Promtail/Vector and LogQL/Lucene.Exposure to AI/ML-driven monitoring, predictive alerting, anomaly detection, and self-healing infrastructure is highly preferred.