Observability Engineer

Apptoza Inc — Canada · Posted ~2 hours ago

Senior Hybrid

Skills

Kubernetes Containerization GitOps Automation Observability Metrics Logging Distributed tracing Alerting Cloud infrastructure Query languages AI/ML Docker Cloud

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A major financial services organization is seeking an experienced Observability Engineer for an enterprise Kubernetes platform. You will own monitoring across more than 50 production clusters, covering metrics, logs, tracing, and alerting, while developing intelligent monitoring, predictive alerting, and self-healing infrastructure with AI/ML capabilities. The role is hybrid with four days on site.

Highlights

Own a large-scale observability environment spanning dozens of production Kubernetes clusters, with a strong focus on reliability, performance, automation, and emerging AI/ML monitoring capabilities.

Description

Observability Engineer (KUBERNETES & CONTAINERS, GITOPS & AUTOMATION, QUERY LANGUAGES, CLOUD & INFRASTRUCTURE) Toronto, On Hybrid - 4 days ABOUT THE ROLE We are seeking an experienced Observability Engineer to join our Enterprise Kubernetes Platform team at a leading financial services organization. You'll own the complete observability stack across 50+ production Kubernetes clusters, providing metrics, logging, tracing, and alerting capabilities that ensure exceptional reliability and performance for mission-critical applications. This role combines deep technical expertise in modern observability tools with emerging AI/ML capabilities to build intelligent monitoring solutions, predictive alerting, and self-healing infrastructure. ======================================================================== WHAT YOU'LL DO ======================================================================== OBSERVABILITY STACK OWNERSHIP -------------------------------------------------------------------------------- • Design, deploy, and maintain enterprise-scale observability infrastructure including Prometheus, Grafana, Thanos, Loki, and modern collection agents • Manage observability deployments using GitOps principles and infrastructure as code • Implement long-term metrics storage solutions with cloud object storage • Maintain and upgrade observability components across development, QA, UAT, production, and DR environments • Configure distributed observability architecture spanning multiple data centers and cloud providers METRICS & MONITORING -------------------------------------------------------------------------------- • Design and implement Prometheus monitoring strategies for Kubernetes infrastructure and containerized applications • Create ServiceMonitors, PodMonitors for automated metrics collection • Develop rules for intelligent alerting with minimal false positives • Configure multi-cluster metrics federation and aggregation • Optimize metrics cardinality, storage e_iciency, and query performance • Implement recording rules for pre-aggregated metrics and SLI calculations DASHBOARDS & VISUALIZATION -------------------------------------------------------------------------------- • Build comprehensive Grafana dashboards for infrastructure health, application performance, and business metrics • Create reusable dashboard templates and libraries for development teams • Implement dashboard-as-code • Configure multiple datasources • Design executive dashboards with SLO/SLI tracking and business KPIs • Implement role-based access control and multi-tenancy in Grafana LOGGING INFRASTRUCTURE -------------------------------------------------------------------------------- • Deploy and manage centralized logging solutions (Loki, ELK, Splunk, or similar) • Configure log collection agents (Promtail, Fluentd, FluentBit, Vector, etc.) • Design log retention policies balancing cost, compliance, and operational needs • Create LogQL/Lucene queries and log-based alerts • Implement log correlation with metrics and traces for unified troubleshooting • Build log aggregation pipelines with parsing, filtering, and enrichment ALERTING & INCIDENT MANAGEMENT -------------------------------------------------------------------------------- • Configure intelligent alerting with Alertmanager or equivalent platforms • Design alert rules with appropriate severity levels, thresholds, and SLOs • Implement alert routing to Slack, PagerDuty, ServiceNow, email, and webhooks • Create automated runbooks and remediation workflows • Develop alert inhibition, silencing, and grouping strategies • Tune alerting to achieve signal-to-noise ratio improvements • Integrate with incident management and on-call rotation systems DISTRIBUTED TRACING -------------------------------------------------------------------------------- • Implement distributed tracing solutions using OpenTelemetry, Jaeger, or Tempo • Instrument applications for trace collection and correlation • Configure trace sampling strategies for cost and performance optimization • Build trace-based dashboards for latency analysis and dependency mapping • Integrate tracing with metrics and logs for comprehensive observability AI/ML FOR OBSERVABILITY -------------------------------------------------------------------------------- • Implement AI-powered anomaly detection for metrics and logs • Build predictive alerting using machine learning models to forecast issues before they occur • Develop intelligent alert correlation and root cause analysis systems • Integrate LLM-based tools for log analysis and troubleshooting assistance • Implement AIOps capabilities for automated incident triage and resolution • Use AI to optimize alert thresholds and reduce false positives • Build natural language query interfaces for observability data • Implement Model Context Protocol (MCP) for AI agent integration with observability platforms AUTOMATION & PLATFORM INTEGRATION -------------------------------------------------------------------------------- • Automate observability deployment using GitOps workflows (FluxCD, ArgoCD) • Integrate with CI/CD pipelines for automated testing and validation • Build self-service portals for teams to create dashboards and alerts • Develop APIs and CLIs for observability automation • Integrate with secrets management solutions (Vault, AWS Secrets Manager) • Configure LDAP/AD/SSO authentication for observability platforms • Automate compliance reporting and audit logging ENABLEMENT & COLLABORATION -------------------------------------------------------------------------------- • Onboard application teams to observability platforms • Provide guidance on instrumentation best practices • Create documentation, training materials, and self-service guides • Conduct workshops • Support development teams during incidents with observability insights • Collaborate with SRE, DevOps, and platform engineering teams PERFORMANCE & COST OPTIMIZATION -------------------------------------------------------------------------------- • Monitor and optimize observability stack resource consumption • Implement autoscaling for stateless observability components • Tune data retention, compaction, and downsampling strategies • Conduct capacity planning for metrics and log storage growth • Optimize query performance and dashboard response times • Implement cost allocation and chargeback for multi-tenant environments ======================================================================= REQUIRED QUALIFICATIONS ======================================================================= EXPERIENCE -------------------------------------------------------------------------------- • 4-6 years of experience in observability, monitoring, SRE, or platform engineering roles • 3+ years hands-on production experience with Prometheus and Grafana • 2+ years working with Kubernetes and containerized environments • Strong expertise with PromQL for metrics querying and alerting • Experience deploying and managing observability stacks at scale (1000+ nodes) • Proven track record of reducing MTTR through e_ective observability CORE TECHNICAL SKILLS -------------------------------------------------------------------------------- OBSERVABILITY PLATFORMS: • Prometheus (including Prometheus Operator) • Grafana (dashboards, alerting, plugins) • Thanos, Cortex, or Mimir for long-term storage • Alertmanager or equivalent alerting platforms LOGGING: • Loki, Elasticsearch/ELK Stack, Splunk, or CloudWatch Logs • Log collection agents (Promtail, Fluentd, FluentBit, Vector) • LogQL, Lucene, or equivalent query languages TRACING: • OpenTelemetry (OTEL Collector, instrumentation) • Jaeger, Zipkin, Tempo, or AWS X-Ray • Trace sampling and correlation strategies KUBERNETES & CONTAINERS: • Kubernetes architecture and operations (1.24+) • Custom Resource Definitions (CRDs) and Operators • ServiceMonitor, PodMonitor, PrometheusRule resources • Kubernetes metrics (kube-state-metrics, node-exporter, cAdvisor) GITOPS & AUTOMATION: • GitOps workflows (FluxCD, ArgoCD, or similar) • Infrastructure as Code (Kustomize, Helm, Terraform) • CI/CD platforms (GitHub Actions, GitLab CI, Jenkins) • Scripting (Python, Bash, Go) QUERY LANGUAGES: • PromQL (expert level required) • LogQL or Lucene • Basic SQL for data analysis CLOUD & INFRASTRUCTURE: • AWS, Azure, or GCP cloud platforms • Object storage (S3, GCS, Azure Blob) • Multi-cloud and hybrid architectures • vSphere or on-premises virtualization (nice to have) AI/ML & EMERGING TECHNOLOGIES -------------------------------------------------------------------------------- • Experience with AI/ML frameworks for observability (Prophet, TensorFlow, PyTorch) • Anomaly detection algorithms and time-series forecasting • LLM integration for log analysis and troubleshooting (GPT, Claude, etc.) • Model Context Protocol (MCP) for AI agent integration • AIOps platforms (Moogsoft, BigPanda, Datadog Watchdog, etc.) • Natural language processing for log parsing and analysis • Familiarity with vector databases for semantic search (Pinecone, Weaviate) • Experience with AI-powered root cause analysis tools • Knowledge of prompt engineering for observability use cases PREFERRED QUALIFICATIONS -------------------------------------------------------------------------------- • Kubernetes certifications (CKA, CKAD, or CKS) • Grafana certification or equivalent training • Knowledge of service mesh observability ( • Experience with SRE practices, SLIs, SLOs, and error budgets • Background in financial services or regulated industries • Familiarity with compliance requirements (SOX, PCI-DSS, etc.) • Contributions to open-source observability projects • Experience with APM tools