SRE / DevOps Engineer - Chaos Engineering

Ktekresourcing — Canada · Posted ~5 hours ago

Senior Hybrid

Skills

Java Python Microservices APIs Kubernetes Docker Cloud platforms AWS FIS AWS CDK Chaos engineering Resilience testing Observability High availability Disaster recovery Automation SLA/SLO/SLI MTTR RTO/RPO AWS AppDynamics Prometheus Grafana

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A hybrid SRE/DevOps opportunity for an experienced engineer specializing in chaos engineering and system resilience. You will design and automate failure scenarios across cloud-native infrastructure, Kubernetes, APIs, microservices, and distributed systems, identify reliability gaps, validate recovery mechanisms, and improve observability and operational resilience.

Highlights

Work on advanced reliability and resilience engineering across cloud-native, Kubernetes, API, microservices, and distributed environments. The role offers cross-functional collaboration, automation, observability, and ownership of critical reliability metrics and recovery practices.

Description

Role: SRE / DevOps Engineer (Chaos Engineering) Location: Toronto ,ON -Hybrid Experience: 5+ years Key Responsibilities Design and execute resiliency and chaos testing scenarios across cloud, Kubernetes, APIs, microservices, and distributed environments.AWS FIS knowledge, setup and Run ExperimentsAWS CDK knowledge.Identify resilience gaps, SPOFs, and operational risks and drive remediation.Validate High Availability (HA), Disaster Recovery (DR), failover, auto-healing, and recovery processes.Automate testing and reliability validation using scripting and cloud-native tools.Monitor application and infrastructure behavior using observability platforms such as AppDynamics, Prometheus, and Grafana.Collaborate with DevOps, SRE, Infrastructure, and Application teams to improve system reliability.Define and track reliability metrics including SLA, SLO, SLI, MTTR, RTO, and RPO. Required Skills Strong experience in Java/Python, Microservices, APIs, Kubernetes, Docker, and Cloud Platforms (Azure/GCP/AWS).Hands-on experience with monitoring and observability tools such as Dynatrace metrics and observability, AppDynamics, Prometheus, and Grafana.Knowledge of Reliability Engineering, Chaos Testing, Incident Analysis, and Resiliency Validation.Failure-as-a-Service platforms to achieve resiliency in infrastructure failures, network and application failures. Chaos Testing mechanism and tools such as AWS-FIS, Lambda testing, On-premise , OpenShift testing.Monitoring tools test/scenario capture and report creation through standardized templates and hypothesis formation.Experience with automation, CI/CD, and cloud-native architectures.Excellent troubleshooting, analytical, and communication skills.