Chaos Engineer (Resilience Automation Specialist)

Ai Talent Au — Australia · Posted ~3 hours ago

Senior Visa Sponsored

Skills

Chaos engineering Resilience testing Fault injection Automation SRE Distributed systems Chaos Engineering Distributed Systems

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A resilience engineering role focused on designing failure experiments, validating distributed systems, and improving reliability through automated testing and fault simulation techniques.

Highlights

Specialized reliability engineering role with visa support, focusing on improving resilience of mission-critical digital systems.

Description

We are partnering with an enterprise client to strengthen platform resilience, validate fault tolerance, and ensure zero-downtime reliability for mission-critical digital systems. We are seeking a skilled Chaos Engineer (Resilience Automation Specialist) to represent our organisation and take technical ownership of fault injection experiments, automated resilience testing, and blast-radius verification. In this role, you will bridge Site Reliability Engineering (SRE) and systems architecture within our client’s environment. You will intentionally simulate real-world failure scenarios—such as network latency, node dropouts, database failovers, and region outages—to uncover hidden failure modes and prove that distributed systems self-heal under duress. 🌏 Visa & Sponsorship OptionsAs the employer of record, we provide full visa and migration support for qualified engineering talent deployed to our clients: 482 On-Hire Sponsorship Transfers: Fully supported for qualified candidates currently in Australia on an existing 482 visa looking to transfer sponsorship to work with our clients.New 482 Visa Sponsorship: Available for qualified candidates meeting the commercial experience and technical requirements.Temporary & Working Visa Holders: Open to all working visa holders seeking a direct pathway to employer sponsorship. Core ResponsibilitiesChaos Experimentation: Design, execute, and automate controlled chaos experiments across distributed microservices and containerized environments.Fault Injection Tooling: Deploy and manage enterprise chaos engineering platforms (e.g., Chaos Mesh, Gremlin, LitmusChaos, AWS Fault Injection Simulator).GameDays & Failure Drills: Organize and facilitate systematic GameDay scenarios to test system recovery, automated failover mechanisms, and operational runbooks.Resilience in CI/CD: Embed automated resilience assertions and steady-state validation directly into continuous deployment pipelines.Observability & Blast Radius Control: Monitor real-time telemetry (metrics, traces, logs) during experiments using tools like Prometheus, Grafana, Datadog, or OpenTelemetry, ensuring safe rollback triggers.Architecture Hardening: Collaborate with software engineering and cloud infrastructure teams to translate experiment findings into actionable architectural improvements (circuit breakers, retry policies, auto-healing).Selection CriteriaSRE & Resilience Mastery: Commercial experience in Site Reliability Engineering (SRE) or platform reliability with a strong focus on distributed systems resilience.Chaos Engineering Tooling: Hands-on experience configuring fault-injection tools such as Chaos Mesh, Gremlin, LitmusChaos, or AWS FIS.Container & Cloud Ecosystems: Deep understanding of Kubernetes architecture, container networking, service meshes, and cloud infrastructure (AWS/Azure).Scripting & Automation: Strong programming or scripting skills in Python, Go, or Bash to build custom failure-injection scenarios and verification logic.Location Requirements: Currently residing in Australia with valid work rights or eligibility for 482 visa sponsorship/transfer. Preferred Qualifications (Nice to Have)Practical experience implementing distributed tracing (OpenTelemetry, Jaeger).Certified Kubernetes Administrator (CKA) or relevant Cloud certifications.Experience in high-availability, low-latency domains (Banking, Fintech, Telecommunications, E-Commerce).