Summary
✨ AI‑Generated
Join a reliability-focused engineering team responsible for keeping critical production systems scalable and highly available. You will troubleshoot complex infrastructure and application incidents, perform root cause analysis, build automation and self-healing capabilities, strengthen monitoring and observability, and partner with engineering teams to improve reliability and performance.
Highlights
Work on critical, highly available systems in a 24/7 engineering environment, with strong opportunities to automate operational work, improve observability, solve complex production problems, and drive long-term reliability improvements.
Description
Core Profile
As a Operations Engineer, you'll help keep Maya's critical platforms reliable, scalable, and always available.
You'll leverage automation, observability, and engineering best practices to minimize operational toil, resolve complex production issues, and improve the overall reliability and performance of our systems.
What You'll Do
Monitor and support highly available production systems in a 24/7 environment.Respond to and troubleshoot complex infrastructure and application incidents.Drive root cause analysis and implement long-term solutions to prevent recurrence.Build automation and self-healing solutions to improve operational efficiency.Design and enhance monitoring, alerting, logging, and observability capabilities.Partner with Engineering teams to improve system reliability, scalability, and performance.Create and maintain runbooks, documentation, and operational best practices.
What We're Looking For
4+ years of experience in Site Reliability Engineering, DevOps, Production Engineering, or Technical Operations.Experience supporting cloud environments (AWS, Azure, or GCP).Strong scripting skills (Python, Bash, Go, or similar).Hands-on experience with monitoring and observability tools such as Datadog, Splunk, Dynatrace, Prometheus, or Grafana.Familiarity with CI/CD, Infrastructure as Code, and automation practices.Strong troubleshooting, problem-solving, and incident management skills.Experience with Kubernetes, Docker, and distributed systems is a plus.
Success in This Role
Improve platform uptime and reliability.Reduce incident recurrence and recovery time.Increase automation and reduce manual operational effort.Enhance system observability and operational excellence.Drive continuous improvements that create a better experience for both customers and engineers.