Operations Engineer

Mayaph — Philippines · Posted ~1 day ago

Senior

Skills

Site Reliability Engineering DevOps Infrastructure operations Production incident response Root cause analysis Automation Monitoring and observability Logging and alerting System reliability Scalability Performance optimization SRE Observability Monitoring Logging

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join a reliability-focused engineering team responsible for keeping critical production systems scalable and highly available. You will troubleshoot complex infrastructure and application incidents, perform root cause analysis, build automation and self-healing capabilities, strengthen monitoring and observability, and partner with engineering teams to improve reliability and performance.

Highlights

Work on critical, highly available systems in a 24/7 engineering environment, with strong opportunities to automate operational work, improve observability, solve complex production problems, and drive long-term reliability improvements.

Description

Core Profile As a Operations Engineer, you'll help keep Maya's critical platforms reliable, scalable, and always available. You'll leverage automation, observability, and engineering best practices to minimize operational toil, resolve complex production issues, and improve the overall reliability and performance of our systems. What You'll Do Monitor and support highly available production systems in a 24/7 environment.Respond to and troubleshoot complex infrastructure and application incidents.Drive root cause analysis and implement long-term solutions to prevent recurrence.Build automation and self-healing solutions to improve operational efficiency.Design and enhance monitoring, alerting, logging, and observability capabilities.Partner with Engineering teams to improve system reliability, scalability, and performance.Create and maintain runbooks, documentation, and operational best practices. What We're Looking For 4+ years of experience in Site Reliability Engineering, DevOps, Production Engineering, or Technical Operations.Experience supporting cloud environments (AWS, Azure, or GCP).Strong scripting skills (Python, Bash, Go, or similar).Hands-on experience with monitoring and observability tools such as Datadog, Splunk, Dynatrace, Prometheus, or Grafana.Familiarity with CI/CD, Infrastructure as Code, and automation practices.Strong troubleshooting, problem-solving, and incident management skills.Experience with Kubernetes, Docker, and distributed systems is a plus. Success in This Role Improve platform uptime and reliability.Reduce incident recurrence and recovery time.Increase automation and reduce manual operational effort.Enhance system observability and operational excellence.Drive continuous improvements that create a better experience for both customers and engineers.