Site Reliability / Cloud Platform Engineer

Global Recruitment And Consultancy Opc — Philippines · Posted ~6 hours ago

Senior

Skills

Site Reliability Engineering Cloud Infrastructure Kubernetes Linux Networking Root Cause Analysis Automation Scripting Monitoring Incident Management

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An organization is seeking an experienced reliability and cloud platform engineer to troubleshoot complex production systems, perform root cause analysis and improve the resilience and scalability of enterprise platforms. You will build automation, monitoring tools and operational solutions while partnering with engineering and infrastructure teams to prevent recurring incidents.

Highlights

Opportunity to solve complex production and infrastructure challenges while improving reliability, scalability and operational efficiency through automation, monitoring and long-term engineering improvements.

Description

We are looking for an experienced Site Reliability / Cloud Platform Engineer to support and improve the reliability, performance, and scalability of enterprise technology platforms. The role involves troubleshooting complex technical issues, performing root cause analysis, improving system reliability, and developing automation and monitoring solutions. You will work closely with engineering, infrastructure, and support teams to resolve critical issues and implement long-term improvements. Key Responsibilities Investigate and resolve complex production, infrastructure, and platform issues through detailed root cause analysis.Troubleshoot issues across Kubernetes, Linux, networking, cloud infrastructure, configurations, and underlying systems.Identify recurring incidents and system weaknesses and implement solutions to prevent future occurrences.Develop automation, scripts, tools, dashboards, and monitoring solutions to improve operational efficiency and supportability.Support and optimize Kubernetes and cloud-native environments, including service mesh technologies.Participate in critical incident response, outage investigations, and post-incident reviews.Analyze system metrics, logs, and performance data to identify reliability and availability risks.Provide technical recommendations for improving platform resilience, scalability, observability, and performance. Required Qualifications 4–8+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Systems Engineering, or a related field.Experience supporting enterprise-scale platforms and production environments.Strong hands-on experience with Kubernetes and Linux.Strong understanding of networking and infrastructure.Experience with cloud-native technologies, automation, and observability.Experience or exposure to AI / Machine Learning technologies. Required Technical Skills: KubernetesLinuxNetworkingService Mesh / IstioPKI / Public Key InfrastructureAI / Machine LearningGoogle Cloud ComputePrometheusGrafanaLokiSplunk