Senior Site Reliability Engineer

Mayaph โ€” Philippines ยท Posted ~1 day ago

Senior Full-time

Skills

Kubernetes Terraform Istio GitLab CI/CD ArgoCD AWS Azure GCP Go Python FluxCD Java Prometheus Grafana OpenTelemetry

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary

Drive reliability engineering by designing resilient cloud infrastructure, automating operations, implementing governance, and building scalable developer platforms across distributed systems.

Highlights

Lead enterprise-scale reliability initiatives, cloud governance, automation, and resilient platform engineering across multi-cloud environments.

Description

NATURE OF WORK Lead architectural design and implementation of fault-tolerant, self-healing infrastructure across cloud and hybrid environmentsDrive organization-wide automation initiatives, eliminating manual operations through advanced IaC and CI/CD frameworksOwn technical program leadership for reliability initiatives spanning multiple teams and servicesStrategic management of OPEX and CAPEX budgets with cost optimization accountabilityDeep expertise in compliance frameworks (CIS, PCI-DSS, BSP) with ability to architect compliant solutionsEstablish and enforce cloud governance policies, account structures, and organizational standards across AWS/Azure/GCP environments REQUIRED QUALIFICATIONS 7-10+ years experience focused on SRE, DevOps, Platform Engineering, or Cloud InfrastructureExpert-level proficiency in Kubernetes (CRDs, Operators, multi-tenancy, advanced scheduling)Advanced Terraform expertise (custom providers, module design, automated testing)Deep Service Mesh knowledge (Istio traffic management, circuit breaking, rate limiting, mTLS)Proven experience building Internal Developer Platforms (IDP) with self-service workflowsAdvanced GitLab CI/CD and GitOps implementation (ArgoCD/FluxCD, multi-project pipelines)Expert-level WAF, API Gateway (Kong, Apigee, AWS APIGW), and network security implementationStrong software development skills in Go, Python, or Java with ability to review code for reliability impactExperience leading technical programs and cross-functional reliability initiativesDeep understanding of observability platforms (Dynatrace, Prometheus, OpenTelemetry) with custom integration experienceProven track record architecting microservices with high-availability and resiliency patternsExperience implementing AWS Organizations, Control Tower, Service Control Policies, and multi-account governance frameworksProficiency in cloud policy-as-code tools (AWS Config, OPA, Sentinel) and compliance automationKnowledge of cloud security standards (CIS Benchmarks, AWS Well-Architected Framework, Azure/GCP best practices)Advanced expertise in Dynatrace, Datadog, or Grafana for building enterprise observability solutionsExperience implementing SLO-based alerting, error budgets, and burn rate monitoring using Prometheus, Grafana, or commercial APM toolsProficiency in distributed tracing (Jaeger, Zipkin, OpenTelemetry) and log aggregation (ELK, Loki)Ability to design custom metrics, synthetic monitoring, and real user monitoring (RUM) strategies