Summary
✨ AI‑Generated
A senior Site Reliability Engineer is sought to own the reliability, scalability, and performance of critical production platforms. The role combines Kubernetes operations, Linux administration, observability, monitoring, automation, and hands-on incident response. You will build dashboards and alerting, improve metrics and centralized logging, troubleshoot infrastructure, and lead root-cause investigations for complex production issues.
Highlights
Senior-level ownership of critical production reliability, with strong exposure to Kubernetes, observability, Linux engineering, automation, and complex incident response.
Description
We are looking for a Senior Site Reliability Engineer to ensure the reliability, scalability, and performance of critical production platforms.
The role combines SRE, Kubernetes operations, observability, Linux engineering, and automation, with responsibility for both building reliability solutions and supporting production operations.
Key Responsibilities
Own SRE, monitoring, and Kubernetes platform reliability across critical systems.Build and maintain observability solutions using Prometheus, VictoriaMetrics, Elasticsearch, Grafana, Kibana, and ITRS Geneos.Develop dashboards, alerting rules, metrics pipelines, and centralized logging solutions.Operate and troubleshoot Kubernetes environments, including workloads, scaling, capacity, and cluster health.Administer and troubleshoot RHEL Linux environments, including patching, upgrades, configuration, and performance tuning.Lead technical investigation and root cause analysis during complex production incidents.Automate operational activities using Python/Bash, Ansible, and CI/CD tooling.Collaborate with application and infrastructure teams to improve availability, resilience, performance, and security.
Requirements
8–10 years of IT experience, ideally in SRE, platform engineering, production support, or financial services.Strong hands-on experience with Kubernetes in production environments.Strong knowledge of Prometheus, Grafana, Elasticsearch/Kibana and modern observability practices.Experience building Prometheus/metrics and Logstash/logging pipelines.Strong Linux (RHEL) administration and troubleshooting experience.Proficiency in automation using Python/Bash, Ansible, and CI/CD tools.Good understanding of SRE principles, HA, scalability, incident management, DR/BCP, and distributed systems.Experience supporting mission-critical production systems and participating in on-call rotations.Strong analytical, troubleshooting, and communication skills in English; Chinese is advantageous.Experience with GPU infrastructure for AI/ML workloads is a plus.