Summary
✨ AI‑Generated
Join an SRE team responsible for highly available, scalable production infrastructure. You will improve reliability and performance, operate distributed systems, define SLIs and SLOs, build monitoring and alerting, automate operational tasks, and participate in production incident management.
Highlights
Work on highly available, large-scale cloud infrastructure, target 99.99%+ service availability, improve observability and automation, and solve challenging reliability problems alongside product engineering teams.
Description
We are looking for a Site Reliability Engineer (SRE) to help maintain and improve the reliability, availability, scalability, and performance of Tencent Cloud's production services and infrastructure.
The ideal candidate will have strong experience in Linux, Kubernetes, cloud infrastructure, distributed systems, monitoring, automation, and production incident management.
You will work closely with product engineering teams to build reliable, scalable, and highly available cloud services.
Key Responsibilities
Maintain the reliability, availability, scalability, and performance of Tencent Cloud production services and infrastructure.Engineer and improve cloud services to achieve 99.99%+ availability, depending on service SLA requirements.Operate and optimize highly distributed production systems supporting large-scale customers.Define and monitor SLIs, SLOs, error budgets, and service health indicators.Build and maintain monitoring, logging, tracing, and automated alerting systems.Detect, troubleshoot, and resolve production incidents across compute, network, storage, and database environments.Lead incident response, root-cause analysis (RCA), and post-mortem activities.Develop self-healing and automated remediation mechanisms to improve system reliability.Automate repetitive operational tasks and reduce engineering toil.Perform capacity planning, performance optimization, and resilience testing.Design and implement failover and disaster recovery (DR) mechanisms.Participate in production/on-call rotations for critical cloud services.Collaborate closely with Tencent Cloud product engineering teams to identify and resolve reliability issues and improve service quality.
Requirements
Around 3+ years of experience in SRE, DevOps, Cloud Engineering, Infrastructure Engineering, Platform Engineering, or a related field.Strong experience with Linux systems and production environments.Hands-on experience with Kubernetes and container technologies.Proficiency in at least one programming or scripting language such as Python, Go, or C/C++.Understanding of distributed systems and cloud infrastructure.Experience with observability, including monitoring, logging, tracing, and alerting.Experience with automation and CI/CD pipelines.Strong troubleshooting and production incident management skills.Knowledge of system performance, scalability, capacity planning, and high availability.Understanding of disaster recovery, failover, and system resilience.Relevant IT qualifications and industry certifications are an advantage.
Preferred Skills
Experience working with large-scale cloud or distributed systems.Experience with AWS, Azure, GCP, Tencent Cloud, or other cloud platforms.Experience with tools such as Prometheus, Grafana, ELK/EFK, Jaeger, or similar observability platforms.Experience with Terraform, Ansible, or other infrastructure automation tools.Familiarity with SLA/SLO/SLI, error budgets, and reliability engineering practices.Experience participating in 24/7 on-call or production support rotations.