Summary
A Site Reliability Engineer position responsible for production system stability, automation, monitoring, incident handling, and continuous improvement. The role involves cloud environments, container platforms, deployment pipelines, and collaboration with engineering teams.
Highlights
Role focused on maintaining highly available systems, automation, cloud technologies, and modern infrastructure practices. Provides opportunities to improve reliability, deployment processes, and operational excellence.
Description
Responsibility
· Maintain and support production systems, ensuring high availability, reliability, and scalability.
· Assist in implementing SRE best practices including monitoring, alerting, SLOs/SLIs, and incident response.
· Deploy, configure, and manage containerized applications using Docker and Kubernetes.
· Contribute to CI/CD pipeline development and configuration to streamline development and deployment processes.
· Automate repetitive tasks and processes using scripting languages such as Python, Bash, or Go.
· Collaborate with development, QA, and operations teams to address issues and drive improvements across the stack.
· Participate in on-call support rotation and help resolve incidents in a timely manner.
· Document procedures, configurations, and post-incident reviews to foster a learning culture.
Qualification
· Bachelor’s degree in Computer Science, Engineering, or a related discipline, or equivalent practical experience.
· At least 3-5 years of hands-on experience in SRE, DevOps, or Production Support roles.
· Working knowledge of containerization technologies (Docker, Kubernetesl,Istio, Helm) and cloud platforms (AWS, Azure, or GCP).
· Experience with monitoring and logging tools (e.g., Prometheus, Grafana, ELK, Zipkin, Jeager, Datadog, etc.).
· Familiarity with CI/CD tools such as Jenkins, GitLab CI, or similar.
· Familiar with kubernetes gitops practice and toolings, such as ArgoCD, FluxCD or Tekton;
· Scripting/programming proficiency in Python, Bash, Go, or comparable languages.
· Strong troubleshooting skills and willingness to dive into complex problems.
Preferred Qualifications:
· Experience with infrastructure-as-code (Terraform, Ansible, or similar tools)
· Exposure to microservices and distributed systems
· Awareness of ITIL, incident/change management
· Experience working in 24x7 or high-availability production environments.