Summary
✨ AI‑Generated
A growing engineering organization is looking for a Site Reliability Engineer to own day-to-day reliability and operations for critical automation and business services. You will operate Kubernetes workloads and supporting infrastructure, build monitoring and alerting, create runbooks, automate recurring operational tasks, and support infrastructure provisioning. The position emphasizes resilient systems, clear operational ownership, and reducing manual operational effort.
Highlights
High-impact reliability role with ownership of production operations, strong opportunities to automate repetitive work, improve observability, establish operational standards, and increase system resilience while reducing operational dependencies.
Description
About the role
Own the day-to-day reliability and operations of the Business Automation platform and KYB services, ensuring stable, resilient systems and reducing dependency on individual team members.
This role strengthens operational ownership across the team while enabling engineering leadership to focus on architecture, design reviews, and quality assurance.
You'll be working on:
Operate and maintain the Automation platform running reliably day to day, including Kubernetes workloads, ingress and load balancing, secrets management, and CI/CD pipelines.Build and tune monitoring and alerting so failures are detected automatically before someone has to notice them manually, and ensure every alert has a clear owner and response path.Write and maintain runbooks and operational documentation, starting with services that currently have little or no documentation.Automate recurring operational work, reducing manual work and repetitive load on the team.Support infrastructure provisioning and configuration through Infrastructure as Code, reviewing changes with the Engineering Manager.Assist root cause analysis after incidents by gathering and reviewing evidence from logs, metrics, and system behaviour.Support backup, restore and disaster recovery exercises, and make sure recovery procedures actually work when needed.Gradually take independent ownership of defined operational areas as you build deeper familiarity with the platform.
Qualifications:
Bachelor’s degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent practical experience.1–2 years of experience in systems administration, DevOps, infrastructure, or IT operations.
Strong internship or project experience may also be considered.Hands-on knowledge of Linux administration, including permissions, networking, storage, performance troubleshooting, and command-line operations.Good understanding of networking fundamentals, including TCP/IP, DNS, routing, load balancing, and firewalls, with the ability to troubleshoot connectivity and latency issues.Scripting skills in Python, Go, or Bash to automate operational and infrastructure tasks.Familiarity with virtualization, Docker, and Kubernetes; experience operating containerized workloads in production is preferred.Exposure to Infrastructure as Code, particularly Terraform or Ansible, and version control using Git.Familiarity with CI/CD pipelines such as GitHub Actions or GitLab CI, including automated testing, builds, and deployment workflows.Exposure to cloud infrastructure, ideally Alibaba Cloud or Azure; experience with AWS or GCP is also relevant.Familiarity with observability and monitoring tools such as Grafana, OpenSearch, or the ELK stack.Strong troubleshooting mindset with a methodical approach to diagnosing production and infrastructure issues.Clear written and verbal communication, including the ability to document systems, share knowledge, collaborate with developers and product teams, and communicate calmly during incidents.