Site Reliability Engineer

Nipa Agency — Thailand · Posted ~3 days ago

Mid

Skills

Linux Networking Troubleshooting Python3 Shell Scripting Root Cause Analysis Ansible Terraform Jenkins Docker Kubernetes Freshdesk Clickup

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary

We are seeking a Site Reliability Engineer to join our technical operations team. In this role, you will collaborate with engineering teams to ensure our applications are built for scale, resilience, and optimal performance. You will take ownership of system design, capacity planning, performance tuning, and the continuous automation of our testing and release pipelines. Responsibilities include resolving technical production issues, performing root cause analysis, managing internal portals, and acting as a prime escalation point for cloud infrastructure. Candidates should have a degree in a related technical field, 1 to 5 years of experience in systems engineering or SRE, strong Linux and networking knowledge, and scripting skills in Python, Shell, Ansible, or Terraform. Hands-on experience with cloud platforms and containerization technologies is considered a strong advantage.

Highlights

Collaborative environment bridging development and operations to build highly scalable and resilient applications. The role encourages continuous learning, innovation in automation, and supports engineers in acquiring new professional certifications. You will act as a key escalation point, shaping CI/CD standards and system architecture.

Description

Responsibilities Collaborate with Product Owner and Tech Lead for planning technical activities. Collaborate with Development and Operations teams to improve CI/CD standards. Partner with development teams to ensure applications are designed with scale, resilience, and performance in mind. Participate in system design consulting, platform management production system through application reviews, testing, capacity planning, and performance tuning. Continuously seek out innovative ways to speed up and automate all aspects of testing, building, and releasing software project versions and updates through our infrastructure. Perform root cause analysis of production errors and resolve technical issues. Support team members throughout the cloud operations team. Act as a prime escalation point and Tier 2 cloud infrastructure engineer ticket (Freshdesk, Clickup). Propose automation ideas and solutions to reduce workload. Control internal portal change. Maintenance of SLA uptime. Study and take tests to acquire relevant certifications for SRE/DevOps. Other ad-hoc tasks as assigned by Manager. Qualifications Bachelor’s degree in Computer Science, Computer Engineering or related technical major, or commensurate experience. 1-5 years in system engineer, systems development, SRE (Site Reliability Engineering), or Resilience Engineering. 1-5 years of server systems debug experience; debugging and root causing complex application platforms. Exposure to cloud or virtualization, Docker, Kubernetes (about AWS, Azure, Google Cloud, OpenStack, OpenShift, VMware vSphere) is plus. Experience with API, OS, Networking troubleshooting (API Request, CPU/Mem profiling, TCP/IP, DNS, routing, switching, firewalls, LAN/WAN or related). Interests in contributing towards increasing durability, security, availability and scalability of systems through exploration, diagnosis and remediation. Systematic thinking - ability to diagnose interactions between discrete components of system and drive product improvements. Familiarity with Linux (Ubuntu, Debian CentOS, RedHat) and Linux namespaces. Scripting in Ansible, Shell, Python3, Terraform, Jenkin to automate and operate system. Ability to handle and resolve critical customer issues/escalations, including on-call job rotations. Hands-on, self-starter and problem-solving attitude. Individual thinker and good team player.