Senior Site Reliability Engineer

Xrexinc β€” Taiwan Β· Posted ~3 weeks ago

Senior

Skills

AWS site reliability engineering infrastructure operations 24/7 system monitoring incident management system performance system stability Zabbix ELK

πŸ”“ Log in to save this job, tailor your resume & track your apply process β€” 7 days free, no card needed.

Log in to add to target list

Summary

Join an experienced infrastructure team as a senior SRE in an English-speaking, fast-moving environment. You will operate AWS infrastructure, maintain 24/7 stability and performance, monitor systems with established observability tools, respond to incidents, and tackle a broad range of technical challenges while continuously expanding your skill set.

Highlights

Senior SRE opportunity in an English-speaking, fast-paced environment with global exposure. The role offers hands-on ownership of AWS infrastructure, 24/7 reliability, monitoring, incident response, and broad technical problem-solving.

Description

About Want to build a worldwide brand from Taiwan, and to communicate our brand story to millions of users worldwide? Want to be based in Taiwan but work in a silicon-valley-like environment, and to build world-class brand and products? Want to participate in the global fintech and blockchain movement, and work at an English-speaking workplace? Come change the world with us! Join this fast-growing startup founded by software veterans and funded by top VCs, Skype co-founders, and the Taiwanese government (NDF)! We’re hiring for an experienced Senior SRE Engineer. The exact mix of other skills does not matter, so long as your tool chest includes a mix of abilities. Be willing to attack anything that comes your way, learn on the fly and get things done. Come talk to us if you want to push your skillset in a dynamic fast-paced environment. Responsibilities Maintain and operate AWS infrastructure to ensure 24/7 stability and performance Monitor systems (Zabbix, ELK), handle incidents, and develop custom scripts as needed Analyze and resolve platform issues; optimize architecture and performance Support high availability, backup, and troubleshooting for applications and databases Collaborate with backend, product, and infra teams on system design and deployment Document operations and incidents; support internal IT needsParticipate in on-call rotation Requirements 5+ years of Linux system administration experience; 24/7 ops experience is a plus Strong hands-on experience with AWS services, including EC2, Lambda, Aurora, ElastiCache (Redis), CloudWatch, CloudFront, EKS, IAM, and more Proficient in scripting (Bash, Python, Golang) and container orchestration (Kubernetes) Experience with infrastructure as code (Terraform, Helm, Kustomize) Familiar with CI/CD tools (Jenkins, GitHub Actions, Argo Workflows/CD) Experience with Airflow and DAG development Knowledge of scalable system design and related tools (MongoDB, Kafka, load balancers, message queues) Solid understanding of information security best practices Strong problem-solving, communication, and teamwork skills; able to work independently under pressure Location: Taipei (check it out on Google Maps!) About XREX Regarding our culture