Automation and Reliability Engineer

Argyll Scott — Thailand · Posted ~2 hours ago

Mid Contract Remote

Skills

automation engineering site reliability engineering production monitoring incident response troubleshooting workflow reliability automation reliability engineering

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Take ownership of the reliability of a business-critical automation platform in a fully remote engineering role. You will monitor production workflows, respond to failed or degraded processes, troubleshoot incidents, and continuously improve platform resilience. The initial 12-month contract may be extended or converted to a permanent position, and the role involves collaboration with engineering and operations teams across time zones.

Highlights

Fully remote 12-month contract with potential extension or permanent conversion, focused on hands-on reliability engineering, automation, production troubleshooting, and collaboration across global time zones.

Description

Argyll Scott is proud to be partnered with a leading global technology business as they're looking to bring onboard an Automation Engineer & Reliability Engineer. This role will start as an initial 12 Month contract, with plans to be extended or be made permanent. Our client, a global organisation with a large-scale security operations function, is looking for an Automation & Reliability Engineer to keep its security automation platform healthy, resilient and continuously improving. This is a hands-on engineering role, not a traditional support position. You will own the day-to-day reliability of a business-critical automation environment that underpins security operations worldwide, providing coverage for the East while working closely with engineering and operational teams across time zones. What you will do Take ownership of platform health during your operational hours, monitoring for failed, stalled, delayed or degraded workflows and responding quicklyTroubleshoot production issues across Cortex XSOAR, integrations, APIs and supporting platformsCarry out root-cause analysis on recurring failures and implement lasting fixes rather than quick patchesBuild and enhance playbooks, integrations and scripts where engineering improvements are neededStrengthen monitoring, alerting, health checks and observability, and analyse performance metrics to spot reliability risks earlySet and contribute to reliability standards covering failover, retries, timeout handling and resilienceWork with SOC analysts, incident responders and engineers to prioritise and resolve operational issues, and take part in post-incident reviewsSupport upgrades, maintenance, deployments and infrastructure changes using GitLab for source control and change managementUse AI-assisted engineering tools to speed up troubleshooting and improve efficiencyWrite and maintain runbooks and documentation, and cross-train colleagues so expertise is sharedJoin a global on-call rotation supporting critical services What you will bring 2-4 years' experience in production engineering, automation engineering, SRE, DevOps, security engineering or a similar fieldStrong Python scripting and automation skills (essential)Hands-on experience with Cortex XSOAR or a comparable automation platformExperience supporting business-critical production systemsA solid track record troubleshooting APIs, integrations, authentication mechanisms and system dependenciesFamiliarity with monitoring, alerting, logging and observability toolsWorking knowledge of Git and GitLab-based workflowsAn understanding of enterprise security operations, incident response and security automationSharp analytical skills, with the ability to investigate complex issues independentlyExcellent documentation and communication skills, and comfort coordinating across global teamsWillingness to take part in on-call coverage Argyll Scott Asia is acting as an Employment Business in relation to this vacancy.