Summary
✨ AI‑Generated
Take ownership of the reliability of a business-critical automation platform in a fully remote engineering role. You will monitor production workflows, respond to failed or degraded processes, troubleshoot incidents, and continuously improve platform resilience. The initial 12-month contract may be extended or converted to a permanent position, and the role involves collaboration with engineering and operations teams across time zones.
Highlights
Fully remote 12-month contract with potential extension or permanent conversion, focused on hands-on reliability engineering, automation, production troubleshooting, and collaboration across global time zones.
Description
Argyll Scott is proud to be partnered with a leading global technology business as they're looking to bring onboard an Automation Engineer & Reliability Engineer.
This role will start as an initial 12 Month contract, with plans to be extended or be made permanent.
Our client, a global organisation with a large-scale security operations function, is looking for an Automation & Reliability Engineer to keep its security automation platform healthy, resilient and continuously improving.
This is a hands-on engineering role, not a traditional support position.
You will own the day-to-day reliability of a business-critical automation environment that underpins security operations worldwide, providing coverage for the East while working closely with engineering and operational teams across time zones.
What you will do
Take ownership of platform health during your operational hours, monitoring for failed, stalled, delayed or degraded workflows and responding quicklyTroubleshoot production issues across Cortex XSOAR, integrations, APIs and supporting platformsCarry out root-cause analysis on recurring failures and implement lasting fixes rather than quick patchesBuild and enhance playbooks, integrations and scripts where engineering improvements are neededStrengthen monitoring, alerting, health checks and observability, and analyse performance metrics to spot reliability risks earlySet and contribute to reliability standards covering failover, retries, timeout handling and resilienceWork with SOC analysts, incident responders and engineers to prioritise and resolve operational issues, and take part in post-incident reviewsSupport upgrades, maintenance, deployments and infrastructure changes using GitLab for source control and change managementUse AI-assisted engineering tools to speed up troubleshooting and improve efficiencyWrite and maintain runbooks and documentation, and cross-train colleagues so expertise is sharedJoin a global on-call rotation supporting critical services
What you will bring
2-4 years' experience in production engineering, automation engineering, SRE, DevOps, security engineering or a similar fieldStrong Python scripting and automation skills (essential)Hands-on experience with Cortex XSOAR or a comparable automation platformExperience supporting business-critical production systemsA solid track record troubleshooting APIs, integrations, authentication mechanisms and system dependenciesFamiliarity with monitoring, alerting, logging and observability toolsWorking knowledge of Git and GitLab-based workflowsAn understanding of enterprise security operations, incident response and security automationSharp analytical skills, with the ability to investigate complex issues independentlyExcellent documentation and communication skills, and comfort coordinating across global teamsWillingness to take part in on-call coverage
Argyll Scott Asia is acting as an Employment Business in relation to this vacancy.