Software Engineer

Russell Tobin & Associates Llc โ€” United Kingdom ยท Posted ~1 day ago

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Description

Software Engineer Location: London, UK Contract Duration: 6 months Start/End Dates: 07/09/2026 - 06/03/2027 Working Hours: 40 hours per week Experience: 8+ years Interview Process: 2 rounds About the Role We are looking for an experienced Software Engineer to provide short-term cover within a Core Engineering team supporting a large-scale User Generated Content (UGC) application in production. The role focuses on building and operating the automation required to keep a live application healthy across multiple surfaces, including VR headsets and mobile/PC through cloud streaming. You will work alongside a runtime and KTLO (Keep The Lights On) team, with a strong focus on reducing manual operational and on-call work through automation, maintaining cloud rendering and deployment pipelines, and keeping CI/CD systems functional as upstream dependencies change. Key Responsibilities Build and maintain automation that keeps a large-scale application healthy in production.Support production health across release pipelines, build health, crash triage and incident detection.Maintain and improve AI-assisted code repair systems capable of creating and landing fixes autonomously.Develop tooling that can automatically identify broken builds and pinpoint the changes responsible for failures.Build systems that can recommend or automatically execute fixes.Monitor production quality and performance metrics and respond to regressions and outages.Reduce manual operational and on-call workload through automation, with a target of significantly reducing recurring operational work.Complete infrastructure and dependency migrations while ensuring downstream CI/CD pipelines continue to function.Maintain backend services supporting the cloud-rendered application.Monitor weekly deployments of the cloud-rendering system and ensure they operate correctly.Ensure compatibility across different platforms and user groups following deployments.Technical Focus This role is primarily focused on backend services and production systems rather than client-side development. The team maintains a remote-rendered backend stack that serves frames to user devices. The focus is on maintaining application stability, performance and reliability, preventing crashes and ensuring deployments work correctly across VR and mobile environments. A major part of the role will involve building AI-driven systems that automate maintenance, monitor performance metrics and proactively help maintain the stability of a live product. Essential Requirements 8+ years of professional software engineering experience, or equivalent experience.Proven experience building and operating CI/CD, build, release and cloud deployment pipelines at scale.Strong experience operating cloud services and server-side production systems, including reliability, capacity and latency considerations.Experience building or operating AI-assisted developer tooling or AI agents that generate or repair code.Experience building autonomous or self-healing systems capable of monitoring performance and determining appropriate actions.Experience developing tooling that detects broken builds and traces failures back to their root-cause change.Experience with production monitoring, crash triage and incident response for large-scale applications.Demonstrable experience reducing operational and on-call workload through automation.Experience completing infrastructure or dependency migrations without breaking downstream CI/CD systems.Strong backend and production engineering experience.Top Non-Negotiable Skills 1. AI-Driven Autonomous / Self-Healing Systems Hands-on experience building systems that use AI to monitor performance metrics and autonomously determine the actions required to maintain system stability and performance. 2. Backend Service Deployment & Maintenance Strong experience deploying and maintaining backend services, ideally involving distributed, cloud-based or remote-rendered systems. 3. Production Automation & Reliability Experience automating production monitoring, maintenance, incident response, build/release processes or other operational workloads. Desirable Experience Experience with cloud gaming, application streaming or remote rendering.Experience with asset delivery or CDN pipelines at scale.Experience monitoring capacity, latency or session orchestration for streamed workloads.Experience operating live-service or large-scale production applications (Live Ops).Familiarity with large monorepo build systems and dependency management.Experience designing self-healing or auto-remediation systems.Day-to-Day Responsibilities The role will involve: Building AI-driven automation to reduce manual maintenance.Monitoring application performance and quality metrics.Maintaining backend services rather than developing new client-side features.Monitoring weekly cloud-rendering deployments.Diagnosing deployment and compatibility issues.Ensuring the backend stack remains stable and performant.Identifying and preventing application crashes.Maintaining CI/CD and deployment infrastructure.Automating recurring operational activities.Investigating production issues and taking ownership of their resolution.Working Environment The successful candidate will need to be comfortable working with a high degree of autonomy and ownership. You will be expected to be self-sufficient, communicate effectively when you encounter roadblocks and take ownership of resolving problems rather than simply identifying them. This is particularly important because the role involves supporting a live product in a maintenance and transition phase, where the focus is on maintaining stability and reliability rather than developing new client-side features. Key Challenges Working with Ambiguity The role requires someone who can work independently, assess problems and determine the appropriate solution without needing constant direction. End-to-End Ownership You will be expected to take responsibility for resolving issues rather than simply escalating them. Production Maintenance The application continues to have a significant daily user base and therefore requires ongoing monitoring, maintenance and reliability work. AI-Driven Automation A core challenge is replacing manual maintenance activities with AI-powered automation and self-healing systems that can monitor application behaviour and proactively take action. Deployment & Compatibility Weekly cloud-rendering deployments need to be monitored carefully to ensure that changes do not break compatibility between different platforms and user groups, such as VR and mobile users. Ideal Candidate Profile The ideal candidate will be a senior/staff-level hands-on Software Engineer, Production Engineer, SRE, Backend Engineer or similar with strong experience across: Backend production systemsCloud infrastructureCI/CD and deployment pipelinesProduction reliabilityAutomationAI-assisted developer toolingAutonomous/self-healing systemsIncident responseBuild and release engineeringCandidates should be comfortable operating complex production systems and solving problems independently. Candidates Less Likely to Be Suitable This role is not primarily focused on client-side development or traditional full-stack development. Candidates whose experience is predominantly focused on frontend/client-side development, or who lack significant backend production and automation experience, are unlikely to be a strong match. The team does not require candidates to have prior experience with internal company-specific tooling. Candidate Value Proposition This role offers the opportunity to work on a challenging AI-driven automation problem within a large-scale live production environment. The key focus is moving beyond manual maintenance by developing systems that can monitor production performance and autonomously determine the actions required to maintain stability. You will work on a live, high-traffic product across VR and mobile and solve complex production problems involving reliability, automation, backend systems and cloud rendering.