Senior Site Reliability Engineer

Csi Global Ltd โ€” United Kingdom ยท Posted ~2 days ago

Senior Contract

Skills

Azure Cloud Operations SRE Azure Monitor Log Analytics Application Insights Datadog CI/CD Terraform Bicep PowerShell Azure CLI Azure DevOps

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary

A technology services team is looking for a senior reliability engineer to manage cloud operations, improve platform stability, and automate infrastructure processes. The role involves observability, incident response, and continuous improvement of production systems.

Highlights

Work on cloud reliability, automation, and production operations for enterprise environments. The role provides opportunities to improve observability, deployment processes, and system resilience.

Description

Title: Site Reliability Engineer Location: 51 Lime street, London, UK Contract: 6 Months Primary Skill :Azure Cloud Operations Site Reliability Engineering SRE Observability Automatio nExperience in Years 7 to 12 Year sMust Have Skills :Strong experience in Azure Cloud Operations Platform Reliability and Production SupportHands-on expertise with Azure Monitor Log Analytics Application Insights Data-dog and Observability PlatformsExperience with CI/CD Pipeline Management using Azure DevOps GitHub Actions and deployment automationStrong scripting and automation skills using PowerShell Azure CLI Azure Automation and Infrastructure as Code Terraform Bicep AR M Secondary Nice To Have Skills NOT Mandato ry Experience with Kubernetes AKS container monitoring and application performance manageme ntExposure to DevSec-Ops practices security baselines and platform hardeningExperience in Azure Cost Optimization Capacity Planning and Performance TuningKnowledge of AIOps, AIdriven Operations anomaly detection correlation and predictive incident manageme nt Soft Ski llsStrong analytical and troubleshooting capabilitiesExcellent stakeholder communication and incident management skillsAbility to work under pressure during critical production incidentsStrong documentation collaboration and continuous improvement mind set 3 Qualifying Quest ionsDescribe your experience implementing and managing observability solutions such as Datadog Azure Monitor Log Analytics and Application Insights What outcomes did you achieveHow have you leveraged automation PowerShell Azure CLI Azure Automation Terraform etc to reduce operational effort and improve reliabilityCan you share an example of a major production incident you handled including the RCA process corrective actions and reliability improvements implemented after ward S killsMandatory Skills : Azure DevOps, Azure Infra Services, Azure Log Analytics, Azure Mo nitorGood to Have Skills : Azure AI, Azure App Se rvice