Cloud Operations Service Reliability Engineer

Aoshearman โ€” United Kingdom ยท Posted ~2 days ago

Mid Full-time

Skills

Microsoft Azure Service Reliability Engineering Infrastructure as Code Bicep Azure DevOps GitHub Monitoring Observability Incident Management Automation Azure Ansible Elastic ITIL

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary

Join a global technology team focused on improving cloud reliability and operational excellence. Build automation, enhance observability, strengthen monitoring, and help deliver resilient cloud services using modern infrastructure and DevOps practices.

Highlights

Work on global cloud services, drive automation and operational resilience, collaborate across international teams, and contribute to continuous service improvement using modern cloud engineering practices.

Description

What you will do The Service Reliability Engineer is accountable for improving the reliability, observability and operational resilience of cloud-hosted services. The role focuses on monitoring, early issue identification, cloud engineering and automation, using tools and practices such as Bicep, Azure DevOps, GitHub and Ansible to support consistent, repeatable and well-governed service operation. Monitoring, observability and alerting across cloud infrastructure, platform services and supported application environments;Cloud engineering and automation, including Infrastructure as Code, deployment pipelines, configuration management and standards-led delivery; andPromoting service resiliency through proactive issue identification, operational insight, automation and continuous improvement.The role involves: Supporting service reliability, observability and cloud engineering across the following areas: Monitoring, observability and alerting for cloud-hosted services, including infrastructure health, service availability, performance signals and operational events - Essential;Azure public cloud engineering, including IaaS, PaaS, networking, identity, RBAC and platform diagnostics - Essential;Infrastructure as Code and automation using Bicep, Azure DevOps pipelines and GitHub-based source control and collaboration - Essential;Configuration management and standards automation using Ansible or equivalent tooling - Preferred;Experience of using or implementing monitoring solutions using Elastic - Preferred;Operational reporting, issue trend analysis and the development of actionable dashboards to support service improvement - Preferred;Experience of working across a broad range of systems, technologies and internal support teams - Preferred.Ensuring that monitoring and operational insight are effectively designed, implemented and understood so that services can be supported, improved and made more resilient.Providing subject matter expertise in cloud operations, observability, automation and reliability engineering practices.Working globally across cloud-hosted services and platform capabilities, independent of location.Support the firm's environmental goals and initiatives. Monitoring, Reliability and Cloud Engineering Works with internal technology teams to improve end-to-end observability for supported services, including: Monitoring coverage for infrastructure, platform services and application components;Actionable alerting that supports early identification of degradation, failure or operational risk;Dashboards and reporting that help teams understand service health, trends and recurring issues;Cloud engineering practices that use Bicep, Azure DevOps, GitHub and Ansible to deliver consistent and repeatable change; andOperational standards that improve service resilience and reduce manual support effort.Maintains appropriate documentation, including monitoring standards, known issues, operational patterns, troubleshooting guidance and support handbooks. Service Delivery Identify, diagnose and support resolution of incidents and problems by interpreting monitoring signals, operational telemetry and service behaviour.Work with service-owning teams to improve the quality, relevance and routing of alerts so that operational issues can be detected and acted upon quickly.Contribute to root cause analysis, problem management and continuous improvement activity by identifying recurring patterns, gaps in observability and opportunities for automation. Build and Implementation Provide specialist guidance to teams adopting cloud engineering patterns, Infrastructure as Code, deployment pipelines and automated configuration management.Support implementation of monitoring and automation standards across new and existing services.Ensure that operational documentation, handover materials and support guidance are created and are suitable for BAU operation. Risk Management Identify operational, reliability and supportability risks arising from gaps in monitoring, alerting, automation or cloud platform standards.Refer to domain experts for guidance on specialised areas such as architecture, security, networking, database platforms and application design.Participate in recovery, resilience and operational readiness activities to help prove that services can be supported effectively. Quality, Methods & Tools Strive for improvements to processes by promoting standardised patterns, automated controls, repeatable engineering practices and effective use of industry best practice.Advocate for the use of source control, pipeline-based delivery, Infrastructure as Code and configuration management to improve quality, auditability and operational reliability. What you will have Business Competencies Strong analytical and problem-solving skills, with a logical approach to issue identification, diagnosis and service improvement.Technically curious, with an enthusiasm for understanding a broad set of systems, technologies and operational domains.Ability to interpret monitoring data, identify patterns and translate operational insight into meaningful improvement activity.Ability to make sound decisions under pressure and support effective incident response.Strong commitment to service reliability, operational resilience and excellent customer service.Commercial acumen, including an understanding of IT service costs, cloud consumption and how technology adds value to the business.Ability to promote technical standards, automation and reliability practices using clear, business-friendly language.Personal credibility; highly self-motivated self-starter who will undertake all activities to the highest professional standards.Excellent communication skills, both orally and written.Ability to operate within a wider team where there may be ambiguity and conflicting priorities.Ability to build effective working relationships across a diverse set of internal teams and influence the adoption of monitoring, automation and cloud engineering standards.Experience of working in a global environment across international locations with an appreciation of multiple cultures. Knowledge Practical knowledge of SRE principles, observability, incident response, problem management and operational resilience.Detailed practical knowledge of Microsoft Azure infrastructure and platform services, including monitoring, diagnostics, RBAC, networking and automation.Knowledge of Infrastructure as Code, source control and pipeline-based delivery using tools such as Bicep, Azure DevOps and GitHub.Knowledge of configuration management and automation tooling such as Ansible.Expected to develop a broad understanding of the technologies, systems and business working practices used by A&O Shearman. Experience Minimum 4-5 years' IT experience with at least 2 years' experience in a cloud operations, infrastructure, platform engineering, SRE or 3rd line support role.Experience of monitoring, alerting, incident investigation and operational issue identification in a complex technology environment.Experience of using or implementing monitoring using Elastic is desirable.Experience using or supporting automation and delivery tooling such as Bicep, Azure DevOps, GitHub and Ansible.Experience working with diverse internal teams to improve service supportability, resilience and operational standards.Experience of working in an ITIL environment. Qualifications Ideally the candidate should have the following, or equivalent: Minimum "A" level standard education or equivalent;Accreditation in relevant technologies - preferred; andITIL Foundation - preferred. Ideal candidate profile The ideal candidate will be technically curious and motivated by understanding how different systems, platforms and operational processes fit together. They will enjoy working across a broad technology estate, using monitoring data and engineering insight to identify issues early, improve service resilience and guide teams towards standardised, automated and supportable ways of working. NO AGENCIES PLEASE - A&O Shearman does not accept unsolicited CVs. For further information, please see our UK Recruitment Agency Policy and our commitment to direct sourcing here.