Site Reliability Engineer (SRE) – Microsoft Hyper-V

Cubestech Ltd — United Kingdom · Posted ~22 hours ago

Senior

Skills

Microsoft Hyper-V Site Reliability Engineering (SRE) Private cloud infrastructure VDI High availability Disaster recovery Backup and business continuity PowerShell Infrastructure as Code (IaC) Hyper-V Failover Clustering Infrastructure monitoring Alerting Logging Observability Storage Networking Capacity management Incident management

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A Site Reliability Engineer is sought to operate and enhance enterprise-scale private cloud infrastructure supporting highly available virtual desktop environments. The role focuses on improving reliability, scalability, resilience, and business continuity through SRE practices, automation, infrastructure as code, disaster recovery, proactive monitoring, and observability. Responsibilities include managing clustered virtualization infrastructure, storage, networking, capacity, patching, backups, and operational workflows while defining service reliability objectives and responding proactively to incidents.

Highlights

Opportunity to operate and improve enterprise-scale private cloud infrastructure, with a strong focus on reliability, scalability, high availability, automation, disaster recovery, and modern observability practices.

Description

Key Responsibilities  Operate enterprise-scale private cloud infrastructure built on Microsoft Hyper-V.  Optimize, and support highly available VDI environments on Hyper-V.  Improve platform reliability, availability, scalability, and resiliency by applying SRE principles and engineering best practices.  Disaster recovery, backup, patch management, and business continuity strategies.  Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational metrics for critical infrastructure services.  Automate infrastructure provisioning, configuration management, and operational workflows using PowerShell and Infrastructure as Code (IaC) principles wherever applicable.  Manage Hyper-V Failover Clusters, host lifecycle, storage, networking, and capacity to ensure high availability and business continuity.  Develop proactive monitoring, alerting, logging, and observability capabilities to detect and prevent service degradation.  Lead incident response for infrastructure-related outages, perform root cause analysis (RCA), and implement preventive actions through post-incident reviews.  Perform capacity planning, performance tuning, and resource optimization across Hyper-V clusters and VDI platforms.  Support infrastructure migration initiatives including P2V, V2V, workload modernization, and private cloud transformations.  Collaborate closely with Security, Networking, Platform Engineering, and Application teams to improve platform reliability and operational efficiency.  Develop and maintain technical documentation, architecture diagrams, operational runbooks, automation scripts, and standard operating procedures.  Mentor junior engineers and promote SRE culture, automation, and operational best practices across the team. Mandatory Skills for Hyper-V SRE Production Support Microsoft Hyper-V Administration (Deployment, Troubleshooting, Optimization) Hyper-V Failover Clustering & High Availability Windows Server 2016/2019/2022 Administration & OS patching Storage Spaces Direct (S2D), CSV, SAN/NAS & Storage Site Reliability Engineering (SRE) Principles, SLI/SLO, Reliability PowerShell Scripting & Automation Disaster Recovery, Backup, Hyper-V Replica & Business Continuity SCVMM (System Center Virtual Machine Manager)