Summary
✨ AI‑Generated
A Site Reliability Engineer is sought to operate and enhance enterprise-scale private cloud infrastructure supporting highly available virtual desktop environments. The role focuses on improving reliability, scalability, resilience, and business continuity through SRE practices, automation, infrastructure as code, disaster recovery, proactive monitoring, and observability. Responsibilities include managing clustered virtualization infrastructure, storage, networking, capacity, patching, backups, and operational workflows while defining service reliability objectives and responding proactively to incidents.
Highlights
Opportunity to operate and improve enterprise-scale private cloud infrastructure, with a strong focus on reliability, scalability, high availability, automation, disaster recovery, and modern observability practices.
Description
Key Responsibilities
Operate enterprise-scale private cloud infrastructure built on Microsoft Hyper-V.
Optimize, and support highly available VDI environments on Hyper-V.
Improve platform reliability, availability, scalability, and resiliency by applying SRE
principles and engineering best practices.
Disaster recovery, backup, patch management, and business continuity
strategies.
Define and maintain Service Level Indicators (SLIs), Service Level Objectives
(SLOs), and operational metrics for critical infrastructure services.
Automate infrastructure provisioning, configuration management, and operational
workflows using PowerShell and Infrastructure as Code (IaC) principles wherever
applicable.
Manage Hyper-V Failover Clusters, host lifecycle, storage, networking, and
capacity to ensure high availability and business continuity.
Develop proactive monitoring, alerting, logging, and observability capabilities to
detect and prevent service degradation.
Lead incident response for infrastructure-related outages, perform root cause
analysis (RCA), and implement preventive actions through post-incident reviews.
Perform capacity planning, performance tuning, and resource optimization across
Hyper-V clusters and VDI platforms.
Support infrastructure migration initiatives including P2V, V2V, workload
modernization, and private cloud transformations.
Collaborate closely with Security, Networking, Platform Engineering, and
Application teams to improve platform reliability and operational efficiency.
Develop and maintain technical documentation, architecture diagrams,
operational runbooks, automation scripts, and standard operating procedures.
Mentor junior engineers and promote SRE culture, automation, and operational
best practices across the team.
Mandatory Skills for Hyper-V SRE Production Support
Microsoft Hyper-V Administration (Deployment, Troubleshooting, Optimization)
Hyper-V Failover Clustering & High Availability
Windows Server 2016/2019/2022 Administration & OS patching
Storage Spaces Direct (S2D), CSV, SAN/NAS & Storage
Site Reliability Engineering (SRE) Principles, SLI/SLO, Reliability
PowerShell Scripting & Automation
Disaster Recovery, Backup, Hyper-V Replica & Business Continuity
SCVMM (System Center Virtual Machine Manager)