Description
Site Reliability Engineer (SRE)
Blue Bell, PA | 6-months Contract
Hybrid role based in Blue Bell, PA (3 days onsite/week).
Candidates must currently live within commuting distance and participate in a rotating on-call/major incident support schedule.
Please send resume to: aghosh@copiastaffing.com
Position Summary
We are seeking an experienced Site Reliability Engineer to improve the availability, resiliency, performance, recoverability, and operational support of business-critical technology services across hybrid on-premises and cloud environments.
This is a hands-on, infrastructure-focused SRE position supporting a broad technology landscape including Azure, Windows/Linux, networking, SQL Server, virtualization, storage, VDI, and vendor-managed services.
The ideal candidate combines strong infrastructure engineering fundamentals with observability, automation, incident/problem management, and disaster recovery experience.
Responsibilities
Improve reliability, availability, resiliency, performance, and recoverability of critical technology services.Implement monitoring, alerting, dashboards, service health indicators, and proactive detection across infrastructure, applications, networks, databases, cloud, storage, and VDI.Provide technical leadership during major P1/P2 incidents and drive root cause analysis and corrective actions.Identify single points of failure, capacity constraints, fragile dependencies, and recurring operational issues.Automate repetitive operational work using PowerShell, Python, REST APIs, Terraform, Ansible, or similar tools.Define and test disaster recovery, failover, high availability, RTO/RPO, backup, replication, and recovery capabilities.Partner with DBA teams on Microsoft SQL Server reliability, clustering, Availability Groups, replication, backups, and recovery.Partner with Network Engineering across WAN, SD-WAN, routing, switching, VPN, DNS, firewalls, BGP, wireless, and remote access.Support Azure, hybrid-cloud services, Azure Virtual Desktop, VMware, Citrix, and enterprise storage environments.Develop runbooks, dependency maps, recovery procedures, technical documentation, and service ownership models.Participate in architecture, change, production-readiness, and operational acceptance reviews.Participate in a rotating on-call and major incident support schedule.
Requirements
5+ years of experience in Site Reliability Engineering, Infrastructure Engineering, Systems Engineering, DevOps, Production Engineering, or enterprise IT operations.Experience supporting business-critical services across hybrid on-premises and cloud environments.Strong hands-on background with monitoring, observability, alerting, performance analysis, and capacity planning.Automation/scripting experience with PowerShell, Python, REST APIs, Terraform, Ansible, or comparable technologies.Experience with major incident management, RCA/problem management, and corrective-action tracking.Strong understanding of disaster recovery, high availability, failover, RTO/RPO, and resiliency.Working knowledge across several infrastructure domains such as Windows/Linux, Azure, networking, virtualization, storage, VDI, and SQL Server.Strong knowledge of networking concepts including TCP/IP, DNS, DHCP, routing, switching, VLANs, VPN, firewalls, WAN, and hybrid-cloud connectivity.Experience creating technical documentation, operational runbooks, dependency maps, and recovery procedures.Knowledge of ITIL/ITSM practices including Incident, Problem, Change, Configuration, and Knowledge Management.Experience with technologies such as Azure, Microsoft SQL Server, VMware, Cisco, Meraki, Juniper, Citrix, Zscaler, Pure Storage, and Ivanti is beneficial.