Summary
✨ AI‑Generated
A reliability engineering position focused on maintaining scalable cloud systems, improving observability, automating operations, and managing production reliability.
Highlights
Full-time reliability engineering role focused on cloud platforms, operational excellence, automation, and system stability.
Description
Role: Site Reliability Engineer
Work location: Hybrid, onsite at least 3 days per week in Midtown, Atlanta, GA.
Duration: Full time (direct hire)
Role Summary
Lead the reliability, scalability, security, and operational excellence of customer-facing platforms across Azure, GCP, and Kubernetes environments.
Drive production stability through automation, observability, incident management, and continuous improvement initiatives.
Key Responsibilities
Lead platform reliability, availability, and performance initiatives.Design and support cloud infrastructure in Azure and GCP.Manage and optimize Kubernetes environments and containerized applications.Implement observability and monitoring using Splunk, AppDynamics, and cloud-native tools.Support Cloudflare, Zscaler, SQL Server, RabbitMQ, and enterprise networking components.Lead major incident response, RCA, and problem management activities.Develop automation and self-healing solutions to improve operational efficiency.Collaborate with Engineering, Product, Security, and Infrastructure teams to enhance customer experience and platform stability.Serve as a technical escalation point for critical production and customer issues.
Required Skills
5+ years of experience in SRE, DevOps, Cloud Operations, or Infrastructure Engineering.Strong expertise in Azure, GCP, Kubernetes, Cloudflare, Splunk, AppDynamics, SQL Server, RabbitMQ, and Zscaler.Solid networking knowledge (DNS, TCP/IP, HTTP/S, CDN, WAF, Load Balancing, SSL/TLS, Firewalls, VPNs).Experience with automation and scripting (Python, PowerShell, Bash, Terraform).Strong customer-facing communication and stakeholder management skills.
Core Principles
Automation First Eliminate manual effort through automation and self-healing systems.Observability-Driven Operations Leverage logs, metrics, traces, and analytics to proactively identify and resolve issues.AI-Powered Reliability Utilize AI and operational intelligence to accelerate detection, diagnosis, and remediation.Customer-Centric Mindset Prioritize customer experience, stability, and business outcomes.Operational Excellence Continuously improve reliability, scalability, and security.
Success Measures
Service Availability & UptimeSLA/SLO ComplianceMTTR ReductionIncident ReductionAutomation AdoptionCustomer Satisfaction (CSAT)Platform Performance & Stability Improvements
Best Regards,
David Roy | Accounts Manager US Staffing | Charter Global Inc.
| https://www.charterglobal.com
LinkedIn