Description
This position is listed on behalf of a partner company, who manages all applications and next steps.
Our partner is looking for a Site Reliability Engineer Technical Lead based in United States.
This is a fully remote, hands-on technical leadership opportunity focused on the reliability and operational excellence of mission-critical production systems.
You’ll apply advanced SRE principles to improve the scalability, resilience, performance, and availability of enterprise environments.
The role combines cloud infrastructure, automation, observability, incident response, and systems engineering to eliminate operational toil and strengthen reliability.
You’ll design innovative solutions, build automation and internal tooling, and develop self-healing capabilities that improve engineering efficiency.
The position requires deep expertise across cloud platforms, microservices, Kubernetes, monitoring, logging, and distributed systems.
You’ll also influence technical direction and collaborate with both technical and non-technical stakeholders, including senior leadership.
It’s an ideal environment for an experienced SRE professional who wants significant technical ownership and the opportunity to shape reliable, scalable systems.
Accountabilities
Drive the operational excellence, reliability, scalability, and performance of critical production systems.Apply Site Reliability Engineering principles to enterprise-level systems and continuously identify opportunities to improve resilience and availability.Lead technical aspects of incident response, troubleshooting complex production issues and developing sustainable solutions to prevent recurrence.Design and implement automation that eliminates repetitive manual work, reduces operational toil, and improves engineering efficiency.Develop and maintain internal tools and automated workflows that support scalable, reliable, and self-healing infrastructure.Work across cloud environments such as AWS, GCP, and Azure, as well as microservices and containerized platforms including Kubernetes.Build, maintain, and optimize observability capabilities covering monitoring, alerting, logging, and distributed tracing.Use platforms and tools such as Dynatrace, Splunk, ELK Stack, or comparable technologies to analyze system behavior and identify potential issues proactively.Analyze operational metrics and system performance data to guide performance tuning, capacity planning, and reliability improvements.Identify opportunities to strengthen system architecture, automation, monitoring, and operational processes.Collaborate effectively with engineering teams and senior stakeholders to communicate technical risks, recommendations, and solutions.Provide technical leadership and influence reliability practices without relying primarily on project-management responsibilities.
Requirements
8+ years of IT experience, including significant experience in a senior-level SRE, infrastructure engineering, systems engineering, or closely related role.Deep understanding and hands-on application of SRE principles within enterprise-scale environments.Strong experience with cloud platforms such as AWS, Google Cloud Platform (GCP), and/or Microsoft Azure.Solid expertise with microservices architectures and container orchestration technologies, particularly Kubernetes.Demonstrated ability to write production-quality code, particularly with Python or comparable programming languages.Proven experience designing and implementing automation to reduce manual operational work and create scalable, resilient, and self-healing systems.Experience building and maintaining internal tooling that improves infrastructure and operational processes.Strong expertise in observability, including monitoring, alerting, logging, and distributed tracing.Experience with technologies such as Dynatrace, Splunk, ELK Stack, or similar observability platforms.Strong analytical skills and the ability to interpret system metrics to proactively identify reliability and performance issues.Experience using operational data to support performance optimization and capacity planning.Excellent troubleshooting and problem-solving abilities in complex production environments.Exceptional communication, leadership, and interpersonal skills, with the ability to influence technical and non-technical stakeholders.Ability to collaborate effectively with senior leadership and communicate complex technical concepts clearly.Comfortable working independently in a fully remote environment while taking ownership of mission-critical technical challenges.Work authorization: U.S.
Citizens, Green Card holders, EAD holders, and candidates with H-1B transfer eligibility are encouraged to apply.
New H-1B sponsorship is not available for this position.
Benefits
Fully remote position within the United States.Full-time, direct W-2 employment.Competitive annual salary range of $100,000–$150,000, depending on qualifications and experience.Opportunity to work on mission-critical enterprise systems and advanced cloud technologies.Significant technical ownership and influence over reliability, automation, and operational engineering practices.Opportunity to collaborate with experienced technical teams and senior stakeholders.Career growth potential within a technology-focused environment.Hands-on exposure to cloud platforms, Kubernetes, automation, observability, and modern SRE practices.
How Jobgether Works
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements.
Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company.
The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer.
This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR).
You may exercise your rights (access, rectification, erasure, objection) at any time.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information.
These tools assist our recruitment team but do not replace human judgment.
Final hiring decisions are ultimately made by humans.
If you would like more information about how your data is processed, please contact us.