Site Reliability Engineer II

About You β€” Germany Β· Posted ~23 hours ago

Mid Visa History βœ“

Skills

Site reliability engineering Automation CI/CD Infrastructure as code Distributed systems Cloud infrastructure On-premises infrastructure Risk management Troubleshooting Cloud

πŸ”“ Log in to save this job, tailor your resume & track your apply process β€” 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join a site reliability engineering team responsible for keeping large-scale services performant, reliable, and available. You will build automation and CI/CD solutions, manage infrastructure as code, troubleshoot distributed systems across cloud and on-prem environments, and proactively reduce operational risks.

Highlights

Work on scalable, fault-tolerant and cost-effective infrastructure while improving service reliability and availability. The role offers cross-functional collaboration, automation opportunities, and meaningful impact on large-scale systems.

Description

Job DescriptionWe are seeking a talented and motivated SRE Engineer II to join our dynamic team. In this role, you will execute a range of site reliability activities, ensuring optimal service performance, reliability, and availability. You will collaborate with cross-functional engineering teams to develop scalable, fault-tolerant, and cost-effective cloud services. If you are passionate about site reliability engineering and ready to make a significant impact, we would love to hear from you! Key Responsibilities: ● Implement automation tools, frameworks, and CI/CD pipelines, promoting best practices and code reusability. ● Enhance site reliability through process automation, reducing mean time to detection, resolution, and repair. ● Identify and manage risks through regular assessments and proactive mitigation strategies. ● Develop and troubleshoot large-scale distributed systems in both on-prem and cloud environments. ● Deliver infrastructure as code to improve service availability, scalability, latency, and efficiency. ● Monitor support processing for early detection of issues and share knowledge on emerging site reliability trends. ● Analyze data to identify improvement areas and optimize system performance through scale testing. ● Take ownership of production issues within assigned domains, performing initial triage and collaborating closely with engineering teams to ensure timely resolution.