Senior Site Reliability Engineer

Adecco โ€” Romania ยท Posted ~1 day ago

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Description

Senior Site Reliability Engineer (SRE) Location: Bucharest, Romania (Hybrid) Adobe is looking for an experienced Senior Site Reliability Engineer (SRE) to join the Cloud Persistence Team, part of the Cloud Foundation organization supporting Adobe Experience Manager (AEM) Cloud Service. This team is responsible for the mission-critical storage layer that powers AEM Cloud Service, ensuring the reliability, durability, and performance of large-scale distributed systems operating across thousands of cloud environments. If you have a strong background in DevOps, cloud infrastructure, Kubernetes, and production reliability, this is an opportunity to work on systems that operate at global scale. Key Responsibilities Own the reliability, availability, and performance of core storage services in production.Design and enhance monitoring, alerting, dashboards, and operational tooling using technologies such as Prometheus, Grafana, and Splunk.Define, measure, and continuously improve Service Level Indicators (SLIs) and Service Level Objectives (SLOs).Lead incident investigations, perform root cause analysis, and implement long-term reliability improvements.Partner with software engineers to design highly observable, scalable, and operationally efficient distributed systems.Automate operational processes to improve efficiency and reduce manual intervention.Contribute to the overall reliability strategy for cloud-native infrastructure running at enterprise scale. Required Experience Bachelor's or Master's degree in Computer Science or a related field, or equivalent practical experience.Minimum 6 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Production Engineering.Strong hands-on experience with Kubernetes in production environments.Experience working with at least one major cloud platform (Azure preferred).Proven expertise with observability tools, including Prometheus, Grafana, Splunk, ELK, or similar technologies.Strong understanding of distributed systems, incident management, system scalability, and operational best practices.Experience defining and managing SLIs, SLOs, and error budgets.Familiarity with Java/JVM-based applications is considered an advantage. Technical Environment KubernetesMicrosoft AzurePrometheusGrafanaSplunkELK StackCloud InfrastructureDistributed SystemsMonitoring & ObservabilityInfrastructure AutomationIncident ManagementSite Reliability Engineering (SRE) This position offers the opportunity to work on highly scalable cloud infrastructure, solve complex reliability challenges, and contribute to the platform that supports Adobe Experience Manager Cloud Service used by organizations worldwide. If you are interested in learning more about this opportunity, feel free to get in touch.