Site Reliability Engineer

Harvey Nash โ€” Ireland ยท Posted ~22 hours ago

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Description

Job Title: Site Reliability Engineer II Location: Dublin, Ireland(Hybrid) Duration: 12 months contract Join Our Team As an SRE, you will collaborate closely with engineering teams to build resilient, highly available systems, improve production readiness, and champion best practices in reliability engineering. You'll play a key role in incident management, root cause analysis, automation, monitoring, and performance optimization while helping shape a proactive and developer-focused operational culture. What You'll Be Doing Ensure the reliability, availability, and performance of critical production platforms.Collaborate with development teams to embed operational excellence throughout the software development lifecycle.Design, implement, and enhance monitoring, alerting, and observability solutions.Develop automation tools and scripts to eliminate manual operational tasks.Support incident response, troubleshooting, root cause analysis, and post-incident reviews.Drive continuous improvements in system resilience, scalability, and fault tolerance.Assist with capacity planning and performance optimization initiatives.Contribute to operational documentation, standards, and best practices.Support change management activities while maintaining platform stability.Work on projects that improve production readiness and operational efficiency. What We're Looking For Experience in Site Reliability Engineering, DevOps, Platform Engineering, or Systems Engineering.Strong knowledge of Linux/Unix systems administration and networking fundamentals.Hands-on experience with cloud platforms, preferably AWS.Proficiency in scripting or programming using Python, Go, Bash, or similar languages.Experience with observability and monitoring tools such as Splunk.Understanding of CI/CD pipelines, automation, containers, and DevOps practices.Knowledge of high availability, scalability, disaster recovery, and performance optimization.Experience with incident, problem, and change management following ITIL principles.Strong troubleshooting, analytical, and problem-solving skills.Excellent communication skills with a proactive and collaborative mindset.