Site Reliability Engineer II

Oliver James โ€” Ireland ยท Posted ~1 day ago

Mid Contract Hybrid

Skills

Site reliability engineering AWS Observability Python Go Bash Linux/Unix administration Networking High availability Scalability Disaster recovery CI/CD Containers Orchestration Troubleshooting Root cause analysis Capacity planning Performance tuning IT service management Splunk Dynatrace Linux Unix

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary

A technology team is seeking a Site Reliability Engineer to improve the availability, scalability, performance, and resilience of critical applications and infrastructure. You will collaborate with development and operations teams, automate operational work, monitor production systems, investigate incidents, and contribute to capacity planning and disaster recovery. Strong AWS, Linux, scripting, observability, DevOps, and troubleshooting skills are required. This is a hybrid 12-month contract with possible extension.

Highlights

A 12-month contract with potential extension, focused on reliability engineering, automation, observability, incident response, and continuous operational improvement.

Description

12-month contract (view to extend)Hybrid 2 days on site (usually Mon & Tue)Day rate Dublin We are looking for a Site Reliability Engineer II (SRE) to help ensure the reliability, availability, scalability, and performance of critical applications and infrastructure. You will collaborate with development and operations teams to improve system resilience, automate operational processes, monitor production environments, and support incident response. Preferred Tools AWSSplunk / DynatraceITIL Framework Key Responsibilities Ensure production systems are reliable, scalable, and highly available.Partner with development teams to build resilient, fault-tolerant applications.Automate operational tasks using scripting and infrastructure tools.Monitor applications and infrastructure using observability solutions.Investigate, troubleshoot, and resolve production issues, conducting root cause analysis and post-incident reviews.Support capacity planning, performance optimisation, and disaster recovery initiatives.Maintain operational documentation and promote best practices.Contribute to change, incident, problem, and risk management processes.Continuously improve system reliability through proactive monitoring and automation. Required Skills Experience with observability tools for monitoring metrics, logs, and traces.Proficiency in scripting/programming (e.g., Python, Go, Bash).Strong Linux/Unix system administration and networking knowledge.Hands-on experience with cloud platforms, particularly AWS.Understanding of high availability, scalability, and disaster recovery.Familiarity with DevOps practices, CI/CD pipelines, containers, and orchestration.Strong troubleshooting and root cause analysis skills.Experience with capacity planning and performance tuning.Knowledge of IT service management concepts, including incident, problem, and change management. Apply Now