Skills
Site reliability engineering
AWS
Observability
Python
Go
Bash
Linux/Unix administration
Networking
High availability
Scalability
Disaster recovery
CI/CD
Containers
Orchestration
Troubleshooting
Root cause analysis
Capacity planning
Performance tuning
IT service management
Splunk
Dynatrace
Linux
Unix
Summary
A technology team is seeking a Site Reliability Engineer to improve the availability, scalability, performance, and resilience of critical applications and infrastructure. You will collaborate with development and operations teams, automate operational work, monitor production systems, investigate incidents, and contribute to capacity planning and disaster recovery. Strong AWS, Linux, scripting, observability, DevOps, and troubleshooting skills are required. This is a hybrid 12-month contract with possible extension.
Highlights
A 12-month contract with potential extension, focused on reliability engineering, automation, observability, incident response, and continuous operational improvement.
Description
12-month contract (view to extend)Hybrid 2 days on site (usually Mon & Tue)Day rate Dublin
We are looking for a Site Reliability Engineer II (SRE) to help ensure the reliability, availability, scalability, and performance of critical applications and infrastructure.
You will collaborate with development and operations teams to improve system resilience, automate operational processes, monitor production environments, and support incident response.
Preferred Tools
AWSSplunk / DynatraceITIL Framework
Key Responsibilities
Ensure production systems are reliable, scalable, and highly available.Partner with development teams to build resilient, fault-tolerant applications.Automate operational tasks using scripting and infrastructure tools.Monitor applications and infrastructure using observability solutions.Investigate, troubleshoot, and resolve production issues, conducting root cause analysis and post-incident reviews.Support capacity planning, performance optimisation, and disaster recovery initiatives.Maintain operational documentation and promote best practices.Contribute to change, incident, problem, and risk management processes.Continuously improve system reliability through proactive monitoring and automation.
Required Skills
Experience with observability tools for monitoring metrics, logs, and traces.Proficiency in scripting/programming (e.g., Python, Go, Bash).Strong Linux/Unix system administration and networking knowledge.Hands-on experience with cloud platforms, particularly AWS.Understanding of high availability, scalability, and disaster recovery.Familiarity with DevOps practices, CI/CD pipelines, containers, and orchestration.Strong troubleshooting and root cause analysis skills.Experience with capacity planning and performance tuning.Knowledge of IT service management concepts, including incident, problem, and change management.
Apply Now