Summary
Join a technology team responsible for building and operating resilient production systems. You will automate infrastructure, improve observability, support CI/CD, and maintain reliable Windows and Linux environments while resolving incidents and enhancing platform stability.
Highlights
Join a newly created reliability engineering role focused on high-availability platforms, automation, observability, and collaboration with engineering teams in a business-critical environment.
Description
Title: Site Reliability Engineer (SRE)
Type: Permanent / onsite 4 days a week
Location: Dublin City Centre
An exciting opportunity for a Site Reliability Engineer (SRE) to join a high-performing technology team responsible for ensuring the availability, performance, and reliability of business-critical platforms.
This role combines infrastructure, automation, and software engineering, working across both Windows and Linux environments to build resilient, scalable, and highly automated production systems.
THIS IS A BRAND NEW LINUX SRE TYPE ROLE
THIS IS A HIGH AVAILABLITY SITE SO THE TEAM IS ONSITE 4 DAYS A WEEK
Key Responsibilities:
Manage and support production platforms, ensuring high availability, performance, and operational stability.Investigate incidents, perform root cause analysis, and implement permanent solutions to prevent recurrence.Drive automation across operational processes, replacing manual tasks with scalable, code-driven solutions.Develop and maintain monitoring, alerting, and observability platforms to improve system health and reduce operational noise.Build and maintain CI/CD pipelines and deployment automation to support reliable software releases.Collaborate with software development and infrastructure teams to improve platform reliability and operational efficiency.Develop internal tools and contribute to continuous improvement initiatives across engineering teams.
Skills & Experience:
Experience in a Site Reliability Engineering, DevOps, Platform Engineering, or Production Engineering role.Strong scripting and automation skills using technologies such as Python, PowerShell, or Bash.Experience administering both Windows and Linux environments.Knowledge of CI/CD pipelines and deployment automation tools.Experience with monitoring and observability platforms such as Grafana, Prometheus, Elasticsearch, or similar.Strong troubleshooting, problem-solving, and stakeholder communication skills.Comfortable working in a shift-based environment with occasional on-call responsibilities.