Site Reliability Engineer - Observability & Monitoring Platform

Hcltech — United Kingdom · Posted ~8 hours ago

Full-time Visa History ✓

Skills

Site reliability engineering Observability Monitoring Linux CI/CD Cloud technologies Grafana Cloud Containers

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An SRE role focused on operating monitoring platforms, improving system reliability, automating operations, and supporting scalable infrastructure.

Highlights

Work with modern observability technologies, improve platform reliability, and collaborate in a global engineering environment.

Description

About the Role We are looking for a talented Site Reliability Engineer (SRE) to join a global Platform Services & Engineering team responsible for delivering enterprise-wide monitoring, observability, and notification services. This role combines production support, platform administration, and reliability engineering. You will help operate and evolve a modern observability ecosystem while partnering with engineering, infrastructure, and support teams to improve platform reliability, scalability, and performance. This is an excellent opportunity for someone passionate about monitoring, telemetry, automation, and operational excellence, working with state-of-the-art technologies in a highly collaborative global environment. Technologies & Platforms Grafana LGTM StackOpen-source observability and telemetry toolsEnterprise incident and notification platformsLinux environmentsCI/CD ecosystemsCloud and container technologiesKey Responsibilities Drive adoption of observability and monitoring solutions across the organization.Administer and support enterprise monitoring, telemetry, and notification platforms.Maintain highly available, resilient, and scalable production environments.Monitor platform health and proactively respond to alerts before they become incidents.Perform incident management, problem management, root cause analysis (RCA), and post-incident reviews.Manage production releases, deployments, and change activities with minimal risk.Troubleshoot platform issues and support internal users with technical guidance.Collaborate with engineering teams to improve operational efficiency and system reliability.Design and implement monitoring solutions that enhance visibility across infrastructure and applications.Contribute to architectural discussions and help shape the future observability strategy.Promote observability best practices and data-driven operational excellence across teams.Participate in Agile ceremonies including sprint planning, reviews, and retrospectives.Leverage AI-powered tools to improve operational efficiency and automation.Mentor junior team members and contribute to knowledge sharing within the team.Work closely with globally distributed teams and stakeholders across multiple regions. Required Skills & Experience Mandatory Minimum 2 years of experience administering Grafana or other enterprise observability/monitoring platforms.At least 2 years of hands-on Linux administration and troubleshooting experience.Experience with Python and/or Ansible.Strong production support background including:Incident ManagementProblem ManagementChange ManagementRelease ManagementOn-call SupportUser Support & Alert ResponseStrong analytical, troubleshooting, and problem-solving skills.Excellent communication and stakeholder management skills.Basic understanding of cloud technologies and services.Familiarity with CI/CD tools such as GitLab, Jenkins, Ansible, Nexus, or similar.Ability to work effectively in a collaborative team environment. Preferred Qualifications Understanding of OpenTelemetry standards and observability concepts.Experience with container and orchestration technologies such as Kubernetes, Docker, EKS, or similar platforms.Experience supporting medium-to-large-scale production environments.Knowledge of AI-assisted operational tools and productivity platforms.Familiarity with ITIL principles and service management practices.Working knowledge of SQL and relational databases (MySQL, MSSQL, Sybase, etc.).Experience with collaboration and project management tools such as Jira and Confluence.Exposure to infrastructure technologies including:Messaging Middleware (ActiveMQ, Solace, EMS, Tibco)Web ServersLoad BalancersDirectory ServicesExperience working in globally distributed teams. What You'll Gain Opportunity to work on large-scale, enterprise observability platforms.Exposure to cutting-edge monitoring and reliability engineering practices.Collaboration with global engineering and operations teams.A culture focused on innovation, automation, operational excellence, and continuous improvement.Career growth within a fast-evolving SRE and observability landscape.