Summary
✨ AI‑Generated
A reliability engineering role focused on cloud-native infrastructure, monitoring, automation, incident response, and improving production system stability.
Highlights
Remote full-time role focused on automation, observability, cloud-native operations, and maintaining highly available services.
Description
Join HCLTech Germany as a Site Reliability Engineer
At HCLTech, we're powering digital transformation for some of the world's largest enterprises.
We are looking for a talented Site Reliability Engineer (SRE) who is passionate about reliability, automation, observability, and cloud-native operations.
If you enjoy solving complex production challenges, automating repetitive tasks, building scalable monitoring solutions, and ensuring business-critical services remain available 24/7, we'd love to hear from you.
Location: Germany (Remote)
Employment Type: Full-Time
Candidates must be based in Germany or willing to relocate to Germany.
What You'll Do:
As a Site Reliability Engineer, you will play a key role in ensuring the stability, availability, performance, and security of mission-critical platforms.
Your responsibilities will include:
Designing and improving monitoring and observability solutions using Grafana, Prometheus, and ELK StackManaging and optimizing Elasticsearch/OpenSearch clustersSupporting and operating containerized workloads in Kubernetes environmentsDeveloping automation solutions using Python, Go, or BashManaging CI/CD pipelines with tools such as Jenkins, Helm, and ArgoCDLeading troubleshooting efforts across infrastructure, platform, and application layersParticipating in 24x7 incident response and on-call rotationsPerforming root cause analysis and driving continuous service improvementsCreating and maintaining operational documentation and runbooksCollaborating with global engineering, cloud, and operations teams
What We're Looking For:
Must-Have Skills:
Grafana & ELK Stack (mandatory)Kubernetes administration and container platformsLinux system administrationElasticsearch and/or OpenSearchPrometheus monitoring ecosystemCI/CD tooling (Jenkins, Helm, ArgoCD)Infrastructure automation and scriptingPython, Bash, or GoUnderstanding of networking concepts and REST APIsStrong troubleshooting and incident management experience
Preferred Qualifications:
AWS and/or Azure experienceKubernetes certifications (CKA/CKAD)Elastic Certified EngineerLPIC Level 2 or equivalent Linux certificationsExperience working in enterprise-scale production environments
Important Requirement:
This role supports environments subject to German security regulations.
Applicants must be eligible for and willing to undergo the German Ü2 Security Clearance process.
EU + NATO citizenship is required.
Why HCLTech?
Be part of a global technology leader
225,000+ employees across 60 countriesWork with enterprise-scale cloud and infrastructure platformsRemote working model within GermanyInternational and diverse teamsContinuous learning and certification opportunitiesLong-term career growth in cloud, platform engineering, and reliability engineering
At HCLTech, we believe technology is powered by people.
We foster an inclusive workplace where innovation, collaboration, and personal growth drive success.
Apply now and help us build reliable, scalable, and secure platforms that power global businesses every day.