Site Reliability Engineer

Spait Infotech2024 — United Kingdom · Posted ~2 hours ago

Mid

Skills

Cloud infrastructure CI/CD System monitoring Logging and alerting Incident management Root cause analysis Production troubleshooting Automation Monitoring Logging Alerting

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A Site Reliability Engineer role focused on maintaining highly available and scalable production systems. You will automate operational processes, manage cloud infrastructure, build CI/CD pipelines, implement monitoring and alerting, troubleshoot incidents, perform root cause analysis, and maintain operational documentation. A relevant bachelor's degree is expected, while candidates ranging from freshers with relevant projects or certifications to experienced professionals may be considered.

Highlights

Opportunity to work on scalable, highly reliable systems with cloud infrastructure, automation, CI/CD, monitoring, and incident management. Both fresh graduates with relevant projects or certifications and experienced candidates are considered.

Description

Key Responsibilities Monitor system performance, availability, reliability, and infrastructure health.Design, implement, and maintain highly scalable and reliable systems.Automate deployment, monitoring, and operational processes.Manage cloud infrastructure and production environments.Troubleshoot and resolve application, infrastructure, and performance issues.Develop and maintain CI/CD pipelines for reliable software delivery.Implement monitoring, logging, alerting, and incident management solutions.Perform root cause analysis and implement preventive measures.Work closely with development, DevOps, and IT teams to improve system reliability.Participate in on-call support and respond to production incidents.Maintain technical documentation, runbooks, and operational procedures.Key Requirements Bachelor's degree in Computer Science, IT, Engineering, or a related field.Freshers with relevant projects or certifications and experienced candidates can apply.Basic to strong understanding of Linux/Unix systems and networking.Knowledge of cloud platforms such as AWS, Azure, or Google Cloud.Familiarity with Docker, Kubernetes, Git, and CI/CD tools.Knowledge of scripting/programming languages such as Python, Bash, or PowerShell.Understanding of monitoring and logging tools such as Prometheus, Grafana, ELK, or similar tools.Good understanding of DevOps, automation, system monitoring, and incident management.Strong troubleshooting and problem-solving skills.Good communication and teamwork skills.