Site Reliability Engineer

Sixgroup — Poland · Posted ~2 hours ago

Mid Full-time

Skills

Site reliability engineering Cloud-native platforms Automation Platform engineering System reliability CI/CD Infrastructure Security collaboration Cloud-native AI

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join an SRE and platform engineering team responsible for designing, automating, and operating resilient cloud-native platforms. You will improve reliability and delivery pipelines through automation while collaborating with development, infrastructure, and security teams and exploring AI-driven operational improvements.

Highlights

Work on scalable, resilient, cloud-native platforms supporting critical services. The environment encourages ownership, experimentation, automation, AI initiatives, continuous learning, certifications, and access to modern training resources.

Description

Do you thrive in highly dynamic, fast-paced environments where reliability, automation, and innovation are at the core of engineering culture? We are looking for a talented Site Reliability Engineer to join our SRE & Platform Engineering team in Warsaw. In this role, you will help design, automate, and operate scalable, resilient, and cloud-native platforms supporting critical business services. You will work closely with development, infrastructure, and security teams to improve system reliability, accelerate delivery pipelines, and drive operational excellence through automation and AI-driven initiatives. At our company, we strongly encourage innovation, ownership, and continuous learning. Engineers are empowered to propose and implement new ideas, experiment with automation and AI technologies, and continuously expand their skills through access to modern training platforms, certifications, and learning resources. What You Will Do Drive automation and reliability initiatives across infrastructure and application platforms to reduce operational overhead and improve system resilience Design, maintain, and optimize CI/CD pipelines and deployment strategies for secure, scalable, and efficient software delivery Manage and improve Kubernetes/OpenShift-based container platforms in production environments. Develop automation solutions and internal tooling using Bash/Shell scripting, Python, and Java (Scala experience is a strong plus). Partner with development teams to improve observability, performance, scalability, and operational readiness of services. Lead incident response activities, perform root cause analysis, and implement preventive measures to improve platform stability. Build and maintain monitoring, alerting, and logging solutions to ensure high availability and operational visibility. Contribute to infrastructure-as-code and GitOps practices using modern DevOps and cloud-native tooling. Participate actively in Agile/Scrum ceremonies and contribute to continuous improvement initiatives within the engineering organization. Explore and adopt AI-powered operational tooling and automation opportunities to improve engineering efficiency and reliability. What You Bring Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field. 5+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering roles within modern containerized/cloud-native environments Strong Linux system administration expertise with advanced Bash/Shell scripting skills. Strong programming experience in Python, experience with Java is required, and Scala knowledge is considered a strong advantage Hands-on experience with container and orchestration technologies including Docker, Podman, Kubernetes, and OpenShift Strong understanding of CI/CD, GitOps, and infrastructure automation practices Experience with tools and technologies such as GitLab CI/CD, ArgoCD, Helm, Ansible, Kafka, and observability/monitoring platforms Solid troubleshooting and problem-solving skills in complex distributed systems environments Experience with production incident management, root cause analysis, and reliability engineering practices. Excellent communication skills in English, with the ability to collaborate effectively across global teams.