Senior Site Reliability Engineer

Ibm — Poland · Posted ~4 hours ago

Senior Visa History ✓

Skills

Site Reliability Engineering Software development lifecycle Infrastructure design Python Bash Terraform Ansible OpenShift Cloud infrastructure Monitoring Incident response

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A Senior Site Reliability Engineer is sought to design, maintain, and modernize scalable and highly available infrastructure supporting business applications. The role combines software engineering with automation, cloud infrastructure management, monitoring, troubleshooting, alerting, and incident response, using scripting and configuration-management tools.

Highlights

Modernize reliable and highly available infrastructure for internal applications in an agile environment. Automate operational work, manage cloud infrastructure, improve observability and reliability, and collaborate closely with software developers.

Description

Introduction IBM Red Hat's IT Developer Enablement Applications team is looking for a Senior Site Reliability Engineer to help us create, modify, and modernize applications for our internal customers. You'll be given unlimited opportunities to make an impact in a global company built on open-source values and principles. Your Role And Responsibilities Execute full software development life (SDLC) cycle within an agile environmentDesign, implement, and maintain scalable, reliable, and highly available infrastructure systems to support business applications and services.Automate repetitive tasks and processes using scripting and configuration management tools (e.g., Python, Bash, Terraform, Ansible).Manage cloud-based infrastructure (OpenShift) and optimize resources for reliability, performance, and security.Monitor system performance, troubleshoot issues, and ensure uptime and reliability through proactive alerting, metrics, and incident response.Collaborate with developers to improve application resilience, deploy code efficiently, and integrate CI/CD pipelines.Conduct root cause analysis (RCA) for incidents, document findings, and implement preventive measures to minimize downtime.Continuously improve system architecture, deployment processes, and tooling to enhance operational efficiency and scalability.Ensure security best practices are applied across infrastructure, including access controls, network configurations, and compliance requirements.Effectively communicate to stakeholders and project team members to ensure proper visibility into development effortsAbility to respond to the customer's tickets and inquiries Preferred Education Bachelor's Degree Required Technical And Professional Expertise 5+ years of experience with Linux Systems Engineering (RHEL, CentOS, etc)5+ years of experience delivering enterprise software applications2+ years of experience with Git (experience with GitLab a plus)Strong knowledge of Linux and networking internalsStrong programming proficiency in Python and experience with the full SDLCExperience with containers and container orchestration tools (Docker, Podman, Kubernetes, OpenShift, etc)Good knowledge of CI/CD pipelines and tooling (GitLab CI, GitHub Actions, etc)Experience with configuration management tools (Ansible  etc)Familiarity with monitoring and logging suites (e.g., Sumo Logic, Prometheus, or ELK).Familiarity with SRE fundamentals and principlesFamiliarity with bare metal environments and APIs such as RedFishData driven and observability first mindset (logs, metrics, traces) Preferred Technical And Professional Experience n/a