Site Reliability Engineer – Enterprise Storage & Python

Principal Engineering — Czechia · Posted ~11 hours ago

Senior Full-time

Skills

Python SRE DevOps enterprise storage infrastructure automation Infrastructure as Code APIs monitoring observability performance troubleshooting automation

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Work as an SRE/DevOps engineer combining enterprise storage expertise with Python automation. You will design and integrate storage solutions, troubleshoot performance and latency, build automation tools, work with APIs and Infrastructure as Code, and improve monitoring and observability. At least three years of practical Python experience is required.

Highlights

Combine enterprise storage engineering, Python development, automation, and SRE practices to improve infrastructure reliability, performance, and operational efficiency.

Description

SRE / DevOps Engineer – Enterprise Storage & Python | Prague, Czech & English language Are you an SRE/DevOps Engineer who understands enterprise storage and can turn repetitive infrastructure work into reliable automation? This role combines enterprise storage engineering, Python development, infrastructure automation and SRE practices in one position. You’ll work on designing and integrating storage solutions, troubleshooting performance and latency issues, and building tools that reduce manual operational work. A key part of the role is hands-on Python development – we’re looking for at least 3+ years of real-world Python experience. You’ll also work with APIs, Infrastructure as Code, monitoring and observability to make infrastructure more reliable and easier to operate. If you have a strong background in enterprise storage / infrastructure engineering and enjoy solving complex operational problems through automation, this could be a strong fit. 🧩 What You’ll Do Design, deploy and maintain enterprise storage solutions, including capacity planning and storage architectureIntegrate storage with virtual infrastructure such as OpenStack and VMwareTroubleshoot storage performance, latency and deployment issuesBuild Python and shell-based automation and self-service toolsIntegrate REST APIs to orchestrate compute, storage and network workflowsAutomate repetitive operational tasks and improve deployment workflowsManage infrastructure and configuration as code using Ansible, Terraform, Puppet or similar toolsDevelop and maintain monitoring, alerting and dashboards using tools such as Prometheus, Grafana, ELK or NagiosTrack infrastructure metrics such as availability, latency, capacity and error ratesParticipate in incident response, root cause analysis and long-term remediationMaintain observability tooling, monitors, alerts and operational runbooksWork with engineering teams to design appropriate infrastructure and storage solutionsParticipate in on-call rotation and planned maintenance windows ✅ Must-Have Experience 5+ years of experience in Infrastructure / Storage Engineering, SRE or DevOpsExperience supporting enterprise storage environments3+ years of hands-on Python development, particularly for:Automation scripts and toolsREST API integrationsData collection, reporting and operational toolingExperience with at least one IaC / configuration management tool: Ansible, Terraform, Puppet, Chef or similarSolid Linux administration skills (Ubuntu, CentOS/RHEL)Good understanding of networking fundamentals: TCP/IP, DNS, DHCP, VLANs and basic routingExperience with infrastructure monitoring and alertingStrong problem-solving skills, ownership and clear communicationExperience with tools such as Jira, Slack and Google Workspace ➕ Nice to Have Familiarity with with SAN/NAS, all-flash arrays and/or object storageKnowledge of Pure Storage FA/FB or Everpure productsStrong understanding of on-premises infrastructure and architecture designSRE practices, Incident Commander methodology and observability principlesExperience with CI/CD tools such as Jenkins, GitHub Actions, ArgoCD or GitLab CIExperience working in globally distributed teams / follow-the-sun support models 📍 Engagement Details Czech + English languageLocation: PragueWork model: Hybrid – 3 days onsite / 2 days remote per weekEngagement: Full-timeStart: September 1, 2026End: September 2036 If enterprise storage, Python automation and infrastructure reliability are areas where you have real hands-on experience, we’d be happy to hear from you. #SRE #DevOps #Python #EnterpriseStorage #Infrastructure #Linux #Terraform #Ansible