SRE/DevOps Engineer - Enterprise Storage

Pgcdigital America Inc — Czechia · Posted ~2 hours ago

Senior Contract Onsite

Skills

Infrastructure engineering Storage engineering SRE/DevOps Python Shell scripting Infrastructure as code Monitoring Observability OpenStack VMware Shell Infrastructure as Code

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A long-term SRE/DevOps opportunity centered on enterprise storage and virtual infrastructure. You will architect and operate robust platforms, automate workflows with Python and shell scripting, implement infrastructure as code, and build monitoring and observability capabilities while solving performance and capacity challenges.

Highlights

Long-term opportunity combining enterprise storage, cloud infrastructure, automation, and SRE practices, with substantial hands-on Python development and architecture responsibilities.

Description

Job Title: SRE/DevOps - Enterprise Storage Platform Location: Prague Duration: Long Term Role Summary This role combines enterprise storage engineering with workflow automation and SRE/DevOps practices to design, deploy, and maintain robust infrastructure. Key responsibilities include architecting storage solutions, automating repetitive tasks via Python and shell scripting, managing infrastructure as code, and implementing comprehensive monitoring and observability systems. Candidates must possess at least 5+ years of experience in Infrastructure/Storage Engineering or SRE/DevOps roles, along with a minimum of 3+ years of hands-on Python development experience. What You Will Do ● Enterprise Storage and Virtual Infrastructure Integration (OpenStack, VMware) ○ Solve customer performance and latency issues, provide workflow-based recommendations, and perform storage architect design and implementation, including capacity planning. ○ Use software development methods to automate workflows and eliminate manual toil. ○ Remotely deploy storage using scripts and internal tooling. ○ Troubleshoot storage issues or script failures. ○ Lead incident response and manage interrupts during critical incidents. ● Build automation and tools in Python and shell scripting ○ Develop Python‑based tools and automation to support self‑service workflows ○ Automate repetitive operational tasks and enhance deployment script efficiency. ○ Integrate with internal and external APIs to orchestrate infrastructure workflows (compute, storage, network) ● Configuration Management, Infrastructure, and Monitoring as Code ○ Use tools such as Ansible, Terraform, Puppet, or similar tools to manage infrastructure declaratively. ○ Maintain reusable playbooks/modules and templates for common infrastructure patterns ○ Enforce configuration standards, security baselines, and repeatable deployments across environments ● Monitoring, observability, and reliability ○ Implement and improve monitoring, alerting, and dashboards for infrastructure health (e.g., Prometheus, Grafana, ELK/Nagios, or similar tools) ○ Define and track key metrics (availability, latency, capacity, error rates), and drive improvements based on data ○ Participate in incident response, perform root cause analysis, and implement long‑term fixes and runbooks ○ Maintain the observability codebase and actively develop new monitors and alerts. ● Collaboration and support ○ Partner with engineering teams to understand their infrastructure needs and design appropriate solutions from a storage perspective. ○ Provide guidance on best practices for using infrastructure platforms (VMs, containers, storage, networking) ○ Participate in an on‑call rotation and planned maintenance windows as needed Minimum Qualifications ● 5+ years of experience in Infrastructure/Storage Engineering or SRE/DevOps roles supporting enterprise storage organizations or teams. Knowledge of Everpure products is a plus. ● 3+ years of hands‑on Python development for: ○ Automation scripts and tools ○ REST API integrations ○ Data collection, reporting, and operational tooling ● Experience with at least one configuration management or IaC tool (e.g., Ansible, Terraform, Puppet, Chef) ● Linux systems administration knowledge (e.g., Ubuntu, CentOS/RHEL) ● Understanding of networking fundamentals (TCP/IP, DNS, DHCP, VLANs, routing basics) ● Experience with office tools such as Jira, Slack, Google Workspace, and others. ● Knowledge of monitoring and alerting on key storage metrics. ● Strong problem‑solving skills, ownership mindset, and clear written and verbal communication Preferred Qualifications ● Familiarity with storage platforms (SAN/NAS), ideally all‑flash arrays and/or object storage ● Knowledge of Pure Storage products FA and FB is a plus. ● Strong knowledge of on-premises infrastructure and architecture design. ● SRE/DevOps knowledge, monitoring and alerting philosophy, and Incident Commander system knowledge. ● Exposure to CI/CD tooling and pipelines (e.g., Jenkins, GitHub Actions, ArgoCD, GitLab CI) ● Experience in globally distributed teams and “follow‑the‑sun” support models is a plus