Senior DevOps Engineer

R Systems — Poland · Posted ~3 hours ago

Senior Full-time

Skills

DevOps CI/CD microservices automation observability infrastructure management GitLab CI

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior DevOps position focused on building deployment ecosystems, managing scalable infrastructure, creating CI/CD pipelines, and improving reliability for microservice-based systems.

Highlights

High-impact role focused on automation, reliability, scalable infrastructure, and improving software delivery processes.

Description

Senior DevOps Engineer About the Role We are seeking a skilled and collaborative Senior DevOps Engineer to join our growing infrastructure team. In this role, you will be a key architect of our deployment ecosystems, ensuring the stability, scalability, and efficiency of our microservices architecture. You will work closely with development teams to bridge the gap between code and production, fostering a culture of automation and reliability. This is a high-impact position designed for a technical expert who thrives in high-load environments and values elegant, automated solutions to complex infrastructural challenges. Responsibilities Maintain and develop robust tools for building and deploying microservices to ensure seamless delivery cycles.Architect and oversee comprehensive observability systems to maintain deep visibility into system health.Design and build sophisticated CI/CD pipelines using GitLab CI to streamline the software development lifecycle.Manage and scale infrastructure as code (IaC) using Ansible and Terraform.Configure and optimize monitoring and logging stacks based on VictoriaMetrics, Grafana, Loki, and Alertmanager.Develop automation scripts using Bash, Python, or Golang to eliminate manual tasks and improve operational efficiency.Collaborate with engineering teams to automate development and testing processes, reducing time-to-market.Participate actively in Site Reliability Engineering (SRE) activities to maintain high system availability.Required Qualifications Professional experience building and maintaining complex CI/CD circuits.Proven track record of managing high-load production environments.Expertise in Infrastructure as Code (IaC), specifically with Terraform, Ansible, or both.Hands-on experience with major cloud providers (AWS, GCP, or Azure).Deep proficiency with Docker and Kubernetes for container orchestration.A strong foundational understanding of network technologies and the TCP/IP suite.Demonstrated ability to write automation scripts in Bash, Python, or Golang.Extensive experience in Linux infrastructure administration.Excellent communication skills with the ability to articulate technical perspectives clearly and constructively.Proficiency in written and spoken English.Preferred Qualifications Familiarity with security frameworks and tools, such as HashiCorp Vault and TLS management.Experience implementing SSDLC (Secure Software Development Life Cycle) practices.Knowledge of SIEM and IDS/IPS systems to enhance infrastructure security.Experience Level (5-8 Years) With 5 to 8 years of dedicated experience in DevOps or System Engineering, you are expected to operate with a high degree of autonomy. You should have a proven history of migrating or scaling infrastructure in production settings. At this level, you will be expected not only to execute tasks but to contribute to the strategic roadmap of our infrastructure, mentor junior colleagues, and drive best practices across the engineering organization. Key Performance Indicators (KPIs) Improve Deployment Frequency: Increase the successful deployment rate by 25% within the first 6 months through CI/CD pipeline optimization, measured by GitLab deployment logs.Infrastructure Automation: Transition 90% of manual infrastructure configurations to Terraform/Ansible within 9 months, measured by the ratio of manual changes to code-based commits.Mean Time to Detect (MTTD): Reduce MTTD for production incidents by 30% within 6 months, measured by Alertmanager and Grafana incident timestamps.System Availability: Maintain a 99.9% uptime for core microservices over the first 12 months, measured by SRE availability dashboards.Scripting Efficiency: Automate at least 5 recurring manual administrative tasks using Python or Golang within the first year, measured by a reduction in manual tickets handled by the team.