Summary
✨ AI‑Generated
Take ownership of infrastructure, reliability, and platform operations as a senior DevOps engineer. You will design and operate scalable systems across AWS and Kubernetes, improve observability and networking, strengthen CI/CD and security, and automate operational processes while partnering closely with software engineers.
Highlights
Hands-on ownership of infrastructure, reliability, and platform operations with broad technical scope across AWS, Kubernetes, observability, networking, CI/CD, security, and automation. The role emphasizes autonomy, problem solving, scalability, and continuous improvement.
Description
About Us
On behalf of our client we are looking for an experienced Senior DevOps Engineer to take ownership of our infrastructure, reliability and platform operations.
This is a hands-on engineering role for someone who enjoys building robust systems, solving complex production problems and continuously improving the way technology is delivered and operated.
You will own the reliability of our platform across infrastructure and application services, from AWS and Kubernetes, to observability, networking, CI/CD, security and automation.
You will work closely with software engineers to ensure our services are scalable, resilient, observable and production-ready.
We are looking for someone with a strong ownership mindset who proactively identifies problems and opportunities, builds solutions and drives improvements rather than waiting for tasks to be assigned.What You’ll Do:
Design, build and operate reliable, scalable infrastructure for our iGaming platform on AWS using Infrastructure as Code.
Take ownership of platform reliability across infrastructure and application services, including Kubernetes, networking, service health and performance.
Define and maintain SLIs and SLOs and establish meaningful reliability targets for critical services.
Build and continuously improve end-to-end observability across metrics, logs and distributed tracing, including dashboards and actionable alerting.
Investigate and resolve production incidents end-to-end, perform root cause analysis and turn findings into permanent reliability improvements.
Design, build and maintain CI/CD and GitOps workflows, with automated validation, safe deployment strategies and reliable rollback mechanisms.
Work directly with application code when required — adding instrumentation, improving health checks, troubleshooting services and collaborating with developers on technical solutions.
Own and improve infrastructure around networking, load balancing, CDN, DNS, TLS and traffic management.
Identify opportunities to automate repetitive operational tasks and reduce manual intervention.
Continuously improve security, performance, scalability, efficiency and cost predictability across the platform.
Work closely with software engineers on architecture, scalability, resilience, failure scenarios and production readiness.
Contribute to technical architecture and help identify potential reliability issues before they reach production.
Develop and maintain practical runbooks, operational procedures and technical documentation.
Build and maintain a technical backlog of infrastructure and reliability improvements and proactively drive initiatives through to completion.
Use modern AI tools and AI agents as part of your daily engineering workflow to accelerate investigation, debugging, automation, code development and problem-solving.
Requirements
4+ years of experience in SRE, DevOps, Platform Engineering or Infrastructure Engineering.
2+ years of hands-on experience running Kubernetes in production.
Strong hands-on experience with AWS and a solid understanding of cloud infrastructure, including compute, networking, IAM, storage, databases, load balancing, CDN and managed services.
Strong experience with Kubernetes, Terraform and Helm.
Practical experience with modern observability tooling such as Prometheus, Grafana and OpenTelemetry, or equivalent technologies.
Strong Linux and networking knowledge, including DNS, TLS, routing and container runtimes.
Experience building and maintaining CI/CD pipelines and GitOps workflows.
Experience with GitHub Actions and ArgoCD or Flux is highly desirable.
Ability to read, understand and modify application code when required.
Strong scripting and automation skills, particularly with Bash and/or Go.
Good understanding of distributed systems, scalability, high availability and common failure modes.
Strong troubleshooting and analytical skills, with the ability to investigate complex issues across multiple layers of a technology stack.
Experience with modern AI-powered engineering tools and agents and a genuine willingness to incorporate them into your everyday workflow.
A strong ownership and continuous-improvement mindset.
You proactively identify reliability, performance, security and operational opportunities, build your own technical backlog and drive initiatives forward.
Comfortable working collaboratively with software engineers and contributing to technical and architectural discussions.
Fluent Ukranian or Russian is a must.
English B1 level or higher is a must.
Benefits
Pay that actually respects your work
Work that grows with you - real learning, mentorship, and fast skill development
Fast career progression - if you grow fast, we grow you fast
Hybrid freedom - office + remote mix that actually fits your life
Flexible Working Hours
21 vacation days + 7 sick days (no questions asked)
Birthday half-day off
Fully stocked snacks & drinks station to keep energy (and morale) high
Brand New MacBook provided
Modern, fast-moving environment - things change fast, and so do opportunities
Ownership from day one - you’re trusted with real responsibility