Summary
✨ AI‑Generated
A telecommunications engineering team is seeking two Senior Site Reliability Engineers to own the infrastructure beneath a cloud-native 4G/5G core. You will operate Linux containers and Kubernetes, build reliable deployment paths, maintain observability, manage network communication layers, and validate recovery and resilience through chaos and scale testing.
Highlights
Senior SRE role owning infrastructure for a cloud-native 4G/5G core network. The position combines Kubernetes and AWS operations with observability, automated recovery, networking, chaos testing, and scale validation in a technically demanding telecom environment.
Description
Job Description
Team: RAN &Core Software Engineering
Experience- 5+ Years
Location: Finland (Hybrid)
About the Role
Customer's 4G/5G converged core runs as a set of stateless network functions on AWS EKS, backed by a two-tier datastore (ElastiCache Valkey and MemoryDB).
We're hiring two Site Reliability Engineers to own the infrastructure this core runs on — from the container and Kubernetes layer up through the observability stack that tells us whether it's actually healthy.
You'll work alongside the engineers building the core itself, but your focus is the platform underneath it: is it deployed correctly, is it observable, does it recover automatically, and does it hold up under chaos and scale testing before it ever sees production traffic.
What You'll Own
Build and operate the Linux-based deployment path for core network functions — containers, orchestration, and the network communication paths (SCTP, HTTP/2, Diameter) those services depend on.
Run and extend our Kubernetes/EKS environment: node group sizing, multi-AZ topology, Helm-based deployment, and the CI/CD pipeline that pushes releases through it.
Own observability for the core: Prometheus/Grafana dashboards,
Open Telemetry-based tracing, and log aggregation — built so an on-call engineer can diagnose a production issue from the dashboard, not by guessing.
Design and run chaos, HA, and disaster-recovery testing (pod eviction, AZ partition, node failure) as a standard part of the release cycle, not a one-off exercise.
Build toward a dedicated AWS environment for the Core team specifically — including Terraform and Flux-based infrastructure-as-code, separate from shared platform resources.
What We're Looking For
Experience building and deploying Linux-based applications that communicate over real network protocols— comfortable below the HTTP layer, not just behind a load balancer.
Proficiency with containerization and orchestration (Kubernetes, Docker), and hands-on experience with at least one major cloud platform (AWS, GCP, or Azure) — AWS strongly preferred given our current stack.
Experience with common monitoring and observability tooling (Prometheus, Grafana, Open Telemetry, or equivalent), and a real understanding of what makes a distributed system observable versus just logged.
Deep understanding of distributed architectures and microservices patterns in a cloud-native environment — you can reason about what happens when one replica of a stateless service dies mid-request.
Comfortable owning infrastructure-as-code (Terraform, Flux/GitOps) rather than clicking through a console.
Bonus: experience supporting telecom or other latency-sensitive, high-availability network infrastructure.
What Success Looks Like
A pod, node, or AZ can fail during a chaos test, and the system recovers automatically — and you already knew it would, because you test it.
When a service degrades in production, the dashboards tell the story before anyone has to SSH into anything.
Infrastructure changes go through Terraform and Flux, reviewed and repeatable — not a one-time manual fix nobody can reproduce.
If interested kindly share your updated CVs to remya-saras@dyneits.com