Senior Platform Engineer (SRE / DevSecOps)

Agion Inc — Finland · Posted ~3 hours ago

Senior

Skills

Site Reliability Engineering DevSecOps Infrastructure automation Observability Telemetry Incident diagnosis Automated recovery Distributed systems Cloud infrastructure SRE Cloud

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior platform engineering role at the intersection of SRE and DevSecOps, focused on making distributed infrastructure diagnosable and recoverable even when direct host access is unavailable. You will use telemetry and observability to identify failures, automate recovery paths, and build reliable operational systems that scale across customer-managed environments.

Highlights

Advanced platform engineering role solving challenging reliability problems across customer-controlled infrastructure. The work emphasizes observability, automated diagnosis and recovery, resilient distributed systems, and designing operational solutions that remain effective as infrastructure scales.

Description

Company description Agion builds infrastructure that lets organisations give AI agents real responsibility while keeping authority, accountability and evidence under their own control. Agion runs agentic workloads in production for public-sector and industrial organisations across Europe, on infrastructure those customers control: their cloud, their data centre, their network. Role description Running agentic workloads on infrastructure our customers control creates an unusual operations problem. When a host misbehaves, you may not be able to log in. You diagnose it from the evidence it emits. The fix has to work without a shell. We want recurring failures to become automated recovery paths rather than better runbooks. Those properties have to hold as the fleet grows. In this role you will: * Make customer-managed systems diagnosable and recoverable without assuming shell access * Distinguish bad deployments, configuration drift, network failures and silent hosts from telemetry alone * Turn recurring failures into automated recovery paths * Keep authority bounded across humans, services and AI agents * Operate Kubernetes and GitOps across customer-controlled environments * Produce evidence that lets us and an outside reviewer establish what happened Infrastructure decisions are owned by the engineer responsible for them. How we work: we reproduce before diagnosing. Missing evidence stays missing rather than becoming an assumption. Risk acceptance has a rationale and an owner. Remote within the EU/EEA, working Helsinki hours (EET/EEST). €75k–€85k base plus stock options. Qualifications * You’ve operated Kubernetes and GitOps in production and owned incidents when things went wrong * You’ve built production observability across metrics, logs and telemetry * You’ve owned production security controls for access, secrets, authentication or authorisation * You understand distributed and multi-tenant systems * You can reason about least privilege and non-human identity when software acts on behalf of humans or organisations * You’re comfortable reading Go Experience with on-premise or hybrid cloud, infrastructure as code and regulated environments where audit evidence matters, such as SOC 2 or ISO 27001, is useful. If you meet most of the requirements, not every line, apply anyway. We welcome applicants regardless of background or personal characteristics. Privacy Policy: ⁠agion.ai/privacy