Platform Engineer - Operations and Reliability

Meiro Customer Data Infrastructure — Czechia · Posted ~3 hours ago

Mid Full-time

Skills

Kubernetes Infrastructure as Code PostgreSQL Cloud Infrastructure Production Troubleshooting TypeScript GCP

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A platform engineering role focused on operating scalable systems, automating infrastructure, improving reliability, and supporting cloud-native applications across distributed environments.

Highlights

Hands-on role combining infrastructure automation, reliability engineering, and collaboration with application teams in a modern technical environment.

Description

👀 Who we are? Meiro builds Pipes, a customer data infrastructure distributed as a single application image written in TypeScript and running on Bun. We operate a multi-tenant managed service across Kubernetes clusters in DigitalOcean, GCP and Hetzner, provisioning isolated namespaces, PostgreSQL databases, and identity configurations for each customer. Our infrastructure is designed for portability, supporting both our managed service and self-hosted customer deployments. ☀️ What's waiting for you? You will run the systems powering Pipes and automate their operations. This is a hands-on role combining production troubleshooting, infrastructure-as-code, and collaboration with application engineers. We are an AI-heavy shop, using coding agents for implementation and routine tasks. You don't need to be an AI expert, but you must be eager to use these tools effectively and remain accountable for their output. Responsibilities We need someone to maintain system health while reducing manual overhead: Production Operations: Monitor customer instances and clusters, investigate incidents using logs and metrics, and manage rolling updates or rollbacks.Lifecycle Management: Automate the provisioning and decommissioning of customer resources, including databases, networking, and identity providers.Platform Maintenance: Manage infrastructure using OpenTofu/Terraform and Ansible. Maintain baseline services like ingress, telemetry, and secrets integration.Data Reliability: Ensure health and performance for PostgreSQL and Elasticsearch-compatible services, including schema migrations and backup/restore drills.Observability & Delivery: Tune alerts in Grafana/VictoriaMetrics to be actionable and maintain GitHub Actions workflows for secure, fast deployments. 🎯 The experience & skills you'll bring We value attitude and a willingness to learn over an exact stack match: Experience: 1–3 years in infrastructure, SRE, or backend engineering with production exposure.Technical Core: Proficiency in Linux, container troubleshooting, and practical Kubernetes (inspecting workloads and logs).Automation & Data: Ability to write SQL and script repetitive tasks (we use TypeScript).Operational Mindset: Careful judgment when applying changes, clear writing for incident reports, and a desire to use AI tools in your daily workflow.Preferred Skills: Familiarity with OpenTofu/Terraform, Ansible, cloud providers (DigitalOcean/Hetzner), or observability stacks (Grafana/Prometheus) is a significant plus. On-Call Reality This role serves as the second-line contact for infrastructure incidents. Our "bugmaster" handles first-line application issues, escalating only when evidence points to the underlying platform. Our goal is stability; every out-of-hours call should result in a fix or better automation to ensure it doesn't happen again. Our Tech Stack Application & Tooling - TypeScript, Bun, React Data Services PostgreSQL, Elasticsearch Orchestration - Docker, Kubernetes (DOKS, K3s, k0s) Infrastructure & Cloud - OpenTofu, Ansible, DigitalOcean, GCP, Hetzner CI/CD & Observability - GitHub Actions, Grafana, VictoriaMetrics, VictoriaLogs Identity & Secrets - Authentik (OIDC), Infisical Ingress - Traefik 🚀 What Success Looks Like We want you to grow into a core member of our engineering culture: Month 1: Become an integral part of the team by learning our deployment model, shadowing incidents, and completing your first operational improvement.Month 3: Independently diagnose provisioning failures and deploy releases with minimal support.Month 6: Own a major reliability project from start to finish, demonstrably reducing avoidable pages and alerts. 🤝 Interview Process Intro Call: 30-minute conversation.Technical Deep-Dive: Discussion about a production incident you investigated.Practical Exercise: A short task resolving a failing deployment.Team Meet: Final conversation with your future colleagues. Why Meiro? Our environment is intense but supportive. It's a place where you are expected to be an "owner" and focus strictly on "impact", and you'll have a team around you to figure it out with.Flexible work environment with the hybrid option to use our office in Prague - KarlínAlthough we mostly work remotely, we love getting together from time to time - for team on-sites, team lunches, or just to catch up over drinks. 💜 What do we stand for? Ownership: you own outcomes, not just tasks.Openness: honest conversations, no politics.Respect & selflessness: strong teams beat strong egos.Growth: personal, professional, and organisational.Impact: quality over quantity, always.