DevOps / Platform Engineer

Yassir — United States · Posted ~1 day ago

Senior

Skills

DevOps Platform engineering Infrastructure as Code Cloud infrastructure Kubernetes CI/CD Reliability engineering Automation Technical documentation

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A platform engineering role with end-to-end ownership of clusters, infrastructure as code, delivery pipelines, and reliability guarantees. You will help a lean engineering team operate resilient systems through automation, rigorous documentation, asynchronous collaboration, and strong quality controls.

Highlights

Own platform and delivery infrastructure end to end in a lean, senior engineering environment, with strong emphasis on resilience, automation, written documentation, predictable delivery, and modern AI-assisted engineering practices.

Description

Context & Vision In 2026, writing code is no longer the primary bottleneck; managing its complexity and ensuring its reliability is. We are building a highly resilient advertising platform with a very lean, senior internal team, and we intend to keep it that way. To achieve scale without the overhead of a large engineering department, we rely on an AI-first development paradigm and a philosophy borrowed from the best large-scale open-source projects. Our code is not public, but our governance is theirs: asynchronous communication, exhaustive written documentation, explicit rules, and uncompromising quality gates. It is the only way a small team of humans and AI agents ships serious systems without accumulating technical debt. This role owns the ground the whole platform runs on: the clusters, the pipelines, and the guarantees that everything ships predictably. The Role You will own the platform and delivery infrastructure end to end — the clusters, the infrastructure-as-code, the delivery pipelines, and the operational guarantees behind them. In a lean team, reliability is not a separate department; it is a discipline you carry for everyone. Infrastructure as Code. Own the cloud footprint through infrastructure-as-code and a GitOps workflow. Infrastructure is declared, reviewed and versioned like any other code — no click-ops, no undocumented state Delivery Pipelines.Own CI/CD. Builds are reproducible, deployments are predictable, and rollbacks are boring. You make shipping a non-event Reliability, Performance & DR. Own observability (metrics, logs, distributed traces), performance testing, and backup / disaster-recovery. You define the SLOs that matter for a real-time serving platform and you make them measurable Quality Gates & AI-First CI. Enforce quality at every passage point in the pipeline. Integrating AI into CI/CD — automated reviews, security and policy checks — is an open frontier here, and we expect you to study it and propose implementations. Everything you build is documented; if it is not written down, it does not exist The Tech Stack Orchestration: managed Kubernetes on Google Cloud Platform IaC & GitOps: infrastructure-as-code and a GitOps workflow CI/CD: modern delivery pipelines (and proposals to evolve them) Observability: metrics, logs, distributed tracing Runtime context: Rust services, a React frontend, streaming, relational databases Cloud:Google Cloud Platform Profile & Requirements We are looking for a platform engineer who treats infrastructure as a product, owns reliability for the whole team, and is comfortable in a lean, high-quality, AI-first environment. This role is not suited to someone who wants to run a ticket queue inside a large ops team. Essential Experience 4+ years in DevOps / platform / SRE roles, running production Kubernetes** for real workloads (GCP preferred) Infrastructure-as-code and a GitOps workflow in production; infrastructure declared and reviewed as code Solid CI/CD ownership and a real observability practice (metrics, logs, distributed tracing) Experience defining and defending SLOs, performance and disaster-recovery for latency-sensitive systems Core Competencies & Mindset AI Development Lifecycle: comfort integrating AI into the delivery workflow, and a point of view on automating quality and security gates in CI Uncompromising Reliability: a maniacal focus on predictability, recoverability and security. Surprises in production are the enemy Written & Asynchronous: exceptional written communication; runbooks and ADRs are part of the job, not a favour. Effective in a distributed, async environment (Paris timezone +/- 3h) Ownership:** you carry reliability for the whole team and raise risks before they become incidents We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.