Site Reliability Engineer

Ct19 — Japan · Posted ~1 week ago

Mid Full-time

Skills

SRE CI/CD monitoring observability deployment automation

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A Site Reliability Engineering role responsible for creating operational foundations, monitoring systems, deployment workflows, and reliable production environments for advanced technology platforms.

Highlights

Build reliability foundations from the ground up while working on advanced technology systems and automation challenges.

Description

Site Reliability Engineer (Quantum Control Systems): Location: Japan Job Summary: As we move from "experiment support" to building an actual machine, the control stack must become reliable enough to run as a product. You will own that reliability - the maintenance, monitoring, deployment, and upgrade path for our quantum control system, across both development and production environments - so that senior engineers stay on the roadmap instead of firefighting, and the machine can move toward continuous, day-long operation. This is a 0→1 operations role. There is no existing ops foundation; you build it. Expect incomplete requirements and shifting constraints. You'll make progress with sound engineering judgment and harden things as the system stabilizes - not wait for a perfect spec. Responsibilities: Build CI/CD, build, and release automation for the control software stack Stand up monitoring, observability, and alerting across the control system and the lab's environmental sensor infrastructure (temperature, vibration, humidity nodes) Build hardware-in-the-loop (HIL) test automation so control-stack changes are validated against the real devices before they ship Drive system reliability: incident response, root-causing, and removing recurring maintenance toil from the engineering teams Own deployment, configuration management, and the upgrade/rollback path Automate regression and validation so changes to the control stack ship safely Scope boundary: partner with the network engineer on lab/infra networking - you own software reliability and deployment, they own the network fabric Partner with the Architecture group on the path to product-level reliability and the next-generation machine's target of continuous operation Required Qualifications: 5+ years in SRE, DevOps, or systems engineering, with demonstrated ownership of production reliability CI/CD and build/release automation Monitoring / observability tooling (metrics, logs, alerting) Python (or equivalent) for automation and tooling Strong Linux fundamentals Business-level English for a global team Self-directed learner comfortable ramping into an unfamiliar (quantum / physics) domain Preferred Qualifications: Infrastructure-as-code (Ansible, Terraform) Containers and orchestration (Docker, Kubernetes) Time-series monitoring stacks (Prometheus / Grafana) and dashboarding Experience operating scientific instruments, lab systems, or hardware-in-the-loop setup Networking fundamentals Experience with high-reliability or 24×7 systems Interest in quantum computing / physics Why this role: Define how a quantum computer is operated and kept alive - from zero Directly unblock the roadmap by taking operational load off senior engineers Own a hard, concrete target: move the machine toward continuous operation Founding-era scope with a clear path into the new quantum computer architecture group