Site Reliability Engineer

Wearequantumpeople — Japan · Posted ~3 days ago

Senior Full-time

Skills

DevOps Site Reliability Engineering Linux CI/CD monitoring automation deployment incident recovery

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An experienced DevOps or Site Reliability Engineer is sought to build a 0-to-1 operational platform for a highly advanced computing system. The role owns how systems are deployed, monitored, tested, recovered, and upgraded across development and production-like environments, offering substantial technical ownership in a new engineering domain.

Highlights

Opportunity to build a new operational foundation from the ground up for advanced computing systems, with broad ownership across deployment, monitoring, testing, recovery, and upgrades.

Description

🚀 DevOps / Site Reliability Engineer Quantum Computer Control Systems Japan | Permanent, Full-time | Confidential Deep-Tech Company Help build the systems that keep a quantum computer running Most DevOps roles improve an existing platform. This one asks you to create the operational foundation for an entirely new kind of machine. We are representing an ambitious deep-tech organisation developing advanced quantum computing technology in Japan. As it moves beyond supporting experiments and towards building an operational machine, its quantum control stack must become reliable enough to run as a product. We are seeking an experienced DevOps or Site Reliability Engineer to take ownership of that challenge. You will define how the control system is built, deployed, monitored, tested, recovered and upgraded across development and live machine environments. This is a genuine 0→1 engineering role. There is no established operations platform waiting for you. You will build it. Your mission Your goal will be to create the operational infrastructure required to move an advanced quantum computing system towards dependable, extended operation. By reducing firefighting and eliminating recurring operational work, you will enable scientists, architects and senior engineers to remain focused on developing the next generation of the machine. You will need to be comfortable working where requirements are incomplete and technical constraints continue to evolve. Success will depend on sound engineering judgement, practical decision-making and the ability to improve systems progressively as the underlying technology matures. What you will do CI/CD and release engineering Design and build CI/CD pipelines for the quantum control software stackAutomate builds, testing, releases and deploymentEstablish reliable promotion paths across development and operational environmentsCreate safe and repeatable upgrade and rollback procedures Monitoring and observability Build monitoring, logging, metrics, dashboards and alerting across the control systemIntegrate monitoring of physical environmental conditions, including temperature, humidity and vibrationImprove the early detection, diagnosis and recovery of degraded system behaviour Hardware-integrated testing Develop Hardware-in-the-Loop testing and validationEnsure changes to the control stack are tested safely against physical devices before releaseAutomate regression testing and operational verificationEstablish dependable processes for validating software changes in hardware-integrated environments Operational reliability Take ownership of incident response and root-cause analysisIdentify and remove recurring sources of operational toilImprove configuration management, deployment safety and system recoveryIntroduce the standards and practices required to move towards product-level reliability Cross-functional engineering Work closely with control software engineers, systems architects, physicists and hardware specialistsPartner with network engineering while retaining ownership of software reliability and deploymentFeed operational learning into the architecture of future quantum computing systems What we are looking for You will bring: At least five years of experience in DevOps, Site Reliability Engineering or systems engineeringDemonstrable ownership of reliability within production or operational environmentsStrong experience designing and operating CI/CD pipelinesPractical expertise in build, release and deployment automationExperience building monitoring and observability platforms across metrics, logs and alertsPython, or an equivalent language, for automation and internal toolingStrong Linux administration, diagnostics and troubleshooting skillsExperience with incident response, root-cause analysis and continuous reliability improvementBusiness-level EnglishJapanese proficiency equivalent to JLPT N2 or aboveThe ability to make progress when requirements and system behaviour are still evolvingThe curiosity and determination to learn an unfamiliar technical domain Previous quantum computing experience is not required. Experience that would make you stand out We would be particularly interested in experience involving: Infrastructure as Code using Terraform, Ansible or comparable technologiesDocker, Kubernetes or other container and orchestration platformsPrometheus, Grafana or related time-series monitoring technologiesScientific instruments, laboratory systems or research infrastructureHardware-in-the-Loop testing or validation against physical devicesHigh-availability, mission-critical or 24/7 operational systemsControl systems, robotics, embedded systems or industrial automationPhotonics, semiconductor equipment or advanced manufacturingEnvironmental monitoring involving temperature, humidity or vibrationAn active interest in quantum computing, physics or frontier technologies Why join? Define operations from the ground up You will not be constrained by a mature platform or inherited operating model. You will establish the foundations, standards and practices for how the system is operated. Solve a genuinely difficult reliability challenge This is not conventional cloud DevOps. The platform combines software, scientific equipment, control systems and physical hardware. Make a direct impact on the technical roadmap Every improvement you make will reduce operational burden and allow senior engineers and scientists to spend more time advancing the machine. Work towards a clear engineering objective Your mission is tangible: help move a sophisticated quantum computing system towards reliable, continuous operation. Influence future architecture The operational knowledge you develop will help shape the architecture of the next generation of quantum computing systems. The opportunity Few engineers will have the chance to define how a quantum computer is deployed, monitored, tested, recovered and kept running. If you are an experienced DevOps or SRE professional who wants to apply your expertise beyond conventional software infrastructure, this is an opportunity to help move one of computing’s most promising technologies towards dependable real-world operation. Apply through the LinkedIn job link or contact Drew Percival at Quantum People for a confidential discussion. #DevOps #SRE #SiteReliabilityEngineering #QuantumComputing #DeepTech #Linux #Python #CICD #Observability #InfrastructureAsCode #EngineeringJobs #JapanJobsbs