Senior Site Reliability Engineer

Buildplusrecruitment — Japan · Posted ~4 hours ago

Senior Visa History ✓

Skills

Site Reliability Engineering Production architecture SLOs SLIs Monitoring Capacity management Incident management Automation Observability SRE SLO SLI AI infrastructure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An innovative technology initiative focused on enterprise AI is seeking a Senior Site Reliability Engineer to design and refine production infrastructure for AI-driven applications. You will define and optimize SLOs and SLIs, improve monitoring and capacity management, establish incident processes, automate operational work, and build strong observability. The position emphasizes secure, scalable, reliable systems and efficient production operations.

Highlights

Work on enterprise AI infrastructure with a strong focus on reliability, scalability, security, observability, and operational excellence. The role combines architecture, automation, performance engineering, and incident management in a technically advanced environment.

Description

Hello all my Site reliability engineers! I am helping my client - an innovative, premier technology initiative delivering state-of-the-art enterprise AI integration across major industry leaders in Japan - to hire a senior SRE. By translating advanced artificial intelligence models into secure, scalable business practices, my client's teams enable complex operational transformation. They bridge modern AI research and mission-critical enterprise systems to establish new standards for digital reliability. Core Responsibilities & Scope of Work Production Architecture: Design, launch, and refine the live operational framework for AI-driven enterprise tools.Reliability & Performance Engineering: Define, measure, and optimize SLOs/SLIs, system health monitoring, capacity management, and incident protocols to preserve platform availability and cost-efficiency.Automation & Observability: Eliminate manual tasks and build robust infrastructure by engineering custom Runbooks, clear alerting flows, and automated deployment processes.Cross-Functional Governance: Drive alignment across internal tech units, legal, auditing, cyber-security teams, external vendor partners, and client operational divisions.AI Security & Lifecycle Management: Implement best practices tailored to advanced language models, AI agents, data protection, compliance, and ongoing model performance evaluations. Key Requirements Mandatory Language Proficiency: high Japanese fluency (oral and written) and Business-level English are required for daily stakeholder engagement and cross-border collaboration.System Operations: 2+ years of hands-on experience in platform reliability, infrastructure operation, or enterprise system management.Cloud Architecture: Proficiency with public cloud ecosystems (AWS, Google Cloud, or Microsoft Azure), including infrastructure provisioning and cost control.AI Lifecycle Awareness: Familiarity with the operational nuances of GenAI/LLM deployments, agent architectures, safety guardrails, and data security standards. Preferred / Nice-to-Have Background in SRE, Platform Engineering, or Production Engineering.Demonstrated ability in client-facing technical consultation and cross-team consensus building.Knowledge of information security compliance standards and enterprise audit frameworks.Experience setting up monitoring or automated testing pipelines for Large Language Models. If you meet the requirements and want to explore this opportunity, please apply for this role and let's have a quick call to discuss it further. I look forward to meeting you!