SRE / Infrastructure Engineer, ML Platform

Active Connector — Japan · Posted ~4 hours ago

Hybrid ¥7008000-¥11000000

Skills

SRE Infrastructure engineering AWS Terraform Cloud infrastructure Network architecture Security Observability High availability Incident response Monitoring SLO/SLA VPC PrivateLink VPC Peering TLS/mTLS

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Take end-to-end ownership of infrastructure supporting machine-learning and financial technology platforms. You will design secure AWS environments with Terraform, build reliable data and ML infrastructure, establish observability and incident-response practices, and improve availability and security.

Highlights

Opportunity to own infrastructure end to end, work across ML, application, and data engineering teams, and focus on reliability, security, observability, and high availability.

Description

SRE / Infrastructure Engineer, ML Platform Location: Tokyo, Japan Work Style: Hybrid / Remote Work Available Salary: ¥7,008,000 - ¥11,000,000 About the Role Our client is looking for an SRE / Infrastructure Engineer to take ownership of infrastructure supporting its ML Ops Platform, Digital Bank, and fintech products. You will play a key role in preparing the infrastructure for the company's planned Digital Bank launch, with a strong focus on reliability, security, observability, and high availability. This is an opportunity to own infrastructure end to end while working closely with application engineers, ML engineers, data teams, and other SREs. Responsibilities Design, build, and operate AWS infrastructure using TerraformBuild reliable infrastructure for serving interfaces, data pipelines, and ML platformsDesign secure network architectures using VPC, PrivateLink, VPC Peering, TLS/mTLSEstablish monitoring, alerting, SLOs/SLA and incident response processesImprove observability using CloudWatchSupport integration between Databricks and the ML Ops PlatformOptimize CI/CD pipelines and container-based deploymentsCollaborate with data, product, ML, and SRE teams across the group Tech Stack AWS, ECS Fargate, API Gateway, VPC, PrivateLink, SageMaker, S3, Glue, CloudWatch, Databricks, Terraform, GitHub Actions, Docker, Python, Shell, SQL The team also actively uses AI development tools such as Claude Code. Requirements Must Have: Around 5+ years of experience as an SRE or Infrastructure EngineerHands-on AWS experience with services such as ECS, API Gateway, VPC, and IAMProfessional experience using Terraform for IaC and automationExperience with networking and security, including VPC Peering, AWS PrivateLink, TLS/mTLSExperience building and operating cloud observability and monitoring systemsExperience with CI/CD and DockerBusiness-level Japanese Basic business-level English (TOEIC 700+ equivalent) Nice to Have: Infrastructure experience supporting MLOps or data platformsExperience with SageMaker, Databricks, or AirflowExperience establishing SLO/SLA, on-call, and incident response processesFinancial services or payments security/audit experienceExperience building shared platforms across multiple teams or organizations Why Join? Own infrastructure supporting a major Digital Bank and fintech initiativeWork with modern AWS, MLOps, cloud security, and observability technologiesStrong ownership and autonomy over infrastructureCollaborate closely with ML, application, data, and product engineersOpportunity to shape architecture and operational processes from an early stageWork in an environment actively adopting AI development tools We also have several other software engineering, SRE, infrastructure, data, and AI roles available, so if this position isn't the right fit, we'd be happy to explore other opportunities with you.