Site Reliability Engineer

Money Forward — Japan · Posted ~2 hours ago

Visa History ✓

Skills

Site reliability engineering Infrastructure engineering Data systems Cloud services System operations

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A site reliability and infrastructure engineering role supporting large-scale data-driven services. You will help design and operate reliable infrastructure for systems that collect and process user data from diverse sources, while contributing to the development of new products and capabilities. The position offers exposure to financial technology, data integration, and large-scale production systems.

Highlights

Opportunity to work on infrastructure supporting widely used financial and data-driven services. The role focuses on reliable systems, large-scale user data infrastructure, and building new capabilities with meaningful real-world user impact.

Description

About the Company Under our mission, "Money Forward. Move your life forward”, Money Forward offers a range of services, including the automatic household accounting and asset management service "Money Forward ME" and the business-oriented cloud service "Money Forward Cloud," which are used by many users. For these services, utilizing various user data collected through a technology called "Account Aggregation" is essential. The Account Aggregation Team develops systems to collect user data. The collected data becomes more valuable information through our products and is returned to the users. However, we believe that the collected user data could have more diverse use cases for each user. The data collected from various sources is also connected to the users' lives. By enabling users to use their data more conveniently, we aim to move their lives forward. With this vision, we are now developing a new product. We are looking for an infrastructure engineer who will work with us to optimize the development and operation of this product. Responsibilities Designing, building, and operating cloud (AWS) and on-premises infrastructure.Establishment of metrics, monitoring, and alerts using tools like Datadog, as well as incident response (including troubleshooting, recovery, incident management, and post-mortems).Performance evaluation of applications and infrastructure.Performance optimization and the creation of systems to support it.Development of systems for incident response, detection, and prevention.Enhancement of service reliability, including the definition and management of SLIs/SLOs for continuous performance improvement.Security operations and compliance management for the entire infrastructure.Designing, building, and operating platforms to maximize the productivity of development teams (CI/CD, development environments, and testing environments).Conducting availability and reliability reviews from the design phase onward.Addressing financial industry-specific regulatory requirements. Required Skills 3+ years of practical experience in design and operation in SRE, DevOps, or infrastructure domains.Basic knowledge and work experience with Linux, Network, Security, etc.Experience in using/designing/operating AWS.Experience with Terraform.Experience with container orchestration systems like Kubernetes or ECS.Experience in building and utilizing monitoring environments with tools for monitoring and observability.Development experience with peer reviews using Git, such as Pull Requests. Preferred Skills Strategic planning skills (Ability to understand company and business challenges, define a clear direction for your organization, and present a rational path to stakeholders).Experience in building and operating Kubernetes in a multi-tenant environment.Experience in handling failures in web services.Experience as an SRE in a web service company.Experience in operating large-scale services on AWS.Experience in building CI/CD pipelines.Experience in implementing and operating monitoring for web services.Experience in capacity planning and tuning for web services.Experience in operating MySQL (experience with version upgrades, knowledge of replication, etc.).Experience with Infrastructure as Code (IaC).Experience in AI development and/or experience in using AI tools to improve development processes. Equal Opportunity Statement Money Forward is at a major turning point, shifting "from Cloud to AI." We are currently driving "AX (AI Transformation)"—the next step beyond DX—with the goal of providing "Digital Workers," where AI agents autonomously execute tasks. As we enter a phase of evolving into Japan's No. 1 back-office AI company by integrating AI agents into all of our products in the future, we are looking for individuals who can contribute to AI-driven development and value creation. (More information here) Language Requirements Japanese: Not required. The ability to grasp the gist of technical discussions in Japanese is a plus. Speaking ability is not required.English: TOEIC 700+ or equivalent. Must be able to handle meetings and text-based communication in English from the start of employment. We will consider other qualifications or experiences that demonstrate your English proficiency. Examples: EIKEN Grade Pre-1, TOEFL iBT 60+, IELTS 5.0+, etc. Candidates who do not have a qualification equivalent to TOEIC 700+ may be asked to take a company-designated English test during the selection process (typically after the first interview). We are looking for someone who Likes Money Forward's services and wants to accelerate the speed of value delivery.Can take abstract challenges or high-level policies, break them down into concrete requirements, select an appropriate architecture, and drive it to completion.Is passionate about expanding business possibilities through technology.Can proactively catch up and tackle unfamiliar technologies.Enjoys building trust with team members and growing together. Technical Stack Database: MySQLMiddleware: Kubernetes, Docker, Nginx, Consul, RedisPlatform: AWS, On-premisesIaC: Terraform/Ansible Tools Repository Management: GitHubMonitoring: DataDog, RollbarCommunication: Slack, ZoomTicket Management: JIRA