Infrastructure Engineer / SRE

Teksystems — Japan · Posted ~1 day ago

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Description

Hybrid Work EnvironmentVisa sponsorship available for overseas candidatesInternational environment We are looking for a mid‑level Site Reliability Engineer who can balance software engineering and Kubernetes platform operations. This role focuses on running and improving production‑grade container platforms while building automation and tooling to support internal engineering teams. Responsibilities Operate and maintain production Kubernetes clusters and nodesRespond to incidents, alerts, and participate in on‑call rotationsPerform OS, middleware, and security‑related updatesExecute high‑risk production changes during off‑peak hoursDesign and develop automation and platform tools using GoCreate and follow strict operational procedures and runbooksSupport internal users and assist with platform and service migrations Required Skills & Experience 3+ years operating Kubernetes in production environments at scale3+ years software development experience with statically typed languages (Go, Java, C++, Rust, etc.)Solid understanding of Linux and basic networking (TCP/IP)Experience working in environments with strong operational processes and controlsWillingness to support late‑night or early‑morning operational work when requiredCKA certification (or ability to obtain within 3 months)English communication skills (business level) Nice to Have Strong proficiency in GoExperience automating large‑scale infrastructurePrivate and/or public cloud experienceMonitoring and observability tools (Prometheus, Grafana, ELK, etc.)Contributions to open‑source projects