Infrastructure Engineer - Observability

Teksystems — Japan · Posted ~2 days ago

Hybrid Visa History ✓

Skills

Infrastructure engineering Observability Kubernetes Grafana Telemetry pipelines Monitoring Logging Tracing Cloud-native environments Automation Cloud Telemetry

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join a global cloud-focused engineering team responsible for building and operating scalable observability platforms. You will develop telemetry pipelines, manage shared infrastructure for large Kubernetes environments, automate operational processes, and improve monitoring, logging, tracing, and observability practices.

Highlights

International and casual work environment with two to three work-from-home days per week. The role focuses on scalable observability infrastructure, cloud-native platforms, automation, and reliable monitoring, logging, and tracing.

Description

International Work Environment: Collaborate with professionals from diverse backgrounds in a global setting.Hybrid work environment, Two to Three days work from home.Casual culture. Infrastructure Engineer (Observability) Location Tokyo, Japan (Hybrid/Remote depending on project requirements) Overview We are seeking an experienced Infrastructure Engineer (Observability) to join a cloud-focused engineering team responsible for building and operating scalable observability platforms. This role will focus on designing and improving telemetry pipelines, supporting engineering teams, and driving observability best practices across cloud-native environments. You will play a key role in managing shared observability infrastructure, developing automation, and ensuring reliable monitoring, logging, and tracing services for large-scale Kubernetes platforms. Responsibilities Develop, operate, and enhance shared observability infrastructure integrated with Kubernetes environments and Grafana Cloud.Design and maintain scalable telemetry collection and routing solutions for metrics, logs, and distributed traces.Build reusable tooling and automation to enable observability-as-code practices.Administer Grafana Cloud, including platform integrations, configuration, monitoring, and lifecycle management.Create and maintain technical documentation, runbooks, onboarding materials, and operational procedures.Provide technical guidance and troubleshooting support to engineering teams on observability-related challenges.Monitor, investigate, and respond to platform incidents and customer requests within established service levels.Collaborate closely with software engineers, platform teams, and stakeholders to improve reliability and operational efficiency. Required Skills & Experience 5+ years of experience in infrastructure, platform engineering, SRE, or observability-focused roles.Strong hands-on experience with Kubernetes in production environments.Deep expertise in observability platforms, monitoring systems, logging, and distributed tracing.Extensive experience administering and supporting Grafana/Grafana Cloud environments.Strong understanding of telemetry collection, routing, and data pipelines.Experience supporting multiple engineering teams through shared platform services.Excellent communication and stakeholder management skills.Proven ability to manage tasks independently while working effectively within a team. Preferred Skills Experience with AWS cloud services.Strong Linux systems administration background.Scripting experience with Bash.Experience building CI/CD workflows using GitHub Actions.Familiarity with infrastructure automation and self-service engineering platforms.Experience working in fast-paced technology organizations. Technology Stack KubernetesGrafana CloudAWSLinuxBashGitHub ActionsObservability & Telemetry Platforms Ideal Candidate Observability specialist rather than a general infrastructure engineer.Strong background supporting cloud-native platforms at scale.Comfortable working directly with engineering teams and customers.Collaborative, proactive, and service-oriented mindset.Passionate about improving reliability, developer experience, and operational excellence.