SRE (Infrastructure) Lead Engineer
Telexistenceinc — Japan · Posted ~4 hours ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
Role Overview
The SRE (Infrastructure) Lead Engineer owns the reliability, operability, and continuous improvement of the platform that powers our retail robotics products.
This is a hands-on technical leadership role for someone who can operate production systems directly, improve infrastructure through code, and guide a small Infra/SRE team toward stronger reliability, security, automation, and cost discipline.
Our current platform is primarily Azure-based, but this role does not require Azure-only experience.
We value strong hands-on cloud infrastructure experience on any major cloud platform, with the ability to learn the specifics of Azure, AKS, and our tooling quickly.
You will work closely with Backend, Frontend, Robotics, Security, Product, and Operations teams to keep our cloud and Kubernetes environments dependable for live store operations, robot and smart shelf workflows, telemetry processing, internal tooling, and customer-facing services.
Current Platform State
We operate a production platform supporting:
Retail robotics SaaS services and internal operations toolsKubernetes-based application workloads across development, staging, infrastructure, and production environmentsRobot, smart shelf, and store operations systems that depend on reliable cloud-to-edge communicationGitOps-style Kubernetes manifests and environment overlaysInfrastructure as Code using Azure Bicep, Terraform, and TerragruntCI/CD automation with GitHub Actions, including cloud authentication through OIDCObservability through Azure Monitor, Application Insights, Log Analytics, Grafana, Loki, Alloy, and alerting workflows
You will inherit systems that are already running and help evolve them into a more standardized, automated, and resilient platform.
Tech Stack
Cloud: Microsoft Azure today; AWS or GCP experience is also welcome if paired with strong cloud fundamentalsCompute & Runtime: AKS, Kubernetes, Docker, App Service, FunctionsInfrastructure as Code: Azure Bicep, Terraform, TerragruntGitOps & Manifests: Argo CD, Kustomize, Helm, Kubernetes YAMLNetworking & Ingress: VNet, subnets, NSG, load balancers, Traefik, ingress-nginx, cert-manager, TLSIdentity & Security: Microsoft Entra ID, Azure RBAC, workload identity, managed identities, Key Vault, GitHub Actions OIDCData & Storage: Azure Cosmos DB, PostgreSQL, Redis / Redis Enterprise, Blob Storage, Storage AccountsObservability: Azure Monitor, Application Insights, Log Analytics, Grafana, Loki, Alloy, Prometheus-compatible metrics, scheduled query alertsLanguages & Scripting: Bash, Python, TypeScript, C# or similar production scripting/application languagesDevelopment Tools: Git/GitHub, GitHub Actions, Azure CLI, kubectl, Helm, Argo CD CLI
Why This Role Matters
Production reliability: our store operations, robot workflows, and customer systems depend on stable infrastructureHands-on platform evolution: this role improves running systems through code, automation, standards, and direct operationCloud-edge complexity: our platform connects cloud services, Kubernetes workloads, store systems, robots, smart shelves, and telemetry flowsHigh business impact: reliability, deployment quality, and cost control directly affect operational efficiency and customer trustTeam leadership: you will shape the practices, rituals, and technical judgment of a small Infra/SRE teamStandardization opportunity: we are actively improving IaC ownership, environment separation, observability, incident response, and deployment guardrails
Key Responsibilities
Reliability & Operations
Own reliability practices for production cloud and Kubernetes systems, including SLOs, SLIs, error budgets, alert quality, and operational readinessLead incident response for infrastructure-related outages, act as Incident Commander when needed, and drive blameless post-incident reviewsBuild and maintain runbooks, dashboards, alerts, and operational tooling that help engineers diagnose and recover systems quicklyImprove backup, restore, disaster recovery, and business continuity practices for critical compute, data, storage, and configuration systemsIdentify recurring operational pain and remove it through automation, better design, or clearer ownership
Infrastructure & Platform Engineering
Design, operate, and improve Kubernetes environments, including AKS clusters, workload scheduling, autoscaling, ingress, TLS, RBAC, and cluster upgradesMaintain and evolve Infrastructure as Code using Bicep, Terraform, and Terragrunt across development, staging, infrastructure, and production environmentsOwn GitOps-style application deployment patterns using Argo CD, Kustomize, Helm, and environment overlaysImprove CI/CD pipelines with safe promotion flows, validation, drift detection, security checks, and rollback strategiesManage cloud networking, identity, secrets, certificates, and access controls in partnership with Security EngineeringSupport application teams with platform guidance for resource requests, scaling, dependency management, release safety, and production readiness
Observability, Security, and Cost
Standardize logs, metrics, traces, dashboards, and alerts across services and infrastructureOperate and improve observability tooling such as Azure Monitor, Application Insights, Log Analytics, Grafana, Loki, Alloy, and scheduled query alertsPartner with Security Engineering on RBAC, workload identity, managed identities, Key Vault usage, secret rotation, cloud access boundaries, and auditabilityDrive cloud cost visibility and optimization through tagging, right-sizing, autoscaling, log retention controls, budget review, and FinOps practicesKeep reliability, security, and cost decisions practical for a growing robotics business
Team Leadership
Lead and mentor 3-6 Infra/SRE/Platform engineers while staying deeply hands-onSet technical direction for infrastructure reliability, deployment automation, observability, and cloud operationsRun design reviews, code reviews, operational reviews, and regular improvement planningBuild healthy on-call practices, escalation paths, incident roles, and postmortem follow-throughCommunicate priorities, risks, and tradeoffs clearly to engineering leadership and cross-functional partners
Cross-Team Collaboration
Work with Backend and Frontend teams to make services easier to deploy, observe, scale, and operateWork with Robotics and Operations teams to understand how infrastructure behavior affects live robot and store workflowsWork with Product and Engineering Management to translate non-functional requirements into roadmap items and engineering guardrailsHelp teams adopt standard platform patterns without blocking delivery
Qualifications
Must Have
5+ years of professional infrastructure, platform, SRE, DevOps, backend infrastructure, or cloud engineering experience2+ years in a senior, lead, or technical leadership role with responsibility for production reliability or infrastructure directionStrong hands-on experience operating production systems on at least one major cloud platform such as Azure, AWS, or GCPProduction Kubernetes experience, including deployments, services, ingress, autoscaling, resource management, RBAC, troubleshooting, and upgradesHands-on Infrastructure as Code experience with Terraform, Bicep, CloudFormation, Pulumi, CDK, or similar toolsExperience designing or operating CI/CD pipelines and deployment workflows for production servicesStrong Linux, networking, DNS, TLS, and cloud identity fundamentalsPractical observability experience with metrics, logs, traces, dashboards, alerting, and incident responseExperience leading incidents, writing postmortems, and driving corrective actions to completionAbility to write scripts or small tools in Bash, Python, TypeScript, Go, C#, or a comparable languageClear communication skills and the ability to work across engineering, operations, security, and product teams
Nice to Have
Azure production experience, especially AKS, Azure Monitor, Application Insights, Log Analytics, Key Vault, Entra ID, managed identities, and Azure RBACExperience with Argo CD, Kustomize, Helm, cert-manager, Traefik, ingress-nginx, Grafana, Loki, Alloy, or Prometheus-style monitoringExperience with GitHub Actions OIDC, workload identity, or federated cloud authentication patternsExperience operating data services such as PostgreSQL, Redis, Cosmos DB, MongoDB-compatible databases, or cloud storage systemsExperience with drift detection, policy-as-code, cloud guardrails, or compliance-oriented infrastructure workflowsBackground in robotics, IoT, retail operations, edge computing, telemetry systems, or other cyber-physical production environmentsExperience building or improving an on-call program for a growing engineering organizationJapanese language skills are helpful but not required
Success Profile
Hands-on operator: comfortable debugging real incidents, reading manifests, reviewing Terraform/Bicep, and using cloud/Kubernetes CLIs directlyReliability-minded: thinks in SLOs, failure modes, blast radius, recovery paths, and operational feedback loopsPragmatic leader: balances engineering quality with business urgency and team capacityAutomation-oriented: turns repeated manual work into durable tools, workflows, or platform patternsSecurity-aware: treats access, secrets, identity, network boundaries, and auditability as part of daily infrastructure workCost-conscious: understands that cloud architecture must be reliable and financially sustainableCollaborative teacher: raises the operational maturity of surrounding teams through guidance, reviews, and shared standards
Vision & Growth
Build a stronger SRE and platform engineering practice for retail roboticsStandardize infrastructure ownership across cloud resources, Kubernetes clusters, manifests, CI/CD, observability, and incident responseHelp scale the platform from today’s operational needs toward enterprise-grade reliability across more stores, robots, smart shelves, and customer environmentsShape the long-term infrastructure roadmap while remaining close to the systems that keep the business running
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information.
These tools assist our recruitment team but do not replace human judgment.
Final hiring decisions are ultimately made by humans.
If you would like more information about how your data is processed, please contact us.
We have 123,872 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 123,872 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume