IT1_Linux & DevOpsインフラエンジニア/Linux & DevOps Infra Engineer
Astroscale — Japan · Posted ~21 hours ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
私たちのリアルな様子はこちらから→ 株式会社アストロスケールの会社情報 - Wantedly
Export Control Laws:
Unless explicitly notified otherwise, our vacancies are covered by Export Control Laws which require candidates to be from an "Export Safe" Country as deemed by the Japanese Government.
The countries are as follows:
Japan, Germany, Australia, Argentina, Italy, USA, France, Netherlands, UK, Austria, Ireland, Czech, Spain, Greece Canada, New Zealand, Belgium, Bulgaria, Sweden, Switzerland, Norway, Finland, Luxembourg, Portugal, Denmark Hungary and Poland.
About The Role
Astroscale is seeking a specialist-level Platform Engineer to design, deploy, and maintain our core Linux/Ubuntu infrastructure, containerized workloads, and internal platform services.
This role blends hands-on DevOps and SRE responsibilities, with ownership over AWS EC2 containers, on-premises Proxmox virtualization, GitLab CI/CD, BookStack, and LiteLLM.
You will operate autonomously, drive automation, ensure platform reliability, and grow toward a Lead DevOps Engineer role as the team and mission scale.
Key Responsibilities
Platform & Virtualization
Deploy, configure, and maintain Ubuntu Linux servers across AWS EC2 and on-premises Proxmox VE environments.
Manage containerized workloads using Docker and Docker Compose; maintain GitLab, BookStack documentation platform, and development VM fleets.
Deploy, configure, and troubleshoot LiteLLM, including LLM routing/proxy infrastructure, and support AI/ML inference environments.
DevOps & Automation
Build and optimize CI/CD pipelines, infrastructure automation, and GitOps workflows using Terraform, Ansible, and Git.
Implement configuration management, provisioning scripts, and version-controlled infrastructure standards.
SRE & Reliability
Establish and maintain monitoring, logging, and alerting for platform health and performance.
Lead incident response, root cause analysis, and post-incident reviews; maintain runbooks and change documentation.
Perform system patching, security hardening, backup/restore, and performance tuning.
Participate in planned weekend or after-hours maintenance windows as required.
Essential Qualifications & Experience
5+ years in IT infrastructure/systems administration, with 3+ years focused on Linux (Ubuntu) production environments.
Proficient in Linux CLI: system management, process control (systemd, cron, ps, top), package management, and log analysis.
Hands-on experience with Docker and Docker Compose in staging or production environments.
Practical knowledge of Infrastructure as Code: Terraform for provisioning, Ansible for configuration management, and Git for version control.
Experience with CI/CD platforms; GitLab CI/CD preferred, with Jenkins, GitHub Actions, or similar platforms also acceptable.
Solid understanding of AWS core services, including EC2, VPC, IAM, S3, security groups, and basic networking.
Experience with monitoring/logging tools such as Prometheus, Grafana, ELK, Datadog, or equivalent, and incident response workflows.
Strong documentation habits, including runbooks, SOPs, change logs, and architecture diagrams.
Language: Professional proficiency in either English or Japanese, written and verbal.
Must also possess at least basic conversational skills in the other language, with a demonstrated passion and commitment to actively learning it.
Availability for planned weekend or late-night maintenance windows as needed.
Desirable Qualifications & Experience
Interest or practical experience in leveraging AI/LLM tools to accelerate and enhance Infrastructure as Code (IaC) workflows, such as AI-assisted Terraform, Ansible, or CI/CD generation.
Bilingual proficiency in English and Japanese, written and verbal.
Experience with Kubernetes, Helm, or container orchestration platforms.
Practical understanding of SRE practices, including SLIs/SLOs, error budgets, capacity planning, and reliability metrics.
Advanced Proxmox VE administration, including ZFS, clustering, HA, and backup/restore strategies.
Experience supporting AI/LLM infrastructure, model serving, or inference proxies.
Leadership, mentoring, or team coordination experience.
AWS Certified Cloud Practitioner or SysOps/DevOps Professional certification.
We have 98,729 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 98,729 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume