Summary
✨ AI‑Generated
A full-time onsite engineering role serving as the technical foundation for advanced AI infrastructure and services. You will work across platform engineering, DevOps, and SRE, handling technical discovery, architecture integration, project delivery, AI tool deployment, GPU cluster monitoring, reliability, scalability, and infrastructure cost management.
Highlights
Work at the intersection of AI infrastructure, platform engineering, DevOps, and SRE. The role provides end-to-end ownership across system architecture, technical delivery, AI deployment, high-performance GPU reliability, monitoring, scalability, and infrastructure cost optimization.
Description
Location: On-site (Taipei, Taiwan)
Employment Type: Full-time
Department: Software Research & Development
About the Role
Are you passionate about bridging the gap between cutting-edge AI technologies and robust enterprise infrastructure? We are looking for a Platform / DevOps / SRE Engineer to serve as the technical bedrock for our AIOps infrastructure and services.
In this end-to-end role, you will span across Platform Engineering, SRE, and DevOps disciplines.
You will drive the entire system lifecycle—from initial requirements gathering and system architecture integration to AI tool deployment, GPU cluster monitoring, and FinOps cost logic implementation.
If you thrive on keeping high-performance GPU environments reliable, scalable, and cost-efficient, this is the role for you.
Key Responsibilities
1.
Platform & Project Integration
Requirements & Delivery: Lead technical discovery interviews for dev projects, define technical specifications, and manage delivery timelines and UAT/acceptance standards.Lifecycle Governance: Oversee end-to-end rollout, feature onboarding, and long-term platform maintenance coordination across cross-functional teams.
2.
AI Toolchain & GPU Infrastructure Operations
AI Tool Packaging & Deployment: Containerize and package AI management platforms (e.g., NVIDIA Run:ai, NVIDIA Dynamo, etc.) on Kubernetes (K8s) and Slurm clusters.Cluster Reliability & Observability: Monitor, maintain, and troubleshoot AI toolchains and GPU acceleration environments to deliver high availability for AI workloads.
3.
System Architecture & Integration Development
Spec Translation: Convert business requirements into system integration specs, testing frameworks, and architectural designs.Integration Operations: Conduct system integration testing, clear acceptance criteria, and ensure seamless system-to-system communication and failover handling.
4.
Application Development & Deployment
Full-Stack Engineering: Execute front-end/back-end development, API integration, deployment pipelines, and database architecture design.Quality Assurance & Maintenance: Draft and execute automated test cases, maintain database health, manage release updates, and perform diagnostic root-cause analysis.
5.
FinOps & Cost Optimization
Billing Models & Metering: Design resource allocation and billing logic, deploying tools to ingest and process Cloud metering datasets.Cost Observability: Implement real-time cost monitoring dashboards and resolve billing anomalies or operational issues within the FinOps pipeline.
Qualifications & Key Skills
Container & Orchestration Expertise: Deep experience with Kubernetes (K8s), container packaging, and infrastructure-as-code (IaC).AI/HPC Infrastructure Experience: Familiarity with HPC/AI scheduling tools such as Slurm, Run:ai, or similar GPU orchestration solutions.Full-Stack Development Capabilities: Hands-on proficiency in application development (Python, Go, Node.js, or similar), database schema design (SQL/NoSQL), and API integrations.SRE & Observability: Strong experience in operational monitoring (Prometheus, Grafana), logging, CI/CD, and systematic troubleshooting.FinOps / Metering Knowledge: Understanding of cloud metering, resource consumption tracking, and billing system implementation.Project Execution: Proven track record in requirement gathering, timeline management, and stakeholder management.
What We Offer
Tech Stack at Scale: Direct access to enterprise-grade AI infrastructure, high-density GPU clusters, and cutting-edge AIOps platforms.Impact & Ownership: High autonomy in defining architecture, FinOps strategies, and toolchain evolution.Growth & Culture: A collaborative environment blending DevOps efficiency, SRE rigor, and Platform Engineering innovation.