Description
Introduction
We are seeking an AI System Engineer to support hands-on implementation, validation, monitoring, and day-to-day technical operations for a large-scale AI Data Center platform.
The role combines NVIDIA GPU infrastructure, virtualization, Kubernetes, monitoring, basic network and storage integration, and structured customer-facing technical support.
The engineer will execute approved designs, test cases, change procedures, and runbooks, and will work with senior architects and platform specialists when issues require deeper escalation.
The role is suitable for an infrastructure engineer with practical GPU, VMware, Kubernetes, monitoring, or cloud-native deployment experience who is ready to expand into AI data center platform operations.
Equivalent experience with comparable technologies may be considered for selected platform tools.
Your Role And Responsibilities
AI / GPU Infrastructure Deployment and Operations
Support installation, configuration, validation, and operational checks for NVIDIA GPU servers and GPU-enabled infrastructure.
Perform GPU node health checks, diagnostic collection, basic topology validation, and performance or functional testing.
Support NVIDIA vGPU or physical GPU environments, including driver and runtime validation under approved procedures.
Assist with server, OOB/BMC, operating system, certificate, DNS, NTP, firewall, and end-to-end connectivity checks.
Collect system, GPU, platform, and application logs and prepare diagnostic evidence for escalation to senior engineers or vendors.
Virtualization, Kubernetes, and Platform Operations
Support VMware-based infrastructure and cloud-native platform deployment, testing, and post-deployment troubleshooting.
Deploy and maintain Kubernetes workloads using YAML and Helm, including namespaces, resource limits, services, and basic GPU scheduling.
Support cluster onboarding, tenant or project setup, user access, quotas, and resource pools through Rafay or comparable Kubernetes management platforms.
Execute approved provisioning, upgrade, retry, recovery, rollback, and configuration-validation procedures.
Assist with platform integration testing across infrastructure, Kubernetes, monitoring, and customer applications.
Network and Storage Integration
Perform first-level TCP/IP, VLAN, routing, DNS, NTP, firewall, and service-connectivity validation using documented procedures.
Support approved tenant network, IP addressing, VLAN/VRF, route, and segmentation changes through Netris or comparable network automation tools.
Assist with storage client installation, mounts, CSI integration, NFS/S3 access, and capacity or latency checks using Weka or comparable storage platforms.
Validate tenant isolation, network reachability, storage connectivity, and change results; escalate advanced fabric or storage issues when required.
Monitoring, Troubleshooting, Incidents, and Changes
Implement and operate monitoring agents, exporters, dashboards, and log collection using Prometheus, Grafana, or comparable tools.
Perform routine GPU, system, Kubernetes, network, and storage health checks and identify abnormal conditions.
Conduct first-level alert triage, problem reproduction, log analysis, ticket updates, troubleshooting, and escalation.
Execute approved standard changes, updates, service restarts, certificate renewals, validation, and rollback during scheduled maintenance windows.
Support incident, request, change, tenant onboarding, offboarding, and decommissioning activities with complete operational records.
Testing, Documentation, and Customer Support
Execute functional, integration, regression, basic performance, recovery, and end-to-end test cases.
Maintain configuration records, test evidence, incident notes, tickets, SOPs, runbooks, and known-error documentation.
Prepare clear technical findings, status updates, and handover materials for internal teams, customers, and partners.
Support technical walkthroughs, operational guidance, and user or partner enablement sessions when required.
Preferred Education
Master's Degree
Required Technical And Professional Expertise
Bachelor’s degree or higher in Computer Science, Information Technology, Engineering, or a related field, or equivalent practical experience.
Approximately 2-5 years of hands-on experience in system engineering, infrastructure support, virtualization, cloud-native platforms, data center operations, or a related technical role.
Hands-on experience with NVIDIA GPU or vGPU infrastructure, GPU-enabled servers, or comparable accelerator environments.
Practical experience with Kubernetes and containerized infrastructure; working knowledge of Helm and YAML is expected.
Experience with VMware virtualization or a comparable enterprise virtualization platform.
Experience using monitoring and troubleshooting tools such as Prometheus, Grafana, or comparable observability platforms.
Basic Linux administration capability and the ability to interpret operating system, platform, and application logs.
Working knowledge of TCP/IP and common infrastructure services such as VLANs, routing, DNS, NTP, and firewalls.
Ability to follow runbooks, test cases, change tickets, and maintenance procedures accurately and to document results clearly.
Strong customer-facing communication and technical coordination skills in project-driven environments.
Full professional proficiency in Chinese for daily communication, issue handling, ticket updates, and operational documentation.
Basic to intermediate English proficiency for technical manuals, product interfaces, vendor documentation, and support cases.
Ability to work primarily on-site in Taiwan and support scheduled maintenance windows and occasional after-hours activities.
Preferred Technical And Professional Experience
Experience managing NVIDIA driver, firmware, CUDA, NCCL, GPU telemetry, or GPU performance benchmarking.
Experience with Rafay, Netris, Weka, or comparable Kubernetes management, network automation, IPAM/SDN, or high-performance storage platforms.
Exposure to NVLink, InfiniBand, Spectrum-X, RoCE, high-speed Ethernet, or other AI cluster interconnect technologies.
Experience with enterprise storage, CSI, NFS, S3, backup/restore, snapshots, or storage performance monitoring.
Experience supporting production environments using incident, problem, change, or ITSM processes, including NOC/SRE-style operations.
Scripting or automation experience using Python, Shell, Ansible, APIs, or similar tools.
Relevant certifications such as CKA, CKAD, NVIDIA AI Infrastructure/Operations, Linux, Kubernetes, networking, or storage certifications.