Summary
✨ AI‑Generated
A hands-on DevOps role focused on building and operating sophisticated on-premise AI infrastructure. You will manage GPU clusters and worker nodes, maintain AI inference environments, deploy containerized systems across segmented client environments, build CI/CD pipelines, and oversee databases, storage, identity, networking, monitoring, and backups. The position suits an engineer who enjoys broad ownership, infrastructure experimentation, and solving complex deployment challenges.
Highlights
Work with modern AI infrastructure, including GPU clusters and LLM inference systems, while having broad ownership across deployment, CI/CD, networking, security, monitoring, and databases. The role offers substantial room for experimentation and hands-on technical responsibility.
Description
Tentang val.id val.id adalah software & AI consultancy yang membangun platform AI/OCR on-premise, document management system, dan sistem enterprise untuk klien pemerintah maupun swasta.
Tim kecil, proyek banyak, dan ruang eksperimen luas.
Yang akan kamu kerjakan
Mengelola infrastruktur AI on-premise: GPU cluster (2–8 server, termasuk GPU kelas B200), worker nodes, dan database nodes.
Mulai dari provisioning, NVIDIA driver & CUDA management, sampai monitoringMengoperasikan inference stack untuk workload AI: vLLM untuk LLM serving dan container OCR on-premDeployment berbasis Docker & Docker Compose ke environment klien, termasuk on-prem dengan arsitektur multi-zone (DMZ, internal, control plane)Membangun dan merawat CI/CD pipeline untuk beberapa proyek yang jalan paralelMengelola layanan pendukung: PostgreSQL, S3-compatible object storage, SSO/OIDC (Casdoor), reverse proxyNetworking: segmentasi jaringan, firewall, VPN antar environmentMonitoring, backup, dan membangun fondasi ke arah container orchestration (Kubernetes) serta high availability
Yang kami cari
Linux system administration yang kuatDocker & Docker Compose sebagai pekerjaan sehari-hariNetworking level menengah: subnetting, firewall, VPNTerbiasa scripting Bash dan/atau PythonPostgreSQL operasional: backup, restore, tuning dasar
Nilai plus
KubernetesPengalaman GPU server / AI workload: NVIDIA driver, CUDA, inference serving (vLLM atau sejenisnya)Pengalaman deployment on-prem di lingkungan enterprise atau pemerintahS3-compatible object storageAI agent expertise
Kenapa val.id
Resource untuk R&D dan eksperimen hampir selalu disetujuiResult oriented, tidak ada micro-management, bila task selesai boleh beristirahatAI-assisted workflow sebagai standar kerja
Kompensasi
Kompetitif untuk level junior–mid, ditentukan dari hasil technical interview.