Summary
✨ AI‑Generated
Join a platform engineering team as an intern and gain hands-on experience with production-grade distributed systems. You will work with Kubernetes, isolated container workloads, multi-cluster management, database replication, automated failover, high availability, and disaster recovery while helping build resilient cloud-native infrastructure.
Highlights
Hands-on internship experience with distributed systems, Kubernetes, high-availability infrastructure, disaster recovery, container isolation, and production-grade cloud-native technologies.
Description
Job Category: Infrastructure & Core Platform Engineering
Job Type: Full Time
Job Location: Nepal
About the Project: 01-Sandbox
01-Sandbox is a high-security, multi-tenant automated static analysis and code execution platform.
Workloads run inside isolated gVisor (runsc and Kata Container) sandbox worker clusters orchestrated across a distributed multi-region environment using Open Cluster Management (OCM).
We are implementing a resilient, zero-touch Multi-Cluster High Availability (HA) & Disaster Recovery (DR) control plane.
The infrastructure features an active-passive dual-hub topology (primaryhub and secondaryhub), real-time PostgreSQL (CloudNativePG) physical streaming replication, Valkey in-memory replication, and an intelligent Envoy Active-Passive Gateway that manages automated failover with zero split-brain and zero downtime.
What You Will Work On (Responsibilities)
As a Platform Engineering Intern, you will work hands-on with real distributed systems and live Kubernetes clusters.
You will:
Kubernetes Multi-Cluster Automation:Assist in modernizing platform scripts into 100% declarative Kubernetes controllers, Custom Resources, and Helm templates.Maintain and test Open Cluster Management (OCM) hub-to-spoke workload distribution (ManifestWork, ManagedCluster lifecycle, taints, and placement rules).Database Replication & Split-Brain Safe Failback:Monitor and validate CloudNativePG (CNPG) physical streaming replication, WAL timeline progression, and schema consistency.Implement declarative Kubernetes Job resources for delta synchronization across divergent cluster states.Gateway & Ingress Routing:Configure, test, and optimize our Envoy Active-Passive Proxy running on WireGuard overlay meshes (ports 80 and 6443).Write automated end-to-end chaos tests (killing primary control planes, simulating network partitions, and verifying sub-second failover).Sandboxed Workload Execution:Test and optimize security sandbox execution on worker spoke nodes using gVisor (runsc), Kata Container and Kubernetes Resource Pools.Observability & Operational Tooling:Build Grafana dashboards and Prometheus metrics for replication lag, proxy latency, and cluster failover triggers.Maintain clear architectural runbooks and automated validation scripts.
Technology Stack You Will Use
Orchestration: Kubernetes, KinD, Open Cluster Management (OCM), Helm 3Networking & Ingress: Envoy Proxy, Gateway API, WireGuard mesh, MetalLBData & Caching: PostgreSQL (CloudNativePG Operator), Valkey / Redis, RabbitMQWorkload Security: gVisor runtime (runsc), Containerd, kata containerOS & Virtualization: Linux, DockerAutomation & Scripting: Bash, Python, YAML, GitOps
Minimum Qualifications
Enrolled in or recently graduated with a degree in Computer Science, Software Engineering, Information Technology, or equivalent practical experience.Solid foundational understanding of Linux operating systems (systemd, process management, file permissions, basic networking like IP, DNS, routing).Familiarity with Docker and container fundamentals (building images, volume mounts, container networking).Basic working knowledge of Kubernetes core concepts (Pods, Deployments, Services, ConfigMaps, Namespaces).Proficiency in Bash scripting and at least one programming language (Python, Go, or JavaScript/TypeScript).Strong debugging mindset, curiosity to understand why distributed systems fail, and a methodical approach to documentation.
Nice-to-Have (Bonus Points)
Prior experience deploying applications with Helm.Hands-on exposure to database replication concepts (master-replica, WAL, failover).Familiarity with reverse proxies (Envoy, NGINX, or HAProxy).Experience working with multi-cluster Kubernetes or KVM virtual machines.
What You Will Learn & Career Growth
True Production Distributed Systems: Move beyond toy single-node projects to production-grade multi-cluster topologies.Disaster Recovery & High Availability: Master how enterprise platforms survive catastrophic outages with zero data loss.Cloud-Native Networking: Deeply understand Layer 4/Layer 7 traffic routing, mTLS, and overlay networks.Mentorship & Portfolio Building: Direct mentorship with senior engineers, working on an architecture that serves as an impressive portfolio piece for top-tier DevOps/SRE roles.