DevOps / IT Infrastructure Engineer

Evollabs Tech — United Arab Emirates · Posted ~16 hours ago

Mid Full-time

Skills

DevOps IT infrastructure Network operating systems Data center infrastructure Telecommunications infrastructure Routing and switching Software engineering Network Operating Systems Routing Switching Telecom

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A DevOps and infrastructure engineer is sought to support critical networking and data-center systems. The role offers exposure to network operating systems, routing and switching infrastructure, telecom environments, and emerging AI infrastructure initiatives.

Highlights

Work on critical networking, data-center, and telecom infrastructure while contributing to ambitious infrastructure and AI initiatives in a rapidly developing technology environment.

Description

About Us We are a young high-tech company incorporated in the heart of one of the world's fastest-growing tech hubs—Dubai, UAE. As the exclusive software partner to one of the world's largest ODMs in the networking equipment space, we develop the Network Operating Systems that power critical data centre and telecom routing & switching infrastructure. Building on this foundation, we've recently launched an AI division focused on designing our own chips to accelerate inference and training workloads. What sets us apart is our unique position at the centre of a historic development: our ODM partner is establishing the first networking equipment factory of its kind in the GCC region, and we're the software engine driving this groundbreaking initiative. We're not just building technology—we're building a true networking vendor that serves regional interests while meeting the growing demand for networking equipment across the MENA region and farther. Our long-term vision extends beyond products to people: creating a thriving ecosystem for embedded systems and ASIC design talent that will produce generations of world-class professionals, establishing our region as a global centre of excellence for Enterprise Compute innovation. As a rapidly growing company at the forefront of AI hardware innovation, we are constantly seeking talented and motivated individuals to join our team. We offer a dynamic and challenging work environment, with opportunities to make a significant impact on the future of AI technology. Your Mission Own the end-to-end design and operation of our on-premise infrastructure for AI and enterprise workloads—built as code, automated, observable, and secure. You will architect and run Kubernetes clusters for training/inference, manage servers, networks, and core services, and enable developers with reliable CI/CD and platform tooling. This is where minutes, time-to-recovery and cost-per-job directly impact AI velocity at scale. Responsibilities Design and operate on-prem infrastructure as code: author reusable Terraform/Ansible/Helm modules; build GitOps workflows for repeatable, audited changes across environments.Build and run Kubernetes for AI: configure multi-tenant GPU clusters (MIG/GPUDirect RDMA, NVIDIA device plugins/DCGM), scheduling/quotas, and workload isolation.Administer servers, networks, and core services: OS lifecycle (Linux), identity/SSO (Keycloak/LDAP), secrets (Vault), DNS/DHCP/NTP, artifact registries, and internal package mirrors.Enable CI/CD: partner with developers to design fast, reproducible pipelines (GitLab CI), caching, and artifact provenance, and to deploy to GPU/CPU nodes.Observability & reliability: build metrics/logs/traces (Prometheus/Grafana), define SLOs/error budgets, incident response/on-call, post-mortems, and runbooks.Backup & disaster recovery: formulate RPO/RTO targets, implement cluster/app backups, test restores and site failover procedures regularly.Document clearly: architecture diagrams, playbooks, and self-service guides for engineers and stakeholders. You'll Collaborate With Platform and ML engineers running training/inference at scale, silicon and systems teams integrating hardware in the lab, security engineers safeguarding credentials and supply chain, application developers delivering services via CI/CD, and site ops supporting data centre deployments — together, we turn infrastructure into a product that accelerates the business. Must-Have Overall Experience: 5+ years in DevOps, SRE, or Platform Engineering with direct hands-on ownership and operation of physical, on-premises environments. Linux Systems Engineering: Advanced troubleshooting across kernel parameters, storage I/O, networking stacks (TCP/IP, VLANs, DNS), and performance tuning. On-Prem Hardware & Datacenter Operations: Hands-on experience managing physical server racks, network switches, patch cabling (Fiber, SFP+/QSFP, Cat6), and hardware lifecycle management. GPU Hardware & Firmware: Practical experience with server GPU installation, driver management (e.g., NVIDIA CUDA), firmware updates, thermal/power considerations, and hardware-level troubleshooting. Air-Gapped Environments & Access Management: Experience designing, operating, and securing air-gapped/isolated networks, enforcing strict access controls alongside FreeIPA (LDAP, Kerberos, SSSD, DNS, and RBAC). Securing inference engines and enforcing strict data residency controls to ensure sensitive corporate data, code, and prompts never leave on-prem or air-gapped boundaries.Observability & Monitoring: Production setup and maintenance of Prometheus and Grafana for metrics collection, system monitoring, dynamic dashboards, and alerting. Proxmox Virtualization: Hands-on experience configuring and operating Proxmox VE clusters, hypervisors, VM storage backends, and virtual networking. Containers & Orchestration: Production experience with Docker and Kubernetes (K8s/K3s) for containerizing workloads, managing pods, and maintaining cluster operations. Scripting & Core Automation: High proficiency in Python and Bash for writing custom administrative scripts, automation tools, and API integrations. CI/CD Pipelines: Experience building and maintaining CI/CD pipelines (GitLab CI, GitHub Actions, Jenkins, etc.) to automate code deployments and operational tasks. Developer AI & vLLM Integration: Experience configuring, supporting, and promoting Cursor/Claude/Copilot and AI-assisted coding tools for development teams. Token Observability & Cost Tracking: Experience setting up proxies, rate-limiting, telemetry, and cost/usage monitoring for LLM endpoints (e.g., via LiteLLM, OpenRouter, or custom API gateways). Agentic Workflows for DevOps: Experience designing, deploying, or utilizing autonomous agent systems (e.g., LangGraph, CrewAI, or custom Python agents) to automate operational workflows like log analysis, issue triage, or system alerts. Nice-to-Have Infrastructure as Code (IaC): Knowledge of Ansible, Terraform, or OpenTofu to help transition the team toward automated infrastructure provisioning in the future. Local LLM Infrastructure: Experience running self-hosted inference engines (vLLM, Ollama) on high-end HPC machines and setting up GPU passthrough (PCIe/vGPU) within Proxmox VMs or containers. Storage Systems: Enterprise experience configuring and managing Ceph, ZFS, NFS, or SAN/NAS storage topologies. Secrets Management: Experience with secrets management solutions (e.g., HashiCorp Vault) for safe handling of API keys, tokens, and system credentials in isolated setups. Internal AI Systems: Basic understanding of vector databases (Qdrant, Pgvector, Milvus) for building retrieval-augmented generation tools over internal infrastructure docs and build AI tooling securely for day to day task automation. What Success Really Means Reproducible environments by default: any engineer can spin up an identical dev/test stack (K8s namespace, storage, secrets, runners) from Git in ≤30 minutes, with audit trails for every change.Solid CI/CD for AI workflows: model/build/test pipelines are deterministic and cache-efficient; median pipeline time down 30–50%, with artifact provenance (SBOM, signatures) and traceable datasets/checkpoints.Lab-to-cluster continuity: hardware bring-up images, drivers, and firmware are versioned and promoted through the same pipelines; new boards/nodes join clusters with push-button automation.Actionable observability: dashboards and alerts reflect SLOs meaningful to researchers (throughput, time-to-first-token, I/O wait, GPU mem pressure);Cost & toil reduction: infra tasks automated to eliminate recurring manual work; fewer "custom one-offs," more reusable modules; quarterly infra spend per GPU hour trends down.Clear docs & self-service: engineers rely on concise runbooks and service catalogs; >80% of routine requests resolved via self-service workflows rather than ad-hoc ops support. Ready to accelerate the future of AI? Submit your application with links to relevant projects, repos, or runbooks. Show us how you design infrastructure as a product, automate everything you can, and keep complex systems reliable—so teams can ship faster with confidence. Join us in our mission to democratize AI compute — where your platform engineering turns ideas into production at scale.