Backend Developer
Claven — Uzbekistan · Posted ~1 day ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
Claven AI (https://claven.ai/) is a Tashkent based startup building a Sovereign AI Infrastructure & Orchestration Platform — a system that turns bare-metal or managed GPU clusters into cloud-native, efficient, and profitable AI clouds.
We make it easy to run inference, fine-tune models, and manage GPU workloads at scale, with a strong focus on data sovereignty for all customers in Uzbekistan and around the globe.
We target enterprise and government customers in regulated industries — banking, finance, healthcare, defense, academia — and already have pilot projects in the pipeline.
Our edge is performance, cost efficiency, modern architecture, and a simplified integration experience across both LLM and traditional ML workloads.
We work in a transparent, low-ego, technology-first way, with a culture of hard work, camaraderie, and customer obsession.
How We Work
We are an AI-native team.
We use Claude Code daily, alongside other agentic AI tools and workflows to ship faster.
This is not a perk, it is how the work happens here.
If you join us, we expect that:
You actively use Claude Code, Cursor, or similar agentic AI in your daily workflow — not as a curiosity, but as your default way of working.
When handed an unfamiliar system (GPU scheduling internals, a model-serving engine, a new operator or controller), you go deep on it with AI as your accelerator and come back with answers, prototypes, and opinions.
You take ownership.
You don't wait to be told what to do once a problem is in your lane.
You maintain curiosity and learning velocity.
What You'll Own
You will own the backend services and Kubernetes machinery that turn a GPU cluster into a multi-tenant AI platform.
This is real distributed-systems and Kubernetes-API work, not CRUD.
Specifically:
Build and evolve our core Go services that orchestrate AI/ML workloads on Kubernetes — routing inference traffic, provisioning and reconciling workloads through the Kubernetes API, and managing their lifecycle.
Write controllers, operators, admission webhooks, and custom resources (controller-runtime / kubebuilder), including GPU-aware scheduling and placement.
Integrate and tune the model-serving stack for cost, latency, and throughput, including multi-GPU (tensor-parallel) serving and the model artifact lifecycle.
Implement multi-tenancy, network policy, and workload isolation in support of our data-sovereignty and air-gapped operating modes — per-tenant isolation, quotas, and RBAC.
Own the API contracts the platform exposes, and partner with our frontend engineers on the seams where the backend meets the UI, including the gateway and identity layers.
Help diagnose and resolve production issues across the stack as we onboard pilot customers.
What We're Looking For
We care more about what you can actually do than how many years it took you to get there.
The bar is:
Strong, production-grade Go — you've built and operated real services, not just scripts.
Real Kubernetes depth: you understand the API and controller/reconciliation model, and you've operated clusters — debugging incidents, not just kubectl apply-ing to a cluster someone else runs.
Strong Linux, networking, and systems fundamentals.
Hands-on with infrastructure-as-code, observability tooling, and CI/CD pipelines.
Comfortable owning an API contract and the boundary it sits on; you write things down and communicate clearly.
Strong written English and clear, low-ego communication.
4+ years in backend, platform, infrastructure, or SRE roles is a useful soft floor, not a hard gate.
Bonus Points
None of these are required.
They will accelerate your ramp-up, and we are happy to hire someone with strong fundamentals and zero GPU experience who is excited to learn.
Writing Kubernetes controllers, operators, admission webhooks, or custom resources (controller-runtime / kubebuilder).
Hands-on with the NVIDIA stack (CUDA, MIG, NCCL, DCGM) or GPU partitioning and virtualization (MIG, time-slicing, MPS).
Serving LLMs in production (vLLM, TGI, SGLang, or similar).
Multi-tenancy, network policy, or compliance work in regulated industries.
Python — some of our services are Python; comfort crossing over helps.
API-gateway or identity/OIDC work.
Background in ML platforms, developer infrastructure, or internal PaaS-style products; familiarity with usage-based metering and quota systems.
What You Get
Ownership of a critical layer of the platform from day one.
Direct work with the founders and a small, AI-native engineering team — no bureaucracy, no political layers.
A real customer pipeline in regulated markets — pilots, not vaporware.
Fully remote, globally.
We hire for talent and attitude.
Working hours should overlap meaningfully with the team.
Compensation is discussed during the hiring process and is competitive for the candidate's market.
We have 93,444 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 93,444 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume