Summary
β¨ AIβGenerated
A senior infrastructure role focused on building, deploying, securing, and operating a cloud platform that supports real-time conversational workloads. You will serve as the most senior infrastructure voice on the engineering team, establish standards for backend, machine-learning, and integration teams, and shape architecture across AWS and Google Cloud with reliability and low-latency operation as core priorities.
Highlights
Take ownership of infrastructure standards and operations for a platform supporting live conversational workloads across multiple industries. The role provides senior-level architectural ownership across AWS and Google Cloud, with strong focus on security, availability, latency, and production reliability.
Description
Hamsa builds AI models that master Arabic dialects.
Our platform covers speech-to-text, text-to-speech, translation and spoken language understanding, and powers voice agents and phone agents deployed by enterprises across MENA β in banking, telecom, government services, call centres, healthcare and hospitality.
That means our infrastructure carries live human conversations.
When a phone agent is mid-call with a customer in Emirati or Levantine dialect, there is no retry, no queue, no "please refresh." Latency, availability and audio quality are product features, not ops metrics.
If that sounds like an interesting engineering problem rather than a stressful one, keep reading.
About the role
We're hiring a Senior DevOps Engineer to own how our platform is built, deployed, secured and operated across AWS and Google Cloud.
You'll be the most senior infrastructure voice on the team β setting the standards that backend, ML and integration engineers build on, and making the architectural calls rather than executing someone else's.
You'll work closely with our ML team on model serving and GPU capacity, with our voice team on real-time media and telephony infrastructure, and with our enterprise clients' technical teams on private cloud and self-hosted deployments.
What you'll do
Own the architecture of our multi-cloud estate across AWS (EKS, VPC, IAM, CloudWatch, GuardDuty) and Google Cloud (GKE, Cloud Run, Vertex AI, Artifact Registry, GPU compute)Design and maintain the Terraform codebase so every environment β including client-dedicated ones β is reproducible and reviewableRun production Kubernetes: cluster architecture, upgrades, node pools, autoscaling for spiky voice traffic, network policies, RBAC and tenant isolationBuild and operate the serving infrastructure for our STT and TTS models, including GPU node management, autoscaling and cost controlSupport the real-time voice stack β media servers, SIP and VoIP integrations, and the network path that keeps call quality highDesign CI/CD and GitOps workflows with proper quality gates and safe rollout strategies: canary, blue/green, fast rollbackOwn observability end to end β metrics, logs, traces and alerting that actually help during a live incident, including call-quality and latency signalsLead incident response for infrastructure issues and write the post-mortem afterwardsPackage and support private cloud and on-premise deployments for enterprise and government clients, including the security documentation those engagements requireDrive cloud cost discipline across compute, GPU and egressImplement security controls across IAM, network segmentation, secrets management and audit logging, and support client security reviews and compliance workMentor mid-level engineers, review infrastructure changes, and document what you build
What we're looking for
5+ years in DevOps, Cloud, Platform or Site Reliability Engineering, with at least 3 years on production systems you were responsible forDeep hands-on experience with AWS or Google Cloud, and working competence in the other β we run both3+ years operating Kubernetes in production, not just deploying to it3+ years with Terraform, including module design and state management across environmentsStrong Linux administration, networking (VPC, DNS, TLS, load balancing) and production troubleshooting under pressureProficiency in Python or Go, plus solid shell scriptingExperience designing CI/CD pipelines for containerised microservicesWorking knowledge of Prometheus, Grafana, or equivalent observability toolingComfortable working in Arabic and English
Nice to have
GPU workloads or ML model serving β Vertex AI, SageMaker, Triton, vLLM, or self-hosted inferenceReal-time or low-latency systems: WebRTC, LiveKit, media servers, streaming, or VoIP and SIPExperience delivering on-premise or air-gapped deployments to enterprise or government clientsExposure to ISO 27001, SOC 2, or regional data residency requirementsCertifications: AWS Solutions Architect or DevOps Engineer Professional, Google Professional Cloud DevOps Engineer, CKA, HashiCorp Terraform Associate
Why Hamsa
Arabic voice AI is being built now, and you'd be building the infrastructure underneath it rather than integrating someone else's APIReal ownership β this is a role where your architectural decisions stickEnterprise-grade problems at startup speed, with clients whose systems people actually depend on