Senior DevOps Engineer

Beanz App — Jordan · Posted ~3 hours ago

Senior

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join a fast-growing digital platform as a Senior DevOps Engineer and take ownership of infrastructure from deployment pipelines through underlying systems. You will improve scalability, reliability, security, performance, and cost efficiency while building systems that remain resilient under real-world traffic and failures.

Highlights

Own infrastructure and reliability across a high-traffic platform, with a strong focus on scalability, security, resilience, performance, cost efficiency, and continuous engineering improvement.

Description

About us: Beanz is a leading coffee tech platform connecting coffee enthusiasts with their favorite specialty coffee shops. Present across the UAE, Oman, Qatar & Bahrain, we partner with over 3,000 coffee shops, providing a seamless pre-ordering, loyalty, and digital experience for customers while supporting local café growth. What we care about: At Beanz we care about the quality of what we ship and how the platform behaves once real traffic reaches it. We want our systems to be durable, to scale with demand, to stay secure, and to absorb a failure without taking the product down. Keeping them that way is steady daily engineering, not a project with an end date. So we are hiring a senior DevOps engineer who genuinely enjoys that work. Someone who likes owning infrastructure, wants to know how each layer works from the pipeline down to the disk, and gets satisfaction from making the platform faster, cheaper, and more reliable than it was last quarter. What follows is the day-to-day of the role. Responsibilities: Deployment pipelines. Run the CI builds, container images, and environment overlays that deploy our services, and keep git as the source of truth for what is running.Cluster operations. Size workloads, set requests and limits, tune autoscaling and disruption budgets, and manage node pool capacity across environments.Data infrastructure. Operate the self-hosted databases, caches, streaming systems and analytics stores our services depend on, including ingestion, retention and compaction.Content delivery and storage. Manage the CDN, object storage and asset delivery path, including caching rules, invalidation and TLS.Upgrades and migrations. Plan and run version upgrades for the open-source components we host, work through breaking changes, and keep a rollback path ready before you start.Backup and restore. Own backups, retention and restore drills for the self-hosted data stores. A backup nobody has restored does not count.Performance tuning. JVM heap and GC settings, storage layout, connection pools and cache sizing, since upstream defaults rarely match our load.Monitoring and alerting. Maintain metrics, dashboards, log pipelines and alert rules, and cut the alerts that nobody acts on.Cost and capacity. Track cloud spend per environment, right-size workloads, plan capacity ahead of growth, and weigh self-hosting against a managed service when the trade is worth revisiting.On-call rotation. Share the rotation with the rest of the platform team, cover your window, and hand over cleanly at the end of a shift.Incident response. Lead the response when the platform needs attention, write up the cause, and merge the fix. Expect to read release notes, issue trackers and upstream source when there is no support line to call.Security and patching. Handle RBAC, secret management, network policy and certificate renewal, and track CVEs in the open source images we build and run.Automation. Write the scripts and tooling that take manual steps out of the platform.AI agents in the workflow. Use modern AI agents to speed up the work, from reviewing configuration changes and pull requests to drafting runbooks, summarising logs and automating repetitive platform tasks. You will also help operate and extend the tooling that runs them. Qualifications & Skills: For each of the areas below, we are looking for real operating experience rather than passing familiarity: 6+ years of professional experience in DevOps or Platform Engineering, with a proven track record of building, operating, and maintaining highly scalable and highly resilient production infrastructure.Strong production experience with Kubernetes, including cluster operations, workload sizing, autoscaling, and deployment rollouts.Hands-on experience with AKS, Kubernetes, Helm, and Kustomize.Strong knowledge of major cloud platforms, with Azure preferred, including networking, identity and access management, storage, and cloud service limits.Hands-on experience with Azure, Azure Monitor, and Azure Storage services.Strong expertise in CI/CD and GitOps practices, including build pipelines, artifact management, environment promotion, rollback strategies, and maintaining Git as the source of truth.Hands-on experience with Argo CD, Jenkins, Docker, Terraform, and Git.Experience operating self-hosted data infrastructure, including managing upgrades, backups, data retention, compaction, and day-to-day operations of data stores.Practical experience with Apache Druid, Kafka, MongoDB, Redis, and Lakehouse technologies.Strong Linux and networking skills, including ingress, DNS, TLS, certificates, proxies, and load balancing.Hands-on experience with Linux, NGINX Ingress, cert-manager, DNS, and TLS.Strong understanding of monitoring and alerting, including designing metrics, dashboards, log pipelines, and alerting rules.Experience with Prometheus, Grafana, and Alertmanager.Strong scripting and automation capabilities, with the ability to build tools and automation rather than relying solely on YAML configuration.Proficiency in Bash, Python, TypeScript, kubectl, jq, and yq.Experience with content delivery and storage, including CDN configuration, object storage, caching rules, and cache invalidation.Hands-on experience with CDN, object storage, Blob Storage, and file shares.Nice to have: Experience with Apache Superset, JVM tuning for data services, FinOps/cost optimization with demonstrated cloud cost reductions, or contributions to the open-source technologies you operate.