Principal DevOps Engineer

Landmark Group โ€” United Arab Emirates ยท Posted ~1 day ago

Lead Full-time

Skills

Cloud Architecture Kubernetes Docker Terraform Ansible CI/CD Linux GCP AWS Azure Infrastructure as Code SRE DevSecOps Prometheus

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary

A senior technical leadership role responsible for cloud infrastructure strategy, DevOps automation, reliability engineering, security practices, and scalable platform development. The position combines architecture ownership with hands-on engineering.

Highlights

Lead large-scale cloud transformation initiatives, define engineering standards, and build reliable automated platforms using modern cloud-native technologies.

Description

We are looking for a Principal DevOps Engineer to lead our cloud infrastructure, platform engineering, DevOps, SRE, MLOps, AIOps and DevSecOps initiatives. This is a hands-on leadership role responsible for building highly scalable, secure, resilient, and automated platforms that enable engineering teams to deliver software faster and more reliably. The ideal candidate combines deep technical expertise with strong architectural and operational leadership, driving platform modernization, automation, cloud transformation, and engineering excellence across the organization. Key Responsibilities Platform Engineering & Architecture โ— Define and drive the organization's DevOps, MLOps, AIOPs, SRE, Platform Engineering, and Infrastructure strategy. โ— Design highly available, scalable, secure, and resilient cloud-native architectures capable of supporting rapid business growth. โ— Build and maintain self-service platform capabilities that improve developer productivity and deployment velocity. โ— Collaborate with architects, engineering leaders, security teams, and product teams to establish engineering best practices and operational standards. โ— Lead cloud transformation programs, and platform optimization efforts. DevOps & Automation โ— Design and implement enterprise-grade CI/CD pipelines supporting multiple development teams and environments. โ— Establish Infrastructure as Code (IaC) standards using Terraform, Ansible, CloudFormation, and similar technologies. โ— Automate infrastructure provisioning, deployments, configuration management, and operational workflows. โ— Drive GitOps adoption and deployment automation across cloud and Kubernetes environments. โ— Standardize release management, environment management, and deployment governance processes. Site Reliability Engineering (SRE) โ— Define and implement Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets. โ— Drive reliability, performance, scalability, availability, and disaster recovery initiatives. โ— Lead incident management, root cause analysis, postmortems, and operational excellence programs. โ— Build proactive monitoring, alerting, observability, and capacity planning frameworks. โ— Optimize infrastructure performance, cost efficiency, and resource utilization across cloud environments. Cloud & Container Platforms โ— Design, build, and manage Kubernetes platforms across cloud environments. โ— Lead containerization initiatives using Kubernetes, Docker, EKS, GKE, ECS, and related technologies. โ— Architect multi-region, highly available cloud infrastructure on GCP, AWS, and Azure. โ— Implement cloud-native networking, service mesh, API gateways, and distributed systems best practices. โ— Establish backup, disaster recovery, business continuity, and resilience strategies. Security & DevSecOps โ— Embed security controls throughout the software delivery lifecycle. โ— Implement DevSecOps practices including vulnerability management, secrets management, policy-as-code, and infrastructure security. โ— Partner with security teams to ensure compliance, governance, and risk mitigation. โ— Drive cloud security best practices, identity and access management, encryption, and threat detection initiatives. Observability & Operations โ— Build centralized observability platforms covering metrics, logs, traces, events, and user experience monitoring. โ— Implement monitoring and logging solutions using Grafana, Prometheus, ELK, Dynatrace, Splunk, OpenTelemetry, and similar tools. โ— Lead production readiness reviews and operational health assessments. โ— Establish operational KPIs and continuously improve platform reliability and efficiency. Leadership & Mentorship โ— Provide technical leadership and mentorship to DevOps, SRE, Platform Engineering, and Infrastructure teams. โ— Define engineering standards, operational practices, and platform governance models. โ— Drive adoption of modern engineering practices across the organization. โ— Act as the technical escalation point for critical production and infrastructure challenges. Required Qualifications โ— Bachelor's or master's degree in computer science, Information Technology, Software Engineering, or related discipline. โ— 10+ years of experience in DevOps, MLOps, AIOps, Site Reliability Engineering, Infrastructure Engineering, Cloud Engineering, or Platform Engineering. โ— Proven experience designing and operating large-scale, mission-critical production platforms. โ— Deep expertise in cloud platforms, with strong hands-on experience in Google Cloud Platform (GCP). Experience with AWS and Azure is desirable. โ— Strong experience with Kubernetes, Docker, container orchestration, and microservices architectures. โ— Extensive experience implementing Infrastructure as Code using Terraform, Ansible, CloudFormation, or similar tools. โ— Strong experience building and managing enterprise CI/CD platforms. โ— Expert-level Linux administration and troubleshooting skills. โ— Experience managing distributed systems, event-driven architectures, and high-volume production workloads. โ— Strong understanding of networking, load balancing, DNS, CDN, security, and distributed computing concepts. โ— Excellent communication, stakeholder management, and leadership skills. Preferred Qualifications โ— Experience building Internal Developer Platforms (IDP) and self-service engineering platforms. โ— Experience implementing SRE frameworks including SLOs, SLIs, error budgets, and reliability engineering practices. โ— Experience with multi-cloud and hybrid-cloud architectures. โ— Experience supporting AI/ML platforms, MLOps, data platforms, and large-scale analytics workloads. โ— Experience implementing service mesh technologies such as Istio or Linkerd. โ— Relevant certifications such as Google Professional Cloud Architect, Google Professional DevOps Engineer, AWS Solutions Architect Professional, AWS DevOps Engineer Professional, Certified Kubernetes Administrator (CKA), or Certified Kubernetes Security Specialist (CKS). What You'll Gain โ— Opportunity to define and lead the DevOps, SRE, and Platform Engineering vision for a rapidly growing technology organization. โ— Ownership of cloud infrastructure supporting large-scale e-commerce, logistics, supply chain and enterprise applications. โ— Exposure to modern cloud-native, AI, platform engineering, and reliability engineering practices. โ— Ability to influence architecture, engineering culture, and operational excellence across the organization. โ— Work alongside senior technology leaders on strategic transformation initiatives.