Senior Platform Engineer

Robert Walters — United Arab Emirates · Posted ~2 hours ago

Senior Full-time Visa History ✓

Skills

Microsoft Azure Kubernetes Terraform Helm cloud infrastructure Azure Prometheus Grafana

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology organization is seeking a Senior Platform Engineer to build and operate cloud-native platforms. Responsibilities include Kubernetes management, infrastructure automation, monitoring, security, disaster recovery, and production support.

Highlights

Advanced cloud engineering role focused on scalable infrastructure, automation, security, observability, and resilient production environments.

Description

The Senior Platform Engineer will design, implement, and operate cloud-native platforms on Microsoft Azure.The role focuses on Kubernetes, cloud infrastructure, automation, observability, security, and Disaster Recovery, ensuring highly available, secure, and resilient production environments.Key Responsibilities: Design, deploy, and manage Azure infrastructure and AKS clusters.Implement and maintain Disaster Recovery, backup, and business continuity solutions.Automate infrastructure and deployments using Terraform, Helm, and Azure DevOps.Manage Kubernetes networking, ingress, storage, and security.Deploy and maintain observability platforms including Prometheus, Grafana, and Loki.Manage TLS certificates, secrets, and platform security.Support production environments, troubleshoot critical incidents, and drive root cause analysis.Plan and execute platform migrations and infrastructure upgrades.Create and maintain technical documentation, architecture diagrams, HLD/LLD, SOPs, operational runbooks, troubleshooting guides, and Disaster Recovery procedures.Define and maintain platform engineering standards, policies, governance frameworks, and best practices across Azure and Kubernetes environments.Establish cloud and Kubernetes governance covering RBAC, naming and tagging standards, resource organization, security controls, namespaces, resource limits, ingress, secrets, and storage.Define Infrastructure-as-Code and CI/CD governance, including reusable Terraform modules, code review standards, state management, pipeline controls, approval gates, environment promotion, and artifact/version management.Participate in architecture and technical design reviews for new platforms, applications, integrations, and infrastructure changes.Drive platform security and compliance readiness through security baselines, vulnerability remediation, access reviews, secrets/certificate management, audit controls, and policy enforcement.Define and track platform availability, SLIs/SLOs, capacity, performance, and operational health.Drive incident and problem management practices, including root cause analysis, corrective actions, and prevention of recurring incidents.Perform capacity planning, performance optimization, and cloud cost optimization across platform infrastructure.Own DR testing, RTO/RPO validation, backup/restore standards, recovery procedures, and evidence from periodic recovery exercises.Evaluate new platform technologies, conduct POCs, and establish approved patterns before production adoption.Provide technical leadership, knowledge sharing, and mentoring to engineers on Azure, Kubernetes, Terraform, CI/CD, security, and platform operations. Required Skills: Microsoft AzureKubernetes (AKS/OpenShift)Docker & HelTerraformAzure DevOps / CI/CDPrometheus, Grafana, LokiAzure Networking (VNets, NSGs, Private Endpoints, Firewall)Linux & Bash scriptingDisaster Recovery, Backup & Restore strategiesExperience with PostgreSQL, MongoDB, MySQL, or Azure SQLCloud & Platform GovernanceKubernetes Security, Governance & Policy EnforcementAzure Policy, RBAC & Security ControlsInfrastructure-as-Code standards and reusable Terraform patternsCI/CD governance, release controls, and environment promotion strategiesTechnical documentation (HLD, LLD, SOPs, runbooks, and architecture diagrams)Architecture design and technical design reviewsIncident, Problem & Root Cause Analysis managementCapacity planning, performance optimization & FinOps / cloud cost optimizationSecurity, compliance, audit controls & operational governanceRTO/RPO planning, DR testing, backup and recovery governanceGit / GitOps practices and source control standardsTechnical leadership, mentoring, and cross-functional collaboration Preferred: Banking or regulated industry experienceStrong understanding of High Availability, Disaster Recovery, and production operationsExperience defining enterprise platform standards, policies, and governance frameworks.Experience working in security- and compliance-controlled environments.Experience leading technical design reviews, platform modernization, and cloud transformation initiatives.