Lead Cloud Infrastructure Engineer

Atomdn โ€” United Arab Emirates ยท Posted ~2 weeks ago

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Description

INTRODUCTION ATOM is the end-to-end insurance operating system, enabling any entity within the insurance lifecycle to operate the entirety of their activities and cooperate/interact with third parties. Its unique advantage is to incorporate not only insurance underwriting and claims handling, but also all associated support and corporate activities. ATOM Technologies is the team behind the platform. Based in the DIFC Innovation Hub, ATOM Technologies is a group of more than 40 passionate and dedicated insurance and software professionals on a mission to upgrade the insurance industry from the inside out. JOB PURPOSE Own the architecture, design, operation and continuous improvement of the organisation's AWS cloud infrastructure, data platforms and resilience capabilities. The role sets the target-state cloud infrastructure architecture and provides deep, hands-on AWS engineering across the full stack governance, networking, compute, storage, security and cost and is the primary owner of disaster recovery, regional failover and business continuity for cloud and data platforms, including the Atom 3 underwriting platform. On the hard-infrastructure side - networking, compute, storage, core AWS, security and resilience the role operates in parallel with the Senior Infrastructure Manager as a co-equal peer, jointly owning the domain to give the function depth, shared load and mutual resilience of cover. It also overlaps with the SAP Basis Lead to provide continuity and additional engineering capacity as platform demand grows. ROLE SCOPE Cloud infrastructure architecture target-state design, patterns, standards and the AWS architecture roadmap.AWS infrastructure account governance, identity, networking, compute, storage, security and change management.Disaster recovery, regional failover and business continuity across cloud and data platforms, including in-house and SaaS platforms.Cloud data platforms and data architecture - Snowflake, databases, data lifecycle, classification and retention.Cloud cost optimisation and performance engineering.Resilience engineering and audit evidence / compliance support. SPECIFIC DUTIES AND RESPONSIBILITIES Cloud Infrastructure Architecture Own the target-state cloud infrastructure architecture reference designs, patterns and standards and the AWS architecture roadmap.Lead architecture and Well-Architected reviews and drive architectural decisions across networking, compute, storage, security and resilience.Own and implement the configuration management database (CMDB), maintaining accurate configuration items and dependency mapping across cloud infrastructure to support change, incident and resilience management. AWS Infrastructure & Governance Operate and govern AWS accounts, AWS Organizations structure and Service Control Policies.Own IAM policy design, role and access provisioning, SSO / Entra ID integration, privileged access and joiner/mover/leaver access changes.Design and manage networking - VPCs, subnets, security groups, routing, VPN / Direct Connect, DNS, load balancers and Transit Gateway.Provision and optimise compute - EC2, Auto Scaling, ECS/EKS, AMI management and OS patching.Manage storage and databases - S3 lifecycle and encryption, EBS, RDS provisioning, backup and replication.Author and maintain Infrastructure as Code (Terraform / CloudFormation); run change control and environment provisioning. Disaster Recovery, Failover & Business Continuity Own DR strategy and architecture for AWS and data platforms.Design and operate cross-region replication and regional failover, with particular focus on the Atom 3 underwriting platform.Define, validate and evidence RPO/RTO; plan and run DR testing and failover exercises.Create and maintain DR runbooks and business continuity plans for cloud and data services.Design and activate Snowflake multi-cloud / multi-region DR and failover a capability the team currently cannot deliver and which is dependent on this hire. Cloud Data Platforms Provide the infrastructure, security - configuration and resilience layer for Snowflake account and warehouse configuration, network policies, MFA and SSO integration. Platform ownership and cost control remain with the data/Atom team (Head of IT receives monthly reporting).Support the access model: access is granted via AD group membership (no direct Snowflake user creation, with AD deprovisioning auto-revoking access); RBAC is administered by the Systems Administrator through the approval/ticketing process, with roles created by the data team.Administer database servers, data networking, storage and capacity planning; provision Dev/Test/Prod data environments.Own data backup, replication and failover, database backup/restore, and data restoration and rollback procedures.Implement data encryption at rest and in transit; support data classification, retention, archival, purging and secure disposal, and time-travel / fail-safe configuration. Security, Monitoring & Compliance Manage KMS, Secrets Manager, WAF/Shield and Config rules; support Security Hub and GuardDuty; contribute to AWS security incident response.Operate CloudWatch monitoring, log aggregation, resource-health and performance tuning; coordinate a single platform observability approach that complements the application-layer monitoring maintained by the data team, avoiding duplicate systems.Provide documentation and evidence to support audits and compliance (UAE PDPL - including 7-day subject-access requests - DFSA / DIFC and GDPR contexts), and support the rollout of data classification, retention and DLP tooling (e.g. AvePoint, Microsoft Purview, OneTrust). Cost Optimisation Take ownership of AWS cost management and budgeting including reserved instances, savings plans and tagging enforcement transitioning this from the current interim arrangement into a permanent responsibility of the role. Parallel Ownership & Shared Capacity Operate as a parallel peer to the Senior Infrastructure Manager, jointly and co-equally owning the hard-infrastructure domain networking, compute, storage, core AWS, security and resilience sharing operational load and providing mutual resilience of cover.Overlap with the SAP Basis Lead on Linux server OS/package patching and Lambda infrastructure providing continuity and additional capacity; SAP Basis patching support where the post-holder's skills allow (SAP knowledge is desirable, not essential).Support pilots and proofs-of-concept and transition them into production operation. KNOWLEDGE AND SKILLS Minimum 15 years of experience in Cloud Engineering.AWS engineering experience across account governance, IAM, networking, compute, storage, security and cost at an architect / senior-engineer level.Proven cloud infrastructure architecture experience designing target-state AWS architectures, patterns and standards.Working knowledge of CMDB / configuration-management practices, with hands on ownership and implementation of a CMDB.Demonstrable experience designing and operating disaster recovery on AWS, including cross-region replication, regional failover and DR testing.Proven delivery of failover and business continuity for business-critical applications, with defined and validated RPO/RTO.Strong Infrastructure as Code (Terraform / CloudFormation) and change-management discipline.Experience with cloud data platforms including Snowflake, and with data security, encryption and DLP concepts.Strong documentation, collaboration and audit-support skills.