Lead AWS Platform Engineer

Creorecruitment — United Kingdom · Posted ~3 hours ago

Lead Full-time Hybrid £100K-£120K

Skills

AWS cloud infrastructure infrastructure as code CI/CD DevOps incident response disaster recovery Aurora Redshift Redis

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A lead cloud platform engineer role responsible for designing, operating, and evolving secure production infrastructure. The position combines hands-on engineering with technical leadership and DevOps best practices.

Highlights

Senior technical leadership role with ownership of scalable cloud infrastructure, architecture decisions, and platform reliability.

Description

Lead AWS Platform Engineer Location: London – Hybrid, 2 days per week in the office Employment: Permanent, Full-Time Salary Range: £100k - £120k The Role: We’re looking for an experienced Lead AWS Platform Engineer to take ownership of the design, operation and ongoing evolution of a production AWS environment. This is a senior, highly hands-on individual contributor role for someone who enjoys taking genuine ownership of infrastructure. You’ll be responsible for ensuring the platform remains scalable, resilient, secure and cost-efficient, while acting as the lead technical authority for cloud infrastructure. You’ll own production data services including Aurora, Redshift and Redis, infrastructure as code, CI/CD pipelines, observability, incident response and disaster recovery. Working closely with Product, Engineering and Data teams, you’ll provide engineering squads with a reliable platform to build on while helping shape infrastructure architecture and DevOps best practice across the organisation. What You'll Be Doing: Own and evolve the AWS infrastructure strategy across development, staging and production.Design for high availability, scalability, resilience and performance.Own production Aurora, Redshift and Redis environments, including availability, failover, backups, upgrades and performance optimisation.Build and maintain infrastructure as code using AWS CDK and Terraform.Develop and maintain GitHub Actions CI/CD pipelines, including safe deployment and rollback strategies.Take ownership of on-call and incident response, leading investigations and producing meaningful postmortems.Design, document and regularly test backup and disaster recovery processes.Build effective observability across metrics, logging, alerting and dashboards.Implement strong cloud security practices including IAM, encryption, patching, vulnerability management and secrets management.Optimise AWS expenditure through rightsizing, Reserved Instances, Savings Plans and ongoing cost analysis.Lead architecture reviews and promote strong DevOps and platform engineering practices.Identify future scale and performance challenges before they become problems.Mentor engineers and improve infrastructure knowledge and practices across engineering teams. What We're Looking For: You’ll have 7+ years' experience across DevOps, SRE, Cloud or Platform Engineering, including significant experience operating production AWS environments. We're particularly interested in people who have worked within start-up or scale-up environments and have previously acted as the lead or sole owner of cloud infrastructure. You'll ideally bring strong hands-on experience across: AWS: VPC, networking, IAM, EC2, ECS/Fargate, Lambda, S3, CloudWatch, KMS and Secrets ManagerInfrastructure as Code: Terraform and AWS CDKData: Aurora, Redshift and RedisCI/CD: GitHub ActionsObservability: metrics, logging, monitoring, alerting and dashboardsSecurity: IAM, encryption, vulnerability management, patching and secrets managementAutomation: Python and/or BashReliability: incident management, on-call, postmortems, backups and disaster recovery You'll have experience making architectural decisions yourself rather than simply implementing designs created by others, and you'll be comfortable balancing reliability, security, scalability, performance and cloud cost. The Experience That Matters: We'd particularly like to speak with engineers who have: 7+ years in DevOps, SRE, Cloud or Platform Engineering.At least 3 years' experience running production AWS data services.Worked as the lead or sole infrastructure owner within a start-up or scale-up.Personally owned production incidents and on-call responsibilities.Led investigations and written postmortems following incidents.Designed and actually tested disaster recovery and backup restores, rather than simply configuring them.A track record of improving AWS cost efficiency and platform performance. Nice to Have: Experience with Databricks or broader data platforms, working within AdTech or another high-volume data environment, and/or an AWS Professional-level certification would all be advantageous, but aren't essential. The Person: This role will suit someone who enjoys ownership and autonomy. You'll be comfortable making decisions around tooling and architecture, taking responsibility for a production platform and explaining technical trade-offs clearly to both engineers and non-technical stakeholders. We're looking for someone pragmatic rather than dogmatic: someone who prefers simple, reliable solutions, improves existing systems where appropriate and thinks carefully about the trade-offs between resilience, cost, security and scalability. You'll also care about good documentation, runbooks and sharing knowledge so that critical infrastructure knowledge doesn't live with one person. Interested? If you're a senior AWS/Platform/DevOps engineer looking for a genuinely hands-on role with significant ownership of a production cloud environment, we'd love to hear from you.