Summary
✨ AI‑Generated
Own reliability and operations for mission-critical technology environments at the intersection of cloud infrastructure, observability, security, and AI. You will work across major cloud platforms and emerging AI-native stacks, helping organizations build secure, governed, highly available systems while gaining exposure to varied technical environments.
Highlights
Hybrid working with only one office day per week, strong ownership, exposure to cutting-edge cloud and AI infrastructure, and opportunities to work across startups and regulated mission-critical environments.
Description
Role: Site Reliability Engineer
Location: Central Cardiff, South Wales
Working pattern: Hybrid (1 day a week in Cardiff office)
Salary: £45,000 - £50,000 per annum
About the Role
This company provides managed AI operations for technology businesses.
The company operates, secures and governs the cloud, observability and AI runtime layer behind mission-critical software.
This is a chance to take real ownership in a Cardiff business on the front line of AI-era cloud operations, with bleeding-edge tech, AI used responsibly, a team worth learning from, and the flexibility to do your best work.
Clients range from start-ups and scale-ups building fast without formal engineering controls, through to regulated, mission-critical platforms in fintech, healthtech and insurance, where the operating model and evidence trail matter as much as uptime.
You'll work across traditional cloud platforms like Azure and AWS, as well as newer, AI-native application stacks such as Lovable, and the modern cloud platforms that often sit behind them, like Vercel and Supabase.
The company is a Datadog Advanced Partner (UK) and holds the accolade of being the world's first accredited MSP powered by Datadog.
Datadog is a Nasdaq-listed observability platform.
The company's technical focus is on observability and LLM observability specifically.
Key Responsibilities
Managed Service Delivery - Support customers across a mix of traditional and modern cloud platforms, Azure and AWS, alongside newer AI-native application stacks and the cloud platforms behind them, such as Lovable, Vercel and Supabase.
Work through an ongoing ticket backlog per customer, balancing reactive support with proactive improvement work, and take the lead on triaging and resolving more complex or ambiguous incidents.
Observability & Datadog - Get up to speed on client environments and the Datadog platform quickly, then work independently: identify improvements and either raise them or implement them directly without needing to be told.
Tickets may also be generated directly by Datadog alerts, covering both backlog items and live incidents, including acting as an escalation point for trickier ones.
Build & Improvement Projects - Deliver end-to-end build-and-run projects for managed service customers, from scoping through to live operation, with minimal oversight.
Design and implement cloud cost optimisation strategies, working through the technical dependencies that can sometimes block them.
Manage monitoring and observability infrastructure as estates grow, keeping platforms performant and well-maintained.
Build and deploy new infrastructure in Azure and AWS, including infrastructure-as-code, Datadog implementation, and APM setup, covering the full lifecycle from onboarding through go-live stabilisation.
AWS Platform Work - Work hands-on and confidently with containers (ECS/EKS), EC2 and RDS instances across a range of customer environments, from day-to-day configuration and troubleshooting through to cost and performance optimisation and architectural recommendations.
Also extends into AWS Bedrock for customers building AI-native products.
Internal Tooling & Process -Contribute to a growing suite of internal AI tools built to improve engineering efficiency, reduce context-switching across platforms, and modernise how we run cloud operations day to day.
Help shape best practice and runbooks as the team scales.
Day in the Life
Review ticket queues across the managed service accounts, taking ownership of the more complex or customer-sensitive items.Project time is spent on a mix of Azure, AWS and Datadog implementation work, alongside newer platforms like Lovable, Vercel and Supabase.
Work includes cost optimisation, infrastructure builds, and Datadog implementation and APM setup on live applications.On AWS, this includes container work, EC2 and RDS configuration, and Bedrock.
Python is used throughout.
Depending on the customer, this work may also involve GitHub Actions, Azure DevOps pipelines, or PowerShell scripting.Time is also allocated to internal work covering internal AI tools and process improvements.The role involves working across managed service delivery, project builds, observability, and internal tooling within the same week, rather than working on a single area on an ongoing basis.
Experience Required
Multiple years of hands-on, production AWS experience, this is not a foundational-familiarity role: you should be comfortable owning EC2, RDS, containerised workloads (ECS/EKS), IAM and networking, and cost/performance optimisation without close supervision, including comfortable Linux troubleshooting (CLI, logs, process/resource issues) on EC2-hosted workloadsDemonstrable experience having worked on AWS builds (standing up new infrastructure, IaC, migrations) as well as ongoing operational/support work (troubleshooting, incident response, day-to-day maintenance)Solid working knowledge of observability: metrics, logs, traces, and how to design or tune alerting and dashboardsComfortable scripting and automating: Python, Bash, or similarExperience troubleshooting live production incidents, ideally including on-callClear written and verbal communication: you'll work directly with customers, including on more technical or sensitive conversations
Nice to have
Hands-on Datadog experienceTerraform or other IaC toolingKubernetes or containerised workload experience beyond the basicsWorking knowledge of Azure alongside AWSExperience in a managed service provider or multi-customer environmentExposure to AI-native platforms or tooling (Bedrock, Vercel, Supabase, Lovable, or similar)Familiarity with ISO 27001 or similar compliance frameworksAWS certification (Solutions Architect, SysOps, or similar)
What's on Offer
Join a growing Cardiff business entering its scaling phase, with real ownership, visibility, and a chance to shape how a fast-moving modern MSP operates AI-era cloud infrastructure.
On-call is documented, structured and remunerated in addition to salary.
Training and tooling costs are paid for by the company.
Attendance at industry events and courses is encouraged and funded as part of ongoing development.
Employees are expected to identify and pursue improvements independently, without pre-defined limits on scope.