AI Platform Engineer

Wavegroupio โ€” United Kingdom ยท Posted ~3 hours ago

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Description

๐Ÿ’ป Job Title: AI Platform Engineer ๐Ÿ’ฐ Salary: up to ยฃ115k (+ very generous early-stage equity, up to ~ยฃ90k) ๐Ÿ“ Location: Central London, EC1 (3 office day/week) ๐Ÿฆ Company: B2B FinTech / Fraud Prevention ๐Ÿ‘ฅ Employees: ~30 ๐Ÿ’ธ Funding: $15m+ (Series A) This London startup is building a new intelligence layer designed to bring more context and security to digital payments. Their technology analyses transactions in real time, gathering signals from multiple sources to determine whether a payment is legitimate or potentially fraudulent. The platform combines distributed data systems, real-time investigations and AI-driven decisioning to help financial institutions detect scams while allowing legitimate payments to flow without unnecessary friction. Within 2 years of being founded, they're working with most Tier 1 banks and payment providers in the UK - and are about to launch in the US! ๐Ÿš€ Hiring an AI Platform Engineer to help build the foundations that let a fast-moving engineering team ship safely to some of the country's largest financial institutions. Production is increasingly powered by non-deterministic AI agents, so this isn't a "keep the lights on" role - you'll be defining what good looks like for platform and infrastructure culture, not inheriting someone else's playbook. Key responsibilities: Owning the reliability and operability of production systems - monitoring, alerting, incident response and post-incident learningMaintaining and evolving the infrastructure-as-code estate, making it easy to ship safely and hard to ship dangerouslySecuring infrastructure defaults so the easy path is the safe pathDesigning observability across the stack - metrics, traces, logs, dashboards and alerts - including for AI agent behaviour, where "correct" isn't always the same twiceDriving incident response maturity from detection through resolution to follow-upBuilding platform capabilities that unblock engineering teams - deployment pipelines, release tooling, developer experienceBuilding the guardrails and automation (including AI-assisted triage and response) that let the wider team move fast without breaking thingsSupporting the platform's evolution from stability through to scalability as the company grows โœ… Must have requirements: At least 3-4 years in platform/DevOps/SRE, ideally with genuine ownership at an early-stage startup - comfortable with ambiguityStrong Terraform/IaC experience on real production infrastructure; GCP a strong plusProven observability track record - monitoring, alerting and dashboards for distributed systemsIncident response experience - on-call, running incidents, and building the processes that make both betterSecurity foundations - least-privilege access, secrets management, secure-by-default infrastructureCI/CD experience - deployment pipelines teams trust, with a focus on deploy velocity and rollback safetyGenuine hands-on exposure to agentic AI frameworks (LangChain, LangGraph or similar) within the last ~12 months - not just conceptual awarenessFintech or other regulated-industry experience ๐Ÿ‘ Bonus points for: Experience building observability/reliability for non-deterministic or ML-powered systems specificallyExposure to compliance frameworks (ISO 27001, SOC 2)Software Engineering background - even better if you're fluent in multiple languagesExperience with workflow orchestration engines (Temporal or similar)A track record of bootstrapping a platform/DevOps function from scratch