Site Reliability Engineer
Amelco Limited — United Kingdom · Posted ~1 day ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
Role: Site Reliability Engineer
Type: Full-time permanent role
Location: Hybrid/ Shoreditch, 3 days per week
About Us
Amelco Ltd are a leading gaming and gambling solution software provider with a strong presence in the USA, UK, and Europe.
Through partnerships with global gaming companies, we build cutting-edge technical platforms across sportsbooks, lottery, casino, virtual gaming, and financial trading.
Our vision is to shape the future of gaming by transforming operations into intelligent, data-driven solutions that deliver exceptional customer experiences and create sustainable value for all stakeholders.
We believe in teamwork, knowledge sharing, and transparency with accountability.
The Role
We’re looking for a hands-on Site Reliability Engineer (SRE) to own the reliability, observability, and cost efficiency of our high-throughput betting platform.
You’ll be embedded in our production operations, working directly with development teams to build resilient systems, implement actionable observability, and drive incident response from detection to remediation.
This role is 60% infrastructure automation and 40% application-focused reliability work — you’ll need to dive deep into both our Java Spring Boot services and Kubernetes deployment patterns to build lasting improvements.
What You’ll Work With
Infrastructure & Platform:
Kubernetes: AWS EKS clusters and on-prem deployments GitOps: FluxCD for declarative Kubernetes management across 20+ environments Infrastructure as Code: Terraform for AWS resources and Kubernetes configurations CI/CD: GitHub Actions pipelines with automated AI code review workflows
Observability Stack:
Monitoring: Prometheus metrics with custom Java Micrometer instrumentation Logging: Loki for distributed log aggregation Tracing: Tempo for distributed tracing AWS: Cloudwatch
Application Environment:
Core Platform: Java Spring Boot microservices for betting, trading, and customer management Event Processing: High-throughput event queues with latency monitoring (Kakfa, JMS) Data Pipeline: Avro-based S3 buffer systems with fault-tolerant write-ahead logs (StatefulS3Buffer)) DB: Postgres DB on RDS/Aurora.
Key Responsibilities
Reliability Engineering (40%):
Partner with development teams to define and manage SLOs/SLIs specific to betting platform services (latency, queue depth, bet processing success rates) Enhance observability of our Java Spring Boot services — ensure metrics, logs, and tracing are actionable for detecting and fixing production issues Implement chaos engineering experiments targeting our event queue systems and high-availability betting services Design and execute pre-deployment readiness checks and post-release validation for risk-critical services
Infrastructure & Platform (40%):
Own Kubernetes cluster reliability across development, QA, and production environments Automate operational processes using Python and Bash scripting within our GitHub Actions ecosystem Optimize infrastructure costs through rightsizing EKS workloads, tuning autoscaling policies, and implementing efficient resource utilization patterns Strengthen platform guardrails through Flux GitOps policies and Terraform module validation
Incident & Operations (20%):
Contribute to major incident response for betting platform outages, providing engineering expertise on Java service behavior and infrastructure dependencies Design and implement automated remediation patterns for common failure modes (event queue backpressure, connection pool exhaustion, database connectivity) Build hourly observability agent rules and thresholds for early-warning detection of production anomalies Develop runbooks and automation for common operational tasks across our multi-environment platform
Required Skills & Experience
3+ years in SRE, Platform Engineering, or DevOps roles with hands-on production experience Strong Kubernetes operational expertise (AWS EKS, GitOps with Flux, Helm packaging) Proven track record building observability for Java Spring Boot applications using Prometheus, Grafana, and Loki Infrastructure as Code proficiency with Terraform and GitOps workflows Python and Bash scripting skills for automation and tooling development Experience designing and operating monitoring for high-throughput event processing systems Demonstrated ability to balance infrastructure cost efficiency with betting platform reliability requirements Excellent communication skills and ability to work across development, platform, and incident management teams
Nice-to-Have Skills
Experience with betting/gaming platform architecture and reliability patterns Familiarity with Avro-based data pipelines and S3 storage optimization AWS Certifications (Solutions Architect, DevOps Engineer, or AWS Certified Kubernetes Administrator) Background in chaos engineering for stateful event processing systems Knowledge of payment processing and risk calculation system reliability patterns Experience with automated incident detection and response workflows
Amelco Benefits
Pension Scheme: Amelco matches up to 7% contributions of your base salary for all staff.
You will automatically be entered at 4%.Staff Benefits Scheme: access to the staff benefits portal after successful completion of a 4-month probation period.Yearly discretionary bonus scheme and pay reviewsOpportunity to travel and visit our office in Poland and Hungary
If you are interested, hit the apply button, and please ensure you upload a copy of your CV.
We look forward to hearing from you!
We have 64,022 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 64,022 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume