Senior Software Engineer (Application Operations)
Reward Gateway โ United Kingdom ยท Posted ~23 hours ago
๐ Log in to save this job, tailor your resume & track your apply process โ 7 days free, no card needed.
Log in to add to target listDescription
Reward Gateway, part of Edenred, helps organisations engage, motivate and retain their people through employee benefits, recognition and wellbeing solutions.
Our platforms support millions of users globally and play a critical role in our clients' day-to-day operations.
As part of our Platform Engineering & Technical Operations team, you'll play a key role in ensuring our business-critical applications remain stable, reliable and performant for millions of users worldwide.
Your Role in Our Mission
We're looking for a hands-on engineer who enjoys solving complex production problems and improving the reliability of business-critical applications.
Working across our PHP, MySQL and AWS estate, you'll investigate incidents, perform root cause analysis, build automation and implement lasting engineering solutions that reduce operational effort and improve customer experience.
This role offers a unique opportunity to combine software engineering, operational excellence and reliability engineering while helping shape the future of Application Operations at Reward Gateway.
Our environment includes:
PHP applications running in AWS
MySQL databases
Datadog observability
Kibana log analysis
Kubernetes (EKS)
Modern engineering and CI/CD practices
Working Pattern & On-Call
Standard hours: Monday to Friday, 9am to 6pm.
1-in-4 on-call rotation.
ยฃ500 on-call allowance per week (in addition to base salary).
Additional hourly payments for call-outs.
Enhanced rates for public holiday coverage.
Historically low call-out volumes, supported by a strong focus on automation, operational maturity and continuous improvement.
This role offers a hybrid work model to be present in our London office twice a week.
Why Join Us?
Application Operations is a growing engineering capability at Reward Gateway, focused on reliability, automation and continuous improvement.
You'll have the opportunity to:
Solve complex production challenges across customer-facing platforms.Influence the reliability and performance of services used by millions of people worldwide.Build automation and tooling that removes repetitive work and delivers measurable business value.Develop expertise across software engineering, cloud operations, observability and platform reliability.Work closely with Engineering, Product, Platform and SRE teams to drive meaningful technical improvements.Join a collaborative team that values ownership, learning and continuous improvement.This removes a lot of the repetition around "improving reliability", "automation", "operational excellence" and "production incidents" while still selling the role.
What You'll Be Doing
Investigate, troubleshoot and resolve production issues across PHP applications, APIs, MySQL databases and AWS-hosted services.Participate in the on-call rota, support incident response activities and contribute to post-incident reviews.Perform root cause analysis, identify recurring issues and implement long-term fixes.Analyse logs, metrics and traces using Datadog, Kibana and supporting observability tools.Assess customer and business impact during incidents and communicate progress to technical and non-technical stakeholders.Write production-quality PHP code, automation and operational tooling to reduce manual effort and improve reliability.Develop and maintain runbooks, playbooks and operational documentation.Partner with Engineering, Platform and SRE teams to improve application performance, resilience and supportability.Enhance monitoring, alerting and operational visibility across services.Contribute to continuous improvement initiatives that reduce operational toil and increase service reliability.
A job description is available on request.
Key Skills & Experience Required
Experience supporting customer-facing production applications in medium to large scale organisations, including incident management, root cause analysis and operational support across complex cloud-based environments.Strong PHP experience, with the ability to investigate issues, troubleshoot code and implement safe fixes where required.Strong MySQL operational knowledge, including query analysis, performance troubleshooting, slow query investigation and database issue diagnosis.Proven experience investigating and resolving complex production issues across applications, APIs, integrations and databases.Experience managing or contributing to the resolution of high-priority production incidents, including root cause analysis and post-incident improvement activities.Hands-on experience using Datadog (or equivalent observability platforms), together with log analysis tools such as Kibana, to investigate and diagnose production issues.Experience assessing customer and business impact during incidents and using this information to prioritise remediation activities.Strong communication and documentation skills, including maintaining runbooks, operational documentation and providing clear stakeholder updates during incidents.Familiarity with ITSM tooling and incident/problem management processes, with a focus on reducing operational toil through automation and continuous improvement.
Interview Process
Screening call with a member of the Talent Acquisition TeamFirst stage interview with Application Operations Leadership and peer (practical scenario or technical assessment relevant to the L2.5 operating model)Final stage interview with Director or VP
At Reward Gateway | Edenred we are committed to ensuring an inclusive and accessible recruitment process for all candidates.
If you have any specific requirements or need reasonable adjustments at any stage of the recruitment journey, please let your Talent Acquisition Partner know.
Your needs are important to us, and we want to ensure an equitable experience for every candidate.
Be Comfortable.
Be You.
At Reward Gateway | Edenred, we want everyone to feel comfortable bringing their passion, creativity and individuality to work.
We value diverse backgrounds, perspectives and experiences because we believe diversity drives innovation.
Join us and help make the world a better place to work.
We have 76,249 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume โ in under a minute we'll analyze all 76,249 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume