Senior Platform/Site Reliability Engineer (SRE)
Newpage Solutions — Canada · Posted ~4 hours ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
Senior Platform/Site Reliability Engineer (SRE)
About Newpage Solutions
Newpage Solutions is a global digital health innovation company helping people live longer, healthier lives.
We partner with life sciences organisations which include, pharmaceutical, biotech and healthcare leaders, to build transformative AI and data driven technologies addressing real-world health challenges.
From strategy and research to UX design and agile development, we deliver and validate impactful solutions using lean, human-centered practices.
We are proud to be a ‘Great Place to Work®’ certified company for the last three consecutive years.
We also hold a top Glassdoor rating and are named among the "Top 50 Most Promising Healthcare Solution Providers" by CIOReview.
As an organisation, we foster creativity, continuous learning and inclusivity, creating an environment where bold ideas thrive and make a measurable difference in people’s lives.
Your Mission
We are looking for a hands-on Senior Site Reliability Engineer (SRE) to own production readiness and day-to-day operations for a self-hosted Langfuse platform running on AWS EKS.
The role will focus on platform reliability, observability, incident management, backup and disaster recovery, infrastructure automation, and scaling of backend and storage components.
The environment includes Langfuse Web/Worker, ClickHouse, PostgreSQL, Redis/Valkey, S3, Kubernetes/EKS, Terraform, and Argo CD.
What You’ll Do
Own production operations and reliability of Langfuse running on AWS EKS.Define and implement SLIs, SLOs, dashboards, alerts, and operational runbooks.Monitor and improve availability, performance, scalability, and platform health.Manage ClickHouse, PostgreSQL, Redis/Valkey, and S3 operational requirements.Implement and validate backup, restore, and disaster-recovery processes.Scale ClickHouse clusters, storage, and Langfuse workers as platform usage grows.Implement Kubernetes autoscaling using technologies such as KEDA/HPA.Manage production deployments and promotions through Argo CD and GitOps.Support incident triage, troubleshooting, rollback, escalation, and post-incident reviews.Work closely with engineering teams to improve production readiness and release safety.Support observability gateway and OpenTelemetry-based trace ingestion.Implement trace governance, routing, and operational controls.
What You Bring
Strong hands-on experience managing production Kubernetes environments, preferably AWS EKS.Strong experience with AWS services, including IAM/IRSA, S3, networking, and secrets management.Hands-on experience operating ClickHouse, including:ReplicationBackup and restorePerformance troubleshootingStorage and capacity scalingGood production experience with PostgreSQL and Redis/Valkey.Experience with Terraform / Infrastructure as Code.Strong experience with Argo CD, GitOps, and CI/CD pipelines.Experience implementing monitoring, alerting, dashboards, and observability solutions.Good understanding of SLIs, SLOs, incident management, and production support practices.Strong troubleshooting and root-cause analysis skillsNice to have Experience with Langfuse or other LLM observability platforms.Nice to have Experience with OpenTelemetry (OTEL) collectors and gateways.Nice to have Experience supporting high-volume tracing or telemetry platforms.Nice to have Experience with KEDA-based Kubernetes autoscaling.Nice to have Experience implementing S3 storage tiering or large-scale data retention strategies.Nice to have Familiarity with AI/LLM production infrastructure and observability.Nice to have Experience working in small, fast-moving platform or infrastructure teams.
Key Deliverables
The engineer will work as part of a small 2–3 member pod responsible for making the self-hosted Langfuse platform production-ready and delivering key platform milestones, including:
OTEL ingest collector/gateway foundationBackend and storage scalingOperational readiness and controlsTrace governance and routing controlsNon-production observability gateway foundation
What We Offer
At Newpage, we’re building a company that works smart and grows with agility, where driven individuals come together to do work that matters.
We offer:
A people-first culture - Supportive peers, open communication and a strong sense of belongingSmart, purposeful collaboration - Work with talented colleagues to create technologies that solve meaningful business challengesBalance that lasts - We respect your time and support a healthy integration of work and lifeRoom to grow - Opportunities for learning, leadership and career development, shaped around youMeaningful rewards - Competitive compensation that recognises both contribution and potential
Ready to Apply?
Let’s build the future of health together.
Apply below or reach out to:
Bhavik.rathod@newpage.io
We have 142,267 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 142,267 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume