Senior Platform/Site Reliability Engineer (SRE)

Newpage Solutions — Canada · Posted ~4 hours ago

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Description

Senior Platform/Site Reliability Engineer (SRE) About Newpage Solutions Newpage Solutions is a global digital health innovation company helping people live longer, healthier lives. We partner with life sciences organisations which include, pharmaceutical, biotech and healthcare leaders, to build transformative AI and data driven technologies addressing real-world health challenges. From strategy and research to UX design and agile development, we deliver and validate impactful solutions using lean, human-centered practices. We are proud to be a ‘Great Place to Work®’ certified company for the last three consecutive years. We also hold a top Glassdoor rating and are named among the "Top 50 Most Promising Healthcare Solution Providers" by CIOReview. As an organisation, we foster creativity, continuous learning and inclusivity, creating an environment where bold ideas thrive and make a measurable difference in people’s lives. Your Mission We are looking for a hands-on Senior Site Reliability Engineer (SRE) to own production readiness and day-to-day operations for a self-hosted Langfuse platform running on AWS EKS. The role will focus on platform reliability, observability, incident management, backup and disaster recovery, infrastructure automation, and scaling of backend and storage components. The environment includes Langfuse Web/Worker, ClickHouse, PostgreSQL, Redis/Valkey, S3, Kubernetes/EKS, Terraform, and Argo CD. What You’ll Do Own production operations and reliability of Langfuse running on AWS EKS.Define and implement SLIs, SLOs, dashboards, alerts, and operational runbooks.Monitor and improve availability, performance, scalability, and platform health.Manage ClickHouse, PostgreSQL, Redis/Valkey, and S3 operational requirements.Implement and validate backup, restore, and disaster-recovery processes.Scale ClickHouse clusters, storage, and Langfuse workers as platform usage grows.Implement Kubernetes autoscaling using technologies such as KEDA/HPA.Manage production deployments and promotions through Argo CD and GitOps.Support incident triage, troubleshooting, rollback, escalation, and post-incident reviews.Work closely with engineering teams to improve production readiness and release safety.Support observability gateway and OpenTelemetry-based trace ingestion.Implement trace governance, routing, and operational controls. What You Bring Strong hands-on experience managing production Kubernetes environments, preferably AWS EKS.Strong experience with AWS services, including IAM/IRSA, S3, networking, and secrets management.Hands-on experience operating ClickHouse, including:ReplicationBackup and restorePerformance troubleshootingStorage and capacity scalingGood production experience with PostgreSQL and Redis/Valkey.Experience with Terraform / Infrastructure as Code.Strong experience with Argo CD, GitOps, and CI/CD pipelines.Experience implementing monitoring, alerting, dashboards, and observability solutions.Good understanding of SLIs, SLOs, incident management, and production support practices.Strong troubleshooting and root-cause analysis skillsNice to have Experience with Langfuse or other LLM observability platforms.Nice to have Experience with OpenTelemetry (OTEL) collectors and gateways.Nice to have Experience supporting high-volume tracing or telemetry platforms.Nice to have Experience with KEDA-based Kubernetes autoscaling.Nice to have Experience implementing S3 storage tiering or large-scale data retention strategies.Nice to have Familiarity with AI/LLM production infrastructure and observability.Nice to have Experience working in small, fast-moving platform or infrastructure teams. Key Deliverables The engineer will work as part of a small 2–3 member pod responsible for making the self-hosted Langfuse platform production-ready and delivering key platform milestones, including: OTEL ingest collector/gateway foundationBackend and storage scalingOperational readiness and controlsTrace governance and routing controlsNon-production observability gateway foundation What We Offer At Newpage, we’re building a company that works smart and grows with agility, where driven individuals come together to do work that matters. We offer: A people-first culture - Supportive peers, open communication and a strong sense of belongingSmart, purposeful collaboration - Work with talented colleagues to create technologies that solve meaningful business challengesBalance that lasts - We respect your time and support a healthy integration of work and lifeRoom to grow - Opportunities for learning, leadership and career development, shaped around youMeaningful rewards - Competitive compensation that recognises both contribution and potential Ready to Apply? Let’s build the future of health together. Apply below or reach out to: Bhavik.rathod@newpage.io