Description
LOCATION: Remote | Type: Contract
About Newpage Solutions
Newpage Solutions is a global digital health innovation company helping people live longer, healthier lives.
We partner with life sciences organisations which include, pharmaceutical, biotech and healthcare leaders, to build transformative AI and data driven technologies addressing real-world health challenges.
From strategy and research to UX design and agile development, we deliver and validate impactful solutions using lean, human-centered practices.
We are proud to be a Great Place to Work certified company for the last three consecutive years.
We also hold a top Glassdoor rating and are named among the "Top 50 Most Promising Healthcare Solution Providers" by CIOReview.
As an organisation, we foster creativity, continuous learning and inclusivity, creating an environment where bold ideas thrive and make a measurable difference in people s lives.
Your Mission
We are looking for an experienced Platform Reliability Engineer to build and manage observability, reliability, and cost-monitoring capabilities for enterprise AI and agent-based platforms.
The role focuses on implementing OpenTelemetry instrumentation, distributed tracing, Service Level Objectives (SLOs), proactive alerting, and monitoring of LLM token consumption and cloud compute costs.
The engineer will work closely with AI platform, cloud infrastructure, and application engineering teams to ensure platform reliability, operational visibility, and cost efficiency.
What You Ll Do
Observability & Instrumentation: Design and implement OpenTelemetry-based instrumentation across AI agents, services, APIs, and platform infrastructure.
Establish centralized logging, metrics, and distributed tracing.AI Agent Monitoring: Implement end-to-end tracing for agent execution, LLM interactions, tool calls, multi-agent workflows, and external API integrations.
Monitor execution failures, latency, token usage, and performance.Platform Reliability: Define and implement Service Level Indicators (SLIs), SLOs, availability monitoring, and reliability dashboards across AI platform components.Alerting & Incident Management: Configure proactive alerts for service degradation, execution failures, latency spikes, and resource utilization.
Support incident investigation, root cause analysis, and operational runbooks.Cost Observability: Build monitoring and reporting capabilities for LLM token consumption, model API costs, AWS compute utilization, and agent execution costs.
Implement cost attribution, budget thresholds, and anomaly alerts.Observability Integration: Integrate telemetry from AI platforms, AWS services, and applications into centralized observability solutions.
Enable consistent monitoring and reporting across development, staging, and production environments.Automation & Continuous Improvement: Automate observability configuration, dashboard provisioning, and alert management using Infrastructure as Code (IaC).
Identify opportunities to improve platform performance, reliability, and operational efficiency.
What You Bring
7+ years of relevant experience in platform engineering, SRE, DevOps, or production infrastructure operations.Hands-on experience with OpenTelemetry SDKs, Collector configuration, instrumentation, and telemetry pipelines.Strong experience implementing distributed tracing, metrics, logging, dashboards, and alerting.Experience defining and implementing reliability metrics, SLOs, error budgets, and availability monitoring.Experience monitoring AWS infrastructure and services, including CloudWatch, EKS, and containerized workloads.Hands-on experience with Prometheus, Grafana, or equivalent enterprise observability platforms.Experience monitoring and analysing cloud compute costs, resource consumption, and usage trends.Proficiency in Python or another scripting language, with experience in Terraform or equivalent IaC tools.Strong production troubleshooting, incident management, and root cause analysis skills.
Nice To Have Skills
Experience implementing observability for LLM applications, AI agents, and multi-agent orchestration.Familiarity with AWS Bedrock, AgentCore, or similar AI runtime platforms.Experience with Langfuse, Lang Smith, or similar LLM observability platforms.Knowledge of LLM token accounting, model pricing, inference latency, and cost attribution.Experience implementing OpenTelemetry GenAI semantic conventions and tracing agent-to-agent interactions.Familiarity with Kubernetes, Docker, GitOps, and CI/CD pipelines.Experience with FinOps practices, cost optimization, and automated budget alerting.Exposure to Grafana Tempo, Loki, Jaeger, or similar observability tools.
Expected Deliverables
OpenTelemetry instrumentation and end-to-end distributed tracing across AI platform components.Centralized dashboards covering platform health, service performance, agent execution, and LLM usage.Defined SLIs, SLOs, error budgets, and proactive alerting for critical platform services.Token and compute cost dashboards with usage attribution, budget monitoring, and anomaly detection.Operational runbooks, incident investigation workflows, and automated observability configurations.
Preferred Skills
An experienced Platform Reliability Engineer with strong hands-on expertise in observability, cloud infrastructure, and production reliability.
The candidate should be comfortable working across AWS-based AI platforms, implementing standardized telemetry, and translating technical operational data into actionable reliability and cost insights.
What We Offer
At Newpage, we re building a company that works smart and grows with agility, where driven individuals come together to do work that matters.
We offer:
A people-first culture - Supportive peers, open communication and a strong sense of belongingSmart, purposeful collaboration - Work with talented colleagues to create technologies that solve meaningful business challengesBalance that lasts - We respect your time and support a healthy integration of work and lifeRoom to grow - Opportunities for learning, leadership and career development, shaped around youMeaningful rewards - Competitive compensation that recognises both contribution and potential
Ready to Apply?
Let s build the future of health together.
Apply below or reach out to:
Ramadevi.bhumireddy@newpage.io