Platform Engineer - Observability Platforms

Msdczech — Czechia · Posted ~1 day ago

Senior

Skills

platform engineering observability infrastructure automation CI/CD Infrastructure as Code DevOps cloud-native systems cloud-native

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An experienced Platform Engineer role focused on designing scalable cloud-native infrastructure and company-wide observability capabilities. You’ll build monitoring, alerting, tracing, and event-management systems, automate CI/CD and infrastructure through IaC, and advise stakeholders on technical direction.

Highlights

Design scalable cloud-native platforms, lead observability capabilities across monitoring and tracing, automate delivery workflows, and influence important technical decisions.

Description

Job Description We are seeking a highly experienced Platform Engineer to join our engineering team. In this role, you will design and implement modern, scalable infrastructure solutions, drive automation across our software delivery lifecycle, and engineer robust cloud-native systems. Beyond technical execution, you will act as a trusted advisor—influencing key technical decisions and guiding stakeholders across the product line. You will collaborate closely with product teams to deliver reliable, secure, and high-performing solutions. Responsibilities Design, build, and operate scalable observability platform capabilities that will enable reliable monitoring, alerting, tracing, and event management across the companyDevelop and automate infrastructure, CI/CD workflows, and platform services using Infrastructure as Code and modern DevOps practices to improve delivery speed, compliance, and operational consistency for the custom solutions our team is responsibleLead engagements with product teams to define and design custom solutions for the observability platforms, enabling engineering capabilities to kick start by the rest of the teamDrive continuous improvement through AIOps, telemetry insights, and incident learnings to reduce alert noise, accelerate response times, and improve overall service resilience Requirements System Design Proven experience in system design and modern custom architecture concepts (microservices, event-driven, serverless, distributed systems)Continuous Delivery & Integration Hands-on expertise building CI/CD pipelines (e.g., GitLab CI, GitHub Actions), designing and building cloud services using IaC then deploying, maintaining those services for compliance and security. (e.g., Terraform, CloudFormation)Cloud Engineering Strong hands-on implementation experience across core AWS services (e.g., EC2, S3, Lambda, ECS/EKS, RDS, IAM, VPC, CloudWatch)API Development Demonstrated ability to design and build APIs using Python (e.g., Flask or FastAPI)Observability & Monitoring Extensive knowledge and hands-on experience of OpenTelemetry, event management, and monitoring systems (e.g., Prometheus, Grafana, Big Panda, Logic Monitor)AIOps Familiarity and exposure on using LLMs via API driven solutions, detect anomalies, and accelerate incident response to improve reliability and reduce alert noiseAgile Ways of Working Deep experience working in an Agile/Scrum environment, with hands-on expertise in managing a product backlog using tools like JIRA and ConfluenceInterpersonal skills Ability to influence key stakeholders and different team members. Providing SME level guidance, removing blockers and enabling other junior team members Preferred 7+ years of relevant professional experienceAbility to influence team members and stakeholders in decision making, ability to convey the right message while working within diver geography and culturesAWS Certification (e.g., AWS Certified Solutions Architect, DevOps Engineer Professional)Hands-on experience with container orchestration (Docker, Kubernetes)Extensive knowledge of monitoring/observability tools (Prometheus, Grafana, ELK, Datadog)Familiarity with security best practices (DevSecOps) and compliance frameworksStrong scripting skills (Python, Bash) and version control (Git)Good understanding of SDLC processesBachelor degree from relevant programmes (Information Technologies, Electronics, etc.) Required Skills AI Ops, Amazon Web Services (AWS), API Development, Atlassian Confluence, Atlassian JIRA, CI/CD, Cloud Engineering, IT Monitoring, Large Language Models (LLMs), Python (Programming Language), Scrum (Agile), Stakeholder Management, System Designs Preferred Skills Bash (Scripting Language), Container Orchestration, Datadog, DevSecOps, Docker (Software), Elastic Stack (ELK), Git, Grafana, Kubernetes, Prometheus (Software), Software Development Life Cycle (SDLC) Current Employees apply HERE Current Contingent Workers apply HERE Search Firm Representatives Please Read Carefully Merck & Co., Inc., Rahway, NJ, USA, also known as Merck Sharp & Dohme LLC, Rahway, NJ, USA, does not accept unsolicited assistance from search firms for employment opportunities. All CVs / resumes submitted by search firms to any employee at our company without a valid written search agreement in place for this position will be deemed the sole property of our company. No fee will be paid in the event a candidate is hired by our company as a result of an agency referral where no pre-existing agreement is in place. Where agency agreements are in place, introductions are position specific. Please, no phone calls or emails. Employee Status Regular Relocation No relocation VISA Sponsorship Yes Travel Requirements 10% Flexible Work Arrangements Hybrid Shift Not Indicated Valid Driving License No Hazardous Material(s) N/A Job Posting End Date 09/11/2026 A job posting is effective until 11 59 59PM on the day BEFORE the listed job posting end date. Please ensure you apply to a job posting no later than the day BEFORE the job posting end date. Requisition ID R415569