Site Reliability Engineer

Cobira โ€” Denmark ยท Posted ~1 day ago

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Description

The Role We're looking for a mid-level SRE to own observability and alerting across our platform. Right now our monitoring lives primarily in New Relic, and while we've built solid foundations - Kafka consumer lag tracking, infrastructure health dashboards, custom NRQL alert policies - we know there's a lot more to do. You'll be the person driving this forward. This isn't a pure ops role. You'll write code, design alerting architectures, and work closely with the engineers shipping the platform. When something is on fire, you'll be one of the people who actually understands why. What You'll Do Own and mature our observability stack โ€” alerting policies, dashboards, on-call runbooks, and incident response workflowsDesign and tune alert conditions in New Relic (NRQL, baseline/anomaly detection, composite conditions) to minimize noise and maximize signalIdentify gaps in our monitoring coverage across services, message queues, infrastructure, and network linksBuild and maintain tooling that helps the team understand system behavior โ€” not just when things break, but before they doScripting ability in Python - enough to automate, glue systems together, and write a useful tool when one doesn't existCollaborate with platform engineers on SLIs, SLOs, and error budgetsParticipate in on-call rotation and drive post-incident improvementsContribute to infrastructure work when needed Must-have What We're Looking For 3โ€“5 years of experience in SRE, platform engineering, or a strong DevOps roleHands-on experience building and maintaining observability systems (alerting, dashboards, tracing, logging) - New Relic, Datadog, Grafana, or similarSolid Linux fundamentals and comfort operating in cloud-hosted VM environmentsExperience with containerised workloads (Docker, Docker Compose)A systematic approach to debugging - you form hypotheses, isolate variables, and document what you findGood written communication; we write things down Nice-to-have Experience with Kafka or other message streaming systemsExperience with Redis or other caching technologiesFamiliarity with network-level infrastructure (VPNs, firewall rules, routing)Exposure to telecom or IoT connectivity domainsExperience with NRQL or another query language for observability platformsIaC experience (Terraform, Ansible, or similar)Familiarity with Kubernetes - we're not there yet, but directionally heading that way What We Offer A technically honest environment - we'll tell you what's messy and where improvement is neededMeaningful ownership from day one; no layers of process between you and the problemA compact, experienced team in CopenhagenCompetitive salary based on experienceFlexible hours and autonomy over how you work, within an on-site team cultureThe chance to shape the reliability culture of a growing IoT and connectivity platform How To Apply Send a short note about yourself and why this role interests you, along with your CV.