DevOps Engineer
Rdt Com โ Denmark ยท Posted ~3 hours ago
๐ Log in to save this job, tailor your resume & track your apply process โ 7 days free, no card needed.
Log in to add to target listDescription
Senior DevOps & Platform Engineer (Internal Developer Platform, DevSecOps, AIOps)
Location: Remote within EU
We are looking for a Senior / Lead Platform & DevOps Engineer with 8+ years of experience to design, build and scale an Internal Developer Platform, enterprise GitOps delivery and automated remediation across multi-cluster Kubernetes environments.
The role sits in Integration Testing & Release DevOps and requires strong hands-on depth from day one.
Responsibilities
Design and standardize a self-service Internal Developer Platform (e.g.
Backstage or Port) with golden deployment paths and secure, ephemeral environmentsArchitect multi-cluster, multi-region GitOps delivery with ArgoCD or Flux, including automated canary analysis, traffic shadowing and automated rollbacksBuild closed-loop remediation and AIOps pipelines using OpenTelemetry, anomaly correlation and workflow/orchestration engines (Temporal, Kubernetes Operators in Go)Implement policy as code and software supply chain security with Kyverno or OPA, SAST/SCA gates, cosign and SBOMsDefine SLO/SLI governance, error budgets and automated chaos testing for distributed servicesEstablish standards for IaC, CI/CD and telemetry, and drive them through cross-team RFCs and architecture reviewsSupport inference and training infrastructure for AI workloads, including GPU autoscaling (KEDA)
Required experience
6+ years of enterprise-scale Kubernetes in production: multi-cluster networking, service meshes (Istio/Linkerd), ingress controllers, CRD designStrong Go (custom operators, platform controllers) and Python (automation, data and AI pipelines)Modular Terraform, OpenTofu or Pulumi, including state management, drift detection and policy testingGitOps and delivery tooling: ArgoCD, Flux, GitHub ActionsObservability: OpenTelemetry collector topologies, tail-based sampling, Prometheus, Grafana (Mimir/Tempo), ClickHouse or DatadogPolicy as code and supply chain security: Kyverno, OPA, SAST/SCA, cosign, SBOMsDeterministic, idempotent workflows with Temporal or event-driven architectures on KafkaTrack record of authoring cross-team RFCs and defining operational SLOs/SLIsCloud: AWS and/or GCPIncident tooling: PagerDuty, Incident.io, Slack API
Nice to have
Integration of ITSM/CMDB platforms (iTop or similar) into automated incident workflowsExperience integrating LLM tool-calling, agentic triage or RAG-based runbook retrieval into incident responseAI inference infrastructure (Ray, vLLM, Triton, Kubeflow)Contributions to CNCF/open-source projects
We have 144,685 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume โ in under a minute we'll analyze all 144,685 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume