DevOps Engineer

Rdt Com โ€” Denmark ยท Posted ~3 hours ago

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Description

Senior DevOps & Platform Engineer (Internal Developer Platform, DevSecOps, AIOps) Location: Remote within EU We are looking for a Senior / Lead Platform & DevOps Engineer with 8+ years of experience to design, build and scale an Internal Developer Platform, enterprise GitOps delivery and automated remediation across multi-cluster Kubernetes environments. The role sits in Integration Testing & Release DevOps and requires strong hands-on depth from day one. Responsibilities Design and standardize a self-service Internal Developer Platform (e.g. Backstage or Port) with golden deployment paths and secure, ephemeral environmentsArchitect multi-cluster, multi-region GitOps delivery with ArgoCD or Flux, including automated canary analysis, traffic shadowing and automated rollbacksBuild closed-loop remediation and AIOps pipelines using OpenTelemetry, anomaly correlation and workflow/orchestration engines (Temporal, Kubernetes Operators in Go)Implement policy as code and software supply chain security with Kyverno or OPA, SAST/SCA gates, cosign and SBOMsDefine SLO/SLI governance, error budgets and automated chaos testing for distributed servicesEstablish standards for IaC, CI/CD and telemetry, and drive them through cross-team RFCs and architecture reviewsSupport inference and training infrastructure for AI workloads, including GPU autoscaling (KEDA) Required experience 6+ years of enterprise-scale Kubernetes in production: multi-cluster networking, service meshes (Istio/Linkerd), ingress controllers, CRD designStrong Go (custom operators, platform controllers) and Python (automation, data and AI pipelines)Modular Terraform, OpenTofu or Pulumi, including state management, drift detection and policy testingGitOps and delivery tooling: ArgoCD, Flux, GitHub ActionsObservability: OpenTelemetry collector topologies, tail-based sampling, Prometheus, Grafana (Mimir/Tempo), ClickHouse or DatadogPolicy as code and supply chain security: Kyverno, OPA, SAST/SCA, cosign, SBOMsDeterministic, idempotent workflows with Temporal or event-driven architectures on KafkaTrack record of authoring cross-team RFCs and defining operational SLOs/SLIsCloud: AWS and/or GCPIncident tooling: PagerDuty, Incident.io, Slack API Nice to have Integration of ITSM/CMDB platforms (iTop or similar) into automated incident workflowsExperience integrating LLM tool-calling, agentic triage or RAG-based runbook retrieval into incident responseAI inference infrastructure (Ray, vLLM, Triton, Kubeflow)Contributions to CNCF/open-source projects