Description
One of the companies we collaborate with at Joppy, a TOP Cybersecurity Startup, is looking for a Senior DevOps Engineer to join the team.
Work modality: HYBRID in Barcelona (3 days/week at office).
Level: Senior (5+ yrs)
Team: Small, high-ownership
We're a fast-growing cybersecurity company building the platform that security teams use to see and act on threats — threat intelligence, compromised-credential monitoring, malware analysis, attack-surface, and automated offensive testing.
It's a data-heavy product: large ingestion pipelines, search at scale, and real-time views customers depend on.
You'd be joining a small platform team that owns the whole stack end to end — from Terraform and Kubernetes to the alert that pages at 3am.
There's no wall to throw things over; you build it, you run it, you make it reliable.
The Environment (The Real Tech Stack)
We'd rather tell you the truth up front than sell you a buzzword.
This is the real stack — and yes, some of it is mid-migration:
Cloud & Compute: AWS multi-account (EC2 / ASG), Google Cloud, Kubernetes / EKS (prod + staging), Docker, Bare metal.IaC & Delivery: Terraform, GitLab CI, Helm, ArgoCD (GitOps).Observability: OpenTelemetry, VictoriaMetrics, Grafana, Loki, Tempo, vmalert / PromQL.Data & Messaging: Elasticsearch / OpenSearch, ClickHouse, MongoDB Atlas, PostgreSQL (RDS), SQS, Redis.Networking & Access: WireGuard mesh, AWS SSM, Cloudflare, API Gateway + Lambda authorizers.AI / Agentic: Agentic AI, LLM orchestration, MCP, Model-provider APIs.Languages: Python, Bash, Go, SQL, PromQL, VRL.
What You'll Do
Own multi-account AWS, Google Cloud, and bare-metal infrastructure as code: Design and evolve Terraform modules; ship changes through pipeline plan/apply; keep drift and blast radius under control across environments.Run Kubernetes for real: Operate EKS (prod + staging), deliver via GitOps (ArgoCD) and manual Helm where it makes sense, and keep workloads healthy, right-sized, and cost-aware.Drive the observability platform: We're migrating off a costly SaaS to a self-hosted OpenTelemetry + VictoriaMetrics + Grafana + Loki + Tempo stack.
Instrument services, and build alerts that actually mean something because silence is not the same as healthy.Keep the data pipelines reliable: SQS-driven ingestion, Elasticsearch/OpenSearch clusters, ClickHouse, MongoDB, and Postgres — capacity, indexing, backlog, and data-integrity are part of the job, not someone else's.Lead incident response with rigour: Chase root cause, not symptoms; write the fix down; make the same incident impossible next time.Cut toil and cost: Automate the manual, delete the redundant, and treat cloud spend as a first-class metric.Raise the security bar: Secrets management, least privilege, vulnerability management, and support for audits and certifications — you're helping secure a security company.
What We're Looking For (Must-haves)
5+ years in DevOps / SRE running production systems you were on the hook for.Deep AWS: IAM, VPC, EKS, across more than one account (we also run workloads on Google Cloud).Comfort with plain Linux / bare-metal ops matters just as much: half of the interesting problems live on bare metal, not in a console.A strong cybersecurity background: You think in threat models and secure-by-default infrastructure, and you're fluent in secrets management, least privilege, and hardening — you're both running and securing a security product.Terraform in anger: Module design, multiple environments, state, and drift.Kubernetes beyond "kubectl apply": Helm, GitOps, debugging what's actually wrong in a cluster.CI/CD with GitLab as the real delivery path, not a demo pipeline.Strong Linux and scripting (Python, Bash).
Go is a plus — some of our tooling and authorizers are written in it.Observability judgement: Prometheus-compatible metrics (VictoriaMetrics/PromQL), Grafana, logs, and traces, and the instinct to build alerts that don't cry wolf or stay silent when it counts.Hands-on with at least one of Elasticsearch/OpenSearch, MongoDB, PostgreSQL, or ClickHouse — enough to reason about performance and failure modes.Calm, methodical problem-solving under pressure: Root cause over quick patch.
Bonus Points (Nice-to-haves)
You've built or operated agentic AI systems (LLM orchestration, agent/workflow pipelines, MCP, model-provider integrations) — and understand their infra and security implications.You've migrated observability to OpenTelemetry / VictoriaMetrics (or run them at scale).Hybrid/multi-site networking — WireGuard, VPNs, cross-cloud connectivity.Event-driven pipelines at scale — SQS, DLQ design, throughput, and backpressure.Elasticsearch cluster operations (sharding, capacity, recovery) under real load.FinOps / cloud cost optimisation.Compliance experience — SOC 2, ISO 27001, Vanta, vendor security assessments.Relevant certifications (e.g., AWS, CKA/CKAD).
Perks
Hybrid Barcelona, remote-friendly and flexible.High ownership on a platform that genuinely matters.A small, senior team that ships and helps each other.Competitive salary and benefits.