Senior AI-Assisted DevOps Engineer

Bullpen Capital — Australia · Posted ~2 hours ago

Senior Full-time

Skills

DevOps SysOps AWS Azure Kubernetes Infrastructure architecture Helm Terraform Monitoring and alerting Infrastructure security IAM Incident response Capacity planning Prometheus Mimir Grafana CloudWatch PostgreSQL ClickHouse Redis RabbitMQ AWS SQS Temporal

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Take ownership of cloud infrastructure for a rapidly scaling AI-focused technology platform. You will manage Kubernetes environments, design reliable infrastructure, build Helm and Terraform configurations, own monitoring and alerting, optimize data stores and messaging systems, and strengthen security through secrets management, IAM, vulnerability scanning, and encryption. This is a high-impact role for an experienced DevOps/SysOps engineer who enjoys owning infrastructure reliability end-to-end.

Highlights

Own and evolve highly scalable cloud infrastructure supporting AI-driven analytics and data processing. The role offers broad responsibility across Kubernetes, infrastructure as code, observability, databases, messaging systems, security, incident response, and reliability engineering in a rapidly scaling environment.

Description

Improvado is an AI-powered marketing intelligence platform trusted by enterprise brands like ASUS, Activision, Docker, and H&R Block. We raised $34M Series A and are scaling fast — which means our infrastructure needs to be rock-solid. We're looking for a DevOps/SysOps engineer who takes ownership and keeps things running. What You'll Do Manage and evolve our cloud infrastructure on AWS/Azure (Kubernetes) - ensure cluster reliability: capacity planning, autoscaling, incident response, and post-mortemsDesign and architect scalable, reliable infrastructure to support AI-driven analytics and data processing at scaleBuild and maintain Helm charts, Terraform, and multi-environment setupsOwn monitoring and alerting across the stack (Prometheus/Mimir, Grafana, CloudWatch)Administer and optimize storages: PostgreSQL, ClickHouse, RedisSupport message brokers: RabbitMQ, AWS SQS, TemporalDrive infrastructure security: secrets management, IAM policies, vulnerability scanning, and encryptionOwn network design, configuration — VPCs, subnets, firewalls, load balancers, VPNsMonitor and optimize cloud costs — identify waste, right-size resources, and report on spend efficiencyParticipate in on-call rotation What We're Looking For 5+ years in a DevOps/SRE roleSolid hands-on experience with most of our stack (80%): AWS/Azure/GCP, Kubernetes (cluster design, capacity planning, and reliability at scale), Helm charts, CI/CD pipelines (GitHub CI), Terraform, PostgreSQL, ClickHouse, Redis, RabbitMQ, AWS SQS, Temporal, and monitoring/alerting stacks (Prometheus/Mimir, Grafana, CloudWatch)AI-assisted development in practice — actively uses Claude Code / agentic coding to ship, knows where AI-generated code needs validationStrong Linux fundamentalsBash and/or Python scriptingComfortable with ambiguity and high-velocity, fast-shifting priorities — we move like a startup, not a committeeDetail-oriented and pedantic, with sharp attention to detail in high-volume, high-stakes systems — accountable and genuinely invested in the quality of your workEST timezone availabilityFluency in English Nice to have Golang or Python development backgroundHashiStack: Vault, PackerClickHouse administrationExperience supporting data-intensive pipelines What We Offer Remote-first environment20 days PTO + US holidaysOptional relocation support to Latin AmericaModern AI-native tech stackStock optionsProfessional development reimbursementA genuinely fun, transparent startup culture