Infrastructure Engineer

Letterboxd — New Zealand · Posted ~5 hours ago

Senior Full-time

Skills

Infrastructure engineering Production operations Performance optimization Reliability engineering Infrastructure planning Security Operational resilience Emergency management

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An infrastructure engineering role responsible for the performance, reliability, security, and long-term health of a high-traffic digital platform. You will lead operational resilience, infrastructure planning, and emergency management while ensuring production systems remain fast, secure, recoverable, and ready to scale.

Highlights

Ownership of production infrastructure, reliability, security, and long-term platform health, with substantial responsibility for resilience, planning, and incident preparedness.

Description

Location: Auckland, New Zealand Reports to: Head of Engineering Letterboxd Letterboxd connects people around the films and shows they love. We spark discovery, conversation and community through a trusted voice and a product that's deeply loved. Behind every review, list, diary entry and late-night rewatch is a platform that millions of members rely on to be fast, secure and always there. As our community grows, it puts more demand on the infrastructure that keeps it running. We're looking for an Infrastructure Engineer to own the performance, reliability, security and long-term health of that platform. This isn't about reinventing what works. It's about making sure Letterboxd can keep growing while staying fast, secure, recoverable and ready for the unexpected. The Role We're hiring an Infrastructure Engineer to lead the operational resilience, infrastructure planning and emergency management of Letterboxd's technology platform. You'll own the health of our production environment, from monitoring and capacity planning to security, incident response and disaster recovery. You'll also bring practical DevOps capability to make releases safer, systems easier to operate and our engineering teams more effective. You'll work closely with software engineers, product teams and leadership, making sure the infrastructure decisions we make today support where Letterboxd is heading tomorrow. You'll also help shape how we support the platform around the clock, including setting up an on-call roster and growing a team of reliability engineers over time. You'll also look after Letterboxd's company IT environment, so our team has the tools, access and security they need to do their best work. This role suits someone who is calm under pressure, enjoys being both strategic and hands-on, and gets satisfaction from building systems that are simple, dependable and built to last. What You'll Do Own the operational health, availability and reliability of Letterboxd's production platformMonitor and improve performance, capacity, scalability and resilience as our community and product growDefine practical service measures, alerts and operational standards that make platform health visible and actionableEstablish 24/7 infrastructure support, including an on-call roster and response processes, and help build out a team of reliability engineersOwn and continuously improve infrastructure security across cloud, networking, access, secrets and deployment environmentsBuild security into infrastructure design and delivery rather than treating it as a final checkpointLead technical responses to major incidents with calm structure and clear decision-makingMaintain pragmatic escalation, on-call, backup and disaster recovery practices, and test them regularlyFacilitate blameless post-incident reviews that turn failures and near misses into measurable improvementsImprove CI/CD pipelines, deployment tooling and release practices so changes ship safely and confidentlyChampion infrastructure as code, repeatable environments and automation-first ways of workingSet a clear infrastructure roadmap covering capacity, lifecycle, cost and technology planning, including the thoughtful retirement of ageing componentsDocument critical systems, decisions and dependencies so knowledge is shared and durableProvide hands-on technical leadership, coach others and influence sound engineering decisions across teamsManage Letterboxd's company IT environment, including hardware, security, and employee onboarding and offboarding Who You Are Reliability Leader Calm, decisive and methodical when systems are under pressureBrings clear structure to incidents without creating bureaucracyCollaborative and low-ego, with a blameless approach to learning and improvementHighly organised, with strong ownership, follow-through and attention to operational detail Practical Technologist Strong technical judgement, balancing speed, cost, security, reliability and complexityProactive, with a habit of finding and fixing risks before they become incidentsPrefers simple, maintainable and automated solutionsComfortable moving between roadmap decisions and hands-on problem-solving Clear Communicator Makes complex technical issues understandable for engineers, product teams and senior stakeholdersInfluences good decisions across teams without needing formal authorityWrites documentation people actually use Infrastructure Experience Significant experience owning cloud infrastructure in production for public-facing digital products; bare metal experience, or interest in it, is a plusStrong hands-on background across DevOps, site reliability, platform engineering or infrastructure engineeringExperience with containers, orchestration, infrastructure as code and modern deployment pipelinesExperience designing monitoring, observability, alerting, backup, recovery and incident management practicesSound knowledge of systems architecture, networking, security and distributed systemsExperience improving performance, resilience and operational maturity in a growing environmentExperience supporting high-traffic platforms is valuablePassion for film, television and the communities that form around them Tools we use Kubernetes (ideally including bare metal), Docker, Linux/Ubuntu, Cloudflare, Grafana, Sensu, Ansible, Git, GitHub and GitHub Actions, CI/CD and shell scripting. You don't need to know every tool on day one. Experience with similar tools counts. Why now We're growing. Fast. New members. New regions. New features. More traffic than ever. Every new member who finds us, logs a film or shares a review relies on a platform that just works. As we scale, the foundations matter more than ever. We've built a product our community loves and trusts. Now we're looking for someone who can make sure it stays fast, secure and resilient, and who can build an infrastructure function that grows alongside the platform. If that sounds like you, we'd love to hear from you.