Summary
✨ AI‑Generated
A remote-first organization is seeking its first dedicated Senior AWS/DevOps Engineer to own infrastructure, reliability, and deployment tooling across multiple live applications. The role includes multi-account cloud architecture, least-privilege IAM, secrets management, compute scaling, load balancing, and infrastructure modernization.
Highlights
Fully remote ownership of infrastructure and reliability across multiple live applications, with significant autonomy and responsibility for modernizing cloud operations.
Description
Reports to: Technical Leadership
Location: Remote
About the role
We're looking for our first dedicated, in-house AWS/DevOps engineer to own infrastructure, reliability, and deployment tooling across a portfolio of live, revenue-generating products.
You'll be the technical owner for cloud infrastructure spanning multiple AWS accounts, supporting several distinct applications with their own architectures and audiences —plus the shared services that cut across all of them.
What you'll own
• Multi-account AWS architecture — administer and harden a multi-account AWS Organization (production, development, and monitoring accounts) with least-privilege IAM, proper root/administrative credential hygiene, and durable secrets management (moving us off ad-hoc password storage and onto a real secrets manager).
• Compute & scaling — manage EC2 fleets behind Application Load Balancers, Auto Scaling Groups, and launch templates backed by a golden-AMI deployment pipeline; modernize legacy standalone instances toward containerized (ECR/ECS) or otherwise more resilient patterns where it makes sense.
• Databases — administer Amazon Aurora MySQL clusters (multi-database, multi-tenant schemas) including read replica strategy, credential scoping (moving application traffic off shared/master-level database users onto least-privilege service accounts), and CDC-based real-time data pipelines (Epsio) feeding BI and marketing systems.
• Networking & access — operate a Tailscale-based zero-trust network (ACLs, subnet routers, tagged service identities) as our primary access layer, including migrating legacy DNS off a legacy registrar onto modern DNS/CDN infrastructure (Cloudflare / Route 53).
• CI/CD — own and improve GitLab CI/CD pipelines (OIDC-based AWS role assumption, environment-scoped deploys, secrets-as-CI-variables) across a PHP monolith and a Laravel application; drive the migration of hardcoded application secrets into properly managed CI/CD variables and parameter stores.
• Observability & incident response — build out real monitoring where gaps exist today (CloudWatch
Logs/Alarms/Synthetics, external uptime monitoring, WAF logging), define on-call/runbook practices, and lead incident response for production issues, including root-cause writeups.
• Security posture — manage AWS WAF rules, resolve longstanding dependency/composer security advisories, own certificate lifecycle management (TLS, Apple Pay/payment-processing certs), and generally reduce the accumulated operational risk of a fast-growing, historically under-resourced infrastructure.
• Cross-team support — work directly with a small in-house engineering team and outside contractors, and interface with third-party MSPs supporting sibling properties.
Our stack (what you'll actually touch)
• AWS: EC2, Aurora MySQL (RDS), Application Load Balancer, Auto Scaling Groups, S3, Storage Gateway, ElastiCache (Memcached), CloudWatch (Logs, Alarms, Synthetics, Internet Monitor), Systems Manager (Session Manager, Parameter Store), IAM & AWS Organizations, SES, WAFv2, CloudFormation
• Application layers: Legacy PHP (no framework) and Laravel (PHP framework); Apache/PHP-FPM and nginx; Composer dependency management
• Data & BI: Aurora MySQL, Epsio (CDC/real-time materialized views), Sigma (BI), Metabase (legacy BI, being retired)
• Networking: Tailscale (zero-trust mesh VPN, ACL policy management), Cloudflare (DNS/CDN), legacy DNS providers (migration in progress)
• CI/CD & tooling: GitLab CI/CD, OIDC-based cloud role assumption, PhpStorm/JetBrains tooling, LastPass (transitional— a real secrets manager is part of the roadmap)
• Third-party integrations: Payment processing (CardPointe/CardConnect, Apple Pay, Google Pay, PayPal), SMS/messaging (Twilio, Plivo), CRM/marketing (HubSpot), accounting exports
What you bring
• 5+ years in a DevOps/SRE/Infrastructure role with deep, hands-on AWS experience (not just "used AWS" — configured
IAM policies, debugged Aurora replication, tuned Auto Scaling, written CloudFormation/Terraform from scratch).
• Real production incident experience: you've diagnosed a live outage under pressure, found the actual root cause (not
just the symptom), and closed every downstream copy of the fix (config, image/AMI, and CI/CD source of truth) — not
just the first thing that unblocked traffic.
• Strong Linux systems administration fundamentals — filesystems, systemd, log management, disk/volume operations,
recovering access to a box when the "normal" path (SSH key, SSM agent) isn't available.
• Experience with CI/CD pipeline design (GitLab CI/CD strongly preferred) and secrets management done right.
• Comfort working across old and new: you can support and gradually modernize a 10+ year old PHP codebase while also
building out reliability tooling for a newer Laravel application, without treating either as beneath you.
• Excellent judgment around production risk — knowing when to move fast, when to get a second pair of eyes, and how to
communicate a live incident clearly to non-technical stakeholders.
Nice to have
• Experience with CDC/real-time data replication tools (Debezium, Epsio, or similar).
• Experience with Tailscale or another zero-trust/mesh VPN platform.
• Payment processing / PCI-adjacent infrastructure experience.
• Prior experience being the first dedicated DevOps hire at a company — you like building the practices, not just following
them.
• Experience using AI coding/ops tools (e.g., Claude Code or similar) as part of your actual workflow — for infrastructure
investigation, incident diagnosis, or accelerating day-to-day DevOps work, not just general familiarity with AI chatbots.