Summary
β¨ AIβGenerated
A senior Site Reliability Engineering role within a large-scale, highly regulated cloud environment. The position suits a systems-first engineer with extensive Linux/UNIX experience who has evolved into SRE, working on mission-critical infrastructure, automation, modernization, and global platform reliability.
Highlights
Highly visible senior infrastructure role with direct collaboration with engineering and information security leadership, focused on large-scale reliability, automation, and infrastructure strategy.
Description
Senior Site Reliability Engineer β Cloud Communications / High-Compliance Enterprise SaaS
Our client is a large-scale, high-compliance cloud communications and healthcare data company that does business with the federal government and requires all employees to be U.S.
citizens.
For over 25 years, they've been a leader in their space and they're continuing to invest heavily in infrastructure modernization, automation, and platform reliability at global scale.
They are adding a 4th member to a tight-knit SRE team supporting mission-critical infrastructure that processes and delivers data across a global, high-volume, always-on platform.
This is a highly visible role β you'll partner directly with Engineering and Information Security leadership on infrastructure strategy, not just execute tickets.
This role is built for a systems-first engineer β ideally a ~15β20 year Linux/UNIX systems engineer who has evolved into an SRE, rather than a developer who transitioned into infrastructure.
This is not a role for someone purely cloud-native with no deep OS/networking background.
The team needs someone who has genuinely run enterprise-scale platforms and lived through the operational fires that come with it β not someone who has only worked greenfield, cloud-native environments.
What you'll own:
Design, write, and maintain Terraform and Terragrunt modules (not just consumption) managing AWS infrastructure at scale, including EC2, S3, RDS, VPC, IAM, Lambda, and both ECS and EKS environmentsOwn and evolve CI/CD pipelines across GitHub Actions and AWS CodePipeline (GitLab experience a plus) β both for infrastructure and application deploymentsTroubleshoot deep networking and mail delivery issues β DNS, IP/sender reputation, deliverability at global scale, and the class of problems that come with high-volume mail transportSupport and maintain Apache and Postfix-based mail infrastructure in productionProvide SRE support into Tomcat/Java-based application environments β diagnosing the class of problems Java/Tomcat teams run into (memory/GC issues, thread pool exhaustion, connection pooling, API contract issues) even without being a full-time Java developerBuild and maintain observability and alerting across the stack using tools like Prometheus, Grafana, OpenTelemetry, and the ELK stack; define and manage SLOs/SLIs for critical servicesSupport and extend configuration automation using tools such as AWS Config, SSM, Ansible, Puppet, or ChefWork hands-on with containerization (Docker) and container orchestration, with particular depth in Amazon ECSUse APM tooling (New Relic, Jaeger, Zipkin, or OpenTelemetry) to identify and resolve performance bottlenecks across distributed systemsOperate in a heavy on-call rotation supporting large-scale, high-availability systems (once fully staffed: 1 week on, 3 weeks off); lead and contribute to blameless postmortems following incidentsDraft and contribute to RFCs and internal standards for IaC, automation, and infrastructure design patternsMentor other engineers β including junior/lower-level ops staff β on Linux systems, IaC, CI/CD, and operational best practicesRamp expectation: contribute to the codebase within 30 days, understand the environment deeply and take on smaller scoped projects by 60 days, own projects independently by 90 daysWhat you bring:
15+ years of Linux/UNIX systems engineering, with meaningful time in SRE/DevOps roles at enterprise scaleExpert-level AWS (EC2, S3, RDS, VPC, IAM, Lambda, ECS/EKS, CloudWatch) and infrastructure design best practicesExpert-level Terraform and Terragrunt (module authorship required β this is not a "consume existing modules" role)Strong production experience in Python and GoDeep enterprise networking fundamentals β DNS, TCP/IP, mail routing and delivery troubleshooting, sender/IP reputation managementHands-on Apache and Postfix administration experiencePrior exposure to Tomcat/Java-based production environments and the operational issues specific to that stackExperience with configuration automation tooling (Ansible, Puppet, Chef, or AWS Config/SSM)Solid containerization background (Docker), with strong familiarity in the ECS ecosystem specificallyExperience with observability/monitoring frameworks at scale (Prometheus, Grafana, OpenTelemetry, ELK, Thanos)Experience operating in large-scale, high-compliance enterprise environments (healthcare, fintech, or similarly regulated industries β PCI/HITRUST/FedRamp not required, but that class of environment strongly preferred)Customer-facing operational support background (UNIX/Windows) is a strong plusComfortable working within Agile/Scrum and Waterfall environments, using Jira/ConfluenceComfortable digesting large amounts of information quickly and moving from context to independent contribution fastA self-starter mentality β able to work independently with minimal oversight while staying aligned with team goalsYou'll stand out if you also have:
Experience mentoring engineers into stronger DevOps/SRE practicesBackground operating in PCI, HITRUST, or FedRamp/GovCloud-adjacent environmentsExperience leading a team or platform through an SDLC/DevOps maturity transitionA GitHub profile or code samples that showcase personal or professional projects