Senior Site Reliability Engineer

Staffingscience β€” United States Β· Posted ~5 hours ago

Senior Full-time No Visa

Skills

Linux/UNIX systems Site reliability engineering Infrastructure engineering Cloud infrastructure Automation Platform reliability Terraform AWS GitHub Actions Java Python Linux UNIX SRE

πŸ”“ Log in to save this job, tailor your resume & track your apply process β€” 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior Site Reliability Engineering role within a large-scale, highly regulated cloud environment. The position suits a systems-first engineer with extensive Linux/UNIX experience who has evolved into SRE, working on mission-critical infrastructure, automation, modernization, and global platform reliability.

Highlights

Highly visible senior infrastructure role with direct collaboration with engineering and information security leadership, focused on large-scale reliability, automation, and infrastructure strategy.

Description

Senior Site Reliability Engineer β€” Cloud Communications / High-Compliance Enterprise SaaS Our client is a large-scale, high-compliance cloud communications and healthcare data company that does business with the federal government and requires all employees to be U.S. citizens. For over 25 years, they've been a leader in their space and they're continuing to invest heavily in infrastructure modernization, automation, and platform reliability at global scale. They are adding a 4th member to a tight-knit SRE team supporting mission-critical infrastructure that processes and delivers data across a global, high-volume, always-on platform. This is a highly visible role β€” you'll partner directly with Engineering and Information Security leadership on infrastructure strategy, not just execute tickets. This role is built for a systems-first engineer β€” ideally a ~15–20 year Linux/UNIX systems engineer who has evolved into an SRE, rather than a developer who transitioned into infrastructure. This is not a role for someone purely cloud-native with no deep OS/networking background. The team needs someone who has genuinely run enterprise-scale platforms and lived through the operational fires that come with it β€” not someone who has only worked greenfield, cloud-native environments. What you'll own: Design, write, and maintain Terraform and Terragrunt modules (not just consumption) managing AWS infrastructure at scale, including EC2, S3, RDS, VPC, IAM, Lambda, and both ECS and EKS environmentsOwn and evolve CI/CD pipelines across GitHub Actions and AWS CodePipeline (GitLab experience a plus) β€” both for infrastructure and application deploymentsTroubleshoot deep networking and mail delivery issues β€” DNS, IP/sender reputation, deliverability at global scale, and the class of problems that come with high-volume mail transportSupport and maintain Apache and Postfix-based mail infrastructure in productionProvide SRE support into Tomcat/Java-based application environments β€” diagnosing the class of problems Java/Tomcat teams run into (memory/GC issues, thread pool exhaustion, connection pooling, API contract issues) even without being a full-time Java developerBuild and maintain observability and alerting across the stack using tools like Prometheus, Grafana, OpenTelemetry, and the ELK stack; define and manage SLOs/SLIs for critical servicesSupport and extend configuration automation using tools such as AWS Config, SSM, Ansible, Puppet, or ChefWork hands-on with containerization (Docker) and container orchestration, with particular depth in Amazon ECSUse APM tooling (New Relic, Jaeger, Zipkin, or OpenTelemetry) to identify and resolve performance bottlenecks across distributed systemsOperate in a heavy on-call rotation supporting large-scale, high-availability systems (once fully staffed: 1 week on, 3 weeks off); lead and contribute to blameless postmortems following incidentsDraft and contribute to RFCs and internal standards for IaC, automation, and infrastructure design patternsMentor other engineers β€” including junior/lower-level ops staff β€” on Linux systems, IaC, CI/CD, and operational best practicesRamp expectation: contribute to the codebase within 30 days, understand the environment deeply and take on smaller scoped projects by 60 days, own projects independently by 90 daysWhat you bring: 15+ years of Linux/UNIX systems engineering, with meaningful time in SRE/DevOps roles at enterprise scaleExpert-level AWS (EC2, S3, RDS, VPC, IAM, Lambda, ECS/EKS, CloudWatch) and infrastructure design best practicesExpert-level Terraform and Terragrunt (module authorship required β€” this is not a "consume existing modules" role)Strong production experience in Python and GoDeep enterprise networking fundamentals β€” DNS, TCP/IP, mail routing and delivery troubleshooting, sender/IP reputation managementHands-on Apache and Postfix administration experiencePrior exposure to Tomcat/Java-based production environments and the operational issues specific to that stackExperience with configuration automation tooling (Ansible, Puppet, Chef, or AWS Config/SSM)Solid containerization background (Docker), with strong familiarity in the ECS ecosystem specificallyExperience with observability/monitoring frameworks at scale (Prometheus, Grafana, OpenTelemetry, ELK, Thanos)Experience operating in large-scale, high-compliance enterprise environments (healthcare, fintech, or similarly regulated industries β€” PCI/HITRUST/FedRamp not required, but that class of environment strongly preferred)Customer-facing operational support background (UNIX/Windows) is a strong plusComfortable working within Agile/Scrum and Waterfall environments, using Jira/ConfluenceComfortable digesting large amounts of information quickly and moving from context to independent contribution fastA self-starter mentality β€” able to work independently with minimal oversight while staying aligned with team goalsYou'll stand out if you also have: Experience mentoring engineers into stronger DevOps/SRE practicesBackground operating in PCI, HITRUST, or FedRamp/GovCloud-adjacent environmentsExperience leading a team or platform through an SDLC/DevOps maturity transitionA GitHub profile or code samples that showcase personal or professional projects