Description
Position Summary
The Effectual SRE Engineer improves the availability, reliability, recoverability, and observability of enterprise applications and services across cloud and hybrid environments.
Working across multiple products and engineering teams in a public sector government contracting environment, this hands-on engineer defines and monitors SLIs, SLOs, and error budgets; improves monitoring and alerting; supports incident response and post-incident reviews; and automates recurring operational work.
Participation in an after-hours support rotation may be required.
Essential Duties
Define, implement, and monitor SLIs, SLOs, error budgets, and other reliability metricsDesign and improve observability — metrics, logging, tracing, dashboards, monitoring, and alertingBuild and maintain automated CI/CD pipelines for large enterprise environmentsImplement DevSecOps practices across the build, test, security-scan, release, and deploy lifecycleDevelop, deploy, and maintain containerized applications using enterprise container platformsAutomate recurring operational work via scripting, infrastructure-as-code, and cloud-native toolingParticipate in incident response, root cause analysis, and post-incident reviewsIdentify reliability and performance risks and implement corrective actionsSupport disaster recovery planning, testing, and implementationTroubleshoot application, infrastructure, deployment, and integration issues across cloud and hybrid environmentsCollaborate with engineering, architecture, security, and operations teams on reliable, secure solutionsProvide technical guidance and operational support to multiple engineering teamsEvaluate processes and identify opportunities to improve automation, reliability, and deployment efficiencyParticipate in after-hours and on-call activities as required by the program
Education & experience (one of the following):
Bachelor's in a related field + 4 years relevant experience and an Associate/Professional-level cloud certification (AWS or Azure), or6 years of on-the-job experience in lieu of degree and certification
Required experience
(4+ years in SRE, DevOps, DevSecOps, cloud, software, or systems engineering), including:Supporting production applications in public or hybrid cloud environmentsCI/CD pipelines, automated builds, testing, release management, and deployment — using GitLab CI/CD or a comparable enterprise platformGit or comparable version controlBuilding and maintaining containerized applications with Docker and Kubernetes, OpenShift, or comparableMonitoring, logging, alerting, and metrics/observability toolingProduction incident response, troubleshooting, and root cause analysisScripting/automation (Python, Bash, PowerShell, or comparable)Infrastructure-as-code and configuration management (Terraform, CloudFormation, Ansible, or comparable)
Nice-to-Have
Active Public Trust or clearance (USDA preferred)Experience with large enterprise environments (hundreds of applications, multiple teams)AWS GovCloud, Azure Government, or other government cloudSystems subject to FISMA, FedRAMP, or NISTProduction Kubernetes administration and orchestrationObservability platforms and distributed tracing (CloudWatch, Prometheus, Grafana, OpenSearch, Elasticsearch, Splunk, OpenTelemetry)Disaster recovery, performance engineering, resilience, and operational readinessSecurity/compliance automation within CI/CD pipelinesMulti-cloud or hybrid-cloud migrations, modernization, and operational transitionsDirect experience with government customers and program leadership
Preferred Certifications
AWS Certified DevOps Engineer – ProfessionalAWS Certified Solutions Architect – Associate or ProfessionalMicrosoft Certified: DevOps Engineer ExpertGoogle Professional Cloud DevOps EngineerCertified Kubernetes Administrator (CKA)FinOps, security, or automation certifications relevant to the role