Site Reliability Engineer/L3 Support
Remotehunter — United States · Posted ~2 hours ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
About Our Client
The organization operates in the financial services and healthcare technology sectors, providing expertise, scale, and technology solutions to more than 20,000 clients worldwide.
With a global workforce of over 27,000 employees across 35 countries, the company supports financial services and healthcare organizations ranging from large multinational corporations to small and mid-market businesses.
About the Opportunity
The Site Reliability Engineer (SRE) / L3 Support Engineer is responsible for maintaining the reliability, availability, security, and operational health of a FedRAMP High cloud platform.
This role combines advanced site reliability engineering with senior-level production support to proactively identify and resolve complex issues, minimize customer impact, and strengthen platform resilience.
The position works closely with engineering and operations teams to improve observability, automation, incident response, and overall system performance.
Responsibilities
• Monitor production services for health, availability, performance, and security.
• Use telemetry, logs, metrics, and distributed tracing to proactively identify emerging issues.
• Troubleshoot and resolve complex production incidents across applications and infrastructure.
• Serve as the L3 escalation point for operational issues beyond L1 and L2 support.
• Participate in on-call rotations for critical production incidents.
• Lead incident response, coordination, communication, and post-incident reviews.
• Conduct root cause analysis and implement corrective actions to prevent recurring issues.
• Develop and maintain operational runbooks, dashboards, alerts, and support procedures.
• Improve platform observability and service-level indicators.
• Collaborate with software engineering teams to improve reliability, scalability, and resilience.
• Automate operational tasks and workflows to reduce manual effort.
• Support production deployments, infrastructure changes, and maintenance activities.
• Assist with disaster recovery, resilience testing, and operational readiness initiatives.
• Ensure platform operations align with FedRAMP High security and compliance requirements.
• Contribute to continuous improvement initiatives focused on reliability, performance, and operational efficiency.
Requirements
• U.S.
citizenship is required.
• 3–6 years of experience in site reliability engineering, production engineering, DevOps, or senior technical support roles.
• Experience supporting mission-critical cloud-based production systems.
• Strong knowledge of Linux operating systems and networking fundamentals.
• Experience troubleshooting distributed applications running in Kubernetes environments.
• Familiarity with public cloud platforms, preferably AWS.
• Experience with infrastructure as code and configuration management.
• Strong scripting or programming skills in languages such as Python, Bash, PowerShell, or Go.
• Experience with monitoring and observability tools such as Prometheus, Grafana, or CloudWatch.
• Ability to analyze logs, metrics, and traces to diagnose complex production issues.
• Knowledge of incident and problem management processes.
• Strong analytical, troubleshooting, communication, and collaboration skills.
Preferred Qualifications
• Experience with FedRAMP High, DoD IL5/IL6, or similar regulated environments.
• Advanced experience operating Kubernetes in production environments.
• Knowledge of AWS services including EKS, RDS, IAM, and networking.
• Experience with CI/CD pipelines and deployment automation.
• Understanding of service mesh technologies such as Istio.
• Familiarity with security best practices and compliance monitoring.
• Experience with incident management platforms such as PagerDuty or Jira Service Management.
• AWS certifications are a plus.
Pay Range and Compensation Package
• The pay range and compensation package for this role will be determined based on the candidate’s experience, skills, qualifications, location, and other relevant factors.
Benefits & Perks
• 401(k) matching program.
• Professional development reimbursement.
• Flexible personal and vacation time off.
• Sick leave.
• Paid holidays.
• Medical, dental, and vision insurance.
• Employee assistance program.
• Parental leave.
• Discounts on fitness clubs and travel.
Equal Opportunity Statement: Our client is an equal opportunity employer.
They celebrate diversity and are committed to creating an inclusive environment for all employees.
All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, or national origin
Note: RemoteHunter is not the Employer of Record (EOR) for this role.
Our purpose in this opportunity is to connect exceptional candidates with leading employers.
We help job seekers worldwide discover roles that match their goals and guide them to complete their full application directly through the hiring company’s career page or ATS.
We have 110,181 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 110,181 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume