Site Reliability Engineer
Vimoinc — United States · Posted ~22 hours ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
Vimo® started as the “Expedia” of health insurance and has evolved into a leader in transforming government IT infrastructure with its proven SaaS and AI technology.
Our innovative approach to health insurance shopping and enrollment has expanded beyond exchanges, and we are now reinventing how states administer safety net programs such as Medicaid, SNAP (food stamps), child care, and unemployment insurance.
With our cutting-edge technology, we are helping agencies serve more people, faster, and transforming healthcare service delivery as we know it.
We are looking for a Site Reliability Engineer (SRE) to join our Vimo team.
About The Role
As a Site Reliability Engineer, you will help ensure the reliability, availability, and performance of Vimo’s production platform.
Our systems power health insurance exchanges, Medicaid enrollment, and other safety-net programs for state governments—the services you support directly impact millions of people’s access to critical benefits.
You will work alongside senior SREs, developers, and infrastructure engineers to build automation, improve observability, respond to incidents, and reduce operational toil.
This role is ideal for an engineer who is passionate about systems thinking, enjoys solving problems at scale, and wants to grow their career in reliability engineering within a mission-driven environment.
Responsibilities
Monitor, maintain, and troubleshoot production services to ensure high availability and performance across Vimo’s SaaS platform.
Respond to production incidents as part of the on-call rotation; triage alerts, coordinate with engineering teams, and drive issues to resolution.
Contribute to blameless postmortem processes by documenting incidents, identifying root causes, and tracking follow-up action items.
Build and maintain CI/CD pipelines to support safe, repeatable, and efficient application deployments.
Write automation scripts and tools (Python, Bash, or Go) to reduce manual operational work and improve reliability.
Support and improve observability infrastructure including monitoring dashboards, log aggregation, distributed tracing, and alerting using tools such as Datadog, Prometheus, Grafana, ELK/OpenSearch, and PagerDuty.
Manage cloud infrastructure on AWS (EC2, EKS, RDS, S3, VPC, CloudFront, Route 53, Lambda) following established standards and best practices.
Work with infrastructure-as-code tools (Terraform) and container orchestration (Docker, Kubernetes/EKS) to provision and manage environments.
Assist with capacity planning, load testing, and performance analysis to prepare for peak traffic periods such as open enrollment seasons.
Collaborate with application development teams to improve service reliability through architecture reviews, production readiness checklists, and resilience patterns.
Support disaster recovery procedures including backup validation, failover testing, and documentation of recovery runbooks.
Help maintain compliance with security and regulatory requirements (HIPAA, FedRAMP, SOC 2) by ensuring infrastructure controls are properly implemented and documented.
Continuously improve on-call processes, runbooks, and operational documentation to reduce mean time to detection (MTTD) and mean time to resolution (MTTR).
Qualifications
Basic Qualifications/Skills
Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
2+ years of experience in Site Reliability Engineering, DevOps, Systems Engineering, or a related operations-focused role.
Proficiency in at least one programming or scripting language (Python, Bash, Go, or similar) for writing automation and tooling.
Hands-on experience with AWS cloud services (EC2, RDS, S3, VPC, IAM, CloudWatch, or equivalent).
Familiarity with container technologies (Docker) and orchestration platforms (Kubernetes).
Experience with infrastructure-as-code tools such as Terraform, or similar.
Understanding of CI/CD concepts and experience with at least one pipeline tool (Jenkins, GitLab CI, GitHub Actions, or similar).
Familiarity with monitoring and observability tools (Loki, Prometheus, VictoriaMetrics, Grafana, CloudWatch, ELK, or PagerDuty).
Solid understanding of Linux/Unix systems administration, including process management, file systems, and networking basics.
Understanding of networking fundamentals: TCP/IP, DNS, HTTP/HTTPS, load balancing, and TLS/SSL.
Strong troubleshooting and analytical skills with the ability to diagnose issues across the application and infrastructure stack.
Good communication skills and the ability to work collaboratively in a team-oriented environment.
Preferred Qualifications/Skills
Experience in healthcare technology, government IT, or benefits administration platforms.
Familiarity with compliance frameworks such as HIPAA, FedRAMP, or SOC 2 and their operational implications.
Experience with PostgreSQL, Mongo or such relational & NoSQL database systems in a production environment.
Exposure to incident management frameworks and blameless postmortem practices.
Experience with GitOps workflows and tools (ArgoCD or similar).
Familiarity with configuration management tools (Puppet, Ansible, Chef or similar).
Experience with log management and analysis at scale.
Exposure to load testing or performance benchmarking tools (k6, Locust, JMeter).
AWS certifications (Cloud Practitioner, Solutions Architect Associate, or SysOps Administrator) are a plus.
Familiarity with SLO/SLI concepts and error budget-driven development practices.
Compensation and Benefits
Competitive compensation - All In range of ($120,000-$165,000).
(Please note that compensation may vary based on factors such as skills, experience, performance and location.)
We offer a comprehensive benefits package, including but not limited to:
Health, Dental, Life, Disability, and Vision insurance Healthcare spending or reimbursement accounts (HSA/FSA) Retirement benefits (401k) Paid time off Holidays: 13 paid days per year Education assistance or tuition reimbursement Employee discounts for Gym memberships & commuting/travel assistance Our Values
We believe that working hard, when it is imbued with purpose, can and should be fun.
You'll find we are a "can do" place where people work together and roll up their sleeves to get the job done.
Everyone has a voice; everyone's ideas count, and everyone is respected.
We have built a company, as well as a community of friends and colleagues, with respect for each other.
We have 77,319 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 77,319 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume