Site Reliability Engineer

Veritas Search Group โ€” United States ยท Posted ~1 hour ago

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Description

This role requires candidates who are currently authorized to work in the U.S. without sponsorship, and C2C arrangements are not accepted. This role is onsite near Tustin, CA. Please submit your resume and send to recruiters@vsearchgroup.com if you are interested. Thank you! Job Description We are seeking an experienced Site Reliability Engineer (SRE) to join our Product Platform team. This role will serve as a critical bridge between product development teams and the platform engineering organization, helping translate application infrastructure needs into scalable, reliable platform solutions. The ideal candidate has deep, hands-on experience with Kubernetes, AWS, and observability, along with strong troubleshooting and cross-functional communication skills. This is an infrastructure-focused engineering role for someone who is comfortable working directly with product teams, diagnosing complex application and platform issues, and driving problems through resolution across multiple technical teams. Responsibilities Build, manage, maintain, and troubleshoot Kubernetes clusters and containerized application environments.Partner closely with product and application development teams to understand infrastructure requirements and translate them into actionable platform engineering needs.Serve as a first point of contact for infrastructure and reliability issues affecting product teams, performing root-cause analysis across application, Kubernetes, cloud, networking, and security layers.Resolve Kubernetes and platform-related issues directly while coordinating with other infrastructure, cloud, network, and security teams when issues fall outside the platform team's ownership.Create, configure, and maintain AWS resources supporting application and platform environments.Support and enhance an internal observability platform used to monitor applications, services, and infrastructure.Onboard new applications and use cases into the observability platform by partnering with technical teams to understand monitoring and telemetry requirements.Develop and maintain Python scripts used for infrastructure automation, troubleshooting, platform operations, and observability.Support CI/CD and GitOps-based deployment processes for Kubernetes environments.Improve platform reliability, scalability, monitoring, operational efficiency, and developer experience.Participate in troubleshooting and root-cause investigations involving multiple engineering teams and drive issues through successful resolution.Document platform standards, troubleshooting procedures, operational processes, and technical solutions. Required Qualifications Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field with 4+ years of relevant experience, or 6+ years of equivalent professional experience in lieu of a degree.Deep hands-on experience with Kubernetes, including building clusters from scratch, cluster administration, deployments, networking, and troubleshooting.Strong hands-on experience with AWS, including creating and maintaining cloud infrastructure and resources.Experience working with AWS services such as S3 and RDS.Strong understanding of observability, monitoring, logging, metrics, and application performance concepts.Experience with observability platforms such as Datadog, Splunk, Grafana, Prometheus, or similar technologies.Experience using Python for scripting, automation, or infrastructure-related tasks.Strong troubleshooting and root-cause analysis skills across complex application and infrastructure environments.Excellent verbal and written communication skills with the ability to work effectively across product, application, infrastructure, cloud, networking, and security teams.Ability and willingness to work in a highly collaborative position that combines hands-on engineering with significant cross-team coordination. Preferred Qualifications Experience with Helm for Kubernetes application packaging and deployment.Experience with Argo CD and GitOps-based deployment practices.Experience building or maintaining CI/CD pipelines.Experience with Prometheus, Grafana, and OpenTelemetry.Experience with Amazon EKS and IAM.Familiarity with Istio or other service mesh technologies.Experience with Kafka or Amazon MSK.Experience with infrastructure-as-code and configuration-management technologies such as Terraform or Ansible.Experience supporting internal developer platforms, infrastructure platforms, or observability products.Experience onboarding development teams or applications onto centralized platform services.Experience working in regulated or highly controlled technology environments.Interest in learning and taking ownership of application code supporting internal platform tooling.