Site Reliability Engineer

Ateeca Inc — United States · Posted ~5 hours ago

Mid

Skills

Kubernetes Prometheus network security Python C C++ monitoring systems Alertmanager

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A research and technology organization is seeking an SRE to improve infrastructure reliability through automation, monitoring, and operational tooling. The role includes developing tools, managing systems, and supporting complex environments.

Highlights

Infrastructure engineering role focused on reliability, automation, monitoring, security, and large-scale technical operations.

Description

Job Description: Work closely with other NERSC groups to coordinate center-wide maintenance activities and manage diagnostic and notification software during maintenance periodsExperience with developing tools using various programming languages such as C, C++, Perl, Java, or Python or a scripting language with knowledge of standard software development practices.Motivated, self-starter who can learn technologies that improve data center management in areas like Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, building management software, evaporative cooling, and power utilization.Experience with network security: configuring/maintaining ACLs, knowledge of firewallsPractical experience in developing and deploying Agentic AI or autonomous automation tools to streamline technical tasks.Experience with ServiceNow implementation is a plus