Site Reliability Engineer

Hmg America Llc — United States · Posted ~21 hours ago

Contract Hybrid

Skills

Python R machine learning scikit-learn TensorFlow PyTorch Pandas NumPy Power BI Tableau SQL Linux/Unix cloud platforms Docker Kubernetes monitoring CI/CD Infrastructure as Code Linux AWS Azure GCP Prometheus Grafana Terraform

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join a long-term hybrid SRE engagement combining reliability engineering with machine learning and data analytics. You will work across Linux, cloud platforms, containers, observability, CI/CD, Infrastructure as Code, databases, and modern ML and visualization tools.

Highlights

Long-term W-2 hybrid opportunity combining SRE, cloud infrastructure, machine learning, data analysis, observability, automation, and DevOps. The role offers broad exposure across modern infrastructure and data technologies.

Description

Role - Site Reliability Engineer Duration: Long term (w2 only) Location - Schaumburg, IL or Secaucus, NJ (Hybrid) Need local Mandatory Skills: Python/R and ML libraries (scikit-learn, TensorFlow, PyTorch), Data analysis and visualization (Pandas, NumPy, Power BI/Tableau), SQL and database management Required Skills • Strong experience with Linux/Unix administration. • Proficiency in scripting languages such as Python, Shell, or PowerShell. • Hands-on experience with cloud platforms (AWS, Azure, or GCP). • Experience with containerization technologies such as Docker and Kubernetes. • Knowledge of monitoring and observability tools such as Prometheus, Grafana, ELK, Splunk, Dynatrace, or Datadog. • Understanding of CI/CD tools such as Jenkins, GitHub Actions, GitLab CI, or Azure DevOps. • Experience with Infrastructure as Code tools such as Terraform, Ansible, or CloudFormation. • Strong troubleshooting, debugging, and problem-solving skills. • Understanding of networking, security, and distributed systems concepts. Experience • 5–10+ years of overall IT experience. • 5+ years of hands-on experience in Site Reliability Engineering, Production Support, DevOps, or Cloud Operations roles. Note - We are seeking a highly motivated Site Reliability Engineer (SRE) to ensure the reliability, scalability, performance, and availability of critical production systems. The ideal candidate will combine software engineering and operations expertise to build automation, improve system resilience, reduce operational toil, and enhance service reliability.