SRE / AWS DevOps Engineer

Saransh Inc Usa β€” United States Β· Posted ~3 hours ago

Full-time Onsite $170000+ per year

Skills

AWS DevOps Site Reliability Engineering Reliability engineering Monitoring and observability MLOps AI observability Terraform CloudFormation Ansible CI/CD Docker Kubernetes Incident response Root cause analysis EC2 EKS ECS Lambda S3 RDS Redshift CloudWatch IAM VPC Jenkins GitHub Actions GitLab CI/CD Helm Arize AI Prometheus Grafana Datadog ELK Stack OpenTelemetry

πŸ”“ Log in to save this job, tailor your resume & track your apply process β€” 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A full-time SRE/DevOps opportunity for an engineer experienced in AWS cloud infrastructure, automation, reliability, and modern observability. You will work across infrastructure as code, CI/CD, containers, Kubernetes, incident response, and monitoring, with additional focus on MLOps and AI model observability. The role suits engineers who enjoy improving reliability and automating complex production environments.

Highlights

High-compensation full-time role focused on cloud reliability, DevOps automation, observability, and AI/MLOps infrastructure. The position provides broad exposure to major cloud, containerization, infrastructure-as-code, and monitoring technologies.

Description

Greetings !!!! Role: SRE AWS DevOps with Arize AI Location: Malvern PA Duration: Full time Salary: $170k/annum+benefits Job Description: Must Have Technical/Functional Skills For an SRE / AWS DevOps Engineer with Arize Observability, the candidate should have knowledge on AWS Cloud, DevOps, Reliability Engineering, Monitoring, MLOps, and AI Observability.AWS Cloud Services – EC2, EKS, ECS, Lambda, S3, RDS, Redshift, CloudWatch, IAM, VPC.Infrastructure as Code (IaC) – Terraform, AWS CloudFormation, Ansible.CI/CD Automation – Jenkins, GitHub Actions, GitLab CI/CD, AWS CodePipeline.Containerization & Orchestration – Docker, Kubernetes (EKS), Helm.Site Reliability Engineering (SRE) – SLI/SLO/SLA management, incident response, root cause analysis (RCA), reliability engineering.Monitoring & Observability – Arize AI, Prometheus, Grafana, Datadog, ELK Stack, OpenTelemetry, CloudWatch.MLOps & AI Observability – Arize platform, model monitoring, drift detection, model performance tracking, data quality monitoring, LLM observability.Programming & Scripting – Python, Bash, PowerShell, SQL.Security & DevSecOps – IAM, Secrets Manager, AWS Security Hub, vulnerability scanning, policy enforcement.Collaboration & Agile Delivery – Scrum, Jira, stakeholder communication, cross-functional incident management, technical documentation. Roles & Responsibilities SRE Lead with 8+ years of experience in building and managing scalable, secure, and highly available AWS cloud platforms leveraging DevOps, Kubernetes, and Infrastructure-as-Code practices. Β· Experienced in implementing CI/CD pipelines, driving SRE best practices, and ensuring platform reliability through proactive monitoring, incident management, and performance optimization using CloudWatch, Prometheus, Grafana, and Arize. Β· Strong collaborator with engineering, data, and ML teams, enabling MLOps, AI model observability, drift detection, and reliable deployment of production-grade AI/GenAI solutions. Generic Managerial Skills, If any β€’ Strong leadership and stakeholder management skills, with experience leading cross-functional teams and driving delivery excellence. β€’ Effective communication with business and technical stakeholders. Suneetha K Team Lead Email: suneetha.k@saranshinc.com 6096590999 Ext 322 Website: www.saranshinc.com Email Is the best way to reach me