Summary
β¨ AIβGenerated
A full-time SRE/DevOps opportunity for an engineer experienced in AWS cloud infrastructure, automation, reliability, and modern observability. You will work across infrastructure as code, CI/CD, containers, Kubernetes, incident response, and monitoring, with additional focus on MLOps and AI model observability. The role suits engineers who enjoy improving reliability and automating complex production environments.
Highlights
High-compensation full-time role focused on cloud reliability, DevOps automation, observability, and AI/MLOps infrastructure. The position provides broad exposure to major cloud, containerization, infrastructure-as-code, and monitoring technologies.
Description
Greetings !!!!
Role: SRE AWS DevOps with Arize AI
Location: Malvern PA
Duration: Full time
Salary: $170k/annum+benefits
Job Description:
Must Have Technical/Functional Skills
For an SRE / AWS DevOps Engineer with Arize Observability, the candidate should have knowledge on AWS Cloud, DevOps, Reliability Engineering, Monitoring, MLOps, and AI Observability.AWS Cloud Services β EC2, EKS, ECS, Lambda, S3, RDS, Redshift, CloudWatch, IAM, VPC.Infrastructure as Code (IaC) β Terraform, AWS CloudFormation, Ansible.CI/CD Automation β Jenkins, GitHub Actions, GitLab CI/CD, AWS CodePipeline.Containerization & Orchestration β Docker, Kubernetes (EKS), Helm.Site Reliability Engineering (SRE) β SLI/SLO/SLA management, incident response, root cause analysis (RCA), reliability engineering.Monitoring & Observability β Arize AI, Prometheus, Grafana, Datadog, ELK Stack, OpenTelemetry, CloudWatch.MLOps & AI Observability β Arize platform, model monitoring, drift detection, model performance tracking, data quality monitoring, LLM observability.Programming & Scripting β Python, Bash, PowerShell, SQL.Security & DevSecOps β IAM, Secrets Manager, AWS Security Hub, vulnerability scanning, policy enforcement.Collaboration & Agile Delivery β Scrum, Jira, stakeholder communication, cross-functional incident management, technical documentation.
Roles & Responsibilities
SRE Lead with 8+ years of experience in building and managing scalable, secure, and highly available AWS cloud platforms leveraging DevOps, Kubernetes, and Infrastructure-as-Code practices.
Β· Experienced in implementing CI/CD pipelines, driving SRE best practices, and ensuring platform reliability through proactive monitoring, incident management, and performance optimization using CloudWatch, Prometheus, Grafana, and Arize.
Β· Strong collaborator with engineering, data, and ML teams, enabling MLOps, AI model observability, drift detection, and reliable deployment of production-grade AI/GenAI solutions.
Generic Managerial Skills, If any
β’ Strong leadership and stakeholder management skills, with experience leading cross-functional teams and driving delivery excellence.
β’ Effective communication with business and technical stakeholders.
Suneetha K
Team Lead
Email: suneetha.k@saranshinc.com
6096590999 Ext 322
Website: www.saranshinc.com
Email Is the best way to reach me