Summary
β¨ AIβGenerated
An established technology organization is seeking a senior SRE/DevOps engineer to build and operate reliable cloud infrastructure and advanced observability capabilities. The role combines AWS, infrastructure as code, CI/CD, containers, Kubernetes, incident management, monitoring, MLOps, and AI observability. You will work with modern tooling to improve system reliability and monitor machine-learning workloads.
Highlights
Highly technical senior role combining cloud infrastructure, reliability engineering, DevOps, observability, and MLOps. Offers broad exposure to modern cloud and AI operations technologies and challenging large-scale engineering responsibilities.
Description
Job Title: SRE AWS DevOps with Arize
Job Location: Malvern, PA-Onsite | 12 - 18 years of experience
Job Description
For an SRE / AWS DevOps Engineer with Arize Observability, the candidate should have knowledge on AWS Cloud, DevOps, Reliability Engineering, Monitoring, MLOps, and AI Observability.AWS Cloud Services β EC2, EKS, ECS, Lambda, S3, RDS, Redshift, CloudWatch, IAM, VPC.Infrastructure as Code (IaC) β Terraform, AWS CloudFormation, Ansible.CI/CD Automation β Jenkins, GitHub Actions, GitLab CI/CD, AWS CodePipeline.Containerization & Orchestration β Docker, Kubernetes (EKS), Helm.Site Reliability Engineering (SRE) β SLI/SLO/SLA management, incident response, root cause analysis (RCA), reliability engineering.Monitoring & Observability β Arize AI, Prometheus, Grafana, Datadog, ELK Stack, Open Telemetry, CloudWatch.MLOps & AI Observability β Arize platform, model monitoring, drift detection, model performance tracking, data quality monitoring, LLM observability.Programming & Scripting β Python, Bash, PowerShell, SQL.Security & DevSecOps β IAM, Secrets Manager, AWS Security Hub, vulnerability scanning, policy enforcement.Collaboration & Agile Delivery β Scrum, Jira, stakeholder communication, cross-functional incident management, technical documentation.
Roles & Responsibilities
SRE Lead with 8+ years of experience in building and managing scalable, secure, and highly available AWS cloud platforms leveraging DevOps, Kubernetes, and Infrastructure-as-Code practices.Experienced in implementing CI/CD pipelines, driving SRE best practices, and ensuring platform reliability through proactive monitoring, incident management, and performance optimization using CloudWatch, Prometheus, Grafana, and Arize.Strong collaborator with engineering, data, and ML teams, enabling MLOps, AI model observability, drift detection, and reliable deployment of production-grade AI/GenAI solutions.