Senior SRE / AWS DevOps Engineer

Envision Tech Sol β€” United States Β· Posted ~3 hours ago

Senior Full-time Onsite

Skills

AWS DevOps Site Reliability Engineering Terraform AWS CloudFormation Ansible Jenkins GitHub Actions GitLab CI/CD AWS CodePipeline Docker Kubernetes Helm Python SLI/SLO/SLA management incident response root cause analysis monitoring and observability MLOps AI observability Arize AI Prometheus Grafana Datadog ELK Stack OpenTelemetry EC2 EKS ECS Lambda S3 RDS Redshift CloudWatch IAM VPC

πŸ”“ Log in to save this job, tailor your resume & track your apply process β€” 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An established technology organization is seeking a senior SRE/DevOps engineer to build and operate reliable cloud infrastructure and advanced observability capabilities. The role combines AWS, infrastructure as code, CI/CD, containers, Kubernetes, incident management, monitoring, MLOps, and AI observability. You will work with modern tooling to improve system reliability and monitor machine-learning workloads.

Highlights

Highly technical senior role combining cloud infrastructure, reliability engineering, DevOps, observability, and MLOps. Offers broad exposure to modern cloud and AI operations technologies and challenging large-scale engineering responsibilities.

Description

Job Title: SRE AWS DevOps with Arize Job Location: Malvern, PA-Onsite | 12 - 18 years of experience Job Description For an SRE / AWS DevOps Engineer with Arize Observability, the candidate should have knowledge on AWS Cloud, DevOps, Reliability Engineering, Monitoring, MLOps, and AI Observability.AWS Cloud Services – EC2, EKS, ECS, Lambda, S3, RDS, Redshift, CloudWatch, IAM, VPC.Infrastructure as Code (IaC) – Terraform, AWS CloudFormation, Ansible.CI/CD Automation – Jenkins, GitHub Actions, GitLab CI/CD, AWS CodePipeline.Containerization & Orchestration – Docker, Kubernetes (EKS), Helm.Site Reliability Engineering (SRE) – SLI/SLO/SLA management, incident response, root cause analysis (RCA), reliability engineering.Monitoring & Observability – Arize AI, Prometheus, Grafana, Datadog, ELK Stack, Open Telemetry, CloudWatch.MLOps & AI Observability – Arize platform, model monitoring, drift detection, model performance tracking, data quality monitoring, LLM observability.Programming & Scripting – Python, Bash, PowerShell, SQL.Security & DevSecOps – IAM, Secrets Manager, AWS Security Hub, vulnerability scanning, policy enforcement.Collaboration & Agile Delivery – Scrum, Jira, stakeholder communication, cross-functional incident management, technical documentation. Roles & Responsibilities SRE Lead with 8+ years of experience in building and managing scalable, secure, and highly available AWS cloud platforms leveraging DevOps, Kubernetes, and Infrastructure-as-Code practices.Experienced in implementing CI/CD pipelines, driving SRE best practices, and ensuring platform reliability through proactive monitoring, incident management, and performance optimization using CloudWatch, Prometheus, Grafana, and Arize.Strong collaborator with engineering, data, and ML teams, enabling MLOps, AI model observability, drift detection, and reliable deployment of production-grade AI/GenAI solutions.