Senior AWS DevOps & Site Reliability Engineer

Sophusitsolutions — United States · Posted ~1 hour ago

Senior Full-time Hybrid

Skills

AWS DevOps Site Reliability Engineering CI/CD Infrastructure Automation Monitoring Logging Alerting Observability Incident Response SLOs SRE Splunk Cloud Infrastructure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join a senior cloud engineering team responsible for building, scaling, and maintaining highly reliable AWS infrastructure. You will design CI/CD pipelines, automate infrastructure, establish observability and SLOs, manage incidents, improve scalability and fault tolerance, and reduce operational toil.

Highlights

Senior hybrid role combining cloud DevOps and SRE responsibilities. Focus on highly available AWS environments, robust CI/CD, infrastructure automation, observability, incident response, service-level objectives, scalability, and reducing operational toil.

Description

Please review the below JD: Role: AWS Devops& SRE Location: Hartford CT Hybrid Position Overview We are seeking a highly experienced Senior Cloud DevOps & Site Reliability Engineer to build, scale, and maintain our cloud infrastructure while ensuring maximum system reliability and uptime. In this hybrid role, you will bridge the gap between software development and IT operations. As a DevOps Engineer, you will accelerate feature delivery by designing robust CI/CD pipelines and automating infrastructure provisioning. As an SRE, you will protect the user experience by enforcing service-level objectives (SLOs), managing incident response, optimizing observability, and aggressively eliminating operational toil. Key Responsibilities Site Reliability & Observability (SRE Focus) Ensure system reliability, high availability, fault tolerance, and scalability across all production AWS environments.Implement and manage comprehensive monitoring, logging, and alerting systems using Splunk, Datadog, Prometheus, Grafana, and AWS CloudWatch.Establish Service Level Indicators (SLIs), Service Level Objectives (SLOs), and manage error budgets to balance feature velocity with platform stability.Lead incident response, perform root cause analysis (RCA), and conduct blameless post-mortems.Design and execute Business Continuity Plans (BCP) and cross-region Disaster Recovery (DR) strategies using Route 53 and optimized data replication. Continuous Delivery & Automation (DevOps Focus) Design, build, and maintain automated CI/CD pipelines using Jenkins, Git, Maven, and Bitbucket for seamless software delivery.Automate operational and infrastructure tasks using Python, Bash, Ruby, and Shell to eliminate manual toil.Manage artifact repositories (Nexus, Artifactory) and container registries (Docker Hub).Collaborate closely with development teams to support Agile operations, unit test automation, and streamline release processes. Cloud Architecture & Infrastructure as Code (IaC) Design, provision, and maintain secure, cost-optimized AWS infrastructure (EC2, VPC, S3, RDS, Auto-Scaling, Elasticache, CloudFront).Implement Infrastructure as Code (IaC) pipelines using Terraform and AWS CloudFormation.Configure configuration management workflows using Ansible, Chef, and Puppet for large-scale application deployments. Container Orchestration & Security Build, containerize, and optimize applications utilizing Docker components (Engine, Compose).Deploy, load-balance, and manage production-ready Kubernetes clusters on AWS (EKS) and leverage Helm for managing Kubernetes charts.Enforce cloud security best practices (DevSecOps), including strict IAM policies, cross-account roles, STS credentialing, MFA, Security Groups, and NACLs. Required Qualifications Experience: 7+ years of IT experience, with 3+ years specifically in DevOps, SRE, or Cloud Engineering environments.Cloud Platform: Extensive hands-on experience with Amazon Web Services (AWS) core services (compute, network, storage, database).Infrastructure & Configuration: Mastery of Terraform, CloudFormation, Ansible, Chef, and Puppet.Containerization: Proven expertise running Docker and Kubernetes (or OpenShift/Mesos) in production at scale.CI/CDTooling: Deep knowledge of Jenkins, Git, Subversion, Maven, and JIRA.Observability: Strong experience setting up and configuring Splunk, Datadog, Prometheus, Grafana, New Relic, and Nagios.Scripting: Advanced proficiency in Bash/Shell, Python, Ruby, and YAML.OS/Systems Admin: Deep understanding of Linux/Unix administration (RHEL, CentOS, Ubuntu, Solaris) and networking protocols (TCP/IP, DNS, HTTP, SSH). Regards, Mohamed Ijaz Sr. US & Canada IT Recruiter mohamed@Sophusinfo.com