Summary
✨ AI‑Generated
An experienced reliability engineer is sought to design, automate, and maintain secure cloud platforms. The role involves improving availability, scalability, monitoring, deployments, and operational excellence in complex technical environments.
Highlights
Opportunity to work on large-scale cloud infrastructure with a focus on reliability, automation, security, and production engineering.
Description
Site Reliability Engineer (SRE) – AWS / Cloud Infrastructure
US Remote (onsite once a month)
MUST BE US CITIZEN
$140,000 - $160,000
We’re looking for an experienced Site Reliability Engineer (SRE) to help build, operate and improve secure, scalable and highly available cloud infrastructure.
This is a hands-on engineering role focused on AWS, infrastructure automation, reliability, networking and security, supporting business-critical applications and platforms within complex enterprise environments.
Key Responsibilities
Build, manage and optimize AWS cloud infrastructure across production environments.Improve platform reliability, availability, scalability and performance.Automate infrastructure provisioning and configuration using Infrastructure as Code.Build and maintain CI/CD pipelines and deployment automation.Implement monitoring, logging, alerting and observability across applications and infrastructure.Troubleshoot complex production, infrastructure and networking issues.Support incident response, root-cause analysis and preventative engineering.Implement cloud security, IAM, networking and infrastructure best practices.Work closely with software, platform, security and DevOps teams to improve engineering standards and operational resilience.
What We’re Looking For
Strong experience in Site Reliability Engineering, DevOps, Platform Engineering or Cloud Infrastructure.Deep hands-on experience with AWS.Strong knowledge of Linux, networking, DNS, load balancing and cloud security.Experience with Terraform, CloudFormation or similar Infrastructure-as-Code tooling.Experience with Docker and Kubernetes/EKS.Strong scripting/automation skills using Python, Bash or similar.Experience with CI/CD tooling such as GitHub Actions, GitLab CI, Jenkins or equivalent.Knowledge of monitoring and observability platforms such as CloudWatch, Datadog, Prometheus or Grafana.Understanding of DevSecOps, IAM and security best practices.Strong troubleshooting skills within complex, distributed environments.
Nice to Have
Experience supporting enterprise, manufacturing or industrial environments.AWS certifications.Experience with cloud-native application architectures.Exposure to infrastructure supporting AI/ML workloads or data platforms.