SRE Engineer

Envision Tech Sol — United States · Posted ~2 hours ago

Mid Full-time Onsite

Skills

Linux AWS EC2 RDS IAM VPC CloudWatch S3 Python Shell PostgreSQL MySQL Git Jenkins Ansible Terraform Monitoring CI/CD

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Seeking an SRE Engineer with 6-8 years of experience to design, implement, and maintain highly available production systems. Responsibilities include automating operational tasks using Shell, Python, Ansible, or Terraform, managing AWS infrastructure (EC2, RDS, VPC, S3), implementing monitoring and alerting solutions, and conducting incident management and RCA processes. Ideal for engineers with strong Linux administration skills and experience with CI/CD pipelines.

Highlights

Join a team managing highly available production systems; work with automation and infrastructure as code; competitive experience in incident management and RCA processes; opportunity to work with modern cloud services and DevOps practices

Description

Role - SRE Engineer Location - Englewood Cliffs, NJ | 6 - 8 years of experience Job Description Must Have Technical/Functional Skills • 6-7 years of experience in Site Reliability Engineering, Production Support, DevOps, or Infrastructure Operations. • Strong understanding of Linux administration and troubleshooting. • Hands-on experience with AWS cloud services (EC2, RDS, IAM, VPC, CloudWatch, S3). • Experience with monitoring, alerting, and observability tools. • Knowledge of incident management, problem management, and RCA processes. • Experience with automation and scripting using Shell and/or Python. • Working knowledge of PostgreSQL and MySQL databases. • Experience with Git version control. • Understanding of CI/CD concepts and tools such as Jenkins. Roles & Responsibilities • Design, implement, and maintain highly available and reliable production systems. • Automate operational tasks and infrastructure management using Shell, Python, Ansible, or Terraform. • Manage and support AWS services including EC2, RDS, S3, IAM, VPC, CloudWatch, and related cloud services. • Perform Linux server administration, troubleshooting, patching, and performance tuning. • Monitor application and infrastructure health using tools such as Grafana, Prometheus, CloudWatch, Datadog, Splunk. • Participate in incident management, root cause analysis (RCA), and problem management activities. • Define and maintain SLIs, SLOs, and SLAs to ensure service reliability. • Support PostgreSQL and MySQL databases for operational and basic administration tasks. • Collaborate with development, QA, cloud, and support teams to improve system reliability and deployment processes. • Drive automation, observability, capacity planning, security, and operational best practices. • Participate in on-call support and production issue resolution.