AWS Platform Engineer / Site Reliability Engineer

Ontrac Solutions — United States · Posted ~3 hours ago

Senior Other Onsite

Skills

AWS EKS Terraform Python cloud infrastructure networking observability security CI/CD site reliability engineering Kubernetes GitLab CI/CD CloudWatch Grafana OpenTelemetry Kibana Dynatrace Apigee

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Work onsite as an experienced AWS platform and site reliability engineer, operating core cloud infrastructure and Kubernetes environments. You will automate infrastructure with Terraform, troubleshoot advanced networking, implement security controls, own observability, and improve release engineering and CI/CD processes.

Highlights

Hands-on ownership of a large AWS environment covering infrastructure automation, Kubernetes, advanced networking, security, observability, and CI/CD, with either FTE or long-term contract engagement.

Description

We are seeking an experienced AWS Platform Engineer / Site Reliability Engineer to join our client's platform team in Las Vegas, NV. This role blends release engineering, observability, and cloud infrastructure automation across a large AWS estate, with hands-on responsibility for EKS, networking, security, and monitoring tooling including Kibana, Dynatrace, and Apigee. This is an onsite role, Monday through Friday, in Las Vegas, NV. Long-term engagement; FTE or contract. What You Will Do Build and operate core AWS infrastructure: VPC, EC2, S3, IAM, Route 53, backups, EKS, SSO, MSK, Security Hub, GuardDuty, and related security/compliance tooling. Design and troubleshoot advanced cloud networking, including AWS Transit Gateway (TGW) and Direct Connect setup and hybrid connectivity. Own observability and monitoring across Amazon CloudWatch, Grafana, and OpenTelemetry (OTEL), including proactive alerting and telemetry configuration. Automate end to end with Terraform, GitLab CI/CD, and shell scripting. Operate Kubernetes (EKS) and Istio service mesh for traffic management, security, and observability. Drive patching and vulnerability remediation across cloud-native and containerized environments in line with compliance requirements. Develop Python tooling for automation, scripting, and infrastructure operations. Support release engineering and monitoring workflows across Kibana, Dynatrace, and Apigee. Required Qualifications Minimum 7+ years of relevant platform, SRE, or cloud infrastructure experience. Strong hands-on experience with core AWS services listed above. Advanced networking knowledge, including Transit Gateway and Direct Connect. Proficiency with Terraform, GitLab CI/CD, and shell scripting. Solid working knowledge of Kubernetes/EKS and Istio. Hands-on Python development for automation and infrastructure tooling. Experience with patching and vulnerability remediation in containerized environments. Problem-solving mindset and the ability to work effectively in fast-paced Agile/Scrum environments. Technical Skills AWS (VPC, EC2, S3, IAM, Route 53, EKS, SSO, MSK, Security Hub, GuardDuty), Kubernetes/EKS, Istio, Transit Gateway, Direct Connect, Terraform, GitLab CI/CD, Python, Shell, CloudWatch, Grafana, OpenTelemetry, Kibana, Dynatrace, Apigee Skills Matrix Applicants will be asked to provide years of experience and a self-rating (1-10) for each of the following: AWS infrastructure Kubernetes / EKS Networking and security Terraform Python development for scripting, automation, and infrastructure tooling AWS services: VPC, EC2, S3, IAM, Route 53, backups, EKS, SSO, MSK, Security Hub, GuardDuty and related security/compliance tooling Release engineering and monitoring (Kibana, Dynatrace, Apigee)