Summary
✨ AI‑Generated
A technology organization is seeking a cloud platform leader to manage AWS operations, improve infrastructure reliability, guide engineers, and handle critical production issues.
Highlights
Leadership opportunity overseeing cloud operations, reliability, automation, and a technical support team in an enterprise environment.
Description
Job Title: AWS Platform Lead
Location: Rahway, NJ (onsite)
Duration: 6+ Months
Job Summary-
We are seeking an experienced AWS Platform Support Lead to oversee cloud platform operations, production support, incident management, and infrastructure reliability across AWS environments.
The ideal candidate will lead a team of support engineers, drive operational excellence, ensure platform stability and security, and act as the primary escalation point for critical incidents.
This role requires strong expertise in AWS cloud services, infrastructure operations, automation, monitoring, and ITIL-based service management processes.
Key Responsibilities
Technical Leadership
• Lead and mentor a team of AWS Platform Support Engineers.
• Provide technical guidance for cloud operations, support activities, and platform improvements.
• Define support standards, operational procedures, and best practices.
• Serve as the technical escalation point for complex production issues.
Platform Operations & Support
• Manage and support AWS infrastructure across multiple environments.
• Ensure high availability, scalability, performance, and reliability of cloud platforms.
• Oversee incident, problem, change, and release management activities.
• Drive major incident management and coordinate recovery efforts.
• Perform Root Cause Analysis (RCA) and implement preventive actions.
Cloud Infrastructure Management
Manage and support AWS services including:
• EC2
• VPC
• IAM
• S3
• RDS
• Lambda
• EKS / ECS
• Route 53
• CloudWatch
• ELB / ALB
• SNS / SQS
• Auto Scaling
Monitoring & Reliability
• Establish proactive monitoring and alerting mechanisms.
• Define and track platform SLAs, KPIs, and operational metrics.
• Analyze trends and recommend improvements for platform stability.
• Ensure capacity planning and performance optimization.
Automation & DevOps
• Drive automation of operational activities using Infrastructure as Code (IaC).
• Support CI/CD processes and release deployments.
• Implement operational automation using:
o Terraform
o CloudFormation
o Ansible
o Python
o Shell Scripting
Security & Compliance
• Ensure cloud environments adhere to security policies and compliance requirements.
• Review IAM configurations, access controls, and overall security posture.
• Collaborate with security teams to remediate vulnerabilities and audit findings.
Stakeholder Management
• Partner with application, DevOps, security, network, and business teams.
• Provide regular operational reports and service review updates.
• Lead customer and stakeholder communications during major incidents and service reviews.
Required Technical Skills
AWS Services
• EC2
• S3
• VPC
• IAM
• RDS
• Lambda
• Route 53
• ELB / ALB / NLB
• CloudWatch
• CloudTrail
• Auto Scaling
• SNS / SQS
• EKS / ECS
Operating Systems
• Linux (RHEL, Amazon Linux, Ubuntu)
• Windows Server
Monitoring & Observability Tools
• CloudWatch
• Splunk
• Datadog
• Dynatrace
• Prometheus
• Grafana
• ELK Stack
Infrastructure as Code (IaC)
• Terraform
• CloudFormation
• Ansible
DevOps Tools
• Jenkins
• GitHub
• GitLab
• Azure DevOps
Scripting & Automation
• Python
• Shell Scripting
• PowerShell
Database Technologies
• Amazon RDS
• MySQL
• PostgreSQL
• DynamoDB
Leadership Responsibilities
• Manage and coach a team of cloud support engineers.
• Conduct performance reviews and development planning.
• Drive shift governance and support coverage planning.
• Ensure adherence to operational SLAs and OLAs.
• Lead major incident bridges and stakeholder communications.
• Drive continuous service improvement initiatives.
Required Qualifications
• Bachelor's degree in Computer Science, Information Technology, or a related discipline.
• 8+ years of IT infrastructure and production support experience.
• 5+ years of hands-on AWS cloud operations experience.
• Experience leading platform support teams in a production environment.
• Strong understanding of cloud networking, security, and infrastructure architecture.
• Experience working within ITIL-based service management environments.
• Excellent communication, leadership, and stakeholder management skills.
Preferred Certifications
• AWS Certified Solutions Architect – Professional
• AWS Certified SysOps Administrator – Associate
• AWS Certified DevOps Engineer – Professional
• AWS Certified Security – Specialty
• ITIL Foundation Certification
Key Competencies
• AWS Cloud Operations Leadership
• Production Support Management
• Incident & Problem Management
• Platform Reliability Engineering
• Site Reliability Engineering (SRE)
• Team Leadership & Mentoring
• Stakeholder Management
• Automation & Continuous Improvement
• Cloud Security & Governance
• Performance Optimization
Nice to Have
• Experience with Kubernetes (EKS) and container platforms.
• Exposure to FinOps and AWS cost optimization practices.
• Multi-cloud experience (Azure and/or GCP).
• Knowledge of SRE practices and observability frameworks.
• Experience supporting global 24x7 operations.
Work Authorization
Applicants must be legally authorized to work in the United States.
Equal Employment Opportunity
Themesoft Inc is an Equal Employment Opportunity employer.
We consider qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other status protected by applicable federal, state, or local law.