Summary
✨ AI‑Generated
A cloud infrastructure company is seeking an AWS administrator and SRE to design, deploy, and maintain highly available and scalable applications on AWS. You will implement SRE principles, manage AWS services including EC2, EKS, and Lambda, build CI/CD pipelines, automate operations, and troubleshoot complex production issues.
Highlights
Deep dive into AWS ecosystem and SRE practices; build and maintain CI/CD pipelines; automate operational tasks; work on highly available and scalable cloud infrastructure; incident management and reliability engineering.
Description
AWS & Site Reliability Engineering
Design, deploy, and maintain highly available, scalable, and fault-tolerant applications and infrastructure on AWS.Implement SRE principles including SLIs, SLOs, SLAs, error budgets, availability, reliability, and capacity planning.Manage AWS services including EC2, EKS, ECS, Lambda, S3, RDS, DynamoDB, API Gateway, CloudFront, Route 53, IAM, VPC, CloudWatch, SNS, SQS, and EventBridge.Troubleshoot complex production issues involving application, infrastructure, networking, database, and cloud components.Participate in incident management, root-cause analysis, problem management, and post-incident reviews.Develop automation to reduce manual operational activities and eliminate repetitive tasks.Perform capacity planning, performance tuning, disaster recovery, and business continuity activities.DevOps & Infrastructure Automation
Build and maintain CI/CD pipelines using tools such as Jenkins, GitHub Actions, GitLab CI/CD, AWS CodePipeline, or Azure DevOps.Implement Infrastructure as Code using Terraform, CloudFormation, or AWS CDK.Automate infrastructure provisioning, configuration management, deployments, and operational processes.Implement containerized workloads using Docker and Kubernetes/Amazon EKS.Develop automated deployment strategies including blue-green, canary, and rolling deployments.Integrate security, compliance, testing, and quality checks into CI/CD pipelines.AIOps & AI-Driven Operations
Implement AIOps solutions to improve monitoring, incident detection, event correlation, root-cause analysis, and automated remediation.Leverage AI/ML and Generative AI capabilities to analyze logs, metrics, traces, alerts, and operational data.Develop intelligent alerting and anomaly-detection mechanisms to identify potential production issues before they impact customers.Build AI-assisted incident investigation and troubleshooting workflows.Integrate LLM/GenAI capabilities into SRE and DevOps workflows for automated log analysis, incident summarization, knowledge retrieval, and remediation recommendations.Develop automated runbooks and self-healing mechanisms using event-driven AWS services and AI-assisted decision making.Integrate AIOps platforms and observability tools such as Datadog, Dynatrace, New Relic, Splunk, CloudWatch, Grafana, and Prometheus.Develop or integrate AI agents/workflows that can assist with incident response, operational diagnostics, and infrastructure management.Monitor AIOps/AI solutions for accuracy, reliability, security, and operational effectiveness.Observability & Monitoring
Implement comprehensive metrics, logs, traces, dashboards, and alerting across cloud and application environments.Work with Prometheus, Grafana, CloudWatch, OpenTelemetry, Datadog, Dynatrace, Splunk, or similar observability platforms.Establish meaningful service-level indicators and operational dashboards.Implement distributed tracing and application performance monitoring.Tune alerts to reduce false positives and alert fatigue.Build proactive monitoring and predictive health checks.