AWS DevOps and Site Reliability Engineer

Carecone — Australia · Posted ~2 hours ago

Senior Other

Skills

AWS DevOps Site Reliability Engineering Cloud infrastructure Infrastructure automation CI/CD GitLab Observability Incident management Cloud migration Cloud SRE

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An experienced AWS DevOps and SRE professional is sought to build, automate, operate, and continuously improve mission-critical cloud platforms. You will develop CI/CD pipelines, strengthen observability and incident response, support modernization initiatives, and improve reliability, performance, scalability, and release quality.

Highlights

Permanent engineering role focused on cloud platforms, automation, reliability, observability, scalability, and operational excellence in a highly regulated enterprise environment.

Description

Job Title : AWS DevOps & Site Reliability Engineer (SRE) Location :Sydney Mode: Permanent Job Description: We are seeking an experienced AWS DevOps & Site Reliability Engineer (SRE) to build, automate, operate, and continuously improve cloud platforms and mission-critical applications. The role focuses on AWS cloud engineering, infrastructure automation, CI/CD, platform reliability, observability, incident management, resiliency, and operational excellence within a highly regulated enterprise environment. The successful candidate will drive automation, reliability, performance, scalability, and availability of cloud-native platforms while partnering with engineering teams to improve release quality, operational stability, and customer experience. Key Responsibilities Cloud & Platform Engineering Design, build, and maintain AWS cloud infrastructure and platform services.Develop and manage CI/CD pipelines using GitLab.Support cloud migration and platform modernization initiatives.Implement Infrastructure as Code (Terraform) and configuration management (Ansible).Automate deployments, environment provisioning, database refreshes, and operational processes.Manage secrets, certificates, access controls, and cloud security controls.Develop reusable infrastructure modules, deployment standards, and platform engineering patterns.Site Reliability Engineering Define and implement reliability engineering practices, operational standards, and platform blueprints.Establish and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs).Improve platform reliability, scalability, resilience, and fault tolerance for mission-critical applications.Drive proactive reliability improvements through automation and elimination of operational toil.Conduct capacity planning, performance tuning, and resource optimization activities.Support disaster recovery planning, backup validation, resilience testing, and business continuity initiatives.Participate in production readiness reviews and ensure operational requirements are embedded into solution designs.Observability & Monitoring Design and implement monitoring, logging, alerting, and observability frameworks.Build and maintain dashboards, health checks, metrics, and operational reporting.Enhance end-to-end system visibility using CloudWatch, Grafana, Splunk, ELK, OpenTelemetry, or similar technologies.Drive alert tuning and noise reduction to improve operational effectiveness.Establish monitoring standards across applications, infrastructure, databases, and integration components.Incident & Operational Management Lead incident triage, troubleshooting, root cause analysis, and post-incident reviews.Support production systems and provide timely resolution of infrastructure and application issues.Drive reduction of Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR).Develop operational runbooks, knowledge articles, and recovery procedures.Collaborate with engineering and support teams to implement preventative actions and continuous improvement initiatives.Participate in on-call and major incident management activities where required.DevOps & Engineering Enablement Promote DevOps, SRE, and Platform Engineering best practices.Collaborate with development teams to improve deployment reliability and automation maturity.Integrate security, observability, and operational controls into CI/CD pipelines.Support release management, change governance, and deployment strategies.Champion Infrastructure as Code, automation, and self-service platform capabilities.Mandatory Skills AWS Cloud AWS EC2, ECS/EKS, Lambda, VPC, IAM, S3, CloudWatchRDS/AuroraRoute53Secrets ManagerAWS Networking and Security ServicesDevOps & Automation GitLab CI/CDTerraformAnsibleBash and Python scriptingLinux/Unix AdministrationInfrastructure AutomationSite Reliability Engineering Production Support and Incident ManagementSLI / SLO / SLA implementation and managementReliability Engineering practicesHigh Availability and Fault-Tolerant Architecture DesignCapacity Planning and Performance OptimizationDisaster Recovery and Business ContinuityRoot Cause Analysis and Problem ManagementObservability Monitoring, Logging and AlertingCloudWatchGrafanaSplunk / ELKOpenTelemetryOperational Metrics and DashboardingPlatform & Containers Kubernetes (EKS preferred)Containerisation (Docker)Networking and Infrastructure TroubleshootingSecurity Secrets Management Solutions (Vault preferred)IAM and Access ControlsCloud Security Best PracticesPreferred Skills AWS Certified Solutions Architect / DevOps Engineer CertificationCertified Kubernetes Administrator (CKA)Experience with Aurora PostgreSQL and Oracle migration programsChaos Engineering and Resilience TestingExperience implementing SRE operating modelsService Mesh technologiesFinancial Services, Banking, Capital Markets, or other regulated industriesITIL processes including Incident, Problem, Change and Release ManagementExperience supporting large-scale cloud-native microservices platforms Interested candidates can send their updated resume to Dhivya.Natarajan@carecone.com.au or reach me @ M: 61283195529