DevOps Cloud Engineer

Glansa Associates โ€” United States ยท Posted ~3 hours ago

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Description

Job Description/ Responsibilities * Design, implement, operate, and support highly available and scalable infrastructure across AliCloud and AWS environments. * Apply SRE principles to improve availability, latency, performance, capacity, scalability, and operational efficiency of production services. * Define, track, and continuously improve SLIs, SLOs, and operational reliability metrics for critical services. * Build and maintain automated CI/CD pipelines for application and infrastructure deployments using modern DevOps practices. * Automate infrastructure provisioning, configuration, deployments, health checks, maintenance, and repetitive operational activities using scripting and Infrastructure as Code. * Develop automation using Python, Bash/Shell, or equivalent scripting languages to eliminate manual effort and improve operational consistency. * Administer and troubleshoot Linux-based production systems, including OS, processes, services, storage, networking, permissions, patching, system performance, and security. * Support cloud services covering compute, storage, networking, IAM/security, load balancing, monitoring, logging, backup, and disaster recovery. * Implement and maintain Infrastructure as Code using tools such as Terraform and configuration automation using tools such as Ansible. * Work with containerized and cloud-native platforms such as Docker and Kubernetes, including deployment, scaling, troubleshooting, and operational support. * Implement comprehensive monitoring, logging, alerting, and dashboards using tools such as CloudWatch, Prometheus, Grafana, ELK/Splunk, or comparable platforms. * Participate in production incident response, perform systematic troubleshooting and root-cause analysis, and drive corrective and preventive actions through blameless post-incident reviews. * Develop and maintain operational runbooks, automation playbooks, recovery procedures, and technical documentation. * Identify recurring operational issues and eliminate toil through engineering and automation. * Support capacity planning, performance tuning, high availability, failover, backup/recovery, and disaster-recovery readiness. * Collaborate closely with development, infrastructure, security, network, database, and application support teams to improve end-to-end platform reliability. * Promote secure DevOps/SRE practices including least-privilege access, secrets management, vulnerability remediation, patching, and compliance controls. * Participate in change, release, incident, and problem-management processes and support production environments as required.