Senior DevOps / Cloud Infrastructure Engineer

Quanxai — Thailand · Posted ~2 hours ago

Senior

Skills

AWS Kubernetes Terraform CI/CD automation cloud infrastructure security reliability

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior DevOps and cloud infrastructure role focused on building and operating secure, scalable production infrastructure. You’ll work extensively with AWS, Kubernetes, Terraform, CI/CD, automation, and reliability while partnering with engineering teams to support demanding AI and enterprise workloads.

Highlights

Own cloud infrastructure for scalable production systems, work on automation and security, and collaborate closely with engineering teams in an AI-focused environment.

Description

QuanXAI is an AI technology firm delivering intelligent, scalable, and secure automation solutions for enterprises. We combine modern software, cloud infrastructure, and industrial data systems to build reliable solutions for real-world business and manufacturing environments. As our AI applications and customer deployments grow, we are looking for a Senior DevOps & Cloud Infrastructure Engineer to build and operate the infrastructure behind our products and services. You will focus on AWS, Kubernetes, Terraform, CI/CD, automation, security, and reliability, working closely with engineering teams to ensure our systems are scalable, secure, and reliable in production. Responsibilities [Be specific when describing each of the responsibilities. Use gender-neutral, inclusive language.] Example: Determine and develop user requirements for systems in production, to ensure maximum usability Qualifications As a Senior DevOps & Cloud Infrastructure Engineer, you will own the infrastructure and platform operations lifecycle end-to-end and take on the following: 1. Cloud Infrastructure & Platform Operations (40%) AWS Infrastructure: Design, provision, and operate production infrastructure on AWS, including compute, networking, storage, databases, load balancing, IAM, monitoring, and related services. Infrastructure Architecture: Design secure, scalable, highly available infrastructure architectures for internal platforms, applications, and customer deployments. Environment Management: Build and maintain development, staging, and production environments with consistent configuration and deployment practices. Infrastructure as Code: Use Terraform to provision and manage infrastructure through version-controlled, reviewable, and repeatable workflows. Capacity & Cost Management: Monitor resource utilization, plan infrastructure capacity, identify bottlenecks, and optimize AWS infrastructure cost without compromising reliability. Security: Implement appropriate IAM, network security, secrets management, certificates, access controls, and infrastructure hardening practices. 2. Container & Kubernetes Engineering (20%) Containerization: Build and maintain Docker-based application environments and standardized container workflows. Kubernetes: Deploy, operate, and troubleshoot Kubernetes clusters and workloads across development, staging, and production environments.● Workload Management: Manage deployments, services, ingress, configuration, secrets, storage, resource limits, and scaling policies. Cluster Reliability: Monitor cluster health and resolve issues involving scheduling, networking, storage, resource utilization, and application availability. Deployment Standards: Establish reusable Kubernetes patterns, Helm charts, templates, and deployment practices for engineering teams. Performance & Scaling: Optimize workloads and infrastructure for performance, availability, scalability, and efficient resource utilization. 3. CI/CD, Automation & Reliability (30%) CI/CD: Build and maintain automated pipelines for application, infrastructure, and service deployment. Release Engineering: Establish reliable build, test, deployment, rollback, and release procedures. Automation: Identify repetitive operational tasks and replace them with sustainable automation and engineering solutions. Observability: Implement monitoring, logging, metrics, dashboards, alerting, and tracing across infrastructure and applications. Incident Response: Participate in production incident response, troubleshooting, root-cause analysis, and post-incident improvements. Reliability Engineering: Define and improve service-level objectives, operational standards, availability, recovery procedures, and system resilience. Documentation: Create infrastructure documentation, architecture diagrams, deployment procedures, troubleshooting guides, and operational runbooks. 4. Engineering Collaboration & Enablement (10%) Work closely with software engineering teams to improve application architecture, deployment processes, and production reliability. Support development teams in diagnosing infrastructure and deployment issues. Participate in system design reviews, infrastructure planning, and technical architecture discussions. Establish reusable infrastructure patterns and DevOps standards across projects. Mentor engineers on cloud infrastructure, CI/CD, containers, Kubernetes, and operational best practices. Help teams balance development velocity with security, reliability, maintainability, and operational risk. Background / Experiences Engineering Foundation: B.Eng. or B.Sc. in Computer Engineering, Computer Science, Information Technology, or a related field — with a solid understanding of operating systems, networking, distributed systems, and system design. Production Infrastructure Experience: 5+ years of hands-on experience designing, deploying, and operating production infrastructure in DevOps, Cloud Infrastructure, SRE, Platform Engineering, or related roles. AWS Experience: Strong practical experience operating production workloads on AWS and understanding how to design secure and reliable cloud architectures. Infrastructure as Code: Strong experience with Terraform and version-controlled infrastructure management. Kubernetes & Containers: Hands-on experience operating Kubernetes and Docker in production environments. Automation: Strong scripting or programming ability using Python, Bash, Go, or similar languages, with an engineering mindset toward automation rather than manual operations. Problem Solving: Demonstrated ability to troubleshoot complex production infrastructure issues, identify root causes, and implement long-term solutions. Engineering Mindset: Comfortable working across infrastructure and application boundaries and understanding how software architecture affects production operations. Ownership: Able to take ownership of infrastructure problems from investigation through implementation, deployment, documentation, and ongoing operation. Knowledge & Skills Cloud: Strong knowledge of AWS services including EC2, VPC, IAM, S3, RDS, EBS, ELB/ALB, Route 53, CloudWatch, and related services. Containers & Orchestration: Docker, Kubernetes, Helm, container networking, storage, resource management, and workload scheduling. Infrastructure as Code: Terraform, reusable modules, state management, environment management, and infrastructure lifecycle management. CI/CD: GitHub Actions, GitLab CI/CD, Jenkins, Argo CD, or equivalent technologies. Linux: Strong Linux administration, system troubleshooting, process management, networking, storage, and performance analysis. Networking: TCP/IP, DNS, HTTP/HTTPS, TLS, routing, load balancing, firewalls, VPNs, and cloud networking. Observability: Prometheus, Grafana, CloudWatch, OpenTelemetry, centralized logging, metrics, tracing, and alerting. Security: AWS IAM, least-privilege access, secrets management, certificates, network security, vulnerability management, and infrastructure hardening. Version Control: Git for daily infrastructure development, code reviews, branching strategies, and collaborative engineering workflows. Databases & Services: Working knowledge of PostgreSQL, Redis, S3, queues, caches, and other common production application dependencies. Reliability: Understanding of availability, scalability, fault tolerance, backup, disaster recovery, incident management, and service-level objectives. Scripting & Programming: programming language for infrastructure automation and tooling. Drop your CV: career@quanxai.com