Description
QuanXAI is an AI technology firm delivering intelligent, scalable, and secure automation
solutions for enterprises.
We combine modern software, cloud infrastructure, and industrial data
systems to build reliable solutions for real-world business and manufacturing environments.
As our AI applications and customer deployments grow, we are looking for a Senior DevOps &
Cloud Infrastructure Engineer to build and operate the infrastructure behind our products and
services.
You will focus on AWS, Kubernetes, Terraform, CI/CD, automation, security, and
reliability, working closely with engineering teams to ensure our systems are scalable, secure,
and reliable in production.
Responsibilities
[Be specific when describing each of the responsibilities.
Use gender-neutral, inclusive language.]
Example: Determine and develop user requirements for systems in production, to ensure maximum usability
Qualifications
As a Senior DevOps & Cloud Infrastructure Engineer, you will own the infrastructure and
platform operations lifecycle end-to-end and take on the following:
1.
Cloud Infrastructure & Platform Operations (40%)
AWS Infrastructure: Design, provision, and operate production infrastructure on AWS,
including compute, networking, storage, databases, load balancing, IAM, monitoring,
and related services.
Infrastructure Architecture: Design secure, scalable, highly available infrastructure
architectures for internal platforms, applications, and customer deployments.
Environment Management: Build and maintain development, staging, and production
environments with consistent configuration and deployment practices.
Infrastructure as Code: Use Terraform to provision and manage infrastructure through
version-controlled, reviewable, and repeatable workflows.
Capacity & Cost Management: Monitor resource utilization, plan infrastructure
capacity, identify bottlenecks, and optimize AWS infrastructure cost without
compromising reliability.
Security: Implement appropriate IAM, network security, secrets management,
certificates, access controls, and infrastructure hardening practices.
2.
Container & Kubernetes Engineering (20%)
Containerization: Build and maintain Docker-based application environments and
standardized container workflows.
Kubernetes: Deploy, operate, and troubleshoot Kubernetes clusters and workloads
across development, staging, and production environments.●
Workload Management: Manage deployments, services, ingress, configuration,
secrets, storage, resource limits, and scaling policies.
Cluster Reliability: Monitor cluster health and resolve issues involving scheduling,
networking, storage, resource utilization, and application availability.
Deployment Standards: Establish reusable Kubernetes patterns, Helm charts,
templates, and deployment practices for engineering teams.
Performance & Scaling: Optimize workloads and infrastructure for performance,
availability, scalability, and efficient resource utilization.
3.
CI/CD, Automation & Reliability (30%)
CI/CD: Build and maintain automated pipelines for application, infrastructure, and
service deployment.
Release Engineering: Establish reliable build, test, deployment, rollback, and release
procedures.
Automation: Identify repetitive operational tasks and replace them with sustainable
automation and engineering solutions.
Observability: Implement monitoring, logging, metrics, dashboards, alerting, and tracing
across infrastructure and applications.
Incident Response: Participate in production incident response, troubleshooting,
root-cause analysis, and post-incident improvements.
Reliability Engineering: Define and improve service-level objectives, operational
standards, availability, recovery procedures, and system resilience.
Documentation: Create infrastructure documentation, architecture diagrams,
deployment procedures, troubleshooting guides, and operational runbooks.
4.
Engineering Collaboration & Enablement (10%)
Work closely with software engineering teams to improve application architecture,
deployment processes, and production reliability.
Support development teams in diagnosing infrastructure and deployment issues.
Participate in system design reviews, infrastructure planning, and technical architecture
discussions.
Establish reusable infrastructure patterns and DevOps standards across projects.
Mentor engineers on cloud infrastructure, CI/CD, containers, Kubernetes, and
operational best practices.
Help teams balance development velocity with security, reliability, maintainability, and
operational risk.
Background / Experiences
Engineering Foundation: B.Eng.
or B.Sc.
in Computer Engineering, Computer
Science, Information Technology, or a related field — with a solid understanding of
operating systems, networking, distributed systems, and system design.
Production Infrastructure Experience: 5+ years of hands-on experience designing,
deploying, and operating production infrastructure in DevOps, Cloud Infrastructure, SRE,
Platform Engineering, or related roles.
AWS Experience: Strong practical experience operating production workloads on AWS
and understanding how to design secure and reliable cloud architectures.
Infrastructure as Code: Strong experience with Terraform and version-controlled
infrastructure management.
Kubernetes & Containers: Hands-on experience operating Kubernetes and Docker in
production environments.
Automation: Strong scripting or programming ability using Python, Bash, Go, or similar
languages, with an engineering mindset toward automation rather than manual
operations.
Problem Solving: Demonstrated ability to troubleshoot complex production
infrastructure issues, identify root causes, and implement long-term solutions.
Engineering Mindset: Comfortable working across infrastructure and application
boundaries and understanding how software architecture affects production operations.
Ownership: Able to take ownership of infrastructure problems from investigation through
implementation, deployment, documentation, and ongoing operation.
Knowledge & Skills
Cloud: Strong knowledge of AWS services including EC2, VPC, IAM, S3, RDS, EBS,
ELB/ALB, Route 53, CloudWatch, and related services.
Containers & Orchestration: Docker, Kubernetes, Helm, container networking, storage,
resource management, and workload scheduling.
Infrastructure as Code: Terraform, reusable modules, state management, environment
management, and infrastructure lifecycle management.
CI/CD: GitHub Actions, GitLab CI/CD, Jenkins, Argo CD, or equivalent technologies.
Linux: Strong Linux administration, system troubleshooting, process management,
networking, storage, and performance analysis.
Networking: TCP/IP, DNS, HTTP/HTTPS, TLS, routing, load balancing, firewalls, VPNs,
and cloud networking.
Observability: Prometheus, Grafana, CloudWatch, OpenTelemetry, centralized logging,
metrics, tracing, and alerting.
Security: AWS IAM, least-privilege access, secrets management, certificates, network
security, vulnerability management, and infrastructure hardening.
Version Control: Git for daily infrastructure development, code reviews, branching
strategies, and collaborative engineering workflows.
Databases & Services: Working knowledge of PostgreSQL, Redis, S3, queues, caches,
and other common production application dependencies.
Reliability: Understanding of availability, scalability, fault tolerance, backup, disaster
recovery, incident management, and service-level objectives.
Scripting & Programming: programming language for infrastructure automation and
tooling.
Drop your CV: career@quanxai.com