Description
About xAQUA:
xAQUA is building a next-generation, AI-native unified data management platform that brings together enterprise data integration, governance, quality, analytics, semantic intelligence, and generative AI capabilities.
We are looking for a highly experienced Senior Cloud DevOps Engineer to design, automate, secure, and operate the cloud infrastructure supporting the xAQUA platform and its customer deployments.
This is an opportunity to work on a modern, cloud-native platform serving complex enterprise and government environments.
Role Overview
The Senior Cloud DevOps Engineer will own cloud architecture, infrastructure automation, Kubernetes operations, CI/CD, observability, security, release engineering, and platform reliability across Google Cloud Platform and AWS.
Strong hands-on expertise in both GCP and AWS is required.
Experience with Microsoft Azure is highly preferred.
This is a senior individual-contributor role requiring strong technical judgment, hands-on implementation ability, ownership, analytical problem-solving, and the ability to guide other engineers.
Key Responsibilities:
Cloud Architecture and Infrastructure
- Design, implement, and manage secure, scalable, highly available cloud infrastructure across GCP and AWS, supporting private, customer-managed, dedicated, and hybrid deployment models.
- Design networking, IAM, compute, storage, database, container, and security architectures across development, QA, staging, production, and customer environments.
- Establish cloud architecture standards, reusable patterns, and operational guardrails; review designs for reliability, security, scalability, performance, and cost.
Kubernetes and Container Platforms
- Deploy and operate Kubernetes platforms using GKE and Amazon EKS, managing cluster configuration, node pools, autoscaling, ingress, networking, storage, secrets, certificates, and workload identity.
- Package and deploy services using Docker and Helm; implement secure workload isolation and environment-specific configuration.
- Troubleshoot Kubernetes networking, scheduling, resource, storage, and deployment issues; improve cluster resilience, availability, performance, and cost efficiency.
Infrastructure as Code
- Build and maintain reusable Infrastructure-as-Code modules using Terraform to automate cloud environments, networking, Kubernetes clusters, databases, storage, IAM, monitoring, and security controls.
- Implement proper state management, module versioning, code review, testing, and promotion practices to prevent unmanaged changes and environment drift.
- Maintain complete infrastructure documentation and deployment runbooks.
CI/CD and Release Engineering
- Design and maintain secure CI/CD pipelines for application, database, infrastructure, and configuration releases, automating build, test, security scanning, packaging, deployment, verification, and rollback.
- Support release promotion across development, QA, staging, production, and customer environments, integrating database migrations, reference-data changes, and configuration validation.
- Establish immutable release-artifact and environment-promotion practices to improve release frequency, predictability, traceability, and recovery.
Security and DevSecOps
- Apply security controls throughout cloud infrastructure and software-delivery pipelines, including least-privilege IAM, workload identity, secrets management, encryption, network controls, vulnerability scanning, and audit logging.
- Support private VPC deployments, VPNs, private endpoints, firewalls, load balancers, DNS, and secure customer connectivity.
- Automate container, dependency, infrastructure, and configuration security scanning; support enterprise and government security, audit, and compliance requirements.
- Participate in security reviews, incident response, vulnerability remediation, and risk reduction.
Reliability, Monitoring, and Operations
- Build observability using cloud-native and open-source logging, monitoring, tracing, and alerting tools; define SLIs, SLOs, dashboards, and actionable alerts.
- Monitor infrastructure, Kubernetes, applications, databases, AI services, and deployment pipelines; lead investigation and resolution of complex production issues.
- Conduct root-cause analysis and implement corrective actions; develop runbooks, disaster-recovery procedures, backup processes, and business-continuity controls.
- Reduce manual operational effort through automation and self-service capabilities.
AI and Data Platform Infrastructure
- Support infrastructure for AI-native services, LLM applications, RAG pipelines, vector databases, data-processing workloads, and model-serving environments, managing CPU- and GPU-based workloads.
- Support PostgreSQL, object storage, distributed data processing, search, caching, and messaging services; optimize infrastructure for high-volume data ingestion, processing, analytics, and AI workloads.
- Implement secure connectivity between data sources, platform services, model endpoints, and customer environments; collaborate with AI, backend, data, quality, and product engineers to productionize new capabilities.
Leadership and Collaboration
- Provide technical leadership and mentoring to DevOps and product engineers; review infrastructure code, architecture, deployment plans, and operational procedures.
- Collaborate with engineering, quality, security, product operations, and customer-delivery teams; participate in architecture reviews, sprint planning, release planning, incident reviews, and go/no-go decisions.
- Communicate technical risks, tradeoffs, dependencies, and recommendations clearly; take end-to-end ownership of platform reliability and cloud-delivery outcomes.
Required Qualifications
- 8+ years in DevOps, Cloud Engineering, SRE, Platform Engineering, or related roles, including 4+ years designing and operating production cloud environments.
- Advanced practical expertise in both GCP and AWS, with strong hands-on experience in GKE, Amazon EKS, Kubernetes, Docker, Helm, Terraform, ArgoCD, Linux, shell scripting, and Python or another automation language.
- Strong knowledge of cloud networking, IAM and workload identity, load balancing, DNS, private connectivity, firewalls, secrets and certificate management, cloud storage, and managed databases.
- Experience operating PostgreSQL and database-deployment processes, implementing logging, metrics, tracing, dashboards, and alerting.
- Experience with high availability, backup, recovery, disaster recovery, and incident response; strong DevSecOps and cloud-security knowledge.
- Demonstrated experience troubleshooting complex distributed-system and cloud-environment problems; excellent analytical, communication, and documentation skills.
- Ability to work independently, make sound technical decisions, and own outcomes from design through production operation.
Required Cloud Certifications:
Candidates must hold current, professional-level or advanced certifications in both GCP and AWS, such as Google Cloud Professional Cloud Architect, Google Cloud Professional Cloud DevOps Engineer, AWS Certified Solutions Architect – Professional, or AWS Certified DevOps Engineer – Professional.
Equivalent advanced certifications may be considered.
Preferred Qualifications:
- Microsoft Azure architecture or DevOps experience, with a relevant certification (e.g., Azure Solutions Architect Expert or DevOps Engineer Expert).
- Experience operating workloads across multiple cloud providers, and with GPU infrastructure and AI/ML workloads.
- Experience with LLM serving, RAG platforms, vector databases, AI observability, PostgreSQL, pgvector, object storage, and data-platform infrastructure.
- Experience with GitHub Actions, GitLab CI, Jenkins, Argo CD, or GitOps deployment practices.
- Experience with service mesh, policy-as-code, Kubernetes security controls, and Flyway or another database migration framework.
- Experience with SOC 2, government, healthcare, financial, or other regulated environments, and supporting private-cloud or customer-hosted deployments.
- Experience with FinOps, capacity planning, and cloud-cost optimization.
What Success Looks Like (First 6 Months):
- Establish clear, reusable GCP and AWS infrastructure standards; improve Kubernetes environment reliability and consistency.
- Strengthen CI/CD, release promotion, database migration, and deployment-verification processes; reduce manual infrastructure and release activities.
- Improve monitoring, alerting, incident response, and operational documentation; eliminate avoidable environment drift and configuration inconsistencies.
- Strengthen security and compliance controls; mentor existing engineers and raise the overall maturity of the DevOps function.
What We Look For:
Engineers who solve problems analytically, automate repeatable work, treat infrastructure and documentation as code, design for failure and observability, understand both architecture and hands-on implementation, communicate risks early, collaborate respectfully, use AI tools effectively while maintaining human accountability, and are comfortable with significant ownership in a fast-moving product company.
Why Join xAQUA:
Work on an AI-native, next-generation enterprise data management platform, designing infrastructure for modern data, analytics, and generative AI workloads across GCP, AWS, Kubernetes, PostgreSQL, and AI services.
Influence cloud architecture and DevOps standards from the ground up, solve complex enterprise and government deployment challenges, and take meaningful technical ownership in a growing product organization.
Employment Type:
Full-time
Application:
Please submit your résumé with details of your GCP and AWS certifications, a brief description of the largest cloud platform you have designed or operated, examples of Kubernetes, Terraform, CI/CD, security, and production-reliability work, and details of any Azure, AI/ML infrastructure, or regulated-environment experience.
Apply: careers@xaqua.io