Summary
✨ AI‑Generated
Join an experienced DevOps/SRE team responsible for designing and operating highly available, scalable cloud infrastructure. You will work extensively with AWS, Kubernetes, infrastructure automation, CI/CD, and reliability engineering while partnering closely with development, platform, and operations teams.
Highlights
Full-time onsite DevOps/SRE role focused on highly available and scalable cloud infrastructure. The position emphasizes AWS, Kubernetes, automation, CI/CD, reliability, performance, and close collaboration with development and operations teams.
Description
Role: Couchbase DevOps Engineer
Location: Austin, TX - Onsite
Job Type: Fulltime / W2
We need someone strong SOFTWARE engineer someone from development background.
Job Description:
We are seeking an experienced DevOps / Site Reliability Engineer (SRE) to design, build, and operate highly available, scalable, and reliable cloud infrastructure and services.
The ideal candidate will bring strong expertise in AWS, Kubernetes, infrastructure automation, CI/CD, and operational excellence, with a focus on improving reliability, automation, and system performance.
This role requires close collaboration with development, platform, and operations teams to ensure resilient production environments, streamline deployments, and implement SRE best practices across the technology landscape.
Key Responsibilities
Infrastructure & Platform Engineering
Design, implement, and maintain highly available, scalable, and secure cloud infrastructure and services.Manage and optimize AWS-based environments to support enterprise-scale applications.Implement Infrastructure as Code (IaC) using Terraform, CloudFormation, or equivalent tools.Support capacity planning and infrastructure readiness for high-volume business events.DevOps & Automation
Develop, maintain, and enhance CI/CD pipelines for application and infrastructure deployments.Automate provisioning, configuration management, patching, upgrades, and release processes.Identify operational inefficiencies and develop automation solutions to reduce manual effort and operational toil.Collaborate with application teams to improve deployment strategies and release reliability.Site Reliability Engineering (SRE)
Define and implement SRE practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and ErrorBudgets.Establish reliability metrics and continuously improve service availability and performance.Develop and maintain operational runbooks, playbooks, and standard operating procedures.Implement proactive monitoring, observability, alerting, logging, and performance management solutions.Incident & Production Operations
Participate in production support, incident response, troubleshooting, and root cause analysis (RCA).Drive preventive and corrective actions to improve platform reliability and resilience.Support critical production events and ensure operational readiness.Work closely with cross-functional teams to resolve complex infrastructure and application issues.
Required Skills & Experience
Skill & Experience
DevOps - 5-10 YearsSite Reliability Engineering (SRE) - 5-10 YearsAWS Cloud - 2-5 YearsKubernetes - 2-5 YearsLinux Administration / Linux SRE - 2-5 YearsPython Scripting & Automation - 2-5 Years
Technical Skills
Strong experience with AWS services and cloud-native architectures.Hands-on expertise with Kubernetes orchestration and containerized workloads.Experience implementing Infrastructure as Code using Terraform, CloudFormation, or similar tools.Strong understanding of DevOps practices, CI/CD pipelines, and deployment automation.Proficiency in Python scripting for automation, monitoring, and operational tooling.Solid understanding of Linux system administration and troubleshooting.Experience with monitoring and observability platforms such as Prometheus, Grafana, Datadog, CloudWatch, ELK, or Splunk.Knowledge of source control and automation tools such as Git, Jenkins, GitHub Actions, GitLab CI/CD, or ArgoCD.Experience with incident management, root cause analysis, and operational excellence practices.Preferred Qualifications
Experience supporting large-scale, mission-critical production environments.Knowledge of security best practices in cloud-native environments.Familiarity with container security, networking, and service mesh technologies.Experience working in Agile and DevSecOps environments.AWS, Kubernetes, Terraform, or Cloud certification(s) preferred.Key Competencies
Problem-solving and analytical thinkingAutomation-first mindsetStrong ownership and accountabilityExcellent collaboration and communication skillsFocus on reliability, scalability, and operational excellenceAbility to perform effectively in high-pressure production environments