Summary
✨ AI‑Generated
An experienced DevOps/SRE Engineer is sought to design, build, and operate highly available, scalable cloud infrastructure and services. The role emphasizes AWS, Kubernetes, infrastructure automation, CI/CD, Linux, Python scripting, reliability engineering, and operational excellence. You will work closely with development, platform, and operations teams to strengthen production resilience, streamline deployments, and improve system performance. Strong software engineering experience is valued.
Highlights
Hands-on SRE/DevOps role focused on highly available cloud infrastructure, automation, reliability, and performance. Offers broad technical ownership across AWS, Kubernetes, CI/CD, Linux, and production operations while collaborating closely with engineering and platform teams.
Description
Position: Couchbase DevOps Engineer
Location: Austin, TX (Onsite)
We need someone strong SOFTWARE engineer someone from development background.
Job Description:
We are seeking an experienced DevOps / Site Reliability Engineer (SRE) to design, build, and operate highly available, scalable, and reliable cloud infrastructure and services.
The ideal candidate will bring strong expertise in AWS, Kubernetes, infrastructure automation, CI/CD, and operational excellence, with a focus on improving reliability, automation, and system performance.
This role requires close collaboration with development, platform, and operations teams to ensure resilient production environments, streamline deployments, and implement SRE best practices across the technology landscape.
Required Skills & Experience
Skill
Experience
DevOps
5-10 Years
Site Reliability Engineering (SRE)
5-10 Years
AWS Cloud
2-5 Years
Kubernetes
2-5 Years
Linux Administration / Linux SRE
2-5 Years
Python Scripting & Automation
2-5 Years
Key Responsibilities
Infrastructure & Platform Engineering
· Design, implement, and maintain highly available, scalable, and secure cloud infrastructure and services.
· Manage and optimize AWS-based environments to support enterprise-scale applications.
· Implement Infrastructure as Code (IaC) using Terraform, CloudFormation, or equivalent tools.
· Support capacity planning and infrastructure readiness for high-volume business events.
DevOps & Automation
· Develop, maintain, and enhance CI/CD pipelines for application and infrastructure deployments.
· Automate provisioning, configuration management, patching, upgrades, and release processes.
· Identify operational inefficiencies and develop automation solutions to reduce manual effort and operational toil.
· Collaborate with application teams to improve deployment strategies and release reliability.
Site Reliability Engineering (SRE)
· Define and implement SRE practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and ErrorBudgets.
· Establish reliability metrics and continuously improve service availability and performance.
· Develop and maintain operational runbooks, playbooks, and standard operating procedures.
· Implement proactive monitoring, observability, alerting, logging, and performance management solutions.
Incident & Production Operations
· Participate in production support, incident response, troubleshooting, and root cause analysis (RCA).
· Drive preventive and corrective actions to improve platform reliability and resilience.
· Support critical production events and ensure operational readiness.
· Work closely with cross-functional teams to resolve complex infrastructure and application issues.
Technical Skills
· Strong experience with AWS services and cloud-native architectures.
· Hands-on expertise with Kubernetes orchestration and containerized workloads.
· Experience implementing Infrastructure as Code using Terraform, CloudFormation, or similar tools.
· Strong understanding of DevOps practices, CI/CD pipelines, and deployment automation.
· Proficiency in Python scripting for automation, monitoring, and operational tooling.
· Solid understanding of Linux system administration and troubleshooting.
· Experience with monitoring and observability platforms such as Prometheus, Grafana, Datadog, CloudWatch, ELK, or Splunk.
· Knowledge of source control and automation tools such as Git, Jenkins, GitHub Actions, GitLab CI/CD, or ArgoCD.
· Experience with incident management, root cause analysis, and operational excellence practices.
Preferred Qualifications
· Experience supporting large-scale, mission-critical production environments.
· Knowledge of security best practices in cloud-native environments.
· Familiarity with container security, networking, and service mesh technologies.
· Experience working in Agile and DevSecOps environments.
· AWS, Kubernetes, Terraform, or Cloud certification(s) preferred.
Key Competencies
· Problem-solving and analytical thinking
· Automation-first mindset
· Strong ownership and accountability
· Excellent collaboration and communication skills
· Focus on reliability, scalability, and operational excellence
· Ability to perform effectively in high-pressure production environments