DevOps/SRE Engineer

Red Oak Technologies โ€” United States ยท Posted ~1 day ago

Full-time Hybrid

Skills

DevOps Site Reliability Engineering Python AWS GCP Infrastructure automation CI/CD Monitoring and alerting Incident response Troubleshooting Docker Kubernetes Cloud infrastructure Security best practices Jenkins Terraform Ansible DataDog Nagios Prometheus Grafana

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary

Join an engineering team responsible for deploying, monitoring, securing, and operating enterprise applications in cloud and hybrid environments. You will automate deployments and alerts, respond to incidents, troubleshoot complex infrastructure problems, and improve reliability and scalability. Strong Python skills, cloud experience, infrastructure automation, CI/CD, monitoring, containers, and Kubernetes are central to the role.

Highlights

Work in a hybrid DevOps/SRE role focused on reliable cloud infrastructure, automation, monitoring, incident response, and scalable application operations. The position offers broad collaboration with engineering and security teams and exposure to modern cloud, container, CI/CD, and AI/ML infrastructure technologies.

Description

Dev Ops/ SRE Engineer Austin, TX (Hybrid) Key Responsibilities Deployment & Integration: Collaborate with software teams to deploy and maintain existing applications and tools, ensuring smooth integration within our infrastructure. Monitoring & Signal Flow: Set up and manage monitoring solutions to track application health, data pipelines, and system signals. Ensure real-time visibility into performance and operational metrics. Incident Response & Troubleshooting: Quickly assess and respond to application outages or issues, identify root causes, and coordinate resolution efforts to minimize downtime. Automation & Optimization: Automate deployment processes, monitoring, and alerting to improve efficiency and reduce manual intervention. Collaboration: Work closely with data engineers, security teams, and software developers to ensure applications are robust, secure, and scalable. Documentation & Best Practices: Maintain comprehensive documentation of deployment processes, system architecture, and incident protocols. Qualifications Proven experience in deploying and managing enterprise software applications in a cloud or hybrid environment. Proficient in Python Knowledge cloud platforms AWS, GCP Strong understanding of infrastructure automation tools (e.g., Jenkins, Terraform, Ansible, or similar). Experience with monitoring and alerting tools (e.g., DataDog, Nagios, Prometheus, Grafana). Familiarity with data pipelines, signal flow, and system architecture related to AI/ML applications. Understanding of CI/CD pipelines and tools Experience with Containerization (Docker, Kubernetes) Skills in incident detection and response Ability to troubleshoot complex system issues quickly and effectively. Excellent collaboration and communication skills. Knowledge of security best practices in deployment and system management. Preferred Skills Experience with AI/ML infrastructure or data platforms. Familiarity with containerization and orchestration (Docker, Kubernetes). Understanding of network protocols and security.