Description
Help power the infrastructure behind large-scale data, analytics, and machine-learning systems.
If you enjoy Kubernetes, automation, and solving complex production problems, this could be the next move worth exploring!
Standardized Job Description:
Optomi, in partnership with a leading technology organization, is seeking a Data Operations Engineer in Austin, Texas.
Position Summary:
This is a hybrid opportunity in Austin, requiring three days onsite and two days remote.
This is an infrastructure-focused role responsible for maintaining, scaling, and troubleshooting the platforms that support enterprise data, analytics, and machine-learning systems across batch, real-time, and hybrid-cloud environments.
This is not a traditional Data Engineer position focused primarily on building ETL pipelines, transforming data, or developing analytics solutions.
We are specifically seeking a Data Operations Engineer with an SRE, infrastructure, platform, or DevOps background who can keep the underlying environment healthy, scalable, monitored, and reliable.
Big Must-Haves:
Strong infrastructure or Site Reliability Engineering experienceStrong hands-on Linux administration and troubleshooting experienceHands-on Kubernetes administration, support, and scaling experienceBasic knowledge of Apache Spark and how Spark environments operate
Additional Experience:
AWS or hybrid-cloud infrastructureAirflow for alerting, health checks, monitoring, or cost reportingCI/CD and continuous-delivery tools such as GitLab or SpinnakerDocker and Kubernetes observabilitySplunk, Grafana, Datadog, or similar monitoring platformsTerraform or other Infrastructure-as-Code toolsAutomation using Python, Go, Rust, Java, Scala, or a similar language
Main Responsibilities:
Provision, maintain, and scale data, analytics, and ML infrastructureAdminister Kubernetes environments and manage platform capacityMonitor overall platform health and investigate infrastructure or job failuresSupport Apache Spark environments from an operational perspectiveTroubleshoot production incidents and participate in post-incident reviewsBuild automation and self-service tools that improve engineering efficiencyDevelop monitoring, alerting, observability, and self-healing solutionsSupport reliable, zero-downtime deployments through CI/CD practicesImprove infrastructure reliability, security, performance, and cost efficiency
What the Right Candidate Will Enjoy:
Supporting modern data, analytics, and machine-learning infrastructureWorking hands-on with Linux, Kubernetes, AWS, Airflow, and Apache SparkScaling environments by managing clusters, workers, nodes, and capacityTroubleshooting complex production and distributed-system issuesBuilding automation, self-service tools, and self-healing capabilitiesImproving observability, alerting, platform reliability, and operational efficiencyCollaborating with infrastructure, platform, product, engineering, and data science teamsTaking ownership and making meaningful technical decisions with limited direction