Data Platform DevOps/SRE Engineer

Envision Tech Sol — United States · Posted ~3 hours ago

Contract Onsite

Skills

Databricks Databricks Genie BI tooling Dynatrace Splunk AWS Python PySpark CI/CD DevOps SRE automation

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Improve delivery automation, environment consistency, and operational readiness across data and analytics platforms. You will build standardized CI/CD frameworks for data-platform assets and Python workloads, with automated promotion, approval gates, security controls, rollback, and progressive delivery.

Highlights

Own automation and operational readiness across modern data and analytics platforms. The role focuses on standardized CI/CD, controlled deployments, security controls, reliable rollback, and progressive delivery while working with widely used cloud and data technologies.

Description

Tittle: Data Platform DevOps/SRE Engineer Location 1: Jersey City, NJ (5 Days Onsite/Week) Job Type: 12+ Months Contract Job Description DevOps/SRE to improve delivery automation, environment consistency, and operational readiness across our data and analytics platforms - Databricks, Databricks Genie, and BI tooling. The role owns automation of CI/CD pipelines for platform components and analytics workloads (Databricks assets, Python-based tooling/jobs, integrations), including controlled promotion across environments, approval gates, embedded security controls, and reliable rollback and progressive-delivery patterns. Required Skills Databricks/Databricks Genie, BI tooling, Dynatrace, Splunk and AWS, Python or PySparkKey Responsibilities Define and drive adoption of standardized CI/CD pipeline frameworks for platform components and analytics workloads (Databricks assets, Python-based tooling/jobs, integrations), including promotion across environments, approval gates, and rollback patterns.Support and manage our CI/CD infrastructure on Spinnaker and Jules, including progressive delivery.Implement and standardize IaC (Terraform Enterprise) for AWS, Databricks, and platform foundations — reusable modules, guardrails, environment baselines, and drift prevention/detection.Embed security into pipelines: dependency/SCA scanning, artifact signing/SBOM, immutable registries, and policy-as-code guardrails across all environments.Improve developer experience and self-service (paved-road/golden-path bootstrap automation, onboarding, diagnostics, access-workflow automation).Enable and maintain observability integrations and operational dashboards/alerts using Splunk (logging) and Dynatrace (APM/telemetry).Define and operate SLOs/SLIs and error budgets; drive reliability reporting, DR/backup-restore validation, and delivery metrics (deploy frequency, lead time, change-fail rate, MTTR).Own cost visibility and optimization (Databricks DBU/cluster and AWS FinOps).Partner with SRE/data engineering teams on production readiness (runbooks and runbook automation, operational checks, change hygiene) and contribute to incident response, on-call, and blameless postmortems as needed.Support evidence-backed change management aligned to firm controls (approvals, segregation of duties, deployment windows, audit trails).Troubleshoot deployment/runtime issues: permissions/access, configuration drift, job failures, and environment inconsistencies; communicate clearly with stakeholders.Required Qualifications Databricks operational hands-on: jobs/workflows, workspace configuration, permissions/access patterns, and troubleshooting common failures.AWS (hands-on): operating or building cloud resources used by data/analytics platforms (IAM concepts, networking basics, logging/monitoring integration).Python (hands-on): automation/scripting (operational tooling, pipeline helpers, integrations); PySpark for job operation/tuning.CI/CD (hands-on): pipeline design, artifact/versioning discipline, promotion across environments, quality gates, and pipeline security scanning.IaC (hands-on): Terraform Enterprise, including module reuse and state/change discipline.Observability (hands-on): Splunk and Dynatrace: building/operating alerts and dashboards and using telemetry for troubleshooting.Secrets management & least-privilege patterns (AWS Secrets Manager, Vault concepts) with rotation and no hard-coded credentials.SLO/SLI and error-budget fluency; release impact measurement.Demonstrated SDLC participation and production support exposure (change management, troubleshooting, post-incident follow-through).Preferred Qualifications Databricks Genie — set up Genie Agent, build metadata/context, and monitor (hands-on) — strongly preferred; expected to ramp within first 90 days.Sigma Computing BI tooling — strongly preferred (BI/dashboard performance and semantic alignment with governed datasets).Ataccama — data quality/governance workflows.Progressive delivery / feature-flag tooling and DR/BCP (RTO/RPO) experience.Containerization (Docker; Kubernetes exposure a plus if relevant).FinOps/cost-optimization experience for cloud data platforms.