Senior DevOps / Cloud Infrastructure Engineer

Open Brand — United States · Posted ~2 hours ago

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Description

Company Overview Be a part of a fast-growing, winning team helping Fortune 1000 consumer brands and retailers leverage AI-driven data insights. OpenBrand is one of the world’s most respected market intelligence companies. OpenBrand’s data and market research products give manufacturers, retailers, and industry players a competitive edge across a wide range of industries (including IT, consumer electronics, home appliances, health, wellness, beauty, small appliances, and other consumer durables) and help marketing, product, sales, and pricing teams make more informed decisions in a rapidly changing market environment. Role Overview This is a permanent, hands-on Senior DevOps / Cloud Infrastructure Engineer role owning the cloud infrastructure, delivery pipelines, and operational reliability behind OpenBrand’s production data platform — the AWS estate, CI/CD and package tooling, observability, and incident response. You will work at the intersection of cloud architecture, automation, and operational excellence, with a strong emphasis on: Owning the AWS estate end to end — architecture, cost, performance, and resilienceBuilding and maintaining reliable CI/CD pipelines, package dependencies, and deployment automationEstablishing real operating control — deploy, observe, troubleshoot, and recover — across the full platformExecuting infrastructure migrations cleanly, on schedule, and without customer or data disruption This is a build-and-run role, not a purely advisory one. You will own infrastructure end to end — design it, automate it, deploy it, monitor it, and improve it — with real accountability for uptime, cost, and delivery velocity. Your first year is an integration year. OpenBrand has grown through acquisition, and the near-term priority is consolidating inherited platform infrastructure onto OpenBrand standards — separating cloud accounts, proving out monitoring and recovery, and retiring legacy dependencies. Success here means being effective inside imperfect inherited systems: making production safe and observable first, then modernizing deliberately. What comes after is the larger half of the job. Once consolidation is complete, this role owns the platform’s forward roadmap — cloud cost re-architecture, moving workloads onto managed and containerized services, deepening automation and observability, and scaling the infrastructure behind a growing data business. We are hiring an owner for the platform, not a migration. You will collaborate closely with Engineering, Data, Security, and IT teams, reporting to the VP of DevOps. This role does not focus on people management, but requires strong judgment, independence, and end-to-end ownership. Key Responsibilities Cloud Infrastructure & Cost Optimization Own the AWS environment across production, QA, and shared services accounts — including account strategy, IAM, networking, EKS/ECS/Lambda, S3, logging, and billing structureLead right-sizing and re-architecture of a large EC2 footprint, moving workloads to appropriate instance families, purchase models, and managed servicesBuild and maintain infrastructure as code (Terraform, CloudFormation, or equivalent) so environments are reproducible and reviewableOwn DNS, certificates, and CDN configuration, including renewal automation and domain transitionsEstablish cost visibility — tagging standards, allocation reporting, budgets, and anomaly alerting — and drive measurable reductions in cloud spendPlan and execute data center and legacy workload decommissioning, including migration sequencing and rollback planningDesign for resilience: multi-AZ posture, documented recovery objectives, and tested backup restores — validated by actual recovery, not by the presence of a backup jobOwn capacity planning and the infrastructure roadmap as the business grows — evaluating managed services, containerization, and architectural changes on their operational and cost merits CI/CD, Observability & Platform Reliability Design, maintain, and improve CI/CD pipelines (GitLab CI, GitHub Actions, or equivalent) from commit through production deployment, including runner fleets and build environmentsContainerize and orchestrate services; manage image registries, build caching, and artifact promotion across environmentsOwn package and artifact dependencies — ECR, JFrog, npm, PyPI, and private mirrors — so builds are reproducible and not silently dependent on external or third-party infrastructureImplement observability — metrics, logging, tracing, dashboards, and actionable alerting (Grafana, Sentry, CloudWatch, or similar) — so failures are detected before customers see themOwn the incident-response path: alert routing and escalation (Opsgenie, PagerDuty, or similar), on-call rotation, P1/P2 severity definitions, and post-incident reviewWrite and test runbooks for critical services, so restart and troubleshooting steps are proven rather than assumedAutomate away manual toil: provisioning, patching, certificate rotation, secrets distribution, and routine operational tasksManage secrets and shared credentials through proper tooling (AWS Secrets Manager, vault-style platforms) rather than ad-hoc practice, including key rotation and access removal after workload transfer Integration, Consolidation & Cross-Functional Collaboration Execute infrastructure separation and integration work arising from acquisitions — account transfers, network segmentation, identity migration, and consolidation onto OpenBrand standardsBuild the infrastructure dependency map: what must run on day one, what is shared, what migrates, what gets replaced, and what retires — with owners and target datesProduce objective evidence of operating independence: validated deploys, working monitoring, tested restores, and an exercised incident path — not just credentialed accessPartner with Data Engineering to keep production data pipelines and analytical platforms (Snowflake, OpenSearch, Airflow, or equivalent) reliable through periods of change, where infrastructure touches data movementPartner with Security and Compliance on IAM design, access reviews, endpoint and network posture, and SOC 2 evidence collectionDocument environments, runbooks, and topology so operational knowledge does not sit with one personSupport Engineering and Product teams with infrastructure input to roadmap decisions and delivery planningCommunicate infrastructure risk clearly to non-technical stakeholders, and help establish best practices for change management, reproducibility, and operational handoff Qualifications Education & Experience 5+ years of experience in DevOps, Site Reliability, Cloud Operations, or Platform Engineering, with significant hands-on responsibility for production infrastructureBS degree in a technical field (e.g., Computer Science, Computer Engineering, Information Systems) or equivalent practical experience Technical Skills Deep, hands-on AWS production operations experience across IAM, VPC/networking, EKS/ECS/Lambda, EC2, S3, RDS, logging, backup/DR, and account strategyStrong proficiency with infrastructure as code (Terraform preferred) and Kubernetes deployment models (Helm, jsonnet, or equivalent)Advanced proficiency in Linux administration and scripting (Python, Bash, or equivalent)Demonstrated experience building and operating CI/CD pipelines for containerized services, including runners, Docker/ECR, and private package management (JFrog, npm, PyPI, artifact mirroring)Practical experience with monitoring and observability platforms (Grafana, Prometheus, Sentry, CloudWatch, Datadog, or similar), plus real on-call and incident-response experience — treating monitoring as operational response, not dashboard creationWorking knowledge of networking fundamentals — routing, VPN, DNS, certificates, firewall rules, and network segmentationSecurity hygiene: secrets management, key rotation, root/admin ownership, least privilege, and access removalProven ability to stabilize and operate inherited or legacy production systems while modernizing them incrementallyAbility to write clear, usable runbooks and communicate technical risk to non-technical stakeholdersExperience using Large Language Models (LLMs) (e.g., GPT-based or similar) to improve productivity and efficiency in engineering workflows, including tasks such as scripting, infrastructure code review, documentation, troubleshooting, and runbook developmentExperience designing workflows that are robust, testable, and maintainable over timeComfort working across the full lifecycle: design → automation → deployment → monitoring → iteration Preferred Experience Experience with M&A carve-out, platform migration, cloud account separation, or data center exit programsExperience driving material AWS cost reduction through right-sizing, reserved capacity, or re-architectureData-platform infrastructure exposure: Snowflake, OpenSearch/Elasticsearch, dbt, MWAA/Airflow, Spark, S3 data lakesFrontend platform and shared-service exposure: identity, OAuth, SSO, module federation / app-shell architectures, CDN, DNS, and certificate managementExperience supporting SOC 2 or comparable compliance programs, including audit evidence and control automationExperience with identity and access platforms (Okta, Entra ID, Google Workspace) and SSO/MFA rolloutExperience with GitLab CI and container registries in a multi-team environmentExperience with SD-WAN or managed network platforms (e.g., Meraki) and site-to-cloud connectivityExperience operating in a private-equity-backed or high-growth environment where speed and cost discipline both matter Benefits Summary Medical, Dental, Vision, and Life InsuranceFlexible Spending Account (FSA) and Health Reimbursement Arrangement (HRA)401(k) Retirement Plan with Company MatchingFlexible Time OffPaid Parental LeaveBase salary range: $115,000-$140,000 annuallyThis role is also eligible for an annual discretionary performance bonus, subject to the terms of the applicable bonus plan