Description
Location
New York
Business Area
Engineering and CTO
Ref #
10053399
Description & Requirements
Who We Are
The CNCI Core & Lifecycle team - part of Bloomberg’s Cloud Native Compute Services group - provides the compute foundations for cloud-native services at Bloomberg scale.
We build the software, automation, and infrastructure capabilities required to operate and evolve Bloomberg’s large-scale physical compute fleet.
Working Below The Application And Orchestration Layers, We Provide The Physical Infrastructure And Lifecycle Systems That Underpin Bloomberg’s Most Critical Stateful Platforms, Including
Bloomberg Search infrastructureHadoop analytics and compute platformsNoSQL database platformsRelational database platformsOn-premises GPU infrastructure supporting AI training and inference
Our mission is to ensure the fleet operates reliably and stays ready for critical workloads by automating provisioning, configuration, upgrades, and remediation across the entire hardware lifecycle.
This demands senior engineering judgment in distributed systems: orchestrating high-stakes changes across large-scale fleets while keeping critical stateful services available, resilient, and performant.
You’ll navigate data placement, replication, capacity, and application-specific constraints to deliver infrastructure improvements safely at scale.
We'll Trust You To
As a Senior Software Engineer On CNCI Core & Lifecycle, You Will Design And Build Systems That Make Bloomberg's Physical Infrastructure Easier, Safer, And More Automated To Operate At Scale.
You Will
Improve the reliability of our foundational infrastructure by safely managing OS, firmware, network policy, and configuration changes across production fleets, while designing lifecycle workflows that enable graceful recovery, failover, and restoration for stateful workloads.Accelerate self-service management by delivering reusable tooling and APIs for platform teams running Hadoop, databases, search, and AI services.Automate compute lifecycles end-to-end, from firmware and OS provisioning to full application-level stack deployment.Build and operate automation for rapid-response CVE patching and proactive security hardening across the fleet.Enhance system observability through robust monitoring, alerting, and cross-team architectural collaboration.Manage incident response and root-cause analysis with your team and end users to drive long-term system stability improvements.
Who You Are
Take ownership: You care about the craft and hold yourself to high engineering standards.
Systems thinking: You see the big picture.
You understand how a firmware change at the node level can ripple through the stack and affect the latency and reliability of a distributed database.
Own the challenge: You’re a self-starter who gets a kick out of untangling tough infrastructure and production problems.
Automate everything: You have an itch to eliminate manual toil; if it’s repetitive, you’ll automate it.
Collaborate openly: You’re a team player who thrives on giving and getting honest feedback - we build better stuff together.
What We're Looking For
We are looking for engineers who enjoy solving infrastructure problems with software and are interested in working at the boundary between software and physical infrastructure.
4+ years of software engineering experience, with proficiency in Go, Python, or a comparable language.BA/BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field.Experience building and maintaining production observability using technologies such as Prometheus, Grafana, and OpenTelemetry.Strong hands on Linux systems expertise, including command-line troubleshooting of processes, storage, networking, system logs, kernel messages, and systemd services on Ubuntu or similar distributions.Experience building infrastructure automation with tools such as Ansible and Terraform.Strong communication and collaboration skills within large-scale production environments.
We'd Love to See
Experience in one or more of the following areas would be particularly valuable:Experience managing large-scale fleets of physical servers using modern infrastructure practices, including immutable and image-based OS frameworks such as Kairos, automated OS management, desired-state configuration, and configuration drift detection.Experience with GPU and HPC infrastructure, including hardware provisioning and lifecycle management, firmware, fleet observability, and high-speed networking such as InfiniBand.Experience building or operating infrastructure that supports large-scale stateful distributed systems such as Hadoop, HBase, NoSQL and relational databases, search platforms, or Kubernetes-based services.Advanced Linux networking expertise (e.g., BGP, eBPF) in security-focused or highly regulated environments.Experience with workload identity, node and workload attestation, and zero-trust infrastructure using SPIFFE/SPIRE or similar technologies.
What's In It For You
You will build the bedrock of Bloomberg’s technology ecosystem.
Your work directly underpins the reliability, security, and performance of our most critical Search, Database, and AI platforms.
You aren't just managing servers; you are building the lifecycle orchestration layer that allows Bloomberg’s physical infrastructure to evolve reliably at scale.
If you want to solve complex problems at the intersection of Linux, hardware, and distributed systems while seeing your code manage thousands of nodes at scale, this is the place for you.
Salary Range = 160,000 - 240,000 USD Annual + Benefits + Bonus
The referenced salary range is based on the Company's good faith belief at the time of posting.
Actual compensation may vary based on factors such as geographic location, work experience, market conditions, education/training and skill level.
We offer one of the most comprehensive and generous benefits plans available and offer a range of total rewards that may include merit increases, incentive compensation (exempt roles only), paid holidays, paid time off, medical, dental, vision, short and long term disability benefits, 401(k) +match, life insurance, and various wellness programs, among others.
The Company does not provide benefits directly to contingent workers/contractors and interns.
Discover what makes Bloomberg unique - watch our for an inside look at our culture, values, and the people behind our success.