Data Center Facilities Engineering Manager

Nebius — Netherlands · Posted ~1 hour ago

Lead Visa History ✓

Skills

Data center facilities management Critical infrastructure Maintenance management Repair management Operations management Service level agreements (SLAs) 24/7 operations Leadership

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A Facilities Engineering Manager is sought to ensure 24/7/365 availability of mission-critical data center operations. The role leads maintenance and repair of critical infrastructure, manages operational performance against SLAs, and provides engineering leadership in a demanding, high-availability environment.

Highlights

Critical leadership position responsible for maintaining continuous data center availability. The role offers ownership of mission-critical infrastructure, maintenance operations, repair programs, and service-level performance in a high-growth technology environment.

Description

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.The Role The Manager, Data Center Facilities Engineering is a critical leadership role responsible for ensuring 24/7/365 availability of data center operations by overseeing maintenance and repair of critical infrastructure in alignment with SLAs. This role leads emergency response, coordinates cross-functional teams, and ensures seamless communication and compliance with operational and safety standards. The manager provides oversight of contracts, maintenance plans, and policy development while driving performance through metrics, continuous improvement, and team development. Your responsibilities will include: Own day-to-day operations of critical electrical and mechanical systems, including UPS, generators, switchgear, PDUs, chillers, and CRAH/CRAC units Ensure high availability and uptime of all facility infrastructure supporting data center operations Lead incident response for facility-related events, including root cause analysis and implementation of corrective actions Monitor performance through BMS/DCIM systems and drive improvements in reliability, efficiency, and capacity utilization Oversee preventive and corrective maintenance programs, ensuring adherence to SOPs, EOPs, and MOPs Manage and hold third-party vendors accountable for service delivery, performance, and compliance with standards Partner with engineering and operations teams on capacity planning, infrastructure scaling, and site optimization initiatives Drive energy efficiency and sustainability efforts, including PUE optimization and support for high-density cooling solutions Ensure compliance with safety, regulatory, and HSE standards, leading audits, inspections, and risk management activities Collaborate cross-functionally with data center operations, network, and infrastructure teams to ensure seamless integration between facilities and IT systems We expect you to have: Bachelor’s degree in Engineering, Mechanical/Electrical Technology, or a related technical field—or equivalent practical experience 10+ years of experience in critical facility or data center operations, or other mission-critical environments 5+ years of experience leading and developing technical teams, including performance management Strong understanding of critical infrastructure systems: electrical distribution (UPS, generators, switchgear), mechanical/HVAC, fire/life safety, and building automation/controls Experience operating in 24/7 high-availability environments with strict uptime requirements Working knowledge of procedure-based operations and safety programs (e.g., LOTO, electrical safety, hazardous energy control) Proven ability to collaborate with cross-functional teams including construction, engineering, and operations Strong communication skills with experience presenting to executive leadership and creating clear, concise reports and documentation Experience managing vendors, maintenance programs, and incident response Familiarity with SOPs, EOPs, MOPs, and operational monitoring systems (BMS/DCIM) It would be an added bonus if you have: Experience in data center or Tier III/IV critical environments, preferably within hyperscale or high-availability operations Advanced education (Master’s or MBA) or relevant trade certification (Electrical, HVAC, Controls) Experience supporting commissioning, new builds, or large-scale facility expansions Strong knowledge of critical infrastructure systems, including power, cooling (CRAH/CRAC, chillers), and exposure to modern technologies such as liquid cooling Proven ability to drive operational excellence through budget management, continuous improvement (Lean/Six Sigma), vendor oversight, and use of tools like CMMS/EAM and AI-enabled workflows Experience operating in hyperscale, AI/ML-driven environments supporting GPU-intensive, high-density workloads Proven ability to scale data center operations to meet rapid growth and next-generation infrastructure demands Strong focus on automation, telemetry, and data-driven decision-making to optimize performance and reliability Experience driving efficiency in advanced cooling environments, including liquid cooling and high-density thermal management Key employee benefits: Health insurance: 100% company-paid medical, dental, and vision coverage for employees and families. 401(k) plan: up to 4% company match with immediate vesting. Parental leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers. Remote work reimbursement: up to $85/month for mobile and internet. Disability & life insurance: company-paid short-term, long-term and life insurance coverage. Compensation We offer competitive salaries ranging from $115K to $275K OTE, which includes base salary and performance bonus. Equity in the form of RSUs may be available at certain salary grades. Join Nebius Today!Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.