Grid Operations Architect

Siemenssoftware — United States · Posted ~2 hours ago

Senior Full-time

Skills

High-Performance Computing Grid Computing Infrastructure Architecture Distributed Systems Job Scheduling Altair Grid Engine Altair Accelerator IBM LSF SLURM

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A global enterprise software organization is seeking a Grid Operations Architect to help design and support large-scale engineering compute infrastructure. The environment includes globally distributed high-performance computing grids serving thousands of engineers, multiple job schedulers, hundreds of thousands of CPU cores, and workloads reaching millions of jobs per day.

Highlights

Architect-level infrastructure role supporting thousands of engineers through globally distributed high-performance computing environments. Offers exposure to large-scale grid operations, multiple scheduling platforms, massive compute capacity, and complex enterprise infrastructure.

Description

We are a leading global software company dedicated to the world of computer aided design, 3D modeling and simulation— helping innovative global manufacturers design better products, faster! With the resources of a large company, and the energy of a software start-up, we have fun together while creating a world class software portfolio. Our culture encourages creativity, welcomes fresh thinking, and focuses on growth, so our people, our business, and our customers can achieve their full potential. Position Overview: Siemens DISW’s Engineering Grid Computing Services team strategically supports over 3000 engineers located in R&D centers worldwide. The Engineering Grid Computing Service (EGCS) within GTI operates 18+ globally distributed high-performance computing grids across 7+ sites, currently provisioning 263,000+ CPU cores across four scheduler platforms (Altair Grid Engine, Altair Accelerator, IBM LSF, and SLURM). This infrastructure processes up to 12 million jobs per day at peak and serves approximately 3,000 engineers — it is the compute backbone that chip design, verification, and validation depend on. We are seeking a Grid Operations Architect to own the architectural vision and technical strategy for the next generation of this computing fabric. The immediate mandate: scale consolidated grid capacity from a quarter million cores to almost one million cores over the next five-six years while simultaneously transforming and modernizing the entire environment— the migration from Altair Grid Engine to Accelerator Plus, a hierarchical scheduler purpose-built for the throughput demands of EDA workloads This is a senior technical leadership role that bridges grid operations, infrastructure engineering, and strategic planning. Today, a lean team of 3 core grid administrators and 4 regional service delivery partners operate this global service — the Architect role is the structural investment that accelerates transformation and reduces key-person risk. You will define how Siemens DISW’s compute infrastructure is designed, deployed, and operated on a world-class scale. The key purposes of this role are based on the following: As EDA workloads grow and push the limits of what performance can be achieved with existing designs, a next generation architecture is the answer to meet the growing computing demands.The Accelerator migration is the architectural response — and it needs a dedicated owner.Key Technical/Strategic leadership need has been identified, providing depth and resilience to teamThe path to 24x5 follow-the-sun coverage, >50% toil reduction through automation, and full-service excellence depends on an architectural foundation that today's operational team does not have the bandwidth to design. Key Responsibilities: Own the multiyear architectural roadmap for scaling grid compute capacity from 263K+ cores to 1M+ cores across 18+ globally distributed grids, consolidating across sites while maintaining availability and performance.Design the phased migration plan from AGE to Accelerator's hierarchical scheduling platformDesign next generation grid topologies, scheduling frameworks, and resource management strategies that optimize utilization, throughput, and cost efficiency at a million core scale. One of the largest grids alone handles 60K+ cores and 2M+ jobs/dayarchitectures must perform at this density.Develop consolidation strategies to reduce operational complexity rationalizing 4 scheduler platforms toward a unified Accelerator based model, standardizing tooling, and harmonizing operational policies across regions.Lead capacity planning and forecasting models that project compute demand growth using linear, exponential, and flat percentage projections. Translate business requirements into infrastructure investment plans aligned with the 80% utilization ceiling and 72% procurement trigger governance model.Architect hybrid and cloud burst strategies (e.g., NavOps ) to augment on premises capacity for peak demand and elastic workloads working with A&S and Networking teams while respecting the latency and data locality requirements of EDA workloadsDesign fault - tolerant, self healing grid architectures that minimize downtime and reduce manual intervention critical for enabling 24x5 follow the sun coverage without proportional headcount growth.Establish architectural standards, reference designs, and best practices for grid deployment, configuration, and lifecycle management across all 18+ grids and 9+ clusters.Design and drive the evolution of the grid automation and tooling ecosystem — modernizing legacy Perl/Shell-based operational tools toward scalable, maintainable platforms (Python, APIs, infrastructure-as-code, GitLab CI/CD). Target: >50% reduction in manual operational toil across 8 identified operational areas.Architect the observability and analytics stack (Elastic Stack with geo-redundant instances across Americas and Asia, Grafana, Power BI) to provide real-time visibility into grid performance, availability, utilization, and capacity across all grids.Design AI-assisted alerting and auto-classification capabilities leveraging Elastic AI Agent and ServiceNow API integration to move from reactive monitoring to predictive operations.Evolve SLA frameworks and operational KPIs from current state to improve reporting, performance: 99.97% availability per quarter, >97% SLO compliance, 30-minute MTTR, job wait time tracking at P50 and P95.Lead proof-of-concept and pilot initiatives for emerging technologies (containerised workloads, GPU/accelerator scheduling, next-gen schedulers) and evaluate their applicability to Siemens EDA workloads.Partner with engineering leadership, R&D teams, and business stakeholders to understand compute demand trends, workload characteristics, and evolving requirements.Collaborate with A&S team, Infrastructure Technology Verticals — networking, storage, virtualization, and data center operations — to ensure end-to-end infrastructure coherence and performance.Serve as the primary technical liaison with grid software vendors (Altair/Siemens, IBM, SchedMD) for roadmap alignment, escalation management, and strategic feature requests. Leverage Siemens' internal ownership of Altair for strategic advantage.Present architectural proposals, trade-off analyses, and investment recommendations to senior IT leadership and executive stakeholders. Connect every infrastructure investment to a business outcome.Mentor and provide technical guidance to grid administrators and operations engineers, elevating the team's architectural thinking and engineering practices. Drive cross-training from 2 to 3+ schedulers per person.Establish and maintain grid operations governance — change management (moving toward Git-based infrastructure-as-code), configuration standards, security policies, and compliance frameworks for large-scale compute environments.Drive continuous improvement through data-driven analysis of incidents, performance trends, and utilization patterns using Elastic dashboards and capacity modelling tools (Griddash).Design self-service onboarding and provisioning workflows (ServiceNow catalogue items) to reduce manual provisioning overhead and improve time-to-compute for engineering teams.Ensure software license compliance across all grid engines and associated tooling — particularly critical during the scheduler consolidation from 3 license agreements to 1.Contribute to disaster recovery and business continuity planning for grid services.Participate in daily operations standups, weekly staff meetings, and bi-weekly cadences with key customer teams. Provide architectural input on operational issues and escalations. Required Qualifications: Bachelor's degree (or equivalent) in Computer Science, Information Technology, Engineering, or related field. Master's degree preferred.15+ years of progressive experience in systems engineering, infrastructure architecture, or HPC/grid computing environments, with at least 5 years in an architecture or technical leadership role.Deep expertise with distributed computing platforms — hands-on experience architecting and operating at least two of: Altair Grid Engine, Altair Accelerator, IBM LSF, SLURM, or equivalent workload schedulers.Demonstrated experience scaling compute infrastructure in large, multi-site, enterprise-wide environments (100K+ cores). Experience with scheduler migrations or platform consolidation at production scale.Strong systems-level knowledge of Linux (RHEL/CentOS), including kernel tuning, performance optimization, and large-scale fleet management.Proficiency in infrastructure automation and tooling — Python, Perl, Shell scripting — with experience in modern approaches (Ansible, Terraform, infrastructure-as-code, CI/CD pipelines).Solid understanding of networking (TCP/IP, DNS, NFS/GPFS, InfiniBand), storage architectures, and virtualisation as they relate to HPC and grid computing.Experience in cloud computing platforms (AWS, Azure, GCP, OCI) and hybrid cloud/cloud-burst architectures for HPC workloads (e.g., NavOps, Cycle Cloud).Proven ability to develop and communicate multi-year technical roadmaps and translate business requirements into architectural decisions with clear investment justification.Strong project management, prioritization, and stakeholder management skills with experience presenting to senior leadership and executive audiences.Excellent verbal and written communication skills with the ability to present complex technical concepts to both technical and non-technical audiences. Preferred Qualifications: Certification in one or more of cloud computing platforms (AWS, Azure, GCP, OCI)Knowledge of containerization technologies (Docker, Kubernetes, Singularity) and their application in HPC environments.Experience with observability and analytics platforms (Elastic Stack, Grafana, Prometheus,Power BI) at scale — particularly geo-distributed deployments.Familiarity with AI/ML operations and GPU/accelerator scheduling for heterogeneous computing architectures.Experience with database systems (MySQL, MSSQL, PostgreSQL) for operational data management and reporting.Understanding ITIL 4 processes — incident management, problem management, change control— in large-scale infrastructure operations.Experience supporting EDA (Electronic Design Automation) workloads — simulation, regression,DRC/LVS, benchmarking, and tape-out compute patterns.Background in capacity planning, financial modelling for infrastructure investments, and TCO analysis. This is not a traditional grid administration role. The Grid Operations Architect will shape the strategic direction of one of the largest compute fabrics in the industry — an infrastructure that processes up to 12 million jobs per day and directly enables the design and verification of next-generation semiconductors. You will lead the most significant technology transformation in the service's history (the Accelerator migration), architect a 4x capacity scale-up across a globally distributed footprint and design the automation and observability foundations that make million-core operations sustainable. If you thrive on solving complex scaling challenges, influencing technology strategy at the intersection of HPC and enterprise IT, and building infrastructure that operates at world-class scale — this role offers an exceptional opportunity to leave a lasting mark on Siemens’ engineering capability. This position will be subject to U.S. export control requirements under the International Traffic in Arms Regulations (ITAR) and/or Export Administration Regulations (EAR). Employment is contingent on either verifying the U.S. Person status or obtaining any necessary export license. Why us? Working at Siemens Software means flexibility - Choosing between working at home and the office at other times is the norm here. We offer great benefits and rewards, as you'd expect from a world leader in industrial software. A collection of over 377,000 minds building the future one day at a time in over 200 countries. We're dedicated to equality, and we welcome applications that reflect the diversity of the communities we work in. All employment decisions at Siemens are based on qualifications, merit, and business need. Bring your curiosity and creativity and help us shape tomorrow! Siemens Software. Transform the Everyday with Us #SWSaaS The pay range for this position is $129,600 - $233,300 annually with a target incentive of 5-10% of the base salary. The actual wage offered may be lower or higher depending on budget and candidate experience, knowledge, skills, qualifications, and premium geographic location. 129,600 233,300 5-10