Vice President, Compliance Engineering DevOps

Goldman Sachs — United States · Posted ~22 hours ago

Head Full-time

Skills

DevOps SRE software engineering systems engineering distributed systems fault-tolerant systems monitoring automation scalability Big Data

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Lead engineering efforts for mission-critical compliance platforms where software and systems engineering meet at scale. You will help build reliable, highly available services, automate operational processes, improve observability, and design systems capable of handling substantial structured and unstructured data.

Highlights

Mission-critical engineering role working with large-scale data, modern technology, distributed systems, automation, and highly scalable platforms. Strong exposure to complex engineering challenges and opportunities to improve production reliability.

Description

Job Description WHAT WE DO: We are Compliance Engineering, a global team of more than 300 engineers and scientists who work on the most complex, mission-critical problems. We build and operate a suite of platforms and applications that prevent, detect, and mitigate regulatory and reputational risk across the firm, have access to the latest technology and to massive amounts of structured and unstructured data, leverage modern frameworks to build responsive and intuitive front end and Big Data applications. The firm is making a significant investment to uplift and rebuild the Compliance application portfolio. SRE at Goldman Sachs combines software and systems engineering to build run, and maintain high performant, distributed, fault tolerant systems. As a SRE Engineer you will be filling a mission-critical role ensuring that our systems are healthy, monitored, automated, and designed to scale. You will collaborate with engineering teams to continually improve our production services, facilitating fast delivery of new services, and reducing downtime. SRE utilizes automation, tools and solid engineering principles to optimize existing systems, build infrastructure and eliminate operational work. We are looking for passionate, curious, driven engineers who thrive on solving operational problems and improving efficiency and would like to apply their skills to Compliance Engineering SRE. Job Responsibilities Proactive management of our production services by measuring and monitoring availability, capacity and overall system health.Shaping software before go-live through activities such as system design consulting, capacity planning and launch reviews.Scaling and evolving systems by pushing for changes that improve capacity and reliability.Practicing sustainable incident management in a blameless postmortem culture.Identifying and building improvements to system behavior, control and monitoring tools.Defining and maintaining Service Level Indicators (SLIs) and Service Level Objectives (SLOs) to quantify and manage service reliability.Engineering solutions to reduce "toil" through advanced automation and self-healing system capabilities.Conducting capacity modeling and performance tuning to ensure systems meet future demand. Technical Skills WHAT WE ARE LOOKING FOR: Experience in one or more of the following: Java, Python and Perl.Strong communication skills and the ability to clearly express ideas and arguments.Solid analytical and problem-solving skills with appreciation of technical risk.Experience with automated testing and SDLC concepts, developing applications in a Linux environment, and sound knowledge of algorithms, data structures and software design.Systematic problem-solving approach and a sense of ownership and drive.Ability to debug and optimize code and to automate routine tasks.Proficiency with Observability stacks, including distributed tracing, logging, and metrics (e.g., Prometheus, Grafana, ELK, or OpenTelemetry).Deep understanding of containerization and orchestration technologies, specifically Docker and Kubernetes (K8s).Knowledge of networking protocols and load balancing strategies in a distributed systems environment. Preferred Qualifications Bachelor’s degree in Computer Science, or a related technical field that involves programming.Experience in some of the following is desired: SRE experience, relational databases, Hadoop and big data technologies, knowledge of the financial industry or compliance/risk functions.Experience with Chaos Engineering principles and fault-injection testing to verify system resilience.Familiarity with cloud-native architecture and managed services (AWS, Azure, or GCP).Understanding of Error Budgeting and its application in balancing feature velocity with system stability.Hands-on experience with Infrastructure as Code (IaC) frameworks such as Terraform, Ansible, or CloudFormation. Interpersonal Skills Passionate about solving operational problems and constant improvement via automation.Highly motivated, pro-active and capable of working under pressure without compromising development processes.Strong, committed and reliable team player and strong communicator, able to take direction but also willing to contribute to discussions on design and strategy.Possess client-facing skills to be able to deal with and form good relationships with the business and other technology groups through day-to-day support and project work.Ability to multi-task and be effective working in a fast-paced environment.Interest in financial markets, technology and the ability to learn. About Goldman Sachs The Goldman Sachs Group, Inc. is a leading global investment banking, securities and investment management firm that provides a wide range of financial services to a substantial and diversified client base that includes corporations, financial institutions, governments and individuals. Founded in 1869, the firm is headquartered in New York and maintains offices in all major financial centers around the world.