Description
Staff DevOps Engineer
Location: Chicago, IL or New York, NY
Job Type: Full-Time / FTE
About the Role
Our Client is looking for a Staff DevOps Engineer to lead the architecture and evolution of our cloud infrastructure, DevOps, and platform engineering capabilities.
This is a highly technical, cross-functional role for an experienced infrastructure engineer who can operate at Staff/Architect level, drive large-scale cloud and CI/CD initiatives, and establish engineering standards across multiple teams.
A key focus of this role will be exploring and implementing AI-driven infrastructure automation and agentic AI capabilities to improve reliability, operational efficiency, and developer productivity.
What You'll Do
Own the technical architecture for complex, cross-team infrastructure and platform initiatives spanning cloud, CI/CD, security, and observability.Design, build, and maintain scalable, resilient cloud-native infrastructure supporting enterprise applications, AI platforms, and core business systems.Architect and integrate AI/LLM and agentic AI workflows into DevOps and platform operations, including automated incident triage, AI-assisted change management, self-healing infrastructure, and operational copilots.Establish and drive best practices for Infrastructure as Code, CI/CD, security, observability, and platform engineering.Lead technical design reviews and provide architectural guidance to engineering and platform teams.Partner with engineering, product, security, and technology leadership to translate business requirements into scalable infrastructure strategies and roadmaps.Identify and resolve complex performance, reliability, scalability, and availability issues across cloud infrastructure and deployment pipelines.Lead incident response, disaster recovery, root-cause analysis, and blameless postmortems.Mentor senior and mid-level engineers and help raise the organization's technical standards.Evaluate emerging technologies and identify opportunities to improve infrastructure automation, developer experience, and system resilience.What We're Looking For
12+ years of experience in infrastructure, DevOps, platform, cloud, SRE, or software engineering.Proven experience architecting and delivering large-scale cloud infrastructure and CI/CD platforms.Expert-level experience with Infrastructure as Code, particularly Terraform, Bicep, or similar technologies.Strong experience with CI/CD platforms such as GitHub Actions, Azure DevOps, Jenkins, or equivalent.Deep knowledge of Kubernetes, containers, distributed systems, and cloud-native architecture.Strong hands-on experience with Azure, AWS, or GCP.Strong understanding of observability, including metrics, logging, tracing, monitoring, and production reliability.Experience with incident response, disaster recovery, performance engineering, and systemic problem solving.Strong understanding of enterprise security practices, including identity/IAM, Zero Trust, secrets management, audit controls, and change management.Experience working in regulated or enterprise environments, including SOX or similar compliance frameworks, is highly valuable.Demonstrated ability to lead technical initiatives across multiple teams and influence engineering direction without direct authority.Strong mentoring, communication, and technical leadership skills.Preferred Experience
Experience with agentic AI, LLM-powered infrastructure automation, or AI-driven DevOps.Experience designing multi-agent systems, tool-calling architectures, or RAG pipelines.Hands-on experience with LangChain, AutoGen, Semantic Kernel, MCP, or similar AI orchestration technologies.Experience integrating AI agents or LLM-powered tooling into DevOps, SRE, or infrastructure workflows.Experience with event-driven architectures and real-time platforms such as Kafka, Azure Event Hubs, or Kinesis.Familiarity with Databricks or modern data platforms.Experience in commercial real estate, fintech, financial services, or transaction/operations platforms.Technical thought leadership through conference presentations, technical writing, or open-source contributions.Why This Role
This is an opportunity to have a significant impact on how a large enterprise builds and operates its cloud, platform, and AI infrastructure.
You'll have the opportunity to shape modern DevOps practices while helping define how AI and autonomous automation can transform infrastructure operations across the organization.