Summary
✨ AI‑Generated
A digital services organization is seeking a Senior SRE to improve service reliability, build automation workflows, enhance monitoring, and support production systems.
Highlights
Senior reliability role focused on cloud operations, automation, observability, and service improvement.
Description
Senior SRE – Job Description
Who We Are
Fulcrum Digital is an agile and next-generation digital accelerating company providing digital transformation and technology services right from ideation to implementation.
These services have applicability across a variety of industries, including banking & financial services, insurance, retail, higher education, food, healthcare, and manufacturing.
Requirements
We are looking for a SRE candidate who must have the below skills:
Hands-on experience with Cloud (AWS preferably)ITSM / BMC Helix ProcessSplunk & Dynatrace monitoring and observabilityOwn and improve the full lifecycle of services, from design and development through deployment, operations, and continuous improvement.Partner with development teams to ensure production readiness through architecture reviews, capacity planning, resiliency design, and launch readiness assessments.Design, build, and evolve CI/CD pipelines with a strong emphasis on automation, quality gating, and operational validation.Lead and influence DevOps automation and SRE best practices across engineering teams.Monitor and maintain production systems by measuring availability, performance, latency, and overall system health.Analyze ITSM events and incidents to identify systemic issues, close operational gaps, and feed actionable improvements back to engineering teams.Drive sustainable reliability through automation, self-healing mechanisms, and reduction of manual toil.Respond to production incidents with a calm, systematic approach; lead troubleshooting across the technology stack to reduce mean time to recovery (MTTR).Facilitate blameless postmortems and ensure long-term corrective actions are implemented.Actively manage operational risk, compliance, and resiliency across all supported environments.Collaborate effectively with global teams across multiple geographies and time zones.Mentor junior engineers and contribute to a strong culture of learning, ownership, and operational excellence.All About You
Strong ownership mentality with a systematic, data-driven approach to problem solving.Proven experience working across development, operations, and product teams to drive alignment and outcomes.Comfortable operating in high-pressure situations, making timely decisions, and managing complex stakeholder expectations.Passion for automation, continuous improvement, and eliminating manual work wherever possible.Required Qualifications
Bachelor’s degree in Computer Science or a related technical field involving software development, or equivalent practical experience.Strong experience with algorithms, data structures, scripting, CI/CD pipeline management, and software design principles.Hands-on experience debugging, optimizing, and automating application and infrastructure workflows.Solid experience in: Scripting, Python, Splunk, Dynatrace, Jenkins, XLR, ITIL process.Experience designing, analyzing, and troubleshooting large-scale distributed systems.Strong communication skills with the ability to influence technical and non-technical stakeholders.Preferred Experience
Experience with industry-standard CI/CD and DevOps tooling such as Git / Bitbucket, Jenkins, Maven, Artifactory, and Chef.Proven ability to design and implement efficient CI/CD workflows that move code from development to production with minimal manual intervention.Demonstrated success driving reliability, availability, and performance improvements in complex production environments.Experience operating in regulated or compliance-focused environments.