Senior Site Reliability Engineer - Data Stores & AI/ML Operations

Adobe — United States · Posted ~8 hours ago

Senior

Skills

site reliability engineering production operations datastore engineering AI/ML operations SLOs incident response post-incident reviews monitoring scalability reliability Datastores AI/ML SRE Cloud infrastructure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A Senior Site Reliability Engineer is sought to ensure reliability, scalability, and operational excellence for globally distributed data services. You will own day-to-day production reliability, participate in on-call and incident response, lead post-incident improvements, strengthen operational readiness, evolve core datastores, and increasingly support the operationalization of AI/ML services and workflows.

Highlights

High-ownership senior SRE role focused on global-scale reliability, scalable data stores, operational excellence, and emerging AI/ML operations, with substantial influence over production readiness and incident management.

Description

Senior SRE — RTCDP Datastores & AI/ML Ops Adobe’s Real-Time Customer Data Platform (RTCDP) powers personalized experiences for some of the world’s largest brands. As a Senior SRE on this team, you’ll be central to keeping RTCDP reliable, scalable, and operationally excellent at global scale. This is a hands-on, high-ownership role at the intersection of production operations (Day 2 ownership) and core datastore engineering, with a growing surface area in operationalizing AI/ML services and workflows. What You’ll Do Own production reliability Own day-to-day reliability for RTCDP services — availability, performance, and durability against SLOsParticipate in on-call rotations and incident response, driving mitigation and recovery through SEV3–SEV1 eventsLead post-incident reviews and follow-up workStrengthen operational readiness, playbooks, and on-call healthPartner with product and platform teams on production-ready launches and regional expansions Operate and evolve core datastores You’ll work across RTCDP’s distributed datastore ecosystem: Aerospike, FoundationDB, Postgres, and CosmosDB/DynamoDB. Drive reliability, scaling, and operational excellence across these platformsOwn upgrades, capacity management, backup/restore, and DR testingBuild automation for provisioning, scaling, and lifecycle managementIdentify and ship cost optimizations (rightsizing, storage/compute efficiency) Drive automation and observability Build automation-first solutions that reduce toil and improve system safetyImprove monitoring, alerting, and observability — anchored to real customer impactEstablish standardized operational patterns across services and regionsSupport the rollout of SLO-driven reliability practices Contribute to AI/ML Ops (emerging area) A complementary part of the role, not the primary focus. Support infrastructure and operational needs for AI/ML-powered services in RTCDPShape operational practices for model serving and data pipelines — reliability, scaling, monitoringHelp land core MLOps patterns where relevant: model deployment workflows, inference observability (latency, errors), data quality and pipeline reliability signalsPartner with ML and data teams to get AI-driven features production-readyLeverage AI-assisted tools (e.g., Copilot, Claude Code, Codex, internal tooling) to accelerate debugging, incident response, and operational workflows Technical leadership and collaboration Operate as a strong IC and technical lead on cross-cutting projectsMentor junior engineers and raise team practicesPartner closely with engineering, infrastructure, and securityLive the core SRE/DevOps principles: ownership, automation, error budgets, continuous improvement Why this role is interesting Deep involvement in production systems at global scaleHands-on ownership spanning operations and datastore platformsExposure to next-generation work: AI/ML systems and AI-assisted engineeringDirect impact on customer reliability, platform scalability, and costA clear path toward architect-level influence over time What We’re Looking For 6–10 years in SRE, infrastructure, or platform engineeringProven track record operating large-scale distributed systems in productionStrong foundation in datastores, reliability engineering, and automationHands-on experience with Kubernetes and containerized environments, a major cloud (AWS, Azure, or GCP), and modern observability tooling (Prometheus, Grafana, OpenTelemetry, or equivalents)Real experience in incident response and driving operational improvements out of itWorking knowledge of — or genuine interest in — AI/ML systems or MLOps (expertise not required)Comfortable with scale, ambiguity, and high ownershipStrong problem-solving instincts and a bias for action About Adobe Adobe empowers everyone to create through innovative platforms and tools that unleash creativity, productivity and personalized customer experiences. Adobe’s industry-leading offerings including Adobe Acrobat Studio, Adobe Express, Adobe Firefly, Creative Cloud, Adobe Experience Platform, Adobe Experience Manager, and GenStudio enable people and businesses to turn ideas into impact, powered by AI and driven by human ingenuity. Our 30,000+ employees worldwide are creating the future and raising the bar as we drive the next decade of growth. We’re on a mission to hire the very best and believe in creating a company culture where all employees are empowered to make an impact. At Adobe, we believe that great ideas can come from anywhere in the organization. The next big idea could be yours. Let’s Adobe together At Adobe, we believe in creating a company culture where all employees are empowered to make an impact. Learn more about Adobe life, including our values and culture, focus on people, purpose and community, Adobe for All, comprehensive benefits programs, the stories we tell, the customers we serve, and how you can help us advance our mission of empowering everyone to create. Adobe is proud to be an Equal Employment Opportunity employer. We do not discriminate based on gender, race or color, ethnicity or national origin, age, disability, religion, sexual orientation, gender identity or expression, veteran status, or any other protected characteristic. Learn more. Adobe aims to make our Careers website and recruiting process accessible to any and all users. If you have a disability or special need that requires accommodation to navigate our website or complete the application process, email accommodations@adobe.com. AI Use Guidelines for Interviews: Our interviews are designed to reflect your own skills and thinking. The use of AI or recording tools during live interviews is not permitted unless explicitly invited by the interviewer or approved in advance as part of a reasonable accommodation. If these tools are used inappropriately or in a way that misrepresents your work, your application may not move forward in the process. At Adobe, we empower employees to innovate with AI — and we look for candidates eager to do the same. As part of the hiring experience, we provide clear guidance on where AI is encouraged during the process and where it’s restricted during live interviews. See how we think about AI in the hiring experience. Expected Pay Range: Our compensation reflects the cost of labor across several  U.S. geographic markets, and we pay differently based on those defined markets. The U.S. pay range for this position is $159,200 - $301,600 annually. Pay within this range varies by work location and may also depend on job-related knowledge, skills, and experience. Your recruiter can share more about the specific salary range for the job location during the hiring process. In California, the pay range for this position is $208,300 - $301,600 At Adobe, for sales roles starting salaries are expressed as total target compensation (TTC = base + commission), and short-term incentives are in the form of sales commission plans. Non-sales roles starting salaries are expressed as base salary and short-term incentives are in the form of the Annual Incentive Plan (AIP). In addition, certain roles may be eligible for long-term incentives in the form of a new hire equity award. State-Specific Notices: California: Fair Chance Ordinances Adobe will consider qualified applicants with arrest or conviction records for employment in accordance with state and local laws and “fair chance” ordinances. Colorado: Application Window Notice If this role is open to hiring in Colorado (as listed on the job posting), the application window will remain open until at least the date and time stated above in Pacific Time, in compliance with Colorado pay transparency regulations. If this role does not have Colorado listed as a hiring location, no specific application window applies, and the posting may close at any time based on hiring needs. Massachusetts: Massachusetts Legal Notice It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.