Site Reliability Engineer

Infosys — Chile · Posted ~1 week ago

Mid

Skills

SRE cloud infrastructure CI/CD automation monitoring incident management cloud

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Work as a reliability-focused engineer responsible for maintaining highly available digital services, improving infrastructure automation, supporting deployments, and driving operational improvements.

Highlights

Focus on improving reliability, automation, operational excellence, and collaboration with engineering teams on large-scale digital services.

Description

Job Description – Site Reliability Engineer (SRE) Role Purpose The Site Reliability Engineer (SRE) is responsible for ensuring the reliability, availability, and performance of digital services in production, balancing service stability with the ability to deliver change at speed. The role focuses on strengthening operational resilience through engineering, automation, and proactive reliability practices, working closely with application and platform teams. Scope of the Role The locally applied SRE role covers: Production digital services (applications, platforms, data products)Associated infrastructure (cloud, CI/CD pipelines, integrations)Continuous operations (24/7 reliability mindset, not necessarily shift-based)Production changes (deployments, configurations)Incidents, problems, and service degradationsContinuous improvement of stability and operational efficiency Key Responsibilities Service Reliability & Availability Define, implement, and maintain SLIs and SLOs(availability, latency, error rates)Continuously monitor service health and anticipate degradationsEnsure services operate within business‑agreed reliability thresholdsManage reliability trade‑offs between speed and stability Incident & Problem Management Lead or coordinate response to relevant incidents (L2/L3)Ensure:Rapid and structured diagnosisSafe service restorationClear and effective communicationFacilitate blameless postmortemsConvert recurring incidents into engineering improvement backlogDrive long‑term remediation rather than reactive firefighting Automation & Operational Excellence Identify repetitive and manual operational tasksDesign and implement automation for:DeploymentsMonitoring and alertingHealth checksBasic recovery and self‑healing (where applicable)Reduce toil and increase system resilience through engineering solutions Change Governance & Production Readiness Support vendor and internal team change trackingEnsure changes:Are traceableHave defined rollback strategiesMinimize operational riskValidate operational readiness before productionParticipate early in solution and architecture design from a reliability perspective (early involvement) Metrics, Observability & Continuous Improvement Define and maintain near real‑time operational KPIs (“service pulse”)Ensure every deviation has:Clear ownershipDefined corrective actionsPrevent reactive operations by driving data‑driven decision makingSupport identification, prioritization, and planning of technical debt remediation What This Role Is Not The SRE role will not be: A dedicated incident operator onlyAn advanced Service DeskThe sole owner of service stability (reliability is shared)A gatekeeper blocking changes without technical justificationThe owner of contractual MOPsA commercial or account management roleThe customer-side account or delivery lead Experience & Profile (Indicative) Proven experience as SRE, Production Engineer, or similar roleStrong background in production systems and reliability engineeringExperience working with:Cloud platformsCI/CD pipelinesMonitoring and observability toolsComfortable operating in product‑oriented or POD-based team modelsStrong problem‑solving, communication, and collaboration skills Operating Model Alignment Works embedded or as an enabling function with PODsFocused on enablement and reliability patterns, not centralized controlPromotes shared ownership of reliability EEO/About Us: About Us Infosys is a global leader in next-generation digital services and consulting. We enable clients in more than 50 countries to navigate their digital transformation. With over four decades of experience in managing the systems and workings of global enterprises, we expertly steer our clients through their digital journey. We do it by enabling the enterprise with an AI-powered core that helps prioritize the execution of change. We also empower the business with agile digital at scale to deliver unprecedented levels of performance and customer delight. Our always-on learning agenda drives their continuous improvement through building and transferring digital skills, expertise, and ideas from our innovation ecosystem. EEO Infosys provides equal employment opportunities to applicants and employees without regard to race; color; sex; gender identity; sexual orientation; religious practices and observances; national origin; pregnancy, childbirth, or related medical conditions; or disability.