Site Reliability Engineer

Synechron — Poland · Posted ~2 hours ago

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Description

About Company: Synechron is a global technology consulting firm that helps leading organizations accelerate digital transformation through innovation, expertise, and agility. With more than 16,500 professionals across around 60 offices in over 20 countries, we combine deep industry knowledge with advanced capabilities in AI, cloud, cybersecurity, and data engineering. Our regional teams, supported by strategic delivery centers, provide scalable, cost-efficient solutions tailored to local markets. Through our award-winning Synechron FinLabs accelerators and strategic partnerships with AWS, Microsoft, Databricks, Salesforce, and ServiceNow, we enable clients to innovate fast and lead with confidence. For more information on the company, please visit our website or LinkedIn community. Diversity, Equity, and Inclusion: Diversity & Inclusion are fundamental to our culture, and Synechron is proud to be an equal opportunity workplace and an affirmative-action employer. Our Diversity, Equity, and Inclusion (DEI) initiative ‘Same Difference’ is committed to fostering an inclusive culture – promoting equality, diversity and an environment that is respectful to all. We strongly believe that a diverse workforce helps build stronger, successful businesses as a global company. We encourage applicants from across diverse backgrounds, race, ethnicities, religion, age, marital status, gender, sexual orientations, or disabilities to apply. We empower our global workforce by offering flexible workplace arrangements, mentoring, internal mobility, learning and development programs, and more. All employment decisions at Synechron are based on business needs, job requirements and individual qualifications, without regard to the applicant’s gender, gender identity, sexual orientation, race, ethnicity, disabled or veteran status, or any other characteristic protected by law. Job Description: We are looking for an experienced Site Reliability Engineer to join our global DevOps team and support highly available, business-critical production services operating 24/7. Key Responsibilities: Ensure the availability, performance, security, and reliability of production services using SRE best practices.Monitor and improve service health by defining SLIs/SLOs and implementing effective observability solutions.Respond to production incidents, perform root cause analysis, and lead post-incident reviews.Develop action plans to address SLO breaches and prevent recurring issues.Contribute to software architecture and technical design discussions.Support the full Software Development Life Cycle (SDLC), from requirements gathering through deployment and maintenance.Plan and execute infrastructure/application migrations, disaster recovery exercises, upgrades, and scheduled maintenance.Automate operational processes and build self-service capabilities to reduce manual effort.Participate in an on-call rotation and provide support during scheduled weekend maintenance when required.Collaborate with globally distributed engineering, product, operations, and vendor teams. Required Qualifications: 5+ years of experience in Production Application Support, Site Reliability Engineering, or a similar role.Strong troubleshooting, incident management, problem-solving, and issue-prevention skills.Hands-on experience with tools such as Ansible, Jenkins, Prometheus, and Grafana.Programming or scripting experience in one or more of Java, Python, or Node.js.Strong SQL knowledge.Good understanding of SDLC principles and practices.Strong analytical, communication, and collaboration skills.Experience working in high-pressure, 24/7 production environments. Preferred Qualifications: Experience supporting large-scale Atlassian Jira and Confluence Data Center environments.Strong knowledge of observability, monitoring, alerting, and performance management.Experience with cloud platforms, infrastructure automation, disaster recovery, and service migrations.Ability to quickly learn and adapt to new technologies.