DevOps & Site Reliability Engineer

Clifyx — United States · Posted ~6 hours ago

Senior Full-time Onsite

Skills

Java Spring Boot Spring MVC Spring Security REST APIs Dynatrace Oracle DB2 MySQL Azure AKS Docker Azure Monitor Log Analytics Enterprise architecture C++ AS400 Python

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Work on enterprise-grade platforms where reliability, observability, and high availability are critical. You will combine Java and Spring development with Azure infrastructure, container orchestration, monitoring, databases, and domain-specific systems supporting retail and payment operations.

Highlights

Permanent onsite engineering role focused on highly available enterprise platforms, observability, Azure cloud infrastructure, containers, and mission-critical payment and retail systems.

Description

Job Title: DevOps & Site Reliability Engineer Location: Deerfield, IL Onsite Duration: Fulltime/Permanent Must Have Technical/Functional Skills Technology and Programming (Expert Level) Strong proficiency in Java full stack developer Object-Oriented programming principles and concepts Hands-on experience on Observability platform Dynatrace Hands-on experience with Spring Framework (Spring Boot, Spring MVC, Spring Security) Knowledge if RESTful API development Experience with database like Oracle, DB2, MySQL Proficiency in Payment Switch BASE24 EPS, C++, AS400 and Python is also added advantage Domain, Cloud & Platform Engineering Must have domain experience on Retail Point of Sale/Payment Systems/Merchandising/Inventory/Logistics area Expertise in Microsoft Azure, including: Compute (VMs, App Services, Azure Container Apps) Containers & Orchestration (AKS, Docker) Storage, Azure Key Vault, Azure Monitor, Log Analytics Proven experience designing enterprise grade, highly available cloud platforms DevOps & Engineering Excellence Advanced experience with Azure DevOps and CI/CD pipeline architecture Strong scripting skills (PowerShell, Bash) GitOps concepts, branching strategies, release orchestration Site Reliability Engineering: Ownership of platform reliability, resiliency, and performance Definition and governance of: SLIs, SLOs, SLAs Error budgets and reliability metrics Advanced observability strategy, designing and implementation: Metrics, logs, traces, alerts, dashboards using Dynatrace Incident response leadership, RCA facilitation, and long term remediation planning Experience operating 99.9%–99.99% availability systems Security, Compliance & Cost Secure cloud design using Key Vault, managed identities, RBAC Cost optimization (FinOps mindset) across cloud infrastructure Roles & Responsibilities Act as SRE Technical Architect (Should be interested to work on Implementations) for clients Retail platforms, owning reliability and stability outcomes Define and enforce SRE standards, best practices, and operating models Architect and govern highly available, scalable cloud platforms Lead the design and implementation of CI/CD and IaC strategies Establish proactive monitoring, alerting, and incident prevention mechanisms Own major incident leadership, RCA execution, and corrective action tracking Partner with application, security, and architecture teams to build reliability by design Drive automation to reduce toil and improve operational efficiency Mentor and coach SRE and DevOps engineers across teams Influence roadmap decisions with a reliability, scalability, and cost lens