Site Reliability Engineer

Shya Workforce Solutions — United States · Posted ~2 hours ago

Mid Full-time

Skills

Observability Kubernetes AWS Azure Application monitoring Log analysis Splunk AppDynamics

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology team is seeking an SRE to build observability solutions, monitor cloud infrastructure, troubleshoot production issues, and improve system reliability.

Highlights

Hands-on reliability engineering role focused on cloud platforms, monitoring, incident response, and performance optimization.

Description

We are looking for a Site Reliability Engineer (SRE) with strong hands-on experience in observability, cloud platforms, Kubernetes, application performance monitoring, and log analysis. The ideal candidate will have experience working with observability platforms such as AppDynamics, Splunk, or equivalent APM/logging tools, along with Kubernetes environments and cloud platforms including AWS and Azure. Responsibilities Design, implement, and maintain observability solutions across applications and infrastructure.Monitor system health, availability, performance, and reliability.Work with AppDynamics, Splunk, or similar APM and log analysis platforms.Analyze application and infrastructure logs to identify performance issues and troubleshoot production incidents.Support and monitor Kubernetes environments, preferably EKS and/or AKS.Implement cloud observability and monitoring across AWS and Azure environments.Participate in incident management, root-cause analysis, and production troubleshooting.Develop and improve monitoring, alerting, dashboards, and reliability metrics.Collaborate with development, infrastructure, DevOps, and cloud teams to improve application reliability.Identify opportunities to automate repetitive operational and monitoring activities.Contribute to SRE best practices around availability, scalability, performance, and resilience. Qualifications 3–5 years of hands-on experience in Site Reliability Engineering, Observability, DevOps, or a similar technical discipline. Required Skills Strong experience with observability and monitoring platforms such as:AppDynamicsSplunkDynatrace, Datadog, New Relic, or equivalent toolsHands-on experience with Kubernetes.Experience with AWS and/or Azure cloud environments.Experience with EKS and/or AKS is preferred.Strong troubleshooting and production-support experience.Experience analyzing application, infrastructure, and system logs.Understanding of monitoring, alerting, application performance, and infrastructure health.Strong communication and problem-solving skills. Preferred Skills Experience supporting large-scale production environments.Experience with cloud-native applications and microservices.Knowledge of CI/CD, Infrastructure as Code, and automation.Experience with scripting using Python, Bash, or similar languages.Understanding of SLI, SLO, SLA, error budgets, and other SRE practices.