Summary
✨ AI‑Generated
A technology team is seeking an SRE to build observability solutions, monitor cloud infrastructure, troubleshoot production issues, and improve system reliability.
Highlights
Hands-on reliability engineering role focused on cloud platforms, monitoring, incident response, and performance optimization.
Description
We are looking for a Site Reliability Engineer (SRE) with strong hands-on experience in observability, cloud platforms, Kubernetes, application performance monitoring, and log analysis.
The ideal candidate will have experience working with observability platforms such as AppDynamics, Splunk, or equivalent APM/logging tools, along with Kubernetes environments and cloud platforms including AWS and Azure.
Responsibilities
Design, implement, and maintain observability solutions across applications and infrastructure.Monitor system health, availability, performance, and reliability.Work with AppDynamics, Splunk, or similar APM and log analysis platforms.Analyze application and infrastructure logs to identify performance issues and troubleshoot production incidents.Support and monitor Kubernetes environments, preferably EKS and/or AKS.Implement cloud observability and monitoring across AWS and Azure environments.Participate in incident management, root-cause analysis, and production troubleshooting.Develop and improve monitoring, alerting, dashboards, and reliability metrics.Collaborate with development, infrastructure, DevOps, and cloud teams to improve application reliability.Identify opportunities to automate repetitive operational and monitoring activities.Contribute to SRE best practices around availability, scalability, performance, and resilience.
Qualifications
3–5 years of hands-on experience in Site Reliability Engineering, Observability, DevOps, or a similar technical discipline.
Required Skills
Strong experience with observability and monitoring platforms such as:AppDynamicsSplunkDynatrace, Datadog, New Relic, or equivalent toolsHands-on experience with Kubernetes.Experience with AWS and/or Azure cloud environments.Experience with EKS and/or AKS is preferred.Strong troubleshooting and production-support experience.Experience analyzing application, infrastructure, and system logs.Understanding of monitoring, alerting, application performance, and infrastructure health.Strong communication and problem-solving skills.
Preferred Skills
Experience supporting large-scale production environments.Experience with cloud-native applications and microservices.Knowledge of CI/CD, Infrastructure as Code, and automation.Experience with scripting using Python, Bash, or similar languages.Understanding of SLI, SLO, SLA, error budgets, and other SRE practices.