Site Reliability Engineer – Observability & Dynatrace

N2Sglobal — Australia · Posted ~9 hours ago

Senior

Skills

Site Reliability Engineering Dynatrace Observability Application Performance Monitoring Cloud Operations Monitoring and Alerting Incident Management Automation Cloud-native environments Distributed Tracing Root Cause Analysis APM Cloud Kubernetes

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join an engineering team responsible for keeping critical enterprise applications and infrastructure reliable, performant, and scalable. You will configure and support observability platforms, build dashboards and alerts, monitor application and business-service health, investigate bottlenecks, automate operational tasks, and lead incident triage and root-cause analysis in cloud-native environments.

Highlights

Play a key role in improving the reliability, availability, performance, and scalability of critical enterprise systems, with hands-on exposure to modern observability, cloud operations, automation, and incident response practices.

Description

We are seeking a highly skilled Site Reliability Engineer (SRE) with strong expertise in Observability, Dynatrace, Application Performance Monitoring (APM), and Cloud Operations. You will play a key role in ensuring the reliability, availability, performance, and scalability of critical enterprise applications and infrastructure. The ideal candidate will have hands-on experience with Dynatrace, monitoring and alerting frameworks, incident management, automation, and cloud-native environments. Key Responsibilities Implement, configure, and support Dynatrace monitoring solutions across applications, infrastructure, and cloud platforms.Monitor system health, application performance, and business services through observability tools.Configure dashboards, alerts, synthetic monitoring, Real User Monitoring (RUM), and distributed tracing.Perform proactive monitoring and identify performance bottlenecks before business impact occurs.Lead incident triage, troubleshooting, root cause analysis (RCA), and problem management activities.Develop automation scripts for operational activities using Python, Shell, or PowerShell.Collaborate with Development, DevOps, Platform Engineering, Cloud, and Infrastructure teams.Support production environments and participate in major incident management processes.Define and maintain SLI, SLO, and SLA metrics.Drive observability best practices and platform reliability improvements.Support capacity planning and performance optimization initiatives.Required Skills Observability & Monitoring Dynatrace AdministrationDynatrace OneAgentActiveGateReal User Monitoring (RUM)Synthetic MonitoringLog MonitoringDistributed TracingApplication Performance Monitoring (APM)SRE & Operations Site Reliability EngineeringIncident ManagementProblem ManagementRoot Cause AnalysisProduction SupportService ReliabilityCloud & Infrastructure AWS, Azure, or GCPLinux/Unix AdministrationDockerKubernetesNetworking FundamentalsAutomation & Scripting PythonShell ScriptingPowerShellMonitoring Tools (Good to Have) GrafanaPrometheusSplunkELK StackDatadogAppDynamicsPreferred Qualifications Dynatrace Associate or Professional CertificationExperience supporting enterprise-scale production environmentsFinancial Services / Banking experienceExperience with DevOps and CI/CD pipelinesExposure to Infrastructure as Code (Terraform, Ansible)