Summary
Help build highly reliable cloud platforms through intelligent observability, infrastructure automation, and cloud-native engineering. Drive proactive monitoring, self-healing systems, and scalable operational practices across distributed environments.
Highlights
Focus on AI-powered observability, automation, multi-cloud infrastructure, and large-scale reliability engineering with modern DevOps tooling.
Description
About the Role
We are seeking a highly skilled SRE / DevOps Engineer with strong expertise in Dynatrace, AI-driven observability (Davis AI), automation (Ansible), and cloud platforms (AWS & Azure).
This role will focus on proactive monitoring, intelligent automation, and reliability engineering, ensuring high system availability and performance across distributed environments.
Responsibilities
Dynatrace & AI-Driven Observability (Primary Focus)
Lead implementation and optimization of Dynatrace platform across applications and infrastructureLeverage Dynatrace Davis AI for:Automated root cause analysisAnomaly detection and event correlationPredictive performance insightsAlert noise reductionConfigure and manage:OneAgent deploymentsSmartscape topology mappingService flow and distributed tracingDefine and monitor SLIs, SLOs, and user experience metricsBuild custom dashboards, alerts, and observability pipelinesIntegrate Dynatrace with:CI/CD pipelines (release validation, performance gating)Incident management tools (PagerDuty, ServiceNow, etc.)Enable self-healing automation using Dynatrace event triggers and AI insights
Automation & Configuration Management (Ansible Focus)
Design and implement automation using Ansible for:Configuration managementApplication deploymentsEnvironment provisioningDevelop reusable playbooks and roles for scalable operationsAutomate operational tasks, patching, and compliance processesIntegrate Ansible with CI/CD pipelines and monitoring systemsImprove system reliability through automated remediation workflows
Cloud & AWS DevOps Tooling
Design and manage cloud-native systems on AWS, with exposure to AzureDevelop infrastructure using Terraform, CloudFormation, or CDKBuild and manage CI/CD pipelines (GitHub Actions, Jenkins, GitLab CI)Develop and deploy serverless architectures (Lambda, API Gateway, Step Functions)Use AWS SDK (boto3) to automate DevOps and operational workflowsDeploy and maintain large-scale production systems via automated pipelinesOptimize cloud infrastructure for cost, performance, and scalability
Monitoring, Logging & Multi-Cloud Observability
Strong experience with:AWS CloudWatch (metrics, logs, alarms, dashboards)Azure Monitor / Log AnalyticsDesign unified observability across multi-cloud environmentsImplement logging and tracing strategies for distributed systems
Containers, Platforms & Reliability Engineering
Work in containerized environments (Docker)Manage orchestration platforms such as Kubernetes, ECS, AKSEnsure high availability using:Fault tolerance designDisaster recovery strategiesSupport incident response, on-call processes, and root cause analysis (RCA)
Qualifications
Proven experience with Dynatrace (APM, RUM, infrastructure monitoring)Strong hands-on experience with Dynatrace Davis AI capabilitiesExperience with Ansible for automation and configuration managementDeep knowledge of AWS services and cloud-native architecturesExperience with Infrastructure as Code tools (Terraform/CloudFormation/CDK)Proficiency in Python (boto3), Bash scriptingExperience working in production-scale environmentsBusiness Analyst experienceScrum Master experience
Required Skills
Dynatrace certification (Associate/Professional)Advanced experience with Dynatrace APIs and automationExperience building self-healing systems using AI-driven triggersFamiliarity with Prometheus, Grafana, ELK stackAzure cloud experience and certificationsExperience with GitOps and platform engineering