Description
Your responsibilities: (Up to 10, Avoid repetition)
Experience: 10+ years of hands-on experience in Site Reliability Engineering (SRE), DevOps, or Systems Architecture, with at least 3+ years specializing deeply in Datadog administration and configuration.Cloud & Container Expertise: Deep professional experience working with AWS, Azure, or GCP, pairedwith heavy production experience managing Kubernetes clusters.Instrumentation & Coding: Proficiency in systems or scripting languages (e.g., Python, Go, Bash, orJavaScript) and experience instrumenting applications for APM.Datadog Mastery: Deep understanding of Datadog’s core pillars—Infrastructure, APM, Logs, Metrics,Synthetics, and Security Monitoring.
Datadog Certifications are a strong plus.Problem-Solving Mindset: Demonstrated ability to debug complex, distributed microservicesarchitectures under high-pressure incident response scenarios.Communication Skills: Excellent interpersonal and stakeholder management skills, with the ability to translate technical telemetry data into actionable business and engineering insights.
Your Profile
Essential skills/knowledge/experience: (Up to 10, Avoid repetition)
Experience: 10+ years of hands-on experience in Site Reliability Engineering (SRE), DevOps, or Systems Architecture, with at least 3+ years specializing deeply in Datadog administration and configuration.Cloud & Container Expertise: Deep professional experience working with AWS, Azure, or GCP, pairedwith heavy production experience managing Kubernetes clusters.Instrumentation & Coding: Proficiency in systems or scripting languages (e.g., Python, Go, Bash, orJavaScript) and experience instrumenting applications for APM.Datadog Mastery: Deep understanding of Datadog’s core pillars—Infrastructure, APM, Logs, Metrics,Synthetics, and Security Monitoring.
Datadog Certifications are a strong plus.Problem-Solving Mindset: Demonstrated ability to debug complex, distributed microservicesarchitectures under high-pressure incident response scenarios.Communication Skills: Excellent interpersonal and stakeholder management skills, with the ability to translate technical telemetry data into actionable business and engineering insights.
Desirable skills/knowledge/experience: (As applicable)
Architecture & ImplementationPlatform Ownership: Design, deploy, and manage Datadog agents, integrations, and custom metricsacross multi-cloud (AWS/Azure/GCP) and containerized (Kubernetes, Docker) environments.Observability Pipelines: Architect and scale high-throughput log processing, routing, and transformationsystems using Datadog.APM & Infrastructure Monitoring: Configure and optimize Application Performance Monitoring (APM),Distributed Tracing, Real User Monitoring (RUM), and Infrastructure metrics.Governance & Best PracticesStandardization: Establish company-wide standards for dashboards, monitors, SLOs/SLIs, and alertrouting (integrating with PagerDuty, Jira, Opsgenie, etc.).Cost & Performance Optimization: Audit and optimize Datadog usage, index management, logretention policies, and custom metric volume to maximize ROI and control licensing costs.Security & Compliance: Leverage Datadog Security products (CSPM, CWPP, Cloud SIEM, ContainerSecurity) to maintain compliance postures and mitigate runtime threats.Collaboration & EnablementCross-Functional Mentorship: Act as the go-to escalation point and technical mentor for DevOps, SRE,and Software Engineering teams regarding troubleshooting and instrumentation.Training & Documentation: Create internal documentation, runbooks, and training modules to elevate organizational proficiency in observability.Vendor Management: Act as the primary technical point of contact for Datadog account teams,