Senior Site Reliability Engineer, Cloud & Domain Services

Dowdisa — United States · Posted ~2 hours ago

Senior Full-time

Skills

Site reliability engineering Cloud platforms Enterprise domain services Automation Observability Performance engineering Infrastructure as Code Identity and directory services Virtual desktop infrastructure High availability Self-healing systems Cloud Virtual desktops

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Lead site reliability engineering for core cloud platforms and enterprise services supporting highly available environments. You will build cloud-native infrastructure automation, improve identity and directory operations, enhance virtual desktop reliability, and define measurable service reliability through automation and observability.

Highlights

Lead reliability engineering for critical cloud and enterprise services, with a strong focus on automation, observability, performance, high availability, and self-healing infrastructure.

Description

About Us: Join the Defense Information Systems Agency (DISA), where we deliver the secure IT and communications capabilities that keep our Nation’s Warfighters and senior leaders connected anytime, anywhere. Within Command, Control, Communications and Computers (C4) Enterprise Directorate (J-6), we empower our Nation’s senior leaders and the warfighter community with assured enterprise services to win in an evolving global landscape. Position Overview: The Senior Site Reliability Engineer (Cloud & Domain Services) leads reliability engineering for core cloud platforms and enterprise domain services, applying automation, observability, and performance engineering to ensure highly available, self‑healing systems. The role builds cloud‑native IaC automation, streamlines identity and directory operations, and enhances the reliability of virtual desktop environments through scalable, automated mechanisms. The Senior SRE defines and measures service reliability across authentication, platform availability, and domain operations, architects deep observability frameworks, and partners with transport SRE counterparts to maintain secure, performant connectivity between cloud services and underlying network infrastructure. Salary: $143,913 to $187,093 per year (inclusive of TLMS and locality adjustments) Location: Fort Meade, MD; Arlington, VA Major Duties: Pioneer SRE for Cloud & Core Domains: Lead the cultural SRE transformation within the J6 operations organization, focusing specifically on automating the reliability of cloud platforms, virtual desktops, and directory services.Engineer Cloud Platform Automation: Write clean, modular infrastructure-as-code (IaC) templates (e.g., Terraform) to automate the deployment, scaling, and configuration of Azure services and Kubernetes (AKS) clusters.Automate Core Directory & Identity Services: Author scripts (PowerShell, Python) to automate Active Directory (AD/Entra ID) administration, group policy enforcement, and identity synchronization, removing manual "toil."Optimize Virtual Desktop Reliability: Engineer automated scaling, performance monitoring, and rapid recovery mechanisms for Azure Virtual Desktop (AVD) environments to ensure a seamless end-user experience.Define Application & Domain SLOs: Collaborate with stakeholders to define and monitor Service Level Indicators (SLIs) and Service Level Objectives (SLOs) specifically for cloud service availability, authentication latency, and system uptime.Architect Cloud-Native Observability: Design and deploy comprehensive monitoring, logging, and tracing frameworks (e.g., Azure Monitor, Prometheus/Grafana) to gain deep visibility into platform health and application performance.Partner with Transport SRE (15%): Collaborate closely with your Transport/Infrastructure counterpart to ensure highly reliable, performant, and secure connectivity between your core cloud platform and the underlying transport networks. Key Skills Site Reliability EngineeringCloud ComputingActive DirectoryPowerShellInfrastructure as CodeKubernetesAzure Virtual DesktopObservabilityAutomationSystems Administration Conditions of Employment: Must be a U.S. Citizen.This national security position, which may require access to classified information, requires a favorable suitability review and security clearance as a condition of employment. Failure to maintain security eligibility may result in termination.Incumbent is required to submit a Financial Disclosure Statement, OGE-450.This is a drug testing designated positionIncumbent is required to travel away from their normal duty station to other locations within the continental United States (CONUS) and/or other locations outside the continental United States (OCONUS) up to 30 percent of the time. Security Clearance: Minimum Requirement: Must possess an active Top Secret (TS) security clearance at the time of application.Preferred Requirement: An active Top Secret / SCI (TS/SCI) clearance is highly preferred.Condition of Employment: Candidates holding a TS clearance must be eligible for, and successfully obtain and maintain, SCI access upon hire. Benefits: A career with the U.S. government provides employees with a comprehensive benefits package. As a federal employee, you and your family will have access to a range of benefits that are designed to make your federal career very rewarding. Learn more about federal benefits. Review our benefits