Description
Job Purpose:
The Senior Specialist - SysOps (DevOps) is responsible for managing, optimizing, and securing enterprise infrastructure, cloud platforms, DevOps toolchains, automation frameworks, and containerized environments.
The role ensures reliability, scalability, performance, and operational excellence across internal and customer-hosted platforms while providing advanced technical support for incidents, service requests, and operational challenges.
Key Responsibilities: -
Infrastructure and Operations
Participate in the design, planning, and implementation of new projects and technologies, ensuring that solutions are highly available, secure, scalable, and performant.Provide day-to-day L2/L3 operational support, including incidents, service requests, problems, changes, and operational tasks.Monitor infrastructure performance, identify and resolve bottlenecks, troubleshoot outages, perform root cause analysis, and recommend improvements.Secure infrastructure by establishing and enforcing policies, defining and monitoring access, and supporting vulnerability and risk remediation.Maintain infrastructure health checks and proactively take action to minimize downtime and performance issues.Ensure services are backed up and recoverable in accordance with approved backup and recovery policies.Report operational infrastructure status, risks, progress, dependencies, and challenges to management.Actively participate in new projects, technology initiatives, and the onboarding of new customers and services.Cloud Platform and Platform Engineering
Design, deploy, administer, and optimize Microsoft Azure infrastructure and platform services.Support cloud migration, modernization, hybrid-cloud integration, governance, security, availability, and cost optimization initiatives.Implement resilient cloud-native architectures and standardized platform services that improve scalability and operational efficiency.Containers, Orchestration, and GitOps
Deploy, manage, upgrade, and troubleshoot containerized workloads using Kubernetes, Docker, and Helm.Implement GitOps-based deployment and configuration management practices for consistent, auditable, and repeatable platform operations.Monitor and optimize cluster availability, performance, resource utilization, capacity, and security.Infrastructure as Code and Configuration Management
Develop and maintain Infrastructure as Code using Terraform, Ansible, and Bicep.Create reusable modules, templates, playbooks, and automated provisioning workflows.Maintain version-controlled infrastructure definitions and configuration standards to ensure consistency across environments.CI/CD and DevSecOps
Design, build, maintain, and optimize CI/CD pipelines using Jenkins, GitHub, GitLab, and Azure DevOps.Automate build, test, security validation, deployment, rollback, and release processes.Collaborate with development, security, and operations teams to embed security, compliance, and quality controls into delivery pipelines.Promote DevOps and GitOps practices to improve deployment speed, consistency, traceability, and reliability.Scripting and Workflow Automation
Develop and maintain automation scripts and configuration files using Bash, PowerShell, Python, and YAML.Automate repetitive infrastructure, deployment, monitoring, reporting, and service-management activities to reduce manual effort and operational risk.Integrate enterprise platforms, APIs, and workflows to streamline service delivery and operational processes.Monitoring and Observability
Implement and maintain monitoring, dashboards, alerting, and observability solutions using Prometheus, Grafana, Azure Monitor, and Log Analytics.Perform proactive monitoring, capacity planning, trend analysis, and performance optimization to maintain service stability and availability.
AI and Automation
Deploy and operationally support Azure AI services and AI-enabled solutions in accordance with security and governance requirements.Design and implement workflow automation that improves operational efficiency, response times, and service quality.Evaluate emerging AI, cloud-native, and automation technologies for practical adoption and service enhancement.Service Management, Documentation, and Stakeholder Support
Ensure compliance with approved ITSM policies, processes, procedures, and service-management guidelines.Write comprehensive technical reports, operational procedures, assessment findings, knowledge-base articles, known-error records, and troubleshooting documentation.
Skills/Certifications (Technical & Non-Technical)
Red Hat Certified System Administrator (RHCSA) or Red Hat Certified Engineer (RHCE)Microsoft Certified: Azure Administrator Associate
Minimum Work Experience: -
5+ Years of Relevant Experience
Education: -
Bachelor’s degree in computer science, Computer Engineering, Information Technology, Information Systems or High School Diploma and Solid Related Work Experience.