Site Reliability Engineer

Smart It Frame Llc — United States · Posted ~7 hours ago

Senior Full-time Hybrid

Skills

Kubernetes cloud infrastructure incident management observability automation cloud

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An SRE role focused on maintaining high-availability applications, managing incidents, improving cloud platforms, and implementing automation for critical systems.

Highlights

Support mission-critical systems while improving reliability, scalability, automation, and operational excellence in enterprise environments.

Description

Dear Candidates, Greetings! We have a Fulltime role with one of our clients. Kindly find the below details and let me know if you are interested. Role: Production Support Engineer Site Reliability Engineer SRE Location: New York - New York (Onsite/ Hybrid) Duration: FTE Job Description: Production Support Engineer Site Reliability Engineer SREExperience 5 YearsDomain Banking Financial Services Capital Markets FinTech Digital AssetsJob Summary We are seeking an experienced Production Support Engineer / Site Reliability Engineer (SRE) to support mission-critical applications and cloud-native platforms. The ideal candidate will have strong expertise in Kubernetes, application support, observability, automation, and cloud infrastructure, with a focus on ensuring high availability, reliability, scalability, and operational excellence across enterprise environments. Key Responsibilities Provide L2/L3 production support and incident management for business-critical applications.Perform troubleshooting, root cause analysis (RCA), and implement preventive measures to reduce recurring incidents.Support and optimize cloud environments across Microsoft Azure and AWS.Manage and troubleshoot Kubernetes (AKS), Docker, virtual machines, and distributed systems.Build and maintain observability and monitoring solutions using Site24x7, Splunk, Prometheus, Grafana, Azure Monitor, and CloudWatch.Develop automation scripts using Python, Bash, or PowerShell to improve operational efficiency.Support CI/CD pipelines, deployment automation, and release management processes.Troubleshoot application, API, database, and infrastructure-related issues.Ensure compliance through IAM management, certificate renewals, vulnerability remediation, and security best practices.Contribute to AI-driven operational automation, intelligent monitoring, and AIOps initiatives.Required Skills Cloud & Infrastructure AzureAWSAKS (Azure Kubernetes Service)Azure MonitorCloudWatchTerraformContainers & Orchestration KubernetesDockerHelmMonitoring & Observability Site24x7 (Mandatory)SplunkPrometheusGrafanaNew RelicDevOps & CI/CD JenkinsAzure DevOpsGitHub ActionsGitLab CI/CDProgramming & Scripting Python (Mandatory)BashPowerShellDatabases PostgreSQLSQL ServerMySQLService Management ServiceNowJiraITIL ProcessesOperating Systems Linux (RHEL, Ubuntu)Preferred Skills Experience in Banking, Financial Services, Capital Markets, FinTech, or Digital Assets domains.Exposure to MLOps, MLflow, AIOps, and AI/LLM-based operational automation.Hands-on experience with Infrastructure as Code (IaC) using Terraform.Experience working in regulated financial services environments.Key Competencies Strong troubleshooting and analytical skills.Incident management and RCA expertise.Reliability engineering mindset.Automation-first approach.Excellent stakeholder management and communication skills.Cross-functional collaboration and customer-focused attitude.Mandatory Skills ✅ Python ✅ Site24x7 Skills Mandatory Skills : Python, SITE24X7