Senior Site Reliability Engineer

Gft Technologies Apac Gcc — Hong Kong Sar · Posted ~2 hours ago

Senior Full-time

Skills

site reliability engineering production operations operational readiness cloud infrastructure monitoring incident management automation

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An experienced Site Reliability Engineer is sought to support production operations and operational readiness for a next-generation digital asset platform. You will focus on reliability, automation, monitoring, infrastructure, and efficient production operations while helping ensure critical services remain secure and resilient.

Highlights

Work on production operations for a next-generation digital asset platform within an international technology environment. The position combines reliability engineering, automation, operational readiness, and modern cloud infrastructure.

Description

Job description: About GFT GFT Technologies is an AI-centric global digital transformation company. We design advanced data and AI transformation solutions, modernize technology architectures and develop next-generation core systems for industry leaders in Banking, Insurance, Manufacturing and Robotics. Partnering closely with our clients, we push boundaries to unlock their full potential. With deep industry expertise, cutting-edge technology, and a strong partner ecosystem, GFT delivers responsible AI-centric solutions that combine engineering excellence, high-performance delivery, and cost efficiency. Our team of 12,000+ technology experts operate in 20+ countries worldwide offering career opportunities at the forefront of software innovation. Overview We are seeking an experienced Senior Site Reliability Engineer (SRE) / Production Operations Engineer to support the production operations and operational readiness of a next-generation Stablecoin and Digital Assets Platform. This role is focused on ensuring the platform operates securely, reliably, and efficiently in production while enhancing observability, monitoring, alerting, and operational support capabilities. The successful candidate will work closely with Product, Engineering, Infrastructure, Security, Compliance, Risk, and Operations teams to support production environments, troubleshoot issues, improve operational processes, and continuously improve platform reliability and resilience. Key Responsibilities Production Operations Own day-to-day production support activities across platform environments. Monitor platform availability, performance, and operational health. Investigate and troubleshoot production issues across application, infrastructure, third-party services, and blockchain components. Coordinate incident management, root cause analysis, and issue resolution. Participate in after-hours production support and critical incident response when required. Observability & Reliability Design and enhance platform observability, monitoring, and alerting capabilities. Implement centralized dashboards, logging, and operational visibility across platform components. Improve incident detection, troubleshooting, and root cause analysis through effective use of logs, metrics, and telemetry. Identify reliability risks and drive improvements in platform resilience and operational maturity. Infrastructure & Platform Operations Support cloud infrastructure and production environments. Support deployment processes, CI/CD pipelines, and infrastructure automation. Work closely with engineering teams to improve operational stability and release quality. Assist in diagnosing infrastructure and performance-related issues. Security, Compliance & Stakeholder Coordination Support operational controls required for a regulated financial platform. Participate in security reviews, audits, disaster recovery exercises, and operational readiness activities. Collaborate with Engineering, Product, Operations, Compliance, Risk, and external technology providers. Provide clear operational reporting and incident updates to stakeholders. Required Experience 7+ years of experience in Site Reliability Engineering (SRE), Production Support, Platform Operations, DevOps, Infrastructure Engineering, or similar roles. Experience supporting business-critical production systems. Strong experience with monitoring, observability, alerting, incident management, and production troubleshooting. Experience with cloud platforms, preferably AWS. Familiarity with CI/CD pipelines and infrastructure automation. Strong troubleshooting and problem-solving skills. Experience working in financial services, fintech, payments, digital banking, or other regulated environments. Preferred Experience Experience supporting blockchain, digital asset, stablecoin, payment, or fintech platforms. Experience with observability tools such as Grafana, OpenTelemetry, CloudWatch, ELK, Datadog, Splunk, or similar platforms. Experience with Kubernetes and containerized workloads. Experience with Linux administration. Experience with scripting and automation (Python, Bash, PowerShell, or similar). Experience operating or supporting 24x7 production services. (Note: Due to the high volume of applications we receive, we are unable to respond to every candidate individually. If you have not received a response from GFT regarding your application within 10 workdays, please consider that we have decided to proceed with other candidates. We truly appreciate your interest in GFT and thank you for your understanding) Profile description: Required Experience 7+ years of experience in Site Reliability Engineering (SRE), Production Support, Platform Operations, DevOps, Infrastructure Engineering, or similar roles. Experience supporting business-critical production systems. Strong experience with monitoring, observability, alerting, incident management, and production troubleshooting. Experience with cloud platforms, preferably AWS. Familiarity with CI/CD pipelines and infrastructure automation. Strong troubleshooting and problem-solving skills. Experience working in financial services, fintech, payments, digital banking, or other regulated environments. Preferred Experience Experience supporting blockchain, digital asset, stablecoin, payment, or fintech platforms. Experience with observability tools such as Grafana, OpenTelemetry, CloudWatch, ELK, Datadog, Splunk, or similar platforms. Experience with Kubernetes and containerized workloads. Experience with Linux administration. Experience with scripting and automation (Python, Bash, PowerShell, or similar). Experience operating or supporting 24x7 production services. We offer: We build a professional & fun working environment. We focus on your growth, yes the long-term growth. We develop the future-ready digital bank platform.