Infrastructure Team Lead

Flowaccount โ€” Thailand ยท Posted ~1 day ago

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Description

About the Role We are looking for an experienced Infrastructure Team Lead to own the reliability, scalability, cost-efficiency, and performance of our infrastructure and platform systems. You will lead a team of infrastructure/DevOps engineers, set technical direction through architecture reviews, and drive a strong developer experience through robust CI/CD practices. This is a hands-on leadership role that blends technical depth with team management and cross-functional collaboration. Key Responsibilities 1. Availability & Reliability (Target: 99.9% Uptime) Own the end-to-end reliability of production infrastructure and critical services.Design and implement high-availability architectures (redundancy, failover, load balancing, disaster recovery).Define and maintain SLIs/SLOs/SLAs; track error budgets and drive incident reduction.Lead incident response, root cause analysis (RCA), and post-mortems for outages or degradations.Establish and maintain on-call rotations, runbooks, and escalation procedures. 2. Architecture Review & Governance Lead and facilitate architecture review sessions for new systems, services, and major changes.Define infrastructure standards, best practices, and reference architectures (cloud, networking, security, data).Evaluate new technologies and tools; make build-vs-buy recommendations.Ensure scalability, security, and maintainability are considered in all infrastructure design decisions.Maintain up-to-date architecture documentation and diagrams. 3. Cost Optimization Monitor and manage cloud/infrastructure spend against budget; identify and eliminate waste.Implement cost governance practices (tagging, budgets, alerts, rightsizing, reserved/spot instance strategies).Regularly report on cost trends, savings initiatives, and ROI of infrastructure investments.Balance cost efficiency with performance and reliability requirements. 4. Performance & Application Performance Monitoring (APM) Implement and maintain observability stack (metrics, logging, tracing) using Grafana and Elastic APM.Proactively identify performance bottlenecks across infrastructure, network, and application layers.Define performance benchmarks and drive continuous improvement initiatives.Partner with engineering teams to optimize application and infrastructure performance. 5. CI/CD & Developer Experience Own and continuously improve CI/CD pipelines (GitHub Actions, Jenkins) to increase deployment speed, frequency, and reliability.Champion Infrastructure as Code (IaC) practices using AWS CDK and Terraform.Reduce friction in the developer workflow โ€” build tools, self-service platforms, and automation that improve engineering velocity.Establish standards for environments (dev/staging/prod), branching strategies, and release management.Gather feedback from engineering teams and iterate on internal developer platforms/tools. 6. Team Leadership Manage, mentor, and grow a team of infrastructure/DevOps/SRE engineers.Set team goals aligned with business priorities; conduct performance reviews and career development planning.Foster a culture of ownership, blameless post-mortems, and continuous improvement.Collaborate with Engineering, Security, Product, and Finance stakeholders. Requirements Experience 5+ years in Infrastructure, Platform Engineering, SRE, IMOC center and including 2+ years in a lead or management role.Proven experience designing and operating highly available, production-grade systems at scale.Track record of driving cost optimization initiatives in cloud environments (AWS/GCP).Hands-on experience with observability/APM tools and performance tuning.Strong background building and maintaining CI/CD pipelines and IaC.Technical Skills Cloud platforms: AWS, GCP (certification a plus).Containerization & orchestration: Docker, Kubernetes.IaC tools: Terraform, CDK or similar.CI/CD tools: GitHub Actions, Jenkins or similar.Monitoring/APM: Prometheus, Grafana, Elasticsearch or similar.Scripting/automation: Python, Bash, or Go.Strong understanding of networking, security, and system architecture principles.Soft Skills Strong leadership, communication, and stakeholder management skills.Ability to balance competing priorities (reliability vs. cost vs. speed).Data-driven decision-making and comfort with metrics/reporting.Mentorship mindset and ability to build strong engineering culture.Nice to Have Experience with FinOps practices/tools.Experience in a regulated or high-scale industry (fintech, e-commerce, SaaS).Relevant certifications (AWS Solutions Architect, CKA, etc.). What We Offer Build products used by tens of thousands of businesses.Ship code that reaches customers quickly.High ownership with minimal bureaucracy.Learn from experienced engineers and product leaders.Flexible work hours and 45 days/year work from anywhere.Competitive compensation, health & accident insurance from day one.17โ€“19 public holidays, 12 personal leave days, and 8 annual leave days.A friendly, collaborative team that genuinely enjoys building together.Clear opportunities to grow your career as the company scales. Location BKK BangRak Office (MRT Samyan or BTS Saladaeng) Why join us: ๐Ÿš€ Work with a modern tech stack | ๐Ÿ’ก Impact thousands of users | ๐Ÿค Collaborative, growth-oriented team