Infrastructure Team Lead
Flowaccount โ Thailand ยท Posted ~1 day ago
๐ Log in to save this job, tailor your resume & track your apply process โ 7 days free, no card needed.
Log in to add to target listDescription
About the Role
We are looking for an experienced Infrastructure Team Lead to own the reliability, scalability, cost-efficiency, and performance of our infrastructure and platform systems.
You will lead a team of infrastructure/DevOps engineers, set technical direction through architecture reviews, and drive a strong developer experience through robust CI/CD practices.
This is a hands-on leadership role that blends technical depth with team management and cross-functional collaboration.
Key Responsibilities
1.
Availability & Reliability (Target: 99.9% Uptime)
Own the end-to-end reliability of production infrastructure and critical services.Design and implement high-availability architectures (redundancy, failover, load balancing, disaster recovery).Define and maintain SLIs/SLOs/SLAs; track error budgets and drive incident reduction.Lead incident response, root cause analysis (RCA), and post-mortems for outages or degradations.Establish and maintain on-call rotations, runbooks, and escalation procedures.
2.
Architecture Review & Governance
Lead and facilitate architecture review sessions for new systems, services, and major changes.Define infrastructure standards, best practices, and reference architectures (cloud, networking, security, data).Evaluate new technologies and tools; make build-vs-buy recommendations.Ensure scalability, security, and maintainability are considered in all infrastructure design decisions.Maintain up-to-date architecture documentation and diagrams.
3.
Cost Optimization
Monitor and manage cloud/infrastructure spend against budget; identify and eliminate waste.Implement cost governance practices (tagging, budgets, alerts, rightsizing, reserved/spot instance strategies).Regularly report on cost trends, savings initiatives, and ROI of infrastructure investments.Balance cost efficiency with performance and reliability requirements.
4.
Performance & Application Performance Monitoring (APM)
Implement and maintain observability stack (metrics, logging, tracing) using Grafana and Elastic APM.Proactively identify performance bottlenecks across infrastructure, network, and application layers.Define performance benchmarks and drive continuous improvement initiatives.Partner with engineering teams to optimize application and infrastructure performance.
5.
CI/CD & Developer Experience
Own and continuously improve CI/CD pipelines (GitHub Actions, Jenkins) to increase deployment speed, frequency, and reliability.Champion Infrastructure as Code (IaC) practices using AWS CDK and Terraform.Reduce friction in the developer workflow โ build tools, self-service platforms, and automation that improve engineering velocity.Establish standards for environments (dev/staging/prod), branching strategies, and release management.Gather feedback from engineering teams and iterate on internal developer platforms/tools.
6.
Team Leadership
Manage, mentor, and grow a team of infrastructure/DevOps/SRE engineers.Set team goals aligned with business priorities; conduct performance reviews and career development planning.Foster a culture of ownership, blameless post-mortems, and continuous improvement.Collaborate with Engineering, Security, Product, and Finance stakeholders.
Requirements
Experience
5+ years in Infrastructure, Platform Engineering, SRE, IMOC center and including 2+ years in a lead or management role.Proven experience designing and operating highly available, production-grade systems at scale.Track record of driving cost optimization initiatives in cloud environments (AWS/GCP).Hands-on experience with observability/APM tools and performance tuning.Strong background building and maintaining CI/CD pipelines and IaC.Technical Skills
Cloud platforms: AWS, GCP (certification a plus).Containerization & orchestration: Docker, Kubernetes.IaC tools: Terraform, CDK or similar.CI/CD tools: GitHub Actions, Jenkins or similar.Monitoring/APM: Prometheus, Grafana, Elasticsearch or similar.Scripting/automation: Python, Bash, or Go.Strong understanding of networking, security, and system architecture principles.Soft Skills
Strong leadership, communication, and stakeholder management skills.Ability to balance competing priorities (reliability vs.
cost vs.
speed).Data-driven decision-making and comfort with metrics/reporting.Mentorship mindset and ability to build strong engineering culture.Nice to Have
Experience with FinOps practices/tools.Experience in a regulated or high-scale industry (fintech, e-commerce, SaaS).Relevant certifications (AWS Solutions Architect, CKA, etc.).
What We Offer
Build products used by tens of thousands of businesses.Ship code that reaches customers quickly.High ownership with minimal bureaucracy.Learn from experienced engineers and product leaders.Flexible work hours and 45 days/year work from anywhere.Competitive compensation, health & accident insurance from day one.17โ19 public holidays, 12 personal leave days, and 8 annual leave days.A friendly, collaborative team that genuinely enjoys building together.Clear opportunities to grow your career as the company scales.
Location
BKK BangRak Office (MRT Samyan or BTS Saladaeng)
Why join us:
๐ Work with a modern tech stack | ๐ก Impact thousands of users | ๐ค Collaborative, growth-oriented team
We have 66,600 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume โ in under a minute we'll analyze all 66,600 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume