Skills
DevOps
SRE
AWS
Python
Terraform
Infrastructure as Code
CI/CD
GitHub Actions
GitLab CI
ArgoCD
Linux
Networking
Monitoring and alerting
Reliability engineering
Performance tuning
Cost optimization
Bash
Go
Cloud infrastructure
Automation
Cloudflare
Prometheus
VictoriaLogs
VictoriaMetrics
VictoriaTraces
Grafana
Summary
Join a high-intensity engineering environment where you’ll architect, automate, and scale cloud infrastructure for systems serving millions of users. You’ll build reliable CI/CD pipelines, strengthen observability and security, optimize performance and costs, and collaborate with software and machine-learning teams to deploy demanding workloads. The position offers broad ownership across the DevOps lifecycle and the opportunity to tackle complex infrastructure challenges at significant scale.
Highlights
High-impact DevOps role focused on scalable cloud infrastructure supporting millions of users. Strong ownership across infrastructure, deployment, monitoring, and incident response, with exposure to complex AI and ML workloads and opportunities to work on reliability, performance, security, and cost optimization.
Description
Why work at Higgsfield AI?
Higgsfield AI is the fastest-scaling generative AI company in history, hitting $500M in annual revenue run rate, 25M+ users worldwide, 6M+ generations per day, and powering 390 of Fortune 500 brands.
We're building at the absolute frontier of AI-powered video creation and next-generation creative tools.
Joining Higgsfield means becoming part of a high-impact team shaping the future of AI-native experiences, at a company that isn't just moving fast, but rewriting what fast looks like.
Who We Are Looking For
Everyone at Higgsfield is an A-player.
We're looking for teammates who bring:
Deep DevOps expertise, with a strong focus on cloud infrastructure and automation.
A systems thinker who can design resilient, scalable, and observable platform Readiness to thrive in a high-intensity startup environment, where priorities shi fast.
Passion for building highly available systems that support millions of users.
What you will work on
Architect, automate, and scale our cloud infrastructure (AWS, Cloudflare)
Build and manage CI/CD pipelines for rapid, reliable deployments Ensure system reliability, security, and observability across all environments.
Optimize performance, cost, and scalability of cloud resources.
Collaborate with backend and ML teams to deploy complex workloads efficiently.
Own the full DevOps lifecycle: infrastructure, deployment, monitoring, and incident response.
Your must haves
3+ years of experience with DevOps / SRE roles.
Strong expertise in AWS, PythonProven track record in infrastructure as code (Terraform).
Experience with CI/CD tools (GitHub Actions, GitLab CI, ArgoCD).
Strong Linux and networking fundamentals.
Knowledge of monitoring & alerting (Prometheus, VictoriaLogs/Metrics/Traces, Grafana).
Performance tuning, reliability engineering, and cost optimization.
Strong skills Python/Bash/Go for automation and tooling.
Independent, able to lead complex infrastructure projects.
Willingness to work in a fast-paced startup environment (80+ hours per week).
The deal
Competitive salary in USD.
On-site role in Almaty office.
Full-time position.