Summary
An AI infrastructure engineer with an SRE mindset is sought to keep user-facing services and production systems reliable, available, and scalable. You will combine pragmatic operations with software engineering, specializing in operating systems, storage, networking, automation, and distributed systems. Expert knowledge of Ansible, Terraform, and Kubernetes and at least seven years of relevant professional experience are required. The role is full-time and on-site in Amsterdam.
Highlights
Work on highly scalable AI infrastructure, combining software engineering with production operations. The role offers significant technical depth across systems, networking, storage, automation, reliability, and distributed systems.
Description
Home Jobs AI infrastructure Engineer (SRE) Amsterdam
T
AI infrastructure Engineer (SRE) Amsterdam
Together AI
On-site Amsterdam Full-time 30 days ago
Apply on Together AI
Engineering
As a AI Infrastructure Engineer (SRE) at Together, you are responsible for keeping all user-facing services and production systems running smoothly.
You are a blend of a pragmatic operator and a software engineer that applies sound engineering principles, operational discipline, and mature automation to our operating environments and codebase.
You specialize in systems (operating systems, storage subsystems, networking), while implementing best practices for availability, reliability and scalability, with varied interests in algorithms and distributed systems.
Requirements
7+ years of professional SRE or related experienceBachelor's degree in Computer Science or a related field or equivalent work experienceExpert knowledge of Ansible (roles, playbooks), Terraform, and KubernetesProficiency in programming/scripting languagesDirect experience in monitoring and observability practicesAdvanced knowledge of cloud servicesAbility to thrive in a collaborative environment involving different stakeholders and subject matter experts
Responsibilities
Be on an on-call (PagerDuty) rotation to respond to incidents that impact availabilityBuild and run our infrastructure with Ansible, Terraform, and Kubernetes to enable scaling to a massive number of concurrent usersBuild monitoring systems to ensure the highest quality service for our customersDesign and implement operational processes (such as deployments and upgrades)Debug production issues across all services and levels of the stackIdentify improvements for the product architecture from the reliability, performance and availability perspectives Plan the growth of Together AIโs infrastructure
About Together AI
Together AI is a research-driven artificial intelligence company.
We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models.
We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancement such as FlashAttention, Hyena, FlexGen, and RedPajama.
We invite you to join a passionate group of researchers and engineers in our journey in building the next generation AI infrastructure.
Equal Opportunity
Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.
Please see our privacy policy at https://www.together.ai/privacy
Interested in this role?
Applications are handled on Together AI's site.
Apply Now
Related Jobs
View all
AI Engineering
S
Senior Product Engineer, Growth & Lifecycle Infrastructure - Music & Audio
Stability AI
Remote
Full-time
T
AI Researcher, Core ML (Turbo)
Together AI
On-site
Full-time
C
Account Solution Architect
CoreWeave
On-site
Full-time
T
Backend Software Engineer - Data Platform & AI Data Products
Together AI
On-site
Full-time
S
AI Applications Ops Lead, GPS
Scale AI
On-site
Full-time
P
Member of Technical Staff (Machine Learning Engineer, Search)
Perplexity
Remote
Full-time
Back to all jobs
Related Jobs
View all
AI Engineering
S
Senior Product Engineer, Growth & Lifecycle Infrastructure - Music & Audio
Stability AI
Remote
Full-time
T
AI Researcher, Core ML (Turbo)
Together AI
On-site
Full-time
C
Account Solution Architect
CoreWeave
On-site
Full-time
T
Backend Software Engineer - Data Platform & AI Data Products
Together AI
On-site
Full-time
S
AI Applications Ops Lead, GPS
Scale AI
On-site
Full-time
P
Member of Technical Staff (Machine Learning Engineer, Search)
Perplexity
Remote
Full-time