Summary
✨ AI‑Generated
An AI technology company is seeking an infrastructure engineer to build and operate the platform behind continuously running AI services. The position combines cloud and reliability engineering with production software development, automation, observability, security, deployment safety, and cost management.
Highlights
Own production infrastructure for AI services while combining cloud engineering, software development, reliability, observability, security, deployment safety, and cost optimization.
Description
Palona's AI agents operate continuously in production, handle real-time guest interactions, integrate with restaurant systems, and face sharp traffic peaks.
Infrastructure is therefore part of the product: latency, reliability, deployment safety, observability, security, and cost directly shape the guest and operator experience.
We are looking for an Infrastructure Engineer who combines cloud and reliability depth with strong software engineering judgment.
You will build and operate the platform beneath Palona's AI products, improve how engineers ship, and turn production signals into durable system improvements.
This is not a ticket-driven IT or operations role.
You will write production code, design systems, automate repetitive work, and own outcomes across the full service lifecycle.
Our current environment includes Python services, Docker, AWS and selected Azure services, ECS and Lambda workloads, API Gateway, load balancers, relational data systems, OpenTofu/Terraform, Datadog, and CI/CD automation.
We value the ability to learn and make sound tradeoffs more than exact tool-for-tool matching.
What you will own:
Design, build, and evolve secure, scalable cloud infrastructure for real-time AI services and customer-facing applicationsImprove service reliability through clear SLOs, actionable observability, capacity planning, failure testing, and pragmatic incident preventionBuild deployment and release systems that make production changes fast, repeatable, auditable, and safeOwn infrastructure as code, environment consistency, and reusable platform patterns across development, staging, and productionPartner with product and AI engineers on architecture, performance, data flows, and operational readiness for new capabilitiesDiagnose complex distributed-system failures across application, network, database, model-provider, and third-party integration boundariesReduce infrastructure and model-serving cost without compromising customer experience or engineering velocityStrengthen secrets management, access controls, backup and recovery, vulnerability management, and other practical security foundationsBuild internal tooling and paved paths that let engineers ship and operate services with less manual workParticipate in incident response and turn incidents into better systems, automation, documentation, and engineering judgment
Requirements
3+ years industrial experience in relevant technical domainStrong software engineering fundamentals and experience building or operating production distributed systemsHands-on experience with a major cloud platform; AWS experience is especially relevantExperience with containers, infrastructure as code, CI/CD, monitoring, alerting, and production debuggingAbility to write reliable automation and services in Python or another modern programming languageSound judgment around availability, latency, scalability, security, and cost tradeoffsA track record of taking ambiguous operational problems from diagnosis through durable resolutionClear communication during architecture reviews, launches, and incidentsAI-native working habits and curiosity about the operational behavior of LLM- and agent-powered systems
Benefits
Competitive Salary and Stock Option PlanMedical, dental, vision, retirement, leave, and disability benefits as applicableFamily LeaveShort Term & Long Term DisabilityPaid time off and company holidaysLearning and development support