AI Infrastructure Engineer

Palona Ai — United States · Posted ~23 hours ago

Full-time

Skills

cloud infrastructure site reliability engineering software engineering Python Docker AWS ECS Lambda API Gateway load balancing relational databases Infrastructure as Code Terraform OpenTofu observability security production operations Azure AWS Lambda Load Balancers Relational Databases Datadog

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An AI infrastructure engineering role focused on building and operating the cloud platform behind continuously running, real-time AI services. You will combine reliability expertise with strong software engineering judgment, developing production code, designing resilient systems, automating repetitive processes, and turning operational signals into lasting improvements. The environment includes Python services, containerized workloads, major cloud platforms, serverless computing, APIs, relational data systems, infrastructure-as-code, and modern observability tooling.

Highlights

High-impact infrastructure role with ownership across the full service lifecycle. Opportunity to write production software, design scalable systems, automate engineering workflows, improve reliability and observability, and directly influence performance, security, deployment safety, and cost.

Description

Palona's AI agents operate continuously in production, handle real-time guest interactions, integrate with restaurant systems, and face sharp traffic peaks. Infrastructure is therefore part of the product: latency, reliability, deployment safety, observability, security, and cost directly shape the guest and operator experience. We are looking for an Infrastructure Engineer who combines cloud and reliability depth with strong software engineering judgment. You will build and operate the platform beneath Palona's AI products, improve how engineers ship, and turn production signals into durable system improvements. This is not a ticket-driven IT or operations role. You will write production code, design systems, automate repetitive work, and own outcomes across the full service lifecycle. Our current environment includes Python services, Docker, AWS and selected Azure services, ECS and Lambda workloads, API Gateway, load balancers, relational data systems, OpenTofu/Terraform, Datadog, and CI/CD automation. We value the ability to learn and make sound tradeoffs more than exact tool-for-tool matching. What you will own: Design, build, and evolve secure, scalable cloud infrastructure for real-time AI services and customer-facing applicationsImprove service reliability through clear SLOs, actionable observability, capacity planning, failure testing, and pragmatic incident preventionBuild deployment and release systems that make production changes fast, repeatable, auditable, and safeOwn infrastructure as code, environment consistency, and reusable platform patterns across development, staging, and productionPartner with product and AI engineers on architecture, performance, data flows, and operational readiness for new capabilitiesDiagnose complex distributed-system failures across application, network, database, model-provider, and third-party integration boundariesReduce infrastructure and model-serving cost without compromising customer experience or engineering velocityStrengthen secrets management, access controls, backup and recovery, vulnerability management, and other practical security foundationsBuild internal tooling and paved paths that let engineers ship and operate services with less manual workParticipate in incident response and turn incidents into better systems, automation, documentation, and engineering judgment Requirements 3+ years industrial experience in relevant technical domainStrong software engineering fundamentals and experience building or operating production distributed systemsHands-on experience with a major cloud platform; AWS experience is especially relevantExperience with containers, infrastructure as code, CI/CD, monitoring, alerting, and production debuggingAbility to write reliable automation and services in Python or another modern programming languageSound judgment around availability, latency, scalability, security, and cost tradeoffsA track record of taking ambiguous operational problems from diagnosis through durable resolutionClear communication during architecture reviews, launches, and incidentsAI-native working habits and curiosity about the operational behavior of LLM- and agent-powered systems Benefits Competitive Salary and Stock Option PlanMedical, dental, vision, retirement, leave, and disability benefits as applicableFamily LeaveShort Term & Long Term DisabilityPaid time off and company holidaysLearning and development support