Senior Software Engineer, Reliability

Berry Ai — Taiwan · Posted ~5 hours ago

Senior

Skills

software architecture monitoring observability production readiness Ansible CI/CD incident management Observability Monitoring Edge Computing Cloud

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join a senior engineering team focused on reliability for a large distributed edge-cloud environment. You’ll review architecture, establish monitoring and observability strategies, improve production readiness, maintain Ansible and CI/CD automation, and strengthen incident response through alerting, runbooks, and postmortems.

Highlights

Senior reliability-focused engineering role centered on architecture, monitoring, production readiness, and large-scale operational systems. The position offers ownership of observability, deployment automation, incident practices, and reliability standards across a large distributed fleet.

Description

Berry AI builds AI-powered operations platforms for QSR restaurants — drive-thru analytics, loss prevention, and store management tooling deployed across thousands of locations across the US — and growing. To help us scale with high reliability, we are seeking a Senior Software Engineer. The role is focused on architecture review, monitoring, and production-readiness. What you'll work on Analyze and optimize our edge-cloud architecture using modern monitoring tools, continuously driven by data insights. Formulate and execute a robust monitoring strategy tailored for our thousands of edge servers and cameras. Establish standards and guide our software development workflows to mitigate or clarify risks before they arise. Contribute to the operational backbone — Ansible playbooks, CI/CD pipelines, and observability — that lets us deploy, monitor, and operate the fleet. Strengthen our on-call practice — leading postmortems, sharpening alerting and runbooks, and turning incidents into durable fixes. You're a strong fit if you have Over 5 years of experience delivering production software, with deep expertise in architecture design and conducting code reviews. Experience mapping out complex data flows within massive codebases. Disciplined practice of DevOps and SRE — metrics, structured logging, tracing, runbooks. Proven ability to validate system reliability by designing and executing load tests, stress tests, and chaos engineering experiments. Strong cloud and system-design fundamentals — designing for AWS/GCP/Azure at scale (compute, storage, autoscaling), clean REST/GraphQL contracts, and async/event-driven pipelines. Strong data modeling experience at scale — relational and NoSQL databases, dimensional modeling for analytics, and robust ETL/ELT pipeline design. Bonus points Experience managing and deploying to massive fleets of on-premise or edge devices (IoT). Hands-on experience working with open-source software, including modifying, extending, and improving existing projects—not just using them. Experience with real-time/streaming systems (RTSP, WebRTC, MediaMTX). Experience Kernel performance insights - scheduling, context switching, hardware acceleration. Deep knowledge of GPU analytics and performance optimization for edge-based AI deployments. Our engineering culture Small team, high ownership, fast feedback from engineer teams — and the operational rigor to make that velocity sustainable. We hack. We drive. We deliver. No technical limitation stands in our way — until we turn what seemed impossible into something real. ============================ Interview Process Online (Google Meet)Team Lead Interview (0.5 - 1 hr)OnsiteTechnical Interview (2.5 hrs)CEO & VP Interview (1.5 hrs)PM/Engineer Interview (0.5 hr)HR Interview (0.5 hr)