Senior Platform Engineer

Gala Solutions Inc — United States · Posted ~4 hours ago

Senior Full-time

Skills

Platform engineering Site reliability engineering DevOps Cloud engineering Systems architecture Automation Observability Incident management Production operations Event-driven architecture SRE Cloud Agentic AI

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Build and operate a modern event-driven platform focused on reliability, automation, observability, and incident management. You will independently make architectural decisions, develop production-ready systems, automate critical capabilities, and prepare the platform for future agentic-AI integrations while partnering with engineering and operations teams.

Highlights

Hands-on senior platform engineering role with substantial architectural ownership and minimal supervision. The position combines reliability, automation, observability, incident management, cloud engineering, and modern AI-assisted development practices.

Description

What We Are Looking For We are looking for a builder—someone who can take ownership of complex platform initiatives, make sound architectural decisions, and deliver production-ready systems with minimal supervision. The ideal candidate combines strong engineering fundamentals, operational excellence, automation expertise, and a passion for leveraging modern AI-assisted development workflows. About the Role We are seeking an experienced platform engineer to design and build a modern, event-driven operational platform that supports reliability, automation, observability, incident management, and future agentic-AI integrations. This is a hands-on engineering role for someone who can independently architect, build, automate, and operationalize critical platform capabilities while collaborating closely with engineering and operations teams. Required Qualifications * 15+ years of experience in platform engineering, site reliability engineering (SRE), DevOps, cloud engineering, or related fields. * Strong experience building production-grade alerting, monitoring, and incident management systems. * Experience with PagerDuty, Fresh Service, ServiceNow, Jira Service Management, or similar platforms. * Deep understanding of event-driven architectures and distributed systems. * Experience designing operational automation and workflow orchestration solutions. * Strong knowledge of cloud platforms such as AWS, Azure, or Google Cloud. * Experience with observability tools, monitoring platforms, logging systems, and operational dashboards. * Strong scripting and software development skills using Python, Go, TypeScript, Java, or similar languages. * Experience implementing secure, scalable, multi-tenant platform solutions. * Excellent documentation and operational process design skills. Preferred Qualifications * Experience building AI-enabled or agentic operational workflows. * Hands-on use of AI-assisted development tools such as Claude Code, Codex, Cursor, Windsurf, or similar platforms. * Experience building internal developer platforms and reliability engineering frameworks. * Knowledge of modern incident management, operational excellence, and service reliability practices. * Experience working in highly regulated or enterprise-scale environments.