Summary
✨ AI‑Generated
Build and operate a modern event-driven platform focused on reliability, automation, observability, and incident management. You will independently make architectural decisions, develop production-ready systems, automate critical capabilities, and prepare the platform for future agentic-AI integrations while partnering with engineering and operations teams.
Highlights
Hands-on senior platform engineering role with substantial architectural ownership and minimal supervision. The position combines reliability, automation, observability, incident management, cloud engineering, and modern AI-assisted development practices.
Description
What We Are Looking For
We are looking for a builder—someone who can take ownership of complex platform initiatives, make sound architectural decisions, and deliver production-ready systems with minimal supervision.
The ideal candidate combines strong engineering fundamentals, operational excellence, automation expertise, and a passion for leveraging modern AI-assisted development workflows.
About the Role
We are seeking an experienced platform engineer to design and build a modern, event-driven operational platform that supports reliability, automation, observability, incident management, and future agentic-AI integrations.
This is a hands-on engineering role for someone who can independently architect, build, automate, and operationalize critical platform capabilities while collaborating closely with engineering and operations teams.
Required Qualifications
* 15+ years of experience in platform engineering, site reliability engineering (SRE), DevOps, cloud engineering, or related fields.
* Strong experience building production-grade alerting, monitoring, and incident management systems.
* Experience with PagerDuty, Fresh Service, ServiceNow, Jira Service Management, or similar platforms.
* Deep understanding of event-driven architectures and distributed systems.
* Experience designing operational automation and workflow orchestration solutions.
* Strong knowledge of cloud platforms such as AWS, Azure, or Google Cloud.
* Experience with observability tools, monitoring platforms, logging systems, and operational dashboards.
* Strong scripting and software development skills using Python, Go, TypeScript, Java, or similar languages.
* Experience implementing secure, scalable, multi-tenant platform solutions.
* Excellent documentation and operational process design skills.
Preferred Qualifications
* Experience building AI-enabled or agentic operational workflows.
* Hands-on use of AI-assisted development tools such as Claude Code, Codex, Cursor, Windsurf, or similar platforms.
* Experience building internal developer platforms and reliability engineering frameworks.
* Knowledge of modern incident management, operational excellence, and service reliability practices.
* Experience working in highly regulated or enterprise-scale environments.