Senior DevOps Engineer

Ignite Tech โ€” Japan ยท Posted ~6 days ago

Senior Remote

Skills

AWS DevOps SRE Cloud Infrastructure Incident Response Infrastructure Automation AI Agents Observability English Grafana Prometheus Datadog PagerDuty OpsGenie Azure

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary

An experienced DevOps engineer will own production operations, automate incident response, build intelligent operational workflows, and improve cloud infrastructure reliability for a large-scale SaaS platform.

Highlights

Build autonomous operational systems, improve reliability through automation, work with large-scale cloud infrastructure, and collaborate with a globally distributed remote engineering team.

Description

Senior DevOps & Autonomous Systems Engineer Your best incident response is the one that never needs you. We run the community and social engagement platform behind some of the world's most recognized brands. A single minute of downtime here doesn't just trigger an internal Slack thread โ€” it puts a Fortune 100 company's customer trust at risk. That reality drives our engineering culture: every human intervention is a prompt to build the automation that makes it unnecessary next time. We're hiring a senior engineer who thrives at the intersection of production operations and intelligent automation. You'll carry real operational responsibility โ€” owning your shift, commanding incidents, safeguarding deployments โ€” while simultaneously engineering the autonomous agents and workflows that absorb more of that operational surface every week. How You'll Spend Your Time Run the shift like it's yours. You're the first voice on the bridge when something breaks. Triage, diagnose, mitigate, restore โ€” and escalate when severity demands it. Uptime on your watch is a personal standard, not a shared abstraction.Engineer the agents that replace the toil. Design, ship, and iterate on autonomous workflows โ€” from alert pre-triage and deployment validation to self-healing sequences and post-incident analysis. When an agent underperforms, you improve the capability, not just the prompt.Guard every production change. Deploys, config updates, cost actions โ€” everything passes through quality gates with a proven rollback. The moment telemetry drifts from plan, you pull back without hesitation.Chase the real root cause and land the fix. Rigorous investigation that separates symptom from cause, followed by the step most orgs skip: shipping the systemic prevention and tracking it to completion. Analysis without action is theater.Turn one-offs into permanent leverage. Every manual fix you apply is raw material for a new agent rule, guardrail, or runbook. Your metric isn't personal throughput โ€” it's the expanding footprint of work the autonomous layer handles without you.Make knowledge outlive the moment. Encode decisions, procedures, and operational context so both agents and humans can retrieve it instantly. On a global async team, institutional memory lives in documentation or it doesn't live at all. What You Bring 5+ years in production-facing roles โ€” SRE, DevOps, Platform Engineering, or Cloud Infrastructure on large-scale SaaS. You've carried a real pager, led real incident bridges, and have the operational scars to show for it.Substantial AWS depth at scale. Multi-AZ, multi-account architectures, infrastructure automation, gated change management, and rollback discipline are second nature. You've navigated major outages and internalized the failure modes.Radical self-direction. You set your own priorities, spot the highest-leverage gap, and close it โ€” without waiting for assignment. When a standard doesn't hold up, you challenge it with a better proposal; you never silently work around it.Hands-on AI-native practice. You delegate real operational work to agents, evaluate output critically, and improve underlying capabilities when results fall short. Proficient with tools like Claude Code, Codex, Warp, or custom agent frameworks โ€” and genuinely excited to test new models the day they drop.AWS Solutions Architect โ€“ Associate or higher โ€” or equivalent battle-proven depth that makes the credential redundant.Fluent, precise English โ€” clear under the pressure of an active incident and equally sharp in long-form post-mortems.Commitment to shift-based coverage. On-call rotations and your designated time-zone window are central to the role, not a footnote.OFAC-clear country of residence. Differentiators We Value Original contributions in agentic operations, AIOps, or intelligent automation โ€” shipped tools, open-source projects, technical posts, or conference talks.Multi-tenant B2B SaaS experience โ€” community platforms, social tools, customer-experience products, or observability systems.Proficiency with observability and alerting stacks: Grafana, Prometheus, Datadog, PagerDuty, OpsGenie. Azure familiarity alongside AWS depth.A story of deep, sustained obsession with a hard technical problem โ€” the kind of focus that reveals how you think, not just what you've shipped. What You'll Take Away You'll help define a new operational discipline โ€” agentic DevOps at enterprise scale โ€” on a platform Fortune 100 brands rely on every day. The agent-engineering instincts, incident patterns, and system-design judgment you develop here are capabilities the industry is still trying to name, let alone hire for. Environment Enterprise gravity, startup clock speed. Contractual SLAs for the world's biggest brands, paired with weekly delivery rhythms, rapid decisions, and a playbook that evolves continuously.Unconstrained tooling. The autonomous harness is the product. Need a larger model, more compute, or a tool we haven't adopted? We invest.Fully remote, distributed global team.