Site Reliability Engineer

Mpch.io โ€” Unknown ยท Posted ~2 days ago

Senior Full-time

Skills

SRE Cloud infrastructure Blockchain infrastructure Monitoring Incident response Automation Cloud Blockchain

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary

A hands-on reliability engineering role maintaining distributed infrastructure, improving resilience, managing production systems, and responding to operational incidents.

Highlights

Work on advanced infrastructure reliability challenges with automation, security, and distributed systems responsibilities.

Description

About the Role MPCH is looking for a hands-on Site Reliability Engineer (SRE) to join our Engineering (InfraOps) team, responsible for the stability, scalability, and security of our blockchain infrastructure. You will be a key operator in a distributed blockchain network environment, maintaining validator nodes and network-critical components across cloud platforms โ€” with a strong bias toward automation, observability, and resilience. This role sits at the intersection of blockchain infrastructure, cloud engineering, and SRE practice. You will be part of an on-call rotation providing 24/7 coverage, and you will be expected to respond to, investigate, and resolve incidents with urgency and rigor. Key Responsibilities Canton Network Operations Operate, maintain, and upgrade Canton Network validator nodes and participant nodes in production environmentsMonitor Canton ledger health, synchronization status, and domain connectivity; respond to degradation eventsManage Canton domain and sequencer configurations, participant onboarding, and topology changesExecute planned maintenance windows including node upgrades, certificate rotations, and key management proceduresParticipate in network governance processes including identity and namespace managementCloud Infrastructure (GCP / AWS) Provision and manage infrastructure on Google Cloud Platform (primary) and AWS (secondary) using both consoles and CLI toolingAdminister compute (GKE, GCE, EKS, EC2), networking (VPCs, firewall rules, load balancers, peering), storage, and IAM across environmentsManage Kubernetes clusters hosting Canton workloads โ€” deployments, statefulsets, resource quotas, PodDisruptionBudgets, and cluster upgradesOwn cloud cost hygiene: right-sizing, reserved capacity planning, and resource tagging disciplineInfrastructure as Code & Automation Author, review, and maintain Terraform modules for all infrastructure provisioning โ€” no manual snowflakesEnforce GitOps workflows; all infrastructure changes go through version control and CI/CD pipelinesLeverage AI-assisted tooling (e.g., GitHub Copilot, Claude, or similar) to accelerate IaC authoring, runbook generation, and incident analysisContinuously improve automation to reduce toil and MTTRSRE / On-Call Duties Participate in a 24/7 on-call rotation with defined escalation pathsRespond to pages, triage alerts, and lead incident response for Canton and infrastructure-layer issuesWrite thorough post-mortems with blameless root cause analysis and tracked action itemsBuild and refine alerting rules, dashboards, and runbooks in partnership with the broader InfraOps teamMaintain and test disaster recovery and business continuity procedures for validator infrastructureSecurity & Compliance Apply hardening standards to all node infrastructure; manage secrets via Vault or equivalent secrets management toolingOwn TLS/mTLS certificate lifecycle for Canton componentsSupport audit and compliance activities including access reviews, change logs, and evidence collection Required Qualifications 3+ years of experience in infrastructure, cloud operations, or SRE roles in production environmentsSolid, hands-on experience with Google Cloud Platform (Compute Engine, GKE, Cloud Networking, IAM, Cloud Monitoring); familiarity with AWS is a plusProficiency with Terraform โ€” you write modules, not just apply them; comfortable with state management, workspaces, and remote backendsStrong Kubernetes operations experience: cluster management, workload troubleshooting, networking (CNI, Ingress, NetworkPolicy), and storageExperience operating blockchain or distributed ledger nodes โ€” Canton / Daml, Hyperledger, Ethereum, or similar; understanding of consensus, finality, and node synchronization conceptsFamiliarity with Canton Network architecture โ€” validators, sequencers, domains, participants, and topology managementProficiency with Linux systems administration, shell scripting (bash/zsh), and standard SRE toolingExperience with observability stacks: metrics (Prometheus/Cloud Monitoring), logging (ELK, Cloud Logging, Loki), tracing, and alertingDemonstrated ability to work an on-call rotation, respond to incidents under pressure, and write high-quality post-mortemsStrong written communication skills โ€” clear incident updates, runbooks, and architecture documentation Foundational IT & General Computing Skills While this is a cloud and blockchain infrastructure role, candidates must also demonstrate solid general computing competency to operate effectively in a distributed, remote-first team environment: Operating Systems: Comfortable working across macOS, Windows, and Linux desktop/laptop environments; able to navigate system settings, manage files, configure network adapters, and perform basic OS-level troubleshooting (e.g., DNS resolution issues, VPN connectivity, permission problems)Hardware Troubleshooting: Ability to diagnose and resolve common computer hardware and peripheral issues โ€” including display, keyboard/mouse, storage, and connectivity problems โ€” and coordinate with IT support or vendors when escalation is neededEmail & Calendar: Proficient with Microsoft Outlook (or equivalent) for email management, calendar scheduling, meeting coordination, and distribution list communication; familiar with shared calendar and resource booking workflowsWord Processing: Able to produce clear, well-formatted documents using Microsoft Word or Google Docs โ€” runbooks, incident reports, onboarding guides, and internal proposalsSpreadsheets: Comfortable with Microsoft Excel or Google Sheets for basic data tasks: building and maintaining tables, applying formulas (SUM, VLOOKUP, IF statements), filtering/sorting, and producing simple charts for capacity tracking or incident metricsRemote Collaboration Tools: Proficient with video conferencing platforms (Zoom, Google Meet, or Teams), chat/messaging tools (Slack or Teams), and shared documentation platforms (Confluence, Notion, or Google Workspace)File & Document Management: Able to organize, version, and share files effectively using cloud storage (Google Drive, OneDrive, or SharePoint); understands basic folder structure hygiene and access permission managementBasic Networking Concepts: Understands home/office network setup sufficiently to self-troubleshoot VPN, Wi-Fi, and connectivity issues that may affect on-call response timeNice to Have Direct production experience running Canton Network validator or domain infrastructureFamiliarity with Daml smart contract concepts and Canton ledger internalsExperience with Helm, Kustomize, or ArgoCD for Kubernetes workload managementExposure to HashiCorp Vault or equivalent secrets/PKI managementWorking knowledge of Python or Go for operational tooling and automationExperience in regulated or compliance-heavy environments (SOC 2, ISO 27001)Familiarity with AI-assisted development workflows and prompt engineering for operational use cases Benefits Salary range: $120-140k, based upon experienceEquity optionsFully remote team