Principal Software Engineer, Distributed Systems

Sigma Automate — United States · Posted ~3 hours ago

Full-time

Skills

Distributed systems Site reliability engineering System architecture Observability Reliability engineering Kubernetes

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A principal engineering role focused on designing resilient distributed systems that automatically handle failures and maintain reliable execution. The position combines architecture, reliability practices, and platform engineering.

Highlights

High-impact technical leadership role focused on building reliable distributed systems and improving large-scale platform resilience.

Description

Background We are hiring a Staff / Principal Site Reliability Engineer focused on distributed systems and platform reliability. This is not a traditional DevOps role. You will own the architecture and engineering practices that ensure Sigma workflows execute reliably even when individual components fail. Processes crash. Networks become unavailable. APIs time out. Workers hang. Messages may be delivered more than once. Databases experience contention. Kubernetes pods restart. Your job is to design Sigma so those conditions are expected, detected, and automatically recovered from. The fundamental principle is simple: An accepted Sigma job may fail, but it should never disappear. You will work closely with engineering leadership to evolve Sigma's execution architecture, improve production reliability, establish observability standards, and eliminate classes of distributed-system failures before they affect customers. What You'll Own Design and improve the systems responsible for executing long-running and scheduled infrastructure workflows. Build patterns for: Durable workflow executionIdempotent operationsAutomatic retries and exponential backoffWorker leases and heartbeatsStuck-job detection and recoveryWorkflow timeouts and cancellationFailure isolationDead-letter handlingCheckpointing and workflow resumptionConcurrency controlDistributed locking and lease managementReconciliation loopsGraceful degradation during dependency failuresEvaluate and implement workflow technologies such as Temporal or equivalent systems where appropriate. Messaging and Event Infrastructure Own and improve Sigma's asynchronous execution architecture. Work extensively with technologies such as: KafkaDistributed consumersConsumer groupsMessage delivery semanticsPartitioningConsumer lagBackpressureRetry strategiesEvent orderingDuplicate message handlingPoison-message isolationEnsure event-processing failures cannot silently cause workflows to stop progressing. Database Reliability Help design application patterns that remain reliable under high concurrency. Deeply understand and troubleshoot: PostgreSQL transactionsRow and advisory lockingDeadlocksLock contentionTransaction isolationConnection poolingLong-running transactionsQuery performanceDatabase failoverSchema and migration safetyPartner with application engineers to identify and eliminate patterns that create production contention or reliability risks. Kubernetes and Platform Reliability Improve the reliability of Sigma deployments running in containerized environments. Responsibilities include: Kubernetes workload architecturePod lifecycle and failure recoveryResource limits and capacity planningAutoscalingHealth checksGraceful shutdownRolling deploymentsHigh availabilityInfrastructure-as-codeDisaster recoveryBackup and restore validationDesign systems that tolerate node, pod, process, and dependency failures without losing customer work. Observability Make the state of Sigma understandable in real time. Build and improve: MetricsDistributed tracingStructured loggingDashboardsAlertingWorkflow-level telemetryKafka consumer monitoringDatabase performance monitoringWorker health monitoringQueue depth and backlog monitoringTechnologies may include: OpenTelemetryGrafanaPrometheusSentryCloudWatchEngineers should be able to answer questions such as: Why is this workflow still running?Which worker owns this task?When did it last make progress?Which dependency is causing the delay?Did this operation execute once or multiple times?Can the workflow safely resume?Are scheduled jobs starting when expected? Reliability Engineering Help establish engineering practices expected of mature enterprise platforms. This includes: Service-level objectives and indicatorsError budgetsCapacity planningLoad testingFailure-mode testingChaos testingIncident responseRoot-cause analysisProduction readiness reviewsRunbooksAutomated recoveryDisaster-recovery testingMove Sigma from detecting incidents to preventing and automatically recovering from them. What We're Looking For You have significant experience operating production distributed systems where reliability matters. Strong candidates will have deep experience with several of the following: KubernetesKafka or similar event-streaming systemsPostgreSQLDistributed systemsAsynchronous worker architecturesWorkflow orchestrationTemporal, Cadence, Step Functions, Conductor, or similar systemsPython backend systemsCloud infrastructureInfrastructure-as-codeObservability platformsHigh-availability architecturesYou understand concepts such as: At-least-once deliveryIdempotencyLeases and fencing tokensDistributed locksLeader electionRetry semanticsBackpressureEventual consistencyTransaction boundariesFailure domainsReconciliationSplit-brain scenariosDistributed tracingMost importantly, you have personally diagnosed and improved production systems experiencing real distributed-system failures. What Success Looks Like Within your first several months, you will help Sigma establish an architecture where: Scheduled workflows reliably begin when expectedAccepted work cannot silently disappearWorker crashes do not result in lost workflowsInfrastructure failures automatically recover where possibleLong-running workflows can safely resumeDuplicate execution is prevented or safely handledStuck workflows are automatically detectedDatabase and messaging contention is visible before becoming an outageEngineers can trace a workflow from request through every execution stepCustomer environments can scale without requiring manual infrastructure babysitting