Lead Site Reliability Engineer

Techholding — United States · Posted ~2 hours ago

Lead Contract Remote

Skills

SRE performance engineering scalability reliability engineering production systems observability

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology services organization is seeking a lead reliability engineer to define performance standards, improve system scalability, and prepare platforms for future growth.

Highlights

Remote contract opportunity focused on improving platform reliability, scalability, and performance through hands-on engineering.

Description

About us: Working at Tech Holding isn't just a job, it's an opportunity to be a part of something bigger. We are a full-service consulting firm that was founded on the premise of delivering predictable outcomes and high-quality solutions to our clients. Our founders and team members have industry experience and have held senior positions in a wide variety of companies – from emerging startups to large Fortune 50 firms – and we have taken our combined experiences and developed a unique approach that is supported by the principles of deep expertise, integrity, transparency, and dependability. The Role: We are looking for a hands-on Lead Site Reliability Engineer for a project based assignment to establish and continuously improve the performance, reliability, and scalability of our platform.This role will define what the platform can reliably sustain today, identify where constraints will emerge, and ensure the organization is prepared to scale before demand arrives. This is not a traditional DevOps role or an advisory architecture position. You will work directly across application services, infrastructure, databases, networking, caching, queues, external dependencies, and operational processes to identify bottlenecks, validate system limits, and lead remediation. You will partner closely with engineering, product, and leadership to provide clear, evidence-based answers around capacity, performance, reliability, risk, and the cost of scaling. Key Responsibilities: Establish performance, throughput, latency, and capacity baselines for critical customer and platform workflowsDefine and maintain SLOs, error budgets, performance budgets, dashboards, alerts, and reliability thresholdsInstrument and analyze the full request path across application services, compute, storage, networking, databases, caches, queues, DNS, registry dependencies, and third-party servicesIdentify system bottlenecks and lead cross-functional remediation efforts with engineering teamsBuild capacity models that show what the platform can sustain, where constraints will emerge, and what additional scale will costLead load, stress, soak, spike, failure, and recovery testing in representative environmentsDevelop realistic demand scenarios for major customers, partnerships, pilots, and high-volume eventsDrive architecture hardening, graceful degradation, dependency-failure planning, and resilience improvementsPartner with Test Automation and Scalability Engineering to establish automated performance testing, regression coverage, and production release gatesOwn technical readiness assessments for major pilots, partnerships, and production launchesCreate operational runbooks for scale-up events, incidents, rollback, recovery, and dependency failuresLead performance and reliability investigations during incidents and ensure lessons are incorporated into future engineering workMake infrastructure cost, performance, and reliability tradeoffs visible to engineering and executive leadershipRecommend capacity and reliability investments before they become production constraints Required Skills: Significant experience in Site Reliability Engineering, performance engineering, platform engineering, distributed systems, or a closely related engineering disciplineExperience supporting production systems with meaningful scale, traffic, latency, or availability requirementsDeep understanding of observability, performance analysis, capacity planning, and reliability engineeringStrong hands-on experience with cloud infrastructure and production distributed systemsDeep knowledge of databases, networking, caching, queueing, compute, storage, and common distributed-system failure modesExperience defining and operating against SLOs, SLIs, error budgets, and production reliability metricsHands-on experience performing load, stress, soak, scalability, and resilience testingAbility to profile systems, diagnose bottlenecks, tune architecture, and work directly with engineering teams to implement improvementsExperience designing for graceful degradation, dependency failures, recovery, and high-demand scenariosStrong incident management and root-cause analysis experienceAbility to translate technical performance and reliability risks into clear business implications for senior leadershipStrong judgment around when systems genuinely require optimization versus when additional complexity is premature Nice to have: Experience operating high-scale SaaS, identity, DNS, registry, infrastructure, or other highly distributed platformsExperience creating capacity-cost models and forecasting infrastructure requirementsExperience building performance and reliability gates into CI/CD pipelinesExperience preparing platforms for significant increases in traffic associated with enterprise customers or strategic partnershipsExperience leading reliability or performance initiatives that span multiple engineering teams What Success Looks Like Within your first several months, you will have: Established measurable throughput, latency, and capacity baselines for critical platform journeysDefined initial SLOs, error budgets, dashboards, alerts, and performance thresholdsIdentified the platform's most significant scalability and reliability constraints and created an actionable remediation roadmapValidated representative high-scale scenarios through load, soak, stress, failure, and recovery testingDeveloped a capacity and cost model showing how the platform can support significant increases in demandEstablished production-readiness criteria and clear go/no-go evidence for major launches and partnershipsCreated repeatable scale-up, incident, rollback, and dependency-failure runbooksGiven leadership a clear, evidence-based understanding of the platform's current capacity envelope and future scaling requirements Employment type: ContractApplicants must be authorized to work for ANY employer in the U.S. We are unable to sponsor or take over sponsorship of an employment Visa at this time Tech Holding is proud to be an Equal Opportunity Employer and is committed to fostering a diverse and inclusive workplace. We welcome applicants from all backgrounds and experiences, and we consider qualified applicants without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, disability, veteran status, or any other legally protected characteristic. If you require accommodation in the application process, please contact our HR