Senior Site Reliability Engineer

Talentdome Staffing — United States · Posted ~2 hours ago

Senior Full-time Remote

Skills

SRE Infrastructure engineering Production systems Cloud operations Automation Cloud Infrastructure Production Systems

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior reliability engineering role focused on building stable production infrastructure, handling large-scale data workloads, and improving system performance.

Highlights

Fully remote senior infrastructure role with high ownership and complex scaling challenges.

Description

About the job Position: Senior Site Reliability Engineer (SRE) / Infrastructure Engineer Location: 100% Remote (US-Based) Company Type: High-Growth, AI-Driven Narrative Intelligence Startup About Our Client Our client is building the "storm tracker for the internet". By developing a cutting-edge platform focused on narrative intelligence, they analyze how data and narratives propagate online, tracking their origins, identifying the entities driving them, and predicting how they will evolve. Having successfully commercialized monetized MVPs, they are currently scaling their architecture into highly stable, enterprise-grade production systems to support massive, real-time data ingestion and search capabilities. They are seeking a highly autonomous, veteran infrastructure specialist who thrives on solving complex scaling and high-throughput challenges. If you are looking to take complete operational ownership of a production environment processing massive data flows without being bogged down by feature coding, this role is built for you. What Our Client Offers True Operational Autonomy: The opportunity to architect and scale greenfield deployments for a rapidly expanding AI data platform.High-Caliber Environment: Collaborate directly with an elite team of backend engineers and machine learning R&D specialists.Flexible, Modern Workspace: Enjoy 100% remote working flexibility across the United States.Competitive Compensation: Top-tier base salary paired with an open, highly competitive equity package. Job Responsibilities Infrastructure Orchestration: Maintain, optimize, and expand the core infrastructure, ensuring everything is cleanly declared via Terraform and managed across high-performance Kubernetes clusters.High-Throughput Scaling: Design and manage environments capable of sustaining immense data ingestion scaling, high-throughput pipelines, and massive search database operations.GPU Application Deployment: Collaborate with the R&D team to successfully deploy, optimize, and manage highly specialized machine learning and AI applications running on GPUs.System Optimization & Reliability: Partner closely with backend teams to heavily optimize production Java deployments and Python workflows, guaranteeing maximum uptime, high availability, and seamless scaling.Technical Leadership: Serve as a foundational pillar for infrastructure architecture, establishing operational best practices without requiring handholding or micro-management. Qualifications and Requirements Experience: 8+ years of dedicated, hands-on experience with Kubernetes and Terraform, with ideally 15+ years of total technical experience in infrastructure or site reliability engineering.Core Focus: Deep architectural mastery of deployment systems, cluster orchestration, and high-availability scaling. (Note: We are strictly seeking an SRE/infrastructure specialist focused on reliable deployments, not a software engineer who builds websites).Cloud Infrastructure: Proven cloud hosting experience, with strong proficiency in AWS. Exposure to or experience with GCP is a significant advantage for supporting R&D workflows.Specialized Skills: Concrete experience deploying and scaling application workflows that interface with GPUs and high-volume data ingestion layers.Application Context: Familiarity with or exposure to optimizing runtime environments for Java and Python applications is highly beneficial.Leadership Traits: Exceptional self-direction and problem-solving capability, with the professional maturity to eventually step into a formal leadership role as the infrastructure team expands. Logistics Location: 100% Remote (US).Compensation: Competitive base salary ranging from $170,000 to $190,000 + open to equity incentives. Please note only US Citizens and GCH will be accepted.