Site Reliability Engineer
Tekclap โ United States ยท Posted ~22 hours ago
๐ Log in to save this job, tailor your resume & track your apply process โ 7 days free, no card needed.
Log in to add to target listDescription
PLEASE only submit candidates that have OpenSearch deployment ON KUBERNETES experience.
Possible convert to hire after 1 year
About the Role
Client is seeking a Site Reliability Engineer โ OpenSearch to help ensure the highest levels of availability, performance, scalability, and Quality of Service (QoS) for mission-critical cloud services.
This role will focus on the reliability, operations, automation, and continuous improvement of distributed search and analytics platforms built on OpenSearch, while working in a diverse, globally distributed team environment.
The ideal candidate brings deep experience in site reliability engineering, DevOps, cloud operations, automation, observability, and distributed systems, with proven hands-on expertise architecting, building, deploying, operating, and optimizing high-performance OpenSearch clusters and platforms from the ground up in production environments.
Responsibilities
Provision, build, deploy, monitor, operate, and support cloud services in a globally distributed team environmentArchitect, build, deploy, and maintain high-performance OpenSearch clusters and platforms from the ground upAdminister and optimize OpenSearch environments for high availability, resiliency, scalability, security, and performanceMonitor and troubleshoot cluster health, node performance, indexing throughput, search latency, shard allocation, replication, and storage utilizationAnalyze and resolve operational issues, platform instability, and production incidents across infrastructure, platform, and application layersConduct incident response, root cause analysis, and post-incident remediation to drive continuous improvementMaintain the integrity and security of servers, systems, and OpenSearch platform infrastructureSupport platform lifecycle activities including installation, configuration, upgrades, patching, hotfixes, backup, restore, and disaster recoveryDevelop and maintain monitoring policies, alerting standards, operational runbooks, and support proceduresAutomate testing, deployment, scaling, recovery, and operational workflows for OpenSearch and related cloud servicesEnsure proper resource allocation and capacity planning across compute, memory, storage, and network resourcesPartner with product development and engineering teams to design and enhance service reliability and operational readinessDevelop and implement testing strategies and document results for platform changes and operational improvementsSupport log ingestion, index management, retention policies, lifecycle management, and search performance tuningWork in a diverse environment and cross-train with other global team membersParticipate in an on-call rotation and support weekend or after-hours operational needs as required
Qualifications
Expert with Kubernetes, including troubleshooting, operations, management, and configuration of complex Kubernetes services.Proven hands-on expertise designing, building, deploying, supporting, and maintaining OpenSearch clusters and platforms from scratch in production environmentsStrong experience with OpenSearch administration, cluster architecture, performance tuning, scaling, upgrades, and troubleshootingExperience with index design, shard and replica strategy, cluster sizing, node management, snapshot/restore, backup, and disaster recoveryStrong understanding of distributed systems, search platforms, indexing pipelines, query optimization, and high-availability architecturesExpertise with GitExpertise with Concourse, including setup, management, and troubleshooting of new pipelinesExpertise with Linux, specifically SUSE and UbuntuExpertise with Kafka, Zookeeper, and Big Data technologiesExpert in development of automation for testing, deployment, scalability, and management of cloud servicesExpertise with building, implementing, and/or supporting cloud monitoring toolsExpert knowledge of cloud computing, infrastructure operations, and databasesExpert understanding of web services, networking, virtualization, and internet protocolsAbility to multitask and handle various projects, deadlines, and changing prioritiesExcellent communication and prioritization skillsExpertise with security fundamentals as they pertain to SaaS multi-tenant application systemsStrong interpersonal, presentation, and customer service skills
Preferred Skills
Experience with AWS services including Route 53, EC2, S3, CloudWatch, DynamoDB, RDS, IAM, ACM, KMS, and VPCExperience deploying and operating OpenSearch in AWS-based environmentsExperience with Cloud Foundry-based environmentsExperience with Jenkins, Chef, and/or TerraformExposure to and understanding of troubleshooting IP networks and application stacksExperience with observability tools such as Prometheus and GrafanaExperience with log ingestion pipelines, index lifecycle management, retention strategies, and search platform security controlsFamiliarity with capacity forecasting, performance benchmarking, and resilience testing for distributed search platforms
We have 77,319 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume โ in under a minute we'll analyze all 77,319 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume