Site Reliability Engineer

Wearehaystack โ€” United States ยท Posted ~2 days ago

Mid Visa History โœ“

Skills

Site reliability engineering Observability Metrics Logs Traces Automation Scripting Incident response Troubleshooting Technical documentation

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary

Join a global-scale technology environment as a Site Reliability Engineer, supporting projects and operational processes that improve system reliability. You will automate workflows, strengthen incident response, troubleshoot routine and complex issues, and use observability solutions to collect and analyze metrics, logs, and traces. The role also involves documentation, knowledge sharing, and close collaboration with development teams and business stakeholders.

Highlights

Work on reliability engineering for a large-scale global technology environment, with exposure to observability, automation, incident response, troubleshooting, and cross-functional collaboration.

Description

We're working with a global technology company that connects consumers, financial institutions, merchants, governments, and businesses worldwide, enabling them to make and receive electronic payments securely and conveniently. The Role Independently execute key elements of projects/processes within Site Reliability Engineering. Assist in evaluating operational requirements and developing technical solutions within existing frameworks. Support automation and scripting efforts to improve operational workflows and incident response processes. Troubleshoot and resolve routine and some complex system issues. Contribute to documentation, knowledge sharing, and best practices to enhance team operational procedures. Collaborate with development teams and stakeholders to ensure reliability solutions align with technical and business needs. What You'll Need Experience with observability solutions, enabling the collection, analysis, and visualization of metrics, logs, and traces. Ability to write and maintain code and scripts to automate tasks and build operational tools. Capability to configure, operate, and troubleshoot Linux/Unix systems and network components. Experience designing, deploying, and managing applications and infrastructure on cloud platforms (e.g., AWS, Azure, GCP). Ability to design and operate systems for high availability, fault tolerance, and disaster recovery. Understanding of DevOps principles and practices, including CI/CD pipelines, containerization, and orchestration. What's On Offer Opportunity to work on essential services that power global operations. Be part of a culture that celebrates individual strengths, views, and experiences. Flexibility to shape a career across disciplines and continents. Opportunity to work alongside experts and leaders at every level of the business. Apply via Haystack today!