Summary
✨ AI‑Generated
A senior Site Reliability Engineering role focused on building and operating reliable, resilient, and highly available distributed systems. You will combine software and systems engineering practices to support cloud services and critical infrastructure, while monitoring reliability, capacity, and performance. The role provides exposure to large-scale systems, modern cloud technologies, and technically challenging reliability work within an international engineering environment.
Highlights
Senior role focused on large-scale reliable systems, cloud infrastructure, resilience, uptime, performance, and capacity. The position offers opportunities to work with modern technologies, contribute to challenging development activities, and strengthen technical expertise in an international environment.
Description
We are looking for a Senior Site Reliability Engineer who is interested in an opportunity to work for an innovative hospitality company with cutting-edge technologies, with new development activities and challenges ahead.
Our international team members share a common desire to develop brilliant products on reliable and resilient systems, along with their own skills.
We run our services in Azure and traditional data centers.
Take a chance to make a valuable contribution and enhance your professional skills.
About the job:
Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems.
SRE ensures that cloud services—both our internally critical and our externally-visible systems—have reliability, uptime appropriate to customer needs, and a fast rate of improvement.
Additionally, SREs will keep an ever-watchful eye on our systems' capacity and performance.
On the SRE team, you’ll have the opportunity to manage the complex challenges of scale that are unique to the project while using your expertise in coding, algorithms, complexity analysis, and large-scale system design.
You will provide scalable, reliable, durable, and secure services using a customer-first approach while innovating technically.
You will understand our customer needs and how we can meet them.
Responsibilities:
Develop and improve the whole lifecycle of servicesEstablish and improve monitoring capabilities to reduce outage frequency and durationCreate sustainable systems through automation and upliftsDevelop and scale systems sustainably through mechanisms such as automation, and evolve systems by pushing for changes that improve reliability and velocity.Lead designs of major software components, systems, and features to improve the availability, scalability, latency, and efficiency of our servicesAnalyze and support services before they go live via system design consulting, developing software platforms and frameworks, capacity planningConduct post-incident analysis and reviews with an attitude of continuous improvement
Requirements:
Ideally, strong experience in Azure Services and capabilities, but other cloud services (AWS, Google Cloud Platform etc.) will be consideredConfidence and strong experience with KubernetesRecent and fluent Terraform and (Chef platform experience nice to have)Extensive expertise in software development/testing, development operations, and site reliability engineeringExperience of Unix/Linux administration - an appreciation of systems internals (e.g., filesystems, system calls) is a bonusExperience with Continuous Integration and Deployment (CI/CD) and release orchestration and Configuration Management of VMsCloud-agnostic approach, with flexibility to work across various cloud platformsExperience programming in one or more of the following languages: C#,, C++, Java, Python, JavaScript, Go, Perl, or Ruby
Nice to have:
Bachelor's degree in Computer Science, similar technical field of study, or equivalent practical experienceExperience in distributed systems, storage systems, or databasesExperience designing, analyzing, and troubleshooting large-scale distributed systemsSystematic problem-solving approach, combined with excellent communication skills and a sense of ownership and driveExperience in configuring application monitoring with Azure Monitor and Application InsightExperience with Service MeshPrevious experience as a DevOps engineer is preferred
We offer*:
Flexible working format - remote, office-based or flexibleA competitive salary and good compensation packagePersonalized career growthProfessional development tools (mentorship program, tech talks and trainings, centers of excellence, and more)Active tech communities with regular knowledge sharingEducation reimbursementMemorable anniversary presentsCorporate events and team buildingsOther location-specific benefitsnot applicable for freelancers