Senior Site Reliability Engineer - Bandung, Indonesia
Ninjaone — Indonesia · Posted ~3 hours ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
Description
About the Role
We are looking for a Senior Site Reliability Engineer to join our existing SRE team and build tools and automation that improve production reliability and day-to-day operations.
As a senior individual contributor, you will turn initial operational needs into clear requirements, practical system designs, working proofs of concept, and maintainable production solutions.
You will work closely with SRE, Engineering, and Production Support teams, quickly learn our products and support model, and contribute to daily SRE activities.
Building new tools and automation is the primary focus of this role, alongside participation in the weekly support rotation.
Location – Hybrid | Bandung, Indonesia (at least 4 days per week in the office)
What You’ll Be Doing
Tooling and Automation
Design, build, and maintain internal tools, services, and scripts using Python, Go, and shell scripting to reduce repetitive work and improve operational efficiency.
Gather and clarify raw requirements from SRE and Engineering teams; define the problem, scope, acceptance criteria, and expected operational outcomes.
Create system designs covering components, integrations, data flows, security, scalability, and failure handling; explain design choices and trade-offs.
Build and demonstrate proofs of concept, incorporate feedback, and own delivery through testing, rollout, documentation, and ongoing maintenance.
Develop and improve CI/CD pipelines for reliable builds, automated testing, deployment, and rollback; automate infrastructure and configuration workflows.
Product Learning and SRE Operations
Quickly learn product architecture, service dependencies, infrastructure, deployment processes, and the production support model.
Work alongside the existing SRE team on daily server and service operations, including monitoring, troubleshooting, maintenance, and controlled changes.
Participate in the weekly rota for SRE operations and on-call support, including out-of-hours incident response when scheduled; follow escalation and handover procedures.
Investigate production incidents with Engineering and Production Support, contribute to root cause analysis, and turn recurring issues into lasting automation improvements.
Collaboration and Engineering Quality
Work closely with SRE, Engineering, and Production Support to prioritize operational needs and deliver tools that fit existing workflows.
Write tested, maintainable code, participate in code and design reviews, and apply secure development practices.
Maintain system designs, runbooks, operating procedures, and user documentation; demonstrate solutions and support their adoption across teams.
Improve monitoring, alerting, and service performance, and evaluate delivered automation by its reliability and reduction in manual effort.
About You
At least 5 years of relevant experience in Site Reliability Engineering, DevOps, platform engineering, systems engineering, or production software engineering with operational responsibilities.
Strong hands-on programming ability in Python or Go, practical shell scripting skills such as Bash, and willingness to work across the team’s languages and tooling.
Demonstrated experience building and maintaining operational tools or automation beyond one-off scripts, including APIs, error handling, testing, and version control.
Hands-on experience designing and maintaining CI/CD automation using tools such as Jenkins, Bitbucket Pipelines, or equivalent platforms.
Ability to independently clarify incomplete requirements, design a suitable solution, and build a proof of concept that demonstrates its value and feasibility.
Experience with cloud infrastructure such as AWS or GCP, Linux server administration, networking fundamentals, and troubleshooting production services.
Experience with infrastructure as code or configuration management, such as Terraform or Ansible, and monitoring and logging platforms such as Datadog, Grafana, or CloudWatch.
Strong ownership, problem-solving, and written and verbal communication skills, including the ability to explain technical decisions to different teams.
Ability to learn unfamiliar products and support processes quickly, collaborate within an established team, and participate reliably in the weekly operational rota.
A degree in Computer Science, Information Technology, or a related field, or equivalent practical experience.
Preferred Experience
Docker and Kubernetes, distributed systems, and reliability practices such as service level objectives, capacity planning, and post-incident reviews.
Cloud backup or SaaS platforms, databases, storage, or messaging systems.
Building self-service operational tools and integrating cloud APIs, monitoring systems, or ticketing workflows.
About Us
NinjaOne unifies IT to simplify work for nearly 40,000 customers in 140+ countries.
The NinjaOne Unified IT Operations Platform delivers endpoint management, autonomous patching, backup, and remote access in a single console to improve efficiency, increase resilience, and reduce spend.
By automating IT and managing all endpoints, organizations give employees a great technology experience at work.
NinjaOne is obsessed with customer success and has retained a 98% customer satisfaction score for more than 5 years.
What You’ll Love
We are a collaborative, kind, and curious community We prioritise your work/life balance offering a hybrid work environment and free in-office lunches throughout the week We reward your work with opportunity for growth and advancement Grow personally and together with one of the fastest growing companies globally Develop your skills through our renowned training platform Receive competitive compensation Collaborate with an amazing international workforce
Additional Information
Applicants must hold valid working rights in Indonesia at the time of application.
We are unable to provide visa sponsorship for this role.
We are an equal opportunity employer in accordance with Indonesia's Manpower Law.
All applicants will receive fair consideration for employment regardless of nationality, gender, age, disability, or background.
We are committed to fostering an inclusive and respectful workplace for all employees.
We have 143,338 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 143,338 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume