Summary
Start your SRE career by helping build and maintain scalable, reliable systems. You will monitor performance and availability, improve alerting and incident response, and assist with CI/CD automation while learning from experienced engineers. The role provides practical exposure to modern cloud and DevOps practices.
Highlights
Designed for early-career engineers, with hands-on exposure to cloud infrastructure, monitoring, automation, CI/CD, and incident response. The role provides close collaboration with senior engineers and opportunities to learn modern DevOps practices.
Description
oin Hiredbuddy as a Junior Site Reliability Engineer to help build and maintain scalable, reliable systems.
Work with a dynamic team on cloud infrastructure, monitoring, and automation.
Perfect for early-career engineers eager to learn modern DevOps practices.
Full Description
At Hiredbuddy, we are dedicated to delivering high-quality digital experiences, and our infrastructure is the backbone of our success.
We are looking for a Junior Site Reliability Engineer to join our growing team.
In this role, you will learn and apply the principles of site reliability engineering (SRE) to ensure our systems are reliable, scalable, and efficient.
You will work closely with senior engineers and developers to improve our monitoring, alerting, and incident response processes.
Your daily responsibilities will include:
Monitoring system performance and availability using tools like Prometheus, Grafana, and Datadog.Assisting in the development and maintenance of CI/CD pipelines using Jenkins, GitLab CI, or similar.Writing scripts to automate routine tasks using Python, Bash, or Go.Collaborating with development teams to ensure smooth deployments and releases.Participating in incident response and post-incident reviews to implement preventive measures.Helping manage cloud infrastructure on AWS or GCP, including EC2, S3, Load Balancers, and Kubernetes clusters.Performing regular health checks and capacity planning.Documenting processes and procedures to improve team knowledge sharing.As an entry-level role, we expect you to have a foundational understanding of Linux systems, networking, and basic scripting.
You should be eager to learn and grow in a fast-paced environment.
We value curiosity and a proactive approach to problem-solving.
You will receive mentorship and have access to learning resources to help you develop your skills.
We are looking for someone who is comfortable working remotely and can communicate effectively with the team.
You will use tools like Slack, Jira, and Confluence for collaboration.
We believe in continuous improvement and encourage everyone to bring new ideas to the table.
If you are passionate about technology and want to start your career in site reliability engineering, we would love to hear from you.
Join us in ensuring our platform runs seamlessly for our customers.
Requirements
0-2 years of experience in a related field (internships, academic projects, or personal projects)Basic understanding of Linux/Unix administrationFamiliarity with scripting languages like Python, Bash, or PowerShellExposure to cloud platforms (AWS, GCP, or Azure) is a plusStrong problem-solving skills and attention to detailExcellent communication and teamwork abilitiesWillingness to learn and growBenefits
Competitive salary and performance bonusesHealth, dental, and vision insurance401(k) retirement plan with company matchFlexible remote work environmentGenerous paid time off and holidaysProfessional development stipendHome office setup allowance