Site Reliability Engineer

Paybyphone Technologies Inc — Canada · Posted ~7 hours ago

Mid Full-time

Skills

Site Reliability Engineering Highly available systems Cloud infrastructure Production operations Monitoring Automation Incident management

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A Site Reliability Engineering opportunity focused on building and operating highly available systems used at significant scale. You will help improve reliability, automation, and production operations while working within a multidisciplinary platform engineering environment that values collaboration, continuous learning, and strong user experiences.

Highlights

Join a growing, collaborative engineering organization focused on highly available technology serving millions of users. The role offers opportunities to build and operate reliable systems, contribute to platform engineering, and grow within a learning-oriented environment.

Description

About PayByPhone At PayByPhone, our strength is our people. Behind our product is a talented, creative, and driven multi-disciplinary team united by a shared ambition: to make everyday mobility simpler. We believe innovation should be collaborative, learning should be constant, and work should be enjoyable. As we grow, we’re looking for people who want to grow with us. Together, we’re on an ambitious mission to create intuitive technology solutions that deliver world-class user experiences. We are a fast-growing, forward-thinking company and already help more than 60 million users across North America and Europe. Our technology helps millions of consumers pay quickly, easily, and securely — without waiting in line, carrying change, or worrying about costly fines. About The Role Location: Vancouver, BC Employment type: Full-Time, Permanent Reports to: Site Reliability Lead, Platform Engineering We're looking for Site Reliability Engineers to help build and operate highly available, secure, and scalable SaaS systems at PayByPhone. This role is open across the experience spectrum — whether you're a few years into an SRE or ops career or a seasoned specialist, we'll shape scope, mentorship, and support around where you're starting from. This role calls for a motivated, quality- and results-oriented person who enjoys collaborating with cross-functional teams of skilled developers. The focus of the role is to: Help ensure the PayByPhone platform meets its availability and stability requirementsDrive continuous improvements in our software quality assurance processes, practices, and culture Working under the Site Reliability Lead (SRL), this role helps drive reliability by ensuring deliverables are implemented on time and on budget, and are operating at the expected levels of 24/7/365 at four nines of availability (target 99.99%) and within specific customer SLAs. This role also supports the SRL's agenda of incident prevention, improved incident response, and timely remediation Key Responsibilities Help ensure platform reliability and stability deliverables are implemented on time and on budget, and are operating at 24/7/365 at four nines of availability, within specific customer SLAsStandardize our quality assurance plans and templates for cross-team projectsAssist in quality assurance tooling selection and operational proceduresHelp ensure continuity across all critical business transactionsGenerate and report on reliability metrics to various stakeholdersWork collaboratively as a member of the team to define, refine, and execute the Platform Reliability / DevOps and SRE roadmapSupport PayByPhone in meeting security and compliance requirements by collaborating with the Security & Compliance team and implementing security and compliance tooling, processes, and policies as part of the CI/CD processWork collaboratively with release & support management leaders on compliant and secure processes for triage, investigation, resolution, and release of softwareAssist the SRL in the quality assurance process within the SDLC, including automation language selection and usageHelp build a strong sense of ownership and accountability across teams and individual contributors, reflected at the code level and in implementation and operationsContribute to high-quality execution and technical and operational excellenceGather and analyze metrics from operating systems and applications to support performance tuning and fault-findingParticipate in system design consulting, assisting the Platform team where needs overlap with reliability and availabilityAssist in scoping, developing, and testing the ongoing disaster recovery plan under the SRL's leadershipHelp own observability, monitoring, and alerting tools under the SRL's leadership — proactively looking for ways to improve, standardize, document, and trainProvide operational support: documentation and debugging of production issues, including —Being available to join Sev-1 pagesAssisting the SRL with responsibilities such as post-mortems and on-call bootcampsAssist with CI/CD toolset setup and support (GitLab runners) from a reliability standpointSupport cost reduction, budget setting, and monitoring across the platform and the tools used within itSupport software reliability practices (logging standards, use of reliable libraries, SLA/SLO goals)Participate in on-call responsibilities when neededMaintain a personal data plan to support your on-call responsibilities Key Requirements 3+ years of experience in software development, delivery, or site reliability / operations for large, complex software systems spanning both legacy and modern stacks. Equivalent experience gained through non-traditional paths is welcome.Bachelor's or higher degree in Computer Science, Computer Engineering, or a related technical field is preferred; equivalent hands-on experience will also be considered.Experience working with high-performing SRE, Ops, or Dev teamsSolid grounding in quality assurance discipline, software quality management, and related frameworks and toolsWorking experience with Amazon Web Services (AWS) solution architectures and technologiesExperience building verification and validation practices into end-to-end delivery pipelines, from business development through the delivery phase, to accelerate product launch to marketFamiliarity with testing techniques such as unit, functional requirement, performance, GUI, regression, integration, system load, vulnerability assessment, security testing, and test automationUnderstanding of SaaS multi-tenant and distributed / micro-service architecturesUnderstanding of DevOps principles, processes, and tools (e.g. IaC, CI/CD, and orchestration)Understanding of cloud computing architecture, services, and platformsUnderstanding of web and/or mobile development technologies and programming/scripting languagesAbility to program (structured and OOP) using one or more high-level languages, such as Python or JavaScript/TypeScriptExperience with distributed storage technologies such as NFS, HDFS, and Amazon S3, as well as dynamic resource management frameworksExperience with Infrastructure as Code (IaC)A proactive approach to identifying problems, performance bottlenecks, and areas for improvementComfortable working with and supporting cross-functional teamsStrong written communication, including technical documentation and training materials What We Offer Compensation: The expected salary range for this role is $110,000 – $120,000 CAD. Final compensation will be based on factors such as experience, skills, qualifications, and internal equity. Retirement Savings Program: Access to our retirement savings program (RRSP for Canada / 401(k) for U.S.-based employees). Vacation: All permanent full-time employees start with 4 weeks of vacation per year. Work from Anywhere: Up to 15 days of work from anywhere subject to management and IT Security approval. Personal Days: We provide 5 personal days annually, in addition to paid sick days, to support flexibility and work-life balance. Comprehensive Medical & Dental Coverage Employee Assistance Program (EAP): Access to confidential support services and resources for you and your family. Career Growth & Learning Support: Opportunities for professional development, continuous learning, and career progression. Working at PayByPhone We Operate In a World That’s Constantly Evolving — And Change Is Something We Embrace. Our Values Guide How We Show Up For One Another And For Our Customers Every Day. In Short, We Make things happenStay curiousWork togetherHave funSee through our customers’ eyes These principles shape how we collaborate, innovate, and deliver on our commitments. We’re also committed to fostering a diverse and representative workforce and an inclusive environment where everyone is treated with respect and fairness. We do not tolerate discrimination or harassment in our workplace or throughout our hiring process. Our hiring decisions are grounded in business needs, role requirements, and individual qualifications — ensuring we reflect the talent and communities we serve. PayByPhone is committed to providing accommodation throughout the recruitment process. If you require accommodation, please reach out to us at askhr@paybyphone.com. Want to see our values in action? Visit our Instagram and LinkedIn. Curious about the story behind our values? Head over to our About Us page to learn more.