Senior Site Reliability Engineer
Accelbyte โ Indonesia ยท Posted ~2 hours ago
๐ Log in to save this job, tailor your resume & track your apply process โ 7 days free, no card needed.
Log in to add to target listDescription
POSITION SUMMARY :
AccelByte is building a 24x7 operations team for AAA multiplayer video games.
In this position, we need a driven Site Reliability Engineer who can actively participate in the day-to-day combat by maintaining high reliability of our service and drive prioritization in fixing what may be broken today, as well as able to envision, design, and implement processes and technologies to improve the ability to identify, isolate, correlate, and mitigate service impacting problems in the system.
The Site Reliability Engineer must also know some coding to automate routine tasks in service metrics gathering, correlating, organizing, and presenting, in addition to detail and in-depth root cause analysis
ESSENTIAL FUNCTIONS/RESPONSIBILITIES:
The Senior Site Reliability Engineer (SRE) is accountable for the following functions and responsibilities:
Design, build and maintain high-performance backend services, tools, and control planes to automate operations and improve system reliability.Architect, implement and maintain a highly scalable deployment framework and tooling that improves our products' stability, reliability, and availability.Build and run service deployment using K8s and other CNCF projectsProvide a secure, high-scalable, and cost-effective cloud platformConstruct and build effective systems to monitor the health of our system/applications, and to handle outagesSolve problems occurring in all our environments and create solutions to prevent them from happening againProduce automation and innovative tools to assist the product development teams and to deliver operational excellenceCreate and maintain infrastructure-related documentation and SRE runbooks Collaborate with other stakeholders to provide cost-effective, operational excellence, and performance-efficient infrastructure solutions to improve our products.Identify technology, process gaps, and opportunities for improvementLiaise, communicate, and work directly with our clientsPerform any other design-related duties as required Envision, design, and implement AIOps solutions to enhance operational efficiency and predictive maintenance.Write clean, maintainable, and well-tested code (primarily in Go/Python) for core platform components.
QUALIFICATIONS/EXPERIENCE REQUIRED
5+ years Cloud Engineering or DevOps experience with AWS, 2+ years Kubernetes, Certification in AWS preferredDegree in Computer Science or equivalent experienceDeep knowledge of cloud service providers and best practices around implementation and configuration, preferably managing AWS and KubernetesFamiliarity with infrastructure management and operations lifecycle concepts and ecosystem, deep understanding of IaC and GitOpsProven track record of building infrastructure as code (Terraform is a must), configuration management, and package manager (eg: Helm Chart)Experience in delivering products against a plan in a fast-paced, multi-disciplined, and often ambiguous environmentExperience working independently to design, plan, and execute technical projectsDemonstrated deep knowledge of technical program management and engineering best practicesInnovative thinking balanced with a strong customer and quality and cost efficiency focusComfort and experience with cross-organizational communication; excellent written and verbal communication skillsWorking experience with some of the following technologies and tools: Docker, Kubernetes, git, Redis, MongoDB, PostgreSQL, ElasticSearch, GitLab CI, Nexus, SonarQube, Terraform, Helm, Prometheus, ELK/EFK, Grafana, CloudWatchSolid security best practices Strong proficiency in Go, including the ability to conduct high-quality code reviews.
Experience with Python and Bash is also required.Keen problem-solving skills with the ability to work under pressure (during a production event)Flexibility in working with people with different timezonesExperience with AIOps and building/harnessing AI tools to automate and optimize operational tasks.
QUALIFICATIONS/EXPERIENCE PREFERRED
Previous experience working in the game industryWorking experience with one or more of the following: Emissary, Linkerd, Istio, Nomad, Kafka, Flux, ArgoCD, GitOps, DevSecOpsFamiliar with web services patterns/architectures, e.g.
REST, SOAP, etc.Experience working with auto-scaling workloads both in containers and VMsExperience with other cloud technologies and infrastructure: GCP, AzureExperience with Confluence, Jira, and BitBucketIT standards, methodologies, Cryptographic key management regulations, and audit experience would be asset(s).
We have 143,338 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume โ in under a minute we'll analyze all 143,338 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume