Site Reliability Engineer

Thrive It Systems Ltd โ€” Poland ยท Posted ~22 hours ago

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Description

Responsibilities: Maintain and support production systems ensuring high availability reliability and scalabilityImplement and support SRE practices including monitoring ing SLOs SLIs and incident responseDeploy configure and manage containerized applications using Docker and KubernetesDevelop and maintain CICD pipelines to improve software delivery and deployment processesAutomate operational tasks using Python Bash Go or similar scripting languagesCollaborate with Development QA and Operations teams to improve platform reliability and performanceParticipate in oncall support rotations and resolve production incidentsConduct troubleshooting root cause analysis and problem resolution activitiesDocument operational procedures system configurations and postincident reviewsSupport continuous improvement initiatives across infrastructure and platform operationsContribute to the management of cloudbased production environmentsSupport GitOps practices and Kubernetes deployment automation To be successful in this role you should have Hands on years ofexperience in SRE DevOps or Production Support rolesStrong knowledge of Docker and KubernetesExperience with cloud platforms such as AWS Azure or GCPExperience implementing monitoring and observability solutionsFamiliarity with SLOs SLIs ing and incident management practicesExperience with CICD tools such as Jenkins or GitLab CIKnowledge of GitOps tools such as ArgoCD FluxCD or TektonStrong scripting skills using Python Bash Go or similar languagesStrong troubleshooting and analytical skillsExperience with Prometheus Grafana ELK Datadog Jaeger or ZipkinKnowledge of Helm and IstioExperience with Infrastructure as Code tools such as Terraform or Ansible is desirableUnderstanding of Microservices and Distributed SystemsAwareness of ITIL Incident Management and Change Management practicesExperience supporting high availability or 24x7 production environments is preferredSkills Mandatory Skills: CI/CD Architecture, Docker, Kubernetes