Summary
✨ AI‑Generated
Manage cloud environments, container platforms, automated delivery pipelines, and operational reliability for large-scale systems.
Highlights
Own cloud infrastructure, improve reliability, and work on automation and scalable platform engineering.
Description
About The Role
We are looking for a dedicated and experienced Cloud Site Reliability Engineer (SRE) to join our technical team in Gothenburg.
In this role, you will take ownership of our public cloud infrastructure, containerized platforms, and automated CI/CD pipelines, ensuring high availability, scalable architecture, and operational efficiency across our cloud environments.
Key Responsibilities
Cloud Infrastructure & Governance: Own end-to-end operation, maintenance support, and architecture governance of public cloud infrastructures, covering core resources including containers, cloud virtual machines, storage, and networks.
Oversee the operational stability and architectural standardization of databases and middleware systems.
CI/CD & Automation: Design, implement, and maintain automated CI/CD pipelines to support continuous integration, continuous delivery, and standardized release workflows for business applications.
Operations & Incident Response: Undertake daily on-call rotation responsibilities.
Proactively troubleshoot and resolve functional defects, resource bottlenecks, and performance anomalies of cloud infrastructure and applications to maintain service stability.
Container & Cloud-Native Management: Manage daily operations and production releases of containerized and cloud-native applications, ensuring reliable and smooth online performance.
Team Collaboration: Collaborate closely with technical teams to formulate and implement unified containerized management specifications, promoting standardized governance, architecture optimization, and application migrations.
Resource Optimization: Continuously analyze cloud and application resource utilization, driving resource scheduling optimization and efficiency improvements to maximize cost-effectiveness.
Compliance & Best Practices: Maintain technical documentation and ensure cloud infrastructure operations strictly comply with data compliance specifications and regional privacy protection regulations.
Requirements
Operating Systems & Networking: Solid mastery of Linux operating system principles and core network protocols (including TCP/IP and HTTP), with proficient hands-on operational capabilities.
Kubernetes Ecosystem: In-depth understanding of the Kubernetes ecosystem and core underlying mechanisms; proficient in operating and maintaining Kubernetes Operators in production environments.
Observability: Proficient in mainstream observability tools (including Prometheus and Grafana), with the ability to build and maintain end-to-end monitoring systems for cloud-native environments.
CI/CD Toolchains: Familiar with mainstream CI/CD toolchains (represented by ArgoCD), with solid capabilities to build, configure, and maintain automated delivery pipelines.
Gateways & Middleware: Familiar with the architecture and operational principles of cloud-native gateway systems (Nginx, APISIX, Envoy) and message queue middleware (RocketMQ, RabbitMQ, Kafka).
Automation & Scripting: Proficient in at least one mainstream scripting language (Python or Shell) for daily operational automation and tool development.
Compliance: Clear understanding of data compliance specifications and privacy protection regulations such as GDPR.
Languages: Fluent in English (both written and verbal).