Description
We are looking for a Senior DevOps / Platform Engineer to develop and operate GoMining's cloud platform.
The primary focus of this role is GCP and Kubernetes, infrastructure and application delivery automation, and the reliability and security of production services.
You will maintain and improve our existing infrastructure, contribute to the development of the next generation of our platform, and help product teams ship changes safely and efficiently.
Responsibilities
Develop and operate infrastructure in GCP, including GKE, virtual machines, networking, load balancers, storage, and managed servicesManage Infrastructure as Code with Terraform: develop modules, manage state and dependencies, ensure reproducible changes, and identify configuration driftOperate Kubernetes: upgrade clusters and platform components, configure autoscaling, workload placement, network policies, and resource allocationDevelop CI/CD and GitOps using GitLab CI and Flux / Argo CD, automating builds, checks, deployments, controlled releases, and rollbacksBuild observability across metrics, logs, distributed tracing, and actionable alerting.
Define and improve SLIs/SLOstogether with engineering teamsTroubleshoot production incidents and performance issues, conduct root cause analysis, and eliminate recurring failure patternsImplement least-privilege access, environment and service isolation, secure secrets management and rotation, and infrastructure audit controlsMaintain backup processes and regularly validate recovery of critical data and infrastructure according to agreed RPO/RTOHelp engineering teams prepare services for production, including health checks, resource configuration, telemetry, and deployment proceduresOptimize cloud infrastructure usage and costs while maintaining reliability requirementsMaintain infrastructure documentation, operational runbooks, and recovery procedures
Requirements
Strong hands-on experience independently operating production infrastructure and Kubernetes, including upgrades, troubleshooting, and disaster recoveryStrong experience with GCP, including GKE, Compute Engine, VPC, IAM, and Cloud Storage; good understanding of Cloud SQL, Cloud Logging, and Cloud MonitoringHands-on experience with Terraform, including modules, remote state, locking, importing existing resources, and safely applying infrastructure changes through CIDeep knowledge of Kubernetes and containerization: Deployments, StatefulSets, Services, Ingress, storage, requests/limits, probes, autoscaling, RBAC, and NetworkPolicyExperience with Helm and KustomizeStrong Linux administration skills, preferably Debian/Ubuntu: processes, systemd, filesystems, disks, and resource troubleshootingStrong networking fundamentals: TCP/IP, DNS, HTTP(S), TLS, routing, NAT, load balancing, and firewallsAbility to troubleshoot connectivity issues across applications, Kubernetes clusters, networks, and cloud servicesExperience building CI/CD pipelines, preferably with GitLab CI, and strong understanding of Git and GitOps principlesUnderstanding of runner permissions, secrets protection, and production deployment controlsAutomation skills with Bash and Python, including maintainable infrastructure scripts and API integrationsExperience with monitoring, logging, and alerting using tools such as Prometheus, Grafana, or Cloud MonitoringPractical experience operating PostgreSQL and understanding Redis and RabbitMQ, including availability, connections, replication or clustering, backup, and recoveryStrong understanding of cloud security principles: least privilege, service accounts, Workload Identity, secrets management, network segmentation, and access auditingAbility to independently deliver infrastructure changes to a validated production result, explain technical trade-offs, and collaborate effectively with Engineering and Security teams
Nice to Have
Experience with Flux CD, Argo CD, Ansible, and HashiCorp VaultExperience with OpenTelemetry and distributed tracingExperience migrating infrastructure between clouds or Kubernetes clusters, including database migrations with limited downtimeExperience designing disaster recovery strategies and conducting recovery exercisesCloud cost optimization and capacity planning experienceExperience with DigitalOcean, Hetzner, or AWSAbility to read and troubleshoot JavaScript / TypeScript applicationsExperience operating financial or payment services with strict requirements around access control, auditing, reliability, and data integrity
Benefits
Professional growth: support for courses, conferences, and English learning (up to 100% coverage).
Work-life fit: remote or hybrid format with flexible hours across international teams.
Paid leave: up to 20 vacation days + 8 company holidays + 5 personal days per year Recognition programs: structured performance reviews and team awards.
Team culture: retreats in international locations (for example, company apartments in Cyprus)