Platform Engineer

Cloudlinux — Bulgaria · Posted ~4 hours ago

Mid Full-time Remote

Skills

platform engineering Linux GitLab CI runners observability infrastructure automation provisioning configuration management self-service platforms cloud infrastructure CI/CD Observability Infrastructure automation

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A remote platform engineering role within a small infrastructure-focused team. You will own a defined set of platform services, keeping them reliable, repeatable, and within agreed service levels while reducing recurring operational work through automation and self-service. Responsibilities span observability, source-control infrastructure, CI runners, provisioning, configuration, network policies, and infrastructure workflows.

Highlights

Remote platform engineering role with strong ownership over infrastructure services, observability, CI runners, provisioning automation, and operational reliability. The position emphasizes autonomy, repeatable operations, and reducing manual work through automation.

Description

CloudLinux builds Linux infrastructure and security products. You will join our Automation & Management Services cell, working closely with Dmitrii Petrov to solve problems across teams and services: cloud-cost data, infrastructure inventory, network policies and capacity workflows. Check out our website for more information https://cloudlinux.com/ Inside the Infrastructure Department, the Platform cell is a small team. We run the observability platform, the company's GitLab and the CI runners behind it, a few smaller engineering services, and the automation the department relies on for provisioning and configuration. We are looking for a Platform Engineer for the PaaS team to take ownership of an agreed set of platform services: keep them reliable and within agreed service levels, make changes and recovery repeatable, and reduce recurring operational work through automation and self-service. You get real freedom in how you implement things, and you own the result: you pick the approach, defend it in review, and answer for how the service behaves afterwards. Most of the job is what site reliability engineering is about: services that run well day after day, and requests from product teams handled properly. Some of it is building. The observability platform and the runner cluster were both built from scratch within the last year, and there will be more of that. Expect the work to split between platform engineering, incident response, and helping engineers in other teams use what we run. What You'll Do Run the observability platform. Keep it healthy, onboard teams, watch cost and capacity, and maintain the alerting that runs on top of itRun GitLab and the CI runner fleet. Upgrades, capacity, access, backups and restore drillsKeep the rest of our services healthy, with the monitoring and runbooks a production service needsDeploy new services when they are requested. Research the options, pick a design, and stand the service up from scratch according to good practice: as code, monitored, backed up, documentedWork with developers' requests. Access, onboarding, pipeline problems, new exporters and dashboards. Answer them, and turn the recurring ones into self-serviceRun incidents. Diagnose and mitigate impact, restore service safely, then complete the root-cause analysis and post-mortem. Deliver the prevention or detection improvements the incident calls forShip everything as code, reviewed in merge requests. Plan and check before every changeWrite for engineers outside the team. Runbooks, onboarding guides, maintenance notices and status updates that people can act onWork with AI agents. Delegate collection and drafting to them, review their output as you would a colleague's merge request, and record what you learn where the team can find it Requirements Must have Senior-level experience in infrastructure, platform or site reliability engineering, including at least one production service you were responsible for keeping up. We will ask you to walk us through it in detail: what broke, how you found out, and what you changed so it would not happen againLinux systems administration and debugging on bare metal and virtual machines. Much of our infrastructure is not KubernetesKubernetes in production delivered through GitOps, including cluster upgrades you performed yourselfInfrastructure as code as your delivery form: Ansible and Terraform or OpenTofu, changes reviewed in merge requestsGitLab administration and GitLab CI in production, self-hosted or SaaS. Deep experience with another CI system is acceptable if you can show the same depthWorking knowledge of the Prometheus and Grafana ecosystem: you have run it for a team, written alert rules and dashboards, and can read PromQL. Depth here is welcome, and learnableWritten technical explanation for engineers outside your team: runbooks, notices, answers to requestsStrong communication and interpersonal skills. This role deals with people at least as much as with servers: most work starts as a conversation with a product team, and you need to understand what they actually need, agree scope, priority and timing with them, push back politely when a request should not be done as asked, and keep everyone informed while the work is in progress. We are looking for someone other teams enjoy working withAdvanced use of AI engineering assistants such as Claude and Codex: providing context, breaking down tasks, designing agent loops, and delegating plans for unattended, end-to-end execution within defined scope and permissions, with clear stop conditions. You can explain, debug and test the resulting automation, and verify generated commands, scripts and conclusions before they touch productionEnglish - upper-intermediate or higher - to ensure clear communication of progress within the teams Nice to have Alerting design: SLOs, burn-rate alerts, thresholds sized from dataMicroVM isolation for CI: Kata Containers, Firecracker or gVisorS3-compatible object storage operations: Ceph RGW or similarAWS with real cost workSelf-hosted Sentry, or another Kafka, ClickHouse and Redis-backed application you have kept alive under loadPython or Go for exporters and small internal services You do not need to have run every system on this list. Solid fundamentals and the judgment to pick up an unfamiliar service, make it observable and hand back a runbook matter more than matching every line. What This Role Is Not Focused On Not a ticket-queue operator. Recurring requests get turned into self-service, not processed one by one foreverNot a pure cloud or Kubernetes role. Bare metal and virtual machines are a large part of our infrastructureNot the DBA, the network engineer or the security engineer. Those teams run their own systems; we provide the platform they monitor them with Benefits What's in it for you? A focus on professional developmentInteresting and challenging projectsFully remote work with flexible working hours, which allows you to schedule your day and work from any location worldwidePaid 24 days of vacation per year, 10 days of national holidays, and unlimited sick leavesCompensation for private medical insuranceCo-working and gym/sports reimbursementBudget for educationThe opportunity to receive a reward for the most innovative idea that the company can patent By applying for this position, you consent to the processing of your personal data as described in our Privacy Policy (https://cloudlinux.com/candidate-privacy-notice), which provides detailed information on how we maintain and handle your data.