Platform Engineer

Trulegal — United Kingdom · Posted ~4 hours ago

Senior Full-time Remote

Skills

Go Python Kubernetes CI/CD Infrastructure as Code DevOps

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An experienced platform engineering role building secure and scalable foundations for software and AI workloads through automation, cloud technologies, and modern engineering practices.

Highlights

Remote-focused platform engineering role with modern cloud technologies, automation, AI infrastructure, and collaborative engineering practices.

Description

Our client, a leading global law firm, is seeking a Platform Engineer to join its technology team in a primarily remote role, (however candidates must be within commutable distance of the London office for occasional onsite visits). This is a highly hands-on, AI/ML-enabled engineering position focused on building the platforms and automation that power modern software and AI workloads across the firm. The successful candidate must be able to code in both Go and Python and will work extensively across pipelines-as-code, infrastructure-as-code, policy-as-code, test automation, and the operationalization of data and AI workflows. It’s an excellent opportunity for an experienced Platform/DevOps Engineer to work with cloud, Kubernetes, CI/CD, observability, and emerging AI infrastructure while helping establish scalable, secure, and reliable engineering practices. This is an opportunity to join an innovative, progressive, and collaborative team. Primary applications and platforms include: CI/CD, Source Control & Test Automation: GitHub Actions, Azure DevOps, GitLab CI, Jenkins; Git, JFrog Artifactory; Playwright, pytest/JUnit Infrastructure & Config as Code: Terraform, Ansible, Bicep/ARM, Helm, Kustomize; GitOps via Argo CD and Flux Cloud & Orchestration: AWS, Azure, Docker, Kubernetes AI Inference & Application Infrastructure: Frontier and open-weight models via Anthropic, Azure OpenAI, and Amazon Bedrock; model gateways and routing, retrieval and hybrid search, document ingestion, tool/function calling, Model Context Protocol (MCP), and agent orchestration AI Evaluation & Quality: Eval harnesses and golden datasets, LLM-as-judge and human-in-the-loop review, regression suites, and red-teaming Observability & Monitoring: Prometheus, Grafana, Datadog, Splunk, Elastic/ELK, OpenTelemetry, including GenAI tracing and token, latency, and cost telemetry Platform Security & Policy-as-Code: HashiCorp Vault, OPA/Conftest, SAST/DAST Developer Portal & Self-Service: Internal developer portal, CLIs/SDKs, and APIs Responsibilities include: Developer Experience & Self-Service Enablement Build and maintain self-service tooling, CLIs, libraries, and templates, as version-controlled, tested code, that make it easy for teams to build, test, and ship software. Contribute features and fixes to the internal developer platform and portal through pull requests and code review. Provide day-to-day support to developers using platform services, triaging and resolving requests and issues. Support governed self-service access to AI platform capabilities, including model access, retrieval tooling, and evaluation workflows, so AI capabilities can move from prototype to production using standard platform patterns. CI/CD, Release Management & Test Automation Implement and maintain pipelines-as-code (e.g., GitHub Actions/Azure DevOps YAML) to established patterns, keeping builds, tests, and deployments reliable and fast. Execute and support release activities, following defined GitOps and change-management processes. Write and maintain automated tests and quality gates within pipelines. Implement evaluation-based quality gates for AI systems, including evals-as-code, regression suites against golden datasets, and human-review thresholds where required. Infrastructure as Code & Cloud Platforms Author and maintain modular, tested infrastructure-as-code (e.g., Terraform, Helm) to provision and configure cloud and on-prem resources. Deploy and operate workloads across cloud platforms (AWS/Azure) and Kubernetes using GitOps (Argo CD/Flux). Follow tagging, cost, and configuration standards when provisioning resources. Help deploy and operate AI workload infrastructure, including model gateways, retrieval services, orchestration components, and supporting cloud or Kubernetes resources. Observability, Monitoring & Site Reliability (SRE) Instrument services and implement monitoring, logging, and alerting as code using standard tooling (Prometheus, Grafana, OpenTelemetry). Participate in the on-call rotation, responding to incidents and helping restore service. Contribute to blameless post-incident reviews and implement follow-up actions in code. Instrument AI services for operational visibility, including tracing across prompts, tools, and agent steps, and monitoring latency, cost, token usage, and quality regressions. Platform Security & Automation Apply platform security controls and remediate vulnerabilities to defined standards, including policy-as-code checks (e.g., OPA/Conftest). Manage secrets, access, and configuration securely using approved tooling (e.g., Vault). Apply security and confidentiality controls for AI workloads, including prompt-injection and data-exfiltration defenses, output filtering, and access controls across prompts, retrieval, and agent tool use. Automate repetitive operational tasks by writing scripts and small services. Software Engineering Standards & Collaboration Follow engineering standards, version-control workflows, code-review, and testing practices. Collaborate with platform and development teams and escalate complex issues appropriately. Create and maintain runbooks, SOPs, and documentation as code alongside the tooling it describes. Qualifications: Proficiency writing production-quality code in Go and Python (plus PowerShell/Bash), using Git, pull requests, code review, and automated testing. Working proficiency building pipelines-as-code (GitHub Actions, Azure DevOps, GitLab CI, Jenkins). Experience with infrastructure as code (Terraform, Ansible, Helm) and GitOps concepts. Hands-on experience with at least one major cloud platform (AWS or Azure) and Kubernetes/Docker. Familiarity with observability tooling (Grafana, Datadog, Splunk, ELK, OpenTelemetry) and basic SRE practices. Exposure to test automation, policy-as-code, and platform security practices. Familiarity with ITIL best practices (incident, change, and problem management) preferred. Experience with Lean or Agile methodologies preferred. Relevant certifications (e.g., AWS/Azure Associate, Certified Kubernetes Administrator) preferred. Experience: Significant experience in platform, DevOps, infrastructure, or software engineering roles. Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field (or equivalent experience). Experience supporting business-critical systems, ideally within professional services, legal, or regulated environments. Job ID: 7604