Senior Software Engineer - Backend & Distributed Systems Engineer

Permutableai — United Kingdom · Posted ~9 hours ago

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Description

At Permutable, we build real-time AI systems that transform large volumes of global news, market and proprietary data into market intelligence for financial institutions. We are moving towards a more real-time, event-driven and agent-driven architecture. We are looking for a Senior Software Engineer to help design and build the backend services, distributed systems and cloud infrastructure behind it. This is a hands-on software engineering role with end-to-end production ownership. You will work directly with our founder and engineering team and have significant influence over how our next generation of technology is designed and built. What You’ll Do Build Production Backend Systems: Design, build and maintain reliable Python services, APIs and data-intensive applications operating at scale. Develop Event-Driven Architecture: Help move our architecture from scheduled pipelines towards real-time services using messaging, queues, asynchronous processing and distributed workflows. Build Systems for AI Agents: Develop the production software required to run increasingly agent-driven AI workflows, including state management, task execution, model routing, retries and observability. Scale our Cloud Platform: Build and evolve highly available production systems across AWS, Kubernetes/EKS, ECS, Docker, RDS/PostgreSQL and S3. Productionise AI Workloads: Work closely with our AI engineers to deploy and operate NLP, LLM and quantitative-model workloads, including workloads using our in-house GPU infrastructure. Own Software Through to Production: Take responsibility for testing, deployment, performance, monitoring and reliability rather than handing code over to a separate operations team. Automate Delivery and Infrastructure: Manage infrastructure using Pulumi or Terraform and improve automated testing and deployment through GitHub Actions. Improve Reliability and Observability: Build effective monitoring, logging, tracing and alerting across our data, service, model and agent infrastructure. What We’re Looking For 7+ years of professional experience in software engineering, backend engineering, distributed systems, platform engineering or a closely related role. Strong Python experience, including building and maintaining production software, services or APIs. Deep AWS experience, ideally including EKS or ECS, RDS/PostgreSQL, S3 and Lambda. Distributed Systems Experience: A strong understanding of event-driven architectures, messaging, asynchronous processing and distributed workflows. Strong Engineering Fundamentals: A good understanding of software architecture, Linux, networking, databases and cloud infrastructure. Production Ownership: Experience taking systems from design and implementation through to deployment, monitoring and ongoing operation. Infrastructure-as-Code: Experience with Pulumi, Terraform or an equivalent framework. CI/CD and Automated Testing: Experience with GitHub Actions or similar tooling, automated testing and reliable deployment processes. Startup Mindset: Proactive, resourceful and comfortable taking responsibility in a small, fast-moving engineering team. Experience with Apache Airflow, Kubernetes, real-time data platforms, AI/ML infrastructure, LLM agents, GPU infrastructure or financial-market systems would be useful, but is not essential. Why Permutable? Real Ownership: Take responsibility for important production systems and influence how they develop. Technically Demanding Work: Build across backend software, distributed computing, cloud infrastructure and production AI. Direct Influence: Work closely with our founder and have a meaningful voice in architecture and technology decisions. Real Production Impact: Your work will quickly reach production and support leading financial institutions. Build the Next Generation: Help shape our move towards a real-time, event-driven and agent-driven architecture. Hybrid Flexibility: Spend at least three days a week in our Vauxhall hub.