Senior Python Full Stack Engineer

Xchange Software β€” United States Β· Posted ~3 hours ago

Senior Contract Hybrid

Skills

Python Vue.js Kubernetes AI platforms Cloud systems API development AI/LLM Cloud

πŸ”“ Log in to save this job, tailor your resume & track your apply process β€” 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior software engineer is sought to build resilient AI platforms using modern backend and frontend technologies. The role includes cloud-native architecture, observability, reliability engineering, and platform development.

Highlights

Opportunity to work on advanced AI platform engineering with modern cloud infrastructure, reliability, and developer tooling challenges.

Description

Software Engineer, AI Platform Engineering 12+ Months contract This is a hybrid role in NYC , NY Interview process: 1st round: 1 hour video interview 2nd round: In-person interview in NY office The Agent Platform team would like to see candidates with Vue.js + Python experience Senior Software Engineer, AI Platform Engineering As a Senior Software Engineer, you will focus on the reliability and resilience of mission-critical AI platforms. You will design systems that withstand provider and dependency failures, establish observability and service-level objectives, and keep the platform current as the AI ecosystem evolves. The team’s scope includes Kubernetes-based platform-as-a-service frameworks, model-agnostic AI/LLM gateways, hybrid networking, observability, and self-service developer tooling. We are looking for a collaborative, self-motivated engineer who is comfortable with ambiguity, takes ownership, and enjoys solving complex infrastructure challenges. Legal AI is one of the most exciting and fast-moving areas in technology today. If you are interested in building the foundational platforms that power the next generation of AI products, we'd love to hear from you. What You’ll Do ● Improve platform reliability and resilience by designing for failure, defining and meeting SLOs, leading incident response, and reducing operational toil. ● Design, build, and operate Kubernetes-based PaaS frameworks, AI/LLM gateways, APIs, and self-service tools for AI applications across Bloomberg. ● Develop model-agnostic gateway capabilities for providers such as OpenAI, Anthropic, Gemini, and AWS Bedrock, including routing, fallback, retries, rate limiting, and cost controls. ● Build observability systems covering metrics, logs, traces, dashboards, and alerting to detect and resolve issues before they affect clients. ● Develop networking solutions that connect applications across public-cloud and on-premises environments. ● Provision and manage cloud infrastructure using Terraform and modern software engineering practices. ● Keep platforms secure and current through dependency patching, runtime upgrades, migrations, and provider-integration updates. ● Create frameworks, templates, and workflows that improve developer productivity and reduce operational overhead. ● Evaluate emerging AI technologies and adapt the platform to support new development patterns and use cases. What You’ll Bring ● 6+ years of professional software engineering experience. ● Strong Python skills and experience developing production-grade backend services and APIs; Java experience is a plus. ● Experience designing and operating distributed systems in public-cloud environments, with a strong understanding of failure modes and resilient design patterns. ● Hands-on AWS experience, including services such as EC2, S3, IAM, and container-based workloads. ● Experience with Infrastructure as Code, preferably Terraform. ● Experience with production operations, including metrics, logging, tracing, alerting, SLOs, and incident response. ● Strong knowledge of software architecture, databases, networking, cloud infrastructure, and modern application development. ● A degree in computer science, engineering, or a related field, or equivalent practical experience. Preferred Qualifications ● Experience building or operating API gateways, LLM gateways, or similar proxy layers with routing, fallback, rate limiting, caching, and cost tracking. ● Experience with Open Telemetry, Prometheus, Grafana, Datadog, or similar observability tools. ● Experience with Kubernetes and autoscaling technologies, preferably Amazon EKS and Karpenter. ● Knowledge of AWS networking and security, including VPC, Direct Connect, IAM, and cloud security controls. ● Experience with chaos engineering, load and failure testing, capacity planning, or disaster recovery. ● Experience developing AI-powered applications, agent-based systems, model-inference services, or AI serving platforms. ● Familiarity with AI development tools such as Claude Code, Cursor, or GitHub Copilot. ● Working knowledge of machine learning concepts and the ML development lifecycle. Experience with SageMaker, Bedrock, PyTorch, TensorFlow, or scikit-learn is a plus. ● The ability to learn quickly and independently lead large technical initiatives from concept through production.