Summary
✨ AI‑Generated
A leading technology company seeks a Senior Software Engineer to optimize large-scale distributed systems. The role involves analyzing performance bottlenecks, redesigning critical paths, and implementing fault-tolerant solutions across microservices and data pipelines. You will work with backend engineers and infrastructure teams to enhance system reliability and efficiency.
Highlights
Work on cutting-edge distributed systems at scale. Design optimizations to improve performance, reliability, and cost efficiency. Collaborate with cross-functional teams to build robust, fault-tolerant solutions.
Description
Job Description
We are seeking a Principal Distributed Systems Optimization Engineer to design, analyze, and optimize large-scale, high-throughput distributed systems.
This role focuses on improving system performance, reliability, and cost efficiency across complex service-oriented architectures operating at significant scale.
You will work closely with backend engineers, data infrastructure teams, and platform reliability groups to identify bottlenecks, redesign critical paths, and implement robust, fault-tolerant solutions.
The ideal candidate has deep expertise in distributed systems theory and hands-on experience tuning production systems under heavy load.
Key Responsibilities
Analyze system performance using metrics, tracing, and profiling tools to identify latency and throughput bottlenecksDesign and implement optimizations across microservices, data pipelines, and storage layersImprove system reliability through fault isolation, graceful degradation, and resilience patternsLead architectural reviews and propose scalable design improvementsCollaborate with cross-functional teams to drive performance best practicesBuild tooling and frameworks for performance benchmarking and capacity planningMentor engineers on distributed systems concepts and performance engineering
Required Qualifications
8+ years of experience in backend or distributed systems engineeringStrong understanding of distributed systems concepts (consensus, partitioning, replication, consistency models)Proficiency in one or more programming languages such as Java, Scala, Go, or PythonExperience with large-scale data systems (e.g., streaming platforms, distributed databases)Deep knowledge of performance tuning, profiling, and observability tools
Preferred Qualifications
Experience with systems like Kafka, Spark, Flink, or similar technologiesFamiliarity with cloud platforms and container orchestration (e.g., Kubernetes)Background in high-QPS, low-latency systemsExperience designing systems with strict SLAs and SLOs
What You’ll Work On
Optimizing services handling millions of requests per secondReducing tail latency and improving system predictabilityEnhancing observability and debugging capabilities in complex environmentsDriving initiatives to reduce infrastructure cost while maintaining performance