Summary
✨ AI‑Generated
Take a hands-on application-layer SRE role focused on making core software more reliable, performant, and supportable. You will handle complex incidents, debug production code, deliver fixes, improve legacy logic, implement observability, and build diagnostic automation that makes support teams more efficient.
Highlights
Focus on application-layer reliability with hands-on development, code-level troubleshooting, observability, performance optimization, and automation that reduces operational toil and incident resolution time.
Description
We are seeking an experienced application developer or core application support professional to improve the reliability, performance, and supportability of core software applications.
This role focuses exclusively on application-layer SRE and hands-on development; DevOps and infrastructure-focused SRE backgrounds are not aligned with the position.
### Responsibilities
- Serve as a primary escalation point for complex tier-2/3 application incidents, using code-level debugging and root cause analysis to resolve issues.
- Develop and deploy patch fixes, maintain and refactor legacy application logic, and optimize production code.
- Implement and maintain application performance monitoring and observability using tools such as Datadog, New Relic, or Dynatrace.
- Build scripts and diagnostic tools that improve troubleshooting efficiency, reduce manual toil, and lower mean time to resolution.
- Collaborate with software development teams, contribute operational insights to product planning, review designs for supportability, and participate in code reviews.
### Qualifications
- Strong coding and debugging experience in one or more languages such as Java, Python, Go, C#, Node.js, or C++.
- Advanced application diagnostics, including analysis of logs, stack traces, memory dumps, query plans, and application metrics.
- Strong SQL and/or NoSQL skills, including complex queries and data-layer troubleshooting.
- Experience diagnosing REST APIs, microservices, web services, and message queues such as Kafka or RabbitMQ.
- Familiarity with application-focused SRE practices, including SLIs, SLOs, error budgets, and blameless post-mortems.
- Typically suited to professionals with 10–12 years of relevant experience.