Summary
✨ AI‑Generated
A growing engineering organization is seeking a Senior Site Reliability Engineer to support the production reliability of a critical cloud-native service. You will work with Kubernetes and modern DevOps practices to improve stability, security, scalability, and operational excellence in a high-performance environment.
Highlights
Senior SRE opportunity focused on production reliability, secure and highly scalable cloud-native systems, and modern DevOps practices. The engagement offers work with experienced engineering teams in a multicultural and growing environment.
Description
Company Description
We are Software Mind, an awesome team of engineers who are ready to ramp up any top-notch company’s projects! Our aim? To always be one step ahead.
Become part of a multicultural company in constant growth with an excellent work environment certified by Great Place To Work!
About The Client
Our client is a leading enterprise software company building highly scalable cloud-native platforms used by organizations around the world.
Their engineering teams focus on delivering reliable, secure, and high-performing services while embracing modern DevOps, Kubernetes, and cloud technologies.
You will join a team responsible for ensuring the stability, reliability, and operational excellence of a critical UI service running in production.
Contract Duration: Initial contract through the end of 2026, extending the engagement to a total 12-month term based on performance.
Job Description
About the Role
This is a Senior SRE role supporting production reliability for a Kubernetes-based UI service / AI experience framework stack.
This is not general infrastructure, and it is not a front-end developer role.
The strongest candidates will have production SRE experience across Kubernetes operations, observability, Node.js runtime troubleshooting, JVM / Java service troubleshooting, Splunk, and incident ownership.
What You’ll Do
Support the deployment, operation, and reliability of production services running on Kubernetes.Monitor service health and investigate production incidents across distributed applications.Participate in on-call support, incident response, root cause analysis, postmortems, and reliability improvements.Troubleshoot application runtime, networking, and service-to-service issues in collaboration with engineering teams.Support CI/CD, GitOps-based deployments, observability, and production monitoring.Work within a client-directed backlog and established priorities.
Qualifications
Required Qualifications
5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Production Engineering, or a closely related role, including strong recent hands-on experience supporting Kubernetes-based production services.3+ years of hands-on production Kubernetes experience strongly preferred.
Kubernetes production operations, including deployment, scaling, rollout / rollback, resource tuning, and service-to-service troubleshootingStrong production incident response experience, including on-call, runbooks, postmortems, and paging hygieneSplunk experience for log aggregation, search, and production troubleshootingPrometheus and Grafana experience, specifically building alert rules and dashboards, not only using existing dashboardsCI/CD and infrastructure-as-code for containerized deployments, including Helm and GitOps tools such as ArgoCD or FluxStrong Linux and networking fundamentals, including DNS, load balancing, TCP / HTTP, HTTP/2, and Kubernetes networkingProduction troubleshooting experience across Node.js and JVM/Java services, with strong depth in at least one runtime environment.
Experience may include Node.js heap snapshots, CPU profiling, event-loop and memory analysis, as well as JVM GC log analysis, thread dumps, JVM tuning, and Java service latency investigation.Service-to-service authentication experience, including mTLS, certificate rotation, certificate format conversion, and JWT-based service authentication
Additional Information
Nice to Have
Web Components / Lit experience, to perform first-level debugging of UI-related issuesServer-side rendering or isomorphic runtime experienceCanary rollout / multi-version production operationsDistributed tracing and request-context correlationKEDA or event-driven autoscalingExperience with enterprise platform integration layers
What We Offer
Competitive salary and laptopProfessional development and training opportunitiesWork with cutting-edge cloud and container technologiesFlexible work arrangements and collaborative team environmentImpact on organization-wide digital transformation initiatives