Summary
✨ AI‑Generated
An enterprise technology team is looking for an experienced Platform and SRE Engineer to build scalable, reliable, and self-service platform capabilities. You will use Python automation, Kubernetes, CI/CD, and observability tooling to reduce operational toil, improve developer productivity, and establish strong reliability practices.
Highlights
Build scalable and reliable internal platforms, automate operational work, improve developer productivity, and apply reliability engineering practices across complex enterprise environments.
Description
Job Description - Platform & SRE Engineer
Experienced Platform & SRE Engineer with expertise in Python automation, Kubernetes, CI/CD, and observability platforms, focused on building scalable, reliable, and self-service platform solutions.
Proven track record of improving developer productivity, reducing operational toil through automation, and driving reliability engineering best practices across complex enterprise environments.
Skilled in platform engineering, cloud-native technologies, and operational excellence with a strong focus on scalability, resilience, and continuous improvement.
Day to Day job Duties: (what this person will do on a daily/weekly basis)
• Build and maintain platform automation solutions using Python and related tooling.
• Design, improve, and operate internal platform capabilities that enhance developer productivity and operational reliability.
• Develop reusable frameworks and tooling for deployments, monitoring, remediation, reporting, and self-service operations.
• Build and maintain CI/CD pipelines to support reliable and repeatable software delivery.
• Drive observability initiatives using metrics, dashboards, logging, monitoring, and alerting platforms.
• Collaborate with application teams to improve platform adoption, deployment standards, and operational excellence.
• Support and enhance Kubernetes-based application platforms and deployment frameworks.
• Work with authentication and authorization technologies including OAuth2, SSO, and JWT.
• Analyze application, infrastructure, and database performance to identify optimization opportunities.
• Contribute to reliability engineering practices including automation, resilience, capacity planning, and operational readiness.
• Partner with engineering teams to reduce operational toil through automation and platform improvements.
Basic Qualifications:
• Minimum 5+ years of experience in Python development, automation, and software engineering.
• Minimum 4+ years of experience building operational tooling, automation frameworks, platform services, or developer productivity solutions.
• Minimum 4+ years of experience administering Linux systems and developing shell scripting solutions.
• Minimum 3+ years of experience with observability platforms including Prometheus, Grafana, logging, monitoring, and alerting solutions.
• Strong understanding of SQL and database fundamentals.
• Minimum 3+ years of experience designing and supporting CI/CD pipelines using Jenkins, GitLab CI, GitHub Actions, or similar platforms.
• Minimum 3+ years of hands-on experience with Kubernetes and Docker in enterprise environments.
• Strong understanding of web applications, APIs, microservices, and distributed service architectures.
• Experience with OAuth2, SSO, JWT, and modern authentication technologies.