Summary
✨ AI‑Generated
A reliability engineering role responsible for maintaining highly available systems through monitoring, incident management, automation, and operational best practices.
Highlights
Opportunity to improve system reliability, manage operational excellence, and collaborate across engineering and infrastructure teams.
Description
We are seeking a detail-oriented Site Reliability Engineer (SRE) to manage and oversee the daily reliability protocols, system availability, and operational quality standards.
In this role, you will act as the vital link between our software engineering teams, infrastructure operation teams, and product managers, ensuring that all services and systems are operated with maximum precision, durability, and adherence to strict industry governance and reliability benchmarks.
Core Responsibilities
Reliability Lifecycle Management: Coordinate daily operational activities, including incident response, service level monitoring (SLOs/SLIs), post-mortems, and root cause analysis (RCA).Regulatory & Technical Compliance: Maintain comprehensive technical documentation and version control for infrastructure and system designs; ensure all operations align with security, compliance, and standard operating procedures (SOPs).Data Integrity & Monitoring: Perform accurate telemetry analysis and system validation; apply sound data integrity principles to system logs, monitoring dashboards, and deployment records.Risk Mitigation: Monitor system performance, latency, and operational bottlenecks; report significant service outages or reliability deviations promptly to management and stakeholders.Quality & Performance Assurance: Author and maintain functional and technical specifications; assist in internal and external reliability and compliance audits to ensure industry standards for production systems are met.Technical Liaison: Serve as the primary point of contact for system audits and third-party infrastructure vendors during performance reviews and technical evaluations.Qualifications
Education: Bachelor’s degree in Computer Science, Software Engineering, Information Technology, or a related field (Required).Experience: 0–5 years of experience in site reliability engineering, DevOps, system administration, or software engineering environments.Technical Skills (Required):Strong understanding of cloud architecture, distributed systems, and Site Reliability Engineering principles.Proficiency in containerization, orchestration (e.g., Kubernetes, Docker), CI/CD pipelines, and Infrastructure as Code (IaC).Familiarity with observability tools (e.g., Prometheus, Grafana, ELK), incident management, and capacity planning.Excellent technical writing and organizational skills.Preferred Skills:Cloud certification (AWS, GCP, Azure) or Kubernetes Certification (CKA/CKAD).Experience with automated testing and deployment in highly scalable production environments.Knowledge of scripting/programming languages (e.g., Python, Go, Bash) and Agile methodologies.What We Offer
Targeted Placement: Direct marketing to our network of hiring managers in the Tech, Manufacturing, and MedTech industries.Technical Resume Rebuild: Optimization of your profile to highlight engineering expertise alongside process efficiency.Interview Coaching: Guidance on technical quality standards and system design case studies.Ready to Take the Next Step?
Email your updated resume to: murali.m@maximatek.com