Site Reliability Engineer
Mphasis โ United Kingdom ยท Posted ~1 day ago
๐ Log in to save this job, tailor your resume & track your apply process โ 7 days free, no card needed.
Log in to add to target listDescription
Job Description
As an SRE, you'll collaborate closely with Application Development and Operations teams to build and maintain scalable systems.
Your core focus will be to automate processes and ensure the highest levels of service reliability, specifically by reducing manual effort (TOIL).
You'll bring a strong passion for continually improving the reliability, availability, and performance of our services.
Primary Skill โ Experience with cloud platforms Primarily in AWS Cloud (e.g., AWS, GCP, Azure) and Container Orchestration (e.g., Kubernetes, Docker).
Proficiency in Monitoring and Logging Tools : Datadog, Splunk, Dynatrace, AppDynamics, Prometheus, Grafana, ELK Stack (Elasticsearch, Logstash, Kibana), Cloude Watch, Gremlin, Thousand Eyes.
Terraform, Jenkins, GitLab CI, PostgreSQL, Redis, Kong API.
Infrastructure skills, Networking and Security Skills, AWS (Atlas), ECS Based internal tooling
Lucidchart, PlantUML
Secondary Skill โSNOW, Jira, Shell Script, Linux, bitbucket, Akamai, DevOps
Expectations :
6+ years of experience with SRE principals and tools (AWS, Monitoring tool, DevOps, CI/CD etc.) should have worked with Toil identificationProven experience as a Site Reliability Engineer, DevOps Engineer, or similar role.Achieve and Maintain the Define SLI, SLO, SLA with business/operations/Engineering team.Monitoring, logging, event detection, Alerting and Error budget (99.9 , 99.99, 99.999 % ) for Cloud or Distributed platforms software, Operations & Business.Good understanding of programming skills in languages such as Python, Go, Java, or Ruby.Experience with cloud platforms Primarily in AWS Cloud (e.g., AWS, GCP, Azure) and container orchestration (e.g., Kubernetes, Docker).Skilled in managing configuration, deployments, observability, handling and resolving incidents, including root cause analysis, managing and operating complex systems for scalability, availability and performance.Experience with CI/CD pipelines and DevOps tools.Good Knowledge of Source Code Repository ToolsGood Knowledge of networking and Security concepts.Proficiency in Monitoring, Logging & Traceability ToolsGood Understanding of Database Systems โ Aurora PostgreSQL, REDISInfrastructure skills: Terraform for infrastructure as code, Skilled in the understanding of use, core cloud application infrastructure services including identity platforms, networking, storage, databases, containers, and serverlessEfficiency in creating Dashboard for Infra / APM / E2E workflows.ITIL โ Incident/ Change, Proficient in Problem management and Jira โ Blameless postmortem, Root Cause Analysis findings, applying permanent fixes, Documentation as Runbooks for lesson learnHands on experience on technical operations application support and stability, reliability and resiliencyProficient in communication and collaboration skills to work effectively with development and operations teams.Willingness to learn new tools, technologies and adapt to changing environments.
Nature of the Job:
SRE end-to-end Support & Monitoring as an Engineer and Architecture.Design, implement to maintain scalable and reliable systems for services.Setup and Monitor system performance and reliability, identifying and resolving issues proactively.Assess current state and mature the SRE function.Collaborate with development teams to improve system architecture and design for reliability.Participate and own the issues and incident response efforts to resolve production issues.
Resolve incidents and provide Incident Resolution with root cause analysis (RCA).Conduct post-mortem analyses of incidents to identify root causes and implement preventive measures.Create and maintain documentation as a Runbook for systems, processes, and procedures.Advocate for best practices in software development, system design, and operational excellence with Tool Implementation.Create and maintain monitoring technologies and processes that improve the visibility of our applications' performance and business metrics and keep operational workload in-check.Establish and ensure the repeatability, traceability, and transparency of application components.Monitoring and optimization of system performance and resource usage, identify and address bottlenecks, and implement best practices for performance tuning.Troubleshoot and resolve critical issues across multiple layers, including CDN, Load balancers, storage, OS, network, K8S, virtualization, and application/DB stack.
We have 66,600 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume โ in under a minute we'll analyze all 66,600 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume