Cloud Operations Engineer

Mongodbinc — Ireland · Posted ~2 weeks ago

Mid Full-time Remote

Skills

Linux administration cloud infrastructure monitoring networking database operations scripting Java Go JavaScript Linux AWS GCP Azure Kubernetes

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary

A global cloud services organization is hiring a Cloud Operations Engineer to monitor systems, troubleshoot incidents, automate operational tasks, and improve reliability for large-scale distributed platforms.

Highlights

Remote opportunity within Ireland with exposure to large-scale cloud operations, incident management, automation, and global engineering collaboration.

Description

MongoDB Atlas is the premier multi-cloud database-as-a-service built and operated by the makers of MongoDB. The Cloud Operations Engineering team at MongoDB is a worldwide team responsible for the consistent operational success of every MongoDB Atlas customer. As a Cloud Operations Engineer, you will help ensure the success of our Atlas customers, whether they are early startups or large multinational companies, cloud-native or just getting started with a digital transformation to the cloud. You are excited about the core mission of MongoDB, and the opportunity to join the team responsible for operating Atlas, the fastest-growing multi-cloud database-as-a-service in the world. You are prepared to be one of the early members of a 24/7/365 global cloud operations team. Cloud Operations Engineers will be responsible for day-to-day duties such as creating and monitoring system’s alert dashboards, reviewing critical events and system logs, accessing customer instances that underpin their production databases and performing server administration duties including performance troubleshooting. Applicants must be critical thinkers who are quick to detect, resolve, or escalate issues that are sometimes broad in scope and difficult to trace. At MongoDB you will grow your career and skills, wear multiple hats, and be part of an operations team that works at the frontier of Cloud services and database systems. We are looking to speak to candidates who will be based remotely in Ireland. Due to the 24/7 nature of our support organization, certain events throughout the year will require volunteering for coverage outside one’s normal work days or work hours (i.e. regional offsites, regional holidays, etc). These are typically announced weeks in advance with a sign-up system that considers equitability. Responsibilities Successfully coordinate and collaborate with a global team of Cloud Operations Engineers who are tasked with ensuring our uptime guarantees to our Atlas customer baseHelp scale the worldwide Cloud Operations Engineering team with the strategic implementation and refinement of new processes and toolsAssist in scoping, designing and deploying systems that reduce Mean Time to Resolve for customer incidentsMonitor and detect emerging customer-facing incidents on the Atlas platform; assist in their proactive resolutionAutomate routine monitoring and troubleshooting tasksDiagnose live incidents, differentiate between platform issues versus usage issues, and take the next steps toward resolutionAssist in performing root cause analysis after incident recovered; identifying any breakdowns in processes or workflows that contributed to the event and what changes need to be made to prevent similar eventsContribute to documentation of corner case scenarios, troubleshooting workflows and SOPs.Work alongside our product management, cloud engineering and support organizations by identifying areas for improvement in the management applications powering the Atlas infrastructureInform executive leadership and escalation management personnel of major outagesCoordinate and participate in a weekly on-call rotation, where you will handle short term customer incidents (proactively from automated monitoring or through reactive alerts via our Technical Services team) Requirements Experience with being an on call DevOps, SRE, or Cloud Operations engineer (at least 2 years)Expertise with Linux system administration, configuration, troubleshootingExperience in monitoring, system performance data collection and analysis, and reportingKnowledge of database operations and conceptsExpertise with networking technologies like DNS, TCP/IP, etc.Familiarity with Amazon Web Services and other Cloud infrastructure platforms (e.g. GCP, Azure)Knowledgeable about a wide range of web and internet technologiesCapability to write small programs/scripts to solve both short-term systems problemsA CS/CE degree or equivalent experienceAt least 1 of the following programming languages: Java, Go, JavascriptA keen interest in learning new things Nice To Have MongoDBSplunkKubernetes Benefits include Competitive salary, equity, pension and health insuranceRegular performance, compensation and development reviews20 weeks Maternity & Paternity leave to spend time with new arrivals About MongoDB MongoDB is built for change, empowering our customers and our people to innovate at the speed of the market. We have redefined the data platform for the AI era, enabling builders to create, transform, and disrupt industries with software. MongoDB’s unified data platform, the most widely available, globally distributed data platform on the market, helps organizations modernize legacy workloads, embrace innovation, and unleash AI. Our cloud-native platform, MongoDB Atlas, is the only globally distributed, multi-cloud data platform and is available across AWS, Google Cloud, and Microsoft Azure. With offices worldwide and over 67,000 customers, including 75% of the Fortune 100 and AI-native startups, relying on MongoDB for their most important applications, we’re powering the next era of software. Our compass at MongoDB is our Leadership Commitment, guiding how and why we make decisions, show up for each other, and win. It’s what makes us MongoDB. To drive the personal growth and business impact of our employees, we’re committed to developing a supportive and enriching culture for everyone. From employee affinity groups, to fertility assistance and a generous parental leave policy, we value our employees’ wellbeing and want to support them along every step of their professional and personal journeys. Learn more about what it’s like to work at MongoDB, and help us make an impact on the world! MongoDB is committed to providing any necessary accommodations for individuals with disabilities within our application and interview process. To request an accommodation due to a disability, please inform your recruiter. MongoDB is an equal opportunities employer. Req ID 4263312827