
Mongodb · Dublin
MongoDB’s mission is to empower innovators to create, transform, and disrupt industries by unleashing the power of software and data. We enable organizations of...
MongoDB’s mission is to empower innovators to create, transform, and disrupt industries by unleashing the power of software and
data. We enable organizations of all sizes to easily build, scale, and run modern applications by helping them modernize legacy
workloads, embrace innovation, and unleash AI. Our industry-leading developer data platform, MongoDB Atlas, is the only globally
distributed, multi-cloud database and is available in more than 115 regions across AWS, Google Cloud, and Microsoft Azure. Atlas
allows customers to build and run applications anywhere—on premises, or across cloud providers. With offices worldwide and over
175,000 new developers signing up to use MongoDB every month, it’s no wonder that leading organizations, like Samsung and Toyota,
trust MongoDB to build next-generation, AI-powered applications.
The Site Reliability Engineering team designs and builds the global infrastructure on which we deploy our services, focusing on
the above mentioned flagship MongoDB Atlas platform. As our customers grow and globalize, our services must satisfy demands for
low-latency requests around the globe, and comply with various data sovereignty requirements. The SRE Team’s mission is to build
this increasingly complex infrastructure, while continually lowering the operational burden associated with it, and increasing our
internal visibility into the health of the system. We are strong believers in infrastructure-as-code and self-healing systems. The
SRE Team is fully integrated with all the other engineering teams, and the teams work closely together with a soft and traversable
boundary between their areas of responsibility.
We are looking to speak to candidates who are based in Dublin for our hybrid working model.
processes a billion metrics per day, and replicates tens of billions of database writes to our backup service
several cloud providers
work on that - participate in a weekly on-call rotation
MongoDB is built for change, empowering our customers and our people to innovate at the speed of the market. We have redefined the
data platform for the AI era, enabling builders to create, transform, and disrupt industries with software. MongoDB’s unified data
platform, the most widely available, globally distributed data platform on the market, helps organizations modernize legacy
workloads, embrace innovation, and unleash AI. Our cloud-native platform, MongoDB Atlas, is the only globally distributed,
multi-cloud data platform and is available across AWS, Google Cloud, and Microsoft Azure.
With offices worldwide and over 67,000 customers, including 75% of the Fortune 100 and AI-native startups, relying on MongoDB for
their most important applications, we’re powering the next era of software.
Our compass at MongoDB is our Leadership Commitment, guiding how and why we make decisions, show up for each other, and win. It’s
what makes us MongoDB.
To drive the personal growth and business impact of our employees, we’re committed to developing a supportive and enriching
culture for everyone. From employee affinity groups, to fertility assistance and a generous parental leave policy, we value our
employees’ wellbeing and want to support them along every step of their professional and personal journeys. Learn more about what
it’s like to work at MongoDB, and help us make an impact on the world!
MongoDB is committed to providing any necessary accommodations for individuals with disabilities within our application and
interview process. To request an accommodation due to a disability, please inform your recruiter.
MongoDB is an equal opportunities employer.
Req ID: 3263213607
Work Mode: Flex 1+ days per week in our Dublin office Department: Engineering ABOUT THE COMPANY LearnUpon partners with over 1,600 organisations globally to unlock the potential of employees, customers & members through learning that’s easy, scalable and focused on results. Read more about life at LearnUpon here. ABOUT THE TEAM Our Engineering organization is dedicated to building robust, scalable infrastructure that handles world-scale platform demands. As part of the Site Reliability Engineering (SRE) team, we focus on system architecture, absolute performance, and technical innovation. Operating with high ownership and technical expertise, we are responsible for the scale-out of the LearnUpon infrastructure, championing internal self-service tooling, and embedding a culture of observability and shared operational responsibility across all engineering squads. ABOUT THE OPPORTUNITY As a Staff Site Reliability Engineer, you will be a principal technical leader and a key catalyst for our infrastructure's evolution. In this role, you will take ownership of our core platform resilience, driving the strategy to build out an advanced, cost-effective observability function spanning metrics, logs, and transaction tracking. This opportunity requires a strategic thinker who can design cross-team SLO/SLI frameworks, navigate complex distributed system requirements, and mentor talent to ensure LearnUpon scales efficiently to support our ambitious global goals. In addition, you’ll be responsible for: * Infrastructure Optimization: Identify opportunities to improve and scale our infrastructure for performance, observability, maintainability, and cost, by creating innovative solutions. * Observability Function Strategy: Lead our efforts to build an observability function that incorporates application metrics, application transaction tracking, and event log management. * Resilience & Scaling: Drive the processes to maintain resilient, scalable, and cost-effective infrastructure while working with other Engineering teams to provide solutions that meet their ongoing requirements. * Tooling & Self-Service: Build tools focused on measuring, monitoring, and alerting, with an eye towards self-service in order to promote Engineers’ ownership of observability. * Operational Agility & Support: React quickly to changing customer and business needs and actively participate in the team's on-call rota. Team Up-Leveling: Mentor junior talent and effectively communicate complex technical ideas to both technical and non-technical peers. SKILLS & EXPERIENCE Must-Haves * 7+ years of experience in a software or Ops role. * 5+ years of cloud engineering experience, with at least 2 years of experience with AWS. * Experience deploying Microservice environments using containerisation technologies such as Kubernetes and Docker. * Experience designing and implementing Observability tech stacks, championing its benefits to Engineering teams, and managing the associated cost analysis of metrics gathering, effort, and tooling. * Ability to architect the design of SLO/SLI implementations that balance the needs of different teams. * Experience building and supporting large-scale distributed systems that back a consumer app or website with associated requirements of performance, security, and disaster recovery. * Experience with implementing IaC (e.g., CloudFormation, Terraform, etc.), automation tooling (e.g., Puppet, Ansible etc.), and CI/CD (e.g., Jenkins, Travis CI, GitLab, etc.). * Experience using AI tools to streamline tasks and improve efficiencies. Nice-to-Haves * Experience with database scaling would be a strong plus. * Certification in AWS, any PaaS, and/or related technologies. *If you don’t tick every box but believe this role is a mutually good fit, please don’t hesitate to apply. We’d love to hear from you. WHY CHOOSE LEARNUPON? From comprehensive rewards and generous time off to meaningful investment in your growth and development, LearnUpon gives you the support, trust, and opportunity to do the most impactful work of your career. Learn more here. HIRING PROCESS * Qualified applicants may be invited to an initial screening call with a member of our TA Team. * Successful candidates will be invited to a series of practical interviews. * Finally, candidates will have an interview with our CTO. * Successful candidates will be contacted with an offer to join our team. Note: At LearnUpon, we utilise AI to enhance the speed and quality of our screening and assessment practices, but our hiring decisions are always human. LearnUpon is an Equal Opportunities Employer. We do not discriminate on the basis of gender, marital status, family status, age disability, sexual orientation, race, religion, membership of the Traveller community, or any other legally protected status. Check out our Careers site and Instagram to learn more about working at LearnUpon. By submitting your application, you agree to LearnUpon's Privacy Policy.
MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. As the Site Reliability Engineering Manager for SLS, you will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll help grow and lead a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. We are looking to speak to candidates who are based in Dublin for our hybrid working model. RESPONSIBILITIES * Build and lead a team of 6-8 engineers, fostering a positive culture, handling career growth and performance conversations, and proactively removing blockers * Define and drive a clear technical vision and comprehensive roadmap for our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs * Contribute through hands-on technical work, such as leading architectural design reviews, reviewing PRs, and stepping in to guide the team through complex operational challenges * Act as the primary liaison for the Storage Layer Services SRE team, collaborating closely with other engineering leaders to ensure platform alignment and manage stakeholder expectations YOU MAY BE A GOOD FIT IF YOU * Have 10+ years of experience working on software and operating distributed systems, with 2+ years managing engineering teams * Possess a customer-focused mindset, treating internal developers as your primary users * Value efficiency in processes and operations, and have a track record of optimizing team workflows * Prefer automation over manual processes, fostering a culture of building software solutions to eliminate toil * Have deep technical familiarity with Kubernetes ecosystems, containerization technologies, and modern IaC tooling (e.g., Terraform, Crossplane, or Operators) so you can effectively guide the team's technical decisions * Have operated or supported stateful storage or database systems at scale and are comfortable with durability, consistency and recovery trade-offs * Excel at translating complex business and engineering requirements into actionable, phased technical roadmaps * Have a high level of empathy, responsibility, ownership, and accountability * Excellent verbal and written technical communication skills STRONG CANDIDATES MAY ALSO HAVE EXPERIENCE WITH * Leading major architectural shifts, such as moving from legacy storage stacks to new multi-tenant storage architectures, including planning and executing large-scale data and workload migrations with tight availability and durability requirements * Managing and scaling infrastructure across multi-cloud environments (AWS, GCP, or Azure) * Designing secure, multi-tenant runtime environments at scale ABOUT MONGODB MongoDB is built for change, empowering our customers and our people to innovate at the speed of the market. We have redefined the data platform for the AI era, enabling builders to create, transform, and disrupt industries with software. MongoDB’s unified data platform, the most widely available, globally distributed data platform on the market, helps organizations modernize legacy workloads, embrace innovation, and unleash AI. Our cloud-native platform, MongoDB Atlas, is the only globally distributed, multi-cloud data platform and is available across AWS, Google Cloud, and Microsoft Azure. With offices worldwide and over 67,000 customers, including 75% of the Fortune 100 and AI-native startups, relying on MongoDB for their most important applications, we’re powering the next era of software. Our compass at MongoDB is our Leadership Commitment, guiding how and why we make decisions, show up for each other, and win. It’s what makes us MongoDB. To drive the personal growth and business impact of our employees, we’re committed to developing a supportive and enriching culture for everyone. From employee affinity groups, to fertility assistance and a generous parental leave policy, we value our employees’ wellbeing and want to support them along every step of their professional and personal journeys. Learn more about what it’s like to work at MongoDB, and help us make an impact on the world! MongoDB is committed to providing any necessary accommodations for individuals with disabilities within our application and interview process. To request an accommodation due to a disability, please inform your recruiter. MongoDB is an equal opportunities employer. Req ID: 1273396229
Work Mode: Hybrid 1+ day per week in our Dublin office Department: Engineering ABOUT THE COMPANY LearnUpon partners with over 1,600 organisations globally to unlock the potential of employees, customers & members through learning that’s easy, scalable and focused on results. Read more about life at LearnUpon here. ABOUT THE TEAM Our Engineering team builds and delivers our best-in-class, world-scale Learning Management System (LMS). We focus on architectural excellence, reliability, and technical innovation across multiple teams and services. Operating with high ownership and deep technical expertise, we champion engineering best practices, seamless automation, and a strong culture of collaborative mentorship to support LearnUpon's global product growth. ABOUT THE OPPORTUNITY As a Staff Engineer, DevOps, you'll be a key technical leader and role model within LearnUpon. You'll shape the DevOps architecture of our best-in-class LMS, influence our technical strategy, and elevate the skills of our entire engineering organization. Your impact will extend beyond our product, contributing to our mission of building robust, lovable learning products that help customers achieve results. In addition, you’ll be responsible for: * Architecture & Strategy: Drive technical innovation, shape our long-term technical roadmap, and establish scalable DevOps architecture across multiple platform services * Automation & CI/CD: Establish and scale robust CI/CD pipelines and infrastructure automation, driving best practices across the engineering organization * Operational Excellence: Champion absolute standards for code quality, system design, scalability, and cloud production reliability * Mentorship & Leadership: Provide strategic technical guidance, acting as a prominent role model to uplevel the skills of our entire engineering and senior technical talent pool * Stakeholder Alignment: Collaborate closely with Product Engineering and cross-functional leadership to align infrastructure capabilities with business goals SKILLS & EXPERIENCE Must-Haves * Proven expertise in cloud infrastructure, particularly within AWS environments, with an advanced understanding of production scalability, security, and reliability * Strong production experience working with Kubernetes and Helm charts as the foundational layer for container platform orchestration * Deep proficiency in Infrastructure as Code (IaC) frameworks like Terraform or Ansible, focusing on modular and reusable architecture patterns * Extensive background designing CI/CD pipelines (GitLab preferred) and engineering automated internal tooling using Bash, Python, or Go * Demonstrated success leading large-scale technical initiatives that span across multiple engineering squads and services * Track record of mentoring, coaching, and elevating senior-level technical talent * Experience using AI tools to streamline tasks and improve efficiencies Nice-to-Haves * Familiarity with GitOps deployment workflows using tools such as ArgoCD or FluxCD * Prior experience optimizing massive multi-tenant platform infrastructure within a high-growth B2B SaaS architecture *If you don’t tick every box but believe this role is a mutually good fit, please don’t hesitate to apply. We’d love to hear from you. WHY CHOOSE LEARNUPON? From comprehensive rewards and generous time off to meaningful investment in your growth and development, LearnUpon gives you the support, trust, and opportunity to do the most impactful work of your career. Learn more here. HIRING PROCESS * Qualified applicants may be invited to an initial screening call with a member of our TA Team. * Successful candidates will be invited to a series of practical interviews. * Finally, candidates will have an interview with our CTO. * Successful candidates will be contacted with an offer to join our team. Note: At LearnUpon, we utilise AI to enhance the speed and quality of our screening and assessment practices, but our hiring decisions are always human. LearnUpon is an Equal Opportunities Employer. We do not discriminate on the basis of gender, marital status, family status, age disability, sexual orientation, race, religion, membership of the Traveller community, or any other legally protected status. Check out our Careers site and Instagram to learn more about working at LearnUpon. By submitting your application, you agree to LearnUpon's Privacy Policy.