
Stripe · Seattle
WHO WE ARE ABOUT STRIPE Stripe is a financial infrastructure platform for businesses. Millions of companies — from the world's largest enterprises to the mos...
Stripe is a financial infrastructure platform for businesses. Millions of companies — from the world's largest enterprises to the
most ambitious startups — use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our
mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an
unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career.
The Core Change Management group is responsible for the systems that let every Stripe engineer ship code, configuration, and
infrastructure changes safely and at high velocity. You will be embedded primarily on the Service Deployments team — the owners of
Stripe's end-to-end code deployment platform — with regular collaboration with the Resource Automation and Feature Deployments
teams.
Service Deployments owns the full lifecycle of software changes at Stripe. The team's mission is to let developers roll out code
and configuration changes safely without sacrificing productivity, with a goal of meaningfully reducing change-related production
incidents year over year. The team operates a meaningful on-call rotation and owns the systems that sit in the critical path of
every engineer's daily workflow at Stripe.
Resource Automation owns the safe-by-default infrastructure change layer: automated remote execution, incremental Infrastructure
as Code tooling, cloud resource inventory, and cloud account governance and IAM role management. You will collaborate with this
team on projects that span the boundary between deployment orchestration and cloud resource management.
Feature Deployments owns Stripe's feature flag system, merchant entitlements, configuration management and distribution, and the
audit log of change-correlated events. You will work with this team when deployment pipelines intersect with feature rollout and
change safety tooling.
workflow at Stripe. The decisions you make affect thousands of deploys per day across hundreds of services, directly
determining how fast and safely Stripe's product evolves.
host-based services at scale, adding intelligent multi-service deploy pipelines, extending real-time anomaly detection to
earlier stages of traffic shifts, and rebuilding deployment event infrastructure on top of a durable message bus. This is not
maintenance work — the architecture is in motion.
business logic to the developer-facing internal platform UI. The problems are multi-layered: reliability, developer experience,
performance, and safety all at once.
better detection, smarter pipelines, and safer defaults. Your technical decisions have a direct and measurable safety impact on
Stripe's reliability.
systems, author designs that span multiple teams, and be the person engineering managers and engineers turn to for the hardest
deployment infrastructure questions.
and long-term reliability. Author the design, sequence the work, unblock the team, and shepherd projects to landed impact.
— including multi-service dependency-aware autodeploy pipelines, Kubernetes-native deployment primitives, and fleetwide
container migration — defining the API contracts, rollout strategies, and operational model that hundreds of teams depend on.
API/method-based regression detection, and build a self-service onboarding system that makes anomaly detection the default for
all supported service types.
Stripe's fleet of host-based services to containerized, fleetwide deployments — keeping the production deployment system
operational while executing the transformation.
operational toil; and make reliability, security, and maintainability first-class properties of the systems you own.
schema, durability model, and integration contracts that downstream systems rely on for observability and automation.
cloud resource management (IAM, account provisioning, infrastructure automation), with Feature Deployments on change-safety
tooling (feature flags, configuration management, change audit logs) that integrates with or depends on the deployment
pipeline, and with the service mesh team on routing capabilities that enable advanced deployment patterns such as canary
rollouts and merchant-priority traffic shaping.
senior engineers through high-stakes architectural decisions, and advocate for the right abstractions — code that consuming
teams can adopt without becoming deployment infrastructure experts.
engineers grow by framing problems clearly and providing decisive technical guidance on the hardest questions.
production infrastructure systems of significant scale and complexity.
managing cross-team dependencies and coordinating migrations across many consuming teams.
and operated at scale, including rollout strategies, staged delivery, and failure modes.
scheduling, and the operational challenges of migrating large fleets from VM-based to containerized infrastructure.
toil, and build systems that are reliable, debuggable, and maintainable by a team.
effect through code review and mentorship, and the ability to set technical direction for a team rather than just execute
within it.
that reduce the blast radius of bad deployments.
notification.
governance and IAM management in AWS or Azure.
abstractions that reduce toil for the engineering teams that depend on your platform.
infrastructure that provides safety guardrails around production changes.
effectively with partner teams on routing capabilities that enable advanced deployment patterns.
Office-assigned Stripes in most of our locations are currently expected to spend at least 50% of the time in a given month in
their local office or with users. This expectation may vary depending on role, team and location. For example, Stripes in Stripe
Delivery Center roles in Mexico City, Mexico, Bengaluru, India, and Dublin, Ireland work 100% from the office. Also, some teams
have greater in-office attendance requirements, to appropriately support our users and workflows, which the hiring manager will
discuss. This approach helps strike a balance between bringing people together for in-person collaboration and learning from each
other, while supporting flexibility when possible.
This new team will architect the platform foundation and infrastructure that enables a "build once, run anywhere" model, ensuring our AI-powered modernisation suite operates seamlessly regardless of a client's security or network constraints. We are looking for engineers to join this high-visibility initiative, where you will solve unique distributed systems puzzles and help shape the future of how global enterprises leverage GenAI. We are looking for an experienced software engineer who thrives on solving infrastructure constraints and building enterprise facing platforms. The ideal candidate is a hands-on technical leader who can architect complex distributed systems, mentor engineers, and collaborate with product teams to deliver a platform that minimizes deployment friction and meets customers’ compliance requirements. This role can be based out of our Gurgaon office or Remotely in India. THE IDEAL CANDIDATE FOR THIS ROLE WILL HAVE * 8+ years of software development and operations experience, with a focus on building platforms and distributable software infrastructure. * 2+ years of experience leading, coaching, and mentoring a team of engineers to achieve high-impact results. * Deep experience designing distributable applications that run in air-gapped or highly restricted network environments. * Strong proficiency in containerization and orchestration, with the ability to design systems where the host executes containerized applications under strict security monitoring. * Experience building Service Host architectures that provide shared routing, proxy layers, and API gateways for multiple underlying services. * Experience managing persistent storage solutions (blob/file storage) within distributed systems. * Understanding of security-first design, specifically regarding the execution of untrusted code, fine-grain access control, and rigorous input validation. * Experience designing auto-update mechanisms for software that cannot directly access the public internet. * Curiosity, a positive attitude, and a drive to continue learning. POSITION EXPECTATIONS * Provide technical leadership and mentorship to a team of engineers, fostering a culture of innovation, quality, and continuous improvement. * Drive the distributed infrastructure for the Application Modernisation Platform, ensuring it supports a self-service SaaS product model while operating within client-controlled environments. * Lead the design and implementation of a sophisticated orchestration layer that manages authentication, telemetry, and the lifecycle of modernization tools. * Architect a client-side application that manages developer workflows and communicates securely with hosted services. * Solve complex "Day 2" operations challenges, such as designing interfaces to solve DevOps challenges to manage users, secrets, and updates. * Collaborate with security and product teams to ensure the platform can execute code generation and LLM-based tasks securely, protecting both client data and MongoDB IP. * Act as a hands-on leader, guiding the team through technical challenges related to container security, local vs. hosted logic distribution, and cross-service data sharing. SUCCESS MEASURES Within the first three months, you will have: * Fully ramped up on our architecture, business domain, and the specific constraints of deploying software to large enterprise client environments. * Established yourself as the technical leader within the team, actively driving technical decisions. * Begun mentoring team members and collaborating with product counterparts to define the integrations between the platform and the tools it hosts. Within six months, you will have: * Successfully led the team to deliver a prototype or initial release of the Deployment Platform that supports the core modernization toolset. * Designed and implemented a streamlined process for platform updates that works within client security constraints. * Established a precise rhythm of execution and served as a role model for a high-performing team. Within 12 months, you will have: * Become recognized as the go-to technical expert for the deployment and operations domain at MongoDB. * Delivered a stable, production-ready platform that allows our Field Engineering teams to deploy tools without significant delays or approvals. * Played a crucial role in helping recruiting and growing the engineering team. ABOUT MONGODB MongoDB is built for change, empowering our customers and our people to innovate at the speed of the market. We have redefined the data platform for the AI era, enabling builders to create, transform, and disrupt industries with software. MongoDB’s unified data platform, the most widely available, globally distributed data platform on the market, helps organizations modernize legacy workloads, embrace innovation, and unleash AI. Our cloud-native platform, MongoDB Atlas, is the only globally distributed, multi-cloud data platform and is available across AWS, Google Cloud, and Microsoft Azure. With offices worldwide and over 67,000 customers, including 75% of the Fortune 100 and AI-native startups, relying on MongoDB for their most important applications, we’re powering the next era of software. Our compass at MongoDB is our Leadership Commitment, guiding how and why we make decisions, show up for each other, and win. It’s what makes us MongoDB. To drive the personal growth and business impact of our employees, we’re committed to developing a supportive and enriching culture for everyone. From employee affinity groups, to fertility assistance and a generous parental leave policy, we value our employees’ wellbeing and want to support them along every step of their professional and personal journeys. Learn more about what it’s like to work at MongoDB, and help us make an impact on the world! MongoDB is committed to providing any necessary accommodations for individuals with disabilities within our application and interview process. To request an accommodation due to a disability, please inform your recruiter. MongoDB is an equal opportunities employer. Requisition ID 1273352650
The Application Modernization Platform (AMP) team is tackling one of the industry's most critical challenges: leveraging Generative AI to transform rigid, legacy applications into modern, microservices-based architectures powered by MongoDB. We are building a comprehensive, SaaS-like platform, encompassing both the "brain" (multi-agent reasoning and orchestration) and the "hands" (the deployment platform and modernization toolset). This solution requires a robust platform foundation and infrastructure designed for a "build once, run anywhere" model, ensuring seamless operation regardless of a client's security or network constraints. A key challenge is balancing the need to tune our tools for each customer's unique tech stack and restrictive environments with making them easily extensible and scalable for common application modernization challenges. We seek an engineering leader for this high-visibility initiative. This role requires defining the high-level strategy and technical direction across all AMP engineering pillars, leading the execution of solving uniquely complex application modernization puzzles, and delivering an enterprise-grade product. The leader will minimize deployment friction, meet customer compliance requirements, and help shape the future of how global enterprises leverage GenAI. The ideal candidate is a hands-on technical leader who excels at leveraging GenAI capabilities, architecting complex distributed systems, and designing the orchestration agents necessary to reliably and fluidly run the entire software development lifecycle. This role will be based in North America's West Coast (PST), and offers a hybrid working model. THE IDEAL CANDIDATE FOR THIS ROLE WILL HAVE * 10+ years of software development and operations experience, with a focus on building platforms and distributable software infrastructure * Deep experience in building data warehouses and core components for data processing systems * Have experience in using GenAI in building complex large-scale systems, and know what GenAI is best suited to do * Understanding of security-first design, specifically regarding the execution of untrusted code, fine-grain access control, and rigorous input validation * 2+ years of experience leading, coaching, and mentoring a team of engineers to achieve high-impact results * Excellent verbal and written technical communication skills and a strong desire to collaborate cross-functions * Excellent time and project management skills including the ability to make realistic assessments of project cost and complexity * Curiosity, a positive attitude, and a drive to continue learning, in particular building AI skillset KEY RESPONSIBILITIES * Provide technical leadership and mentorship to a team of engineers, fostering a culture of innovation, quality, and continuous improvement * Provide thought leadership, and execute AMP roadmaps to influence and shape the future of how global enterprises leverage GenAI * Lead the design and implementation of a sophisticated orchestration layer that manages authentication, telemetry, and the workflow of application modernization stages * Lead the team design and leverage GenAI to build reliable and scalable building blocks that are extendable and easy to assemble to an autonomous workflow * Collaborate with MongoDB product teams to ensure the modernized applications integrate seamlessly with customers’ database ecosystem * Act as a hands-on leader, guiding the team through technical challenges related to where and how to leverage GenAI to build the critical building blocks of application modernization toolset SUCCESS MEASURES Within the first three months, you will have: * Fully ramped up on our architecture, business domain, and the specific constraints of deploying software to large enterprise client environments * Established yourself as the technical leader within the team, actively driving engineering strategy and key technical decisions * Begun collaborating with team leaders (TLs, managers) of AMP org and product counterparts on designing/improving the orchestration framework and the workflow components it manages * Developed a comprehensive grasp of all AMP pillars and their roadmap Within six months, you will have: * Influenced AMP product and engineering leads on key technical decisions, and set clear direction for the team * Successfully led the team to making substantial progress on the “big rock” deliverables and critical milestones of AMP product, enabling the team to smoothly land the planned roadmap * Owned and delivered some key pieces of AMP orchestration framework and the workflow components it manages * Served as a role model and strategic leader for the AMP org * Contributed to establish engineering excellence standards for the team, and began mentoring and providing technical guidance to team members Within 12 months, you will have: * Become recognized as the strategic engineering leader both within and outside of the AMP organization * Helped senior leadership (e.g., Directors/VP) on defining the technical strategy and roadmap for the future * Enabled the team to deliver a stable, production-ready platform that supports autonomous application modernization workflow for our Field Engineering teams * Drove the adoption of engineering excellence, making it a core tenet of the team's culture and operations * Contributed to the development of tech leads within the team, fostering growth opportunities. Played a key role in shaping the AMP organization into a self-driving high-performing engineering team ABOUT MONGODB MongoDB is built for change, empowering our customers and our people to innovate at the speed of the market. We have redefined the data platform for the AI era, enabling builders to create, transform, and disrupt industries with software. MongoDB’s unified data platform, the most widely available, globally distributed data platform on the market, helps organizations modernize legacy workloads, embrace innovation, and unleash AI. Our cloud-native platform, MongoDB Atlas, is the only globally distributed, multi-cloud data platform and is available across AWS, Google Cloud, and Microsoft Azure. With offices worldwide and over 67,000 customers, including 75% of the Fortune 100 and AI-native startups, relying on MongoDB for their most important applications, we’re powering the next era of software. Our compass at MongoDB is our Leadership Commitment, guiding how and why we make decisions, show up for each other, and win. It’s what makes us MongoDB. To drive the personal growth and business impact of our employees, we’re committed to developing a supportive and enriching culture for everyone. From employee affinity groups, to fertility assistance and a generous parental leave policy, we value our employees’ wellbeing and want to support them along every step of their professional and personal journeys. Learn more about what it’s like to work at MongoDB, and help us make an impact on the world! MongoDB is committed to providing any necessary accommodations for individuals with disabilities within our application and interview process. To request an accommodation due to a disability, please inform your recruiter. MongoDB, Inc. provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type and makes all hiring decisions without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws. Req. ID: 1273362664 MongoDB’s base salary range for this role is posted below. Compensation at the time of offer is unique to each candidate and based on a variety of factors such as skill set, experience, qualifications, and work location. Salary is one part of MongoDB’s total compensation and benefits package. Other benefits for eligible employees may include: equity, participation in the employee stock purchase program, flexible paid time off, 20 weeks fully-paid gender-neutral parental leave, fertility and adoption assistance, 401(k) plan, mental health counseling, access to transgender-inclusive health insurance coverage, and health benefits offerings. Please note, the base salary range listed below and the benefits in this paragraph are only applicable to U.S.-based candidates. MongoDB’s base salary range for this role in the U.S. is: $211,000—$363,000 USD
Reddit is a community of communities. It’s built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 126 million daily active unique visitors, Reddit is one of the internet’s largest sources of information. For more information, visit www.redditinc.com. As Reddit continues to scale globally, reliability and performance are more critical than ever. The Site Experience SRE team sits at the intersection of infrastructure, product engineering, and user experience - ensuring that every interaction across web, mobile, APIs, feeds, media delivery, and real time systems is fast, reliable, and resilient. We are looking for a Staff Site Reliability Engineer to lead reliability engineering initiatives for critical user facing systems at internet scale. In this role, you will partner closely with product and infrastructure teams to improve availability, latency, scalability, and operational excellence across Reddit’s most business critical experiences. This is a highly technical leadership role for someone who thrives in large-scale distributed systems, enjoys solving complex reliability challenges, and can influence engineering culture across the organization. WHAT YOU’LL DO: * Lead Reliability Engineering for User Experience * Drive reliability, scalability, and operational excellence for critical user facing systems and services. Improve performance and resiliency across APIs, content delivery, feed generation, search, messaging, and real-time experiences. * Architect for Scale * Partner with product and infrastructure engineering teams to design systems that remain highly available and performant under massive global load. Guide architectural decisions around failover, redundancy, graceful degradation, traffic management, and capacity planning. * Reduce Operational Risk * Identify systemic risks and reliability bottlenecks across services, dependencies, deployments, and infrastructure. Build proactive mitigation strategies and drive engineering improvements that reduce incidents and improve service health. * Drive Automation * Eliminate repetitive operational work through automation and tooling. Build systems that improve deployment safety, incident response, remediation workflows, and reliability guardrails * Incident Management * Lead complex incident response efforts across engineering teams. Drive blameless postmortems, identify root causes, and ensure sustainable long-term fixes are implemented. Influence Engineering Standards * Define and champion best practices around reliability engineering, SLIs/SLOs, capacity management, release engineering, and operational maturity across the company. Mentor and Multiply Impact * Provide technical leadership and mentorship to engineers across SRE and software engineering teams. Help shape reliability culture and raise the operational excellence bar across the organization. WHAT WE’RE LOOKING FOR * 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or related roles operating large scale distributed systems. * Strong collaboration and communication skills with the ability to influence technical direction across teams. * Strong experience supporting high traffic, user facing production environments. * Deep understanding of one or more: distributed systems, networking, Linux systems, cloud native architectures. * Experience designing highly available systems with strong operational and reliability practices. * Strong programming skills in languages such as Go, Python, or similar. * Strong understanding of observability systems including metrics, logging, tracing, and alerting. * Experience improving reliability through SLOs, automation, incident management, and performance optimization. * Demonstrated ability to troubleshoot complex issues across applications, infrastructure, networking, and services. Nice to Have * Experience operating systems at internet scale traffic volumes. * Experience with Kubernetes, containers, cloud infrastructure, and modern deployment platforms. * Familiarity with technologies such as Prometheus, Grafana, OpenTelemetry, Envoy, Kafka, ClickHouse, Cassandra, Redis, or similar distributed infrastructure technologies. * Experience with CDN optimization, edge reliability, traffic engineering, or global infrastructure. * Contributions to open source software or participation in technical communities. * Experience leading large scale incident response and operational transformation initiatives. Why Join Reddit? You’ll help shape the reliability and performance of one of the internet’s largest platforms, influencing experiences used by millions of people every day. This is an opportunity to solve deeply complex engineering problems at massive scale while helping define the future of reliability engineering for a modern consumer platform. Benefits * Global Benefit programs that fit your lifestyle, from workspace to professional development to caregiving support * Family Planning Support * Gender-Affirming Care * Mental Health & Coaching Benefits * Group Personal Pension Scheme with Employer match * Private Medical and Dental Scheme * Income Replacement Programs * Bike to Work scheme * Flexible Vacation & Paid Volunteer Time Off * Generous Paid Parental Leave * * Generous Paid Parental Leave * * Generous Paid Parental Leave * In select roles and locations, the interviews will be recorded, transcribed and summarized by artificial intelligence (AI). You will have the opportunity to opt out of recording, transcription and summarization prior to any scheduled interviews. During the interview, we will collect the following categories of personal information: Identifiers, Professional and Employment-Related Information, Sensory Information (audio/video recording), and any other categories of personal information you choose to share with us. We will use this information to evaluate your application for employment or an independent contractor role, as applicable. We will not sell your personal information or disclose it to any third party for their marketing purposes. We will delete any recording of your interview promptly after making a hiring decision. For more information about how we will handle your personal information, including our retention of it, please refer to our Candidate Privacy Policy for Potential Employees and Contractors. Reddit is proud to be an equal opportunity employer, and is committed to building a workforce representative of the diverse communities we serve. Reddit is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If, due to a disability, you need an accommodation during the interview process, please let your recruiter know.