
Stripe · San Francisco
WHO WE ARE ABOUT STRIPE Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ...
Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the
most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission
is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented
opportunity to put the global economy within everyone's reach while doing the most important work of your career.
Stripe processes over $1.9T in payments volume per year, which is roughly 1.6% of the world's GDP, for millions of customers from
startups to enterprises. The tremendous amount of data makes Stripe one of the best places to do machine learning. While being an
integral part of almost every product line at Stripe (e.g., Payments, Radar, Capital, Billing, etc.), we have lots of exciting
opportunities to innovate in ML Platform at Stripe.
The ML Platform team builds the platforms and services that enable ML engineers and data scientists across Stripe to take data and
build features and models from prototype to production—reliably, at low latency, and at scale. Our scope spans ML training
infrastructure, model serving and deployment, feature computation and online serving, observability and monitoring, and agentic AI
capabilities. We work closely with product teams, data scientists, and platform infrastructure teams to build powerful, flexible,
and user-friendly systems that substantially increase ML velocity across the company.
You'll serve as a technical lead across the ML Platform space and a key contributor to the evolution of the platforms that power
Stripe's ML-driven products. As a Staff Engineer, you'll make decisions with a large impact on Stripe. You'll influence our
investments and strategy while making our systems more reliable, secure, and a delight to use. You'll work cross-functionally with
other technical staff, data science, product, and senior leadership to increase the impact of ML at Stripe.
You'll help define the long-term strategy and lead the technical direction for the next generation of ML infrastructure that
powers Stripe's ML-driven products.
orchestration, scalable CPU and GPU compute infrastructure, model training, LLM fine-tuning, low-latency model inference,
large-scale feature stores, real-time monitoring, and LLM and agent orchestration.
excellence with creative problem-solving.
and production operation.
scalable technical solutions.
constraints.
technical considerations related to the end-to-end ML lifecycle.
We're looking for someone who meets the minimum requirements to be considered for the role. If you meet these requirements, you
are encouraged to apply. The preferred qualifications are a bonus, not a requirement.
service-oriented architecture and large-scale distributed systems.
mentor team members.
orchestration, or ML data systems, with requirements for performance, reliability, scalability, and cost efficiency.
stakeholders.
engineers, product managers, and business stakeholders.
distributed training, model inference, feature stores, real-time feature computation, and model registries.
model evaluation.
retrieval-augmented generation).
WHO WE ARE ABOUT STRIPE Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. ABOUT THE TEAM Stripe processes over $1.9T in payments volume per year, which is roughly 1.6% of the world's GDP, for millions of customers from startups to enterprises. The tremendous amount of data makes Stripe one of the best places to do machine learning. While being an integral part of almost every product line at Stripe (e.g., Payments, Radar, Capital, Billing, etc.), we have lots of exciting opportunities to innovate in ML Platform at Stripe. The ML Platform team builds the platforms and services that enable ML engineers and data scientists across Stripe to take data and build features and models from prototype to production—reliably, at low latency, and at scale. Our scope spans ML training infrastructure, model serving and deployment, feature computation and online serving, observability and monitoring, and agentic AI capabilities. We work closely with product teams, data scientists, and platform infrastructure teams to build powerful, flexible, and user-friendly systems that substantially increase ML velocity across the company. WHAT YOU'LL DO You'll serve as a technical lead across the ML Platform space and a key contributor to the evolution of the platforms that power Stripe's ML-driven products. As a Staff Engineer, you'll make decisions with a large impact on Stripe. You'll influence our investments and strategy while making our systems more reliable, secure, and a delight to use. You'll work cross-functionally with other technical staff, data science, product, and senior leadership to increase the impact of ML at Stripe. You'll help define the long-term strategy and lead the technical direction for the next generation of ML infrastructure that powers Stripe's ML-driven products. RESPONSIBILITIES * Take ownership of end-to-end architecture and system design for large, complex projects across ML Platform. * Define technical direction for highly ambiguous projects, transforming complex user needs into long-lasting platform strategy. * Design system architectures for the most challenging ML Platform problems in one or more areas, including AI and ML workflow orchestration, scalable CPU and GPU compute infrastructure, model training, LLM fine-tuning, low-latency model inference, large-scale feature stores, real-time monitoring, and LLM and agent orchestration. * Turn high-leverage ideas into tangible, robust solutions that shape platform and product roadmap, combining technical excellence with creative problem-solving. * Scope and lead large projects with significant business impact, driving them from requirements through design, implementation, and production operation. * Work with ML engineers, data scientists, and product teams directly to translate their needs into functional requirements and scalable technical solutions. * Arbitrate critical decisions that balance competing priorities while meeting latency, reliability, cost, and security constraints. * Serve as a key engineering representative, engaging senior leaders across Stripe and advising the leadership team on key technical considerations related to the end-to-end ML lifecycle. * Drive cross-team technical initiatives that improve ML development velocity and MLOps maturity across the company. * Mentor and grow other engineers. Serve as a role model for designing, implementing, and operating great software systems. WHO YOU ARE We're looking for someone who meets the minimum requirements to be considered for the role. If you meet these requirements, you are encouraged to apply. The preferred qualifications are a bonus, not a requirement. MINIMUM REQUIREMENTS * 10+ years of professional software development experience, or equivalent domain expertise, with a solid background in service-oriented architecture and large-scale distributed systems. * Track record of serving as a technical lead, with the ability to provide technical direction, lead multi-team initiatives, and mentor team members. * Experience building and operating production ML platform in one or more areas such as model training, model serving, orchestration, or ML data systems, with requirements for performance, reliability, scalability, and cost efficiency. * Strong product instincts and a deep understanding of the business context in which you operate. * Strong communication skills with the ability to explain complex technical concepts to both technical and non-technical stakeholders. * Demonstrated ability to work cross-functionally, collaborating effectively with ML engineers, data scientists, software engineers, product managers, and business stakeholders. * The ability to thrive on a high level of autonomy and responsibility, and comfort operating in ambiguous environments. * Hands-on experience using AI tools to accelerate how you work. PREFERRED QUALIFICATIONS * Experience building large-scale ML training, serving, or data infrastructure for machine learning use cases, such as distributed training, model inference, feature stores, real-time feature computation, and model registries. * Experience with distributed ML training systems, accelerator-backed compute, training data pipelines, experiment tracking, and model evaluation. * Experience rapidly developing prototypes and iterating based on user feedback. * Experience training and shipping machine learning models to production to solve critical business problems. * Familiarity with LLMs, LLM application frameworks, and agentic AI patterns (e.g., tool use, multi-agent orchestration, retrieval-augmented generation). * Familiarity with cloud services (e.g., AWS) and cloud-based AI and ML services (e.g., SageMaker, Bedrock, Databricks, OpenAI). * Ability to synthesize ideas across the organization while setting a compelling technical vision. * Comfortable working with geographically distributed teams. * Passion for side projects, open source, or self-driven technical initiatives.
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Build the future of data. Join the Snowflake team. The Snowflake Machine Learning Platform team’s mission is to enable customers to bring their machine learning and deep learning workloads to Snowflake. Our customers want to build powerful models with the ever-increasing data in Snowflake but face several challenges including infrastructure optimizations, orchestration, performance, and security. The team aims to solve these challenges by building highly integrated platform solutions that are simple, secure, and enable end-to-end ML workflows. We are on an early journey to build the most scalable machine learning and data platform without sacrificing the benefits of a single platform and governance. We are looking for outstanding technical leaders who will join our ML Platform team to build the next-generation platform and play a pivotal role in this journey by understanding Snowflake’s core platform architecture and evolving it to enable state-of-the-art machine learning and LLM workloads. Join us to define strategies, set technical directions, design and execute, engage and deliver innovation, and unlock the power of AI for thousands of enterprise customers. This position is based in Menlo Park, CA, and Bellevue, WA. RESPONSIBILITIES: * Help define and own the roadmap, working collaboratively and proactively with senior architects, PMs, and team leadership. The initiatives include platforms and tools that enable customers to do state-of-the-art machine learning on Snowflake natively. * Collaboratively build and execute a vision for incorporating new advances in machine learning in ways that best achieve the team’s business objectives. * Ensure operational excellence of the services and meet the commitments to our customers regarding reliability, availability, and performance. * Collaborate across other ML partner teams to continuously improve ML development velocity and capabilities at Snowflake. * Support team members in delivering a high level of technical quality. IDEAL REQUIREMENTS & QUALIFICATIONS: * Have 7+ years of industry experience designing, building, and supporting Internet serving infrastructure, machine learning platforms, machine learning services, and frameworks. * Strong track record of working with machine learning systems and/or platforms. * Experience in serving LLMs using inference engines like vLLM, TensorRT-LLM, TEI, SGLang, and knowing tradeoffs between them. * Experience serving fine-tuned LLMs (PEFT, DPO, RL). * Experience with several of the following frameworks: SKLearn, XGBoost, PyTorch, Tensorflow, MLflow is a plus. * Previous experience in building batch and real-time ML serving systems preferred. * Have built a roadmap and vision around machine learning teams, and led technical decision making with help of architects and PMs and team. * BS/MS/PhD in Computer Science or related majors, or equivalent experience Snowflake employee is expected to follow the company’s confidentiality and security standards for handling sensitive data. Snowflake employees must abide by the company’s data security plan as an essential part of their duties. It is every employee's duty to keep customer information secure and confidential. Snowflake is growing fast, and we’re scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake. How do you want to make your impact? For jobs located in the United States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake.com
Reddit is a community of communities. It’s built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 126 million daily active unique visitors, Reddit is one of the internet’s largest sources of information. For more information, visit www.redditinc.com. Who We Are: The Machine Learning Platform team at Reddit is a high-impact team that owns the infrastructure that powers recommendations, content discovery, user and content quantification, while directly impacting other teams such as Growth, Ads, Feeds, and Core Machine Learning teams. What You’ll Do: As a Senior Staff Software Engineer, you will help define and lead the vision for Reddit’s large-scale GenAI Platform, shaping the strategy, architecture, and operating model that enable teams across the company to build, deploy, and scale generative AI products with confidence. Contribute to the design, implementation, and maintenance of the LLM Gateway, focusing on features like unified API endpoints for internal/externally hosted LLM, rate/token limit management, and intelligent failover mechanisms to boost uptime and reliability. * Lead and execute the vision, strategy, and roadmap for Reddit’s large-scale GenAI Platform. * Define the platform architecture and operating model that enable teams to build, deploy, and scale GenAI products reliably. * Drive the strategy for a unified LAG Gateway supporting internally and externally hosted LLMs through consistent APIs and abstractions. * Set the direction for core platform capabilities such as rate and token limit management, intelligent failover, and production resilience. * Shape Reddit’s approach to an enterprise-grade RAG system * Establish the strategic direction for agentic AI workflows and tool-use patterns across the platform. * Own the end-to-end platform strategy from concept through production adoption and long-term evolution. * Drive MLOps and LLMOps standards across CI/CD, testing, versioning, evaluation, and lifecycle management. * Define best practices for observability, monitoring, governance, and operational excellence across GenAI systems. * Partner across engineering, product, and leadership to align platform investments with company priorities and user needs. * Champion platform thinking with a strong focus on scalability, reliability, performance, and developer experience. * Influence technical direction across teams by turning emerging AI capabilities into a scalable platform strategy. Who You Might Be: * 10+ years of experience in ML Engineering, AI Platform Engineering, or Cloud AI Deployment roles. * Have a track record of leading technical strategy and delivering AI platforms in cloud-based production environments at scale. * Demonstrate strong execution by turning strategy into action, driving complex initiatives end to end, and consistently delivering high-quality platform outcomes. * Bring deep experience operating Kubernetes and other orchestration systems in large-scale production environments. * Deep experience with cloud-based technologies for supporting an ML platform, including tools like AWS, Google Cloud Storage, infrastructure-as-code (Terraform), and more * Proficiency with the common programming languages and frameworks of ML, such as Go, Python, etc. * Excellent communication skills with the ability to articulate technical AI concepts to non-technical stakeholders * Strong focus on scalability, reliability, performance, and developer experience. You are an undying advocate for platform users and have a deep intuition for the genAI product development lifecycle. * Strong knowledge of model serving, inference pipelines, monitoring, and observability for AI systems is a plus Benefits: * Comprehensive Healthcare Benefits and Income Replacement Programs * 401k with Employer Match * Global Benefit programs that fit your lifestyle, from workspace to professional development to caregiving support * Family Planning Support * Gender-Affirming Care * Mental Health & Coaching Benefits * Flexible Vacation & Paid Volunteer Time Off * Generous Paid Parental Leave Pay Transparency: This job posting may span more than one career level. In addition to base salary, this job is eligible to receive equity in the form of restricted stock units, and depending on the position offered, it may also be eligible to receive a commission. Additionally, Reddit offers a wide range of benefits to U.S.-based employees, including medical, dental, and vision insurance, 401(k) program with employer match, generous time off for vacation, and parental leave. To learn more, please visit https://www.redditinc.com/careers/. To provide greater transparency to candidates, we share base salary ranges for all US-based job postings regardless of state. We set standard base pay ranges for all roles based on function, level, and country location, benchmarked against similar stage growth companies. Final offer amounts are determined by multiple factors including, skills, depth of work experience and relevant licenses/credentials, and may vary from the amounts listed below. The base salary range for this position is: $292,500—$409,500 USD In select roles and locations, the interviews will be recorded, transcribed and summarized by artificial intelligence (AI). You will have the opportunity to opt out of recording, transcription and summarization prior to any scheduled interviews. During the interview, we will collect the following categories of personal information: Identifiers, Professional and Employment-Related Information, Sensory Information (audio/video recording), and any other categories of personal information you choose to share with us. We will use this information to evaluate your application for employment or an independent contractor role, as applicable. We will not sell your personal information or disclose it to any third party for their marketing purposes. We will delete any recording of your interview promptly after making a hiring decision. For more information about how we will handle your personal information, including our retention of it, please refer to our Candidate Privacy Policy for Potential Employees and Contractors. Reddit is proud to be an equal opportunity employer, and is committed to building a workforce representative of the diverse communities we serve. Reddit is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If, due to a disability, you need an accommodation during the interview process, please let your recruiter know.