
Mongodb · Palo Alto
ABOUT VOYAGE AI TEAM AT MONGODB Voyage AI team in MongoDB is building a best-in-class, general-purpose, domain-specific, and fine-tuned embedding models and re...
Voyage AI team in MongoDB is building a best-in-class, general-purpose, domain-specific, and fine-tuned embedding models and
rerankers to enable accurate, efficient unstructured data search and retrieval for RAG, recommendation, semantic search, and more.
It is backed by a strong team of AI researchers from Stanford, MIT, Berkeley, Princeton, and CMU, who have conducted over five
years of cutting-edge research on training embedding models. Voyage AI was acquired by MongoDB recently, and is now integrating
the SOTA embedding models with MongoDB's data platform to create powerful end-to-end solutions.
We are seeking a Staff Research Scientist to join our team and contribute to the development of next-generation AI models. This
position offers a unique opportunity to work on challenging problems at the intersection of machine learning research and
practical deployment of large neural networks.
This role can be based out of our Palo Alto office, or remotely in the United States.
accomplishments such as first author publications in top venues
MongoDB is built for change, empowering our customers and our people to innovate at the speed of the market. We have redefined the
data platform for the AI era, enabling builders to create, transform, and disrupt industries with software. MongoDB’s unified data
platform, the most widely available, globally distributed data platform on the market, helps organizations modernize legacy
workloads, embrace innovation, and unleash AI. Our cloud-native platform, MongoDB Atlas, is the only globally distributed,
multi-cloud data platform and is available across AWS, Google Cloud, and Microsoft Azure.
With offices worldwide and over 67,000 customers, including 75% of the Fortune 100 and AI-native startups, relying on MongoDB for
their most important applications, we’re powering the next era of software.
Our compass at MongoDB is our Leadership Commitment, guiding how and why we make decisions, show up for each other, and win. It’s
what makes us MongoDB.
To drive the personal growth and business impact of our employees, we’re committed to developing a supportive and enriching
culture for everyone. From employee affinity groups, to fertility assistance and a generous parental leave policy, we value our
employees’ wellbeing and want to support them along every step of their professional and personal journeys. Learn more about what
it’s like to work at MongoDB, and help us make an impact on the world!
MongoDB is committed to providing any necessary accommodations for individuals with disabilities within our application and
interview process. To request an accommodation due to a disability, please inform your recruiter.
MongoDB, Inc. provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination
and harassment of any type and makes all hiring decisions without regard to race, color, religion, age, sex, national origin,
disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other
characteristic protected by federal, state or local laws.
Req ID: 2273454547
MongoDB’s base salary range for this role is posted below. Compensation at the time of offer is unique to each candidate and based
on a variety of factors such as skill set, experience, qualifications, and work location. Salary is one part of MongoDB’s total
compensation and benefits package. Other benefits for eligible employees may include: equity, participation in the employee stock
purchase program, flexible paid time off, 20 weeks fully-paid gender-neutral parental leave, fertility and adoption assistance,
401(k) plan, mental health counseling, access to transgender-inclusive health insurance coverage, and health benefits offerings.
Please note, the base salary range listed below and the benefits in this paragraph are only applicable to U.S.-based candidates.
WHAT YOU WILL DO At Doctolib, we're revolutionizing healthcare delivery through advanced AI systems focused on medical reasoning. As a Senior Staff Research Scientist, you'll develop cutting-edge AI solutions that enhance clinical decision-making and support healthcare professionals in their daily practice. You'll work at the intersection of machine learning and healthcare, creating models that can understand medical knowledge, reason through complex clinical scenarios, and assist medical practitioners in providing better care. Your responsibilities include (but are not limited to): * Solve Real-World Healthcare Challenges: Design and build advanced machine learning models, especially in NLP and large language models, to improve healthcare workflows, enhance patient experiences, and power intelligent health solutions at scale. * Innovate & Experiment: Research and prototype novel algorithms and architectures, staying at the forefront of AI/ML developments. Bring new ideas from concept to working prototypes, leveraging the latest advances in deep learning and generative AI. * Collaborate Across Teams: Work closely with product, engineering, medical experts, and other stakeholders to translate business needs into impactful research projects. Communicate technical concepts and results clearly to both technical and non-technical audiences. * Drive Impact: Analyze large, complex data sets to extract actionable insights, define key metrics, and rigorously evaluate model performance. * Share & Learn: Advance Doctolib’s AI expertise by sharing findings within the team and with the wider ML community through publications, talks, and workshops. Mentor and support other scientists and engineers in adopting best practices. WHO YOU ARE * Technical Expertise: * PhD or Master’s degree (plus significant experience) in Computer Science, Machine Learning, Mathematics, or a related field; Publications as first author in top-tier AI/ML conferences such as NeurIPS, ICLR, ICML, ACL, EMNLP, or relevant medical informatics venues are a plus; * Deep understanding of machine learning and deep learning concepts, with hands-on experience in NLP, LLMs, or generative AI. * Strong programming skills in Python (and ideally experience with ML frameworks such as PyTorch or JAX). * Familiarity with modern data tools and scalable training on large datasets. * Innovator & Problem Solver: * Track record of independently driving research and translating it into real-world applications. * Experience designing experiments, evaluating intrinsic and extrinsic metrics, and iterating quickly. * Curiosity and drive to learn, research, and apply new machine learning techniques. * Collaborative & Communicative: * Comfortable working in highly cross-functional teams, with the ability to clearly communicate complex concepts to diverse audiences. * Passionate about sharing knowledge and contributing to a culture of learning and innovation. * Mission-Driven: * Motivated by Doctolib’s mission to make healthcare better for all. Eager to solve meaningful problems with technology. * Detail-oriented, rigorous, and committed to delivering high-quality, production-ready solutions. What we offer * Free comprehensive health insurance for you and your children * 25 days of paid vacation per year, plus up to 14 days of RTT * Free mental health and coaching services through our partner Moka.care * Work from abroad for up to 10 days per year thanks to our flexibility days policy * Lunch vouchers (Swile card) worth €8.50 per working day, with €4.50 covered by Doctolib * A subsidy from the work council to refund part of the membership to a sport club or a creative class * 50% reimbursement of your public transport subscription * Parent Care Program: receive one additional month of leave on top of the legal parental leave * For caregivers and workers with disabilities, a package including an adaptation of the remote policy, extra days off for medical reasons, and psychological support * Relocation support in case of international mobility * Access to the best AI tools for coding, development and dedicated training The interview process * HR Screen * Call with Research Team Member 30 mins * Scientific Presentation + Case Study 1 hour 30mins * Behavioral Interview 1 hour 15mins * At least one Reference check Job details * Permanent position * Full Time * Location: Doctolib Paris office in Levallois Perret * Work mode: Hybrid (3 days/week in the office) * Start date: ASAP If you would like to find out more about tech life at Doctolib, feel free to read our latest Medium blog articles! At Doctolib, we are committed to improving access to healthcare for everyone. This translates into our recruitment process. We evaluate candidates based solely on qualifications and motivation, without any form of discrimination. The more diverse ideas are heard, the more our product will truly improve healthcare for all. You are welcome to apply to Doctolib, regardless of your gender, religion, age, sexual orientation, ethnicity, disability. To ensure equal opportunities, we invite you to exclude personal information (e.g. pictures, age) from your applications. If you require any accommodation, please let us know for support during the hiring process. Join us in building the healthcare we all dream of! All information provided is processed by Doctolib for application management. For data processing details, click here. Please contact hr.dataprivacy(at)doctolib.com for inquiries or to exercise your rights.
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake's Data Engineering organization builds the platform that ingests, transforms, and stores data for modern lakehouse architectures — powering billions of queries, DML, and DDL operations with industry-leading price-performance. We lead the industry's shift to open data lakes through our work on Iceberg and Polaris, and we deliver capabilities like Snowpark, Dynamic Tables, cross-region replication, time travel, and zero-copy cloning at enterprise scale. We are investing in a new line of applied research — building toward verified data infrastructure and trustworthy data systems — that brings formal methods, automated reasoning, and modern AI techniques to bear on the hardest problems in our distributed systems and developer tooling. The goal is to improve correctness, reliability, and engineering velocity at a scale very few platforms operate at. We're hiring at both the Staff and Principal level; we'll calibrate the offer to the candidate's experience and scope of impact. WHAT YOU'LL DO * Lead research projects that apply formal methods, program analysis, automated reasoning, and AI-driven techniques (including code generation and modeling) to real problems in our cloud data platform. * Translate research ideas into prototypes, then into shipped capabilities that move concrete business metrics — quality, velocity, reliability, and operational performance at scale. * Partner closely with engineering leaders, product managers, and key customers to identify high-leverage opportunities and turn them into deliverables. * Influence the engineering and product roadmap; advise leaders on which research directions are pragmatic and which are not. * Train and uplevel engineering teams on new methods, and scale those methods across the organization. * Maintain expertise at the frontier of the field through publications, conference participation, open-source contributions, and patent filings. WHAT WE'RE LOOKING FOR * PhD (or equivalent research experience) in Computer Science or a closely related field. * Depth across the areas this role sits at the intersection of: * Formal methods — e.g., model checking, theorem proving, SAT/SMT, program verification, type systems, or program analysis. * Distributed systems — designing, reasoning about, or verifying large-scale concurrent and distributed systems. * Software engineering — strong fundamentals; able to go from a research idea to production-quality code in collaboration with engineering teams. * AI / ML — practical experience applying modern ML, including LLMs, to systems problems such as code generation, synthesis, or automated reasoning. * 8+ years applying theoretical computer science to large-scale software systems — ideally cloud data platforms, distributed systems, or developer infrastructure. * Demonstrated ability to drive company-level initiatives in partnership with engineering and product leadership. (Weighted more heavily for Principal-level candidates.) * Track record of technical contribution to the field — publications, open-source work, patents, or comparable evidence of impact. * Comfortable in a fast-paced, ambiguous environment where impact is measured by what ships. ABOUT WORKING HERE Every Snowflake employee is expected to follow the company's confidentiality and security standards for handling sensitive data, and to keep customer information secure and confidential as an essential part of their duties. Snowflake is growing fast, and we’re scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake. How do you want to make your impact? For jobs located in the United States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake.com
Reddit is a community of communities. It’s built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 126 million daily active unique visitors, Reddit is one of the internet’s largest sources of information. For more information, visit www.redditinc.com. Reddit is continuing to grow our teams with the best talent. This role is completely remote friendly within the United States. If you happen to live close to one of our physical office locations (San Francisco, Los Angeles, New York City & Chicago) our doors are open for you to come into the office as often as you'd like. The AI Engineering team at Reddit is building our own Reddit-native foundational Large Language Models (LLMs). This team sits at the intersection of applied research and massive-scale infrastructure, training models that truly understand the unique culture, language, and structure of Reddit communities. You'll join a team of distinguished engineers and researchers building the "engine room" of Reddit's AI future — the foundational models that power Safety & Moderation, Search, Ads, and the next generation of consumer products. As a Staff Research Engineer for Post-Training & Evaluation Science, you will own the science of our model development "feedback loop." While pre-training builds the base models, you define how we measure whether those models are safe, smart, and "Reddit-native," and you set the post-training methodology that turns base checkpoints into high-performing endpoints. You will define the Reddit Benchmark — our internal standard for rigorous model quality across both generation and representation — and own the evaluation science that the rest of the org's iteration depends on. RESPONSIBILITIES * Define the "Reddit Benchmark" evaluation standard: Own the methodology — not just the harness — for rigorously measuring model quality across Safety, Reasoning, representation/retrieval, and Reddit-specific knowledge. Decide what "Reddit-native" means in measurable terms and set the bar the org trains against. * Own evaluation reliability and statistical rigor: Establish the science behind trustworthy evals — judge variance, multi-sample scoring, inter-rater/inter-sample agreement, sampling and temperature effects, and calibration of automated judges. You are accountable for whether a benchmark delta is real or noise. Drive the practice of evaluation as a release gate — offline against frozen datasets, and pre-merge in CI/CD — so regressions are caught before endpoints ship. * Design model-as-a-judge methodology: Own judge selection, prompt design, calibration, and reliability for automated evaluation using frontier external models, enabling rapid, trustworthy iteration cycles. * Set post-training recipes and strategy: Design SFT recipes (data mixtures, curriculum, ablation strategy) that convert base models into helpful, well-aligned endpoints; partner with engineering to scale them. * Evaluate base and CPT checkpoints, not just endpoints: Design checkpoint-selection methodology across CPT experiments and LR studies, so we pick the right base before committing post-training compute. * Drive synthetic data generation strategy: Define and curate high-quality instruction and evaluation sets to improve generalization where human data is scarce. * Partner with Safety Engineering: Translate high-level safety policy into concrete classification metrics, probe sets, and CI/CD unit tests — including precision/recall at threshold, label-noise handling, and false-positive taxonomy for abuse detection (HHV). * Diagnose post-training instability: Dive into loss curves and eval logs to identify alignment tax and capability degradation, and recommend the fix. * Lead research direction: Set technical direction for evaluation and post-training across the team, mentor engineers and scientists, and represent the work internally (and externally where appropriate). REQUIRED QUALIFICATIONS * 6+ years of professional ML experience (or PhD + 4+) with a direct focus on LLM post-training and evaluation. * PhD or MS in CS, ML, NLP, IR, or a related quantitative field — or equivalent industry research experience. * Deep expertise in evaluation reliability: judge/sample variance, multi-sample scoring, calibration, statistical significance, and the failure modes of automated evaluation. * Strong experience building custom, domain-specific evaluation harnesses (e.g., lm-eval-harness, Inspect AI, LightEval) — you know the strengths and limits of benchmarks like MMLU and GSM8K and when they don't apply, and you treat eval sets as versioned, frozen, regression-tracked code. * Experience evaluating both generation and representation/classification: model-as-a-judge for generative quality and precision/recall, PR-AUC, retrieval/MTEB-style metrics, gold-label denoising, and label-noise handling. * Deep understanding of Continuous Pre-training (CPT), Instruction Tuning (SFT), and how data quality shapes model behavior. * Fluency in Python; strong data-pipeline and eval-harness engineering (e.g., Hugging Face Transformers, vLLM, lm-eval-harness). Working knowledge of PyTorch and distributed training (FSDP2, DeepSpeed ZeRO-3) sufficient to direct and debug post-training runs. NICE TO HAVE * Experience with MLflow or similar experiment-tracking frameworks. * Familiarity with modern fine-tuning frameworks (Axolotl, TorchTune) and PyTorch-native training stacks (TorchTitan). * Synthetic data generation techniques (e.g., Self-Instruct). * Experience with preference optimization (DPO, RLHF, RLAIF, GRPO). * Publications in NLP/ML/FAccT or related venues, or other evidence of research leadership. * Experience evaluating multimodal models (embeddings, hateful-memes-style classification). Benefits: * Comprehensive Healthcare Benefits and Income Replacement Programs * 401k with Employer Match * Global Benefit programs that fit your lifestyle, from workspace to professional development to caregiving support * Family Planning Support * Gender-Affirming Care * Mental Health & Coaching Benefits * Flexible Vacation & Paid Volunteer Time Off * Generous Paid Parental Leave #LI-SP1 Pay Transparency: This job posting may span more than one career level. In addition to base salary, this job is eligible to receive equity in the form of restricted stock units, and depending on the position offered, it may also be eligible to receive a commission. Additionally, Reddit offers a wide range of benefits to U.S.-based employees, including medical, dental, and vision insurance, 401(k) program with employer match, generous time off for vacation, and parental leave. To learn more, please visit https://www.redditinc.com/careers/. To provide greater transparency to candidates, we share base salary ranges for all US-based job postings regardless of state. We set standard base pay ranges for all roles based on function, level, and country location, benchmarked against similar stage growth companies. Final offer amounts are determined by multiple factors including, skills, depth of work experience and relevant licenses/credentials, and may vary from the amounts listed below. The base salary range for this position is: $230,000—$322,000 USD In select roles and locations, the interviews will be recorded, transcribed and summarized by artificial intelligence (AI). You will have the opportunity to opt out of recording, transcription and summarization prior to any scheduled interviews. During the interview, we will collect the following categories of personal information: Identifiers, Professional and Employment-Related Information, Sensory Information (audio/video recording), and any other categories of personal information you choose to share with us. We will use this information to evaluate your application for employment or an independent contractor role, as applicable. We will not sell your personal information or disclose it to any third party for their marketing purposes. We will delete any recording of your interview promptly after making a hiring decision. For more information about how we will handle your personal information, including our retention of it, please refer to our Candidate Privacy Policy for Potential Employees and Contractors. Reddit is proud to be an equal opportunity employer, and is committed to building a workforce representative of the diverse communities we serve. Reddit is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If, due to a disability, you need an accommodation during the interview process, please let your recruiter know.