
Manychat · Barcelona, Spain
WHO WE ARE 🌍 We help creators get more out of every conversation with Instagram-focused automations and support for other channels like Messenger, WhatsApp, a...
We help creators get more out of every conversation with Instagram-focused automations and support for other channels like
Messenger, WhatsApp, and TikTok. The result? Better engagement, more sales, and real, sustainable growth.
With a diverse team of 350+ people spread across three continents, we’re building the leading Chat Marketing platform that is used
— and loved — by more than 1.5 million customers worldwide.
We’re looking for a Senior Site Reliability Engineer who thrives at the crossroads of classic Linux and AWS infrastructure and
modern Site Reliability Engineering. This is a high-impact, hybrid role designed for someone who can manage cloud resources,
harden Kubernetes clusters, and shape a more reliable and developer-friendly platform.
We need you not just to maintain but to rethink and evolve our infrastructure, balancing hands-on operations with strategic
improvements that future-proof our growing AI product landscape.
You’ll take over key responsibilities from our current Infra Lead who is transitioning to a software-focused role, giving you
immediate ownership and space to shine.
You won’t be a cog in a massive SRE org. You’ll be the bridge between Infrastructure and Engineering, shaping how we scale
Kubernetes, how we approach platform reliability, and how developers ship fast without fear. You’ll get autonomy, ownership, and a
smart, humble team excited to learn with you.
Manychat is an Equal Opportunity Employer. We’re committed to building a diverse and inclusive team. We do not discriminate
against qualified employees or applicants because of race, color, religion, gender identity, sex, sexual preference, sexual
identity, pregnancy, national origin, ancestry, citizenship, age, marital status, physical disability, mental disability, medical
condition, military status, or any other characteristic protected by local law or ordinance.
This commitment is also reflected through our candidate experience. If you have individual needs that may require an accommodation
during the interview process, please indicate this in your application. We will do our best to provide assistance throughout your
interview process to ensure you’re set up for success.
With my application, I accept the Manychat Privacy Policy [https://manychat.com/legal/privacy].
WHO WE ARE 🌍 Creating content that resonates is great — turning that attention into growth is even better. That's what Manychat does. Our AI-powered automations help creators and brands engage with audiences across Instagram, Messenger, WhatsApp, and TikTok — at the right moment, with the right message, minus the manual work. We're 400+ people across three continents behind the leading platform for conversational growth, with AI skills shaping how we build, how we work, and how we grow. WHO WE’RE LOOKING FOR 🌟 Want to shape an AI platform's architecture from day one, instead of just maintaining what someone else built? Manychat runs AI features for businesses worldwide, and we're hiring an SRE to own the reliability, performance, and cost of that AI infrastructure, and to raise the bar for how our whole engineering org builds on LLMs. This isn't a classical SRE role with a bit of AI sprinkled on top. We need someone AI-native: you already understand how modern LLM systems behave in production, things like token throughput, provider rate limits, degraded model quality, and inference latency tails, and you treat all of that as a first-class reliability concern, not an afterthought. Why this role is worth your time: * The AI Platform is still young, so you'll shape its architecture, standards, and roadmap from the ground up. * The stakes are real. AI features sit right in the critical path of customer-facing automation, so your work actually matters to the business, not just to a dashboard. * You'll partner directly with the Head of Infrastructure, with real autonomy and visibility into how your decisions play out. Sound like the kind of ownership you've been looking for? WHAT YOU'LL DO 🚀 * Own reliability and performance of our AI infrastructure: AI Gateway, inference services, and integrations with Amazon Bedrock, Azure OpenAI, and other LLM providers. * Design and evolve the AI Gateway: routing, failover between providers, rate limiting, caching, and guardrails. * Build observability for AI systems: latency/throughput/error SLOs per model and provider, token-level metrics, quality and drift signals. * Drive cost optimization and FinOps for AI workloads: per-feature cost visibility, model right-sizing, caching strategies, provider mix. * Run capacity planning and incident response for inference services; write and improve runbooks and postmortems. * Scale AI expertise across the org: set standards, review designs, and coach teams shipping LLM-backed features. TO SHINE IN THIS ROLE 💥 You’ll need: * 5+ years in SRE / platform / infrastructure engineering, including production ownership at significant scale. * Hands-on experience operating LLM-backed systems in production: provider APIs (Bedrock, OpenAI, Anthropic, or similar), inference pipelines, self-hosted or managed model serving. * Deep cloud-native background: AWS, Kubernetes, Terraform/IaC, CI/CD. * Strong observability practice (Prometheus/Grafana, OpenTelemetry, or equivalent) and experience defining SLOs for non-deterministic systems. * Proven cost-optimization work: you can show where you cut cloud or inference spend and how you made cost visible. * Staff-level influence: you've set technical direction beyond your own team and brought others along It would be great if you have: * Experience building or operating an LLM gateway/proxy (e.g., LiteLLM, Kong AI Gateway, custom). * Experience with GPU workload optimization, quantization, or serving frameworks (vLLM, TGI, Triton). * Experience with eval pipelines and quality monitoring for LLM outputs. WHY THIS ROLE * Green-field ownership: the AI Platform is young; you'll shape its architecture, standards, and roadmap. * Real scale and real stakes: AI features sit in the critical path of customer-facing automation. * Direct partnership with the Head of Infrastructure; high autonomy and visibility. WHAT WE OFFER 🤗 We care deeply about your growth, well-being, and comfort: * 🌍 Hybrid onboarding to start work remotely and relocation support for you and your family. * 💙 Comprehensive health insurance for both you and your family. * 📚 Professional development budget for conference tickets, online courses, and other relevant resources to help you grow. * 🫶 Flexible benefits package: no one-size-fits-all perks here. You get a budget and you decide where it goes, from health and wellbeing to family and setting up your home office. * 🪴 Hybrid work and generous, flexible time off — planned with your team, not rationed by a rigid quota. * 🍽️ In-office perks, including free meals and snacks. * 🤝 Company-funded sport activities, annual offsites and team-building events. * 🤖 AI isn't a perk here — it's how we work. Claude, OpenAI, and more are on by default from day one, and we back teams in adopting whatever makes them faster. Manychat is an Equal Opportunity Employer. We’re committed to building a diverse and inclusive team. We do not discriminate against qualified employees or applicants because of race, color, religion, gender identity, sex, sexual preference, sexual identity, pregnancy, national origin, ancestry, citizenship, age, marital status, physical disability, mental disability, medical condition, military status, or any other characteristic protected by local law or ordinance. This commitment is also reflected through our candidate experience. If you have individual needs that may require an accommodation during the interview process, please indicate this in your application. We will do our best to provide assistance throughout your interview process to ensure you’re set up for success.
YOUR IMPACT We are looking for a Staff DataOps Engineer to join the Data and AI Platform team. Your mission will be to shape our data platform strategy and architecture, driving enterprise-scale solutions that accelerate machine learning initiatives, enable engineering excellence, and unlock business insights. You will work in a team at the heart of Doctolib's data-driven transformation, enabling innovation through robust, scalable data infrastructure that empowers engineers, AI teams, and business stakeholders across the organization. Working in the tech team at Doctolib means building innovative products and features to improve the daily lives of care teams and patients. WHAT YOU'LL BUILD Your responsibilities include but are not limited to: * Design and implement enterprise-scale data infrastructure strategies, conducting thorough impact and cost analysis for major technical decisions, and establishing architectural standards across the organization * Build and optimize complex, multi-region data pipelines handling petabyte-scale datasets, ensuring 99.9% reliability and implementing advanced monitoring and alerting systems * Lead cost analysis initiatives, identify optimization opportunities across our data stack, and implement solutions that reduce infrastructure spend while improving performance and reliability * Provide technical guidance to data engineers and cross-functional teams, conduct architecture reviews, and drive adoption of best practices in DataOps, security, and governance * Evaluate emerging technologies, conduct proof-of-concepts for new data tools and platforms, and lead the technical roadmap for data infrastructure modernization WHAT YOU'LL BRING Before you read on: if you don't have the exact profile described below, but you feel this job description matches your skill set, we still encourage you to apply. You'll be a great fit if you: * You have 7+ years of experience after graduation as a Staff Data Platform Engineer, Staff DataOps, Staff Site Reliability Engineer, or in a similar role, with a history of architecting and scaling robust data platforms * You have extensive experience with Google Cloud Platform and a command of Kubernetes & Terraform for automated deployments, and you are an authority on implementing network and IAM security best practices * You have deep technical proficiency in orchestrating data pipelines using Airflow or Dagster, deploying applications to the cloud, and leveraging modern data warehouses such as BigQuery * You are highly skilled in programming with Python, and have a solid understanding of software development principles * You are an excellent troubleshooter who excels at diagnosing and fixing data infrastructure and identifying performance bottlenecks, and a strong communicator who can articulate complex technical concepts to both technical and non-technical audiences It would be fantastic if you: * Have hands-on experience building and deploying APIs, with a strong preference for frameworks like FastAPI, and are skilled at designing APIs that provide reliable and efficient access to data * Have actively applied data governance principles to manage data quality, security, and compliance, and understand how to implement controls that protect sensitive information * Are experienced with CI/CD tools and methodologies for data-related projects, and can build automated pipelines that streamline development, testing, and deployment * Have practical experience with cloud cost optimization and FinOps principles, actively experimenting with cost-reduction strategies, and are comfortable analyzing infrastructure spending and optimizing resource allocation LIFE AT DOCTOLIB TECH * Our solutions are built on a single fully cloud-native platform that supports web and mobile app interfaces, multiple languages, and is adapted to country and healthcare specialty requirements. * Our stack is composed of Rails, TypeScript, Java, Python, Kotlin, Swift, and React Native. * We leverage AI ethically across our products to empower patients and health professionals. Discover our AI vision here. Want to learn more about our tech culture and environment? Visit the Doctolib Tech site. WHAT WE OFFER * Free comprehensive health insurance for you and your children * 25 days of paid vacation per year, plus up to 14 days of RTT * Free mental health and coaching services through our partner Moka.care * Work from abroad for up to 10 days per year thanks to our flexibility days policy * Lunch vouchers (Swile card) worth €8.50 per working day, with €4.50 covered by Doctolib * A subsidy from the work council to refund part of the membership to a sport club or a creative class * 50% reimbursement of your public transport subscription * Parent Care Program: receive one additional month of leave on top of the legal parental leave * For caregivers and workers with disabilities, a package including an adaptation of the remote policy, extra days off for medical reasons, and psychological support * Relocation support in case of international mobility * Access to the best AI tools for coding, development and dedicated training OUR INTERVIEW PROCESS * Recruiter Interview * Case Study * System Design Interview * Behavioral interview with Hiring Manager * At least one reference check We want your experience to be clear, respectful, and transparent. Learn more about our hiring process on our candidate experience page. JOB DETAILS * Permanent position * Tech stack: Google Cloud Platform, Kubernetes, Terraform, Dagster, BigQuery, dbt * Full-time * Paris * Start date: as soon as possible WE WELCOME EVERYONE At Doctolib, we are committed to improving access to healthcare for everyone. This translates into our recruitment process. We evaluate candidates based solely on qualifications and motivation, without any form of discrimination. The more diverse ideas are heard, the more our product will truly improve healthcare for all. You are welcome to apply to Doctolib, regardless of your gender, religion, age, sexual orientation, ethnicity, or disability. To ensure equal opportunities, we invite you to exclude personal information (e.g., pictures, age) from your applications. If you require any accommodation, please let us know for support during the hiring process. Join us in building the healthcare we all dream of! YOUR DATA PRIVACY All information provided is processed by Doctolib for application management. For data processing details, click here: France. Please contact hr.dataprivacy(at)doctolib.com for inquiries or to exercise your rights.
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Workforce Identity Cloud Okta Workforce Identity Cloud (WIC) provides easy, secure access for your workforce so you can focus on other strategic priorities, such as reducing costs and doing more for your customers. If you like to be challenged and have a passion for solving large-scale automation, testing, and tuning problems, we would love to hear from you. The ideal candidate is someone who exemplifies the ethics of, “If you have to do something more than once, automate it” and who can rapidly self-educate on new concepts and tools. Position Overview: The Staff Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services. This position focuses on architecting and managing reliable, scalable, and secure Kubernetes-based platforms on AWS, ensuring high availability and performance while optimising costs and automation. The ideal candidate will have hands-on experience with AWS infrastructure, Kubernetes platform creation, Helm charts, Karpenter scaling, and Istio service mesh. Key Responsibilities: * Kubernetes Platform Creation: Design, implement, and maintain highly available, scalable, and fault-tolerant Kubernetes platforms. Ensure clusters are optimised for production workloads, providing high resilience and operational efficiency. * AWS Infrastructure Management: Build, manage, and optimise AWS cloud infrastructure, including EKS, ECS, S3, VPCS, RDS, IAM, and more. Implement best practices for cost management, scaling, and security within AWS. * Helm Management: Utilise Helm to automate and streamline the deployment of applications and services to Kubernetes clusters. Create, maintain, and manage Helm charts for production-ready deployments. * Karpenter Implementation: Implement and manage Karpenter to dynamically scale Kubernetes clusters in response to workload demands. * Istio Service Mesh Management: Configure and manage Istio to provide service-to-service communication, security, and observability within the Kubernetes clusters. Enable fine-grained traffic management, service discovery, and policy enforcement. * Platform Automation & Scaling: Automate the deployment, scaling, and management of infrastructure and applications. Work with CI/CD pipelines to ensure a seamless flow from development to production with minimal downtime. * Incident Management & Troubleshooting: Respond to incidents, troubleshoot, and resolve system issues related to performance, availability, and security in a timely and effective manner. * Security & Compliance: Design and implement secure cloud infrastructure with appropriate access controls, network security, and compliance frameworks. * Documentation & Knowledge Sharing: Create and maintain detailed documentation for Kubernetes platform setup, operational procedures, and best practices. Promote knowledge sharing across teams. Required Qualifications: * 5+ years of experience with Kubernetes/ K8s, Helm,Karpenter,Istio; * 8+ years of Experience with infrastructure-as-code tools like Terraform, Chef or Ansible * 8+ years of Experience with serverless computing (AWS Lambda, API Gateway) and microservices architecture. * Proven experience with AWS (EKS, ECS, RDS, S3, CloudFormation, IAM, etc.) and solid understanding of cloud-native architectures. * Strong expertise in Kubernetes platform creation, management, and optimisation (e.g., setting up highly available clusters, networking, and storage). * Hands-on experience with Helm for Kubernetes application deployment and management. * Practical experience with Karpenter for dynamic scaling of Kubernetes clusters and optimising resource usage. * Expertise in managing and securing Istio for service mesh, including traffic management, security, and observability features. * Proficiency in CI/CD pipelines and automation tools (e.g., Jenkins, GitLab, CircleCI, Terraform, Spinnaker, Ansible). * Strong scripting and automation skills in Python or Go for infrastructure management and platform automation. * Experience with monitoring, logging, and alerting tools such as Prometheus, Grafana, CloudWatch, and ELK Stack. Preferred Qualifications: * Experience with multi-region cloud environments. * Understanding of security best practices for cloud platforms and Kubernetes (e.g., role-based access control (RBAC), encryption, and compliance frameworks). * Familiarity with Docker and containerization principles. * Bachelor’s degree in Computer Science, Engineering, or related field (or equivalent professional experience). * Certifications (Preferred): CKA (Certified Kubernetes Administrator), CKAD (Certified Kubernetes Application Developer), or AWS Certified DevOps Engineer are highly desirable. #LI_Hybrid P25021_3418720 The Okta Experience * Supporting Your Well-Being * Driving Social Impact * Developing Talent and Fostering Connection + Community We are intentional about connection. Our global community, spanning over 20 offices worldwide, is united by a drive to innovate. Your journey begins with an immersive, in-person onboarding experience designed to accelerate your impact and connect you to our mission and team from day one. Okta is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, ancestry, marital status, age, physical or mental disability, or status as a protected veteran. We also consider for employment qualified applicants with arrest and convictions records, consistent with applicable laws. If reasonable accommodation is needed to complete any part of the job application, interview process, or onboarding please use this Form to request an accommodation. Notice for New York City Applicants & Employees: Okta may use Automated Employment Decision Tools (AEDT), as defined by New York City Local Law 144, that use artificial intelligence, machine learning, or other automated processes to assist in our recruitment and hiring process. In accordance with NYC Local Law 144, if you are an applicant or employee residing in New York City, please click here to view our full NYC AEDT Notice.