
Scaleway · Paris
OUR STORY: 🇪🇺 Join Scaleway and shape the sovereign cloud of tomorrow ! Since 1999, we have been designing secure, sustainable infrastructures aimed at supp...
🇪🇺 Join Scaleway and shape the sovereign cloud of tomorrow !
Since 1999, we have been designing secure, sustainable infrastructures aimed at supporting the most ambitious companies.
Historically known for our dedicated servers (Dedibox), we made a strategic shift to cloud computing in 2015. Staying true to our principles of simplicity, flexibility, and technical excellence, we have become one of the leading players in Europe in the sector.
With the rise of artificial intelligence, we have strengthened our commitment, supported by the Iliad Group, which is investing €3 billion to develop a serious, sovereign AI alternative to American and Asian giants.
Every day, thanks to our fast-growing portfolio of cloud and AI products (bare metal, containerization, serverless, AI, etc.), Scaleway proudly serves thousands of customer across the private and public sector, from corporations like France Télévisions or Hachette Livre, to fast-growing startups like Photoroom and Biolevate, to institutions like the City of Copenhagen.
📍 Our offices are located in Paris, Lille, Toulouse, Rennes, Rouen, Bordeaux and Lyon.
Our growth is driving us to strengthen our SRE team to support and scale our production environments.
Your mission will be to build and maintain reliable, observable, and secure infrastructure in order to ensure optimal service availability for our customers around the world.
We work in a collaborative and international environment where the diversity of Scalers, combined with a spirit of sharing, helps bring new projects to life every day, advancing our ambitions together.
You will join a newly formed team dedicated to building and operating Scaleway’s future AI infrastructure. As part of this group, you will design, maintain, and scale core systems and observability tools, partner with product teams, and ensure the reliability and performance of AI services across Scaleway.
Hybrid work: We offer up to 3 days of remote work per week.
Offices: Our offices are spacious, dynamic workspaces with bold design, conveniently located near public transport. Most of our offices feature outdoor spaces (terraces) and bike parking facilities.
Dining: Our chef provides a healthy meal service at the headquarters, and breakfast is available across all our sites year-round. Scalers working from regional sites enjoy a Swile card for lunches.
Well-being commitments: Whether it’s access to a gym, daycare places, or discounted services for caring services, Scaleway is committed to supporting Scalers in maintaining a balanced life.
International environment: With dozens of nationalities, Scaleway offers a stimulating environment where English is as widely spoken as French.
Career & Mobility: Our managers value internal mobility, and opportunities to transition to other entities within the Iliad Group are accessible to all Scalers.
🚀 Why join the Scaleway adventure ?
✔ A rich and diverse product offering: Scaleway offers over 100 public cloud products in IaaS, PaaS, and AI.
✔ A cutting-edge technical environment: Scaleway provides modern infrastructures, including high-performance bare metal servers, to tackle exciting technical challenges.
✔ Commitment to responsible cloud: Scaleway is dedicated to a more responsible cloud, with data centers powered solely by renewable energy since 2017, minimizing our ecological footprint and holding top-level certification.
About Man Group Man Group is a global alternative investment management firm focused on pursuing outperformance for sophisticated clients via our Systematic, Discretionary and Solutions offerings. Powered by talent and advanced technology, our single and multi-manager investment strategies are underpinned by deep research and span public and private markets, across all major asset classes, with a significant focus on alternatives. Man Group takes a partnership approach to working with clients, establishing deep connections and creating tailored solutions to meet their investment goals and those of the millions of retirees and savers they represent. Headquartered in London, we manage $228.7 billion* and operate across multiple offices globally. Man Group plc is listed on the London Stock Exchange under the ticker EMG.LN and is a constituent of the FTSE 250 Index. Further information can be found at www.man.com At Man Group, we respect your privacy and we are committed to protecting and safeguarding your Personal Data. We have developed policies and processes which are designed to provide for the security and integrity of your Personal Data. We are committed to Processing your Personal Data fairly and lawfully, and being open and transparent about such Processing. For further information on how we process your data, please see the privacy notice for applicants here * As at 31 March 2026 The Role Join our high-performing Site Reliability Engineering (SRE) team and play a pivotal role in ensuring the reliability, scalability, and performance of the technology powering Man Group’s hedge funds. You’ll have the autonomy, tools, and support to innovate and shape the future of our platform. This is an opportunity to work on cutting-edge projects, gain mentorship from senior leaders, and develop a deep understanding of both technology and the business. As an SRE, you’ll take ownership of service reliability and deliver solutions that make a real impact. Your initial focus will include leveraging AI to accelerate incident diagnosis and resolution, improving observability, capacity planning, and automation. Over time, you’ll work across our entire infrastructure stack, operating at scale and driving continuous improvement. Role Responsibilities * Ensure reliability and performance of critical systems across global infrastructure through proactive monitoring and rapid incident response. * Design and implement observability solutions using tools like Prometheus, Grafana, ELK, and Loki to provide deep insights into system health. * Automate operational tasks and build self-service capabilities to eliminate toil and improve efficiency. * Develop and maintain SLIs, SLOs, and error budgets to guide reliability improvements and inform engineering priorities. * Participate in incident response efforts, blameless post-mortems, and implement preventive measures to reduce recurrence. * Collaborate with development teams to improve system design, deployment practices, and operational excellence. * Operate at scale, managing petabyte-level storage, large CPU/GPU deployments, and high-throughput distributed systems. * Contribute to capacity planning and performance tuning, ensuring systems meet business demands. * Manage multiple ELK clusters hosting hundreds of terabytes of logs, telemetry, and APM data. Key competencies Required * Strong understanding of SRE principles, including SLIs, SLOs, error budgets, and reliability best practices. * Hands-on experience with observability and monitoring tools (Prometheus, Grafana, ELK, Loki, or similar). * Proficiency with automation tools (Ansible, Terraform) and scripting/programming languages (Python, Go, PowerShell). * Strong troubleshooting and debugging skills across distributed systems, with the ability to diagnose complex production issues under pressure. * Experience with incident management, on-call rotations, and post-incident reviews. * Familiarity with Kubernetes and container orchestration. * A proactive mindset and ability to take ownership of reliability initiatives. Advantageous * Experience with CI/CD pipelines and source control workflows (Git, Jenkins, TeamCity). * Administration of Linux and Windows systems and exposure to cloud technologies (AWS/Azure). * Understanding of networking concepts, load balancing, and distributed architectures. * Knowledge of AI/LLM concepts (context windows, prompt tuning, MCP servers). * Interest in FinOps principles, desire to understand the true cost of our decisions. * Excellent communication and collaboration skills. Benefits * Modern office located in the OfficeX campus with easy access to transport and amenities. * Hybrid working model * Competitive compensation package * 25 days holiday allowance * Premium Health insurance * Employee Assistance program * Referral Bonus * Additional days off for long service and volunteering * Multisport card * Opportunities for professional development including internal tech talks * Conference attendance, and engagement with the open-source community. Inclusion, Work-Life Balance and Benefits at Man Group You'll thrive in our working environment that champions equality of opportunity. Your unique perspective will contribute to our success, joining a workplace where inclusion is fundamental and deeply embedded in our culture and values. Through our external and internal initiatives, partnerships and programmes, you'll find opportunities to grow, develop your talents, and help foster an inclusive environment for all across our firm and industry. Learn more at www.man.com/diversity. You'll have opportunities to make a difference through our charitable and global initiatives, while advancing your career through professional development, and with flexible working arrangements available too. Like all our people, you'll receive two annual 'Mankind' days of paid leave for community volunteering. Our comprehensive benefits package includes competitive holiday entitlements, pension/401k, life and long-term disability coverage, group sick pay, enhanced parental leave and long-service leave. Depending on your location, you may also enjoy additional benefits such as private medical coverage, discounted gym membership options and pet insurance. Equal Employment Opportunity Policy Man Group provides equal employment opportunities to all applicants and all employees without regard to race, color, creed, national origin, ancestry, religion, disability, sex, gender identity and expression, marital status, sexual orientation, military or veteran status, age or any other legally protected category or status in accordance with applicable federal, state and local laws. Man Group is a Disability Confident Committed employer; if you require help or information on reasonable adjustments as you apply for roles with us, please contact TalentAcquisition@man.com.
We are looking for a Senior Site Reliability Engineer to join our Site Reliability Engineering (SRE) team. In this role, you'll drive the reliability, scalability, and performance of our platform, ensuring our systems remain stable as we grow. We value innovation and are seeking someone eager to bring fresh ideas – especially around building automation that reduces manual effort and improving distributed systems resilience. This isn't a top-down organization; our engineers are the ones who flag technical challenges and design the solutions. You will collaborate closely with Platform Engineering, Security, AI Platform, and Product teams to design durable systems and make data-driven operational decisions. WHAT YOU'LL DO * Collaborate with Engineering, Platform, and Security teams to embed SRE best practices early in system design. * Lead advancements in observability, monitoring, alerting, and incident-response workflows. * Analyze platform performance to contribute to cost-optimization, performance tuning, and resilience planning. * Build infrastructure and automation tooling that improves platform reliability and enhances deployment safety. * Diagnose and resolve complex production issues across distributed systems, and drive open post-incident reviews so failures translate into durable improvements. * Strengthen system consistency and author clear, concise documentation for runbooks and operational processes. WHO YOU ARE * 4+ years of experience in SRE, DevOps, platform engineering, or similar production-facing roles. * Strong problem-solving and debugging skills in distributed systems to maintain higher platform stability. * Eager to share operational guidelines, champion SRE practices across teams, and openly discuss what we can learn from system failures. * Excellent communication skills (English is our default language) with a genuine, collaborative approach to working across diverse engineering teams. * Strong hands-on experience with cloud environments (AWS, GCP, or similar) and proficiency with infrastructure-as-code and CI/CD pipelines. * Familiarity with Kubernetes (or container orchestration), event-driven architectures, or supporting ML/AI workloads and GPU infrastructure. WHAT SUCCESS LOOKS LIKE: Within 3 Months: * Fully onboarded into the Rossum ecosystem, gaining a deep understanding of our infrastructure, observability stack, and SRE processes while building relationships across the team. * Gaining a deep understanding of our synergy with Coupa and our shared roadmap. * Initial Impact Goal: Improve a small reliability issue or add value to an existing automation or monitoring area. Within 6 Months: * Independently managing key responsibilities, owning recurring reliability tasks, and identifying areas for strategic improvement. * Actively participating in the alignment of processes within the new Coupa organizational structure. * Operational KPI: Implement measurable enhancements to alert quality, CI/CD reliability, or service health metrics. Within 12 Months: * Recognized as a subject matter expert within the team, navigating the global Coupa ecosystem. * Successfully contributing to Rossum's mission at a massive scale using new global resources. * Long-Term Strategic Goal: Lead a major reliability or infrastructure initiative, providing technical recommendations to guide our long-term reliability strategy. WHY JOIN US? At Rossum, we're on a mission to free the world from boring manual data entry. Our AI platform helps companies save millions of hours, allowing professionals to focus on creative, impactful work. In an exciting move for our future, we have joined forces with Coupa, the world's leading unified platform for Business Spend Management. By combining Rossum's cutting-edge document AI with Coupa's global ecosystem, we are uniquely positioned to redefine how businesses operate at a massive scale. You can read more about this exciting milestone and our shared vision in the official announcement here. What sets us apart? * Cutting-edge AI technology reshaping how businesses operate globally. * A collaborative, supportive environment where autonomy thrives. * Opportunities to grow in a fast-scaling company. * A culture that values diversity, empathy, and genuine connection. As part of the Coupa family, you'll enjoy the agility of a fast-moving, innovation-focused team with the stability and reach of a global market leader. For you, this means an even greater opportunity to make an impact, access new global markets, and grow your career within a collaborative culture that values autonomy, diversity, and genuine connection. Together, we're not just automating data—we're giving time back to the world's professionals. WHAT WE OFFER (BENEFITS) * Work-Life Harmony: 5 weeks of vacation, 5 sick/personal days, and a Birthday Day Off to celebrate you. For new parents, we offer an extra 18 weeks of fully paid Maternity leave and 8 weeks of fully paid Paternity leave. * Flexibility: Hybrid work model with flexible hours—find the balance that works best for you and work how you work best. * Tech & Setup: High-end laptop (MacBook) and tech setup. We also support you with personal development, including language lessons (English & Czech, on all levels). * Community & Well-being: Wellness Days (company-wide days to unplug, reset, and recharge), MultSport cards, regular team offsites, and meetups, in a friendly, ambitious team environment. Ready to make an impact in your next role? Apply now!
WHO WE ARE 🌍 Creating content that resonates is great — turning that attention into growth is even better. That's what Manychat does. Our AI-powered automations help creators and brands engage with audiences across Instagram, Messenger, WhatsApp, and TikTok — at the right moment, with the right message, minus the manual work. We're 400+ people across three continents behind the leading platform for conversational growth, with AI skills shaping how we build, how we work, and how we grow. WHO WE’RE LOOKING FOR 🌟 Want to shape an AI platform's architecture from day one, instead of just maintaining what someone else built? Manychat runs AI features for businesses worldwide, and we're hiring an SRE to own the reliability, performance, and cost of that AI infrastructure, and to raise the bar for how our whole engineering org builds on LLMs. This isn't a classical SRE role with a bit of AI sprinkled on top. We need someone AI-native: you already understand how modern LLM systems behave in production, things like token throughput, provider rate limits, degraded model quality, and inference latency tails, and you treat all of that as a first-class reliability concern, not an afterthought. Why this role is worth your time: * The AI Platform is still young, so you'll shape its architecture, standards, and roadmap from the ground up. * The stakes are real. AI features sit right in the critical path of customer-facing automation, so your work actually matters to the business, not just to a dashboard. * You'll partner directly with the Head of Infrastructure, with real autonomy and visibility into how your decisions play out. Sound like the kind of ownership you've been looking for? WHAT YOU'LL DO 🚀 * Own reliability and performance of our AI infrastructure: AI Gateway, inference services, and integrations with Amazon Bedrock, Azure OpenAI, and other LLM providers. * Design and evolve the AI Gateway: routing, failover between providers, rate limiting, caching, and guardrails. * Build observability for AI systems: latency/throughput/error SLOs per model and provider, token-level metrics, quality and drift signals. * Drive cost optimization and FinOps for AI workloads: per-feature cost visibility, model right-sizing, caching strategies, provider mix. * Run capacity planning and incident response for inference services; write and improve runbooks and postmortems. * Scale AI expertise across the org: set standards, review designs, and coach teams shipping LLM-backed features. TO SHINE IN THIS ROLE 💥 You’ll need: * 5+ years in SRE / platform / infrastructure engineering, including production ownership at significant scale. * Hands-on experience operating LLM-backed systems in production: provider APIs (Bedrock, OpenAI, Anthropic, or similar), inference pipelines, self-hosted or managed model serving. * Deep cloud-native background: AWS, Kubernetes, Terraform/IaC, CI/CD. * Strong observability practice (Prometheus/Grafana, OpenTelemetry, or equivalent) and experience defining SLOs for non-deterministic systems. * Proven cost-optimization work: you can show where you cut cloud or inference spend and how you made cost visible. * Staff-level influence: you've set technical direction beyond your own team and brought others along It would be great if you have: * Experience building or operating an LLM gateway/proxy (e.g., LiteLLM, Kong AI Gateway, custom). * Experience with GPU workload optimization, quantization, or serving frameworks (vLLM, TGI, Triton). * Experience with eval pipelines and quality monitoring for LLM outputs. WHY THIS ROLE * Green-field ownership: the AI Platform is young; you'll shape its architecture, standards, and roadmap. * Real scale and real stakes: AI features sit in the critical path of customer-facing automation. * Direct partnership with the Head of Infrastructure; high autonomy and visibility. WHAT WE OFFER 🤗 We care deeply about your growth, well-being, and comfort: * 🌍 Hybrid onboarding to start work remotely and relocation support for you and your family. * 💙 Comprehensive health insurance for both you and your family. * 📚 Professional development budget for conference tickets, online courses, and other relevant resources to help you grow. * 🫶 Flexible benefits package: no one-size-fits-all perks here. You get a budget and you decide where it goes, from health and wellbeing to family and setting up your home office. * 🪴 Hybrid work and generous, flexible time off — planned with your team, not rationed by a rigid quota. * 🍽️ In-office perks, including free meals and snacks. * 🤝 Company-funded sport activities, annual offsites and team-building events. * 🤖 AI isn't a perk here — it's how we work. Claude, OpenAI, and more are on by default from day one, and we back teams in adopting whatever makes them faster. Manychat is an Equal Opportunity Employer. We’re committed to building a diverse and inclusive team. We do not discriminate against qualified employees or applicants because of race, color, religion, gender identity, sex, sexual preference, sexual identity, pregnancy, national origin, ancestry, citizenship, age, marital status, physical disability, mental disability, medical condition, military status, or any other characteristic protected by local law or ordinance. This commitment is also reflected through our candidate experience. If you have individual needs that may require an accommodation during the interview process, please indicate this in your application. We will do our best to provide assistance throughout your interview process to ensure you’re set up for success.