
Douro Labs · South America - Remote
We are looking for a Data Systems Operations Lead to architect and run mission-critical operations for Pyth's Price Feeds, our real-time, distributed market dat...
We are looking for a Data Systems Operations Lead to architect and run mission-critical operations for Pyth's Price Feeds, our
real-time, distributed market data infrastructure. You'll ensure our feeds operate flawlessly 24/7 while translating operational
challenges into strategic improvements that measurably reduce risk and improve service quality.
Our goal is to hire someone who can help us uplift the service management aspects of Pyth's Price Feeds. You'll investigate and
lead implementation of automation around price feeds, own the runtime requirements of Pyth Pro (monitoring, provisioning, risk
metrics), and when incidents occur, lead data forensics to quantify impact and drive solutions that measurably reduce recurrence.
This role is for a high-level operator who understands distributed systems at scale and can identify automation opportunities
while managing relationships with institutional users. You'll make high-stakes operational decisions, architect systematic
solutions, and own projects end-to-end.
Location: Americas / Remote-first
Language: English (Fluent)
About Our Team and Your Role
We are a well-rounded team: half of us are tech whizzes, while the other half excel in building partnerships with institutional
clients, prime brokers, and the global financial infrastructure community. Communication is key to our network-driven approach.
Remote Work: Our team is spread across the globe, from the US and South America to Europe and Asia. We are a digital-first
organization where remote work is the norm.
Startup-level Speed: We thrive in the dynamic DeFi space and love adaptable problem solvers who are eager to meet the evolving
needs of the market. We're building an Amsterdam hub with strong technical and business development talent.
The Mission: You will be architecting partnerships that define the future of market data. Your focus will be Pyth Pro, our
institutional-grade data solution designed for sophisticated buyers who require the highest fidelity, lowest latency, and most
reliable feeds directly from primary sources.
Own Pyth Pro Operations & Service Management
Investigate and Lead Automation Implementation
Lead Incident Response & Data Forensics
Bridge Technical & Business Operations
Nice to Have
You're a problem-solver who owns complex challenges from zero to one. You've earned trust in previous roles, which is why people
listen to your solutions. You understand that operations serve customers and users - you're driven by reliability that matters,
not perfection for its own sake. You're technically deep enough to debug distributed systems, but mature enough to know that
systems design beats heroic firefighting.
That sounds like you?
Apply directly !
SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: As a Facilities Operations Technician at SpaceXAI, you'll dive in, operating and maintaining the critical facility systems that power our supercomputing data centers. Working hands-on with MEP systems, you’ll tackle maintenance, operations, and repairs, ensuring peak performance and reliability. Expect to troubleshoot complex issues, streamline processes, and collaborate with teams to keep our infrastructure running smoothly, directly supporting SpaceXAI's mission to accelerate human scientific discovery through AI. This role includes the following positions, each with increasing responsibility: * Facilities Operations Technician: Independently operate and maintain critical mechanical, electrical, and plumbing (MEP) systems in supercomputing data centers, ensuring high reliability and performance through hands-on maintenance, troubleshooting, and process optimization. * Sr Facilities Operations Technician: Perform advanced maintenance and troubleshooting of critical data center systems like HVAC, power, and environmental controls, while mentoring junior technicians to ensure operational reliability. * Facilities Operations Lead: Oversee the maintenance and optimization of critical data center infrastructure, including power, cooling, and environmental systems, while leading a team of technicians to ensure operational reliability and efficiency. RESPONSIBILITIES: * Responsible for MEP Infrastructure site safety and environmental compliance, ensuring the sites align with all local and regional health and safety requirements. * Provide operational expertise for Mechanical, Electrical, and Plumbing (MEP) systems at the datacenter facility. The systems and equipment can include pumps, piping, chillers, cooling towers, CRAH units, wiring, lighting systems, generators, motors, water supply, UPS and switchgear, etc. * Routinely inspect all areas to ensure performance measures are being maintained and proactively self-report the problems of facilities. * Experience with writing and executing MOPs, SOPs, EOPs. * Perform daily inspections of critical systems. * Identify and maintain spare parts inventory to minimize downtime. * Coordinate and supervise scheduled system shutdowns. * Develop and implement energy monitoring metrics. * Support commissioning efforts by witnessing equipment startup, assisting with functional testing, and monitoring construction progress. BASIC QUALIFICATIONS: * Bachelor’s Degree in Mechanical or Electrical or 2+ years of hands-on operations/facilities work experience in lieu of degree (data center experience preferred). * Knowledge of mechanical or electrical engineering principles and systems (both would be a huge plus). * Experience and knowledge of all critical systems documentation including Plans, Procedures, and Operations Manuals. * Experience and knowledge of Building and Electrical Power Monitoring Systems and Operations. * Strong troubleshooting skills. * Experience with PLC-based control systems and/or with Building Management Systems (Client). * Excellent written and verbal communication and technical writing skills in English. * Physical Requirements: * Ability to lift up to 50 lbs. * Ability to utilize hand tools and power tools as needed. * Work is often performed in tight quarters, and physical dexterity is necessary to perform job functions. * Ability to lift up to 35 lbs. unassisted. * Comfortable working at elevated heights (up to 100 feet) with appropriate safety gear. * Comfortable working in extreme outdoor environments—heat, cold, rain, etc. * Comfortable working in an environment requiring exposure to fumes, odors, and noise. * Available to work flexible shifts providing 24x7 coverage as needed, including evenings, weekends, and holidays. * Position is subject to pre-employment drug and random drug and alcohol testing. * Position is subject to pre-employment and annual post-employment background checks. * Willing to travel as needed (up to 10%). * Valid driver’s license. SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice [https://x.ai/legal/recruitment-privacy-notice].
Mercury is building a whole stack of financial tools for startups. We work hard to create dashboards with thought and simplicity. You can check out our demo dashboard at www.demo.mercury.com [https://www.demo.mercury.com]. Behind every product Mercury ships is a fast-growing company: dozens of teams across engineering, product, design, data, and business functions, and hundreds of projects that require alignment across compliance, legal, partnerships, customer support, finance, and leadership. AI Ops builds the systems that keep that organization moving, helping teams move quickly without losing shared context. In this role, you'll own Mercury's internal knowledge infrastructure: the systems and standards that make company information accurate, discoverable, and useful. You'll build and maintain a trusted context layer—a structured, living record of what teams own, are building, and know—and ensure it stays current automatically. That context layer powers leadership reporting, planning, operational reviews, and the internal AI agents employees use every day. You'll define how information is organized, validated, and maintained across systems like Linear and Mercury's internal platforms so they function as a single source of truth. Working closely with Engineering, who own the underlying infrastructure, you'll design the knowledge layer above it: the taxonomies, schemas, validation workflows, and automations that make company knowledge reliable for both people and AI systems. This role sits at the intersection of systems operations, knowledge architecture, and product thinking. Success isn't measured by collecting more information, but by creating a high-signal knowledge system that helps employees find answers quickly, enables leaders to make decisions from shared context, and gives AI systems the foundation they need to operate effectively. Mercury aims to make banking* feel secure, reliable, thoughtful, and perhaps even magical. Your job is to make the company's internal knowledge systems just as dependable. *Mercury is a fintech company, not an FDIC-insured bank. Banking services provided through Choice Financial Group and Column N.A., Members FDIC. YOU WILL: * Own Mercury's knowledge infrastructure: the trusted context layer that captures what every team owns, is building, and knows, along with the information architecture, taxonomy, and governance that keep it accurate, current, and useful. * Build the knowledge layer on top of Mercury's AI infrastructure by partnering with AI Engineering to design the schemas, automations, and validation workflows that allow people and AI agents to reliably retrieve and act on company knowledge. * Own the reporting layer that turns shared context into operational insight, including leadership reporting, roadmap views, planning dashboards, and the reporting that powers company operating cadences. * Drive company-wide adoption of standardized systems and practices by partnering across Engineering, Product, Design, Data, Compliance, Legal, Finance, Partnerships, Customer Support, and other teams to replace fragmented documentation with trusted, structured sources of truth. * Continuously improve how Mercury captures, organizes, and uses knowledge by identifying operational friction, building better workflows, and ensuring employees have the tooling and enablement they need to effectively work with AI. YOU SHOULD: * Have 5–8 years of experience in program or product operations, technical program management, product management, data, or similar roles where you drove company-wide systems or operational improvements. * Think like a systems designer and knowledge architect, able to turn messy, distributed information into simple, scalable structures that people and AI systems can easily understand and trust. * Be comfortable working with technical systems, including APIs, data models, analytics, and tools like Linear, GitHub, Metabase, and modern AI platforms, even if you aren't building the underlying infrastructure yourself. * Have hands-on experience using AI to create leverage through workflows, automations, agents, or other practical applications, along with a solid understanding of how LLMs retrieve and consume information. * Influence organizations through strong judgment, clear communication, and thoughtful execution, thriving in ambiguous environments where the right systems have to be invented rather than inherited. The total rewards package at Mercury includes base salary, equity (stock options), and benefits. Our salary and equity ranges are highly competitive within the SaaS and fintech industry and are updated regularly using the most reliable compensation survey data for our industry. New hire offers are made based on a candidate's experience, expertise, geographic location, and internal pay equity relative to peers. Our target new hire base salary ranges for this role are the following: * US employees in New York City, Los Angeles, Seattle, or the San Francisco Bay Area: $163,000 - $203,800 * US employees outside of the New York City, Los Angeles, Seattle, or the San Francisco Bay Area: $146,700 - $183,400 * Canadian employees (any location): CAD $154,100 - $192,600 *Mercury is a fintech company, not an FDIC-insured bank [https://mercury.com/how-mercury-works]. Banking services provided through Choice Financial Group and Column N.A., Members FDIC. Mercury values diversity & belonging and is proud to be an Equal Employment Opportunity employer. All individuals seeking employment at Mercury are considered without regard to race, color, religion, national origin, age, sex, marital status, ancestry, physical or mental disability, veteran status, gender identity, sexual orientation, or any other legally protected characteristic. We are committed to providing reasonable accommodations throughout the recruitment process for applicants with disabilities or special needs. If you need assistance, or an accommodation, please let your recruiter know once you are contacted about a role. We use Covey as part of our hiring and / or promotional process for jobs in NYC and certain features may qualify it as an AEDT. As part of the evaluation process we provide Covey with job requirements and candidate submitted applications. [Please see the independent bias audit report covering our use of Covey for more information.] [https://getcovey.com/nyc-local-law-144] #LI-JD1
WHO WE ARE ABOUT STRIPE Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. ABOUT THE TEAM The Stream Compute team at Stripe builds and operates the infrastructure, tooling, and systems behind our Flink-powered stream processing systems. We're at the heart of several core asynchronous workflows, operating at significant scale and handling vast amounts of sensitive financial data. Our work powers intricate processes involving various critical financial operations and real-time analytics. We run globally distributed systems with high reliability and performance to meet Stripe's scaling, availability, and product needs, and we continually reduce operational toil by investing in automation and self-service tooling for upgrades, maintenance, and day-to-day operations. The team is distributed between Seattle, Toronto, and remote locations. What makes our team truly exciting is our commitment to our users: we ensure no event is dropped, state integrity is preserved, and exactly-once processing is supported as a first-class feature. Working at the intersection of real-time data processing and fintech innovation, we continuously push the boundaries of what's possible. Our focus on innovation, user experience, reliability, and compliance drives increased ROI and operational excellence, making us a crucial part of Stripe's success. WHAT YOU'LL DO You'll help define and deliver the next generation of Stripe's Flink-first stream compute infrastructure—driving innovation to meet extremely high availability targets at global scale. Partnering with infrastructure engineers, adjacent platform teams, and the product orgs that depend on Flink every day, you'll set a long-term technical direction that scales with Stripe's growth while enabling reliable, efficient operations for years to come. You'll work on the hardest problems in operating Flink in production—state management, exactly-once processing, performance isolation, and automated recovery—so teams across Stripe can confidently build stateful stream processing applications on top of it. RESPONSIBILITIES * Design, build, and operate stream compute infrastructure with Apache Flink at the center, alongside technologies like Kafka, Temporal, and AWS services * Partner with product and platform teams across Stripe to understand requirements, unblock Flink adoption, and improve how stream processing infrastructure is used end-to-end * Define and implement operational best practices (e.g., shuffle sharding, cellular architecture, load shedding, automated state recovery) to improve resilience and reliability at scale * Drive fleet-level automation and standardization ("pets" to "cattle") through self-service workflows, safer rollouts, and self-healing systems that reduce manual operations * Lead initiatives that raise the bar on Flink availability and state durability (e.g., multi-region strategies, disaster recovery readiness, operational readiness reviews, incident learning) * Evaluate and productionize Flink ecosystem capabilities (e.g., SQL, connectors, state backends) to improve developer experience and scalability without compromising reliability * Work closely with the open-source community to identify opportunities for adopting new open-source features and contributing back to OSS WHO YOU ARE We're looking for someone who meets the minimum requirements to be considered for the role. If you meet these requirements, you are encouraged to apply. The preferred qualifications are a bonus, not a requirement. MINIMUM REQUIREMENTS * This is a Staff-level role—that typically means 10+ years of experience building, operating, and evolving large-scale production systems. * Experience as a technical lead for team(s) working on distributed systems, including scaling them in fast-moving environments * Hands-on experience with big data technologies such as Flink, Spark, Kafka, Pulsar, or Pinot * Experience developing, maintaining, and debugging distributed systems built with open-source tools * Experience building and scaling infrastructure as a product • Strong software engineering skills and a passion for big data distributed systems * Ability to write high-quality code (in programming languages like Go, Java, Scala, etc.) * Comfortable operating with high autonomy and ownership * Growth mindset and a willingness to learn quickly, explore ambiguous problem spaces, and dive deep when needed * Strong written and verbal communication skills, including the ability to produce clear technical documentation PREFERRED QUALIFICATIONS * Experience operating streaming infrastructure as a platform (e.g., Flink clusters, Kafka, Pulsar) for internal customers at scale * Deep hands-on experience authoring, optimizing, and operating real-time processing frameworks such as Flink, Spark Streaming, Storm, or Kafka Streams in production * Experience building or operating control planes for managing large-scale infrastructure * Open-source contributions to data processing or big data systems (Hadoop, Spark, Celeborn, Flink, etc.)