
Coupang · Bengaluru
Company Introduction We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did I ever live without Coupang?...
Company Introduction
We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did I ever live without
Coupang?” Born out of an obsession to make shopping, eating, and living easier than ever, we’re collectively disrupting the
multi-billion-dollar e-commerce industry from the ground up. We are one of the fastest-growing e-commerce companies that
established an unparalleled reputation for being a dominant and reliable force in South Korean commerce.
We are proud to have the best of both worlds — a startup culture with the resources of a large global public company. This fuels
us to continue our growth and launch new services at the speed we have been since our inception. We are all entrepreneurs
surrounded by opportunities to drive new initiatives and innovations. At our core, we are bold and ambitious people that like to
get our hands dirty and make a hands-on impact. At Coupang, you will see yourself, your colleagues, your team, and the company
grow every day.
Our mission to build the future of commerce is real. We push the boundaries of what’s possible to solve problems and break
traditional trade-offs. Join Coupang now to create an epic experience in this always-on, high-tech, and hyper-connected world.
Role Overview
We are seeking a Sr Staff System Engineer, GPU Fleet for our Coupang Intelligent Cloud (CIC) team, to serve as the senior
technical owner for our hyperscale GPU compute infrastructure. In this role, you will define fleet architecture, drive reliability
and automation at scale, and lead the operation and evolution of GPU systems supporting large‑scale AI training and inference
workloads. This is a hands‑on, staff‑level individual contributor role with broad technical ownership, high operational impact,
and significant cross‑functional influence across hardware, infrastructure, and datacenter operations.
CIC builds the infrastructure for abundant intelligence. We partner with leading AI labs, governments, and enterprises to deliver
hyperscale GPU compute with high reliability, performance, and efficiency. Our infrastructure supports some of the most demanding
AI training and inference workloads in production today.
We operate with urgency, deep ownership, and a strong bias toward execution. Reliability, operational excellence, and rigorous
systems engineering are core to our business.
What You Will Do
As a Sr Staff System Engineer, GPU Fleet, you will be the senior technical owner for CIC’s large‑scale GPU compute infrastructure.
This is a hands‑on senior individual contributor role with fleet‑level responsibility and broad cross‑functional influence.
You will define the technical direction for how GPU fleets are architected, operated, automated, and evolved across multiple
generations of hardware. Your work will directly affect fleet reliability, operating efficiency, scalability, and customer
success.
This role does not involve people management, but it carries principal‑level scope, autonomy, and decision‑making authority across
infrastructure, hardware, and operations.
Fleet Architecture & Technical Ownership
OS configuration, drivers, networking, and observability.
redesigns.
ambiguity.
Reliability, Availability & Performance
datacenters.
processes.
Automation & Systems Engineering
Operational Leadership
improvements.
Technical Influence & Mentorship
Basic Qualifications
or datacenter operations, operating production environments with strict uptime and performance requirements.
networking, kernel behavior, and system performance analysis.
drivers, OS configuration, and user‑space services.
diagnostics, remediation, and fleet operations.
manual intervention.
Preferred Qualifications
workloads.
Ethernet.
reliability transformations.
platform designs.
systems.
urgency, and technical leadership.
Type of work model
Hybrid
Details to consider
treatment for employment in accordance with applicable laws.
Privacy Notice
Company Introduction We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did we ever live without Coupang?” Born out of an obsession to make shopping, eating, and living easier than ever, we’re collectively disrupting the multi-billion-dollar e-commerce industry from the ground up. We are one of the fastest-growing e-commerce companies that established an unparalleled reputation for being a dominant and reliable force in South Korean commerce. We are proud to have the best of both worlds — a startup culture with the resources of a large global public company. This fuels us to continue our growth and launch new services at the speed we have been since our inception. We are all entrepreneurs surrounded by opportunities to drive new initiatives and innovations. At our core, we are bold and ambitious people that like to get our hands dirty and make a hands-on impact. At Coupang, you will see yourself, your colleagues, your team, and the company grow every day. Our mission to build the future of commerce is real. We push the boundaries of what’s possible to solve problems and break traditional tradeoffs. Join Coupang now to create an epic experience in this always-on, high-tech, and hyper-connected world. Role Overview As a Senior Staff Data Centre Observability and Site Reliability Engineer, you will design, build, and operate scalable observability and reliability solutions for large-scale datacenter infrastructure. This role focuses on developing high-performance monitoring and telemetry platforms, ensuring system reliability, and driving operational excellence through automation, performance optimization, and SRE best practices. The ideal candidate will work across the full service lifecycle—design, deployment, and continuous improvement—while collaborating with cross-functional teams to enhance visibility, resilience, and efficiency of critical systems. What You Will Do OBSERVABILITY AND MONITORING * Design, implement, and maintain observability solutions for datacenter infrastructure, including monitoring, logging, alerting, and telemetry systems. * Develop, deploy, and operate large-scale observability and telemetry platforms with a focus on real-time monitoring, high performance, and scalability. * Own and contribute to the full lifecycle of observability services—from design and development to deployment and ongoing optimization. * Build and enhance monitoring systems to ensure high availability, reliability, and performance of infrastructure. * Create and manage dashboards, alerts, and reports to provide clear visibility into system health, performance, and capacity trends. SITE RELIABILITY ENGINEERING (SRE) * Apply SRE principles and best practices to improve reliability, scalability, and operational efficiency of datacenter services. * Develop and maintain automation for infrastructure provisioning, monitoring, and system management. * Lead root cause analysis (RCA) and post-incident reviews, driving corrective actions to prevent recurrence and improve system resilience. PERFORMANCE OPTIMIZATION * Analyze system and application performance across the datacenter infrastructure to identify bottlenecks and improvement areas. * Implement optimization strategies to enhance performance, efficiency, and resource utilization. COLLABORATION * Partner with cross-functional engineering teams to understand observability and reliability requirements and deliver effective solutions. * Collaborate with hardware and software vendors to evaluate, integrate, and optimize new technologies within the ecosystem. SECURITY AND COMPLIANCE * Ensure observability and reliability solutions adhere to organizational security policies and industry standards. * Implement and maintain appropriate security controls to safeguard infrastructure, systems, and data. TROUBLESHOOTING AND SUPPORT * Provide hands-on support for observability and reliability issues, including debugging complex hardware and software problems. * Develop and maintain documentation, including troubleshooting guides and operational best practices, to support efficient issue resolution. CONTINUOUS IMPROVEMENT * Stay current with emerging trends, tools, and technologies in observability and SRE, and incorporate them into the platform. * Continuously enhance the scalability, reliability, and operational efficiency of datacenter services through proactive improvements. Basic Qualifications * Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field. * 12+ years of progressive software engineering experience, with a heavy emphasis on distributed systems, cloud-native architectures, or platform operations. * Proven experience in managing and optimizing large-scale datacenter environments * Strong proficiency in Go or Python, with a deep understanding of networked systems and performance optimization. * Expert-level knowledge of Kubernetes internals (scheduling, controllers) and containerization ecosystems. * Proven experience with load balancing, service mesh, and request routing at scale. * Proficiency in observability tools and technologies (e.g., Prometheus, Grafana, ELK Stack). * Experience with SRE practices and tools (e.g., Kubernetes, Docker, Terraform). * Familiarity with cloud platforms (AWS, Azure, GCP) and their observability and reliability services Preferred Qualifications * Prior experience building infrastructure specifically for LLM inference or large-scale training clusters. * Familiarity with inference, including mixed precision, kernel tuning, or custom hardware accelerators. * Experience managing hybrid-cloud or multi-AZ deployments across AWS, Azure, or GCP. * Experience operating in regulated environments with strict security and compliance requirements Type of work model * Hybrid Our Hybrid work model: Coupang hybrid work model is designed to enable a culture of collaboration that acts a catalyst to enrich the experience of employees. Employees are required to work at least 3 days in the office per week, with the flexibility to work from home 2 days a week, depending on the role requirement. Some businesses may require more time in office due to nature of work. Details to consider * Those eligible for employment protection (recipients of veteran’s benefits, the disabled, etc.) may receive preferential treatment for employment in accordance with applicable laws. Privacy Notice * Your personal information will be collected and managed by Coupang as stated in the Application Privacy Notice located below. https://privacy.coupang.com/en/land/jobs/
By bringing together next-gen technology and the finest live data available, Genius Sports is enabling a new era of sports for fans worldwide, delivering experiences that are more immersive, interactive and personalized than ever before. Learn more at geniussports.com. The Role: Senior Site Reliability Engineer – Edge Computing We are seeking a Senior Site Reliability Engineer to be part of the Edge Computing team in the Data&AI group. As we deliver real-time player tracking, sport analytics and broadcast augmentation to more customers worldwide, we are looking to scale from hundreds of sport venues to thousands. Specifically, you can expect to: * Design and code end-to-end processes enabling operational staff to autonomously prepare, install and monitor all Linux servers, networking devices and cameras installed in 300+ sport venues across the world * Design and code end-to-end processes enabling developers to autonomously deploy and monitor our player tracking and augmentation applications * Take ownership of long-term technical efforts and articulate design choices to technical and non-technical people * Collaborate closely with teammates to solve problems, share knowledge and provide actionable feedback * Participate in an on-call rotation that emphasizes eliminating repeating escalations * Visit our wonderful Lausanne Jordils office 4 times per week, with flexible hours Minimum Qualifications * Swiss/EU/EFTA citizen or residency permit in Switzerland * 5+ years experience in SRE with Linux * Strong understanding of the entire Linux server stack: OS boot and installation, systemd, networking, container deployment, logging, metrics & monitoring, out-of-band management, etc... * Strong experience designing robust automation processes for a large inventory of on-premises servers * Proficient with Python programming and Bash scripting * Ability to communicate efficiently and articulate concepts based on the audience Preferred Qualifications * Experience with remote fleet management without easy physical access * Experience designing large-scale automation processes for network routers and switches * Strong understanding of OSI network layers 2 and 3, ability to assess network conditions at customer sites: explain packet loss or fragmentation, cabling or NIC defects, bandwidth evaluation locally and to cloud via intercontinental transit * Strong experience deploying container applications on premises with Kubernetes * Experience with AWS EC2, S3, VPC, IAM * Experience with Nvidia GPU driver installation and monitoring Our Stack: * Languages and frameworks: Python, Rust, Bash * Servers: Ubuntu/Linux, bootc, Nvidia GPUs * Networking: Mikrotik, FS * Cloud: Tailscale, Netbox, Docker, Kubernetes, Prometheus, Grafana, AWS Cloud Our Work Environment and What You Will Benefit From * Become part of the world of elite sports, build systems that support data analytics and augmentation for the best leagues * Enjoy an innovative and dynamic environment that encourages self-development * Develop automation solutions that directly improve the daily work of dozens of operational staff and developers The base salary range for this role is CHF 140,000 - CHF 155,000. There are a number of factors which affect what the specific pay for this role is, including seniority and relevant educational and working experience. We enjoy an ‘office-first’ culture and maximize opportunities to collaborate, connect and learn together. Our hybrid working models differ depending on your role and location. Occasional travel may be required. As well as a competitive salary and range of benefits, we’re committed to supporting employee wellbeing and helping you grow your skills, experience and career. Learn more about how rewarding life at Genius can be at Reward | Genius Sports. One team, being brave, driving change We strive to create an inclusive working environment, where everyone feels a sense of belonging and the ability to make a difference. Learn more about our values and culture at Culture | Genius Sports. Let us know when you apply if you need any assistance during the recruiting process due to a disability.
By bringing together next-gen technology and the finest live data available, Genius Sports is enabling a new era of sports for fans worldwide, delivering experiences that are more immersive, interactive and personalized than ever before. Learn more at geniussports.com. The Role: Senior Software Engineer – Edge Computing We are seeking a Senior Software Engineer to be part of the Edge Computing team in the Data&AI group. As we deliver real-time player tracking, sport analytics and broadcast augmentation to more customers worldwide, we are looking to scale from hundreds of sport venues to thousands. Specifically, you can expect to: * Design and code end-to-end processes enabling operational staff to autonomously prepare, install and monitor all Linux servers, networking devices and cameras installed in 300+ sport venues across the world * Design and code end-to-end processes enabling developers to autonomously deploy and monitor our player tracking and augmentation applications * Take ownership of long-term technical efforts and articulate design choices to technical and non-technical people * Collaborate closely with teammates to solve problems, share knowledge and provide actionable feedback * Participate in an on-call rotation that emphasizes eliminating repeating escalations * Visit our wonderful Lausanne Jordils office 4 times per week, with flexible hours Minimum Qualifications * Swiss/EU/EFTA citizen or residency permit in Switzerland * 5+ years experience in SRE with Linux * Strong understanding of the entire Linux server stack: OS boot and installation, systemd, networking, container deployment, logging, metrics & monitoring, out-of-band management, etc... * Strong experience designing robust automation processes for a large inventory of on-premises servers * Proficient with Python programming and Bash scripting * Ability to communicate efficiently and articulate concepts based on the audience Preferred Qualifications * Experience with remote fleet management without easy physical access * Experience designing large-scale automation processes for network routers and switches * Strong understanding of OSI network layers 2 and 3, ability to assess network conditions at customer sites: explain packet loss or fragmentation, cabling or NIC defects, bandwidth evaluation locally and to cloud via intercontinental transit * Strong experience deploying container applications on premises with Kubernetes * Experience with AWS EC2, S3, VPC, IAM * Experience with Nvidia GPU driver installation and monitoring Our Stack: * Languages and frameworks: Python, Rust, Bash * Servers: Ubuntu/Linux, bootc, Nvidia GPUs * Networking: Mikrotik, FS * Cloud: Tailscale, Netbox, Docker, Kubernetes, Prometheus, Grafana, AWS Cloud Our Work Environment and What You Will Benefit From * Become part of the world of elite sports, build systems that support data analytics and augmentation for the best leagues * Enjoy an innovative and dynamic environment that encourages self-development * Develop automation solutions that directly improve the daily work of dozens of operational staff and developers The base salary range for this role is CHF 140,000 - CHF 155,000. There are a number of factors which affect what the specific pay for this role is, including seniority and relevant educational and working experience. We enjoy an ‘office-first’ culture and maximize opportunities to collaborate, connect and learn together. Our hybrid working models differ depending on your role and location. Occasional travel may be required. As well as a competitive salary and range of benefits, we’re committed to supporting employee wellbeing and helping you grow your skills, experience and career. Learn more about how rewarding life at Genius can be at Reward | Genius Sports. One team, being brave, driving change We strive to create an inclusive working environment, where everyone feels a sense of belonging and the ability to make a difference. Learn more about our values and culture at Culture | Genius Sports. Let us know when you apply if you need any assistance during the recruiting process due to a disability.