
Coupang · Bengaluru
Company Introduction We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did we ever live without Coupang...
Company Introduction
We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did we ever live without
Coupang?” Born out of an obsession to make shopping, eating, and living easier than ever, we’re collectively disrupting the
multi-billion-dollar e-commerce industry from the ground up. We are one of the fastest-growing e-commerce companies that
established an unparalleled reputation for being a dominant and reliable force in South Korean commerce.
We are proud to have the best of both worlds — a startup culture with the resources of a large global public company. This fuels
us to continue our growth and launch new services at the speed we have been since our inception. We are all entrepreneurs
surrounded by opportunities to drive new initiatives and innovations. At our core, we are bold and ambitious people that like to
get our hands dirty and make a hands-on impact. At Coupang, you will see yourself, your colleagues, your team, and the company
grow every day.
Our mission to build the future of commerce is real. We push the boundaries of what’s possible to solve problems and break
traditional tradeoffs. Join Coupang now to create an epic experience in this always-on, high-tech, and hyper-connected world.
Role Overview
As a Senior Staff Data Centre Observability and Site Reliability Engineer, you will design, build, and operate scalable
observability and reliability solutions for large-scale datacenter infrastructure. This role focuses on developing
high-performance monitoring and telemetry platforms, ensuring system reliability, and driving operational excellence through
automation, performance optimization, and SRE best practices. The ideal candidate will work across the full service
lifecycle—design, deployment, and continuous improvement—while collaborating with cross-functional teams to enhance visibility,
resilience, and efficiency of critical systems.
What You Will Do
and telemetry systems.
performance, and scalability.
optimization.
trends.
resilience.
solutions.
resolution.
improvements.
Basic Qualifications
architectures, or platform operations.
Preferred Qualifications
Type of work model
Our Hybrid work model: Coupang hybrid work model is designed to enable a culture of collaboration that acts a catalyst to
enrich the experience of employees. Employees are required to work at least 3 days in the office per week, with the flexibility
to work from home 2 days a week, depending on the role requirement. Some businesses may require more time in office due to
nature of work.
Details to consider
treatment for employment in accordance with applicable laws.
Privacy Notice
Company Introduction We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did I ever live without Coupang?” Born out of an obsession to make shopping, eating, and living easier than ever, we’re collectively disrupting the multi-billion-dollar e-commerce industry from the ground up. We are one of the fastest-growing e-commerce companies that established an unparalleled reputation for being a dominant and reliable force in South Korean commerce. We are proud to have the best of both worlds — a startup culture with the resources of a large global public company. This fuels us to continue our growth and launch new services at the speed we have been since our inception. We are all entrepreneurs surrounded by opportunities to drive new initiatives and innovations. At our core, we are bold and ambitious people that like to get our hands dirty and make a hands-on impact. At Coupang, you will see yourself, your colleagues, your team, and the company grow every day. Our mission to build the future of commerce is real. We push the boundaries of what’s possible to solve problems and break traditional trade-offs. Join Coupang now to create an epic experience in this always-on, high-tech, and hyper-connected world. Role Overview We are seeking a Sr Staff System Engineer, GPU Fleet for our Coupang Intelligent Cloud (CIC) team, to serve as the senior technical owner for our hyperscale GPU compute infrastructure. In this role, you will define fleet architecture, drive reliability and automation at scale, and lead the operation and evolution of GPU systems supporting large‑scale AI training and inference workloads. This is a hands‑on, staff‑level individual contributor role with broad technical ownership, high operational impact, and significant cross‑functional influence across hardware, infrastructure, and datacenter operations. CIC builds the infrastructure for abundant intelligence. We partner with leading AI labs, governments, and enterprises to deliver hyperscale GPU compute with high reliability, performance, and efficiency. Our infrastructure supports some of the most demanding AI training and inference workloads in production today. We operate with urgency, deep ownership, and a strong bias toward execution. Reliability, operational excellence, and rigorous systems engineering are core to our business. What You Will Do As a Sr Staff System Engineer, GPU Fleet, you will be the senior technical owner for CIC’s large‑scale GPU compute infrastructure. This is a hands‑on senior individual contributor role with fleet‑level responsibility and broad cross‑functional influence. You will define the technical direction for how GPU fleets are architected, operated, automated, and evolved across multiple generations of hardware. Your work will directly affect fleet reliability, operating efficiency, scalability, and customer success. This role does not involve people management, but it carries principal‑level scope, autonomy, and decision‑making authority across infrastructure, hardware, and operations. Key Responsibilities: Fleet Architecture & Technical Ownership * Own the end‑to‑end technical architecture of hyperscale GPU fleets, including hardware platform selection, firmware strategy, OS configuration, drivers, networking, and observability. * Define and enforce technical standards and best practices for fleet reliability, availability, performance, and operability. * Lead major fleet‑wide initiatives such as new GPU platform bring‑ups, multi‑generation hardware transitions, and architectural redesigns. * Evaluate trade‑offs across cost, performance, reliability, and time‑to‑deploy, and make technically sound decisions under ambiguity. Reliability, Availability & Performance * Set and drive fleet‑level reliability, availability, and performance objectives. * Lead root‑cause analysis and resolution of complex, systemic failures affecting large portions of the fleet or multiple datacenters. * Identify recurring failure patterns and drive long‑term fixes spanning hardware, software, automation, and operational processes. * Work directly with hardware vendors and partners to resolve platform‑level issues and influence future hardware designs. Automation & Systems Engineering * Design and build large‑scale automation systems for: * GPU fleet provisioning and lifecycle management * GPU health validation, diagnostics, and certification * Automated remediation, recovery, and replacement workflows * Eliminate manual operational toil through durable, well‑designed tooling that scales with fleet growth. * Ensure all fleet systems are observable, testable, and resilient under failure conditions. Operational Leadership * Act as a senior escalation point for critical production incidents impacting GPU availability or customer workloads. * Participate in on‑call rotations with a strong emphasis on preventing future incidents, not just responding to them. * Lead high‑severity post‑incident reviews and ensure learnings are translated into concrete engineering and process improvements. Technical Influence & Mentorship * Provide technical mentorship and guidance to system and infrastructure engineers across the organization. * Serve as a trusted technical partner to platform engineering, networking, datacenter operations, and leadership teams. * Influence CIC’s long‑term infrastructure roadmap through strong technical judgment and data‑driven recommendations. Basic Qualifications * 12+ Years of overall experience with at least 8+ years of experience in Linux systems engineering, infrastructure engineering, or datacenter operations, operating production environments with strict uptime and performance requirements. * Deep, hands‑on expertise in Linux system internals, including process scheduling, memory management, filesystem behavior, networking, kernel behavior, and system performance analysis. * Demonstrated experience operating hardware‑intensive infrastructure in production, including bare‑metal servers at scale. * Proven ability to debug complex issues across multiple system layers, including hardware components, firmware/BIOS, kernel drivers, OS configuration, and user‑space services. * Extensive experience writing production‑grade automation using Python and Bash for provisioning, configuration management, diagnostics, remediation, and fleet operations. * Strong understanding of how to design systems that are observable, resilient, and safe under failure, rather than reliant on manual intervention. Preferred Qualifications * Direct experience operating large‑scale GPU fleets supporting AI/ML training and/or inference workloads in production. * Familiarity with modern GPU platforms and ecosystems, including GPU drivers, CUDA, NCCL, and high‑performance compute workloads. * Experience with high‑speed interconnects and datacenter networking, such as NVLink, InfiniBand, RDMA, and high‑throughput Ethernet. * Prior ownership of fleet‑wide or platform‑wide initiatives, such as new hardware bring‑ups, major architectural changes, or reliability transformations. * Experience partnering directly with hardware vendors or manufacturers to troubleshoot systemic issues or influence future platform designs. * Strong intuition for failure modes at scale, including cascading failures, correlated faults, and second‑order effects across systems. * History of acting as a technical authority or escalation point for ambiguous, high‑impact production problems. * Ability to mentor engineers through design reviews, technical problem solving, and modelling strong operational ownership. * Experience participating in on‑call rotations and responding to high‑severity production incidents with clear ownership, urgency, and technical leadership. * Strong written and verbal communication skills, including clear post‑incident reviews and technical documentation. Type of work model Hybrid Details to consider * Those eligible for employment protection (recipients of veteran’s benefits, the disabled, etc.) may receive preferential treatment for employment in accordance with applicable laws. Privacy Notice * Your personal information will be collected and managed by Coupang as stated in the Application Privacy Notice located below. https://privacy.coupang.com/en/land/jobs/
Wrike is the most powerful work management platform. Built for teams and organizations looking to collaborate, create, and exceed every day, Wrike brings everyone and all work into a single place to remove complexity, increase productivity, and free people up to focus on their most purposeful work. Our vision: A world where everyone is free to focus on their most purposeful work, together. ABOUT THE ROLE: Wrike’s Backend Reliability (BRE) team is the backbone of our backend infrastructure and the guardian of our uptime. Our mission is to achieve and sustain 99.99% availability while building the tools, components, and safety nets that the entire engineering organization relies on. As a Senior / Staff Backend Engineer on this team, you won’t just close tickets - you’ll architect core reliability solutions that shape how Wrike scales, performs, and recovers from failure. YOUR IMPACT: * Design, build, and maintain critical reliability components such as HTTP rate limiters, internal DB schema migration tools, circuit breakers, and distributed Redis-based caching. * Troubleshoot complex production issues, optimize PostgreSQL usage, and ensure our distributed systems remain performant and stable under high load. * Lead preliminary investigations during severe production incidents: identify likely root causes, assess impact, and propose mitigation options. The long-term fixes are then implemented by the owning team, based on your findings. * Create scalable, reusable tools and frameworks that help other engineering teams build more resilient services. * Leverage AI-powered tools and coding agents to accelerate development, analyze architectures, and automate repetitive or error-prone tasks. * Influence reliability best practices across engineering by sharing knowledge, reviewing designs, and setting high technical standards. YOUR QUALIFICATIONS: * Strong expertise with Java/JVM, building scalable, high-performance backend systems; open to leveraging other languages when appropriate. * Solid understanding of distributed systems concepts, including high availability, CAP theorem, and fault tolerance. * Deep experience with relational databases (PostgreSQL) and key–value / non-relational storages (Redis). * Practical experience with containerization and cloud-native environments, including Docker and Kubernetes. * Hands-on experience with message brokers such as RabbitMQ or Kafka. * Ability to work independently with minimal supervision, using critical thinking to question assumptions and validate your own decisions. * Strong written and spoken English skills suitable for collaborating in an international engineering environment. STANDOUT QUALITIES: * Background in infrastructure engineering or Site Reliability Engineering (SRE), including infrastructure-as-code practices. * Experience leading technical initiatives, driving cross-team projects, and mentoring other engineers while remaining an individual contributor. * Familiarity with observability and monitoring stacks (e.g., Graylog, Zabbix, Grafana) and/or data analytics tools such as BigQuery. * A strong interest in how complex systems fail and a track record of designing them to recover gracefully. TEAM DYNAMICS: You will join the Backend Reliability (BRE) team, a small, highly specialized, senior group focused solely on Wrike’s reliability. The team operates as an internal “reliability task force,” partnering closely with product and platform engineering teams across the company. You’ll collaborate with other senior engineers who value autonomy, deep technical discussions, and rigorous engineering practices. The culture is ownership-driven: you are trusted to manage your time, make architectural decisions, and drive initiatives to completion. OUR WORK STYLE: The BRE team works on core backend and infrastructure services that support millions of users. We operate in a collaborative environment that values clear communication, thoughtful design, and fast feedback loops. * Tech focus areas include: Java/JVM-based services, PostgreSQL, Redis, Docker, Kubernetes, RabbitMQ/Kafka, and observability/monitoring tools. * We encourage the use of modern AI-based tooling and coding agents as part of daily development and troubleshooting workflows. * Work is organized with an emphasis on impact and reliability goals rather than ticket volume, giving you room for deep work and long-term improvements. * Hybrid work mode (Prague, Czech Republic / Nicosia, Cyprus), with flexibility to balance focused individual work and collaborative sessions. WHAT’S NEXT? * Interview with a Recruiter * Technical interview * System Design Interview * Cultural interview Your recruitment buddy will be Aleksandar Chernev [https://www.linkedin.com/in/aleksandar-chernev-479785214/], Senior Technical Recruiter. #LI-AC1 WHO IS WRIKE AND OUR CULTURE We’re a team of innovators and creators who solve the complex work problems of today and tomorrow. Hybrid work mode Wrike is our people, not a place. With 1,000+ employees collaborating across nearly every time zone, we support talent through 10 global hubs — Australia, Costa Rica, Cyprus, Czechia, Estonia, France, India, Ireland, Japan, and the United States — offering flexible ways of working that include remote work, hybrid environments, and co-working spaces across many locations. While flexibility looks different across teams and regions, employees located near certain hubs — particularly in Prague (CZ), Nicosia (CY), Bangalore (IN), and Rennes (FR) — are generally expected to collaborate in person around 2–3 days per week, balancing the flexibility of distributed work with opportunities for in-person collaboration and connection. OUR PERSONA 💡 Smart: We love what we do, and we’re great at it because this is our domain. Our combined knowledge in this space is unmatched. 💚 Dedicated: We get up every day focused on helping our customers win. We’re committed to helping our teammates win, too! 🤗 Approachable: We're friendly, easy to get along with, considerate, and helpful. OUR CULTURE AND VALUES 🤩 Customer-Focused We care about our customers. We understand the customer journey, experience, and value derived from Wrike. Decision-making and action-taking are done with the customer in mind. 🤝 Collaborative We work as one and win together, each bringing unique strengths that contribute to diversity of thought for better outcomes. Leveraging our own work management platform, we foster an environment of creative collaboration and shared achievement. 🎨 Creative We strive to succeed through continuous innovation. It’s our pursuit of novel concepts that helped us create a market category. We continue to cultivate a workplace that fosters creative thinking as a means of transcending conventional boundaries and empowers us to break new ground to deliver extraordinary work management solutions. 💪 Committed We believe in ownership at all levels of the organization, by owning workflows from start to finish. Each member of our team is an integral part of this commitment, establishing work as a platform for personal growth and transformation, as well as collective success and growth. Check out our LinkedIn Life Page [https://www.linkedin.com/company/wrike/life/3fd588bf-73e8-47be-9b0a-1933d404ea88/], Company culture page [https://www.wrike.com/wrike-company-culture/], Instagram [https://www.instagram.com/wriketeam], Wrike Engineering Team [https://www.wrike.com/wrike-engineering/], Medium [https://medium.com/wriketechclub], Meetup.com [https://www.meetup.com/WrikeTechClub/?_cookie-check=wtgfN9ARYGPSGd3e], Youtube [https://www.youtube.com/c/wriketechclub] for a feel for what life is like at Wrike. Check us out on Glassdoor. [https://www.glassdoor.com/pc-app/static/img/partnerCenter/badges/eng_CHECK_US_273x90.png]https://www.glassdoor.com/Overview/Working-at-Wrike-EI_IE420969.11,16.htm
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Okta Identity Governance (OIG) organization is looking for a Principal Engineer to join our team — OIG is Okta’s Identity Governance and Administration solution that is directly responsible for one of the most critical and visible workflows in enterprise identity: how employees request, approve, and gain access to the resources they need. Opportunity As a Principal Engineer on the OIG team, you will be involved into development, design, and maintenance of our product to serve enterprise customers at scale. You will involve the architecture and evolution of our product features — spanning the across different access governance personas — and work collaboratively across engineering, product, and design to deliver features that are secure, scalable, and delightful to use. This is a role for someone who thinks in systems, leads with craft, and is energized by the challenge of building enterprise-grade software that millions of users depend on. This is a rare opportunity to join a team where your technical decisions will directly shape how thousands of enterprises manage access governance at scale, where the technical challenges are genuinely hard, and where the work you ship will be seen and felt by real users every day. Job Duties And Responsibilities * Take end-to-end ownership of feature areas — from technical design through implementation, testing, deployment, and post-launch monitoring * Design, build, and ship high-quality, production-ready features across the OIG system, with a focus on correctness, reliability, and performance * Quickly deliver high-quality bug fixes and handle customer-reported issues * Write clean, well-tested, maintainable code and hold yourself and teammates to a high bar for engineering quality * Partner with our Product Development, QA, and Site Reliability Engineering teams for scoping the development and deployment work Required Knowledge, Skills, And Abilities * 8+ years of software engineering experience, with at least 3+ years in a senior or staff-level technical leadership role * Proven track record of architecting and delivering large-scale, distributed systems in production environments serving enterprise customers * Deep expertise in microservices, event-driven architectures, APIs, and data modeling at scale * Strong proficiency in one or more of: Java, Kotlin, Go, or equivalent JVM/cloud-native languages used in Okta's backend stack * Demonstrated ability to influence without authority — driving technical alignment across multiple teams and stakeholders * Experience designing systems that meet enterprise reliability, scalability, and security requirements (think: multi-tenancy, high availability, audit logging, RBAC/ABAC) * Excellent written and verbal communication skills, with the ability to make complex technical topics accessible to both engineers and non-engineers PREFERRED IF YOU HAVE EXPERIENCE IN ANY OF THE FOLLOWING! * Experience building identity, access management, or governance (IAM/IGA) products or platforms * Familiarity with workflow engines or approval/routing systems (e.g., finite state machines, BPM-style systems) * Experience with enterprise SaaS platforms serving large, complex customer organizations * Hands-on experience with cloud infrastructure (AWS, GCP, or Azure) and modern observability tooling Education * B.S. Computer Science or equivalent #LI-Remote #P25268_3467543 Below is the annual base salary range for candidates located in California (excluding San Francisco Bay Area), Colorado, Illinois, New York and Washington. Your actual base salary will depend on factors such as your skills, qualifications, experience, and work location. In addition, Okta offers equity (where applicable), bonus, and benefits, including health, dental and vision insurance, 401(k), flexible spending account, and paid leave (including PTO and parental leave) in accordance with our applicable plans and policies. To learn more about our Total Rewards program please visit: https://rewards.okta.com/us. The annual base salary range for this position for candidates located in California (excluding San Francisco Bay Area), Colorado, Illinois, New York, and Washington is between: $184,000—$276,000 USD The Okta Experience * Supporting Your Well-Being * Driving Social Impact * Developing Talent and Fostering Connection + Community We are intentional about connection. Our global community, spanning over 20 offices worldwide, is united by a drive to innovate. Your journey begins with an immersive, in-person onboarding experience designed to accelerate your impact and connect you to our mission and team from day one. Okta is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, ancestry, marital status, age, physical or mental disability, or status as a protected veteran. We also consider for employment qualified applicants with arrest and convictions records, consistent with applicable laws. If reasonable accommodation is needed to complete any part of the job application, interview process, or onboarding please use this Form to request an accommodation. Notice for New York City Applicants & Employees: Okta may use Automated Employment Decision Tools (AEDT), as defined by New York City Local Law 144, that use artificial intelligence, machine learning, or other automated processes to assist in our recruitment and hiring process. In accordance with NYC Local Law 144, if you are an applicant or employee residing in New York City, please click here to view our full NYC AEDT Notice.