
Farfetch · Porto
Farfetch is a leading global marketplace for the luxury fashion industry. The Farfetch Marketplace connects customers in over 190 countries and territories with...
Farfetch is a leading global marketplace for the luxury fashion industry. The Farfetch Marketplace connects customers in over 190 countries and territories with items from more than 50 countries and over 1,400 of the world’s best brands, boutiques, and department stores, delivering a truly unique shopping experience and access to the most extensive selection of luxury on a global marketplace.
We're on a mission to build end-to-end products and technology that powers the an incredible e-commerce experience for luxury customers everywhere, understanding the motivations and needs of our customers and partners, to designing and testing hypotheses, to creating industry-leading experiences for luxury customers.
Our office is near Porto, in the north of Portugal, and is located in a vibrant business hub. It offers a dynamic and welcoming environment where our employees can connect and network with a large community of tech professionals.
As a Senior Platform Engineer in the Tech Platform team, you will play a critical role in building the foundation of our entire technology ecosystem. Your mission is to engineer a seamless, scalable, and reliable platform that empowers engineering teams to deliver software faster and with higher quality.
In this role, you will be a key technical contributor dedicated to evolving our cloud-native infrastructure. You will focus on driving the adoption of Self-Service capabilities and Infrastructure as Code to reduce cognitive load for our developers, while maintaining a high bar for security, resilience, and operational excellence.
THE WOW As an Observability DevOps Engineer specializing in observability platforms, you will play a pivotal role in enhancing the visibility, reliability, and performance of our systems and applications. Leveraging your expertise in modern technologies and protocols, you will be responsible for implementing, and maintaining robust observability solutions to monitor, analyse, and troubleshoot our distributed systems. This position requires passion about leveraging observability platforms to drive operational excellence and thrive in a dynamic, collaborative environment. It also requires strong technical skills in AWS, Terraform, Git workflows, and a deep understanding of DevOps principles. Our Product Development organisation is truly Global with cross functional teams spanning 6 Tech Hubs – Malta, Budapest, Stockholm, Tallinn, Kyiv and Athens. With around 1,000 professionals, the technology organisation is spear-headed by our CTO-CPO with all our talented Area Teams working together. A TASTER OF WHAT YOU WILL BE INVOLVED WITH: * Design, deploy, and maintain observability platforms and tools such as Prometheus, Grafana LGTM+ stack, OTel and others to provide comprehensive insights into system behaviour and performance. * Collaborate with development, operations, and other cross-functional teams to integrate observability solutions into the CI/CD pipeline and automate monitoring and alerting processes. * Develop custom monitoring solutions and instrumentation libraries to capture relevant metrics, logs, traces, and events across microservices architectures. * Configure and optimize telemetry collection, storage, and visualization components to ensure scalability, reliability, and cost-effectiveness. * Implement anomaly detection algorithms and predictive analytics to proactively identify and mitigate potential issues before they impact users. * Conduct thorough root cause analysis of incidents and performance bottlenecks, leveraging observability data to drive continuous improvement initiatives. * Stay abreast of emerging trends, best practices, and industry standards in observability, and assess their applicability to our environment. * Provide mentorship and technical guidance to junior team members, fostering a culture of knowledge sharing and collaboration. * You will be expected to form part of a 24/7 on call roster to support your engineering team during out of office incidents calls * Drive automation initiatives using Terraform, Git workflows, and other DevOps tools to streamline deployment processes and improve operational efficiency. WHAT WE ARE LOOKING FOR Hands on experience with most of the following technologies: * * AWS * Kubernetes * Helm * Ansible * Proficiency in scripting languages such as Python, Bash, or Go for automation and tooling development. * Hands-on experience with containerization and orchestration technologies (e.g., Docker, Kubernetes) in cloud-native environments. * Strong understanding of distributed systems architecture, networking concepts, and security principles. * Familiarity with modern observability protocols and standards, including OpenTelemetry for unified observability. * Expertise in configuring and managing monitoring solutions like Prometheus, Grafana LGTM+ stack, and APM tools. * Experience with infrastructure as code (IaC) tools such as Terraform or Ansible for provisioning and configuration management. * Knowledge of eBPF (extended Berkeley Packet Filter) for deep observability into the Linux kernel and user-space applications. WHAT WE OFFER Much like riding a rollercoaster, sometimes life at Betsson can be lightning fast with twists and turns but always FUN! Then again, what else would you expect from a business 75% millennial and 3,000+ strong, spread across 7 offices with 1,500 based out of our Malta HQ alone! We recognise it may not be for the faint-hearted, but if you’re a go-getter, initiator and adrenaline junkie, always striving to push the boundaries and challenge yourself, then you’ll fit right in. CHALLENGE ACCEPTED? By submitting your application, you understand that your personal data will be processed as set out in our Privacy Policy
Salary Range: PLN 260,400 - 352,200 + Benefits + Equity Subject to alignment to the responsibilities and duties of the role. Location: Gdańsk - Hybrid Working Policy - 2-3 Days per week in office About Graphcore At Graphcore, we’re building the future of AI compute.We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale.As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence. Job Summary As a Senior QA Engineer within the Management & Observability team, you will be responsible for validating Graphcore's end-to-end telemetry and observability platform. Working closely with Telemetry and Observability engineers, you will design, develop and automate comprehensive test strategies covering telemetry generation, data collection, processing, storage, visualization and alerting. Your work will ensure that Graphcore's observability solutions are reliable, scalable and production-ready for both internal engineering teams and customers. You will contribute throughout the software development lifecycle by defining quality standards, building automated test frameworks and validating distributed systems operating at scale. Responsibilities and Duties * Define and implement the end-to-end quality strategy for Graphcore's telemetry and observability platform. * Design, develop and maintain automated functional, integration, system and regression tests covering the complete telemetry lifecycle - from telemetry generation to dashboards, APIs and alerting. * Develop automated validation frameworks for telemetry pipelines, data quality, metrics, logs, traces and time-series data. * Design realistic test environments capable of validating large-scale distributed deployments and production-like workloads. * Work closely with software engineers throughout design and implementation to ensure testability, reliability and quality are built into every component. * Validate performance, scalability, resilience and fault recovery of telemetry and observability solutions under realistic operating conditions. * Integrate automated testing into CI/CD pipelines and continuously improve test coverage, execution time and release quality. * Investigate defects through root-cause analysis, working with engineering teams to resolve complex system-level issues. * Develop quality metrics, test reports and release readiness criteria to support engineering and product decisions. * Contribute to continuous improvement of testing methodologies, automation frameworks and engineering best practices. Skills and Experience * BSc or MSc degree in Computer Science, Computer Engineering or equivalent practical experience. * 5–8 years of experience in Software QA, Test Automation or Software Engineering. * Experience designing automated test frameworks for distributed systems. * Experience testing cloud-native or infrastructure software running on Linux. * Experience with Python programming. * Experience building automated integration and system tests. * Familiarity with CI/CD platforms and automated testing pipelines. * Experience with Kubernetes, Docker and containerized environments. * Understanding of distributed systems, networking and API testing. * Experience validating REST and gRPC APIs. * Strong debugging and root-cause analysis skills. * Excellent written and verbal communication skills. Desirable: * Experience testing observability platforms based on Prometheus, Grafana, OpenTelemetry, ClickHouse, Kafka or Elastic Stack. * Experience validating telemetry pipelines and large-scale time-series data. * Experience with performance, scalability and resiliency testing. * Familiarity with Infrastructure as Code technologies such as Terraform or Ansible. * Experience testing AI infrastructure, HPC platforms or cloud infrastructure. * Experience with one additional programming language such as Go or C++. * Knowledge of modern observability practices including metrics, logs and distributed tracing. BENEFITS In addition to a competitive salary, annual leave policy, medical and dental health plans, a gym card and employee pension (matched up to 4%). We review our benefits on a yearly basis to ensure we offer a valuable and rewarding benefits programme to our employees. We welcome people of different backgrounds and experiences; we’re committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments. SPONSORSHIP Applicants for this position must hold the right to work in the Poland. Unfortunately at this time, we are unable to provide visa sponsorship or support for visa applications.
About Coupang We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did we ever live without Coupang?” Born out of an obsession to make shopping, eating, and living easier than ever, we’re collectively disrupting the multi-billion-dollar e-commerce industry from the ground up. We are one of the fastest-growing e-commerce companies that established an unparalleled reputation for being a dominant and reliable force in South Korean commerce. We are proud to have the best of both worlds — a startup culture with the resources of a large global public company. This fuels us to continue our growth and launch new services at the speed we have been since our inception. We are all entrepreneurial surrounded by opportunities to drive new initiatives and innovations. At our core, we are bold and ambitious people that like to get our hands dirty and make a hands-on impact. At Coupang, you will see yourself, your colleagues, your team, and the company grow every day. Our mission to build the future of commerce is real. We push the boundaries of what’s possible to solve problems and break traditional tradeoffs. Join Coupang now to create an epic experience in this always-on, high-tech, and hyper-connected world. Role Overview We are seeking a Sr. Staff Observability Engineer to lead the design and evolution of our observability platform for a GPU-as-a-Service (GPUaaS) infrastructure. This role will own the end-to-end telemetry strategy—from high-throughput metric ingestion to log pipelines and real-time visualization—powering deep insights into GPU clusters, datacenter systems, and distributed workloads. You will architect and operate planet-scale telemetry pipelines leveraging Grafana Alloy, Mimir, Loki, and Vector, ensuring high-fidelity observability across GPU workloads, Kubernetes clusters, and datacenter infrastructure. ---------------------------------------------------------------------------------------------------------------------------------- Key Responsibilities * End-to-End Observability Platform Ownership: Design and scale telemetry pipelines using: * Grafana Alloy for metrics collection (Prometheus-compatible pipelines) * Datadog Vector for high-throughput log ingestion and transformation * Grafana Mimir for scalable time-series storage * Grafana Loki for log aggregation and querying * Strategic Roadmap: Define the multi-year vision for GPU infrastructure observability, transitioning from reactive monitoring to SLO-driven, predictive, and automated observability. * High-Cardinality Telemetry Design: Optimize pipelines for GPU workloads characterized by: * High-cardinality labels (GPU IDs, tenants, workloads) * Burst-heavy workloads (ML training, inference spikes) * Multi-tenant isolation requirements * Architect low-latency, high-throughput pipelines capable of ingesting: * GPU metrics (utilization, memory, thermals, MIG partitions) * Kubernetes and container telemetry * Distributed system logs and traces * Build and optimize metric pipelines (Alloy → Mimir) ensuring: * Efficient remote_write tuning * Cost-effective retention strategies * Horizontal scalability and compaction tuning * Design log pipelines (Vector → Loki) with: * Structured logging and enrichment * Intelligent filtering/sampling * Stream partitioning for high-ingest environments * Establish deep observability into: * GPU hardware (NVIDIA DCGM, MIG, NVLink, PCIe) * Kubernetes GPU operators and scheduling behavior * Network fabric (RDMA, InfiniBand, TCP performance) * Define GPU-specific SLIs/SLOs such as: * GPU utilization efficiency * Job scheduling latency * Cluster fragmentation * Thermal and power anomalies * Build rich Grafana dashboards for: * Real-time GPU fleet health * Tenant-level usage and billing insights * Capacity planning and forecasting * Standardize dashboard frameworks and reusable panels across teams * Enable self-service observability for platform and ML engineering teams * Drive adoption of SRE principles: * SLIs, SLOs, error budgets tailored to GPU workloads * Integrate observability into CI/CD and IaC pipelines (Terraform/Kubernetes): * Automated canary analysis * Observability-driven rollbacks * Build automation (Go/Python) for: * Pipeline health monitoring * Dynamic routing and scaling of telemetry workloads * Develop tooling and practices for cross-layer correlation: * GPU → Node → Kubernetes → Application → Network * Lead deep RCA efforts for: * GPU contention issues * Performance degradation in ML workloads * Telemetry pipeline backpressure/failures * Enable “needle-in-a-haystack” debugging using unified logs + metrics * Mentor engineers and lead design reviews for observability systems * Act as a force multiplier across SRE, Infra, and ML platform teams * Promote Observability-by-Design in all new GPU cluster deployments * Drive adoption and contribution to: * Grafana stack (Alloy, Mimir, Loki, Tempo) * OpenTelemetry ecosystem * Define build vs. buy decisions (Datadog vs OSS vs hybrid approaches) * Optimize interoperability between Vector and OTEL pipelines * Architect secure telemetry pipelines with: * Encryption in transit and at rest * Multi-tenant isolation and RBAC * Data residency compliance * Implement Zero Trust observability patterns ---------------------------------------------------------------------------------------------------------------------------------- Qualifications & Requirements * BS/MS in Computer Science or equivalent practical experience * Extensive experience in Observability, SRE, or Distributed Infrastructure * Proven track record building large-scale telemetry pipelines (metrics/logs) * Observability Stack: * Grafana Alloy / Prometheus ecosystem * Grafana Mimir (or Cortex/Thanos) * Grafana Loki * Datadog Vector (or similar log pipelines) * Programming: * Strong in Go or Python * Data Systems: * TSDBs and log storage at scale * Infrastructure: * Kubernetes, Linux internals * GPU systems (NVIDIA DCGM, CUDA ecosystem) * High-performance networking (RDMA, InfiniBand preferred) * Cloud & Hybrid: * Experience building observability across: * Bare-metal GPU clusters * Hybrid cloud environments ---------------------------------------------------------------------------------------------------------------------------------- Core Impact Success in this role is measured by: * A highly reliable, scalable observability platform powering GPU infrastructure * Ability to diagnose complex GPU and distributed system issues in minutes * Enabling data-driven optimization of GPU utilization and cost efficiency * Building systems that proactively detect and mitigate failures before user impact ---------------------------------------------------------------------------------------------------------------------------------- Recruitment Process and Others Recruitment Process * Application Review - 1st Virtual Interview - 2nd Virtual Interview - Offer * The exact nature of the recruitment process may vary according to the specific job and may be changed due to scheduling or other circumstances. * Interview schedules and the results will be informed to the applicant via the e-mail address submitted at the application stage. Details to Consider * This job posting may be closed prior to the stated end date for application if all openings are filled. * Coupang has the right to rescind an offer of employment if a candidate is found to have submitted false information as part of the application process. * Those eligible for employment protection (recipients of veteran’s benefits, the disabled, etc.) may receive preferential treatment for employment in accordance with applicable laws. * Job titles and responsibilities may be subject to change depending on the candidate's overall experience, etc. this will be communicated to the candidate at the appropriate time before the offer. * Hiring may be restricted in case the legal qualifications required for hiring and work performance is not met. * This is a full-time regular position and includes 12 weeks of probation period; provided, however, the probationary period may be either skipped, shortened or extended if necessary for business purposes. Privacy Notice * Your personal information will be collected and managed by Coupang as stated in the Application Privacy Notice is located below. * https://privacy.coupang.com/en/land/jobs/ Document Return Policy 1. This notification is given pursuant to Article 11 (6) of the Fair Hiring Procedure Act. 2. A job applicant, who has applied but not been finally selected for a position at Coupang (the “Company”), may request the Company to return his/her hiring documents submitted pursuant to the Fair Hiring Procedure Act. However, this will not apply where the hiring documents were submitted via the website of the Company or e-mail, or where the job applicant submitted those documents voluntarily without a request from the Company. In addition, if the hiring documents were destroyed due to a natural disaster or any other reasons not attributable to the Company, such documents will be deemed to have been returned to the job applicant. 3. A job applicant who wishes to request the return of his/her hiring documents pursuant to the main sentence of paragraph 2 above should fill out a “Request for Return of Hiring Documents” [Annex Form No. 3 in the Enforcement Rule of the Fair Hiring Procedure Act] and submit It by email (recruitingops@coupang.com). In such case, within fourteen (14) days from the date of identifying the receipt of the request, the Company will send the hiring documents to the job applicant’s designated address via registered mail. Please be informed that the job applicant is required to pay the postage on the registered mail. 4. In preparation for a job applicant’s request for the return of hiring documents pursuant to the main sentence of paragraph 2 above, the Company shall retain the original hiring documents submitted by the job applicant for 180 days from the completion of the recruiting process. If no request is made until the end of this period, all his/her hiring documents will be destroyed immediately in accordance with the Personal Information Protection Act. 5. The above paragraphs 1 - 4 shall only apply when the labor-related laws of Korea govern the application. They are otherwise not applicable.