Steadcast
NVIDIA AI Podcast cover art
NVIDIA AI Podcast

Snap’s Secret to Processing 10 Petabytes a Day: GPU-Accelerated Spark | NVIDIA AI Podcast Ep. 298

May 13, 202623 min · 3,767 words

Show notes

Snap processes more than 10 petabytes of experimentation data every single morning—and with NVIDIA GPU-accelerated Apache Spark on Google Cloud, Snap cut job costs by 76%, reduced memory usage by 80%, and eliminated 120 terabytes of disk spill from its pipelines.

Highlighted moments

We were able to cut almost about 76% of our job costs as a result of this migration. 76? 76.
with Spark Rapids, I want to mention it, we didn't have to change a single thing about how we ran the jobs.
when some of our biggest markets went to bed, a lot of our online inference GPU capacity was sitting idle.
we had to figure out how to gracefully fall back from GPUs to CPUs, right? And then, if the shared GKE resources itself was the constraint, then we had to gracefully fall back from CPUs to data proc clusters.

Transcript

Transcript not available for this episode yet.

More from NVIDIA AI Podcast

Inside Instacart's AI-Powered Smart Shopping Cart | NVIDIA AI Podcast Ep. 302

Jun 24, 202639 min

How Mistral Is Building Frontier AI for the Enterprise | NVIDIA AI Podcast Ep. 301

Jun 10, 202621 min

Everyone Can Build a Robot: Open Source Embodied AI With Seeed Studio | NVIDIA AI Podcast Ep. 300

May 27, 202629 min

Inside AI Tokenomics: How to Profitably Turn Tokens Into Business Value | NVIDIA AI Podcast Ep. 299

May 21, 202633 min

Harrison Chase of LangChain on Deep Agents, LangSmith, and Earning Trust | NVIDIA AI Podcast Ep. 297

May 6, 202624 min