[BidClub_]

PERSON DIRECTORY

Reiner Pope

Reiner Pope appears in 2 indexed conversations across Dwarkesh Podcast. This directory brings every appearance, source, TL;DR, digest, and transcript into one searchable feed.

2 EPISODES1 SHOW
1 episode2 active
Language
Dwarkesh PodcastEN · 134 min

How GPT, Claude, and Gemini are actually trained and served – Reiner Pope

Dwarkesh PatelReiner Pope

Pope’s roofline analysis shows inference economics are governed by batch size, memory bandwidth and active parameters, with optimal batching around 300 times the model’s sparsity ratio and unbatched serving potentially a thousand times worse.The memory wall now constrains context length and scale-up design more than weight capacity, while API pricing reveals bandwidth bottlenecks and keeps HBM demand tied to low-latency inference.