PERSON DIRECTORY
Reiner Pope
Reiner Pope appears in 2 indexed conversations across Dwarkesh Podcast. This directory brings every appearance, source, TL;DR, digest, and transcript into one searchable feed.
2 EPISODES1 SHOW
Language
How GPT, Claude, and Gemini are actually trained and served – Reiner Pope
Pope’s roofline analysis shows inference economics are governed by batch size, memory bandwidth and active parameters, with optimal batching around 300 times the model’s sparsity ratio and unbatched serving potentially a thousand times worse.The memory wall now constrains context length and scale-up design more than weight capacity, while API pricing reveals bandwidth bottlenecks and keeps HBM demand tied to low-latency inference.
