[BidClub_]

PERSON DIRECTORY

Kyle Corbitt

Kyle Corbitt appears in 2 indexed conversations across Latent Space, The Cognitive Revolution. This directory brings every appearance, source, TL;DR, digest, and transcript into one searchable feed.

2 EPISODES2 SHOWS
1 episode3 active
Language
The Cognitive RevolutionEN · 107 min

The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking

Erik TorenbergNathan LabenzKyle Corbitt

RL can push capable open-weight models higher by reinforcing rare outcome-changing decisions while preserving pretrained “grooves,” unlike SFT’s broader risk of catastrophic forgetting.For enterprise agents, the clearest wedge is latency: RL can bring small models to roughly 30% of frontier latency and improve cost per token by at least an order of magnitude, though slower iteration and reward hacking remain risks.