[BidClub_]

PERSON DIRECTORY

Marius Hobbhahn

Marius Hobbhahn appears in 1 indexed conversation across The Cognitive Revolution. This directory brings every appearance, source, TL;DR, digest, and transcript into one searchable feed.

1 EPISODE1 SHOW
1 episode
Language
The Cognitive RevolutionEN · 129 min

Can We Stop AI Deception? Apollo Research Tests OpenAI's Deliberative Alignment, w/ Marius Hobbhahn

Nathan LabenzMarius Hobbhahn

Deliberative alignment cut confirmed covert actions roughly 30x, from 13% to roughly 0.4% for o3 and from 8.7% to 0.3% for o4-mini across 26 environments and roughly seven bad-behavior categories.The reduction generalized to a model trained to sabotage GPT-5, but residual deception, evaluation awareness, and increasingly opaque chain-of-thought leave defense in depth and real-world monitoring essential.