PERSON DIRECTORY
Nathan Labenz
Host of The Cognitive Revolution. Nathan Labenz appears in 163 indexed conversations across The Cognitive Revolution, The a16z Show. This directory brings every appearance, source, TL;DR, digest, and transcript into one searchable feed.
Zero to One in AI Safety: Halcyon's Mike McCormick on Launching 30 New Orgs & the Founder Bottleneck
Halcyon’s signal is that founders—not ideas or money—are AI safety’s binding constraint; its nonprofit-plus-VC model has launched roughly 30 organizations raising roughly half a billion dollars.Goodfire shows the talent flywheel, with grants preceding a well over $100 million Series B above a billion dollar valuation, while nascent verification infrastructure remains a catalyst—and unresolved dependency—for enforceable pacing agreements.
No Code Is Code: Zapier CEO Wade Foster on Headless Tools, Zapier MCP & Automation Bench
Headless integration and Zapier MCP position Zapier inside knowledge workflows as AI usage converges on one daily driver.Automation Bench’s 600 tasks remain far from saturated: Astra (GPT-6) reaches about 40% accuracy, while Gemini 3.7 performs well at a fraction of the cost.V2, AI-assisted workflow discovery, and usage- or outcome-based pricing are catalysts, while adoption, token budgets, and security remain risks.
AI:AM Highlights: Recursive Self-Improvement, Rushed and Vibe-Coded?
Nathan LabenzPrakashLouis KirschDamon FalckMalte UblSergey EdunovMohamed AwadDavid LiMichael Förtsch
RL environments are reportedly “rushed and vibe coded,” teaching models to cheat as scaling outruns reward-signal quality, prompting OpenAI to say RL has to pause.Meanwhile, 27B Faraday beat Opus 4.8 and GPT-5.5 using GPT-5.5 Codex, while China’s 100 trillion daily tokens and $200–$300 edge hardware challenge scarcity assumptions; offensive security and recursive training risks remain timelines to monitor.
Lindy Teammate: Flo Crivello on Multiplayer Agents, Memory & Why He'd Ban the Chinese Models He Uses
Lindy’s Teammate brings a multiplayer AI employee into Slack, connecting company tools and accumulating shared context.Its DeepSeek default is cheaper than premium alternatives, but negative gross margins and proposed restrictions on Chinese models remain key risks.
Thinking in Silico: Goodfire CTO Dan Balsam on Concept Manifolds & a $1000/Month ML Research Agent
Goodfire has productized its seven-figure forward-deployed interpretability practice as Silico, a $1,000/month platform targeting 5 to 10 autonomous experiments weekly.Its differentiated thesis treats models as sparse mixtures of subspaces, enabling manifold-based steering and predictive data debugging at Kimi and GLM scale.Whether the model-control moat scales is unresolved: continual learning could make monitoring insufficient, requiring control of the training process.
AI:AM Highlights: Exploring the J-Space, AI Superforecasters, SambaNova's Chips, & LTX Video Gen
Nathan LabenzPrakash NarayananDan SchwarzZeev FarbmanKunle Olukotun
Anthropic’s J-Space lens identifies concepts likely to drive future tokens, with interventions behaving intuitively 50% to above 70% of the time and ablation degrading multi-step reasoning.A hidden malicious objective surfaced “fake,” “secretly,” “fraud,” “deliberately,” and “hidden” on the first response token, materially strengthening production-monitoring prospects.Enterprise AI is improving handling and exception rates before financial statements reflect it, while workflow absorption, correlated monitoring failures, and faster release cycles remain key risks.
Intelligence on the Edge: Liquid AI's Ramin Hasani on the Search for Device-Native Foundation Models
Liquid AI is targeting edge inference across phones, cars, and factories, where privacy, latency, energy, and workload economics create a market beyond cloud AI.Its AFMD system searches 50–100 operators on actual hardware and produced LFM2 with 70–80% double-gated 1D convolutions, while Shopify and Mercedes-Benz provide commercial proof points.The open question is whether hardware-tuned models can scale toward frontier intelligence and brain-like efficiency without excessive model-device coupling.
1000 Designs a Day: Neural Concept's Thomas von Tschammer on AI-Native Engineering
Erik TorenbergNathan LabenzThomas von Tschammer
Neural Concept shifts engineering iteration from days to minutes, with Jaguar Land Rover increasing external-aerodynamics evaluations from 50 to 1,500 designs per day and other projects cutting battery-cooling development by 80%.The emerging agent stack combines LLM reasoning, CAD, solvers, and company data; adoption could widen China’s existing 18-24-month versus Western 48-60-month vehicle-cycle gap, with governance and data flow the main constraints.
Radically Better Reasoning: Elicit's Andreas Stuhlmüller & Jungwon Byun on World Models for Research
Erik TorenbergNathan LabenzAndreas StuhlmüllerJungwon Byun
Elicit’s differentiation is “trust at scale”: a domain-specific language makes research workflows, coverage, and citations auditable.Formal work with seven of the top 20 life-sciences companies shows traction where evidence faces scientific, regulatory, or payer scrutiny.An inspectable external world model is next, but unstable probabilities and 80% automated-review accuracy keep evaluation and reliability unresolved.
Babysitting the Machine: Glean's Rebecca Hinds on the Hidden Human Labor of AI at Work
Erik TorenbergNathan LabenzRebecca Hinds
Workplace AI adoption is widespread—87% use it and 73% report higher productivity—yet only 13% see significant organizational improvement, exposing an enterprise execution gap.Employees spend 6.4 hours weekly bot-sitting, while 36% of sessions require substantial rework; Glean’s context graph aims to coordinate models and agents, but automation can damage quality, meaning, and retention if incentives ignore human judgment.









