[BidClub_]

PERSON DIRECTORY

Vibhu

Vibhu appears in 17 indexed conversations across Latent Space. This directory brings every appearance, source, TL;DR, digest, and transcript into one searchable feed.

17 EPISODES1 SHOW
4 episodes2 active
Language
Latent SpaceEN · 101 min

Recursive Language Models — Alex Zhang, MIT PhD

swyxVibhuAlex Zhang

Alex Zhang sees openings for smaller AI labs in overlooked research bets and output architectures, with JEPA potentially cutting binary-classification inference costs by 400 times while expert verification remains scarce.RLM and Prime Agent replace trajectory-as-a-prompt loops with persistent state and code-native delegation, but reported swarm economics—10,000 agents, 88 hours, 130 billion output tokens, and about $40 million—make orchestration and reliability the key watchpoints.

Latent SpaceEN · 101 min

Inference Is the New Training — Philip Kiely and Ali Taha, Basten

swyxVibhuPhilip KielyAli Taha

Inference remains an early optimization market: on identical hardware, quantization, caching, speculation and traffic tuning can typically deliver 2–4X gains, while production support requires weeks of debugging.The economics move customers from pay-per-token trials to dedicated capacity, as reliability, isolation and workload-specific tuning justify self-saturating infrastructure; faster interconnects and training-integrated optimization remain the next catalysts.

Latent SpaceEN · 58 min

The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO

swyxVibhuAkshat Bubna

Modal is repositioning infrastructure around agent experience, using decorators, CLI observability, and a 17-provider footprint instead of owning data centers.Its investment case rests on bursty workloads and elastic orchestration: speculative decoding may deliver 2× to 4× speedups, while batch pricing and reliability determine whether compute planning converts into margins.

Latent SpaceEN · 84 min

Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup

swyxVibhuNader KhalilKyle Kranen

NVIDIA’s Dynamo treats agent inference as a data-center-scale system, coordinating KV-cache-aware scheduling, disaggregated prefill and decode, and heterogeneous fleets above vLLM, SGLang, and TensorRT-LLM.The economic frontier is quality, cost, and latency across the entire workflow, with GB200 NVL72 cited as roughly 35 times cheaper per token than Hopper over much of one comparison curve.Brev’s DGX Spark integration and NVIDIA’s agent rollout could expand GPU access, but security controls remain unresolved as agents combine file access, internet access, and arbitrary code execution.