SHOW DIRECTORY
Latent Space
Technical conversations for AI engineers and builders covering models, agents, developer tools, inference, data, and production infrastructure.
BIDCLUB DESCRIPTION
Recursive Language Models — Alex Zhang, MIT PhD
Alex Zhang sees openings for smaller AI labs in overlooked research bets and output architectures, with JEPA potentially cutting binary-classification inference costs by 400 times while expert verification remains scarce.RLM and Prime Agent replace trajectory-as-a-prompt loops with persistent state and code-native delegation, but reported swarm economics—10,000 agents, 88 hours, 130 billion output tokens, and about $40 million—make orchestration and reliability the key watchpoints.
Inference Is the New Training — Philip Kiely and Ali Taha, Basten
Inference remains an early optimization market: on identical hardware, quantization, caching, speculation and traffic tuning can typically deliver 2–4X gains, while production support requires weeks of debugging.The economics move customers from pay-per-token trials to dedicated capacity, as reliability, isolation and workload-specific tuning justify self-saturating infrastructure; faster interconnects and training-integrated optimization remain the next catalysts.
The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO
Modal is repositioning infrastructure around agent experience, using decorators, CLI observability, and a 17-provider footprint instead of owning data centers.Its investment case rests on bursty workloads and elastic orchestration: speculative decoding may deliver 2× to 4× speedups, while batch pricing and reliability determine whether compute planning converts into margins.
AI Agents Need Computers: 74% MoM Growth, 850K/Day Runs, & New Agent Cloud — Ivan Burazin, Daytona
Daytona is repositioning the sandbox as an API-addressable, stateful computer for AI agents, reporting 74% month-over-month growth versus roughly 40% for the broader infrastructure market.Bare-metal scheduling, local NVMe snapshots and pause-resume state deliver 60-millisecond starts, but RL/eval bursts can drive 100,000 CPUs at once, leaving 15% average utilization and making committed capacity and GPU-CPU coordination central to scale.
The Agent-Native Cloud: 3M Users, 100K Signups/Wk, Data Centers, & Death PRs — Jake Cooper, Railway
Alessio FanelliswyxJake Cooper
Railway is betting agents will dominate software building, using bare metal and a versioned, forkable application layer to turn deployment into continuous, reversible evolution.Hardware reportedly pays back in about three months with roughly 70% margins, but CI/CD may melt under parallel workloads; autonomous remediation, eventual GPUs, and disciplined capacity financing remain key watchpoints.
Marc Andreessen introspects on Death of the Browser, Pi + OpenClaw, and Why "This Time Is Different"
Marc AndreessenswyxAlessio Fanelli
Andreessen argues AI has crossed from fluent generation into consequential reasoning, coding, agents, and recursive self-improvement, making the technical progress real despite possible boom-bust cycles.GPU overbuild remains a dot-com-style risk, but strong balance sheets, sold-out supply through the next three to four years, and software-driven chip value make betting against infrastructure dangerous while security, consolidation, and adoption remain unresolved.
Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
swyxVibhuNader KhalilKyle Kranen
NVIDIA’s Dynamo treats agent inference as a data-center-scale system, coordinating KV-cache-aware scheduling, disaggregated prefill and decode, and heterogeneous fleets above vLLM, SGLang, and TensorRT-LLM.The economic frontier is quality, cost, and latency across the entire workflow, with GB200 NVL72 cited as roughly 35 times cheaper per token than Hopper over much of one comparison curve.Brev’s DGX Spark integration and NVIDIA’s agent rollout could expand GPU access, but security controls remain unresolved as agents combine file access, internet access, and arbitrary code execution.
Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
Claude Code with Opus 4.5 now behaves like a junior analyst, compressing PhD-scale projects into days while still requiring expert review of its frequent mistakes.Adoption near 4–5% of public GitHub and a memory shortage that could lift DRAM prices another 100% point to accelerating agent demand, constrained infrastructure, and exposed information-work roles.
SF Compute: Commoditizing Compute to solve the GPU Bubble forever
Alessio FanelliswyxEvan Conrad
SF Compute treats GPUs as tradeable, financeable capacity rather than software infrastructure, turning unused commitments into a spot market and separating hardware ownership from higher-margin software businesses.Long-term contracts with creditworthy customers can make GPU ownership resemble real estate, while hourly resale improves utilization and flexibility; a future cash-settled GPU index could reduce financing risk, but standardization, chip supply and physical cluster reliability remain unresolved.








