PERSON DIRECTORY
Kimbo Chen
Kimbo Chen appears in 3 indexed conversations across SemiAnalysis. This directory brings every appearance, source, TL;DR, digest, and transcript into one searchable feed.
Training a 400B Model on 2,048 Blackwell GPUs for $20M | Researcher Conversations at GTC
Arcee is moving into pre-training to break the sub-20B ceiling imposed by Llama, Mistral and Qwen while giving regulated customers control over data and model provenance.Trinity’s Arcee–DatologyAI–Prime Intellect structure and B300 availability target one month of pre-training, but immature sparse tooling leaves execution and compute economics as key risks.
Ep. 017 - DeepSeek V4 and Huawei Ascend NPU Performance (InferenceX) | Kimbo Chen, Cam Quilici, Bryan Shan, Jordan Nanos
Kimbo ChenCam QuiliciBryan ShanJordan Nanos
DeepSeek V4 changes inference economics by combining million-token context with roughly 100X lower KV-cache usage, while MegaMoE claims 1.5-1.73X speedups through communication-computation overlap.Huawei Ascend delivered credible release-day performance, but the larger catalyst is software iteration: AgentX will test realistic Claude Code traces, caching, and prefill-decode systems beyond synthetic chip benchmarks.
How Makora Generates CUDA Kernels That Beat Hand-Tuned Code | Researcher Conversations at GTC
Makora is expanding from kernel generation into a foundation-model-agnostic deployment engine that sells end-to-end performance across inference, training, reinforcement learning, numerics, and heterogeneous hardware.Sequential Monte Carlo speculative decoding reached roughly 5× the SGLang baseline and 2× SGLang speculative decoding with the experimental overlap scheduler in batch-size-one, low-latency tests, but it uses more compute, is fundamentally lossy, and may hit limits at larger batches.


