[BidClub_]
Frictionless · · 73 分钟

AI能预测未来吗?——与Gensyn联合创始人兼CEO Ben Fielding 对谈 | EP 171

Logan JastremskiBen Fielding

AI与软件区块链技术企业经营
YouTube ↗
TL;DR
  • Ben Fielding 的核心判断是,前沿实验室已被困在一场资本密集型扩模竞赛中,而这些模型很可能仅靠使用收入无法收回开发成本。 蒸馏和开放权重发布会迅速复现以数百亿美元开发的能力,几乎「没有可防守的护城河」(no protective moat)。因此,持久价值应归属于拥有大量付费且长期留存客户的产品,而不是暂时领先的基准测试成绩。

  • AI的约束条件已从算力转向及时、可信的信息。 2021年,互联网数据供给充足,瓶颈在算力;如今实验室可以搭建超大规模集群,却在「几乎把整个地球翻遍」寻找新的训练材料和可靠的真实信号。Jastremski 将不断上升的赌注概括为每GW约500亿–700亿美元,今年可能新增20 GW产能,明年30 GW,2028年50 GW。

  • Gensyn 已证明去中心化训练的若干关键环节可行,但 Fielding 不再认为分布式算力是最高杠杆的业务。 其验证技术可以在 NVIDIA、AMD、x86和MacBook硬件上复现并检查机器学习操作,RL Swarm 据称还协调了约40,000台设备。但集中式预训练仍然更快,算力供给也已扩张,目前没有哪类天然客户需要去中心化训练所要求的那种截然不同的模型架构。

  • 公司的「中间道路」保留去中心化的优势,却不要求每次计算都经过去中心化网络。 Open1B 在集中式数据中心高效训练完成,但任何人都可以复现并审计每一步操作,而不只是查看开放权重或训练配方。Fielding 的关键区分是:在必要之处去中心化验证和信任,同时把算力放在最适合完成任务的地方。

  • Fielding 认为,预测可能成为AI的下一轮大浪潮,尽管它获得的关注度「大概只有10%」的LLM水平。 LLM能够压缩并重排信息,但涉及重大后果的决策仍需人类判断;预测系统则可以进一步估计这些决策的结果。他对产品的比喻是「一台扑克式的万事赔率计算器」(a poker-like odds calculator for everything)——一台比人类更擅长预测未来、并持续改善决策的机器。

  • 预测最难的部分,是在结果揭晓前收集证据,再在现实落地后对证据进行奖励。 事后重建历史会受到 hindsight、编辑和搜索结果更新的污染,因此 Gensyn 的 Deli 项目采用不可篡改的区块链记录、无许可市场和可靠且经过测试的AI裁决机制,把带时间戳的主张与之后的结果连接起来。证据市场还可以直接购买预测背后的信息,而不只是观察一笔10美元的押注。

  • Fielding 认为,只有在去中心化能换来不可替代的东西时,crypto与AI才会真正汇流。 预测市场交易量、代币叙事或哲学信念都不够;「哲学付不了账单」(philosophy doesn’t pay bills)。真正有护城河的交集在于全球信任、自动结算、微支付,以及直接奖励充当信息传感器的人类——他认为这些能力无法由中心化运营者复刻。

摘要 · 为研究而整理的核心内容

1. Fielding走向去中心化,起点是并行搜索。

  • Fielding 于2015年进入深度学习领域,当时GPU加速正把神经网络从理论上可扩展的机制变成实用系统。他的第一篇论文描述了一款配备摄像头的3D聊天机器人,但目标识别大约需要5分钟,因此系统讨论的其实是5分钟前出现在视野里的内容。这是一个很糟糕的系统,却展示了这项技术的潜力。

  • 他的博士论文从计算机视觉转向神经网络架构搜索。与其让研究人员一次手工寻找一套更好的层级组合,他使用进化算法和群体算法并行探索大量候选网络;论文背后的信念是,机器学习应被视为一组计算原理,而不应被简化成「我们永远使用反向传播」。

  • 这种框架自然导向去中心化:地理上分散的机器能够提供规模最大的并行工作环境。Fielding 的现实动机同样朴素——他桌下有几块GPU,想要更多,却很难获得使用权限。这说明资源约束本身就是研究障碍,而不只是基础设施不够方便。

2. 横向扩展要求改变算法,而不是把机器做大。

  • Fielding 描述了人们熟悉的纵向扩展路径:从 GeForce GTX 980 到 GTX 1080,再从1块GPU扩展到4块,最后进入集群节点阶段。由于算法保持不变,每一步起初都几乎感觉不到成本;但很快,内存、主板总线、互联带宽和散热会依次成为新的瓶颈。

  • 他偏好的类比,是 Google 早期从纵向扩展 PageRank 转向 MapReduce。把任务重写成彼此可分离的部分,可以使其「近乎无需协调即可并行」,让新增机器并排扩展任务,而不是继续把越来越昂贵的能力堆到单台系统上。

  • 机器学习同样包含大量可并行化问题,架构搜索就是 Fielding 直接熟悉的例子。他认为,分布式数据中心可以共同容纳比任何单一设施大10倍、100倍甚至1,000倍的模型,但要兑现这一上行空间,需要新算法,也需要解决集中式集群中根本不存在的问题。

  • Jastremski 的观察值得保留:现代机架、NVLink以及400Gb、800Gb乃至Tb级互联,正是为了让集群表现得像一台机器。Fielding 重申,纵向扩展能带来更快的即时回报;只有当效率下降到迫使架构改变时,横向方案才真正有吸引力。

3. 可验证的执行,比验证GPU身份更重要。

  • Gensyn 最初的问题是:能否协调全球各地闲置的消费级GPU。Jastremski 用一台声称是B300的设备与一台Mac mini作对比,说明其中的信任难题;Fielding 的答案是,支付体系和经济网络都不能建立在不可信的执行之上。客户必须确认所要求的工作确实完成了,而不管服务提供者声称使用了什么硬件。

  • 因此,他偏好的信任原语是证明操作确实发生,而不是证明硬件规格、服务商声誉或服务等级协议。定期检查服务商是否真的拥有其声称的GPU,只会制造不断作弊的「打地鼠」激励;证明精确执行则会消除欺骗的价值,同时让服务商可以自由优化基础设施。

  • Gensyn 表示,它可以在 NVIDIA GPU、AMD硬件、x86机器和MacBook处理器上复现机器学习操作。这一精确复现基础支持乐观证明,并进一步支持速度更慢的零知识证明;正是这一技术结果让真正的无许可算力成为可能,尽管公司后来改变了商业重点。

4. 集中式算力赢得了短期市场,迫使Gensyn转向。

  • Fielding 明确改变了看法:去中心化训练在技术上可行,但集中式AI吸引资源的速度太快,要证明去中心化的决定性优势极其困难。即使前沿系统面临边际回报递减,实验室仍可以用去中心化竞争者无法匹敌的投入规模突破障碍。

  • Gensyn 2021年的判断之后,算力市场也有所改善。芯片数量增加、neocloud供应商增多、竞争加剧,都缓解了资源瓶颈;与此同时,大多数GPU消耗集中在执行特定任务的中心化超大规模云服务商手中。

  • 因此,去中心化算力网络面临的是客户缺位问题。任何要搭建这类网络的团队,都必须同时成为一家设计全新模型的研究实验室;但 Fielding 尚未见过去中心化AI公司筹集到他认为参与前沿竞争所需的数百亿美元。业余爱好者、小型GPU用户和部分主权工作负载只能构成细分市场,不是Gensyn想要的规模。

  • 只要预训练在速度和成本上仍有实质优势,就应继续保持集中式。Gensyn 的 Open1B体现了这条「中间道路」:模型在数据中心训练,但每一步训练操作都可以由任何人独立重复和检查。它提供的远不只是开放权重或公开配方,而是把审计去中心化,而非把执行去中心化。

5. RL Swarm已经跑通,但分布式后训练仍需要合适的产品。

  • Fielding 认为,强化学习后训练是去中心化算力剩下的最强机会。一个产品可以基于获得同意的客户在各自设备上的活动进行训练,让用户群同时成为信息来源和分布式学习网络。

  • RL Swarm 已经用约40,000台设备共同训练同一任务上的模型,证明了这条研究路径的可行性。尚未解决的是产品设计:编程助手或 Hermes kernels 这类开源项目可能受益,但应用和学习系统必须一起构建。

  • 这种组合「有点像奇迹」,因为同一家公司必须同时解决产品问题和研究问题。Gensyn 认为自己并不适合两头兼顾,于是把基础设施转向去中心化更有把握解锁的资源:信息。

6. 开放模型让研究领先变成易逝优势。

  • Fielding 将机器学习的约束分为算力、数据,以及真实度或标注。互联网提供的文本多到研究人员处理不过来时,算力是主导瓶颈;随后,自监督的下一词预测缓解了人工标注瓶颈,而强化学习则通过奖励函数生成标签,不再依赖大批标注人员。

  • 如今约束条件再次转移。实验室增加算力的速度快于寻找有用信息的速度,因此开始在全球范围内搜寻数据。Fielding 提到,有报告称有人把书籍实体拆开后扫描;与此同时,Lean这类基于规则的环境可以充当奖励函数。

  • 蒸馏意味着,任何通过可访问模型暴露出来的能力,都比当初创造它更容易被竞争者复现。Fielding 预计,开源仍将是战略均衡器:落后的实验室可以通过发布模型削弱领先者的优势;不过他怀疑,如果今天的开源倡导者成为前沿领导者,立场可能会反转。

  • 在他看来,持久护城河是「庞大的客户群」,并且这些客户会反复为有用的产品付费。他认为 OpenAI 和 Anthropic 从研究向更广泛业务扩张,正是对这一现实的承认;他也反驳 Meta 输掉只是因为 Llama 4令人失望的说法:Meta 仍然拥有规模庞大、深度嵌入用户生活的生态。

7. Neocloud经济学可能把公司困在昨天的瓶颈里。

  • Fielding 承认,neocloud运营商可以利用算力市场的阶段性低效,其中一家也可能成为大赢家。但他的战略异议是:归根结底,这类供应商仍是「大宗商品的经销商」,会暴露在规模效应之下,而规模效应最终应使价格逐步正常化。

  • Gensyn 曾考虑通过半集中式算力市场创造收入,同时为研究提供资金,最终放弃了这一方案。收入一旦到手,投资人就会追问如何扩大收入,公司也会围绕服务这项收入进行优化;5年后,团队可能发现自己已经非常擅长一个从未真正想做的细分领域。

  • Fielding 用「掉进沼泽」来比喻这种结果。他关注的是能够扩展机器学习的变革性机制,而不是一个看起来很有吸引力、却会让未来转型更困难的过渡业务,尤其是在产品分发比增量式研究优势更重要的情况下。

8. 预测可以让AI从信息检索走向判断。

  • Fielding 将今天的LLM描述为一名能力极强的图书管理员:它压缩海量多模态信息,重新组织后返回一个听起来很有说服力的结果。但涉及重大后果的输出仍需人类验证,因为模型还没有足够的机制去理解现实世界中的因果影响。

  • 预测获得的关注度「大概只有10%」的LLM水平,但在 Fielding 看来,它所处的位置大致相当于2021年的下一词预测。实现规模化所需的要素已经存在,剩下的是工程和资源问题,这为把学术预测任务转化为通用的「决策与判断智能」创造了机会。

  • 他最有力的类比是线上扑克与现场扑克:前者的软件会持续显示每手牌的概率,后者则要求玩家自行计算。一个充分规模化的预测系统将成为「扑克式的万事赔率计算器」(a poker-like odds calculator for everything)——它不只是预测事件,还会展示日常选择可能带来的后果。

9. Deli把带时间戳的人类证据变成预测飞轮。

  • 预测数据无法在事件发生后安全重建。研究人员1个月后进行搜索时,无法知道某个页面是否以原始形式提前存在,还是在结果出现后被编辑过;唯一可靠的方法,是在当下捕捉信息、保存信息,并在未来到来时进行评分。

  • Jastremski 把这一过程比作奖励周期异常漫长的强化学习。系统记录关于一个事件可能如何影响另一个事件的连续主张,等待终局结果出现,再返回这段序列并发放奖励。

  • Gensyn 的 Deli 是一个无许可信息市场:任何人都可以为未来事件创建市场,参与者可以用资金押注自己的判断,再由可靠且经过测试的AI模型自动裁决。不可篡改的账本记录了某项主张在何时出现,而押注本身提供了信号:一个理性的人如果愿意押上10美元,想必认为自己掌握相关信息。

  • 传统预测市场往往围绕交易量展开;Jastremski 指出,手续费让这种模式在商业上顺理成章,但 Fielding 认为,某些鼓励交易量的机制反而会主动损害信息聚合。Gensyn 提出的证据市场,则购买预测背后的信息,并在预测兑现时奖励提供者,让人成为获得报酬的「现实世界传感器」,而不是无偿贡献训练数据。

10. Crypto与AI将在信任无法被中心化消除之处汇流。

  • Fielding 区分了对去中心化的哲学支持与可行的商业模式。去中心化系统通常成本更高,因此创业者必须明确回答「你到底买的是什么」;那些为了关注度或资本而采用crypto的公司,最终都会在发现「哲学付不了账单」后离开。

  • Gensyn 的标准是现实中的必要性。全球市场结算、不可篡改的主张,以及陌生参与者之间的信任,需要去中心化,而普通企业采用往往不需要。钱包、微支付和链上支付流也都是有价值的额外收益,但信任才是用户真正付费购买的核心价值。

  • 这种系统可能把模型训练的经济利益还给用户。某人提供的信息如果后来被证明正确,就可以直接获得部分强化学习价值,从而形成一个大型离策略训练闭环,让人类和机器共同贡献、共同分享奖励。

  • Fielding 最后的预测是确定的,尽管落地仍然困难:一个能比人类更准确预测未来的「盒子」可以存在。正如2021年的聊天机器人后来演变成能够完成大量知识工作的工具,他相信LLM会成为更大预测系统的组成部分,而预测将成为AI的下一轮重大浪潮。

完整逐字稿
Ben Fielding

I believe that centralized AI faces a serious problem, and I think the labs know it. Given the amounts they invest in developing their models, will they really be able to pay for them through the use of these models in products? There is no defensible moat.

The bottleneck we’re focused on now is information flow: how can we receive true, verified information from people and, let’s say, sensors all over the world, and constantly collect it into the system to make a model as good as possible? That’s what we’re focused on now, because I think it’s the most effective thing we can do.

Forecasting—the idea of predicting future events—gets probably 10% of the attention now given to large language models. But it’s actually in a very similar situation to where LLMs were in 2021. If you’re a poker player, you’re used to playing online with an odds calculator, where you know the probability of each hand in real time because it shows it to you. When you play live, you don’t have that, and it’s much more difficult. You have to calculate it independently.

Imagine if your whole life were like this: a poker calculator with probabilities that worked for everything. You could have this, and in fact, this “box” that predicts the future can exist.

1. Ben Fielding and the Origins of Gensyn

Logan Jastremski

Hi, Ben. Thank you for coming to the podcast. I was looking forward to this. I think it’s very timely. As we just discussed, the world changes a lot. AI safety is at the top of all the headlines, but there’s also this feeling that AI will be in every aspect of our lives, which, in my opinion, completely deserves a lot of attention.

2. Rewarding the Evidence Behind a Prediction

I’ve wanted to chat with you for a long time, and I’m grateful that you came to the podcast so we can see how you’re thinking about this. I think you have a unique perspective on the world in all respects. What are you doing with Gensyn?

Ben Fielding

Yes, thank you for inviting me, and I agree that the timing is very interesting. A lot is happening, and I know that we’ve talked a lot in the past about AI, perhaps even before it attracted such huge attention. It’s interesting to watch some predictions we talked about for years come true, while others are implemented into something else or become things we didn’t expect. It’s pleasant to get together again and talk about what is actually happening.

Logan Jastremski

100%. So maybe that’s why we should begin there. How did you become interested in Gensyn? How did this idea come to you? Tell me briefly, for those who aren’t very familiar with you and what Gensyn has done over the last few years.

Ben Fielding

Of course. My specialty is computer science, specifically research in machine learning. I started my postgraduate studies in 2015, so over 10 years ago, specifically in AI and, more precisely, deep learning. These are neural networks, which had only just started showing significant promise in an applied sense. If anyone knows the story of deep learning, AlexNet and accelerating neural networks on GPUs gave them a real boost.

They existed long before that, but they were rather interesting, peculiar mechanisms that theoretically could scale but, practically, were impossible to scale until we began accelerating them on processors. That started the wave of applying these methods to other problems, to see what we could really do with them. At first, it was computer vision, and that was my field, so I focused directly on deep learning in computer vision.

Interestingly, I don’t know if many people know this, but my first scientific article, which I wrote in 2015, was about a chatbot. It was a 3D chatbot that you could communicate with; it used a camera, recognized what was happening around it, and could talk to you about those things. It worked very badly because the methods at the time took a long time to process what was happening in front of the camera. It took us about 5 minutes to localize objects, recognize them, and then say something. So you talked to the chatbot, and it told you about something that had happened in front of the camera 5 minutes earlier.

The point was to show: look, this is possible; all these technologies are already here. It was a huge mistake in terms of how the system worked internally, but it was the beginning of my research. I eventually delved into the theoretical part, especially focusing on neural network architectures.

The idea was that a neural network is basically just a big graph. At one point, many of the advances people found came from developing new structures for this graph. Someone would develop a certain set of layers for a network, launch it, and it would work better than the previous version. That was considered a form of progress in research. But this quickly became outdated because it was simply a search performed by people.

So I focused my dissertation and my general research on automating this process. Most relevant to what I do now and to what we’re doing at Gensyn, I have a deep belief that we can parallelize this search process. I turned to evolutionary algorithms, in particular swarm algorithms. The title of my dissertation was “The Use of Swarm Algorithms for Optimization of Neural Network Architectures.”

Many years ago, I understood how it was possible to use swarms of many different candidate networks to explore the space faster than with a manual or linear approach, or even with backpropagation. Even then, this showed me that significant parallelization was possible. We could achieve significant gains if we considered machine learning not as one specific method where “backpropagation is always what we use,” but as a set of computer science principles for problem-solving.

It seems to me that the world constantly hesitates between these 2 positions. A lot of developments in scaling large language models suffer from tunnel vision in many companies and among individual researchers who are trying to scale only one thing. In a sense, they’ve lost sight of the bigger picture, where, by taking a step back and using a parallelizable method, you can get better results.

I think this is a classic pattern in technologies that we’ve repeatedly observed. This is, in fact, my short summary of what brought me to decentralization, especially because it provides the most scalable environment for parallelizing anything. But this is accompanied by a whole series of challenges, and these are the very problems we’ve been fighting at Gensyn for close to 7 years now.

Ultimately, the opportunities for parallelization and an understanding of resource constraints in machine learning led me to what I’m doing now. When I talk about resource constraints, I mean it literally: while writing my dissertation and carrying out experiments, I used multiple graphics cards in a computer that I built under my desk, and I wanted more, but getting them was hard. It was hard to get access, and it was difficult to trust the people who could give me access.

3. Scaling AI Beyond Bigger Clusters

This taught me that we need to solve these resource problems. If we did, we could scale parallel machine training much further than we did then—and further than we do now. So I think there are many ways to figure this out.

Logan Jastremski

Maybe we can start with the fact that scaling clusters became popular around 2017, or that’s when it started. Then, in 2019, the approach was, “Okay, there’s a Bitter Lesson: just throw more computational capacity at it and start to scale.” This worked for a while, until we reached reinforcement learning for specific tasks.

Can you tell me about the approach you’ve observed so far—simply adding computational capacity compared with what you’re doing at Gensyn—and say a little more about the algorithmic part?

Ben Fielding

Yes. I think this applies to the problem of vertical scaling versus horizontal scaling, which is inherent to technologies in general. In the case of vertical scaling, you develop a certain function, program, or algorithm that performs a task for you. You look at this and say, “If only this thing had more resources, it would work faster.” And you just give it more resources because, at first, most likely, you can do that. You can simply scale the resource.

A good example is that in 2015 and 2016, I spent a lot of time thinking about which graphics card to run my own developments on. I started with consumer models. I remember I had a GeForce GTX 980, and then the 10-series came out, and I thought, “If I get a GTX 1080, I can do more.” It has more memory and more CUDA cores, so I can scale the task without changing anything. I’m just using this new, more powerful device, and it’s almost free because I don’t need to change any code or algorithm.

As people went through this process, algorithms had some minor changes added. At first, sometime in 2015, the question was: is it possible to switch from 1 GPU in the machine to 4? Yes, we can use more memory. You can move data between GPUs, and then the buses become their own limiting factor.

You start to think: do you have a suitable motherboard with a wide enough bus to ensure data exchange? Then the question arose: can we move to nodes in a cluster? Yes, but then all these problems with communication within the cluster arise. We were looking for the next bottleneck from an equipment point of view and trying to eliminate it, just to accelerate the same algorithms.

This process continued for long enough. When it came to clusters, it was possible to achieve significant scaling. But lately—although this had already started a few years ago—it became clear that, at some point, you encounter problems that are impossible to solve by simply adding resources.

Physical limitations arise, for example, with cooling, and you invent new, cunning ways to bypass them. But at a certain moment, efficiency drops, and you have to make dramatic changes in the architecture of the building systems, up to creating new conditions to force them to work.

For us, it is obvious that vertical scaling is the fastest way to scale. You can just throw resources at the task and receive an immediate return. But when efficiency begins to fall, you need to look for a completely different approach.

I always remember the MapReduce example from Google in the early stages of development. At first, PageRank scaled vertically, until it reached the limit when further increases in server capacity gave less and less increase in productivity. Then Google thought about how to achieve this: Can we achieve it with the help of another algorithm that will change the resource constraints underlying it? They invented MapReduce as a method of parallelization, essentially transforming the task into something embarrassingly parallel.

This is something that can scale by adding resources, not by building up capacity from above. Can you split an algorithm into 2 parts, run them on 2 different devices, and get the same result? In many cases, yes, but for this you need to reformulate the problem to a certain extent, although the solution is 100% complete. Therefore, I always remember this lesson.

I believe machine learning has huge potential to apply this in many ways. I know at least 1 of them: search architecture, which is easy to make embarrassingly parallel. This is what simply works exactly like that. Therefore, I know at least 1 problem in machine learning that can be solved in this way. I know there are others, and I want the world to discover them and scale horizontally, not vertically, simply by adding more resources. You just need to change the algorithm and use existing or parallel resources.

Logan Jastremski

That was interesting. I spend more time studying how data centers are arranged. In some respects, I observe a lot of similar problems with bandwidth that concern the blockchain industry within data centers. Even a server has a certain bandwidth capability. This concerns both the hardware—GPU-to-memory communication—and the interconnects between racks, which also have a certain bandwidth capability.

Then the question arises: A rack can accommodate only 72 GPUs, so it is necessary to go beyond its limits. I wonder how to continue scaling, trying to force the entire cluster to work as a single whole, and it seems we will try to scale this as hard as possible. But as you rightly noticed, you eventually run into certain limits, whether of placement or something else.

Whatever it is, perhaps the computational power is in different parts of the world and is not so concentrated in one place. How can you make use of this? Horizontal scaling is another direction that, I think, is less appreciated today.

4. Why Startups Can Take a Different Approach

Ben Fielding

Yeah, I think this is one of those things where it is, to some extent, a motivation problem. Imagine a company that has a product—for example, a scalable LLM that works in a certain way—and they want to train the next, bigger version. They actually ask the company, “We’re going to make a larger version of this. For this, we need to build a larger data center,” and so on. Then you get a bunch of people focused only on this problem.

But you don’t necessarily get people focused on the problem of what this network was originally intended to do—that is, what fundamentally determines it. That company has, in some sense, already passed this stage and is currently in the scaling phase of its solution. But then startups may appear that are capable of looking at this much more broadly than the company and saying, “Okay, yes, the company is scaling its solution, but actually there could be a whole new solution space for this or adjacent problems that could be implemented differently.”

We think about it at the company level, but you can think about it at a technical level, asking, “What is actually fundamentally taking place in machine learning?” The lens I always come back to, which in my opinion dispels a lot of this fog, is to say, “Okay, forget about all this magic, internal processes, and details. What you actually have is a huge amount of data in the world and a certain signal on top of this data, which I usually call information. This is, in essence, data plus labels in the context of supervised learning.”

You’re trying to squeeze that into a representation that is smaller, faster to access, and multimodal. You can use this differently; it’s a different representation, but fundamentally this is just a great compression mechanism. That’s what’s happening, and neural networks are universal function approximators. They are capable of learning these compression mechanisms, which are actually very powerful, and that’s why they are so good: We don’t need to identify them through expert systems or something like that.

But you could actually do this compression in various ways, and it could be compression by several parts, not 1 big block in 1 data center. While you can find ways, as before, to carry out search in the representation, make the necessary internal associations, and so on, it is just as convenient as that massive single compression you would make in a big data center. Therefore, if you look more broadly, of course, there are various reasons why it is difficult to do, but none of them are decisive for me.

It’s just, “Okay, yeah, you will have to solve a pile of problems that aren’t there in the vertically scalable option.” But I think it’s worth solving these problems, because I think the achievements you get as a result simply fully unlock certain opportunities. You no longer need to build an absolutely huge single data center. You could actually return and say that all the distributed data centers we have now can collectively accommodate a model 10, 100, or 1,000 times larger than the one data center that you can build.

This is such a breakthrough that otherwise it would simply be impossible. It requires a certain degree of optimism about the results of such approaches. But I think that’s what this is all about, and that’s why there are startups. You’re engaging in something risky, but it brings higher returns than any other approach you could choose.

And I think that’s right. The startup sphere and researchers in the academic environment are worth engaging in this. Companies that already have their own solution and product and are scaling it are much less likely to do that. This is, in essence, again, the innovator’s dilemma.

Logan Jastremski

Yes, yes. So, developing this topic, it was interesting to listen to statements from companies, mainly from SpaceX, but also to what Sam Altman and Dario Amodei generally talk publicly about: the necessity of continuing to scale the computational capacity to which they have access. It seems that SpaceX’s data center in Memphis said, “We have 2.5 gigawatts as of this year, or maybe by the end of the year,” and next year SpaceX plans to achieve a level of 10 gigawatts.

I think we used to perceive graphics processors like—remember, just like the 4080 or 4090, these consumer video cards—and then they, of course, migrated to data centers, and the question arose: “Okay, how many of these devices can we network?” But it was within limits, let’s say, not gigawatts but rather megawatts. So the order of magnitude here is completely different: We spend 50 to 70 billion for each separate gigawatt.

It seems like a race has started, and everyone is trying to get access to larger computational capacities. So my general question is this: If you look around, there are many computational resources in the world that may not be integrated or centralized in the form historically needed for model training. Perhaps these capacities are enough for inference, but how, in your opinion, can we better optimize computational resources that already exist in the world but are not used completely or have less than 100% utilization?

5. Trust and Verifying Machine Learning

Ben Fielding

This is a very appropriate question, because we fought with this problem for a long time. In our 2021 concept, we presented the Gensyn’s World approach, talking about the problem of computation and what decentralization could do, in particular about resolving issues of trust so that you could theoretically trust a random GPU located somewhere in another corner of the world.

Many people thought, “Hey, can we unite all these random consumer graphics cards in the world and use them?” There are certain things that can be done to involve them in the work, and we can discuss this a little subsequently. But if you go further, it becomes clear: Even if we solve these problems, if I can’t trust the system, I can’t build an economic model on it. I can’t pay the equipment owner, because I can’t trust him. So you can’t build anything on this foundation if there is no system of trust.

Logan Jastremski

Maybe I’ll quickly interject. When you say “trust,” do you mean the physical aspect—for example, the contracts and SLAs needed to guarantee uptime? Or do you mean that you tell me you have a B300, but in fact it’s some kind of Mac mini and you just lied to me? What is the difference, actually, in the concept of trust?

Ben Fielding

Ultimately, for us, these things are secondary to the real problem, which is that if I want to use your device, I want it to do something, and I need to know that it will do what I expect. Maybe it happens that you lied to me, saying that you have a different device from the one you actually have. But, honestly speaking, as a user of your device, I shouldn’t be concerned with its details. It’s enough for me to know only whether it can complete the task that I need.

So our approach to solving this question is that we already have it. We created all this technology, and it is already completely functional.

For various reasons, compute is less in the focus of our attention, and I'll explain why. But in the final result of this research, we reached the conclusion that you only need to know that what you sent to someone's device was completed. I don't need to know which device you have, or what your SLAs are. I only need to know that when I sent you a task, you really did it, and that I can be sure of this.

I don't have to rely on trusting you, your reputation, or something similar. I get cryptographic confirmation that says, “Yes, you sent this set of operations to Logan's GPU, and it fulfilled them.” Here is the result. After that, I can easily continue working, and the system has satisfied everything that interested me at that moment.

We saw how some protocols focus on periodically checking what is really in someone's specific GPU, the GPU they declared. We believe that this is treating the symptoms, not the root causes, and that you'll just be playing whack-a-mole, because there will always be an incentive for someone to bypass these limitations. They might inspect a user's graphics processor or something similar. Everyone can come up with ways to get around these systems.

But if you're able to prove the correctness of the operations performed, then you can no longer manipulate them. You have to perform exactly those operations. As long as you can perform them in the required time, it doesn't matter how you do it. That is already your concern. You are the infrastructure provider, so you can decide for yourself how to arrange it.

That gives you the freedom to buy different graphics processors and look for ways to increase efficiency, because you only have to satisfy the needs of the end user, who just needs you to complete the task. That's what we're aiming for. We consider this the ultimate state of trust, and we decided that our technology could help us achieve it.

We can reproduce machine-learning operations theoretically on any device. Right now, we do this on NVIDIA GPUs, AMD GPUs, x86 architecture, and MacBook processors. We can compare all of them. We can build optimistic proofs on this basis, and, if you want, you can even create zero-knowledge proofs from this evidence. It's just very slow.

But thanks to the ability to reproduce operations exactly, you can build all these proof systems, and then you don't need to know what GPU someone has because you can be sure that the operations were completed.

I touched on the topic of our research into the computations themselves. We thought a lot about this, and our idea was that if we solve the issue of trust, we can also solve the problem of efficiency. We would be able to build neural networks on distributed sets of GPUs and provide something that the centralized world cannot necessarily offer.

In practice, we found that continuing this work requires a lot of research, while the resources invested in centralized AI are enormous. Those systems can scale incredibly fast. It's extremely difficult to prove the advantage of a decentralized approach if you can't attract the same level of resources.

6. The Economics of Centralized AI

Even if centralized models are facing diminishing returns, their ability to keep increasing their capabilities is so great that they overcome these barriers. This brings me to the conclusion that serious problems await centralized AI. I think the labs know that the amounts they're investing in developing their models will hardly pay for themselves through their use in products.

The protective moat simply doesn't exist. For example, we constantly see successful distillation of open-source models. There are no real advantages to these huge models, whose creation cost tens of billions of dollars. At a certain point, continuing down this path becomes financially irresponsible.

Of course, the labs have ways around this, but they're locked in a race against one another. It looks like a game of survival: theoretically, the winner is the one who can last the longest while continuing to absorb resources and scale. But if you look at the big picture, there isn't much financial sense in it. It also has no environmental value.

There are other approaches, and if those resources were directed toward them, I'm sure they would work. But I'm not the one stuck in this race, right? From the outside, it's easy to say that. Apparently, it's very difficult when you have a company that you want to take to an IPO at the highest valuation in history.

Logan Jastremski

So, to summarize for our listeners: what you did and tried to do was extraordinarily difficult. Honestly, I was skeptical when we first talked about uniting various graphics processors all over the world, connecting them to one another, and performing interesting tasks when those GPUs aren't in the same place because of latency and bandwidth problems.

The world is big, and when you look at these data centers, as you already mentioned, many of them have interconnects for 400 gigabits or 800 gigabits, and now they're even reaching terabits. For example, NVLink provides many terabits of bandwidth, and it seems that all of this was intentionally designed to support high-speed data transfer inside the data center.

But I have to give you credit: you actually managed to do it. My second, more general question was: how much computational capacity could these networks provide if the technology really worked? How do you look at it from your perspective?

Because, according to your view and what you just mentioned, the centralized players are using brute force, and that will continue. Trillions of dollars are sent every year toward capital expenditures. All these new gigawatts are coming online. It seems that this year it was about 20 gigawatts, 30 gigawatts of new compute capacity will appear next year, and then 50 gigawatts in 2028.

That's an enormous amount of money. What do you think about this? Do you think there will be some kind of residual graphics-processing capacity? Do you think you need to be able to work with hyperscalers, in the sense that the main customers, at least for the moment, are not Groq, but OpenAI, Claude, and XAI? How do you evaluate the customer landscape? It seems that everything is being consolidated around those with the most computational resources, and that they obviously receive larger and larger profits over time.

7. Why Gensyn Moved Away from Compute

Ben Fielding

Yes, good question. Apparently, this leads us to the changes we've been talking about publicly, although perhaps we weren't entirely frank about them. As a company, we're no longer focused on compute. We solved a set of trust issues underlying machine execution and training on arbitrary devices around the world, theoretically on any device.

But if you return to our Litepaper from 2021, we clearly identified 3 basic resources for machine learning that, in our opinion, must scale for machine learning itself to scale. Compute was severely constrained when we first focused on it, and you could get a lot of value simply by giving someone access to another GPU somewhere in the world, because the huge cost of the centralized approach could be significantly reduced with decentralized compute.

I think the world of centralized compute has largely solved that issue. We've produced many more chips. More companies similar to neoclouds have appeared, and there's more competition. There wasn't much competition in this area before, but competition has increased, and we've seen the shift we described: most GPU consumption is concentrated in centralized hyperscalers performing specific tasks.

Decentralized compute simply won't scale in that way. It doesn't serve that purpose. If you're training a massive centralized model, you simply aren't going to do that with decentralized compute. You would have to train a completely different kind of model.

For now, there isn't really a client for such a completely different model. Therefore, anyone building a decentralized compute network must also be a client themselves and a research lab that builds the model. As a company—as a decentralized AI company—I think you could potentially attract the multibillion-dollar funding needed to become a lab that competes at the frontier with a decentralized model, but I still haven't seen anyone like that.

Honestly, I don't see much chance that someone will be able to attract the tens of billions of dollars needed to secure those resources through the centralized war in which the large, closed laboratories are engaged. The distributors of capital are concentrated in those large laboratories, and they continue allocating more capital to them. There's a concentration of resources.

Maybe a particular competitor will emerge, but that's a difficult path, and I still haven't seen anyone navigate it properly.

Logan Jastremski

So if you're not your own client, who is your customer segment?

Ben Fielding

For example, if you're doing decentralized compute, I think there's a segment for individuals. There are hobbyists and people who need access to a single GPU or a small number of GPUs. There are sovereign use cases where, if for some reason you can provide trust to those people and they still need to outsource computation, they'll use your system.

But for me, that's not enough at a large scale. It doesn't solve a big enough problem. It's more of an isolated, on-demand use case, and it doesn't scale machine learning, which is ultimately our main purpose.

The only place where I see an opportunity—and this is where we're directing our efforts—is the RL post-training wave, as you mentioned. I think there are ways to scale centralized compute in a decentralized way during post-training.

Pretraining is inherently faster if you execute it centrally. I think that's obvious. You can do this up to a certain size, where you're not seeing diminishing returns yet and you're receiving all the advantages of centralized compute, so it would be foolish to do otherwise. Another way is simply more expensive and doesn't offer any real advantages. Therefore, I believe it makes sense to do this centrally.

8. open-1b and Auditable AI Training

It's a small digression, but one important thing that we understood is that there are many advantages to decentralized compute that we and other companies talk about: transparency, trust, and opportunities for verification that you get with a decentralized approach. But you don't have to build everything in a decentralized way to get them.

Recently, we launched Open1B, which I think was mentioned earlier. Open1B is a completely transparent, verifiable, and reproducible model. It is trained so that anyone can check any part of the pretraining and see exactly how it was trained. It's the most transparent model ever released. It goes far beyond the boundaries of simple open weights.

It's much more than an open recipe: literally every operation can be reproduced and checked by anyone on any device. But it was trained in a data center, on a centralized cluster, because that's simply more efficient. So we separated out many of the advantages of decentralization that everyone was talking about and decided, indeed, that we could get these advantages without decentralizing the training itself. We only decentralize verification and auditing.

I think this is just the industry growing up and understanding what we're actually paying for. We understood that if someone wants to pay a little more for a reliable model, they can do so with centralized models using decentralized principles. This is a middle way that allows you to achieve your goals without resorting to full decentralization.

Again, I already said it was a small digression, but it led to a peculiar realization: we make these promises, but can we really fulfill them more effectively? Yes, we can. This is not a rejection of the principles of decentralization, because we keep them all—transparency and the opportunity for auditing—we simply use resources in the most appropriate manner.

I believe that decentralization still has plenty of space for what can be done in this direction. But going back to what I think is possible in decentralization, and to the thesis that stands behind this around compute, I think reinforcement learning after model training can quite possibly scale in this way.

A company could have a model that works on its clients' devices and use the benefits of learning from its customers' information, of course with their permission, using the equipment of its user base. The problem here is the product: it has to be designed so that such an approach makes sense. The technology already exists. We proved this using RL Swarm.

You can train models this way, using tens of thousands of devices. I think we had 40,000 devices that trained models together to solve 1 task in RL Swarm. But you need a product that can benefit from it.

At the moment, there isn't actually a product into which you can simply integrate this technology. You need a product where the technology is developed in parallel with it. It's still unclear what this product will be. Maybe it's programming tools, assistants, or something like open-source projects such as Hermes kernels.

I think there are possibilities here, but this should be built very consciously together with the use case. I believe that the company that handles this will create the functionality itself and develop a suitable product. You need to choose a task that seems a bit like a miracle. You have to solve the product problem and the technology research problem simultaneously.

We understood as a company that with RL Swarm, we're not in the best position for this. We're in a better position for something else: unlocking a resource that scales with decentralization—namely, information and data. We have a whole series of pipelines and infrastructure that facilitates this, and we're incredibly fascinated by it. But it doesn't really rely on decentralized compute anymore. It relies on decentralized information aggregation.

Logan Jastremski

So, maybe to reformulate, the initial thesis was somewhat refuted in the sense that distributed learning worked, but compared with centralized training, it was hard to get enough computational resources. It was probably a little more expensive in terms of access costs, or maybe even a little more expensive in terms of time, simply because it wasn't as fast as centralized training. And because the labs received more and more resources over time, training became centralized. They had higher throughput, so they could iterate a little faster and create smarter models.

9. Open Models and Closed Labs

Then you decided to change focus. That makes sense. How do you generally assess the landscape, even in the field of reinforcement learning, considering open-weight models and closed models? I think the main discussion now is that these advanced labs all have computational capacity. We have open-weight models that are actually pretty good; you can take these weights and apply reinforcement learning to them. But in general, unless you're running quite small models or models that have already undergone reinforcement learning, I think so far we're seeing that you still need a lot of compute.

When I look, for example, at OpenRouter and choose the 10 best models, even for the simplest model in the top 10, you need several H100s to actually run it with open weights. It's not even something you can buy a MacBook Pro—even the most expensive one—and run locally.

What if you're going to newer models in the GLM class? You understand: okay, you need about 40–50 H100 equivalents just to fit the weights, and that's even without taking context into account. What about agents? The context is also becoming longer and longer. Maybe you need to do some complicated things on the software side, or maybe you're just adding more hardware for problem-solving. So how do you see the development of open source and open weights compared to closed models, and what can you do with this?

Ben Fielding

Yes. Good question. As I said, large labs are locked in this race to improve capabilities. For a long time, and to this day, I suppose, they have continued to demonstrate ever-greater capabilities, holding the world captive in anticipation: “Oh my God, what's next?” But this is certainly significantly slowing down large-scale change, because in 2021 people didn't even realize that machines would be able to perform intellectual work at today's level.

The prediction was that next-token prediction could scale if we added more compute. In academic circles, we knew that we were using significantly less available data than we could have because we were compute-constrained. So it all came down to bottlenecks. When we promoted Gensyn's thesis about compute in 2021, that's exactly how it was. We understood that there was a lot of data, but our bottleneck was compute.

Therefore, if we could scale it, even if decentralized compute was more expensive, that would lead to huge gains because there would be more data to consume. But since then, the limiting factor has changed. Now labs are literally combing the planet in search of new data to train their models. They're building a terafactory in Texas, apparently trying to increase the world's computing power 10×.

You hear about secret facilities at Amazon where they buy books, cut off the spines, and scan everything, because they can get more data for training models. But what is it based on? What do you train on? You need data. You need information. Therefore, bottlenecks and constraints in machine learning are constantly shifting, and that's why in our litepaper in 2021, we divided the problem into 3 specific limitations.

The first is compute: your machine needs to be able to perform operations, and for model training you just need a bunch of operations. If you can't perform these operations, you're limited. You need more capacity to execute operations. But you also need data, which actually needs to be compressed. For some time, the whole internet was available for this. We said there was a lot of data to compress, and we just needed more compute.

There's also human labeling—or, generally speaking, ground truth. How do I put it? It accompanies the data; you need a certain signal together with the data for effective compression. There was a breakthrough in scaling the original GPT models and transformers in general, where we could use unsupervised learning to achieve huge results. Therefore, next-token prediction is quite a big breakthrough, because you don't actually need a label. In fact, it is its own label.

We had to employ people to label everything manually, and that was a bottleneck. You had a lot of data. You could probably get a bunch of GPUs in academic labs, but you had to gather people so they could sit and label all the data. Mechanical Turk was incredible at the time, sometime in 2019 or around then. I remember that.

Logan Jastremski

That's right.

Ben Fielding

And it was a bottleneck. It's interesting that we pass through waves where each of these 3 things, at different times, becomes an essential constraint on scaling machine learning. Our view in the litepaper in 2021 was that decentralization could solve each of these 3 problems.

We thought compute would be the first thing to scale quickly. That didn’t happen because the centralized world solved a lot of problems with compute. So, as a company, we simply switched to another one of these problems. We believe that combining data with an element of truth is what decentralization can solve now, and we believe this is the best direction to focus on.

At the same time, the machine-learning world passed through the problem of truth and moved from supervised learning to unsupervised learning and reinforcement learning. Instead of creating a bunch of labels, you give the system the possibility to create its own labels, defining reward signals and reward functions.

Now this is clearly visible in the idea that Lean itself becomes a reward function. You can teach a model to work in a structured system based on rules, and it will be able to do a huge number of rollouts if it has access to data in the environment. All of this changes things from the point of view of research, products, and technology, and every company is being guided by it in real time. I think it’s a wave.

Open source cuts across this by essentially destroying any moat that some people think it is possible to build for the long term. If you think, “Hey, now compute is the bottleneck. If I have a lot of compute, I can create the best model of all the others, and this will be enough for me for some time before the next bottleneck appears,” the problem with machine learning is that it is incredibly easy to distill and copy.

If you provide access to this model in any way, you can extract information from it, and that information has great value through the process of compression. It allows someone else to replicate the model much faster than you did it the first time. Open source actually acts as a competitive mechanism for anyone who is not a leader: It brings down the price, opens things up, destroys leadership, and lets them get ahead. Every lab now realizes this.

We hear a lot that China is creating these models from open source. I think Chinese companies understand that, while they are lagging behind the most advanced labs in the United States, it is profitable for them to open-source their code. As soon as they pull ahead, I think the open-sourcing will suddenly stop. Then American labs will suddenly start saying, “Hey, we’re really into open source. We’ve always been like that. We took a short pause, but we adore open source.” The incentives will change.

All of this prevents anyone from creating too large an artificial barrier, which, in my opinion, is actually very good for the rest of the world. This means that startups and others can stay just a little behind the advanced labs, because those labs spend billions creating models that then instantly become open.

10. AI Products, Customers, and Neoclouds

Logan Jastremski

So maybe you can forecast? You sound skeptical about closed advanced laboratories and more optimistic that models themselves will become publicly available goods, as you say, because they can be distilled and you won’t be too far behind the frontier. Maybe they will be closed for some time, and then someone else will release an equivalent model into open access. Are you more interested in things like neoclouds, or in people who have access to compute and perhaps the basic resources to start an LLM or an AI company, as opposed to those who build these clusters?

Ben Fielding

In the end, I think we’ll look back at this period, and it will be incredibly obvious what actually won. Ultimately, I think the best product and the best company will win—the one that can serve ordinary customers—and the race to prove some technical advantage or research advantage will be short-lived. Truly reliable protection is having a huge customer base that is ready to pay you constantly for your product.

I think OpenAI and Anthropic have clearly realized this and impressively changed the direction of their companies: from purely research breakthroughs to serving a new customer base all over the world that is not interested in the model itself. They don’t care about the architecture or how it works. They just want to use the thing to solve their everyday problems.

I think we’ll see this happen more and more often. The companies that will really succeed are the ones that manage to build and retain this customer base. I think this is one of the reasons why Meta is in such an incredibly advantageous position. In artificial intelligence, people say, “Meta had Llama, but then Llama 4 became a catastrophe. What’s going on? It seems like Meta is losing the race.” No. Meta has a huge customer base—an absolutely huge one.

These users are locked into the ecosystem. Meta made huge efforts to keep these people on its platform, and it has now proven that advantage with Muse, having released an attractive product for many people. You can see this in its communications: Meta doesn’t feel the need to do what I think other labs do, namely constantly emphasize, “Look how amazing this model is on research benchmarks,” because, in the end, that won’t bring victory to companies with market caps above $1 trillion.

That doesn’t bring victory anymore. It just attracts attention to you at the beginning. What wins is whether billions of customers are ready to continue paying you. One of the things you mentioned is neoclouds. I think neoclouds will become something that, when we look back on all this, will seem almost accidental.

You can exploit some inefficiencies in the computational-capacity market simply by being a neocloud right now. But realistically, I think this is a product, and it is leveling off because of economies of scale. We thought about it because it was a way we could have pursued our focus on compute. Instead of direct decentralization, we could have started with semi-centralized trading of computational capacity to generate income and develop as a company.

We considered it and told ourselves that we would fall into a trap. This is not a very good business. We didn’t even want to start this business, because when you do it, your company gradually begins to organize itself around solving that problem. You get stuck as a company, and after 5 years you wake up and realize, “We’re very good at serving a segment that we don’t need, and now we have to get out of this somehow.”

I think many companies that chose this attractive path of neoclouds already see that they can’t raise funding for another plan. Investors look at them and ask, “How are you going to increase this revenue? What is the total addressable market in all of neoclouds?” Because this is actually your main source of income, and you do all this other work that is more like research. But why do that if you have to increase your revenue from neoclouds? It becomes harder and harder. It’s like falling into a swamp.

We consciously didn’t enter this niche because it doesn’t attract us. It leads to a worse outcome. I’ve done podcasts for many years about why we don’t do it, and I’m very glad that we avoided it.

Logan Jastremski

This is funny. Maybe my algorithm on X is different from yours, but all I see in the feed are people who believe heavily in neoclouds. As usually happens on X, maybe it’s worth ignoring what everyone says. Maybe I’ll be a contrarian here.

Ben Fielding

Of course, someone can build a breakthrough neocloud. They can become winners. But then you’re just a commodity merchant. That may not be a bad place to be, but it’s not the kind of business I’m interested in.

I’m interested in scaling machine learning much further than that, perhaps beyond where we are now. I think we need to search for a specific transformational breakthrough. For a while, it seemed like that would be compute, although I don’t think it is compute.

I believe there are structural problems with how to provide machine learning to customers—for example, product problems, which I mentioned. There are many research problems that, even if they are solved, lead to product problems, because machine learning is now in the process of becoming a product. The way you bring it to the world is probably more important now than the advantages that can be obtained from various research breakthroughs.

Logan Jastremski

Yes, let’s talk about this. Then, perhaps, the last thing I’d like to touch on before we finish is the state of cryptography and XAI. But let’s talk first about data flows, as you mentioned, and why you’re focused on them now, Ben, because I think this is very interesting, and you have generally been at the vanguard of what is possible here.

Ben Fielding

Yes, thank you. No, we’re completely satisfied with how we managed this. For example, with the wave of reinforcement learning in post-training, I think with RL Swarm we saw what was happening, and we created a lot of technologies that could do what we saw was possible from a research point of view.

But we determined that this was impossible to turn into a product. There was no clear way to use it, so we went further. But I’m quite satisfied with what we saw and the potential it has.

11. From LLMs to Forecasting

The potential that I see now is precisely in predictive analytics—or, as I formulated it, intelligence for decision-making and judgment. I look at it this way: as I said before, in 2021, LLMs were machines for predicting the next token. That was the academic standard. You tried to improve next-token prediction, and it was mostly a research, theoretical, academic problem. OpenAI and others then understood that if you added scale, something very good would happen. That opened opportunities for them to create very good products.

Is it a coincidence that they came across these products? However you want to look at it, they turned it into a product. They transferred LLMs from the regime of next-token prediction to general intellectual work. So now, if you ask an ordinary ChatGPT user, they do not think about next-token prediction; that is a technical detail. They think, “This can perform intellectual work for me. It can create a program, write code, or write a recipe, and I can follow this recipe.” It can write a research paper for me.

It can collect a lot of information from the world, structure it, and put it in a form that is convincing to me. But if you look at what people use ChatGPT, Claude, and other systems for, they use them for things that need to be checked and require a person’s decision-making. You have to apply your own judgment to what it gives you. So it is not very capable of independently making important decisions. The reason for this, in our opinion, is that it does not have a proper mechanism for understanding cause-and-effect connections in the world.

It has a wonderful mechanism for assimilating huge volumes of data—text data, as well as other data that can be compressed into its latent space—but then it simply reformats that data and gives it to you. This is extremely valuable from a librarian’s point of view, but that is where its capabilities end. Our idea is: what if you stop thinking only about scaling LLMs and whether we can do more intellectual work, and think about what would now be beneficial to humanity? Then, in my opinion, it is the ability to navigate specific decisions.

Instead of simply reflecting recycled information back to you, a system could provide more information about the consequences of those decisions. If you think about such a problem, you see a field of research that receives much less attention: forecasting. Forecasting now—the idea of predicting future events—gets probably 10% of the attention that LLMs currently receive, but it is actually in a very similar position to where LLMs were in 2021. The things necessary for scaling already exist. There are no theoretical problems; there are engineering problems and problems with resources. If you can scale this, you can transform forecasting from an academic exercise into a much larger solution for the world, namely judgment and decision-making. This is what we are focusing on now.

12. Training Models to Make Better Predictions

Logan Jastremski

It is funny that you say that, because I heard Elon say that a real test of intelligence is, let’s say, accurate forecasting of the future. And, of course, the ability to think about complicated things and have a higher percentage of accuracy than others—that is what true intelligence is. I think Elon, despite everything, has a fairly high percentage of accuracy in his forecasts. Although his timeframe may need recalibration, he hits the mark often enough. So it is funny that you say that. How do you actually get data to have higher intelligence, in the sense of better decision-making or higher accuracy regarding future events?

Ben Fielding

That is exactly the right question, because the main problem with forecasting now is access to quality data to make the right forecast and then to train the system. The problem here is not the LLM, as it was in 2021, when it seemed that you simply needed more computational capacity to scale the system: you get capital and scale. The problem—and here it gets really interesting—lies in something else. The problem is that you need access to data in time, data that you can trust, but you do not know whether it is correct until the future comes. This data is very difficult to collect, because you cannot go back; the world has already changed.

Therefore, much forecasting research is unable to collect data in advance so that it can later find out whether it was correct. So you have to look at what has already happened and try to use it to confirm events that have just taken place. But when you do that, you encounter a difficult problem: understanding what was revised. Do I get access to the data that was there a month ago, or was it already edited, updated, changed, and distorted by the event itself?

This is one of the huge problems: cleaning information and data. Academic forecasting teams are working very hard and persistently on this large manual process—cleaning data and setting limits. For example: “If I use a search engine, can I limit it to only the data that existed a month ago?” Maybe a search engine will give you this, but is it true? You do not know. The only way to know for sure is to collect that information at that moment and then compare it with reliable data in the future.

Logan Jastremski

It is fascinating that this is incredibly similar to reinforcement learning. In fact, you have a very long reward cycle, where you generate sequences—fragments of information concerning a final result that will happen at a certain moment. You want a chain of these events. You want to say, “I think this event will affect that event in the future.” So you put this together now, then collect the next one, another one, and another one. When the event happens, you go back and get a reward.

This is a very complicated task because it is a global problem. This does not apply to the results of one company’s products or to some separate product lines. It takes place worldwide, for all information. So how do you actually establish the truth regarding the results of events in the world, trace that truth based on things that happened before, and relate it to these events?

Ben Fielding

It turns out that decentralization is the perfect mechanism for this, because within decentralization you have to solve a lot of problems around truth, transparency, and reliability so that people can trust it. By solving these problems and placing them on a blockchain, you create an immutable register that you can then use as a data source in the future. So you have created something that is impossible to fake. It is impossible to edit. It is impossible to change.

13. Delphi and Information Markets

Our position is that all the work we have done to establish truth during the execution of machine learning workloads, we are now directing toward establishing the truth regarding events that are happening in the world. The embodiment of this is a bit difficult, but for the outside world, the embodiment is that we have a project, Deli, which is an open, permissionless information market. Essentially, this is a market platform for forecasts. It is fully operational and decentralized. It allows anyone to create a market for a future event completely online, without anyone controlling it.

This market will be resolved fully automatically by means of a reliable, tested AI model when the event happens. Anyone else in the world can trade based on this event. People can make predictions and bet money that, in their opinion, it will happen. The money they invest is a kind of distinguishing factor between signal and noise. If you invest money in something, you must sincerely believe in it if you are a financially rational player. So we built a system that collects this information.

Logan Jastremski

So it is more of a prediction market, so to speak?

Ben Fielding

Essentially, yes. It is like a very, very big market-shaped forecast, but there is much more to it. That is why Deli is a market for information in itself. There will be an evolution of Deli that will make what it does much clearer to people.

In essence, we are saying that a lot of research is being done around the world on the elicitation and aggregation of information, especially in economics, at the intersection of economics and computer science. There are many methods we can use for information aggregation in the world, and then ultimately compare it with ground truth. We do this with the help of verifiable machine learning technology, which we have been building for many, many years. We can establish a source of truth that allows us to create much more effective, fully automated markets, but it also gives us a reinforcement-learning reward signal for a system that is training to forecast the future.

The whole system actually becomes one huge flywheel for improving predictive intelligence, and that is exactly what we are fascinated by. Similar to our early thesis about computation, we did not build a compute network because we wanted it to be a computer. As I already said, we could have just built a neocloud, downloaded a bunch of GPUs, and started giving access to them if we wanted to.

The reason we solved the problem of decentralized computation was that we wanted to scale machine learning. We wanted to build our own models based on it. We understood that this was not the best place to do that. So we moved to solving information problems, but the reason we care about the information problem is not to launch the next generation of prediction markets.

The mechanism seems to exist for prediction. It has market mechanisms, but we do this because it scales predictive intelligence into an entirely new era, because, in our opinion, this is the bottleneck for the next stage of machine learning.

Training these machine learning models will go beyond forecasting. They won’t just predict the future; they’ll actually become decision models that people use as a reference in everyday life.

Imagine you’re a poker player who’s used to playing online with a calculator. Chances are, you know the probability of each hand in real time because the site shows it to you. In real life, you don’t have that, and it’s much more difficult. You have to count everything by yourself.

Imagine that throughout your life, you had a poker-like calculator with probabilities for everything around you. You can have that, and that’s, in essence, a device that predicts the future. It can exist. The only problem is information, but we have a solution to it.

The blockchain aspect here acts as a source of truth for canonical events in the past, so that you couldn’t come back a month later and rewrite history, saying, “I actually said it was true,” even though you said it was a lie. Yes, this is a reward mechanism based on the source of truth. It turns out that it’s useful as a training system for machine intelligence, but it’s also useful for attracting people.

Why should anyone enter information into this system? Eventually, you can trace it back and ask, if you’re very cynical, “What’s the role of humans in this system?” In a certain measure, a person is like a sensor in the world who collects information about what’s happening and makes a judgment. You can encourage people to share this information, and blockchain is very good at creating those kinds of incentive mechanisms.

This essentially means that we can create a system where you bet on an event and receive a reward in the future if you’re right. We have a huge amount of internal, ongoing research that goes much further than simple prediction markets. I think prediction markets only superficially touch on what’s possible in eliciting and aggregating information.

Famous prediction markets focus on increasing trading volume, and the way they do that actually begins to undermine the elicitation and aggregation of information. There are incentives for trading that contradict the essence of data aggregation.

Our view is similar to our “neocloud” approach. The reason we didn’t go the way of computational neoclouds is that we didn’t want to do only that business. We don’t just want volume in prediction markets. We could scale that, but I’m not interested in it. I’m interested in information aggregation. To scale that, completely different market mechanisms are needed.

You’ll see how we’re launching new things. We’ve published articles, including a good article that has already come out and can be read. It’s our work on Evidence Markets, which shows that someone has actually made a forecast.

Logan Jastremski

Of course, this is interesting.

Ben Fielding

If you make a forecast and bet $10 on it, I know that, for some reason, you think it’s true. So I can assume that you have information confirming this event, but I don’t get that information from you.

Our article about Evidence Markets says that we can also get from you the information behind your forecast, and you can receive a reward for providing that data. There’s much more we can do: we can tell people that if, for some reason, they think something will happen and they have evidence, there’s a system that will actually buy that information from them.

That’s an absolutely fair exchange. You can participate in it. If you’re right, you’ll be paid a reward proportional to the information you provide. It’s a kind of fair system, but in the end, it buys information that no one else could buy. This is impossible to make centralized. It should be a decentralized system through which access can be obtained to information that would otherwise be unavailable.

Another way to look at this is to think about frontier labs. Many people know that when you interact with models from frontier labs, they use your activity to become better, and they don’t pay you for it. It’s just part of the deal. They provide access, and you use your data, like in the old SaaS era or the era of social networks. You use the product free of charge because they monetize your use in one way or another.

In the future, you can build systems where monetization returns directly to humans. If you inform the system of something that later turns out to be true, this human-generated reinforcement-learning process may reward that person directly. Therefore, it becomes a quasi-human-and-machine massive off-policy reinforcement-learning system that can share results between the people and machines learning there.

14. Where Crypto and AI Actually Overlap

Logan Jastremski

Again, that’s probably too abstract, but I think it’s fascinating. That’s very interesting. I think this leads me to my last question for you.

To a large extent, it seems that these markets bifurcate. It seems that many companies that may have started as crypto-AI companies have mostly reoriented themselves exclusively toward AI. What I’ve generally seen is that crypto is largely the next-stage evolution of fintech and finance.

You noted that many of these prediction markets are really focused on increasing volume because they take commissions from transactions. In a sense, I think that, from the cryptocurrency side, that’s perfectly normal. If crypto becomes the next evolution of capital markets and turns into a global foundation, that will be very interesting.

I’m also very optimistic about AI, in the sense that we’ll continue to see these larger clusters of intelligence. Instead of just developing software, you’ll have agents that perform operational tasks, accounting, and everything else we usually don’t want to engage in.

But at least at the moment, it seems that the intersection is shrinking, except perhaps in the information and coordination aspects. So my question is: how do you generally see the intersection of these industries?

From an investor’s perspective, many people will tell you they believe in the intersection of these 2 things, but when I ask them, they have difficulty determining what actually works. How do you, as someone who was truly a leader in this field and was at the vanguard for a certain period, think these 2 areas continue to come together? Is it perhaps a little different from what people first thought, or are they becoming more divergent?

Ben Fielding

Very, very good question. I think they’re getting closer, but in a very mature way—namely, by cutting off the things that don’t make sense and really narrowing the focus to where there is actually a decision at the intersection.

We’ve seen many companies with a philosophical affection for decentralization and a narrative attachment to it. At the core of all this lies a philosophical aspect: some companies—or, more precisely, the people who manage them—simply support decentralization philosophically, but they don’t understand where to apply it. They just think it should be there, and frankly, you can’t build a successful business on that.

I like that these people exist. I think the more people who share the philosophy of decentralization, the better. But if you want to build a successful business, there must be a real reason to use it, because it’s usually more expensive in many ways, but you get something you won’t find anywhere else.

Let’s say that, if you consider yourself the owner of a company, you specifically implement decentralization as a technology to do something and get something from it. If you don’t know what the hell you’re buying, why are you doing this?

I think a lot of companies have gone through this. Pragmatically, they may even reject decentralization: “In fact, we don’t want anything to do with decentralization.” Of course, it was our philosophy, but philosophy doesn’t pay the bills, so let’s go where the money is. Then you see how they completely leave this space.

What I don’t like is when companies never even felt this philosophy. They were there simply because there was a lot of noise around it. They seemed to demonstrate the philosophy only on the surface; actually, they didn’t even know that they had it. They used it to gain access to capital and attention, and then they left.

It’s hard to distinguish the real believers from those who aren’t, and in the end, to some extent, it doesn’t matter, since neither of these approaches achieves great results. But our view was, and always will be, that we consciously chose decentralization because we needed to solve a problem.

My co-founder and I deeply believe in the philosophy of decentralization, but we don’t allow it to interfere with pragmatically doing business. Here’s the point: we turned to decentralization because we saw a problem with trust. We had to solve it.

We tried to unite computational power between isolated systems that existed in companies and organizations. First, as a company, we created distributed machine training for banks, and then we said that we wanted to go beyond 1 bank. This problem arose: we couldn’t trust individuals. How does this work? That’s when we discovered blockchain and methods for using it.

We still adhere to this. We say that in situations where we need to solve global problems of trust, it’s the best solution we have, and we’ll use it. We also found many other things during the work—for example, the financial aspect. As you said, there are different reasons, from a financial point of view, why using decentralized networks and cryptocurrencies is very profitable.

But there actually are solutions that can take advantage of this without being related to cryptocurrency, for solving many such problems. For some companies, I would say that the financial aspects so far aren't enough to push them toward implementation; you also need to have a trust issue or something else. Right here, I think we see that the crypto narrative, when it penetrates the old company world, usually doesn't happen because it really solves their problems, but because the narrative became so popular that companies decided it solved enough problems that they were ready to try it.

But I don't think we've seen any real progress, except for stablecoins and things like that. We didn't see anything really better than a purely financial solution in many areas. At the end, to summarize why we still believe in decentralization: it's because it solves trust problems in a way we can't solve them otherwise.

For example, settlements in the markets, such as the way in which we are conducting calculations in Deli—you couldn't do this without decentralized networks. This should be decentralized to scale and function as we want. In addition to this, it gives us huge benefits from payment flows: opportunities to work with micropayments and opportunities for people to have an on-chain wallet for transfers. All this is an advantage for us, but it probably wouldn't be enough if we didn't also have trust issues.

So, in general, we are pragmatic: we consciously use this technology to solve problems that otherwise we can't solve. At the same time, as I already said, philosophically, I think that's right. The world must continue to evolve. But I look at things realistically: I think the world must do a lot of things. This doesn't mean that it will be so if there's no financial incentive.

Logan Jastremski

That is very true. Perfect. Well, probably this is what we are doing. Let's finish, Ben. I feel like we could continue for an hour, but it was an extremely interesting conversation, and I am grateful you came on the podcast to tell us about your worldview, about what Gensyn does, and also about the state of AI and crypto. It's very interesting; there's something to work on and think about, especially regarding predictions of the future.

15. The Next Wave of AI

Ben Fielding

Hmm, yes, there are many opinions here. Certainly. Absolutely. If I could leave people with the same opinion, I would want them to understand this: it seems like magic, but it is quite possible to have a “box” that predicts the future. And I mean literally, a model able to predict the future is better than people.

Before that, there are structural problems—systemic problems of engineering and scaling. But in essence, these are crypto-AI problems, because for real scaling and access to information, you need to understand decentralization and use it as infrastructure, and we are very happy to have accepted it.

But yes, that's right, maybe, although it sounds crazy. It's like if you told someone in 2021, “What is this cool new thing with which you can communicate on the internet that one day will perform a huge part of intellectual work?” The majority of people would say, “No, it's just something interesting. It's a cool chatbot, but it's not going to be able to do this.” But it was able to.

As for prediction of the future, we will come to this. It's completely, completely possible. We just need to zoom. I think it's an extremely exciting time for AI.

It also seems to me that I finally see how thinking in the spirit of “LLM— that's absolutely everything” is gradually receding into the past. These are really useful tools. They will be used in the sense of predicting the future. Without them, it would be impossible to build this. They're just a step; we will continue to build on their basis.

The LLM wave was not the only wave in AI. There will be others. And I think that prediction of the future is the next big wave. The future is exciting.

Logan Jastremski

Yes. Perfect. Thank you.

Ben Fielding

Thank you for having me. It was fun, as always.

Logan Jastremski

Likewise. Thank you.