[BidClub_]
Latent Space · · 69 分钟

Bee AI:可穿戴式环境智能体

Alessio FanelliswyxMaria de Lourdes ZolloEthan Sutin

YouTube
TL;DR
  • Bee 的核心押注是:个人 AI 的价值,来自被动积累第一视角上下文,而不是让用户反复向另一个聊天机器人解释自己。 记忆只是“最基础的一类用例”;更大的目标,是让助手充分理解用户的关系、态度和偏好,做到“已经知道”用户想要什么。Ethan 估算,一个人每天大约能产生 200,000个 token,并称自己的个人语料库约有 5,000万个 token。

  • 这款售价 49.99美元的可穿戴设备之所以存在,是因为极小的交互成本也会破坏始终在线的体验。 手机会独占麦克风,也会被来电打断;Apple Watch 用户在通话和充电后必须重新启动 Bee,因此有人几天内就买了专用设备。Bee 的续航最长可达 7天,同时作为手机的补充;Ethan 预计,未来大约 5年内手机仍将占据主导,直到类似 Orion 的廉价轻量眼镜出现。

  • Bee 的核心产品层是上下文引擎和执行能力,而不只是手环或转录功能。 它能够识别说话人、通过语义判断对话结束、生成摘要、提取语气和行动项,并将语音与 Gmail、Google Calendar、位置数据及网页搜索连接起来。WhatsApp 测试版可以识别可执行消息、提出帮助建议,并操作一部 Android 云手机;目前 API 提供只读和实时 socket 访问,外部行动能力则将在“很快”上线。

  • 尽管 Bee 并非为垂直工作场景设计,专业用户却已经成为最清晰的早期切入口。 公司称日发货量已超过 100台;Maria 认为 Texas 可能是其最大的州,Florida 也很突出。用户包括复盘对话的白领、询问面试表现的求职者,以及检查经营表现的餐厅运营者。Bee 希望提供通用的“理解型引擎”,再让第三方构建专业教练,因为“没有哪家创业公司能把所有垂直领域都做好”。

  • 隐私既是用户采用产品时的核心顾虑,也是尚未定论的法律问题和产品设计约束。 Ethan 区分了单方同意的 Nevada 与双方同意的 California;Maria 明确说明自己不是律师,但她认为,经过处理且不持久保存的音频,以及围绕用户生成的摘要,未必等同于保留录音,这在法律上是一个“灰色地带,尚未经过检验”。Bee 正在加入基于位置和上下文的禁用功能,因为用户可能喜欢产品,却仍会说:“我不能在工作时戴着它。”

  • 硬件路线优先考虑不显眼和续航,而不是技术上炫目的功能。 社区拒绝笨重的吊坠后,产品转向模块化手环或夹子;视觉功能则被推迟,因为图像采集和无线传输耗电过高,胸前摄像头又常常拍不到真正相关的视野,而且任何戴在脸上的东西“都必须够酷”。从原型走向量产仍是一次“巨大跨越”,涉及供应商、模具、法规,以及 4-6周的开模周期。

  • Token 价格下降削弱了最初的“剃须刀与刀片”商业逻辑,但语音处理和个人记忆仍是困难的技术问题。 语音活动检测可以避免转录无人说话的时段;摘要并不需要 Sonnet 级模型,但代理执行可能需要,并足以支撑订阅收费,R1 尚未完成评估。团队自行托管模型、微调 ASR,并使用定制的大规模并行检索,因为面对不断增长且充满噪声、事实会衰减或相互矛盾的语料库,“不存在一种通用的 RAG 方法能够奏效”。

  • 长期上行空间,是一个经授权、由个人智能体代表用户彼此协作的网络。 Bee 智能体已经展示过如何结合两名用户的位置、日历和偏好,选出一家法国餐厅;但一次送出约 20支蜡烛的实验,也说明授权和消费控制为何重要。Ethan 最有把握的判断是,“始终在线的 AI 真的会爆发”,创业公司和大公司都会逐渐发现它将如何改变日常行为。

摘要 · 为研究而整理的核心内容

1. 个人上下文,而不只是回忆,是 Bee 的产品核心

  • Maria 将 Bee 定义为“以第一视角与你共同生活的 AI”,持续捕捉现实生活中的上下文,从而能够回忆、反思,并最终在不反复询问用户偏好的情况下提供帮助。记忆“其实只是最基础的一类用例”;理想模型还应理解态度、欲望、关系和不断变化的处境。

  • Maria 表示,目前最明显的使用行为集中在陪伴和职业辅助上。她偏好的类比,是一个非常了解用户的最好朋友;主持人则认为,如果把这一愿景压缩成一张“用例”清单,产品听起来会比设想中的体验干燥得多。

  • Ethan 现在会随口统计“我今天输出了多少 token”,得出的结果约为 200,000个 token;他积累的个人上下文可能已有约 5,000万个 token。即便只是单日摘要,也让这个不写日记的人感到意外;一份“我的 2024 Spotify Wrapped”还提取出访问过 3个国家、发出的测试设备数量等指标。

2. 2016年失败的聊天机器人,让被动学习成为必选项

  • Ethan 第一次尝试个人 AI 始于 2016年,当时“还没有 transformer,甚至没有 BERT,只有 RNN”,正值 F8 之后人们短暂相信机器人会取代应用的时期。要实现有说服力的对话几乎不可能,但他和前联合创始人已经在尝试通过一款会提问并返回反馈的应用,模拟一个动态的人。

  • 这家公司加入了 Betaworks 的 Botcamp;当时最初的 Hugging Face 也在其中,只是一款会说“嘿,朋友,今天在学校怎么样?我们交换自拍吧”之类话术的青少年聊天应用。Ethan 认为,Hugging Face 为改进这款产品而构建了 Transformers 库,随后将其开源,最终这个库变成了更大的机会。

  • 青少年有时愿意没完没了地回答关于自己的问题,但大多数用户不会主动承担手把手教 AI 所需的劳动。Ethan 的团队后来通过 YC 转向病毒式消费视频,打造 Squad,将其出售给 Twitter,并帮助推出 Twitter Spaces;在一系列应用实验暴露早期方案的局限后,他们在 ChatGPT 出现前不久重新开始开发 Bee。

3. 删除日常微小摩擦,专用硬件才有价值

  • Ethan 判断任何可穿戴设备的标准是:“为什么这不能做成一个应用?”如果由手机持续采集,Bee 会独占麦克风、在通话时停止,还要求用户记得重新启动。那“一点点摩擦”足以实质性破坏一种智能体正与你共同生活的感觉。

  • 因此,Bee 支持 Apple Watch,作为无需额外硬件的即时试用方案,并通过后台运行设计控制耗电。但通话会打断采集,每日充电又会制造一次重新打开 Bee 的提醒和摩擦;Maria 表示,用户体验到价值后,会厌倦反复打开应用,有人仅仅用了几天就买了 Bee。

  • 拥有专用硬件后,Bee 可以控制增益、采样率和其他 Apple 有限框架无法提供的音频参数。挑战不在于录音棚级音质,而在于构建一个足够灵活的系统:它既要能应对嘈杂餐厅里的晚餐场景,又不能为了适配一个环境而调得过于激进,导致另一个环境无法使用。

  • 这款设备的定位是与手机协同,而不是复制 Rabbit 或 Humane 建立全新硬件平台的路线。Ethan 称这款可穿戴设备是“AI 的耳朵”,并预计手机会继续占据主导,直到一种廉价、轻量的新一代界面出现——可能是多年后的 Orion 类眼镜。

4. 只有能够在别处执行行动,采集才真正有价值

  • Bee 的应用会持续展示当下发生的事情,然后将一天的内容转换成可读的对话片段。说话人识别可以避免把他人的话误记成用户的事实;通过语音教学,系统能够识别亲近联系人;语义终点检测则用来划分边界本就模糊的不同对话。

  • 对话结束后,更大的模型会进行分析和摘要,提取要点、氛围、语气和可能的行动;日终视图再将这些内容与位置数据结合起来。用户事实始终可见、可编辑,在不确定观察结果变成可信个人上下文之前,保留了一层人在回路中的校验机制。

  • 回忆界面可以搜索记忆、Gmail、Google Calendar 和网页,并将综合结果关联回原始对话。Ethan 演示了询问一次 Taiwan 制造之旅,系统返回了相关细节、生产问题和讨论内容,而不是一段没有依据的回忆。

  • 在 WhatsApp 测试版中,Bee 会监测通知并提出两个问题:这件事是否重要到值得提醒,以及它能否提供帮助?一次餐厅请求触发了建议,用户可以通过聊天或按住说话接受;随后,一部 Android 云手机找到 WhatsApp 对话并发出了答案。手机处于非活跃状态时,可以通过“Hey Alfred”唤醒。

5. API 将个人上下文变成基础设施

  • Alessio 对开发者场景的判断很直接:没有任何一款消费应用能够预判他的全部需求,因此用户必须能够拥有、纠正并重新处理自己的数据。他将 Bee 与其他没有 API 的可穿戴设备作了对比,swyx 则宣布:“我们家都是 API 爱好者。”

  • Bee 默认不保存音频,但其 API 可以暴露经过处理的上下文,并提供实时 socket 访问。对话期间访问权限仍是只读;Ethan 表示,写入支持和外部行动能力会随更广泛的行动系统一起“很快”全面开放。

  • 更大的产品愿景“并不只是一个工具”:Bee 应该在合适的时机主动提出有用行动。Ethan 承认其中的门槛很难把握——不断插话的助手令人无法忍受,而每次重要机会都错过的助手又几乎没有价值;他认为,必须将强大的推理能力与个人上下文结合起来,才能找到平衡。

6. 隐私既是法律灰区,也是社会转型

  • 当被问及即时回放能否解决争论或揭穿谎言时,Maria 强调,AI 可以成为对事情经过以及对话情绪的“客观视角”。相比打造一个追究每个矛盾之处的机器,她更关注如何支持那些被愤怒或悲伤扭曲判断的用户。

  • 讨论也区分了合理变化与欺骗:人在不同语境下会说不同的话,也会随着时间改变。Bee 目前允许用户批准、拒绝或编辑系统提出的事实;团队希望最终实现自动化,但并不假装仅凭充满噪声的语音就能确立持久事实。

  • Ethan 区分了单方同意的 Nevada 与双方同意的 California,并指出公共场所具有不同的隐私预期。Maria 明确说明自己不是律师,但她表示,由于 Bee 会处理音频却不持久保存,并储存围绕用户生成的摘要而非声音,是否在法律上构成录音仍是“一个灰色地带,尚未经过检验”。

  • 在录音合法的地方,伦理问题仍然存在。Maria 表示,包括注重隐私的 Italy 在内,披露 Bee 后没有人要求将其关闭,但用户会提到不自在的伴侣和机密工作场所。计划中的防护措施包括地理围栏和“上下文围栏”;Maria 预计,随着社会经历从现在到未来 5年左右的转型,相关规范会像门铃摄像头那样逐渐改变。

7. 没有垂直市场路线,专业用户却自然出现

  • Maria 表示,Bee 每天发货超过 100台,客户并不主要集中在 San Francisco 的早期采用者圈层。她认为 Texas 可能是最大的州,Florida 也很突出,同时发现大量需求来自“靠说话谋生”的白领专业人士。

  • 一名用户在求职面试中使用 Bee,随后询问:“你觉得我的面试表现怎么样?我应该怎样改进?”其他场景还包括餐厅从业者检查经营表现,以及用户复盘自己是否发挥良好——这是一种环境式个人教练,而不只是会议转录工具。

  • 当主持人将这一切入口与 Gong 等垂直工具进行比较时,Maria 不愿把 Bee 变成每一种专业应用。公司希望构建底层理解引擎,再通过 API 让第三方开发;消费行为过于不可预测,无法提前断言最终的杀手级功能会是什么。

8. 续航和可穿戴性优先于摄像头与炫技

  • 在社区成员,尤其是已经佩戴项链的女性表示不会使用后,Bee 放弃了笨重的圆形吊坠。最终形成的模块既可以做成手环,也可以做成夹子,目前有黄色和黑色,未来还可加入其他颜色,甚至重新推出吊坠设计。

  • Maria 认为,早期版本的续航约为 35小时,而当前设备可达 7天。她强调产品应带来这样的反应:“我喜欢戴着它,然后忘记它的存在”;因此,讨论优先考虑更小、更轻、更省电的硬件,而不是为了追求类似 Apple 的薄度而追求薄度。

  • swyx 欣赏 Humane 的工程能力,但认为它的重量、发热、可更换电池和激光界面,并没有比手机更好地解决一个问题。MagSafe 录音设备如 Plaud 可以搭载大容量电池,但把麦克风连接在口袋里的手机上,会带来另一种采集难题。

  • 视觉功能被推迟,是因为即便低帧率采集和无线传输也会消耗大量电量,而胸前佩戴的位置常常拍不到佩戴者正在看到的内容。swyx 将显眼的 Snap Spectacles 与几乎像普通眼镜的 Meta Ray-Bans 作对比;Ethan 补充说,目前创业公司的眼镜仍面临续航过短的问题,还没有达到人们愿意每天佩戴的程度。

9. 制造让硬件迭代从根本上变慢

  • Ethan 称,从原型到量产产品的跨越“巨大”。固件和电子部分越来越容易上手——即便是陌生的固件,也能在 Claude Sonnet 的帮助下开始开发——但量产还要增加采购、法规、物料清单、成本控制和外壳模具等环节,而模具可能需要 4-6周才能完成,届时缺陷才会暴露。

  • Maria 建议通过其他创始人选择供应商,让塑料、PCB 等不同需求匹配各自的专业厂商,并亲自拜访以建立关系。Alessio 介绍说,团队早期为了快速制作电路板,使用了成本高昂的 Los Angeles 本地生产;他认为中国原型制造在时间和价格上越来越有竞争力,而 Bee 最终在 Taiwan 完成生产和组装。

  • Bee 参加 CES 时计划停留 3天,但没有购买展位,而是利用约 80,000-90,000人的人流接触媒体、合作伙伴和供应商。一个 10×10的展位在展示成本之外约需 5,000美元;如果想脱颖而出,成本可能达到 6位数。不过,供应商会面仍带来了体温采集、太阳能和动能等方案,其中太阳能看起来更符合 Bee 的电力需求。

10. 推理成本下降后,难题转移到了记忆

  • Ethan 认为,约 250,000个输入 token 已经不再天然昂贵,因为 token 价格仍在持续下降。语音转文字更难:它既需要实时处理,也需要随后交给更大的模型处理,因此系统会先用廉价的语音活动检测,过滤掉一天中绝大多数无人说话的时间。

  • 摘要不需要 Sonnet 级模型,但可靠的代理执行可能需要;Bee 预计会针对这些昂贵能力收取订阅费。团队自行托管模型并微调 ASR;R1 尚未完成评估。

  • Alessio 承认自己对商业模式的判断错了:他原本预计通过低价硬件和持续性的“剃须刀与刀片”收入获得补贴,但他提到 Friend 和 Limitless 的一次性售价分别为 99美元,而 Bee 的售价为 49.99美元。推理成本快速下降,使消费硬件的经济模型问题与他最初形成判断时相比发生了实质变化。

  • Ethan 预计,“所有 ASR,所有语音转文字”很快都会过时,因为端到端模型会吸收今天这条繁琐的处理链路;不过,在蒸馏之前,它们可能需要更多算力。Bee 同样发现,通用 RAG 不够用:记忆会衰减,传统 embedding 和 RAG 在不断增长的个人语料库上表现不佳,而且“没有一种通用的 RAG 方法能够奏效”,不可能脱离数据本身独立工作。

11. 经授权的智能体交换,是长期上行空间

  • Bee 构建了定制化的小模型大规模并行检索,并将用户确认的事实与更模糊的推断分开存储。Ethan 认为,知识图谱可能是保存这些数据的正确方式,但如何在不让 LLM 负担过重或产生混淆的情况下,检索并格式化图谱数据,仍然困难。

  • 理想系统必须推断用户从未明确说出口的行为。例如,它可以注意到收据和重复订单,从而理解“从这家店点点东西”其实意味着购买用户平时会点的那份餐。要做到这一点,系统必须跨越对话、邮件、日历和行动进行推理,而不是简单地将明确偏好追加到记忆清单中。

  • 2个 Bee 智能体已经结合日历、位置以及双方对法国菜的共同偏好,在 Pacific Heights 附近选出一家新餐厅。未来还可能出现社交交换、家庭动态、人格推断和约会应用;但一次向 Maria 发送约 20支蜡烛的智能体实验说明,必须明确控制智能体可以披露什么、购买什么。

Alessio Fanelli

Hey everyone, welcome to the Latent Space podcast. This is Alessio, partner and CTO at Decibel Partners, and I'm joined by my co-host, swyx, founder of Smol AI.

swyx

Hey, and today we are very honored to have in the studio Maria and Ethan from Bee. Welcome.

Maria de Lourdes Zollo

Hi. Thank you for having us.

Alessio Fanelli

You are, I think, the first hardware founders we've had on the podcast. We've been looking to have an AI wearable hardware founder on for a while. I think we're going to have 2 or 3 of them this year, and you're the ones whose product I wear every day.

swyx

Thank you for making Bee.

Maria de Lourdes Zollo

Thank you for all the feedback and usage.

swyx

I've been a big fan. You were the speaker gift for the Engine No. 1 Sphere. Let's start from the beginning. What is Bee Computer?

Maria de Lourdes Zollo

Bee Computer is a personal AI system. You can think of it as AI living alongside you in the first person. It can capture your real-life context, and with that understanding, it can help you in significant ways.

The obvious one is memory, but that's really just the base use case: recalling things and reflecting. I know, swyx, that you like the idea of journaling, but you don't want to do it, while still having some kind of reflective summary of what you experienced in real life.

It's also about having the whole context of a human being and giving the machine the ability to understand what's going on in your life—your attitudes, your desires, and specifics about your preferences.

Ethan Sutin

That means it can not only help you with recall, but also with anything else you need it to do, because it already knows what you would want. Think about somebody you've worked with or lived with for a long time. They just know, without having to ask you, what you would want.

swyx

It's clear that this is the future. Personal AI is just going to be much more valuable with personal context.

Maria de Lourdes Zollo

One of the things we're really passionate about is understanding this personal context, because it will make AI more useful. Think about a best friend who knows you so well. That's one of the things we're seeing from users: they use Bee from a companionship standpoint and for professional use cases. There are many ways to use Bee, but companionship and professional use are the ones we're seeing the most right now.

swyx

It feels so dry to talk about use cases.

Alessio Fanelli

Yeah, well.

swyx

It's really an investor question. But at its best, don't you want your AI to know everything you've said and everywhere you've been? Wouldn't you want that? You shouldn't have to repeat every time what you like. It already knows that, and it does things for you based on that. I think that's really cool.

Great. Do you want to jump into a demo? Do you have any other final questions?

Alessio Fanelli

Before we do that, maybe we should cover the origin story. How did you two meet? Was this the first idea you started on, or was there something else before?

Maria de Lourdes Zollo

I can start. Ethan and I have known each other for 6 years. He had a company called Squad, and before that it was called All About, which was a personal AI company.

Ethan Sutin

Yeah, maybe you should start with that.

Maria de Lourdes Zollo

That's how I know Ethan. He was pivoting from personal AI to Squad, which was a co-watching-with-friends product. I had experience working with TikTok and video content, so I helped with the pivot, and we launched Squad.

That was really successful. In the end, the founders decided to sell it to Twitter, now X. Both of us joined X. We launched Twitter Spaces and many other products, and we continued working together until we started Bee.

Ethan Sutin

The interesting thing is that this isn't our first attempt at personal AI. In 2016, when I started my first company, it started out as a personal AI company. This was before transformers—before BERT even, just RNNs. You couldn't really do any convincing dialogue at all.

I met Esther, who was my previous co-founder, and we were both really interested in the idea of having a machine model and understand a dynamic human. We wanted to make personal AI. This was more geared toward younger people, because we had much more limited tools at the time.

I don't know if you remember, but in 2016 there was a brief chatbot boom. It was way premature, but it was when Zuckerberg went up at F8 and talked about the Messenger platform. People thought, “Oh, bots are going to replace apps.”

That lasted for about 6 months, and then everybody realized that these things were terrible and weren't replacing apps. But that's when we got excited and tried to make something where you could teach the AI about yourself. It was just an app that you chatted with. It would ask you questions and then give you some feedback.

What we learned was that some people really loved chatting and answering questions, but it was a lot of work just to manually teach an AI all these things about you. We had sentence similarity and other things we could try, but the technology was premature at the time, so we pivoted.

Alessio Fanelli

How did the first version get launched?

Ethan Sutin

We started in the same office as Hugging Face because Betaworks was our investor. They had a program called Botcamp. Betaworks is a really cool VC because they invest in things that are out there and way ahead of everybody else.

At the time, Botcamp had 6 companies. It was us and Hugging Face, and I think the other 4 are dead. Hugging Face was the one that really took off. A 30% success rate is pretty good.

It was just the 2 founders. It was a chat app for teenagers. A lot of people don't know that Hugging Face was originally like, “Hey friend, how was school? Let's trade selfies.” They built the Transformers library, I believe, to help make their chat app better. Then they open-sourced it, and it blew up. They realized that this was the opportunity, and now they're Hugging Face.

Maria de Lourdes Zollo

There were some people who were super passionate about it. Teenagers, for example, really like to talk about themselves, so they would reply to a lot of questions and talk about themselves. But most people don't really want to spend that much time talking.

Ethan Sutin

We went through Y Combinator. Long story short, we pivoted to consumer video, and that went really viral and got a lot of usage quickly. We ended up selling it to Twitter, worked there, and left before Elon—not related to Elon, but we left before that.

Alessio Fanelli

I should mention that this was the famous time when Elon had just come in. Esther was the famous one.

Ethan Sutin

Yes, she was my former co-founder. She was the one who was sleeping in the office. She stayed. We had left by that time, but she stayed for a while. Later, she left, or I think she got laid off. I think the whole product team got laid off. She was a director of product.

Then we thought, “Oh my God, things are different now.” We really started working on Bee again right before ChatGPT came out. We had an app version and were trying different things around it, but ultimately it was clear that there were some limitations.

A good question to ask any wearable company is: why isn't this an app?

Maria de Lourdes Zollo

Because we tried the app at the beginning.

Ethan Sutin

The idea behind Bee comes from “ambient.” If it were just around you all the time, rather than requiring you to open the app and make the effort to enter data, that led us down the path of hardware.

The sensors on this are microphones, so it's capturing and understanding audio. We started with hardware that also had a vision component, and we can talk about why we're not doing that right now.

If you wanted continuous audio understanding with your phone, it would monopolize your microphone. It would get interrupted by calls, and you'd have to remember to turn it on. That little bit of friction is actually a substantial barrier to the experience of having it with you all the time and living alongside you.

We do have Apple Watch support, so anybody with an Apple Watch can use Bee right away without buying any hardware. We worked really hard to make a version for the watch that can run in the background without draining your battery too much.

Even with the watch, there's still friction because you have to remember to turn it on, and it still gets interrupted if somebody calls you. You have to remember to go back and turn it on. We send a notification, but you still have to go back and turn it on because that's just the way watchOS works.

Maria de Lourdes Zollo

One of the things we're seeing from our Apple Watch users is that they like the Apple Watch integration. Many people start using Bee from the Apple Watch, and after a couple of days they buy the Bee because they just like wearing it.

We're learning from that, and it's really cool.

Ethan Sutin

Fundamentally, we like to think that personal AI is the mission. It's about understanding, connecting the dots, and making use of the data to provide some value.

The hardware is the ears of the AI. It's about integrating the incoming sensor data, and that's really what we focus on. If we can do it well and have a great experience on the Apple Watch, that's just great.

Ethan Sutin

I mean, but there are some platform restrictions that existing hardware makes it hard to overcome.

Alessio Fanelli

What do people do in 2 or 3 days that then convinces them to buy the product? This was a product where, after you use it for a while, you have enough data to start to get a lot of insights.

Ethan Sutin

For Apple Watch users, I believe that's because every time they receive a call, they need to go back to Bee and open it again. Or, for example, every day they need to charge the Apple Watch, and it reminds them to open the app every day. They feel like, “Okay, maybe this is too much work. I just want to wear the Bee, keep it open, and that's it. I don't need to think about it.”

I think they see the potential just from the watch, because even if you wear it for a day, we send a summary notification at the end of the day about the key things that happened to you during your day. I didn't even think—I’m not a journaling-type person. I was like, “Oh, I just lived the day. Why do I need to think about it?” But it's actually pretty interesting. Sometimes I'm surprised by how interesting it is to me just to be like, “Oh, yeah, that,” and see how it fits together. I think that's something people get immediately with the watch, but they're like, “Oh, I'd like an easier way to do this.”

swyx

It's surprising because I only know about the hardware. But I use the watch as a backup when I don't have the hardware. I feel like, because now you're beamforming and all that, this is significantly better.

Ethan Sutin

Yeah, that's the other thing. We have way more control over the hardware. With the Apple Watch, you're limited: you can't set the gain, you can't change the sample rate, and there's very limited framework support for doing anything with audio. Whereas if you control the hardware, you can optimize it for your use case.

The Apple Watch isn't meant to record this, and we can talk, when we get to the part about audio, about why it's so hard. This is audio at the hardest level because you don't know what environment it has to work in. This environment is great—we're in a studio—but afterwards, at dinner in a restaurant, it's a totally different audio environment. There are a lot of challenges with that. Having really good source audio helps, but there's still a lot that has to be done with machine learning to account for it. You can tune something for one environment or another, but it will make one good and the other bad. Making something flexible enough is really challenging.

swyx

Do we want to do a demo just to set the stage, and then kind of talk about it?

Ethan Sutin

Yeah, I think we can walk through the product.

For listeners, we'll be switching to video that will superimpose onto this video. If you want to see, go to our YouTube. Like and subscribe as always. And buy the B. Yes, and buy the B. While you wait for the video to come out, buy the B. Maybe we should have a discount code just for the listeners. Sure. If you want to offer it, I'll take it. Yeah. All right. Discount code SWX.

An important thing to mention is that the hardware is meant to work with the phone. If you look at Rabbit or Humane, they're trying to create a new hardware platform. We think that the phone is just so dominant, and it will be until we have the next generation, which is not going to be for 5 years. Maybe some Orion-type glasses that are cheap enough and light enough will come along, but that's going to take a long time. So we work with the phone rather than trying to replace it.

In the app, we have a summary of your days, but at the top is what's going on now, and that's updating continuously. Right now, it's saying that I'm discussing the development of personal AI, and that's just the ongoing conversation. Then we give you a readable form with little segments of the important parts of the conversations.

We do speaker identification, which is really important because you don't want your personal AI thinking you said something and attributing it to you when it was just somebody else in the conversation. You can also teach it other people's voices, so if there's somebody close to you, it can start to understand your relationships a little better.

We do conversation endpointing, which is a task that didn't even exist before because nobody needed to do this. If you had somebody's whole day, how do you break it into logical pieces? We use not just voice activity, but other signals to try to split it up, because conversations are a little fuzzy. They can lead into one another, and one can start before the next. We also use the semantic content of the conversation.

When a conversation ends, we run it through larger models to try to get a better sense of what was actually said, then summarize it, provide key points, describe the general atmosphere and tone of the conversation, and identify potential action items that might have come from it. At the end of the day, we give you a summary of your whole day, where you were, and a step-by-step walkthrough of what happened and what the key points were.

That's the base capture layer. If you just want to get a glimpse, recall something, or reflect, that's there. But really, the key is that all of this now feeds into generating personal context about you. We generate key facts known to be true about you, and there's a human-in-the-loop aspect: you have visibility into that. I have a lot of facts about technology because that's basically what I talk about all the time, but I also have some hobbies that show up.

I measure my day now by asking, “What is my token output of the day?” As a human, how much information do I produce? It's measured in tokens, and it turns out to be around 200,000 tokens a day.

In the recall case, we have a chat interface. The key here is recall: I probably have 50 million tokens of personal context, so how do you make sense of that and make it useful? I can ask simple recall questions, like details about a recent trip I took to Taiwan, where we were with our manufacturer. In real time, it has various capabilities, such as searching through my memories, searching the web, or looking at my calendar. We have integrations with Gmail and Google Calendar, so it can connect the dots between in-real-life and digital life. I just asked it about my Taiwan trip, and it gives me a breakdown of the details, what happened, and the issues we had around certain manufacturing problems. It also goes back and references the conversation, so I can return to the source.

Maria de Lourdes Zollo

Yeah, and not just the conversation, but the integrations as well. We also have Gmail and Google Calendar. If there's something there that's useful for more context, we can see that.

I never use the word “agentic” because it's strange, but it can search through your conversations, search through email, and look at the calendar. If I'm brainstorming about something that spans across all of those, it can search through my conversations, search through email, look at the calendar, and then, depending on what's needed, synthesize something with all that context.

swyx

I love that you did Spotify Wrapped. That was pretty cool.

Ethan Sutin

Yeah, one thing I did was make a Spotify Wrapped for my life in 2024. It's surprisingly good. It gave me metrics like, “You visited 3 countries and shipped X many beta devices.” It gives a lot of personal insights and reflection points.

swyx

That's fascinating. So that's the demo. We can show something that's in beta if we want to.

Ethan Sutin

The vision is not just about AI being with you, passively understanding you through your experiences, but also proactively suggesting things to you at the appropriate time. It's not just a tool; it can step in and suggest things to you.

Speaker 2

You’re asking for a recommendation for an Italian restaurant. Would you like me to look up some highly rated Italian restaurants nearby and send her a suggestion?

Maria de Lourdes Zollo

What I did was just send Ethan a message through WhatsApp on his personal phone.

Ethan Sutin

Basically, Bee is watching all my incoming notifications, and if a notification meets 2 criteria—is it important enough for me to raise a suggestion to the user, and is there something I could potentially help with?—this is where the actions come into play.

Because Maria is my co-founder and because it was a restaurant recommendation, something Bee could probably help with, it proposed that to me. I can respond either through the chat or with another push-to-talk, walkie-talkie-style button. It's actually a multipurpose button to toggle Bee on or off, but if you push and hold it, you can talk. So I can say, “Yes, find one and send it to her on WhatsApp.”

Bee is an Android cloud phone, so it has access to all my accounts.

We're going to abstract this away, and the execution environment isn't really important. We can go into technically why Android is actually a pretty good one right now. But it's searching for Italian restaurants, and we don't have to watch this. I could have my AirPods in my ears and my phone in my pocket. It's going to go to WhatsApp, find Maria's thread, send her the response, and then let us know.

Alessio Fanelli

Oh God. What's a good suggestion? I mean, it's not an Italian restaurant.

Ethan Sutin

Yeah. What did you say?

Alessio Fanelli

It's easy to say.

Ethan Sutin

Exactly. It's easy to say. Successfully found and shared. Let's see what the AI says.

Alessio Fanelli

What does the AI say?

Ethan Sutin

Bottega. I think it's called Bottega. It said it twice. I've been to one called La Cocina, I think. That was good.

Alessio Fanelli

Beretta's on Valencia Street. It's fine. The pizza isn't good.

swyx

It's not good?

Alessio Fanelli

Some of the pastas are good. I'm sorry—Beretta's, sorry. But there's this place, Delfina. Everybody here says La Pizzeria Delfina is amazing. I'm like, this is not—I don't know. It's great. North Beach Cafe.

swyx

For the record, since you're all Italians, what's the best Italian restaurant in SF?

Maria de Lourdes Zollo

Oh my God. I feel like I don't have one. No. I'm not sure.

Alessio Fanelli

The place you took us with Michele last night.

Maria de Lourdes Zollo

Vega.

Alessio Fanelli

The guy at Vega just happened to be Italian. It's in Bernal Heights.

Maria de Lourdes Zollo

He's nice.

swyx

What's the name of the place?

Alessio Fanelli

Vega.

swyx

Vega.

Alessio Fanelli

Okay. Cool, cool. We got the name.

swyx

But Vega—is that in Italian? What does it mean in Italian?

Maria de Lourdes Zollo

Vega means... It doesn't mean anything to me.

Ethan Sutin

The last thing I'll mention is also the wake-word detection. The phone can be off, and you can just say, “Hey, Alfred,” and then it just...

Alessio Fanelli

Yeah. Being able to have wake words enables some form of voice-agent features that even ChatGPT can never have, because they don't have the recording layer.

Ethan Sutin

Yeah, I think we have some other ideas, even beyond wake words, but I think it's interesting to see how people use the voice side of it. We're going to see a lot of innovation around hardware and stuff, but the real core is being able to do something useful with the personal context.

You always had the ability to capture everything. We've always had recorders, camcorders, body cameras, stuff like that. But what's different now is that we can actually make sense of and find the important parts in all of that context.

Alessio Fanelli

And then one last thing—I'm just doing this for you—is that you also have an API, which I think I'm the first developer advocate, because I had to build my own app.

Ethan Sutin

We need to hire a developer advocate.

swyx

Or just—yeah, hire AI engineers.

Ethan Sutin

The point is that you should be able to program your own assistant.

Alessio Fanelli

I tried Omi, the former Friend, the knockoff Friend. Real Friend doesn't have an API, and Limitless also doesn't have an API. I think it's very important to own your data and be able to reprocess your audio, although by default you don't store audio.

Ethan Sutin

Yeah.

Alessio Fanelli

And then also just to do any corrections. There's no way that my needs can be fully met by you.

Ethan Sutin

Yeah.

Alessio Fanelli

The API is very important.

Ethan Sutin

Yeah. I've always been a consumer of APIs in all my products.

swyx

We are API enjoyers in this house.

Ethan Sutin

Yeah, I am. It's very frustrating when you have to go build a scraper.

Alessio Fanelli

This whole combination—you have my location, my calendar—it's really, for me, the sort of personal assistant: just to write into it or to have it take action on external systems?

Ethan Sutin

We're expanding it. It's right now read-only. In the future, very soon, when the actions are more generally available, it'll be fully supported in the API.

Alessio Fanelli

Mhm. Nice. I'll buy one after the episode. The API thing, to me, is the most interesting.

Ethan Sutin

We do have real-time APIs, so you can even connect a socket and connect it to whatever you want it to take actions with.

Alessio Fanelli

Yeah. When I look at these apps—and there are so many of these products being launched—it's great that I can go on this app and do things, but most of my work and personal life is managed somewhere else. So being able to plug into it is nice.

swyx

I have a bunch of more human questions that I think people might have. One is: Is it good to have instant replay for any argument that you have? I can imagine arguing with my wife about something. There are these commercials now where it's basically 2 people arguing, and they can throw a flag, like in football, and have an instant replay of their conversation.

I think this is similar, where people can almost not argue anymore or lie to each other, because in a world in which everybody adopts this. I don't know if you've thought about it. And also, all the lies that all of us tell—there are sometimes things that contradict each other, because I might say something publicly and think something privately that I tell someone else. How do you handle that when you think about building a product like this?

Maria de Lourdes Zollo

I would say that I like the fact that the AI is an objective point of view. I don't care too much about the lies, but I care more about the fact that it can help me understand what happened and the emotions in a really objective way—a really critical and objective way.

If you think about humans, they have so many emotions. Sometimes something that happened to me—I don't know—I will feel really upset about it, or really angry, or really emotional. But the AI doesn't have those emotions. I can read the conversation, understand what happened, and be objective. I think that level of support is the one that I like more, instead of, “Did this guy tell me a lie?” I think that's not exactly what I find curious for me in terms of opportunity.

swyx

Are you going to interject in real time? Say I'm arguing with somebody. Will the AI say, “Hey, no, you're wrong. That person actually said...”?

Ethan Sutin

Proactivity is something we're very interested in. Maybe not specifically for settling arguments, but more generally. A lot of the challenge here is that you need really good reasoning to pull that off, because you don't want it constantly interjecting—that would be super annoying. You also don't want it to miss things that it should be interjecting.

It would be a hard task even for a human, to come in at the right times when it's appropriate. With the personal context, it's going to be a lot better, because if somebody knows about you. But even still, it requires really good reasoning to not be too much or too little, and just right.

swyx

The part about some things is that you say something to somebody else, but afterward I change my mind and send something. Every time I have a different type of conversation and data about me.

Maria de Lourdes Zollo

One of the things that we're learning is that humans evolve over time. For us, one of the challenges is actually understanding whether this is a real fact. So far, we have a human in the loop who can say, “Yes, this is true. This is not,” or they can edit their own fact. In the future, we want to have all of that automated inside the product.

Your question also hits on privacy, which I know we'll talk about. If you have some memory and you want to confirm it with somebody else, that's one thing. But it's for sure going to be true that in the future—not even that far into the future—it's just going to be normalized.

We're in a transitional period now, and I think one of the key things for us is to navigate that and make sure we're thinking of all the consequences and how to make the right choices in the way that everything's designed. It's more beneficial than it could be harmful, but it's just too valuable for your AI to understand you.

If it's Meta Ray-Bans or Google Astra, I think people are going to be more used to it. People's behaviors and expectations will change. Whether that's something that is going to happen now or in 5 years, it's probably in that range.

We adapt to new technologies all the time. When the Ring cameras came out, that was quite controversial. But now people understand that a lot of people have cameras on their doors. We're in that transitional period for sure.

swyx

I'll press on the privacy issue, because that's the number 1 thing that everyone talks about. Obviously, I think in Silicon Valley people are a little more tech-forward and experimental, whatever. But you want to go mainstream. You want to sell to consumers, and we have to worry about this stuff.

The baseline question—the hardest version of this—is the law. There are 1-party-consent states where this is perfectly legal. Then there are 2-party-consent states where it's not.

Ethan Sutin

Yeah, the EU is a totally different regulatory environment. But in the US, it's basically on a state-by-state level. In Nevada, it's single-party; in California, it's 2-party.

But it's kind of untested. There are different laws depending on whether it's a phone call or whether it's in person. In a state like California, anytime you're in public, consent doesn't come into play, because the expectation of privacy is that you're in public.

Maria de Lourdes Zollo

But we process the audio, and nothing is persisted. Then it’s summarized with speaker identification focused on the user. Now, it’s kind of untested legally—and I’m not a lawyer—but does that constitute the same as a recording? So it’s kind of a gray area and untested in law right now.

I think the bigger question is, because if you had your Ray-Bans on and were recording, then you have a video of something that happened. That’s different from having an AI give you a summary focused on you that’s not really capturing anybody’s voice. I think the bigger question, regardless of the legal status, is: What’s the ethical situation with that? Even in Nevada—or many other U.S. states where you can record everything and don’t have to have consent—is it still the right thing to do?

The way we think about it is that we take a lot of precautions not to capture personal information from people around you, both through speaker identification, through the pipeline, and then through the prompts and the way we store the information, to be really focused on the user. We know that’s not going to satisfy a lot of people, but if you do try it and wear it, it’s very hard for me to see anything that somebody wearing a Bee around me could capture that I would ever object to as a third party.

We’re in this transitional period where the expectation will become more normalized that it’s an AI. It’s not capturing a full audio recording of what you said; everything is fully geared toward helping the person understand their state and providing valuable information to them, not logging details about people they encounter.

swyx

You know, I’ve had the same question with Zoom meeting transcribers. I think there’s kind of a personal impact. There’s a Fireflies.ai recorder, and I just know that it’s being recorded. It’s not like I know whether I’m going to say anything different, but intrinsically, you kind of feel different because it’s not pervasive.

I’m curious, especially in your investor meetings, whether people feel differently. Have you had people ask you to turn it off in a business meeting and not record? I’m curious whether you’ve run into any of these behaviors.

Maria de Lourdes Zollo

What’s funny is, on my end, I wear it all the time. I take my coffee at Blue Bottle with it, or I work with it. Obviously, I’m working on it, so I wear it all the time. So far, I don’t think anybody has asked me to turn it off.

I’m not sure whether it’s because they’re really friendly with me or because they know I’m working on it, but nobody really cared.

swyx

This is because you live in San Francisco.

Maria de Lourdes Zollo

Actually, I’ve been in Italy as well, and Italy is super privacy-concerned. Europe is super privacy-concerned. Again, nothing. I don’t know. That, for me, was interesting. Nobody’s ever asked me to turn it off, even after giving them full demos and disclosing.

I think some people have said, “Well, in a personal relationship, my partner was initially kind of uncomfortable about it.” We heard that from a few users, and that was more in a personal relationship situation. The other big one is people say, “I do like it, but I cannot wear this at work,” because they think they’ll get in trouble based on policies.

If you’re wearing it inside a research lab or somewhere where you’re working on things that are sensitive, we’re adding certain features like geofencing, so you can set it to never be active at a particular location. We’re also looking at context fencing, so you can say, “If these topics come up, don’t capture anything.”

I’ve often explained it the other way: Maybe you only want it at work. So you never take it from work, and it’s just a work device, like your Zoom meeting recorder is your work device.

Professionals have been big early adopters. You say San Francisco, but our daily shipments of over 100 are going to addresses in Texas, which I think is our biggest state, and Florida—just the biggest states. A lot of professionals talk for a living, and we didn’t set out to build it for that use case, but I think there’s a lot of demand from white-collar people who talk for a living. We’re just starting to talk with them. I think they want to be able to improve their performance, understand what they were doing, and figure out how they can do better.

swyx

How do you think about Gong.io and some of these sales-training tools where you put on a sales call and then it coaches you through it, which are more verticalized, versus having a more horizontal platform?

Maria de Lourdes Zollo

I’m not super familiar with the space. Like I said, we weren’t building for that use case, so it’s kind of a surprise to us. But I think those are interesting. I’ve seen there are a bunch of them now, right? It kind of makes sense.

I’m terrible at sales, so I could probably use one, but it’s not my job fundamentally. Maybe in a lot of situations, it could be useful. We’ve also heard from people with restaurants that, if they’re able to understand whether they’re doing well, that can be valuable.

In general, I think a lot of people like to have that double-check: Did it do this well, or can you suggest how I can do better? We had a user who said he used Bee for job interviews and then asked Bee, “Actually, how do you think my interview went? What should I do better?” I like that. I’m like, “Oh, that’s actually a personal coach in a way.”

swyx

But I guess the question is: Do you want to build all of those use cases, or do you see Bee more as a platform where somebody’s going to build the sales coach that connects to Bee, so you’re kind of the data feed into it?

Maria de Lourdes Zollo

I don’t think this is just a data feed. It’s more of an understanding engine. Definitely in the future, having third parties use the API and build out all the different use cases is something that we want to do, but the initial case we’re trying to build is that layer for all of it to work.

We’re not trying to build all those verticals, because no startup could do that well. But it’s been fascinating to see. I’ve done consumer for a long time, and consumer is very hard to predict. It’s hard to know what’s going to be the killer feature. We really believe that’s the future, but we don’t know exactly what process it will take to gain mass adoption.

swyx

The killer consumer feature is whatever Nikita Bier does. Social apps for teens. Yeah, well, I like Nikita, but he’s good at building bootstrap companies and getting them very viral, then selling them, and then they shut down.

Okay, so you just came back from CES?

Maria de Lourdes Zollo

Yeah. Crazy. It was my first time in Vegas and my first time at CES. Both were overwhelming.

swyx

First of all, did you feel like you had to do it because you’re in consumer hardware?

Maria de Lourdes Zollo

We decided to be there and have a lot of partner and media meetings, but we didn’t have our own booth, so we decided to skip that. But we decided to be there and have a presence, even just us, and speak with people.

swyx

Hard to stand out.

Maria de Lourdes Zollo

Yeah, I think it depends on what type of booth you have. I think if you can prepare a really cool booth, it can be pretty cool.

swyx

Have you been to CES? It can be pretty cool. It’s massive. It’s like 80,000 or 90,000 people across the Venetian and the convention center. To me, I always wanted to go, even just as a fan of—

Maria de Lourdes Zollo

Yeah, you wanted to go. Growing up, I think CES kind of peaked for a while, and it was like, “Oh, I want to go there. That’s where all the cool gadgets, everything, is.”

swyx

There are a lot of cool vacuums and pet tech, and dog stuff.

Maria de Lourdes Zollo

There’s a lot of robot stuff, new TVs, and new cars that never ship.

swyx

I think this time last year was when Rabbit and Humane launched at CES, and Rabbit kind of won CES. Now, this year, there were no wearables except for you guys.

Maria de Lourdes Zollo

It’s funny because it’s obviously AI everything. Every single product—like, a toothbrush with AI. We saw a hair blower, literally a hair dryer with AI. That was cool.

But another difference around us is that we didn’t want to do a big, overhyped, promised kind of Rabbit launch. I mean, hats off to them on the presentation and everything, obviously, but we wanted to let the product speak for itself and get it out there. We were really happy with the interest we got from the media and some of the partners there, so it was definitely worth going.

I would say that if you’re in hardware, it’s just about how you make use of it. To do a big Rabbit-style launch or have a huge show there, you need to plan that 6 months in advance, and it’s very expensive. But if you go there, everybody’s there. All the media are there, and there are a lot of pre-show events where it’s great to talk to people in the industry.

All the manufacturers and suppliers are there, too, so we learned about some really cool stuff. We met with them; they have thermal energy capture.

And it's like, oh, could you maybe not need to charge it because they have a thermal that can capture your body heat.

swyx

What?

Maria de Lourdes Zollo

Yeah, they're here. They're actually here in Palo Alto. They have a Fitbit thing that you don't have to charge.

swyx

How much power can you get from that? What's the power draw for this thing?

Ethan Sutin

It's more than you could get from the body heat, it turns out, but it's quite small. I don't disclose the technical details, but I think solar is still more realistic. They also have one where the face of it is just a solar cell, and that is more realistic. Or kinetic. Kinetic, apparently—they seem to think it wouldn't be enough. Kinetic capture is quite small, I guess.

swyx

Well, I mean, watchmakers have been powering things with kinetic energy for a long time. We don't have to talk about that. I just wanted to get a sense of CES. Would you do it again?

Maria de Lourdes Zollo

I definitely would. Okay, you're just a fan of CES. From a business point of view, it doesn't make sense.

swyx

I happen to be in the conference business, right? So I'm kind of just curious.

Maria de Lourdes Zollo

Yeah, so I would say that, without the booth and with really straightforward conversations that were already planned, 3 days was okay. I think it was okay. But if you need to invest in a booth, that's not cheap.

swyx

Which is how much?

Maria de Lourdes Zollo

A 10-by-10 is $5,000. But on top of that, you have the financial costs. A 10-by-10 is like one of the super-fancy things, and some companies have—I think we'd probably be more in the 6-figure range to get a booth.

I mean, I think that, yeah, it's very noisy. We heard that it's very, very noisy. Obviously, everything is being launched there, from cars to cell phones, so it's hard to stand out. But I think going in with a plan of who you want to talk to was worth it. We had a lot of really positive media coverage from it, and we got the word out, so I think we accomplished what we wanted to do.

swyx

I mean, there's some world in which my conference is kind of the CES of whatever AI becomes.

Maria de Lourdes Zollo

Yeah, I think that—don't do it in Vegas. Don't do it in Vegas. That's the only thing I didn't really like there.

swyx

Okay, that's great. Amazing. Those are my favorite ones. You cannot fit 90,000 people in San Francisco.

Maria de Lourdes Zollo

That's really the problem. You need to do multiple locations. You can do Moscone and then have one—

swyx

That's what the Salesforce conference is. GDC is how many?

Maria de Lourdes Zollo

That might be 50,000, right?

swyx

Okay, form factor, right? My way to introduce this idea was that I was at the launch in Solaris. What's the old name of it? Newton? Of Tab when Avi first launched it. He was like, “I thought through every form factor. Pendant is the thing.” And then we got the pendant for the original one—the first one, which was a pendant—and I took it off and forgot to put it back on.

So you went through pendant, pin, and now bracelet, and maybe there are AirPods—a sort of earphone—in the future. What was your iteration to that?

Maria de Lourdes Zollo

Yeah, so we had, I believe, 3 or 4 iterations, and one of the things that we learned is that people don't like the pendant. In particular, women don't want to have anything here on the chest because maybe they have another necklace or other stuff.

swyx

You just ship a premium one that's gold. We're talking about some fashion—some big fashion. There's something there. This is where it helps to have an Italian.

Maria de Lourdes Zollo

Exactly. Some big Italian luxury. I can't say anything, sorry.

The bracelet actually came from the community because they were like, “I don't want to wear anything as a necklace or a pendant.” Also, the one that we had—I don't know if you remember—was a circle and was really bulky. People didn't like it.

And I actually don't dislike it. We were running fast when we did that. Our thing was that we wanted to ship them as soon as possible, so we weren't overthinking the form factor or the material. We just wanted to be out.

But after the community organically told us, basically, all of them were like, “Why don't you just do the bracelet? It's way better. I will just wear it, and that's it.” So that's how we ended up with the bracelet, but it's still modular. I still want to play around with the fact that it's modular, and you can take it off and wear it as a clip. In the future, maybe we will bring back the pendant, but I like the fact that there is some personalization.

Right now, we have 2 colors, yellow and black. Soon we will have other ones, so we can play a lot around that.

swyx

I think the goal for the form factor is for it to be not super invasive, right? Something that's easy. In the future, smaller and thinner—not like Apple's obsession with thinness, but it does matter, the size and weight. We would love to have more context because that will help, but to make it work, I think it really needs to have good power consumption and good battery life.

With the Humane, swapping the batteries—I have one. I think the Humane is pretty incredible, some of the engineering they did, but it wasn't geared toward solving the problem. It was just too heavy. The swappable batteries are too much to manage—the heat, the thermals. It's too much for a light-interface thing.

Maria de Lourdes Zollo

Yeah, like that. That was cool.

swyx

It's cool. It's cool, but if you have your hand out here and you want to use your phone, it's not really solving a problem because you know how to use your phone. It's got a brilliant display, but you have to learn how to gesture with this low-resolution laser. The laser is cool—the fact that they got it working in that thing, even though it did overheat—but it's too heavy, too cumbersome, and too complicated with the multiple batteries. So something that's power-efficient and thin, both in the physical sense and in the edge-compute kind of way, so that it can be as unobtrusive as possible.

Maria de Lourdes Zollo

Users really like it. I like when they say, “Yes, I like to wear it and forget about it,” because I don't need to charge it every single day. On the other version, I believe we had 35 hours or something, which was okay, but people just prefer the 7-day battery life.

swyx

Oh, this is 7 days?

Maria de Lourdes Zollo

Yeah.

swyx

Oh, I've been charging it every 3 days.

Maria de Lourdes Zollo

Yeah, you can keep it full for 7 days.

swyx

The other thing that occurs to me is maybe there's an Apple Watch strap.

Maria de Lourdes Zollo

Yeah.

swyx

So that I don't have to double-watch. I have—

Maria de Lourdes Zollo

Yeah, that's the other one.

swyx

Yeah, I thought about it. I also saw the ones that you can put back on the phone. There are a lot of possibilities.

There's a competitor called PLAUD. It's not really a competitor; they only transcribe, right?

Maria de Lourdes Zollo

Transcribe, but they're very good at it.

swyx

Yeah, no, they're great. Their hardware is really good, too, and he just launched the pin, too.

Maria de Lourdes Zollo

Yeah, I think the MagSafe kind of form factor has a lot of advantages, but also some disadvantages. You can definitely put a very large battery on that, and so the power consumption isn't as much of a concern. The downside is that the phone is in your pocket. I think form factors will continue to evolve, with more sensors, less obtrusiveness, and easier use. We have a new version.

swyx

Okay, looking forward to that. Whenever we launch this, we'll try to show whatever, but I'm sure you're going to keep iterating.

Last thing on hardware, and then we'll go on to the software side, because I think that's where you guys are also really strong. Vision—you wanted to talk about why no vision?

Ethan Sutin

Yeah, I think it comes down to the fact that when you're a startup, especially in hardware, you work within the constraints. Vision is super useful and super interesting, and it's what we actually started with. There are 2 issues with vision that make it not the place we decided to start.

One is power consumption. You have to trade off your power budget. Capturing, even at a low frame rate, and transmitting over the radio actually takes up the majority of the power. So you would really have to have an unacceptably large and heavy battery to do it continuously all day. We have some novel alternative ways that might allow us to do that, and we have some prototypes.

The other issue is form factor. Even with a wide field of view, if you're wearing something on your chest, it's obviously not going to capture the field of view of what's interesting to you. The wrist isn't really that much of an option, and if you're wearing it on your chest, it's often not going to capture what's interesting to you. That leaves you with your head and face.

Anything that goes on the face has to look cool. I don't know if you remember Snap Spectacles. That was kind of like the first—

swyx

Yeah, but they weren't very successful, and I think one of the reasons is that they were so weird-looking. Their camera was so big on the side.

If you look at the Meta Ray-Bans, where they're way more successful, they look almost indistinguishable from a Ray-Ban. Meta invested a lot into that, and they have a partnership with Qualcomm to develop custom silicon. They also have a stake in Luxottica now, so they're coming from all angles to make glasses.

I don't know if you know Brilliant Labs. They're a cool company. They make Frame, which is a cool, hackable pair of glasses, and they're really good on hardware.

Ethan Sutin

But even if you look at the frames, which I would say are from the most advanced kind of startup, there was one that launched at CES, but it isn’t shipping yet. The one that you can buy now still isn’t something you’d wear every day, and the battery life is super short. So I think the challenge of doing vision right off the bat would require quite a bit more resources. Audio is such a good entry point, and there’s also the privacy around audio. If you had images, that’s another huge challenge to overcome. Ideally, personally, I would have all the senses, and we’ll get there.

Alessio Fanelli

Okay, one last hardware thing, because I have to ask this before we move to software. Were either of you electrical engineers?

Ethan Sutin

No, I’m CS, and I’ve taken some EE courses, but prior to working on the hardware here, I had done a little bit of embedded systems—very little firmware—but we luckily have somebody on the team with deep experience.

Maria de Lourdes Zollo

Yeah. I’m just like, you have to become hardware people. I learned how to worry about supply chain, power, and radio. I would say this about hardware—and I know it’s been said before—but building a prototype, learning how the electronics work, learning about firmware, and developing this is fun for a lot of engineers, and it’s all totally achievable, especially now with the tools we have. Stuff you might have been intimidated by, like, “How do I write this firmware?”—now, with Claude Sonnet, you can get going and actually see results quickly. But I think going from a prototype to actually making something manufactured is an enormous jump, and it’s not all about technology: supply chain, procurement, regulations, costs, and tooling.

Alessio Fanelli

The thing about software that I’m used to is that it’s funny: you can make changes all along the way and ship them. But when you have to buy tooling for an enclosure, that’s expensive. Own tooling?

Maria de Lourdes Zollo

You have to.

Alessio Fanelli

Don’t you just subcontract out to someone in China?

Maria de Lourdes Zollo

We make the tooling? No, no. You have to have CNC and a bunch of machines. Nobody makes their own tooling, but you have to design it and submit it, and then they go 4–6 weeks later. If there’s a problem with it—

Alessio Fanelli

Right.

Maria de Lourdes Zollo

Well, then you’re not making any of your enclosures.

Alessio Fanelli

What resources or websites are most helpful in your manufacturing journey?

Ethan Sutin

I think it’s different depending on the product, because hardware is so specialized in different ways. I would say that, for example, you should choose a manufacturing company and speak with other founders. They’ll give you the language and some tips about who is good and who is not, who specializes in something versus somebody else. Some people are good in plastics, and some people are good in PCBs. For us, it really helped at the beginning to speak with others and understand who is around. I worked in Shenzhen and lived almost 2 years in China, so I have an idea about different hardware manufacturers and all of that. Soon I’ll go back to check things out. I think it’s good also to go in person and check and see the right people.

Alessio Fanelli

We did some stuff domestically, and if you have that ability—the reason I say ability is that it’s very expensive—you can build out some proofs of concept and do field testing before taking it to a manufacturer. Despite what people say, there’s really good domestic manufacturing for small quantities, at extremely high prices. We got our first PCB and assembly done in LA, because the defense industry supports a lot of good manufacturing that can do quick turns. It’s like, we need this board, we need to find out if it’s working, we have this deadline, and we want to start; if you want to have it done and fabricated in a week, they can do it for a price. But I think everybody’s trending, even for prototyping, toward moving offshore now, because in China you can do prototyping and get it within almost the same timeline. The thing is, with manufacturing, it really helps to go there and establish the relationship.

Maria de Lourdes Zollo

Yeah, my first company was a hardware company, and we did our PCBs in China. It took a long time. Now things are better, but this was, I don’t know, 10 years ago or something like that. I’ve heard this too: if it’s something where you don’t have the relationships, they don’t see you, they don’t know you, you might get subcontracted out, or they’re not paying attention. But if you have the relationship and a priority, it’s usually good.

Ethan Sutin

We ended up doing the fabrication and assembly in Taiwan for various reasons, but it really helped that you went there at some point.

Maria de Lourdes Zollo

Yeah, we’re really happy with the process. But the whole process of choosing the right people—choose the right people—but also just sourcing the bill of materials and all of that stuff: I guess if you have time, it’s not that bad, but if you’re trying to really push the speed, it’s incredibly stressful.

Alessio Fanelli

Okay, we’ve got to move to software. The hardware may be hard for people to understand, but what software people can understand is that running transcription and summarization, all these things, in real time every day for 24 hours a day, isn’t easy. So you mentioned 200,000 tokens for a day. How do you make it basically free to run all of this for the consumer?

Ethan Sutin

Well, I think that the pipeline and the inference—people think about all these tokens, but as you know, the price of tokens is dramatically dropping. You guys probably have some charts somewhere that you’ve posted, and if you see that trend, 250,000 input tokens isn’t really that much, right? The output layers?

Alessio Fanelli

You do live?

Ethan Sutin

Yeah. Speech-to-text is the most challenging part, actually, because it requires real-time processing and then later processing with a larger model. One thing that’s fairly obvious is that you don’t need to transcribe things that don’t have any voice in them, right? Good voice activity detection is key, because the majority of most people’s day isn’t spent with voice activity. That’s the first step to cutting down the amount of compute you have to do, and voice activity detection is a fairly cheap thing to do—very, very cheap.

The models that need to summarize don’t need a Claude Sonnet-level model. You do need a Claude Sonnet-level model to execute things like the agent, and we’ll be having a subscription for features like that because, although now with DeepSeek R1, we’ll see—we haven’t evaluated it. There are already models that can perform at that level. I was going to say in 6 months, but—

Alessio Fanelli

DeepSeek R1?

Ethan Sutin

Yeah, I mean, not that one in particular, but they’re already there and can perform at that level. Self-hosted models help with the things where you can.

Alessio Fanelli

So you’re self-hosting models? You’re fine-tuning your own ASR?

Ethan Sutin

Yes. I will say that I see everything trending down in the future, although I think there might be an intermediary step where things become expensive, which we’re really interested in because the pipeline is very tedious and requires a lot of tuning. That’s brutal because it’s just a lot of trial and error. Wouldn’t it be nice if an end-to-end model could just do all of this and learn it? If we could do transcription with an LLM, there are so many advantages to that, but it’s going to be a larger model and hence more compute. We’re optimistic that maybe we could distill something down, and we kind of focus more on reducing the cost of the existing pipeline or trying the next generation, because it’s very clear that all ASR, all speech-to-text, is going to be pretty obsolete pretty soon. Investing in that is probably a dead end because it’s just going to be obsolete.

Alessio Fanelli

It’s interesting. I think when I initially invested in Tab—this shows you how wrong I was—I thought, “Oh, this is a sort of razor-and-blades model where you sell cheap hardware and make up a subscription, a monthly subscription.” Now I just checked: Friend is a 1-time sale, $99; Limitless is a 1-time sale, $99; these guys are a 1-time sale, $49. And Friend is free? What? When you probably invested, how much was 1,000,000 input tokens at that time, and what is it now? It’s a fascinating business, and there’s a lot to dig into there, but just getting that perspective out there is not something that people think about a lot. You obviously have thought a lot about it.

What about memory? I think this is something we go back and forth on: you’re just memorizing facts and then understanding what is a preference, and adjusting the facts that you think about a person? Any learnings from that? I know there are a lot of open-source frameworks now that do it. Did you build all of your own infrastructure internally?

Ethan Sutin

Yeah, we did. I evaluated and used a lot of them in other projects. I think there are a few different tasks or things that revolve around memory. One is retrieval, obviously, and when you need to find something—even if you have a large corpus—how do you find it? I think existing RAG pipelines will probably also be obsoleted. The frameworks—I have not found one. There’s no general way to do RAG that works; it’s really highly dependent on the data.

swyx

So, if you're going to be customizing something that much, you get more bang for the buck from designing it all yourself. A lot of these frameworks are great for getting going quickly, but I think it's really interesting when you're trying to do memory for a person, because memories decay, right? I'm going to London, then I come back; I'm not going to London anymore.

Ethan Sutin

What we've learned is that doing traditional embedding and RAG is suboptimal. We built our own using small models to do massively parallel retrieval, which I think is going to be more common in the future. To represent a person, we still require some human loop. This is an ongoing project, and we're learning every day. How do you correct the model when it gets something wrong about you?

Right now, we have things that are super confirmed—ground truth about you because a human accepted it—but ideally that step wouldn't be necessary. Then we have things that are fuzzier. The more stuff that we know is true, the more accurate we are when we're trying to decide whether this fuzzy stuff is true, because if you have the context, it's probably not true.

So, I think one of the core challenges is how to handle both retrieval and modeling, especially when you're dealing with noisy source data. Even in an ideal world, if you had perfect transcription and were going off that, that's still not enough information, right? Even if you had visuals, it's still not enough. There's still going to be some misunderstandings, so how do you not let that damage the value of it and make it unrecoverable and uncorrectable?

swyx

Yeah, one way I think about it is that I usually like to order the same thing from the same restaurant if I like it, but I'm not saying that out loud. Are these the types of behaviors that you can capture? When you ask about a favorite restaurant, I would want it to give me restaurants I've already been to and liked. Or if I'm like, “Hey, just order something from this place,” it should reorder the same thing because it knows that I like to get the same thing again. But I feel like today, most agent memory things that I see people publish are just, you know, write down the data thing.

Ethan Sutin

Yeah, I think that's why the reasoning—in our case, giving it time to consider all of the sources it has—is really important. Look at the emails, see the receipts, and then look at the conversations to see what I've mentioned. Then be able to take enough time to search through all the context and connect the dots is really important. I don't know; some of the agent memory stuff is like key-value with RAG on top, and the results are just not complete enough when you have a growing corpus, managing decay, and hallucinations that might be in the source material.

Alessio Fanelli

So, this is where people usually bring in knowledge graphs.

Ethan Sutin

Yes.

Alessio Fanelli

And do you use them? Is it for speed, or what are the issues?

Ethan Sutin

We don't extensively use knowledge graphs. It's something—we didn't talk also about the potential future social aspects—but the problem with knowledge graphs that we found is that they're great for representing the data, but using them at inference time is challenging. It's just the LLM understanding the graph input. It's not in the training data, for sure. I think the graph is the right kind of way to store the data, but then you need to have the right retrieval and format it in a way that doesn't overwhelm or confuse what you're trying to do.

swyx

Should we ask about social?

Alessio Fanelli

Yeah, I know. I thought you were going to go into it.

swyx

Yeah, what's that?

Ethan Sutin

Not directly related to graph retrieval or graph knowledge bases, but the idea is that you have your personal context, and then other people can query it. It can divulge some things that you would have full control over. Then Maria and I are trying to negotiate where we're going to have dinner. There can be an exchange between the agents.

Alessio Fanelli

Yeah, there can be an exchange between the agents.

Maria de Lourdes Zollo

So, how can my agent speak with Ethan's agent? Both of them know our location, what we like, and where we went in the past. Even if we have our calendars integrated, they know when we're free. They can interact with each other, have a conversation, and decide on a place to go for us.

Alessio Fanelli

Wow.

Ethan Sutin

And we did that. It was really cool for me because they suggested a nice French restaurant that we went to in the end.

swyx

That you'd never been to?

Maria de Lourdes Zollo

That we'd never been to. They saw that we both liked French food, and we were both in Pacific Heights. This was really trivial.

Alessio Fanelli

Yeah, it's a trivial toy use case, but I guess, in terms of having used it for a while, if I wanted to buy you a gift—

Maria de Lourdes Zollo

Oh my God, you bought me a bunch of candles, now that I think about it. This is another use case. I was like, “Yeah.”

Ethan Sutin

When we were testing the agent, a bunch of candles from Amazon showed up at her door.

Maria de Lourdes Zollo

Yeah, because I really love candles, but I didn't expect 20.

Ethan Sutin

Yeah, it's a lot—extreme, but how do you manage that? What's okay for your Bee to divulge, and to whom? Shouldn't you get an authorization request every time for personal context?

Maria de Lourdes Zollo

Yeah, yeah, yeah. A human would have to sign off on it, but then I wouldn't have to guess.

swyx

There's this culture that's very alien to everyone else outside of San Francisco, and outside of the Gen Z bubble in San Francisco, which is sharing location. I can tell exactly where my close friends are right now in the city. It's normal here, and it freaked out everyone who's not here.

Maria de Lourdes Zollo

Yeah.

swyx

So maybe we can share preferences—who we like, or even small updates about your day. My parents would love that because I don't do that.

Ethan Sutin

Yeah, so now there's no friction. It can just be more or less automatic.

Maria de Lourdes Zollo

Dating? I was always trying to avoid dating as a startup founder.

swyx

Everyone hates it?

Maria de Lourdes Zollo

We thought about it. Sometimes people ask us, “Oh, you know so much about me. Can you measure compatibility with somebody else or something like that?”

swyx

Yeah, probably there is a future, and maybe somebody should build that.

Alessio Fanelli

I think on our end, we were like, “No, this is—I’ll build on your API.”

swyx

My sister's actually a personality psychology professor, and she studies personality. We were at Thanksgiving with my parents, and I was like, “Give me my Big Five,” which is the personality type. “Does it know my Big Five?” You just ask it to consider everything and give you your Big Five. My sister said it was pretty accurate. I didn't agree with it because it said I was disagreeable, but she seems to think I'm agreeable.

Alessio Fanelli

You disagree that you're disagreeable?

swyx

Yeah. What other proof do you need, then? I think I'm very agreeable.

Maria de Lourdes Zollo

But I think we did get some users who were like, “Oh, if we're a couple—” We had couples actually buy the product together. Both members of a couple bought our hardware, so there is something there.

swyx

Another test is the Myers-Briggs. I know you don't like that one.

Maria de Lourdes Zollo

No, no, no. OCEAN is cooler than Myers-Briggs. Everyone stop using my MBTI. Use my OCEAN.

Ethan Sutin

For me, it was on point every time.

Maria de Lourdes Zollo

Awesome.

swyx

Anything else that we didn't cover? Any cool underrated things?

Ethan Sutin

Go to B. dot computer $49.99 and buy the device. That's the call to action. You're hiring? We are hiring—for sure, AI engineers. Nice. What is an AI engineer? Somebody who's scrappy and willing to work with us. I think you coined the term, right? So you can tell us. People have different perspectives, and what is useful for you is different from what is useful for me.

So, anyway, it's useful. I think always-on AI is really going to explode. There's going to be a lot from startups, but incumbents, too, and there are going to be all kinds of new things that we're going to learn about how it's going to change all of our lives. I think that's the thing I'm most certain about. And be in the AI. Well, thanks very much. Yeah, this was a pleasure. Yeah, we'll see you launch whenever that launch is happening. Yeah, thanks. Thank you.