[BidClub_]
Latent Space · · 69 min

Bee AI: The Wearable Ambient Agent

Alessio FanelliswyxMaria de Lourdes ZolloEthan Sutin

YouTube
TL;DR
  • Bee’s central bet is that personal AI becomes valuable when it passively accumulates first-person context, not when users repeatedly explain themselves to another chatbot. Memory is only “the base kind of use case”: the larger ambition is an assistant that understands relationships, attitudes and preferences well enough to “already know” what its user wants. Ethan estimates a person can produce roughly 200,000 tokens daily and says his own personal corpus is about 50 million tokens.

  • The $49.99 wearable exists because tiny interaction costs undermine an always-on experience. A phone monopolizes its microphone and gets interrupted by calls; Apple Watch users must restart Bee after calls and charging, and some consequently buy the dedicated device within days. Bee gets up to seven days of battery life while complementing the phone, which Ethan expects to remain dominant for perhaps five years until something like affordable, lightweight Orion-type glasses arrives.

  • Bee’s core product layer is its context engine and ability to act, not merely its bracelet or transcription. It performs speaker identification, semantic conversation endpointing, summaries, tone and action-item extraction, then connects speech with Gmail, Google Calendar, location and web search. A WhatsApp beta can notice an actionable message, propose help and operate an Android cloud phone, while the API currently offers read-only and real-time socket access with external actions promised “very soon.”

  • Professional users have emerged as the clearest early wedge even though Bee was not designed as a vertical work product. The company says daily shipments exceed 100, with Texas—Maria thinks—its biggest state and Florida also prominent; users include white-collar professionals reviewing conversations, job candidates asking how interviews went and restaurant operators checking performance. Bee intends to provide the general “understanding kind of engine” and let third parties build specialist coaches because “no startup could do all those verticals” well.

  • Privacy is a central adoption concern, an unsettled legal question and a product-design constraint. Ethan distinguishes one-party-consent Nevada from two-party-consent California, while Maria—explicitly noting she is not a lawyer—says processed, unpersisted audio and user-focused summaries may not be equivalent to retaining a recording, an explicitly “gray area and untested in law.” Bee is adding location and context-based disabling, because users may like the product yet say, “I cannot wear this at work.”

  • The hardware roadmap prioritizes unobtrusiveness and battery life over technically impressive spectacle. Community rejection of a bulky pendant led to a modular bracelet or clip, while vision was deferred because image capture and radio transmission consume too much power, chest cameras miss the relevant field of view and anything worn on the face “has to look cool.” Moving from a prototype to production remains an “enormous jump” involving suppliers, tooling, regulations and four-to-six-week tooling lead times.

  • Falling token prices undermine the original razor-and-blades thesis, but speech processing and personal memory remain hard technical problems. Voice-activity detection avoids transcribing silent hours; small models can summarize, while Sonnet-level agent execution may justify a subscription and R1 had not yet been evaluated. The team self-hosts models, fine-tunes ASR and uses custom massively parallel retrieval because “there’s no general way to do RAG that works” across a growing, noisy corpus whose facts decay or contradict one another.

  • The long upside case is a permissioned network of personal agents coordinating on their users’ behalf. Bee agents already demonstrated choosing a French restaurant from two users’ locations, calendars and preferences, although an experiment that delivered roughly 20 candles illustrates why authorization and spending controls matter. Ethan’s highest-conviction forecast is that “always-on AI is really going to explode,” with startups and incumbents discovering how it changes daily behavior.

Digest · the substance, structured for research

1. Personal context—not recall alone—is Bee’s product

  • Maria defines Bee as “AI living alongside you in first person,” capturing real-life context so it can recall, reflect and eventually help without repeatedly asking about the user’s preferences. Memory is “really just the base kind of use case”; the intended model includes attitudes, desires, relationships and changing circumstances.

  • Maria says the most visible behavior currently clusters around companionship and professional assistance. Her preferred analogy is a best friend who knows the user extremely well, while the hosts argue that reducing this ambition to a list of “use cases” makes the product sound drier than the envisioned experience.

  • Ethan now informally measures “what is my token output of the day?” and arrives at roughly 200,000 tokens; his accumulated personal context may be around 50 million. Even a single-day summary surprised him as a non-journaler, while a “Spotify Wrapped for my 2024” surfaced metrics such as three countries visited and beta devices shipped.

2. A failed 2016 chatbot made passive learning non-negotiable

  • Ethan’s first personal-AI attempt began in 2016, “before transformers—no BERT even—just RNNs,” during the brief post-F8 belief that bots would replace apps. Convincing dialogue was impossible, but he and his previous co-founder were already trying to model a dynamic person through an app that asked questions and returned feedback.

  • The company joined Betaworks’ Botcamp alongside the original Hugging Face, then a teenage chat app saying things like, “Hey friend, how was school? Let’s trade selfies.” Ethan believes Hugging Face built the Transformers library to improve that product, then open-sourced it, after which the library became the larger opportunity.

  • Teenagers sometimes enjoyed answering endless questions about themselves, but most users would not do the labor required to teach an AI manually. Ethan’s team pivoted through YC into viral consumer video, created Squad, sold it to Twitter and helped launch Twitter Spaces; they started working on Bee again shortly before ChatGPT appeared, after app experiments exposed the limitations of the earlier approach.

3. Dedicated hardware wins by deleting tiny daily frictions

  • Ethan’s test for any wearable is, “Why isn’t this an app?” Continuous phone capture would monopolize the microphone, stop during calls and require users to remember to restart it. That “little bit of friction” materially breaks the feeling of an intelligence simply living alongside its user.

  • Bee therefore supports Apple Watch as an immediate, hardware-free trial, with background operation engineered to limit battery drain. Yet calls interrupt capture, and daily charging creates another reminder and friction to reopen Bee; Maria says users experience the value, tire of reopening the app and sometimes buy Bee after only a couple of days.

  • Owning the hardware gives Bee control over gain, sample rate and other audio parameters unavailable through Apple’s limited framework. The challenge is not studio audio but a system flexible enough to work at dinner in a noisy restaurant without tuning one environment so aggressively that another becomes unusable.

  • The device is meant to work with the phone, not recreate the Rabbit or Humane strategy of establishing a new hardware platform. Ethan calls the wearable “the ears of the AI” and expects phones to remain dominant until a cheap, light next-generation interface—perhaps Orion-type glasses—arrives years from now.

4. Capture becomes valuable only when it can act elsewhere

  • Bee’s app continuously shows what is happening now, then converts the day into readable conversation segments. Speaker identification prevents another person’s statement from becoming a fact about the user, while voice teaching helps identify close contacts and semantic endpointing separates conversations whose boundaries are naturally fuzzy.

  • Once a conversation ends, larger models analyze and summarize it, extracting key points, atmosphere, tone and possible actions; an end-of-day view combines those pieces with location. User facts remain visible and editable, creating a human-in-the-loop layer before uncertain observations become trusted personal context.

  • The recall interface can search memories, Gmail, Google Calendar and the web, then link its synthesis back to source conversations. Ethan demonstrated asking about a Taiwan manufacturing trip and receiving its details, production issues and relevant discussions rather than an ungrounded recollection.

  • In the WhatsApp beta, Bee watches notifications and asks two questions: is this important enough to surface, and can it help? A restaurant request triggered a suggestion, accepted through chat or push-to-talk; an Android cloud phone then found the WhatsApp thread and sent the answer. “Hey Alfred” supplies wake-word access when the phone is inactive.

5. An API turns personal context into infrastructure

  • Alessio’s developer case was blunt: no single consumer app can anticipate all his needs, so users must be able to own, correct and reprocess their data. He contrasted Bee with other wearables that lacked APIs, while swyx declared, “We are API enjoyers in this house.”

  • Bee does not store audio by default, but its API can expose processed context and provide real-time socket access. Access was read-only during the conversation; Ethan said write support and external actions would become fully available alongside the broader action system “very soon.”

  • The larger product vision is “not just a tool”: Bee should proactively suggest useful actions at an appropriate moment. Ethan acknowledges that the threshold is difficult—an assistant that constantly interjects is intolerable, while one that misses every meaningful opening contributes little—and believes strong reasoning plus personal context is necessary to find the balance.

6. Privacy remains both a legal gray zone and a social transition

  • Asked whether instant replay could settle arguments or expose lies, Maria emphasized AI as an “objective point of view” on what happened and on the emotions in a conversation. She is more interested in support for a user distorted by anger or sadness than in creating a machine that prosecutes every contradiction.

  • The discussion also distinguishes legitimate change from deception: humans say different things in different contexts and evolve over time. Bee currently lets users approve, reject or edit asserted facts; the team wants eventual automation but does not pretend that noisy speech alone establishes durable truth.

  • Ethan distinguishes single-party-consent Nevada from two-party-consent California and notes that public settings have different privacy expectations. Maria, explicitly noting she is not a lawyer, says that because Bee processes audio without persisting it and stores user-centered summaries rather than voices, whether it legally constitutes a recording remains “a gray area and untested.”

  • Ethics still matter where recording is legal. Maria reports no requests to disable Bee after disclosure, including in privacy-conscious Italy, but users cite uncomfortable partners and confidential workplaces. Planned safeguards include geofencing and “context fencing,” while Maria expects norms to shift over roughly the present-to-five-year transition, as they did around doorbell cameras.

7. Professional users emerged without a vertical go-to-market

  • Maria says Bee ships more than 100 devices daily, not chiefly to an early-adopter enclave in San Francisco. She thinks Texas is its largest state, with Florida also prominent, and sees strong demand from white-collar professionals “who talk for a living.”

  • One user used Bee during job interviews and then asked, “How do you think my interview went? What should I do better?” Other possibilities include restaurant professionals examining performance and people checking whether they performed well—an ambient version of personal coaching rather than only meeting transcription.

  • When the hosts compared this wedge with vertical tools such as Gong, Maria resisted turning Bee into every specialist application. The company wants to build the underlying understanding engine and enable third parties through its API; consumer behavior is too unpredictable to declare the eventual killer feature in advance.

8. Battery and wearability outrank cameras and spectacle

  • Bee moved from a bulky circular pendant after community members—particularly women already wearing necklaces—said they would not use it. The resulting module can become a bracelet or clip, currently comes in yellow or black and leaves room for other colors and possibly a future pendant design.

  • Maria believes an earlier version had roughly 35 hours of battery life, while the current device offers seven days. She emphasizes the desired reaction: “I like to wear it and forget about it”; the discussion accordingly prioritizes smaller, lighter and power-efficient hardware over an Apple-like pursuit of thinness for its own sake.

  • swyx admires Humane’s engineering but argues that its weight, heat, swappable batteries and laser interface did not solve a problem better than a phone. MagSafe recorders such as Plaud can carry large batteries, but a microphone attached to a phone inside a pocket creates a different capture problem.

  • Vision was deferred because even low-frame-rate capture and radio transmission consume much of the power budget, while chest placement often misses what the wearer sees. swyx contrasted conspicuous Snap Spectacles with nearly ordinary Meta Ray-Bans; Ethan added that current startup glasses still face short battery life and are not yet something people would wear every day.

9. Manufacturing makes hardware iteration fundamentally slower

  • Ethan calls the jump from prototype to manufactured product “enormous.” Firmware and electronics are increasingly approachable—even with Claude Sonnet helping someone get started on unfamiliar firmware—but production adds procurement, regulation, bill-of-materials sourcing, cost control and enclosure tooling that can take four to six weeks before a flaw becomes visible.

  • Maria recommends choosing suppliers through other founders, matching specialists to plastics, PCBs or other needs, and visiting them to establish relationships. Alessio described using costly domestic Los Angeles production for rapid early boards and said Chinese prototyping is increasingly competitive on time and price; Bee ultimately performed fabrication and assembly in Taiwan.

  • Bee attended CES for three planned days without buying a booth, using the roughly 80,000–90,000-person gathering for media, partners and suppliers. A 10-by-10 space costs about $5,000 before presentation costs, while standing out may require six figures; supplier meetings nevertheless surfaced body-heat capture, solar and kinetic options, with solar appearing more realistic for Bee’s power needs.

10. Falling inference costs push the hard problem into memory

  • Ethan argues that roughly 250,000 input tokens are no longer inherently expensive because token prices keep falling. Speech-to-text remains harder: it needs real-time processing and later processing by a larger model, so cheap voice-activity detection first removes the majority of a day in which nobody is speaking.

  • Summaries do not require a Sonnet-level model, whereas reliable agent execution may—and Bee expects to charge a subscription for those expensive capabilities. The team self-hosts models and fine-tunes ASR; R1 had not yet been evaluated.

  • Alessio acknowledged getting the business model wrong: he expected cheap hardware subsidized by recurring “razors and blades” revenue, yet cited Friend and Limitless at one-time prices of $99 and Bee at $49.99. Rapidly declining inference costs make that consumer-hardware economics question materially different from when he formed the thesis.

  • Ethan expects “all ASR, all speech to text” to become obsolete relatively soon as end-to-end models absorb today’s tedious pipeline, though they may initially require more compute before distillation. Bee likewise found generic RAG inadequate: memories decay, traditional embeddings and RAG underperform on a growing personal corpus, and “there’s no general way to do RAG that works” independently of the data.

11. Permissioned agent exchange is the long upside case

  • Bee built custom, small-model, massively parallel retrieval and separates user-confirmed ground truth from fuzzier inferences. Knowledge graphs may be the right way to store the data, Ethan says, but retrieving and formatting graph data for an LLM without overwhelming or confusing it remains difficult.

  • The desired system must infer behaviors never spoken aloud—for example, noticing receipts and repeated orders so “order something from this place” means buying the user’s usual meal. That requires reasoning across conversations, emails, calendars and actions rather than appending explicit preferences to a memory list.

  • Two Bee agents already combined calendars, locations and shared preference for French food to choose a new restaurant near Pacific Heights. Future ideas included social exchanges, family updates, personality inference and dating applications, but an agent experiment that sent Maria about 20 candles illustrates why explicit control is necessary over what an agent may disclose or purchase.

Alessio Fanelli

Hey everyone, welcome to the Latent Space podcast. This is Alessio, partner and CTO at Decibel Partners, and I'm joined by my co-host, swyx, founder of Smol AI.

swyx

Hey, and today we are very honored to have in the studio Maria and Ethan from Bee. Welcome.

Maria de Lourdes Zollo

Hi. Thank you for having us.

Alessio Fanelli

You are, I think, the first hardware founders we've had on the podcast. We've been looking to have an AI wearable hardware founder on for a while. I think we're going to have 2 or 3 of them this year, and you're the ones whose product I wear every day.

swyx

Thank you for making Bee.

Maria de Lourdes Zollo

Thank you for all the feedback and usage.

swyx

I've been a big fan. You were the speaker gift for the Engine No. 1 Sphere. Let's start from the beginning. What is Bee Computer?

Maria de Lourdes Zollo

Bee Computer is a personal AI system. You can think of it as AI living alongside you in the first person. It can capture your real-life context, and with that understanding, it can help you in significant ways.

The obvious one is memory, but that's really just the base use case: recalling things and reflecting. I know, swyx, that you like the idea of journaling, but you don't want to do it, while still having some kind of reflective summary of what you experienced in real life.

It's also about having the whole context of a human being and giving the machine the ability to understand what's going on in your life—your attitudes, your desires, and specifics about your preferences.

Ethan Sutin

That means it can not only help you with recall, but also with anything else you need it to do, because it already knows what you would want. Think about somebody you've worked with or lived with for a long time. They just know, without having to ask you, what you would want.

swyx

It's clear that this is the future. Personal AI is just going to be much more valuable with personal context.

Maria de Lourdes Zollo

One of the things we're really passionate about is understanding this personal context, because it will make AI more useful. Think about a best friend who knows you so well. That's one of the things we're seeing from users: they use Bee from a companionship standpoint and for professional use cases. There are many ways to use Bee, but companionship and professional use are the ones we're seeing the most right now.

swyx

It feels so dry to talk about use cases.

Alessio Fanelli

Yeah, well.

swyx

It's really an investor question. But at its best, don't you want your AI to know everything you've said and everywhere you've been? Wouldn't you want that? You shouldn't have to repeat every time what you like. It already knows that, and it does things for you based on that. I think that's really cool.

Great. Do you want to jump into a demo? Do you have any other final questions?

Alessio Fanelli

Before we do that, maybe we should cover the origin story. How did you two meet? Was this the first idea you started on, or was there something else before?

Maria de Lourdes Zollo

I can start. Ethan and I have known each other for 6 years. He had a company called Squad, and before that it was called All About, which was a personal AI company.

Ethan Sutin

Yeah, maybe you should start with that.

Maria de Lourdes Zollo

That's how I know Ethan. He was pivoting from personal AI to Squad, which was a co-watching-with-friends product. I had experience working with TikTok and video content, so I helped with the pivot, and we launched Squad.

That was really successful. In the end, the founders decided to sell it to Twitter, now X. Both of us joined X. We launched Twitter Spaces and many other products, and we continued working together until we started Bee.

Ethan Sutin

The interesting thing is that this isn't our first attempt at personal AI. In 2016, when I started my first company, it started out as a personal AI company. This was before transformers—before BERT even, just RNNs. You couldn't really do any convincing dialogue at all.

I met Esther, who was my previous co-founder, and we were both really interested in the idea of having a machine model and understand a dynamic human. We wanted to make personal AI. This was more geared toward younger people, because we had much more limited tools at the time.

I don't know if you remember, but in 2016 there was a brief chatbot boom. It was way premature, but it was when Zuckerberg went up at F8 and talked about the Messenger platform. People thought, “Oh, bots are going to replace apps.”

That lasted for about 6 months, and then everybody realized that these things were terrible and weren't replacing apps. But that's when we got excited and tried to make something where you could teach the AI about yourself. It was just an app that you chatted with. It would ask you questions and then give you some feedback.

What we learned was that some people really loved chatting and answering questions, but it was a lot of work just to manually teach an AI all these things about you. We had sentence similarity and other things we could try, but the technology was premature at the time, so we pivoted.

Alessio Fanelli

How did the first version get launched?

Ethan Sutin

We started in the same office as Hugging Face because Betaworks was our investor. They had a program called Botcamp. Betaworks is a really cool VC because they invest in things that are out there and way ahead of everybody else.

At the time, Botcamp had 6 companies. It was us and Hugging Face, and I think the other 4 are dead. Hugging Face was the one that really took off. A 30% success rate is pretty good.

It was just the 2 founders. It was a chat app for teenagers. A lot of people don't know that Hugging Face was originally like, “Hey friend, how was school? Let's trade selfies.” They built the Transformers library, I believe, to help make their chat app better. Then they open-sourced it, and it blew up. They realized that this was the opportunity, and now they're Hugging Face.

Maria de Lourdes Zollo

There were some people who were super passionate about it. Teenagers, for example, really like to talk about themselves, so they would reply to a lot of questions and talk about themselves. But most people don't really want to spend that much time talking.

Ethan Sutin

We went through Y Combinator. Long story short, we pivoted to consumer video, and that went really viral and got a lot of usage quickly. We ended up selling it to Twitter, worked there, and left before Elon—not related to Elon, but we left before that.

Alessio Fanelli

I should mention that this was the famous time when Elon had just come in. Esther was the famous one.

Ethan Sutin

Yes, she was my former co-founder. She was the one who was sleeping in the office. She stayed. We had left by that time, but she stayed for a while. Later, she left, or I think she got laid off. I think the whole product team got laid off. She was a director of product.

Then we thought, “Oh my God, things are different now.” We really started working on Bee again right before ChatGPT came out. We had an app version and were trying different things around it, but ultimately it was clear that there were some limitations.

A good question to ask any wearable company is: why isn't this an app?

Maria de Lourdes Zollo

Because we tried the app at the beginning.

Ethan Sutin

The idea behind Bee comes from “ambient.” If it were just around you all the time, rather than requiring you to open the app and make the effort to enter data, that led us down the path of hardware.

The sensors on this are microphones, so it's capturing and understanding audio. We started with hardware that also had a vision component, and we can talk about why we're not doing that right now.

If you wanted continuous audio understanding with your phone, it would monopolize your microphone. It would get interrupted by calls, and you'd have to remember to turn it on. That little bit of friction is actually a substantial barrier to the experience of having it with you all the time and living alongside you.

We do have Apple Watch support, so anybody with an Apple Watch can use Bee right away without buying any hardware. We worked really hard to make a version for the watch that can run in the background without draining your battery too much.

Even with the watch, there's still friction because you have to remember to turn it on, and it still gets interrupted if somebody calls you. You have to remember to go back and turn it on. We send a notification, but you still have to go back and turn it on because that's just the way watchOS works.

Maria de Lourdes Zollo

One of the things we're seeing from our Apple Watch users is that they like the Apple Watch integration. Many people start using Bee from the Apple Watch, and after a couple of days they buy the Bee because they just like wearing it.

We're learning from that, and it's really cool.

Ethan Sutin

Fundamentally, we like to think that personal AI is the mission. It's about understanding, connecting the dots, and making use of the data to provide some value.

The hardware is the ears of the AI. It's about integrating the incoming sensor data, and that's really what we focus on. If we can do it well and have a great experience on the Apple Watch, that's just great.

Ethan Sutin

I mean, but there are some platform restrictions that existing hardware makes it hard to overcome.

Alessio Fanelli

What do people do in 2 or 3 days that then convinces them to buy the product? This was a product where, after you use it for a while, you have enough data to start to get a lot of insights.

Ethan Sutin

For Apple Watch users, I believe that's because every time they receive a call, they need to go back to Bee and open it again. Or, for example, every day they need to charge the Apple Watch, and it reminds them to open the app every day. They feel like, “Okay, maybe this is too much work. I just want to wear the Bee, keep it open, and that's it. I don't need to think about it.”

I think they see the potential just from the watch, because even if you wear it for a day, we send a summary notification at the end of the day about the key things that happened to you during your day. I didn't even think—I’m not a journaling-type person. I was like, “Oh, I just lived the day. Why do I need to think about it?” But it's actually pretty interesting. Sometimes I'm surprised by how interesting it is to me just to be like, “Oh, yeah, that,” and see how it fits together. I think that's something people get immediately with the watch, but they're like, “Oh, I'd like an easier way to do this.”

swyx

It's surprising because I only know about the hardware. But I use the watch as a backup when I don't have the hardware. I feel like, because now you're beamforming and all that, this is significantly better.

Ethan Sutin

Yeah, that's the other thing. We have way more control over the hardware. With the Apple Watch, you're limited: you can't set the gain, you can't change the sample rate, and there's very limited framework support for doing anything with audio. Whereas if you control the hardware, you can optimize it for your use case.

The Apple Watch isn't meant to record this, and we can talk, when we get to the part about audio, about why it's so hard. This is audio at the hardest level because you don't know what environment it has to work in. This environment is great—we're in a studio—but afterwards, at dinner in a restaurant, it's a totally different audio environment. There are a lot of challenges with that. Having really good source audio helps, but there's still a lot that has to be done with machine learning to account for it. You can tune something for one environment or another, but it will make one good and the other bad. Making something flexible enough is really challenging.

swyx

Do we want to do a demo just to set the stage, and then kind of talk about it?

Ethan Sutin

Yeah, I think we can walk through the product.

For listeners, we'll be switching to video that will superimpose onto this video. If you want to see, go to our YouTube. Like and subscribe as always. And buy the B. Yes, and buy the B. While you wait for the video to come out, buy the B. Maybe we should have a discount code just for the listeners. Sure. If you want to offer it, I'll take it. Yeah. All right. Discount code SWX.

An important thing to mention is that the hardware is meant to work with the phone. If you look at Rabbit or Humane, they're trying to create a new hardware platform. We think that the phone is just so dominant, and it will be until we have the next generation, which is not going to be for 5 years. Maybe some Orion-type glasses that are cheap enough and light enough will come along, but that's going to take a long time. So we work with the phone rather than trying to replace it.

In the app, we have a summary of your days, but at the top is what's going on now, and that's updating continuously. Right now, it's saying that I'm discussing the development of personal AI, and that's just the ongoing conversation. Then we give you a readable form with little segments of the important parts of the conversations.

We do speaker identification, which is really important because you don't want your personal AI thinking you said something and attributing it to you when it was just somebody else in the conversation. You can also teach it other people's voices, so if there's somebody close to you, it can start to understand your relationships a little better.

We do conversation endpointing, which is a task that didn't even exist before because nobody needed to do this. If you had somebody's whole day, how do you break it into logical pieces? We use not just voice activity, but other signals to try to split it up, because conversations are a little fuzzy. They can lead into one another, and one can start before the next. We also use the semantic content of the conversation.

When a conversation ends, we run it through larger models to try to get a better sense of what was actually said, then summarize it, provide key points, describe the general atmosphere and tone of the conversation, and identify potential action items that might have come from it. At the end of the day, we give you a summary of your whole day, where you were, and a step-by-step walkthrough of what happened and what the key points were.

That's the base capture layer. If you just want to get a glimpse, recall something, or reflect, that's there. But really, the key is that all of this now feeds into generating personal context about you. We generate key facts known to be true about you, and there's a human-in-the-loop aspect: you have visibility into that. I have a lot of facts about technology because that's basically what I talk about all the time, but I also have some hobbies that show up.

I measure my day now by asking, “What is my token output of the day?” As a human, how much information do I produce? It's measured in tokens, and it turns out to be around 200,000 tokens a day.

In the recall case, we have a chat interface. The key here is recall: I probably have 50 million tokens of personal context, so how do you make sense of that and make it useful? I can ask simple recall questions, like details about a recent trip I took to Taiwan, where we were with our manufacturer. In real time, it has various capabilities, such as searching through my memories, searching the web, or looking at my calendar. We have integrations with Gmail and Google Calendar, so it can connect the dots between in-real-life and digital life. I just asked it about my Taiwan trip, and it gives me a breakdown of the details, what happened, and the issues we had around certain manufacturing problems. It also goes back and references the conversation, so I can return to the source.

Maria de Lourdes Zollo

Yeah, and not just the conversation, but the integrations as well. We also have Gmail and Google Calendar. If there's something there that's useful for more context, we can see that.

I never use the word “agentic” because it's strange, but it can search through your conversations, search through email, and look at the calendar. If I'm brainstorming about something that spans across all of those, it can search through my conversations, search through email, look at the calendar, and then, depending on what's needed, synthesize something with all that context.

swyx

I love that you did Spotify Wrapped. That was pretty cool.

Ethan Sutin

Yeah, one thing I did was make a Spotify Wrapped for my life in 2024. It's surprisingly good. It gave me metrics like, “You visited 3 countries and shipped X many beta devices.” It gives a lot of personal insights and reflection points.

swyx

That's fascinating. So that's the demo. We can show something that's in beta if we want to.

Ethan Sutin

The vision is not just about AI being with you, passively understanding you through your experiences, but also proactively suggesting things to you at the appropriate time. It's not just a tool; it can step in and suggest things to you.

Speaker 2

You’re asking for a recommendation for an Italian restaurant. Would you like me to look up some highly rated Italian restaurants nearby and send her a suggestion?

Maria de Lourdes Zollo

What I did was just send Ethan a message through WhatsApp on his personal phone.

Ethan Sutin

Basically, Bee is watching all my incoming notifications, and if a notification meets 2 criteria—is it important enough for me to raise a suggestion to the user, and is there something I could potentially help with?—this is where the actions come into play.

Because Maria is my co-founder and because it was a restaurant recommendation, something Bee could probably help with, it proposed that to me. I can respond either through the chat or with another push-to-talk, walkie-talkie-style button. It's actually a multipurpose button to toggle Bee on or off, but if you push and hold it, you can talk. So I can say, “Yes, find one and send it to her on WhatsApp.”

Bee is an Android cloud phone, so it has access to all my accounts.

We're going to abstract this away, and the execution environment isn't really important. We can go into technically why Android is actually a pretty good one right now. But it's searching for Italian restaurants, and we don't have to watch this. I could have my AirPods in my ears and my phone in my pocket. It's going to go to WhatsApp, find Maria's thread, send her the response, and then let us know.

Alessio Fanelli

Oh God. What's a good suggestion? I mean, it's not an Italian restaurant.

Ethan Sutin

Yeah. What did you say?

Alessio Fanelli

It's easy to say.

Ethan Sutin

Exactly. It's easy to say. Successfully found and shared. Let's see what the AI says.

Alessio Fanelli

What does the AI say?

Ethan Sutin

Bottega. I think it's called Bottega. It said it twice. I've been to one called La Cocina, I think. That was good.

Alessio Fanelli

Beretta's on Valencia Street. It's fine. The pizza isn't good.

swyx

It's not good?

Alessio Fanelli

Some of the pastas are good. I'm sorry—Beretta's, sorry. But there's this place, Delfina. Everybody here says La Pizzeria Delfina is amazing. I'm like, this is not—I don't know. It's great. North Beach Cafe.

swyx

For the record, since you're all Italians, what's the best Italian restaurant in SF?

Maria de Lourdes Zollo

Oh my God. I feel like I don't have one. No. I'm not sure.

Alessio Fanelli

The place you took us with Michele last night.

Maria de Lourdes Zollo

Vega.

Alessio Fanelli

The guy at Vega just happened to be Italian. It's in Bernal Heights.

Maria de Lourdes Zollo

He's nice.

swyx

What's the name of the place?

Alessio Fanelli

Vega.

swyx

Vega.

Alessio Fanelli

Okay. Cool, cool. We got the name.

swyx

But Vega—is that in Italian? What does it mean in Italian?

Maria de Lourdes Zollo

Vega means... It doesn't mean anything to me.

Ethan Sutin

The last thing I'll mention is also the wake-word detection. The phone can be off, and you can just say, “Hey, Alfred,” and then it just...

Alessio Fanelli

Yeah. Being able to have wake words enables some form of voice-agent features that even ChatGPT can never have, because they don't have the recording layer.

Ethan Sutin

Yeah, I think we have some other ideas, even beyond wake words, but I think it's interesting to see how people use the voice side of it. We're going to see a lot of innovation around hardware and stuff, but the real core is being able to do something useful with the personal context.

You always had the ability to capture everything. We've always had recorders, camcorders, body cameras, stuff like that. But what's different now is that we can actually make sense of and find the important parts in all of that context.

Alessio Fanelli

And then one last thing—I'm just doing this for you—is that you also have an API, which I think I'm the first developer advocate, because I had to build my own app.

Ethan Sutin

We need to hire a developer advocate.

swyx

Or just—yeah, hire AI engineers.

Ethan Sutin

The point is that you should be able to program your own assistant.

Alessio Fanelli

I tried Omi, the former Friend, the knockoff Friend. Real Friend doesn't have an API, and Limitless also doesn't have an API. I think it's very important to own your data and be able to reprocess your audio, although by default you don't store audio.

Ethan Sutin

Yeah.

Alessio Fanelli

And then also just to do any corrections. There's no way that my needs can be fully met by you.

Ethan Sutin

Yeah.

Alessio Fanelli

The API is very important.

Ethan Sutin

Yeah. I've always been a consumer of APIs in all my products.

swyx

We are API enjoyers in this house.

Ethan Sutin

Yeah, I am. It's very frustrating when you have to go build a scraper.

Alessio Fanelli

This whole combination—you have my location, my calendar—it's really, for me, the sort of personal assistant: just to write into it or to have it take action on external systems?

Ethan Sutin

We're expanding it. It's right now read-only. In the future, very soon, when the actions are more generally available, it'll be fully supported in the API.

Alessio Fanelli

Mhm. Nice. I'll buy one after the episode. The API thing, to me, is the most interesting.

Ethan Sutin

We do have real-time APIs, so you can even connect a socket and connect it to whatever you want it to take actions with.

Alessio Fanelli

Yeah. When I look at these apps—and there are so many of these products being launched—it's great that I can go on this app and do things, but most of my work and personal life is managed somewhere else. So being able to plug into it is nice.

swyx

I have a bunch of more human questions that I think people might have. One is: Is it good to have instant replay for any argument that you have? I can imagine arguing with my wife about something. There are these commercials now where it's basically 2 people arguing, and they can throw a flag, like in football, and have an instant replay of their conversation.

I think this is similar, where people can almost not argue anymore or lie to each other, because in a world in which everybody adopts this. I don't know if you've thought about it. And also, all the lies that all of us tell—there are sometimes things that contradict each other, because I might say something publicly and think something privately that I tell someone else. How do you handle that when you think about building a product like this?

Maria de Lourdes Zollo

I would say that I like the fact that the AI is an objective point of view. I don't care too much about the lies, but I care more about the fact that it can help me understand what happened and the emotions in a really objective way—a really critical and objective way.

If you think about humans, they have so many emotions. Sometimes something that happened to me—I don't know—I will feel really upset about it, or really angry, or really emotional. But the AI doesn't have those emotions. I can read the conversation, understand what happened, and be objective. I think that level of support is the one that I like more, instead of, “Did this guy tell me a lie?” I think that's not exactly what I find curious for me in terms of opportunity.

swyx

Are you going to interject in real time? Say I'm arguing with somebody. Will the AI say, “Hey, no, you're wrong. That person actually said...”?

Ethan Sutin

Proactivity is something we're very interested in. Maybe not specifically for settling arguments, but more generally. A lot of the challenge here is that you need really good reasoning to pull that off, because you don't want it constantly interjecting—that would be super annoying. You also don't want it to miss things that it should be interjecting.

It would be a hard task even for a human, to come in at the right times when it's appropriate. With the personal context, it's going to be a lot better, because if somebody knows about you. But even still, it requires really good reasoning to not be too much or too little, and just right.

swyx

The part about some things is that you say something to somebody else, but afterward I change my mind and send something. Every time I have a different type of conversation and data about me.

Maria de Lourdes Zollo

One of the things that we're learning is that humans evolve over time. For us, one of the challenges is actually understanding whether this is a real fact. So far, we have a human in the loop who can say, “Yes, this is true. This is not,” or they can edit their own fact. In the future, we want to have all of that automated inside the product.

Your question also hits on privacy, which I know we'll talk about. If you have some memory and you want to confirm it with somebody else, that's one thing. But it's for sure going to be true that in the future—not even that far into the future—it's just going to be normalized.

We're in a transitional period now, and I think one of the key things for us is to navigate that and make sure we're thinking of all the consequences and how to make the right choices in the way that everything's designed. It's more beneficial than it could be harmful, but it's just too valuable for your AI to understand you.

If it's Meta Ray-Bans or Google Astra, I think people are going to be more used to it. People's behaviors and expectations will change. Whether that's something that is going to happen now or in 5 years, it's probably in that range.

We adapt to new technologies all the time. When the Ring cameras came out, that was quite controversial. But now people understand that a lot of people have cameras on their doors. We're in that transitional period for sure.

swyx

I'll press on the privacy issue, because that's the number 1 thing that everyone talks about. Obviously, I think in Silicon Valley people are a little more tech-forward and experimental, whatever. But you want to go mainstream. You want to sell to consumers, and we have to worry about this stuff.

The baseline question—the hardest version of this—is the law. There are 1-party-consent states where this is perfectly legal. Then there are 2-party-consent states where it's not.

Ethan Sutin

Yeah, the EU is a totally different regulatory environment. But in the US, it's basically on a state-by-state level. In Nevada, it's single-party; in California, it's 2-party.

But it's kind of untested. There are different laws depending on whether it's a phone call or whether it's in person. In a state like California, anytime you're in public, consent doesn't come into play, because the expectation of privacy is that you're in public.

Maria de Lourdes Zollo

But we process the audio, and nothing is persisted. Then it’s summarized with speaker identification focused on the user. Now, it’s kind of untested legally—and I’m not a lawyer—but does that constitute the same as a recording? So it’s kind of a gray area and untested in law right now.

I think the bigger question is, because if you had your Ray-Bans on and were recording, then you have a video of something that happened. That’s different from having an AI give you a summary focused on you that’s not really capturing anybody’s voice. I think the bigger question, regardless of the legal status, is: What’s the ethical situation with that? Even in Nevada—or many other U.S. states where you can record everything and don’t have to have consent—is it still the right thing to do?

The way we think about it is that we take a lot of precautions not to capture personal information from people around you, both through speaker identification, through the pipeline, and then through the prompts and the way we store the information, to be really focused on the user. We know that’s not going to satisfy a lot of people, but if you do try it and wear it, it’s very hard for me to see anything that somebody wearing a Bee around me could capture that I would ever object to as a third party.

We’re in this transitional period where the expectation will become more normalized that it’s an AI. It’s not capturing a full audio recording of what you said; everything is fully geared toward helping the person understand their state and providing valuable information to them, not logging details about people they encounter.

swyx

You know, I’ve had the same question with Zoom meeting transcribers. I think there’s kind of a personal impact. There’s a Fireflies.ai recorder, and I just know that it’s being recorded. It’s not like I know whether I’m going to say anything different, but intrinsically, you kind of feel different because it’s not pervasive.

I’m curious, especially in your investor meetings, whether people feel differently. Have you had people ask you to turn it off in a business meeting and not record? I’m curious whether you’ve run into any of these behaviors.

Maria de Lourdes Zollo

What’s funny is, on my end, I wear it all the time. I take my coffee at Blue Bottle with it, or I work with it. Obviously, I’m working on it, so I wear it all the time. So far, I don’t think anybody has asked me to turn it off.

I’m not sure whether it’s because they’re really friendly with me or because they know I’m working on it, but nobody really cared.

swyx

This is because you live in San Francisco.

Maria de Lourdes Zollo

Actually, I’ve been in Italy as well, and Italy is super privacy-concerned. Europe is super privacy-concerned. Again, nothing. I don’t know. That, for me, was interesting. Nobody’s ever asked me to turn it off, even after giving them full demos and disclosing.

I think some people have said, “Well, in a personal relationship, my partner was initially kind of uncomfortable about it.” We heard that from a few users, and that was more in a personal relationship situation. The other big one is people say, “I do like it, but I cannot wear this at work,” because they think they’ll get in trouble based on policies.

If you’re wearing it inside a research lab or somewhere where you’re working on things that are sensitive, we’re adding certain features like geofencing, so you can set it to never be active at a particular location. We’re also looking at context fencing, so you can say, “If these topics come up, don’t capture anything.”

I’ve often explained it the other way: Maybe you only want it at work. So you never take it from work, and it’s just a work device, like your Zoom meeting recorder is your work device.

Professionals have been big early adopters. You say San Francisco, but our daily shipments of over 100 are going to addresses in Texas, which I think is our biggest state, and Florida—just the biggest states. A lot of professionals talk for a living, and we didn’t set out to build it for that use case, but I think there’s a lot of demand from white-collar people who talk for a living. We’re just starting to talk with them. I think they want to be able to improve their performance, understand what they were doing, and figure out how they can do better.

swyx

How do you think about Gong.io and some of these sales-training tools where you put on a sales call and then it coaches you through it, which are more verticalized, versus having a more horizontal platform?

Maria de Lourdes Zollo

I’m not super familiar with the space. Like I said, we weren’t building for that use case, so it’s kind of a surprise to us. But I think those are interesting. I’ve seen there are a bunch of them now, right? It kind of makes sense.

I’m terrible at sales, so I could probably use one, but it’s not my job fundamentally. Maybe in a lot of situations, it could be useful. We’ve also heard from people with restaurants that, if they’re able to understand whether they’re doing well, that can be valuable.

In general, I think a lot of people like to have that double-check: Did it do this well, or can you suggest how I can do better? We had a user who said he used Bee for job interviews and then asked Bee, “Actually, how do you think my interview went? What should I do better?” I like that. I’m like, “Oh, that’s actually a personal coach in a way.”

swyx

But I guess the question is: Do you want to build all of those use cases, or do you see Bee more as a platform where somebody’s going to build the sales coach that connects to Bee, so you’re kind of the data feed into it?

Maria de Lourdes Zollo

I don’t think this is just a data feed. It’s more of an understanding engine. Definitely in the future, having third parties use the API and build out all the different use cases is something that we want to do, but the initial case we’re trying to build is that layer for all of it to work.

We’re not trying to build all those verticals, because no startup could do that well. But it’s been fascinating to see. I’ve done consumer for a long time, and consumer is very hard to predict. It’s hard to know what’s going to be the killer feature. We really believe that’s the future, but we don’t know exactly what process it will take to gain mass adoption.

swyx

The killer consumer feature is whatever Nikita Bier does. Social apps for teens. Yeah, well, I like Nikita, but he’s good at building bootstrap companies and getting them very viral, then selling them, and then they shut down.

Okay, so you just came back from CES?

Maria de Lourdes Zollo

Yeah. Crazy. It was my first time in Vegas and my first time at CES. Both were overwhelming.

swyx

First of all, did you feel like you had to do it because you’re in consumer hardware?

Maria de Lourdes Zollo

We decided to be there and have a lot of partner and media meetings, but we didn’t have our own booth, so we decided to skip that. But we decided to be there and have a presence, even just us, and speak with people.

swyx

Hard to stand out.

Maria de Lourdes Zollo

Yeah, I think it depends on what type of booth you have. I think if you can prepare a really cool booth, it can be pretty cool.

swyx

Have you been to CES? It can be pretty cool. It’s massive. It’s like 80,000 or 90,000 people across the Venetian and the convention center. To me, I always wanted to go, even just as a fan of—

Maria de Lourdes Zollo

Yeah, you wanted to go. Growing up, I think CES kind of peaked for a while, and it was like, “Oh, I want to go there. That’s where all the cool gadgets, everything, is.”

swyx

There are a lot of cool vacuums and pet tech, and dog stuff.

Maria de Lourdes Zollo

There’s a lot of robot stuff, new TVs, and new cars that never ship.

swyx

I think this time last year was when Rabbit and Humane launched at CES, and Rabbit kind of won CES. Now, this year, there were no wearables except for you guys.

Maria de Lourdes Zollo

It’s funny because it’s obviously AI everything. Every single product—like, a toothbrush with AI. We saw a hair blower, literally a hair dryer with AI. That was cool.

But another difference around us is that we didn’t want to do a big, overhyped, promised kind of Rabbit launch. I mean, hats off to them on the presentation and everything, obviously, but we wanted to let the product speak for itself and get it out there. We were really happy with the interest we got from the media and some of the partners there, so it was definitely worth going.

I would say that if you’re in hardware, it’s just about how you make use of it. To do a big Rabbit-style launch or have a huge show there, you need to plan that 6 months in advance, and it’s very expensive. But if you go there, everybody’s there. All the media are there, and there are a lot of pre-show events where it’s great to talk to people in the industry.

All the manufacturers and suppliers are there, too, so we learned about some really cool stuff. We met with them; they have thermal energy capture.

And it's like, oh, could you maybe not need to charge it because they have a thermal that can capture your body heat.

swyx

What?

Maria de Lourdes Zollo

Yeah, they're here. They're actually here in Palo Alto. They have a Fitbit thing that you don't have to charge.

swyx

How much power can you get from that? What's the power draw for this thing?

Ethan Sutin

It's more than you could get from the body heat, it turns out, but it's quite small. I don't disclose the technical details, but I think solar is still more realistic. They also have one where the face of it is just a solar cell, and that is more realistic. Or kinetic. Kinetic, apparently—they seem to think it wouldn't be enough. Kinetic capture is quite small, I guess.

swyx

Well, I mean, watchmakers have been powering things with kinetic energy for a long time. We don't have to talk about that. I just wanted to get a sense of CES. Would you do it again?

Maria de Lourdes Zollo

I definitely would. Okay, you're just a fan of CES. From a business point of view, it doesn't make sense.

swyx

I happen to be in the conference business, right? So I'm kind of just curious.

Maria de Lourdes Zollo

Yeah, so I would say that, without the booth and with really straightforward conversations that were already planned, 3 days was okay. I think it was okay. But if you need to invest in a booth, that's not cheap.

swyx

Which is how much?

Maria de Lourdes Zollo

A 10-by-10 is $5,000. But on top of that, you have the financial costs. A 10-by-10 is like one of the super-fancy things, and some companies have—I think we'd probably be more in the 6-figure range to get a booth.

I mean, I think that, yeah, it's very noisy. We heard that it's very, very noisy. Obviously, everything is being launched there, from cars to cell phones, so it's hard to stand out. But I think going in with a plan of who you want to talk to was worth it. We had a lot of really positive media coverage from it, and we got the word out, so I think we accomplished what we wanted to do.

swyx

I mean, there's some world in which my conference is kind of the CES of whatever AI becomes.

Maria de Lourdes Zollo

Yeah, I think that—don't do it in Vegas. Don't do it in Vegas. That's the only thing I didn't really like there.

swyx

Okay, that's great. Amazing. Those are my favorite ones. You cannot fit 90,000 people in San Francisco.

Maria de Lourdes Zollo

That's really the problem. You need to do multiple locations. You can do Moscone and then have one—

swyx

That's what the Salesforce conference is. GDC is how many?

Maria de Lourdes Zollo

That might be 50,000, right?

swyx

Okay, form factor, right? My way to introduce this idea was that I was at the launch in Solaris. What's the old name of it? Newton? Of Tab when Avi first launched it. He was like, “I thought through every form factor. Pendant is the thing.” And then we got the pendant for the original one—the first one, which was a pendant—and I took it off and forgot to put it back on.

So you went through pendant, pin, and now bracelet, and maybe there are AirPods—a sort of earphone—in the future. What was your iteration to that?

Maria de Lourdes Zollo

Yeah, so we had, I believe, 3 or 4 iterations, and one of the things that we learned is that people don't like the pendant. In particular, women don't want to have anything here on the chest because maybe they have another necklace or other stuff.

swyx

You just ship a premium one that's gold. We're talking about some fashion—some big fashion. There's something there. This is where it helps to have an Italian.

Maria de Lourdes Zollo

Exactly. Some big Italian luxury. I can't say anything, sorry.

The bracelet actually came from the community because they were like, “I don't want to wear anything as a necklace or a pendant.” Also, the one that we had—I don't know if you remember—was a circle and was really bulky. People didn't like it.

And I actually don't dislike it. We were running fast when we did that. Our thing was that we wanted to ship them as soon as possible, so we weren't overthinking the form factor or the material. We just wanted to be out.

But after the community organically told us, basically, all of them were like, “Why don't you just do the bracelet? It's way better. I will just wear it, and that's it.” So that's how we ended up with the bracelet, but it's still modular. I still want to play around with the fact that it's modular, and you can take it off and wear it as a clip. In the future, maybe we will bring back the pendant, but I like the fact that there is some personalization.

Right now, we have 2 colors, yellow and black. Soon we will have other ones, so we can play a lot around that.

swyx

I think the goal for the form factor is for it to be not super invasive, right? Something that's easy. In the future, smaller and thinner—not like Apple's obsession with thinness, but it does matter, the size and weight. We would love to have more context because that will help, but to make it work, I think it really needs to have good power consumption and good battery life.

With the Humane, swapping the batteries—I have one. I think the Humane is pretty incredible, some of the engineering they did, but it wasn't geared toward solving the problem. It was just too heavy. The swappable batteries are too much to manage—the heat, the thermals. It's too much for a light-interface thing.

Maria de Lourdes Zollo

Yeah, like that. That was cool.

swyx

It's cool. It's cool, but if you have your hand out here and you want to use your phone, it's not really solving a problem because you know how to use your phone. It's got a brilliant display, but you have to learn how to gesture with this low-resolution laser. The laser is cool—the fact that they got it working in that thing, even though it did overheat—but it's too heavy, too cumbersome, and too complicated with the multiple batteries. So something that's power-efficient and thin, both in the physical sense and in the edge-compute kind of way, so that it can be as unobtrusive as possible.

Maria de Lourdes Zollo

Users really like it. I like when they say, “Yes, I like to wear it and forget about it,” because I don't need to charge it every single day. On the other version, I believe we had 35 hours or something, which was okay, but people just prefer the 7-day battery life.

swyx

Oh, this is 7 days?

Maria de Lourdes Zollo

Yeah.

swyx

Oh, I've been charging it every 3 days.

Maria de Lourdes Zollo

Yeah, you can keep it full for 7 days.

swyx

The other thing that occurs to me is maybe there's an Apple Watch strap.

Maria de Lourdes Zollo

Yeah.

swyx

So that I don't have to double-watch. I have—

Maria de Lourdes Zollo

Yeah, that's the other one.

swyx

Yeah, I thought about it. I also saw the ones that you can put back on the phone. There are a lot of possibilities.

There's a competitor called PLAUD. It's not really a competitor; they only transcribe, right?

Maria de Lourdes Zollo

Transcribe, but they're very good at it.

swyx

Yeah, no, they're great. Their hardware is really good, too, and he just launched the pin, too.

Maria de Lourdes Zollo

Yeah, I think the MagSafe kind of form factor has a lot of advantages, but also some disadvantages. You can definitely put a very large battery on that, and so the power consumption isn't as much of a concern. The downside is that the phone is in your pocket. I think form factors will continue to evolve, with more sensors, less obtrusiveness, and easier use. We have a new version.

swyx

Okay, looking forward to that. Whenever we launch this, we'll try to show whatever, but I'm sure you're going to keep iterating.

Last thing on hardware, and then we'll go on to the software side, because I think that's where you guys are also really strong. Vision—you wanted to talk about why no vision?

Ethan Sutin

Yeah, I think it comes down to the fact that when you're a startup, especially in hardware, you work within the constraints. Vision is super useful and super interesting, and it's what we actually started with. There are 2 issues with vision that make it not the place we decided to start.

One is power consumption. You have to trade off your power budget. Capturing, even at a low frame rate, and transmitting over the radio actually takes up the majority of the power. So you would really have to have an unacceptably large and heavy battery to do it continuously all day. We have some novel alternative ways that might allow us to do that, and we have some prototypes.

The other issue is form factor. Even with a wide field of view, if you're wearing something on your chest, it's obviously not going to capture the field of view of what's interesting to you. The wrist isn't really that much of an option, and if you're wearing it on your chest, it's often not going to capture what's interesting to you. That leaves you with your head and face.

Anything that goes on the face has to look cool. I don't know if you remember Snap Spectacles. That was kind of like the first—

swyx

Yeah, but they weren't very successful, and I think one of the reasons is that they were so weird-looking. Their camera was so big on the side.

If you look at the Meta Ray-Bans, where they're way more successful, they look almost indistinguishable from a Ray-Ban. Meta invested a lot into that, and they have a partnership with Qualcomm to develop custom silicon. They also have a stake in Luxottica now, so they're coming from all angles to make glasses.

I don't know if you know Brilliant Labs. They're a cool company. They make Frame, which is a cool, hackable pair of glasses, and they're really good on hardware.

Ethan Sutin

But even if you look at the frames, which I would say are from the most advanced kind of startup, there was one that launched at CES, but it isn’t shipping yet. The one that you can buy now still isn’t something you’d wear every day, and the battery life is super short. So I think the challenge of doing vision right off the bat would require quite a bit more resources. Audio is such a good entry point, and there’s also the privacy around audio. If you had images, that’s another huge challenge to overcome. Ideally, personally, I would have all the senses, and we’ll get there.

Alessio Fanelli

Okay, one last hardware thing, because I have to ask this before we move to software. Were either of you electrical engineers?

Ethan Sutin

No, I’m CS, and I’ve taken some EE courses, but prior to working on the hardware here, I had done a little bit of embedded systems—very little firmware—but we luckily have somebody on the team with deep experience.

Maria de Lourdes Zollo

Yeah. I’m just like, you have to become hardware people. I learned how to worry about supply chain, power, and radio. I would say this about hardware—and I know it’s been said before—but building a prototype, learning how the electronics work, learning about firmware, and developing this is fun for a lot of engineers, and it’s all totally achievable, especially now with the tools we have. Stuff you might have been intimidated by, like, “How do I write this firmware?”—now, with Claude Sonnet, you can get going and actually see results quickly. But I think going from a prototype to actually making something manufactured is an enormous jump, and it’s not all about technology: supply chain, procurement, regulations, costs, and tooling.

Alessio Fanelli

The thing about software that I’m used to is that it’s funny: you can make changes all along the way and ship them. But when you have to buy tooling for an enclosure, that’s expensive. Own tooling?

Maria de Lourdes Zollo

You have to.

Alessio Fanelli

Don’t you just subcontract out to someone in China?

Maria de Lourdes Zollo

We make the tooling? No, no. You have to have CNC and a bunch of machines. Nobody makes their own tooling, but you have to design it and submit it, and then they go 4–6 weeks later. If there’s a problem with it—

Alessio Fanelli

Right.

Maria de Lourdes Zollo

Well, then you’re not making any of your enclosures.

Alessio Fanelli

What resources or websites are most helpful in your manufacturing journey?

Ethan Sutin

I think it’s different depending on the product, because hardware is so specialized in different ways. I would say that, for example, you should choose a manufacturing company and speak with other founders. They’ll give you the language and some tips about who is good and who is not, who specializes in something versus somebody else. Some people are good in plastics, and some people are good in PCBs. For us, it really helped at the beginning to speak with others and understand who is around. I worked in Shenzhen and lived almost 2 years in China, so I have an idea about different hardware manufacturers and all of that. Soon I’ll go back to check things out. I think it’s good also to go in person and check and see the right people.

Alessio Fanelli

We did some stuff domestically, and if you have that ability—the reason I say ability is that it’s very expensive—you can build out some proofs of concept and do field testing before taking it to a manufacturer. Despite what people say, there’s really good domestic manufacturing for small quantities, at extremely high prices. We got our first PCB and assembly done in LA, because the defense industry supports a lot of good manufacturing that can do quick turns. It’s like, we need this board, we need to find out if it’s working, we have this deadline, and we want to start; if you want to have it done and fabricated in a week, they can do it for a price. But I think everybody’s trending, even for prototyping, toward moving offshore now, because in China you can do prototyping and get it within almost the same timeline. The thing is, with manufacturing, it really helps to go there and establish the relationship.

Maria de Lourdes Zollo

Yeah, my first company was a hardware company, and we did our PCBs in China. It took a long time. Now things are better, but this was, I don’t know, 10 years ago or something like that. I’ve heard this too: if it’s something where you don’t have the relationships, they don’t see you, they don’t know you, you might get subcontracted out, or they’re not paying attention. But if you have the relationship and a priority, it’s usually good.

Ethan Sutin

We ended up doing the fabrication and assembly in Taiwan for various reasons, but it really helped that you went there at some point.

Maria de Lourdes Zollo

Yeah, we’re really happy with the process. But the whole process of choosing the right people—choose the right people—but also just sourcing the bill of materials and all of that stuff: I guess if you have time, it’s not that bad, but if you’re trying to really push the speed, it’s incredibly stressful.

Alessio Fanelli

Okay, we’ve got to move to software. The hardware may be hard for people to understand, but what software people can understand is that running transcription and summarization, all these things, in real time every day for 24 hours a day, isn’t easy. So you mentioned 200,000 tokens for a day. How do you make it basically free to run all of this for the consumer?

Ethan Sutin

Well, I think that the pipeline and the inference—people think about all these tokens, but as you know, the price of tokens is dramatically dropping. You guys probably have some charts somewhere that you’ve posted, and if you see that trend, 250,000 input tokens isn’t really that much, right? The output layers?

Alessio Fanelli

You do live?

Ethan Sutin

Yeah. Speech-to-text is the most challenging part, actually, because it requires real-time processing and then later processing with a larger model. One thing that’s fairly obvious is that you don’t need to transcribe things that don’t have any voice in them, right? Good voice activity detection is key, because the majority of most people’s day isn’t spent with voice activity. That’s the first step to cutting down the amount of compute you have to do, and voice activity detection is a fairly cheap thing to do—very, very cheap.

The models that need to summarize don’t need a Claude Sonnet-level model. You do need a Claude Sonnet-level model to execute things like the agent, and we’ll be having a subscription for features like that because, although now with DeepSeek R1, we’ll see—we haven’t evaluated it. There are already models that can perform at that level. I was going to say in 6 months, but—

Alessio Fanelli

DeepSeek R1?

Ethan Sutin

Yeah, I mean, not that one in particular, but they’re already there and can perform at that level. Self-hosted models help with the things where you can.

Alessio Fanelli

So you’re self-hosting models? You’re fine-tuning your own ASR?

Ethan Sutin

Yes. I will say that I see everything trending down in the future, although I think there might be an intermediary step where things become expensive, which we’re really interested in because the pipeline is very tedious and requires a lot of tuning. That’s brutal because it’s just a lot of trial and error. Wouldn’t it be nice if an end-to-end model could just do all of this and learn it? If we could do transcription with an LLM, there are so many advantages to that, but it’s going to be a larger model and hence more compute. We’re optimistic that maybe we could distill something down, and we kind of focus more on reducing the cost of the existing pipeline or trying the next generation, because it’s very clear that all ASR, all speech-to-text, is going to be pretty obsolete pretty soon. Investing in that is probably a dead end because it’s just going to be obsolete.

Alessio Fanelli

It’s interesting. I think when I initially invested in Tab—this shows you how wrong I was—I thought, “Oh, this is a sort of razor-and-blades model where you sell cheap hardware and make up a subscription, a monthly subscription.” Now I just checked: Friend is a 1-time sale, $99; Limitless is a 1-time sale, $99; these guys are a 1-time sale, $49. And Friend is free? What? When you probably invested, how much was 1,000,000 input tokens at that time, and what is it now? It’s a fascinating business, and there’s a lot to dig into there, but just getting that perspective out there is not something that people think about a lot. You obviously have thought a lot about it.

What about memory? I think this is something we go back and forth on: you’re just memorizing facts and then understanding what is a preference, and adjusting the facts that you think about a person? Any learnings from that? I know there are a lot of open-source frameworks now that do it. Did you build all of your own infrastructure internally?

Ethan Sutin

Yeah, we did. I evaluated and used a lot of them in other projects. I think there are a few different tasks or things that revolve around memory. One is retrieval, obviously, and when you need to find something—even if you have a large corpus—how do you find it? I think existing RAG pipelines will probably also be obsoleted. The frameworks—I have not found one. There’s no general way to do RAG that works; it’s really highly dependent on the data.

swyx

So, if you're going to be customizing something that much, you get more bang for the buck from designing it all yourself. A lot of these frameworks are great for getting going quickly, but I think it's really interesting when you're trying to do memory for a person, because memories decay, right? I'm going to London, then I come back; I'm not going to London anymore.

Ethan Sutin

What we've learned is that doing traditional embedding and RAG is suboptimal. We built our own using small models to do massively parallel retrieval, which I think is going to be more common in the future. To represent a person, we still require some human loop. This is an ongoing project, and we're learning every day. How do you correct the model when it gets something wrong about you?

Right now, we have things that are super confirmed—ground truth about you because a human accepted it—but ideally that step wouldn't be necessary. Then we have things that are fuzzier. The more stuff that we know is true, the more accurate we are when we're trying to decide whether this fuzzy stuff is true, because if you have the context, it's probably not true.

So, I think one of the core challenges is how to handle both retrieval and modeling, especially when you're dealing with noisy source data. Even in an ideal world, if you had perfect transcription and were going off that, that's still not enough information, right? Even if you had visuals, it's still not enough. There's still going to be some misunderstandings, so how do you not let that damage the value of it and make it unrecoverable and uncorrectable?

swyx

Yeah, one way I think about it is that I usually like to order the same thing from the same restaurant if I like it, but I'm not saying that out loud. Are these the types of behaviors that you can capture? When you ask about a favorite restaurant, I would want it to give me restaurants I've already been to and liked. Or if I'm like, “Hey, just order something from this place,” it should reorder the same thing because it knows that I like to get the same thing again. But I feel like today, most agent memory things that I see people publish are just, you know, write down the data thing.

Ethan Sutin

Yeah, I think that's why the reasoning—in our case, giving it time to consider all of the sources it has—is really important. Look at the emails, see the receipts, and then look at the conversations to see what I've mentioned. Then be able to take enough time to search through all the context and connect the dots is really important. I don't know; some of the agent memory stuff is like key-value with RAG on top, and the results are just not complete enough when you have a growing corpus, managing decay, and hallucinations that might be in the source material.

Alessio Fanelli

So, this is where people usually bring in knowledge graphs.

Ethan Sutin

Yes.

Alessio Fanelli

And do you use them? Is it for speed, or what are the issues?

Ethan Sutin

We don't extensively use knowledge graphs. It's something—we didn't talk also about the potential future social aspects—but the problem with knowledge graphs that we found is that they're great for representing the data, but using them at inference time is challenging. It's just the LLM understanding the graph input. It's not in the training data, for sure. I think the graph is the right kind of way to store the data, but then you need to have the right retrieval and format it in a way that doesn't overwhelm or confuse what you're trying to do.

swyx

Should we ask about social?

Alessio Fanelli

Yeah, I know. I thought you were going to go into it.

swyx

Yeah, what's that?

Ethan Sutin

Not directly related to graph retrieval or graph knowledge bases, but the idea is that you have your personal context, and then other people can query it. It can divulge some things that you would have full control over. Then Maria and I are trying to negotiate where we're going to have dinner. There can be an exchange between the agents.

Alessio Fanelli

Yeah, there can be an exchange between the agents.

Maria de Lourdes Zollo

So, how can my agent speak with Ethan's agent? Both of them know our location, what we like, and where we went in the past. Even if we have our calendars integrated, they know when we're free. They can interact with each other, have a conversation, and decide on a place to go for us.

Alessio Fanelli

Wow.

Ethan Sutin

And we did that. It was really cool for me because they suggested a nice French restaurant that we went to in the end.

swyx

That you'd never been to?

Maria de Lourdes Zollo

That we'd never been to. They saw that we both liked French food, and we were both in Pacific Heights. This was really trivial.

Alessio Fanelli

Yeah, it's a trivial toy use case, but I guess, in terms of having used it for a while, if I wanted to buy you a gift—

Maria de Lourdes Zollo

Oh my God, you bought me a bunch of candles, now that I think about it. This is another use case. I was like, “Yeah.”

Ethan Sutin

When we were testing the agent, a bunch of candles from Amazon showed up at her door.

Maria de Lourdes Zollo

Yeah, because I really love candles, but I didn't expect 20.

Ethan Sutin

Yeah, it's a lot—extreme, but how do you manage that? What's okay for your Bee to divulge, and to whom? Shouldn't you get an authorization request every time for personal context?

Maria de Lourdes Zollo

Yeah, yeah, yeah. A human would have to sign off on it, but then I wouldn't have to guess.

swyx

There's this culture that's very alien to everyone else outside of San Francisco, and outside of the Gen Z bubble in San Francisco, which is sharing location. I can tell exactly where my close friends are right now in the city. It's normal here, and it freaked out everyone who's not here.

Maria de Lourdes Zollo

Yeah.

swyx

So maybe we can share preferences—who we like, or even small updates about your day. My parents would love that because I don't do that.

Ethan Sutin

Yeah, so now there's no friction. It can just be more or less automatic.

Maria de Lourdes Zollo

Dating? I was always trying to avoid dating as a startup founder.

swyx

Everyone hates it?

Maria de Lourdes Zollo

We thought about it. Sometimes people ask us, “Oh, you know so much about me. Can you measure compatibility with somebody else or something like that?”

swyx

Yeah, probably there is a future, and maybe somebody should build that.

Alessio Fanelli

I think on our end, we were like, “No, this is—I’ll build on your API.”

swyx

My sister's actually a personality psychology professor, and she studies personality. We were at Thanksgiving with my parents, and I was like, “Give me my Big Five,” which is the personality type. “Does it know my Big Five?” You just ask it to consider everything and give you your Big Five. My sister said it was pretty accurate. I didn't agree with it because it said I was disagreeable, but she seems to think I'm agreeable.

Alessio Fanelli

You disagree that you're disagreeable?

swyx

Yeah. What other proof do you need, then? I think I'm very agreeable.

Maria de Lourdes Zollo

But I think we did get some users who were like, “Oh, if we're a couple—” We had couples actually buy the product together. Both members of a couple bought our hardware, so there is something there.

swyx

Another test is the Myers-Briggs. I know you don't like that one.

Maria de Lourdes Zollo

No, no, no. OCEAN is cooler than Myers-Briggs. Everyone stop using my MBTI. Use my OCEAN.

Ethan Sutin

For me, it was on point every time.

Maria de Lourdes Zollo

Awesome.

swyx

Anything else that we didn't cover? Any cool underrated things?

Ethan Sutin

Go to B. dot computer $49.99 and buy the device. That's the call to action. You're hiring? We are hiring—for sure, AI engineers. Nice. What is an AI engineer? Somebody who's scrappy and willing to work with us. I think you coined the term, right? So you can tell us. People have different perspectives, and what is useful for you is different from what is useful for me.

So, anyway, it's useful. I think always-on AI is really going to explode. There's going to be a lot from startups, but incumbents, too, and there are going to be all kinds of new things that we're going to learn about how it's going to change all of our lives. I think that's the thing I'm most certain about. And be in the AI. Well, thanks very much. Yeah, this was a pleasure. Yeah, we'll see you launch whenever that launch is happening. Yeah, thanks. Thank you.

Bee AI: The Wearable Ambient Agent | BidClub