[BidClub_]
Latent Space · · 64 分钟

个性化 AI 语言教育——Andrew Hsu 与 Speak

Andrew Hsu

YouTube
TL;DR
  • Speak 的核心押注——语音和语言模型将在5到10年内达到超人水平——如今已成长为一家 ARR 远超5,000万美元的企业。Andrew Hsu 表示,“现在已经有80%到90%的技术到位了”,这让公司最初的2016年愿景成为可能:用纯软件打造一名能比真人教师更快帮助学习者达到流利水平的导师。由于模型演进路径始终足够有吸引力,公司在没有转型的情况下熬过了痛苦的4到5年。
  • 韩国的产品市场匹配,来自把产品、市场和客户都收窄,而不是把它们做大。2018年,Speak 放弃免费、多语言内容目录,转向面向单一市场的引导式英语课程并改为收费,意识到“人们不想做选择”。如今 Speak 已是韩国最大的英语应用,约有6%的人口试用过,公司称其在日本和台湾也已走上正轨。
  • Hsu 将 Speak 定义为语言学习的“Gen 3”:追求功能性流利,而不是游戏化学习或电子化教材。学习者反复练习句型和真实场景——比如和 Uber 司机交谈——直到说话变得自发,“几乎就像在健身房训练”。产品的私密环境也消除了在人类面前犯错的心理成本。
  • Speak 的技术栈是混合式的,并不依赖单一前沿 API。其定制流式 ASR 基于大量非母语语音训练,负责对延迟敏感的练习;Whisper 和 LLM 则处理开放式表达、语义反馈与辅导。Hsu 的运营模式是先用产品把模型能力“榨干”,前沿能力进步后再重复这一过程。
  • 下一阶段的规模化引擎,是由人类教学法约束、并由持续演进的个体学习者知识图谱驱动的 AI 生成课程。Speak 希望获得“100倍的内容”、10倍的语言,以及最终100倍的语言对,借助导师和课程编写 agent,同时由人类审核教学大纲和课程。词汇、句型与聚类错误最终应汇总为一个完整的 Speak Score:即便不围绕考试训练,54分与5分之间的方向性差异也足够有意义。
  • 实时语音已经近在眼前,但瓶颈在单位经济和交互设计,而非模型的原始智能。OpenAI 的实时 API 定价更适合替代按小时计费的人工劳动,而不是让消费者连续对话数小时;在 Speak 的规模下,一个错误可能带来“数百万美元”的成本。Hsu 还称,除非把轮次检测纳入其中,否则首段音频延迟只是“虚荣指标”,因为学习者可能在回答中途犹豫10秒。
  • Hsu 认为实时翻译更像补充,而非生死威胁;语言则是更广泛学习平台的滩头阵地。他指出,德语翻译成英语必须等到句末动词出现,认为 Speak 的亚洲用户想要的是直接的人际连接,而不只是一条巴别鱼;但他仍预计 Speak 会将翻译纳入产品。B2B 已开始向沟通、管理和酒店服务技能扩展,支撑了他的更大判断:AI 将重塑学习,尽管“现实世界的惯性极其强大”。
摘要 · 为研究而整理的核心内容

1. Speak 能够存活,是因为最初的模型判断从未改变

  • 2011年,19岁的 Hsu 进入第一届 Thiel Fellowship,当时他已经通过加速教育进入研究生阶段。该项目为20岁以下、愿意离开学校追求几乎任何目标的人提供10万美元,这一机会“改变了人生”;项目也让他结识了后来的 Speak 联合创始人,对方加入了第二届。

  • 2016年,创始人休假整整一年研究 AI,包括与当时即将完成研究生学业的 Andrej Karpathy 交流。他们的判断异常具体:未来5到10年,语音和语言模型将达到超人水平,因此由“纯软件、纯 AI”构成的语言导师将成为可能。

  • 在迎来正确时点之前,时机错得很严重。早期4到5年痛苦不堪,创始人一次次问自己:“我们为什么要做这个?”但他们始终没有转型;公司后来在台北重新播放最初提交给 YC 的申请材料时,里面的论断与 Hsu 今天的说法几乎逐字一致。他如今估计,“现在已经有80%到90%的技术到位了”。

2. 聚焦与高价策略在韩国转化为产品市场匹配

  • Speak 最初的“红色应用”提供跨多种语言的免费内容包,让用户自行选择学习内容,结果失败。2018年,团队推倒重来,只服务学习英语的韩国人,设计课程和新的课型,并把学习路径铺设好:“他们打开应用已经要消耗一部分动力……不想再做另一个选择。”

  • 放弃免费同样关键。高级订阅价格筛选出了本来就有动力学英语的人,绕开了部分动机问题。没有哪一项改变是银弹;Hsu 将转机归因于3到4年间不断累积的经验教训。

  • 韩国并非一开始就注定胜出,团队当时差点选择台湾。这个决定带有一定偶然性。落地之后,首尔的补习班和“挤满教室的摩天大楼”展现出强烈需求。团队的押注是:如果能在一个充斥真人竞品、且人们真正关心流利度的市场胜出,就能建立强大且可迁移的产品市场匹配:“如果我们真的能在这个市场取得进展并赢下来……那我们大概就拥有真正扎实的 PMF。”

  • 本地化带来了可信度。Speak 的第一名员工 Sungjae 帮助团队把细节打磨到按钮文案,早期用户得知公司来自美国时都十分震惊。如今 Speak 已是韩国最大的英语应用,约6%的人口试用过,业务 ARR 超过5,000万美元。

3. “Gen 3”语言学习训练的是自动表达,而不是考试知识

  • Hsu 将市场分成 Rosetta Stone CD-ROM 代表的“Gen 1”、移动优先的“Gen 2”,以及原生于 AI 的第三代。他认可 Duolingo 的游戏化,但认为它最接近一款有生产力的手游;Speak 的目标则是通过辅导和开放式练习,培养功能性流利。

  • Speak 的核心方法并不围绕词汇和语法组织。它先教授句型,再让用户组合、重复这些句型,“几乎就像在健身房训练”,直到说话变得自发。角色扮演强调与 Uber 司机交谈之类的场景,而不是回忆教材规则。

  • “魔法入门”希望用户一开始就感知这一新品类:导师询问学习目标,随后由另一个 LLM 将回答转化为抽象摘要,而不是展示完整转录。结果仍未定型——从安装到注册的转化率更低,因为开口说话比点击更难;但试用启动率更高。Hsu 明确表示:“我们还不知道。”

4. 定制 ASR 与前沿模型分处不同的延迟区间

  • 在 LLM 出现之前,Speak 就构建了定制语音识别模型,并从日常课程中积累了大量非母语英语音频。这套系统依然有价值,因为核心录音环节必须快速流式处理;公司利用自身数据微调模型、理解自己的用户,而不是把通用转录当作终点。

  • Whisper 在2022年的出现跨过了期待已久的门槛。4名员工在闭眼状态下测试一名韩国初学者的音频,所有人都听不懂;Whisper 却正确完成了转录。对 Hsu 而言,这不只是一次改进,而是语音识别达到“超人水平”的直接证据。

  • 随后,Whisper 与 GPT-3.5 Turbo 让 Speak 超越了听说跟读。产品可以理解自发回答、解释错误,并指出母语者会选择不同的词或表达方式。在此之前,更简单的 Speak 产品已经在韩国做到数百万美元 ARR;这些模型则把补充式口语练习变成了更完整的导师。

  • Hsu 不接受“更好的基础模型让 Speak 早期工作失去意义”的说法。定制 ASR 仍服务于快速、受约束的交互;前沿模型则处理语义和开放式表达。路线图纪律是:先把产品做到能够“榨干模型能力”,下一次能力跃迁后再重新构建。

5. AI 生成课程是杠杆化尝试,而不是无人参与的流水线

  • Speak 内部的“Speak method”、教师、脚本和制作工具带来了沉重的运营负担。如今扩张需要“100倍的内容”、10倍的语言,以及最终100倍的母语到目标语言组合,远远超出其洛杉矶工作室和手工写作流程的供给能力。

  • 公司正在构建导师 agent、课程编写 agent,以及基于大型 LLM 的流水线,用于搭建课程框架和撰写课程。Hsu 谨慎使用“agent”一词,但目标很明确:在进入更多语言和市场的同时,保持组织规模小巧。

  • 评估仍是最难的部分。即便是向一名新的人类作者解释,为什么某一节略有不同的课程比另一节更符合 Speak method,也很难说清楚。Speak 高度依赖内容团队,正在开发由模型评分的评估体系,未来或许会基于内部数据对课程 agent 做强化微调;Hsu 称这项工作“还处在相当早期”。

  • 主持人将其描述为工作转型,而非简单消除岗位:一个人不再独立编写2门课程,而是可能审核机器生成的50门课程。Hsu 认同这种杠杆化判断——人类仍会审核教学大纲、课程以及具体句子,但同一支团队最终应能推出“100倍的课程”。

6. 流利度是崎岖、多维且扎根于真实任务的能力

  • Hsu 举的核心例子非常务实:学习者能否去墨西哥城,在街边塔可摊点单?具备这项能力,并不意味着他能谈论家庭,因此“流利度的前沿非常崎岖”。Speak 因而将语言能力建模为多个子分数,而不是一条均匀向上的单一阶梯。

  • Speak 正在开发领域专用知识图谱,用来组织学习者长期积累的词汇、句型和错误集群。这些维度最终应汇总为 Speak Score:Hsu 承认,54分单独看意义不大,但它与5分之间的差距是直观可感的,也能对应学习者实际可以完成的具体事情。

  • 主持人质疑 Speak 没有对齐任何公认考试目标。Hsu 的回答是,Speak 不“围绕考试教学”;未来可能提供备考服务,但当前优化目标是现实世界的熟练度。从 A0 到 B1,学习者共享一套相对线性的概念骨架;到了中高级阶段,路径会明显分化,知识图谱则会围绕个人弱项调整这套基础。

  • Speak 教授的是日常语言,而不是“教材英语”,但当前规模迫使产品采用诸如标准美式英语这样的务实默认值。眼下每种语言都选择一个标准,Hsu 预计未来会出现区分更明确的口音和方言。发音的重要性低于表达想法:用户应该“真的动嘴……发出这些声音”。一款仅支持英语的教练目前基于 Speak 的音素数据、使用 wav2vec 微调,负责给单词评分,后续计划扩展到句子和更多语言。

7. 实时语音受制于代码切换、轮次检测和成本

  • 真正的语言导师必须支持代码切换:学习西班牙语的英语使用者,即使在同一句话中,也应能自然地在两种语言之间切换。能够正确发出这种混合语音的 TTS 模型很少。主持人建议根据检测到的语言进行路由,但 Hsu 指出,子词级切换会让简单路由和拼接失效,因为生成结果听起来不再像人声。

  • Speak 很早就获得了 OpenAI 实时 API 的访问权限,但 Hsu 澄清说,它当时还没有进入生产环境。这样的定价更适合替代按小时计费的客服 agent,而不是让消费者练习数小时。Speak 已经接近目标,但定制 WebRTC 基础设施加上业务规模意味着,架构上的错误“将让我们付出数百万美元”。

  • 首个应用是一个3到5分钟的教学课程,在听或观看内容与短时互动对话之间交替切换。课程采取半引导式设计,Speak 构建了脚手架,以便在这些模式间稳定切换;它的定位是增强现有教师视频,而不是立即替代每一节课。

  • 按传统方式衡量的请求到首段音频时间,用 Hsu 的话说,“其实是一个虚荣指标”。真正有意义的区间始于学习者停止说话之后,而轮次检测可能额外增加1秒甚至更久。语义 VAD 适合流利对话,却无法应对学习者在句中停顿10秒的情况,因此定制轮次检测成为影响感知延迟的主导问题。

8. 语言是更广泛 AI 学习公司的滩头阵地

  • Speak 如今已在另外40个国家教授英语,西班牙语和法语版本已经上线,并计划在当年推出数种新增语言。B2B 业务大约在一年前作为副项目启动,随后“突然开始奏效”;Hsu 预计它会与仍以消费者业务为主的公司一起,逐渐变得有意义。

  • Hsu 认为翻译威胁在技术和情感上都不完整。德语句末的动词意味着英语翻译在整句到达前无法推进,因此他认为翻译永远不可能“真正、真正完美”。更重要的是,亚洲用户想要自我提升和直接连接:“他们希望能够看着你的眼睛说英语。”不过 Hsu 仍预计 Speak 最终会把翻译纳入学习体验。

  • Speak for Business 未来或许可以通过 Mac 应用或浏览器集成接入工作文档,但 Hsu 称这“会带来一大堆麻烦”。当主持人提醒员工可能不愿使用一款暗中评估管理技能的语言工具时,Hsu 立即承认,这些产品应该分开。

  • 更长远的路径将从语言延伸到沟通、酒店服务、管理,最终进入数学或生物学等学科。Hsu 的收束判断带有张力:AI “可能是我们迄今创造的最具变革性的技术”,但湾区之外的人们的生活“几乎没有变化”。他的处方是增加应用构建者和规模化消费产品;当 AI 焦虑出现时,他会回到一个可控变量上:“我只关注我们的用户。”

Speaker 1

Hey everyone, welcome to the Latent Space podcast. This is Alessio, partner and CTO at Decibel, and I'm joined by my co-host, swyx, founder of Small AI.

Speaker 2

Hello, hello. We're back in the studio with Andrew Hsu of Speak. Welcome.

Andrew Hsu

Thank you for having me.

Speaker 1

I have to start this off. I didn't prep you on this at all, but you were a Thiel Fellow in 2011.

Andrew Hsu

First class.

Speaker 1

First class. Yeah. Is that the one with SBF?

Andrew Hsu

No, he was, I think, several years later, actually.

Speaker 1

What was it like? Just talk about the—

Andrew Hsu

That's a good question. I haven't been asked that one in a while. It was a really crazy idea at the time and very controversial. I think the first few years of the fellowship were definitely, "Let's just find 20 people under 20 and give them $100,000 to drop out of college." It was no holds barred. You could do anything. You could be doing some crazy research idea, a startup, anything.

I actually met my current co-founder at Speak. He was in the second year of the fellowship, and I made many close friends from the first few years. For me, it was life-changing. I had a very unusual path where I did finish college. Unfortunately, I was in grad school at the time because I went to school really early.

Speaker 2

Yeah, I was like, weren't you too old? You know, Thiel likes them young.

Andrew Hsu

I was 19 at the time and in grad school. It was a very accelerated path, but I knew at the time that I was going to leave grad school and do startups anyway, and the timing lined up really well.

Speaker 1

Yeah, yeah. Vitalik, I think, was also in a later year.

Andrew Hsu

Ah, yeah. Damn. Okay, anyway. But the first 2 years had some crazy successes. You know, Dylan from Figma—I mean, a lot of people.

Speaker 1

Awesome. Feel free to bring in those stories as and when, because obviously only you know those kinds of people. You are now CTO and co-founder of Speak. I would say, from a very early stage, Speak was one of the most successful and prominent AI products that anyone would know was doing well, and teaching English to Koreans was your core remit at the time. How did that all come about?

Andrew Hsu

It's funny that you say that because, despite our current revenue scale and objectively how successful we are, we've always operated in a market—at least initially—on the other side of the world. We've been much more popular in the Eastern world and a bunch of Asian markets, and relatively unknown in the West. It hasn't really felt like we've had that sort of awareness until the past few years, really.

The brief story is that my co-founder and I, back in 2016, were fascinated by the promise of AI, and we spent a year-long sabbatical basically learning everything we could. We talked to Karpathy back then, when he was just finishing grad school, and did a lot of self-study and research. We were so convinced that speech models were going to develop rapidly, language models were going to develop rapidly, and that in a 5- to 10-year span they would become superhuman.

We were utterly convinced of this future, and we saw that the way people learn things—and specifically learn languages, which was a very human-based thing if you really care about fluency—would completely change. We'd be able to build language tutors that were pure software, pure AI. That was the genesis story of Speak.

It took much longer than we expected to build a great product and find good product-market fit. The first few years were very painful, and without this really compelling vision of the future, we would have quit. We actually never pivoted.

Last year, we brought the entire company to Taipei. We do this company trip every year, and we played our original Y Combinator application video on screen. It was really funny because the things we were saying in that video were the exact same things that I still say today about the long-term vision and what we're building toward. That was really cool to see.

Speaker 2

Can you summarize the long-term vision again?

Andrew Hsu

As speech models and language models become superhuman, that would let us create an AI language tutor that would help you become fluent faster than any human could. I think 80% to 90% of the technology is here now.

Speaker 2

And you have this big focus on speaking, obviously. It's in the name of the company.

Andrew Hsu

Yeah, that's right.

Speaker 2

I think the speech models were maybe a little delayed compared to the text models. Did you ever think, "Okay, maybe speech is just not going to work for this use case"? What were the valleys of discomfort, and what were some of the pivotal releases and models where you thought, "Okay, it's going to work. It might take a little longer, but it's going to work"?

Andrew Hsu

We've always done custom speech stuff. The first act of the company, if you will, was before LLMs—before 2022, when Whisper came out and when ChatGPT came out.

In the 2 to 3 years before that, we felt like we found product-market fit in South Korea and then started growing, still only in that market and still only teaching English. We developed custom speech-recognition models, and users were speaking into the app all day. We had a ton of this non-native English-speaker data, and we would use that to fine-tune models and understand our users better. We still do that today.

It's important for us that the core recording loop in many of our lessons is extremely fast, so we're very latency-sensitive. There are many other product surfaces within the app today that are more LLM-powered, where it's more open-ended, real tutoring, and where we actually give you feedback on what you said semantically and so on. That stuff is more Whisper-powered and more LLM-powered. But we've always had a very fast core ASR loop that's been fully custom.

Speaker 2

I just onboarded to the app earlier today. Unlike other apps, there's this tutor conversation that you do for onboarding. I'm guessing that it's mostly LLM-based, and then you're judging the person's response. I selected Spanish, and the conversation was in Spanish via text to start. From there, it started to create lessons for me.

Was that all unlocked from LLMs, or can you now have these conversations and then bring people into the speech flow?

Andrew Hsu

Before that, we had much more traditional app onboarding. There's still a lot of open questions around what the proper onboarding UX is, because a lot of people start using Speak and aren't in a situation where they can actually speak out loud. We have fallback flows and so on, but it's something we're actively experimenting with.

We call that Magic Onboarding, and it was a new thing we built that was more conversational. We wanted it to feel more like you were talking with a tutor, and that the tutor was learning things about you. We would use that later to personalize the experience.

Speaker 1

Is there a structured output behind that? Is there anything that you found implementing Magic Onboarding? People always want to improve onboarding. What was the uplift, or was there one?

Andrew Hsu

We still don't know yet. The interesting thing is that, in general, because it's speaking-based—which is a much higher barrier than just tapping a multiple-choice button—we see that the install-to-sign-up rate is a decent amount lower, but the trial-start rate is higher.

It's still an active experiment that's running, and we're trying to be very agile about testing many different formats. I don't think I have the final answer yet, but I think the intent—the real vision we're going for here—is that as soon as you download the app from the App Store, maybe you see it in an ad, the first interaction when you have a fresh open of the app should feel pretty futuristic.

It should feel like, "Okay, this is the new, AI-native, next-generation way of learning a language to fluency." That's always been our ambition. We want to build something that wasn't possible before LLM and AI technology.

Speaker 2

I think I wanted to go back to the onboarding. There's a general idea that when you replace a form with a voice bot, you need to have some kind of state machine behind the hood to drive questions like, "What else don't I know about you? Let me proactively ask that." I'm just wondering if you had any insights there, or is it literally just a state machine?

Andrew Hsu

We tried both, actually. Right now, I think probably what you saw is a state machine, but I think things should move in a direction where it's much more of a natural conversation.

There is a general sense of a goal in the prompt that you can specify, and part of the hard thing here is all the guardrails. Once you start talking about what you had for breakfast yesterday and trying to be antagonistic toward the system, things start really going off the rails.

For a bunch of these experiences, we're pretty careful about the fallbacks, and we have a lot of evals around that. But I think where it should end up is just feeling like you have a quick 3- to 5-minute conversation with your tutor, then it knows a lot about you, and then you create your account.

Speaker 2

And you create memories?

Andrew Hsu

Yeah. We store what you're saying and summarize it. In the experience, the tutor will ask you a question like, "What are your goals around learning English or the language?" Then we'll use a separate LLM prompt to summarize.

It's not the full transcript of what you said that you see. It's more of an abstracted, "Okay, here's what you care about."

And we think that's a better product experience.

Speaker 1

What were some of the other key tenets of the product? Obviously, language learning is one of those consumer markets where dozens of companies always try to get started, and you have these old companies like Babbel and Duolingo. Speaking—the act of speaking—was a big part of it. I think this memory stuff is great. If you try some of the other apps, they always start to re-ask you the same things that you got wrong before, but you're not really learning. Is there anything else that is maybe not as obvious from the outside in the design of the app and the product that you think is really different?

Andrew Hsu

I would say, from a macro level, this is actually a pretty new product category: AI-powered language learning. All these apps that you mentioned—Duolingo, Babbel, and so on—are more like the Gen 2 of language learning. If you think about it, Gen 1 was Rosetta Stone, if you remember, right? CD-ROMs in airports. Then Gen 2 was basically mobile.

You have these very casual, massively popular mobile apps like Duolingo, where I think the comp is probably closer to a mobile game—something that feels productive, something that's very engaging and very gamified. Duolingo is really leaning into the gamification, and they've done an amazing job of that, to be clear.

Speaker 1

Yeah, they might be the world's best people at it.

Andrew Hsu

Our view is that LLMs and AI now enable Gen 3 of language learning, which is very AI-native and focused on functional fluency. That's why we do all these role-plays and let you practice Spanish by talking to your Uber driver.

We don't teach vocabulary and grammar; we teach sentence patterns. We try to get you to repeat and drill and drill and drill, almost like you're in a gym, until it's automatic, because that's what speaking is, right? It has to be spontaneous and automatic.

In terms of the other aspects of the design, though, we went through many, many iterations over the first few years of starting the company. This is what I was mentioning about being really painful in the first 4 or 5 years. In fact, the current version of the Speak app is not the first thing that we launched.

We had something that we called internally the Red app, which had a red app icon and still had a similar logo. It was more about packs of content instead of courses, where you could choose any topic that you wanted to learn. It was for learning many different languages. It was essentially not a very directed experience, and it didn't really work. It was free and very basic, but in 2018 we tore everything down and realized that we had to fully change what we were doing.

That's when we decided to focus on South Korea, specifically on teaching English. We built a bunch of new lesson types and created our courses so that the experience was much more on rails. We realized people don't want to choose. They're already using some of their motivation on a daily basis just to open the app. They don't want to make another choice after that, right? Just tell me what to do. Give me a big button, and then I can tap it and just start a video lesson or whatever.

We also, pretty critically, abandoned the free version and went straight premium. We kind of sidestepped the motivation question that way because we knew that there were a ton of users who really wanted to learn English and were already really motivated. We wanted to basically filter for those users.

I wouldn't say there was one silver bullet. It was the combination of many learnings over 3 or 4 years. Then that started really growing in South Korea. From there, I guess phase 2 was really 2022, when LLMs came out and Whisper came out.

That allowed us to go from this more supplemental speaking-practice tool to more full-featured language tutoring, where we could use LLMs like GPT-3.5 Turbo back then to give you direct feedback on your wording and tell you, "That was kind of a weird thing to say. A native speaker would say it this way," or suggest using a different word.

Speaker 1

I always do a poor job of doing this, but can we get some headline numbers, just to get a sense of scale? I think maybe some audiences don't know. Where are you now in terms of your reach?

Andrew Hsu

We're now the biggest English app in South Korea. We do billboards and big celebrity campaigns—that sort of scale. We're very popular there. I think 6% of the Korean population has tried us.

We're well on the way in a bunch of other Asian markets like Japan and Taiwan. The Asian markets are currently our mainstay. We also teach English in 40 more countries. We're coming to the US as well. We have Spanish and French live, and several more languages are coming this year. That's a huge focus of the company right now.

In terms of revenue scale, we're well over $50 million ARR. It's a pretty simple business model. It's mostly consumer. The B2B stuff is super exciting, and that's also growing really fast. I think it'll be a really meaningful part of the business.

Speaker 1

Did you just start B2B?

Andrew Hsu

About a year ago. It was very much a side bet or experiment at first, and then it just started working.

Speaker 1

Of course it's going to work.

Andrew Hsu

Yeah. Now it's like, okay, this is part of the future, right? This is a real thing.

Speaker 1

What's the B2B relationship between language learning and real-time AI translation? At Google I/O, they had one of those Google Beam things for conferencing that does real-time translation.

Andrew Hsu

People always ask this, right? They're always like, "What happens when the Babel fish comes?"—when real-time translation comes. The Babel fish is from The Hitchhiker's Guide to the Galaxy, right?

Speaker 1

Right. Yes, exactly.

Andrew Hsu

The counterexample that I always have, which I think is quite illustrative, is that in German, the verb is at the end of the sentence, right? So if you're trying to do real-time translation from German to English, as an example, you can't actually make any progress on the English until you hear the whole German sentence and know what the verb is at the end, right?

The minimum latency there is the full sentence. That's an example of the technical blocker for why it'll never be truly, truly perfect.

But also, I think besides that, if you talk to all of our users in Asia, they don't want a translator. The reason they're trying to learn English is to make themselves a better person and to connect with other people. They want to be able to look you in the eye and speak English—speak the same language as you, right? So it's actually a very different thing.

I think what will end up happening is that we will build a real-time translation feature into Speak and have it integrated into the learning experience. There's always that human side, right?

Speaker 1

I'm dating a Romanian woman.

Andrew Hsu

Yeah. My wife is trying to learn Italian. There's always that.

Speaker 1

Yeah, absolutely. I want to double-click on Korea. It's a very insightful, smart decision. Maybe people only know Korea through K-pop, but actually I think a lot of Americans learn Korean because of K-pop. That's a side thing.

You could have done Taiwan; you could have done China. I remember seeing a documentary about how China was crazy about English, or "Mad About English." I think that was the title of the documentary. Was it obvious? Were you sure when you went into Korea, or was it just a test?

Andrew Hsu

We visited a bunch of Asian countries when we were thinking about how to relaunch things and focus in. We almost chose Taiwan, actually. But I think it was a little bit serendipitous.

Our first employee is Korean and was my cofounder's college roommate, actually. When my cofounder visited Seoul to check out the market, he asked Sungjae to come along as essentially a translator and to facilitate. I think that just went really well, and it was very obvious from being on the ground in the market that Korea is pretty obsessed with learning English.

There is every human-based solution possible: English academies, classes, skyscrapers full of classrooms, stuff like that. Our logic was basically, if we can really make headway and win this market that's chock-full of these human competitor products and all these people who fundamentally care about fluency, then we probably have something pretty real and strong PMF that we could win other markets with.

That was the original logic. So far, it's been working.

In hindsight, it's super weird, right? We were definitely sitting in an office here in San Francisco, operating with users in a market all the way on the other side of the world. It would not have worked without Sungjae.

I have to give him a lot of credit here because we paid a lot of attention to the specific wording of button text in the app and localized strings. We had a lot of reports from users pretty early on that they were shocked that it was an American company. They thought it was a Korean company, right? Because you can always tell. There's always some weird wording or whatever, but there wasn't in Speak, and I think that probably had a large, intangible effect.

Speaker 1

Yeah. Focus—attention to detail. Tech stack: this is 2018. What were you rolling? You just did ASR, and there was no LM, so BERT maybe?

Andrew Hsu

We actually had really no LM component of it. So all of the content—oh, yeah, another thing we did that I forgot to mention was that we decided we needed to fully own all the content. The way that we teach is all in-house, all sort of thought through from first principles. We built this thing called the Speak method, which is basically a pedagogical philosophy around teaching sentence patterns that you drill and then sort of combine into higher-order patterns.

And all of that was in-house with our content team and our teachers. We built a lot of internal tooling to make this possible. There was just a lot of operational overhead, I would say. This is something we've struggled with in scaling to many more languages, and that's a big research effort within the company right now.

We're building a mobile product, right? My co-founder and I have always just loved apps and been big iPhone users. So we cared a lot about the app being native, feeling great, and being high-performance. The DNA of the company was always consumer.

Frankly, my co-founder and I had never worked in a real company. I dropped out of grad school, had a few failed startups, and then eventually started Speak. Kia never worked in a real company either; he had just worked at startups in the past. So I think, frankly, consumer was the only path. I don't think we could have done anything else. We just didn't know enough.

I think that has served us well, though, in terms of really caring about the craft of it and wanting to build something that felt not 90% to 95%, but 95% to 100% in terms of polish.

Speaker 1

Was it hard to build an engineering team that did that at the time? ML engineering was very academia-driven back then, and then you had the more consumer stuff, which was maybe more nascent.

Andrew Hsu

And it's mobile. I'm now realizing that our story is very weird. In addition to having a market on the other side of the world, our first iOS engineer, whom we hired through a YC referral, was in Slovenia. If you don't know where Slovenia is, look it up on Google Maps, but it's a pretty obscure little country, right?

Speaker 1

Nowhere. Yeah.

Andrew Hsu

And then we needed to hire a back-end engineer, and one of his best friends was a great back-end engineer, so we hired him. Then this happened 4 more times, all in the same city. Then we were like, “Okay, we should probably just open a physical office.”

Speaker 1

So, for Slovenia?

Andrew Hsu

Yes. So for several years, we had an engineering office in Slovenia.

Speaker 1

What?

Andrew Hsu

And then a few people here in San Francisco. We still do. Now we have 90% of our core product development team here in San Francisco, in an office in FiDi. We're really only hiring here, but for the first several years, that was another very interesting cultural aspect of the company, I guess.

Speaker 1

I think a lot of early-stage founders have to do that. It's the only people they can afford or whatever.

Andrew Hsu

Yeah.

Speaker 1

What are your tips that make that remote stage work?

Andrew Hsu

For us, it wasn't really a price thing. I think we legitimately thought he was the best person that we interviewed, and then it just kind of happened that way.

Speaker 1

Yeah, it's not a price thing. It's more about remote work, right? A distributed team, early-stage. A lot of people say, “No, you have to move everyone to SF or your startup will die.”

Andrew Hsu

Yeah. I don't think that we were good at remote work. I don't think that my personality or my co-founder's personality is inherently very good at async, just to be perfectly frank. I actually think that, almost in spite of it, we made it work. It was a little bit brute force: I would just sync with them every single day, right?

There was pain because of the time-zone overlaps. It was exactly the most inconvenient. But for several years, we did that and got really good at the cadence of it. I think they were excellent engineers as well, so it worked out. But if I had to do it over again, I probably wouldn't do it.

Speaker 1

It's hard to say. Yeah. Shall we move to phase 2 on the LLM side? That's when the AI started opening up. And when did you invest?

Andrew Hsu

This was 2022. That was also when Whisper dropped. Whisper was a really exciting moment for us. Since we started the company and made that prediction—“Okay, in 5 or 10 years, speech models and language models will become superhuman”—Whisper was really that magic moment where we were like, “Oh, I think what we predicted is here.”

I pretty distinctly remember this moment in the office when we got access to the model and were testing it on an audio clip of a very beginner English learner in Korea saying something. If you closed your eyes as a human, you'd have no idea what they were saying. There were 4 of us in the room; we all closed our eyes, and none of us had any idea, and the model got it right. So, I mean, superhuman.

I think that was the moment we'd been waiting for. At the same time, LLMs were on the ascent, and ChatGPT would come out, I think, on Thanksgiving of 2022. GPT-3.5 Turbo came out, and I think we realized very quickly that all the pieces were clicking now. We had what we needed at our fingertips to go from something that was listen-and-repeat, where the user would see something on screen, hear a reference of the teacher saying the thing, and then just repeat the thing. It was very simple.

Still a great product, by the way. It still grew to several million ARR in South Korea, so clearly there was a big market need for that.

Speaker 1

Pre-Whisper. Yes. Wow. This is from 2019 through 2022. Yeah, that's the grind. You needed to hang in there.

Andrew Hsu

Yeah. And again, I think there were many moments when things weren't working from 2017 through 2019. We were looking in the mirror and saying, “Why are we doing this? This is crazy.” But I think we were so convinced about the vision, we just couldn't believe that the vision would not—

Speaker 1

Yep.

Andrew Hsu

—come true. So, fast-forward to 2022, the pieces started coming together. We realized that we could start building something that felt more like a language tutor that could give you feedback, that could start explaining to you why you did something wrong. That was Act 2 of Speak: a true English tutor.

Speaker 1

This is something that a lot of founders struggle with today. It's like, “I'm kind of building something hoping that the models get better later.” How did you feel once the models got better? Did you feel like, “Okay, I am ahead of the curve because I built all this history of building product and doing all this work”? Or did you almost feel like, “Okay, we spent all this money and time building these models, and now we're just going to use Whisper”?

Andrew Hsu

It was purely positive for us. We still kept using our custom ASR system because it was streaming in real time, really fast, and really well fine-tuned. Whisper wasn't streaming, so it was a different use case. We used it for the more spontaneous stuff.

And I think in almost every way, we were just really excited because, pretty directly, as the frontier of model intelligence improved, it would just unlock things on our roadmap that were locked before, if that makes sense. We still operate in that mode today, where we take a model and then try to think about, “Okay, how do we saturate model capability by building product on top of it?” And then it happens again, right? Then we build and saturate the model capability again. I think that's a really cool paradigm to think about.

But all the LLM stuff basically allowed us to build a tutor for English, and we still didn't have real-time voice, for example. But the barriers are coming down now. Obviously, it's a really hot topic. We're actively building out a real-time voice platform that we can build a lot more verticalized, specific lesson experiences on top of, and I'm super, super excited about.

I don't think they're going to replace our current lessons. They're going to be more immersive, just a different thing, probably for more advanced learners. Still language learning, though, not broadening out from language.

Speaker 1

Yeah, so I think that language learning is interesting because it is so universal. 99% of people you know have certainly tried to learn a language, and it's so hard, right? Becoming fluent just has a huge failure rate, and it's something people are willing to pay for. So I think that has been just an amazing beachhead for us, and I think we'll be doing language learning for a long time.

There's a huge company to be built here, but our even longer-term ambition is really this idea that, even beyond language, we think AI will reinvent how people learn anything, right? It already has for me. I use ChatGPT to learn things every 10 minutes, and I think I'm just naturally a very curious person, so whenever I'm thinking about something, I want to know more about it, and then I'll naturally go to ChatGPT, and then I'll learn about it.

It's unlocked this entirely new dimension of learning, and I'm spending way more time learning as an adult, which is really cool. I want to bring that in a more structured, systematic way to everyone. So I think that's the vision behind the language-learning product. I'm curious to double-click on just the tech side.

Andrew Hsu

Mm-hm.

Speaker 1

We talked a little bit about the content that you own and develop in-house, and we talked a little bit about the onboarding and memory. I assume that you have conversational memory as you go, right? And any other major pieces of the puzzle that really unlocked it for you?

Andrew Hsu

So, there are a few things I can talk about.

I think one thing is that, in order to go from teaching English to teaching a bunch more languages, we needed to really figure out more direct AI content generation. That was pretty hard because it's hard to scale our little studio in LA, where we shoot a lot of the video lessons. All of the scripts were written manually before by our content team, but we want 100× more content, right? And 10× more languages. Eventually, 100× more language pairs, which is how we think about it. It's like, what's your native language, and then what language are you learning?

Really, the only way to do that is to make it more AI-generated. Very much like an AI-native company, we want to be at the frontier here. We want to keep a small team and have as much leverage as possible through these types of tools. That's a big active area where we're building out—I think people overuse the word “agent,” but we have a tutor agent, a curriculum-writing agent, and a giant LLM-based pipeline that creates curriculum, scaffolds it in the right way, and writes the lessons themselves. That's a big active area that will basically help us scale to a lot more markets and a lot more languages.

So, that's one big thing. Another big thing is that we care a lot about fluency, obviously. Specifically, we want to be able to quantify how fluent you are. If you're learning Spanish, it's like, okay, what does it mean to be fluent, right?

Speaker 1

And a real-world test for that.

Andrew Hsu

We care about real-world fluency: your ability to go to Mexico City, go to a street taco stand, and actually order, right? That's very functional fluency in one aspect. You might be really good at that but be completely unable to talk about your family, right? So, the frontier of fluency is very jagged, but we're very pragmatic, and we care a lot about meeting user goals and helping them become fluent at what they care about.

We're thinking a lot about, okay, how do you quantify that? How do you actually store a knowledge graph of everything you know about Spanish, in terms of the vocabulary you know or don't know, the patterns that you know or don't know, and the mistakes you made using Speak over the last month that are clustered?

Speaker 1

You said the magic word: knowledge graphs. Is that live? Is that experimental?

Andrew Hsu

There are aspects of it that are live, and it's a very multidimensional system. We think of it as there being many aspects of fluency, right? There are many subscores. We have a few of them that are currently live, and we're actively developing other aspects of it. Then all of those will fold up into a more holistic fluency score.

The idea is that eventually, once we have a complete enough picture, everything will fold up into a number that we call the Speak Score—a very holistic measure of just how good you are at Spanish, right? Obviously, 54 is kind of meaningless by itself, but it does give you a general sense. Being at 54 versus being at 5 is very different, right? I think everyone can intuitively understand that.

Speaker 1

And surprisingly, I would have grounded it more in the real world. We’ll get you to pass this exam that is a standard, like the CEFR standard or whatever.

Andrew Hsu

The way that we think about that is that we don't really teach for the test. I think it's possible in the future that we'll do a test-prep product, but in general, we care about real-world proficiency in various functional situations. The way that we think about it is, if you're at this level, then these are the things you can do.

We have that a lot in Italy. I grew up in Italy, so English is my second language. There are a lot of people who pass a lot of tests and get high grades in all the classes, and then they travel to the US and the UK, and it's hard to speak because they don't—I feel like the hard part is being in the conversation. When I started, my writing and reading were much higher than my conversation, which doesn't really help you if you're traveling somewhere.

Speaker 1

That's me for Chinese, because my parents spoke Mandarin to me growing up. I can understand a nontrivial amount, but I'm very bad at speaking. Mhm. Yeah. I heard there's a good language-learning product. I have one question on the course generation. How do you evaluate the product? When you're asking the AI to generate courses, how do you figure out whether the courses are going to be good?

Andrew Hsu

We rely very heavily on our content team, and we're trying to build out an eval suite. It's really hard. The illustrative example here is that, as we try to hire and train new content writers on our content team, it's so nuanced. There are many different aspects of training them in the Speak method, how to write the right types of lessons, and articulating why this form of lesson, which is subtly different from this other form of lesson, is better, right? We try as hard as we can to articulate that.

I think forming a sense of evals using model-graded evals is one piece of it. I also think that, in the future, a really good curriculum or lesson-writing agent will probably be reinforcement fine-tuned on a lot of our internal data as well. That's something we're experimenting with, but it's still pretty early.

Speaker 1

This seems like a great example of AI removing jobs. It's like, oh, you're creating the courses with AI, so you don't have a person. But it's actually, instead of one person creating 2 courses, having that person review 50 courses that the AI generates. That's kind of how you're seeing the content.

Andrew Hsu

The way that we see it, not just for our content team members but also, I think, perfectly applicable to engineering, is that it's leverage. It just allows you to do 100× as much in the same amount of time. We still need human review of the syllabus, the curriculum, the specific lines, and so on. But the hope is that this will allow us to launch 100× more courses.

Speaker 1

A lot of language is colloquial. I think the way that you put it on one of our episodes one time was that the Italian taught in school is not the Italian Italians speak. How much do you adjust for informal versus formal?

Andrew Hsu

Entirely. That's one of our fundamental tenets: We don't teach textbook English or textbook language. We try very hard to teach Gen Z slang. We don't go quite that far, but we try to teach slangy, very casual conversational language that is actually what real people use. Like you said, that's usually very, very different. If you pick up a typical English textbook in Korea, it's all really traditional and weird formulations, and it's not how people actually speak.

Speaker 1

I know you're going to release Italian soon, so I can give you a hand on that. I know in the US there aren't that many dialects. There are accents, but most of the language—the words that people use—are similar. Spanish, for example, is very different when spoken in Argentina than when spoken in Mexico. How are you going to adjust for that? Or maybe you don't?

Andrew Hsu

I would say that currently, for example, we teach American English—standard American English. We don't really teach other accents or other dialects. For now, given how small we are, we just have to be pragmatic and teach in the direction that most people want and most of our users know.

We've made those decisions in the content, and we said that, for now, for every language we're teaching, we're going to choose a standard. But I do expect that in the future, we're going to get a lot more sharply differentiated. If you want to learn British English, then we'll teach you British English and teach you how to pronounce it. I think all of that feels like something that a superhuman language tutor should be able to do.

Speaker 1

I just think it'd be very funny if all the Koreans had a very distinct Southern accent.

Andrew Hsu

Yeah. It'd be great to make that happen.

Speaker 1

I do think about this because, obviously, there's a moving of the goalposts. Now that we have this, we want the next thing. People who speak English as a second language always have an accent. I haven't—A lot of people think I don't have an accent, but if you know any Singaporeans, you know I'm Singaporean. How important is accent training? I think that actually does help a lot for people. And you cannot tokenize accents yet.

Andrew Hsu

Yes, that's right. I have 2 main thoughts on this. I think the first one is that communication and your ability to speak spontaneously and get a concept across—an idea across—is almost fully orthogonal to pronunciation. You can be really bad at pronunciation but still communicate effectively.

A lot of the current core product experience is about speaking as much as possible, making mistakes, and not worrying about screwing something up on the accent or pronunciation side. The important thing is that you literally move your mouth and make the sounds, right? It turns out there's a really key psychological barrier there where people are just not willing to do this in front of a human, even if it's a human teacher that you're paying.

A lot of the core message of our marketing campaigns in many of our biggest markets is along the lines of, “You can make mistakes in this private space with Speak.” I think psychologically that's extremely powerful. Then you can go and get it right more confidently in the real world after you practice with Speak.

Having said that, people do care about their pronunciation and their accent, right? For English only right now, we have a pronunciation coach that is basically a fine-tuned version of wav2vec, which is a Meta model. We fine-tune it on a bunch of our own phonetic transcripts and fine-tuning data.

It works pretty well. It's currently for single words. We're going to expand it to full sentences, to more languages, and so on. But I think that if you look at the pure market opportunity, our sense is that we really want to push people to speak as freely as possible and just get that volume up.

Speaker 1

In terms of immersing language learning in the real world, one of the more interesting approaches that people keep trying is to have, let's say, a Chrome extension or something on top of a page. I think Toucan was doing this.

Andrew Hsu

There's a bunch of those, yeah.

Speaker 1

And then there was another one I saw recently, which is like: watch a YouTube video and it'll transcribe for you, but randomly mask out words.

Andrew Hsu

I saw that too.

Speaker 1

Yeah. That was like a Show HN. Do those work? There's kind of the question of whether that's the right product, right?

Andrew Hsu

Yeah. Basically, the difference is your content or real-world content, right? Obviously, you want real-world content. I think that for work—for Speak for Business, for the B2B product—another part of the vision is really: what should a superhuman language tutor be able to do? It should probably be able to handle kids as well as a Samsung employee who wants to transfer to the U.S. office and wants to use Speak for work, right?

Our view there is that it's the same product. It's a different distribution mechanism, right? Consumer versus B2B. I think that we will eventually build something like a Mac app. Maybe it'll be integrated with the browser in some way. We're not really sure yet. But obviously, in order to apply it to your day-to-day, there needs to be some way to hook into your actual work documents or whatever. That's a whole can of worms.

We are actively thinking about it, but I think my sense is that it's not clear to me that any of these products have really taken off. There are many other approaches that are possible. I don't have the answer, but another example—a very hypothetical future world—is maybe OpenAI, with the new Jony Ive thing, will come out with some hardware that will be listening to you all day. Then we can give you some sort of very deep analysis that's integrated with the Speak app at the end of the day or the end of the week, whatever. I don't know.

Speaker 1

Okay, one more time, since you brought that up. I'm sure you don't—I actually haven't told you anything.

Andrew Hsu

I don't know anything.

Speaker 1

What's it going to be?

Andrew Hsu

It's the number one topic at all the parties I go to now.

Speaker 1

Really?

Andrew Hsu

Yeah.

Speaker 1

What's the most compelling idea you've heard?

Andrew Hsu

Okay, there are people who say Jony Ive hates wearables.

Speaker 1

Yeah, I've heard that, too.

Andrew Hsu

If it's not a wearable, then you've just made a second phone. In that case, just make a phone.

Speaker 1

Yeah, but I thought that he said it was going to—I mean, didn't Sam Altman say that he wanted to do a phone in the past?

Andrew Hsu

That was in the far past. He says a lot of things.

Speaker 1

Yeah, they'll say a lot of things.

Andrew Hsu

Yes. Anyway, I think a wearable makes sense. I think the race is to capture context.

Speaker 1

I have a wearable on.

Andrew Hsu

Yeah, we have a wearable here, too. It transcribes everything?

Speaker 1

Yeah, that's cool. It's from a previous episode of ours with Omi. I can hook you up if you want.

Andrew Hsu

Yeah. I think it's something that a lot of people are interested in, obviously, because it's a huge bet by them. I'm curious.

Speaker 1

You mentioned video. I just wanted to double-click on that a little bit. I'm sure engagement is very high for video because people love to watch video. I thought that Speak would be one of those places where you just leave it in your pocket, take a walk, and learn to speak. Probably that's not true.

Andrew Hsu

What we've done so far is make part of the course experience a teacher video. We've tested other, more audio-forward types as well. We found that, of course, like you said, video is very engaging, but at the same time, we have a lot of users who do want to be able to walk around with the phone locked in their pocket. So doing something that's more like voice mode with optional visuals, I think, is really good.

I think there's a huge opportunity for a better way to learn things like listening comprehension. I took German in grad school for 2 years, and I thought I was getting somewhere, but anytime I listened to a native German speaker, it was so fast. It was completely on a different level.

You can imagine a plethora of really cool experiences that feel kind of like you're listening to a podcast, but it's all AI-generated, fully controllable, and integrated with the app. There's something there for sure.

Speaker 1

Yeah. I don't want to do an AI podcast, man. We're cooked. It's okay. I mean, when that happens, we just end the show.

Andrew Hsu

To zoom out a little bit, in the pretty near future, multimodal models will cross the threshold where they'll be able to generate images a lot faster than they currently are, maybe somewhat close to real time even, right? And audio at the same time, text at the same time. You can imagine a very powerful multimodal tutor that can do it all at once, where there's an audio track, and then, if the teacher is teaching you something with the right timing, it uses that timing: “Okay, at this point, I'm about to introduce a new concept, so I'm going to show the word on screen so that the user can see how it's spelled,” right?

There's a lot there. You can do generative UI. There's a lot of nuance there, where it's easy to do it badly, but to do it well requires a fair amount of reasoning and mental modeling of what the user knows.

Speaker 1

Yeah.

Andrew Hsu

Which feeds into what you need to show at what time. So that's probably going to have to be a pretty parallel set of systems.

Speaker 1

Have you spent any time looking at Veo 3, where you do video plus audio at the same time, and how you can tweak the audio part versus the video part? I can imagine you might work on the video part and then want to change the audio-generation model. I don't actually know how the model works internally or how much you can tweak just the—

Andrew Hsu

We haven't really looked at the video stuff much. We basically think that we're very bandwidth-constrained, right? We're just scaling and trying to hire as fast as possible, like everyone else is. As a result, we're really focusing on just the most in-reach, highest-impact things.

I do think that the barriers are coming down very fast for all of this stuff. I'm just so excited about multimodality and where things are going here. Imagine if you're learning Spanish, being able to look at an image that the model generates for you and then doing Q&A on it, right? Like a beach scene, and then the model will ask you, “Oh, how many people are running on the beach?” Then you have to respond in the target language that you're learning.

It's a very traditional language-learning exercise, but you can imagine it being fully generative, which is really cool.

Speaker 1

Awesome. Lots of stuff like that. The engineer in me worries about inference costs, but I think you can just sweep that under the rug for now. See if it works first, and then you can worry about cost.

Andrew Hsu

Yes.

Speaker 1

You mentioned your real-time voice platform. I just wanted to give you the platform to talk more about that. You mentioned, for example, that you're a very heavy user of the real-time API from OpenAI and that you built a bunch of tooling around it.

Andrew Hsu

Last year, we had early access to the real-time API, and there's a very obvious use case for language learning. One common theme that has just been pretty awesome since LLMs came out is that language learning as an application is just a really good fit for LLMs and all these model types in almost every way, which has been really great for Speak specifically.

For real-time, I think the audio piece promises to really infuse almost every surface in the app. You can imagine this being the primary way that you talk to your tutor, right? An initial complication is that it needs to be multilingual, and there needs to be code-switching. So that's a pretty frontier problem right now.

Like, I should be able, if I'm learning Spanish, to speak both English and Spanish, and vice versa with the model. That's a pretty hard TTS problem today. Actually, only a few models are able to speak 2 languages in the same sentence and then pronounce them properly.

Speaker 1

Sorry. You can have a router model, like a tiny little router model, guess which language comes first, and then route it.

Andrew Hsu

Well, the problem is that you could have a subword in a single sentence in a different language. So you can't just concatenate either.

Speaker 1

It won't sound right, right? It won't sound natural. That's not how humans do it.

Andrew Hsu

So there seems to be a very native, controllable audio function. We are in the process of building a variety of experiences on top of the real-time API. I want to clarify that nothing is in production yet, mostly for price reasons, frankly.

The pricing model of the real-time API makes more sense for something like a customer-support agent, where you're very directly replacing somebody whom you would pay hourly otherwise. That's how you're seeing the pricing model for a lot of these initial agents work out.

For us, we want our users to be able to do these real-time role-plays and have these conversations for many hours a day, if they want. Getting the cost under control is definitely a pretty key consideration right now. But we are pretty close. Maybe even by the time this episode is released, we'll have something live.

But we have what I think is a really cool application of the Realtime API: a new instructional lesson where the model is actually teaching you something, like a new language concept. It’s intended to augment—or play the same role as—our current video lessons, which are the instructional-lesson type. And it’s interactive, obviously. At certain points in the 3-to-5-minute lesson, you’re interacting with the Realtime API.

It’s semi-on-rails. There was a lot of scaffolding we needed to build to properly switch between the interactive and noninteractive portions of the lesson, if that makes sense. There are some portions where you’re just listening or looking, and then some portions where you’re actively in a short conversation. We swap back and forth, and we have a bunch of custom architecture and infrastructure around that.

Then there’s also the challenge of making the cost make sense, or at least semi-make sense. There’s a bunch of WebRTC infrastructure at a scale that’s not huge, but is still nontrivial. It’ll definitely cost us millions of dollars if we do something wrong.

Speaker 1

Do you do inference in Korea because of the latency and all that?

Andrew Hsu

It’s something that we’ve been increasingly paying attention to for all the real-time paths. I would say that 2 or 3 years ago, when real-time stuff was still quite nascent, users didn’t really care as much. But I think now the standards have risen. Latency has to be low. Everyone cares.

Speaker 1

Do you have a hard latency budget for responses, or do you just kind of work it out? For example, you have a knowledge graph that you’re accessing, and you have content that you’re retrieving. There’s a lot of stuff there. Maybe you’re using a reasoning model—probably not—but that all eats into the budget.

Andrew Hsu

I will say that from the real-time engineering side, everyone talks about, “Okay, submit a user request to get an agent audio response—first bytes, right? First audio bytes. What’s that latency?” Then we try to get that as low as possible. I would argue that’s actually a vanity metric, because what you don’t take into account is how the VAD works. How do you do turn detection to detect when the user is finished speaking? That can easily add another second if you do it badly.

Nobody talks about that for some reason. What you need to measure is actually the time from when the user stops talking to when the model’s first audio comes, and usually that number is much larger.

That’s a very domain-specific problem. You can use the semantic VAD on the Realtime API for regular English conversation, and it will basically classify at every token how likely it is that you’re done speaking as a normal conversational English speaker, like in this conversation. That’s fine, but it doesn’t work at all for language learners. If I’m trying to respond in a language I’m learning, I’m going to be hesitating halfway through for 10 seconds or more. It needs to be fully custom, probably. It’s something that we’re also actively working on, but that is actually the dominating factor in perceived latency.

Speaker 1

Coding-wise, do you use Cursor, Windsurf, and other autonomous agents?

Andrew Hsu

It’s kind of all of the above. I think, as the CTO, I view it as part of my responsibility to really set expectations, push everyone on the team, and show them what’s possible. We’ve been trying everything.

I think we’ve tried to basically set the expectation that the frontier is moving so fast, it’s deeply nonintuitive. If you tried coding tools 6 months ago and they weren’t that great, especially if it wasn’t TypeScript or Python, right?

Speaker 1

The 2 most popular languages.

Andrew Hsu

That’s all it is. We try to set a culture in the engineering team where usage of these tools as much as possible, and as a default path, is the expectation. In hiring, we’re now explicitly asking about this a lot and thinking about what types of people are going to have higher agency in trying these types of tools. It’s so important.

Speaker 1

Before we zoom out, is there anything we missed about Speak that you really want to highlight, or something that people underrate about it?

Andrew Hsu

One thing that I’ve always been really excited about is that I feel like a lot of the foundational pieces that we’re building around the knowledge graph, for example, should be applicable to not just learning a language, but also other things in the future. We’re already starting to see the very beginnings of this on the B2B side, where a lot of it is more like management skills, hospitality skills, and communication skills—more like true L&D for enterprise, less like core English proficiency.

So I think, you know, that’s obviously the immediate neighborhood. But you can imagine many academic subjects—math, biology, and so on—could also work. I’m super excited about that.

Speaker 1

If I knew my employer was giving me a language tool, but then it was evaluating me on my management skills while learning a language, I might use it less.

Andrew Hsu

Fair. Fair. Fair.

Speaker 1

You want to separate that out.

Andrew Hsu

Yeah, very fair.

Speaker 1

I agree overall that the knowledge graph problem is very important. We have a whole track on it for the conference. And I think that the amount of data can be so high. Actually, you want to generate relevant triplets. I assume you use the normal subject-predicate-object type.

Andrew Hsu

It’s a bit more custom than that, because it’s more domain-specific around the way that we conceptualize the vocabulary, the sentence patterns, and so on. It’s more specifically around language-learning concepts, if you will.

Speaker 1

But what I think we can extract from Speak, or generalize as a framework, is what I’ve been calling the Bloom’s 2 Sigma Problem type of thing: the level-adjusting tutor. Where are you at? Let me adjust my thing to where you’re at, and then I’ll push you up to the next level. I think the knowledge graph is a part of it, but I don’t know if that’s all of it. I’ve never seen a working example.

Andrew Hsu

We’re approaching that problem from a few different angles. I think part of it is the knowledge graph. Part of it is being very careful in how we structure the curriculum so that you’re placed at the right level, so that the learning path itself has a foundational backbone. Beginner to intermediate English learners actually all need to know a bunch of similar concepts. It isn’t really until you get to intermediate and more advanced levels that things start to diverge more sharply.

From A0 through B1, I would say there’s a pretty well-defined, sort of linear path, actually. A lot of the deep thinking that we’ve done around how to structure the pedagogy is also super useful in terms of just matching people to the right level. Then you can take this backbone and basically modify it based on the knowledge graph and your system’s knowledge of what the user is bad at versus good at.

I think, for a lot of startups—especially in EdTech—that is the core engine. Once you do that, you can kind of teach anything.

Speaker 1

Totally. We have a few more broader, fun questions. Speak.com—great domain. I looked it up. Voice.com got bought for $30 million.

Andrew Hsu

Oh my god.

Speaker 1

When?

Andrew Hsu

2019.

Speaker 1

Okay. I don’t know if you want to share how much you paid for it, but—

Andrew Hsu

A lot less. It was a lot less than that.

Speaker 1

I figured it would be a lot less, but I’m curious. My estimate was $100,000, but—

Andrew Hsu

It was more than that.

Speaker 1

More than that. Wow. Okay, I’m not going to say anything more about the numbers. What was the sort of process? Was it easy? Did you use a broker? We had Dharmesh Shah from HubSpot, who sold Chat.com to OpenAI, and he has a lot of very fancy people—

Andrew Hsu

A multimillion-dollar deal or something, right? That was very big.

Speaker 1

No, that was AI.com. Wait—Chat.com. Chat.com, okay.

Andrew Hsu

Yeah. We bought it several years ago. It felt very expensive for us at the time. It was a little bit of a crazy move, but I think we were very convinced that we needed a super-strong consumer brand that was scalable globally. That was always our ambition. We want to be the way the next 1 billion people learn languages, and we need Speak.com.

So we don’t regret it. It’s such a great word. It makes for great swag.

Speaker 1

Very nice decision. A couple of other fun questions: any fun Korean celebrity stories? You work with so many influencers.

Andrew Hsu

We have a bunch baking right now, but I think something more generally that has just been so fun on the journey is that we would visit Seoul every year. Seeing Speak go from nothing to the first time we saw somebody on the street using Speak—and now our main teacher in the app is like a mini-celebrity. People come up to her on the street as she’s just walking around Seoul and recognize her from the app, which is really cool.

Now we do a lot of advertising. We do billboards and TV commercials, and we work with big influencers and so on. Seeing the scale of that has me kind of in awe. It’s really cool to see something that used to be nothing.

Speaker 1

I wanted you to name-drop, like BLACKPINK or something.

Andrew Hsu

Look, there’s some stuff baking right now.

Speaker 1

Yeah, okay. All right. We talked about the Thiel Fellowship. On your LinkedIn, you’ve got this hole between 2012 and 2016, which you talked about. You did some startups. Any of them that you want to share—ideas that you worked on that you thought were cool?

Andrew Hsu

Early, but, you know—

Speaker 1

What projects would you revisit?

Andrew Hsu

I’ve always been interested in learning and education.

One of the other Thiel startups that I did during that time—it feels silly to even talk about this because it amounted to nothing—was called Bloom. You know, Bloom’s 2 Sigma Problem. It was actually named after that. We were trying to basically build a better adult-learning platform and have really cool interactive JavaScript widgets for various concepts that you can learn.

We didn’t find product-market fit. I was young and didn’t really know anything about business at the time, either. But I think the common thread through everything that I’ve been interested in since leaving grad school has been: How do we build software and tools that help people learn things more effectively, better, and faster?

And now I feel very lucky to be in this position because, obviously, AI is the ultimate version of that, right? It’s been completely transformative for me personally because I get a lot of inherent fun and pleasure out of being able to think of a concept and then talk to this omniscient LLM that can tell me more about it. I’m really good at asking the right follow-up questions about what I want to know, so that’s been completely transformative for me.

Speaker 1

Do you get a lot of people using Speak for therapy? It’s not meant to be that, but since you have an interface, they will use it.

Andrew Hsu

In 2023, when we first launched our AI role-plays using GPT-4, people were way more concerned about safety, right? Obviously, the models now are much better at refusals, and the line is sharper between what’s appropriate and what’s not. But we did see a lot of our first users start to put in pretty questionable custom scenarios. You can probably guess.

This was something we expected, but I think seeing the logs in person is very different.

Speaker 1

Got it.

Speaker 2

Some shocking stuff in there. Last couple of questions. One on Andrej Karpathy: You talked to him in your machine-learning journey.

Andrew Hsu

That was a long time ago, yeah.

Speaker 2

He’s also working on EdTech now.

Andrew Hsu

Yeah.

Speaker 2

I don’t know if you’ve ever had conversations with him on Speak.

Andrew Hsu

He’s also interested in language learning, by the way. One thing that I think we didn’t really realize early on, or fully internalize at least, was just how deep the market is.

Speaker 2

Say more.

Andrew Hsu

It was so universal that we really struggled to do some of the basic startup stuff around defining your ICP—

Speaker 1

Ideal customer profile.

Andrew Hsu

—and segmenting your users, because our users were everyone. We had parents using it with their kids. We had really old people using it. We had people using it for work. That was kind of mind-boggling.

Speaker 2

You still did customer segmentation, or are you saying it doesn’t matter?

Andrew Hsu

I’m saying it was hard to do. We tried, and we have a sweet spot. In Korea, it’s 25 to 45—more professional, more white-collar—but it’s a very long tail on either side.

Speaker 2

Yeah, I think it’s a huge market, and I think it’s a very special moment in time right now where it’s obvious that a lot of the tech is here. I think it’s really good for humanity if we make a lot of progress here. So I’m really excited for this company, too.

We started by asking about the Thiel Fellowship, so maybe we can wrap with one of Thiel’s favorite questions: What’s something you believe in today that most people would not agree with you on?

Andrew Hsu

I think that people, if you recall, expected the world to kind of explode when GPT-4 came out and everything would change. If you go to another state outside of the Bay Area—probably even in California outside of the Bay Area—and ask somebody how much their life has materially changed, it’s pretty close to 0. Real-world inertia is enormous.

Obviously, AI is probably the most transformative technology we’ve ever built, but I think in a very real sense, the world hasn’t changed that much, either. That’s a really weird thing, right? So I think we need more builders. We need more people building applications.

It’s weird to me that there are actually not that many net-new, consumer, AI-native applications at scale. There should be way more. I would love for there to be way more.

Speaker 2

Consumer is hard.

Andrew Hsu

Yeah, I’m intimidated, but there was just never any alternative for us.

Speaker 1

Yeah, like I said before.

Speaker 2

You didn’t have a choice. But also, you’re very smart. Maybe you have some growth-hack things that you can advise people on that they could learn from.

Andrew Hsu

But yeah, I agree. I think the general take actually is this is what we want, which is a slow takeoff, short timeline.

Speaker 2

That’s fair, right? This is the 2-by-2 that everyone always talks about in AI safety. You see slow takeoff and maybe don’t complain. Maybe we have a heads-up, because Dario’s right and half of us lose our jobs in the next 2 years.

Andrew Hsu

Yeah. It’s so hard to predict. Sometimes I get AI anxiety.

Speaker 1

Uh-huh.

Speaker 2

You get anxiety?

Andrew Hsu

Yeah.

Speaker 1

Okay.

Andrew Hsu

And then I just focus on our users.

Speaker 2

That’s a perfect place to wrap. Thank you so much for taking the time.

Andrew Hsu

Yeah. Thank you both so much. This is great.

个性化 AI 语言教育——Andrew Hsu 与 Speak — 文字稿与摘要 | BidClub