[BidClub_]
Machine Learning Street Talk · · 55 分钟

John Palazza——CentML 全球销售副总裁(赞助)

John Palazza

播客
TL;DR
  • AI基础设施交易正从购买最大GPU转向从已部署算力中榨取更多工作量。 Palazza表示,客户通常只使用GPU 30%-40%的能力,而理论上可达到80%-100%;Tim援引一项调查称,到2026年,所需GPU数量可能增加“约”33%。“始终会需要更多算力”;效率提升会扩大有效供给,但不会终结需求增长。
  • 企业AI真正的成本危机始于试点进入生产环境。 创新预算可以承受一次8张H100 GPU的实验,但面向45,000名用户上线后,企业突然开始担心这项能力“会让我们破产”。针对CentML宣称最高可节省60%的说法,Palazza称算力是初创公司的“最大单项成本”;随后一位未标注身份的发言者表示,算力削减60%可能让资金续航延长6个月。
  • 更高效率很可能带来更多AI消费,而非降低总需求。 Tim的框架是,呼叫中心、软件开发或销售环节节省下来的成本,将为更多功能和自动化买单;Palazza认同,释放出来的资金、开发者时间和创造力会打开下一轮增长周期。“效率提升才能放大兴奋感,也才能放大创新。”
  • 企业采用AI,既是模型问题,也是组织设计问题。 在大约5,000次AI/ML客户合作之后,Palazza认为最典型的反模式是一家金融机构的8位CIO各自开发情绪分析系统——没有任何协作,最终变成“一系列不同的雪花”。他更认可的方案是由最高层负责推动,再由贴近实际痛点的技术负责人提供支持。
  • CentML将自己定位为不绑定硬件和云厂商的AI工作负载平台。 CentML Serve、CentML Train和CentML Cluster覆盖推理服务、训练、集群编排、编译器、网络,以及跨云和本地基础设施的无服务器Llama端点。其运营承诺是让“合适的工作负载……在合适的时间”运行于A100、H100、L4、A10及其他GPU之上,无需针对每种配置重写模型。
  • 开放权重模型似乎正接近企业所需的“够用”门槛。 Tim认为,Claude 3.5 Sonnet处理歧义的能力已经足以支持有用的零样本任务;更小的Llama模型则可通过更多提示、示例和目标限定变得可靠。Palazza表示,如今他的客户大多在评估开放权重模型。“眼前最好的机会……将来自那些开放权重模型。”
  • Agent被视为从惊艳聊天机器人走向可量化工作流自动化的桥梁。 Palazza预计,Agent将把生成式AI从“第一步或第二步”推进到第三至第五步,覆盖故障排查、工单解决、数据中心操作和医疗信息工作流。对投资者而言,基础设施抽象与Agent相互强化:更低的执行成本让更深层的自动化具备经济性。
  • CentML将超大规模云厂商和NVIDIA视为互补方,而非必然侵蚀利润的对手。 Palazza认为,更高效的工作负载会让云客户更满意,也让他们能够推出更多用例;NVIDIA则会受益于其技术栈获得更高采用率和利用率。其护城河主张建立在可移植性之上——在模型、AWS、GCP和本地系统之间保持“自由移动”——以及这家初创公司跟上架构快速变化的能力。
摘要 · 为研究而整理的核心内容

1. 算力稀缺使利用率成为第一种扩容方式

  • CentML于2022年从多伦多大学孵化而出。当时科研经费有限,研究人员无法简单地选择“最大、最好的硬件”。Palazza的起点是一个由现实逼出来的原则:使用最有效的算力,再把这套纪律扩展成一个平台,让“合适的算力在合适的时间服务合适的工作负载”。

  • Tim援引一项调查称,到2026年,所需GPU数量将增加约33%。Palazza不认为需求会下降:云计算普及后,市场原本以为存储会被淘汰,但一位EMC高管告诉他,云和存储支出都在上升——“事情总会继续向前。这就是规律,对吧?”(Things just keep going. It’s the law, right?)

  • 眼下最直接的机会在于盘活闲置容量。Palazza表示,企业通常只消耗一块可用GPU 30%-40%的能力;在继续购买稀缺硬件之前,先把利用率提升到80%-100%,是解决算力短缺的“第一步”。

  • 他对市场时机的判断,是从追求最大规模向务实主义发生的“橡皮筋回弹”。当最大的基础设施无法使用或不具备经济性时,规模本身就会变成负担,客户不得不追问:如何把理论上的机器学习变成一个在财务和技术层面都经过优化的实用系统。

2. 企业AI要先找到负责人,平台化才有回报

  • 根据过去10年中接近5,000次AI和ML客户合作的经验,Palazza回忆起一家大型金融机构,内部有8位CIO。每个团队都在独立开发情绪分析系统,彼此没有协作,留下一个个无人负责的模型;在本应共享的企业目标下,最终却成了“一系列不同的雪花”。

  • Tim提出了一个值得保留的反驳:类似Amazon“两张披萨”团队的组织模式,可以通过自治换取速度;而LLM让情绪分析等任务变得如此简单,重复使用可能并没有多少技术优势。但同样的自治也会造成基础设施和支出的重复,再次把组织拉回去中心化与集中化之间的摆动。

  • Palazza的解法是自上而下的文化负责、自下而上的落地执行。高管愿景让AI成为标准工作方式,并使其“渗透到业务的每个层面”;与此同时,解决眼前集成问题的资深ML工程师,也可能成为企业内部最有力的推动者。

3. 只有生产规模才能显现单位经济性

  • Palazza将创新预算与运营经济性分开。一位CTO可能愿意投入一笔固定金额,去探索生成式AI能做什么;但从“创新到生产再到规模化”的转变会消除这种预算弹性:一次8张H100 GPU的测试,在目标用户扩大到45,000人后,立刻变得令人警惕。

  • 假设中的客户反应非常直接:试点成功了,但规模化部署可能“会让我们破产”。正是在这一刻,而不是最初的实验阶段,具有前瞻性的CTO才会开始把推理效率、模型部署位置和利用率视为业务约束,而不只是工程优化项。

  • Tim援引CentML网站上最高可节省60%的说法。Palazza提醒,“把统计数字放在网站上是危险的”,随后又表示,网站公布的数据相较客户实际结果仍偏低。Palazza称算力是初创公司的“最大单项成本”;随后一位未标注身份的发言者表示,算力减少60%可能让资金续航再延长6个月。

  • 效率提升可能带来更多总算力消耗,而不是降低需求。Tim回忆Bain的估算——软件开发节省15%或30%、销售节省30%、呼叫中心节省25%,这些数字是他“凭印象说的”——并认为释放出来的预算会为更多功能买单。Palazza认同:节省下来的成本会解锁下一阶段的采用,而不是终结支出周期。

4. CentML把硬件选择变成工作负载层面的决策

  • 产品栈从CentML Serve开始,用于配置、部署服务和估算模型在特定硬件上的表现,包括成本与效率。CentML Train负责训练;CentML Cluster增加集群、编译器和网络优化;底层平台还通过API提供针对Llama模型优化的无服务器端点。

  • CentML的抽象层可运行在主要云平台和本地基础设施之上,原本通常分配给A100或H100的工作负载,也可以转由L4、A10或其他类型的GPU承载。其Snowflake产品通过Snowpark Container Services运行,将这一模式延伸到传统云部署之外。

  • Tim描述的终点,是一个可自我修复、自我扩展、可暂停的算力集群,用户提交任务即可,无需指定NVIDIA H100。Tim指出,GPU还没有实现完全虚拟化;Palazza则表示,CentML通过按照延迟、性能、容量和成本要求调度工作负载,已经实现了实际上的核心收益。

  • 他的比喻是开着法拉利去买菜:如果不考虑成本,所有人都会选择最快的机器,但这通常并不高效。如果一个工作负载有“16种不同的实现方式”,平台就应该在其中无缝选择,而不是让开发者针对每种GPU型号把模型重写“150次”。

5. 产品化替代超大规模工程团队

  • Palazza的职业经历横跨服兵役、4或5家初创公司的软件销售、被Apple收购的GraphLab,以及CentML之前在Algorithmia和Converge从事的早期MLOps工作。

  • 一次向Salesforce推销产品的经历让他认识到专业化的边界:他们带着25名专职工程师前去提案,却听说Salesforce有380人在开发类似系统。顶尖企业会继续自建内部平台,但包括大型零售商在内的大多数公司,都无法组建如此规模或深度的团队。

  • 可投资的缺口在于产品化:大型工程团队内部开发出的能力,可以被转化为拥有约12名工程师的公司也能使用的产品。Uber的Michelangelo展示了MLOps平台能够做到什么;Algorithmia和Converge等产品,则服务于需要同等结果、却没有Uber基础设施的组织。

  • 界面设计是另一层抽象。Tim提到,产品正从聊天界面转向画布界面,并描述了自己用Cursor在约30秒内构建Interview Notes应用的过程:应用带有计时器和“Good Bit、Bad Bit、Reference”字段,还能生成供编辑使用的SRT字幕文件,把一闪而过的想法变成可运行的工具。

  • Palazza表示,企业仍需要在RAG、微调、嵌入式界面、独立前端和Agent之间做选择。决策变得更容易,采用速度就会加快;最终,Agent将负责采取行动,而不只是返回文本。

6. Agent让生成式AI突破聊天机器人上限

  • Palazza认为,软件故障排查是Agent化最自然的第一步:先诊断故障,再执行修复动作。同样的模式也可以用于处理故障工单和数据中心工作,让对话模型成为运营流程中的实际参与者。

  • 他更谨慎地将这一机制延伸到医疗信息、医生和处方领域——在这些场景中,Agent可以沿着既定流程执行一系列逻辑步骤。本期节目并未声称这些应用已经完成,而是将其作为CentML探索或参与过的方向。

  • 他强调的核心差别在于成熟度:许多企业部署仍停留在“第一步或第二步”,而Agent可能将其加速推进到“第三步、第四步、第五步”。他承认聊天机器人有价值,但“可实现的业务影响远不止聊天机器人”。

7. 开放权重模型正接近企业所需的“够用”

  • Palazza回顾了市场演进:企业最初期待训练自有LLM,随后迎来ChatGPT带来的“顿悟时刻”,如今越来越关注Llama。随着Llama在准确性和能力上逐渐接近闭源模型,他表示,企业讨论正主要转向开放权重模型,把它作为内部开发者可以继续扩展的基础。

  • 他描述了一条渐进路径:先把Llama模型作为端点使用,再将推理迁入AWS、GCP或本地GPU,最终,一些公司会训练自己的大型语言模型。

  • Tim补充说,闭源模型的表现仍然重要:Claude 3.5 Sonnet处理歧义的能力已经跨过他的实用门槛,足以支持有用的零样本工作。更小的Llama模型可能需要更多提示工程、示例和更窄的目标范围,但更强的控制力可以支持路由、Agent、分层优化,并让更多运营环节留在组织内部。

  • Palazza坦诚地描述销售现实:“初创公司有时最糟糕的敌人是什么都不做,第二糟糕的敌人就是够用。”一家公司不应只是为了让法拉利每小时快1英里而存在;但尽管提出这一警告,他最终仍认为,开放权重模型如今拥有最强的机会。

8. 可移植性与架构迭代速度决定护城河

  • NVIDIA和Deloitte既是投资者,也是用户。Deloitte提供客户项目、行业优先级和创新实验室方面的可见性;NVIDIA则提供技术反馈,并会因优化提升其硬件技术栈的采用率和利用率而受益。

  • AWS、GCP或Azure是否会复刻这项能力,并没有让Palazza夜不能寐。CentML已通过AWS和GCP市场提供服务;他认为,超大规模云厂商更希望拥有高效且满意的客户,让客户把节省下来的成本投入新工作负载,而不是让客户把自己的平台与失控的支出联系在一起。

  • 云和模型之间的可移植性仍不完善。Tim指出,跨云迁移往往需要修改代码,而输入token、输出token的模型也并不能真正做到热切换。CentML给出的答案是“自由移动”(freedom of movement):切换Transformer模型、基础设施,或在AWS、GCP和本地系统之间迁移,都不应令人恐惧。

  • 架构能否经久耐用,需要Palazza所说的“初创公司的灵魂”。Tim认为,状态空间模型、minLSTM,甚至RNN,都可能最快在次年挑战效率低下的Transformer;CentML的应对方式,是支持不断变化的区域模型偏好,并在市场演进过程中维持自身的“炼金术”。

Speaker 0

There’s always a pendulum swinging between decentralization and centralization. But then there’s the question of who owns AI in an organization. If we’re going to leverage AI more efficiently, we need to have a joined-up approach, surely.

John Palazza

An executive with a vision that has ownership from the top level creates a culture that allows the adoption of machine learning to permeate through each facet of their business. I think it’s been proven that the companies doing that are more successful, both in adoption and in their future vision of where things are going.

Even though there’s a lot of excitement, I think there’s a lot of individual excitement, not necessarily that next phase of collaboration. But this transition from innovation to production to scale is when the looseness that happens in innovation needs to start changing.

All of a sudden, they’re saying, “You know what? It’s great that we’re running this solution on an 8-1 H-100 that we’ve been testing, but now we’re going to roll it out to 45,000 users, and we’re concerned that this capability is going to bankrupt us.”

As Llama has progressed, so many of the conversations I have with customers are predominantly about looking at the more open weights that are being made available to them.

Speaker 0

John, welcome to MLST. It’s great to have you here.

John Palazza

It’s a pleasure to be here, Tim. Thank you so much.

Speaker 0

Tell me about your background.

John Palazza

Wow. So, I started a small farm. No, I’m just kidding.

I originally started out as an officer in the military, so I went to the Naval Academy in the US. I was an officer in the Navy and started exploring the early fundamentals of machine learning and aspects of technology right then. I ended up selling software and went through a startup cycle for a long period of time. I was the first salesperson at about 4 or 5 startups.

I really enjoyed building product-market fit and understanding where and how new technology was being used when I transitioned from the military to the civilian sector. For the last 10 years or so, I’ve been working through the challenges around machine learning and artificial intelligence, first with a startup that I advised called GraphLab, which was acquired by Apple, and then early on in the MLOps space with companies called Algorithmia and Converge.

We were dealing with the challenges and adoption of getting started with machine learning, and that’s what drove me to CentML. We’re going through a very similar narrative now when it comes to generative AI and large language models: understanding how people use them and why they do, and the excitement that was being built here at CentML was very attractive to me. That’s where we are today.

Speaker 0

Amazing. CentML was founded at the University of Toronto in 2022, if I understand correctly.

John Palazza

That is correct.

Speaker 0

What is it? What are your products? What do you guys do?

John Palazza

We’re really the first platform to look at optimization across multiple levels. Optimization is such a clichéd term, so I apologize for using it, but let’s look at it a better way.

The hardest thing to do sometimes is to actually get started with machine learning—or, in this case, generative AI and large language models. CentML realizes that some of those burdens or challenges are with the infrastructure itself. How is it possible to use so much compute and such heavy and large amounts of that environment?

We said, “What if there was a better way? What if there was a way that we could, first, meet the customer where they are in their journey”—and we can go through what that means—“and also, once they are at that stage or thinking about it, make sure that they’re using the right compute at the right time for the right workloads to help people optimize and adopt large language models?”

At the core, that’s really where CentML is and what it is. I always like to say that academics sometimes have the biggest challenges to face. For the people at CentML and the University of Toronto, one of their biggest challenges was, “How do we actually use our grant money efficiently to test all of this machine learning?”

In that case, we couldn’t just use the biggest and best hardware. We had to use the most effective hardware. They always say necessity is the mother of invention. In this case, the necessity of having the right compute to generate these large language models and work with large workloads for artificial intelligence is what led to the foundation of CentML.

Speaker 0

So do you think we’re starting to move toward more pragmatism with compute? There was this modus operandi of, “Let’s just build bigger and bigger models,” and there was no concern whatsoever about efficiency and the environment, and so on.

It’s even more of a problem when we deploy into the enterprise, because of course there are heterogeneous compute environments. There are lots of different applications being used by different people on different cost centers, so it’s really becoming an issue now.

John Palazza

Yeah, I completely agree with you on that. I think there’s always this desire—and I would say almost fear of missing out—where people want the largest infrastructure they can get. “Let’s get bigger, let’s get faster, let’s get larger.”

But then all of a sudden, it becomes a question of how usable this is. I always think it’s wonderful to have the biggest, baddest, largest thing you can have, but when it’s not usable and actually starts becoming a burden and a hindrance, I feel like that’s the transitional moment where we are now.

People are saying, “How do we go from a theoretical build of machine learning to a practical build that’s both financially and technically optimized?” That’s what the conversations we’re having on a daily basis at CentML focus on.

It’s this elastic-band snapback, where people are saying, “Sure, there are opportunities for it, but I see that there’s a shift coming. How do we take advantage of that shift? How do we start using things in the right way, in the most efficient way?”

Speaker 0

Demand for compute is increasing significantly. I read a survey the other day that said by 2026, we’re going to need something like 33% more GPUs. There’s a GPU shortage, and generally speaking, the demand for compute is increasing.

The applications are increasing. Soon, people will be generating real-time virtual games and virtual worlds on their VR headsets, and so on. What’s your perspective on those kinds of industry trends?

John Palazza

First, it’s an exciting time. I would love to see some of those things come to fruition. I benefit from the use of AI on a daily basis already, so the more I get to see of it, the more excited I get.

I would say, though, that it’s interesting. This shortage of compute isn’t changing. I always think back to some of the stories going back to the storage days. I was acquired many moons ago by EMC, which is a pretty large storage company, and the big concern at EMC at that time was that they were going to go away. Storage was going to go away. Everything was going to be 100% in the cloud.

I remember 2 or 3 years after I left, I met with a VP over there and asked casually how things had progressed. He said, “It’s been our best year in history.” I said, “But that’s impossible. All this cloud spend has gone up.” He said, “Yeah, you know what? So has the storage spend. Things just keep going. It’s the law, right?”

I think that’s true with compute. I do think that demand isn’t going to decrease, but I think one of the things we need is better efficiency. So much of the compute available today is not 100% utilized or utilized correctly. We’re seeing people consistently use 30% or 40% of an available GPU and then move on.

I believe there’s so much value in that space. Being able to get 80% or 100% utilization and efficiency from a GPU is, to me, step 1 in solving some of this compute crisis that we’re dealing with. That’s also foundational to where we feel CentML fits in the space.

There will always be a need for more compute. Our goal is to help people use the compute they have as efficiently as possible as that demand constantly ramps up.

Speaker 0

There are a few things here. First of all, GPUs are more expensive than CPUs, so even now on the cloud, CPU utilization isn’t that great.

John Palazza

Yeah.

Speaker 0

It’s better if you use PaaS services, of course, but there are still plenty of people using IaaS services. The other thing is that GPUs aren’t yet virtualized. I think there are some startups doing this, and maybe the cloud vendors will start to do something like it, but at the moment it’s a little bit like the 1990s, where we’re talking directly to the metal.

The other thing is that we’re moving toward this platformification of data science. Five years ago, individual contributors were building models, tinkering around, and putting things into production. Now we need economies of scale, simply because we can’t afford to have all these random GPU machines burning money everywhere.

How do you see that?

John Palazza

It's such an interesting question because I feel that, meeting with so many customers all the time, I've probably done close to 5,000 customer engagements in AI and ML over 10 years, which is pretty hefty when you think about the conversations we're having out there.

Even though everyone wants to move closer and closer toward platforms, the ability for individuals to still make direct decisions is very practical and very much happening, whether we want it to or not. When you give too many keys to the kingdom to too many people, you have situations where not always the right compute is used for the right reasons, and not necessarily the right models. There are orphan models.

I always like to say I had a meeting, probably about 5 years ago, with a very, very large financial services company, and we were discussing the state of machine learning. In that meeting, there were about 8 different CIOs in the same organization, but they all had a CIO title, and we were discussing use cases for machine learning.

One of the really basic examples I'd used was sentiment analysis. It's such an easy use case, especially for understanding customer feedback. All of them were working on their own projects for sentiment analysis. There was no collaboration among teams, no centralization. There was a series of different snowflakes.

Even though there's a lot of excitement right now, I think there's a lot of individual excitement, not necessarily that next phase of collaboration and unification at large enterprise customers. I think we're still a little bit away from that, and I think that's part of what we're seeing. Even though the platforms are where we're going, adoption needs to accelerate.

We need to see people start looking at it from more of a unification or best-in-breed perspective, actually leveraging those workloads instead of individuals working in a silo. I think those things should change, and I think we're moving in that direction. It hasn't happened yet, but we are moving in that direction.

Speaker 0

Yeah, it's so interesting contrasting the 2 approaches because there's something to be said for the Amazon culture, where they have these 2-pizza teams. They have autonomy, and it's great because you can have a flywheel and people doing things. But then you get a lot of people duplicating the same effort.

Also, with language models, the cool thing is that sentiment is so easy. You just ask the language model, “What's the sentiment of this thing?” You can describe it using natural language, and there's not really much of a technical lift, so you don't need to reuse something that someone has done before.

But then there's the question of who owns AI in an organization. Before, for a while, we had these chief data officers and chief AI officers and so on, and there's always a pendulum swinging between decentralization and centralization.

John Palazza

That's right.

Speaker 0

But I genuinely believe that we need to have a top-down approach in these large enterprises. If we're going to leverage AI more efficiently, we need to have a joined-up approach, surely.

John Palazza

I agree with you completely. I think companies that have adopted machine learning most effectively are companies that start from the very top, as a cultural decision, as a unification or company decision, and then have an implementation that trickles down from that top-down approach.

You always like to think you start from the top, but on a lot of levels, you start with the people who are engaged in the moment, working with them and helping build credibility for the technology. But within AI and the wider space, I really do feel that an executive with a vision and ownership from the top level creates a culture that allows the adoption of machine learning to permeate through each facet of the business.

That's really where I think it needs to go, and I think it's been proven that the companies doing that are more successful, both in adoption and in their future vision. To your point, it's a lot easier to make a decision on some of the more complex models when it comes from a vision of the company saying that we're adopting, enabling, or making this part of our standard and practice. I think that's what really pushes the boundaries of where and how we can be successful with this.

Speaker 0

What kind of conversations do you have with your customers? For example, when you're talking with prospective customers, who are you talking to? Are you talking at all levels of the organization, or do you tend to focus on a particular role?

John Palazza

No, that's a great question. We're speaking to all areas within the organization. In an ideal setup, we're speaking to the senior executives and senior leadership within an organization, explaining and understanding what their strategy is toward generative AI and large language models.

Are they building their own? Are they leveraging open source? Are they leveraging closed models? Are they incorporating a structure where they're consuming, or are they building? Every customer's journey is a little bit different.

I always like to say that, at every startup, the perfect world is that there is a director or a VP who's in charge of your product. The VP of sentiment at any company would be a very easy person to call upon for the conversation. Unfortunately, that hasn't quite existed yet, so we do tend to have our conversations with people who are in the generative AI space or with executive leadership.

Some of our best sponsors have been people who have been dealing with a problem. That problem may be integrating a large language model into an application and figuring out how to use it at scale. Those people may just be senior ML engineers. Having a conversation with them and showing them what we can build and how it can be a little bit easier for them is fantastic, and it also resonates because it's a technical solution. So it's definitely across the organization, but there are different value and conversation points for each group.

Speaker 0

How do you see this working out? To give you an example, Oscar Wilde said that fashion was a form of ugliness which was so intolerable that we have to change it every 6 months, and data platform architecture is very much like that.

We had the monolithic data lake, then we had the slow path and the fast path, and then we've got the data mesh, data lakes, the lakehouse, and all this kind of stuff. These are very expensive projects, and they tend to slowly get buy-in from various different parts of the business. You've got the finance people over here, and you've got the retail people over there.

In a way, it's good to have a joined-up approach, but it also means that there are so many stakeholders you need to get on board. A lot of people say, “You can't use my data for this thing. You need to pass me a token. You can only use my data when someone is actually authenticated,” or whatever. You just get this explosion of complexity. How do you bite that off?

John Palazza

Sometimes what's hot or what's now is even less than 6 months in technology or in AI. A solution that could be of the moment when it's funded is no longer of the moment.

I joke that one of my first startups was a company called Backweb, and we had developed something called polite push technology. I'm dating myself here, but the big competitor at that moment was a product called PointCast. At one point, it was the largest user on the internet. It was bigger than Yahoo or AOL, whatever it was at that moment.

It was great until corporate networks realized it was unbelievably invasive, and then it became banned. All this technology was there, so what was hot one moment was not hot the next. We went public, and that technology ended up not necessarily landing as well 2 years later.

Now, with AI and machine learning, I think one of the things we're encountering is that there's a need that's been shifting over time. The universal concept of adoption, integration, and leveraging into production is still there. Some companies are achieving it, some companies are looking to achieve it, and some are just getting started on their journey, so there are still a lot of companies with different evolutionary traits.

I also think the approach is what's changing. If you looked 7 or 8 years ago, a lot of companies were working independently and trying to work on it. You found companies that would start their projects living on the cloud, then move from the cloud to on-premises, then to a hybrid cloud and to hybrid on-premises. There were all these different levels of adoption and different approaches, all of which had their own merit and some of which had their own challenges.

I think today what we're seeing is that there's a universal need, or understanding, that large language models and generative AI have a tremendous impact both for business and socially, and it's going to happen. The question now, and what we're dealing with, is how our customers and companies in the enterprise space see the best path forward.

A lot of companies start by saying, “We just want to put our toe in the water and start with generative AI, maybe serving it up as an endpoint. Is there a way we could try a Llama model as an endpoint?”

However, that journey doesn’t begin and end there. When a company starts there, we’ve been having conversations with other companies that say, “That’s wonderful, but what we really want is to bring some of this in-house. We want to run this on our AWS environment or GCP, and we have some on-premises GPUs that we want to leverage. Our biggest struggle is around the inference capabilities, where inference has been growing. Can you help us solve that?”

The answer is, “Yes, we can.” We focus on optimizing inference and building that process. Then we get to the next phase, and we’ve met with some companies that want to build and train their own large language models. They want to completely move away from leveraging the existing architectures and build something new. I think, to your point, where people are on that journey is what’s shifting every 6 months.

Speaker 0

Yeah. The opportunity is certainly huge. I was reading the Bain report a couple of days ago, and they were saying that there are huge efficiencies that AI could bring to businesses. In software development, I think they estimated it to be 15% or 30% if you do it correctly. Across sales, for example, it could be about 30%; call centers, about 25%, just off the top of my head.

The opportunity is huge, but there is reticence. I think it’s a little bit like the cloud was about 10 years ago: no one wanted to be first. You’ve seen so many customers that you must see patterns. There’s a certain type of culture, a certain type of organization that is adopting AI better.

Google, for example, has done lots of centralized optimization. They’ve got a centralized monorepo and a centralized build system. They’ve got all these economies of scale because they really managed to work together and figure the thing out, and I’m not yet seeing that in a lot of other businesses.

Does that imply that we need amazing software engineers, or do you think there’ll be some increasing level of abstraction that will allow us to build AI applications more easily?

John Palazza

Yeah, I think that’s a great question, and I also think it’s a trend that can’t be dismissed. Companies that have great engineers and great software engineers are pushing the boundaries of what’s available, and they’re setting an expectation of what’s capable. I think that’s amazing.

But for the majority of companies, it’s very difficult to build an engineering team that large with that much skill. I remember going on a sales call with a startup that was very successful. It had fantastic technology. We met with Salesforce, and when we had that conversation, Salesforce was like, “Wow, this is an unbelievable product. We’re building something similar in-house.”

As any good salesperson would say, I responded, “Absolutely, but we have a team of dedicated engineers who are building just this. We’re specialized in it, and we have a team of 25 engineers.” They said, “That’s awesome. We have 380 on our team at Salesforce.” I thought, “Okay, so my 25 specialized engineers are a little less than what you have in your capability.”

But I think what you’ll see is that those unique snowflakes who are working at the engineering level in the enterprise, who can develop and work in these larger teams and large organizations, will always exist at those organizations. The work that comes out of those teams can influence what a smaller company with a team of 12 engineers can benefit from, because some of those functionalities can become productized.

If you look at the MLOps space at Uber, they created a solution called Michelangelo. It was very obvious that that was their MLOps solution. It didn’t stop customers from benefiting from using Algorithmia or Converge, which are 2 MLOps platforms that I had led sales for.

The reason was simple: people needed it, but they didn’t have the infrastructure to build their own. Rather, they wanted to benefit from solving those same problems. We’re dealing with that now in the optimization space. If you go to a larger company or enterprise where they have 300 engineers doing that every day, tuning it and working on it, absolutely. But most companies don’t have that.

Utilizing an open-source product, which still takes a huge amount of work, is not always the right answer for everybody. Solutions such as CentML and other products in different spaces help fill in those gaps when you don’t have the large, centralized engineering team that Google has at a local big-box retailer.

Speaker 0

Yeah, I remember Michelangelo from Uber. That was wonderful. I know some of the people who used to work there, actually. It was really inspirational for a lot of folks in large enterprises to start thinking about platformification.

John Palazza

Agreed.

Speaker 0

Building data products and creating this kind of information architecture. You were saying that there’s all this complexity to deal with, but the smart people can create a schema. They can create ways of doing things with data connectors—

John Palazza

Mm-hmm.

Speaker 0

—and standard forms of processing, and whatnot. Of course, these can be templatized and reused downstream.

I’m also seeing a lot of innovation around the user experience. GenAI is so new that it took a long time to move from the chat interface, and now we seem to be progressing to the Canvas interface. ChatGPT, of course, has this Canvas interface, and Salesforce has just released a similar thing today.

I think over time we’re going to see this evolution of interface design with respect to generative AI, and that might be the secret to creating that low-code/no-code accessibility in the enterprise.

John Palazza

I also think taking it the next step is what will help drive adoption. So many companies want to benefit from it and use it, but they’re not 100% sure what the right path is to use it.

There are some basic questions: whether you’re using RAG or fine-tuning, whether you integrate it into something or create a different front end. Do you take it to an agent level and deliver it there? I think that’s really where everything will end up going very shortly.

Those types of innovations and changes drive adoption and allow companies to become more pervasive in their use of it. To your point, I agree. The multiple options to engage, multiple formats, and that evolution keep pushing adoption and making it a little bit faster, a little bit easier, and hopefully a little bit more pervasive, because I’d love to see every company being able to benefit from it.

Speaker 0

So you used the magic agent word, which is one of my favorite words. It’s so exciting when we think about these levels of abstraction. Of course, you folks have done incredible work optimizing the compilers, the kernel, all of the hardware, and so on.

Then we’re just going to build on top of that foundation with things like agents. I think that might itself become a new user-experience concept, because right now in ChatGPT or whatever, you do a search, and what you actually want to do is a kind of hierarchical search.

You want to say, “I’ve got these 5 PDFs, and I want you to go away and do some research and summarize each of them. All of the respective results should be injected into my thing here.” You’re doing this compositional processing, because this is how we think.

Currently, weirdly, it’s not really possible to do that kind of thing, and I think there are a few reasons for that. But what do you see the future being with agents? How are we actually going to start leveraging that?

John Palazza

We at CentML are actually working in multiple areas around extending our solution to engage at the agent level. I think the first logical step would be around troubleshooting and engaging within a software model itself. It’s such an easy way to diagnose and deliver.

What would be the next step of taking an action and enabling those actions to be run? The impact that you could have on troubleshooting, trouble tickets, and working in a data center would be amazing.

You can take those same use cases where we’ve worked on similar areas and take that step toward the healthcare industry—working with doctor information, information around prescriptions, and taking the logical steps through those processes or areas that we’ve both engaged on and worked on.

So much of what we’re seeing with generative AI in the business space today is step 1 or step 2, where agents can take it to full usability—step 3, step 4, step 5—with an accelerated rate and actually start having a higher business impact.

Many of the use cases we deal with today, where customers see the most impactful results, are around chatbots specifically. I just feel like so much more business impact is available than chatbots.

Not that they’re not valuable, because they are, but it’s that next phase, and I think agents are really what’s going to be the focal point to deliver that value to the enterprise.

By the way, as an aside, I got Cursor, the generative AI coding program, to generate me an app that I now use for interviews. It’s called Interview Notes. It’s got a timer, and I’ve got Good Bit, Bad Bit, Reference, and a text box. So when I type something in—

It will generate an SRT captions file, and then my editor can overlay that on the recording, and they’ll just know where everything is.

John Palazza

Nice.

Speaker 0

It’s so cool with generative AI because you get an idea that comes into your mind and you’re like, “Oh.”

John Palazza

Yep.

Speaker 0

And before, you just wouldn’t have been bothered to do it, but now there’s so much innovation because you’re like, “Oh, I could try this or I could try that,” and it just takes 30 seconds, so you just do it.

John Palazza

It’s awesome.

Speaker 0

It really is. Okay, cool. So, CentML, you guys have been cooking some very interesting stuff. Can you tell me about your core product stack?

John Palazza

We have CentML Serve, which is really a trailblazer in the LLMOps space, allowing you to very easily integrate and serve, preconfigure, and even get insight into how models will run on specific hardware, giving ideas around cost, optimization, and efficiency. And then we also have another platform called CentML Train, focused specifically on the training aspects within each of those areas. We then have, underneath that, a product called CentML Cluster, all of which fits into what makes our platform solution unique, and that incorporates our components around compiler and network capabilities, really providing yet another level of optimization and enhancement to that experience.

Then, at the very bottom, our platform is what enables us to provide our serverless endpoints on a variety of our Llama models that have a highly efficient and optimized endpoint, where customers can engage on an API level and get their first taste or even scale out to production-level hosted large language models.

Speaker 0

Let’s talk about the cluster management first.

John Palazza

Sure.

Speaker 0

I remember, I used to use Databricks, for example, and that was quite similar in the sense that it was a standalone startup, and it became the incumbent in many of the cloud providers, like GCP and Azure, and so on.

The weird thing is that Databricks was moving so quickly that no one could out-innovate them, and everyone loved using them. In a way, I see a similar thing with what you guys are doing. You have so many smart people working really, really quickly, optimizing these things, and you’ve almost become a kind of incumbent in many of the clouds. People can just set up and build stuff with CentML.

The cluster management is really interesting because it’s an example of what we were saying: rather than going straight to the metal, you now have this self-healing, self-scaling cluster that automatically does a whole bunch of optimization and runs jobs for you. And, of course, you can pause it when you don’t want to pay for it. But it’s just this increasing virtualization of compute.

John Palazza

Right. And to the earlier part of our conversations, so much of the GPU environment has not yet been virtualized. But what we believe is that our approach allows the right workloads to run at the right time, in the right fashion. So it’s a very similar concept of the benefits that are received from it, and I think what’s also pervasive about it is we can run that anywhere. It can run on any of the major cloud providers, and it can run on-premises as well.

Regardless of where customers are running their workloads, they have the ability to run them and actually see the difference between running them in their own environments and on their own infrastructure, and I think that’s very unique. It also allows us to run on a variety of different GPU types, both manufacturers and levels. With our platform, we actually have the ability to leverage and run some of those workloads that you would traditionally see running on an A100 or an H100, augmented by your L4s or A10s or other environments, and we’re actually giving that insight.

So, to your point, I think that’s very valuable. We have a solution that’s available on the Snowflake App Store, and what’s unique to it is that we run on Snowpark Container Services, which is something you don’t see every day. So our functionality extends not just to traditional cloud providers, but even to some of these other customers, and our goal is to become that standard to allow efficient workloads to flow to where they can run most efficiently.

Speaker 0

Yeah, I love this abstraction of a job or a workload. We were talking about having a new interface, a new platform, and in the olden days we would just say, “Okay, well, it has to run on an NVIDIA H100,” or whatever.

And now we’ve got this interface where we say, “I want to have a job, and just do what you predict will be best for this job.” The reason I’m saying this is that I think there’s this notion of “good enough” with ML models, right? Sometimes people are massively overestimating the hardware they need, maybe even the model that they need for a particular job.

Having some kind of virtualization layer in the middle means that, through experience perhaps, we can cleverly do some routing and mapping. We can run it on optimized hardware, maybe even swap out a model and share another model that someone else is using. I really think we need to do that.

John Palazza

I think what prevents things from happening in that way—and I agree with you, by the way—is if it’s not easy to do and it actually requires a decision, an effort, and a movement, people are reticent to take it. What it has to be is seamless.

Our approach is that we don’t want to have a customer rewrite their model 150 times to take advantage of every different flavor of GPU that’s available. What we want them to do is have the ability to make educated decisions that allow them to make that change.

If you give everybody the ability to run their workloads on an H100, they’re going to run it on that if they could, right? If they didn’t care about cost and they could just run it on the fastest, it’s like, if I could drive a Ferrari to go get groceries every time and they fit on the front seat, I would do it, right? But that’s not always the most efficient way to do it.

With CentML, the customer doesn’t have to worry about the effort it would take to choose between a Tesla or a Ferrari based on gas consumption to get groceries 20 miles away. They can just automatically have that chosen for them based on how long they want it to take. If it’s, “I want to deliver this model with this latency within this performance zone,” there are 16 different ways to do it, and you can enable CentML to decide the most cost-efficient or the most financially capable option based on a cost parameter with other platforms.

It starts allowing you to make business decisions without having to choose the lesser of 2 evils. I think ultimately that’s what enables people to make the right decision. If it’s seamless, painless, and doesn’t require a huge amount of change, then you can benefit from those infrastructure pieces.

And to your point, there are an unbelievable number of models and infrastructure that aren’t married correctly. There are a lot of people that should probably not be married together in those environments, but because they don’t have a path to change them or an understanding of what would happen if they did run them, and they didn’t have insight into it, they don’t do it. With CentML, we hope we can start having those workloads run on what makes the most sense, both financially and performance-wise, given some of the higher-capacity workloads and the necessity to run on those environments.

Speaker 0

Yeah, and there’s a lot of money to be saved here as well. I read on your website that, in some circumstances, I think you can save up to 60%. If you think about it, large enterprises must be spending an incredible amount of money on cloud compute.

John Palazza

I think NVIDIA might say, “Not enough.” There’s no such thing as an insane amount of money. There are so many startups that are coming out today that are trying to be a large language model for [insert generic use case here]. I think some of them have the most creative ideas and are unbelievably changing the way business will get done in these areas, but their single biggest cost is compute. Period.

John Palazza

Mm.

Speaker 1

So the startups are working on it and saying, “If there’s a way that we can extend our runway by another 6 months because we can reduce our compute cost by 60%...” So many GPUs are underutilized to begin with, right? People aren’t running at 80% or 90% utilization rates. And through no fault of their own, it’s just the way things were programmed and coded, and the way that individual developers work versus collaborative developers.

John Palazza

CentML looks to, first, increase efficiency and actual utilization rates, but then also increase efficiency in both where those models run and how they run. With a lot of our unique and patented data and approach, we're actually able to drive that consumption. I always say a website is a dangerous place to put statistics because you can always try to prove it. I actually see that our numbers are low based on what our actual customers are saying. So it's a multifold approach, and it's been very, very impactful and very, very beneficial to both our enterprise customers, and it's significantly impactful for the startups.

Speaker 0

How much of this is in the consciousness of a typical CTO in a large enterprise? I can imagine that, as the years roll on, we're talking about climate change because the compute cost will just continue to increase. But I have an intuition at the moment, certainly in my recent experience, that a lot of senior leaders aren't really thinking enough about utilization and efficiency.

John Palazza

I think that the CTOs and leaders who listen to your podcast are thinking about efficiency because I think you bring awareness to some areas that people aren't paying attention to. But unfortunately, I tend to agree with what you're saying. I think so many CTOs have yet to feel the bite of that problem. I say this having been in startups for almost 25 years. It's an interesting kind of adoption curve, and I think where we are in so much of this is innovation right now.

A lot of companies, when a CTO looks at it, look at an innovation budget and say, “I'm willing to spend X amount of money to see where this goes,” right? So it's almost like an experiment. But this transition from innovation to production to scale is when that looseness that happens in innovation needs to start changing. The customers that we're dealing with, and the customers that we're having our best conversations and best success with as a company, are customers that were in that innovation stage and are moving into production and scale.

All of a sudden, they're saying, “You know what? It's great that we're running this solution on an 8-GPU H100 setup that we've been testing, but now we're going to roll it out to 45,000 users, and we're concerned that this capability is going to bankrupt us and impact a huge amount of our costs.” That's when it becomes real, right? It becomes real when it goes from innovation to production. I think good CTOs and CTOs that are forward-leaning are thinking about that today, and that's why we're having so many great conversations. The ones that haven't gotten to that point, I anticipate having those conversations with them in the coming months and years as they get closer and closer to that moment from innovation to production to scale.

Speaker 0

Interesting. Do you think that when we make this more efficient, it will lead to less compute being used overall? The reason I say that is I had a bit of a weird thought the other day: in this Bain report, they said that we're going to be spending 25% less on call centers. That creates a bit of margin because we were spending this amount of money before, and now we've got some spare money in the budget. If it were me, I would be adding more features.

John Palazza

That's right.

Speaker 0

I've now got all of these LLMs, and I've got generative AI, and we can write code more quickly. Instead of just doing what we were doing before with call-center management, I would be adding many more automations. So it's almost like it'll just keep blowing up, but at least we can get more done with the money.

John Palazza

And that's exactly what I think is going to happen. I think it's been played out over time throughout all innovations. Efficiency is what scales excitement, it's what scales innovation, and it scales the next generation of movement. Right now, people are stuck in the mud a little bit with the amount of cost it takes to run this.

If you're able to drive higher efficiency and free up some of that capability, even free up some of that time, right? If you're able to free up some of the developer resources that are used to get to that innovation period and go to production and scale—free up those individuals, the developers; free up the creative-content capabilities; free up the cost—that opens up the next stage for growth, the next stage for efficiency, and the next stage to expand upon your value.

To me, that's been a consistent thread. I always feel that when you have a solution that can help drive innovation as it goes forward, and you drive innovation by driving higher adoption and optimization or efficiency, you're providing value today, but what you really do is unlock the value for this technology in the future. I think there's no more valuable technology in the future than generative AI and LLMs today. I think this is the biggest area of potential impact that we have in the next 3 to 5 years. I'm sure there will be more that comes out from there, but in the short horizon, if we can free that up and accelerate it, I'm very proud to do that at SentML.

Speaker 0

So, John, I want to get your thoughts on the open-weights situation. I should say open weights—it's not technically open source. But folks like Meta, for example, are giving away Llama, and in the large enterprise, they are reticent to send their data to the likes of OpenAI and Anthropic for obvious reasons. They want to have control, and they want to have pragmatism. But then there's the thing of, well, you could argue that the capabilities are better on these proprietary models, but the gap seems to be shrinking. In many ways, it's better to have the flexibility, right? Because I can do my agents and I can do all my optimization that you guys are talking about. So how do you see that playing out?

John Palazza

I think it's an interesting narrative. A few years back, my previous company, Converge, was acquired, but we were working under an understanding of what LLMs were. We had an MLOps platform and were starting to understand what we thought would be the next direction for our platform and product.

When we investigated LLMs at that time, so much of the focus was on building your own large language model. It wasn't ChatGPT or OpenAI; that wasn't what was being thought about. Meta wasn't there. People were really just thinking, “We're going to have to build our own large language model.” The companies I engaged with at that time were trying to build it. It was a very time-consuming, very hefty lift.

All of a sudden, this eureka moment came with ChatGPT, which I thought was just really going to be used for my kids' homework. Then it turned out that it was being used in corporate environments, and it's pervasive, and everyone's hearing about it. Then it shifts, and it's very much a conversation around Llama.

I think that's the right direction to go, and I think it's the direction that most enterprise companies seem to be leaning. To your point, a lot of companies that want to think about how to utilize their generative AI solutions or large language model solutions in the best way for their environment and their company think about it as a base, and then you use that base to expand and extend your reach.

I feel like it aligns very well with the developer community and the community at these companies to innovate along the lines of something that's open. As Llama has progressed in its capabilities and accuracy, achieving parity with what's available today from Anthropic and ChatGPT, I think it allows people a little easier use, and I don't think that's going to change. I actually think that it's going to be the standard, and I think that the companies that are currently engaged today will utilize that as a base.

From my conversations, so many customers are predominantly looking at the more open weights that are being made available to them, and I think that will be consistent as it goes forward as well.

Speaker 0

It's so interesting because I think we're starting to have more awareness of how to develop applications—

John Palazza

Mm-hmm.

Speaker 0

—on language models. A lot of people just go and use Claude 3.5 Sonnet or whatever, and to give it its due, I think that's a threshold for performance. I've always been quite skeptical about LLMs, but I think Claude 3.5 Sonnet is just good enough to do a lot of things because, for me, the difference in capabilities between the different models is their ability to deal with ambiguity.

John Palazza

Yep.

Speaker 0

You can zero-shot with Claude 3.5, do some very useful stuff, and it works a lot of the time. That's not to say that you couldn't take any smaller Llama model. You just need to do a bit more work. You need to do more prompt engineering. You need to give more examples. The more specific it is, and the more specifically you target it, the more reliable it is, and that's just an architectural concern. There are so many reasons why it's better to actually control the whole thing inside your own organization, because you can do all the optimizations you're talking about.

You can do all sorts of interesting routing and agents and layers and layers and layers. You can build this thing from the ground up. But it's just interesting that, at some point, there's such a thing as good enough, and I think we're very close to good enough already, if that makes sense.

John Palazza

I always like to say that the worst enemy to a startup sometimes is doing nothing, and the second-worst enemy is good enough, right? Because there's always some solution out there or some capability that's good enough. Do we really need to change it?

You never want to be at a startup that makes a Ferrari go 1 mile per hour faster. You want to be at a startup that makes it go 400 miles per hour faster, or some order of magnitude of craziness. To your point, I totally agree with you on that, and I think that we're really at this moment now where the best opportunities in front of us, I feel, are going to come from those open-weight models.

Speaker 0

You folks are partnered with NVIDIA, if I understand correctly, and Deloitte. How does that affect your go-to-market strategy?

John Palazza

We're fortunate to have some wonderful investors in our company. In addition to companies like Deloitte and NVIDIA, which both use and invest in our company, it allows us opportunities to work with them, to partner with them, to get engagement, as well as to get a pulse of where directions are going.

When you work with one of the larger service providers in the world, you definitely get a feeling about what projects or areas of focus and attention are hot buttons, both for their customers and for the space. Working within their innovation labs and their teams has really given us insight into the thought patterns of what could happen within verticals or specific companies.

That's been an unbelievably wonderful partnership for us, as well as a wonderful opportunity for us to partner technically on some solutions that they're bringing to market. I think that's been a very, very valuable partnership for us.

With NVIDIA, the opportunity to work with them, to understand and get feedback on our solutions, has been unbelievably innovative. I also think that they're invested in the success of the space, not just the success of SentML, but the success of optimization and the success of adoption and utilization rates on their platforms, because they do see products like ours being such an additive value to their stack and to their solutions for their customers.

That partnership has been wonderful. It wasn't just financially, but rather technically, and the overall support of both companies has been tremendous in accelerating both our development of a solution and our adoption with our customer base.

Speaker 0

Are you concerned at all about Azure, GCP, or AWS building something similar into their own platforms?

John Palazza

It doesn't keep me up at night, and not because they're not capable, but because of our relationships with each one of those as well. We're a partner within GCP.

Speaker 0

Mm.

John Palazza

We're available within their community. We're making it easier to work with GCP, right? Our platform makes it a very easy process. The one thing that I think permeates through each one of those vendors—and, by the way, we're also working with AWS; we're available in the marketplaces for both AWS and GCP—is that every one of those companies, GCP, AWS, and Azure, cares about its customers a lot. They care about customer satisfaction.

They realize that adoption sometimes can be limited by a feeling. If a customer feels that they're spending an unbelievable amount of money for the infrastructure they're providing, or that there are efficiencies that could be gained or utilization that could be increased, they don't see that if we can run a model more efficiently on their infrastructure, the customer will spend less money.

I think what they realize is that the customer will be happier for the efficiency that's gained, and they'll find other, more exciting and innovative use cases to leverage on their platform. We find that companies that are using SentML have a higher satisfaction rate, and I think that extends as well to the cloud providers that they're using and partnering with us on.

Speaker 0

In a way, that's one of the great things about the whole cloud-provider thing. I used to work at Microsoft, and there's an old story that, back in the Ballmer days, if you were caught holding an iPhone, he would threaten to fire you or something like that. Then they embraced the cloud, and of course they're just making money on the hardware.

John Palazza

Yeah.

Speaker 0

At some point, they started embracing open source, and they're like, “Yeah, I don't care if you run Linux,” and they just became a much more open, cool company, actually. That's reflected across the board.

By the same token, there are some negatives to these big behemoths, these hyperscalers, that control everything, right? It's almost like you're an incumbent in their sovereign cloud. Would you ever start your own cloud? How do you feel about that kind of relationship?

John Palazza

I think there's always a desire to see where and how you can expand your footprint and how and where you can help customers better. If you see a gap, I think it would be remiss for any startup not to try to step into that gap to provide value if you can.

But I feel that the cloud space itself is such a dynamic space. You had made the Oscar Wilde reference around fashion being every 6 months. I think cloud might be every 3 months, because if you Google anything—and who knows what sponsors this podcast may have, or others—it seems that there's a new AI-focused cloud coming out every 3 weeks.

I don't see a need for us to step into that area, but I do think that there's a need or a larger growth that can happen. I was actually listening to a podcast a couple of weeks ago and heard an advertisement on it. It wasn't even a technically focused podcast, and they were mentioning the type of hardware that they had in their cloud, and I was like, “That is a very specific audience.”

There were 4 people in my car, and only 1 of them knew what those references were. I think that there's definitely growth that's going to happen anyway. From my own conversations, we've had multiple conversations with newer cloud providers or unique specialized cloud providers that are looking at building out or expanding those scenarios, and I do think that's something that's relevant.

Speaker 0

It's interesting because part of the lock-in on traditional cloud is that you would need to rewrite most of your code to port it over. Of course, there are all of these portability things, but you know what it's like.

What they do is give you all of these free credits and say, “Come and build on my cloud for free,” and then they know that you're not going anywhere. But it's kind of like the same thing with LLMs. Let me think out loud here a little bit, because there's this perception that these models are just tokens in and tokens out, and you can just hot-swap them. It's not really like that, is it?

John Palazza

In so many things, people start down a path and get stuck with that path. I don't know if it's maybe my mom being a hippie or how I was raised, but I don't like to work at companies that lock people in.

Every startup I've been at has been about freedom of movement. Our platform lets people move, right? People can change models, people can change infrastructures, and they shouldn't be burdened by doing so.

Sentimel actually makes it easy to even change your model and change the architecture of the model that you're using. Going from one transformer model to a different transformer model, going from AWS to GCP or to on-prem, shouldn't be something that's so scary. It should be something that's useful, makes sense, and is usable, because this is about the customer and how the best way a customer can engage.

That freedom of movement is paramount to Sentimel, and it's built into it. To your point, I think everyone should embrace that. No one likes to be in a relationship where the other person is trying to lock you in. You want to be in a relationship for the best service, the best experience, and the right way to feel. I think anyone that takes a different approach is not necessarily the best one.

Speaker 0

I was reflecting back on this thing we were talking about earlier, that when you have more margin, you can do more. I was thinking that right now transformer models are hideously inefficient. It's insane, right? They've got this quadratic—

John Palazza

Yep.

Speaker 0

—layer-wise time complexity. That means that even if we want to do open-source stuff, we're getting the crumbs off the table.

We wait for Meta or, you know, Cohere or something like that. We take one of their models, do a bunch of fine-tuning, and start building the whole thing out. And, of course, your technology allows us to do that much more efficiently.

John Palazza

Sure.

Speaker 0

But possibly next year we might see these state-space models, and we might see these minLSTMs. It might go back to RNNs because what we've lost a little bit is the alchemy that we had 5 years ago. We want to have—and we do have—people out there just building stuff, but we want to have more people out there just building stuff. Do you know what I mean?

John Palazza

I think that's one of the things that keeps the soul of a startup, right? You have to be aware that innovation doesn't stop just because you decided to build something today. And I think the key to a successful startup is our ability to have our own alchemy that's able to keep pace with the changing market, because we're focused on that and we're looking at the same trends that you're mentioning now for the models that we're supporting and the way that we're optimizing.

Even some regions have different approaches to models that are their favorites, right? And so, as we sell in multiple regions globally, we have to be prepared for each one of those model types, challenges, and infrastructures. I think it's that flexibility, as well as the soul of a startup, that keeps Sentimel so relevant and fresh. All good startups have to have that DNA. Otherwise, we'd be kicking people out for using an iPhone, right? We don't want to be that.

Speaker 0

Final question. Just casting the gaze outwards a little bit, obviously you're an incredibly exciting startup. I think you had, was it a 27 million funding round in October 2023? What do you think about the other startups in the space? Not necessarily in your direct space, but what excites you in the broader AI space at the moment?

John Palazza

I think it's a wonderfully exciting time. There are a lot of companies being funded and created around this emerging new space, and everyone's take is a little bit different around what they're solving. Some people are looking at enterprise customers, like we are. Some people are looking at people who are just getting started—individual users, business-to-consumer.

This space is so dynamic, and some of the brightest minds in the world are coming at it to try to solve how we optimize, how we engage, and how we drive adoption and enhancement for large language models and generative AI. At any moment, I wouldn't be surprised by any direction a startup could take, and I think that's a great thing.

I also think that one of the unique traits is that our team is such a cohesive unit. Coming from the University of Toronto, we've been able to recruit some of the best and brightest minds from the University of Toronto to join our team—people who have established working relationships. So it's like a company within a company, and I think that's fantastic.

In the nature of what they've done before, we've been able to grow at a very quick pace, building a very large and diverse engineering team, and I think that is very exciting. If you look at the other startups in the industry, they're going down similar paths, right? It's so exciting to see so many teams of developers coming together to build, develop, and advance this cause.

I think we're at the beginning of what's going to be a very exciting and dynamic run, and I'm very excited to be part of it at Sentimel.

Speaker 0

Amazing. John, thank you so much, man. This has been really great.

John Palazza

A pleasure. I love it. I look forward to it. Thank you, Tim.

Speaker 0

Yeah, absolutely love it.