投资者如何使用 AI [Business Breakdowns:第 240 集]
- David Plon——前 Baupost 和 Slate Path 的综合型投资人、现 Portrait Analytics 创始人——认为,AI 应该攻克投资人的信息瓶颈,而不是消除那些帮助投资人建立确信的摩擦。 他列出 3 类目标工作流:想法生成、初步背景搭建和持续监测;对综合型投资人而言,最后一项尤其关键——他曾不止一个财报季因为最后才发现终端市场需求走弱而挨了重击。
- 更广撒网的监测问题,如今是 AI 最具象的一项胜场:它能在持仓生态系统的全部数据之上“加一层聪明的过滤器”。 持有 Expedia 时,你不需要 Marriott 相比 Hilton 的每条新闻,只需要那些能拼出消费者旅行需求、定价、OTA 分销以及市场份额变化图景的数据点;过去这类工作只有专门的行业分析师才有能力完成。
- 买入前的一大优势,是更快杀死投资想法,并把深度分析提前排进研究流程。 5 年代理可比指标的映射,如今在 Portrait 上点击一下按钮即可完成;重建 3-4 年的硬性和软性指引,则能建立管理层可信度画像——一个年年承诺利润率扩张却从未兑现的团队,可以在 12 小时阅读马拉松开始前就终结一项困境反转论点。
- 想法生成既适合捕捉趋势敞口(关税、供应链的二阶影响),也适合编码细腻的心智模型——但难点在于把模型说清楚。 Plon 自己的模型是:此前表现优异的业务在经历一次可能只是暂时的磕绊后被重新估值——“很多时候更像是一种感觉,对吧?你看到它就知道。” 真正命中时,他会说:“清空日程,接下来一周我就专门把这件事搞清楚。”
- Prompt 要像写给海外一名聪明的隔夜分析师的邮件:交代背景、任务及原因、输出形式、准则和领域知识,并要求它对一贯正面的管理层评论“保持怀疑”。 需要按任务校准准确率容忍度:“有些任务即使做到 99% 准确,也等于 0% 有用”(比如建模);而处于初筛阶段的行业调研则可以容忍偶发错误。
- 把约 15% 的时间花在实验上:保留一套模型尚不能稳定完成的 10 项任务,每次新版本发布都重新跑一遍,因为能力前沿“相当不平滑”。 使用原始模型时,不要把上下文窗口塞得过满;用到约 70% 时,“任何复杂任务都会出现性能下降”(Gemini 约 1M tokens、GPT 约 400K、Claude/Opus 约 200K)。
- 复利式押注在于现在就记录思考、决策和研究,因为模型的价值“会随着上下文数量增加而呈指数上升”;但日常采用必须自下而上,不能只靠强制推行。 Plon 短期内对记忆功能偏悲观,认为相较于直接粘贴上下文,它“有点像走捷径”;但长期看,他设想模型能够“经历一家公司每一笔投资”。代理式研究正沿着 coding 的路径前进——“这个 buzzword 背后的实质,可能比 1 年前多得多。”
1. 从投资人到产品构建者:瞄准信息瓶颈,而非确信
- Plon 的职业路径始于 Barclays 特殊机会投资业务,随后在多空基金 Slate Path Capital 担任综合型投资人,之后进入 Baupost 的公开市场团队,覆盖股票和困境债务。这个念头可以追溯到他 2015-2017 年在斯坦福商学院读书时:深度学习“在当时非常符合时代潮流”,虽然 Transformer 尚未出现,但 AI 用于研究“似乎只是迟早的事”。
- 他坚持区分研究过程与研究产出:研究的产出“希望是一项高质量决策”,而有些摩擦本身会建立确信——“我永远不可能真正把建模外包出去”。因此,AI 应该被用于那些受每日时间限制的环节,而不是替代那些通过艰难过程建立归属感的工作。
- 他将工作分为 3 类:想法生成——任何时候可能都有几十个处于甜蜜点的想法,但很难知道自己到底在研究哪一个;作为综合型投资人搭建初步背景——掌握行业运作方式属于“基本门槛”;以及持续监测客户、供应商和竞争对手组成的周边生态系统。
2. 监测:覆盖整个生态系统的智能过滤器,而不只是盯着股票代码
- 监测单一标的已经有很多成熟方案,真正的突破在于面对更稀疏的信息流“撒更大的网”,寻找相关数据。以 Expedia 为例,投资人并不关心 Marriott 的门店增长相对 Hilton 如何,只关心那些能够拼出消费者旅行需求、定价、OTA 分销策略和市场份额变化图景的数据点。过去“没有一个聪明的过滤器覆盖所有这些数据”,而 AI 让这件事简单得多。
- Matt 以自己做运输分析师时的经历印证了这一点:周四深夜,别人都在曼哈顿享受欢乐时光,他还在 CPG 电话会文字稿里按 Ctrl+F 搜索运费成本相关表述。
3. 买入前工作:更快杀死想法,把深度研究提前
- Plon 对自己投资工作的年终复盘是:真正进入深度研究的想法“少得惊人”,其中很多本来可以更早被淘汰——在投入“12 小时读完所有历史材料”之前,致命风险或可比公司错配就已经能够浮现。
- 被提前拉到买入前的分析主要有两类:可模板化的筛选,包括 5 年代理可比指标映射、激进收入确认风险标记;以及模式搜寻——重建 3-4 年的硬性与软性指引(比如“我们预计收入将在下半年某个时候加速”),以刻画管理层可信度。如果一家困境反转公司的管理层连续 3 年承诺利润率扩张,却一次都没有兑现,“这或许就足以杀死这个想法”。
- 值得保留的细节是:Bloomberg 筛选器可能显示管理层“每个季度都超额完成指引”,但深入查看会发现,Q1 给出的全年指引在之后一整年里不断下调——“你可能应该给管理层说的话打个折”。Matt 提供了另一面:把所有坏消息一次性塞进某个季度,有时“可能更好”,胜过连续不断地下调指引。
4. 想法生成:把“一种感觉”变成查询
- 第一种模式是捕捉趋势敞口:关税首次宣布时——“我想应该是去年 4 月”——投资人会问,哪些公司的供应链主要位于美国,而竞争对手的供应链则分布在国际市场。AI 擅长这种二阶映射,现代模型还拥有大量关于哪些公司可能受到影响的世界知识。
- 第二种模式更难:把细腻的心智模型编码进去。Plon 自己关注的是,曾经表现优异的业务遭遇某种磕绊,例如宏观问题、糟糕的产品周期或执行混乱,市场随之重新定价这项业务。逆风可能只是暂时的,但“无论是不是暂时的, headline numbers 看起来都会很糟”。
- 真正的障碍在于表达:有些投资人可以列出 A、B、C 等属性,但“很多时候更像是一种感觉”。Portrait 会从交易历史和机构语境倒推,定义“对你而言,10 分想法是什么样”。一旦奏效,“这个行业里很少有比这更令人兴奋的感受……我会清空日程”。
5. Prompt:像写邮件给海外的聪明分析师
- 即便有效的 prompting 每 3 个月都在变化,最持久的心智模型仍然是:把邮件写给一个聪明、正在海外熬夜工作、但不了解你背景的人。邮件应包括背景、任务及其原因(“因为我认为成本曲线可能正在变化,所以请把它建出来”)、可选的输出格式、任务准则和领域知识。
- 有时应该不提供输出规格:限制格式可能损害结果,就像给人类分析师布置任务时,也要“把绳子放松一点”,给他留出创造性综合的空间。
- 他最喜欢的一句领域经验是:管理层评论“总是带有正面偏见……保持怀疑非常重要”。这是必要的,因为有用性训练会把模型推向乐观,就像“一个热切的大学生,既兴奋又相信一切都会向好”。
- 需要根据任务 stakes 和结构校准标准:“如果你做到 99% 准确,也等于 0% 有用”,这适用于建模等高精度任务;早期初筛式调研则可以容忍偶发的事实错误。结构化任务应上传文件,因为现成的网络抓取并不完美,可能产生幻觉;探索型任务则要指示模型“沿着线索继续挖,追踪潜在线索”。此外还要持续迭代:Matt 指出回复即时、查询成本近乎为零,Plon 则把这个过程称为“一场双向舞蹈”。
6. 把能力视为一条崎岖且不断移动的前沿
- 把约 15% 的时间用于实验,因为能力前沿上存在“你可以远早于他人发现并拿到手的未发掘能力”。Plon 每次新版本发布,都会重新运行一套模型尚不能稳定完成的 10 项任务,以衡量这次跃迁并重新标定能力边界;一旦某项任务——比如成本曲线——变得可行,就保存模板并复用于研究流程。
- Matt 自己做过一项 prompt 实验:逐步提高同一 prompt 的具体程度,观察输出如何随问题顺序和上下文加载而变化。这项技术“没有多少按钮会告诉你:哦,现在这是四轮驱动模式”。
- 过去大约 1 年,Gemini、GPT 和 Opus/Claude 的上下文窗口大小基本没有变化,分别约为 1M、400K 和 200K tokens,但使用效果有所改善。模型过去能够在海量材料中找到“藏在干草堆里的针”,却难以跨 5 份 10-K 建立“三表模型”。在 Portrait 或 NotebookLM 这类托管工具里,Plon“基本不会刻意控制输入量”;但对原始模型而言,使用约 70% 的上下文窗口就可能导致复杂任务性能下降。
7. 复利式押注:现在记录,让代理慢慢到位
- “你必须使用某某工具”之类的自上而下指令可能适得其反。有效的做法是推出不要求强行改变流程的全公司项目:定制化想法生成筛选器像“外包分析师”一样运行,投资论点则可以从股票代码出发进行监测;至于日常使用,应当逐个人、由下而上地建立信任。
- 体育分析带来的启示是,模型的价值“会随着上下文数量增加而呈指数上升”。因此,记录交易为何发生的备忘录和短段落会成为有价值的知识产权——“很难想象 2-3 年后,仍有任何一条数据没有被纳入一个处于基金语境中的模型”。Matt 分享了自己的入职经历:追踪正在浮现的 Amazon 威胁的快速财报后短评,比经过精心打磨的买入和卖出备忘录教会他的东西更多。
- 关于记忆功能,他给出了坦诚的保留意见:短期看,它只是重复输入上下文时“有点像走捷径”,相较于直接粘贴内容,他“有点悲观”;但长期看,他设想模型能够“经历一家公司每一笔投资”,吸收优秀投资人直觉判断背后的经验和“伤疤”,最终可能比任何单个人类都更强大。他提出的 AGI 测试是:“它能预测未来吗?”
- 代理式 AI——围绕一个目标进行推理、反思、采取行动,再反思这些行动——已经从 buzzword 变成了可运行的系统。Portrait 最初的 GPT-4 agent 需要约 30,000 token 的系统提示词,因为它无法可靠地自我纠错;如今 Claude Code 和 Codex 已经可以被放置数小时而不受干预。代码是最理想的试验场:上下文局部明确,而且“代码要么能运行,要么不能运行”。底层的迭代式推理能力如今已经可用,把它迁移到投资研究仍是工程问题——“归根结底就是工程工作”。
完整逐字稿
I'm Matt Reustle, and today we have a special episode breaking down how investors are using AI. This is a question I get from many of you. While there is no shortage of content on the implications of AI, I know there's an appetite to learn more about tangible use cases, how to make sure you're getting the most out of these tools, how to think about advancements in the technology, and how to keep pace with the innovation curve.
My guest today is David Plon, founder of Portrait Analytics. Now, David and Portrait have been a partner of Business Breakdowns since last year, but I specifically asked David to do this episode because, first, he is front and center in how investors are using and applying AI. Second, and maybe more importantly, he and his team come with a background in investing. While the conversation doesn't focus on Portrait, you'll hear references to what he and his team are building and how they've shaped it for investors. You'll understand when you hear David talk that he is someone who understands the pain points of an investor.
I think everyone will find something in this episode that will benefit them in their day-to-day work. David, I'm excited to finally do this recording. Over the holiday break, many investors were spending time trying to figure out how they could get more out of AI tools. It's constantly evolving, but at least for a snapshot in time, I think you'll be able to give us some really good perspective on use cases and some of the skill sets that can help people get the most out of the technology.
I wanted to start with your background because I think it's really unique and speaks to why I find you such a valuable resource when it comes to thinking about this through an investor's mindset. Can you bring us back in time, prior to Portrait, what you were doing, and how that led to starting this business?
Yeah, absolutely. It's a real pleasure to be on here. I spent my career as an investor in a few different seats before starting Portrait. I started at Barclays on the trading floor in the special situations group. I was then at a long-short hedge fund called Slate Path Capital, where I was a generalist, and most recently was at Baupost, which is a large hedge fund based in Boston, where I was a generalist on the public markets team doing both public equities and distressed credit.
I think I probably first got interested in the idea of using AI within the investment research process when I was at business school at Stanford, from 2015 to 2017. Deep learning was very much in the zeitgeist. We were still pre-transformer right up until I graduated. It just seemed inevitable to me that, at some point, this technology would be really impactful for a lot of the research workflows that I went through.
After spending a lot of time studying, exploring, and prototyping, I got the conviction to throw myself full-time into this project. So that's my background leading into starting Portrait.
When you think back to some of those stints, whether it's Baupost, Slate Path, or even potentially on the trading floor, can you think of any scenarios where you were spending a significant amount of time on pain points that might be easily solved by AI today?
We'll get into some of the very specific use cases, but I'm curious, when you look back at that time and some of the workflows you had, where you could see the biggest lift today versus where it was 10 years ago. It's very different today.
1. The Investment Research Bottleneck
It's interesting because the output of an investment research process is ultimately a decision, and hopefully a high-quality decision. It's very unlike other fields where there is actually a widget or a service being provided.
When I thought about the investment research process, there were aspects of it that had friction, but that friction was helpful because it helped me build conviction personally. I was one of these guys who could never really outsource model building. I always had to build a model myself in order to feel conviction in recommending a position based on that company.
There were certainly aspects that limited me in terms of how productive I could be in finding and researching the best ideas. Essentially, I would bucket it all into the category of there being a ton of information out there that I could potentially consume. Consuming it in a way that was efficient and additive to the research process was challenging, given the number of hours in a day.
I would put it into 3 main categories. There's idea generation: finding the ideas that best fit my mental models. I knew there were probably dozens out there that were totally in my sweet spot, but at any given time, it was very hard to know if I was working on one of those.
There's also the initial context-building that goes into a new idea, especially as a generalist. It's not even about developing a differentiated insight, but just getting the table stakes on how a company and an industry work, the narrative, and the business breakdowns. Business Breakdowns has always been a very useful resource. Whenever I was working on something where there was an episode, that was always a huge help.
Lastly, there's monitoring an investment idea. As a generalist, it was often challenging to stay on top of not just the companies, but the surrounding ecosystems—the customers, suppliers, and competitors. If someone lived in a given sector, they were usually able to put together a real-time mosaic based on all the data points they were absorbing about what was happening in the industry. I always found that really challenging as a generalist.
There was more than 1 earnings season when I got smacked because I was the last person to realize that end-market demand was weakening or that suppliers were pushing prices through. Ultimately, I would say those are the types of workflows—high-volume information processing to help triage new data points, broadly speaking—where AI can be really useful. We can get into some specific applications of that, but that's how I would think about where AI can fit in a research process without replacing the parts of the process that are important for building conviction.
That really resonates and seems to align with the people I've seen getting the most out of it. A lot of it is productivity and efficiency, which on the surface level sometimes people can dismiss, but when it comes to making decisions, having clarity of thought, and being able to take in the most valuable information, I think there's a lot more to it.
I like the idea of getting into some of the use cases. I wanted to start backward in terms of the buying process and reference what you just brought up in terms of position monitoring. What are some of the tangible use cases of AI today that you think investors are using, or that you would recommend they use, to monitor what's in their portfolio and stay on top of it as effectively and efficiently as possible?
2. The Portfolio Monitoring Mosaic
I think there are many good solutions today for staying on top of what's happening with an individual name. There are plenty of ways you can add it to a watchlist on any number of services and stay on top of all the new filings, transcripts, and news related to that specific company.
What I think has historically been hard, but is now much simpler, is being able to cast a wider net and pull in relevant data points from a much sparser set of data. For instance, let's say you own Expedia. One of the key investment factors for Expedia is always going to be what's happening in the hotel ecosystem: how volume is trending, how pricing is trending, what OTAs are doing with their distribution strategy, and how it relates to other OTAs.
The challenge historically was that if you're an internet analyst or a journalist following Expedia, you probably aren't going to want to subscribe to every piece of news about what Marriott is saying. You don't care what their profitability necessarily looks like or whether they're adding new units to take share from Hilton. What you really care about is any specific data point that would add to the mosaic around what's happening in the hotel landscape with respect to consumer demand for travel, or whether they're seeing trends and market-share shifts with respect to OTAs.
Historically, it was impossible to pick up all those data points unless you really consumed every piece of news. There was no smart filter on top of all that data. I think with AI today—and obviously the company I run, Portrait, has solutions that target this—but even outside of Portrait, there are plenty of ways today to pick up those data points much more efficiently across that surface area. What was once a level of understanding unique to someone who just covered travel, you're now able to pick up much more efficiently as it relates to your investment thesis.
I can definitely speak to that as a transport analyst trying to monitor truckers and knowing that there was valuable information coming out of CPG companies talking about freight costs.
I spent many late Thursday nights while the rest of the world was enjoying Manhattan happy hour Ctrl+F-ing through these various transcripts and earnings releases. I think that is a major area that has evolved, just in the ability to gather all of that information in a much more efficient way and not Ctrl+F the heck out of things.
When you next move into building up research on a name, or the pre-buy research process of getting up to speed, this is another one where I think it's the widest net of possibilities. When you think about how you used to approach it versus how you're seeing other investors approach it, what are some of the use cases that you're seeing people implement to really get the most out of it and maybe build better conviction or gather more information, whatever it might be?
3. Triage Before Deep Research
It's interesting because I was someone who—and still do—love doing the pre-buy work. My wife likes to joke that my favorite Saturday morning activity is reading a 10-K. The thing that's interesting about what I mentioned before is that, to me, that's an important part of the conviction-building process: building up that context and feeling like you've read all the material and can start piecing together the various data points.
Where I think AI can be really helpful in that process is in 2 things. One is getting you enough information to know whether an idea should be killed. I used to do this accounting at the end of every year when I was an investor: How many ideas did I really look at? A remarkably low number made it through to the deep-research stage. A lot of times, I could have killed an idea much quicker once I had surfaced enough information to understand that there was maybe an existential risk here that I was never going to be able to get past, or that management compensation was structured in such a way that I was never going to feel comfortable that their incentives were aligned.
I think where AI can be really helpful in that process is giving you enough baseline context to be able to triage the idea: “Oh, yeah, this is passing the initial sniff test. I want to spend more time on it and maybe spend 12 hours reading through all the historical content.” The other thing that has been really helpful is expanding what gets moved up earlier in the pipeline in terms of types of analyses.
For instance, I mentioned earlier CEO compensation. One of the things that I would do if I was really getting into a name is go through the last 5 proxies and try to map out the metrics that the CEO is being compensated on, how those have changed, and how the weighting of those metrics has changed. A lot of times, that's a useful signal into what the board is thinking about. That's a bit of a pain. Now, on Portrait, it's the click of a button.
Things that historically would be more in the deeper-dive category, I can now move up. That, to me, is one of the most powerful use cases of AI. As an analyst, you're able to turn over far more rocks in a given period of time than you would otherwise and spend more of your deep-research creative time on ideas that you know have already passed your initial sniff test and are worth that time.
When thinking about really tapping into that, would that involve having a list of non-negotiables in terms of whether it's compensation elements or something else that you can screen through quickly? Or do you not necessarily know what you're going to hit that might break it, but you're just going to get there faster through the various iterations of getting up to speed? I know it's a little bit technical, but I'm curious mostly if there is a way that you think you could templatize that versus not.
Oh, yeah, there are 2 categories. There are definitely things that you can templatize. For instance, I mentioned CEO compensation. There are certain accounting things I would look for with respect to aggressive revenue recognition or changes in certain assumptions that would inflate figures. I'd say there were also patterns that, if I saw them, would turn me off.
One analysis that is now very trivial but took a lot of time historically was going back through the last 3 or 4 years and laying out every piece of guidance that the management team had given—both the hard, specific guidance and anything qualitative or soft, like, “We expect revenue to accelerate sometime in the second half of the year.” Building up a picture of management's guidance style and credibility is really important, and that's the type of context that someone who's followed a name for a long time intuitively has that, being fresh to a name, I wouldn't have.
That's a situation where, if my thesis is based on some sort of turnaround execution plan that the management team is undertaking, but I can do the work and say, “Actually, over the last 3 years, every year they've said they expect margins to expand, and it never happens,” that would be enough to maybe kill the idea, depending on the other circumstances.
I think that's another cool example of where AI can be really useful: surfacing patterns that historically were pretty painstaking to put together and now can happen far more easily.
Yeah, 100%. It's a very valid point, too, on credibility. I think there's a trust level, and once investors show they don't really trust a management team or a stock, it could be a free fall. You see that quite frequently.
Yeah, and it can be really subtle, too. There are times where, if you just look on the Bloomberg screen, it says management beats guidance every quarter. Then you actually dig into it and realize that when they give full-year guidance in the Q1 call, they end up revising it down quarter after quarter after quarter. By the time they get to Q4, sure, they're going to beat their numbers. But if you're building your model looking out a year, you probably should subtly shade whatever management is saying at that point.
There is effectiveness in just kitchen-sinking and taking all the pain in 1 quarter, and that can be better than just consistently revising down. There are all types of things that, to your point, feel like they're definitely quantitative in nature, but the ability to extract that information was not always easy. I think that's an entire category where AI is showing up.
The last section is in sourcing new ideas, and maybe it's something you alluded to before, where you have certain frameworks or characteristics of businesses that you really like. It's not always easy to find all the companies out there that fit that, or to know if you're looking at a company that might fit that. Where, in idea sourcing and idea generation, are you seeing AI be most effective? I will just mention that, to me personally, it feels like there's always a step before you get to sourcing. You're usually introduced to a name for other reasons. But I'm curious where you're seeing this come into play.
4. Finding Ideas Through Mental Models
There are 2 main ways I'm seeing investors use AI for this type of use case. One is understanding companies that are exposed to a given trend or a given development. For instance, when tariffs were first announced, I guess in April of last year, there were a lot of investors, especially on Portrait, trying to figure out which companies were going to be most exposed.
The second-order thinking is even trickier. Which companies, for instance, have a predominantly U.S.-based supply chain while their competitors have an international supply chain? That type of analysis is where AI is really useful. Even out of the box, there's just a lot of world knowledge built into a modern, state-of-the-art model about first- and second-order effects, who might be affected, and so forth.
The more nuanced and, I'd say, difficult case—and where we spend a lot of time—has been working with firms to find ideas that fit a nuanced definition of a mental model. For instance, in my past life, I had this mental model of companies that were previously high-performing, where there were really attractive elements of the business model that you could believe in and underwrite, and then there was some sort of hiccup.
It could be something macro-related, a really bad product cycle, or some sort of execution mess-up, and the market is now revaluing the franchise value of that business. A lot of times, I could isolate the research to that one specific headwind and figure out if it was temporary or not. If it was temporary, then you could buy it.
That's a pretty nuanced thing to find. It's really qualitative in nature because the headline numbers are going to look bad in the near term whether it's temporary or not. Finding ideas that fit the mold of that is challenging for 2 reasons. One is that it requires a lot of qualitative reasoning along with the figures to do that.
But, 2, I'd say even enunciating that clearly is hard. When I ask a lot of investors, “What are the mental models that you think about when finding a really attractive idea?” some of them have been very explicit about writing down, “I'm looking for attributes A, B, and C.” But a lot of times, it's more of a feeling, right? You know it when you see it. Figuring out how to translate that into a query, for lack of a better term, is challenging.
That's one area where we spend a lot of time at Portrait as well: helping folks look at their history and figuring out, “What does a 10-out-of-10 idea look like for you, given your historical trading context and firm and all of that?” I think the combination of those 2 things is difficult, but when it works, it's incredible.
There are few feelings as exciting in this business as getting served up a pitch where you're reading it and you're like, “Oh, my God, yeah, this is really exciting. Clear my calendar. I'm going to spend the next week just figuring this out.” That's a really magical moment, and it's something AI has been really helpful with.
Yeah, I can tell you from doing the show, I get to experience all these different businesses, and I'm trying to draw the patterns and match them up. Some of them come more simply than others.
The mission-critical part of the supply chain that's a small percentage of the overall cost, but they're a dominant player. They have all the market share. Pricing power should be there. That's one that's emerged, and I think gained a lot of appreciation over time.
And then there are the more subtle ones like what you described before, where really articulating them can be more challenging. But that is an interesting approach, and actually something that I've seen a few investors who have a differentiated framework try to lean into: finding ways to take that proprietary knowledge or pattern recognition that they developed over time and use this tool to get the most out of it in terms of uncovering new opportunities.
I think one of the things that you've made clear to me in other appearances you've made, and as we've talked, is that there are skill sets to get the most out of AI. I don't think they're always that obvious, or at least they weren't obvious to me. But I was hoping we could just talk through some of the things that seem to make a massive difference in terms of whether investors are getting a better-quality output than a Google search, versus really building on what they're doing and becoming massively more efficient.
The first is just prompt writing. How would you articulate how important prompt writing is? What are some of the basic things that you think benefit investors when they're putting together prompts? Feel free to wax poetic on it.
5. Writing Better Investment Prompts
It's an interesting question. The answer changes every 3 months in terms of what makes an effective prompt because the model capabilities change, as do the tools available to the models. They're constantly changing in the investment research context, and I think this is true for other contexts as well, but particularly research-based workflows with large language models.
The mental model that I think still works really well is to imagine you were writing an email to somebody, maybe overseas, who's going to work overnight and is going to be doing a task for you. Assume they're smart but maybe lack a lot of context on you. What information would you want them to have to be able to do a good job? Then whatever you end up writing is probably a pretty good starting point for an effective prompt.
To dig a little bit deeper, some of the subcomponents, when I write prompts today, are that I usually outline a specific task and why the task is happening. In the same way, if you give an analyst, “Build this cost curve,” they say, “Okay.” If you say, “Hey, build this cost curve because I think it might be shifting and that could imply something about future changes in the pricing,” that's helpful.
I usually provide some background context. I provide a task. Depending on the nature of what I'm looking for, I might specify an output. I might say, “Hey, I want your output to look like this. Have an intro paragraph, then outline these data points. Show me a table that has X, Y, and Z.” That sometimes is useful. Sometimes that can be hurtful, in that you're constraining the model's output. In the same way as with a human analyst, you might want to give them some slack on the rope to be able to be creative in how they want to synthesize the information.
I usually have a section of my prompts where I have specific task guidelines and then also just general domain knowledge. For instance, in terms of guidelines, these are things that I would just think about: “Hey, maybe I'm doing that guidance prompt. Make sure you capture soft guidance that doesn't have a specific figure on it.” Otherwise, if you say just capture all guidance, it's only going to capture specific quantitative guidance.
In terms of transferring domain knowledge, I have a set of bullet points that I try to encapsulate. Here's how I think about being an analyst and some of the things you might learn through your experience. One of those is, I always remind the model, “Look, you're going to largely be reading commentary from management teams. They are always biased positively. It's important you take a skeptical eye to anything they're saying.”
Even something that simple really helps because the models have been trained to want to be helpful, and that usually influences them to be more positive than they otherwise should be. You could imagine taking an eager college student who's excited and believes in the good and trying to impart a little wisdom: management teams tend to inflate things a bit.
Just to wrap that up, defining the task, the context behind the task, some specific outputs if I want them, some guidelines, and some kind of domain context—I think that's a really good starting point. If you can send that to a human and they read it and understand the task, you're probably going to do pretty well with the model as well.
Yeah, the level of instruction—if you almost humanize it—that was one of the initial things that I think you mentioned months and months ago that made a material change for me. And then, on the idea of context and where we'll use an LLM, assume an off-the-shelf LLM for this question: when you're thinking about when to just give them massive freedom versus when to load in documents, how does that change the scale and the output quality in your experience?
So I think it's very task-specific. There are a few different dimensions on which it depends. One is the nature of the task: is it something defined and structured and maybe more quantitative? In which case, yeah, it can be a pain if you're using an off-the-shelf solution, but you should upload the documents, because otherwise, generally speaking, an off-the-shelf model is going to have to crawl the web. That's an imperfect, inefficient way to gather a bunch of structured information, and you'll most likely get hallucinations, which is why, of course, companies like Portrait exist: all that stuff's preloaded.
To the extent that the goal is something more exploratory or creative, it can still be really useful to upload the relevant documents. But I would explicitly instruct the model: “Hey, I don't really know which way this answer is going to go. So pull on threads, follow it, chase down leads.” In that case, you can steer the model, especially modern ones that are quite responsive, to do a lot of web searching.
Where this matters as well is how important the task is with respect to accuracy versus usefulness. There are certain tasks where if you're 99% accurate, you're 0% useful. That's like self-driving cars, where the cost of an error is exceptionally high.
I think of model building: if I presented a model that had an obvious flaw in it to my boss in any of my past roles, that would be pretty bad. I certainly did that on occasion. [Laughter]
Two of us.
Yeah. Maybe you're just doing an industry survey: What's the history of this industry, and how has it evolved to these 3 players? If there's a factual error here or there, and it's earlier in the research process and it's more of a triage exercise, that's okay. I'm fine with that. Having some calibration depending on where the task is in the process, the importance of accuracy, and how structured you want the output to be, I would vary my approach and experiment.
Yeah, it's quite interesting to see where things matter based on precision versus being directionally right, and where all the value comes from. On your point about experimentation, I'm curious how much this matters, and this is something that I mentally struggle with sometimes as well. You mentioned the technology is constantly changing.
By experimenting or getting better with certain skill sets that eventually the models are going to be able to do on their own, what difference do you see in investors, or in yourself, when you experiment with different things in creative ways, when you know that there's a standard, decent-quality way that you can solve it already? I know it's a broad question just about experimentation as a skill set, but I am curious if you've noticed any patterns with the clients that you work with, or with yourself and some of your teammates.
In an evolving technology like this, I think spending some 15% of your time on experimentation is really important because, unlike past technologies, one, the underlying models are changing, but two, the capability frontier is evolving as well and is quite jagged. There are a lot of undiscovered, for lack of a better term, capabilities that you can pick up on well before others do.
Having a process by which folks experiment and really spend some time pushing the models can be helpful in its own right, but it also gives you a better intuition of where the models are today and maybe where they're going. To add a little specificity to that, I—and we at Portrait—have plenty of tasks like this.
Even personally, I have some tasks where I know the models can't do this today, or maybe they can do it a small percentage of the time. Anytime there's a new model that comes out, I run my suite of 10 different exercises across it. I, for one, can get a feel for how much of a leap this model is, but I might also now need to update where I think the edge of that capability frontier is.
A lot of times, once you start doing that process, it kind of builds on itself. You might have an experiment—maybe it's building a cost curve and seeing how that has shifted. That's a type of task that can take a long time and be challenging, and now the model can start doing it. Well, great. Now you start having ideas of, “Well, I wonder if I layer in a long-term cost curve versus a marginal cost curve.”
There are different things you can start coming up with and different ways you could actually apply it within the research process today. I would say for anybody—certainly companies do this, but individual investors can have a suite of 10 different things where, if the models could do them, it would be great—and just constantly run these against the model. Once they're able to do them, save that template and reuse it as part of your research process when working on a name.
So I think those are some of the things where experimentation can pay a lot of dividends. I would bucket this as experimentation, but one of the things that you laid out when you talked to Brett from Fundamental Edge was just writing prompts in different forms. Start out with the simplest prompt that you know is probably going to give you a very broad, mediocre answer. Then take it down in more detail, specify more detail, and see how the responses change based on how you're instructing it via prompts.
I did it a bit, and it made me appreciate how sensitive it is to maybe the ordering of questions, or whether I specify the ordering of questions and when I'm loading into the context window. It doesn't just impact that single exercise. It makes you appreciate that this is a technology that really doesn't have many buttons that tell you, “Oh, this is four-wheel drive that I use in this condition.” It's just a chat, and you have to play around with it yourself.
And just on that, I think one thing that's really cool that I mentioned earlier with the way I write prompts is I think about sending an email to someone overseas. The difference with an LLM is that the response is, for all intents and purposes, instant. The real insight, I think, comes from doing exactly what you just described, which is iterating on this.
The cost of sending a single query is trivial. Start simple. Start adding complexity as it's helpful.
It's really like a two-way dance to end up in a spot where the AI really is acting like a useful analyst.
You don't need to wait overnight to get the content and then see how you give feedback. You can give that feedback in real time and adjust things. That becomes super powerful.
Your prompts are loaded into Portrait, whether it's primers or whatnot, and you see the levels of detail that go into those. It's actually quite helpful. It made me feel a little down on myself in terms of what I was prompting prior to seeing those, but I think it definitely rings true.
Taking it up a level, I have seen quite a few individuals who are getting a lot of value out of AI, but it almost feels like it's siloed. You could have 2 individuals at the same fund using entirely different tools and getting value from them. Obviously, the enterprise or the fund is benefiting from all of it, but it doesn't necessarily feel strategically aligned, or like there's much collaboration around it.
Have you noticed anything in terms of best practices for funds to get some type of adoption that's aligned or collaborative? Investing is ultimately about building conviction, and people build that in different ways. It's challenging to mandate changes in a research process because if I was at my last job and the head of the firm said, “Okay, now you must use only sell-side models for your models,” I would really struggle.
6. Scaling AI Across Investment Funds
I can certainly sympathize. I know a lot of firms have been trying to take a top-down approach and say, “Hey, we don't want to fall behind on AI. You need to start using such-and-such tool.” I think that can sometimes work in certain ways, but in other ways it can backfire and people push back.
What I've seen from the firms with the most adoption and that are taking the most advantage of this, I'd say, is finding the right balance between having firm-wide initiatives that don't force people to change their behavior, while letting individuals experiment and figure out where AI is going to be additive and reduce friction without negatively impacting conviction.
For instance, just using Portrait as an example, there are firms where, at the firm level, we spend a lot of time building idea-generation screens that are super bespoke to them and their process, and we just run those like an outsourced analyst. That's really valuable because all we're doing is giving them a bunch of really interesting pitches that they can spend time on if they want to. We're not asking anyone to change their process.
Similarly, thesis monitoring has become a really big thing. The nice thing about thesis monitoring is that it doesn't require a change in process. All we're doing is taking a ticker and maybe your investment thesis, and then pulling in data points that at least our system believes could be really helpful and incremental to your thinking. Again, that doesn't require anything new. It just requires looking at a new insight when it arrives and ignoring it if it's not relevant.
Whereas, for the actual day-to-day work of asking queries and building outputs, it really just depends on the user. Obviously, working with a firm that specializes in this and has built software specifically for this vertical can be really helpful for driving adoption. Unlike a general tool like ChatGPT, someone building software in this space can customize the user experience around specific workflows.
Ultimately, I think adoption has to happen bottom-up. People need to feel comfortable with it to be able to continue making high-quality investment decisions. To the extent that people have changed their process—and many people have—they've done so because they trust it. Building that trust has to happen on the individual level.
Yeah, it's actually an interesting perspective that I somehow hadn't thought about. When it comes to any type of investment research and investment process, people do it all differently within the same fund. They might have the same core things that they look for, but the path to get there is quite different, including the format of their models. I can remember many arguments over having a single specified template and people moving off of that.
So, yeah, I think there's a lot to that, and maybe it opened my eyes a little bit to the practices that firms or individuals can adopt in terms of investing in things that may not have obvious value today, but as the technology evolves, maybe they'll be able to leverage them more. I mentioned before that we talked about the famous example of sports analytics teams that had been tracking data well before sports analytics were a thing. They were only able to really benefit from all those data practices decades into the future, but for the teams that never kept up with it, it's really hard to go back in time.
Are there exercises or practices that you think funds or individuals can start to implement that will pay off over time?
Obviously, everything we spoke about earlier with experimentation certainly applies here. But I think one thing that's going to be really important—and already was historically, but I think carries a lot more weight going forward—is the importance of documenting thinking, decisions, and research.
Essentially, the usefulness of these models rises exponentially with the amount of context they're given. As the model context length expands, and as their agentic reasoning gets more heavily utilized to do longer-running tasks, they will ultimately shift from being a helpful research tool to something that lives within a firm and can execute a research process. Their ability to do that is going to be a function of how much data they have on how you and your firm operate.
I think firms where people have documented memos and even written up short paragraphs on why they made a trading decision will benefit. It's hard to know exactly which data is going to be used and how, but I think it's a pretty reasonable bet that having that data, at a minimum, is helpful for humans and will certainly be helpful for machines.
A moment ago, we were speaking about adoption. A challenge is loading in that context. It's the same challenge you face when you hire a junior analyst. You need to teach them everything about how you think and your past trading history with a given name, and all that information takes a ton of time to convey.
To the extent that data can be captured in real time, that forms an enormously valuable source of intellectual property that will make AI unique to you. Again, it's a little hard to know upfront how that's going to be used, and it probably takes some creativity to figure it out. But it's hard to imagine that 2 or 3 years from now, every piece of data won't be used within a model operating in the context of a fund.
It's a perfect example. I could speak to onboarding at places where all they had were the memos, which were the buy and sell decisions for positions. Those are valuable in some ways, but they're usually overly polished. That's different from places where, after earnings, they would have a quick blurb: Maybe Amazon entered the space, and it was something to monitor. Then, a month later, they had a meeting with the management team, and it made them feel a little worse off. Another quarter comes out, and it's clear Amazon is impacting them. The memo might say, “Amazon's movement,” but seeing the thought process and how the team makes decisions was very valuable in helping me get up to speed.
Again, I think it goes back to that lens of almost thinking of it like a team member, in the same way that you would give them the most ammunition to align with how you work and what you do.
Going back to some of the quick-hitting questions and thinking about where we're going, I already brought up context windows. With LLMs, again, just thinking about some of the most-used tools, where do you personally find that the context-window advantage has the most impact? Or how do you adjust for the challenge that I think many investors face, which is, “I want to get as much information in there as possible, but sometimes I'm limited by it”? How do you approach that, and how much do you think you should worry about that as an investor?
It's super dependent on what tool you're using, the context in which you're using it, and the task. To go a little deeper on that, the context windows obviously have been expanding, although really in the past year they've been flat, with Gemini at around 1,000,000, GPT at around 400,000, and Opus or Claude at around 200,000.
Where the models have gotten more capable is in using that context. Historically, models struggled. If you loaded in, I don't know, the last 5 10-Ks, and let's say a 10-K is 100,000 tokens, that's 500,000 tokens. The models would historically be really good if you asked, “What was SG&A in 2023?”
They called that a needle-in-a-haystack problem. They would struggle if you said, “Build a three-statement model using these.” The model needs to apply its attention across lots of different pieces of the context window simultaneously. I think that's still a roughly helpful rule of thumb: If you're using a ton of context, you do better if you're loading that in because you're looking for very specific data points.
That being said, I think a change versus if we had this conversation 3 months ago is that models have gotten very good at using the context intelligently. If you're thinking about building a project in Portrait or NotebookLM or something like that, I wouldn't really hold back. I would give it pretty much the entire corpus, and you're seeing this a lot in software engineering, where Claude Code and other agentic coding tools are very capable of exploring super-long context in a way that is efficient and focused on the task at hand.
In the same way, I wouldn't give a junior analyst a task and say, “These are the only 4 documents you can use.” You'd want them to ultimately have the agency to choose what content they're going to use, and I think that holds true. All that being said, if you're going to use the raw model itself, you do not want to overload the context window and use 70% of the context window. You will see a degradation in any complex task.
In that case, you should break it down. But hopefully, the much easier way to do this is to use a tool that manages that complexity already. You don't need to worry about, “Okay, what percentage of the context window am I utilizing?” That's how I think about that today.
On memory, I think it's obvious. Whereas memory has now been implemented in a lot of these tools, it customizes to you and that iteration. Again, I'm going to overuse the lens of the employee or user: You're creating a better, more aligned working environment. Are there any non-obvious impacts that you think memory will have?
I think the interesting thing about memory is, today as it's implemented, it's far worse than a human's memory. It is super limited. It's expensive in the sense that having deep memories takes up a lot of context. It lacks a lot of the nuance of our memories. Our memories are really useful both for specific facts, but more so for the abstract concepts that they create and that we learn. When you think of those memories, a lot of times that connection is where they become really powerful, and that's what becomes pattern recognition for investments.
In the near term, I think memory is a useful tool for all the obvious reasons that you can think about in terms of just adding more context that persists in the context window. Longer term, everyone knows the fallibility of human memory. This is a little bit above my pay grade in terms of the actual ML research and how memory is being incorporated in different ways.
But you could imagine any one of the architectures that are being experimented with today with memory, having short-term memory and longer-term memory, and the ability to look up memories if implemented correctly, both in terms of how the models are trained as well as the harnesses around them. Imagine having a model that has lived every single one of a firm's investments and maybe things they even passed on. There's a world where it becomes so much more powerful than any given human.
I look at some of the investors I've worked for, and they just have this insanely, eerily intuitive judgment about things. A lot of that is not just memory of specific facts, but memory that comes from experience and scar tissue. There's a world where that certainly could exist in models.
In some sense, people talk about AGI being the ability to be generally intelligent. I think a really interesting test for a model, if it's AGI, is: Can it predict the future? You need a really good world model and an understanding of lots of variables to predict the future, and that's what an investor does. Memory is clearly a really important ingredient to that.
I'd say, in the near term, I'm a little bearish on just how big of an impact it's going to be relative to simply copy-pasting that context into a window. But longer term, I think that's the unlock where this goes from being a useful tool to the real core driver of a lot of the research a firm is doing.
Well, I'll separate it. Near term, I'm almost taking away that it might be more valuable to limit the memory, therefore reducing the context window and investing more into the prompt that I'm using. I know they're not like for like, but ultimately memory is customizing to you. If you could just put that in the prompt, that might be more beneficial than sucking up memory.
It's a near-term convenience. In other words, if there's knowledge you want to impart and every single time you interact with the model, that should live in memory. Or if there are aspects of working on a name and you've maybe done 20 queries on a particular name, you may want some sense of that memory living in the model. So if you reference something back, or it already knows it did XYZ analysis, it can do that.
I think that's where it's useful. It's a bit of a shortcut for just loading in context that you don't have to repeat. I definitely would recommend spending a lot of time being thoughtful about the extent to which you can manage the system prompt in an off-the-shelf tool or anything like that. I think we're just very early innings from a technical perspective in tapping what it means to really utilize memory in a thoughtful way.
To illustrate that, these models have such incredible world knowledge by essentially reading the entire internet in their pretraining. They clearly have the capability to store a lot of knowledge in their weights and then learn important concepts from that knowledge that lead to things like reasoning. So it's just a matter of time before that same paradigm applies to continual learning on the job. To me, that's going to be really, really fun.
Yeah, I do think about if Bridgewater's really been recording all those meetings for all of this time, what all of that input could do eventually and what could happen with that technology. I don't even have an idea of what it could be, but it just feels pretty unique, seeing how people are testing this even with their own personal recordings and creating their own things.
The last thing I wanted to talk about is agentic AI. It gets brought up, it's a buzzword, and it has very practical applications. The definition gets a little bit fuzzy sometimes, but just thinking about agentic AI as a topic and as it applies to investors, could you talk about where that's going, how you would even categorize that within the various tools, and anything else you would mention on that?
7. Agentic AI Enters Research
There are lots of definitions of agentic AI. I think of it essentially as the model having the capability to reason, reflect, take action, and reflect on those actions, ultimately in pursuit of a goal. There are obviously different levels of scaffolding and constraint around what the model can and can't do.
What's really exciting, especially now, is that agents have started to work. Interestingly, when we first started Portrait, we did actually have an agent that was running using GPT-4, and it was so hard to get to work. The way you got it to work was that you had to be so prescriptive about what the model could and couldn't do. I think our system prompt was 30,000 tokens or something crazy.
You really had to limit it because, at any given time, if the model started going down the wrong direction, it didn't have the ability to self-correct and be thoughtful about what it had learned and how it should adjust its plan. That required just a ton of prompting.
What's cool today, and I think what's happening in software engineering is the leading edge of what's eventually going to happen in investment research and really any other knowledge field, is that the models have now become smart enough to do longer-running tasks by using tools and updating their thinking process as they go.
You're seeing this with things like Claude Code and Codex, where people are leaving these models alone for multiple hours, and they are dynamically operating like a software engineer to understand a codebase, make changes to it, test things out, see an error, go look at the trace and figure out why there's a problem, and make a fix. That type of thinking is certainly really challenging, and that now works.
The nice thing about code is that it's the perfect environment for these models to figure this type of capability out because, first, it's all just text files at the end of the day. It's all contained, usually within a repo, so all the context is locally available, and it's super verifiable. The code either runs or it doesn't.
Investment research is obviously much different. None of those things are true. The context is in many different forms and very broad, and some is more accessible than others. It's also much more qualitative in terms of output. But the question of whether these models can reason iteratively and arrive at complex answers that require on-the-go changes of plan—the answer is yes.
It's just a matter of the engineering work to get those models into the right context so that they can operate like a junior analyst and ultimately like a senior analyst doing long-running research. It's definitely still a buzzword, but there's a lot more meat behind that buzzword than there was maybe a year ago.
The point on code—and I often bring up chess—these are constrained environments that are still very complex, but they're constrained by certain things. You tend to see more of a clean iteration cycle and more productivity that comes out of it. The more complex and adaptive the system, the harder it is to control the evolution, so it's one that I think will be interesting to see how it all evolves.
But this has been fascinating. Thank you for going on some more philosophical tangents in addition to the very tangible use cases.
It's been a pleasure. Thank you for sharing your knowledge with us.
Thank you. It's been really fun.