2026-09-06
每日一文 · 长文精读
Presentation: A Few Predicted Talks From QConAI 2030
QConAI 2030 精选预测演讲
作者:Meryem Arik · InfoQ 原文
摘要:Meryem Arik 在 QConAI 2030 上大胆预测未来,讨论 AI 领域的 token 支出、基础设施、智能体及监管趋势。她指出 token 成本持续下降,但总支出因 Jevons 悖论飙升,预计到 2030 年 agent 的 token 用量将增长 24 倍。她还驳斥了 AI 实验室补贴 token 的说法,认为其推理业务已高度盈利,并强调开源模型对成本降低的推动作用。
Meryem Arik: I'm going to be doing something very risky today, and I'm going to be predicting the future.
Meryem Arik:今天我要做一件非常冒险的事——预测未来。
You should famously never, ever make predictions, especially about the future, because it always goes wrong.
众所周知,你永远不该做预测,尤其是关于未来的,因为总会出错。
A very wise man said this.
一位非常睿智的人说过这话。
He actually was a baseball player in Boston.
他其实是波士顿的一名棒球运动员。
Luckily, I'm not a wise man.
幸运的是,我不是聪明人。
I'm Meryem. I'm the cofounder of Doubleword. I work on high throughput inference.
我是 Meryem,Doubleword 的联合创始人,从事高吞吐量推理工作。
The reason why I felt somewhat confident to make vague predictions is I actually have a fairly good track record of it.
我之所以有信心做出模糊预测,是因为我在这方面的记录相当不错。
A couple months after ChatGPT came out, I did a TEDx talk called, AI's Looming Hardware Crisis, where I argued that inference would be very important and that we would not have enough GPUs to be able to serve all the inference that we needed.
ChatGPT 发布几个月后,我做了一场 TEDx 演讲,题为《AI 迫在眉睫的硬件危机》,其中指出推理将非常重要,且我们没有足够的 GPU 来满足所有推理需求。
This was three years ago, and in that time, NVIDIA's stock price has gone up 10x.
那是三年前,其间 NVIDIA 的股价上涨了 10 倍。
It turns out I was right.
事实证明我是对的。
I was right once, I might not be right again.
我对过一次,也可能不再正确。
I'm going to give you a brief overview of the structure of this talk.
我来简要介绍一下这次演讲的结构。
I'm going to present my worldview, which I think is a fairly coherent worldview.
我将阐述我的世界观,我认为这是一个相当连贯的世界观。
I'm going to make a bunch of predictions, and they hopefully shouldn't contradict each other too drastically.
我会做出一系列预测,希望它们彼此之间不会过于矛盾。
You can think of the sections of my talk as standalone mini-talks.
你可以把我的演讲各部分看作独立的小演讲。
The reason I structured it like this is so I can clip it up in five years' time if one of them is right and none of the others are right, and I can claim that I predicted the future really well.
我这样组织的原因是,如果五年后其中一个是正确的而其他都不对,我可以剪裁出那一段,声称自己预测未来非常准确。
It's aimed at you, so specifically, I'm going to be talking about the types of things that I think enterprise early adopters and early majority will be caring about in 2030.
它针对的是你们,具体来说,我将讨论我认为企业早期采用者和早期大众在 2030 年将会关心的那些事情。
Finally, you might not agree with all of my predictions, and that's totally fine.
最后,你可能不会同意我的所有预测,这完全没问题。
I tend to think of myself as a semi-logical person, and so hopefully the reasoning should follow.
我倾向于认为自己是一个半逻辑性的人,所以希望推理是连贯的。
If you disagree with my conclusions, I'd recommend that you think about what assumptions that I've made that you disagree with, and then you can follow the logic of where you go from there.
如果你不同意我的结论,我建议你思考一下我做出的哪些假设你不认同,然后顺着逻辑推导出你自己的结论。
As I mentioned, this is aimed at the InfoQ audience.
正如我提到的,这针对的是 InfoQ 的观众。
I'm only going to be talking about things that I think are going to matter to the early adopters and early majority.
我只谈我认为对早期采用者和早期大众重要的事情。
For the laggards, the stuff that we talked about at this conference probably still won't be relevant even then.
对于落后者,我们在这次会议上讨论的内容即便到了那时可能仍然不相关。
Also, I'll make a note on the likely direction of error.
另外,我会指出可能出现的误差方向。
I'm self-aware enough of myself to know where I tend to overestimate and underestimate things.
我有足够的自知之明,知道自己在哪些地方倾向于高估或低估。
I tend to overestimate how quickly enterprises will change, which I think a lot of people do.
我倾向于高估企业变革的速度,我觉得很多人都是这样。
I tend to underestimate how quickly the frontier has improved.
我倾向于低估前沿技术提升的速度。
This didn't used to be a problem, but actually the frontier has been improving so quickly that it's possible that we're actually way further than I'm predicting.
这以前不是问题,但事实上前沿进展如此之快,以至于我们可能比我的预测走得更远。
I'm pretty good at spotting bottlenecks.
我很擅长发现瓶颈。
We decided that inference would be a bottleneck in 2021, and that turned out to be pretty good.
我们在 2021 年判断推理会成为瓶颈,结果证明这个判断相当准确。
I would also caveat that if we achieve AGI, all of the bets are off, and then I'll be doing my power washing business or something like that.
我还要补充说明,如果实现了 AGI,所有赌注都不作数了,到时候我可能就去经营我的高压清洗生意了。
Prediction Framework
预测框架
What am I going to be making predictions about?
我将对哪些方面做出预测?
I've tried to pick a range of topics.
我尝试选择了一系列话题。
I'm going to talk about money and how much money we are spending on AI.
我会谈到钱,以及我们在 AI 上花了多少钱。
I'm going to be talking about what happens when we have more code and more builders within our organizations.
我会讨论当组织内部有更多代码和更多构建者时会发生什么。
I'm going to be talking about infrastructure, what's the kind of infrastructure we're going to be building on top of.
我会讨论基础设施,我们将在什么样的基础设施之上构建。
I'm going to be talking about agents, obviously.
显然,我会讨论智能体。
I'll talk also about regulation and how, if at all, it's going to impact people here.
我还会谈到监管,以及它(如果有的话)将如何影响这里的人。
Then, finally, I'm going to talk a bit about what I think careers look like for AI engineers and software engineers.
最后,我会简单谈谈我认为 AI 工程师和软件工程师的职业前景。
How We Manage and Attribute Token Spend
我们如何管理和分配 Token 支出
Let's start with money.
让我们从钱开始。
How much do I think we're going to be spending in 2030?
我认为我们在 2030 年将花费多少?
Do I think it's going to be a topic of conversation?
我认为它会成为讨论话题吗?
I've been in the inference space since 2022, so a relatively long time.
我从 2022 年就进入推理领域,所以时间相对较长。
I've always thought that inference cost was going to become a problem.
我一直认为推理成本会成为问题。
We would talk about it with customers, that inference cost is a really big issue.
我们和客户讨论过,推理成本确实是个大问题。
Actually, it wasn't really a big issue until about three months ago, when it seems to be a huge issue.
实际上,直到大约三个月前它才成为大问题,现在似乎成了巨大问题。
I'll talk about why it's become an issue in that very short period of time.
我会讨论为什么在这么短的时间内它成了问题。
It was released via Twitter that some unnamed company blew through half a billion dollars' worth of Claude credits in one month.
有消息通过 Twitter 透露,某家未具名公司在一个月内用掉了价值 5 亿美元的 Claude 积分。
Uber has burned through their entire token budget by the end of April.
Uber 在四月底之前就用光了整个 token 预算。
J.P. Morgan is spending $2 billion more on IT this year than it did last year.
摩根大通今年在 IT 上的支出比去年多了 20 亿美元。
Token spending is becoming a very big issue now.
Token 支出现在正成为一个非常大的问题。
We've seen talks in the conference about this particular issue.
我们已经在会议上看到了关于这个具体问题的讨论。
It's an issue, and we are still so early.
这是一个问题,而我们仍然处于早期阶段。
This is a source from Deloitte, 23% of companies say they use agentic AI at least moderately.
这是德勤提供的数据,23% 的公司表示他们至少中度使用智能体 AI。
When we think about what moderately means, it typically means that some of the organizations are using it on some of their use cases.
当我们思考“中度”意味着什么,它通常表示一些组织在其部分用例中使用。
We are still so early, and token spend is already an issue.
我们还处于早期,token 支出就已经是个问题了。
There's a myth that I want to dispel.
有一个我想澄清的迷思。
Actually, we heard this as well in Jordan's talk.
实际上,我们在 Jordan 的演讲中也听到了这一点。
There's this myth that I've actually heard from a lot of the lunchtime conversations that AI labs are subsidizing tokens.
有一个我经常从午餐对话中听到的迷思,说 AI 实验室在补贴 token。
That they're eventually going to jack up the price of these tokens, because they're subsidizing it.
它们最终会大幅提高这些 token 的价格,因为它们在补贴。
That's not true.
这不是真的。
Their inference part of their businesses are highly profitable.
它们业务中的推理部分利润很高。
Yes, these organizations do lose money, although I think actually Anthropic now makes money.
是的,这些组织确实亏损,虽然我认为实际上 Anthropic 现在已经开始盈利。
The reason they are not as highly profitable on their bottom line is because they are spending so much money on infrastructure and so much money on training.
它们的净利润不够高的原因是它们在基础设施和训练上投入了大量资金。
I'm going to firstly dispel this myth that I don't think we're going to see AI labs jacking up token prices for their existing models because they're subsidizing it.
我首先要澄清这个迷思:我认为我们不会看到 AI 实验室因为补贴而提高现有模型的 token 价格。
In fact, I think tokens are getting cheaper.
事实上,我认为 token 正在变得更便宜。
Empirically, we've seen that tokens have gotten much cheaper.
根据经验,我们看到 token 已经便宜了很多。
They are going to continue to.
它们还会继续便宜下去。
If you take a fixed intelligence point, so let's say a fixed IQ point, that fixed IQ point over the last three years has gone down and cost roughly 10x every single year.
如果你取一个固定的智能水平,比如说固定的 IQ 分值,过去三年里这个固定 IQ 分值的成本每年大约下降 10 倍。
If you were using GPT-4 in 2023, that price has gone down 10x year over year in that time.
如果你在 2023 年使用 GPT-4,其价格在那段时间里每年下降 10 倍。
That is largely driven thanks to open-source model pressures.
这很大程度上是开源模型压力推动的。
The open-source model providers have gotten very good at taking an intelligence level provided by the frontier labs and distilling them and building more efficient architectures to deliver them.
开源模型提供商非常擅长获取前沿实验室提供的智能水平,进行蒸馏,并构建更高效的架构来交付。
Tokens for a fixed intelligence level are going to continue to get cheaper.
固定智能水平的 token 将继续变得更便宜。
We'll also see a rise of open-source models.
我们还将看到开源模型的兴起。
This is taken from OpenRouter, which is a place where you can get access to a lot of different models.
这张图来自 OpenRouter,你可以在那里访问许多不同的模型。
The share of people using open models versus closed is growing significantly, in part because of cost pressures.
使用开源模型与闭源模型的人数比例正在显著增长,部分原因是成本压力。
Despite that, spend is still skyrocketing, even with cheaper tokens.
尽管如此,即使 token 更便宜,支出仍在飙升。
This is the very classic Jevons paradox.
这正是经典的杰文斯悖论。
I think we've had a couple talks mention this exactly.
我想我们已经有几个演讲提到了这一点。
Token prices are plummeting, but we're finding more and more use cases to apply them to.
Token 价格在暴跌,但我们发现了越来越多的应用场景。
Spend is continuing to rise.
支出持续上升。
This is taken from Goldman Sachs.
这张图来自高盛。
I actually think it's a fairly conservative estimate, but they expect that token usage by agents will multiply 24 times by 2030.
我实际上认为这是一个相当保守的估计,但他们预计到 2030 年,智能体的 token 使用量将增长 24 倍。
I actually think it'll be much more than this.
我实际上认为会远高于此。
The percentage of headcount budget, which I think is like a way to think about it, that is spent on tokens will go up significantly and will be a very big pain point in 2030.
人员预算中用于 token 的百分比(我认为这是一种思考方式)将大幅上升,并在 2030 年成为一个非常大的痛点。
We're going to have more use cases that are unlocked.
我们将解锁更多用例。
As tokens get cheaper, we find more things to apply them to.
随着 token 变得更便宜,我们找到更多应用它们的地方。
We also are seeing more and more token-hungry use cases.
我们也在看到越来越多消耗 token 的用例。
Unfortunately, or fortunately, every single generation of AI use cases that have come out, so starting with simple chatbots, then moving to RAG, then moving to reasoning models, then moving to agents, and then increasingly complex agents, every single generation is getting more and more token-hungry, and we should expect that to continue.
不幸或幸运的是,每一代出现的 AI 用例——从简单的聊天机器人,到 RAG,再到推理模型,再到智能体,以及越来越复杂的智能体——每一代都越来越消耗 token,我们应该预期这种情况会继续。
阅读理解
1. According to Meryem Arik, what is the main reason token prices for a fixed intelligence level have dropped significantly over the past three years?
2. What does Meryem Arik say about the Jevons paradox in the context of AI token usage?
3. What does Meryem Arik predict about token usage by agents by 2030, according to Goldman Sachs?