2026-08-27
每日一文 · 长文精读

Presentation: Can Claude Fix Itself? Using LLMs for Incident Response

演示:克劳德能自我修复吗?利用大语言模型进行事件响应

作者:Alex Palcuie · InfoQ 原文

摘要:Anthropic 的 AI 可靠性团队成员 Alex Palcuie 分享了他在事件响应中使用 LLM 的亲身经历。他坦言,尽管 Claude 无法完全取代人工值班,但在某些环节能提供帮助。他批评了 AI SRE 领域的过度炒作,指出真实事件远比基准测试复杂。同时,他提出一个愿景:如果 LLM 能处理常规修复,人类就能专注于预防性工作,从而提升系统可靠性。

Alex Palcuie: I'm Alex. I'm on the AI reliability team at Anthropic, which means my job is keeping Claude up.
我是 Alex,在 Anthropic 的 AI 可靠性团队工作,我的职责就是确保 Claude 正常运行。
I've been using LLMs as part of actual incident response, and I'd like to share a candid and humorous discussion about what works and what doesn't.
我一直在将 LLM 用于实际的事件响应,想坦诚且幽默地聊聊哪些方法有效、哪些无效。
I'm one of the two people that joined the reliability team in London.
我是加入伦敦可靠性团队的两人之一。
I did on-call for Claude's serving stack, Solo, for my first three months, which is a very effective and fast way, but really stressful, to learn about the system.
头三个月我负责 Claude 服务栈 Solo 的值班,这是了解系统非常高效快速但也压力巨大的方式。
Then I onboarded everyone else so I could stop being on-call forever.
之后我培训了其他人,这样我就永远不用值班了。
Before that, I was at Google, SRE on Google Cloud's compute product, GCE.
在此之前,我在 Google 担任 Google Cloud 计算产品 GCE 的 SRE。
I was on the SRE for SRE team, which is the escalation layer for when an outage is bad enough that even normal SREs want to back up.
我曾在 SRE 的 SRE 团队工作,那是当故障严重到普通 SRE 都想退缩时的升级层。
I've been carrying pages for a while.
我处理告警已经有一段时间了。
I do have opinions of what makes an incident response process good and what makes one bad.
对于什么造就好的事件响应流程、什么导致差的流程,我确实有自己的看法。
While naturally skeptical in the beginning, since about January this year, I've started doing something that feels slightly transgressive to admit, which is, I start reaching out for Claude before I reach out to my monitoring dashboards.
虽然一开始自然持怀疑态度,但从今年一月左右开始,我开始做一些承认起来有点越界的事情——我在查看监控仪表盘之前,会先向 Claude 求助。
Then you're asking yourselves the one question, and I get asked all the time at dinner parties, someone finds me at this conference hallway, my friends at my old company where their VPs are telling them to do something with AI, does Claude really fix your incidents?
那么你们都在问同一个问题,我在晚宴上总被问到,有人在会议走廊找到我,我在老公司的朋友们,他们的副总裁让他们用 AI 做点什么——Claude 真的能修复你们的事件吗?
Is the on-call rotation just Claude right now?
现在值班轮换就是 Claude 自己吗?
Have you basically automated yourself out of the job, just like the original SRE book wrote more than 10 years ago?
你们是不是基本上把自己自动化出局了,就像十多年前那本 SRE 原版书写的那样?
I understand why people ask this.
我理解人们为什么这么问。
If anyone has this working, it's the company that makes the model.
如果有人能做到这一点,那一定是制造这个模型的公司。
We have unlimited tokens.
我们有无限的 token。
Researchers sit a desk away from me.
研究人员就坐在我隔壁桌。
I get involved in training the models.
我参与模型的训练。
Surely, if it's anyone, it's us.
当然,如果有人能做到,那就是我们。
Let me get the answer out of the way.
让我先把答案说了吧。
It's a no.
答案是否定的。
I want to sit with the no for a second.
我想先停留在这个“不”上。
It would be genuinely hypocritical for me to stand up here and tell you that Claude fixes everything.
如果我站在这里告诉你 Claude 能修复一切,那真是虚伪至极。
My team just shortly will have its one-year anniversary.
我的团队即将迎来成立一周年。
If an LLM could carry a pager, we might not need to hire so much.
如果 LLM 能携带寻呼机,我们可能就不需要招这么多人了。
The fact that my team exists, the fact that we're hiring for many positions, in London, in Dublin, and the U.S., and we have staff positions, this should show to you that, no, it doesn't work.
我的团队存在的事实,我们在伦敦、都柏林和美国招聘许多职位的事实,以及我们还有员工岗位的事实,应该向你表明:不,它行不通。
However, there's an asterisk over there.
不过,这里有个星号。
The asterisk is related to the timelines.
这个星号与时间线有关。
I think many of us would not be surprised if somewhere in the future we would be able to do such a thing.
我想我们很多人不会惊讶于未来某个时候我们能够做到这样的事情。
Today, I also point out the useful ways in which Claude helps me during my on-call.
今天,我还要指出 Claude 在我值班时帮助我的有用方式。
Going back, Claude is down more often than any of us would like.
回想起来,Claude 宕机的频率比我们任何人希望的都要高。
Earlier, I was involved in an incident, even if I'm at a conference.
早些时候,我参与了一个事件,即使我当时在开会。
You may have noticed, and some of you might have tweeted, and I have the same sentiment when Claude goes down.
你可能注意到了,有些人可能发了推文,当 Claude 宕机时我也有同样的感受。
Full of opinions, people are tagging me.
人们充满意见,都在 @ 我。
Please continue doing these.
请继续这样做。
These are the positive tweets.
这些是积极的推文。
I also got the mean ones, but that's social media for you nowadays.
我也收到了恶意的,但这就是如今的社交媒体。
There's even more.
还有更多。
I know at least 10 companies right now whose entire pitch is some version of AI SRE.
我知道目前至少有 10 家公司,它们的整个宣传就是某种版本的 AI SRE。
Since I'm also pretty open about angel investing, and somehow people think I would be less skeptical about them, and I'm more skeptical about them.
由于我对天使投资也很开放,不知为何人们认为我对它们不会那么怀疑,但实际上我更怀疑它们。
There are now benchmarks.
现在有基准测试了。
There are curated datasets.
有精心整理的数据集。
There are historical incidents.
有历史事件。
There's the game of how much percentage can your app solve of these incidents.
有那种游戏,看你的应用能解决这些事件的百分之多少。
There are papers.
有论文。
There's like the whole on-call industry.
还有整个值班行业。
The cynical take that we could have in a room full of senior engineers who've seen hype cycles before, is, of course, this will be it.
在一个充满见过炒作周期的资深工程师的房间里,我们可能会有的愤世嫉俗的看法是:当然,就是这样。
There's VC money sloshing around.
有风险投资在四处流动。
Someone was always going to slap AI on ops, say that the TAM is billions, that it will swallow the observability industry, and our above-average compensation.
总会有人把 AI 贴到运维上,说 TAM 有数十亿,它会吞噬可观测性行业和我们高于平均的薪酬。
They go, they raise a Series A from investors.
他们就去从投资者那里融 A 轮。
Those investors are building a diversified portfolio.
那些投资者正在构建多元化投资组合。
There's AI law, there's AI healthcare, and there's AIOps.
有 AI 法律、AI 医疗保健,还有 AIOps。
Yes, it's a lot.
是的,很多。
This is not an exhaustive list.
这不是一个详尽的列表。
By the time I was writing the slides, at least two other companies appeared.
在我写幻灯片的时候,至少又出现了两家公司。
The cynicism is warranted.
这种愤世嫉俗是有道理的。
Some of the benchmarks are like, can the model solve these pre-packaged problems into a clean prompt?
有些基准测试就像:模型能否解决这些预先打包好的、放入一个干净提示中的问题?
This is not what something at 3 a.m. when you get paged looks like.
这不是你在凌晨 3 点收到告警时遇到的情况。
Actual incidents are not well-formatted problems.
实际的事件不是格式良好的问题。
I'm not cynical about the goal.
我对目标并不愤世嫉俗。
Genuinely, unironically, I'm cheering for these people.
真诚地、毫不讽刺地说,我在为这些人加油。
I want them to practice.
我希望他们实践。
It's not for the reasons that people typically assume.
原因并非人们通常假设的那样。
The reason is that on-call is a tax we levy on humans because our systems are not good enough to look after themselves.
原因是值班是我们向人类征收的一种税,因为我们的系统还不够好,无法自我照料。
How many people here have been on-call?
在座有多少人值过班?
You know the physical reality of it.
你知道它的物理现实。
Your phone buzzes, there's a half a second where you go from asleep to probably incident commander mode.
你的手机震动,半秒钟内你就从睡眠状态进入可能的事件指挥官模式。
It's not that any individual page is bad.
并不是说某一次告警很糟糕。
I like a good incident.
我喜欢一个好的事件。
Every other day, it's the cumulative weight of this problem.
每隔一天,就是这个问题的累积重量。
It's never quite relaxing.
它从来都不太放松。
By on-call standards, I'm one of the lucky ones.
按照值班标准,我是幸运儿之一。
I have followed the sun rotation.
我遵循的是日不落轮换。
London hands off to the West Coast at 6 p.m., so I can go to the pub.
伦敦在下午 6 点交接给西海岸,所以我可以去酒吧。
The worst case for me is I'm up at 6 a.m.
对我来说最坏的情况是早上 6 点起床。
Like some of you, you're up at 3 a.m. or 1 a.m.
像你们中的一些人,凌晨 3 点或凌晨 1 点起床。
You're on a 24/7 on-call rotation.
你处于 24/7 的值班轮换中。
There's probably four people.
大概有四个人。
You get paged at 3 a.m.
你凌晨 3 点收到告警。
You go to sleep, get paged at 4.30 a.m. because the other database is broken.
你刚睡着,凌晨 4 点 30 分又收到告警,因为另一个数据库坏了。
Then at 9 a.m., you show up at work and you have to be at your stand-up and look professional and presentable.
然后早上 9 点,你出现在工作岗位上,必须参加站会,看起来专业且得体。
I've been there in my career.
我在职业生涯中经历过这些。
Also, I want to be clear, this should not be a badge of honor.
另外,我想说清楚,这不应该是荣誉的徽章。
Our industry sometimes treats it like this, like we call them war stories.
我们的行业有时会这样对待它,比如我们称之为战争故事。
In panic rooms, we use combat metaphors.
在应急指挥室,我们使用战斗隐喻。
It is a cost in sleep, in our attention, in our relationships with the people.
这是睡眠、注意力以及我们与他人关系的代价。
When someone tells me, we're building AI SRE, my first reaction could be, that'll never work.
当有人告诉我,我们在构建 AI SRE,我的第一反应可能是:那永远行不通。
I'm like, I hope you get it right.
我就像,我希望你们能成功。
There's a second reason, which is more cold-blooded.
还有第二个原因,更加冷酷。
Some of the best early career advice that I got from someone who was more senior than me was that you don't get promoted for fixing incidents.
我从一位比我资深的人那里得到的一些最好的早期职业建议是:你不会因为修复事件而获得晋升。
I remember thinking that was wrong the first time I heard it, because it feels like it should.
我记得第一次听到时觉得那是错的,因为它感觉上应该是这样的。
You're the hero.
你是英雄。
You got paged.
你收到了告警。
You figured it out.
你搞定了。
You fixed it.
你修复了它。
The service came back up and people said thank you on the Slack channel.
服务恢复了,人们在 Slack 频道上说谢谢。
It does count at first when you're a junior.
当你还是初级工程师时,这确实算数。
Like, can I fix a production system on fire?
比如,我能修复一个着火的线上系统吗?
That's a real signal that you're on your way to be a good engineer.
这是一个真实的信号,表明你正在成为一名优秀工程师的路上。
It stops counting quite quickly.
但它很快就变得不算数了。
By the time you're a senior, of course you can fix this.
当你成为资深工程师时,你当然能修复这个。
That's table stakes.
那是基本要求。
It's in your job description.
它就在你的职位描述里。
What actually moves the needle in big tech parlance, what gets you promoted, what makes you the tech lead, is when people are like, they're operating at a different level.
用大科技公司的行话来说,真正能改变局面、让你获得晋升、让你成为技术负责人的,是当人们觉得,他们在以不同的层次运作。
It's building the thing that makes the whole class of incidents just not happen.
是构建那个让整类事件根本不会发生的东西。
Not fixing the one fire, but making it structurally so that those fires will never happen again.
不是修复一场火灾,而是从结构上确保那些火灾永远不会再发生。
Here's where the AI on-call dream happens for me, if, and big if, an LLM can handle the generic mitigations, the obvious rollbacks, the have you tried, the standard remediation stuff.
对我来说,AI 值班的梦想就在这里实现,如果——一个大大的如果——LLM 能够处理通用的缓解措施、明显的回滚、你试过没有之类的标准修复工作。
Then, us, the humans, can go and work on the prevention part, the platform part, the stuff that scales.
那么,我们人类就可以去从事预防部分、平台部分、那些可以扩展的工作。
That's a dream, and I want it to work.
那是一个梦想,我希望它能实现。
Here's the framework that I use to explain incident management.
这是我用来解释事件管理的框架。
I'll be using it across, for a good chunk of the presentation.
我会在演讲的大部分内容中使用它。
It's from a U.S. Air Force colonel that used to train fighter pilots.
它来自一位曾训练战斗机飞行员的美国空军上校。
He was trying to explain why some pilots consistently win against other pilots, even though they're flying less better aircraft.
他试图解释为什么有些飞行员能持续战胜其他飞行员,即使他们驾驶的飞机不那么好。
The answer was, it wasn't the aircraft.
答案是,不在于飞机。
The winner was the person who would cycle the most through this loop of observe, what's happening?
赢家是那个最能循环这个循环的人:观察,发生了什么?
Orient, what does it mean?
定向,这意味着什么?
What's my mental model?
我的心智模型是什么?
Decide, what am I going to do about it?
决定,我打算怎么做?
Act, taking the decision.
行动,执行决定。
Then, again, taking the fe
然后,再次,获取反

阅读理解

1. According to Alex, what is the main reason he starts reaching out to Claude before checking his monitoring dashboards?

2. Why does Alex say that 'you don't get promoted for fixing incidents'?

3. What is the OODA Loop framework used for in the context of incident management?

温故复习 →每日一句 →