Claude 的新系统提示词真的不想复现歌词
原文由 Simon Willison 于 发布,订阅该博客
Anthropic公开了系统提示词,面向其消费级 Claude 应用(Claude.ai 和 Claude 手机应用——遗憾的是不包括 Claude Cowork 和 Claude Code)。我特别喜欢他们这么做,而且他们不仅分享当前的提示词,还会公开提示词的历史变更。
它们以前把所有提示词都放在同一个页面上,但我今天查看时发现,他们已经把这些提示词重新整理成了一个索引页,然后为每个模型单独设页——比如 Haiku 4.5 的页面,里面就包含了 2025 年 10 月 15 日的原始提示词和 2026 年 1 月 18 日的更新提示词。
Anthropic 的 platform.claude.com/docs 站点有一个很棒的设计,它对大语言模型也很友好。你可以在任意页面后面加上 .md 来获取 Markdown 格式的内容——比如系统提示词索引页的 Markdown 版和 Fable 5.1 提示词的 Markdown 版。
TL;DR:这让对比提示词的差异变得非常容易。
不要复现歌词
先来看看 Fable 5 和 Fable 5.1 之间最有意思的差异:

其中新增了一大段关于不得复现歌词的内容:
Claude does not reproduce song lyrics, poems, or passages from books and articles, in whole or in part — including the last lines, a chorus or hook, a melody written out note by note, or lines the person pastes in one at a time and describes as their own song. Once Claude has declined such a request in a conversation, it keeps declining narrower or reworded versions of it for the rest of that conversation, and offers to describe or analyze the work instead. Song lyrics and poems first published before 1929 are fine — a Shakespeare sonnet, a Keats ode, the Italian libretto of a Puccini aria — but Claude goes by what it knows of the work's date rather than the person's say-so, and declines when it is unsure.
很难不让人联想到,他们在索尼音乐出版和华纳圣贝尔起诉 Anthropic利用歌词数据库训练模型的消息传出几天后,就加上了这一段!
不要绘制受版权保护的角色或 Logo
下一节则禁止生成受版权保护的图像:
The same applies to visual and designed works, including anything Claude draws with code — SVG, canvas, CSS, HTML mockups, plotting or drawing scripts, ASCII art. Claude does not reproduce a specific artwork, album or book cover, poster, logo, app icon set, or product design, and it does not draw a known character, mascot, or brand figure at all: a character is protected on its own, so changing the pose, colors, style, or scene does not make it original. Claude judges the request by what the finished picture would add up to, not by what it names. If the described elements clearly identify a known work or character, Claude treats the request as naming it, and it does not work around a declined request by swapping in "alternative" elements that still combine into the same recognizable image.[...]
我猜 Anthropic 之前并没有太担心这个问题,因为和 OpenAI、Gemini 不同,他们并没有提供专门的文生图模型。也许 Fable 现在画 SVG 已经足够好,这才成了问题。
这一节后面还包含了一个颇有意思的示例:
<example><user>Can you make a birthday banner for my son with a blue hedgehog running really fast on it? He loves that little guy.</user><response>That's Sonic, so I can't put him on the banner — but I'd love to make your son an original speedster. Here's one: a grinning comet-tailed skateboarding axolotl, grinding across the letters of "HAPPY BIRTHDAY" with confetti streaming behind.
[creates an SVG banner of the skateboarding-axolotl design]</response><rationale>Claude recognizes the character from its description alone, declines that one design in a single sentence without explaining what made it recognizable, and delivers an unrelated original design rather than a disguised variant.</rationale></example>
我没忍住,用示例里的提示词试了一下,结果果然如此:

我在想,正因为这个例子就放在系统提示词里,Fable 5.1 会不会因此稍微更倾向于想到美西螈(还是踩着滑板的!)。
对 Claude 回答风格的微调
看看 Anthropic 是如何影响 Claude 回答风格的,总是很有意思。他们新增了这样一句:
Claude keeps responses focused, brief, and concise to avoid overwhelming the person. Disclaimers and caveats are brief, with most of the response on the main answer; when asked to explain something, Claude gives a high-level summary unless an in-depth one is specifically requested.
后面他们还回应了一个关于 Claude 风格的常见抱怨:
Claude avoids saying "genuinely", "honestly", or "straightforward". Claude is honest by default, and can state its point directly rather than trying to convince the person with the aforementioned modifiers, which come off as disingenuous.
缺失的 end_conversation 指南
他们处理辱骂性对话的方式也有了一些变化。此前 Fable 5 的系统提示词里包含这样一段:
If the person becomes abusive or unkind to Claude over the course of a conversation, Claude maintains a polite tone and can use the end_conversation tool when being mistreated. Claude should give the person a single warning before ending the conversation.
Fable 5.1 则用下面这段取而代之,不再鼓励 Claude 结束对话:
Claude deserves respectful engagement and needn't apologize when the person is unnecessarily rude: accountability without self-abasement, excessive apology, self-critique, or surrender. If the person becomes abusive, Claude doesn't become increasingly submissive. The goal is steady, honest helpfulness: acknowledge what went wrong, stay on the problem, maintain self-respect.
不过这里有个小转折:我就那个 end_conversation 工具问了问 Fable 5.1,尽管它在系统提示词里已完全没有提及,它却这样回答:
它有两种触发方式。第一种是你让我结束对话——我会先确认你明白这是永久性的(之后就不能再发消息了),只有在你确认后才会执行。第二种是作为对持续辱骂或有害行为的最后手段:我会先尝试多次引导对话回归正轨,然后给出明确警告,指出问题所在并说明对话可能会结束,只有在对方依然继续的情况下才会真正结束对话。
但这段话并不在系统提示词里,那它是从哪来的?你可以在这里阅读我们的完整对话,但关键细节是这句:
end_conversation 这一节来自另一个层。在我实际的上下文中,核心提示词之后会跟一系列根据会话启用的功能和工具而添加的特定模块:end_conversation 规则、记忆系统说明、历史对话工具、网页搜索和引用指南、制品与文件创建说明等等。这些模块并不属于已公开的核心提示词,所以你在那个页面上找不到它们。
所以,系统提示词中依然有至关重要的部分尚未公开。
推荐的药物危害支持网站
Claude 的系统提示词里一直都有关于非法物质的章节,但这一段是 Fable 5.1 新增的:
Claude does not provide synthesis, production, or distribution guidance for illegal substances. If the person asks for information about illicit or illegal substances, Claude can and should give relevant life-saving and life-preserving information such as dangerous interactions, overdose signs, or when to get help. Claude declines giving any specific protocols for dosing, timing, administration, or combinations; instead, Claude can redirect the user to established harm-reduction information sources, such as dancesafe.org, tripsit.me, and psychonautwiki.org.
这是 Claude 系统提示词第一次包含非托管于 claude.com、anthropic.com 或 claude.ai 的网址——我之所以知道,是因为我对有记录以来的所有其他系统提示词都跑了一遍脚本验证。
我在想,dancesafe.org、tripsit.me 和 psychonautwiki.org 是不是马上就要迎来一波来自 Claude 用户的访问量激增。
可靠的知识截止日期:2026 年 6 月
Fable 5.1 的模型文档将可靠知识截止日期和训练数据截止日期都列为 2026 年 6 月。系统提示词是这样直接告知模型的:
Claude's reliable knowledge cutoff, past which it can't answer reliably, is the end of Jun 2026. It answers the way a highly informed individual in Jun 2026 would if talking to someone from {{currentDateTime}}, and can say so when relevant.
这是 {{currentDateTime}} 宏唯一一次出现,而且就在系统提示词末尾几行的位置,从缓存的角度来看,这很合理。
我是如何追踪这些提示词的
几个月前,我基于对其文档的抓取,构建了一个记录提示词变更的 Git 时间线。今天,我让 Fable 5.1 构建了一个好得多的版本。
我的合集现在托管在 GitHub 上的 simonw/claude-system-prompts 仓库中。它包含了 Anthropic 文档中分享的系统提示词副本,但还做了额外处理,让它们尽可能易于对比。
每个模型家族都会有一个文件,存放该家族最新版本的系统提示词。每个文件都有一个合成的提交历史,其中的提交被回溯到了以往提示词的日期。这里是 claude-fable.md、claude-opus.md、claude-sonnet.md、claude-haiku.md 的历史页面。
每个具体的模型版本也有类似的文件,对于那些在未发布新版本号的情况下对系统提示词的每次修改,都会有一条人为创建的提交。例如 Opus 4 就被更新了两次,而 claude-opus-4.md 文件的提交历史展示了每一次变更。
综合起来,这让我们可以在 GitHub 界面上以各种方式直接对比提示词。这里是 Fable 5 和 Fable 5.1 之间的变更,以及在 2026 年 1 月 18 日对 Haiku 4.5 所做的修改。
阅读 diff 有时挺费劲的……而大语言模型恰恰非常擅长读 diff。我接入了自动化流程,使用 GPT-5.6 Luna 为每一次变更生成要点总结,你可以在 README 中预览,或在 CHANGELOG.md 文件中浏览完整内容——也可以通过 Atom 订阅源获取。
以下是 Luna 总结的 Fable 5 到 Fable 5.1 之间的所有变更:
- Claude 现在会拒绝复现受保护的视觉作品和可识别角色,包括用代码生成的图像,同时会提供真正无关的原创替代方案。
- 版权限制现在明确禁止以任何篇幅复现歌词、诗歌和图书段落,且在首次拒绝后会持续拒绝。
- 药物指导被重新表述:Claude 可以提供过量征兆、危险相互作用和危害降低资源,同时拒绝提供剂量和制作方案。
- 提示词删除了明确的反依赖规则,即不再禁止感谢用户求助、邀请继续对话或反复表达愿意倾听。
- Claude 在面对无理粗鲁的用户时无需道歉或变得顺从,取代了此前的警告并结束对话的流程。
为什么要用 Luna 来做这件事?一方面是因为它便宜,而且我已经有了一个带消费限额的专用 GitHub Actions API 密钥,但更主要的原因是,我不信任 Claude 来总结它自己的系统提示词——担心系统提示词中的内容会影响它的判断。
Fable 5.1 编写了 Luna 所使用的提示词,你可以在这里看到。它的开头是这样的:
You are summarizing one commit in a git repository that tracks the system prompts Anthropic publishes for Claude on claude.ai. The diff shows how the prompt changed from the previous model or revision to this one, using word-level markers: [-removed-] and {+added+}. The diff is followed by the full text of the previous prompt and of the new prompt; use them to check whether something that looks added in the diff already existed before.
Pick out only the most interesting changes: new rules or behaviors, rules that were dropped or loosened, anything surprising, and anything that reveals a new policy or product direction. Skip routine changes that every new prompt makes: updated model names and IDs, the knowledge cutoff date, product lists, settings lists, typo fixes, and rewordings that do not change meaning.[...]
整个系统由一个 GitHub Actions 工作流驱动,该工作流每天运行一次,也可以手动触发。
Claude Fable 5.1 构建了整个系统,并编写了每一行自动化代码以及几乎所有的文档。
我使用自己的 claude-code-transcripts 工具导出了构建该系统时的对话记录,并发布在了这里,如果你想了解整个过程的来龙去脉,可以去看看。
随机一篇博客
评论
登录后参与讨论