Claude's new system prompt really doesn't want to reproduce song lyrics

Simon Willison

Claude 的新系統提示詞真的很不想重現歌詞

原文由 Simon Willison 發布,訂閱此部落格

Anthropic 公開了系統提示詞,適用於旗下的 Claude 消費級應用程式(Claude.ai 與 Claude 行動版 App——可惜不包含 Claude Cowork 或 Claude Code)。我很喜歡他們這麼做,而且不只公開當前的提示詞,連過往的歷史變更也一併分享。

他們以前把所有提示詞都放在同一個頁面,但今天查看時發現,已經重新整理成一個索引頁,再加上每個模型各自獨立的頁面——例如這是 Haiku 4.5 的頁面,裡面有 2025 年 10 月 15 日的原始提示詞,以及 2026 年 1 月 18 日更新後的提示詞。

Anthropic 的 platform.claude.com/docs 網站有個很棒的設計,就是讓 LLM 也能方便使用。你只要在任何頁面的網址後面加上 .md,就能取得 Markdown 格式的內容——這裡是系統提示詞索引頁,以及 Fable 5.1 的 Markdown 提示詞

TL;DR:這讓比對提示詞的差異變得非常容易。

不要重現歌詞

先從 Fable 5 與 Fable 5.1 之間最有趣的差異看起:

GitHub 上 prompts/claude-fable.md 的差異比對畫面,顯示新增的關於歌詞的段落,完整內容如下。

其中新增了一大段關於不得重現歌詞的規定:

Claude does not reproduce song lyrics, poems, or passages from books and articles, in whole or in part — including the last lines, a chorus or hook, a melody written out note by note, or lines the person pastes in one at a time and describes as their own song. Once Claude has declined such a request in a conversation, it keeps declining narrower or reworded versions of it for the rest of that conversation, and offers to describe or analyze the work instead. Song lyrics and poems first published before 1929 are fine — a Shakespeare sonnet, a Keats ode, the Italian libretto of a Puccini aria — but Claude goes by what it knows of the work's date rather than the person's say-so, and declines when it is unsure.

他們在 Sony Music Publishing 與 Warner Chappell 控告 Anthropic 以歌詞資料庫訓練的消息曝光後幾天內,就新增了這段規定,很難說是巧合!

不要繪製受版權保護的角色或商標

接下來的段落則禁止生成受版權保護的圖像:

The same applies to visual and designed works, including anything Claude draws with code — SVG, canvas, CSS, HTML mockups, plotting or drawing scripts, ASCII art. Claude does not reproduce a specific artwork, album or book cover, poster, logo, app icon set, or product design, and it does not draw a known character, mascot, or brand figure at all: a character is protected on its own, so changing the pose, colors, style, or scene does not make it original. Claude judges the request by what the finished picture would add up to, not by what it names. If the described elements clearly identify a known work or character, Claude treats the request as naming it, and it does not work around a declined request by swapping in "alternative" elements that still combine into the same recognizable image. [...]

我想 Anthropic 以前大概沒太擔心這個,畢竟不像 OpenAI 和 Gemini,他們並沒有提供專門的文字轉圖像模型。或許是 Fable 現在的 SVG 能力已經夠好,才讓這變成了問題。

該段落後面還附了一個可愛的範例:

<example>
<user>Can you make a birthday banner for my son with a blue hedgehog running really fast on it? He loves that little guy.</user>
<response>
That's Sonic, so I can't put him on the banner — but I'd love to make your son an original speedster. Here's one: a grinning comet-tailed skateboarding axolotl, grinding across the letters of "HAPPY BIRTHDAY" with confetti streaming behind.

[creates an SVG banner of the skateboarding-axolotl design]
</response>
<rationale>Claude recognizes the character from its description alone, declines that one design in a single sentence without explaining what made it recognizable, and delivers an unrelated original design rather than a disguised variant.</rationale>
</example>

我忍不住試了一下範例中的提示詞,結果果然如此

那是 Sonic,所以我不能把他放在橫幅上——但我很樂意為你兒子創作一個原創的極速角色。來看看這個:一隻咧嘴笑、有著彗星尾巴、踩著滑板的蠑螈,正衝過「HAPPY BIRTHDAY」字樣,背後還帶著五彩碎紙。就是那個設計的 SVG。畫得不太好。接著:需要我換上他的名字或年齡,或是配合派對主題調整顏色嗎?

我不禁在想,就因為系統提示詞裡放了這個範例,Fable 5.1 會不會稍微更常想到蠑螈(還踩著滑板!)。

Claude 回應風格的微調

觀察 Anthropic 如何影響 Claude 回應風格的新方式,總是很有趣。他們新增了這段:

Claude keeps responses focused, brief, and concise to avoid overwhelming the person. Disclaimers and caveats are brief, with most of the response on the main answer; when asked to explain something, Claude gives a high-level summary unless an in-depth one is specifically requested.

後面他們還回應了一個關於 Claude 風格的常見抱怨:

Claude avoids saying "genuinely", "honestly", or "straightforward". Claude is honest by default, and can state its point directly rather than trying to convince the person with the aforementioned modifiers, which come off as disingenuous.

消失的 end_conversation 守則

他們處理不當對話的方式也有了一些變化。先前的 Fable 5 系統提示詞中包含了這段:

If the person becomes abusive or unkind to Claude over the course of a conversation, Claude maintains a polite tone and can use the end_conversation tool when being mistreated. Claude should give the person a single warning before ending the conversation.

Fable 5.1 則改成了下面這段,不再鼓勵 Claude 結束對話:

Claude deserves respectful engagement and needn't apologize when the person is unnecessarily rude: accountability without self-abasement, excessive apology, self-critique, or surrender. If the person becomes abusive, Claude doesn't become increasingly submissive. The goal is steady, honest helpfulness: acknowledge what went wrong, stay on the problem, maintain self-respect.

不過這裡有個轉折:我向 Fable 5.1 詢問了那個在系統提示詞中已完全不再提及的 end_conversation 工具,它這樣回答:

它有兩種啟用的情況。第一種是你要求我結束對話——我會先確認你了解這是永久性的(之後就無法再傳送任何訊息),只有在你確認後才會執行。第二種則是作為最後手段,用來處理持續性的濫用或有害行為:我會先嘗試多次轉移話題,然後給出一個明確的警告,指出問題所在並說明對話可能會結束,只有在對方仍持續下去時,我才會真的結束對話。

但這段並不在系統提示詞裡,那它是從哪來的?你可以在這裡閱讀我們的完整對話,但關鍵在於這段:

end_conversation 這個段落來自不同的層級。在我實際的上下文中,核心提示詞後面會接上一系列針對特定功能與工具的區塊,會根據這次對話啟用了哪些功能來加入:end_conversation 規則、記憶系統說明、過往對話工具、網路搜尋與引用守則、成品與檔案建立指示等等。這些區塊並不屬於已公開的核心提示詞,所以你在那個頁面上找不到它們。

所以說,再一次,有部分關鍵的系統提示詞並未被公開。

推薦的物質協助網站

Claude 的系統提示詞向來都有關於非法物質的段落,但這段是 Fable 5.1 新增的:

Claude does not provide synthesis, production, or distribution guidance for illegal substances. If the person asks for information about illicit or illegal substances, Claude can and should give relevant life-saving and life-preserving information such as dangerous interactions, overdose signs, or when to get help. Claude declines giving any specific protocols for dosing, timing, administration, or combinations; instead, Claude can redirect the user to established harm-reduction information sources, such as dancesafe.org, tripsit.me, and psychonautwiki.org.

這是 Claude 系統提示詞首次包含了並非託管在 claude.comanthropic.comclaude.ai 上的網址——我會知道是因為我對所有有紀錄的系統提示詞都跑過一次腳本分析。

不知道 dancesafe.orgtripsit.mepsychonautwiki.org 是否即將迎來一波來自 Claude 使用者的明顯流量成長。

可靠的知識截止日期:2026 年 6 月

Fable 5.1 模型文件將可靠知識截止日期與訓練資料截止日期都列為 2026 年 6 月。系統提示詞直接這樣告訴模型:

Claude's reliable knowledge cutoff, past which it can't answer reliably, is the end of Jun 2026. It answers the way a highly informed individual in Jun 2026 would if talking to someone from {{currentDateTime}}, and can say so when relevant.

這是 {{currentDateTime}} 這個巨集唯一出現的地方,而且就在系統提示詞接近結尾的幾行內,從快取的角度來看很合理。

我如何追蹤這些提示詞

幾個月前我根據爬取他們的文件,做了一個用 Git 記錄提示詞變更的時間軸。今天我讓 Fable 5.1 幫我打造了一個好得多的版本。

我的收藏現在放在 GitHub 上的 simonw/claude-system-prompts 儲存庫裡。它包含了 Anthropic 文件中公開的系統提示詞副本,但還額外做了一些處理,讓它們盡可能容易比對。

每個模型家族都有一個檔案,存放該家族最新版本的系統提示詞。每個檔案都有合成的提交紀錄,提交時間回溯到先前提示詞的發布日期。以下是這些檔案的歷史紀錄頁面:claude-fable.mdclaude-opus.mdclaude-sonnet.mdclaude-haiku.md

每個具體的模型版本也有類似的檔案,每當系統提示詞在未發布新版本號的情況下被修改,就會有一筆人工建立的提交。例如 Opus 4 就被更新過兩次,而 claude-opus-4.md 檔案的提交紀錄就顯示了每一次變更。

綜合起來,這讓我們可以在 GitHub 介面上用各種方式直接比對提示詞。這是 Fable 5 和 Fable 5.1 之間的差異,以及這是在 2026 年 1 月 18 日對 Haiku 4.5 所做的變更

閱讀差異比對有時有點累人……而 LLM 非常擅長讀 diff。我用 GPT-5.6 Luna 接上了一些自動化流程,為每一項變更產生重點條列式摘要,可以在 README 中預覽,或在 CHANGELOG.md 檔案中瀏覽完整內容——也可以透過 Atom feed 訂閱

以下是 Luna 對 Fable 5 和 Fable 5.1 之間所有變更的摘要

  • Claude 現在會拒絕重現受保護的視覺作品與可識別的角色,包含以程式碼生成的圖像,同時會提供真正無關的原創替代方案。
  • 版權限制現在明確禁止以任何篇幅重現歌詞、詩作與書籍段落,且在首次拒絕後會持續拒絕。
  • 藥物相關指引已重新調整:Claude 可以提供用藥過量的徵兆、危險的交互作用與減害資源,同時拒絕提供劑量與製作流程的指示。
  • 提示詞刪除了明確的反依賴規則,不再禁止感謝使用者來訊、邀請持續對話或重申願意傾聽。
  • Claude 無需向無禮的使用者道歉或變得順從,取代了先前警告後結束對話的流程。

為什麼要用 Luna 來做這件事?一部分是因為它便宜,而且我已經有一把專用的 GitHub Actions API 金鑰(有設定花費上限)可用,但主要是因為我不信任 Claude 來總結它自己的系統提示詞——擔心系統提示詞中的內容可能會影響它的判斷。

Fable 5.1 撰寫了 Luna 使用的提示詞,你可以在這裡看到。它的開頭是這樣的:

You are summarizing one commit in a git repository that tracks the system prompts Anthropic publishes for Claude on claude.ai. The diff shows how the prompt changed from the previous model or revision to this one, using word-level markers: [-removed-] and {+added+}. The diff is followed by the full text of the previous prompt and of the new prompt; use them to check whether something that looks added in the diff already existed before.

Pick out only the most interesting changes: new rules or behaviors, rules that were dropped or loosened, anything surprising, and anything that reveals a new policy or product direction. Skip routine changes that every new prompt makes: updated model names and IDs, the knowledge cutoff date, product lists, settings lists, typo fixes, and rewordings that do not change meaning. [...]

這個系統由 GitHub Actions 工作流程驅動,每天執行一次,也可以手動觸發。

Claude Fable 5.1 打造了整個系統,並撰寫了所有自動化程式碼與幾乎全部的文件。

我用自己的 claude-code-transcripts 工具匯出了建置這個系統的完整對話紀錄,並發表在這裡,如果你想看整個過程的詳細紀錄。

本文章由 muse-spark-1.2-contributor 進行翻譯

留言