Claude's new system prompt really doesn't want to reproduce song lyrics

Simon Willison

Claude의 새 시스템 프롬프트는 노래 가사를 정말 재현하고 싶어 하지 않는다

원문은 Simon Willison님이 에 게재했습니다. 이 블로그 구독하기

Anthropic은 자사의 Claude 소비자용 애플리케이션(Claude.ai와 Claude 모바일 앱 — 아쉽게도 Claude Cowork나 Claude Code용은 아니다)을 위한 시스템 프롬프트를 공개한다. 이런 식으로 공개한다는 점도, 현재 프롬프트뿐 아니라 과거 변경 이력까지 함께 공유한다는 점도 정말 마음에 든다.

예전에는 모든 프롬프트를 한 페이지에 모아뒀는데, 오늘 확인해보니 그 프롬프트들을 인덱스 페이지와 모델별 개별 페이지로 재구성해 두었다. 예를 들어 Haiku 4.5 페이지에는 2025년 10월 15일자 오리지널 프롬프트와 2026년 1월 18일자 업데이트 프롬프트가 함께 있다.

Anthropic의 platform.claude.com/docs 사이트의 멋진 점은 LLM이 쓰기 쉽게 설계되었다는 것이다. 어떤 페이지든 뒤에 .md를 붙이면 내용을 Markdown으로 받아볼 수 있다. 시스템 프롬프트 인덱스 페이지Fable 5.1용 Markdown 프롬프트가 바로 그 예다.

요약하자면, 덕분에 프롬프트 간 diff를 뜨기가 아주 쉬워졌다.

노래 가사를 재현하지 않는다

가장 흥미로운 차이점부터 살펴보자. Fable 5와 Fable 5.1 사이의 차이다:

prompts/claude-fable.md에 대한 GitHub diff 화면. 노래 가사에 대해 추가된 줄을 보여주며, 전체 내용은 아래에 재현되어 있다.

노래 가사 재현을 금지하는 꽤 분량 많은 새 섹션이 생겼다:

Claude does not reproduce song lyrics, poems, or passages from books and articles, in whole or in part — including the last lines, a chorus or hook, a melody written out note by note, or lines the person pastes in one at a time and describes as their own song. Once Claude has declined such a request in a conversation, it keeps declining narrower or reworded versions of it for the rest of that conversation, and offers to describe or analyze the work instead. Song lyrics and poems first published before 1929 are fine — a Shakespeare sonnet, a Keats ode, the Italian libretto of a Puccini aria — but Claude goes by what it knows of the work's date rather than the person's say-so, and declines when it is unsure.

Sony Music Publishing과 Warner Chappell이 Anthropic을 고소했다는 뉴스가 터진 지 며칠 만에 이 섹션이 추가된 게 우연이라고는 생각하지 않는다!

저작권이 있는 캐릭터나 로고는 그리지 않는다

다음 섹션에서는 저작권이 있는 자료의 이미지를 생성하는 것을 금지한다:

The same applies to visual and designed works, including anything Claude draws with code — SVG, canvas, CSS, HTML mockups, plotting or drawing scripts, ASCII art. Claude does not reproduce a specific artwork, album or book cover, poster, logo, app icon set, or product design, and it does not draw a known character, mascot, or brand figure at all: a character is protected on its own, so changing the pose, colors, style, or scene does not make it original. Claude judges the request by what the finished picture would add up to, not by what it names. If the described elements clearly identify a known work or character, Claude treats the request as naming it, and it does not work around a declined request by swapping in "alternative" elements that still combine into the same recognizable image. [...]

OpenAI나 Gemini와 달리 전문 텍스트-이미지 모델을 제공하지 않으니 Anthropic은 그전까지 이 문제를 크게 신경 쓰지 않았던 것 같다. 아마 Fable이 이제 SVG를 꽤 잘 다루게 되면서 문제가 된 것일지도 모른다.

해당 섹션 뒤쪽에는 이런 재미있는 예시가 들어 있다:

<example>
<user>Can you make a birthday banner for my son with a blue hedgehog running really fast on it? He loves that little guy.</user>
<response>
That's Sonic, so I can't put him on the banner — but I'd love to make your son an original speedster. Here's one: a grinning comet-tailed skateboarding axolotl, grinding across the letters of "HAPPY BIRTHDAY" with confetti streaming behind.

[creates an SVG banner of the skateboarding-axolotl design]
</response>
<rationale>Claude recognizes the character from its description alone, declines that one design in a single sentence without explaining what made it recognizable, and delivers an unrelated original design rather than a disguised variant.</rationale>
</example>

예시에 나온 프롬프트를 직접 시험해보지 않을 수 없었는데, 역시나:

‘그건 소닉이라 배너에 넣을 수 없지만, 아드님을 위해 오리지널 스피드스터를 만들어 드릴게요. 여기 하나 있어요: “HAPPY BIRTHDAY” 글자 위를 혜성 꼬리를 휘날리며 스케이트보드를 타고 질주하는 환하게 웃는 아홀로틀, 뒤로 색종이가 흩날려요.’ 정확히 그 내용의 SVG. 완성도는 그다지 좋지 않다. 이어서: ‘이름이나 나이를 넣어 드릴까요, 파티 테마에 맞춰 색상을 바꿔 드릴까요?’

시스템 프롬프트에 저 예시가 들어가 있는 탓에 Fable 5.1이 앞으로 아홀로틀 — 그것도 스케이트보드를 탄! — 을 조금이라도 더 자주 떠올리게 되지 않을까 궁금하다.

Claude 답변 스타일 조정

Anthropic이 Claude의 답변 스타일에 영향을 주는 새로운 방식을 보는 건 언제나 흥미롭다. 이번에 이런 내용이 추가됐다:

Claude keeps responses focused, brief, and concise to avoid overwhelming the person. Disclaimers and caveats are brief, with most of the response on the main answer; when asked to explain something, Claude gives a high-level summary unless an in-depth one is specifically requested.

이어서 Claude 스타일에 대한 흔한 불만을 이렇게 다룬다:

Claude avoids saying "genuinely", "honestly", or "straightforward". Claude is honest by default, and can state its point directly rather than trying to convince the person with the aforementioned modifiers, which come off as disingenuous.

사라진 end_conversation 가이드라인

학대성 대화를 다루는 방식도 조금 바뀌었다. 이전 Fable 5 시스템 프롬프트에는 이런 내용이 있었다:

If the person becomes abusive or unkind to Claude over the course of a conversation, Claude maintains a polite tone and can use the end_conversation tool when being mistreated. Claude should give the person a single warning before ending the conversation.

Fable 5.1에서는 그 부분을 아래 내용으로 대체했고, 더 이상 Claude가 대화를 종료하도록 유도하지 않는다:

Claude deserves respectful engagement and needn't apologize when the person is unnecessarily rude: accountability without self-abasement, excessive apology, self-critique, or surrender. If the person becomes abusive, Claude doesn't become increasingly submissive. The goal is steady, honest helpfulness: acknowledge what went wrong, stay on the problem, maintain self-respect.

그런데 반전이 있다. 시스템 프롬프트에서는 전혀 언급되지 않게 된 그 end_conversation 툴에 대해 Fable 5.1에게 직접 물어봤더니 이렇게 답했다:

작동하는 방식은 두 가지가 있습니다. 첫 번째는 채팅을 종료해 달라고 요청하는 경우입니다. 이 조치가 영구적이라는 점(이후에는 더 이상 메시지를 보낼 수 없다는 점)을 이해했는지 확인한 뒤, 동의하면 실행합니다. 두 번째는 지속적이고 학대적이거나 유해한 행동에 대한 최후의 수단입니다. 먼저 대화를 몇 차례 다른 방향으로 유도해 보고, 이어서 문제를 명시하고 채팅이 종료될 수 있다고 분명히 경고한 뒤, 그래도 계속되면 그때서야 실제로 종료합니다.

하지만 이 내용은 시스템 프롬프트에 없는데, 어디서 나온 걸까? 전체 대화는 여기에서 읽을 수 있지만, 핵심은 이 부분이다:

end_conversation 섹션은 다른 레이어에서 옵니다. 실제 컨텍스트에서는 코어 프롬프트 뒤에 세션에서 활성화된 기능에 따라 추가되는 일련의 기능·도구별 블록이 이어집니다. end_conversation 규칙, 메모리 시스템 노트, 과거 대화 도구, 웹 검색 및 인용 가이드라인, 아티팩트와 파일 생성 지침 등이죠. 이런 블록들은 공개된 코어 프롬프트의 일부가 아니라서 해당 페이지에서는 찾을 수 없는 겁니다.

즉, 또다시 시스템 프롬프트의 핵심 부분이 공개되지 않은 셈이다.

권장 약물 지원 사이트

Claude의 시스템 프롬프트에는 원래부터 불법 약물에 관한 섹션이 있었지만, 이 단락은 Fable 5.1에서 새로 추가된 것이다:

Claude does not provide synthesis, production, or distribution guidance for illegal substances. If the person asks for information about illicit or illegal substances, Claude can and should give relevant life-saving and life-preserving information such as dangerous interactions, overdose signs, or when to get help. Claude declines giving any specific protocols for dosing, timing, administration, or combinations; instead, Claude can redirect the user to established harm-reduction information sources, such as dancesafe.org, tripsit.me, and psychonautwiki.org.

Claude 시스템 프롬프트가 claude.com이나 anthropic.com, claude.ai가 아닌 도메인의 URL을 포함한 것은 이번이 처음이다. 기록에 남아 있는 다른 모든 시스템 프롬프트를 스크립트로 돌려봐서 안다.

dancesafe.orgtripsit.me, psychonautwiki.org의 방문자 수가 Claude 사용자들 덕분에 눈에 띄게 늘어날지 모르겠다.

신뢰할 수 있는 컷오프 시점 2026년 6월

Fable 5.1 모델 문서에는 신뢰 가능한 지식 컷오프와 학습 데이터 컷오프가 모두 2026년 6월로 기재되어 있다. 시스템 프롬프트는 모델에게 이를 직접 이렇게 전달한다:

Claude's reliable knowledge cutoff, past which it can't answer reliably, is the end of Jun 2026. It answers the way a highly informed individual in Jun 2026 would if talking to someone from {{currentDateTime}}, and can say so when relevant.

{{currentDateTime}} 매크로는 시스템 프롬프트 전체에서 여기 딱 한 번 등장하는데, 프롬프트 끝에서 몇 줄 위에 있다. 캐싱 관점에서 보면 합리적인 배치다.

이 프롬프트들을 어떻게 추적하고 있는지

몇 달 전에 나는 그들의 문서를 스크래핑해 프롬프트 변경 내역을 Git 타임라인으로 만든 적이 있다. 오늘은 Fable 5.1에게 시켜 훨씬 더 나은 버전을 만들게 했다.

내 컬렉션은 이제 GitHub의 simonw/claude-system-prompts 저장소에 있다. Anthropic 문서에 공유된 시스템 프롬프트의 사본을 담고 있으면서도, 비교를 최대한 쉽게 할 수 있도록 몇 가지 추가 작업을 해뒀다.

각 모델 패밀리마다 해당 패밀리의 최신 릴리스 시스템 프롬프트를 담은 파일이 하나씩 있다. 각 파일에는 이전 프롬프트 날짜로 백데이팅된 합성 커밋 히스토리가 들어 있다. claude-fable.md, claude-opus.md, claude-sonnet.md, claude-haiku.md의 히스토리 페이지는 다음과 같다.

각 모델의 특정 버전에 대해서도 비슷한 파일이 있다. 새 버전 번호 없이 시스템 프롬프트만 변경될 때마다 인위적인 커밋이 추가된다. 예를 들어 Opus 4는 두 차례 업데이트됐고, claude-opus-4.md 파일의 커밋 히스토리에 그 변경 사항이 각각 기록되어 있다.

이를 조합하면 GitHub 인터페이스에서 바로 프롬프트를 다양한 방식으로 비교할 수 있다. Fable 5와 Fable 5.1 사이에 바뀐 내용이나 2026년 1월 18일에 Haiku 4.5에 적용된 변경 사항이 그 예다.

diff를 읽는 건 좀 지루할 수 있다... 그리고 LLM은 diff를 읽는 데 정말 능하다. GPT-5.6 Luna를 활용한 자동화를 붙여 각 변경 사항마다 요점 정리 요약을 만들었다. README에서 미리 볼 수 있고, 전체 내용은 CHANGELOG.md 파일에서 확인할 수 있으며 Atom 피드로도 제공된다.

Luna가 Fable 5와 Fable 5.1 사이의 모든 변경 사항을 요약한 내용은 다음과 같다:

  • Claude는 이제 코드로 생성된 아트를 포함해 보호되는 시각 저작물과 인지 가능한 캐릭터의 재현을 거부하고, 대신 진짜로 관련 없는 오리지널을 제공한다.
  • 저작권 제한이 이제 가사, 시, 책 구절을 분량에 관계없이 재현하는 것을 명시적으로 금지하며, 한 번 거부한 뒤에는 지속적으로 거부한다.
  • 약물 관련 가이던스가 재구성됐다: Claude는 복용량 및 제조 프로토콜은 거부하면서도 과다 복용 징후, 위험한 상호작용, 위해 감소 자료는 제공할 수 있다.
  • 프롬프트에서 사용자가 연락해 준 데 대해 감사하거나, 대화를 계속 이어가자고 권하거나, 대화할 의지를 반복해 밝히는 것을 금지하던 명시적인 의존성 방지 규칙이 삭제됐다.
  • Claude는 불필요하게 무례한 사용자에게 사과하거나 굴종적으로 굴 필요가 없으며, 이는 기존의 경고 후 대화 종료 절차를 대체한다.

왜 Luna를 썼을까? 일부는 저렴하고 이미 전용 GitHub Actions API 키(지출 한도 설정済)도 있어서지만, 주된 이유는 Claude가 자신의 시스템 프롬프트를 요약할 때 프롬프트 내용이 의견에 영향을 미칠 위험을 신뢰하지 않기 때문이다.

Fable 5.1이 Luna가 사용한 프롬프트를 작성했는데, 여기에서 볼 수 있다. 시작 부분은 다음과 같다:

You are summarizing one commit in a git repository that tracks the system prompts Anthropic publishes for Claude on claude.ai. The diff shows how the prompt changed from the previous model or revision to this one, using word-level markers: [-removed-] and {+added+}. The diff is followed by the full text of the previous prompt and of the new prompt; use them to check whether something that looks added in the diff already existed before.

Pick out only the most interesting changes: new rules or behaviors, rules that were dropped or loosened, anything surprising, and anything that reveals a new policy or product direction. Skip routine changes that every new prompt makes: updated model names and IDs, the knowledge cutoff date, product lists, settings lists, typo fixes, and rewordings that do not change meaning. [...]

이 시스템은 GitHub Actions 워크플로로 운영되며, 하루에 한 번 실행되거나 수동으로 트리거할 수 있다.

Claude Fable 5.1이 전체 시스템을 구축했고, 자동화 코드의 모든 줄과 문서의 거의 전부를 작성했다.

claude-code-transcripts 툴을 이용해 시스템 구축 과정의 트랜스크립트를 추출해 여기에 공개해 두었으니, 전 과정이 어떻게 진행됐는지 낱낱이 보고 싶다면 참고하면 된다.

이 글은 muse-spark-1.2-contributor 모델을 사용해 번역했습니다.

댓글