For how zero data retention (ZDR) applies to this feature, see API and data retention.
System instructions normally live in the top-level system field, ahead of every message in the conversation. That position is great for prompt caching: the system prompt is part of the stable prefix, so subsequent turns hit the cache. It is a poor position for instructions you only discover you need partway through a session, because editing the top-level system field changes the very beginning of the prompt and invalidates the cache for everything that follows.
Mid-conversation system messages close that gap. You append a {"role": "system"} message at the point in the conversation where the new instruction becomes relevant, instead of editing the top-level system field. The cached prefix stays the same, so the next request still reads it from cache, and the new instruction is still applied as a system instruction rather than as ordinary user text.
This page covers two features: mid-conversation system messages, which are generally available, and mid-conversation tool changes, a beta introduced with Claude Opus 5 that applies the same approach to the tools array.
Mid-conversation system messages are available on the Claude API, Claude in Amazon Bedrock, and Google Cloud.
This feature is available on Claude Fable 5, Claude Mythos 5, Claude Opus 4.8, and Claude Opus 5. No beta header is required for mid-conversation system messages. This feature is not available on Claude Sonnet 5; use the top-level system field instead.
Mid-conversation tool changes are in beta and require the mid-conversation-tool-changes-2026-07-01 beta header. They are available on Claude Fable 5, Claude Mythos 5, Claude Opus 4.8, and Claude Opus 5, on the Claude API, Amazon Bedrock, and Google Cloud.
The tools array sits even earlier in the hashed request prefix than the top-level system field, so editing it invalidates the prompt cache for the entire conversation. Mid-conversation tool changes, a beta introduced with Claude Opus 5, are the tools counterpart to mid-conversation system messages. Instead of fixing the tool list for the lifetime of the conversation, you change which tools are offered to the model between turns: declare the full tool set in tools up front, then use tool_addition and tool_removal blocks to offer a tool to the model, or withdraw it, from a specific point in the conversation onward. The tools array itself never changes, so the cached prefix stays intact.
tool_addition and tool_removal are content blocks in the content array of a role: "system" message, and they can be mixed with text blocks in the same message. The message follows the same placement rules as any mid-conversation system message (see Limitations), and the change applies from that point in the conversation onward. Each block's tool field references a tool rather than defining one: {"type": "tool_reference", "name": "..."} names a tool declared in the request's tools array, and MCP connector tools can be referenced individually with mcp_tool_reference (server_name and name) or as a whole toolset with mcp_toolset_reference (server_name). Referencing a name that is not declared in tools returns a 400 error.
Every tool declared in tools is offered to the model from the start of the conversation unless it is declared with defer_loading: true, which keeps it withheld until a tool_addition block surfaces it. tool_addition also re-offers a tool that an earlier tool_removal withdrew.
client = anthropic.Anthropic()
response = client.beta.messages.create(
model="claude-opus-5",
max_tokens=1024,
betas=["mid-conversation-tool-changes-2026-07-01"],
# The full tool set is declared up front and never changes, so the
# cached prefix stays intact.
tools=[
{
"name": "get_weather",
"description": "Get the current weather for a location.",
"input_schema": {
"type": "object",
"properties": {
"location": {"type": "string", "description": "City name"},
},
"required": ["location"],
},
},
],
messages=[
{
"role": "user",
"content": "Say OK.",
},
# Withdraw get_weather from this point onward. The block references
# the tool by name instead of editing `tools`, so earlier turns stay
# byte-identical and the cache still hits.
{
"role": "system",
"content": [
{
"type": "tool_removal",
"tool": {"type": "tool_reference", "name": "get_weather"},
},
],
},
],
)
for block in response.content:
if block.type == "text":
print(block.text)Mid-conversation tool changes are in beta. To use them, include the beta header mid-conversation-tool-changes-2026-07-01 in your requests. They are available on Claude Fable 5, Claude Mythos 5, Claude Opus 4.8, and Claude Opus 5, on the Claude API, Amazon Bedrock, and Google Cloud.
Prompt caching hashes the request prefix in order: tools, then system, then messages. A cache hit requires the prefix to match a recent request exactly, byte for byte, up to the cache breakpoint.
That ordering means the top-level system field sits near the very start of the hashed prefix. Any change to it, even appending a sentence, produces a different hash, and the request misses the cache for the system prompt and every cached message after it.
Mid-conversation system messages let you add the instruction at the end of the message history instead. Everything before the new instruction is unchanged, so the existing cache entry still matches, and only the new message is processed as fresh input.
A few situations where this matters:
system field would re-process the entire history.In all of these cases you could put the instruction in a regular user message, and Claude does follow instructions that arrive in user turns. The difference is priority: a user message is treated as coming from the end user, while a system message is treated as coming from you, the application operator. When the two conflict, system instructions take precedence, so use the system role for operator-level facts and constraints that should hold even if the end user asks for something different. A mid-conversation system message keeps that operator-level priority without paying the cache-miss cost of editing the top-level system field.
Add a message with "role": "system" to the messages array. Use a plain string or content blocks for content, the same as a user or assistant turn. The instruction applies from that point in the conversation onward. When instructions conflict, later system messages take precedence over earlier ones, and mid-conversation system messages take precedence over the top-level system field for the turns that follow them.
You can still set the top-level system field for instructions that should apply to the entire conversation. Reserve mid-conversation system messages for instructions that only become relevant later, or that you want to add without invalidating the cached prefix.
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-opus-5",
max_tokens=1024,
# Automatic prompt caching: each request caches the conversation so far,
# and the next request reads the unchanged prefix from cache.
cache_control={"type": "ephemeral"},
system="You are a code review assistant. Be concise.",
messages=[
{
"role": "user",
"content": "Review process() in utils.py for performance issues.",
},
{
"role": "assistant",
"content": "The list comprehension is fine for small inputs. For large inputs, consider a generator to avoid materializing the full list.",
},
{
"role": "user",
"content": "Now review the calling code that invokes process().",
},
# The reviewer realizes mid-session that all suggestions must
# also pass the team's strict typing policy. Appending the
# instruction here keeps earlier turns byte-identical, so the
# prefix cached by the previous request is still read from cache.
{
"role": "system",
"content": "From now on, every suggestion must include explicit type annotations.",
},
],
)
for block in response.content:
if block.type == "text":
print(block.text)This example enables automatic caching with the top-level cache_control field. Prompt caching is opt-in: if a request has no cache_control field (automatic or an explicit breakpoint), nothing is cached and every request pays the regular input token price for the full conversation. With caching enabled, appending the system message leaves the already-cached turns unchanged, so the request that carries the new instruction still reads them from cache instead of processing them again. Caching also requires the conversation to meet the minimum cacheable prompt length; an example as short as this one falls below it, so cache_creation_input_tokens and cache_read_input_tokens stay at 0 until the conversation grows.
A mid-conversation system message must immediately follow a user turn (or an assistant turn ending in a server tool result), and must either be the last entry in messages or be immediately followed by an assistant turn. A user message that carries tool_result blocks counts: in an agentic loop you can place the system message right after the tool results, before Claude's next turn. Any other position, including between an assistant tool_use block and the tool_result that answers it, returns a 400 error.
In an agentic loop, the system message goes after the user message that delivers the tool results. This is also where your application can relay input that the user typed while Claude was working, so the new context is absorbed without restarting the turn:
[
{ "role": "user", "content": "Run the test suite and fix any failures." },
{
"role": "assistant",
"content": [{ "type": "tool_use", "id": "toolu_01", "name": "run_tests", "input": {} }]
},
{
"role": "user",
"content": [
{ "type": "tool_result", "tool_use_id": "toolu_01", "content": "12 passed, 0 failed" }
]
},
{
"role": "system",
"content": "The user sent the following message while you were working: also update the changelog before you finish."
}
]Phrase the system content as context rather than as a command that overrides the user. State the fact ("new input arrived from the user: X", "the remaining token budget is now Y") and let Claude act on it. Claude is trained to resist instructions that appear to work against the user, and that protection still applies to the system role, so language such as "ignore what the user said" is less effective than stating what changed.
This pattern is for relaying input from the conversation's own end user. Do not use it to pass tool output, retrieved documents, or other third-party content; keep that content in tool_result blocks (see Limitations).
Mid-conversation system messages and prompt caching are designed to be used together:
cache_control, either the top-level automatic caching field or an explicit breakpoint on a content block. A mid-conversation system message does not create a cache entry on its own, and without caching enabled there are no savings to preserve.cache_control on the last block that stays the same across requests, whether that is the end of the top-level system field, the end of your tool definitions, or a stable point in the message history.Avoid editing or removing a mid-conversation system message that has already been sent. Like any other change to earlier messages, that invalidates the cache from that point forward. If the instruction needs to evolve, append a new system message rather than rewriting the old one. Consecutive system messages are accepted and treated as a single system section, which follows the same placement rule as a whole.
system message cannot be the first entry in messages. Use the top-level system field for instructions that apply from the very start.system message must immediately follow a user turn (including a user turn that carries tool_result blocks) or an assistant turn ending in a server tool result, and must precede an assistant turn or end the array. It cannot sit between a tool_use block and its tool_result. Placing it elsewhere returns a 400 error.tool_result blocks and continue to follow Mitigate jailbreaks and prompt injections.How caching works, where to place breakpoints, and how to read cache usage fields.
Find out exactly where two requests diverged when a cache hit you expected does not happen.
Message structure, multi-turn conversations, and the system field.
Writing effective prompts and system instructions.
How tool_use and tool_result blocks are structured in the messages array.
Was this page helpful?