A code-backed guide to inference state, API cache rules, SDK translation, and context management in coding agents.

Thoughts by human, co-written by AI

35
live Sonnet 5.5 requests across two small experiments
7,870
cached tokens preserved after an appended system update
$0.21
estimated inference spend; not total research cost

I'm building an agent harness, and I want to change its instructions, tools, and retrieved documents without paying to process the whole conversation again. The catch is correctness. A cheap request is no use if the model is still working from an outdated instruction or a document I meant to remove.

My starting picture was a Jenga tower: change something near the bottom of the prompt and everything above it needs rebuilding. That is mostly right for an edit to earlier text. But an agent has other ways to handle change. It can append a new instruction, return to an earlier conversation branch, or deliberately replace a long history with a shorter one.

In a small Sonnet 5.5 test, changing an instruction near the start caused a complete cache miss. Appending the new instruction instead preserved 7,870 cached tokens. Both requests produced the expected answer. Understanding why those requests differ is the key to building a flexible harness without treating every change as a fresh conversation.


What the model saves

A prompt first becomes tokens: numbers representing pieces of text and the message structure around them. Tokens are often shorter than words. The model turns those numbers into vectors, which are lists of values, then processes them through a stack of layers.

At an attention layer, it creates three kinds of vectors. A query is used to decide which earlier positions to attend to. Each earlier position has a key used in that comparison and a value carrying information to combine into the result. The names are usually shortened to Q, K, and V. These are learned numerical representations, not a database of facts.

When generating the next token, the model needs keys and values for the text already processed. Saving them avoids computing them again. That saved state is the KV cache. It exists during ordinary generation even when no provider offers a prompt-caching discount. Reusing compatible state in a later request is the additional step called prompt caching. Attention mechanism, walkthrough of the serving code.

There are two stages to keep separate:

  • Prefill: process the input and create its KV state. A prompt-cache hit can skip much of this work for an unchanged beginning.
  • Decode: produce new output tokens. With ordinary full attention, each new token still attends over the earlier keys and values. Keeping 100,000 tokens cached does not make them disappear from the next token's computation or from the context window.
Inference · illustrative token sequence Build once. Read at every next step. Prefill processes the known input. Decode adds one generated token at a time.
01

Prefill the prompt

Known inputA B C

All three input tokens are available.

Model layersCreate keys and values

Each position can attend to itself and earlier positions.

First outputD

Scores at C select D. D’s state is built next.

Save for reuseK/V [ A B C ]Separate keys and values at each attention layer.
02

Decode the next token

Next inputD

Feed the selected token back through the model.

Model layersRead old state + add D

D’s query attends to saved A/B/C and D’s new keys and values.

Next outputE

New scores select E. Repeat for the next token.

Retain and extendK/V [ A B C D ]Read the saved prefix; retain the new state alongside it.

A prompt-cache hit skips rebuilding compatible old state. Reading that state, processing new input, and generating output still take work.

That helps explain why a cache hit can be cheap without making the whole response fast. The server may still queue the request, move saved state into GPU memory, process new input, and generate a long answer.


Why one earlier edit affects later text

A prefix is the beginning of the input. A suffix is what follows it. For ordinary causal attention, each position can use earlier positions but cannot use future ones. Appending a message therefore leaves the earlier computation valid. Changing an earlier message can change what later positions compute.

Imagine an instruction followed by a document:

Original:  [Report prices in USD.] [A long document about prices…]
Edited:    [Report prices in EUR.] [The same document…]

The document's words are unchanged. Its deeper-layer representations need not be: those layers processed the document with the earlier instruction available. Reusing the old document state after replacing USD with EUR can feed the model numbers from a different input than the one requested. Similar meaning, a small text diff, or an unchanged document hash cannot prove the computation is interchangeable.

The serving code tracks this dependency. In vLLM, the identifier for a cached block combines its tokens with the previous block's identifier and other relevant inputs. In simplified pseudocode:

block_id = hash(previous_block_id, token_ids, extra_keys)

Changing an early block changes the identifiers of later blocks even when their own tokens stay the same. SGLang organizes cached prefixes as a tree, so requests can share the beginning and branch where they differ. vLLM block hashing, SGLang prefix matching.

A new branch does not have to destroy the old one. If I later send the original USD request again, the server may still have its cached state. The limit is availability: entries can expire or be evicted to make room for other work.

A worked example · six illustrative tokens What changes when token D changes? Replace D with D′. Keep every other token and its position fixed.
Asame Bsame Csame D′edit Esame Fsame
01

An earlier edit can change F’s state

Follow F through a conventional causal transformer. Its text stays the same; what it reads changes.

Entering layer 1 F at position 6

Same token, same position. Its first-layer key and value can remain the same.

Layer 1 attention F reads A B C D′ E F

The changed earlier token can change the result of this attention calculation.

Entering layer 2 F uses the attention result

Its deeper keys and values are computed from this contextual state. Old values may no longer be valid.

A, B, and C cannot read future token D. Their state remains valid under this same-length edit. E and F can depend on it, even though their token IDs did not change.

What exactly is being reused?

At each layer, the model stores keys and values used by attention. In a conventional decoder, layer 1 starts from each token’s own embedding and position; deeper layers start from the previous layer’s contextual result. Reusing the suffix must preserve those dependencies, not merely its text. This assumes unchanged weights, positions, masks, and other relevant settings.

Attention Is All You Need, sections 3.1–3.2 ↗
02

Block hashes include the earlier prefix

Suppose this engine groups two tokens per block. The block containing E F includes the preceding block’s identity in its hash.

A BShared prefix · same hash H₁
Original branch · reusable if retained
C DH₂ = hash(H₁, C D)
E FH₃ = hash(H₂, E F)

The old branch can still be reused if it remains resident.

Edited branch · compute new state
C D′H₂′ = hash(H₁, C D′)
E FH₃′ = hash(H₂′, E F)

Same E F tokens; different parent hash. The old suffix is not a match.

This simplified chain follows vLLM’s block-hash implementation ↗, which also includes extra keys. Creating the edited branch does not, by itself, erase the original.

03

Reuse stops at an available checkpoint

A B C is still mathematically valid. But if the engine can resume only after complete two-token blocks, the available matching checkpoint is after B.

A BRead from cacheCheckpoint matches
C D′Compute againC is unchanged, but shares the changed block
E FCompute againState can depend on D′

Unchanged does not always mean reusable at the available boundary. Actual engines and APIs choose their own checkpoint rules; some support finer matching. vLLM’s versioned matching options ↗

The measured request is a separate example.

Our Anthropic middle edit read 5,371 tokens and wrote 2,499. Those are API counters for explicit prompt boundaries—not evidence of two-token blocks or Anthropic’s internal hash scheme. See the requests and measurements ↗

A small numerical example shows the difference. It uses an untrained three-layer transformer and replaces one early token without changing the input length. The table compares its final output scores, called logits, with a fresh calculation:

ComputationMaximum difference from fresh final logits
Reuse the valid prefix, recompute the rest0
Recompute the changed token, splice the old suffix state back in0.339935
Append new context to a valid cached prefix0

Reusing the valid beginning gives the same result. Reattaching the old later state gives a different result. This does not measure answer quality: the tiny model has never learned language. It also shows why “every value after an edit changes” is too strong. At an unchanged later token, the first layer’s keys and values stayed equal; deeper layers changed.

The exact zeros come from running the same arithmetic in the same order. Real serving systems can use different numerical precision or GPU execution orders, so this is not a promise of bit-for-bit identical outputs. Recorded result and reproduction limits.


Follow one request through the four layers

Suppose the user asks an agent to inspect a file. The harness chooses the messages and tools, an SDK turns them into an API request, and the provider runs the model. When the model asks to read the file, the harness executes that tool and sends another request with the result attached.

Follow the requestFrom conversation history to saved computation
  1. 01

    Harness

    Chooses the context

    Instructions, messages, tools, and retained history

    Does the next request keep the earlier messages?

  2. 02

    SDK / adapter

    Builds the API request

    Turns the harness’s objects into provider fields

    Does the outgoing JSON contain the intended update?

  3. 03

    Provider API

    Sets the caching rules

    Allowed boundaries, lifetime, updates, and prices

    Where may this request reuse saved input?

  4. 04

    Inference engine

    Finds and uses saved state

    Looks up earlier computations and manages memory

    Is the matching state still available?

The harness controls what is sent. The serving system controls what can be reused.

The second request contains the earlier instructions and conversation, followed by the tool call and result. That repeated beginning is a candidate for reuse. The provider still has to recognize an allowed cache boundary and find the saved state. A conversation in the UI is not proof that either happened.

The API's JSON is also not the model's literal input. Providers render roles, tool definitions, images, and messages into their own model input. An SDK can move instructions or transform a tool schema before that rendering happens. To explain a miss, inspect what was actually sent, then compare the returned usage counters.

A conversation ID may tell a provider which history to restore. A named cache object may refer to a fixed saved prefix. A response cache may return a previous answer without running the model. None of these names, on its own, means that arbitrary edited text can reuse old KV state.


Changing the request in six different ways

The test asked Sonnet 5.5 to return two fields: a currency from the system instruction and a value from record 137. The input contained instructions, reference text, and two blocks of records. Each block ended at an explicit cache boundary: a place where the API was asked to save the preceding input for five minutes. The nine-case sequence ran three times, for 27 requests costing about $0.18. Six of the cases show the useful contrasts:

Observed · Sonnet 5.5 · 3 requests per variantWhere the request changedOpen a row to see what changed in the request.
Cache readCache writeOrdinary input
Edit the early instruction$0.0199mean / request 0 read · 7,871 written · 13 ordinaryInspect +

The original system instruction changes. USD → EUR inside the original system block. The edited prefix is written again.

system: Currency policy: EUR.
        [same reference text]
user:   [same record blocks]

This is a change to the historical prompt. Its repeated request later hit the newly cached branch.

First visible text: median 1.432 s; range 1.407–1.447 s. This small, ordered sample does not establish a speedup.

Replace a record in the middle$0.0076mean / request 5,371 read · 2,499 written · 13 ordinaryInspect +

Reuse stops at the saved boundary. RECORD-137 changes in the second record block. The preceding system and first record block remain reusable.

system: [unchanged]                  ▼
user:   records 000–099 [unchanged]  ▼
        records 100–199 [137 edited] ▼

▼ marks an explicit cache boundary. Even unchanged records before 137 inside the second block are processed again.

First visible text: median 1.442 s; range 1.184–1.492 s. This small, ordered sample does not establish a speedup.

Append a system update$0.0025mean / request 7,870 read · 0 written · 46 ordinaryInspect +

New policy, preserved history. Keep the original USD policy; append a system instruction that sets EUR from this point onward.

system: Currency policy: USD. [same text]
user:   [same record blocks]         ▼
system: From now on, use EUR.

Both expected answer fields passed. This does not prove equivalence to replacing the original instruction. The new suffix itself was not cache-marked.

First visible text: median 1.761 s; range 1.459–2.116 s. This small, ordered sample does not establish a speedup.

Add a tool inline$0.0020mean / request 7,870 read · 0 written · 122 ordinaryInspect +

The definition arrives after the prefix. Append a native tool_addition block. The initial lookup tool and previous text stay unchanged.

tools:  lookup [unchanged]
        [same system and user]      ▼
system: tool_addition → lookup_extra

Request acceptance and reuse were observed. The prompt prohibited tool calls, so tool execution was not tested.

First visible text: median 1.289 s; range 1.218–1.369 s. This small, ordered sample does not establish a speedup.

Return to the original branch$0.0018mean / request 7,870 read · 0 written · 13 ordinaryInspect +

The earlier entry still exists. Send the original USD request after testing the EUR edit. The old branch was still reusable.

A: USD → original records  [cached]
B: EUR → original records  [cached]
A: USD → original records  [reused]

Observed within the short test window. Residency after expiry, routing changes, or eviction is not established.

First visible text: median 1.209 s; range 1.124–1.332 s. This small, ordered sample does not establish a speedup.

Repeat the same request$0.0018mean / request 7,870 read · 0 written · 13 ordinaryInspect +

The unchanged control. Repeat the original request without changing instructions, tools, or record content.

request A
request A again → same prefix

A matched prefix still needs an eligible, available cache entry. This is not an indefinite retention guarantee.

First visible text: median 1.214 s; range 1.143–1.297 s. This small, ordered sample does not establish a speedup.

Bars show each request’s input composition. Costs include output. Each case ran three times. These are measured costs, not a latency prediction.

Read/write counts were the same in all three repetitions of each case. Every answer matched its expected currency and record value. The early edit and appended instruction both changed the expected currency from USD to EUR; the middle edit changed the expected record value.

Changing record 137 preserved the earlier 5,371-token boundary. The entire second record block was processed again, including its unchanged records before 137. The provider had a saved entry at the block boundary; it did not expose reuse at every unchanged token. “How much text is unchanged?” and “where can this API resume?” can have different answers.

The inline tool request was accepted and kept the prefix hit. It did not test whether the model could select or correctly call that tool: the prompt prohibited tool calls. Nor did the test cache the appended control messages for a future turn; all explicit markers were before them.

Cheap also did not mean faster in this sample. The appended-system case had median time to first visible text of 1.761 seconds, versus 1.432 seconds for the early edit. Two appended requests generated extra thinking tokens. Timing includes transport, queueing and reasoning; with three ordered samples per case, it would be misleading to claim a general latency result. Appending the update preserved input reuse and cost less here. The two answer fields passed; broader instruction-following behavior was not tested. Complete method, timing ranges, and counters.

For budgeting, separate the stable prefix from everything else. If it contains P tokens, appears in N requests, and is written once then successfully read on every repeat, its cost is P × (write price + (N − 1) × read price), with prices expressed per token. The uncached comparison is P × N × ordinary input price. Add new input, output, and any storage charges separately.

At the tested Sonnet rates, a five-minute write costs 1.25 times ordinary input and a read costs 0.1 times. One successful reuse already pays back the write premium: 1.25 + 0.1 is less than two ordinary prefills. A one-hour write at twice ordinary input would need two successful reuses: 2 + 0.1 is still more than two prefills, but 2 + 0.1 + 0.1 is less than three. These are prefix-only calculations assuming hits inside retention, not measured whole-task savings. Keeping irrelevant text merely to improve hit rate can still cost more than shortening the context. Dated prices and counter rules.


Add an instruction for the next turn

Anthropic's current API supports system messages within the conversation on selected models, including the tested Sonnet 5.5. A new instruction retains system authority while leaving the preceding prefix intact. Tool additions and removals have their own protocol, including beta inline definitions. This is useful for tools unknown at session start or a schema that changes later. The inline path has constraints, including an initially present non-deferred tool to avoid changing the rendered head when the first new tool appears. Official update contract.

The important change is where the new instruction goes. This shortened request shows the shape; the omitted document must be long enough to meet the model's caching threshold:

{
  "system": [{"type": "text", "text": "Report prices in USD."}],
  "messages": [
    {"role": "user", "content": [{
      "type": "text",
      "text": "The original document…",
      "cache_control": {"type": "ephemeral", "ttl": "5m"}
    }]},
    {"role": "system", "content": "From now on, report prices in EUR."}
  ]
}

The USD instruction and document remain where they were. The new system message changes the instruction for what follows. This is a supported API operation on the tested model; putting the same words in an ordinary user message would give them a different role.

OpenAI also documents controls for later turns: restrict callable tools while keeping their definitions stable, load deferred tools later, or append supported reasoning-configuration updates. Its current caching behavior differs by model generation; older advice about automatic interval caching and retention keys does not fully describe newer explicit boundaries. OpenAI cache controls.

The details differ across APIs:

SurfaceDistinction that changes harness design
OpenAISome models select boundaries automatically; newer models also let callers mark them explicitly. Supported configuration updates can be appended.
AnthropicMark blocks yourself or use an automatically advancing boundary. Supported models accept later system and tool updates.
GeminiAutomatic reuse and named cache objects are different features. A named object’s stored content cannot be edited.
DeepSeekMatching text is reusable only where the service has stored a suitable prefix unit.
OpenRouterRoutes requests to other providers. Staying with the same provider helps locality but does not prove a cache hit.

The dated provider report records endpoints, limits, counters, hosted-platform distinctions, and authoritative links. In particular, current Gemini Interactions and GenerateContent expose different caching surfaces. Do not transfer the contract of one endpoint to another because both use the same model family. Gemini caching, explicit resources, DeepSeek persistence, OpenRouter routing.

For my harness, apply this policy from now on is often enough. A supported system message can express that without changing earlier messages. Removing earlier information is a separate operation: an appended correction leaves the old text and its saved representations in place.


Check what the SDK actually sends

An API feature is useful only if the SDK sends the right fields.

The mocked SDK experiment executed published Vercel provider packages and captured their outgoing requests. OpenAI explicit breakpoints and TTL controls survived. An Anthropic mid-conversation system message and tool-removal reference survived too. Those controls reached the outgoing JSON.

There was a smaller gap: the tested OpenAI provider-options schema accepted mode and ttl, but had no diagnostics comparison-ID field. Supplying an invented JavaScript option did not forward it. The official OpenAI SDK exposed the underlying field. Adding a property to a JavaScript object does not guarantee that an adapter will forward it. This particular gap affected diagnostics; the tested caching controls worked. Versioned SDK findings.

When an adapter cannot express a requested update, the caller needs to know what it did instead. Rebuilding the leading prompt may apply the new instruction but change the cost of every following token. A capture of the outgoing request makes that fallback visible.


How existing harnesses handle changes

The implementations already contain useful ideas to borrow. They also show why a single cache: true option cannot describe every operation.

HarnessFinding at the audited revision
PiPreserves native system updates when supported; otherwise collapses them into the leading prompt. Tool redefinitions trigger a separate fallback.
Oh My PiImplements explicit OpenAI cache anchors, Anthropic tool-control history, stable MCP ordering, and cache-warming policy.
Codex CLICan pin a context window's request-level reasoning effort and append trusted configuration updates on supported models.
OpenCodeApplies provider-specific cache markers through AI SDK, with distinct automatic-caching and runtime paths.
Gemini CLIAppends selected retry nudges rather than rewriting system instructions; authentication routes differ.
Claude CodeOfficial documentation describes stable prefixes, appended updates, cache settings, and compaction; this was not a source audit of its production request builder.

The harness audit, Codex audit, and Claude Code documentation provide the versions and evidence boundaries. Default-branch code is not proof of the behavior of an older installed release.

Pi supplied the clearest local reproduction. Using its actual pinned transcript functions, the same system-section change kept the original prefix when native mid-conversation support was enabled. With that capability disabled or unspecified, it replaced the section in the leading system message. Both representations replayed to the same current system text in Pi's helper, but that does not make them identical model inputs. Fixture and returned transcripts.

A harness can expose that result directly: the update was appended, earlier history was rewritten, or the provider does not support the requested operation. That lets callers choose whether the fallback is acceptable before sending the request.


Replacing a document can be cheaper than appending a correction

Instructions are only half the problem. An agent also needs to replace evidence: a file changes, a search result is outdated, or a tool returns a newer record. Keeping the old version can preserve cache hits while leaving more work for the model.

A second experiment used an inventory record. Version 1 said 6 items at $17 each, for a total of $102. The model answered from that record. Version 2 then changed the values to 9 items at $23, for a total of $207. The next request either replaced the old source and dropped its answer, or retained both and appended an explicit correction.

Next requestCache readsCache writesFull request cost
Replace the source and drop the old answer3,51076$0.0016640
Keep the old source and answer; append the correction3,586290$0.0022142

The appended correction read 76 more cached tokens but cost about 33% more in this fixture. A cache boundary immediately before the small source let the replacement keep the long background prefix. Only 76 tokens needed rewriting. Appending instead retained the previous question and answer and added the correction, producing a larger new suffix.

Both layouts returned the correct current price, quantity, total, and source identifier in both repetitions. The instructions explicitly told the model to use the highest source revision and recompute from it. There was no answer failure here, and no test of ambiguous source precedence. A larger document, an earlier edit, or a longer remaining conversation could change the cost comparison. Requests, results, and all eight calls.

A third branch replaced the source but kept the actual old assistant answer. The model corrected the answer and identified it as stale. That branch was already warm from the earlier replacement, so its cache hit is evidence of returning to a computed branch—not of preserving the old source's state through an edit.

The important design issue is that a source and the work based on it are separate objects. Updating a file does not update an earlier summary, plan, or proposed patch. Removing a source does not remove facts copied into those objects. If the operation really means “do not include this old value in the next request,” the harness has to check those copies too.


What pruning and compaction actually change

Truncation shortens an item before or while it enters the context. Pruning removes selected older items. Compaction builds a smaller continuation, often from a summary plus recent messages. They can all reduce token count, but they do different things to the next request.

Before:
  instructions → long build log → assistant diagnosis → recent work

After pruning:
  instructions → [old log cleared] → assistant diagnosis → recent work

After summarizing:
  instructions → summary of the earlier work → recent work

In both rewritten histories, the text of “recent work” may be identical. Its preceding context is different, so unchanged recent messages do not make the old later KV state reusable. The earlier stable instructions may still hit.

OpenCode makes the distinction between stored history and active context easy to see. Its pruning pass stamps an old tool result as compacted. Later, the request converter replaces that result with [Old tool result content cleared] and drops its attachments. The original output string remains in the stored record on that path. Pruning, request conversion.

Oh My Pi also considers how much already-sent conversation follows a pruning candidate. Deleting a small old result can disturb a much larger cached suffix. Its pruning configuration can protect such a result rather than treating every removed token as an immediate saving. When it does change a message, it invalidates local cached token estimates and message conversions too. Provider caching is only one cache that must stay consistent. OMP pruning and local invalidation.

The runnable local examples execute Pi's session-context projection and OMP's pruning functions. Pi drops the old source from active context while retaining a supplied summary that contains its conclusion. OMP removes old tool text while retaining an assistant statement based on it. Neither function has erased the information. That is usually the purpose of summarization, but it matters when the original evidence was wrong or must be removed. Upstream functions and returned messages, recorded output.

Compaction also has a cost of its own. It may pay for a summary and a new cache write now to reduce later reads. Under illustrative prices of 1.25 units per written token and 0.1 per cached-read token, retaining 100,000 tokens costs 10,000 units each turn. Replacing them with 20,000 tokens costs 25,000 units once, then 2,000 per later turn. Over three turns those input costs are 30,000 versus 29,000 units, before paying to generate the summary. A summary that loses the key error can erase those savings by causing another investigation. Calculation and code walkthrough.


Can the engine repair and reuse later chunks?

There are research systems that assemble separately cached chunks or repair only some of the state after a change. Their tradeoff is how closely the reused calculation matches processing the complete new input.

Prompt Cache uses schema-defined modules and positions. Its paper explicitly describes an attention approximation: independent modules omit some dependencies that ordinary concatenated attention would have computed. CacheBlend instead reuses chunks and selectively recomputes tokens to reduce deviation from full recomputation. PIE, designed for code edits, repairs positional effects while treating suffix reuse as an approximation. Prompt Cache, CacheBlend, PIE.

A system can answer a test set just as well while producing different token probabilities. That is weaker than preserving the calculation a fresh request would perform. A position correction is not a recomputation of what the token learned from the old context. Moving or compressing a cache also does not establish that its contents remain valid after an edit.

There are architecture-specific exceptions. A deliberately independent attention mask changes which dependencies exist. Pure local attention can bound their reach, while hybrid or recurrent models require different checkpoint reasoning. None supplies a universal hosted-API operation for editing arbitrary old text while retaining the entire suffix unchanged. Inference report, implementation details, and research limits.

Approximate repair may be useful when a workload can measure and accept its errors. The papers do not establish a general way to replace any earlier text while preserving exactly the computation that a fresh request would perform.


What I would build into a harness

I would make the requested change explicit before choosing how to preserve the cache:

Requested operationCandidate implementationRequired correctness check
Change a preference from this turn onwardSupported authoritative message appended to historyNew behavior applies with the intended scope
Add or change a toolNative tool update where supported; otherwise deliberate rebuildCorrect schema/version and executor permissions
Replace stale evidenceNew version with explicit precedence, or rebuild the relevant contextThe answer uses the intended evidence; old content may still influence appended form
Erase or redact earlier contentRebuild without it and its derived context; address provider retention separatelyNo claim that an appended correction removes prior information
Compact a long sessionSummarize and begin a new context versionTask-relevant facts survive; lower hit rate may still reduce total cost

Stable instructions, deterministic tool ordering, versioned context, and unchanged history avoid accidental misses. Supported prospective controls provide flexibility. Meaningful rewrites should remain visible, even when they cost more.

To diagnose a miss, record the model and serving route, the adapter version, the selected cache boundaries, and a sanitized comparison of consecutive requests. Keep the provider's original usage counters alongside any normalized totals: some APIs include cached tokens in total input and others report separate buckets. Diagnostics can help identify changed instructions or tools, but an eligible request can still miss because its saved state is unavailable. OpenAI diagnostics, Anthropic diagnostics.

For the harness I'm building, the goal is to preserve useful work while making real changes correctly. Appended instructions and old conversation branches can reuse saved work. Replacing evidence or compacting a session may require a new prefix. That cost can be worth paying when it gives the model a shorter, more accurate context.



Glossary

Term / claim Primary source Date
Causal attention and state dependencies Attention Is All You Need 2017
Prefix hash provenance Pinned vLLM implementation Inspected 2026-09-28
Mid-conversation system/tool changes Anthropic contract Checked 2026-09-28
Provider-specific caching controls OpenAI, Anthropic, Gemini Checked 2026-09-28
Modular and selectively repaired reuse Prompt Cache, CacheBlend 2023 / 2024 preprints; versions reviewed in research
Source audits, reproduction and measurements Research and runnable labs 2026-09-28 local / 2026-09-29 UTC
Sources & Evidence
Run and inspect the examples Request construction, returned usage, answers, and limits
Research reading guide
Follow the implementations Pinned revisions for eleven cloned repositories
Source manifest
Check what each result establishes Measurements, source findings, and open questions kept separate
Claim ledger