Pattern · Cost

Cache-Preserving Injection

Split injected memory by how often it changes, so a per-turn recall block cannot invalidate the provider's prompt-prefix cache on every request.

This is a cost pattern, and it is here on a narrow argument. The patterns index declines context-window pruning as a prompt-assembly concern rather than a memory one, and that decision stands. This page is the part of that territory which is not about token budgets: where you inject memory constrains what your memory system is allowed to be, and systems in the atlas have shaped their write paths around that constraint rather than around anything about recall.

Intent

Place injected memory according to how often it changes, so that the stable part of the prompt stays byte-identical between turns and the provider's prefix cache survives the session.

The problem

The obvious place to put memory is the system prompt, near the top, where the model is most likely to attend to it. Every provider that offers prompt caching keys that cache on an exact prefix match. The two facts collide.

A retrieval block is a function of the current message. Put it in the system prompt and the system prompt differs every turn, so the cached prefix misses on every request — and the miss is not partial. Everything from the first differing byte onward is re-processed, which for a memory-augmented agent is usually the whole prompt.

This is invisible in exactly the way that matters. Retrieval quality is unaffected, tests pass, and the only signal is the bill. These reports record a system in this state, and none of the projects appears to have noticed:

  • Helm — "Prompt-prefix caching is invalidated on every turn, by construction." INDEX.md is stable and arrives through an import, but --append-system-prompt carries buildPersona(mode) + recallMemories(prompt), and the recall block is a function of the current message.
  • CSM — the per-turn injection invalidates the prefix cache every turn, and the report notes that nothing measures the cost.
  • RisuAI — the injected block varies per turn, including randomly, which defeats caching by construction.
  • SillyTavern — memory is injected at an unstable position rather than in a stable prefix.
  • OpenCode — no memory of its own, but everything a plugin injects goes through experimental.chat.system.transform into the system prompt, so the contract hands every plugin author this failure.
  • MemoraX Code, OpenMasq, PLUR1BUS and Mnemosyne — per-turn recall injected into the prompt, each recorded in its report as invalidating the prefix cache on every turn whose recalled set changes.

The cost is not a rounding error, and it compounds with the thing memory systems are for. The more memory you inject, the larger the prefix you are invalidating.

The pattern

Sort injected material by volatility and give each class a position that matches:

system prompt (cached prefix)
  ├─ persona, policy, instructions        — changes ~never
  └─ memory INDEX / stable working set    — changes between sessions
─────────────────────────────────── cache boundary
user turn (never cached)
  └─ query-specific recall results        — changes every turn

Three shapes implement this, and they differ in what they give up:

Split by position. Stable memory stays in the prefix; query-specific recall moves into the user turn. Nothing is lost — the model sees the same tokens, in a different envelope. Helm's report notes this arrangement is "one line away" and that the system's own static channel already demonstrates it.

Freeze the snapshot. Render memory into the system prompt once, at session start, and refuse to update it mid-session. Writes still land on disk immediately and durably; they simply do not reach the model until the next session begins.

Split the memory, not the position. Put a one-line index of the store in the prefix and load an item's body on demand. This is the cheapest of the three where the unit is large and individually addressable, and it is described under Why it works below because its cost curve differs from the other two.

Diagram — stable memory is rendered once into the system prompt and turn-varying material is appended to the user turn, so a mid-session write persists to disk without invalidating the prefix cache
Diagram source
%% caption: stable memory is rendered once into the system prompt and turn-varying material is appended to the user turn, so a mid-session write persists to disk without invalidating the prefix cache
flowchart TD
    A["session start"] --> B["render stable memory<br/>into system prompt"]
    B --> C["prefix cached"]
    D["turn N"] --> E{"material varies<br/>with this turn?"}
    E -- "no" --> C
    E -- "yes" --> F["append to the user turn"]
    F --> G["prefix cache still hits"]
    H["mid-session write"] --> I["persist to disk"]
    I -.->|"deliberately not re-rendered"| C

Why it works

Caching is a prefix property, so the only thing that matters is where the first difference falls. Moving the volatile block after every stable byte converts a guaranteed miss into a guaranteed hit without changing what the model reads.

The second-order effect is the interesting one. Once prompt cost is a static, known quantity, a memory budget becomes enforceable — which is what lets Hermes Agent take the position that memory is hard-bounded and frozen, because the prompt cache matters more than completeness. MEMORY.md is capped at 2,200 characters and USER.md at 1,375. When an add would exceed the cap the write is refused, and the tool returns the current entries with an instruction to consolidate and retry within the same turn. Compaction is not a background worker; it is a synchronous obligation handed to the model at the moment of overflow.

That is a write path, a correction path and a capacity policy all derived from a caching constraint. It is why this belongs in a memory pattern library rather than in a prompt-engineering note.

Ollama's built-in agent was the third shape, and the cheapest of the three, until Ollama removed the agent in September 2026. Rather than choosing where to put the memory, it split the memory itself: SkillCatalog.SystemContext() renders one - name: description line per skill into the system prompt, and the body of a skill loads only when something calls the skill tool, arriving as a tool result in the message history. The comment gives the reasoning as a token argument rather than a caching one — "advertises the catalog without expanding full instructions in every request. The skill call is the explicit loading boundary" — and the caching property falls out for free, because the index is computed once at start from a name-sorted list and is therefore byte-identical for the session.

Call it index in the prefix, body after it. It costs one line per item always and the full item only when used, which inverts the usual scaling: a store that grows in item size costs nothing extra, and only a store that grows in item count pushes on the prefix. It applies wherever memories are large, individually addressable and rarely all needed at once — documents, playbooks, runbooks — and not at all where the unit is a one-line fact, since there the index and the body are the same size.

The catch is that it moves retrieval into the model. Nothing here scores a match; the model reads forty descriptions and picks. That is fine for forty and not for four thousand, and it makes the quality of the description the whole retrieval system — which is why Ollama's bundled skill-creator spends most of its length on how to write one.

Tradeoffs

  • Freshness is the price of freezing. Under the frozen-snapshot shape, a memory written at turn 3 is invisible to the model until the next session. Hermes accepts this explicitly. If your agent must act on what it just learned, take the split-by-position shape instead, which costs nothing in freshness.
  • The bounded variant hands editorial control to the model. Hermes's refusal loop makes the model choose what to discard under time pressure, with no review and no record of what was dropped — which is the failure append-only memory audit exists to prevent.
  • Attention position is not free. Material moved out of the system prompt and into the user turn is in a different position, and whether that changes how the model weighs it is not something this atlas has measured.
  • It only pays where caching does. A local model, a provider without prefix caching, or sessions of one or two turns will not repay the restructuring.
  • Splitting has a floor. If the "stable" block is itself regenerated by a nightly consolidation pass, the cache dies once a day rather than once a turn — better, but the boundary should be drawn where the regeneration happens, not where the taxonomy suggests.

Seen in the atlas

  • Hermes Agent — the frozen snapshot, with a capacity policy and a write-refusal path that follow from the cache constraint.
  • Nuum — the memory section frozen per epoch in prompt-cache.json so the system-prompt prefix stays byte-identical between turns. An explicit update_state write bumps the epoch; the extractor that runs after every turn does not, and run.test.ts pins both halves. The cost is the same boundary seen from the store: what the extractor writes or removes changes the files at once and the frozen prompt only at compaction or the next explicit write.
  • Reasonix — split by position, with the cache as the stated design constraint. The cached system prompt carries a memory protocol that is byte-identical across projects plus the pinned fact bodies, and a test fails if the saved-fact index reaches it; relevant facts arrive in a recall block after the user's text, and standing instructions ride the turn. The cost is on correction: a pinned fact forgotten mid-session stays in the prompt until the next session.
  • Hipocampus — size caps on the always-loaded files justified in the spec as a caching decision: "stable content maximizes prompt cache hit rate", with about 3K tokens of ROOT.md per session chosen to sit inside the prompt cache.
  • Memobase — the profile is injected at the front, which the report notes is friendlier to prefix caching than the alternative.
  • Graphify — assembles context in a way the report contrasts explicitly with "invalidating a prefix cache on every request".
  • Ollama — the index/body split, with a sorted catalog line in the prefix and the instruction body arriving as a tool result; removed with the built-in agent in September 2026.
  • Context Mode — the boundary shape by accident of lifecycle: its <session_knowledge> block is emitted once at SessionStart and once at PreCompact and never re-rendered mid-session, so it lands in the cached prefix. Nothing in that tree says this was reasoned about, so it is an arrangement that could silently stop being true if a maintainer added a per-turn refresh.
  • Khabeer — that failure, with the intent written down. A port of Hermes whose specification states the frozen-snapshot rule in so many words, and whose runtime rebuilds the system prompt from the memory files inside every provider request builder — once per tool step on the Anthropic path. Most of the Hermes store came across; the one property this pattern names did not, because the function that renders the block is called systemPromptSnapshot and holds no state.
  • Helm, CSM, RisuAI, SillyTavern — the counter-examples, each invalidating on every turn.
  • OpenCode — the contract-level version: a host whose only injection seam is the cache-sensitive one.
  • Lobu — the deliberate case and its own counterexample, in one product. Org guidance is injected into every managed agent's system prompt under a hard 3,000-character cap, ordered by event id with no timestamps or counts and a deterministic truncation marker, explicitly so that it "counts as stable data" inside the cached prefix. The per-turn memory plugin does the opposite, prepending a freshly recalled memory block ahead of the user message on every turn; the turn machinery keys transient context per message rather than per turn for exactly this reason, naming "a prompt-cache prefix that changes on each call" as the failure.
  • dsh-continual-harness — the session-stable anchor as the default, with a measured reason and a visible price. The block is ranked against the session's opening request and cwd, republished only on a first injection, a new refinement, an emptied store or a new matched key, and replaced in place; the config comment reports a 95.4% against 81.2% cache-read hit rate over one six-message session. To stay byte-stable the default block drops entry content and lists 15 ids, under a header telling the model to read entries on demand, which no plugin tool can do.
  • Orgtree — two cache decisions about one memory. The breadcrumb tail is spliced into a cheap-compacted successor's system prompt for its first turn only, and _retire_breadcrumb_splice clears the marker at the first successful result, because re-splicing a file the agent appends to every turn would rewrite the prefix each turn; the docstring puts the stake at roughly 24% against 61% of cold starts prevented. The standing notes go the other way: they sit in the hashed identity, so an edit deliberately forces a cold respawn rather than letting a warm process serve the old notes. The cost is that turn two onward sees the breadcrumbs only through conversation history.

Note that MemOS's "activation memory: KV/prefix cache" is a different mechanism — reusing model state rather than positioning text — and falls under the KV-cache scope boundary in the comparative report.

Signet AI splits the session-start payload rather than the memory design. Its session-start handler returns a stable system prompt separately from a dynamic context block, with identity files selected by preset and budgeted per file. A second session-start for the same session hits a dedupe guard and returns a minimal stub rather than re-injecting, on the stated ground that the identity files are already in the context. The invariant block sits ahead of the variable one, and per-turn recall arrives through the prompt-submit hook and the MCP tools instead of being woven into the system prompt.

The same constraint, one layer down

Everything above is about where a memory system puts its recall block. The same prefix property is being defended one layer below, at the proxy, by projects with no memory at all — and one of them has written the argument out more carefully than anything in this atlas.

llmtrim is an MPL-2.0 compression proxy that sits between a coding agent and its provider, examined on 2026-08-09 at d7fd2c4e…. It gets no report here, because its store is in-process only, size-capped and explicitly never written to disk — nothing survives the session, which is the inclusion bar. crates/llmtrim-core/src/memo.rs matters to this page for three reasons.

It names the failure precisely, and the failure is caused by compression itself. Its stages are deterministic per request but read the whole conversation, so "the compressed form of an old message can change when a new turn arrives. Two consecutive turns then serialize a divergent prefix → the provider cache is busted → the product's headline savings leak silently on exactly the highest-traffic (agent) shape." That is the same silence the top of this page describes — quality unaffected, tests pass, only the bill moves — found in a system whose entire purpose is reducing the bill.

It supplies the number this page otherwise cannot. The module's opening cites a 2026 measurement putting 85–95% of an agentic request's prompt tokens at unchanged turn-to-turn. That is a claim about agent traffic rather than a cache hit rate, and it is the reason the boundary's position matters so much for exactly the shape memory systems serve.

Its fix is the split-by-position shape applied to history rather than to memory. A cumulative 128-bit hash chain over the original bytes fingerprints each message prefix; the longest run whose boundaries are all present in the store is the frozen prefix, and its slots in the compressed output are overwritten with the bytes emitted last turn. Appending a turn leaves every earlier boundary hash unchanged; editing one byte of an old message invalidates that boundary and every one after it — which is the prefix-cache semantics restated as a data structure. Writes are first-write-wins, so a frozen message is never re-mutated.

Two details are worth stealing regardless of layer. The memo is "an optimization, never a correctness dependency" — any doubt (a restructured message array, a cold prefix, an unexpected count delta) falls back to full recomputation. And it carves out the one stage where splicing is unsound: the n-gram stage assigns placeholders from frequencies across the whole conversation, so reuse is disabled whenever that stage is on, rather than reaching into the stage to freeze its dictionary. A caching optimisation that knows which of its own components it cannot safely apply to is rarer than it should be.

A second proxy shows the other half of the trade. headroom, examined at 675d13f0…, implements what it calls CCR — Compress-Cache-Retrieve — where a dropped payload is stashed in SQLite keyed by the hash that goes into the prompt, and a retrieval tool call trades the hash back for the original: "lossy on the wire, lossless end-to-end." CCR is a dereference table rather than a memory. The Headroom report covers the repository's separate cross-agent memory subsystem, and places CCR in the compression product beside a CacheAligner that flags content which would bust a provider's prefix cache and never rewrites the prompt. The CCR shape is the one a memory system reaches for when a recalled item is too big to inject — put the handle in the prefix and let the model ask for the body, which is what Ollama's agent did above with a name instead of a hash.

Tests before relying on it

  • Capture the exact bytes of two consecutive requests and diff them. The first differing byte is the cache boundary; assert it falls after the stable block.
  • Assert a mid-session memory write does not change the system prompt, under the frozen shape — and that it does reach disk.
  • Assert the stable block is byte-identical across turns, including ordering of anything assembled from a set or a dict.
  • Under the bounded variant, assert a write that exceeds the cap is refused and returns the current entries, rather than silently truncating.
  • Measure it. Every claim on this page is about a mechanism read in code, and the one hit rate quoted on it, dsh-continual-harness's, comes from a config comment over a corpus that project's tree excludes.