1. Executive Summary
ostk-recall is roughly 90,000 lines of Rust across ten crates, dual-licensed MIT and Apache-2.0, shipping one binary that indexes notes, source trees and assistant session logs into a local corpus and serves them over MCP. Data stays on the machine. The README describes it as pre-alpha and says the maintainer runs it daily.
Two stores sit behind it. LanceDB holds chunk vectors and a Tantivy BM25 index; SQLite holds everything with an identity — the claim table, the concept ledger, the audit tables and an append-only access chain log. That split is the design's spine: the corpus is bulk and rebuildable, and the ledger is the part that can be wrong.
It carries five of this atlas's seven marks. The two it does not carry are the report.
forget tells the caller it wrote a tombstone,
and it did not. The MCP handler returns the warning "claim
suppressed with an anti-resurrection tombstone; hard purge was not
performed", and the call behind it is
transition_claim(id, ClaimState::Suppressed, …) — keyed on
the claim's autoincrement id. Three things follow, each of them
checkable in one file. record_claim inserts
unconditionally; its only pre-write lookup is a
memory_mutation_receipts idempotency check, which guards a
replayed request rather than a re-asserted value.
Conflict detection selects
WHERE project=? AND claim_key=? AND conflict_eligible=1 AND state IN ('active','disputed'),
so the suppressed row is invisible to the one mechanism that would
notice its value coming back. And reads default to the same two states.
So assert the fact again and a fresh row lands in active
and is recalled, with the suppressed row sitting beside it
unconsulted.
That is read-side suppression, which is a real protection for the
agent and no protection for the store. The mark is withheld on the
atlas's definition — a tombstone is keyed on the value — and
the near-miss is worth stating precisely because the vocabulary is
exactly right and the mechanism is not. Adding it is cheap here in a way
it is not elsewhere: claim_key, subject,
predicate and value_json are already columns,
so a digest over the value already exists in all but name.
And the only thing making forget mean anything
is untested. The read filter is what stops a suppressed claim
being recalled, and no committed case asserts that it holds.
split_is_atomic_replayable_and_keep_false_suppresses_parent
asserts the parent's state, not its absence from a result.
Against that, three mechanisms here are ones the atlas has not found elsewhere, and they are all in the concept ledger rather than the claim table.
2. Mental Model
A thing becomes a belief here along one of two paths that never merge.
Ingested material is chunked, embedded and indexed.
It has provenance and a project, and it is bulk: nothing
about it is asserted, and rebuilding the corpus from source loses
nothing.
A claim is asserted, by an agent through
remember or by promotion, and it is the part with an
epistemic life. It enters active. If another claim shares
its project and claim_key and both are
conflict_eligible, an automatic detector opens a conflict
and moves both to disputed — no human in the loop, and the
reverse transition (conflict_cleared) is automatic too. If
every support behind a claim is invalidated it becomes
unsupported. A human or agent can supersede it, retract it,
or forget it.
The state is a field, not a score, and a separate
confidence float lives beside it — the distinction the trust-state machine
pattern argues for. is_current() is
Active | Disputed, and that pair is what every default read
filters to.
The gap is on the write side, and the diagram is drawn on it.
Diagram source
%% caption: reads and conflict detection filter the same way, so a suppressed claim is neither recalled nor compared against — and asserting the same value again lands in active under a new id
stateDiagram-v2
[*] --> active: record_claim inserts unconditionally
active --> disputed: conflict detector on project + claim_key
disputed --> active: conflict_cleared
active --> superseded: supersede_claim sets superseded_by
active --> unsupported: last support invalidated
active --> retracted: remember/retract
active --> suppressed: remember/forget
suppressed --> active: remember/restore
note right of suppressed
Reads filter to active + disputed, so a suppressed
claim stops being recalled.
Conflict detection filters the same way, so nothing
compares a new assertion against it.
Assert the same value again and it lands in active
under a new id.
end note3. Architecture
One binary, no services. crates/store owns SQLite and
every table with an identity; crates/query owns LanceDB,
the hybrid lanes and the diffusion walk; crates/pipeline
owns scan, chunk and embed; crates/attention is a runtime
of timer- and event-driven passes; crates/mcp and
crates/attention-mcp are the agent-facing surfaces, served
over stdio or a shared local socket.
Embeddings are static (model2vec), which is the decision that makes the whole thing a single binary with no GPU and no service to stand up. The optional cross-encoder reranker is fastembed. An operator's cost is a config file and a scan; there is nothing to operate beyond the daemon.
4. Essential Implementation Paths
- Claim write —
crates/store/src/claims.rs:record_claim→ idempotency receipt lookup →insert_claim→ support rows →insert_event, all in one transaction. - Conflict detection — same file: the
project/claim_key/conflict_eligiblequery, opening a conflict and moving members todisputed, or clearing it and writing aconflict_clearedevent. - State transitions —
transition_claimandsupersede_claim, both guarded on the prior state beingactiveordisputedand both bumpingrevision. - Hybrid retrieval —
crates/query/src/hybrid.rs: filter construction withsql_escape, dense and BM25 lanes, RRF fusion. - Edge conductance —
crates/store/src/activation.rs:conductance_of. - Attention passes —
crates/attention/src/{observer,weaver,curator}.rs.
5. Memory Data Model
memory_claims carries id,
project, kind, claim_key,
subject, predicate, value_json,
text, polarity, state,
origin, actor, confidence,
valid_from, valid_to,
superseded_by, revision,
conflict_eligible, created_at and
updated_at, indexed on
(project, claim_key, state).
Two things there are worth naming.
valid_from/valid_to beside
created_at/updated_at is genuine
bi-temporality, and the part that earns the mark is what the validity
interval decides rather than that it exists.
recompute_conflict pulls every current claim sharing a key
and pairs them through
intervals_overlap(a.valid_from, a.valid_to, b.valid_from, b.valid_to);
two claims with the same key and incompatible values are only a
contradiction if their validity windows meet. "I lived in Berlin" and "I
live in Lisbon" do not fight when the first one ended. Record time is
untouched by that test and carries the revision history instead.
And polarity lets a claim assert that something is
not the case, which is rarer than it should be — most stores in
this atlas can only record presence.
The concept ledger is the other half: concepts,
concept_edges, concept_aliases,
concept_evidence, concept_notes. Every edge
records an origin of authored, observed or promoted.
6. Retrieval Mechanics
model2vec dense vectors and Tantivy BM25, fused by reciprocal rank
fusion, with an optional cross-encoder rerank pass. project
is compiled into the LanceDB filter predicate —
stale = false AND project = 'a''b' in the committed test,
with sql_escape doing the quoting — which is what earns
scope_enforced.
The filter can be widened by source:
project = 'p' OR source = 'code' is a shape the tests
exercise. That does not cost the mark, which certifies the key reaches
the query rather than that a caller cannot pass a different argument,
but it is worth knowing before treating project as a
boundary.
Diffusion walks both halves of the concept graph — the reified edges and the latent vector-similarity neighbourhood. An off-diagonal bridge walked during consolidation is promoted into a weak reified edge, which then has to earn its conductance through use or decay away.
7. Write Mechanics
record_claim is synchronous and transactional: the
claim, its support rows, its concept links and its event all land or
none do, and the caller gets the claim back. There is no queue in front
of it, so a claim is retrievable through the claim reads
immediately.
Embeddings are the asynchronous part.
upsert_claim_embeddings is a separate model-scoped call —
"replays are idempotent and model changes replace the old coordinate
atomically" — so a claim exists before its vector does, and the lag
before it is reachable through vector recall is the lag of
whatever drives that call.
Ingest is a different path entirely: scan, chunk, embed, upsert into LanceDB, with a manifest in SQLite. The orphan sweep tombstones ingest chunks by path — that is the ordinary record-keyed kind, and unrelated to the claim story above.
A claim can be decomposed rather than replaced, and the
decomposition has a shape the store enforces.
split_claim creates bounded child claims and links each one
back with a part_of edge carrying a
sequence_index and a sequence_total in its
evidence; the children are additionally chained to each other by
continues links. The parent stays current by default, and
when keep_parent is false it is moved to
suppressed rather than superseded, under a comment giving
the reason: "no single child is misrepresented as its
superseded_by successor." That is a distinction most
systems here collapse — a claim broken into parts has no single
successor, and saying it does would make the supersession chain lie.
claim_continuity reads the structure back and refuses
anything malformed: more than one ordered part_of parent is
an error, and so is a branching continues topology. A
decomposition is a list, and the reader will not accept a graph
pretending to be one.
No background pass rewrites the claim store. The curator fades threads, and the docstring is explicit that this is not forgetting: "the substrate doesn't forget, but the surfacer stops shouting about threads whose score has fallen below the archive line." Hysteresis around each threshold keeps a thread near a boundary from flipping every tick.
8. Agent Integration
Two tools — recall and remember — with
historical names kept as hidden aliases for one transition cycle.
remember carries eleven verbs: record, supersede, retract,
forget, restore, resolve, relate, split, focus, track and consolidate.
Beside them is a resources surface including an ambient "memory lens"
aligned to the current attention vector, which is an unusual thing to
expose: a resource whose content tracks what the agent is attending to
rather than what it asked for.
Served over stdio or a shared local-socket daemon, so several clients can share one index.
9. Reliability, Safety, and Trust
The audit surface is genuinely strong.
memory_claim_events and
memory_claim_link_events are documented append-only
histories, mutations carry receipts in
memory_mutation_receipts keyed by idempotency key with
payload comparison — a replay with a different payload is
rejected rather than silently accepted — and chain_log is
an indexed access ledger. Every state-changing verb takes an actor and a
reason.
resolve_conflict_with_receipt records a
resolution_kind and a resolution_reason
alongside the actor, which is what earns human_review: a
person adjudicates a conflict the detector opened, and the adjudication
is itself a durable record rather than a mutation.
The weaknesses are the two withheld marks, and they compound. The
forget warning asserts an anti-resurrection property, the
mechanism behind it is a status flip on one row, and nothing committed
asserts even the read-side half holds. A reader who takes the warning at
face value will believe a value cannot come back when it can, and the
test suite will not tell them otherwise.
10. Tests, Evals, and Benchmarks
1,054 #[test] and #[tokio::test] functions
across the crates over 93,936 lines of Rust, plus a tests/
directory with fixtures and a queries.yaml. Nothing was run
for this review.
The suite is dense where the design is careful.
claim_link_lifecycle_is_idempotent_audited_and_scoped
exercises the three properties in its name at once.
build_filter_project_and_source pins the scope predicate
including SQL escaping.
unstructured_notes_are_not_conflict_eligible pins the gate
that keeps free text out of the conflict detector.
orphan_marking_gates_not_deletes and
ensure_is_idempotent_and_never_downgrades are both the
shape of assertion this atlas asks for.
What is missing is the one the design most needs: no committed case
asserts that a suppressed or retracted claim is absent from a recall
result. The filter is state IN ('active','disputed'),
repeated in eight separate queries in claims.rs, six of
them behind an (? OR ...) override flag — and a ninth query
written without the predicate would pass every test here. The number is
the argument: a rule copied by hand into eight places, guarded by none
of them, is one refactor away from a suppressed claim coming back
through whichever copy was missed.
11. For Your Own Build
Steal
- Derive the edge weight instead of storing it.
conductance_of(confidence, last_seen_at, now)isconfidence × recency_lift(…), clamped. Every other decay-and-reinforcement implementation in this atlas stores a mutable score and updates it on a schedule; deriving it means there is no weight to drift out of step with its inputs, no migration when the formula changes, and no background pass whose failure silently freezes the graph. - Make a promoted edge earn its place. A bridge found in the latent neighbourhood is reified as a weak edge that must earn conductance through use or decay away. That is evidence-before-belief applied to graph structure: the promotion is visible, provisional, and reversible by inaction.
- Put the origin on the edge.
authored/observed/promotedanswers "why does this link exist" without a join. - Compare the payload on an idempotency replay. Accepting a replayed key with a different body is the failure this guards, and most implementations do not.
- Say what fading is. "The substrate doesn't forget, but the surfacer stops shouting" is the clearest one-line statement of decay-as-ranking in this atlas, and the hysteresis beside it is the detail most implementations skip.
Avoid
- A warning that asserts a property the code does not have. The cost is not the missing tombstone — plenty of systems here lack one. It is that the string tells the caller the value cannot come back, which is worse than silence, because it forecloses the question.
- A read-path filter with no test. Four queries carry
state IN ('active','disputed')and an override flag; nothing asserts the fifth will.
Fit
Take this if you want one local binary over your own files and sessions, you are comfortable on a pre-alpha daily driver, and the concept ledger is the part you actually want — that is where the original thinking is. It suits a single operator with one machine and several MCP clients sharing a socket.
Walk away if you need a correction to hold against re-extraction. Everything else here is careful, which makes the gap easy to miss: the states are rich, the audit is real, the scoping reaches the query, and none of that stops a forgotten value returning under a new id on the next ingest.
12. Open Questions
- Does anything drive
upsert_claim_embeddingson therememberpath, or is a claim reachable only by claim reads until an ingest pass runs? - The
polaritycolumn can express a negative claim. Does the conflict detector treat a positive and a negative claim on the sameclaim_keyas conflicting, which is the case that would make polarity load-bearing? conflict_eligibledefaults to0. What proportion of real claims are eligible, given that ineligible ones never reachdisputed?
Appendix: File Index
| Path | Role |
|---|---|
crates/store/src/claims.rs |
Claim table, states, conflicts, audit events, mutation receipts |
crates/store/src/concepts.rs |
Concept ledger, aliases, merge and canonicalization |
crates/store/src/activation.rs |
conductance_of, activation reads, promoted-edge
audit |
crates/store/src/threads.rs |
Threads, thread links, chain_log |
crates/store/src/events.rs |
audit_events |
crates/query/src/hybrid.rs |
Filter construction, dense and BM25 lanes, RRF |
crates/query/src/lanes.rs |
Lane execution and the scope vector |
crates/attention/src/curator.rs |
Idle fade with hysteresis |
crates/attention/src/weaver.rs |
Auto-weaver, ProposedWeave |
crates/pipeline/src/lib.rs |
Scan, chunk, embed, orphan sweep |
crates/mcp/src/claims.rs |
remember verbs and the soft_forget
warning |
Appendix: Recorded Searches
Run from the root of the checkout at the pinned commit.
| Claim | Command | Result at this pin |
|---|---|---|
| Nothing consults the suppression before a write | read record_claim at
crates/store/src/claims.rs:772, then
grep -n "claim_key" crates/store/src/claims.rs |
The only pre-insert lookup is the idempotency receipt keyed on the
request; claim_key appears in the conflict path only, which
filters to state IN ('active','disputed') |
| The state filter is repeated, not centralised | grep -rn "state IN ('active','disputed')" --include="*.rs" crates |
Eight occurrences, all in claims.rs, six behind an
(? OR ...) override |
| No test asserts a suppressed claim is absent from a recall | grep -rn "fn .*suppress|fn .*retract|fn .*forget" --include="*.rs" crates |
Four test functions, none of them a recall-exclusion case |
| Validity intervals decide contradiction | grep -n "intervals_overlap" crates/store/src/claims.rs |
Used in recompute_conflict to pair same-key claims |
| Seven claim states | sed -n '/pub enum ClaimState/,/^}/p' crates/store/src/claims.rs |
Seven variants; is_current() admits two |
| Test and tree size | grep -rc "#\[test\]|#\[tokio::test\]" --include="*.rs" crates
summed, and find . -name "*.rs" | xargs wc -l |
1,054 test functions over 93,936 lines |
History
2026-09-11 — 4c75f920…
— re-read, 36 files and 4,672 insertions past the previous pin, most of
it in crates/store/src/claims.rs and
crates/pipeline/src/lib.rs. The headline finding is
unchanged and re-verified: record_claim still
inserts without consulting the suppression,
recompute_conflict still filters to
state IN ('active','disputed'), and no committed case
asserts a suppressed claim is absent from a recall result. The criticism
sharpened on one axis — the read filter is repeated in eight queries
rather than the four the first reading counted, six of them behind an
override flag, none of them guarded by a test. One first-reading
error, in the report's own favour to correct:
ClaimState has seven variants, not six;
Expired was present at both pins and was missed. The
bitemporal evidence was thin and is now specific:
intervals_overlap gates whether two same-key claims
contradict at all, so the validity window decides an outcome rather than
merely being stored beside record time. New since the first
reading: split_claim and
claim_continuity — a claim decomposed into ordered children
linked part_of with a sequence index and chained by
continues, the parent suppressed rather than superseded
because no single child is its successor, and a reader that refuses a
branching topology; lens_candidate_claims and
claim_lens_generation behind the memory-lens resource; and
backfill_transcript_projections in the pipeline.
remember now carries eleven verbs. Test count 1,028 to
1,054 over 93,936 lines of Rust. Marks unchanged at five. Screened
before reading: one build-time exec path (Makefile), no
auto-run surface, the lockfile unchanged for 29 days; nothing was built
or run.
2026-08-08 — 5f25e844…
— first reading, at the v0.9.3 release commit. Screened
before reading: 0 auto-run surfaces, 1 build-time exec
path (Makefile), and 3 dependency surfaces changed inside
the seven-day cooldown — Cargo.lock,
Cargo.toml and crates/pipeline/Cargo.toml, all
changed the same day, because the release landed the same day.
Nothing was executed, and every claim here is
established by reading.
The claim lifecycle was traced end to end: record_claim
and its idempotency receipt, insert_claim, the conflict
detector's (project, claim_key, conflict_eligible) query
and the disputed transitions it drives in both directions,
supersede_claim, and transition_claim behind
retract, forget and restore. Marks: trust_state for the
six-value state with is_current() and a
separate confidence float; bitemporal for
valid_from/valid_to beside
created_at/updated_at;
scope_enforced for the project predicate
compiled into the LanceDB filter with sql_escape, pinned by
build_filter_project_and_source; audit_log for
the documented append-only memory_claim_events and
memory_claim_link_events with receipts beside them; and
human_review for resolve_conflict_with_receipt
carrying actor, reason and resolution kind.
tombstone is withheld and the near-miss is the report.
The remember/forget handler returns "claim suppressed
with an anti-resurrection tombstone", and the suppression is keyed
on the claim's autoincrement id: record_claim inserts
without consulting it, and conflict detection filters to
state IN ('active','disputed'), so the suppressed row
cannot see its own value return. negative_eval is withheld
for the matching reason — the read filter is the only thing the warning
rests on, and nothing committed asserts it holds.