Forgetting that fuses rather than deletes

MemoryBear

Low-activation memories are merged into a summary node that keeps DERIVED_FROM edges to what it replaced — and each forgetting cycle records how many fusions failed.

Carries 2 of 7 rubric mechanisms. Most systems here carry none or one (41%), and a dash means the mechanism was not found at this commit — not that the system needed it. Each mark is one LLM reviewer's reading of the code at this commit rather than a run of it — known limits.

  • Tombstone
  • Trust state
  • Bi-temporal
  • Scope enforced
  • Mutation audit
  • Human review
  • Negative evals

1. Executive Summary

MemoryBear is a memory service built on Neo4j and FastAPI with a web console — "Perceive · Extract · Associate · Forget" — Apache-2.0, roughly 410,000 lines across an API, a web app and sandbox infrastructure, bilingual in English and Chinese throughout.

The mechanism worth the report is what its forgetting does instead of deleting.

forgetting_engine/actr_calculator.py implements a real ACT-R base-level activation model, with the formula written out in the header and Anderson (2007) cited:

R(i) = offset + (1-offset) * exp(-λ*t / Σ(I·t_k^(-d)))

where offset is a "minimum retention rate (prevents complete forgetting)", I is importance, t_k is time since the k-th access and d is a decay constant around 0.5. Recency and frequency combine into one activation value rather than being separate ranking terms.

When activation falls below a threshold, forgetting_strategy.py does not delete. It identifies low-activation Statement–Entity node pairs, fuses them into a MemorySummary node, and wires DERIVED_FROM edges from the sources to the summary before removing the originals — the module's own list of responsibilities ends with "保留溯源信息并删除原始节点" (preserve provenance information and delete the original nodes), and step 4 of the fusion is "溯源保留: 记录原始节点 ID,保持可追溯性" (provenance retention: record the original node ids, maintain traceability).

So forgetting here is lossy compression with a typed edge back to what was lost, and the summary keeps original_statement_id and original_entity_id. A reader who later asks "what happened to this fact" gets a summary and a pointer rather than an absence — which is the property most decay-based systems in this atlas give up.

And the forgetting cycle audits itself. forgetting_cycle_history records, per run and per user: merged_count, failed_count, average_activation_value, total_nodes, low_activation_nodes, duration_seconds and trigger_type (manual or scheduled), indexed on (end_user_id, execution_time).

failed_count is the column to notice. A forgetting pass that records the node pairs it could not fuse is admitting that fusion can fail and making the failure countable — most background passes in this corpus report what they did and are silent about what they could not do.

2. Mental Model

A conversation is perceived, extracted into Statements and Entities, and associated into a Neo4j graph. Every node and every relationship carries end_user_id.

Access history accumulates per node (access_history_manager.py), activation is computed from it, and the scheduler runs cycles that fuse the bottom of the distribution.

Diagram — ACT-R activation decides what falls below the forgetting threshold, and fusion writes a summary node with DERIVED_FROM edges before deleting the originals — with each cycle's counts recorded
Diagram source
%% caption: ACT-R activation decides what falls below the forgetting threshold, and fusion writes a summary node with DERIVED_FROM edges before deleting the originals — with each cycle's counts recorded
flowchart TD
    P["perceive: conversation in"] --> X["extract: Statements + Entities"]
    X --> A["associate: graph edges, end_user_id stamped"]
    A --> AH["access history per node"]
    AH --> ACTR["ACT-R activation:<br/>offset + (1-offset)·exp(-λt / Σ I·t_k^-d)"]
    ACTR --> LOW{"below the forgetting threshold?"}
    LOW -->|no| KEEP["retained; offset floors the decay"]
    LOW -->|yes| PAIR["identify Statement-Entity pairs"]
    PAIR --> FUSE["fuse into a MemorySummary node,<br/>optional LLM-written summary"]
    FUSE --> DF["MERGE (source)-[:DERIVED_FROM]->(ms)<br/>original ids recorded"]
    DF --> DEL["originals deleted"]
    FUSE -.->|"fusion fails"| FC["counted in failed_count"]
    DEL --> H["forgetting_cycle_history row:<br/>merged, failed, avg activation, duration, trigger"]

The offset term is worth its own sentence. Because activation is floored, a memory never decays to zero — the curve asymptotes to a minimum retention rate. Combined with fusion rather than deletion, the design's position is that nothing is ever entirely gone, only progressively coarser.

3. Architecture

Neo4j holds the graph; Postgres holds configuration, cycle history and operational state; Redis caches. FastAPI serves the API, a React web console sits over it, and an e2b-infra directory plus a sandbox directory support sandboxed execution. Docker Compose is the documented deployment.

The forgetting engine is its own package with nine modules — the ACT-R calculator, an access-history manager, a memory-strength module, a strategy, a scheduler, a service and configuration utilities — which is more separation than most decay implementations here get, and it is why the strategy can be read independently of the maths.

Configuration is loaded from the database (load_actr_config_from_db), so the decay constant, forgetting rate and threshold are operator-tunable per deployment rather than compiled in.

4. Essential Implementation Paths

Activationforgetting_engine/actr_calculator.py, with access_history_manager.py supplying the t_k series.

Cyclememory_forget_service.pyForgettingSchedulerForgettingStrategy.identify → fuse → ForgettingCycleHistoryRepository.

Nothing starts that chain on a clock. The entry point is memory_forget_controller.py:59, an HTTP trigger_forgetting_cycle calling the service at :109; the periodic trigger that would call it unattended is commented out in three places. tasks.py:4539-4574 holds the whole run_forgetting_cycle_task definition behind #, including its trigger_forgetting_cycle call. celery_app.py:283-285 comments out the beat entry that would schedule it. And celery_app.py:153 comments out its queue route with the note "已废弃,保留路由防 unregistered" — deprecated, the route kept only so an unregistered-task error cannot fire.

What survives is the configuration around the hole. celery_app.py:240 still builds forgetting_cycle_schedule = crontab(hour=settings.FORGETTING_CYCLE_HOUR, minute=settings.FORGETTING_CYCLE_MINUTE) at import, and config.py:414-417 still parses and validates both environment variables with a default of 18:00. The crontab is constructed on every start and referenced by nothing. An operator setting FORGETTING_CYCLE_HOUR gets no error, no warning and no cycle.

ForgettingScheduler itself is not the dead part — memory_forget_service.py still uses it inside a triggered run to decide which users are due. The class that decides when is live; the thing that would have asked it is commented out. So the trigger_type column, which distinguishes manual from scheduled, can only be written with one of its two values.

Fusionforgetting_strategy.py:355-395: the Cypher OPTIONAL MATCH over inbound relationships and MERGE (source)-[:DERIVED_FROM]->(ms) for both the statement and the entity side, so edges into the originals are rerouted to the summary rather than orphaned.

Scope — the predicate lives in the search templates, not in the connector. cypher_queries.py holds one full-text and one embedding template per node type, and FULLTEXT_QUERY_CYPHER_MAPPING picks the one search_by_fulltext executes. Six of the seven AND the key in unconditionally, in the form

CALL db.index.fulltext.queryNodes("statementsFulltext", $query) YIELD node AS s, score
WHERE s.end_user_id = $end_user_id
  AND s.delete_at IS NULL

neo4j_connector.py:219 and :232 carry the same predicate over nodes and relationships, but they are the body of delete_group(end_user_id) — a scoped deletion, not a read filter — and create_indexes.py builds indexes that require the property.

The seventh template is the one to watch. SEARCH_DIALOGUE_BY_FULLTEXT is written permissively:

WHERE ($end_user_id IS NULL OR d.end_user_id = $end_user_id)

A null scope key does not fail there; it matches every user's dialogue. The question is whether null can arrive, and today it cannot: search_graph declares end_user_id: Optional[str] = None and documents it as an "Optional group filter", but the live caller is ContentSearch._keyword_search, which passes self.ctx.end_user_id from a MemoryContext — a Pydantic BaseModel whose end_user_id: str is required and validated at construction. ContentSearch also puts Neo4jNodeType.DIALOGUE in its default includes, so this template does run on the ordinary path; it is simply never handed a null.

That is a real guarantee and it is in the wrong place. The isolation of the dialogue store depends on a model definition two layers up rather than on the query, and search_graph's own optional-with-default signature is an invitation to call it from somewhere that has no MemoryContext. The other six templates would return nothing in that case. This one would return everything.

5. Memory Data Model

The graph is the model: Statement and Entity nodes with a MemorySummary tier above them, DERIVED_FROM edges recording fusion lineage, and end_user_id on everything.

forgetting_cycle_history is the relational side and it is the more interesting table, described in section 1. Recording average_activation_value per cycle means an operator can watch the distribution move over time — whether the store is drifting toward everything being cold, which is the failure mode a decay system needs to detect and almost none instruments.

trigger_type distinguishing manual from scheduled matters for reading that history: a manually-triggered cycle during a demo and a nightly one produce very different numbers, and separating them is one column.

6. Retrieval Mechanics

Graph traversal plus vector search, with activation available as a ranking signal, and a "vector version (non-graph)" mode described in the README as trading some accuracy for latency.

Scope is enforced and it is enforced everywhere, which is the correct answer for a multi-user service on a single graph database: end_user_id is a predicate in the connector's queries, a required property in the index definitions, and the key of the delete path (add_nodes.py:16 deletes a user's entire subgraph with MATCH (n {end_user_id: ...}) DETACH DELETE n).

That last line is worth flagging as a hazard as well as a feature: it is an f-string interpolating end_user_id directly into Cypher rather than binding a parameter, in a function that deletes everything matching. The surrounding calls bind parameters properly; this one does not.

7. Write Mechanics

Extraction is an LLM pipeline behind the API, so writes are not instantaneous and the graph lags the conversation.

Correction is not a first-class operation. There is no supersession pointer, no contradiction detection surfaced in the forgetting engine, and no rejected-value record. What exists is decay plus fusion: a memory that stops being reinforced becomes part of a summary. A memory that is wrong and frequently accessed will be reinforced and retained, which is the standard weakness of usage-driven retention and is worth stating plainly for a system whose whole lifecycle is activation-driven.

8. Agent Integration

A FastAPI service with an MCP surface, a web console, a sandbox tier and Docker Compose. The console is where the forgetting curve and the cycle history are surfaced — the service layer exposes "遗忘曲线生成" (forgetting-curve generation) as an API concern, so an operator can see the decay model they configured.

9. Reliability, Safety, and Trust

Scope — awarded, per section 6, with the interpolation caveat.

Audit log — awarded, and scoped precisely. forgetting_cycle_history is an append-only per-run record with counts, an average, a duration and a trigger type. ForgettingCycleHistoryRepository exposes create, create_async and three getters and nothing that updates or deletes, and no UPDATE or DELETE against the table exists anywhere outside the migration that creates it; memory_forget_service.py:404 writes the row and :741 reads it back. It audits the forgetting subsystem, not every mutation — an edit or an association does not appear — so it is a background-pass ledger rather than a full mutation log, and it is a good one.

Trust state — no. importance feeds activation; nothing records belief.

Tombstone — no. The DERIVED_FROM edge is lineage, not refusal: the same statement can be re-extracted and will enter as a new node.

Bitemporal, human review, negative eval — no on what was inspected.

Two cautions. The Cypher interpolation in the user-delete path noted above. And the offset floor means a memory can never decay out entirely — which is the design's intent, and it also means an operator who wants something gone must delete rather than wait, and the fusion path is not a deletion path.

10. Tests, Evals, and Benchmarks

Three papers are cited, which is more than most systems here have: a core technical report hosted on the project's own site, a multimodal affective memory engine report (arXiv:2603.22306), and A-MBER, an affective memory benchmark (arXiv:2604.07017) — which the atlas's own benchmarks page already tracks. Publishing a benchmark dataset alongside the system it measures is a real contribution and also a conflict of interest worth naming.

The benchmark claims cannot be checked from this tree. The README reports F1, BLEU-1 and LLM-as-a-Judge scores, states that MemoryBear "consistently outperforms competing systems including Mem0, Zep, and LangMem across all four task categories", and gives 72.90 ± 0.19% for the vector version and 75.00 ± 0.20% for the graph version. Every one of those figures is inside a PNG. There is no harness, no result file, no dataset reference and no run configuration in the repository — the numbers are images with error bars, and an error bar in an image is not reproducible.

That is a different failure from the three systems in earlier batches whose benchmarks live in a sibling repository: there, a reader can go and look. Here there is nowhere to go.

14 test files against 410,000 lines. I ran nothing.

11. For Your Own Build

Steal

  • Fuse instead of deleting, and keep the edge. A MemorySummary with DERIVED_FROM edges from what it replaced, holding the original ids, turns forgetting from an absence into a coarser record. "What happened to this fact" stays answerable.
  • Reroute inbound edges to the summary. The OPTIONAL MATCH plus MERGE is what stops fusion orphaning the graph around the nodes it removed.
  • Count what your background pass could not do. failed_count beside merged_count is one column and it is the difference between a pass that reports its work and one that reports its success rate.
  • Record the average activation per cycle. It is how you notice a store drifting cold before every memory is a summary.
  • Separate manual from scheduled runs in the history. Otherwise a demo contaminates the trend.
  • Floor the decay curve. An offset term means activation asymptotes to a minimum rather than reaching zero, which is a deliberate position on whether forgetting should ever be total.
  • Load the decay parameters from the database. Decay constants are exactly the thing an operator needs to tune per deployment without a rebuild.
  • Give the forgetting engine its own package. Nine modules — calculator, history, strength, strategy, scheduler, service, config — means the maths can be read separately from the policy.

Avoid

  • Do not publish benchmark numbers only as images. Error bars in a PNG cannot be checked, reproduced or cited, and a reader has nowhere to look.
  • Do not interpolate an identifier into a DETACH DELETE query. The surrounding code binds parameters; add_nodes.py:16 does not, and it is the most destructive statement in the repository.
  • Do not expect usage-driven retention to correct anything. A wrong memory that is frequently retrieved is reinforced by exactly the mechanism that keeps a right one.

Fit

This suits a team wanting a hosted multi-user memory service with a console, a graph backend and a principled decay model, comfortable with Neo4j plus Postgres plus Redis and with a codebase whose comments are largely Chinese.

The transferable idea is small and separable: fusion-with-provenance as the forgetting operation, and a cycle-history table that counts its own failures. Both are worth lifting into a system with a different decay model entirely.

12. Open Questions

  • Where are the benchmark numbers from? No harness, dataset or result file is in the tree, and the three papers are hosted elsewhere.
  • What happens when fusion fails? failed_count is recorded; whether the pair is retried next cycle, skipped permanently, or left below threshold forever was not traced.
  • Can a MemorySummary itself be forgotten? The strategy fuses Statement–Entity pairs; whether summaries participate in later cycles decides whether the store converges or accumulates a summary tier.
  • Is the A-MBER benchmark's evaluation of MemoryBear independent? The benchmark and the system share authors, which the reader should know.

Appendix: File Index

Activationapi/app/core/memory/storage_services/forgetting_engine/actr_calculator.py (the formula and the Anderson citation :1-23), access_history_manager.py, memory_strength.py, config_utils.py

Fusionforgetting_engine/forgetting_strategy.py (the responsibilities :1-12, provenance retention :38 and :191, the DERIVED_FROM rewiring :355-395), forget_service.py, forgetting_engine.py

Cyclesforgetting_engine/forgetting_scheduler.py, api/app/services/memory_forget_service.py, api/app/models/forgetting_cycle_history_model.py:15-35, api/app/repositories/forgetting_cycle_history_repository.py

Scopeapi/app/repositories/neo4j/neo4j_connector.py:219, :232, create_indexes.py:348-364, add_nodes.py:16 (the interpolated delete)

Retrievalapi/app/core/memory/src/search.py

Service and consoleapi/, web/, sandbox/, e2b-infra/

ClaimsREADME.md §Benchmarks (images only), §Papers (three, all hosted outside the repository)

History

2026-09-14e5087b10… — second reading, 782 commits on. Screened again: 0 auto-run surfaces, 0 build-time exec paths, 0 dependency surfaces inside the cooldown and 6 unpinned manifests; nothing was installed and nothing was run. Both marks were re-tested at the producer and both hold, and each now carries the evidence record it had been asserted without. The scope_enforced citation was wrong in a way worth recording: neo4j_connector.py moved to api/app/repositories/neo4j/, and lines :219 and :232 are the body of delete_group — a scoped deletion rather than a read filter. The read-path predicate is in the per-node-type Cypher templates, where six of seven AND the key in unconditionally and the DIALOGUE template makes it conditional on the key being non-null; the null branch is unreachable today only because the live caller's MemoryContext is a Pydantic model requiring end_user_id. The scheduled forgetting cycle is gone: the Celery task, its beat entry and its queue route are all commented out, the last marked 已废弃, while the crontab built from FORGETTING_CYCLE_HOUR and FORGETTING_CYCLE_MINUTE is still constructed at import and referenced by nothing. The cycle runs on demand through memory_forget_controller.py instead.

2026-08-09857bb5b4… — first reading. Screened before reading; the tree was read, never installed, and no test or benchmark was run.