A decision ledger that validates before it appends

Agent Mesh

A hash-chained event log whose decisions carry tiers, supersession and executable verification commands, with a reviewer quorum that gates promotion — a decision_accepted event that misses quorum lands in the log and leaves the record proposed — while the changed-path check is advisory and read-only by committed contract, and the grounding packet an agent receives still never reads the decision store at all.

Carries 3 of 7 rubric mechanisms. Most systems here carry none or one (44%), and a dash means the mechanism was not found at this commit — not that the system needed it. Each mark is one LLM reviewer's reading of the code at this commit rather than a run of it — known limits.

  • Tombstone
  • Trust state
  • Bi-temporal
  • Scope enforced
  • Mutation audit
  • Human review
  • Negative evals

1. Executive Summary

Agent Mesh is a project-local coordination substrate for human-plus-agent teams: 59,028 lines of Python across src/agent_mesh/, MIT, release v0.4.2 (PyPI my-agent-mesh), on a repository whose first commit is chore: publish clean agent-mesh surface — a curated publish of work developed elsewhere, which is why the history says nothing about how the design arrived. It has no third-party dependencies: pyproject.toml declares an empty dependencies list under the comment "Stdlib-only by design for core." That is rare enough in this corpus to state first, and it means the supply-chain surface of adopting it is the Python standard library and nothing else.

The substrate is a hash-chained append-only events.jsonl with SQLite as a declared projection of it. Most of what flows through it — requests, responses, backlog items, dispatch runs and leases — is coordination state rather than memory: a request is an act, not a claim that can turn out false. The memory is the decision log, and it is a serious one. A decision carries a human ID with aliases, a tier, an externalized Markdown body addressed by SHA, an owner, a status, an enforcement mode, affected-code globs, exemptions, required checks, assumptions, evidence references, tags, and a list of verification commands with expected signals. Supersession detects cycles and refuses a target that was never accepted. Three tiers skip accepted and go straight to in_force.

The best thing here is the revision rule, and it is a mechanism rather than a convention. Editing a decision that is accepted or in_force through the Workbench requires a reason, emits a decision_revisited event, and sets fields_changed["status"] = [status, "proposed"] — the projection then clears accepted_utc and the record has to be accepted again. A revised decision cannot inherit the approval of the thing it replaced. Very little in this atlas closes that loop.

Two fields of the decision payload have no producer. cmd_decision_propose (cli/mail.py) and create_decision (workbench.py) both fill verification, required_checks, affected_code_globs and tags from their arguments, decision amend edits those four after the fact, and agent-q decisions verify executes what they hold: authoring parses each command into argv and rejects shell operators and env-assignments (core/decision_schema.py, reject_unsafe=True), execution is subprocess.Popen(argv, …, shell=False), and a verification on a decision that is not accepted/in_force is refused.

The same two payloads also carry exemptions, generated_artifact_paths, assumptions, evidence and review_policy from real arguments, normalized through core/decision_schema.py, and all five sit in DECISION_REVISION_AUTHORITY_FIELDS — the set whose revision forces re-approval. What remains hardcoded is exactly "rejected_alternatives": [] and "consequences": [], at workbench.py:2245-2246 and cli/mail.py:2275-2276. Both are projected (store/rebuild.py:3883-3884) and both are rendered by agent-q decisions show and the Workbench detail view, so a reader is shown two sections that nothing can fill. Neither appears in the revision-authority set, which is consistent: there is nothing to revise.

The reviewer quorum is a gate on promotion, and it binds to a revision. review_policy is authored at both write paths and normalized by normalize_decision_review_policy, which rejects an approval_quorum without reviewers and requires the quorum to lie between 1 and the number of reviewers. _decision_quorum_reached (store/rebuild.py:4380) counts distinct decision_accepted actors that appear in required_reviewers, and the projector that consumes it does the load-bearing thing: when the quorum is not met it returns before the UPDATE decisions SET status=…, so the acceptance is in the log and the record is still proposed. Partial approval is durable and visible without being effective, which is the property a two-of-three review needs.

The binding is the careful part. Approvals count only when their approved_revision_sha matches the revision being approved, so an approval of an earlier draft does not carry forward; decision_review_progress surfaces approved_reviewers, remaining_reviewers and an approval_binding that reads legacy_pre_authoring_digest when an approval predates that binding. Combined with the revision rule below — editing an accepted decision returns it to proposed — an approval in this store is an approval of a specific version by a named participant, and nothing inherits it. Section 9 has the caveat on who a participant may be.

The gap is between the ledger and the agent. dispatch/grounding.py assembles a section called prior-decisions, and it is not the decision store: prior_decisions_text regex-matches a verdict pattern against the summary and first two hundred characters of res_posted message bodies in the same thread. Beside it, message_packet.py builds the packet an agent receives over HTTP, and the string decision does not occur in that file at all — it is body, thread, and a grounding block declaring "default_scope": "thread" with the instruction "Use this packet/body/thread data for automated dispatch grounding." No decision record reaches an agent's context through any automatic path; the contract installed into AGENTS.md tells the model to go and look.

Enforcement is advisory by construction, and a test says so. decision_applicability.py computes an effective_enforcement from the stored enforcement_mode, and the first rule it applies is if configured == "required": effective = "advisory" — nothing is ever enforcing — with a further downgrade to none for an invalid context or a decision still at proposed. Beside it in the same payload, "evaluation_status": "not_run" and "would_block": None are assigned once and never reassigned anywhere in the package. check decisions therefore reports which decisions touch the changed paths and declines to say whether they would block, which test_changed_path_decision_check_is_complete_advisory_and_read_only pins: it asserts both constants and then asserts the event log and database are byte-identical before and after the check. Testing that a read path does not write is two lines and it is the kind of assertion most suites leave implicit.

The read path has been rebuilt around a new module: store/read_model.py, "mutation-free reads over one verified canonical event-log snapshot", entered at thirty-two call sites inside cli/ and warmed on a thread for the Workbench's decision view. Section 6 has the shape of what that leaves behind.

The published tree ships one test file, tests/public/test_public_contract.py, a fourteen-test behaviour contract, with CI.

2. Mental Model

A memory here is a decision: a durable, human-identified record (D001, D038-S1, D076-E) of something the project settled, with the argument in an externalized Markdown body addressed by body_sha and the structure in the event payload. Everything else the log carries is a different kind of thing. A request is a speech act; a backlog item is workflow state; a dispatch lease is a concurrency primitive. None of them can be wrong in the way a decision can, and the design is right not to treat them alike.

Nothing extracts a decision. There is no model anywhere in the package — an agent or a human names the decision, and the substrate's job is to make the naming durable, ordered and hash-linked. Capture is therefore a deliberate act with a form, which puts Agent Mesh at the far explicit end of zero-LLM capture.

The state machine is a status column and it is real. decision_proposed inserts at proposed. decision_accepted first checks the reviewer quorum (_decision_quorum_reached, which passes trivially when required_reviewers is empty and otherwise blocks promotion until enough named reviewers have accepted this revision), then promotes to in_force if the tier is architecture_contract, production_invariant or compliance_security, and to accepted otherwise. decision_rejected and decision_retired are terminal marks; decision_superseded points the old record at a successor and the successor back at the old one, after _ensure_supersede_target_valid refuses a target that is not accepted or in_force and _ensure_no_supersede_cycle walks the chain.

The transition worth studying is the reverse one. Editing an accepted or in-force decision in the Workbench is refused without a revision_reason, then emits decision_revisited and folds status: [old, "proposed"] into the metadata update. _project_decision_metadata_updated reads that and sets accepted_utc = NULL. Approval is attached to a version of the content, not to the record, and the code enforces the difference.

Two states in the vocabulary have no producer either, and they are the two that would police the fields nothing fills. decision_assumption_violated would flip a decision_assumptions row to violated and stamp the invalidating event; decision_check_failed would append to the decision's log. Both appear in the accepted-kinds set and in the projection's elif chain, and neither is emitted by any code path in the package. decision_drift_detected is emitted, by one command.

Nothing forgets. There is no delete, no redaction, no expiry and no retention policy over decisions or messages — the only DELETE statements in the tree clear derived SQLite rows before reprojection. body_fidelity admits the value redacted, and it is a label a caller may pass to agent-mesh request --body-fidelity; no code path produces it, and nothing removes the bytes it would describe.

Diagram — the decision lifecycle, where a quorum that is not reached logs the acceptance and leaves the record proposed, rejection is keyed on the record rather than the value, and the enforcement mode is downgraded to advisory before any read path sees it
Diagram source
%% caption: the decision lifecycle, where a quorum that is not reached logs the acceptance and leaves the record proposed, rejection is keyed on the record rather than the value, and the enforcement mode is downgraded to advisory before any read path sees it
stateDiagram-v2
    [*] --> Proposed: decision_proposed — verification, checks, globs, assumptions, evidence and review_policy from arguments, while rejected_alternatives and consequences are hardcoded empty
    Proposed --> Proposed: decision_accepted below quorum — appended to the log, no promotion
    Proposed --> Accepted: decision_accepted, ordinary tier — quorum reached for this revision_sha
    Proposed --> InForce: decision_accepted — architecture_contract, production_invariant or compliance_security
    Proposed --> Rejected: decision_rejected
    Accepted --> InForce: status set through decision_metadata_updated
    Accepted --> Proposed: Workbench edit — reason required, emits decision_revisited, clears accepted_utc
    InForce --> Proposed: same rule — approval belongs to the content, not the record
    Accepted --> Superseded: decision_superseded — target must be accepted or in_force, cycles refused
    InForce --> Superseded
    Accepted --> Retired: decision_retired
    InForce --> Retired

    note right of Rejected
        Keyed on the record, not the value.
        Nothing stops the same content
        being proposed again as D-next.
    end note

    note right of InForce
        enforcement_mode is stored per tier,
        then required is downgraded to advisory
        before any consumer sees it, and
        would_block stays None by construction.
    end note

3. Architecture

Nothing has to be running to store or read anything. A project is a .agent-mesh/ directory holding config.toml, events.jsonl, externalized bodies, a SQLite file, generated Markdown views and a lock.

The log is the store. core/events.py defines a frozen Event envelope — event_id, schema_version, occurred_utc, event_seq, actor, kind, entity_id, thread_id, prev_event_hash, payload — serialized through canonical_json with sorted keys and no whitespace, hashed with SHA-256 over the exact newline-terminated bytes. core/chain.py verifies the chain in two modes and its docstring states the limit of the cheap one precisely: an anchored walk "does NOT prove that an arbitrary earlier prefix line was not rewritten in place at the same byte length; that requires the full walk". The chain is unkeyed, so it is tamper-evident against accident and against an editor who does not recompute, and not against anyone who does.

SQLite is a projection and says so. store/sqlite.py creates roughly thirty tables; store/rebuild.py replays the log into them. reset_schema wipes, apply_record refuses a gap in event_seq, and table_hashes_for canonically dumps and hashes every table so two machines can compare projections. This is the cleanest instance in the corpus of the log-and-projection split the atlas records in its own notes: the derived store is disposable by construction, and the code treats it that way.

Concurrency is an owner-aware file lock (core/lock.py) carrying a PID, a boot ID and a hostname, with OwnerStatus distinguishing live, stale_pid_dead, stale_age_exceeded and cannot_verify, and a 600-second staleness bound. Interrupted writes are recovered idempotently (core/recovery.py), and dispatch/atomic.py journals a set of replace and delete operations so a crash rolls forward rather than leaving a half-applied unit.

The Workbench (workbench.py, 4,411 lines) is a loopback HTTP server plus a single-file HTML page, with a per-server access token carried in a URL fragment, restricted browser origins, a 0600 machine-local bookmark stored outside any repository, and retry-safe receipts so an uncertain feedback retry returns the original request instead of creating a second one. workbench_service.py installs it as a user-level service under launchd, systemd --user or Task Scheduler, serving every repository in a machine-local registry from one process.

Deployment and ergonomics

pip install agent-mesh, then agent-mesh init in a repository. No database server, no key, no network, no model. Python 3.11+ and the standard library.

The store is repairable by hand in the ordinary sense — events.jsonl is one canonical JSON object per line and the bodies are Markdown files — but with a sharp caveat the design creates: editing a line invalidates every prev_event_hash after it, so hand-repair means rewriting the chain, and there is no tool for that. The delete-free, redaction-free design means the recovery story for a mistaken write is append a corrective event, and for a decision that is decision_revisited or decision_superseded. For anything the schema cannot express, there is no story.

Privacy is the part of the operator experience that has had the most thought. New projects default to local-only, whose generated .agent-mesh/.gitignore is deny-all and ignores itself, so a plain git add -A selects nothing. git-shared is an explicit opt-in allowlisting exactly the config, the event log and externalized bodies — and docs/privacy.md refuses to oversell it: "This is an allowlist of canonical state, not a promise that the allowed files are public." It also states that configs predating the setting are read as git-shared rather than silently made private, and that changing .gitignore "does not erase prior commits, forks, caches, or clones."

4. Essential Implementation Paths

  • Append. core/events.py:append_event — assigns event_seq and prev_event_hash from the tail, validates provenance (core/provenance.py:validate_event_provenance), dispatch payloads (core/dispatch_schema.py) and decision stop lines (_validate_stateful_event_before_append → validate_decision_event, :150, :428–:468), then writes the canonical line. An invalid decision event fails before it is journalled.
  • Project. store/rebuild.py:rebuild_all → reset_schema → apply_record per event → _project_record. Decision handling is _project_decision_event around :1099 and _project_decision_proposed at :1179.
  • Decision write surfaces. cli/mail.py:cmd_decision_propose (:838), cmd_decision_accept, cmd_decision_revisit, supersede, retire; and workbench.py:create_decision (:500) and the edit path (:680).
  • Decision read surfaces. cli/q.py:cmd_decisions_list/show/log/search/at/verify (:733–:914) and the Workbench Decisions tab.
  • Verification runner. cli/q.py:cmd_decisions_verify — subprocess.Popen(argv, …, shell=False) (:1324), emitting decision_drift_detected on non-zero exit; refuses a verification on a non-accepted decision.
  • Grounding. dispatch/grounding.py — sections for referenced rows, git state, prior-decisions (prior_decisions_text, :57) and the full thread, each tagged stable or volatile and hashed.
  • Packet. cli/q.py:cmd_packet → message_packet.py:build_message_packet; the word decision does not occur in that module.
  • Chain integrity. core/chain.py, exposed as agent-q verify-chain.
  • Locking and recovery. core/lock.py, core/recovery.py, dispatch/atomic.py.
  • Agent contract. skill/render.py:CANONICAL_SKILL_BODY, installed by adoption.py into AGENTS.md and, when the repository shows signs of it, CLAUDE.md.
  • Tests. tests/public/test_public_contract.py, four cases, plus a CI workflow. Nothing else in the tree is a test.

5. Memory Data Model

decisions is the table that matters: dec_ulid primary key, human_id, parent_human_id, title, tier, status, enforcement_mode, owner, body_sha, body_path, body_bytes, body_media_type, superseded_by, supersedes, proposed_utc, accepted_utc, in_force_utc, retired_utc, last_verified_utc, drift_risk, event_seq, meta_json. Around it sit decision_aliases (a renamed decision keeps its old ID resolvable, with one primary), decision_globs (kinds affected, exempt, generated), decision_checks, decision_verifications, decision_assumptions, decision_evidence, decision_references_in_code and decision_tags.

Separating the body from the record by SHA is the right call and it pays off twice: body_sha is a field decision_metadata_updated can change, so a revision is a content-hash change with an event behind it, and the alias table means the identifier a human uses is decoupled from the identity the projection keys on.

Temporal fields are all record time. occurred_utc on the event, event_seq for order, and four transition stamps on the decision. There is nothing recording when a decision was true of the world as distinct from when the project wrote it down, and no as-of query — agent-q decisions at takes a file path, not a date. Single axis, so not bi-temporal, though the log makes the history reconstructible by replay in a way most single-axis stores cannot manage.

There is no scope key. The boundary is .agent-mesh/ in a directory, the same filesystem-boundary answer as Graphify and Klypix MCP. The near-miss is worth naming because it looks like more: the multi-repo Workbench server "resolves its opaque repo ID before feedback, request-status, backlog, attachment, or decision operations", which is a real access-control indirection over a machine-local registry. But it selects which store to open rather than filtering rows inside one, and no record carries a tenant, user or workspace key. config.participants is the closest thing to an identity list, and it gates who may act, not what may be read.

Provenance is the model's other strength, and it is on the message side rather than the decision side. core/provenance.py defines closed vocabularies for body_authority (human_chat, agent_mail, tool_payload, agent_summary, recovery_artifact, unknown), body_fidelity (full, metadata_only, reconstructed, inferred, redacted, missing), source-selection modes and eight causal-edge relations. Distinguishing what a human said from what an agent summarised, as a validated enum on the record, is a thing this atlas asks about constantly and finds rarely — and the grounding path uses it, refusing to call a packet complete when a thread event's fidelity is not full.

6. Retrieval Mechanics

Retrieval is deliberate and thin. agent-q decisions search selects every row from decisions, builds a haystack from human_id, title, and the context and decision strings inside meta_json, and prints rows whose lowercased haystack contains the lowercased query. No index, no ranking, no scoring, no limit, and the body — the file where the actual argument lives — is not searched at all. agent-q decisions at <path> is the more interesting query: it joins decision_globs where kind='affected' and fnmatches the relative path, which answers which decisions govern this file. That is the right question, and both write paths can supply the globs it needs — --affects on the CLI, the affected_code_globs argument in create_decision. The exempt and generated glob kinds the same table carries have no producer, so the join is populated in one of its three kinds.

Neither query filters on status. A retired, rejected or superseded decision prints alongside a live one with its status in the second column, which is the recall-first-and-label choice and is defensible on a surface a human reads.

The automatic path does not retrieve decisions at all. dispatch/grounding.py assembles a packet with a stable prefix and volatile suffix — a cache-shaped design, and it hashes every section so a caller can see what changed — but its prior-decisions section comes from prior_decisions_text, which scans res_posted events in the current thread for the regex \b(APPROVE_WITH_CHANGES|APPROVE|REQUEST_CHANGES|REJECT|NO-GO|GO)\b in the summary and the first 200 characters of the body. Those are review verdicts inside messages, not records in the decision store. build_message_packet, which backs agent-q packet, contains no reference to decisions either.

So the substrate's durable, tiered, supersession-aware memory is reachable by an agent only if the agent runs a query, and the thing that tells it to is prose: the contract installed in AGENTS.md says "Run agent-q status and targeted agent-q list/locate/body before responding." This is the shape the atlas records as the guidance was already in context — an instruction to consult standing in for a mechanism that consults.

The other retrieval cost is structural, and it is the part of the system that has moved furthest. store/read_model.py — "mutation-free reads over one verified canonical event-log snapshot" — is the path most agent-q reads take: open_read_model is entered twenty-nine times in cli/q.py alone, against two remaining _rebuild_all_locked call sites, where the wipe-and-replay was once the default. The full replay is not gone — rebuild_all is mentioned twenty-five times in q.py and thirty-one in cli/mail.py — and projection_is_current, which compares the log's SHA-256 against the value stamped in the projection's meta table and would let a reader skip a replay entirely, has exactly one caller, the Workbench refresh at workbench.py:4594. Three read strategies coexist: a verified snapshot, an unconditional full replay, and a staleness check used once.

7. Write Mechanics

Writes are synchronous, deterministic, and cheap in model terms because no model is involved. The path is: acquire the project lock, read the tail line for event_seq and prev_event_hash, validate, append the canonical line, project. A new decision is retrievable immediately — the next reader replays the log and sees it.

Idempotence and crash safety got real attention. AGENT_MESH_FAULT_AFTER is a fault-injection hook in the append protocol, core/recovery.py handles interrupted writes, dispatch/atomic.py journals its file operations for roll-forward, and the Workbench's feedback receipts make an uncertain retry return the original request ID.

Decision invariants are checked at both boundaries. A superseded target must be accepted or in force, supersession may not cycle, a human ID may not collide, an alias may not fork, a parent may not be missing — and apply_record raises DecisionStopLine at projection time when one is broken, which alone would make a bad event permanent poison in a log with no delete: appended durably, then aborting every subsequent replay. append_event therefore runs the same check first, through _validate_stateful_event_before_append → validate_decision_event (events.py:150,428-468), so an invalid decision event fails before it is journalled. The public-contract test asserts the bad event never lands and events.jsonl stays byte-identical, and agent-q decisions diagnose reports replay health. Validating at the write boundary as well as the replay boundary is the standard defence, and it is a defence a store this shape needs, because the replay-only version of it is a durable denial of service against yourself.

The other write-side finding is narrower than the payload shape suggests. Both propose paths hardcode exactly two values:

"rejected_alternatives": [],
"consequences": [],

Everything else the payload declares is authored. exemptions, generated_artifact_paths, assumptions, evidence and review_policy all arrive from arguments through core/decision_schema.py, which normalizes and validates each — a review_policy with an approval_quorum and no reviewers is rejected, and a quorum outside 1..len(reviewers) is rejected. All five sit in DECISION_REVISION_AUTHORITY_FIELDS, so revising any of them returns an accepted decision to proposed.

The two that remain are projected into the decisions meta (store/rebuild.py:3883-3884), listed in the metadata-update meta_fields set, and rendered as "Rejected Alternatives" and "Consequences" sections by both agent-q decisions show and the Workbench detail view. So the reader is shown two headings that no shipped command can fill, and neither field appears in the revision-authority set — consistent, because there is nothing to revise. The honest framing is that the substrate is a library — the README says "Project-specific importers should live in the consumer repository" — so a consumer can import append_event and construct a fuller payload. What is missing is two rendered sections, not the argument of the decision itself.

Input handling on the fields that are writable is careful, and the shape of the residual risk is worth keeping in view. agent-q decisions verify parses each stored command into argv at authoring time, rejecting shell operators and env-assignments (core/decision_schema.py, reject_unsafe=True), executes via shell=False (q.py:1324), and refuses both a legacy shell-dependent string and any verification on a non-accepted decision. The command is still memory — it arrives in an event payload written by a participant, and the participant set includes agents — so a store whose records feed a later executor stays a path worth watching. What the argv boundary buys is that the executor takes a vector, not a string, so the field can hold a command and cannot hold a shell.

8. Agent Integration

There is no MCP server and no SDK. The integration is two CLIs and a contract.

agent-mesh writes: init, request, reply, respond, resolve, reopen, decision propose|accept|revisit|supersede|retire, backlog upsert|link, skill render|install, adopt, workbench. agent-q reads: list, locate, body, packet, thread, trace, render, rebuild, recover, verify-chain, status, events, backlog, decisions list|show|log|search|at|verify, dispatches.

adoption.py installs a versioned managed block into AGENTS.md, and into CLAUDE.md when the repository already has one or a .claude/ directory. agent-mesh adopt --check detects a stale contract and conflicting legacy decision-write guidance, and the contract text tells the agent that finding one "is an adoption defect; report it instead of silently choosing a second source of truth." Treating a second writable surface for the same facts as a defect the tool detects, rather than a documentation problem, is the right instinct.

The contract itself is the best-written agent-facing prose in this part of the corpus, and one passage is worth copying outright. Under Quality Discipline it names event kinds that do not exist yet and instructs against inventing them:

Until quality/investigation events exist, treat this as advisory procedure. Future event names (do not invent today): quality_bar_declared, quality_bar_updated, quality_gate_evaluated, investigation_opened, …

A model asked to record something for which no verb exists will invent one. Naming the reserved vocabulary in advance is a cheap defence against a schema being polluted by plausible guesses, and nothing else in this atlas does it.

The contract's own wording invites a weaker reading than the code supports: "Run agent-mesh decision accept only after explicit human approval and name the approving human with --by." That sounds like an honour system, and cmd_decision_accept (cli/mail.py:2549-2615) is not one. It calls _ensure_participant with the role string approving human, then refuses any --by outside config.decision_approval_identities with DECISION_APPROVER_UNAUTHORIZED, requires a non-empty approval note, and then does the thing that actually separates a person from a process:

if not sys.stdin.isatty():
    raise ConfigError(
        "decision acceptance requires direct human action in an interactive terminal; "
        "use the Workbench Approve and accept control instead"
    )

A dispatched agent runs without a terminal, so the command is closed to it before any identity question arises. What follows is stricter still: the approver must be in required_reviewers if that list is non-empty, the whole decision is printed — body and revision hashes, tier, owner, scope, globs, required checks, verification commands, assumptions, evidence, reviewers and quorum — and the person has to type ACCEPT <id> exactly. After the typed confirmation the snapshot is re-read and the append is refused if the dec_ulid or revision_sha moved while it was on screen. The CLI path is the more demanding of the two surfaces, not the looser one.

9. Reliability, Safety, and Trust

Integrity. The hash chain is the strongest property and its own docstrings scope it correctly: full walks catch a same-length rewrite of an earlier line, anchored walks do not, and the anchored mode is offered as an incremental gate rather than as the audit. table_hashes_for extends the idea to the projection so two replicas can compare derived state. What none of it provides is authenticity: SHA-256 with no key means the chain proves the log has not been carelessly edited, and anyone who can write the file can produce a consistent forgery. For a project-local file that is a reasonable place to stop, but "tamper-evident" in the README is doing work that a signature would do properly.

Trust states are genuine and applied. status gates supersession (_ensure_supersede_target_valid), drives the tier promotion, is reset by the revision path, and gates the promotion itself: a decision_accepted event that does not reach the configured quorum is appended and the record stays proposed. What is absent is any gate on reading: an agent that queries decisions search gets proposed records beside in-force ones.

enforcement_mode is the field that does not do what its name says. It is stored per tier, and core/decision_applicability.py reads it into an effective_enforcement — which is the first real consumer — but the first rule applied is if configured == "required": effective = "advisory", with a further downgrade to none for an invalid context or a proposed decision. Nothing is ever enforcing. In the same payload, "evaluation_status": "not_run" and "would_block": None are assigned at one line each and reassigned nowhere in the package, so the two fields that would say whether a decision blocks a change are constants — and that is a pinned contract rather than a gap a reader discovers. test_changed_path_decision_check_is_complete_advisory_and_read_only asserts both constants, and asserts that running the check leaves events.jsonl and the SQLite file byte-identical. A read path tested not to mutate the store is rare, and pinning the non-enforcement as a contract is the honest way to ship a field you have not wired.

Audit is the capability this system has most completely. The event log is not a sidecar record of mutations; it is the store, append-only, ordered, hash-linked and schema-versioned, with the queryable form defined as derived from it. _append_decision_log additionally keeps a per-decision event trail, and agent-q decisions log prints it. Nothing in this corpus makes the mutation record more load-bearing.

Human review is a place and a rule. The Workbench Decisions tab creates proposals, appends revisions with a required reason, and records acceptance, with the re-approval rule enforced in code. Beside it, review_policy names required reviewers and a quorum, both write paths author it, and the projector refuses to promote a decision until enough of those reviewers have accepted this revision — approvals are matched on approved_revision_sha, so an approval of a superseded draft does not carry. decision_review_progress reports approved_reviewers, remaining_reviewers and an approval_binding that flags legacy_pre_authoring_digest where an approval predates the binding.

One caveat survives, and it is a configuration default rather than a missing check. Both surfaces resolve the approver against decision_approval_identities — _human_decision_actor in the Workbench, an explicit DECISION_APPROVER_UNAUTHORIZED in the CLI — and that list is decision_approval.human_approvers when a project has written a [decision_approval] table, and tuple(self.participants) when it has not. Participants are the agents as well as the person; the documentation's own example is ["human", "builder", "reviewer", "observer"]. So until the table is added, "a reviewer" need not be a person as far as the identity check is concerned — though the CLI's terminal requirement still stands in the way of a dispatched one, and the Workbench control is a browser surface rather than an agent-callable API.

The project does not hide this. The unconfigured state is named unmigrated_participant_compatibility, the Workbench reports the diagnosis, the configuration guide says the table "does not grant agents or dispatched reviewers human approval authority" and tells the operator to add it deliberately, and every acceptance event binds the approval_authority_mode and revision that were in force when it was made — a stale binding raising REVIEW_ASSURANCE_HUMAN_AUTHORITY_STALE — so a later reader can tell which regime an approval was made under instead of assuming today's.

The reference scanner checks the inverse of what the schema suggests. agent-mesh decision refs walks the tree for D001-shaped tokens and reports the ones that resolve to nothing — a dangling reference, a citation of a decision that does not exist — recording the list in a decision_scanner_run_completed payload with --record-scan. It never writes decision_references_in_code; that table is filled by decision_reference_resolved events, which no code path emits, so the join between a decision and the lines of code that cite it stays empty. The check that ships is the cheap and useful half, and no other system here checks that a memory's identifier is citable at all — but it runs the opposite direction to verify memory against its subject, which asks whether the cited code moved rather than whether the citation resolves. Both would be worth having; one is here.

Failure modes worth naming. Full-replay-per-read is the standing one: correctness is excellent — the projection cannot drift, because it is rebuilt — and the cost is linear in total history on every query. Behind it sits the failure the write-time validator exists to prevent, which is worth understanding even though the validator stands in front of it: a log that can hold a state the projection will refuse is a store that can be bricked by one append, and append_event is the only thing standing between a caller and that. A consumer importing append_event gets the check; a consumer writing a line to events.jsonl by hand does not, and there is no repair tool, because repairing means rewriting every prev_event_hash after the bad line.

Privacy and deletion. Local-only by default, deny-all .gitignore including itself, an explicit allowlist for the shared mode, and a publish checklist that tells the reader to inspect commit authors, screenshots and package artifacts. Against that: there is no delete and no redaction. If a request body captured a secret or a personal detail, the substrate's answer is the same one docs/privacy.md gives for Git — rotate it, because the record is not coming out. For a design that stores human chat verbatim under a human_chat authority label, an intentional redaction path is the obvious next mechanism, and the redacted fidelity value is already sitting in the enum waiting for it.

10. Tests, Evals, and Benchmarks

The published tree ships a curated public verification pack: tests/public/test_public_contract.py, four tests, with a CI workflow. The only test-related line in .gitignore is .pytest_cache/, which is an artifact path and not evidence of a suite kept elsewhere. The four are behaviour contracts and they are pointed: request→respond→packet round-trips and the packet carries the bodies; a hash-chain tamper is detected (prev_event_hash mismatch); an invalid decision transition raises DecisionStopLine at append and leaves events.jsonl byte-identical (the poison-event fix, asserted); and the git-shared allowlist tracks exactly config/events/gitignore and not a private attachment. They assert must-detect and must-not-write, not must-not-retrieve, so they do not earn negative_eval, but they turn "what a reader can check is nothing" into a real, if small, checkable surface.

That matters more here than in most reports, because the design's claims are exactly the kind that only a test can support, and four cases reach one of them. Idempotent crash recovery, a fault-injection hook (AGENT_MESH_FAULT_AFTER, core/events.py:29) that exists specifically to be driven by a test, roll-forward of a journaled atomic apply, owner-aware lock staleness across a reboot, a projection that must exactly reproduce the log, and the re-approval rule that is the best thing in the design — every one of these is asserted by a docstring and by no executable. The fault-injection environment variable is the clearest signal: somebody wrote a seam for a test harness, the seam is exported from core/events.py, and nothing in the tree drives it.

What ships instead is two runnable examples, examples/solo-project/run.sh and a parameterized N=3 examples/n-agent/run.sh, which exercise the flow and check nothing. agent-q verify-chain, agent-q audit-recovered-sources and agent-mesh adopt --check are operator-facing verification commands and a reasonable substitute for runtime assurance, but they verify a live store, not the code's behaviour.

There is no paper, no benchmark, and no performance claim — grepping the README, docs/ and the source for arxiv, bibtex, @article, @misc, Citation and doi returns nothing, and there is no CITATION.cff. The README makes no quantitative claim at all, which given how thin the suite is makes for correct restraint: nothing here is asserted that a reader is invited to trust on numbers.

Before relying on this I would want, in order: a test asserting that a revised decision loses its accepted_utc, because that is the mechanism the design is best at and nothing protects it from a refactor; a test that the projection of a fixed log hashes to a fixed value; and a test driving AGENT_MESH_FAULT_AFTER through the recovery path.

11. For Your Own Build

Steal

  • Make approval belong to the content, not the record. Requiring a reason to edit an approved record, emitting a distinct revision event, and clearing the approval stamp so it must be granted again is a handful of lines and closes the hole where a memory keeps its blessing through a rewrite.
  • Name your reserved vocabulary to the model before you implement it. A contract that lists the event names a future version will use, under do not invent today, costs one paragraph and prevents a schema being polluted by plausible invented verbs.
  • Treat a second writable surface for the same facts as a detectable defect. adopt --check looks for conflicting legacy decision-write guidance and reports it. Most projects document the migration and hope; a check that fails is better.
  • Put provenance in a closed enum and validate it on append. Separating who authored this from how faithful is this copy — human_chat versus agent_summary, full versus reconstructed — and rejecting values outside the set makes the difference queryable instead of conventional.
  • Hash the projection, not just the log. A canonical per-table dump hashed after replay turns "did these two machines derive the same state" into a comparison rather than an argument.
  • State the limit of the cheap integrity check in the code. The anchored-walk docstring explains exactly what it does not prove and keeps the full walk as the audit. Every incremental verifier should carry that paragraph.

Avoid

  • Do not enforce invariants only at replay. If a record can be durably accepted at write time and rejected at read time, you have built a store that can hold a state your reader will never accept — and in an append-only log with no delete, that state is permanent. Validate at both boundaries, or make the replay skip and quarantine rather than abort.
  • Do not ship a consumer with no producer. Two payload fields here — rejected_alternatives and consequences — are projected, listed in the metadata-update field set, and rendered as headings by two views, with [] hardcoded at both write paths. A reader sees two empty sections and cannot tell whether the author had no alternatives or no way to record them. If a field is not writable yet, leave the reader out or make it fail loudly.
  • Do not execute memory through a shell. A stored field a maintenance command runs is a memory store with a code-execution path, and the writer of that field is whoever can append an event. The defence here is worth copying in both halves: parse to argv at authoring time so an unsafe command cannot be stored, and execute with shell=False so a stored string cannot become a shell line. Rejecting at authoring is the half most designs skip, and it is the half that keeps a bad value out of an append-only log.
  • Do not let the derived index be rebuilt from scratch on every read. The correctness argument for full replay is good and the cost is unbounded in history. The skip check here is written and correct; it is simply not called from the reader that runs most often.
  • Do not let a field's name outrun its wiring — and if you must, pin it. enforcement_mode is computed, stored and non-null, and the one consumer that reads it downgrades required to advisory unconditionally while shipping would_block: None and evaluation_status: "not_run" as constants. The redeeming move is the test: a public contract asserting both constants, and asserting the check leaves the log and the database byte-identical, turns an unfinished feature into a documented posture a caller can rely on.

Fit

This suits a small team that wants coordination between humans and several coding agents to leave a record it can audit, in a repository it controls, with no service and no vendor. The install cost is genuinely near zero — standard library only, one directory, and a privacy default that errs toward not committing anything — and the log-and-projection architecture means the parts most likely to be wrong are the disposable ones. If what you want is an inspectable history of who asked for what and what the project decided, this is a more careful foundation than most.

It is not a memory layer for an agent's working knowledge, and reading it as one will disappoint. Nothing is retrieved automatically, nothing is ranked, nothing is summarised, and the only thing standing between a decision and the model that should honour it is an instruction to go and query. Walk away entirely if you need multi-user or multi-tenant boundaries, or if you need to delete or redact anything you have stored. An approval that more than one person has to give is available: review_policy names reviewers and a quorum, and the projector withholds promotion until they have accepted the current revision — with the caveat that the participant list which bounds "a reviewer" routinely contains agents. At v0.4.2, with fourteen contract tests against a 59,028-line surface, the right posture is to read the design for its ideas — several of which are better than what surrounds them — while checking any specific guarantee against the code rather than the docstring.

12. Open Questions

  • Does the upstream development repository hold the suite AGENT_MESH_FAULT_AFTER implies? The seam is exported for a harness that is not in the published tree, whose history is eight commits beginning with a curated publish, so this cannot be settled from what is here.
  • What is the intended producer for rejected_alternatives and consequences? Both are rendered as headings and neither can be written by a shipped command, which is the remaining instance of a pattern the rest of the payload has grown out of.
  • The route back for a log that stricter validation now rejects exists and is deliberately walled off: store/decision_recovery.py is "reviewed migrations for narrowly identified legacy decision events … deliberately not part of normal append or replay … for operator-authorized recovery when stricter replay validation identifies an already-canonical legacy event and no valid log backup is available." Whether an operator can actually complete that recovery needs the tool run against a constructed log, which this reading did not do.
  • _decision_quorum_reached counts distinct accepting actors and the participant list routinely includes agents. With the quorum configurable, would two agents accepting satisfy it?

Appendix: File Index

  • Log and integrity — core/events.py, core/hashing.py, core/chain.py, core/ids.py, core/lock.py, core/recovery.py, core/source_recovery*.py, core/external_recovery_plan.py.
  • Schema and projection — store/sqlite.py, store/rebuild.py.
  • Decision model — _project_decision_proposed and _project_decision_metadata_updated in store/rebuild.py; enforcement_for_tier, _decision_quorum_reached (:4380, gating the promotion at :3728), _ensure_supersede_target_valid, _ensure_no_supersede_cycle; core/decision_schema.py (normalize_decision_review_policy :286-327, decision_review_progress :355-400); workbench.py:203-222 (DECISION_REVISION_AUTHORITY_FIELDS), :2245-2246 and cli/mail.py:2275-2276 (the two hardcoded fields).
  • Read path and recovery — store/read_model.py ("mutation-free reads over one verified canonical event-log snapshot", entered at thirty-two call sites in cli/), store/decision_recovery.py (operator-authorized legacy migration, outside normal append and replay), store/rebuild.py:1063 (projection_is_current, one caller).
  • Enforcement — core/decision_applicability.py:284-312 (the required → advisory downgrade, evaluation_status and would_block as constants), tests/public/test_public_contract.py::test_changed_path_decision_check_is_complete_advisory_and_read_only.

Recorded searches. Run from a checkout at a8187089….

  • Only two payload fields lack a writer. for f in assumptions evidence review_policy rejected_alternatives consequences exemptions generated_artifact_paths; do grep -rn "\"$f\"" --include="*.py" src; done — five resolve to core/decision_schema.py normalizers and authored Workbench fields; rejected_alternatives and consequences resolve only to [] literals, projection reads and renderers.
  • Nothing decides whether a decision blocks. grep -rn "would_block\|evaluation_status" --include="*.py" src — one assignment each, both in core/decision_applicability.py, neither reassigned.
  • The grounding packet holds no decision. grep -n "decision" src/agent_mesh/message_packet.py — nothing.
  • The skip check has one caller. grep -rn "projection_is_current" --include="*.py" src — the definition, one import, and workbench.py:4594.
  • Provenance — core/provenance.py.
  • Write surface — cli/mail.py, workbench.py.
  • Read surface — cli/q.py, message_packet.py, views/rendering.py, views/inbox.py, views/outbox.py, views/archive.py, views/log.py.
  • Grounding and dispatch — dispatch/grounding.py, dispatch/dispatch.py, dispatch/atomic.py, dispatch/guard.py, dispatch/runtime.py, core/dispatch_schema.py.
  • Agent integration — skill/render.py, adoption.py, project_registry.py, workbench_service.py.
  • Docs — README.md, docs/adoption.md, docs/configuration.md, docs/migration.md, docs/privacy.md.

History

2026-09-19 — a8187089… — trust_state re-tested at the unchanged pin. Every writer the record cited is exact, and the record was all writers: it proved the status is set carefully and never said where it is read. The reader is evaluate_decision_applicability in core/decision_applicability.py, the engine that decides which decisions apply to a set of paths — and the predicate is in Python rather than SQL, which is why a grep for a status clause misses it. The shape is the strong one: a status set defaulting to {"accepted", "in_force"} and a continue on anything outside it, so it is an allowlist and an unanticipated value is withheld rather than admitted. Widening is explicit and named — include_proposed adds exactly one value, lifecycle_statuses replaces the set — and the agent-facing context builder (core/decision_context.py:54) passes the narrow default through; the only caller that widens is the CLI, behind --include-proposed. adoption.py:325 applies the same pair as SQL for its migration report. One detail is added on the write side too: folding a decision back to proposed clears in_force_utc as well as accepted_utc, in the same statement, so neither timestamp can outlive the standing it records. Nothing was installed and no suite was run.

2026-09-18 — a8187089… — re-read at the same commit; nothing upstream moved, so the corrections are this report's. The human_review account was wrong in the direction that matters. It said cmd_decision_accept "takes --by as a free string and appends the event" and called CLI acceptance an honour system; at this same pin the command calls _ensure_participant with the role approving human, refuses any approver outside decision_approval_identities with DECISION_APPROVER_UNAUTHORIZED, refuses an empty note, and then refuses to run at all unless sys.stdin.isatty() — "decision acceptance requires direct human action in an interactive terminal" — which closes the path to a dispatched agent before any identity question arises. It then prints the whole decision, requires a typed ACCEPT <id>, and re-reads the snapshot to refuse the append if the revision moved while it was on screen. The CLI is the stricter of the two surfaces. The report also named the Workbench guard _decision_actor and credited it with one check; the function is _human_decision_actor and it refuses twice, the second time against the configured direct-human approval authority. The caveat that survives is narrower and is a default rather than an omission: decision_approval.human_approvers is None until a project writes a [decision_approval] table, and the property then falls back to the participant list, which includes the agents — a state the project names unmigrated_participant_compatibility, reports as a diagnosis, and binds into every acceptance event as the authority mode in force.

2026-09-11 — a8187089… — re-read at v0.4.2, eighty-one files and 40,943 insertions past the previous pin. Screened before reading: no auto-run surface, pyproject.toml two days old inside the cooldown and declaring dependencies with no lockfile beside it. The tree was read, never installed, and nothing was run. Marks unchanged at trust_state, audit_log and human_review, and the evidence records for trust_state and human_review were rewritten, because this report's central criticism is closed. review_policy is authored at both write paths through a normalizer that rejects a quorum without reviewers and a quorum outside 1..len(reviewers); it sits in DECISION_REVISION_AUTHORITY_FIELDS; approvals count only when their approved_revision_sha matches the revision being approved; and _decision_quorum_reached gates the projector, which returns before the promotion UPDATE when the quorum is not met — so a decision_accepted event below quorum is logged and the record stays proposed. Of the seven fields the previous reading found hardcoded empty, five are now authored and two remain: rejected_alternatives and consequences, still [] at workbench.py:2245-2246 and cli/mail.py:2275-2276, still rendered as headings by two views. enforcement_mode gained its first consumer in core/decision_applicability.py, which downgrades required to advisory unconditionally and ships evaluation_status and would_block as constants — pinned by a contract test that also asserts the check leaves the event log and database byte-identical. Two new store modules: read_model.py for mutation-free reads over a verified snapshot, entered at thirty-two call sites in cli/, and decision_recovery.py, an operator-authorized legacy migration deliberately outside normal append and replay. The grounding claim was re-run and holds: dispatch/grounding.py still derives prior-decisions from a verdict regex over thread messages, and message_packet.py mentions no decision at all. Counts corrected to 59,028 lines of Python and fourteen contract tests.

2026-08-17 — 43bfe5cc… — read again at the same commit: upstream main has not moved, there are no other branches, and v0.3.0 is still the only tag. Screened again before reading — pyproject.toml inside the seven-day cooldown, no auto-run surface, nothing installed or run; the dependency list it declares is empty, so the cooldown has nothing to hold back. Two published claims were wrong and are corrected in the body. The reviewer quorum has no producer: review_policy is hardcoded {} by cmd_decision_propose (cli/mail.py:1568) and create_decision (workbench.py:655), decision amend has no flag for it, no Workbench field sets it, and _decision_quorum_reached returns True on an empty required_reviewers — so acceptance is single-actor in every store the shipped verbs can build, and the previous entry's "optional reviewer quorum" overstated a gate that cannot be switched on. The same check applied to the whole payload puts the count of fields with a projection and no write surface at seven, not four. Second: .gitignore was read as evidence of a suite held back in an internal repository; its only test-related line is .pytest_cache/, an artifact path, at this commit and at the first reading's. Three sections still carried the pre-v0.3.0 state beside the corrected summary — Tests. None. in the path list, "decision stop lines are not checked here" on append_event, a code block listing verification and required_checks among the hardcoded-empty fields, and "the Workbench cannot create the globs it needs" — all four are rewritten to the current tree. Marks unchanged and now carrying evidence records. Verified afresh at this pin: enforcement_mode is stored, projected and rendered by views/rendering.py:169 and read by nothing; assumptions and evidence reach decision_assumptions/decision_evidence only from decision_proposed, which hardcodes both; events.py:150,428-468 and q.py:1324 are still the validator and the shell=False executor. No paper, no CITATION.cff.

2026-08-15 — 43bfe5cc… — re-pinned at release v0.3.0 (24,653 lines, eight commits, PyPI distribution my-agent-mesh). Screened again; a manifest inside the cooldown, nothing installed or run. The release fixed all three of this report's central negative findings: the verification apparatus, which both write paths hardcoded empty, is now writable through --verification, the Workbench and a new decision amend verb (three of the five fields — verification, required_checks, affected_code_globs — now populate; assumptions and evidence still have no producer); the poison-event hazard is closed by _validate_stateful_event_before_append → validate_decision_event in append_event (events.py:150,428-468), so an invalid decision event fails before it is journalled; and the shell=True verification runner is now argv/shell=False with authoring-time rejection of shell operators (core/decision_schema.py, q.py:1324). A public-contract test suite ships (tests/public/test_public_contract.py, four tests, plus CI) where there were none. The eyebrow and description are rewritten accordingly. Marks are unchanged — trust_state, audit_log and human_review all hold and are lightly strengthened (a completeness gate on acceptance, receipt↔︎decision binding, direct human approval of revisions); no new mark is earned (the contract tests are must-detect/must-not-write, not must-not-retrieve). What still holds: the grounding packet never auto-reads the decision store (it regexes posted result bodies), enforcement_mode is printed and gates nothing, and there is no scope key, no validity-time axis and no delete. No paper.

2026-08-10 — 258a1eed… — first reading, at 5 commits. Screened before reading: 0 auto-run surfaces, 0 dependency surfaces inside the seven-day cooldown, 1 unpinned manifest with no lockfile — which is a declared-but-empty dependencies list, so there is no third-party surface to pin. Nothing was installed and nothing was executed; the two shipped examples were read rather than run.