Evidence-tiered pattern learning

OmniIntelligence

A pattern-learning service whose evidence ladder and hysteresis band are argued in code, and whose cold-start path selects exactly what its gate refuses.

LicenceMIT
Size176,327 lines of Python in 1,088 files under src, 66 ONEX node packages and 27 Postgres migrations
Activity925 commits on dev by six author names, two of them bots, 13 November 2025 – 28 September 2026
Tests7,228 test functions in 174,097 lines, under tests and the per-node node_tests

Carries 3 of 7 rubric mechanisms. Most systems here carry none or one (44%), and a dash means the mechanism was not found at this commit — not that the system needed it. Each mark is one LLM reviewer's reading of the code at this commit rather than a run of it — known limits.

  • Tombstone
  • Trust state
  • Bi-temporal
  • Scope enforced
  • Mutation audit
  • Human review
  • Negative evals

1. Executive Summary

OmniIntelligence is the memory half of a two-repository loop: it learns patterns from coding-session events, grades them on an evidence ladder, and serves the survivors over HTTP to OmniClaude, which injects them into sessions and reports back what happened. Its lifecycle thresholds are argued in the code, with demotion deliberately harder than promotion. Its weakness is wiring: the cold-start promotion path selects exactly the rows its own reducer refuses.

The reasoning is in the code rather than in a design document. Demotion is deliberately harder than promotion — ten injections rather than five, a failure streak of five rather than three, a 40% success floor against a 60% ceiling — and the docstring beside each constant says why, in the terms an experimentalist would use: "The 20% gap between promotion (60%) and demotion (40%) ensures ... random variance doesn't cause flip-flopping between states." Every transition is written to an append-only table carrying a gate_snapshot of the conditions that justified it. The evidence tier can only increase, and the guarantee is not a convention but the WHERE clause of the UPDATE.

Three things are broken in the way this atlas exists to find, and all three are invisible from the outside.

The cold-start path selects exactly what the gate refuses. SQL_FETCH_CANDIDATE_PATTERNS deliberately admits evidence_tier = 'unmeasured' rows for bootstrap promotion; apply_transition — which the same handler then calls — rejects any transition to PROVISIONAL from a tier below OBSERVED. The one test named "full promotion lifecycle" passes a mock in place of apply_transition that returns success, and asserts a promotion the real reducer would refuse.

The top evidence tier is unreachable. verified is a valid column value, is gated on, and is written by nothing: compute_evidence_tier computes only OBSERVED or MEASURED and otherwise keeps the current tier, and its docstring says "VERIFIED requires independent validation (not computed here)."

The manual kill switch reads a stale view. Disabling a pattern is a hard override that bypasses the cooldown — and it is read from disabled_patterns_current, a materialized view whose only REFRESH statements in the tree are inside integration tests.

The Goodhart and reward-hacking guardrails are real, pure, and tested, and nothing calls them.

The framework beneath the mechanism is public and locked by version. The ONEX runtime, the repository runtime that executes the SQL contracts, and EnumEvidenceTier itself come from omnibase-core, omnibase-infra, omnibase-spi and omnimarket. pyproject.toml declares them as version ranges, and uv.lock resolves each from pypi.org — 0.47.24, 0.38.59, 0.23.5 and 0.4.261 at this pin (uv.lock:2883, 2907, 2969, 2983). All four are MIT repositories under github.com/OmniNode-ai, tagged per release. The evidence guard's verdict depends on the enum's ordering, which section 4 reads at the locked tag.

2. Mental Model

A memory here is a learned pattern: a signature clustered from session events, with a domain, keywords, a confidence float, and two independent axes of standing.

The first axis is lifecycle status — candidate → provisional → validated → deprecated — and it decides whether the pattern may be injected into a session. The second is evidence tier — unmeasured → observed → measured → verified — and it decides whether the pattern is allowed to advance. Status is what the pattern is; tier is what is known about it.

Tier is computed from one question: was there a measured pipeline run, and did it succeed?

if run_id is None:
    computed = EnumEvidenceTier.OBSERVED          # anecdotal
elif run_result == "success":
    computed = EnumEvidenceTier.MEASURED          # quantitative
else:
    computed = EnumEvidenceTier.OBSERVED          # a run happened and didn't succeed

and it can only go up. The monotonicity is not a Python convention — it is the predicate of the statement that writes it:

UPDATE learned_patterns SET evidence_tier = $2, updated_at = NOW()
WHERE id = $1
  AND (CASE evidence_tier WHEN 'unmeasured' THEN 0 WHEN 'observed' THEN 10
       WHEN 'measured' THEN 20 WHEN 'verified' THEN 30 ELSE 0 END)
    < (CASE $2 WHEN 'unmeasured' THEN 0 WHEN 'observed' THEN 10
       WHEN 'measured' THEN 20 WHEN 'verified' THEN 30 ELSE 0 END)

A concurrent writer, a replayed Kafka message and a buggy caller all fail the same way: the statement matches no rows. The migration names one writer — "The attribution binder is the SOLE writer of this column" — and the guard holds even if that stops being true for updates. It does not cover inserts: an operator seeding script creates rows born measured (section 7).

At runtime, status moves only through apply_transition, which the demotion handler calls the single source of truth, and which refuses → PROVISIONAL below OBSERVED and → VALIDATED below MEASURED. So the two axes are wired together in one direction: evidence gates status, and status never touches evidence.

The enum those gates compare is one rung stricter in its own docstring. omnibase-core calls MEASURED sufficient for PROVISIONAL and VERIFIED "Required for VALIDATED lifecycle state in production" (enum_evidence_tier.py:100, :107 at the locked tag). apply_transition asks OBSERVED and MEASURED. Read against that docstring, the tier nothing writes is the one it names for production validation.

What kills a pattern is asymmetric on purpose. Promotion is optimistic, demotion conservative, and between them sits a 20-point band of success rate that belongs to neither.

Diagram — promotion needs injections, a success rate and an evidence tier together, and the band between forty and sixty per cent is hysteresis where nothing moves
Diagram source
%% caption: promotion needs injections, a success rate and an evidence tier together, and the band between forty and sixty per cent is hysteresis where nothing moves
stateDiagram-v2
  [*] --> candidate: clustered from session events
  candidate --> provisional: 5 injections · 60% success · tier >= observed
  candidate --> candidate: bootstrap path selects unmeasured, reducer refuses it
  provisional --> validated: 5 injections · 60% success · tier >= measured
  validated --> deprecated: 10 injections · under 40% success · after 24h cooldown
  validated --> deprecated: 5 consecutive failures
  validated --> deprecated: manual disable, cooldown bypassed
  validated --> validated: between 40% and 60% — the hysteresis band
  deprecated --> [*]: excluded from every injection query

3. Architecture

An event-driven Python service. ONEX nodes — 66 package directories, each with a contract.yaml, handlers, models and its own node_tests — are dispatched from Kafka topics through a runtime plugin. Postgres is the only store this repository runs.

  • deployment/database/migrations/ — 27 forward migrations, numbered 000 to 028 with 026 and 027 absent, and 20 rollback scripts. 005 is learned_patterns, 006 disable events, 007 injections, 010 the lifecycle audit, 011 the evidence tier, 012 measured attributions, 022 project scope.
  • src/omniintelligence/nodes/node_pattern_* — learning, extraction, storage, promotion, demotion, feedback, lifecycle, projection, compliance, matching, assembly.
  • src/omniintelligence/repositories/learned_patterns.repository.yaml — the read SQL, declared as data rather than embedded in Python.
  • src/omniintelligence/api/router_patterns.py — the FastAPI surface that serves patterns to clients.
  • src/omniintelligence/runtime/dispatch_handler_*.py — the wiring from Kafka envelope to handler.

Deployment and ergonomics

This is the heavy end of the atlas. An operator needs Postgres with 27 migrations applied, Kafka with the topics the node contracts declare, the ONEX runtime, and — for the loop to close — a separate repository running in the agent's session. It cannot run offline as a library; there is no embedded mode and no file-backed fallback.

The store is inspectable and repairable in the way a SQL store is: every state is a row, the audit is a table, and an operator who knows SQL can answer any question this report asks. That is a genuine advantage over the file-and-vector stores that dominate this corpus, and it is bought with an operational floor most single-user memory systems would not accept.

The framework is the larger adoption cost, and it is an open one. The four framework packages are version ranges in pyproject.toml, locked to PyPI releases, and each release has a matching tag in a public MIT repository. The block declaring them carries the comment "ONEX private ecosystem dependencies (NOT on public PyPI)" (pyproject.toml:164), which the lockfile beside it contradicts. Until 20 August 2026 omnibase-core was a git-rev override on an unreleased commit, and ac6817ad… removed it. Only the dev-group onex-change-control stays a git source, at a public commit.

4. Essential Implementation Paths

Learning — node_pattern_learning_compute

Session events are clustered, and each cluster is scored by handler_confidence_scoring.py, which is explicit that a single number is not the answer:

"Confidence scores are DECOMPOSED, not monolithic. A single confidence value cannot answer 'why did this pattern score high?' — but component scores can."

Three components with fixed weights — label_agreement 0.40, cluster_cohesion 0.30, frequency_factor 0.30 — asserted at import time to sum to 1.0, and a warning in capitals: "NEVER make downstream decisions solely on the confidence field. The rolled-up score is for convenience only. Inspect components."

learned_patterns has one column, confidence FLOAT. The three components appear in no migration. The advice is correct and the schema makes it impossible to follow — and the bootstrap promotion gate, described below, decides on lp.confidence >= 0.8.

Promotion — handler_auto_promote.py, and the path that cannot work

Two phases per run: CANDIDATE → PROVISIONAL, then PROVISIONAL → VALIDATED. The candidate query admits two kinds of row:

AND (
  lp.evidence_tier IN ('observed', 'measured', 'verified')
  OR (
    lp.evidence_tier = 'unmeasured'
    AND COALESCE(lp.injection_count_rolling_20, 0) = 0
    AND lp.confidence >= 0.8
    AND lp.recurrence_count >= 1
    AND lp.distinct_days_seen >= 1
  )
)

The second branch is the cold-start path, and its thresholds carry the clearest record in this repository of a gate being loosened under supply pressure:

"Lowered from 2 to 1: existing candidate patterns were initialized with recurrence_count=1 and distinct_days_seen=1. The confidence gate (>= 0.8) is the primary quality signal for the bootstrap path; requiring recurrence >= 2 blocked all 5,384 high-confidence candidates that had never been re-observed."

Each selected row is then passed to apply_transition_fn with to_status=PROVISIONAL. In production that function is the real apply_transition — dispatch_handler_promotion_check.py imports it directly — and it contains:

if (to_status == EnumPatternLifecycleStatus.PROVISIONAL
        and current_evidence_tier < EnumEvidenceTier.OBSERVED):
    return ModelTransitionResult(success=False, …,
        reason="Evidence tier guard: insufficient evidence for PROVISIONAL", …)

There is no bootstrap exemption in the guard. Every row admitted by the second branch is unmeasured, and every unmeasured row is refused. The handler records promoted=False, logs a warning, and continues; nothing raises.

The refusal rests on the framework's comparison, and it holds. EnumEvidenceTier is a str enum in omnibase-core, and a plain one would compare lexicographically, where "unmeasured" < "observed" is false and every bootstrap row would pass. It overrides all four ordering operators to compare weights of 0, 10, 20 and 30, coercing a raw string operand first (enum_evidence_tier.py:27, :126). The framework's tests assert the non-lexicographic order against raw strings (tests/unit/enums/pattern_learning/test_enum_evidence_tier.py:216).

Both halves are tested, separately. test_candidate_to_provisional_with_unmeasured_rejected calls the real apply_transition with an unmeasured candidate and asserts failure (tests/unit/nodes/node_pattern_lifecycle_effect/test_handler_transition_evidence_guards.py:320). Every test that calls the auto-promote handler passes mock_apply_transition, eleven call sites in all, so no test joins the selection to the gate.

So the loosening cannot have unblocked the 5,384 candidates, because the reducer refuses them for a reason the loosened thresholds do not touch. The two halves are individually correct and disagree about what the cold-start rule is.

Demotion — handler_demotion.py

The forgetting half, and the best-argued code in either repository. Three gates, any one of which fires: manual disable (a hard trigger that bypasses the cooldown), a failure streak of five, or a success rate under 40% with at least ten injections in the rolling window. Two eligibility checks precede gates two and three: the pattern must be validated, and 24 hours must have passed since promotion.

The override bounds are the part worth copying. An operator may tune the thresholds per request, and the permitted range is itself argued: SUCCESS_RATE_THRESHOLD_MAX = 0.60 exists to "prevent setting demotion threshold at or above promotion threshold (60%) which would cause immediate demotion of marginal patterns", and FAILURE_STREAK_THRESHOLD_MIN = 3 prevents demoting "on small runs of bad luck". The hysteresis band cannot be configured away.

The handler also refuses to run without Kafka rather than degrading, and says so: "Without Kafka, lifecycle events cannot be emitted, and demotions will not occur. This is intentional — the reducer is the single source of truth." A caller detects it by checking reason == "kafka_producer_unavailable". Failing loudly on the forgetting path is the right direction for the failure to point.

The kill switch that does not reach the reader

Manual disable is an append-only event — pattern_disable_events, with event_type IN ('disabled', 're_enabled'), a reason TEXT NOT NULL and an actor VARCHAR(100) NOT NULL. Requiring a reason and an actor on a memory override is rare here and right.

The demotion and promotion queries do not read that table. They join disabled_patterns_current, a materialized view created in migration 008, whose own comment carries the operating instruction: "Refresh with: REFRESH MATERIALIZED VIEW CONCURRENTLY disabled_patterns_current;". Searching the tree for that statement returns three lines of migration 008 — two SQL comments and the view's COMMENT ON text — and four calls inside tests/integration/nodes/node_pattern_promotion_effect/test_promotion_integration.py. No scheduler, no job, no dispatch handler refreshes it.

So the strongest correction the system offers — a person naming a pattern and a reason and turning it off — takes effect when somebody remembers a maintenance command that is documented in a SQL comment. The integration tests pass because they run it themselves.

That section is also why this report does not carry human_review. Three things stand against it and any one would be enough. The event is a disable, applied to a pattern already promoted and already injected, so nothing waits in it. The actor is a string the writer chooses, admitting "user, system, or automated process" by the schema's own comment at 006_create_pattern_disable_events.sql:39. And the disable does not reach a reader at all until the materialized view is refreshed by hand. Requiring a reason and an actor on an override remains rare and right, and it is recorded here for that rather than for a gate.

Outcome feedback — node_enforcement_feedback_effect

The rule here is the one this atlas asks for and finds stated in code nowhere else. A violation only counts as negative evidence when the agent was both warned and observed to correct:

return [v for v in violations if v.was_warned and v.was_corrected]

"If an agent was warned of a violation but did NOT correct it, we cannot confirm the violation was real. The warning might have been a false positive."

That is a real epistemic distinction — between the memory fired and the memory was right — and most feedback loops in this corpus collapse it, treating every advisory as a labelled example and learning from their own false positives.

Audit — pattern_lifecycle_transitions

Every status change writes a row: request_id for idempotency, from_status, to_status, transition_trigger, correlation_id, actor, reason, and a gate_snapshot JSONB holding the gate conditions at transition time. The foreign key is ON DELETE RESTRICT with the reason in a comment: "Audit records must never be silently deleted when parent patterns are removed."

An audit that stores the evidence a decision was made on, rather than only the decision, is the version of this pattern worth having.

The guardrails nothing calls

node_anti_gaming_guardrails_compute implements four defences as pure functions with 452 lines of tests: Goodhart detection over correlated metric pairs, reward-hacking detection when a score improves without matching human acceptance, distributional-shift detection by symmetric KL approximation, and a diversity constraint that is a veto. Its contract declares an operation and no topics. No module under src/ outside its own package calls run_all_guardrails; the one other mention is an allowlist entry in a contract-validation test (tests/unit/test_contract_validation.py:77).

The project records this itself. Its node inventory, linked from the README and kept in the project's public knowledge base, lists seven node directories as "(unregistered)", including this node's companion alerter and node_objective_ab_framework_compute (reference/omniintelligence-node-inventory.md). A published list of what is not wired is a better artifact than most projects offer, and it is the reason this finding is a citation rather than an accusation.

5. Memory Data Model

learned_patterns carries identity (pattern_signature, signature_hash, is_current, version), a domain foreign key with domain_candidates JSONB, keywords TEXT[], confidence FLOAT CHECK (confidence >= 0.5), the lifecycle columns, provenance (source_session_ids UUID[], recurrence_count, first_seen_at, last_seen_at, distinct_days_seen), and rolling metrics over a window of 20.

The constraints are unusually load-bearing. Every rolling counter is bounded to [0, 20], and:

CONSTRAINT check_rolling_metrics_sum
    CHECK (success_count_rolling_20 + failure_count_rolling_20 <= injection_count_rolling_20)

An arithmetic invariant of the metrics that drive promotion, enforced by the database rather than by whoever writes next.

What is absent:

  • No bi-temporal axis. first_seen_at and last_seen_at are observation times of the pattern, not validity times of a claim, and nothing separates when a pattern was true from when the row learned it.
  • No rejected-value tombstone. deprecated is keyed on the pattern row. The clustering path that produced the signature will produce it again from new sessions, as a new candidate, and nothing consults the deprecation.
  • No confidence components, as above.
  • project_scope is a column nothing can pass. Migration 022 adds it with two indexes, and learned_patterns.repository.yaml applies it: AND ($6::text IS NULL OR project_scope IS NULL OR project_scope = $6::text). The FastAPI endpoint that serves patterns declares domain, language, min_confidence, limit and offset — and no project parameter — so $6 is never bound on the network read path, and every project is served every project's patterns.

6. Retrieval Mechanics

Retrieval is SQL. list_validated_patterns filters on status IN ('validated', 'provisional'), is_current = TRUE, a confidence floor, an optional domain, an optional keyword against keywords, and the unreachable project predicate. Ordering puts validated before provisional, then confidence descending, then id — a stable order, which matters when the consumer takes a prefix.

The SQL in the contract is the SQL Postgres runs. omnibase-infra's PostgresRepositoryRuntime leaves a contract's WHERE clause untouched; for a multi-row read it appends ORDER BY the primary key when none is present, and LIMIT its max_row_limit when no limit is present (postgres_repository_runtime.py:584, :731). The default is ten rows, and this repository constructs the runtime without a config. Every multi-row read in learned_patterns.repository.yaml binds its own LIMIT $n, so the default never truncates a pattern read.

The framework also documents a boundary this read path leaves to the client. omnibase-core's EnumPatternLifecycleState gives PROVISIONAL "Limited eligibility (testing/staging only)" (enum_pattern_lifecycle_state.py:46 at the locked tag). This repository defines its own EnumPatternLifecycleStatus, and list_validated_patterns serves provisional beside validated with no environment predicate, as the bootstrap fallback its comment names.

There is no vector search, no embedding, no reranker and no query rewriting in this repository's read path for patterns. For a store of a few thousand short signatures selected by domain and confidence, that is a defensible choice rather than a missing feature, and it is the reason the retrieval stack reads as lexical alone.

The failure mode is over-supply at the client. The endpoint caps limit at 200, and OmniClaude requests ten times its injection budget precisely because its own filters "each of which can eliminate the majority of candidates" run after the fetch. The filtering that decides what an agent sees happens on the far side of an HTTP boundary from the store that knows the evidence.

7. Write Mechanics

Patterns are born from clustering over session events, dispatched through Kafka. Outcomes arrive separately: injections are recorded by the client, run attributions by the binder, violations by the enforcement path. pattern_measured_attributions is the evidence table, with a constraint tying tier to the presence of a run:

CHECK ((run_id IS NULL AND evidence_tier = 'observed')
    OR (run_id IS NOT NULL AND evidence_tier IN ('measured', 'verified')))

Attribution inserts are idempotent against Kafka redelivery by INSERT … WHERE NOT EXISTS — chosen over ON CONFLICT because the unique indexes are partial, with the reasoning written beside it.

A second writer enters beside the ladder rather than through it. scripts/seed_patterns_from_families.py --execute inserts one row per ONEX node family found on disk, built by node_family_to_pattern_row, born status = 'validated' and evidence_tier = 'measured' with no pipeline run and no row in pattern_lifecycle_transitions (handler_store_node_families.py:67, :75). On conflict it refreshes confidence and the compiled snippet and leaves status and tier alone, so a deprecated family stays deprecated. It is an operator script, not a runtime node, and the migration's "SOLE writer" comment describes the runtime.

Operational cost

  • Nothing on the agent's turn blocks on this repository's writes. Learning, attribution and lifecycle work are all consumers of events; the agent's session has moved on.
  • The lag from a session to an injectable pattern is unbounded by anything in the tree — it is clustering cadence, plus a promotion check, plus a demotion sweep, plus the tier the attribution binder can reach. Nothing measures it.
  • The rolling window is 20 and the counters are capped at 20, so a pattern's standing is always computed over its last twenty injections rather than its history. That is a decay policy expressed as a bound, and a cheap one.
  • No background pass rewrites the store. Promotion and demotion touch one row each, and the audit only grows.

8. Agent Integration

GET /api/v1/patterns and Kafka topics. The model never talks to this repository; the client does, and the model sees the result as text in its context. There are no MCP tools and no agent-callable write path — an agent cannot save a memory here, which is a deliberate shape: what becomes a pattern is decided by clustering over what agents did, not by an agent deciding something is worth remembering.

Adapting this to another host means reimplementing the client, which is most of OmniClaude's injection module, against an endpoint that is stable and small.

9. Reliability, Safety, and Trust

Strengths:

  • Monotonicity enforced in the statement, not by the writer's discipline.
  • Asymmetric promotion and demotion with a hysteresis band, argued in the constants themselves and bounded against operator override.
  • An audit that records the gate conditions, not only the outcome, and a foreign key that refuses to lose it.
  • Confirmed-only negative feedback — warned and corrected — so the loop does not learn from its own unconfirmed advisories.
  • Arithmetic invariants in the schema, including the one that keeps the promotion metrics coherent.
  • Idempotency against redelivery on both the attribution insert and the transition, with the choice of mechanism explained.
  • Failing closed on the forgetting path when Kafka is absent, rather than silently not demoting.
  • A published list of unregistered nodes, which is how the guardrail finding below is checkable at all.

Gaps:

  • The bootstrap promotion path is refused by the reducer it calls.
  • verified is unreachable, so a four-tier ladder is a three-tier ladder with a gate nothing can satisfy.
  • The manual kill switch is behind an unrefreshed materialized view.
  • The anti-gaming guardrails are not wired to anything.
  • project_scope cannot be passed by the API that serves the injector.
  • Confidence components are computed, warned about, and not persisted.
  • A deprecated pattern's signature can be relearned as a new candidate from new sessions.
  • An operator seeding script inserts rows born validated and measured, outside both the evidence ladder and the transition audit.

10. Tests, Evals, and Benchmarks

I ran nothing. The screen found no auto-executing surfaces, 29 conftest.py files that execute at pytest collection, and uv.lock inside the seven-day cooldown. The findings above are all static, and so are the framework readings, taken from the tags the lockfile names.

Per-node suites sit under node_tests/ alongside shared tests/unit, tests/integration and tests/integration/e2e.

The retrieval eval, and what it refuses to claim

src/omniintelligence/code_projection/retrieval_eval/ is a committed harness over a committed corpus — a golden query set, metrics, a replay driver, a scorecard and a baseline — and the interesting part is what it says about its own limits rather than what it scores.

The golden set is ten rows: six positives and four negatives. Its docstring gives the reason for the size rather than the number alone — the corpus exposes two retrievable chunk keys, so "six positives — two documents times three genuinely distinct intents — exhausts the honest positive space; a fourth phrasing per document would paraphrase one of those three rather than add a new correct answer." Volume comes from replaying each row at every state where it has a deterministic expected result, not from padding the set.

metrics.py then states which of its own metrics cannot discriminate on that corpus. recall@10 is pinned at 1.00 for every N ≤ 10, "so it cannot express a failure below N = 11", and the module records the chance floors that would apply at N = 20 — 0.50 for recall@10, 0.25 for recall@5, 0.05 for recall@1 and 0.18 for MRR — as an ARMING_MINIMUM_DISTINCT_DOCUMENTS threshold the eval must clear before its ranking numbers mean anything.

A harness that ships the conditions under which its own headline metric is uninformative is rare in this corpus, and it is the opposite failure mode to the one the benchmarks page spends its length on. The committed baseline_scorecard.json carries the corpus manifest id, the distinct chunk keys, an embedding compatibility key and an assertion count, so a later run can be compared against it or refused as incomparable.

The negative retrieval assertion is tests/unit/repositories/test_contract_lifecycle_filter.py, and it is an unusual shape. Rather than asserting about a result set, it loads the repository contract and greps the SQL of every injection query for a status it must not permit:

"DEPRECATED patterns have been demoted and should no longer be used. They MUST NOT be injected."

with a companion asserting the same for candidate. Because it asserts over the query text, it covers every caller of those operations at once, which an example-based test cannot, and the framework's runtime executes that text with its WHERE clause unchanged (section 6). It is also a regex over SQL, so a differently-phrased predicate — a negated filter, a status passed as a parameter — would satisfy it without meaning the same thing. Both halves are worth knowing.

tests/unit/enums/test_enum_injection.py is the other test worth naming: it asserts in both directions that the Python enum and the database CHECK constraint contain the same values, so the two cannot drift.

What is missing is the measurement the design implies. The evidence tiers, the hysteresis band and the rolling windows are an argument that this loop improves outcomes, and no committed result evaluates it; the A/B framework node that would is documented as unregistered, and the guardrails that would watch for the metrics being gamed are not called. The most valuable missing test is narrower and would have caught the headline finding: run the bootstrap path through the real apply_transition and assert what happens. A guard test asserts the refusal; the join between the two is what is absent.

No paper, arXiv reference or citation file exists in this repository.

11. For Your Own Build

Steal

  • Put monotonicity in the WHERE clause. A tier, a version or a state that must only advance should be guarded by the statement that writes it, so a concurrent writer and a redelivered message fail identically and silently-correctly.
  • Make demotion harder than promotion, and write down the band. More observations, a longer streak, a lower floor, and a cooldown after promotion. The 20-point gap between 60% and 40% is a design parameter that deserves a name and a comment.
  • Bound the operator's override. If thresholds are tunable, refuse a demotion threshold at or above the promotion threshold. A knob that can erase a hysteresis band will.
  • Snapshot the gates into the audit row. "Why was this promoted" is answerable from the audit only if the audit holds the conditions, not just the verdict.
  • Count a violation only when it was surfaced and then corrected. An advisory the agent ignored is not evidence the advisory was right, and a loop that treats it as one trains on its own false positives.
  • Require a reason and an actor on a manual override, as a NOT NULL column rather than a convention.
  • Assert enum-versus-constraint parity in both directions. It is ten lines and it catches the drift that produces a runtime constraint violation months later.
  • Enforce metric arithmetic in the schema — successes plus failures cannot exceed injections — so the numbers that drive promotion cannot go incoherent.
  • Publish which of your components are unregistered.

Avoid

  • A selection query and an enforcement gate with different rules. If one layer chooses rows and another refuses them, the system reports zero and nobody looks; put the gate where the selection happens, or make the selection call the gate.
  • Mocking the component under test in the test named for the whole flow. An integration test that replaces the gate with a stub returning success asserts only that the caller is wired, and reads as though it asserts the behaviour.
  • A state at the top of a ladder that nothing writes. Either something reaches verified or the ladder has three rungs; a permanently empty top tier makes every gate that mentions it unsatisfiable and every reader assume it is reachable.
  • A kill switch behind a manual refresh. Any correction a human can make must reach the read path without a second, undocumented human action.
  • Decomposing a score and persisting only the roll-up. The warning not to decide on the composite is unfollowable if the components are gone by the time anyone can decide.
  • A scope column the serving API cannot accept. Storing a boundary that no caller can pass is the same as not having one, and it reads like having one.

Fit

This suits a team, not a person: a platform group running Kafka and Postgres who want memory to be an auditable process with a lifecycle, and who will staff the operations that lifecycle implies. The design assumes a maintainer who thinks in gates, evidence and windows, and it rewards that — the argued constants and the gate snapshots are worth more than most of the retrieval sophistication elsewhere in this corpus.

Walk away if you want memory as a library. There is no embedded mode, the framework beneath it is four packages it cannot run without, and the smallest useful deployment is a message bus plus a database plus a second service inside the agent.

The uncomfortable judgement is about the gap between the design and its wiring. Four of the mechanisms this report praises — the bootstrap path, the top evidence tier, the kill switch, the guardrails — are specified more completely than they are connected, and each was found by following one call rather than by reading a diagram. A reader borrowing from here should borrow the reasoning, which is excellent and portable, and verify the wiring in their own tree rather than assuming it.

12. Open Questions

  • Has the bootstrap path ever promoted a pattern in production, and what did the 5,384 candidates do after the thresholds were lowered?
  • What is intended to write verified — an independent validator, a human, a second pipeline — and does the design want the tier gated on something that does not exist yet?
  • Is disabled_patterns_current refreshed by a deployment job outside the public repositories? omnibase_infra at the locked tag carries a copy of migration 008 and no refresh of it.
  • What consumes the anti-gaming guardrails in the intended design, and what topic would carry their alerts?
  • How long is the median path from a session event to an injectable validated pattern?
  • Which tier is VALIDATED meant to require: MEASURED, as apply_transition enforces, or VERIFIED, as the omnibase-core enum's docstring states?

Appendix: File Index

  • Schema and constraints: deployment/database/migrations/005_create_learned_patterns.sql, 006_create_pattern_disable_events.sql, 007_create_pattern_injections.sql, 008_create_disabled_patterns_current_view.sql, 010_create_pattern_lifecycle_audit.sql, 011_add_evidence_tier_to_learned_patterns.sql, 012_create_pattern_measured_attributions.sql, 022_add_project_scope_to_learned_patterns.sql.
  • Evidence tier and attribution: src/omniintelligence/nodes/node_pattern_feedback_effect/handlers/handler_attribution_binder.py.
  • Lifecycle gates: src/omniintelligence/nodes/node_pattern_lifecycle_effect/handlers/handler_transition.py.
  • Promotion: src/omniintelligence/nodes/node_pattern_promotion_effect/handlers/handler_auto_promote.py; wiring in src/omniintelligence/runtime/dispatch_handler_promotion_check.py.
  • Demotion: src/omniintelligence/nodes/node_pattern_demotion_effect/handlers/handler_demotion.py.
  • Confidence decomposition: src/omniintelligence/nodes/node_pattern_learning_compute/handlers/handler_confidence_scoring.py.
  • Confirmed-only feedback: src/omniintelligence/nodes/node_enforcement_feedback_effect/handlers/handler_enforcement_feedback.py.
  • Guardrails: src/omniintelligence/nodes/node_anti_gaming_guardrails_compute/handlers/handler_guardrails.py.
  • Read SQL and API: src/omniintelligence/repositories/learned_patterns.repository.yaml, src/omniintelligence/api/router_patterns.py.
  • Seeding writer: scripts/seed_patterns_from_families.py, src/omniintelligence/nodes/node_pattern_extraction_compute/handlers/handler_store_node_families.py.
  • Framework versions: pyproject.toml (ranges, and the removed git-rev overrides in comments), uv.lock.
  • Tests cited: tests/unit/repositories/test_contract_lifecycle_filter.py, tests/unit/enums/test_enum_injection.py, tests/unit/nodes/node_pattern_lifecycle_effect/test_handler_transition_evidence_guards.py, tests/unit/nodes/node_pattern_promotion_effect/test_handler_auto_promote.py, tests/integration/test_promotion_lifecycle_integration.py, tests/integration/nodes/node_pattern_promotion_effect/test_promotion_integration.py.
  • Framework, at the locked tags: omnibase_core v0.47.24 — src/omnibase_core/enums/pattern_learning/enum_evidence_tier.py, enum_pattern_lifecycle_state.py, tests/unit/enums/pattern_learning/test_enum_evidence_tier.py; omnibase_infra v0.38.59 — src/omnibase_infra/runtime/db/postgres_repository_runtime.py, src/omnibase_infra/runtime/db/models/model_repository_runtime_config.py, docker/migrations/intelligence/008_create_disabled_patterns_current_view.sql.
  • Self-documented wiring: the node inventory in the project's knowledge base, reference/omniintelligence-node-inventory.md, linked from README.md.

Recorded searches

Run from the repository root at the pinned commit unless a framework checkout is named.

git grep -n -A2 -E '^name = "(omnibase-core|omnibase-infra|omnibase-spi|omnimarket)"' -- uv.lock   # four registry sources, no git
git grep -n -E '^(omnibase|omnimarket)[a-z_-]* *= *\{' -- pyproject.toml   # nothing: no framework git source
git grep -n -E 'EnumEvidenceTier\.VERIFIED|evidence_tier *= *.verified' -- src   # nothing writes verified
git grep -n -E 'REFRESH MATERIALIZED VIEW' -- .   # migration 008 comments and one integration test
git grep -n -E 'run_all_guardrails' -- . ':!src/omniintelligence/nodes/node_anti_gaming_guardrails_compute'   # one test allowlist entry
git grep -n -E 'apply_transition_fn=' -- tests 'src/*/node_tests/*'   # eleven sites, all mock_apply_transition
git grep -n -E 'label_agreement|cluster_cohesion|frequency_factor' -- deployment   # nothing: components not persisted
git grep -n -E 'project_scope' -- src/omniintelligence/api/router_patterns.py   # nothing: no project parameter
git grep -n -E 'pattern_lifecycle_transitions' -- scripts/seed_patterns_from_families.py   # nothing: seeding writes no audit row
git grep -n -E 'def __(lt|le|gt|ge)__' -- src/omnibase_core/enums/pattern_learning/enum_evidence_tier.py   # in omnibase_core v0.47.24: four overrides
git grep -n -E 'REFRESH MATERIALIZED VIEW.*disabled_patterns_current' -- .   # in omnibase_infra v0.38.59: comment in the 008 copy only

History

2026-09-28 — re-pinned to 25ee01e7…, 37 commits of CI and dependency bumps; no mechanism file changed. The framework was described as three private repositories, and omnibase-core as a git-rev pin. Both were wrong at the previous pin: the override was removed on 20 August 2026, and the repositories were public, with release sources on PyPI. Read at the locked tags, EnumEvidenceTier compares by weight, so the headline refusal holds (section 4). Added: a seeding script that inserts validated/measured rows outside the ladder, and a guard test asserting the refusal. NODE_INVENTORY.md left the tree on 2 September 2026 and is cited in the knowledge base. Marks unchanged. Screened: no auto-run surface, 29 build-time execution points, uv.lock inside the cooldown. Nothing was installed, built or run.

2026-09-19 — re-pinned to 274f097f…. human_review is withdrawn on three grounds, each already written in this report. The record itself supplied the second: the actor column "does not distinguish a human actor from an automated one — the column admits user, system, or automated process by its own comment", so the mark rested on the surface existing rather than on who used it. The first is that a pattern_disable_events row is a disable applied to a pattern already promoted and already injected, so no memory waits in it. The third is section 4's own finding: the disable does not reach any reader until REFRESH MATERIALIZED VIEW CONCURRENTLY disabled_patterns_current is run by hand, and nothing in the tree runs it outside the integration tests. Requiring a reason and a named actor on an override is still rare and right, and it keeps that credit. trust_state, audit_log and negative_eval stand. Screened again first; nothing was installed, built or run.

2026-09-16 — 5039c8a3… — re-read at a commit dated 2026-09-16, 11 commits past the previous pin. All four marks re-tested and held; each anchored file has the same blob at both commits, so no line number moved. Screened before reading: no auto-run surface, 29 build-time execution points, no unpinned dependency surface and one dependency file inside the seven-day cooldown; an agent-addressed instruction file was recorded as data. Nothing was installed, built or run.

2026-09-09 — ce104631… — second reading, 59 commits along the default dev branch: 226 files, 23,928 insertions. Screened before reading; nothing was installed and no suite was run.

The four marks are untouched. Every path this report's appendix names still exists, and the diff across all of them is 71 files and 172 insertions, most of it one-line contract.yaml edits — so the lifecycle audit table, the disable-events table, the transition gate, the guardrails and the contract-filter test are all as described.

The 24,000 insertions landed almost entirely in a subsystem the report did not cover: code_projection, with context_serving and retrieval_eval beside it. Section 10 now covers the eval, because it is the part with something to say to this atlas. It is a committed harness over a committed corpus whose golden set is ten rows, and both the set size and the metric selection are argued rather than asserted — the docstring explains why a seventh positive would be a paraphrase rather than a new answer, and metrics.py records that recall@10 is pinned at 1.00 below eleven distinct documents and therefore "cannot express a failure", along with the chance floors that would apply once the corpus is large enough to arm it.

Six new migrations arrived (020 through 028, with debug-intelligence tables, LLM routing decisions, review calibration runs and an agent-actions retention policy); none of them touches the pattern lifecycle the marks rest on.

2026-08-11 — 8c67665a… — first reading, on the dev default branch. Screened before reading: 0 auto-run surfaces, 29 build-time exec surfaces (all conftest.py), 0 unpinned manifests, and a uv.lock unchanged for 16 days; nothing was installed and nothing was executed. The three omnibase-*/omnimarket git dependencies were not readable at this reading, so no claim here rests on them.