1. Executive Summary
OmniIntelligence is the memory half of a two-repository loop: it learns patterns from coding-session events, grades them on an evidence ladder, and serves the survivors over HTTP to OmniClaude, which injects them into sessions and reports back what happened. Its lifecycle thresholds are argued in the code, with demotion deliberately harder than promotion. Its weakness is wiring: the cold-start promotion path selects exactly the rows its own reducer refuses.
The reasoning is in the code rather than in a design document.
Demotion is deliberately harder than promotion — ten injections rather
than five, a failure streak of five rather than three, a 40% success
floor against a 60% ceiling — and the docstring beside each constant
says why, in the terms an experimentalist would use: "The 20% gap
between promotion (60%) and demotion (40%) ensures ... random variance
doesn't cause flip-flopping between states." Every transition is
written to an append-only table carrying a gate_snapshot of
the conditions that justified it. The evidence tier can only increase,
and the guarantee is not a convention but the WHERE clause
of the UPDATE.
Three things are broken in the way this atlas exists to find, and all three are invisible from the outside.
The cold-start path selects exactly what the gate
refuses. SQL_FETCH_CANDIDATE_PATTERNS deliberately
admits evidence_tier = 'unmeasured' rows for bootstrap
promotion; apply_transition — which the same handler then
calls — rejects any transition to PROVISIONAL from a tier
below OBSERVED. The one test named "full promotion
lifecycle" passes a mock in place of apply_transition that
returns success, and asserts a promotion the real reducer would
refuse.
The top evidence tier is unreachable.
verified is a valid column value, is gated on, and is
written by nothing: compute_evidence_tier computes only
OBSERVED or MEASURED and otherwise keeps the
current tier, and its docstring says "VERIFIED requires independent
validation (not computed here)."
The manual kill switch reads a stale view. Disabling
a pattern is a hard override that bypasses the cooldown — and it is read
from disabled_patterns_current, a materialized view whose
only REFRESH statements in the tree are inside integration
tests.
The Goodhart and reward-hacking guardrails are real, pure, and tested, and nothing calls them.
The framework beneath the mechanism is public and locked by
version. The ONEX runtime, the repository runtime that executes
the SQL contracts, and EnumEvidenceTier itself come from
omnibase-core, omnibase-infra,
omnibase-spi and omnimarket.
pyproject.toml declares them as version ranges, and
uv.lock resolves each from pypi.org — 0.47.24,
0.38.59, 0.23.5 and 0.4.261 at
this pin (uv.lock:2883, 2907,
2969, 2983). All four are MIT repositories
under github.com/OmniNode-ai, tagged per release. The
evidence guard's verdict depends on the enum's ordering, which section 4
reads at the locked tag.
2. Mental Model
A memory here is a learned pattern: a signature clustered from session events, with a domain, keywords, a confidence float, and two independent axes of standing.
The first axis is lifecycle status —
candidate → provisional → validated → deprecated — and it
decides whether the pattern may be injected into a session. The second
is evidence tier —
unmeasured → observed → measured → verified — and it
decides whether the pattern is allowed to advance. Status is
what the pattern is; tier is what is known about it.
Tier is computed from one question: was there a measured pipeline run, and did it succeed?
if run_id is None:
computed = EnumEvidenceTier.OBSERVED # anecdotal
elif run_result == "success":
computed = EnumEvidenceTier.MEASURED # quantitative
else:
computed = EnumEvidenceTier.OBSERVED # a run happened and didn't succeed
and it can only go up. The monotonicity is not a Python convention — it is the predicate of the statement that writes it:
UPDATE learned_patterns SET evidence_tier = $2, updated_at = NOW()
WHERE id = $1
AND (CASE evidence_tier WHEN 'unmeasured' THEN 0 WHEN 'observed' THEN 10
WHEN 'measured' THEN 20 WHEN 'verified' THEN 30 ELSE 0 END)
< (CASE $2 WHEN 'unmeasured' THEN 0 WHEN 'observed' THEN 10
WHEN 'measured' THEN 20 WHEN 'verified' THEN 30 ELSE 0 END)
A concurrent writer, a replayed Kafka message and a buggy caller all
fail the same way: the statement matches no rows. The migration names
one writer — "The attribution binder is the SOLE writer of this
column" — and the guard holds even if that stops being true for
updates. It does not cover inserts: an operator seeding script creates
rows born measured (section 7).
At runtime, status moves only through apply_transition,
which the demotion handler calls the single source of truth, and which
refuses → PROVISIONAL below OBSERVED and
→ VALIDATED below MEASURED. So the two axes
are wired together in one direction: evidence gates status, and status
never touches evidence.
The enum those gates compare is one rung stricter in its own
docstring. omnibase-core calls MEASURED
sufficient for PROVISIONAL and VERIFIED
"Required for VALIDATED lifecycle state in production"
(enum_evidence_tier.py:100, :107 at the locked
tag). apply_transition asks OBSERVED and
MEASURED. Read against that docstring, the tier nothing
writes is the one it names for production validation.
What kills a pattern is asymmetric on purpose. Promotion is optimistic, demotion conservative, and between them sits a 20-point band of success rate that belongs to neither.
Diagram source
%% caption: promotion needs injections, a success rate and an evidence tier together, and the band between forty and sixty per cent is hysteresis where nothing moves
stateDiagram-v2
[*] --> candidate: clustered from session events
candidate --> provisional: 5 injections · 60% success · tier >= observed
candidate --> candidate: bootstrap path selects unmeasured, reducer refuses it
provisional --> validated: 5 injections · 60% success · tier >= measured
validated --> deprecated: 10 injections · under 40% success · after 24h cooldown
validated --> deprecated: 5 consecutive failures
validated --> deprecated: manual disable, cooldown bypassed
validated --> validated: between 40% and 60% — the hysteresis band
deprecated --> [*]: excluded from every injection query3. Architecture
An event-driven Python service. ONEX nodes — 66 package directories,
each with a contract.yaml, handlers, models and its own
node_tests — are dispatched from Kafka topics through a
runtime plugin. Postgres is the only store this repository runs.
deployment/database/migrations/— 27 forward migrations, numbered000to028with026and027absent, and 20 rollback scripts.005islearned_patterns,006disable events,007injections,010the lifecycle audit,011the evidence tier,012measured attributions,022project scope.src/omniintelligence/nodes/node_pattern_*— learning, extraction, storage, promotion, demotion, feedback, lifecycle, projection, compliance, matching, assembly.src/omniintelligence/repositories/learned_patterns.repository.yaml— the read SQL, declared as data rather than embedded in Python.src/omniintelligence/api/router_patterns.py— the FastAPI surface that serves patterns to clients.src/omniintelligence/runtime/dispatch_handler_*.py— the wiring from Kafka envelope to handler.
Deployment and ergonomics
This is the heavy end of the atlas. An operator needs Postgres with 27 migrations applied, Kafka with the topics the node contracts declare, the ONEX runtime, and — for the loop to close — a separate repository running in the agent's session. It cannot run offline as a library; there is no embedded mode and no file-backed fallback.
The store is inspectable and repairable in the way a SQL store is: every state is a row, the audit is a table, and an operator who knows SQL can answer any question this report asks. That is a genuine advantage over the file-and-vector stores that dominate this corpus, and it is bought with an operational floor most single-user memory systems would not accept.
The framework is the larger adoption cost, and it is an open one. The
four framework packages are version ranges in
pyproject.toml, locked to PyPI releases, and each release
has a matching tag in a public MIT repository. The block declaring them
carries the comment "ONEX private ecosystem dependencies (NOT on
public PyPI)" (pyproject.toml:164), which the lockfile
beside it contradicts. Until 20 August 2026 omnibase-core
was a git-rev override on an unreleased commit, and ac6817ad…
removed it. Only the dev-group onex-change-control stays a
git source, at a public commit.
4. Essential Implementation Paths
Learning —
node_pattern_learning_compute
Session events are clustered, and each cluster is scored by
handler_confidence_scoring.py, which is explicit that a
single number is not the answer:
"Confidence scores are DECOMPOSED, not monolithic. A single confidence value cannot answer 'why did this pattern score high?' — but component scores can."
Three components with fixed weights — label_agreement
0.40, cluster_cohesion 0.30, frequency_factor
0.30 — asserted at import time to sum to 1.0, and a warning in capitals:
"NEVER make downstream decisions solely on the confidence field. The
rolled-up score is for convenience only. Inspect components."
learned_patterns has one column,
confidence FLOAT. The three components appear in no
migration. The advice is correct and the schema makes it impossible to
follow — and the bootstrap promotion gate, described below, decides on
lp.confidence >= 0.8.
Promotion
— handler_auto_promote.py, and the path that cannot
work
Two phases per run: CANDIDATE → PROVISIONAL, then
PROVISIONAL → VALIDATED. The candidate query admits two
kinds of row:
AND (
lp.evidence_tier IN ('observed', 'measured', 'verified')
OR (
lp.evidence_tier = 'unmeasured'
AND COALESCE(lp.injection_count_rolling_20, 0) = 0
AND lp.confidence >= 0.8
AND lp.recurrence_count >= 1
AND lp.distinct_days_seen >= 1
)
)
The second branch is the cold-start path, and its thresholds carry the clearest record in this repository of a gate being loosened under supply pressure:
"Lowered from 2 to 1: existing candidate patterns were initialized with recurrence_count=1 and distinct_days_seen=1. The confidence gate (>= 0.8) is the primary quality signal for the bootstrap path; requiring recurrence >= 2 blocked all 5,384 high-confidence candidates that had never been re-observed."
Each selected row is then passed to apply_transition_fn
with to_status=PROVISIONAL. In production that function is
the real apply_transition —
dispatch_handler_promotion_check.py imports it directly —
and it contains:
if (to_status == EnumPatternLifecycleStatus.PROVISIONAL
and current_evidence_tier < EnumEvidenceTier.OBSERVED):
return ModelTransitionResult(success=False, …,
reason="Evidence tier guard: insufficient evidence for PROVISIONAL", …)
There is no bootstrap exemption in the guard. Every row admitted by
the second branch is unmeasured, and every
unmeasured row is refused. The handler records
promoted=False, logs a warning, and continues; nothing
raises.
The refusal rests on the framework's comparison, and it holds.
EnumEvidenceTier is a str enum in
omnibase-core, and a plain one would compare
lexicographically, where "unmeasured" < "observed" is
false and every bootstrap row would pass. It overrides all four ordering
operators to compare weights of 0, 10, 20 and 30, coercing a raw string
operand first (enum_evidence_tier.py:27,
:126). The framework's tests assert the
non-lexicographic order against raw strings
(tests/unit/enums/pattern_learning/test_enum_evidence_tier.py:216).
Both halves are tested, separately.
test_candidate_to_provisional_with_unmeasured_rejected
calls the real apply_transition with an
unmeasured candidate and asserts failure
(tests/unit/nodes/node_pattern_lifecycle_effect/test_handler_transition_evidence_guards.py:320).
Every test that calls the auto-promote handler passes
mock_apply_transition, eleven call sites in all, so no test
joins the selection to the gate.
So the loosening cannot have unblocked the 5,384 candidates, because the reducer refuses them for a reason the loosened thresholds do not touch. The two halves are individually correct and disagree about what the cold-start rule is.
Demotion —
handler_demotion.py
The forgetting half, and the best-argued code in either repository.
Three gates, any one of which fires: manual disable (a hard trigger that
bypasses the cooldown), a failure streak of five, or a success rate
under 40% with at least ten injections in the rolling window. Two
eligibility checks precede gates two and three: the pattern must be
validated, and 24 hours must have passed since
promotion.
The override bounds are the part worth copying. An operator may tune
the thresholds per request, and the permitted range is itself argued:
SUCCESS_RATE_THRESHOLD_MAX = 0.60 exists to "prevent
setting demotion threshold at or above promotion threshold (60%) which
would cause immediate demotion of marginal patterns", and
FAILURE_STREAK_THRESHOLD_MIN = 3 prevents demoting "on
small runs of bad luck". The hysteresis band cannot be configured
away.
The handler also refuses to run without Kafka rather than degrading,
and says so: "Without Kafka, lifecycle events cannot be emitted, and
demotions will not occur. This is intentional — the reducer is the
single source of truth." A caller detects it by checking
reason == "kafka_producer_unavailable". Failing loudly on
the forgetting path is the right direction for the failure to
point.
The kill switch that does not reach the reader
Manual disable is an append-only event —
pattern_disable_events, with
event_type IN ('disabled', 're_enabled'), a
reason TEXT NOT NULL and an
actor VARCHAR(100) NOT NULL. Requiring a reason and an
actor on a memory override is rare here and right.
The demotion and promotion queries do not read that table. They join
disabled_patterns_current, a materialized
view created in migration 008, whose own comment
carries the operating instruction: "Refresh with: REFRESH
MATERIALIZED VIEW CONCURRENTLY disabled_patterns_current;".
Searching the tree for that statement returns three lines of migration
008 — two SQL comments and the view's
COMMENT ON text — and four calls inside
tests/integration/nodes/node_pattern_promotion_effect/test_promotion_integration.py.
No scheduler, no job, no dispatch handler refreshes it.
So the strongest correction the system offers — a person naming a pattern and a reason and turning it off — takes effect when somebody remembers a maintenance command that is documented in a SQL comment. The integration tests pass because they run it themselves.
That section is also why this report does not carry
human_review. Three things stand against it and any one
would be enough. The event is a disable, applied to a pattern
already promoted and already injected, so nothing waits in it. The actor
is a string the writer chooses, admitting "user, system, or
automated process" by the schema's own comment at
006_create_pattern_disable_events.sql:39. And the disable
does not reach a reader at all until the materialized view is refreshed
by hand. Requiring a reason and an actor on an override remains rare and
right, and it is recorded here for that rather than for a gate.
Outcome
feedback — node_enforcement_feedback_effect
The rule here is the one this atlas asks for and finds stated in code nowhere else. A violation only counts as negative evidence when the agent was both warned and observed to correct:
return [v for v in violations if v.was_warned and v.was_corrected]
"If an agent was warned of a violation but did NOT correct it, we cannot confirm the violation was real. The warning might have been a false positive."
That is a real epistemic distinction — between the memory fired and the memory was right — and most feedback loops in this corpus collapse it, treating every advisory as a labelled example and learning from their own false positives.
Audit —
pattern_lifecycle_transitions
Every status change writes a row: request_id for
idempotency, from_status, to_status,
transition_trigger, correlation_id,
actor, reason, and a
gate_snapshot JSONB holding the gate conditions at
transition time. The foreign key is ON DELETE RESTRICT
with the reason in a comment: "Audit records must never be silently
deleted when parent patterns are removed."
An audit that stores the evidence a decision was made on, rather than only the decision, is the version of this pattern worth having.
The guardrails nothing calls
node_anti_gaming_guardrails_compute implements four
defences as pure functions with 452 lines of tests: Goodhart detection
over correlated metric pairs, reward-hacking detection when a score
improves without matching human acceptance, distributional-shift
detection by symmetric KL approximation, and a diversity constraint that
is a veto. Its contract declares an operation and no topics. No module
under src/ outside its own package calls
run_all_guardrails; the one other mention is an allowlist
entry in a contract-validation test
(tests/unit/test_contract_validation.py:77).
The project records this itself. Its node inventory, linked from the
README and kept in the project's public knowledge base, lists seven node
directories as "(unregistered)", including this node's
companion alerter and node_objective_ab_framework_compute
(reference/omniintelligence-node-inventory.md).
A published list of what is not wired is a better artifact than most
projects offer, and it is the reason this finding is a citation rather
than an accusation.
5. Memory Data Model
learned_patterns carries identity
(pattern_signature, signature_hash,
is_current, version), a domain foreign key
with domain_candidates JSONB, keywords TEXT[],
confidence FLOAT CHECK (confidence >= 0.5), the
lifecycle columns, provenance (source_session_ids UUID[],
recurrence_count, first_seen_at,
last_seen_at, distinct_days_seen), and rolling
metrics over a window of 20.
The constraints are unusually load-bearing. Every rolling counter is
bounded to [0, 20], and:
CONSTRAINT check_rolling_metrics_sum
CHECK (success_count_rolling_20 + failure_count_rolling_20 <= injection_count_rolling_20)
An arithmetic invariant of the metrics that drive promotion, enforced by the database rather than by whoever writes next.
What is absent:
- No bi-temporal axis.
first_seen_atandlast_seen_atare observation times of the pattern, not validity times of a claim, and nothing separates when a pattern was true from when the row learned it. - No rejected-value tombstone.
deprecatedis keyed on the pattern row. The clustering path that produced the signature will produce it again from new sessions, as a newcandidate, and nothing consults the deprecation. - No confidence components, as above.
project_scopeis a column nothing can pass. Migration022adds it with two indexes, andlearned_patterns.repository.yamlapplies it:AND ($6::text IS NULL OR project_scope IS NULL OR project_scope = $6::text). The FastAPI endpoint that serves patterns declaresdomain,language,min_confidence,limitandoffset— and no project parameter — so$6is never bound on the network read path, and every project is served every project's patterns.
6. Retrieval Mechanics
Retrieval is SQL. list_validated_patterns filters on
status IN ('validated', 'provisional'),
is_current = TRUE, a confidence floor, an optional domain,
an optional keyword against keywords, and the unreachable
project predicate. Ordering puts validated before
provisional, then confidence descending, then id — a stable
order, which matters when the consumer takes a prefix.
The SQL in the contract is the SQL Postgres runs.
omnibase-infra's PostgresRepositoryRuntime
leaves a contract's WHERE clause untouched; for a multi-row
read it appends ORDER BY the primary key when none is
present, and LIMIT its max_row_limit when no
limit is present (postgres_repository_runtime.py:584,
:731). The default is ten rows, and this repository
constructs the runtime without a config. Every multi-row read in
learned_patterns.repository.yaml binds its own
LIMIT $n, so the default never truncates a pattern
read.
The framework also documents a boundary this read path leaves to the
client. omnibase-core's
EnumPatternLifecycleState gives PROVISIONAL
"Limited eligibility (testing/staging only)"
(enum_pattern_lifecycle_state.py:46 at the locked tag).
This repository defines its own EnumPatternLifecycleStatus,
and list_validated_patterns serves provisional
beside validated with no environment predicate, as the
bootstrap fallback its comment names.
There is no vector search, no embedding, no reranker and no query
rewriting in this repository's read path for patterns. For a store of a
few thousand short signatures selected by domain and confidence, that is
a defensible choice rather than a missing feature, and it is the reason
the retrieval stack reads as lexical alone.
The failure mode is over-supply at the client. The endpoint caps
limit at 200, and OmniClaude
requests ten times its injection budget precisely because its own
filters "each of which can eliminate the majority of
candidates" run after the fetch. The filtering that decides what an
agent sees happens on the far side of an HTTP boundary from the store
that knows the evidence.
7. Write Mechanics
Patterns are born from clustering over session events, dispatched
through Kafka. Outcomes arrive separately: injections are recorded by
the client, run attributions by the binder, violations by the
enforcement path. pattern_measured_attributions is the
evidence table, with a constraint tying tier to the presence of a
run:
CHECK ((run_id IS NULL AND evidence_tier = 'observed')
OR (run_id IS NOT NULL AND evidence_tier IN ('measured', 'verified')))
Attribution inserts are idempotent against Kafka redelivery by
INSERT … WHERE NOT EXISTS — chosen over
ON CONFLICT because the unique indexes are partial, with
the reasoning written beside it.
A second writer enters beside the ladder rather than through it.
scripts/seed_patterns_from_families.py --execute inserts
one row per ONEX node family found on disk, built by
node_family_to_pattern_row, born
status = 'validated' and
evidence_tier = 'measured' with no pipeline run and no row
in pattern_lifecycle_transitions
(handler_store_node_families.py:67, :75). On
conflict it refreshes confidence and the compiled snippet and leaves
status and tier alone, so a deprecated family stays deprecated. It is an
operator script, not a runtime node, and the migration's "SOLE
writer" comment describes the runtime.
Operational cost
- Nothing on the agent's turn blocks on this repository's writes. Learning, attribution and lifecycle work are all consumers of events; the agent's session has moved on.
- The lag from a session to an injectable pattern is unbounded by anything in the tree — it is clustering cadence, plus a promotion check, plus a demotion sweep, plus the tier the attribution binder can reach. Nothing measures it.
- The rolling window is 20 and the counters are capped at 20, so a pattern's standing is always computed over its last twenty injections rather than its history. That is a decay policy expressed as a bound, and a cheap one.
- No background pass rewrites the store. Promotion and demotion touch one row each, and the audit only grows.
8. Agent Integration
GET /api/v1/patterns and Kafka topics. The model never
talks to this repository; the client does, and the model sees the result
as text in its context. There are no MCP tools and no agent-callable
write path — an agent cannot save a memory here, which is a deliberate
shape: what becomes a pattern is decided by clustering over what agents
did, not by an agent deciding something is worth
remembering.
Adapting this to another host means reimplementing the client, which is most of OmniClaude's injection module, against an endpoint that is stable and small.
9. Reliability, Safety, and Trust
Strengths:
- Monotonicity enforced in the statement, not by the writer's discipline.
- Asymmetric promotion and demotion with a hysteresis band, argued in the constants themselves and bounded against operator override.
- An audit that records the gate conditions, not only the outcome, and a foreign key that refuses to lose it.
- Confirmed-only negative feedback — warned and corrected — so the loop does not learn from its own unconfirmed advisories.
- Arithmetic invariants in the schema, including the one that keeps the promotion metrics coherent.
- Idempotency against redelivery on both the attribution insert and the transition, with the choice of mechanism explained.
- Failing closed on the forgetting path when Kafka is absent, rather than silently not demoting.
- A published list of unregistered nodes, which is how the guardrail finding below is checkable at all.
Gaps:
- The bootstrap promotion path is refused by the reducer it calls.
verifiedis unreachable, so a four-tier ladder is a three-tier ladder with a gate nothing can satisfy.- The manual kill switch is behind an unrefreshed materialized view.
- The anti-gaming guardrails are not wired to anything.
project_scopecannot be passed by the API that serves the injector.- Confidence components are computed, warned about, and not persisted.
- A deprecated pattern's signature can be relearned as a new candidate from new sessions.
- An operator seeding script inserts rows born
validatedandmeasured, outside both the evidence ladder and the transition audit.
10. Tests, Evals, and Benchmarks
I ran nothing. The screen found no auto-executing
surfaces, 29 conftest.py files that execute at pytest
collection, and uv.lock inside the seven-day cooldown. The
findings above are all static, and so are the framework readings, taken
from the tags the lockfile names.
Per-node suites sit under node_tests/ alongside shared
tests/unit, tests/integration and
tests/integration/e2e.
The retrieval eval, and what it refuses to claim
src/omniintelligence/code_projection/retrieval_eval/ is
a committed harness over a committed corpus — a golden query set,
metrics, a replay driver, a scorecard and a baseline — and the
interesting part is what it says about its own limits rather than what
it scores.
The golden set is ten rows: six positives and four negatives. Its docstring gives the reason for the size rather than the number alone — the corpus exposes two retrievable chunk keys, so "six positives — two documents times three genuinely distinct intents — exhausts the honest positive space; a fourth phrasing per document would paraphrase one of those three rather than add a new correct answer." Volume comes from replaying each row at every state where it has a deterministic expected result, not from padding the set.
metrics.py then states which of its own metrics cannot
discriminate on that corpus. recall@10 is pinned at 1.00
for every N ≤ 10, "so it cannot express a failure below N =
11", and the module records the chance floors that would apply at N
= 20 — 0.50 for recall@10, 0.25 for recall@5, 0.05 for recall@1 and 0.18
for MRR — as an ARMING_MINIMUM_DISTINCT_DOCUMENTS threshold
the eval must clear before its ranking numbers mean anything.
A harness that ships the conditions under which its own headline
metric is uninformative is rare in this corpus, and it is the opposite
failure mode to the one the benchmarks
page spends its length on. The committed
baseline_scorecard.json carries the corpus manifest id, the
distinct chunk keys, an embedding compatibility key and an assertion
count, so a later run can be compared against it or refused as
incomparable.
The negative retrieval assertion is
tests/unit/repositories/test_contract_lifecycle_filter.py,
and it is an unusual shape. Rather than asserting about a result set, it
loads the repository contract and greps the SQL of every injection query
for a status it must not permit:
"DEPRECATED patterns have been demoted and should no longer be used. They MUST NOT be injected."
with a companion asserting the same for candidate.
Because it asserts over the query text, it covers every caller
of those operations at once, which an example-based test cannot, and the
framework's runtime executes that text with its WHERE
clause unchanged (section 6). It is also a regex over SQL, so a
differently-phrased predicate — a negated filter, a status passed as a
parameter — would satisfy it without meaning the same thing. Both halves
are worth knowing.
tests/unit/enums/test_enum_injection.py is the other
test worth naming: it asserts in both directions that the Python enum
and the database CHECK constraint contain the same values,
so the two cannot drift.
What is missing is the measurement the design implies. The evidence
tiers, the hysteresis band and the rolling windows are an argument that
this loop improves outcomes, and no committed result evaluates it; the
A/B framework node that would is documented as unregistered, and the
guardrails that would watch for the metrics being gamed are not called.
The most valuable missing test is narrower and would have caught the
headline finding: run the bootstrap path through the real
apply_transition and assert what happens. A guard test
asserts the refusal; the join between the two is what is absent.
No paper, arXiv reference or citation file exists in this repository.
11. For Your Own Build
Steal
- Put monotonicity in the
WHEREclause. A tier, a version or a state that must only advance should be guarded by the statement that writes it, so a concurrent writer and a redelivered message fail identically and silently-correctly. - Make demotion harder than promotion, and write down the band. More observations, a longer streak, a lower floor, and a cooldown after promotion. The 20-point gap between 60% and 40% is a design parameter that deserves a name and a comment.
- Bound the operator's override. If thresholds are tunable, refuse a demotion threshold at or above the promotion threshold. A knob that can erase a hysteresis band will.
- Snapshot the gates into the audit row. "Why was this promoted" is answerable from the audit only if the audit holds the conditions, not just the verdict.
- Count a violation only when it was surfaced and then corrected. An advisory the agent ignored is not evidence the advisory was right, and a loop that treats it as one trains on its own false positives.
- Require a reason and an actor on a manual override,
as a
NOT NULLcolumn rather than a convention. - Assert enum-versus-constraint parity in both directions. It is ten lines and it catches the drift that produces a runtime constraint violation months later.
- Enforce metric arithmetic in the schema — successes plus failures cannot exceed injections — so the numbers that drive promotion cannot go incoherent.
- Publish which of your components are unregistered.
Avoid
- A selection query and an enforcement gate with different rules. If one layer chooses rows and another refuses them, the system reports zero and nobody looks; put the gate where the selection happens, or make the selection call the gate.
- Mocking the component under test in the test named for the whole flow. An integration test that replaces the gate with a stub returning success asserts only that the caller is wired, and reads as though it asserts the behaviour.
- A state at the top of a ladder that nothing writes.
Either something reaches
verifiedor the ladder has three rungs; a permanently empty top tier makes every gate that mentions it unsatisfiable and every reader assume it is reachable. - A kill switch behind a manual refresh. Any correction a human can make must reach the read path without a second, undocumented human action.
- Decomposing a score and persisting only the roll-up. The warning not to decide on the composite is unfollowable if the components are gone by the time anyone can decide.
- A scope column the serving API cannot accept. Storing a boundary that no caller can pass is the same as not having one, and it reads like having one.
Fit
This suits a team, not a person: a platform group running Kafka and Postgres who want memory to be an auditable process with a lifecycle, and who will staff the operations that lifecycle implies. The design assumes a maintainer who thinks in gates, evidence and windows, and it rewards that — the argued constants and the gate snapshots are worth more than most of the retrieval sophistication elsewhere in this corpus.
Walk away if you want memory as a library. There is no embedded mode, the framework beneath it is four packages it cannot run without, and the smallest useful deployment is a message bus plus a database plus a second service inside the agent.
The uncomfortable judgement is about the gap between the design and its wiring. Four of the mechanisms this report praises — the bootstrap path, the top evidence tier, the kill switch, the guardrails — are specified more completely than they are connected, and each was found by following one call rather than by reading a diagram. A reader borrowing from here should borrow the reasoning, which is excellent and portable, and verify the wiring in their own tree rather than assuming it.
12. Open Questions
- Has the bootstrap path ever promoted a pattern in production, and what did the 5,384 candidates do after the thresholds were lowered?
- What is intended to write
verified— an independent validator, a human, a second pipeline — and does the design want the tier gated on something that does not exist yet? - Is
disabled_patterns_currentrefreshed by a deployment job outside the public repositories?omnibase_infraat the locked tag carries a copy of migration008and no refresh of it. - What consumes the anti-gaming guardrails in the intended design, and what topic would carry their alerts?
- How long is the median path from a session event to an injectable validated pattern?
- Which tier is
VALIDATEDmeant to require:MEASURED, asapply_transitionenforces, orVERIFIED, as theomnibase-coreenum's docstring states?
Appendix: File Index
- Schema and constraints:
deployment/database/migrations/005_create_learned_patterns.sql,006_create_pattern_disable_events.sql,007_create_pattern_injections.sql,008_create_disabled_patterns_current_view.sql,010_create_pattern_lifecycle_audit.sql,011_add_evidence_tier_to_learned_patterns.sql,012_create_pattern_measured_attributions.sql,022_add_project_scope_to_learned_patterns.sql. - Evidence tier and attribution:
src/omniintelligence/nodes/node_pattern_feedback_effect/handlers/handler_attribution_binder.py. - Lifecycle gates:
src/omniintelligence/nodes/node_pattern_lifecycle_effect/handlers/handler_transition.py. - Promotion:
src/omniintelligence/nodes/node_pattern_promotion_effect/handlers/handler_auto_promote.py; wiring insrc/omniintelligence/runtime/dispatch_handler_promotion_check.py. - Demotion:
src/omniintelligence/nodes/node_pattern_demotion_effect/handlers/handler_demotion.py. - Confidence decomposition:
src/omniintelligence/nodes/node_pattern_learning_compute/handlers/handler_confidence_scoring.py. - Confirmed-only feedback:
src/omniintelligence/nodes/node_enforcement_feedback_effect/handlers/handler_enforcement_feedback.py. - Guardrails:
src/omniintelligence/nodes/node_anti_gaming_guardrails_compute/handlers/handler_guardrails.py. - Read SQL and API:
src/omniintelligence/repositories/learned_patterns.repository.yaml,src/omniintelligence/api/router_patterns.py. - Seeding writer:
scripts/seed_patterns_from_families.py,src/omniintelligence/nodes/node_pattern_extraction_compute/handlers/handler_store_node_families.py. - Framework versions:
pyproject.toml(ranges, and the removed git-rev overrides in comments),uv.lock. - Tests cited:
tests/unit/repositories/test_contract_lifecycle_filter.py,tests/unit/enums/test_enum_injection.py,tests/unit/nodes/node_pattern_lifecycle_effect/test_handler_transition_evidence_guards.py,tests/unit/nodes/node_pattern_promotion_effect/test_handler_auto_promote.py,tests/integration/test_promotion_lifecycle_integration.py,tests/integration/nodes/node_pattern_promotion_effect/test_promotion_integration.py. - Framework, at the locked tags:
omnibase_corev0.47.24—src/omnibase_core/enums/pattern_learning/enum_evidence_tier.py,enum_pattern_lifecycle_state.py,tests/unit/enums/pattern_learning/test_enum_evidence_tier.py;omnibase_infrav0.38.59—src/omnibase_infra/runtime/db/postgres_repository_runtime.py,src/omnibase_infra/runtime/db/models/model_repository_runtime_config.py,docker/migrations/intelligence/008_create_disabled_patterns_current_view.sql. - Self-documented wiring: the node inventory in the project's
knowledge base,
reference/omniintelligence-node-inventory.md, linked fromREADME.md.
Recorded searches
Run from the repository root at the pinned commit unless a framework checkout is named.
git grep -n -A2 -E '^name = "(omnibase-core|omnibase-infra|omnibase-spi|omnimarket)"' -- uv.lock # four registry sources, no git
git grep -n -E '^(omnibase|omnimarket)[a-z_-]* *= *\{' -- pyproject.toml # nothing: no framework git source
git grep -n -E 'EnumEvidenceTier\.VERIFIED|evidence_tier *= *.verified' -- src # nothing writes verified
git grep -n -E 'REFRESH MATERIALIZED VIEW' -- . # migration 008 comments and one integration test
git grep -n -E 'run_all_guardrails' -- . ':!src/omniintelligence/nodes/node_anti_gaming_guardrails_compute' # one test allowlist entry
git grep -n -E 'apply_transition_fn=' -- tests 'src/*/node_tests/*' # eleven sites, all mock_apply_transition
git grep -n -E 'label_agreement|cluster_cohesion|frequency_factor' -- deployment # nothing: components not persisted
git grep -n -E 'project_scope' -- src/omniintelligence/api/router_patterns.py # nothing: no project parameter
git grep -n -E 'pattern_lifecycle_transitions' -- scripts/seed_patterns_from_families.py # nothing: seeding writes no audit row
git grep -n -E 'def __(lt|le|gt|ge)__' -- src/omnibase_core/enums/pattern_learning/enum_evidence_tier.py # in omnibase_core v0.47.24: four overrides
git grep -n -E 'REFRESH MATERIALIZED VIEW.*disabled_patterns_current' -- . # in omnibase_infra v0.38.59: comment in the 008 copy only
History
2026-09-28 — re-pinned to 25ee01e7…,
37 commits of CI and dependency bumps; no mechanism file changed. The
framework was described as three private repositories, and
omnibase-core as a git-rev pin. Both were wrong at the
previous pin: the override was removed on 20 August 2026, and the
repositories were public, with release sources on PyPI. Read at the
locked tags, EnumEvidenceTier compares by weight, so the
headline refusal holds (section 4). Added: a
seeding script that inserts validated/measured
rows outside the ladder, and a guard test asserting the refusal.
NODE_INVENTORY.md left the tree on 2 September 2026 and is
cited in the knowledge base. Marks unchanged. Screened: no auto-run
surface, 29 build-time execution points, uv.lock inside the
cooldown. Nothing was installed, built or run.
2026-09-19 — re-pinned to 274f097f….
human_review is withdrawn on three
grounds, each already written in this report. The record itself supplied
the second: the actor column "does not distinguish a human actor
from an automated one — the column admits
user, system, or automated process by its own
comment", so the mark rested on the surface existing rather than on
who used it. The first is that a pattern_disable_events row
is a disable applied to a pattern already promoted and already injected,
so no memory waits in it. The third is section 4's own finding: the
disable does not reach any reader until
REFRESH MATERIALIZED VIEW CONCURRENTLY disabled_patterns_current
is run by hand, and nothing in the tree runs it outside the integration
tests. Requiring a reason and a named actor on an override is still rare
and right, and it keeps that credit. trust_state,
audit_log and negative_eval stand. Screened
again first; nothing was installed, built or run.
2026-09-16 — 5039c8a3…
— re-read at a commit dated 2026-09-16, 11 commits past the previous
pin. All four marks re-tested and held; each anchored file has the same
blob at both commits, so no line number moved. Screened before reading:
no auto-run surface, 29 build-time execution points, no unpinned
dependency surface and one dependency file inside the seven-day
cooldown; an agent-addressed instruction file was recorded as data.
Nothing was installed, built or run.
2026-09-09 — ce104631…
— second reading, 59 commits along the default dev branch:
226 files, 23,928 insertions. Screened before reading; nothing was
installed and no suite was run.
The four marks are untouched. Every path this report's appendix names
still exists, and the diff across all of them is 71 files and 172
insertions, most of it one-line contract.yaml edits — so
the lifecycle audit table, the disable-events table, the transition
gate, the guardrails and the contract-filter test are all as
described.
The 24,000 insertions landed almost entirely in a subsystem the
report did not cover: code_projection, with
context_serving and retrieval_eval beside it.
Section 10 now covers the eval, because it is the part with something to
say to this atlas. It is a committed harness over a committed corpus
whose golden set is ten rows, and both the set size and the metric
selection are argued rather than asserted — the docstring explains why a
seventh positive would be a paraphrase rather than a new answer, and
metrics.py records that recall@10 is pinned at
1.00 below eleven distinct documents and therefore "cannot express a
failure", along with the chance floors that would apply once the corpus
is large enough to arm it.
Six new migrations arrived (020 through 028, with debug-intelligence tables, LLM routing decisions, review calibration runs and an agent-actions retention policy); none of them touches the pattern lifecycle the marks rest on.
2026-08-11 — 8c67665a…
— first reading, on the dev default branch. Screened before
reading: 0 auto-run surfaces, 29 build-time exec surfaces (all
conftest.py), 0 unpinned manifests, and a
uv.lock unchanged for 16 days; nothing was installed and
nothing was executed. The three
omnibase-*/omnimarket git dependencies were
not readable at this reading, so no claim here rests on them.