MEMORY / COGNITIVE MEMORY

Remembered, not assumed.

CognitiveMemory learns durable facts from your turns and puts the relevant ones in front of the model. It is deterministic where it can be: what gets stored, what gets pre-staged, and what a recall returns do not depend on which model you configured.

Four tiers, one prompt block

TierNameContents
L0RegistersActive tensions, the proprioceptive self-model, and the fast-gate notice. Bodies only.
L1Hot cachePre-staged memories. Rendered as a one-line index; a body is spent only when a trigger earns it.
L2Warm storeIndexed candidates offered to the arbiter for promotion, capped at 20 per turn.
L3Cold archiveHistorical memories that stay out of the prompt until something promotes or recalls them.

Nothing here is required for the agent loop to run — memory is an additive layer you construct yourself.

What gets learned, and what does not

Naive memory only catches phrasings like 'always use kebab-case' and silently drops everything else. These rules exist because that made memory look broken: a stated fact simply never reached the tiers.

Deterministic first

Pattern matching runs before any model call, so a plainly-stated fact is captured even when the model refuses, hedges, or returns nothing.

Then the model

The same model that ran the turn extracts facts, preferences, and conventions in the user's own words.

Never from questions

A lookup is not a lesson. Extraction is skipped when your message contains a question mark.

Interaction-scoped text is dropped

Instructions about how to behave right now, such as 'do not verify this against the repo', are not durable project facts.

Paraphrases collapse

Two restatements of one fact become one entry, so the index cannot fill with copies of itself.

A newly taught memory lands in L1 immediately. Waiting for the arbiter to promote it added a turn of latency, so a fact you taught on one turn was still missing from the next prompt.

Index by default, bodies on demand

Always injecting every memory body is expensive and injects distractors by construction. Instead each turn gets a one-line index of everything remembered, and a full body is spent only where a deterministic signal earns it.

The trigger is the same idea Aider uses for its repo map: pull identifiers out of your message — URLs, paths, SCREAMING_SNAKE, camelCase, long kebab-case, hex-ish codes — and keep only those absent from the visible transcript. If you name something concrete the model cannot already see, and a memory mentions it, that memory's body is included. No model decision is involved.

ReasonWhat was included
indexOne gist line. The default, and what nearly everything costs.
triggerFull body — you named an identifier absent from the visible transcript and a memory matched it.
tensionFull body. Unresolved contradictions earn the tokens.
guardrailFull body. A domain this agent is unreliable in.
textnah
## Cognitive Memory State
### Memory index — established earlier in this project
- staging build ID (deployment)
- file naming: kebab-case (naming)
- internal staging host is internal-hbr-2291.example (infra)

maxTotalTokens caps everything injected in one prompt, index and bodies together. When the cap bites the report says so rather than silently dropping the tail.

The recall tool

Pre-staging is a best-effort optimisation: a model may not read the index, and a question sharing no words with a memory will not have it staged. The recall tool lets the agent ask directly. Ranking runs in-process with the same token-overlap scoring used for promotion, so recall quality does not depend on model quality.

textnah
recall { query: 'staging build id', limit: 6 }

// Remembered (1 match):
// - The staging build ID is ZQ7X4M2K. (deployment) [L1, relevance 0.75]

When nothing matches it says so, and tells the model to say it was not told rather than guessing. It is read-only, so it is not approval-gated. The CLI registers it automatically alongside the coding tools.

Seeing what it cost

There is no published benchmark for pre-inject versus on-demand retrieval for coding-agent project memory, so the only way to tune the tradeoff is to watch it. Every injection is recorded with its reason and token cost.

textnah
/memory

Prompt injection log (what memory added, and why)
  7 turn(s) · avg 58 tokens · max 83
  index=25
  turn 7: 58 tokens
    index [index/L1] staging build ID is ZQ7X4M2K
    index [index/L1] internal staging host is internal-hbr-2291.example

Use it to decide whether the budget is too tight, whether the trigger is firing when it should, and whether an index line is enough or the model keeps reaching for recall.

Promotion and the arbiter

After extraction, warm candidates are scored against the current turn by content-word overlap, and the best few are promoted. Scoring used to compare a single domain tag against the turn text with a substring match, which almost never fired — a memory tagged naming-conventions cannot match a question that says 'naming'. Promotion was effectively random.

Supply arbiter to decide tiers with a model instead. An arbiter is just a function receiving the turn text, the assistant reply, the L0 prompt, the L1 summaries and the L2 candidates, returning a structured ArbiterEvaluationResult. One that throws returns an empty result rather than failing the turn.

tsnah
import { CognitiveMemory } from '@astracollab/not-another-harness';

const memory = new CognitiveMemory({
  extract: myExtractor,       // omit to use only the built-in patterns
  arbiter: createModelArbiter({ model: arbiterModel }),
  maxTotalTokens: 2000,
});

Self-model and knowledge tensions

The self-model tracks per-domain reliability as a moving average, plus known failure patterns and recommended strategies. Any active domain below 75% reliability is rendered into L0 as a guardrail, so weak areas get explicit attention instead of confident guesses.

Active tensions are pinned into every prompt block until resolved, each paired with an actionable question so the agent asks rather than guesses. addTension, resolveTension and recordDomainOutcome drive both.

Persistence

The CLI stores memory per working directory under ~/.nah/memory and reloads it on the next run, so a fact you taught yesterday is available in a session that starts today. Set NAH_MEMORY_NOPERSIST=1 to keep memory in-process only.

Memory is entirely in-process plus the JSON snapshot; the agent is never asked to author or maintain memory files by hand.

API surface

planInjection({ userMessage?, forceFull? })Returns text, entries, totalTokens and truncated. The primary entry point.
getPromptContext(message?)Prompt text only. Prefer planInjection when you want to know what was included.
search(query, limit?)Deterministic ranked lookup across all tiers. Backs the recall tool.
postTurnAsync({ userMessage, assistantResponse })Extraction, arbitration, budget enforcement, and persistence.
addMemory(item, targetTier)Adds or promotes a MemoryItem into a tier. Defaults to L2.
addTension(tension) / resolveTension(id, resolution)Registers or closes a contradiction pinned in L0.
recordDomainOutcome(domain, success, failurePattern?)Updates the moving-average reliability score for a domain.
getSnapshot() / loadSnapshot(snapshot)Serialize or restore the full L0–L3 state and stats.

Types for every item above — MemoryItem, MemoryInjectionReport, MemoryInclusionReason, KnowledgeTension, CognitiveMemoryOptions — are exported from the package root alongside the class.

There is also packages/nah/eval-memory.mts, which teaches facts containing unguessable tokens, wipes the transcript, and asks again — so anything the model can still answer came out of memory rather than context.

How memory fits into the loop →