Structured long-term memory per persona. Backed by Postgres with tsvector full-text search always-on and pgvector semantic search when an OpenAI key is configured.
Memory is for journal/events only: timestamped occurrences, project milestones, and specific conversations worth recalling. Stable facts about who the user is or how they work belong in the user model.
memories
id uuid primary key
persona string not null -- "main", "trader", etc.
category string -- "preference", "position", …
content text not null
importance int 1..10 default 5
day date not null -- for `recall {date: …}`
embedding vector(1536) -- nullable, filled async
search_vector tsvector GENERATED ALWAYS -- weighted: category A, content B
inserted_at timestampIndexes: GIN on search_vector, HNSW on embedding (cosine ops),
btree on (persona, importance, inserted_at), btree on (day).
Tool-driven:
memory {action: "write", category: "journal", content: "…", importance: 7}
or from Elixir:
Neoharness.Memory.Memories.write(%{
persona: "trader",
category: "project",
content: "On 2026-05-08, the deploy milestone finished cleanly.",
importance: 9
})After insert, an Oban job (Neoharness.Memory.EmbeddingWorker)
fetches an embedding and stores it. The database contract is always
1536 dimensions: text-embedding-3-* requests set dimensions: 1536,
and responses of any other size are rejected. If the key isn't set, the
row keeps embedding = NULL and search falls back to tsvector-only.
a memory write only fires when the agent decides to call it, and an agent left purely to its own judgment drifts — durable facts slip by unsaved. Three triggers keep capture reliable without doubling the foreground turn cost:
- Immediate explicit capture. The foreground agent calls
memory(actionwrite) oruser_model_writeimmediately when the current user message contains an obviously durable fact: explicit "remember this" instructions, stable preferences, corrections, goals, working patterns, timestamped events, project milestones, or specific conversations worth recalling. This is the fast path for a single important message — the agent should not wait for periodic review. - Periodic background review on user turns. Every
@memory_review_every_user_msgssuccessful user turns in a conversation (default 8),Agent.ServerenqueuesNeoharness.Scheduler.MemoryReviewerwith a compact recent transcript slice. The counter lives onconversations.user_msgs_since_nudgeand is bumped/reset atomically. The review is a single structured LLM call on the:memory_reviewOban queue; it dedupes against the current user model and top memories, then writes high-confidence user-model updates and event memories directly. No[memory check]system message is staged into the live user turn. - Pre-compaction flush. Before
Summarizer.compact/3folds old messages into a prose summary (which truncates to ~4k chars and drops "small talk"),CompactorWorkerruns a singleMemoryFlush.flush/2LLM call on the exact slice about to be folded. The call asks the model to return a JSON array of{category, content, importance}event memories; we parse it and callMemories.write/1per memory. Single-shot by design — no tool loop, no iterative roundtrips, bounded cost. Failures (LLM error, malformed JSON) are non-fatal; summarization always proceeds.CompactorWorkercomputes the fold plan once and hands it to both the flush and the compaction so we don't double-query.
Tool-driven:
recall {query: "what does the user drink?", limit: 5}
From Elixir:
Neoharness.Memory.Memories.search(
"what does the user drink?",
persona: "main", limit: 5
)When the query can be embedded (OpenAI configured) and the row has an embedding:
score =
0.3 * ts_rank(search_vector, to_tsquery(query))
+ 0.5 * (1 - cosine_distance(embedding, query_vec))
+ 0.2 * importance / 10
multiplied by
1 / (1 + days_since_creation)
Without OpenAI, the vector term collapses and text-only ranking kicks in — same formula, coefficient 0.3 becomes effectively 1.0 for that term.
Stored: "The user enjoys quiet mornings with black coffee and jazz."
Query: "what does the user like to drink in the morning?"
No token overlap on "drink"/"coffee" — tsvector alone scores zero. Cosine similarity on the embeddings surfaces the row cleanly. See the integration test in README for the exact trace.
The top 20 memories by importance+recency for the current persona are injected into the system prompt:
## Top memories for this persona
- [journal, imp=9] On 2026-05-08, the user asked to migrate USER.md into user_model.
- [project, imp=9] The neoharness memory contract changed to make memories event-only.
- …
This means the agent often has the event context it needs without calling
recall — but for anything not in the top 20, the tool's there.
The top-20 list is snapshotted per conversation, not recomputed
each turn. On first use of a conv, the rendered memories block is
written to conversations.memory_snapshot and reused verbatim on
every subsequent turn — so the full system prompt is byte-for-byte
stable and the provider's prefix cache serves it for free.
If we rebuilt the block each turn, every memory write (from immediate foreground capture, periodic background review, or pre-compaction flush) would reshuffle the ranking and invalidate the prefix cache — paying full system-prompt tokens on the next turn. The snapshot fixes that.
The snapshot is refreshed at exactly two moments:
- After compaction.
CompactorWorkernulls the cached snapshot after a successful fold. The next turn renders a fresh top-20 from current memories (including any freshly-flushed ones from the pre-compaction MemoryFlush) and caches it again. - On
/clear. The purge path clears the canonical conv row's snapshot (along with messages + rolling summary), so the new transcript starts with a fresh anchor.
Between those moments, newly-written memories are still accessible
live via recall — they just don't bubble into the system
prompt until the next refresh. This is the right tradeoff for a 200k
context budget: cache hits on the fixed prefix dominate.
If you wrote memories before configuring OPENAI_API_KEY:
Neoharness.Memory.Memories.backfill_missing_embeddings()
# => N (number of jobs enqueued)Each job retries up to 5 times with exponential backoff; a transient OpenAI outage doesn't need manual intervention.
Alongside the append-only memory stream, each persona keeps a compact,
upsert-in-place user model — the source of truth for who the user
is and how they work. Stable facts, preferences, working patterns,
feedback, and goals go here, not in memory. One row per
(persona, dimension); writing the same
dimension twice replaces the previous value.
user_models
id uuid primary key
persona string not null
dimension string not null -- e.g. "baseline"
content text not null
confidence int 1..10 default 5
inserted_at timestamp
updated_at timestampUnique index on (persona, dimension).
Free-form names are allowed, but the agent is nudged toward these:
baseline— name, birthday, nationality, work, home location, routinepreferences— evolving likes, dislikes, tools, formats, defaultsworking_patterns— when and how the user works bestfeedback— explicit "do X" / "don't Y" performance signalsgoals— what the user is driving toward right now
user_model_write {dimension, content, confidence?} → upsert
(no read tool — the full model is injected into the prompt every turn)
Every turn, the full user model for the active persona is rendered verbatim into the system prompt as:
## User model (call `user_model_write` to update a dimension)
- [baseline, conf=9] lives in Seoul and works on neoharness
- [working_patterns, conf=8] deep focus 9-12, flexible after lunch
- [preferences, conf=7] markdown for long-form, plain text for chat
Unlike the top-20 memories block (which is snapshotted per conv for prefix-cache stability), the user model is read fresh on every turn. Writes are rare and small (at most one row per dimension), so rebuild cost is trivial — and a just-written dimension shows up on the next turn without waiting for a compaction.
| You want to remember… | Write to… |
|---|---|
| Stable facts about who the user is | user model |
| How the user works or wants to be supported | user model |
| Preferences, feedback, or goals | user model |
| Timestamped occurrence or project milestone | memory |
| A specific conversation worth recalling | memory |
import Ecto.Query
from(m in Neoharness.Memory.MemoryEntry, where: m.persona == "trader")
|> Neoharness.Repo.delete_all()Tool-driven equivalent: recall to find ids, then memory action delete
one by one. (Intentionally no bulk-delete tool — an LLM mass-deleting
memory without explicit user confirmation is how you lose months of
state.)
The third long-term store, next to memories and the user model:
reference documents (specs, manuals, articles) ingested as chunks into
knowledge_chunks and retrieved through the same recall tool.
knowledge_chunks
id uuid pk
persona text not null default 'main'
source text not null -- stable document name
chunk_index int not null -- order within the document
content text not null
embedding vector(1536) -- nullable; same model as memories
search_vector tsvector (generated) -- source weighted A, content B
- Ingest — the
knowledge_ingesttool (workspace file, PDF, or raw text + asourcename). Documents split into ~3000-char paragraph-aligned chunks with a 200-char overlap; all chunks are embedded in one batched call. Re-ingesting a source replaces its chunks, so refreshed documents never duplicate. - Read —
recallsearches knowledge alongside memories and tags results[knowledge <source>#<chunk>]. Ranking is0.4 × ts_rank + 0.6 × cosine— no importance or recency factors, since document chunks don't age the way journal memories do. Without an embeddings key, ingest stores text-only and search degrades to full-text, same contract as memories. - Not prompt-injected — unlike the top-memories block, knowledge is retrieved only on demand, so a large corpus costs nothing per turn.
- Routing rule of thumb — chat-derived facts go to memory or user_model; documents worth consulting later go to knowledge.