Skip to content

Latest commit

 

History

History
306 lines (242 loc) · 11.2 KB

File metadata and controls

306 lines (242 loc) · 11.2 KB

Memory

Structured long-term memory per persona. Backed by Postgres with tsvector full-text search always-on and pgvector semantic search when an OpenAI key is configured.

Memory is for journal/events only: timestamped occurrences, project milestones, and specific conversations worth recalling. Stable facts about who the user is or how they work belong in the user model.

Schema

memories
  id              uuid primary key
  persona         string not null            -- "main", "trader", etc.
  category        string                     -- "preference", "position", …
  content         text not null
  importance      int 1..10 default 5
  day             date not null              -- for `recall {date: …}`
  embedding       vector(1536)               -- nullable, filled async
  search_vector   tsvector GENERATED ALWAYS  -- weighted: category A, content B
  inserted_at     timestamp

Indexes: GIN on search_vector, HNSW on embedding (cosine ops), btree on (persona, importance, inserted_at), btree on (day).

Writes

Tool-driven:

memory {action: "write", category: "journal", content: "…", importance: 7}

or from Elixir:

Neoharness.Memory.Memories.write(%{
  persona: "trader",
  category: "project",
  content: "On 2026-05-08, the deploy milestone finished cleanly.",
  importance: 9
})

After insert, an Oban job (Neoharness.Memory.EmbeddingWorker) fetches an embedding and stores it. The database contract is always 1536 dimensions: text-embedding-3-* requests set dimensions: 1536, and responses of any other size are rejected. If the key isn't set, the row keeps embedding = NULL and search falls back to tsvector-only.

Save triggers

a memory write only fires when the agent decides to call it, and an agent left purely to its own judgment drifts — durable facts slip by unsaved. Three triggers keep capture reliable without doubling the foreground turn cost:

  1. Immediate explicit capture. The foreground agent calls memory (action write) or user_model_write immediately when the current user message contains an obviously durable fact: explicit "remember this" instructions, stable preferences, corrections, goals, working patterns, timestamped events, project milestones, or specific conversations worth recalling. This is the fast path for a single important message — the agent should not wait for periodic review.
  2. Periodic background review on user turns. Every @memory_review_every_user_msgs successful user turns in a conversation (default 8), Agent.Server enqueues Neoharness.Scheduler.MemoryReviewer with a compact recent transcript slice. The counter lives on conversations.user_msgs_since_nudge and is bumped/reset atomically. The review is a single structured LLM call on the :memory_review Oban queue; it dedupes against the current user model and top memories, then writes high-confidence user-model updates and event memories directly. No [memory check] system message is staged into the live user turn.
  3. Pre-compaction flush. Before Summarizer.compact/3 folds old messages into a prose summary (which truncates to ~4k chars and drops "small talk"), CompactorWorker runs a single MemoryFlush.flush/2 LLM call on the exact slice about to be folded. The call asks the model to return a JSON array of {category, content, importance} event memories; we parse it and call Memories.write/1 per memory. Single-shot by design — no tool loop, no iterative roundtrips, bounded cost. Failures (LLM error, malformed JSON) are non-fatal; summarization always proceeds. CompactorWorker computes the fold plan once and hands it to both the flush and the compaction so we don't double-query.

Reads

Tool-driven:

recall {query: "what does the user drink?", limit: 5}

From Elixir:

Neoharness.Memory.Memories.search(
  "what does the user drink?",
  persona: "main", limit: 5
)

Ranking formula

When the query can be embedded (OpenAI configured) and the row has an embedding:

score =
  0.3 * ts_rank(search_vector, to_tsquery(query))
+ 0.5 * (1 - cosine_distance(embedding, query_vec))
+ 0.2 * importance / 10
  multiplied by
  1 / (1 + days_since_creation)

Without OpenAI, the vector term collapses and text-only ranking kicks in — same formula, coefficient 0.3 becomes effectively 1.0 for that term.

Semantic search worked example

Stored: "The user enjoys quiet mornings with black coffee and jazz."

Query: "what does the user like to drink in the morning?"

No token overlap on "drink"/"coffee" — tsvector alone scores zero. Cosine similarity on the embeddings surfaces the row cleanly. See the integration test in README for the exact trace.

System prompt injection

The top 20 memories by importance+recency for the current persona are injected into the system prompt:

## Top memories for this persona
- [journal, imp=9] On 2026-05-08, the user asked to migrate USER.md into user_model.
- [project, imp=9] The neoharness memory contract changed to make memories event-only.
- …

This means the agent often has the event context it needs without calling recall — but for anything not in the top 20, the tool's there.

Snapshotting for cache stability

The top-20 list is snapshotted per conversation, not recomputed each turn. On first use of a conv, the rendered memories block is written to conversations.memory_snapshot and reused verbatim on every subsequent turn — so the full system prompt is byte-for-byte stable and the provider's prefix cache serves it for free.

If we rebuilt the block each turn, every memory write (from immediate foreground capture, periodic background review, or pre-compaction flush) would reshuffle the ranking and invalidate the prefix cache — paying full system-prompt tokens on the next turn. The snapshot fixes that.

The snapshot is refreshed at exactly two moments:

  1. After compaction. CompactorWorker nulls the cached snapshot after a successful fold. The next turn renders a fresh top-20 from current memories (including any freshly-flushed ones from the pre-compaction MemoryFlush) and caches it again.
  2. On /clear. The purge path clears the canonical conv row's snapshot (along with messages + rolling summary), so the new transcript starts with a fresh anchor.

Between those moments, newly-written memories are still accessible live via recall — they just don't bubble into the system prompt until the next refresh. This is the right tradeoff for a 200k context budget: cache hits on the fixed prefix dominate.

Backfilling embeddings

If you wrote memories before configuring OPENAI_API_KEY:

Neoharness.Memory.Memories.backfill_missing_embeddings()
# => N   (number of jobs enqueued)

Each job retries up to 5 times with exponential backoff; a transient OpenAI outage doesn't need manual intervention.

User model

Alongside the append-only memory stream, each persona keeps a compact, upsert-in-place user model — the source of truth for who the user is and how they work. Stable facts, preferences, working patterns, feedback, and goals go here, not in memory. One row per (persona, dimension); writing the same dimension twice replaces the previous value.

Schema

user_models
  id            uuid primary key
  persona       string not null
  dimension     string not null     -- e.g. "baseline"
  content       text not null
  confidence    int 1..10 default 5
  inserted_at   timestamp
  updated_at    timestamp

Unique index on (persona, dimension).

Recommended dimensions

Free-form names are allowed, but the agent is nudged toward these:

  • baseline — name, birthday, nationality, work, home location, routine
  • preferences — evolving likes, dislikes, tools, formats, defaults
  • working_patterns — when and how the user works best
  • feedback — explicit "do X" / "don't Y" performance signals
  • goals — what the user is driving toward right now

Tools

user_model_write {dimension, content, confidence?}   → upsert
(no read tool — the full model is injected into the prompt every turn)

System-prompt injection

Every turn, the full user model for the active persona is rendered verbatim into the system prompt as:

## User model (call `user_model_write` to update a dimension)
- [baseline, conf=9] lives in Seoul and works on neoharness
- [working_patterns, conf=8] deep focus 9-12, flexible after lunch
- [preferences, conf=7] markdown for long-form, plain text for chat

Unlike the top-20 memories block (which is snapshotted per conv for prefix-cache stability), the user model is read fresh on every turn. Writes are rare and small (at most one row per dimension), so rebuild cost is trivial — and a just-written dimension shows up on the next turn without waiting for a compaction.

When to use memory vs user model

You want to remember… Write to…
Stable facts about who the user is user model
How the user works or wants to be supported user model
Preferences, feedback, or goals user model
Timestamped occurrence or project milestone memory
A specific conversation worth recalling memory

Deleting everything for a persona

import Ecto.Query
from(m in Neoharness.Memory.MemoryEntry, where: m.persona == "trader")
|> Neoharness.Repo.delete_all()

Tool-driven equivalent: recall to find ids, then memory action delete one by one. (Intentionally no bulk-delete tool — an LLM mass-deleting memory without explicit user confirmation is how you lose months of state.)

Knowledge base

The third long-term store, next to memories and the user model: reference documents (specs, manuals, articles) ingested as chunks into knowledge_chunks and retrieved through the same recall tool.

knowledge_chunks
id              uuid pk
persona         text not null default 'main'
source          text not null              -- stable document name
chunk_index     int  not null              -- order within the document
content         text not null
embedding       vector(1536)               -- nullable; same model as memories
search_vector   tsvector (generated)       -- source weighted A, content B
  • Ingest — the knowledge_ingest tool (workspace file, PDF, or raw text + a source name). Documents split into ~3000-char paragraph-aligned chunks with a 200-char overlap; all chunks are embedded in one batched call. Re-ingesting a source replaces its chunks, so refreshed documents never duplicate.
  • Read — recall searches knowledge alongside memories and tags results [knowledge <source>#<chunk>]. Ranking is 0.4 × ts_rank + 0.6 × cosine — no importance or recency factors, since document chunks don't age the way journal memories do. Without an embeddings key, ingest stores text-only and search degrades to full-text, same contract as memories.
  • Not prompt-injected — unlike the top-memories block, knowledge is retrieved only on demand, so a large corpus costs nothing per turn.
  • Routing rule of thumb — chat-derived facts go to memory or user_model; documents worth consulting later go to knowledge.