Skip to content

Latest commit

 

History

History
executable file
·
413 lines (303 loc) · 21.2 KB

File metadata and controls

executable file
·
413 lines (303 loc) · 21.2 KB

Surreal-Memory Roadmap

Forward-looking vision. What's next, what's possible, where we're going. ZERO LLM dependency — pure algorithmic, regex, graph-based.

Current state: v3.11.0 — 57 MCP tools, 7900+ tests, neuroscience engine (10 brain-inspired algorithms), tiered memory loading (HOT/WARM/COLD).

Storage: SurrealDB is the only persistent backend — schema v12 — and storage_backend defaults to it. #141 removed the SQLite backend entirely; selecting storage_backend = "sqlite" is now a hard ValueError, not a fallback. InMemory remains, opt-in only, for trying the tool without running a database — it has no schema to version. Earlier revisions of this line described SQLite as still selectable with its own schema number; that stopped being true when #141 merged. Architecture: Spreading activation reflex engine, biological memory model, MCP standard. Community: Thanks to WebBrain for the project's first outside contribution — a full Spanish translation, README.es-ES.md (#127).


What We've Built (v1.0 → v4.11)

Capability Version Brain Test
Spreading activation (4 depth levels + RRF score fusion) v1.0–v2.29 Associative reflex
14 memory types, 24 synapse types v1.0 Typed memory
Hebbian learning + memory decay (type-aware) v1.0 Use it or lose it
Sleep consolidation (13 strategies: prune/merge/dream/mature/infer/...) v1.0 Sleep replay
Multi-format KB training (PDF/DOCX/PPTX/HTML/JSON/XLSX/CSV) v2.0 Learning from documents
Pinned KB memories (skip decay/prune/compress) v2.0 Core knowledge
Tool memory (PostToolUse → neuron clusters) v2.25 Procedural memory
Error resolution learning (RESOLVED_BY synapses) v2.0 Learning from mistakes
Multi-device sync (hub-spoke, 4 conflict strategies) v2.0 —
Fernet encryption + sensitive content auto-detect v2.0 —
VS Code extension (status bar, graph explorer, CodeLens) v2.10 —
REST API + WebSocket dashboard (7 pages) v2.0 —
Telegram backup integration v2.0 —
Brain versioning + transplant + merge v2.0 Portable consciousness
Algorithmic sufficiency gate (8-gate retrieval validator) v2.0 Attention filter
Codebase indexing + code-aware recall v2.0 —
SimHash deduplication + graph query expansion v2.29 —
Personalized PageRank activation (opt-in) v2.29 Hub dampening
RRF multi-retriever score fusion v2.29 —
Cognitive reasoning (hypothesize/evidence/predict/verify/gaps/schema) v2.27 Scientific reasoning
Source-Aware Memory (registry, exact recall, citations, audit) v3.1 Source memory
Structured encoding (tables, CSV, JSON arrays) v3.1 Structured recall
Cloud Sync Hub (Cloudflare Workers + D1, API key auth) v3.3 —
Session intelligence (topic EMA, auto-expiry, SQLite persist) v3.2 Working memory
Adaptive depth selection (calibration-driven, session-aware) v3.4 Efficient recall
Predictive priming (4-source: cache, topic, habit, co-activation) v3.5 Priming
Semantic drift detection (tag co-occurrence, Union-Find clustering) v4.0 Concept merging
Diminishing returns gate (stop traversal when no new signal) v4.11 Attention economy
Brain Quality Track A: Smart instructions, Knowledge Surface (.nm), reflection engine v4.8–v4.9 Proactive memory
Brain Quality Track B: Auto-consolidation, Hebbian retrieval, IDF keywords, adaptive decay v4.8 Graph quality
Lazy entity promotion (2+ mentions before neuron creation) v4.8 Selective encoding
Auto-importance scoring (heuristic priority from content signals) v4.8 Salience detection
Context merger (structured context dict in remember) v4.5 —
Quality scorer (per-memory quality hints) v4.5 —
Onboarding overhaul (smem init --full, smem doctor) v4.10 —
IDE rules generator (Cursor, Windsurf, Cline, Gemini, AGENTS.md) v4.6 —
Cascading retrieval with fiber summary tier v4.3 —
HuggingFace Spaces chatbot (ReflexPipeline, no LLM) v4.3 —
Config-driven cross-encoder reranking (HTTP or in-process, blended with SA score) v2.7.0 Attention refinement

Phase A: Production Hardening (v5.0)

Ship quality. Fix gaps. Make existing features bulletproof.

A1. SurrealDB Backend Parity

Problem: Some cognitive-layer features were implemented first for SQLite; the SurrealDB mixins need a parity pass before SurrealDB can be the default backend.

Scope:

  • Cognitive tables parity (cognitive_state, hot_index, knowledge_gaps) in SurrealDB
  • pin_fibers() implementation for SurrealDB storage — shipped with get_pinned_neuron_ids(), list_pinned_fibers(), graph density and document-training file tracking, all promoted to the NeuralStorage interface so a backend can no longer drop a capability silently
  • smem_edit type/priority changes persisted via SurrealDB
  • Parity test suite: run SQLite test matrix against SurrealDB

A2. File Watcher Ingestion — Issue #66 ✅

Problem: Users manually run smem train on files. Should be automatic: drop file → auto-memorize.

Scope: 3 phases (plan: .rune/plan-file-watcher.md)

  • Phase 1: Core FileWatcher class, watchdog integration, state tracking (mtime + simhash)
  • Phase 2: smem watch CLI + smem_watch MCP tool + config
  • Phase 3: smem serve integration, debounce (2s), metrics
  • Brain test: The brain absorbs information from its environment on its own → Yes

A3. Brain Quality Track C — Vertical Intelligence

Problem: Brain treats all content the same. Domain-specific entities (financial amounts, legal references) deserve specialized extraction and encoding.

Scope: 3 sub-phases (plan: .rune/plan-brain-quality.md)

  • C1+C2: Domain entity types + structured data encoding (regex-based, no LLM)
  • C3: Cross-encoder reranking (optional bge-reranker-v2-m3 post-SA refinement, HTTP or in-process) — shipped v2.7.0
  • C4: Agent visualization (smem_visualize → Vega-Lite/markdown/ASCII charts)
  • Brain test: Kế toán nhớ "ROE" khác "Paris" → Yes

A4. Stability & Polish

  • Pre-ship smoke test automation (scripts/pre_ship.py → CI)
  • E2E test coverage for dashboard (Playwright)
  • Schema migration rollback testing (v29 → v28 → v27)
  • Performance benchmarks: recall latency at 10K/50K/100K neurons
  • Consolidation performance fixes (v4.20.2-v4.20.4: timeouts, O(N²) caps, async yields)
  • InfinityDB integration fixes (7 bugs: singleton, list_brains, set_brain, WAL fallback, migrator)

A5. Neuroscience Engine ✅ (plan: .rune/plan-neuro-engine.md)

Problem: Brain metaphor stops at storage/retrieval. Real brains have lateral inhibition, reconsolidation, prediction error, context-dependent recall, and tiered access patterns. NM treats all memories equally at encoding and retrieval time.

Scope: 4 phases, 10 improvements (~1600 LOC total) — shipped v4.21.0

  • Phase 1: Lateral Inhibition + Temporal Binding + Emotional Valence (~250 LOC)
  • Phase 2: Prediction Error Encoding + Retrieval Reconsolidation (~350 LOC)
  • Phase 3: Context-Dependent Retrieval + Hippocampal Replay + Working Memory Chunking (~470 LOC)
  • Phase 4: Schema Assimilation + Interference Forgetting (~550 LOC)
  • v4.21.1: input firewall noise stripping, clean_for_prompt recall mode. Note: the multilingual (en/vi) extraction layer originally shipped here was later removed — extraction is now English-only; embedding-level semantic multilingual recall remains possible via the embedding model.
  • Brain test: ALL 10 improvements map to documented neuroscience principles → Yes
  • Zero LLM: Pure algorithmic (regex, SimHash, graph ops). No embeddings required.
  • 107 new tests, post-encode hooks, paginated tag fetch, real activation scores

A6. Tiered Memory Loading — Issue #111 ✅ (plan: .rune/plan-tiered-memory.md)

Problem: All memories have equal access priority. Real brains have fast-access working memory vs long-term storage. Safety rules and user preferences should always be available, not just when semantically matched.

Scope: Logical tiers on neurons (prerequisite for C1 physical storage tiers) — shipped v4.22.0, fixes v4.22.1

  • Schema migration v37: tier TEXT DEFAULT 'warm' on typed_memories + index
  • HOT tier: always injected into context, decay floor = 0.5, MAX_HOT_CONTEXT_MEMORIES = 50
  • WARM tier: default behavior (semantic match, normal decay)
  • COLD tier: explicit smem_recall only, 2× decay rate, excluded from auto-context
  • smem_remember(..., tier="hot") + smem_edit + smem_recall(tier=...) filter
  • smem_pin → auto-promote to HOT
  • Safety boundaries: type=boundary → always HOT, enforced in create/edit/pin/decay
  • Context optimizer: HOT +0.3 score boost, COLD excluded by default
  • Dashboard: TierDistribution card (progress bars), count_typed_memories() SQL COUNT
  • v4.22.1: 6 review fixes — with_priority data loss, boundary migration v38, case-insensitive tier, broader exception handling
  • 42 new tests across 4 phase files
  • Brain test: The brain has working memory (fast) vs long-term memory (slow) → Yes
  • Backward compatible: default warm → existing memories unchanged

Target: v5.0 = "production-ready for teams" release.


Phase B: Monetization & Growth (v5.x → v6.0)

From open-source tool to sustainable product. Revenue enables long-term development.

B1. Sync Hub: Landing + Payment (plan: .rune/plan-sync-hub-phase3.md)

Problem: Cloud sync works but has no billing. Need landing page + payment flow.

Scope:

  • Landing page (Cloudflare Pages) — features, pricing, signup
  • SePay integration (Vietnam, 0% fee) for domestic users
  • Stripe integration for global users
  • Free tier (100 neurons synced) → Pro tier ($5/mo, unlimited)
  • Usage dashboard: sync history, storage used, device count

B2. Sync Hub: Team Sharing (plan: .rune/plan-sync-hub-phase4.md)

Problem: Each agent has its own brain. Knowledge doesn't flow between team members.

Scope:

  • Team brain: shared namespace with per-user attribution
  • Roles: owner, editor, viewer
  • Activity feed: "Agent B learned about React hooks 2 hours ago"
  • Audit log: who changed what, when
  • Brain test: Collective memory (team knowledge) → Yes

B3. Distribution & Discoverability

Problem: Users don't know Surreal-Memory exists. Need to be where they search.

Scope:

  • MCP Registry listing (modelcontextprotocol.io)
  • awesome-mcp-servers PR (punkpeye/awesome-mcp-servers)
  • PyPI package optimization (description, classifiers, keywords)
  • npm package for OpenClaw plugin
  • Blog posts: "Surreal-Memory vs Mem0", "Why spreading activation beats RAG"
  • HuggingFace Spaces demo polished + promoted

B4. Brain Marketplace v1

Problem: Expert knowledge is siloed. A React expert's brain could help thousands of developers.

Scope:

  • smem brain publish --name "react-19-patterns" --tags react,hooks,rsc
  • smem brain install react-19-patterns --merge
  • Brain packages: versioned, with metadata (description, tags, size, neuron count)
  • Discovery: browse/search on sync hub landing page
  • Free tier: publish up to 3 brains. Premium: unlimited + featured listing
  • Brain test: Humans learn from books/teachers (external knowledge) → Yes

B5. Pro: Smart Tiers & Decision Intelligence — Issue #112

Problem: Free tier gives manual HOT/WARM/COLD control (#111). Pro makes it intelligent — auto-promote/demote by usage, structured decision matching, domain-scoped safety boundaries.

Scope: Requires #111 (A6) first

  • Auto-tier promotion/demotion: access patterns → auto WARM↔HOT, configurable thresholds
  • Decision component matching: structured components metadata on decisions, overlap scoring
  • Domain-filtered boundaries: boundary:financial, boundary:external, smem_boundaries(domain=...)
  • Tier analytics dashboard: distribution chart, decay curves, coverage gap detection, promotion history
  • Foundation: A5 hippocampal replay (LTP/LTD) already provides strengthen/weaken mechanism
  • Monetization gate: Intelligence, not data access — users always see all memories, Pro makes management smarter
  • Brain test: The brain self-adjusts priorities based on habit → Yes

Target: v6.0 = "Surreal-Memory as a service" with revenue stream.


Phase C: Scale & Enterprise (v6.x → v7.0)

From laptop brain to production brain. Handle millions of neurons.

C1. Tiered Storage Architecture

Problem: SQLite great for <500K neurons. Beyond that, graph queries slow down.

Vision: Hybrid storage — hot data in memory + SurrealDB, warm in SQLite WAL, cold in compressed archives. Prerequisite: A6 Tiered Memory Loading (logical tiers) + B5 Auto-tier (intelligent placement). C1 adds the physical storage layer underneath.

Hot tier (in-memory + optional SurrealDB)— recent + frequently activated
  ↕ auto-promote/demote
Warm tier (SQLite WAL)                   — moderate activity, queryable
  ↕ auto-archive
Cold tier (SQLite read-only, compressed) — archived, rarely accessed
  • Access frequency drives tier placement (already tracked in NeuronState)
  • KB (pinned) memories stay in hot tier permanently
  • Single query interface — storage layer handles tier routing transparently
  • Target: Sub-100ms recall at 1M+ neurons
  • Brain test: The brain has a fast-memory region (working memory) vs long-term memory → Yes

C2. Approximate Nearest Neighbor Index

Problem: SimHash dedup is O(n) scan. At 500K+ neurons, embedding-based recall bottlenecks.

Vision: ANN index (sqlite-vec or HNSW) for embedding pre-filtering, spreading activation refines within candidate set.

  • ANN narrows 500K → 500 candidates, SA refines final ranking
  • Index rebuilds async during consolidation (not on hot path)
  • Important: Acceleration, not replacement. Spreading activation remains central.
  • Brain test: The brain has a region that filters quickly before deep reflex (thalamus) → Yes

C3. Partitioned Brain Sharding

Problem: Single brain file grows unbounded. At GB scale, VACUUM takes minutes.

Vision: Auto-shard by domain_tag or time window. Each shard is independent SQLite file.

  • Domain shards: brain-kb-react.db, brain-kb-python.db, brain-organic-2026-Q1.db
  • Query router fans out to relevant shards only
  • Cross-shard synapses: (shard_id, neuron_id) tuple reference
  • Target: Individual shard stays <200MB, total brain can be 10GB+

C4. Self-Hosted Brain Hub (Production Docker)

Problem: Current sync is Cloudflare-hosted. Enterprises need self-hosted option.

Scope:

  • Docker one-liner: docker run -p 8080:8080 surrealmemory/hub
  • Admin dashboard: connected devices, sync status, brain health
  • Backup: automatic daily snapshots to configurable storage (S3/GCS/local)
  • Rate limiting + connection pooling
  • Deployment targets: Docker, Kubernetes, Railway, Fly.io

Target: v7.0 = "enterprise-ready" with million-neuron scale.


Phase D: Platform & Ecosystem (v7.0+)

From tool to platform. Surreal-Memory as the memory standard for AI.

D1. Brain Protocol Specification

Vision: Publish formal spec for how AI memory systems should work. Any vendor can implement it.

  • Core spec: neuron/synapse/fiber model, spreading activation algorithm, consolidation rules
  • Transport: MCP (primary), REST, gRPC
  • Serialization: brain export format (JSON + binary embeddings)
  • Compliance test suite: "Does your memory system pass Brain Protocol tests?"

D2. Plugin Architecture

Vision: Plugin hooks at every lifecycle stage. Community extends NM without forking.

Lifecycle hooks:
  on_encode    → custom extraction, enrichment, tagging
  on_recall    → custom ranking, filtering, augmentation
  on_consolidate → custom pruning, merging, summarization
  on_decay     → custom decay curves, preservation rules
  on_sync      → custom conflict resolution, transformation
  • Plugin registry: smem plugin install sentiment-boost
  • Sandboxed execution: plugins can't break core

D3. Multi-Modal Memory

Vision: Extend neuron types beyond text — images, code AST, audio.

  • Image neurons: store image embeddings, activate on visual similarity
  • Code neurons: AST-aware storage, activate on structural similarity
  • Audio neurons: voice memo → transcription + audio embedding
  • Cross-modal synapses: screenshot → error message → fix
  • Brain test: The brain stores memories multi-modally → Yes

D4. Federation Protocol

Vision: Brain Hubs peer with each other. Selective knowledge sharing across organizations.

  • Federation handshake: Hub A ↔ Hub B establish trust
  • Selective sync: share only neurons tagged with specific domains
  • Discovery: brain directory service (like DNS for brains)

Phase E: Intelligence Frontier (v8.0+)

Where Surreal-Memory goes beyond current AI memory paradigms.

E1. Dream Engine v2 (Insight Generation) — partially pulled to A5 Phase 3

Vision: During consolidation, detect patterns across unrelated memories → surface non-obvious connections.

  • Cross-domain pattern detection: "auth tokens expire" + "memory decay" → pattern
  • Anomaly detection: memories that should be connected but aren't
  • Weekly "dream report": "Your brain discovered 3 new connections this week"
  • Pulled forward: Hippocampal Replay (biased LTP/LTD) → A5 Phase 3
  • Brain test: Dreams create unexpected associations → Yes

E2. Forgetting Curves & Spaced Repetition — partially pulled to A5 Phase 4

Vision: Integrate Ebbinghaus curves into recall loop. Foundation exists (smem_review with Leitner boxes).

  • Auto-schedule review for important memories approaching decay threshold
  • Agent hints: "You haven't recalled 'deployment checklist' in 14 days. Review?"
  • Memories surviving multiple reviews → lower decay rate automatically
  • Pulled forward: Interference Forgetting (retroactive/proactive/fan effect) → A5 Phase 4
  • Brain test: The brain needs review to remember long-term → Yes

E3. Contextual Personality — partially pulled to A5 Phase 3

Vision: Brain adapts retrieval based on agent persona, task context, user preferences.

  • "Security expert" persona → boost security-related synapses
  • "Quick chat" context → shallow depth; "code review" → deep
  • Personality profiles stored as brain metadata
  • Pulled forward: Context-Dependent Retrieval (project/topic fingerprint) → A5 Phase 3
  • Brain test: Context influences how the brain remembers → Yes

E4. Causal Reasoning Engine — partially pulled to A5 Phase 2

Vision: Detect implicit causality from temporal patterns. "X always happens before Y" → auto-create causal synapse.

  • Temporal co-occurrence mining (existing sequence_mining foundation)
  • Confidence scoring (correlation ≠ causation guard)
  • Counterfactual queries: "What would have happened if X didn't occur?"
  • Causal graph visualization in dashboard
  • Pulled forward: Prediction Error Encoding (surprise signal) → A5 Phase 2
  • Brain test: The brain infers causality from experience → Yes

Stretch Goals (Exploratory)

Ideas worth tracking. May never ship, but inform direction.

Idea Brain Test Feasibility Impact
Voice interface — speak memories, hear recalls Yes (auditory) Medium High UX
Spatial memory — memories tied to locations/projects Yes (hippocampus) Medium Medium
Sleep mode — agent idle → deep consolidation Yes (sleep cycle) Easy High quality
Brain aging — long-lived brains develop "wisdom" Yes (wisdom) Hard High value
Memory palace — spatial organization of knowledge Yes (method of loci) Hard Novel
Neuroplasticity — brain structure adapts to usage Yes (plasticity) Medium High
Mirror neurons — learn by observing other agents Yes (mirror system) Hard Team AI

Guiding Principles

Every roadmap item must pass:

  1. Activation, not search — Does this make recall more like reflex, not query?
  2. Spreading activation stays central — Is graph traversal still the core mechanism?
  3. Works without embeddings — Would this work with pure graph + SimHash?
  4. Detailed query = faster recall — Does specificity still help?
  5. Brain test — Does a real brain do something analogous?
  6. Zero LLM dependency — Pure algorithmic. LLM is optional enhancement, never requirement.

Priority Signal

Phase Timeline Risk Value
Phase A: Production Hardening Now → v5.0 Low High — stability unlocks adoption
Phase B: Monetization & Growth v5.x → v6.0 Medium Critical — revenue sustains development
Phase C: Scale & Enterprise v6.x → v7.0 Medium High — unlocks enterprise use cases
Phase D: Platform & Ecosystem v7.0+ High Transformative — memory standard for AI
Phase E: Intelligence Frontier v8.0+ High Moonshot — novel AI memory paradigm

Last updated: 2026-03-28