A guide for anyone wiring an agent host to TinyMemory's memory API, such as OpenHuman or another runtime. It covers what to call, when, and what each call costs.
The contract behind it is specs/agent-memory.md.
The design is architecture/lifecycle.md, and
how well it works is in evals/agent-memory.md.
- One object per agent (
AgentMemory) with four calls that fit the agent loop:start_session, when a session starts or resumes;pre_turn, before the model runs;post_turn, after the model replies;recall_for_compaction, when the prompt overflows.
- A shared brain (
Brain) of documents filed by source type (pdf, markdown, notion, github, web), visible to every agent. - Context packs. Every read returns one token-budgeted markdown block for the prompt, drawn from learnings, the brain, this agent's history and the team's conversations. The beliefs the engine built are merged into the learnings.
- Background work as values. Belief builds and deferred ingests come
back as
BackgroundJobvalues the host runs when it likes. The library never spawns a task. - Engine independence. Everything is written against
Arc<dyn MemoryEngine>. Moving from the reference engine to CortexDB (either wire), or to an engine of your own, changes one line.
Nothing is published to crates.io. Take the crates by git, pinned to a release tag:
[dependencies]
tinymemory-api = { git = "https://github.com/tinyhumansai/tinymemory", tag = "vX.Y.Z" }
tinymemory-tools = { git = "https://github.com/tinyhumansai/tinymemory", tag = "vX.Y.Z" }
tinymemory-integrations = { git = "https://github.com/tinyhumansai/tinymemory", tag = "vX.Y.Z", features = ["cortex", "brain", "documents-office"] }Feature (tinymemory-integrations) |
Gives you |
|---|---|
cortex (default) |
the CortexDB engine, both wires, plus registry and config |
brain |
brain_document: a file converted and filed under its source |
documents-office |
PDF, DOCX and XLSX conversion for the brain |
full |
everything |
For tests and offline development, enable tinymemory-api's conformance
feature as a dev-dependency to get ReferenceEngine, an in-memory engine
that needs nothing.
use std::sync::Arc;
use tinymemory_api::MemoryEngine;
use tinymemory_integrations::cortex::{CortexCredential, CortexEngine};
// A CortexDB server, direct.
let engine: Arc<dyn MemoryEngine> =
Arc::new(CortexEngine::direct("https://cortex.example.com", CortexCredential::api_key(key))?);Or build it from configuration, which keeps the engine choice out of code:
use tinymemory_integrations::{EngineCredential, MemoryConfig};
let config: MemoryConfig = serde_json::from_value(serde_json::json!({
"engine": "cortexdb", // or "tinyhumans"
"engines": { "cortexdb": { "endpoint": "https://cortex.example.com" } }
}))?;
let engine = config.build(EngineCredential::Static(key))?;A host that must attribute its requests (the TinyHumans backend expects an
x-sdk-name header naming the product) fixes them per engine, in
configuration ("engines": { "tinyhumans": { "headers": { "x-sdk-name": "my-product" } } }) or with CortexEngine::with_default_headers. They ride
every request; the transport refuses Authorization, Idempotency-Key, the
actor header and the headers the HTTP stack sets, so the credential stays
the EngineCredential's alone.
| Engine id | Where it runs | Consolidation (belief builds) |
|---|---|---|
cortexdb |
a CortexDB server, v1/* routes |
on demand: v1/beliefs/build, built within the request |
tinyhumans |
the hosted TinyHumans backend, memory/* routes |
on the server's own schedule |
ReferenceEngine |
in process (tests) | on demand, a deterministic toy |
engine.descriptor() tells you, at run time, which fetch modes an engine
serves and how it consolidates.
One MemoryLayout per tenant or workspace. Everything lives below its root:
use tinymemory_tools::MemoryLayout;
let layout = MemoryLayout::new("team:acme".parse()?)?; // or MemoryLayout::default()team:acme
├── source:pdf, source:notion, … the brain: documents, no agent id
├── agent:support-01 one agent's conversations (a turn per item)
├── agent:coder-42
└── (the root itself) shared learnings; beliefs are built in every scope
Two layouts with different roots never see each other's memory. The eval
checks this (isolation, 0 leaks). Use one root per tenant.
use tinymemory_tools::{AgentMemory, PostTurn, PreTurn, RecallPolicy, SessionStart};
let memory = AgentMemory::new(engine.clone(), layout.clone(), "support-01")?
.with_policy(RecallPolicy::default());
// Session start or resume: put the pack in the system prompt.
let initial = memory
.start_session(SessionStart { thread_id: Some(thread.clone()), focus: None })
.await?;
// Every user turn.
let mut pre = PreTurn::new(&thread, index, &user_text);
pre.in_prompt_from = first_turn_still_in_prompt; // turns from here on are left out
pre.at = Some(message_sent_at); // the message's own time
let context = memory.pre_turn(pre).await?; // never fails on an engine error
let reply = model.generate(&system, &context.pack.markdown, &user_text).await?;
let mut post = PostTurn::new(&thread, index + 1, &reply.text);
post.tool_calls = reply.tool_calls; // name and id only, see the rules
post.at = Some(reply_at);
let report = memory.post_turn(post).await?;
queue.extend(report.jobs); // belief builds, run off the turn| Call | Writes | Reads | Typical cost |
|---|---|---|---|
start_session |
— | thread turns (if resuming), then the standard sections, newest first or ranked for focus |
about 8 ms |
pre_turn |
the user turn, accepted but not yet indexed, concurrently with the read | the standard sections, ranked for the turn | 13–30 ms with local embeddings; one embedding round trip per scope with hosted ones |
post_turn |
the reply, accepted | — | about 2 ms |
recall_for_compaction |
— | an answered summary of the thread (a model runs), then the standard sections | seconds; off the hot path |
recall(query) |
— | the pre_turn read without logging |
as pre_turn |
The standard sections of a pack, highest priority first, are: Learnings (stored learnings and the engine's beliefs), Brain, This agent's history, and Team conversations. An item appears once, in its first section. When the pack is over budget, trimming starts from the end.
When the host drops turns from the prompt, pass them in:
let carried = memory
.recall_for_compaction(Compaction { thread_id: thread.clone(), dropped, focus: None })
.await?;The summary section answers from the thread's turns. It reads as many as were dropped, up to 24. The standard sections follow, ranked for a sample of every dropped turn.
use tinymemory_tools::{Brain, BrainDocument, BrainSource};
use tinymemory_integrations::brain::brain_document;
let brain = Brain::new(engine.clone(), layout.clone());
// Text you already have:
let ingested = brain
.ingest(BrainDocument::new(BrainSource::Notion, text).titled("Refund policy"))
.await?;
queue.push(ingested.job); // build this source's beliefs later
// A file: converted, its source picked from the format (PDF → pdf, md → markdown).
let document = brain_document(&converters, &raw, None, MemoryMeta::default()).await?;
brain.ingest(document).await?;ingest waits until the document is readable, which takes about 70 ms with
local embeddings and 2 s with hosted ones. Use ingest_with(doc, WaitFor::Accepted), or queue a BackgroundJob::IngestBrain, to skip the
wait. search and forget work per source or across the whole brain.
BackgroundJob is plain, serializable data: queue it, persist it, merge
duplicates, and run it when convenient:
let runner = memory.background();
for job in queue.drain(..) {
let report = runner.run(job).await?; // Done, Started, Scheduled or Skipped
}post_turnhands back aBuildBeliefsjob everyRecallPolicy::build_beliefs_everyturns (10 by default).Brain::ingesthands one back per ingest.- On CortexDB, a build is only as good as the extraction before it. Beliefs are built from facts, which CortexDB extracts from every event with a model, in the background. With a hosted model that takes minutes. A build run straight after a write builds from nothing, so schedule builds a few minutes behind the writes. A periodic sweep works well.
- An engine that cannot consolidate answers
Skipped, so the same host code runs on any engine.
These come from the eval; each one cost accuracy when ignored.
- Pass the message's own time as
PreTurn::atandPostTurn::at. Packs show a turn's date ([2026-09-15 09:01] user: …). Without dates, a corrected value and its correction look alike. With them, a model picked the current value in 6 of 6 contradiction probes. Use the message time, not the call time, so a retried turn is recognised as a duplicate. - Put tool results in the reply text.
ToolCallRefkeeps a tool's name and id. A result the assistant does not mention is lost, and a long reply dilutes what it does mention. A short "tool → result" line per call is what the eval's agent does. - Set
in_prompt_fromto the first turn of the thread still in the prompt. Turns from there on are left out of the pack, so the model never sees the same turn twice. - Keep thread ids and turn indices stable. A turn's identity is its content, thread and index. A retried call with the same input is recognised as a duplicate and not stored twice.
- Never block a turn on memory.
pre_turnreturns its pack even when logging fails; the reason is inTurnContext::log_error. Log it, and carry on with the turn. - Run builds off the turn. A build takes seconds per scope with a real model.
- Use one root per tenant. Agents under one root share learnings, the brain and each other's conversations by design.
RecallPolicy sets the shape of every pack:
| Field | Default | Meaning |
|---|---|---|
budget_tokens |
1200 | the pack's size, at about 4 characters per token |
learnings_limit |
8 | learnings and beliefs together |
brain_limit |
6 | brain documents |
history_limit |
6 | this agent's turns |
team_limit |
3 | other agents' turns; 0 leaves the section out |
build_beliefs_every |
Some(10) |
turns between belief builds; None turns them off |
For a different shape altogether, build a HolisticRecall with your own
ScopeSections and call holistic_recall. context.md
(tinymemory_tools::context) is one preset of it, with briefs answered by
the engine's model. It suits the top of a long-lived system prompt.
agent.tools() gives MemoryTools scoped to that agent. It has seven
tools with JSON Schemas (specs()), called by name (call(name, args)).
Writes land at the agent's node; reads see that node and the shared root.
The model never chooses whose memory it touches.
Implement tinymemory_api::MemoryEngine.
Required:
descriptorhealthrecallfetchstoreforgetlist
Optional, with defaults:
| Method | Default | Override when |
|---|---|---|
store_with |
behaves as store |
the engine can acknowledge a write before indexing it |
consolidate |
Unsupported |
the engine builds beliefs (declare it in descriptor().consolidation) |
beliefs |
none | the engine keeps beliefs apart from its stored items |
fetch beliefs (FetchRequest::beliefs) |
none | as above, served from the same read as the fetch |
get, explore, store_many |
built from list and store |
the engine can do them more cheaply |
Then run the conformance suite, the same one CortexDB passes:
#[tokio::test]
async fn my_engine_upholds_the_contract() {
tinymemory_api::conformance::run(&MyEngine::new()).await.unwrap();
}It checks every promise the lifecycle relies on:
- an accepted store answers with the item's own id;
- consolidation matches what the descriptor declares;
- beliefs are learnings within the reach asked for;
- reach never leaks across siblings.
- Unit tests: run against
ReferenceEngine, with no network. - Against a real CortexDB:
integration/cortexdb/runs a pinned server with deterministic mock models../scripts/cortexdb-live.shruns the live suite against it. - Measuring accuracy and latency:
./scripts/memory-eval.shruns 12 scenarios through a scripted agent, among them restarts, contradictions, tool-heavy runs, compaction, isolation, learnings, surprises and conflicts.MODELS=openrouteruses real models; seeevals/README.mdfor cost and model requirements.
- The engine is built from config, and the key is not in code.
- There is one
MemoryLayoutroot per tenant. - There is one
AgentMemoryper agent id. -
start_sessionfeeds the system prompt. -
pre_turnruns before every model call, withatandin_prompt_fromset. -
post_turnruns after every reply, with tool results in the text. -
log_erroris logged, never raised. - Jobs are queued and run off the turn, a few minutes behind the writes.
-
recall_for_compactionis wired to prompt truncation. - The brain is ingested per source.
- The integration passes on
ReferenceEnginein CI and on the live harness before release.