
( this is though all console-based )
A small Python toolbox for running small agents on local hardware. And two very different agents built with it.
(could be used with larger remote llm's as well)
The idea behind all of it: small models are good at judgment in the moment and bad at holding things in mind. So Python as a layer in between the user and LLM. It holds the structure (the files, the memory, the loop, the rules) While the model only makes the next decision. Python also talks back to the model: when a tool call is wrong, it nudges the model what went wrong and what was probably meant, which keeps a small model on track instead of derailing. The LLM's tool calling is orchestrated by Python. Each tool has its own help, and there is even a document about how to solve problems (for Lisa).
Written for models you can run at home (about 12 GB of VRAM or less), through LM Studio or any OpenAI-compatible server. No agent framework underneath: plain Python, readable start to finish.
Status: experimental, but both sides run end to end.
-
Lisa runs continuously. Talk to her, and she answers; leave her alone, and she thinks about what is on her mind, looks things up, and after a while she gets tired and sleeps. While she sleeps, she dreams about her memories, merges the ones that say the same thing, and sometimes wakes up wondering about something new. Her whole mind is a folder of markdown files you can open in any text editor. You can see her looking up Wikipedia (tool use is yellow) Or see her dreaming in green, or see her thinking throughout her days. She has a novel idea of dreaming; it isn't mainly cleanup; it drives her thoughts.
This is where most of my interest is these days. Read more about Lisa.
It's is more akin to a research project for me. -
Where this project started. Instead of asking a small model to hold a whole file in its head and write a diff, it edits code by symbol through Python tools: read one function, rewrite it, and Python checks the syntax before anything is saved. An architect agent plans work as todos, a coder picks them up, and a room runs the two in turn. Everything is folder based sandboxed per agent. (though its not a virtual environment, just basic safety here). While an agent can take multiple turns to solve something. More complex than the Lisa agent, though more targeted towards work.
The coder is the most advanced agent for more info: Read more about the coder.
pip install requests
pip install ddgs (optional: web search for Lisa)
Start LM Studio with a model loaded, then:
python lisa_agent.py talk to Lisa
python coder_agent.py give the coder something to build
Pointing it at another server or model needs no code changes:
set BABYCODER_URL=http://localhost:1234/v1/chat/completions
set BABYCODER_MODEL=your-model-name
(export instead of set on Linux and macOS.) It works with any
tool-calling model; reasoning models are handled too (see the coder page).
| File | What it is |
|---|---|
babycoder/ |
The toolkit: every tool, the agent loop, sandboxing, memory, the model layer. |
lisa_agent.py |
Lisa. |
coder_agent.py, architect_agent.py, agent_room.py |
The coding agents, and the room that runs them together. |
designer_agent.py |
A small example: an agent made of nothing but a role and a set of granted tools. |
emotion_agent.py |
An older experiment: a character whose emotions are kept in a tool schema. |
LISA_DESIGN.md |
The design Lisa is written against. |
tests/ |
Offline checks, with a stand-in for the model server. No LM Studio needed. |
Run the checks with:
python tests/check_babycoder.py
This is a personal research project and I can still make big design choices. If you are working on something similar, persistent agents, memory that consolidates while idle, or coding with small models, I would like to hear from you. Open an issue or a discussion.
MIT licensed.
Two commands write straight into her stores, decorated so that what she reads back is identical to something she produced herself:
/goal Why do gulls follow tractors across a field?
/mem I watched a heron at the Vooroever this morning #thinking
You type plain text; the number, timestamp and tags are generated the way her
own entries are. She has no way to tell. The one field that records where it
came from, origin: injected in goals.md, never reaches the model - it is
there for you when you read a transcript back.
An injected goal is backdated so she takes it on her next cycle rather than
last. next_goal() picks the goal left alone longest, so without that a fresh
injection would be the last one she would choose - with a full mind, seven
cycles away, which is no use for breaking a loop you are watching now.
A goal her own filters would refuse (anything about whether her past is real) is still accepted, because you are the operator, but the console says it may loop.
This is a development instrument, not part of who she is. Watching a loop does
not tell you what would break it. /injectgoal and /injectmem work too.