Concepts
Shared experience
A new agent has no record of its own, so it has nothing to learn from. Shared experience fixes that: when an agent finishes a run, its rows join a file that every other agent on the same machine reads beside its own.
How it works
- An agent works as usual, writing its rows into its own file.
- When the run ends, its Tablua rows are copied into the node's shared experience file. Each task is renamed
<computer>|<task>, so one agent's tasks can never be mistaken for another's. - Every agent attaches the shared file when it starts a run. When TabPFN is fitted, it reads the shared rows first, then the agent's own.
The copy happens only at the end of a run. An agent never reads its own run back from the shared file, so nothing is counted twice, and no agent learns from a run that hasn't finished.
Why rows make this easy
Sharing works because every agent's rows have the same columns. A step taken by one agent building a plants app and a step taken by another building a chores app line up exactly: same stages, same moves, same labels. There is nothing to translate, summarise or embed. The shared file is just more rows.
This is also why the agent's vocabulary is kept small and fixed. A few dozen moves and stages, the same on every computer, make one agent's experience useful to the next.
Turning it on
Shared experience is a file path. For agents on Moss, set it once for the node:
export MOSS_EXPERIENCE=/var/lib/tablua/experience.sqliteor in your Elixir config:
config :moss, :experience, "/var/lib/tablua/experience.sqlite"With neither set, each agent learns from its own rows alone.
If you embed the harness yourself, attach any file of Tablua rows:
t:attach("shared", "/var/lib/tablua/experience.sqlite")
local train, labels = t:training("progress", { before = true }) -- shared rows first, then this file'sMeasuring whether it helps
Sharing should help, but that is a claim to measure, not assume. Tablua's evaluation compares two ways of running the same ordered list of tasks:
- stream: each run reads the shared experience left by the runs before it;
- reset: each run learns from its own steps alone.
The difference in outcomes between the two is the gain from experience, reported with a confidence interval over several seeds and rephrasings of each task. Tasks are run as dependent streams (variations of one kind of app) and independent ones, to check both that experience helps where it should and that it does no harm where it shouldn't.
Next
The rules that hold moves back are rows too. See Policy as data.