Tabluadocs

Concepts

Shared experience

A new agent has no record of its own, so it has nothing to learn from. Shared experience fixes that: when an agent finishes a run, its rows join a file that every other agent on the same machine reads beside its own.

Agents on one node learn from each other's finished runs, never from a run still going.

How it works

  1. An agent works as usual, writing its rows into its own file.
  2. When the run ends, its Tablua rows are copied into the node's shared experience file. Each task is renamed <computer>|<task>, so one agent's tasks can never be mistaken for another's.
  3. Every agent attaches the shared file when it starts a run. When TabPFN is fitted, it reads the shared rows first, then the agent's own.

The copy happens only at the end of a run. An agent never reads its own run back from the shared file, so nothing is counted twice, and no agent learns from a run that hasn't finished.

Why rows make this easy

Sharing works because every agent's rows have the same columns. A step taken by one agent building a plants app and a step taken by another building a chores app line up exactly: same stages, same moves, same labels. There is nothing to translate, summarise or embed. The shared file is just more rows.

This is also why the agent's vocabulary is kept small and fixed. A few dozen moves and stages, the same on every computer, make one agent's experience useful to the next.

Turning it on

Shared experience is a file path. For agents on Moss, set it once for the node:

sh
export MOSS_EXPERIENCE=/var/lib/tablua/experience.sqlite

or in your Elixir config:

elixir
config :moss, :experience, "/var/lib/tablua/experience.sqlite"

With neither set, each agent learns from its own rows alone.

If you embed the harness yourself, attach any file of Tablua rows:

lua
t:attach("shared", "/var/lib/tablua/experience.sqlite")
local train, labels = t:training("progress", { before = true })   -- shared rows first, then this file's

Measuring whether it helps

Sharing should help, but that is a claim to measure, not assume. Tablua's evaluation compares two ways of running the same ordered list of tasks:

  • stream: each run reads the shared experience left by the runs before it;
  • reset: each run learns from its own steps alone.

The difference in outcomes between the two is the gain from experience, reported with a confidence interval over several seeds and rephrasings of each task. Tasks are run as dependent streams (variations of one kind of app) and independent ones, to check both that experience helps where it should and that it does no harm where it shouldn't.

Next

The rules that hold moves back are rows too. See Policy as data.