Start here
A tour of one step
This page follows an agent building a small app, step by step, and shows each row it writes. You don't need to run anything. The values are illustrative, but the shape of every row is exactly Tablua's.
The task
A person asks: "I'd like a little app for my house plants. It should list each plant with when I last watered it, let me add a plant, and let me mark one as watered today."
The agent builds apps the way a careful developer does. First it writes the person's ask as test scenarios, in Gherkin. Once the person agrees to them, it writes the code, the test steps and the pages until every scenario passes, then asks the person before publishing.
Step 1: write the feature
Where the work stands. Before anything is decided, the host writes a state row from facts it can check: there is no feature file yet.
state n=1 stage=no_feature passed=0/0 stalls=0What could be done. In the no_feature stage only a few moves are allowed. Jev is asked to pick one and gives a probability for each. Each option becomes a candidate row.
candidate write_feature jev_p=.82What was done, and by whom.
decision chosen=write_feature by=jev
action write features/plants.feature (4 scenarios) ← written by Mercury
outcome complete progress=1 passed=0/4 ← checked by the hostThe feature now exists, with four scenarios that all fail, because there is no code yet. That is expected. The step still counts as progress: it did its job.
Step 2: write the test steps
The person agrees to the feature. Now the stage is building, and many more moves are allowed. This time TabPFN has past rows to learn from, so each candidate also gets p_progress: TabPFN's estimate of the chance that this move makes progress here.
state n=2 stage=building passed=0/4 last=write_feature
candidate write_steps jev_p=.61 p_progress=.52
candidate write_page jev_p=.27 p_progress=.31
candidate rewrite jev_p=.12 p_progress=.07
decision chosen=write_steps by=jev propensity=1
action write code/steps/plants.lua (38 lines)
outcome complete progress=1 passed=2/4Two of four scenarios pass. Jev and TabPFN agreed on the best move, and the outcome bears them out.
Step 3: a step that makes things worse
state n=3 stage=building passed=2/4 last=write_steps
candidate rewrite jev_p=.55 p_progress=.09
decision chosen=rewrite by=jev
action rewrite code/plants.lua (52 lines)
outcome broken progress=0 passed=1/4 regressed=1Jev was fairly sure about rewrite. TabPFN, having seen how rewrites in this stage tend to go, gave it a 9% chance. The rewrite broke a scenario that was passing. The outcome row records it plainly: regressed=1.
This is the kind of disagreement the table makes visible. Over many runs, rows like this show whether Jev's confidence or TabPFN's estimate deserves more trust in which situations. Tablua can be set to act on that. See Turn learning on.
Step 4: undo
After a regression, the undo move becomes available. It puts back the files the last change wrote.
state n=4 stage=building passed=1/4 last=rewrite
decision chosen=undo by=jev
outcome complete progress=1 passed=2/4Steps 5 and 6: finish and publish
state n=5 stage=building passed=2/4 last=undo
candidate write_page jev_p=.70 p_progress=.66
decision chosen=write_page by=jev
action write ui/index.org (61 lines)
outcome complete progress=1 passed=4/4
state n=6 stage=ready passed=4/4 pages_ok=1
decision chosen=publish by=jev
run shipped=1 works=1 steps=6All four scenarios pass, so the stage becomes ready and publish is allowed. Publishing waits for the person's yes. The run row records how the whole run ended.
What to notice
- Facts come from the host, choices from the models. The stage, the pass counts and the regression are measured. No model is asked whether its own work succeeded.
- Every option is recorded, not just the one taken. That is what lets a model later learn what would likely have happened with the others.
- Who decided is recorded.
by=jevorby=tabpfn. When you learn from your own decisions, knowing who made them, and how likely the choice was, keeps you honest about cause and effect. - Every step is labelled for free.
progressis worked out from the rows themselves, so the agent's own work becomes training data without anyone writing labels.
Next
Try it yourself in the Quickstart, or read The step loop for the full picture.