Tabluadocs

Concepts

The step loop

A Tablua agent works one step at a time. Each step has the same shape: read where the work stands, list the moves allowed, choose one, do it, and check how it went. Every part of that is written down as a row.

  1. stateWhere the work standswritten by the host from facts
  2. candidateEvery move it could makewith each model's number for it
  3. decisionThe move it tookand who chose it
  4. actionWhat the move didcalls Mercury filled in
  5. outcomeHow it turned outchecked by the host, not a model

The outcome is the next step's starting point, and a labelled row the tabular model learns from.

One step of the loop. Each box is a row in a table of the agent's file.

Code owns the workflow, models make the choices

Tablua draws a clear line between what is decided by code and what is decided by a model.

Code decides where the work stands. The host reads facts it can check: which feature files exist, whether the person has agreed to them, how many test scenarios pass, which pages answer, whether the app was published. From those facts it works out the stage, such as building or ready. No model is asked.

Code decides which moves are allowed. Each stage has a fixed list of moves. In no_feature, the agent can write a feature, read the help, think, or say it is blocked. In ready, it can publish. A model can never pick a move that is not on the list.

Models choose among the allowed moves. Jev picks one, with a probability for each option. TabPFN, when it is on, adds its estimate of each move's chance of making progress. Mercury then fills in what the chosen move needs.

This split keeps the agent's progress measurable. A model can be wrong about what to do next, but it can't be wrong about whether the tests pass, because it is never asked.

The stages

StageWhat it meansExample moves allowed
no_featureNothing written yetwrite_feature, read_help, think
awaiting_agreementA feature is written; the person hasn't agreedwait_for_agreement, write_feature
buildingScenarios are agreed and some failwrite_steps, write_code, write_page, run_test, fix_failure, undo, rewrite
readyEvery scenario passes and every page answerspublish, look_at_app, fix_failure
awaiting_yesPublishing waits for the person's yeswait_for_yes
shippedThe app is publishedanswer_task
answeredThe task is doneanswer
changingThe task asks to change an app that already shippedwrite_feature, write_code, write_page

The full list is in Stages and moves.

The rows of one step

RowWritten byWhenWhat it holds
tablua_statehostbefore decidingstage, scenarios passed and total, how long it has stalled, the last move and its outcome, the suspected cause of a failure
tablua_candidateJev and TabPFNwhile decidingone row per allowed move: Jev's probability, TabPFN's chance of progress
tablua_decisionthe harnesswhen decidedthe move taken, who chose it (jev or tabpfn), how likely the choice was
tablua_actionthe harnesswhile actingeach call the move made: the command, the file and its kind, bytes written, exit code
tablua_outcomehostafter actinghow it turned out: complete, broken or no_effect; whether more scenarios pass; whether anything regressed
tablua_effecthostafter actingwhat changed, as keywords such as More Passing, Regressed or Page Fixed

At the end of a run, one tablua_run row records whether the app shipped and works, how many steps it took, and what it cost.

Progress, worked out from the rows

Each outcome gets a progress label of 1 or 0, from a simple rule:

  1. If more scenarios pass than before the step, it made progress.
  2. Otherwise, if the step didn't complete, it didn't.
  3. If it was a change to the app, made while scenarios were failing, and no more pass afterwards, it didn't.
  4. Any other completed step did.

That rule turns every step the agent takes into a labelled example, at no cost. It is short-sighted on purpose, and Tablua adds a longer view after each run. See How Tablua learns.

One step per call

The host doesn't keep the agent in memory between steps. Each call reads what the last one saved, takes one step and saves again. This is why the agent can run on a computer that sleeps between steps, and why the whole agent fits in one file you can move.

The person's part, agreeing to a feature or saying yes to a publish, is never done by the agent. The run waits for it.

Next

Three models, one table explains who writes which columns, and why.