Token ergonomics¶
Model tokens are the one running cost of this project. They are spent in two places:
- by the plane (the system that runs scans as small tasks) when a Run verifies candidates or compiles the decision engine;
- by the maintainers' own agentic coding sessions.
The maintainer's framing (2026-09-29): "the token saving strategy is a routing concern. ax can handle that and structure our services in small components." Here ax is Google's controller that runs agent tasks on Kubernetes. So the two halves look like this:
- Runtime is routing: one Model per Task with its own budget, and reuse of what was already paid for.
- Development is tooling that keeps bulk file contents out of the session.
Every number on this page comes from a committed record, named next to it. Records hold counts only. None holds code, prompts or responses.
The before-and-after summary is on Where we stand. It compares the script pipeline with the plane and gives the development share. This page holds the mechanisms and the detail table.
A. Runtime: the plane's token economy¶
flowchart LR
subgraph budget ["Budget and attribution"]
start["start request (credentials only here)"] --> task["Task: one Model, budget usd and calls"]
task --> meter["metered client refuses the next call at the ceiling"]
meter --> summary["summary.json: calls, tokens, usd"]
summary --> status["ousast plane status: attribution.json per task and per Model"]
end
subgraph reuse ["Pay once, reuse"]
facts["repo-facts, model-free"] -- "same repo, pin, candidates, image" --> seeded["seeded: calls 0, no usd"]
blobs["excerpts, embeddings, responses, programs by sha256"] -- "replay" --> zero["rerun at $0"]
end
subgraph order ["Cheap before expensive"]
triage["recorded triage per file"] --> hunt["one hunt per file, up to 6 candidates"]
hunt --> agree["agree: a and b"]
agree -- "disputed only" --> tiebreak["pass c, 2-of-3"]
end
One Model per Task, one budget per Task¶
A Run is a DAG (a graph of steps with no cycles) of ax Tasks. Each Task binds at most one Model.
Each Task has its own budget: {usd, calls} (The plane on ax).
The metered client enforces the budget (src/openultrasast/plane/budget.py, MeteredClient):
check()runs before every call.- It raises
BudgetExhaustedonce the spentusdor the call count reaches the ceiling. - So the call that crosses the ceiling completes, and the next one is refused.
- The task then ends
unfinished. A rerun with a larger budget resumes from itsunits.jsonl.
An HTTP 401/402 or "Insufficient Balance" is an AccountError. It fails the task and starts
nothing further.
Usage is attributed per call:
- The wrapped client appends one usage row per call (
prompt_tokens,prompt_cache_hit_tokens,completion_tokens). - Rows are absorbed even when the call raises.
- Rows are priced by the Model's price list.
- An unpriced Model cannot run under a
usdbudget.summary()reportsusd: Nonefor it, never 0. - Model-free tasks run with
{usd: 0, calls: 0}, so a stray call fails loudly.
ousast plane status RUN prints the attribution table. It reads each task's summary.json
and shows:
- one row per task with
calls,prompt,cache_hit,outputandusd; - a
subtotalrow per Model; - a
total. A missing value makes a sumn/a, never a smaller number.
--units adds the per-unit rows. The same table is written as attribution.json in the run
directory (src/openultrasast/plane/reconciler.py, attribution()/status()).
Provider credentials reach a task only in the start request:
router.pyreads the variable that the Model'ssecretKey.keynames.- The runner keeps the value in memory and passes it to the command's environment.
- The runner redacts it from echoed stderr.
- It never appears in a manifest, a rendered Task or a file.
Pay once, reuse¶
- Repo facts are computed model-free and reused.
repo-factsreads source text only (no model, no git).- Its
facts.jsonis byte-identical across runs over the same tree. - The memory store keeps a facts entry under the key
(repo, pin, candidates digest, runner image digest)(src/openultrasast/plane/memory.py,memory_key()). - A later Run's task may carry the same key. Then
seed()writes the storedfacts.jsonand adonesummary withcalls: 0and nousd. It marks the task done without starting an actor. - A changed pin, candidate set or image finds nothing and recomputes.
- Content-addressed memory (
src/openultrasast/learn/). Each item is stored under a hash of its content. - Excerpts are stored as
excerpts/<sha256 of the normalised text>.txt. - An embedding lives at
embeddings/<model>/<excerpt sha>.json. The source says: "an excerpt is embedded once; a rerun makes no call" (embeddings.py). - Model responses are cached under the sha256 of
{model, messages, params, temperature, sample}(program.py,request_key()). - A compiled program is
programs/<sha256 of its artifact>.json. It records the digests of the memory snapshot and of the response cache it was built from (compile.py). - Replay is free. With no client, the response cache replays at $0. A miss raises
ReplayMissinstead of calling. The injection-slice record states it: "evaluate re-run replay-only on the committed code: 550 responses replayed, $0, identical overall and calibration" (benchmarks/measurements/2026-10-01-decision-engine-injection-slice/record.json,replay_check).
Cheap before expensive¶
- Triage first. Verification runs only over the candidates the recorded triage kept.
- The plane applies that recording without a model call.
- It prices the recording as a separate line (
recorded_triage_usd). - So every per-candidate cost below is reported both with and without it
(
src/openultrasast/plane/tasks/agree.py). - Hunts are batched per file.
verifyruns one tool hunt per file for up toPER_HUNT = 6candidates. The prompt includes the candidates' known callers, capped at 8 (src/openultrasast/plane/tasks/verify.py,units_of()). - The tie-break pass runs only on disputes.
agreecompares passes a and b.- Pass c starts with
inputs.onlyset to the case'sdisputed.json. - An empty list ends it
donewith no model call. - The final verdict is 2-of-3 ("pass c alone never decides",
agree.py). - Cache-friendly prompt order. The decision engine's prompt has three parts, in this order
(
learn/program.py): - The compiled prefix: instruction and fixed demonstrations. It is identical for every candidate, so the provider's prefix cache hits.
- The retrieved examples.
- The candidate.
The self-consistency k (the number of samples per candidate) is tied to that cache: "k = 5
when samples 2..k hit the prefix cache for >= 0.8 of their prompt tokens". The measured
repeat-hit share was 0.96 (injection-slice record, smoke.k_rule,
smoke.k5_repeat_hit_share_measured).
The verify hunts share a prefix too. In the first plane increment, 2,385,152 of 3,820,395
prompt tokens were cache hits (benchmarks/independent/plane-increment-1.json,
detail.cache_hit_tokens, detail.prompt_tokens). The record's reading: "Cost fell 17% (62%
of prompt tokens were cache hits)".
Measured costs¶
| Measurement | Spend | Record |
|---|---|---|
| Plane increment 2: verify a, b and the tie-break on the 46-candidate validation set | $0.0205 per candidate over the 46 including the recorded triage, against the $0.022 reference (gate.cost_per_candidate_with_recorded_triage_usd); totals: plane $0.8567, recorded triage $0.0872, with triage $0.9439, reference $1.0123 (detail.usd); the tie-break asked 10 disputed candidates and flipped 5 to agreed |
benchmarks/independent/plane-increment-2.json |
| Plane increment 1: verify a and b alone | 357 calls, 54 hunts at 5.37 turns per hunt, $0.7471 plus $0.0872 recorded triage (detail) |
benchmarks/independent/plane-increment-1.json |
| Decision engine, injection slice: evaluation at k = 5 | $0.003215 per candidate (spend.cost_per_candidate_evaluate_usd); metered at list price: canary $0.010158, compile $0.429194, a first compile attempt that failed with HTTP 400 $0.049047, evaluate $0.353626, smoke $0.070812, in all $0.912838, plus $0.003976 of embeddings; 8,601,840 of 9,936,337 prompt tokens were cache hits (spend) |
benchmarks/measurements/2026-10-01-decision-engine-injection-slice/record.json |
| Decision engine, slice 2 (six families): compile, canary and evaluation per family | $2.8832 metered at list price for DeepSeek under an $8 ceiling (spend.metered_usd_deepseek_total), $0.000868 of embeddings; $0.0027 to $0.0032 per evaluated candidate (families.<family>.spend.cost_per_candidate_usd); 25,633,349 of 29,779,297 prompt tokens were cache hits (spend.deepseek_tokens); balance $17.73 to $16.73 |
benchmarks/measurements/2026-10-02-decision-engine-slice-2/record.json |
Decision engine harvest: verify part |
$2.5373 over 1,961 calls, 474 tasks done, 10,921,167 prompt tokens of which 6,275,060 cache hits, 306,941 output tokens, task budgets summing to $8.624 under a $10 ceiling (verify) |
benchmarks/measurements/2026-10-01-decision-engine-harvest/record.json |
Decision engine harvest: roles part |
$1.4036 over 1,034 calls, 151 tasks done, 2,398,669 prompt tokens of which 329,600 cache hits, 370,155 output tokens, task budgets summing to $4.917 under a $10 ceiling (roles) |
same record |
The increment-2 record notes that the two denominators differ: "the reference divides by the 46 candidates before triage; on that basis the plane is 7% cheaper, on the 43 after triage it is 0.2% under".
The meter prices at list; the account moved less
The attribution table prices every call at the Model's list price
(plane/models/deepseek-flash.yaml). The harvest record says: "Spend is the plane's
attribution table (ousast plane status, task meters priced at
plane/models/deepseek-flash.yaml); the DeepSeek account moved less (balance below)".
The balance was $19.63 before, $18.75 after verify and $18.23 after roles.
The injection-slice record names the ratio: "the provider charged ~1/3 of the list-price
meter (balance moved $0.30 for $0.91 metered): DeepSeek's off-peak discount window,
presumably; the meter prices at list" (spend.balance_note, balance $18.2 to $17.9).
Budgets and gates are set against the meter, never against the balance.
Not done yet¶
- No model-side token budget in ax itself.
- The plane's requirements state what ax does not yet carry: "task dependencies, artifacts, token budgets (on ax's roadmap), a DeepSeek provider".
- The thin
Runlayer carries them. It dissolves into ax as ax gains them. - Until then the ceiling lives in the metered client, not in the executor.
- Triage is applied, not run.
- The plane has no triage task.
- It filters by the recorded triage and adds that recording's cost as its own line.
- A triage task is planned as a later specification.
kand the contrast examples are not knobs.kis a compiled setting in {1, 3, 5} (learn/compile.toml, "k and lam remain A/B variants").- It changes only through the decision engine's A/B experiments. Requirement 4 of its specification asks that "changes to prompts, models, pass counts, rule sets and features [are] compared by controlled experiments, so that nothing is adopted on one run's number".
- Retrieval adds at most two contrast examples from other families. This cap is a module
constant (
learn/retrieve.py,CONTRAST = 2). - The cap is a maintainer decision, not an experiment variant yet.
B. Development: the maintainers' own sessions¶
flowchart LR
session["agentic coding session"] --> read["Read of a whole file"]
read --> guard{"read_guard: over 350 lines?"}
guard -- "no" --> allow["read proceeds"]
guard -- "yes" --> deny["denied: read a range or ask the bulk reader"]
deny --> bulk["bulk_read.py: cheap model answers with path:line under --max-usd"]
session --> worktree["subagent in a worktree: its tool output never enters the main context"]
session --> files["scripts in files, terse output"]
transcript["session transcript .jsonl"] --> report["token_report.py: share per content kind"]
report --> record["counts-only dev_token_report in the measurement record"]
The measurement: tool I/O dominates¶
benchmarks/dev/token_report.py reads a Claude Code session transcript (.jsonl). Without an
argument it reads the newest one under ~/.claude/projects/.
It reports characters and share per content kind:
tool_use_input(the tool call's input as JSON);tool_result;thinking;assistant:text;user:text.
It also reports the tool-call count and the top five tools. Tokens are approximated at four characters each. It prints to stdout and writes nothing. Only the counts-only summary is committed.
The committed summary is in benchmarks/independent/plane-increment-1.json
(dev_token_report). It covers "the whole development session (c09e30b2), not the increment
alone". It counts 7.3M unique characters (~1.81M tokens). The shares
are tool_use_input_pct 45.4, tool_result_pct 40.1, and tool_io_pct 85.5 against
previous_tool_io_pct 90, over 2,701 tool calls.
The reasoning the session is paid for (assistant text and thinking) is the small remainder. That is why the remedies below target file contents, not prose.
The remedies in benchmarks/dev/¶
read_guard.py is a PreToolUse hook (a script that runs before each tool call) for
agentic coding sessions.
- It reads the hook event on stdin.
- For the
Readtool only, it denies a whole-file read of a file longer than 350 lines (LIMIT = 350). - These pass: a read with
offsetorlimit, a missing path, an unreadable file, or a file within the limit. - It always exits 0 ("never fails the tool").
- A denial is one JSON line. Its
permissionDecisionReasontells the agent to read a range or to ask the bulk reader.
The script's docstring records how it is wired. It sits in .claude/settings.local.json (a
local, untracked file) as PreToolUse with matcher Read. The setting has this shape. The
command is repo-relative because hooks run from the project directory:
{
"hooks": {
"PreToolUse": [
{
"matcher": "Read",
"hooks": [
{"type": "command", "command": ".venv/bin/python benchmarks/dev/read_guard.py"}
]
}
]
}
}
bulk_read.py follows the portal pattern. A cheap model reads the files and answers one
question. It cites path:line for every claim. It reports what the files say and decides
nothing.
.venv/bin/python benchmarks/dev/bulk_read.py "<question>" FILE [FILE ...] [--max-usd 0.05]
How it works:
- Files go to the model in 400-line chunks, with a
--- path (lines a-b)header and numbered lines. - The client is the project's detector client:
deepseek-flashthroughDEEPSEEK_API_KEY, otherwise OpenRouter. - Each call has a 90 s timeout and no tools.
--max-usd(default 0.05) is checked before every chunk. At the ceiling the answer ends with[stopped at the $X ceiling before: <chunk header>].- The answer goes to stdout. The summary
[bulk_read: N chunk(s), $X.XXXX]goes to stderr. - Without a chat client it exits 2.
token_report.py: python benchmarks/dev/token_report.py [TRANSCRIPT.jsonl], described
above.
Practices¶
- Scripts into files, not heredocs. A heredoc (a script typed inline in a shell command) is
tool_use_input, and its output istool_result. Both are paid again at every later turn of the session. A script on disk is read once, by the interpreter. - Terse outputs. Commands print counts, exit codes and the lines that matter. A harness fails loudly with the exit code. It does not print a table that has to be read back.
- Subagents with worktrees for work whose output need not enter the main context.
- A subagent's reads and command output stay in its own context.
- Only its conclusion comes back.
- A git worktree (a separate checkout of the same repository) keeps its edits off the main checkout.
- Counts-only measurement records. A record under
benchmarks/measurements/orbenchmarks/independent/holds numbers and their keys. It never holds code, prompts or responses. So it can be quoted on a page like this one. A model can reread it at a few hundred tokens.
Related:
- The plane on ax for the Run and Task contract;
- Memory for the store the reuse keys live in;
- Evaluation for the gates these costs are measured against.