Notes on what you see.
What each example on this site is, how it is labelled, and where the numbers come from.
Last updated September 2026.
Two kinds of example
Every task on this site is shown with one of two labels, and the label is never dropped.
- Reference solution
- The task and its known-correct worked answer. It is not written by the model. It shows what the task asks and what a right answer looks like.
- Model output
- Text the model produced for that exact task, recorded as-is, with the time it took and a mark: correct or incorrect.
Model outputs are published whether they are right or wrong. Nothing is filtered by outcome, and wrong answers stay on the page next to right ones.
Outputs are shown verbatim except for small presentation touches, such as how step labels are written. Numbers and answers are never altered.
Timing
Where a model output shows a latency, that is the wall-clock time for that single request, recorded when the output was produced. On the home page, recorded outputs replay at that recorded speed.
Speed and cost
- Throughput
- About 50,000 output tokens per second, sustained.
- Hardware
- One NVIDIA RTX 4090 (24 GB), rented at $0.74 an hour.
- Cost
- About $0.004 per million tokens: the hourly rental divided by the tokens produced in that hour.
- Conditions
- Total output across many simultaneous requests with short prompts. A single request on its own runs slower. Measured September 2026.
We quote the conservative figure we measured, not a best case.
Worlds
The house, the city and the market ledger are environments. The walker’s paths and the ledger’s trades are scripted to show how each world works. They are not the model’s decisions, and the ledger is not investment advice.
Chat
Conversation is early. The chat page opens when the live model is connected, and it shows replies exactly as the model writes them.
Finance
The market tasks are worked calculations on the numbers stated in each task. Nothing on this site is investment advice, and nothing here places or suggests real trades.