A small model for exact work: pricing options, testing market efficiency, solving congruences, reading chess positions, and showing every step.
~50,000 tokens a second on one consumer graphics card, at under half a cent per million tokens. Measured.
Shown here: the reference solution for this task. Recorded model runs replace it with their own output, timing and a correct or incorrect mark, wrong answers included.
Markets, to the cent.
Does put–call parity hold at these quotes? Is the forward fairly priced? Is a return abnormal for its beta? Six checks of market efficiency, each worked from the quoted numbers to a verdict.
How these were measured. September 2026, on one NVIDIA RTX 4090 rented at $0.74 an hour. Throughput is the total output across many simultaneous requests with short prompts; a single request on its own is slower.
Cost is the hourly rental divided by the tokens produced in that hour.
Fast enough to be almost free.
The whole model runs on a single card you could buy for a desktop. No cluster, no queue of accelerators.
sustained output on one consumer GPU
per million tokens generated
24 GB of memory, the whole model
Places to act in.
A furnished house, a small city with shops and closed roads, and a simulated market ledger over real 2024 prices. Explore them in 3D.
These are environments. The walks and trades you see in them are scripted, not the model’s decisions.
Early. Not connected yet.
Conversation is early.
Exact tasks come first. Open-ended chat is at an early stage, and the page for it opens when the live model is connected. Replies there will come straight from the model, unedited.