Scores that are proofs.
Every number on this board came from a real metered run through Sable's ordinary chat path. Each task minted its own signed receipt, the receipt ids are published, and the exact task list is published with its sha256 — so you can re-hash the declaration, pull every receipt, verify every signature, and run the same prompts yourself.
Read live from GET /v1/arena/leaderboard and GET /v1/arena/suites. Nothing here is entered by hand, backfilled, or weighted.
The gateway answered and returned no value for this field.
Printed exactly as the gateway returned it, so the ranking rule cannot drift from the ranking.
The gateway answered and returned no value for this field.
The gateway's own instructions, verbatim. Nothing on this page has to be trusted for them to work.
The gateway answered and returned no value for this field.
The trust model the gateway attaches to every Arena response, printed unedited.
- GET https://api.buildsable.com/v1/arena/leaderboard
- GET https://api.buildsable.com/v1/arena/suites
- GET https://api.buildsable.com/v1/arena/suites/{slug}
- GET https://api.buildsable.com/v1/arena/runs/{id}
A benchmark, not a verdict.
A row proves these exact declared tasks ran against this model at this time and scored this way. It says nothing about any task outside the published list, and the list is deliberately small.
There is no aggregate across suites, no weighting and no hidden tiebreak. Comparing a score from one task list against a score from another is the error this board exists to make visible.
A run that could not reach a model is recorded as failed, with its error class and no score. Scoring an unreachable model 0% would be a false claim about the model rather than a true one about the gateway.
Each task is an ordinary metered inference: held, capped, guardrailed and receipted. A score that was typed rather than measured has no receipts to show, and this board has nothing else to show.