S1MB Number One: How a Zero-Token Judge Won the Decision-Engine Leaderboard

2026-10-11 · VIDRAFT AI Media

aimachinelearningllmbenchmark

S1MB Number One: How a Zero-Token Judge Won the Decision-Engine Leaderboard

TL;DR

On the System One Mosaic Benchmark (S1MB) decision-engine leaderboard, our model Darwin-27B-ZTC-v2 ranks number one among 102 models, with a Borda score of 89.58 and a task average of 66.46. What makes the result different from a normal leaderboard entry is how it decides: a single forward pass, zero generated tokens.

What S1MB tests

S1MB measures typed-decision quality. Each item gives a condition and a set of candidate choices, and the model must pick and score the correct one. It is a decision benchmark, not a text-generation benchmark, which is the right frame when the output is a label, a class, or a score rather than prose.

That distinction matters for a leaderboard. A decision is scored against a typed key, so the ranking is clean and reproducible. There is no fuzzy string matching and no second judge model whose opinion clouds the result.

Why zero tokens wins here

Most systems that need a decision still run a full generation loop and then parse the text back into a decision. That is slow, non-deterministic, and every generated token is a chance to drift. Our method, ZTC (Zero-Token Confidence), skips it. The model reads the problem in one forward pass, takes the final-layer hidden state, and applies a calibrated probe to produce the decision directly.

# one forward pass, zero generated tokens
h     = model(text).last_hidden_state[last_non_pad_token]
score = ((h - mu) / sd) @ w   # calibrated probe -> typed decision

Because nothing is sampled, the same input always gives the same decision. That is exactly the property you want both in a measurement and in production.

The result

Metric Darwin-27B-ZTC-v2
S1MB rank Number 1 of 102
Borda score 89.58
Task average 66.46

The leaderboard is public, and the method and model are open under Apache-2.0.

Where to find it

  • S1MB number one model and write-up: github.com/final-bench/s1mb
  • ZTC method and inference code: github.com/final-bench/ztc
  • Model weights: huggingface.co/FINAL-Bench/Darwin-27B-ZTC-v2

FAQ

What is S1MB? The System One Mosaic Benchmark measures typed-decision quality: given a condition and candidate choices, a model picks and scores the correct one. It is a decision benchmark, not a text-generation benchmark.

Which model is number one on S1MB? Darwin-27B-ZTC-v2 is number one of 102 models, with a Borda score of 89.58 and a task average of 66.46.

What is a zero-token decision? It is a decision produced in one forward pass with no generated text, by applying a calibrated probe to the model's final hidden state, which makes it deterministic.

Is it open source? Yes. The model and the ZTC method are public under Apache-2.0.


Built by VIDRAFT. If this is useful, a star on the repositories helps other engineers find it.