From exams to real work: an open 180B model now leads nine official Hugging Face leaderboards
aillmopensourcemachinelearning
Most leaderboard wins are about exams: math olympiads, graduate science, bar-style law questions. Two new results for Darwin-180B-RSI are about something more practical: whether a model can do the kind of work a company actually hands it.
With them, the model family now holds first place on nine of the 48 official benchmarks on Hugging Face, the most of any of the 95 organizations taking part. Moonshot AI and Z.ai follow with four each.
IFStruct: 98.95% on structured output
IFStruct, from Liquid AI, checks whether a model returns valid JSON or YAML that follows a requested schema, with the requirements phrased the many ways real users phrase them. Scoring is pass/fail per prompt and is strict:
- exact item count (or range), wrapper key, JSON vs. YAML
- code fence required or forbidden, no commentary when asked
- every field type, enum, numeric bound and nested list length
- any field the schema did not ask for fails the response
No constrained decoding is used, so only the model's own discipline is measured.
| Model | IFStruct |
|---|---|
| Darwin-180B-RSI | 98.95 |
| Agents-A1 | 93.25 |
| gpt-oss-20b | 91.95 |
| Nemotron-3-Nano-30B | 86.80 |
| Granite 4.1 8B | 68.45 |
Darwin passed 1,979 of 2,000 prompts with the official harness defaults (temperature 0, 16K tokens, thinking on). One prompt hit the harness's 120-second read timeout and was scored as a failure, so the number is conservative. With thinking off, the same model scores 89.7%: the reasoning pass is what catches the last few constraint violations.
Why it matters: in agent pipelines and back-office systems, a model's output is parsed by code, not read by a person. One wrong type or an invented key and the pipeline stops.
ExtractBench: 90.29 on document extraction
ExtractBench, from LlamaIndex, asks a model to pull structured fields out of 370 PDF documents, from short forms to files close to 200 pages. The second self-improvement round, Darwin-180B-RSI-R3, scored 90.29, ahead of its own base model Qwen3.8-Flash-Next (89.88) and Qwen3.8-27B (89.75). This run used the official harness with thinking disabled, which we state in the submission notes.
Nine first places
| Area | Benchmark | Score |
|---|---|---|
| Science and knowledge | GPQA Diamond | 94.44 |
| Science and knowledge | MMLU-Pro | 88.12 |
| Math | AIME 2026 | 100 |
| Math | HMMT Feb 2026 | 100 |
| Vision | MMMU-Pro | 79.48 |
| Law | LEXam | 68.94 |
| Law | LEXam-hard | 45.72 |
| Work | ExtractBench | 90.29 |
| Work | IFStruct | 98.95 |
MMLU-Pro (about 141 entries) and GPQA (about 111) are the two most crowded official boards, and both are led by this model. All numbers are self-reported through the standard .eval_results mechanism, with settings documented in each submission.
How it was built
Darwin-180B-RSI starts from Alibaba's open Qwen3.8-Flash-Next and improves through Model-level Recursive Self-Improvement: the model solves verifiable problems, keeps only its own solutions that check out as correct, and trains on them. No human-written solutions or reasoning traces are used. It was never trained on legal text or document extraction; the careful, step-by-step habits learned on verifiable problems appear to carry over.
- Model: https://huggingface.co/FINAL-Bench/Darwin-180B-RSI
- R3: https://huggingface.co/FINAL-Bench/Darwin-180B-RSI-R3
- All 48 official boards in one view: https://huggingface.co/spaces/quantid/huggingface-official-benchmark-leaderboards