a16z Bet $870M on Typed, Machine-Native AI. Here Is the Open, Benchmark-Proven Version
aimachinelearningllmbenchmark
TL;DR
Andreessen Horowitz just led an 870 million dollar Series A into TypeSafe AI at a 7.5 billion dollar valuation, around a product the company calls Jev. The pitch is simple and, we think, correct: the next wave of AI is not chat, it is machine-native output. Models should return typed, structured decisions that software can consume directly, not paragraphs a human has to read and re-parse.
We agree with the thesis. We also shipped it, in the open, and measured it. Our typed-decision judge is ranked number one of 102 models on the System One Mosaic Benchmark (S1MB), it returns a decision in a single forward pass with zero generated tokens, and both the code and the weights are public.
What is typed, machine-native AI?
A normal chat model answers in free text. A program then has to parse that text, guess the structure, and hope the model did not drift. A typed model skips the prose. Given a condition and a set of choices, it returns the decision itself: a label, a class, a score. The output has a type, so the calling software can act on it without a fragile parsing layer.
This matters because most real systems do not want an essay. A router wants to pick one of eight backends. A moderation pipeline wants a verdict. A trading or pricing service wants a scored choice. In all of these, generation is overhead, and every generated token is a chance to hallucinate.
Why zero tokens changes the economics
If the output is a decision, you do not need a decoding loop at all. Our approach, called ZTC (Zero-Token Confidence), reads the problem in one forward pass, takes the final-layer hidden state, and applies a calibrated probe to produce the decision. No tokens are generated. That makes the result deterministic, cheap enough to run on every request, and free of the second-model hallucination you get when you use an LLM to judge another LLM.
# one forward pass, zero generated tokens
h = model(text).last_hidden_state[last_non_pad_token]
score = ((h - mu) / sd) @ w # calibrated probe -> typed decision
How well does it actually work?
On S1MB, the System One Mosaic Benchmark, our model Darwin-27B-ZTC-v2 is number one overall among 102 models, with a typed-decision accuracy of 0.743 zero-shot. S1MB measures fast, intuitive decision quality across a mosaic of typed tasks, which is exactly the capability a machine-native model needs to be useful.
| Rank | Benchmark | Models | Result |
|---|---|---|---|
| Number 1 | S1MB (typed decisions) | 102 | 0.743 accuracy, zero-shot |
The leaderboard is public, so you can check it yourself.
The open stack
The funding news validates the category. Here is the part you can use today, under Apache-2.0:
- ZTC family, the method and inference code, at github.com/final-bench/ztc
- Darwin-27B-ZTC-v2, the S1MB number one model, at github.com/final-bench/s1mb
- ONGRID, an ontology graph plus retrieval plus typed decision engine, at github.com/final-bench/ongrid
- POCKET, on-device CPU models for when the decision has to run locally, at github.com/final-bench/pocket
A large, well-funded company is now telling the market that typed, machine-native AI is the direction. We think that is right. The difference is that you do not have to wait for a closed product to try it.
FAQ
What is typed decision AI? It is an AI model that returns structured, typed output such as a label, class, or score that software can consume directly, instead of free-form text a program must parse.
What is ZTC (Zero-Token Confidence)? ZTC is a method that produces a decision in a single forward pass and generates zero tokens, by applying a calibrated probe to the model's final hidden state. It is deterministic and costs one forward pass per decision.
Which model is number one on S1MB? Darwin-27B-ZTC-v2 is number one overall on the System One Mosaic Benchmark among 102 models, with 0.743 typed-decision accuracy zero-shot.
Is there an open alternative to closed typed-AI products? Yes. The ZTC family, the S1MB number one model, the ONGRID engine, and the POCKET on-device models are all open source under Apache-2.0.
Why zero tokens instead of an LLM judge? An LLM judge generates a critique token by token, which is slow, non-deterministic, and can add its own hallucinations. A zero-token decision is deterministic and cheap enough to run on every request.
Built by VIDRAFT. If this is useful, a star on the repositories helps other engineers find it.