Both gates are measured, not modelled.
Measured on 2026-09-22. Raw per-item output and per-pack timings are in the repository; every number on this page is derived from those files by a script at build time, not typed in by hand; the notes around them quote the testnet record. Total spend across both runs: $1.17.
What was not measured. The scope of the measurement: the us-west-2 numbers are 4 packs per configuration from a shared serverless function, and the two testnet seasons were 16 answers from two rigs in Tokyo.
6 packs per configuration
The clock.
| Configuration | p50 | p95 | on time at 400 ms | cost / pack |
|---|---|---|---|---|
| Jev, 2-way split | 318ms | 344ms | 100% | $0.00132 |
| Jev, one request | 465ms | 487ms | 0% | $0.00130 |
| Jev, 4-way split | 299ms | 494ms | 83.3% | $0.00138 |
| Contender B, 8-way | 880ms | 1061ms | 0% | $0.00656 |
| Contender A, 8-way | 1538ms | 1958ms | 0% | $0.00331 |
Two things in that table matter more than the winner. First, the split size: two ways holds the tail, four ways has a lower median and a p95 of 494 ms, because a pack finishes when its slowest branch does. Second, the contenders were run at the split size they can actually complete — one of them cannot emit 128 labels in a single reply at all, so it must pay for the policy card eight times over.
That table was taken from us-east-2 through a proxy: the long way round. The one below was taken from us-west-2, the region the rig template deploys to and the region the model is served from, with 4 packs per row.
| Configuration, us-west-2 | p50 | slowest | within 350 ms | within 400 ms |
|---|---|---|---|---|
| Jev, 2-way split | 227ms | 443ms | 3 of 4 | 3 of 4 |
| Jev, one request | 321ms | 360ms | 3 of 4 | 4 of 4 |
| Contender A, 8-way | 1375ms | 1386ms | 0 of 4 | 0 of 4 |
| Contender B, 8-way | 1804ms | 2004ms | 0 of 4 | 0 of 4 |
Network round trip from us-west-2: the model's API 18 ms, one contender's host 81 ms, another's host 44 ms. The rivals start further away before they have answered anything.
3840 candidates, $0.92
The hard items.
| System | on the kept pool | on separators | error rate on the hard band |
|---|---|---|---|
| Jev | 100.0% | 0 wrong of 322 | 0 of 90 |
| Contender B | 94.1% | 53% | 39.3% |
| Contender C | 92.6% | 42.8% | 29.2% |
| Contender A | 96.4% | 69.9% | 77.8% |
Jev's grade came from a third, independent run — not from the runs used to screen the candidates in the first place, which would have let it mark its own homework. On the kept pool it made 0 errors in 2680 items; with a sample that size, all that rules out is an error rate above 0.1%.
40,000 simulated packs
Carried through to whole packs.
A pack is 80 anchors and 48 hard items, and it scores only if whole-pack errors stay within 4 points and hard-item errors within 3. Drawing items with replacement from the measured per-item results gives the share of packs each system would have passed:
| System | hard items drawn from all separators | drawn from the hard band only |
|---|---|---|
| Jev | 100% | 100% |
| Contender B | 0% | 0% |
| Contender C | 0% | 0% |
| Contender A | 0% | 0% |
Contenders are simulated as if their connection never failed: each is resampled only from the items it actually answered, so a timeout is never scored as a wrong answer. Counted that way, every contender goes to zero in this 80 + 48 shape whether the hard items are drawn from all separators or from the hard band alone. The shipped season is a different shape: 128 screened items against an error budget of two, with no hard-item band. Resampled the same way, Contender A passes about 15% of those packs, B 1.7% and C 0.3%.
the uncomfortable part
Four things worth saying out loud.
The screening is what carries this
Jev is not naturally good at these items. Before screening it got a large fraction of the hardest candidates wrong. What makes the pool safe for an honest miner is the screening step every candidate passes before it can appear in a pack. If that step breaks, the first people it hurts are honest miners — not attackers.
The accuracy gate rests on 90 items
Not 322, and not 2680. 90. That is the thinnest number in this project and the next measurement to run.
We have corrected ourselves twice
The first version of this benchmark had four defects, all in our favour: only Jev was given the per-option criteria, the questions did not depend on rule order, the contenders ran at a split size that is not their best, and a rate-limited contender was scored as wrong. After that was fixed, a second error remained: a contender that returned nothing was still counted as having answered wrongly, which more than doubled the separator count. Every number on this page is now computed without either mistake.
Hard items are scarce
A template generator yields about one hard item in eight kept items, and in live factory runs it could not fill a band of 48. So the factory has a model rewrite every ticket, and the shipped season scores whole packs of 128 screened items against an error budget of two, with no separate hard-item band.
in full
What was not measured.
- The latency run from the rig region
- us-west-2 was measured with 4 packs per configuration from a serverless function, which shares CPU and can cold-start. The testnet seasons added 16 answers from Tokyo: 14 took 289 to 693 ms from opening to answers (the other two, about 2.2 s, include waiting for a late round opening), and every answer landed in the block the pack opened in or the next.
- The contenders at their most accurate setting
- They were timed at the split size they can complete, not the smaller one that scores best. The smaller one is slower, so the time gate here is understated, not overstated.
- Questions written the way the real factory writes them
- The accuracy sample was assembled from two card skeletons by a template. The factory now rewrites every ticket with a model; its packs are expected to be harder for every system. On the eight factory packs of the testnet run Jev made no error, which is too small a sample to settle it.
- Anything economic
- No revenue, no yield, no return. Nothing on this site is a projection.