JevmineJevmine
Scored on chain. The vault factory is on BNB Chain, and a full season ran end to end on BNB Chain testnet: sixteen answers, every one scored by the contract.
Bar chart of median pack latency measured from us-west-2 for two Jev configurations and two contenders, with each configuration's slowest pack marked, against a 350 millisecond line, the design's off-chain target; only the Jev bars finish before the line.pack of 128, us-west-2, medianJev · 2-way split227Jev · one request321Contender A1375Contender B1804350 msmilliseconds
Median wall-clock time for one pack of 128, measured from us-west-2, where the rigs run. The short tick on each bar is that configuration's slowest pack. The line is 350 ms: both contenders finish on the wrong side of it, several times over.

Both gates are measured, not modelled.

Measured on 2026-09-22. Raw per-item output and per-pack timings are in the repository; every number on this page is derived from those files by a script at build time, not typed in by hand; the notes around them quote the testnet record. Total spend across both runs: $1.17.

What was not measured. The scope of the measurement: the us-west-2 numbers are 4 packs per configuration from a shared serverless function, and the two testnet seasons were 16 answers from two rigs in Tokyo.

6 packs per configuration

The clock.

Configurationp50p95on time at 400 mscost / pack
Jev, 2-way split318ms344ms100%$0.00132
Jev, one request465ms487ms0%$0.00130
Jev, 4-way split299ms494ms83.3%$0.00138
Contender B, 8-way880ms1061ms0%$0.00656
Contender A, 8-way1538ms1958ms0%$0.00331

Two things in that table matter more than the winner. First, the split size: two ways holds the tail, four ways has a lower median and a p95 of 494 ms, because a pack finishes when its slowest branch does. Second, the contenders were run at the split size they can actually complete — one of them cannot emit 128 labels in a single reply at all, so it must pay for the policy card eight times over.

That table was taken from us-east-2 through a proxy: the long way round. The one below was taken from us-west-2, the region the rig template deploys to and the region the model is served from, with 4 packs per row.

Configuration, us-west-2p50slowestwithin 350 mswithin 400 ms
Jev, 2-way split227ms443ms3 of 43 of 4
Jev, one request321ms360ms3 of 44 of 4
Contender A, 8-way1375ms1386ms0 of 40 of 4
Contender B, 8-way1804ms2004ms0 of 40 of 4

Network round trip from us-west-2: the model's API 18 ms, one contender's host 81 ms, another's host 44 ms. The rivals start further away before they have answered anything.

3840 candidates, $0.92

The hard items.

A four-step funnel from 3840 generated candidates down to 90 items in the hard band.what the sample narrowed to3840candidates generated2680retained · 69.8%screened before any pack can use them322separators · 12% of retainedat least one contender got it wrong90the hard bandthe whole accuracy gate rests herethe screening rule is a defence parameter and is not published; its counts are
3840 generated, 2680 kept (69.8%), 322 of those separate one system from another, and 90 of those are the band the gate actually rests on.
Systemon the kept poolon separatorserror rate on the hard band
Jev100.0%0 wrong of 3220 of 90
Contender B94.1%53%39.3%
Contender C92.6%42.8%29.2%
Contender A96.4%69.9%77.8%
Bar chart of per-item error rates on the 90-item hard band: Jev at zero, the three contenders between 29.2 and 77.8 percent.per-item error rate on the 90-item hard bandJev0 of 90Contender C29.2%Contender B39.3%Contender A77.8%Jev made 0 errors here; with 90 items the most that rules out is an error rate above 3.3%
On the hard band the contenders were wrong on 29.2% to 77.8% of items. Jev was wrong on 0 of 90.

Jev's grade came from a third, independent run — not from the runs used to screen the candidates in the first place, which would have let it mark its own homework. On the kept pool it made 0 errors in 2680 items; with a sample that size, all that rules out is an error rate above 0.1%.

40,000 simulated packs

Carried through to whole packs.

A pack is 80 anchors and 48 hard items, and it scores only if whole-pack errors stay within 4 points and hard-item errors within 3. Drawing items with replacement from the measured per-item results gives the share of packs each system would have passed:

Systemhard items drawn from all separatorsdrawn from the hard band only
Jev100%100%
Contender B0%0%
Contender C0%0%
Contender A0%0%

Contenders are simulated as if their connection never failed: each is resampled only from the items it actually answered, so a timeout is never scored as a wrong answer. Counted that way, every contender goes to zero in this 80 + 48 shape whether the hard items are drawn from all separators or from the hard band alone. The shipped season is a different shape: 128 screened items against an error budget of two, with no hard-item band. Resampled the same way, Contender A passes about 15% of those packs, B 1.7% and C 0.3%.

the uncomfortable part

Four things worth saying out loud.

The screening is what carries this

Jev is not naturally good at these items. Before screening it got a large fraction of the hardest candidates wrong. What makes the pool safe for an honest miner is the screening step every candidate passes before it can appear in a pack. If that step breaks, the first people it hurts are honest miners — not attackers.

The accuracy gate rests on 90 items

Not 322, and not 2680. 90. That is the thinnest number in this project and the next measurement to run.

We have corrected ourselves twice

The first version of this benchmark had four defects, all in our favour: only Jev was given the per-option criteria, the questions did not depend on rule order, the contenders ran at a split size that is not their best, and a rate-limited contender was scored as wrong. After that was fixed, a second error remained: a contender that returned nothing was still counted as having answered wrongly, which more than doubled the separator count. Every number on this page is now computed without either mistake.

Hard items are scarce

A template generator yields about one hard item in eight kept items, and in live factory runs it could not fill a band of 48. So the factory has a model rewrite every ticket, and the shipped season scores whole packs of 128 screened items against an error budget of two, with no separate hard-item band.

in full

What was not measured.

The latency run from the rig region
us-west-2 was measured with 4 packs per configuration from a serverless function, which shares CPU and can cold-start. The testnet seasons added 16 answers from Tokyo: 14 took 289 to 693 ms from opening to answers (the other two, about 2.2 s, include waiting for a late round opening), and every answer landed in the block the pack opened in or the next.
The contenders at their most accurate setting
They were timed at the split size they can complete, not the smaller one that scores best. The smaller one is slower, so the time gate here is understated, not overstated.
Questions written the way the real factory writes them
The accuracy sample was assembled from two card skeletons by a template. The factory now rewrites every ticket with a model; its packs are expected to be harder for every system. On the eight factory packs of the testnet run Jev made no error, which is too small a sample to settle it.
Anything economic
No revenue, no yield, no return. Nothing on this site is a projection.