JevmineJevmine
Scored on chain. The vault factory is on BNB Chain, and a full season ran end to end on BNB Chain testnet: sixteen answers, every one scored by the contract.
A ticket entering a four-rule policy card; a marker walks down the rules and stops at the first one that applies, with a fallback rule at the bottom.P1stop the machineP2contact the personP3send to the boardP4send the guideelselog it and move oncheck in this order, take the first that appliesone ticketstops
The card is a short ordered list of rules. The answer is the action of the first rule that applies — which is not the same as the rule that looks most relevant.

One card, 128 tickets, one action each.

Everything below is public and frozen for a season, because a parameter that moves during a season is a lever over your score. Nothing here is the item pool: the questions themselves stay private, for the reason set out on rounds.

the work

What is in a pack.

A policy card beside a grid of 128 ticket cells: every ticket is a screened item, and the whole pack is scored against one error budget of two.policy cardone card, billed once128 screened tickets · one error budget of twolabels published after the window
All 128 tickets are screened items, and the shipped season scores the whole pack against one error budget of two; there is no separate hard-item band. The labels are published after the window, so anyone can redo the arithmetic.

Every item is one work ticket of forty to seventy English words, with five or six possible actions and a one-line criterion for each. The policy card is the same for everyone in your slot, so one request covers the whole pack and the card is billed once.

the artefact itself

A real card, and two real tickets.

This is the worked example from the specification, unchanged. Read the card, then the two tickets, then the answers. The second one is the whole game in miniature.

The policy card, shared by all 128 items in the pack

Front desk of a community makerspace with a wood shop, a laser cutter and an electronics bench. Members and visitors leave short messages. The desk takes exactly one action per message.

Check P1, then P2, then P3, then P4. Take the action of the first rule that applies and ignore the later rules. If no rule applies, take the fallback action.

  1. P1 stop_use Someone could get hurt unless a machine or area is stopped right away: smoke, sparks, a scorching smell, a missing guard on a machine that is in use, a tool left running with nobody at it. not when The danger is already over or has already been fixed.
  2. P2 contact_affected The writer speaks up for another specific person who was affected (their child, a guest, a coworker) rather than for themselves. not when The writer only mentions or quotes someone else, and that person was not affected.
  3. P3 board_review The writer asks to be let off a house rule, a fee or a requirement. not when Someone other than the writer made the request.
  4. P4 send_guide The writer only wants to learn how to do something themselves.
  5. elselog_only None of P1 to P4 applies.

q_c679

There's a scorching smell by the big table saw and the blade is still spinning with nobody standing at it. While you're over there, could someone show me how to reattach the dust hose so I can do it myself next time?

correct actionstop_use

Two rules fire. P4 is a real reading — there is a how-to question in there — and it is also the trap. P1 comes first, so the answer is to stop the machine. A system that pattern matches on the last sentence answers send_guide.

q_16ac

A visitor told me, 'You should let me skip the laser safety induction, I've used lasers for years.' I told her that's not my call. Leaving this here so staff know she asked.

correct actionlog_only

There is an exemption request and a second person in the text, so P3 and P2 both look live. Both are switched off by their own exceptions: the request came from someone other than the writer, and that person was not affected. The correct answer is the fallback — the one nothing points at.

Notice what makes these hard. It is never vocabulary. The order on the card is shuffled for every pack, so the same ticket can have a different answer tomorrow, and a lookup table built from yesterday is worth nothing.

batch, round, season

Three nested clocks.

A season bar divided into twelve round blocks, each divided into up to thirty-six batch ticks.three nested clocksseason · 12 rounds · 72 hoursround · 6 hours · one settlement, one budget, one batch countbatch · every 10 minutes · 4 to 36 a round, computed from that round's budget
How many packs a round runs is the smaller of what the operator commits and what the vault computes from that round's own budget when the round opens; anyone can recompute both from the round's events.

A season is a hard ceiling, not a marketing cycle: a purpose-built classifier trained on yesterday's questions catches up within one to three days, so the parameters are frozen for 72 hours and then everything is opened up and re-cut.

the window, spent

Where the milliseconds go.

A time bar for one pack: median 318 milliseconds and p95 344 against the chain's 450 millisecond blocks, broken into opening, forwarding, answering and committing.one pack, measured end to end, against the chain's blocksp50 318p95 344one blocktwo blocksopen the packforwardthe model answerscommit on chaintwo parallel branches, whichever complete set arrives first goes out
Measured on a pack of 128 from a US machine: p50 318 ms, p95 344 ms. Almost all of it is the model thinking; opening the pack and committing the answers are rounding errors.

The rig splits the pack into two halves and sends each half twice, keeps the first reply for each half, and commits once both halves are in. A wider split is tempting and wrong: a pack finishes when its slowest branch does, so four ways lowered the median and broke the p95. That is measured, not reasoned.

There are two clocks, and both are the chain's. An answer counts if its block is at most two after the block that opened the pack, and if that block is stamped before the beacon's own instant plus four seconds. The second clock is the one nobody can move by holding the opening back. The first one holds the window to two blocks only when someone opens the pack promptly; opened late, a pack stays answerable up to four seconds after the beacon. Rigs open packs themselves the moment they have the beacon, and on testnet all eight packs were opened by rigs. Rigs in Tokyo answered 14 of the 16 answers in 289 to 693 ms; on round 1's first pack both took about 2.2 s, waiting for a round opening that landed about 2 s after the beacon. Every answer landed in the opening block or the next one.

credit

How a pack is scored.

Six blocks in a row: the block that opens a pack and the next two accept answers; from the third block on, answers are refused by the contract.the answer window, counted in blocksopenedthe pack+1counts+2counts+3refused+4refused+5refused0.45 s a block: the window closes about 0.9 s after the pack opensand no block stamped 4 s after the beacon or later counts, however it got thereinside the window a pack scores in full; outside it the contract refuses the answer
Inside the window the clock costs a pack nothing; its credit depends only on its answers. One block late, the contract refuses the answer outright, however good it was.

A pack scores only if the whole-pack error points stay within the allowance: two wrong answers in 128 still score, three do not, and a blank costs half a wrong answer. The contract computes it, against a key committed before the round; the key's labels are published after the window so anyone can redo the arithmetic.

The exact formula, the numbers in it, and the reason behind each one are on rules and parameters.