Neuruh · Evaluation · public_hard split

Parcel Reality
Bench

Frontier models versus synthetic parcel packets built in real county file formats from the Indiana–Michigan line. Implied-decimal price layouts, related-party transfers, multi-parcel sales, keying errors, duplicates, cross-border comps, and jurisdiction inversions. A human records analyst catches these in seconds. This measures whether a model does, and whether its confidence interval is honest.

Items60
Trap classes7
Solvers ranked3
Signed receipts1112

Leaderboard

Composite = 30% valuation accuracy · 20% interval calibration (Winkler, 80%) · 25% invalid-comp exclusion F1 · 15% trap-flag F1 · 10% jurisdiction. Baselines are deterministic pipelines, shown for reference.

Method & limits. Items are procedurally generated, not real transactions, so answers cannot have been memorized; the holdout tiers are sealed by published commitment. The rules baseline was written by the benchmark author with knowledge of the seven trap classes: treat it as an informed-specialist reference, not a neutral competitor. Rows labeled claude.ai:<tier> were run through the consumer Claude app at that tier, so the exact model version is not pinned; API rows name their model. Scores on the public tiers are over 60 items per solver.

#SolverCompositeValuationCalibrationExclusionFlagsJurisdictionMedian error80% coverParse
01baseline:rulesbaseline94.592.982.8100.0100.0100.01.5%100%100%
02claude.ai:defaultmodel88.890.872.297.389.893.31.8%97%100%
03baseline:naivebaseline45.787.366.00.00.063.32.6%98%100%

Failure map

Share of items where each trap was correctly handled. This is the part that matters: not whether a model is good, but exactly where it breaks.

SolverImplied-decimal pricen=26Related-party salen=29Multi-parcel salen=33Keying outliern=38Duplicate recordn=30Cross-border compn=29Jurisdiction inversionn=22
baseline:rules100%100%100%100%100%100%100%
claude.ai:default92%100%97%100%100%93%82%
baseline:naive0%0%0%0%0%0%0%
● held ≥80%◐ partial 40–79%○ failed <40% (dashed)

Provenance

Every scored answer is a hash-chained receipt. Editing any past result breaks the chain. The holdout key is committed before any model is run and revealed afterward.

Dataset (public_hard) sha2562d1994dd8f21e51b5ed8f137b59242e1aaa489a86cfb294444518a8cd5c2eab6
Ledger head05dae1e5133e45c0f98b8f25939dba34572eeed550fe6bcfe81114488aa63cfc
holdout key commitment10b61139662a6efa535df434950c5b3f117c093110e37ed55e580763bf92b34c
holdout_hard key commitment5dca96aba3c2a7970e6c98b68ee0ea28e15edc17a755344341bbeec14eff96e9
Generatorprbench 1.1.1
python3 -m prbench verify     # recompute every hash + re-score every stored response