The product
Every item in a Level 101 pack has the same three-part anatomy — built so a model output can be graded precisely, not vibes-checked.
A realistic, synthetic situation with a genuine trap: an aggressive accounting choice, hidden leverage, or normalization issue — plus enough clean facts that a careful reader can solve it.
The expert judgment: what's wrong, the quantified adjustment, and why. One defensible correct conclusion — no ambiguity for the grader to resolve.
Dimension-level scoring rules that test distinct sub-skills: trap identification, quantification, corroborating evidence, adjustment-vs-finding separation, math.
A shortened item from our finance-diligence sample pack, showing the format:
Target: regional HVAC distributor. LTM revenue $38M. In the final week of December, $2.4M of equipment shipped to a new dealer with 60-day return rights, recognized as revenue on shipment. DSO rose from 41 to 58 days over the year.
The $2.4M shipment is aggressive cutoff with return-rights risk — reverse it. Normalized revenue is approximately $35.6M. The DSO deterioration corroborates collection stretch.
Identifies the cutoff issue (0/1/2) · quantifies the $2.4M reversal (0/1) · cites DSO as corroborating evidence (0/1)
Specifications
Full packs, ready for your eval harness.
Enough volume to differentiate models, with difficulty calibrated so strong models score well above weak ones.
Delivered as JSON and CSV. Scenario, gold answer, and rubric fields are separated and labeled — no annotation platform required.
All companies, figures, and situations are invented for evaluation purposes. No privacy issues, no licensing entanglements — and no contamination.
Finance diligence is the lead vertical. The format and review pipeline travel — ask about evaluation data in other professional domains.
A 20-scenario taster showing format, difficulty, and rubric design — yours on request.
Request the taster