Dimgrade
Four hundred quantitative estimation problems, graded on the number and on whether
the unit is the right KIND of thing. Answer one below and the grader runs on this page.
The unit box takes an expression, not a spelling from a list.
kJ/s, kilowatts, hp, kN m/s and
J/s are all the same kind of thing as kW; kWh is not.
Leave the unit box empty to see what a bare number scores.
- what was graded
- your unit is
- the question wants
- your answer in SI
- the answer key
Those buttons fill the boxes with answers built from this question's own key. Each
one is a way of being right or wrong that this grader separates.
What the grader is separating
Three questions from the corpus, each answered several ways. Every verdict and every sentence
in these tables was produced by dimgrade.grading.grade in Python when this page was
built, so they are the tool's output rather than a description of it.
power_budget-121: A roof carries 22 m^2 of panels at 18 percent efficiency, in full sun of 1,000 W/m^2. Answer in kW.
the key: 3.96 kW, a power, band 1.01
estimation-366: An emergency store has to cover 50 people for 7 days of drinking, cooking and washing water. Roughly what volume is that? Answer in m^3.
the key: 5.25 m^3, a volume, band 10.0
temperature-091: A thermostat reads 14 degrees Celsius. Answer in degF.
the key: 57.2 degF, a temperature, band 1.001
Why a spelling is never compared
Every unit expression is reduced to a factor that takes it to SI base units and a vector of
exponents saying which base quantities it is made of. Comparison is over the vector. This is
python3 -m dimgrade unit, run when the page was built:
Litres per hundred kilometres is an area because volume per distance is an area.
Milligrams per kilogram per day is a frequency because a mass fraction is dimensionless. Both are
physics rather than defects, and the second one is a limit worth knowing: a dimension is not an
identity, so an answer of 3 Hz to a dosing rate question would pass the dimensional
check. The corpus asks those questions in a form whose answer is a mass or a volume.
What the baselines score
A benchmark with no baseline is a file. None of these calls a model; every one is a pure
function of the corpus. The third row is the one the project exists for, and the row under it is
its mirror.
Twenty five miles an hour IS 11.176 metres per second, so a grader that compares physical
quantities has to accept either. 90 of the
400 problems are single step conversions with that property. They are flagged in the
corpus and every row is reported a second time with them set aside.
The tolerance is per problem and says why
A band is a multiplicative window: an answer counts if it lands between the truth divided by
the band and the truth times the band. One global tolerance would make half the corpus free and
the other half impossible.
The ten families
The grading on this page is a port, and the port is checked
A page cannot run Python, so the quiz above runs a JavaScript translation of
dimgrade/units.py and dimgrade/grading.py. Two implementations of a
grader are two sets of scores waiting to disagree, so scripts/check_js_port.py runs
both of them over every problem in the corpus against 30 ways of answering each, plus
40 real replies from a model frozen on disk, and requires the verdict,
the number and unit read out of the reply, the SI value, the dimension vector and the sentence of
explanation to be identical in every case.
The spread produces 4928 correct, 1273 numeric_off, 2373 dimension_wrong, 412 unit_missing, 630 unit_unreadable, 800 unit_ambiguous, 447 number_unreadable, 1177 no_answer, so every verdict this grader can reach is
compared rather than only the common ones. The check also runs a deliberately broken copy of the
port, with a watt hour worth 3601 joules, and fails if that copy is NOT caught, because a
comparison that passes while comparing nothing looks exactly like one that passes because two
things agree. This page is built by a script that runs the check first and refuses to write the
file if anything disagrees.
One divergence is known and stated rather than papered over: Python's
\d matches any Unicode decimal digit and JavaScript's matches 0 to 9, so an answer
typed in Devanagari digits would be a number to the Python grader and unreadable to this one. No
problem in the corpus and no answer in the spread uses them.
What is not here
Running an actual model is scripts/run_model.py, deliberately outside everything
the tests cover, because model output is not deterministic and nothing that varies between runs
belongs on a page that has to rebuild to the same bytes. One run against a forty question sample
is frozen under fixtures/ and graded by the test suite, and those forty replies are
part of the comparison above.