Each model is built and handed to a solver — run() and
optimize() are never called. Every rival model is hand-written in
bench/models/ beside ours and is solved against ours in CI, so
nothing here is fast because it built a different model. The legend carries
six lines for the five: gurobipy appears twice, the same
library written as a loop and as a matrix.
Every number here is the median of a measurement's rounds, and every band is the first to the third quartile of the same rounds. Nine rounds is the floor; a quick cell takes many more.
Log on both axes: the slope is the claim. The line is the median of the rounds — the number in the table — and the band is the middle half of them, first quartile to third. Its height is what the machine did to the same work, not a confidence interval: two lines whose bands overlap are two numbers this run cannot tell apart, and a band you cannot see is a measurement that barely moved. Median rather than the fastest round, because the fastest is a best-of-n and n is not equal — a quick cell here took 84 rounds and a slow one 9; rather than the mean, because one round in forty of a 20 ms measurement took 1.5 s, and one scheduler hiccup should not set a published number.
>30 s is a measurement the harness refused — projected past
its time budget and skipped. An em dash is a cell with no number for another
reason: gurobipy has no HiGHS.