Skip to main content
T·M TOMFORMER live research instrument Patent pending

Run the bench yourself.

A fixed dev slice, streamed through all three arms. Every result is reported as precision @ coverage. A precision figure never appears here without its coverage. An arm that declines the questions it cannot answer will always score higher precision than one that guesses at them, while attempting fewer. Both halves of that trade have to be on screen for either number to mean anything.

Slice
MEASURE · BENCH

Shows: precision reported at its coverage, broken out per query family, with latency percentiles from a monotonic clock around the forward and warmup excluded and counted. Does not show: a single overall winner, or valid comparative RAG deficits unless the N11 corrected artifact/build basis is public beside the RAG figures. The families are different tasks won by different architectures, so any blended figure measures the probe's family mix rather than capability, and an arm that abstains scores higher precision while attempting fewer questions.