Skip to main content
T·M TOMFORMER live research instrument Patent pending

ToMFormer / inspect the result

Correct it. Teach it. Ask it to reason. Verify the decision. Inspect what it takes to scale.

Five narrow proofs. Each starts from this page, states what you should observe, and says what the result does not establish. Live measurements identify the running deployment; banked results carry their own date and scope.

Correct → Teach → Reason → Verify → Scale

Five visitor-verifiable proofs

A mock response demonstrates the interface only. It does not complete a live proof.

  1. 01 / Correct LIVE

    Change the answer. Put it back.

    Ask one question, apply one fact correction, ask the identical question again, then revert.

    Observe
    The next answer changes; revert restores the session’s prior observable state.
    Limit
    A temporary, expiring session overlay. The shared base is never retrained and never permanently changed.
    Run the correction proof →
  2. 02 / Teach LIVE

    Restore knowledge. Check what moved.

    Withhold a curated document, measure its target questions, restore it, and rerun a fixed prior probe.

    Observe
    Target knowledge returns; the same retention rows are compared before and after.
    Limit
    Bounded withhold-and-restore, not arbitrary novel-subject learning.
    Run the no-forgetting proof →
  3. 03 / Reason BANKED + LIVE

    Follow the relation hops.

    Inspect explicit two-hop and three-hop questions where the answer differs from the starting document entity.

    Observe
    Hop count, relation sequence, answer identity, and top-k position.
    Limit
    One live example is an anecdote. The 0.4192 @1 result is banked: family 0, n=167, 2026-07-30, gate off.
    Open the relation-hop proof →
  4. 04 / Verify BANKED

    Move the serve threshold.

    Inspect precision and coverage together as TOMX serves fewer answers under recorded operating points.

    Observe
    Precision 0.4192 → 0.5636 while coverage 1.000 → 0.329 on family 0.
    Limit
    Four banked operating points, not proof that every unsupported input is detected.
    Run the abstention experiment →
  5. 05 / Scale LIVE

    Measure the physical system.

    Read allocated model bytes, served disk, shipped-but-unserved data, and process memory from this deployment.

    Observe
    6,460,208 parameters / 24.644 MB allocated; 841.965 MB served knowledge.
    Limit
    Allocated model bytes are not total RAM; the larger shipped world is not all queryable.
    Open the physical-scale proof →

Research-grade generative view

Scoped 5M speech with silence when unsupported

The speech panel emits only licensed entity names from the 5M-world candidate namespace, reports spoken rate with precision, and keeps candidate-absent rows silent instead of free-running an invented name.

Boundary: text-only name emission, not voice synthesis, celebrity impersonation, open-ended chat, or unlicensed generation.

Open scoped speech proof

GL-2 conversational readout

Bounded talking-Floyd with visible abstention

The Floyd panel shows a fixed GL-2 transcript with support-gated text emission, declined duplicates, and explicit ABSTAIN rows when a request crosses the displayed support boundary.

Boundaries: self-speech is only a card readout; cross-entity speech is blocked; ambiguous coreference abstains; coherence is local, not unrestricted dialogue planning.

Open bounded Floyd

Separate device / catalog deployed 2026-08-18

A handheld sensing unit that names what it is pointed at

The same idea in hardware. A battery-powered unit reads a surface and answers with every label it has evidence for, each carrying its own score and its own threshold, while its catalog sits on a memory card. 1,219 gate reads produced no wrong name, the resident file has stayed the same size to the byte across every deployed catalog from 37 labels to 50, and a surface it has never seen can be taught in the field in seconds with nothing retrained.

Boundary: none of this runs on this web deployment and your browser cannot reach the device. It does not detect trace contamination and it does not identify pathogens. It never reports a surface as safe.

Open the scanner record

Separate system / recorded 2026-08-19

A voice system that knows who is talking and waits to be addressed

The same idea in conversation. Every person or agent in the room is one editable record: enrolled voices are named live, a stranger reads NOVEL-UNKNOWN until enrolled, and an agent speaks only when a read of its own record licenses it, with silence as the default. On the hardest recorded meeting bench it named wrongly once in 1,175 display events at coverage 0.209, and its agents produced zero false interruptions through real speech recognition. The page carries a recorded run you can play.

Boundary: none of this runs on this web deployment; the page replays a recording of the real system, labeled as a recording. It is not spoof-proof and never claims to be, and it never claims prompted agents cannot be silent.

Open the voice record

How to read the evidence

LIVE

A current request against the running driver. Capture and deployment context belong with the result.

BANKED

A prior measurement with a stated date, scoring unit, sample count, and available deployment pin.

RESEARCH-ONLY

A different artifact or harness. It cannot complete one of these deployed proofs.

BLOCKED

A validity or provenance gap prevents comparative use. N11: the deployed RAG text path remains blocked from comparative claims until the corrected artifact and build basis are publicly pinned.

Reproduce supported operations through the API

A trial token exposes scoped ask, edit, grow, and revert operations. It expires and excludes bench runs. Read the contract or inspect the machine-readable OpenAPI document.

curl -s -X POST https://tomformer.com/api/v1/trial-token
# scopes: ask, edit, grow; bench excluded

Agent-readable versions: /agents.md and /llms.txt.

Every panel

Each page states what it shows and what it does not in its own footer. That footer is the scope of its claim.

Answer

Audit

Change

Emit

Sense

Hear

Measure

OVERVIEW

Shows: what this deployment measures, where each number comes from, dated results from the research program labeled as research results, dated readings from a separate handheld sensing device labeled as a separate device, and the N11 RAG validity/provenance block wherever RAG numbers appear. Live panels identify current measurements; banked numbers show date and scope. Does not show: a general assistant (TOMX answers over its fixed corpus and declines only under the displayed margin gate), a single overall winner across query families, valid comparative RAG deficits until the corrected artifact/build basis is public, or how any of it works inside.