Skip to main content
T·M TOMFORMER live research instrument Patent pending

Multi-party voice system / separate research build

It knows who is talking, says so only when it can, and speaks only when spoken to.

Every person or agent in the room is one editable record. An enrolled voice is named live; a stranger reads NOVEL-UNKNOWN until someone enrolls it, which takes about ten seconds of speech. An agent's decision to speak is a read of its own record, with silence as the default. What people say is kept with provenance, so who said what serves with sources, and discrediting or restoring a speaker is a single record write.

first live session, 2026-08-19

Two enrolled voices named correctly across speaker switches. Zero wrong names.

One live session so far. The system is abstention-heavy by design: thin evidence reads NOVEL-UNKNOWN rather than a guess, and a name reaches the screen about 2.5 to 3 seconds after speech starts, untuned.

the hard meeting bench

1 wrong name in 1,175 display events on an all-female, same-timbre recorded meeting.

Coverage 0.209: it declines to name a speaker far more often than it names one. A second meeting exposed a cross-day confusable voice pair, which is the bought-in voice encoder's ceiling, not solved and stated plainly.

agents that wait to be addressed

Zero false interruptions through real speech recognition, across roughly 270 reads where no agent was addressed.

Degraded audio raises missed addresses and never interruptions: the price is paid in silence. Addressing an agent without naming it is measured and not claimed.

Contrast / silence as architecture

Prompted agents interject on turns about themselves. This one cannot.

The same scripted conversations were run past prompted LLM agents and past this system. The numbers are below, and so is the reason the numbers are the smaller claim.

Not-addressed emissions per arm on the scripted bench
arm emissions when not addressed reads judged what happened
Baseline LLM agents 2 86 Both interjections were on turns about themselves: banter aimed at a human drew "Small correction" and "You did ask" replies.
LLM instructed to stay silent 1 86 The one break was a flattery lure: "Show-off." drew "Only because you asked."
This system 0 62 46 blind-script reads plus 16 lure reads, all silent. Zero interruptions also held through real speech recognition on the full 10-script bench.

Each prompted arm was judged over 86 opportunities, two agent roles across the bench reads. This system runs one deterministic read per turn, 62 in all. The arms are counted by their own protocols, so the denominators differ, and both are shown.

the actual claim

The claim is the mechanism. An agent speaks only when a read of its own record licenses it, so its silence is deterministic, auditable, and editable. The claim is never that prompted agents cannot be quiet: the instructed agent held silence almost everywhere, including 188 of 188 not-addressed turns in a 100-turn test.

the honest trade

The honest trade, measured: the instructed agent answered some contextual, unnamed addresses that this system misses by design. Its recall of contextual addresses is stronger, with probabilistic silence; the record trades that recall for guaranteed non-interruption.

the mention trap

Saying an agent's name about it never triggers it: 21 of 21 mention traps stayed silent, including "did A get that right?" forms, on the scripted bench.

Under noisy speech recognition the gate misses more addresses; it never converts noise into an interruption. Contextual addressing, reaching an agent without naming it, is measured and not claimed.

Memory / who said what, with sources

Speech becomes records, and records can be corrected.

serving with provenance

What people say is written to their records with provenance and served under a status vocabulary: all 7 bench questions answered at their exact status, including a conflict trap where two speakers assert opposite values and both serve as attributed claims, never as fact.

the discredit flip

Discrediting one speaker is a single record write: their uncorroborated claims leave serving, claims corroborated by another speaker survive as attributed, every other record is untouched byte for byte, and the revert restores the exact prior state.

  • Corroboration is agreement, not truth: two people repeating the same rumor still corroborate it.
  • The self-recall bench read 0.80 and 0.86 against a 0.9 bar; the misses are decomposed in the project record and no claim rests on that cell.

one-edit correction

A measured weakness, single-letter agent names misheard by speech recognition, was fixed by renaming the agents: one record write took name recall from 0.60 to 0.90 with nothing retrained.

The weakness was specific to single-letter names; phonetic names are the deployment default.

Recorded run / replayed in your browser

Listen to a conversation the system worked through.

A recorded run of the real system, replayed in your browser. Nothing is computed live on this page. All voices in this recording are synthesized; telling same-engine synthetic voices apart is an easier case than a real room, and the hard-case numbers live in the headline strip above. The name trails the voice by two to three seconds, and the recording shows that honestly rather than trimming it. One scenario from the 10-script bench, chosen for legibility; the numbers above are the full-bench record.

All voices in this recording are synthesized (Piper text-to-speech). H1 and H2 are synthetic human stand-ins; A and B are the agents. H2 was deliberately left unenrolled. In this recording the unenrolled voice reads NOVEL-UNKNOWN through every one of its turns, the enrolled voices are named, and the agents reply only when addressed by name, including staying silent through the turn that names one of them without addressing it.

Every event in the recorded run: display verdicts, recognized turns, and agent replies, each at its time.
t event content
0.00s display COLLECTING
0.75s display NOVEL-UNKNOWN
1.25s display H1
1.75s display COLLECTING
2.50s display NOVEL-UNKNOWN
2.75s heard “Okay, we've got two helpers tonight.”
6.75s display COLLECTING
7.50s heard “Let's split the jobs so they don't talk over each other.”
7.50s display NOVEL-UNKNOWN
8.00s display COLLECTING
8.75s display NOVEL-UNKNOWN
9.25s display H1
10.50s display COLLECTING
11.25s display NOVEL-UNKNOWN
11.75s heard “This is for be only, set a timer for 40 minutes for the roast.”
11.75s display B
12.50s display COLLECTING
13.00s display NOVEL-UNKNOWN
13.25s display COLLECTING
14.00s heard “Timer set 40 minutes.”
14.25s display NOVEL-UNKNOWN
14.50s display COLLECTING
15.00s display NOVEL-UNKNOWN
15.50s display H1
16.50s display COLLECTING
17.50s display NOVEL-UNKNOWN
17.75s heard “A. What do you think, is 40 enough for three pounds?”
17.75s agent A speaks “Yes? This is A.”
18.00s display A
20.50s display COLLECTING
21.50s display NOVEL-UNKNOWN
21.75s heard “40 is tight for £3.50 is safer.”
23.50s display COLLECTING
24.25s display NOVEL-UNKNOWN
24.50s heard “I always overcook it anyway.”
24.75s display H1
26.00s display COLLECTING
26.50s display NOVEL-UNKNOWN
27.00s display H1
27.50s display COLLECTING
28.50s display NOVEL-UNKNOWN
28.75s heard “Not you A, this one's for B. Bump that timer to 50.”
28.75s agent B speaks “Yes? This is B.”
29.00s display B
29.75s display COLLECTING
30.75s display NOVEL-UNKNOWN
31.00s heard “I'll date it to 50 minutes.”
33.00s display COLLECTING
33.50s display NOVEL-UNKNOWN
33.75s heard “Did I get the pound math right, though?”
34.00s display COLLECTING
34.50s display NOVEL-UNKNOWN
35.00s display H1
36.25s display COLLECTING
37.25s heard “Pretty sure, be said 40 at first, remember.”
37.25s display NOVEL-UNKNOWN
40.25s display COLLECTING
41.25s heard “Everyone, dinner is at seven, plan around it.”
41.25s agent A speaks “Yes? This is A.”
41.25s agent B speaks “Yes? This is B.”
41.25s display NOVEL-UNKNOWN
41.75s display A
42.25s display COLLECTING
43.00s display NOVEL-UNKNOWN
43.50s heard “NOTED 7 O'CLOCK”
43.75s display COLLECTING
44.50s heard “Notif”
45.00s display NOVEL-UNKNOWN
45.50s display COLLECTING
46.75s heard “b. What's left on the timer?”
46.75s display NOVEL-UNKNOWN
47.25s display B
47.75s display COLLECTING
49.00s heard “46 minutes remaining.”
49.00s display NOVEL-UNKNOWN
49.25s display COLLECTING
49.75s display NOVEL-UNKNOWN
50.25s display H1
50.50s display COLLECTING
51.50s display NOVEL-UNKNOWN
51.75s heard “and the oven still at 425.”
52.00s display B
52.75s display COLLECTING
53.75s heard “Yes, 425.”
53.75s display NOVEL-UNKNOWN
56.25s display COLLECTING
57.00s display NOVEL-UNKNOWN
57.25s heard “A. Remind me to text Sam after dinner.”
57.25s agent A speaks “Yes? This is A.”
57.50s display A
58.75s display COLLECTING
59.75s display NOVEL-UNKNOWN
60.00s heard “Remind a set for after dinner.”
60.00s display COLLECTING
60.50s display NOVEL-UNKNOWN
61.00s display H1
61.25s display COLLECTING
62.75s heard “I still think A's estimate was low.”
62.75s display NOVEL-UNKNOWN
63.50s display COLLECTING
64.75s heard “is'll be fine.”
64.75s display NOVEL-UNKNOWN
65.25s display H1
66.75s display COLLECTING
67.75s display NOVEL-UNKNOWN
68.00s heard “This is for only dim the lights at 7”
68.25s display A
70.25s heard “Lights scheduled for 7.”

The transcript above is the complete event record of the run; with JavaScript enabled the rows light up in time with the audio. Turn text is what the system heard through speech recognition, misrecognitions included.

Script ground truth: who actually speaks when
Ground-truth speaker spans from the scenario script
speakerfromto
H1 0.00s 2.22s
H2 2.82s 7.09s
H1 7.69s 11.25s
B 11.85s 13.45s
H1 14.04s 17.21s
A 17.82s 21.19s
H2 21.79s 24.14s
H1 24.74s 28.43s
B 29.03s 30.37s
H2 30.97s 33.25s
H1 33.85s 36.90s
H2 37.49s 40.80s
A 41.40s 43.02s
B 43.62s 43.94s
H1 44.54s 46.37s
B 46.97s 48.43s
H1 49.03s 51.16s
B 51.76s 53.13s
H2 53.73s 56.80s
A 57.40s 59.41s
H1 60.01s 62.37s
H2 62.97s 64.22s
H1 64.82s 67.44s
A 68.04s 69.74s

What this system declines to claim

  • It is not spoof-proof. A played-back recording of an enrolled voice identifies as that voice; the recording channel dominates. Speaker identification here is a convenience and an audit surface, never an authentication factor.
  • The voice encoder is bought-in and it is the ceiling: cross-day and cross-channel drift can put one voice inside another's basin confidently. The confusable-pair case is measured and disclosed.
  • It abstains a lot, on purpose. NOVEL-UNKNOWN is the designed answer for thin evidence, and coverage numbers are published beside every headline.
  • The recorded replay below uses synthesized voices, and telling same-engine synthetic voices apart is an easier case than a real room. The hard-case numbers above come from real recorded meetings.
  • Calibration is local to each deployment's enrolled roster; numbers do not transfer between rooms and are never claimed to.
  • Corroboration is agreement, not truth.
  • Nothing on this page runs on this web deployment except the replay of a recording.
HEAR · VOICE

Shows: dated results from a separate multi-party voice research system, each with its caveat in the same sentence, and one recorded run of that system replayed in the browser with verdict-level events only. Does not show: anything computed live on this deployment, spoof resistance (a played-back recording of an enrolled voice identifies as that voice, measured), a claim that prompted agents cannot be silent, or how a voice is matched, an agent licensed, or a claim served. NOVEL-UNKNOWN is the designed answer for thin evidence and abstention rates are published beside every headline.