01 · comprehension
Did it notice confusion?
The person hesitated or reconsidered. Did the system notice before it acted?
in development · goes public this autumn · early access open
the Human Input Benchmark
Words, tone, timing, gaze and gesture. The Human Input Benchmark measures how well AI systems understand what people mean, graded against human judgement on consented interaction sessions. Systems that take it earn an xevall score.
the dimensions
Every dimension grades the machine. Human annotations establish what the person meant so the system's understanding can be scored against it.
01 · comprehension
The person hesitated or reconsidered. Did the system notice before it acted?
02 · timing
Did the system jump in mid-thought, or wait for the thought to land?
03 · attention
Did the system action match what the person was attending to?
04 · coherence
When words, tone and expression disagreed, did the system understand what was meant?
05 · calibration
Did the system state uncertainty when its interpretation was likely to be wrong?
the leaderboard
Human reviewers agree only so often. That measured ceiling is published by dimension, and each system result is reported relative to it.
the Human Input Benchmark · v0 draft
xevall score · higher is better
scroll for dimensions
| rank | system | xevall score | comprehension | timing | attention | coherence | calibration | on-device | execution location |
|---|---|---|---|---|---|---|---|---|---|
| · | human reviewers (measured ceiling) | 94 | 0.94 | 0.96 | 0.93 | 0.91 | 0.95 | not applicable | human annotation panel |
| 1 | omni backbone · A | 78 | 0.81 | 0.77 | 0.79 | 0.71 | 0.82 | yes | evaluation device |
| 2 | hybrid pipeline · B | 74 | 0.72 | 0.80 | 0.75 | 0.66 | 0.77 | yes | evaluation device |
| 3 | remote judge · C | 69 | 0.74 | 0.61 | 0.70 | 0.68 | 0.72 | no | provider infrastructure |
| 4 | voice-only stack · D | 55 | 0.62 | 0.58 | 0.44 | 0.39 | 0.72 | yes | evaluation device |
illustrative · first results this autumn
Execution location is disclosed for every row. On-device operation is a measured capability reported for each system, not a universal benchmark method and not a claim that every evaluated system runs locally.
methodology
Consent, measurable agreement and open replay make each result inspectable. Execution location is reported separately for every system.
01
Consented interaction sessions with release, never scraped and never resold.
02
Independent annotators establish meaning. Their measured agreement is published as the ceiling for each dimension.
03
The open harness gives every system the same sessions, making results reproducible by strangers.
04
Every leaderboard row states where it ran so capability and privacy claims stay specific to that system.
in preparation · this autumn
early access
we score machines, never people