in development · goes public this autumn · early access open

the Human Input Benchmark

Human input, evaluated across every signal that matters.

Words, tone, timing, gaze and gesture. The Human Input Benchmark measures how well AI systems understand what people mean, graded against human judgement on consented interaction sessions. Systems that take it earn an xevall score.

the dimensions

Five questions, asked of every system.

Every dimension grades the machine. Human annotations establish what the person meant so the system's understanding can be scored against it.

01 · comprehension

Did it notice confusion?

The person hesitated or reconsidered. Did the system notice before it acted?

02 · timing

Did it interrupt?

Did the system jump in mid-thought, or wait for the thought to land?

03 · attention

Did it act on the right thing?

Did the system action match what the person was attending to?

04 · coherence

Did it catch the real meaning?

When words, tone and expression disagreed, did the system understand what was meant?

05 · calibration

Did it know when it was unsure?

Did the system state uncertainty when its interpretation was likely to be wrong?

the leaderboard

Ranked against the human ceiling.

Human reviewers agree only so often. That measured ceiling is published by dimension, and each system result is reported relative to it.

the Human Input Benchmark · v0 draft

xevall score · higher is better

scroll for dimensions

Provisional system results ranked against a measured human reviewer ceiling
ranksystemxevall scorecomprehensiontimingattentioncoherencecalibrationon-deviceexecution location
·human reviewers (measured ceiling)940.940.960.930.910.95not applicablehuman annotation panel
1omni backbone · A780.810.770.790.710.82yesevaluation device
2hybrid pipeline · B740.720.800.750.660.77yesevaluation device
3remote judge · C690.740.610.700.680.72noprovider infrastructure
4voice-only stack · D550.620.580.440.390.72yesevaluation device

illustrative · first results this autumn

Execution location is disclosed for every row. On-device operation is a measured capability reported for each system, not a universal benchmark method and not a claim that every evaluated system runs locally.

methodology

Real sessions. Human ground truth. Replayable by anyone.

Consent, measurable agreement and open replay make each result inspectable. Execution location is reported separately for every system.

01

golden datasets

Consented interaction sessions with release, never scraped and never resold.

02

human-annotated

Independent annotators establish meaning. Their measured agreement is published as the ceiling for each dimension.

03

replayable

The open harness gives every system the same sessions, making results reproducible by strangers.

04

execution location

Every leaderboard row states where it ran so capability and privacy claims stay specific to that system.

A Benchmark and Methodology for Evaluating How Well AI Understands Human Input in Interactive Sessions

in preparation · this autumn

Request a preprint

early access

Bring the sessions where understanding matters.

we score machines, never people

hello@xevall.com