Oliver Hazard

Research note

The announcer test

A vision model looked like it could referee. Then I muted the broadcast.

The result

I gave Gemini 50 curated NBA broadcast clips and asked it to judge the call. With audio on, its output agreed with the on-court call 74% of the time. In a later saved run, I replaced the broadcast audio with silence. That run used 40 clips, and agreement fell to 20%.

These were successive founder-produced trials, not a paired cohort. The saved summaries do not say why the later run contained fewer clips. The drop is directional evidence of an audio confound, not a controlled estimate of audio's causal effect. The label was the on-court call rather than expert consensus, so agreement here is a research result, not product accuracy.

What the model heard

The transcripts showed the shortcut directly: the model quoted announcer language in its answers. Broadcast commentary had supplied evidence about the ruling. I made audio stripping the default for later evaluations because a vision result should survive without that cue.

What the literature says

RefereeBench (arXiv, April 2026) names broadcast commentary, replay selection, and officials' gestures as likely shortcuts, and reports that a frontier model misidentifies a foul class on 63.6% of its hand-selected negative clips, rising to 76.8% under suggestive wording. It scores the same model higher when clips are submitted as original video with audio rather than sampled frames, but that comparison changes modality and presentation together, so it is not an isolated audio ablation and I do not read it as replicating mine. What it does corroborate is the failure mode: these systems invent incidents and lean on cues outside the play.