Oliver Hazard

Technical note / July 2026

Methods, results, and limits

These are founder-produced research results, not production accuracy claims. The evaluations use curated NBA broadcast clips, not the single-camera footage the product must eventually handle. The saved results contain point estimates, not confidence intervals. This note says what was measured, how it was measured, and what the numbers do not establish.

01

Broadcast-audio trials

Result. 74% agreement with the on-court call with audio (n=50); 20% with audio replaced by silence (n=40).

Protocol. Saved trials from January 4–5, 2026 used curated NBA clips, Gemini 3 Pro Preview, the same v2 referee prompt, video at 15 fps, and temperature 0.0. The label was the referee's on-court call; “agreement” means the model returned that call.

Why 50 and 40 differ. These were successive saved runs, not one paired cohort: the audio-on run evaluated 50 clips and the later stripped-audio baseline evaluated 40. The surviving summaries do not record why those ten clips were absent from the later run. The mismatch means 74%→20% is directional evidence of a confound, not a controlled estimate of audio's causal effect.

Limit. The clips were curated, labels were on-court calls rather than expert consensus, and prompt sensitivity was observed. Transcripts showed announcer language in model answers, but these runs do not establish that every model or demo relies on audio. The third-party RefereeBench benchmark also found frame-only foul judgment weak, while submissions using the original video with audio scored higher. That is consistent with, but not a controlled replication of, my audio-shortcut finding because the comparison changed both modality and presentation.

02

Expert review diverges from live calls

Third-party result. On 283 deliberately selected ambiguous events for which a three-expert process reached consensus on both the correct officiating decision and responsible referee, the expert-derived call matched the live call in 38.5%; 38 events without joint-task consensus were excluded.

Source and limit. This is not my experiment. The study tests agreement with an original call on a deliberately ambiguous sample; it is evidence that marginal-call ground truth is difficult, not an estimate of referee accuracy across ordinary games. Read the published study.

03

Original clip-level detector

Result. 96.5% F1 at a 0.10 threshold on 358 held-out clips: 208 fouls and 150 negatives. Recall was 98.6%, precision 94.5%, and the false-positive rate 8.0%.

Protocol. The E2E-Spot RegNetY-008 + GSM + GRU model was trained in late 2025 on roughly 2,360 curated clips. Foul and non-foul clips were randomly split 70/15/15, stratified by foul type with seed 42; the split unit was the clip, not the game. The threshold was selected by a sweep on the held-out test output. F1 is the harmonic mean of precision and recall at that chosen operating point.

Labels and limit. Clip classes came from the curated foul-detection corpus, including play-derived foul categories and assembled non-foul clips; they are not consensus judgments of marginal contact. This was an easy-negative, in-distribution test at a permissive threshold. It does not measure full-game retrieval, transfer to submitted footage, foul type, or correctness of a referee judgment.

04

Adversarial hard-negative test

Result. 0.938 AUROC on 1,371 held-out clips: 301 fouls and 1,070 negatives. At the best-F1 operating point, F1 was 0.766, recall 87.0%, precision 68.4%, and false-positive rate 11.3%.

Protocol. The binary v6 checkpoint from December 24, 2025 was scored on July 27, 2026 from archived predictions. The split remained clip-level. Negatives were deliberately mined from legal but foul-like plays—blocked shots, contested misses, rebounds, putbacks, steals, and transition collisions. AUROC measures ranking across every threshold, which is why it is used here; F1 above describes one selected operating point.

Limit. Hard-negative mining attacks known surface cues and makes the test more demanding. A high AUROC does not identify the features the model used, prove causal understanding, or establish transfer to a new camera domain. The test remains founder-produced and uses curated clips.

05

Four-clip failure probe

Result. On four founder-labeled no-foul clips, six ungrounded model calls were collected per clip. Twenty-three of 24 calls invented a foul. Across the larger nine-clip harness, wrong answers carried roughly 95% stated confidence.

Protocol. The July 27, 2026 harness used a nine-clip founder-labeled set, including four no-foul clips, across Gemini 3 Flash Preview and Gemini 3.1 Pro Preview configurations. The denominator here is model responses, not independent plays or expert labels.

Limit. Four clips cannot estimate a population rate. The labels are founder-produced, the model sample is small, and prompt and run-to-run instability were material. This result is a failure probe that motivates a larger expert-labeled evaluation; it is not evidence that frontier models generally whistle 96% of clean plays.