Oliver Hazard

Benchmark card

We attacked our own metric first

The original score looked too clean. I rebuilt the test around legal plays chosen to look like fouls, then measured what survived.

01

The first score

The v2 model reached 96.5% F1 on held-out clip-level foul detection (n=358 clips). It was a curated, easy-negative test at a permissive threshold. That result measured whether each short clip contained a foul. It did not measure full-game retrieval or referee judgment.

Label provenance. The clip classes came from the curated foul-detection corpus, with founder/human labels: play-derived foul categories and assembled non-foul clips. They were not independently expert-confirmed consensus judgments. These are founder-produced research results, not product accuracy claims.

02

Make it adversarial

I responded by mining hard negatives from legal plays that shared the surface cues of fouls. On the rebuilt test, 1,070 negatives outnumbered 301 fouls by 3.6:1. The same clip-level task had become about 7× more adversarial.

03

What held

The successor v6 model held 0.94 AUROC. At the selected operating point, recall was 87% and precision was 68%. AUROC describes ranking across thresholds; recall and precision describe one chosen point.

The honest successor headline is narrow: after I made the benchmark adversarial, discrimination held. The test remains clip-level and uses curated clips. It does not establish transfer to a new camera domain or production accuracy.