Benchmark card
We attacked our own metric first
The original score looked too clean. I rebuilt the test around legal plays chosen to look like fouls, then measured what survived.
The first score
The v2 model reached 96.5% F1 on held-out clip-level foul detection (n=358 clips). It was a curated, easy-negative test at a permissive threshold. That result measured whether each short clip contained a foul. It did not measure full-game retrieval or referee judgment.
Label provenance. The clip classes came from the curated foul-detection corpus, with founder/human labels: play-derived foul categories and assembled non-foul clips. They were not independently expert-confirmed consensus judgments. These are founder-produced research results, not product accuracy claims.
Make it adversarial
I responded by mining hard negatives from legal plays that shared the surface cues of fouls. On the rebuilt test, 1,070 negatives outnumbered 301 fouls by 3.6:1. The same clip-level task had become about 7× more adversarial.
What held
The successor v6 model held 0.94 AUROC. At the selected operating point, recall was 87% and precision was 68%. AUROC describes ranking across thresholds; recall and precision describe one chosen point.
The honest successor headline is narrow: after I made the benchmark adversarial, discrimination held. The test remains clip-level and uses curated clips. It does not establish transfer to a new camera domain or production accuracy.