Accuracy, honestly

Our numbers, published

A referee — human or AI — earns trust by being right, and by being accountable when it isn't. Here is exactly how good Verdict is today, how we measure it, and what it can't call yet.

Numbers current as of July 2026. This page is updated as models improve — measured results only, never projections.

How we test

Verdict is trained and evaluated on 23,515 quality-labeled sabre touches from real competition video, each labeled by experienced fencers and referees with the winner and the action. Accuracy is always reported on held-out footage the model has never seen — including a separate benchmark drawn from international-level competition, which is deliberately harder than typical club footage.

Seeing the fencers: 94% blade tracking

The foundation of every call is knowing where each fencer — and each blade — is in every frame. Verdict's pose model tracks 16 keypoints per fencer, including the blade guard, midpoint, and tip, at 94% keypoint accuracy on held-out competition footage. Blade tips at full lunge speed are the hardest single thing to track in fencing video; this is where most of our engineering effort has gone.

Making the call: two-light accuracy

One-light touches are unambiguous — the light decides, and Verdict simply reads the box. The AI's job is the two-light touch, where right-of-way determines the winner. On those:

MeasurementAccuracyNotes
All two-light calls (club-level held-out set) 71% Every call, including ones the model is unsure about
All two-light calls (international competition benchmark) 69% Harder footage, faster actions
High-confidence calls (top 20% by model confidence) 92% When Verdict is sure, it's almost always right

Confidence matters as much as accuracy: when Verdict isn't sure, it says so rather than guessing. In the app, low-confidence calls are flagged as such — the same way a good referee acknowledges a close call.

By action type

Not all touches are equally hard to call. Our current per-action accuracy on held-out two-light touches:

ActionAccuracyStatus
Attack vs. counterattack81%Strongest — the most common two-light situation
Attack in preparation78%Strong
Remise / reprise74%Solid
Parry-riposte49%Our hardest problem — see below

What Verdict can't call well yet

  • Parry-riposte. Telling a real parry from a blade graze requires knowing exactly when and how blades touched — information that's genuinely ambiguous in 2D video. We're attacking this from two directions: audio analysis (the sound of blade contact) and 3D pose reconstruction. This is the single biggest driver of our accuracy roadmap.
  • Rare and messy actions. Simultaneous attacks and unusual phrases are underrepresented in training data and called less reliably. These are flagged as low-confidence.
  • Poor footage. Verdict abstains rather than guesses when it can't see the scoring box or track the fencers reliably — severe backlight, heavy occlusion, or a camera angle that hides one fencer.

What about speed?

We're benchmarking end-to-end analysis time now and will publish measured numbers here — the same rule applies to speed as to accuracy: we only claim what we've measured.

The trajectory

Every number on this page is a snapshot, not a ceiling. The dataset grows weekly through our expert annotation team, and the two biggest known gains — blade-contact audio and 3D pose — are in active development. When the numbers change, this page changes.

Watch the numbers climb

Early-access members get accuracy updates as they happen.

Free tier at launch · One email when your spot opens — no spam.