The coach assistant: 32 sources, 70 questions, 77.4%
The corpus failed its internal check before the public re-run: 19 correct answers out of 31, or 61%. Every question tied to two video seminars ended in a refusal. The assistant said the material was not there.
The material existed. It was missing from the merged corpus.
On July 10, we added 16 public RFB Academy seminars to the original nine methodology documents. The merge completed without an error, so we treated all 16 as present. Only 14 had actually made it across. A seminar on shooting technique and another on physical preparation for young players were left behind in a backup corpus. Nobody noticed until the full run produced the same refusal pattern across every question tied to those two sources.
That was our mistake. We checked whether the operation succeeded, but did not reconcile fragment counts source by source.
We restored the missing 115 fragments — 76 from one seminar and 39 from the other. The internal gold run moved to 27/31, or 87%. A corpus merge now passes only after its per-source fragment counts match the source manifest. A green operation result is no longer enough.
As with the first release, this is an independent public evaluation built from open materials. It is not a joint release with the RFB and does not imply federation endorsement.
Real coach questions changed the exam
Before the corpus expansion, the assistant’s certificate showed 75.2% on 60 questions over nine documents.
The first pilot user — a federation program lead — tested the assistant with real questions coaches had sent by email. Most were about everyday processes: how to register, what to do after forgetting a password, and which course fits an under-15 team.
Methodology books cannot answer those questions on their own. We added the official registration guides for the PRO system and Online Academy, the step-by-step admission guide, the course catalogue, a declaration, and application forms. With the restored seminars, the corpus now contains 32 public sources and 4,809 fragments.
The benchmark grew as well: from 60 to 70 questions, with a new module for registration and administrative processes.
| Measure | Before | Now |
|---|---|---|
| Questions | 60 | 70 |
| Sources | 9 | 32 |
| Composite | 75.2% (88/117) | 77.4% (106/137) |
| Attestation and regulations | 50.0% | 75.0% |
| Registration and administrative processes | — | 80.0% |
In the new module, the assistant grounded 10 answers out of 10 in federation sources. Across all answerable modules, it scored 58/60 on source matching. On questions where the corpus provides no basis and the assistant should stop, it made 10/10 correct refusals.
The move from 50.0% to 75.0% in attestation and regulations also has a plain explanation: the corpus now includes the admission steps and the required forms. The measurement identified a specific gap that had been closed. It did not prove that the assistant had become generally “smarter.”
Without federation documents: 0/10
We gave the same ten process questions to two control AI systems from the same model families, but without the federation corpus. Both scored 0/10 on source grounding.
They can produce plausible instructions. But registration steps, account recovery, and a federation’s course structure do not live in general model knowledge. Without the documents, the system has to guess — and fluent guesses turn into invented procedure very quickly.
For me, the important number in this run is not only 77.4%. It is the gap between 10/10 and 0/10 on questions that came from actual coaching work.
The 77.4% score describes one defined exam: 70 public questions against the current 32-source corpus. It is not a universal rating for the assistant, and it does not promise an answer to every coaching question. New emails and new failures should become new benchmark items instead of disappearing behind the top-line score.
One operating rule has already changed: every corpus merge now ends with a source-by-source fragment reconciliation. The full run was what proved that our “16 seminars” had in fact been fourteen.