An apology to DeepSeek
We owe DeepSeek an apology. We published a number against their model that we could not substantiate. We should have been more careful before publishing it. We will learn from this and will do better in the future.
What we published
An ARC-AGI-2 result of 22.8% attributed to DeepSeek V4 Pro. It carried no source at all: no URL, no named evaluator, no record of where the number came from or when it was measured.
What was wrong with it
Two things, and the second is what makes the first serious.
It was unsourced. Our own published rule is that every cell on the board cites where it came from. This one did not, and it had been sitting on the site regardless.
It was identical to another model’s score, to three decimal places. GLM 5.2 holds exactly 0.228 on ARC-AGI-2, properly sourced to the ARC Prize leaderboard. Two different models holding the same value to three decimals, with one of them sourceless, is not a coincidence we are willing to explain away. It is the same duplication pattern we had found in an earlier audit of a secondary aggregator.
We cannot establish that DeepSeek V4 Pro was ever evaluated on ARC-AGI-2. That is the honest statement of our position: not that the number is wrong, but that we do not know what it is, and we published it anyway.
What it affected
No ranking changed, on this board or the previous one. The cell had already been excluded from scoring before we found it, so no published score moved when we voided it.
We are saying so plainly because it cuts both ways. It means the practical harm to DeepSeek was limited to a number displayed against their model that should never have been displayed. It also means this correction costs us nothing, and a correction that costs nothing is the easiest kind to stay quiet about. Staying quiet is precisely how a value with no source survives long enough to matter.
What we did
The cell is voided. The record is retained rather than deleted, as with every other correction we publish, so the mistake stays auditable instead of disappearing from the history.
We found it while auditing every cell on the board for methodology v2. Four further cells were voided in the same pass, for weaker provenance defects: one whose contest year had never been pinned down, two sourced to a vendor’s own model page rather than the evaluation board that produced the result, and one sourced to a commentary article about a different model entirely. None of the five score on the current board.
What we changed so it cannot happen again
An apology that does not change a process is just an apology.
Methodology v2 makes a source a condition of scoring rather than an expectation. A cell without a recorded origin cannot enter the board. All 156 scored cells on the current board carry a source you can open, and that is checked on every release rather than trusted.
The same audit enforced a provenance gate we had already published but were not applying strictly enough, which removed the last laboratory self-report from the board. Every scored cell now comes from an independent evaluator or a benchmark owner’s own leaderboard.
We would rather be the site that tells you it got something wrong than the site that never appears to.