A leaderboard scored against humans, not against another model.
Accuracy is claimed only on rows a person judged. No anchor model, no "agreement with GPT" dressed up as accuracy. The referee is published with the ranking.
| # | Model | Accuracy vs human labels | Cost / 1K | |
|---|---|---|---|---|
| 1 | Claude Fable 5 | 83.5% | $128.90 | |
| 2 | Gemini 3.5 Flash | 75.2% | $22.17 | |
| 3 | Kimi K3open | 73.5% | $45.14 | |
| 4 | GPT-5.6 Sol | 72.9% | $35.15 | |
| 5 | Claude Sonnet 5 | 71.4% | $38.38 | |
| 6 | Claude Opus 4.8 | 69.9% | $76.34 | |
| 7 | Gemini 3.1 Pro | 68.4% | $27.02 | |
| 8 | GLM-5.2open | 65.4% | $12.88 | |
| 9 | GPT-5.6 Luna | 60.2% | $7.43 | |
| 10 | Claude Haiku 4.5 | 58.6% | $17.95 | |
| 11 | GPT-5.6 Terra | 57.1% | $17.51 |