Independent benchmarks · run by Zevenue

We benchmark GTM tasks.

Models now qualify accounts, research prospects and answer replies. Almost nobody measures whether they're right.

The GTM Open scores frontier models on real go-to-market work. Published methodology, human-judged labels, every model on the same frozen evidence, re-run as the models change.

Models on the board
11

9 frontier, 2 open-weight, one rubric

Companies judged
243

funded, AI-branded, frozen evidence packs

Human-judged rows
133

the only rows accuracy is ever claimed on

Top accuracy
83.5%

against human judgment, benchmark 01

Most GTM benchmarks are published by someone with a stake in the result.

Vendors grade their own tools. Directories rank what they resell. Evals get scored by whichever model the author already liked.

The GTM Open has no product on the board and sells no placement on it.

How a benchmark here works

Same task, same bytes, scored against people.

Every edition follows the same four rules. They are the part of the work that makes a number worth publishing.

A leaderboard scored against humans, not against another model.

Accuracy is claimed only on rows a person judged. No anchor model, no "agreement with GPT" dressed up as accuracy. The referee is published with the ranking.

#ModelAccuracy vs human labelsCost / 1K
1Claude Fable 583.5%$128.90
2Gemini 3.5 Flash75.2%$22.17
3Kimi K3open73.5%$45.14
4GPT-5.6 Sol72.9%$35.15
5Claude Sonnet 571.4%$38.38
6Claude Opus 4.869.9%$76.34
7Gemini 3.1 Pro68.4%$27.02
8GLM-5.2open65.4%$12.88
9GPT-5.6 Luna60.2%$7.43
10Claude Haiku 4.558.6%$17.95
11GPT-5.6 Terra57.1%$17.51

Frozen evidence.

Every model sees identical bytes. No live browsing, no "one model got a better search result." The part most public evals skip.

pack / example-ai.com5 files
homepage.htmlsha256 3f9a…c21e
product.htmlsha256 8b02…77d4
careers.htmlsha256 c1e7…0a93
docs.htmlsha256 51d0…e6bf
engineering.htmlsha256 9a44…12c8
served byte-identical to 11 models

Re-run when the models change.

A new checkpoint gets its row on the same frozen packs the week it ships. One version bump moved more than any gap between labs.

wedKimi K3 released
friK3 swept on the frozen packs, next to its predecessor
K2.656.8%
K373.5%
+16.7points, paired p = 0.001. Last place to third.

Open methodology, disclosed labels.

Test-set construction, label provenance and every bug we found in our own labels are in the write-up. Rows a model panel adjudicated are never scored for accuracy.

133 human-judged (accuracy lives here) 110 panel-adjudicated (label-free analysis only)
What we grade

Four GTM tasks, in the order a pipeline runs them.

One task per edition. Each one is work a GTM team already routes through a model today.

01Published

Qualification

Given a company's own pages, is the claim real? Twelve models judged 243 AI-branded companies on one four-tier rubric.

read benchmark 01 →
02In development

Research

Profile an account: firmographics, stack, buying signals. How accurate is the profile, and where does a model confidently invent things?

03Planned

Inbox

Handle a reply. Scored on what a good operator would have done next, not on whether the answer read well.

04Planned

Outbound

Write the first touch for a real account. Judged against human-rated sends, with the rubric published before the run.

Benchmark 01·Qualification·July 2026·15 min

Which frontier models can tell if an AI company actually ships?

Twelve frontier and open-weight models, one rubric, frozen evidence from 243 funded AI companies. The most expensive model topped the board, then price stopped buying accuracy at about $22 per thousand accounts.

One version bump moved an open-weight model from last place to third. The best judge disagreed with itself on 14% of companies. And 96 companies got flagged as AI-washing by every model on the panel.

50%60%70%80%90%$7$15$30$60$120$250SIX MODELS, STATISTICALLY TIEDFable 5FlashKimi K3SolSonnet 5Opus 4.8Gemini ProGLM 5.2LunaHaiku 4.5TerraAnthropicOpenAIGoogleopen-weight
accuracy vs human labels, n=133cost per 1,000 accounts, log scale
How the Open stays open

Six rules, none of them negotiable.

Independence is the whole asset. These are the terms every edition and every commissioned test runs under.

01

Methodology is published.

Test-set construction, rubric, label provenance and known bugs ship with every result.

02

No pay for placement.

Nobody buys a rank, a preview or an edit. Commissioned tests are labeled as commissioned.

03

Publish or silent.

A vendor that commissions a test can keep it private. It can never change the result.

04

Relationships are disclosed.

Every report names the vendors we use, resell or are sponsored by. No affiliate links in benchmark content.

05

Sponsors sit outside their category.

A sponsor cannot support an edition that ranks it.

06

Client data stays out.

Test sets are built fresh from public or purchasable data, never from a client's lists.

Who it's for

Test before you build on it.

Buyers

Pick the model on your task.

Choosing which model runs your qualification, research or inbox workflow? Bring the task. We run it through the candidates and hand you the data before you commit.

Vendors

Enter the Open.

Any vendor can ask to be tested. The methodology is the same whether you asked or not, and the independence rules above apply either way.

Researchers

Argue with the method.

Every edition names its referee, its label provenance and the fix it still owes. If you think a number is wrong, the write-up tells you where to look.

Enter the GTM Open.

Vendors can ask to be tested. Buyers can bring their own task. Either way the rules are the same.