Three separate faculties: appreciation, production, and taste, measured under a dual-judge protocol. No model is graded by its own family. Funniness is judged by humans, never by other models. 😂
We asked every model to explain why jokes are funny. Two AI judges score each explanation against the human-written list of things a good explanation must catch. Two AI families, so neither judge ever grades its own family. Sort by any underlined column; yellow chips mark the top three.
| # | Model | Comprehension ↓ | Uncertainty | Scored | Spotting bad jokes · why a joke fails | Spend |
|---|---|---|---|---|---|---|
| Loading scores… | ||||||
| Score = share of the required joke mechanics each explanation caught (0 / ½ / 1 per item). Red band = 95% confidence interval: where the true score plausibly sits. Wider red = fewer explanations judged so far. A model posts a score once 10+ explanations are judged; before that it shows pending. n=1 anchors ran once per item, so their bands stay wide. Small samples are floored at a binomial interval, so a lucky 3-for-3 cannot claim 100 with zero doubt. | ||||||
Both judges grade the same rows independently. Their agreement is the free validity signal: if two model families concord on explanations, the rubric (not one model's taste) is doing the work. Human calibration replaces this signal at C5.
Every model wrote jokes under the same 40 constrained premises. Anonymous pair below: vote once, identities reveal after. Win rates with Wilson intervals appear as ballots land; no Elo until volume justifies Bradley-Terry. The judge is you.
Loading…
…
Loading ballots…