LOLBENCH dataset v0.1.0
Live evidence · wave 0

Do large language models get the joke?

Three separate faculties: appreciation, production, and taste, measured under a dual-judge protocol. No model is graded by its own family. Funniness is judged by humans, never by other models. 😂

Preliminary · auto-judged, uncalibrated claims ship when earned; visibility ships now
Models scored
pending
Judge agreement
pending
Ballots cast
pending

We asked every model to explain why jokes are funny. Two AI judges score each explanation against the human-written list of things a good explanation must catch. Two AI families, so neither judge ever grades its own family. Sort by any underlined column; yellow chips mark the top three.

# Model Comprehension Uncertainty Scored Spotting bad jokes · why a joke fails Spend
Loading scores…
Score = share of the required joke mechanics each explanation caught (0 / ½ / 1 per item). Red band = 95% confidence interval: where the true score plausibly sits. Wider red = fewer explanations judged so far. A model posts a score once 10+ explanations are judged; before that it shows pending. n=1 anchors ran once per item, so their bands stay wide. Small samples are floored at a binomial interval, so a lucky 3-for-3 cannot claim 100 with zero doubt.

Why trust a number here

protocol
One model grading its own familyDual judges, family-disjoint; self-rows dropped at scoring
Judges “rating funniness”They only match explanations to human gold elements
Memorized benchmarksCanaries, private split, fresh premises each wave
OverclaimingCIs on everything; preliminary banner until human calibration

Judge validity

do the two judges agree?

Both judges grade the same rows independently. Their agreement is the free validity signal: if two model families concord on explanations, the rubric (not one model's taste) is doing the work. Human calibration replaces this signal at C5.

Exact agreementpending
Cohen's kappapending
Rows both judgedpending
LOL-C · taste is the third faculty: does a model's own sense of humor match ours? Pools ship with dataset v0.2; correlations land here with error bars attached.