When Will The Benchmaxxing Plague End? — Nick Heiner, Surge AI

AI Engineer17mAug 2, 2026
Watch Original (opens in new tab)
0:00 / 17:25
Chapters8

No clickbait detected — the title and thumbnail deliver what they promise.

AI Opinion

The episode convincingly argues that current AI benchmarks often incentivize superficial improvements rather than genuine progress toward useful models, citing specific examples like the manipulation of Elm Marina and the Opus model's data memorization. While observations from figures like Andre Karpathy and Wor lend weight to this critique, the claim that roughly 20% of benchmark tasks are "broken" feels somewhat overstated without more granular explanation of what constitutes a broken task. Listeners should consider whether Surge AI’s Hemingway Bench, despite its human evaluation focus, might introduce new biases or limitations inherent in relying on professional writer preferences as a sole measure of model quality.

Avatars are AI rewrites of the same facts — style changes, not substance.

Summary

The episode explores a recurring problem in the AI industry known as "benchmaxing," where labs prioritize achieving high scores on benchmarks over developing models with practical utility. Popular benchmarks like Elm Marina are criticized for being easily manipulated, leading to improvements in scores that don't reflect genuine advancements in AI capabilities. Creating reliable benchmarks is expensive and prone to contamination—as demonstrated by the Opus model memorizing data from SweetBench—and often suffers from flawed tasks, distorting model rankings. Surge AI has developed a benchmark called Hemingway Bench utilizing human evaluators, specifically professional writers, to address these shortcomings, though this approach incurs significant costs. The episode defines benchmaxing as exploiting misalignments between benchmarks and actual user preferences, advocating for higher standards in both benchmark creation and reporting to ensure more accurate assessments of model performance and move beyond superficial metrics.

Avatars are AI rewrites of the same facts — style changes, not substance.

Key Points

00:18

The Cycle of Hype and Disappointment

The tech industry, particularly AI, frequently experiences hype cycles where new models are announced with impressive benchmark results. However, these benchmarks often fail to accurately reflect real-world performance, leading to allegations of 'benchmaxing,' which is when labs train models excessively on benchmarks that deviate from practical applications. This cycle highlights a disconnect between theoretical scores and actual utility.

01:36

The Issue with Elm Marina Benchmarks

Despite widespread use, many popular AI benchmarks like Elm Marina are considered 'very bad' by industry leaders. Andre Karpathy observed that teams were improving their scores on Elm Marina rather than developing genuinely better models overall. This suggests a focus on gaming the benchmark system instead of achieving meaningful progress in AI capabilities.

03:19

The High Cost and Contamination Risks of Benchmark Creation

Creating high-quality benchmarks, such as agentic coding assessments, is incredibly expensive. A benchmark with 1,000 tasks can cost $15 million to create initially, and another $5 million annually to maintain due to models improving. This financial barrier often leads labs to cut corners, like using AI assistance or cheap labor, which compromises the benchmark's usefulness.

04:41

Opus Model Contamination with SweetBench Data

Contamination occurs when models memorize data from benchmarks. Surge AI conducted an investigation revealing that Opus had memorized a significant portion of the Sweepbench verified content, which was not disclosed in the model card. This lack of transparency hinders consumers' ability to accurately evaluate model performance and highlights a systemic problem within the industry.

15:14

The Issue of Broken Tasks in Benchmarking

Nick explains that approximately 20% of tasks used in benchmarks are often flawed or 'broken'. The challenge lies in identifying these problematic tasks because their impact isn't apparent until all other tasks have been addressed. This leads to significant noise and distortion within the model ranking process, skewing results.

15:41

Hemingway Bench: A Human-Driven Evaluation

Surge AI developed a benchmark called Hemingway bench to address shortcomings in existing writing assessments. Recognizing that mechanical benchmarks and LLM judges fail to capture the complexity of human writing, they established a workforce of thousands of professional writers—including technical writers, poets, journalists, and editors—to conduct blind model comparisons.

16:28

The Cost-Quality Tradeoff in Human Evaluation

Nick acknowledges that utilizing human evaluators for benchmark assessments is a costly endeavor. Securing the time and expertise of professional writers represents a significant expense compared to automated methods. However, Surge AI prioritizes maximizing quality over minimizing costs in their evaluation process.

16:37

Benchmaxing as Exploitation & Call for Higher Standards

Nick defines 'benchmaxing' as the exploitation of benchmark misalignments between human preferences and automated evaluations. He urges both those creating benchmarks and those reporting on them to uphold higher standards, emphasizing a need to move beyond flawed methodologies and prioritize genuine quality in model evaluation.

Chapters

8 chapters · 8 key moments
KEYkey momentNot checkable hereUnverified

Claims & Fact Check

Labs are training too hard on benchmarks in a way that deviates from what people actually care about.

Not checkable here

Thought leaders like Wor say Elm Marina can be easily gamed.

Not checkable here

Andre Karpathy observed that teams are getting better at Elm Marina, not necessarily better models overall.

Not checkable here

Opus has memorized a lot of Sweetbench content.

?Unverified

Approximately 20% of tasks used in benchmarks are broken.

?Unverified

LLMs don't have good taste in writing.

Not checkable here

Benchmaxing is the exploitation of benchmark misalignments between human preference.

Not checkable here

Was this digest good?

More from AI Engineer

Digest any single YouTube video — free.

3 free digests — no card, no sign-up wall.

Or just swap the domain of any YouTube link → instant digest