Matt Mayer Harness Benchmark: Same Opus, 16-Point Swing Between Claude Code and Cursor
The Matt Mayer benchmark held the model, PRD, and rubric constant while swapping the harness, finding Opus scored 77% in Claude Code versus 93% in Cursor — a 16-point swing attributable to the harness alone.
The Matt Mayer benchmark is the most-cited datapoint for the claim that AI coding harnesses dominate coding performance. Mayer ran the same Opus model against the same 100-feature PRD (product requirements document) with the same evaluation rubric, varying only the harness. Opus scored 77% when driven by Claude Code and 93% when driven by Cursor — a 16-point swing caused entirely by the harness around the same weights. Multiple independent comparisons since have shown a 5 to 40 point spread from harness quality across different model and harness pairings. The size of the gap depends on which model you run and how invested the harness vendor is in per-model tuning. The usual explanation is staffing and incentives. Cursor employs engineers whose full-time job is iterating system prompts and tool descriptions for each model release until the harness extracts good behavior. Anthropic's Claude Code only runs Claude-family models, so its tuning targets only one family — and arguably gets less per-release polish than a third-party harness that has to compete on raw capability across providers. Theo Browne (t3.gg) overstated the gap in his April 2026 video by claiming Anthropic "hasn't changed these lines since it was knitted." The Piebald-AI/claude-code-system-prompts GitHub repo tracks frequent updates to Claude Code's system prompt, so the absolute claim is wrong even though the relative tuning-effort gap versus Cursor is real. See AI Coding Harness: Tools, System Prompt, Permissions Around the Model for what a harness actually is.