We put our hooks up against ChatGPT, Gemini, and Claude. Blind. 600 times.

Every AI tool claims to be better. We wanted a number — a real one, with a confidence interval, not a vibe. Here's how we got it, and what it does and doesn't prove.

The question

ConvertHooks runs on the same class of AI you can use for free. So why would anyone use us? Our bet is that hooks fail for predictable reasons — they resolve their own tension, they run long, they overclaim, they could open anyone's post — and that a tool built to catch those failures beats a blank chat box. That's a testable claim. So we tested it, twice: a small manual pass first, then an automated version built to remove every source of bias the manual version couldn't control for.

The setup

We built a harness that runs the whole test end to end. It draws from a corpus of 500 passages spanning 10 genres — news, memoir, children's books, product specs, B2B copy, health, finance, food, travel, opinion — and sampled 100 of them, evenly balanced, for this run. If a tool only writes good hooks for one kind of content, ten genres will expose it.

Each passage went to four sources with identical input, a plain "write 5 hooks" request — no taxonomy, no gates, no coaching: ConvertHooks (our production pipeline), ChatGPT (latest model), Gemini, and Claude Opus 4.8 — Anthropic's most powerful model, run raw. That's 2,000 hooks in total.

Every set of four was anonymized, randomly relabeled per passage, and scored blind by three AI judges — ChatGPT, Gemini, and Claude Opus 4.8 — on curiosity, specificity, credibility, cadence, and stopping power, with a forced 1st-through-4th ranking and no ties. Each judge scored every passage, including the ones where its own raw entry was in the field — no exclusions, no carve-outs. Then each matchup was judged a second time with the label order reversed, specifically to cancel out any position bias a judge might carry. Three judges × two orders × 100 passages = 600 independent blind judgments.

The results

ConvertHooks won 558 of 600 blind judgments — a 93.0% win rate, 95% confidence interval 90.7%–94.8%. It placed first in every one of the 10 genres tested, never dropping below 78%, and cleared 90% in nine of them. Final standings by average rank: ConvertHooks first, raw Claude Opus 4.8 a distant second (5.7%), Gemini third (1.0%), ChatGPT last, winning just 2 of the 600.

This scaled result confirms what a smaller manual test found first: in an earlier 5-passage, 10-round pass judged by hand, ConvertHooks placed first 9 times out of 10. The automated version repeats that finding at 60× the sample size, with bias controls a manual test can't easily apply — and with every judge grading its own raw output alongside everyone else's, blind, exactly like the others.

Why it matters

The most useful finding isn't the score — it's who lost. Second place went to raw Claude Opus 4.8, a strictly more powerful model than the one ConvertHooks runs. A bigger engine lost to a smaller engine with better instructions. That's the whole thesis of this tool: the advantage isn't the AI, it's what we've built around it — a technique taxonomy distilled from a corpus of 233 real, high-performing hooks, six quality gates every line must pass (open the loop, stay tight, one metaphor, no overclaims, ground it in your text, read clean aloud), and a curation step that drafts roughly twelve candidates and ships only the five survivors.

You can't get that from a blank prompt box, no matter which model is behind it.

What this doesn't prove

Honesty makes the number worth something, so: the judges were AI models, not human audiences — the only verdict that finally matters is your retention graph. And a test like this measures the version of each tool on that day; models change. Our commitment is simple: we re-run the panel whenever we change how ConvertHooks generates, and we publish the new number — whatever it says.

Test conducted July 2026. 100 passages, 10 genres, 600 blind judgments by ChatGPT, Gemini, and Claude Opus 4.8 on anonymized, order-randomized sets.

[Ad space]