A benchmark for paid social teams

71.1% of estimated human time can be saved using AI.

The rest still needs a human.

I've tested AI on 18 tasks across paid social. See which models I'd use for each job, what they cost, and where they still need a human.

Still in beta 605 scored AI runs · 18 creative tasks tracked · updated Sep 9, 2026
Explore the full benchmark ↗ Free access by email. Updates are optional.
See the step-by-step view ↓

Estimated time savings

benmark score over time

benmark estimated human time saved over time14 dated frontier points from Apr 30, 2026, shown through Sep 9, 2026. 5 points record no score change.0%25%50%75%100%Apr 30, 2026Jun 2Jul 5Aug 7Sep 9, 2026Initial measurement, Apr 30, 2026: 46.6%, initial scoreGemini 3.1 Flash-Lite · high, May 7, 2026: 52%, changed by +5.4 ptsGemini 3.1 Flash-Lite · highGemini 3.5 Flash · high, May 19, 2026: 55%, changed by +3 ptsGemini 3.5 Flash · high3 model launches, May 28, 2026: 55%, no score changeClaude Opus 4.8 · maxGemini 3 Pro Image · defaultSeedance 2 · defaultClaude Fable 5 · max, Jun 9, 2026: 58.1%, changed by +3.1 ptsClaude Fable 5 · maxGemini Omni Flash Preview · default, Jun 30, 2026: 60.8%, changed by +2.7 ptsGemini Omni Flash Preview · defaultSeedream 5 Pro · default, Jul 8, 2026: 61.1%, changed by +0.3 ptsSeedream 5 Pro · defaultGPT-5.6 Sol · xhigh, Jul 9, 2026: 61.1%, no score changeGPT-5.6 Sol · xhighGemini 3.6 Flash · high, Jul 21, 2026: 61.1%, no score changeGemini 3.6 Flash · highClaude Opus 5 · max, Jul 24, 2026: 61.1%, no score changeClaude Opus 5 · maxGemini 3.7 Flash · high, Aug 13, 2026: 61.9%, changed by +0.8 ptsGemini 3.7 Flash · highClaude Fable 5.1 · max, Sep 1, 2026: 70.1%, changed by +8.2 ptsClaude Fable 5.1 · max2 model launches, Sep 2, 2026: 71.1%, changed by +1 ptsGemini 3.8 Flash · highMeta Muse Spark 1.3GPT-6 Astra · xhigh, Sep 3, 2026: 71.1%, no score changeGPT-6 Astra · xhigh
Sep 3, 2026GPT-6 Astra · xhigh71.1% · no score change
Interactive benmark score history14 dated frontier points from Apr 30, 2026, shown through Sep 9, 2026. 5 points record no score change. Select any point for its model, date and score change.0%50%100%Apr 30, 2026Jul 5Sep 9, 2026Initial measurement, Apr 30, 2026: 46.6%, initial scoreGemini 3.1 Flash-Lite · high, May 7, 2026: 52%, changed by +5.4 ptsGemini 3.5 Flash · high, May 19, 2026: 55%, changed by +3 pts3 model launches, May 28, 2026: 55%, no score changeClaude Fable 5 · max, Jun 9, 2026: 58.1%, changed by +3.1 ptsGemini Omni Flash Preview · default, Jun 30, 2026: 60.8%, changed by +2.7 ptsSeedream 5 Pro · default, Jul 8, 2026: 61.1%, changed by +0.3 ptsGPT-5.6 Sol · xhigh, Jul 9, 2026: 61.1%, no score changeGemini 3.6 Flash · high, Jul 21, 2026: 61.1%, no score changeClaude Opus 5 · max, Jul 24, 2026: 61.1%, no score changeGemini 3.7 Flash · high, Aug 13, 2026: 61.9%, changed by +0.8 ptsClaude Fable 5.1 · max, Sep 1, 2026: 70.1%, changed by +8.2 pts2 model launches, Sep 2, 2026: 71.1%, changed by +1 ptsGPT-6 Astra · xhigh, Sep 3, 2026: 71.1%, no score change

Rebuilt from today's tests and time estimates, using the models available at each date.

The step-by-step view

How much of the work AI can take on.

Each percentage is the share of the estimated workload at that step that AI can take on from humans. Put simply: the amount of human time that can be saved.

0178%

Ideation

Coming up with strong concepts.

Based on 4 tasks
0289%

Planning

Choosing what deserves making.

Based on 1 task
0395%

Briefing

Turning the concept into a clear plan.

Based on 3 tasks
0448%

Production

Creating and reviewing the assets.

Based on 3 tasks
0572%

Launch

Preparing the work for market.

Based on 3 tasks
0676%

Analysis

Understanding what happened.

Based on 3 tasks
0728%

Insight

Carrying the learning forward.

Based on 1 task

The averages hide a bigger story. On some tasks the gap between frontier models is huge. Just because it scores the highest on coding benchmarks, it doesn't mean it is the best for performance marketing.

Get the data

Find the right AI for the job.

01

Which model should I use?

Compare models by quality and cost, with a winner for each task.

02

Where could I put it to work?

Explore the scores, estimated time savings and my take on all 18 tasks.

03

Does a better setup help?

See where extra guidance improves the result and where it adds little.

I wouldn't choose one model for every job. The breakdown shows you where each one earns its place.

Methodology

How we get to the number.

Every task follows a declared test and scoring method. The number is built from actual AI runs, human judgement and objective checks designed by a human expert.

  1. 01Map 18 tasks across the selected creative journey.
  2. 02Run each task through leading AI models two ways: with a plain prompt (the baseline) and with a purpose-built prompt and workflow (the scaffold).
  3. 03Score every output blind, with model names hidden: judgement-led work is assessed by a human; objective work is checked against human-designed answer keys and tests.
  4. 04Estimate the work still needed on each output: individual fixes, human checking and, where necessary, a complete rebuild. A high average cannot cancel out a serious failure.
  5. 05Compare that work with estimated manual task time. Missing repair evidence receives no saving. Combine the best tested model and prompt approach for each task; business approval gates remain separate.

The person behind benmark

Hi, I'm Ben.

I've spent 15+ years in paid social, built and sold one of the three largest independent media agencies on Meta in EMEA, and worked with 300+ businesses across more than $1bn in ad spend.

Then AI turned up and I got curious. A couple of years of testing later, that curiosity has become benmark. I set the tests, blind-score the creative work and keep asking: “Yes, but is it actually useful?” Apparently, this is what I do for fun.