Estimated time savings
benmark score over time
Rebuilt from today's tests and time estimates, using the models available at each date.
A benchmark for paid social teams
The rest still needs a human.
I've tested AI on 18 tasks across paid social. See which models I'd use for each job, what they cost, and where they still need a human.
Estimated time savings
Rebuilt from today's tests and time estimates, using the models available at each date.
The step-by-step view
Each percentage is the share of the estimated workload at that step that AI can take on from humans. Put simply: the amount of human time that can be saved.
Coming up with strong concepts.
Based on 4 tasksChoosing what deserves making.
Based on 1 taskTurning the concept into a clear plan.
Based on 3 tasksCreating and reviewing the assets.
Based on 3 tasksPreparing the work for market.
Based on 3 tasksUnderstanding what happened.
Based on 3 tasksCarrying the learning forward.
Based on 1 taskThe averages hide a bigger story. On some tasks the gap between frontier models is huge. Just because it scores the highest on coding benchmarks, it doesn't mean it is the best for performance marketing.
Get the data
Compare models by quality and cost, with a winner for each task.
Explore the scores, estimated time savings and my take on all 18 tasks.
See where extra guidance improves the result and where it adds little.
Methodology
Every task follows a declared test and scoring method. The number is built from actual AI runs, human judgement and objective checks designed by a human expert.
The person behind benmark
I've spent 15+ years in paid social, built and sold one of the three largest independent media agencies on Meta in EMEA, and worked with 300+ businesses across more than $1bn in ad spend.
Then AI turned up and I got curious. A couple of years of testing later, that curiosity has become benmark. I set the tests, blind-score the creative work and keep asking: “Yes, but is it actually useful?” Apparently, this is what I do for fun.