Get your free SEO audit today Call 91 060 30 90
Home / Blog / Online Advertising
Online Advertising

How many ad creatives should you test in a campaign? The answer nobody gives with an exact number

There's a very common and very wrong intuition in online advertising: that the more creative variants you throw into a campaign (more images, more videos, more copy), the faster you'll find the winner. In practice, the opposite happens. Every new variant you add to a campaign splits the budget and the data volume across more pieces, which means each one individually takes longer to accumulate enough impressions and clicks for the difference between "this performs better" and "this is statistical noise" to be reliable. Testing 12 creatives with a budget sized for 3 isn't running a more thorough test, it's running 4 worse tests at the same time.

There's a very concrete statistical basis for this that's worth understanding even if you're not an analytics person: for the difference in results between two variants to be significant (meaning it isn't just sampling coincidence), you need a minimum volume of data per variant. The more variants competing for that same budget, the less data each one gets, and the longer it takes before you can confidently say which one wins.

The number actually worth using

For most small businesses managing their own ad budget (between 300 and 3,000 euros a month on the platform), the range that works well in practice is 3 to 5 creatives per campaign or ad set. Below 3 there's not enough variation to learn anything useful (if you're only testing 2 very similar versions, any difference you see could just be chance). Above 6-7, for most small-business budgets, the spend gets spread so thin that no variant ever reaches a reliable data volume, and the campaign takes weeks to "learn" (in platform jargon, exit the learning phase) because the algorithm doesn't have enough signal per variant to optimise well either.

This isn't a rigid rule: a big brand spending 50,000 euros a month can afford to test 15 variants because each one still gets enough volume. The rule isn't "always use 4 creatives," it's "match the number of variants to the available budget, not the other way round." If your budget is 500 euros a month, 4 creatives is reasonable; if it's 5,000, you can afford 6-8 without diluting the learning too much.

What "statistical sample fatigue" actually means

This concept, widely cited and rarely explained, comes down to this: when you split a data sample (impressions, clicks, conversions) across too many groups, each individual group ends up with so little data that the differences you observe between them stop being reliable, even if they look big in percentage terms. If one creative has 3 conversions out of 200 clicks and another has 1 conversion out of 180 clicks, it looks like the first one "triples" the second, but at those low volumes that difference could easily be down to chance, not to the creative actually being better. You'd need a lot more data on both to state anything with confidence. The more variants you throw into competition for the same budget, the easier it is to fall into this trap: drawing conclusions from differences that are actually noise.

What to vary between creatives so the test actually tells you something

Generating random variants isn't enough; you need to vary specific elements so that, once the test ends, you know what to learn from it. The dimensions genuinely worth testing, in typical order of impact:

  • The opening hook (the first line of copy or the first 2-3 seconds of a video): by far the variable with the biggest impact on results, because it decides whether the person keeps paying attention or scrolls past.
  • The format (static image vs. video vs. carousel): different formats perform differently depending on the platform and the funnel stage.
  • The call to action (buy now, learn more, book an appointment, download free): small changes here sometimes move conversion more than you'd expect.
  • The social proof or central argument (price, testimonial, guarantee, scarcity): each one speaks to a different kind of hesitation the potential buyer has.

If you're going to test 4 creatives, make sure each one clearly and deliberately differs on at least one of these dimensions, not 4 slightly different photos of the same product with the same copy. That kind of test tells you nothing, because there's no clear hypothesis behind each variant.

How long to let it run before drawing conclusions

Beyond the number of variants, the other common mistake is checking results too early and killing whichever one "seems" to be losing on day two or three. Most platforms need at least 3-7 days of stable data (no changes to budget or to the creative itself) before comparisons between variants start being reliable, and in low-budget campaigns it can take even longer to accumulate the minimum volume. The right discipline is to decide in advance how long you'll wait and how many minimum conversions you want to see per variant before deciding, instead of checking the dashboard every morning and acting on impulse over numbers that still aren't reliable.

How to set up a genuinely clean test

A well-run creative test isn't just about uploading several variants and waiting; it needs some structure for the comparison to mean anything. First, keep everything else fixed while the creatives vary: the same starting budget per variant, the same audience, the same time period, the same landing page. If you change targeting mid-test or shift budget from one variant to another as the first results trickle in, you're introducing a new variable that contaminates the comparison and makes it impossible to know whether the final difference is down to the creative or to the change you made halfway through.

Second, decide your decision criteria in advance: not just which metric you'll look at (cost per conversion, click-through rate, cost per thousand impressions), but what minimum difference between variants you'll treat as significant before acting. If one creative has a cost per conversion of 12 euros and another has 13.50, that difference on a small sample may mean nothing; deciding in advance that you need, say, a sustained 20-25% gap over several days prevents you from making decisions based on the ad system's normal fluctuations.

Finally, it's worth documenting every test, even in a simple spreadsheet: what you varied, what result it produced, what hypothesis you confirmed or ruled out. Without that record, every new campaign starts from zero and repeats tests you already ran months earlier, without ever accumulating real learning about what works with your specific audience.

Frequently asked questions

Is it better to test many cheap creatives or few well-crafted ones?

Few, but with a clear, deliberate difference between them. A creative that's cheap to produce but has a genuinely different, well-thought-out hook tells you more than ten minor variations on the same idea backed by heavy production.

What if no creative clearly stands out from the rest?

That's a valid signal in itself: the problem is probably not the creative but the targeting, the price, or the offer. Before generating more creative variants, check whether the audience you're showing the ad to is actually the right one.

Should I turn off underperforming creatives as soon as I see it?

Not immediately. Wait until you have a minimum volume of data per variant (at least a few hundred impressions and, if possible, several conversions) before pausing anything; turning things off too soon is the most common way to accidentally kill a creative that would have performed well given more time.

More on Online Advertising

Shall we talk about online advertising for your business?

Tell us about your project and we'll tell you how we can help, no strings attached.

Call 91 060 30 90