For most of the last decade, the number of ad concepts a DTC brand could test was decided by how many it could physically make. A studio shoot ran into the thousands and produced a handful of usable cuts. Commissioning UGC cost somewhere between £150 and £500 an asset and took two to three weeks from brief to delivery, assuming the creator shipped on time and the product arrived. So the brand budgeted for four concepts a month, the account needed forty, and the gap got filled with recycled files and audience tinkering.
That constraint is gone. We now produce hundreds of new ad concepts a week for the brands we work with, without a studio, without a crew, and without a real creator on camera. Not variants of the same headline. Different angles, different settings, different openings, different objections answered.
The interesting part is what happened next. Removing the production bottleneck did not remove the bottleneck. It moved it somewhere most brands have never measured.
The rule that decides everything: 2x CPA per concept
A concept is not tested because it ran. It is tested when enough money has passed through it that the result stops being noise. The number we hold to is simple: every ad concept needs at least two times your current cost per acquisition behind it before you are allowed to have an opinion about it.
If your CPA is £40, a concept needs roughly £80 before you read it. Spend £30 and kill it, and you have not learned that the concept failed. You have learned that a concept which converts at your account average would be expected to produce well under one sale at that spend, and it produced none. That is not a signal. That is arithmetic doing what arithmetic does.
The same maths runs the other way and does more damage. A concept that gets lucky and lands two conversions on £30 looks like a hero. It gets scaled. It reverts to the mean three days later with real budget behind it, and now the account is down and nobody knows why. Most of the mystery losses I get called in to diagnose are winners that were never winners, promoted on samples too small to mean anything.
Two times CPA is a floor, not a target. It is the point at which you can start to distinguish a concept from randomness. On higher-consideration products with longer purchase cycles it needs to be more.
So how many concepts can you actually test?
Take your monthly testing budget and divide it by two times your CPA. That is your real creative ceiling, and it has nothing to do with how fast you can make things.
A brand spending £30,000 a month that ringfences 20% for testing has £6,000. At a £40 CPA, each concept needs £80. That is 75 concepts a month, or roughly 17 a week. A brand spending £100,000 a month with the same CPA and the same 20% allocation can test around 250 a month.
Which means we can now produce more concepts in a week than most brands can afford to test in a month. That sounds like a problem. It is the entire advantage, and only if you understand what it is actually for.
Volume is not for running. It is for choosing.
When you could only make four concepts a month, all four ran. There was no selection step, because there was nothing to select from. Every asset you made went live regardless of whether you believed in it, and your testing budget got spread across whatever you happened to produce.
Producing hundreds a week changes the job from making to filtering. You generate against every angle the reviews gave you, then you kill most of them before a penny of media touches them, and the seventeen that survive get a proper 2x CPA behind them instead of seventy getting a quarter of what they need each.
That is the actual mechanism. Not more ads. Better odds on the ads you fund, because they were selected from a much larger pool, and funded properly because you stopped spreading the budget thin.
The production line, end to end
The input is the angle, not the image. Nothing gets generated until we know what it is testing, and that comes out of review mining rather than a brainstorm. If the recurring objection for a supplement brand is that customers do not believe they will remember to take it daily, that objection is the brief. The visual is downstream of it. This is also why there is always a queue: the reviews produce more hypotheses than anyone can shoot.
The prompt carries the brand rules, not just the scene. Every generation call includes the palette, the lighting register, the setting, the demographic, and a banned list. The banned list matters more than the creative direction. Ours includes visible hands where possible, jewellery, on-screen text, and any attempt to render the actual product label. Those are the four things that fail most often, and excluding them at the prompt stage saves more time than fixing them afterwards.
Generation runs in parallel, in batches. Never one asset evaluated at a time. The marginal cost of the fourth generation is close to zero and the marginal value of having options is high. This is the step that turns tens into hundreds.
The product goes on afterwards. This is the step most people skip and it decides whether the ad looks like a brand or a hallucination. The generated asset is the environment, the model, the lighting and the mood. The real product is composited in from an actual product shot. The AI never renders the packaging.
A review gate stands between generation and the ad account. Every surviving asset is checked before it goes near a live campaign: does the product read correctly, are the hands right, is there any invented text, does the claim in the copy still match what the visual implies. Nothing reaches a live ad set on trust.
Where it genuinely falls over
Product fidelity is the hard limit. The model will redesign your packaging without telling you. It will move the logo, invent a label, change the cap colour, and produce something that is recognisably almost your product. For a category where the pack is the brand asset, that is disqualifying on its own. The compositing step is not optional, it is the price of working this way at all.
Hands, jewellery and text fail repeatedly. The jewellery problem comes up often enough that it now lives permanently in the banned list, and the workaround, keeping hands out of frame or in pockets, is a real constraint on what you can shoot. If your concept depends on someone demonstrating the product with their hands, this is not the tool for that shot.
It does not close high-consideration sales on its own. AI-generated creative lifts click-through by around 12% on Meta, and drops conversion by roughly 8% on higher-consideration purchases above a £100 order value. That is not a reason to avoid it. It is a reason to use it where it wins, which is at the top of the funnel and in the testing layer, and to put real creator footage and real product photography where the buyer has to trust you.
And volume without selection is worse than no volume. If you produce hundreds a week and run them all on a testing budget that supports seventeen, you have not scaled your creative programme. You have starved every concept in it and guaranteed you will read noise on all of them. The 2x CPA floor is what stops the volume becoming a liability.
What this looks like in practice
Work out your number before you change anything else. Monthly testing budget, divided by two times your CPA. That single figure tells you how many concepts you are entitled to an opinion about this month, and for most brands it is far smaller than the number they are currently running.
Then fix the funding before you fix the volume. Running seventeen concepts properly funded beats running seventy underfunded, every time, and it costs the same. Most accounts can get a real lift from that change alone, before a single new asset is produced.
Only then does production volume pay. Pull the angles out of your reviews so every concept tests a hypothesis you can name, build the constraint layer that keeps the output looking like your brand, and keep a human review gate in front of the ad account. A hundred beautiful images testing nothing is worse than seventeen testing something, properly funded.
Inside the system
How we build this for brands
Creative generation is the last step, not the first. The system we deploy starts with a VOC engine that mines customer reviews and support messages into positioning, so the angles being tested come from the customer rather than from a brainstorm. That feeds a production layer where generation, compositing and the review gate run as one pipeline, and the output lands in the ad account already tagged against the hypothesis it is testing and the budget it needs to clear the 2x CPA floor.
Behind it sits the reporting layer: profit and cash-flow dashboards built from live Shopify and ad data, with an agent that surfaces weekly which concepts are carrying the account, which are quietly draining it, and which were killed before they had enough spend to be judged. Part of this runs live for portfolio brands today; the full system is what we deploy when we take a brand on.
Creative Audit
Find out how many of your ad concepts were ever actually tested
We will work out your 2x CPA floor, count how many concepts in your account ever cleared it, and show you what your testing budget is really buying you each month.
Book Your Creative AuditFrequently asked questions
How much budget does each ad concept need to be tested properly?
At least two times your current cost per acquisition. If your CPA is £40, a concept needs roughly £80 behind it before the result means anything. Below that you are almost certainly reading noise: one or two conversions landing by chance will make a bad concept look like a winner and a good one look dead. This is the number that decides your real testing capacity, not your production capacity.
How many ad concepts can a DTC brand actually test each month?
Divide your monthly testing budget by two times your CPA. A brand spending £30,000 a month that allocates 20% to testing has £6,000. At a £40 CPA, each concept needs £80, which gives 75 concepts a month, or roughly 17 a week. That is your ceiling. Producing more than that is only useful if it raises the quality of the ones you choose to put money behind.
How do you produce hundreds of ad concepts a week without a studio or creators?
The angles come out of customer reviews rather than a brainstorm, so there is always a queue of hypotheses waiting. Generation runs in parallel batches rather than one asset at a time. The real product is composited in from a genuine product shot rather than rendered. And a review gate sits between generation and the ad account. Once that pipeline exists, producing a concept is a briefing job measured in minutes rather than a production job measured in weeks.
Does AI-generated ad creative actually perform?
It performs well at the top of the funnel and less well at the bottom. Across large datasets AI-generated ads have shown around 12% higher click-through rate on Meta, while conversion rate has been reported to drop by about 8% on higher-consideration purchases above roughly a £100 order value. Use it where it wins, which is earning the click and filling the testing layer, and put real creator footage and real product photography where the buyer has to trust you.
Can AI creative replace real UGC creators?
No, and treating it as a replacement is the fastest way to burn the account. It replaces the twenty briefs a month you were never going to commission anyway. Real creators still win on genuine product interaction, on unscripted delivery, and on anything where the claim depends on the person being real. Use AI creative to find which angle earns attention cheaply, then commission a real creator against only the angles that already proved out.
What are the real limitations of AI ad creative production?
Three that cost time every week. Product fidelity is the big one: the model will quietly redesign your packaging, move your logo, or invent a label, so anything showing the actual product needs a compositing step. Hands, jewellery and text-on-screen fail more often than anything else, and prompt fixes only partly solve it. And generation is non-deterministic, so the same prompt gives a different result each run, which means a human review gate before anything reaches a live campaign.
About the author
Caner Veli built Liquiproof to global distribution across 3,000+ retailers, then exited. He now runs Purposeful Profits using a combination of operator strategy and AI-powered systems he has built and uses daily, having 10x'd monthly revenue in his own business in the last 90 days.
