An ad set test only means something when the two sides differ in one thing, get the same money, each collect enough results to measure, and differ by more than chance could explain; when that bar is not met, the honest answer is a tie, and the original stays.
Every agency tests. Duplicate the ad set, change something, let it run, keep whichever did better. It feels rigorous. Most of the time it is a coin toss with a spreadsheet attached.
Not because testing is wrong, but because the four things that make a test mean something are the four things that get skipped when a client is waiting for an answer.
1. Change one thing
Copy the ad set exactly: same audience, same placements, same optimisation, same ads. Then change one thing. A different creative, or a narrower age range, or different interests, or fewer placements. One.
Change the creative and the audience together and a win tells you nothing you can reuse. You will not know which half did it, and next month you will copy the wrong half.
2. Give both sides the same money
Split the original’s daily budget evenly between the two, so the account spends what it spent before and each side gets an equal chance.
This is where campaign-level budgets quietly ruin tests. With Advantage campaign budget, Meta decides how much each ad set gets, and it moves money towards whichever looks better early. The side that gets starved never had a chance to show what it could do. A test needs budgets set on the ad sets.
3. Make sure each side can actually get results
This is the one people skip, and it decides everything else. Before starting, work out what each side can expect in a week. Half a ₹1,000 daily budget is ₹500 a day, ₹3,500 a week. At ₹150 a lead that is about 23 leads each side. Workable.
Half of ₹400 at ₹300 a lead is about five leads each side in a week. Nothing you see at that size is a finding. Raise the budget first, or do not run the test.
The reason is how noisy small counts are. A cost of ₹400 per lead measured over 8 leads could plausibly be anywhere from about ₹220 to about ₹800. Two sides at 20 and 22 leads on the same spend are not different; they are the same campaign having a slightly different week.
4. Only call a winner when it is clear
Two conditions, both required. First, the difference has to survive the uncertainty: the better side’s worst plausible cost should still beat the other side’s best plausible cost. Second, it should be big enough to matter, say at least 15% cheaper per result. A statistically real 3% is not worth restarting anything for.
Give it at least a week before judging, because days of the week behave differently and each new ad set spends its first days in Meta’s learning phase. If it is still unclear, let it run up to two weeks.
If it is still unclear after that, call it a tie and keep the original. A change has to earn its place. “Slightly better, probably” is how accounts drift into worse setups one reasonable-looking decision at a time.
A tie is a result. It means the thing you changed does not matter much, which is worth knowing.
Then act on it, and write it down
Pause the loser, give the winner the whole budget back, and write down what you tested and what happened. The note is the part that compounds: in six months, “we tested narrower ages on this client twice and it never won” is worth more than any single result.
One more trap: judge the test on the result that matters. If you are testing something meant to improve lead quality, like Meta’s Conversion leads goal, cost per lead will look worse by design. Judge it on cost per good lead, or you will kill the better campaign for being honest about its price.