Why channel tests stall after month three
Most paid channel tests do not end with a decision. They end when the budget runs out, the attribution model changes, or everyone quietly stops looking at the dashboard.
Most paid channel tests do not end with a decision. They end when the budget runs out, the attribution model changes, or everyone quietly stops looking at the dashboard.

Channel tests stall for four repeatable reasons. They start without a decision rule, so nobody knows what result would end them. Budgets are set too small to reach a result that is actually distinguishable from noise, so the team reads random variation as a trend. Attribution definitions change partway through, breaking the comparison between week one and week ten. And when a channel underperforms, there is no named owner for reporting that result, so it goes quiet instead of getting killed. A test that survives past month three has a written hypothesis, a decision rule set before launch, a budget floor sized to reach a real result, and a fixed measurement window. When the channel loses, the same person who proposed it reports why and what happens to the budget next, on a set date, whether or not the news is good.
A channel test rarely dies in a meeting where someone decides to kill it. It dies by attrition: the dashboard gets checked less often, the budget quietly rolls back into a channel everyone already trusts, and three months later nobody can say whether the test worked because nobody agreed in advance on what “worked” meant.
A channel test is a bounded trial of a new acquisition channel, run against a stated hypothesis and a fixed budget, designed to produce a keep-or-kill decision by a set date.
Four causes show up again and again in growth teams running these tests, in SaaS and in fintech alike. They compound: a test with no decision rule is also usually a test with no budget floor, because nobody sized the spend against a target the rule would need.
Most tests start with a hypothesis of a sort: try a new channel and see how it goes. That is a direction, not a decision rule. A decision rule states, before the first dollar is spent, what result ends the test and what action follows each outcome. Without it, a result at the edge of acceptable gets argued over for another month, and a genuinely bad result gets reframed as still early.
Writing the rule before launch is uncomfortable because it forces a real commitment: stop the test if cost per qualified lead exceeds a set figure once a set sample size is reached. That statement is falsifiable. Seeing how a channel performs is not, and a team cannot lose an argument it never framed as a bet.
A small test budget against a channel with an expensive qualified lead produces a small sample, and a small sample is not enough to tell a genuinely good channel from a genuinely mediocre one. The random swing in a small sample’s average cost per lead can easily run 30 or 40 percent in either direction, and a team reading that swing as a trend will make the wrong call in both directions over time.
This is not a fintech-specific rule, but fintech teams should note that campaign creative and claims tested under a small, fast-moving budget are still subject to the same compliance review as a full campaign. A test is not exempt from sign-off because it is small.
Evan Miller’s widely cited essay “How Not To Run an A/B Test” makes the underlying point directly: significance calculations assume the sample size was fixed in advance, and if a team instead runs a test until it sees a difference, the reported significance becomes meaningless. The same logic applies to channel tests read on cost per lead or cost per conversion. Set the budget floor from the sample size the decision rule needs, then treat that floor as a minimum commitment, not a ceiling to negotiate down when the quarter gets tight.
A test that starts under one attribution model and finishes under another is not one test, it is two tests stitched together and read as one. This happens more often from small changes than large ones: a marketing team adjusts what counts as a qualified lead, a new tracking parameter changes how a channel’s conversions get attributed, or a CRM update changes when a lead is marked as sales-qualified.
A test comparing paid social against paid search for a lending product ran for ten weeks. In week six, the sales team tightened the definition of a qualified lead to require a completed application rather than a submitted form. Cost per qualified lead on both channels jumped, and the team read it as both channels underperforming, when the real change was the yardstick. This pattern is common enough to describe in general terms. It is not a specific client case.
The fix is procedural. Freeze the conversion definition, the attribution window and the tracking setup for the duration of the test, and if a change is unavoidable, restart the measurement window rather than splicing the before and after together.
Longer sales cycles make this worse, not better. A fintech product with underwriting or a multi-step approval can take weeks between the first click and a funded account, so a test measured only on early-funnel signups will report a result before the channel’s true conversion has had time to show up. Decide upfront whether the test is being read on a leading metric, such as qualified applications, or a lagging one, such as funded accounts, and hold that choice for the whole window rather than switching to whichever metric looks better partway through.
A losing channel does not report itself. If nobody is named as responsible for writing up a negative result, the most common outcome is that the test simply stops being discussed, the budget drifts back to the familiar channel, and the same untested idea resurfaces in six months with nobody remembering it already failed once.
The fix costs nothing but discipline: whoever proposes the channel test also owns reporting its outcome, on the date the decision rule specified, to whoever approved the budget. That report is two paragraphs: what the decision rule said, what the actual result was, and what happens to the budget now. Ron Kohavi, Alex Deng, Roger Longbotham and Ya Xu’s “Seven Rules of Thumb for Web Site Experimenters,” presented at the 2014 KDD conference, makes a related point about organizational discipline in testing: the value of a testing program depends as much on consistently acting on results, including negative ones, as on running the tests correctly in the first place.
| Element | What it should say | Set before launch |
|---|---|---|
| Hypothesis | The specific claim being tested, stated so it can be wrong | Yes |
| Decision rule | The exact metric, threshold and sample size that ends the test, and the action for each outcome | Yes |
| Budget floor | Spend sized to reach the sample the decision rule needs, treated as a floor, not a ceiling | Yes |
| Measurement window | A fixed start and end date, or a fixed sample size, whichever the decision rule is built around | Yes |
| Result owner | The person who reports the outcome on the set date, win or lose | Yes |
This table is the whole design. Nothing here requires new tooling. It requires writing five things down before the first dollar goes to the channel, and holding the team to the fifth line even when the answer is unwelcome.
Killing a channel is a specific action, not a vague deprioritization. The budget is reallocated on the date the decision rule specified, the channel is logged as tested with its actual result so it is not retested blind in six months, and the reason for the loss, cost, targeting, creative or a genuine mismatch with the audience, gets written down so the next test in a similar channel starts smarter.
This discipline sits inside the same approach we cover on the growth systems hub and in the Growth Lab practice: a hypothesis, an action, data and an interpretation that leads to the next decision, on a schedule, whether the news is good or not. If your channel tests keep losing momentum around month two or three, our marketing operations audit checklist and the ownership questions in the revenue operations manager role are a good next read. If you want a second set of eyes on a specific test, get in touch.
Set the window from the sample size the decision rule needs, not from a calendar habit like running it for a quarter. A test that needs several hundred conversions per arm to be readable should run until it reaches that number or a hard budget ceiling, whichever comes first.
A statement written before launch of what result ends the test and what happens next: for example, if cost per qualified lead is above a set figure once the target sample is reached, cut the channel; if it is at or below that figure, scale it. Without it, any result can be argued either way.
Checking significance repeatedly and stopping as soon as a result looks good inflates the false positive rate well above the stated significance level, a problem documented in Evan Miller's widely cited analysis and in Kohavi and colleagues' experimentation research. The fix is to decide the sample size in advance and look once.
The person who proposed the channel writes a short report: what was tested, what the result was against the decision rule, and what happens to the budget. That report goes to the reviewer who approved the test, on the date set when the test started, not on request.