Stop choosing an advertising test winner from daily peeking
Opening a test report every morning is useful for detecting broken pages and collection failures. Declaring a winner whenever one variant briefly looks best is another practice. It creates repeated opportunities to select a favourable fluctuation. The solution is not to stop monitoring; it is to separate operational checks, exploratory observations and the planned evaluation that supports the business decision.
A dashboard check is not a decision rule
A rule designed for one final assessment should not be assumed valid when reused until it returns the desired answer. NIST discusses why repeated comparisons need an appropriate overall comparison procedure. Daily observations of one accumulating test are correlated, so they require their own design; a multiple-comparison method for another setting is not automatically a sequential-testing recipe. If formal interim decisions matter, specify a suitable sequential method before observing the outcomes.
Educational probability illustration, not a claim about your daily dashboard: imagine twenty independent decision opportunities, each with a 5% false-positive probability under its stated assumptions. The probability of at least one false positive is 1 − 0.95^20, approximately 64%. Daily looks at the same campaign are not independent, so 64% is not their actual error rate. The illustration only shows why repeatedly offering chances to declare a result changes the overall question.
Maintain separate records. An operational log records outages, overspending against authorized exposure and treatment inconsistencies. An exploratory log records interesting subgroup movements without treating them as confirmations. A decision log records the planned endpoint, mature primary outcome, comparison method and uncertainty. Finding a new hypothesis in exploration is useful; testing it on the same conveniently selected observations does not create independent confirmation.
Repeated opportunities change the question
- Preselect one primary business outcome and a review schedule. A result that disappears when the metric or date range changes should be described as unstable, not hidden.
- Keep emergency conditions active between evaluations. Waiting for a scheduled analysis does not require tolerating a broken customer journey or an unauthorized exposure.
- Label post hoc segments explicitly. A promising region discovered during a test can motivate another test, but it should not silently replace the original whole-population question.
Use an evaluation calendar
- Write the experiment's decision and measurement plan before results arrive. Include the primary outcome, assignment unit, maturity rule and intended interpretation method.
- Schedule a formal evaluation at a prespecified endpoint. If more than one decision look is necessary, agree on a method that accounts for those looks rather than repeatedly using an unadjusted final-test criterion.
- Assign daily checks to operational health: coverage, treatment delivery, destination availability and observed exposure. Make the resulting incident actions distinct from choosing a winner.
- Save exploratory notes with dates and clear labels. Do not extend, shorten or segment the test solely because a favourable number has appeared; document any justified protocol amendment.
- At evaluation, report all prespecified outcomes and relevant limitations, including an inconclusive result. AdAce Ads can provide available stored evidence for read-only review; request a textual assessment without propose_change or other tool writes.
Monitoring must remain possible
- A fixed end date alone does not guarantee a sound test. Allocation, sample size, measurement and the statistical assumptions still matter.
- The educational independence calculation is not an estimate of advertising false-positive frequency and should not be copied into a client report as such.
- An exploratory insight can be commercially valuable without being statistically confirmed. Keep its label and choose a proportionate next decision instead of claiming proof.
Sources and further reading



Try it on your own accounts
Create a workspace, connect Google or Meta in a couple of clicks and see your accounts clearly. Changes follow your approvals or the policy you configure.
Create your workspace