Change control & safety

Write an advertising experiment hypothesis that can change a decision

Try a different ad is an activity, not a hypothesis. A useful experiment card states the decision you will make, why a change might influence a particular outcome and what observations would support or weaken that explanation. The card belongs before implementation. It does not replace an experimental design, authorize spending or promise that the platform can run the exact comparison you have imagined.

A hypothesis card is a decision contract

Write a decision sentence first: choose between two clearly defined ways of acquiring the same business outcome. Then name the proposed mechanism. For example, showing a service restriction earlier may discourage ineligible enquiries while preserving qualified demand. This is more specific than expecting better performance, because it predicts which intermediate signal should change and which business outcome must remain acceptable.

Educational example, not a live test: a qualification-message hypothesis compares two otherwise comparable variants. Each illustrative variant spends 500 units. A receives 25 raw enquiries and 10 qualified enquiries; B receives 20 raw and 12 qualified. Raw CPL worsens from 20 to 25, while qualified CPL improves from 50 to 41.67. A card that preselected qualified CPL asks a different question from one that chooses raw CPL after seeing the results. The example demonstrates metric conflict, not a statistically proven winner.

Include primary outcome, diagnostic outcomes, unit of comparison, exposure rule, observation window and risk limits. List plausible alternative explanations: inconsistent delivery, different weekdays, delayed qualification or uneven sales handling. A hypothesis is not strengthened by ignoring these possibilities. Google offers campaign experiments, but the design must fit the supported configuration and actual business question; a pair of unrelated campaigns is not automatically a randomized comparison.

Make the mechanism falsifiable

  • A predicted mechanism connects the treatment to observable changes. If the expected intermediate change is absent, a favourable headline metric may need another explanation.
  • A preselected business metric protects against switching goals after seeing the data. Diagnostics remain useful, but their role is interpretation rather than selecting a convenient winner.
  • An implementation owner, measurement owner and decision owner make the card actionable. The same person can hold several roles, provided responsibilities are explicit.

Complete the card before implementation

  1. Write the precise decision and the smallest change needed to test it. Avoid changing audience, offer and landing page simultaneously when the question concerns message qualification alone.
  2. State the mechanism and one primary business outcome. Define eligibility, deduplication, maturity and the minimum economically meaningful improvement for that outcome.
  3. Choose a comparison design and identify the assignment unit. Document how exposure, contamination and unequal treatment will be detected; get methodological help when inference matters.
  4. Set the planned evaluation date, data-quality checks and separate emergency-stop conditions. Determine budget exposure through an authorized planning process, not through the article's example.
  5. Save the card with version and owner before implementation. AdAce Ads can supply accessible stored evidence for a read-only planning discussion; keep any campaign modification outside that analysis request.

A tidy card is not a valid experiment by itself

  • A hypothesis card cannot guarantee enough observations or remove allocation bias. The comparison design and data process still determine what can be learned.
  • A commercial decision threshold is not a statistical significance threshold. Both may matter, but they answer different questions.
  • If implementation changes mid-test, record the change and reconsider the interpretation. Do not preserve the original confident claim while quietly redefining what was tested. Retain failed implementation checks too: knowing that the planned treatment was never delivered can explain why an apparently neutral result does not actually test the mechanism.

Sources and further reading

Ace, the AdAce Ads mascot

Try it on your own accounts

Create a workspace, connect Google or Meta in a couple of clicks and see your accounts clearly. Changes follow your approvals or the policy you configure.

Create your workspace