Skip to content

Run marketing experiments

Every marketing plan runs on hunches. Should the launch post lead with the hours saved, or with the AI? Do people actually read feature lists? Is Tuesday morning really better than Thursday? Usually the most confident hunch wins, the post goes out, and nobody ever finds out whether it was right.

An experiment turns a hunch into an answer. You write the claim down, publish a version of the content for each way it could go, and choose the number that will decide it before you see any results. When enough data has come in, you get a verdict. The verdict sticks around as a learning: six months from now, an agent drafting for the same audience starts from the answer instead of the guess, and nobody has to wonder twice.

One thing to internalize early: a disproven claim counts as a win here. Finding out cheaply that fear-based hooks do not work on your audience is worth more than three months of publishing them because someone was sure they would. The dashboard scores it as progress, and so should you.

On the marketing home, open the experiments view. The strip at the top is the whole program at a glance:

  • Running, shown as “2 of 5”, because a project runs at most five experiments at a time by default. The cap forces prioritization; it can be adjusted per project between 1 and 10 in the marketing context.
  • Validated this month and Invalidated, both counted from concluded verdicts. Invalidated carries the caption “assumptions disproved, cheaply” and is styled as information, never as an error.
  • Learnings captured, the all-time count across the whole workspace, linking to the learnings browser.

Click any experiment for its detail page: the arms with their numbers, the stop rule’s progress, and after conclusion, the verdict.

PartWhat it means
HypothesisOne falsifiable sentence. “Founders respond better to outcome-focused messaging than AI buzzwords”, not “improve engagement”.
ArmsTwo or more content assets, exactly one marked as the control. The others are the variants competing against it.
Success metricThe one number that settles the claim, declared before any result exists.
Stop ruleHow much data is enough (a minimum number of observations per arm) and how long the experiment may run (one to thirty days).

Start stays disabled until the experiment is complete, and the form lists every missing piece at once rather than one at a time: a real hypothesis, at least two arms with one control, a metric, a stop rule, a free slot under the cap.

Every experiment declares how its metric arrives. Collected automatically by an integration is the default: the metric select then offers only what an integration binding can actually collect, and a metric with no binding behind it appears disabled with a note saying which category of integration would light it up.

I will record the numbers myself is the other option, and it needs no integration at all. Every engagement and demand metric becomes available, because you are the one reading them off the platform. See recording results by hand.

Google Analytics 4 (pageviews and sessions) and Mailchimp (email opens and clicks) ship in the integration catalogue ready to import, which covers landing-page and email experiments out of the box. X and Threads ship there too, for social engagement: both need a developer account of your own and X charges per post, so connect social platforms walks each one through step by step. Any other platform is registered by its API through the same page.

If the arms differ in channel or asset type, the form warns that the result will conflate the message with the medium. It is a warning rather than a block: comparing channels is a legitimate experiment, as long as you know that is the one you are running.

You can write one by hand with New experiment, but four other sources feed the queue:

  • Huddles. When a huddle states a comparative claim (“I think concrete demos beat feature lists for this audience”), the draft review that already proposes tasks also proposes an experiment, with the claim as its hypothesis and a checklist of what is still missing (usually the metric and the arms). Accept it and it lands in proposed, ready to be completed.
  • The weekly review. The same review that drafts marketing proposals can propose at most one experiment per week: a change it is not confident enough to recommend outright, reframed as something to test. Its expected outcome is recorded, and the verdict later reports whether the prediction was met.
  • Near misses. A metric that lands noticeably off its baseline, but not far enough to become evidence, is a question worth asking. At most one such candidate per project per week becomes a proposed experiment pointing at the observation that raised it.
  • Search demand. With Search Console connected, the weekly pass also reads the queries your site is shown for and proposes the strongest missed demand: a query you rank just off page one for (average position 11 to 20 with at least 50 impressions over the trailing 28 days), or one that is shown often but rarely clicked (200 or more impressions with a click-through rate below 1%). The proposal names the query, its numbers, and the page that currently ranks, and it picks the action to test: improve that page, or write something new when no page ranks. Candidates are ranked by impressions, share the same one-per-project-per-week cap, and the same query is not proposed twice within 30 days. Without Search Console this source is simply absent, and the experiments panel invites you to connect it.

Every source produces a proposal. Nothing starts running without a person completing the gate and pressing Start.

An experiment set to I will record the numbers myself needs no integration and no developer account. The loop runs like this:

  1. Approve the arm assets as usual.
  2. On each asset, Copy for posting puts the post text on your clipboard and shows a brand check beside it. The check informs and never blocks: you are posting from your own account, so the platform’s job is to make any risk visible, not to pretend it can stop an external paste.
  3. Post it yourself, then Mark as published and paste the post URL if the platform gives you one.
  4. Open the experiment and use Record results: one row per arm, one column per metric, one submission. Read both posts’ numbers on the platform and type them in together.

Hand-entered numbers are first-class. They feed baselines, evidence, and the verdict exactly like collected ones, arm by arm. What changes is that the result says where the numbers came from: a manual experiment’s verdict is labeled “measured from manually recorded numbers”, in the same neutral register as its directional or suggestive confidence label.

A running manual experiment with nothing recorded for three days raises one reminder task linking straight back to the entry form, and no more until you record something and another quiet stretch begins.

Articles get a different treatment from social posts, on purpose. Two posts saying the same thing two ways are a fair A/B test. Two articles are not: they cover different topics, and topic decides search demand by orders of magnitude, so any A/B verdict would be measuring the topics rather than the writing. The platform will not let you set two articles against each other as arms. What it offers instead are the two comparisons that are honest for content:

Before and after. You changed an article: a rewrite, a new title, a refresh proposed by decay detection. The comparison takes the change date and measures the same number of days after the change against the days before it, side by side, adjacent and never overlapping. The default is 28 days each way, and the result concludes on its own once the after window has fully elapsed.

Read a before-and-after result as evidence, not proof. The comparison sets an article against its own recent past, and an article’s traffic can move for reasons no comparison can remove: seasonality, a competitor publishing, a search algorithm update. Your change is one candidate explanation, not the established cause. The result page states this caveat right next to the numbers, every time, because a number without it would claim more than the data can.

Cohort. A set of articles that received a treatment (say, every post that got a custom cover image) measured against a set that did not, over the same calendar window for both. Never different windows: comparing this month’s cohort against last month’s would measure the season, not the treatment. Cohorts are compared on the metric per asset, so a five-article cohort and an eight-article cohort are on the same scale.

A cohort needs at least five assets on each side before any number is computed. Below that the result says “not enough assets to compare” and shows how many more each side needs, rather than reporting a figure with no power behind it. A small corpus is a reason to publish more, not a reason to trust a number that should not exist.

Comparison success metrics are the impression and click families (pageviews, impressions, clicks, search impressions, search clicks). Average search position is excluded here for the same reason it is excluded from experiment arms: lower is better there, and every comparison here reads higher as better.

Once every arm has enough observations, or time runs out, the daily evaluation concludes the experiment:

  • Validated: the best variant beat the control by at least 20 percent on the declared metric.
  • Invalidated: the variant lost by at least 20 percent. The assumption is disproved, which is the cheap version of finding out.
  • Inconclusive: the difference stayed inside that band, or the experiment hit its time limit without enough data. Also a result; it tells you the effect, if any, is small.

The verdict shows each arm’s observation count and the measured lift, labeled directional below thirty observations per arm and suggestive above. There are no p-values and nothing here is called statistically significant, because at typical volumes that would be theater. A verdict on nine observations says so.

Each verdict carries a default decision you can override with a note: validated suggests scale, invalidated suggests kill, inconclusive suggests iterate.

Conclusion drafts a learning: one plain sentence pairing the hypothesis with what the numbers showed, citing the arm observations as evidence. It arrives as a marketing proposal, so a person accepts every learning before it becomes part of the strategy. What happens after you accept it is the subject of the learnings page.

An experiment that should stop early can be abandoned with a reason. Abandoned experiments keep their record; nothing is deleted, because the point of the system is that the record survives.