Nothing compounds. Month twelve starts like month one.
Most marketing cannot test at its own volume, and what worked is never written where the next job will read it.
Most tests lose, and most teams cannot run them
Across 127,000 experiments, Optimizely found only 12% won on the primary metric. Most B2B sites do not have the traffic to reach a reliable answer at all, so teams either guess or run tests that can never conclude.
The ground moves under the measurement
Ask an AI engine the same question twice and the recommendation list rarely repeats, with agreement under one in a hundred in SparkToro's 2026 research. Any tool selling you a single AI rank is selling noise, and decisions built on it are guesses wearing a number.
Simulate, then prove
Ideas are generated as a set rather than one at a time. A panel of simulated buyers kills the weak ones, and is only ever allowed to kill, never to crown, because simulated preference is not purchase. One idea the panel rated poorly always goes to the real market anyway, so the simulator can be proven wrong.
Then it keeps what survived
Survivors run in the market against a measure agreed beforehand and a point at which they are stopped. What won moves a written belief, and that belief seeds the next round, which is what makes month twelve start from everything the first eleven learned.
Everything we have written on this
Questions
We get very little traffic. Can we test at all?
Not in the classic A/B sense, and pretending otherwise wastes months. We test on the things that happen often enough to read, such as replies, chats started and clicks, and use simulation to remove the weakest ideas before spending.
Are simulated buyers trustworthy?
Only for removing weak options, and we treat them as advisory until they have been checked against real outcomes. The first such check is scheduled and will be published whichever way it goes.
What stops the same idea being retried forever?
Every test is registered with an owner, a date to check it and a condition that ends it.