Forthcoming. This describes where the Scale practice is going rather than what runs today. We are building toward it deliberately, and we would rather show you the destination than imply we are already there.
Every period we run a small number of deliberate experiments, and every one of them is declared in writing before it starts.
The declaration names what we believe will happen, which piece of work is the test, what we will measure, the number that counts as success, how long we will wait, and—the part that matters most—what we will do if it fails. All of that is on record in the audit before any result exists.
Why declare first
Because an undeclared experiment cannot fail.
Without a threshold written down in advance, any outcome can be narrated as a partial success. Engagement was flat but reach was up. Conversions did not move but the audience got more qualified. These sentences are not lies, exactly. They are what happens when a result arrives and a story gets constructed to fit it, and every agency in the world does this, mostly without noticing.
Writing the number down first makes retrofitting harder than reporting honestly. That is the entire mechanism, and it is a discipline imposed on us rather than on you.
Failures get reported as plainly as successes
A month where two of three experiments missed their threshold is reported as a month where two of three experiments missed their threshold. We revert or move on according to what we said we would do, and the audit says what we learned.
This is uncomfortable to write and it is the point. A report that is always green is a report nobody can use, and a partner who never hears about a failure has no reason to believe the successes.
What counts as an experiment
Something bounded and answerable within a defined window. A change to the newsletter’s structure. A different opening format for episodes. A landing page tested against its predecessor. A shift in publishing time. A new topic area, run long enough to be judged.
What does not count is anything too diffuse to threshold. “Improve the brand” is not an experiment. Neither is a change measured over a window too short to distinguish signal from an ordinary bad week.
We run a small number at a time, because a month with six simultaneous changes teaches you nothing about which one mattered.
Where these sit relative to the thesis
The thesis is the engagement-level claim, and it carries its own falsification line—one sentence naming what would prove the whole bet wrong.
Monthly experiments are subordinate bets underneath that claim. A failed experiment is a local correction. A pattern of failures in the same direction is evidence about the thesis itself, and the audit will say so rather than letting a hundred small corrections accumulate into an unexamined position.
Falsifiability at the experiment level is fairly common. Falsifiability at the engagement level is not, and it is the more consequential of the two.
The rigor is not in the metrics
Anyone can produce a dashboard. The thing that separates real measurement from a well-designed report is the willingness to call something a failure when it was your own idea.
That is what this section exists to enforce, and it is a standard we hold ourselves to whether or not you are checking.