Six steps in a fixed order. Any competent team already does some of them. The discipline is doing all six, in this order, on every pass, and the order is where a large share of the arguments about marketing performance originate.

Each step carries a rule and the specific failure that rule exists to prevent.
Write the output the system has to produce, in figures, with a date on it, before anything is built.
Fit the measurement, then watch a real event travel from the page to the report and confirm it arrived with its fields intact.
Assemble the thinnest complete path from first contact to recorded outcome, then widen it.
Take the reading on a fixed cadence, from the instruments, never from recollection.
Compare the reading to the pass band that was written before the reading existed.
Change one variable, log it against the reading, and run the sequence again.
A large share of the disagreements about marketing performance trace back to instrumentation fitted after the thing it was meant to measure.
Fit the instruments after launch and the first weeks of data record a configuration in progress. Those weeks then become the baseline every later result is compared against, so the programme is measured against its own worst period and nobody notices for two quarters.
Fit them first, prove a test event arrives, and the first reading is a reading. Everything downstream depends on that single ordering choice, which is why it is step two rather than step five.
No campaign launches until somebody has watched a real event travel from the page to the report
The same rule extends past tracking. It covers a form posting into a routing rule, a webhook, a sending domain, a tag moving between accounts, and every connector anybody ever describes as configured. A configured integration and a working one are different states, and only one of them has been observed.
Demand volume, changes on the answer surfaces, and the rate at each junction through conversion. Read for a step change, never against the tolerance.
Cost per qualified opportunity, instrumentation integrity, and the variable log. This is the only reading compared against the written tolerance.
The specification itself, and whether any subsystem should be rebuilt rather than tuned.
A weekly figure on a business to business programme is mostly noise, and treating it as a verdict produces panic changes that break the one variable per pass rule inside a fortnight.
The machine prepares evidence at scale. A person keeps the call. At each of these, the work stops, the evidence is presented, and somebody named signs it.

Step six feeds step one on the next programme. A rule proven on one system starts the next one at a higher standard of evidence, so the sequence is never finished and that is the point of running it.
The practical form of this is a ledger. Every test that produced a clean read gets written down with the client, the date and the condition it was true under. The ledger is what turns a decade of separate engagements into a discipline instead of a set of anecdotes, and it is the reason the same mistake stops being made.
What each one produces, the three readings taken on it, and the way each one fails.