Field Notes 05 · EVIDENCE

You cannot prove ROI with a survey

The CFO asks what the programme returned. You show a rising NPS line. He asks what would have happened anyway. That question is the entire discipline.

Here is a meeting that happens every year, in every large retailer, in more or less the same words.

The CFO asks what the customer-experience programme returned. The CX director presents a chart: satisfaction is up eleven points since the programme began. It's a good chart. Then the CFO asks the only question he was ever going to ask.

"And what would have happened if we'd done nothing?"

The honest answer, in almost every case, is that nobody knows. And the chart cannot help.

The blind spot

Nearly every VoC business case rests on a before-and-after comparison. We ran the programme; the number improved; the improvement is the return.

This is the oldest error in empirical work, and — this is the uncomfortable part — your CFO already knows it is an error. He has seen numbers go up because a competitor closed. Because the weather was mild. Because a rival's app broke. Because the category as a whole drifted upward and carried everyone with it. He does not distrust your chart because he distrusts you. He distrusts it because a trend line has no control group, and he learned to ask for one long before you walked in.

A before-and-after chart cannot distinguish between a programme that worked and a year that was going to be good regardless. Neither can you. Neither can the vendor who sold it to you.
The principle

The reframe

Proving that an intervention worked is not an analytics problem. It is a causal inference problem, and causal inference requires one thing above all others: a credible counterfactual. A version of the world where you did not act, running in parallel, against which the version where you did can be measured.

Which sounds impossible — until you notice that a large retail network is one of the finest natural experiments in commercial life, and that it is sitting unused in almost every company that has one.

Five hundred stores. Broadly comparable formats, comparable assortments, comparable customers, all absorbing the same macroeconomic weather at the same time. You do not have one business. You have five hundred near-replicates of the same business — which is to say, you have a laboratory.

So run the fix as an experiment. Roll it out to a hundred stores and not to the other four hundred. Stagger the rollout by region over eight weeks. Compare treated against untreated. Difference-in-differences is not exotic mathematics — it is a technique a competent analyst can execute in an afternoon, and it converts "satisfaction went up" into "satisfaction went up because of us, by this much, with this confidence interval." Those are different sentences. Only one of them survives contact with a CFO.

Figure 5 — The only comparison that means anything
STORES THAT GOT THE FIX vs STORES THAT DIDN'T FIX ROLLED OUT Control Treated THE EFFECT COUNTERFACTUAL — WHAT WOULD HAVE HAPPENED ANYWAY OUTCOME BEFORE AFTER Both groups moved together before. Only the gap after the fix is yours.
Before the intervention, treated and control stores move together — and that parallel movement is what earns you the right to make the claim at all. After it, they separate. The dashed line is the counterfactual: what the treated stores would have done had you left them alone, inferred from what the control stores actually did. The shaded gap between them is the effect. It is the only part of the improvement you are entitled to take credit for — and, not coincidentally, the only part your CFO will fund again.

The uncomfortable condition

There is a price of entry, and it would be dishonest to bury it in a footnote.

This only works if you are willing to share the outcome variable — revenue, footfall, basket, retention, per store, per week. A satisfaction score cannot validate itself. Something has to sit on the other side of the equation, and it has to be something the business already cares about.

Most organisations flinch at this, and the flinch is understandable. But notice what it costs. Refusing to connect experience data to commercial data is precisely what condemns customer experience to being a soft function forever: perpetually asking for budget on the strength of a trend line, perpetually unable to answer the one question that would settle the argument in its favour.

What this means for everyday decisions

Three consequences
  • Never roll out a fix everywhere at once. A full simultaneous rollout is the deliberate destruction of your only control group. Stagger it, and the evidence comes free.
  • Your network is the asset. Hundreds of comparable sites is not merely an operational burden — it is the counterfactual machine that single-site businesses would pay almost anything to own.
  • Insist on the outcome variable. If experience data never touches commercial data, the ROI question will remain unanswerable forever, and it will be unanswerable by design.

The takeaway

The CX function has spent twenty years trying to prove its worth with better surveys, and it has lost that argument twenty times, because the problem was never the survey. It was the absence of a comparison.

The counterfactual is invisible. It is the world where you did nothing, and it is the only thing your improvement can honestly be measured against. You cannot survey your way to it. But if you have a network, you can build it.

Method & scope

Position Magicalia treats VoC ROI as a causal-inference problem. Where a client's network permits it, we recommend staggered rollout and difference-in-differences over before/after reporting.

Condition This method requires the client to share an outcome variable. We state that requirement up front rather than discovering it at the business-case stage.

Figure 5 Illustrative, drawn to demonstrate the method. Parallel pre-trends are an assumption that must be tested, not asserted, in any real application.

— · —

We'll show you something you don't know about your own customers — before you sign anything.

A public-signal diagnostic uses only what your customers have already said in public. No data handoff, no NDA, no procurement. You'll see what the market sees, and where the gap is.

Request a diagnostic →