Field Notes 02 · METHOD

Stop averaging your reviews

Your support tickets and your public reviews are different instruments, pointed at different people. The average of the two describes nobody.

Somewhere in your business there is a slide with a single number on it. Overall customer sentiment: 3.9. It was produced by taking everything customers said — the support tickets, the survey responses, the public reviews — and averaging it.

It is a tidy number. It is also, in a precise and demonstrable sense, a number describing a person who does not exist.

The blind spot

Every feedback channel is an instrument, and every instrument selects its own population.

People open a support case because something went wrong and they needed it fixed. That channel is selected, almost by definition, for friction. It over-samples problems, because a problem is the entry condition.

People post a public review for the opposite reason: they felt something strongly enough to tell strangers about it. That channel is selected for intensity, in both directions. It over-samples the extremes — the delighted and the furious — and it barely hears from the vast, quiet middle who had a perfectly fine time and never thought about you again.

These are not two samples of the same thing. They are two different measurements, of two different populations, with two different biases. Averaging them is not aggregation. It is a category error.

The same operator, the same period: 1.2 on the public review platform, 4.2 on the app stores, 4.1 to 4.6 from local users in its own cities. Three instruments. Three truths. Their average is true of nobody.
From a public-signal diagnostic, 2026

That spread is not noise, and it is not a data-quality problem to be cleaned away. Each instrument is faithfully measuring a different population. The public review platform caught the people whose trust had been broken — a specific, furious cohort with a specific grievance. The app stores caught the daily users for whom the core product works well. The local city ratings caught the residents who use it as ordinary infrastructure and think nothing of it. Every one of those numbers is correct. The average of them is the only figure in the set that corresponds to no living customer.

The reframe

Keep the instruments apart. Always. Report them side by side and let the divergence between them become the finding — because the gap between what your customers tell you privately and what they tell the world is one of the most diagnostic signals you own, and averaging is the single most efficient way to destroy it.

An operator whose internal cases look healthy and whose public reviews are savage has a very specific problem: the people who are angriest are not bothering to complain to you. They are going straight to the internet. That is a different disease, with a different treatment, from an operator with the reverse profile — and both of them can average out to 3.9.

Figure 2 — Two instruments, two populations, one meaningless mean
INSTRUMENT A — INTERNAL CASES (PEOPLE WITH A PROBLEM) 1★5★ INSTRUMENT B — PUBLIC REVIEWS (MOTIVATED TO POST) 1★5★ AVERAGE THEM AND YOU GET… 3,9 a score held by almost none of your actual customers !
Internal cases skew toward problems, because having a problem is how you get into the dataset. Public reviews go bimodal — a wall of 1★ and a wall of 5★, with a hollow middle. Average the two and you land at 3.9: a value sitting almost exactly in the trough where the fewest real customers are. The mean is not a summary of this data. It is an artefact of it.

What this means for everyday decisions

Three consequences
  • Never average across sources. One score per instrument, reported alongside its siblings. If a dashboard shows you one blended sentiment number, it is hiding something from you.
  • Read the gap, not the mean. The distance between your private and public signal is a finding in its own right — often the most actionable one on the page.
  • Distributions over averages. A bimodal 3.9 and a flat 3.9 demand opposite responses. The mean cannot tell them apart. The histogram takes two seconds.

The takeaway

Averaging feels like rigour. It has the shape of rigour — it reduces, it summarises, it produces a single defensible-looking figure for the top of the deck. But a mean is only meaningful when the things being meaned are the same kind of thing.

Your channels are not the same kind of thing. The blend is not a summary of your customers. It is a summary of your instruments.

Method & scope

Position Source-provenance separation is a locked design decision in Magicalia Insights: signals from distinct instruments are never merged into a composite score.

Grounding The 1.2 / 4.2 / 4.1–4.6 spread is a real observation from a Mode A diagnostic on a European parking operator — built entirely from public signals (372 reviews on a public review platform, 764+ app-store reviews, a national complaints portal), in the same measurement period. The operator is not named: the finding is structural, and the identity is not ours to publish.

Figure 2 Illustrative distributions, drawn to demonstrate the mechanism.

— · —

We'll show you something you don't know about your own customers — before you sign anything.

A public-signal diagnostic uses only what your customers have already said in public. No data handoff, no NDA, no procurement. You'll see what the market sees, and where the gap is.

Request a diagnostic →