
Comparing sites: the honest denominator - and why a ranking without it is worse than none
A ranking that mixes opening days and sensor outages is worse than no ranking at all. Nobody is misled by a ranking that does not exist. A wrong one tempts you to visit the wrong site, call the wrong store manager and budget the wrong refit.
The fault rarely lies in the raw data. It lies in a question hardly anyone asks before reading a league table: what was it divided by?
This article explains why site comparisons are harder than they look, which denominators exist and which question each of them answers. It builds on Measuring occupancy, which covered day boundaries and time zones, and on The half outage, which describes how a damaged day is caught.
Three ways to make a site comparison useless
Anyone running a chain of fifty or two hundred sites compares, whether they mean to or not. At that size comparability is no longer one question among many - it is the question.
Three effects make a comparison useless, and none of them has anything to do with performance:
- The opening day. A site open six days, seven days in the denominator: it looks weaker than it is. Divide only by days with figures and it looks stronger. Same raw data, two opposite distortions - depending on what you divide by.
- The silent sensor. Two days of outage look exactly like two days without visitors in a daily series. Fill the gap with zero and you invent a slump, then average it in.
- The mixed ranking basis. One page ranks per reporting day, the next per site, the third by raw total. Laid side by side without the basis stated, you are comparing answers to different questions.
A worked example shows how large the first effect is. Two sites have exactly 1,000 visitors on every day they open. Site A opens seven days, site B six.
| Denominator | Site A | Site B | What it looks like |
|---|---|---|---|
| Calendar days (7) | 1,000 | 857 | B is 14 per cent weaker |
| Reporting days (7 and 6) | 1,000 | 1,000 | both equally strong |
Both rows are correctly calculated. They just answer different questions - and anyone who sees only one of them concludes from a closed Sunday that a store is doing badly.
Two denominators, two questions
The answer is not to pick the right denominator. There isn't one. The answer is to show both and tie each to the question it answers:
| Metric | Denominator | Answers the question |
|---|---|---|
| Avg visitors per day | all calendar days in the window | How much footfall was there per calendar day? |
| Avg visitors per reporting day | days with a figure above zero | How strong is the site when it reports? |
| Coverage | reporting days divided by calendar days | How many days does the figure rest on? |
| Deviation from the mean | mean of per-reporting-day values in the cluster | Where does the site stand in its group? |
| Uniformity | standard deviation divided by the mean | Does the group run evenly, or does it carry a gradient? |
The overview shows the calendar day because it answers the question of volume. Ranking and rating work per reporting day because they answer the question of strength. A closed Sunday therefore does not drag the ranking down - and it still does not vanish from the volume.
The same idea already appears in Measuring dwell time: there, a zone's peak factor divides by days with data, because days without data would inflate every peak artificially. It is the same rule in a different place. The denominator follows the question, not habit.
Coverage: how many days the figure rests on
The third row of the table is the one missing from most league tables, and the most important. A site that reported on seven of fourteen days can have an excellent value per reporting day. The question is whether you believe it.
So every row in the ranking carries its coverage as a column of its own, and below 70 per cent a warning appears. The figure stays visible - it is not wrong, after all - but it carries a sign that it stands on thin ground.
Coverage can only be read correctly with the opening hours alongside. A site closed on Sundays has, over two weeks, coverage of twelve out of fourteen days, and that is its normal state. Maintain opening hours per site and you read that figure as normal rather than as an outage.
When an analysis would rather show no number
There are situations in which the most honest output is not a number. A good analysis recognises them and says so, instead of inventing a position, a trend or a deviation:
| Situation | Threshold | What appears instead |
|---|---|---|
| Too few sites | fewer than 4 reporting sites | no percentile rank |
| Trend base too short | fewer than 8 days with data | “insufficient” instead of up or down |
| Trend base too patchy | under half the days with data | “insufficient” |
| No usable mean | mean of zero or below | no deviation instead of 0 per cent |
| Day without a report | no record | the line visibly breaks off |
The first row deserves an example. A percentile rank across three sites says: “site B is at the 50th percentile.” That sounds like statistics and only means it sits in the middle of three. A suppressed number is a statement in itself: the data will not support one.
“Did not report” is not a bad result
This is the distinction the whole comparison turns on. A site whose sensor was silent during the reporting period is not weak. It is unknown.
What a comparison usually does with it: counts the missing days as zero, keeps the site in the same ranking as everyone else, and there it lands at the bottom - next to the genuine problem cases. The store manager gets a call about a sensor cable.
The right approach is the opposite, in three steps:
- A missing day stays a missing day - empty, not zero.
- Sites that did not report in the period sit in a list of their own next to the ranking, not in it.
- The daily series breaks off at the gap instead of drawing a line through as if nothing happened.
Days that are damaged but not empty are a separate case. How they are caught and removed from daily statistics is covered in The half outage.
Within the cluster: mean, median and gradient
An average across all the sites of a chain says little, because a city-centre store and a site by a motorway exit have nothing to do with each other. So comparison happens in clusters: groups of comparable sites, put together by hand, because only somebody who knows the sites knows what is comparable.
A cluster of unlike sites produces correct answers to the wrong question. That is the one decision no system takes off your hands.
Within a cluster, four values carry the most weight:
- Median with P25 and P75. The median is immune to the one outlier that drags the mean. P25 and P75 show how wide the middle is.
- Deviation from the mean. From 20 per cent below the mean a flag appears - it names exactly the sites that need an explanation, and no others.
- Uniformity. A single value for whether a group runs evenly or carries a gradient: the standard deviation divided by the mean. Two clusters with the same average can look entirely different - one with ten similar sites, one with five strong and five weak.
- Performance tiers plus the strongest and weakest sites, so a long list becomes readable.
Size against strength: the view across countries
Anyone operating in several countries gets either a grand total in which the structure disappears, or a list of sites nobody can take in. And a country with many sites beats a country with few in any total - without that saying anything about the sites.
The answer is two readings of the same country figure. Total shows weight, per site shows strength. The normalisation divides by the sites that reported, not by all that exist - a silent site does not dilute the country figure.
Countries roll up into regions and break down to their sites, and the daily series shows gaps as gaps. If data is missing, an error appears rather than a row of confident zeros.
A note on reading, so two percentages are not confused: in the cluster comparison the deviation is calculated per reporting day, in the country comparison per reporting site. Both are correct. They answer different questions.
The last seven days: a ranking with a fixed window
The third view is the quickest: who was ahead over the last seven days, who behind? It deliberately calculates differently from the cluster, and that is exactly what you need to know.
Here the divisor is always the full window - seven days, never the days with data. The average per day is the total divided by seven. That is the exact reverse of the cluster, and it is right: the question is not “how strong when open” but “how much this week”.
Two properties make the ranking reliable. The comparison period is the same length and directly adjacent: the seven days against the seven before, not against a monthly average. And the trend across all sites is weighted by visitors - not the mean of per-site percentages, in which a small store at +40 per cent counts as much as the flagship at −5.
Each site is cut into days in its own time zone. The top and bottom of the list come only from sites with actual data days, and whoever did not report sits - as in the cluster - in the separate list alongside.
Only sites set up the same way are comparable
One last effect arises not in the analysis but in the setup. Whoever opens a store every month sets up the same configuration by hand again and again. Each repetition produces a different threshold, a different zone, a different alert recipient - and the comparison then compares differently configured sites.
What helps is to clone a configured site instead of setting each up from scratch. Zones with capacity and warning level, thresholds, alert recipients and channels, opening hours and the report schedule are carried over - according to a fixed list that is disclosed beforehand. Display keys are issued afresh rather than copied; a cloned site inherits no public access.
The twelfth site is then set up exactly like the first - and only then genuinely comparable with it.
Five decisions before the first comparison holds
- Country per site. Decided once, it carries the entire country and region comparison.
- Time zone per site. If it is wrong, the day boundary shifts relative to its neighbours.
- Opening hours. With them, coverage of six out of seven reads as normal rather than as an outage.
- Cluster membership. Comparable together, incomparable apart.
- The comparison analyses. We switch the cluster and country comparison on during onboarding.
All five are decided once and hardly touched again. They decide, though, whether every later league table is a statement or a claim.
How we do it
The ANALYSIT Counting System compares sites in clusters with two stated denominators: average visitors per calendar day in the overview, average per reporting day in ranking, tiers and deviation. Every ranking row carries its coverage, with a warning below 70 per cent; a flag appears from 20 per cent below the mean. Plus median with P25, P75 and P90, uniformity, and export as CSV and XLSX.
Below four reporting sites there is no percentile rank, below eight days or half coverage no trend. Sites that did not report sit in a list of their own, and daily series break off at gaps.
The country comparison rolls sites up into countries and regions, either as a total or normalised per reporting site. The 7-day ranking uses raw totals over a fixed window against the equally long previous period, with a visitor-weighted trend. Configured sites can be cloned as a template.
If your league tables divide by calendar days today without saying so, talk to us. One site that closes one day a week is usually enough to see the difference.
