Skip to content
An almost empty colonnaded hall with a shining floor, two small silhouettes at the bright far end - no individual is identifiable.

Footfall data quality: a silent sensor must never make a quiet day

Footfall data quality is not decided on the day a sensor goes completely silent, but on the day it half fails: its daily figure is smaller but plausible, and it enters every statistic as a quiet day. That day can be caught against the site itself - by its minutes with data, measured against the 90th percentile of its own days.

The outage people fear is the complete one: the sensor goes quiet, the display shows a dash, somebody picks up the phone. That outage is harmless, because it gets noticed.

The dangerous one is the half outage. A sensor drops out for four hours in the morning and then delivers again. In the evening there is a daily figure in the report - smaller than usual, but there. No row is missing. It is simply wrong.

This article is about that day, and about a second, equally inconspicuous error: the day that starts in the wrong place. It builds on Measuring queues and wait times, which describes the three states measured, estimated and not measured, and on Measuring occupancy, which covered the dash instead of the zero. Both are about the value that is missing. This one is about the value that is there and is not right.

Why a small day is worse than an empty one

A sensor writes by the minute. On a normal trading day a site may have around 900 minutes with counts. If the sensor is silent for most of the day, perhaps 40 remain by the evening.

Every statistic built on daily values reads that day as a real day with little traffic. Why wouldn't it - it has a date, it has a number, and the number sits in a plausible range. A check for missing values finds nothing, because nothing is missing.

The result is not a gap but a false conclusion, and one that carries forward: into a weekday comparison, into a correlation with the weather, into a staffing recommendation. All plausible, all wrong - and nowhere a hint that it might be.

A worked example: a 3 per cent weekend uplift becomes 20

How far a few small days shift an analysis can be worked out. A site has 1,000 visitors on weekdays and 1,030 at the weekend - a weekend uplift of 3 per cent.

In a month with 20 weekdays, the sensor is silent for most of the day on four of them; those days report around 300 visitors each. The weekday average drops to (16 × 1,000 + 4 × 300) / 20 = 860. Against 1,030 at the weekend, that is a weekend uplift of around 20 per cent instead of 3 - from four days on which nothing happened except an outage.

The same happens to a correlation with the weather: if the outage days happen to cluster on rainy days, the correlation measures the sensor, not the weather.

A staffing recommendation can hang on each of those numbers. And hardly anyone reviewing them would recognise either as wrong, because both numbers tell a story you can picture.

The yardstick is the site itself

The obvious fix would be a minimum from the manual: a day only counts if it has at least so many minutes of data. It fails on the variety of sites. An airport terminal has traffic in almost every minute of every day; a boutique has eight hours of it and then none.

So every site is measured against itself. From its own days comes an expectation of how many minutes with data a complete day has. A day that falls clearly short does not qualify and enters no statistic.

StepHowWhy
Expectation90th percentile of minutes with data across all days with dataA complete day of this site, not of an average site
Threshold80 per cent of that expectationNormal fluctuation stays in, an outage of hours drops out
Resultqualifying and excluded days, kept separateBoth are visible, neither disappears quietly

Why the 90th percentile and not the median

The choice looks like a detail and decides whether the check works at all in the worst case.

The median is the middle day. If a sensor partially fails on more than half the days in a window, the median is itself an outage day. The expectation drops to the level of the fault, and the check declares the fault normal. It fails precisely when it is needed most.

A worked example with ten days. Six had a partial outage with around 300 minutes of data each, four were complete with 900 each:

YardstickExpectationThreshold (80 %)The six outage days
Median300 minutes240 minutespass - and enter the statistics as quiet days
90th percentile900 minutes720 minutesdrop out - and are reported as excluded

The 90th percentile sits with the most complete days of the window. It stays right even when most days are damaged - as long as a few good ones remain. A yardstick that even many bad days cannot drag down is the basis an outage check needs.

Why the same set for every metric

A second detail concerns analyses that relate two quantities - visitors against temperature, visitors against opening hours.

If outage days are determined separately for each quantity, two slightly different sets of days emerge. The pairs slip: Tuesday's visitor figure suddenly sits next to Wednesday's weather. So every series is derived from the same set of qualifying days. It is unspectacular, and without it any correlation can become a product of chance.

Excluded does not mean hidden

A day that drops out of an analysis changes it. So the number of excluded days belongs on the result, not in a log nobody opens.

“Correlation with the weather, 58 of 62 days” is a different statement from “correlation with the weather”. The first says what it rests on. The second asks for trust.

That is why the check belongs in front of every analysis built on daily values: anomaly detection and visitor forecasting, influencing factors, scenarios and staffing needs, correlations, opening-hours and weather analyses, campaign and country comparisons. And each of them should show the number of excluded days directly.

And one consequence that is easy to overlook: when days are missing, groupings shift. A weekly seasonality that counts by position rather than by calendar assigns every value after a removed Monday to the wrong weekday. So the gaps are mapped back to real calendar days before a weekly pattern is formed.

Remove, don't fill in

Once an outage day has been found, there is a temptation: to fill it in. Put in the average of the other Tuesdays, smooth the gap, close the curve. The chart looks more complete afterwards.

It is still the wrong path, for the very reason the check exists. A filled-in value looks like a measured one. It flows into forecasts, comparisons and correlations, and there it can no longer be told apart from a real measurement. You would have replaced the small, wrong day with a plausible, invented one.

So an outage day is removed and reported as removed. The dataset gets smaller, and that is honest: an analysis over 58 real days is more reliable than one over 62, four of which were made up.

A day starts where the site is

The second inconspicuous error is not an outage but a clock. That daily and hourly totals are formed in the site's time zone is covered in Measuring occupancy. How hard that is to get right is not, and it is worth explaining, because it can go wrong in more places than you would think.

The anchor on the wrong day

A portfolio across Zurich, Dubai and Auckland has no common day. That much is known. Less well known is how easily a day still slips, even when somebody has thought about time zones.

A common approach is to fix a day by a point in time - noon UTC, say - and then convert it to local time. In Zurich that works. In Auckland, at UTC plus 13 hours in the New Zealand summer, noon UTC becomes one in the morning on the following day. The site gets the wrong day's figures without an error appearing anywhere.

So day ranges are anchored on the date itself, not on a point in time. A date has no time zone that could shift it.

Daylight saving time: the hour that happens twice

On the last Sunday in October, Switzerland has the hour from two to three in the morning twice. On the last Sunday in March it is missing. An hourly curve that does not know this shows a jump on those days - and nobody can say whether it was the business or the clock.

Hour ranges are therefore built from wall-clock time, not by adding hours to a starting point. That way 2 pm stays 2 pm on the day the clocks change, and a curve across that weekend needs no footnote.

Axis and divisor from one source

The least conspicuous of the three errors: a chart labels 31 days on its axis, the average above it divides by 30, because the two numbers were computed in different places. The difference is small and therefore never noticed.

Axis and divisor come from the same array of dates. An average then divides by exactly the days visible on the axis. Invalid, reversed and over-366-day ranges are rejected before they produce a number.

Three in the morning or midnight

One distinction belongs here explicitly, because otherwise it surfaces months later as an unexplained difference. Running occupancy - for the site and for zones - cuts the day at three in the morning local time, for the reasons set out in Measuring occupancy. All other daily totals start at local midnight.

Both are right, because they answer different questions: “who is here now” needs a zero point when nobody is there. “How many came on Tuesday” needs the calendar day. You only need to know that the two views describe days offset by three hours.

Checking footfall data quality: how to recognise a number you can rely on

Five questions every analysis should be able to answer before anyone builds a decision on it:

  1. What happens to a day on which the sensor was silent for six hours? Does it enter the statistics as a small day, or is it caught?
  2. What is an outage day measured against? A fixed number, the median, or a measure that stays right even with many bad days?
  3. Is the number of excluded days on the result? Or do you have to trust that none are missing?
  4. Where does a day start? In the site's local time, anchored on the date - and what happens on the day the clocks change?
  5. What does the average divide by? The days on the axis, or a number that came about somewhere else?

None of these questions is exotic. But each of them separates a number you believe from one you can believe.

How we do it

In the ANALYSIT Counting System, every day is checked before it enters any of the 14 statistical analyses built on daily values. The expectation is formed per site from its own days, as the 90th percentile of minutes with data; a day below 80 per cent of that counts as an outage day and is excluded. All series in an analysis come from the same set of qualifying days.

These 14 analyses include anomaly detection, visitor forecasting, influencing factors, scenarios, staffing needs, correlation, opening-hours analysis, weather correlation and the campaign comparison. Each shows the number of excluded days directly on the result.

Daily and hourly aggregations are formed per site in its own IANA time zone, day ranges are anchored on the date, hour ranges on wall-clock time, and axis and divisor come from the same array of dates.

If you want to hold existing analyses against these five questions, talk to us. A single day with a partial outage is often enough to see how a system handles it.