
800 visitors is not an answer - how weather, holidays and staffing explain a number
“800 visitors” is not an answer. It is a question.
The same daily figure means something different in steady rain, on a public holiday, in the middle of a campaign, or with three staff on the floor instead of six. Look only at the number and you see four different days that look alike - and draw the same conclusion from each.
This article is about putting context next to the number, and about where a relationship is actually evidenced rather than merely plausible. It builds on The half outage, which describes how damaged days drop out of every statistic, and on Comparing sites, which covered the denominator.
The year-on-year comparison that says nothing
The favourite comparison in operations is against the previous year. It is also the one that misleads most easily.
Easter falls in March one year and April the next. Autumn school holidays in the canton of Zurich do not necessarily fall in the same weeks as in Bern, nor in the same weeks as the year before. A holiday that fell on a Sunday last year falls on a Monday this year. And whether October was wet, nobody remembers twelve months later.
A month against the same month last year therefore compares not two trading years but two calendars and two sets of weather - and the difference that remains looks like performance. Without context, a deviation is a question, not an explanation.
The answer is not to calculate the context away as if it did not exist. It is to place it next to the count, so that the question “was it the rain or us?” gets an answer that can be evidenced.
Weather: a relationship with evidence
For every site, the daily weather summary at its coordinates is fetched and set next to the daily count. That gives four views on one page: a scatter of temperature against visitors, bars by weather condition, a rain-versus-dry comparison and a day-by-day table.
The arithmetic behind it is the Pearson correlation, and three points decide whether it stays honest:
- The p-value is exact. It is calculated via the regularised incomplete beta function, not a rule of thumb. That sounds like mathematics for mathematicians and is a practical difference: with 30 or 40 days, approximations can be far enough off right around 0.05 to turn “just not significant” into “just significant”.
- Rain has a threshold. A day counts as rainy above 0.5 millimetres. Below that it is drizzle that stops hardly anyone shopping, and counting it dilutes the comparison.
- Complete days only. A day with a sensor outage looks like a quiet day and quietly pulls every correlation in one direction. How few such days it takes to shift an analysis is worked out in Footfall data quality: the half sensor outage. Damaged days therefore drop out of every statistic, and their number appears on the result.
An empty box is a result too
The analysis comes with a planning hint, and it has a strict condition: it only appears if the relationship is significant - p at most 0.05 - and the effect is at least 10 per cent. Otherwise the box stays empty.
The second condition matters more. With enough days almost any relationship becomes significant, including one that means nothing in day-to-day operations. A two-per-cent effect of rain may be statistically provable and worthless for staff planning. The threshold separates what is true from what changes something.
The same attitude applies to the details: if the chosen window has not a single rainy day, the rain-versus-dry comparison shows a dash instead of an invented difference. Below two points there is no trend line. An analysis that stays silent when it has no basis is the only one you believe when it speaks.
And one note that belongs with every correlation: it describes a relationship, not a cause. Weather explains part of the day, not the day.
The lead time belongs in the project plan
A weather correlation needs weather history, and for a new site that has to be fetched first. It takes around 1.2 days per site; ten new sites therefore take about twelve days. It also needs valid coordinates and at least seven usable days before the page calculates at all.
That is a matter of planning: the weather correlation is an onboarding topic, not a day-one feature. Promise it for the first board meeting after installation and you promise something that has no basis yet.
The outlier usually has a name
An unexplained daily peak ends up in a discussion: was it the holiday in the neighbouring country, the start of the school holidays or our own event? The answer is almost always in a calendar - just not the one next to the visitor figures.
The calendar brings three sources into one monthly grid:
- Public holidays for 136 countries, from public sources and without a licence. Regional holidays are marked as such.
- School holidays as a type of their own, so a holiday week is not mistaken for a normal one.
- Your own dates - events, promotions, refits - single or multi-day, with tags, colours and a type filter.
In November and December the following year is fetched in advance, so January does not start with an empty calendar. And because the calendar works per country, the country belongs to every site - the same onboarding step that carries the country comparison in Comparing sites.
For operators across several countries the calendar is more than convenience. A week with a public holiday in Germany and none in Switzerland is simply not a comparable week between two sites on either side of the border.
Campaigns: two safeguards against the spectacular uplift
How a campaign comparison is built and why a significance block belongs to it is covered in The visitors who don't buy. This is about two safeguards you rarely see, which protect every comparison from its most common error.
First: no percentages out of nothing. If the baseline is zero, the change shows a dash - not “+100 per cent”, let alone “+infinity”. Otherwise a promotion on an area nobody entered before produces the biggest success of the year.
Second: no statement from too little or from the incomparable. Hints only appear if both periods have at least 30 visitors and their daily means lie no more than a factor of 10 apart. Beyond that you are not comparing two phases of the same business but two different businesses.
Both windows can be up to 731 days wide. That is enough to set a promotion against the same period a year earlier - and so to take the calendar out of the comparison, as far as any comparison can.
Staffing: the ratio of the sums, not the mean of the ratios
“We were understaffed” stays an opinion as long as staffing does not sit next to footfall. So the headcount per shift is recorded per site and day, and from it comes the metric visitors per employee. How it feeds into staff planning is covered in The visitors who don't buy. This is about a calculation error made almost everywhere at this point.
Over a period, the metric is formed as the ratio of the sums: all visitors divided by all staff-days. Not as the mean of the daily ratios. An example shows why:
| Day | Visitors | Staff | Ratio |
|---|---|---|---|
| Saturday | 2,000 | 10 | 200 |
| Monday | 100 | 1 | 100 |
| Mean of the ratios | 150 | ||
| Ratio of the sums | 2,100 | 11 | 191 |
The mean of the ratios lets Monday with a single person weigh as much as Saturday with ten. The ratio of the sums weights each day by its actual staffing - and only that answers how many visitors there were per staff-day.
Under- and overstaffing are measured against adjustable thresholds, by default above 50 visitors per employee as understaffed and below 10 as overstaffed, with exceptions for individual dates. And a coverage note names not only the recorded days but explicitly the days with traffic but no entry - the analysis says itself how reliable it is.
Correcting with a reason
The half outage sets out the principle that a damaged day is removed, not filled in automatically. There is a different case, though: the known incident. The scaffolding in front of the entrance, the blocked door, the system's test run. Everyone in the room knows those days are wrong, and they still distort every weekly and monthly figure.
Here a correction is right - but only under conditions that distinguish it from quiet data tidying:
- Record with a reason. Day or hour, metric, target value, plus a reason code and a note. A preview shows beforehand what would change.
- Review. The correction sits as “pending” in the history - visible, but with no effect on charts, exports or alerts.
- Apply or reject. Both are decisions, and both land in the history and the log.
- Revert. An applied correction can be undone.
The difference from automatic filling is not the technology but the responsibility: a correction is a documented decision by a person, with a reason. A filled-in value is an assumption by a program, without a trace.
Which days might need a correction is shown by highlighting unusual days: a day counts as unusual if, within a 28-day window, it lies more than two standard deviations from the mean. The day itself is left out of that calculation - otherwise an extreme day would widen its own spread so much that it made itself look ordinary.
Opening hours: the cheapest context source
The least conspicuous source has the greatest effect. A site's opening hours are on the notice in every store and in hardly any analysis system.
They are maintained as a weekly grid with time windows per day and as dated special hours. Input is checked strictly - no overlapping windows, no duplicate special dates, “24:00” only as a closing time - because a wrong window has effects everywhere.
And the effects reach far: alerts can be suppressed outside opening hours, so that no message at three in the morning costs trust in the whole alerting system. The staffing heatmap limits itself to open hours. And in a site comparison, only the opening hours make coverage of six out of seven days readable as normal. Maintained once, they then work quietly in half the analysis.
How we do it
The ANALYSIT Counting System sets the daily weather next to the daily count for each site and calculates the Pearson correlation with an exact two-sided p-value, a rain-versus-dry comparison above 0.5 millimetres and a trend line over complete days. A planning hint appears only at p of at most 0.05 and an effect of at least 10 per cent; excluded days are shown on the result.
The calendar holds public holidays and school holidays for 136 countries plus your own dates in one monthly grid. The campaign comparison sets two periods of up to 731 days against each other, with a dash instead of a percentage when the baseline is zero and hints only from 30 visitors per period.
Staffing holds the headcount per shift, calculates visitors per employee as the ratio of the sums and shows days with traffic but no entry. Data correction covers 21 metrics per day or hour, with reason code, approval and history. Opening hours are available for every site and feed into alerts, the heatmap and comparison.
If your year-on-year comparison is once again raising more questions than it answers, talk to us. The answer is usually in the calendar.
