
Finding a pattern is easy - when a deviation is really a signal
Finding a pattern is easy. Knowing whether it is real is the work.
Two curves move alike, and a course of action is on the table. A day looks unusual, and somebody goes hunting for the cause. A forecast points upwards, and staff get scheduled. In all three cases the same step is missing: checking whether what was found survives a second look.
This article covers how to calculate relationships, outliers, forecasts and influencing factors so that they can carry a decision. It builds on The half outage, which dealt with damaged days, and on 800 visitors is not an answer, which covers the weather correlation and the exact p-value.
Fifteen pairs, one chance hit
Check six metrics of a site for relationships - visitors, customers, passengers, dwell time, gender distribution, group share - and you are checking 15 pairs. At the usual five-per-cent threshold per test, a chance hit among them is not the exception but to be expected - even in data made of pure noise.
The usual approach turns that into a finding. Each pair is tested as if it were the only test. The report shows the one hit; nobody mentions the 14 unremarkable pairs. Chance becomes a refit, a campaign, a changed rota.
The remedy is a correction for the number of pairs tested at once. The most widely used is the Benjamini-Hochberg procedure: all computable pairs are considered together, sorted by p-value and adjusted so that the share of false discoveries among all hits stays under control. Significant then means the adjusted value q stays below 0.05 - not that p is below 0.05.
The result is fewer hits. But ones you can base an investment on.
Strength follows fixed bands: from an absolute value of 0.7 strong, from 0.4 moderate, from 0.2 weak. The list of the strongest positive and negative relationships contains only pairs that survived the correction.
Six safeguards before the first coefficient
Most false relationships arise not in the formula but in the data before it. So six checks run before anything is calculated:
- Damaged days drop out, as described in The half outage.
- A metric only takes part if it is defined on at least half the days - and on at least seven per pair.
- A metric without variance drops out. A constant can relate to nothing, but it can still produce a number.
- Identities never enter the matrix. The female share correlates at −1.0 with the male share by definition; that is arithmetic, not a finding.
- Missing values stay empty. No filling with zero or a mean.
- One list of days for every series, so the same index is the same day everywhere.
Metrics that fail one of these checks do not vanish quietly. They appear in a list of their own, “no usable data”, on the same page.
Which day demands an explanation?
An unusual day in a line chart is usually only noticed when someone is looking for it. The obvious automation - flag anything more than two or three standard deviations from the mean - has a built-in flaw: the outlier distorts exactly the figures it is measured against. An extreme day lifts the mean and widens the standard deviation, and together that makes it look less unusual.
Robust statistics get round this by using quantities a single outlier barely moves:
- The modified z-score measures the deviation from the median, divided by the median absolute deviation (MAD), with a factor of 0.6745. Median and MAD stay stable even when one day is completely out of line.
- The interquartile fence runs from the lower quartile minus 1.5 times the interquartile range to the upper quartile plus the same distance.
A day is anomalous from a modified z-score of 2.5 in absolute terms or outside the fence, critical from 3.5. Hits reported only by the fence count as information.
A second detail makes detection usable in operation. At most sites weekdays and weekends sit at different levels, and an ordinary Saturday always looks unusual against a weekday profile. So both segments are assessed separately as soon as each has at least eight observations.
Detection starts from 14 days of data. Metrics with too few observations are explicitly listed as skipped. And the sensitivity is deliberately fixed: a threshold that can be turned down at will eventually finds whatever anomaly somebody is looking for.
An anomaly is a prompt, not an answer. It says which day needs explaining. The explanation is usually in the calendar or in a note on that day.
The forecast: your own weekly pattern, carried forward
Staff, deliveries and opening hours are planned for the weeks ahead - mostly from memory of the last one. Nobody carries a weekly pattern over months in their head, and a single public holiday shifts the memory permanently in the wrong direction.
The method for this is Holt-Winters, a form of triple exponential smoothing. It estimates three things and carries them forward: level, trend and weekly seasonality - how much a Tuesday brings relative to the week. Multiplicative means Tuesday is not “200 visitors fewer” but “15 per cent below the weekly mean”, and that stays right even as the level rises.
How strongly each of the three components reacts to new data is set by three smoothing parameters. They are not guessed but searched across 80 combinations, and the best one wins. Three properties turn that into a forecast you can rely on rather than one that merely looks good:
- Gaps drive the calculation but do not count. Missing days are interpolated only to keep the recursion going and are excluded from every error measure. The series is spread over real calendar days, so a missing day does not shift the weekday.
- Diverging combinations are discarded. A parameter choice that produces impossible values is thrown out.
- The fallback model is named. If no combination holds, the forecast falls back to the median per weekday - with its own reason in the result and a visible label. No quiet degradation.
Closed weekdays are shown as closed, not as a very small number. The horizon runs from one to 30 days, and weekly totals mark part-weeks. The diagnostics are open: the three smoothing parameters, the seasonal period, the mean absolute error and the root mean squared error - not hidden behind a traffic light.
The ramp-up curve
A forecast is only as good as the history it carries forward. It is calculated on days with actual data, not on the calendar range selected:
| From day | What the forecast can do |
|---|---|
| 14 | Minimum. A weekly pattern needs two complete cycles; before that the result is deliberately empty |
| 28 | The weekly pattern is stable enough to support planning conversations |
| 60 | Single outlier days no longer dominate the seasonal factors. Only now is a statement about trend worthwhile |
This curve belongs in every rollout plan. A site that goes live today has a usable forecast signal in four weeks and a good one in two months.
Driver or companion
This is the distinction missing from most analyses, and it is the real value of this part.
“Influencing factor” charts are easy to produce and easy to misread. The most common error: quantities measured on the same day as the target end up on top and “explain” it. Passenger count, for instance, correlates at r ≈ 0.999 with visitor count - both come from the same counting line on the same day. As a bar it would be the strongest driver on the page, and entirely without planning value.
| Companions - measured on the same day | Drivers - known before the day |
|---|---|
| Passenger count from the same counting line | Weekend yes or no |
| Dwell time | Previous day's visitor count |
| Gender distribution | Visitor count on the same weekday last week |
| Group share | Temperature and precipitation |
The left column only comes into being with the visit itself. It describes the day; it does not predict it. So it sits in a table of its own, never in the ranking of influencing factors and never in the model. Only what is known before the day can explain the day.
For the right column there are two calculations. Numeric factors are each assessed with a simple regression. Categorical factors such as weekday use a bias-corrected epsilon squared - otherwise the mere number of groups creates apparent strength. The headline figure is a genuine multivariate adjusted R²; if the model is collinear or singular, it explicitly returns no result rather than a reassuring number.
Lagged factors - previous day, same weekday last week - are looked up by calendar date, not by position in a list. Otherwise, with a day missing, “the previous day” would suddenly be the day before that.
What if: a calculation, not an oracle
“What if we opened an hour longer?” gets estimated in many meetings, usually by whoever estimates most convincingly. What is missing is not an oracle but a traceable calculation that is the same for everyone, with its assumptions set out alongside.
The baseline comes from your own data: the historical mean per weekday and hour. Four levers act on it, and each is shown in the scenario with its value:
| Lever | Effect on the baseline | Source |
|---|---|---|
| Weather | sunny ×1.12, rain ×0.83, cold ×0.88, hot ×0.93 | industry assumption |
| Public holiday | visitors ×1.38 | industry assumption |
| Staffing | elasticity 0.28 on the customer count | industry assumption |
| New opening hour | 30 per cent of the neighbouring hour's level | from your own data |
That three of the four levers are industry assumptions is not in the small print but in plain sight in the scenario. That is the point: a scenario is a structured thought experiment, and everyone in the room sees which assumptions it uses. Anyone who thinks a different assumption is right for their site can argue about that - instead of about the result.
The contributions of the individual levers are attributed step by step and add up exactly to the headline figure. Below 100 groups of date and hour, the scenario stays empty instead of showing a row of zeros.
The first and the last hour
Opening hours have usually just evolved over time. Whether the first hour pays for itself, whether people are already there before opening, whether the last hour still has traffic - a site guesses, often for years.
The analysis for this is deliberately descriptive, with its thresholds in the open: open later if the first hour carries less than 2 per cent of daily traffic; open earlier if more than 2 per cent is there before opening; close later if the last hour carries more than 5 per cent. Weekdays and weekends are looked at separately.
One calculation detail decides whether it is usable: the mean per hour is divided by the number of days analysed, not by the days on which that hour had data. Otherwise the hourly shares add up to more than 100 per cent of the daily mean - and a recommendation rests on an hour that never existed in that form. Below three analysed days per weekday the value stays empty.
How we do it
The ANALYSIT Counting System calculates correlations between six metrics with an exact p-value and a Benjamini-Hochberg correction; temperature, humidity and precipitation join once weather data covers more than half the days. Anomaly detection works with the modified z-score and the interquartile fence, weekdays and weekends separately, from 14 days.
The Holt-Winters visitor forecast covers one to 30 days, with a search across 80 parameter combinations, a named fallback model and open diagnostics. Influencing factors separate drivers from companions and report an adjusted R².
Scenarios run on your own baseline with four openly stated levers. The opening-hours analysis shows the distribution across and before trading hours with fixed thresholds, and the staffing heatmap sets planned against recommended, hour by hour.
If a decision in your organisation currently rests on a single correlation, talk to us. The question of how many pairs were checked to find it is usually the most revealing one.
