
Measuring visitor demographics: what a counter can tell apart - and why every number needs its denominator
A counter that counts answers one question: how many. A counter that distinguishes answers a dozen - child or adult, with a pram or without, alone or in a group, hurrying or strolling, and where the gaze goes. Measuring visitor demographics means getting those answers from a 3D sensor: it tells children apart by body height, trolleys and prams by object detection, and groups by a group count of its own. Age, gender and height hold as distributions - as a share of the people the attribute was assigned to, not as a ratio against the visitor count.
Every one of those answers is useful. And every one of them has a property that often gets lost in a presentation: its own denominator.
This article walks through what a 3D sensor can tell apart, how each of those quantities is calculated and what it refers to. It builds on People counting: technologies compared, which covered the methods, and on Measuring occupancy, which covered the difference between a count and a balance.
The distribution holds. The ratio against visitors does not.
That is the sentence this whole subject turns on, and it sounds pickier than it is.
An age distribution answers the question: how does the group that was assigned an age break down? It does not answer: what percentage of my visitors were under 18? The two questions sound alike and have different denominators.
Why the difference is not academic shows up in one detail of how counting works. An age counter records people in both directions, on the way in and on the way out. The visitor counter only counts the entry. Divide one by the other and you get a share of children that is roughly twice as high as reality - and that still looks plausible, because nobody knows what it ought to be.
Within the age distribution, by contrast, numerator and denominator come about the same way. The skew hits both equally and the proportion stays right. That is why the distribution is the statement, and why every number in this article comes with a note of what it refers to.
Anyone reviewing an analysis should therefore ask one question first: what was it divided by? If the answer is not stated, the number is not finished.
Gender distribution: a statistic about many, not a determination
The sensor assigns a person to one gender or the other on the basis of easily recognisable visual cues. A determination of biological sex is not the intention; the counts are a statistical measurement across a large number of people. The accurate term for this is gender expression.
We use it because it is the most honest one. It describes a distribution and promises nothing about any individual.
The share is calculated as assigned male divided by assigned male plus assigned female, to one decimal place. The denominator is the gender-assigned population, not the visitor count. Where the sensor makes no assignment, the person drops out of numerator and denominator alike: the distribution describes the people who were assigned a gender.
The figure becomes useful as a pattern over time: how the split shifts through the day, across weekdays, and whether a promotion changes it in particular hours. That is a different source from the loyalty card, which is only collected at the till and only from those who take part.
Age estimation: seven bands, apparent age
The wording is precise here too: what is estimated is apparent age, which can differ from actual age. The estimate runs on the sensor; what gets analysed is counts per age band, not an image.
What is reported is the distribution across seven bands: 0 to 17, 18 to 24, 25 to 34, 35 to 44, 45 to 54, 55 to 64, and 65 and over. The share per band is that band's value divided by the sum of all seven - the age-attributed population.
The typical use is the test. A change of range, a new window display, altered opening hours: setting the distribution before and after side by side shows whether the make-up of the people present has shifted. That is a question a till system only answers when it is too late.
One precondition applies to any comparison between sites: identical sensors and identical mounting. A different model or a different height changes what gets recognised. Two distributions from differently equipped sites are two measurements, not a comparison.
Object detection: sightings are not entries
The sensor recognises four kinds of object: bicycle, pram, wheelchair and shopping trolley. None of them appears in a transaction report, and each shapes the flow of people - trolleys make a queue look longer because they take up more floor than a basket, prams and wheelchairs decide how wide a route has to be.
The number behind them is a different one from the visitor counter, and you need to know that before you read it. Objects are sighted, not counted on entry. Within one transmission from the sensor each object is counted once; across several transmissions it is seen again.
A bicycle parked in the field of view therefore produces around 120 sightings in five minutes. That is not an error but the nature of the quantity: it measures presence over time, not arrivals.
Which gives the right way to read it. “Trolley use rises on Thursdays from 4 pm” makes sense - a pattern over time. “This many trolleys entered the store” does not - that would be a count, which this quantity is not.
Read correctly, it is still a strong basis: for providing and returning trolleys, for the case for wider routes, and for the question of whether the entrance is short of bicycle parking.
Groups as a presence index
Groups shop differently from individuals. They move through the floor together, and often one person pays for everyone. Anyone who plans tills and staff around the individual visitor therefore misses the hours in which the floor is full and the till is still quiet.
What is measured is the presence of groups per transmission, capped at the visitor count. The result is a presence index, not a number of distinct groups. It shows when an area sees groups: peak hour, peak weekday, trend over weeks.
The metric “group entries per 100 visitors” is an index, and by construction it can exceed 100. That is not a malfunction but the correct output of a measure that tracks presence rather than heads - much like capture rate, which may go above one hundred per cent in Measuring dwell time because it counts zone visits.
The same holds for the ratio of trolleys to groups. It relates two presence indices to each other and can be read as a time series against itself - not as the share of shoppers who take a trolley.
Group counting is included in the sensors we use, with no additional licence.
View direction: a rose, not a map
Security cameras rarely hang where they could measure what people are interested in. A 3D sensor with the relevant add-on supplies a gaze vector per person, and from that a directional profile emerges.
It is shown as a rose diagram: a histogram of the recorded glances over twelve slices, adjustable between one and 72. Zero degrees is north. The rose shows directions, not places.
One detail of the arithmetic shows why you cannot simply average here. Two glances, one at 350 degrees and one at 10 degrees, both point almost exactly north. Their arithmetic mean, however, is 180 degrees - south. The dominant direction is therefore calculated as a circular mean, which handles the wrap-around at zero degrees correctly. A tool that averages angles like ordinary numbers points the wrong way at every north-facing window.
Two further properties make the rose readable. Whatever cannot be assigned to a zone is shown as an area of its own rather than quietly added to one. And the attention rate - viewers divided by visitors - is reported with its denominator. Viewers are the identifiers distinct within each day, summed across days; someone who comes on several days counts on each of them.
The strongest use is before-and-after: does the rose shift after a change, and by how many degrees? That makes a layout hypothesis testable instead of a matter of feel.
Walking speed: strolling or passing through
Who strolls and who hurries decides routing, placement and staffing. To this day it is mostly established by observation - sampled, expensive and rarely repeatable.
Walking speed is calculated from the change in position between successive frames. It needs at least two positions per person. Anyone standing still - below 0.05 metres per second - drops out, as do outliers above running speed. The default bands:
| Band | Speed | What it describes |
|---|---|---|
| Slow browser | up to 0.5 m/s | Looking, comparing, lingering on the move |
| Normal | up to 1.5 m/s | Ordinary walking pace |
| Fast walker | up to 2.5 m/s | Purposeful, on the way somewhere else |
The statement lies in the distribution, not the mean. The share of slow browsers says more about the quality of an area than any average speed; its denominator is the number of valid samples, not of people. And a trend is only shown once both comparison periods have at least 30 samples - below that, an arrow would too often be mere chance.
Height and the share of children
Whether an area sees families decides range, route width and fittings. Telling adults from children runs on body height, which the sensor supplies without an additional licence.
Height is converted to centimetres, held between 30 and 250 centimetres and split into five bands: under 100, under 130, under 160, under 190, and 190 and over. On top come a fine-grained histogram and the share of children - of the height-attributed population, and so again with its denominator.
Five bands rather than two is not a detail. Where the line between child and adult sits is a choice, not a law of nature. Draw it at 130 centimetres and you measure a different group than at 160.
Boundaries you draw later
This is a property you cannot see from the outside of a counting system, and in operation it turns out to be particularly valuable: the band boundaries for speed and height can be changed afterwards, and the new split applies to the entire stored history.
That works because the fine-grained distribution is stored, not just the bands. The split is a matter of analysis, not of capture. What that means in a project:
- The first analysis runs on the default boundaries. It can be read at once, with no prior knowledge of the floor.
- After a few weeks, the real distribution is known. Only now is it visible where the interesting boundaries lie at this site.
- The boundaries move to where they belong. For speed anywhere between 0.1 and 5.0 metres per second, for height between 50 and 250 centimetres.
- The whole history is read with the new boundaries. No data lost, no recapture, no year of waiting.
The dwell-time tiers, by contrast, are deliberately fixed. That is not a contradiction: the tiers are there to compare sites, and a tier cut differently at every site is useless for comparison. A child threshold, on the other hand, should fit your own customers.
What the sensor has to be able to do
None of these quantities comes as standard: each needs a function on the sensor. That belongs in the planning, not in the surprise after installation:
| Quantity | On the sensor |
|---|---|
| Gender distribution | Gender add-on, licensed |
| Age distribution | Age estimation add-on, licensed |
| Objects | Object detection add-on, licensed |
| Groups | Built into the sensor, no additional licence |
| View direction | View direction add-on, licensed |
| Speed and height | Position data; height without additional licence |
On top of that, each of these functions only works within a certain range of mounting heights, depending on the sensor model. The question of which ones a project needs therefore belongs before the choice of sensor model - not after it.
On data protection, none of these quantities changes the principle: image processing stays in the sensor, and in counting operation what leaves it is counts, positions and attributes, not an image. Why no identifying data arises from that is covered in People counting and privacy.
How we do it
The ANALYSIT Counting System reports gender and age distribution hourly and per site, each as a share of the assigned population and with its denominator stated. The seven age bands run from 0 to 17 up to 65 and over.
Bicycle, pram, wheelchair and shopping trolley appear as a series of their own over time. Groups appear as a presence index with daily trend, peak hour, peak weekday and group entries per 100 visitors. View direction shows a rose diagram with the dominant direction as a circular mean, an hourly list, a daily trend and the attention rate.
Walking speed and height are reported in bands, with a histogram and the share of children of the height-attributed population. The band boundaries of both can be moved afterwards and apply to the entire history. We switch the analyses on during onboarding, matched to the functions of the sensors.
If you want to know which of these distinctions actually decides something on your floor - and which sensor function it takes - talk to us. The answer affects the choice of model, and that is made before installation.
