
Open interfaces - who owns the data and how to get it out
An analytics system you cannot leave is not a tool. It is a trap.
The question of who owns the data is rarely asked when buying a counting system, and usually only when it is too late: when changing supplier, when building a data warehouse, when finance asks for the figures to go into its own reporting system. That is when it shows whether “we have an interface” was a commitment or a sales line.
This article covers how to recognise an open architecture - on the way out and on the way in. It builds on From number to action, which dealt with calls to external systems, and on Passenger counting in public transport, which covers the standard formats of public transport.
Four ways out
An open architecture has not one exit but several, independent of one another. Each answers a different question:
| Route | How | What for |
|---|---|---|
| Pull | REST interface with a key | BI tool, data warehouse, your own dashboard |
| Push | signed events to your own endpoints | ticketing system, control room, chat channel |
| File | CSV and XLSX from every major analysis | spreadsheet, archive, auditor |
| Standard | ITxPT and VDV 301 with checksum | transport network, authority, tender |
Add to that the report that arrives by itself, and the intake through which data comes in. Both belong to the question of openness, and both are described below.
An interface you can build an integration on
“We have an API” often means: a single report as an endpoint, a password in plain text, no error catalogue. Anyone trying to fill a data warehouse with that ends up writing a manual export after all.
The difference shows in four properties you can ask about in any tender:
- The whole analysis surface is available, not an excerpt. Live occupancy, historical series from one minute to one day, zones and dwell time, queues, sensor analytics, vehicle counting, clusters, campaigns, alerts and the log - the same data that is on the dashboard.
- One response format. On success, success and data; on failure, success, error and a code. An integration then has exactly one path to handle.
- Stable error codes - more than 40, and stable means: a program can tell “permission missing” from “site unknown” without reading error texts that change with the next translation.
- A version header on every response. An integration notices when something changes, instead of noticing it through wrong numbers.
Every request is checked against a schema before it reaches the database. That is invisible while everything is right, and decisive the moment a script sends a wrong parameter.
A key can only do what it was issued for
Access to the interface runs on keys, and the question is not whether there are keys but what a single key may do. For that there are seven scopes, set at issue and checked on every call:
| Scope | Grants | Typical use |
|---|---|---|
| Read data | count, zone, dwell and analysis data | BI connection, data warehouse |
| Write data | write operations on data | corrections and context data from your own systems |
| Read sites | master data and configuration | matching against store master data |
| Write sites | create and change sites | automated onboarding of new sites |
| Read alerts | rules and history | ticketing system, evidence |
| Write alerts | create and change rules | set thresholds from your own planning |
| Administration | manage keys and endpoints | provisioning by IT, rotation |
A reporting tool thus gets a key that can read and nothing else. If it is compromised, no site can be created with it and no threshold changed.
What happens to a key
The mechanics of a key decide whether a leak is an incident or a disaster:
- Only the hash is stored. The key itself is visible exactly once, when it is created. After that nobody can read it - not even the supplier.
- There are no immortal keys. Lifetime is capped at one year, and anyone who gives no expiry date when creating one gets one anyway.
- Rotation with overlap. When switching, the old key stays valid for 15 minutes so an integration can change over without an outage window.
- The type is in the key. Live, test and sensor keys can be told apart by their prefix - a test key in a production configuration stands out at a glance.
- Every use is logged, and so is every rejection. And every key carries its own rate limit, 1,000 requests per minute by default.
Events that arrive - even if the recipient is away
Polling is expensive and slow. Ask every five minutes whether an alert has fired and you create thousands of empty requests and still find out two and a half minutes late on average. Hence the reverse route: the system delivers events - fired alerts, changes to sites, zone capacity reached, queue thresholds exceeded, sensor outage and return.
There is a difference here from the calls to external systems in From number to action, and it is deliberate. A switching command has to arrive now or not at all - a traffic light turning red half an hour late does damage. An event has to arrive reliably, even if it takes time - a ticketing system that was down for maintenance should still get the alert afterwards.
So an event that is not accepted is delivered again after one minute, then after five and 30 minutes, after two and after 24 hours - six attempts over around 26 hours. Individual deliveries can then be triggered again by hand.
Four properties make this route dependable:
- Signed over timestamp and raw body. Every delivery carries an HMAC-SHA256 signature over the time and the unaltered content, not over re-formatted JSON. Any silent change on the way is detectable, and the timestamp lets a stale delivery be rejected before its content is even read.
- The target address is checked three times - at registration, at every change and at the moment of delivery, each time with name resolution. Private address ranges are refused, redirects are never followed.
- No double delivery. Every due delivery is claimed exclusively before it is sent, and carries a unique identifier. A recipient seeing a delivery twice can recognise it as the same one.
- Self-healing. A delivery stuck in processing for more than ten minutes goes back into the retry cycle. A crashed process does not lose an event. And an endpoint that fails ten times in a row is switched off rather than continuing to load the recipient's server.
The delivery history shows status, attempt count and last response code per endpoint. And it shows an honest success rate: as long as no delivery has completed, it shows nothing - not a flattering 100 per cent.
Files you can still open in ten years
Finance wants the figures in Excel, the data warehouse wants a file, the auditor wants a format that will still open in ten years. A screenshot from the dashboard meets none of the three.
So every major analysis can be exported as CSV and XLSX, cluster and country reports as a workbook with several sheets. For some areas the CSV file is also available directly through the interface, without a browser, from a script.
Two details decide whether an exported file is safe and usable:
- Hardening against formula injection. A cell starting with an equals sign, plus, minus or at sign is executed as a formula by a spreadsheet. If its content comes from a free-text field - a note, a site name -, opening the file can trigger something nobody intended. Such cells are therefore defused with a leading apostrophe, on every export path.
- Special characters that survive. CSV files carry a UTF-8 marker so that “Zürich” does not turn into gibberish in Excel. XLSX files come with fitted column widths and are readable without rework.
For reports someone reads rather than processes, the server produces genuine PDF files - for clusters, countries, fleet, campaigns and the management summary.
The report that arrives by itself
Whoever needs the same analysis every week opens the same page every week - or forgets. And management does not read dashboards. It reads what is in the inbox.
A scheduled report on site performance or demographics therefore goes out on schedule as PDF, XLSX or CSV to a recipient list, with copy recipients, site selection and send history. Before the first delivery a test send and a sample file can be checked. Every send passes a consent check per recipient, and anyone who unsubscribes drops out of the list without the report ending for everyone else.
The way in is as open as the way out
Openness has a second direction that is almost never asked about: how does data get in? If only the manufacturer knows how a sensor delivers its data, nobody else can connect a gateway, run a test setup or add a second fleet. A closed intake is the second form of lock-in.
So the intake is described just like the exits: a single endpoint per site, a key of its own per site, a fixed size limit per transmission and a validated data structure. How that intake recognises duplicate transmissions and counts atomically is covered in People counting and privacy.
And there is a second intake, for counting systems from other makers. Many sites already have a turnstile, a light barrier or a third-party system. Running a second analysis tool for it means comparing two numbers that never meet. Through an endpoint of its own with a key scope of its own, counts from such sources go into the same minute values as the sensor data - and so into the same series, alerts, exports and comparisons.
One rule prevents double counting: one source per site. A site that already delivers sensor data does not accept a third-party count in addition. Otherwise the same minute counts twice, and nobody knows which.
What belongs in the tender
- Is the entire analysis surface available through an interface - or a single report?
- Is there one response format, stable error codes and a version header?
- Can a key be restricted to specific rights, and does it expire on its own?
- Are events delivered, signed and retried - and for how long?
- Are exports hardened against formula injection?
- Is it documented how data gets in - and can counting systems from other manufacturers be connected?
How we do it
The ANALYSIT Counting System makes its analysis surface available through a REST interface with one response format, more than 40 stable error codes, schema validation and a version header. Keys are restricted to seven scopes, stored only as a hash, expire after one year at most, rotate with a 15-minute overlap and carry their own rate limit.
Events are delivered HMAC-signed to up to five endpoints per tenant, with target-address checks, six attempts over around 26 hours, manual redelivery and a delivery history. Every major analysis exports CSV and XLSX hardened against formula injection; cluster, country, fleet and campaign reports are available as PDF. Scheduled reports go out by email on schedule.
The intake for sensor data is openly documented, and counting systems from other manufacturers can feed the same analysis through an endpoint of their own. For public transport, ITxPT and VDV 301 with checksum are added.
If you want to know how your figures get into your own system, talk to us. The answer should be a key, not a project.
