Skip to content

Engineering

Relative to what? The question that decides whether your data survives

A measurement is a number and a frame. Instrumentation projects reliably ship the number and leave the frame implicit, and a frame that was never recorded cannot be recovered afterwards by anyone, at any price.

8 min readComputing America

In short

  • A number without its frame is not a measurement, it is a reading. Position relative to wherever the device powered on is internally consistent and mutually incomparable, and no post-processing recovers a reference that was never written down.
  • Every instrumented system has four frames (spatial, temporal, scale and identity) and each one is a decision. Left implicit, each is decided by whichever component happened to be written first.
  • Timestamp at the sensor, not at the collector. A timestamp applied on arrival silently encodes queue depth as measurement error, and it stays monotonic while it does it, so nothing looks wrong.
  • “We will normalize it later” is the most expensive sentence in an instrumentation project. Data you can re-derive is a processing problem; a frame you did not capture is not recoverable at any budget.
  • When the defect is in somebody else’s firmware, fixing it upstream costs more this month and less forever: a local workaround becomes part of your frame, and nobody remembers it is there.

An autonomous vehicle test range was producing exactly the data its program needed and keeping none of it in a usable form. The vehicle reported its position and rotation continuously, and it reported them relative to wherever that vehicle happened to be sitting when it powered on.

Every run was therefore internally perfect. The trajectory was smooth, the numbers were consistent, and anyone looking at a single run would have said the instrumentation was working. Two runs of the same maneuver produced two coordinate frames with nothing in common, and the fix, referencing the whole rig to magnetic north, is not a signal-processing improvement. It is the difference between a log and a dataset.

That failure has a general shape and it is the most expensive one in this discipline, because it is silent while it is happening and permanent once it has happened. A measurement is a number and a frame. Ship the number, leave the frame implicit, and you have built something that looks correct on the day and is worth nothing the moment somebody wants to compare two of its outputs.

Four frames, and what happens when each is left implicit

Every instrumented system carries four, whether or not anyone decided them. Unstated, each one gets decided by whichever component was written first, which is a decision, made by nobody, that the rest of the system then inherits.

FrameThe question it answersWhat an implicit answer costs
SpatialRelative to what origin, along which axes, in which handednessThe failure above. Runs that cannot be compared to each other, to a map, or to the physical thing they happened on, and it is unrecoverable, because the reference existed only at capture time.
TemporalWhich clock, at what resolution, stamped whereTwo sensors that disagree by an amount nobody can characterize, because one is stamped at the device and one on arrival. Covered on its own below; it is the most common of the four.
ScaleRaw counts or engineering units, and where the conversion happensA calibration constant applied twice, or not at all, discovered when a value is off by exactly the gain. Worse: a conversion that lives in a script rather than in the record, so historical data cannot be re-derived when the constant is corrected.
IdentityWhich physical device produced this, and does its id survive replacementA sensor swapped under warranty, the id reused, and eighteen months of data attributed to a device that has been in a drawer since March. Every drift analysis over that period is now fiction.
The four frames every measurement carries

Notice what those four failures have in common. None of them produces an error, an alarm or a gap. The system keeps running, the numbers keep looking plausible, and the damage is only visible at the moment somebody tries to do the analysis the system was built for, which is usually months later, and usually by somebody who was not there.

The temporal frame, because it is the one everybody gets wrong

The common architecture stamps a reading when it arrives at the collector. It is easy, it needs no clock on the device, and it is wrong in a way that hides itself: the value recorded is not when the measurement was taken but when the collector got round to it, so queue depth is silently encoded as measurement error.

What makes it pernicious rather than merely inaccurate is that the result stays monotonic. Timestamps still increase. Nothing in the data looks corrupted. The error is proportional to how busy the system was, which means it is smallest when nothing interesting is happening and largest during exactly the events the instrumentation was installed to capture. An intermittent fault on a machine is characterized by the microsecond relationship between two signals; measure that through a shared queue under load and you can produce a confident, repeatable, wrong answer about which one came first.

The corollary is that clock discipline is part of the instrument, not part of the network. A device whose clock is set once at commissioning and then drifts is producing timestamps whose accuracy degrades continuously, and nothing downstream can tell. If the analysis needs sub-second alignment between two devices, that requirement lands on the hardware design, and it is far cheaper to notice while the enclosure is still on the bench.

Why “we will normalize it later” is the expensive sentence

Deferring normalization is reasonable in most software contexts. Data engineering routinely fixes shapes, fills gaps and reconciles schemas after the fact, and the instinct that a messy pipeline can be cleaned downstream is usually correct. It fails here for a specific reason: the two problems are different in kind.

Data you can re-derive is a processing problem, and processing problems yield to time and money. A frame you did not capture is not a processing problem; it is information that never existed. There is no budget, no model and no amount of cleverness that recovers which direction the vehicle was pointing when a run began if nothing recorded it. The only place that fact was available was the moment of capture, and that moment is gone.

This is the argument for spending the first week of an instrumentation project on things that produce no visible output. Deciding the origin, the axis convention, the units, the clock and the identity scheme is unglamorous, and there is real pressure to defer it in favor of getting a reading on a screen, because a reading on a screen is what everyone can see. The reading is the easy half and it is reversible. The frame is the hard half and it is not.

Metrology settled this vocabulary long before instrumentation projects existed, and it is worth borrowing because it says what to write down. NIST defines metrological traceability (opens in a new tab) as a property of a measurement result, not of an instrument or a laboratory: the result has to be relatable to a stated reference through a documented unbroken chain of calibrations, each contributing to the uncertainty. A calibrated sensor does not make a number traceable. A recorded reference does, and the record is the deliverable.

When the defect is in somebody else’s firmware

Instrumentation work is nearly always work against a device somebody else built, and eventually the device is wrong. On that vehicle platform it was pulse-width modulation faults that only surfaced under real test conditions, not on a bench, not in the vendor’s lab, and not in any way that could be reproduced by describing it.

There are two available responses, and the cheap one this month is the expensive one over the system’s life. Compensating locally is fast: you know the shape of the defect, you correct for it in your own layer, and the readings come out right. What you have actually done is fold somebody else’s bug into your frame. It is now part of how your system interprets the world; it is documented nowhere the vendor can see, and it will be discovered by whoever is holding the system on the day the vendor ships a firmware update that fixes the original defect, at which point your correction is applied to data that no longer needs it, and everything is wrong in a new direction.

Diagnosing it and contributing the fix upstream costs more the first time and nothing thereafter. The fix belongs to the device, it survives your involvement, and, the part that is easy to undervalue, it means your system’s behavior can still be explained by reading the two things it is made of, rather than by knowing a story about a bug in 2024.

The test to run before anything is built

One question settles most of this, and it can be asked at a whiteboard before a single component is specified. Take two measurements from this system: one taken today, one taken a year from now, after the device has been replaced under warranty and the firmware has been updated twice. Can they be plotted on the same axes, and would anybody be able to defend that plot?

  1. 1.If the answer is no, name the missing frame. It will be one of the four, and it will be the one nobody has written down anywhere.
  2. 2.For each frame, say where it is recorded. Not where it is known, where it is written, in the record, beside the value it applies to.
  3. 3.For the spatial frame specifically, say what physical thing the origin is attached to and what happens if that thing moves. An origin defined by a device’s own power-on state is not an origin.
  4. 4.For units, put the conversion in the record rather than in the code that reads it, so a corrected constant can be applied to history rather than only to the future.
  5. 5.For identity, decide now what happens when a device is replaced, and prefer a scheme where the new unit gets a new id. Reusing one is the cheapest possible way to corrupt a year of trend analysis.

None of that requires knowing which sensor you are buying, and all of it is harder to change afterwards than the sensor is.

Where this argument stops

Frame discipline is necessary and it is not sufficient. It does nothing about a sensor sited where it cannot see the phenomenon, a signal conditioned badly, or an instrument specified for a range the process routinely exceeds. A perfectly framed measurement of the wrong thing is still the wrong thing, and there is a version of this argument that becomes an excuse for a metadata schema in place of a working instrument.

But those failures are visible. Somebody looks at the data and it is obviously wrong, and the problem gets fixed, because a system producing obviously wrong readings does not get handed over. The frame failures are the ones that get handed over, run for two years, and are discovered by the person finally asked to answer the question the whole thing was installed to answer.

Sources

  1. 1.Metrological Traceability: Frequently Asked Questions and NIST Policy (opens in a new tab), NIST

Next step

Send us two runs.

Two exports of the same thing measured on different days, or one export and the schema behind it. We will tell you which of the four frames is missing and whether it can still be recovered.

Reply
A person replies, not a sequence: within one business day, from someone who would be on the engagement.