Skip to content

Systems

The update path is the product

A device that cannot be updated safely is a device with a fixed expiry date, and the date is set by the first security advisory nobody can act on. The update path is decided in the first two weeks of a hardware project and paid for over its whole life.

5 min readComputing America

In short

  • The unit that matters is not the device, it is the truck. A fleet's real update cost is the number of site visits an update requires, and that number is fixed by a design decision made early.
  • An update must be authenticated before it is installed. RFC 9124 names unauthenticated images as a threat whose consequence is complete control of the device, and requires the manifest to be cryptographically bound to the image.
  • Rollback is an attack, not just an inconvenience: a valid but old image with known vulnerabilities is a documented threat, answered by monotonic sequence numbers rather than by hoping nobody keeps the file.
  • An interrupted update must leave a bootable device. The A/B slot design writes to the unused slot while the system runs and boots the old one if the new one fails, which is why it is the default on shipping consumer fleets.
  • The rehearsal that matters is the failed update, not the successful one. A fleet whose recovery path has never been executed does not have one.

Hardware programs are scoped around the first unit and paid for around the two hundredth. The demonstration is a device on a bench doing the thing, and the conversation that follows is about enclosure, cost and lead time. The question that decides what the program costs over five years is usually asked late, quietly, by whoever will be holding it: how does one of these get new software once it is bolted to a wall in another state.

The answer determines whether a security advisory is a Tuesday afternoon or a quarter. It determines whether a fault found in the field can be fixed at all, or whether the workaround becomes permanent because the fix cannot be delivered. And it is not a thing that can be added later at reasonable cost, because two of its four properties are decisions about how storage is partitioned and how the bootloader behaves, taken before the first board is populated.

The unit of cost is the site visit

Before any of the engineering, it is worth stating the economics, because they are what make this an owner’s decision rather than an engineer’s preference. A fleet’s update cost is not the size of the image or the bandwidth. It is the number of times a person has to travel to a device, multiplied by what that travel costs, multiplied by how often software changes over the life of the fleet.

Two hundred devices across four states, updated twice a year, with a technician visit costing what a technician visit costs, is a recurring line that dwarfs the hardware. The same fleet with a working remote path is a person watching a rollout. That difference is decided in the first two weeks of the project and is nearly impossible to renegotiate afterwards, which is exactly the shape of decision that deserves an explicit conversation rather than a default.

Four properties, and the ones nobody adds later

The IETF has already done the enumeration, in public, for constrained devices. RFC 9019 (opens in a new tab) sets out the architecture and RFC 9124 (opens in a new tab) sets out the threats and the requirements that answer them. What follows is those requirements in the order a buyer meets them.

PropertyWhat it preventsWhen it must be decided
AuthenticityInstalling an image from anybody who can reach the device. RFC 9124 names unauthenticated images as a threat whose consequence is complete control, and requires the manifest to be signed and cryptographically bound to the image it describes.Before the bootloader is chosen, because verification happens beneath the operating system.
Rollback resistanceAn attacker replaying a genuine older build with a known vulnerability. The answer is a monotonic sequence number the device refuses to go below, not the absence of old files.With the manifest format, at the start.
AtomicityA power cut or a dropped link leaving an unbootable device. The device must end up running either the old software or the new one, never a mixture.At the storage layout. This is the one that cannot be retrofitted.
RecoverabilityA build that installs correctly and then fails to work, which is the common case and is not a security event at all.With the boot logic, and it needs a health check the device itself can fail.
What an update path has to guarantee

Atomicity is where the flash budget gets spent, and the shipping answer is the A/B slot. Android documents it in terms a buyer can check: updates are applied to the unused slot while the system is running (opens in a new tab), and if an update is applied but fails to boot, the device reboots back into the old partition and remains usable. The stated purpose is exactly the commercial one above, fewer replacements and fewer reflashes at repair centers. The cost is that the device carries two copies of its system, which is a bill of materials decision made on day one.

The offline device, which is the case that is always underestimated

RFC 9124 names a threat worth quoting to anybody who thinks fleet management is a solved problem: a device that has been offline for a long period, presented on reconnection with firmware that is newer than what it runs and older than what is current. Its rollback counter does not stop it, because the image genuinely is newer than the one installed. The suggested answer is expiration in the manifest, checked against a trustworthy time source, which quietly imports a requirement most device designs never wrote down: the device needs to know what day it is, and to know it in a way an adversary cannot set.

This is characteristic of the whole area. Each requirement is individually reasonable and each one lands on a component chosen for another reason entirely. Rollback protection needs persistent monotonic storage. Expiry needs a clock and a source of trusted time. Authenticity needs a place to keep a public key that the running software cannot rewrite. None of the three is expensive when it is on the schematic and all three are structural afterwards.

The rehearsal

The pilot every fleet program runs is the successful update. It is the wrong rehearsal, because success is not the state anybody needs to be prepared for. Before a fleet is committed, we would rather see four failures performed deliberately on real units in the real deployment posture: power removed mid-write, the link dropped mid-transfer, a build that boots and does not work, and a build that is correctly signed by the wrong key.

The value is not that the failures are surprising. It is that each one produces a number: how long until the device is usable again, and whether anybody had to travel. Those two numbers are the fleet’s real operating characteristics, and a program that has never measured them is quoting a support model from a diagram.

Where this argument stops

A real category of device should not be field-updatable at all, and pretending otherwise adds an attack surface for the sake of a principle. A sensor with no writable firmware, doing one job, in a place where a technician visits monthly anyway, is legitimately better as a sealed unit; every update mechanism is also a way in, and RFC 9019 exists partly to enumerate how. The decision is a trade, and the honest form of it is to say which of the two risks the operation would rather carry.

It is also worth saying that an update path is not a substitute for software that is finished. A fleet that requires monthly updates to stay working is not benefiting from a good update path, it is being kept alive by one, and the correct response is to fix the thing that keeps breaking rather than to celebrate the speed of the remedy. The path is insurance. Consuming it constantly means something else is wrong.

Sources

  1. 1.A Firmware Update Architecture for Internet of Things (opens in a new tab), IETF RFC 9019
  2. 2.A Manifest Information Model for Firmware Updates in Internet of Things (IoT) Devices (opens in a new tab), IETF RFC 9124
  3. 3.A/B (seamless) system updates (opens in a new tab), Android Open Source Project

Next step

Tell us how a device gets a new build today.

Describe what happens now when a device in the field needs new software, including who drives where. We will tell you which of the four properties below your path already has, and which one is the cheapest to add next.

Reply
A person replies, not a sequence: within one business day, from someone who would be on the engagement.