ProcessEdge AI & IoT

Over-the-air model updates for edge AI fleets, from signed bundle to rollback

Shipping OTA model updates to an edge AI fleet is a release-engineering problem first: the model travels with preprocessing code, a runtime and firmware it must match, over links that drop. A safe process versions all of that as one signed bundle, proves it in shadow mode on a small cohort, widens exposure only when health gates pass, and keeps an automatic way back on every device.

Reviewed 7 min read

On this page
  1. Why a new model is a fleet operation, not a notebook result
  2. Who sends what during one staged model release
  3. Seven stages of a safe edge model release
  4. Health gates to read at shadow, canary and fleet-wide stages
  5. Halt, exclude or proceed: reading a rollout that looks odd
  6. Telemetry and drift when links come and go
  7. Security baselines that shape how devices receive updates
  8. Questions and answers
  9. Sources

Why a new model is a fleet operation, not a notebook result

A model that scores better offline can still make a fleet worse. It may need an operator the older runtime lacks, use more memory than devices with a busy firmware build can spare, or behave differently at sites whose sensors were never in the training data. And unlike a cloud deployment, you cannot simply redeploy in seconds: some devices are asleep, some are behind a cellular link with a data cap, and some will lose power halfway through the download.

That is why the hub's third stage, operating a fleet, is about secure updates, version tracking and delayed synchronization5. The process below treats every model release the way firmware teams treat a firmware release: as an artifact with a manifest, a signature, a staged rollout and a tested path back.

Who sends what during one staged model release

signed bundleoffer to cohortverify, check compatinstall to slot Bshadow resultsgate metricspromote or haltactivate or revert01Model registry02Update service03Edge device04Telemetry store05Release owner
  1. Model registry

    Holds approved bundles and their manifests.

  2. Update service

    Targets cohorts and tracks which bundle each device runs.

  3. Edge device

    Verifies, installs to an inactive slot and reports.

  4. Telemetry store

    Receives buffered results when links return.

  5. Release owner

    Reads gate metrics and promotes or halts.

  1. Model registry to Update servicesigned bundle
  2. Update service to Edge deviceoffer to cohort
  3. Edge device to Edge deviceverify, check compat
  4. Edge device to Edge deviceinstall to slot B
  5. Edge device to Telemetry storeshadow results
  6. Telemetry store to Release ownergate metrics
  7. Release owner to Update servicepromote or halt
  8. Update service to Edge deviceactivate or revert
Conceptual message sequence for one cohort of a staged model release. It illustrates a design pattern, not a specific product or deployment.

Seven stages of a safe edge model release

  1. Define the bundle and its compatibility

    Package the model file with the preprocessing code that produces its inputs, the runtime version it was compiled for, the minimum firmware and the hardware revisions it was tested on. A model without its preprocessing is the most common source of silent errors after an update.

    Output
    Release manifest with a unique bundle version
    Owner
    ML engineer with firmware lead
  2. Sign the bundle and protect the channel

    Sign the manifest and artifacts with keys held outside the build system, and have the device check the signature against a trust anchor installed before deployment. The IETF firmware update architecture describes this manifest-and-signature model, including sequence numbers that stop a device being tricked into installing an older, vulnerable image1.

    Output
    Signed bundle in the registry
    Owner
    Release manager and security
  3. Prove it on a bench fleet

    Install on at least one device of every hardware revision in service, with production firmware. Interrupt power and network mid-install on purpose, and confirm the device boots the previous slot.

    Output
    Install, boot and recovery test record
    Owner
    Test engineer
  4. Shadow on a small cohort

    Run the new model beside the current one on a few representative devices. The current model still decides; the new one only logs its outputs, latency and memory so you can compare both on live inputs.

    Output
    Shadow comparison by site and condition
    Owner
    ML engineer
  5. Activate on a canary cohort

    Let the new model decide on a slightly larger group that spans sites, hardware revisions and connectivity types. Agree the health gates and the halt criteria before activation, not after the first odd chart.

    Output
    Gate review with promote or halt decision
    Owner
    Release owner
  6. Ramp by cohort

    Widen exposure in steps, pausing after each for enough operating time to include the conditions that matter: night shifts, weekend loads, cold starts. Devices that were offline join later cohorts rather than receiving the newest bundle on reconnection unannounced.

    Output
    Fleet inventory showing bundle per device
    Owner
    Operations
  7. Close out and keep watching

    Retire the previous bundle only once rollback is no longer needed, and keep comparing input distributions per site so drift is noticed before outcomes degrade.

    Output
    Release record and drift baseline
    Owner
    ML engineer with operations

Health gates to read at shadow, canary and fleet-wide stages

SignalShadowCanaryFleet-wide
Install and bootEvery device verifies and installs to the inactive slotNo device stuck between slotsFailed installs revert automatically and are counted
Latency and memoryNew model within budget alongside the old oneWithin budget under real loadWatched per hardware revision
Agreement with current modelDisagreements reviewed by a domain expertDisagreements traced to expected improvementsTracked as a trend, not a gate
Field outcomesNot yet measurableConfirmed events compared with the previous periodFeeds the next training set
Telemetry completenessBuffered results arrive from every shadow deviceSilent devices are chased, not assumed healthyCoverage reported alongside every metric

Thresholds depend on the decision the device makes. Set them with the people who own that decision.

Halt, exclude or proceed: reading a rollout that looks odd

  • If

    Devices fail signature checks or cannot boot the new slot.

    Then

    Halt the release and investigate the build and keys before anything else.

    Signature or boot failures point to the artifact or the update channel, which affects every device, not just one site.

  • If

    Latency exceeds budget on one hardware revision only.

    Then

    Exclude that revision from the release and ship it a variant compiled for its runtime.

    Mixed hardware generations are normal in long-lived fleets; one bundle rarely suits all of them.

  • If

    Disagreement with the old model clusters at a few sites.

    Then

    Inspect the inputs from those sites before judging the model.

    A sensor change, mounting difference or new operating mode often explains it, and it is evidence for the next training set.

  • If

    A cohort reports almost nothing after activation.

    Then

    Pause promotion until telemetry returns or devices are checked in person.

    In intermittent networks silence is missing evidence, not a sign that everything works.

Security baselines that shape how devices receive updates

Regulation (EU) 2024/2847 (Cyber Resilience Act)

EU

Applies whenProducts with digital elements are made available on the EU market; most obligations apply from 11 December 2027 and the Article 14 reporting obligations from 11 September 20264.

  • Products must be able to receive security updates, with automatic security updates enabled by default where applicable and a way to opt out4.
  • Manufacturers must provide mechanisms to distribute updates securely and, where technically feasible, ship security updates separately from functionality updates4.

ETSI EN 303 645 V3.1.3, Cyber Security for Consumer Internet of Things: Baseline Requirements[^3]

European standard, used internationally

Applies whenA manufacturer designs network-connected consumer IoT devices and chooses to demonstrate conformance with this baseline3.

  • The device must have a secure update mechanism unless resource constraints make updating impossible3.
  • Where updates arrive over a network interface, the device must verify the authenticity and integrity of each update via a trust relationship3.

NIST IR 8259 Rev. 1, Foundational Cybersecurity Activities for IoT Product Manufacturers[^2]

US (voluntary guidance)

Applies whenA manufacturer plans pre-market and post-market cybersecurity activities for an IoT product2.

  • Recommends building cybersecurity, including how products will be maintained after sale, into the development process rather than adding it afterwards2.

Questions and answers

How often should edge AI models be updated?

On evidence, not on a calendar. Update when drift monitoring or confirmed outcomes show the current model is degrading, when a new failure mode needs to be recognized, or when a security fix in the runtime forces a rebuild. Security updates to firmware and runtimes follow their own, faster path and should not wait for a model release.

How do we update models over cellular links with tight data limits?

Send binary deltas against the installed bundle where your update tooling supports them, schedule downloads for off-peak windows, and quantize models so the artifact itself is small. Let devices resume interrupted downloads rather than restart them, and verify the reassembled bundle's signature before installing anything.

What if our fleet mixes several hardware generations?

Treat each hardware revision as a separate release target with its own compiled bundle and test devices, tracked under one logical model version. The manifest should state which revisions a bundle supports, and the update service should refuse to offer it to anything else.

Is shadow mode always possible on constrained devices?

Not always. If a device cannot hold two models in memory, run shadow evaluation on a small set of higher-specification units at representative sites, or replay recorded field inputs through the new bundle on bench hardware. Both are weaker evidence than on-device shadowing, so keep the canary cohort small for longer.

Sources

  1. RFC 9019: A Firmware Update Architecture for Internet of Things — IETF / RFC Editor · checked 10 October 2026
  2. NIST IR 8259 Rev. 1: Foundational Cybersecurity Activities for IoT Product Manufacturers — NIST · checked 10 October 2026
  3. ETSI EN 303 645 V3.1.3 (2024-09) Cyber Security for Consumer Internet of Things: Baseline Requirements — ETSI · checked 10 October 2026
  4. Regulation (EU) 2024/2847 on horizontal cybersecurity requirements for products with digital elements (Cyber Resilience Act) — EUR-Lex · checked 10 October 2026
  5. Edge AI & IoT — ColdAI

More in Edge AI & IoT

Back to Edge AI & IoT

Next step

Map your current update path before the next model ships

Send a short description of your fleet: hardware revisions, connectivity and how updates reach devices today. We will reply with the gaps we would close first and the gates we would add.

Review my update path