ArchitectureAerospace & Defense
Edge AI for DDIL environments: a reference architecture for the tactical edge
In a denied, disrupted, intermittent or limited (DDIL) network, the link to the cloud is a bonus, not a dependency. This guide sets out a reference architecture for tactical-edge AI: what runs on the platform, at a local fusion node and at reach-back, how data synchronizes over a failing link, which degradation states the system must show its operators, and how to test behavior under emulated jamming.
On this page
- What DDIL means operationally and why cloud-first patterns fail
- Reference architecture from sensor to reach-back
- What runs where in a DDIL design, and why
- Fitting models to size, weight and power limits
- Synchronizing data over intermittent links
- Degradation modes the system must show its operators
- Human judgment and autonomy when the link drops
- Updating models across a partly disconnected fleet
- Failure modes to test before fielding
- A maritime ISR detector with hours-long gaps
- Questions and answers
- Sources
What DDIL means operationally and why cloud-first patterns fail
DDIL describes four different conditions. Denied means an adversary is jamming or the platform must stay silent. Disrupted means the link drops without warning. Intermittent means it returns on a schedule nobody controls, such as a satellite pass. Limited means it exists but carries a trickle. A tactical system can move through all four within one mission.
Most enterprise AI quietly assumes the opposite. Inference calls a hosted endpoint, models are pulled from a central registry at start-up, identity checks a remote service, and telemetry streams home continuously. At the edge, each of those assumptions becomes a failure point, and together they produce systems that perform well in a connected demonstration and stop working at exactly the moment they are needed.
So design the zero-connectivity case first. Everything the mission depends on runs locally, with local identity, storage and decision logic; connectivity then speeds up sharing, retraining and oversight instead of being a prerequisite.
Reference architecture from sensor to reach-back
- Sensor payload
Cameras, radar or RF receivers producing raw data at a rate no tactical link could carry.
- On-platform inference
Compact models detect and classify locally, so only results need to travel.
- Local fusion node
A ruggedized node at the edge correlates detections from several platforms into tracks.
- Operator display
Shows tracks, model confidence and the current connectivity state side by side.
- Sync gateway
Queues outbound data by priority and releases it whenever any link is available.
- Reach-back enclave
Central infrastructure for retraining, archive and analysis, plus signing new model packages.
What runs where in a DDIL design, and why
| Function | On the platform | Local fusion node | Reach-back |
|---|---|---|---|
| Detection and classification | Yes: closest to the signal, works with no link | Second-stage checks on uncertain detections | Re-analysis of archived data only |
| Track fusion and correlation | Only for the platform's own sensors | Yes: combines platforms within local radio range | Theater-wide picture when links allow |
| Human review and labeling | Not practical | Quick confirm or reject by operators | Full labeling and adjudication |
| Training and retraining | No | Rarely; only small calibration updates | Yes, with full evaluation before release |
| Model signing and registry | Verifies signatures before loading | Caches approved versions for nearby platforms | Holds signing keys and the master registry |
The fusion node repays most design effort: it keeps a shared local picture alive when reach-back is gone.
Fitting models to size, weight and power limits
Edge hardware is bounded by size, weight and power (SWaP) and by heat. Quantization lowers numeric precision, pruning removes low-value parameters, and distillation trains a small model to imitate a large one. All three trade some accuracy for speed and footprint, and the loss is rarely uniform: it tends to concentrate on rare classes and hard conditions such as low light, clutter or unusual aspect angles, which are often the cases the mission cares about most.
Evaluate the compressed model on those slices specifically, not on the overall score, and keep the full-precision model at reach-back as the reference. The trade-offs of each technique are covered in more depth by ColdAI's edge AI and IoT practice.
Synchronizing data over intermittent links
Treat bandwidth as a budget allocated by mission value, and decide the rules before deployment rather than at the moment a link opens.
- If
A link opens for an unknown, possibly short, window.
ThenSend alerts and new tracks first, then detection metadata, then thumbnails; full-resolution data waits for a long or wideband window.
A short window spent on one image can starve every alert queued behind it.
- If
Two nodes have updated the same track while disconnected.
ThenMerge append-only observation logs with source and timestamp rather than overwriting the track state.
Keeping every observation lets fusion be recomputed; last-writer-wins silently discards evidence.
- If
The outbound queue reaches its storage limit.
ThenDrop or downsample the lowest-priority class first and record what was discarded.
Operators and analysts need to know the archive is incomplete, and where.
- If
Clocks on platforms have drifted during a long disconnection.
ThenCarry a time-quality flag with every record and correct offsets at the fusion node before correlation.
Track correlation built on wrong timestamps produces confident, wrong tracks.
Degradation modes the system must show its operators
| Capability | Full connectivity | Degraded link | No link |
|---|---|---|---|
| Detection | Local, with reach-back re-analysis | Local only | Local only |
| Shared picture | Theater-wide | Local fusion plus delayed updates | Local fusion within radio range |
| Model version | Latest approved release | Pinned; updates deferred | Pinned; age of model displayed |
| Operator indication | Normal status | Banner showing what is delayed | Banner showing data age and missing feeds |
Each state should be visible on the display and recorded in the log, so post-mission review can tell a model error from a connectivity gap.
Human judgment and autonomy when the link drops
Updating models across a partly disconnected fleet
Model updates in DDIL settings arrive unevenly: some platforms receive a release within hours, others weeks later, and some never until they return to base. Treat every model as a signed, versioned package that the platform verifies before loading, keep the previous version on board for rollback, and pin versions explicitly so the fusion node knows which model produced each detection.
Fusion logic must tolerate mixed versions for long periods, and configuration such as thresholds and class lists should travel inside the same signed package so it cannot drift separately. Fleet update mechanics are covered in depth by the edge AI and IoT practice; the defense-specific addition is that a field update is also a change to an accredited system.
Failure modes to test before fielding
Test behavior under realistic conditions, not benchmark accuracy on clean data. Adversarial techniques are catalogued in the NIST adversarial machine learning taxonomy2 and in MITRE ATLAS3.
Spoofed or degraded sensors
Early signalConfident detections that do not agree with other sensors or with physical motion limits.
MitigationCross-check sensors at the fusion node and replay spoofed and degraded inputs in test.
Evasion through camouflage or adversarial patterns
Early signalSharp drops in detection on objects with unusual markings or coverings.
MitigationRed-team with physical and digital evasion attempts and record the results in the evaluation pack.
Poisoned retraining data from the field
Early signalNew labels that shift class boundaries more than expected after a deployment.
MitigationQuarantine field data at reach-back and review it before it can enter a training set.
Priority inversion in the sync queue
Early signalAlerts arriving at reach-back after bulk imagery from the same window.
MitigationEmulate short and lossy link windows in test and assert the order of delivery.
Stale models after long disconnection
Early signalPerformance falls on new target types while the platform cannot receive updates.
MitigationDisplay model age to operators and set a policy for when an old model's outputs need extra confirmation.
A maritime ISR detector with hours-long gaps
Questions and answers
What should we prototype first for a DDIL system?
Prototype the disconnected loop before the model: local inference, the local track store, the priority queue and the operator's state display, all running with the network unplugged. A simple baseline detector is enough at this stage. Once that loop behaves correctly under emulated link loss, improving the model is incremental work rather than a redesign.
Can large language models run at the tactical edge?
Small and compressed language models can run on capable edge hardware for tasks such as summarizing reports or querying local documents. Expect tighter limits on context and quality than a hosted model, and evaluate them on your own vocabulary and formats. For most sensor tasks, purpose-built detection and classification models remain the better fit for size, weight and power budgets.
How do we keep edge models consistent when some nodes are offline for weeks?
Accept that they will not be consistent and design for it. Pin and display versions, attach the model version to every detection, keep fusion logic tolerant of mixed versions, and define a maximum model age after which outputs need extra confirmation. Consistency is restored when platforms reconnect, through signed packages verified on board.
How should success be measured if benchmark accuracy is not enough?
Measure behavior under conditions you expect to meet: detection on difficult slices, time taken to switch degradation states, whether high-priority data arrives first in short link windows, how the system handles spoofed inputs, and whether operators correctly understood the state shown to them. A model with a slightly lower benchmark score that degrades predictably is usually the better field system.
Sources
- DoD Directive 3000.09, Autonomy in Weapon Systems — US Department of Defense, Executive Services Directorate · checked 10 October 2026
- NIST AI 100-2 E2025, Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations — National Institute of Standards and Technology · checked 10 October 2026
- MITRE ATLAS: Adversarial Threat Landscape for Artificial-Intelligence Systems — MITRE · checked 10 October 2026