Superpose

Dataset definition / Counter-UAS v1

Counter-UAS v1

A synthetic dataset definition for representation learning, classification, and scene understanding around drone links and co-band interference. Isolated records contain 16,384 complex samples; the scene configuration composes labeled wideband examples.

ConfigurationsSingle signal · wideband scenes
Validation maturitySpec + decode + scenes validated · transfer failed
Corpus availability200M paired records
Technical reportFormat and accessHugging Face repository · forthcoming

01 / Intended use

Research scope

Counter-UAS captures are scarce and operationally difficult to collect. This synthetic capability provides controlled labels and published-spec structure, while the real-capture comparison shows exactly where that structure is insufficient.

02 / Collection composition

Configurations

Track A · classification

Fixed 16,384-sample windows aligned with Communications v1 for joint pretraining.

Track B · scenes

2–4 emitters per 240,000-sample scene at 160 Msps, spanning a 1.5 ms snapshot with one time–frequency box per signal.

03 / Signal taxonomy

Ten drone classes and four confusers.

Drone linksDroneID ×2FHSS RC ×3Analog FPV ×2Remote IDOcuSync surrogate ×2
ConfusersOFDMQPSKFSKTone

The class balance supports drone-versus-confuser and fine-grained synthetic classification studies. OcuSync and generic confusers remain labeled abstractions, not evidence of protocol decoding.

04 / Generation conditions

Published structure with explicit surrogates.

Window16,384 complex samplesTrack A
Scene span240,000 samples · 160 Msps1.5 ms · ≈80 MHz usable band
NoiseDrawn SNR via AWGNApplied axis
OcuSyncWideband OFDM surrogateNot proprietary reconstruction

Use Track A for isolated-signal learning and the scene track for short wideband detection contexts. Field environments with hardware effects, multipath, and real interference require separate data.

05 / Labels and provenance

Reproduction truth accompanies each generated record.

Records pair numeric I/Q with sibling JSON Lines provenance, including reproduction seed and stored-I/Q binding. Scene records add structured time–frequency boxes. Surrogate status is part of the scientific contract, not hidden in implementation notes.

06 / Text annotations

Signal and scene descriptions

Text annotations identify drone-link and confuser roles, scene composition, time-frequency extents, surrogate status, and realized generation conditions. Protocol surrogates are named explicitly.

Current limitation

A domain-specific Counter-UAS annotation example will be published after it passes the same live qualification. No generated example is presented here as a verified record yet.

The annotation contract is intended to expose roles, relationships, and surrogate boundaries to a model. Until a qualified domain example is published, assess the schema rather than inferred prose quality for this domain.

07 / Validation evidence

Validation summary

Published physical-layer quantities and an independent DroneID receiver support structural fidelity. A 50,000-record diagnostic tests class separability, and a 20,000-scene pilot verifies density, overlap, power closure, and box bounds. Real-capture comparisons remain mixed.

33.3% accuracy · approximately zero real-drone recall.
Specification fidelity
DroneID, FHSS, and analog-video checkpoints measured over thousands of renders
Synthetic DroneID decode
ZC root correlations ≈1.000; wrong-root correlations ≈0.039
Weak 14-class probe
76.8% overall; 93.6% at SNR ≥10 dB; no zero-recall class
Drone / confuser
96.4% binary accuracy in the synthetic pilot
Scene integrity
60,009 / 60,009 boxes valid; power closure p95 error 0.019 dB
Real transfer
33.3% overall; zero recall for both real-drone classes

The real-transfer test uses real FHSS and OcuSync captures but a synthetic AWGN noise class. It is decisive for that controlled baseline, not a complete field benchmark.

Counter-UAS validation checkpoints: four published-spec measurements are within 1.2 percent of target; independent synthetic DroneID decoding gives approximately 1.0 expected-root correlation and 0.039 wrong-root correlation; a weak 14-class probe reaches 76.8 percent overall, 93.6 percent above 10 dB SNR, and 96.4 percent drone-versus-confuser accuracy.
Specification
Four checkpoints within 1.2% of target
Synthetic decode
Expected-root correlation ≈1.000; wrong-root ≈0.039
Weak probe
76.8% overall · 93.6% at SNR ≥10 dB · 96.4% drone/confuser
Figure 03

Specification, decode, and separability checkpoints

Method. Thousands of spec-fidelity renders, independent synthetic DroneID correlation checks, and a 50,000-record weak-probe pilot.

Interpretation. The generated structure is measurable and the 14 classes are diagnostically separable, with harder within-drone boundaries retained.

Does not establish. A full real DroneID decode, device identification, or field-detection performance.

Counter-UAS 20,000-scene pilot: emitter counts are 6,705, 6,581, and 6,714 for two, three, and four emitters; 85.6 percent of scenes contain a drone, 54.1 percent contain overlap, 29.5 percent of emitter pairs overlap, and all 60,009 checked boxes are valid.
Emitter counts
6,705 two-emitter · 6,581 three-emitter · 6,714 four-emitter scenes
Scene structure
85.6% with drone · 54.1% with overlap · 29.5% overlapping pairs
Box validity
60,009 / 60,009 valid
Figure 04

Counter-UAS scene and label realization

Method. 20,000 stored scenes, each 240,000 samples at 160 Msps, with 60,009 emitter boxes checked.

Interpretation. The configured density and class coverage are realized, overlap is measured, and every checked box remains inside time and Nyquist bounds.

Does not establish. Detector-level box utility or pixel-level segmentation-mask correctness.

External-validity boundary

The clean-synthetic baseline did not transfer.

The next result limits the preceding specification, decode, separability, and scene evidence. Those internal checks do not establish performance on real captures.

Three-by-three confusion matrix for 450 balanced windows: all 150 real FHSS windows and 138 of 150 real OcuSync windows were predicted as synthetic noise; 12 OcuSync windows were predicted FHSS; overall accuracy was 33.3 percent.
Overall
33.3% accuracy across 450 balanced windows
Real FHSS
0 / 150 recalled; all predicted as noise
Real OcuSync
0 / 150 recalled; 138 predicted as noise, 12 as FHSS
Figure 05

The controlled synthetic-to-real transfer test failed

Method. Weak classifier trained on synthetic FHSS, OcuSync-surrogate, and AWGN windows; evaluated on real FHSS and OcuSync captures plus synthetic AWGN.

Interpretation. Both real-drone recalls were zero and overall accuracy equaled balanced three-class chance, 33.3 percent.

Does not establish. A permanent transfer limit with real negatives, capture augmentation, hardware effects, or stronger learners.

08 / Known limitations

Evidence boundaries

  • The evaluated clean-synthetic baseline does not support direct field detection.
  • OcuSync is an explicit surrogate, not a proprietary-protocol reconstruction.
  • Generic OFDM is not protocol-exact Wi-Fi.
  • Real spectral fine structure matches only weakly; the real DJI footprint is much wider.
  • A real DJI matched-filter peak was detected, but a full real DroneID decode remains future work.
  • Remote ID over BLE is not distinguishable from generic BLE without payload decoding; Wi-Fi/NaN Remote ID is absent.
  • Scene masks are existence-checked only, and no detector-level box benchmark has been run.

09 / Access and citation

Data format and availability

Signal storage
HDF5 signal arrays with sibling JSONL provenance
Single-signal track
complex64 [N, 16,384], SNR, sample index, class, and per-class sample rate
Scene track
complex64 [N, 240,000] at 160 Msps with scene index
Scene labels
Per-emitter class, role, time-frequency box, seed, and signal binding

ML integration

Fixed arrays can feed batched training directly. The JSONL sidecar carries class, scene geometry, reproducibility, and explicit surrogate status without expanding the signal tensor.

Publication status

The validated definition, negative result, and reproduction path are documented in the technical report. The public dataset repository will publish deterministic splits, license, shard manifest, storage footprint, and citation text on Hugging Face. Access conditions and citation support are available from Superpose while publication is in progress.

Technical reportHugging Face repository · forthcomingRequest access