Dataset definition / Counter-UAS v1
Counter-UAS v1
A synthetic dataset definition for representation learning, classification, and scene understanding around drone links and co-band interference. Isolated records contain 16,384 complex samples; the scene configuration composes labeled wideband examples.
01 / Intended use
Research scope
Counter-UAS captures are scarce and operationally difficult to collect. This synthetic capability provides controlled labels and published-spec structure, while the real-capture comparison shows exactly where that structure is insufficient.
02 / Collection composition
Configurations
Fixed 16,384-sample windows aligned with Communications v1 for joint pretraining.
2–4 emitters per 240,000-sample scene at 160 Msps, spanning a 1.5 ms snapshot with one time–frequency box per signal.
03 / Signal taxonomy
Ten drone classes and four confusers.
The class balance supports drone-versus-confuser and fine-grained synthetic classification studies. OcuSync and generic confusers remain labeled abstractions, not evidence of protocol decoding.
04 / Generation conditions
Published structure with explicit surrogates.
Use Track A for isolated-signal learning and the scene track for short wideband detection contexts. Field environments with hardware effects, multipath, and real interference require separate data.
05 / Labels and provenance
Reproduction truth accompanies each generated record.
Records pair numeric I/Q with sibling JSON Lines provenance, including reproduction seed and stored-I/Q binding. Scene records add structured time–frequency boxes. Surrogate status is part of the scientific contract, not hidden in implementation notes.
06 / Text annotations
Signal and scene descriptions
Text annotations identify drone-link and confuser roles, scene composition, time-frequency extents, surrogate status, and realized generation conditions. Protocol surrogates are named explicitly.
A domain-specific Counter-UAS annotation example will be published after it passes the same live qualification. No generated example is presented here as a verified record yet.
The annotation contract is intended to expose roles, relationships, and surrogate boundaries to a model. Until a qualified domain example is published, assess the schema rather than inferred prose quality for this domain.
07 / Validation evidence
Validation summary
Published physical-layer quantities and an independent DroneID receiver support structural fidelity. A 50,000-record diagnostic tests class separability, and a 20,000-scene pilot verifies density, overlap, power closure, and box bounds. Real-capture comparisons remain mixed.
33.3% accuracy · approximately zero real-drone recall.- Specification fidelity
- DroneID, FHSS, and analog-video checkpoints measured over thousands of renders
- Synthetic DroneID decode
- ZC root correlations ≈1.000; wrong-root correlations ≈0.039
- Weak 14-class probe
- 76.8% overall; 93.6% at SNR ≥10 dB; no zero-recall class
- Drone / confuser
- 96.4% binary accuracy in the synthetic pilot
- Scene integrity
- 60,009 / 60,009 boxes valid; power closure p95 error 0.019 dB
- Real transfer
- 33.3% overall; zero recall for both real-drone classes
The real-transfer test uses real FHSS and OcuSync captures but a synthetic AWGN noise class. It is decisive for that controlled baseline, not a complete field benchmark.

- Specification
- Four checkpoints within 1.2% of target
- Synthetic decode
- Expected-root correlation ≈1.000; wrong-root ≈0.039
- Weak probe
- 76.8% overall · 93.6% at SNR ≥10 dB · 96.4% drone/confuser
Specification, decode, and separability checkpoints
Method. Thousands of spec-fidelity renders, independent synthetic DroneID correlation checks, and a 50,000-record weak-probe pilot.
Interpretation. The generated structure is measurable and the 14 classes are diagnostically separable, with harder within-drone boundaries retained.
Does not establish. A full real DroneID decode, device identification, or field-detection performance.

- Emitter counts
- 6,705 two-emitter · 6,581 three-emitter · 6,714 four-emitter scenes
- Scene structure
- 85.6% with drone · 54.1% with overlap · 29.5% overlapping pairs
- Box validity
- 60,009 / 60,009 valid
Counter-UAS scene and label realization
Method. 20,000 stored scenes, each 240,000 samples at 160 Msps, with 60,009 emitter boxes checked.
Interpretation. The configured density and class coverage are realized, overlap is measured, and every checked box remains inside time and Nyquist bounds.
Does not establish. Detector-level box utility or pixel-level segmentation-mask correctness.
External-validity boundary
The clean-synthetic baseline did not transfer.
The next result limits the preceding specification, decode, separability, and scene evidence. Those internal checks do not establish performance on real captures.

- Overall
- 33.3% accuracy across 450 balanced windows
- Real FHSS
- 0 / 150 recalled; all predicted as noise
- Real OcuSync
- 0 / 150 recalled; 138 predicted as noise, 12 as FHSS
The controlled synthetic-to-real transfer test failed
Method. Weak classifier trained on synthetic FHSS, OcuSync-surrogate, and AWGN windows; evaluated on real FHSS and OcuSync captures plus synthetic AWGN.
Interpretation. Both real-drone recalls were zero and overall accuracy equaled balanced three-class chance, 33.3 percent.
Does not establish. A permanent transfer limit with real negatives, capture augmentation, hardware effects, or stronger learners.
08 / Known limitations
Evidence boundaries
- The evaluated clean-synthetic baseline does not support direct field detection.
- OcuSync is an explicit surrogate, not a proprietary-protocol reconstruction.
- Generic OFDM is not protocol-exact Wi-Fi.
- Real spectral fine structure matches only weakly; the real DJI footprint is much wider.
- A real DJI matched-filter peak was detected, but a full real DroneID decode remains future work.
- Remote ID over BLE is not distinguishable from generic BLE without payload decoding; Wi-Fi/NaN Remote ID is absent.
- Scene masks are existence-checked only, and no detector-level box benchmark has been run.
09 / Access and citation
Data format and availability
- Signal storage
- HDF5 signal arrays with sibling JSONL provenance
- Single-signal track
- complex64 [N, 16,384], SNR, sample index, class, and per-class sample rate
- Scene track
- complex64 [N, 240,000] at 160 Msps with scene index
- Scene labels
- Per-emitter class, role, time-frequency box, seed, and signal binding
ML integration
Fixed arrays can feed batched training directly. The JSONL sidecar carries class, scene geometry, reproducibility, and explicit surrogate status without expanding the signal tensor.
Publication status
The validated definition, negative result, and reproduction path are documented in the technical report. The public dataset repository will publish deterministic splits, license, shard manifest, storage footprint, and citation text on Hugging Face. Access conditions and citation support are available from Superpose while publication is in progress.