Dataset definition / Communications v1
Communications v1
A synthetic dataset definition for representation learning across waveform class, propagation, mobility, bandwidth, and hardware conditions. Records contain 16,384 complex baseband samples from one of 66 classes; the scene configuration contains 1–6 emitters with structured labels.
01 / Intended use
Research scope
The dataset is intended for pretraining and controlled benchmarking where class truth and condition metadata are required. Recorded fields are limited to class labels, causally active draws, realized read-backs, or explicitly not-applicable values.
02 / Collection composition
Configurations
66 modulation and PHY classes. Exactly one emitter per 16,384-sample complex-baseband record. SNR from −20 to 30 dB.
1–6 emitters per 262,144-sample scene at 40 Msps, with time–frequency boxes and multi-label masks.
03 / Signal taxonomy
Class families and structured PHYs.
Exhaustive class-by-class analysis remains in progress; the report contains the authoritative evaluated catalog.
Use this taxonomy for broad representation learning across analog, digital, and structured PHY families. Do not assume every named class is separable: the weak probe exposes substantial family-level confusion.
04 / Generation conditions
Condition ranges
Compare these ranges with your deployment envelope. They support controlled robustness studies inside the declared SNR, mobility, bandwidth, and impairment space, not conditions outside it.
05 / Labels and provenance
Record-level provenance
Per-record roles distinguish class labels, drawn axes, derived read-backs, and not-applicable fields. Generation is seed-deterministic from shipped configuration directories and the stock rfgen CLI. The source report also documents content-addressed shard and definition identifiers.
06 / Text annotations
Signal and scene descriptions
Text annotations describe signal identity, scene composition, time-frequency relationships, and realized generation conditions. The narrative is generated from declared scene evidence embedded beside the output, rather than from an unobserved real-world source.
The collection overview contains the canonical declared-scene qualification fixture. Review the RF-to-text qualification example for the generated-language boundary used here.
The text is useful as supervision for declared signal identity, composition, and relationships. It should not be treated as an independent interpretation of an unknown real capture.
07 / Validation evidence
Validation summary
A pre-fix audit found seven corpus defects. Five shared one mechanism: a value was drawn and recorded but ignored. Each defect was pinned, fixed, and rerun through controlled render-twice causality tests. Validation then measured the stored pilot rather than trusting configuration alone.
Post-fix: zero identical pairs on applicable drawn axes in the 25-seed sweep.- Causal axes
- 0 / 25 identical render pairs on every applicable axis
- Delay spread
- KS 0.0087, below the 99% critical value 0.0121
- Cyclic prefix
- 9,120 / 9,120 OFDM records have a nonzero CP
- Weak probe
- 31.3% overall; 42.7% at SNR ≥10 dB
- Probe failures
- 12 zero-recall classes; 33 effective classes at the 30% merge
- Scene labels
- 200 / 200 checked boxes contained in time and frequency
The classifier is a deliberately weak diagnostic, not a benchmark. Its confusion structure shows that high-SNR family boundaries remain difficult and that the current SSB classes behave as proxies.

- Delay spread
- KS 0.0087; 99% critical value 0.0121
- Mobility mix
- 14.9% near-static · 74.9% mobile · 10.2% high-speed
- Time variation
- Median envelope CV 0.060 → 0.094 → 0.131
Communications v1 stored-waveform physics
Method. 50,000-record pilot measured after rendering and storage.
Interpretation. Delay spread and mobility match their designed distributions, while measured time variation rises with speed.
Does not establish. Over-the-air channel realism or receiver performance.

- Emitter counts
- 3,285–3,401 scenes for each count from 1 to 6
- Median occupancy
- 4.1%
- Overlap
- 31.1% of scenes · 8.2% of emitter pairs
Communications v1 scene realization
Method. 20,000 scenes measured from stored outputs.
Interpretation. Emitter density and calibration match the definition; the planned 70/30 occupancy mixture was falsified and removed.
Does not establish. Real spectrum occupancy, box utility, or detector accuracy.
08 / Known limitations
Evidence boundaries
- No over-the-air comparison has been run.
- Documented library channel models remain realism proxies.
- Scenes use a thermal floor and co-channel interference without the single-emitter fading or hardware-impairment chain.
- Scene masks were checked for presence and configured resolution, not pixel-level energy alignment.
- Detector-level validation has not been run for masks or boxes.
- SSB labels are generation proxies; integrated sideband power does not establish physical sideband separation.
09 / Access and citation
Data format and availability
- Signal storage
- Content-addressed HDF5 with complex-valued records
- Record structure
- 16,384-sample records; 262,144-sample scenes at 40 Msps
- Metadata
- Per-record IDs, labels, causal draws, realized read-backs, and definition digest
- Scene labels
- Time-frequency boxes and multi-label masks
ML integration
HDF5 arrays provide the training path. Record identifiers and structured scene labels join each signal to provenance and language sidecars.
Publication status
Use the technical report to assess fitness. The public dataset repository, deterministic splits, license, shard manifest, storage footprint, and citation text will be published together on Hugging Face. Access conditions and citation support are available from Superpose while publication is in progress.