Superpose

Dataset definition / Communications v1

Communications v1

A synthetic dataset definition for representation learning across waveform class, propagation, mobility, bandwidth, and hardware conditions. Records contain 16,384 complex baseband samples from one of 66 classes; the scene configuration contains 1–6 emitters with structured labels.

ConfigurationsSingle emitter · wideband scenes
Validation maturityDefinition and pilots validated
Corpus availability200M paired records
Technical reportFormat and accessHugging Face repository · forthcoming

01 / Intended use

Research scope

The dataset is intended for pretraining and controlled benchmarking where class truth and condition metadata are required. Recorded fields are limited to class labels, causally active draws, realized read-backs, or explicitly not-applicable values.

02 / Collection composition

Configurations

Single emitter

66 modulation and PHY classes. Exactly one emitter per 16,384-sample complex-baseband record. SNR from −20 to 30 dB.

Wideband scenes

1–6 emitters per 262,144-sample scene at 40 Msps, with time–frequency boxes and multi-label masks.

03 / Signal taxonomy

Class families and structured PHYs.

AnalogAMFMCW
Digital modulationFSKPSKQAMAPSKOQPSK
Structured PHYOFDM5G NR PUSCHFEC-coded constellationChirp spread spectrum

Exhaustive class-by-class analysis remains in progress; the report contains the authoritative evaluated catalog.

Use this taxonomy for broad representation learning across analog, digital, and structured PHY families. Do not assume every named class is separable: the weak probe exposes substantial family-level confusion.

04 / Generation conditions

Condition ranges

SNR−20…30 dBDrawn axis
Channel6 familiesDrawn axis
Mobility0…300 m/sApplicable to TDL/CDL
Bandwidth0.2 / 1 / 5 / 20 MHzDrawn axis
HardwarePA, quantization, IQ, phase noise, AGCDrawn subset

Compare these ranges with your deployment envelope. They support controlled robustness studies inside the declared SNR, mobility, bandwidth, and impairment space, not conditions outside it.

05 / Labels and provenance

Record-level provenance

Per-record roles distinguish class labels, drawn axes, derived read-backs, and not-applicable fields. Generation is seed-deterministic from shipped configuration directories and the stock rfgen CLI. The source report also documents content-addressed shard and definition identifiers.

06 / Text annotations

Signal and scene descriptions

Text annotations describe signal identity, scene composition, time-frequency relationships, and realized generation conditions. The narrative is generated from declared scene evidence embedded beside the output, rather than from an unobserved real-world source.

The collection overview contains the canonical declared-scene qualification fixture. Review the RF-to-text qualification example for the generated-language boundary used here.

The text is useful as supervision for declared signal identity, composition, and relationships. It should not be treated as an independent interpretation of an unknown real capture.

07 / Validation evidence

Validation summary

A pre-fix audit found seven corpus defects. Five shared one mechanism: a value was drawn and recorded but ignored. Each defect was pinned, fixed, and rerun through controlled render-twice causality tests. Validation then measured the stored pilot rather than trusting configuration alone.

Post-fix: zero identical pairs on applicable drawn axes in the 25-seed sweep.
Causal axes
0 / 25 identical render pairs on every applicable axis
Delay spread
KS 0.0087, below the 99% critical value 0.0121
Cyclic prefix
9,120 / 9,120 OFDM records have a nonzero CP
Weak probe
31.3% overall; 42.7% at SNR ≥10 dB
Probe failures
12 zero-recall classes; 33 effective classes at the 30% merge
Scene labels
200 / 200 checked boxes contained in time and frequency

The classifier is a deliberately weak diagnostic, not a benchmark. Its confusion structure shows that high-SNR family boundaries remain difficult and that the current SSB classes behave as proxies.

Communications physics pilot: delay-spread KS statistic 0.0087 is below the 99 percent critical value 0.0121; mobility realizes 14.9, 74.9, and 10.2 percent against a 15, 75, and 10 percent design; received-power variation increases from 0.060 to 0.094 to 0.131 across speed bands.
Delay spread
KS 0.0087; 99% critical value 0.0121
Mobility mix
14.9% near-static · 74.9% mobile · 10.2% high-speed
Time variation
Median envelope CV 0.060 → 0.094 → 0.131
Figure 01

Communications v1 stored-waveform physics

Method. 50,000-record pilot measured after rendering and storage.

Interpretation. Delay spread and mobility match their designed distributions, while measured time variation rises with speed.

Does not establish. Over-the-air channel realism or receiver performance.

Communications scene pilot: emitter counts range from 3,285 to 3,401 across one to six emitters; median time-frequency occupancy is 4.1 percent, 31.1 percent of scenes contain overlap, and 8.2 percent of emitter pairs overlap.
Emitter counts
3,285–3,401 scenes for each count from 1 to 6
Median occupancy
4.1%
Overlap
31.1% of scenes · 8.2% of emitter pairs
Figure 02

Communications v1 scene realization

Method. 20,000 scenes measured from stored outputs.

Interpretation. Emitter density and calibration match the definition; the planned 70/30 occupancy mixture was falsified and removed.

Does not establish. Real spectrum occupancy, box utility, or detector accuracy.

08 / Known limitations

Evidence boundaries

  • No over-the-air comparison has been run.
  • Documented library channel models remain realism proxies.
  • Scenes use a thermal floor and co-channel interference without the single-emitter fading or hardware-impairment chain.
  • Scene masks were checked for presence and configured resolution, not pixel-level energy alignment.
  • Detector-level validation has not been run for masks or boxes.
  • SSB labels are generation proxies; integrated sideband power does not establish physical sideband separation.

09 / Access and citation

Data format and availability

Signal storage
Content-addressed HDF5 with complex-valued records
Record structure
16,384-sample records; 262,144-sample scenes at 40 Msps
Metadata
Per-record IDs, labels, causal draws, realized read-backs, and definition digest
Scene labels
Time-frequency boxes and multi-label masks

ML integration

HDF5 arrays provide the training path. Record identifiers and structured scene labels join each signal to provenance and language sidecars.

Publication status

Use the technical report to assess fitness. The public dataset repository, deterministic splits, license, shard manifest, storage footprint, and citation text will be published together on Hugging Face. Access conditions and citation support are available from Superpose while publication is in progress.

Technical reportHugging Face repository · forthcomingRequest access