The world’s largest multimodal dataset for Signal Foundation Models.
Signal ATLAS is a dataset of radio-frequency signals paired with text annotations. It contains 200 million paired records per domain and is designed to train foundation models that understand electromagnetic signals alongside language.
How declared RF scenes become languageQualification fixture using declared scene metadata. Not a stored Signal ATLAS record.
Qualification context
Illustrative multi-emitter scene · not the stored record
Dense scene · 10 declared emitters · full 1.000 ms duration
Generated annotation
The captured scene contains ten distinct communication signals, all uniformly present throughout the 1.000 millisecond capture duration. Each signal is characterized by an occupied bandwidth of approximately 200 kHz and an estimated signal-to-noise ratio (SNR) of 20.0 dB. The receiver center frequency for this observation was 3.500 GHz, with a sample rate of 2.000 MHz.
Two RF domains with record-level RF and text pairing
The collection provides isolated-signal and multi-emitter records for model training and evaluation. Scientific validation is reported separately from corpus scale.
Dataset definitions
02 validated domains
Signal coverage
66 communications + 14 Counter-UAS classes
Paired-record scale
200M records per domain
Language is paired with RF at record level in both domains.
Dataset definitions and validation
Choose a domain by its training unit and measured evidence. Each definition contains 200 million paired RF-text records; the pilots below establish generated-signal properties, not corpus scale.
01Communications v1
Physics and scene realization measured
A 50,000-record pilot tests causal axes and stored-waveform physics. A separate 20,000-scene pilot measures density, calibration, overlap, boxes, and a falsified occupancy assumption.
Single emitters and multi-emitter scenes · access supported
Specification evidence with a failed transfer test
Published PHY checkpoints, independent synthetic decoding, separability, and scene labels are measured. The controlled synthetic-to-real baseline fails at chance.
Drone links and co-band confusers · access supported
Superpose uses this term for general-purpose models trained on signal measurements and structured context, with representations intended to transfer across classification, detection, and scene-understanding tasks. Signal ATLAS is a dataset and validation program supporting that research objective; it is not a trained model.