Define
Declare class catalogs, condition ranges, scene composition, labels, and intended use before production scale.
Methods
ATLAS definitions are generated through a reproducible framework, inspected at stored-output level, and published with limitations intact.
Declare class catalogs, condition ranges, scene composition, labels, and intended use before production scale.
Use shipped configuration directories and the stock rfgen CLI. Record seeds, realized values, and provenance roles.
Test causal axes, physics, scene statistics, separability, robustness, and—where evidence exists—real-capture behavior.
Pin defects, rerun experiments, preserve negative results, and narrow claims when assumptions fail.
Publish the definition, evidence, source pin, corpus status, and explicit non-claims together.
Provenance roles
Label Class or scene truth.
Drawn axis A sampled value proven to affect rendering.
Realized read-back A value measured after rendering.
Not applicable An explicit absence—not a plausible but inert number.
Corpus scale / 10 August 2026
“Largest” is measured by paired RF-text records in a released multimodal dataset. One record is a bounded RF example with one or more grounded text annotations. Individual complex-valued signal samples are excluded from the count. Signal ATLAS contains 200 million paired records in each released domain, or 400 million across the two current domains.
| Dataset | Reported scale | Normalization decision |
|---|---|---|
| Signal ATLAS | 400M paired RF-text records | Included. Two released domains at 200M records each. |
| RF-Lang ↗ | 288,000 I/Q-language examples | Included. Direct RF and structured-language pairing. |
| RF-Behavior ↗ | 44 participants across gesture, activity, and emotion tasks | Excluded from normalized count. RF is paired with sensor modalities, not a reported RF-text record corpus. |
| RVTALL ↗ | 20 participants; RF, visual, text, audio, laser, and landmarks | Excluded from normalized count. No released paired-record total comparable to the declared unit. |
| OPERAnet ↗ | Approximately 8 hours of annotated measurements | Excluded from normalized count. RF and vision activity data, not RF-language records. |
| NIST semantic RF proposal ↗ | Proposed collection; no released record count | Excluded. No released corpus at the snapshot date. |
The comparison is versioned by snapshot date because corpus sizes and access conditions change. Different task definitions are not treated as evidence of equivalent scientific scope.