Generation contract
Versioned configuration defines class catalogs, condition ranges, scene composition, labels, and intended use. Seeds and realized values make individual records reproducible.
Methods
ATLAS definitions are generated through a reproducible framework, inspected at stored-output level, and published with limitations intact.
The methodology connects each released record to a declared generation contract, measurements of stored output, and a published account of what the evidence does and does not support. Evaluate a release by reading its definition, pinned source, and validation report together.
Versioned configuration defines class catalogs, condition ranges, scene composition, labels, and intended use. Seeds and realized values make individual records reproducible.
Validation measures rendered records rather than assuming the configuration was realized. Tests cover causal axes, signal properties, scene statistics, labels, and failure cases.
The technical report preserves negative results, source lineage, corpus status, and explicit non-claims. Failed assumptions narrow the release claim instead of disappearing from the record.
These roles tell you which metadata can be used as training truth, which values drove generation, and which values were observed only after rendering.
This comparison is current as of 10 August 2026. “Largest” is measured by paired RF-text records in a released multimodal dataset. One record is a bounded RF example with one or more grounded text annotations. Individual complex-valued signal samples are excluded from the count. Signal ATLAS contains 200 million paired records in each released domain, or 400 million across the two current domains.
| Dataset | Reported scale | Normalization decision |
|---|---|---|
| Signal ATLAS | 400M paired RF-text records | Included. Two released domains at 200M records each. |
| RF-Lang ↗ | 288,000 I/Q-language examples | Included. Direct RF and structured-language pairing. |
| RF-Behavior ↗ | 44 participants across gesture, activity, and emotion tasks | Excluded from normalized count. RF is paired with sensor modalities, not a reported RF-text record corpus. |
| RVTALL ↗ | 20 participants; RF, visual, text, audio, laser, and landmarks | Excluded from normalized count. No released paired-record total comparable to the declared unit. |
| OPERAnet ↗ | Approximately 8 hours of annotated measurements | Excluded from normalized count. RF and vision activity data, not RF-language records. |
| NIST semantic RF proposal ↗ | Proposed collection; no released record count | Excluded. No released corpus at the snapshot date. |
The comparison is versioned by snapshot date because corpus sizes and access conditions change. Different task definitions are not treated as evidence of equivalent scientific scope.