Superpose

Methods

Generation and validation

ATLAS definitions are generated through a reproducible framework, inspected at stored-output level, and published with limitations intact.

What makes an ATLAS definition reproducible

The methodology connects each released record to a declared generation contract, measurements of stored output, and a published account of what the evidence does and does not support. Evaluate a release by reading its definition, pinned source, and validation report together.

01

Generation contract

Versioned configuration defines class catalogs, condition ranges, scene composition, labels, and intended use. Seeds and realized values make individual records reproducible.

02

Stored-output verification

Validation measures rendered records rather than assuming the configuration was realized. Tests cover causal axes, signal properties, scene statistics, labels, and failure cases.

03

Published evidence boundary

The technical report preserves negative results, source lineage, corpus status, and explicit non-claims. Failed assumptions narrow the release claim instead of disappearing from the record.

Provenance field roles

These roles tell you which metadata can be used as training truth, which values drove generation, and which values were observed only after rendering.

Label
Class or scene truth.
Drawn axis
A sampled value proven to affect rendering.
Realized read-back
A value measured after rendering.
Not applicable
An explicit absence, not a plausible but inert number.

Basis for the largest-dataset claim

This comparison is current as of 10 August 2026. “Largest” is measured by paired RF-text records in a released multimodal dataset. One record is a bounded RF example with one or more grounded text annotations. Individual complex-valued signal samples are excluded from the count. Signal ATLAS contains 200 million paired records in each released domain, or 400 million across the two current domains.

RF multimodal dataset comparison, snapshot 10 August 2026
DatasetReported scaleNormalization decision
Signal ATLAS400M paired RF-text recordsIncluded. Two released domains at 200M records each.
RF-Lang ↗288,000 I/Q-language examplesIncluded. Direct RF and structured-language pairing.
RF-Behavior ↗44 participants across gesture, activity, and emotion tasksExcluded from normalized count. RF is paired with sensor modalities, not a reported RF-text record corpus.
RVTALL ↗20 participants; RF, visual, text, audio, laser, and landmarksExcluded from normalized count. No released paired-record total comparable to the declared unit.
OPERAnet ↗Approximately 8 hours of annotated measurementsExcluded from normalized count. RF and vision activity data, not RF-language records.
NIST semantic RF proposal ↗Proposed collection; no released record countExcluded. No released corpus at the snapshot date.

The comparison is versioned by snapshot date because corpus sizes and access conditions change. Different task definitions are not treated as evidence of equivalent scientific scope.