Research agenda
Turn physical footage into testable evidence
Four baseline protocols proposed
Abstract
A research agenda covering visible capture evidence, relations over time, evaluation provenance and physical grounding. Each question has a bounded protocol and a failure condition. The aim is to measure whether structure and physical signals improve a defined task before scaling collection. Strategic opinions, author-reported results and our untested proposals are kept distinct.
In this paper
DVIDIA should test the usefulness of its evidence before expanding collection volume. The immediate opportunity is to make demonstrations easier to inspect, compare and reproduce, then measure whether added structure or physical signals help a defined task. Sam Padilla's 5 October essay provides a strategic prompt; the technical sources support narrower, testable questions. The four baseline protocols proposed in this agenda have not run, and neither customer demand nor robot transfer has been established. Four contributor-sized studies cover capture quality, relations over time, provenance and physical grounding.
The essay motivates a hypothesis, not a market forecast
In Robotics Data: Bet Hard Or Get Out, Sam Padilla, 5 October 2026, argues for exceptional collection scale, distinctive physical information or synthetic generation. He warns that buyers can internalize embodiment-specific collection and points to opportunities beyond raw data supply. His conversations and volume anecdote are not independently verified demand evidence. (Original essay).
Our inference is a practical test: does an attributable demonstration with observable state changes improve a specific evaluation more than a clip with coarse tags? This would establish a useful property of a research artifact. It would not establish willingness to pay, a durable business model or an executable robot skill. The first deliverable should therefore be a small benchmark with a clear failure condition, rather than a volume target.
Observable interactions come before richer labels
MediaPipe Hands provides an accessible RGB hand-tracking baseline. Ego4D's hand–object work asks a different question: what interaction occurred, where and when, and what state changed. That distinction motivates evaluating whether footage exposes useful evidence instead of treating detected hands as a quality certificate. (Zhang et al., 2020; Ego4D FHO documentation).
The proposed capture-quality study begins with a small annotation pilot and expands only if two annotators can apply the definitions consistently. It compares hand presence, image-quality features and their combination against observable interaction intervals. Its preregistered continuation target is false acceptance of non-observable intervals at most 10%, while retaining at least 60% of observable intervals, with contributor-grouped uncertainty estimates. These are proposed targets, not measured outcomes. Sustained capture performance is tested separately on named devices. See the capture protocol.
The skill-relations study separates perception error from temporal reasoning. Maëlic Neau's September 2026 RelateAnything consumes externally supplied regions and predicate strings; it does not itself detect objects. Its AGPL-3.0-only code plus NOTICE and DINOv3-derived weight terms must be distinguished from the older EvolvingLMMs-Lab RAM repository's Apache-2.0 code. These are two different projects. (Neau, 2026; 2026 repository; model card; RAM repository).
First compare human-provided and detected regions for four visual relations, retaining an explicit unknown state. Then compare independent frame predictions with simple temporal rules, attaching every proposed event to its source interval. Action Genome motivates this decomposition into changing relations, but provides no validation of our particular procedure. The proposal measures relation precision/coverage, event F1, boundary errors and false events; a diagram alone is not a result. See the relation protocol. (Ji et al., 2019).
Physical grounding requires narrower comparisons and honest costs
T-Rex combines human-video pretraining with instrumented robot data and task demonstrations. Its authors report 100 hours of tactile robot midtraining, while the released dataset is an approximately 50-hour subset. The dataset's 23 video streams include RGB, raw tactile and deformation streams; they are not 23 ordinary camera views. Its preprocessing also includes imputation, making revision and transformation records material to any comparison. (T-Rex paper; dataset card).
The proposed physical-grounding baseline uses a capped metadata/subset inspection to compare coarse tags, temporal state descriptions and measured contact descriptors for retrieval. It holds out objects and whole episodes, reports Recall@5 and state compatibility, and stops if richer metadata does not earn its annotation cost. The first experiment tests retrieval, not robot execution. A later simulator comparison would keep training updates and held-out scene seeds fixed, then add generated trajectories as one intervention. See the physical-grounding protocol.
NVIDIA's March 2025 article reports 780,000 simulated trajectories generated in 11 hours and a 40% performance improvement. The latter claim's exact denominator remains unresolved in this review. GR00T N1's separately reported neural-video process used about 105,000 L40 GPU-hours for 827 hours of video; it is a different process from the 11-hour simulation trajectory result. These figures cannot be merged into a single cost or used as a forecast for this project. (NVIDIA article; GR00T N1 paper).
The documented NVIDIA blueprint specifies an RTX A6000 with 48 GB VRAM, with an optional separate 80 GB H100-class Cosmos workflow. That is a reason to start with capped metadata/retrieval work, not a universal hardware requirement for all synthetic experiments. A simulation proposal needs its own resource budget before execution; appearance generation must be separately costed. (Blueprint repository).
Provenance makes small results worth sharing
RLDS documents episode boundaries and optional action fields; LeRobot v3 separates video, tabular records and episode metadata. Neither format establishes that a human video contains robot actions. The proposed evaluation-provenance study preserves absent signals, records originals and derivatives separately, and tests actual reader behavior. Datasheets supplies a framework for describing collection and intended use. (RLDS; LeRobot v3; Gebru et al., 2021).
A tiny permission-cleared fixture will contain deliberate corruption, timestamp errors and duplicate derivatives. The acceptance criterion is reproducible inventory/split hashes, detection of all seeded structural faults, and readback within declared alignment tolerances. Near-duplicate detection is evaluated separately; a byte hash does not identify semantic equivalence. All derivatives from an original recording share its split. See the provenance protocol.
Contribution value comes from resolved uncertainty
The most useful early result can be a failed hypothesis: a visibility score that accepts obscured interactions, a relation model that cannot calibrate contact, or metadata whose annotation cost exceeds its retrieval benefit. Publishing the protocol and failure evidence makes that result reusable. Expanding collection or instrumentation should follow a demonstrated bottleneck and a measured benefit, with physical execution evaluated as a separate claim.