# Make robot skills earn their installation

11 October 2026 · Native simulation benchmark and next learning experiment

DVIDIA's native contact benchmark completed **6/6 grasp, carry and release placements** across six frozen synthetic layouts; matched open-jaw replay and hold controls completed **0/12**. All 18 attempts were numerically valid. The six-axis reference arm, coupled parallel-jaw gripper and freely moving rigid block provide an executable task for a future learned action policy. The first baseline is authored and uses privileged simulator state; no policy has learned this task from footage. CPU execution also showed substantial timing variation, including a 29.51-second attempt for 20 seconds of simulated motion, so the result establishes task behavior under these settings rather than real-time control. Complete receipts, unsuccessful controls and source identities accompany the replay. Shoe organization and physical robot qualification require their own evidence. ([Frozen confirmation receipt](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-11-native-grasp/replay/receipt/report.json))

## Native contact makes the task measurable

The reference combines **six UR5e joint targets and one Robotiq 2F-85 actuator command**. Its eight gripper linkage joints move through the upstream coupled mechanism; the action interface is one native jaw target in 0–255, rather than eight independent finger commands or a calibrated force request. The task adds a free-joint block and a support table while preserving the pinned arm/gripper geometry, inertias, active collision proxies and linkage constraints. An authored position controller and gravity feedforward act through the bounded native servos. During the rollout, the object moves through dynamics and contact rather than a weld or runtime pose assignment. The openly licensed simulation assets do not imply open-source UR5e or Robotiq physical hardware. ([Frozen task source](https://github.com/Dvidia-Inference/dvidia-training/blob/ac920cbf70147a7edeb0c706a9b71047f318e12f/src/dvidia_training/reference_task.py), [hardware reference](https://github.com/Dvidia-Inference/dvidia-training/blob/ac920cbf70147a7edeb0c706a9b71047f318e12f/docs/hardware-reference.md))

The baseline reaches above the block, descends, closes the jaws, lifts, carries, places, opens, retreats and waits for independent outcome verification. Its inputs are privileged simulator state, including tool/object poses and pad contacts. These are useful for an initial control benchmark because they make perception errors a separate question; their availability on physical hardware must subsequently be demonstrated. The object-specific contact configuration is an explicit experiment parameter, not a measurement of a real shoe material. The declared six-joint/one-actuator profile also keeps this environment separate from the earlier authored arm with a metric jaw-gap interface.

The scorer checks object behavior independently of the controller's named phase. The protocol requires a **60-mm table-free bilateral lift sustained for 100 ms** and retained bilateral contact until the still-held block rests within **25 mm** of the destination. Before that carry-completion latch, **40 ms of consecutive bilateral-contact loss permanently fails retention**; contact counts require normal force above 0.01 N. Released placement requires table support, no remaining pad contact, both measured gripper driver coordinates below 0.1 rad, linear speed below 0.035 m/s, angular speed below 0.3 rad/s, the pinch point at least **100 mm above the object center**, and a **250-ms final dwell**. Unexpected robot–table or non-pad robot–block contact with normal force above 0.01 N disqualifies the result. These are chosen simulation thresholds, not physical calibration tolerances; the destination criterion scores position, not a required final object orientation. The fixed base has no collision proxy, so the check covers the declared active geometry. ([Frozen task source](https://github.com/Dvidia-Inference/dvidia-training/blob/ac920cbf70147a7edeb0c706a9b71047f318e12f/src/dvidia_training/reference_task.py))

Numerical validity remains a prerequisite for completion. Invalid state, solver warnings and prohibited contact cannot be converted into success because an earlier moment looked correct. The simulator uses a pinned MuJoCo version, integrator and command clock; timestep behavior is a separate measurement from task success. MuJoCo documents both the growth of small numerical differences during contact and the version/architecture limits of exact replay. ([MuJoCo computation](https://mujoco.readthedocs.io/en/stable/computation/))

## Frozen attempts distinguish retention from a convincing replay

The test records its scenes, object properties, timeout, controller revision, action limits and outcome thresholds before execution. Its controls include authored placement, open-jaw replay and holding the initial arm posture. The open-jaw condition reuses the **exact issued placement arm-joint command tape** from the same scene while requesting open jaws. Replanning a different trajectory after grasp failure would answer a different question. A complete placement tape is necessary for this paired comparison; an unavailable or partial tape creates an invalid dependent attempt that remains in the denominator.

This control tests a concrete failure mode in demonstration software: an arm can execute a plausible path while the object never becomes supported by the gripper. Independent lift, carry, release and settled-placement evidence distinguishes that path from successful manipulation. Low-friction scenes stress retention without changing the success definition. The benchmark retains unsuccessful and numerically invalid attempts rather than showing only favorable replays. Its seeded layouts are synthetic test groups, and repetitions or related controls do not become independent physical trials.

The confirmation protocol used seed **521047**, four nominal layouts at friction coefficient **0.8** and two low-friction layouts at **0.15**, with independent ±15-mm start/goal x/y jitter. Each scene supplied three complete 20-second attempts at a 20-ms command cadence and 1-ms physics timestep. This local run used an operating-system network-denial profile. The all-attempt results were **4/4 nominal placements, 2/2 low-friction placements, 0/6 open-jaw controls and 0/6 hold controls**. Both control families timed out without acquiring or lifting the object. The six successful placements reported carry completion, no retention-loss event, no prohibited contact and final position errors of **0.489–0.551 mm**; first verified completion occurred at **11.230–11.611 simulated seconds**. No attempt reported solver warnings. These narrow synthetic results do not estimate performance across general objects or hardware. ([Frozen confirmation receipt](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-11-native-grasp/replay/receipt/report.json))

All attempts below contain 1,000 command boundaries. Wall time includes setup and trace construction; the p95 and maximum columns measure narrower per-boundary work. A boundary over 20 ms is recorded as an over-budget computation, rather than a measured physical control deadline miss. No slow attempt is excluded.

| Scene | Control | Success | Attempt wall, s | Boundary p95, ms | Boundary max, ms | Over 20 ms |
| --- | --- | --- | ---: | ---: | ---: | ---: |
| Nominal 00 | Placement | Yes | 5.35 | 12.15 | 124.19 | 31 |
| Nominal 00 | Open jaw | No | 2.37 | 3.09 | 11.71 | 0 |
| Nominal 00 | Hold | No | 3.36 | 5.06 | 26.66 | 1 |
| Nominal 01 | Placement | Yes | 2.97 | 4.00 | 13.81 | 0 |
| Nominal 01 | Open jaw | No | 3.70 | 4.74 | 516.21 | 2 |
| Nominal 01 | Hold | No | 2.83 | 4.03 | 7.01 | 0 |
| Nominal 02 | Placement | Yes | 3.11 | 4.45 | 24.43 | 1 |
| Nominal 02 | Open jaw | No | 8.38 | 14.12 | 438.94 | 32 |
| Nominal 02 | Hold | No | 15.84 | 38.84 | 742.74 | 139 |
| Nominal 03 | Placement | Yes | 29.51 | 78.14 | 1479.25 | 251 |
| Nominal 03 | Open jaw | No | 27.04 | 83.60 | 514.66 | 277 |
| Nominal 03 | Hold | No | 5.64 | 9.78 | 67.48 | 7 |
| Low friction 00 | Placement | Yes | 6.61 | 9.58 | 433.41 | 9 |
| Low friction 00 | Open jaw | No | 3.72 | 5.30 | 33.20 | 2 |
| Low friction 00 | Hold | No | 4.27 | 7.17 | 20.26 | 1 |
| Low friction 01 | Placement | Yes | 4.66 | 7.60 | 60.28 | 5 |
| Low friction 01 | Open jaw | No | 6.83 | 12.23 | 71.08 | 26 |
| Low friction 01 | Hold | No | 4.95 | 8.05 | 45.63 | 8 |

The measurements used **macOS 26.5.2 arm64, Python 3.12.9 and MuJoCo 3.15.0**, with ten logical CPUs reported and no GPU. Process lifetime peak resident memory was **531,054,592 bytes**, recorded before final bundle inspection and replay export. It includes imports, models and accumulated traces; it is not per-episode memory or a minimum RAM specification. Host load was uncontrolled and the large latency variation remains unattributed. The nominal-03 placement exceeded real-time duration and recorded a **1.479-second maximum boundary**. Other placement p95 values were 4.00–12.15 ms, but they do not erase this outlier or establish camera-to-actuator performance. Profiling under controlled load is the next performance experiment. ([Frozen confirmation receipt](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-11-native-grasp/replay/receipt/report.json))

The frozen task source is commit **ac920cbf70147a7edeb0c706a9b71047f318e12f**. The receipt binds protocol SHA-256 **313d0be81ad13a0e20c218f88983650f44a65ea590d955c8d30adda47231c0a6** and task implementation SHA-256 **bbc768ccc202b8dd41cbd235b206e6478806ce47c1f30ae15dc662c0c1edcc3e**, with separate runner, simulator and asset-lock digests. This confirmation uses one physics timestep; loaded 2/1/0.5-ms comparisons and physical parameter identification remain additional experiments. The [recorded task replay](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-11-native-grasp/replay/) preserves the complete [source receipt](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-11-native-grasp/replay/receipt/report.json), including attempts outside the displayed excerpt. ([Frozen benchmark source](https://github.com/Dvidia-Inference/dvidia-training/blob/ac920cbf70147a7edeb0c706a9b71047f318e12f/src/dvidia_training/reference_task_benchmark.py))

Timing needs the same care as success. Complete-attempt wall time includes asset checks, compilation, reset, planning, native stepping, observations and trace construction. A control-boundary distribution measures a narrower portion of that work. Output serialization, browser playback and physical sensing/transport are separate costs. Hardware metadata describes the host used for this workload; installed memory is not a minimum RAM requirement. The benchmark contains no learned inference call, so it cannot supply a camera-to-actuator latency or policy speed claim.

Artifact inspection checks bounded files, digests, planned identities and consistency between declared scenes and saved outcomes. It rejects disagreement even when a changed report and changed trace have both been resealed. This is a reproducibility boundary, not authentication of the measurements or an independent physics recomputation. The browser reconstructs recorded geometry and poses and carries forward the original receipt; it does not turn a displayed motion into fresh task evidence.

## Seven causal action outputs are the next learned component

The current compact model forecasts the next 20-ms joint/jaw state change, and the existing movement head supplies learned motion under an authored task supervisor. The next experiment must give a small DVIDIA policy responsibility for **all seven issued task commands**. Behavioral cloning supplies a direct observations-to-actions baseline, while recurrent and transformer variants offer separately testable ways to represent temporal state. Research on offline manipulation demonstrations also shows that data quality and model-selection criteria materially affect actual task performance. ([robomimic algorithms](https://github.com/ARISE-Initiative/robomimic/blob/master/docs/introduction/implemented_algorithms.md), [offline-demonstration study](https://arxiv.org/abs/2108.03298))

Our proposed first comparison uses low-dimensional native observations and a two-layer, 64-unit MLP with an eight-boundary causal window. At the declared 50-Hz cadence, its eight samples are spaced 20 ms apart and span 140 ms from first to last. A single-boundary ablation tests whether history contributes useful information. A compact four-action head tests short-horizon chunking while replanning each boundary and initially executing only the first action. These are proposed DVIDIA designs. ACT motivates the action-sequence comparison through its treatment of compounding error; its published results do not establish DVIDIA throughput or data requirements. ([ACT project](https://tonyzhaozh.github.io/aloha/))

The observation contract will name ordered joint positions/velocities, gripper state, tool and object/goal relative poses, previous issued commands, acquisition ages and history-validity masks. Authored phase, terminal outcome, future state, known material parameters, teacher IK goal and case identity remain excluded. Optional contact or velocity inputs require explicit ablations and sensing-origin declarations. The teacher provides issued joint targets and native jaw commands as action labels; next measured positions and applied servo forces remain distinct evidence. An explicit low-level limiter can protect the experiment, but its interventions must be counted, and no hidden task supervisor may rescue a failed learner.

Successful nominal attempts supply imitation targets. Failed and intervention attempts stay in the collection manifest and can support separate failure/outcome models. All derivatives from a recording or reset group stay in one partition. Training fits both weights and normalization, development rollouts choose the checkpoint, and a fresh confirmation cohort tests the frozen candidate. Joint-radian and native jaw targets need separate normalization, and a declared sampler should prevent stationary post-success timeout frames from overwhelming acquisition and release examples; complete original tapes remain evidence and sampling labels never enter deployment inputs. A proposed 8/16/32-group development curve investigates data efficiency rather than declaring a universal clip count. Complete closed-loop placement, retention, drops, false completion, unplanned contact and intervention rates determine whether the candidate learned the task; low offline action error alone does not.

Local deployment is an engineering target to measure. The artifact should include bounded weights, processors, units/frames, history/reset behavior, profile/controller compatibility, data/protocol identities, failure handling and the evaluation receipt. We propose measuring cold load, warm inference tails, memory and the full 20-ms loop; a warm p95 below 2 ms is an unmeasured goal. LeRobot's custom-policy interface permits an independently maintained DVIDIA policy and consistent saved processors, allowing collection infrastructure and the learned model to evolve separately. ([LeRobot custom policies](https://huggingface.co/docs/lerobot/main/en/bring_your_own_policies))

## Footage and footwear require additional evidence

A useful skill-space recording captures more than a short visual gesture. Video provides task context, object/goal observations and reviewable actions; robot deployment also requires a metric, robot-feasible action representation and aligned observation/command timing. UMI demonstrates an information-rich human-to-robot bridge built around handheld grippers, recoverable trajectories, width information and a latency-aware policy interface. That precedent supports DVIDIA's goal while identifying what ordinary object boxes and 2D hand tracks still need before becoming actuator labels. ([UMI paper](https://arxiv.org/abs/2402.10329))

Our pipeline should accept synchronized teleoperation first and preserve a reviewed action-reconstruction route for human footage. Reconstruction needs explicit scale/geometry uncertainty, kinematic validity and versioned transformations into the robot's command interface. Actual acquisition and command clocks belong beside video metadata: nominal frame cadence can differ from the achieved recording rate and distort estimated movement speed. ([LeRobot cadence reporting](https://huggingface.co/docs/lerobot/main/en/il_robots#cadence-reporting))

Shoe organization needs its own benchmark. Soft uppers, soles, loose laces, asymmetric geometry, occluded grasp regions, pair identity and toe orientation change both observation and contact requirements. A first shoe experiment should narrow the task to one measured shoe, a fixed table/camera, reachable starts, a reviewed grasp region and tucked laces, then include complete attempts through final placement. Calibrated physical command response, encoder/gap mappings, contact/slip and sensor age remain separate qualification evidence. Each additional arm should receive its own verified model pack, command/sensing contract and native benchmark before inheriting an installation claim. Adding bimanual arms or five-finger hands introduces new transmissions, action dimensions, tactile interfaces and cross-arm collision/handoff tests; the present seven-command capsule does not establish that transfer.

## Conclusion

An installable skill should carry a testable claim about where, on which profile and under which observations it works. The immediate benchmark turns that claim into an independently scored task with reproducible controls. This creates a practical place to compare small learned action models, discover missing coverage and reject attractive demonstrations that fail complete attempts.

The product opportunity is the loop connecting reviewed skill-space evidence, bounded local learning, frozen task receipts and compatible deployment. Each new skill can inherit that loop while supplying its own task conditions and failures. That is a concrete route toward fast downloadable robot behavior, with uncertainty exposed early enough to improve the data and simulation instead of hiding it in the replay.

Run the benchmark using the [installation and evaluation guide](https://github.com/Dvidia-Inference/dvidia-training/blob/ac920cbf70147a7edeb0c706a9b71047f318e12f/docs/native-task-benchmark.md). The implementation and review are available in [GitHub PR #11](https://github.com/Dvidia-Inference/dvidia-training/pull/11).

Published by **DVIDIA** · [hello@dvidia.org](mailto:hello@dvidia.org) · [@iammrriver on X](https://x.com/iammrriver)

The [public Robot Simulation Lab](https://huggingface.co/spaces/Dvidia/robot-simulation-lab) also provides separate unloaded PiPER and SO-101 receipts and installation instructions. [Browser acceptance evidence](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-11-native-grasp/replay-browser/acceptance.json) checks the recorded task replay on desktop and narrow screens.
