# Skillspace arm benchmarks and evolution

**Research and implementation assessment · 7 October 2026 · Simulation-only evidence**

**Skillspace → install → simulated arm makes sense as a grounded, compatible execution workflow.** The present implementation imports task metadata, binds an explicitly authored adapter, supplies a metric scene, executes native contact dynamics, and checks completion. Two engineering rounds increased development success from **14/20 to 16/20 to 17/20**, but the frozen held-out comparison declined from **30/32 to 29/32**: both controllers passed all 24 nominal cases, while actuator-stress success fell from 6/8 to 5/8. There were zero held-out improvements and one regression. Observed instrumented CPU throughput also declined from **12.283× to 11.388×** simulated seconds per wall second. The result establishes an auditable simulation workflow and useful failure diagnostics; it establishes neither improved held-out reliability nor calibrated physical accuracy. Skillspaces should evolve toward immutable releases that bind task, embodiment, observations, controller, and independently evaluated evidence. ([Frozen assessment](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-07-arm-benchmark-v1/benchmark/assessment.json)).

## Installation supplies an explicit execution bridge

The public `place_cup` envelope contains a task label, robot identifier, recording format, and demonstration count. It does not contain learned weights, actions, object dimensions, metric target coordinates, or camera calibration. The local bridge supplies a declared rigid-box proxy and `contact-pick-place-box-v0` adapter for `dvidia-authored-6dof-parallel-jaw-v0`. Thus installation selects existing authored behavior; it does not infer cup manipulation from the referenced demonstrations. This distinction preserves the useful architecture: language identifies the goal, grounding supplies the state and physical assumptions, a compatible controller acts, and a verifier judges the result. ([Public metadata](https://dvidia.org/packs/dvidia/place_cup/skill.json), [execution guide](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-07-arm-benchmark-v1/PRODUCT.md)).

The installation receipt now uses schema 2 to preserve exact source and grounding separately, bind **seven execution files**, and pin MuJoCo/NumPy versions. Execution and installation reads reject stale installation receipts or live runtime edits that break those bindings; reinstallation creates a new record. Export separately rejects a stale run/runtime pairing after code changes; it does not receive the installation receipt. The interface exposes dimensions, mass/material parameters, and actuator capacities rather than hiding their effect behind a task name. These checks establish content consistency. They do not authenticate a publisher or prove the declared task, permissions, or physical outcome. ([Installer](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-07-arm-benchmark-v1/simlab/installer.py), [pack exporter](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-07-arm-benchmark-v1/simlab/arm_pack.py)).

The mechanism has **six arm hinge DOF and two physical jaw slide DOF**, but only **seven independent external commands**: six joint-position targets and one coupled total jaw opening. The free object contributes another six unactuated world DOF. Consequently `nv=14`, `nq=15`, and `nu=8`; the quaternion adds a position coordinate, not an additional velocity DOF. The controller requests XYZ tool position while fixing upright orientation. It is not an arbitrary six-dimensional pose/perception interface. ([Native environment](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-07-arm-benchmark-v1/simlab/arm_env.py)).

| Required execution input | Current contract |
| --- | --- |
| Task and outcome | Pick, retain a lift, transport, release, and settle a box at a world-frame target |
| Arm action | Six radians-valued joint targets; limits respectively ±2.6, [−2.2, 1.6], ±2.7, ±3.0, ±2.8, and ±3.1 rad |
| Gripper action | One total gap, 0–80 mm; each physical slide receives half the commanded gap |
| Geometry and scene | Object/target XYZ in metres; box widths independently 30–50 mm in the benchmark; table height 0.29 m |
| Mass and contact | 20–80 g; raw object friction 0.4–1.5; explicit jaw/table contact law |
| Actuation and timing | Default caps 40 N·m per arm actuator and 15 N per jaw; 1 ms physics and 20 ms control intervals |
| Observations | Privileged joint/TCP/object poses and velocities, jaw gap, and native contacts/normal forces |

These inputs come from authored grounding rather than measured reconstruction. Held-out coordinates were proposed over X [0.251, 0.551] m and Y [−0.198, 0.198] m, filtered by radial distance ≤0.555 m and transport ≥0.09 m. Scene jitter was zero; changing seed labels alone would repeat a layout. There is no calibrated vision or tactile stream. ([Frozen protocol](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-07-arm-benchmark-v1/benchmark/protocol.json)).

## Two rounds improved development behavior but exposed a regression

The benchmark froze 20 development cases, 32 distinct held-out layouts, and seven separately counted rejection probes. Only development outcomes guided the two revisions. Exact sources were copied into isolated snapshots, and the final source was frozen before the baseline/final held-out comparison. The contact law, original actuator capacities, and task success predicate remained unchanged. These controls prevent scoring gains through easier success rules or stronger motors. ([Benchmark method and results](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-07-arm-benchmark-v1/benchmark/RESULTS.md)).

Round 1 replaced a fixed 28 mm closure with geometry/load-aware jaw commands, native normal-force feedback, and a sustained bilateral-contact gate before lifting. It resolved the narrow/heavy nominal failures but introduced a development stress regression. Round 2 separated required preload from available actuator force, added target ramps and phase/capacity diagnostics, and implemented one bounded recovery: release, wait for actual table support and settling, then regrasp the object's observed new position within the upright-box assumptions. Recovery does not edit object coordinates or attach it to the tool. These are authored controller revisions, not learned-from-video or reinforcement-learning updates. ([Round 1 records](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-07-arm-benchmark-v1/benchmark/round1/dev.json), [Round 2 records](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-07-arm-benchmark-v1/benchmark/round2/dev.json), [controller](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-07-arm-benchmark-v1/simlab/arm_policy.py)).

| Evaluation | Baseline | Round 1 | Round 2/final |
| --- | ---: | ---: | ---: |
| Development nominal | 12/14 | 14/14 | 14/14 |
| Development actuator stress | 2/6 | 2/6 | 3/6 |
| Development total | 14/20 | 16/20 | 17/20 |
| Held-out nominal | 24/24 | Not evaluated | 24/24 |
| Held-out actuator stress | 6/8 | Not evaluated | 5/8 |
| Held-out total | 30/32 | Not evaluated | 29/32 |

The paired held-out outcomes were **29 successes for both, two failures for both, one regression, and no improvements**. The regressed `held-29` uses a 50 mm cube, 40 g mass, effective jaw friction 1.0, 2 N·m arm caps, and 0.25 N jaw caps. Baseline completed in 15.969 simulated seconds with 6.692 mm error. Final recorded a 116.95 mm retained lift and one recovery, then exhausted the unchanged 18-second horizon in `place`, still gripping an object unsupported by the table, with 80.70 mm error. Release and settled dwell had not occurred. Recovery and qualification used time, but **no causal ablation isolates which revision caused the regression**. ([Baseline receipt](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-07-arm-benchmark-v1/benchmark/v0/heldout.json), [final receipt](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-07-arm-benchmark-v1/benchmark/final/heldout.json)).

The nominal result is encouraging within this small synthetic suite, but the stress result rejects a general improvement claim. For scale, 24/24 gives a **95% Wilson binomial reference interval of 86.2–100%**. This fixed, heterogeneous suite is not an independent, identically distributed population sample; the interval is not a real-world reliability guarantee. A future cycle must use new development cases and a newly frozen evaluation; these now-visible held-out cases should remain an immutable published result. ([Assessment](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-07-arm-benchmark-v1/benchmark/assessment.json), [NIST interval method](https://www.itl.nist.gov/div898/handbook/prc/section2/prc241.htm)).

## Contact dependence and numerical validity do not establish physical accuracy

MuJoCo evolves a dynamic free box through gravity and contact; there is no weld, mocap attachment, or supplied object trajectory. The verifier requires a bilateral loaded grasp and table-free lift ≥55 mm retained for 100 ms, released jaws, tool clearance ≥100 mm, table support, target error within 25 mm, linear speed <35 mm/s, angular speed <0.3 rad/s, and 250 ms settled dwell. Idle/open-jaw and identical-arm-target/open-jaw controls produced **zero successes across twelve episodes**. All 76 held-out task/control episodes completed with valid state and no captured engine warnings. Those results support contact-dependent execution and numerical checks. ([Assessment](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-07-arm-benchmark-v1/benchmark/assessment.json)).

The model's boundaries are substantial. Only jaws, object, and table collide; links and palm are excluded. Reset boxes are axis-aligned and the tool remains upright. Obstacles, self-collision, handles, arbitrary orientation, deformables, and physical transfer were not qualified. Equal-priority geometry mixing gives jaw friction `max(1.0, object_friction)` and table friction `max(0.8, object_friction)`, so low raw coefficients do not test a slippery grasp. This follows MuJoCo's contact rules, not measured material behavior. ([MuJoCo contact parameters](https://mujoco.readthedocs.io/en/stable/modeling.html#contact-parameters)).

**Maximum contact penetration sampled at 20 ms rose from 1.614 to 3.497 mm.** The 1 ms-sampled maximum normal-force sum on a single pad declined from **2.923 to 1.393 N**; this statistic takes the larger of the two pad sums rather than adding both pads. Neither measurement is a continuous physical peak. Actuator drive limits and contact forces are different quantities, and numerical penetration reflects the modeled compliant constraint response. Zero warnings and a successful placement do not calibrate stiffness, friction, sensor noise, or transient force accuracy. ([Baseline measurements](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-07-arm-benchmark-v1/benchmark/v0/heldout.json), [final measurements](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-07-arm-benchmark-v1/benchmark/final/heldout.json)).

Round 2 ramps applied target trajectories at **3 rad/s per hinge and 0.16 m/s total jaw width**, each native tick before gravity/bias compensation. These are command-trajectory limits, not physical speed caps or a guarantee on the final compensated control's slew. Actual joint velocity is measured separately. Hardware claims require identified actuators, kinematics, contact/material traces, observation latency, and independent physical tests.

## The CPU and offline results describe a bounded installed runtime

Baseline processed 417.094 simulated episode seconds in 33.956 wall seconds; final processed 432.874 in 38.012, giving **12.283× versus 11.388×**. The workload ran one CPU world on Apple M5, ten logical CPUs and 24 GiB RAM, with Python 3.12.9, MuJoCo 3.15.0, and NumPy 2.5.3. Timing includes compilation, reset/settling, controller/IK, native stepping, observations, and diagnostics; it excludes startup/imports, installation, copying/JSON output, rendering, and UI. Simulated settling is excluded from the numerator although its wall cost is included. ([Timing receipts](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-07-arm-benchmark-v1/benchmark/RESULTS.md)).

This is one execution per version/case rather than a repeated, thermally controlled speed campaign. Extra instrumentation contributes to cost; Round 2 development timing overlapped another regression run and is excluded from speed comparisons. No GPU throughput, minimum VRAM, cross-engine ranking, or physics-core acceleration follows. A GPU backend should earn its place through feature-matched, batched measurements; MuJoCo's documentation separates JAX/Warp support and performance tradeoffs. ([Official MJX documentation](https://mujoco.readthedocs.io/en/latest/mjx.html)).

A fresh process-level network-denial check completed **3/3 known layouts**, after native libc connection returned `EPERM`. It verifies installed execution on this macOS host, not air-gapped dependency installation or unseen task capability. Verification also covered 100 distinct tests: a 99-test full run plus a targeted 26-test subset containing one additional test. Test coverage and native offline execution complement benchmark evidence without replacing it. ([Offline receipt](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-07-arm-benchmark-v1/native-offline/native-policy.json), [execution guide](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-07-arm-benchmark-v1/PRODUCT.md)).

A browser check after runtime restart installed the public source and executed a changed **30 × 40 × 30 mm, 80 g** box scene, from XY (0.31, 0.05) m to (0.52, 0.18) m. It completed in **13.636 simulated seconds with 0.277 mm target error**, showing native joint/gripper feedback and the unchanged held-out benchmark disclosure. Raw friction 0.4 still produced effective jaw friction 1.0. This verifies the local interface-to-native-job path. ([Recorded run](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-07-arm-benchmark-v1/demo/run.json), [recorded replay](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-07-arm-benchmark-v1/demo/replay.html)).

The final source archive was freshly extracted with preinstalled pinned dependencies. Three visible development cases (`dev-00`, `dev-10`, `dev-16`) reproduced exact control steps, final state, object position, metrics, and policy diagnostics, excluding wall timing; **85 manifest files were verified** and the receipt records `all_match: true`. These are package-integrity and reproduction checks, not new held-out coverage or cross-platform qualification. ([Reproduction receipt](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-07-arm-benchmark-v1/release/reproduction.json)).

## Media collections should graduate through evidence-bearing releases

Folders, videos, images, and curation remain useful source organization. Their next role is to preserve immutable recording/segment identities, modality and clock provenance, reviewed states, uncertainty, purpose-specific permissions, and grouping. A thumbnail is not a calibrated observation; a video count is not action supervision. A learned branch needs temporally aligned robot observations/actions, an actual tested dataset reader, frozen group splits, training records, and checkpoint identity. LeRobot v3 explicitly separates state/action/timestamp tables, video, and episode/feature metadata. ([Official dataset documentation](https://huggingface.co/docs/lerobot/lerobot-dataset-v3)).

Maintain two explicit routes: an authored adapter grounded to a scenario, and a policy learned from declared supervision. Both require exact executable/runtime identity and evaluation. A proposed release inspector now checks role-correct artifact hashes, profile/action/sensor declarations, stage bindings, baseline/test episode matching, and declared group overlap. It is separate from current metadata exports and always returns **`simulation_release_ready: false` and `hardware_ready: false` for drafts**, even when every hash agrees. It neither executes lifecycle predicates nor authenticates rights/evaluators or detects semantic duplicates. A trusted executor/publisher remains necessary. ([Proposed inspector](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-07-arm-benchmark-v1/simlab/skill_release.py)).

Future compatibility must bind ordered joints and units, kinematics/limits, TCP/base frames, gripper geometry/aperture/force, control mode/rate, sensing shape/frame/latency, material/state envelope, initiation/termination/recovery, and evaluation evidence. Equal joint counts do not supply that mapping. MuJoCo also does not guarantee numerical reproducibility across releases, so engine identity belongs in the receipt. ([Hardware integration contract](https://huggingface.co/docs/lerobot/integrate_hardware), [MuJoCo versioning](https://github.com/google-deepmind/mujoco/blob/main/VERSIONING.md)).

Expand through distinct families: rigid manipulation adds collision-enabled geometry and release checks; articulated tasks add axes, limits, handles and latches; insertion/wiping adds clearance and force/compliance verification; deformables add deformation/topology and calibrated material/contact observations; bimanual tasks add synchronized profiles, relative-base transforms, and inter-arm collision/load sharing. Each needs its own success predicate, failure cases, and independent grouped evaluation. Box placement cannot qualify these by inheritance.

The next **new development protocol** should isolate weak-force transport, remaining-time budgets, recovery costs, and contact/perception errors through separate ablations. Preserve this held-out release. Calibrate measured material/sensor responses before claiming physical fidelity; keep calibration episodes and all their derivatives outside independent tests. Two Grok 4.7 consultations supplied adversarial critiques of stale receipts, declaration-only validation, and group leakage. Those opinions were checked against primary documentation; they are not experimental evidence, and private transcripts are not part of this public report.

## Conclusion

The useful advance is that a skill label now leads to reproducible, contact-dependent execution with an explicit failure boundary. The preserved regression makes the release more informative: iteration improved selected development behavior while failing to improve the sealed comparison. Future Skillspace value should be measured by fewer resources to an independently qualified capability, with unsupported inputs and failed recoveries visible.

The next download should carry the exact controller, grounded scenario, observation contract, runtime versions, and evaluation receipt together. Start from the [source release](https://research.dvidia.org/downloads/topics/physical-grounding/runs/2026-10-07-arm-benchmark-v1/release/simlab-v0.3.zip) and its declared simulation scope, then qualify new task families and hardware through new evidence rather than extending the meaning of a familiar task name.
