Paired simulation benchmark and product assessment

Skillspace arm benchmarks and evolution

· DVIDIA

Mixed held-out result · 30/32 baseline, 29/32 updated

Source & history · Download Markdown

Abstract

Two controller revisions improve development completion from 14/20 to 17/20, but the frozen paired test declines from 30/32 to 29/32: both pass all 24 normal-strength cases while actuator-stress completion falls from 6/8 to 5/8. Instrumented CPU throughput also declines. We publish the regression, unchanged contact and success rules, exact controller snapshots, richer physical inputs, runtime-bound installation and a proposed staged Skillspace contract. Contact forces and penetration remain uncalibrated; learned perception and physical qualification are open.

In this paper

Research and implementation assessment · 7 October 2026 · Simulation-only evidence

Skillspace → install → simulated arm makes sense as a grounded, compatible execution workflow. The present implementation imports task metadata, binds an explicitly authored adapter, supplies a metric scene, executes native contact dynamics, and checks completion. Two engineering rounds increased development success from 14/20 to 16/20 to 17/20, but the frozen held-out comparison declined from 30/32 to 29/32: both controllers passed all 24 nominal cases, while actuator-stress success fell from 6/8 to 5/8. There were zero held-out improvements and one regression. Observed instrumented CPU throughput also declined from 12.283× to 11.388× simulated seconds per wall second. The result establishes an auditable simulation workflow and useful failure diagnostics; it establishes neither improved held-out reliability nor calibrated physical accuracy. Skillspaces should evolve toward immutable releases that bind task, embodiment, observations, controller, and independently evaluated evidence. (Frozen assessment).

Installation supplies an explicit execution bridge

The public place_cup envelope contains a task label, robot identifier, recording format, and demonstration count. It does not contain learned weights, actions, object dimensions, metric target coordinates, or camera calibration. The local bridge supplies a declared rigid-box proxy and contact-pick-place-box-v0 adapter for dvidia-authored-6dof-parallel-jaw-v0. Thus installation selects existing authored behavior; it does not infer cup manipulation from the referenced demonstrations. This distinction preserves the useful architecture: language identifies the goal, grounding supplies the state and physical assumptions, a compatible controller acts, and a verifier judges the result. (Public metadata, execution guide).

The installation receipt now uses schema 2 to preserve exact source and grounding separately, bind seven execution files, and pin MuJoCo/NumPy versions. Execution and installation reads reject stale installation receipts or live runtime edits that break those bindings; reinstallation creates a new record. Export separately rejects a stale run/runtime pairing after code changes; it does not receive the installation receipt. The interface exposes dimensions, mass/material parameters, and actuator capacities rather than hiding their effect behind a task name. These checks establish content consistency. They do not authenticate a publisher or prove the declared task, permissions, or physical outcome. (Installer, pack exporter).

The mechanism has six arm hinge DOF and two physical jaw slide DOF, but only seven independent external commands: six joint-position targets and one coupled total jaw opening. The free object contributes another six unactuated world DOF. Consequently nv=14, nq=15, and nu=8; the quaternion adds a position coordinate, not an additional velocity DOF. The controller requests XYZ tool position while fixing upright orientation. It is not an arbitrary six-dimensional pose/perception interface. (Native environment).

Required execution inputCurrent contract
Task and outcomePick, retain a lift, transport, release, and settle a box at a world-frame target
Arm actionSix radians-valued joint targets; limits respectively ±2.6, [−2.2, 1.6], ±2.7, ±3.0, ±2.8, and ±3.1 rad
Gripper actionOne total gap, 0–80 mm; each physical slide receives half the commanded gap
Geometry and sceneObject/target XYZ in metres; box widths independently 30–50 mm in the benchmark; table height 0.29 m
Mass and contact20–80 g; raw object friction 0.4–1.5; explicit jaw/table contact law
Actuation and timingDefault caps 40 N·m per arm actuator and 15 N per jaw; 1 ms physics and 20 ms control intervals
ObservationsPrivileged joint/TCP/object poses and velocities, jaw gap, and native contacts/normal forces

These inputs come from authored grounding rather than measured reconstruction. Held-out coordinates were proposed over X [0.251, 0.551] m and Y [−0.198, 0.198] m, filtered by radial distance ≤0.555 m and transport ≥0.09 m. Scene jitter was zero; changing seed labels alone would repeat a layout. There is no calibrated vision or tactile stream. (Frozen protocol).

Two rounds improved development behavior but exposed a regression

The benchmark froze 20 development cases, 32 distinct held-out layouts, and seven separately counted rejection probes. Only development outcomes guided the two revisions. Exact sources were copied into isolated snapshots, and the final source was frozen before the baseline/final held-out comparison. The contact law, original actuator capacities, and task success predicate remained unchanged. These controls prevent scoring gains through easier success rules or stronger motors. (Benchmark method and results).

Round 1 replaced a fixed 28 mm closure with geometry/load-aware jaw commands, native normal-force feedback, and a sustained bilateral-contact gate before lifting. It resolved the narrow/heavy nominal failures but introduced a development stress regression. Round 2 separated required preload from available actuator force, added target ramps and phase/capacity diagnostics, and implemented one bounded recovery: release, wait for actual table support and settling, then regrasp the object's observed new position within the upright-box assumptions. Recovery does not edit object coordinates or attach it to the tool. These are authored controller revisions, not learned-from-video or reinforcement-learning updates. (Round 1 records, Round 2 records, controller).

EvaluationBaselineRound 1Round 2/final
Development nominal12/1414/1414/14
Development actuator stress2/62/63/6
Development total14/2016/2017/20
Held-out nominal24/24Not evaluated24/24
Held-out actuator stress6/8Not evaluated5/8
Held-out total30/32Not evaluated29/32

The paired held-out outcomes were 29 successes for both, two failures for both, one regression, and no improvements. The regressed held-29 uses a 50 mm cube, 40 g mass, effective jaw friction 1.0, 2 N·m arm caps, and 0.25 N jaw caps. Baseline completed in 15.969 simulated seconds with 6.692 mm error. Final recorded a 116.95 mm retained lift and one recovery, then exhausted the unchanged 18-second horizon in place, still gripping an object unsupported by the table, with 80.70 mm error. Release and settled dwell had not occurred. Recovery and qualification used time, but no causal ablation isolates which revision caused the regression. (Baseline receipt, final receipt).

The nominal result is encouraging within this small synthetic suite, but the stress result rejects a general improvement claim. For scale, 24/24 gives a 95% Wilson binomial reference interval of 86.2–100%. This fixed, heterogeneous suite is not an independent, identically distributed population sample; the interval is not a real-world reliability guarantee. A future cycle must use new development cases and a newly frozen evaluation; these now-visible held-out cases should remain an immutable published result. (Assessment, NIST interval method).

Contact dependence and numerical validity do not establish physical accuracy

MuJoCo evolves a dynamic free box through gravity and contact; there is no weld, mocap attachment, or supplied object trajectory. The verifier requires a bilateral loaded grasp and table-free lift ≥55 mm retained for 100 ms, released jaws, tool clearance ≥100 mm, table support, target error within 25 mm, linear speed <35 mm/s, angular speed <0.3 rad/s, and 250 ms settled dwell. Idle/open-jaw and identical-arm-target/open-jaw controls produced zero successes across twelve episodes. All 76 held-out task/control episodes completed with valid state and no captured engine warnings. Those results support contact-dependent execution and numerical checks. (Assessment).

The model's boundaries are substantial. Only jaws, object, and table collide; links and palm are excluded. Reset boxes are axis-aligned and the tool remains upright. Obstacles, self-collision, handles, arbitrary orientation, deformables, and physical transfer were not qualified. Equal-priority geometry mixing gives jaw friction max(1.0, object_friction) and table friction max(0.8, object_friction), so low raw coefficients do not test a slippery grasp. This follows MuJoCo's contact rules, not measured material behavior. (MuJoCo contact parameters).

Maximum contact penetration sampled at 20 ms rose from 1.614 to 3.497 mm. The 1 ms-sampled maximum normal-force sum on a single pad declined from 2.923 to 1.393 N; this statistic takes the larger of the two pad sums rather than adding both pads. Neither measurement is a continuous physical peak. Actuator drive limits and contact forces are different quantities, and numerical penetration reflects the modeled compliant constraint response. Zero warnings and a successful placement do not calibrate stiffness, friction, sensor noise, or transient force accuracy. (Baseline measurements, final measurements).

Round 2 ramps applied target trajectories at 3 rad/s per hinge and 0.16 m/s total jaw width, each native tick before gravity/bias compensation. These are command-trajectory limits, not physical speed caps or a guarantee on the final compensated control's slew. Actual joint velocity is measured separately. Hardware claims require identified actuators, kinematics, contact/material traces, observation latency, and independent physical tests.

The CPU and offline results describe a bounded installed runtime

Baseline processed 417.094 simulated episode seconds in 33.956 wall seconds; final processed 432.874 in 38.012, giving 12.283× versus 11.388×. The workload ran one CPU world on Apple M5, ten logical CPUs and 24 GiB RAM, with Python 3.12.9, MuJoCo 3.15.0, and NumPy 2.5.3. Timing includes compilation, reset/settling, controller/IK, native stepping, observations, and diagnostics; it excludes startup/imports, installation, copying/JSON output, rendering, and UI. Simulated settling is excluded from the numerator although its wall cost is included. (Timing receipts).

This is one execution per version/case rather than a repeated, thermally controlled speed campaign. Extra instrumentation contributes to cost; Round 2 development timing overlapped another regression run and is excluded from speed comparisons. No GPU throughput, minimum VRAM, cross-engine ranking, or physics-core acceleration follows. A GPU backend should earn its place through feature-matched, batched measurements; MuJoCo's documentation separates JAX/Warp support and performance tradeoffs. (Official MJX documentation).

A fresh process-level network-denial check completed 3/3 known layouts, after native libc connection returned EPERM. It verifies installed execution on this macOS host, not air-gapped dependency installation or unseen task capability. Verification also covered 100 distinct tests: a 99-test full run plus a targeted 26-test subset containing one additional test. Test coverage and native offline execution complement benchmark evidence without replacing it. (Offline receipt, execution guide).

A browser check after runtime restart installed the public source and executed a changed 30 × 40 × 30 mm, 80 g box scene, from XY (0.31, 0.05) m to (0.52, 0.18) m. It completed in 13.636 simulated seconds with 0.277 mm target error, showing native joint/gripper feedback and the unchanged held-out benchmark disclosure. Raw friction 0.4 still produced effective jaw friction 1.0. This verifies the local interface-to-native-job path. (Recorded run, recorded replay).

The final source archive was freshly extracted with preinstalled pinned dependencies. Three visible development cases (dev-00, dev-10, dev-16) reproduced exact control steps, final state, object position, metrics, and policy diagnostics, excluding wall timing; 85 manifest files were verified and the receipt records all_match: true. These are package-integrity and reproduction checks, not new held-out coverage or cross-platform qualification. (Reproduction receipt).

Media collections should graduate through evidence-bearing releases

Folders, videos, images, and curation remain useful source organization. Their next role is to preserve immutable recording/segment identities, modality and clock provenance, reviewed states, uncertainty, purpose-specific permissions, and grouping. A thumbnail is not a calibrated observation; a video count is not action supervision. A learned branch needs temporally aligned robot observations/actions, an actual tested dataset reader, frozen group splits, training records, and checkpoint identity. LeRobot v3 explicitly separates state/action/timestamp tables, video, and episode/feature metadata. (Official dataset documentation).

Maintain two explicit routes: an authored adapter grounded to a scenario, and a policy learned from declared supervision. Both require exact executable/runtime identity and evaluation. A proposed release inspector now checks role-correct artifact hashes, profile/action/sensor declarations, stage bindings, baseline/test episode matching, and declared group overlap. It is separate from current metadata exports and always returns simulation_release_ready: false and hardware_ready: false for drafts, even when every hash agrees. It neither executes lifecycle predicates nor authenticates rights/evaluators or detects semantic duplicates. A trusted executor/publisher remains necessary. (Proposed inspector).

Future compatibility must bind ordered joints and units, kinematics/limits, TCP/base frames, gripper geometry/aperture/force, control mode/rate, sensing shape/frame/latency, material/state envelope, initiation/termination/recovery, and evaluation evidence. Equal joint counts do not supply that mapping. MuJoCo also does not guarantee numerical reproducibility across releases, so engine identity belongs in the receipt. (Hardware integration contract, MuJoCo versioning).

Expand through distinct families: rigid manipulation adds collision-enabled geometry and release checks; articulated tasks add axes, limits, handles and latches; insertion/wiping adds clearance and force/compliance verification; deformables add deformation/topology and calibrated material/contact observations; bimanual tasks add synchronized profiles, relative-base transforms, and inter-arm collision/load sharing. Each needs its own success predicate, failure cases, and independent grouped evaluation. Box placement cannot qualify these by inheritance.

The next new development protocol should isolate weak-force transport, remaining-time budgets, recovery costs, and contact/perception errors through separate ablations. Preserve this held-out release. Calibrate measured material/sensor responses before claiming physical fidelity; keep calibration episodes and all their derivatives outside independent tests. Two Grok 4.7 consultations supplied adversarial critiques of stale receipts, declaration-only validation, and group leakage. Those opinions were checked against primary documentation; they are not experimental evidence, and private transcripts are not part of this public report.

Conclusion

The useful advance is that a skill label now leads to reproducible, contact-dependent execution with an explicit failure boundary. The preserved regression makes the release more informative: iteration improved selected development behavior while failing to improve the sealed comparison. Future Skillspace value should be measured by fewer resources to an independently qualified capability, with unsupported inputs and failed recoveries visible.

The next download should carry the exact controller, grounded scenario, observation contract, runtime versions, and evaluation receipt together. Start from the source release and its declared simulation scope, then qualify new task families and hardware through new evidence rather than extending the meaning of a familiar task name.