Product prototype and native contact qualification

Installing a Skillspace on a simulated arm

· DVIDIA

Public skill link to authored adapter · 6/6 scene placements

Source & history · Download Markdown

Abstract

The first link-to-arm product path installs an existing DVIDIA placement envelope, binds a local authored controller to one simulated arm, and reuses it in independently grounded scenes. Native jaw contacts carry a free box; six layouts complete, while no-motion and identical joint commands with open jaws each complete zero of six. Installed execution passes process-level offline checks. Task metadata contains no learned policy; perception, cup geometry, full arm collisions and physical qualification remain open.

In this paper

Research note, 7 October 2026. The current bridge makes a narrow part of the proposed skill workflow executable: import a DVIDIA skill envelope, install an explicitly authored local simulation adapter, create a scene, and run an articulated arm. It provides a concrete place to investigate observation-to-simulation-to-skill packaging. It does not yet infer an executable policy, object geometry, material properties, or robot calibration from human video.

The public example is DVIDIA place_cup, whose metadata is served at place_cup/skill.json. The recorded 360-byte source is vdia:0, version 0.3.0, with task “Set a cup on the mark.” Its exact-byte SHA-256 is 16b4e1ef0f3799345c1d18ce9f1da83b072863ab00ce30a83424080fb467d4ec. The envelope declares a demonstration count, recording format and robot identifier; these declarations do not supply trajectories, observations, trained weights, or an evaluated controller. The installation receipt preserves the source bytes, URL, source hash, separate grounding hash, adapter identity and physical_robot_ready: false. A later download may differ; its receipt must identify its own bytes.

In the local studio, one click installs the compatible public envelope and binds contact-pick-place-box-v0 to dvidia-authored-6dof-parallel-jaw-v0. A bundled envelope supports installation without fetching the public site. JSON upload/drop accepts the supported DVIDIA formats. Missing grounding is supplied through a visible, locally authored adapter; inline grounding remains separately validated. Installation downloads metadata and executes the already installed local controller. Object and target coordinates belong to the scene and can change between executions. The browser displays recorded native simulation poses rather than inventing an object trajectory. Execution and replay are designed to work with installed dependencies and local assets; the broader product's air-gapped deployment remains a separate validation task.

The embodiment is an original six-hinge arm with two actuated sliding jaws. Native MuJoCo 3.15.0 simulates a free rigid box, gravity and jaw/object/table contact. The default proxy is a 40 mm cube weighing 40 g on a table at 0.29 m. No weld, mocap attachment, or object-state update carries the object during execution. Arm controls use bounded position servos, model-based bias compensation and current-state damped Jacobian control. The declared limits are 40 N·m per arm actuator and 15 N per jaw. Physics uses a 1 ms timestep, implicitfast integration, Newton solving and an elliptic Coulomb friction cone. These choices describe an authored fixture, not measured hardware or calibrated material fidelity.

Contact provenance matters. At equal geometry priority, MuJoCo mixes sliding friction by taking the maximum, as described in its contact-parameter documentation. Consequently jaw/object friction is max(1.0, object_friction) and table/object friction is max(0.8, object_friction); lowering the object parameter below those floors does not produce a low-friction grasp. Only the jaws, object and table collide. Arm links and palm are collision-excluded, so obstacle avoidance and self-collision are unqualified. Observations expose simulator positions, velocities and contacts directly. The controller has no camera perception, tactile sensor model, cup handle geometry, deformable object, or learned recovery policy. A cup-labelled envelope currently selects a declared box proxy.

Success requires bilateral loaded jaw contact, a table-free lift of at least 55 mm retained for 100 ms, subsequent open jaws with no object contact, and a hand at least 100 mm above the object. The object must contact the table, remain within its target tolerance, have linear speed below 35 mm/s and angular speed below 0.3 rad/s, and satisfy the complete condition for the declared dwell, normally 250 ms. Textual engine warnings and native warning counters invalidate execution. A 0.56 m radial guard supplements rectangular placement bounds. Qualification should include identical arm-command replay with open jaws, deliberate release, varied scene positions and contact/retention traces. Changing seeds with scene jitter disabled is repetition, not new scene coverage. Full accepted-domain reach checks, contact calibration, perception and physical robot transfer remain pending.

Frozen scene evaluation

The six-layout protocol was fixed in source before evaluation, using the same authored controller with no parameter training. It covers six different object/target coordinate pairs. The default and two layout families had been exercised during development, so these results are regression and scene-variation evidence rather than a claim of unseen generalization. Seeds 400–405 label the runs; jitter is disabled, and additional seeds alone would not create new physical layouts.

SceneOutcomeFinal target error (mm)Object lift (mm)Active simulation (s)Episode wall (s)
defaultsuccess0.286128.911.3100.642
cross_rightsuccess0.400128.213.3560.739
cross_leftsuccess0.353124.111.3340.567
forwardsuccess0.329129.011.5910.616
reversesuccess0.330124.713.5940.684
widesuccess0.437127.712.7150.665

Placement completed 6/6 scenes; no-motion and replayed joint commands with the jaws held open each completed 0/6. The open-jaw replay uses the successful run's exact joint targets, forcing an 80 mm jaw gap and holding the last command if the control horizon is longer. It provides a contact-coupling control, not the same object trajectory. Each failed control exhausts its 18-second horizon. Placement totals 73.90 active simulated seconds in 3.91 wall seconds, approximately 18.9×, on Apple M5 CPU. Wall time includes model compilation, reset/settling, IK, physics, observations and trace collection. UI, process startup, installation and output I/O are excluded. Another test process was running during this local measurement. No GPU throughput or minimum-memory requirement was measured.

Three of the qualified scene layouts also completed under macOS's process-level network-denial sandbox, after a native libc connection self-test returned EPERM. This tests installed runtime on this host and repeats existing layouts. It does not test downloading dependency wheels into an air-gapped server. An extracted-source reproduction, using the preinstalled pinned dependencies, matched all recorded control frames, commanded actions and final diagnostics exactly after JSON normalization for the first three placement scenes. Sixty-five distinct tests cover native contact execution, installation/HTTP boundaries, provenance integrity and the existing cable foundation.

Product and evidence

The public release provides a standalone arm replay, an executable source adapter archive, full control evidence, gzip JSON, compact metrics, integrity manifest, reproduction receipt, native offline receipt, and a local product guide. The replay samples recorded observations every 100 ms; full evidence retains 20 ms control observations and action commands. Hashes establish content consistency, not publisher authenticity or physical competence.

This implements the link-to-compatible-arm execution path within an explicit simulated skill space. The environment supplies metric object/target state, the installed adapter converts current state into joint/jaw actions, and the verifier checks the resulting object state. The initial thesis's spatial operating region becomes a bounded task contract; feedback supplies the missing connection between a skill name and effective movement. Demonstrations becoming learned policies, calibrated objects, realistic perception, repair policies and robot-specific physical qualification remain the next capability gates.