Pinned mechanics implementation and CPU numerical benchmark
Hardware matched robot simulation
18/18 numerical runs valid · physical fidelity unmeasured
Abstract
A separate UR5e/Robotiq reference preserves pinned geometry, inertias, moving-link collision proxies and coupled gripper structure: fourteen modeled coordinates and seven commands. Eighteen deterministic unloaded motion/closure runs are numerically valid; at a 1 ms step, CPU physics throughput measures 40.80–47.29 simulated seconds per measured stepping wall second on one host, excluding loading, sensing, rendering, policy and real actuation. Adjacent timestep comparisons and a matched static overlap expose numerical sensitivity and collision participation. A pure-Python paired trace scorer measures joint, tool and jaw mismatch without automatic qualification. Left/right five-finger references are separately audited; LeRobot collection with an original DVIDIA policy is proposed. Serial calibration, physical response, shoe task learning and bimanual execution remain unmeasured.
In this paper
DVIDIA research · 11 October 2026 · Compiled reference assembly and unloaded CPU diagnostics; physical transfer unmeasured
Build the first simulator around a specific six-axis arm and gripper, then measure the physical response that the learner will encounter. DVIDIA's vision remains a Skillspace that supplies demonstrations, local training and an installable skill. Rendered images support visual learning, but matching a room's appearance and the arm's number of joints does not establish that a learned action will produce the same motion, grip or outcome on hardware. We have compiled a pinned UR5e and Robotiq 2F-85 reference and run eighteen unloaded CPU diagnostics; this does not select the user's eventual purchase or reproduce that robot's controller. Later, two arms with five-finger hands need a different command, sensing and contact contract. Imported structure, numerical characterization and held-out physical agreement remain separate evidence states. Useful speed is the cost of acquiring a qualified capability, measured alongside errors and failures.
Six arm joints do not specify the whole robot
DVIDIA's existing dvidia-authored-6dof-parallel-jaw-v0 environment is an authored rigid-box contact experiment. It has six arm hinges, two simulated sliding jaws and eight native position actuators, exposed through six joint targets and one total-gap target. Arm links, base and palm have collision masks disabled; the experiment therefore does not evaluate whole-arm avoidance or self-collision. Exact object state, contact counts and pad forces are native simulator observations, not measured camera or fingertip signals. Preserve this profile and its frozen results while adding a separately named hardware reference. (Current arm source)
The UR5e is a suitable six-axis software reference. Its manufacturer lists six rotating joints, 850 mm reach and 5 kg payload. Those specifications describe hardware capability; the listed 0.03 mm repeatability does not mean that a generic simulator predicts tool position to 0.03 mm. Menagerie's UR5e derives from a URDF, uses simplified collision shapes and adds position actuators. It supplies a reproducible mechanical description rather than the proprietary controller's complete response. (UR5e specifications, Pinned model derivation)
The new reference attaches the gripper through MuJoCo's spec interface with a stable prefix and the upstream attachment transform. The pinned original 2F-85 model has eight linkage hinges, a split tendon and loop-closing constraints, driven by one actuator with a 0–255 command range. The compiled assembly confirms fourteen position coordinates, fourteen velocity coordinates and seven actuator channels, with three equality objects and one tendon. Gripper coordinates are coupled linkage motion, not eight independent finger commands. A requested physical jaw gap requires a calibrated conversion, and the model's tendon force limit is not the physical gripper's fingertip force setting. (MuJoCo attachment API, Pinned arm XML, Pinned gripper XML, DVIDIA reference tools)
The assembly has 24 active collision geoms. Each of its six moving arm links has a collision proxy; the fixed base has none, and exclusions/filtering still affect self-contact. It has zero native sensor channels. Recorded poses, velocities and contact counts are privileged simulator state. This structural improvement over the authored box experiment does not establish complete collision geometry, physical sensing or a shoe task. (Reference audit receipt)
Pin the asset revision 0059d4335f8156206f63a35662313385f7ad6d74, model and mesh hashes, derivation notes and per-directory licenses. The arm's BSD-3-Clause and gripper's BSD-2-Clause notices accompany any redistributed assets; DVIDIA's original MIT software license does not replace them. A cached dependency bundle supports local operation after provisioning. Downloading the bundle is a separate setup step. (Pinned revision, Arm license, Gripper license)
Match observable mechanics before calling it fidelity
The first physical calibration should fit base and tool transforms and serial-specific kinematics on measured poses, then score untouched poses throughout the intended tabletop workspace. Universal Robots warns that using generic calibration can introduce end-effector discrepancies of centimeters. Report independently measured tool-position and orientation residuals, reference uncertainty, operating conditions and the calibration revision. Joint encoder agreement alone cannot reveal a wrong base or tool transform. (UR calibration documentation)
Next, record issued commands and actual joint motion during bounded unloaded trajectories, then repeat under declared payloads. Compare onset delay, tracking residuals, overshoot and settling. UR's RTDE interface supplies timestamps, target and actual joint states, TCP data and safety status; its e-series output can reach 500 Hz, while busy controllers can skip packages. That interface rate does not establish a 2 ms application reaction time. The servoj command's gain and lookahead change response, so the reference model's affine servos cannot silently stand in for every controller setting. (RTDE guide, Official servoj manual)
Gripper calibration adds opening versus encoder value, opening/closing hysteresis, speed, response delay and contact behavior. Robotiq's hardware interface requests position, speed and force and reports encoder position, current, object detection and faults. Its current-driven detection can miss thin objects; current is not calibrated tactile force. Measure force against a suitable reference and retain fingertip material, wear and command settings. Menagerie's derivation describes increased pad friction and historical solver choices. The actual pinned XML and DVIDIA import use an elliptic contact cone and impratio=10; DVIDIA retains the default zero no-slip iterations. Numerically stable gripping still needs independent slip and release tests. (Robotiq control manual, Pinned gripper derivation, Reference implementation)
DVIDIA's proposed fidelity gate compares aligned simulator and physical traces under the same issued-command history, with calibration and confirmation sessions separated. The new pure-Python paired-trace scorer measures named joint, tool translation, rotation and explicitly defined inner-gap residuals. Its synthetic fixture tests the scoring contract; it is not a hardware dataset. The importer must establish clock/frame alignment and command provenance, which the scorer does not authenticate or infer. Contact onset, slip and load retention still need task-specific evidence. A timestep comparison between h and h/2 tests numerical sensitivity; it cannot establish agreement with the purchased robot. Hardware accuracy remains unmeasured until physical reference data exists. (Paired scoring contract, Scorer source)
Five fingers require a new command and feedback contract
Finger count does not determine independent motion. Shadow's technical specification describes 24 movements and 20 motor modules, coupled distal finger joints, two wrist joints and a 4.3 kg hand/forearm assembly. Its host and internal motor control rates differ, as do position, tactile, tendon-force and electrical sensor update rates. DVIDIA has separately compiled and audited the pinned left and right Shadow references: each has 24 joint coordinates, 20 actuators, four fixed tendons and zero native sensor channels. A summed-angle tendon is not a rigid equal-angle joint coupling or a complete model of the physical transmission. The hands are not yet attached to a bimanual DVIDIA assembly; their Apache-2.0 asset license and notices remain separate from DVIDIA software. (Shadow specification, Pinned Shadow model, Asset license, Audit receipt)
The later embodiment should describe each arm and hand separately: component identity, side, attachment frame, mounting transform, named joints, independent actuators, transmissions, passive/coupled motion and sensors. Artifact compatibility binds that ordering and profile hash. The existing six-joint/one-gap transition model must reject an incompatible hand contract rather than pad or truncate its input. Two six-joint arms plus two Shadow model hands would imply 60 articulated coordinates and 52 actuator channels, including the hands' wrist joints; this is conditional arithmetic, not a compiled or qualified DVIDIA bimanual system. Hand weight also changes the available payload and mounting requirements, which must be checked on the actual hardware.
Build shared-frame and time contracts before coordination. Both arms need explicit transforms into the table/world frame, while observations retain acquisition time, arrival time, clock domain, offset uncertainty and age. Qualify each component first, then arm-arm, arm-hand, hand-hand and environment collisions. Start shared-object tasks with hold-and-reposition, handoff and coordinated placement. Include internal loads, clearance, failed handoffs, release synchronization and stale-sensor response. Shadow's official description supports configurable bimanual arm/hand assemblies, but that composition support is not DVIDIA transfer evidence. (Shadow descriptions, ROS frame conventions)
Keep solver contact truth separate from policy observations. MuJoCo's touch sensor sums normal contact forces inside a site; it does not reproduce a spatial force/shear array. A hardware sensor adapter needs calibrated response, update interval, delay, filtering, quantization, saturation and missing-data behavior. Neither adding a hand mesh nor adding a scalar contact sensor validates shoe or lace deformation. (MuJoCo touch reference)
Offline training needs actions and separate speed clocks
For the first shoe skill, define one bounded goal: acquire one shoe and place it stably in a marked slot. Skillspace footage supplies appearances, phases, outcomes and coverage gaps. Robot action supervision needs synchronized commands and feedback, or an independently evaluated action bridge. Complete attempts retain failed grasps, interventions and recoveries. Separate physical shoe-pair/session groups for training, development and confirmation; cut clips and synthetic descendants retain their parent lineage. Once a confirmation set has informed model changes, it becomes regression evidence and fresh confirmation is needed. These are proposed DVIDIA protocol requirements, not a claim that a video quota establishes skill coverage.
Train a small action policy on the declared observation and command interface, then compare it with simple matched baselines in a shoe-specific environment. An offline simulator can expand practice, but the shoe's geometry, compliance, friction and laces need their own identified envelope. Image-only practice does not uniquely establish hidden friction or grip stability. Work on uncertainty-aware parameter inference and simulation calibration provides precedents for identifying plausible dynamics and refining them with physical observations. (BayesSim, SimOpt)
The first performance report should separate cold loading/compilation, warm physics, rendered observations, policy computation, transport and actuator delay. Report named hardware, threads, engine settings, contact counts, memory, warnings, missed deadlines and p50/p95/worst timing. A fast next-state predictor is not a measurement of hand physics, video generation or sensor-to-motion response. CPU operation is a useful first target; batched GPU practice can follow a matched task benchmark. Additional servers add rollout capacity rather than automatically pooling memory or accelerating one contact-rich scene.
Use fixed issued-command replays to expose timestep sensitivity, then measure complete task attempts with resets, observations and failures charged to the cost. A physics-only benchmark supports a narrow simulator-speed statement. A physical skill claim additionally needs fresh hardware trials, independently observable success, critical-error and recovery criteria and a declared operating envelope. No hardware minimum or transfer-success target can be justified from rendered-image count alone.
LeRobot can collect evidence while DVIDIA owns the model
LeRobot offers hardware/teleoperation interfaces, cameras and dataset tooling without requiring a supplied policy architecture. Its official custom-policy interface supports independently packaged models with training, reset and action-selection methods. DVIDIA can reuse that infrastructure while developing its own objectives, model and skill artifacts. No DVIDIA/LeRobot integration or robot action policy is built by this reference release. The earlier compact transition predictor forecasts motion conditional on a command; it does not choose the command. A proposed first seam is a validated local dataset reader and strict observation/action adapter, followed by an original DVIDIA trainer. (Custom hardware guide, Custom policy guide)
Hardware selection remains explicit. SO-101's six motor channels are shoulder pan, shoulder lift, elbow flex, wrist flex, wrist roll and gripper: five arm axes plus a gripper, not six arm axes plus a gripper. Its body positions can use degrees or normalized values, and the jaw uses a normalized command. A separate named profile must specify ordering, units, calibration, action representation and requested versus actually sent targets; it cannot accept DVIDIA's current six-axis/radian artifact unchanged. (SO-101 guide, Follower implementation, Action representations)
LeRobotDataset v3 stores state/action rows, camera video and episode/schema metadata, and recordings can stay local. However, the recording guide warns that frame-index/FPS timestamps can make motion appear faster when acquisition misses the requested cadence. DVIDIA needs a sidecar with actual monotonic acquisition, arrival, command and feedback clocks, camera calibration, missed samples, controller settings and physical source groups. Nominal timestamps alone cannot fit actuator delay or validate physical speed. Pin and test the chosen LeRobot version and review datasets/assets separately from its Apache-2.0 software license. (Dataset format, Recording and cadence guide, Software license)
Eighteen unloaded CPU runs characterize this reference
The new native benchmark attempts three synthetic cases—hold-open, base-joint sweep and jaw cycle—at 2, 1 and 0.5 ms timesteps, with two deterministic repeats and four simulated seconds per attempt. All 18 attempts are numerically valid with no captured warnings. Commands update at 20 ms boundaries. The measured platform is macOS 26.5.2 arm64, Python 3.12.9 and MuJoCo 3.15.0; these runs used CPU physics. They contain no task object, rendering, perception, learned policy or hardware actuation. Repeats are repeated numerical/performance diagnostics, not independent physical trials. (Interactive report, Exact receipt)
At the 1 ms setting, the six runs produce the following measurements. Each 20 ms chunk includes native stepping, per-tick instrumentation and a boundary mj_forward. The throughput ratio divides four simulated seconds by accumulated chunk wall time; it excludes asset verification/compilation, settling, sampling, rendering, perception, policy, transport and physical actuation. (Timing receipt)
| Case | Repeat index | Simulated seconds / chunk wall second | 20 ms chunk p95 | Worst chunk |
|---|---|---|---|---|
| Hold-open | 0 | 44.41× | 0.574 ms | 0.807 ms |
| Hold-open | 1 | 47.29× | 0.562 ms | 0.852 ms |
| Arm sweep | 0 | 44.32× | 0.712 ms | 1.137 ms |
| Arm sweep | 1 | 40.80× | 0.709 ms | 0.954 ms |
| Jaw cycle | 0 | 41.79× | 0.686 ms | 0.941 ms |
| Jaw cycle | 1 | 42.83× | 0.683 ms | 1.408 ms |
Adjacent-timestep comparisons align 201 elapsed boundaries per comparison, including the post-settling start. For repeat index 0, halving the timestep reduces arm-sweep tool-translation and jaw-cycle pad-center differences as shown below. Pad-center separation is not calibrated inner jaw gap. These residuals compare two numerical resolutions of the same model, not the model with reality. Both deterministic repeats are retained in the receipt. (Numerical comparisons)
| Case and measured quantity | Timesteps compared | RMSE | Worst difference |
|---|---|---|---|
| Arm sweep, reference tool translation | 2 versus 1 ms | 0.07424 mm | 0.11504 mm |
| Arm sweep, reference tool translation | 1 versus 0.5 ms | 0.03702 mm | 0.05738 mm |
| Jaw cycle, pad-center separation | 2 versus 1 ms | 0.03448 mm | 0.05043 mm |
| Jaw cycle, pad-center separation | 1 versus 0.5 ms | 0.01733 mm | 0.02538 mm |
A separate authored sphere-overlap diagnostic produces one arm contact with active collision masks and zero with those masks disabled. It checks that the selected proxy participates in collision detection; it is not a moving collision-response or avoidance test. The software checkpoint also passed 488 local tests, including native contract checks. CPU-only headless characterization is demonstrated for this unloaded reference workload; shoe contact, complete training throughput, minimum hardware, physical accuracy and the composed two-arm system remain unmeasured. (Collision receipt, Reference instructions)
Release an evidence-bearing skill, not an appearance match
The installable unit should carry the policy, skill goal, observations/actions and units, named embodiment ordering, adapter, dependency hashes, calibration requirements, supported conditions, refusal/recovery behavior and benchmark receipts. Imported asset structure, characterized simulation and measured hardware agreement advance separately. A reference-only or simulation-qualified package remains useful for research while declaring physical execution unqualified. This follows the current DVIDIA roadmap's separation of collection, learning, simulation and physical qualification. (Execution map)
The product opportunity is an efficient route from missing evidence to a tested capability. A hardware-matched reference gives that route a concrete target; measured residuals identify what to fix or collect next. When simulation succeeds and hardware fails, the discrepancy becomes a reproducible investigation rather than a reason to gather undirected images. That feedback loop is how Skillspace can accumulate reusable knowledge across skills without pretending that one visual dataset or one embodiment solves every interaction.
Published by DVIDIA. Contact hello@dvidia.org and follow @iammrriver.