Product implementation and measured synthetic pilot

Skillspace footage training pipeline

· DVIDIA

Public CPU package · one native scene qualified · simulation-only

Source & history · Download Markdown

Abstract

The public Dvidia-Inference/dvidia-training repository provides a standalone CPU pipeline that audits Skillspace media, freezes grouped splits and exports inspectable artifacts. Ten synthetic six-second clips train a visual predictor with a 0.5106-second median pipeline time across three repeats on one macOS host. Aligned native simulator telemetry supplies 4,029 training labels for a learned joint-delta head, evaluated on one held-out recording. Its capsule passes native placement qualification in that recording's exact scene; identical joint targets with open jaws fail. Installation and hardware guides accompany the package. Human-video action recovery, shoe organization, demonstration sufficiency, independent scene generalization and physical robot competence remain unestablished.

In this paper

The public, standalone DVIDIA training repository provides a local CPU pipeline that turns actual Skillspace media into a fitted visual model and, when aligned numeric actions are supplied, a learned movement head. The useful advance is a reproducible path from source files through provenance checks, grouped splits, measured fitting and portable artifacts. Video prediction and robot movement remain separate claims. Ten synthetic six-second clips demonstrate visual fitting; simulator telemetry supplies the movement pilot's action labels. Its exported capsule installs and passes native qualification in one exact scene, while identical joint targets with open jaws fail. Neither experiment learns shoe organization or reconstructs robot actions from human footage. This report documents the implemented path and its limited measured evidence, rather than establishing a sufficient demonstration count or hardware competence. (Standalone source archive, Qualification receipt)

Actual media and connected source groups gate training

Start with the repository's installation guide and hardware guide. The pipeline accepts a local folder, ZIP or prepared dataset; explicitly enabled online mode also accepts a supported public DVIDIA manifest and its declared media. Visual training requires Python 3.12+, NumPy and FFmpeg/ffprobe; arm simulation adds MuJoCo. A Skillspace description with no accessible video does not become a dataset. Images alone are rejected. An audit-only preparation step checks media, copies the exact bytes, records hashes and freezes train, development and test assignments before fitting. Fresh output directories prevent silent replacement of earlier runs. The command-line interface and a loopback training studio expose this same path. (Implementation)

Split integrity depends on connected source groups, rather than independent filenames. Clips sharing a source recording, session, shoe-pair identifier or identical media bytes stay together. At least three groups are required to populate all three splits. Plain video folders can enter with explicit lineage warnings, but their per-file groups do not prove that recordings are independent. A supplied “complete” flag records the contributor's declaration; it does not establish that the entire task appears or that the demonstration is correct. Thus the implementation prevents declared lineage from crossing splits while retaining the unresolved need for semantic and provenance review. (Example visual Skillspace)

Ten synthetic clips establish visual fitting, not a shoe skill

The visual pilot uses ten encoded six-second moving-disc clips, with seven training, two development and one test recording. Each contributes twelve sampled frames. The learner projects low-resolution RGB through a basis fitted only on training data, then fits an affine ridge transition between successive representations. Development error selects regularization; the untouched test recording measures prediction against persistence and a training-mean baseline. This compact CPU learner checks the training path without requiring a large pretrained vision model. (Visual measurements)

Three standalone repeats on a macOS Apple M5 host with 24 GiB RAM take 0.5315, 0.5106 and 0.5102 seconds for intake, fitting and export: median 0.5106 seconds. Dependency installation, imports and fixture creation are excluded; integrity inspection is timed separately. The Python process lifetime peak RSS is 48.13 MiB, excluding FFmpeg/ffprobe children. Repeats reuse the same test recording and establish repeatability, not additional success evidence. These measurements describe this small default learner; they do not establish minimum hardware or scaling to larger video models. (CPU benchmark receipt)

Test measurementLearned modelBaselineEvidence unit
Normalized RGB mean squared error0.00030390.0017714, persistenceOne video, eleven transitions
Joint-delta mean squared error, rad²5.768 × 10⁻⁷8.173 × 10⁻⁵, zero deltaOne telemetry recording, 621 samples

The visual result supports a narrow conclusion: the fitted transition predicts these synthetic image samples better than copying the preceding frame. Eleven correlated transitions are not eleven independent task trials. Pixel error measures neither grasping nor placement success, and the model observes no calibrated depth, contact force or human action. No shoe demonstration was used. The exported visual artifact contains numeric weights, a model contract, measurements and a dataset receipt; original media remain in the local prepared dataset. An integrity inspection checks the exported files and ZIP hashes. (Visual measurements, Training implementation)

Executable movement requires aligned actions and an adapter

Movement fitting is conditional on action sidecars with timestamped context, target error and six joint deltas in the existing feature contract. Each sidecar binds its media and samples by hash and declares an alignment method and supporting evidence. Accepted origins include teleoperation, a validated video-action bridge and native simulator telemetry. These fields check internal consistency; a declaration does not independently validate action recovery from pixels. The standalone pilot uses the third origin: ten separately executed native simulator rollouts with state and authored targets, accompanied by schematic footage. (Native telemetry Skillspace)

Its grouped split contains seven training, two development and one test recording, totaling 116.5 seconds of footage. A normalized radial-basis ridge model fits 4,029 training labels, chooses regularization on development data and reaches the joint-delta error shown above on 621 test samples. This is an offline prediction result from one held-out recording. It does not measure closed-loop success, independent scene generalization or improvement over the earlier controller. The learned head maps privileged simulator context and a requested goal to joint deltas; phases, waypoints, grasping and recovery remain authored. This scope follows the earlier skill-capsule prototype. (Movement measurements, Run receipt, Skill capsule acquisition)

An explicit supported task source permits export of a runtime-bound simulation capsule. The standalone run records 1.482 seconds on its CPU pipeline timer, excluding final ZIP creation and integrity verification. It exports a 14,843-byte candidate, passes inspection and installs it. This describes one small run, not a throughput comparison. An earlier filename-handoff failure remains recorded in the release ledger. (Run receipt, Release ledger)

Installation alone does not enable qualified execution. A separate native test in the exact scene of the held-out recording completes placement in 12.952 simulated seconds with 0.311 mm target error. Replaying the same joint targets with open jaws fails at the 18-second horizon. This supports the contact-dependent execution path in that scene; it adds no independent scene sample or hardware evidence. Artifact hashes establish which candidate and runtime were tested, not physical accuracy. (Qualification receipt, Native replay)

The next experiment should measure the missing bridge

The current local path avoids API dependence once its tools and data are installed. Local media decoding restricts FFmpeg protocols, while optional remote intake is explicit. Separate continuous-integration checks passed visual fitting and studio tests on Linux, macOS and Windows with Python 3.12 and 3.13, the optional Linux arm pilot, and an actual Docker example, training run and inspection under --network none. The image build required network access; this does not establish installation on a new air-gapped machine or minimum hardware. No GPU scaling, throughput improvement over another implementation or hardware transfer follows from the pilots. The package makes experiments inspectable and repeatable; its scientific value depends on what demonstrations, labels and task environments enter it. (Passing software checks, Source and training guide)

For shoe organization, the next decisive step is a reviewed corpus of complete one-shoe-to-marked-slot attempts and a validated observation-to-action bridge for a defined arm, gripper and sensing arrangement. Sample-count comparisons should hold architecture, task definition and evaluation fixed, group related originals, and report independent closed-loop episodes alongside prediction error. Until those components exist, a shorter visual training run cannot answer how many human videos make a robot competent. The pipeline's contribution is to preserve those distinctions in executable product gates, making the missing evidence measurable before a candidate becomes a downloadable skill.