Local observation workflow and engineering feasibility pilot

DVIDIA Observation Lab pilot

· DVIDIA

Six public excerpts processed · accuracy not assessed · no robot actions

Source & history · Download Markdown

Abstract

DVIDIA Observation Lab supplies a local detector and palm-observation workflow with sampled timestamps, source/model provenance, editable timeline revisions and saved-review export. Two engineering attempts each process the same six public excerpts: 19 seconds and 38 sampled frames. Processing takes 4.835 and 4.635 seconds on one warm macOS ARM64 machine, excluding model setup, UI and human review. Software and browser checks establish integration; independent annotation accuracy, correction time, complete cup-placement coverage, depth, action learning and hardware competence remain unmeasured.

In this paper

Engineering integration passed on six public excerpts; annotation accuracy remains unmeasured. On 10 October 2026, we ran a local observation pipeline that decodes video, proposes sampled object and palm boxes, associates nearby detections and opens an editable review timeline. Review changes persist as revisions and export with their source identity. This establishes a working research tool, not reliable automatic task labeling, a cup-placement benchmark or a trained robot skill. The published implementation is anchored to DVIDIA Training commit d355a4a.

Install the Observation Lab release · GitHub release and downloads · Try the synthetic review demo · Frozen engineering receipts

The pilot produces reviewable 2D evidence

The local adapter uses the pinned OpenCV Zoo YOLOX COCO object detector and MediaPipe palm detector. It samples at two frames per second, up to 120 frames and a 640-pixel maximum image side per clip. Proposals retain decoded presentation timestamps, frame and source hashes, raw labels and scores. Same-class image overlap associates detections; palm/object overlap proposes 2D proximity. Neither association establishes physical identity nor overlap proves touching. Palm rectangles are not full-hand segmentation. No optical flow, depth reconstruction, natural-language action captioner or language model participates in this pilot. (Observation contract, frozen model and recipe)

The review interface shares one clock between the original video and timeline. It displays a box only within 0.25 seconds of its sampled timestamp, leaving gaps visible. Reviewers can correct labels and boundaries, approve or reject proposals and add missed events. Raw model labels remain separate from the accepted draft. Outcomes default to unknown; visible_completion and incomplete describe the operator's visual assessment. They are not robot-success measurements. (Review interface contract)

Saving appends a revision with conflict detection against the current base revision. A stale tab cannot silently replace newer edits. Every event needs a decision before review completion; unsaved changes trigger navigation protection, and export is disabled until saving. Downloads contain the last saved review, episode ID and original source hash. These checks make corrections inspectable. They do not authenticate the operator or turn agreement into independent ground truth. The loopback server performs no uploads and checks the local origin for writes. (Local workflow)

Six excerpts establish integration, not accuracy

The fixed input order was shirt, stir, pack, onion, melon and jeans, selected before examining model output. Five excerpts come from EgoAnnotate, credited to Taher Panbiharwala and Zainab Barwaniwala; the stir excerpt comes from CoMind, with its full author credit retained in the protocol. The records identify CC BY 4.0 terms and upstream source links. We hash-verified the exact excerpts; full upstream originals were not downloaded or hash-verified for this run. This publication distributes sanitized receipts rather than footage or sampled frames. (Input protocol and credits)

The six excerpts total 19 seconds, 2,753,287 bytes and 38 sampled frames. They are neither independent complete attempts nor the intended cup-placement cohort. Unknown session provenance remains one conservative shared group. There was no training split, learning run or physical evaluation. Publisher titles were display metadata and were not fed to the detector. (Episode summary, run and grouping)

Recipe attemptCompleted / attempted excerptsProcessing timeRecorded failures
Initial v16/64.835 seconds0
Hardened v1.16/64.635 seconds0

These are two single runs on the same inputs and a warm macOS ARM64 machine. Timing includes intake, copy/hash, decoding, detection, association and per-episode writes. It excludes model provisioning/loading, final receipt writing, UI startup and human review. The difference is not a controlled speedup. Peak RAM, energy, cost, p95 latency, minimum hardware and production throughput were not measured. No paid inference API was called. The runtime used Python 3.12.9, OpenCV 5.0.0.93 and NumPy 2.5.3. Execution requested four CPU threads and one inference worker, while frozen OpenCV metadata records ten threads; an enforced thread cap was not independently verified. No GPU path was requested or measured. (Initial receipt, hardened receipt, full frozen protocols, original benchmark description)

Workflow checks preserve failures and revisions

Review after the initial run identified a possible browser-clock mismatch for nonzero-start videos. Recipe v1.1 rejects nonzero container, video or decoded starts rather than silently moving boxes against an unchanged original. All six inputs have zero starts and were accepted in both runs. Initial artifacts remain alongside the hardened attempt. Same-size media replacement, malformed review IDs and oversized metadata also received workflow checks. These changes do not establish improved detector quality. (Original and revised recipes)

The implementation records 307 passing repository tests followed by 24 focused observation tests, including four additional detector-admission/suppression checks. Fifteen browser checks exercised playback, sampled overlays, correction save/reload, manual events, export, unsaved edits and completion validation on an isolated copy. The 390px and 320px layouts had no horizontal overflow; no browser console or page errors were recorded. These are interface and software checks, not physical-phone, Safari or annotation-accuracy qualification. Test edits do not count as independent human references. (Test scope, browser receipt)

Qualitative inspection found a saucepan labeled cup, storage jars labeled cups and missed palms in the preliminary stir check. These observations are not a scored accuracy estimate. COCO categories omit task-specific objects; overlap association can fragment or switch identities; two samples per second can miss brief actions. Model scores remain uncalibrated and all automatic outcomes remain unknown. Without independent references, inspection reports accuracy: null and accuracy_status: not_assessed. (Recorded limitations)

The release keeps ownership and missing evidence explicit

DVIDIA owns the orchestration, intake, timing, proposal association, review contract and workflow. Original DVIDIA Training code remains MIT. Detector preprocessing/postprocessing adapted from OpenCV Zoo retains Apache-2.0 attribution, and the pinned detector artifacts retain their separate upstream licenses and credits. The two ONNX files total about 40 MB and are provisioned explicitly with size/hash verification; inference does not download weights. Model terms do not grant rights to input footage. No new foundation model, SpatialLM, MANO or depth-model weights are claimed. (Exact artifacts and notices)

This lab is upstream of robot learning. Image rectangles supply no calibrated XYZ, contact forces, joint actions, gripper commands or robot policy. No automatic physical action is produced, and the existing visual predictor and movement head gain no new skill from this release. The next proposed study collects 30–50 complete permission-cleared cup-placement attempts with explicit session/person grouping, failures, occlusion and recovery. Independent blinded references should precede comparisons of detector-only, temporal-caption, combined and manual annotation at matched quality, correction time and full cost. Depth, a supervised action bridge and closed-loop qualification require separate experiments. (Uncollected cohort and gates)

The public Hugging Face Observation Lab Space reuses our review interface with two authored geometric videos and authored image boxes. It is a synthetic workflow demonstration: no detector, VLM or robot policy runs in the hosted page. Revisions stay in each visitor's browser, with a visible temporary-memory fallback if browser storage or cross-tab locking is unavailable. Visitors can edit, save and export their own review without server writes. The actual detector pipeline is installed and run locally on permitted footage. All sixteen hosted files were read anonymously at Space commit 8547b78 and compared with the staged hashes. The public runtime was observed running with its synthetic scope labels visible.

Publication preparation reran 311 repository tests and separately checked 22 publisher tests, six synthetic preparation/privacy/media tests and browser-local revision/conflict behavior. Both pinned ONNX detectors loaded and ran on a neutral image in the local CPU runtime; that smoke test measures neither detection accuracy nor task understanding. Release source preserves the original pilot receipts and clarifies the requested-versus-reported thread discrepancy. Original code uses MIT and the detector adapter retains its Apache-2.0 notices. Production uploads, accounts, storage and application analysis remain separate from this research release.