Decision report and local experiment
Fast accurate robot training simulation
54 CPU fixture trials · GPU and physical transfer untested
Abstract
A design for an offline, self-hosted training environment that turns demonstrations into calibrated practice and qualified robot skills. We examine contact, deformation, sensor feedback and the performance–accuracy tradeoff, then report 54 deterministic native MuJoCo CPU fixture trials. Fine endpoint agreement did not establish accurate contact transients, and coarse timesteps missed collisions. Minimum GPU memory, rope accuracy and physical transfer remain unmeasured.
In this paper
Decision report · 7 October 2026 · Grok 4.7 research plus primary-source verification · 54 local CPU fixture trials; no GPU training or physical robot tests
Build an offline, self-hosted environment that converts demonstrations into calibrated practice and qualified robot skills. Start with existing physics engines and concentrate the product on task construction, uncertainty, failure/recovery coverage, efficient scheduling and qualification. Use native MuJoCo for local numerical checks, test MuJoCo Warp for batched rigid contact, and qualify Newton rods or another deformable backend against the actual rope task before selecting it. The microscopic intuition explains why objects resist and deform; the useful simulation interface is the resulting force, motion, friction and sensor response at robot scale. The first product vertical should be bounded rope grasping, pulling, threading and regrasping, with bow tying a later capability. One owned GPU is the starting design target; minimum useful VRAM and an “ultrafast” training rate remain unmeasured. Adding GPUs or servers should scale independent practice worlds, while a released skill retains a declared hardware profile and physical evidence.
Reproduce the response that changes the robot's decision
When a fingertip presses a table, resistance ultimately comes from electromagnetic interactions and quantum constraints on matter. “Electrons reject each other” is a useful intuition but incomplete: atoms also attract and bond, and Pauli exclusion restricts allowed electron states rather than adding a separate classical repulsive force. Molecular structure and collective interactions produce elasticity. We do not need to track electrons to calculate whether a gripper holds a lace; we need the effective mechanical behavior at the force, length and time scales the task uses. (Feynman on forces and elasticity, exclusion and matter)
The computational sequence is action → contact and deformation → new state → sensed feedback → corrected action. A rigid model approximates bodies whose deformation is negligible for the chosen tolerance. Compliant contact produces force as surfaces compress. A rod model represents a slender object's centerline, stretch, bend and twist; a volumetric model spends more computation on internal deformation. Contact force can be understood as stress integrated over the contacting surface. Perfect rigidity is an approximation, and rigid Coulomb contact can have ambiguous or nonexistent mathematical solutions. Compliance serves physical and numerical purposes. (Drake contact modeling)
For laces, the required distinctions include usable length, thickness, bending/twisting, tension, pad/lace and lace/lace friction, grasp slip, crossings and hidden segments. Tight knots can introduce compression, hysteresis and internal yarn slip that a simple elastic rod misses. Add detail when experiments show it changes the next action or final outcome. Discrete Elastic Rods demonstrates a principled reduction, keeping dynamic bending while treating fast twist quasistatically; that does not establish every textile contact law. (Original DER paper)
Visible motion does not uniquely identify stiffness, preload or friction. Preserve uncertainty in hidden state separately from uncertainty in material parameters; inspection helps the former, calibrated probing constrains the latter. Parameter-inference work explicitly permits multiple explanations of the same observed dynamics. (BayesSim)
Fingertip feedback also requires a sensor model: what force, image, shear, delay and noise reach the controller? A convincing tactile render is insufficient. TACTO explicitly delegates dynamics and disclaims accurate deformation/friction, whereas DiffTactile couples elastomer deformation, object dynamics and optical feedback. These represent different fidelity/cost choices. (TACTO, DiffTactile)
“Correct physics” therefore means bounded error over declared interventions and operating conditions. Test force/displacement, slip onset, contact timing, impulse, rope shape, tension and the actual sensor stream. Smaller timesteps check numerical convergence; a converged model can still describe the wrong physical behavior. Physical rod-validation protocols find contact formulations that remain inaccurate despite refinement. A finer simulator is a numerical reference; measured material and robot experiments are the physical reference. (Validation protocols)
Preserve the observation-to-skill thesis inside the product
The proposed sequence becomes a concrete contract. Observation supplies hypotheses about goals and phases; language retrieves them; spatial expectations direct inspection; physics expands executable practice; feedback handles failures; qualification determines what can be downloaded. The command-conditioned spatial prior ranks where evidence is likely, then broadens after contradictory observation. XYZ helps locate the task but cannot identify contact mode, orientation, tension or knot security by itself.
| Original idea | Proposed product component | What must be established |
|---|---|---|
| Learn by watching | Demonstration intake, effects/phases and uncertainty | What was observed versus inferred; robot-action bridge |
| Name the skill | Goal predicates and language aliases | Intended object, completion and already-completed states |
| Shoes below, tables near torso | Body/object-relative inspection prior | Total latency including estimation; unusual placements |
| Each skill has an operating space | State/material/embodiment envelope | Reach, contacts, sensing and unsupported starts |
| Simulate variants | Parameterized scenes and feasible execution | Accepted coverage, numerical validity and generation cost |
| Trial, solve and remember | Feedback policy and validated recovery branches | Outcome detection, retry limits and recovery success |
| Download the skill | Policy/procedure, compatibility manifest and evidence | Qualification on the named hardware profile |
The first engineering vertical is one gripper, one rope class and one known fixture: grasp, pull, pass an end through an eyelet, detect lost purchase and regrip. A fixture supplies the support that another hand might provide; it is part of the profile. Extend later to self-contact, loops, knots and secure releasable bows. An overhand knot surviving a pull is a different goal from a shoe bow. Demonstration count is an adaptation budget to investigate, not a declaration that all possibilities have been covered.
Close prior art shows why this scope is useful. SoftMimicGen expands teleoperated demonstrations through non-rigid registration; its physical rope result is 33.3% from 1,000 simulated demonstrations, 46.7% from 30 physical demonstrations, and 76.6% from combined training. Its simulated rope benchmark is near 100% under a different protocol. It assumes fixed subtask sequences and explicitly leaves conditional transitions for future work. Recovery coverage and qualification are therefore consequential product work. Its public implementation is a build-on resource, not evidence we reproduced it. (Paper, code)
SILO uses fast articulated-capsule approximations, localized routing RL and known fixtures, reporting 18/24 three-harness successes on nylon rope. Deployment synchronizes a simulated robot with the real robot, while real perception supplies cable observations. Its start-stop execution is quasi-static; its older-baseline comparison is expressly unmatched. This supports reduced-model practice plus feedback, with a runtime simulator dependency. A simulator-free controller is a separate export route to validate. (SILO) DLO-Lab instead calibrates a differentiable rod model and reports 7/12 physical ring-threading successes, using calibrated colored point observations and proprioception. Neither establishes universal bow competence. (DLO-Lab)
Select a backend through task gates and pin its dependencies
The reversible decision is native MuJoCo for analytical/numerical fixtures, direct MuJoCo Warp for the first NVIDIA batch candidate, and a separate deformable qualification gate. This avoids writing a new general physics engine before establishing the product advantage. Newton VBD is the leading rod challenger to test; Genesis supplies another deformable/platform option. Isaac Lab 3 already provides substantial training infrastructure, so a new product should reuse suitable plumbing. No engine is proven fastest or physically sufficient for our rope workload.
| Candidate and version anchor | Dependency/platform boundary | Selection condition |
|---|---|---|
| MuJoCo / MJWarp 3.15.0, Oct 5 | Direct MJWarp: Python ≥3.10, MuJoCo ≥3.12, NumPy, Warp ≥1.15; NVIDIA batch path | Fixed rigid-contact tapes, capacity checks and latency/throughput measurement. (MuJoCo release, MJWarp release/manifest) |
| Newton 1.6.1, Oct 5 | Warp ≥1.17; its sim dependency line uses MuJoCo/MJWarp 3.12; macOS CPU only | Rod/contact/rigid coupling and topology qualification; VBD remains experimental. (Releases, solver matrix) |
| Isaac Lab 3.0.0-EA, Sep 16 | Tested Python 3.12, PyTorch 2.11, Warp 1.16, Newton 1.5.2, Isaac Sim 6.1; kit-less supported paths | Use supported backend/task combinations when they reduce integration work; retain that dependency set. (EA release) |
| Isaac Lab 2.3.2, Feb 2 | Latest retrieved non-EA release | Stable alternative without assigning it 3.0's new interface/features. (Release) |
| Genesis World 1.4.3, Sep 30 | CPU/CUDA/AMD/Metal base paths; optional IPC requires x86 Linux/Windows NVIDIA | Matched contact/deformable tests and per-solver platform manifest. (Releases, README) |
These are compatibility anchors, not interchangeable ingredients. Updating standalone MuJoCo to 3.15 does not establish a tested Newton/Isaac Lab stack. MuJoCo's CPU cable plugin models inextensible bending/twist; GPU parity must be exercised, and the MJX table distinguishes JAX from Warp flex support. Newton's VBD exposes stretch/shear/bend/twist but its self-contact buffers can silently omit pairs when undersized. Its optional external-rigid integration is one-way, distinct from internal VBD arrangements; verify force exchange for the chosen arrangement. (Cable plugin, MJX matrix, VBD API)
Fallback decisions stay task-specific. If an MJWarp fixture fails numerical checks, retain CPU verification and revise the GPU configuration; do not enlarge tolerated contact errors to win throughput. If a reduced rope model fails physical slip/tension tests, qualify another representation or backend. ManiSkill is a useful manipulation harness, but its Linux/NVIDIA GPU path and CC BY-NC asset terms require an explicit content choice. (ManiSkill matrix)
Make useful practice fast on one owned GPU before stacking servers
The product should operate fully locally after setup, including training, qualification and export. Package runtime dependencies, assets, material priors, datasets, model weights/tokenizers and compiled-kernel caches with hashes. Test a single-node run with network disabled and a multi-server run with external internet blocked while local cluster communication remains available; initial downloads or pre-staged installation are a separate step. Hosted AI is unnecessary during execution: explicit goals and annotations remain available, with optional cached local models for assistance. Grok supports online research here; it is not a product runtime dependency.
The minimum candidate is a headless state-based worker on one owned NVIDIA GPU, with a small policy and CPU authoring/audits. Direct MJWarp physics does not inherently require JAX, PyTorch, Isaac Sim or a viewer; training adds a learner separately. Its docs target aggregate throughput and warn that a single world can be slower than native MuJoCo. (Package manifest, performance guide) 8, 12 and 16 GB are memory tiers to test, not verified minima or purchase recommendations.
State-only training is an initial performance profile. A teacher can inspect privileged simulator state, but physical qualification requires a calibrated perception/sensor bridge or a student policy trained on the declared observations. Exact simulated rope points, contact flags or tension do not become camera/tactile measurements merely by exporting weights. Charge that bridge's errors, latency and memory to the deployment profile.
Record resident engine/scene allocations, world/contact/constraint buffers, observation tensors, policy/optimizer weights, learner activations, replay and temporary compile/render peaks. A proposed governor grows batch size using measured peak allocation and a declared reserve; on capacity pressure it reduces worlds or schedules learning/rendering separately. It must never silently truncate contacts to fit memory. Keep simulation and policy tensors resident; synthesize observations only at required rates. A state profile, point-cloud profile and RGB/tactile profile receive different memory and performance contracts.
On multiple machines, a local queue leases immutable task/scene/seed packs to node agents. Independent rollout shards tag experience with policy versions; learner workers consume it and publish checkpoints. Adding servers adds worlds, not automatically pooled VRAM or a distributed solver for one world. Measure communication, stale-policy experience and learner contention. MJWarp documents per-device graph creation; cluster scheduling and qualification remain proposed product components. (Multi-GPU guide)
Speed claims require normalized clocks. Genesis's 43-million-FPS comment accompanies a script with 30,000 worlds, 0.01-second steps, a plane/Franka, one position command and no policy-training or sensor loop. Arithmetic gives roughly 14.3× real time per world under those assumed conditions; the huge aggregate sums worlds. It is not lace throughput. (Benchmark source)
| Clock | Frozen workload and charged costs |
|---|---|
| Cold setup | Generation/reconstruction, rejected candidates, loading and compilation |
| Warm physics | Robot/action semantics, material/contact model, timestep/substeps, solver tolerances, precision, contact capacities and batch |
| Interactive latency | One-world p50/p95 with the chosen control/observation rates |
| Valid experience | Physics, observations, inference, goal checks, resets and serialization; invalid runs still consume cost |
| Learning | Updates, optimizer/activations, synchronization and audit replays |
| Product outcome | End-to-end time/compute/physical data to a frozen held-out capability |
Cheap broad practice, targeted boundary/recovery starts and expensive audits form a proposed scheduler. Active Domain Randomization and multifidelity RL are relevant precedents. Maintain an independent fixed evaluation distribution; training's selected hard states do not define reported deployment success. Check alternate-action consequences and model disagreement near slip/crossings. Fidelity changes need state mapping; start with whole-episode selection or short replays, since switching representations mid-contact can inject artifacts. (Active Domain Randomization, MIT multifidelity work)
Qualification supplies the evidence that a download can carry
The first local scaffold tests settled normal load, Coulomb stick/slip and thin-wall collision detection across timestep settings. An analytical acceleration/stopping law verifies a declared approximation; a fine-step replay checks numerical consistency. Neither establishes real fingertip or lace behavior. The next fixture should measure rope tension/extension and bend/twist before adding grasp/slip and self-contact. Axial validation needs a declared extensible law; the native inextensible cable abstraction cannot supply that test.
We built and independently reviewed a local foundation: 54 deterministic trials, 18 cases × three repeats, using MuJoCo 3.15.0 and Python 3.12.9 on Apple M5. Nine scoped semantic checks pass, including detection of deliberately unsupported coarse cases. The benchmark is executable and preserves configurations, warnings and per-trial diagnostics. (Run instructions, results, raw JSON)
| Measured fixture | Numerical result | Meaning |
|---|---|---|
| Sliding block, 0.25 ms | Final velocity error 0.45%; stopping-distance error 0.72% | Agreement with declared Coulomb endpoint references |
| Same sliding case | Contact on 326/1,200 sampled steps; peak normal force 43.58 N versus 9.81 N weight | Endpoint agreement does not validate contact transients or tactile feedback |
| 20 m/s sphere, 2 mm wall, 10 µs | First contact 3.46 ms versus geometric 3.45 ms; 1.679 mm penetration; no full crossing | Compliant numerical response, not a precise hard wall |
| Wall, 1 ms / 5–20 ms | 11 mm penetration at 1 ms; full crossing with zero contacts at 5/20 ms | Faster settings can lose the interaction entirely |
Fine sliding and impact median realtime factors are about 59.7× and 3.5×, respectively. Timing includes Python diagnostics; impact additionally calls mj_forward, so these are separate workloads. Compilation, settling and output are excluded. The coarse 20 ms friction ramp emits a captured solver warning. A separate one-repeat sweep passed with Python socket operations denied and zero attempts; it does not test native C networking or air-gapped installation. No minimum VRAM, GPU scaling, learned policy, deformable material or physical transfer has been demonstrated. (Independent review, offline-check record)
Qualification then proceeds from analytical checks to measured material/sensor coupons, fixed-action intervention replay, closed-loop tests with recoveries and held-out hardware. Compare identical task/controller/observation budgets; keep calibration and final testing separate. Assess false success, unsupported starts, failure recovery, force limits, completion time and material/state slices. SimOpt supplies iterative parameter-distribution calibration precedent; arbitrary wide randomization can include infeasible instances. Ranking agreement, such as SIMPLER's MMRV, complements absolute success calibration but does not certify a new rope policy. (SimOpt, SIMPLER)
| Tentative engineering phase | Reviewable deliverable and proposed acceptance gate |
|---|---|
| 7 days | Network-disabled CPU fixture, reproducible error-versus-speed plots and explicit failure boundaries; no NaNs or silently missing contacts in accepted settings |
| 30 days | Cached single-GPU headless rope/fixture workload, measured VRAM tiers, held-out seeds and slip/regrasp; compare replay, feedback and recovery at matched budgets |
| 60 days | Offline intake-to-skill-pack loop, targeted-curriculum comparison and frozen qualification report; test multi-node scaling if owned capacity exists, physical calibration/holdouts if compatible hardware becomes available |
These are proposed sequencing targets, not delivery estimates or guaranteed success rates. Thresholds must follow task clearance, force tolerance, sensor resolution and pilot variability; any initial percentage is an engineering assumption. A software-only study can establish local numerical behavior, task coverage and simulator performance while marking exported packs simulation-only. Physical competence requires physical evidence.
A pack includes goal/aliases, policy/procedure, robot/end-effector/controller revision, sensors/frames/rates/calibration, preconditions, material/state limits, verifiers, recovery branches and retry budgets. It carries asset/model/dependency hashes, seeds, sampling decisions, qualification splits/results, memory requirements and runtime components. SILO-style packs retain the local simulator; simulator-free packs require separately tested control. New hardware changes compatibility even if the language label is unchanged.
The defensible product advantage is fewer resources to a reliably qualified capability, with difficult cases visible. Existing stacks already offer batching, deformables and training infrastructure. Our contribution must be demonstrated in the calibrated observation-to-practice-to-release loop: which missing decisions to train, where fast approximations remain useful, and when evidence justifies a skill claim.
Research provenance
Three completed native Grok 4.7 consultations ran through Pi and its existing xAI login with hosted X/web search: 23 X calls, 95 fetched post items and 45 web calls. Items can repeat and are not independently verified posts. Checked primary papers/documentation supply the claims above; results are author-reported unless this report labels them as our local measurements. The adapter preserved completed native responses and blocked Pi's redundant normalization request. Research used online search; the proposed product requires offline operation. No cloud GPU spend or robot deployment occurred. (Public method summary)