Skip to main content

Real-World Validation

Evaluate whether the trained policy completes the physical task before selecting a checkpoint for deployment. Use repeated trials and recorded outcomes to compare models.

Prepare the inference setup​

  1. Confirm the Follower configuration and calibration. The Leader is not required for inference.
  2. Check the camera connections and enabled state. Keep camera positions and viewing angles consistent with recording, and verify the actual feeds.
  3. Select the checkpoint to evaluate and use a task instruction consistent with the demonstration. Match the inference FPS to the recording FPS, following inference settings.
  4. Clear the arm's working area and keep the stop control accessible during initial trials.

Follow Inference & Model Conversion to start inference. Where a CUDA host is available, verify the original model first, then compare the converted model on the intended device.

Record a comparable result​

Define the task's acceptance criteria before testing. Evaluate task completion, motion path, alignment, and response to object positions that differ from the demonstrations.

RecordInclude
Model identificationDataset, policy, checkpoint, and training step
RuntimeHost, execution device, model format, FPS, and action chunk setting
Physical conditionsCamera placement, task instruction, object positions, and lighting
Task resultSuccessful attempts / total attempts, plus failure descriptions

Use the same acceptance criteria and comparable conditions when comparing checkpoints. A lower training loss alone does not show that task performance improved.

Improve the next training run​

Use the inference diagnosis table to identify where to investigate.

ObservationNext check
Arm approaches the object but alignment is poorCheck the Wrist and Top viewpoints and their consistency with recording
Motion hesitates or repeatedly correctsReview demonstrations for pauses and repeated corrections
Demonstrated positions work but other positions failRecord additional complete demonstrations covering the required position variation
Motion stops before completing the taskCheck for truncated episodes or pauses in the recordings
Original model works but converted model differsRecheck the target device, conversion, and runtime settings

When the data needs improvement, return to physical recording, review the new episodes in dataset management, and start a new training run. Keep the previous dataset and model identification with their results so that the change remains traceable.

Next: Once a checkpoint meets the task's acceptance criteria, continue to Model Deployment.