Real-World Validation
Evaluate whether the trained policy completes the physical task before selecting a checkpoint for deployment. Use repeated trials and recorded outcomes to compare models.
Prepare the inference setup
- Confirm the Follower configuration and calibration. The Leader is not required for inference.
- Check the camera connections and enabled state. Keep camera positions and viewing angles consistent with recording, and verify the actual feeds.
- Select the checkpoint to evaluate and use a task instruction consistent with the demonstration. Match the inference FPS to the recording FPS, following inference settings.
- Clear the arm's working area and keep the stop control accessible during initial trials.
Follow Inference & Model Conversion to start inference. Where a CUDA host is available, verify the original model first, then compare the converted model on the intended device.
Record a comparable result
Define the task's acceptance criteria before testing. Evaluate task completion, motion path, alignment, and response to object positions that differ from the demonstrations.
| Record | Include |
|---|---|
| Model identification | Dataset, policy, checkpoint, and training step |
| Runtime | Host, execution device, model format, FPS, and action chunk setting |
| Physical conditions | Camera placement, task instruction, object positions, and lighting |
| Task result | Successful attempts / total attempts, plus failure descriptions |
Use the same acceptance criteria and comparable conditions when comparing checkpoints. A lower training loss alone does not show that task performance improved.
Improve the next training run
Use the inference diagnosis table to identify where to investigate.
| Observation | Next check |
|---|---|
| Arm approaches the object but alignment is poor | Check the Wrist and Top viewpoints and their consistency with recording |
| Motion hesitates or repeatedly corrects | Review demonstrations for pauses and repeated corrections |
| Demonstrated positions work but other positions fail | Record additional complete demonstrations covering the required position variation |
| Motion stops before completing the task | Check for truncated episodes or pauses in the recordings |
| Original model works but converted model differs | Recheck the target device, conversion, and runtime settings |
When the data needs improvement, return to physical recording, review the new episodes in dataset management, and start a new training run. Keep the previous dataset and model identification with their results so that the change remains traceable.
Next: Once a checkpoint meets the task's acceptance criteria, continue to Model Deployment.