Policy Selection: ACT and GR00T N1.5
The "Policy Selection" control on the training page is a dropdown offering two options, ACT and GR00T N1.5, with the interface hint "Choose the training policy algorithm; different policies suit different task types".
The consequences of this selection extend beyond the training stage itself, simultaneously determining three subsequent conditions: the number of demonstrations required, the order of magnitude of the training steps, and the deployment hardware ultimately available. This chapter sets out the criteria for the decision.
1. Assessment by the degree of task variation
The primary criterion for policy selection is the degree of variation the task exhibits at the inference stage.
Fixed motion and stable environment: ACT
Applicable where the workpiece is fixed in a specific fixture position, lighting is supplied by fixed factory fittings, and the arm follows the same trajectory on each cycle. Pick-and-place stations, insertion operations, and positioning prior to screwdriving fall into this category.
What such tasks require is repeatable stability of motion. ACT achieves this by learning the demonstrated trajectory, without semantic-level interpretation of the scene.
Variable object position or appearance: GR00T N1.5
Applicable where parts are loaded manually and their positions are not fixed, where a single station handles components of several sizes, or where ambient lighting varies by time of day.
What such tasks require is the capacity to interpret the scene. GR00T N1.5 is a foundation model pretrained by NVIDIA on large-scale robot data. It possesses general robotic motion capability and can interpret a textual instruction together with the image content before determining its motion, and therefore exhibits greater tolerance of variation in object position and appearance.
The criterion is the degree of variation at the inference stage, not at the recording stage. If the workpiece is deliberately fixed in one position during demonstration but is placed manually in actual operation, GR00T should be selected according to the operational condition, and the positional variation should be incorporated into the demonstrations during recording.
2. Assessment by available demonstration effort
Demonstration recording constitutes a labour cost: the operator must guide the Leader Arm through the task on each occasion, and the process cannot be automated.
| Item | ACT | GR00T N1.5 |
|---|---|---|
| Learning method | Learns the motion from scratch | Fine-tunes on existing general capability |
| Demonstrations required | Higher; dozens are required to achieve stability | Lower; a small number suffices for fine-tuning |
| Recording effort | High | Low |
| Requirement for consistency | High; excessive variation between iterations affects the learning outcome | Comparatively relaxed |
It should be noted that the smaller number of demonstrations required by GR00T does not imply a lower requirement for demonstration quality. As the number of demonstrations decreases, the weight of each individual demonstration within the dataset increases correspondingly: where one of five demonstrations exhibits hesitation, twenty percent of the data is of impaired quality.
Where labour resources are insufficient to support dozens of physical recordings, synthetic data generation in the simulation environment may be considered.
3. Assessment by deployment target
This section describes the criterion with the furthest-reaching consequences among those governing policy selection.
The platform provides two deployment conversion paths, and each path accepts only one policy:
The badge beside each device indicates whether conversion is required: DIRECT denotes direct execution, and CONVERT denotes that model conversion must be completed first. GPU (CUDA) is the sole point at which the two paths meet — both policies execute on it directly — and is the device recommended as the starting point for inference verification.
Apart from that point, no intersection exists between the two paths: the platform provides no path for converting an ACT model to TensorRT, nor for converting a GR00T model to OpenVINO. Any combination without a connecting line is unsupported.
The correspondence between deployment target and policy is therefore as follows:
| Deployment target | Conversion path | Required policy |
|---|---|---|
| Intel platform edge devices (no discrete GPU; cost- and power-sensitive) | OpenVINO | ACT |
| Edge platforms such as NVIDIA Jetson AGX Thor (stringent cycle times, multiple video feeds) | TensorRT | GR00T N1.5 |
| Workstations equipped with an NVIDIA GPU (edge deployment not yet in scope) | None required | Either |
The training and inference device lists differ
The training page lists only GPU (CUDA) and CPU; where the host lacks CUDA, GPU (OpenVINO) and CPU are listed instead. GPU (TensorRT) appears only on the inference page, as it belongs to the deployment stage. Its absence from the training page is expected behaviour.
4. Decision reference table
| Condition | Recommended policy |
|---|---|
| Fixed workpiece position, highly repetitive motion | ACT |
| Workpiece position or appearance differs on each cycle | GR00T N1.5 |
| Site permits only an Intel edge device | ACT (the only viable OpenVINO path) |
| Site uses an NVIDIA Jetson platform with stringent cycle times | GR00T N1.5 |
| Labour resources sufficient for dozens of demonstrations | Either; determine by the preceding conditions |
| Only a small number of demonstrations is feasible | GR00T N1.5 |
| Deployment evaluation stage | ACT; shorter training time and lower hardware requirements make it suitable for an initial pass |
| Comparison of both policies on the same dataset is required | Train both once |
Comparing the two policies on a single dataset
The platform imposes no restriction on this approach. Datasets are in the LeRobot standard format and both policies read the same data; it is necessary only to change the policy and output folder name on the training page and re-execute. For projects where the deployment hardware has not yet been determined, this constitutes the lowest-cost evaluation method, at the expense only of additional training time and disk space.
Summary
ACT and GR00T N1.5 are not different capability tiers of the same approach, but two distinct means of addressing uncertainty: the former emphasises repeatable precision of the motion trajectory, the latter emphasises decision-making following scene interpretation. The degree of task variation determines which is applicable, while the hardware available on site constitutes a hard constraint.
As production hardware specifications are generally determined during the early stages of a project, policy selection is in many cases not an open choice but is determined in reverse by the deployment conditions. Where the deployment hardware has not yet been determined, it is recommended that a complete workflow first be executed using ACT, in order to obtain empirical evidence regarding the complexity of the task.