Skip to main content

Policy Selection: ACT and GR00T N1.5

The "Policy Selection" control on the training page is a dropdown offering two options, ACT and GR00T N1.5, with the interface hint "Choose the training policy algorithm; different policies suit different task types".

The consequences of this selection extend beyond the training stage itself, simultaneously determining three subsequent conditions: the number of demonstrations required, the order of magnitude of the training steps, and the deployment hardware ultimately available. This chapter sets out the criteria for the decision.

1. Assessment by the degree of task variation​

The primary criterion for policy selection is the degree of variation the task exhibits at the inference stage.

Fixed motion and stable environment: ACT​

Applicable where the workpiece is fixed in a specific fixture position, lighting is supplied by fixed factory fittings, and the arm follows the same trajectory on each cycle. Pick-and-place stations, insertion operations, and positioning prior to screwdriving fall into this category.

What such tasks require is repeatable stability of motion. ACT achieves this by learning the demonstrated trajectory, without semantic-level interpretation of the scene.

Variable object position or appearance: GR00T N1.5​

Applicable where parts are loaded manually and their positions are not fixed, where a single station handles components of several sizes, or where ambient lighting varies by time of day.

What such tasks require is the capacity to interpret the scene. GR00T N1.5 is a foundation model pretrained by NVIDIA on large-scale robot data. It possesses general robotic motion capability and can interpret a textual instruction together with the image content before determining its motion, and therefore exhibits greater tolerance of variation in object position and appearance.

caution

The criterion is the degree of variation at the inference stage, not at the recording stage. If the workpiece is deliberately fixed in one position during demonstration but is placed manually in actual operation, GR00T should be selected according to the operational condition, and the positional variation should be incorporated into the demonstrations during recording.

2. Assessment by available demonstration effort​

Demonstration recording constitutes a labour cost: the operator must guide the Leader Arm through the task on each occasion, and the process cannot be automated.

ItemACTGR00T N1.5
Learning methodLearns the motion from scratchFine-tunes on existing general capability
Demonstrations requiredHigher; dozens are required to achieve stabilityLower; a small number suffices for fine-tuning
Recording effortHighLow
Requirement for consistencyHigh; excessive variation between iterations affects the learning outcomeComparatively relaxed

It should be noted that the smaller number of demonstrations required by GR00T does not imply a lower requirement for demonstration quality. As the number of demonstrations decreases, the weight of each individual demonstration within the dataset increases correspondingly: where one of five demonstrations exhibits hesitation, twenty percent of the data is of impaired quality.

Where labour resources are insufficient to support dozens of physical recordings, synthetic data generation in the simulation environment may be considered.

3. Assessment by deployment target​

This section describes the criterion with the furthest-reaching consequences among those governing policy selection.

The platform provides two deployment conversion paths, and each path accepts only one policy:

The mapping between policies and deployment devices

The badge beside each device indicates whether conversion is required: DIRECT denotes direct execution, and CONVERT denotes that model conversion must be completed first. GPU (CUDA) is the sole point at which the two paths meet — both policies execute on it directly — and is the device recommended as the starting point for inference verification.

Apart from that point, no intersection exists between the two paths: the platform provides no path for converting an ACT model to TensorRT, nor for converting a GR00T model to OpenVINO. Any combination without a connecting line is unsupported.

The correspondence between deployment target and policy is therefore as follows:

Deployment targetConversion pathRequired policy
Intel platform edge devices (no discrete GPU; cost- and power-sensitive)OpenVINOACT
Edge platforms such as NVIDIA Jetson AGX Thor (stringent cycle times, multiple video feeds)TensorRTGR00T N1.5
Workstations equipped with an NVIDIA GPU (edge deployment not yet in scope)None requiredEither
note

The training and inference device lists differ The training page lists only GPU (CUDA) and CPU; where the host lacks CUDA, GPU (OpenVINO) and CPU are listed instead. GPU (TensorRT) appears only on the inference page, as it belongs to the deployment stage. Its absence from the training page is expected behaviour.

4. Decision reference table​

ConditionRecommended policy
Fixed workpiece position, highly repetitive motionACT
Workpiece position or appearance differs on each cycleGR00T N1.5
Site permits only an Intel edge deviceACT (the only viable OpenVINO path)
Site uses an NVIDIA Jetson platform with stringent cycle timesGR00T N1.5
Labour resources sufficient for dozens of demonstrationsEither; determine by the preceding conditions
Only a small number of demonstrations is feasibleGR00T N1.5
Deployment evaluation stageACT; shorter training time and lower hardware requirements make it suitable for an initial pass
Comparison of both policies on the same dataset is requiredTrain both once

Comparing the two policies on a single dataset​

The platform imposes no restriction on this approach. Datasets are in the LeRobot standard format and both policies read the same data; it is necessary only to change the policy and output folder name on the training page and re-execute. For projects where the deployment hardware has not yet been determined, this constitutes the lowest-cost evaluation method, at the expense only of additional training time and disk space.

Summary​

ACT and GR00T N1.5 are not different capability tiers of the same approach, but two distinct means of addressing uncertainty: the former emphasises repeatable precision of the motion trajectory, the latter emphasises decision-making following scene interpretation. The degree of task variation determines which is applicable, while the hardware available on site constitutes a hard constraint.

As production hardware specifications are generally determined during the early stages of a project, policy selection is in many cases not an open choice but is determined in reverse by the deployment conditions. Where the deployment hardware has not yet been determined, it is recommended that a complete workflow first be executed using ACT, in order to obtain empirical evidence regarding the complexity of the task.


Next: Physical demonstration recording.