Prioritized Rollouts for Efficient World Model-based Vision-Language-Action Policy Optimization
arXiv preprint, 2026
U-GROW uses policy uncertainty to focus world-model rollouts on states with greater potential for improvement, enabling more efficient and effective reinforcement learning for VLA policies.