QF3: Fast Flow RL with Filtered Q-Gradients
An off-policy method combines flow matching with filtered critic gradients to train robot policies, with the authors reporting faster humanoid-policy training and hardware transfer.
TL;DR
- QF3 applies critic gradients only to action dimensions that remain near replay actions while training flow policies off-policy.
- The authors report training humanoid locomotion and motion-tracking policies from scratch, with zero-shot transfer to hardware.
- They report a 10x wall-clock speedup over FPO++ in their humanoid experiments; the result is a preprint claim, not an independent replication.
The paper introduces QF3, an online off-policy reinforcement-learning algorithm that combines flow matching with a critic’s action gradient. It filters updates to dimensions that stay near replay actions, which the authors say keeps the critic gradient within a more reliable region. The authors report a 10x wall-clock speedup over FPO++ for humanoid locomotion and motion-tracking policies. [arXiv; QF3 project.] [1] [2]
The paper also reports zero-shot transfer of humanoid policies to hardware and fine-tuning of pretrained manipulation policies. These are results from the authors’ experiments; the collected coverage does not provide an independent replication. [arXiv; QF3 project.] [1] [2]
Why it matters
If the reported results hold beyond the authors’ setup, QF3 could make reinforcement learning a more practical route for improving flow-based robot policies. The evidence remains a new preprint with no independent replication in the collected sources.
Editor's note
Preprint; performance and hardware-transfer claims are author-reported. TODO-review: confirm Korean term for “flow policy.”