VOTE
Efficient fine-tuning and parallel action prediction for Vision-Language-Action models. ICANN 2026.
Vision-Language-Action (VLA) Model · Northeastern University, funded by EmbodyX · Jan 2025 – Sept 2025 · ICANN 2026
- Proposed VOTE, an efficient fine-tuning framework for parallel action prediction in VLA models, reducing computational overhead and accelerating inference. Adopted by Cisco and EmbodyX.
- Proposed an ensemble voting strategy for action sampling, improving performance and generalization across diverse tasks.
- Improved OpenVLA’s average success rate by over 20% across four LIBERO task suites, surpassed the state-of-the-art VLA model by 7% average success rate on the SimplerEnv WidowX robot, and accelerated action generation throughput by 39× on the NVIDIA Jetson Orin edge device.
- Pretrained Llama3.2-1B into a vision-language model and fine-tuned it with VOTE, surpassing the 7B OpenVLA’s average success rate by 16% across four LIBERO task suites.