CVPR 2026 · Collaborative research · Multimodal efficiency
Read the paper · Project repository
The question
Long visual token sequences make multimodal language models expensive to run. Removing tokens early can reduce computation, but it can also discard useful visual evidence before the model has absorbed it.
The idea
Hi-Lo Prune follows a three-stage process:
- Select: use a coarse-to-fine process to identify retained tokens and pruning candidates.
- Fuse: transfer information from the candidates to retained tokens through attention before removal.
- Prune: remove candidates at a designated Transformer layer to reduce subsequent computation.
The method requires no additional training. The paper evaluates it on Qwen2-VL, Qwen2.5-VL, and Qwen3-VL across multiple vision-language benchmarks.
Paper and resources
Full title: Hi-Lo Prune: Look at What You’ll Lose before Pruning with Hierarchical Token Selection
Authors: Zixun Sun, Yubo Dong (董玉博), Hehe Fan, Yi Yang
Publication: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026, pp. 31941–31951.
The CVF paper page provides the paper, supplementary material, and citation. The authors’ repository currently contains a placeholder README; implementation availability should be checked there.