VLDB 2026 Research / reviewers in the wild / expert
Zhaoye Fei
dblp:304/3217
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Planning, search and constraint satisfaction · 18% Language models and text generation · 14% Vision and language · 14% |
Topics — the 17 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › agent planning
embodied planning |
1.7 | 2 | 2025 | World-aware Planning Narratives Enhance Large Vision-Language Model Planner · NeurIPS 2025 World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning · ACL (1) 2025 |
Natural language and speech › Speech recognition and synthesis › speech coding
low-bit-rate speech coding |
1.0 | 1 | 2026 | XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs · ACL (1) 2026 |
Natural language and speech › Speech recognition and synthesis
speech coding |
1.0 | 1 | 2026 | XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs · ACL (1) 2026 |
Robotics › Robot navigation and mapping
embodied instruction following |
0.9 | 1 | 2025 | World-aware Planning Narratives Enhance Large Vision-Language Model Planner · NeurIPS 2025 |
Robotics › Motion planning and robot control › robot learning › manipulation learning
language-conditioned manipulation |
0.9 | 1 | 2025 | VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks · ICCV 2025 |
Computer vision › Vision and language
multimodal reasoning |
0.9 | 1 | 2025 | VisuoThink: Empowering LVLM Reasoning with Multimodal Tree Search · ACL (1) 2025 |
Machine learning › Trustworthy machine learning › robustness
overfitting mitigation |
0.9 | 1 | 2025 | How to Mitigate Overfitting in Weak-to-strong Generalization? · ACL (1) 2025 |
Machine learning › Reinforcement learning
preference learning |
0.9 | 1 | 2025 | World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning · ACL (1) 2025 |
Natural language and speech › Language models and text generation › alignment
super-alignment |
0.9 | 1 | 2025 | How to Mitigate Overfitting in Weak-to-strong Generalization? · ACL (1) 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
tree search |
0.9 | 1 | 2025 | VisuoThink: Empowering LVLM Reasoning with Multimodal Tree Search · ACL (1) 2025 |
Robotics › Robot manipulation › embodied foundation models
vision-language-action model |
0.9 | 1 | 2025 | VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks · ICCV 2025 |
Computer vision › Vision and language › multimodal reasoning
vision-language model reasoning |
0.9 | 1 | 2025 | VisuoThink: Empowering LVLM Reasoning with Multimodal Tree Search · ACL (1) 2025 |
Machine learning › Learning paradigms
weakly supervised learning |
0.9 | 1 | 2025 | How to Mitigate Overfitting in Weak-to-strong Generalization? · ACL (1) 2025 |
Natural language and speech › Language models and text generation › alignment › scalable oversight
weak-to-strong generalization |
0.9 | 1 | 2025 | How to Mitigate Overfitting in Weak-to-strong Generalization? · ACL (1) 2025 |
Machine learning › Deep learning architectures and training
mixture of experts |
0.8 | 1 | 2024 | Turn Waste into Worth: Rectifying Top-k Router of MoE · EMNLP 2024 |
Natural language and speech › Language models and text generation › large language model reasoning › multi-step reasoning
long-horizon reasoning |
0.3 | 1 | 2025 | VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks · ICCV 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.3 | 1 | 2025 | World-aware Planning Narratives Enhance Large Vision-Language Model Planner · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
tree search · 1.7tokenizer design · 1.0semantic-acoustic disentanglement · 1.0world modeling · 0.9vision-language-action model · 0.9vision-language model · 0.9two-stage framework · 0.9test-time scaling · 0.9label filtering · 0.9dual preference optimization · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech CodecsabstractYitian Gong, Luozhijie Jin, Kuangwei Chen, Dong Zhang, Ruifan Deng, Xiaogui Yang, Xin Zhang, Zhaoye Fei, Qinyuan Cheng, Shimin Li, Xipeng Qiu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yitian Gong, Luozhijie Jin, Kuangwei Chen, Ruifan Deng, Xiaogui Yang, Zhaoye Fei, Qinyuan Cheng, Xipeng Qiu |
ACL (1) | 8 |
| 2025 | How to Mitigate Overfitting in Weak-to-strong Generalization?abstractAligning powerful AI models on tasks that surpass human evaluation capabilities is the central problem of superalignment. To address this problem, weak-to-strong generalization aims to elicit the capabilities of strong models through weak supervisors and ensure that the behavior of strong models aligns with the intentions of weak supervisors without unsafe behaviors such as deception. Although weak-to-strong generalization exhibiting certain generalization capabilities, strong models exhibit significant overfitting in weak-to-strong generalization: Due to the strong fit ability of strong models, erroneous labels from weak supervisors may lead to overfitting in strong models. In addition, simply filtering out incorrect labels may lead to a degeneration in question quality, resulting in a weak generalization ability of strong models on hard questions. To mitigate overfitting in weak-to-strong generalization, we propose a two-stage framework that simultaneously improves the quality of supervision signals and the quality of input questions. Experimental results in three series of large language models and two mathematical benchmarks demonstrate that our framework significantly improves PGR (Performance Gap Recovered) compared to naive weak-to-strong generalization, even achieving up to 100% PGR on some models. Junhao Shi, Qinyuan Cheng, Zhaoye Fei, Yining Zheng, Qipeng Guo, Xipeng Qiu |
ACL (1) | 3 |
| 2025 | World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task PlanningabstractRecent advances in large vision-language models (LVLMs) have shown promise for embodied task planning, yet they struggle with fundamental challenges like dependency constraints and efficiency. Existing approaches either solely optimize action selection or directly leverage pre-trained models as world models during inference, overlooking the benefits of learning to model the world as a way to enhance planning capabilities. We propose Dual Preference Optimization (D^2PO), a new learning framework that jointly optimizes state prediction and action selection through preference learning, enabling LVLMs to understand environment dynamics for better planning. To automatically collect trajectories and stepwise preference data without human annotation, we introduce a tree search mechanism for extensive exploration via trial-and-error. Extensive experiments on VoTa-Bench demonstrate that our D^2PO-based method significantly outperforms existing methods and GPT-4o when applied to Qwen2-VL (7B), LLaVA-1.6 (7B), and LLaMA-3.2 (11B), achieving superior task success rates with more efficient execution paths. Siyin Wang, Zhaoye Fei, Qinyuan Cheng, Shiduo Zhang, Panpan Cai, Jinlan Fu, Xipeng Qiu |
ACL (1) | 2 |
| 2025 | VisuoThink: Empowering LVLM Reasoning with Multimodal Tree SearchabstractRecent advancements in Large Vision-Language Models have showcased remarkable capabilities. However, they often falter when confronted with complex reasoning tasks that humans typically address through visual aids and deliberate, step-by-step thinking. While existing methods have explored text-based slow thinking or rudimentary visual assistance, they fall short of capturing the intricate, interleaved nature of human visual-verbal reasoning processes. To overcome these limitations and inspired by the mechanisms of slow thinking in human cognition, we introduce VisuoThink, a novel framework that seamlessly integrates visuospatial and linguistic domains. VisuoThink facilitates multimodal slow thinking by enabling progressive visual-textual reasoning and incorporates test-time scaling through look-ahead tree search. Extensive experiments demonstrate that VisuoThink significantly enhances reasoning capabilities via inference-time scaling, even without fine-tuning, achieving state-of-the-art performance in tasks involving geometry and spatial reasoning. Yikun Wang 0001, Siyin Wang, Qinyuan Cheng, Zhaoye Fei, Liang Ding 0006, Qipeng Guo, Dacheng Tao, Xipeng Qiu |
ACL (1) | 4 |
| 2025 | VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning TasksabstractGeneral-purposed embodied agents are designed to understand the users' natural instructions or intentions and act precisely to complete universal tasks. Recently, methods based on foundation models especially Vision-Language-Action models (VLAs) have shown a substantial potential to solve language-conditioned manipulation (LCM) tasks well. However, existing benchmarks do not adequately meet the needs of VLAs and relative algorithms. To better define such general-purpose tasks in the context of LLMs and advance the research in VLAs, we present VLABench, an open-source benchmark for evaluating universal LCM task learning. VLABench provides 100 carefully designed categories of tasks, with strong randomization in each category of task and a total of 2000+ objects. VLABench stands out from previous benchmarks in four key aspects: 1) tasks requiring world knowledge and common sense transfer, 2) natural language instructions with implicit human intentions rather than templates, 3) long-horizon tasks demanding multi-step reasoning, and 4) evaluation of both action policies and language model capabilities. The benchmark assesses multiple competencies including understanding of mesh\&texture, spatial relationship, semantic instruction, physical laws, knowledge transfer and reasoning, etc. To support the downstream finetuning, we provide high-quality training data collected via an automated framework incorporating heuristic skills and prior information. The experimental results indicate that both the current state-of-the-art pretrained VLAs and the workflow based on VLMs face challenges in our tasks. Shiduo Zhang, Peiju Liu, Qinghui Gao, Zhaoye Fei, Zhangyue Yin, Zuxuan Wu, Yu-Gang Jiang 0001, Xipeng Qiu |
ICCV | 7 |
| 2025 | World-aware Planning Narratives Enhance Large Vision-Language Model PlannerabstractLarge Vision-Language Models (LVLMs) show promise for embodied planning tasks but struggle with complex scenarios involving unfamiliar environments and multi-step goals.
Current approaches rely on environment-agnostic imitation learning that disconnects instructions from environmental contexts, causing models to struggle with context-sensitive instructions and rely on supplementary cues rather than visual reasoning during long-horizon interactions.
In this work, we propose World-Aware Planning Narrative Enhancement (WAP), a framework that infuses LVLMs with comprehensive environmental understanding through four cognitive capabilities (visual appearance modeling, spatial reasoning, functional abstraction, and syntactic grounding) while developing and evaluating models using only raw visual observations through curriculum learning.
Evaluations on the EB-ALFRED benchmark demonstrate substantial improvements, with Qwen2.5-VL achieving a 60.7 absolute improvement in task success rates—particularly in commonsense reasoning (+60.0) and long-horizon planning (+70.0). Notably, our enhanced open-source models outperform proprietary systems like GPT-4o and Claude-3.5-Sonnet by a large margin. Junhao Shi, Zhaoye Fei, Siyin Wang, Qipeng Guo, Jingjing Gong, Xipeng Qiu |
NeurIPS | 2 |
| 2024 | Turn Waste into Worth: Rectifying Top-k Router of MoEabstractZhiyuan Zeng, Qipeng Guo, Zhaoye Fei, Zhangyue Yin, Yunhua Zhou, Linyang Li, Tianxiang Sun, Hang Yan, Dahua Lin, Xipeng Qiu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Zhiyuan Zeng 0004, Qipeng Guo, Zhaoye Fei, Zhangyue Yin, Yunhua Zhou, Linyang Li, Tianxiang Sun, Hang Yan 0001, Dahua Lin, Xipeng Qiu |
EMNLP | 3 |
| 2022 | Coarse-to-Fine: Hierarchical Multi-task Learning for Natural Language UnderstandingabstractGeneralized text representations are the foundation of many natural language understanding tasks. To fully utilize the different corpus, it is inevitable that models need to understand the relevance among them. However, many methods ignore the relevance and adopt a single-channel model (a coarse paradigm) directly for all tasks, which lacks enough rationality and interpretation. In addition, some existing works learn downstream tasks by stitches skill block (a fine paradigm), which might cause irrational results due to its redundancy and noise. In this work, we first analyze the task correlation through three different perspectives, , data property, manual design, and model-based relevance, based on which the similar tasks are grouped together. Then, we propose a hierarchical framework with a coarse-to-fine paradigm, with the bottom level shared to all the tasks, the mid-level divided to different groups, and the top-level assigned to each of the tasks. This allows our model to learn basic language properties from all tasks, boost performance on relevant tasks, and reduce the negative impact from irrelevant tasks. Our experiments on 13 benchmark datasets across five natural language understanding tasks demonstrate the superiority of our method. Zhaoye Fei, Yongkang Wu, Xinyu Zhang 0019, Yutao Zhu 0001, Zheng Liu 0011, Jiawen Wu 0002, Dejiang Kong, Ruofei Lai, Zhao Cao, Zhicheng Dou, Xipeng Qiu |
COLING | 1 |