VLDB 2026 Research / reviewers in the wild / expert
Wang Zhu 0001
dblp:223/4711-1 · also Wang Bill Zhu
· DBLP profile ↗
11ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-6821-4115ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World TasksabstractWe present MEGA-Bench, an evaluation suite that scales multimodal evaluation to over 500 real-world tasks, to address the highly heterogeneous daily use cases of end users.
Our objective is to optimize for a set of high-quality data samples that cover a highly diverse and rich set of multimodal tasks, while enabling cost-effective and accurate model evaluation.
In particular, we collected 505 realistic tasks encompassing over 8,000 samples from 16 expert annotators to extensively cover the multimodal task space. Instead of unifying these problems into standard multi-choice questions (like MMMU, MM-Bench, and MMT-Bench), we embrace a wide range of output formats like numbers, phrases, code, \LaTeX, coordinates, JSON, free-form, etc. To accommodate these formats, we developed over 40 metrics to evaluate these tasks.
Unlike existing benchmarks, MEGA-Bench offers a fine-grained capability report across multiple dimensions (e.g., application, input type, output format, skill), allowing users to interact with and visualize model capabilities in depth. We evaluate a wide variety of frontier vision-language models on MEGA-Bench to understand their capabilities across these dimensions. Tianhao Liang, Sherman Siu, Zhengqing Wang, Kai Wang 0068, Yubo Wang 0019, Yuansheng Ni, Ziyan Jiang, Wang Zhu 0001, Bohan Lyu 0001, Dongfu Jiang, Hexiang Hu, Xiang Yue, Wenhu Chen |
ICLR | 9 |
| 2025 | TLDR: Token-Level Detective Reward Model for Large Vision Language ModelsabstractAlthough reward models have been successful in improving multimodal large language models, the reward models themselves remain brutal and contain minimal information. Notably, existing reward models only mimic human annotations by assigning only one feedback to any text, no matter how long the text is. In the realm of multimodal language models, where models are required to process both images and texts, a naive reward model may learn implicit biases toward texts and become less grounded in images. In this paper, we propose a **T**oken-**L**evel **D**etective **R**eward Model (**TLDR**) to provide fine-grained annotations to each text token. We first introduce a perturbation-based method to generate synthetic hard negatives and their token-level labels to train TLDR models. Then we show the rich usefulness of TLDR models both in assisting off-the-shelf models to self-correct their generations, and in serving as a hallucination evaluation tool. We show that TLDR automatically trains a token-level likelihood optimization, and can improve the base model's performance significantly. Finally, we show that TLDR models can significantly speed up human annotation by 3 times to acquire a broader range of high-quality vision language data. Deqing Fu, Tong Xiao 0003, Wang Zhu 0001, Pengchuan Zhang, Guan Pang, Robin Jia, Lawrence Chen 0002 |
ICLR | 4 |
| 2025 | Language Models Can Infer Action Semantics for Symbolic Planners from Environment FeedbackabstractWang Bill Zhu, Ishika Singh, Robin Jia, Jesse Thomason. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Wang Zhu 0001, Ishika Singh, Robin Jia, Jesse Thomason |
NAACL (Long Papers) | 1 |
| 2025 | To Think or Not To Think: A Study of Thinking in Rule-Based Visual Reinforcement Fine-TuningabstractThis paper investigates the role of explicit thinking process in rule-based reinforcement fine-tuning (RFT) for multi-modal large language models (MLLMs). We first extend \textit{Thinking-RFT} to image classification task, using verifiable rewards for fine-tuning~(FT). Experiments show {Thinking-RFT} significantly outperforms supervised FT and yields a cross-dataset generalization effect. We then rethink and question whether explicit thinking in RFT is always necessary and beneficial. Challenging the convention that explicit thinking is crucial for the success of RFT, we introduce \textit{No-Thinking-RFT}, exploring RFT without thinking by introducing a simple equality accuracy reward. We evaluate No-Thinking-RFT on six diverse tasks across different model sizes and types. Experiment results reveal four key findings: \textbf{(1).} Visual perception tasks do not require thinking during RFT, as No-Thinking-RFT consistently outperforms or matches Thinking-RFT across model sizes and types. \textbf{(2).} Models with limited capabilities struggle to generate high-quality CoT for RFT, making Thinking-RFT less effective than No-Thinking-RFT. \textbf{(3).} There are inconsistencies between the answers in the thinking tags and answer tags for some responses of Thinking-RFT, which show lower average accuracy than the overall accuracy. \textbf{(4).} The performance gain of No-Thinking-RFT mainly stems from improved learning during no thinking FT and the avoidance of inference overthinking, as evidenced by the partial gains from appending empty thinking tags at inference time of Thinking-RFT. We hypothesize that explicit thinking before verifiable answers may hinder reward convergence and reduce performance in certain scenarios. To test this, we propose \textit{Think-After-Answer}, which places thinking after the answer to mitigate this effect for experimental verification. Lastly, we conduct a pilot study to explore whether MLLMs can learn when to think during RFT, introducing an \textit{Adaptive-Thinking} method. Experiments show that model converges to either thinking or not depending on model capability, achieving comparable or better performance than both Thinking and No-Thinking-RFT. Our findings suggest MLLMs can adaptively decide to think or not based on their capabilities and task complexity, offering insights into the thinking process in RFT. Jike Zhong, Shitian Zhao, Yuxiang Lai, Haoquan Zhang, Wang Zhu 0001, Kaipeng Zhang |
NeurIPS | 6 |
| 2025 | VisualLens: Personalization through Task-Agnostic Visual HistoryabstractExisting recommendation systems either rely on user interaction logs, such as online shopping history for shopping recommendations, or focus on text signals.
However, item-based histories are not always accessible and generalizable for multimodal recommendation.
We hypothesize that a user's visual history --- comprising images from daily life --- can offer rich, task-agnostic insights into their interests and preferences, and thus be leveraged for effective personalization.
To this end, we propose VisualLens, a novel framework that leverages multimodal large language models (MLLMs) to enable personalization using task-agnostic visual history.
VisualLens extracts, filters, and refines a spectrum user profile from the visual history to support personalized recommendation.
We created two new benchmarks, Google-Review-V and Yelp-V, with task-agnostic visual histories, and show that VisualLens improves over state-of-the-art item-based multimodal recommendations by 5-10\% on Hit@3, and outperforms GPT-4o by 2-5\%.
Further analysis shows that VisualLens is robust across varying history lengths and excels at adapting to both longer histories and unseen content categories. Wang Zhu 0001, Deqing Fu, Kai Sun 0006, Zhaojiang Lin, Seungwhan Moon, Kanika Narang, Mustafa Canim, Xin Dong 0001 |
NeurIPS | 1 |
| 2024 | Efficient End-to-End Visual Document Understanding with Rationale DistillationabstractWang Zhu, Alekh Agarwal, Mandar Joshi, Robin Jia, Jesse Thomason, Kristina Toutanova. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Wang Zhu 0001, Alekh Agarwal, Mandar Joshi, Robin Jia, Jesse Thomason, Kristina Toutanova |
NAACL-HLT | 1 |
| 2023 | Iterative Vision-and-Language NavigationabstractWe present Iterative Vision-and-Language Navigation (IVLN), a paradigm for evaluating language-guided agents navigating in a persistent environment over time. Existing Vision-and-Language Navigation (VLN) benchmarks erase the agent's memory at the beginning of every episode, testing the ability to perform cold-start navigation with no prior information. However, deployed robots occupy the same environment for long periods of time. The IVLN paradigm addresses this disparity by training and evaluating VLN agents that maintain memory across tours of scenes that consist of up to 100 ordered instruction-following Room-to-Room (R2R) episodes, each defined by an individual language instruction and a target path. We present discrete and continuous Iterative Room-to-Room (IR2R) benchmarks comprising about 400 tours each in 80 indoor scenes. We find that extending the implicit memory of high-performing transformer VLN agents is not sufficient for IVLN, but agents that build maps can benefit from environment persistence, motivating a renewed focus on map-building agents in VLN. Jacob Krantz, Shurjo Banerjee, Wang Zhu 0001, Jason J. Corso, Stefan Lee, Jesse Thomason |
CVPR | 3 |
| 2023 | Chain-of-Questions Training with Latent Answers for Robust Multistep Question AnsweringabstractWe propose Chain-of-Questions, a framework that trains a model to robustly answer multistep questions by generating and answering sub-questions.We obtain supervision for subquestions from human-annotated question decomposition meaning representation (QDMR), but QDMR does not include annotated answers to sub-questions.To overcome this technical challenge, we treat sub-answers as latent variables and infer them with a novel dynamic mixture of Hard-EM and MAPO.Chain-of-Questions is effective and robust, greatly outperforming strong neuro-symbolic methods by 9.0 F1 on a DROP contrast set and GPT-3.5 by 24.3 F1 on a HOTPOTQA adversarial set. Wang Zhu 0001, Jesse Thomason, Robin Jia |
EMNLP | 1 |
| 2022 | Navigating Memory Construction by Global Pseudo-Task Simulation for Continual LearningabstractContinual learning faces a crucial challenge of catastrophic forgetting. To address this challenge, experience replay (ER) that maintains a tiny subset of samples from previous tasks has been commonly used. Existing ER works usually focus on refining the learning objective for each task with a static memory construction policy. In this paper, we formulate the dynamic memory construction in ER as a combinatorial optimization problem, which aims at directly minimizing the global loss across all experienced tasks. We first apply three tactics to solve the problem in the offline setting as a starting point. To provide an approximate solution to this problem under the online continual learning setting, we further propose the Global Pseudo-task Simulation (GPS), which mimics future catastrophic forgetting of the current task by permutation. Our empirical results and analyses suggest that the GPS consistently improves accuracy across four commonly used vision benchmarks. We have also shown that our GPS can serve as the unified framework for integrating various memory construction policies in existing ER works. Yejia Liu, Wang Zhu 0001, Shaolei Ren |
NeurIPS | 2 |
| 2020 | BabyWalk: Going Farther in Vision-and-Language Navigation by Taking Baby StepsabstractLearning to follow instructions is of fundamental importance to autonomous agents for vision-and-language navigation (VLN).In this paper, we study how an agent can navigate long paths when learning from a corpus that consists of shorter ones.We show that existing state-of-the-art agents do not generalize well.To this end, we propose BabyWalk, a new VLN agent that is learned to navigate by decomposing long instructions into shorter ones (BabySteps) and completing them sequentially.A special design memory buffer is used by the agent to turn its past experiences into contexts for future steps.The learning process is composed of two phases.In the first phase, the agent uses imitation learning from demonstration to accomplish BabySteps.In the second phase, the agent uses curriculum-based reinforcement learning to maximize rewards on navigation tasks with increasingly longer instructions.We create two new benchmark datasets (of long navigation tasks) and use them in conjunction with existing ones to examine BabyWalk's generalization ability.Empirical results show that BabyWalk achieves state-of-the-art results on several metrics, in particular, is able to follow long instructions better.The codes and the datasets are released on our project page https://github.com/Sha-Lab/babywalk. Wang Zhu 0001, Hexiang Hu, Zhiwei Deng, Vihan Jain, Eugene Ie, Fei Sha |
ACL | 1 |
| 2018 | Toward Interpretable Deep Reinforcement Learning with Linear Model U-Trees
Guiliang Liu, Oliver Schulte, Wang Zhu 0001, Qingcan Li |
ECML/PKDD (2) | 3 |