VLDB 2026 Research / reviewers in the wild / expert
Diji Yang
dblp:234/1212
· DBLP profile ↗
12ranked-venue papers
4as first author
9since 2021 · last 2026
0009-0005-1591-4846ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Guiding Generative Recommender Systems with Structured Human Priors via Multi-head DecodingabstractOptimizing recommender systems for objectives beyond accuracy, such as diversity, novelty, and personalization, is crucial for long-term user satisfaction. To this end, industrial practitioners have accumulated vast amounts of structured domain knowledge, which we term human priors (e.g., item taxonomies, temporal patterns). This knowledge is typically applied through post-hoc adjustments during ranking or post-ranking. However, this approach remains decoupled from the core model learning, which is particularly undesirable as the industry shifts to end-to-end generative recommendation foundation models. On the other hand, many methods targeting these beyond-accuracy objectives often require architecture-specific modifications and discard these valuable human priors by learning user intent in a fully unsupervised manner. Instead of discarding the human priors accumulated over years of practice, we introduce a backbone-agnostic framework that seamlessly integrates these human priors directly into the end-to-end training of generative recommenders. With lightweight, prior-conditioned adapter heads inspired by efficient LLM decoding strategies, our approach guides the model to disentangle user intent along human-understandable axes (e.g., interaction types, long- vs. short-term interests). We also introduce a hierarchical composition strategy for modeling complex interactions across different prior types. Extensive experiments on three large-scale datasets demonstrate that our method significantly enhances both accuracy and beyond-accuracy objectives. We also show that human priors allow the backbone model to more effectively leverage longer context lengths and larger model sizes. Yunkai Zhang 0002, Diji Yang, Ryan Lin, Ruizhong Qiu, Benyu Zhang, Hanchao Yu, Yinglong Xia, Zhuokai Zhao, Lizhu Zhang, Xiangjun Fan, Zhuoran Yu, Zeyu Zheng 0002 |
WWW | 3 |
| 2025 | GRIT: Teaching MLLMs to Think with ImagesabstractRecent studies have demonstrated the efficacy of using Reinforcement Learning (RL) in building reasoning models that articulate chains of thoughts prior to producing final answers. However, despite ongoing advances that aim at enabling reasoning for vision-language tasks, existing open-source visual reasoning models typically generate reasoning content with pure natural language, lacking explicit integration of visual information. This limits their ability to produce clearly articulated and visually grounded reasoning chains. To this end, we propose Grounded Reasoning with Images and Texts (GRIT), a novel method for training MLLMs to think with images. GRIT introduces a grounded reasoning paradigm, in which models generate reasoning chains that interleave natural language and explicit bounding box coordinates. These coordinates point to regions of the input image that the model consults during its reasoning process. Additionally, GRIT is equipped with a reinforcement learning approach, GRPO-GR, built upon the GRPO algorithm. GRPO-GR employs robust rewards focused on the final answer accuracy and format of the grounded reasoning output, which eliminates the need for data with reasoning chain annotations or explicit bounding box labels. As a result, GRIT achieves exceptional data efficiency, requiring as few as 20 image-question-answer triplets from existing datasets. Comprehensive evaluations demonstrate that GRIT effectively trains MLLMs to produce coherent and visually grounded reasoning chains, showing a successful unification of reasoning and grounding abilities. All code, data, and checkpoints will be released. Xuehai He, Diji Yang, Kaizhi Zheng, Ching-Chen Kuo, Xinze Guan, Xin Wang 0061 |
NeurIPS | 3 |
| 2025 | GenIR: Generative Visual Feedback for Mental Image RetrievalabstractVision-language models (VLMs) have shown strong performance on text-to-image retrieval benchmarks. However, bridging this success to real-world applications remains a challenge. In practice, human search behavior is rarely a one-shot action. Instead, it is often a multi-round process guided by clues in mind. That is, a mental image ranging from vague recollections to vivid mental representations of the target image. Motivated by this gap, we study the task of Mental Image Retrieval (MIR), which targets the realistic yet underexplored setting where users refine their search for a mentally envisioned image through multi-round interactions with an image search engine. Central to successful interactive retrieval is the capability of machines to provide users with clear, actionable feedback; however, existing methods rely on indirect or abstract verbal feedback, which can be ambiguous, misleading, or ineffective for users to refine the query. To overcome this, we propose GenIR, a generative multi-round retrieval paradigm leveraging diffusion-based image generation to explicitly reify the AI system's understanding at each round. These synthetic visual representations provide clear, interpretable feedback, enabling users to refine their queries intuitively and effectively. We further introduce a fully automated pipeline to generate a high-quality multi-round MIR dataset. Experimental results demonstrate that GenIR significantly outperforms existing interactive methods in the MIR scenario. This work establishes a new task with a dataset and an effective generative retrieval method, providing a foundation for future research in this direction Diji Yang, Minghao Liu 0009, Chung-Hsiang Lo, Yi Zhang 0001, James Davis 0001 |
NeurIPS | 1 |
| 2025 | Worse than Zero-shot? A Fact-Checking Dataset for Evaluating the Robustness of RAG Against Misleading RetrievalsabstractRetrieval-augmented generation (RAG) has shown impressive capabilities in mitigating hallucinations in large language models (LLMs). However, LLMs struggle to maintain consistent reasoning when exposed to misleading or conflicting evidence, especially in real-world domains such as politics, where information is polarized or selectively framed. Mainstream RAG benchmarks evaluate models under clean retrieval settings, where systems generate answers from gold-standard documents, or under synthetically perturbed settings, where documents are artificially injected with noise. These assumptions fail to reflect real-world conditions, often leading to an overestimation of RAG system performance. To address this gap, we introduce \textsc{RAGuard}, the first benchmark to evaluate the robustness of RAG systems against \textit{misleading} retrievals. Unlike prior benchmarks that rely on synthetic noise, our fact-checking dataset captures naturally occurring misinformation by constructing its retrieval corpus from Reddit discussions. It categorizes retrieved evidence into three types: \textit{supporting}, \textit{misleading}, and \textit{unrelated}, providing a realistic and challenging testbed for assessing how well RAG systems navigate different types of evidence. Our experiments reveal that, when exposed to potentially misleading retrievals, all tested LLM-powered RAG systems perform worse than their zero-shot baselines (i.e., no retrieval at all), while human annotators consistently perform better, highlighting LLMs' susceptibility to noisy environments. To our knowledge, \textsc{RAGuard} is the first benchmark to systematically assess the robustness of the RAG against misleading evidence.We expect this benchmark to drive future research toward improving RAG systems beyond idealized datasets, making them more reliable for real-world applications. Linda Zeng, Rithwik Gupta, Divij Motwani, Yi Zhang 0001, Diji Yang |
NeurIPS | 5 |
| 2025 | Knowing You Don't Know: Learning When to Continue Search in Multi-round RAG through Self-PracticingabstractRetrieval Augmented Generation (RAG) has shown strong capability in enhancing language models' knowledge and reducing AI generative hallucinations, driving its widespread use. However, complex tasks requiring multi-round retrieval remain challenging, and early attempts tend to be overly optimistic without a good sense of self-skepticism. Current multi-round RAG systems may continue searching even when enough information has already been retrieved, or they may provide incorrect answers without having sufficient information or knowledge. Existing solutions either require large amounts of expensive human-labeled process supervision data or lead to subpar performance. Diji Yang, Linda Zeng, Jinmeng Rao, Yi Zhang 0001 |
SIGIR | 1 |
| 2024 | Tackling Vision Language Tasks through Learning Inner MonologuesabstractVisual language tasks such as Visual Question Answering (VQA) or Visual Entailment (VE) require AI models to comprehend and reason with both visual and textual content. Driven by the power of Large Language Models (LLMs), two prominent methods have emerged: (1) the hybrid integration between LLMs and Vision-Language Models (VLMs), where visual inputs are firstly converted into language descriptions by VLMs, serving as inputs for LLMs to generate final answer(s); (2) visual feature alignment in language space, where visual inputs are encoded as embeddings and projected to LLMs' language space via further supervised fine-tuning. The first approach provides light training costs and interpretability but is hard to be optimized in an end-to-end fashion. The second approach presents decent performance, but feature alignment usually requires large amounts of training data and lacks interpretability. To tackle this dilemma, we propose a novel approach, Inner Monologue Multi-Modal Optimization (IMMO), to solve complex vision language problems by simulating Inner Monologue, a cognitive process in which an individual engages in silent verbal communication with themselves. More specifically, we enable LLMs and VLMs to interact through natural language conversation (i.e., Inner Monologue) and propose to use a two-stage training process to learn how to do Inner Monologue (self-asking questions and answering questions). IMMO is evaluated on two popular tasks and achieves competitive performance with less training data when compared with state-of-the-art models while concurrently keeping the interpretability. The results suggest that by emulating the cognitive phenomenon of internal dialogue, our approach can enhance reasoning and explanation abilities, contributing to the more effective fusion of vision and language models. More importantly, instead of using predefined human-crafted monologues, IMMO learns this process within the deep learning models, broadening its potential applications across various AI challenges beyond vision and language tasks. Diji Yang, Kezhen Chen, Jinmeng Rao, Xiaoyuan Guo, Jie Yang 0002, Yi Zhang 0001 |
AAAI | 1 |
| 2024 | Right this way: Can VLMs Guide Us to See More to Answer Questions?abstractIn question-answering scenarios, humans can assess whether the available information is sufficient and seek additional information if necessary, rather than providing a forced answer. In contrast, Vision Language Models (VLMs) typically generate direct, one-shot responses without evaluating the sufficiency of the information. To investigate this gap, we identify a critical and challenging task in the Visual Question Answering (VQA) scenario: can VLMs indicate how to adjust an image when the visual information is insufficient to answer a question? This capability is especially valuable for assisting visually impaired individuals who often need guidance to capture images correctly. To evaluate this capability of current VLMs, we introduce a human-labeled dataset as a benchmark for this task. Additionally, we present an automated framework that generates synthetic training data by simulating ``where to know'' scenarios. Our empirical results show significant performance improvements in mainstream VLMs when fine-tuned with this synthetic data. This study demonstrates the potential to narrow the gap between information assessment and acquisition in VLMs, bringing their performance closer to humans. Li Liu 0046, Diji Yang, Sijia Zhong, Kalyana Suma Sree Tholeti, Yi Zhang 0001, Leilani H. Gilpin |
NeurIPS | 2 |
| 2024 | IM-RAG: Multi-Round Retrieval-Augmented Generation Through Learning Inner MonologuesabstractAlthough the Retrieval-Augmented Generation (RAG) paradigms can use external knowledge to enhance and ground the outputs of Large Language Models (LLMs) to mitigate generative hallucinations and static knowledge base problems, they still suffer from limited flexibility in adopting Information Retrieval (IR) systems with varying capabilities, constrained interpretability during the multi-round retrieval process, and a lack of end-to-end optimization. To address these challenges, we propose a novel LLM-centric approach, IM-RAG, that integrates IR systems with LLMs to support multi-round RAG through learning Inner Monologues (IM, i.e., the human inner voice that narrates one's thoughts). During the IM process, the LLM serves as the core reasoning model (i.e., Reasoner ) to either propose queries to collect more information via the Retriever or to provide a final answer based on the conversational context. We also introduce a Refiner that improves the outputs from the Retriever, effectively bridging the gap between the Reasoner and IR modules with varying capabilities and fostering multi-round communications. The entire IM process is optimized via Reinforcement Learning (RL) where a Progress Tracker is incorporated to provide mid-step rewards, and the answer prediction is further separately optimized via Supervised Fine-Tuning (SFT). We conduct extensive experiments with the HotPotQA dataset, a popular benchmark for retrieval-based, multi-step question-answering. The results show that our approach achieves state-of-the-art (SOTA) performance while providing high flexibility in integrating IR modules as well as strong interpretability exhibited in the learned inner monologue. Diji Yang, Jinmeng Rao, Kezhen Chen, Xiaoyuan Guo, Jie Yang 0002, Yi Zhang 0001 |
SIGIR | 1 |
| 2022 | CPL: Counterfactual Prompt Learning for Vision and Language ModelsabstractXuehai He, Diji Yang, Weixi Feng, Tsu-Jui Fu, Arjun Akula, Varun Jampani, Pradyumna Narayana, Sugato Basu, William Yang Wang, Xin Wang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Xuehai He, Diji Yang, Weixi Feng, Tsu-Jui Fu, Arjun R. Akula, Varun Jampani, Pradyumna Narayana, Sugato Basu, William Yang Wang, Xin Wang 0061 |
EMNLP | 2 |
| 2020 | Hybrid multiobjective evolutionary algorithm with fast sampling strategy-based global search and route sequence difference-based local search for VRPTW
Diji Yang, Guohui Zhang 0002, Mitsuo Gen |
Expert Syst. Appl. | 2 |
| 2019 | Multiobjective Evolutionary Algorithm based on Fast Elite Sampling Strategy and Difference-based Local Search for VRPTWabstractThis paper addresses the vehicle routing problem with time windows (VRPTW), aiming to reduce the vehicles number and to minimize the time-wasting during the delivery process caused by early arrival. For solving VRPTW, a multiobjective evolutionary algorithm based on fast elite sampling strategy and difference-based local search (MOEAFESS/DLS) is proposed. The special strategy in MOEAFESS/DLS is fast elite sampling strategy (FESS) which consists of two parts. First, using Pareto dominating and dominated relationship-based fitness function (PDDR-FF) to evaluate individuals, which can easily select nondominated individuals and the individuals which have larger domination area. Then, mixing with the sampling strategy of vector evaluated genetic algorithm (VEGA) can achieve fast convergence and sufficient diversity. The evolution according to FESS will be used as a global search strategy for MOEA-FESS/DLS. In addition, a local search based on differences between individuals called difference-based local search (DLS) is used in MOEA-FESS/DLS. In this way, the individuals with poor performance in the population generated by global search are guided to move closer to those who perform better. So it can further enhance the search ability of MOEA-FESS/DLS. Experimental results on Solomon benchmark demonstrate that the proposed MOEA-FESS/DLS is effective, which the performance outperforms NSGA-II, SPEA2 and MOEA/D. Diji Yang, Zhan Qian, Mitsuo Gen |
SMC | 2 |
| 2018 | Hybrid Multiobjective Differential Evolution Based on Positions of Individuals in Multiobjective OptimizationabstractThis paper proposes a hybrid multiobjective differential evolution (HMODE) framework based on positions of individuals in Pareto frontier (PF) to solve multiobjective optimization problem. Firstly, the hybrid multiobjective evolutionary algorithm (HMOEA) in HMODE is designed as the global search strategy to explore the entire solution space. HMOEA uses Pareto dominating and dominated relationship-based fitness function (PDDR-FF) to distinguish the nondominated and dominated individuals. The location information of individual can be determined by its PDDR-FF value. The elitist maintenance strategy based on PDDR-FF and simple selection strategy guarantee the capability of converging to the multiple directions of PF quickly and distributing along the PF uniformly. Secondly, differential evolution (DE) is combined with HMOEA as the local search to enhance the convergence and distribution performances on the elite population derived from HMOEA. Different PDDR-FF values of individuals clearly reflect their dominance relationships and position information in the PF. DE uses the position information of alternative individuals to finely tune the search directions to multiple areas of PF. Numerical comparisons indicate that the efficacy of HMODE outperforms HMOEA in convergence and distribution performances. Diji Yang, Zhan Qian, Heyang Xu, Mitsuo Gen |
SMC | 2 |