VLDB 2026 Research / reviewers in the wild / expert
Haochun Wang
dblp:329/5284
· DBLP profile ↗
14ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0003-2908-9750ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Collaborative Chain-of-Agents for Parametric-Retrieved Knowledge SynergyabstractRetrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs), especially for knowledge-intensive tasks.Despite its advantages, current RAG methods often struggle to fully exploit knowledge during generation.In particular, the synergy between the model's internal parametric knowledge and external retrieved knowledge remains limited.Retrieved contents may sometimes mislead generation, while certain generated content can guide the model toward more accurate outputs.In this work, we propose Collaborative Chainof-Agents, a framework designed to enhance explicitly synergy over both parametric and retrieved knowledge.Specifically, we first introduce CoCoA-zero, a multi-agent RAG framework that first performs conditional knowledge induction and then reasons answers.Building on this, we develop CoCoA, a long-chain training strategy that synthesizes extended multiagent reasoning trajectories from CoCoA-zero to fine-tune the LLM.This strategy enhances the model's capability to explicitly integrate and jointly leverage parametric and retrieved knowledge.Experimental results demonstrate the superiority of CoCoA in open-domain QA and multi-hop QA.Code is public 1 . Sendong Zhao, Haochun Wang, Lizhe Zhang, Bing Qin 0001 |
ACL (1) | 4 |
| 2026 | When Correct Beliefs Collapse: Epistemic Resilience of LLMs under Clinical PressureabstractBoyu Xiao, Xiuqi Tian, Xuwen Song, Haochun Wang, Guanchun Song, Sendong Zhao, Bing Qin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Boyu Xiao, Xiuqi Tian, Xuwen Song, Haochun Wang, Guanchun Song, Sendong Zhao, Bing Qin 0001 |
ACL (1) | 4 |
| 2025 | GainRAG: Preference Alignment in Retrieval-Augmented Generation through Gain Signal SynthesisabstractThe Retrieval-Augmented Generation (RAG) framework introduces a retrieval module to dynamically inject retrieved information into the input context of large language models (LLMs), and has demonstrated significant success in various NLP tasks.However, the current study points out that there is a preference gap between retrievers and LLMs in the RAG framework, which limit the further improvement of system performance.Some highly relevant passages may interfere with LLM reasoning because they contain complex or contradictory information; while some indirectly related or even inaccurate content may help LLM generate more accurate answers by providing suggestive information or logical clues.To solve this, we propose GainRAG, a novel approach that aligns the retriever's and LLM's preferences by defining a new metric, "gain", which measure how well an input passage contributes to correct outputs.Specifically, we propose a method to estimate these gain signals and train a middleware that aligns the preferences of the retriever and the LLM using only limited data.In addition, we introduce a pseudo-passage strategy to mitigate degradation.The experimental results on 6 datasets verify the effectiveness of GainRAG 1 . Sendong Zhao, Haochun Wang, Bing Qin 0001 |
ACL (1) | 4 |
| 2025 | Beyond Frameworks: Unpacking Collaboration Strategies in Multi-Agent SystemsabstractMulti-agent collaboration has emerged as a pivotal paradigm for addressing complex, distributed tasks in large language model (LLM)-driven applications. While prior research has focused on high-level architectural frameworks, the granular mechanisms governing agents—critical to performance and scalability—remain underexplored. This study systematically investigates four dimensions of collaboration strategies: (1) agent governance, (2) participation control, (3) interaction dynamics, and (4) dialogue history management. Through rigorous experimentation under two context-dependent scenarios—Distributed Evidence Integration (DEI) and Structured Evidence Synthesis (SES)—we quantify the impact of these strategies on both task accuracy and computational efficiency. Our findings reveal that centralized governance, instructor-led participation, ordered interaction patterns, and instructor-curated context summarization collectively optimize the trade-off between decision quality and resource utilization with the support of the proposed Token-Accuracy Ratio (TAR). This work establishes a foundation for designing adaptive, scalable multi-agent systems, shifting the focus from structural novelty to strategic interaction mechanics. Haochun Wang, Sendong Zhao, Zewen Qiang, Bing Qin 0001, Ting Liu 0001 |
ACL (1) | 1 |
| 2025 | LLMs May Perform MCQA by Selecting the Least Incorrect OptionabstractIn the field of NLP, Large Language Models (LLMs) have markedly enhanced performance across a variety of tasks. However, the comprehensive evaluation of LLMs remains an inevitable challenge for the community. Recently, the adoption of Multiple Choice Question Answering (MCQA) as a benchmark for assessing LLMs has gained considerable traction. However, concerns regarding the robustness of this evaluative method persist. Building upon previous discussions on the issue of variability, we reveal an additional dimension of concern: LLMs may perform MCQA by selecting the least incorrect option rather than distinctly correct. This observation suggests that LLMs might regard multiple options as correct, which could undermine the reliability of MCQA as a metric for evaluating LLMs. To address this challenge, we introduce an enhanced dataset augmentation method for MCQA, termed MCQA+, to provide a more accurate reflection of the performance, thereby highlighting the necessity for more sophisticated evaluation mechanisms in the assessment of LLM capabilities. Haochun Wang, Sendong Zhao, Zewen Qiang, Nuwa Xi, Bing Qin 0001, Ting Liu 0001 |
COLING | 1 |
| 2025 | MolFusion: Multimodal Fusion Learning for Molecular Representations via Multi-granularity ViewsabstractArtificial intelligence advances drug design by predicting drug properties through the encoding of drug molecules. Since different molecular representations contain complementary information, a large amount of research is currently dedicated to the fusion of these representations. However, current multimodal molecular studies rely more on single-granularity methods, resulting in the loss of atomic-level information between representations. Inspired by the success of multi-granularity approaches, we introduce MolFusion, an innovative method for multi-granularity molecular representation fusion. MolFusion comprises two core components: MolSim for molecular-level fusion and AtomAlign for atomic-level fusion. Comprehensive experiments show that our method outperforms all powerful baselines in terms of average performance across both classification and regression tasks, with a particular enhancement in regression tasks. Our code and dataset will be available on https://github.com/Mengqi97/MolFusion. Muzhen Cai, Sendong Zhao, Haochun Wang, Haoqiang Guo, Yanrui Du, Zewen Qiang, Bing Qin 0001, Ting Liu 0001 |
IJCNN | 3 |
| 2025 | Knowledge-tuning Large Language Models with Structured Medical Knowledge Bases for Trustworthy Response Generation in ChineseabstractLarge Language Models (LLMs) have demonstrated remarkable success in diverse natural language processing (NLP) tasks in general domains. However, LLMs sometimes generate responses with the hallucination about medical facts due to limited domain knowledge. Such shortcomings pose potential risks in the utilization of LLMs within medical contexts. To address this challenge, we propose knowledge-tuning, which leverages structured medical knowledge bases for the LLMs to grasp domain knowledge efficiently and facilitate trustworthy response generation. We also release cMedKnowQA, a Chinese medical knowledge question-answering dataset constructed from medical knowledge bases to assess the medical knowledge proficiency of LLMs. Experimental results show that the LLMs which are knowledge-tuned with cMedKnowQA can exhibit higher levels of accuracy in response generation compared with vanilla instruction-tuning and offer a new trustworthy way for the domain adaptation of LLMs. We release our code and data at https://github.com/SCIR-HI/Huatuo-Llama-Med-Chinese . Haochun Wang, Sendong Zhao, Zewen Qiang, Zijian Li 0020, Chi Liu 0003, Nuwa Xi, Yanrui Du, Bing Qin 0001, Ting Liu 0001 |
ACM Trans. Knowl. Discov. Data | 1 |
| 2024 | From Artificially Real to Real: Leveraging Pseudo Data from Large Language Models for Low-Resource Molecule DiscoveryabstractMolecule discovery serves as a cornerstone in numerous scientific domains, fueling the development of new materials and innovative drug designs. Recent developments of in-silico molecule discovery have highlighted the promising results of cross-modal techniques, which bridge molecular structures with their descriptive annotations. However, these cross-modal methods frequently encounter the issue of data scarcity, hampering their performance and application. In this paper, we address the low-resource challenge by utilizing artificially-real data generated by Large Language Models (LLMs). We first introduce a retrieval-based prompting strategy to construct high-quality pseudo data, then explore the optimal method to effectively leverage this pseudo data. Experiments show that using pseudo data for domain adaptation outperforms all existing methods, while also requiring a smaller model scale, reduced data size and lower training cost, highlighting its efficiency. Furthermore, our method shows a sustained improvement as the volume of pseudo data increases, revealing the great potential of pseudo data in advancing low-resource cross-modal molecule discovery. Yuhan Chen 0002, Nuwa Xi, Yanrui Du, Haochun Wang, Sendong Zhao, Bing Qin 0001 |
AAAI | 4 |
| 2024 | MolTailor: Tailoring Chemical Molecular Representation to Specific Tasks via Text PromptsabstractDeep learning is now widely used in drug discovery, providing significant acceleration and cost reduction. As the most fundamental building block, molecular representation is essential for predicting molecular properties to enable various downstream applications. Most existing methods attempt to incorporate more information to learn better representations. However, not all features are equally important for a specific task. Ignoring this would potentially compromise the training efficiency and predictive accuracy. To address this issue, we propose a novel approach, which treats language models as an agent and molecular pretraining models as a knowledge base. The agent accentuates task-relevant features in the molecular representation by understanding the natural language description of the task, just as a tailor customizes clothes for clients. Thus, we call this approach MolTailor. Evaluations demonstrate MolTailor's superior performance over baselines, validating the efficacy of enhancing relevance for molecular representation learning. This illustrates the potential of language model guided optimization to better exploit and unleash the capabilities of existing powerful molecular representation methods. Our code and appendix are available at https://github.com/SCIR-HI/MolTailor. Haoqiang Guo, Sendong Zhao, Haochun Wang, Yanrui Du, Bing Qin 0001 |
AAAI | 3 |
| 2024 | Manifold-Based Verbalizer Space Re-embedding for Tuning-Free Prompt-Based ClassificationabstractPrompt-based classification adapts tasks to a cloze question format utilizing the [MASK] token and the filled tokens are then mapped to labels through pre-defined verbalizers. Recent studies have explored the use of verbalizer embeddings to reduce labor in this process. However, all existing studies require a tuning process for either the pre-trained models or additional trainable embeddings. Meanwhile, the distance between high-dimensional verbalizer embeddings should not be measured by Euclidean distance due to the potential for non-linear manifolds in the representation space. In this study, we propose a tuning-free manifold-based space re-embedding method called Locally Linear Embedding with Intra-class Neighborhood Constraint (LLE-INC) for verbalizer embeddings, which preserves local properties within the same class as guidance for classification. Experimental results indicate that even without tuning any parameters, our LLE-INC is on par with automated verbalizers with parameter tuning. And with the parameter updating, our approach further enhances prompt-based tuning by up to 3.2%. Furthermore, experiments with the LLaMA-7B&13B indicate that LLE-INC is an efficient tuning-free classification approach for the hyper-scale language models. Haochun Wang, Sendong Zhao, Chi Liu 0003, Nuwa Xi, Muzhen Cai, Bing Qin 0001, Ting Liu 0001 |
AAAI | 1 |
| 2024 | CMCOQA: A Chinese Medical Complex Open-Question Answering BenchmarkabstractWith the development of Large Language Models (LLMs), many Chinese medical benchmarks have emerged. These benchmarks have primarily used multiple-choice questions and open-ended questions as test items. However, our experimental results indicate that using multiple-choice questions to test the capabilities of LLMs is not very reasonable. Additionally, relatively simple open-ended questions do not effectively assess LLMs’ actual grasp of medical knowledge. Therefore, we propose the Chinese Medical Complex Open-Question Answering Benchmark (CMCOQA), designed to more accurately and efficiently evaluate the true medical proficiency of LLMs by constructing complex open-ended questions within medical scenarios. Our proposed benchmark involves three evaluation dimensions: Completeness, Depth, and Professionalism. Starting with 100 manually generated complex questions as seeds, we expand the set to 1,200 using the Self-Instruct method with GPT-4o. We then have GPT-4o self-check the questions, followed by a manual screening process to ensure a broad coverage and a certain level of depth. We have both humans and GPT-4o score from these three dimensions, while also employing automated metrics. We also calculate correlations between these metrics and human scores to validate the results. Through this work, CMCOQA can further promote the development of Chinese medical LLMs in terms of medical professionalism. Zijian Li 0020, Sendong Zhao, Haochun Wang, Bing Qin 0001, Ting Liu 0001 |
BIBM | 3 |
| 2023 | UniCoRN: Unified Cognitive Signal ReconstructioN bridging cognitive signals and human languageabstractDecoding text stimuli from cognitive signals (e.g.fMRI) enhances our understanding of the human language system, paving the way for building versatile Brain-Computer Interface.However, existing studies largely focus on decoding individual word-level fMRI volumes from a restricted vocabulary, which is far too idealized for real-world application.In this paper, we propose fMRI2text, the first openvocabulary task aiming to bridge fMRI time series and human language.Furthermore, to explore the potential of this new task, we present a baseline solution, UniCoRN: the Unified Cognitive Signal ReconstructioN for Brain Decoding.By reconstructing both individual time points and time series, UniCoRN establishes a robust encoder for cognitive signals (fMRI & EEG).Leveraging a pre-trained language model as decoder, UniCoRN proves its efficacy in decoding coherent text from fMRI series across various split settings.Our model achieves a 34.77%BLEU score on fMRI2text, and a 37.04% BLEU when generalized to EEGto-text decoding, thereby surpassing the former baseline.Experimental results indicate the feasibility of decoding consecutive fMRI volumes, and the effectiveness of decoding different cognitive signals using a unified structure. Nuwa Xi, Sendong Zhao, Haochun Wang, Chi Liu 0003, Bing Qin 0001, Ting Liu 0001 |
ACL (1) | 3 |
| 2023 | Global Prompt Cell: A Portable Control Module for Effective Prompt Tuning
Chi Liu 0003, Haochun Wang, Nuwa Xi, Sendong Zhao, Bing Qin 0001 |
NLPCC (1) | 2 |
| 2022 | Prompt Combines Paraphrase: Teaching Pre-trained Models to Understand Rare Biomedical WordsabstractPrompt-based fine-tuning for pre-trained models has proven effective for many natural language processing tasks under few-shot settings in general domain. However, tuning with prompt in biomedical domain has not been investigated thoroughly. Biomedical words are often rare in general domain, but quite ubiquitous in biomedical contexts, which dramatically deteriorates the performance of pre-trained models on downstream biomedical applications even after fine-tuning, especially in low-resource scenarios. We propose a simple yet effective approach to helping models learn rare biomedical words during tuning with prompt. Experimental results show that our method can achieve up to 6% improvement in biomedical natural language inference task without any extra parameters or training steps using few-shot vanilla prompt settings. Haochun Wang, Chi Liu 0003, Nuwa Xi, Sendong Zhao, Meizhi Ju, Yefeng Zheng 0001, Bing Qin 0001, Ting Liu 0001 |
COLING | 1 |