VLDB 2026 Research / reviewers in the wild / expert
Haotian Wang 0007
dblp:63/11345-7
· DBLP profile ↗
17ranked-venue papers
3as first author
17since 2021 · last 2026
0009-0001-5363-3886ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AgriEval: A Comprehensive Chinese Agricultural Benchmark for Large Language Modelsabstractn the agricultural domain, the deployment of large language models (LLMs) is hindered by the lack of training data and evaluation benchmarks. To mitigate this issue, we propose AgriEval, the first comprehensive Chinese agricultural benchmark with three main characteristics: (1) Comprehensive Capability Evaluation. AgriEval covers six major agriculture categories and 29 subcategories within agriculture, addressing four core cognitive scenarios—memorization, understanding, inference, and generation. (2) High-Quality Data. The dataset is curated from university-level examinations and assignments, providing a natural and robust benchmark for assessing the capacity of LLMs to apply knowledge and make expert-like decisions. (3) Diverse Formats and Extensive Scale. AgriEval comprises 14,697 multiple-choice questions and 2,167 open-ended question-and-answer questions, establishing it as the most extensive agricultural benchmark available to date. We also present comprehensive experimental results over 51 open-source and commercial LLMs. The experimental results reveal that most existing LLMs struggle to achieve 60 percent accuracy, underscoring the developmental potential in agricultural LLMs. Additionally, we conduct extensive experiments to investigate factors influencing model performance and propose strategies for enhancement. Lian Yan, Haotian Wang 0007, Tianyang Sun, Liangliang Liu 0002, Yi Guan, Jingchi Jiang |
AAAI | 2 |
| 2026 | Towards Efficient and Generalizable Retrieval: Adaptive Semantic Quantization and Residual Knowledge TransferabstractWhile semantic ID-based generative retrieval enables efficient end-to-end modeling in industrial applications, these methods face a persistent trade-off. On one hand, data-rich head items often suffer from ID collisions, which blur their distinct features and degrade downstream tasks. On the other hand, data-sparse tail items especially cold-start items are prone to semantic fragmentation during quantization; they are often mapped as isolated discrete points, which severely hinders their ability to generalize. To address this issue, we propose the Anchored Curriculum with Sequential Adaptive Quantization (SA2CRQ) framework. The framework introduces Sequential Adaptive Residual Quantization (SARQ) to dynamically allocate code lengths based on item path entropy, assigning longer, discriminative IDs to head items and shorter, generalizable IDs to tail items. To mitigate data sparsity, the Anchored Curriculum Residual Quantization (ACRQ) component utilizes a frozen semantic manifold learned from head items to regularize and accelerate the representation learning of tail items. Experimental results from a large-scale industrial search system and multiple public datasets indicate that SA2CRQ yields consistent improvements over existing baselines, particularly in cold-start retrieval scenarios. Huimu Wang, Xingzhi Yao, Yiming Qiu 0003, Qinghong Zhang, Haotian Wang 0007, Yufan Cui, Songlin Wang, Sulong Xu |
SIGIR | 5 |
| 2026 | PestScope: Exclusion-Aware Large Multimodal Model for Fine-Grained Agricultural Pest SegmentationabstractReasoning segmentation (RS) interprets implicit textual instructions to accurately segment target regions. This reasoning capability transforms ambiguous non-expert queries into precise pixel-level masks, thereby enabling downstream tasks like area measurement and density analysis with a level of precision unattainable by detection methods. However, existing RS models are not tailored for agriculture and lack domain-specific knowledge, which poses challenges in handling similar pest appearances and small target scales. To bridge this gap, we introduce a fine-grained pest RS task with two subtasks: Pest Discriminative Referring Expression Segmentation (PDRES) and Pest Exclusion Reasoning Segmentation (PERS). Based on this, we propose PestScope, which integrates vision, language, and reasoning for fine-grained pest segmentation. To tackle the exclusion of small non-target pests, we introduce a dedicated [NON] token alongside the standard [SEG] token for target pests. This guides the model to prioritize small target pests and suppress non-target background regions. To further address pest similarity, we propose an Exclusivity Suppression Loss, applying differentiated supervision to [SEG] and [NON] tokens to better separate target and non-target pests. Additionally, we develop an automated dataset construction pipeline to address the scarcity of fine-grained, difficulty-controllable pest RS datasets. It produces 45k and 27.6k image-text-mask samples for the PDRES and PERS tasks, respectively, covering 18 pest categories. Experiments show that in small and similar pest scenarios, integrating PestScope into mainstream models improves average gIoU by 4.28% on PDRES and 6.49% on PERS. For unseen pest categories, gIoU increases by 21.72% and 8.66%, respectively, demonstrating strong generalization. Code and datasets will be available at: https://github.com/aluodaydayup/PestScope. Yang Yang 0137, Huibin Luo, Haotian Wang 0007, Jingchi Jiang, Jie Liu 0001, Ming Fang 0005 |
IEEE Trans. Image Process. | 3 |
| 2025 | Agri-CM³: A Chinese Massive Multi-modal, Multi-level Benchmark for Agricultural Understanding and ReasoningabstractMulti-modal Large Language Models (MLLMs) integrating images, text, and speech can provide farmers with accurate diagnoses and treatment of pests and diseases, enhancing agricultural efficiency and sustainability. However, existing benchmarks lack comprehensive evaluations, particularly in multi-level reasoning, making it challenging to identify model limitations. To address this issue, we introduce Agri-CM^3, an expert-validated benchmark assessing MLLMs’ understanding and reasoning in agricultural management. It includes 3,939 images and 15,901 multi-level multiple-choice questions with detailed explanations. Evaluations of 45 MLLMs reveal significant gaps. Even GPT-4o achieves only 63.64% accuracy, falling short in fine-grained reasoning tasks. Analysis across three reasoning levels and seven compositional abilities highlights key challenges in accuracy and cognitive understanding. Our study provides insights for advancing MLLMs in agricultural management, driving their development and application. Code and data are available at https://github.com/HIT-Kwoo/Agri-CM3. Haotian Wang 0007, Yi Guan, Fanshu Meng, Chao Zhao 0002, Lian Yan, Yang Yang 0041, Jingchi Jiang |
ACL (1) | 1 |
| 2025 | Modeling clinical thinking based on knowledge hypergraph attention network and prompt learning for disease prediction
Yang Yang 0137, Xin Li 0012, Haotian Wang 0007, Yi Guan, Jingchi Jiang |
Expert Syst. Appl. | 3 |
| 2025 | Learning to break: Knowledge-enhanced reasoning in multi-agent debate system
Haotian Wang 0007, Xiyuan Du, Weijiang Yu, Qianglong Chen, Kun Zhu 0025, Lian Yan, Yi Guan |
Neurocomputing | 1 |
| 2025 | Quality-Controllable automatic construction method of Chinese knowledge graph for medical decision-making applications
Yang Yang 0137, Yi Guan, Haotian Wang 0007, Jingchi Jiang, Huaizhang Shi, Xiguang Liu |
Inf. Process. Manag. | 5 |
| 2025 | Knowledge assimilation: Implementing knowledge-guided agricultural large language model
Jingchi Jiang, Lian Yan, Zhenbo Xia, Haotian Wang 0007, Yang Yang 0137, Yi Guan |
Knowl. Based Syst. | 5 |
| 2025 | A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open QuestionsabstractThe emergence of large language models (LLMs) has marked a significant breakthrough in natural language processing (NLP), fueling a paradigm shift in information acquisition. Nevertheless, LLMs are prone to hallucination, generating plausible yet nonfactual content. This phenomenon raises significant concerns over the reliability of LLMs in real-world information retrieval (IR) systems and has attracted intensive research to detect and mitigate such hallucinations. Given the open-ended general-purpose attributes inherent to LLMs, LLM hallucinations present distinct challenges that diverge from prior task-specific models. This divergence highlights the urgency for a nuanced understanding and comprehensive overview of recent advances in LLM hallucinations. In this survey, we begin with an innovative taxonomy of hallucination in the era of LLM and then delve into the factors contributing to hallucinations. Subsequently, we present a thorough overview of hallucination detection methods and benchmarks. Our discussion then transfers to representative methodologies for mitigating LLM hallucinations. Additionally, we delve into the current limitations faced by retrieval-augmented LLMs in combating hallucinations, offering insights for developing more robust IR systems. Finally, we highlight the promising research directions on LLM hallucinations, including hallucination in large vision-language models and understanding of knowledge boundaries in LLM hallucinations. Lei Huang 0021, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang 0007, Qianglong Chen, Weihua Peng, Bing Qin 0001, Ting Liu 0001 |
ACM Trans. Inf. Syst. | 6 |
| 2024 | An Information Bottleneck Perspective for Effective Noise Filtering on Retrieval-Augmented GenerationabstractKun Zhu, Xiaocheng Feng, Xiyuan Du, Yuxuan Gu, Weijiang Yu, Haotian Wang, Qianglong Chen, Zheng Chu, Jingchang Chen, Bing Qin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Kun Zhu 0025, Xiyuan Du, Yuxuan Gu 0004, Weijiang Yu, Haotian Wang 0007, Qianglong Chen, Jingchang Chen, Bing Qin 0001 |
ACL (1) | 6 |
| 2024 | BeamAggR: Beam Aggregation Reasoning over Multi-source Knowledge for Multi-hop Question AnsweringabstractZheng Chu, Jingchang Chen, Qianglong Chen, Haotian Wang, Kun Zhu, Xiyuan Du, Weijiang Yu, Ming Liu, Bing Qin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Jingchang Chen, Qianglong Chen, Haotian Wang 0007, Kun Zhu 0025, Xiyuan Du, Weijiang Yu, Ming Liu 0004, Bing Qin 0001 |
ACL (1) | 4 |
| 2024 | TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language ModelsabstractZheng Chu, Jingchang Chen, Qianglong Chen, Weijiang Yu, Haotian Wang, Ming Liu, Bing Qin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Jingchang Chen, Qianglong Chen, Weijiang Yu, Haotian Wang 0007, Ming Liu 0004, Bing Qin 0001 |
ACL (1) | 5 |
| 2024 | Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and FutureabstractZheng Chu, Jingchang Chen, Qianglong Chen, Weijiang Yu, Tao He, Haotian Wang, Weihua Peng, Ming Liu, Bing Qin, Ting Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Jingchang Chen, Qianglong Chen, Weijiang Yu, Tao He 0014, Haotian Wang 0007, Weihua Peng, Ming Liu 0004, Bing Qin 0001, Ting Liu 0001 |
ACL (1) | 6 |
| 2024 | EIRAD: An Evidence-Based Dialogue System With Highly Interpretable Reasoning Path for Automatic DiagnosisabstractDialogue System for Medical Diagnosis (DSMD) based on reinforcement learning (RL) can simulate patient-doctor interactions, playing a crucial role in clinical diagnosis. However, due to the complexity of disease etiology, DSMD faces the challenges of low efficiency in diagnostic evidence search. Moreover, solely RL-based DSMS, without the constraints of professional medical knowledge, often generates irrational, meaningless, or even erroneous symptom inquiries, leading to poor interpretability of diagnostic path and high misdiagnosis rates. To address these issues, we propose anEvidence-based dialogue system with highlyInterpretableReasoning path forAutomaticDiagnosis (EIRAD) grounded in medical knowledge graph (MKG). Specifically, our automated diagnostic model captures key symptoms for suspected diseases by explicitly leveraging the topology of MKG, enhancing the interpretability and accuracy of diagnosis. To expedite the retrieval of factual evidence, we develop two mechanisms: 1) Mapping mechanism between the entity set of MKG and DSMD's diagnostic evidence and diseases. According to the patient's symptoms, EIRAD prunes irrelevant disease and symptom nodes from the MKG, which can truncate the invalid action of RL-based DSMD. 2) Reward Mechanism of integrating the effectiveness of symptom inquiry and the accuracy of disease diagnosis. The comprehensive reward system is suitable for intelligent consultation, which can effectively drive DSMD to accelerate evidence collection. Experimental results demonstrate that our model significantly outperforms competitive benchmark methods in symptom inquiry efficiency and diagnostic accuracy. Lian Yan, Yi Guan, Haotian Wang 0007, Yang Yang 0137, Boran Wang, Jingchi Jiang |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | Efficient Evidence-Based Dialogue System for Medical DiagnosisabstractWith the rise of intelligent medical assistance, the Dialogue System for Medical Diagnosis(DSMD) guided by reinforcement learning(RL) has gained much attention. However, currently available medical dialogue datasets suffer from insufficient diagnostic evidence caused by sparse symptoms, making it difficult to reproduce the evidence-based process of doctors in differential diagnosis and disease confirmation. Moreover, purely data-driven RL often involves extensive and blind trial-and-error, leading to inquiries about irrelevant symptoms to the patient’s chief complaints in limited dialogue turns, further exacerbating the issue of inadequate diagnostic evidence. To enhance the quantity and effectiveness of potential symptom collection in DSMD, we first construct a more comprehensive medical dialogue dataset CMD based on electronic medical records. The diversity of diseases and symptoms mentioned in the dialogue context of CMD surpasses that of existing public datasets. Furthermore, to enhance the efficiency of diagnostic evidence collection in DSMD, inspired by the logic of symptom inquiries in doctor-patient interactions, we combine experiential diagnostic knowledge with a specialized medical knowledge graph to constrain the inquiry of symptoms via RL, eliminating the introduction of symptoms unrelated to the patient. Experimental results demonstrate that our model significantly outperforms competitive benchmark methods in terms of diagnostic accuracy and the efficiency of symptom inquiries. Our codes and the CMD dataset are available at https://github.com/YanPioneer/EBAD. Lian Yan, Yi Guan, Haotian Wang 0007, Jingchi Jiang |
BIBM | 3 |
| 2023 | LHP: Logical hypergraph link prediction
Yang Yang 0137, Yi Guan, Haotian Wang 0007, Chaoran Kong, Jingchi Jiang |
Expert Syst. Appl. | 4 |
| 2022 | Multi-scale Label Attention Network based on Abductive Causal Graph for Disease DiagnosisabstractThe auxiliary disease diagnosis based on electronic medical records is of great significance, providing doctors with diagnostic advice and avoiding misdiagnosis. Existing work on disease diagnosis mainly utilizes deep learning models to extract sequence information in electronic medical records, ignoring the interpretability of results and the structural knowledge, especially causal knowledge. In our work, we propose a multiscale label attention network based on abductive causal graph (MSLAN-ACG) to improve model accuracy and interpretability of results. First, we construct multiple encoders in the multiscale label attention network, which can extract n-gram segment information of different lengths for each disease. Meanwhile, to enhance the interpretability of results, we visualize the weight score of different segments for disease results. Second, we propose a disease representation method by defining an abductive causal graph and then using graph convolutional network for knowledge fusion on this graph. The information propagation based on abductive causal graph is consistent with the actual abductive reasoning process from symptoms to diseases, making the model more reasonable. The effectiveness of our model is demonstrated by achieving state-of-the-art results on MIMICIII-50 and ChineseEMR datasets. Haotian Wang 0007, Yi Guan, Linjiang Ma, Xin Li 0012, Jing Xie 0012, Jingchi Jiang |
BIBM | 1 |