EDBT 2026 Demo / reviewers in the wild / expert
Weihong Zhong
dblp:166/5143
· DBLP profile ↗
16ranked-venue papers
2as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 2 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic Interpretability and Hallucination MitigationabstractDespite the remarkable advancements of Large Vision-Language Models (LVLMs), the mechanistic interpretability remains underexplored. Existing analyses are insufficiently comprehensive and lack examination covering visual and textual tokens, model components, and the full range of layers. This limitation restricts actionable insights to improve the faithfulness of model output and the development of downstream tasks, such as hallucination mitigation. To address this limitation, we introduce Fine-grained Cross-modal Causal Tracing (FCCT) framework, which systematically quantifies the causal effects on visual object perception. FCCT conducts fine-grained analysis covering the full range of visual and textual tokens, three core model components including multi-head self-attention (MHSA), feed-forward networks (FFNs), and hidden states, across all decoder layers. Our analysis is the first to demonstrate that MHSAs of the last token in middle layers play a critical role in aggregating cross-modal information, while FFNs exhibit a three-stage hierarchical progression for the storage and transfer of visual object representations. Building on these insights, we propose Intermediate Representation Injection (IRI), a training-free inference-time technique that reinforces visual object information flow by precisely intervening on cross-modal representations at specific components and layers, thereby enhancing perception and mitigating hallucination. Consistent improvements across five widely used benchmarks and LVLMs demonstrate IRI achieves state-of-the-art performance, while preserving inference speed and other foundational performance. Zekai Ye, Weihong Zhong, Weitao Ma, Xiachong Feng |
AAAI | 4 |
| 2026 | MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI AgentsabstractRuihan Chen, Qiming Li, Xiaocheng Feng, Weihong Zhong, Xiaoliang Yang, Yuxuan Gu, Zekun Zhou, Yunfei Lu, Haoyu Ren, Kun Chen, Dandan Tu, Bing Qin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ruihan Chen 0001, Weihong Zhong, Xiaoliang Yang, Yuxuan Gu 0004, Ze-kun Zhou, Yunfei Lu, Dandan Tu, Bing Qin 0001 |
ACL (1) | 4 |
| 2026 | Question Tells You Where the Answer Is: Intention-aware Long-Context KV Cache CompressionabstractLiang Zhao, Xiaocheng Feng, Weihong Zhong, Lei Huang, Kun Zhu, Baoxin Wang, Dayong Wu, Guoping Hu, Ting Liu, Bing Qin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Weihong Zhong, Lei Huang 0021, Kun Zhu 0025, Baoxin Wang, Dayong Wu, Ting Liu 0001, Bing Qin 0001 |
ACL (1) | 3 |
| 2025 | Cross-Lingual Text-Rich Visual Comprehension: An Information Theory PerspectiveabstractRecent Large Vision-Language Models (LVLMs) have shown promising reasoning capabilities on text-rich images from charts, tables, and documents. However, the abundant text within such images may increase the model's sensitivity to language. This raises the need to evaluate LVLM performance on cross-lingual text-rich visual inputs, where the language in the image differs from the language of the instructions. To address this, we introduce XT-VQA (Cross-Lingual Text-Rich Visual Question Answering), a benchmark designed to assess how LVLMs handle language inconsistency between image text and questions. XT-VQA integrates five existing text-rich VQA datasets and a newly collected dataset, XPaperQA, covering diverse scenarios that require faithful recognition and comprehension of visual information despite language inconsistency. Our evaluation of prominent LVLMs on XT-VQA reveals a significant drop in performance for cross-lingual scenarios, even for models with multilingual capabilities. A mutual information analysis suggests that this performance gap stems from cross-lingual questions failing to adequately activate relevant visual information. To mitigate this issue, we propose MVCL-MI (Maximization of Vision-Language Cross-Lingual Mutual Information), where a visual-text cross-lingual alignment is built by maximizing mutual information between the model's outputs and visual information. This is achieved by distilling knowledge from monolingual to cross-lingual settings through KL divergence minimization, where monolingual output logits serve as a teacher. Experimental results on the XT-VQA demonstrate that MVCL-MI effectively reduces the visual-text cross-lingual performance disparity while preserving the inherent capabilities of LVLMs, shedding new light on the potential practice for improving LVLMs. Xinmiao Yu, Minghui Liao, Ya-Qi Yu, Xiachong Feng, Weihong Zhong, Ruihan Chen 0001, Mengkang Hu, Jihao Wu, Duyu Tang, Dandan Tu, Bing Qin 0001 |
AAAI | 7 |
| 2025 | Length Controlled Generation for Black-box LLMsabstractYuxuan Gu, Wenjie Wang, Xiaocheng Feng, Weihong Zhong, Kun Zhu, Lei Huang, Ting Liu, Bing Qin, Tat-Seng Chua. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yuxuan Gu 0004, Wenjie Wang 0007, Weihong Zhong, Kun Zhu 0025, Lei Huang 0021, Ting Liu 0001, Bing Qin 0001, Tat-Seng Chua |
ACL (1) | 4 |
| 2025 | Alleviating Hallucinations from Knowledge Misalignment in Large Language Models via Selective Abstention LearningabstractLei Huang, Xiaocheng Feng, Weitao Ma, Yuchun Fan, Xiachong Feng, Yuxuan Gu, Yangfan Ye, Liang Zhao, Weihong Zhong, Baoxin Wang, Dayong Wu, Guoping Hu, Lingpeng Kong, Tong Xiao, Ting Liu, Bing Qin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Lei Huang 0021, Weitao Ma, Yuchun Fan, Xiachong Feng, Yuxuan Gu 0004, Yangfan Ye, Weihong Zhong, Baoxin Wang, Dayong Wu, Lingpeng Kong, Tong Xiao 0001, Ting Liu 0001, Bing Qin 0001 |
ACL (1) | 9 |
| 2025 | Improving Contextual Faithfulness of Large Language Models via Retrieval Heads-Induced OptimizationabstractLei Huang, Xiaocheng Feng, Weitao Ma, Yuchun Fan, Xiachong Feng, Yangfan Ye, Weihong Zhong, Yuxuan Gu, Baoxin Wang, Dayong Wu, Guoping Hu, Bing Qin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Lei Huang 0021, Weitao Ma, Yuchun Fan, Xiachong Feng, Yangfan Ye, Weihong Zhong, Yuxuan Gu 0004, Baoxin Wang, Dayong Wu, Bing Qin 0001 |
ACL (1) | 7 |
| 2025 | Unveiling Entity-Level Unlearning for Large Language Models: A Comprehensive AnalysisabstractLarge language model unlearning has garnered increasing attention due to its potential to address security and privacy concerns, leading to extensive research in the field. However, existing studies have predominantly focused on instance-level unlearning, specifically targeting the removal of predefined instances containing sensitive content. This focus has left a gap in the exploration of removing an entire entity, which is critical in real-world scenarios such as copyright protection. To close this gap, we propose a novel task named Entity-level unlearning, which aims to erase entity-related knowledge from the target model completely. To investigate this task, we systematically evaluate popular unlearning algorithms, revealing that current methods struggle to achieve effective entity-level unlearning. Then, we further explore the factors that influence the performance of unlearning algorithms, identifying that the knowledge coverage of the forget set and its size play pivotal roles. Notably, our analysis also uncovers that entities introduced through fine-tuning are more vulnerable than pre-trained entities during unlearning. We hope these findings can inspire future improvements in entity-level unlearning for LLMs. Weitao Ma, Weihong Zhong, Lei Huang 0021, Yangfan Ye, Xiachong Feng, Bing Qin 0001 |
COLING | 3 |
| 2025 | A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open QuestionsabstractThe emergence of large language models (LLMs) has marked a significant breakthrough in natural language processing (NLP), fueling a paradigm shift in information acquisition. Nevertheless, LLMs are prone to hallucination, generating plausible yet nonfactual content. This phenomenon raises significant concerns over the reliability of LLMs in real-world information retrieval (IR) systems and has attracted intensive research to detect and mitigate such hallucinations. Given the open-ended general-purpose attributes inherent to LLMs, LLM hallucinations present distinct challenges that diverge from prior task-specific models. This divergence highlights the urgency for a nuanced understanding and comprehensive overview of recent advances in LLM hallucinations. In this survey, we begin with an innovative taxonomy of hallucination in the era of LLM and then delve into the factors contributing to hallucinations. Subsequently, we present a thorough overview of hallucination detection methods and benchmarks. Our discussion then transfers to representative methodologies for mitigating LLM hallucinations. Additionally, we delve into the current limitations faced by retrieval-augmented LLMs in combating hallucinations, offering insights for developing more robust IR systems. Finally, we highlight the promising research directions on LLM hallucinations, including hallucination in large vision-language models and understanding of knowledge boundaries in LLM hallucinations. Lei Huang 0021, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang 0007, Qianglong Chen, Weihua Peng, Bing Qin 0001, Ting Liu 0001 |
ACM Trans. Inf. Syst. | 4 |
| 2024 | Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language ModelsabstractWeihong Zhong, Xiaocheng Feng, Liang Zhao, Qiming Li, Lei Huang, Yuxuan Gu, Weitao Ma, Yuan Xu, Bing Qin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Weihong Zhong, Lei Huang 0021, Yuxuan Gu 0004, Weitao Ma, Bing Qin 0001 |
ACL (1) | 1 |
| 2024 | Advancing Large Language Model Attribution through Self-ImprovingabstractLei Huang, Xiaocheng Feng, Weitao Ma, Liang Zhao, Yuchun Fan, Weihong Zhong, Dongliang Xu, Qing Yang, Hongtao Liu, Bing Qin. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Lei Huang 0021, Weitao Ma, Yuchun Fan, Weihong Zhong, Dongliang Xu, Qing Yang 0033, Hongtao Liu 0008, Bing Qin 0001 |
EMNLP | 6 |
| 2024 | Extending Context Window of Large Language Models from a Distributional PerspectiveabstractScaling the rotary position embedding (RoPE) has become a common method for extending the context window of RoPE-based large language models (LLMs).However, existing scaling methods often rely on empirical approaches and lack a profound understanding of the internal distribution within RoPE, resulting in suboptimal performance in extending the context window length.In this paper, we propose to optimize the context window extending task from the view of rotary angle distribution.Specifically, we first estimate the distribution of the rotary angles within the model and analyze the extent to which length extension perturbs this distribution.Then, we present a novel extension strategy that minimizes the disturbance between rotary angle distributions to maintain consistency with the pre-training phase, enhancing the model's capability to generalize to longer sequences.Experimental results compared to the strong baseline methods demonstrate that our approach reduces by up to 72% of the distributional disturbance when extending LLaMA2's context window to 8k, and reduces by up to 32% when extending to 16k.On the LongBench-E benchmark, our method achieves an average improvement of up to 4.33% over existing state-of-the-art methods.Furthermore, our method maintains the model's performance on the Hugging Face Open LLM benchmark after context window extension, with only an average performance fluctuation ranging from -0.12 to +0.22.Our code is available at https: //github.com/1180301012/DPRoPE. Yingsheng Wu, Yuxuan Gu 0004, Weihong Zhong, Dongliang Xu, Qing Yang 0033, Hongtao Liu 0008, Bing Qin 0001 |
EMNLP | 4 |
| 2024 | Discrete Modeling via Boundary Conditional Diffusion ProcessesabstractWe present an novel framework for efficiently and effectively extending the powerful continuous diffusion processes to discrete modeling.
Previous approaches have suffered from the discrepancy between discrete data and continuous modeling.
Our study reveals that the absence of guidance from discrete boundaries in learning probability contours is one of the main reasons.
To address this issue, we propose a two-step forward process that first estimates the boundary as a prior distribution and then rescales the forward trajectory to construct a boundary conditional diffusion model.
The reverse process is proportionally adjusted to guarantee that the learned contours yield more precise discrete data.
Experimental results indicate that our approach achieves strong performance in both language modeling and discrete image generation tasks.
In language modeling, our approach surpasses previous state-of-the-art continuous diffusion language models in three translation tasks and a summarization task, while also demonstrating competitive performance compared to auto-regressive transformers. Moreover, our method achieves comparable results to continuous diffusion models when using discrete ordinal pixels and establishes a new state-of-the-art for categorical image generation on the Cifar-10 dataset. Yuxuan Gu 0004, Lei Huang 0021, Yingsheng Wu, Ze-kun Zhou, Weihong Zhong, Kun Zhu 0025, Bing Qin 0001 |
NeurIPS | 6 |
| 2023 | STOA-VLP: Spatial-Temporal Modeling of Object and Action for Video-Language Pre-trainingabstractAlthough large-scale video-language pre-training models, which usually build a global alignment between the video and the text, have achieved remarkable progress on various downstream tasks, the idea of adopting fine-grained information during the pre-training stage is not well explored. In this work, we propose STOA-VLP, a pre-training framework that jointly models object and action information across spatial and temporal dimensions. More specifically, the model regards object trajectories across frames and multiple action features from the video as fine-grained features. Besides, We design two auxiliary tasks to better incorporate both kinds of information into the pre-training process of the video-language model. The first is the dynamic object-text alignment task, which builds a better connection between object trajectories and the relevant noun tokens. The second is the spatial-temporal action set prediction, which guides the model to generate consistent action features by predicting actions found in the text. Extensive experiments on three downstream tasks (video captioning, text-video retrieval, and video question answering) demonstrate the effectiveness of our proposed STOA-VLP (e.g. 3.7 Rouge-L improvements on MSR-VTT video captioning benchmark, 2.9% accuracy improvements on MSVD video question answering benchmark, compared to previous approaches). Weihong Zhong, Mao Zheng, Duyu Tang, Heng Gong, Bing Qin 0001 |
AAAI | 1 |
| 2023 | Controllable Text Generation via Probability Density Estimation in the Latent SpaceabstractYuxuan Gu, Xiaocheng Feng, Sicheng Ma, Lingyuan Zhang, Heng Gong, Weihong Zhong, Bing Qin. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yuxuan Gu 0004, Sicheng Ma, Lingyuan Zhang, Heng Gong, Weihong Zhong, Bing Qin 0001 |
ACL (1) | 6 |
| 2018 | VRprofile: gene-cluster-detection-based profiling of virulence and antibiotic resistance traits encoded within genome sequences of pathogenic bacteriaabstractVRprofile is a Web server that facilitates rapid investigation of virulence and antibiotic resistance genes, as well as extends these trait transfer-related genetic contexts, in newly sequenced pathogenic bacterial genomes. The used backend database MobilomeDB was firstly built on sets of known gene cluster loci of bacterial type III/IV/VI/VII secretion systems and mobile genetic elements, including integrative and conjugative elements, prophages, class I integrons, IS elements and pathogenicity/antibiotic resistance islands. VRprofile is thus able to co-localize the homologs of these conserved gene clusters using HMMer or BLASTp searches. With the integration of the homologous gene cluster search module with a sequence composition module, VRprofile has exhibited better performance for island-like region predictions than the other widely used methods. In addition, VRprofile also provides an integrated Web interface for aligning and visualizing identified gene clusters with MobilomeDB-archived gene clusters, or a variety set of bacterial genomes. VRprofile might contribute to meet the increasing demands of re-annotations of bacterial variable regions, and aid in the real-time definitions of disease-relevant gene clusters in pathogenic bacteria of interest. VRprofile is freely available at http://bioinfo-mml.sjtu.edu.cn/VRprofile. Cui Tai, Zixin Deng, Weihong Zhong, Yongqun He, Hong-Yu Ou |
Briefings Bioinform. | 4 |