EDBT 2026 Demo / reviewers in the wild / expert
Lei Huang 0021
dblp:18/1763-21
· DBLP profile ↗
16ranked-venue papers
4as first author
16since 2021 · last 2026
0009-0002-9650-8353ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 3 first-author · 15 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LangGPS: Language Separability Guided Data Pre-Selection for Joint Multilingual Instruction TuningabstractJoint multilingual instruction tuning is a widely adopted approach to improve the multilingual instruction-following ability and downstream performance of large language models (LLMs), but the resulting multilingual capability remains highly sensitive to the composition and selection of the training data. Existing selection methods, often based on features like text quality, diversity, or task relevance, typically overlook the intrinsic linguistic structure of multilingual data. In this paper, we propose LangGPS, a lightweight two-stage pre-selection framework guided by language separability—a signal that quantifies how well samples in different languages can be distinguished in the model’s representation space. LangGPS first filters training data based on separability scores and then refines the subset using existing selection methods. Extensive experiments across six benchmarks and 22 languages demonstrate that applying LangGPS on top of existing selection methods improves their effectiveness and generalizability in multilingual training, especially for understanding tasks and low-resource languages. Further analysis reveals that highly separable samples facilitate the formation of clearer language boundaries and support faster adaptation, while low-separability samples tend to function as bridges for cross-lingual alignment. Besides, we also find that language separability can serves as an effective signal for multilingual curriculum learning, where interleaving samples with diverse separability levels yields stable and generalizable gains. Together, we hope our work offers a new perspective on data utility in multilingual contexts and support the development of more linguistically informed LLMs. Yangfan Ye, Xiachong Feng, Lei Huang 0021, Weitao Ma, Qichen Hong, Yunfei Lu, Duyu Tang, Dandan Tu, Bing Qin 0001 |
AAAI | 4 |
| 2026 | Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-PlayabstractXiachong Feng, Deyi Yin, Xiaocheng Feng, Yi Jiang, Libo Qin, Yangfan Ye, Lei Huang, Weitao Ma, Qiming Li, Yuxuan Gu, Bing Qin, Lingpeng Kong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xiachong Feng, Deyi Yin, Libo Qin 0001, Yangfan Ye, Lei Huang 0021, Weitao Ma, Yuxuan Gu 0004, Bing Qin 0001, Lingpeng Kong |
ACL (1) | 7 |
| 2026 | Fine-Mem: Fine-Grained Feedback Alignment for Long-Horizon Memory ManagementabstractWeitao Ma, Xiaocheng Feng, Lei Huang, Xiachong Feng, Zhanyu Ma, Jun Xu, Jiuchong Gao, Jinghua Hao, Renqing He, Bing Qin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Weitao Ma, Lei Huang 0021, Xiachong Feng, Zhanyu Ma, Jun Xu 0001, Jiuchong Gao, Jinghua Hao, Renqing He, Bing Qin 0001 |
ACL (1) | 3 |
| 2026 | Question Tells You Where the Answer Is: Intention-aware Long-Context KV Cache CompressionabstractLiang Zhao, Xiaocheng Feng, Weihong Zhong, Lei Huang, Kun Zhu, Baoxin Wang, Dayong Wu, Guoping Hu, Ting Liu, Bing Qin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Weihong Zhong, Lei Huang 0021, Kun Zhu 0025, Baoxin Wang, Dayong Wu, Ting Liu 0001, Bing Qin 0001 |
ACL (1) | 4 |
| 2025 | Length Controlled Generation for Black-box LLMsabstractYuxuan Gu, Wenjie Wang, Xiaocheng Feng, Weihong Zhong, Kun Zhu, Lei Huang, Ting Liu, Bing Qin, Tat-Seng Chua. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yuxuan Gu 0004, Wenjie Wang 0007, Weihong Zhong, Kun Zhu 0025, Lei Huang 0021, Ting Liu 0001, Bing Qin 0001, Tat-Seng Chua |
ACL (1) | 6 |
| 2025 | Alleviating Hallucinations from Knowledge Misalignment in Large Language Models via Selective Abstention LearningabstractLei Huang, Xiaocheng Feng, Weitao Ma, Yuchun Fan, Xiachong Feng, Yuxuan Gu, Yangfan Ye, Liang Zhao, Weihong Zhong, Baoxin Wang, Dayong Wu, Guoping Hu, Lingpeng Kong, Tong Xiao, Ting Liu, Bing Qin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Lei Huang 0021, Weitao Ma, Yuchun Fan, Xiachong Feng, Yuxuan Gu 0004, Yangfan Ye, Weihong Zhong, Baoxin Wang, Dayong Wu, Lingpeng Kong, Tong Xiao 0001, Ting Liu 0001, Bing Qin 0001 |
ACL (1) | 1 |
| 2025 | Improving Contextual Faithfulness of Large Language Models via Retrieval Heads-Induced OptimizationabstractLei Huang, Xiaocheng Feng, Weitao Ma, Yuchun Fan, Xiachong Feng, Yangfan Ye, Weihong Zhong, Yuxuan Gu, Baoxin Wang, Dayong Wu, Guoping Hu, Bing Qin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Lei Huang 0021, Weitao Ma, Yuchun Fan, Xiachong Feng, Yangfan Ye, Weihong Zhong, Yuxuan Gu 0004, Baoxin Wang, Dayong Wu, Bing Qin 0001 |
ACL (1) | 1 |
| 2025 | One for All: Update Parameterized Knowledge Across Multiple Models with Once EditabstractWeitao Ma, Xiyuan Du, Xiaocheng Feng, Lei Huang, Yichong Huang, Huiyi Zhang, Xiaoliang Yang, Baohang Li, Xiachong Feng, Ting Liu, Bing Qin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Weitao Ma, Xiyuan Du, Lei Huang 0021, Yichong Huang, Huiyi Zhang, Xiaoliang Yang, Baohang Li, Xiachong Feng, Ting Liu 0001, Bing Qin 0001 |
ACL (1) | 4 |
| 2025 | CC-Tuning: A Cross-Lingual Connection Mechanism for Improving Joint Multilingual Supervised Fine-TuningabstractYangfan Ye, Xiaocheng Feng, Zekun Yuan, Xiachong Feng, Libo Qin, Lei Huang, Weitao Ma, Yichong Huang, Zhirui Zhang, Yunfei Lu, Xiaohui Yan, Duyu Tang, Dandan Tu, Bing Qin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yangfan Ye, Zekun Yuan, Xiachong Feng, Libo Qin 0001, Lei Huang 0021, Weitao Ma, Yichong Huang, Zhirui Zhang, Yunfei Lu, Duyu Tang, Dandan Tu, Bing Qin 0001 |
ACL (1) | 6 |
| 2025 | SLAM: Towards Efficient Multilingual Reasoning via Selective Language AlignmentabstractDespite the significant improvements achieved by large language models (LLMs) in English reasoning tasks, these models continue to struggle with multilingual reasoning. Recent studies leverage a full-parameter and two-stage training paradigm to teach models to first understand non-English questions and then reason. However, this method suffers from both substantial computational resource computing and catastrophic forgetting. The fundamental cause is that, with the primary goal of enhancing multilingual comprehension, an excessive number of irrelevant layers and parameters are tuned during the first stage. Given our findings that the representation learning of languages is merely conducted in lower-level layers, we propose an efficient multilingual reasoning alignment approach that precisely identifies and fine-tunes the layers responsible for handling multilingualism. Experimental results show that our method, SLAM, only tunes 6 layers’ feed-forward sub-layers including 6.5-8% of all parameters within 7B and 13B LLMs, achieving superior average performance than all strong baselines across 10 languages. Meanwhile, SLAM only involves one training stage, reducing training time by 4.1-11.9× compared to the two-stage method. Yuchun Fan, Yongyu Mu, Lei Huang 0021, Junhao Ruan, Tong Xiao 0001, Shujian Huang |
COLING | 4 |
| 2025 | Unveiling Entity-Level Unlearning for Large Language Models: A Comprehensive AnalysisabstractLarge language model unlearning has garnered increasing attention due to its potential to address security and privacy concerns, leading to extensive research in the field. However, existing studies have predominantly focused on instance-level unlearning, specifically targeting the removal of predefined instances containing sensitive content. This focus has left a gap in the exploration of removing an entire entity, which is critical in real-world scenarios such as copyright protection. To close this gap, we propose a novel task named Entity-level unlearning, which aims to erase entity-related knowledge from the target model completely. To investigate this task, we systematically evaluate popular unlearning algorithms, revealing that current methods struggle to achieve effective entity-level unlearning. Then, we further explore the factors that influence the performance of unlearning algorithms, identifying that the knowledge coverage of the forget set and its size play pivotal roles. Notably, our analysis also uncovers that entities introduced through fine-tuning are more vulnerable than pre-trained entities during unlearning. We hope these findings can inspire future improvements in entity-level unlearning for LLMs. Weitao Ma, Weihong Zhong, Lei Huang 0021, Yangfan Ye, Xiachong Feng, Bing Qin 0001 |
COLING | 4 |
| 2025 | Context-Aware Hierarchical Taxonomy Generation for Scientific Papers via LLM-Guided Multi-Aspect ClusteringabstractThe rapid growth of scientific literature demands efficient methods to organize and synthesize research findings.Existing taxonomy construction methods, leveraging unsupervised clustering or direct prompting of large language models (LLMs), often lack coherence and granularity.We propose a novel context-aware hierarchical taxonomy generation framework that integrates LLM-guided multi-aspect encoding with dynamic clustering.Our method leverages LLMs to identify key aspects of each paper (e.g., methodology, dataset, evaluation) and generates aspect-specific paper summaries, which are then encoded and clustered along each aspect to form a coherent hierarchy.In addition, we introduce a new benchmark of 156 expert-crafted taxonomies encompassing 11.6 k papers, providing the first naturally annotated dataset for this task.Experimental results demonstrate that our method significantly outperforms prior approaches, achieving stateof-the-art performance in taxonomy coherence, granularity, and interpretability. 1 Kun Zhu 0025, Lizi Liao, Yuxuan Gu 0004, Lei Huang 0021, Bing Qin 0001 |
EMNLP | 4 |
| 2025 | A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open QuestionsabstractThe emergence of large language models (LLMs) has marked a significant breakthrough in natural language processing (NLP), fueling a paradigm shift in information acquisition. Nevertheless, LLMs are prone to hallucination, generating plausible yet nonfactual content. This phenomenon raises significant concerns over the reliability of LLMs in real-world information retrieval (IR) systems and has attracted intensive research to detect and mitigate such hallucinations. Given the open-ended general-purpose attributes inherent to LLMs, LLM hallucinations present distinct challenges that diverge from prior task-specific models. This divergence highlights the urgency for a nuanced understanding and comprehensive overview of recent advances in LLM hallucinations. In this survey, we begin with an innovative taxonomy of hallucination in the era of LLM and then delve into the factors contributing to hallucinations. Subsequently, we present a thorough overview of hallucination detection methods and benchmarks. Our discussion then transfers to representative methodologies for mitigating LLM hallucinations. Additionally, we delve into the current limitations faced by retrieval-augmented LLMs in combating hallucinations, offering insights for developing more robust IR systems. Finally, we highlight the promising research directions on LLM hallucinations, including hallucination in large vision-language models and understanding of knowledge boundaries in LLM hallucinations. Lei Huang 0021, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang 0007, Qianglong Chen, Weihua Peng, Bing Qin 0001, Ting Liu 0001 |
ACM Trans. Inf. Syst. | 1 |
| 2024 | Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language ModelsabstractWeihong Zhong, Xiaocheng Feng, Liang Zhao, Qiming Li, Lei Huang, Yuxuan Gu, Weitao Ma, Yuan Xu, Bing Qin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Weihong Zhong, Lei Huang 0021, Yuxuan Gu 0004, Weitao Ma, Bing Qin 0001 |
ACL (1) | 5 |
| 2024 | Advancing Large Language Model Attribution through Self-ImprovingabstractLei Huang, Xiaocheng Feng, Weitao Ma, Liang Zhao, Yuchun Fan, Weihong Zhong, Dongliang Xu, Qing Yang, Hongtao Liu, Bing Qin. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Lei Huang 0021, Weitao Ma, Yuchun Fan, Weihong Zhong, Dongliang Xu, Qing Yang 0033, Hongtao Liu 0008, Bing Qin 0001 |
EMNLP | 1 |
| 2024 | Discrete Modeling via Boundary Conditional Diffusion ProcessesabstractWe present an novel framework for efficiently and effectively extending the powerful continuous diffusion processes to discrete modeling.
Previous approaches have suffered from the discrepancy between discrete data and continuous modeling.
Our study reveals that the absence of guidance from discrete boundaries in learning probability contours is one of the main reasons.
To address this issue, we propose a two-step forward process that first estimates the boundary as a prior distribution and then rescales the forward trajectory to construct a boundary conditional diffusion model.
The reverse process is proportionally adjusted to guarantee that the learned contours yield more precise discrete data.
Experimental results indicate that our approach achieves strong performance in both language modeling and discrete image generation tasks.
In language modeling, our approach surpasses previous state-of-the-art continuous diffusion language models in three translation tasks and a summarization task, while also demonstrating competitive performance compared to auto-regressive transformers. Moreover, our method achieves comparable results to continuous diffusion models when using discrete ordinal pixels and establishes a new state-of-the-art for categorical image generation on the Cifar-10 dataset. Yuxuan Gu 0004, Lei Huang 0021, Yingsheng Wu, Ze-kun Zhou, Weihong Zhong, Kun Zhu 0025, Bing Qin 0001 |
NeurIPS | 3 |