VLDB 2026 Research / reviewers in the wild / expert
Zexue He
dblp:215/4688
· DBLP profile ↗
17ranked-venue papers
4as first author
11since 2021 · last 2026
0009-0001-9733-0545ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WildFeedback: Aligning LLMs With In-situ User Interactions And FeedbackabstractTaiwei Shi, Zhuoer Wang, Longqi Yang, Ying-Chun Lin, Zexue He, Mengting Wan, Pei Zhou, Sujay Kumar Jauhar, Sihao Chen, Shan Xia, Hongfei Zhang, Jieyu Zhao, Xiaofeng Xu, Xia Song, Jennifer Neville. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Taiwei Shi, Zhuoer Wang, Longqi Yang 0001, Ying-Chun Lin, Zexue He, Mengting Wan, Sujay Kumar Jauhar, Shan Xia, Jieyu Zhao 0001, Jennifer Neville |
ACL (1) | 5 |
| 2025 | Large Scale Knowledge WashingabstractLarge language models show impressive abilities in memorizing world knowledge, which leads to concerns regarding memorization of private information, toxic or sensitive knowledge, and copyrighted content. We introduce the problem of Large Scale Knowledge Washing, focusing on unlearning an extensive amount of factual knowledge. Previous unlearning methods usually define the reverse loss and update the model via backpropagation, which may affect the model's fluency and reasoning ability or even destroy the model due to extensive training with the reverse loss. Existing works introduce additional data from downstream tasks to prevent the model from losing capabilities, which requires downstream task awareness. Controlling the tradeoff of unlearning existing knowledge while maintaining existing capabilities is also challenging. To this end, we propose LaW (Large Scale Washing), where we update the MLP layers in decoder-only large language models to perform knowledge washing, as inspired by model editing methods. We derive a new objective with the knowledge to be unlearned to update the weights of certain MLP layers. Experimental results demonstrate the effectiveness of LaW in forgetting target knowledge while maximally maintaining reasoning ability. The code will be open-sourced. Yu Wang 0170, Ruihan Wu, Zexue He, Xiusi Chen, Julian J. McAuley |
ICLR | 3 |
| 2025 | M+: Extending MemoryLLM with Scalable Long-Term MemoryabstractEquipping large language models (LLMs) with latent-space memory has attracted increasing attention as they can extend the context window of existing language models. However, retaining information from the distant past remains a challenge. For example, MemoryLLM (Wang et al., 2024a), as a representative work with latent-space memory, compresses past information into hidden states across all layers, forming a memory pool of 1B parameters. While effective for sequence lengths up to 16k tokens, it struggles to retain knowledge beyond 20k tokens. In this work, we address this limitation by introducing M+, a memory-augmented model based on MemoryLLM that significantly enhances long-term information retention. M+ integrates a long-term memory mechanism with a co-trained retriever, dynamically retrieving relevant information during text generation. We evaluate M+ on diverse benchmarks, including long-context understanding and knowledge retention tasks. Experimental results show that M+ significantly outperforms MemoryLLM and recent strong baselines, extending knowledge retention from under 20k to over 160k tokens with similar GPU memory overhead. Yu Wang 0170, Dmitry Krotov, Yifan Gao 0001, Wangchunshu Zhou, Julian J. McAuley, Dan Gutfreund, Rogério Feris, Zexue He |
ICML | 9 |
| 2024 | Deciphering Compatibility Relationships with Textual Descriptions via Extraction and ExplanationabstractUnderstanding and accurately explaining compatibility relationships between fashion items is a challenging problem in the burgeoning domain of AI-driven outfit recommendations. Present models, while making strides in this area, still occasionally fall short, offering explanations that can be elementary and repetitive. This work aims to address these shortcomings by introducing the Pair Fashion Explanation (PFE) dataset, a unique resource that has been curated to illuminate these compatibility relationships. Furthermore, we propose an innovative two stage pipeline model that leverages this dataset. This fine-tuning allows the model to generate explanations that convey the compatibility relationships between items. Our experiments showcase the model's potential in crafting descriptions that are knowledgeable, aligned with ground-truth matching correlations, and that produce understandable and informative descriptions, as assessed by both automatic metrics and human evaluation. Our code and data are released at https://github.com/wangyu-ustc/PairFashionExplanation. Yu Wang 0170, Zexue He, Zhankui He, Julian J. McAuley |
AAAI | 2 |
| 2024 | InfoRank: Unbiased Learning-to-Rank via Conditional Mutual Information MinimizationabstractRanking items regarding individual user interests is a core technique of multiple downstream tasks such as recommender systems. Learning such a personalized ranker typically relies on the implicit feedback from users' past click-through behaviors. However, collected feedback is biased toward previously highly-ranked items and directly learning from it would result in "rich-get-richer" phenomena. In this paper, we propose a simple yet sufficient unbiased learning-to-rank paradigm named InfoRank that aims to simultaneously address both position and popularity biases. We begin by consolidating the impacts of those biases into a single observation factor, thereby providing a unified approach to addressing bias-related issues. Subsequently, we minimize the mutual information between the observation estimation and the relevance estimation conditioned on the input features. By doing so, our relevance estimation can be proved to be free of bias. To implement InfoRank, we first incorporate an attention mechanism to capture latent correlations within user-item features, thereby generating estimations of observation and relevance. We then introduce a regularization term, grounded in conditional mutual information, to promote conditional independence between relevance estimation and observation estimation. Experimental evaluations conducted across three extensive recommendation and search datasets reveal that InfoRank learns more precise and unbiased ranking strategies. Jiarui Jin, Zexue He, Mengyue Yang, Weinan Zhang 0001, Yong Yu 0001, Jun Wang 0012, Julian J. McAuley |
WWW | 2 |
| 2023 | "Nothing Abnormal": Disambiguating Medical Reports via Contrastive Knowledge InfusionabstractSharing medical reports is essential for patient-centered care. A recent line of work has focused on automatically generating reports with NLP methods. However, different audiences have different purposes when writing/reading medical reports – for example, healthcare professionals care more about pathology, whereas patients are more concerned with the diagnosis ("Is there any abnormality?"). The expectation gap results in a common situation where patients find their medical reports to be ambiguous and therefore unsure about the next steps. In this work, we explore the audience expectation gap in healthcare and summarize common ambiguities that lead patients to be confused about their diagnosis into three categories: medical jargon, contradictory findings, and misleading grammatical errors. Based on our analysis, we define a disambiguation rewriting task to regenerate an input to be unambiguous while preserving information about the original content. We further propose a rewriting algorithm based on contrastive pretraining and perturbation-based rewriting. In addition, we create two datasets, OpenI-Annotated based on chest reports and VA-Annotated based on general medical reports, with available binary labels for ambiguity and abnormality presence annotated by radiology specialists. Experimental results on these datasets show that our proposed algorithm effectively rewrites input sentences in a less ambiguous way with high content fidelity. Our code and annotated data will be released to facilitate future research. Zexue He, An Yan 0003, Amilcare Gentili, Julian J. McAuley, Chun-Nan Hsu |
AAAI | 1 |
| 2023 | Targeted Data Generation: Finding and Fixing Model WeaknessesabstractEven when aggregate accuracy is high, stateof-the-art NLP models often fail systematically on specific subgroups of data, resulting in unfair outcomes and eroding user trust.Additional data collection may not help in addressing these weaknesses, as such challenging subgroups may be unknown to users, and underrepresented in the existing and new data.We propose Targeted Data Generation (TDG), a framework that automatically identifies challenging subgroups, and generates new data for those subgroups using large language models (LLMs) with a human in the loop.TDG estimates the expected benefit and potential harm of data augmentation for each subgroup, and selects the ones most likely to improve withingroup performance without hurting overall performance.In our experiments, TDG 1 significantly improves the accuracy on challenging subgroups for state-of-the-art sentiment analysis and natural language inference models, while also improving overall test accuracy. Zexue He, Marco Túlio Ribeiro, Fereshte Khani |
ACL (1) | 1 |
| 2023 | MedEval: A Multi-Level, Multi-Task, and Multi-Domain Medical Benchmark for Language Model EvaluationabstractCurated datasets for healthcare are often limited due to the need of human annotations from experts.In this paper, we present MEDEVAL, a multi-level, multi-task, and multi-domain medical benchmark to facilitate the development of language models for healthcare.MEDEVAL is comprehensive and consists of data from several healthcare systems and spans 35 human body regions from 8 examination modalities.With 22,779 collected sentences and 21,228 reports, we provide expert annotations at multiple levels, offering a granular potential usage of the data and supporting a wide range of tasks.Moreover, we systematically evaluated 10 generic and domain-specific language models under zero-shot and finetuning settings, from domain-adapted baselines in healthcare to general-purposed state-of-the-art large language models (e.g., ChatGPT).Our evaluations reveal varying effectiveness of the two categories of language models across different tasks, from which we notice the importance of instruction tuning for few-shot usage of large language models.Our investigation paves the way toward benchmarking language models for healthcare and provides valuable insights into the strengths and limitations of adopting large language models in medical domains, informing their practical applications and future advancements 1 . Chest Zexue He, Yu Wang 0170, An Yan 0003, Yao Liu 0017, Eric Y. Chang, Amilcare Gentili, Julian J. McAuley, Chun-Nan Hsu |
EMNLP | 1 |
| 2023 | InterFair: Debiasing with Natural Language Feedback for Fair Interpretable PredictionsabstractDebiasing methods in NLP models traditionally focus on isolating information related to a sensitive attribute (e.g.gender or race).We instead argue that a favorable debiasing method should use sensitive information 'fairly,' with explanations, rather than blindly eliminating it.This fair balance is often subjective and can be challenging to achieve algorithmically.We explore two interactive setups with a frozen predictive model and show that users able to provide feedback can achieve a better and fairer balance between task performance and bias mitigation.In one setup, users, by interacting with test examples, further decreased bias in the explanations (5-8%) while maintaining the same prediction accuracy.In the other setup, human feedback was able to disentangle associated bias and predictive information from the input leading to superior bias mitigation and improved task performance (4-5%) simultaneously. Bodhisattwa Prasad Majumder, Zexue He, Julian J. McAuley |
EMNLP | 2 |
| 2023 | Learning Concise and Descriptive Attributes for Visual RecognitionabstractRecent advances in foundation models present new opportunities for interpretable visual recognition – one can first query Large Language Models (LLMs) to obtain a set of attributes that describe each class, then apply vision-language models to classify images via these attributes. Pioneering work shows that querying thousands of attributes can achieve performance competitive with image features. However, our further investigation on 8 datasets reveals that LLM-generated attributes in a large quantity perform almost the same as random words. This surprising finding suggests that significant noise may be present in these attributes. We hypothesize that there exist subsets of attributes that can maintain the classification performance with much smaller sizes, and propose a novel learning-to-search method to discover those concise sets of attributes. As a result, on the CUB dataset, our method achieves performance close to that of massive LLM-generated attributes (e.g., 10k attributes for CUB), yet using only 32 attributes in total to distinguish 200 bird species. Furthermore, our new paradigm demonstrates several additional benefits: higher interpretability and interactivity for humans, and the ability to summarize knowledge for a recognition task. An Yan 0003, Yu Wang 0170, Yiwu Zhong, Chengyu Dong, Zexue He, William Yang Wang, Jingbo Shang, Julian J. McAuley |
ICCV | 5 |
| 2022 | Leashing the Inner Demons: Self-Detoxification for Language ModelsabstractLanguage models (LMs) can reproduce (or amplify) toxic language seen during training, which poses a risk to their practical application. In this paper, we conduct extensive experiments to study this phenomenon. We analyze the impact of prompts, decoding strategies and training corpora on the output toxicity. Based on our findings, we propose a simple yet effective unsupervised method for language models to ``detoxify'' themselves without an additional large corpus or external discriminator. Compared to a supervised baseline, our proposed method shows better toxicity reduction with good generation quality in the generated content under multiple settings. Warning: some examples shown in the paper may contain uncensored offensive content. Canwen Xu, Zexue He, Zhankui He, Julian J. McAuley |
AAAI | 2 |
| 2019 | Learning Robust Representations by Projecting Superficial Statistics Out
Haohan Wang, Zexue He, Zachary C. Lipton, Eric P. Xing |
ICLR | 2 |
| 2019 | Rapid and high-quality 3D fusion of heterogeneous CT and MRI data for the human brain
Zexue He, Minjie Li, Jinyao Li, Yiran Chen 0006, Yanlin Luo |
Sci. China Inf. Sci. | 1 |
| 2018 | A Two-Stage Model for User's Examination Behavior in Mobile SearchabstractWith the rapid growth of mobile search, it is important to understand how users browse the mobile SERPs and allocate their limited attention to each result. To address this problem, we introduce a two-stage examination model that can separately capture the position bias with a skimming model and the attractiveness bias with an attractiveness model. The effectiveness of the proposed model is validated by using a dataset that contains explicit examination feedbacks from users. We further investigate user»s examination behaviors by analyzing the model parameters learned via EM algorithm. The results reveal some interesting findings such as how the skimming behavior is dependent on the previous examination sequence and what factors are associated with the attractiveness of search results on mobile SERPs. Jiaxin Mao, Yiqun Liu 0001, Noriko Kando, Zexue He, Min Zhang 0006, Shaoping Ma |
CHIIR | 4 |
| 2018 | Understanding Reading Attention Distribution during Relevance JudgementabstractReading is a complex cognitive activity in many information retrieval related scenarios, such as relevance judgement and question answering. There exists plenty of works which model these processes as a matching problem, which focuses on how to estimate the relevance score between a document and a query. However, little is known about what happened during the reading process, i.e., how users allocate their attention while reading a document during a specific information retrieval task. We believe that a better understanding of this process can help us design better weighting functions inside the document and contributes to the improvement of ranking performance. In this paper, we focus on the reading process during relevance judgement task. We designed a lab-based user study to investigate human reading patterns in assessing a document, where users' eye movements and their labeled relevant text were collected, respectively. Through a systematic analysis into the collected data, we propose a two-stage reading model which consists of a preliminary relevance judgement stage (Stage 1) and a reading with preliminary relevance stage (Stage 2). In addition, we investigate how different behavior biases affect users' reading behaviors in these two stages. Taking these biases into consideration, we further build prediction models for user's reading attention. Experiment results show that query independent features outperform query dependent features, which indicates that users allocate attentions based on many signals other than query terms in this process. Our study sheds light on the understanding of users' attention allocation during relevance judgement and provides implications for improving the design of existing ranking models. Xiangsheng Li, Yiqun Liu 0001, Jiaxin Mao, Zexue He, Min Zhang 0006, Shaoping Ma |
CIKM | 4 |
| 2018 | Leveraging Gloss Knowledge in Neural Word Sense Disambiguation by Hierarchical Co-AttentionabstractThe goal of Word Sense Disambiguation (WSD) is to identify the correct meaning of a word in the particular context.Traditional supervised methods only use labeled data (context), while missing rich lexical knowledge such as the gloss which defines the meaning of a word sense.Recent studies have shown that incorporating glosses into neural networks for WSD has made significant improvement.However, the previous models usually build the context representation and gloss representation separately.In this paper, we find that the learning for the context and gloss representation can benefit from each other.Gloss can help to highlight the important words in the context, thus building a better context representation.Context can also help to locate the key words in the gloss of the correct word sense.Therefore, we introduce a co-attention mechanism to generate co-dependent representations for the context and gloss.Furthermore, in order to capture both word-level and sentence-level information, we extend the attention mechanism in a hierarchical fashion.Experimental results show that our model achieves the state-of-the-art results on several standard English all-words WSD test datasets. Fuli Luo, Tianyu Liu 0001, Zexue He, Qiaolin Xia, Zhifang Sui, Baobao Chang |
EMNLP | 3 |
| 2018 | A Large-Scale Study of Mobile Search Examination BehaviorabstractWith the rapid growth of mobile web search, it is necessary and important to understand user's examination behavior on mobile devices in the absence of clicks. Previous studies used viewport metrics to estimate user's attention. However, there still lacks an in-depth understanding of how search users examine and interact with the mobile SERP. In this work, based on the large-scale real search log collected from a popular commercial mobile search engine, we present a comprehensive analysis of examination behavior. Specifically, we analyze the position bias, the relationship with click behavior, and examination's change as the session continues. The findings shed new light on the understanding of user's examination behavior, and also provide some implication for the improvement and evaluation of mobile search engine. Zexue He, Yiqun Liu 0001, Shaoping Ma |
SIGIR | 3 |