Li Zhou 0010

dblp:54/40-10 · DBLP profile ↗
← Back
22ranked-venue papers
6as first author
21since 2021 · last 2026
0000-0002-5510-6406ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 5 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ScholarLens: Tracking the growing penetration of Large Language Models in scholarly writing and peer review
abstract
Although the widespread use of Large Language Models (LLMs) brings convenience, it also raises concerns about the credibility of academic research and scholarly processes. To better understand the extent and characteristics of LLM use in scholarly writing and peer review, the penetration of LLMs across academic workflows is evaluated from multiple perspectives and dimensions, providing compelling evidence of their growing influence. A framework consisting of two components is proposed: ScholarLens , a curated dataset of human-written and LLM-generated content across scholarly writing and peer review for multi-perspective evaluation, and LLMetrica , a tool for assessing LLM penetration using rule-based metrics and model-based detectors for multi-dimensional evaluation. The effectiveness of LLMetrica is demonstrated through experiments, revealing the increasing role of LLMs in scholarly processes. These findings emphasize the need for transparency, accountability, and ethical practices in the use of LLMs to maintain academic credibility.
Li Zhou 0010, Xunlian Dai, Daniel Hershcovich, Lihui Wang 0002
Eng. Appl. Artif. Intell.1
2025 Know You First and Be You Better: Modeling Human-Like User Simulators via Implicit Profiles
abstract
User simulators are crucial for replicating human interactions with dialogue systems, supporting both collaborative training and automatic evaluation, especially for large language models (LLMs). However, current role-playing methods face challenges such as a lack of utterance-level authenticity and user-level diversity, often hindered by role confusion and dependence on predefined profiles of well-known figures. In contrast, direct simulation focuses solely on text, neglecting implicit user traits like personality and conversation-level consistency. To address these issues, we introduce the User Simulator with Implicit Profiles (USP), a framework that infers implicit user profiles from human-machine interactions to simulate personalized and realistic dialogues. We first develop an LLM-driven extractor with a comprehensive profile schema, then refine the simulation using conditional supervised fine-tuning and reinforcement learning with cycle consistency, optimizing at both the utterance and conversation levels. Finally, a diverse profile sampler captures the distribution of real-world user profiles. Experimental results show that USP outperforms strong baselines in terms of authenticity and diversity while maintaining comparable consistency. Additionally, using USP to evaluate LLM on dynamic multi-turn aligns well with mainstream benchmarks, demonstrating its effectiveness in real-world applications.
Kuang Wang, Xianfei Li, Shenghao Yang 0001, Li Zhou 0010, Feng Jiang 0007, Haizhou Li 0001
ACL (1)4
2025 A Compressive Memory-based Retrieval Approach for Event Argument Extraction
abstract
Recent works have demonstrated the effectiveness of retrieval augmentation in the Event Argument Extraction (EAE) task. However, existing retrieval-based EAE methods have two main limitations: (1) input length constraints and (2) the gap between the retriever and the inference model. These issues limit the diversity and quality of the retrieved information. In this paper, we propose a Compressive Memory-based Retrieval (CMR) mechanism for EAE, which addresses the two limitations mentioned above. Our compressive memory, designed as a dynamic matrix that effectively caches retrieved information and supports continuous updates, overcomes the limitations of input length. Additionally, after pre-loading all candidate demonstrations into the compressive memory, the model further retrieves and filters relevant information from the memory based on the input query, bridging the gap between the retriever and the inference model. Extensive experiments show that our method achieves new state-of-the-art performance on three public datasets (RAMS, WikiEvents, ACE05), significantly outperforming existing retrieval-based EAE methods.
Wanlong Liu, Enqi Zhang, Shaohuan Cheng, Dingyi Zeng, Li Zhou 0010, Chen Zhang 0020, Malu Zhang, Wenyu Chen 0001
COLING5
2025 From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Test
abstract
The human-centered word association test (WAT) serves as a cognitive proxy, revealing sociocultural variations through culturally shared semantic expectations and implicit linguistic patterns shaped by lived experiences.We extend this test into an LLM-adaptive, freerelation task to assess the alignment of large language models (LLMs) with cross-cultural cognition.To address culture preference, we propose CultureSteer, an innovative approach that moves beyond superficial cultural prompting by embedding cultural-specific semantic associations directly within the model's internal representation space.Experiments show that current LLMs exhibit significant bias toward Western (notably American) schemas at the word association level.In contrast, our model substantially improves cross-cultural alignment, capturing diverse semantic associations.Further validation on culture-sensitive downstream tasks confirms its efficacy in fostering cognitive alignment across cultures.This work contributes a novel methodological paradigm for enhancing cultural awareness in LLMs, advancing the development of more inclusive language technologies.
Xunlian Dai, Li Zhou 0010, Benyou Wang, Haizhou Li 0001
EMNLP2
2025 RAG-Instruct: Boosting LLMs with Diverse Retrieval-Augmented Instructions
abstract
Retrieval-Augmented Generation (RAG) has emerged as a key paradigm for enhancing large language models by incorporating external knowledge.However, current RAG methods exhibit limited capabilities in complex RAG scenarios and suffer from limited task diversity.To address these limitations, we propose RAG-Instruct, a general method for synthesizing diverse and high-quality RAG instruction data based on any source corpus.Our approach leverages (1) five RAG paradigms, which encompass diverse query-document relationships, and (2) instruction simulation, which enhances instruction diversity and quality by utilizing the strengths of existing instruction datasets.Using this method, we construct a 40K instruction dataset from Wikipedia, comprehensively covering diverse RAG scenarios and tasks.Experiments demonstrate that RAG-Instruct effectively enhances LLMs' RAG capabilities, achieving strong zero-shot performance and outperforming various RAG baselines.The code is publicly available at https://github.com/FreedomIntelligence/RAG- Instruct.
Wanlong Liu, Ke Ji, Li Zhou 0010, Wenyu Chen 0001, Benyou Wang
EMNLP4
2025 Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation
abstract
Culture is a rich and dynamic domain that evolves across both geography and time.However, existing studies on cultural understanding with vision-language models (VLMs) primarily emphasize geographic diversity, often overlooking the critical temporal dimensions.To bridge this gap, we introduce Hanfu-Bench, a novel, expert-curated multimodal dataset.Hanfu, a traditional garment spanning ancient Chinese dynasties, serves as a representative cultural heritage that reflects the profound temporal aspects of Chinese culture while remaining highly popular in Chinese contemporary society.Hanfu-Bench comprises two core tasks: cultural visual understanding and cultural image transcreation.The former task examines temporal-cultural feature recognition based on single-or multi-image inputs through multiplechoice visual question answering, while the latter focuses on transforming traditional attire into modern designs through cultural element inheritance and modern context adaptation.Our evaluation shows that closed VLMs perform comparably to non-experts on visual cutural understanding but fall short by 10% to human experts, while open VLMs lags further behind non-experts.For the transcreation task, multi-faceted human evaluation indicates that the best-performing model achieves a success rate of only 42%.Our benchmark provides an essential testbed, revealing significant challenges in this new direction of temporal cultural understanding and creative adaptation.
Li Zhou 0010, Lutong Yu, Dongchu Xie, Shaohuan Cheng, Wenyan Li 0001, Haizhou Li 0001
EMNLP1
2025 Enhancing Document-Level Relation Extraction through Entity-Pair-Level Interaction Modeling
abstract
Document-level relation extraction aims at extracting relational facts between two entities in a document. Existing approaches mainly focus on target entities, utilizing techniques such as graph neural networks to enhance their representations. However, they ignore the rich semantic correlations among entity pairs which provide wider and multifaceted information at a higher level. In this paper, we propose the Relation-based Entity-pair-level Inference (REI) model, which facilitates information interaction at the entity-pair level, enhancing logical reasoning among entities and capturing semantic correlations among entity pairs. Our REI model comprises two modules: Relation-based Information Aggregation (RIA) and Entity-pair-level Information Interaction (EII). The RIA module builds and integrates relation representations to filter out distractions from unrelated entity pairs, while the EII module models entity-pair-level information interaction through multi-head attentions. Extensive experiments on the DocRED, DWIE, CDR, and GDA datasets demonstrate the superiority of the proposed REI model, outperforming previous state-of-the-art approaches. Furthermore, we provide detailed experimental analyses based on the performance gains and illustrate the interpretability.
Wanlong Liu, Dingyi Zeng, Li Zhou 0010, Yichen Xiao, Malu Zhang, Wenyu Chen 0001
ICASSP3
2025 Does Mapo Tofu Contain Coffee? Probing LLMs for Food-related Cultural Knowledge
abstract
Li Zhou, Taelin Karidi, Wanlong Liu, Nicolas Garneau, Yong Cao, Wenyu Chen, Haizhou Li, Daniel Hershcovich. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Li Zhou 0010, Taelin Karidi, Wanlong Liu, Nicolas Garneau, Yong Cao 0001, Wenyu Chen 0001, Haizhou Li 0001, Daniel Hershcovich
NAACL (Long Papers)1
2025 Beyond Binary: Towards Fine-Grained LLM-Generated Text Detection via Role Recognition and Involvement Measurement
abstract
The rapid development of large language models (LLMs), like ChatGPT, has resulted in the widespread presence of LLM-generated content on social media platforms, raising concerns about misinformation, data biases, and privacy violations, which can undermine trust in online discourse. While detecting LLM-generated content is crucial for mitigating these risks, current methods often focus on binary classification, failing to address the complexities of real-world scenarios like human-LLM collaboration. To move beyond binary classification and address these challenges, we propose a new paradigm for detecting LLM-generated content. This approach introduces two novel tasks: LLM Role Recognition (LLM-RR), a multi-class classification task that identifies specific roles of an LLM in content generation, and LLM Involvement Measurement (LLM-IM), a regression task that quantifies the extent of LLM involvement in content creation. To support these tasks, we propose LLMDetect, a benchmark designed to evaluate detectors' performance on these new tasks. LLMDetect includes the Hybrid News Detection Corpus (HNDC) for training detectors, as well as DetectEval, a comprehensive evaluation suite that considers five distinct cross-context variations and two multi-intensity variations within the same LLM role. This allows for a thorough assessment of detectors' generalization and robustness across diverse contexts. Our empirical validation of 10 baseline detection methods demonstrates that fine-tuned Pre-trained Language Model (PLM)-based models consistently outperform others on both tasks, while advanced LLMs face challenges in accurately detecting their own generated content. Our experimental results and analysis offer insights for developing more effective detection models for LLM-generated content. This research enhances the understanding of LLM-generated content and establishes a foundation for more nuanced detection methodologies.
Li Zhou 0010, Feng Jiang 0007, Benyou Wang, Haizhou Li 0001
WWW2
2025 Document-level relation extraction with structural encoding and entity-pair-level information interaction
Wanlong Liu, Yichen Xiao, Shaohuan Cheng, Dingyi Zeng, Li Zhou 0010, Weishan Kong, Malu Zhang, Wenyu Chen 0001
Expert Syst. Appl.5
2024 FoodieQA: A Multimodal Dataset for Fine-Grained Understanding of Chinese Food Culture
abstract
Wenyan Li, Crystina Zhang, Jiaang Li, Qiwei Peng, Raphael Tang, Li Zhou, Weijia Zhang, Guimin Hu, Yifei Yuan, Anders Søgaard, Daniel Hershcovich, Desmond Elliott. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Wenyan Li 0001, Xinyu Zhang 0018, Jiaang Li 0002, Qiwei Peng 0003, Raphael Tang, Li Zhou 0010, Weijia Zhang 0004, Guimin Hu, Yifei Yuan 0002, Anders Søgaard, Daniel Hershcovich, Desmond Elliott
EMNLP6
2024 MLPs Compass: What is Learned When MLPs are Combined with PLMs?
abstract
While Transformer-based pre-trained language models and their variants exhibit strong semantic representation capabilities, the question of comprehending the information gain derived from the additional components of PLMs remains an open question in this field. Motivated by recent efforts that prove Multilayer-Perceptrons (MLPs) modules achieving robust structural capture capabilities, even outperforming Graph Neural Networks (GNNs), this paper aims to quantify whether simple MLPs can further enhance the already potent ability of PLMs to capture linguistic information. Specifically, we design a simple yet effective probing framework containing MLPs components based on BERT structure and conduct extensive experiments encompassing 10 probing tasks spanning three distinct linguistic levels. The experimental results demonstrate that MLPs can indeed enhance the comprehension of linguistic structure by PLMs. Our research provides interpretable and valuable insights into crafting variations of PLMs utilizing MLPs for tasks that emphasize diverse linguistic structures.
Li Zhou 0010, Wenyu Chen 0001, Yong Cao 0001, Dingyi Zeng, Wanlong Liu, Hong Qu 0002
ICASSP1
2024 Dynamic training for handling textual label noise
Shaohuan Cheng, Wenyu Chen 0001, Wanlong Liu, Li Zhou 0010, Honglin Zhao, Weishan Kong, Hong Qu 0002, Mingsheng Fu
Appl. Intell.4
2024 Cultural Adaptation of Recipes
abstract
Abstract Building upon the considerable advances in Large Language Models (LLMs), we are now equipped to address more sophisticated tasks demanding a nuanced understanding of cross-cultural contexts. A key example is recipe adaptation, which goes beyond simple translation to include a grasp of ingredients, culinary techniques, and dietary preferences specific to a given culture. We introduce a new task involving the translation and cultural adaptation of recipes between Chinese- and English-speaking cuisines. To support this investigation, we present CulturalRecipes, a unique dataset composed of automatically paired recipes written in Mandarin Chinese and English. This dataset is further enriched with a human-written and curated test set. In this intricate task of cross-cultural recipe adaptation, we evaluate the performance of various methods, including GPT-4 and other LLMs, traditional machine translation, and information retrieval techniques. Our comprehensive analysis includes both automatic and human evaluation metrics. While GPT-4 exhibits impressive abilities in adapting Chinese recipes into English, it still lags behind human expertise when translating English recipes into Chinese. This underscores the multifaceted nature of cultural adaptations. We anticipate that these insights will significantly contribute to future research on culturally aware language models and their practical application in culturally diverse contexts.
Yong Cao 0001, Yova Kementchedjhieva, Ruixiang Cui, Antonia Karamolegkou, Li Zhou 0010, Megan Dare, Lucia Donatelli, Daniel Hershcovich
Trans. Assoc. Comput. Linguistics5
2024 CreoleVal: Multilingual Multitask Benchmarks for Creoles
abstract
Abstract Creoles represent an under-explored and marginalized group of languages, with few available resources for NLP research. While the genealogical ties between Creoles and a number of highly resourced languages imply a significant potential for transfer learning, this potential is hampered due to this lack of annotated data. In this work we present CreoleVal, a collection of benchmark datasets spanning 8 different NLP tasks, covering up to 28 Creole languages; it is an aggregate of novel development datasets for reading comprehension relation classification, and machine translation for Creoles, in addition to a practical gateway to a handful of preexisting benchmarks. For each benchmark, we conduct baseline experiments in a zero-shot setting in order to further ascertain the capabilities and limitations of transfer learning for Creoles. Ultimately, we see CreoleVal as an opportunity to empower research on Creoles in NLP and computational linguistics, and in general, a step towards more equitable language technology around the globe.
Heather C. Lent, Kushal Tatariya, Raj Dabre, Yiyi Chen 0002, Marcell Fekete, Esther Ploeger, Li Zhou 0010, Ruth-Ann Armstrong, Abee Eijansantos, Catriona Malau, Hans Erik Heje, Ernests Lavrinovics, Diptesh Kanojia, Paul Belony, Marcel Bollmann, Loïc Grobol, Miryam de Lhoneux, Daniel Hershcovich, Michel DeGraff, Anders Søgaard, Johannes Bjerva
Trans. Assoc. Comput. Linguistics7
2023 Substructure Aware Graph Neural Networks
abstract
Despite the great achievements of Graph Neural Networks (GNNs) in graph learning, conventional GNNs struggle to break through the upper limit of the expressiveness of first-order Weisfeiler-Leman graph isomorphism test algorithm (1-WL) due to the consistency of the propagation paradigm of GNNs with the 1-WL.Based on the fact that it is easier to distinguish the original graph through subgraphs, we propose a novel framework neural network framework called Substructure Aware Graph Neural Networks (SAGNN) to address these issues. We first propose a Cut subgraph which can be obtained from the original graph by continuously and selectively removing edges. Then we extend the random walk encoding paradigm to the return probability of the rooted node on the subgraph to capture the structural information and use it as a node feature to improve the expressiveness of GNNs. We theoretically prove that our framework is more powerful than 1-WL, and is superior in structure perception. Our extensive experiments demonstrate the effectiveness of our framework, achieving state-of-the-art performance on a variety of well-proven graph tasks, and GNNs equipped with our framework perform flawlessly even in 3-WL failed graphs. Specifically, our framework achieves a maximum performance improvement of 83% compared to the base models and 32% compared to the previous state-of-the-art methods.
Dingyi Zeng, Wanlong Liu, Wenyu Chen 0001, Li Zhou 0010, Malu Zhang, Hong Qu 0002
AAAI4
2023 Copyright Violations and Large Language Models
abstract
Language models may memorize more than just facts, including entire chunks of texts seen during training.Fair use exemptions to copyright laws typically allow for limited use of copyrighted material without permission from the copyright holder, but typically for extraction of information from copyrighted materials, rather than verbatim reproduction.This work explores the issue of copyright violations and large language models through the lens of verbatim memorization, focusing on possible redistribution of copyrighted text.We present experiments with a range of language models over a collection of popular books and coding problems, providing a conservative characterization of the extent to which language models can redistribute these materials.Overall, this research highlights the need for further examination and the potential impact on future developments in natural language processing to ensure adherence to copyright regulations.
Antonia Karamolegkou, Jiaang Li 0002, Li Zhou 0010, Anders Søgaard
EMNLP3
2023 Rethinking Random Walk in Graph Representation Learning
abstract
With the help of deep learning, Graph Neural Networks (GNNs) have achieved remarkable progress in various fields. However, due to the limitation of the message passing mechanism of GNNs, there exists an upper limit on its expressiveness. Some high-order GNNs have achieved good results in expressiveness, but they also have shortcomings in complexity and real-world performance. In this paper, we attempt to provide a graph neural network architecture that simultaneously addresses expressiveness, complexity and real-world performance. To this end, we propose Spatially constrained Random walk diffusion structural Encoding (SRE) to encode structural information and can be used for any GNN under our architecture. Our extensive and diverse experiments on datasets of different types and sizes demonstrate the superior expressiveness and state-of-the-art performance of our architecture on real-world tasks.
Dingyi Zeng, Wenyu Chen 0001, Wanlong Liu, Li Zhou 0010, Hong Qu 0002
ICASSP4
2023 DPGNN: Dual-perception graph neural network for representation learning
Li Zhou 0010, Wenyu Chen 0001, Dingyi Zeng, Shaohuan Cheng, Wanlong Liu, Malu Zhang, Hong Qu 0002
Knowl. Based Syst.1
2022 A Simple Graph Neural Network via Layer Sniffer
abstract
Due to the success of Graph Neural Networks(GNNs) in graph-structure data, many efforts have been devoted to enhancing the propagation ability and alleviating the over-smoothing problem of GNNs. However, from the perspective of closeness extent of node representations, most existing GNNs pay less attention to the attributes of node representation space. In light of this, we design a Layer Sniffer module that can combine the effects of the local node-level representation closeness extent and the global layer-level information attention. On this basis, we propose a simple Layer Sniffer Graph Neural Network (LSGNN) with a propagation scheme that can fuse neighborhood information of different receptive fields densely and adaptively. Our extensive experiments on three public node classification datasets demonstrate the superior performance and stability of our proposed model.
Dingyi Zeng, Li Zhou 0010, Wanlong Liu, Hong Qu 0002, Wenyu Chen 0001
ICASSP2
2022 Document-Level Relation Extraction with Structure Enhanced Transformer Encoder
abstract
Document-level relation extraction aims at discovering relational facts among entity pairs in a document, which has attracted more and more attention in recent years. Most existing methods are mainly summarized as graph-based and transformer-based methods. However, previous transformer-based methods neglect structural information between entities, while graph-based methods are unable to extract structural information effectively on account that they isolate the en-coding stage and structure reasoning stage. In this paper, we propose an effective structure enhanced transformer encoder model (SETE), integrating entity structural information into the transformer encoder. We first define a mention-level graph based on mention dependencies and convert it to a token-level graph. Then we design a dual self-attention mechanism, which enriches the structural and contextual information between entities to increase the vanilla transformer encoder inferential capability. Experiments on three public datasets show that the proposed SETE outperforms previous state-of-the-art methods and further analyses illustrate the interpretability of our model.
Wanlong Liu, Li Zhou 0010, Dingyi Zeng, Hong Qu 0002
IJCNN2
2020 A Weighted GCN with Logical Adjacency Matrix for Relation Extraction
abstract
Graph convolutional network (GCN), with its capability to update the current node features according to the features of its first-order adjacent nodes and edges, has achieved impressive performance in dependency capturing. But some important nodes from which we should figure out the dependencies are not first-order reachable, which calls for multi-layer GCNs for indirect relevance capturing. In this paper, we propose a novel weighted graph convolutional network by constructing a logical adjacency matrix which effectively solves the feature fusion of multi-hop relation without additional layers and parameters for relation extraction task. And we apply an Entity-Attention mechanism to enrich the entity pairs with more focused semantic information. Experimental results on TACRED and SemEval 2010 task 8 show that our model can take better advantage of the structural information in the dependency tree and produce better results than previous models.
Li Zhou 0010, Hong Qu 0002, Li Huang 0002, Yuguo Liu
ECAI1