Jiawei Chen 0011

dblp:03/1390-11 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0002-9759-9747ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 7 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 CoCoNUTS: Concentrating on Content while Neglecting Uninformative Textual Styles for AI-Generated Peer Review Detection
abstract
The growing use of large language models (LLMs) in peer review threatens scholarly integrity.Recent conference policies allow AI tools for language polishing but prohibit their use for generating substantive content.However, existing detectors mainly rely on stylistic cues, making it difficult to distinguish between surface-level language refinement and genuine content generation.To address this, we advocate a content-based detection paradigm and introduce CoCoNUTS, a comprehensive benchmark containing 315,535 reviews covering leading AI conferences and six human-AI collaboration modes.Our evaluation shows that current detectors struggle to handle these nuanced settings.Consequently, we propose CoCoDet, an AI review detector designed to identify substantive AI-generation.Experiments demonstrate that CoCoDet achieves a macro F1-score of 98.24%.Crucially, on permissible machinepolished reviews, it maintains a low false positive rate of 3.89%, substantially outperforming the strongest baseline (7.84%).Examination on real-world reviews using CoCoDet reveals an escalating trend of substantive AI generation.Our work exposes the inadequacy of current detectors, underscoring the importance of domainspecific solutions.Our code is available at https://github.com/icip-cas/CoCoNUTS.
Jiawei Chen 0011, Guozhao Mo, Xuanang Chen, Ben He 0001, Xianpei Han, Le Sun 0001
ACL (1)2
2025 ConsistentChat: Building Skeleton-Guided Consistent Multi-Turn Dialogues for Large Language Models from Scratch
abstract
Jiawei Chen, Xinyan Guan, Qianhao Yuan, Mo Guozhao, Weixiang Zhou, Yaojie Lu, Hongyu Lin, Ben He, Le Sun, Xianpei Han. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Jiawei Chen 0011, Xinyan Guan, Qianhao Yuan, Guozhao Mo, Weixiang Zhou, Yaojie Lu 0001, Ben He 0001, Le Sun 0001, Xianpei Han
EMNLP1
2025 ShortV: Efficient Multimodal Large Language Models by Freezing Visual Tokens in Ineffective Layers
abstract
Multimodal Large Language Models (MLLMs) suffer from high computational costs due to their massive size and the large number of visual tokens. In this paper, we investigate layer-wise redundancy in MLLMs by introducing a novel metric, Layer Contribution (LC), which quantifies the impact of a layer's transformations on visual and text tokens, respectively. The calculation of LC involves measuring the divergence in model output that results from removing the layer's transformations on the specified tokens. Our pilot experiment reveals that many layers of MLLMs exhibit minimal contribution during the processing of visual tokens. Motivated by this observation, we propose ShortV, a training-free method that leverages LC to identify ineffective layers, and freezes visual token updates in these layers. Experiments show that ShortV can freeze visual token in approximately 60\% of the MLLM layers, thereby dramatically reducing computational costs related to updating visual tokens. For example, it achieves a 50\% reduction in FLOPs on LLaVA-NeXT-13B while maintaining superior performance. The code will be publicly available at https://github.com/icip-cas/ShortV
Qianhao Yuan, Yanjiang Liu, Jiawei Chen 0011, Yaojie Lu 0001, Jia Zheng 0009, Xianpei Han, Le Sun 0001
ICCV4
2025 The Rise and Down of Babel Tower: Investigating the Evolution Process of Multilingual Code Large Language Model
abstract
Large language models (LLMs) have shown significant multilingual capabilities. However, the mechanisms underlying the development of these capabilities during pre-training are not well understood. In this paper, we use code LLMs as an experimental platform to explore the evolution of multilingual capabilities in LLMs during the pre-training process. Based on our observations, we propose the Babel Tower Hypothesis, which describes the entire process of LLMs acquiring new language capabilities. During the learning process, multiple languages initially share a single knowledge system dominated by the primary language and gradually develop language-specific knowledge systems. We then validate the above hypothesis by tracking the internal states of the LLM using specific methods. Experimental results show that the internal state changes of the LLM are consistent with our Babel Tower Hypothesis. Building on these insights, we propose a novel method to construct an optimized pre-training corpus for multilingual code LLMs, which significantly outperforms LLMs trained on the original corpus. The proposed Babel Tower Hypothesis provides new insights into designing pre-training data distributions to achieve optimal multilingual capabilities in LLMs.
Jiawei Chen 0011, Mengjie Ren, Yaojie Lu 0001, Xianpei Han, Le Sun 0001
ICLR1
2024 Benchmarking Large Language Models in Retrieval-Augmented Generation
abstract
Retrieval-Augmented Generation (RAG) is a promising approach for mitigating the hallucination of large language models (LLMs). However, existing research lacks rigorous evaluation of the impact of retrieval-augmented generation on different large language models, which make it challenging to identify the potential bottlenecks in the capabilities of RAG for different LLMs. In this paper, we systematically investigate the impact of Retrieval-Augmented Generation on large language models. We analyze the performance of different large language models in 4 fundamental abilities required for RAG, including noise robustness, negative rejection, information integration, and counterfactual robustness. To this end, we establish Retrieval-Augmented Generation Benchmark (RGB), a new corpus for RAG evaluation in both English and Chinese. RGB divides the instances within the benchmark into 4 separate testbeds based on the aforementioned fundamental abilities required to resolve the case. Then we evaluate 6 representative LLMs on RGB to diagnose the challenges of current LLMs when applying RAG. Evaluation reveals that while LLMs exhibit a certain degree of noise robustness, they still struggle significantly in terms of negative rejection, information integration, and dealing with false information. The aforementioned assessment outcomes indicate that there is still a considerable journey ahead to effectively apply RAG to LLMs.
Jiawei Chen 0011, Xianpei Han, Le Sun 0001
AAAI1
2024 Few-shot Named Entity Recognition via Superposition Concept Discrimination
abstract
Few-shot NER aims to identify entities of target types with only limited number of illustrative instances. Unfortunately, few-shot NER is severely challenged by the intrinsic precise generalization problem, i.e., it is hard to accurately determine the desired target type due to the ambiguity stemming from information deficiency. In this paper, we propose Superposition Concept Discriminator (SuperCD), which resolves the above challenge via an active learning paradigm. Specifically, a concept extractor is first introduced to identify superposition concepts from illustrative instances, with each concept corresponding to a possible generalization boundary. Then a superposition instance retriever is applied to retrieve corresponding instances of these superposition concepts from large-scale text corpus. Finally, annotators are asked to annotate the retrieved instances and these annotated instances together with original illustrative instances are used to learn FS-NER models. To this end, we learn a universal concept extractor and superposition instance retriever using a large-scale openly available knowledge bases. Experiments show that SuperCD can effectively identify superposition concepts from illustrative instances, retrieve superposition instances from large-scale corpus, and significantly improve the few-shot NER performance with minimal additional efforts.
Jiawei Chen 0011, Xianpei Han, Yaojie Lu 0001, Shanshan Jiang 0001, Bin Dong 0003, Le Sun 0001
LREC/COLING1
2024 Retentive or Forgetful? Diving into the Knowledge Memorizing Mechanism of Language Models
abstract
Memory is one of the most essential cognitive functions serving as a repository of world knowledge and episodes of activities. In recent years, large-scale pre-trained language models have shown remarkable memorizing ability. On the contrary, vanilla neural networks without pre-training have been long observed suffering from the catastrophic forgetting problem. To investigate such a retentive-forgetful contradiction and understand the memorizing dynamic mechanism of language models, we conduct thorough experiments by controlling the target knowledge types, the learning strategies and the learning schedules. We find that: 1) Vanilla language models without pre-training are forgetful; 2) Pre-training leads to retentive language models; 3) Knowledge relevance and diversification significantly influence the memory formation. These conclusions are useful for understanding the abilities of pre-trained language models and shed light on designing and evaluating new learning and inference algorithms of language models.
Boxi Cao, Qiaoyu Tang, Shanshan Jiang 0001, Bin Dong 0003, Xianpei Han, Jiawei Chen 0011, Tianshu Wang 0002, Le Sun 0001
LREC/COLING7
2024 Self-Retrieval: End-to-End Information Retrieval with One Large Language Model
abstract
The rise of large language models (LLMs) has significantly transformed both the construction and application of information retrieval (IR) systems. However, current interactions between IR systems and LLMs remain limited, with LLMs merely serving as part of components within IR systems, and IR systems being constructed independently of LLMs. This separated architecture restricts knowledge sharing and deep collaboration between them. In this paper, we introduce Self-Retrieval, a novel end-to-end LLM-driven information retrieval architecture. Self-Retrieval unifies all essential IR functions within a single LLM, leveraging the inherent capabilities of LLMs throughout the IR process. Specifically, Self-Retrieval internalizes the retrieval corpus through self-supervised learning, transforms the retrieval process into sequential passage generation, and performs relevance assessment for reranking. Experimental results demonstrate that Self-Retrieval not only outperforms existing retrieval approaches by a significant margin, but also substantially enhances the performance of LLM-driven downstream applications like retrieval-augmented generation.
Qiaoyu Tang, Jiawei Chen 0011, Bowen Yu 0002, Yaojie Lu 0001, Cheng Fu 0003, Haiyang Yu 0003, Fei Huang 0002, Ben He 0001, Xianpei Han, Le Sun 0001, Yongbin Li 0001
NeurIPS2
2023 Learning In-context Learning for Named Entity Recognition
abstract
Jiawei Chen, Yaojie Lu, Hongyu Lin, Jie Lou, Wei Jia, Dai Dai, Hua Wu, Boxi Cao, Xianpei Han, Le Sun. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Jiawei Chen 0011, Yaojie Lu 0001, Jie Lou, Dai Dai, Hua Wu 0003, Boxi Cao, Xianpei Han, Le Sun 0001
ACL (1)1
2022 Few-shot Named Entity Recognition with Self-describing Networks
abstract
Few-shot NER needs to effectively capture information from limited instances and transfer useful knowledge from external resources.In this paper, we propose a self-describing mechanism for few-shot NER, which can effectively leverage illustrative instances and precisely transfer knowledge from external resources by describing both entity types and mentions using a universal concept set.Specifically, we design Self-describing Networks (SDNet), a Seq2Seq generation model which can universally describe mentions using concepts, automatically map novel entity types to concepts, and adaptively recognize entities ondemand.We pre-train SDNet with large-scale corpus, and conduct experiments on 8 benchmarks from different domains.Experiments show that SDNet achieves competitive performances on all benchmarks and achieves the new state-of-the-art on 6 benchmarks, which demonstrates its effectiveness and robustness.
Jiawei Chen 0011, Xianpei Han, Le Sun 0001
ACL (1)1
2021 Honey or Poison? Solving the Trigger Curse in Few-shot Event Detection via Causal Intervention
abstract
Event detection has long been troubled by the trigger curse: overfitting the trigger will harm the generalization ability while underfitting it will hurt the detection performance.This problem is even more severe in few-shot scenario.In this paper, we identify and solve the trigger curse problem in few-shot event detection (FSED) from a causal view.By formulating FSED with a structural causal model (SCM), we found that the trigger is a confounder of the context and the result, which makes previous FSED methods much easier to overfit triggers.To resolve this problem, we propose to intervene on the context via backdoor adjustment during training.Experiments show that our method significantly improves the FSED on ACE05, MAVEN and KBP17 datasets.
Jiawei Chen 0011, Xianpei Han, Le Sun 0001
EMNLP (1)1