Liwen Peng

dblp:240/9462 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0003-1202-9583ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 A unified multimodal classification framework based on deep metric learning
Liwen Peng, Songlei Jian, Minne Li, Zhigang Kan, Linbo Qiao, Dongsheng Li 0001
Neural Networks1
2024 Emancipating Event Extraction from the Constraints of Long-Tailed Distribution Data Utilizing Large Language Models
abstract
Event Extraction (EE) is a challenging task that aims to extract structural event-related information from unstructured text. Traditional methods for EE depend on manual annotations, which are both expensive and scarce. Furthermore, the existing datasets mostly follow the long-tail distribution, severely hindering the previous methods of modeling tail types. Two techniques can address this issue: transfer learning and data generation. However, the existing methods based on transfer learning still rely on pre-training with a large amount of labeled data in the source domain. Additionally, the quality of data generated by previous data generation methods is difficult to control. In this paper, leveraging Large Language Models (LLMs), we propose novel methods for event extraction and generation based on dialogues, overcoming the problems of relying on source domain data and maintaining data quality. Specifically, this paper innovatively transforms the EE task into multi-turn dialogues, guiding LLMs to learn event schemas from historical dialogue information and output structural events. Furthermore, we introduce a novel LLM-based method for generating high-quality data, significantly improving traditional models’ performance with various paradigms and structures, especially on tail types. Adequate experiments on real-world datasets demonstrate the effectiveness of the proposed event extraction and data generation methods.
Zhigang Kan, Liwen Peng, Linbo Qiao, Dongsheng Li 0001
LREC/COLING2
2024 LFDe: A Lighter, Faster and More Data-Efficient Pre-training Framework for Event Extraction
abstract
Pre-training Event Extraction (EE) models on unlabeled data is an effective strategy that frees researchers from costly and labor-intensive data annotation. However, existing pre-training methods necessitate substantial computational resources, requiring high-performance hardware infrastructure and extensive training duration. In response to these challenges, this paper proposes a Lighter, Faster, and more Data-efficient pre-training framework for EE, named LFDe. Distinct from existing methods that strive to establish a comprehensive representation space during pre-training, our framework focuses on quickly familiarizing with the task format from a small amount of automatically constructed pseudo-events. It comprises three stages: weak-label data construction, pre-training, and fine-tuning. Specifically, during the first stage, LFDe first automatically designates pseudo-triggers and arguments based on the characteristics of real events to form pre-training samples. In the processes of pre-training and fine-tuning, the framework reframes EE as the identification of tokens semantically closest to the prompt within the given sentence. This paper also introduces a novel prompt-based sequence labeling model for EE to accommodate this reframing. Experiments on real-world datasets show that compared to similar models, our framework requires fewer pre-training data (only about 0.04%), a shorter pre-training period (about 0.03%), and lower memory requirements (about 57.6%). Simultaneously, our framework significantly improves performance in various data-scarce scenarios.
Zhigang Kan, Liwen Peng, Yifu Gao, Ning Liu 0015, Linbo Qiao, Dongsheng Li 0001
WWW2
2024 Not all fake news is semantically similar: Contextual semantic representation learning for multimodal fake news detection
Liwen Peng, Songlei Jian, Zhigang Kan, Linbo Qiao, Dongsheng Li 0001
Inf. Process. Manag.1
2023 MRML: Multimodal Rumor Detection by Deep Metric Learning
abstract
Multimodal rumor detection aims at detecting rumors using information from textual and visual modalities. The most critical difficulty in multimodal rumor detection lies in capturing both the intra-modal and inter-modal relationships from multimodal data. However, existing methods mainly focus on the multimodal fusion process while paying little attention to the intra-modal relationships. To address these limitations, we propose a multimodal rumor detection method with deep metric learning (MRML) to effectively extract multimodal relationships of news for detecting rumors. Specifically, we design the metric-based triplet learning to extract the intra-modal relationships between rumors and non-rumors in every modality and the contrastive pairwise learning to capture the inter-modal relationships across multimodal. Extensive experiments on two real-world multimodal datasets show the superior performance of our rumor detection method.
Liwen Peng, Songlei Jian, Dongsheng Li 0001
ICASSP1
2023 An anchor-guided sequence labeling model for event detection in both data-abundant and data-scarce scenarios
Zhigang Kan, Yanqi Shi, Zhangyue Yin, Liwen Peng, Linbo Qiao, Xipeng Qiu, Dongsheng Li 0001
Inf. Sci.4
2019 Author Disambiguation through Adversarial Network Representation Learning
abstract
Many persons share with the same name. Distinguishing different persons with the same name is important but challenging. Albeit much work has been proposed for author disambiguation, most of them do not adequately consider the heterogeneous relationships among authors and papers. In our work, ambiguous names and their related information, such as papers, conferences, titles, abstracts, etc., are constructed into a heterogeneous network which consists of different edge types. To fully incorporate all the information of the constructed network, we use Generative Adversarial Networks (GAN) to learn the network representation of the heterogeneous network. Although GAN has been used in many fields such as image generation, it hasn't been used to obtain representations for the heterogeneous network. As far as we know, our work is the first work which use adversarial training to learn heterogeneous network representation. After the representations are learned, they are partitioned into different groups each representing distinct authors. After extensive experiments on three major author disambiguation datasets, we demonstrate that our method outperforms several state-of-the-art baselines in author disambiguation problem.
Liwen Peng, Dongsheng Li 0001, Yongquan Fu, Huayou Su
IJCNN1