VLDB 2026 Research / reviewers in the wild / expert
Shuaipeng Liu
dblp:289/1075
· DBLP profile ↗
7ranked-venue papers
0as first author
7since 2021 · last 2025
0009-0009-3958-9020ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | OmniNER2025: Diverse and Comprehensive Fine-Grained NER Dataset and Benchmark for ChineseabstractAs Named Entity Recognition (NER) tasks have evolved, artificial intelligence has been widely applied in this field. However, most benchmarks are limited to English, making it challenging to replicate successful experiences in other languages. To expand NER to informal and diverse Chinese text scenarios, we have proposed a new large-scale Chinese NER dataset, OmniNER2025. This dataset, obtained from user posts on a popular Chinese social media platform Xiaohongshu, contains 195,568 samples and 89 categories, all manually annotated. To our knowledge, it is currently the largest Chinese open-source NER dataset in terms of sample size, category diversity, and domain coverage. This dataset is more challenging than existing Chinese NER datasets and better reflects real-world applications. The large sample size and diverse entity types provide valuable research resources. Additionally, we introduced the ERRTA tool for error analysis and teacher model guidance, significantly reducing model errors and improving performance. In the future, we will refine the ERRTA framework and explore optimization strategies to enhance the practical value of NER models. By releasing the OmniNER2025 dataset and introducing the ERRTA tool, we have advanced fine-grained NER research and improved model performance, promoting its application and development in real-world scenarios. Shuaipeng Liu, Mengting Hu 0002, Wen Dai, Xiaowei Zhao 0003, Xiujuan Xu |
SIGIR | 2 |
| 2024 | Type-Specific Modality Alignment for Multi-Modal Information ExtractionabstractMulti-modal information extraction aims to identify structured information, such as entities or relations between entities, from text with the help of visual clues. Although existing studies have achieved great progress, they mainly focused on modality interactions in the global space while neglecting fine-grained modality alignment under the semantic subspace specific to each entity type or relation type. To solve this problem, we propose a multi-space modality alignment method (MSMA) in this letter. The core of our model is a typespecific modality interaction module (TMI), which constructs a unique semantic subspace for each entity/relation type and independently performs type-specific modality alignments under each subspace. To enable mutual promotion between different types, a global modality integration module (GMI) is designed to learn the associations between different subspaces. Furthermore, we execute these two modules iteratively for high-level semantic fusion. Extensive experiments on three benchmark datasets show that our model significantly outperforms advanced methods. Shaowei Chen, Shuaipeng Liu, Jie Liu 0007 |
IEEE Signal Process. Lett. | 2 |
| 2023 | Exploiting Pseudo Future Contexts for Emotion Recognition in Conversations
Yinyi Wei, Shuaipeng Liu, Hailei Yan, Wei Ye 0004, Tong Mo, Guanglu Wan |
ADMA (1) | 2 |
| 2023 | ECOD: A Multi-modal Dataset for Intelligent Adjudication of E-Commerce Order Disputes
Liyi Chen 0003, Shuaipeng Liu, Hailei Yan, Jie Liu 0007, Lijie Wen 0001, Guanglu Wan |
NLPCC (1) | 2 |
| 2022 | DESED: Dialogue-based Explanation for Sentence-level Event DetectionabstractMany recent sentence-level event detection efforts focus on enriching sentence semantics, e.g., via multi-task or prompt-based learning. Despite the promising performance, these methods commonly depend on label-extensive manual annotations or require domain expertise to design sophisticated templates and rules. This paper proposes a new paradigm, named dialogue-based explanation, to enhance sentence semantics for event detection. By saying dialogue-based explanation of an event, we mean explaining it through a consistent information-intensive dialogue, with the original event description as the start utterance. We propose three simple dialogue generation methods, whose outputs are then fed into a hybrid attention mechanism to characterize the complementary event semantics. Extensive experimental results on two event detection datasets verify the effectiveness of our method and suggest promising research opportunities in the dialogue-based explanation paradigm. Yinyi Wei, Shuaipeng Liu, Jianwei Lv, Xiangyu Xi, Hailei Yan, Wei Ye 0004, Tong Mo, Fan Yang 0087, Guanglu Wan |
COLING | 2 |
| 2022 | MUSIED: A Benchmark for Event Detection from Multi-Source Heterogeneous Informal TextsabstractEvent detection (ED) identifies and classifies event triggers from unstructured texts, serving as a fundamental task for information extraction.Despite the remarkable progress achieved in the past several years, most research efforts focus on detecting events from formal texts (e.g., news articles, Wikipedia documents, financial announcements).Moreover, the texts in each dataset are either from a single source or multiple yet relatively homogeneous sources.With massive amounts of user-generated text accumulating on the Web and inside enterprises, identifying meaningful events in these informal texts, usually from multiple heterogeneous sources, has become a problem of significant practical value.As a pioneering exploration that expands event detection to the scenarios involving informal and heterogeneous texts, we propose a new large-scale Chinese event detection dataset based on user reviews, text conversations, and phone conversations in a leading e-commerce platform for food service.We carefully investigate the proposed dataset's textual informality and multi-source heterogeneity characteristics by inspecting data samples quantitatively and qualitatively.Extensive experiments with state-of-the-art event detection methods verify the unique challenges posed by these characteristics, indicating that multisource informal event detection remains an open problem and requires further efforts.Our benchmark and code are released at https: //github.com/myeclipse/MUSIED. Xiangyu Xi, Jianwei Lv, Shuaipeng Liu, Wei Ye 0004, Fan Yang 0087, Guanglu Wan |
EMNLP | 3 |
| 2022 | A Low-Cost, Controllable and Interpretable Task-Oriented Chatbot: With Real-World After-Sale Services as ExampleabstractThough widely used in industry, traditional task-oriented dialogue systems suffer from three bottlenecks: (i) difficult ontology construction (e.g., intents and slots); (ii) poor controllability and interpretability; (iii) annotation-hungry. In this paper, we propose to represent utterance with a simpler concept named Dialogue Action, upon which we construct a tree-structured TaskFlow and further build task-oriented chatbot with TaskFlow as core component. A framework is presented to automatically construct TaskFlow from large-scale dialogues and deploy online. Our experiments on real-world after-sale customer services show TaskFlow can satisfy the major needs, as well as reduce the developer burden effectively. Xiangyu Xi, Chenxu Lv, Yuncheng Hua, Wei Ye 0004, Chaobo Sun, Shuaipeng Liu, Fan Yang 0087, Guanglu Wan |
SIGIR | 6 |