VLDB 2026 Research / reviewers in the wild / expert
Shaoru Guo
dblp:190/7914
· DBLP profile ↗
13ranked-venue papers
5as first author
11since 2021 · last 2027
0000-0003-4130-3924ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Learning to encode and activate internal knowledge for retrieval and generation
Shaoru Guo |
Inf. Process. Manag. | 2 |
| 2026 | Debate to Extract: Enhancing Event Relation Extraction through Collaborative Debate and Agent OptimizationabstractEvent Relation Extraction (ERE) plays a crucial role in understanding document structures by identifying relations between events. However, most existing methods either rely on single-model instances, which often suffer from overconfidence, or adopt multi-agent frameworks that rely on manually designed prompts and heuristics to define agents, making effective optimization difficult. In this article, we propose Debate to Extract (D2E), a novel multi-agent optimization framework for ERE that leverages structured multi-turn debates and specialized agent training to enhance performance. Specifically, to organize the debate, participants are divided into multiple groups, each assigned its own debate topic. This process effectively integrates both cooperation and confrontation. We also incorporate an audience as a crucial participant, whose conclusions, from an observer’s perspective, tend to be more objective. Building on this debate framework, D2E further optimizes the debate participants, combining structured multi-turn debates with agent training. During the debate, agents refine their initial opinions through collaborative interactions. This iterative process generates valuable supervision signals for training, allowing agents to improve their responses progressively. To address the issue of diminishing returns from data diversity, each agent is trained on distinct subsets of generated data, promoting specialization across different task dimensions. Experimental results on the MAVEN-ERE and EventStoryLine datasets show that D2E achieves significant improvements in causal relation extraction, outperforming baseline methods by 12.1% and 4.69%, respectively. Through further analysis of the debate participants’ performance before and after the debate, we find that participation in the debate generally leads to improved ERE performance. This work demonstrates that combining collaborative debate with agent specialization leads to substantial performance gains in event relation extraction tasks. Shaoru Guo, Lei Hou 0001, Juan-Zi Li |
ACM Trans. Inf. Syst. | 2 |
| 2025 | Structure-to-word dynamic interaction model for abstractive sentence summarization
Shaoru Guo, Ru Li 0001 |
Neural Comput. Appl. | 2 |
| 2024 | NutFrame: Frame-based Conceptual Structure Induction with LLMsabstractConceptual structure is fundamental to human cognition and natural language understanding. It is significant to explore whether Large Language Models (LLMs) understand such knowledge. Since FrameNet serves as a well-defined conceptual structure knowledge resource, with meaningful frames, fine-grained frame elements, and rich frame relations, we construct a benchmark for coNceptual structure induction based on FrameNet, called NutFrame. It contains three sub-tasks: Frame Induction, Frame Element Induction, and Frame Relation Induction. In addition, we utilize prompts to induce conceptual structure of Framenet with LLMs. Furthermore, we conduct extensive experiments on NutFrame to evaluate various widely-used LLMs. Experimental results demonstrate that FrameNet induction remains a challenge for LLMs. Shaoru Guo, Yubo Chen 0001, Kang Liu 0001, Ru Li 0001, Jun Zhao 0001 |
LREC/COLING | 1 |
| 2024 | Improving Implicit Discourse Relation Recognition via Connective Prediction and Dependency-weighted Label HierarchyabstractImplicit discourse relation recognition aims to identify logical relations between two arguments without explicit connectives and is a challenging task in discourse analysis. Recent methods tend to leverage the label hierarchy to enhance discourse relation representations. However, they fail to fully utilize the connective information. Specifically, the methods overlook the guiding role of connectives in discourse relation classification by treating them as the last-level labels in the label hierarchy to leverage connective information, whereas it would be more appropriate to exploit connective information prior to relation classification. Moreover, these methods ignore the dependency degree of labels between different levels in the label hierarchy. In other words, they consider the label hierarchy as an unweighted undirected graph, and assume that the path weights between high-level labels and their corresponding low-level labels are the same, which leads to an insufficient construction of the label hierarchy. To overcome these issues, we propose a method for implicit discourse relation recognition (IDRR) utilizing Connective Prediction and Dependency-weighted Label Hierarchy (CP-DLH). Experimental results on PDTB 2.0 dataset show that our model achieves the state-of-the-art performance at all hierarchical levels. Xianzhi Liu, Shaoru Guo, Juncai Li, Zhichao Yan 0002, Xuefeng Su, Boxiang Ma, Yuzhi Wang, Ru Li 0001 |
IJCNN | 2 |
| 2024 | Multi-Granularity Dual-Aware Contrastive Learning for Few-shot Named Entity RecognitionabstractFew-shot Named Entity Recognition aims to identify named entities from unstructured texts in various domains using a minimal amount of training samples and classify them into predefined categories. Many popular approaches decompose this process into two tasks: Span Detection and Entity Classification. However, they still have some issues: (1) Neglecting presentation optimization. Most of these methods emphasize classification, neglecting the optimization of span presentation during Span Detection. (2) Missing label semantics. They have not fully leveraged semantic information of entity type labels during Entity Classification. To address these issues, this paper proposes Multi-Granularity Dual-Aware Contrastive Learning (MGDAC) for few-shot NER. Specifically, we introduce multi-granularity contrastive learning for solve neglecting presentation optimization, focusing on both the overall vector and internal vector granularity to enhance features beneficial for span detection in token vector representations. Additionally, we design dual-aware contrastive learning for solve missing label semantics, effectively utilizing semantic information from both entity tokens and entity type labels during prototype construction to jointly optimize prototype representations. Finally, extensive experiments on the FewNERD dataset demonstrate that our proposed method exhibits improvements in both Span Detection and Entity Classification, outperforming other competitive baseline methods. Boxiang Ma, Changzheng Wang, Shaoru Guo, Xuefeng Su, Zhichao Yan 0002, Wenyuan Shao, Zezheng Zhang, Ru Li 0001 |
IJCNN | 3 |
| 2021 | Frame Semantic-Enhanced Sentence Modeling for Sentence-level Extractive Text SummarizationabstractSentence-level extractive text summarization aims to select important sentences from a given document. However, it is very challenging to model the importance of sentences. In this paper, we propose a novel Frame Semantic-Enhanced Sentence Modeling for Extractive Summarization, which leverages Frame semantics to model sentences from both intra-sentence level and inter-sentence level, facilitating the text summarization task. In particular, intra-sentence level semantics leverage Frames and Frame Elements to model internal semantic structure within a sentence, while inter-sentence level semantics leverage Frame-to-Frame relations to model relationships among sentences. Extensive experiments on two benchmark corpus CNN/DM and NYT demonstrate that our model outperforms six state-of-the-art methods significantly. Shaoru Guo, Ru Li 0001, Xiaoli Li 0001, Hongye Tan |
EMNLP (1) | 2 |
| 2021 | Integrating Semantic Scenario and Word Relations for Abstractive Sentence SummarizationabstractRecently graph-based methods have been adopted for Abstractive Text Summarization. However, existing graph-based methods only consider either word relations or structure information, which neglect the correlation between them. To simultaneously capture the word relations and structure information from sentences, we propose a novel Dual Graph network for Abstractive Sentence Summarization. Specifically, we first construct semantic scenario graph and semantic word relation graph based on FrameNet, and subsequently learn their representations and design graph fusion method to enhance their correlation and obtain better semantic representation for summary generation. Experimental results show our model outperforms existing state-of-the-art methods on two popular benchmark datasets, i.e., Gigaword and DUC 2004. Shaoru Guo, Ru Li 0001, Xiaoli Li 0001, Hu Zhang 0003 |
EMNLP (1) | 2 |
| 2021 | Frame Semantics guided network for Abstractive Sentence Summarization
Shaoru Guo, Ru Li 0001, Xiaoli Li 0001, Hu Zhang 0003 |
Knowl. Based Syst. | 2 |
| 2021 | Frame-based Multi-level Semantics Representation for text matching
Shaoru Guo, Ru Li 0001, Xiaoli Li 0001, Hongye Tan |
Knowl. Based Syst. | 1 |
| 2021 | Frame-based Neural Network for Machine Reading Comprehension
Shaoru Guo, Hongye Tan, Ru Li 0001, Xiaoli Li 0001 |
Knowl. Based Syst. | 1 |
| 2020 | A Frame-based Sentence Representation for Machine Reading ComprehensionabstractSentence representation (SR) is the most crucial and challenging task in Machine Reading Comprehension (MRC).MRC systems typically only utilize the information contained in the sentence itself, while human beings can leverage their semantic knowledge.To bridge the gap, we proposed a novel Frame-based Sentence Representation (FSR) method, which employs frame semantic knowledge to facilitate sentence modelling.Specifically, different from existing methods that only model lexical units (LUs), Frame Representation Models, which utilize both LUs in frame and Frame-to-Frame (F-to-F) relations, are designed to model frames and sentences with attention schema.Our proposed FSR method is able to integrate multiple-frame semantic information to get much better sentence representations.Our extensive experimental results show that it performs better than state-of-the-art technologies on machine reading comprehension task. Shaoru Guo, Ru Li 0001, Hongye Tan, Xiaoli Li 0001, Yueping Zhang |
ACL | 1 |
| 2020 | Incorporating Syntax and Frame Semantics in Neural Network for Machine Reading ComprehensionabstractMachine reading comprehension (MRC) is one of the most critical yet challenging tasks in natural language understanding(NLU), where both syntax and semantics information of text are essential components for text understanding.It is surprising that jointly considering syntax and semantics in neural networks was never formally reported in literature.This paper makes the first attempt by proposing a novel Syntax and Frame Semantics model for Machine Reading Comprehension (SS-MRC), which takes full advantage of syntax and frame semantics to get richer text representation.Our extensive experimental results demonstrate that SS-MRC performs better than ten state-of-the-art technologies on machine reading comprehension task. Shaoru Guo, Ru Li 0001, Xiaoli Li 0001, Hongye Tan |
COLING | 1 |