VLDB 2026 Research / reviewers in the wild / expert
Shaojuan Wu
dblp:281/9022
· DBLP profile ↗
14ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0002-7407-485XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Existence to Exhaustiveness: Unveiling the Compounding Failures of LLMs in Multi-answer Event Temporal ReasoningabstractLarge Language Models (LLMs) have achieved remarkable success in temporal reasoning. However, existing benchmarks predominantly adopt a "single-answer" paradigm, focusing on verifying the existence of a specific fact while overlooking the challenge of exhaustiveness. In real-world scenarios, entities often simultaneously play multiple roles or exist in multiple states within the same timeframe. To bridge this gap, we introduce MulTR, a comprehensive benchmark designed for multi-answer temporal reasoning from long unstructured contexts. Specifically, MulTR integrates structured temporal facts from Wikidata and natural language text from Wikipedia MulTR integrates structured temporal facts from Wikidata and natural language text from Wikipedia through a logic-driven synthesis process. Notably, we formulate two distinct settings, question-dependent and document-dependent, based on the presence of cue words in the question. It is designed to systematically decouple temporal reasoning capabilities from the uncertainty of the number of answers. Experiment results demonstrate that state-of-the-art models suffer from retrieval laziness, terminating the search process prematurely after locating the first valid piece of evidence. Consequently, their performance drops sharply when evaluated on strict exact match metrics. MulTR, as a diagnostic testing platform, reveal these defects and establish the rigorous standard for future research in dynamic knowledge processing. The MulTR benchmark and evaluation prompt are publicly available at https://github.com/TemporalNLP/MulTR. Shaojuan Wu |
SIGIR | 1 |
| 2025 | Temporal-based graph reasoning for Visual Commonsense Reasoning
Shaojuan Wu, Jitong Li, Peng Chen 0053, Xiaowang Zhang, Zhiyong Feng 0002 |
Knowl. Based Syst. | 1 |
| 2024 | An Event-based Abductive Learning for Hard Time-sensitive Question AnsweringabstractTime-Sensitive Question Answering (TSQA) is to answer questions qualified for a certain timestamp based on the given document. It is split into easy and hard modes depending on whether the document contain time qualifiers mentioned in the question. While existing models have performed well on easy mode, their performance is significant reduced for answering hard time-sensitive questions, whose time qualifiers are implicit in the document. An intuitive idea is to match temporal events in the given document by treating time-sensitive question as a temporal event of missing objects. However, not all temporal events extracted from the document have explicit time qualifiers. In this paper, we propose an Event-AL framework, in which a graph pruning model is designed to locate the timespan of implicit temporal events by capturing temporal relation between events. Moreover, we present an abductive reasoning module to determine proper objects while providing explanations. Besides, as the same relation may be scattered throughout the document in diverse expressions, a relation-based prompt is introduced to instructs LLMs in extracting candidate temporal events. We conduct extensive experiment and results show that Event-AL outperforms strong baselines for hard time-sensitive questions, with a 12.7% improvement in EM scores. In addition, it also exhibits great superiority for multi-answer and beyond hard time-sensitive questions. Shaojuan Wu, Jitong Li, Xiaowang Zhang, Zhiyong Feng 0002 |
LREC/COLING | 1 |
| 2024 | Diversity-Enhanced Learning for Unsupervised Syntactically Controlled Paraphrase GenerationabstractSyntactically controlled paraphrase generation is to generate diverse sentences that have the same semantics as the given original sentence but conform to the target syntactic structure. An optimal opportunity to enhance diversity is to make word substitutions during rephrasing based on syntactic control. Existing unsupervised methods have made great progress in syntactic control, but the generated paraphrases rarely have substitutions due to the limitation of training data. In this paper, we propose a Diversity syntactically controlled Paraphrase generation framework (DiPara), in which a novel training strategy is designed to obtain semantic sentences while using the given sentence as training objects. As diverse words vary the syntactic structure around them, we propose a phrase-aware attention mechanism to capture the syntactic structure associated with the current word. To achieve it, the linearized triple sequence is introduced to represent structure singly. Experiment results on two datasets show that DiPara outperforms strong baselines, especially diversity (Self-BLEU4) is improved by 10.18% in ParaNMT-Small. Shaojuan Wu, Jitong Li, Xiaowang Zhang, Zhiyong Feng 0002 |
ECAI | 1 |
| 2023 | Multi-relation Identification for Few-Shot Document-Level Relation Extraction
Dazhuang Wang, Shaojuan Wu, Xiaowang Zhang, Zhiyong Feng 0002 |
ICANN (9) | 2 |
| 2023 | LenANet: A Length-Controllable Attention Network for Source Code Summarization
Peng Chen 0053, Shaojuan Wu, Jiarui Zhang 0005, Xiaowang Zhang, Zhiyong Feng 0002 |
ICONIP (11) | 2 |
| 2023 | Phrase-level attention network for few-shot inverse relation classification in knowledge graph
Shaojuan Wu, Chunliu Dou, Dazhuang Wang, Jitong Li, Xiaowang Zhang, Zhiyong Feng 0002, Kewen Wang 0001, Sofonias Yitagesu |
World Wide Web (WWW) | 1 |
| 2022 | Learning Disentangled Semantic Representations for Zero-Shot Cross-Lingual Transfer in Multilingual Machine Reading ComprehensionabstractLinjuan Wu, Shaojuan Wu, Xiaowang Zhang, Deyi Xiong, Shizhan Chen, Zhiqiang Zhuang, Zhiyong Feng. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Linjuan Wu, Shaojuan Wu, Xiaowang Zhang, Deyi Xiong, Shizhan Chen, Zhiqiang Zhuang, Zhiyong Feng 0002 |
ACL (1) | 2 |
| 2022 | Inter-subtask Consistent Representation Learning for Visual Commonsense Reasoning
Shaojuan Wu, Xiaowang Zhang |
ICANN (3) | 2 |
| 2022 | Function-words Adaptively Enhanced Attention Networks for Few-Shot Inverse Relation ClassificationabstractThe relation classification is to identify semantic relations between two entities in a given text. While existing models perform well for classifying inverse relations with large datasets, their performance is significantly reduced for few-shot learning. In this paper, we propose a function words adaptively enhanced attention framework (FAEA) for few-shot inverse relation classification, in which a hybrid attention model is designed to attend class-related function words based on meta-learning. As the involvement of function words brings in significant intra-class redundancy, an adaptive message passing mechanism is introduced to capture and transfer inter-class differences.We mathematically analyze the negative impact of function words from dot-product measurement, which explains why the message passing mechanism effectively reduces the impact. Our experimental results show that FAEA outperforms strong baselines, especially the inverse relation accuracy is improved by 14.33% under 1-shot setting in FewRel1.0. Chunliu Dou, Shaojuan Wu, Xiaowang Zhang, Zhiyong Feng 0002, Kewen Wang 0001 |
IJCAI | 2 |
| 2022 | Structure-sensitive semantic matching for aggregate question answering over knowledge base
Shaojuan Wu, Yunjie Wu, Linyi Han, Jiarui Zhang 0005, Xiaowang Zhang, Zhiyong Feng 0002 |
J. Web Semant. | 1 |
| 2021 | Re-embedding Difficult Samples via Mutual Information Constrained Semantically Oversampling for Imbalanced Text ClassificationabstractDifficult samples of the minority class in imbalanced text classification are usually hard to be classified as they are embedded into an overlapping semantic region with the majority class.In this paper, we propose a Mutual Information constrained Semantically Oversampling framework (MISO) that can generate anchor instances to help the backbone network determine the re-embedding position of a non-overlapping representation for each difficult sample.MISO consists of (1) a semantic fusion module that learns entangled semantics among difficult and majority samples with an adaptive multi-head attention mechanism, (2) a mutual information loss that forces our model to learn new representations of entangled semantics in the non-overlapping region of the minority class, and (3) a coupled adversarial encoder-decoder that fine-tunes disentangled semantic representations to remain their correlations with the minority class, and then using these disentangled semantic representations to generate anchor instances for each difficult sample.Experiments on a variety of imbalanced text classification tasks demonstrate that anchor instances help classifiers achieve significant improvements over strong baselines. Shizhan Chen, Xiaowang Zhang, Zhiyong Feng 0002, Deyi Xiong, Shaojuan Wu, Chunliu Dou |
EMNLP (1) | 6 |
| 2021 | A Mutual Information-Based Disentanglement Framework for Cross-Modal Retrieval
Xiaowang Zhang, Shaojuan Wu, Chunliu Dou, Zhiyong Feng 0002 |
ICONIP (4) | 4 |
| 2020 | EmoEM: Emotional Expression in a Multi-turn Dialogue ModelabstractEmotional intelligence is a crucial part for human-machine dialogue system. However, the existing research on dialogue mainly faces three problems: (1) focus on the content level of each response while ignoring the impact of emotional factors in the multi-turn dialogue; (2) lacking scalability and adaptability is that only the emotion categories specified by users are generated in a single-turn dialogue; (3) it is difficult to capture and perceive fine-grained emotions and the speaker's emotional state according to the emotional context. To address these problems, we propose an emotional expression model in multi-turn dialogue (EmoEM), which combines emotion-semantic graph with multitask learning mechanism, applying the dialogue generator based on seq2seq network and graph convolution network (GCN) to generate more natural and personalized emotional responses in a structured manner. Generally, EmoEM considers constructing emotion-semantic graph to describe explicit and implicit emotions dynamically. Then, the emotion-semantic graph is applied to the dialogue generator based on the seq2seq neural network, mainly to improve the semantic consistency and text quality in multi-turn dialogue. Moreover, multi-task learning mechanism is introduced to enhance the emotional expression of the text and obtain expected emotional responses. The experimental results show that EmoEM outperforms several baselines in BLEU, diversity and emotional expression. Shaojuan Wu, Xiaowang Zhang, Shizhan Chen, Yuchun Shu, Zhiyong Feng 0002 |
ICTAI | 2 |