VLDB 2026 Research / reviewers in the wild / expert
Dongsheng Zou
dblp:59/7694
· DBLP profile ↗
15ranked-venue papers
1as first author
11since 2021 · last 2026
0000-0002-8001-6461ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 6 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MIRAGE: Scaling Test-Time Inference with Parallel Graph-Retrieval-Augmented Reasoning ChainsabstractLarge reasoning models (LRMs) have shown significant progress in test-time scaling through chain-of-thought prompting. Current approaches like search-o1 integrate retrieval augmented generation (RAG) into multi-step reasoning processes but rely on a single, linear reasoning path while incorporating unstructured textual information in a flat, context-agnostic manner. As a result, these approaches can lead to error accumulation throughout the reasoning chain, which significantly limits its effectiveness in medical question-answering (QA) tasks where both accuracy and traceability are critical requirements. To address these challenges, we propose MIRAGE (Multi-path Inference with Retrieval-Augmented Graph Exploration), a novel test-time scalable reasoning framework that performs dynamic multi-path inference over structured medical knowledge graphs. Specifically, MIRAGE 1) decomposes complex queries into entity-grounded sub-questions, 2) executes parallel inference paths, 3) retrieves evidence adaptively via neighbor expansion and multi-hop traversal, and 4) integrates answers using cross-path verification to resolve contradictions. Experiments on three medical QA benchmarks (GenMedGPT-5k, CMCQA, and ExplainCPE) show that MIRAGE consistently outperforms GPT-4o, Tree-of-Thought variants, and other retrieval-augmented baselines in both automatic and human evaluations. Additionally, MIRAGE improves interpretability by generating explicit reasoning chains that trace each factual claim to concrete paths within the knowledge graph, making it especially suitable for complex medical reasoning scenarios. Kaiwen Wei, Rui Shan, Dongsheng Zou, Jianzhong Yang, Bi Zhao, Junnan Zhu |
AAAI | 3 |
| 2024 | MRC-FEE: Machine Reading Comprehension for Chinese Financial Event ExtractionabstractThe event extraction problem involves detecting event trigger words and extracting their corresponding event arguments. In contrast to general event extraction, financial event extraction focuses on financial texts, primarily at the document level. Existing methods mainly rely on a uniform sequence labeling model to identify event triggers and event arguments. However, this approach struggles with identifying nested and long entities and faces difficulties in handling complex event arguments. To address these challenges, we propose MRCFEE, a Chinese financial event extraction model based on machine reading comprehension (MRC). In the multi-turn MRC, the input to the model incorporates external knowledge and historical answers based on question templates and previous extraction results. Subsequently, we utilize a large-scale pre-trained language model in the financial field as the embedding layer and employ BiGRU to extract contextual information further. Then, we employ two binary classifiers to identify the probabilities of the starting and ending positions. Finally, we utilize the dynamic thresholding method to assess the rationality of the boundary of the obtained results. Experimental results on the DuEE-Fin dataset demonstrate that our model outperforms the previous methods, obtaining 84.1% and 75.6% F1 values for event detection and event argument extraction, respectively. Dongsheng Zou, Xinyi Song, Kang Xi |
CSCWD | 2 |
| 2024 | AdaStyleSpeech: A Fast Stylized Speech Synthesis Model Based on Adaptive Instance NormalizationabstractStylized speech synthesis transforms text into a specific style of speech guided by reference speech. Despite recent advancements in speech synthesis, challenges persist in this domain, including limitations in quality, speed, and similarity. To address these issues, we introduce AdaStyleSpeech, a novel model for stylized speech synthesis. This model can directly extract the style vector from reference speech. By combining textual information and style vectors, AdaStyleSpeech effectively transfers content into stylized speech using adaptive instance normalization. Additionally, we present AdaGANSpeech, a multi-style synthesis model based on stylistic mutual information and generative adversarial networks. Unlike AdaStyleSpeech, it works faster and can generate more diverse speech without the need for a reference. Experimental results demonstrate that AdaStyleSpeech attains remarkable outcomes in synthesis quality, making it a State-of-the-Art solution in stylized speech synthesis. AdaGANSpeech addresses the AdaStyleSpeech’s reliance on reference speech guidance during the generation phase and exhibits notable advantages in speech diversity, clarity, and synthesis speed. Dongsheng Zou |
ICME | 2 |
| 2024 | RDLinear: A Novel Time Series Forecasting Model Based on Decomposition with RevINabstractTime series forecasting, with its wide range of practical applications such as power load and weather prediction, has become a pivotal field of research. Over the past few years, neural network models have made remarkable progress in this domain. Many time series forecasting models now employ sequence decomposition techniques to enhance forecasting accuracy, including Autoformer, DLinear, and MICN. These techniques break down the original time series data into two components: trend and seasonal term, to facilitate more accurate predictions. However, a significant limitation of existing models that utilize sequence decomposition is their incomplete exploitation of the trend component. To address this issue, we introduce RDLinear, a structurally simple model designed to fully leverage the unique attributes of sequence decomposition. RDLinear employs distinct forecasting strategies, with a primary focus on utilizing the RevIN method to predict the trend component. In this paper, we present extensive experimental results on multiple real-world datasets. Our findings demonstrate that RDLinear outperforms other time-series forecasting models, particularly in long-term forecasting. Furthermore, ablation experiments confirm the effectiveness of our proposed method. Dongsheng Zou, Bi Zhao, Jiyuan Liu 0011, Naiquan Chai, Xinyi Song |
IJCNN | 2 |
| 2024 | Entity and Evidence Guided Attention for Document-Level Relation ExtractionabstractDocument-level relation extraction (DRE) aims to extract relations between entities in unstructured documents. Unlike sentence-level relation extraction, DRE introduces complexities associated with entities that appear in multiple sentences with different mentions, and where the head and tail entities of a given relation triple can be situated in diverse sentences. Consequently, the aggregation of semantic information for entity pairs is a paramount challenge in DRE. To address this challenge, we introduce a novel framework, Entity and Evidence Guided Attention (EEGA). This framework employs a pre-trained language model as an encoder, crafting rich contextual representations for entity pairs through the integration of both relation extraction and evidence retrieval. Initially, we guide the model’s attention towards the contextual information surrounding an entity pair using evidence. We then introduce an attention mechanism that assigns weights to words, guiding the extraction of semantic information at different levels: sentence, document, and evidence. An adaptive fusion module dynamically amalgamates the semantic information of the entity pairs at various granularities to obtain context-aware entity pair representations with rich semantics. Additionally, we propose self-training with relation labels and evidence attention on distantly supervised data to enhance DocRE performance. Experimental results on a benchmark dataset demonstrate the superiority of EEGA over strong baselines. Dongsheng Zou, Xinyi Song, Bi Zhao |
IJCNN | 2 |
| 2024 | RAVL: A Retrieval-Augmented Visual Language Model Framework for Knowledge-Based Visual Question Answering
Naiquan Chai, Dongsheng Zou, Jiyuan Liu 0011, Xinyi Song |
NLPCC (3) | 2 |
| 2024 | PqE: Zero-Shot Document Expansion for Dense Retrieval with Large Language Models
Jiyuan Liu 0011, Dongsheng Zou, Naiquan Chai, Xinyi Song |
NLPCC (1) | 2 |
| 2023 | FW-ECPE: An Emotion-Cause Pair Extraction Model Based on Fusion Word VectorsabstractEmotion-Cause Pair Extraction (ECPE) aims to extract potential emotion-cause pairs from text without emotion labels. It lays an important foundation for downstream research such as causal reasoning, public opinion prediction, and reason detection. However, the ECPE task now faces two dilemmas: 1) insufficient utilization of word sequence information, and 2) inadequate use of position information between clauses. To address the above problems, we proposed an Emotion-Cause Pair Extraction model based on Fusion Word Vectors named FW-ECPE. It is a two-stage model that first extracts emotion clauses and cause clauses respectively then combines them into pairs and filters out the right emotion-cause pairs. The Fusion Word Vector is reflected in two aspects. Firstly, we integrate the clause context vectors and the emotion clauses prediction results with cause context vectors in cause clauses extraction. Secondly, in the emotion-cause pair extraction stage, we fuse the position information between clauses and contextual information. Finally, we extend Easy Data Augmentation, a corpus enhancement algorithm, to enlarge the amount of data and alleviate the risk of overfitting. The experiment results show that our proposed approach outperforms the previous methods on a benchmark dataset. Xinyi Song, Dongsheng Zou |
IJCNN | 2 |
| 2023 | DehazeDM: Image Dehazing via Patch Autoencoder Based on Diffusion ModelsabstractImage dehazing is a crucial computer vision application with the primary objective of estimating haze-free images from hazy images. Deep neural network architectures have emerged as the dominant approaches and achieved remarkable progress. However, due to the intricacy, existing dehazing methods need help to train large deep learning networks. This work proposes a novel image dehazing network based on Diffusion Model (DehazeDM). Firstly, by segmenting the image into patches during the sampling procedure, we can dehaze images of arbitrary size. Then we compress the image into the latent space via the auto-encoder model and conduct the diffusion operation in the latent space, significantly decreasing the computational complexity associated with the task while exhibiting negligible effects on the perceptual fidelity of the resultant images. Extensive experiments verify the effectiveness and the superior performance of DehazeDM in image dehazing. Dongsheng Zou, Xinyi Song |
SMC | 2 |
| 2022 | Multiway Bidirectional Attention and External Knowledge for Multiple-choice Reading ComprehensionabstractTeaching machines to understand human language is one of the most elusive challenges in artificial intelligence. Machine reading comprehension is a crucial task in evaluating how computer systems understand natural language. This study presents a machine reading comprehension model based on external knowledge. We use a framework named K-Adapter to infuse two kinds of external knowledge with two specific adapters. This model can capture richer semantic information, which is more suitable for real application scenarios. The proposed model is evaluated on the COSMOS QA dataset and outperforms the competitive baselines. Dongsheng Zou, Xinyi Song, Kang Xi |
SMC | 1 |
| 2021 | Four-way Bidirectional Attention for Multiple-choice Reading ComprehensionabstractAs one of the crucial tasks of natural language processing, machine reading comprehension has gained increased attention in recent years. In this paper, we propose a four-way bidirectional attention network for a multiple-choice reading comprehension task, where every question comes with a set of candidate options and only one correct answer. Current methods on such tasks usually judge options independently and ignore their relations. Thus, this work designs a four-way bidirectional attention strategy to formulate the interactions among the passage, questions and candidate options. In particular, the relations among options are well represented. This enables the model to leverage the option correlation information for inferring the final answer accurately. The experimental evaluations on the CosmosQA dataset demonstrate the competitive performance of our model, and confirm the effectiveness of the option comparison strategy. Dongsheng Zou, Xiwang Guo 0001, Liang Qi 0001, Ying Tang 0001, Jieying Yuan |
SMC | 2 |
| 2020 | Embedding Compression with Right Triangle Similarity Transformations
Dongsheng Zou, Jieying Yuan |
ICANN (2) | 2 |
| 2020 | Learning Discrete Sentence Representations via Construction & Decomposition
Dongsheng Zou |
ICONIP (1) | 2 |
| 2020 | A two-stage approach for automatic liver segmentation with Faster R-CNN and DeepLab
Dongsheng Zou, Jingpei Dan, Guowu Song |
Neural Comput. Appl. | 2 |
| 2018 | DSL: Automatic Liver Segmentation with Faster R-CNN and DeepLab
Dongsheng Zou |
ICANN (2) | 2 |