VLDB 2026 Research / reviewers in the wild / expert
Dongning Rao
dblp:83/2236
· DBLP profile ↗
14ranked-venue papers
8as first author
12since 2021 · last 2026
0000-0002-6306-6811ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 7 first-author · 11 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Text-Routed Sparse Mixture-of-Experts Model with Explanation and Temporal Alignment for Multi-Modal Sentiment AnalysisabstractHuman-interaction-involved applications underscore the need for Multi-modal Sentiment Analysis (MSA). Although many approaches have been proposed to address the subtle emotions in different modalities, the power of explanations and temporal alignments is still underexplored. Thus, this paper proposes the Text-routed sparse mixture-of-Experts model with eXplanation and Temporal alignment for MSA (TEXT). TEXT first augments explanations for MSA via Multi-modal Large Language Models (MLLM), and then novelly aligns the representations of audio and video through a temporality-oriented neural network block. TEXT aligns different modalities with explanations and facilitates a new text-routed sparse mixture-of-experts with gate fusion. Our temporal alignment block merges the benefits of Mamba and temporal cross-attention. As a result, TEXT achieves the best performance across four datasets among all tested models, including three recently proposed approaches and three MLLMs. TEXT wins on at least four metrics out of all six metrics. For example, TEXT decreases the mean absolute error to 0.353 on the CH-SIMS dataset, which signifies a 13.5% decrement compared with recently proposed approaches. Dongning Rao, Yunbiao Zeng, Zhihua Jiang, Jujian Lv |
AAAI | 1 |
| 2026 | Leveraging dynamic few-shot prompting and ensemble method for task-oriented dialogue with subjective knowledgeabstractSubjective knowledge is key to meeting customer needs. Thus, the Subjective Knowledge-grounded Task-oriented Dialogue (SK-TOD) task tries to accommodate subjective user requests like “Does the restaurant have a good atmosphere?” by choosing relevant subjective knowledge snippets and generating appropriate responses. However, unlike existing methods like retrieval-augmented generation using external objective knowledge, selecting subjective knowledge and summarizing opinions from reviews in a specified scope pose new challenges. Therefore, this paper proposes the DESIGN ( D ynamic f E w- S hot prompt I n G and e N semble) method for SK-TOD. Specifically, DESIGN first adopts Aspect-Based Sentiment Analysis (ABSA) to enhance subjective knowledge snippets and then builds an ensemble composed of diverse base models for knowledge selection (KS). Here, the base models include both classification models and generative models. At last, for response generation (RG), DESIGN employs generative models conditioned on dialogue context and ABSA-enhanced knowledge. Particularly, we devise the sample selection via the similarity-alignment algorithm to choose similar samples dynamically for the few-shot prompting of KS and RG. We experiment on the 11th Dialog System Technology Challenge (DSTC11) SK-TOD benchmark and an extended dataset, ReDial, with 6147 instances. For KS, we beat the winner of DSTC11 and boosted the F1 for 7% regarding the baseline and achieved 86.16%. For RG, DESIGN outperforms baselines and the DSTC11 winner across eight metrics.E.g., DESIGN improves entailment performance by 5% over the DSTC11 winner and 10% over the baseline. 1 Dongning Rao, Jietao Zhuang, Zhihua Jiang |
Inf. Process. Manag. | 1 |
| 2025 | A Comprehensive Literary Chinese Reading Comprehension Dataset with an Evidence Curation Based SolutionabstractLow-resource language understanding is a challenging task, even for large language models (LLMs).An epitome of this problem is the CompRehensive lIterary chineSe readIng comprehenSion (CRISIS), whose difficulties include limited linguistic data, long input, and insight-required questions.Besides the compelling need to provide a larger dataset for CRISIS, excessive information, order bias, and entangled conundrums still plague the CRISIS solutions.Thus, we present the eVIdence cuRation with opTion shUffling and Abstract meaning representation-based cLauses segmenting (VIRTUAL) procedure for CRISIS, with the most extensive dataset.While the dataset is also named CRISIS, it results from a three-phase construction process, including question selection, data cleaning, and a silver-standard data augmentation step, which augments translations, celebrity profiles, government jobs, reign mottos, and dynasty to CRISIS.The six steps of VIRTUAL include embedding, shuffling, abstract meaning representation-based option segmenting, evidence extraction, solving, and voting.Notably, the evidence extraction algorithm facilitates the extraction of literary Chinese evidence sentences, translated evidence sentences, and annotations of keywords using a similaritybased ranking strategy.While CRISIS compiles understanding-required questions from seven sources, the experiments on CRISIS substantiate the effectiveness of VIRTUAL, with a 7 percent increase in accuracy compared to the baseline.Interestingly, both non-LLMs and LLMs exhibit order bias, and abstract meaning representation-based option segmenting is beneficial for CRISIS. Dongning Rao, Rongchu Zhou, Zhihua Jiang |
EMNLP | 1 |
| 2025 | Leveraging meta-data of code for adapting prompt tuning for code summarization
Zhihua Jiang, Dongning Rao |
Appl. Intell. | 3 |
| 2025 | A robust dialogue evaluation metric exploiting denoising, pre-training and ensembling
Dongning Rao, Lianyong Ling, Zhihua Jiang |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | Leveraging Context-Aware Prompting for Commit Message GenerationabstractWriting comprehensive commit messages is tedious yet important, because these messages describe changes of code, such as fixing bugs or adding new features.However, most existing methods focus on either only the changed lines or nearest context lines, without considering the effectiveness of selecting useful contexts.On the other hand, it is possible that introducing excessive contexts can lead to noise.To this end, we propose a code model COMMIT (Context-aware prOMpting based comMIt-message generaTion) in conjunction with a code dataset CODEC (COntext and metaData Enhanced Code dataset).Leveraging program slicing, CODEC consolidates code changes along with related contexts via property graph analysis.Further, utilizing CodeT5+ as the backbone model, we train COMMIT via context-aware prompt on CODEC.Experiments show that COMMIT can surpass all compared models including pre-trained language models for code (code-PLMs) such as Com-mitBART and large language models for code (code-LLMs) such as Code-LlaMa.Besides, we investigate several research questions (RQs), further verifying the effectiveness of our approach.We release the data and code at: https: //github.com/Jnunlplab/COMMIT.git. Zhihua Jiang, Dongning Rao, Guanghui Ye |
EMNLP | 3 |
| 2024 | An Empirical Study of Leveraging PLMs and LLMs for Long-Text Summarization
Zhihua Jiang, Junzhan Yang, Dongning Rao |
PRICAI (2) | 3 |
| 2023 | Ancient Chinese Machine Reading Comprehension Exception Question Dataset with a Non-trivial Model
Dongning Rao, Guanju Huang, Zhihua Jiang |
PRICAI (2) | 1 |
| 2022 | IM⌃2: an Interpretable and Multi-category Integrated Metric Framework for Automatic Dialogue EvaluationabstractEvaluation metrics shine the light on the best models and thus strongly influence the research directions, such as the recently developed dialogue metrics USR, FED, and GRADE.However, most current metrics evaluate the dialogue data as isolated and static because they only focus on a single quality or several qualities.To mitigate the problem, this paper proposes an interpretable, multi-faceted, and controllable framework IM 2 (Interpretable and M ulti-category Integrated M etric) to combine a large number of metrics which are good at measuring different qualities.The IM 2 framework first divides current popular dialogue qualities into different categories and then applies or proposes dialogue metrics to measure the qualities within each category and finally generates an overall IM 2 score.An initial version of IM 2 was submitted to the AAAI 2022 Track5.1@DSTC10challenge 1 and took the 2 nd place on both of the development and test leaderboard.After the competition, we develop more metrics and improve the performance of our model.We compare IM 2 with other 13 current dialogue metrics and experimental results show that IM 2 correlates more strongly with human judgments than any of them on each evaluated dataset 2 . Zhihua Jiang, Guanghui Ye, Dongning Rao |
EMNLP | 3 |
| 2021 | STANKER: Stacking Network based on Level-grained Attention-masked BERT for Rumor Detection on Social MediaabstractRumor detection on social media puts pretrained language models (LMs), such as BERT, and auxiliary features, such as comments, into use.However, on the one hand, rumor detection datasets in Chinese companies with comments are rare; on the other hand, intensive interaction of attention on Transformer-based models like BERT may hinder performance improvement.To alleviate these problems, we build a new Chinese microblog dataset named Weibo20 1 by collecting posts and associated comments from Sina Weibo and propose a new ensemble named STANKER (Stacking neTwork bAsed-on atteNtion-masKed BERT).STANKER adopts two level-grained attentionmasked BERT (LGAM-BERT) models as base encoders.Unlike the original BERT, our new LGAM-BERT model takes comments as important auxiliary features and masks coattention between posts and comments on lower-layers.Experiments on Weibo20 and three existing social media datasets showed that STANKER outperformed all compared models, especially beating the old state-of-theart on Weibo dataset. Dongning Rao, Zhihua Jiang |
EMNLP (1) | 1 |
| 2021 | Syntax and Sentiment Enhanced BERT for Earliest Rumor Detection
Dongning Rao, Zhihua Jiang |
NLPCC (1) | 2 |
| 2021 | A dual deep neural network with phrase structure and attention mechanism for sentiment analysis
Dongning Rao, Sihong Huang, Zhihua Jiang, Ganesh Gopal Devarajan, Rizwan Patan |
Neural Comput. Appl. | 1 |
| 2019 | Scalable and optimal planning based on PregelabstractSummary Automated planning generates plans for specific tasks. Optimal planning aims at generating optimal plans under global constraints. As a result, the divide‐and‐conquer method is not applicable for optimal planning. Therefore, engineering applications of optimal planning face the scalability issue. Fortunately, cloud computing tools are on the shelf. For example, the Apache Spark is an engine for big data processing. It supports the Pregel for scalable computing. Therefore, we proposed an optimal Planning method based on the Pregel, called the PbP. Unlike classical planning, the PbP method uses the Pregel as the computation model, instead of the traditional state‐space searching. The core idea is to transform planning problems into graph processing problems. Specifically, actions are mapped into vertices, partial orders between actions are mapped into edges between vertices, and states are mapped into messages. Furthermore, the planning is viewed as message propagating in the graph, and plan traces are stored as attributes of vertices. Experimental results showed the feasibility of the proposed method PbP. Moreover, compared with state‐of‐the‐art optimal planners, our approach is more scalable and faster. Zhihua Jiang, Dongning Rao |
Concurr. Comput. Pract. Exp. | 2 |
| 2016 | Cost-Sensitive Action Model LearningabstractAction model learning can relieve people from writing planning domain descriptions from scratch. Real-world learners need to be sensitive to all kinds of expenses which it will spend in the learning. However, most of previous studies in this research line only considered the running time as the learning cost. In real-world applications, we will spend extra expense when we carry out actions or get observations, particularly for online learning. The learning algorithm should apply more techniques for saving the total cost when keeping a high rate of accuracy. The cost of carrying out actions and getting observations is the dominated expense in online learning. Therefore, we design a cost-sensitive algorithm to learn action models under partial observability. It combines three techniques to lessen the total cost: constraints, filtering and active learning. These techniques are used in observation reduction in action model learning. First, the algorithm uses constraints to confine the observation space. Second, it removes unnecessary observations by belief state filtering. Third, it actively picks up observations based on the results of the previous two techniques. This paper also designs strategies to reduce the amount of plan steps used in the learning. We performed experiments on some benchmark domains. It shows two results. For one thing, the learning accuracy is high in most cases. For the other, the algorithm dramatically reduces the total cost according to the definition of cost in this paper. Therefore, it is significant for real-world learners, especially, when long plans are unavailable or observations are expensive. Dongning Rao, Zhihua Jiang |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 1 |