EDBT 2026 Demo / reviewers in the wild / expert
Di Wu 0088
dblp:52/328-88
· DBLP profile ↗
14ranked-venue papers
3as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Introducing Visual Scenes and Reasoning: A More Realistic Benchmark for Spoken Language UnderstandingabstractSpoken Language Understanding (SLU) consists of two sub-tasks: intent detection (ID) and slot filling (SF). Given its broad range of real-world applications, enhancing SLU for practical deployment is increasingly critical. Profile-based SLU addresses ambiguous user utterances by incorporating context awareness (CA), user profiles (UP), and knowledge graphs (KG) to support disambiguation, thereby advancing SLU research toward real-world applicability. However, existing SLU datasets still fall short in representing real-world scenarios. Specifically, (1) CA uses one-hot vectors for representation, which is overly idealized, and (2) models typically focuses solely on predicting intents and slot labels, neglecting the reasoning process that could enhance performance and interpretability. To overcome these limitations, we introduce VRSLU, a novel SLU dataset that integrates both Visual images and explicit Reasoning. For over-idealized CA, we use GPT-4o and FLUX.1-dev to generate images reflecting users’ environments and statuses, followed by human verification to ensure quality. For reasoning, GPT-4o is employed to generate explanations for predicted labels, which are then refined by human annotators to ensure accuracy and coherence. Additionally, we propose an instructional template, LR-Instruct, which first predicts labels and then generates corresponding reasoning. This two-step approach helps mitigate the influence of reasoning bias on label prediction. Experimental results confirm the effectiveness of incorporating visual information and highlight the promise of explicit reasoning in advancing SLU. Di Wu 0088, Liting Jiang, Ruiyu Fang, Bianjing, Hongyan Xie, Haoxiang Su, Hao Huang 0009, Zhongjiang He, Shuangyong Song, Xuelong Li 0001 |
AAAI | 1 |
| 2026 | Direct preference optimization with Pareto dominance constraint for online multi-objective alignment
Hongyan Xie, Yikun Ban, Ruiyu Fang, Di Wu 0088, Zixuan Huang 0012, Deqing Wang 0001, Jianxin Li 0002, Shuangyong Song |
Neurocomputing | 4 |
| 2025 | Utterance as A Bridge: Few-shot Joint Learning of Empathy Detection and Empathy Intent ClassificationabstractEmpathy detection (ED) and empathy intent classification (EIC) aim to identify the empathy direction expressed in user utterances and the underlying empathy intent behind them. Previous studies show that facilitating information transfer between tasks can enhance model performance. However, the interaction between ED and EIC in few-shot learning remains underexplored. To this end, we identify the challenges in jointly training ED and EIC in a few-shot setting: establishing effective information transfer between them and improving the model’s generalization capability. We propose a novel model called USB. For information transfer, the interactive module maps empathy and empathy intent labels through utterances to model task correlations. For generalization capability, after capturing empathy and empathy intent representations with an adaptive fusion module, we introduce a multi-level contrastive learning strategy to optimize representations at task and label levels, enhancing generalization. Experimental results on two public datasets show that our model outperforms all baselines. Liting Jiang, Di Wu 0088, Shuangyong Song, Yanbing Li, Hao Huang 0009 |
ICASSP | 2 |
| 2025 | A Label Co-occurrence Transformation Network for Joint Empathy Detection and Empathy Intent ClassificationabstractEmpathy detection (ED) aims to understand the user’s empathy direction, while empathy intent classification (EIC) focuses on identifying the empathy intent behind the user’s utterance. Both tasks have garnered significant attention. Recent studies have shown that jointly training these tasks can improve model performance, as their correlation enhances the diversity of information. However, previous studies have relied solely on shallow information transfer between two task representations, failing to fully leverage the inter-task correlation, thus limiting performance. To this end, we propose a novel Label Co-occurrence Transformation Network (LCoT-Net), which models the correlation between the two tasks using the co-occurrence matrix of empathy and empathy intent labels as a medium. By performing category feature transformation at both the label and utterance levels, we achieve two-level mutual task guidance. Experimental results demonstrate that our model achieves competitive performance across various settings on two public datasets. Liting Jiang, Di Wu 0088, Haoxiang Su, Xiaoyong Guo, Shuangyong Song, Yanbing Li |
ICASSP | 2 |
| 2025 | RAICL-DSC: Retrieval-Augmented In-Context Learning for Dialogue State Correction
Haoxiang Su, Hongyan Xie, Di Wu 0088, Liting Jiang, Hao Huang 0009, Zhongjiang He, Ruiyu Fang, Shuangyong Song |
Knowl. Based Syst. | 4 |
| 2025 | A classifier expansion framework with dual knowledge distillation and dynamic weighting for continual relation extraction
Aonan Mao, Di Wu 0088, Liting Jiang, Shuangyong Song, Yanbing Li, Hao Huang 0009, Wushour Slamu |
J. Supercomput. | 2 |
| 2024 | Fact-Aware Summarization with Contrastive Learning for Few-Shot Dialogue State TrackingabstractDialogue state tracking (DST) is a crucial component of task-oriented dialogue systems, as it aims to accurately track the user’s goals throughout the dialogue history. However, DST models struggle with new domains due to limited annotated data, leading to poor performance. To solve this key challenge in DST, we propose a model called Fact-aware Summarization model for few-shot DST (FaS-DST), which introduces a "Summarize, Extract, and Select" pattern. Specifically, we decompose DST into three sub-tasks: generating candidate summaries, extracting dialogue states, and scoring the candidates to select the most accurate one. Contrastive learning is incorporated to train a candidate scorer, which improves faithfulness and factuality in dialogue summarization. Additionally, we employ two strategies namely data augmentation and summary & state concatenation to improve the model’s training effectiveness. Experimental results demonstrate that FaS-DST outperforms state-of-the-art models on both MultiWOZ 2.0 and MultiWOZ 2.1 datasets in few-shot settings. Sijie Feng, Haoxiang Su, Hongyan Xie, Di Wu 0088, Hao Huang 0009, Wushour Slamu |
ICASSP | 4 |
| 2024 | Domain-Slot Aware Contrastive Learning for Improved Dialogue State TrackingabstractLarge-scale pre-trained neural language model has facilitated to achieve the state-of-the-art performance on Dialogue State Tracking (DST) tasks. One of the existing works models the semantic correlation between the dialogue context and (domain, slot) pair encoded by BERT and make the prediction. Despite the effectiveness, they ignore the fact that there is no perfect semantic correspondence between (domain, slot) pair and the dialogue context. In this paper, we propose a domain-slot aware contrastive learning framework to solve this problem, which proposes three methods to bridge the semantic gap between the dialogue context and the (domain, slot) by constructing training sample pairs to fine-tune the BERT model and use it for base DST model. The experiments demonstrate that our proposed method has improved the performance of the baseline model on the MultiWOZ2.1 and MultiWOZ2.4 datasets, yielding competitive results. Haoxiang Su, Sijie Feng, Hongyan Xie, Di Wu 0088, Hao Huang 0009, Zhongjiang He, Shuangyong Song, Ruiyu Fang, Xiaomeng Huang, Wushour Slamu |
ICASSP | 4 |
| 2024 | Dual Level Intent-Slot Interaction for Improved Multi-Intent Spoken Language UnderstandingabstractMulti-intent spoken language understanding consists of two typical subtasks: multi-intent detection and slot filling. Existing approach suffers from two limitations: (1) It fails to explicitly model the information transfer between slots associated within the same intent clause; (2) Using a co-occurrence matrix of both label encodings introduces needless slot positional information such as the prefix ‘B-’ or ‘I-’. For (1), we propose a Gaussian Graph Attention Network that allows interaction to focus not only on the connection between slots within the current intent clause, but also on the connection between intent, and between intent and slot. For (2), we use a co-occurrence matrix of intent categories and slot types to model the knowledge transfer between the two subtasks in the corpus-level interaction, bypassing the introduction of slot positional information. Our framework achieves significant accuracy gains on both the MixATIS and MixSNIPS datasets. Di Wu 0088, Liting Jiang, Haoxiang Su, Hao Huang 0009 |
ICASSP | 1 |
| 2024 | An Uyghur Extension to the MASSIVE Multi-lingual Spoken Language Understanding Corpus with Comprehensive Evaluations
Ainikaerjiang Aimaiti, Di Wu 0088, Liting Jiang, Gulinigeer Abudouwaili, Hao Huang 0009, Wushour Slamu |
INTERSPEECH | 2 |
| 2024 | CEA-Net: a co-interactive external attention network for joint intent detection and slot filling
Di Wu 0088, Liting Jiang, Hao Huang 0009 |
Neural Comput. Appl. | 1 |
| 2023 | Empathy Intent Drives Empathy DetectionabstractEmpathy plays an important role in the human dialogue.Detecting the empathetic direction expressed by the user is necessary for empathetic dialogue systems because it is highly relevant to understanding the user's needs.Several studies have shown that empathy intent information improves the ability to response capacity of empathetic dialogue.However, the interaction between empathy detection and empathy intent recognition has not been explored.To this end, we invite 3 experts to manually annotate the healthy empathy detection datasets IEMPATHIZE and TwittEmp with 8 empathy intent labels, and perform joint training for the two tasks.Empirical study has shown that the introduction of empathy intent recognition task can improve the accuracy of empathy detection task, and we analyze possible reasons for this improvement.To make joint training of the two tasks more challenging, we propose a novel framework, Cascaded Label Signal Network, which uses the cascaded interactive attention module and the label signal enhancement module to capture feature exchange information between empathy and empathy intent representations.Experimental results show that our framework outperforms all baselines under both settings on the two datasets.1 Liting Jiang, Di Wu 0088, Bohui Mao, Yanbing Li, Wushour Slamu |
EMNLP | 2 |
| 2023 | Mitigating Domain Dependency for Improved Speech Enhancement Via SNR Loss BoostingabstractCurrent supervised speech enhancement methods based on deep learning typically utilize amplitude-based loss functions for optimization, such as Mean Absolute Error (MAE) or Mean Square Error (MSE) loss, which measures the difference between the amplitudes of the estimated and clean speech signals. However, models trained with these losses heavily depend on specific domain properties, i.e. speaker, noise type, and signal-to-noise ratio (SNR). In this paper, we first validate this assumption by visually analyzing the model’s internal representation, and these dependencies result in severe performance degradation in unseen situations. Given that the SNR is irrelevant to speakers and noise types, we propose a simple but effective novel objective function by minimizing the discrepancy between indirectly estimated SNR and true SNR over time-frequency units to alleviate the model’s reliance on those domain properties. Experimental results demonstrate that our proposed method outperforms other prevalent loss functions in terms of both performance gain and generalization capability. Di Wu 0088, Zhibin Qiu, Hao Huang 0009 |
ICASSP | 2 |
| 2023 | Emp-USIR: A Unidirectional Synchronous Interactive Reasoning Model for Empathetic DialogueabstractThe goal of an empathetic dialogue generative sys-tem is to generate coherently and relate emotional responses after finite turns of dialogue. Most of the previous work guided the dialogue system to generate empathetic responses from the dialogue history or emotional reasons. However, in people's daily communication, the empathetic response is often generated by deep reasoning through dialogue history and existing commonsense knowledge. To fill this gap, we propose an empathetic dialogue model of unidirectional synchronous interactive reasoning, Emp- Usir.Specifically, according to the existing commonsense knowledge base, we extract commonsense knowledge with emotional representation from it, and use it as the connection feature of dialogue history and remaining common-sense knowledge to guide the interactive reasoning between multi-turns dialogue history and commonsense knowledge, to imitate human reasoning based on commonsense to generate responses. Then, we propose a cross-token level attention mechanism to learn emotional dependence from reasoning features, to produce more smooth and diverse empathy responses. The experimental results and manual evaluation prove the effectiveness of the Emp-USIR model proposed in this paper on the widely-used benchmark dataset. Liting Jiang, Di Wu 0088, Yanbing Li, Wushour Slamu |
IJCNN | 2 |