Haoxiang Su

dblp:305/1874 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Speech recognition and synthesis · 40% Question answering and dialogue systems · 24% Vision and language · 20%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
dialogue state tracking
1.222023
Scalable-DSC: A Structural Template Prompt Approach to Scalable Dialogue State Correction · EMNLP 2023
Correctable-DST: Mitigating Historical Context Mismatch between Training and Inference for Improved Dialogue State Tracking · EMNLP 2022
Natural language and speech › Speech recognition and synthesis › spoken language understanding
intent detection and slot filling
1.012026
Introducing Visual Scenes and Reasoning: A More Realistic Benchmark for Spoken Language Understanding · AAAI 2026
Natural language and speech › Speech recognition and synthesis
spoken language understanding
1.012026
Introducing Visual Scenes and Reasoning: A More Realistic Benchmark for Spoken Language Understanding · AAAI 2026
Natural language and speech › Language models and text generation › text generation › large language model generation
prompt-based generation
0.712023
Scalable-DSC: A Structural Template Prompt Approach to Scalable Dialogue State Correction · EMNLP 2023

Methods — techniques the papers use, named apart from their topics

reasoning generation · 1.0instruction tuning · 1.0image generation · 1.0predictive state simulator · 0.7large language model · 0.7
YearPublicationVenuePosition
2026 Introducing Visual Scenes and Reasoning: A More Realistic Benchmark for Spoken Language Understanding
abstract
Spoken Language Understanding (SLU) consists of two sub-tasks: intent detection (ID) and slot filling (SF). Given its broad range of real-world applications, enhancing SLU for practical deployment is increasingly critical. Profile-based SLU addresses ambiguous user utterances by incorporating context awareness (CA), user profiles (UP), and knowledge graphs (KG) to support disambiguation, thereby advancing SLU research toward real-world applicability. However, existing SLU datasets still fall short in representing real-world scenarios. Specifically, (1) CA uses one-hot vectors for representation, which is overly idealized, and (2) models typically focuses solely on predicting intents and slot labels, neglecting the reasoning process that could enhance performance and interpretability. To overcome these limitations, we introduce VRSLU, a novel SLU dataset that integrates both Visual images and explicit Reasoning. For over-idealized CA, we use GPT-4o and FLUX.1-dev to generate images reflecting users’ environments and statuses, followed by human verification to ensure quality. For reasoning, GPT-4o is employed to generate explanations for predicted labels, which are then refined by human annotators to ensure accuracy and coherence. Additionally, we propose an instructional template, LR-Instruct, which first predicts labels and then generates corresponding reasoning. This two-step approach helps mitigate the influence of reasoning bias on label prediction. Experimental results confirm the effectiveness of incorporating visual information and highlight the promise of explicit reasoning in advancing SLU.
Di Wu 0088, Liting Jiang, Ruiyu Fang, Bianjing, Hongyan Xie, Haoxiang Su, Hao Huang 0009, Zhongjiang He, Shuangyong Song, Xuelong Li 0001
AAAI6
2026 A Cooperative Model Predictive Control Approach for Virtual Coupling Train Arrival in Metro Railways
abstract
The synchronous arrival of trains in a virtually coupled train set (VCTS) is crucial to ensure efficient operations. The primary objective of VCTS is to decrease the inter-train spacing while stopping, aiming to minimize the time difference between arrivals. In this paper, we propose a distributed cooperative model predictive control-based (DCMPC) approach for synchronous arrival of VCTS. Firstly, the dynamics model of the unit train in VCTS is constructed, which considers external uncertainties. Then, a novel dynamic penalty matrix updating method is introduced to improve the accuracy and efficiency of unit train stopping. We also present the design of terminal constraints on the DCMPC control problem and the stability proof to ensure the feasibility of the controller. Finally, the effectiveness of the proposed method is verified through simulation experiments with concrete data from the Chengdu Metro Line NO.8 in China. Compared to conventional control approaches, the proposed DCMPC significantly reduces the arrival time difference while increasing the average stopping speed, resulting in a considerable 4% increase in the passing capacity of the line.
Jialei Liang, Ming Chai, Haoxiang Su, Tao Tang 0004
IEEE Trans. Intell. Transp. Syst.3
2025 A Label Co-occurrence Transformation Network for Joint Empathy Detection and Empathy Intent Classification
abstract
Empathy detection (ED) aims to understand the user’s empathy direction, while empathy intent classification (EIC) focuses on identifying the empathy intent behind the user’s utterance. Both tasks have garnered significant attention. Recent studies have shown that jointly training these tasks can improve model performance, as their correlation enhances the diversity of information. However, previous studies have relied solely on shallow information transfer between two task representations, failing to fully leverage the inter-task correlation, thus limiting performance. To this end, we propose a novel Label Co-occurrence Transformation Network (LCoT-Net), which models the correlation between the two tasks using the co-occurrence matrix of empathy and empathy intent labels as a medium. By performing category feature transformation at both the label and utterance levels, we achieve two-level mutual task guidance. Experimental results demonstrate that our model achieves competitive performance across various settings on two public datasets.
Liting Jiang, Di Wu 0088, Haoxiang Su, Xiaoyong Guo, Shuangyong Song, Yanbing Li
ICASSP3
2025 RAICL-DSC: Retrieval-Augmented In-Context Learning for Dialogue State Correction
Haoxiang Su, Hongyan Xie, Di Wu 0088, Liting Jiang, Hao Huang 0009, Zhongjiang He, Ruiyu Fang, Shuangyong Song
Knowl. Based Syst.1
2024 Fact-Aware Summarization with Contrastive Learning for Few-Shot Dialogue State Tracking
abstract
Dialogue state tracking (DST) is a crucial component of task-oriented dialogue systems, as it aims to accurately track the user’s goals throughout the dialogue history. However, DST models struggle with new domains due to limited annotated data, leading to poor performance. To solve this key challenge in DST, we propose a model called Fact-aware Summarization model for few-shot DST (FaS-DST), which introduces a "Summarize, Extract, and Select" pattern. Specifically, we decompose DST into three sub-tasks: generating candidate summaries, extracting dialogue states, and scoring the candidates to select the most accurate one. Contrastive learning is incorporated to train a candidate scorer, which improves faithfulness and factuality in dialogue summarization. Additionally, we employ two strategies namely data augmentation and summary & state concatenation to improve the model’s training effectiveness. Experimental results demonstrate that FaS-DST outperforms state-of-the-art models on both MultiWOZ 2.0 and MultiWOZ 2.1 datasets in few-shot settings.
Sijie Feng, Haoxiang Su, Hongyan Xie, Di Wu 0088, Hao Huang 0009, Wushour Slamu
ICASSP2
2024 Domain-Slot Aware Contrastive Learning for Improved Dialogue State Tracking
abstract
Large-scale pre-trained neural language model has facilitated to achieve the state-of-the-art performance on Dialogue State Tracking (DST) tasks. One of the existing works models the semantic correlation between the dialogue context and (domain, slot) pair encoded by BERT and make the prediction. Despite the effectiveness, they ignore the fact that there is no perfect semantic correspondence between (domain, slot) pair and the dialogue context. In this paper, we propose a domain-slot aware contrastive learning framework to solve this problem, which proposes three methods to bridge the semantic gap between the dialogue context and the (domain, slot) by constructing training sample pairs to fine-tune the BERT model and use it for base DST model. The experiments demonstrate that our proposed method has improved the performance of the baseline model on the MultiWOZ2.1 and MultiWOZ2.4 datasets, yielding competitive results.
Haoxiang Su, Sijie Feng, Hongyan Xie, Di Wu 0088, Hao Huang 0009, Zhongjiang He, Shuangyong Song, Ruiyu Fang, Xiaomeng Huang, Wushour Slamu
ICASSP1
2024 Dual Level Intent-Slot Interaction for Improved Multi-Intent Spoken Language Understanding
abstract
Multi-intent spoken language understanding consists of two typical subtasks: multi-intent detection and slot filling. Existing approach suffers from two limitations: (1) It fails to explicitly model the information transfer between slots associated within the same intent clause; (2) Using a co-occurrence matrix of both label encodings introduces needless slot positional information such as the prefix ‘B-’ or ‘I-’. For (1), we propose a Gaussian Graph Attention Network that allows interaction to focus not only on the connection between slots within the current intent clause, but also on the connection between intent, and between intent and slot. For (2), we use a co-occurrence matrix of intent categories and slot types to model the knowledge transfer between the two subtasks in the corpus-level interaction, bypassing the introduction of slot positional information. Our framework achieves significant accuracy gains on both the MixATIS and MixSNIPS datasets.
Di Wu 0088, Liting Jiang, Haoxiang Su, Hao Huang 0009
ICASSP5
2024 Improving Pointer Network based Dialogue State Tracking via Dual Hierarchical Selective Augmentation
abstract
Dialogue state tracking is responsible for predicting the user’s dialogue state during the whole dialogue process. In practical applications, values for different slots exist in individual utterances of the dialog history. With the accumulation of the dialogue history, it becomes extremely difficult to accurately predict slots and corresponding values from the lengthy dialogue history. To solve the problem of the interference caused by lengthy dialogue history, we propose a dual hierarchical selective augmentation method, which makes use of two hierarchical level information selection strategy to generate slot values. In the encoding phase, we first extract word-level matching features between the slot and each dialogue turn, and then build turn-level context relevance. In the decoding phase, first of all, from a global perspective, the dialogue turn information is selected multiple according to the dialogue context and slot, so that the model focuses more on the turn containing slot value. Secondly, our model performs weighted context attention to capture the critical words of dialogue turn from the local view. This dual hierarchical context selection alleviates the interference caused by excessive redundant information in the dialogue history and enhances the judgment ability of the model for vital turns and words. Furthermore, to enhance the copying ability of the model, we use the turn selection-guided pointer network to copy slot values from the dialogue. Experimental results show that our model significantly outperforms multiple baselines on the released MultiWOZ benchmark.
Shuangyong Song, Hongyan Xie, Haoxiang Su, Hao Huang 0009, Mengxiang Li, Zhongjiang He, Ruiyu Fang
IJCNN3
2024 Graph-based Dynamic Domain Selection for Dialogue State Tracking
abstract
The Dialogue State Tracking (DST) module tracks the user’s intent by populating multiple predefined slots related to the dialogue task. In recent years, various graph neural network-based DST methods have been proposed to establish graph structures capturing the correlations between domains and slots, thereby enhancing model performance. However, these methods may involve redundant connections in the graph structure. To better construct relationships between domains and slots, we introduce a graph neural network-based dialogue state tracking method called Dynamic Domain Selection Graph DST (DDSG-DST). Specifically, (1) we employ Graphormer to establish hierarchical relationships between domains and slots; (2) we propose an additional domain prediction auxiliary task to predict the domain relevant to the dialogue context; (3) based on the predicted relevant domain from the auxiliary task, we dynamically select domain node information in the graph and perform dialogue state prediction. Experimental results demonstrate that we effectively establish hierarchical relationships between domains and slots, mitigate the negative impact of redundant connections in the graph structure, and enhance model performance.
Shuangyong Song, Hao Huang 0009, Hongyan Xie, Haoxiang Su, Mengxiang Li, Zhongjiang He, Ruiyu Fang
IJCNN5
2024 Monocular Ranging Based on Camera Pose in Visual Train Positioning
abstract
Real-time and accurate positioning is the key to ensuring the safe operation of trains. Traditional train positioning technology relies on track-side equipment, and there are problems such as high construction and maintenance costs and difficulties in obtaining the initial train positions. With the rapid development of enabling technologies such as AI and image recognition, visual perception-based train autonomous positioning technology has attracted widespread attention in recent years. To address the problem of large measurement errors due to camera pose changes in train visual positioning, this article proposes a monocular ranging method based on visual beacon imaging size correction. A kilometer post is used as a visual beacon, and a deep neural network is used to detect the visual beacon and extract its feature points' pixel coordinates. The PnP algorithm is used to estimate the camera pose and correct the visual beacon imaging size through the camera pose angle, the average relative error of ranging is stabilized at 2% ~ 3%, and the ranging effect is improved by about 60%, which effectively reducing the ranging error of the monocular ranging model based on known references.
Ming Chai, Jinke Shi, Jidong Lv, Haoxiang Su
SMC7
2023 Scalable-DSC: A Structural Template Prompt Approach to Scalable Dialogue State Correction
abstract
Dialogue state error correction has recently been proposed to correct wrong slot values in predicted dialogue states, thereby mitigating the error propagation problem for dialogue state tracking (DST).These approaches, though effective, are heavily intertwined with specific DST models, limiting their applicability to other DST models.To solve this problem, we propose Scalable Dialogue State Correction (Scalable-DSC), which can correct wrong slot values in the dialogue state predicted by any DST model.Specifically, we propose a Structural Template Prompt (STP) that converts predicted dialogue state from any DST models into a standardized natural language sequence as a part of the historical context, associates them with dialogue history information, and generates a corrected dialogue state sequence based on predefined template options.We further enhance Scalable-DSC by introducing two training strategies.The first employs a predictive state simulator to simulate the predicted dialogue states as the training data to enhance the generalization ability of the model.The second involves using the dialogue state predicted by DST as the training data, aiming at mitigating the inconsistent error type distribution between the training and inference.Experiments confirm that our model achieves state-of-the-art results on MultiWOZ 2.0-2.4 △ .
Haoxiang Su, Hongyan Xie, Shuangyong Song, Ruiyu Fang, Xiaomeng Huang, Sijie Feng
EMNLP1
2022 Correctable-DST: Mitigating Historical Context Mismatch between Training and Inference for Improved Dialogue State Tracking
abstract
Hongyan Xie, Haoxiang Su, Shuangyong Song, Hao Huang, Bo Zou, Kun Deng, Jianghua Lin, Zhihui Zhang, Xiaodong He. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Hongyan Xie, Haoxiang Su, Shuangyong Song, Jianghua Lin, Xiaodong He 0001
EMNLP2