VLDB 2026 Research / reviewers in the wild / expert
Yuke Si
dblp:199/7616
· DBLP profile ↗
12ranked-venue papers
3as first author
9since 2021 · last 2024
0009-0009-7170-1916ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Prompt-Based Hierarchical Pipeline for Cross-Domain Slot FillingabstractIn task-oriented dialogue systems, slot filling aims to identify the semantic slot types of each token in user utterances. Due to the lack of sufficient supervised data in many scenarios, it is necessary to transfer relevant knowledge by using cross-domain slot filling. Previous studies rely on manually additional meta-information to build the relationships among similar slots across domains, yet not fully utilizing the knowledge learned by language models in the pre-training stage. In this study, we propose a prompt-based hierarchical pipeline (PHP) with three innovations. First, we design a hierarchical pipeline to separately model domain-independent syntactic structures and domain-specific semantic structures, i.e., span detection and slot prediction. Second, we improve the prompt paradigm with discriminative structure to fully utilize pre-trained language models, which reformulates downstream tasks into pre-trained tasks. Finally, we polish our template and verbalizer to effectively utilize task-specific prior knowledge, adding some meta-information and updating their additional trainable parameters. We conducted extensive experiments on three datasets to evaluate our method, and experimental results show that our method significantly outperforms the previous state-of-the-art results. Yuke Si, Longbiao Wang, Xiaobao Wang, Jianwu Dang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Local and Global Context Modeling with Relation Matching Task for Dialog Act RecognitionabstractIn dialog act recognition (DAR) of an utterance in a conversation, the prior studies have focused either on the global context using the whole utterances in the dialog, or the local context using the neighbouring utterance flow in the dialog. However, their methods attempt to deal with all types of dialogs indiscriminately. In this study, we propose a model to extract the local context information by an inter-utterance relation matching task (RMT), and a DAR framework to incorporate the local context information into a hierarchical network to fulfil both local and global context modeling. Extensive evaluations were conducted on a Mandarin dialog corpus and two benchmark English corpora. It is found that the different dialog types possess different window lengths for RMT, which is related to the length of subtopics in a given type of dialog. According to ablation experiments, the global information contributed more to the DAR in the hierarchical framework, while the contribution ratio of the local to the global context information was larger than 0.1. The results demonstrated that the proposed RMT and DAR framework significantly improved the DAR performance. Yuke Si, Yan Zhang 0004, Xiaobao Wang, Longbiao Wang, Jianwu Dang 0001, Chng Eng Siong, Haizhou Li 0001 |
IJCNN | 1 |
| 2023 | Improving Zero-shot Cross-domain Slot Filling via Transformer-based Slot Semantics Fusion
Yuke Si, Longbiao Wang, Xiaobao Wang, Jianwu Dang 0001 |
INTERSPEECH | 3 |
| 2022 | Cache: Modeling Contribution-Aware Context Hierarchically for Long-Range Dialogue State TrackingabstractRecently, many studies on dialogue state tracking (DST) based on the copy-augmented encoder-decoder framework have been proposed and have achieved encouraging performance. However, these studies commonly lose earlier information during encoding the long dialogues with RNNs, and have difficulty for the decoder to focus on specific dialogue turns from lengthy context, which causes decreased performance as the dialogue gets longer. In this work, we propose a novel method to model Contribution-Aware Context HiErarchically (CACHE) with a hierarchical encoder and a slot-turn attention module. The hierarchical encoder is designed to prevent information loss by reducing the length of the sequence sent to each encoder. The slot-turn attention module is explored to help the decoder focus on the slotrelated dialogue turn information. To evaluate models more appropriately, we introduce a new metric continued joint accuracy considering the prediction accuracy of both current and historical dialogue turns. Experiments on MultiWOZ 2.0 show that CACHE is an effective model for tracking states especially in long context. Jianshu Qi, Yuke Si, Longbiao Wang, Jianwu Dang 0001 |
ICASSP | 2 |
| 2022 | Speech Emotion Recognition with Co-Attention Based Multi-Level Acoustic InformationabstractSpeech Emotion Recognition (SER) aims to help the machine to understand human’s subjective emotion from only audio in-formation. However, extracting and utilizing comprehensive in-depth audio information is still a challenging task. In this paper, we propose an end-to-end speech emotion recognition system using multi-level acoustic information with a newly designed co-attention module. We firstly extract multi-level acoustic information, including MFCC, spectrogram, and the embedded high-level acoustic information with CNN, BiL-STM and wav2vec2, respectively. Then these extracted features are treated as multimodal inputs and fused by the pro-posed co-attention mechanism. Experiments are carried on the IEMOCAP dataset, and our model achieves competitive performance with two different speaker-independent cross-validation strategies. Our code is available on GitHub. Heqing Zou, Yuke Si, Chen Chen 0075, Deepu Rajan, Chng Eng Siong |
ICASSP | 2 |
| 2022 | TopicKS: Topic-driven Knowledge Selection for Knowledge-grounded Dialogue Generation
Shiquan Wang, Yuke Si, Longbiao Wang, Zhiqiang Zhuang, Xiaowang Zhang, Jianwu Dang 0001 |
INTERSPEECH | 2 |
| 2022 | Hierarchical Tagger with Multi-task Learning for Cross-domain Slot Filling
Yuke Si, Shiquan Wang, Longbiao Wang, Jianwu Dang 0001 |
INTERSPEECH | 2 |
| 2021 | Domain-Specific Multi-Agent Dialog Policy Learning in Multi-Domain Task-Oriented Scenarios
Yuke Si, Longbiao Wang, Jianwu Dang 0001 |
Interspeech | 2 |
| 2021 | Exploiting Explicit and Inferred Implicit Personas for Multi-turn Dialogue Generation
Ruifang He, Longbiao Wang, Yuke Si, Jianwu Dang 0001 |
NLPCC (1) | 4 |
| 2020 | A Hierarchical Model for Dialog Act Recognition Considering Acoustic and Lexical Context InformationabstractDialog act recognition (DAR) is important to capture speakers' intention in a dialog system. Traditional methods commonly use the lexical information from transcripts, acoustic information from speech, and dialog context information to do DAR. However, in these methods, textual context information may be considered, whereas acoustic context information is ignored, which leads to ambiguity in certain DAs especially in Mandarin. To solve the problem, we propose a hierarchical model for DAR considering context information of both lexical and acoustic prosody. The experimental results on a Mandarin dialog corpus demonstrate that the contextual-acoustic information is helpful for recognizing DAs. The contextually specific prosodies involved in the utterances such as the echo question and open-end question are beneficial to identify the users' intention. We also investigate the effect of the context length on the DAR. The proper context length is approximately equal to the length of the entire subtopics. Yuke Si, Longbiao Wang, Jianwu Dang 0001, Mengfei Wu |
ICASSP | 1 |
| 2020 | Adversarial Shared-Private Attention Network for Joint Slot Filling and Intent Detection
Mengfei Wu, Longbiao Wang, Yuke Si, Jianwu Dang 0001 |
ICONIP (4) | 3 |
| 2019 | CNN-BLSTM Based Question Detection from Dialogs Considering Phase and Context Information
Yuke Si, Longbiao Wang, Jianwu Dang 0001, Mengfei Wu |
INTERSPEECH | 1 |