VLDB 2026 Research / reviewers in the wild / expert
Yajing Sun
dblp:08/10952
· DBLP profile ↗
19ranked-venue papers
4as first author
13since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Learning to Know Myself: A Coarse-to-Fine Persona-Aware Training Framework for Personalized Dialogue GenerationabstractA critical challenge for open-domain dialogue agents is to generate persona-relevant and consistent responses. Due to the nature of persona sparsity in conversation scenarios, previous persona-based dialogue agents trained with Maximum Likelihood Estimation tend to overlook the given personas and generate responses irrelevant or inconsistent with personas. To address this problem, we propose a two-stage coarse-to-fine persona-aware training framework to improve the persona consistency of a dialogue agent progressively. Specifically, our framework first trains the dialogue agent to answer the constructed persona-aware questions, making it highly sensitive to the personas to generate persona-relevant responses. Then the dialogue agent is further trained with a contrastive learning paradigm by explicitly perceiving the difference between the consistent and the generated inconsistent responses, forcing it to pay more attention to the key persona information to generate consistent responses. By applying our proposed training framework to several representative baseline models, experimental results show significant boosts on both automatic and human evaluation metrics, especially the consistency of generated responses. Yunpeng Li 0006, Yue Hu 0002, Yajing Sun, Luxi Xing, Ping Guo 0002, Yuqiang Xie, Wei Peng 0008 |
AAAI | 3 |
| 2023 | Knowledge-Augmented Frame Semantic Parsing with Hybrid Prompt-TuningabstractFrame semantics-based approaches have been widely used in semantic parsing tasks and have become mainstream. It remains challenging to disambiguate frame representations evoked by target lexical units under different contexts. Pre-trained Language Models (PLMs) have been used in semantic parsing and significantly improve the accuracy of neural parsers. However, the PLMs-based approaches tend to favor collocated patterns presented in the training data, leading to inaccurate outcomes. The intuition here is to design a mechanism to optimally use knowledge captured in semantic frames in conjunction with PLMs to disambiguate frames. We propose a novel Knowledge-Augmented Frame Semantic Parsing Architecture (KAF-SPA) to enhance semantic representation by incorporating accurate frame knowledge into PLMs during frame semantic parsing. Specifically, a Memory-based Knowledge Extraction Module (MKEM) is devised to select accurate frame knowledge and construct the continuous templates in the high dimensional vector space. Moreover, we design a Task-oriented Knowledge Probing Module (TKPM) using hybrid prompts (in terms of continuous and discrete prompts) to incorporate the selected knowledge into the PLMs and adapt PLMs to the tasks of frame and argument identification. Experimental results on two public FrameNet datasets demonstrate that our method significantly outperforms strong baselines (by more than +3% in F1), achieving state-of-art results on the current benchmark. Ablation studies verify the effectiveness of KAF-SPA. Yajing Sun, Jingyuan Yang 0008, Wei Peng 0011 |
ICASSP | 2 |
| 2022 | CogIntAc: Modeling the Relationships between Intention, Emotion and Action in Interactive Process from Cognitive PerspectiveabstractIntention, emotion and action are important psychological factors in human activities, which play an important role in the interaction between individuals. How to model the interaction process between individuals by analyzing the relationship of their intentions, emotions, and actions at the cognitive level is challenging. In this paper, we propose a novel cognitive framework of individual interaction. The core of the framework is that individuals achieve interaction through external action driven by their inner intention. Based on this idea, the interactions between individuals can be constructed by establishing relationships between the intention, emotion and action. Furthermore, we conduct analysis on the interaction between individuals and give a reasonable explanation for the predicting results. To verify the effectiveness of the framework, we reconstruct a dataset and propose three tasks as well as the corresponding baseline models, including action abduction, emotion prediction and action generation. The novel framework shows an interesting perspective on mimicking the mental state of human beings in cognitive science. Wei Peng 0008, Yue Hu 0002, Yuqiang Xie, Luxi Xing, Yajing Sun |
CEC | 5 |
| 2022 | Modeling Intention, Emotion and External World in Dialogue SystemsabstractIntention, emotion and action are important elements in human activities. Modeling the interaction process between individuals by analyzing the relationships between these elements is a challenging task. However, previous work mainly focused on modeling intention and emotion independently, and neglected of exploring the mutual relationships between intention and emotion. In this paper, we propose a RelAtion Interaction Network (RAIN), consisting of Intention Relation Module and Emotion Relation Module, to jointly model mutual relationships and explicitly integrate historical intention information. The experiments on the dataset show that our model can take full advantage of the intention, emotion and action between individuals and achieve a remarkable improvement over BERT-style baselines. Qualitative analysis verifies the importance of the mutual interaction between the intention and emotion. Wei Peng 0008, Yue Hu 0002, Luxi Xing, Yuqiang Xie, Xingsheng Zhang, Yajing Sun |
ICASSP | 6 |
| 2022 | Control Globally, Understand Locally: A Global-to-Local Hierarchical Graph Network for Emotional Support ConversationabstractEmotional support conversation aims at reducing the emotional distress of the help-seeker, which is a new and challenging task. It requires the system to explore the cause of help-seeker's emotional distress and understand their psychological intention to provide supportive responses. However, existing methods mainly focus on the sequential contextual information, ignoring the hierarchical relationships with the global cause and local psychological intention behind conversations, thus leads to a weak ability of emotional support. In this paper, we propose a Global-to-Local Hierarchical Graph Network to capture the multi-source information (global cause, local intentions and dialog history) and model hierarchical relationships between them, which consists of a multi-source encoder, a hierarchical graph reasoner, and a global-guide decoder. Furthermore, a novel training objective is designed to monitor semantic information of the global cause. Experimental results on the emotional support conversation dataset, ESConv, confirm that the proposed GLHG has achieved the state-of-the-art performance on the automatic and human evaluations. Wei Peng 0008, Yue Hu 0002, Luxi Xing, Yuqiang Xie, Yajing Sun, Yunpeng Li 0006 |
IJCAI | 5 |
| 2022 | CogIntAc: Modeling the Relationships between Intention, Emotion and Action in Interactive Process from Cognitive PerspectiveabstractIntention, emotion and action are important psychological factors in human activities, which play an important role in the interaction between individuals. How to model the interaction process between individuals by analyzing the relationship of their intentions, emotions, and actions at the cognitive level is challenging. In this paper, we propose a novel cognitive framework of individual interaction. The core of the framework is that individuals achieve interaction through external action driven by their inner intention. Based on this idea, the interactions between individuals can be constructed by establishing relationships between the intention, emotion and action. Furthermore, we conduct analysis on the interaction between individuals and give a reasonable explanation for the predicting results. To verify the effectiveness of the framework, we reconstruct a dataset and propose three tasks as well as the corresponding baseline models, including action abduction, emotion prediction and action generation. The novel framework shows an interesting perspective on mimicking the mental state of human beings in cognitive science. Wei Peng 0008, Yue Hu 0002, Yuqiang Xie, Luxi Xing, Yajing Sun |
IJCNN | 5 |
| 2022 | KC2UM: Knowledge-Conversation Cyclic Utilization Mechanism for Knowledge-Grounded Dialogue GenerationabstractEnd-to-End open-domain dialogue systems suffer from the issues of generating inconsistent and repetitive responses. Existing dialogue models pay attention to unilaterally incorporating personalized knowledge into the dialogue to enhance the quality of generated response. However, they ignore that incorporating the personality-related information from dialogue history into personalized knowledge can boost the subsequent dialogue quality. In this paper, A Knowledge-Conversation Cyclic Utilization Mechanism (KC2UM) is proposed to enhance the dialogue quality. Specifically, A novel cyclic interaction module is designed to iteratively incorporate personalized knowledge into each turn conversation and capture the personality-related conversation information to enhance personalized knowledge semantic representation. We represent the knowledge with semantic and utilization representations to keep track of the personalized knowledge utilization. Experiments on two knowledge-grounded dialogue datasets show that our approach manages to select knowledge more accurately and generates more informative responses. Yajing Sun, Yue Hu 0002, Luxi Xing, Wei Peng 0008, Yuqiang Xie, Xingsheng Zhang |
IJCNN | 1 |
| 2022 | Exploiting Semantic and Syntactic Diversity for Diverse Task-oriented DialogueabstractTask-oriented dialogues have one-to-many property from semantic and syntactic perspectives, with many suitable dialogue acts and syntactic forms for a given post. However, current state-of-the-art task-oriented dialogue systems attempt to improve the quality of dialogues in terms of the most popular metrics (i.e., BLEU and entity F1), measuring the similarity between the generated responses and the human annotations, while the diversity of task-oriented dialogues remains less explored. This paper aims to improve the diversity of task-oriented dialogues from both semantic and syntactic perspectives by proposing a structural causal model to learn the causality composition of the dialogue acts and syntactic forms. Specifically, the disentangled understanding module decouples the dialogue into semantic and syntactic spaces and learns one-to-many property with multiple reference training. Then the casual collaboration generation module is proposed to apply Structural Causal Mechanism (SCM) to learn the causality composition relationship of the semantic and syntactic representations to generate the diverse response. Extensive experiments on the MultiWOZ datasets demonstrate that the proposed method achieves significantly better diversity than solid competitors. Yajing Sun, Yue Hu 0002, Luxi Xing, Yuqiang Xie, Wei Peng 0008, Yunpeng Li 0006 |
IJCNN | 1 |
| 2022 | Calibration of the Multiple Choice Machine Reading ComprehensionabstractPrediction calibration devotes to making the model produce correct prediction probability, corresponding with the model's empirical measurement (i.e., the accuracy on benchmark). It plays a crucial role in the Multiple-Choice based Machine Reading Comprehension (MCRC) task. Once the model gives the wrong candidate answer with a high probability, users or downstream applications will not trust the model easily. However, few works pay attention to the prediction calibration of MCRC models. In this paper, we study the prediction calibration of the MCRC models and introduce the self-supervised target label softening (SS-TLS) training method to develop a well-calibrated MRC model while improving its performance. Specifically, the proposed SS-TLS method softens the target label to train the MCRC model instead of the standard cross-entropy objective. It employs the self-supervised confidence signal to monitor the soften scale adaptively at the instance level. Experimental results on several multiple-Choice style MRC datasets illustrate that the proposed method can improve both model prediction calibration and performance. Luxi Xing, Yue Hu 0002, Yuqiang Xie, Wei Peng 0008, Yajing Sun, Chen Zhang 0001 |
IJCNN | 5 |
| 2022 | Document-Level Multi-event Extraction via Event Ontology Guiding
Xingsheng Zhang, Yue Hu 0002, Yajing Sun, Luxi Xing, Yuqiang Xie, Yunpeng Li 0006, Wei Peng 0008 |
KSEM (2) | 3 |
| 2022 | Do You Know My Emotion? Emotion-Aware Strategy Recognition Towards a Persuasive Dialogue System
Wei Peng 0008, Yue Hu 0002, Luxi Xing, Yuqiang Xie, Yajing Sun |
ECML/PKDD (2) | 5 |
| 2021 | Unsupervised Learning of Deterministic Dialogue Structure with Edge-Enhanced Graph Auto-EncoderabstractIt is important for task-oriented dialogue systems to discover the dialogue structure (i.e. the general dialogue flow) from dialogue corpora automatically. Previous work models dialogue structure by extracting latent states for each utterance first and then calculating the transition probabilities among states. These two-stage methods ignore the contextual information when calculating the probabilities, which makes the transitions between the states ambiguous. This paper proposes a conversational graph (CG) to represent deterministic dialogue structure where nodes and edges represent the utterance and context information respectively. An unsupervised Edge-Enhanced Graph Auto-Encoder (EGAE) architecture is designed to model local-contextual and global-structural information for conversational graph learning. Furthermore, a self-supervised objective is introduced with the response selection task to guide the unsupervised learning of the dialogue structure. Experimental results on several public datasets demonstrate that the novel model outperforms several alternatives in aggregating utterances with similar semantics. The effectiveness of the learned dialogue structured is also verified by more than 5\% joint accuracy improvement in the downstream task of low resource dialogue state tracking. Yajing Sun, Yong Shan, Chengguang Tang, Yue Hu 0002, Yinpei Dai, Jing Yu 0007, Jian Sun 0021, Fei Huang 0002, Luo Si |
AAAI | 1 |
| 2021 | MCR-NET: A Multi-Step Co-Interactive Relation Network for Unanswerable Questions on Machine Reading ComprehensionabstractQuestion answering systems usually use keyword searches to retrieve potential passages related to a question, and then extract the answer from passages with the machine reading comprehension methods. However, many questions tend to be unanswerable in the real world. In this case, it is significant and challenging how the model determines when no answer is supported by the passage and abstains from answering. Most of the existing systems design a simple classifier to determine answerability implicitly without explicitly modeling mutual interaction and relation between the question and passage, leading to the poor performance for determining the unanswerable questions. To tackle this problem, we propose a Multi-Step Co-Interactive Relation Network (MCR-Net) to explicitly model the mutual interaction and locate key clues from coarse to fine by introducing a co-interactive relation module. The co-interactive relation module contains a stack of interaction and fusion blocks to continuously integrate and fuse history-guided and current-query-guided clues in an explicit way. Experiments on the SQuAD 2.0 and DuReader datasets show that our model achieves a remarkable improvement, outperforming the BERT-style baselines in literature. Visualization analysis also verifies the importance of the mutual interaction between the question and passage. Wei Peng 0008, Yue Hu 0002, Jing Yu 0007, Luxi Xing, Yuqiang Xie, Yajing Sun |
ICASSP | 7 |
| 2020 | History-Adaption Knowledge Incorporation Mechanism for Multi-Turn Dialogue SystemabstractKeeping the conversation consistent and avoiding its repetition are two key factors to construct an intelligent multi-turn knowledge-grounded dialogue system. Although some works tend to combine history with external knowledge such as personal background information to boost dialogue quality, they are prone to ignore the fact that incorporating the same knowledge multiple times into the conversation leads to repetition. The main reason is the lack of effective control over the use of knowledge on the conversation level. So we design a history-adaption knowledge incorporation mechanism to build an effective multi-turn dialogue model. Our proposed model addresses repetition by recurrently updating the knowledge from the conversation level and progressively incorporating it into the history step-by-step. And the knowledge-grounded history representation also enhances the conversation consistency. Experimental results show that our proposed model significantly outperforms several retrieval-based models on some benchmark datasets. The human evaluation demonstrates that our model can maintain conversation consistent and reduce conversation repetition. Yajing Sun, Yue Hu 0002, Luxi Xing, Jing Yu 0007, Yuqiang Xie |
AAAI | 1 |
| 2020 | Bi-directional CognitiveThinking Network for Machine Reading ComprehensionabstractWe propose a novel Bi-directional Cognitive Knowledge Framework (BCKF) for reading comprehension from the perspective of complementary learning systems theory. It aims to simulate two ways of thinking in the brain to answer questions, including reverse thinking and inertial thinking. To validate the effectiveness of our framework, we design a corresponding Bi-directional Cognitive Thinking Network (BCTN) to encode the passage and generate a question (answer) given an answer (question) and decouple the bi-directional knowledge. The model has the ability to reverse reasoning questions which can assist inertial thinking to generate more accurate answers. Competitive improvement is observed in DuReader dataset, confirming our hypothesis that bi-directional knowledge helps the QA task. The novel framework shows an interesting perspective on machine reading comprehension and cognitive science. Wei Peng 0008, Yue Hu 0002, Luxi Xing, Yuqiang Xie, Jing Yu 0007, Yajing Sun, Xiangpeng Wei |
COLING | 6 |
| 2020 | DAM: Deliberation, Abandon and Memory Networks for Generating Detailed and Non-repetitive Responses in Visual DialogueabstractVisual Dialogue task requires an agent to be engaged in a conversation with human about an image. The ability of generating detailed and non-repetitive responses is crucial for the agent to achieve human-like conversation. In this paper, we propose a novel generative decoding architecture to generate high-quality responses, which moves away from decoding the whole encoded semantics towards the design that advocates both transparency and flexibility. In this architecture, word generation is decomposed into a series of attention-based information selection steps, performed by the novel recurrent Deliberation, Abandon and Memory (DAM) module. Each DAM module performs an adaptive combination of the response-level semantics captured from the encoder and the word-level semantics specifically selected for generating each word. Therefore, the responses contain more detailed and non-repetitive descriptions while maintaining the semantic accuracy. Furthermore, DAM is flexible to cooperate with existing visual dialogue encoders and adaptive to the encoder structures by constraining the information selection mode in DAM. We apply DAM to three typical encoders and verify the performance on the VisDial v1.0 dataset. Experimental results show that the proposed models achieve new state-of-the-art performance with high-quality responses. The code is available at https://github.com/JXZe/DAM. Xiaoze Jiang, Jing Yu 0007, Yajing Sun, Zengchang Qin, Yue Hu 0002, Qi Wu 0001 |
IJCAI | 3 |
| 2020 | Mucko: Multi-Layer Cross-Modal Knowledge Reasoning for Fact-based Visual Question AnsweringabstractFact-based Visual Question Answering (FVQA) requires external knowledge beyond the visible content to answer questions about an image. This ability is challenging but indispensable to achieve general VQA. One limitation of existing FVQA solutions is that they jointly embed all kinds of information without fine-grained selection, which introduces unexpected noises for reasoning the final answer. How to capture the question-oriented and information-complementary evidence remains a key challenge to solve the problem. In this paper, we depict an image by a multi-modal heterogeneous graph, which contains multiple layers of information corresponding to the visual, semantic and factual features. On top of the multi-layer graph representations, we propose a modality-aware heterogeneous graph convolutional network to capture evidence from different layers that is most relevant to the given question. Specifically, the intra-modal graph convolution selects evidence from each modality and cross-modal graph convolution aggregates relevant information across different graph layers. By stacking this process multiple times, our model performs iterative reasoning across three modalities and predicts the optimal answer by analyzing all question-oriented evidence. We achieve a new state-of-the-art performance on the FVQA task and demonstrate the effectiveness and interpretability of our model with extensive experiments. Jing Yu 0007, Yujing Wang 0002, Yajing Sun, Yue Hu 0002, Qi Wu 0001 |
IJCAI | 4 |
| 2020 | Enhancing Pre-trained Language Models by Self-supervised Learning for Story Cloze Test
Yuqiang Xie, Yue Hu 0002, Luxi Xing, Xiangpeng Wei, Yajing Sun |
KSEM (1) | 7 |
| 2020 | KBGN: Knowledge-Bridge Graph Network for Adaptive Vision-Text Reasoning in Visual DialogueabstractVisual dialogue is a challenging task that needs to extract implicit information from both visual (image) and textual (dialogue history) contexts. Classical approaches pay more attention to the integration of the current question, vision knowledge and text knowledge, despising the heterogeneous semantic gaps between the cross-modal information. In the meantime, the concatenation operation has become de-facto standard to the cross-modal information fusion, which has a limited ability in information retrieval. In this paper, we propose a novel Knowledge-Bridge Graph Network (KBGN) model by using graph to bridge the cross-modal semantic relations between vision and text knowledge in fine granularity, as well as retrieving required knowledge via an adaptive information selection mode. Moreover, the reasoning clues for visual dialogue can be clearly drawn from intra-modal entities and inter-modal bridges. Experimental results on VisDial v1.0 and VisDial-Q datasets demonstrate that our model outperforms existing models with state-of-the-art results. Xiaoze Jiang, Siyi Du, Zengchang Qin, Yajing Sun, Jing Yu 0007 |
ACM Multimedia | 4 |