VLDB 2026 Research / reviewers in the wild / expert
Xiuyi Chen
dblp:218/7190
· DBLP profile ↗
23ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0002-9351-4160ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TinyChemVL: Advancing Chemical Vision-Language Models via Efficient Visual Token Reduction and Complex Reaction TasksabstractWhile Vision Language Models (VLMs) have demonstrated remarkable capabilities in general visual understanding, their application in the chemical domain has been limited, with previous works predominantly focusing on text and thus overlooking critical visual information, such as molecular structures. Current approaches that directly adopt standard VLMs for chemical tasks suffer from two primary issues: (i) computational inefficiency of processing entire chemical images with non-informative backgrounds. (ii) a narrow scope on molecular-level tasks that restricts progress in chemical reasoning. In this work, we propose TinyChemVL, an efficient and powerful chemical VLM that leverages visual token reduction and reaction-level tasks to improve model efficiency and reasoning capacity. Also, we propose ChemRxn-V, a reaction-level benchmark for assessing vision-based reaction recognition and prediction tasks. Directly predicting reaction products from molecular images poses a non-trivial challenge, as it requires models to integrate both recognition and reasoning capacities. Our results demonstrate that, with only 4B parameters, TinyChemVL achieves superior performance on both molecular and reaction tasks, while also demonstrating faster inference and training speeds compared to existing models. Notably, TinyChemVL outperforms ChemVLM while utilizing only 1/16th of the visual tokens. This work builds efficient yet powerful VLMs for chemical domains by co-designing model architecture and task complexity. Xuanle Zhao, Shuxin Zeng, Xinyuan Cai, Duzhen Zhang, Xiuyi Chen, Bo Xu 0002 |
AAAI | 6 |
| 2026 | From System 1 to System 2: A Survey of Reasoning Large Language ModelsabstractAchieving human-level intelligence requires refining the transition from the fast, intuitive System 1 to the slower, more deliberate System 2 reasoning. While System 1 excels in quick, heuristic decisions, System 2 relies on logical reasoning for more accurate judgments and reduced biases. Foundational Large Language Models (LLMs) excel at fast decision-making but lack the depth for complex reasoning, as they have not yet fully embraced the step-by-step analysis characteristic of true System 2 thinking. Recently, reasoning LLMs like OpenAI's o1/o3 and DeepSeek's R1 have demonstrated expert-level performance in fields such as mathematics and coding, closely mimicking the deliberate reasoning of System 2 and showcasing human-like cognitive abilities. This survey begins with a brief overview of the progress in foundational LLMs and the early development of System 2 technologies, exploring how their combination has paved the way for reasoning LLMs. Next, we discuss how to construct reasoning LLMs, trace the evolution of various reasoning models, and examine the core methods that enable advanced reasoning behind them. Additionally, we provide an overview of reasoning benchmarks, offering an in-depth comparison of the performance of representative reasoning LLMs. Finally, we explore promising directions for advancing reasoning LLMs and maintain a real-time GitHub Repository to track the latest developments. We hope this survey will serve as a valuable resource to inspire innovation and drive progress in this rapidly evolving field. Duzhen Zhang, Zhongzhi Li, Jiaxin Zhang 0024, Zengyan Liu, Junhao Zheng, Xiuyi Chen, Jiahua Dong 0001, Zhijiang Guo, Cheng-Lin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2025 | Tool Learning in the Wild: Empowering Language Models as Automatic Tool AgentsabstractAugmenting large language models (LLMs) with external tools has emerged as a promising approach to extend their utility, enabling them to solve practical tasks.Previous methods manually parse tool documentation and create in-context demonstrations, transforming tools into structured formats for LLMs to use in their step-by-step reasoning.However, this manual process requires domain expertise and struggles to scale to large toolsets.Additionally, these methods rely heavily on ad-hoc inference techniques or special tokens to integrate free-form LLM generation with tool-calling actions, limiting the LLM's flexibility in handling diverse tool specifications and integrating multiple tools.In this work, we propose AutoTools, a framework that enables LLMs to automate the tool-use workflow.Specifically, the LLM automatically transforms tool documentation into callable functions, verifying syntax and runtime correctness.Then, the LLM integrates these functions into executable programs to solve practical tasks, flexibly grounding tool-use actions into its reasoning processes.Extensive experiments on existing and newly collected, more challenging benchmarks illustrate the superiority of our framework.Inspired by these promising results, we further investigate how to improve the expertise of LLMs, especially opensource LLMs with fewer parameters, within AutoTools.Thus, we propose the AutoTools-Learning approach, training the LLMs with three learning tasks on 34k instances of high-quality synthetic data, including documentation understanding, relevance learning, and function programming.Fine-grained results validate the effectiveness of our overall training approach and each individual task. Zhengliang Shi, Shen Gao, Lingyong Yan, Yue Feng 0002, Xiuyi Chen, Zhumin Chen, Dawei Yin 0001, Suzan Verberne, Zhaochun Ren |
WWW | 5 |
| 2024 | Consecutive knowledge meta-adaptation learning for unsupervised medical diagnosis
Hongliu Li, Yawen Hou, Xiuyi Chen, Hongyuan Yu |
Knowl. Based Syst. | 4 |
| 2024 | Category-Contextual Relation Encoding Network for Few-Shot Object DetectionabstractFew-shot object detection (FSOD) has brought increasing academic interest by recognizing previously unseen novel classes with very limited well-labeled samples. However, most existing methods identify novel classes via some object-specific characteristics in the few provided samples rather than intrinsic inter-class relations between base and novel classes, which heavily degrades the detection performance on novel classes. Moreover, they cannot learn discriminative proposal representations to distinguish base and novel classes, and thus misclassify novel objects as confusable base classes. To tackle the above challenges, we develop a novel Category-contextual Relation Encoding Network (CRE-Net), which is an early attempt to reason inter-class context relationships for FSOD task. To be specific, we propose a novel category-contextual relation encoding mechanism to capture intrinsic inter-class relations between base and novel classes via knowledge aggregation from global category-contextual descriptors. It utilizes intrinsic inter-class contextual relations to adaptively refine the convolution kernel, thus encoding the local semantic context of query image with category-contextual relation as guidance. Furthermore, to explore discriminative representations for base and novel classes, we develop a scarcity-compensatory contrastive proposal loss by incorporating data scarcity of novel classes and proposal semantic consistency with high confidence. This loss could compact object instances from the same category to a tighter cluster, and enhance the space separability of different classes. Extensive experiments on Pascal VOC and COCO datasets verify the state-of-the-art detection performance of our CRE-Net model when compared with other baseline methods. Ating Yin, Yaonan Wang 0001, Jianxu Mao, Hui Zhang 0023, Xiuyi Chen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Structure Aware Multi-Graph Network for Multi-Modal Emotion Recognition in ConversationsabstractMulti-Modal Emotion Recognition in Conversations (MMERC) is an increasingly active research field that leverages multi-modal signals to understand the feelings behind each utterance. Modeling contextual interactions and multi-modal fusion lie at the heart of this field, with graph-based models recently being widely used for MMERC to capture global multi-modal contextual information. However, these models generally mix all modality representations in a single graph, and utterances in each modality are fully connected, potentially ignoring three problems: 1) the heterogeneity of the multi-modal context, 2) the redundancy of contextual information, and 3) over-smoothing of the graph networks. To address these problems, we propose a Structure Aware Multi-Graph Network (SAMGN) for MMERC. Specifically, we construct multiple modality-specific graphs to model the heterogeneity of the multi-modal context. Instead of fully connecting the utterances in each modality, we design a structure learning module that determines whether edges exist between the utterances. This module reduces redundancy by forcing each utterance to focus on the contextual ones that contribute to its emotion recognition, acting like a message propagating reducer to alleviate over-smoothing. Then, we develop the SAMGN via Dual-Stream Propagation (DSP), which contains two propagation streams, i.e., intra- and inter-modal, performed in parallel to aggregate the heterogeneous modality information from multi-graphs. DSP also contains a gating unit that adaptively integrates the co-occurrence information from the above two propagations for emotion recognition. Experiments on two popular MMERC datasets demonstrate that SAMGN achieves new State-Of-The-Art (SOTA) results. Duzhen Zhang, Jianlong Chang, Xiuyi Chen, Qi Tian 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | DualGATs: Dual Graph Attention Networks for Emotion Recognition in ConversationsabstractCapturing complex contextual dependencies plays a vital role in Emotion Recognition in Conversations (ERC).Previous studies have predominantly focused on speaker-aware context modeling, overlooking the discourse structure of the conversation.In this paper, we introduce Dual Graph ATtention networks (Dual-GATs) to concurrently consider the complementary aspects of discourse structure and speaker-aware context, aiming for more precise ERC.Specifically, we devise a Discourseaware GAT (DisGAT) module to incorporate discourse structural information by analyzing the discourse dependencies between utterances.Additionally, we develop a Speakeraware GAT (SpkGAT) module to incorporate speaker-aware contextual information by considering the speaker dependencies between utterances.Furthermore, we design an interaction module that facilitates the integration of the DisGAT and SpkGAT modules, enabling the effective interchange of relevant information between the two modules.We extensively evaluate our method on four datasets, and experimental results demonstrate that our proposed DualGATs surpass state-of-the-art baselines on the majority of the datasets.1 Duzhen Zhang, Xiuyi Chen |
ACL (1) | 3 |
| 2023 | Task Relation Distillation and Prototypical Pseudo Label for Incremental Named Entity RecognitionabstractIncremental Named Entity Recognition (INER) involves the sequential learning of new entity types without accessing the training data of previously learned types. However, INER faces the challenge of catastrophic forgetting specific for incremental learning, further aggravated by background shift (i.e., old and future entity types are labeled as the non-entity type in the current task). To address these challenges, we propose a method called task Relation Distillation and Prototypical pseudo label (RDP) for INER. Specifically, to tackle catastrophic forgetting, we introduce a task relation distillation scheme that serves two purposes: 1) ensuring inter-task semantic consistency across different incremental learning tasks by minimizing inter-task relation distillation loss, and 2) enhancing the model's prediction confidence by minimizing intra-task self-entropy loss. Simultaneously, to mitigate background shift, we develop a prototypical pseudo label strategy that distinguishes old entity types from the current non-entity type using the old model. This strategy generates high-quality pseudo labels by measuring the distances between token embeddings and type-wise prototypes. We conducted extensive experiments on ten INER settings of three benchmark datasets (i.e., CoNLL2003, I2B2, and OntoNotes5). The results demonstrate that our method achieves significant improvements over the previous state-of-the-art methods, with an average increase of 6.08% in Micro F1 score and 7.71% in Macro F1 score. Duzhen Zhang, Hongliu Li, Wei Cong, Rongtao Xu, Jiahua Dong 0001, Xiuyi Chen |
CIKM | 6 |
| 2023 | Continual Named Entity Recognition without Catastrophic ForgettingabstractContinual Named Entity Recognition (CNER) is a burgeoning area, which involves updating an existing model by incorporating new entity types sequentially.Nevertheless, continual learning approaches are often severely afflicted by catastrophic forgetting.This issue is intensified in CNER due to the consolidation of old entity types from previous steps into the non-entity type at each step, leading to what is known as the semantic shift problem of the non-entity type.In this paper, we introduce a pooled feature distillation loss that skillfully navigates the trade-off between retaining knowledge of old entity types and acquiring new ones, thereby more effectively mitigating the problem of catastrophic forgetting.Additionally, we develop a confidence-based pseudo-labeling for the non-entity type, i.e., predicting entity types using the old model to handle the semantic shift of the non-entity type.Following the pseudo-labeling process, we suggest an adaptive re-weighting type-balanced learning strategy to handle the issue of biased type distribution.We carried out comprehensive experiments on ten CNER settings using three different datasets.The results illustrate that our method significantly outperforms prior state-of-the-art approaches, registering an average improvement of 6.3% and 8.0% in Micro and Macro F1 scores, respectively.1 * Equal contributions.† The corresponding author is Dr. Duzhen Zhang, Wei Cong, Jiahua Dong 0001, Yahan Yu, Xiuyi Chen, Yonggang Zhang 0003, Zhen Fang 0001 |
EMNLP | 5 |
| 2023 | Matching-Based Term Semantics Pre-Training for Spoken Patient Query UnderstandingabstractMedical Slot Filling (MSF) task aims to convert medical queries into structured information, playing an essential role in diagnosis dialogue systems. However, the lack of sufficient term semantics learning makes existing approaches hard to capture semantically identical but colloquial expressions of terms in medical conversations. In this work, we formalize MSF into a matching problem and propose a Term Semantics Pre-trained Matching Network (TSPMN) that takes both terms and queries as input to model their semantic inter-action. To learn term semantics better, we further design two self-supervised objectives, including Contrastive Term Discrimination (CTD) and Matching-based Mask Term Modeling (MMTM). CTD determines whether it is the masked term in the dialogue for each given term, while MMTM directly predicts the masked ones. Experimental results on two Chinese benchmarks show that TSPMN outperforms strong baselines, especially in few-shot settings1. Zefa Hu, Xiuyi Chen, Minglun Han, Ziyi Ni, Jing Shi 0003, Bo Xu 0002 |
ICASSP | 2 |
| 2023 | Crucial Semantic Classifier-based Adversarial Learning for Unsupervised Domain AdaptationabstractUnsupervised Domain Adaptation (UDA), which aims to explore the transferrable features from a well-labeled source domain to a related unlabeled target domain, has been widely progressed. Nevertheless, as one of the mainstream, existing adversarial-based methods neglect to filter the irrelevant semantic knowledge, hindering adaptation performance improvement. Besides, they require an additional domain discriminator that strives extractor to generate confused representations, but discrete designing may cause model collapse. To tackle the above issues, we propose Crucial Semantic Classifier-based Adversarial Learning (CSCAL), which pays more attention to crucial semantic knowledge transferring and leverages the classifier to implicitly play the role of domain discriminator without extra network designing. CSCAL effectively mitigates distribution shifts between the source and target domains from both intra- and inter-class perspectives. Specifically, in intra-class-wise alignment, a Paired-Level Discrepancy (PLD) is designed to transfer crucial semantic knowledge. Additionally, based on classifier predictions, a Nuclear Norm-based Discrepancy (NND) is formed that considers inter-class-wise information and improves the adaptation performance. Moreover, CSCAL can be effortlessly merged into different UDA methods as a regularizer and dramatically promote their performance. Yajun Gao, Hongliu Li, Ating Yin, Duzhen Zhang, Xiuyi Chen |
IJCNN | 6 |
| 2023 | Decomposing Logits Distillation for Incremental Named Entity RecognitionabstractIncremental Named Entity Recognition (INER) aims to continually train a model with new data, recognizing emerging entity types without forgetting previously learned ones. Prior INER methods have shown that Logits Distillation (LD), which involves preserving predicted logits via knowledge distillation, effectively alleviates this challenging issue. In this paper, we discover that a predicted logit can be decomposed into two terms that measure the likelihood of an input token belonging to a specific entity type or not. However, the traditional LD only preserves the sum of these two terms without considering the change in each component. To explicitly constrain each term, we propose a novel Decomposing Logits Distillation (DLD) method, enhancing the model's ability to retain old knowledge and mitigate catastrophic forgetting. Moreover, DLD is model-agnostic and easy to implement. Extensive experiments show that DLD consistently improves the performance of state-of-the-art INER methods across ten INER settings in three datasets. Duzhen Zhang, Yahan Yu, Xiuyi Chen |
SIGIR | 4 |
| 2023 | Lifelong Text-Audio Sentiment Analysis learning
Xiuyi Chen, Zhongshi He |
Neural Networks | 3 |
| 2022 | TSAM: A Two-Stream Attention Model for Causal Emotion EntailmentabstractCausal Emotion Entailment (CEE) aims to discover the potential causes behind an emotion in a conversational utterance. Previous works formalize CEE as independent utterance pair classification problems, with emotion and speaker information neglected. From a new perspective, this paper considers CEE in a joint framework. We classify multiple utterances synchronously to capture the correlations between utterances in a global view and propose a Two-Stream Attention Model (TSAM) to effectively model the speaker’s emotional influences in the conversational history. Specifically, the TSAM comprises three modules: Emotion Attention Network (EAN), Speaker Attention Network (SAN), and interaction module. The EAN and SAN incorporate emotion and speaker information in parallel, and the subsequent interaction module effectively interchanges relevant information between the EAN and SAN via a mutual BiAffine transformation. Extensive experimental results demonstrate that our model achieves new State-Of-The-Art (SOTA) performance and outperforms baselines remarkably. Duzhen Zhang, Fandong Meng, Xiuyi Chen, Jie Zhou 0016 |
COLING | 4 |
| 2022 | A Multi Domain Knowledge Enhanced Matching Network for Response Selection in Retrieval-Based Dialogue SystemsabstractBuilding a human-machine conversational agent is a core problem in Artificial Intelligence, where knowledge has to be integrated into the model effectively. In this paper, we propose a Multi Domain Knowledge Enhanced Matching Network (MDKEMN) to build retrievalbased dialogue systems that could leverage both explicit knowledge graph and implicit domain knowledge for response selection. Specifically, our MDKEMN leverages the self-attention mechanism of a single-stream Transformer to make deep interactions among the dialogue context, response candidate and external knowledge graph, and finally returns the matching degree of each context-response pair under the external knowledge. Furthermore, to leverage the implicit domain knowledge from all domains to improve the performance of each domain, we combine the multi-domain datasets for training and then finetune the pretrained model on each domain. Experimental results show (1) the effectiveness of both explicit and implicit knowledge incorporating and (2) the superiority of our approach over previous baselines on a Chinese multi-domain knowledge-driven dialogue dataset. Xiuyi Chen, Bo Xu 0002 |
ICASSP | 1 |
| 2022 | Improving Cross-Modal Understanding in Visual Dialog Via Contrastive LearningabstractVisual Dialog is a challenging vision-language task since the visual dialog agent needs to answer a series of questions after reasoning over both the image content and dialog history. Though existing methods try to deal with the cross-modal understanding in visual dialog, they are still not enough in ranking candidate answers based on their understanding of visual and textual contexts. In this paper, we analyze the cross-modal understanding in visual dialog based on the vision-language pre-training model VD-BERT and propose a novel approach to improve the cross-modal understanding for visual dialog, named ICMU. ICMU enhances cross-modal understanding by distinguishing different pulled inputs (i.e. pulled images, questions or answers) based on four-way contrastive learning. In addition, ICMU exploits the single-turn visual question answering to enhance the visual dialog model’s cross-modal understanding to handle a multi-turn visually-grounded conversation. Experiments show that the proposed approach improves the visual dialog model’s cross-modal understanding and brings satisfactory gain to the Vis-Dial dataset. Xiuyi Chen, Bo Xu 0002 |
ICASSP | 2 |
| 2022 | Unsupervised and Pseudo-Supervised Vision-Language Alignment in Visual DialogabstractVisual dialog requires models to give reasonable answers according to a series of coherent questions and related visual concepts in images. However, most current work either focuses on attention-based fusion or pre-training on large-scale image-text pairs, ignoring the critical role of explicit vision-language alignment in visual dialog. To remedy this defect, we propose a novel unsupervised and pseudo-supervised vision-language alignment approach for visual dialog (AlignVD). Firstly, AlginVD utilizes the visual and dialog encoder to represent images and dialogs. Then, it explicitly aligns visual concepts with textual semantics via unsupervised and pseudo-supervised vision-language alignment (UVLA and PVLA). Specifically, UVLA utilizes a graph autoencoder, while PVLA uses dialog-guided visual grounding to conduct alignment. Finally, based on the aligned visual and textual representations, AlignVD gives a reasonable answer to the question via the cross-modal decoder. Extensive experiments on two large-scale visual dialog datasets have demonstrated the effectiveness of vision-language alignment, and our proposed AlignVD achieves new state-of-the-art results. In addition, our single model has won first place on the visual dialog challenge leaderboard with a NDCG metric of 78.70, surpassing the previous best ensemble model by about 1 point. Duzhen Zhang, Xiuyi Chen, Jing Shi 0003, Bo Xu 0002 |
ACM Multimedia | 3 |
| 2020 | Knowledge Aware Emotion Recognition in Textual Conversations via Multi-Task Incremental TransformerabstractEmotion recognition in textual conversations (ERTC) plays an important role in a wide range of applications, such as opinion mining, recommender systems, and so on.ERTC, however, is a challenging task.For one thing, speakers often rely on the context and commonsense knowledge to express emotions; for another, most utterances contain neutral emotion in conversations, as a result, the confusion between a few non-neutral utterances and much more neutral ones restrains the emotion recognition performance.In this paper, we propose a novel Knowledge Aware Incremental Transformer with Multi-task Learning (KAITML) to address these challenges.Firstly, we devise a dual-level graph attention mechanism to leverage commonsense knowledge, which augments the semantic information of the utterance.Then we apply the Incremental Transformer to encode multi-turn contextual utterances.Moreover, we are the first to introduce multi-task learning to alleviate the aforementioned confusion and thus further improve the emotion recognition performance.Extensive experimental results show that our KAITML model outperforms the state-of-the-art models across five benchmark datasets. Duzhen Zhang, Xiuyi Chen, Bo Xu 0002 |
COLING | 2 |
| 2020 | Bridging the Gap between Prior and Posterior Knowledge Selection for Knowledge-Grounded Dialogue GenerationabstractKnowledge selection plays an important role in knowledge-grounded dialogue, which is a challenging task to generate more informative responses by leveraging external knowledge.Recently, latent variable models have been proposed to deal with the diversity of knowledge selection by using both prior and posterior distributions over knowledge and achieve promising performance.However, these models suffer from a huge gap between prior and posterior knowledge selection.Firstly, the prior selection module may not learn to select knowledge properly because of lacking the necessary posterior information.Secondly, latent variable models suffer from the exposure bias that dialogue generation is based on the knowledge selected from the posterior distribution at training but from the prior distribution at inference.Here, we deal with these issues on two aspects: (1) We enhance the prior selection module with the necessary posterior information obtained from the specially designed Posterior Information Prediction Module (PIPM); (2) We propose a Knowledge Distillation Based Training Strategy (KDBTS) to train the decoder with the knowledge selected from the prior distribution, removing the exposure bias of knowledge selection.Experimental results on two knowledge-grounded dialogue datasets show that both PIPM and KDBTS achieve performance improvement over the state-of-theart latent variable model and their combination shows further improvement. Xiuyi Chen, Fandong Meng, Peng Li 0030, Bo Xu 0002, Jie Zhou 0016 |
EMNLP (1) | 1 |
| 2019 | A Working Memory Model for Task-oriented Dialog Response GenerationabstractRecently, to incorporate external Knowledge Base (KB) information, one form of world knowledge, several end-to-end task-oriented dialog systems have been proposed.These models, however, tend to confound the dialog history with KB tuples and simply store them into one memory.Inspired by the psychological studies on working memory, we propose a working memory model ( WMM2Seq) for dialog response generation.Our WMM2Seq adopts a working memory to interact with two separated long-term memories, which are the episodic memory for memorizing dialog history and the semantic memory for storing KB tuples.The working memory consists of a central executive to attend to the aforementioned memories, and a short-term storage system to store the "activated" contents from the longterm memories.Furthermore, we introduce a context-sensitive perceptual process for the token representations of the dialog history, and then feed them into the episodic memory.Extensive experiments on two task-oriented dialog datasets demonstrate that our WMM2Seq significantly outperforms the state-of-the-art results in several evaluation metrics. Xiuyi Chen, Jiaming Xu 0001, Bo Xu 0002 |
ACL (1) | 1 |
| 2018 | Modeling Attention and Memory for Auditory Selection in a Cocktail Party EnvironmentabstractDeveloping a computational auditory model to solve the cocktail party problem has long bedeviled scientists, especially for a single microphone recording. Although recent deep learning based frameworks have made significant progress in multi-talker mixed speech separation, most existing deep learning based methods, focusing on separating all the speech channels rather than selectively attending the target speech and ignoring other sounds, may fail to offer a satisfactory solution in a complex auditory scene where the number of input sounds is usually uncertain and even dynamic. In this work, we employ ideas from auditory selective attention of behavioral and cognitive neurosciences and from recent advances of memory-augmented neural networks. Specifically, a unified Auditory Selection framework with Attention and Memory (dubbed ASAM) is proposed. Our ASAM first accumulates the prior knowledge (that is the acoustic feature to one specific speaker) into a life-long memory during the training phase, meanwhile a speech perceptor is trained to extract the temporal acoustic feature and update the memory online when a salient speech is given. Then, the learned memory is utilized to interact with the mixture input to attend and filter the target frequency out from the mixture stream. Finally, the network is trained to minimize the reconstruction error of the attended speech. We evaluate the proposed approach on WSJ0 and THCHS-30 datasets and the experimental results demonstrate that our approach successfully conducts two auditory selection tasks: the top-down task-specific attention (e.g. to follow a conversation with friend) and the bottom-up stimulus-driven attention (e.g. be attracted by a salient speech). Compared with deep clustering based methods, our method conducts competitive advantages especially in a real noise environment (e.g. street junction). Our code is available at https://github.com/jacoxu/ASAM. Jiaming Xu 0001, Jing Shi 0003, Guangcan Liu, Xiuyi Chen, Bo Xu 0002 |
AAAI | 4 |
| 2018 | Distilled Binary Neural Network for Monaural Speech SeparationabstractMonaural speech separation, aiming at solving the cocktail party problem, has many important application scenarios, most of which ask for the real-time response, high energy efficiency and efficient storage. However, the state-of-the-art Deep Neural Network based separation models usually require huge memory and computation for the 32-bit floating point multiply accumulations, hence most of them cannot meet those requirements. Recently, there are many methods proposed to solve the problem, and binary neural networks have drawn many attentions for they compress and speed up its counterparts at the cost of some performance. Hence, in this paper, we binarize Deep Neural Network based separation models, aiming to deploy them on embedded devices for real-time applications. Furthermore, we improve the separation performance by integrating knowledge distillation into the training phase of binary neural network based models, which is referred as Distilled Binary Neural Network (DBNN). To the best of our knowledge, DBNN is the first attempt to integrate two types of model compression. In the experiments, we demonstrate the effectiveness of our proposed method, which successfully binarizes the Deep Neural Network based separation models with a comparable performance. Xiuyi Chen, Guangcan Liu, Jing Shi 0003, Jiaming Xu 0001, Bo Xu 0002 |
IJCNN | 1 |
| 2018 | Improving Speech Separation with Adversarial Network and Reinforcement LearningabstractIn contrast to the conventional deep neural network for single-channel speech separation, we propose a separation framework based on adversarial network and reinforcement learning. The purpose of the adversarial network inspired by the generative adversarial network is to make the separated result and ground-truth with the same data distribution by evaluating the discrepancy between them. Meanwhile, in order to enable the model to bias the generation towards desirable metrics and reduce the discrepancy between training loss (such as mean squared error) and testing metric (such as SDR), we present the future success based on reinforcement learning. We directly optimize the performance metric to accomplish exactly that. With the combination of adversarial network and reinforcement learning, our model is able to improve the performance of single-channel speech separation. Guangcan Liu, Jing Shi 0003, Xiuyi Chen, Jiaming Xu 0001, Bo Xu 0002 |
IJCNN | 3 |