VLDB 2026 Research / reviewers in the wild / expert
Bowen Zhou 0001
dblp:61/5024-1
· DBLP profile ↗
38ranked-venue papers
0as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The 1st SpeechWellness Challenge: Detecting Suicide Risk Among AdolescentsabstractThe 1st SpeechWellness Challenge (SW1) aims to advance methods for detecting current suicide risk in adolescents using speech analysis techniques. Suicide among adolescents is a critical public health issue globally. Early detection of suicidal tendencies can lead to timely intervention and potentially save lives. Traditional methods of assessment often rely on self-reporting or clinical interviews, which may not always be accessible. The SW1 challenge addresses this gap by exploring speech as a non-invasive and readily available indicator of mental health. We release the SW1 dataset which contains speech recordings from 600 adolescents aged 10-18 years. By focusing on speech generated from natural tasks, the challenge seeks to uncover patterns and markers that correlate with current suicide risk. Wen Wu 0007, Ziyun Cui, Chang Lei, Yinan Duan, Diyang Qu, Ji Wu 0002, Bowen Zhou 0001, Runsen Chen, Chao Zhang 0031 |
INTERSPEECH | 7 |
| 2025 | BrainOmni: A Brain Foundation Model for Unified EEG and MEG SignalsabstractElectroencephalography (EEG) and magnetoencephalography (MEG) measure neural activity non-invasively by capturing electromagnetic fields generated by dendritic currents.
Although rooted in the same biophysics, EEG and MEG exhibit distinct signal patterns, further complicated by variations in sensor configurations across modalities and recording devices.
Existing approaches typically rely on separate, modality- and dataset-specific models, which limits the performance and cross-domain scalability.
This paper proposes BrainOmni, the first brain foundation model that generalises across heterogeneous EEG and MEG recordings.
To unify diverse data sources, we introduce BrainTokenizer, the first tokeniser that quantises spatiotemporal brain activity into discrete representations.
Central to BrainTokenizer is a novel Sensor Encoder that encodes sensor properties such as spatial layout, orientation, and type, enabling compatibility across devices and modalities.
Building upon the discrete representations, BrainOmni learns unified semantic embeddings of brain signals by self-supervised pretraining. To the best of our knowledge, it is the first foundation model to support both EEG and MEG signals, as well as the first to incorporate large-scale MEG pretraining.
A total of 1,997 hours of EEG and 656 hours of MEG data are curated and standardised from publicly available sources for pretraining.
Experiments show that BrainOmni outperforms both existing foundation models and state-of-the-art task-specific models on a range of downstream tasks. It also demonstrates strong generalisation to unseen EEG and MEG devices. Further analysis reveals that joint EEG-MEG (EMEG) training yields consistent improvements across both modalities. Code and checkpoints are publicly available at https://github.com/OpenTSLab/BrainOmni Qinfan Xiao, Ziyun Cui, Wen Wu 0007, Andrew Thwaites, Alexandra Woolgar, Bowen Zhou 0001, Chao Zhang 0031 |
NeurIPS | 8 |
| 2023 | Federated User Modeling from Hierarchical InformationabstractThe generation of large amounts of personal data provides data centers with sufficient resources to mine idiosyncrasy from private records. User modeling has long been a fundamental task with the goal of capturing the latent characteristics of users from their behaviors. However, centralized user modeling on collected data has raised concerns about the risk of data misuse and privacy leakage. As a result, federated user modeling has come into favor, since it expects to provide secure multi-client collaboration for user modeling through federated learning. Unfortunately, to the best of our knowledge, existing federated learning methods that ignore the inconsistency among clients cannot be applied directly to practical user modeling scenarios, and moreover, they meet the following critical challenges: 1) Statistical heterogeneity . The distributions of user data in different clients are not always independently identically distributed (IID), which leads to unique clients with needful personalized information; 2) Privacy heterogeneity . User data contains both public and private information, which have different levels of privacy, indicating that we should balance different information shared and protected; 3) Model heterogeneity . The local user models trained with client records are heterogeneous, and thus require a flexible aggregation in the server; 4) Quality heterogeneity . Low-quality information from inconsistent clients poisons the reliability of user models and offsets the benefit from high-quality ones, meaning that we should augment the high-quality information during the process. To address the challenges, in this paper, we first propose a novel client-server architecture framework, namely Hierarchical Personalized Federated Learning (HPFL), with a primary goal of serving federated learning for user modeling in inconsistent clients. More specifically, the client train and deliver the local user model via the hierarchical components containing hierarchical information from privacy heterogeneity to join collaboration in federated learning. Moreover, the client updates the personalized user model with a fine-grained personalized update strategy for statistical heterogeneity. Correspondingly, the server flexibly aggregates hierarchical components from heterogeneous user models in the case of privacy and model heterogeneity with a differentiated component aggregation strategy. In order to augment high-quality information and generate high-quality user models, we expand HPFL to the Augmented-HPFL (AHPFL) framework by incorporating the augmented mechanisms, which filters out low-quality information such as noise, sparse information and redundant information. Specially, we construct two implementations of AHPFL, i.e., AHPFL-SVD and AHPFL-AE, where the augmented mechanisms follow SVD (singular value decomposition) and AE (autoencoder), respectively. Finally, we conduct extensive experiments on real-world datasets, which demonstrate the effectiveness of both HPFL and AHPFL frameworks. Qi Liu 0003, Zhenya Huang, Hao Wang 0076, Yuting Ning, Enhong Chen, Jinfeng Yi, Bowen Zhou 0001 |
ACM Trans. Inf. Syst. | 9 |
| 2021 | Incremental Learning for End-to-End Automatic Speech RecognitionabstractIn this paper, we propose an incremental learning method for end-to-end Automatic Speech Recognition (ASR) which enables an ASR system to perform well on new tasks while maintaining the performance on its originally learned ones. To mitigate catastrophic forgetting during incremental learning, we design a novel explainability-based knowledge distillation for ASR models, which is combined with a response-based knowledge distillation to maintain the original model's predictions and the “reason” for the predictions. Our method works without access to the training data of original tasks, which addresses the cases where the previous data is no longer available or joint training is costly. Results on a multi-stage sequential training task show that our method outperforms existing ones in mitigating forgetting. Furthermore, in two practical scenarios, compared to the target-reference joint training method, the performance drop of our method is 0.02% Character Error Rate (CER), which is 97% smaller than the drops of the baseline methods. Libo Zi, Zhengchen Zhang, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
ASRU | 7 |
| 2021 | Learn to Copy from the Copying History: Correlational Copy Network for Abstractive SummarizationabstractThe copying mechanism has had considerable success in abstractive summarization, facilitating models to directly copy words from the input text to the output summary.Existing works mostly employ encoder-decoder attention, which applies copying at each time step independently of the former ones.However, this may sometimes lead to incomplete copying.In this paper, we propose a novel copying scheme named Correlational Copying Network (CoCoNet) that enhances the standard copying mechanism by keeping track of the copying history.It thereby takes advantage of prior copying distributions and, at each time step, explicitly encourages the model to copy the input word that is relevant to the previously copied one.In addition, we strengthen CoCoNet through pretraining with suitable corpora that simulate the copying behaviors.Experimental results show that CoCoNet can copy more accurately and achieves new state-of-the-art performances on summarization benchmarks, including CNN/DailyMail for news summarization and SAMSum for dialogue summarization.Our code is available at https:// github.com/hrlinlp/coconet. Haoran Li 0001, Song Xu 0002, Peng Yuan 0002, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
EMNLP (1) | 7 |
| 2021 | Conversational Query Rewriting with Self-Supervised LearningabstractContext modeling plays a critical role in building multi-turn dialogue systems. Conversational Query Rewriting (CQR) aims to simplify the multi-turn dialogue modeling into a single-turn problem by explicitly rewriting the conversational query into a self-contained utterance. However, existing approaches rely on massive supervised training data, which is labor-intensive to annotate. And the detection of the omitted important information from context can be further improved. Besides, intent consistency constraint between contextual query and rewritten query is also ignored. To tackle these issues, we first propose to construct a large-scale CQR dataset automatically via self-supervised learning, which does not need human annotation. Then we introduce a novel CQR model Teresa based on Transformer, which is enhanced by self-attentive keywords detection and intent consistency constraint. Finally, we conduct extensive experiments on two public datasets. Experimental results demonstrate that our proposed model outperforms existing CQR baselines significantly, and also prove the effectiveness of self-supervised learning on improving the CQR performance. Hang Liu 0005, Meng Chen 0006, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
ICASSP | 5 |
| 2021 | Dian: Duration Informed Auto-Regressive Network for Voice CloningabstractIn this paper, we propose a novel end-to-end speech synthesis approach, Duration Informed Auto-regressive Network (DIAN), which consists of an acoustic model and a separate duration model. Un-like other auto-regressive TTS methods, the duration information of phonemes is provided as part of the input to the acoustic model, which enables the removal of the attention mechanism between its encoder and decoder parts. This eliminates the common seen skipping and repeating issues and improves speech intelligibility while ensuring high speech quality. A Transformer-based duration model is used to predict the duration of each phoneme for the attention-free acoustic model. We developed our TTS systems for the multi-speaker multi-style voice cloning challenge (M2VoC) using the proposed DIAN approach. In our procedure, a multi-speaker attention-free acoustic model and its Transformer-based duration model are first separately trained based on the training data released by M2VoC. Next, the multi-speaker models are adapted to form the speaker-specific models with the speaker-dependent data and transfer learning. At last, a speaker-specific LPCNet is estimated and used to synthesize the speech of the corresponding speaker. The M2VoC results showed that our proposed approach achieved the 3rd-place in the speech quality ranking and the 4th-place in the speaker similarity and style similarity ranking in the Track1-a task. Zhengchen Zhang, Chao Zhang 0031, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
ICASSP | 7 |
| 2021 | Improving Prosody Modelling with Cross-Utterance Bert Embeddings for End-to-End Speech SynthesisabstractAlthough speech prosody is related to the linguistic information up to the discourse structure, most text-to-speech (TTS) systems only take into account the information within each sentence. This makes it challenging when converting a paragraph of text into natural and expressive speech. In this paper, we propose to use the text embeddings of the neighboring sentences to improve the prosody generation for each utterance of a paragraph in an end-to-end fashion without using any explicit prosody features. More specifically, cross-utterance (CU) context vectors, which are produced by an additional CU encoder based on the sentence embeddings extracted by a pretrained BERT model, are used to augment the input of the Tacotron2 decoder. Two types of BERT embeddings are investigated, which leads to the use of different CU encoder structures. Experimental results on a Mandarin audiobook dataset and the LJ-Speech English audiobook dataset demonstrate the use of CU information can improve the naturalness and expressiveness of the synthesized speech. Subjective listening testing shows most of the participants prefer the voice generated using the CU encoder over that generated using standard Tacotron2. It is also found that the prosody can be controlled indirectly by changing the neighbouring sentences. Zhengchen Zhang, Chao Zhang 0031, Xiaodong He 0001, Bowen Zhou 0001 |
ICASSP | 6 |
| 2021 | Neural Kalman Filtering for Speech EnhancementabstractConventional learning-based speech enhancement methods usually utilize existing building blocks to design the deep neural networks (DNNs), while how to effectively integrate the statistical signal processing based schemes, which are expert-knowledge driven and could ameliorate the over-fitting problem, into the network design remains an open issue. In this paper, we extend the conventional Kalman filtering (KF) and propose a supervised-learning based neural Kalman filter (NKF) for speech enhancement. Similar to KF, the proposed method first obtains a prediction from the speech evolution model and then integrates the short-term instantaneous observation by linear weighting, and the weights are calculated by comparing between the speech prediction residual error and the environmental noise level. An end-to-end network is designed to convert the speech linear prediction model in KF to non-linear, and to compact all other conventional linear filtering operations. Different with other DNN based methods, the proposed method provides a specialized network design inspired from the conventional signal processing, the backpropagation can be directly applied on the linear filtering operations integrated from KF. We conduct experiments in different noisy conditions, and the results demonstrate that the proposed method outperforms the baseline methods which are based on either signal processing or DNNs. Wei Xue 0002, Gang Quan, Chao Zhang 0031, Guo-Hong Ding, Xiaodong He 0001, Bowen Zhou 0001 |
ICASSP | 6 |
| 2021 | Graph Ensemble Learning over Multiple Dependency Trees for Aspect-level Sentiment ClassificationabstractXiaochen Hou, Peng Qi, Guangtao Wang, Rex Ying, Jing Huang, Xiaodong He, Bowen Zhou. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Xiaochen Hou, Peng Qi 0003, Guangtao Wang, Rex Ying, Jing Huang 0019, Xiaodong He 0001, Bowen Zhou 0001 |
NAACL-HLT | 7 |
| 2021 | SGG: Learning to Select, Guide, and Generate for Keyphrase GenerationabstractJing Zhao, Junwei Bao, Yifan Wang, Youzheng Wu, Xiaodong He, Bowen Zhou. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Junwei Bao 0001, Yifan Wang 0016, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
NAACL-HLT | 6 |
| 2021 | CUSTOM: Aspect-Oriented Product Summarization for E-Commerce
Jiahui Liang, Junwei Bao 0001, Yifan Wang 0016, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
NLPCC (2) | 6 |
| 2021 | EviDR: Evidence-Emphasized Discrete Reasoning for Reasoning Machine Reading Comprehension
Yongwei Zhou, Junwei Bao 0001, Haipeng Sun, Jiahui Liang, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001, Tiejun Zhao |
NLPCC (1) | 7 |
| 2021 | Hierarchical Personalized Federated Learning for User ModelingabstractUser modeling aims to capture the latent characteristics of users from their behaviors, and is widely applied in numerous applications. Usually, centralized user modeling suffers from the risk of privacy leakage. Instead, federated user modeling expects to provide a secure multi-client collaboration for user modeling through federated learning. Existing federated learning methods are mainly designed for consistent clients, which cannot be directly applied to practical scenarios, where different clients usually store inconsistent user data. Therefore, it is a crucial demand to design an appropriate federated solution that can better adapt to user modeling tasks, and however, meets following critical challenges: 1) Statistical heterogeneity. The distributions of user data in different clients are not always independently identically distributed which leads to personalized clients; 2) Privacy heterogeneity. User data contains both public and private information, which have different levels of privacy. It means we should balance different information to be shared and protected; 3) Model heterogeneity. The local user models trained with client records are heterogeneous which need flexible aggregation in the server. In this paper, we propose a novel client-server architecture framework, namely Hierarchical Personalized Federated Learning (HPFL) to serve federated learning in user modeling with inconsistent clients. In the framework, we first define hierarchical information to finely partition the data with privacy heterogeneity. On this basis, the client trains a user model which contains different components designed for hierarchical information. Moreover, client processes a fine-grained personalized update strategy to update personalized user model for statistical heterogeneity. Correspondingly, the server completes a differentiated component aggregation strategy to flexibly aggregate heterogeneous user models in the case of privacy and model heterogeneity. Finally, we conduct extensive experiments on real-world datasets, which demonstrate the effectiveness of the HPFL framework. Qi Liu 0003, Zhenya Huang, Yuting Ning, Hao Wang 0076, Enhong Chen, Jinfeng Yi, Bowen Zhou 0001 |
WWW | 8 |
| 2020 | Zero-Shot Text-to-SQL Learning with Auxiliary TaskabstractRecent years have seen great success in the use of neural seq2seq models on the text-to-SQL task. However, little work has paid attention to how these models generalize to realistic unseen data, which naturally raises a question: does this impressive performance signify a perfect generalization model, or are there still some limitations?In this paper, we first diagnose the bottleneck of the text-to-SQL task by providing a new testbed, in which we observe that existing models present poor generalization ability on rarely-seen data. The above analysis encourages us to design a simple but effective auxiliary task, which serves as a supportive model as well as a regularization term to the generation task to increase the models' generalization. Experimentally, We evaluate our models on a large text-to-SQL dataset WikiSQL. Compared to a strong baseline coarse-to-fine model, our models improve over the baseline by more than 3% absolute in accuracy on the whole dataset. More interestingly, on a zero-shot subset test of WikiSQL, our models achieve 5% absolute accuracy gain over the baseline, clearly demonstrating its superior generalizability. Shuaichen Chang, Yun Tang 0002, Jing Huang 0019, Xiaodong He 0001, Bowen Zhou 0001 |
AAAI | 6 |
| 2020 | Aspect-Aware Multimodal Summarization for Chinese E-Commerce ProductsabstractWe present an abstractive summarization system that produces summary for Chinese e-commerce products. This task is more challenging than general text summarization. First, the appearance of a product typically plays a significant role in customers' decisions to buy the product or not, which requires that the summarization model effectively use the visual information of the product. Furthermore, different products have remarkable features in various aspects, such as “energy efficiency” and “large capacity” for refrigerators. Meanwhile, different customers may care about different aspects. Thus, the summarizer needs to capture the most attractive aspects of a product that resonate with potential purchasers. We propose an aspect-aware multimodal summarization model that can effectively incorporate the visual information and also determine the most salient aspects of a product. We construct a large-scale Chinese e-commerce product summarization dataset that contains approximately 1.4 million manually created product summaries that are paired with detailed product information, including an image, a title, and other textual descriptions for each product. The experimental results on this dataset demonstrate that our models significantly outperform the comparative methods in terms of both the ROUGE score and manual evaluations. Haoran Li 0001, Peng Yuan 0002, Song Xu 0002, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
AAAI | 6 |
| 2020 | Select, Answer and Explain: Interpretable Multi-Hop Reading Comprehension over Multiple DocumentsabstractInterpretable multi-hop reading comprehension (RC) over multiple documents is a challenging problem because it demands reasoning over multiple information sources and explaining the answer prediction by providing supporting evidences. In this paper, we propose an effective and interpretable Select, Answer and Explain (SAE) system to solve the multi-document RC problem. Our system first filters out answer-unrelated documents and thus reduce the amount of distraction information. This is achieved by a document classifier trained with a novel pairwise learning-to-rank loss. The selected answer-related documents are then input to a model to jointly predict the answer and supporting sentences. The model is optimized with a multi-task learning objective on both token level for answer prediction and sentence level for supporting sentences prediction, together with an attention-based interaction between these two tasks. Evaluated on HotpotQA, a challenging multi-hop RC data set, the proposed SAE system achieves top competitive performance in distractor setting compared to other existing systems on the leaderboard. Kevin Huang 0002, Guangtao Wang, Jing Huang 0019, Xiaodong He 0001, Bowen Zhou 0001 |
AAAI | 6 |
| 2020 | Orthogonal Relation Transforms with Graph Context Modeling for Knowledge Graph EmbeddingabstractDistance-based knowledge graph embeddings have shown substantial improvement on the knowledge graph link prediction task, from TransE to the latest state-of-the-art RotatE.However, complex relations such as N-to-1, 1-to-N and N-to-N still remain challenging to predict.In this work, we propose a novel distance-based approach for knowledge graph link prediction.First we extend the RotatE from 2D complex domain to high dimensional space with orthogonal transforms to model relations.The orthogonal transform embedding for relations keeps the capability for modeling symmetric/anti-symmetric, inverse and compositional relations while achieves better modeling capacity.Second, the graph context is integrated into distance scoring functions directly.Specifically, graph context is explicitly modeled via two directed context representations.Each node embedding in knowledge graph is augmented with two context representations, which are computed from the neighboring outgoing and incoming nodes/edges respectively.The proposed approach improves prediction accuracy on the difficult N-to-1, 1-to-N and N-to-N cases.Our experimental results show that it achieves state-of-the-art results on two common benchmarks FB15k-237 and WNRR-18, especially on FB15k-237 which has many high in-degree nodes.Code available at https://github. com/JD-AI-Research-Silicon-Valley/ KGEmbedding-OTE. Yun Tang 0002, Jing Huang 0019, Guangtao Wang, Xiaodong He 0001, Bowen Zhou 0001 |
ACL | 5 |
| 2020 | Self-Attention Guided Copy Mechanism for Abstractive SummarizationabstractCopy module has been widely equipped in the recent abstractive summarization models, which facilitates the decoder to extract words from the source into the summary.Generally, the encoder-decoder attention is served as the copy distribution, while how to guarantee that important words in the source are copied remains a challenge.In this work, we propose a Transformer-based model to enhance the copy mechanism.Specifically, we identify the importance of each source word based on the degree centrality with a directed graph built by the self-attention layer in the Transformer.We use the centrality of each source word to guide the copy process explicitly.Experimental results show that the self-attention graph provides useful guidance for the copy distribution.Our proposed models significantly outperform the baseline methods on the CNN/Daily Mail dataset and the Gigaword dataset. Song Xu 0002, Haoran Li 0001, Peng Yuan 0002, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
ACL | 6 |
| 2020 | Learning to Decouple Relations: Few-Shot Relation Classification with Entity-Guided Attention and Confusion-Aware TrainingabstractThis paper aims to enhance the few-shot relation classification especially for sentences that jointly describe multiple relations.Due to the fact that some relations usually keep high cooccurrence in the same context, previous few-shot relation classifiers struggle to distinguish them with few annotated instances.To alleviate the above relation confusion problem, we propose CTEG, a model equipped with two mechanisms to learn to decouple these easily-confused relations.On the one hand, an Entity-Guided Attention (EGA) mechanism, which leverages the syntactic relations and relative positions between each word and the specified entity pair, is introduced to guide the attention to filter out information causing confusion.On the other hand, a Confusion-Aware Training (CAT) method is proposed to explicitly learn to distinguish relations by playing a pushing-away game between classifying a sentence into a true relation and its confusing relation.Extensive experiments are conducted on the FewRel dataset, and the results show that our proposed model achieves comparable and even much better results to strong baselines in terms of accuracy.Furthermore, the ablation test and case study verify the effectiveness of our proposed EGA and CAT, especially in addressing the relation confusion problem. Yingyao Wang, Junwei Bao 0001, Guangyi Liu 0005, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001, Tiejun Zhao |
COLING | 6 |
| 2020 | On the Faithfulness for E-commerce Product SummarizationabstractIn this work, we present a model to generate e-commerce product summaries.The consistency between the generated summary and the product attributes is an essential criterion for the ecommerce product summarization task.To enhance the consistency, first, we encode the product attribute table to guide the process of summary generation.Second, we identify the attribute words from the vocabulary, and we constrain these attribute words can be presented in the summaries only through copying from the source, i.e., the attribute words not in the source cannot be generated.We construct a Chinese e-commerce product summarization dataset, and the experimental results on this dataset demonstrate that our models significantly improve the faithfulness. Peng Yuan 0002, Haoran Li 0001, Song Xu 0002, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
COLING | 6 |
| 2020 | Learning to Predict Charges for Legal Judgment via Self-Attentive Capsule NetworkabstractWith the rapid development of deep learning technology, more and more traditional industries are changed by Artificial Intelligence. The legal industry is such a popular scenario which attracts lots of researchers' interests. In this work, we focus on automatic charge prediction, which predicts the final charges according to the given fact descriptions in criminal cases. It is crucial for legal assistant systems and can help the judges improve work efficiency greatly. However, extremely imbalanced data distribution and lengthy fact descriptions make this task especially challenging. To tackle these two issues, we propose a novel model, namely Self-Attentive Capsule Network (dubbed as SAttCaps). In particular, we devise a self-attentive dynamic routing, which can not only capture long-range dependency more directly than vanilla dynamic routing, but also learn the high-level generalized features better. The experimental results on three real-world datasets demonstrate that our model significantly outperforms the baselines and creates new state-of-the-art performance. Moreover, our model performs much better than the baselines especially in the low-frequency charges and can bring 5.7% absolute improvement under F1 score. Yuquan Le, Congqing He, Meng Chen 0006, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
ECAI | 6 |
| 2020 | Multimodal Joint Attribute Prediction and Value Extraction for E-commerce ProductabstractProduct attribute values are essential in many e-commerce scenarios, such as customer service robots, product recommendations, and product retrieval.While in the real world, the attribute values of a product are usually incomplete and vary over time, which greatly hinders the practical applications.In this paper, we propose a multimodal method to jointly predict product attributes and extract values from textual product descriptions with the help of the product images.We argue that product attributes and values are highly correlated, e.g., it will be easier to extract the values on condition that the product attributes are given.Thus, we jointly model the attribute prediction and value extraction tasks from multiple aspects towards the interactions between attributes and values.Moreover, product images have distinct effects on our tasks for different product attributes and values.Thus, we selectively draw useful visual information from product images to enhance our model.We annotate a multimodal product attribute value dataset that contains 87,194 instances, and the experimental results on this dataset demonstrate that explicitly modeling the relationship between attributes and values facilitates our method to establish the correspondence between them, and selectively utilizing visual product information is necessary for the task.Our code and dataset are available 1 . Tiangang Zhu, Haoran Li 0001, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
EMNLP (1) | 6 |
| 2020 | Efficient WaveGlow: An Improved WaveGlow Vocoder with Enhanced Speed
Zhengchen Zhang, Chao Zhang 0031, Xiaodong He 0001, Bowen Zhou 0001 |
INTERSPEECH | 6 |
| 2020 | Sound Event Localization and Detection Based on Multiple DOA Beamforming and Multi-Task LearningabstractThe performance of sound event localization and detection (SELD) degrades in source-overlapping cases since features of different sources collapse with each other, and the network tends to fail to learn to separate these features effectively. In this paper, by leveraging the conventional microphone array signal processing to generate comprehensive representations for SELD, we propose a new SELD method based on multiple direction of arrival (DOA) beamforming and multi-task learning. By using multiple beamformers to extract the signals from different DOAs, the sound field is more diversely described, and specialised representations of target source and noises can be obtained. With labelled training data, the steering vector is estimated based on the cross-power spectra (CPS) and the signal presence probability (SPP), which eliminates the need of knowing the array geometry. We design two networks for sound event localization (SED) and sound source localization (SSL) and use a multi-task learning scheme for SED, in which the SSL-related task act as a regularization. Experimental results using the database of DCASE2019 SELD task show that the proposed method achieves the state-of-art performance. Wei Xue 0002, Ying Tong, Chao Zhang 0031, Guo-Hong Ding, Xiaodong He 0001, Bowen Zhou 0001 |
INTERSPEECH | 6 |
| 2020 | The JDDC Corpus: A Large-Scale Multi-Turn Chinese Dialogue Dataset for E-commerce Customer ServiceabstractHuman conversations are complicated and building a human-like dialogue agent is an extremely challenging task. With the rapid development of deep learning techniques, data-driven models become more and more prevalent which need a huge amount of real conversation data. In this paper, we construct a large-scale real scenario Chinese E-commerce conversation corpus, JDDC, with more than 1 million multi-turn dialogues, 20 million utterances, and 150 million words. The dataset reflects several characteristics of human-human conversations, e.g., goal-driven, and long-term dependency among the context. It also covers various dialogue types including task-oriented, chitchat and question-answering. Extra intent information and three well-annotated challenge sets are also provided. Then, we evaluate several retrieval-based and generative models to provide basic benchmark performance on the JDDC corpus. And we hope JDDC can serve as an effective testbed and benefit the development of fundamental research in dialogue task. Meng Chen 0006, Ruixue Liu, Lei Shen 0001, Shaozu Yuan, Jingyan Zhou, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
LREC | 8 |
| 2019 | End-to-End Structure-Aware Convolutional Networks for Knowledge Base CompletionabstractKnowledge graph embedding has been an active research topic for knowledge base completion, with progressive improvement from the initial TransE, TransH, DistMult et al to the current state-of-the-art ConvE. ConvE uses 2D convolution over embeddings and multiple layers of nonlinear features to model knowledge graphs. The model can be efficiently trained and scalable to large knowledge graphs. However, there is no structure enforcement in the embedding space of ConvE. The recent graph convolutional network (GCN) provides another way of learning graph node embedding by successfully utilizing graph connectivity structure. In this work, we propose a novel end-to-end StructureAware Convolutional Network (SACN) that takes the benefit of GCN and ConvE together. SACN consists of an encoder of a weighted graph convolutional network (WGCN), and a decoder of a convolutional network called Conv-TransE. WGCN utilizes knowledge graph node structure, node attributes and edge relation types. It has learnable weights that adapt the amount of information from neighbors used in local aggregation, leading to more accurate embeddings of graph nodes. Node attributes in the graph are represented as additional nodes in the WGCN. The decoder Conv-TransE enables the state-of-the-art ConvE to be translational between entities and relations while keeps the same link prediction performance as ConvE. We demonstrate the effectiveness of the proposed SACN on standard FB15k-237 and WN18RR datasets, and it gives about 10% relative improvement over the state-of-theart ConvE in terms of HITS@1, HITS@3 and HITS@10. Yun Tang 0002, Jing Huang 0019, Jinbo Bi, Xiaodong He 0001, Bowen Zhou 0001 |
AAAI | 6 |
| 2019 | Multi-hop Reading Comprehension across Multiple Documents by Reasoning over Heterogeneous GraphsabstractMulti-hop reading comprehension (RC) across documents poses new challenge over single-document RC because it requires reasoning over multiple documents to reach the final answer. In this paper, we propose a new model to tackle the multi-hop RC problem. We introduce a heterogeneous graph with different types of nodes and edges, which is named as Heterogeneous Document-Entity (HDE) graph. The advantage of HDE graph is that it contains different granularity levels of information including candidates, documents and entities in specific document contexts. Our proposed model can do reasoning over the HDE graph with nodes representation initialized with co-attention and self-attention based context encoders. We employ Graph Neural Networks (GNN) based message passing algorithms to accumulate evidences on the proposed HDE graph. Evaluated on the blind test set of the Qangaroo WikiHop data set, our HDE graph based single model delivers competitive result, and the ensemble model achieves the state-of-the-art performance. Guangtao Wang, Jing Huang 0019, Yun Tang 0002, Xiaodong He 0001, Bowen Zhou 0001 |
ACL (1) | 6 |
| 2019 | Relation Module for Non-Answerable Predictions on Reading ComprehensionabstractMachine reading comprehension (MRC) has attracted significant amounts of research attention recently, due to an increase of challenging reading comprehension datasets.In this paper, we aim to improve a MRC model's ability to determine whether a question has an answer in a given context (e.g. the recently proposed SQuAD 2.0 task).Our solution is a relation module that is adaptable to any MRC model.The relation module consists of both semantic extraction and relational information.We first extract high level semantics as objects from both question and context with multihead self-attentive pooling.These semantic objects are then passed to a relation network, which generates relationship scores for each object pair in a sentence.These scores are used to determine whether a question is nonanswerable.We test the relation module on the SQuAD 2.0 dataset using both the BiDAF and BERT models as baseline readers.We obtain 1.8% gain of F1 accuracy on top of the BiDAF reader, and 1.0% on top of the BERT base model.These results show the effectiveness of our relation module on MRC. Kevin Huang 0002, Yun Tang 0002, Jing Huang 0019, Xiaodong He 0001, Bowen Zhou 0001 |
CoNLL | 5 |
| 2019 | Deep Speaker Embedding Learning with Multi-level Pooling for Text-independent Speaker VerificationabstractThis paper aims to improve the widely used deep speaker embedding x-vector model. We propose the following improvements: (1) a hybrid neural network structure using both time delay neural network (TDNN) and long short-term memory neural networks (LSTM) to generate complementary speaker information at different levels; (2) a multi-level pooling strategy to collect speaker information from both TDNN and LSTM layers; (3) a regularization scheme on the speaker embedding extraction layer to make the extracted embeddings suitable for the following fusion step. The synergy of these improvements are shown on the NIST SRE 2016 eval test (with a 19% EER reduction) and SRE 2018 dev test (with a 9% EER reduction), as well as more than 10% DCF scores reduction on these two test sets over the x-vector baseline. Yun Tang 0002, Guo-Hong Ding, Jing Huang 0019, Xiaodong He 0001, Bowen Zhou 0001 |
ICASSP | 5 |
| 2019 | Universal Stagewise Learning for Non-Convex Problems with Convergence on Averaged Solutions
Zaiyi Chen, Zhuoning Yuan, Jinfeng Yi, Bowen Zhou 0001, Enhong Chen, Tianbao Yang |
ICLR (Poster) | 4 |
| 2019 | On the Convergence and Robustness of Adversarial TrainingabstractImproving the robustness of deep neural networks (DNNs) to adversarial examples is an important yet challenging problem for secure deep learning. Across existing defense techniques, adversarial training with Projected Gradient Decent (PGD) is amongst the most effective. Adversarial training solves a min-max optimization problem, with the inner maximization generating adversarial examples by maximizing the classification loss, and the outer minimization finding model parameters by minimizing the loss on adversarial examples generated from the inner maximization. A criterion that measures how well the inner maximization is solved is therefore crucial for adversarial training. In this paper, we propose such a criterion, namely First-Order Stationary Condition for constrained optimization (FOSC), to quantitatively evaluate the convergence quality of adversarial examples found in the inner maximization. With FOSC, we find that to ensure better robustness, it is essential to use adversarial examples with better convergence quality at the later stages of training. Yet at the early stages, high convergence quality adversarial examples are not necessary and may even lead to poor robustness. Based on these observations, we propose a dynamic training strategy to gradually increase the convergence quality of the generated adversarial examples, which significantly improves the robustness of adversarial training. Our theoretical and empirical results show the effectiveness of the proposed method. Yisen Wang 0001, Xingjun Ma, James Bailey 0001, Jinfeng Yi, Bowen Zhou 0001, Quanquan Gu |
ICML | 5 |
| 2019 | Improving the Robustness of Deep Neural Networks via Adversarial Training with Triplet LossabstractRecent studies have highlighted that deep neural networks (DNNs) are vulnerable to adversarial examples. In this paper, we improve the robustness of DNNs by utilizing techniques of Distance Metric Learning. Specifically, we incorporate Triplet Loss, one of the most popular Distance Metric Learning methods, into the framework of adversarial training. Our proposed algorithm, Adversarial Training with Triplet Loss (AT2L), substitutes the adversarial example against the current model for the anchor of triplet loss to effectively smooth the classification boundary. Furthermore, we propose an ensemble version of AT2L, which aggregates different attack methods and model structures for better defense effects. Our empirical studies verify that the proposed approach can significantly improve the robustness of DNNs without sacrificing accuracy. Finally, we demonstrate that our specially designed triplet loss can also be used as a regularization term to enhance other defense methods. Jinfeng Yi, Bowen Zhou 0001, Lijun Zhang 0005 |
IJCAI | 3 |
| 2019 | Multi-Stride Self-Attention for Speech Recognition
Kyu Jeong Han, Jing Huang 0019, Yun Tang 0002, Xiaodong He 0001, Bowen Zhou 0001 |
INTERSPEECH | 5 |
| 2019 | Speaker Diarization with Lexical InformationabstractThis work presents a novel approach for speaker diarization to leverage lexical information provided by automatic speech recognition. We propose a speaker diarization system that can incorporate word-level speaker turn probabilities with speaker embeddings into a speaker clustering process to improve the overall diarization accuracy. To integrate lexical and acoustic information in a comprehensive way during clustering, we introduce an adjacency matrix integration for spectral clustering. Since words and word boundary information for word-level speaker turn probability estimation are provided by a speech recognition system, our proposed method works without any human intervention for manual transcriptions. We show that the proposed method improves diarization performance on various evaluation datasets compared to the baseline diarization system using acoustic information only in speaker embeddings. Tae Jin Park, Kyu Jeong Han, Jing Huang 0019, Xiaodong He 0001, Bowen Zhou 0001, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 5 |
| 2019 | Direct-Path Signal Cross-Correlation Estimation for Sound Source Localization in ReverberationabstractSound source localization (SSL) is challenging in presence of reverberation since the cross-correlation between the direct-path signals in different microphones, which indicates the spatial information of the sound source, is interfered by the reverberation signal components. A novel algorithm is proposed in this paper to estimate the cross-correlation of the direct-path speech signals, such that the robustness of SSL to reverberation can be improved. The proposed method follows a similar scheme to the multichannel linear prediction (MCLP), which is commonly used for speech dereverberation, while avoids the explicit estimation of the direct-path signal of each channel. This is achieved by revealing the relationship between the direct-path signal cross-correlation (DPCC) and the MCLP coefficient vector, and finally deriving the DPCC by using only the multichannel reverberant signals. It is also shown that the pre-whitening operation, which is widely used for SSL, can be inherently integrated into the estimated DPCC. An adaptive method is further derived to facilitate online frame-level SSL. The proposed method can be easily applied to conventional cross-correlation based SSL methods by using the DPCC rather than the full cross-correlation. Experiments conducted in various reverberant conditions demonstrate the effectiveness of the proposed method. Wei Xue 0002, Ying Tong, Guo-Hong Ding, Chao Zhang 0031, Xiaodong He 0001, Bowen Zhou 0001 |
INTERSPEECH | 7 |
| 2019 | Automated Thematic and Emotional Modern Chinese Poetry Composition
Meng Chen 0006, Yang Song 0008, Xiaodong He 0001, Bowen Zhou 0001 |
NLPCC (1) | 5 |
| 2019 | Fast Unsupervised Location Category Inference from Highly Inaccurate Mobility DataabstractUnderstanding a mobile user's behavior, e.g., to infer if she is exercising in a gym or dining in a restaurant, is the key to a variety of applications. However, in many real-world scenarios, precisely determining user visitation is extremely challenging due to the uncertainty present in mobile location updates, where errors can be hundreds of meters or even more. We consider the location uncertainty circle determined by the reported location coordinates as the center and the associated location error as the radius. Such a location uncertainty circle is likely to cover multiple location categories, especially in densely populated areas. Worse still, in many cases, mobile users are anonymous, and we have no access to their personal information or other labeled data, which compels us to develop an unsupervised learning approach to solve this problem. Using a user-time-location category tensor, we capture the user behavior and propose a novel tensor factorization framework to accurately infer the location categories visited by mobile users. This framework leverages several key observations including the negative-unlabeled nature of the data and the intrinsic correlations between users. Also, the proposed algorithm can predict where users are even in the absence of location information. To efficiently solve the proposed framework, we propose a parameter-free and scalable optimization algorithm by effectively exploring the sparse and low-rank structure of the tensor. Our empirical studies show that the proposed algorithm is both effective and scalable: it can solve problems with millions of users and billions of location updates, and also provide superior prediction accuracies on real-world location update and check-in datasets. Jinfeng Yi, Wesley M. Gifford, Junchi Yan, Bowen Zhou 0001 |
SDM | 6 |