VLDB 2026 Research / reviewers in the wild / expert
Yannan Wang
dblp:153/0750
· DBLP profile ↗
40ranked-venue papers
6as first author
29since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 2 first-author · 21 since 2021Artificial intelligence and machine learning · 17 · 4 first-author · 11 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Maintaining Content Strong Consistency From an Edge Caching Perspective: An Offline Reinforcement Learning ApproachabstractIn mobile edge computing (MEC), edge caching (EC), i.e., deploying data closer to end users at the network edge layer, has become one of the most active research areas for mitigating the high latency caused by massive data in centralized storage systems. However, when caching data that are sensitive to freshness, distributed EC systems often face severe strong content consistency challenges, as such data undergo frequent updates and modifications over long usage cycles. To address the strong consistency issue faced by freshness-sensitive data in resource-constrained MEC scenarios, we innovatively reformulate this challenge as a content replacement problem. The objective is to minimize long-term content transmission costs while maintaining strong content consistency. More specifically, we propose the Content Strong Consistency Decision Transformer (CSC-DT) algorithm to maintain strong consistency. The cache replacement problem is modeled as a Markov Decision Process (MDP), with decisions constructed as trajectories. These trajectories are processed using the Transformer architecture to capture the long-term dependencies between the temporal dynamics of content requests and updates, optimizing cache replacement decisions. Validated with railway engineering data, which undergo frequent updates during the design phase, our method effectively maintains strong content consistency and reduces content transmission costs. Yannan Wang, Zhen Liu 0052, Chong Geng, Feng Liu 0061 |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | Multi-Level Speaker Representation for Target Speaker ExtractionabstractTarget speaker extraction (TSE) relies on a reference cue of the target to extract the target speech from a speech mixture. While a speaker embedding is commonly used as the reference cue, such embedding pre-trained with a large number of speakers may suffer from confusion of speaker identity. In this work, we propose a multi-level speaker representation approach, from raw features to neural embeddings, to serve as the speaker reference cue. We generate a spectral-level representation from the enrollment magnitude spectrogram as a raw, low-level feature, which significantly improves the model’s generalization capability. Additionally, we propose a contextual embedding feature based on cross-attention mechanisms that integrate frame-level embeddings from a pre-trained speaker encoder. By incorporating speaker features across multiple levels, we significantly enhance the performance of the TSE model. Our approach achieves a 2.74 dB improvement and a 4.94% increase in extraction accuracy on Libri2mix test set over the baseline. Shuai Wang 0016, Yangjie Wei, Yannan Wang, Haizhou Li 0001 |
ICASSP | 6 |
| 2025 | Distributed Cloud-Edge Scheduling for Multimedia Data Requests: A MARL ApproachabstractHigh-frequency, large-volume, and latency-sensitive multimedia data requests present challenges to the load capacity of edge caching-based content distribution networks. However, existing research primarily focuses on optimizing caching algorithms or employing centralized scheduling based on global system states, often overlooking content delivery and resource allocation tailored to request characteristics. In addition, centralized scheduling is vulnerable to network congestion and single points of failure, while distributed scheduling introduces high information synchronization overhead and difficulties in node collaboration. To address these issues, we propose a novel distributed cloud-edge request scheduling framework that decomposes the scheduling problem into multiple sub-problems for efficient algorithmic optimization. We propose DEPPO, which is an improved Multi-Agent Reinforcement Learning (MARL) algorithm designed to address the challenges of distributed scheduling while also tackling the issue of dimensional uncertainty in deep reinforcement learning(DRL). Experiments conducted on real-world datasets demonstrate that our approach significantly enhances request success rates and resource utilization. Chong Geng, Yannan Wang |
ICME | 3 |
| 2025 | Accent Normalization Using Self-Supervised Discrete Tokens with Non-Parallel Data
Qibing Bai, Sho Inoue, Zhongjie Jiang, Yannan Wang, Haizhou Li 0001 |
INTERSPEECH | 5 |
| 2025 | SpeechRefiner: Towards Perceptual Quality Refinement for Front-End Algorithms
Zhongjie Jiang, Yannan Wang, Haizhou Li 0001 |
INTERSPEECH | 5 |
| 2025 | Multi-teacher knowledge distillation for debiasing recommendation with uniform data
Zhen Liu 0052, Yafan Yuan, Yannan Wang |
Expert Syst. Appl. | 5 |
| 2025 | Distributed Multi-Agent Reinforcement Learning on a Hierarchical Game Model for Railway Engineering Data Collaborative Edge CachingabstractThe rapid expansion and intelligent development of railway infrastructure are driving significant growth in railway engineering data. For the dispersed users across railway networks’ complex topology, traditional centralized storage systems are insufficient for their low-latency, cost-efficient data retrieval. Existing edge caching solutions based on multi-agent reinforcement learning fail to address the asymmetric relationships among railway nodes, such as data centers, stations, and sections, etc. Besides, the complexity of computing Nash equilibrium points also gets higher as the number of agents (edge caching servers) increases. This study introduces a Hierarchical Game model-based MADRL-driven Collaborative Edge Caching method(HG-MCEC) tailored for railway engineering data. By considering the distribution characteristics and caching strategy games among railway nodes, a hierarchical game model for collaborative edge caching is constructed. This model treats the railway edge caching as a multi-agent system, in which each railway node server is regarded as an agent. HG-MCEC utilizes deep learning to mitigate computational complexity and recognize agents’ asymmetry. Upper-level agents adjust cache replacement strategies according to environmental changes and decisionmaking experience. Lower-level agents, under the guidance of upper-level decisions, optimize collaborative caching strategies toward achieving hierarchical game equilibrium. Using a highspeed railway building information modeling data for validation, the method significantly outperforms existing approaches by enhancing content hit rates and reducing latency at edge caching servers while decreasing system content transmission costs. Yannan Wang, Zhen Liu 0052, Chong Geng, Yidong Li |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | Global Route Planning for Large-Scale Requests on Traffic-Aware Road Network
Yannan Wang |
DASFAA (6) | 2 |
| 2024 | Disentangled causal representation learning for debiasing recommendation with uniform data
Zhen Liu 0052, Yannan Wang, Sibo Lu, Feng Liu 0061 |
Appl. Intell. | 4 |
| 2024 | Real-Time Adaptive Partition and Resource Allocation for Multi-User End-Cloud Inference Collaboration in Mobile EnvironmentabstractThe deployment of Deep Neural Networks (DNNs) requires significant computational and storage resources, which is challenging for resource-constrained end devices. To this end, collaborative deep inference is proposed, in which the DNN is divided into two parts and executed on the end device and cloud respectively. The selection of DNN partition point is the key challenge to realize end-cloud collaborative deep inference, especially in mobile environments with unstable networks. In this paper, we propose a Real-time Adaptive Partition (RAP) framework, in which a fast split point decision algorithm is proposed to realize real-time adaptive DNN model partition in the mobile network. A weighted joint optimization of DNN quantization loss, inference and transmission latency is performed. We further propose a Joint Multi-user Model Partition and Resource Allocation (JM-MPRA) algorithm under RAP framework. JM-MPRA aims to guarantee the optimized latency, accuracy and resource utilization in the multi-user scene. Experimental evaluations have demonstrated the effectiveness of RAP with JM-MPRA in improving the performance of real-time end-cloud collaborative inference in both stable and unstable mobile networks. Compared with the state-of-the-art methods, the proposed approaches can achieve up to 5.06x decrease in inference latency and bring performance improvement of 1.52% in inference accuracy. Zhen Liu 0052, Ze Kou, Yannan Wang, Yidong Li, Yongqi Sun |
IEEE Trans. Mob. Comput. | 4 |
| 2023 | Efficient Routing Algorithm for Large-Scale Query Requests in LEO Satellite NetworksabstractLow Earth Orbit (LEO) satellites have become important means of communication, and more and more transmission data flow query requests may arrive simultaneously due to the increasing number of users. Existing routing algorithms often prioritize individual data flow efficiency, neglecting satellite occupation and link utilization. This exacerbates queuing time and transmission delay. In this paper, Software Defined Network (SDN) is employed to obtain the information of satellite networks and focus on the overall transmission efficiency of a batch of data flows. A Large-Scale Query satellite routing Algorithm (LSQA) is proposed, which estimates the resources required for each data flow first and intends to find an optimal data flow query execution order to reduce satellite congestion. To speed up the estimation, we construct the node labels, so that the shortest path between satellites can be obtained quickly. Furthermore, we propose a strategy based on threshold filtering to obtain the optimal execution order more efficiently by finding out data flows whose execution order does not affect the overall transmission delay. Extensive experiments conducted on satellite network simulation show that LSQA has the superiority in terms of queuing delay and load-balancing compared with counterparts. Jiajia Li 0003, Yannan Wang, Liang Zhao 0004 |
GLOBECOM | 2 |
| 2023 | Inter-Subnet: Speech Enhancement with Subband InteractionabstractSubband-based approaches process subbands in parallel through the model with shared parameters to learn the commonality of local spectrums for noise reduction. In this way, they have achieved remarkable results with fewer parameters. However, in some complex environments, the lack of global spectral information has a negative impact on the performance of these subband-based approaches. To this end, this paper introduces the subband interaction as a new way to complement the subband model with the global spectral information such as cross-band dependencies and global spectral patterns, and proposes a new lightweight single-channel speech enhancement framework called Interactive Subband Network (Inter-SubNet). Experimental results on DNS Challenge - Interspeech 2021 dataset show that the proposed Inter-SubNet yields a significant improvement over the subband model and outperforms other state-of-the-art speech enhancement approaches, which demonstrate the effectiveness of subband interaction. Jun Chen 0024, Wei Rao 0002, Zilin Wang 0002, Jiuxin Lin, Zhiyong Wu 0001, Yannan Wang, Shidong Shang, Helen M. Meng |
ICASSP | 6 |
| 2023 | Gesper: A Unified Framework for General Speech RestorationabstractThis paper describes the legends-tencent team’s real-time General Speech Restoration (Gesper) system submitted to the ICASSP 2023 Speech Signal Improvement (SSI) Challenge. This newly proposed system is a two-stage architecture, in which the speech restoration is performed, and then followed by speech enhancement. We propose a complex spectral mapping-based generative adversarial network (CSM-GAN) as the speech restoration module for the first time. For noise suppression and dereverberation, the enhancement module is presented with fullband-wideband parallel processing. On the blind test set of ICASSP 2023 SSI Challenge, the proposed Gesper system, which satisfies the real-time condition, achieves 3.27 P.804 overall mean opinion score (MOS) and 3.35 P.835 overall MOS, ranked 1st in both track 1 and track 2. Jun Chen 0024, Yupeng Shi, Wei Rao 0002, Shulin He, Andong Li, Yannan Wang, Zhiyong Wu 0001, Shidong Shang, Chengshi Zheng |
ICASSP | 7 |
| 2023 | Speech Enhancement with Intelligent Neural Homomorphic SynthesisabstractMost neural network speech enhancement models ignore speech production mathematical models by directly mapping Fourier transform spectrums or waveforms. In this work, we propose a neural source filter network for speech enhancement. Specifically, we use homomorphic signal processing and cepstral analysis to obtain noisy speech’s excitation and vocal tract. Unlike traditional signal processing, we use an attentive recurrent network (ARN) model predicted ratio mask to replace the liftering separation function. Then two convolutional attentive recurrent network (CARN) networks are used to predict the excitation and vocal tract of clean speech, respectively. The system’s output is synthesized from the estimated excitation and vocal. Experiments prove that our proposed method performs better, with SI-SNR improving by 1.363dB compared to FullSubNet. Shulin He, Wei Rao 0002, Jun Chen 0024, Yukai Jv, Xueliang Zhang 0001, Yannan Wang, Shidong Shang |
ICASSP | 7 |
| 2023 | TEA-PSE 3.0: Tencent-Ethereal-Audio-Lab Personalized Speech Enhancement System For ICASSP 2023 Dns-ChallengeabstractThis paper introduces the Unbeatable Team’s submission to the ICASSP 2023 Deep Noise Suppression (DNS) Challenge. We expand our previous work, TEA-PSE, to its upgraded version – TEA-PSE 3.0. Specifically, TEA-PSE 3.0 incorporates a residual LSTM after squeezed temporal convolution network (S-TCN) to enhance sequence modeling capabilities. Additionally, the local-global representation (LGR) structure is introduced to boost speaker information extraction, and multi-STFT resolution loss is used to effectively capture the time-frequency characteristics of the speech signals. Moreover, retraining methods are employed based on the freeze training strategy to fine-tune the system. According to the official results, TEA-PSE 3.0 ranks 1st in both ICASSP 2023 DNS-Challenge track 1 and track 2. Yukai Jv, Jun Chen 0024, Shulin He, Wei Rao 0002, Weixin Zhu, Yannan Wang, Shidong Shang |
ICASSP | 7 |
| 2023 | A Multi-Scale Feature Aggregation Based Lightweight Network for Audio-Visual Speech EnhancementabstractAudio-visual speech enhancement (AVSE) was shown to be superior over conventional audio-only counterpart for improving the speech quality. However, most existing AVSE models are heavyweight in the sense of parameter count, which is inappropriate for the deployment and practical applications. In this paper, we therefore present a lightweight AVSE approach (called M3Net) by incorporating several multi-modality, multi-scale and multi-branch strategies. Three multi-scale techniques are designed for the visual and audio streams, including multi-scale average pooling (MSAP), multi-scale ResNet (MSResNet) and multi-scale short time Fourier transform (MSSTFT). It is shown that each multi-scale module positively contributes to the performance. Also, we consider four skip connections for the audio-visual feature aggregation, which have a great complementary effect on the designed multi-scale techniques. Experimental results show that these techniques are flexible in combination with existing approaches, and more importantly obtain a comparable performance with a smaller model size compared to the heavyweight networks. Liangfa Wei, Jie Zhang 0042, Jianming Yang, Yannan Wang, Tian Gao 0005, Li-Rong Dai 0001 |
ICASSP | 5 |
| 2023 | Distance-Based Weight Transfer for Fine-Tuning From Near-Field to Far-Field Speaker VerificationabstractThe scarcity of labeled far-field speech is a constraint for training superior far-field speaker verification systems. In general, fine-tuning the model pre-trained on large-scale near- field speech through a small amount of far-field speech substantially outperforms training from scratch. However, the vanilla fine-tuning suffers from two limitations – catastrophic forgetting and overfitting. In this paper, we propose a weight transfer regularization (WTR) loss to constrain the distance of the weights between the pre-trained model and the fine-tuned model. With the WTR loss, the fine-tuning process takes advantage of the previously acquired discriminative ability from the large-scale near-field speech and avoids catastrophic for- getting. Meanwhile, the analysis based on the PAC-Bayes generalization theory indicates that the WTR loss makes the fine-tuned model have a tighter generalization bound, thus mitigating the overfitting problem. Moreover, three different norm distances for weight transfer are explored, which are L1-norm distance, L2-norm distance, and Max-norm distance. We evaluate the effectiveness of the WTR loss on VoxCeleb (pre-trained) and FFSVC (fine-tuned) datasets. Experimental results show that the distance-based weight transfer fine-tuning strategy significantly outperforms vanilla fine- tuning and other competitive domain adaptation methods. Li Zhang 0084, Qing Wang 0039, Wei Rao 0002, Yannan Wang, Lei Xie 0001 |
ICASSP | 6 |
| 2023 | MC-SpEx: Towards Effective Speaker Extraction with Multi-Scale Interfusion and Conditional Speaker Modulation
Jun Chen 0024, Wei Rao 0002, Zilin Wang 0002, Jiuxin Lin, Yukai Jv, Shulin He, Yannan Wang, Zhiyong Wu 0001 |
INTERSPEECH | 7 |
| 2023 | Gesper: A Restoration-Enhancement Framework for General Speech Reconstruction
Yupeng Shi, Jun Chen 0024, Wei Rao 0002, Shulin He, Andong Li, Yannan Wang, Zhiyong Wu 0001 |
INTERSPEECH | 7 |
| 2022 | TEA-PSE: Tencent-Ethereal-Audio-Lab Personalized Speech Enhancement System for ICASSP 2022 DNS ChallengeabstractThis paper describes Tencent Ethereal Audio Lab – Northwestern Polytechnical University personalized speech enhancement (TEA-PSE) system submitted to track 2 of the ICASSP 2022 Deep Noise Suppression (DNS) challenge. Our system specifically combines the dual-stage network which is a superior real-time speech enhancement framework with the ECAPA-TDNN speaker embedding network which achieves state-of-the-art performance in speaker verification. The dual-stage network aims to decouple the primal speech enhancement problem into multiple easier sub-problems. Specifically, in stage 1, only the magnitude of the target speech is estimated, which is incorporated with the noisy phase to obtain a coarse complex spectrum estimation. To facilitate the formal estimation, in stage 2, an auxiliary network serves as a post-processing module, where residual noise and interfering speech are further suppressed and the phase information is effectively modified. With the asymmetric loss function to penalize over-suppression, more target speech is preserved, which is helpful for both speech recognition performance and subjective sense of hearing. Our system reaches 3.97 in overall audio quality (OVRL) MOS and 0.69 in word accuracy (WAcc) on the blind test set of the challenge, which outperforms the DNS baseline by 0.57 OVRL and ranks 1st in track 2. Yukai Jv, Wei Rao 0002, Xiaopeng Yan, Yihui Fu, Shubo Lv, Luyao Cheng, Yannan Wang, Lei Xie 0001, Shidong Shang |
ICASSP | 7 |
| 2022 | S-DCCRN: Super Wide Band DCCRN with Learnable Complex Feature for Speech EnhancementabstractIn speech enhancement, complex neural network has shown promising performance due to their effectiveness in processing complex-valued spectrum. Most of the recent speech enhancement approaches mainly focus on wide-band signal with a sampling rate of 16K Hz. However, research on super wide band (e.g., 32K Hz) or even full-band (48K) denoising using deep learning is still in its infancy due to the difficulty of modeling more frequency bands and particularly high frequency components. In this paper, we extend our previous deep complex convolution recurrent neural network (DCCRN) substantially to a super wide band version–S-DCCRN, to perform speech denoising on speech of 32K Hz sampling rate. We first employ a cascaded sub-band and full-band processing module, which consists of two small-footprint DCCRNs–one operates on sub-band signal and one operates on full-band signal, aiming at benefiting from both local and global frequency information. Moreover, instead of simply adopting the STFT feature as input, we use a complex feature encoder trained in an end-to-end manner to refine the information of different frequency bands. We also use a complex feature decoder to revert the feature to time-frequency domain. Finally, a learnable spectrum compression method is adopted to adjust the energy of different frequency bands, which is beneficial for neural network learning. The proposed model, S-DCCRN, has surpassed PercepNet as well as several competitive models and achieves state-of-the-art performance in terms of speech quality and intelligibility. Ablation studies also demonstrate the effectiveness of different contributions. Shubo Lv, Yihui Fu, Mengtao Xing, Jiayao Sun, Lei Xie 0001, Yannan Wang |
ICASSP | 7 |
| 2022 | Speech Enhancement with Fullband-Subband Cross-Attention Network
Jun Chen 0024, Wei Rao 0002, Zilin Wang 0002, Zhiyong Wu 0001, Yannan Wang, Shidong Shang, Helen M. Meng |
INTERSPEECH | 5 |
| 2022 | A study of production error analysis for Mandarin-speaking Children with Hearing Impairment
Jingwen Cheng, Yingming Gao, Xiaoli Feng, Yannan Wang, Jinsong Zhang 0001 |
INTERSPEECH | 5 |
| 2022 | TEA-PSE 2.0: Sub-Band Network for Real-Time Personalized Speech EnhancementabstractPersonalized speech enhancement (PSE) utilizes additional cues like speaker embeddings to remove background noise and interfering speech and extract the speech from target speaker. Previous work, the Tencent-Ethereal-Audio-Lab personalized speech enhancement (TEA-PSE) system, ranked 1st in the ICASSP 2022 deep noise suppression (DNS2022) challenge. In this paper, we expand TEA-PSE to its sub-band version - TEA-PSE 2.0, to reduce computational complexity as well as further improve performance. Specifically, we adopt finite impulse response filter banks and spectrum splitting to reduce computational complexity. We introduce a time frequency convolution module (TFCM) to the system for increasing the receptive field with small convolution kernels. Besides, we explore several training strategies to optimize the two-stage network and investigate various loss functions in the PSE task. TEA-PSE 2.0 significantly outperforms TEA-PSE in both speech enhancement performance and computation complexity. Experimental results on the DNS2022 blind test set show that TEA-PSE 2.0 brings 0.102 OVRL personalized DNSMOS improvement with only 21.9% multiply-accumulate operations compared with the previous TEA-PSE. Yukai Jv, Wei Rao 0002, Yannan Wang, Lei Xie 0001, Shidong Shang |
SLT | 4 |
| 2022 | Spatial-DCCRN: DCCRN Equipped with Frame-Level Angle Feature and Hybrid Filtering for Multi-Channel Speech EnhancementabstractRecently, multi-channel speech enhancement has drawn much interest due to the use of spatial information to distinguish target speech from interfering signal. To make full use of spatial information and neural network based masking estimation, we propose a multi-channel denoising neural network - Spatial DCCRN. Firstly, we extend S-DCCRN to multi -channel scenario, aiming at performing cascaded sub-channel and full-channel processing strategy, which can model different channels separately. Moreover, instead of only adopting multi-channel spectrum or concatenating first-channel's magnitude and IPD as the model's inputs, we apply an angle feature extraction module (AFE) to extract frame-level angle feature embeddings, which can help the model to apparently perceive spatial information. Finally, since the phenomenon of residual noise will be more serious when the noise and speech exist in the same time frequency (TF) bin, we particularly design a masking and mapping filtering method to substitute the traditional filter-and-sum operation, with the purpose of cascading coarsely denoising, dereverberation and residual noise suppression. The proposed model, Spatial-DCCRN, has surpassed EaBNet, FasNet as well as several competitive models on the L3DAS22 Challenge dataset. Not only the 3D scenario, Spatial-DCCRN outperforms state-of-the-art (SOTA) model MIMO-UNet by a large margin in multiple evaluation metrics on the multi-channel ConferencingSpeech2021 Challenge dataset. Ablation studies also demonstrate the effectiveness of different contributions. Shubo Lv, Yihui Fu, Yukai Jv, Lei Xie 0001, Weixin Zhu, Wei Rao 0002, Yannan Wang |
SLT | 7 |
| 2021 | Conferencingspeech Challenge: Towards Far-Field Multi-Channel Speech Enhancement for Video ConferencingabstractThe ConferencingSpeech 2021 challenge is proposed to stimulate research on far-field multi-channel speech enhancement for video conferencing. The challenge consists of two separate tasks: 1) Task 1 is multi-channel speech enhancement with single microphone array and focusing on practical application with real-time requirement and 2) Task 2 is multi-channel speech enhancement with multiple distributed micro-phone arrays, which is a non-real-time track and does not have any constraints so that participants could explore any algorithms to obtain high speech quality. Targeting the real video conferencing room application, the challenge database was recorded from real speakers and all recording facilities were located by following the real setup of conferencing room. In this challenge, we open-sourced the list of open source clean speech and noise datasets, simulation scripts, and a baseline system for participants to develop their own system. The final ranking of the challenge will be decided by the subjective evaluation which is performed using Absolute Category Ratings (ACR) to estimate Mean Opinion Score (MOS), speech MOS (S-MOS), and noise MOS (N-MOS). This paper describes the challenge, tasks, datasets, subjective evaluation, and challenge results. The baseline system which is a complex ratio mask based neural network and its experimental results are also presented. Wei Rao 0002, Yihui Fu, Yanxin Hu, Yvkai Jv, Jiangyu Han, Zhongjie Jiang, Lei Xie 0001, Yannan Wang, Shinji Watanabe 0001, Zheng-Hua Tan, Hui Bu, Shidong Shang |
ASRU | 9 |
| 2021 | A Two-Stage Approach to Device-Robust Acoustic Scene ClassificationabstractTo improve device robustness, a highly desirable key feature of a competitive data-driven acoustic scene classification (ASC) system, a novel two-stage system based on fully convolutional neural networks (CNNs) is proposed. Our two-stage system leverages on an ad-hoc score combination based on two CNN classifiers: (i) the first CNN classifies acoustic inputs into one of three broad classes, and (ii) the second CNN classifies the same inputs into one of ten finergrained classes. Three different CNN architectures are explored to implement the two-stage classifiers, and a frequency sub-sampling scheme is investigated. Moreover, novel data augmentation schemes for ASC are also investigated. Evaluated on DCASE 2020 Task 1a, our results show that the proposed ASC system attains a state-of-the-art accuracy on the development set, where our best system, a two-stage fusion of CNN ensembles, delivers a 81.9% average accuracy among multi-device test data, and it obtains a significant improvement on unseen devices. Finally, neural saliency analysis with class activation mapping (CAM) gives new insights on the patterns learnt by our models. Hu Hu, Chao-Han Huck Yang, Xianjun Xia, Yajian Wang, Shutong Niu, Li Chai 0002, Juanjuan Li, Hongning Zhu, Sabato Marco Siniscalchi, Yannan Wang, Jun Du 0002, Chin-Hui Lee 0001 |
ICASSP | 14 |
| 2021 | Embedding Bottleneck Gated Recurrent Unit Network for Radar Signal RecognitionabstractRadar signal recognition plays an significant role in civil applications. Corresponding to two types of intentional modulation signal and unintentional fingerprint signal, radar signal recognition has two kinds of tasks—automatic modulation classification and radar emitter identification. In this paper, we propose a Embedding Bottleneck Gated Recurrent Unit (EBGRU) network that can handle these two tasks separately. The EBGRU consists of three main processing steps. Firstly, the normalized signal pulses are trained in pulse embedding network containing several embedding methods: Pulse2Vec, GloVeP and EPMo, during which we regard the radar signal pulses as radar signal-linguistic sequences for the first time. Then, pulses embeddings are added to original pulses and are sampled to form latent representations of pulses through information bottleneck. Finally, the gated recurrent unit network is utilized to predict radar signal labels. Experiment results show that the proposed method has reached 95.33% on simulated modulation signals and 94.67% at real intercepted emitter signals with relatively less network parameters. Yannan Wang, Guitao Cao, Danning Su |
IJCNN | 1 |
| 2021 | Improving Channel Decorrelation for Multi-Channel Target Speech ExtractionabstractTarget speech extraction has attracted widespread attention. When microphone arrays are available, the additional spatial information can be helpful in extracting the target speech. We have recently proposed a channel decorrelation (CD) mechanism to extract the inter-channel differential information to enhance the reference channel encoder representation. Although the proposed mechanism has shown promising results for extracting the target speech from mixtures, the extraction performance is still limited by the nature of the original decorrelation theory. In this paper, we propose two methods to broaden the horizon of the original channel decorrelation, by replacing the original softmax-based inter-channel similarity between encoder representations, using an unrolled probability and a normalized cosine-based similarity at the dimensional-level. Moreover, new combination strategies of the CD-based spatial information and target speaker adaptation of parallel encoder outputs are also investigated. Experiments on the reverberant WSJ0 2-mix show that the improved CD can result in more discriminative differential information and the new adaptation strategy is also very effective to improve the target speech extraction. Jiangyu Han, Wei Rao 0002, Yannan Wang, Yanhua Long |
Interspeech | 3 |
| 2020 | Tensor-To-Vector Regression for Multi-Channel Speech Enhancement Based on Tensor-Train NetworkabstractWe propose a tensor-to-vector regression approach to multi-channel speech enhancement in order to address the issue of input size explosion and hidden-layer size expansion. The key idea is to cast the conventional deep neural network (DNN) based vector-to-vector regression formulation under a tensor-train network (TTN) framework. TTN is a recently emerged solution for compact representation of deep models with fully connected hidden layers. Thus TTN maintains DNN's expressive power yet involves a much smaller amount of trainable parameters. Furthermore, TTN can handle a multi-dimensional tensor input by design, which exactly matches the desired setting in multi-channel speech enhancement. We first provide a theoretical extension from DNN to TTN based regression. Next, we show that TTN can attain speech enhancement quality comparable with that for DNN but with much fewer parameters, e.g., a reduction from 27 million to only 5 million parameters is observed in a single-channel scenario. TTN also improves PESQ over DNN from 2.86 to 2.96 by slightly increasing the number of trainable parameters. Finally, in 8-channel conditions, a PESQ of 3.12 is achieved using 20 million parameters for TTN, whereas a DNN with 68 million parameters can only attain a PESQ of 3.06. Jun Qi 0002, Hu Hu, Yannan Wang, Chao-Han Huck Yang, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
ICASSP | 3 |
| 2020 | Geometry Constrained Progressive Learning for Lstm-Based Speech EnhancementabstractIn our previous work, a progressive learning framework for long short-term memory (LSTM)-based speech enhancement was proposed to improve the performance in low SNR environment, where each LSTM layer is guided to learn an intermediate target with a specific SNR gain via the MMSE criterion. However, the constraint relationship among these targets is not considered in the objective function. In this paper, we incorporate two kinds of geometric constraints among these targets into the objective function to help LSTM achieve better training. One constraint is edge constraint and the other is the centroid constraint. In addition, we propose a method for constructing the intermediate targets online. It saves device storage space and alleviates the trouble of manually constructing intermediate targets. Experiment results demonstrate these geometric constraints can bring remarkable improvements in low SNR environments. Jun Du 0002, Li Chai 0002, Yannan Wang, Qing Wang 0008, Chin-Hui Lee 0001 |
ICASSP | 4 |
| 2020 | Audio Sound Determination Using Feature Space Attention Based Convolution Recurrent Neural NetworkabstractThe classification framework has been popularly adopted to perform sound event detection. However, the existing neural network based classification based approaches treat each feature dimension equally and the varying influence of feature dimensions has not been taken into consideration. To deal with this, we propose a feature space attention based convolution recurrent neural network approach utilizing the varying importance of each feature dimension to perform acoustic event detection. The convolution layers are used to extract the high level information from the audio signals. Then the feature space attention scheme is applied to the extracted features to automatically determine the importance of each feature dimension. Experimental results on the latest TUT Sound Event 2017 dataset demonstrate the improved performance of the proposed approach compared to the existing acoustic event detection systems. Xianjun Xia, Jingjing Pan, Yannan Wang |
ICASSP | 3 |
| 2020 | An Acoustic Segment Model Based Segment Unit Selection Approach to Acoustic Scene Classification with Partial UtterancesabstractIn this paper, we propose a sub-utterance unit selection framework to remove acoustic segments in audio recordings that carry little information for acoustic scene classification (ASC). Our approach is built upon a universal set of acoustic segment units covering the overall acoustic scene space. First, those units are modeled with acoustic segment models (ASMs) used to tokenize acoustic scene utterances into sequences of acoustic segment units. Next, paralleling the idea of stop words in information retrieval, stop ASMs are automatically detected. Finally, acoustic segments associated with the stop ASMs are blocked, because of their low indexing power in retrieval of most acoustic scenes. In contrast to building scene models with whole utterances, the ASM-removed sub-utterances, i.e., acoustic utterances without stop acoustic segments, are then used as inputs to the AlexNet-L back-end for final classification. On the DCASE 2018 dataset, scene classification accuracy increases from 68%, with whole utterances, to 72.1%, with segment selection. This represents a competitive accuracy without any data augmentation, and/or ensemble strategy. Moreover, our approach compares favourably to AlexNet-L with attention. Hu Hu, Sabato Marco Siniscalchi, Yannan Wang, Jun Du 0002, Chin-Hui Lee 0001 |
INTERSPEECH | 3 |
| 2020 | Relational Teacher Student Learning with Neural Label Embedding for Device Adaptation in Acoustic Scene ClassificationabstractIn this paper, we propose a domain adaptation framework to address the device mismatch issue in acoustic scene classification leveraging upon neural label embedding (NLE) and relational teacher student learning (RTSL). Taking into account the structural relationships between acoustic scene classes, our proposed framework captures such relationships which are intrinsically device-independent. In the training stage, transferable knowledge is condensed in NLE from the source domain. Next in the adaptation stage, a novel RTSL strategy is adopted to learn adapted target models without using paired source-target data often required in conventional teacher student learning. The proposed framework is evaluated on the DCASE 2018 Task1b data set. Experimental results based on AlexNet-L deep classification models confirm the effectiveness of our proposed approach for mismatch situations. NLE-alone adaptation compares favourably with the conventional device adaptation and teacher student based adaptation techniques. NLE with RTSL further improves the classification accuracy Hu Hu, Sabato Marco Siniscalchi, Yannan Wang, Chin-Hui Lee 0001 |
INTERSPEECH | 3 |
| 2020 | Exploring Deep Hybrid Tensor-to-Vector Network Architectures for Regression Based Speech EnhancementabstractThis paper investigates different trade-offs between the number of model parameters and enhanced speech qualities by employing several deep tensor-to-vector regression models for speech enhancement. We find that a hybrid architecture, namely CNN-TT, is capable of maintaining a good quality performance with a reduced model parameter size. CNN-TT is composed of several convolutional layers at the bottom for feature extraction to improve speech quality and a tensor-train (TT) output layer on the top to reduce model parameters. We first derive a new upper bound on the generalization power of the convolutional neural network (CNN) based vector-to-vector regression models. Then, we provide experimental evidence on the Edinburgh noisy speech corpus to demonstrate that, in single-channel speech enhancement, CNN outperforms DNN at the expense of a small increment of model sizes. Besides, CNN-TT slightly outperforms the CNN counterpart by utilizing only 32% of the CNN model parameters. Besides, further performance improvement can be attained if the number of CNN-TT parameters is increased to 44% of the CNN model size. Finally, our experiments of multi-channel speech enhancement on a simulated noisy WSJ0 corpus demonstrate that our proposed hybrid CNN-TT architecture achieves better results than both DNN and CNN models in terms of better-enhanced speech qualities and smaller parameter sizes. Jun Qi 0002, Hu Hu, Yannan Wang, Chao-Han Huck Yang, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
INTERSPEECH | 3 |
| 2020 | An Attention Based Speaker-Independent Audio-Visual Deep Learning Model for Speech Enhancement
Yannan Wang |
MMM (2) | 2 |
| 2020 | Locality-constrained feature space learning for cross-resolution sketch-photo face recognition
Guangwei Gao, Yannan Wang, Heyou Chang, Huimin Lu 0001, Dong Yue 0001 |
Multim. Tools Appl. | 2 |
| 2017 | A Maximum Likelihood Approach to Deep Neural Network Based Nonlinear Spectral Mapping for Single-Channel Speech Separation
Yannan Wang, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 1 |
| 2017 | A Gender Mixture Detection Approach to Unsupervised Single-Channel Speech Separation Based on Deep Neural NetworksabstractWe propose an unsupervised speech separation framework for mixtures of two unseen speakers in a single-channel setting based on deep neural networks (DNNs). We rely on a key assumption that two speakers could be well segregated if they are not too similar to each other. A dissimilarity measure between two speakers is first proposed to characterize the separation ability between competing speakers. We then show that speakers with the same or different genders can often be separated if two speaker clusters, with large enough distances between them, for each gender group could be established, resulting in four speaker clusters. Next, a DNN-based gender mixture detection algorithm is proposed to determine whether the two speakers in the mixture are females, males, or from different genders. This detector is based on a newly proposed DNN architecture with four outputs, two of them representing the female speaker clusters and the other two characterizing the male groups. Finally, we propose to construct three independent speech separation DNN systems, one for each of the female-female, male-male, and female-male mixture situations. Each DNN gives dual outputs, one representing the target speaker group and the other characterizing the interfering speaker cluster. Trained and tested on the speech separation challenge corpus, our experimental results indicate that the proposed DNN-based approach achieves large performance gains over the state-of-the-art unsupervised techniques without using any specific knowledge about the mixed target and interfering speakers being segregated. Yannan Wang, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2015 | High-resolution acoustic modeling and compact language modeling of language-universal speech attributes for spoken language identification
Yannan Wang, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 1 |