Shuai Nie 0001

dblp:42/10187-1 · DBLP profile ↗
← Back
22ranked-venue papers
6as first author
7since 2021 · last 2024
0000-0002-8078-6829ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 3 since 2021
YearPublicationVenuePosition
2024 BViT: Broad Attention-Based Vision Transformer
abstract
Recent works have demonstrated that transformer can achieve promising performance in computer vision, by exploiting the relationship among image patches with self-attention. They only consider the attention in a single feature layer, but ignore the complementarity of attention in different layers. In this article, we propose broad attention to improve the performance by incorporating the attention relationship of different layers for vision transformer (ViT), which is called BViT. The broad attention is implemented by broad connection and parameter-free attention. Broad connection of each transformer layer promotes the transmission and integration of information for BViT. Without introducing additional trainable parameters, parameter-free attention jointly focuses on the already available attention information in different layers for extracting useful information and building their relationship. Experiments on image classification tasks demonstrate that BViT delivers superior accuracy of 75.0%/81.6% top-1 accuracy on ImageNet with 5M/22M parameters. Moreover, we transfer BViT to downstream object recognition benchmarks to achieve 98.9% and 89.9% on CIFAR10 and CIFAR100, respectively, that exceed ViT with fewer parameters. For the generalization test, the broad attention in Swin Transformer, T2T-ViT and LVT also brings an improvement of more than 1%. To sum up, broad attention is promising to promote the performance of attention-based models. Code and pretrained models are available at https://github.com/DRL/BViT.
Nannan Li 0003, Yaran Chen, Weifan Li, Zixiang Ding, Dongbin Zhao, Shuai Nie 0001
IEEE Trans. Neural Networks Learn. Syst.6
2022 ADD 2022: the first Audio Deep Synthesis Detection Challenge
abstract
Audio deepfake detection is an emerging topic, which was included in the ASVspoof 2021. However, the recent shared tasks have not covered many real-life and challenging scenarios. The first Audio Deep synthesis Detection challenge (ADD) was motivated to fill in the gap. The ADD 2022 includes three tracks: low-quality fake audio detection (LF), partially fake audio detection (PF) and audio fake game (FG). The LF track focuses on dealing with bona fide and fully fake utterances with various real-world noises etc. The PF track aims to distinguish the partially fake audio from the real. The FG track is a rivalry game, which includes two tasks: an audio generation task and an audio fake detection task. In this paper, we describe the datasets, evaluation metrics, and protocols. We also report major findings that reflect the recent advances in audio deepfake detection tasks.
Jiangyan Yi, Ruibo Fu, Jianhua Tao 0001, Shuai Nie 0001, Haoxin Ma, Chenglong Wang 0001, Tao Wang 0074, Zhengkun Tian, Ye Bai 0001, Cunhang Fan, Shan Liang 0007, Shuai Zhang 0014, Xinrui Yan, Zhengqi Wen, Haizhou Li 0001
ICASSP4
2022 A Robust Deep Audio Splicing Detection Method via Singularity Detection Feature
abstract
There are many methods for detecting forged audio produced by conversion and synthesis. However, as a simpler method of forgery, splicing has not attracted widespread attention. Based on the characteristic that the tampering operation will cause singularities at high-frequency components, we propose a high-frequency singularity detection feature obtained by wavelet transform. The proposed feature can explicitly show the location of the tampering operation on the waveform. Moreover, the long short-term memory (LSTM) is introduced to the CNN-architecture LCNN to ensure that the sequence information can be fully learned. The proposed feature is sent to the improved RNN-architecture LCNN together with the widely used linear frequency cepstral coefficients (LFCC) to learn forgery characteristics where the LFCC is used as a supplement. Systematic evaluation and comparison show that the proposed method has greatly improved the accuracy and generalization.
Kanghao Zhang, Shan Liang 0007, Shuai Nie 0001, Shulin He, Xueliang Zhang 0001, Haoxin Ma, Jiangyan Yi
ICASSP3
2022 Speaker recognition-assisted robust audio deepfake detection
Shuai Nie 0001, Hui Zhang 0031, Shulin He, Kanghao Zhang, Shan Liang 0007, Xueliang Zhang 0001, Jianhua Tao 0001
INTERSPEECH2
2021 Deep neural network-based generalized sidelobe canceller for dual-channel far-field speech recognition
Guanjun Li, Shan Liang 0001, Shuai Nie 0001, Zhanlei Yang
Neural Networks3
2021 Exploiting the directional coherence function for multichannel source extraction
Shan Liang 0007, Guanjun Li, Shuai Nie 0001, Zhanlei Yang, Jianhua Tao 0001
Speech Commun.3
2021 Robust Text Image Recognition via Adversarial Sequence-to-Sequence Domain Adaptation
abstract
Robust text reading is a very challenging problem, due to the distribution of text images changing significantly in real-world scenarios. One effective solution is to align the distribution between different domains by domain adaptation methods. However, we found that these methods might struggle when dealing sequence-like text images. An important reason is that conventional domain adaptation methods strive to align images as a whole, while text images consist of variable-length fine-grained character information. To address this issue, we propose a novel Adversarial Sequence-to-Sequence Domain Adaptation (ASSDA) method to learn "where to adapt" and "how to align" the sequential image. Our key idea is to mine the local regions that contain characters, and focus on aligning them across domains in an adversarial manner. Extensive text recognition experiments show the ASSDA could efficiently transfer sequence knowledge and validate the promising power towards the various domain shift in the real world applications.
Shuai Nie 0001, Shan Liang 0001
IEEE Trans. Image Process.2
2019 Loss and Double-edge-triggered Detector for Robust Small-footprint Keyword Spotting
abstract
Keyword spotting (KWS) system constitutes a critical component of human-computer interfaces, which detects the specific keyword from a continuous stream of audio. The goal of KWS is providing a high detection accuracy at a low false alarm rate while having small memory and computation requirements. The DNN-based KWS system faces a large class imbalance during training because the amount of data available for the keyword is usually much less than the background speech, which overwhelms training and leads to a degenerate model. In this paper, we explore the focal loss for the training of a small-footprint KWS system. It can automatically down-weight the contribution of easy samples during training and focus the model on hard samples, which naturally solves the class imbalance and allows us to efficiently utilize all data available. Furthermore, many keywords of Chinese conversational assistants are repeated words due to the idiomatic usage, such as `XIAO DU XIAO DU'. We propose a double-edge-triggered detecting method for the repeated keyword, which significantly reduces the false alarm rate relative to the single threshold method. Systematic experiments demonstrate significant further improvements compared to the baseline system.
Bin Liu 0041, Shuai Nie 0001, Shan Liang 0007, Zhanlei Yang
ICASSP2
2019 Direction-Aware Speaker Beam for Multi-Channel Speaker Extraction
Guanjun Li, Shan Liang 0007, Shuai Nie 0001, Meng Yu 0003, Lianwu Chen, Shouye Peng, Changliang Li
INTERSPEECH3
2019 Jointly Adversarial Enhancement Training for Robust End-to-End Speech Recognition
Bin Liu 0041, Shuai Nie 0001, Shan Liang 0007, Meng Yu 0003, Lianwu Chen, Shouye Peng, Changliang Li
INTERSPEECH2
2018 Boosting Noise Robustness of Acoustic Model via Deep Adversarial Training
abstract
In realistic environments, speech is usually interfered by various noise and reverberation, which dramatically degrades the performance of automatic speech recognition (ASR) systems. To alleviate this issue, the commonest way is to use a well-designed speech enhancement approach as the front-end of ASR. However, more complex pipelines, more computations and even higher hardware costs (microphone array) are additionally consumed for this kind of methods. In addition, speech enhancement would result in speech distortions and mismatches to training. In this paper, we propose an adversarial training method to directly boost noise robustness of acoustic model. Specifically, a jointly compositional scheme of generative adversarial net (GAN) and neural network-based acoustic model (AM) is used in the training phase. GAN is used to generate clean feature representations from noisy features by the guidance of a discriminator that tries to distinguish between the true clean signals and generated signals. The joint optimization of generator, discriminator and AM concentrates the strengths of both GAN and AM for speech recognition. Systematic experiments on CHiME-4 show that the proposed method significantly improves the noise robustness of AM and achieves the average relative error rate reduction of 23.38% and 11.54% on the development and test set, respectively.
Bin Liu 0041, Shuai Nie 0001, Dengfeng Ke, Shan Liang 0007
ICASSP2
2018 Stochastic Multiple Choice Learning for Acoustic Modeling
abstract
Even for deep neural networks, it is still a challenging task to indiscriminately model thousands of fine-grained senones only by one model. Ensemble learning is a well-known technique that is capable of concentrating the strengths of different models to facilitate the complex task. In addition, the phones may be spontaneously aggregated into several clusters due to the intuitive perceptual properties of speech, such as vowels and consonants. However, a typical ensemble learning scheme usually trains each submodular independently and doesn't explicitly consider the internal relation of data, which is hardly expected to improve the classification performance of fine-grained senones. In this paper, we use a novel training schedule for DNN-based ensemble acoustic model. In the proposed training schedule, all submodels are jointly trained to cooperatively optimize the loss objective by a Stochastic Multiple Choice Learning approach. It results in that different submodels have specialty capacities for modeling senones with different properties. Systematic experiments show that the proposed model is competitive with the dominant DNN-based acoustic models in the TIMIT and THCHS-30 recognition tasks.
Bin Liu 0041, Shuai Nie 0001, Shan Liang 0007, Zhanlei Yang
IJCNN2
2018 Deep Noise Tracking Network: A Hybrid Signal Processing/Deep Learning Approach to Speech Enhancement
Shuai Nie 0001, Shan Liang 0007, Bin Liu 0041, Jianhua Tao 0001
INTERSPEECH1
2018 Robust offline handwritten character recognition through exploring writer-independent features under the guidance of printed data
Shuai Nie 0001, Shouye Peng
Pattern Recognit. Lett.3
2018 Deep Learning Based Speech Separation via NMF-Style Reconstructions
abstract
Deep learning based speech separation usually uses a supervised algorithm to learn a mapping function from noisy features to separation targets. These separation targets, either ideal masks or magnitude spectrograms, have prominent spectro-temporal structures. Nonnegative matrix factorization (NMF) is a well-known representation learning technique that is capable of capturing the basic spectral structures. Therefore, the combination of deep learning and NMF as an organic whole is a smart strategy. However, previous methods typically use deep neural networks (DNN) and NMF for speech separation in a separate manner. In this paper, we propose a jointly combinatorial scheme to concentrate the strengths of both DNN and NMF for speech separation. NMF is used to learn the basis spectra that then are integrated into a DNN to directly reconstruct the magnitude spectrograms of speech and noise. Instead of predicting activation coefficients inferred by NMF, which is used as an intermediate target by the previous methods, DNN directly optimizes an actual separation objective in our system, so that the accumulated errors could be alleviated. Moreover, we explore a discriminative training objective with sparsity constraints to suppress noise and preserve more speech components further. Systematic experiments show that the proposed models are competitive with the previous methods.
Shuai Nie 0001, Shan Liang 0001, Xueliang Zhang 0001, Jianhua Tao 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2016 Exploiting spectro-temporal structures using NMF for DNN-based supervised speech separation
abstract
The targets of speech separation, whether ideal masks or magnitude spectrograms of interest, have prominent spectro-temporal structures. These characteristics are very worthy to be exploited for speech separation, however, they are usually ignored in previous works. In this paper, we use nonnegative matrix factorization (NMF) to exploit the spectro-temporal structures of magnitude spectrograms. With nonnegative constrains, NMF can capture the basis spectra patterns of speech and noise. Then the learned basis spectra are integrated into a deep neural network (DNN) to reconstruct the magnitude spectrograms of speech and noise with their nonnegative linear combination. Using the reconstructed spectrograms, we further explore a discriminative training objective and a joint optimization framework for the proposed model. Systematic experiments show that the proposed model is competitive with the previous methods in monaural speech separation tasks.
Shuai Nie 0001, Hao Li 0046, Xueliang Zhang 0001, Zhanlei Yang, Like Dong
ICASSP1
2016 Jointly Optimizing Activation Coefficients of Convolutive NMF Using DNN for Speech Separation
Hao Li 0046, Shuai Nie 0001, Xueliang Zhang 0001, Hui Zhang 0031
INTERSPEECH2
2016 A Pairwise Algorithm Using the Deep Stacking Network for Speech Separation and Pitch Estimation
abstract
Speech separation and pitch estimation in noisy conditions are considered to be a “chicken-and-egg” problem. On one hand, pitch information is an important cue for speech separation. On the other hand, speech separation makes pitch estimation easier when background noise is removed. In this paper, we propose a supervised learning architecture to solve these two problems iteratively. The proposed algorithm is based on the deep stacking network (DSN), which provides a method for stacking simple processing modules to build deep architectures. Each module is a classifier whose target is the ideal binary mask (IBM), and the input vector includes spectral features, pitch-based features and the output from the previous module. During the testing stage, we estimate the pitch using the separation results and update the pitch-based features to the next module. When embedded into the DSN, pitch estimation and speech separation each run several times. We obtain the final results from the last module. Systematic evaluations show that the proposed system results in both a high quality estimated binary mask and accurate pitch estimation and outperforms recent systems in its generalization ability.
Xueliang Zhang 0001, Hui Zhang 0031, Shuai Nie 0001, Guanglai Gao
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 A pairwise algorithm for pitch estimation and speech separation using deep stacking network
abstract
Pitch information is an important cue for speech separation. However, pitch estimation in noisy condition is also a task as challenging as speech separation. In this paper, we propose a supervised learning architecture which combines these two problems concisely. The proposed algorithm is based on deep stacking network (DSN) which provides a method of stacking simple processing modules in building deep architecture. In the training stage, an ideal binary mask is used as target. The input vector includes the outputs of lower module and frame-level features which consist of spectral and pitch-based features. In the testing stage, each module provides an estimated binary mask which is employed to re-estimate pitch. Then we update the pitch-based features to the next module. This procedure is embedded iteratively in DSN, and we obtain the final separation results from the last module of DSN. Systematic evaluations show that the proposed approach produces high quality estimated binary mask and outperforms recent systems in generalization.
Hui Zhang 0031, Xueliang Zhang 0001, Shuai Nie 0001, Guanglai Gao
ICASSP3
2015 Two-stage multi-target joint learning for monaural speech separation
abstract
Recently, supervised speech separation has been extensively studied and shown considerable promise. Due to the temporal continuity of speech, speech auditory features and separation targets present prominent spectro-temporal structures and strong correlations over the time-frequency (T-F) domain, which can be exploited for speech separation. However, many supervised speech separation methods independently model each T-F unit with only one target and much ignore these useful information. In this paper, we propose a two-stage multi-target joint learning method to jointly model the related speech separation targets at the frame level. Systematic experiments show that the proposed approach consistently achieves better separation and generalization performances in the low signal-to-noise ratio(SNR) conditions.
Shuai Nie 0001, Xueliang Zhang 0001, Like Dong
INTERSPEECH1
2015 Joint optimization of recurrent networks exploiting source auto-regression for source separation
abstract
In music interferences condition, source separation is very difficult. In this paper, we propose a novel recurrent network exploiting the auto-regressions of speech and music interference for source separation. An auto-regression can capture the shortterm temporal dependencies in data to help the source separation. For the separation, we independently separate the magnitude spectra of speech and interference from the mixture spectra by including an extra masking layer in the recurrent network. Compared to directly evaluating the ideal mask, the extra masking layer relaxes the assumption of independence between speech and interference which is more suitable for the realworld environments. Using the separated spectra of speech and interference, we further explore a discriminative training objective and joint optimization framework for the proposed network, which incorporates the correlations and spectral dependencies of speech and interference into the separation. Systematic experiments show that the proposed model is competitive with the state-of-the-art method in singing-voice separations.
Shuai Nie 0001, Xueliang Zhang 0001, Liwei Qiao
INTERSPEECH1
2014 Deep stacking networks with time series for speech separation
abstract
In many present speech separation approaches, the separation task is formulated as a binary classification problem. Several classification-based approaches have been proposed and performed satisfactorily. However, they do not explicitly model the correlation in time and each time-frequency (T-F) unit is still classified individually. As we know, the speech signal has a very rich time series and temporal dynamic information that can be exploited for speech separation. In this study, we incorporate the correlation in time into classification. Compared with the previous approaches, the proposed approach achieves better separation and generalization performance by using deep stacking networks (DSN) with time series and re-threshold method.
Shuai Nie 0001, Hui Zhang 0031, Xueliang Zhang 0001
ICASSP1