Dengfeng Ke

dblp:62/8051 · DBLP profile ↗
← Back
29ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0001-8459-0412ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 DDP-Unet: A mapping neural network for single-channel speech enhancement
Haoxiang Chen 0009, Yanyan Xu 0001, Dengfeng Ke, Kaile Su
Comput. Speech Lang.3
2024 Concss: Contrastive-based Context Comprehension for Dialogue-Appropriate Prosody in Conversational Speech Synthesis
abstract
Conversational speech synthesis (CSS) incorporates historical dialogue as supplementary information with the aim of generating speech that has dialogue-appropriate prosody. While previous methods have already delved into enhancing context comprehension, context representation still lacks effective representation capabilities and context-sensitive discriminability. In this paper, we introduce a contrastive learning-based CSS framework, CONCSS. Within this framework, we define an innovative pretext task specific to CSS that enables the model to perform self-supervised learning on unlabeled conversational datasets to boost the model’s context understanding. Additionally, we introduce a sampling strategy for negative sample augmentation to enhance context vectors’ discriminability. This is the first attempt to integrate contrastive learning into CSS. We conduct ablation studies on different contrastive learning strategies and comprehensive experiments in comparison with prior CSS systems. Results demonstrate that the synthesized speech from our proposed method exhibits more contextually appropriate and sensitive prosody.
Yayue Deng, Jinlong Xue, Yukang Jia, Yichen Han, Fengping Wang, Yingming Gao, Dengfeng Ke, Ya Li 0001
ICASSP8
2024 PLDE: A lightweight pooling layer for spoken language recognition
abstract
In recent years, the transfer learning method of replacing acoustic features with phonetic features has become a new paradigm for end-to-end spoken language recognition. However, these larger transfer learning models always encode too much redundant information. In this paper, we propose a lightweight language recognition decoder based on a phonetic learnable dictionary encoding (PLDE) layer, which is more suitable for phonetic features and achieves better recognition performances while significantly reducing the number of parameters. The lightweight decoder consists of three main parts: (1) a phonetic learnable dictionary with ghost clusters, which improves the traditional LDE pooling layer and enhances the model’s ability to model noise with ghost clusters; (2) coarse-grained chunk-level pooling, which can highlight the phone sequence and suppress noise around ghost clusters, and hence reduce their influence to the subsequent network; (3) fine-grained chunk-level projection, which enables the discriminative network to obtain more linguistic information and hence improve the model’s modelling ability. These three parts simplify the language recognition decoder into a PLDE pooling layer, reducing the parameter size of the decoder by at least one order of magnitude while achieving better recognition performances. In experiments on the OLR2020 dataset, the C a v g of the proposed method exceeds that of the current state-of-the-art language recognition system, achieving 24.68% and 42.24% improvements on the cross-channel test set and unknown noise test set, respectively. Furthermore, experimental results on the OLR2021 dataset also demonstrate the effectiveness of PLDE.
Zimu Li, Yanyan Xu 0001, Dengfeng Ke, Kaile Su
Speech Commun.3
2023 LIMI-VC: A Light Weight Voice Conversion Model with Mutual Information Disentanglement
abstract
Voice conversion(VC) model aims to convert the source timbre to the target one. Recently, many VC models utilize pre-trained models to enhance the performance and achieve good results. However, pre-trained models could not somehow disentangle the timbre and linguistic information, thus resulting in a redundancy, which may hurt the conversion performance. In this paper we proposed LIMI-VC, reducing the redundancy between the linguistic content and the timbre information with mutual information disentanglement. We design the model in a light weight form, for the sake of parameter and computation efficiency when pre-trained models are commonly used nowadays. Experiments show that the proposed model can still improve the performance, with 15 times smaller size, compared to baseline. An out-of-domain cross-lingual inference also shows that our model greatly outperforms the baseline. Our source code and audio examples will be available at: https://github.com/WongLaw/LIMI-VC.
Liangjie Huang, Yunming Liang, Can Wen, Yanlu Xie, Jinsong Zhang 0001, Dengfeng Ke
ICASSP8
2023 Dual Audio Encoders Based Mandarin Prosodic Boundary Prediction by Using Multi-Granularity Prosodic Representations
Ruishan Li, Yingming Gao, Yanlu Xie, Dengfeng Ke, Jinsong Zhang 0001
INTERSPEECH4
2022 StyleFormerGAN-VC:Improving Effect of few shot Cross-Lingual Voice Conversion Using VAE-StarGAN and Attention-AdaIN
abstract
Voice Conversion (VC) aims to transfer the speaker timbre while retaining the lexical content of the source speech and has attracted much attention lately. Although previous VC models have achieved good performance, unstability can not be avoided when it comes cross-lingual scenario. In this paper, we propose the StyleFormerGAN-VC to achieve better cross language speech conversion, where variational auto-encoder is introduced to model the feature distribution of the cross-lingual utterances and adversarial training is applied to elevate the speech quality. In addition, we combine the Attention mechanism and AdaIN to make our model more generalized to unseen speaker with long utterance. Experiments show that our model performs stably in the cross-lingual scenario and gains well MOS evaluation scores.
Dengfeng Ke, Wenhan Yao, Ruixin Hu, Liangjie Huang, Wentao Shu
SNPD1
2021 A Study on Fine-Tuning wav2vec2.0 Model for the Task of Mispronunciation Detection and Diagnosis
Linkai Peng, Kaiqi Fu, Binghuai Lin, Dengfeng Ke, Jinsong Zhang 0001
Interspeech4
2021 WINVC: One-Shot Voice Conversion with Weight Adaptive Instance Normalization
Shengjie Huang, Yanyan Xu 0001, Dengfeng Ke, Thomas Hain
PRICAI (2)4
2021 μ-law SGAN for generating spectra with more details in speech enhancement
Yanyan Xu 0001, Dengfeng Ke, Kaile Su
Neural Networks3
2020 Dynamically Mitigating Data Discrepancy with Balanced Focal Loss for Replay Attack Detection
abstract
It becomes urgent to design effective anti-spoofing algorithms for vulnerable automatic speaker verification systems due to the advancement of high-quality playback devices. Current studies mainly treat anti-spoofing as a binary classification problem between bonafide and spoofed utterances, while lack of indistinguishable samples makes it difficult to train a robust spoofing detector. In this paper, we argue that for anti-spoofing, it needs more attention for indistinguishable samples over easily-classified ones in the modeling process, to make correct discrimination a top priority. Therefore, to mitigate the data discrepancy between training and inference, we propose to leverage a balanced focal loss function as the training objective to dynamically scale the loss based on the traits of the sample itself. Besides, in the experiments, we select three kinds of features that contain both magnitude-based and phase-based information to form complementary and informative features. Experimental results on the ASVspoof2019 dataset demonstrate the superiority of the proposed methods by comparison between our systems and top-performing ones. Systems trained with the balanced focal loss perform significantly better than conventional cross-entropy loss. With complementary features, our fusion system with only three kinds of features outperforms other systems containing five or more complex single models by 22.5% for min-tDCF and 7% for EER, achieving a min-tDCF and an EER of 0.0124 and 0.55% respectively. Furthermore, we present and discuss the evaluation results on real replay data apart from the simulated ASVspoof2019 data, indicating that research for anti-spoofing still has a long way to go.
Yongqiang Dou, Yanyan Xu 0001, Dengfeng Ke
ICPR5
2020 Formant Tracking Using Dilated Convolutional Networks Through Dense Connection with Gating Mechanism
abstract
Formant tracking is one of the most fundamental problems in speech processing.Traditionally, formants are estimated using signal processing methods.Recent studies showed that generic convolutional architectures can outperform recurrent networks on temporal tasks such as speech synthesis and machine translation.In this paper, we explored the use of Temporal Convolutional Network (TCN) for formant tracking.In addition to the conventional implementation, we modified the architecture from three aspects.First, we turned off the "causal" mode of dilated convolution, making the dilated convolution see the future speech frames.Second, each hidden layer reused the output information from all the previous layers through dense connection.Third, we also adopted a gating mechanism to alleviate the problem of gradient disappearance by selectively forgetting unimportant information.The model was validated on the open access formant database VTR.The experiment showed that our proposed model was easy to converge and achieved an overall mean absolute percent error (MAPE) of 8.2% on speech-labeled frames, compared to three competitive baselines of 9.4% (LSTM), 9.1% (Bi-LSTM) and 8.9% (TCN).
Wang Dai, Jinsong Zhang 0001, Yingming Gao, Dengfeng Ke, Binghuai Lin, Yanlu Xie
INTERSPEECH5
2020 Improving speech enhancement by focusing on smaller values using relative loss
abstract
The task of single‐channel speech enhancement is to restore clean speech from noisy speech. Recently, speech enhancement has been greatly improved with the introduction of deep learning. Previous work proved that using ideal ratio mask or phase‐sensitive mask as intermediation to recover clean speech can yield better performance. In this case, the mean square error is usually selected as the loss function. However, after conducting experiments, the authors find that the mean square error has a problem. It considers absolute error values, meaning that the gradients of the network depend on absolute differences between estimated values and true values, so the points in magnitude spectra with smaller values contribute little to the gradients. To solve this problem, they propose relative loss, which pays more attention to relative differences between magnitude spectra, rather than the absolute differences, and is more in accordance with human sensory characteristics. The perceptual evaluation of speech quality, the short‐time objective intelligibility, the signal‐to‐distortion ratio, and the segmental signal‐to‐noise ratio are used to evaluate the performance of the relative loss. Experimental results show that it can greatly improve speech enhancement by focusing on smaller values.
Yanyan Xu 0001, Dengfeng Ke, Kaile Su
IET Signal Process.3
2019 Trainable back-propagated functional transfer matrices
Yanyan Xu 0001, Dengfeng Ke, Kaile Su, Jing Sun 0002
Appl. Intell.3
2018 Boosting Noise Robustness of Acoustic Model via Deep Adversarial Training
abstract
In realistic environments, speech is usually interfered by various noise and reverberation, which dramatically degrades the performance of automatic speech recognition (ASR) systems. To alleviate this issue, the commonest way is to use a well-designed speech enhancement approach as the front-end of ASR. However, more complex pipelines, more computations and even higher hardware costs (microphone array) are additionally consumed for this kind of methods. In addition, speech enhancement would result in speech distortions and mismatches to training. In this paper, we propose an adversarial training method to directly boost noise robustness of acoustic model. Specifically, a jointly compositional scheme of generative adversarial net (GAN) and neural network-based acoustic model (AM) is used in the training phase. GAN is used to generate clean feature representations from noisy features by the guidance of a discriminator that tries to distinguish between the true clean signals and generated signals. The joint optimization of generator, discriminator and AM concentrates the strengths of both GAN and AM for speech recognition. Systematic experiments on CHiME-4 show that the proposed method significantly improves the noise robustness of AM and achieves the average relative error rate reduction of 23.38% and 11.54% on the development and test set, respectively.
Bin Liu 0041, Shuai Nie 0001, Dengfeng Ke, Shan Liang 0007
ICASSP4
2018 Mutual-optimization Towards Generative Adversarial Networks For Robust Speech Recognition
abstract
In the context of Automatic Speech Recognition (ASR), improving the noise robustness remains an intractable task. Speech enhancement, combined with Generative Adversarial Networks (GAN), such as SEGAN, has effective performance in denoising raw waveform speech signals. Instead of waveforms, using Mel filterbank spectra in GAN is proposed, which has better performance in the task of ASR. However, these techniques will still miss useful information when GAN is used in them. In this paper, we investigate to protect the useful information in GAN, and propose a novel model, called Discriminator Generator Classifier-GAN (DGC-GAN). While normal GAN combining just two networks will lead the model to denoising rather than recognition, DGC-GAN has another network called classifier, which is an ASR system that will tune GAN to be recognized easier. By adding a classifier into previous GAN to get DGC-GAN, we achieve 29.1% Phone Error Rate (PER) relative improvement in a tiny dataset and 47.4% PER relative improvement in a large dataset.
Ne Luo, Yanyan Xu 0001, Dengfeng Ke, Kaile Su
ICPR4
2018 Fine-Grained Air Quality Prediction using Attention Based Neural Network
abstract
We present a new two-stage fine-grained air Particulate Matter (PM) prediction system using a variety of deep memory networks. Our model is significantly simpler than traditional weather report systems, which rely heavily on the atmospheric reaction equations and pollutant emissions inventories. Pollutant inventories are notoriously difficult to obtain and are often packed with declination for multiple reasons, while reaction equations are impossible to exhaust. These traditional models also tend to perform poorly when affected by strong convective weather. In contrast, our model does not need precise hand collected inventories of pollution sources. It can utilize the potential of Deep-Neural-Networks (DNN) to reveal the relationship among different locations and even find relations between known social events and air quality. Both potentially provide a valid path to air pollution control. The key to our approach is a well-tuned sophisticated attention based network that uses multiple GPUs, allowing us to transform a traditional sparse prediction problem into a sequence-to-sequence learning problem and train it end-to-end. By evaluating the model on over 1,625 instances of data of Northern China collected by our team, we show that it is not only computationally efficient but also accurately feasible compared with other optional models.
Yongzhi Ying, Yanyan Xu 0001, Dengfeng Ke, Kaile Su
IJCNN4
2017 Robust Math Formula Recognition in Degraded Chinese Document Images
abstract
In this paper, we study the problem of math formula recognition (MFR) in degraded Chinese document images. Compared to traditional optical character recognition (OCR), the MFR problem brings new challenges in terms of character segmentation and structural analysis, especially in degraded images. To tackle these issues, we propose an over-segmentation strategy to split and recognize adhesive formula elements based on convolutional neural network (CNN). In addition, we propose a hierarchical framework for formula structure analysis that constructs the formula in a top-down manner to iteratively split the regions into recognizable units. Due to the lack of degraded Chinese document images with math formulas in the community, we also harvest a diverse ground-truth dataset containing 100 images submitted from our system users. Extended experiments demonstrate the effectiveness and robustness of our proposed method in comparison with state-of-the-art methods.
Dongxiang Zhang, Long Guo, Lijiang Chen, Dengfeng Ke
ICDAR7
2017 An Iterative Refinement Framework for Image Document Binarization with Bhattacharyya Similarity Measure
abstract
Background noise and illumination condition are two primary factors degrading the performance of document image binarization. In this paper, we propose an iterative refinement framework to support robust binarization. Initially, an input image is transformed into a Bhattacharyya similarity matrix with Gaussian kernel, which is subsequently converted into a binary image using maximum entropy classifier. Then, we adopt the run-length histogram to estimate the character stroke width, an important indicator to determine the length of filter window. After noise elimination, the output image is used for the next round of refinement and the process terminates when the estimated stroke width is stable. Extensive experiments were conducted on the standard DIBCO datasets as well as a new benchmark harvested from our user query log. Results show that our proposed method outperforms state-of-the-art methods and is more robust to handle low-quality images.
Dongxiang Zhang, Dengfeng Ke, Long Guo, Shengkun Shi, Lijiang Chen
ICDAR5
2017 Symbolic manipulation based on deep neural networks and its application to axiom discovery
abstract
Symbolic reasoning is difficult for neural networks. Especially, reasoning with variables can be a challenging task for them. In this paper, a symbolic reasoning method based on deep neural networks is proposed, and this method is applied to axiom discovery. This method makes use of the concept of “symbolic manipulation”. Specifically, it relies on the learning ability of the deep neural networks and the reasoning ability of a logical system: The logical system generates training examples, which indicate how to manipulate symbols, from given data, and then the deep neural networks try to learn these examples, score them and abstract possible axioms from them. In particular, this method enables the deep neural networks to realise simple reasoning with variables in predicate logic. In experiments, we demonstrate that the deep neural networks are able to learn to copy and generate symbols from a certain form of rules produced by the logical system. Moreover, we find that the more hidden layers usually mean the stronger learning ability of symbolic manipulation: An increasing number of hidden layers usually bring about a higher rule acceptance rate. Also, we find that the more hidden layers can bring about better results on axiom discovery tasks, and we show that the deep neural networks can discover some useful axioms in mathematics.
Dengfeng Ke, Yanyan Xu 0001, Kaile Su
IJCNN2
2017 Deep neural network bottleneck features for bird species verification
abstract
Recently, bottleneck features as effective representations have been successfully used in Speaker Recognition (SR) and Language Recognition (LR), but little work has focused on bottleneck features for Bird Species Verification (BSV). In SR, LR and BSR tasks, using short-time spectra features may be insufficient, so it need some more abstract and discriminative representations as complementation to conventional spectra features. Some SR and LR work shows that bottleneck features can form a low-dimension representation of the original inputs with a powerful descriptive and discriminative capability. Due to the general audio representation principles of speakers, language and birds being similar, we propose a hypothesis: the bottleneck features are also useful for BSV. Therefore, in this paper, we use the bottleneck feature framework based on the standard i-vector framework to deal with crucial problems in conventional methods of BSV, such as the session variability and insufficient features. Moreover, we make no distinction between bird calls and bird songs in the evaluation phase. Experimental results show that the standard i-vector system and the bottleneck feature system gain 3.39% and 0.85% Equal Error Rate (EER) respectively. The bottleneck feature system obtains 75% relative improvement over the standard i-vector system, meaning that the bottleneck features as a complementation to spectra features are significantly useful for BSV. The deep feature system, which is an another state-of-the-art framework based on deep features used in SR, however, only results in 18.64% EER, which is much worse than the other two systems, and a brief explanation is provided in this paper.
Jinming Zhao, Yanyan Xu 0001, Dengfeng Ke, Kaile Su
IJCNN3
2016 SCESS: a WFSA-based automated simplified chinese essay scoring system with incremental latent semantic analysis
abstract
Abstract Writing in language tests is regarded as an important indicator for assessing language skills of test takers. As Chinese language tests become popular, scoring a large number of essays becomes a heavy and expensive task for the organizers of these tests. In the past several years, some efforts have been made to develop automated simplified Chinese essay scoring systems, reducing both costs and evaluation time. In this paper, we introduce a system called SCESS (automated Simplified Chinese Essay Scoring System) based on Weighted Finite State Automata (WFSA) and using Incremental Latent Semantic Analysis (ILSA) to deal with a large number of essays. First, SCESS uses ann-gram language model to construct a WFSA to perform text pre-processing. At this stage, the system integrates a Confusing-Character Table, a Part-Of-Speech Table, beam search and heuristic search to perform automated word segmentation and correction of essays. Experimental results show that this pre-processing procedure is effective, with a Recall Rate of 88.50%, a Detection Precision of 92.31% and a Correction Precision of 88.46%. After text pre-processing, SCESS uses ILSA to perform automated essay scoring. We have carried out experiments to compare the ILSA method with the traditional LSA method on the corpora of essays from the MHK test (the Chinese proficiency test for minorities). Experimental results indicate that ILSA has a significant advantage over LSA, in terms of both running time and memory usage. Furthermore, experimental results also show that SCESS is quite effective with a scoring performance of 89.50%.
Shudong Hao, Yanyan Xu 0001, Dengfeng Ke, Kaile Su, Hengli Peng
Nat. Lang. Eng.3
2015 A Combination of Multi-state Activation Functions, Mean-normalisation and Singular Value Decomposition for learning Deep Neural Networks
abstract
In this paper, we propose Multi-state Activation Functions (MSAFs) for Deep Neural Networks (DNNs). These multi-state functions do extra classification based on the 2-state Logistic function. Discussions on the MSAFs reveal that these activation functions have potentials for altering the parameter distribution of the DNN models, improving model performances and reducing model sizes. Meanwhile, an extension of the XOR problem indicates how neural networks with the multistate functions facilitate classifying patterns. Furthermore, basing on running average mean-normalisation rules, we actualise a combination of mean-normalised optimisation with the MSAFs as well as Singular Value Decomposition (SVD). Experimental results on TIMIT reveal that acoustic models based on DNNs can be improved by applying the MSAFs. The models obtain better phone error rates when the Logistic function is replaced with the multi-state functions. Further experiments on large vocabulary continuous speech recognition tasks reveal that the MSAFs and mean-normalised Stochastic Gradient Descent (MN-SGD) bring better recognition performances for DNNs in comparison with the conventional Logistic function and SGD learning method. Beyond this, the combination of the MSAFs, the SVD method and MN-SGD shrinks the parameter scales of DNNs to 44% approximately, leading to considerable increasing on decoding speed and decreasing on model sizes without any loss of recognition performances.
Dengfeng Ke, Yanyan Xu 0001, Kaile Su
IJCNN2
2015 Multi-task learning deep neural networks for speech feature denoising
abstract
Traditional automatic speech recognition (ASR) systems usually get a sharp performance drop when noise presents in speech. To make a robust ASR, we introduce a new model using the multi-task learning deep neural networks (MTL-DNN) to solve the speech denoising task in feature level. In this model, the networks are initialized by pre-training restricted Boltzmann machines (RBM) and fine-tuned by jointly learning multiple interactive tasks using a shared representation. In multi-task learning, we choose a noisy-clean speech pair fitting task as the primary task and separately explore two constraints as the secondary tasks: phone label and phone cluster. In experiments, the denoised speech is reconstructed by the MTL-DNN using the noisy speech as input and it is respectively evaluated by the DNN-hidden Markov model (HMM) based and the Gaussian Mixture Model (GMM)-HMM based ASR systems. Results show that, using the denoised speech, the word error rate (WER) is respectively reduced by 53.14% and 34.84% compared with baselines. The MTL-DNN model also outperforms the general single-task learning deep neural networks (STL-DNN) model with a performance improvement of 4.93% and 3.88% respectively.
Dengfeng Ke, Hao Zheng 0009, Bo Xu 0002, Yanyan Xu 0001, Kaile Su
INTERSPEECH2
2014 Automated Chinese Essay Scoring from Topic Perspective Using Regularized Latent Semantic Indexing
abstract
Finding out an effective way to score Chinese written essays automatically remains challenging for researchers. Several methods have been proposed and developed but limited in the character and word usage levels. As one of the scoring standards, however, content or topic perspective is also an important and necessary indicator to assess an essay. Therefore, in this paper, we propose a novel perspective -- topic, and a new method integrating topic modeling strategy called Regularized Latent Semantic Indexing to recognize the latent topics and Support Vector Machines to train the scoring model. Experimental results show that automated Chinese essay scoring from topic perspective is effective which can improve the rating agreement to 89%.
Shudong Hao, Yanyan Xu 0001, Hengli Peng, Kaile Su, Dengfeng Ke
ICPR5
2014 Fast Learning of Deep Neural Networks via Singular Value Decomposition
Dengfeng Ke, Yanyan Xu 0001, Kaile Su
PRICAI2
2013 Automated Error Detection and Correction of Chinese Characters in Written Essays Based on Weighted Finite-State Transducer
abstract
Chinese text error detection and correction is widely applicable, but the methods so far are not robust enough for industrial use. In this paper, a new method is proposed based on Tri-gram modeled-Weighted Finite-State Transducer (WFST). By integrating confusing-character table, beam search and A* search, we evaluate the performance on real test essays. Various experiments have been conducted to prove that the proposed method is effective with the recall rate of 85.68%, the detection accuracy of 91.22% and the correction accuracy of 87.30%.
Shudong Hao, Zongtian Gao, Yanyan Xu 0001, Hengli Peng, Kaile Su, Dengfeng Ke
ICDAR7
2012 Automated Essay Scoring Based on Finite State Transducer: towards ASR Transcription of Oral English Speech
Xingyuan Peng, Dengfeng Ke, Bo Xu 0002
ACL (1)2
2010 A new approach for automatic tone error detection in strong accented Mandarin based on dominant set
Taotao Zhu, Dengfeng Ke, Zhenbiao Chen, Bo Xu 0002
INTERSPEECH2
2009 Chinese intonation assessment using SEV features
abstract
Intonation assessment is an important part of Chinese CALL system. Nowadays, most systems use the correlation and RMSE features to assess the quality of the intonation of a given speech. As correlation and RMSE assign unoptimized weights to different degrees of mismatching errors, they may lead to performance degradation. In this paper, we propose a new feature called sorted error vector (SEV) for intonation assessment. The basic idea is to calculate mismatching quantities, sort them with ascending order, and then re-sample them to a K-points vector. This feature has four benefits: first, it is text-length independent; second, weights are let to train by classifiers; third, the relationship between the errors and the final results is not limited to any assumption; fourth, SEV is not sensitive to the performance of different pitch extracting algorithms. Experiments show that no matter in which case, SEV feature performs the best.
Dengfeng Ke, Bo Xu 0002
ICASSP1