EDBT 2026 Demo / reviewers in the wild / expert
Heinrich Dinkel
dblp:173/6647
· DBLP profile ↗
32ranked-venue papers
12as first author
20since 2021 · last 2025
0000-0003-4330-8980ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 8 first-author · 18 since 2021Artificial intelligence and machine learning · 21 · 7 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GLCLAP: A Novel Contrastive Learning Pre-trained Model for Contextual Biasing in ASR
Yuxiang Kong, Fan Cui, Liyong Guo, Heinrich Dinkel, Lichun Fan, Jian Luan 0001 |
INTERSPEECH | 4 |
| 2025 | Efficient Speech Enhancement via Embeddings from Pre-trained Generative Audioencoders
Xingwei Sun, Heinrich Dinkel, Yadong Niu, Linzhang Wang, Jian Luan 0001 |
INTERSPEECH | 2 |
| 2025 | X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance
Heinrich Dinkel, Yadong Niu, Anbei Zhao, Jian Luan 0001 |
INTERSPEECH | 2 |
| 2024 | CED: Consistent Ensemble Distillation for Audio TaggingabstractAugmentation and knowledge distillation (KD) are well-established techniques employed in audio classification tasks, aimed at enhancing performance and reducing model sizes on the widely recognized Audioset (AS) benchmark. Although both techniques are effective individually, their combined use, called consistent teaching, hasn’t been explored before. This paper proposes CED, a simple training framework that distils student models from large teacher ensembles with consistent teaching. To achieve this, CED efficiently stores logits as well as the augmentation methods on disk, making it scalable to large-scale datasets. Central to CED’s efficacy is its label-free nature, meaning that only the stored logits are used for the optimization of a student model only requiring 0.3% additional disk space for AS. The study trains various transformer-based models, including a 10M parameter model achieving a 49.0 mean average precision (mAP) on AS. Pretrained models and code are available online. Heinrich Dinkel, Zhiyong Yan |
ICASSP | 1 |
| 2024 | Scaling up masked audio encoder learning for general audio classification
Heinrich Dinkel, Zhiyong Yan, Bin Wang 0004 |
INTERSPEECH | 1 |
| 2024 | Streaming Audio Transformers for Online Audio Tagging
Heinrich Dinkel, Zhiyong Yan |
INTERSPEECH | 1 |
| 2024 | Enhancing Automated Audio Captioning via Large Language Models with Optimized Audio Encoding
Jizhong Liu, Heinrich Dinkel, Zhiyong Yan |
INTERSPEECH | 4 |
| 2024 | Bridging Language Gaps in Audio-Text Retrieval
Zhiyong Yan, Heinrich Dinkel, Jizhong Liu |
INTERSPEECH | 2 |
| 2023 | Unified Keyword Spotting and Audio Tagging on Mobile Devices with TransformersabstractKeyword spotting (KWS) is a core human-machine-interaction front-end task for most modern intelligent assistants. Recently, a unified (UniKW-AT) framework has been proposed that adds additional capabilities in the form of audio tagging (AT) to a KWS model. However, previous work did not consider the real-world deployment of a UniKW-AT model, where factors such as model size and inference speed are more important than performance alone. This work introduces three mobile-device deployable models named Unified Transformers (UiT). Our best model achieves an mAP of 34.09 on Audioset, and an accuracy of 97.76 on the public Google Speech Commands V1 dataset. Further, we benchmark our proposed approaches on four mobile platforms, revealing that the proposed UiT models can achieve a speedup of 2 - 6 times against a competitive MobileNetV2. Heinrich Dinkel, Zhiyong Yan |
ICASSP | 1 |
| 2023 | Av-Sepformer: Cross-Attention Sepformer for Audio-Visual Target Speaker ExtractionabstractVisual information can serve as an effective cue for target speaker extraction (TSE) and is vital to improving extraction performance. In this paper, we propose AV-SepFormer, a SepFormer-based attention dual-scale model that utilizes cross- and self-attention to fuse and model features from audio and visual. AV-SepFormer splits the audio feature into a number of chunks, equivalent to the length of the visual feature. Then self- and cross-attention are employed to model the multi-modal features. Furthermore, we use a novel 2D positional encoding, that introduces the positional information between and within chunks and provides significant gains over the traditional positional encoding. Our model has two key advantages: the time granularity of audio chunked feature is synchronized to the visual feature, which alleviates the harm caused by the inconsistency of audio and video sampling rate; by combining self- and cross-attention, feature fusion and speech extraction processes are unified within an attention paradigm. The experimental results show that AV-SepFormer significantly outperforms other existing methods. Jiuxin Lin, Xinyu Cai, Heinrich Dinkel, Jun Chen 0024, Zhiyong Yan, Zhiyong Wu 0001, Helen M. Meng |
ICASSP | 3 |
| 2023 | Focus on the Sound around You: Monaural Target Speaker Extraction via Distance and Speaker Information
Jiuxin Lin, Heinrich Dinkel, Jun Chen 0024, Zhiyong Wu 0001, Zhiyong Yan |
INTERSPEECH | 3 |
| 2022 | Pseudo Strong Labels for Large Scale Weakly Supervised Audio TaggingabstractLarge-scale audio tagging datasets inevitably contain imperfect labels, such as clip-wise annotated (temporally weak) tags with no exact on- and offsets, due to a high manual labeling cost. This work proposes pseudo strong labels (PSL), a simple label augmentation framework that enhances the supervision quality for large-scale weakly supervised audio tagging. A machine annotator is first trained on a large weakly supervised dataset, which then provides finer supervision for a student model. Using PSL we achieve an mAP of 35.95 balanced train subsets of Audioset using a MobileNetV2 backend, significantly outperforming approaches without PSL. An analysis is provided which reveals that PSL mitigates missing labels. Lastly, we show that models trained with PSL are also superior at generalizing to the Freesound datasets (FSD) than their weakly trained counterparts. Heinrich Dinkel, Zhiyong Yan |
ICASSP | 1 |
| 2022 | Category-Adapted Sound Event Enhancement with Weakly Labeled DataabstractPrevious audio enhancement training usually requires clean signals with additive noises; hence commonly focuses on speech enhancement, where clean speech is easy to access. This paper goes beyond a broader sound event enhancement by using a weakly supervised approach via sound event detection (SED) to approximate the location and presence of a specific sound event. We propose a category-adapted system to enable enhancement on any selected sound category, where we first familiarize the model to all common sound classes and followed by a category-specific fine-tune procedure to enhance the targeted sound class. Evaluation is conducted on ten common sound classes, with a comparison to traditional and weakly supervised enhancement methods. Results indicate an average 2.86 dB SDR increase, with more significant improvement on speech (9.15 dB), music (5.01 dB), and typewriter (3.68 dB) under SNR of 0 dB. All enhancement metrics outperform previous weakly supervised methods and achieve comparable results to the state-of-the-art method that requires clean signals. Guangwei Li, Xuenan Xu, Heinrich Dinkel, Mengyue Wu, Kai Yu 0004 |
ICASSP | 3 |
| 2022 | UniKW-AT: Unified Keyword Spotting and Audio TaggingabstractWithin the audio research community and the industry, keyword spotting (KWS) and audio tagging (AT) are seen as two distinct tasks and research fields. However, from a technical point of view, both of these tasks are identical: they predict a label (keyword in KWS, sound event in AT) for some fixed-sized input audio segment. This work proposes UniKW-AT: An initial approach for jointly training both KWS and AT. UniKW-AT enhances the noise-robustness for KWS, while also being able to predict specific sound events and enabling conditional wake-ups on sound events. Our approach extends the AT pipeline with additional labels describing the presence of a keyword. Experiments are conducted on the Google Speech Commands V1 (GSCV1) and the balanced Audioset (AS) datasets. The proposed MobileNetV2 model achieves an accuracy of 97.53% on the GSCV1 dataset and an mAP of 33.4 on the AS evaluation set. Further, we show that significant noise-robustness gains can be observed on a real-world KWS dataset, greatly outperforming standard KWS approaches. Our study shows that KWS and AT can be merged into a single framework without significant performance degradation. Heinrich Dinkel, Zhiyong Yan |
INTERSPEECH | 1 |
| 2021 | Text-to-Audio Grounding: Building Correspondence Between Captions and Sound EventsabstractAutomated Audio Captioning is a cross-modal task, generating natural language descriptions to summarize the audio clips’ sound events. However, grounding the actual sound events in the given audio based on its corresponding caption has not been investigated. This paper contributes an Audio-Grounding dataset1, which provides the correspondence be-tween sound events and the captions provided in Audiocaps, along with the location (timestamps) of each present sound event. Based on such, we propose the text-to-audio grounding (TAG) task, which interactively considers the relationship be-tween audio processing and language understanding. A base-line approach is provided, resulting in an event-F1 score of 28.3% and a Polyphonic Sound Detection Score (PSDS) score of 14.7%. Xuenan Xu, Heinrich Dinkel, Mengyue Wu, Kai Yu 0004 |
ICASSP | 2 |
| 2021 | Investigating Local and Global Information for Automated Audio Captioning with Transfer LearningabstractAutomated audio captioning (AAC) aims at generating summarizing descriptions for audio clips. Multitudinous concepts are described in an audio caption, ranging from local information such as sound events to global information like acoustic scenery. Currently, the mainstream paradigm for AAC is the end-to-end encoder-decoder architecture, expecting the encoder to learn all levels of concepts embedded in the audio automatically. This paper first proposes a topic model for audio descriptions, comprehensively analyzing the hierarchical audio topics that are commonly covered. We then explore a transfer learning scheme to access local and global information. Two source tasks are identified to respectively represent local and global information, being Audio Tagging (AT) and Acoustic Scene Classification (ASC). Experiments are conducted on the AAC benchmark dataset Clotho and Audiocaps, amounting to a vast increase in all eight metrics with topic transfer learning. Further, it is discovered that local information and abstract representation learning are more crucial to AAC than global information and temporal relationship learning. Xuenan Xu, Heinrich Dinkel, Mengyue Wu, Zeyu Xie, Kai Yu 0004 |
ICASSP | 2 |
| 2021 | A Lightweight Framework for Online Voice Activity Detection in the WildabstractVoice activity detection (VAD) is an essential pre-processing component for speech-related tasks such as automatic speech recognition (ASR).Traditional VAD systems require strong frame-level supervision for training, inhibiting their performance in real-world test scenarios.Previously, the generalpurpose VAD (GPVAD) framework has been proposed to enhance noise robustness significantly.However, GPVAD models are comparatively large and only work for offline evaluation.This work proposes the use of a knowledge distillation framework, where a (large, offline) teacher model provides framelevel supervision to a (light, online) student model.Our experiments verify that our proposed lightweight student models outperform GPVAD on all test sets, including clean, synthetic and real-world scenarios.Our smallest student model only uses 2.2% of the parameters and 15.9% duration cost of our teacher model for inference when evaluated on a Raspberry Pi. Xuenan Xu, Heinrich Dinkel, Mengyue Wu, Kai Yu 0004 |
Interspeech | 2 |
| 2021 | DEPA: Self-Supervised Audio Embedding for Depression DetectionabstractDepression detection research has increased over the last few decades, one major bottleneck of which is the limited data availability and representation learning. Recently, self-supervised learning has seen success in pretraining text embeddings and has been applied broadly on related tasks with sparse data, while pretrained audio embeddings based on self-supervised learning are rarely investigated. This paper proposes DEPA, a self-supervised, pretrained dep ression a udio embedding method for depression detection. An encoder-decoder network is used to extract DEPA on in-domain depressed datasets (DAIC and MDD) and out-domain (Switchboard, Alzheimer's) datasets. With DEPA as the audio embedding extracted at response-level, a significant performance gain is achieved on downstream tasks, evaluated on both sparse datasets like DAIC and large major depression disorder dataset (MDD). This paper not only exhibits itself as a novel embedding extracting method capturing response-level representation for depression detection but more significantly, is an exploration of self-supervised learning in a specific task within audio processing. Pingyue Zhang, Mengyue Wu, Heinrich Dinkel, Kai Yu 0004 |
ACM Multimedia | 3 |
| 2021 | Towards Duration Robust Weakly Supervised Sound Event DetectionabstractSound event detection (SED) is the task of tagging the absence or presence of audio events and their corresponding interval within a given audio clip. While SED can be done using supervised machine learning, where training data is fully labeled with access to per event timestamps and duration, our work focuses on weakly-supervised sound event detection (WSSED), where prior knowledge about an event's duration is unavailable. Recent research within the field focuses on improving segmentand eventlevel localization performance for specific datasets regarding specific evaluation metrics. Specifically, well-performing event-level localization requires fully labeled development subsets to obtain event duration estimates, which significantly benefits localization performance. Moreover, well-performing segment-level localization models output predictions at a coarse-scale (e.g.,1 second), hindering their deployment on datasets containing very short events (<; 1second). This work proposes a duration robust CRNN (CDur) framework, which aims to achieve competitive performance in terms of segmentand event-level localization. This paper proposes a new post-processing strategy named “Triple Threshold” and investigates two data augmentation methods along with a label smoothing method within the scope of WSSED. Evaluation of our model is done on the DCASE2017 and 2018 Task 4 datasets, and URBAN-SED. Our model outperforms other approaches on the DCASE2018 and URBAN-SED datasets without requiring prior duration knowledge. In particular, our model is capable of similar performance to strongly-labeled supervised models on the URBANSED dataset. Lastly, ablation experiments to reveal that without post-processing, our model's localization performance drop is significantly lower compared with other approaches. Heinrich Dinkel, Mengyue Wu, Kai Yu 0004 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2021 | Voice Activity Detection in the Wild: A Data-Driven Approach Using Teacher-Student TrainingabstractVoice activity detection is an essential pre-processing component for speech-related tasks such as automatic speech recognition (ASR). Traditional supervised VAD systems obtain frame-level labels from an ASR pipeline by using, e.g., a Hidden Markov model. These ASR models are commonly trained on clean and fully transcribed data, limiting VAD systems to be trained on clean or synthetically noised datasets. Therefore, a major challenge for supervised VAD systems is their generalization towards noisy, real-world data. This work proposes a data-driven teacher-student approach for VAD, which utilizes vast and unconstrained audio data for training. Unlike previous approaches, only weak labels during teacher training are required, enabling the utilization of any real-world, potentially noisy dataset. Our approach firstly trains a teacher model on a source dataset (Audioset) using clip-level supervision. After training, the teacher provides frame-level guidance to a student model on an unlabeled, target dataset. A multitude of student models trained on mid- to large-sized datasets are investigated (Audioset, Voxceleb, NIST SRE). Our approach is then respectively evaluated on clean, artificially noised, and real-world data. We observe significant performance gains in artificially noised and real-world scenarios. Lastly, we compare our approach against other unsupervised and supervised VAD methods, demonstrating our method's superiority. Heinrich Dinkel, Shuai Wang 0016, Xuenan Xu, Mengyue Wu, Kai Yu 0004 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | Multiple Sound Sources Localization from Coarse to Fine
Rui Qian 0001, Di Hu 0001, Heinrich Dinkel, Mengyue Wu, Ning Xu 0007, Weiyao Lin |
ECCV (20) | 3 |
| 2020 | Duration Robust Weakly Supervised Sound Event DetectionabstractTask 4 of the DCASE2018 challenge demonstrated that substantially more research is needed for a real-world application of sound event detection. Analyzing the challenge results it can be seen that most successful models are biased towards predicting long (e.g., over 5s) clips. This work aims to investigate the performance impact of fixed-sized window median filter post-processing and advocate the use of double thresholding as a more robust and predictable post-processing method. Further, four different temporal subsampling methods within the CRNN framework are proposed: mean-max, α-mean-max, Lp-norm and convolutional. We show that for this task subsampling the temporal resolution by a neural network enhances the F1 score as well as its robustness towards short, sporadic sound events. Our best single model achieves 30.1% F1 on the evaluation set and the best fusion model 32.5%, while being robust to event length variations. Heinrich Dinkel, Kai Yu 0004 |
ICASSP | 1 |
| 2020 | Voice Activity Detection in the Wild via Weakly Supervised Sound Event DetectionabstractTraditional supervised voice activity detection (VAD) methods work well in clean and controlled scenarios, with performance severely degrading in real-world applications.One possible bottleneck is that speech in the wild contains unpredictable noise types, hence frame-level label prediction is difficult, which is required for traditional supervised VAD training.In contrast, we propose a general-purpose VAD (GPVAD) framework, which can be easily trained from noisy data in a weakly supervised fashion, requiring only clip-level labels.We proposed two GP-VAD models, one full (GPV-F), trained on 527 Audioset sound events, and one binary (GPV-B), only distinguishing speech and noise.We evaluate the two GPV models against a CRNN based standard VAD model (VAD-C) on three different evaluation protocols (clean, synthetic noise, real data).Results show that our proposed GPV-F demonstrates competitive performance in clean and synthetic scenarios compared to traditional VAD-C.Further, in real-world evaluation, GPV-F largely outperforms VAD-C in terms of frame-level evaluation metrics as well as segment-level ones.With a much lower requirement for framelabeled data, the naive binary clip-level GPV-B model can still achieve comparable performance to VAD-C in real-world scenarios. Yefei Chen, Heinrich Dinkel, Mengyue Wu, Kai Yu 0004 |
INTERSPEECH | 2 |
| 2020 | Dual-Adversarial Domain Adaptation for Generalized Replay Attack Detection
Heinrich Dinkel, Shuai Wang 0016, Yanmin Qian, Kai Yu 0004 |
INTERSPEECH | 2 |
| 2019 | Audio Caption: Listen and TellabstractIncreasing amount of research has shed light on machine perception of audio events, most of which concerns detection and classification tasks. However, human-like perception of audio scenes involves not only detecting and classifying audio sounds, but also summarizing the relationship between different audio events. Comparable research such as image caption has been conducted, yet the audio field is still quite barren. This paper introduces a manually-annotated dataset for audio caption. The purpose is to automatically generate natural sentences for audio scene description and to bridge the gap between machine perception of audio and image. The whole dataset is labelled in Mandarin and we also include translated English annotations. A baseline encoder-decoder model is provided for both English and Mandarin. Similar BLEU scores are derived for both languages: our model can generate understandable and data-related captions based on the dataset. Mengyue Wu, Heinrich Dinkel, Kai Yu 0004 |
ICASSP | 2 |
| 2019 | Cross-Domain Replay Spoofing Attack Detection Using Domain Adversarial Training
Heinrich Dinkel, Shuai Wang 0016, Yanmin Qian, Kai Yu 0004 |
INTERSPEECH | 2 |
| 2019 | The SJTU Robust Anti-Spoofing System for the ASVspoof 2019 Challenge
Yexin Yang, Heinrich Dinkel, Zhengyang Chen, Shuai Wang 0016, Yanmin Qian, Kai Yu 0004 |
INTERSPEECH | 3 |
| 2018 | Investigating Raw Wave Deep Neural Networks for End-to-End Speaker Spoofing DetectionabstractRecent advances in automatic speaker verification (ASV) lead to an increased interest in securing these systems for real-world applications. Malicious spoofing attempts against ASV systems can lead to serious security breaches. A spoofing attack within the context of ASV is a condition in which a (potentially harmful) person successfully masks as another, to the ASV system already known person by falsifying or manipulating data. While most previous work focuses on enhanced, spoof-aware features, end-to-end models can be a potential alternative. In this paper, we investigate the training of a raw wave front-ends for deep convolutional, long short-term memory (LSTM) and vanilla neural networks, which are analyzed for their suitability toward spoofing detection, regarding the influence of frame size, number of output neurons, and sequence length. A joint convolutional LSTM neural network (CLDNN) is proposed, which outperforms previous attempts on the BTAS2016 dataset (0.82% → 0.19% HTER), placing itself as the current state-of-the-art model for the dataset. We show that end-to-end approaches are appropriate for the important replay detection task and show that the proposed model is capable of distinguishing device-invariant spoofing attempts. Regarding the ASVspoof2015 dataset, the end-to-end solution achieves an equal error rate (EER) of 0.00% for the S1-S9 conditions. We show that the end-to-end approach based on a raw waveform input can outperform common cepstral features, without the use of context-dependent frame extensions. In addition, a cross-database (domain mismatch) scenario is also evaluated, which shows that the proposed CLDNN model trained on the BTAS2016 dataset achieves an EER of 25.7% on the ASVspoof2015 dataset. Heinrich Dinkel, Yanmin Qian, Kai Yu 0004 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2017 | End-to-end spoofing detection with raw waveform CLDNNSabstractAlbeit recent progress in speaker verification generates powerful models, malicious attacks in the form of spoofed speech, are generally not coped with. Recent results in ASVSpoof2015 and BTAS2016 challenges indicate that spoof-aware features are a possible solution to this problem. Most successful methods in both challenges focus on spoof-aware features, rather than focusing on a powerful classifier. In this paper we present a novel raw waveform based deep model for spoofing detection, which jointly acts as a feature extractor and classifier, thus allowing it to directly classify speech signals. This approach can be considered as an end-to-end classifier, which removes the need for any pre- or post-processing on the data, making training and evaluation a streamlined process, consuming less time than other neural-network based approaches. The experiments on the BTAS2016 dataset show that the system performance is significantly improved by the proposed raw waveform convolutional long short term neural network (CLDNN), from the previous best published 1.26% half total error rate (HTER) to the current 0.82% HTER. Moreover it shows that the proposed system also performs well under the unknown (RE-PH2-PH3,RE-LPPH2-PH3) conditions. Heinrich Dinkel, Nanxin Chen, Yanmin Qian, Kai Yu 0004 |
ICASSP | 1 |
| 2017 | Small-footprint convolutional neural network for spoofing detectionabstractAlbeit recent progress in speaker verification engendered powerful models, malicious attacks in the form of spoofed speech, are generally not coped with. In previous attempts, deep neural networks were used to extract high dimensional features which were later classified using an independent classifier. Even though the results of this approach are promising, this architecture's disadvantage is it's complexity of optimizing both, neural network and back-end classifier. In this paper we present a simplified neural network approach to address this problem based on the convolutional neural network architecture. Our model concatenates the output of all abstract convolutional representations within the network into a single high-dimensional vector. By preserving all the information within the network, the networks generalization capabilities are greatly enhanced, resulting in an favorable error rate of 5.4 % on the S10 condition. Scores are frame wise obtained by directly extracting the posteriors from the output neurons and further reduced to an utterance score by the use of variance reduction. We show that by using variance posterior score reduction, large performance gains can be achieved. This model outperforms standard feature extracting neural network approaches, in addition on being more versatile, robust and faster to train. Our best model achieves an error rate of 0.7% on the ASVspoof corpus, utilizing common PLP features. It significantly outperforms conventional feature extraction neural networks, while only having 100k parameters. Heinrich Dinkel, Yanmin Qian, Kai Yu 0004 |
IJCNN | 1 |
| 2017 | Deep Feature Engineering for Noise Robust Spoofing DetectionabstractSpoofing detection for automatic speaker verification (ASV) aims to discriminate between genuine and spoofed speech. This topic has received increased attentions recently due to safety concerns with deploying an ASV system. While the performance of spoofing detection has improved significantly in clean condition in recent studies, the performance degrades dramatically in noisy conditions. To address this issue, in this paper, we propose to extract robust and discriminative deep features by using deep learning techniques for spoofing detection. In particular, we employ deep feedforward, recurrent, and convolutional neural networks to extract discriminative features. We also introduce multicondition training, noise-aware training, and annealed dropout training to make neural networks more robust against noise and to avoid overfitting to specific spoofing attacks and noise types. The proposed neural networks and training techniques are combined into a single framework for spoofing detection. Experimental evaluation is carried out on a noisy version of the standard ASVspoof 2015 corpus, including both additive noisy and reverberant scenarios. Experimental results confirm that the proposed system dramatically decreases averaged equal error rates from 19.1% and 22.6% to 3.2% and 5.1% for seen and unseen noisy conditions, respectively. Yanmin Qian, Nanxin Chen, Heinrich Dinkel, Zhizheng Wu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | Robust deep feature for spoofing detection - the SJTU system for ASVspoof 2015 challengeabstractRecently there have been wide interests in speaker verification for various applications. Although the reported equal error rate (EER) is relatively low, many evidences show that the present speaker verification technologies can be susceptible to malicious spoofing attacks. Inspired by the great success of deep learning in the automatic speech recognition, deep neural network (DNN) based approaches are developed on the spoofing detection for the first time. In this paper, a novel DNN based robust representation is proposed for the spoofing detection to extract the representative spoofing-vector (s-vector). Then the mahalanobis distance and appropriate normalization methods are investigated to get the best system performance. Using the designed deep learning based strategy, our team obtained an impressive result on spoofing detection task, and achieved the 3 rd position in the first spoofing detection challenge evaluation, i.e. ASVspoof 2015 Challenge. Index Terms: Automatic speaker verification, Spoofing attack, Anti-Spoofing, Spoofing detection, Deep learning Nanxin Chen, Yanmin Qian, Heinrich Dinkel, Kai Yu 0004 |
INTERSPEECH | 3 |