EDBT 2026 Demo / reviewers in the wild / expert
Tran Huy Dat
dblp:05/2859 · also Huy Dat Tran
· DBLP profile ↗
50ranked-venue papers
13as first author
9since 2021 · last 2025
0000-0003-0223-6140ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 45 · 13 first-author · 9 since 2021Artificial intelligence and machine learning · 29 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | KAN-AST: Kolmogorov-Arnold Network based Audio Spectrogram Transformer for Audio ClassificationabstractRecently, Kolmogorov-Arnold Networks (KANs) have emerged as a promising alternative to enhance the performance of Multi-Layer Perceptrons (MLPs). Instead of relying on traditional linear weights, KANs utilize spline-parametrized univariate functions, significantly improving interpretability and enabling dynamic learning of activation patterns. This novel approach has garnered considerable interest from the AI community and has seen rapid global adoption. While research on applying KANs to Machine Learning and Computer Vision has expanded, their application in audio processing remains largely unexplored. In this paper, we investigate the untapped potential of KANs in improving models for audio tasks. We propose the first Kolmogorov-Arnold Network-based Audio Spectrogram Transformer (KAN-AST) for audio classification. Our study demonstrates how replacing MLPs with KANs enhances the performance of the AST model, highlighting the potential of KANs in advancing audio classification. Phuong Tuan Dat, Tran Huy Dat |
ASRU | 2 |
| 2025 | Utilizing Kolmogorov-Arnold Network in Self-Supervised Learning for Speaker DiarizationabstractSpeaker diarization is essential in speech processing but remains difficult due to challenges like background noise, far-field audio, overlapping speech, and unknown numbers of speakers. Self-supervised learning (SSL) has recently shown strong performance in neural speaker segmentation (NSS), typically using pre-trained models with Transformer or Conformer decoders connected via a Multi-Layer Perceptron (MLP) adapter. In this work, we enhance the WavLM-Conformer architecture by replacing the MLP with a Kolmogorov-Arnold Network (KAN), which offers better approximation of multivariate functions. Our KAN-based adapter improves the integration of SSL features into downstream tasks. Experiments show that our model achieves consistent gains over existing approaches, setting new state-of-the-art results on challenging diarization benchmarks such as AMI and AISHELL-4. These results demonstrate the promise of KANs in advancing deep learning architectures for speaker diarization. Phuong Tuan Dat, Kah Kuan Teh, Tran Huy Dat |
ASRU | 5 |
| 2025 | XLSR-Kanformer: A KAN-Intergrated model for Synthetic Speech DetectionabstractRecent advancements in speech synthesis technologies have led to increasingly sophisticated spoofing attacks, posing significant challenges for automatic speaker verification systems. While systems based on self-supervised learning (SSL) models, particularly the XLSR-Conformer architecture, have demonstrated remarkable performance in synthetic speech detection, there remains room for architectural improvements. In this paper, we propose a novel approach that replaces the traditional Multi-Layer Perceptron (MLP) in the XLSR-Conformer model with a Kolmogorov-Arnold Network (KAN), a powerful universal approximator based on the Kolmogorov-Arnold representation theorem. Our experimental results on ASVspoof2021 demonstrate that the integration of KAN to XLSR-Conformer model can improve the performance by 60.55% relatively in Equal Error Rate (EER) LA and DF sets, further achieving 0.70% EER on the 21LA set. Besides, the proposed replacement is also robust to various SSL architectures. These findings suggest that incorporating KAN into SSL-based models is a promising direction for advances in synthetic speech detection. Tuan Dat Phuong, Tran Huy Dat |
AVSS | 2 |
| 2025 | Automatic Speech Recognition and Spoken Language Understanding of Maritime Radio Communications: A case study with Singapore dataabstractSpeech communication has been a major part of maritime transportation, particularly in Vessel Traffic Management (VTM). Automatic Speech Recognition (ASR) and Spoken Language Understanding (SLU) of VTM communications provide shipping traffic information in digital forms which improves the productivity and efficiency of VTM operations. The task is, however, challenging due to specific conditions including low-quality audio signals from maritime radio communication channels, non-native accents from ship masters and constrained maritime spoken lingo which is different from natural spoken languages. This paper reports a pilot development of ASR and SLU on Singapore maritime radio communication data. We provide an overview of the dataset and block diagram processing of ASRU and SLU, respectively. This paper presents several contributions designed to improve the ASR and the SLU systems by releasing a dataset for ASR and SLU task in maritime domain. Firstly, we introduce an ASR dataset in the maritime domain which has the purpose of promoting research in the demanding field of maritime. Secondly, we provide experimental results on the efficacy of the different state-of-the-art ASR systems on the maritime dataset. Finally, we evaluate numerous SLU models with maritime dataset for SLU task and also provide some ideas to improve the capabilities of this system for use in the maritime domain in the future. Phuong Dat, Jayakrishnan Melur Madhathil, Tran Huy Dat |
ICASSP | 3 |
| 2025 | Acoustic scattering AI for non-invasive object classifications: A case study on hair assessment
Long-Vu Hoang, Tran Huy Dat |
INTERSPEECH | 3 |
| 2025 | Can we train ASR systems on Code-switch without real code-switch data? Case study for Singapore's languages
Tran Huy Dat |
INTERSPEECH | 2 |
| 2025 | Pushing the Performance of Synthetic Speech Detection with Kolmogorov-Arnold Networks and Self-Supervised Learning Models
Tuan Dat Phuong, Long-Vu Hoang, Tran Huy Dat |
INTERSPEECH | 3 |
| 2024 | LingWav2Vec2: Linguistic-augmented wav2vec 2.0 for Vietnamese Mispronunciation Detection
Tran Huy Dat |
INTERSPEECH | 2 |
| 2021 | Open-Set Audio Classification with Limited Training Resources Based on Augmentation Enhanced Variational Auto-Encoder GAN with Detection-Classification Joint Training
Kah Kuan Teh, Tran Huy Dat |
Interspeech | 2 |
| 2019 | Embedding Physical Augmentation and Wavelet Scattering Transform to Generative Adversarial Networks for Audio Classification with Limited Training ResourcesabstractThis paper addresses audio classification with limited training resources. We first investigate different types of data augmentation including physical modeling, wavelet scattering transform and Generative Adversarial Networks (GAN). We than propose a novel GAN method to embed physical augmentation and wavelet scattering transform in processing. The experimental results on Google Speech Command show significant improvements of the proposed method when training with limited resources. It could lift up classification accuracy from the best baselines of 62.06% and 77.29% on ResNet, to as far as 91.96% and 93.38%, when training with 10% and 25% training data, respectively. Kah Kuan Teh, Tran Huy Dat |
ICASSP | 2 |
| 2019 | The I2R's ASR System for the VOiCES from a Distance Challenge 2019
Tze Yuang Chong, Kye Min Tan, Kah Kuan Teh, Chang Huai You, Hanwu Sun, Tran Huy Dat |
INTERSPEECH | 6 |
| 2019 | The I2R's ASR System for the VOiCES from a Distance Challenge 2019
Tze Yuang Chong, Kye Min Tan, Kah Kuan Teh, Chang Huai You, Hanwu Sun, Tran Huy Dat |
INTERSPEECH | 6 |
| 2019 | Semi-Supervised Audio Classification with Consistency-Based Regularization
Kangkang Lu 0001, Chuan-Sheng Foo, Kah Kuan Teh, Tran Huy Dat, Vijay Chandrasekhar 0001 |
INTERSPEECH | 4 |
| 2019 | The I2R's Submission to VOiCES Distance Speaker Recognition Challenge 2019
Hanwu Sun, Kah Kuan Teh, Ivan Kukanov, Tran Huy Dat |
INTERSPEECH | 4 |
| 2019 | Device Feature Extractor for Replay Spoofing Detection
Chang Huai You, Tran Huy Dat |
INTERSPEECH | 3 |
| 2017 | Data Augmentation, Missing Feature Mask and Kernel Classification for Through-the-Wall Acoustic Surveillance
Tran Huy Dat, Wen Zheng Terence Ng, Yi Ren Leng |
INTERSPEECH | 1 |
| 2017 | An Integrated Solution for Snoring Sound Classification Using Bhattacharyya Distance Based GMM Supervectors with SVM, Feature Selection with Random Forest and Spectrogram with CNN
Tin Lay Nwe, Tran Huy Dat, Wen Zheng Terence Ng, Bin Ma 0001 |
INTERSPEECH | 2 |
| 2016 | A comparative study of multi-channel processing methods for noisy automatic speech recognition in urban environmentsabstractFor the distant speech recognition, the multi-channel processing has been proven to significantly improve the ASR performances compared to the single channel approaches. However, there is very little work has done to provide a comparative evaluation of the approaches, particularly with the modern Deep Neural Network (DNN) recognizers. In this paper, we address the above problem by evaluating the most recently reported mutti-channel methods for the distant speech recognition under urban environments using the 3rd CHiME Challenge database. Particularly, we analyse the effects of each stage of processing of beamforming, adaptive noise cancellation and dereverberation. The back-end processing components are also investigated. We further describe in details our best performing system which combines a harmonic to subharmonic ratio (SHR) voice activity detection, and correlative beamforming with adaptive channel selection in the from-end; semi-supervised DNN adaptation and RNN language model rescoring in the back-end. The system achieved impressive 60% and 55% relative WER reductions on the development set, as well as 65% and 60% of the same on the test set, for real and simulated data sets, respectively. Tran Huy Dat, Jonathan William Dennis, Yi Ren Leng, Wen Zheng Terence Ng |
ICASSP | 1 |
| 2015 | Single and multi-channel approaches for distant speech recognition under noisy reverberant conditions: I2R'S system description for the ASpIRE challengeabstractIn this paper, we introduce the system developed at the Institute for Infocomm Research (I2 R) for the ASpIRE (Automatic Speech recognition In Reverberant Environments) challenge. The main components of the system are a front-end processing system consisting of a distributed beam-forming algorithm, that performs adaptive weighting and channel elimination, a speech dereverberation approach using a maximum-kurtosis criteria, and a robust voice activity detection (VAD) module based on using the sub-harmonic ratio (SHR). The acoustic back-end consists of a multi-conditional Deep Neural Network (DNN) model that uses speaker adapted features combined with a decoding strategy that performs semi-supervised DNN model adaptation using weighted labels generated by the first-pass decoding output. On the single-microphone evaluation, our system achieved a word error rate (WER) of 44.8%. With the incorporation of beamforming on the multi-microphone evaluation, our system achieved an improvement in WER of over 6% to give the best evaluation result of 38.5%. Jonathan William Dennis, Tran Huy Dat |
ASRU | 2 |
| 2015 | Combining robust spike coding with spiking neural networks for sound event classificationabstractThis paper proposes a novel biologically inspired method for sound event classification which combines spike coding with a spiking neural network (SNN). Our spike coding extracts keypoints that represent the local maxima components of the sound spectrogram, and are encoded based on their local time-frequency information; hence both location and spectral information are being extracted. We then design a modified tempotron SNN that, unlike the original tempotron, allows the network to learn the temporal distributions of spike coding input, in an analogous way to the generalized Hough transform. The proposed method simultaneously enhances the sparsity of the sound event spectrogram, producing a representation which is robust against noise, as well as maximises the discriminability of the spike coding input in terms of its temporal information, which is important for sound event classification. Experimental results on a large dataset of 50 environment sound events show the superiority of both the spike coding versus the raw spectrogram and the SNN versus conventional cross-entropy neural networks. Jonathan William Dennis, Tran Huy Dat, Haizhou Li 0001 |
ICASSP | 2 |
| 2015 | Spiking neural networks and the generalised hough transform for speech pattern detection
Jonathan William Dennis, Tran Huy Dat, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2015 | Generalized Hough Transform for Speech Pattern ClassificationabstractWhile typical hybrid neural network architectures for automatic speech recognition (ASR) use a context window of frame-based features, this may not be the best approach to capture the wider temporal context, which contains phonetic and linguistic information that is equally important. In this paper, we introduce a system that integrates both the spectral and geometrical shape information from the acoustic spectrum, inspired by research in the field of machine vision. In particular, we focus on the Generalized Hough Transform (GHT), which is a sophisticated technique that can model the geometrical distribution of speech information over the wider temporal context. To integrate the GHT as part of a hybrid-ASR system, we propose to use a neural network, with features derived from the probabilistic Hough voting step of the GHT, to implement an improved version of the GHT where the output of the network represents the conventional target class posteriors. A major advantage of our approach is that each step of the GHT is highly interpretable, particularly compared to deep neural network (DNN) systems which are commonly treated as powerful black-box classifiers that give little insight into how the output is achieved. Experiments are carried out on two speech pattern classification tasks. The first is the TIMIT phoneme classification, which demonstrates the performance of the approach on a standard ASR task. The second is a spoken word recognition challenge, which highlights the flexibility of the approach to capture phonetic information within a longer temporal context. Jonathan William Dennis, Tran Huy Dat, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Generalized Gaussian Distribution Kullback-Leibler kernel for robust sound event recognitionabstractIn previous works, we have developed a spectrogram image feature extraction framework for robust sound event recognition. The basic idea here is to extract useful information from the 2D time-frequency representation of the sound signal to build up specific feature extractions and classifier under noisy conditions. In this paper, we propose a novel robust spectrogram image method where the key is the observed sparsity of the sound spectrogram image in wavelet representations, which is modeled by the Generalized Gaussian Distributions modeling. Furthermore, the Generalized Gaussian Distribution Kullback-Leibler (GGD-KL) kernel SVM is developed to embed the given probabilistic distance into the quadratic programming machine to optimize the classification The experimental result shows the superiority of the proposed method to the previous works and the state-of-the-art in the field. Tran Huy Dat, Wen Zheng Terence Ng, Jonathan William Dennis, Yi Ren Leng |
ICASSP | 1 |
| 2014 | A discriminatively trained Hough Transform for frame-level phoneme recognitionabstractDespite recent advances in the use of Artificial Neural Network (ANN) architectures for automatic speech recognition (ASR), relatively little attention has been given to using feature inputs beyond MFCCs in such systems. In this paper, we propose an alternative to conventional MFCC or filterbank features, using an approach based on the Generalised Hough Transform (GHT). The GHT is a common approach used in the field of image processing for the task of object detection, where the idea is to learn the spatial distribution of a codebook of feature information relative to the location of the target class. During recognition, a simple weighted summation of the codebook activations is commonly used to detect the presence of the target classes. Here we propose to learn the weighting discriminatively in an ANN, where the aim is to optimise the static phone classification error at the output of the network. As such an ANN is common to hybrid ASR architectures, the output activations from the GHT can be considered as a novel feature for ASR. Experimental results on the TIMIT phoneme recognition task demonstrate the state-of-the-art performance of the approach. Jonathan William Dennis, Tran Huy Dat, Haizhou Li 0001, Chng Eng Siong |
ICASSP | 2 |
| 2014 | Analysis of spectrogram image methods for sound event classification
Jonathan William Dennis, Tran Huy Dat, Chng Eng Siong |
INTERSPEECH | 2 |
| 2013 | Temporal coding of local spectrogram features for robust sound recognitionabstractThere is much evidence to suggest that the human auditory system uses localised time-frequency information for the robust recognition of sounds. Despite this, conventional systems typically rely on features extracted from short windowed frames over time, covering the whole frequency spectrum. Such approaches are not inherently robust to noise, as each frame will contain a mixture of the spectral information from noise and signal. Here, we propose a novel approach based on the temporal coding of Local Spectrogram Features (LSFs), which generate spikes that are used to train a Spiking Neural Network (SNN) with temporal learning. LSFs represent robust location information in the spectrogram surrounding keypoints, which are detected in a signal-driven manner such that the effect of noise on the temporal coding is reduced. Our experiments demonstrate the robust performance of our approach across a variety of noise conditions, such that it is able to outperform the conventional frame-based baseline methods. Jonathan William Dennis, Qiang Yu 0005, Huajin Tang, Tran Huy Dat, Haizhou Li 0001 |
ICASSP | 4 |
| 2013 | Evaluation of the Pet Robot CuDDler Using Godspeed Questionnaire
Yeow Kee Tan, Alvin Hong Yee Wong, Chern Yuen Anthony Wong, Tran Anh Dung, Adrian Hwang Jian Tay, Dilip Kumar Limbu, Tran Huy Dat, Weng Zheng Ng, Benedict Tay Tiong Chee |
ICOST | 7 |
| 2013 | Overlapping sound event recognition using local spectrogram features and the generalised hough transform
Jonathan William Dennis, Tran Huy Dat, Chng Eng Siong |
Pattern Recognit. Lett. | 2 |
| 2013 | Image Feature Representation of the Subband Power Distribution for Robust Sound Event ClassificationabstractThe ability to automatically recognize a wide range of sound events in real-world conditions is an important part of applications such as acoustic surveillance and machine hearing. Our approach takes inspiration from both audio and image processing fields, and is based on transforming the sound into a two-dimensional representation, then extracting an image feature for classification. This provided the motivation for our previous work on the spectrogram image feature (SIF). In this paper, we propose a novel method to improve the sound event classification performance in severe mismatched noise conditions. This is based on the subband power distribution (SPD) image - a novel two-dimensional representation that characterizes the spectral power distribution over time in each frequency subband. Here, the high-powered reliable elements of the spectrogram are transformed to a localized region of the SPD, hence can be easily separated from the noise. We then extract an image feature from the SPD, using the same approach as for the SIF, and develop a novel missing feature classification approach based on a nearest neighbor classifier (kNN). We carry out comprehensive experiments on a database of 50 environmental sound classes over a range of challenging noise conditions. The results demonstrate that the SPD-IF is both discriminative over the broad range of sound classes, and robust in severe non-stationary noise. Jonathan William Dennis, Tran Huy Dat, Chng Eng Siong |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Overlapping Sound Event Recognition using Local Spectrogram Features with the Generalised Hough Transform
Jonathan William Dennis, Tran Huy Dat, Chng Eng Siong |
INTERSPEECH | 2 |
| 2012 | Using Blob Detection in Missing Feature Linear-Frequency Cepstral Coefficients for Robust Sound Event Recognition
Yi Ren Leng, Tran Huy Dat |
INTERSPEECH | 2 |
| 2011 | Probabilistic distance SVM with Hellinger-Exponential Kernel for sound event classificationabstractThis paper presents a novel method for sound event classification based on probabilistic distance SVM. The basic idea is to embed probabilistic distances into classical SVM to classify the sound events. The main point of this method is that the long-term characterization of sound events are better used in the classification compared to conventional method. Furthermore, taking into account the relative short time span of sound events, we develop a probabilistic distance SVM approach based on Hellinger distance from exponential modeling of temporal subband envelopes. An experiment on classifying 10 types of sound events was carried out and showed promising results of the proposed method compared to conventional methods. Tran Huy Dat, Haizhou Li 0001 |
ICASSP | 1 |
| 2011 | Jump Function Kolmogorov for overlapping audio event classificationabstractThis paper presents a novel method for audio event classification in overlapping conditions. The method is based on Jump Function Kolmogorov (JFK), a stochastic representation, which is (a) additive, thus the sum of signal and noise yields the sum of their JFKs; (b) sparse, therefore audio events are separable in this domain. The proposed method is an extension of our previous works for classification under noise-mismatch conditions. Similar to that approach, the robustness of the JFK feature is obtained by limiting them within confidence intervals, which can be learned in advance. However, in order to classify overlapped events, we design the classification system as a set of event detectors and develop a novel approach which maps JFKs to a specific feature for each detector. The experiment shows that the proposed method achieves promising results in very challenging overlapping conditions. Tran Huy Dat, Haizhou Li 0001 |
ICASSP | 1 |
| 2011 | Image Representation of the Subband Power Distribution for Robust Sound Classification
Jonathan William Dennis, Tran Huy Dat, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2011 | Semi-Supervised Tree Support Vector Machine for Online Cough Recognition
Huynh Thai Hoa, An Vu Tran, Tran Huy Dat |
INTERSPEECH | 3 |
| 2011 | Alternative Frequency Scale Cepstral Coefficient for Robust Sound Event Recognition
Yi Ren Leng, Tran Huy Dat, Norihide Kitaoka, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2011 | Spectrogram Image Feature for Sound Event Classification in Mismatched ConditionsabstractIn this letter, we present a novel feature extraction method for sound event classification, based on the visual signature extracted from the sound's time-frequency representation. The motivation stems from the fact that spectrograms form recognisable images, that can be identified by a human reader, with perception enhanced by pseudo-coloration of the image. The signal processing in our method is as follows. 1) The spectrogram is normalised into greyscale with a fixed range. 2) The dynamic range is quantized into regions, each of which is then mapped to form a monochrome image. 3) The monochrome images are partitioned into blocks, and the distribution statistics in each block are extracted to form the feature. The robustness of the proposed method comes from the fact that the noise is normally more diffuse than the signal and therefore the effect of the noise is limited to a particular quantization region, leaving the other regions less changed. The method is tested on a database of 60 sound classes containing a mixture of collision, action and characteristic sounds and shows a significant improvement over other methods in mismatched conditions, without the need for noise reduction. Jonathan William Dennis, Tran Huy Dat, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 2 |
| 2010 | Feature integration for heart sound biometricsabstractThis paper proposes a feature integration framework for heart sound biometric applications. The method selects the best features of different sound classification systems into a unique heart sound biometric system. The framework is developed and tested for both user identification and verification tasks. The experimental results show significant improvements in performance of the proposed system over methods adopting single feature extraction. Among the investigated feature extraction methods, the linear frequency band cepstral coefficients (LFCC) and the GMM super vector are shown to be the best complementary methods. Tran Huy Dat, Yi Ren Leng, Haizhou Li 0001 |
ICASSP | 1 |
| 2010 | Selective gammatone filterbank feature for robust sound event recognition
Yi Ren Leng, Tran Huy Dat, Norihide Kitaoka, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2009 | Sound event classification based on Feature Integration, Recursive Feature Elimination and Structured ClassificationabstractThis paper proposes a novel system for sound event classification based on Feature Integration, Recursive Feature Elimination Support Vector Machine (RFESVM) and Structured Classification. The key points of the proposed method can be summarized as follows: 1) the integration of various feature extraction methods coming from different research communities in one system; 2) the use of feature selection to analyze and select the optimal subset of the integrated features; 3) the adoption of a knowledge-based taxonomic structured classification scheme. Particularly, six groups of features including temporal shape, spectral shape, spectrogram, perceptual cepstral coefficients, harmonic and rhythmic feature sets are investigated in this paper. For the feature selection, the employed RFESVM method enables to select the optimal feature subset taking into account their mutual information. We further develop different feature elimination strategies for RFESVM depending on the requirements of complexity. The RFESVM is combined with a structured classification designed for our task in surveillance and security applications. The proposed method is tested in two realistic environments and the experimental results show good improvements of the classification performance compared to the conventional method. Tran Huy Dat, Haizhou Li 0001 |
ICASSP | 1 |
| 2008 | Jump function komogorov and its application for audio stream segmentation and classificationabstractThis paper proposes a new similarity measurement based on Jump Function Komogorov (JFK) and presents its application for audio content analysis. This is done by means of comparing JFK, a stochastic representation which is (a) additive, so a sum of sources yields a sum of JFK’s, and (b) sparse, so the signal and noise are better separated in the JFK domain. The properties of JFK make it more robust than the probability density function when comparing the signal distributions. In the application, we use the JFK in wavelet domain for the audio stream segmentation and classification. The experimental results show that the proposed method is comparable to the conventional methods under normal condition but significantly outperformed them under miss-match conditions. Tran Huy Dat, Haizhou Li 0001 |
ICASSP | 1 |
| 2008 | Speaker identification in noise mismatch conditions based on jump function Kolmogorov analysis in wavelet domain
Tran Huy Dat, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2008 | Heart sound as a biometric
Koksoon Phua, Tran Huy Dat, Louis Shue |
Pattern Recognit. | 3 |
| 2007 | Feature Selection Based on Fisher Ratio and Mutual Information Analyses for Robust Brain Computer InterfaceabstractThis paper proposes a novel feature selection method based on two-stage analysis of Fisher ratio and mutual information for robust brain computer interface. This method decomposes multichannel brain signals into subbands. The spatial filtering and feature extraction is then processed in each subband. The two-stage analysis of Fisher ratio and mutual information is carried out in the feature domain to reject the noisy feature indexes and select the most informative combination from the remaining. In the approach, we develop two practical solutions, avoiding the difficulties of using high dimensional mutual information in the application, that are the feature indexes clustering using cross mutual information and the latter estimation based on conditional empirical PDF. We test the proposed feature selection method on two BCI data sets and the results are at least comparable to the best results in the literature. The main advantage of proposed method is that the method is free from any time-consuming parameter tweaking and therefore suitable for the BCI system design. Tran Huy Dat, Cuntai Guan |
ICASSP (1) | 1 |
| 2007 | Using Keyword Spotting and Replacement for Speech AnonymizationabstractPrivacy protection issue introduces numerous challenges in the multimedia processing domain. In this paper, we propose an anonymization framework for audio clinical data. The HMM based keyword recognition technique is used to locate the predefined sensitive keywords, which are identified by the users or patients in advance. These keywords will then be substituted by the synthesized nominal words of the similar nature and voice characteristics. The ultimate goal is to protect the privacy information as much as possible, while trying to preserve the speech properties, especially the disease-related symptoms, such as the loudness, the rhythm, the emotion, etc. A preliminary system is presented to demonstrate the usage of the process. Tran Huy Dat, Koksoon Phua, Jit Biswas, Maniyeri Jayachandran |
ICME | 2 |
| 2006 | Multichannel Speech Enhancement Based on Speech Spectral Magnitude Estimation Using Generalized Gamma Prior DistributionabstractWe present multichannel speech enhancement method based on MAP speech spectral magnitude estimation using a generalized gamma model of speech prior distribution, where the model parameters are adapted from actual noisy speech in a frame-by-frame manner. The utilization of a more general prior distribution with its online estimation is shown to be effective for speech spectral estimation. We tested the proposed algorithm in an in-car speech database and obtained significant improvements on the speech recognition performance, particularly under nonstationary noise conditions such as music, air-conditioner and open window. Tran Huy Dat, Kazuya Takeda, Fumitada Itakura |
ICASSP (4) | 1 |
| 2006 | On-line Gaussian mixture modeling in the log-power domain for signal-to-noise ratio estimation and speech enhancement
Tran Huy Dat, Kazuya Takeda, Fumitada Itakura |
Speech Commun. | 1 |
| 2005 | Generalized gamma modeling of speech and its online estimation for speech enhancementabstractGeneralized gamma modeling and its online method of parameter estimation of speech spectral magnitude are proposed for MAP based speech enhancement systems. Generalized gamma modeling is shown to be a natural extension of the Gaussian modeling of speech spectral component distribution, and is therefore, able to fit the prior distribution better than the conventional method. An online parameter estimation method for the gamma distribution, based on a moment matching method, is then proposed. The effectiveness of the proposed methods are confirmed by improvement in both SNR and ASR using the AURORA2 standard database, where about 4 dB improvement in SNR and 20% improvement in relative ASR performance are obtained. Tran Huy Dat, Kazuya Takeda, Fumitada Itakura |
ICASSP (4) | 1 |
| 2005 | SNR and Local Noise Power Estimations Based on Gaussian Mixture Modeling on the Log-Power DomainabstractWe propose a flexible and robust SNR estimation method for the real conditions, when neither clean reference signal nor speech activity is available. This method is based on Gaussian mixture modeling on the log-power domain of the noisy speech and use the estimated subspace distribution parameters to derive the SNR measures. The experimental results show better performance in estimating both the segmental and global SNR compared to the conventional method based on voice activity detection (VAD). The second application presented in this work is local noise power estimation, where the same model is applied to each frequency bin. Furthermore, an empirical MAP solution using second order statistics is applied to estimate the local noise powers in order to implement a Wiener filtering system. The evaluation experiments show the improvements of the proposed speech enhancement method in both segmental SNR and automatic speech recognition (ASR) performance. Kazuya Takeda, Tran Huy Dat, Hiroshi Fujimura, Fumitada Itakura |
ICASSP (1) | 2 |
| 2004 | Speech enhancement based on magnitude estimation using the gamma priorabstractIn this paper, we propose a speech enhancement method based on spectral magnitude estimation. We modify the noise estimation from the minimum statistics method and combine with a maximum a posterior (MAP) decomposition, using the Rice-conditional probability and a non-Gaussian statistic model of the speech. We derive two versions of magnitude decomposition and magnitude-phase decomposition and compare to spectral subtraction and other MAP methods based on the Gaussian statistic (MMSE, LSA). The experiments show the advantage of the proposed method in the improvement of both SNR (up to 12 dB) and recognition accuracy rate (up to 21 % to base line). Weifeng Li 0001, Kazuya Takeda, Fumitada Itakura, Tran Huy Dat |
INTERSPEECH | 4 |