VLDB 2026 Research / reviewers in the wild / expert
Hakan Erdogan
dblp:09/6903
· DBLP profile ↗
72ranked-venue papers
13as first author
13since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 61 · 11 first-author · 13 since 2021Artificial intelligence and machine learning · 33 · 8 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SoundSculpt: Direction and Semantics Driven Ambisonic Target Sound Extraction
Tuochao Chen, D. Shin, Hakan Erdogan, Sinan Hersek |
INTERSPEECH | 3 |
| 2024 | Binaural Angular Separation NetworkabstractWe propose a neural network model that can separate target speech sources from interfering sources at different angular regions using two microphones. The model is trained with simulated room impulse responses (RIRs) using omnidirectional microphones without needing to collect real RIRs. By relying on specific angular regions and multiple room simulations, the model utilizes consistent time difference of arrival (TDOA) cues, or what we call delay contrast, to separate target and interference sources while remaining robust in various reverberation environments. We demonstrate the model is not only generalizable to a commercially available device with a slightly different microphone geometry, but also outperforms our previous work which uses one additional microphone on the same device. The model runs in real-time on-device and is suitable for low-latency streaming applications such as telephony and video conferencing. Yang Yang 0010, George Sung, Shao-Fu Shih, Hakan Erdogan, Chehung Lee, Matthias Grundmann 0002 |
ICASSP | 4 |
| 2024 | Quantifying The Effect Of Simulator-Based Data Augmentation For Speech Recognition On Augmented Reality GlassesabstractAugmented reality (AR) glasses have an immense potential for enhancing conversations by leveraging speech recognition to display real-time transcription or translation, for example, to assist people with hearing impairments or for people conversing in a non-native language. For deployment in real environments, such systems, however, need to be able to separate the speech of interest from noise and other speakers. In this paper, we evaluate the effectiveness of leveraging a room simulator to generate large amounts of simulated training data for such front-end sound separation models, to complement the ideal, but costly, collection of real-world data recorded on the device. Using both recorded and simulated impulse responses (IRs), we demonstrate that the use of simulation data is an effective method for training models that can ultimately enhance speech recognition performance in real-world settings. Furthermore, we show that performance can be further improved by adding microphone directivity in the room simulation, and by fusing synthetic data with a small amount of real IRs. Our results also suggest that existing room simulators would benefit from incorporating the head shadow effect, given its significant impact on multi-microphone recordings on AR glasses. Riku Arakawa, Mathieu Parvaix, Chiong Lai, Hakan Erdogan, Alex Olwal |
ICASSP | 4 |
| 2023 | Guided Speech Enhancement NetworkabstractHigh quality speech capture has been widely studied for both voice communication and human computer interface reasons. To improve the capture performance, we can often find multi-microphone speech enhancement techniques deployed on various devices. Multi-microphone speech enhancement problem is often decomposed into two decoupled steps: a beamformer that provides spatial filtering and a single-channel speech enhancement model that cleans up the beamformer output. In this work, we propose a speech enhancement solution that takes both the raw microphone and beamformer outputs as the input for an ML model. We devise a simple yet effective training scheme that allows the model to learn from the cues of the beamformer by contrasting the two inputs and greatly boost its capability in spatial rejection, while conducting the general tasks of denoising and dereverberation. The proposed solution takes advantage of classical spatial filtering algorithms instead of competing with them. By design, the beamformer module then could be selected separately and does not require a large amount of data to be optimized for a given form factor, and the network model can be considered as a standalone module which is highly transferable independently from the microphone array. We name the ML module in our solution as GSENet, short for Guided Speech Enhancement Network. We demonstrate its effectiveness on real world data collected on multi-microphone devices in terms of the suppression of noise and interfering speech. Yang Yang 0010, Shao-Fu Shih, Hakan Erdogan, Jamie Menjay Lin, Chehung Lee, George Sung, Matthias Grundmann 0002 |
ICASSP | 3 |
| 2023 | TokenSplit: Using Discrete Speech Representations for Direct, Refined, and Transcript-Conditioned Speech Separation and Recognition
Hakan Erdogan, Scott Wisdom, Xuankai Chang, Zalan Borsos, Marco Tagliasacchi, Neil Zeghidour, John R. Hershey |
INTERSPEECH | 1 |
| 2022 | Adapting Speech Separation to Real-World Meetings using Mixture Invariant TrainingabstractThe recently-proposed mixture invariant training (MixIT) is an unsupervised method for training single-channel sound separation models because it does not require ground-truth isolated reference sources. In this paper, we investigate using MixIT to adapt a separation model on real far-field overlapping reverberant and noisy speech data from the AMI Corpus. The models are tested on real AMI recordings containing overlapping speech, and are evaluated subjectively by human listeners. To objectively evaluate our models, we also devise a synthetic AMI test set. For human evaluations on real recordings, we also propose a modification of the standard MUSHRA protocol to handle imperfect reference signals, which we call MUSHIRA. Holding network architectures constant, we find that a fine-tuned semi-supervised model yields the largest SI-SNR improvement, PESQ scores, and human listening ratings across synthetic and real datasets, outperforming unadapted generalist models trained on orders of magnitude more data. Our results show that unsupervised learning through MixIT enables model adaptation on real-world unlabeled spontaneous speech recordings. Aswin Sivaraman, Scott Wisdom, Hakan Erdogan, John R. Hershey |
ICASSP | 3 |
| 2022 | CycleGAN-based Unpaired Speech DereverberationabstractTypically, neural network-based speech dereverberation models are trained on paired data, composed of a dry utterance and its corresponding reverberant utterance.The main limitation of this approach is that such models can only be trained on large amounts of data and a variety of room impulse responses when the data is synthetically reverberated, since acquiring real paired data is costly.In this paper we propose a CycleGAN-based approach that enables dereverberation models to be trained on unpaired data.We quantify the impact of using unpaired data by comparing the proposed unpaired model to a paired model with the same architecture and trained on the paired version of the same dataset.We show that the performance of the unpaired model is comparable to the performance of the paired model on two different datasets, according to objective evaluation metrics.Furthermore, we run two subjective evaluations and show that both models achieve comparable subjective quality on the AMI dataset, which was not seen during training. Hannah Muckenhirn, Aleksandr Safin, Hakan Erdogan, Félix de Chaumont Quitry, Marco Tagliasacchi, Scott Wisdom, John R. Hershey |
INTERSPEECH | 3 |
| 2021 | End-To-End Diarization for Variable Number of Speakers with Local-Global Networks and Discriminative Speaker EmbeddingsabstractWe present an end-to-end deep network model that performs meeting diarization from single-channel audio recordings. End-to-end diarization models have the advantage of handling speaker overlap and enabling straightforward handling of discriminative training, unlike traditional clustering-based diarization methods. The proposed system is designed to handle meetings with unknown numbers of speakers, using variable-number permutation-invariant cross-entropy based loss functions. We introduce several components that appear to help with diarization performance, including a local convolutional network followed by a global self-attention module, multitask transfer learning using a speaker identification component, and a sequential approach where the model is refined with a second stage. These are trained and validated on simulated meeting data based on LibriSpeech and LibriTTS datasets; final evaluations are done using LibriCSS, which consists of simulated meetings recorded using real acoustics via loudspeaker playback. The proposed model performs better than previously proposed end-to-end diarization models on these data. Soumi Maiti, Hakan Erdogan, Kevin W. Wilson, Scott Wisdom, Shinji Watanabe 0001, John R. Hershey |
ICASSP | 2 |
| 2021 | Sound Event Detection and Separation: A Benchmark on Desed Synthetic SoundscapesabstractWe propose a benchmark of state-of-the-art sound event detection systems (SED). We design synthetic evaluation sets to focus on specific sound event detection challenges. We analyze the performance of the submissions to DCASE 2020 Task 4 as a function of time-related modifications (time position of an event and length of clips) and study the impact of non-target sound events and reverberation. We show that temporal localization of sound events remains a challenge for SED systems. We also show that reverberation and non-target sound events severely degrade system performance. In the latter case, sound separation seems like a promising solution. Nicolas Turpault, Romain Serizel, Scott Wisdom, Hakan Erdogan, John R. Hershey, Eduardo Fonseca, Prem Seetharaman, Justin Salamon |
ICASSP | 4 |
| 2021 | What's all the Fuss about Free Universal Sound Separation Data?abstractWe introduce the Free Universal Sound Separation (FUSS) dataset, a new corpus for experiments in separating mixtures of an unknown number of sounds from an open domain of sound types. The dataset consists of 23 hours of single-source audio data drawn from 357 classes, which are used to create mixtures of one to four sources. To simulate reverberation, an acoustic room simulator is used to generate impulse responses of box-shaped rooms with frequency-dependent reflective walls. Additional open-source data augmentation tools are also provided to produce new mixtures with different combinations of sources and room simulations. Finally, we introduce an open-source baseline separation model, based on an improved time-domain convolutional network (TDCN++), that can separate a variable number of sources in a mixture. This model achieves 9.8 dB of scale-invariant signal-to-noise ratio improvement (SI-SNRi) on mixtures with two to four sources, while reconstructing single-source inputs with 35.8 dB absolute SI-SNR. We hope this dataset will lower the barrier to new research and allow for fast iteration and application of novel techniques from other machine learning domains to the sound separation challenge. Scott Wisdom, Hakan Erdogan, Daniel P. W. Ellis, Romain Serizel, Nicolas Turpault, Eduardo Fonseca, Justin Salamon, Prem Seetharaman, John R. Hershey |
ICASSP | 2 |
| 2021 | Continuous Speech Separation Using Speaker Inventory for Long Recording
Cong Han 0001, Yi Luo 0004, Chenda Li, Tianyan Zhou, Keisuke Kinoshita, Shinji Watanabe 0001, Marc Delcroix, Hakan Erdogan, John R. Hershey, Nima Mesgarani, Zhuo Chen 0006 |
Interspeech | 8 |
| 2021 | Integration of Speech Separation, Diarization, and Recognition for Multi-Speaker Meetings: System Description, Comparison, and AnalysisabstractMulti-speaker speech recognition of unsegmented recordings has diverse applications such as meeting transcription and automatic subtitle generation. With technical advances in systems dealing with speech separation, speaker diarization, and automatic speech recognition (ASR) in the last decade, it has become possible to build pipelines that achieve reasonable error rates on this task. In this paper, we propose an end-to-end modular system for the LibriCSS meeting data, which combines independently trained separation, diarization, and recognition components, in that order. We study the effect of different state-of-the-art methods at each stage of the pipeline, and report results using task-specific metrics like SDR and DER, as well as downstream WER. Experiments indicate that the problem of overlapping speech for diarization and ASR can be effectively mitigated with the presence of a well-trained separation module. Our best system achieves a speaker-attributed WER of 12.7%, which is close to that of a non-overlapping ASR. Desh Raj, Pavel Denisov, Zhuo Chen 0006, Hakan Erdogan, Zili Huang, Maokui He, Shinji Watanabe 0001, Jun Du 0002, Takuya Yoshioka, Yi Luo 0004, Naoyuki Kanda, Jinyu Li 0001, Scott Wisdom, John R. Hershey |
SLT | 4 |
| 2021 | Sequential Multi-Frame Neural Beamforming for Speech Separation and EnhancementabstractThis work introduces sequential neural beamforming, which alternates between neural network based spectral separation and beamforming based spatial separation. Our neural networks for separation use an advanced convolutional architecture trained with a novel stabilized signal-to-noise ratio loss function. For beamforming, we explore multiple ways of computing time-varying covariance matrices, including factorizing the spatial covariance into a time-varying amplitude component and a time-invariant spatial component, as well as using block-based techniques. In addition, we introduce a multi-frame beamforming method which improves the results significantly by adding contextual frames to the beamforming formulations. We extensively evaluate and analyze the effects of window size, block size, and multi-frame context size for these methods. Our best method utilizes a sequence of three neural separation and multi-frame time-invariant spatial beamforming stages, and demonstrates an average improvement of 2.75 dB in scale-invariant signal-to-noise ratio and 14.2% absolute reduction in a comparative speech recognition metric across four challenging reverberant speech enhancement and separation tasks. We also use our three-speaker separation model to separate real recordings in the LibriCSS evaluation set into non-overlapping tracks, and achieve a better word error rate as compared to a baseline mask based beamformer. Zhongqiu Wang 0001, Hakan Erdogan, Scott Wisdom, Kevin W. Wilson, Desh Raj, Shinji Watanabe 0001, Zhuo Chen 0006, John R. Hershey |
SLT | 2 |
| 2020 | Performance Study of a Convolutional Time-Domain Audio Separation Network for Real-Time Speech DenoisingabstractTime-domain audio separation networks based on dilated temporal convolutions have recently been shown to perform very well compared to methods that are based on a time-frequency representation in speech separation tasks, even outperforming an oracle binary time-frequency mask of the speakers. This paper investigates the performance of such a time-domain network (Conv-TasNet) for speech denoising in a real-time setting, comparing various parameter settings. Most importantly, different amounts of lookahead are evaluated and compared to the baseline of a fully causal model. We show that a large part of the increase in performance between a causal and non-causal model is achieved with a lookahead of only 20 milliseconds, demonstrating the usefulness of even small lookaheads for many real-time applications. Samuel Sonning, Christian Schüldt, Hakan Erdogan, Scott Wisdom |
ICASSP | 3 |
| 2020 | Unsupervised Sound Separation Using Mixture Invariant TrainingabstractIn recent years, rapid progress has been made on the problem of single-channel sound separation using supervised training of deep neural networks. In such supervised approaches, a model is trained to predict the component sources from synthetic mixtures created by adding up isolated ground-truth sources. Reliance on this synthetic training data is problematic because good performance depends upon the degree of match between the training data and real-world audio, especially in terms of the acoustic conditions and distribution of sources. The acoustic properties can be challenging to accurately simulate, and the distribution of sound types may be hard to replicate. In this paper, we propose a completely unsupervised method, mixture invariant training (MixIT), that requires only single-channel acoustic mixtures. In MixIT, training examples are constructed by mixing together existing mixtures, and the model separates them into a variable number of latent sources, such that the separated sources can be remixed to approximate the original mixtures. We show that MixIT can achieve competitive performance compared to supervised methods on speech separation. Using MixIT in a semi-supervised learning setting enables unsupervised domain adaptation and learning from large amounts of real-world data without ground-truth source waveforms. In particular, we significantly improve reverberant speech separation performance by incorporating reverberant mixtures, train a speech enhancement system from noisy mixtures, and improve universal sound separation by incorporating a large amount of in-the-wild data. Scott Wisdom, Efthymios Tzinis, Hakan Erdogan, Ron J. Weiss, Kevin W. Wilson, John R. Hershey |
NeurIPS | 3 |
| 2019 | SDR - Half-baked or Well Done?abstractIn speech enhancement and source separation, signal-to-noise ratio is a ubiquitous objective measure of denoising/separation quality. A decade ago, the BSS_eval toolkit was developed to give researchers worldwide a way to evaluate the quality of their algorithms in a simple, fair, and hopefully insightful way: it attempted to account for channel variations, and to not only evaluate the total distortion in the estimated signal but also split it in terms of various factors such as remaining interference, newly added artifacts, and channel errors. In recent years, hundreds of papers have been relying on this toolkit to evaluate their proposed methods and compare them to previous works, often arguing that differences on the order of 0.1 dB proved the effectiveness of a method over others. We argue here that the signal-to-distortion ratio (SDR) implemented in the BSS_eval toolkit has generally been improperly used and abused, especially in the case of single-channel separation, resulting in misleading results. We propose to use a slightly modified definition, resulting in a simpler, more robust measure, called scale-invariant SDR (SI-SDR). We present various examples of critical failure of the original SDR that SI-SDR overcomes. Jonathan Le Roux, Scott Wisdom, Hakan Erdogan, John R. Hershey |
ICASSP | 3 |
| 2019 | Single-channel Speech Extraction Using Speaker Inventory and Attention NetworkabstractNeural network-based speech separation has received a surge of interest in recent years. Previously proposed methods either are speaker independent or extract a target speaker's voice by using his or her voice snippet. In applications such as home devices or office meeting transcriptions, a possible speaker list is available, which can be leveraged for speech separation. This paper proposes a novel speech extraction method that utilizes an inventory of voice snippets of possible interfering speakers, or speaker enrollment data, in addition to that of the target speaker. Furthermore, an attention-based network architecture is proposed to form time-varying masks for both the target and other speakers during the separation process. This architecture does not reduce the enrollment audio of each speaker into a single vector, thereby allowing each short time frame of the input mixture signal to be aligned and accurately compared with the enrollment signals. We evaluate the proposed system on a speaker extraction task derived from the Libri corpus and show the effectiveness of the method. Zhuo Chen 0006, Takuya Yoshioka, Hakan Erdogan, Changliang Liu, Dimitrios Dimitriadis, Jasha Droppo, Yifan Gong 0001 |
ICASSP | 4 |
| 2019 | Low-latency Speaker-independent Continuous Speech SeparationabstractSpeaker independent continuous speech separation (SI-CSS) is a task of converting a continuous audio stream, which may contain overlapping voices of unknown speakers, into a fixed number of continuous signals each of which contains no overlapping speech segment. A separated, or cleaned, version of each utterance is generated from one of SI-CSS's output channels nondeterministically without being split up and distributed to multiple channels. A typical application scenario is transcribing multi-party conversations, such as meetings, recorded with microphone arrays. The output signals can be simply sent to a speech recognition engine because they do not include speech overlaps. The previous SI-CSS method uses a neural network trained with permutation invariant training and a data-driven beamformer and thus requires much processing latency. This paper proposes a low-latency SI-CSS method whose performance is comparable to that of the previous method in a microphone array-based meeting transcription task. This is achieved (1) by using a new speech separation network architecture combined with a double buffering scheme and (2) by performing enhancement with a set of fixed beamformers followed by a neural post-filter. Takuya Yoshioka, Zhuo Chen 0006, Changliang Liu, Hakan Erdogan, Dimitrios Dimitriadis |
ICASSP | 5 |
| 2019 | Fixed-length asymmetric binary hashing for fingerprint verification through GMM-SVM based representations
Berkay Topcu, Hakan Erdogan |
Pattern Recognit. | 2 |
| 2018 | Exploring Practical Aspects of Neural Mask-Based Beamforming for Far-Field Speech RecognitionabstractThis work examines acoustic beamformers employing neural networks (NNs) for mask prediction as front -end for automatic speech recognition (ASR) systems for practical scenarios like voice-enabled home devices. To test the versatility of the mask predicting network, the system is evaluated with different recording hardware, different microphone array designs, and different acoustic models of the downstream ASR system. Significant gains in recognition accuracy are obtained in all configurations despite the fact that the NN had been trained on mismatched data. Unlike previous work, the NN is trained on a feature level objective, which gives some performance advantage over a mask related criterion. Furthermore, different approaches for realizing online, or adaptive, NN-based beamforming are explored, where the online algorithms still show significant gains compared to the baseline performance. Christoph Böddeker, Hakan Erdogan, Takuya Yoshioka, Reinhold Häb-Umbach |
ICASSP | 2 |
| 2018 | Multi-Microphone Neural Speech Separation for Far-Field Multi-Talker Speech RecognitionabstractThis paper describes a neural network approach to far-field speech separation using multiple microphones. Our proposed approach is speaker-independent and can learn to implicitly figure out the number of speakers constituting an input speech mixture. This is realized by utilizing the permutation invariant training (PIT) framework, which was recently proposed for single-microphone speech separation. In this paper, PIT is extended to effectively leverage multi-microphone input. It is also combined with beamforming for better recognition accuracy. The effectiveness of the proposed approach is investigated by multi-talker speech recognition experiments that use a large quantity of training data and encompass a range of mixing conditions. Our multi-microphone speech separation system significantly outperforms the single-microphone PIT. Several aspects of the proposed approach are experimentally investigated. Takuya Yoshioka, Hakan Erdogan, Zhuo Chen 0006, Fil Alleva |
ICASSP | 2 |
| 2018 | Investigations on Data Augmentation and Loss Functions for Deep Learning Based Speech-Background SeparationabstractA successful deep learning-based method for separation of a speech signal from an interfering background audio signal is based on neural network prediction of time-frequency masks which multiply noisy signal’s short-time Fourier transform (STFT) to yield the STFT of an enhanced signal. In this paper, we investigate training strategies for mask-prediction based speech-background separation systems. First, we examine the impact of mixing speech and noise files on the fly during training, which enables models to be trained on virtually infinite amount of data. We also investigate the effect of using a novel signal-to-noise ratio related loss function, instead of mean-squared error which is prone to scaling differences among utterances. We evaluate bi-directional long-short term memory (BLSTM) networks as well as a combination of convolutional and BLSTM (CNN+BLSTM) networks for mask prediction and compare performances of real and complex-valued mask prediction. Data-augmented training combined with a novel loss function yields significant improvements in signal to distortion ratio (SDR)and perceptual evaluation of speech quality (PESQ) as compared to the best published result on CHiME-2 medium vocabulary data set when using a CNN+BLSTM network. Hakan Erdogan, Takuya Yoshioka |
INTERSPEECH | 1 |
| 2018 | Recognizing Overlapped Speech in Meetings: A Multichannel Separation Approach Using Neural NetworksabstractThe goal of this work is to develop a meeting transcription system that can recognize speech even when utterances of different speakers are overlapped. While speech overlaps have been regarded as a major obstacle in accurately transcribing meetings, a traditional beamformer with a single output has been exclusively used because previously proposed speech separation techniques have critical constraints for application to real meetings. This paper proposes a new signal processing module, called an unmixing transducer, and describes its implementation using a windowed BLSTM. The unmixing transducer has a fixed number, say J, of output channels, where J may be different from the number of meeting attendees, and transforms an input multi-channel acoustic signal into J time-synchronous audio streams. Each utterance in the meeting is separated and emitted from one of the output channels. Then, each output signal can be simply fed to a speech recognition back-end for segmentation and transcription. Our meeting transcription system using the unmixing transducer outperforms a system based on a state-of-the-art neural mask-based beamformer by 10.8%. Significant improvements are observed in overlapped segments. To the best of our knowledge, this is the first report that applies overlapped speech recognition to unconstrained real meeting audio. Takuya Yoshioka, Hakan Erdogan, Zhuo Chen 0006, Fil Alleva |
INTERSPEECH | 2 |
| 2018 | Multi-Channel Overlapped Speech Recognition with Location Guided Speech Extraction NetworkabstractAlthough advances in close-talk speech recognition have resulted in relatively low error rates, the recognition performance in far-field environments is still limited due to low signal-to-noise ratio, reverberation, and overlapped speech from simultaneous speakers which is especially more difficult. To solve these problems, beamforming and speech separation networks were previously proposed. However, they tend to suffer from leakage of interfering speech or limited generalizability. In this work, we propose a simple yet effective method for multi-channel far-field overlapped speech recognition. In the proposed system, three different features are formed for each target speaker, namely, spectral, spatial, and angle features. Then a neural network is trained using all features with a target of the clean speech of the required speaker. An iterative update procedure is proposed in which the mask-based beamforming and mask estimation are performed alternatively. The proposed system were evaluated with real recorded meetings with different levels of overlapping ratios. The results show that the proposed system achieves more than 24% relative word error rate (WER) reduction than fixed beamforming with oracle selection. Moreover, as overlap ratio rises from 20% to 70+%, only 3.8% WER increase is observed for the proposed system. Zhuo Chen 0006, Takuya Yoshioka, Hakan Erdogan, Jinyu Li 0001, Yifan Gong 0001 |
SLT | 4 |
| 2017 | Deep long short-term memory adaptive beamforming networks for multichannel robust speech recognitionabstractFar-field speech recognition in noisy and reverberant conditions remains a challenging problem despite recent deep learning breakthroughs. This problem is commonly addressed by acquiring a speech signal from multiple microphones and performing beamforming over them. In this paper, we propose to use a recurrent neural network with long short-term memory (LSTM) architecture to adaptively estimate real-time beamforming filter coefficients to cope with non-stationary environmental noise and dynamic nature of source and microphones positions which results in a set of timevarying room impulse responses. The LSTM adaptive beamformer is jointly trained with a deep LSTM acoustic model to predict senone labels. Further, we use hidden units in the deep LSTM acoustic model to assist in predicting the beamforming filter coefficients. The proposed system achieves 7.97% absolute gain over baseline systems with no beamforming on CHiME-3 real evaluation set. Zhong Meng, Shinji Watanabe 0001, John R. Hershey, Hakan Erdogan |
ICASSP | 4 |
| 2017 | Multi-microphone speech recognition integrating beamforming, robust feature extraction, and advanced DNN/RNN backend
Takaaki Hori, Zhuo Chen 0006, Hakan Erdogan, John R. Hershey, Jonathan Le Roux, Vikramjit Mitra, Shinji Watanabe 0001 |
Comput. Speech Lang. | 3 |
| 2017 | Image noise level estimation based on higher-order statistics
Mostafa Mehdipour-Ghazi, Hakan Erdogan |
Multim. Tools Appl. | 2 |
| 2016 | Deep beamforming networks for multi-channel speech recognitionabstractDespite the significant progress in speech recognition enabled by deep neural networks, poor performance persists in some scenarios. In this work, we focus on far-field speech recognition which remains challenging due to high levels of noise and reverberation in the captured speech signals. We propose to represent the stages of acoustic processing including beamforming, feature extraction, and acoustic modeling, as three components of a single unified computational network. The parameters of a frequency-domain beam-former are first estimated by a network based on features derived from the microphone channels. These filter coefficients are then applied to the array signals to form an enhanced signal. Conventional features are then extracted from this signal and passed to a second network that performs acoustic modeling for classification. The parameters of both the beamforming and acoustic modeling networks are trained jointly using back-propagation with a common cross-entropy objective function. In experiments on the AMI meeting corpus, we observed improvements by pre-training each sub-network with a network-specific objective function before joint training of both networks. The proposed method obtained a 3.2% absolute word error rate reduction compared to a conventional pipeline of independent processing stages. Shinji Watanabe 0001, Hakan Erdogan, Liang Lu 0001, John R. Hershey, Michael L. Seltzer, Guoguo Chen, Yu Zhang 0033, Michael I. Mandel, Dong Yu 0001 |
ICASSP | 3 |
| 2016 | Improved MVDR Beamforming Using Single-Channel Mask Prediction Networks
Hakan Erdogan, John R. Hershey, Shinji Watanabe 0001, Michael I. Mandel, Jonathan Le Roux |
INTERSPEECH | 1 |
| 2016 | Improving A⋆ OMP: Theoretical and empirical analyses with a novel dynamic cost model
Nazim Burak Karahanoglu, Hakan Erdogan |
Signal Process. | 2 |
| 2015 | The MERL/SRI system for the 3RD CHiME challenge using beamforming, robust feature extraction, and advanced speech recognitionabstractThis paper introduces the MERL/SRI system designed for the 3rd CHiME speech separation and recognition challenge (CHiME-3). Our proposed system takes advantage of recurrent neural networks (RNNs) throughout the model from the front speech enhancement to the language modeling. Two different types of beamforming are used to combine multi-microphone signals to obtain a single higher quality signal. Beamformed signal is further processed by a single-channel bi-directional long short-term memory (LSTM) enhancement network which is used to extract stacked mel-frequency cepstral coefficients (MFCC) features. In addition, two proposed noise-robust feature extraction methods are used with the beamformed signal. The features are used for decoding in speech recognition systems with deep neural network (DNN) based acoustic models and large-scale RNN language models to achieve high recognition accuracy in noisy environments. Our training methodology includes data augmentation and speaker adaptive training, whereas at test time model combination is used to improve generalization. Results on the CHiME-3 benchmark show that the full cadre of techniques substantially reduced the word error rate (WER). Combining hypotheses from different robust-feature systems ultimately achieved 9.10% WER for the real test data, a 72.4% reduction relative to the baseline of 32.99% WER. Takaaki Hori, Zhuo Chen 0006, Hakan Erdogan, John R. Hershey, Jonathan Le Roux, Vikramjit Mitra, Shinji Watanabe 0001 |
ASRU | 3 |
| 2015 | PLDA-based diarization of telephone conversationsabstractThis paper investigates the application of the probabilistic linear discriminant analysis (PLDA) to speaker diarization of telephone conversations. We introduce using a variational Bayes (VB) approach for inference under a PLDA model for modelling segmental i-vectors in speaker diarization. Deterministic annealing (DA) algorithm is imposed in order to avoid local optimal solutions in VB iterations. We compare our proposed system with a well-known system that applies k-means clustering on principal component analysis (PCA) coefficients of segmental i-vectors. We used summed channel telephone data from the National Institute of Standards and Technology (NIST) 2008 Speaker Recognition Evaluation (SRE) as the test set in order to evaluate the performance of the proposed system. We achieve about 20% relative improvement in Diarization Error Rate (DER) compared to the baseline system. Ahmet Emin Bulut, Hakan Demir, Yusuf Ziya Isik, Hakan Erdogan |
ICASSP | 4 |
| 2015 | Phase-sensitive and recognition-boosted speech separation using deep recurrent neural networksabstractSeparation of speech embedded in non-stationary interference is a challenging problem that has recently seen dramatic improvements using deep network-based methods. Previous work has shown that estimating a masking function to be applied to the noisy spectrum is a viable approach that can be improved by using a signal-approximation based objective function. Better modeling of dynamics through deep recurrent networks has also been shown to improve performance. Here we pursue both of these directions. We develop a phase-sensitive objective function based on the signal-to-noise ratio (SNR) of the reconstructed signal, and show that in experiments it yields uniformly better results in terms of signal-to-distortion ratio (SDR). We also investigate improvements to the modeling of dynamics, using bidirectional recurrent networks, as well as by incorporating speech recognition outputs in the form of alignment vectors concatenated with the spectral input features. Both methods yield further improvements, pointing to tighter integration of recognition with separation as a promising future direction. Hakan Erdogan, John R. Hershey, Shinji Watanabe 0001, Jonathan Le Roux |
ICASSP | 1 |
| 2015 | Speech enhancement and recognition using multi-task learning of long short-term memory recurrent neural networksabstractLong Short-Term Memory (LSTM) recurrent neural network has proven effective in modeling speech and has achieved outstanding performance in both speech enhancement (SE) and automatic speech recognition (ASR). To further improve the performance of noise-robust speech recognition, a combination of speech enhancement and recognition was shown to be promising in earlier work. This paper aims to explore options for consistent integration of SE and ASR using LSTM networks. Since SE and ASR have different objective criteria, it is not clear what kind of integration would finally lead to the best word error rate for noise-robust ASR tasks. In this work, several integration architectures are proposed and tested, including: (1) a pipeline architecture of LSTM-based SE and ASR with sequence training, (2) an alternating estimation architecture, and (3) a multi-task hybrid LSTM network architecture. The proposed models were evaluated on the 2nd CHiME speech separation and recognition challenge task, and show significant improvements relative to prior results. Zhuo Chen 0006, Shinji Watanabe 0001, Hakan Erdogan, John R. Hershey |
INTERSPEECH | 3 |
| 2014 | Counting people by clustering person detector outputsabstractWe present a people counting system that estimates the number of people in a scene by employing a clustering scheme based on Dirichlet Process Mixture Models (DP-MMs) which takes outputs of a person detector system as input. For each frame, we run a person detector on the frame, take its output as a set of detection areas and define a set of features based on spatial, color and temporal information for each detection. Then using these features, we cluster the detections using DPMMs and Gibbs sampling while having no restriction on the number of clusters, thus can estimate an arbitrary number of people or groups of people. We finally define a measure to calculate the actual number of people within each cluster to infer the final estimation of the number of people in the scene. Ibrahim Saygin Topkaya, Hakan Erdogan, Fatih Porikli |
AVSS | 2 |
| 2014 | Deep neural networks for single channel source separationabstractIn this paper, a novel approach for single channel source separation (SCSS) using a deep neural network (DNN) architecture is introduced. Unlike previous studies in which DNN and other classifiers were used for classifying time-frequency bins to obtain hard masks for each source, we use the DNN to classify estimated source spectra to check for their validity during separation. In the training stage, the training data for the source signals are used to train a DNN. In the separation stage, the trained DNN is utilized to aid in estimation of each source in the mixed signal. Single channel source separation problem is formulated as an energy minimization problem where each source spectra estimate is encouraged to fit the trained DNN model and the mixed signal spectrum is encouraged to be written as a weighted sum of the estimated source spectra. The proposed approach works regardless of the energy scale differences between the source signals in the training and separation stages. Nonnegative matrix factorization (NMF) is used to initialize the DNN estimate for each source. The experimental results show that using DNN initialized by NMF for source separation improves the quality of the separated signal compared with using NMF for source separation. Emad M. Grais, Mehmet Umut Sen, Hakan Erdogan |
ICASSP | 3 |
| 2013 | A mixed integer linear programming formulation for the sparse recovery problem in compressed sensingabstractWe propose a new mixed integer linear programming (MILP) formulation of the sparse signal recovery problem in compressed sensing (CS). This formulation is obtained by introduction of an auxiliary binary vector, where ones locate the recovered nonzero indices. Joint optimization for finding this auxiliary vector together with the underlying sparse vector leads to the proposed MILP formulation. By addition of a few appropriate constraints, this problem can be solved by existing MILP solvers. In contrast to other methods, this formulation is not an approximation of the sparse optimization problem, but is its equivalent. Hence, its solution is exactly equal to the optimal solution of the original sparse recovery problem, once it is feasible. We demonstrate this by recovery simulations involving different sparse signal types. The proposed scheme improves recovery over the mainstream CS recovery methods especially when the underlying sparse signals have constant amplitude nonzero elements. Nazim Burak Karahanoglu, Hakan Erdogan, S. Ilker Birbil |
ICASSP | 2 |
| 2013 | Discriminative nonnegative dictionary learning using cross-coherence penalties for single channel source separationabstractIn this work, we introduce a new discriminative training method for nonnegative dictionary learning. The new method can be used in single channel source separation (SCSS) applications. In SCSS, nonnegative matrix factorization (NMF) is used to learn a dictionary (a set of basis vectors) for each source in the magnitude spectrum domain. The trained dictionaries are then used in decomposing the mixed signal to find the estimate for each source. Learning discriminative dictionaries for the source signals can improve the separation performance. To achieve discriminative dictionaries, we try to avoid the bases set of one source dictionary from representing the other source signals. We propose to minimize cross-coherence between the dictionaries of all sources in the mixed signal. We incorporate a simplified cross-coherence penalty using a regularized NMF cost function to simultaneously learn discriminative and reconstructive dictionaries. The new regularized NMF update rules that are used to discriminatively train the dictionaries are introduced in this work. Experimental results show that using discriminative training gives better separation results than using conventional NMF. Copyright © 2013 ISCA. Emad M. Grais, Hakan Erdogan |
INTERSPEECH | 2 |
| 2013 | Spectro-temporal post-enhancement using MMSE estimation in NMF based single-channel source separationabstractWe propose to use minimum mean squared error (MMSE) esti-mates to enhance the signals that are separated by nonnegative matrix factorization (NMF). In single channel source separa-tion (SCSS), NMF is used to train a set of basis vectors for each source from their training spectrograms. Then NMF is used to decompose the mixed signal spectrogram as a weighted linear combination of the trained basis vectors from which estimates of each corresponding source can be obtained. In this work, we deal with the spectrogram of each separated signal as a 2D distorted signal that needs to be restored. A multiplicative dis-tortion model is assumed where the logarithm of the true signal distribution is modeled with a Gaussian mixture model (GMM) and the distortion is modeled as having a log-normal distribu-tion. The parameters of the GMM are learned from training data whereas the distortion parameters are learned online from each separated signal. The initial source estimates are improved and replaced with their MMSE estimates under this new probabilis-tic framework. The experimental results show that using the proposed MMSE estimation technique as a post enhancement after NMF improves the quality of the separated signal. Index Terms: Single channel source separation, nonnegative matrix factorization, Minimum mean square error estimates, and Gaussian mixture models. 1. Emad M. Grais, Hakan Erdogan |
INTERSPEECH | 2 |
| 2013 | Regularized nonnegative matrix factorization using Gaussian mixture priors for supervised single channel source separation
Emad M. Grais, Hakan Erdogan |
Comput. Speech Lang. | 2 |
| 2013 | Linear classifier combination and selection using group sparse regularization and hinge loss
Mehmet Umut Sen, Hakan Erdogan |
Pattern Recognit. Lett. | 2 |
| 2012 | Gaussian Mixture Gain Priors for Regularized Nonnegative Matrix Factorization in Single-Channel Source SeparationabstractWe propose a new method to incorporate statistical priors on the solution of the nonnegative matrix factorization (NMF) for single-channel source separation (SCSS) applications. The Gaussian mixture model (GMM) is used as a log-normalized gain prior model for the NMF solution. The normalization makes the prior models energy independent. In NMF based SCSS, NMF is used to decompose the spectra of the observed mixed signal as a weighted linear combination of a set of trained basis vectors. In this work, the NMF decomposition weights are enforced to consider statistical prior information on the weight combination patterns that the trained basis vectors can jointly receive for each source in the observed mixed signal. The NMF solutions for the weights are encouraged to increase the loglikelihood with the trained gain prior GMMs while reducing the NMF reconstruction error at the same time. Emad M. Grais, Hakan Erdogan |
INTERSPEECH | 2 |
| 2012 | Hidden Markov Models as Priors for Regularized Nonnegative Matrix Factorization in Single-Channel Source SeparationabstractWe propose a new method to incorporate rich statistical priors, modeling temporal gain sequences in the solutions of nonnegative matrix factorization (NMF). The proposed method can be used for single-channel source separation (SCSS) applications. In NMF based SCSS, NMF is used to decompose the spectra of the observed mixed signal as a weighted linear combination of a set of trained basis vectors. In this work, the NMF decomposition weights are enforced to consider statistical and temporal prior information on the weight combination patterns that the trained basis vectors can jointly receive for each source in the observed mixed signal. The Hidden Markov Model (HMM) is used as a log-normalized gains (weights) prior model for the NMF solution. The normalization makes the prior models energy independent. HMM is used as a rich model that characterizes the statistics of sequential data. The NMF solutions for the weights are encouraged to increase the log-likelihood with the trained gain prior HMMs while reducing the NMF reconstruction error at the same time. Emad M. Grais, Hakan Erdogan |
INTERSPEECH | 2 |
| 2012 | SUTAV: A Turkish Audio-Visual Database
Ibrahim Saygin Topkaya, Hakan Erdogan |
LREC | 2 |
| 2012 | Facial feature extraction using a probabilistic approach
Mustafa Berkay Yilmaz, Hakan Erdogan, Mustafa Unel |
Signal Process. Image Commun. | 2 |
| 2011 | Compressed sensing signal recovery via A* Orthogonal Matching PursuitabstractReconstruction of sparse signals acquired in reduced dimensions requires the solution with minimum ℓ0norm. As solving the ℓ0minimization directly is unpractical, a number of algorithms have appeared for finding an indirect solution. A semi-greedy approach, A* Orthogonal Matching Pursuit (A*OMP), is proposed in [1] where the solution is searched on several paths of a search tree. Paths of the tree are evaluated and extended according to some cost function, for which novel dynamic auxiliary cost functions are suggested. This paper describes the A*OMP algorithm and the proposed cost functions briefly. The novel dynamic auxiliary cost functions are shown to provide improved results as compared to a conventional choice. Reconstruction performance is illustrated on both synthetically generated data and real images, which show that the proposed scheme outperforms well-known CS reconstruction methods. Nazim Burak Karahanoglu, Hakan Erdogan |
ICASSP | 2 |
| 2011 | Using multiple visual tandem streams in audio-visual speech recognitionabstractThe method which is called the "tandem approach" in speech recognition has been shown to increase performance by using classifier posterior probabilities as observations in a hidden Markov model. We study the effect of using visual tandem features in audio-visual speech recognition using a novel setup which uses multiple classifiers to obtain multiple visual tandem features. We adopt the approach of multi-stream hidden Markov models where visual tandem features from two different classifiers are considered as additional streams in the model. It is shown in our experiments that using multiple visual tandem features improve the recognition accuracy in various noise conditions. In addition, in order to handle asynchrony between audio and visual observations, we employ coupled hidden Markov models and obtain improved performance as compared to the synchronous model. Ibrahim Saygin Topkaya, Hakan Erdogan |
ICASSP | 2 |
| 2011 | Adaptation of Speaker-Specific Bases in Non-Negative Matrix Factorization for Single Channel Speech-Music SeparationabstractThis paper introduces a speaker adaptation algorithm for nonnegative matrix factorization (NMF) models. The proposed adaptation algorithm is a combination of Bayesian and subspace model adaptation. The adapted model is used to separate speech signal from a background music signal in a single record. Training speech data for multiple speakers is used with NMF to train a set of basis vectors as a general model for speech signals. The probabilistic interpretation of NMF is used to achieve Bayesian adaptation to adjust the general model with respect to the actual properties of the speech signals that is observed in the mixed signal. The Bayesian adapted model is adapted again by a linear transform, which changes the subspace that the Bayesian adapted model spans to better match the speech signal that is in the mixed signal. The experimental results show that combining Bayesian with linear transform adaptation improves the separation results. Copyright © 2011 ISCA. Emad M. Grais, Hakan Erdogan |
INTERSPEECH | 2 |
| 2011 | Single Channel Speech Music Separation Using Nonnegative Matrix Factorization with Sliding Windows and Spectral MasksabstractA single channel speech-music separation algorithm based on nonnegative matrix factorization (NMF) with sliding windows and spectral masks is proposed in this work. We train a set of basis vectors for each source signal using NMF in the magnitude spectral domain. Rather than forming the columns of the matrices to be decomposed by NMF of a single spectral frame, we build them with multiple spectral frames stacked in one column. After observing the mixed signal, NMF is used to decompose its magnitude spectra into a weighted linear combination of the trained basis vectors for both sources. An initial spectrogram estimate for each source is found, and a spectral mask is built using these initial estimates. This mask is used to weight the mixed signal spectrogram to find the contributions of each source signal in the mixed signal. The method is shown to perform better than the conventional NMF approach. Copyright © 2011 ISCA. Emad M. Grais, Hakan Erdogan |
INTERSPEECH | 2 |
| 2011 | Bayesian Models and Algorithms for Protein β-Sheet PredictionabstractPrediction of the 3D structure greatly benefits from the information related to secondary structure, solvent accessibility, and nonlocal contacts that stabilize a protein's structure. We address the problem of \beta-sheet prediction defined as the prediction of \beta--strand pairings, interaction types (parallel or antiparallel), and \beta-residue interactions (or contact maps). We introduce a Bayesian approach for proteins with six or less \beta-strands in which we model the conformational features in a probabilistic framework by combining the amino acid pairing potentials with a priori knowledge of \beta-strand organizations. To select the optimum \beta-sheet architecture, we significantly reduce the search space by heuristics that enforce the amino acid pairs with strong interaction potentials. In addition, we find the optimum pairwise alignment between \beta-strands using dynamic programming in which we allow any number of gaps in an alignment to model \beta-bulges more effectively. For proteins with more than six \beta-strands, we first compute \beta-strand pairings using the BetaPro method. Then, we compute gapped alignments of the paired \beta-strands and choose the interaction types and \beta--residue pairings with maximum alignment scores. We performed a 10-fold cross-validation experiment on the BetaSheet916 set and obtained significant improvements in the prediction accuracy. Zafer Aydin, Yücel Altunbasak, Hakan Erdogan |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2010 | Semi-blind Speech-Music Separation Using Sparsity and Continuity PriorsabstractIn this paper we propose an approach for the problem of single channel source separation of speech and music signals. Our approach is based on representing each source's power spectral density using dictionaries and nonlinearly projecting the mixture signal spectrum onto the combined span of the dictionary entries. We encourage sparsity and continuity of the dictionary coefficients using penalty terms (or log-priors) in an optimization framework. We propose to use a novel coordinate descent technique for optimization, which nicely handles nonnegativity constraints and nonquadratic penalty terms. We use an adaptive Wiener filter, and spectral subtraction to reconstruct both of the sources from the mixture data after corresponding power spectral densities (PSDs) are estimated for each source. Using conventional metrics, we measure the performance of the system on simulated mixtures of single person speech and piano music sources. The results indicate that the proposed method is a promising technique for low speech-to-music ratio conditions and that sparsity and continuity priors help improve the performance of the proposed system. Hakan Erdogan, Emad M. Grais |
ICPR | 1 |
| 2010 | A Unifying Framework for Learning the Linear Combiners for Classifier EnsemblesabstractFor classifier ensembles, an effective combination method is to combine the outputs of each classifier using a linearly weighted combination rule. There are multiple ways to linearly combine classifier outputs and it is beneficial to analyze them as a whole. We present a unifying framework for multiple linear combination types in this paper. This unification enables using the same learning algorithms for different types of linear combiners. We present various ways to train the weights using regularized empirical loss minimization. We propose using the hinge loss for better performance as compared to the conventional least-squares loss. We analyze the effects of using hinge loss for various types of linear weight training by running experiments on three different databases. We show that, in certain problems, linear combiners with fewer parameters may perform as well as the ones with much larger number of parameters even in the presence of regularization. Hakan Erdogan, Mehmet Umut Sen |
ICPR | 1 |
| 2010 | Decision Fusion for Patch-Based Face RecognitionabstractPatch-based face recognition is a recent method which uses the idea of analyzing face images locally, in order to reduce the effects of illumination changes and partial occlusions. Feature fusion and decision fusion are two distinct ways to make use of the extracted local features. Apart from the well-known decision fusion methods, a novel approach for calculating weights for the weighted sum rule is proposed in this paper. Improvements in recognition accuracies are shown and superiority of decision fusion over feature fusion is advocated. In the challenging AR database, we obtain significantly better results using decision fusion as compared to conventional methods and feature fusion methods by using validation accuracy weighting scheme and nearest-neighbor discriminant analysis dimension reduction method. Berkay Topcu, Hakan Erdogan |
ICPR | 2 |
| 2008 | Evolving Implicit Polynomial InterfacesabstractAlthough algebraic or so-called “implicit polynomial ” curves have been studied rather extensively for several decades, to the best of our knowledge, a dynamic formulation of them, similar to active contours, has not been done yet. This paper develops a dynamic formulation for implicit polynomial curves based on level set formalism. In particular, it is shown that utilization of an implicit polynomial distance function in the level set equation yields an ordinary differential equation (ODE) for the temporal behavior of the polynomial coefficients. Using a control theoretic approach, several problems such as curve morphing, dynamic conic fitting without and with constraint, i.e. dynamic ellipse fit, and dynamic curve fitting can be tackled within this new framework. Results are verified by several examples on real images. 1 Erol Ozgur, Mustafa Unel, Hakan Erdogan, Aytül Erçil |
BMVC | 3 |
| 2008 | Using local temporal features of bounding boxes for walking/running classificationabstractFor intelligent surveillance, one of the major tasks to achieve is to recognize activities present in the scene of interest. Human subjects are the most important elements in a surveillance system and it is crucial to classify human actions. In this paper, we tackle the problem of classifying human actions as running or walking in videos. We propose using local temporal features extracted from rectangular boxes that surround the subject of interest in each frame. We test the system using a database of hand-labeled walking and running videos. Our experiments yield a low 2.5% classification error rate using period-based features and the local speed computed using a range of frames around the current frame. Shorter range time-derivative features are not very useful since they are highly variable. Our results show that the system is able to correctly recognize running or walking activities despite differences in appearance and clothing of subjects. Berkay Topcu, Hakan Erdogan |
ICASSP | 2 |
| 2008 | Joint Morphological-Lexical Language Modeling for Processing Morphologically Rich Languages With Application to Dialectal ArabicabstractLanguage modeling for an inflected language such as Arabic poses new challenges for speech recognition and machine translation due to its rich morphology. Rich morphology results in large increases in out-of-vocabulary (OOV) rate and poor language model parameter estimation in the absence of large quantities of data. In this study, we present a joint morphological-lexical language model (JMLLM) that takes advantage of Arabic morphology. JMLLM combines morphological segments with the underlying lexical items and additional available information sources with regards to morphological segments and lexical items in a single joint model. Joint representation and modeling of morphological and lexical items reduces the OOV rate and provides smooth probability estimates while keeping the predictive power of whole words. Speech recognition and machine translation experiments in dialectal-Arabic show improvements over word and morpheme based trigram language models. We also show that as the tightness of integration between different information sources increases, both speech recognition and machine translation performances improve. Ruhi Sarikaya, Mohamed Afify, Yonggang Deng, Hakan Erdogan |
IEEE Trans. Speech Audio Process. | 4 |
| 2007 | Protein Fold Recognition using Residue-Based Alignments of Sequence and Secondary StructureabstractProtein structure prediction aims to determine the three-dimensional structure of proteins form their amino acid sequences. When a protein does not have similarity (homology) to any known fold, threading or fold recognition methods are used to predict structure. Fold recognition methods frequently employ secondary structure, solvent accessibility, and evolutionary information to enhance the accuracy and the quality of the predictions. In this paper, we present a residue based alignment method as an alternative to the state-of-the-art SSEA method, originally introduced by Przytycka et al., and further modified by McGuffin et al. We introduce a residue-based score function, which can incorporate amino acid similarity matrices such as BLOSUM into secondary structure similarity scoring and compute joint alignments. We show that the power of the SSEA method comes from the length normalization instead of the element alignment technique and similar performance can be achieved using residue-based alignments of secondary structures by optimizing gap costs. In simulations with the two benchmark datasets, our method performs slightly better than the SSEA in terms of the fold recognition accuracy. When the secondary structure similarity matrix is combined with the amino acid based BLOSUM30 matrix, the accuracy of our method improves further (4% for the McGuffin set and 10% for the Ding and Dubchak set). The availability of aligning the amino acid and secondary structure sequences in a joint manner offers a better starting point for more elaborate techniques that employ profile-profile alignments and machine learning methods. Zafer Aydin, Hakan Erdogan, Yücel Altunbasak |
ICASSP (1) | 2 |
| 2005 | Regularizing linear discriminant analysis for speech recognitionabstractFeature extraction is an essential first step in speech recognition applications. In addition to static features extracted from each frame of speech data, it is beneficial to use dynamic features (called Δ and ΔΔ coefficients) that use information from neighboring frames. Linear Discriminant Analysis (LDA) followed by a diagonalizing maximum likelihood linear transform (MLLT) applied to spliced static MFCC features yields important performance gains as compared to MFCC+Δ+ΔΔ features in most tasks. However, since LDA is obtained using statistical averages trained on limited data, it is reasonable to regularize LDA transform computation by using prior information and experience. In this paper, we regularize LDA and heteroschedastic LDA transforms using two methods: (1) Using statistical priors for the transform in a MAP formulation (2) Using structural constraints on the transform. As prior, we use a transform that computes static+Δ+ΔΔ coefficients. Our structural constraint is in the form of a block structured LDA transform where each block acts on the same cepstral parameters across frames. The second approach suggests using new coefficients for static, first difference and second difference operators as compared to the standard ones to improve performance. We test the new algorithms on two different tasks, namely TIMIT phone recognition and AURORA2 digit sequence recognition in noise. We obtain consistent improvement in our experiments as compared to MFCC features. In addition, we obtain encouraging results in some AURORA2 tests as compared to LDA+MLLT features. Hakan Erdogan |
INTERSPEECH | 1 |
| 2005 | Using semantic analysis to improve speech recognition performance
Hakan Erdogan, Ruhi Sarikaya, Stanley F. Chen, Michael Picheny |
Comput. Speech Lang. | 1 |
| 2005 | Semantic confidence measurement for spoken dialog systemsabstractThis paper proposes two methods to incorporate semantic information into word and concept level confidence measurement. The first method uses tag and extension probabilities obtained from a statistical classer and parser. The second method uses a maximum entropy based semantic structured language model to assign probabilities to each word. Incorporation of semantic features into a lattice posterior probability based confidence measure provides significant improvements compared to posterior probability when used together in an air travel reservation task. At 5% False Alarm (FA) rate relative improvements of 28% and 61% in Correct Acceptance (CA) rate are achieved for word level and concept level confidence measurements, respectively. Ruhi Sarikaya, Michael Picheny, Hakan Erdogan |
IEEE Trans. Speech Audio Process. | 4 |
| 2004 | Filler model based confidence measures for spoken dialogue systems: a case study for TurkishabstractBecause of the inadequate performance of speech recognition systems, an accurate confidence scoring mechanism should be employed to understand user requests correctly. To determine a confidence score for a hypothesis, certain confidence features are combined. The performance of filler-model based confidence features have been investigated. Five types of filler model networks were defined: triphone-network; phone-network; phone-class network; 5-state catch-all model; 3-state catch-all model. First, all models were evaluated in a Turkish speech recognition task in terms of their ability to tag correctly (recognition-error or correct) recognition hypotheses. The best performance was obtained from the triphone recognition network. Then, the performances of reliable combinations of these models were investigated and it was observed that certain combinations of filler models could significantly improve the accuracy of the confidence annotation. Aydin Akyol, Hakan Erdogan |
ICASSP (1) | 2 |
| 2002 | Turn-Based Language Modeling for spoken dialog systemsabstractIn this paper I we propose a turn-based language modeling (TurnLM) technique for spoken dialog systems. This technique utilizes the time dependent nature of a dialog aimed at accomplishing a task. As opposed to the dialog state based language modeling techniques which depend on the information in the system prompt, TurnLM does not require any information from the dialog manager. As such, TurnLM can be used not only for human-machine dialogs but also human-human dialogs. We report performance improvement compared to the baseline system on the IBM DARPA Communicator spoken dialog system. Experimental results also suggest that TurnLM is a viable alternative to dialog state based language modeling technique. Ruhi Sarikaya, Hakan Erdogan, Michael Picheny |
ICASSP | 3 |
| 2002 | Semantic structured language modelsabstractIn this study, we propose two novel semantic language modeling techniques for spoken dialog systems. These methods are called semantic concept based language modeling and semantic structured language modeling. In the concept based language modeling, we propose to use long span semantic units to model meaning sequences in spoken utterances. In the latter technique, we use statistical semantic parsers to extract information from a sentence. This information is then utilized in a maximum entropy based language model. The language models are trained and evaluated in the air travel reservation domain. We obtain improvement over a sophisticated class based N-gram language model both in terms of recognition accuracy and perplexity. Interpolation of the proposed techniques with the class-based N-gram LM provides additional improvement. 1. Hakan Erdogan, Ruhi Sarikaya, Michael Picheny |
INTERSPEECH | 1 |
| 2002 | Incremental on-line feature space MLLR adaptation for telephony speech recognitionabstractIn this paper, we present a method for incremental on-line adaptation based on feature space Maximum Likelihood Linear Regression (FMLLR) for telephony speech recognition applications. We explain how to incorporate a feature space MLLR transform into a stack decoder and perform on-line adaptation. The issues discussed are as follows: collecting adaptation data on-line and in real time; mapping adaptation data from previous feature space to the present feature space; and smoothing adaptation statistics with initial statistics based on original acoustical model to achieve stability. Testing results on various systems demonstrate that on-line incremental FM-LLR adaptation could be an effective and stable method when the adaptation statistics are mapped and smoothed. 1. Hakan Erdogan, Etienne Marcheret |
INTERSPEECH | 2 |
| 2001 | Rapid adaptation using penalized-likelihood methodsabstractWe introduce rapid adaptation techniques that extend and improve two successful methods previously introduced, cluster weighting (CW) and MAPLR. First, we introduce an adaptation scheme called CWB which extends the cluster weighting adaptation method by including a bias term and a reference speaker model. CWB is shown to improve the adaptation performance as compared to CW. Second, we introduce an extension of cluster weighting that uses penalized-likelihood objective functions to stabilize the estimation and provide soft constraints. Third, we propose a variant of MAPLR adaptation that uses prior speaker information. Previously, prior distributions of transforms in MAPLR were obtained using the same adaptation data, speaker independent HMM means or by some heuristics. We propose to use the prior information of speaker variability to obtain the priors, by using CW or CWB weights. Penalized-likelihood or Bayesian theory serves as a tool to combine transformation based and prior speaker information based adaptation methods resulting in effective rapid adaptation techniques. The techniques are shown to outperform full, block diagonal and diagonal MLLR as well as some other recently proposed methods for rapid adaptation. Hakan Erdogan, Michael Picheny |
ICASSP | 1 |
| 2001 | Innovative approaches for large vocabulary name recognitionabstractAutomatic name dialing is a practical and interesting application of speech recognition on telephony systems. The IBM name recognition system is a large vocabulary, speaker independent system currently in use for reaching IBM employees in the United States. We present some innovative algorithms that improve name recognition accuracy. Unlike transcription tasks, such as the Switchboard task, recognition of names poses a variety of different problems. Several of these problems arise from the fact that foreign names are hard to pronounce for speakers who are not familiar with the names and that there are no standardized methods for pronouncing proper names. Noise robustness is another very important factor as these calls are typically made in noisy environments, such as from a car, cafeteria, airport, etc. and over different kinds of cellular and land-line telephone channels. We have performed a systematic analysis of the speech recognition errors and tackled the issues separately with techniques ranging from weighted speaker clustering, massive adaptation, rapid and unsupervised adaptation methods to pronunciation modeling methods. We find that the decoding accuracy can be improved significantly (28% relative) in this manner. Bhuvana Ramabhadran, C. Julian Chen, Hakan Erdogan, Michael Picheny |
ICASSP | 4 |
| 2001 | Recent advances in speech recognition system for IBM DARPA communicatorabstractIn this paper, we present methods to improve speech recognition performance of the IBM DARPA Communicator system. Our efforts for acoustic modeling include training a domain specific yet broad acoustic model, speaker clustering and speaker adaptation using feature space transforms. For language modeling, we achieved improvements by using compound words, carefully designed LM classes and adjusting the within class probabilities, using NLU state information to enhance the language model and building a language model with embedded grammar objects. Our efforts produced a relative error rate reduction of 34.6 % on the test set that consists of 1173 utterances that IBM received during the NIST evaluation of the DARPA Communicator systems in June 2000. We also tested our decoding on the data from some other sites to further demonstrate the robustness of the system improvements. 1. Hakan Erdogan, Vaibhava Goel, Michael Picheny |
INTERSPEECH | 2 |
| 2000 | Algorithms for joint estimation of attenuation and emission images in PETabstractIn positron emission tomography (PET), positron emission from radiolabeled compounds yields two high energy photons emitted in opposing directions. However, often the photons are not detected due to attenuation within the patient. This attenuation is nonuniform and must be corrected to obtain quantitatively accurate emission images. To measure attenuation effects, one typically acquires a PET transmission scan before or after the injection of radiotracer. In commercially available PET scanners, image reconstruction is performed sequentially in two steps regardless of the reconstruction method: 1. Attenuation correction factor computation (ACF) from transmission scans, 2. Emission image reconstruction using the computed ACFs. This two-step reconstruction scheme does not use all the information in the transmission and emission scans. Postinjection transmission scans contain emission contamination that includes information about emission parameters. Similarly, emission scans contain information about the attenuating medium. To use all the available information, we propose a joint estimation approach that estimates the attenuation map and the emission image simultaneously from these two scans. The penalized-likelihood objective function is nonconvex for this problem. We propose an algorithm based on paraboloidal surrogates that alternates between updating emission and attenuation parameters and is guaranteed to monotonically decrease the objective function. Hakan Erdogan, Jeffrey A. Fessler |
ICASSP | 1 |
| 2000 | Weighted pairwise scatter to improve linear discriminant analysisabstractLinear Discriminant Analysis (LDA) aims to transform an original feature space to a lower dimensional space with as little loss in discrimination as possible. We introduce a novel LDA matrix computation that incorporates confusability information between classes into the transform. Our goal is to improve discrimination in LDA. In conventional LDA, a between class covariance matrix that is based on the scatter of class means around the global mean is used. By rewriting the between class covariance expression in a more revealing way, we unveil that each class pair is considered equally confusable in the conventional LDA. We introduce a weighting factor for each pairwise scatter that enables to integrate the confusability information into the between class covariance matrix. There are many possibilities to choose the weighting factors. We consider few of them that depend on Euclidean and Kullback-Leibler distances between classes when a single Gaussian approximation is used for each class. The method combined with speaker cluster based transformation decreases the error rate by about relative 10 % on a large vocabulary speech recognition task using IBM’s speech recognition engine. 1. Hakan Erdogan |
INTERSPEECH | 3 |
| 2000 | Exact distribution of edge-preserving MAP estimators for linear signal models with Gaussian measurement noiseabstractWe derive the exact statistical distribution of maximum a posteriori (MAP) estimators having edge-preserving nonGaussian priors. Such estimators have been widely advocated for image restoration and reconstruction problems. Previous investigations of these image recovery methods have been primarily empirical; the distribution we derive enables theoretical analysis. The signal model is linear with Gaussian measurement noise. We assume that the energy function of the prior distribution is chosen to ensure a unimodal posterior distribution (for which convexity of the energy function is sufficient), and that the energy function satisfies a uniform Lipschitz regularity condition. The regularity conditions are sufficiently general to encompass popular priors such as the generalized Gaussian Markov random field prior and the Huber prior, even though those priors are not everywhere twice continuously differentiable. Jeffrey A. Fessler, Hakan Erdogan, Wei Biao Wu |
IEEE Trans. Image Process. | 2 |
| 1999 | Fast Monotonic Algorithms for Transmission TomographyabstractWe present a framework for designing fast and monotonic algorithms for transmission tomography penalized-likelihood image reconstruction. The new algorithms are based on paraboloidal surrogate functions for the log likelihood. Due to the form of the log-likelihood function it is possible to find low curvature surrogate functions that guarantee monotonicity. Unlike previous methods, the proposed surrogate functions lead to monotonic algorithms even for the nonconvex log likelihood that arises due to background events, such as scatter and random coincidences. The gradient and the curvature of the likelihood terms are evaluated only once per iteration. Since the problem is simplified at each iteration, the CPU time is less than that of current algorithms which directly minimize the objective, yet the convergence rate is comparable. The simplicity, monotonicity, and speed of the new algorithms are quite attractive. The convergence rates of the algorithms are demonstrated using real and simulated PET transmission scans. Hakan Erdogan, Jeffrey A. Fessler |
IEEE Trans. Medical Imaging | 1 |
| 1998 | Accelerated Monotonic Algorithms for Transmission TomographyabstractWe present a framework for designing fast and monotonic algorithms for transmission tomography penalized likelihood image reconstruction. The new algorithms are based on paraboloidal surrogate functions for the log-likelihood. Due to the form of the log-likelihood function, it is possible to find low curvature surrogate functions that guarantee monotonicity. Unlike previous methods, the proposed surrogate functions lead to monotonic algorithms even for the nonconvex log-likelihood that arises due to background events such as scatter and random coincidences. The gradient and the curvature of the likelihood terms are evaluated only once per iteration. Since the problem is simplified, the CPU time per iteration is less than that of current algorithms which directly minimize the objective, yet the convergence rate is comparable. The simplicity, monotonicity and speed of the new algorithms are quite attractive. The convergence rates of the algorithms are demonstrated using real PET transmission scans. Hakan Erdogan, Jeffrey A. Fessler |
ICIP (2) | 1 |