VLDB 2026 Research / reviewers in the wild / expert
Kadri Hacioglu
dblp:02/4285
· DBLP profile ↗
31ranked-venue papers
14as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 10 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 7 first-author · 4 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Information extraction and text analysis · 79% Vision and language · 10% Face, body and person analysis · 10% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-AI interaction · 50% Learning and educational technologies · 50% | |
| Computer networks
1 paper |
Physical-layer communications · 100% | |
| Computer graphics and multimedia
1 paper |
Audio and music processing · 100% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis
semantic role labeling |
0.1 | 2 | 2005 | Semantic Role Labeling Using Different Syntactic Views · ACL 2005 Semantic Role Parsing: Adding Semantic Structure to Unstructured Text · ICDM 2003 |
Learning and educational technologies › pedagogical agents
animated pedagogical agents |
0.0 | 1 | 2003 | Perceptive animated interfaces: first steps toward a new paradigm for human-computer interaction · Proc. IEEE 2003 |
Human-AI interaction › conversational agents
embodied conversational agents |
0.0 | 1 | 2003 | Perceptive animated interfaces: first steps toward a new paradigm for human-computer interaction · Proc. IEEE 2003 |
Physical-layer communications
code-division multiple access |
0.0 | 1 | 2000 | Multiuser detection using a genetic algorithm in CDMA communications systems · IEEE Trans. Commun. 2000 |
Physical-layer communications › signal detection › multiuser detection
multistage detection |
0.0 | 1 | 2000 | Multiuser detection using a genetic algorithm in CDMA communications systems · IEEE Trans. Commun. 2000 |
Physical-layer communications › signal detection
multiuser detection |
0.0 | 1 | 2000 | Multiuser detection using a genetic algorithm in CDMA communications systems · IEEE Trans. Commun. 2000 |
Audio and music processing › speech coding
linear predictive coding |
0.0 | 1 | 1998 | Pulse-by-pulse reoptimization of the synthesis filter in pulse-based coders · IEEE Trans. Speech Audio Process. 1998 |
Audio and music processing
speech coding |
0.0 | 1 | 1998 | Pulse-by-pulse reoptimization of the synthesis filter in pulse-based coders · IEEE Trans. Speech Audio Process. 1998 |
Computer vision › Face, body and person analysis
facial expression analysis |
0.0 | 1 | 2003 | Perceptive animated interfaces: first steps toward a new paradigm for human-computer interaction · Proc. IEEE 2003 |
Computer vision › Vision and language
multimodal dialogue |
0.0 | 1 | 2003 | Perceptive animated interfaces: first steps toward a new paradigm for human-computer interaction · Proc. IEEE 2003 |
Methods — techniques the papers use, named apart from their topics
support vector machine · 0.1dialogue system · 0.1computer vision · 0.1computer animation · 0.1parse combination · 0.1feature selection · 0.1feature engineering · 0.0hybrid optimization · 0.0genetic algorithm · 0.0pulse amplitude estimation · 0.0iterative optimization · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task ActivationabstractToken-based multitasking frameworks like TokenVerse require all training utterances to have labels for all tasks, hindering their ability to leverage partially annotated datasets and scale effectively. We propose TokenVerse++, which introduces learnable vectors in the acoustic embedding space of the XLSR-Transducer ASR model for dynamic task activation. This core mechanism enables training with utterances labeled for only a subset of tasks, a key advantage over TokenVerse. We demonstrate this by successfully integrating a dataset with partial labels, specifically for ASR and an additional task, language identification, improving overall performance. TokenVerse++ achieves results on par with or exceeding TokenVerse across multiple tasks, establishing it as a more practical multitask alternative without sacrificing ASR performance. Shashi Kumar, Srikanth R. Madikeri, Esaú Villatoro-Tello, Sergio Burdisso, Pradeep Rangappa, Roberto Andrés Vasco Carofilis, Petr Motlícek, D. S. Karthik Pandia, Shankar Venkatesan, Kadri Hacioglu, Andreas Stolcke |
ASRU | 10 |
| 2025 | Better Semi-supervised Learning for Multi-domain ASR Through Incremental Retraining and Data FilteringabstractFine-tuning pretrained ASR models for specific domains is challenging when labeled data is scarce. But unlabeled audio and labeled data from related domains are often available. We propose an incremental semi-supervised learning pipeline that first integrates a small in-domain labeled set and an auxiliary dataset from a closely related domain, achieving a relative improvement of 4% over no auxiliary data. Filtering based on multi-model consensus or named entity recognition (NER) is then applied to select and iteratively refine pseudo-labels, showing slower performance saturation compared to random selection. Evaluated on the multi-domain Wow call center and Fisher English corpora, it outperforms single-step fine-tuning. Consensus-based filtering outperforms other methods, providing up to 22.3% relative improvement on Wow and 24.8% on Fisher over single-step fine-tuning with random selection. NER is the second-best filter, providing competitive performance at a lower computational cost. Roberto Andrés Vasco Carofilis, Pradeep Rangappa, Srikanth R. Madikeri, Shashi Kumar, Sergio Burdisso, Jeena J. Prakash, Esaú Villatoro-Tello, Petr Motlícek, Bidisha Sharma, Kadri Hacioglu, Shankar Venkatesan, Saurabh Vyas, Andreas Stolcke |
INTERSPEECH | 10 |
| 2025 | Better Pseudo-labeling with Multi-ASR Fusion and Error Correction by SpeechLLMabstractAutomatic speech recognition (ASR) models rely on high-quality transcribed data for effective training. Generating pseudo-labels for large unlabeled audio datasets often relies on complex pipelines that combine multiple ASR outputs through multi-stage processing, leading to error propagation, information loss and disjoint optimization. We propose a unified multi-ASR prompt-driven framework using postprocessing by either textual or speech-based large language models (LLMs), replacing voting or other arbitration logic for reconciling the ensemble outputs. We perform a comparative study of multiple architectures with and without LLMs, showing significant improvements in transcription accuracy compared to traditional methods. Furthermore, we use the pseudo-labels generated by the various approaches to train semi-supervised ASR models for different datasets, again showing improved performance with textual and speechLLM transcriptions compared to baselines. Jeena J. Prakash, Blessingh Kumar, Kadri Hacioglu, Bidisha Sharma, Sindhuja Gopalan, Malolan Chetlur, Shankar Venkatesan, Andreas Stolcke |
INTERSPEECH | 3 |
| 2025 | Efficient Data Selection for Domain Adaptation of ASR Using Pseudo-Labels and Multi-Stage FilteringabstractFine-tuning pretrained ASR models for specific domains is challenging for small organizations with limited labeled data and computational resources. Here, we explore different data selection pipelines and propose a robust approach that improves ASR adaptation by filtering pseudo-labels generated using Whisper (encoder-decoder) and Zipformer (transducer) models. Our approach integrates multiple selection strategies -- including word error rate (WER) prediction, named entity recognition (NER), and character error rate (CER) analysis -- to extract high-quality training segments. We evaluate our method on Whisper and Zipformer using a 7500-hour baseline, comparing it to a CER-based approach relying on hypotheses from three ASR systems. Fine-tuning on 7500 hours of pseudo-labeled call center data achieves 12.3% WER, while our filtering reduces the dataset to 100 hours (1.4%) with similar performance; a similar trend is observed on Fisher English. Pradeep Rangappa, Roberto Andrés Vasco Carofilis, Jeena J. Prakash, Shashi Kumar, Sergio Burdisso, Srikanth R. Madikeri, Esaú Villatoro-Tello, Bidisha Sharma, Petr Motlícek, Kadri Hacioglu, Shankar Venkatesan, Saurabh Vyas, Andreas Stolcke |
INTERSPEECH | 10 |
| 2012 | Improving L1-Specific Phonological Error Diagnosis in Computer Assisted Pronunciation Training
Theban Stanley, Kadri Hacioglu |
INTERSPEECH | 2 |
| 2005 | Semantic Role Labeling Using Different Syntactic ViewsabstractSemantic role labeling is the process of annotating the predicate-argument structure in text with semantic labels. In this paper we present a state-of-the-art baseline semantic role labeling system based on Support Vector Machine classifiers. We show improvements on this system by: i) adding new features including features extracted from dependency parses, ii) performing feature selection and calibration and iii) combining parses obtained from semantic parsers trained using different syntactic views. Error analysis of the baseline system showed that approximately half of the argument identification errors resulted from parse errors in which there was no syntactic constituent that aligned with the correct argument. In order to address this problem, we combined semantic parses from a Minipar syntactic parse and from a chunked syntactic representation with our original baseline system which was based on Charniak parses. All of the reported techniques resulted in performance improvements. Sameer Pradhan, Wayne H. Ward, Kadri Hacioglu, James H. Martin, Daniel Jurafsky |
ACL | 3 |
| 2005 | Automatic Time Expression Labeling for English and Chinese Text
Kadri Hacioglu, Benjamin Douglas |
CICLing | 1 |
| 2005 | Semantic Role Chunking Combining Complementary Syntactic Views
Sameer Pradhan, Kadri Hacioglu, Wayne H. Ward, James H. Martin, Daniel Jurafsky |
CoNLL | 2 |
| 2005 | Support Vector Learning for Semantic Argument Classification
Sameer Pradhan, Kadri Hacioglu, Valerie Krugler, Wayne H. Ward, James H. Martin, Daniel Jurafsky |
Mach. Learn. | 2 |
| 2004 | Semantic Role Labeling Using Dependency Trees
Kadri Hacioglu |
COLING | 1 |
| 2004 | Semantic Role Labeling by Tagging Syntactic Chunks
Kadri Hacioglu, Sameer Pradhan, Wayne H. Ward, James H. Martin, Daniel Jurafsky |
CoNLL | 1 |
| 2004 | Parsing speech into articulatory eventsabstractIn this paper, the states in the speech production process are defined by a number of categorical articulatory features. We describe a detector that outputs a stream (sequence of classes) for each articulatory feature given the Mel frequency cepstral coefficient (MFCC) representation of the input speech. The detector consists of a bank of recurrent neural network (RNN) classifiers, a variable depth lattice generator and Viterbi decoder. A bank of classifiers has been previously used for articulatory feature detection by many researchers. We extend their work first by creating variable depth lattices for each feature and then by combining them into product lattices for rescoring using the Viterbi algorithm. During the rescoring we incorporate language and duration constraints along with the posterior probabilities of classes provided by the RNN classifiers. We present our results for the place and manner features using TIMIT data, and compare the results to a baseline system. We report performance improvements both at the frame and segment levels. Kadri Hacioglu, Bryan L. Pellom, Wayne H. Ward |
ICASSP (1) | 1 |
| 2004 | Shallow Semantic Parsing using Support Vector Machines
Sameer Pradhan, Wayne H. Ward, Kadri Hacioglu, James H. Martin, Daniel Jurafsky |
HLT-NAACL | 3 |
| 2004 | The adaptive path selective decorrelating detector: performance analysis with channel estimation errors
Ali Hakan Ulusoy, Ahmet Rizaner, Kadri Hacioglu, Hasan Amca |
Signal Process. | 3 |
| 2003 | A distributed architecture for robust automatic speech recognitionabstractIn this paper, we attempt to decompose a state-of-the-art speech recognition system into its components and define an infrastructure that allows a flexible, efficient and effective interaction among the components. Motivated by the success of DARPA Communicator program, we select the open source Galaxy architecture as our development test bed. It consists of a hub that allows communication among servers connected to it by message passing and supports the plug-and-play paradigm. In addition to message passing it supports high bandwidth data (binary or audio) transfer between servers via a brokering scheme. For several reasons, we believe that it is the right time to start developing a distributed framework for speech recognition along with data and protocol standards supporting interoperability. We present our work towards that goal using the Colorado University (CU) Sonic recognizer. We divide Sonic into a number of components and structure it around the Hub. We describe the system in some detail and report on its present status with some possibilities for future development. Kadri Hacioglu, Bryan L. Pellom |
ICASSP (1) | 1 |
| 2003 | Recent improvements in the CU Sonic ASR system for noisy speech: the SPINE taskabstractWe report on recent improvements in the University of Colorado system for the DARPA/NRL Speech in Noisy Environments (SPINE) task. In particular, we describe our efforts on improving acoustic and language modeling for the task and investigate methods for unsupervised speaker and environment adaptation from limited data. We show that the MAPLR adaptation method outperforms single and multiple regression class MLLR on the SPINE task. Our current SPINE system uses the Sonic speech recognition engine that was developed at the University of Colorado. This system is shown to have a word error rate of 31.5% on the SPINE-2 evaluation data. These improvements amount to a 16% reduction in relative word error rate compared to our previous SPINE-2 system fielded in the November 2001 DARPA/NRL evaluation. Bryan L. Pellom, Kadri Hacioglu |
ICASSP (1) | 2 |
| 2003 | Semantic Role Parsing: Adding Semantic Structure to Unstructured TextabstractThere is an ever-growing need to add structure in the form of semantic markup to the huge amounts of unstructured text data now available. We present the technique of shallow semantic parsing, the process of assigning a simple WHO did WHAT to WHOM, etc., structure to sentences in text, as a useful tool in achieving this goal. We formulate the semantic parsing problem as a classification problem using support vector machines. Using a hand-labeled training set and a set of features drawn from earlier work together with some feature enhancements, we demonstrate a system that performs better than all other published results on shallow semantic parsing. Sameer Pradhan, Kadri Hacioglu, Wayne H. Ward, James H. Martin, Daniel Jurafsky |
ICDM | 2 |
| 2003 | On lexicon creation for turkish LVCSRabstractIn this paper, we address the lexicon design problem in Turkish large vocabulary speech recognition. Although we focus only on Turkish, the methods described here are general enough that they can be considered for other agglutinative languages like Finnish, Korean etc. In an agglutinative language, several words can be created from a single root word using a rich collection of morphological rules. So, a virtually infinite size lexicon is required to cover the language if words are used as the basic units. The standard approach to this problem is to discover a number of primitive units so that a large set of words can be created by compounding those units. Two broad classes of methods are available for splitting words into their sub-units; morphology-based and data-driven methods. Although the word splitting significantly reduces the out of vocabulary rate, it shrinks the context and increases acoustic confusibility. We have used two methods to address the latter. In one method, we use word counts to avoid splitting of high frequency lexical units, and in the other method, we recompound splits according to a probabilistic measure. We present experimental results that show the methods are very effective to lower the word error rate at the expense of lexicon size. 1. Kadri Hacioglu, Bryan L. Pellom, Tolga Çiloglu, Özlem Öztürk, Mikko Kurimo, Mathias Creutz |
INTERSPEECH | 1 |
| 2003 | Target Word Detection and Semantic Role Chunking using Support Vector Machines
Kadri Hacioglu, Wayne H. Ward |
HLT-NAACL | 1 |
| 2003 | Question Classification with Support Vector Machines and Error Correcting Codes
Kadri Hacioglu, Wayne H. Ward |
HLT-NAACL | 1 |
| 2003 | Perceptive animated interfaces: first steps toward a new paradigm for human-computer interactionabstractThis paper presents a vision of the near future in which computer interaction is characterized by natural face-to-face conversations with lifelike characters that speak, emote, and gesture. These animated agents will converse with people much like people converse effectively with assistants in a variety of focused applications. Despite the research advances required to realize this vision, and the lack of strong experimental evidence that animated agents improve human-computer interaction, we argue that initial prototypes of perceptive animated interfaces can be developed today, and that the resulting systems will provide more effective and engaging communication experiences than existing systems. In support of this hypothesis, we first describe initial experiments using an animated character to teach speech and language skills to children with hearing problems, and classroom subjects and social skills to children with autistic spectrum disorder. We then show how existing dialogue system architectures can be transformed into perceptive animated interfaces by integrating computer vision and animation capabilities. We conclude by describing the Colorado Literacy Tutor, a computer-based literacy program that provides an ideal testbed for research and development of perceptive animated interfaces, and consider next steps required to realize the vision. Ronald A. Cole, Sarel van Vuuren, Bryan L. Pellom, Kadri Hacioglu, Jiyong Ma, Javier Movellan, Scott Schwartz, David Wade-Stein, Wayne H. Ward |
Proc. IEEE | 4 |
| 2002 | A concept graph based confidence measureabstractIn this paper, the confidence measure of a hypothesized word is derived from its posterior probability. In contrast to common approaches, in which N-best lists or word graphs/lattices are used, the posterior probabilities are derived from a concept graph. The concept graph is obtained from a word graph through a partial parsing process using semantic grammars. This approach allows us to use relatively complex and better language models along with acoustic models to compute word posterior probabilities. The language model used is comprised of stochastic context free grammars (one for each concept) and an n-gram concept language model. We show that the posterior probabilities computed on concept graphs outperform those computed on word graphs when used as confidence measures. Results are presented within the context of Colorado University (CU) Communicator System; a telephone-based dialog system for making travel plans by accessing information about flights, hotels and car rentals. Kadri Hacioglu, Wayne H. Ward |
ICASSP | 1 |
| 2002 | A figure of merit for the analysis of spoken dialog systemsabstractIn this paper, a single metric, which we will call the figure of merit , for the quantitative analysis and comparison of spoken dialog systems is introduced. This figure of merit is the product of the weighted dialog accuracy (expressed as the rate of success) and the weighted dialog efficiency (expressed as the average number of concepts per turn). Actually, it is highly desirable to have a quick and accurate dialog. However, these two requirements are conflicting. That is, an improvement in efficiency is accomplished at the expense of accuracy or vice versa. This makes difficult to compare two different spoken dialog systems or tune a particular system. We believe that this figure of merit would avoid those difficulties. To illustrate its use, we consider spoken dialog systems with different dialog strategies and compare them by performing quantitative analysis based on the finate state models of information items using the proposed metric. Kadri Hacioglu, Wayne H. Ward |
INTERSPEECH | 1 |
| 2002 | On developing new text and audio corpora and speech recognition tools for the turkish languageabstractThis paper describes recent work towards development of new corpora and tools for Turkish speech research. This effort represents an on-going collaboration between the Center for Özgül Salor-Durna, Bryan L. Pellom, Tolga Çiloglu, Kadri Hacioglu, Mübeccel Demirekler |
INTERSPEECH | 4 |
| 2001 | Dialog-context dependent language modeling combining n-grams and stochastic context-free grammarsabstractWe present our research on dialog dependent language modeling. In accordance with a speech (or sentence) production model in a discourse we split language modeling into two components; namely, dialog dependent concept modeling and syntactic modeling. The concept model is conditioned on the last question prompted by the dialog system and it is structured using n-grams. The syntactic model, which consists of a collection of stochastic context-free grammars one for each concept, describes word sequences that may be used to express the concepts. The resulting LM is evaluated by rescoring N-best lists. We report significant perplexity improvement with moderate word error rate drop within the context of the CU Communicator System; a dialog system for making travel plans by accessing information about flights, hotels and car rentals. Kadri Hacioglu, Wayne H. Ward |
ICASSP | 1 |
| 2001 | Confidence measures for spoken dialogue systemsabstractThis paper provides improved confidence assessment for detection of word-level speech recognition errors, out of domain utterances and incorrect concepts in the CU Communicator system. New features from the speech understanding component are proposed for confidence annotation at the utterance and concept levels. We consider a neural network to combine all features in each level. Using the data collected from a live telephony system, it is shown that 53.2% of incorrectly recognized words, 53.2% of out of domain utterances and 50.1% of incorrect concepts are detected at a 5% false rejection rate. In addition, the confidence measures are used to improve the word recognition accuracy. Several hypotheses from different speech recognizers are compiled into a word-graph. The word-graph is searched for the hypothesis with the best confidence. We report a 14.0% relative word error rate reduction after this confidence rescoring. Rubén San-Segundo-Hernández, Bryan L. Pellom, Kadri Hacioglu, Wayne H. Ward, José Manuel Pardo |
ICASSP | 3 |
| 2001 | A word graph interface for a flexible concept based speech understanding framework
Kadri Hacioglu, Wayne H. Ward |
INTERSPEECH | 1 |
| 2001 | Improvements in audio processing and language modeling in the CU communicatorabstractThis paper presents some up-to-date audio processing techniques which have been developed and integrated into the University of Colorado (CU) communicator system. The CU Communicator is an interactive human-machine dialogue system for airline, hotel and rental car information. The baseline system was fully functional in June 1999. Since then, many improvements have been made. The paper will concentrate on acoustic echo cancellation, voice activity detection (VAD) and language modeling techniques and provide a paradigm for speech and audio processing in a dialog system with barge-in capabilities. Specifically, a real-time block least-mean-square (LMS) algorithm is discussed. A robust voice activity detector using energy threshold is applied to detect user voice. Experimental results are presented and some real-time implementation issues are addressed. Wayne H. Ward, Bryan L. Pellom, Xiuyang Yu, Kadri Hacioglu |
INTERSPEECH | 5 |
| 2000 | Multiuser detection using a genetic algorithm in CDMA communications systemsabstractIn this study, a hybrid approach that employs a genetic algorithm (GA) and a multistage detector (MSD) for the multiuser detection problem in a code-division multiple-access communications system is proposed. Using this approach: (1) the GA is used as the first stage of the MSD to provide a good initial point for successive stages of the MSD and (2) the MSD is embedded into the GA as a "genetic operator" to improve further the fitness of the population at each generation. Such a hybridization of the GA with the MSD reduces its computational complexity by providing faster convergence. In addition, a better initial data estimate supplied by the GA improves the performance of the MSD, and the embedded MSD improves the performance of the GA. Simulation results for the synchronous and asynchronous cases are provided to show that the approach is promising. Cem Ergün, Kadri Hacioglu |
IEEE Trans. Commun. | 2 |
| 1998 | Pulse-by-pulse reoptimization of the synthesis filter in pulse-based codersabstractTwo iterative algorithms for pulse-based linear predictive (PB-LP) analysis are developed. At the kth step, one of the algorithms assumes that the pulse locations of k pulses are known and calculates the LP coefficients jointly with k pulse amplitudes. In contrast, the other algorithm assumes that the previous k-1 pulse amplitudes are known and jointly estimates the LP filter coefficients and the kth pulse amplitude. These algorithms are proposed for synthesis filter reoptimization in pulse-based LP coders, and their effectiveness in multipulse coding is demonstrated through several simulations. Kadri Hacioglu, Allam Hasib |
IEEE Trans. Speech Audio Process. | 1 |
| 1997 | An improved recurrent neural network for M-PAM symbol detectionabstractIn this paper, a fully connected recurrent neural network (RNN) is presented for the recovery of M-ary pulse amplitude modulated (M-PAM) signals in the presence of intersymbol interference and additive white Gaussian noise. The network makes use of two different activation functions. One is the traditional two-level sigmoid function, which is used at its hidden nodes, and the other is the M-level sigmoid function (MSF), which is used at the output node. The shape of the M-level activation function is controlled by two parameters: the slope and shifting parameters. The effect of these parameters on the learning performance is investigated through extensive simulations. In addition, the network is compared with a linear transversal equalizer, a decision feedback equalizer and a recently proposed RNN equalizer which has used a scaled sigmoid function (SSF) at its output node. Comparisons are made in terms of their learning properties and symbol error rates. It is demonstrated that the proposed RNN equalizer performs better, provided that the MSF parameters are properly selected. Kadri Hacioglu |
IEEE Trans. Neural Networks | 1 |