Kadri Hacioglu

dblp:02/4285 · DBLP profile ↗
← Back
31ranked-venue papers
14as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 10 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 7 first-author · 4 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Information extraction and text analysis · 79% Vision and language · 10% Face, body and person analysis · 10%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 50% Learning and educational technologies · 50%
Computer networks
1 paper
Physical-layer communications · 100%
Computer graphics and multimedia
1 paper
Audio and music processing · 100%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
semantic role labeling
0.122005
Semantic Role Labeling Using Different Syntactic Views · ACL 2005
Semantic Role Parsing: Adding Semantic Structure to Unstructured Text · ICDM 2003
Learning and educational technologies › pedagogical agents
animated pedagogical agents
0.012003
Perceptive animated interfaces: first steps toward a new paradigm for human-computer interaction · Proc. IEEE 2003
Human-AI interaction › conversational agents
embodied conversational agents
0.012003
Perceptive animated interfaces: first steps toward a new paradigm for human-computer interaction · Proc. IEEE 2003
Physical-layer communications
code-division multiple access
0.012000
Multiuser detection using a genetic algorithm in CDMA communications systems · IEEE Trans. Commun. 2000
Physical-layer communications › signal detection › multiuser detection
multistage detection
0.012000
Multiuser detection using a genetic algorithm in CDMA communications systems · IEEE Trans. Commun. 2000
Physical-layer communications › signal detection
multiuser detection
0.012000
Multiuser detection using a genetic algorithm in CDMA communications systems · IEEE Trans. Commun. 2000
Audio and music processing › speech coding
linear predictive coding
0.011998
Pulse-by-pulse reoptimization of the synthesis filter in pulse-based coders · IEEE Trans. Speech Audio Process. 1998
Audio and music processing
speech coding
0.011998
Pulse-by-pulse reoptimization of the synthesis filter in pulse-based coders · IEEE Trans. Speech Audio Process. 1998
Computer vision › Face, body and person analysis
facial expression analysis
0.012003
Perceptive animated interfaces: first steps toward a new paradigm for human-computer interaction · Proc. IEEE 2003
Computer vision › Vision and language
multimodal dialogue
0.012003
Perceptive animated interfaces: first steps toward a new paradigm for human-computer interaction · Proc. IEEE 2003

Methods — techniques the papers use, named apart from their topics

support vector machine · 0.1dialogue system · 0.1computer vision · 0.1computer animation · 0.1parse combination · 0.1feature selection · 0.1feature engineering · 0.0hybrid optimization · 0.0genetic algorithm · 0.0pulse amplitude estimation · 0.0iterative optimization · 0.0
YearPublicationVenuePosition
2025 TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation
abstract
Token-based multitasking frameworks like TokenVerse require all training utterances to have labels for all tasks, hindering their ability to leverage partially annotated datasets and scale effectively. We propose TokenVerse++, which introduces learnable vectors in the acoustic embedding space of the XLSR-Transducer ASR model for dynamic task activation. This core mechanism enables training with utterances labeled for only a subset of tasks, a key advantage over TokenVerse. We demonstrate this by successfully integrating a dataset with partial labels, specifically for ASR and an additional task, language identification, improving overall performance. TokenVerse++ achieves results on par with or exceeding TokenVerse across multiple tasks, establishing it as a more practical multitask alternative without sacrificing ASR performance.
Shashi Kumar, Srikanth R. Madikeri, Esaú Villatoro-Tello, Sergio Burdisso, Pradeep Rangappa, Roberto Andrés Vasco Carofilis, Petr Motlícek, D. S. Karthik Pandia, Shankar Venkatesan, Kadri Hacioglu, Andreas Stolcke
ASRU10
2025 Better Semi-supervised Learning for Multi-domain ASR Through Incremental Retraining and Data Filtering
abstract
Fine-tuning pretrained ASR models for specific domains is challenging when labeled data is scarce. But unlabeled audio and labeled data from related domains are often available. We propose an incremental semi-supervised learning pipeline that first integrates a small in-domain labeled set and an auxiliary dataset from a closely related domain, achieving a relative improvement of 4% over no auxiliary data. Filtering based on multi-model consensus or named entity recognition (NER) is then applied to select and iteratively refine pseudo-labels, showing slower performance saturation compared to random selection. Evaluated on the multi-domain Wow call center and Fisher English corpora, it outperforms single-step fine-tuning. Consensus-based filtering outperforms other methods, providing up to 22.3% relative improvement on Wow and 24.8% on Fisher over single-step fine-tuning with random selection. NER is the second-best filter, providing competitive performance at a lower computational cost.
Roberto Andrés Vasco Carofilis, Pradeep Rangappa, Srikanth R. Madikeri, Shashi Kumar, Sergio Burdisso, Jeena J. Prakash, Esaú Villatoro-Tello, Petr Motlícek, Bidisha Sharma, Kadri Hacioglu, Shankar Venkatesan, Saurabh Vyas, Andreas Stolcke
INTERSPEECH10
2025 Better Pseudo-labeling with Multi-ASR Fusion and Error Correction by SpeechLLM
abstract
Automatic speech recognition (ASR) models rely on high-quality transcribed data for effective training. Generating pseudo-labels for large unlabeled audio datasets often relies on complex pipelines that combine multiple ASR outputs through multi-stage processing, leading to error propagation, information loss and disjoint optimization. We propose a unified multi-ASR prompt-driven framework using postprocessing by either textual or speech-based large language models (LLMs), replacing voting or other arbitration logic for reconciling the ensemble outputs. We perform a comparative study of multiple architectures with and without LLMs, showing significant improvements in transcription accuracy compared to traditional methods. Furthermore, we use the pseudo-labels generated by the various approaches to train semi-supervised ASR models for different datasets, again showing improved performance with textual and speechLLM transcriptions compared to baselines.
Jeena J. Prakash, Blessingh Kumar, Kadri Hacioglu, Bidisha Sharma, Sindhuja Gopalan, Malolan Chetlur, Shankar Venkatesan, Andreas Stolcke
INTERSPEECH3
2025 Efficient Data Selection for Domain Adaptation of ASR Using Pseudo-Labels and Multi-Stage Filtering
abstract
Fine-tuning pretrained ASR models for specific domains is challenging for small organizations with limited labeled data and computational resources. Here, we explore different data selection pipelines and propose a robust approach that improves ASR adaptation by filtering pseudo-labels generated using Whisper (encoder-decoder) and Zipformer (transducer) models. Our approach integrates multiple selection strategies -- including word error rate (WER) prediction, named entity recognition (NER), and character error rate (CER) analysis -- to extract high-quality training segments. We evaluate our method on Whisper and Zipformer using a 7500-hour baseline, comparing it to a CER-based approach relying on hypotheses from three ASR systems. Fine-tuning on 7500 hours of pseudo-labeled call center data achieves 12.3% WER, while our filtering reduces the dataset to 100 hours (1.4%) with similar performance; a similar trend is observed on Fisher English.
Pradeep Rangappa, Roberto Andrés Vasco Carofilis, Jeena J. Prakash, Shashi Kumar, Sergio Burdisso, Srikanth R. Madikeri, Esaú Villatoro-Tello, Bidisha Sharma, Petr Motlícek, Kadri Hacioglu, Shankar Venkatesan, Saurabh Vyas, Andreas Stolcke
INTERSPEECH10
2012 Improving L1-Specific Phonological Error Diagnosis in Computer Assisted Pronunciation Training
Theban Stanley, Kadri Hacioglu
INTERSPEECH2
2005 Semantic Role Labeling Using Different Syntactic Views
abstract
Semantic role labeling is the process of annotating the predicate-argument structure in text with semantic labels. In this paper we present a state-of-the-art baseline semantic role labeling system based on Support Vector Machine classifiers. We show improvements on this system by: i) adding new features including features extracted from dependency parses, ii) performing feature selection and calibration and iii) combining parses obtained from semantic parsers trained using different syntactic views. Error analysis of the baseline system showed that approximately half of the argument identification errors resulted from parse errors in which there was no syntactic constituent that aligned with the correct argument. In order to address this problem, we combined semantic parses from a Minipar syntactic parse and from a chunked syntactic representation with our original baseline system which was based on Charniak parses. All of the reported techniques resulted in performance improvements.
Sameer Pradhan, Wayne H. Ward, Kadri Hacioglu, James H. Martin, Daniel Jurafsky
ACL3
2005 Automatic Time Expression Labeling for English and Chinese Text
Kadri Hacioglu, Benjamin Douglas
CICLing1
2005 Semantic Role Chunking Combining Complementary Syntactic Views
Sameer Pradhan, Kadri Hacioglu, Wayne H. Ward, James H. Martin, Daniel Jurafsky
CoNLL2
2005 Support Vector Learning for Semantic Argument Classification
Sameer Pradhan, Kadri Hacioglu, Valerie Krugler, Wayne H. Ward, James H. Martin, Daniel Jurafsky
Mach. Learn.2
2004 Semantic Role Labeling Using Dependency Trees
Kadri Hacioglu
COLING1
2004 Semantic Role Labeling by Tagging Syntactic Chunks
Kadri Hacioglu, Sameer Pradhan, Wayne H. Ward, James H. Martin, Daniel Jurafsky
CoNLL1
2004 Parsing speech into articulatory events
abstract
In this paper, the states in the speech production process are defined by a number of categorical articulatory features. We describe a detector that outputs a stream (sequence of classes) for each articulatory feature given the Mel frequency cepstral coefficient (MFCC) representation of the input speech. The detector consists of a bank of recurrent neural network (RNN) classifiers, a variable depth lattice generator and Viterbi decoder. A bank of classifiers has been previously used for articulatory feature detection by many researchers. We extend their work first by creating variable depth lattices for each feature and then by combining them into product lattices for rescoring using the Viterbi algorithm. During the rescoring we incorporate language and duration constraints along with the posterior probabilities of classes provided by the RNN classifiers. We present our results for the place and manner features using TIMIT data, and compare the results to a baseline system. We report performance improvements both at the frame and segment levels.
Kadri Hacioglu, Bryan L. Pellom, Wayne H. Ward
ICASSP (1)1
2004 Shallow Semantic Parsing using Support Vector Machines
Sameer Pradhan, Wayne H. Ward, Kadri Hacioglu, James H. Martin, Daniel Jurafsky
HLT-NAACL3
2004 The adaptive path selective decorrelating detector: performance analysis with channel estimation errors
Ali Hakan Ulusoy, Ahmet Rizaner, Kadri Hacioglu, Hasan Amca
Signal Process.3
2003 A distributed architecture for robust automatic speech recognition
abstract
In this paper, we attempt to decompose a state-of-the-art speech recognition system into its components and define an infrastructure that allows a flexible, efficient and effective interaction among the components. Motivated by the success of DARPA Communicator program, we select the open source Galaxy architecture as our development test bed. It consists of a hub that allows communication among servers connected to it by message passing and supports the plug-and-play paradigm. In addition to message passing it supports high bandwidth data (binary or audio) transfer between servers via a brokering scheme. For several reasons, we believe that it is the right time to start developing a distributed framework for speech recognition along with data and protocol standards supporting interoperability. We present our work towards that goal using the Colorado University (CU) Sonic recognizer. We divide Sonic into a number of components and structure it around the Hub. We describe the system in some detail and report on its present status with some possibilities for future development.
Kadri Hacioglu, Bryan L. Pellom
ICASSP (1)1
2003 Recent improvements in the CU Sonic ASR system for noisy speech: the SPINE task
abstract
We report on recent improvements in the University of Colorado system for the DARPA/NRL Speech in Noisy Environments (SPINE) task. In particular, we describe our efforts on improving acoustic and language modeling for the task and investigate methods for unsupervised speaker and environment adaptation from limited data. We show that the MAPLR adaptation method outperforms single and multiple regression class MLLR on the SPINE task. Our current SPINE system uses the Sonic speech recognition engine that was developed at the University of Colorado. This system is shown to have a word error rate of 31.5% on the SPINE-2 evaluation data. These improvements amount to a 16% reduction in relative word error rate compared to our previous SPINE-2 system fielded in the November 2001 DARPA/NRL evaluation.
Bryan L. Pellom, Kadri Hacioglu
ICASSP (1)2
2003 Semantic Role Parsing: Adding Semantic Structure to Unstructured Text
abstract
There is an ever-growing need to add structure in the form of semantic markup to the huge amounts of unstructured text data now available. We present the technique of shallow semantic parsing, the process of assigning a simple WHO did WHAT to WHOM, etc., structure to sentences in text, as a useful tool in achieving this goal. We formulate the semantic parsing problem as a classification problem using support vector machines. Using a hand-labeled training set and a set of features drawn from earlier work together with some feature enhancements, we demonstrate a system that performs better than all other published results on shallow semantic parsing.
Sameer Pradhan, Kadri Hacioglu, Wayne H. Ward, James H. Martin, Daniel Jurafsky
ICDM2
2003 On lexicon creation for turkish LVCSR
abstract
In this paper, we address the lexicon design problem in Turkish large vocabulary speech recognition. Although we focus only on Turkish, the methods described here are general enough that they can be considered for other agglutinative languages like Finnish, Korean etc. In an agglutinative language, several words can be created from a single root word using a rich collection of morphological rules. So, a virtually infinite size lexicon is required to cover the language if words are used as the basic units. The standard approach to this problem is to discover a number of primitive units so that a large set of words can be created by compounding those units. Two broad classes of methods are available for splitting words into their sub-units; morphology-based and data-driven methods. Although the word splitting significantly reduces the out of vocabulary rate, it shrinks the context and increases acoustic confusibility. We have used two methods to address the latter. In one method, we use word counts to avoid splitting of high frequency lexical units, and in the other method, we recompound splits according to a probabilistic measure. We present experimental results that show the methods are very effective to lower the word error rate at the expense of lexicon size. 1.
Kadri Hacioglu, Bryan L. Pellom, Tolga Çiloglu, Özlem Öztürk, Mikko Kurimo, Mathias Creutz
INTERSPEECH1
2003 Target Word Detection and Semantic Role Chunking using Support Vector Machines
Kadri Hacioglu, Wayne H. Ward
HLT-NAACL1
2003 Question Classification with Support Vector Machines and Error Correcting Codes
Kadri Hacioglu, Wayne H. Ward
HLT-NAACL1
2003 Perceptive animated interfaces: first steps toward a new paradigm for human-computer interaction
abstract
This paper presents a vision of the near future in which computer interaction is characterized by natural face-to-face conversations with lifelike characters that speak, emote, and gesture. These animated agents will converse with people much like people converse effectively with assistants in a variety of focused applications. Despite the research advances required to realize this vision, and the lack of strong experimental evidence that animated agents improve human-computer interaction, we argue that initial prototypes of perceptive animated interfaces can be developed today, and that the resulting systems will provide more effective and engaging communication experiences than existing systems. In support of this hypothesis, we first describe initial experiments using an animated character to teach speech and language skills to children with hearing problems, and classroom subjects and social skills to children with autistic spectrum disorder. We then show how existing dialogue system architectures can be transformed into perceptive animated interfaces by integrating computer vision and animation capabilities. We conclude by describing the Colorado Literacy Tutor, a computer-based literacy program that provides an ideal testbed for research and development of perceptive animated interfaces, and consider next steps required to realize the vision.
Ronald A. Cole, Sarel van Vuuren, Bryan L. Pellom, Kadri Hacioglu, Jiyong Ma, Javier Movellan, Scott Schwartz, David Wade-Stein, Wayne H. Ward
Proc. IEEE4
2002 A concept graph based confidence measure
abstract
In this paper, the confidence measure of a hypothesized word is derived from its posterior probability. In contrast to common approaches, in which N-best lists or word graphs/lattices are used, the posterior probabilities are derived from a concept graph. The concept graph is obtained from a word graph through a partial parsing process using semantic grammars. This approach allows us to use relatively complex and better language models along with acoustic models to compute word posterior probabilities. The language model used is comprised of stochastic context free grammars (one for each concept) and an n-gram concept language model. We show that the posterior probabilities computed on concept graphs outperform those computed on word graphs when used as confidence measures. Results are presented within the context of Colorado University (CU) Communicator System; a telephone-based dialog system for making travel plans by accessing information about flights, hotels and car rentals.
Kadri Hacioglu, Wayne H. Ward
ICASSP1
2002 A figure of merit for the analysis of spoken dialog systems
abstract
In this paper, a single metric, which we will call the figure of merit , for the quantitative analysis and comparison of spoken dialog systems is introduced. This figure of merit is the product of the weighted dialog accuracy (expressed as the rate of success) and the weighted dialog efficiency (expressed as the average number of concepts per turn). Actually, it is highly desirable to have a quick and accurate dialog. However, these two requirements are conflicting. That is, an improvement in efficiency is accomplished at the expense of accuracy or vice versa. This makes difficult to compare two different spoken dialog systems or tune a particular system. We believe that this figure of merit would avoid those difficulties. To illustrate its use, we consider spoken dialog systems with different dialog strategies and compare them by performing quantitative analysis based on the finate state models of information items using the proposed metric.
Kadri Hacioglu, Wayne H. Ward
INTERSPEECH1
2002 On developing new text and audio corpora and speech recognition tools for the turkish language
abstract
This paper describes recent work towards development of new corpora and tools for Turkish speech research. This effort represents an on-going collaboration between the Center for
Özgül Salor-Durna, Bryan L. Pellom, Tolga Çiloglu, Kadri Hacioglu, Mübeccel Demirekler
INTERSPEECH4
2001 Dialog-context dependent language modeling combining n-grams and stochastic context-free grammars
abstract
We present our research on dialog dependent language modeling. In accordance with a speech (or sentence) production model in a discourse we split language modeling into two components; namely, dialog dependent concept modeling and syntactic modeling. The concept model is conditioned on the last question prompted by the dialog system and it is structured using n-grams. The syntactic model, which consists of a collection of stochastic context-free grammars one for each concept, describes word sequences that may be used to express the concepts. The resulting LM is evaluated by rescoring N-best lists. We report significant perplexity improvement with moderate word error rate drop within the context of the CU Communicator System; a dialog system for making travel plans by accessing information about flights, hotels and car rentals.
Kadri Hacioglu, Wayne H. Ward
ICASSP1
2001 Confidence measures for spoken dialogue systems
abstract
This paper provides improved confidence assessment for detection of word-level speech recognition errors, out of domain utterances and incorrect concepts in the CU Communicator system. New features from the speech understanding component are proposed for confidence annotation at the utterance and concept levels. We consider a neural network to combine all features in each level. Using the data collected from a live telephony system, it is shown that 53.2% of incorrectly recognized words, 53.2% of out of domain utterances and 50.1% of incorrect concepts are detected at a 5% false rejection rate. In addition, the confidence measures are used to improve the word recognition accuracy. Several hypotheses from different speech recognizers are compiled into a word-graph. The word-graph is searched for the hypothesis with the best confidence. We report a 14.0% relative word error rate reduction after this confidence rescoring.
Rubén San-Segundo-Hernández, Bryan L. Pellom, Kadri Hacioglu, Wayne H. Ward, José Manuel Pardo
ICASSP3
2001 A word graph interface for a flexible concept based speech understanding framework
Kadri Hacioglu, Wayne H. Ward
INTERSPEECH1
2001 Improvements in audio processing and language modeling in the CU communicator
abstract
This paper presents some up-to-date audio processing techniques which have been developed and integrated into the University of Colorado (CU) communicator system. The CU Communicator is an interactive human-machine dialogue system for airline, hotel and rental car information. The baseline system was fully functional in June 1999. Since then, many improvements have been made. The paper will concentrate on acoustic echo cancellation, voice activity detection (VAD) and language modeling techniques and provide a paradigm for speech and audio processing in a dialog system with barge-in capabilities. Specifically, a real-time block least-mean-square (LMS) algorithm is discussed. A robust voice activity detector using energy threshold is applied to detect user voice. Experimental results are presented and some real-time implementation issues are addressed.
Wayne H. Ward, Bryan L. Pellom, Xiuyang Yu, Kadri Hacioglu
INTERSPEECH5
2000 Multiuser detection using a genetic algorithm in CDMA communications systems
abstract
In this study, a hybrid approach that employs a genetic algorithm (GA) and a multistage detector (MSD) for the multiuser detection problem in a code-division multiple-access communications system is proposed. Using this approach: (1) the GA is used as the first stage of the MSD to provide a good initial point for successive stages of the MSD and (2) the MSD is embedded into the GA as a "genetic operator" to improve further the fitness of the population at each generation. Such a hybridization of the GA with the MSD reduces its computational complexity by providing faster convergence. In addition, a better initial data estimate supplied by the GA improves the performance of the MSD, and the embedded MSD improves the performance of the GA. Simulation results for the synchronous and asynchronous cases are provided to show that the approach is promising.
Cem Ergün, Kadri Hacioglu
IEEE Trans. Commun.2
1998 Pulse-by-pulse reoptimization of the synthesis filter in pulse-based coders
abstract
Two iterative algorithms for pulse-based linear predictive (PB-LP) analysis are developed. At the kth step, one of the algorithms assumes that the pulse locations of k pulses are known and calculates the LP coefficients jointly with k pulse amplitudes. In contrast, the other algorithm assumes that the previous k-1 pulse amplitudes are known and jointly estimates the LP filter coefficients and the kth pulse amplitude. These algorithms are proposed for synthesis filter reoptimization in pulse-based LP coders, and their effectiveness in multipulse coding is demonstrated through several simulations.
Kadri Hacioglu, Allam Hasib
IEEE Trans. Speech Audio Process.1
1997 An improved recurrent neural network for M-PAM symbol detection
abstract
In this paper, a fully connected recurrent neural network (RNN) is presented for the recovery of M-ary pulse amplitude modulated (M-PAM) signals in the presence of intersymbol interference and additive white Gaussian noise. The network makes use of two different activation functions. One is the traditional two-level sigmoid function, which is used at its hidden nodes, and the other is the M-level sigmoid function (MSF), which is used at the output node. The shape of the M-level activation function is controlled by two parameters: the slope and shifting parameters. The effect of these parameters on the learning performance is investigated through extensive simulations. In addition, the network is compared with a linear transversal equalizer, a decision feedback equalizer and a recently proposed RNN equalizer which has used a scaled sigmoid function (SSF) at its output node. Comparisons are made in terms of their learning properties and symbol error rates. It is demonstrated that the proposed RNN equalizer performs better, provided that the MSF parameters are properly selected.
Kadri Hacioglu
IEEE Trans. Neural Networks1