Imed Zitouni

dblp:26/1674 · DBLP profile ↗
← Back
77ranked-venue papers
25as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 58 · 22 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 13 first-authorDatabases, data management, data science and information retrieval · 22Applied, interdisciplinary, general and emerging computing · 3Computer networks · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
14 papers
Information retrieval · 95% Data mining · 4% Knowledge graphs · 1%
Artificial intelligence
14 papers
Information extraction and text analysis · 55% Reinforcement learning · 18% Question answering and dialogue systems · 14%
Software engineering, system software, and programming languages
1 paper
Services computing and microservices · 100%

Topics — the 30 heaviest of 46, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
evaluation
0.942016
Quality Management in Crowdsourcing using Gold Judges Behavior · WSDM 2016
Predicting User Satisfaction with Intelligent Assistants · SIGIR 2016
Contextual and dimensional relevance judgments for reusable SERP-level evaluation · WWW 2014
Information retrieval › evaluation
user satisfaction prediction
0.842017
User Interaction Sequences for Search Satisfaction Prediction · SIGIR 2017
Predicting User Satisfaction with Intelligent Assistants · SIGIR 2016
Comparing client and server dwell time estimates for click-level satisfaction prediction · SIGIR 2014
Natural language and speech › Information extraction and text analysis › named entity recognition
mention detection
0.652014
Aligned-Parallel-Corpora Based Semi-Supervised Learning for Arabic Mention Detection · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Improving Mention Detection Robustness to Noisy Input · EMNLP 2010
Enhancing Mention Detection Using Projection via Aligned Corpora · EMNLP 2010
Information retrieval › retrieval models › neural retrieval › dense retrieval
bi-encoder retrieval
0.612022
Exploring Dual Encoder Architectures for Question Answering · EMNLP 2022
Information retrieval › retrieval models › neural retrieval
dense retrieval
0.612022
Exploring Dual Encoder Architectures for Question Answering · EMNLP 2022
Information retrieval › user behavior
search satisfaction
0.422016
Detecting Good Abandonment in Mobile Search · WWW 2016
Modeling dwell time to predict click-level satisfaction · WSDM 2014
Information retrieval › interactive information retrieval › conversational information seeking
conversational search
0.312018
Conversational Semantic Search: Looking Beyond Web Search, Q&A and Dialog Systems · WSDM 2018
Information retrieval
question answering and dialogue systems
0.312018
Conversational Semantic Search: Looking Beyond Web Search, Q&A and Dialog Systems · WSDM 2018
Information retrieval › search engines
semantic search
0.312018
Conversational Semantic Search: Looking Beyond Web Search, Q&A and Dialog Systems · WSDM 2018
Machine learning › Reinforcement learning
off-policy evaluation
0.312017
Off-policy evaluation for slate recommendation · NIPS 2017
Data mining › probabilistic model › temporal point process
hawkes process
0.312017
User Interaction Sequences for Search Satisfaction Prediction · SIGIR 2017
Information retrieval › ranking
learning to rank
0.312017
Off-policy evaluation for slate recommendation · NIPS 2017
Information retrieval
retrieval evaluation
0.212016
Is This Your Final Answer?: Evaluating the Effect of Answers on Good Abandonment in Mobile Search · SIGIR 2016
Information retrieval › user behavior
search abandonment
0.212016
Is This Your Final Answer?: Evaluating the Effect of Answers on Good Abandonment in Mobile Search · SIGIR 2016
Information retrieval › evaluation
offline evaluation
0.212015
Toward Predicting the Outcome of an A/B Experiment for Search Relevance · WSDM 2015
Information retrieval › retrieval evaluation
ranking evaluation
0.212015
Toward Predicting the Outcome of an A/B Experiment for Search Relevance · WSDM 2015
Information retrieval › evaluation
effectiveness metrics
0.212014
Contextual and dimensional relevance judgments for reusable SERP-level evaluation · WWW 2014
Information retrieval › evaluation
relevance judgment
0.212014
Contextual and dimensional relevance judgments for reusable SERP-level evaluation · WWW 2014
Natural language and speech › Question answering and dialogue systems
open-domain question answering
0.212022
Exploring Dual Encoder Architectures for Question Answering · EMNLP 2022
Information retrieval › evaluation › user-oriented evaluation
preference-based evaluation
0.212013
Relevance dimensions in preference-based IR evaluation · SIGIR 2013
Information retrieval
search result diversification
0.212013
Relevance dimensions in preference-based IR evaluation · SIGIR 2013
Information retrieval › user behavior › search behavior
click model
0.122016
Predicting User Satisfaction with Intelligent Assistants · SIGIR 2016
Modeling dwell time to predict click-level satisfaction · WSDM 2014
Information retrieval
user behavior
0.122016
Predicting User Satisfaction with Intelligent Assistants · SIGIR 2016
Modeling dwell time to predict click-level satisfaction · WSDM 2014
Natural language and speech › Question answering and dialogue systems › conversational agents
intelligent assistants
0.112019
Automatic Task Completion Flows from Web APIs · SIGIR 2019
Knowledge graphs › knowledge graph querying
knowledge graph question answering
0.112018
Conversational Semantic Search: Looking Beyond Web Search, Q&A and Dialog Systems · WSDM 2018
Machine learning › Reinforcement learning › multi-armed bandit
combinatorial bandits
0.112017
Off-policy evaluation for slate recommendation · NIPS 2017
Information retrieval › interactive information retrieval
search result interaction
0.112017
User Interaction Sequences for Search Satisfaction Prediction · SIGIR 2017
Machine learning › Kernel, tree and ensemble methods
classifier combination
0.112008
Constrained Minimization and Discriminative Training for Natural Language Call Routing · IEEE Trans. Speech Audio Process. 2008
Natural language and speech › Information extraction and text analysis › multilingual NLP
cross-lingual NLP
0.112008
Mention Detection Crossing the Language Barrier · EMNLP 2008
Information retrieval › web search
mobile search
0.112016
Is This Your Final Answer?: Evaluating the Effect of Answers on Good Abandonment in Mobile Search · SIGIR 2016

Methods — techniques the papers use, named apart from their topics

t-SNE · 1.1graph path extraction · 0.8unbiased estimator · 0.6implicit feedback · 0.4action sequence modeling · 0.4acoustic features · 0.4semantic functional unit composition · 0.3subsequence mining · 0.3hawkes process · 0.3combinatorial bandits · 0.3combinatorial bandit · 0.3predictive modeling · 0.2log analysis · 0.2gold tasks · 0.2gesture feature analysis · 0.2classifier · 0.2behavioral signals · 0.2semi-supervised learning · 0.2
YearPublicationVenuePosition
2026 Entity Image and Mixed-Modal Image Retrieval Datasets
abstract
Despite advances in multimodal learning, challenging benchmarks for mixed-modal image retrieval that combines visual and textual information are lacking. This paper introduces a novel benchmark to rigorously evaluate image retrieval that demands deep cross-modal contextual understanding. We present two new datasets: the Entity Image Dataset (EI), providing canonical images for Wikipedia entities, and the Mixed-Modal Image Retrieval Dataset (MMIR), derived from the WIT dataset. The MMIR benchmark features two challenging query types requiring models to ground textual descriptions in the context of provided visual entities: single entity-image queries (one entity image with descriptive text) and multi-entity-image queries (multiple entity images with relational text). We empirically validate the benchmark's utility as both a training corpus and an evaluation set for mixed-modal retrieval. The quality of both datasets is further affirmed through crowd-sourced human annotations. The datasets are accessible through the GitHub page: https://github.com/google-research-datasets/wit-retrieval.
Cristian-Ioan Blaga, Paul Suganthan G. C., Sahil Dua, Krishna Srinivasan, Enrique Alfonseca, Péter Dornbach, Tom Duerig, Imed Zitouni
LREC8
2022 Exploring Dual Encoder Architectures for Question Answering
abstract
Dual encoders have been used for questionanswering (QA) and information retrieval (IR) tasks with good results.Previous research focuses on two major types of dual encoders, Siamese Dual Encoder (SDE), with parameters shared across two encoders, and Asymmetric Dual Encoder (ADE), with two distinctly parameterized encoders.In this work, we explore different ways in which the dual encoder can be structured, and evaluate how these differences can affect their efficacy in terms of QA retrieval tasks.By evaluating on MS MARCO, open domain NQ and the Mul-tiReQA benchmarks, we show that SDE performs significantly better than ADE.We further propose three different improved versions of ADEs by sharing or freezing parts of the architectures between two encoder towers.We find that sharing parameters in projection layers would enable ADEs to perform competitively with or outperform SDEs.We further explore and explain why parameter sharing in projection layer significantly improves the efficacy of the dual encoders, by directly probing the embedding spaces of the two encoder towers with t-SNE algorithm (van der Maaten and Hinton, 2008).
Jianmo Ni, Dan Bikel, Enrique Alfonseca, Chen Qu 0001, Imed Zitouni
EMNLP7
2020 Wasf-Vec: Topology-based Word Embedding for Modern Standard Arabic and Iraqi Dialect Ontology
abstract
Word clustering is a serious challenge in low-resource languages. Since words that share semantics are expected to be clustered together, it is common to use a feature vector representation generated from a distributional theory-based word embedding method. The goal of this work is to utilize Modern Standard Arabic (MSA) for better clustering performance of the low-resource Iraqi vocabulary. We began with a new Dialect Fast Stemming Algorithm (DFSA) that utilizes the MSA data. The proposed algorithm achieved 0.85 accuracy measured by the F1 score. Then, the distributional theory-based word embedding method and a new simple, yet effective, feature vector named Wasf-Vec word embedding are tested. Wasf-Vec word representation utilizes a word’s topology features. The difference between Wasf-Vec and distributional theory-based word embedding is that Wasf-Vec captures relations that are not contextually based. The embedding is followed by an analysis of how the dialect words are clustered within other MSA words. The analysis is based on the word semantic relations that are well supported by solid linguistic theories to shed light on the strong and weak word relation representations identified by each embedding method. The analysis is handled by visualizing the feature vector in two-dimensional (2D) space. The feature vectors of the distributional theory-based word embedding method are plotted in 2D space using the t-sne algorithm, while the Wasf-Vec feature vectors are plotted directly in 2D space. A word’s nearest neighbors and the distance-histograms of the plotted words are examined. For validation purpose of the word classification used in this article, the produced classes are employed in Class-based Language Modeling (CBLM). Wasf-Vec CBLM achieved a 7% lower perplexity (pp) than the distributional theory-based word embedding method CBLM. This result is significant when working with low-resource languages.
Tiba Zaki Abdulhameed, Imed Zitouni, Ikhlas Abdel-Qader
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2020 Editorial from the New Editor-in-Chief: the Era of Natural Language Processing Innovations on Asian and Low-Resource Languages
abstract
No abstract available.
Imed Zitouni
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2019 Slot Tagging for Task Oriented Spoken Language Understanding in Human-to-Human Conversation Scenarios
abstract
Task oriented language understanding (LU) in human-to-machine (H2M) conversations has been extensively studied for personal digital assistants.In this work, we extend the task oriented LU problem to human-to-human (H2H) conversations, focusing on the slot tagging task.Recent advances on LU in H2M conversations have shown accuracy improvements by adding encoded knowledge from different sources.Inspired by this, we explore several variants of a bidirectional LSTM architecture that relies on different knowledge sources, such as Web data, search engine click logs, expert feedback from H2M models, as well as previous utterances in the conversation.We also propose ensemble techniques that aggregate these different knowledge sources into a single model.Experimental evaluation on a four-turn Twitter dataset in the restaurant and music domains shows improvements in the slot tagging F1-score of up to 6.09% compared to existing approaches.
Kunho Kim, Rahul Jha, Kyle Williams 0003, Alex Marin, Imed Zitouni
CoNLL5
2019 Automatic Task Completion Flows from Web APIs
abstract
The Web contains many APIs that could be combined in countless ways to enable Intelligent Assistants to complete all sorts of tasks. We propose a method to automatically produce task completion flows from a collection of these APIs by combining them in a graph and automatically extracting paths from the graph for task completion. These paths chain together API calls and use the output of executed APIs as inputs to others. We automatically extract these paths from an API graph in response to a user query and then rank the paths by the likelihood of them leading to user satisfaction. We apply our approach for task completion in the email and calendar domains and show how it can be used to automatically create task completion flows.
Kyle Williams 0003, Seyyed Hadi Hashemi, Imed Zitouni
SIGIR3
2018 Measuring User Satisfaction on Smart Speaker Intelligent Assistants Using Intent Sensitive Query Embeddings
abstract
Intelligent assistants are increasingly being used on smart speaker devices, such as Amazon Echo, Google Home, Apple Homepod, and Harmon Kardon Invoke with Cortana. Typically, user satisfaction measurement relies on user interaction signals, such as clicks and scroll movements, in order to determine if a user was satisfied. However, these signals do not exist for smart speakers, which creates a challenge for user satisfaction evaluation on these devices. In this paper, we propose a new signal, user intent, as a means to measure user satisfaction. We propose to use this signal to model user satisfaction in two ways: 1) by developing intent sensitive word embeddings and then using sequences of these intent sensitive query representations to measure user satisfaction; 2) by representing a user's interactions with a smart speaker as a sequence of user intents and thus using this sequence to identify user satisfaction. Our experimental results indicate that our proposed user satisfaction models based on the intent-sensitive query representations have statistically significant improvements over several baselines in terms of common classification evaluation metrics. In particular, our proposed task satisfaction prediction model based on intent-sensitive word embeddings has a 11.81% improvement over a generative model baseline and 6.63% improvement over a user satisfaction prediction model based on Skip-gram word embeddings in terms of the F1 metric. Our findings have implications for the evaluation of Intelligent Assistant systems.
Seyyed Hadi Hashemi, Kyle Williams 0003, Ahmed El Kholy, Imed Zitouni, Paul A. Crook
CIKM4
2018 Impact of Domain and User's Learning Phase on Task and Session Identification in Smart Speaker Intelligent Assistants
abstract
Task and session identification is a key element of system evaluation and user behavior modeling in Intelligent Assistant (IA) systems. However, identifying task and sessions for IAs is challenging due to the multi-task nature of IAs and the differences in the ways they are used on different platforms, such as smart-phones, cars, and smart speakers. Furthermore, usage behavior may differ among users depending on their expertise with the system and the tasks they are interested in performing. In this study, we investigate how to identify tasks and sessions in IAs given these differences. To do this, we analyze data based on the interaction logs of two IAs integrated with smart-speakers. We fit Gaussian Mixture Models to estimate task and session boundaries and show how a model with 3 components models user interactivity time better than a model with 2 components. We then show how session boundaries differ for users depending on whether they are in a learning-phase or not. Finally, we study how user inter-activity times differs depending on the task that the user is trying to perform. Our findings show that there is no single task or session boundary that can be used for IA evaluation. Instead, these boundaries are influenced by the experience of the user and the task they are trying to perform. Our findings have implications for the study and evaluation of Intelligent Agent Systems.
Seyyed Hadi Hashemi, Kyle Williams 0003, Ahmed El Kholy, Imed Zitouni, Paul A. Crook
CIKM4
2018 Conversational Semantic Search: Looking Beyond Web Search, Q&A and Dialog Systems
abstract
User expectations of web search are changing. They are expecting search engines to answer questions, to be more conversational, and to offer means to complete tasks on their behalf. At the same time, to increase the breadth of tasks that personal digital assistants (PDAs), such as Microsoft»s Cortana or Amazon»s Alexa, are capable of, PDAs need to better utilize information about the world, a significant amount of which is available in the knowledge bases and answers built for search engines. It thus seems likely that the underlying systems that power web search and PDAs will converge. This demonstration presents a system that merges elements of traditional multi-turn dialog systems with web based question answering. This demo focuses on the automatic composition of semantic functional units, Botlets, to generate responses to user»s natural language (NL) queries. We show that such a system can be trained to combine information from search engine answers with PDA tasks to enable new user experiences.
Paul A. Crook, Alex Marin, Vipul Agarwal, Samantha Anderson, Ohyoung Jang, Aliasgar Lanewala, Karthik Tangirala, Imed Zitouni
WSDM8
2017 Beyond Success Rate: Utility as a Search Quality Metric for Online Experiments
abstract
User satisfaction metrics are an integral part of search engine development as they help system developers to understand and evaluate the quality of the user experience. Research to date has mostly focused on predicting success or frustration as a proxy for satisfaction. However, users' search experience is more complex than merely being either successful or not. As such, using success rate as a measure of satisfaction can be limiting. In this work, we propose the use of utility as a measure of searcher satisfaction. This concept represents the fulfillment a user receives from con-suming a service and explains how users aim to gain optimal overall satisfaction. Our utility metrics measure the user satisfac-tion by aggregating all their interaction with the search engine. These interactions are represented as a timeline of actions and their dwelltimes, where each action is classified as having a posi-tive or negative effect on the user. We examine sessions mined from Bing logs, with multi-point scale assessment of searcher satisfaction and show that utility is a better proxy for satisfaction compared to success. Leveraging that data, we design metrics of searcher satisfaction that assess the overall utility accumulated by a user during her search session. We use real user traffic from millions of users in an A/B setting to compare utility metrics to success rate metrics. We show that utility is a better metric for evaluating searcher satisfaction with the search engine, and a more sensitive and accurate metric when compared to predicting success. These metrics are currently adopted as the top-level met-ric for evaluating the thousands of A/B experiments that are run on Bing each year.
Widad Machmouchi, Ahmed Awadallah 0001, Imed Zitouni, Georg Buscher
CIKM3
2017 Deep Sequential Models for Task Satisfaction Prediction
abstract
Detecting and understanding implicit signals of user satisfaction are essential for experimentation aimed at predicting searcher satisfaction. As retrieval systems have advanced, search tasks have steadily emerged as accurate units not only to capture searcher's goals but also in understanding how well a system is able to help the user achieve that goal. However, a major portion of existing work on modeling searcher satisfaction has focused on query level satisfaction. The few existing approaches for task satisfaction prediction have narrowly focused on simple tasks aimed at solving atomic information needs.
Rishabh Mehrotra, Ahmed Awadallah 0001, Milad Shokouhi, Emine Yilmaz, Imed Zitouni, Ahmed El Kholy, Madian Khabsa
CIKM5
2017 Does That Mean You're Happy?: RNN-based Modeling of User Interaction Sequences to Detect Good Abandonment
abstract
Queries for which there are no clicks are known as abandoned queries. Differentiating between good and bad abandonment queries has become an important task in search engine evaluation since it allows for better measurement of search engine features that do not require users to click. Examples of these features include answers on the SERP and detailed Web result snippets. In this paper, we investigate how sequences of user interactions on the SERP differ between good and bad abandonment. To do this, we study the behavior patterns on a labeled dataset of abandoned queries and find that they differ in several ways, such as in the number of user interactions and the nature of those interactions. Based on this insight, we frame good abandonment detection as a sequence classification problem. We use a Long Short-Term Memory (LSTM) Recurrent Neural Network (RNN) to model the sequence of user interactions and show that it performs significantly better than other baselines when detecting good abandonment, achieving 71% accuracy. Our findings have implications for search engine evaluation.
Kyle Williams 0003, Imed Zitouni
CIKM2
2017 Hyperarticulation detection in repetitive voice queries using pairwise comparison for improved speech recognition
abstract
Automatic speech recognition systems can benefit from cues in user voice such as hyperarticulation. Traditional approaches typically attempt to define and detect an absolute state of hyperarticulation, which is very difficult, especially on short voice queries. We present a novel approach for hyperarticulation detection using pairwise comparisons and demonstrate its application in a real-world speech recognition system. Our approach uses delta features extracted from a pair of repetitive user utterances. Results show significant improvements in WER (word error rate) by using hyperarticulation information as a feature in a second pass N-best hypotheses rescoring setup.
Ranjitha Gurunath Kulkarni, Ahmed El Kholy, Ziad Al Bawab, Noha Alon, Imed Zitouni, Umut Ozertem, Shuangyu Chang
ICASSP5
2017 Off-policy evaluation for slate recommendation
abstract
This paper studies the evaluation of policies that recommend an ordered set of items (e.g., a ranking) based on some context---a common scenario in web search, ads, and recommendation. We build on techniques from combinatorial bandits to introduce a new practical estimator that uses logged data to estimate a policy's performance. A thorough empirical evaluation on real-world data reveals that our estimator is accurate in a variety of settings, including as a subroutine in a learning-to-rank task, where it achieves competitive performance. We derive conditions under which our estimator is unbiased---these conditions are weaker than prior heuristics for slate evaluation---and experimentally demonstrate a smaller bias than parametric approaches, even when these conditions are violated. Finally, our theory and experiments also show exponential savings in the amount of required data compared with general unbiased estimators.
Adith Swaminathan, Akshay Krishnamurthy, Alekh Agarwal, Miroslav Dudík, John Langford 0001, Damien Jose, Imed Zitouni
NIPS7
2017 User Interaction Sequences for Search Satisfaction Prediction
abstract
Detecting and understanding implicit measures of user satisfaction are essential for meaningful experimentation aimed at enhancing web search quality. While most existing studies on satisfaction prediction rely on users' click activity and query reformulation behavior, often such signals are not available for all search sessions and as a result, not useful in predicting satisfaction. On the other hand, user interaction data (such as mouse cursor movement) is far richer than just click data and can provide useful signals for predicting user satisfaction. In this work, we focus on considering holistic view of user interaction with the search engine result page (SERP) and construct detailed universal interaction sequences of their activity. We propose novel ways of leveraging the universal interaction sequences to automatically extract informative, interpretable subsequences. In addition to extracting frequent, discriminatory and interleaved subsequences, we propose a Hawkes process model to incorporate temporal aspects of user interaction. Through extensive experimentation we show that encoding the extracted subsequences as features enables us to achieve significant improvements in predicting user satisfaction. We additionally present an analysis of the correlation between various subsequences and user satisfaction. Finally, we demonstrate the usefulness of the proposed approach in covering abandonment cases. Our findings provide a valuable tool for fine-grained analysis of user interaction behavior for metric development.
Rishabh Mehrotra, Imed Zitouni, Ahmed Awadallah 0001, Ahmed El Kholy, Madian Khabsa
SIGIR2
2016 Understanding User Satisfaction with Intelligent Assistants
abstract
Voice-controlled intelligent personal assistants, such as Cortana, Google Now, Siri and Alexa, are increasingly becoming a part of users' daily lives, especially on mobile devices. They introduce a significant change in information access, not only by introducing voice control and touch gestures but also by enabling dialogues where the context is preserved. This raises the need for evaluation of their effectiveness in assisting users with their tasks. However, in order to understand which type of user interactions reflect different degrees of user satisfaction we need explicit judgements. In this paper, we describe a user study that was designed to measure user satisfaction over a range of typical scenarios of use: controlling a device, web search, and structured search dialogue. Using this data, we study how user satisfaction varied with different usage scenarios and what signals can be used for modeling satisfaction in the different scenarios. We find that the notion of satisfaction varies across different scenarios, and show that, in some scenarios (e.g. making a phone call), task completion is very important while for others (e.g. planning a night out), the amount of effort spent is key. We also study how the nature and complexity of the task at hand affects user satisfaction, and find that preserving the conversation context is essential and that overall task-level satisfaction cannot be reduced to query-level satisfaction alone. Finally, we shed light on the relative effectiveness and usefulness of voice-controlled intelligent agents, explaining their increasing popularity and uptake relative to the traditional query-response interaction.
Julia Kiseleva, Kyle Williams 0001, Jiepu Jiang, Ahmed Awadallah 0001, Aidan C. Crook, Imed Zitouni, Tasos Anastasakos
CHIIR6
2016 Learning to Account for Good Abandonment in Search Success Metrics
abstract
Abandonment in web search has been widely used as a proxy to measure user satisfaction. Initially it was considered a signal of dissatisfaction, however with search engines moving towards providing answer-like results, a new category of abandonment was introduced and referred to as Good Abandonment. Predicting good abandonment is a hard problem and it was the subject of several previous studies. All those studies have focused, though, on predicting good abandonment in offline settings using manually labeled data. Thus, it remained a challenge how to have an online metric that accounts for good abandonment. In this work we describe how a search success metric can be augmented to account for good abandonment sessions using a machine learned metric that depends on user's viewport information. We use real user traffic from millions of users to evaluate the proposed metric in an A/B experiment. We show that taking good abandonment into consideration has a significant effect on the overall performance of the online metric.
Madian Khabsa, Aidan C. Crook, Ahmed Awadallah 0001, Imed Zitouni, Tasos Anastasakos, Kyle Williams 0001
CIKM4
2016 Predicting User Satisfaction with Intelligent Assistants
abstract
There is a rapid growth in the use of voice-controlled intelligent personal assistants on mobile devices, such as Microsoft's Cortana, Google Now, and Apple's Siri. They significantly change the way users interact with search systems, not only because of the voice control use and touch gestures, but also due to the dialogue-style nature of the interactions and their ability to preserve context across different queries. Predicting success and failure of such search dialogues is a new problem, and an important one for evaluating and further improving intelligent assistants. While clicks in web search have been extensively used to infer user satisfaction, their significance in search dialogues is lower due to the partial replacement of clicks with voice control, direct and voice answers, and touch gestures.
Julia Kiseleva, Kyle Williams 0001, Ahmed Awadallah 0001, Aidan C. Crook, Imed Zitouni, Tasos Anastasakos
SIGIR5
2016 Is This Your Final Answer?: Evaluating the Effect of Answers on Good Abandonment in Mobile Search
abstract
Answers on mobile search result pages have become a common way to attempt to satisfy users without them needing to click on search results. Many different types of answers exist, such as weather, flight and currency answers. Understanding the effect that these different answer types have on mobile user behavior and how they contribute to satisfaction is important for search engine evaluation. We study these two aspects by analyzing the logs of a commercial search engine and through a user study. Our results show that user click, abandonment and engagement behavior differs depending on the answer types present on a page. Furthermore, we find that satisfaction rates differ in the presence of different answer types with simple answer types, such as time zone answers, leading to more satisfaction than more complex answers, such as news answers. Our findings have implications for the study and application of user satisfaction for search systems.
Kyle Williams 0001, Julia Kiseleva, Aidan C. Crook, Imed Zitouni, Ahmed Awadallah 0001, Madian Khabsa
SIGIR4
2016 Quality Management in Crowdsourcing using Gold Judges Behavior
abstract
Crowdsourcing relevance labels has become an accepted practice for the evaluation of IR systems, where the task of constructing a test collection is distributed over large populations of unknown users with widely varied skills and motivations. Typical methods to check and ensure the quality of the crowd's output is to inject work tasks with known answers (gold tasks) on which workers' performance can be measured. However, gold tasks are expensive to create and have limited application. A more recent trend is to monitor the workers' interactions during a task and estimate their work quality based on their behavior. In this paper, we show that without gold behavior signals that reflect trusted interaction patterns, classifiers can perform poorly, especially for complex tasks, which can lead to high quality crowd workers getting blocked while poorly performing workers remain undetected. Through a series of crowdsourcing experiments, we compare the behaviors of trained professional judges and crowd workers and then use the trained judges' behavior signals as gold behavior to train a classifier to detect poorly performing crowd workers. Our experiments show that classification accuracy almost doubles in some tasks with the use of gold behavior data.
Gabriella Kazai, Imed Zitouni
WSDM2
2016 Detecting Good Abandonment in Mobile Search
abstract
Web search queries for which there are no clicks are referred to as abandoned queries and are usually considered as leading to user dissatisfaction. However, there are many cases where a user may not click on any search result page (SERP) but still be satisfied. This scenario is referred to as good abandonment and presents a challenge for most approaches measuring search satisfaction, which are usually based on clicks and dwell time. The problem is exacerbated further on mobile devices where search providers try to increase the likelihood of users being satisfied directly by the SERP. This paper proposes a solution to this problem using gesture interactions, such as reading times and touch actions, as signals for differentiating between good and bad abandonment. These signals go beyond clicks and characterize user behavior in cases where clicks are not needed to achieve satisfaction. We study different good abandonment scenarios and investigate the different elements on a SERP that may lead to good abandonment. We also present an analysis of the correlation between user gesture features and satisfaction. Finally, we use this analysis to build models to automatically identify good abandonment in mobile search achieving an accuracy of 75%, which is significantly better than considering query and session signals alone. Our findings have implications for the study and application of user satisfaction in search systems.
Kyle Williams 0001, Julia Kiseleva, Aidan C. Crook, Imed Zitouni, Ahmed Awadallah 0001, Madian Khabsa
WWW4
2015 Toward Predicting the Outcome of an A/B Experiment for Search Relevance
abstract
A standard approach to estimating online click-based metrics of a ranking function is to run it in a controlled experiment on live users. While reliable and popular in practice, configuring and running an online experiment is cumbersome and time-intensive. In this work, inspired by recent successes of offline evaluation techniques for recommender systems, we study an alternative that uses historical search log to reliably predict online click-based metrics of a \emph{new} ranking function, without actually running it on live users. To tackle novel challenges encountered in Web search, variations of the basic techniques are proposed. The first is to take advantage of diversified behavior of a search engine over a long period of time to simulate randomized data collection, so that our approach can be used at very low cost. The second is to replace exact matching (of recommended items in previous work) by \emph{fuzzy} matching (of search result pages) to increase data efficiency, via a better trade-off of bias and variance. Extensive experimental results based on large-scale real search data from a major commercial search engine in the US market demonstrate our approach is promising and has potential for wide use in Web search.
Lihong Li 0001, Jin Young Kim 0005, Imed Zitouni
WSDM3
2015 Automatic Online Evaluation of Intelligent Assistants
abstract
Voice-activated intelligent assistants, such as Siri, Google Now, and Cortana, are prevalent on mobile devices. However, it is challenging to evaluate them due to the varied and evolving number of tasks supported, e.g., voice command, web search, and chat. Since each task may have its own procedure and a unique form of correct answers, it is expensive to evaluate each task individually. This paper is the first attempt to solve this challenge. We develop consistent and automatic approaches that can evaluate different tasks in voice-activated intelligent assistants. We use implicit feedback from users to predict whether users are satisfied with the intelligent assistant as well as its components, i.e., speech recognition and intent classification. Using this approach, we can potentially evaluate and compare different tasks within and across intelligent assistants ac-cording to the predicted user satisfaction rates. Our approach is characterized by an automatic scheme of categorizing user-system interaction into task-independent dialog actions, e.g., the user is commanding, selecting, or confirming an action. We use the action sequence in a session to predict user satisfaction and the quality of speech recognition and intent classification. We also incorporate other features to further improve our approach, including features derived from previous work on web search satisfaction prediction, and those utilizing acoustic characteristics of voice requests. We evaluate our approach using data collected from a user study. Results show our approach can accurately identify satisfactory and unsatisfactory sessions.
Jiepu Jiang, Ahmed Awadallah 0001, Rosie Jones, Umut Ozertem, Imed Zitouni, Ranjitha Gurunath Kulkarni, Omar Zia Khan
WWW5
2014 Machine-Assisted Search Preference Evaluation
abstract
Information Retrieval systems are traditionally evaluated using the relevance of web pages to individual queries. Other work on IR evaluation has focused on exploring the use of preference judgments over two search result lists. Unlike traditional query-document evaluation, collecting preference judgments over two search result-lists takes the context of documents, and hence takes the interaction between search results, into consideration. Moreover, preference judgments have been shown to produce more accurate results compared to absolute judgment. On the other hand result list preference judgments have very high annotation cost. In this work, we investigate how machine learned models can assist human judges in order to collect reliable result list preference judgments at large scale with lower judgment-cost. We build novel models that can predict user preference automatically. We investigate the effect of different features on the prediction quality. We focus on predicting preferences with high confidence and show that these models can be effectively used to assist human judges resulting in significant reduction in annotation cost.
Ahmed Awadallah 0001, Imed Zitouni
CIKM2
2014 Comparing client and server dwell time estimates for click-level satisfaction prediction
abstract
Click dwell time is the amount of time that a user spends on a clicked search result. Many previous studies have shown that click dwell time is strongly correlated with result-level satisfaction and document relevance. Accurate estimates of dwell time are therefore important for applications such as search satisfaction prediction and result ranking. However, dwell time can be estimated in different ways according to the information available about the search process. For example, a result reached for the query [Garfield] may involve 145s of "server-side" dwell time (observable to the search engine) and 40s of "client-side" dwell time (observable from the browser). Since search engines can only observe server-side actions (i.e., activity on the search engine result page), server-side dwell times are estimated by measuring the time between a search result click and the next search event (click or query). Conversely, more detailed information about page dwell times can be obtained via client-side methods such as Web browser toolbars. The client-side information enables the estimation of more accurate dwell times by measuring the amount of time that a user spends on pages of interest (either the landing page, or pages on the full navigation trail). In this paper, we define three different dwell times, i.e., server-side, client-side, and trail dwell time, and examine their effectiveness for predicting click satisfaction. For this, we collect toolbar and search engine logs from real users, and provide an analysis of dwell times for improving prediction performance. Moreover, we show further improvements in predicting click-level satisfaction by combining dwell times with other query features (e.g., query clarity).
Ahmed Awadallah 0001, Ryen W. White, Imed Zitouni
SIGIR4
2014 Modeling dwell time to predict click-level satisfaction
abstract
Clicks on search results are the most widely used behavioral signals for predicting search satisfaction. Even though clicks are correlated with satisfaction, they can also be noisy. Previous work has shown that clicks are affected by position bias, caption bias, and other factors. A popular heuristic for reducing this noise is to only consider clicks with long dwell time, usually equaling or exceeding 30 seconds. The rationale is that the more time a searcher spends on a page, the more likely they are to be satisfied with its contents. However, having a single threshold value assumes that users need a fixed amount of time to be satisfied with any result click, irrespective of the page chosen. In reality, clicked pages can differ significantly. Pages have different topics, readability levels, content lengths, etc. All of these factors may affect the amount of time spent by the user on the page. In this paper, we study the effect of different page characteristics on the time needed to achieve search satisfaction. We show that the topic of the page, its length and its readability level are critical in determining the amount of dwell time needed to predict whether any click is associated with satisfaction. We propose a method to model and provide a better understanding of click dwell time. We estimate click dwell time distributions for SAT (satisfied) or DSAT (dissatisfied) clicks for different click segments and use them to derive features to train a click-level satisfaction model. We compare the proposed model to baseline methods that use dwell time and other search performance predictors as features, and demonstrate that the proposed model achieves significant improvements.
Ahmed Awadallah 0001, Ryen W. White, Imed Zitouni
WSDM4
2014 Contextual and dimensional relevance judgments for reusable SERP-level evaluation
abstract
Document-level relevance judgments are a major component in the calculation of effectiveness metrics. Collecting high-quality judgments is therefore a critical step in information retrieval evaluation. However, the nature of and the assumptions underlying relevance judgment collection have not received much attention. In particular, relevance judgments are typically collected for each document in isolation, although users read each document in the context of other documents. In this work, we aim to investigate the nature of relevance judgment collection. We collect relevance labels in both isolated and conditional setting, and ask for judgments in various dimensions of relevance as well as overall relevance. Then we compare the relevance metrics based on various types of judgments with other metrics of quality such as user preference. Our analyses illuminate how these settings for judgment collection affect the quality and the characteristics of the judgments. We also find that the metrics based on conditional judgments show higher correlation with user preference than isolated judgments.
Peter B. Golbus, Imed Zitouni, Jin Young Kim 0005, Ahmed Awadallah 0001, Fernando Diaz 0001
WWW2
2014 Aligned-Parallel-Corpora Based Semi-Supervised Learning for Arabic Mention Detection
abstract
In the last two decades, significant effort has been put into annotating linguistic resources in several languages. Despite this valiant effort, there are still many languages left that have only small amounts of such resources. The goal of this article is to present and investigate a method of propagating information (specifically mentions) from a resource-rich language such as English into a relatively less-resource language such as Arabic. We compare also this approach to its equivalent counterpart using monolingual resources. Part of the investigation is to quantify the contribution of propagating information in different conditions - based on the availability of resources in the target language. Experiments on the language pair Arabic-English show that one can achieve relatively decent performance by propagating information from a language with richer resources such as English into Arabic alone (no resources or models in the source language Arabic). Furthermore, results show that propagated features from English do help improve the Arabic system performance even when used in conjunction with all feature types built from the source language. Experiments also show that using propagated features in conjunction with lexically-derived features only (as can be obtained directly from a mention annotated corpus) brings the system performance at the one obtained in the target language by using feature derived from many linguistic resources, therefore improving the system when such resources are not available.
Imed Zitouni, Yassine Benajiba
IEEE ACM Trans. Audio Speech Lang. Process.1
2013 Relevance dimensions in preference-based IR evaluation
abstract
Evaluation of information retrieval (IR) systems has recently been exploring the use of preference judgments over two search result lists. Unlike the traditional method of collecting relevance labels per single result, this method allows to consider the interaction between search results as part of the judging criteria. For example, one result list may be preferred over another if it has a more diverse set of relevant results, covering a wider range of user intents. In this paper, we investigate how assessors determine their preference for one list of results over another with the aim to understand the role of various relevance dimensions in preference-based evaluation. We run a series of experiments and collect preference judgments over different relevance dimensions in side-by-side comparisons of two search result lists, as well as relevance judgments for the individual documents. Our analysis of the collected judgments reveals that preference judgments combine multiple dimensions of relevance that go beyond the traditional notion of relevance centered on topicality. Measuring performance based on single document judgments and NDCG aligns well with topicality based preferences, but shows misalignment with judges' overall preferences, largely due to the diversity dimension. As a judging method, dimensional preference judging is found to lead to improved judgment quality.
Jin Young Kim 0005, Gabriella Kazai, Imed Zitouni
SIGIR3
2011 Introduction to Arabic Natural Language Processing Nizar Y. Habash (Columbia University) Morgan & Claypool (Synthesis Lectures on Human Language Technologies, edited by Graeme Hirst, volume 10), 2010, xvii+167 pp; paperbound, ISBN 978-1-59829-795-9, $40.00; ebook, ISBN 978-1-59829-796-6, $30.00 or by subscription
Imed Zitouni
Comput. Linguistics1
2010 Enhancing Mention Detection Using Projection via Aligned Corpora
Yassine Benajiba, Imed Zitouni
EMNLP2
2010 Improving Mention Detection Robustness to Noisy Input
Radu Florian, John F. Pitrelli, Salim Roukos, Imed Zitouni
EMNLP4
2010 Morphological and syntactic features for Arabic speech recognition
abstract
In this paper, we study the use of morphological and syntactic context features to improve speech recognition of a morphologically rich language like Arabic. We examine a variety of syntactic features, including part-of-speech tags, shallow parse tags, and exposed head words and their non-terminal labels both before and after the word to be predicted. Neural network LMs are used to model these features since they generalize better to unseen events by modeling words and other context features in continuous space. Using morphological and syntactic features, we can improve the word error rate (WER) significantly on various test sets, including EVAL'08U, the unsequestered portion of the DARPA GALE Phase 3 evaluation test set.
Hong-Kwang Jeff Kuo, Lidia Mangu, Ahmad Emami, Imed Zitouni
ICASSP4
2010 Augmented context features for Arabic speech recognition
Ahmad Emami, Hong-Kwang Jeff Kuo, Imed Zitouni, Lidia Mangu
INTERSPEECH3
2010 Arabic Word Segmentation for Better Unit of Analysis
Yassine Benajiba, Imed Zitouni
LREC2
2010 Arabic Mention Detection: Toward Better Unit of Analysis
Yassine Benajiba, Imed Zitouni
HLT-NAACL2
2009 Syntactic features for Arabic speech recognition
abstract
We report word error rate improvements with syntactic features using a neural probabilistic language model through N-best re-scoring. The syntactic features we use include exposed head words and their non-terminal labels both before and after the predicted word. Neural network LMs generalize better to unseen events by modeling words and other context features in continuous space. They are suitable for incorporating many different types of features, including syntactic features, where there is no pre-defined back-off order. We choose an N-best re-scoring framework to be able to take full advantage of the complete parse tree of the entire sentence. Using syntactic features, along with morphological features, improves the word error rate (WER) by up to 5.5% relative, from 9.4% to 8.6%, on the latest GALE evaluation test set.
Hong-Kwang Jeff Kuo, Lidia Mangu, Ahmad Emami, Imed Zitouni, Young-Suk Lee 0001
ASRU4
2009 Arabic diacritic restoration approach based on maximum entropy models
Imed Zitouni, Ruhi Sarikaya
Comput. Speech Lang.1
2009 Morphology-Based Segmentation Combination for Arabic Mention Detection
abstract
The Arabic language has a very rich/complex morphology. Each Arabic word is composed of zero or more prefixes , one stem and zero or more suffixes . Consequently, the Arabic data is sparse compared to other languages such as English, and it is necessary to conduct word segmentation before any natural language processing task. Therefore, the word-segmentation step is worth a deeper study since it is a preprocessing step which shall have a significant impact on all the steps coming afterward. In this article, we present an Arabic mention detection system that has very competitive results in the recent Automatic Content Extraction (ACE) evaluation campaign. We investigate the impact of different segmentation schemes on Arabic mention detection systems and we show how these systems may benefit from more than one segmentation scheme. We report the performance of several mention detection models using different kinds of possible and known segmentation schemes for Arabic text: punctuation separation, Arabic Treebank, and morphological and character-level segmentations. We show that the combination of competitive segmentation styles leads to a better performance. Results indicate a statistically significant improvement when Arabic Treebank and morphological segmentations are combined.
Yassine Benajiba, Imed Zitouni
ACM Trans. Asian Lang. Inf. Process.2
2009 Cross-Language Information Propagation for Arabic Mention Detection
abstract
In the last two decades, significant effort has been put into annotating linguistic resources in several languages. Despite this valiant effort, there are still many languages left that have only small amounts of such resources. The goal of this article is to present and investigate a method of propagating information (specifically mention detection) from a resource-rich language into a relatively resource-poor language such as Arabic. Part of the investigation is to quantify the contribution of propagating information in different conditions based on the availability of resources in the target language. Experiments on the language pair Arabic-English show that one can achieve relatively decent performance by propagating information from a language with richer resources such as English into Arabic alone (no resources or models in the source language Arabic). Furthermore, results show that propagated features from English do help improve the Arabic system performance even when used in conjunction with all feature types built from the source language. Experiments also show that using propagated features in conjunction with lexically derived features only (as can be obtained directly from a mention annotated corpus) brings the system performance at the one obtained in the target language by using feature derived from many linguistic resources, therefore improving the system when such resources are not available. In addition to Arabic-English language pair, we investigate the effectiveness of our approach on other language pairs such as Chinese-English and Spanish-English.
Imed Zitouni, Radu Florian
ACM Trans. Asian Lang. Inf. Process.1
2009 A Cascaded Approach to Mention Detection and Chaining in Arabic
abstract
This paper presents a fully statistical approach to Arabic mention detection and chaining system, built around the maximum entropy principle. The presented system takes a cascade approach to processing an input document, by first detecting mentions in the document and then chaining the identified mentions into entities. Both system components use a common maximum entropy framework, which allows the integration of a large array of feature types, including lexical, morphological, syntactic, and semantic features. Arabic offers additional challenges for this task (when compared with English, for example), as segmentation is a needed processing step, so one can correctly identify and resolve enclitic pronouns. The system presented has obtained very competitive performance in the automatic content extraction (ACE) evaluation program.
Imed Zitouni, Xiaoqiang Luo, Radu Florian
IEEE Trans. Speech Audio Process.1
2008 When Harry Met Harri: Cross-lingual Name Spelling Normalization
Fei Huang 0002, Ahmad Emami, Imed Zitouni
EMNLP3
2008 Mention Detection Crossing the Language Barrier
Imed Zitouni, Radu Florian
EMNLP1
2008 Hierarchical linear discounting class N-gram language models: A multilevel class hierarchy approach
abstract
We introduce in this paper a hierarchical linear discounting class n-gram language modeling technique that has the advantage of combining several language models, trained at different nodes in a class hierarchy. The approach hierarchically clusters the word vocabulary into a word-tree. The closer a tree node is to the leaves, the more specific the corresponding word class is. The tree is used to balance generalization ability and word specificity when estimating the likelihood of an n-gram event. Experiments are conducted on Wall Street Journal corpus using a vocabulary of 20,000 words. Results show a reduction on the test perplexity over the standard n-gram approaches by 10%. We also report considerable improvement in the accuracy of the speech recognition task.
Imed Zitouni, Qiru Zhou
ICASSP1
2008 Rich morphology based n-gram language models for Arabic
Ahmad Emami, Imed Zitouni, Lidia Mangu
INTERSPEECH2
2008 Constrained Minimization and Discriminative Training for Natural Language Call Routing
abstract
This paper presents a combination strategy of multiple individual routing classifiers to improve classification accuracy in natural language call routing applications. Since errors of individual classifiers in the ensemble should somehow be uncorrelated, we propose a combination strategy where the combined classifier accuracy is a function of the accuracy of individual classifiers and also the correlation between their classification errors. We show theoretically and empirically that our combination strategy, named the constrained minimization technique, has a good potential in improving the classification accuracy of single classifiers. We also show how discriminative training, more specifically the generalized probabilistic descent (GPD) algorithm, can be of benefit to further boost the performance of routing classifiers. The GPD algorithm has the potential to consider both positive and negative examples during training to minimize the classification error and increase the score separation of the correct from competing hypotheses. Some parameters become negative when using the GPD algorithm, resulting from suppressive learning not traditionally possible; important antifeatures are thus obtained. Experimental evaluation is carried on a banking call routing task and on switchboard databases with a set of 23 and 67 destinations, respectively. Results show either the GPD or constrained minimization technique outperform the accuracy of baseline classifiers by 44% when applied separately. When the constrained minimization technique is added on top of GPD, we show an additional 15% reduction in the classification error rate.
Imed Zitouni
IEEE Trans. Speech Audio Process.1
2007 Backoff hierarchical class n-gram language models: effectiveness to model unseen events in speech recognition
Imed Zitouni
Comput. Speech Lang.1
2006 Factorizing Complex Models: A Case Study in Mention Detection
abstract
As natural language understanding research advances towards deeper knowledge modeling, the tasks become more and more complex: we are interested in more nuanced word characteristics, more linguistic properties, deeper semantic and syntactic features. One such example, explored in this article, is the mention detection and recognition task in the Automatic Content Extraction project, with the goal of identifying named, nominal or pronominal references to real-world entities---mentions---and labeling them with three types of information: entity type, entity subtype and mention type. In this article, we investigate three methods of assigning these related tags and compare them on several data sets. A system based on the methods presented in this article participated and ranked very competitively in the ACE'04 evaluation.
Radu Florian, Hongyan Jing, Nanda Kambhatla, Imed Zitouni
ACL4
2006 Maximum Entropy Based Restoration of Arabic Diacritics
abstract
Short vowels and other diacritics are not part of written Arabic scripts. Exceptions are made for important political and religious texts and in scripts for beginning students of Arabic. Script without diacritics have considerable ambiguity because many words with different diacritic patterns appear identical in a diacritic-less setting. We propose in this paper a maximum entropy approach for restoring diacritics in a document. The approach can easily integrate and make effective use of diverse types of information; the model we propose integrates a wide array of lexical, segment-based and part-of-speech tag features. The combination of these feature types leads to a state-of-the-art diacritization model. Using a publicly available corpus (LDC's Arabic Treebank Part 3), we achieve a diacritic error rate of 5.1%, a segment error rate 8.5%, and a word error rate of 17.3%. In case-ending-less setting, we obtain a diacritic error rate of 2.2%, a segment error rate 4.0%, and a word error rate of 7.2%.
Imed Zitouni, Jeffrey S. Sorensen, Ruhi Sarikaya
ACL1
2006 Maximum entropy modeling for diacritization of Arabic text
Ruhi Sarikaya, Ossama Emam, Imed Zitouni
INTERSPEECH3
2005 Discriminative training and support vector machine for natural language call routing
Imed Zitouni, Hui Jiang 0001, Qiru Zhou
INTERSPEECH1
2004 Prediction-based packet loss concealment for voice over IP: a statistical n-gram approach
abstract
We investigate the possibility of predicting lost packets for packet loss concealment using n-gram predictive models. Unlike the conventional repetition-based algorithms, the proposed algorithm is based on the Shannon game, which serves as a principle for predicting the speech parameters of lost packets using the previously received parameters. During the training phase, we construct statistical backoff n-gram models. In the test phase, the models are used to predict the speech parameters of lost packets. Experiments were performed on a switchboard telephone speech database and the proposed algorithm is compared with the conventional repetition-based algorithm. The performance is evaluated in terms of the spectral distortion between the original and the predicted (or repeated) speech. The algorithm based on the back-off n-gram models reduces the spectral distortion by 8.7% over the conventional repetition-based algorithm for the first lost packet after receiving one. Further, it maintains about 6.2% improvement for up to six consecutive lost packets. In terms of perplexity of the predictive models, the backoff n-gram approach outperforms the repetition-based algorithm by 8.65%, which is very close to the improvement rate obtained from the spectral distortion measurement.
Imed Zitouni, Qiru Zhou
GLOBECOM2
2004 Discriminative training of naive Bayes classifiers for natural language call routing
abstract
In this paper, we propose to use a discriminative training(DT) method to improve naive Bayes classifiers in context of natural language call routing. As opposed to the traditional maximum likelihood estimation, all conditional probabilties in Naive Bayes classifers (NBC) are estimated discriminatively based on the minimum classification error (MCE) criterion. A smoothed classification error rate in training set is formulated as an objective function and the GPD (generalized probabilistic descent) method is used to minimize the objective function with respect to all conditional probabilities in NBCs. Two versions of NBC are used in this work. In the first version all NBCs corresponding to various destinations use the same word feature set while destination-dependent feature set is chosen for each destination in the second version. Experimental results on a banking call routing task show that the discriminative training method can achieve up to about 31% error reduction over our best MLtrained system. The proposed formulation is applicable to other algorithms addressing a wide range of tasks, such as topic identification, information retrieval and speech understanding.
Hui Jiang 0001, Imed Zitouni
INTERSPEECH3
2004 On a n-gram model approach for packet loss concealment
abstract
In this paper, we investigate the possibility of predicting lost packets for packet loss concealment using n-gram predictive models. Unlike the conventional repetition-based algorithms, the proposed algorithm is based on the Shannon game, which serves as a principle for predicting the speech parameters of lost packets using the previously received parameters. During training phase, we construct statistical backoff n-gram models. In test phase, the models are used to predict the speech parameters of lost packets. Experiments were performed on switchboard telephone speech database and the proposed algorithm is compared with the conventional repetitionbased algorithm. The performance is evaluated in terms of the spectral distortion between the original and the predicted (or repeated) speech. The algorithm based on the backoff n-gram models reduces the spectral distortion by 8.7% over the conventional repetition-based algorithm for the first lost packet after receiving one. Further it maintains about 6.2% improvement up to six consecutive lost packets. In terms of perplexity of the predictive models, backoff n-gram approach outperforms the repetition-based algorithm by 8.65%, which is very close to the improvement rate obtained from the spectral distortion measurement.
Imed Zitouni, Qiru Zhou
INTERSPEECH2
2004 Constrained minimization technique for topic identification using discriminative training and support vector machines
abstract
This paper describes the constrained minimization approach to combine multiple classifiers in order to improve classification accuracy. Since errors of individual classifiers in the ensemble should somehow be uncorrelated to yield higher classification accuracy, we propose a combination strategy where the combined classifier accuracy is a function of the correlation between classification errors of the individual classifiers. To obtain powerful single classifiers, different techniques are investigated including support vector machines and latent semantic indexing (LSI) matrix, which is a popular vector-space model. We also investigate discriminative training (DT) of the LSI matrix on constrained minimization approach. DT minimizes the classification error by increasing the score separation of the correct from competing documents. Experimental evaluation is carried out on a banking call routing and on switchboard databases with a set of 23 and 67 topics respectively. Results show that the combined classifier we propose outperforms the accuracy of individual baseline classifiers by 44%.
Imed Zitouni, Hui Jiang 0001
INTERSPEECH1
2004 OrienTel - Telephony Databases Across Northern Africa and the Middle East
Dorota J. Iskra, Rainer Siemund, Jamal Borno, Asunción Moreno, Ossama Emam, Khalid Choukri, Oren Gedge, Herbert S. Tropf, Albino Nogueiras, Imed Zitouni, Anastasios Tsopanoglou, Nikos Fakotakis
LREC10
2003 Minimum verification error training for topic verification
abstract
We propose a new formulation of minimum verification error training and apply it to the problem of topic verification as an example. In topic verification, a decision is made as to whether a document truly belongs to a particular topic of interest. Such a decision typically depends on a comparison between a model for the desired topic and a model for background topics, using a decision threshold. We propose modeling the background topics as a cohort model consisting of a weighted combination of the M closest topics discovered from the training data. The weights and the decision threshold are optimized using the generalized probabilistic descent algorithm to explicitly minimize the verification error rate, which is defined to be a weighted sum of the Type I (false rejection) and Type II (false acceptance) errors.
Hong-Kwang Jeff Kuo, Imed Zitouni, Eric Fosler-Lussier
ICASSP (1)3
2003 Hierarchical class n-gram language models: towards better estimation of unseen events in speech recognition
Imed Zitouni, Olivier Siohan
INTERSPEECH1
2003 Statistical language modeling based on variable-length sequences
Imed Zitouni, Kamel Smaïli, Jean Paul Haton
Comput. Speech Lang.1
2003 Boosting and combination of classifiers for natural language call routing systems
Imed Zitouni, Hong-Kwang Jeff Kuo
Speech Commun.1
2002 Adaptive language models for spoken dialogue systems
abstract
In this paper, we investigate both generative and statistical approaches for language modeling in spoken dialogue systems. Semantic class-based finite state and n-gram grammars are used for improving coverage and modeling accuracy when little training data is available. We have implemented dialogue-state specific language model adaptation to reduce perplexity and improve the efficiency of grammars for spoken dialogue systems. A novel algorithm for combining state-independent n-gram and state-dependent finite state grammars using acoustic confidence scores is proposed. Using this combination strategy, a relative word error reduction of 12% is achieved for certain dialogue states within a travel reservation task. Finally, semantic class multigrams are proposed and briefly evaluated for language modeling in dialogue systems.
Roger Argiles Solsona, Eric Fosler-Lussier, Hong-Kwang Jeff Kuo, Alexandros Potamianos, Imed Zitouni
ICASSP5
2002 Combination of boosting and discriminative training for natural language call steering systems
abstract
In this paper, we describe the combination of two different techniques to improve natural language call routing: boosting and discriminative training. The goal of boosting is to re-weight the data in order to train a set of classifiers whose errors may be uncorrelated so that when combined, the classification error rate (CER) can be reduced. We propose using discriminative training to improve the individual classifier accuracy at each iteration of the boosting algorithm. Compared to the baseline classifiers, an improvement in the CER of 41–50% was observed on call routing for a banking task. More importantly, synergistic effects of discriminative training on the boosting algorithm were demonstrated: more iterations were possible because discriminative training reduced the CER of individual classifiers trained on re-weighted data by an average of 72%.
Imed Zitouni, Hong-Kwang Jeff Kuo
ICASSP1
2002 Discriminative training for call classification and routing
Hong-Kwang Jeff Kuo, Imed Zitouni, Eric Fosler-Lussier, Egbert Ammicht
INTERSPEECH3
2002 Orientel: speech-based interactive communication applications for the mediterranean and the middle east
abstract
In this paper, we introduce a new European project named OrienTel. The aim of OrienTel is to enable the project's participants to design and develop multilingual interactive communication services for the Mediterranean and the Middle East, ranging from Morocco in the West to the Gulf states in the East, including Turkey and Cyprus. These multilingual applications will be largely speech-based and will typically be implemented on mobile and multi-modal platforms such as cellular GSM or UMTS phones, personal digital assistants (PDAs) or combinations of the two. Applications of the kind targeted are unified messaging, information retrieval, customer care, banking, WAP and service portals. To achieve this aim, the consortium will produce various surveys of the OrienTel region, compile a set of 22 linguistic databases, conduct research into ASR-related problems the OrienTel languages hold and develop demonstrator applications bearing evidence of OrienTel's multilingual orientation.
Imed Zitouni, Joseph P. Olive, Dorota J. Iskra, Khalid Choukri, Ossama Emam, Oren Gedge, Emmanuel Maragoudakis, Herbert S. Tropf, Asunción Moreno, Albino Nogueiras, Barbara Heuft, Rainer Siemund
INTERSPEECH1
2002 Backoff hierarchical class n-gram language modelling for automatic speech recognition systems
Imed Zitouni, Olivier Siohan, Hong-Kwang Jeff Kuo
INTERSPEECH1
2002 OrienTel - Multilingual access to interactive communication services for the Mediterranean and the Middle East
Rainer Siemund, Barbara Heuft, Khalid Choukri, Ossama Emam, Emmanuel Maragoudakis, Herbert S. Tropf, Oren Gedge, Sherrie Shammass, Asunción Moreno, Albino Nogueiras, Imed Zitouni, Dorota J. Iskra
LREC11
2002 A hierarchical language model based on variable-length class sequences: the MCnnu approach
abstract
We propose a new language model which represents long-term dependencies between word sequences using a multilevel hierarchy. We call this model MC/sub n//sup /spl nu//, where n is the maximum number of words in a sequence and /spl nu/ is the maximum number of levels. The originality of this model, which is an extension of the multigrams, is its ability to take into account long distance dependencies according to dependent variable-length sequences. In order to discover the variable-length sequences and to build the hierarchy, we use a set of 233 syntactic classes extracted from eight elementary grammatical classes of French. The MC/sub n//sup /spl nu// model learns hierarchical word patterns and uses them to reevaluate and filter the n-best utterance hypotheses output by our speech recognizer MAUD. The model has been trained on a corpus of 43 million words extracted from the French newspaper "Le Monde" and uses a vocabulary of 20 000 words. Tests have been conducted on 300 sentences. Compared to the class trigram and the baseline multigrams approach, we report a perplexity reduction of 17% and 20%, respectively. Rescoring the original n-best hypotheses resulted in an improvement of the word error rate: 7% and 2% compared to the class trigram and multigrams, respectively.
Imed Zitouni
IEEE Trans. Speech Audio Process.1
2001 Statistical language model based on a hierarchical approach: MCnv
abstract
Colloque avec actes et comité de lecture. internationale.
Imed Zitouni, Kamel Smaïli, Jean Paul Haton
INTERSPEECH1
2001 A Comparative Study of Topic Identification on Newspaper and E-mail
abstract
This work presents several statistical methods for topic identification on two kinds of textual data: newspaper articles and e-mails. Five methods are tested on these two corpora: topic unigrams, cache model, TFIDF classijier, topic peqdexity, and weighted model. Our work aims to study these methods by confronting them to very diferent data. This study is very fruitful for our research. Statistical topic identiJication methods depend not only on a corpus, but also on its type. One of the methods achieves a topic identiJcation of 80% on a general newspaper corpus but does not exceed 30% on e-mail corpus. Another method gives the best result on e-mails, but has not the same behavior on a newspaper corpus. We also show in this paper that almost all our methods achieve good results in retrieving the first two manually annotated labels.
Brigitte Bigi, Armelle Brun, Jean Paul Haton, Kamel Smaïli, Imed Zitouni
SPIRE5
2001 An alternative scheme for perplexity estimation and its assessment for the evaluation of language models
Frédéric Bimbot, Marc El-Bèze, Stéphane Igounet, Michèle Jardino, Kamel Smaïli, Imed Zitouni
Comput. Speech Lang.6
2000 Beyond the conventional statistical language models: the variable-length sequences approach
abstract
Colloque avec actes et comité de lecture. internationale.
Imed Zitouni, Kamel Smaïli, Jean Paul Haton
INTERSPEECH1
1999 Automatic and manual clustering for large vocabulary speech recognition: a comparative study
abstract
An ubiquitous phenomenon in psychology is the `repetition effect': a repeated stimulus is processed better on the second occurrence than on the first. Yet, what counts as a repetition? When a spoken word is repeated, is it the acoustic shape or the linguistic type that matters? In the present study, we contrasted the contribution of acoustic and phonological features by using participants with different linguistic backgrounds: they came from two populations sharing a common vocabulary (Catalan) yet possessing different phonemic systems. They performed a lexical decision task with lists containing words that were repeated verbatim, as well as words that were repeated with one phonetic feature changed. The feature changes were phonemic, i.e. linguistically relevant, for one population, but not for the other. The results revealed that the repetition effect was modulated by linguistic, not acoustic, similarity: it depended on the subjects' phonemic system. \n
Kamel Smaïli, Armelle Brun, Imed Zitouni, Jean Paul Haton
EUROSPEECH3
1999 Variable-length sequence language model for large vocabulary continuous dictation machine
abstract
Colloque avec actes et comité de lecture.
Imed Zitouni, Jean-François Mari, Kamel Smaïli, Jean Paul Haton
EUROSPEECH1
1998 A language modeling based on a hierarchical approach: m_n^v
abstract
In this work, we introduce the concept of hierarchical M n language model and we compare it to the based class multigram and interpolated class n-gram model. The originality of our approach is its capability to parse a string of class/tags into variable length dependent sequences. A few experimental tests were carried out on a class corpus extracted from the French “Le Monde” word corpus labeled automatically. In our experiments, M n outperforms based class multigram and interpolated class bigram but are comparable to the interpolated class trigram model.
Imed Zitouni
ICSLP1
1998 A comparative study between polyclass and multiclass language models
abstract
International audience
Imed Zitouni, Kamel Smaïli, Jean Paul Haton, Sabine Deligne, Frédéric Bimbot
ICSLP1
1998 A first evaluation campaign for language models
Michèle Jardino, Frédéric Bimbot, Stéphane Igounet, Kamel Smaïli, Imed Zitouni, Marc El-Bèze
LREC5
1997 An hybrid language model for a continuous dictation prototype
abstract
International audience
Kamel Smaïli, Imed Zitouni, François Charpillet, Jean Paul Haton
EUROSPEECH2