Salim Roukos

dblp:01/1417 · DBLP profile ↗
← Back
75ranked-venue papers
2as first author
14since 2021 · last 2025
0000-0003-2140-4349ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 57 · 2 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 3 since 2021Databases, data management, data science and information retrieval · 6 · 2 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Automated Single-Turn Solution Recommendation System for Software IT Support Tickets
Paulina Toro Isaza, Michael Nidd, Noah Zheutlin, Jae-wook Ahn, Chidansh Amitkumar Bhatt, Yu Deng 0004, Ruchi Mahindru, Martin Franz, Hans Florian, Salim Roukos
IEEE Big Data10
2025 From Multiple-Choice to Extractive QA: A Case Study for English and Arabic
abstract
The rapid evolution of Natural Language Processing (NLP) has favoured major languages such as English, leaving a significant gap for many others due to limited resources. This is especially evident in the context of data annotation, a task whose importance cannot be underestimated, but which is time-consuming and costly. Thus, any dataset for resource-poor languages is precious, in particular when it is task-specific. Here, we explore the feasibility of repurposing an existing multilingual dataset for a new NLP task: we repurpose a subset of the BELEBELE dataset (Bandarkar et al., 2023), which was designed for multiple-choice question answering (MCQA), to enable the more practical task of extractive QA (EQA) in the style of machine reading comprehension. We present annotation guidelines and a parallel EQA dataset for English and Modern Standard Arabic (MSA). We also present QA evaluation results for several monolingual and cross-lingual QA pairs including English, MSA, and five Arabic dialects. We aim to help others adapt our approach for the remaining 120 BELEBELE language variants, many of which are deemed under-resourced. We also provide a thorough analysis and share insights to deepen understanding of the challenges and opportunities in NLP task reformulation.
Teresa Lynn, Malik H. Altakrori, Samar Mohamed Magdy, Rocktim Jyoti Das, Chenyang Lyu, Mohamed Nasr, Younes Samih, Kirill Chirkunov, Alham Fikri Aji, Preslav Nakov, Shantanu Godbole, Salim Roukos, Radu Florian, Nizar Habash
COLING12
2025 CLAPnq: Cohesive Long-form Answers from Passages in Natural Questions for RAG systems
abstract
Abstract Retrieval Augmented Generation (RAG) has become a popular application for large language models. It is preferable that successful RAG systems provide accurate answers that are supported by being grounded in a passage without any hallucinations. While considerable work is required for building a full RAG pipeline, being able to benchmark performance is also necessary. We present CLAPnq, a benchmark Long-form Question Answering dataset for the full RAG pipeline. CLAPnq includes long answers with grounded gold passages from Natural Questions (NQ) and a corpus to perform either retrieval, generation, or the full RAG pipeline. The CLAPnq answers are concise, 3x smaller than the full passage, and cohesive, meaning that the answer is composed fluently, often by integrating multiple pieces of the passage that are not contiguous. RAG models must adapt to these properties to be successful at CLAPnq. We present baseline experiments and analysis for CLAPnq that highlight areas where there is still significant room for improvement in grounded RAG. CLAPnq is publicly available at https://github.com/primeqa/clapnq.
Sara Rosenthal, Avirup Sil, Radu Florian, Salim Roukos
Trans. Assoc. Comput. Linguistics4
2024 CHRONOS: A Schema-Based Event Understanding and Prediction System
abstract
Chronological and Hierarchical Reasoning Over Naturally Occurring Schemas (CHRONOS) is a system that combines language model-based natural language processing with symbolic knowledge representations to analyze and make predictions about newsworthy events. CHRONOS consists of an event-centric information extraction pipeline and a complex event schema instantiation and prediction system. Resulting predictions are detailed with arguments, event types from Wikidata, schema-based justifications, and source document provenance. We evaluate our system by its ability to capture the structure of unseen events described in news articles and make plausible predictions as judged by human annotators.
Maria Chang 0001, Achille Fokoue, Rosario Uceda-Sosa, Parul Awasthy, Ken Barker 0002, Sadhana Kumaravel, Oktie Hassanzadeh, Elton F. S. Soares, Debarun Bhattacharjya, Radu Florian, Salim Roukos
AAAI12
2024 Graph-based Uncertainty Metrics for Long-form Language Model Generations
abstract
Recent advancements in Large Language Models (LLMs) have significantly improved text generation capabilities, but these systems are still known to hallucinate, and granular uncertainty estimation for long-form LLM generations remains challenging. In this work, we propose Graph Uncertainty -- which represents the relationship between LLM generations and claims within them as a bipartite graph and estimates the claim-level uncertainty with a family of graph centrality metrics. Under this view, existing uncertainty estimation methods based on the concept of self-consistency can be viewed as using degree centrality as an uncertainty measure, and we show that more sophisticated alternatives such as closeness centrality provide consistent gains at claim-level uncertainty estimation. Moreover, we present uncertainty-aware decoding techniques that leverage both the graph structure and uncertainty estimates to improve the factuality of LLM generations by preserving only the most reliable claims. Compared to existing methods, our graph-based uncertainty metrics lead to an average of 6.8% relative gains on AUPRC across various long-form generation settings, and our end-to-end system provides consistent 2-4% gains in factuality over existing decoding techniques while significantly improving the informativeness of generated responses.
Mingjian Jiang, Yangjun Ruan, Prasanna Sattigeri, Salim Roukos, Tatsunori B. Hashimoto
NeurIPS4
2023 UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers
abstract
Jon Saad-Falcon, Omar Khattab, Keshav Santhanam, Radu Florian, Martin Franz, Salim Roukos, Avirup Sil, Md Sultan, Christopher Potts. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Jon Saad-Falcon, Omar Khattab, Keshav Santhanam, Radu Florian, Martin Franz, Salim Roukos, Avirup Sil, Md. Arafat Sultan, Christopher Potts
EMNLP6
2022 Logical Neural Networks for Knowledge Base Completion with Embeddings & Rules
abstract
Prithviraj Sen, Breno William Carvalho, Ibrahim Abdelaziz, Pavan Kapanipathi, Salim Roukos, Alexander Gray. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Prithviraj Sen, Breno W. Carvalho, Ibrahim Abdelaziz, Pavan Kapanipathi, Salim Roukos, Alexander G. Gray
EMNLP5
2022 Maximum Bayes Smatch Ensemble Distillation for AMR Parsing
abstract
Young-Suk Lee, Ramón Astudillo, Hoang Thanh Lam, Tahira Naseem, Radu Florian, Salim Roukos. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Young-Suk Lee 0001, Ramón Fernandez Astudillo, Hoang Thanh Lam, Tahira Naseem, Radu Florian, Salim Roukos
NAACL-HLT6
2022 DocAMR: Multi-Sentence AMR Representation and Evaluation
abstract
Tahira Naseem, Austin Blodgett, Sadhana Kumaravel, Tim O’Gorman, Young-Suk Lee, Jeffrey Flanigan, Ramón Astudillo, Radu Florian, Salim Roukos, Nathan Schneider. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Tahira Naseem, Austin Blodgett, Sadhana Kumaravel, Tim O'Gorman, Young-Suk Lee 0001, Jeffrey Flanigan, Ramón Fernandez Astudillo, Radu Florian, Salim Roukos, Nathan Schneider 0001
NAACL-HLT9
2021 A Semantic Parsing and Reasoning-Based Approach to Knowledge Base Question Answering
abstract
Knowledge Base Question Answering (KBQA) is a task where existing techniques have faced significant challenges, such as the need for complex question understanding, reasoning, and large training datasets. In this work, we demonstrate Deep Thinking Question Answering (DTQA), a semantic parsing and reasoning-based KBQA system. DTQA (1) integrates multiple, reusable modules that are trained specifically for their individual tasks (e.g. semantic parsing, entity linking, and relationship linking), eliminating the need for end-to-end KBQA training data; (2) leverages semantic parsing and a reasoner for improved question understanding. DTQA is a system of systems that achieves state-of-the-art performance on two popular KBQA datasets.
Ibrahim Abdelaziz, Srinivas Ravishankar, Pavan Kapanipathi, Salim Roukos, Alexander G. Gray
AAAI4
2021 KAAPA: Knowledge Aware Answers from PDF Analysis
abstract
We present KaaPa (Knowledge Aware Answers from Pdf Analysis), an integrated solution for machine reading comprehension over both text and tables extracted from PDFs. KaaPa enables interactive question refinement using facets generated from an automatically induced Knowledge Graph. In addition it provides a concise summary of the supporting evidence for the provided answers by aggregating information across multiple sources. KaaPa can be applied consistently to any collection of documents in English with zero domain adaptation effort. We showcase the use of KaaPa for QA on scientific literature using the COVID-19 Open Research Dataset.
Nicolas R. Fauceglia, Mustafa Canim, Alfio Massimiliano Gliozzo, Jennifer J. Liang, Nancy Xin Ru Wang, Douglas Burdick, Nandana Mihindukulasooriya, Vittorio Castelli, Guy Feigenblat, David Konopnicki, Yannis Katsis, Radu Florian, Yunyao Li 0001, Salim Roukos, Avirup Sil
AAAI14
2021 Bootstrapping Multilingual AMR with Contextual Word Alignments
abstract
Janaki Sheth, Young-Suk Lee, Ramón Fernandez Astudillo, Tahira Naseem, Radu Florian, Salim Roukos, Todd Ward. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Janaki Sheth, Young-Suk Lee 0001, Ramón Fernandez Astudillo, Tahira Naseem, Radu Florian, Salim Roukos, Todd Ward
EACL6
2021 Structure-aware Fine-tuning of Sequence-to-sequence Transformers for Transition-based AMR Parsing
abstract
Predicting linearized Abstract Meaning Representation (AMR) graphs using pre-trained sequence-to-sequence Transformer models has recently led to large improvements on AMR parsing benchmarks.These parsers are simple and avoid explicit modeling of structure but lack desirable properties such as graph well-formedness guarantees or built-in graph-sentence alignments.In this work we explore the integration of general pre-trained sequence-to-sequence language models and a structure-aware transition-based approach.We depart from a pointer-based transition system and propose a simplified transition set, designed to better exploit pre-trained language models for structured fine-tuning.We also explore modeling the parser state within the pre-trained encoder-decoder architecture and different vocabulary strategies for the same purpose.We provide a detailed comparison with recent progress in AMR parsing and show that the proposed parser retains the desirable properties of previous transition-based approaches, while being simpler and reaching the new parsing state of the art for AMR 2.0, without the need for graph re-categorization.
Jiawei Zhou 0001, Tahira Naseem, Ramón Fernandez Astudillo, Young-Suk Lee 0001, Radu Florian, Salim Roukos
EMNLP (1)6
2021 Synthetic Target Domain Supervision for Open Retrieval QA
abstract
Neural passage retrieval is a new and promising approach in open retrieval question answering. In this work, we stress-test the Dense Passage Retriever (DPR)---a state-of-the-art (SOTA) open domain neural retrieval model---on closed and specialized target domains such as COVID-19, and find that it lags behind standard BM25 in this important real-world setting. To make DPR more robust under domain shift, we explore its fine-tuning with synthetic training examples, which we generate from unlabeled target domain text using a text-to-text generator. In our experiments, this noisy but fully automated target domain supervision gives DPR a sizable advantage over BM25 in out-of-domain settings, making it a more viable model in practice. Finally, an ensemble of BM25 and our improved DPR model yields the best results, further pushing the SOTA for open retrieval QA on multiple out-of-domain test sets.
Revanth Gangi Reddy, Bhavani Iyer, Md. Arafat Sultan, Rong Zhang 0010, Avirup Sil, Vittorio Castelli, Radu Florian, Salim Roukos
SIGIR8
2020 The TechQA Dataset
abstract
Vittorio Castelli, Rishav Chakravarti, Saswati Dana, Anthony Ferritto, Radu Florian, Martin Franz, Dinesh Garg, Dinesh Khandelwal, Scott McCarley, Michael McCawley, Mohamed Nasr, Lin Pan, Cezar Pendus, John Pitrelli, Saurabh Pujar, Salim Roukos, Andrzej Sakrajda, Avi Sil, Rosario Uceda-Sosa, Todd Ward, Rong Zhang. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Vittorio Castelli, Rishav Chakravarti, Saswati Dana, Anthony Ferritto, Radu Florian, Martin Franz, Dinesh Garg, Dinesh Khandelwal, J. Scott McCarley, Mike McCawley, Mohamed Nasr, Lin Pan 0003, Cezar Pendus, John F. Pitrelli, Saurabh Pujar, Salim Roukos, Andrej Sakrajda, Avirup Sil, Rosario Uceda-Sosa, Todd Ward, Rong Zhang 0010
ACL16
2020 GPT-too: A Language-Model-First Approach for AMR-to-Text Generation
abstract
Manuel Mager, Ramón Fernandez Astudillo, Tahira Naseem, Md Arafat Sultan, Young-Suk Lee, Radu Florian, Salim Roukos. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Manuel Mager, Ramón Fernandez Astudillo, Tahira Naseem, Md. Arafat Sultan, Young-Suk Lee 0001, Radu Florian, Salim Roukos
ACL7
2020 Multi-Stage Pre-training for Low-Resource Domain Adaptation
abstract
Rong Zhang, Revanth Gangi Reddy, Md Arafat Sultan, Vittorio Castelli, Anthony Ferritto, Radu Florian, Efsun Sarioglu Kayi, Salim Roukos, Avi Sil, Todd Ward. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Rong Zhang 0010, Revanth Gangi Reddy, Md. Arafat Sultan, Vittorio Castelli, Anthony Ferritto, Radu Florian, Efsun Sarioglu Kayi, Salim Roukos, Avirup Sil, Todd Ward
EMNLP (1)8
2020 Leveraging Semantic Parsing for Relation Linking over Knowledge Bases
Nandana Mihindukulasooriya, Gaetano Rossiello, Pavan Kapanipathi, Ibrahim Abdelaziz, Srinivas Ravishankar, Mo Yu, Alfio Massimiliano Gliozzo, Salim Roukos, Alexander G. Gray
ISWC (1)8
2019 Rewarding Smatch: Transition-Based AMR Parsing with Reinforcement Learning
abstract
Our work involves enriching the Stack-LSTM transition-based AMR parser (Ballesteros and Al-Onaizan, 2017) by augmenting training with Policy Learning and rewarding the Smatch score of sampled graphs.In addition, we also combined several AMR-to-text alignments with an attention mechanism and we supplemented the parser with pre-processed concept identification, named entities and contextualized embeddings.We achieve a highly competitive performance that is comparable to the best published results.We show an indepth study ablating each of the new components of the parser.
Tahira Naseem, Abhishek Shah, Hui Wan 0001, Radu Florian, Salim Roukos, Miguel Ballesteros
ACL (1)5
2014 Adaptive HTER Estimation for Document-Specific MT Post-Editing
abstract
We present an adaptive translation quality estimation (QE) method to predict the human-targeted translation error rate (HTER) for a document-specific machine translation model.We first introduce features derived internal to the translation decoding process as well as externally from the source sentence analysis.We show the effectiveness of such features in both classification and regression of MT quality.By dynamically training the QE model for the document-specific MT model, we are able to achieve consistency and prediction quality across multiple documents, demonstrated by the higher correlation coefficient and F-scores in finding Good sentences.Additionally, the proposed method is applied to IBM English-to-Japanese MT post editing field study and we observe strong correlation with human preference, with a 10% increase in human translators' productivity.
Fei Huang 0002, Jian-Ming Xu, Abraham Ittycheriah, Salim Roukos
ACL (1)4
2014 Invited Talk: IBM Cognitive Computing - An NLP Renaissance!
Salim Roukos
EMNLP1
2014 Improving MT post-editing productivity with adaptive confidence estimation for document-specific translation model
Fei Huang 0002, Jian-Ming Xu, Abraham Ittycheriah, Salim Roukos
Mach. Transl.4
2012 Document-Specific Statistical Machine Translation for Improving Human Translation Productivity
Salim Roukos, Abraham Ittycheriah, Jian-Ming Xu
CICLing (2)1
2012 Distilling and exploring nuggets from a corpus
abstract
This paper describes a live and scalable system that automatically extracts information nuggets for entities/topics from a continuously updated corpus for effective exploration and analysis. A nugget is a piece of semantic information that (1) must be mapped semantically to the transitive closure of a pre-defined ontology, (2) is explicitly supported by text, and (3) has a natural language description that completely conveys its semantic to a user. Fig. 1 shows a type of nugget "involvement in events" for a person entity (Leon Panetta): each nugget has a short description ("meeting", "news conference") with a list of supporting passages.
Vittorio Castelli, Hema Raghavan, Radu Florian, Ding-Jung Han, Xiaoqiang Luo, Salim Roukos
SIGIR6
2011 A Correction Model for Word Alignments
J. Scott McCarley, Abraham Ittycheriah, Salim Roukos, Bing Xiang, Jian-Ming Xu
EMNLP3
2010 Learning to Predict Readability using Diverse Linguistic Features
Rohit J. Kate, Xiaoqiang Luo, Siddharth Patwardhan, Martin Franz, Radu Florian, Raymond J. Mooney, Salim Roukos, Christopher A. Welty
COLING7
2010 Improving Mention Detection Robustness to Noisy Input
Radu Florian, John F. Pitrelli, Salim Roukos, Imed Zitouni
EMNLP3
2009 Iterative sentence-pair extraction from quasi-parallel corpora for machine translation
abstract
This paper addresses parallel data extraction from the quasi–parallel corpora generated in a crowd-sourcing project where ordinary people watch tv shows and movies and transcribe/translate what they hear, creating document pools in different languages. Since they do not have guidelines for naming and performing translations, it is often not clear which documents are the translations of the same show/movie and which sentences are the translations of the each other in a given document pair. We introduce a method for automatically pairing documents in two languages and extracting parallel sentences from the paired documents. The method consists of three steps: i) document pairing, ii) sentence pair alignment of the paired documents, and iii) context extrapolation to boost the sentence pair coverage. Human evaluation of the extracted data shows that 95 % of the extracted sentences carry useful information for translation. Experimental results also show that using the extracted data provides significant gains over the baseline statistical machine translation system built with manually annotated data. Index Terms: data extraction, comparable data, machine translation
Ruhi Sarikaya, Sameer Maskey, Ea-Ee Jan, Bhuvana Ramabhadran, Salim Roukos
INTERSPEECH7
2009 Real Time Translation Services at IBM
David M. Lubensky, Salim Roukos
MTSummit2
2008 System Combination for Machine Translation of Spoken and Written Language
abstract
This paper describes an approach for computing a consensus translation from the outputs of multiple machine translation (MT) systems. The consensus translation is computed by weighted majority voting on a confusion network, similarly to the well-established ROVER approach of Fiscus for combining speech recognition hypotheses. To create the confusion network, pairwise word alignments of the original MT hypotheses are learned using an enhanced statistical alignment algorithm that explicitly models word reordering. The context of a whole corpus of automatic translations rather than a single sentence is taken into account in order to achieve high alignment quality. The confusion network is rescored with a special language model, and the consensus translation is extracted as the best path. The proposed system combination approach was evaluated in the framework of the TC-STAR speech translation project. Up to six state-of-the-art statistical phrase-based translation systems from different project partners were combined in the experiments. Significant improvements in translation quality from Spanish to English and from English to Spanish in comparison with the best of the individual MT systems were achieved under official evaluation conditions.
Evgeny Matusov, Gregor Leusch, Rafael E. Banchs, Nicola Bertoldi, Daniel Déchelotte, Marcello Federico, Muntsin Kolss, Young-Suk Lee 0001, José B. Mariño, Matthias Paulik, Salim Roukos, Holger Schwenk, Hermann Ney
IEEE Trans. Speech Audio Process.11
2007 Extracting Social Networks and Biographical Facts From Conversational Speech Transcripts
Hongyan Jing, Nanda Kambhatla, Salim Roukos
ACL3
2007 Direct Translation Model 2
Abraham Ittycheriah, Salim Roukos
HLT-NAACL2
2004 A Mention-Synchronous Coreference Resolution Algorithm Based On the Bell Tree
abstract
This paper proposes a new approach for coreference resolution which uses the Bell tree to represent the search space and casts the coreference resolution problem as finding the best path from the root of the Bell tree to the leaf nodes. A Maximum Entropy model is used to rank these paths. The coreference performance on the 2002 and 2003 Automatic Content Extraction (ACE) data will be reported. We also train a coreference system using the MUC6 data and competitive results are obtained.
Xiaoqiang Luo, Abraham Ittycheriah, Hongyan Jing, Nanda Kambhatla, Salim Roukos
ACL5
2004 A Statistical Model for Multilingual Entity Detection and Tracking
Radu Florian, Hany Hassan, Abraham Ittycheriah, Hongyan Jing, Nanda Kambhatla, Xiaoqiang Luo, Nicolas Nicolov, Salim Roukos
HLT-NAACL8
2003 Language Model Based Arabic Word Segmentation
abstract
We approximate Arabic's rich morphology by a model that a word consists of a sequence of morphemes in the pattern prefix*-stem-suffix* (* denotes zero or more occurrences of a morpheme). Our method is seeded by a small manually segmented Arabic corpus and uses it to bootstrap an unsupervised algorithm to build the Arabic word segmenter from a large unsegmented Arabic corpus. The algorithm uses a trigram language model to determine the most probable morpheme sequence for a given input. The language model is initially estimated from a small manually segmented corpus of about 110,000 words. To improve the segmentation accuracy, we use an unsupervised algorithm for automatically acquiring new stems from a 155 million word unsegmented corpus, and re-estimate the model parameters with the expanded vocabulary and training corpus. The resulting Arabic word segmentation system achieves around 97% exact match accuracy on a test corpus containing 28,449 word tokens. We believe this is a state-of-the-art performance and the algorithm can be used for many highly inflected languages provided that one can create a small manually segmented corpus of the language of interest.
Young-Suk Lee 0001, Kishore Papineni, Salim Roukos, Ossama Emam, Hany Hassan
ACL3
2003 tRuEcasIng
abstract
Truecasing is the process of restoring case information to badly-cased or non-cased text. This paper explores truecasing issues and proposes a statistical, language modeling based truecaser which achieves an accuracy of ~98% on news articles. Task based evaluation shows a 26% F-measure improvement in named entity recognition when using truecasing. In the context of automatic content extraction, mention detection on automatic speech recognition text is also improved by a factor of 8. Truecasing also enhances machine translation output legibility and yields a BLEU score improvement of 80.2%. This paper argues for the use of truecasing as a valuable component in text processing applications.
Lucian Vlad Lita, Abraham Ittycheriah, Salim Roukos, Nanda Kambhatla
ACL3
2003 TIPS: A Translingual Information Processing System
Yaser Al-Onaizan, Radu Florian, Martin Franz, Hany Hassan, Young-Suk Lee 0001, J. Scott McCarley, Kishore Papineni, Salim Roukos, Jeffrey S. Sorensen, Christoph Tillmann, Todd Ward
HLT-NAACL8
2003 dentifying and Tracking Entity Mentions in a Maximum Entropy Framework
Abraham Ittycheriah, Lucian Vlad Lita, Nanda Kambhatla, Nicolas Nicolov, Salim Roukos, Margo Stys
HLT-NAACL5
2003 Automatic Derivation of Surface Text Patterns for a Maximum Entropy Based Question Answering System
Deepak Ravichandran, Abraham Ittycheriah, Salim Roukos
HLT-NAACL3
2002 Bleu: a Method for Automatic Evaluation of Machine Translation
abstract
Human evaluations of machine translation are extensive but expensive. Human evaluations can take months to finish and involve human labor that can not be reused. We propose a method of automatic machine translation evaluation that is quick, inexpensive, and language-independent, that correlates highly with human evaluation, and that has little marginal cost per run. We present this method as an automated understudy to skilled human judges which substitutes for them when there is need for quick or frequent evaluations.
Kishore Papineni, Salim Roukos, Todd Ward, Wei-Jing Zhu
ACL2
2002 Active Learning for Statistical Natural Language Parsing
abstract
It is necessary to have a (large) annotated corpus to build a statistical parser. Acquisition of such a corpus is costly and time-consuming. This paper presents a method to reduce this demand using active learning, which selects what samples to annotate, instead of annotating blindly the whole training corpus.Sample selection for annotation is based upon "representativeness" and "usefulness". A model-based distance is proposed to measure the difference of two sentences and their most likely parse trees. Based on this distance, the active learning process analyzes the sample distribution by clustering and calculates the density of each sample to quantify its representativeness. Further more, a sentence is deemed as useful if the existing model is highly uncertain about its parses, where uncertainty is measured by various entropy-based scores.Experiments are carried out in the shallow semantic parser of an air travel dialog system. Our result shows that for about the same parsing accuracy, we only need to annotate a third of the samples as compared to the usual random selection method.
Xiaoqiang Luo, Salim Roukos
ACL3
2002 DARPA communicator evaluation: progress from 2000 to 2001
abstract
This paper describes the evaluation methodology and results of the DARPA Communicator spoken dialog system evaluation experiments in 2000 and 2001. Nine spoken dialog systems in the travel planning domain participated in the experiments resulting in a total corpus of 1904 dialogs. We describe and compare the experimental design of the 2000 and 2001 DARPA evaluations. We describe how we established a performance baseline in 2001 for complex tasks. We present our overall approach to data collection, the metrics collected, and the application of PARADISE to these data sets. We compare the results we achieved in 2000 for a number of core metrics with those for 2001. These results demonstrate large performance improvements from 2000 to 2001 and show that the Communicator program goal of conversational interaction for complex tasks has been achieved.
Marilyn A. Walker, Alexander I. Rudnicky, John S. Aberdeen, Elizabeth Owen Bratt, John S. Garofolo, Helen Hastie, Audrey N. Le, Bryan L. Pellom, Alexandros Potamianos, Rebecca J. Passonneau, Rashmi Prasad, Salim Roukos, Gregory A. Sanders, Stephanie Seneff, David Stallard
INTERSPEECH12
2002 DARPA communicator: cross-system results for the 2001 evaluation
abstract
This paper describes the evaluation methodology and results of the 2001 DARPA Communicator evaluation. The experiment spanned 6 months of 2001 and involved eight DARPA Communicator systems in the travel planning domain. It resulted in a corpus of 1242 dialogs which include many more dialogues for complex tasks than the 2000 evaluation. We describe the experimental design, the approach to data collection, and the results. We compare the results by the type of travel plan and by system. The results demonstrate some large differences across sites and show that the complex trips are clearly more difficult.
Marilyn A. Walker, Alexander I. Rudnicky, Rashmi Prasad, John S. Aberdeen, Elizabeth Owen Bratt, John S. Garofolo, Helen Hastie, Audrey N. Le, Bryan L. Pellom, Alexandros Potamianos, Rebecca J. Passonneau, Salim Roukos, Gregory A. Sanders, Stephanie Seneff, David Stallard
INTERSPEECH12
2002 A multistage algorithm for spotting new words in speech
abstract
In this paper, we present a fast, vocabulary independent, algorithm for spotting words in speech. The algorithm consists of a phone-ngram representation (indexing) stage and a coarse-to-detailed search stage for spotting a word/phone sequence in speech. The phone-ngram representation stage provides a phoneme-level representation of the speech that can be searched efficiently. We present a novel method for phoneme-recognition using a vocabulary prefix tree to guide the creation of the phone-ngram index. The coarse search, consisting of phone-ngram matching, identifies regions of speech as putative word hits. The detailed acoustic match is then conducted only at the putative hits identified in the coarse match. This gives us vocabulary independence and the desired accuracy and speed in wordspotting. Current lattice-based phoneme-matching algorithms are similar to the coarse-match step of our algorithm. We show that our combined algorithm gives a factor of two improvement over the coarse match. The algorithm has wide-ranging use in distributed and pervasive speech recognition applications such as audio-indexing, spoken message retrieval and video-browsing.
Satya Dharanipragada, Salim Roukos
IEEE Trans. Speech Audio Process.2
2000 Statistical methods for topic segmentation
Satya Dharanipragada, Martin Franz, J. Scott McCarley, Kishore Papineni, Salim Roukos, Todd Ward, Wei-Jing Zhu
INTERSPEECH5
2000 Real-time multilingual HMM training robust to channel variations
Ea-Ee Jan, Jaime Botella Ordinas, George Saon, Salim Roukos
INTERSPEECH4
1999 Phrase splicing and variable substitution using the IBM trainable speech synthesis system
abstract
This paper describes a phrase splicing and variable substitution system which offers an intermediate form of automated speech production lying in-between the extremes of recorded utterance playback and full text-to-speech synthesis. The system incorporates a trainable speech synthesiser and an application specific set of pre-recorded phrases. The text to be synthesised is converted to a phone sequence using phone sequences present in the pre-recorded phrases wherever possible, and a pronunciation dictionary elsewhere. The synthesis inventory of the synthesiser is augmented with the synthesis information associated with the pre-recorded phrases used to construct the phone sequence. The synthesiser then performs a dynamic programming search over the augmented inventory to select a segment sequence to produce the output speech. The system enables the seamless splicing of pre-recorded phrases both with other phrases and with synthetic speech. It enables very high quality speech to be produced automatically within a limited domain.
Robert E. Donovan, Martin Franz, Jeffrey S. Sorensen, Salim Roukos
ICASSP4
1999 The IBM conversational telephony system for financial applications
abstract
We describe our development work on a telephonebased conversational system in the domain of mutual fund transactions. This system uses several components including robust large vocabulary continuous speech recognition, natural language understanding, dialog management, and text-to-speech synthesis technologies.
K. Davies, Robert E. Donovan, Mark Epstein, Martin Franz, Abraham Ittycheriah, Ea-Ee Jan, Jean-Michel LeRoux, David M. Lubensky, Chalapathy Neti, Mukund Padmanabhan, Kishore Papineni, Salim Roukos, Andrej Sakrajda, Jeffrey S. Sorensen, Borivoj Tydlitát, Todd Ward
EUROSPEECH12
1999 Story segmentation and topic detection for recognized speech
abstract
We present a technique for the segmention of a sound track into two classes of segments. Each frame of signal is preprocessed by extracting cepstral coefficients and their first order derivatives. For each class, the distribution of the frame parameter vectors is modeled by a Gaussian Mixture Model (GMM). GMM order is selected using two criteria : the Minimum Description Length (MDL) criterion and the Akaike Information Criterion (AIC). Frame score is based on a weighted loglikelihood ratio in a window around the frame. Decision for each frame is taken by comparing its score to a threshold. Experiments are presented on speech / music segmentation in audio tracks. In these experiments, the MDL criterion leads to a reasonable GMM order. Using the MDL criterion for GMM order selection, frame classification error rate is around 20%. However, using GMMs with much lower orders, only decreases marginally performances.
Satya Dharanipragada, Martin Franz, J. Scott McCarley, Salim Roukos, Todd Ward
EUROSPEECH4
1999 Use of recursive mumble models for confidence measuring
abstract
In many speech recognition applications such as name dialing, it is necessary to have the ability to know when a recognition error has occurred so that undesired or unpredicted system behavior can be minimized. Con dence measure is usually used for detection of probable errors. In this paper, a new method for measuring condence is presented. The method is based on use of recursive mumble models. During a regular decoding from which word hypotheses and word boundaries are known, the score of recursive mumblemodels is then determined. The (weighted) di erence between the word detail-match score and the mumble score is used as the con dence measure. It is next compared to a prede ned threshold to decide whether the decoded result is con dently correct or not. The method has been evaluated with two di erent databases. The results show that the new method outperforms our previous method solely based on the word detail-match scores. In particular, the results show that the new method is able to reduce the equal error rate from 32% to 23% and that it rejects far more (78% versus 35%) out-of-domain sentences at the xed 5% false rejection rate.
Qiguang Lin, David M. Lubensky, Salim Roukos
EUROSPEECH3
1999 Free-flow dialog management using forms
abstract
In natural language, some sequences of words are very frequent.A classical language model, like n-gram, does not adequately take into account such sequences, because it underestimates their probabilities.A better approach consists in modeling word sequences as if they were individual dictionary elements.Sequences are considered as additional entries of the word lexicon, on which language models are computed.In this paper, we present two methods for automatically determining frequent phrases in unlabeled corpora of written sentences.These methods are based on information theoretic criteria which insure a high statistical consistency.Our models reach their local optimum since they minimize the perplexity.One procedure is based only on the n-gram language model to extract word sequences.The second one is based on a class n-gram model trained on 233 classes extracted from the eight grammatical classes of French.Experimental tests, in terms of perplexity and recognition rate, are carried out on a vocabulary of 20000 words and a corpus of 43 million words extracted from the "Le Monde" newspaper.Our models reduce perplexity by more than 20% compared with n-gram (nR3) and multigram models.In terms of recognition rate, our models outperform n-gram and multigram models.
Kishore Papineni, Salim Roukos, Todd Ward
EUROSPEECH2
1998 A fast vocabulary independent algorithm for spotting words in speech
abstract
In applications such as audio-indexing, spoken message retrieval and video-browsing, it is necessary to have the ability to detect spoken words that are outside the vocabulary of the speech recognizer used in these systems, in large amounts of speech at speeds many times faster than real-time. We present a fast, vocabulary independent, algorithm for spotting words in speech. The algorithm consists of a preprocessing stage and a coarse-to-detailed search strategy for spotting a word/phone sequence in speech. The preprocessing method provides a phone-level representation of the speech that can be searched efficiently. The coarse search, consisting of phone-ngram matching, identifies regions of speech as putative word hits. The detailed acoustic match is then conducted only at the putative hits identified in the coarse match. This gives us the desired accuracy and speed in word spotting. Overall, the algorithm has a speed of execution that is 2400 times faster than real-time.
Satya Dharanipragada, Salim Roukos
ICASSP2
1998 Maximum likelihood and discriminative training of direct translation models
abstract
We consider translating natural language sentences into a formal language using direct translation models built automatically from training data. Direct translation models have three components: an arbitrary prior conditional probability distribution, features that capture correlations between automatically determined key phrases or sets of words in both languages, and weights associated with these features. The features and the weights are selected using a training corpus of matched pairs of source and target language sentences to maximize the entropy or a new discrimination measure of the resulting conditional probability model. We report results in the air travel information system domain and compare the two methods of training.
Kishore Papineni, Salim Roukos, Todd Ward
ICASSP2
1998 Towards speech understanding across multiple languages
abstract
In this paper we describe our initial eorts in build-ing a natural language understanding (NLU) system across multiple languages. The system allows users to switch lan-guages seamlessly in a single session without requiring any switch in the speech recognition system. Context depen-dence is maintained across sentences, even when the user changes languages. Towards this end we have begun build-ing a universal speech recognizer for English and French languages. We experiment with a universal phonology for both French and English with a novel mechanism to han-dle language dependent variations. Our best results so far show about 5 % relative performance degradation for Eng-lish relative to a unilingual English system and a 9 % rela-tive degradation in French relative to a unilingual French system. The NLU system uses the same statistical un-derstanding algorithms for each language, making system development, maintenance, and portability vastly superior to systems built customly for each language. 1.
Todd Ward, Salim Roukos, Chalapathy Neti, Jerome Gros, Mark Epstein, Satya Dharanipragada
ICSLP2
1998 Speech Research: Near and Not-so-near Results and What They Might Mean for IUI (Panel)
abstract
No abstract available.
Candace L. Sidner, Alex Acero, Janet E. Cahn, Julia Hirschberg, Salim Roukos
IUI6
1998 Probabilistic Modeling for Information Retrieval with Unsupervised Training Data
Ernest P. Chan, Santiago Garcia, Salim Roukos
KDD3
1998 A Method for Scoring Correlated Features in Query Expansion
abstract
No abstract available.
Martin Franz, Salim Roukos
SIGIR2
1997 Fertility Models for Statistical Natural Language Understanding
abstract
Several recent efforts in statistical natural language understanding (NLU) have focused on generating clumps of English words from semantic meaning concepts (Miller et al., 1995; Levin and Pieracini, 1995; Epstein et al., 1996; Epstein, 1996). This paper extends the IBM Machine Translation Group's concept of fertility (Brown et al., 1993) to the generation of clumps for natural language understanding. The basic underlying intuition is that a single concept may be expressed in English as many disjoint clump of words. We present two fertility models which attempt to capture this phenomenon. The first is a Poisson model which leads to appealing computational simplicity. The second is a general nonparametric fertility model. The general model's parameters are boot-strapped from the Poisson model and updated by the EM algorithm. These fertility models can be used to impose clump fertility structure on top of preexisting clump generation models. Here, we present results for adding fertility structure to unigram, bigram, and headword clump generation models on ARPA's Air Travel Information Service (ATIS) domain.
Stephen Della Pietra, Mark Epstein, Salim Roukos, Todd Ward
ACL3
1997 Word-based confidence measures as a guide for stack search in speech recognition
abstract
The maximum a posteriori hypothesis is treated as the decoded truth in speech recognition. However, since the word recognition accuracy is not 100%, it is desirable to have an independent confidence measure on how good the maximum a posteriori hypothesis is relative to the spoken truth for some applications. Efforts are in progress to develop such confidence measures with the intent of applying them to the assessment of the confidence of whole utterances, rescoring of N-best lists, etc. In this paper, we explore the use of word-based confidence measures to adaptively modify the hypothesis score during searches in continuous speech recognition: specifically, based on the confidence of the current sequence of hypothesized words during the search, the weight of its prediction is changed as a function of the confidence. Experimental results are described for ATIS and SwitchBoard tasks. About 8% relative reduction in word error is obtained for ATIS.
Chalapathy Neti, Salim Roukos, Ellen Eide
ICASSP2
1997 Feature-based language understanding
Kishore Papineni, Salim Roukos, Todd Ward
EUROSPEECH2
1997 MDI adaptation of language models across corpora
P. Srinivasa Rao, Satya Dharanipragada, Salim Roukos
EUROSPEECH3
1996 An Iterative Algorithm to Build Chinese Language Models
abstract
We present an iterative procedure to build a Chinese language model (LM). We segment Chinese text into words based on a word-based Chinese language model. However, the construction of a Chinese LM itself requires word boundaries. To get out of the chicken-and-egg problem, we propose an iterative procedure that alternates two operations: segmenting text into words and building an LM. Starting with an initial segmented corpus and an LM based upon it, we use a Viterbi-liek algorithm to segment another set of data. Then, we build an LM based on the second set and use the resulting LM to segment again the first corpus. The alternating procedure provides a self-organized way for the segmenter to detect automatically unseen words and correct segmentation errors. Our preliminary experiment shows that the alternating procedure not only improves the accuracy of our segmentation, but discovers unseen words suprisingly well. The resulting word-based LM has a perplexity of 188 for a general Chinese corpus.
Xiaoqiang Luo, Salim Roukos
ACL2
1996 Statistical natural language understanding using hidden clumpings
abstract
We present a new approach to natural language understanding (NLU) based on the source-channel paradigm, and apply it to ARPA's Air Travel Information Service (ATIS) domain. The model uses techniques similar to those used by IBM in statistical machine translation. The parameters are trained using the exact match algorithm; a hierarchy of models is used to facilitate the bootstrapping of more complex models from simpler models.
Mark Epstein, Kishore Papineni, Salim Roukos, Todd Ward, Stephen Della Pietra
ICASSP3
1995 Performance of the IBM large vocabulary continuous speech recognition system on the ARPA Wall Street Journal task
abstract
In this paper we discuss various experimental results using our continuous speech recognition system on the Wall Street Journal task. Experiments with different feature extraction methods, varying amounts and type of training data, and different vocabulary sizes are reported.
Lalit R. Bahl, S. Balakrishnan-Aiyer, Jerome R. Bellegarda, Martin Franz, Ponani S. Gopalakrishnan, David Nahamoo, Miroslav Novak, Mukund Padmanabhan, Michael Picheny, Salim Roukos
ICASSP10
1995 Language model adaptation via minimum discrimination information
abstract
Statistical language models improve the performance of speech recognition systems by providing estimates of a priori probabilities of word sequences. The commonly used trigram language models obtain the conditional probability estimate of a word given the previous two words, from a large corpus of text. The text corpus is often a collection of several small diverse segments such as newspaper articles, or conversations on different topics. Knowledge of the current topic could be utilized to adapt the general trigram language models to match that topic closely. For example, an interpolation of the general language model with one built on the topic data could be used. The authors first discuss the adaptation of general trigram language models to a known topic using the minimum discrimination information (MDI) method. They then present results on the switchboard corpus which consists of telephone conversations on several topics.
P. Srinivasa Rao, Michael D. Monkowski, Salim Roukos
ICASSP3
1995 A statistical approach to language modelling for the ATIS task
Joshua Koppelman, Stephen Della Pietra, Mark Epstein, Salim Roukos, Todd Ward
EUROSPEECH4
1994 A maximum entropy model for parsing
abstract
this paper, we present a method where more of the tree structure is used in the parsing model. We define a set of features that capture long distance dependency such as parallelism in coordination. These features are then integrated with a Maximum Entropy model into an overall probabilistic model for parsing. We introduce the decision tree parser in Section 2, describe the Maximum Entropy model in Section 3, describe the feature extraction algorithm in Section 4, give experimental results in Section 5, and present our conclusions in Section 6.
Adwait Ratnaparkhi, Salim Roukos, Todd Ward
ICSLP2
1993 Towards History-Based Grammars: Using Richer Models for Probabilistic Parsing
abstract
We describe a generative probabilistic model of natural language, which we call HBG, that takes advantage of detailed linguistic information to resolve ambiguity. HBG incorporates lexical, syntactic, semantic, and structural information from the parse tree into the disambiguation process in a novel way. We use a corpus of bracketed sentences, called a Treebank, in combination with decision tree building to tease out the relevant aspects of a parse tree that will determine the correct parse of a sentence. This stands in contrast to the usual approach of further grammar tailoring via the usual linguistic introspection in the hope of generating the correct parse. In head-to-head tests against one of the best existing robust probabilistic parsing models, which we call P-CFG, the HBG model significantly outperforms P-CFG, increasing the parsing accuracy rate from 60% to 75%, a 37% reduction in error.
Ezra Black, Frederick Jelinek, John D. Lafferty, David M. Magerman, Robert L. Mercer, Salim Roukos
ACL6
1993 Trigger-based language models: a maximum entropy approach
Raymond Y. K. Lau, Ronald Rosenfeld, Salim Roukos
ICASSP (2)3
1992 Development and Evaluation of a Broad-Coverage Probabilistic Grammar of English-Language Computer Manuals
abstract
Americanae nace como un proyecto conjunto que surge dentro de la Red Europea de Información y Documentación sobre América Latina (REDIAL), y que ha afrontado la Biblioteca de la Agencia Española de Cooperación Internacional para el Desarrollo (AECID). Esta nueva biblioteca virtual hace más accesibles los libros digitales de tema americanista a los investigadores y usuarios interesados de cualquier parte del mundo.
Ezra Black, John D. Lafferty, Salim Roukos
ACL3
1992 Adaptation of large vocabulary recognition system parameters
abstract
The authors report on a series of experiments in which the hidden Markov model baseforms and the language model probabilities were updated from spontaneously dictated speech captured during recognition sessions with the IBM Tangora system. The basic technique for baseform modification consisted of constructing new fenonic baseforms for all recognized words. To modify the language model probabilities, a simplified version of a cache language model was implemented. The word error rate across six talkers was 3.7%. Baseform adaptation reduced the average error rate to 3.5%, and using the cache language model reduced the error rate to 3.2%. Combining both techniques further reduced the error rate to 3.1%-a respectable improvement over the original error rate, especially given that the system was speaker-trained prior to adaptation.>
Lalit R. Bahl, Peter V. de Souza, David Nahamoo, Michael Picheny, Salim Roukos
ICASSP5
1992 Adaptive language modeling using minimum discriminant estimation
abstract
The authors present an algorithm to adapt a n-gram language model to a document as it is dictated. The observed partial document is used to estimate a unigram distribution for the words that already occurred. Then, they find the closest n-gram distribution to the static n-gram distribution (using the discrimination information distance measure) that satisfies the marginal constraints derived from the document. The resulting minimum discrimination information model results in a perplexity of 208 instead of 290 for the static trigram model on a document of 321 words.>
Stephen Della Pietra, Vincent J. Della Pietra, Robert L. Mercer, Salim Roukos
ICASSP4
1990 Classifying words for improved statistical language models
abstract
A method for assigning a word to many classes based on the context in which the word occurs is presented. A trigram language model is used to determine the classes which are called statistical synonyms for that word. This classification method is used to build an adaptive language model that incorporates unknown words after their first occurrence by using their statistical synonyms in determining the model's probabilities for the added words. It is shown that the dynamic coverage of the language model increases significantly with a rather low perplexity on the added words.>
Frederick Jelinek, Robert L. Mercer, Salim Roukos
ICASSP3
1989 Speech understanding using a unification grammar
abstract
The authors describe a system for speech understanding that uses hidden Markov models for acoustic modeling, a unification grammar for the syntax of English, and a higher order intensional logic for semantic representation. To maximize speech understanding performance, the constraints of the linguistic models must be used in the search for the most likely interpretation of the spoken message. The basic approach is to perform parsing (as in natural language processing) on the word lattice input produced by a lattice engine. The authors present two lattice parsing algorithms that determine a list (ordered by acoustic likelihood) of word sequences allowed by the grammar and present in the lattice. They describe how semantic constraints can be applied a posteriori to find the most likely interpretation of the input speech. They give speech understanding results on the standard DARPA resource management speech database.>
Yen-Lu Chow, Salim Roukos
ICASSP2
1989 Continuous hidden Markov modeling for speaker-independent word spotting
abstract
A word-spotting system using Gaussian hidden Markov models is presented. Several aspects of this problem are investigated. Specifically, results are reported on the use of various signal processing and feature transformation techniques. The authors have observed that performance can be greatly affected by the choice of features used, the covariance structure of the Gaussian models, and transformations based on energy and feature distributions. Due to the open-set nature of the problem, the specific techniques for modeling out-of-vocabulary speech and the choice of scoring metric can have a significant effect on performance.>
Jan Robin Rohlicek, William Russell, Salim Roukos, Herbert Gish
ICASSP3