Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

J. Scott McCarley

dblp:80/3522 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
1since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Question answering and dialogue systems · 71% Deep learning architectures and training · 11% Machine translation · 9%
Databases, data mining, and information retrieval
9 papers
Information retrieval · 90% Data mining · 10%

Topics — the 18 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
extractive question answering
0.712023
GAAMA 2.0: An Integrated System That Answers Boolean and Extractive Questions · AAAI 2023
Natural language and speech › Question answering and dialogue systems
machine reading comprehension
0.712023
GAAMA 2.0: An Integrated System That Answers Boolean and Extractive Questions · AAAI 2023
Machine learning › Deep learning architectures and training
transformer
0.212023
GAAMA 2.0: An Integrated System That Answers Boolean and Extractive Questions · AAAI 2023
Information retrieval
cross-language information retrieval
0.242009
Cross language name matching · SIGIR 2009
Quantifying the Utility of Parallel Corpora · SIGIR 2001
Machine Translation and Monolingual Information Retrieval (poster abstract) · SIGIR 1999
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.112020
The TechQA Dataset · ACL 2020
Natural language and speech › Machine translation › statistical machine translation
word alignment
0.112011
A Correction Model for Word Alignments · EMNLP 2011
Information retrieval
evaluation
0.132007
User-oriented text segmentation evaluation measure · SIGIR 2007
Influence of speech recognition errors on topic detection · SIGIR 2000
Quantifying the Utility of Parallel Corpora · SIGIR 2001
Information retrieval › cross-language information retrieval
transliteration
0.112009
Cross language name matching · SIGIR 2009
Information retrieval › text analysis › topic analysis
topic detection and tracking
0.022001
Influence of speech recognition errors on topic detection · SIGIR 2000
Unsupervised and Supervised Clustering for Topic Tracking · SIGIR 2001
Information retrieval › indexing
index compression
0.012002
How Many Bits are Needed to Store Term Frequencies? · SIGIR 2002
Data mining › clustering
document clustering
0.012001
Unsupervised and Supervised Clustering for Topic Tracking · SIGIR 2001
Information retrieval › information filtering
topic tracking
0.012001
Unsupervised and Supervised Clustering for Topic Tracking · SIGIR 2001
Information retrieval › ranking › relevance estimation
relevance scoring
0.012000
Word document density and relevance scoring · SIGIR 2000
Information retrieval
retrieval models
0.012000
Word document density and relevance scoring · SIGIR 2000
Data mining › text mining
topic detection
0.012000
Influence of speech recognition errors on topic detection · SIGIR 2000
Natural language and speech › Machine translation
cross-lingual retrieval
0.011999
Should we Translate the Documents or the Queries in Cross-language Information Retrieval? · ACL 1999
Information retrieval
term frequency
0.012002
How Many Bits are Needed to Store Term Frequencies? · SIGIR 2002
Information retrieval › text analysis › text corpus analysis
word frequency distribution
0.012000
Word document density and relevance scoring · SIGIR 2000

Methods — techniques the papers use, named apart from their topics

transformer models · 0.7adapter · 0.7dataset construction · 0.4correction model · 0.1statistical machine translation · 0.1word-based translation model · 0.1unsupervised learning · 0.1user-oriented evaluation · 0.1unsupervised clustering · 0.1supervised clustering · 0.0out-of-vocabulary rate analysis · 0.0cross-language IR · 0.0
YearPublicationVenuePosition
2023 GAAMA 2.0: An Integrated System That Answers Boolean and Extractive Questions
abstract
Recent machine reading comprehension datasets include extractive and boolean questions but current approaches do not offer integrated support for answering both question types. We present a front-end demo to a multilingual machine reading comprehension system that handles boolean and extractive questions. It provides a yes/no answer and highlights the supporting evidence for boolean questions. It provides an answer for extractive questions and highlights the answer in the passage. Our system, GAAMA 2.0, achieved first place on the TyDi QA leaderboard at the time of submission. We contrast two different implementations of our approach: including multiple transformer models for easy deployment, and a shared transformer model utilizing adapters to reduce GPU memory footprint for a resource-constrained environment.
J. Scott McCarley, Mihaela A. Bornea, Sara Rosenthal, Anthony Ferritto, Md. Arafat Sultan, Avirup Sil, Radu Florian
AAAI1
2020 The TechQA Dataset
abstract
Vittorio Castelli, Rishav Chakravarti, Saswati Dana, Anthony Ferritto, Radu Florian, Martin Franz, Dinesh Garg, Dinesh Khandelwal, Scott McCarley, Michael McCawley, Mohamed Nasr, Lin Pan, Cezar Pendus, John Pitrelli, Saurabh Pujar, Salim Roukos, Andrzej Sakrajda, Avi Sil, Rosario Uceda-Sosa, Todd Ward, Rong Zhang. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Vittorio Castelli, Rishav Chakravarti, Saswati Dana, Anthony Ferritto, Radu Florian, Martin Franz, Dinesh Garg, Dinesh Khandelwal, J. Scott McCarley, Mike McCawley, Mohamed Nasr, Lin Pan 0003, Cezar Pendus, John F. Pitrelli, Saurabh Pujar, Salim Roukos, Andrej Sakrajda, Avirup Sil, Rosario Uceda-Sosa, Todd Ward, Rong Zhang 0010
ACL9
2011 A Correction Model for Word Alignments
J. Scott McCarley, Abraham Ittycheriah, Salim Roukos, Bing Xiang, Jian-Ming Xu
EMNLP1
2009 Cross language name matching
abstract
Cross language information retrieval methods are used to determine which segments of Arabic language documents match name-based English queries. We investigate and contrast a word-based translation model with a character-based transliteration model in order to handle spelling variation and previously unseen names. We measure performance by making a novel use of the training data from the 2007 ACE Entity Translation
J. Scott McCarley
SIGIR1
2007 User-oriented text segmentation evaluation measure
abstract
The paper describes a user oriented performance evaluation measure for text segmentation. Experiments show that the proposed measure differentiates well between error distributions with varying user impact.
Martin Franz, J. Scott McCarley, Jian-Ming Xu
SIGIR2
2003 Unsupervised Learning of Arabic Stemming Using a Parallel Corpus
abstract
This paper presents an unsupervised learning approach to building a non-English (Arabic) stemmer. The stemming model is based on statistical machine translation and it uses an English stemmer and a small (10 K sentences) parallel corpus as its sole training resources. No parallel text is needed after the training phase. Monolingual, unannotated text can be used to further improve the stemmer by allowing it to adapt to a desired domain or genre. Examples and results will be given for Arabic, but the approach is applicable to any language that needs affix removal. Our resource-frugal approach results in 87.5% agreement with a state of the art, proprietary Arabic stemmer built using rules, affix lists, and human annotated text, in addition to an unsupervised component. Task-based evaluation using Arabic information retrieval indicates an improvement of 22-38% in average precision over unstemmed text, and 96% of the performance of the proprietary stemmer above.
Monica Rogati, J. Scott McCarley, Yiming Yang 0002
ACL2
2003 TIPS: A Translingual Information Processing System
Yaser Al-Onaizan, Radu Florian, Martin Franz, Hany Hassan, Young-Suk Lee 0001, J. Scott McCarley, Kishore Papineni, Salim Roukos, Jeffrey S. Sorensen, Christoph Tillmann, Todd Ward
HLT-NAACL6
2002 How Many Bits are Needed to Store Term Frequencies?
abstract
Search algorithms in most current text retrieval systems use index data structures extracted from the original text documents. In this paper we focus on reducing the size of the indices by reducing the amount of space dedicated to store term frequencies. In experiments using TREC Ad Hoc [2, 3] corpora and query sets, we show that it is possible to store the term frequency in only two bits without decreasing retrieval performance.
Martin Franz, J. Scott McCarley
SIGIR2
2001 Topic styles in IR and TDT: effect on system behavior
abstract
The TREC Spoken Document Retrieval Track (SDR) and the Topic Detection and Tracking (TDT) project have annotated the same corpus with difference styles of relevance judgements, using differenct notions of topic. We compare the behavior of a topic tracking system using relevance judgements from TDT with that of the same system using relevance from the SDR in order to investigate the influence of differences document relevance judgements on the behavior of the tracking system.
Martin Franz, J. Scott McCarley, Todd Ward, Wei-Jing Zhu
INTERSPEECH2
2001 Unsupervised and Supervised Clustering for Topic Tracking
abstract
We investigate important differences between two styles of document clustering in the context of Topic Detection and Tracking. Converting a Topic Detection system into a Topic Tracking system exposes fundamental differences between these two tasks that are important to consider in both the design and the evaluation of TDT systems. We also identify features that can be used in systems for both tasks.
Martin Franz, J. Scott McCarley, Todd Ward, Wei-Jing Zhu
SIGIR2
2001 Quantifying the Utility of Parallel Corpora
abstract
Our English-Chinese cross-language IR system is trained from parallel corpora; we investigate its performance as a function of training corpus size for three different training corpora. We find that the performance of the system as trained on the three parallel corpora can be related by a simple measure, namely the out-of-vocabulary rate of query words.
Martin Franz, J. Scott McCarley, Todd Ward, Wei-Jing Zhu
SIGIR2
2000 Statistical methods for topic segmentation
Satya Dharanipragada, Martin Franz, J. Scott McCarley, Kishore Papineni, Salim Roukos, Todd Ward, Wei-Jing Zhu
INTERSPEECH3
2000 Word document density and relevance scoring
abstract
Previous work addressing the issue of word distribution in documents has shown the importance of Word repetitiveness as an indicator of the word content-bearing characteristics. In this paper we propose a simple method using a measure of the tendency of words to repeat within a document to separate the words with similar document frequencies, but different topic discriminating characteristics. We describe the application of the new measure in query-document relevance scoring. Experiments on the TREC Ad Hoc and Spoken Document Retrieval tasks [7] show useful performance improvements.
Martin Franz, J. Scott McCarley
SIGIR2
2000 Influence of speech recognition errors on topic detection
abstract
We investigate the effect of speech-recognition errors on a system for the unsupervised, nearly synchronous clustering of broadcast news stories, using the TDT (Topic Detection and Tracking) Corpora. Two questions are addressed: (1) Are speech recognition errors detrimental to the performance of the system? (2) Can a background collection of contemporaneous clean text improve performance? We investigate both the large-cluster and small-cluster limits.
J. Scott McCarley, Martin Franz
SIGIR1
1999 Should we Translate the Documents or the Queries in Cross-language Information Retrieval?
abstract
Previous comparisons of document and query translation suffered difficulty due to differing quality of machine translation in these two opposite directions. We avoid this difficulty by training identical statistical translation models for both translation directions using the same training data. We investigate information retrieval between English and French, incorporating both translations directions into both document translation and query translation-based information retrieval, as well as into hybrid systems. We find that hybrids of document and query translation-based systems out-perform query translation systems, even human-quality query translation systems.
J. Scott McCarley
ACL1
1999 Story segmentation and topic detection for recognized speech
abstract
We present a technique for the segmention of a sound track into two classes of segments. Each frame of signal is preprocessed by extracting cepstral coefficients and their first order derivatives. For each class, the distribution of the frame parameter vectors is modeled by a Gaussian Mixture Model (GMM). GMM order is selected using two criteria : the Minimum Description Length (MDL) criterion and the Akaike Information Criterion (AIC). Frame score is based on a weighted loglikelihood ratio in a window around the frame. Decision for each frame is taken by comparing its score to a threshold. Experiments are presented on speech / music segmentation in audio tracks. In these experiments, the MDL criterion leads to a reasonable GMM order. Using the MDL criterion for GMM order selection, frame classification error rate is around 20%. However, using GMMs with much lower orders, only decreases marginally performances.
Satya Dharanipragada, Martin Franz, J. Scott McCarley, Salim Roukos, Todd Ward
EUROSPEECH3
1999 Machine Translation and Monolingual Information Retrieval (poster abstract)
abstract
No abstract available.
Martin Franz, J. Scott McCarley
SIGIR2