Joshua Goodman 0001

dblp:g/JoshuaGoodman · also Joshua T. Goodman · DBLP profile ↗
← Back
24ranked-venue papers
16as first author
0since 2021 · last 2007
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 12 first-authorGraphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-authorSystems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Language models and text generation · 43% Information extraction and text analysis · 27% Probabilistic and Bayesian machine learning · 11%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Network and information security
1 paper
Network security · 100%
Theoretical computer science
2 papers
Automata and formal languages · 100%

Topics — the 13 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › text summarization
multi-document summarization
0.112007
Multi-Document Summarization by Maximizing Informative Content-Words · IJCAI 2007
Information retrieval › online advertising
contextual advertising
0.112006
Finding advertising keywords on web pages · WWW 2006
Information retrieval › text analysis
keyword extraction
0.112006
Finding advertising keywords on web pages · WWW 2006
Natural language and speech › Information extraction and text analysis
syntactic parsing
0.031997
Global Thresholding and Multiple-Pass Parsing · EMNLP 1997
Efficient Algorithms for Parsing the DOP Model · EMNLP 1996
Parsing Algorithms and Metrics · ACL 1996
Network security
spam mitigation
0.012004
Stopping outgoing spam · EC 2004
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › exponential family
maximum entropy models
0.012002
Sequential Conditional Generalized Iterative Scaling · ACL 2002
Natural language and speech › Language models and text generation › language modeling
statistical language modeling
0.012002
Exploring Asymmetric Clustering for Statistical Language Modeling · ACL 2002
Automata and formal languages
parsing algorithms
0.021996
Efficient Algorithms for Parsing the DOP Model · EMNLP 1996
Parsing Algorithms and Metrics · ACL 1996
Information retrieval
query log analysis
0.012006
Finding advertising keywords on web pages · WWW 2006
Machine learning › Trustworthy machine learning
thresholding
0.011997
Global Thresholding and Multiple-Pass Parsing · EMNLP 1997
Natural language and speech › Information extraction and text analysis › syntactic parsing › statistical parsing
data-oriented parsing
0.011996
Efficient Algorithms for Parsing the DOP Model · EMNLP 1996
Natural language and speech › Language models and text generation
language modeling
0.011996
An Empirical Study of Smoothing Techniques for Language Modeling · ACL 1996
Natural language and speech › Language models and text generation › language modeling
n-gram language model
0.012002
Exploring Asymmetric Clustering for Statistical Language Modeling · ACL 2002

Methods — techniques the papers use, named apart from their topics

integer linear programming · 0.1feature-based learning · 0.1paid subscriptions · 0.0computational challenges · 0.0HIP challenges · 0.0sequential parameter training · 0.0maximum likelihood estimation · 0.0generalized iterative scaling · 0.0asymmetric clustering · 0.0bracketed recall · 0.0global thresholding · 0.0viterbi algorithm · 0.0linear interpolation · 0.0labelled recall · 0.0katz smoothing · 0.0jelinek-mercer smoothing · 0.0
YearPublicationVenuePosition
2007 Multi-Document Summarization by Maximizing Informative Content-Words
Scott Yih, Joshua Goodman 0001, Lucy Vanderwende, Hisami Suzuki
IJCAI2
2006 Finding advertising keywords on web pages
abstract
A large and growing number of web pages display contextual advertising based on keywords automatically extracted from the text of the page, and this is a substantial source of revenue supporting the web today. Despite the importance of this area, little formal, published research exists. We describe a system that learns how to extract keywords from web pages for advertisement targeting. The system uses a number of features, such as term frequency of each potential keyword, inverse document frequency, presence in meta-data, and how often the term occurs in search query logs. The system is trained with a set of example pages that have been hand-labeled with "relevant" keywords. Based on this training, it can then extract new keywords from previously unseen pages. Accuracy is substantially better than several baseline systems.
Scott Yih, Joshua Goodman 0001, Vitor R. Carvalho
WWW2
2004 Filtering, Stamping, Blocking, Anti-Spoofing: How to Stop the Spam
Joshua Goodman 0001
LISA1
2004 Exponential Priors for Maximum Entropy Models
Joshua Goodman 0001
HLT-NAACL1
2004 Stopping outgoing spam
abstract
We analyze the problem of preventing outgoing spam. We show that some conventional techniques for limiting outgoing spam are likely to be ineffective. We show that while imposing per message costs would work, less annoying techniques also work. In particular, it is only necessary that the average cost to the spammer over the lifetime of an account exceed his profits, meaning that not every message need be challenged. We develop three techniques, one based on additional HIP challenges, one based on computational challenges, and one based on paid subscriptions. Each system is designed to impose minimal costs on legitimate users, while being too costly for spammers. We also show that maximizing complaint rates is a key factor, and suggest new standards to encourage high complaint rates.
Joshua Goodman 0001, Robert Rounthwaite
EC1
2003 The State of the Art in Language Modeling
Joshua Goodman 0001
HLT-NAACL1
2002 Exploring Asymmetric Clustering for Statistical Language Modeling
abstract
The n-gram model is a stochastic model, which predicts the next word (predicted word) given the previous words (conditional words) in a word sequence. The cluster n-gram model is a variant of the n-gram model in which similar words are classified in the same cluster. It has been demonstrated that using different clusters for predicted and conditional words leads to cluster models that are superior to classical cluster models which use the same clusters for both words. This is the basis of the asymmetric cluster model (ACM) discussed in our study. In this paper, we first present a formal definition of the ACM. We then describe in detail the methodology of constructing the ACM. The effectiveness of the ACM is evaluated on a realistic application, namely Japanese Kana-Kanji conversion. Experimental results show substantial improvements of the ACM in comparison with classical cluster models and word n-gram models at the same model size. Our analysis shows that the high-performance of the ACM lies in the asymmetry of the model.
Jianfeng Gao 0001, Joshua Goodman 0001, Guihong Cao, Hang Li 0001
ACL2
2002 Sequential Conditional Generalized Iterative Scaling
abstract
We describe a speedup for training conditional maximum entropy models. The algorithm is a simple variation on Generalized Iterative Scaling, but converges roughly an order of magnitude faster, depending on the number of constraints, and the way speed is measured. Rather than attempting to train all model parameters simultaneously, the algorithm trains them sequentially. The algorithm is easy to implement, typically uses only slightly more memory, and will lead to improvements for most maximum entropy problems.
Joshua Goodman 0001
ACL1
2002 An Incremental Decision List Learner
abstract
We demonstrate a problem with the standard technique for learning probabilistic decision lists. We describe a simple, incremental algorithm that avoids this problem, and show how to implement it efficiently. We also show
Joshua Goodman 0001
EMNLP1
2002 Language modeling for soft keyboards
abstract
Language models predict the probability of letter sequences. Soft keyboards are images of keyboards on a touch screen for input on Personal Digital Assistants. When a soft keyboard user hits a key near the boundary of a key position, the language model and key press model are combined to select the most probable key sequence. This leads to an overall error rate reduction by a factor of 1.67 to 1.87. An extended version of this paper [4] is available.
Joshua Goodman 0001, Gina Venolia, Keith Steury, Chauncey Parker
IUI1
2002 Reduction of Maximum Entropy Models to Hidden Markov Models
Joshua Goodman 0001
UAI1
2002 Toward a unified approach to statistical language modeling for Chinese
abstract
This article presents a unified approach to Chinese statistical language modeling (SLM). Applying SLM techniques like trigram language models to Chinese is challenging because (1) there is no standard definition of words in Chinese; (2) word boundaries are not marked by spaces; and (3) there is a dearth of training data. Our unified approach automatically and consistently gathers a high-quality training data set from the Web, creates a high-quality lexicon, segments the training data using this lexicon, and compresses the language model, all by using the maximum likelihood principle, which is consistent with trigram model training. We show that each of the methods leads to improvements over standard SLM, and that the combined method yields the best pinyin conversion result reported.
Jianfeng Gao 0001, Joshua Goodman 0001, Mingjing Li, Kai-Fu Lee
ACM Trans. Asian Lang. Inf. Process.2
2001 Classes for fast maximum entropy training
abstract
Maximum entropy models are considered by many to be one of the most promising avenues of language modeling research. Unfortunately, long training times make maximum entropy research difficult. We present a speedup technique: we change the form of the model to use classes. Our speedup works by creating two maximum entropy models, the first of which predicts the class of each word, and the second of which predicts the word itself. This factoring of the model leads to fewer nonzero indicator functions, and faster normalization, achieving speedups of up to a factor of 35 over one of the best previous techniques. It also results in typically slightly lower perplexities. The same trick can be used to speed training of other machine learning techniques, e.g. neural networks, applied to any problem with a large number of outputs, such as language modeling.
Joshua Goodman 0001
ICASSP1
2001 MiPad: a multimodal interaction prototype
abstract
Dr. Who is a Microsoft research project aiming at creating a speech-centric multimodal interaction framework, which serves as the foundation for the NET natural user interface. MiPad is the application prototype that demonstrates compelling user advantages for wireless personal digital assistant (PDA) devices, MiPad fully integrates continuous speech recognition (CSR) and spoken language understanding (SLU) to enable users to accomplish many common tasks using a multimodal interface and wireless technologies. It tries to solve the problem of pecking with tiny styluses or typing on minuscule keyboards in today's PDAs. Unlike a cellular phone, MiPad avoids speech-only interaction. It incorporates a built-in microphone that activates whenever a field is selected. As a user taps the screen or uses a built in roller to navigate, the tapping action narrows the number of possible instructions for spoken word understanding. MiPad currently runs on a Windows CE Pocket PC with a Windows 2000 machine where speech recognition is performed. The Dr Who CSR engine uses a unified CFG and n-gram language model. The Dr Who SLU engine is based on a robust chart parser and a plan-based dialog manager. The paper discusses MiPad's design, implementation work in progress, and preliminary user study in comparison to the existing pen-based PDA interface.
Xuedong Huang 0001, Alex Acero, Ciprian Chelba, Li Deng 0001, Jasha Droppo, Doug Duchene, Joshua Goodman 0001, Hsiao-Wuen Hon, Derek Jacoby, Ricky Loynd, Milind Mahajan, Peter Mau, Scott Meredith, Salman Mughal, Salvado Neto, Mike Plumpe, Kuansan Steury, Gina Venolia, Kuansan Wang, Ye-Yi Wang
ICASSP7
2001 A bit of progress in language modeling
Joshua Goodman 0001
Comput. Speech Lang.1
2000 Putting it all together: language model combination
abstract
In the past several years, a number of different language modeling improvements over simple trigram models have been found, including caching, higher-order n-grams, skipping, modified Kneser-Ney smoothing and clustering. While all of these techniques have been studied separately, they have rarely been studied in combination. We find some significant interactions, especially with smoothing techniques. The combination of all techniques leads to up to a 45% perplexity reduction over a Katz (1987) smoothed trigram model with no count cutoffs, the highest such perplexity reduction reported.
Joshua Goodman 0001
ICASSP1
2000 Language model size reduction by pruning and clustering
abstract
Several techniques are known for reducing the size of language models, including count cutoffs [1], Weighted Difference pruning [2], Stolcke pruning [3], and clustering [4]. We compare all of these techniques and show some surprising results. For instance, at low pruning thresholds, Weighted Difference and Stolcke pruning underperform count cutoffs. We then show novel clustering techniques that can be combined with Stolcke pruning to produce the smallest models at a given perplexity. The resulting models can be a factor of three or more smaller than models pruned with Stolcke pruning, at the same perplexity. The technique creates clustered models that are often larger than the unclustered models, but which can be pruned to models that are smaller than unclustered models with the same perplexity.
Joshua Goodman 0001, Jianfeng Gao 0001
INTERSPEECH1
2000 Mipad: a next generation PDA prototype
abstract
MiPad is one of the application prototypes in a project codenamed Dr Who. As a wireless Personal Digital Assistant (PDA), MiPad fully integrates continuous speech recognition (CSR) and spoken language understanding (SLU) to enable users to accomplish many common tasks using a multimodal interface and wireless technologies. It tries to solve the problem of pecking with tiny styluses or typing on minuscule keyboards in today’s PDAs or smart phones. It also avoids the problem of being a cellular telephone that depends on speech-only interaction. MiPad incorporates a built-in microphone that activates whenever a field is selected. As a user taps the screen or uses a built-in roller to navigate, the tapping action narrows the number of possible instructions for spoken language processing. MiPad currently runs on a Windows CE Pocket PC with a Windows 2000 Server where speech recognition is performed. The Dr Who CSR engine has a 64k word vocabulary with a unified context-free grammar and n-gram language model. The Dr Who SLU engine is based on a robust chart parser and a plan-based dialog manager. This paper discusses MiPad’s design, implementation work in progress, and preliminary user study in comparison to the existing pen-based PDA interface. 1.
Xuedong Huang 0001, Alex Acero, Ciprian Chelba, Li Deng 0001, Doug Duchene, Joshua Goodman 0001, Hsiao-Wuen Hon, Derek Jacoby, Ricky Loynd, Milind Mahajan, Peter Mau, Scott Meredith, Salman Mughal, Salvado Neto, Mike Plumpe, Kuansan Wang, Ye-Yi Wang
INTERSPEECH6
1999 Semiring Parsing
Joshua Goodman 0001
Comput. Linguistics1
1999 An empirical study of smoothing techniques for language modeling
Stanley F. Chen, Joshua Goodman 0001
Comput. Speech Lang.2
1997 Global Thresholding and Multiple-Pass Parsing
Joshua Goodman 0001
EMNLP1
1996 An Empirical Study of Smoothing Techniques for Language Modeling
abstract
We present an extensive empirical comparison of several smoothing techniques in the domain of language modeling, including those described by Jelinek and Mercer (1980), Katz (1987), and Church and Gale (1991). We investigate for the first time how factors such as training data size, corpus (e.g., Brown versus Wall Street Journal), and n-gram order (bigram versus trigram) affect the relative performance of these methods, which we measure through the cross-entropy of test data. In addition, we introduce two novel smoothing techniques, one a variation of Jelinek-Mercer smoothing and one a very simple linear interpolation technique, both of which outperform existing methods.
Stanley F. Chen, Joshua Goodman 0001
ACL2
1996 Parsing Algorithms and Metrics
abstract
Many different metrics exist for evaluating parsing results, including Viterbi, Crossing Brackets Rate, Zero Crossing Brackets Rate, and several others. However, most parsing algorithms, including the Viterbi algorithm, attempt to optimize the same metric, namely the probability of getting the correct labelled tree. By choosing a parsing algorithm appropriate for the evaluation metric, better performance can be achieved. We present two new algorithms: the "Labelled Recall Algorithm," which maximizes the expected Labelled Recall Rate, and the "Bracketed Recall Algorithm," which maximizes the Bracketed Recall Rate. Experimental results are given, showing that the two new algorithms have improved performance over the Viterbi algorithm on many criteria, especially the ones that they optimize.
Joshua Goodman 0001
ACL1
1996 Efficient Algorithms for Parsing the DOP Model
Joshua Goodman 0001
EMNLP1