Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

John Blitzer

dblp:48/2163 · also John C. Blitzer · DBLP profile ↗
← Back
23ranked-venue papers
4as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 4 first-authorHuman-computer interaction and ubiquitous computing · 2Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
14 papers
Information extraction and text analysis · 35% Machine translation · 12% Transfer learning and domain adaptation · 12%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%

Topics — the 30 heaviest of 39, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.342011
Co-Training for Domain Adaptation · NIPS 2011
Learning Bounds for Domain Adaptation · NIPS 2007
Analysis of Representations for Domain Adaptation · NIPS 2006
Natural language and speech › Information extraction and text analysis › syntactic parsing › grammar-based parsing
combinatory categorial grammar parsing
0.212016
Evaluating Induced CCG Parsers on Grounded Semantic Parsing · EMNLP 2016
Natural language and speech › Information extraction and text analysis
semantic parsing
0.212016
Evaluating Induced CCG Parsers on Grounded Semantic Parsing · EMNLP 2016
Natural language and speech › Information extraction and text analysis
slot filling
0.212016
Evaluating Induced CCG Parsers on Grounded Semantic Parsing · EMNLP 2016
Natural language and speech › Information extraction and text analysis
syntactic parsing
0.212016
Evaluating Induced CCG Parsers on Grounded Semantic Parsing · EMNLP 2016
Machine learning › Learning theory › generalization bounds
domain adaptation bound
0.122007
Learning Bounds for Domain Adaptation · NIPS 2007
Analysis of Representations for Domain Adaptation · NIPS 2006
Machine learning › Graph learning › limited supervision › multi-view semi-supervised learning
co-training
0.112011
Co-Training for Domain Adaptation · NIPS 2011
Machine learning › Learning paradigms
semi-supervised learning
0.112011
Co-Training for Domain Adaptation · NIPS 2011
Natural language and speech › Machine translation › synchronous grammar
inversion transduction grammar
0.112009
Better Word Alignments with Supervised ITG Models · ACL/IJCNLP 2009
Natural language and speech › Machine translation
statistical machine translation
0.112009
Better Word Alignments with Supervised ITG Models · ACL/IJCNLP 2009
Natural language and speech › Machine translation › statistical machine translation
word alignment
0.112009
Better Word Alignments with Supervised ITG Models · ACL/IJCNLP 2009
Information retrieval
cross-language information retrieval
0.112009
Exploiting Bilingual Information to Improve Web Search · ACL/IJCNLP 2009
Computer vision › Image recognition and object detection
feature similarity
0.112008
Regularized Learning with Networks of Features · NIPS 2008
Machine learning › Graph learning
graph regularization
0.112008
Regularized Learning with Networks of Features · NIPS 2008
Machine learning › Optimization for machine learning
regularized learning
0.112008
Regularized Learning with Networks of Features · NIPS 2008
Human-AI interaction
intelligent user interfaces
0.112008
Intelligent Email: Aiding Users with AI · AAAI 2008
Natural language and speech › Language models and text generation
grammar induction
0.112016
Evaluating Induced CCG Parsers on Grounded Semantic Parsing · EMNLP 2016
Natural language and speech › Language models and text generation › grammar induction
unsupervised grammar induction
0.112016
Evaluating Induced CCG Parsers on Grounded Semantic Parsing · EMNLP 2016
Natural language and speech › Information extraction and text analysis › sentiment analysis
sentiment classification
0.112007
Biographies, Bollywood, Boom-boxes and Blenders: Domain Adaptation for Sentiment Classification · ACL 2007
Machine learning › Learning theory › generalization bounds
uniform convergence
0.112007
Learning Bounds for Domain Adaptation · NIPS 2007
Machine learning › Learning theory
generalization bounds
0.112006
Analysis of Representations for Domain Adaptation · NIPS 2006
Machine learning › Transfer learning and domain adaptation › domain adaptation
representation learning for domain adaptation
0.112006
Analysis of Representations for Domain Adaptation · NIPS 2006
Machine learning › Transfer learning and domain adaptation › feature-based transfer learning
structural correspondence learning
0.112006
Domain Adaptation with Structural Correspondence Learning · EMNLP 2006
Machine learning › Representation and self-supervised learning › representation learning › metric learning
mahalanobis distance metric learning
0.112005
Distance Metric Learning for Large Margin Nearest Neighbor Classification · NIPS 2005
Machine learning › Representation and self-supervised learning › representation learning
metric learning
0.112005
Distance Metric Learning for Large Margin Nearest Neighbor Classification · NIPS 2005
Machine learning › Representation and self-supervised learning › word representation
distributed representation
0.012004
Hierarchical Distributed Representations for Statistical Language Modeling · NIPS 2004
Machine learning › Representation and self-supervised learning › hierarchical representation
hierarchical representation learning
0.012004
Hierarchical Distributed Representations for Statistical Language Modeling · NIPS 2004
Natural language and speech › Language models and text generation › language modeling
statistical language modeling
0.012004
Hierarchical Distributed Representations for Statistical Language Modeling · NIPS 2004
Natural language and speech › Language models and text generation › text summarization
summarization evaluation
0.012003
Evaluation Challenges in Large-Scale Document Summarization · ACL 2003
Natural language and speech › Language models and text generation
text summarization
0.012003
Evaluation Challenges in Large-Scale Document Summarization · ACL 2003

Methods — techniques the papers use, named apart from their topics

extrinsic task-based evaluation · 0.2CCG parsing · 0.2query expansion · 0.2bilingual term translation · 0.2AI assistance · 0.2joint optimization · 0.1feature space view splitting · 0.1supervised ITG · 0.1manifold learning · 0.1graph regularization · 0.1meta-evaluation · 0.0
YearPublicationVenuePosition
2016 Evaluating Induced CCG Parsers on Grounded Semantic Parsing
abstract
We compare the effectiveness of four different syntactic CCG parsers for a semantic slotfilling task to explore how much syntactic supervision is required for downstream semantic analysis.This extrinsic, task-based evaluation also provides a unique window into the semantics captured (or missed) by unsupervised grammar induction systems.
Yonatan Bisk, Siva Reddy, John Blitzer, Julia Hockenmaier, Mark Steedman
EMNLP3
2012 Latent Structured Ranking
Jason Weston, John Blitzer
UAI2
2011 Co-Training for Domain Adaptation
abstract
Domain adaptation algorithms seek to generalize a model trained in a source domain to a new target domain. In many practical cases, the source and target distributions can differ substantially, and in some cases crucial target features may not have support in the source domain. In this paper we introduce an algorithm that bridges the gap between source and target domains by slowly adding both the target features and instances in which the current algorithm is the most confident. Our algorithm is a variant of co-training, and we name it CODA (Co-training for domain adaptation). Unlike the original co-training work, we do not assume a particular feature split. Instead, for each iteration of co-training, we add target features and formulate a single optimization problem which simultaneously learns a target predictor, a split of the feature space into views, and a shared subset of source and target features to include in the predictor. CODA significantly out-performs the state-of-the-art on the 12-domain benchmark data set of Blitzer et al.. Indeed, over a wide range (65 of 84 comparisons) of target supervision, ranging from no labeled target data to a relatively large number of target labels, CODA achieves the best performance.
Minmin Chen, Kilian Q. Weinberger, John Blitzer
NIPS3
2010 Learning Better Monolingual Models with Unannotated Bilingual Text
David Burkett, Slav Petrov, John Blitzer, Daniel Klein 0001
CoNLL3
2010 Joint Parsing and Alignment with Weakly Synchronized Grammars
David Burkett, John Blitzer, Daniel Klein 0001
HLT-NAACL2
2010 A theory of learning from different domains
abstract
Discriminative learning methods for classification perform well when training and test data are drawn from the same distribution. Often, however, we have plentiful labeled training data from a source domain but wish to learn a classifier which performs well on a target domain with a different distribution and little or no labeled training data. In this work we investigate two questions. First, under what conditions can a classifier trained from source data be expected to perform well on target data? Second, given a small amount of labeled target data, how should we combine it during training with the large amount of labeled source data to achieve the lowest target error at test time? We address the first question by bounding a classifier’s target error in terms of its source error and the divergence between the two domains. We give a classifier-induced divergence measure that can be estimated from finite, unlabeled samples from the domains. Under the assumption that there exists some hypothesis that performs well in both domains, we show that this quantity together with the empirical source error characterize the target error of a source-trained classifier. We answer the second question by bounding the target error of a model which minimizes a convex combination of the empirical source and target errors. Previous theoretical work has considered minimizing just the source error, just the target error, or weighting instances from the two domains equally. We show how to choose the optimal combination of source and target error as a function of the divergence, the sample sizes of both domains, and the complexity of the hypothesis class. The resulting bound generalizes the previously studied cases and is always at least as tight as a bound which considers minimizing only the target error or an equal weighting of source and target errors.
Shai Ben-David, John Blitzer, Koby Crammer, Alex Kulesza, Fernando Pereira 0003, Jennifer Wortman Vaughan
Mach. Learn.2
2009 Exploiting Bilingual Information to Improve Web Search
Wei Gao 0001, John Blitzer, Ming Zhou 0001, Kam-Fai Wong
ACL/IJCNLP2
2009 Better Word Alignments with Supervised ITG Models
Aria Haghighi, John Blitzer, John DeNero, Daniel Klein 0001
ACL/IJCNLP2
2008 Intelligent Email: Aiding Users with AI
Mark Dredze, Hanna M. Wallach, Danny Puller, Tova Brooks, Josh Carroll, Joshua Magarick, John Blitzer, Fernando Pereira 0003
AAAI7
2008 Intelligent email: reply and attachment prediction
abstract
We present two prediction problems under the rubric of Intelligent Email that are designed to support enhanced email interfaces that relieve the stress of email overload. Reply prediction alerts users when an email requires a response and facilitates email response management. Attachment prediction alerts users when they are about to send an email missing an attachment or triggers a document recommendation system, which can catch missing attachment emails before they are sent. Both problems use the same underlying email classification system and task specific features. Each task is evaluated for both single-user and cross-user settings.
Mark Dredze, Tova Brooks, Josh Carroll, Joshua Magarick, John Blitzer, Fernando Pereira 0003
IUI5
2008 Regularized Learning with Networks of Features
abstract
For many supervised learning problems, we possess prior knowledge about which features yield similar information about the target variable. In predicting the topic of a document, we might know that two words are synonyms, or when performing image recognition, we know which pixels are adjacent. Such synonymous or neighboring features are near-duplicates and should therefore be expected to have similar weights in a good model. Here we present a framework for regularized learning in settings where one has prior knowledge about which features are expected to have similar and dissimilar weights. This prior knowledge is encoded as a graph whose vertices represent features and whose edges represent similarities and dissimilarities between them. During learning, each feature's weight is penalized by the amount it differs from the average weight of its neighbors. For text classification, regularization using graphs of word co-occurrences outperforms manifold learning and compares favorably to other recently proposed semi-supervised learning methods. For sentiment analysis, feature graphs constructed from declarative human knowledge, as well as from auxiliary task learning, significantly improve prediction accuracy.
Ted Sandler, John Blitzer, Partha P. Talukdar, Lyle H. Ungar
NIPS2
2008 Multi-View Learning over Structured and Non-Identical Outputs
Kuzman Ganchev, João Graça, John Blitzer, Ben Taskar
UAI3
2008 Multi-View Learning over Structured and Non-Identical Outputs
Kuzman Ganchev, João Graça, John Blitzer, Ben Taskar
UAI3
2007 Biographies, Bollywood, Boom-boxes and Blenders: Domain Adaptation for Sentiment Classification
John Blitzer, Mark Dredze, Fernando Pereira 0003
ACL1
2007 Frustratingly Hard Domain Adaptation for Dependency Parsing
Mark Dredze, John Blitzer, Partha P. Talukdar, Kuzman Ganchev, João Graça, Fernando Pereira 0003
EMNLP-CoNLL2
2007 Learning Bounds for Domain Adaptation
abstract
Empirical risk minimization offers well-known learning guarantees when training and test data come from the same domain. In the real world, though, we often wish to adapt a classifier from a source domain with a large amount of training data to different target domain with very little training data. In this work we give uniform convergence bounds for algorithms that minimize a convex combination of source and target empirical risk. The bounds explicitly model the inherent trade-off between training on a large but inaccurate source data set and a small but accurate target training set. Our theory also gives results when we have multiple source domains, each of which may have a different number of instances, and we exhibit cases in which minimizing a non-uniform combination of source risks can achieve much lower target error than standard empirical risk minimization.
John Blitzer, Koby Crammer, Alex Kulesza, Fernando Pereira 0003, Jennifer Wortman Vaughan
NIPS1
2006 Domain Adaptation with Structural Correspondence Learning
John Blitzer, Ryan T. McDonald, Fernando Pereira 0003
EMNLP1
2006 Analysis of Representations for Domain Adaptation
abstract
Discriminative learning methods for classification perform well when training and test data are drawn from the same distribution. In many situations, though, we have labeled training data for a source domain, and we wish to learn a classifier which performs well on a target domain with a different distribution. Under what conditions can we adapt a classifier trained on the source domain for use in the target domain? Intuitively, a good feature representation is a crucial factor in the success of domain adaptation. We formalize this intuition theoretically with a generalization bound for domain adaption. Our theory illustrates the tradeoffs inherent in designing a representation for domain adaptation and gives a new justification for a recently proposed model. It also points toward a promising new model for domain adaptation: one which explicitly minimizes the difference between the source and target domains, while at the same time maximizing the margin of the training set.
Shai Ben-David, John Blitzer, Koby Crammer, Fernando Pereira 0003
NIPS2
2005 Distance Metric Learning for Large Margin Nearest Neighbor Classification
abstract
We show how to learn a Mahanalobis distance metric for k -nearest neighbor (kNN) classification by semidefinite programming. The metric is trained with the goal that the k -nearest neighbors always belong to the same class while examples from different classes are separated by a large margin. On seven data sets of varying size and difficulty, we find that metrics trained in this way lead to significant improvements in kNN classification--for example, achieving a test error rate of 1.3% on the MNIST handwritten digits. As in support vector machines (SVMs), the learning problem reduces to a convex optimization based on the hinge loss. Unlike learning in SVMs, however, our framework requires no modification or extension for problems in multiway (as opposed to binary) classification.
Kilian Q. Weinberger, John Blitzer, Lawrence K. Saul
NIPS2
2004 MEAD - A Platform for Multidocument Multilingual Text Summarization
Dragomir R. Radev, Timothy Allison, Sasha Blair-Goldensohn, John Blitzer, Arda Çelebi, Stanko Dimitrov, Elliott Drábek, Ali Hakim, Wai Lam, Danyu Liu, Jahna Otterbacher, Horacio Saggion, Simone Teufel, Michael Topper, Adam Winkel
LREC4
2004 Hierarchical Distributed Representations for Statistical Language Modeling
abstract
Statistical language models estimate the probability of a word occurring in a given context. The most common language models rely on a discrete enumeration of predictive contexts (e.g., n-grams) and consequently fail to capture and exploit statistical regularities across these contexts. In this paper, we show how to learn hierarchical, distributed representations of word contexts that maximize the predictive value of a statistical language model. The representations are initialized by unsupervised algorithms for linear and nonlinear dimensionality reduction [14], then fed as input into a hierarchical mixture of experts, where each expert is a multinomial dis- tribution over predicted words [12]. While the distributed representations in our model are inspired by the neural probabilistic language model of Bengio et al. [2, 3], our particular architecture enables us to work with significantly larger vocabularies and training corpora. For example, on a large-scale bigram modeling task involving a sixty thousand word vocab- ulary and a training corpus of three million sentences, we demonstrate consistent improvement over class-based bigram models [10, 13]. We also discuss extensions of our approach to longer multiword contexts. 1 Introduction Statistical language models are essential components of natural language systems for human-computer interaction. They play a central role in automatic speech recognition [11], machine translation [5], statistical parsing [8], and information retrieval [15]. These mod- els estimate the probability that a word will occur in a given context, where in general a context specifies a relationship to one or more words that have already been observed. The simplest, most studied case is that of n-gram language modeling, where each word is predicted from the preceding n-1 words. The main problem in building these models is that the vast majority of word combinations occur very infrequently, making it difficult to estimate accurate probabilities of words in most contexts. Researchers in statistical language modeling have developed a variety of smoothing tech- niques to alleviate this problem of data sparseness. Most smoothing methods are based on simple back-off formulas or interpolation schemes that discount the probability of observed events and assign the "leftover" probability mass to events unseen in training [7]. Unfortu- nately, these methods do not typically represent or take advantage of statistical regularities among contexts. One expects the probabilities of rare or unseen events in one context to be related to their probabilities in statistically similar contexts. Thus, it should be possible to estimate more accurate probabilities by exploiting these regularities. Several approaches have been suggested for sharing statistical information across contexts. The aggregate Markov model (AMM) of Saul and Pereira [13] (also discussed by Hofmann and Puzicha [10] as a special case of the aspect model) factors the conditional probability table of a word given its context by a latent variable representing context "classes". How- ever, this latent variable approach is difficult to generalize to multiword contexts, as the size of the conditional probability table for class given context grows exponentially with the context length. The neural probabilistic language model (NPLM) of Bengio et al. [2, 3] achieved signifi- cant improvements over state-of-the-art smoothed n-gram models [6]. The NPLM encodes contexts as low-dimensional continuous vectors. These are fed to a multilayer neural net- work that outputs a probability distribution over words. The low-dimensional vectors and the parameters of the network are trained simultaneously to minimize the perplexity of the language model. This model has no difficulty encoding multiword contexts, but its training and application are very costly because of the need to compute a separate normalization for the conditional probabilities associated to each context. In this paper, we introduce and evaluate a statistical language model that combines the advantages of the AMM and NPLM. Like the NPLM, it can be used for multiword con- texts, and like the AMM it avoids per-context normalization. In our model, contexts are represented as low-dimensional real vectors initialized by unsupervised algorithms for di- mensionality reduction [14]. The probabilities of words given contexts are represented by a hierarchical mixture of experts (HME) [12], where each expert is a multinomial distri- bution over predicted words. This tree-structured mixture model allows a rich dependency on context without expensive per-context normalization. Proper initialization of the dis- tributed representations is crucial; in particular, we find that initializations from the results of linear and nonlinear dimensionality reduction algorithms lead to better models (with significantly lower test perplexities) than random initialization. In practice our model is several orders of magnitude faster to train and apply than the NPLM, enabling us to work with larger vocabularies and training corpora. We present re- sults on a large-scale bigram modeling task, showing that our model also leads to significant improvements over comparable AMMs. 2 Distributed representations of words Natural language has complex, multidimensional semantics. As a trivial example, consider the following four sentences: The vase broke. The vase contains water. The window broke. The window contains water. The bottom right sentence is syntactically valid but semantically meaningless. As shown by the table, a two-bit distributed representation of the words "vase" and "window" suffices to express that a vase is both a container and breakable, while a window is breakable but can- not be a container. More generally, we expect low dimensional continuous representations of words to be even more effective at capturing semantic regularities. Distributed representations of words can be derived in several ways. In a given corpus of text, for example, consider the matrix of bigram counts whose element Cij records the number of times that word wj follows word wi. Further, let pij = Cij/ C k ik denote the conditional frequencies derived from these counts, and let pi denote the V -dimensional frequency vector with elements pij, where V is the vocabulary size. Note that the vectors pi themselves provide a distributed representation of the words wi in the corpus. For large vocabularies and training corpora, however, this is an extremely unwieldy representation, tantamount to storing the full matrix of bigram counts. Thus, it is natural to seek a lower dimensional representation that captures the same information. To this end, we need to map each vector pi to some d-dimensional vector xi, with d V . We consider two methods in dimensionality reduction for this problem. The results from these methods are then used to initialize the HME architecture in the next section. 2.1 Linear dimensionality reduction The simplest form of dimensionality reduction is principal component analysis (PCA). PCA computes a linear projection of the frequency vectors pi into the low dimensional subspace that maximizes their variance. The variance-maximizing subspace of dimension- ality d is spanned by the top d eigenvectors of the frequency vector covariance matrix. The eigenvalues of the covariance matrix measure the variance captured by each axis of the subspace. The effect of PCA can also be understood as a translation and rotation of the frequency vectors pi, followed by a truncation that preserves only their first d elements. 2.2 Nonlinear dimensionality reduction Intuitively, we would like to map the vectors pi into a low dimensional space where se- mantically similar words remain close together and semantically dissimilar words are far apart. Can we find a nonlinear mapping that does this better than PCA? Weinberger et al. recently proposed a new solution to this problem based on semidefinite programming [14]. Let xi denote the image of pi under this mapping. The mapping is discovered by first learning the V V matrix of Euclidean squared distances [1] given by Dij = |xi - xj|2. This is done by balancing two competing goals: (i) to co-locate semantically similar words, and (ii) to separate semantically dissimilar words. The first goal is achieved by fixing the distances between words with similar frequency vectors to their original values. In particu- lar, if pj and pk lie within some small neighborhood of each other, then the corresponding element Djk in the distance matrix is fixed to the value |pj - pk|2. The second goal is achieved by maximizing the sum of pairwise squared distances ijDij. Thus, we push the words in the vocabulary as far apart as possible subject to the constraint that the distances between semantically similar words do not change. The only freedom in this optimization is the criterion for judging that two words are se- mantically similar. In practice, we adopt a simple criterion such as k-nearest neighbors in the space of frequency vectors pi and choose k as small as possible so that the resulting neighborhood graph is connected [14]. The optimization is performed over the space of Euclidean squared distance matrices [1]. Necessary and sufficient conditions for the matrix D to be interpretable as a Euclidean squared distance matrix are that D is symmetric and that the Gram matrix1 derived from G = - 1 HDHT is semipositive definite, where H = I - 1 11T. The optimization can 2 V thus be formulated as the semidefinite programming problem: Maximize ijDij subject to: (i) DT = D, (ii) - 1 HDH 0, and 2 (iii) Dij = |pi - pj|2 for all neighboring vectors pi and pj. 1Assuming without loss of generality that the vectors xi are centered on the origin, the dot prod- ucts Gij = xi xj are related to the pairwise squared distances Dij = |xi - xj |2 as stated above.
John Blitzer, Kilian Q. Weinberger, Lawrence K. Saul, Fernando Pereira 0003
NIPS1
2003 Evaluation Challenges in Large-Scale Document Summarization
abstract
We present a large-scale meta evaluation of eight evaluation measures for both single-document and multi-document summarizers. To this end we built a corpus consisting of (a) 100 Million automatic summaries using six summarizers and baselines at ten summary lengths in both English and Chinese, (b) more than 10,000 manual abstracts and extracts, and (c) 200 Million automatic document and summary retrievals using 20 queries. We present both qualitative and quantitative results showing the strengths and draw-backs of all evaluation methods and how they rank the different summarizers.
Dragomir R. Radev, Simone Teufel, Horacio Saggion, Wai Lam, John Blitzer, Arda Çelebi, Danyu Liu, Elliott Drábek
ACL5
2003 Summarizing archived discussions: a beginning
abstract
This paper describes an approach to digesting threads of archived discussion lists by clustering messages into approximate topical groups, and then extracting shorter overviews, and longer summaries for each group.
Paula S. Newman, John Blitzer
IUI2