Tom Minka

dblp:m/ThomasPMinka · also Thomas P. Minka · DBLP profile ↗
← Back
37ranked-venue papers
8as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 8 first-authorGraphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-authorDatabases, data management, data science and information retrieval · 5Computer networks · 4Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
16 papers
Probabilistic and Bayesian machine learning · 83% Segmentation and scene understanding · 5% Image recognition and object detection · 5%
Databases, data mining, and information retrieval
7 papers
Information retrieval · 92% Data mining · 4% Data integration and cleaning · 4%
Computer networks
2 papers
Wireless sensing and localization · 67% Wireless networking · 27% Internet of things and sensor networks · 6%
Computer graphics and multimedia
4 papers
Computational photography and imaging · 70% Image and video processing · 15% Multimedia analysis and retrieval · 15%

Topics — the 30 heaviest of 63, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Wireless sensing and localization
indoor localization
0.322012
You are facing the Mona Lisa: spot localization using PHY layer information · MobiSys 2012
Precise indoor localization using PHY information · MobiSys 2011
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
variational message passing
0.222011
Non-conjugate Variational Message Passing for Multinomial and Binary Regression · NIPS 2011
Gates · NIPS 2008
Information retrieval › user behavior › search behavior
click model
0.222010
A novel click model and its applications to online advertising · WSDM 2010
Click chain model in web search · WWW 2009
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
approximate inference
0.242008
Gates · NIPS 2008
Predictive automatic relevance determination by expectation propagation · ICML 2004
Tree-structured Approximations by Expectation Propagation · NIPS 2003
Machine learning › Probabilistic and Bayesian machine learning
monte carlo methods
0.212014
A* Sampling · NIPS 2014
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
bayesian graphical model
0.112012
How To Grade a Test Without Knowing the Answers - A Bayesian Graphical Model for Adaptive Crowdsourcing and Aptitude Testing · ICML 2012
Wireless sensing and localization › indoor localization › wifi fingerprinting
CSI fingerprinting
0.112012
You are facing the Mona Lisa: spot localization using PHY layer information · MobiSys 2012
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models
0.122008
Gates · NIPS 2008
Diagram Structure Recognition by Bayesian Conditional Random Fields · CVPR (2) 2005
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression › probabilistic regression
bayesian regression
0.112011
Non-conjugate Variational Message Passing for Multinomial and Binary Regression · NIPS 2011
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression › generalized linear model › logistic regression
multinomial logistic regression
0.112011
Non-conjugate Variational Message Passing for Multinomial and Binary Regression · NIPS 2011
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.112011
Non-conjugate Variational Message Passing for Multinomial and Binary Regression · NIPS 2011
Wireless networking › WLAN
physical-layer wi-fi
0.112011
Precise indoor localization using PHY information · MobiSys 2011
Computational photography and imaging
color constancy
0.122008
Bayesian color constancy revisited · CVPR 2008
Bayesian Color Constancy with Non-Gaussian Models · NIPS 2003
Computational photography and imaging › color constancy
illuminant estimation
0.122008
Bayesian color constancy revisited · CVPR 2008
Bayesian Color Constancy with Non-Gaussian Models · NIPS 2003
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
expectation propagation
0.132007
Predictive automatic relevance determination by expectation propagation · ICML 2004
Tree-structured Approximations by Expectation Propagation · NIPS 2003
TrueSkill Through Time: Revisiting the History of Chess · NIPS 2007
Information retrieval › ranking
relevance estimation
0.112009
Click chain model in web search · WWW 2009
Information retrieval
retrieval models
0.122008
SoftRank: optimizing non-smooth rank metrics · WSDM 2008
The Bayesian image retrieval system, PicHunter: theory, implementation, and psychophysical experiments · IEEE Trans. Image Process. 2000
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
factor graphs
0.112008
Gates · NIPS 2008
Information retrieval
evaluation
0.112008
SoftRank: optimizing non-smooth rank metrics · WSDM 2008
Information retrieval › ranking
learning to rank
0.112008
SoftRank: optimizing non-smooth rank metrics · WSDM 2008
Information retrieval › retrieval evaluation
ranking evaluation
0.112008
SoftRank: optimizing non-smooth rank metrics · WSDM 2008
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.132008
Bayesian Color Constancy with Non-Gaussian Models · NIPS 2003
Bayesian color constancy revisited · CVPR 2008
Diagram Structure Recognition by Bayesian Conditional Random Fields · CVPR (2) 2005
Computer vision › Segmentation and scene understanding › image segmentation
co-segmentation
0.112006
Cosegmentation of Image Pairs by Histogram Matching - Incorporating a Global Constraint into MRFs · CVPR (1) 2006
Machine learning › Probabilistic and Bayesian machine learning
generative and discriminative models
0.112006
Principled Hybrids of Generative and Discriminative Models · CVPR (1) 2006
Computer vision › Segmentation and scene understanding
image segmentation
0.112006
Cosegmentation of Image Pairs by Histogram Matching - Incorporating a Global Constraint into MRFs · CVPR (1) 2006
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
markov random field inference
0.112006
Cosegmentation of Image Pairs by Histogram Matching - Incorporating a Global Constraint into MRFs · CVPR (1) 2006
Computer vision › Image recognition and object detection
object recognition
0.112006
Principled Hybrids of Generative and Discriminative Models · CVPR (1) 2006
Machine learning › Learning paradigms
semi-supervised learning
0.112006
Principled Hybrids of Generative and Discriminative Models · CVPR (1) 2006
Machine learning › Probabilistic and Bayesian machine learning › structured prediction
conditional random field
0.112005
Diagram Structure Recognition by Bayesian Conditional Random Fields · CVPR (2) 2005
Computer vision › Image recognition and object detection › image classification
object classification
0.112005
Object Categorization by Learned Universal Visual Dictionary · ICCV 2005

Methods — techniques the papers use, named apart from their topics

bayesian inference · 0.6rejection sampling · 0.4a* sampling · 0.4graphical model · 0.3expectation propagation · 0.2bayesian network · 0.2variational message passing · 0.2grey-world algorithm · 0.2bayesian estimation · 0.2channel response analysis · 0.1OFDM subcarrier signature extraction · 0.1variational bayes · 0.1softmax bound · 0.1location fingerprinting · 0.1OFDM subcarrier aggregation · 0.1smoothed approximation · 0.1gradient optimization · 0.1visual dictionary · 0.1
YearPublicationVenuePosition
2014 A* Sampling
Chris J. Maddison, Daniel Tarlow, Tom Minka
NIPS3
2012 How To Grade a Test Without Knowing the Answers - A Bayesian Graphical Model for Adaptive Crowdsourcing and Aptitude Testing
Yoram Bachrach, Thore Graepel, Tom Minka, John Guiver
ICML3
2012 You are facing the Mona Lisa: spot localization using PHY layer information
abstract
This paper explores the viability of precise indoor localization using physical layer information in WiFi systems. We find evidence that channel responses from multiple OFDM subcarriers can be a promising location signature. While these signatures certainly vary over time and environmental mobility, we notice that their core structure preserves certain properties that are amenable to localization. We attempt to harness these opportunities through a functional system called PinLoc, implemented on off-the-shelf Intel 5300 cards. We evaluate the system in a busy engineering building, a crowded student center, a cafeteria, and at the Duke University museum, and demonstrate localization accuracies in the granularity of 1m x 1m boxes, called "spots". Results from 100 spots show that PinLoc is able to localize users to the correct spot with 89% mean accuracy, while incurring less than 6% false positives. We believe this is an important step forward, compared to the best indoor localization schemes of today, such as Horus.
Souvik Sen, Bozidar Radunovic, Romit Roy Choudhury, Tom Minka
MobiSys4
2011 Precise indoor localization using PHY layer information
abstract
This paper shows the viability of precise indoor localization using physical layer information in WiFi systems. We find that channel frequency responses across multiple OFDM sub-carriers can be suitably aggregated into a location fingerprint. While these fingerprints vary over time and environmental mobility, we notice that their core structure preserves certain properties that are amenable to localization. We demonstrate these ideas through a functional prototype, implemented on off-the-shelf Intel 5300 cards (that export per-subcarrier information to the driver). We evaluate the prototype using the existing APs inside a busy building, a cafeteria, and a museum, and demonstrate localization accuracies in the granularity of 1m × 1m boxes, called spots. Results show that our system, PinLoc, is able to localize users to a spot with 90% mean accuracy, while incurring less than 6% false positives. We believe this holds promise towards an important development in indoor localization.
Souvik Sen, Romit Roy Choudhury, Bozidar Radunovic, Tom Minka
HotNets4
2011 Precise indoor localization using PHY information
abstract
This poster shows the viability of precise indoor localization using physical layer information in WiFi systems. We find that channel frequency responses across multiple OFDM subcarriers can be suitably aggregated into a location signature. While these signatures vary over time and environmental mobility, we notice that their core structure preserves certain properties that are amenable to localization. We demonstrate these ideas through a functional system, implemented on off-the-shelf Intel 5300 cards (that are designed to export per-subcarrier information to the driver). We evaluate the system in a busy engineering building, a cafeteria, and the university museum, and demonstrate localization accuracies in the granularity of 1m x 1m boxes, called spots. Results show that our system, PinLoc, is able to localize users to a spot with 90% mean accuracy, while incurring less than 6% false positives. We believe this is an important step forward, compared to the best indoor localization schemes of today.
Souvik Sen, Bozidar Radunovic, Romit Roy Choudhury, Tom Minka
MobiSys4
2011 Non-conjugate Variational Message Passing for Multinomial and Binary Regression
abstract
Variational Message Passing (VMP) is an algorithmic implementation of the Variational Bayes (VB) method which applies only in the special case of conjugate exponential family models. We propose an extension to VMP, which we refer to as Non-conjugate Variational Message Passing (NCVMP) which aims to alleviate this restriction while maintaining modularity, allowing choice in how expectations are calculated, and integrating into an existing message-passing framework: Infer.NET. We demonstrate NCVMP on logistic binary and multinomial regression. In the multinomial case we introduce a novel variational bound for the softmax factor which is tighter than other commonly used bounds whilst maintaining computational tractability.
David A. Knowles, Tom Minka
NIPS2
2010 Sparse-posterior Gaussian Processes for general likelihoods
Yuan Qi 0001, Ahmed H. Abdel-Gawad, Tom Minka
UAI3
2010 A novel click model and its applications to online advertising
abstract
Recent advances in click model have positioned it as an attractive method for representing user preferences in web search and online advertising. Yet, most of the existing works focus on training the click model for individual queries, and cannot accurately model the tail queries due to the lack of training data. Simultaneously, most of the existing works consider the query, url and position, neglecting some other important attributes in click log data, such as the local time. Obviously, the click through rate is different between daytime and midnight. In this paper, we propose a novel click model based on Bayesian network, which is capable of modeling the tail queries because it builds the click model on attribute values, with those values being shared across queries. We called our work General Click Model (GCM) as we found that most of the existing works can be special cases of GCM by assigning different parameters. Experimental results on a large-scale commercial advertisement dataset show that GCM can significantly and consistently lead to better results as compared to the state-of-the-art works.
Zeyuan Allen Zhu, Weizhu Chen, Tom Minka, Zheng Chen 0001
WSDM3
2009 Virtual Vector Machine for Bayesian Online Classification
Tom Minka, Rongjing Xiang, Yuan Qi 0001
UAI1
2009 Click chain model in web search
abstract
Given a terabyte click log, can we build an efficient and effective click model? It is commonly believed that web search click logs are a gold mine for search business, because they reflect users' preference over web documents presented by the search engine. Click models provide a principled approach to inferring user-perceived relevance of web documents, which can be leveraged in numerous applications in search businesses. Due to the huge volume of click data, scalability is a must.We present the click chain model (CCM), which is based on a solid, Bayesian framework. It is both scalable and incremental, perfectly meeting the computational challenges imposed by the voluminous click logs that constantly grow. We conduct an extensive experimental study on a data set containing 8.8 million query sessions obtained in July 2008 from a commercial search engine. CCM consistently outperforms two state-of-the-art competitors in a number of metrics, with over 9.7% better log-likelihood, over 6.2% better click perplexity and much more robust (up to 30%) prediction of the first and the last clicked position.
Fan Guo 0006, Chao Liu 0001, Anitha Kannan, Tom Minka, Michael J. Taylor 0001, Yi Min Wang, Christos Faloutsos
WWW4
2008 Bayesian color constancy revisited
abstract
Computational color constancy is the task of estimating the true reflectances of visible surfaces in an image. In this paper we follow a line of research that assumes uniform illumination of a scene, and that the principal step in estimating reflectances is the estimation of the scene illuminant. We review recent approaches to illuminant estimation, firstly those based on formulae for normalisation of the reflectance distribution in an image - so-called grey-world algorithms, and those based on a Bayesian formulation of image formation. In evaluating these previous approaches we introduce a new tool in the form of a database of 568 high-quality, indoor and outdoor images, accurately labelled with illuminant, and preserved in their raw form, free of correction or normalisation. This has enabled us to establish several properties experimentally. Firstly automatic selection of grey-world algorithms according to image properties is not nearly so effective as has been thought. Secondly, it is shown that Bayesian illuminant estimation is significantly improved by the improved accuracy of priors for illuminant and reflectance that are obtained from the new dataset.
Peter V. Gehler, Carsten Rother, Andrew Blake 0001, Tom Minka, Toby Sharp
CVPR4
2008 Gates
abstract
Gates are a new notation for representing mixture models and context-sensitive independence in factor graphs. Factor graphs provide a natural representation for message-passing algorithms, such as expectation propagation. However, message passing in mixture models is not well captured by factor graphs unless the entire mixture is represented by one factor, because the message equations have a containment structure. Gates capture this containment structure graphically, allowing both the independences and the message-passing equations for a model to be readily visualized. Different variational approximations for mixture models can be understood as different ways of drawing the gates in a model. We present general equations for expectation propagation and variational message passing in the presence of gates.
Tom Minka, John M. Winn
NIPS1
2008 SoftRank: optimizing non-smooth rank metrics
abstract
We address the problem of learning large complex ranking functions. Most IR applications use evaluation metrics that depend only upon the ranks of documents. However, most ranking functions generate document scores, which are sorted to produce a ranking. Hence IR metrics are innately non-smooth with respect to the scores, due to the sort. Unfortunately, many machine learning algorithms require the gradient of a training objective in order to perform the optimization of the model parameters,and because IR metrics are non-smooth,we need to find a smooth proxy objective that can be used for training. We present a new family of training objectives that are derived from the rank distributions of documents, induced by smoothed scores. We call this approach SoftRank. We focus on a smoothed approximation to Normalized Discounted Cumulative Gain (NDCG), called SoftNDCG and we compare it with three other training objectives in the recent literature. We present two main results. First, SoftRank yields a very good way of optimizing NDCG. Second, we show that it is possible to achieve state of the art test set NDCG results by optimizing a soft NDCG objective on the training set with a different discount function
Michael J. Taylor 0001, John Guiver, Stephen E. Robertson, Tom Minka
WSDM4
2007 TrueSkill Through Time: Revisiting the History of Chess
abstract
We extend the Bayesian skill rating system TrueSkill to infer entire time series of skills of players by smoothing through time instead of (cid:12)ltering. The skill of each participating player, say, every year is represented by a latent skill variable which is a(cid:11)ected by the relevant game outcomes that year, and coupled with the skill variables of the previous and subsequent year. Inference in the resulting factor graph is carried out by approximate message passing (EP) along the time series of skills. As before the system tracks the uncertainty about player skills, explicitly models draws, can deal with any number of competing entities and can infer individual skills from team results. We extend the system to estimate player-speci(cid:12)c draw mar- gins. Based on these models we present an analysis of the skill curves of important players in the history of chess over the past 150 years. Results include plots of players’ lifetime skill development as well as the ability to compare the skills of di(cid:11)erent players across time. Our results indicate that a) the overall playing strength has increased over the past 150 years, and b) that modelling a player’s ability to force a draw provides signi(cid:12)cantly better predictive power.
Pierre Dangauthier, Ralf Herbrich, Tom Minka, Thore Graepel
NIPS3
2007 Window-based expectation propagation for adaptive signal detection in flat-fading channels
abstract
In this paper, we propose a new Bayesian receiver for signal detection in flat-fading channels. First, the detection problem is formulated as an inference problem in a graphical model that models a hybrid dynamic system with both continuous and discrete variables. Then, based on the expectation propagation (EP) framework, we develop a smoothing algorithm to address the inference problem and visualize this algorithm using factor graphs. As a generalization of loopy belief propagation, EP efficiently approximates Bayesian estimation by iteratively propagating information between different nodes in the graphical model and projecting the posterior distributions into the exponential family. We use window-based EP smoothing for online estimation as in the signal detection problem. Window-based EP smoothing achieves accuracy similar to that obtained by batch EP smoothing, as shown in our simulations, while reducing delay time. Compared to sequential Monte Carlo filters and smoothers, the new method has lower computational complexity since it makes analytically deterministic approximation instead of Monte Carlo approximations. Our simulations demonstrate that the new receiver achieves accurate detection without the aid of any training symbols or decision feedbacks. Furthermore, the new receiver achieves accuracy comparable to that achieved by sequential Monte Carlo methods, but with less than one-tenth computational cost.
Yuan Qi 0001, Tom Minka
IEEE Trans. Wirel. Commun.2
2006 Principled Hybrids of Generative and Discriminative Models
abstract
When labelled training data is plentiful, discriminative techniques are widely used since they give excellent generalization performance. However, for large-scale applications such as object recognition, hand labelling of data is expensive, and there is much interest in semi-supervised techniques based on generative models in which the majority of the training data is unlabelled. Although the generalization performance of generative models can often be improved by ‘training them discriminatively’, they can then no longer make use of unlabelled data. In an attempt to gain the benefit of both generative and discriminative approaches, heuristic procedure have been proposed [2, 3] which interpolate between these two extremes by taking a convex combination of the generative and discriminative objective functions. In this paper we adopt a new perspective which says that there is only one correct way to train a given model, and that a ‘discriminatively trained’ generative model is fundamentally a new model [7]. From this viewpoint, generative and discriminative models correspond to specific choices for the prior over parameters. As well as giving a principled interpretation of ‘discriminative training’, this approach opens door to very general ways of interpolating between generative and discriminative extremes through alternative choices of prior. We illustrate this framework using both synthetic data and a practical example in the domain of multi-class object recognition. Our results show that, when the supply of labelled training data is limited, the optimum performance corresponds to a balance between the purely generative and the purely discriminative.
Julia A. Lasserre, Christopher M. Bishop, Tom Minka
CVPR (1)3
2006 Cosegmentation of Image Pairs by Histogram Matching - Incorporating a Global Constraint into MRFs
abstract
We introduce the term cosegmentation which denotes the task of segmenting simultaneously the common parts of an image pair. A generative model for cosegmentation is presented. Inference in the model leads to minimizing an energy with an MRF term encoding spatial coherency and a global constraint which attempts to match the appearance histograms of the common parts. This energy has not been proposed previously and its optimization is challenging and NP-hard. For this problem a novel optimization scheme which we call trust region graph cuts is presented. We demonstrate that this framework has the potential to improve a wide range of research: Object driven image retrieval, video tracking and segmentation, and interactive image editing. The power of the framework lies in its generality, the common part can be a rigid/non-rigid object (or scene), observed from different viewpoints or even similar objects of the same class.
Carsten Rother, Tom Minka, Andrew Blake 0001, Vladimir Kolmogorov
CVPR (1)2
2006 TrueSkillTM: A Bayesian Skill Rating System
Ralf Herbrich, Tom Minka, Thore Graepel
NIPS2
2005 Diagram Structure Recognition by Bayesian Conditional Random Fields
abstract
Hand-drawn diagrams present a complex recognition problem. Elements of the diagram are often individually ambiguous, and require context to be interpreted. We present a recognition method based on Bayesian conditional random fields (BCRFs) that jointly analyzes all drawing elements in order to incorporate contextual cues. The classification of each object affects the classification of its neighbors. BCRFs allow flexible and correlated features, and take both spatial and temporal information into account. BCRFs estimate the posterior distribution of parameters during training, and average predictions over the posterior for testing. As a result of model averaging, BCRFs avoid the overfitting problems associated with maximum likelihood training. We also incorporate automatic relevance determination (ARD), a Bayesian feature selection technique, into BCRFs. The result is significantly lower error rates compared to ML- and MAP-trained CRFs.
Yuan Qi 0001, Martin Szummer, Tom Minka
CVPR (2)3
2005 Object Categorization by Learned Universal Visual Dictionary
abstract
This paper presents a new algorithm for the automatic recognition of object classes from images (categorization). Compact and yet discriminative appearance-based object class models are automatically learned from a set of training images. The method is simple and extremely fast, making it suitable for many applications such as semantic image retrieval, Web search, and interactive image editing. It classifies a region according to the proportions of different visual words (clusters in feature space). The specific visual words and the typical proportions in each object are learned from a segmented training set. The main contribution of this paper is twofold: i) an optimally compact visual dictionary is learned by pair-wise merging of visual words from an initially large dictionary. The final visual words are described by GMMs. ii) A novel statistical measure of discrimination is proposed which is optimized by each merge operation. High classification accuracy is demonstrated for nine object classes on photographs of real objects viewed under general lighting conditions, poses and viewpoints. The set of test images used for validation comprise: i) photographs acquired by us, ii) images from the Web and iii) images from the recently released Pascal dataset. The proposed algorithm performs well on both texture-rich objects (e.g. grass, sky, trees) and structure-rich ones (e.g. cars, bikes, planes)
John M. Winn, Antonio Criminisi, Tom Minka
ICCV3
2005 Structured Region Graphs: Morphing EP into GBP
Max Welling, Tom Minka, Yee Whye Teh
UAI2
2004 Predictive automatic relevance determination by expectation propagation
abstract
In many real-world classification problems the input contains a large number of potentially ir-relevant features. This paper proposes a new Bayesian framework for determining the rele-vance of input features. This approach extends one of the most successful Bayesian methods for feature selection and sparse learning, known as Automatic Relevance Determination (ARD). ARD finds the relevance of features by optimiz-ing the model marginal likelihood, also known as the evidence. We show that this can lead to over-fitting. To address this problem, we propose Pre-dictive ARD based on estimating the predictive performance of the classifier. While the actual leave-one-out predictive performance is generally very costly to compute, the expectation propaga-tion (EP) algorithm proposed by Minka provides an estimate of this predictive performance as a side-effect of its iterations. We exploit this in our algorithm to do feature selection, and to select data points in a sparse Bayesian kernel classifier. Moreover, we provide two other improvements to previous algorithms, by replacing Laplace’s approximation with the generally more accurate EP, and by incorporating the fast optimization algorithm proposed by Faul and Tipping. Our experiments show that our method based on the EP estimate of predictive performance is more accurate on test data than relevance determina-tion by optimizing the evidence.
Yuan Qi 0001, Tom Minka, Rosalind W. Picard, Zoubin Ghahramani
ICML2
2003 Tree-structured Approximations by Expectation Propagation
abstract
Approximation structure plays an important role in inference on loopy graphs. As a tractable structure, tree approximations have been utilized in the variational method of Ghahramani & Jordan (1997) and the se- quential projection method of Frey et al. (2000). However, belief propa- gation represents each factor of the graph with a product of single-node messages. In this paper, belief propagation is extended to represent fac- tors with tree approximations, by way of the expectation propagation framework. That is, each factor sends a “message” to all pairs of nodes in a tree structure. The result is more accurate inferences and more fre- quent convergence than ordinary belief propagation, at a lower cost than variational trees or double-loop algorithms.
Tom Minka, Yuan Qi 0001
NIPS1
2003 Bayesian Color Constancy with Non-Gaussian Models
abstract
We present a Bayesian approach to color constancy which utilizes a non- Gaussian probabilistic model of the image formation process. The pa- rameters of this model are estimated directly from an uncalibrated image set and a small number of additional algorithmic parameters are chosen using cross validation. The algorithm is empirically shown to exhibit RMS error lower than other color constancy algorithms based on the Lambertian surface reflectance model when estimating the illuminants of a set of test images. This is demonstrated via a direct performance comparison utilizing a publicly available set of real world test images and code base.
Charles R. Rosenberg, Tom Minka, Alok Ladsariya
NIPS2
2002 Bayesian spectrum estimation of unevenly sampled nonstationary data
abstract
Spectral estimation methods typically assume stationarity and uniform spacing between samples of data. The non-stationarity of real data is usually accommodated by windowing methods, while the lack of uniformly-spaced samples is typically addressed by methods that “fill in” the data in some way. This paper presents a new approach to both of these problems: We use a non-stationary Kalman filter within a Bayesian framework to jointly estimate all spectral coefficients instantaneously. The new method works regardless of how the signal samples are spaced. We illustrate the method on several data sets, showing that it provides more accurate estimation than the Lomb-Scargle method and several classical spectral estimation methods.
Yuan Qi 0001, Tom Minka, Rosalind W. Picard
ICASSP2
2002 Novelty and redundancy detection in adaptive filtering
abstract
This paper addresses the problem of extending an adaptive information filtering system to make decisions about the novelty and redundancy of relevant documents. It argues that relevance and redundance should each be modelled explicitly and separately. A set of five redundancy measures are proposed and evaluated in experiments with and without redundancy thresholds. The experimental results demonstrate that the cosine similarity metric and a redundancy measure based on a mixture of language models are both effective for identifying redundant documents.
Yi Zhang 0001, Jamie Callan, Tom Minka
SIGIR3
2002 Expectation-Propogation for the Generative Aspect Model
Tom Minka, John D. Lafferty
UAI1
2001 Document Image Decoding Using Iterated Complete Path Search with Subsampled Heuristic Scoring
abstract
It has been shown that the computation time of document image decoding can be significantly reduced by employing heuristics in the search for the best decoding of a text line. In the Iterated Complete Path (ICP) method, template matches are performed only along the best path found by dynamic programming on each iteration. When the best path stabilizes, the decoding is optimal and no more template matches need to be performed. In this way, only a tiny fraction of potential template matches must be evaluated, and the computation time is typically dominated by the evaluation of the initial heuristic upper bound for each template at each location in the image. The time to compute this bound depends on the resolution at which the matching scores are found. At lower resolution, the heuristic computation is reduced, but because a weaker bound is used, the number of Viterbi iterations is increased. We present the optimal (lowest upper-bound) heuristic for any degree of subsampling of multilevel template and/or interpolation, for use in text line decoding with ICP. The optimal degree of subsampling depends on image quality, but it is typically found that a small amount of template subsampling is effective in reducing the overall decoding time.
Dan S. Bloomberg, Kris Popat, Tom Minka
ICDAR3
2001 Expectation Propagation for approximate Bayesian inference
Tom Minka
UAI1
2000 Automatic Choice of Dimensionality for PCA
abstract
A central issue in principal component analysis (PCA) is choosing the number of principal components to be retained. By interpreting PCA as density estimation, we show how to use Bayesian model selection to es(cid:173) timate the true dimensionality of the data. The resulting estimate is sim(cid:173) ple to compute yet guaranteed to pick the correct dimensionality, given enough data. The estimate involves an integral over the Steifel manifold of k-frames, which is difficult to compute exactly. But after choosing an appropriate parameterization and applying Laplace's method, an accu(cid:173) rate and practical estimator is obtained. In simulations, it is convincingly better than cross-validation and other proposed algorithms, plus it runs much faster.
Tom Minka
NIPS1
2000 The Bayesian image retrieval system, PicHunter: theory, implementation, and psychophysical experiments
abstract
This paper presents the theory, design principles, implementation and performance results of PicHunter, a prototype content-based image retrieval (CBIR) system. In addition, this document presents the rationale, design and results of psychophysical experiments that were conducted to address some key issues that arose during PicHunter's development. The PicHunter project makes four primary contributions to research on CBIR. First, PicHunter represents a simple instance of a general Bayesian framework which we describe for using relevance feedback to direct a search. With an explicit model of what users would do, given the target image they want, PicHunter uses Bayes's rule to predict the target they want, given their actions. This is done via a probability distribution over possible image targets, rather than by refining a query. Second, an entropy-minimizing display algorithm is described that attempts to maximize the information obtained from a user at each iteration of the search. Third, PicHunter makes use of hidden annotation rather than a possibly inaccurate/inconsistent annotation structure that the user must learn and make queries in. Finally, PicHunter introduces two experimental paradigms to quantitatively evaluate the performance of the system, and psychophysical experiments are presented that support the theoretical claims.
Ingemar J. Cox, Matthew L. Miller, Tom Minka, Thomas V. Papathomas, Peter N. Yianilos
IEEE Trans. Image Process.3
2000 Correction to "the Bayesian image retrieval system, pichunter: theory, implementation, and psychophysical experiments"
Ingemar J. Cox, Matthew L. Miller, Tom Minka, Thomas V. Papathomas, Peter N. Yianilos
IEEE Trans. Image Process.3
1998 An Optimized Interaction Strategy for Bayesian Relevance Feedback
abstract
A new algorithm and systematic evaluation is presented for searching a database via relevance feedback. It represents a new image display strategy for the PicHunter system. The algorithm takes feedback in the form of relative judgments ("item A is more relevant than item B") as opposed to the stronger assumption of categorical relevance judgments ("item A is relevant but item B is not"). It also exploits a learned probabilistic model of human behavior to make better use of the feedback it obtains. The algorithm can be viewed as an extension of indexing schemes like the k-d tree to a stochastic setting, hence the name "stochastic-comparison search." In simulations, the amount of feedback required for the new algorithm scales like log/sub 2/ |D|, where |D| is the size of the database, while a simple query-by-example approach scales like |D|/sup /spl alpha//, where /spl alpha/<1 depends on the structure of the database. This theoretical advantage is reflected by experiments with real users on a database of 1500 stock photographs.
Ingemar J. Cox, Matthew L. Miller, Tom Minka, Peter N. Yianilos
CVPR3
1997 Interactive learning with a "society of models"
Tom Minka, Rosalind W. Picard
Pattern Recognit.1
1996 Interactive Learning with a "Society of Models"
abstract
Digital library access is driven by features, but the relevance of a feature for a query is not always obvious. This paper describes an approach for integrating a large number of context-dependent features into a semi-automated tool. Instead of requiring universal similarity measures or manual selection of relevant features, the approach provides a learning algorithm for selecting and combining groupings of the data, where groupings can be induced by highly specialized features. The selection process is guided by positive and negative examples from the user. The inherent combinatorics of using multiple features is reduced by a multistage grouping generation, weighting, and collection process. The stages closest to the user are trained fastest and slowly propagate their adaptations back to earlier stages. The weighting stage adapts the collection stage's search space across uses, so that, in later interactions, good groupings are found given few examples from the user.
Tom Minka, Rosalind W. Picard
CVPR1
1996 Modeling user subjectivity in image libraries
abstract
In addition to the problem of which image analysis models to use in digital libraries, e.g. wavelet, Wold, color histograms, is the problem of how to combine these models with their different strengths. Most present systems place the burden of combination on the user, e.g. the user specifies 50% texture features, 20% color features, etc. This is a problem since most users do not know how to best pick the settings for the given data and search problem. The paper addresses this problem, describing research in progress for a system that: (1) automatically infers which combination of models best represents the data of interest to the user; and (2) learns continuously during interaction with each user. In particular, these two components-inference and learning-provide a solution that adapts to the subjective and hard to predict behaviors frequently seen when people query or browse image libraries.
Rosalind W. Picard, Tom Minka, Martin Szummer
ICIP (2)2
1995 Vision Texture for Annotation
Rosalind W. Picard, Tom Minka
Multim. Syst.2