Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Scott Gaffney

dblp:31/486 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
0since 2021 · last 2010
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-authorDatabases, data management, data science and information retrieval · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Data mining · 80% Machine learning and data management · 20%
Artificial intelligence
3 papers
Probabilistic and Bayesian machine learning · 100%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning and data management
active learning
0.112010
A large-scale active learning system for topical categorization on the web · WWW 2010
Data mining › predictive modeling
classification
0.112010
A large-scale active learning system for topical categorization on the web · WWW 2010
Data mining › text mining › text classification
web page classification
0.112010
A large-scale active learning system for topical categorization on the web · WWW 2010
Data mining
clustering
0.132003
Translation-invariant mixture models for curve clustering · KDD 2003
A general probabilistic framework for clustering individuals and objects · KDD 2000
Trajectory Clustering with Mixtures of Regression Models · KDD 1999
Machine learning › Probabilistic and Bayesian machine learning › clustering
model-based clustering
0.012004
Joint Probabilistic Curve Clustering and Alignment · NIPS 2004
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.012003
Translation-invariant mixture models for curve clustering · KDD 2003
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
expectation-maximization
0.012003
Translation-invariant mixture models for curve clustering · KDD 2003
Data mining › clustering › model-based clustering
mixture model clustering
0.012003
Translation-invariant mixture models for curve clustering · KDD 2003
Data mining › clustering
model-based clustering
0.012000
A general probabilistic framework for clustering individuals and objects · KDD 2000
Data mining › clustering
probabilistic clustering
0.012000
A general probabilistic framework for clustering individuals and objects · KDD 2000
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
mixture model
0.011999
Trajectory Clustering with Mixtures of Regression Models · KDD 1999
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
mixture of linear regressions
0.011999
Trajectory Clustering with Mixtures of Regression Models · KDD 1999
Data mining › clustering › sequence clustering
trajectory clustering
0.011999
Trajectory Clustering with Mixtures of Regression Models · KDD 1999
Bioinformatics and computational biology
gene expression analysis
0.012004
Joint Probabilistic Curve Clustering and Alignment · NIPS 2004

Methods — techniques the papers use, named apart from their topics

expectation-maximization · 0.2supervised learning · 0.1active learning · 0.1probabilistic alignment model · 0.1EM algorithm · 0.1bayesian estimation · 0.1kernel regression · 0.0
YearPublicationVenuePosition
2010 Resolving Surface Forms to Wikipedia Topics
Yiping Zhou, Lan Nie, Omid Rouhani-Kalleh, Flavian Vasile, Scott Gaffney
COLING5
2010 A large-scale active learning system for topical categorization on the web
abstract
Many web applications such as ad matching systems, vertical search engines, and page categorization systems require the identification of a particular type or class of pages on the Web. The sheer number and diversity of the pages on the Web, however, makes the problem of obtaining a good sample of the class of interest hard. In this paper, we describe a successfully deployed end-to-end system that starts from a biased training sample and makes use of several state-of-the-art machine learning algorithms working in tandem, including a powerful active learning component, in order to achieve a good classification system. The system is evaluated on traffic from a real-world ad-matching platform and is shown to achieve high categorization effectiveness with a significant reduction in editorial effort and labeling time.
Suju Rajan, Dragomir Yankov, Scott Gaffney, Adwait Ratnaparkhi
WWW3
2009 Improving web page classification by label-propagation over click graphs
abstract
In this paper, we present a semi-supervised learning method for web page classification, leveraging click logs to augment training data by propagating class labels to unlabeled similar documents. Current state-of-the-art classifiers are supervised and require large amounts of manually labeled data. We hypothesize that unlabeled documents similar to our positive and negative labeled documents tend to be clicked through by the same user queries. Our proposed method leverages this hypothesis and augments our training set by modeling the similarity between documents in a click graph. We experiment with three different web page classifiers and show empirical evidence that our proposed approach outperforms state-of-the-art methods and reduces the amount of human effort to label training data.
Patrick Pantel, Lei Duan, Scott Gaffney
CIKM4
2004 Joint Probabilistic Curve Clustering and Alignment
abstract
Clustering and prediction of sets of curves is an important problem in many areas of science and engineering. It is often the case that curves tend to be misaligned from each other in a continuous manner, either in space (across the measurements) or in time. We develop a probabilistic framework that allows for joint clustering and continuous alignment of sets of curves in curve space (as opposed to a fixed-dimensional feature- vector space). The proposed methodology integrates new probabilistic alignment models with model-based curve clustering algorithms. The probabilistic approach allows for the derivation of consistent EM learn- ing algorithms for the joint clustering-alignment problem. Experimental results are shown for alignment of human growth data, and joint cluster- ing and alignment of gene expression time-course data.
Scott Gaffney, Padhraic Smyth
NIPS1
2003 Translation-invariant mixture models for curve clustering
abstract
In this paper we present a family of algorithms that can simultaneously align and cluster sets of multidimensional curves defined on a discrete time grid. Our approach uses the Expectation-Maximization (EM) algorithm to recover both the mean curve shapes for each cluster, and the most likely shifts, offsets, and cluster memberships for each curve. We demonstrate how Bayesian estimation methods can improve the results for small sample sizes by enforcing smoothness in the cluster mean curves. We evaluate the methodology on two real-world data sets, time-course gene expression data and storm trajectory data. Experimental results show that models that incorporate curve alignment systematically provide improvements in predictive power and within-cluster variance on test data sets. The proposed approach provides a non-parametric, computationally efficient, and robust methodology for clustering broad classes of curve data.
Darya Chudova, Scott Gaffney, Eric Mjolsness, Padhraic Smyth
KDD2
2003 Probabilistic Models For Joint Clustering And Time-Warping Of Multidimensional Curves
Darya Chudova, Scott Gaffney, Padhraic Smyth
UAI2
2000 A general probabilistic framework for clustering individuals and objects
abstract
This paper presents a unifying probabilistic framework for clustering individuals or systems into groups when the available data measurements are not multiv ariate v ectors of xed dimensionality.For example, one might h a ve data from a set of medical patien ts,where for each patien tone has a set of of observed time-series, each time-series of potentially dierent length and dierent sampling rate.We propose a general model-based probabilistic framework for clustering data types of this form whic hare non-v ectorin nature and may vary in size from individual to individual.The Expectation-Maximization (EM) procedure for clustering within this framework is discussed and w e discuss ho w it be applied in a general manner to clustering of sequences, time-series, trajectories, and other non-vector data.We sho w that a number of earlier algorithms can be viewed as special cases within this unifying framework.The paper concludes with several illustrations of the method, including clustering of red blood cell data in a medical diagnosis context, clustering of proteins from curves of gene expression data, and clustering of individuals based on their sequences of Web na vigation.
Igor V. Cadez, Scott Gaffney, Padhraic Smyth
KDD2
1999 Trajectory Clustering with Mixtures of Regression Models
abstract
In this paper we address the problem of clustering trajectories, namely sets of short sequences of data measured as a function of a dependent variable such as time.Examples include storm path trajectories, longitudinal data such as drug therapy response, functional expression data in computational biology, and movements of objects or individuals in video sequences.Our clustering algorithm is based on a principled method for probabilistic modelhng of a set of trajectories as individual sequences of points generated from a finite mixture model consisting of regression model components.Unsupervised learning is carried out using maximum likelihood principles.Specifically, the EM algorithm is used to cope with the hidden data problem (i.e., the cluster memberships).We also develop generalizations of the method to handle non-parametric (kernel) regression components as well as multi-dimensional outputs.Simulation results comparing our method with other clustering methods such as K-means and Gaussian mixtures are presented as well as experimental results on real data sets.
Scott Gaffney, Padhraic Smyth
KDD1