VLDB 2026 Research / reviewers in the wild / expert
Scott Gaffney
dblp:31/486
· DBLP profile ↗
8ranked-venue papers
2as first author
0since 2021 · last 2010
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-authorDatabases, data management, data science and information retrieval · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
4 papers |
Data mining · 80% Machine learning and data management · 20% | |
| Artificial intelligence
3 papers |
Probabilistic and Bayesian machine learning · 100% |
Topics — the 14 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning and data management
active learning |
0.1 | 1 | 2010 | A large-scale active learning system for topical categorization on the web · WWW 2010 |
Data mining › predictive modeling
classification |
0.1 | 1 | 2010 | A large-scale active learning system for topical categorization on the web · WWW 2010 |
Data mining › text mining › text classification
web page classification |
0.1 | 1 | 2010 | A large-scale active learning system for topical categorization on the web · WWW 2010 |
Data mining
clustering |
0.1 | 3 | 2003 | Translation-invariant mixture models for curve clustering · KDD 2003 A general probabilistic framework for clustering individuals and objects · KDD 2000 Trajectory Clustering with Mixtures of Regression Models · KDD 1999 |
Machine learning › Probabilistic and Bayesian machine learning › clustering
model-based clustering |
0.0 | 1 | 2004 | Joint Probabilistic Curve Clustering and Alignment · NIPS 2004 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference |
0.0 | 1 | 2003 | Translation-invariant mixture models for curve clustering · KDD 2003 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
expectation-maximization |
0.0 | 1 | 2003 | Translation-invariant mixture models for curve clustering · KDD 2003 |
Data mining › clustering › model-based clustering
mixture model clustering |
0.0 | 1 | 2003 | Translation-invariant mixture models for curve clustering · KDD 2003 |
Data mining › clustering
model-based clustering |
0.0 | 1 | 2000 | A general probabilistic framework for clustering individuals and objects · KDD 2000 |
Data mining › clustering
probabilistic clustering |
0.0 | 1 | 2000 | A general probabilistic framework for clustering individuals and objects · KDD 2000 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
mixture model |
0.0 | 1 | 1999 | Trajectory Clustering with Mixtures of Regression Models · KDD 1999 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
mixture of linear regressions |
0.0 | 1 | 1999 | Trajectory Clustering with Mixtures of Regression Models · KDD 1999 |
Data mining › clustering › sequence clustering
trajectory clustering |
0.0 | 1 | 1999 | Trajectory Clustering with Mixtures of Regression Models · KDD 1999 |
Bioinformatics and computational biology
gene expression analysis |
0.0 | 1 | 2004 | Joint Probabilistic Curve Clustering and Alignment · NIPS 2004 |
Methods — techniques the papers use, named apart from their topics
expectation-maximization · 0.2supervised learning · 0.1active learning · 0.1probabilistic alignment model · 0.1EM algorithm · 0.1bayesian estimation · 0.1kernel regression · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2010 | Resolving Surface Forms to Wikipedia Topics
Yiping Zhou, Lan Nie, Omid Rouhani-Kalleh, Flavian Vasile, Scott Gaffney |
COLING | 5 |
| 2010 | A large-scale active learning system for topical categorization on the webabstractMany web applications such as ad matching systems, vertical search engines, and page categorization systems require the identification of a particular type or class of pages on the Web. The sheer number and diversity of the pages on the Web, however, makes the problem of obtaining a good sample of the class of interest hard. In this paper, we describe a successfully deployed end-to-end system that starts from a biased training sample and makes use of several state-of-the-art machine learning algorithms working in tandem, including a powerful active learning component, in order to achieve a good classification system. The system is evaluated on traffic from a real-world ad-matching platform and is shown to achieve high categorization effectiveness with a significant reduction in editorial effort and labeling time. Suju Rajan, Dragomir Yankov, Scott Gaffney, Adwait Ratnaparkhi |
WWW | 3 |
| 2009 | Improving web page classification by label-propagation over click graphsabstractIn this paper, we present a semi-supervised learning method for web page classification, leveraging click logs to augment training data by propagating class labels to unlabeled similar documents. Current state-of-the-art classifiers are supervised and require large amounts of manually labeled data. We hypothesize that unlabeled documents similar to our positive and negative labeled documents tend to be clicked through by the same user queries. Our proposed method leverages this hypothesis and augments our training set by modeling the similarity between documents in a click graph. We experiment with three different web page classifiers and show empirical evidence that our proposed approach outperforms state-of-the-art methods and reduces the amount of human effort to label training data. Patrick Pantel, Lei Duan, Scott Gaffney |
CIKM | 4 |
| 2004 | Joint Probabilistic Curve Clustering and AlignmentabstractClustering and prediction of sets of curves is an important problem in many areas of science and engineering. It is often the case that curves tend to be misaligned from each other in a continuous manner, either in space (across the measurements) or in time. We develop a probabilistic framework that allows for joint clustering and continuous alignment of sets of curves in curve space (as opposed to a fixed-dimensional feature- vector space). The proposed methodology integrates new probabilistic alignment models with model-based curve clustering algorithms. The probabilistic approach allows for the derivation of consistent EM learn- ing algorithms for the joint clustering-alignment problem. Experimental results are shown for alignment of human growth data, and joint cluster- ing and alignment of gene expression time-course data. Scott Gaffney, Padhraic Smyth |
NIPS | 1 |
| 2003 | Translation-invariant mixture models for curve clusteringabstractIn this paper we present a family of algorithms that can simultaneously align and cluster sets of multidimensional curves defined on a discrete time grid. Our approach uses the Expectation-Maximization (EM) algorithm to recover both the mean curve shapes for each cluster, and the most likely shifts, offsets, and cluster memberships for each curve. We demonstrate how Bayesian estimation methods can improve the results for small sample sizes by enforcing smoothness in the cluster mean curves. We evaluate the methodology on two real-world data sets, time-course gene expression data and storm trajectory data. Experimental results show that models that incorporate curve alignment systematically provide improvements in predictive power and within-cluster variance on test data sets. The proposed approach provides a non-parametric, computationally efficient, and robust methodology for clustering broad classes of curve data. Darya Chudova, Scott Gaffney, Eric Mjolsness, Padhraic Smyth |
KDD | 2 |
| 2003 | Probabilistic Models For Joint Clustering And Time-Warping Of Multidimensional Curves
Darya Chudova, Scott Gaffney, Padhraic Smyth |
UAI | 2 |
| 2000 | A general probabilistic framework for clustering individuals and objectsabstractThis paper presents a unifying probabilistic framework for clustering individuals or systems into groups when the available data measurements are not multiv ariate v ectors of xed dimensionality.For example, one might h a ve data from a set of medical patien ts,where for each patien tone has a set of of observed time-series, each time-series of potentially dierent length and dierent sampling rate.We propose a general model-based probabilistic framework for clustering data types of this form whic hare non-v ectorin nature and may vary in size from individual to individual.The Expectation-Maximization (EM) procedure for clustering within this framework is discussed and w e discuss ho w it be applied in a general manner to clustering of sequences, time-series, trajectories, and other non-vector data.We sho w that a number of earlier algorithms can be viewed as special cases within this unifying framework.The paper concludes with several illustrations of the method, including clustering of red blood cell data in a medical diagnosis context, clustering of proteins from curves of gene expression data, and clustering of individuals based on their sequences of Web na vigation. Igor V. Cadez, Scott Gaffney, Padhraic Smyth |
KDD | 2 |
| 1999 | Trajectory Clustering with Mixtures of Regression ModelsabstractIn this paper we address the problem of clustering trajectories, namely sets of short sequences of data measured as a function of a dependent variable such as time.Examples include storm path trajectories, longitudinal data such as drug therapy response, functional expression data in computational biology, and movements of objects or individuals in video sequences.Our clustering algorithm is based on a principled method for probabilistic modelhng of a set of trajectories as individual sequences of points generated from a finite mixture model consisting of regression model components.Unsupervised learning is carried out using maximum likelihood principles.Specifically, the EM algorithm is used to cope with the hidden data problem (i.e., the cluster memberships).We also develop generalizations of the method to handle non-parametric (kernel) regression components as well as multi-dimensional outputs.Simulation results comparing our method with other clustering methods such as K-means and Gaussian mixtures are presented as well as experimental results on real data sets. Scott Gaffney, Padhraic Smyth |
KDD | 1 |