Martin H. C. Law

dblp:68/2450 · DBLP profile ↗
← Back
11ranked-venue papers
8as first author
0since 2021 · last 2006
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-authorDatabases, data management, data science and information retrieval · 4 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Data mining · 100%
Artificial intelligence
4 papers
Probabilistic and Bayesian machine learning · 73% Representation and self-supervised learning · 21% Image recognition and object detection · 6%

Topics — the 17 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
clustering
0.132004
Simultaneous Feature Selection and Clustering Using Mixture Models · IEEE Trans. Pattern Anal. Mach. Intell. 2004
Analysis of Consensus Partition in Cluster Ensemble · ICDM 2004
Multiobjective Data Clustering · CVPR (2) 2004
Data mining
dimensionality reduction
0.112006
Incremental Nonlinear Dimensionality Reduction by Manifold Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2006
Data mining › dimensionality reduction
manifold learning
0.112006
Incremental Nonlinear Dimensionality Reduction by Manifold Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2006
Machine learning › Probabilistic and Bayesian machine learning › clustering
constrained clustering
0.112005
Learning with Constrained and Unlabelled Data · CVPR (1) 2005
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › exponential family
maximum entropy models
0.112005
Learning with Constrained and Unlabelled Data · CVPR (1) 2005
Data mining › clustering › clustering evaluation
cluster number estimation
0.012004
Simultaneous Feature Selection and Clustering Using Mixture Models · IEEE Trans. Pattern Anal. Mach. Intell. 2004
Data mining › clustering
ensemble clustering
0.012004
Analysis of Consensus Partition in Cluster Ensemble · ICDM 2004
Data mining › dimensionality reduction
feature selection
0.012004
Simultaneous Feature Selection and Clustering Using Mixture Models · IEEE Trans. Pattern Anal. Mach. Intell. 2004
Data mining › clustering › model-based clustering
mixture model clustering
0.012004
Simultaneous Feature Selection and Clustering Using Mixture Models · IEEE Trans. Pattern Anal. Mach. Intell. 2004
Data mining › clustering
multi-objective clustering
0.012004
Multiobjective Data Clustering · CVPR (2) 2004
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
expectation-maximization
0.012002
Feature Selection in Mixture-Based Clustering · NIPS 2002
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
feature selection
0.012002
Feature Selection in Mixture-Based Clustering · NIPS 2002
Machine learning › Probabilistic and Bayesian machine learning › clustering
gaussian mixture clustering
0.012002
Feature Selection in Mixture-Based Clustering · NIPS 2002
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
nonlinear dimensionality reduction
0.012006
Incremental Nonlinear Dimensionality Reduction by Manifold Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2006
Computer vision › Image recognition and object detection › image classification › object classification
face classification
0.012005
Learning with Constrained and Unlabelled Data · CVPR (1) 2005
Image and video processing
image segmentation
0.012005
Learning with Constrained and Unlabelled Data · CVPR (1) 2005
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
mixture model
0.012004
Simultaneous Feature Selection and Clustering Using Mixture Models · IEEE Trans. Pattern Anal. Mach. Intell. 2004

Methods — techniques the papers use, named apart from their topics

expectation-maximization · 0.1manifold learning · 0.1isomap · 0.1pairwise constraints · 0.1maximum entropy principle · 0.1minimum message length · 0.1stochastic partition model · 0.0plurality voting · 0.0multi-objective optimization · 0.0evolutionary algorithm · 0.0mutual-information-based feature relevance · 0.0
YearPublicationVenuePosition
2006 Incremental Nonlinear Dimensionality Reduction by Manifold Learning
abstract
Understanding the structure of multidimensional patterns, especially in unsupervised cases, is of fundamental importance in data mining, pattern recognition, and machine learning. Several algorithms have been proposed to analyze the structure of high-dimensional data based on the notion of manifold learning. These algorithms have been used to extract the intrinsic characteristics of different types of high-dimensional data by performing nonlinear dimensionality reduction. Most of these algorithms operate in a "batch" mode and cannot be efficiently applied when data are collected sequentially. In this paper, we describe an incremental version of ISOMAP, one of the key manifold learning algorithms. Our experiments on synthetic data as well as real world images demonstrate that our modified algorithm can maintain an accurate low-dimensional representation of the data in an efficient manner.
Martin H. C. Law, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2005 Learning with Constrained and Unlabelled Data
abstract
Classification problems abundantly arise in many computer vision tasks eing of supervised, semi-supervised or unsupervised nature. Even when class labels are not available, a user still might favor certain grouping solutions over others. This bias can be expressed either by providing a clustering criterion or cost function and, in addition to that, by specifying pairwise constraints on the assignment of objects to classes. In this work, we discuss a unifying formulation for labelled and unlabelled data that can incorporate constrained data for model fitting. Our approach models the constraint information by the maximum entropy principle. This modeling strategy allows us (i) to handle constraint violations and soft constraints, and, at the same time, (ii) to speed up the optimization process. Experimental results on face classification and image segmentation indicates that the proposed algorithm is computationally efficient and generates superior groupings when compared with alternative techniques.
Tilman Lange, Martin H. C. Law, Anil K. Jain 0001, Joachim M. Buhmann
CVPR (1)2
2005 Model-based Clustering With Probabilistic Constraints
abstract
The problem of clustering with constraints is receiving increasing attention. Many existing algorithms assume the specified constraints are correct and consistent. We take a new approach and model the uncertainty of constraints in a principled manner by treating the constraints as random variables. The effect of specified constraints on a subset of points is propagated to other data points by biasing the search for cluster boundaries. By combining the a posteriori enforcement of constraints with the log-likelihood, we obtain a new objective function. An EM-type algorithm derived by variational method is used for efficient parameter estimation. Experimental results demonstrate the usefulness of the proposed algorithm. In particular, our approach can identify the desired clusters even when only a small portion of data participates in constraints.
Martin H. C. Law, Alexander P. Topchy, Anil K. Jain 0001
SDM1
2004 Multiobjective Data Clustering
Martin H. C. Law, Alexander P. Topchy, Anil K. Jain 0001
CVPR (2)1
2004 Analysis of Consensus Partition in Cluster Ensemble
abstract
In combination of multiple partitions, one is usually interested in deriving a consensus solution with a quality better than that of given partitions. Several recent studies have empirically demonstrated improved accuracy of clustering ensembles on a number of artificial and real-world data sets. Unlike certain multiple supervised classifier systems, convergence properties of unsupervised clustering ensembles remain unknown for conventional combination schemes. In this paper, we present formal arguments on the effectiveness of cluster ensemble from two perspectives. The first is based on a stochastic partition generation model related to re-labeling and consensus function with plurality voting. The second is to study the property of the "mean" partition of an ensemble with respect to a metric on the space of all possible partitions. In both the cases, the consensus solution can be shown to converge to a true underlying clustering solution as the number of partitions in the ensemble increases. This paper provides a rigorous justification for the use of cluster ensemble.
Alexander P. Topchy, Martin H. C. Law, Anil K. Jain 0001, Ana Fred
ICDM2
2004 Nonlinear Manifold Learning for Data Stream
abstract
There has been a renewed interest in understanding the structure of high dimensional data set based on manifold learning. Examples include ISOMAP [25], LLE [20] and Laplacian Eigenmap [2] algorithms. Most of these algorithms operate in a “batch” mode and cannot be applied efficiently for a data stream. We propose an incremental version of ISOMAP. Our experiments not only demonstrate the accuracy and efficiency of the proposed algorithm, but also reveal interesting behavior of the ISOMAP as the size of available data increases.
Martin H. C. Law, Nan Zhang 0002, Anil K. Jain 0001
SDM1
2004 Simultaneous Feature Selection and Clustering Using Mixture Models
abstract
Clustering is a common unsupervised learning technique used to discover group structure in a set of data. While there exist many algorithms for clustering, the important issue of feature selection, that is, what attributes of the data should be used by the clustering algorithms, is rarely touched upon. Feature selection for clustering is difficult because, unlike in supervised learning, there are no class labels for the data and, thus, no obvious criteria to guide the search. Another important problem in clustering is the determination of the number of clusters, which clearly impacts and is influenced by the feature selection issue. In this paper, we propose the concept of feature saliency and introduce an expectation-maximization (EM) algorithm to estimate it, in the context of mixture-based clustering. Due to the introduction of a minimum message length model selection criterion, the saliency of irrelevant features is driven toward zero, which corresponds to performing feature selection. The criterion and algorithm are then extended to simultaneously estimate the feature saliencies and the number of clusters.
Martin H. C. Law, Mário A. T. Figueiredo, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2003 Mining customer product ratings for personalized marketing
William Kwok-Wai Cheung, James T. Kwok, Martin H. C. Law, Kwok Ching Tsui
Decis. Support Syst.3
2002 Feature Selection in Mixture-Based Clustering
abstract
There exist many approaches to clustering, but the important issue of feature selection, i.e., selecting the data attributes that are relevant for clustering, is rarely addressed. Feature selection for clustering is difficult due to the absence of class labels. We propose two approaches to feature selection in the context of Gaussian mixture-based clustering. In the first one, instead of making hard selections, we estimate feature saliencies. An expectation-maximization (EM) algorithm is derived for this task. The second approach extends Koller and Sahami’s mutual-information- based feature relevance criterion to the unsupervised case. Feature selec- tion is then carried out by a backward search scheme. This scheme can be classified as a “wrapper”, since it wraps mixture estimation in an outer layer that performs feature selection. Experimental results on synthetic and real data show that both methods have promising performance.
Martin H. C. Law, Anil K. Jain 0001, Mário A. T. Figueiredo
NIPS1
2001 Applying the Bayesian Evidence Framework to \nu -Support Vector Regression
Martin H. C. Law, James T. Kwok
ECML1
2000 Rival Penalized Competitive Learning for Model-Based Sequence Clustering
abstract
We propose a model-based, competitive learning procedure for the clustering of variable-length sequences. Hidden Markov models (HMMs) are used as representations for the cluster centers, and rival penalized competitive learning (RPCL), originally developed for domains with static, fixed-dimensional features, is extended. State merging operations are also incorporated to favor the discovery of smaller HMMs. Simulation results show that our extended version of RPCL can produce a more accurate cluster structure than k-means clustering.
Martin H. C. Law, James T. Kwok
ICPR1