Robert P. W. Duin

dblp:86/3985 · DBLP profile ↗
← Back
132ranked-venue papers
24as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 116 · 18 first-authorGraphics, computer vision, multimedia, augmented reality and games · 50 · 8 first-authorApplied, interdisciplinary, general and emerging computing · 5Theory of computation · 3 · 3 first-authorSystems, architecture and hardware · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
12 papers
Representation and self-supervised learning · 62% Learning theory · 14% 3D vision · 8%
Databases, data mining, and information retrieval
4 papers
Data mining · 92% Machine learning and data management · 8%
Theoretical computer science
7 papers
Computational geometry · 47% Algorithms and data structures · 39% Mathematical optimization · 9%

Topics — the 30 heaviest of 38, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
manifold learning
0.322014
Spherical and Hyperbolic Embeddings of Data · IEEE Trans. Pattern Anal. Mach. Intell. 2014
Spherical embeddings for non-Euclidean dissimilarities · CVPR 2010
Machine learning › Representation and self-supervised learning › representation learning › embedding learning › geometric embedding
non-euclidean embedding
0.212014
Spherical and Hyperbolic Embeddings of Data · IEEE Trans. Pattern Anal. Mach. Intell. 2014
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction
0.132004
Linear Dimensionality Reduction via a Heteroscedastic Extension of LDA: The Chernoff Criterion · IEEE Trans. Pattern Anal. Mach. Intell. 2004
Dimensionality Reduction by Canonical Contextual Correlation Projections · ECCV (1) 2004
Multiclass Linear Dimension Reduction by Weighted Pairwise Fisher Criteria · IEEE Trans. Pattern Anal. Mach. Intell. 2001
Data mining › predictive modeling
classification
0.122008
Subclass Problem-Dependent Design for Error-Correcting Output Codes · IEEE Trans. Pattern Anal. Mach. Intell. 2008
A Generalized Kernel Approach to Dissimilarity-based Classification · J. Mach. Learn. Res. 2001
Computer vision › 3D vision
3d shape analysis
0.112010
Spherical embeddings for non-Euclidean dissimilarities · CVPR 2010
Machine learning › Representation and self-supervised learning › representation learning › embedding learning › geometric embedding
spherical embedding
0.112010
Spherical embeddings for non-Euclidean dissimilarities · CVPR 2010
Machine learning › Learning theory › classification
classifier evaluation
0.112008
Efficient Multiclass ROC Approximation by Decomposition via Confusion Matrix Perturbation Analysis · IEEE Trans. Pattern Anal. Mach. Intell. 2008
Machine learning › Learning paradigms › cost-sensitive learning
cost-sensitive classification
0.112008
Efficient Multiclass ROC Approximation by Decomposition via Confusion Matrix Perturbation Analysis · IEEE Trans. Pattern Anal. Mach. Intell. 2008
Data mining › predictive modeling › classification
error-correcting output codes
0.112008
Subclass Problem-Dependent Design for Error-Correcting Output Codes · IEEE Trans. Pattern Anal. Mach. Intell. 2008
Data mining › predictive modeling › classification
multiclass classification
0.112008
Subclass Problem-Dependent Design for Error-Correcting Output Codes · IEEE Trans. Pattern Anal. Mach. Intell. 2008
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › discriminant analysis
linear discriminant analysis
0.122004
Linear Dimensionality Reduction via a Heteroscedastic Extension of LDA: The Chernoff Criterion · IEEE Trans. Pattern Anal. Mach. Intell. 2004
Multiclass Linear Dimension Reduction by Weighted Pairwise Fisher Criteria · IEEE Trans. Pattern Anal. Mach. Intell. 2001
Machine learning › Learning theory
statistical pattern recognition
0.122004
Linear Dimensionality Reduction via a Heteroscedastic Extension of LDA: The Chernoff Criterion · IEEE Trans. Pattern Anal. Mach. Intell. 2004
Statistical Pattern Recognition: A Review · IEEE Trans. Pattern Anal. Mach. Intell. 2000
Machine learning › Time series and sequential data › anomaly detection
one-class classification
0.122002
One-Class LP Classifiers for Dissimilarity Representations · NIPS 2002
Uniform Object Generation for Optimizing One-class Classifiers · J. Mach. Learn. Res. 2001
Algorithms and data structures › numerical linear algebra › dimensionality reduction
canonical correlation analysis
0.012004
Dimensionality Reduction by Canonical Contextual Correlation Projections · ECCV (1) 2004
Machine learning › Time series and sequential data
anomaly detection
0.012002
One-Class LP Classifiers for Dissimilarity Representations · NIPS 2002
Machine learning › Representation and self-supervised learning › prototype learning
prototype-based representation
0.012002
One-Class LP Classifiers for Dissimilarity Representations · NIPS 2002
Data mining › predictive modeling › classification › nearest neighbor classification
distance-based classification
0.012001
A Generalized Kernel Approach to Dissimilarity-based Classification · J. Mach. Learn. Res. 2001
Machine learning and data management
kernel methods
0.012001
A Generalized Kernel Approach to Dissimilarity-based Classification · J. Mach. Learn. Res. 2001
Machine learning › Learning theory › classification
classifier design
0.012000
Statistical Pattern Recognition: A Review · IEEE Trans. Pattern Anal. Mach. Intell. 2000
Machine learning › Kernel, tree and ensemble methods
classifier combination
0.011998
On Combining Classifiers · IEEE Trans. Pattern Anal. Mach. Intell. 1998
Data mining › model selection
classifier selection
0.011995
An Evaluation of Intrinsic Dimensionality Estimators · IEEE Trans. Pattern Anal. Mach. Intell. 1995
Data mining
dimensionality reduction
0.011995
An Evaluation of Intrinsic Dimensionality Estimators · IEEE Trans. Pattern Anal. Mach. Intell. 1995
Data mining › dimensionality reduction
intrinsic dimensionality estimation
0.011995
An Evaluation of Intrinsic Dimensionality Estimators · IEEE Trans. Pattern Anal. Mach. Intell. 1995
Data mining › predictive modeling › classification
pattern classification
0.011995
An Evaluation of Intrinsic Dimensionality Estimators · IEEE Trans. Pattern Anal. Mach. Intell. 1995
Mathematical optimization
linear programming
0.012002
One-Class LP Classifiers for Dissimilarity Representations · NIPS 2002
Machine learning › Kernel, tree and ensemble methods
kernel methods
0.012001
A Generalized Kernel Approach to Dissimilarity-based Classification · J. Mach. Learn. Res. 2001
Data mining
pattern mining
0.012000
Statistical Pattern Recognition: A Review · IEEE Trans. Pattern Anal. Mach. Intell. 2000
Machine learning › Learning theory › statistical estimation
error estimation
0.011998
On Combining Classifiers · IEEE Trans. Pattern Anal. Mach. Intell. 1998
Information theory › pattern recognition
statistical pattern recognition
0.021978
On the evaluation of independent binary features (Corresp.) · IEEE Trans. Inf. Theory 1978
The mean recognition performance for independent distributions (Corresp.) · IEEE Trans. Inf. Theory 1978
Coding theory
channel coding
0.011978
The mean recognition performance for independent distributions (Corresp.) · IEEE Trans. Inf. Theory 1978

Methods — techniques the papers use, named apart from their topics

tangent space optimization · 0.4exponential map · 0.4hyperbolic embedding · 0.1elliptic embedding · 0.1contextual correlation projections · 0.1problem-dependent ECOC design · 0.1neyman-pearson optimization · 0.1confusion matrix perturbation · 0.1ROC decomposition · 0.1eigenvector decomposition · 0.0chernoff criterion · 0.0linear programming · 0.0dissimilarity transformation · 0.0kernel · 0.0dissimilarity representation · 0.0statistical learning theory · 0.0neural network · 0.0cluster analysis · 0.0
YearPublicationVenuePosition
2016 A Compact Representation of Multiscale Dissimilarity Data by Prototype Selection
Yenisel Plasencia, Yan Li 0009, Robert P. W. Duin, Mauricio Orozco-Alzate, Marco Loog, Edel B. García Reyes
CIARP3
2015 The dissimilarity representation for finding universals from particulars by an anti-essentialist approach
Robert P. W. Duin
Pattern Recognit. Lett.1
2014 Improving Hyperspectral Pixel Classification With Unsupervised Training Data Selection
abstract
An unsupervised method for selecting training data is suggested here. The method is tested by applying it to hyperspectral land-use classification. The data set is reduced using an unsupervised band selection method and then clustered with a nonparametric cluster technique. The cluster technique provides centers of the clusters, and those are the samples selected to compose the training set. Both the band selection and the clustering are unsupervised techniques. Afterward, an expert labels those samples, and the rest of unlabeled data can be classified. The inclusion of the selection step, although unsupervised, allows to select automatically the most suitable pixels to build the classifier. This reduces the expert effort because less pixels need to be labeled. However, the classification results are significantly improved in comparison with the results obtained by a random selection of training samples, in particular for very small training sets.
Olga Rajadell, Pedro García-Sevilla, Viet Cuong Dinh, Robert P. W. Duin
IEEE Geosci. Remote. Sens. Lett.4
2014 Spherical and Hyperbolic Embeddings of Data
abstract
Many computer vision and pattern recognition problems may be posed as the analysis of a set of dissimilarities between objects. For many types of data, these dissimilarities are not euclidean (i.e., they do not represent the distances between points in a euclidean space), and therefore cannot be isometrically embedded in a euclidean space. Examples include shape-dissimilarities, graph distances and mesh geodesic distances. In this paper, we provide a means of embedding such non-euclidean data onto surfaces of constant curvature. We aim to embed the data on a space whose radius of curvature is determined by the dissimilarity data. The space can be either of positive curvature (spherical) or of negative curvature (hyperbolic). We give an efficient method for solving the spherical and hyperbolic embedding problems on symmetric dissimilarity data. Our approach gives the radius of curvature and a method for approximating the objects as points on a hyperspherical manifold without optimisation. For objects which do not reside exactly on the manifold, we develop a optimisation-based procedure for approximate embedding on a hyperspherical manifold. We use the exponential map between the manifold and its local tangent space to solve the optimisation problem locally in the euclidean tangent space. This process is efficient enough to allow us to embed data sets of several thousand objects. We apply our method to a variety of data including time warping functions, shape similarities, graph similarity and gesture similarity data. In each case the embedding maintains the local structure of the data while placing the points in a metric space.
Richard C. Wilson 0001, Edwin R. Hancock, Elzbieta Pekalska, Robert P. W. Duin
IEEE Trans. Pattern Anal. Mach. Intell.4
2013 Towards Cluster-Based Prototype Sets for Classification in the Dissimilarity Space
Yenisel Plasencia, Mauricio Orozco-Alzate, Edel B. García Reyes, Robert P. W. Duin
CIARP (1)4
2013 Missing Values in Dissimilarity-Based Classification of Multi-way Data
Diana Porro-Muñoz, Robert P. W. Duin, Isneri Talavera-Bustamante
CIARP (1)2
2013 The Area under the ROC Curve as a Criterion for Clustering Evaluation
Helena Aidos, Robert P. W. Duin, Ana Fred
ICPRAM2
2013 FIDOS: A generalized Fisher based feature extraction method for domain shift
Viet Cuong Dinh, Robert P. W. Duin, Ignacio Piqueras-Salazar, Marco Loog
Pattern Recognit.2
2013 Multiple-instance learning as a classifier combining problem
Yan Li 0009, David M. J. Tax, Robert P. W. Duin, Marco Loog
Pattern Recognit.3
2013 Multi-spectral video endoscopy system for the detection of cancerous tissue
Raimund Leitner, Martin De Biasio, Viet Cuong Dinh, Marco Loog, Robert P. W. Duin
Pattern Recognit. Lett.6
2012 On Using Asymmetry Information for Classification in Extended Dissimilarity Spaces
Yenisel Plasencia, Edel B. García Reyes, Robert P. W. Duin, Mauricio Orozco-Alzate
CIARP3
2012 Continuous Multi-way Shape Measure for Dissimilarity Representation
Diana Porro-Muñoz, Robert P. W. Duin, Mauricio Orozco-Alzate, Isneri Talavera-Bustamante
CIARP2
2012 A study on semi-supervised dissimilarity representation
Viet Cuong Dinh, Robert P. W. Duin, Marco Loog
ICPR2
2012 Training data selection for cancer detection in multispectral endoscopy images
Viet Cuong Dinh, Marco Loog, Raimund Leitner, Olga Rajadell, Robert P. W. Duin
ICPR5
2012 Automated classification of local patches in colon histopathology
Habil Kalkan, Marius Nap, Robert P. W. Duin, Marco Loog
ICPR3
2012 Combining multi-scale dissimilarities for image classification
Yan Li 0009, Robert P. W. Duin, Marco Loog
ICPR2
2012 Classification using High Order Dissimilarities in Non-euclidean Spaces
Helena Aidos, Ana Fred, Robert P. W. Duin
ICPRAM (1)3
2012 Automated Colorectal Cancer Diagnosis for Whole-Slice Histopathology
Habil Kalkan, Marius Nap, Robert P. W. Duin, Marco Loog
MICCAI (3)3
2012 Bridging Structure and Feature Representations in Graph Matching
abstract
Structures and features are opposite approaches in building representations for object recognition. Bridging the two is an essential problem in pattern recognition as the two opposite types of information are fundamentally different. As dissimilarities can be computed for both the dissimilarity representation can be used to combine the two. Attributed graphs contain structural as well as feature-based information. Neglecting the attributes yields a pure structural description. Isolating the features and neglecting the structure represents objects by a bag of features. In this paper we will show that weighted combinations of dissimilarities may perform better than these two extremes, indicating that these two types of information are essentially different and strengthen each other. In addition we present two more advanced integrations than weighted combining and show that these may improve the classification performances even further.
Wan-Jui Lee, Veronika Cheplygina, David M. J. Tax, Marco Loog, Robert P. W. Duin
Int. J. Pattern Recognit. Artif. Intell.5
2012 The dissimilarity space: Bridging structural and statistical pattern recognition
Robert P. W. Duin, Elzbieta Pekalska
Pattern Recognit. Lett.1
2011 The Dissimilarity Representation for Structural Pattern Recognition
Robert P. W. Duin, Elzbieta Pekalska
CIARP1
2011 Dissimilarity-Based Classifications in Eigenspaces
Sang-Woon Kim, Robert P. W. Duin
CIARP2
2011 SEDMI: Saliency based edge detection in multispectral images
Viet Cuong Dinh, Raimund Leitner, Pavel Paclík, Marco Loog, Robert P. W. Duin
Image Vis. Comput.5
2011 An experimental study of one- and two-level classifier fusion for different sample sizes
Chunxia Zhang 0002, Robert P. W. Duin
Pattern Recognit. Lett.2
2011 Classification of three-way data by the dissimilarity representation
Diana Porro-Muñoz, Robert P. W. Duin, Isneri Talavera-Bustamante, Mauricio Orozco-Alzate
Signal Process.2
2010 On Improving Dissimilarity-Based Classifications Using a Statistical Similarity Measure
Sang-Woon Kim, Robert P. W. Duin
CIARP2
2010 Spherical embeddings for non-Euclidean dissimilarities
abstract
Many computer vision and pattern recognition problems may be posed by defining a way of measuring dissimilarities between patterns. For many types of data, these dissimilarities are not Euclidean, and may not be metric. In this paper, we provide a means of embedding such data. We aim to embed the data on a hypersphere whose radius of curvature is determined by the dissimilarity data. The hypersphere can be either of positive curvature (elliptic) or of negative curvature (hyperbolic). We give an efficient method for solving the elliptic and hyperbolic embedding problems on symmetric dissimilarity data. This method gives the radius of curvature and a method for approximating the objects as points on a hyperspherical manifold. We apply our method to a variety of data including shape-similarities, graph-similarity and gesture-similarity data. In each case the embedding maintains the local structure of the data while placing the points in a metric space.
Richard C. Wilson 0001, Edwin R. Hancock, Elzbieta Pekalska, Robert P. W. Duin
CVPR4
2010 Prototype Selection for Dissimilarity Representation by a Genetic Algorithm
abstract
Dissimilarities can be a powerful way to represent objects like strings, graphs and images for which it is difficult to find good features. The resulting dissimilarity space may be used to train any classifier appropriate for feature spaces. There is, however, a strong need for dimension reduction. Straightforward procedures for prototype selection as well as feature selection have been used for this in the past. Complicated sets of objects may need more advanced procedures to overcome local minima. In this paper it is shown that genetic algorithms, previously used for feature selection, may be used for building good dissimilarity spaces as well, especially when small sets of prototypes are needed for computational reasons.
Yenisel Plasencia, Edel B. García Reyes, Mauricio Orozco-Alzate, Robert P. W. Duin
ICPR4
2010 Classification of Volcano Events Observed by Multiple Seismic Stations
abstract
Seismic events in and around volcanos, like tremors, earth quakes, ice quakes and strokes of lightning, are usually observed by multiple stations. The question rises whether classifiers trained for one seismic station can be used for classifying observations by other stations, and, moreover, whether a combination of station signals improves the classification performances for a single station. We study this for seismic time signals represented by spectra and spectrograms obtained from 5 seismic stations on the Nevado del Ruiz in Colombia.
Robert P. W. Duin, Mauricio Orozco-Alzate, John Makario Londoño-Bonilla
ICPR1
2010 Random Subspace Method in Text Categorization
abstract
In text categorization (TC), which is a supervised technique, a feature vector of terms or phrases is usually used to represent the documents. Due to the huge number of terms in even a moderate-size text corpus, high dimensional feature space is an intrinsic problem in TC. Random subspace method (RSM), a technique that divides the feature space to smaller ones each submitted to a (base) classifier (BC) in an ensemble, can be an effective approach to reduce the dimensionality of the feature space. Inspired by a similar research on functional magnetic resonance imaging (fMRI) of brain, here we address the estimation of ensemble parameters, i.e., the ensemble size (L) and the dimensionality of feature subsets (M) by defining three criteria: usability, coverage, and diversity of the ensemble. We will show that relatively medium M and small L yield an ensemble that improves the performance of a single support vector machine, which is considered as the state-of-the-art in TC.
Mehrdad J. Gangeh, Mohamed S. Kamel, Robert P. W. Duin
ICPR3
2010 A Study on Combining Sets of Differently Measured Dissimilarities
abstract
The ways distances are computed or measured enable us to have different representations of the same objects. In this paper we want to discuss possible ways of merging different sources of information given by differently measured dissimilarity representations. We compare here a simple averaging scheme [1] with dissimilarity forward selection and other techniques based on the learning of weights of linear and quadratic forms. Our general conclusion is that, although the more advanced forms of combination cannot always lead to better classification accuracies, combining given distance matrices prior to training is always worthwhile. We can thereby suggest which combination schemes are preferable with respect to the problem data.
Alessandro Ibba, Robert P. W. Duin, Wan-Jui Lee
ICPR2
2010 ROC Analysis and Cost-Sensitive Optimization for Hierarchical Classifiers
abstract
Instead of solving complex pattern recognition problems using a single complicated classifier, it is often beneficial to leverage our prior knowledge and decompose the problem into parts. These may be tackled using specific feature subsets and simpler classifiers resulting in a hierarchical system. In this paper, we propose an efficient and scalable approach for cost-sensitive optimization of a general hierarchical classifier using ROC analysis. This allows the designer to view the hierarchy of trained classifiers as a system, and tune it according to the application needs.
Pavel Paclík, Carmen Lai, Thomas C. W. Landgrebe, Robert P. W. Duin
ICPR4
2010 Classifying Three-way Seismic Volcanic Data by Dissimilarity Representation
abstract
Multi-way data analysis is a multivariate data analysis technique having a wide application in some fields. Nevertheless, the development of classification tools for this type of representation is incipient yet. In this paper we study the dissimilarity representation for the classification of three-way data, as dissimilarities allow the representation of multi-dimensional objects in a natural way. As an example, the classification of seismic volcanic events is used. It is shown that in this application classification based on 2D spectrograms, dissimilarities perform better than on 1D spectral features.
Diana Porro-Muñoz, Robert P. W. Duin, Mauricio Orozco-Alzate, Isneri Talavera-Bustamante, John Makario Londoño-Bonilla
ICPR2
2010 Image Dissimilarity-Based Quantification of Lung Disease from CT
Lauge Sørensen, Marco Loog, Pechin Lo, Haseem Ashraf, Asger Dirksen, Robert P. W. Duin, Marleen de Bruijne
MICCAI (1)6
2010 Award winning papers from the 19th International Conference on Pattern Recognition (ICPR)
Robert P. W. Duin, Denis Laurendeau, Brian C. Lovell
Pattern Recognit. Lett.1
2010 A multi-classifier for grading knee osteoarthritis using gait analysis
Nigar Sen Köktas, Nese Yalabik, Günes Yavuzer, Robert P. W. Duin
Pattern Recognit. Lett.4
2009 A Combine-Correct-Combine Scheme for Optimizing Dissimilarity-Based Classifiers
Sang-Woon Kim, Robert P. W. Duin
CIARP2
2009 A Study on Representations for Face Recognition from Thermal Images
Yenisel Plasencia, Edel B. García Reyes, Robert P. W. Duin, Heydi Mendez Vazquez, César San-Martín, Claudio Soto
CIARP3
2009 The Representation of Chemical Spectral Data for Classification
Diana Porro-Muñoz, Robert P. W. Duin, Isneri Talavera-Bustamante, Noslén Hernández-González
CIARP2
2009 Minimum spanning tree based one-class classifier
Piotr Juszczak, David M. J. Tax, Elzbieta Pekalska, Robert P. W. Duin
Neurocomputing4
2009 Component-based discriminative classification for hidden Markov models
Manuele Bicego, Elzbieta Pekalska, David M. J. Tax, Robert P. W. Duin
Pattern Recognit.4
2009 A generalization of dissimilarity representations using feature lines and feature planes
Mauricio Orozco-Alzate, Robert P. W. Duin, Germán Castellanos-Domínguez
Pattern Recognit. Lett.2
2009 Dimensionality Reduction of Hyperspectral Data via Spectral Feature Extraction
abstract
This paper proposes an innovative spectral feature extraction (SFE) method called prototype space (PS) feature extraction (PSFE) based only on class spectra. The main novelties of the proposed SFE lie in the following: representing channels in a new space called PS, where they are characterized in terms of reflection properties of classes; and proposing uncertainty, angle, and distance measures to distinguish highly correlated and informative channels. Having clustered the channels by Fuzzy C-Means (FCM) in PS, highly correlated and isolated channels are separated by an uncertainty measure. Consequently, PSFE is built by a linear combination of spectra weighted by class membership values of channels that fall in each cluster. Furthermore, we will enrich PSFE with informative channels which are outlier channels identified through their angle and distance with respect to diagonal and cluster centers in PS. In contrast to the previous SFE methods, PSFE substitutes the search strategies with FCM clustering to find relevant channels. Moreover, instead of optimizing separability criteria, the accuracy of classification over a subset of training data is used to decide which disjoint optical region yields maximum accuracy. According to how class spectra are obtained, PSFE incorporates four approaches: knowledge based, supervised, semisupervised, and unsupervised. The latter three PSFE approaches are assessed in two main cases including with and without informative channels and compared with the conventional feature extraction methods. Experimental results demonstrated higher overall accuracy of PSFE compared to its conventional counterparts with limited sample sizes.
Barat Mojaradi, Hamid Abrishami Moghaddam, Mohammad Javad Valadan Zoej, Robert P. W. Duin
IEEE Trans. Geosci. Remote. Sens.4
2008 Spectral Characterization of Volcanic Earthquakes at Nevado del Ruiz Volcano Using Spectral Band Selection/Extraction Techniques
Mauricio Orozco-Alzate, Marina Skurichina, Robert P. W. Duin
CIARP3
2008 On refining dissimilarity matrices for an improved NN learning
abstract
Application-specific dissimilarity functions can be used for learning from a set of objects represented by pairwise dissimilarity matrices in this context. These dissimilarities may, however, suffer from various defects, e.g. when derived from a suboptimal optimization or by the use of non-metric or noisy measures. In this paper, we study procedures for refining such dissimilarities. These methods work in a representation space, either a dissimilarity space or a pseudo-Euclidean embedded space. On a series of experiments we show that refining may significantly improve the nearest neighbor classifications of dissimilarity measurements.
Robert P. W. Duin, Elzbieta Pekalska
ICPR1
2008 Variance estimation for two-class and multi-class ROC analysis using operating point averaging
abstract
Receiver Operating Characteristic (ROC) analysis enables fine-tuning of a trained classifier to a desired performance trade-off situation. ROC estimated from a finite test set is, however, insufficient for the sake of classifier comparison as it neglects performance variances. This research presents a practical algorithm for variance estimation at individual operating points of ROC curves or surfaces. It generalizes the threshold averaging of Fawcett et.al. to arbitrary operating point definition including the weighting-based formulation used in multi-class ROC analysis. The statistical test comparing performance differences between operating points of the same curve is illustrated for two-class and multi-class ROC.
Pavel Paclík, Carmen Lai, Jana Novovicová, Robert P. W. Duin
ICPR4
2008 Subclass Problem-Dependent Design for Error-Correcting Output Codes
abstract
A common way to model multi-class classification problems is by means of Error-Correcting Output Codes (ECOC). Given a multi-class problem, the ECOC technique designs a code word for each class, where each position of the code identifies the membership of the class for a given binary problem. A classification decision is obtained by assigning the label of the class with the closest code. One of the main requirements of the ECOC design is that the base classifier is capable of splitting each sub-group of classes from each binary problem. However, we can not guarantee that a linear classifier model convex regions. Furthermore, non-linear classifiers also fail to manage some type of surfaces. In this paper, we present a novel strategy to model multi-class classification problems using sub-class information in the ECOC framework. Complex problems are solved by splitting the original set of classes into sub-classes, and embedding the binary problems in a problem-dependent ECOC design. Experimental results show that the proposed splitting procedure yields a better performance when the class overlap or the distribution of the training objects conceil the decision boundaries for the base classifier. The results are even more significant when one has a sufficiently large training size.
Sergio Escalera, David M. J. Tax, Oriol Pujol, Petia Radeva, Robert P. W. Duin
IEEE Trans. Pattern Anal. Mach. Intell.5
2008 Efficient Multiclass ROC Approximation by Decomposition via Confusion Matrix Perturbation Analysis
abstract
ROC analysis has become a standard tool in the design and evaluation of 2-class classification problems. It allows for an analysis that incorporates all possible priors, costs, and operating points, which is important in many real problems, where conditions are often nonideal. Extending this to the multiclass case is attractive, conferring the benefits of ROC analysis to a multitude of new problems. Even though ROC analysis does extend theoretically to the multiclass case, the exponential computational complexity as a function of the number of classes is restrictive. In this paper we show that the multiclass ROC can often be simplified considerably because some ROC dimensions are independent of each other. We present an algorithm that analyses interactions between various ROC dimensions, identifying independent classes, and groups of interacting classes, allowing the ROC to be decomposed. The resultant decomposed ROC hypersurface can be interrogated in a similar fashion to the ideal case, allowing for approaches such as cost-sensitive and Neyman-Pearson optimisation, as well as the volume under the ROC. An extensive bouquet of examples and experiments demonstrates the potential of this methodology.
Thomas C. W. Landgrebe, Robert P. W. Duin
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 Integration of prior knowledge of measurement noise in kernel density classification
Yunlei Li, Dick de Ridder, Robert P. W. Duin, Marcel J. T. Reinders
Pattern Recognit.3
2008 Maximizing the area under the ROC curve by pairwise feature combination
Claudio Marrocco, Robert P. W. Duin, Francesco Tortorella
Pattern Recognit.2
2008 Growing a multi-class classifier with a reject option
David M. J. Tax, Robert P. W. Duin
Pattern Recognit. Lett.2
2008 Beyond Traditional Kernels: Classification in Two Dissimilarity-Based Representation Spaces
abstract
Proximity captures the degree of similarity between examples and is thereby fundamental in learning. Learning from pairwise proximity data usually relies on either kernel methods for specifically designed kernels or the nearest neighbor (NN) rule. Kernel methods are powerful, but often cannot handle arbitrary proximities without necessary corrections. The NN rule can work well in such cases, but suffers from local decisions. The aim of this paper is to provide an indispensable explanation and insights about two simple yet powerful alternatives when neither conventional kernel methods nor the NN rule can perform best. These strategies use two proximity-based representation spaces (RSs) in which accurate classifiers are trained on all training objects and demand comparisons to a small set of prototypes. They can handle all meaningful dissimilarity measures, including non-Euclidean and nonmetric ones. Practical examples illustrate that these RSs can be highly advantageous in supervised learning. Simple classifiers built there tend to outperform the NN rule. Moreover, computational complexity may be controlled. Consequently, these approaches offer an appealing alternative to learn from proximity data for which kernel methods cannot directly be applied, are too costly or impractical, while the NN rule leads to noisy results.
Elzbieta Pekalska, Robert P. W. Duin
IEEE Trans. Syst. Man Cybern. Part C2
2007 On Using a Pre-clustering Technique to Optimize LDA-Based Classifiers for Appearance-Based Face Recognition
Sang-Woon Kim, Robert P. W. Duin
CIARP2
2007 Generalizing Dissimilarity Representations Using Feature Lines
Mauricio Orozco-Alzate, Robert P. W. Duin, Germán Castellanos-Domínguez
CIARP2
2007 Pairwise feature evaluation for constructing reduced representations
Artsiom Harol, Carmen Lai, Elzbieta Pekalska, Robert P. W. Duin
Pattern Anal. Appl.4
2007 Approximating the multiclass ROC by pairwise analysis
Thomas C. W. Landgrebe, Robert P. W. Duin
Pattern Recognit. Lett.2
2006 Experimental study on prototype optimisation algorithms for prototype-based classification in vector spaces
Mayte Lozano, José Martínez Sotoca, J. Salvador Sánchez 0001, Filiberto Pla, Elzbieta Pekalska, Robert P. W. Duin
Pattern Recognit.6
2006 Prototype selection for dissimilarity-based classifiers
Elzbieta Pekalska, Robert P. W. Duin, Pavel Paclík
Pattern Recognit.2
2006 Recent submissions in linear dimensionality reduction and face recognition
Robert P. W. Duin, Marco Loog, Tin Kam Ho
Pattern Recognit. Lett.1
2006 The interaction between classification and reject performance for distance-based reject-option classifiers
Thomas C. W. Landgrebe, David M. J. Tax, Pavel Paclík, Robert P. W. Duin
Pattern Recognit. Lett.4
2006 Building Road-Sign Classifiers Using a Trainable Similarity Measure
abstract
Deriving an informative data representation is an important prerequisite when designing road-sign classifiers. A frequently used strategy for road-sign classification is based on the normalized cross correlation similarity to class prototypes followed by the nearest neighbor classifier. Because of the global nature of the cross correlation similarity, this method suffers from presence of uninformative pixels (caused, e.g., by occlusions) and is computationally demanding. In this paper, a novel concept of a trainable similarity measure is introduced, which alleviates these shortcomings. The similarity is based on individual matches in a set of local image regions. The set of regions that are relevant for a particular similarity assessment is refined by the training process. It is illustrated on a set of experiments with road-sign-classification problems that the trainable similarity yields high-performance data representations and classifiers. Apart from a multiclass classification accuracy, nonsign rejection capability and computational demands in execution are also discussed. It appears that the trainable similarity representation alleviates some difficulties of other algorithms that are currently used in road-sign classification.
Pavel Paclík, Jana Novovicová, Robert P. W. Duin
IEEE Trans. Intell. Transp. Syst.3
2005 Dimensionality reduction of image features using the canonical contextual correlation projection
Marco Loog, Bram van Ginneken, Robert P. W. Duin
Pattern Recognit.3
2004 Dimensionality Reduction by Canonical Contextual Correlation Projections
Marco Loog, Bram van Ginneken, Robert P. W. Duin
ECCV (1)3
2004 A Study On Combining Image Representations For Image Classification And Retrieval
abstract
A flexible description of images is offered by a cloud of points in a feature space. In the context of image retrieval such clouds can be represented in a number of ways. Two approaches are here considered. The first approach is based on the assumption of a normal distribution, hence homogeneous clouds, while the second one focuses on the boundary description, which is more suitable for multimodal clouds. The images are then compared either by using the Mahalanobis distance or by the support vector data description (SVDD), respectively. The paper investigates some possibilities of combining the image clouds based on the idea that responses of several cloud descriptions may convey a pattern, specific for semantically similar images. A ranking of image dissimilarities is used as a comparison for two image databases targeting image classification and retrieval problems. We show that combining of the SVDD descriptions improves the retrieval performance with respect to ranking, on the contrary to the Mahalanobis case. Surprisingly, it turns out that the ranking of the Mahalanobis distances works well also for inhomogeneous images.
Carmen Lai, David M. J. Tax, Robert P. W. Duin, Elzbieta Pekalska, Pavel Paclík
Int. J. Pattern Recognit. Artif. Intell.3
2004 Support Vector Data Description
David M. J. Tax, Robert P. W. Duin
Mach. Learn.2
2004 The MDF discrimination measure: Fisher in disguise
Marco Loog, Robert P. W. Duin, Max A. Viergever
Neural Networks2
2004 Linear Dimensionality Reduction via a Heteroscedastic Extension of LDA: The Chernoff Criterion
abstract
We propose an eigenvector-based heteroscedastic linear dimension reduction (LDR) technique for multiclass data. The technique is based on a heteroscedastic two-class technique which utilizes the so-called Chernoff criterion, and successfully extends the well-known linear discriminant analysis (LDA). The latter, which is based on the Fisher criterion, is incapable of dealing with heteroscedastic data in a proper way. For the two-class case, the between-class scatter is generalized so to capture differences in (co)variances. It is shown that the classical notion of between-class scatter can be associated with Euclidean distances between class means. From this viewpoint, the between-class scatter is generalized by employing the Chernoff distance measure, leading to our proposed heteroscedastic measure. Finally, using the results from the two-class case, a multiclass extension of the Chernoff criterion is proposed. This criterion combines separation information present in the class mean as well as the class covariance matrices. Extensive experiments and a comparison with similar dimension reduction techniques are presented.
Marco Loog, Robert P. W. Duin
IEEE Trans. Pattern Anal. Mach. Intell.2
2004 Almost autonomous training of mixtures of principal component analyzers
Mohamed E. M. Musa, Dick de Ridder, Robert P. W. Duin, Volkan Atalay
Pattern Recognit. Lett.3
2003 Selective Sampling Methods in One-Class Classification Problems
Piotr Juszczak, Robert P. W. Duin
ICANN2
2003 Supervised Locally Linear Embedding
Dick de Ridder, Olga Kouropteva, Oleg Okun, Matti Pietikäinen, Robert P. W. Duin
ICANN5
2003 Segmentation of multi-spectral images using the combined classifier approach
Pavel Paclík, Robert P. W. Duin, Geert M. P. van Kempen, Reinhard Kohlus
Image Vis. Comput.2
2003 Limits on the majority vote accuracy in classifier fusion
Ludmila I. Kuncheva, Christopher J. Whitaker, Catherine A. Shipp, Robert P. W. Duin
Pattern Anal. Appl.4
2002 One-Class LP Classifiers for Dissimilarity Representations
abstract
Problems in which abnormal or novel situations should be detected can be approached by describing the domain of the class of typical exam- ples. These applications come from the areas of machine diagnostics, fault detection, illness identification or, in principle, refer to any prob- lem where little knowledge is available outside the typical class. In this paper we explain why proximities are natural representations for domain descriptors and we propose a simple one-class classifier for dissimilarity representations. By the use of linear programming an efficient one-class description can be found, based on a small number of prototype objects. This classifier can be made (1) more robust by transforming the dissimi- larities and (2) cheaper to compute by using a reduced representation set. Finally, a comparison to a comparable one-class classifier by Campbell and Bennett is given.
Elzbieta Pekalska, David M. J. Tax, Robert P. W. Duin
NIPS3
2002 Blind separation of rotating machine sources: bilinear forms and convolutive mixtures
Alexander Ypma, Amir Leshem, Robert P. W. Duin
Neurocomputing3
2002 Bagging, Boosting and the Random Subspace Method for Linear Classifiers
Marina Skurichina, Robert P. W. Duin
Pattern Anal. Appl.2
2002 A note on core research issues for statistical pattern recognition
Robert P. W. Duin, Fabio Roli, Dick de Ridder
Pattern Recognit. Lett.1
2002 Dissimilarity representations allow for building good classifiers
Elzbieta Pekalska, Robert P. W. Duin
Pattern Recognit. Lett.2
2001 Health Monitoring with Learning Methods
Alexander Ypma, Co Melissant, Ole Baunbæk-Jensen, Robert P. W. Duin
ICANN4
2001 A Generalized Kernel Approach to Dissimilarity-based Classification
Elzbieta Pekalska, Pavel Paclík, Robert P. W. Duin
J. Mach. Learn. Res.3
2001 Uniform Object Generation for Optimizing One-class Classifiers
David M. J. Tax, Robert P. W. Duin
J. Mach. Learn. Res.2
2001 Multiclass Linear Dimension Reduction by Weighted Pairwise Fisher Criteria
abstract
We derive a class of computationally inexpensive linear dimension reduction criteria by introducing a weighted variant of the well-known K-class Fisher criterion associated with linear discriminant analysis (LDA). It can be seen that LDA weights contributions of individual class pairs according to the Euclidean distance of the respective class means. We generalize upon LDA by introducing a different weighting function.
Marco Loog, Robert P. W. Duin, Reinhold Häb-Umbach
IEEE Trans. Pattern Anal. Mach. Intell.2
2001 Decision templates for multiple classifier fusion: an experimental comparison
Ludmila I. Kuncheva, James C. Bezdek, Robert P. W. Duin
Pattern Recognit.3
2000 Probabilistic PCA and ICA Subspace Mixture Models for Image Segmentation
abstract
High-dimensional data, such as images represented as points in the space spanned by their pixel values, can often be described in a significantly smaller number of dimensions than the original. One of the ways of finding lowdimensional representations is to train a mixture model of principal component analysers (PCA) on the data. However, some types of data do not fulfill the assumptions of PCA, calling for application of different subspace methods. One such a method is ICA, which has been shown in recent years to be able to find interesting basis vectors (features) in signal and image data. In this paper, a mixture model of ICA subspaces is developed similar to a mixture model of PCA subspaces proposed by others. The new algorithm is applied to a natural texture segmentation problem and is shown to give encouraging results.
Dick de Ridder, Josef Kittler, Robert P. W. Duin
BMVC3
2000 Classifiers in Almost Empty Spaces
abstract
Recent developments in defining and training statistical classifiers make it possible to build reliable classifiers in very small sample size problems. Using these techniques advanced problems may be tackled, such as pixel based image recognition and dissimilarity based object classification. It can be explained and illustrated how recognition systems based on support vector machines and subspace classifiers circumvent the curse of dimensionality, and even may find nonlinear decision boundaries for small training sets represented in Hilbert space.
Robert P. W. Duin
ICPR1
2000 Multi-Class Linear Feature Extraction by Nonlinear PCA
abstract
The traditional way to find a linear solution to feature extraction problems is based on the maximization of the class-between scatter over the class-within scatter (Fisher's mapping). For the multi-class problem this is sub-optimal due to class conjunctions, even for the simple situation of normal distributed classes with identical covariance matrices. We propose a novel, equally fast method, based on nonlinear principal component analysis (PCA). Although still sub-optimal, it may avoid the class conjunction. The proposed method is experimentally compared with Fisher's mapping and with a neural network based approach to nonlinear PCA. It appears to outperform the both methods.
Robert P. W. Duin, Marco Loog, Reinhold Häb-Umbach
ICPR1
2000 Is Independence Good For Combining Classifiers?
abstract
Independence between individual classifiers is typically viewed as an asset in classifier fusion. We study the limits on the majority vote accuracy when combining dependent classifiers. Q statistics are used to measure the dependence between classifiers. We show that dependent classifiers could offer a dramatic improvement over the individual accuracy. However, the relationship between dependency and accuracy of the pool is ambivalent. A synthetic experiment demonstrates the intuitive result that, in general, negative dependence is preferable.
Ludmila I. Kuncheva, Christopher J. Whitaker, Catherine A. Shipp, Robert P. W. Duin
ICPR4
2000 Classifiers for Dissimilarity-Based Pattern Recognition
abstract
In the traditional way of learning from examples of objects the classifiers are built in a feature space. However, alternative ways can be found by constructing decision rules on dissimilarity (distance) representations, instead. In such a recognition process a new object is described by its distances to (a subset of) the training samples. In this paper, a number of methods to tackle this type of classification problem are investigated: the feature-based (interpreting the distance representation as a feature space) and rank-based (considering the given relations) decision rules. The experiments demonstrate that the feature-based (especially normal-based) classifiers often outperform the rank-based ones. This is to be expected, since summation-based distances are, under general conditions, approximately normally distributed. In addition, the support vector classifier also achieves a high accuracy.
Elzbieta Pekalska, Robert P. W. Duin
ICPR2
2000 The Adaptive Subspace Map for Texture Segmentation
abstract
A nonlinear mixture-of-subspaces model is proposed to describe images. Images or image patches, when translated, rotated or scaled, lie in low-dimensional subspaces of the high-dimensional space spanned by the grey values. These manifolds can locally be approximated by a linear subspace. The adaptive subspace map is a method to learn such a mixture-of-subspaces from the data. Due to its general nature, various clustering and subspace-finding algorithms can be used. In the paper, two clustering algorithms are compared in an application to some texture segmentation problems. It is shown to compare well to a standard Gabor filter bank approach.
Dick de Ridder, Josef Kittler, Olaf Lemmers, Robert P. W. Duin
ICPR4
2000 The Role of Subclasses in Machine Diagnostics
abstract
In machine diagnostics it is difficult to collect for learning all possible operating modes of machine functioning. Some of the operating modes are often missing. In these circumstances, it is important to know which modes (subclasses) are the most valuable for successful machine diagnosis. It is also of interest to investigate the usefulness of noise injection to cover the missing operating modes in the data. In this paper, we study the importance of selecting different operating modes of a water-pump and using them for learning in both 2-class and 4-class problems. We show that the operating modes representing different running speeds are more valuable than those representing machine loads. We also demonstrate that the 2-nearest neighbours directed noise injection is useful when filing in missing operating modes in the data.
Marina Skurichina, Alexander Ypma, Robert P. W. Duin
ICPR3
2000 Data Description in Subspaces
abstract
We investigate how the boundary of a data set can be obtained in case of (very) low sample sizes. This boundary can be used to detect if new objects resemble the data set and therefore make the subsequent classification more confident. When a large number of training objects is available it is possible to directly estimate the density. After thresholding the probability density a boundary around the data is obtained. However, in the case of very low sample sizes, extrapolations have to be performed. In this paper we propose a simple method based on nearest neighbor distances which is capable of finding data boundaries in these low sample sizes. It appears to be especially useful when the data is distributed in subspaces.
David M. J. Tax, Robert P. W. Duin
ICPR2
2000 Statistical Pattern Recognition: A Review
abstract
The primary goal of pattern recognition is supervised or unsupervised classification. Among the various frameworks in which pattern recognition has been traditionally formulated, the statistical approach has been most intensively studied and used in practice. More recently, neural network techniques and methods imported from statistical learning theory have been receiving increasing attention. The design of a recognition system requires careful attention to the following issues: definition of pattern classes, sensing environment, pattern representation, feature extraction and selection, cluster analysis, classifier design and learning, selection of training and test samples, and performance evaluation. In spite of almost 50 years of research and development in this field, the general problem of recognizing complex patterns with arbitrary orientation, location, and scale remains unsolved. New and emerging applications, such as data mining, web searching, retrieval of multimedia data, face recognition, and cursive handwriting recognition, require robust and efficient pattern recognition techniques. The objective of this review paper is to summarize and compare some of the well-known methods used in various stages of a pattern recognition system and identify research topics and applications which are at the forefront of this exciting and challenging field.
Anil K. Jain 0001, Robert P. W. Duin, Jianchang Mao
IEEE Trans. Pattern Anal. Mach. Intell.2
2000 Combining multiple classifiers by averaging or by multiplying?
David M. J. Tax, Martijn van Breukelen, Robert P. W. Duin, Josef Kittler
Pattern Recognit.3
2000 k-nearest neighbors directed noise injection in multilayer perceptron training
abstract
The relation between classifier complexity and learning set size is very important in discriminant analysis. One of the ways to overcome the complexity control problem is to add noise to the training objects, increasing in this way the size of the training set. Both the amount and the directions of noise injection are important factors which determine the effectiveness for classifier training. In this paper the effect is studied of the injection of Gaussian spherical noise and -nearest neighbors directed noise on the performance of multilayer perceptrons. As it is impossible to provide an analytical investigation for multilayer perceptrons, a theoretical analysis is made for statistical classifiers. The goal is to get a better understanding of the effect of noise injection on the accuracy of sample-based classifiers. By both empirical as well as theoretical studies, it is shown that the -nearest neighbors directed noise injection is preferable over the Gaussian spherical noise injection for data with low intrinsic dimensionality.
Marina Skurichina, Sarunas Raudys, Robert P. W. Duin
IEEE Trans. Neural Networks Learn. Syst.3
1999 Data domain description using support vectors
David M. J. Tax, Robert P. W. Duin
ESANN2
1999 Pump Failure Detection Using Support Vector Data Descriptions
David M. J. Tax, Alexander Ypma, Robert P. W. Duin
IDA3
1999 The Applicability of Neural Networks to Non-linear Image Processing
Dick de Ridder, Robert P. W. Duin, Piet W. Verbeek
Pattern Anal. Appl.2
1999 Regularisation of Linear Classifiers by Adding Redundant Features
Marina Skurichina, Robert P. W. Duin
Pattern Anal. Appl.2
1999 Relational discriminant analysis
Robert P. W. Duin, Elzbieta Pekalska, Dick de Ridder
Pattern Recognit. Lett.1
1999 Support vector domain description
David M. J. Tax, Robert P. W. Duin
Pattern Recognit. Lett.2
1998 On the Application of Neural Networks to Non-Linear Image Processing Tasks
Dick de Ridder, Robert P. W. Duin, Piet W. Verbeek, Lucas J. van Vliet
ICONIP2
1998 Neural network initialization by combined classifiers
abstract
If a set of linear classifiers in the same feature spaces is combined by a linear output classifier and if each of these classifiers has a sigmoid output-function then this set of classifiers has the same architecture as a feedforward neural network. A combined set of classifiers, however is trained in an entirely different way. In this paper it is shown that it can be advantageous to use such a set as an initialization for a neural network.
Martijn van Breukelen, Robert P. W. Duin
ICPR2
1998 Relational discriminant analysis and its large sample size problem
abstract
Relational discriminant analysis is based on a similarity matrix of the training set. It is able to construct reliable nonlinear discriminants in infinite dimensional feature spaces based on small training sets. This technique has a large sample size problem as the size of the similarity matrix equals the square of the number of objects in the training set. We discuss and initially evaluate a solution that drastically decreases training times and memory demands.
Robert P. W. Duin
ICPR1
1998 On Combining Classifiers
abstract
We develop a common theoretical framework for combining classifiers which use distinct pattern representations and show that many existing schemes can be considered as special cases of compound classification where all the pattern representations are used jointly to make a decision. An experimental comparison of various classifier combination schemes demonstrates that the combination rule developed under the most restrictive assumptions-the sum rule-outperforms other classifier combinations schemes. A sensitivity analysis of the various schemes to estimation errors is carried out to show that this finding can be justified theoretically.
Josef Kittler, Mohamad Hatef, Robert P. W. Duin, Jiri Matas
IEEE Trans. Pattern Anal. Mach. Intell.3
1998 Bagging for linear classifiers
Marina Skurichina, Robert P. W. Duin
Pattern Recognit.2
1998 Expected classification error of the Fisher linear classifier with pseudo-inverse covariance matrix
Sarunas Raudys, Robert P. W. Duin
Pattern Recognit. Lett.2
1997 Neural network experiences between perceptrons and support vectors
Robert P. W. Duin, Dick de Ridder
BMVC1
1997 Novelty Detection Using Self-Organizing Maps
Alexander Ypma, Robert P. W. Duin
ICONIP (2)2
1997 Experiments with a featureless approach to pattern recognition
Robert P. W. Duin, Dick de Ridder, David M. J. Tax
Pattern Recognit. Lett.1
1997 Investigating redundancy in feed-forward neural classifiers
Aarnoud Hoekstra, Robert P. W. Duin
Pattern Recognit. Lett.2
1997 Sammon's mapping using neural networks: A comparison
Dick de Ridder, Robert P. W. Duin
Pattern Recognit. Lett.2
1996 Estimating the Reliability of Neural Network Classifications
Aarnoud Hoekstra, Servaes A. Tholen, Robert P. W. Duin
ICANN3
1996 On the nonlinearity of pattern classifiers
abstract
This paper presents a novel approach to the analysis of the overtraining phenomenon in pattern classifiers. A nonlinearity measure /spl Nscr/ is introduced which relates the shape of the classification function to the generalization capability of a classifier. Experiments using the k-nearest neighbour rule, a neural classifier and the quadratic classifier show that the introduced measure /spl Nscr/ can be used to study the overtraining behaviour of a classifier. Moreover /spl Nscr/ shows to be a predictor for the local sensitivity of a classifier. Classifiers that have a small local sensitivity are shown to have a low nonlinearity whereas an increased nonlinearity indicates an increase in local sensitivity.
Aarnoud Hoekstra, Robert P. W. Duin
ICPR2
1996 Combining classifiers
abstract
We develop a common theoretical framework for combining classifiers which use distinct pattern representations and show that many existing schemes can be considered as special cases of compound classification where all the pattern representations are used jointly to make a decision. An experimental comparison of various classifier combination schemes demonstrates that the combination rule developed under the most restrictive assumptions-the sum rule-and its derivatives consistently outperform other classifier combinations schemes. A sensitivity analysis of the various schemes to estimation errors is carried out to show that this finding can be justified theoretically.
Josef Kittler, Mohamad Hatef, Robert P. W. Duin
ICPR3
1996 Stabilizing classifiers for very small sample sizes
abstract
In this paper the possibilities for constructing linear classifiers are considered for very small sample sizes. We propose a stability measure and present a study on the performance and stability of the following techniques: regularization by the ridge-estimate of the covariance matrix, bootstrapping followed by aggregation ("bagging") and editing combined with pseudo-inversion. It is shown that by these techniques a smooth transition can be made between the nearest mean classifier and the Fisher discriminant (1936, 1940) based on large samples sizes. Especially for highly correlated data very good results are obtained compared with the nearest mean method.
Marina Skurichina, Robert P. W. Duin
ICPR2
1996 A note on comparing classifiers
Robert P. W. Duin
Pattern Recognit. Lett.1
1995 SMD position measurement by a Kohonen network compared with image processing
Robert P. W. Duin, Edwin Th. G. Hoek
CAIP1
1995 An Evaluation of Intrinsic Dimensionality Estimators
abstract
The intrinsic dimensionality of a data set may be useful for understanding the properties of classifiers applied to it and thereby for the selection of an optimal classifier. In this paper the authors compare the algorithms for two estimators of the intrinsic dimensionality of a given data set and extend their capabilities. One algorithm is based on the local eigenvalues of the covariance matrix in several small regions in the feature space. The other estimates the intrinsic dimensionality from the distribution of the distances from an arbitrary data vector to a selection of its neighbors. The characteristics of the two estimators are investigated and the results are compared. It is found that both can be applied successfully, but that they might fail in certain cases. The estimators are compared and illustrated using data generated from chromosome banding profiles.>
Peter J. Verveer, Robert P. W. Duin
IEEE Trans. Pattern Anal. Mach. Intell.2
1994 The effective capacity of multilayer feedforward network classifiers
abstract
Theoretical results on the capacity, or Vapnik-Chervonenkis dimension, of a multilayer feedforward (neural) network classifier leads to much larger training sets than is used in many applications. In this paper it is shown that the effective capacity, that takes into account the training rule, is much smaller than the upper bounds on the capacity that are derived from these theoretical considerations. The success of many network applications can thereby be understood from the restricted possibilities of the optimization technique that is used for training the network.
Martin A. Kraaijveld, Robert P. W. Duin
ICPR (2)2
1994 Superlearning and neural network magic
Robert P. W. Duin
Pattern Recognit. Lett.1
1993 Evaluation method for an automatic map interpretation system for cadastral maps
abstract
A base line system for automatic interpretation of Dutch cadastral maps is discussed. Knowledge of the rules for drawing these maps is incorporated in the system. Also, a method for evaluating this system is presented. This method consists of different parts which subsequently evaluate the vectorization, the performance on finding parcels, and the performance on finding parcel numbers. Parts of this evaluation method may also be applicable for other line drawing interpretation systems.>
Rik D. T. Janssen, Robert P. W. Duin, Albert M. Vossepoel
ICDAR2
1992 Feedforward neural networks with random weights
abstract
In the field of neural network research a number of experiments described seem to be in contradiction with the classical pattern recognition or statistical estimation theory. The authors attempt to give some experimental understanding why this could be possible by showing that a large fraction of the parameters (the weights of neural networks) are of less importance and do not need to be measured with high accuracy. The remaining part is capable to implement the desired classifier and because this is only a small fraction of the total number of weights, the reported experiments seem to be more realistic from a classical point of view.>
Wouter F. Schmidt, Martin A. Kraaijveld, Robert P. W. Duin
ICPR (2)3
1990 An algorithm for benchmarking an SIMD pyramid with the Abingdon Cross
Wouter B. Teeuw, Robert P. W. Duin
Pattern Recognit. Lett.2
1986 Fast percentile filtering
Robert P. W. Duin, H. Haringa, R. Zeelen
Pattern Recognit. Lett.1
1984 Spirograph Theory: A Framework for Calculations on Digitized Straight Lines
abstract
Using diagrams called ``spirographs'' a general theory is developed with which one can easily perform calculations on various aspects of digitized straight lines. The mathematics of the theory establishes a link between digitized straight lines and the theory of numbers (Farey series, continued fractions). To show that spirograph theory is a useful unification, we derive two previously known advanced results within the framework of the theory, and new results concerning the accuracy in position of a digitized straight line as a function of its slope and length.
Leo Dorst, Robert P. W. Duin
IEEE Trans. Pattern Anal. Mach. Intell.2
1982 The use of continuous variables for labeling objects
Robert P. W. Duin
Pattern Recognit. Lett.1
1981 Dealing with a priori knowledge by fuzzy labels
Frits T. Beukema toe Water, Robert P. W. Duin
Pattern Recognit.2
1978 The mean recognition performance for independent distributions (Corresp.)
abstract
Conditions are given for the mean recognition performance over a class of independent distributions to approach unity when the dimensionality is raised to infinity.
Robert P. W. Duin
IEEE Trans. Inf. Theory1
1978 Correction to 'The Mean Recognition performance for Independent Distribution'
Robert P. W. Duin
IEEE Trans. Inf. Theory1
1978 On the evaluation of independent binary features (Corresp.)
abstract
248 IEEE TRANSACTIONS ON INFORMATION THEORY, VOL. IT-24, NO. 2, MARCH 1978 automaton that is close to optimal and eliminates the need for artificial randomization was also provided. This automaton is close to optimal in the sense that it requires at most 2 extra bits of memory, independent of m, to match the performance of the optimal randomized m-state automaton for all PA and PB. Both the problems studied here, however, involve only 2 coins. How to extend the results of this paper to situations where more than 2 coins are involved is an open question. Some ad hoc expedient automata are available in the literature [6], [7]. Before an optimal solution to the many-armed bandit problem is possible, the problem of multiple hypothesis testing with finite memory needs to be solved. For some recent results concerning this problem, see [13]. Further, finite time finite memory solutions to these problems are of interest. Vasilev [14] and Witten [15] studied the finite time behavior of some solutions to the TABP. No optimal solution, however, is available. Some recent progress has been reported by Cover et al. [16]. ACKNOWLEDGMENT The authors thank the referees for comments which helped to improve the paper. APPENDIX Denote by p(a;pA,PB) the asymptotic proportion of heads achieved, given the coins A and B and the automaton a. Even though p (a;pA,PB) is maximized over all m-state automata if and only if r (a;pA,PB) is maximized, maximizing {inf p (a;pA,PB)} is not necessarily equivalent to maximizing linf r(a;pA,Ps)} where the infimum is over {(PA,PS)}. In fact, for the TABPO, where PS is known precisely, an automaton that maximizes {infp(a;pA,PB)} tosses coin B exclusively. Furthermore, this automaton is not even expedient, and thus in some sense this solution is unsatisfactory. REFERENCES [1] H. Robbins, "Some Aspects of the sequential design of experiments," Bull. Am. Math. Soc., vol. 58, pp. 527-535, 1952. [2] H. Robbins, "A sequential decision problem with a finite memory," Proc. Nat'l. Acad. Sci., vol. 42, pp. 920-923, 1956. [3] I. H. Witten, "The apparent conflict between estimation and control--A survey of the two-armed bandit problem," J. Franklin Institute, vol. 301, no. 1-2, pp. 161-190, Jan.-Feb. 1976. [4] T. Cover and M. E. Hellman, "The two-armed bandit problem with time-invariant finite memory," IEEE Trans. Inform. Theory, vol. IT-16, No. 2, pp. 185-195, Mar. 1970. [5] M.E. Hellman and T. Cover, "Learning with finite memory," Ann. Math. Stat., vol. 41, pp. 765-782, June 1970. [6] M. L. Tsetlin, Automaton Theory and Modeling of Biological Systems. New York: Academic, 1973. [7] K.S. Fu and T. J. Li, "Formulation of learning automata and automata games," Information Sciences, vol. 1, no. 3, pp. 237-256, July 1969. [8] H. Chernoff, "Approaches in sequential design of experiments," in Statistical Design and Linear Models, J. N. Srivastava (Ed.). New York: American Elsevier, 1975, pp. 67-90. [9] S. J. Yakowitz, Mathematics of Adaptive Control Processes. New York: American Elsevier, 1969. [10] M. H. DeGroot, Optimal Statistical Decisions. New York: McGraw-Hin, 1970, Ch. 14. [11] K. B. Lakshmanan and B. Chandrasekaran, "Compound hypothesis testing with finite memory," submitted for publication. [12] A.A. Milyutin, "On automata with optimal expedient behavior in stationary media," Automation and Remote Control, vol. 26, pp. 116-131, 1965. [13] B. Chandrasekaran and K. B. Lakshmanan, "Multiple hypothesis testing with finite memory," Cybernetics and Information Science, 1977, to appear. [14] N.B. Vasilev and I. I. Pyatetskii-Shapiro, "The time for an automaton to adapt to the external medium," Automation and Remote Control, pp. 1100-1103, 1967. [15] I.H. Witten, "Finite time performance of some two-armed bandit controllers," IEEE Trans. Systems, Man, Cybern., vol. SMC-3, no. 3, pp. 194-197, Mar. 1973. [16]T. Cover, M. A. Freedman, and M. E. Hellman, "Optimal finite memory learning algorithms for the finite sample problem," Information and Control, vol. 30, pp. 49-85, Jan. 1976. On the Evaluation of Independent Binary Features ROBERT P. W. DUIN, CHRIS E. VAN HAERSMA BUMA, AND LUITZEN ROOSMA For the case of independent binary features, conditions under which the addition of a new feature does not decrease the Bayes error are derived. These conditions lead to illustrations of families of distributions for which the best two independent measurements are not the two best. I. INTRODUCTION We consider the problem of classifying the K-dimensional binary vector x = (x 1,x 2,... ,x K) into one of the two classes A and B. The features are assumed to be statistically independent for both classes so that the probability distribution of x, given class i, can be written K Fi(x): I~ {Pi xJ ']- (1 - pi)(1 -- x J)}, (1) j=l where i = A,B and Pi = Prob (xJ = l Ix e class i). The Bayes error e made by using (1) for classification is = ~. min {cFA(X), (1 -- c)FB(x)} (2) x in which c is the a priori probability for class A. The main purpose of this note is to investigate the effect upon ~ of the addition of a K + 1st feature. Conditions under which e does not decrease will be given, that is, cases in which the addition of a new feature does not result in an improvement of the probability of classification. The Bayes error (2) can be expressed in the contributions ex of all points x by = ~ exF(x), (3) x where F(x) = CFA (x) + (1 -- c)FB (x) is the probability of x. We will write the error d, if the dimensionality is raised from K to K + 1, as a sum over all points x of the K-dimensional space, let us say ~' = Y~ e'x F(x), (4) x where e'x can be interpreted as the probability of error for a given K-dimensional point :~ if the additional K + 1st feature is used. The probabilities e and e' will be compared by comparing ex and e'x for all points x. Let a(x) = (1 - c)FB(X)/IcFA(.X)} (5) be the probability ratio of class B to class A for a given point x. It can be shown (see [5[) that e'x -- ex when both pAK+i/pBK+l and (1 - pAK+i)/(1 -- pBK+l) are either simultaneously larger than a(x) or smaller than a(x). If this is valid for all x, then e' -- e, and the addition of the new feature give no improvement. When features with probabilities pA~+l and pBK+I are plotted in a (PA, PB) plane, there is an area (shaded in Fig. 1) where these conditions apply simultaneously. For the proof, see [4]. A feature in the shaded area of Fig. 1 therefore gives no improvement when it is added to the feature set. The important thing to note is that such a feature is not necessarily a feature such that pAK+l ---- p~+l. Manuscript received July 1, 1975; revised April 26, 1977. R. P. W. Duin and L. Roosma are with the Department of Applied Physics, Delft University of Technology, Delft, The Netherlands. C. E. van Haersma Buma is with Philips Audio Division, Eindhoven, The For the case of independent binary features, conditions under which the addition of a new feature does not decrease the Bayes error are derived. These conditions lead to illustrations of families of distributions for which the best two independent measurements are not the two best.
Robert P. W. Duin, Chris E. van Haersma Buma, L. Roosma
IEEE Trans. Inf. Theory1
1977 Comments on "Independence, Measurement Complexity, and Classification Performance"
abstract
Chandrasekaran and Jain [1] presented conditions under which the mean correct classification probability has a limit of one. Their conditions, however, are not sufficient to guarantee that this limit is reached monotonically. Thus local peaking is possible. Additional comments are made on their example on binary measurements.
Robert P. W. Duin
IEEE Trans. Syst. Man Cybern.1
1976 Computation of Concave Piecewise Linear Discriminant Functions Using Chebyshev Polynomials
abstract
In two-class pattern classification problems, the use of piecewise linear discriminant functions (PLDF's) is often encountered. Following a consideration of the relative advantages of a PLDF above an optimal-high degree-discriminant function, a new procedure to compute a concave PLDF is presented. The characteristic property of this procedure is that the number of linear functions necessary for adequate results of the PLDF is determined adaptively during the procedure. The linear functions are computed with Chebyshev polynominals.
Chris E. van Haersma Buma, Robert P. W. Duin
IEEE Trans. Computers2
1976 On the Choice of Smoothing Parameters for Parzen Estimators of Probability Density Functions
abstract
Parzen estimators are often used for nonparametric estimation of probability density functions. The smoothness of such an estimation is controlled by the smoothing parameter. A problem-dependent criterion for its value is proposed and illustrated by some examples. Especially in multimodal situations, this criterion led to good results.
Robert P. W. Duin
IEEE Trans. Computers1