EDBT 2026 Demo / reviewers in the wild / expert
Michel Verleysen
dblp:v/MichelVerleysen
· DBLP profile ↗
183ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0003-4366-6155ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 161 · 6 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13Databases, data management, data science and information retrieval · 7Systems, architecture and hardware · 5Applied, interdisciplinary, general and emerging computing · 2Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Interpretable Parametric Neighbour Embedding
Edouard Couplet, Pierre Lambert, Michel Verleysen, John A. Lee 0001, Cyril de Bodt |
ESANN | 3 |
| 2026 | Multi-Scale Stochastic Neighbor Embedding with Twice Adaptive BandwidthsabstractNeighbor embedding has been a quantum leap in nonlinear dimensionality reduction, revolutionizing the way data can be visualized.Neighbor embedding typically adapts to the local density in the highdimensional data space with adaptive bandwidths in entropic affinities, while it resolves scale indeterminacies by having unit bandwidths in the low-dimensional embedding space.In this paper, multi-scale stochastic neighbor embedding (Ms.SNE) is improved by allowing it to adapt lowdimensional bandwidths in a data-driven way instead of having fixed ones.In practice, Ms.SNE goes through a multi-scale optimization process; coordinates and bandwidths are optimized separately, in an alternate fashion, to avoid interferences: (i) bandwidths are optimized from previous coordinates and (ii) coordinates are optimized given the new bandwidths.Experimentally, twice adaptive bandwidths improve Ms.SNE's capability to preserve neighborhoods on all scales, i.e., local and global data structure; this claim is supported with quantitative results on several benchmarks. Neighbor embedding for data visualizationDimensionality reduction (DR) [1] yields nonlinear embeddings [2] that allow for visualization and exploratory analysis of data in many domains, such as computational biology [3], to cite just one example.Modern DR involves mostly methods of neighbor embedding (NE) [4], like Student t-distributed stochastic NE (t-SNE) [5] or uniform manifold approximation and projection (UMAP) [6].These methods are very robust to the curse of dimensionalty [7] and produce local embeddings; sparsity of small-size neighborhoods is also the key to accelerate these methods [8,6,9,10].However, sparsity might also cause the loss of the global structure of data [11,6,12,10,2,13,14].This depends on how the final embedding does reminisce [12] about its initialization with PCA [15] or Laplacian eigenmaps [16], either due to early stopping [5] or explicit regularization [13].Another workaround consists in having neighborhoods on two [9, 10] or more scales [11], even though acceleration can become more difficult.A less investigated feature of NE is a form of uniformization of data density in the low-dimensional (LD) embedding.It results from the use of entropic affinities John A. Lee 0001, Pierre Lambert, Edouard Couplet, Pierre Merveille, Dounia Mulders, Cyril de Bodt, Michel Verleysen |
ESANN | 7 |
| 2026 | Improving on early exaggeration in t -SNE: Early hierarchization better preserves global structure
John A. Lee 0001, Edouard Couplet, Pierre Lambert, Pierre Merveille, Ludovic Journaux, Dounia Mulders, Cyril de Bodt, Michel Verleysen |
Neurocomputing | 8 |
| 2025 | Can MDS rival with t-SNE by using the symmetric Kullback-Leibler divergence\\ across neighborhoods as a pseudo-distance?abstractLocal methods of dimensionality reduction like neighborhood embedding (NE) and t-SNE in particular outperform older global approaches such as stress-based multi-dimensional scaling (MDS).Stochastic neighborhoods are less sensitive than distances to statistical variations between spaces with strongly different dimensionalities, making a match across them very difficult.Here, we take inspiration from those stochastic neighborhoods in order to devise a pseudo-distance that is less prone to concentration than the Euclidean distance.For two points in the high-dimensional data space, it is defined as the symmetrized Kullback-Leibler divergence across the (stochastic) neighborhoods of the two points (SKLAN).Plugging the SKLAN in a method of stress-based MDS, we compare quantitatively t-SNE, MDS with all Euclidean distances, and MDS with SKLAN & Euclidean distances on several data sets.The results show that SKLAN allows MDS to perform competitively with t-SNE. John A. Lee 0001, Pierre Lambert, Edouard Couplet, Pierre Merveille, Ludovic Journaux, Dounia Mulders, Cyril de Bodt, Michel Verleysen |
ESANN | 8 |
| 2024 | Forget early exaggeration in t-SNE: early hierarchization preserves global structureabstractAs a local method of dimensionality reduction, t-SNE requires careful initialization in order to preserve the data global structure to the best extent.In regular t-SNE, the low-dimensional embedding is initialized either randomly or with PCA; next, gradient descent refines the embedding coordinates in two phases.In the first one, called early exaggeration, attractive forces between points are artificially strengthened to delay any detrimental effect of repulsive forces while points are still poorly organized.In this paper, a novel initialization of t-SNE is proposed.It works by hierarchizing the data points into a space-partitioning binary tree and successive runs of t-SNE with 4, 8, 16, ..., N points.Between two runs, the prototypical point in each tree branch is split into its two children prototypes, with some little random noise, and the embedding is rescaled to account for the increased population.Experimental results show the effectiveness of the method.The proposed method is compatible with any method of neighbor embedding (t-SNE, UMAP, etc.) provided early exaggeration can be disabled and initial coordinates can be fed into. John A. Lee 0001, Edouard Couplet, Pierre Lambert, Ludovic Journaux, Dounia Mulders, Cyril de Bodt, Michel Verleysen |
ESANN | 7 |
| 2024 | Investigating latent representations and generalization in deep neural networks for tabular data
Edouard Couplet, Pierre Lambert, Michel Verleysen, John A. Lee 0001, Cyril de Bodt |
Neurocomputing | 3 |
| 2023 | On the number of latent representations in deep neural networks for tabular dataabstractMost recent deep neural network architectures for tabular data operate at the feature level and process multiple latent representations simultaneously.While the dimension of these representations is set through hyper-parameter tuning, their number is typically fixed and equal to the number of features in the original data.In this paper, we explore the impact of varying the number of latent representations on model performance.Our results suggest that increasing the number of representations beyond the number of features can help capture more complex interactions, whereas reducing their number can improve performance in cases where there are many uninformative features. Edouard Couplet, Pierre Lambert, Michel Verleysen, John A. Lee 0001, Cyril de Bodt |
ESANN | 3 |
| 2023 | Semi-supervised t-SNE with multi-scale neighborhood preservation
Walter Serna-Serna, Cyril de Bodt, Andrés Marino Álvarez-Meza, John A. Lee 0001, Michel Verleysen, Álvaro-Ángel Orozco-Gutiérrez |
Neurocomputing | 5 |
| 2022 | SQuadMDS: A lean Stochastic Quartet MDS improving global structure preservation in neighbor embedding like t-SNE and UMAP
Pierre Lambert, Cyril de Bodt, Michel Verleysen, John A. Lee 0001 |
Neurocomputing | 3 |
| 2022 | Fast Multiscale Neighbor EmbeddingabstractDimension reduction (DR) computes faithful low-dimensional (LD) representations of high-dimensional (HD) data. Outstanding performances are achieved by recent neighbor embedding (NE) algorithms such as t -SNE, which mitigate the curse of dimensionality. The single-scale or multiscale nature of NE schemes drives the HD neighborhood preservation in the LD space (LDS). While single-scale methods focus on single-sized neighborhoods through the concept of perplexity, multiscale ones preserve neighborhoods in a broader range of sizes and account for the global HD organization to define the LDS. For both single-scale and multiscale methods, however, their time complexity in the number of samples is unaffordable for big data sets. Single-scale methods can be accelerated by relying on the inherent sparsity of the HD similarities they involve. On the other hand, the dense structure of the multiscale HD similarities prevents developing fast multiscale schemes in a similar way. This article addresses this difficulty by designing randomized accelerations of the multiscale methods. To account for all levels of interactions, the HD data are first subsampled at different scales, enabling to identify small and relevant neighbor sets for each data point thanks to vantage-point trees. Afterward, these sets are employed with a Barnes-Hut algorithm to cheaply evaluate the considered cost function and its gradient, enabling large-scale use of multiscale NE schemes. Extensive experiments demonstrate that the proposed accelerations are, statistically significantly, both faster than the original multiscale methods by orders of magnitude, and better preserving the HD neighborhoods than state-of-the-art single-scale schemes, leading to high-quality LD embeddings. Public codes are freely available at https://github.com/cdebodt. Cyril de Bodt, Dounia Mulders, Michel Verleysen, John A. Lee 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Stochastic quartet approach for fast multidimensional scalingabstractMultidimensional scaling is a statistical process that aims to embed high-dimensional data into a lower-dimensional, more manageable space.Common MDS algorithms tend to have some limitations when facing large data sets due to their high time and spatial complexities.This paper attempts to tackle the problem by using a stochastic approach to MDS which uses gradient descent to optimise a loss function defined on randomly designated quartets of points.This method mitigates the quadratic memory usage by computing distances on the fly, and has iterations in O(N ) time complexity, with N samples.Experiments show that the proposed method provides competitive results in reasonable time.Public codes are available at https://github.com/PierreLambert3/SQuaD-MDS.git. Multidimensional scaling and its limitationsDimensionality reduction (DR) is the process of mapping high-dimensional (HD) observations into a lower-dimensional (LD) space such that the LD embedding is a faithful representation of the HD data.The main DR uses are in machine learning, to curb the curse of dimensionality, and in visualisation.Mapped data can reveal structures that would lay hidden from the human perception if left in HD.Typically, some information is lost by the DR and, therefore, each DR method has a take on what kind of information should be preserved and what can be lost.Used frequently in visualisation, t-SNE [1] aims at retaining the neighbourhood of each point according to a distance metric and a perplexity, which reflects the size of the neighbourhood to preserve.While t-SNE excels at retaining local structures, sufficiently remote points tend to be considered equally distant by the algorithm and, therefore, the larger-scale structures can be distorted.Such distortions can lead to erroneous conclusions by the human user, who might overestimate the dissimilarity between two clusters that are distant in the LD embedding.For this reason, using multiple DR paradigms in conjunction is a good practice in visualisation: another embedding that preserves distances instead of neighbourhoods would have prevented this erroneous conclusion.This paper considers metric multidimensional scaling (MDS): a DR technique that produces a LD embedding such that the pairwise distances in LD reflect those in HD.MDS minimises a cost function which, in its simplest form, is the sum of the squared differences between distances in HD and the Euclidean distances in LD.A common strategy to optimize this cost function is based on 417 Pierre Lambert, Cyril de Bodt, Michel Verleysen, John A. Lee 0001 |
ESANN | 3 |
| 2021 | Impact of data subsamplings in Fast Multi-Scale Neighbor EmbeddingabstractFast multi-scale neighbor embedding (f-ms-NE) is an algorithm that maps high-dimensional data to a low-dimensional space by preserving the multi-scale data neighborhoods.To lower its time complexity, f-ms-NE uses random subsamplings to estimate the data properties at multiple scales.To improve this estimation and study the f-ms-NE sensitivity to randomness, this paper generalizes the f-ms-NE cost function by averaging several subsamplings.Experiments reveal that this can slightly improve the quality of the embeddings while maintaining reasonable computation times.Codes are available at https://github.com/cdebodt/Fast_Multi-scale_NE. Pierre Lambert, John A. Lee 0001, Michel Verleysen, Cyril de Bodt |
ESANN | 3 |
| 2021 | Reading grid for feature selection relevance criteria in regression
Alexandra Degeest, Benoît Frénay, Michel Verleysen |
Pattern Recognit. Lett. | 3 |
| 2020 | Perplexity-free Parametric t-SNE
Francesco Crecchi, Cyril de Bodt, Michel Verleysen, John A. Lee 0001, Davide Bacciu |
ESANN | 3 |
| 2020 | Data Augmentation and Text Recognition on Khmer Historical ManuscriptsabstractAnalysis and recognition of historical documents faces many challenges, one of which is the scarcity of the ground truth data needed for most machine learning techniques, deep learning in particular. In this paper, we present a novel approach which significantly augments the word image samples generated from an existing dataset of Khmer ancient palm leaf manuscripts. Instead of segmenting real Khmer words, we combine the annotated glyphs into groups called sub-syllables. A new text recognition method is also proposed to take into account the spatially complex structure of Khmer writing. The proposed method is composed of two main modules: a feature generator and a decoder. The generator utilizes convolutional blocks, inception blocks, and also a bi-directional LSTM to encode information extracted from the input image so that it can be decoded by the attention-based decoder to predict the final text transcription. Experiments are conducted on a new dataset of groups of sub-syllables constructed from annotated glyphs of the SleukRith Set. Dona Valy, Michel Verleysen, Sophea Chhun |
ICFHR | 2 |
| 2020 | Machine Learning with Limited Size Datasets
Michel Verleysen |
IJCCI | 1 |
| 2020 | Inference of node attributes from social network assortativity
Dounia Mulders, Cyril de Bodt, Johannes Bjelland, Alex Pentland, Michel Verleysen, Yves-Alexandre de Montjoye |
Neural Comput. Appl. | 5 |
| 2020 | LASSO multi-objective learning algorithm for feature selection
Frederico Gualberto Ferreira Coelho, Marcelo Costa, Michel Verleysen, Antônio de Pádua Braga |
Soft Comput. | 3 |
| 2019 | Class-aware t-SNE: cat-SNE
Cyril de Bodt, Dounia Mulders, Daniel López Sánchez, Michel Verleysen, John A. Lee 0001 |
ESANN | 4 |
| 2019 | MAP best performances prediction for endurance runners
Ángel Campo, Marc Francaux, Laurent Baijot, Michel Verleysen |
ESANN | 4 |
| 2019 | Tensor factorization to extract patterns in multimodal EEG data
Dounia Mulders, Cyril de Bodt, Nicolas Lejeune, John A. Lee 0001, André Mouraux, Michel Verleysen |
ESANN | 6 |
| 2019 | Comparison Between Filter Criteria for Feature Selection in Regression
Alexandra Degeest, Michel Verleysen, Benoît Frénay |
ICANN (2) | 2 |
| 2019 | Text Recognition on Khmer Historical Documents using Glyph Class Map Generation with Encoder-Decoder ModelabstractIn this paper, we propose a handwritten text recognition approach on word image patches extracted from Khmer historical documents. The network consists of two main modules composing of deep convolutional and multi-dimensional recurrent blocks. We utilize the annotated information of glyph components in the word image to build a glyph class map which is to be predicted by the first module of the network call glyph class map generator. The second module of the network encodes the generated glyph class map and transform it into a context vector which is to be decoded to produce the final word transcription. We also adapt an attention mechanism to the decoder to take advantage of local contexts which are also provided by the encoder. Experiments on a publicly available dataset of digitized Khmer palm leaf manuscripts called SleukRith set are conducted. Dona Valy, Michel Verleysen, Sophea Chhun |
ICPRAM | 2 |
| 2019 | Semi-supervised relevance index for feature selection
Frederico Gualberto Ferreira Coelho, Cristiano Leite Castro, Antônio de Pádua Braga, Michel Verleysen |
Neural Comput. Appl. | 4 |
| 2019 | Nonlinear Dimensionality Reduction With Missing Data Using Parametric Multiple ImputationsabstractDimensionality reduction (DR) aims at faithfully and meaningfully representing high-dimensional (HD) data into a low-dimensional (LD) space. Recently developed neighbor embedding DR methods lead to outstanding performances, thanks to their ability to foil the curse of dimensionality. Unfortunately, they cannot be directly employed on incomplete data sets, which become ubiquitous in machine learning. Discarding samples with missing features prevents their LD coordinates computation and deteriorates the complete samples treatment. Common missing data imputation schemes are not appropriate in the nonlinear DR context either. Indeed, even if they model the data distribution in the feature space, they can, at best, enable the application of a DR scheme on the expected data set. In practice, one would, instead, like to obtain the LD embedding with the closest cost function value on average with respect to the complete data case. As the state-of-the-art DR techniques are nonlinear, the latter embedding results from minimizing the expected cost function on the incomplete database, not from considering the expected data set. This paper addresses these limitations by developing a general methodology for nonlinear DR with missing data, being directly applicable with any DR scheme optimizing some criterion. In order to model the feature dependences, an HD extension of Gaussian mixture models is first fitted on the incomplete data set. It is afterward employed under the multiple imputation paradigms to obtain a single relevant LD embedding, thus minimizing the cost function expectation. Extensive experiments demonstrate the superiority of the suggested framework over alternative approaches. Cyril de Bodt, Dounia Mulders, Michel Verleysen, John A. Lee 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | Perplexity-free t-SNE and twice Student tt-SNE
Cyril de Bodt, Dounia Mulders, Michel Verleysen, John A. Lee 0001 |
ESANN | 3 |
| 2018 | Extensive assessment of Barnes-Hut t-SNE
Cyril de Bodt, Dounia Mulders, Michel Verleysen, John A. Lee 0001 |
ESANN | 3 |
| 2018 | ICFHR 2018 Competition On Document Image Analysis Tasks for Southeast Asian Palm Leaf ManuscriptsabstractThis paper presents the results of the Competition on Document Image Analysis Tasks for Southeast Asian Palm Leaf Manuscripts that was organized in the context of the 16th International Conference on Frontiers in Handwriting Recognition (ICFHR-2018). For this competition, three different corpus of palm leaf manuscripts written in three different scripts and languages (Balinese, Sundanese and Khmer) are used. Four Document Image Analysis (DIA) tasks are proposed as the challenges in this competition: binarization, text line segmentation, solated character/glyph recognition, and word transliteration. The results of this competition will be very useful in benchmarking analysis for the collection of palm leaf manuscripts, accelerating, evaluating and improving the performance of existing DIA system for a new type of document collection. This paper describes the competition details including the dataset, the evaluation measures used, a short description of each participant as well as the performance of the all submitted methods Made Windu Antara Kesiman, Dona Valy, Jean-Christophe Burie, Erick Paulus, Mira Suryani, Setiawan Hadi, Michel Verleysen, Sophea Chhun, Jean-Marc Ogier |
ICFHR | 7 |
| 2018 | Character and Text Recognition of Khmer Historical Palm Leaf ManuscriptsabstractThis paper presents methods for two historical document analysis tasks on digitized Khmer palm leaf manuscripts. The first task consisting of isolated character recognition is conducted utilizing different types of neural network architectures such as CNN, LSTM-RNN, and a combination of both. The second task focuses on recognizing word/text image patches of variable length and simultaneously localizing each glyph in the text image. For this task, according to the characteristic of Khmer writing system, both one-dimensional and two-dimensional RNN are used. Dona Valy, Michel Verleysen, Sophea Chhun, Jean-Christophe Burie |
ICFHR | 2 |
| 2018 | Linear Periodic Discriminant Analysis of Multidimensional Signals
Dounia Mulders, Cyril de Bodt, Nicolas Lejeune, André Mouraux, Michel Verleysen |
ICONIP (6) | 5 |
| 2017 | Kernel-based dimensionality reduction using Renyi's α-entropy measures of similarity
Andrés Marino Álvarez-Meza, John A. Lee 0001, Michel Verleysen, Germán Castellanos-Domínguez |
Neurocomputing | 3 |
| 2017 | Clustering Smart Card Data for Urban Mobility AnalysisabstractSmart card data gathered by automated fare collection (AFC) systems are valuable resources for studying urban mobility. In this paper, we propose two approaches to cluster smart card data, which can be used to extract mobility patterns in a public transportation system. Two complementary standpoints are considered: a station-oriented operational point of view and a passenger-focused one. The first approach clusters stations based on when their activity occurs, i.e., how trips made at the stations are distributed over time. The second approach makes it possible to identify groups of passengers that have similar boarding times aggregated into weekly profiles. By applying our approaches to a real data set issued from the metropolitan area of Rennes, France, we illustrate how they can help reveal valuable insights about urban mobility, such as the presence of different station key roles, including residential stations used mostly in the mornings and work stations used only in the evening and almost exclusively during weekdays, as well as different passenger behaviors ranging from the sporadic and diffuse usage to typical commute practices. By cross comparing passenger clusters with fare types, we also highlight how certain usages are more specific to particular types of passengers. Mohamed Khalil El Mahrsi, Etienne Côme, Latifa Oukhellou, Michel Verleysen |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2016 | A state-space model on interactive dimensionality reduction
Ignacio Díaz Blanco, Abel Alberto Cuadrado Vega, Michel Verleysen |
ESANN | 3 |
| 2016 | Line Segmentation Approach for Ancient Palm Leaf Manuscripts Using Competitive Learning AlgorithmabstractLine segmentation is very crucial in handwritten text recognition/analysis task. A new text line extraction scheme based on a data clustering algorithm is proposed. Our approach starts by determining the number of lines and setting up text line mid points' initial positions using a modified piece-wise projection profile technique. We apply afterwards competitive learning algorithm to adaptively move those mid points according to the geometrical information of connected components in the document page to form lines. Borders between text lines are defined so that they can be used to separate touching components that spread over multiple lines. The proposed method is robust in handling documents with skewed, fluctuated, or discontinued text lines. Experimental evaluations were made on a data set of Khmer ancient palm leaf manuscripts. Dona Valy, Michel Verleysen, Kimheng Sok |
ICFHR | 2 |
| 2016 | Multi-step-ahead forecasting using kernel adaptive filteringabstractAccurate prediction of time series is of great interest because it can guide decisions in many economical or industrial fields. Methods can forecast either one or several steps ahead. The former is simpler and more common in many applications, while the latter is more challenging. The literature describes mainly three strategies of multi-step-ahead prediction: iterated (repeated one-step-ahead), direct, and MIMO (multiple input, multiple output). This paper proposes a MIMO strategy based on kernel adaptive filtering, which we named MSAKAF. The proposed MSAKAF has shown to be a more effective method in short, medium and long-term forecast. The proposed approach is validated on two real-world datasets. The results show that our proposal outperforms the compared baseline methods in terms of prediction accuracy. Sergio García-Vega, Germán Castellanos-Domínguez, Michel Verleysen, John A. Lee 0001 |
IJCNN | 3 |
| 2016 | Reinforced Extreme Learning Machines for Fast Robust Regression in the Presence of OutliersabstractExtreme learning machines (ELMs) are fast methods that obtain state-of-the-art results in regression. However, they are not robust to outliers and their meta-parameter (i.e., the number of neurons for standard ELMs and the regularization constant of output weights for L2 -regularized ELMs) selection is biased by such instances. This paper proposes a new robust inference algorithm for ELMs which is based on the pointwise probability reinforcement methodology. Experiments show that the proposed approach produces results which are comparable to the state of the art, while being often faster. Benoît Frénay, Michel Verleysen |
IEEE Trans. Cybern. | 2 |
| 2015 | Geometrical homotopy for data visualization
Diego Hernán Peluffo-Ordóñez, Juan C. Alvarado-Pérez, John A. Lee 0001, Michel Verleysen |
ESANN | 4 |
| 2015 | Feature ranking in changing environments where new features are introducedabstractFeature selection and taking into account dynamic environments are two important aspects of modern data analysis and machine learning. In particular, performing feature selection on datasets where the latest instances contain more features than the initial ones is a problem that may be encountered in many application areas where new sensors are acquired. This paper proposes a method for incremental feature selection with rankings combining the information extracted before and after the introduction of new features, even when the number of instances that include these new features is small. Results on three real-world datasets show that using the ranking of features on the original, smaller-dimensional dataset improves the feature selection results performed on the new, larger-dimensional dataset. Alexandra Degeest, Michel Verleysen, Benoît Frénay |
IJCNN | 2 |
| 2015 | Incremental classification of objects in scenes: Application to the delineation of images
Guillaume Bernard 0001, Michel Verleysen, John A. Lee 0001 |
Neurocomputing | 2 |
| 2015 | Multi-scale similarities in stochastic neighbour embedding: Reducing dimensionality while preserving both local and global structure
John A. Lee 0001, Diego Hernán Peluffo-Ordóñez, Michel Verleysen |
Neurocomputing | 3 |
| 2014 | Two key properties of dimensionality reduction methodsabstractDimensionality reduction aims at providing faithful low-dimensional representations of high-dimensional data. Its general principle is to attempt to reproduce in a low-dimensional space the salient characteristics of data, such as proximities. A large variety of methods exist in the literature, ranging from principal component analysis to deep neural networks with a bottleneck layer. In this cornucopia, it is rather difficult to find out why a few methods clearly outperform others. This paper identifies two important properties that enable some recent methods like stochastic neighborhood embedding and its variants to produce improved visualizations of high-dimensional data. The first property is a low sensitivity to the phenomenon of distance concentration. The second one is plasticity, that is, the capability to forget about some data characteristics to better reproduce the other ones. In a manifold learning perspective, breaking some proximities typically allow for a better unfolding of data. Theoretical developments as well as experiments support our claim that both properties have a strong impact. In particular, we show that equipping classical methods with the missing properties significantly improves their results. John A. Lee 0001, Michel Verleysen |
CIDM | 2 |
| 2014 | Generalized kernel framework for unsupervised spectral methods of dimensionality reductionabstractThis work introduces a generalized kernel perspective for spectral dimensionality reduction approaches. Firstly, an elegant matrix view of kernel principal component analysis (PCA) is described. We show the relationship between kernel PCA, and conventional PCA using a parametric distance. Secondly, we introduce a weighted kernel PCA framework followed from least-squares support vector machines (LS-SVM). This approach starts with a latent variable that allows to write a relaxed LS-SVM problem. Such a problem is addressed by a primal-dual formulation. As a result, we provide kernel alternatives to spectral methods for dimensionality reduction such as multidimensional scaling, locally linear embedding, and laplacian eigenmaps; as well as a versatile framework to explain weighted PCA approaches. Experimentally, we prove that the incorporation of a SVM model improves the performance of kernel PCA. Diego Hernán Peluffo-Ordóñez, John A. Lee 0001, Michel Verleysen |
CIDM | 3 |
| 2014 | Interactive dimensionality reduction for visual analytics
Ignacio Díaz Blanco, Abel Alberto Cuadrado Vega, Daniel Pérez 0001, Francisco J. García-Fernández, Michel Verleysen |
ESANN | 5 |
| 2014 | Multiscale stochastic neighbor embedding: Towards parameter-free dimensionality reduction
John A. Lee 0001, Diego Hernán Peluffo-Ordóñez, Michel Verleysen |
ESANN | 3 |
| 2014 | Recent methods for dimensionality reduction: A brief comparative analysis
Diego Hernán Peluffo-Ordóñez, John A. Lee 0001, Michel Verleysen |
ESANN | 3 |
| 2014 | Unsupervised Relevance Analysis for Feature Extraction and Selection - A Distance-based Approach for Feature Relevance
Diego Hernán Peluffo-Ordóñez, John A. Lee 0001, Michel Verleysen, José L. Rodríguez, Germán Castellanos-Domínguez |
ICPRAM | 3 |
| 2014 | The delta test: The 1-NN estimator as a feature selection criterionabstractFeature selection is essential in many machine learning problem, but it is often not clear on which grounds variables should be included or excluded. This paper shows that the mean squared leave-one-out error of the first-nearest-neighbour estimator is effective as a cost function when selecting input variables for regression tasks. A theoretical analysis of the estimator's properties is presented to support its use for feature selection. An experimental comparison to alternative selection criteria (including mutual information, least angle regression, and the RReliefF algorithm) demonstrates reliable performance on several regression tasks. Emil Eirola, Amaury Lendasse, Francesco Corona, Michel Verleysen |
IJCNN | 4 |
| 2014 | Pointwise probability reinforcements for robust statistical inference
Benoît Frénay, Michel Verleysen |
Neural Networks | 2 |
| 2014 | Classification in the Presence of Label Noise: A SurveyabstractLabel noise is an important issue in classification, with many potential negative consequences. For example, the accuracy of predictions may decrease, whereas the complexity of inferred models and the number of necessary training samples may increase. Many works in the literature have been devoted to the study of label noise and the development of techniques to deal with label noise. However, the field lacks a comprehensive survey on the different types of label noise, their consequences and the algorithms that consider label noise. This paper proposes to fill this gap. First, the definitions and sources of label noise are considered and a taxonomy of the types of label noise is proposed. Second, the potential consequences of label noise are discussed. Third, label noise-robust, label noise cleansing, and label noise-tolerant algorithms are reviewed. For each category of approaches, a short discussion is proposed to help the practitioner to choose the most suitable technique in its own particular field of application. Eventually, the design of experiments is also discussed, what may interest the researchers who would like to test their own algorithms. In this paper, label noise consists of mislabeled instances: no additional information is assumed to be available like e.g., confidences on labels. Benoît Frénay, Michel Verleysen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2013 | Risk Estimation and Feature Selection
Gauthier Doquire, Benoît Frénay, Michel Verleysen |
ESANN | 3 |
| 2013 | Sensitivity to parameter and data variations in dimensionality reduction techniques
Francisco J. García-Fernández, Michel Verleysen, John A. Lee 0001, Ignacio Díaz Blanco |
ESANN | 2 |
| 2013 | Nonlinear Dimensionality Reduction for Visualization
Michel Verleysen, John A. Lee 0001 |
ICONIP (1) | 1 |
| 2013 | A graph Laplacian based approach to semi-supervised feature selection for regression problems
Gauthier Doquire, Michel Verleysen |
Neurocomputing | 2 |
| 2013 | Mutual information-based feature selection for multilabel classification
Gauthier Doquire, Michel Verleysen |
Neurocomputing | 2 |
| 2013 | Theoretical and empirical study on the potential inadequacy of mutual information for feature selection in classification
Benoît Frénay, Gauthier Doquire, Michel Verleysen |
Neurocomputing | 3 |
| 2013 | Feature selection for nonlinear models with extreme learning machines
Benoît Frénay, Mark van Heeswijk, Yoan Miché, Michel Verleysen, Amaury Lendasse |
Neurocomputing | 4 |
| 2013 | Type 1 and 2 mixtures of Kullback-Leibler divergences as cost functions in dimensionality reduction based on similarity preservation
John A. Lee 0001, Emilie Renard, Guillaume Bernard 0001, Pierre Dupont, Michel Verleysen |
Neurocomputing | 5 |
| 2013 | Distance estimation in numerical data sets with missing values
Emil Eirola, Gauthier Doquire, Michel Verleysen, Amaury Lendasse |
Inf. Sci. | 3 |
| 2013 | Is mutual information adequate for feature selection in regression?
Benoît Frénay, Gauthier Doquire, Michel Verleysen |
Neural Networks | 3 |
| 2012 | Incremental feature building and classification for image segmentation
Guillaume Bernard 0001, Michel Verleysen, John A. Lee 0001 |
ESANN | 2 |
| 2012 | Cluster homogeneity as a semi-supervised principle for feature selection using mutual information
Frederico Gualberto Ferreira Coelho, Antônio de Pádua Braga, Michel Verleysen |
ESANN | 3 |
| 2012 | On the Potential Inadequacy of Mutual Information for Feature Selection
Benoît Frénay, Gauthier Doquire, Michel Verleysen |
ESANN | 3 |
| 2012 | The stability of feature selection and class prediction from ensemble tree classifiers
Jérôme Paul, Michel Verleysen, Pierre Dupont |
ESANN | 2 |
| 2012 | Handling Imprecise Labels in Feature Selection with Graph Laplacian
Gauthier Doquire, Michel Verleysen |
ICPRAM (1) | 2 |
| 2012 | A Comparison of Multivariate Mutual Information Estimators for Feature Selection
Gauthier Doquire, Michel Verleysen |
ICPRAM (1) | 2 |
| 2012 | Feature selection with missing data using mutual information estimators
Gauthier Doquire, Michel Verleysen |
Neurocomputing | 2 |
| 2011 | Feature Selection with Mutual Information for Uncertain Data
Gauthier Doquire, Michel Verleysen |
DaWaK | 2 |
| 2011 | Mutual information for feature selection with missing data
Gauthier Doquire, Michel Verleysen |
ESANN | 2 |
| 2011 | Mutual information based feature selection for mixed data
Gauthier Doquire, Michel Verleysen |
ESANN | 2 |
| 2011 | Class-Specific Feature Selection for One-Against-All Multiclass SVMs
Gaël de Lannoy, Damien François, Michel Verleysen |
ESANN | 3 |
| 2011 | Label Noise-Tolerant Hidden Markov Models for Segmentation: Application to ECGs
Benoît Frénay, Gaël de Lannoy, Michel Verleysen |
ECML/PKDD (1) | 3 |
| 2011 | Mode estimation in high-dimensional spaces with flat-top kernels: Application to image denoising
Arnaud de Decker, Damien François, Michel Verleysen, John A. Lee 0001 |
Neurocomputing | 3 |
| 2011 | Parameter-insensitive kernel in extreme learning for non-linear support vector regression
Benoît Frénay, Michel Verleysen |
Neurocomputing | 2 |
| 2010 | Multi-Objective Semi-Supervised Feature Selection and Model Selection Based on Pearson's Correlation Coefficient
Frederico Gualberto Ferreira Coelho, Antônio de Pádua Braga, Michel Verleysen |
CIARP | 3 |
| 2010 | Self Organizing Star (SOS) for health monitoring
Etienne Côme, Marie Cottrell, Michel Verleysen, Jérôme Lacaille |
ESANN | 3 |
| 2010 | Mode estimation in high-dimensional spaces with flat-top kernels: application to image denoising
Arnaud de Decker, John A. Lee 0001, Damien François, Michel Verleysen |
ESANN | 4 |
| 2010 | Using SVMs with randomised feature spaces: an extreme learning approach
Benoît Frénay, Michel Verleysen |
ESANN | 2 |
| 2010 | Ensemble Modeling with a Constrained Linear System of Leave-One-Out Outputs
Yoan Miché, Emil Eirola, Patrick Bas, Olli Simula, Christian Jutten, Amaury Lendasse, Michel Verleysen |
ESANN | 7 |
| 2010 | Recent Advances in Nonlinear Dimensionality Reduction, Manifold and Topological Learning
Axel Wismüller, Michel Verleysen, Michaël Aupetit 0001, John A. Lee 0001 |
ESANN | 2 |
| 2010 | Unsupervised dimensionality reduction: Overview and recent advancesabstractUnsupervised dimensionality reduction aims at representing high-dimensional data in lower-dimensional spaces in a faithful way. Dimensionality reduction can be used for compression or denoising purposes, but data visualization remains one its most prominent applications. This paper attempts to give a broad overview of the domain. Past developments are briefly introduced and pinned up on the time line of the last eleven decades. Next, the principles and techniques involved in the major methods are described. A taxonomy of the methods is suggested, taking into account various properties. Finally, the issue of quality assessment is briefly dealt with. John A. Lee 0001, Michel Verleysen |
IJCNN | 2 |
| 2010 | Dimensionality reduction by rank preservationabstractDimensionality reduction techniques aim at representing high-dimensional data in low-dimensional spaces. To be faithful and reliable, the representation is usually required to preserve proximity relationships. In practice, methods like multidimensional scaling try to fulfill this requirement by preserving pairwise distances in the low-dimensional representation. However, such a simplification does not easily allow for local scalings in the representation. It also makes these methods suboptimal with respect to recent quality criteria that are based on distance rankings. This paper addresses this issue by introducing a dimensionality reduction method that works with ranks. Appropriate hypotheses enable the minimization of a rank-based cost function. In particular, the scale indeterminacy that is inherent to ranks is circumvented by representing data on a space with a spherical topology. Victor Onclinx, John A. Lee 0001, Vincent Wertz, Michel Verleysen |
IJCNN | 4 |
| 2010 | A principled approach to image denoising with similarity kernels involving patches
Arnaud de Decker, John A. Lee 0001, Michel Verleysen |
Neurocomputing | 3 |
| 2010 | Scale-independent quality criteria for dimensionality reduction
John A. Lee 0001, Michel Verleysen |
Pattern Recognit. Lett. | 2 |
| 2009 | Patch-based bilateral filter and local m-smoother for image denoising
Arnaud de Decker, John A. Lee 0001, Michel Verleysen |
ESANN | 3 |
| 2009 | Improving the transition modelling in hidden Markov models for ECG segmentation
Benoît Frénay, Gaël de Lannoy, Michel Verleysen |
ESANN | 3 |
| 2009 | Supervised variable clustering for classification of NIR spectra
Catherine Krier, Damien François, Fabrice Rossi, Michel Verleysen |
ESANN | 4 |
| 2009 | Adaptive anisotropic denoising: a bootstrapped procedure
John A. Lee 0001, Arnaud de Decker, Michel Verleysen |
ESANN | 3 |
| 2009 | Simbed: Similarity-Based Embedding
John A. Lee 0001, Michel Verleysen |
ICANN (2) | 2 |
| 2009 | K nearest neighbours with mutual information for simultaneous classification and missing data imputation
Pedro J. García-Laencina, José-Luis Sancho-Gómez, Aníbal R. Figueiras-Vidal, Michel Verleysen |
Neurocomputing | 4 |
| 2009 | Information-theoretic feature selection for functional data classification
Vanessa Gómez-Verdejo, Michel Verleysen, Jérôme Fleury |
Neurocomputing | 2 |
| 2009 | Quality assessment of dimensionality reduction: Rank-based criteria
John A. Lee 0001, Michel Verleysen |
Neurocomputing | 2 |
| 2009 | Residual variance estimation in machine learning
Elia Liitiäinen, Michel Verleysen, Francesco Corona, Amaury Lendasse |
Neurocomputing | 2 |
| 2009 | Nonlinear data projection on non-Euclidean manifolds with controlled trade-off between trustworthiness and continuity
Victor Onclinx, Vincent Wertz, Michel Verleysen |
Neurocomputing | 3 |
| 2008 | Using the Delta Test for Variable Selection
Emil Eirola, Elia Liitiäinen, Amaury Lendasse, Francesco Corona, Michel Verleysen |
ESANN | 5 |
| 2008 | K-nearest neighbours based on mutual information for incomplete data classification
Pedro J. García-Laencina, José-Luis Sancho-Gómez, Aníbal R. Figueiras-Vidal, Michel Verleysen |
ESANN | 4 |
| 2008 | Rank-based quality assessment of nonlinear dimensionality reduction
John A. Lee 0001, Michel Verleysen |
ESANN | 2 |
| 2008 | Nonlinear data projection on a sphere with controlled trade-off between trustworthiness and continuity
Victor Onclinx, Vincent Wertz, Michel Verleysen |
ESANN | 3 |
| 2008 | An Alternative to Center-Based Clustering Algorithm Via Statistical Learning Analysis
Rui Nian, Guangrong Ji, Michel Verleysen |
ICIC (2) | 3 |
| 2008 | Improving the Robustness to Outliers of Mixtures of Probabilistic PCAs
Nicolas Delannay, Cédric Archambeau, Michel Verleysen |
PAKDD | 3 |
| 2008 | Mixtures of robust probabilistic principal component analyzers
Cédric Archambeau, Nicolas Delannay, Michel Verleysen |
Neurocomputing | 3 |
| 2008 | Collaborative filtering with interlaced generalized linear models
Nicolas Delannay, Michel Verleysen |
Neurocomputing | 2 |
| 2008 | Blind source separation based on endpoint estimation with application to the MLSP 2006 data competition
John A. Lee 0001, Frédéric Vrins, Michel Verleysen |
Neurocomputing | 3 |
| 2007 | Mixtures of robust probabilistic principal component analyzers
Cédric Archambeau, Nicolas Delannay, Michel Verleysen |
ESANN | 3 |
| 2007 | Collaborative Filtering with interlaced Generalized Linear Models
Nicolas Delannay, Michel Verleysen |
ESANN | 2 |
| 2007 | Feature clustering and mutual information for the selection of variables in spectral data
Catherine Krier, Damien François, Fabrice Rossi, Michel Verleysen |
ESANN | 4 |
| 2007 | Resampling methods for parameter-free and robust feature selection with mutual information
Damien François, Fabrice Rossi, Vincent Wertz, Michel Verleysen |
Neurocomputing | 4 |
| 2007 | Time series prediction competition: The CATS benchmark
Amaury Lendasse, Erkki Oja, Olli Simula, Michel Verleysen |
Neurocomputing | 4 |
| 2007 | Forecasting the CATS benchmark with the Double Vector Quantization method
Geoffroy Simon, John A. Lee 0001, Marie Cottrell, Michel Verleysen |
Neurocomputing | 4 |
| 2007 | High-dimensional delay selection for regression models with mutual information and distance-to-diagonal criteria
Geoffroy Simon, Michel Verleysen |
Neurocomputing | 2 |
| 2007 | Robust Bayesian clustering
Cédric Archambeau, Michel Verleysen |
Neural Networks | 2 |
| 2007 | Mixing and Non-Mixing Local Minima of the Entropy Contrast for Blind Source SeparationabstractIn this paper, both non-mixing and mixing local minima of the entropy are analyzed from the viewpoint of blind source separation (BSS); they correspond respectively to acceptable and spurious solutions of the BSS problem. The contribution of this work is twofold. First, a Taylor development is used to show that the exact output entropy cost function has a non-mixing minimum when this output is proportional to any of the non-Gaussian sources, and not only when the output is proportional to the lowest entropic source. Second, in order to prove that mixing entropy minima exist when the source densities are strongly multimodal, an entropy approximator is proposed. The latter has the major advantage that an error bound can be provided. Even if this approximator (and the associated bound) is used here in the BSS context, it can be applied for estimating the entropy of any random variable with multimodal density Frédéric Vrins, Dinh-Tuan Pham, Michel Verleysen |
IEEE Trans. Inf. Theory | 3 |
| 2007 | The Concentration of Fractional DistancesabstractNearest neighbor search and many other numerical data analysis tools most often rely on the use of the euclidean distance. When data are high dimensional, however, the euclidean distances seem to concentrate; all distances between pairs of data elements seem to be very similar. Therefore, the relevance of the euclidean distance has been questioned in the past, and fractional norms (Minkowski-like norms with an exponent less than one) were introduced to fight the concentration phenomenon. This paper justifies the use of alternative distances to fight concentration by showing that the concentration is indeed an intrinsic property of the distances and not an artifact from a finite sample. Furthermore, an estimation of the concentration as a function of the exponent of the distance and of the distribution of the data is given. It leads to the conclusion that, contrary to what is generally admitted, fractional norms are not always less concentrated than the euclidean norm; a counterexample is given to prove this claim. Theoretical arguments are presented, which show that the concentration phenomenon can appear for real data that do not match the hypotheses of the theorems, in particular, the assumption of independent and identically distributed variables. Finally, some insights about how to choose an optimal metric are given. Damien François, Vincent Wertz, Michel Verleysen |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2007 | DD-HDS: A Method for Visualization and Exploration of High-Dimensional DataabstractMapping high-dimensional data in a low-dimensional space, for example, for visualization, is a problem of increasingly major concern in data analysis. This paper presents data-driven high-dimensional scaling (DD-HDS), a nonlinear mapping method that follows the line of multidimensional scaling (MDS) approach, based on the preservation of distances between pairs of data. It improves the performance of existing competitors with respect to the representation of high-dimensional data, in two ways. It introduces (1) a specific weighting of distances between data taking into account the concentration of measure phenomenon and (2) a symmetric handling of short distances in the original and output spaces, avoiding false neighbor representations while still allowing some necessary tears in the original distribution. More precisely, the weighting is set according to the effective distribution of distances in the data set, with the exception of a single user-defined parameter setting the tradeoff between local neighborhood preservation and global mapping. The optimization of the stress criterion designed for the mapping is realized by "force-directed placement" (FDP). The mappings of low- and high-dimensional data sets are presented as illustrations of the features and advantages of the proposed algorithm. The weighting function specific to high-dimensional data and the symmetric handling of short distances can be easily incorporated in most distance preservation-based nonlinear dimensionality reduction methods. Sylvain Lespinats, Michel Verleysen, Alain Giron, Bernard Fertil |
IEEE Trans. Neural Networks | 2 |
| 2007 | A Minimum-Range Approach to Blind Extraction of Bounded SourcesabstractIn spite of the numerous approaches that have been derived for solving the independent component analysis (ICA) problem, it is still interesting to develop new methods when, among other reasons, specific a priori knowledge may help to further improve the separation performances. In this paper, the minimum-range approach to blind extraction of bounded source is investigated. The relationship with other existing well-known criteria is established. It is proved that the minimum-range approach is a contrast, and that the criterion is discriminant in the sense that it is free of spurious maxima. The practical issues are also discussed, and a range measure estimation is proposed based on the order statistics. An algorithm for contrast maximization over the group of special orthogonal matrices is proposed. Simulation results illustrate the performances of the algorithm when using the proposed range estimation criterion. Frédéric Vrins, John A. Lee 0001, Michel Verleysen |
IEEE Trans. Neural Networks | 3 |
| 2006 | The permutation test for feature selection by mutual information
Damien François, Vincent Wertz, Michel Verleysen |
ESANN | 3 |
| 2006 | Non-orthogonal Support Width ICA
John A. Lee 0001, Frédéric Vrins, Michel Verleysen |
ESANN | 3 |
| 2006 | Determination of the Mahalanobis matrix using nonparametric noise estimations
Amaury Lendasse, Francesco Corona, Jin Hao, Nima Reyhani, Michel Verleysen |
ESANN | 5 |
| 2006 | Lag selection for regression models using high-dimensional mutual information
Geoffroy Simon, Michel Verleysen |
ESANN | 2 |
| 2006 | Effective Input Variable Selection for Function Approximation
Luis Javier Herrera, Héctor Pomares, Ignacio Rojas, Michel Verleysen, Alberto Guillén |
ICANN (1) | 4 |
| 2006 | A Functional Approach to Variable Selection in Spectrometric Problems
Fabrice Rossi, Damien François, Vincent Wertz, Michel Verleysen |
ICANN (1) | 4 |
| 2006 | Robust probabilistic projectionsabstractPrincipal components and canonical correlations are at the root of many exploratory data mining techniques and provide standard pre-processing tools in machine learning. Lately, probabilistic reformulations of these methods have been proposed (Roweis, 1998; Tipping & Bishop, 1999b; Bach & Jordan, 2005). They are based on a Gaussian density model and are therefore, like their non-probabilistic counterpart, very sensitive to atypical observations. In this paper, we introduce robust probabilistic principal component analysis and robust probabilistic canonical correlation analysis. Both are based on a Student-t density model. The resulting probabilistic reformulations are more suitable in practice as they handle outliers in a natural way. We compute maximum likelihood estimates of the parameters by means of the EM algorithm. Cédric Archambeau, Nicolas Delannay, Michel Verleysen |
ICML | 3 |
| 2006 | Assessment of probability density estimation methods: Parzen window and finite Gaussian mixturesabstractProbability density function (PDF) estimation is a very critical task in many applications of data analysis. For example in the Bayesian framework decisions are taken according to Bayes' rule, which directly involves the evaluation of the PDF. Many methods are available to this aim, but there is no consensus in the literature about which to use, nor about the pros and cons of each of them. In this paper, we present a thorough and extensive experimental comparison between two of the most popular methods: Parzen window and finite Gaussian mixture. Extended experimental results and application development guidelines are reported Cédric Archambeau, Maurizio Valle, Alex Assenza, Michel Verleysen |
ISCAS | 4 |
| 2006 | Advances in Self-Organizing Maps
Marie Cottrell, Michel Verleysen |
Neural Networks | 2 |
| 2006 | Unfolding preprocessing for meaningful time series clustering
Geoffroy Simon, John A. Lee 0001, Michel Verleysen |
Neural Networks | 3 |
| 2005 | Non-Euclidean metrics for similarity search in noisy datasets
Damien François, Vincent Wertz, Michel Verleysen |
ESANN | 3 |
| 2005 | Pruned lazy learning models for time series prediction
Antti Sorjamaa, Amaury Lendasse, Michel Verleysen |
ESANN | 3 |
| 2005 | clustering using a random walk based distance measure
Luh Yen, Denis Vanvyve, Fabien Wouters, François Fouss, Michel Verleysen, Marco Saerens |
ESANN | 5 |
| 2005 | Manifold Constrained Variational Mixtures
Cédric Archambeau, Michel Verleysen |
ICANN (2) | 2 |
| 2005 | LS-SVM Hyperparameter Selection with a Nonparametric Noise Estimator
Amaury Lendasse, Yongnan Ji, Nima Reyhani, Michel Verleysen |
ICANN (2) | 4 |
| 2005 | SWM: a class of convex contrasts for source separationabstractWe derive a class of contrasts for blind source separation (BSS) to separate bounded sources (or more generally, finite sources), based on support width measures (SWM) of the marginal output distributions. These contrasts are shown to have no spurious local maxima, i.e., all the local maxima are relevant from the source separation point of view; they all correspond to non-mixing BSS solutions so that a gradient-ascent method can be used. Frédéric Vrins, Michel Verleysen, Christian Jutten |
ICASSP (5) | 2 |
| 2005 | Vector quantization: a weighted version for time-series forecasting
Amaury Lendasse, Damien François, Vincent Wertz, Michel Verleysen |
Future Gener. Comput. Syst. | 4 |
| 2005 | Nonlinear dimensionality reduction of data manifolds with essential loops
John A. Lee 0001, Michel Verleysen |
Neurocomputing | 2 |
| 2005 | Fast bootstrap methodology for regression model selection
Amaury Lendasse, Geoffroy Simon, Vincent Wertz, Michel Verleysen |
Neurocomputing | 4 |
| 2005 | Representation of functional data in neural networks
Fabrice Rossi, Nicolas Delannay, Brieuc Conan-Guez, Michel Verleysen |
Neurocomputing | 4 |
| 2005 | Time series forecasting: Obtaining long term trends with self-organizing maps
Geoffroy Simon, Amaury Lendasse, Marie Cottrell, Jean-Claude Fort, Michel Verleysen |
Pattern Recognit. Lett. | 5 |
| 2005 | On the entropy minimization of a linear mixture of variables for source separation
Frédéric Vrins, Michel Verleysen |
Signal Process. | 2 |
| 2005 | Information theoretic versus cumulant-based contrasts for multimodal source separationabstractRecently, several authors have emphasized the existence of spurious maxima in usual contrast functions for source separation (e.g., the likelihood and the mutual information) when several sources have multimodal distributions. The aim of this letter is to compare the information theoretic contrasts to cumulant-based ones from the robustness to spurious maxima point of view. Even if all of them tend to measure, in some way, the same quantity, which is the output independence (or equivalently, the output non-Gaussianity), it is shown that in the case of a mixture involving two sources, the kurtosis-based contrast functions are more robust than the information theoretic ones when the source distributions are multimodal. Frédéric Vrins, Michel Verleysen |
IEEE Signal Process. Lett. | 2 |
| 2004 | Flexible and Robust Bayesian Classification by Finite Mixture Models
Cédric Archambeau, Frédéric Vrins, Michel Verleysen |
ESANN | 3 |
| 2004 | Functional radial basis function networks
Nicolas Delannay, Fabrice Rossi, Brieuc Conan-Guez, Michel Verleysen |
ESANN | 4 |
| 2004 | How to project 'circular' manifolds using geodesic distances?
John A. Lee 0001, Michel Verleysen |
ESANN | 2 |
| 2004 | Fast bootstrap for least-square support vector machines
Amaury Lendasse, Geoffroy Simon, Vincent Wertz, Michel Verleysen |
ESANN | 4 |
| 2004 | Towards a Local Separation Performances Estimator Using Common ICA Contrast Functions?
Frédéric Vrins, Cédric Archambeau, Michel Verleysen |
ESANN | 3 |
| 2004 | Fast bootstrap applied to LS-SVM for long term prediction of time seriesabstractTime series forecasting is usually limited to one-step ahead prediction. This goal is extended here to longer-term prediction, obtained using the least-square support vector machines model. The influence of the model parameters is observed when the time horizon of the prediction is increased and for various prediction methods. The model selection to optimize the design parameters is performed using the fast bootstrap methodology introduced in previous works. Amaury Lendasse, Vincent Wertz, Geoffroy Simon, Michel Verleysen |
IJCNN | 4 |
| 2004 | Double quantization forecasting method for filling missing data in the CATS time seriesabstractThe double vector quantization forecasting method based on Kohonen self-organizing maps is applied to predict the missing values of the CATS competition data set. As one of the features of the method is the ability to predict vectors instead of scalar values in a single step, the compromise between the size of the vector prediction and the number of repetitions needed to reach the required prediction horizon is studied. The long-term stability of the double vector quantization method makes it possible to obtain reliable values on a rather long-term forecasting horizon. Geoffroy Simon, John A. Lee 0001, Michel Verleysen, Marie Cottrell |
IJCNN | 3 |
| 2004 | Prediction of visual perceptions with artificial neural networks in a visual prosthesis for the blind
Cédric Archambeau, Jean Delbeke, Claude Veraart, Michel Verleysen |
Artif. Intell. Medicine | 4 |
| 2004 | On the use of self-organizing maps to accelerate vector quantization
Eric de Bodt, Marie Cottrell, Patrick Letrémy, Michel Verleysen |
Neurocomputing | 4 |
| 2004 | Nonlinear projection with curvilinear distances: Isomap versus curvilinear distance analysis
John A. Lee 0001, Amaury Lendasse, Michel Verleysen |
Neurocomputing | 3 |
| 2004 | Double quantization of the regressor space for long-term time series prediction: method and proof of stability
Geoffroy Simon, Amaury Lendasse, Marie Cottrell, Jean-Claude Fort, Michel Verleysen |
Neural Networks | 5 |
| 2003 | On Convergence Problems of the EM Algorithm for Finite Gaussian Mixtures
Cédric Archambeau, John A. Lee 0001, Michel Verleysen |
ESANN | 3 |
| 2003 | Locally Linear Embedding versus Isotop
John A. Lee 0001, Cédric Archambeau, Michel Verleysen |
ESANN | 3 |
| 2003 | Fast approximation of the bootstrap for model selection
Geoffroy Simon, Amaury Lendasse, Vincent Wertz, Michel Verleysen |
ESANN | 4 |
| 2003 | Model Selection with Cross-Validations and Bootstraps - Application to Time Series Prediction with RBFN Models
Amaury Lendasse, Vincent Wertz, Michel Verleysen |
ICANN | 3 |
| 2003 | On the Kernel Widths in Radial-Basis Function Networks
Nabil Benoudjit, Michel Verleysen |
Neural Process. Lett. | 2 |
| 2002 | Width optimization of the Gaussian kernels in Radial Basis Function Networks
Nabil Benoudjit, Cédric Archambeau, Amaury Lendasse, John A. Lee 0001, Michel Verleysen |
ESANN | 5 |
| 2002 | Curvilinear Distance Analysis versus Isomap
John A. Lee 0001, Amaury Lendasse, Michel Verleysen |
ESANN | 3 |
| 2002 | Nonlinear Projection with the Isotop Method
John A. Lee 0001, Michel Verleysen |
ICANN | 2 |
| 2002 | Forecasting electricity consumption using nonlinear projection and self-organizing maps
Amaury Lendasse, John A. Lee 0001, Vincent Wertz, Michel Verleysen |
Neurocomputing | 4 |
| 2002 | Special issue on fundamental and information processing aspects of neurocomputing
Michel Verleysen, Joos Vandewalle |
Neurocomputing | 1 |
| 2002 | Statistical tools to assess the reliability of self-organizing maps
Eric de Bodt, Marie Cottrell, Michel Verleysen |
Neural Networks | 3 |
| 2002 | Self-organizing maps with recursive neighborhood adaptation
John A. Lee 0001, Michel Verleysen |
Neural Networks | 2 |
| 2001 | Input data reduction for the prediction of financial time series
Amaury Lendasse, John A. Lee 0001, Eric de Bodt, Vincent Wertz, Michel Verleysen |
ESANN | 5 |
| 2000 | A robust non-linear projection method
John A. Lee 0001, Amaury Lendasse, Nicolas Donckers, Michel Verleysen |
ESANN | 4 |
| 2000 | Time series forecasting using CCA and Kohonen maps - application to electricity consumption
Amaury Lendasse, John A. Lee 0001, Vincent Wertz, Michel Verleysen |
ESANN | 4 |
| 2000 | A current-mode CMOS loser-take-all with minimum function for neural computationsabstractA novel architecture for loser-take-all functions is proposed. Inputs and outputs of the circuit are currents, which make the circuit appropriated for low-voltage neural hardware computation. In contrast to most existing realisations the circuit does not require subtraction from a fixed reference which decreases accuracy and input dynamic. Moreover, in addition to the loser, it also outputs the minimum input current. The circuit was synthesized using a SOI (silicon on insulator) technology and optimised to work with 1.5 V voltage supply showing improved speed and accuracy for a very low power consumption (typically 5 /spl mu/W per cell when the input current is 1 /spl mu/A). Nicolas Donckers, Carlos Dualibe, Michel Verleysen |
ISCAS | 3 |
| 2000 | A 5.26 Mflips programmable analogue fuzzy logic controller in a standard CMOS 2.4 μ technologyabstractA complete digitally-programmable analogue fuzzy logic controller (FLC) is presented. The design of some new functional blocks and the improvement of others aim towards speed optimisation with a reasonable accuracy, as it is needed in several analogue signal processing applications. A nine-rules, two-inputs and one-output prototype was fabricated and successfully tested using a standard CMOS 2.4 /spl mu/ technology showing good agreement with the expected performances, namely: 5.26 Mflips (mega fuzzy logic inferences per second) at the pin terminal (@CL=13 pF), 933 /spl mu/W power consumption per rule (@Vdd=5V) and 5 to 6 bits of precision. Since the circuit is intended for a subsystem embedded in an application chip (@CL/spl les/15 pF) over 8 Mflips may be expected. Carlos Dualibe, Paul G. A. Jespers, Michel Verleysen |
ISCAS | 3 |
| 2000 | A low-power silicon-on-insulator PWM discriminator for biomedical applicationsabstractA CMOS/SOI circuit to decode PWM signals is presented as part of a body-implanted neurostimulator for visual prosthesis. Since encoded data is the sole input to the circuit, the decoding technique is based on a double-integration concept and does not require DC filtering. Non-overlapping control phases are internally derived from the incoming pulses and a fast-settling comparator ensures good discrimination accuracy in the megahertz range. The circuit was integrated on a 2 /spl mu/m single-metal SOI fabrication process and has an effective area of 2mm/sup 2/. Typically, the measured resolution of the encoding parameter /spl alpha/ was better than 10% at 6 MHz and V/sub DD/=3.3 V. Stand-by consumption is around 340 /spl mu/W. Pulses with frequencies up to 15 MHz and /spl alpha/=10% can be discriminated for V/sub DD/ spanning from 2.3 V to 3.3 V. Such an excellent immunity to V/sub DD/ deviations meets a design specification with respect to inherent coupling losses on transmitting data and power by means of a transcutaneous link. Jader A. De Lima, Sidnei F. Silva, Adriano S. Cordeiro, Alexandro C. Araujo, Michel Verleysen |
ISCAS | 5 |
| 1999 | Using the Kohonen algorithm for quick initialization of Simple Competitive Learning algorithm
Eric de Bodt, Marie Cottrell, Michel Verleysen |
ESANN | 3 |
| 1999 | Extraction of intrinsic dimension using CCA - Application to blind sources separation
Nicolas Donckers, Amaury Lendasse, Vincent Wertz, Michel Verleysen |
ESANN | 4 |
| 1998 | Forecasting time-series by Kohonen classification
Amaury Lendasse, Michel Verleysen, Eric de Bodt, Marie Cottrell, Philippe Grégoire |
ESANN | 2 |
| 1998 | Image compression by self-organized Kohonen mapabstractThis paper presents a compression scheme for digital still images, by using the Kohonen's neural network algorithm, not only for its vector quantization feature, but also for its topological property. This property allows an increase of about 80% for the compression rate. Compared to the JPEG standard, this compression scheme shows better performances (in terms of PSNR) for compression rates higher than 30. C. Amerijckx, Michel Verleysen, Philippe Thissen, Jean-Didier Legat |
IEEE Trans. Neural Networks | 2 |
| 1997 | Kohonen maps versus vector quantization for data analysis
Eric de Bodt, Michel Verleysen, Marie Cottrell |
ESANN | 2 |
| 1997 | Placing spline knots in neural networks using splines as activation functions
Katerina Hlavácková-Schindler, Michel Verleysen |
Neurocomputing | 2 |
| 1995 | Suboptimal Bayesian classification by vector quantization with small clusters
Jean-Luc Voz, Michel Verleysen, Philippe Thissen, Jean-Didier Legat |
ESANN | 2 |
| 1995 | Editorial
François Blayo, Michel Verleysen |
Neural Process. Lett. | 2 |
| 1994 | Estimation of performance bounds in supervised classification
Pierre Comon, Jean-Luc Voz, Michel Verleysen |
ESANN | 3 |
| 1994 | An optimized RBF network for approximation of functions
Michel Verleysen, Katerina Hlavácková-Schindler |
ESANN | 1 |
| 1994 | Editorial
François Blayo, Michel Verleysen |
Neural Process. Lett. | 2 |
| 1993 | Laplacian pyramid with multilayer perceptrons interpolators
Benoît Simon, Benoît Macq, Michel Verleysen |
ESANN | 3 |
| 1993 | Optimal decision surfaces in LVQ1 classiffication of patterns
Michel Verleysen, Philippe Thissen, Jean-Didier Legat |
ESANN | 1 |
| 1993 | Analog implementation of a Kohonen map with on-chip learningabstractKohonen maps are self-organizing neural networks that classify and quantify n-dimensional data into a one- or two-dimensional array of neurons. Most applications of Kohonen maps use simulations on conventional computers, eventually coupled to hardware accelerators or dedicated neural computers. The small number of different operations involved in the combined learning and classification process, however, makes the Kohonen model particularly suited to a dedicated VLSI implementation, taking full advantage of the parallelism and speed that can be obtained on the chip. A fully analog implementation of a one-dimensional Kohonen map, with on-chip learning and refreshment of on-chip analog synaptic weights, is proposed. The small number of transistors in each cell allows a high degree of parallelism in the operations, which greatly improves the computation speed compared to other implementations. The storage of analog synaptic weights, based on the principle of current copiers, is emphasized. It is shown that this technique can be used successfully for the realization of VLSI Kohonen maps. Damien Macq, Michel Verleysen, Paul G. A. Jespers, Jean-Didier Legat |
IEEE Trans. Neural Networks | 2 |
| 1992 | A real-time VLSI-based architecture for multi-motion estimationabstractThis paper describes a new parallel architecture dedicated to multi-motion estimation. The input image is scanned by a standard video camera with 256 grey levels. Motion computing is based on the optical flow determination. Some constraints are proposed to allow multi-motion evaluation. The algorithm is presented and the main features of a 1-D systolic architecture which is based on a custom VLSI chip is given. This architecture allows a real-time implementation of the multi-motion estimation algorithm.> Jean-Didier Legat, J. P. Cornil, Damien Macq, Michel Verleysen |
ICPR (4) | 4 |
| 1988 | An algorithm for pattern recognition with VLSI neural networks
Bruno Sirletti, Michel Verleysen, Andre M. Vandemeulebroecke, Paul G. A. Jespers |
Neural Networks | 2 |
| 1988 | A large VLSI Hopfield network for pattern recognition problems
Michel Verleysen, Bruno Sirletti, Paul G. A. Jespers |
Neural Networks | 1 |