John A. Lee 0001

dblp:30/5819 · also John Aldo Lee · DBLP profile ↗
← Back
92ranked-venue papers
29as first author
23since 2021 · last 2026
0000-0001-5218-759XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 85 · 29 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Fairness in machine learning: A Compact Survey
abstract
Considering, assessing, and ensuring fairness is key when relying on machine learning (ML) for sensitive decision making.Yet, despite growing attention, multiple fairness definitions are currently adopted and consensual guidelines across learning paradigms are still lacking.Therefore, this review first describes sources of bias and focuses on outcome fairness, comparing popular criterion such as independence, separation, and sufficiency.Fairness interventions are then examined at pre-, in-, and post-processing stages and across supervised, semi-supervised, unsupervised, and self-supervised settings.
Jeremy de Bodt, Dounia Mulders, Cyril de Bodt, John A. Lee 0001, Marco Saerens
ESANN4
2026 Interpretable Parametric Neighbour Embedding
Edouard Couplet, Pierre Lambert, Michel Verleysen, John A. Lee 0001, Cyril de Bodt
ESANN4
2026 Multi-Scale Stochastic Neighbor Embedding with Twice Adaptive Bandwidths
abstract
Neighbor embedding has been a quantum leap in nonlinear dimensionality reduction, revolutionizing the way data can be visualized.Neighbor embedding typically adapts to the local density in the highdimensional data space with adaptive bandwidths in entropic affinities, while it resolves scale indeterminacies by having unit bandwidths in the low-dimensional embedding space.In this paper, multi-scale stochastic neighbor embedding (Ms.SNE) is improved by allowing it to adapt lowdimensional bandwidths in a data-driven way instead of having fixed ones.In practice, Ms.SNE goes through a multi-scale optimization process; coordinates and bandwidths are optimized separately, in an alternate fashion, to avoid interferences: (i) bandwidths are optimized from previous coordinates and (ii) coordinates are optimized given the new bandwidths.Experimentally, twice adaptive bandwidths improve Ms.SNE's capability to preserve neighborhoods on all scales, i.e., local and global data structure; this claim is supported with quantitative results on several benchmarks. Neighbor embedding for data visualizationDimensionality reduction (DR) [1] yields nonlinear embeddings [2] that allow for visualization and exploratory analysis of data in many domains, such as computational biology [3], to cite just one example.Modern DR involves mostly methods of neighbor embedding (NE) [4], like Student t-distributed stochastic NE (t-SNE) [5] or uniform manifold approximation and projection (UMAP) [6].These methods are very robust to the curse of dimensionalty [7] and produce local embeddings; sparsity of small-size neighborhoods is also the key to accelerate these methods [8,6,9,10].However, sparsity might also cause the loss of the global structure of data [11,6,12,10,2,13,14].This depends on how the final embedding does reminisce [12] about its initialization with PCA [15] or Laplacian eigenmaps [16], either due to early stopping [5] or explicit regularization [13].Another workaround consists in having neighborhoods on two [9, 10] or more scales [11], even though acceleration can become more difficult.A less investigated feature of NE is a form of uniformization of data density in the low-dimensional (LD) embedding.It results from the use of entropic affinities
John A. Lee 0001, Pierre Lambert, Edouard Couplet, Pierre Merveille, Dounia Mulders, Cyril de Bodt, Michel Verleysen
ESANN1
2026 Fast denoising of low-count Monte Carlo proton therapy dose distributions with ResUNet
Pierre Merveille, Ana M. Barragan-Montero, Kevin Souris, John A. Lee 0001
ESANN4
2026 Towards meaningful evaluation of uncertainty-aware segmentation workflows for medical applications
abstract
In radiation oncology, image segmentation with deep learning models must reduce clinician workload without compromising patient safety.Uncertainty quantification is therefore essential to provide reliable error estimates and to determine which segmentations require correction.Despite progress, we argue that current evaluation protocols enable only general comparison and ignore the definition of decision thresholds used in practice.We propose a new paradigm for the evaluation of such automated systems through the quantification of clinical outcomes.We compute the fraction of high-confidence segmentation prediction meeting quality standards and their corresponding average performance, based on a decision threshold.We calibrate this threshold to bound the amount of segmentations inferior to standards below a risk tolerance specified by clinicians.Through experiments across four medical datasets, we show our approach delivers meaningful performance guarantees essential for regulatory compliance and building trust in automated systems.
Dany Rimez, John A. Lee 0001, Ana M. Barragan-Montero
ESANN2
2026 Improving on early exaggeration in t -SNE: Early hierarchization better preserves global structure
John A. Lee 0001, Edouard Couplet, Pierre Lambert, Pierre Merveille, Ludovic Journaux, Dounia Mulders, Cyril de Bodt, Michel Verleysen
Neurocomputing1
2025 Can MDS rival with t-SNE by using the symmetric Kullback-Leibler divergence\\ across neighborhoods as a pseudo-distance?
abstract
Local methods of dimensionality reduction like neighborhood embedding (NE) and t-SNE in particular outperform older global approaches such as stress-based multi-dimensional scaling (MDS).Stochastic neighborhoods are less sensitive than distances to statistical variations between spaces with strongly different dimensionalities, making a match across them very difficult.Here, we take inspiration from those stochastic neighborhoods in order to devise a pseudo-distance that is less prone to concentration than the Euclidean distance.For two points in the high-dimensional data space, it is defined as the symmetrized Kullback-Leibler divergence across the (stochastic) neighborhoods of the two points (SKLAN).Plugging the SKLAN in a method of stress-based MDS, we compare quantitatively t-SNE, MDS with all Euclidean distances, and MDS with SKLAN & Euclidean distances on several data sets.The results show that SKLAN allows MDS to perform competitively with t-SNE.
John A. Lee 0001, Pierre Lambert, Edouard Couplet, Pierre Merveille, Ludovic Journaux, Dounia Mulders, Cyril de Bodt, Michel Verleysen
ESANN1
2024 Forget early exaggeration in t-SNE: early hierarchization preserves global structure
abstract
As a local method of dimensionality reduction, t-SNE requires careful initialization in order to preserve the data global structure to the best extent.In regular t-SNE, the low-dimensional embedding is initialized either randomly or with PCA; next, gradient descent refines the embedding coordinates in two phases.In the first one, called early exaggeration, attractive forces between points are artificially strengthened to delay any detrimental effect of repulsive forces while points are still poorly organized.In this paper, a novel initialization of t-SNE is proposed.It works by hierarchizing the data points into a space-partitioning binary tree and successive runs of t-SNE with 4, 8, 16, ..., N points.Between two runs, the prototypical point in each tree branch is split into its two children prototypes, with some little random noise, and the embedding is rescaled to account for the increased population.Experimental results show the effectiveness of the method.The proposed method is compatible with any method of neighbor embedding (t-SNE, UMAP, etc.) provided early exaggeration can be disabled and initial coordinates can be fed into.
John A. Lee 0001, Edouard Couplet, Pierre Lambert, Ludovic Journaux, Dounia Mulders, Cyril de Bodt, Michel Verleysen
ESANN1
2024 Estimated neighbour sets and smoothed sampled global interactions are sufficient for a fast approximate tSNE
abstract
To minimise its loss function, the popular method of nonlinear dimensionality reduction t-SNE requires O(N 2 ) computations.As its applications often involve large datasets, fast approximations have been developed, such as Barnes-Hut t-SNE and FIt-SNE.Most fast approximations to t-SNE require the embedding dimensionality to be small, typically 2 or 3, limiting the use of t-SNE to data visualisation.Additionally, the effective computation time of the current accelerated t-SNE algorithms stays too high for a comfortable interactive visual exploration of data.This paper proposes an accelerated approximation to t-SNE with iterations of complexity O(N K), which does not rely on the use of a model to capture information about the low-dimensional space, relieving the computational burden of high dimensionality of the embedding space.For this purpose, the proposed method approximates neighbour sets and keeps track of smoothed estimations of long-range interactions in O(N K) time.The method is qualitatively tested on a handful of datasets and shows comparable results to existing fast neighbour embedding methods in the context of data visualisation.Code is available at https://github.com/PierreLambert3/c_fast_hSNE.git.
Pierre Lambert, Edouard Couplet, Cyril de Bodt, John A. Lee 0001
ESANN4
2024 SDE U-Net: Disentangling Aleatoric and Epistemic Uncertainties in Medical Image Segmentation
abstract
Quantifying uncertainty is crucial in artifical intelligence (AI) applications, particularly in high-stakes healthcare settings.This paper introduces SDE U-Net, a novel architecture that integrates stochastic differential equations (SDEs) with the U-Net framework, effectively distinguishing between aleatoric and epistemic uncertainties.By incorporating a randomness component, SDE U-Net directly captures and quantifies aleatoric uncertainty, while epistemic uncertainty is assessed through multiple forward passes.Comparative results show that SDE U-Net not only matches but also exceeds benchmark performance, achieving similar results in just 500 epochs, half the epochs required by the benchmark.This approach enhances the reliability of AI in medical decision-making by providing a clear, comprehensive representation of uncertainty, marking a significant advancement in the field of medical image segmentation.
Chuxin Zhang, Ana M. Barragan-Montero, John A. Lee 0001
ESANN3
2024 Investigating latent representations and generalization in deep neural networks for tabular data
Edouard Couplet, Pierre Lambert, Michel Verleysen, John A. Lee 0001, Cyril de Bodt
Neurocomputing4
2023 Single-pass uncertainty estimation with layer ensembling for regression: application to proton therapy dose prediction for head and neck cancer
abstract
We developed a new uncertainty quantification method for deep learning regression models, based on Layer Ensembles [1], which is competitive with state-of-the-art ensembling and Monte Carlo (MC) dropout techniques.The method was implemented in a UNet-like architecture and applied to predicting 3D dose maps for head and neck cancer patients who are treated with proton therapy.The new approach runs approximately 8 times faster than MC Dropout.Our statistical analysis showed no significant difference in prediction accuracy between the two different methods (p-value = 0.09).Moreover, the correlation uncertainty/error in the body is only -3%.These findings demonstrate the potential of the new method in enabling fast and accurate uncertainty quantification for regression problems and, in particular, for proton therapy dose prediction.
Ana M. Barragan-Montero, Robin Tilman, Margerie Huet-Dastarac, John A. Lee 0001
ESANN4
2023 On the number of latent representations in deep neural networks for tabular data
abstract
Most recent deep neural network architectures for tabular data operate at the feature level and process multiple latent representations simultaneously.While the dimension of these representations is set through hyper-parameter tuning, their number is typically fixed and equal to the number of features in the original data.In this paper, we explore the impact of varying the number of latent representations on model performance.Our results suggest that increasing the number of representations beyond the number of features can help capture more complex interactions, whereas reducing their number can improve performance in cases where there are many uninformative features.
Edouard Couplet, Pierre Lambert, Michel Verleysen, John A. Lee 0001, Cyril de Bodt
ESANN4
2023 Nesterov momentum and gradient normalization to improve t-SNE convergence and neighborhood preservation, without early exaggeration
abstract
Student t-distributed stochastic neighbor embedding (t-SNE) finds low-dimensional data representations allowing visual exploration of data sets.t-SNE minimises a cost function with a custom two-phase gradient descent.The first phase is called early exaggeration and involves a hyper-parameter whose value can be tricky and time-consuming to set.This paper proposes another way to optimise the cost function without early exaggeration.Empirical evaluation shows that the proposed method of optimization converges faster and yields competitive results in terms of neighborhood preservation.
Pierre Lambert, John A. Lee 0001, Edouard Couplet, Cyril de Bodt
ESANN2
2023 Semi-supervised t-SNE with multi-scale neighborhood preservation
Walter Serna-Serna, Cyril de Bodt, Andrés Marino Álvarez-Meza, John A. Lee 0001, Michel Verleysen, Álvaro-Ángel Orozco-Gutiérrez
Neurocomputing4
2022 Tuning Database-Friendly Random Projection Matrices for Improved Distance Preservation on Specific Data
abstract
Abstract Random Projection is one of the most popular and successful dimensionality reduction algorithms for large volumes of data. However, given its stochastic nature, different initializations of the projection matrix can lead to very different levels of performance. This paper presents a guided random search algorithm to mitigate this problem. The proposed method uses a small number of training data samples to iteratively adjust a projection matrix, improving its performance on similarly distributed data. Experimental results show that projection matrices generated with the proposed method result in a better preservation of distances between data samples. Conveniently, this is achieved while preserving the database-friendliness of the projection matrix, as it remains sparse and comprised exclusively of integers after being tuned with our algorithm. Moreover, running the proposed algorithm on a consumer-grade CPU requires only a few seconds.
Daniel López Sánchez, Cyril de Bodt, John A. Lee 0001, Angélica González Arrieta, Juan M. Corchado
Appl. Intell.3
2022 Deep learning to detect bacterial colonies for the production of vaccines
Thomas Beznik, Paul Smyth, Gaël de Lannoy, John A. Lee 0001
Neurocomputing4
2022 SQuadMDS: A lean Stochastic Quartet MDS improving global structure preservation in neighbor embedding like t-SNE and UMAP
Pierre Lambert, Cyril de Bodt, Michel Verleysen, John A. Lee 0001
Neurocomputing4
2022 Compressive Imaging Through Optical Fiber with Partial Speckle Scanning
abstract
Fluorescence imaging through ultrathin fibers is a promising approach to obtain high-resolution imaging with molecular specificity at depths much larger than the scattering mean-free paths of biological tissues. Such imaging techniques, generally termed lensless endoscopy, rely upon the wavefront control at the distal end of a fiber to coherently combine multiple spatial modes of a multicore (MCF) or multimode fiber (MMF). Typically, a spatial light modulator (SLM) is employed to combine hundreds of modes by phase-matching to generate a high-intensity focal spot. This spot is subsequently scanned across the sample to obtain an image. We propose here a novel scanning scheme, partial speckle scanning (PSS), inspired by compressive sensing theory, that avoids the use of an SLM to perform fluorescent imaging with optical fibers with reduced acquisition time. Such a strategy avoids photo-bleaching while keeping high reconstruction quality. We develop our approach on two key properties of the MCF: (i) the ability to easily generate speckles, and (ii) the memory effect that allows one to use fast scan mirrors to shift light patterns. First, we show that speckles are subexponential random fields. Despite their granular structure, an appropriate choice of the reconstruction parameters makes them good candidates to build efficient sensing matrices. Then, we numerically validate our approach and apply it on experimental data. The proposed sensing technique outperforms conventional raster scanning: higher reconstruction quality is achieved with far fewer observations. For a fixed reconstruction quality, our speckle scanning approach is faster than compressive sensing schemes which require changing the speckle pattern for each observation.
Stéphanie Guérit, Siddharth Sivankutty, John A. Lee 0001, Hervé Rigneault, Laurent Jacques
SIAM J. Imaging Sci.3
2022 Fast Multiscale Neighbor Embedding
abstract
Dimension reduction (DR) computes faithful low-dimensional (LD) representations of high-dimensional (HD) data. Outstanding performances are achieved by recent neighbor embedding (NE) algorithms such as t -SNE, which mitigate the curse of dimensionality. The single-scale or multiscale nature of NE schemes drives the HD neighborhood preservation in the LD space (LDS). While single-scale methods focus on single-sized neighborhoods through the concept of perplexity, multiscale ones preserve neighborhoods in a broader range of sizes and account for the global HD organization to define the LDS. For both single-scale and multiscale methods, however, their time complexity in the number of samples is unaffordable for big data sets. Single-scale methods can be accelerated by relying on the inherent sparsity of the HD similarities they involve. On the other hand, the dense structure of the multiscale HD similarities prevents developing fast multiscale schemes in a similar way. This article addresses this difficulty by designing randomized accelerations of the multiscale methods. To account for all levels of interactions, the HD data are first subsampled at different scales, enabling to identify small and relevant neighbor sets for each data point thanks to vantage-point trees. Afterward, these sets are employed with a Barnes-Hut algorithm to cheaply evaluate the considered cost function and its gradient, enabling large-scale use of multiscale NE schemes. Extensive experiments demonstrate that the proposed accelerations are, statistically significantly, both faster than the original multiscale methods by orders of magnitude, and better preserving the HD neighborhoods than state-of-the-art single-scale schemes, leading to high-quality LD embeddings. Public codes are freely available at https://github.com/cdebodt.
Cyril de Bodt, Dounia Mulders, Michel Verleysen, John A. Lee 0001
IEEE Trans. Neural Networks Learn. Syst.4
2021 Stochastic quartet approach for fast multidimensional scaling
abstract
Multidimensional scaling is a statistical process that aims to embed high-dimensional data into a lower-dimensional, more manageable space.Common MDS algorithms tend to have some limitations when facing large data sets due to their high time and spatial complexities.This paper attempts to tackle the problem by using a stochastic approach to MDS which uses gradient descent to optimise a loss function defined on randomly designated quartets of points.This method mitigates the quadratic memory usage by computing distances on the fly, and has iterations in O(N ) time complexity, with N samples.Experiments show that the proposed method provides competitive results in reasonable time.Public codes are available at https://github.com/PierreLambert3/SQuaD-MDS.git. Multidimensional scaling and its limitationsDimensionality reduction (DR) is the process of mapping high-dimensional (HD) observations into a lower-dimensional (LD) space such that the LD embedding is a faithful representation of the HD data.The main DR uses are in machine learning, to curb the curse of dimensionality, and in visualisation.Mapped data can reveal structures that would lay hidden from the human perception if left in HD.Typically, some information is lost by the DR and, therefore, each DR method has a take on what kind of information should be preserved and what can be lost.Used frequently in visualisation, t-SNE [1] aims at retaining the neighbourhood of each point according to a distance metric and a perplexity, which reflects the size of the neighbourhood to preserve.While t-SNE excels at retaining local structures, sufficiently remote points tend to be considered equally distant by the algorithm and, therefore, the larger-scale structures can be distorted.Such distortions can lead to erroneous conclusions by the human user, who might overestimate the dissimilarity between two clusters that are distant in the LD embedding.For this reason, using multiple DR paradigms in conjunction is a good practice in visualisation: another embedding that preserves distances instead of neighbourhoods would have prevented this erroneous conclusion.This paper considers metric multidimensional scaling (MDS): a DR technique that produces a LD embedding such that the pairwise distances in LD reflect those in HD.MDS minimises a cost function which, in its simplest form, is the sum of the squared differences between distances in HD and the Euclidean distances in LD.A common strategy to optimize this cost function is based on 417
Pierre Lambert, Cyril de Bodt, Michel Verleysen, John A. Lee 0001
ESANN4
2021 Impact of data subsamplings in Fast Multi-Scale Neighbor Embedding
abstract
Fast multi-scale neighbor embedding (f-ms-NE) is an algorithm that maps high-dimensional data to a low-dimensional space by preserving the multi-scale data neighborhoods.To lower its time complexity, f-ms-NE uses random subsamplings to estimate the data properties at multiple scales.To improve this estimation and study the f-ms-NE sensitivity to randomness, this paper generalizes the f-ms-NE cost function by averaging several subsamplings.Experiments reveal that this can slightly improve the quality of the embeddings while maintaining reasonable computation times.Codes are available at https://github.com/cdebodt/Fast_Multi-scale_NE.
Pierre Lambert, John A. Lee 0001, Michel Verleysen, Cyril de Bodt
ESANN2
2021 Estimating uncertainty in radiation oncology dose prediction with dropout and bootstrap in U-Net models
abstract
Deep learning models, such as U-Net, can be used to efficiently predict the optimal dose distribution in radiotherapy treatment planning.In this work, we want to supplement the prediction model with a measurement of its uncertainty at each voxel.For this purpose, a full Bayesian approach would, however, be too costly.Instead, we compare, based on their correlation with the actual error, three simpler methods, namely, the dropout, the bootstrap and a modification of the U-Net.These methods can be easily adapted to other architectures.200 patients with head and neck cancer were used in this work.
John A. Lee 0001, Alyssa Vanginderdeuren, Margerie Huet-Dastarac, Ana M. Barragan-Montero
ESANN1
2020 Perplexity-free Parametric t-SNE
Francesco Crecchi, Cyril de Bodt, Michel Verleysen, John A. Lee 0001, Davide Bacciu
ESANN4
2020 Deep Learning to Detect Bacterial Colonies for the Production of Vaccines
Paul Smyth, John A. Lee 0001, Gaël de Lannoy, Thomas Beznik
ESANN2
2019 Class-aware t-SNE: cat-SNE
Cyril de Bodt, Dounia Mulders, Daniel López Sánchez, Michel Verleysen, John A. Lee 0001
ESANN5
2019 Tensor factorization to extract patterns in multimodal EEG data
Dounia Mulders, Cyril de Bodt, Nicolas Lejeune, John A. Lee 0001, André Mouraux, Michel Verleysen
ESANN4
2019 Nonlinear Dimensionality Reduction With Missing Data Using Parametric Multiple Imputations
abstract
Dimensionality reduction (DR) aims at faithfully and meaningfully representing high-dimensional (HD) data into a low-dimensional (LD) space. Recently developed neighbor embedding DR methods lead to outstanding performances, thanks to their ability to foil the curse of dimensionality. Unfortunately, they cannot be directly employed on incomplete data sets, which become ubiquitous in machine learning. Discarding samples with missing features prevents their LD coordinates computation and deteriorates the complete samples treatment. Common missing data imputation schemes are not appropriate in the nonlinear DR context either. Indeed, even if they model the data distribution in the feature space, they can, at best, enable the application of a DR scheme on the expected data set. In practice, one would, instead, like to obtain the LD embedding with the closest cost function value on average with respect to the complete data case. As the state-of-the-art DR techniques are nonlinear, the latter embedding results from minimizing the expected cost function on the incomplete database, not from considering the expected data set. This paper addresses these limitations by developing a general methodology for nonlinear DR with missing data, being directly applicable with any DR scheme optimizing some criterion. In order to model the feature dependences, an HD extension of Gaussian mixture models is first fitted on the incomplete data set. It is afterward employed under the multiple imputation paradigms to obtain a single relevant LD embedding, thus minimizing the cost function expectation. Extensive experiments demonstrate the superiority of the suggested framework over alternative approaches.
Cyril de Bodt, Dounia Mulders, Michel Verleysen, John A. Lee 0001
IEEE Trans. Neural Networks Learn. Syst.4
2018 Multi-organ Segmentation of Chest CT Images in Radiation Oncology: Comparison of Standard and Dilated UNet
Umair Javaid, Damien Dasnoy, John A. Lee 0001
ACIVS3
2018 Contour Propagation in CT Scans with Convolutional Neural Networks
Jean Léger, Eliott Brion, Umair Javaid, John A. Lee 0001, Christophe De Vleeschouwer, Benoît Macq
ACIVS4
2018 Perplexity-free t-SNE and twice Student tt-SNE
Cyril de Bodt, Dounia Mulders, Michel Verleysen, John A. Lee 0001
ESANN4
2018 Extensive assessment of Barnes-Hut t-SNE
Cyril de Bodt, Dounia Mulders, Michel Verleysen, John A. Lee 0001
ESANN4
2018 Information visualisation and machine learning: latest trends towards convergence
Benoît Frénay, Bruno Dumas, John A. Lee 0001
ESANN3
2018 Capturing variabilities from Computed Tomography images with Generative Adversarial Networks (GANs)
Umair Javaid, John A. Lee 0001
ESANN2
2017 Large-scale nonlinear dimensionality reduction for network intrusion detection
Yasir Hamid, Ludovic Journaux, John A. Lee 0001, Lucile Sautot, Nabi Bushra, M. Sugumaran
ESANN3
2017 Comparing dynamics of fluency and inter-limb coordination in climbing activities using multi-scale Jensen-Shannon embedding and clustering
Romain Hérault, Dominic Orth, Ludovic Seifert, Jérémie Boulanger, John A. Lee 0001
Data Min. Knowl. Discov.5
2017 Kernel-based dimensionality reduction using Renyi's α-entropy measures of similarity
Andrés Marino Álvarez-Meza, John A. Lee 0001, Michel Verleysen, Germán Castellanos-Domínguez
Neurocomputing2
2017 What you see is what you can change: Human-centered machine learning by interactive visualization
Dominik Sacha, Michael Sedlmair, Leishi Zhang, John A. Lee 0001, Jaakko Peltonen, Daniel Weiskopf, Stephen C. North, Daniel A. Keim
Neurocomputing4
2017 Visual Interaction with Dimensionality Reduction: A Structured Literature Analysis
abstract
Dimensionality Reduction (DR) is a core building block in visualizing multidimensional data. For DR techniques to be useful in exploratory data analysis, they need to be adapted to human needs and domain-specific problems, ideally, interactively, and on-the-fly. Many visual analytics systems have already demonstrated the benefits of tightly integrating DR with interactive visualizations. Nevertheless, a general, structured understanding of this integration is missing. To address this, we systematically studied the visual analytics and visualization literature to investigate how analysts interact with automatic DR techniques. The results reveal seven common interaction scenarios that are amenable to interactive control such as specifying algorithmic constraints, selecting relevant features, or choosing among several DR algorithms. We investigate specific implementations of visual analysis systems integrating DR, and analyze ways that other machine learning methods have been combined with DR. Summarizing the results in a "human in the loop" process model provides a general lens for the evaluation of visual interactive DR systems. We apply the proposed model to study and classify several systems previously described in the literature, and to derive future research opportunities.
Dominik Sacha, Leishi Zhang, Michael Sedlmair, John A. Lee 0001, Jaakko Peltonen, Daniel Weiskopf, Stephen C. North, Daniel A. Keim
IEEE Trans. Vis. Comput. Graph.4
2016 Human-centered machine learning through interactive visualization: review and open challenges
Dominik Sacha, Michael Sedlmair, Leishi Zhang, John A. Lee 0001, Daniel Weiskopf, Stephen C. North, Daniel A. Keim
ESANN4
2016 Multi-step-ahead forecasting using kernel adaptive filtering
abstract
Accurate prediction of time series is of great interest because it can guide decisions in many economical or industrial fields. Methods can forecast either one or several steps ahead. The former is simpler and more common in many applications, while the latter is more challenging. The literature describes mainly three strategies of multi-step-ahead prediction: iterated (repeated one-step-ahead), direct, and MIMO (multiple input, multiple output). This paper proposes a MIMO strategy based on kernel adaptive filtering, which we named MSAKAF. The proposed MSAKAF has shown to be a more effective method in short, medium and long-term forecast. The proposed approach is validated on two real-world datasets. The results show that our proposal outperforms the compared baseline methods in terms of prediction accuracy.
Sergio García-Vega, Germán Castellanos-Domínguez, Michel Verleysen, John A. Lee 0001
IJCNN4
2015 Unsupervised dimensionality reduction: the challenge of big data visualization
Kerstin Bunte, John A. Lee 0001
ESANN2
2015 Geometrical homotopy for data visualization
Diego Hernán Peluffo-Ordóñez, Juan C. Alvarado-Pérez, John A. Lee 0001, Michel Verleysen
ESANN3
2015 Incremental classification of objects in scenes: Application to the delineation of images
Guillaume Bernard 0001, Michel Verleysen, John A. Lee 0001
Neurocomputing3
2015 Multi-scale similarities in stochastic neighbour embedding: Reducing dimensionality while preserving both local and global structure
John A. Lee 0001, Diego Hernán Peluffo-Ordóñez, Michel Verleysen
Neurocomputing1
2014 Two key properties of dimensionality reduction methods
abstract
Dimensionality reduction aims at providing faithful low-dimensional representations of high-dimensional data. Its general principle is to attempt to reproduce in a low-dimensional space the salient characteristics of data, such as proximities. A large variety of methods exist in the literature, ranging from principal component analysis to deep neural networks with a bottleneck layer. In this cornucopia, it is rather difficult to find out why a few methods clearly outperform others. This paper identifies two important properties that enable some recent methods like stochastic neighborhood embedding and its variants to produce improved visualizations of high-dimensional data. The first property is a low sensitivity to the phenomenon of distance concentration. The second one is plasticity, that is, the capability to forget about some data characteristics to better reproduce the other ones. In a manifold learning perspective, breaking some proximities typically allow for a better unfolding of data. Theoretical developments as well as experiments support our claim that both properties have a strong impact. In particular, we show that equipping classical methods with the missing properties significantly improves their results.
John A. Lee 0001, Michel Verleysen
CIDM1
2014 Generalized kernel framework for unsupervised spectral methods of dimensionality reduction
abstract
This work introduces a generalized kernel perspective for spectral dimensionality reduction approaches. Firstly, an elegant matrix view of kernel principal component analysis (PCA) is described. We show the relationship between kernel PCA, and conventional PCA using a parametric distance. Secondly, we introduce a weighted kernel PCA framework followed from least-squares support vector machines (LS-SVM). This approach starts with a latent variable that allows to write a relaxed LS-SVM problem. Such a problem is addressed by a primal-dual formulation. As a result, we provide kernel alternatives to spectral methods for dimensionality reduction such as multidimensional scaling, locally linear embedding, and laplacian eigenmaps; as well as a versatile framework to explain weighted PCA approaches. Experimentally, we prove that the incorporation of a SVM model improves the performance of kernel PCA.
Diego Hernán Peluffo-Ordóñez, John A. Lee 0001, Michel Verleysen
CIDM2
2014 Multiscale stochastic neighbor embedding: Towards parameter-free dimensionality reduction
John A. Lee 0001, Diego Hernán Peluffo-Ordóñez, Michel Verleysen
ESANN1
2014 Recent methods for dimensionality reduction: A brief comparative analysis
Diego Hernán Peluffo-Ordóñez, John A. Lee 0001, Michel Verleysen
ESANN2
2014 Unsupervised Relevance Analysis for Feature Extraction and Selection - A Distance-based Approach for Feature Relevance
Diego Hernán Peluffo-Ordóñez, John A. Lee 0001, Michel Verleysen, José L. Rodríguez, Germán Castellanos-Domínguez
ICPRAM2
2014 Advances in artificial neural networks, machine learning, and computational intelligence (ESANN 2013)
Mark J. Embrechts, Fabrice Rossi, Frank-Michael Schleif, John A. Lee 0001
Neurocomputing4
2013 Sensitivity to parameter and data variations in dimensionality reduction techniques
Francisco J. García-Fernández, Michel Verleysen, John A. Lee 0001, Ignacio Díaz Blanco
ESANN3
2013 Nonlinear Dimensionality Reduction for Visualization
Michel Verleysen, John A. Lee 0001
ICONIP (1)2
2013 Type 1 and 2 mixtures of Kullback-Leibler divergences as cost functions in dimensionality reduction based on similarity preservation
John A. Lee 0001, Emilie Renard, Guillaume Bernard 0001, Pierre Dupont, Michel Verleysen
Neurocomputing1
2012 Incremental feature building and classification for image segmentation
Guillaume Bernard 0001, Michel Verleysen, John A. Lee 0001
ESANN3
2012 Type 1 and 2 symmetric divergences for stochastic neighbor embedding
John A. Lee 0001
ESANN1
2012 Advances in artificial neural networks, machine learning, and computational intelligence (ESANN 2011)
John A. Lee 0001, Petra Schneider, John A. Quinn
Neurocomputing1
2012 Comparative Study With New Accuracy Metrics for Target Volume Contouring in PET Image Guided Radiation Therapy
abstract
The impact of PET on radiation therapy is held back by poor methods of defining functional volumes of interest. Many new software tools are being proposed for contouring target volumes but the different approaches are not adequately compared and their accuracy is poorly evaluated due to the illdefinition of ground truth. This paper compares the largest cohort to date of established, emerging and proposed PET contouring methods, in terms of accuracy and variability. We emphasise spatial accuracy and present a new metric that addresses the lack of unique ground truth. 30 methods are used at 13 different institutions to contour functional VOIs in clinical PET/CT and a custom-built PET phantom representing typical problems in image guided radiotherapy. Contouring methods are grouped according to algorithmic type, level of interactivity and how they exploit structural information in hybrid images. Experiments reveal benefits of high levels of user interaction, as well as simultaneous visualisation of CT images and PET gradients to guide interactive procedures. Method-wise evaluation identifies the danger of over-automation and the value of prior knowledge built into an algorithm.
Tony Shepherd, Mika Teräs, Reinhard Beichel, Ronald Boellaard, Michel Bruynooghe, Volker Dicken, Mark J. Gooding, Peter J. Julyan, John A. Lee 0001, Sébastien Lefèvre, Michael Mix, Valery Naranjo, Habib Zaidi, Heikki Minn
IEEE Trans. Medical Imaging9
2011 Mode estimation in high-dimensional spaces with flat-top kernels: Application to image denoising
Arnaud de Decker, Damien François, Michel Verleysen, John A. Lee 0001
Neurocomputing4
2011 Advances in artificial neural networks, machine learning, and computational intelligence
John A. Lee 0001, Frank-Michael Schleif, Thomas Martinetz
Neurocomputing1
2010 Mode estimation in high-dimensional spaces with flat-top kernels: application to image denoising
Arnaud de Decker, John A. Lee 0001, Damien François, Michel Verleysen
ESANN2
2010 Recent Advances in Nonlinear Dimensionality Reduction, Manifold and Topological Learning
Axel Wismüller, Michel Verleysen, Michaël Aupetit 0001, John A. Lee 0001
ESANN4
2010 Unsupervised dimensionality reduction: Overview and recent advances
abstract
Unsupervised dimensionality reduction aims at representing high-dimensional data in lower-dimensional spaces in a faithful way. Dimensionality reduction can be used for compression or denoising purposes, but data visualization remains one its most prominent applications. This paper attempts to give a broad overview of the domain. Past developments are briefly introduced and pinned up on the time line of the last eleven decades. Next, the principles and techniques involved in the major methods are described. A taxonomy of the methods is suggested, taking into account various properties. Finally, the issue of quality assessment is briefly dealt with.
John A. Lee 0001, Michel Verleysen
IJCNN1
2010 Dimensionality reduction by rank preservation
abstract
Dimensionality reduction techniques aim at representing high-dimensional data in low-dimensional spaces. To be faithful and reliable, the representation is usually required to preserve proximity relationships. In practice, methods like multidimensional scaling try to fulfill this requirement by preserving pairwise distances in the low-dimensional representation. However, such a simplification does not easily allow for local scalings in the representation. It also makes these methods suboptimal with respect to recent quality criteria that are based on distance rankings. This paper addresses this issue by introducing a dimensionality reduction method that works with ranks. Appropriate hypotheses enable the minimization of a rank-based cost function. In particular, the scale indeterminacy that is inherent to ranks is circumvented by representing data on a space with a spherical topology.
Victor Onclinx, John A. Lee 0001, Vincent Wertz, Michel Verleysen
IJCNN2
2010 Advances in computational intelligence and learning (ESANN 2009)
Cecilio Angulo, John A. Lee 0001, Frank-Michael Schleif
Neurocomputing2
2010 A principled approach to image denoising with similarity kernels involving patches
Arnaud de Decker, John A. Lee 0001, Michel Verleysen
Neurocomputing2
2010 Scale-independent quality criteria for dimensionality reduction
John A. Lee 0001, Michel Verleysen
Pattern Recognit. Lett.1
2009 Patch-based bilateral filter and local m-smoother for image denoising
Arnaud de Decker, John A. Lee 0001, Michel Verleysen
ESANN2
2009 Adaptive anisotropic denoising: a bootstrapped procedure
John A. Lee 0001, Arnaud de Decker, Michel Verleysen
ESANN1
2009 Simbed: Similarity-Based Embedding
John A. Lee 0001, Michel Verleysen
ICANN (2)1
2009 Quality assessment of dimensionality reduction: Rank-based criteria
John A. Lee 0001, Michel Verleysen
Neurocomputing1
2008 Rank-based quality assessment of nonlinear dimensionality reduction
John A. Lee 0001, Michel Verleysen
ESANN1
2008 Blind source separation based on endpoint estimation with application to the MLSP 2006 data competition
John A. Lee 0001, Frédéric Vrins, Michel Verleysen
Neurocomputing1
2008 Edge-Preserving Filtering of Images with Low Photon Counts
abstract
Edge-preserving filters such as local M-smoothers or bilateral filtering are usually designed for Gaussian noise. This paper investigates how these filters can be adapted in order to efficiently deal with Poissonian noise. In addition, the issue of photometry invariance is addressed by changing the way filter coefficients are normalized. The proposed normalization is additive, instead of being multiplicative, and leads to a strong connection with anisotropic diffusion. Experiments show that ensuring the photometry invariance leads to comparable denoising performances in terms of the root mean square error computed on the signal.
John A. Lee 0001, Xavier Geets, Vincent Grégoire, Anne Bol
IEEE Trans. Pattern Anal. Mach. Intell.1
2007 Forecasting the CATS benchmark with the Double Vector Quantization method
Geoffroy Simon, John A. Lee 0001, Marie Cottrell, Michel Verleysen
Neurocomputing2
2007 A Minimum-Range Approach to Blind Extraction of Bounded Sources
abstract
In spite of the numerous approaches that have been derived for solving the independent component analysis (ICA) problem, it is still interesting to develop new methods when, among other reasons, specific a priori knowledge may help to further improve the separation performances. In this paper, the minimum-range approach to blind extraction of bounded source is investigated. The relationship with other existing well-known criteria is established. It is proved that the minimum-range approach is a contrast, and that the criterion is discriminant in the sense that it is free of spurious maxima. The practical issues are also discussed, and a range measure estimation is proposed based on the order statistics. An algorithm for contrast maximization over the group of special orthogonal matrices is proposed. Simulation results illustrate the performances of the algorithm when using the proposed range estimation criterion.
Frédéric Vrins, John A. Lee 0001, Michel Verleysen
IEEE Trans. Neural Networks2
2006 Non-orthogonal Support Width ICA
John A. Lee 0001, Frédéric Vrins, Michel Verleysen
ESANN1
2006 Unfolding preprocessing for meaningful time series clustering
Geoffroy Simon, John A. Lee 0001, Michel Verleysen
Neural Networks2
2005 Nonlinear dimensionality reduction of data manifolds with essential loops
John A. Lee 0001, Michel Verleysen
Neurocomputing1
2004 How to project 'circular' manifolds using geodesic distances?
John A. Lee 0001, Michel Verleysen
ESANN1
2004 Double quantization forecasting method for filling missing data in the CATS time series
abstract
The double vector quantization forecasting method based on Kohonen self-organizing maps is applied to predict the missing values of the CATS competition data set. As one of the features of the method is the ability to predict vectors instead of scalar values in a single step, the compromise between the size of the vector prediction and the number of repetitions needed to reach the required prediction horizon is studied. The long-term stability of the double vector quantization method makes it possible to obtain reliable values on a rather long-term forecasting horizon.
Geoffroy Simon, John A. Lee 0001, Michel Verleysen, Marie Cottrell
IJCNN2
2004 Nonlinear projection with curvilinear distances: Isomap versus curvilinear distance analysis
John A. Lee 0001, Amaury Lendasse, Michel Verleysen
Neurocomputing1
2003 On Convergence Problems of the EM Algorithm for Finite Gaussian Mixtures
Cédric Archambeau, John A. Lee 0001, Michel Verleysen
ESANN2
2003 Locally Linear Embedding versus Isotop
John A. Lee 0001, Cédric Archambeau, Michel Verleysen
ESANN1
2002 Width optimization of the Gaussian kernels in Radial Basis Function Networks
Nabil Benoudjit, Cédric Archambeau, Amaury Lendasse, John A. Lee 0001, Michel Verleysen
ESANN4
2002 Curvilinear Distance Analysis versus Isomap
John A. Lee 0001, Amaury Lendasse, Michel Verleysen
ESANN1
2002 Nonlinear Projection with the Isotop Method
John A. Lee 0001, Michel Verleysen
ICANN1
2002 Forecasting electricity consumption using nonlinear projection and self-organizing maps
Amaury Lendasse, John A. Lee 0001, Vincent Wertz, Michel Verleysen
Neurocomputing2
2002 Self-organizing maps with recursive neighborhood adaptation
John A. Lee 0001, Michel Verleysen
Neural Networks1
2001 Input data reduction for the prediction of financial time series
Amaury Lendasse, John A. Lee 0001, Eric de Bodt, Vincent Wertz, Michel Verleysen
ESANN2
2000 A robust non-linear projection method
John A. Lee 0001, Amaury Lendasse, Nicolas Donckers, Michel Verleysen
ESANN1
2000 Time series forecasting using CCA and Kohonen maps - application to electricity consumption
Amaury Lendasse, John A. Lee 0001, Vincent Wertz, Michel Verleysen
ESANN2