EDBT 2026 Demo / reviewers in the wild / expert
Gianni Franchi
dblp:155/3061
· DBLP profile ↗
33ranked-venue papers
14as first author
25since 2021 · last 2026
0000-0002-2184-1381ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 9 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 12 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Benchmarking XAI Explanations with Human-Aligned EvaluationsabstractWe introduce PASTA (Perceptual Assessment System for explanaTion of Artificial Intelligence), a novel human-centric framework for evaluating eXplainable AI (XAI) techniques in computer vision. Our first contribution is the creation of the PASTA-dataset, the first large-scale benchmark that spans a diverse set of models and both saliency-based and concept-based explanation methods. This dataset enables robust, comparative analysis of XAI techniques based on human judgment. Our second contribution is an automated, data-driven benchmark that predicts human preferences using the PASTA-dataset. This scoring called PASTA-score method offers scalable, reliable, and consistent evaluation aligned with human perception. Additionally, our benchmark allows for comparisons between explanations across different modalities, an aspect previously unaddressed. We then propose to apply our scoring method to probe the interpretability of existing models and to build more human interpretable XAI methods. Rémi Kazmierczak, Steve Azzolin, Eloïse Berthier, Anna Hedström, Patricia Delhomme, David Filliat, Nicolas Bousquet 0001, Goran Frehse, Massimiliano Mancini, Baptiste Caramiaux, Andrea Passerini, Gianni Franchi |
AAAI | 12 |
| 2025 | Towards Understanding and Quantifying Uncertainty for Text-to-Image GenerationabstractUncertainty quantification in text-to-image (T2I) generative models is crucial for understanding model behavior and improving output reliability. In this paper, we are the first to quantify and evaluate the uncertainty of T2I models with respect to the prompt. Alongside adapting existing approaches designed to measure uncertainty in the image space, we also introduce Prompt-based UNCertainty Estimation for T2I models (PUNC), a novel method leveraging Large Vision-Language Models (LVLMs) to better address uncertainties arising from the semantics of the prompt and generated images. PUNC utilizes a LVLM to caption a generated image, and then compares the caption with the original prompt in the more semantically meaningful text space. PUNC also enables the disentanglement of both aleatoric and epistemic uncertainties via precision and recall, which image-space approaches are unable to do. Extensive experiments demonstrate that PUNC outperforms state-of-the-art uncertainty estimation techniques across various settings. Uncertainty quantification in text-to-image generation models can be used on various applications including bias detection, copyright protection, and OOD detection. We also introduce a comprehensive dataset of text prompts and generation pairs to foster further research in uncertainty quantification for generative models. Our findings illustrate that PUNC not only achieves competitive performance but also enables novel applications in evaluating and improving the trustworthiness of text-to-image models. The code is available at https://github.com/ENSTA-U2IS-AI/Uncertainty_diffusion Gianni Franchi, Nacim Belkhir, Dat Nguyen Trong, Guoxuan Xia, Andrea Pilzer |
CVPR | 1 |
| 2025 | Collaborative Instance Object Navigation: Leveraging Uncertainty-Awareness to Minimize Human-Agent Dialogues
Francesco Taioli, Edoardo Zorzi, Gianni Franchi, Alberto Castellini, Alessandro Farinelli, Marco Cristani, Yiming Wang 0002 |
ICCV | 3 |
| 2025 | Towards Understanding Why Label Smoothing Degrades Selective Classification and How to Fix ItabstractLabel smoothing (LS) is a popular regularisation method for training neural networks as it is effective in improving test accuracy and is simple to implement. ''Hard'' one-hot labels are ''smoothed'' by uniformly distributing probability mass to other classes, reducing overfitting. Prior work has shown that in some cases *LS can degrade selective classification (SC)* -- where the aim is to reject misclassifications using a model's uncertainty. In this work, we first demonstrate empirically across an extended range of large-scale tasks and architectures that LS *consistently* degrades SC.
We then address a gap in existing knowledge, providing an *explanation* for this behaviour by analysing logit-level gradients: LS degrades the uncertainty rank ordering of correct vs incorrect predictions by regularising the max logit *more* when a prediction is likely to be correct, and *less* when it is likely to be wrong.
This elucidates previously reported experimental results where strong classifiers underperform in SC.
We then demonstrate the empirical effectiveness of post-hoc *logit normalisation* for recovering lost SC performance caused by LS. Furthermore, linking back to our gradient analysis, we again provide an explanation for why such normalisation is effective. Guoxuan Xia, Olivier Laurent 0002, Gianni Franchi, Christos-Savvas Bouganis |
ICLR | 3 |
| 2025 | Double Descent Meets Out-of-Distribution Detection: Theoretical Insights and Empirical Analysis on the Role of Model Complexityabstract**Out-of-distribution (OOD) detection** is essential for ensuring the reliability and safety of machine learning systems. In recent years, it has received increasing attention, particularly through post-hoc detection and training-based methods. In this paper, we focus on **post-hoc OOD detection**, which enables identifying OOD samples without altering the model's training procedure or objective. Our primary goal is to investigate the relationship between **model capacity** and its OOD detection performance. Specifically, we aim to answer the following question:
*Does the Double Descent phenomenon manifest in post-hoc OOD detection?* This question is crucial, as it can reveal whether overparameterization, which is already known to benefit generalization, can also enhance OOD detection.
Despite the growing interest in these topics by the classic supervised machine learning community, this intersection remains unexplored for OOD detection.
We empirically demonstrate that the Double Descent effect does indeed appear in post-hoc OOD detection. Furthermore, we provide theoretical insights to explain why this phenomenon emerges in such setting. Finally, we show that the overparameterized regime does not yield superior results consistently, and we propose a method to identify the optimal regime for OOD detection based on our observations. Mouïn Ben Ammar, David Brellmann, Arturo Mendoza, Antoine Manzanera, Gianni Franchi |
NeurIPS | 5 |
| 2025 | Torch-Uncertainty: Deep Learning Uncertainty QuantificationabstractDeep Neural Networks (DNNs) have demonstrated remarkable performance across various domains, including computer vision and natural language processing. However, they often struggle to accurately quantify their predictions' uncertainty, limiting their broader adoption in critical industrial applications. Uncertainty Quantification (UQ) for Deep Learning seeks to address this challenge by providing methodologies to improve the reliability of uncertainty estimates. While numerous techniques have been proposed, a unified tool remains lacking that offers a seamless workflow for evaluating and integrating these methods. To bridge this gap, we introduce Torch-Uncertainty, a PyTorch and Lightning framework designed to streamline the training and evaluation of DNNs with UQ techniques. In this paper, we outline the foundational principles of our library and present comprehensive experimental results that benchmark a diverse set of UQ methods across classification, segmentation, and regression tasks. Our library is available at: https://github.com/ENSTA-U2IS-AI/torch-uncertainty. Adrien Lafage, Olivier Laurent 0002, Firas Gabetni, Gianni Franchi |
NeurIPS | 4 |
| 2025 | Hierarchical Light Transformer Ensembles for Multimodal Trajectory ForecastingabstractAccurate trajectory forecasting is crucial for the performance of various systems, such as advanced driver-assistance systems and self-driving vehicles. These forecasts allow us to anticipate events that lead to collisions and, therefore, to mitigate them. Deep Neural Networks have excelled in motion forecasting, but overconfidence and weak uncertainty quantification persist. Deep Ensembles address these concerns, yet applying them to multimodal distributions remains challenging. In this paper, we propose a novel approach named Hierarchical Light Transformer Ensembles (HLT-Ens) aimed at efficiently training an ensemble of Transformer architectures using a novel hierarchical loss function. HLT-Ens leverages grouped fully connected layers, inspired by grouped convolution techniques, to capture multimodal distributions effectively. We demonstrate that HLT-Ens achieves state-of-the-art performance levels through extensive experimentation, offering a promising avenue for improving trajectory forecasting techniques. We make our code available at github.com/alafage/hlt-ens. Adrien Lafage, Mathieu Barbier, Gianni Franchi, David Filliat |
WACV | 3 |
| 2025 | Robust trajectory forecasting in autonomous systems using mixtures of Student's T-distributions with T-DistNet
Adrien Lafage, Gianni Franchi, Mathieu Barbier, David Filliat |
Pattern Recognit. | 2 |
| 2024 | Discretization-Induced Dirichlet Posterior for Robust Uncertainty Quantification on RegressionabstractUncertainty quantification is critical for deploying deep neural networks (DNNs) in real-world applications. An Auxiliary Uncertainty Estimator (AuxUE) is one of the most effective means to estimate the uncertainty of the main task prediction without modifying the main task model. To be considered robust, an AuxUE must be capable of maintaining its performance and triggering higher uncertainties while encountering Out-of-Distribution (OOD) inputs, i.e., to provide robust aleatoric and epistemic uncertainty. However, for vision regression tasks, current AuxUE designs are mainly adopted for aleatoric uncertainty estimates, and AuxUE robustness has not been explored. In this work, we propose a generalized AuxUE scheme for more robust uncertainty quantification on regression tasks. Concretely, to achieve a more robust aleatoric uncertainty estimation, different distribution assumptions are considered for heteroscedastic noise, and Laplace distribution is finally chosen to approximate the prediction error. For epistemic uncertainty, we propose a novel solution named Discretization-Induced Dirichlet pOsterior (DIDO), which models the Dirichlet posterior on the discretized prediction error. Extensive experiments on age estimation, monocular depth estimation, and super-resolution tasks show that our proposed method can provide robust uncertainty estimates in the face of noisy inputs and that it can be scalable to both image-level and pixel-wise tasks. Xuanlong Yu, Gianni Franchi, Jindong Gu, Emanuel Aldea |
AAAI | 2 |
| 2024 | Make Me a BNN: A Simple Strategy for Estimating Bayesian Uncertainty from Pre-trained ModelsabstractDeep Neural Networks (DNNs) are powerful tools for various computer vision tasks, yet they often struggle with reliable uncertainty quantification — a critical requirement for real-world applications. Bayesian Neural Networks (BNN) are equipped for uncertainty estimation but cannot scale to large DNNs where they are highly unstable to train. To address this challenge, we introduce the Adaptable Bayesian Neural Network (ABNN), a simple and scalable strategy to seamlessly transform DNNs into BNNs in a post-hoc manner with minimal computational and training overheads. ABNN preserves the main predictive properties of DNNs while enhancing their uncertainty quantification abilities through simple BNN adaptation layers (attached to normalization layers) and a few fine-tuning steps on pretrained models. We conduct extensive experiments across multiple datasets for image classification and semantic segmentation tasks, and our results demonstrate that ABNN achieves state-of-the-art performance without the computational budget typically associated with ensemble methods. Gianni Franchi, Olivier Laurent 0002, Maxence Leguéry, Andrei Bursuc, Andrea Pilzer, Angela Yao |
CVPR | 1 |
| 2024 | A Symmetry-Aware Exploration of Bayesian Neural Network PosteriorsabstractThe distribution of modern deep neural networks (DNNs) weights -- crucial for uncertainty quantification and robustness -- is an eminently complex object due to its extremely high dimensionality. This paper presents one of the first large-scale explorations of the posterior distribution of deep Bayesian Neural Networks (BNNs), expanding its study to real-world vision tasks and architectures. Specifically, we investigate the optimal approach for approximating the posterior, analyze the connection between posterior quality and uncertainty quantification, delve into the impact of modes on the posterior, and explore methods for visualizing the posterior. Moreover, we uncover weight-space symmetries as a critical aspect for understanding the posterior. To this extent, we develop an in-depth assessment of the impact of both permutation and scaling symmetries that tend to obfuscate the Bayesian posterior. While the first type of transformation is known for duplicating modes, we explore the relationship between the latter and L2 regularization, challenging previous misconceptions. Finally, to help the community improve our understanding of the Bayesian posterior, we release the first large-scale checkpoint dataset, including thousands of real-world models, along with our code. Olivier Laurent 0002, Emanuel Aldea, Gianni Franchi |
ICLR | 3 |
| 2024 | NECO: NEural Collapse Based Out-of-distribution detectionabstractDetecting out-of-distribution (OOD) data is a critical challenge in machine learning due to model overconfidence, often without awareness of their epistemological limits. We hypothesize that "neural collapse", a phenomenon affecting in-distribution data for models trained beyond loss convergence, also influences OOD data. To benefit from this interplay, we introduce NECO, a novel post-hoc method for OOD detection, which leverages the geometric properties of “neural collapse” and of principal component spaces to identify OOD data. Our extensive experiments demonstrate that NECO achieves state-of-the-art results on both small and large-scale OOD detection tasks while exhibiting strong generalization capabilities across different network architectures. Furthermore, we provide a theoretical explanation for the effectiveness of our method in OOD detection. We plan to release the code after the anonymity period. Mouïn Ben Ammar, Nacim Belkhir, Sebastian Popescu, Antoine Manzanera, Gianni Franchi |
ICLR | 5 |
| 2024 | Scaling for Training Time and Post-hoc Out-of-distribution Detection EnhancementabstractActivation shaping has proven highly effective for identifying out-of-distribution (OOD) samples post-hoc. Activation shaping prunes and scales network activations before estimating the OOD energy score; such an extremely simple approach achieves state-of-the-art OOD detection with minimal in-distribution (ID) accuracy drops. This paper analyzes the working mechanism behind activation shaping. We directly show that the benefits for OOD detection derive only from scaling, while pruning is detrimental. Based on our analysis, we propose SCALE, an even simpler yet more effective post-hoc network enhancement method for OOD detection. SCALE attains state-of-the-art OOD detection performance without any compromises on ID accuracy. Furthermore, we integrate scaling concepts into learning and propose Intermediate Tensor SHaping (ISH) for training-time OOD detection enhancement. ISH achieves significant AUROC improvements for both near- and far-OOD, highlighting the importance of activation distributions in emphasizing ID data characteristics. Our code and models are available at https://github.com/kai422/SCALE. Rongyu Chen, Gianni Franchi, Angela Yao |
ICLR | 3 |
| 2024 | Frustratingly Easy Test-Time Adaptation of Vision-Language ModelsabstractVision-Language Models seamlessly discriminate among arbitrary semantic categories, yet they still suffer from poor generalization when presented with challenging examples. For this reason, Episodic Test-Time Adaptation (TTA) strategies have recently emerged as powerful techniques to adapt VLMs in the presence of a single unlabeled image. The recent literature on TTA is dominated by the paradigm of prompt tuning by Marginal Entropy Minimization, which, relying on online backpropagation, inevitably slows down inference while increasing memory. In this work, we theoretically investigate the properties of this approach and unveil that a surprisingly strong TTA method lies dormant and hidden within it. We term this approach ZERO (TTA with “zero” temperature), whose design is both incredibly effective and frustratingly simple: augment N times, predict, retain the most confident predictions, and marginalize after setting the Softmax temperature to zero. Remarkably, ZERO requires a single batched forward pass through the vision encoder only and no backward passes. We thoroughly evaluate our approach following the experimental protocol established in the literature and show that ZERO largely surpasses or compares favorably w.r.t. the state-of-the-art while being almost 10× faster and 13× more memory friendly than standard Test-Time Prompt Tuning. Thanks to its simplicity and comparatively negligible computation, ZERO can serve as a strong baseline for future work in this field. Code will be available. Matteo Farina, Gianni Franchi, Giovanni Iacca, Massimiliano Mancini, Elisa Ricci 0001 |
NeurIPS | 2 |
| 2024 | InfraParis: A multi-modal and multi-task autonomous driving datasetabstractCurrent deep neural networks (DNNs) for autonomous driving computer vision are typically trained on specific datasets that only involve a single type of data and urban scenes. Consequently, these models struggle to handle new objects, noise, nighttime conditions, and diverse scenarios, which is essential for safety-critical applications. Despite ongoing efforts to enhance the resilience of computer vision DNNs, progress has been sluggish, partly due to the absence of benchmarks featuring multiple modalities. We introduce a novel and versatile dataset named InfraParis that supports multiple tasks across three modalities: RGB, depth, and infrared. We assess various state-of-the-art baseline techniques, encompassing models for the tasks of semantic segmentation, object detection, and depth estimation. More visualizations and the download link for InfraParis are available at https://enstau2is.github.io/infraParis/. Gianni Franchi, Marwane Hariat, Xuanlong Yu, Nacim Belkhir, Antoine Manzanera, David Filliat |
WACV | 1 |
| 2024 | Learning to generate training datasets for robust semantic segmentationabstractSemantic segmentation methods have advanced significantly. Still, their robustness to real-world perturbations and object types not seen during training remains a challenge, particularly in safety-critical applications. We propose a novel approach to improve the robustness of semantic segmentation techniques by leveraging the synergy between label-to-image generators and image-to-label segmentation models. Specifically, we design Robusta, a novel robust conditional generative adversarial network to generate realistic and plausible perturbed images that can be used to train reliable segmentation models. We conduct in-depth studies of the proposed generative model, assess the performance and robustness of the downstream segmentation network, and demonstrate that our approach can significantly enhance the robustness in the face of real-world perturbations, distribution shifts, and out-of-distribution samples. Our results suggest that this approach could be valuable in safety-critical applications, where the reliability of perception modules such as semantic segmentation is of utmost importance and comes with a limited computational budget in inference. We release our code at github.com/ENSTA-U2IS/robusta. Marwane Hariat, Olivier Laurent 0002, Rémi Kazmierczak, Andrei Bursuc, Angela Yao, Gianni Franchi |
WACV | 7 |
| 2024 | Encoding the Latent Posterior of Bayesian Neural Networks for Uncertainty QuantificationabstractBayesian Neural Networks (BNNs) have long been considered an ideal, yet unscalable solution for improving the robustness and the predictive uncertainty of deep neural networks. While they could capture more accurately the posterior distribution of the network parameters, most BNN approaches are either limited to small networks or rely on constraining assumptions, e.g., parameter independence. These drawbacks have enabled prominence of simple, but computationally heavy approaches such as Deep Ensembles, whose training and testing costs increase linearly with the number of networks. In this work we aim for efficient deep BNNs amenable to complex computer vision architectures, e.g., ResNet-50 DeepLabv3+, and tasks, e.g., semantic segmentation and image classification, with fewer assumptions on the parameters. We achieve this by leveraging variational autoencoders (VAEs) to learn the interaction and the latent distribution of the parameters at each network layer. Our approach, called Latent-Posterior BNN (LP-BNN), is compatible with the recent BatchEnsemble method, leading to highly efficient (in terms of computation and memory during both training and testing) ensembles. LP-BNNs attain competitive results across multiple metrics in several challenging benchmarks for image classification, semantic segmentation, and out-of-distribution detection. Gianni Franchi, Andrei Bursuc, Emanuel Aldea, Séverine Dubuisson, Isabelle Bloch |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Packed Ensembles for efficient uncertainty estimation
Olivier Laurent 0002, Adrien Lafage, Enzo Tartaglione, Geoffrey Daniel, Jean-Marc Martinez, Andrei Bursuc, Gianni Franchi |
ICLR | 7 |
| 2022 | MUAD: Multiple Uncertainties for Autonomous Driving, a benchmark for multiple uncertainty types and tasks
Gianni Franchi, Xuanlong Yu, Andrei Bursuc, Ángel Tena, Rémi Kazmierczak, Séverine Dubuisson, Emanuel Aldea, David Filliat |
BMVC | 1 |
| 2022 | Latent Discriminant Deterministic Uncertainty
Gianni Franchi, Xuanlong Yu, Andrei Bursuc, Emanuel Aldea, Séverine Dubuisson, David Filliat |
ECCV (12) | 1 |
| 2022 | On Monocular Depth Estimation and Uncertainty Quantification Using Classification Approaches for RegressionabstractMonocular depth is important in many tasks, such as 3D reconstruction and autonomous driving. Deep learning based models achieve state-of-the-art performance in this field. A set of novel approaches for estimating monocular depth consists of transforming the regression task into a classification one. However, there is a lack of detailed descriptions and comparisons for Classification Approaches for Regression (CAR) in the community and no in-depth exploration of their potential for uncertainty estimation. To this end, this paper will introduce a taxonomy and summary of CAR approaches, a new uncertainty estimation solution for CAR, and a set of experiments on depth accuracy and uncertainty quantification for CAR-based models on KITTI dataset. The experiments reflect the differences in the portability of various CAR methods on two backbones. Meanwhile, the newly proposed method for uncertainty estimation can outperform the ensembling method with only one forward propagation. Xuanlong Yu, Gianni Franchi, Emanuel Aldea |
ICIP | 2 |
| 2022 | Greybox XAI: A Neural-Symbolic learning framework to produce interpretable predictions for image classification
Adrien Bennetot, Gianni Franchi, Javier Del Ser, Raja Chatila 0001, Natalia Díaz Rodríguez |
Knowl. Based Syst. | 2 |
| 2022 | Learning deep morphological networks with neural architecture search
Yufei Hu, Nacim Belkhir, Jesús Angulo, Angela Yao, Gianni Franchi |
Pattern Recognit. | 5 |
| 2021 | Robust Semantic Segmentation with Superpixel-Mix
Gianni Franchi, Nacim Belkhir, Mai Lan Ha, Yufei Hu, Andrei Bursuc, Volker Blanz, Angela Yao |
BMVC | 1 |
| 2021 | SLURP: Side Learning Uncertainty for Regression Problems
Xuanlong Yu, Gianni Franchi, Emanuel Aldea |
BMVC | 2 |
| 2020 | TRADI: Tracking Deep Neural Network Weight Distributions
Gianni Franchi, Andrei Bursuc, Emanuel Aldea, Séverine Dubuisson, Isabelle Bloch |
ECCV (17) | 1 |
| 2020 | Tracking Hundreds of People in Densely Crowded Scenes With Particle Filtering Supervising Deep Convolutional Neural NetworksabstractTracking an entire high-density crowd composed of more than five hundred individuals is a difficult task that has not yet been accomplished. In this article, we propose to track pedestrians using a model composed of a Particle Filter (PF) and three Deep Convolutional Neural Networks (DCNN). The first network is a detector that learns to localize the persons. The second one is a pretrained network that estimates the optical flow, and the last one corrects the flow. Our contribution resides in the way we train this last network by PF supervision, and in Markov Random Field linking the different tracks. Gianni Franchi, Emanuel Aldea, Séverine Dubuisson, Isabelle Bloch |
ICIP | 1 |
| 2020 | Deep morphological networks
Gianni Franchi, Amin Fehri, Angela Yao |
Pattern Recognit. | 1 |
| 2019 | Crowd Behavior Characterization for Scene TrackingabstractIn this work, we perform an in-depth analysis of the specific difficulties a crowded scene dataset raises for tracking algorithms. Starting from the standard characteristics depicting the crowd and their limitations, we introduce six entropy measures related to the motion patterns and to the appearance variability of the individuals forming the crowd, and one appearance measure based on Principal Component Analysis. The proposed measures are discussed on synthetic configurations and on multiple real datasets. These criteria are able to characterize the crowd behavior at a more detailed level and may be helpful for evaluating the tracking difficulty of different datasets. The results are in agreement with the perceived difficulty of the scenes. Gianni Franchi, Emanuel Aldea, Séverine Dubuisson, Isabelle Bloch |
AVSS | 1 |
| 2018 | Segmentation and Shape Extraction from Convolutional Neural NetworksabstractWe propose a novel method for creating high-resolution class activation maps from a given deep convolutional neural network which was trained for image classification. The resulting class activation maps not only provide information about the localization of the main objects and their instances in the image, but are also accurate enough to predict their shapes. Rather than pursuing a weakly supervised learning strategy, the proposed algorithm is a multiscale extension of the classical class activation maps using a principal component analysis of the classification network feature maps, guided filtering, and a conditional random field. Nevertheless, the resulting shape information is competitive with state-of-the-art weakly supervised segmentation methods on datasets on which the latter have been trained, while being significantly better at generalizing to other datasets and unknown classes. Mai Lan Ha, Gianni Franchi, Michael Möller 0001, Andreas Kolb 0001, Volker Blanz |
WACV | 2 |
| 2016 | A deep spatial/spectral descriptor of hyperspectral texture using scattering transformabstractA technique to describe the spatial / spectral features of hyperspectral images is introduced. These descriptors aim at representing the content of the image while considering invariances related to the texture and to its geometric transformations, so called spatial invariances. Moreover, we also consider spectral invariances which are related to the composition of the pixels. Our approach is based on the scattering transform, which provides an useful framework for deep learning classification. The goal through these descriptors is to improve pixel-wise classification of hyperspectral images. Gianni Franchi, Jesús Angulo |
ICIP | 1 |
| 2016 | Hyperspectral image classification with support vector machines on kernel distribution embeddingsabstractWe propose a novel approach for pixel classification in hyperspectral images, leveraging on both the spatial and spectral information in the data. The introduced method relies on a recently proposed framework for learning on distributions - by representing them with mean elements in reproducing kernel Hilbert spaces (RKHS) and formulating a classification algorithm therein. In particular, we associate each pixel to an empirical distribution of its neighbouring pixels, a judicious representation of which in an RKHS, in conjunction with the spectral information contained in the pixel itself, give a new explicit set of features that can be fed into a suite of standard classification techniques - we opt for a well established framework of support vector machines (SVM). Furthermore, the computational complexity is reduced via random Fourier features formalism. We study the consistency and the convergence rates of the proposed method and the experiments demonstrate strong performance on hyperspectral data with gains in comparison to the state-of-the-art results. Gianni Franchi, Jesús Angulo, Dino Sejdinovic |
ICIP | 1 |
| 2014 | Spatially-Variant Area Openings for Reference-Driven Adaptive Contour Preserving FilteringabstractClassical adaptive mathematical morphology is based on operators which locally adapt the structuring elements to the image properties. Connected morphological operators act on the level of the flat zones of an image, such that only flat zones are filtered out, and hence the object edges are preserved. Area opening (resp. area closing) is one of the most useful connected operators, which filters out the bright (resp. dark) regions. It intrinsically involves the adaptation of the shape of the structuring element parameterized by its area. In this paper, we introduce the notion of reference-driven adaptive area opening according to two spatially-variant paradigms. First, the parameter of area is locally adapted by the reference image. This approach is applied to processing intensity depth images where the depth image is used to adapt the scale-size processing. Second, a self-dual area opening, where the reference image determines if the area filter is an opening or a closing with respect to the relationship between the image and the reference. Its natural application domain are the video sequences. Gianni Franchi, Jesús Angulo |
ICPR | 1 |