Minh-Tan Pham

dblp:153/9498 · DBLP profile ↗
← Back
39ranked-venue papers
13as first author
23since 2021 · last 2025
0000-0003-0266-767XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 30 · 13 first-author · 14 since 2021Artificial intelligence and machine learning · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021
YearPublicationVenuePosition
2025 RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling
abstract
International audience
Long-Khanh Pham, Thanh V. T. Tran, Minh-Tan Pham
INTERSPEECH3
2025 Multi-Prototype Hyperbolic Learning Guided by Class Hierarchy
abstract
Abstract In many computer vision applications, datasets often exhibit an underlying taxonomy within the label space. To adhere to this hierarchical structure, hyperbolic spaces have emerged as an effective manifold for representation learning, thanks to their ability to encode hierarchical relationships, with little distortion, even for low-dimensional embeddings. Hyperbolic prototypical learning, where class labels are represented by prototypes, has recently demonstrated strong potential in this setting. However, existing methods generally assume a uniform distribution of prototypes, overlooking the hierarchical organization of labels that may be available for a given task. To better exploit this prior knowledge, we propose a hierarchically informed method for prototype positioning. Our approach leverages the Gromov-Wasserstein distance to align the hierarchical relationships between labels with the initial uniform spherical distribution of prototypes, leading to more structured and semantically meaningful representations. Additionally, within a deep learning framework, we propose an alternative characterization of decision boundaries using horospheres, which are level sets of the Busemann function. Geometrically, horospheres correspond to spheres tangent to the boundary of hyperbolic space at a virtual point analogous to a prototype, which makes them a compliant tool in the prototypical learning context. Accordingly, we define a new horospherical layer that can be adapted to any neural network backbone. This layer is particularly advantageous when the number of prototypes exceeds the number of classes, offering enhanced flexibility to the classifier. Through our experiments, we demonstrate that the combination of proper initialization and optimized prototype positioning significantly enhances baseline performance for image classification on hierarchical datasets. Additionally, we validate our approach in two semantic segmentation tasks, using both image and point cloud datasets, confirming its effectiveness in leveraging hierarchical label structures for improved classification performance.
Paul Berg, Léo Buecher, Björn Michele, Minh-Tan Pham, Laetitia Chapel, Nicolas Courty
Int. J. Comput. Vis.4
2025 YOLO-G3CF: Gaussian Contrastive Cross-Channel Fusion for Multimodal Object Detection
abstract
Object detection is a crucial task in both computer vision and remote sensing. The performance of object detectors can vary across different modalities depending on lighting and weather conditions. To address these challenges, we propose a fusion module based on contrastive learning and Gaussian cross-channel attention, called Gaussian Contrastive Cross-Channel Fusion (G3CF). We integrate this module into a dual-YOLO architecture, forming YOLO-G3CF. The contrastive loss enforces similarity between the features sent to the detection head from both modality branches, as they should lead to the same detections. The Gaussian attention mechanism enables the model to fuse features in a higher-dimensional space, enhancing discriminative power. Extensive experiments on VEDAI, GeoImageNet, VTUAV-det, and FLIR demonstrate that G3CF improves detection performance, achieving a mAP increase of up to 6.64% over the best single-modality baselines and outperforming prior multimodal fusion methods. Regarding model complexity, our fusion method operates at a late stage, increasing the computational cost of single-modality YOLO by approximately 150% in terms of GFLOPs. For instance, YOLOv8 requires 52.84 GFLOPs, whereas YOLOv8-G3CF, due to its dual architecture and three G3CF modules, increases this to 131.22 GFLOPs. However, a single G3CF module requires only ~ 15 GFLOPs. Despite this overhead, our approach remains computationally less expensive than transformer-based models, e.g., ICAFusion requires 284.80 GFLOPs. Moreover, the proposed method still operates in real-time, achieving ~ 19 FPS on an NVIDIA RTX 2080. The code is available at https://github.com/abelmouhcine/YOLO-G3CF.
Abdelbadie Belmouhcine, Minh-Tan Pham, Sébastien Lefèvre
IEEE Geosci. Remote. Sens. Lett.2
2025 FedShip: Federated Learning for Ship Detection From Multi-Source Satellite Images
abstract
Detecting ships from satellite imagery is vital for maritime surveillance. Most current methods rely on deep learning (DL), which requires a large number of high-quality annotated images to train accurate models. Since satellite imagery comes from various sensors, DL-based ship detection algorithms need to perform well across different sensor types. However, privacy concerns, especially with commercial images, limit data, and annotation sharing. Federated learning (FL) offers a promising solution for collaborative learning while addressing these concerns. Despite its potential, research on FL for ship detection is still sparse. This study implements and evaluates three FL models for detecting ships using multi-source optical satellite images, spanning high to low resolution. Our experiments on two distinct datasets demonstrate that FL models significantly enhance detection performance without centralizing data. Source codes are publicly available athttps://github.com/ffyyytt/FLYOLO.
Anh-Kiet Duong, Tran Vu La, Hoàng-Ân Lê, Minh-Tan Pham
IEEE Geosci. Remote. Sens. Lett.4
2024 Horospherical Learning with Smart Prototypes
Paul Berg, Björn Michele, Minh-Tan Pham, Laetitia Chapel, Nicolas Courty
BMVC3
2024 Box for Mask and Mask for Box: weak losses for multi-task partially supervised learning
Hoàng-Ân Lê, Paul Berg, Minh-Tan Pham
BMVC3
2024 Leveraging Feature Communication in Federated Learning for Remote Sensing Image Classification
abstract
In the realm of Federated Learning (FL) applied to remote sensing image classification, this study introduces and assesses several innovative communication strategies. Our exploration includes feature-centric communication, pseudo-weight amalgamation, and a combined method utilizing both weights and features. Experiments conducted on two public scene classification datasets unveil the effectiveness of these strategies, showcasing accelerated convergence, heightened privacy, and reduced network information exchange. This research provides valuable insights into the implications of feature-centric communication in FL, offering potential applications tailored for remote sensing scenarios.
Anh-Kiet Duong, Hoàng-Ân Lê, Minh-Tan Pham
IGARSS3
2024 Insight into the Collocation of Multi-Source Satellite Imagery for Multi-Scale Vessel Detection
abstract
Ship detection from satellite imagery using Deep Learning (DL) is an indispensable solution for maritime surveillance. However, applying DL models trained on one dataset to others having differences in spatial resolution and radiometric features requires many adjustments. To overcome this issue, this paper focused on the DL models trained on datasets that consist of different optical images and a combination of radar and optical data. When dealing with a limited number of training images, the performance of DL models via this approach was satisfactory. They could improve 5–20% of average precision, depending on the optical images tested. Likewise, DL models trained on the combined optical and radar dataset could be applied to both optical and radar images. Our experiments showed that the models trained on an optical dataset could be used for radar images, while those trained on a radar dataset offered very poor scores when applied to optical images.
Tran Vu La, Minh-Tan Pham, Marco Chini
IGARSS2
2024 Leveraging Knowledge Distillation for Partial Multi-Task Learning from Multiple Remote Sensing Datasets
abstract
Partial multi-task learning where training examples are annotated for one of the target tasks is a promising idea in remote sensing as it allows combining datasets annotated for different tasks and predicting more tasks with fewer network parameters. The naïve approach to partial multi-task learning is sub-optimal due to the lack of all-task annotations for learning joint representations. This paper proposes using knowledge distillation to replace the need of ground truths for the alternate task and enhance the performance of such approach. Experiments conducted on the public ISPRS 2D Semantic Labeling Contest dataset show the effectiveness of the proposed idea on partial multi-task learning for semantic tasks including object detection and semantic segmentation in aerial images. Multi-task learning source codes are available at https://github.com/lhoangan/multas.
Hoàng-Ân Lê, Minh-Tan Pham
IGARSS2
2024 Plant Detection from Ultra High Resolution Remote Sensing Images: A Semantic Segmentation Approach Based on Fuzzy Loss
abstract
In this study, we tackle the challenge of identifying plant species from ultra high resolution (UHR) remote sensing images. Our approach involves introducing an RGB remote sensing dataset, characterized by millimeter-level spatial resolution, meticulously curated through several field expeditions across a mountainous region in France covering various landscapes. The task of plant species identification is framed as a semantic segmentation problem for its practical and efficient implementation across vast geographical areas. However, when dealing with segmentation masks, we confront instances where distinguishing boundaries between plant species and their background is challenging. We tackle this issue by introducing a fuzzy loss within the segmentation model. Instead of utilizing one-hot encoded ground truth (GT), our model incorporates Gaussian filter refined GT, introducing stochasticity during training. First experimental results obtained on both our UHR dataset and a public dataset are presented, showing the relevance of the proposed methodology, as well as the need for future improvement.
Shivam Pande, Baki Uzun, Florent Guiotte, Minh-Tan Pham, Thomas Corpetti, Florian Delerue, Sébastien Lefèvre
IGARSS4
2024 Mapping Earth Mounds from Space
abstract
Regular patterns of vegetation are considered widespread landscapes, although their global extent has never been estimated. Among them, spotted landscapes are of particular interest in the context of climate change. Indeed, regularly spaced vegetation spots in semi-arid shrublands result from extreme resource depletion and prefigure catastrophic shift of the ecosystem to a homogeneous desert, while termite mounds also producing spotted landscapes were shown to increase robustness to climate change. Yet, their identification at large scale calls for automatic methods, for instance using the popular deep learning framework, able to cope with a vast amount of remote sensing data, e.g., optical satellite imagery. In this paper, we tackle this problem and benchmark some state-of-the-art deep networks on several landscapes and geographical areas. Despite the promising results we obtained, we found that more research is needed to be able to map automatically these earth mounds from space.
Baki Uzun, Shivam Pande, Gwendal Cachin-Bernard, Minh-Tan Pham, Sébastien Lefèvre, Rumsais Blatrix, Doyle McKey
IGARSS4
2024 Multimodal Supervised Contrastive Learning in Remote Sensing Downstream Tasks
abstract
To leverage the large amount of unlabelled data available in remote sensing datasets, self-supervised learning (SSL) methods have recently emerged as an ubiquitous tool to pre-train robust image encoder models from unlabelled images. However, when used in a downstream setting, these models often need to be finetuned for a specific task after their pre-training. This finetuning still requires labelling information in order to train a classifier on top of the encoder while also updating the encoder weights. In this paper, we investigate the specific task of multimodal scene classification where a sample is composed of multiple views from multiple heterogeneous satellite sensors. We propose a method to improve the categorical cross-entropy finetuning process which is often used to specify the model for this downstream task. Our approach, based on the Supervised Contrastive Learning, uses the label information available to train an image encoder in a contrastive manner from multiple modalities while also training the task-specific classifier online. Such a multimodal supervised contrastive loss helps to better align representations from samples coming from multiple sensors but having the same class labels, thus improving the performance of the finetuning process. Our experiments on two public datasets including DFC2020 and Meter-ML with Sentinel-1/Sentinel-2 images show a significant gain over the baseline multimodal cross-entropy loss. All codes and datasets will be made publicly available for reproducibility upon acceptance.
Paul Berg, Baki Uzun, Minh-Tan Pham, Nicolas Courty
IEEE Geosci. Remote. Sens. Lett.3
2024 Hyperbolic prototypical network for few shot remote sensing scene classification
Manal Hamzaoui, Laetitia Chapel, Minh-Tan Pham, Sébastien Lefèvre
Pattern Recognit. Lett.3
2023 Data exploitation: multi-task learning of object detection and semantic segmentation on partially annotated data
Hoàng-Ân Lê, Minh-Tan Pham
BMVC2
2023 Unsupervised Anomaly Detection Using Variational Autoencoder with Gaussian Random Field Prior
abstract
We propose a new model of Variational Autoencoder (VAE) for Anomaly Detection (AD) with improved modeling power. More precisely, we introduce a VAE model with a Gaussian Random Field (GRF) prior, namely VAE-GRF, which generalizes the classical VAE model. We show that, under some assumptions, the VAE-GRF largely outperforms the traditional VAE and some other probabilistic models developed for AD. Our experimental results suggest that the VAE-GRF could be used as a relevant VAE baseline in place of the traditional VAE with very limited additional computational cost. We provide competitive results on the public MVTec benchmark dataset for visual inspection, as well as on the public Livestock dataset dedicated to the task of unsupervised animal detection from aerial images.
Hugo Gangloff, Minh-Tan Pham, Luc Courtrai, Sébastien Lefèvre
ICIP2
2023 Spherical Sliced-Wasserstein
Clément Bonet, Paul Berg, Nicolas Courty, François Septier, Lucas Drumetz, Minh-Tan Pham
ICLR6
2023 Multimodal Object Detection in Remote Sensing
abstract
Object detection in remote sensing is a crucial computer vision task that has seen significant advancements with deep learning techniques. However, most existing works in this area focus on the use of generic object detection and do not leverage the potential of multimodal data fusion. In this paper, we present a comparison of methods for multimodal object detection in remote sensing, survey available multimodal datasets suitable for evaluation, and discuss future directions.
Abdelbadie Belmouhcine, Jean-Christophe Burnel, Luc Courtrai, Minh-Tan Pham, Sébastien Lefèvre
IGARSS4
2023 Joint Multi-Modal Self-Supervised Pre-Training in Remote Sensing: Application to Methane Source Classification
abstract
With the current ubiquity of deep learning methods to solve computer vision and remote sensing specific tasks, the need for labelled data is growing constantly. However, in many cases, the annotation process can be long and tedious depending on the expertise needed to perform reliable annotations. In order to alleviate this need for annotations, several self-supervised methods have recently been proposed in the literature. The core principle behind these methods is to learn an image encoder using solely unlabelled data samples. In earth observation, there are opportunities to exploit domain-specific remote sensing image data in order to improve these methods. Specifically, by leveraging the geographical position associated with each image, it is possible to cross reference a location captured from multiple sensors, leading to multiple views of the same locations. In this paper, we briefly review the core principles behind so-called joint-embeddings methods and investigate the usage of multiple remote sensing modalities in self-supervised pre-training. We evaluate the final performance of the resulting encoders on the task of methane source classification.
Paul Berg, Minh-Tan Pham, Nicolas Courty
IGARSS2
2023 Hyperbolic Variational Auto-Encoder for Remote Sensing Scene Embeddings
abstract
The computer vision community is increasingly interested in exploring hyperbolic space for image representation, as hyperbolic approaches have demonstrated outstanding results in efficiently representing data with an underlying hierarchy. This interest arises from the intrinsic hierarchical nature among images. However, despite the hierarchical nature of remote sensing (RS) images, the investigation of hyperbolic spaces within the RS community has been relatively limited. The objective of this study is therefore to examine the relevance of hyperbolic embeddings of RS data, focusing on scene embedding. Using a Variational Auto-Encoder, we project the data into a hyperbolic latent space while ensuring numerical stability with a feature clipping technique. Experiments conducted on the NWPU-RESISC45 image dataset demonstrate the superiority of hyperbolic embeddings over the Euclidean counterparts in a classification task. Our study highlights the potential of operating in hyperbolic space as a promising approach for embedding RS data.
Manal Hamzaoui, Laetitia Chapel, Minh-Tan Pham, Sébastien Lefèvre
IGARSS3
2023 Knowledge Distillation for Object Detection: From Generic To Remote Sensing Datasets
abstract
Knowledge distillation, a well-known model compression technique, is an active research area in both computer vision and remote sensing communities. In this paper, we evaluate in a remote sensing context various off-the-shelf object detection knowledge distillation methods which have been originally developed on generic computer vision datasets such as Pascal VOC. In particular, methods covering both logit mimicking and feature imitation approaches are applied for vehicle detection using the well-known benchmarks such as xView and VEDAI datasets. Extensive experiments are performed to compare the relative performance and interrelationships of the methods. Experimental results show high variations and confirm the importance of result aggregation and cross validation on remote sensing datasets.
Hoàng-Ân Lê, Minh-Tan Pham
IGARSS2
2023 Weakly Supervised Marine Animal Detection from Remote Sensing Images Using Vector-Quantized Variational Autoencoder
abstract
This paper studies a reconstruction-based approach for weakly-supervised animal detection from aerial images in marine environments. Such an approach leverages an anomaly detection framework that computes metrics directly on the input space, enhancing interpretability and anomaly localization compared to feature embedding methods. Building upon the success of Vector-Quantized Variational Autoencoders in anomaly detection on computer vision datasets, we adapt them to the marine animal detection domain and address the challenge of handling noisy data. To evaluate our approach, we compare it with existing methods in the context of marine animal detection from aerial image data. Experiments conducted on two dedicated datasets demonstrate the superior performance of the proposed method over recent studies in the literature. Our framework offers improved interpretability and localization of anomalies, providing valuable insights for monitoring marine ecosystems and mitigating the impact of human activities on marine animals.
Minh-Tan Pham, Hugo Gangloff, Sébastien Lefèvre
IGARSS1
2023 Object Counting from Aerial Remote Sensing Images: Application to Wildlife and Marine Mammals
abstract
Anthropogenic activities pose threats to wildlife and marine fauna, prompting the need for efficient animal counting methods. This research study utilizes deep learning techniques to automate counting tasks. Inspired by previous studies on crowd and animal counting, a UNet model with various backbones is implemented, which uses Gaussian density maps for training, bypassing the need of training a detector. The new model is applied to the task of counting dolphins and elephants in aerial images. Quantitative evaluation shows promising results, with the EfficientNet-B5 backbone achieving the best performance for African elephants and the ResNet18 backbone for dolphins. The model accurately locates animals despite complex image background conditions. By leveraging artificial intelligence, this research contributes to wildlife conservation efforts and enhances coexistence between humans and wildlife through efficient object counting without detection from aerial remote sensing.
Tanya Singh, Hugo Gangloff, Minh-Tan Pham
IGARSS3
2022 Leveraging Vector-Quantized Variational Autoencoder Inner Metrics for Anomaly Detection
abstract
Anomaly Detection (AD) is an important research topic, with very diverse applications such as industrial defect detection, medical diagnosis, fraud detection, intrusion detection, etc. Within the last few years, deep learning-based methods have become the standard approach for AD. In many practical cases, the anomalies are unknown in advance. Therefore, most of challenging AD problems need to be addressed in an unsupervised or weakly supervised framework. In this context, deep generative models are widely used, in particular Variational Autoencoder (VAE) models. VAEs have been extended to Vector-Quantized VAEs (VQ-VAEs), a model increasingly popular because of its versatility enabled by the discrete latent space. We present for the first time a robust approach which takes advantage of the inner metrics of VQ-VAEs for AD. We show that the distance between the output of the encoder and the codebook vectors of a VQ-VAE provides a valuable information which can be used to localize the anomalies. In our approach, this metric complements a reconstruction-based metric to improve AD results. We compare our model with state-of-the-art AD models on three standards datasets, including the MVTec, UCSD-Ped1 and CIFAR-10 datasets. Experiments show that the proposed method yields high competitive results.
Hugo Gangloff, Minh-Tan Pham, Luc Courtrai, Sébastien Lefèvre
ICPR2
2020 Small Object Detection from Remote Sensing Images with the Help of Object-Focused Super-Resolution Using Wasserstein GANs
abstract
In this paper, we investigate and improve the use of a super-resolution approach to benefit the detection of small objects from aerial and satellite remote sensing images. The main idea is to focus the super-resolution on target objects within the training phase. Such a technique requires a reduced number of network layers depending on the desired scale factor and the reduced size of the target objects. The learning of our super-resolution network is performed using deep residual blocks integrated in a Wasserstein Generative adversarial network. Then, detection task is performed by exploiting two state-of-the-art detectors including Faster-RCNN and YOLOv3. Experiments were conducted on small vehicle detection from both aerial and satellite images from the VEDAI and xView data sets. Results showed that object-focused super-resolution improves the detection performance and facilitates the transfer learning from one data set to another.
Luc Courtrai, Minh-Tan Pham, Chloé Friguet, Sébastien Lefèvre
IGARSS2
2020 Vehicle Detection and Counting from VHR Satellite Images: Efforts and Open Issues
abstract
Detection of new infrastructures (commercial, logistics, industrial or residential) from satellite images constitutes a proven method to investigate and follow economic and urban growth. The level of activities or exploitation of these sites may be hardly determined by building inspection, but could be inferred from vehicle presence from nearby streets and parking lots. We present in this paper two deep learning-based models for vehicle counting from optical satellite images coming from the Pleiades sensor at 50-cm spatial resolution. Both segmentation (Tiramisu) and detection (YOLO, You Only Look Once) architectures were investigated. These networks were adapted, trained and validated on a data set including 87k vehicles, annotated using an interactive semi-automatic tool developed by the authors. Experimental results show that both segmentation and detection models could achieve a precision rate higher than 85 % with a recall rate also high (76.4 % and 71.9 % for Tiramisu and YOLO respectively).
Alice Froidevaux, Andréa Julier, Agustin Lifschitz, Minh-Tan Pham, Romain Dambreville, Sébastien Lefèvre, Pierre Lassalle, Thanh-Long Huynh
IGARSS4
2020 A Compound Polarimetric-Textural Approach for Unsupervised Change Detection in Multi-Temporal Full-Pol SAR Imagery
abstract
Change Detection represents a relevant topic for the analysis of multi-temporal analysis of Polarimetric SAR (PolSAR) data. However, most of the CD approaches for PolSAR imagery do not take into account textural information, which can be useful for have larger performance robustness. In this work, we propose a novel approach for unsupervised change detection considering polarimetric and textural information from multi-temporal PolSAR imagery. The approach is based on the joint use of features from coherency matrix and gradient tensor and the definition of a multi-temporal distance. A binary unsupervised thresholding is used for discriminating change and no-change classes. Experimental results obtained on a multi-temporal PolSAR dataset over Los Angeles area illustrate the effectiveness of the proposed approach.
Davide Pirrone, Minh-Tan Pham
IGARSS2
2020 Semantic Segmentation of LiDAR Points Clouds: Rasterization Beyond Digital Elevation Models
abstract
LiDAR point clouds are receiving a growing interest in remote sensing as they provide rich information to be used independently or together with optical data sources, such as aerial imagery. However, their nonstructured and sparse nature make them difficult to handle, conversely to raw imagery for which many efficient tools are available. To overcome this specific nature of LiDAR point clouds, the standard approach relies on converting the point cloud into a digital elevation model, represented as a 2-D raster. Such a raster can then be used similarly as optical images, e.g., with 2-D convolutional neural networks (CNNs) for semantic segmentation. In this letter, we show that LiDAR point clouds provide more information than only the digital elevation model and that considering alternative rasterization strategies helps to achieve better semantic segmentation results. We illustrate our findings on the IEEE Data Fusion Contest (DFC) 2018 data set.
Florent Guiotte, Minh-Tan Pham, Romain Dambreville, Thomas Corpetti, Sébastien Lefèvre
IEEE Geosci. Remote. Sens. Lett.2
2018 Classification of Remote Sensing Images Using Attribute Profiles and Feature Profiles from Different Trees: A Comparative Study
abstract
The motivation of this paper is to conduct a comparative study on remote sensing image classification using the morphological attribute profiles (APs) and feature profiles (FPs) generated from different types of tree structures. Over the past few years, APs have been among the most effective methods to model the image's spatial and contextual information. Recently, a novel extension of APs called FPs has been proposed by replacing pixel gray-levels with some statistical and geometrical features when forming the output profiles. FPs have been proved to be more efficient than the standard APs when generated from component trees (max-tree and min-tree). In this work, we investigate their performance on the inclusion tree (tree of shapes) and partition trees (alpha tree and omega tree). Experimental results from both panchromatic and hyperspectral images again confirm the efficiency of FPs compared to APs.
Minh-Tan Pham, Erchan Aptoula, Sébastien Lefèvre
IGARSS1
2018 Buried Object Detection from B-Scan Ground Penetrating Radar Data Using Faster-RCNN
abstract
In this paper, we adapt the Faster-RCNN framework for the detection of underground buried objects (i.e. hyperbola reflections) in B-scan ground penetrating radar (GPR) images. Due to the lack of real data for training, we propose to incorporate more simulated radargrams generated from different configurations using the gprMax toolbox. Our designed CNN is first pre-trained on the grayscale Cifar-10 database. Then, the Faster-RCNN framework based on the pre-trained CNN is trained and fine-tuned on both real and simulated GPR data. Preliminary detection results show that the proposed technique can provide significant improvements compared to classical computer vision methods and hence becomes quite promising to deal with this kind of specific GPR data even with few training samples.
Minh-Tan Pham, Sébastien Lefèvre
IGARSS1
2018 Attribute Profiles on Derived Textural Features for Highly Textured Optical Image Classification
abstract
Morphological attribute profiles (APs) have thus far been proven effective for remote sensing image classification by several research studies. However, recent studies have shown that a direct application of APs to highly textured and structured images, especially in very high-resolution (VHR) optical imagery, may be insufficient. Some solutions have been proposed to deal with this issue, such as to extract the local histograms and the local features of AP images [histogram-based APs (HAPs) and local feature-based APs (LFAP), respectively], or to combine APs with different textural features. In this letter, we review these approaches and then propose a novel strategy, which directly generates APs on some derived textural features instead of separately combining them. Experimental results from both natural textures and VHR optical remotely sensed images show that the proposed approach can produce superior classification performance than the standard APs, HAPs, LFAPs, as well as the classical combination of APs with textural features.
Minh-Tan Pham, Sébastien Lefèvre, François Merciol
IEEE Geosci. Remote. Sens. Lett.1
2018 Local Feature-Based Attribute Profiles for Optical Remote Sensing Image Classification
abstract
This paper introduces an extension of morphological attribute profiles (APs) by extracting their local features. The so-called local feature-based APs (LFAPs) are expected to provide a better characterization of each APs' filtered pixel (i.e., APs' sample) within its neighborhood, and hence better deal with local texture information from the image content. In this paper, LFAPs are constructed by extracting some simple first-order statistical features of the local patch around each APs' sample such as mean, standard deviation, and range. Then, the final feature vector characterizing each image pixel is formed by combining all local features extracted from APs of that pixel. In addition, since the self-dual APs (SDAPs) have been proved to outperform the APs in recent years, a similar process will be applied to form the local feature-based SDAPs (LFSDAPs). In order to evaluate the effectiveness of LFAPs and LFSDAPs, supervised classification using both the random forest and the support vector machine classifiers is performed on the very high resolution Reykjavik image as well as the hyperspectral Pavia University data. Experimental results show that LFAPs (respectively, LFSDAPs) can considerably improve the classification accuracy of the standard APs (respectively, SDAPs) and the recently proposed histogram-based APs.
Minh-Tan Pham, Sébastien Lefèvre, Erchan Aptoula
IEEE Trans. Geosci. Remote. Sens.1
2017 Classification of VHR remote sensing images using local feature-based attribute profiles
abstract
The present paper introduces an extension of attribute profiles (APs) by extracting their local features. The so-called local feature-based attribute profiles (LFAPs) are expected to provide a better characterization of each APs' filtered pixel (i.e. APs' sample) within its neighborhood, hence better deal with local texture information from the image's content. In this work, LFAP is constructed by extracting some simple first-order statistical features of the local patch around each APs' sample such as mean, standard deviation, range, etc. Then, the final feature vector characterizing each image pixel is formed by combining all local features extracted from APs of that pixel. In order to evaluate the effectiveness of the proposed technique, supervised classification using Random Forest classifier is performed on the VHR panchromatic Reykjavik image. Experimental results show that LFAPs can considerably improve the classification accuracy of the standard APs and the recently proposed histogram-based APs.
Minh-Tan Pham, Sébastien Lefèvre, Erchan Aptoula, Bharath Bhushan Damodaran
IGARSS1
2017 SAR image texture tracking using a pointwise graph-based model for glacier displacement measurement
abstract
This paper investigates the problem of glacier flow estimation using Synthetic Aperture Radar (SAR) image data. Our motivation is to exploit a weighted graph model constructed from characteristic points (i.e. keypoints) to measure the displacement vectors located at their positions. In fact, characteristic points are capable of capturing the image's radiometric and contextual information. Then, by encoding their interaction and inter-connection, a graph model is able to characterize both intensity and geometry information from the image content, which is relevant for texture tracking task. In this work, we employ a graph-based similarity measure to track the local texture information around each keypoint in order to figure out its correspondence from the other image and calculate the associated displacement. The proposed approach is tested and evaluated using high resolution TerraSAR-X images acquired from the Argentiere Glacier located in the French Alps. Our preliminary experimental results show the algorithm's capacity to provide a fast and reliable estimation of glacier flows, especially over highly textured and structured regions.
Minh-Tan Pham, Grégoire Mercier, Emmanuel Trouvé, Sébastien Lefèvre
IGARSS1
2016 Texture retrieval from very high resolution remote sensing images using local extrema-based descriptors
abstract
This paper proposes a novel approach for texture-based image indexing and retrieval in the scope of very high resolution (VHR) optical imagery. Our motivation is to take into account local textural features and structures inside each image to measure its similarity to other images. These local features are extracted for a set of characteristic points from the image using the local extrema-based descriptors (LED) from which the radiometric, spatial and gradient features of the local extrema pixels (i.e. maximums and minimums) are integrated to characterize local textures. Due to the fact that VHR images usually involve a variety of local textures which may weakly verify the stationarity hypothesis, an approach based on characteristic points like extrema pixels becomes relevant and effective. We perform our experimentation using texture databases extracted from VHR Pleiades images within the application of vineyard cultivation and oyster farming study. Retrieval results yielded by the proposed strategy are very promising and competitive compared to reference methods.
Minh-Tan Pham, Grégoire Mercier, Olivier Regniers, Lionel Bombrun, Julien Michel
IGARSS1
2016 Change Detection Between SAR Images Using a Pointwise Approach and Graph Theory
abstract
This paper investigates the problem of change detection in multitemporal synthetic aperture radar (SAR) images. Our motivation is to avoid using a large-size dense neighborhood around each pixel to measure its change level, which is usually considered by classical methods in order to perform their accurate detectors. Therefore, we propose to develop a pointwise approach to detect land-cover changes between two SAR images employing the principle of signal processing on graphs. First, a set of characteristic points is extracted from one of the two images to capture the image's significant contextual information. A weighted graph is then constructed to encode the interaction among these keypoints and hence capture the local geometric structure of this first image. With regard to this graph, the coherence of the information carried by the two images is considered for measuring changes between them. In other words, the change level will depend on how much the second image still conforms to the graph structure constructed from the first image. Additionally, due to the presence of speckle noise in SAR imaging, the log-ratio operator will be exploited to perform the image comparison measure. Experimental results performed on real SAR images show the effectiveness of the proposed algorithm, in terms of detection performance and computational complexity, compared to classical methods.
Minh-Tan Pham, Grégoire Mercier, Julien Michel
IEEE Trans. Geosci. Remote. Sens.1
2016 PW-COG: An Effective Texture Descriptor for VHR Satellite Imagery Using a Pointwise Approach on Covariance Matrix of Oriented Gradients
abstract
In this paper, a novel algorithm for textural feature description in very high resolution (VHR) satellite imagery is developed. It is based on a pointwise (PW) approach on the feature covariance matrix. The main motivation of this work is to construct the covariance matrix of oriented gradients (COG) using a nondense approach based on characteristic points (i.e., keypoints) extracted from the image. The proposed descriptor, which is named PW-COG, is expected to be effective when applied to VHR images. First, a COG-based descriptor is capable of not only capturing both radiometric and local geometric information (given by gradient features) from the image but also encoding their joint distribution and correlation, which are effectively relevant for texture characterization and discrimination. Second, by employing a keypoint-based approach, the proposed method is able to deal with large amount of VHR image data, since we do not take into consideration all pixels of the image, without requiring the stationarity hypothesis. In order to demonstrate the effectiveness of the proposed descriptor, texture-based image classification is carried out. Experimental study on both Brodatz texture database and VHR satellite images using the proposed algorithm provides very competitive results, in terms of texture discrimination and algorithm complexity, compared to reference methods.
Minh-Tan Pham, Grégoire Mercier, Julien Michel
IEEE Trans. Geosci. Remote. Sens.1
2015 Pointwise approach on covariance matrix of oriented gradients for very high resolution image texture segmentation
abstract
The present study involves an investigation of a pointwise approach on the feature covariance matrix to extract textu-ral features for very high resolution (VHR) satellite images. Indeed, our proposition is to construct the covariance matrix of oriented gradients using a non-dense approach based on characteristic points extracted from the image. This novel non-dense covariance descriptor is capable of not only capturing both radiometric and local geometric information from the image, but also encoding their joint distribution and correlation, which are effectively relevant for texture characterization and discrimination. In order to demonstrate the efficiency of the proposed descriptor, a texture-based image segmentation stage is carried out. First efforts on VHR panchromatic images using the proposed algorithm provide very promising and competitive results compared to classical methods.
Minh-Tan Pham, Grégoire Mercier, Julien Michel
IGARSS1
2015 Covariance-based texture description from weighted coherency matrix and gradient tensors for polarimetric SAR image classification
abstract
The present paper proposes a texture-based unsupervised classification algorithm for fully polarimetric SAR (PolSAR) images. Here, the main motivation is to combine polarimetric information and local structure gradients from PolSAR image data to describe textural features and then use them for classification purpose. In this work, the notion of PolSAR image textures is characterized by two key features. First, the polarimetric coherency matrix is estimated using a weighted averaging operator based on patch similarity. Second, the image local geometry is taken into account by exploiting the structure gradient tensors. These characteristics are then integrated into texture descriptors via the approach of covariance matrix. Unsupervised classification stage is finally achieved by employing an adapted distance measure for covariance-based descriptors. Experiments performed on very high resolution complex PolSAR images using the proposed algorithm provide very promising results in terms of terrain classification and discrimination.
Minh-Tan Pham, Grégoire Mercier, Julien Michel
IGARSS1
2014 Wavelets on graphs for very high resolution multispectral image texture segmentation
abstract
This paper proposes a texture-based segmentation method for very high spatial resolution imagery. Indeed, our main objective is to perform a sparse image representation modeled by a graph and then to exploit the wavelet transform on graph for the final purpose of image segmentation. Here, a set of pixels of interest, called representative pixels, is first extracted from the image and considered as vertices for constructing a weighted graph. Once the wavelet transform on graph is generated, their coefficients serve as textural features and will be exploited for unsupervised segmentation. Experimental results show the effectiveness of the proposed method when applied for very high spatial resolution multi-spectral images in terms of good segmentation precision as well as low complexity requirement.
Minh-Tan Pham, Grégoire Mercier, Julien Michel
IGARSS1