Jhony-Heriberto Giraldo-Zuluaga

dblp:170/0084 · also Jhony H. Giraldo · DBLP profile ↗
← Back
30ranked-venue papers
9as first author
25since 2021 · last 2026
0000-0002-0039-1270ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 5 first-author · 14 since 2021Artificial intelligence and machine learning · 16 · 5 first-author · 13 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 WildIng: A Wildlife Image Invariant Representation Model for Geographical Domain Shift
abstract
Abstract Wildlife monitoring is crucial for studying biodiversity loss and climate change. Camera trap images provide a non-intrusive method for analyzing animal populations and identifying ecological patterns over time. However, manual analysis is time-consuming and resource-intensive. Deep learning, particularly foundation models, has been applied to automate wildlife identification, achieving strong performance when tested on data from the same geographical locations as their training sets. Yet, despite their promise, these models struggle to generalize to new geographical areas, leading to significant performance drops. For example, training an advanced vision-language model, such as CLIP with an adapter, on an African dataset achieves an accuracy of 84.77%. However, this performance drops significantly to 16.17% when the model is tested on an American dataset. This limitation partly arises because existing models rely predominantly on image-based representations, making them sensitive to geographical data distribution shifts, such as variation in background, lighting, and environmental conditions. To address this, we introduce WildIng, a Wild life image In variant representation model for g eographical domain shift. WildIng integrates text descriptions with image features, creating a more robust representation to geographical domain shifts. By leveraging textual descriptions, our approach captures consistent semantic information, such as detailed descriptions of the appearance of the species, improving generalization across different geographical locations. Experiments show that WildIng enhances the accuracy of foundation models such as BioCLIP by 30% under geographical domain shift conditions. We evaluate WildIng on two datasets collected from different regions, namely America and Africa. The code and models are publicly available at https://github.com/Julian075/CATALOG/tree/WildIng .
Julian D. Santamaria, Claudia Isaza, Jhony-Heriberto Giraldo-Zuluaga
Int. J. Comput. Vis.3
2026 TEP-ones: A simple yet effective approach for transferability estimation of pruned backbones
abstract
In deep learning, the conventional transfer learning paradigm involves fine-tuning a model pre-trained on a complex source task to adapt it to a simpler target task, capitalizing on abundant training data. Concurrently, the paradigm of neural network pruning has emerged as a powerful strategy for enhancing model efficiency, reducing complexity, and optimizing resource utilization. This paper focuses on pruned model transferability estimation for resource-constraint scenarios, where the goal is to rank the performance of pruned pre-trained models on a downstream task without fine-tuning. To this end, from a formal analysis of the intra-class mutual information between samples belonging to the same target class, we observe that, as pruning increases, a sweet phase naturally rises, where the model benefits from better features at the encoder’s output. From this, we derive a Transferability Estimation for Pruned Backbones (TEP-ones) that eases the choice of which pruned model (without the need to train the classifier) is the best candidate for transfer learning.
Gabriele Spadaro, Andrea Bragagnolo, Riccardo Renzulli, Marco Grangetto, Jhony-Heriberto Giraldo-Zuluaga, Attilio Fiandrotti, Enzo Tartaglione
Neurocomputing5
2026 CALICE: Continuous Bitrate Control with Adapted LIC Model
abstract
Learned image compression (LIC) has drawn much attention recently as it outperforms standardized codecs in rate-distortion (RD) efficiency. However, an LIC model is typically trained for a specific RD tradeoff, and achieving a different target rate requires retraining the model and storing the weights as a whole, limiting the practical applicability of LIC. In this article, we introduce CALICE, a framework for achieving continuous bitrate control by plugging into a pre-trained LIC model a set of modular adapters. Unlike similar methods that require a distinct set of adapters for each target rate, our method achieves continuous bitrate control by modulating a single set of adapters via a scalar parameter \(\boldsymbol{\alpha}\) , with a total overhead of less than \(\mathbf{0.35}\boldsymbol{\%}\) of the parameters of the LIC model. This design enables efficient support for multiple distortion objectives by learning lightweight, distortion-aware adapters. We also extend our strategy beyond rate control, demonstrating its ability to provide fine-grained adaptation of perceptual quality along the distortion–perception tradeoff. To our knowledge, this is the first method that jointly addresses rate and perceptual control using a unified, low-cost strategy. We publicly released the code at https://github.com/EIDOSLAB/CALICE .
Gabriele Spadaro, Alberto Presta, Jhony-Heriberto Giraldo-Zuluaga, Attilio Fiandrotti, Marco Grangetto, Enzo Tartaglione
ACM Trans. Multim. Comput. Commun. Appl.3
2025 HYGENE: A Diffusion-Based Hypergraph Generation Method
abstract
Hypergraphs are powerful mathematical structures that can model complex, high-order relationships in various domains, including social networks, bioinformatics, and recommender systems. However, generating realistic and diverse hypergraphs remains challenging due to their inherent complexity and lack of effective generative models. In this paper, we introduce a diffusion-based Hypergraph Generation (HYGENE) method that addresses these challenges through a progressive local expansion approach. HYGENE works on the bipartite representation of hypergraphs, starting with a single pair of connected nodes and iteratively expanding it to form the target hypergraph. At each step, nodes and hyperedges are added in a localized manner using a denoising diffusion process, which allows for the construction of the global structure before refining local details. Our experiments demonstrated the effectiveness of HYGENE, proving its ability to closely mimic a variety of properties in hypergraphs. To the best of our knowledge, this is the first attempt to employ diffusion models for hypergraph generation.
Dorian Gailhard, Enzo Tartaglione, Lirida A. B. Naviner, Jhony-Heriberto Giraldo-Zuluaga
AAAI4
2025 Denoising Diffusion Probabilistic Model for Point Cloud Compression at Low Bit-Rates
abstract
Efficient compression of low-bit-rate point clouds is critical for bandwidth-constrained applications. However, existing techniques mainly focus on high-fidelity reconstruction, requiring many bits for compression. This paper proposes a "Denoising Diffusion Probabilistic Model" (DDPM) architecture for point cloud compression (DDPM-PCC) at low bit-rates. A PointNet encoder produces the condition vector for the generation, which is then quantized via a learnable vector quantizer. This configuration allows to achieve a low bitrates while preserving quality. Experiments on ShapeNet and ModelNet40 show improved rate-distortion at low rates compared to standardized and state-of-the-art approaches. We publicly released the code at https://github.com/EIDOSLAB/DDPM-PCC.
Gabriele Spadaro, Alberto Presta, Jhony-Heriberto Giraldo-Zuluaga, Marco Grangetto, Giuseppe Valenzise, Attilio Fiandrotti, Enzo Tartaglione
ICME3
2025 Continuous Simplicial Neural Networks
abstract
Simplicial complexes provide a powerful framework for modeling higher-order interactions in structured data, making them particularly suitable for applications such as trajectory prediction and mesh processing. However, existing simplicial neural networks (SNNs), whether convolutional or attention-based, rely primarily on discrete filtering techniques, which can be restrictive. In contrast, partial differential equations (PDEs) on simplicial complexes offer a principled approach to capture continuous dynamics in such structures. In this work, we introduce continuous simplicial neural network (COSIMO), a novel SNN architecture derived from PDEs on simplicial complexes. We provide theoretical and experimental justifications of COSIMO's stability under simplicial perturbations. Furthermore, we investigate the over-smoothing phenomenon—a common issue in geometric deep learning—demonstrating that COSIMO offers better control over this effect than discrete SNNs. Our experiments on real-world datasets demonstrate that COSIMO achieves competitive performance compared to state-of-the-art SNNs in complex and noisy environments. The implementation codes are available in https://github.com/ArefEinizade2/COSIMO.
Aref Einizade, Dorina Thanou, Fragkiskos D. Malliaros, Jhony-Heriberto Giraldo-Zuluaga
NeurIPS4
2025 Subgraph Gaussian Embedding Contrast for Self-supervised Graph Representation Learning
Shifeng Xie, Aref Einizade, Jhony-Heriberto Giraldo-Zuluaga
ECML/PKDD (6)3
2025 CATALOG: A Camera Trap Language-Guided Contrastive Learning Model
abstract
Foundation Models (FMs) have been successful in various computer vision tasks like image classification, object detection and image segmentation. However, these tasks remain challenging when these models are tested on datasets with different distributions from the training dataset, a problem known as domain shift. This is especially problematic for recognizing animal species in camera-trap images where we have variability in factors like lighting, camouflage and occlusions. In this paper, we propose the Camera Trap Language-guided Contrastive Learning (CATALOG) model to address these issues. Our approach combines multiple FMs to extract visual and textual features from camera-trap data and uses a contrastive loss function to train the model. We evaluate CATALOG on two benchmark datasets and show that it outperforms previous state-of-the-art methods in camera-trap image recognition, especially when the training and testing data have different animal species or come from different geographical areas. Our approach demonstrates the potential of using FMs in combination with multi-modal fusion and contrastive learning for addressing domain shifts in camera-trap image recognition. The code of CATALOG is publicly available at https://github.com/Julian075/CATALOG.
Julian D. Santamaria, Claudia Isaza, Jhony-Heriberto Giraldo-Zuluaga
WACV3
2025 WiGNet: Windowed Vision Graph Neural Network
abstract
In recent years, Graph Neural Networks (GNNs) have demonstrated strong adaptability to various real-world challenges, with architectures such as Vision GNN (ViG) achieving state-of-the-art performance in several computer vision tasks. However, their practical applicability is hindered by the computational complexity of constructing the graph, which scales quadratically with the image size. In this paper, we introduce a novel Windowed vision Graph neural Network (WiGNet) model for efficient image processing. WiGNet explores a different strategy from previous works by partitioning the image into windows and constructing a graph within each window. Therefore, our model uses graph convolutions instead of the typical 2D convolution or self-attention mechanism. WiGNet effectively manages computational and memory complexity for large image sizes. We evaluate our method in the ImageNet-1k benchmark dataset and test the adaptability of WiGNet using the CelebA-HQ dataset as a downstream task with higher-resolution images. In both of these scenarios, our method achieves competitive results compared to previous vision GNNs while keeping memory and computational complexity at bay. WiGNet offers a promising solution toward the deployment of vision GNNs in real-world applications. We publicly released the code and pre-trained models at https://github.com/EIDOSLAB/WiGNet.
Gabriele Spadaro, Marco Grangetto, Attilio Fiandrotti, Enzo Tartaglione, Jhony-Heriberto Giraldo-Zuluaga
WACV5
2025 Graph-based Moving Object Segmentation for underwater videos using semi-supervised learning
abstract
International audience
Meghna Kapoor, Wieke Prummel, Jhony-Heriberto Giraldo-Zuluaga, Badri N. Subudhi, Anastasia Zakharova, Thierry Bouwmans, Ankur Bansal
Comput. Vis. Image Underst.3
2024 Privacy-Preserving Adaptive Re-Identification Without Image Transfer
Hamza Rami, Jhony-Heriberto Giraldo-Zuluaga, Nicolas Winckler, Stéphane Lathuilière
ECCV (52)2
2024 Gabic: Graph-Based Attention Block for Image Compression
abstract
While standardized codecs like JPEG and HEVC-intra represent the industry standard in image compression, neural Learned Image Compression (LIC) codecs represent a promising alternative. In detail, integrating attention mechanisms from Vision Transformers into LIC models has shown improved compression efficiency. However, extra efficiency often comes at the cost of aggregating redundant features. This work proposes a Graph-based Attention Block for Image Compression (GABIC), a method to reduce feature redundancy based on a k-Nearest Neighbors enhanced attention mechanism. Our experiments show that GABIC outperforms comparable methods, particularly at high bit rates, enhancing compression performance.
Gabriele Spadaro, Alberto Presta, Enzo Tartaglione, Jhony-Heriberto Giraldo-Zuluaga, Marco Grangetto, Attilio Fiandrotti
ICIP4
2024 OVOSE: Open-Vocabulary Semantic Segmentation in Event-Based Cameras
Muhammad Rameez Ur Rahman, Jhony-Heriberto Giraldo-Zuluaga, Indro Spinelli, Stéphane Lathuilière, Fabio Galasso
ICPR (16)2
2024 Continuous Product Graph Neural Networks
abstract
Processing multidomain data defined on multiple graphs holds significant potential in various practical applications in computer science. However, current methods are mostly limited to discrete graph filtering operations. Tensorial partial differential equations on graphs (TPDEGs) provide a principled framework for modeling structured data across multiple interacting graphs, addressing the limitations of the existing discrete methodologies. In this paper, we introduce Continuous Product Graph Neural Networks (CITRUS) that emerge as a natural solution to the TPDEG. CITRUS leverages the separability of continuous heat kernels from Cartesian graph products to efficiently implement graph spectral decomposition. We conduct thorough theoretical analyses of the stability and over-smoothing properties of CITRUS in response to domain-specific graph perturbations and graph spectra effects on the performance. We evaluate CITRUS on well-known traffic and weather spatiotemporal forecasting datasets, demonstrating superior performance over existing approaches. The implementation codes are available at https://github.com/ArefEinizade2/CITRUS.
Aref Einizade, Fragkiskos D. Malliaros, Jhony-Heriberto Giraldo-Zuluaga
NeurIPS3
2024 ALICE: Adapt your Learnable Image Compression modEl for variable bitrates
abstract
When training a Learned Image Compression model, the loss function is minimized such that the encoder and the decoder attain a target Rate-Distorsion trade-off. Therefore, a distinct model shall be trained and stored at the transmitter and receiver for each target rate, fostering the quest for efficient variable bitrate compression schemes. This paper proposes plugging Low-Rank Adapters into a transformer-based pre-trained LIC model and training them to meet different target rates. With our method, encoding an image at a variable rate is as simple as training the corresponding adapters and plugging them into the frozen pre-trained model. Our experiments show performance comparable with state-of-the-art fixed-rate LIC models at a fraction of the training and deployment cost. We publicly released the code at https://github.com/EIDOSLAB/ALICE.
Gabriele Spadaro, Muhammad Salman Ali, Alberto Presta, Giommaria Pilo, Sung-Ho Bae, Jhony-Heriberto Giraldo-Zuluaga, Attilio Fiandrotti, Marco Grangetto, Enzo Tartaglione
VCIP6
2024 Source-Guided Similarity Preservation for Online Person Re-Identification
abstract
Online Unsupervised Domain Adaptation (OUDA) for person Re-Identification (Re-ID) is the task of continuously adapting a model trained on a well-annotated source-domain dataset to a target domain observed as a data stream. In OUDA, person Re-ID models face two main challenges: catastrophic forgetting and domain shift. In this work, we propose a new Source-guided Similarity Preservation (S2P) framework to alleviate these two problems. Our framework is based on the extraction of a support set composed of source images that maximizes the similarity with the target data. This support set is used to identify feature similarities that must be preserved during the learning process. S2P can incorporate multiple existing UDA methods to mitigate catastrophic forgetting. Our experiments show that S2P outperforms previous state-of-the-art methods on multiple real-to-real and synthetic-to-real challenging OUDA benchmarks.
Hamza Rami, Jhony-Heriberto Giraldo-Zuluaga, Nicolas Winckler, Stéphane Lathuilière
WACV2
2024 Gegenbauer Graph Neural Networks for Time-Varying Signal Reconstruction
abstract
Reconstructing time-varying graph signals (or graph time-series imputation) is a critical problem in machine learning and signal processing with broad applications, ranging from missing data imputation in sensor networks to time-series forecasting. Accurately capturing the spatio-temporal information inherent in these signals is crucial for effectively addressing these tasks. However, existing approaches relying on smoothness assumptions of temporal differences and simple convex optimization techniques that have inherent limitations. To address these challenges, we propose a novel approach that incorporates a learning module to enhance the accuracy of the downstream task. To this end, we introduce the Gegenbauer-based graph convolutional (GegenConv) operator, which is a generalization of the conventional Chebyshev graph convolution by leveraging the theory of Gegenbauer polynomials. By deviating from traditional convex problems, we expand the complexity of the model and offer a more accurate solution for recovering time-varying graph signals. Building upon GegenConv, we design the Gegenbauer-based time graph neural network (GegenGNN) architecture, which adopts an encoder-decoder structure. Likewise, our approach also uses a dedicated loss function that incorporates a mean squared error (MSE) component alongside Sobolev smoothness regularization. This combination enables GegenGNN to capture both the fidelity to ground truth and the underlying smoothness properties of the signals, enhancing the reconstruction performance. We conduct extensive experiments on real datasets to evaluate the effectiveness of our proposed approach. The experimental results demonstrate that GegenGNN outperforms state-of-the-art methods, showcasing its superior capability in recovering time-varying graph signals.
Jhon A. Castro-Correa, Jhony-Heriberto Giraldo-Zuluaga, Mohsen Badiey, Fragkiskos D. Malliaros
IEEE Trans. Neural Networks Learn. Syst.2
2023 On the Trade-off between Over-smoothing and Over-squashing in Deep Graph Neural Networks
abstract
Graph Neural Networks (GNNs) have succeeded in various computer science applications, yet deep GNNs underperform their shallow counterparts despite deep learning's success in other domains. Over-smoothing and over-squashing are key challenges when stacking graph convolutional layers, hindering deep representation learning and information propagation from distant nodes. Our work reveals that over-smoothing and over-squashing are intrinsically related to the spectral gap of the graph Laplacian, resulting in an inevitable trade-off between these two issues, as they cannot be alleviated simultaneously. To achieve a suitable compromise, we propose adding and removing edges as a viable approach. We introduce the Stochastic Jost and Liu Curvature Rewiring (SJLR) algorithm, which is computationally efficient and preserves fundamental properties compared to previous curvature-based methods. Unlike existing approaches, SJLR performs edge addition and removal during GNN training while maintaining the graph unchanged during testing. Comprehensive comparisons demonstrate SJLR's competitive performance in addressing over-smoothing and over-squashing.
Jhony-Heriberto Giraldo-Zuluaga, Konstantinos Skianis, Thierry Bouwmans, Fragkiskos D. Malliaros
CIKM1
2023 Time-Varying Signals Recovery Via Graph Neural Networks
abstract
The recovery of time-varying graph signals is a fundamental problem with numerous applications in sensor networks and forecasting in time series. Effectively capturing the spatiotemporal information in these signals is essential for the downstream tasks. Previous studies have used the smoothness of the temporal differences of such graph signals as an initial assumption. Nevertheless, this smoothness assumption could result in a degradation of performance in the corresponding application when the prior does not hold. In this work, we relax the requirement of this hypothesis by including a learning module. We propose a Time Graph Neural Network (TimeGNN) for the recovery of time-varying graph signals. Our algorithm uses an encoder-decoder architecture with a specialized loss composed of a mean squared error function and a Sobolev smoothness operator. TimeGNN shows competitive performance against previous methods in real datasets.
Jhon A. Castro-Correa, Jhony-Heriberto Giraldo-Zuluaga, Anindya Mondal, Mohsen Badiey, Thierry Bouwmans, Fragkiskos D. Malliaros
ICASSP2
2023 Higher-Order Sparse Convolutions in Graph Neural Networks
abstract
Graph Neural Networks (GNNs) have been applied to many problems in computer sciences. Capturing higher-order relationships between nodes is crucial to increase the expressive power of GNNs. However, existing methods to capture these relationships could be infeasible for large-scale graphs. In this work, we introduce a new higher-order sparse convolution based on the Sobolev norm of graph signals. Our Sparse Sobolev GNN (S-SobGNN) computes a cascade of filters on each layer with increasing Hadamard powers to get a more diverse set of functions, and then a linear combination layer weights the embeddings of each filter. We evaluate S-SobGNN in several applications of semi-supervised learning. S-SobGNN shows competitive performance in all applications as compared to several state-of-the-art methods.
Jhony-Heriberto Giraldo-Zuluaga, Sajid Javed, Arif Mahmood, Fragkiskos D. Malliaros, Thierry Bouwmans
ICASSP1
2023 Inductive Graph Neural Networks for Moving Object Segmentation
abstract
Moving Object Segmentation (MOS) is a challenging problem in computer vision, particularly in scenarios with dynamic backgrounds, abrupt lighting changes, shadows, camouflage, and moving cameras. While graph-based methods have shown promising results in MOS, they have mainly relied on transductive learning which assumes access to the entire training and testing data for evaluation. However, this assumption is not realistic in real-world applications where the system needs to handle new data during deployment. In this paper, we propose a novel Graph Inductive Moving Object Segmentation (GraphIMOS) algorithm based on a Graph Neural Network (GNN) architecture. Our approach builds a generic model capable of performing prediction on newly added data frames using the already trained model. GraphI-MOS outperforms previous inductive learning methods and is more generic than previous transductive techniques. Our proposed algorithm enables the deployment of graph-based MOS models in real-world applications.
Wieke Prummel, Jhony-Heriberto Giraldo-Zuluaga, Anastasia Zakharova, Thierry Bouwmans
ICIP2
2023 Uncertainty clustering internal validity assessment using Fréchet distance for unsupervised learning
abstract
Knowing the number of clusters a priori is one of the most challenging aspects of unsupervised learning. Clustering Internal Validity Indices (CIVIs) evaluate partitions in unsupervised algorithms based on metrics like compactness, separation, and density. However, specialized CIVIs for specific applications have been designed, and there is no general CIVI that works in all scenarios. The absence of CIVIs based on crisp uncertainty metrics is especially critical in decision-making processes that involve ambiguity, non-convex distributions, outliers, and overlapping data. To address this problem, we propose a novel Uncertainty Fréchet (UF) CIVI that assesses the certainty of a well-defined partition. UF leverages uncertainty fingerprints based on Type-2 fuzzy Gaussian Mixture Models (T2FGMM) and the Fréchet distance between clusters to introduce a metric that evaluates partition quality. We integrate UF into a merging methodology that combines similar clusters within a partition, allowing us to determine the number of clusters without the need to run the clustering algorithms iteratively as other CIVIs require. We undertake a comprehensive evaluation of our proposal on 5,250 convex, 36 non-convex synthetic datasets, and five benchmark real datasets. In addition, we apply UF in a real-world scenario that involves high uncertainty: Passive Acoustic Monitoring (PAM) of ecosystems, which aims to study ecological transformations through acoustic recordings. The results show that UF exhibits notable performance in synthetic and real-world scenarios, obtaining an Adjusted Mutual Information (AMI) score higher than 0.88 for normal, uniform, gamma, and triangular distribution datasets. In the PAM application, UF identifies the transformation of ecosystems through sound using clustering algorithms and UF, achieving an F1 score of 0.84. Therefore, results show that the UF index is a suitable tool for researchers and practitioners working with highly uncertain data.
Nestor Rendon, Jhony-Heriberto Giraldo-Zuluaga, Thierry Bouwmans, Susana Rodríguez-Buritica, Edison Ramirez, Claudia Isaza
Eng. Appl. Artif. Intell.2
2022 Hypergraph Convolutional Networks for Weakly-Supervised Semantic Segmentation
abstract
Semantic segmentation is a fundamental topic in computer vision. Several deep learning methods have been proposed for semantic segmentation with outstanding results. However, these models require a lot of densely annotated images. To address this problem, we propose a new algorithm that uses Hy-perGraph Convolutional Networks for Weakly-supervised Semantic Segmentation (HyperGCN-WSS). Our algorithm constructs spatial and k-Nearest Neighbor (k-NN) graphs from the images in the dataset to generate the hypergraphs. Then, we train a specialized HyperGraph Convolutional Network (HyperGCN) architecture using some weak signals. The outputs of the HyperGCN are denominated pseudo-labels, which are later used to train a DeepLab model for semantic segmentation. HyperGCN-WSS is evaluated on the PASCAL VOC 2012 dataset for semantic segmentation, using scribbles or clicks as weak signals. Our algorithm shows competitive performance against previous methods.
Jhony-Heriberto Giraldo-Zuluaga, Vincenzo Mariano Scarrica, Antonino Staiano, Francesco Camastra, Thierry Bouwmans
ICIP1
2022 SemiSegSAR: A Semi-Supervised Segmentation Algorithm for Ship SAR Images
abstract
Automatic ship segmentation from high-resolution Synthetic Aperture Radar (SAR) remote sensing images has been a topic of interest that has gradually gained attention over the years due to the abundance of earth observation sensors. Recently, deep learning methods have provided a breakthrough increasing the performance greatly by using large amount of labeled data. Yet, the high cost related to the samples labeling and their scarcity result in significant limitation of their wide use. Therefore, it is crucial to overcome the unlabeled inputs challenge and develop semi-supervised learning approaches to enhance the machine learning models capacity. Our letter proposes a semi-supervised segmentation algorithm for SAR images named SemiSegSAR based on the use of Graph Signal Processing. This method includes instance segmentation; texture and statistical SAR features to represent the nodes of the graph; K-nearest neighbors to construct the graph; and Sobolev minimization algorithm to tackle the problem of semi-supervised semantic segmentation. The proposed algorithm is trained and tested using the publicly available SSDD and HRSID ship detection datasets. Experiments show that SemiSegSAR outperforms the current state-of-the-art semi-supervised and supervised methods while requiring only few labeled data.
Marwa Chendeb, Jhony-Heriberto Giraldo-Zuluaga, Mina Al-Saad, Muna Darweesh, Thierry Bouwmans
IEEE Geosci. Remote. Sens. Lett.2
2022 Graph Moving Object Segmentation
abstract
Moving Object Segmentation (MOS) is a fundamental task in computer vision. Due to undesirable variations in the background scene, MOS becomes very challenging for static and moving camera sequences. Several deep learning methods have been proposed for MOS with impressive performance. However, these methods show performance degradation in the presence of unseen videos; and usually, deep learning models require large amounts of data to avoid overfitting. Recently, graph learning has attracted significant attention in many computer vision applications since they provide tools to exploit the geometrical structure of data. In this work, concepts of graph signal processing are introduced for MOS. First, we propose a new algorithm that is composed of segmentation, background initialization, graph construction, unseen sampling, and a semi-supervised learning method inspired by the theory of recovery of graph signals. Second, theoretical developments are introduced, showing one bound for the sample complexity in semi-supervised learning, and two bounds for the condition number of the Sobolev norm. Our algorithm has the advantage of requiring less labeled data than deep learning methods while having competitive results on both static and moving camera videos. Our algorithm is also adapted for Video Object Segmentation (VOS) tasks and is evaluated on six publicly available datasets outperforming several state-of-the-art methods in challenging conditions.
Jhony-Heriberto Giraldo-Zuluaga, Sajid Javed, Thierry Bouwmans
IEEE Trans. Pattern Anal. Mach. Intell.1
2020 Semi-Supervised Background Subtraction Of Unseen Videos: Minimization Of The Total Variation Of Graph Signals
abstract
Recently, several successful methods based on deep neural networks have been proposed for background subtraction. These deep neural algorithms have almost perfect performance, relying in the availability of ground-truth frames of the tested videos during the training step. However, the performance of some of these algorithms drops significantly when tested on unseen videos. In this paper, concepts of semi-supervised learning are introduced in the problem of background subtraction for unseen videos. We propose a new algorithm named Graph-BGS-TV, this method uses: Mask R-CNN for instances segmentation; temporal median filter for background initialization; motion, texture, and intensity features for representing the nodes of a graph; k-nearest neighbors for the construction of the graph; and finally a total variation minimization algorithm to solve the problem of background subtraction. GraphBGS-TV is tested in the change detection dataset, outperforming unsupervised and supervised methods in the challenges “PTZ” and “shadows”.
Jhony-Heriberto Giraldo-Zuluaga, Thierry Bouwmans
ICIP1
2020 GraphBGS: Background Subtraction via Recovery of Graph Signals
abstract
Background subtraction is a fundamental preprocessing task in computer vision. This task becomes challenging in real scenarios due to variations in the background for both static and moving camera sequences. Several deep learning methods for background subtraction have been proposed in the literature with competitive performances. However, these models show performance degradation when tested on unseen videos; and they require huge amount of data to avoid overfitting. Recently, graph-based algorithms have been successful approaching unsupervised and semi-supervised learning problems. Furthermore, the theory of graph signal processing and semi-supervised learning have been combined leading to new insights in the field of machine learning. In this paper, concepts of recovery of graph signals are introduced in the problem of background subtraction. We propose a new algorithm called Graph BackGround Subtraction (GraphBGS), which is composed of: instance segmentation, background initialization, graph construction, graph sampling, and a semi-supervised algorithm inspired from the theory of recovery of graph signals. Our algorithm has the advantage of requiring less labeled data than deep learning methods while having competitive results on both: static and moving camera videos. GraphBGS outperforms unsupervised and supervised methods in several challenging conditions on the publicly available Change Detection (CDNet2014), and UCSD background subtraction databases.
Jhony-Heriberto Giraldo-Zuluaga, Thierry Bouwmans
ICPR1
2019 Camera-trap images segmentation using multi-layer robust principal component analysis
Jhony-Heriberto Giraldo-Zuluaga, Augusto Salazar, Alexandra Gomez-Villa, Angélica Diaz-Pulido
Vis. Comput.1
2018 Automatic identification of Scenedesmus polymorphic microalgae from microscopic images
Jhony-Heriberto Giraldo-Zuluaga, Augusto Salazar, German Díez, Alexandra Gomez-Villa, Tatiana Martinez, Jesús Francisco Vargas-Bonilla, Mariana Peñuela Vasquez
Pattern Anal. Appl.1
2017 Recognition of Mammal Genera on Camera-Trap Images Using Multi-layer Robust Principal Component Analysis and Mixture Neural Networks
abstract
The segmentation and classification of animals from camera-trap images is a difficult task due to the conditions under which the images are taken. This work presents a method for recognizing mammal genera from camera-trap images. Our method uses Multi-Layer Robust Principal Component Analysis (RPCA) for segmenting, Convolutional Neural Networks (CNNs) for extracting features, Least Absolute Shrinkage and Selection Operator (LASSO) for selecting features, and Artificial Neural Networks (ANNs) or Support Vector Machines (SVM) for classifying mammal genera present in the Colombian forest. Our classification method mixes the features of several CNNs. We evaluated our method with the camera-trap images from the Instituto de Investigación de Recursos Biológicos Alexander von Humboldt. We obtained an accuracy of 92.65% classifying 8 mammal genera and a False Positive (FP) class, using automatic-segmented images. On the other hand, we reached 90.32% of accuracy classifying 10 mammal genera, using ground-truth images only. Unlike all previous works, we confront the animal segmentation and genera classification on the camera-trap framework. This method shows a new approach toward a fully-automatic detection of animals from camera-trap images.
Jhony-Heriberto Giraldo-Zuluaga, Augusto Salazar, Alexandra Gomez-Villa, Angélica Diaz-Pulido
ICTAI1