Eric Granger

dblp:86/2306 · also Éric Granger · DBLP profile ↗
← Back
165ranked-venue papers
8as first author
67since 2021 · last 2026
0000-0001-6116-7945ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 102 · 7 first-author · 39 since 2021Graphics, computer vision, multimedia, augmented reality and games · 81 · 1 first-author · 40 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 since 2021Databases, data management, data science and information retrieval · 6 · 1 since 2021Systems, architecture and hardware · 2Security and privacy · 2 · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 Beyond Patches: Mining Interpretable Part-Prototypes for Explainable AI
abstract
As AI systems become more capable, it is important that their decisions are understandable and aligned with human expectations. A key challenge is the lack of interpretability in deep models. Existing methods such as GradCAM generate heatmaps but provide limited conceptual insight, while prototype-based approaches offer example-based explanations but often rely on rigid region selection and lack semantic consistency. To address these limitations, we propose PCMNet, a Part-Prototypical Concept Mining Network that learns human-comprehensible prototypes from meaningful regions without extra supervision. By clustering these into concept groups and extracting concept activation vectors, PCMNet provides structured, concept-level explanations and enhances robustness under occlusion and adversarial conditions, which are both critical for building reliable and aligned AI systems. Experiments across multiple benchmarks show that PCMNet outperforms state-of-the-art methods in interpretability, stability, and robustness. This work contributes to AI alignment by enhancing transparency, controllability, and trustworthiness in modern AI systems.
Mahdi Alehdaghi, Rajarshi Bhattacharya, Pourya Shamsolmoali, Rafael M. O. Cruz, Eric Granger
AAAI5
2026 DogFit: Domain-guided Fine-tuning for Efficient Transfer Learning of Diffusion Models
abstract
Transfer learning of diffusion models to smaller target domains is challenging, as naively fine-tuning the model often results in poor generalization. Test-time guidance methods help mitigate this by offering controllable improvements in image fidelity through a trade-off with sample diversity. However, this benefit comes at a high computational cost, typically requiring dual forward passes during sampling. We propose the Domain-guided Fine-tuning (DogFit) method, an effective guidance mechanism for diffusion transfer learning that maintains controllability without incurring additional computational overhead. DogFit injects a domain-aware guidance offset into the training loss, effectively internalizing the guided behavior during the fine-tuning process. The domain-aware design is motivated by our observation that during fine-tuning, the unconditional source model offers a stronger marginal estimate than the target model. To support efficient controllable fidelity–diversity trade-offs at inference, we encode the guidance strength value as an additional model input through a lightweight conditioning mechanism. We further investigate the optimal placement and timing of the guidance offset during training and propose two simple scheduling strategies, i.e., late-start and cut-off, which improve generation quality and training stability. Experiments on DiT and SiT backbones across six diverse target domains show that DogFit can outperform prior guidance methods in transfer learning in terms of FID and FD DINOV2 while requiring up to 2x fewer sampling TFLOPS.
Yara Bahram, Mohammadhadi Shateri, Eric Granger
AAAI3
2026 MixER: From Cross-Modal to Mixed-Modal Visible-Infrared Re-Identification
Mahdi Alehdaghi, Rajarshi Bhattacharya, Pourya Shamsolmoali, Rafael M. O. Cruz, Eric Granger
WACV5
2026 High-Rate Mixout: Revisiting Mixout for Robust Domain Generalization
abstract
Ensembling fine-tuned models initialized from powerful pre-trained weights is a common strategy to improve robustness under distribution shifts, but it comes with substantial computational costs due to the need to train and store multiple models. Dropout offers a lightweight alternative by simulating ensembles through random neuron deactivation; however, when applied to pre-trained models, it tends to over-regularize and disrupt critical representations necessary for generalization. In this work, we investigate Mixout, a stochastic regularization technique that provides an alternative to Dropout for domain generalization. Rather than deactivating neurons, Mixout mitigates overfitting by probabilistically swapping a subset of fine-tuned weights with their pre-trained counterparts during training, thereby maintaining a balance between adaptation and retention of prior knowledge. Our study reveals that achieving strong performance with Mixout on domain generalization benchmarks requires a notably high masking probability of 0.9 for ViTs and 0.8 for ResNets. While this may seem like a simple adjustment, it yields two key advantages for domain generalization: (1) higher masking rates more strongly penalize deviations from the pre-trained parameters, promoting better generalization to unseen domains; and (2) high-rate masking substantially reduces computational overhead, cutting gradient computation by up to 45% and gradient memory usage by up to 90%. Experiments across five domain generalization benchmarks—PACS, VLCS, OfficeHome, TerraIncognita, and DomainNet—using ResNet and ViT architectures show that our approach, High-rate Mixout, achieves out-of-domain accuracy comparable to ensemble-based methods while significantly reducing training costs.
Masih Aminbeidokhti, Heitor Rapela Medeiros, Srikanth Muralidharan, Eric Granger, Marco Pedersoli
WACV4
2026 CLIP-IT: CLIP-based Pairing of Histology Images with Privileged Textual Information
abstract
Multimodal learning has shown promise in medical imaging, combining complementary modalities like images and text. Vision-language models (VLMs) capture rich diagnostic cues but often require large paired datasets and promptor text-based inference. Their practicality is therefore limited due to annotation cost, privacy, and compute demands. Unpaired external text, like pathology reports, can still provide complementary diagnostic cues if semantically relevant content is retrievable per image. To address this, we introduce CLIP-IT, a novel framework that relies on rich unpaired text reports. Specifically, CLIP-IT uses a CLIP model pre-trained on histology image–text pairs from a separate dataset to retrieve the most relevant unpaired textual report for each image in the downstream unimodal dataset. These reports, sourced from the same disease domain and tissue type, form pseudo-pairs that reflect shared clinical semantics rather than exact alignment. Knowledge from these texts is distilled into the vision model during training, while LoRA-based adaptation mitigates the semantic gap between unaligned modalities. At inference, only the vision model is used, maintaining low overhead while still benefiting from multimodal training without requiring paired data in the downstream dataset. Experiments1show that CLIP-IT consistently improves classification accuracy over both unimodal and multimodal CLIP-based baselines in most cases, without requiring paired annotations per dataset or incurring additional inference-time complexity.
Banafsheh Karimian, Giulia Avanzato, Soufiane Belharbi, Alexis Guichemerre, Luke McCaffrey, Mohammadhadi Shateri, Eric Granger
WACV7
2026 WiSE-OD: Benchmarking Robustness in Infrared Object Detection
abstract
Object detection (OD) in infrared (IR) imagery is critical for low-light and nighttime applications. However, the scarcity of large-scale IR datasets forces models to rely on weights pre-trained on RGB images. While fine-tuning on IR improves accuracy, it often compromises robustness under distribution shifts due to the inherent modality gap between RGB and IR. To address this, we introduce LLVIP-C and FLIR-C, two cross-modality out-of-distribution (OOD) benchmarks built by applying corruptions to standard IR datasets. Additionally, to fully leverage the complementary knowledge from RGB and infrared-trained models, we propose WiSE-OD, a weight-space ensembling method with two variants: WiSE-ODZS, which combines RGB zero-shot and IR fine-tuned weights, and WiSE-ODLP, which blends zero-shot and linear probing. Evaluated using four RGB-pretrained detectors and two robust baselines on our benchmark and in the real-world out-of-distribution M3FD dataset, our WiSE-OD improves robustness across modalities and to corruption in synthetic and real-world distribution shifts without any additional training or inference costs. Our code is available at: https://github.com/heitorrapela/wiseod.
Heitor Rapela Medeiros, Atif Belal, Masih Aminbeidokhti, Eric Granger, Marco Pedersoli
WACV4
2026 Pretraining Helps When Capacity Allows: Evidence from Ultra-Small ConvNets
abstract
Robust visual recognition on embedded platforms requires models that both generalize out-of-distribution (OOD) and fit into tiny compute/memory budgets. While pre-training is a standard route to robustness for mid/large backbones, its value in the ultra-small regime remains unclear. We present a capacity-aware study of pre-training for two efficient ConvNet families (EfficientNet and MobileNetV3) scaled from "small" to "ultra-small" via a simple, reproducible recipe. We compare three initializations — ImageNet→COCO pretraining, ImageNet classification pretraining, and training from scratch—on two axes of distribution shift: (i) cross-dataset RGB→RGB transfer between LLVIP and FLIR (ii) cross-modality detection where models are fine-tuned on RGB and evaluated on infrared (IR). A complementary classification study on DomainNet probes whether the trends extend beyond detection. Across settings, we find that pretraining’s benefit is conditional on both backbone capacity and shift difficulty. Task-aligned Imagenet→COCO pretraining is the most reliable starting point at moderate sizes and for the easier transfer direction. In the low-capacity regimes, differences are typically within run-to-run variation, and training from scratch can match or surpass pre-training. Classification mirrors this capacity gating. Our results test the premise "pretraining always helps" and instead quantify when task-aligned pretraining pays off for ultra-small backbones and when it likely does not1.
Srikanth Muralidharan, Heitor Rapela Medeiros, Masih Aminbeidokhti, Eric Granger, Marco Pedersoli
WACV4
2026 Low-Rank Expert Merging for Multi-Source Domain Adaptation in Person Re-Identification
abstract
Adapting person re-identification (reID) models to new target environments remains a challenging problem that is typically addressed using unsupervised domain adaptation (UDA) methods. Recent works show that when labeled data originates from several distinct sources (e.g., datasets and cameras), considering each source separately and applying multi-source domain adaptation (MSDA) typically yields higher accuracy and robustness compared to blending the sources and performing conventional UDA. However, state-of-the-art MSDA methods learn domain-specific backbone models or require access to source domain data during adaptation, resulting in significant growth in training parameters and computational cost. In this paper, a Source-free Adaptive Gated Experts (SAGE-reID) method is introduced for person reID. Our SAGE-reID is a cost-effective, source-free MSDA method that first trains individual source-specific low-rank adapters (LoRA) through source-free UDA. Then, a lightweight gating network is introduced and trained to dynamically assign optimal merging weights for fusion of LoRA experts, enabling effective cross-domain knowledge transfer. While the number of backbone parameters remains constant across source domains, LoRA experts scale linearly but remain negligible in size (ď 2% per source), reducing both the memory consumption and risk of overfitting. Extensive experiments conducted on three challenging benchmarks – Market-1501, DukeMTMC-reID, and MSMT17 – indicate that SAGE-reID can outperform state-of-the-art methods while remaining computationally efficient. Our code is available: https://github.com/nehdiii/SAGE-reID
Taha Mustapha Nehdi, Nairouz Mrabah, Atif Belal, Marco Pedersoli, Eric Granger
WACV5
2026 MuSACo: Multimodal Subject-Specific Selection and Adaptation for Expression Recognition with Co-Training
abstract
Personalized expression recognition (ER) involves adapting a machine learning model to subject-specific data for improved recognition of expressions with considerable inter-personal variability. Subject-specific ER can benefit significantly from multi-source domain adaptation (MSDA) methods – where each domain corresponds to a specific subject – to improve model accuracy and robustness. Despite promising results, state-of-the-art MSDA approaches often overlook multimodal information or blend sources into a single domain, limiting subject diversity and failing to explicitly capture unique subject-specific characteristics. To address these limitations, we introduce MuSACo, a multi-modal subject-specific selection and adaptation method for ER based on co-training. It leverages complementary information across multiple modalities and multiple source domains for subject-specific adaptation. This makes MuSACo particularly relevant for affective computing applications in digital health, such as patient-specific assessment for stress or pain, where subject-level nuances are crucial. MuSACo selects source subjects relevant to the target and generates pseudo-labels using the dominant modality for class-aware learning, in conjunction with a class-agnostic loss to learn from less confident target samples. Finally, source features from each modality are aligned, while only confident target features are combined. Experimental results on challenging multimodal ER datasets – BioVid, StressID, and BAH – show that MuSACo outperforms UDA (blending) and state-of-the-art MSDA methods. Our code is available: https://github.com/osamazeeshan/MuSACo
Muhammad Osama Zeeshan, Natacha Gillet, Alessandro L. Koerich, Marco Pedersoli, François Brémond, Eric Granger
WACV6
2026 Weakly Supervised Learning for Facial Affective Behavior Analysis: A Review
abstract
Recent advances in deep learning (DL) and computational capacity have enabled facial affective behavior analysis (FABA) to progress from static images captured in controlled settings to fine-grained analysis of facial expressions in real world video data. However, training accurate DL models for FABA typically requires large-scale, expert-annotated datasets, which are costly to obtain and inherently noisy due to the ambiguity of labeling subtle facial expressions and action units (AUs). To mitigate these challenges, weakly supervised learning (WSL) has emerged as a promising paradigm for training models with weak annotations. In this paper, we present a structured taxonomy of WSL scenarios for FABA, organized according to the type of weak annotation and the specific affective task. Building on this taxonomy, we provide a critical synthesis of representative WSL methods for both classification (expression and AU recognition) and regression (expression and AU intensity estimation) tasks, focusing on their core methodological ideas, strengths, and limitations. Furthermore, we systematically summarize the comparative performance of WSL approaches along with widely adopted experimental setups and evaluation proto cols. Our critical assessment identifies key challenges and future research directions, including the need for efficient adaptation of foundation models and for the development of robust, scalable FABA systems suitable for real-world applications.
Gnana Praveen Rajasekhar, Patrick Cardinal, Eric Granger
IEEE Trans. Affect. Comput.3
2026 Progressive Multi-Source Domain Adaptation for Personalized Facial Expression Recognition
abstract
Personalized facial expression recognition (FER) involves adapting a machine learning model using samples from labeled sources and unlabeled target domains. Given the challenges of recognizing subtle expressions with considerable interpersonal variability, state-of-the-art unsupervised domain adaptation (UDA) methods focus on the multi-source UDA (MSDA) setting, where each domain corresponds to a specific subject, and improve model accuracy and robustness. However, when adapting to a specific target, the diverse nature of multiple source domains translates to a large shift between source and target data. State-of-the-art MSDA methods for FER address this domain shift by considering all the sources to adapt to the target representations. Nevertheless, adapting to a target subject presents significant challenges due to large distributional differences between source and target domains, often resulting in negative transfer. In addition, integrating all sources simultaneously increases computational costs and causes misalignment with the target. To address these issues, we propose a progressive MSDA approach that gradually introduces information from subjects (source domains) based on their similarity to the target subject. This will ensure that only the most relevant sources from the target are selected, which helps avoid the negative transfer caused by dissimilar sources. During adaptation, the source domains are introduced in a curriculum manner. We first exploit the closest sources to reduce the distribution shift with the target and then move towards the furthest while only considering the most relevant sources based on the predetermined threshold. Furthermore, to mitigate catastrophic forgetting caused by the incremental introduction of source subjects, we implemented a density-based memory mechanism that preserves the most relevant historical source samples for adaptation. Our extensive experiments 1 show the effectiveness of our proposed method on challenging FER datasets: Biovid, UNBC-McMaster, Aff-Wild2, and BAH. Further, performance is evaluated on a cross-dataset setting (UNBC-McMaster → BioVid), showing the importance of gradually adapting to source subjects.
Muhammad Osama Zeeshan, Marco Pedersoli, Alessandro L. Koerich, Eric Granger
IEEE Trans. Affect. Comput.4
2025 Disentangled Source-Free Personalization for Facial Expression Recognition with Neutral Target Data
abstract
Facial Expression Recognition (FER) from videos is a crucial task in various application areas, such as human-computer interaction and health diagnosis and monitoring (e.g., assessing pain and depression). Beyond the challenges of recognizing subtle emotional or health states, the effectiveness of deep FER models is often hindered by the considerable inter-subject variability in expressions. Source-free (unsupervised) domain adaptation (SFDA) methods may be employed to adapt a pre-trained source model using only unlabeled target domain data, thereby avoiding data privacy, storage, and transmission issues. Typically, SFDA methods adapt to a target domain dataset corresponding to an entire population and assume it includes data from all recognition classes. However, collecting such comprehensive target data can be difficult or even impossible for FER in healthcare applications. In many real-world scenarios, it may be feasible to collect a short neutral control video (which displays only neutral expressions) from target subjects before deployment. These videos can be used to adapt a model to better handle the variability of expressions among subjects. This paper introduces the Disentangled SFDA (DSFDA) method to address the challenge posed by adapting models with missing target expression data. DSFDA leverages data from a neutral target control video for end-to-end generation and adaptation of target data with missing non-neutral data. Our method learns to disentangle features related to expressions and identity while generating the missing non-neutral expression data for the target subject, thereby enhancing model accuracy. Additionally, our self-supervision strategy improves model adaptation by reconstructing target images that maintain the same identity and source expression. Experimental results1on the challenging BioVid, UNBC-McMaster and StressID datasets indicate that our DSFDA approach can outperform state-of-the-art adaptation methods.1https://github.com/MasoumehSharafi/DSFDA/
Masoumeh Sharafi, Emma Ollivier, Muhammad Osama Zeeshan, Soufiane Belharbi, Alessandro L. Koerich, Marco Pedersoli, Simon Bacon, Eric Granger
FG8
2025 Visual Modality Prompt for Adapting Vision-Language Object Detectors
abstract
The zero-shot performance of object detectors degrades when tested on different modalities, such as infrared and depth. While recent work has explored image translation techniques to adapt detectors to new modalities, these methods are limited to a single modality and apply only to traditional detectors. Recently, vision-language detectors, such as YOLO-World and Grounding DINO, have shown promising zero-shot capabilities, however, they have not yet been adapted for other visual modalities. Traditional fine-tuning approaches compromise the zero-shot capabilities of the detectors. The visual prompt strategies commonly used for classification with vision-language models apply the same linear prompt translation to each image, making them less effective. To address these limitations, we propose ModPrompt, a visual prompt strategy to adapt vision-language detectors to new modalities without degrading zero-shot performance. In particular, an encoder-decoder visual prompt strategy is proposed, further enhanced by the integration of inference-friendly modality prompt decoupled residual, facilitating a more robust adaptation. Empirical benchmarking results show our method for modality adaptation on two vision-language detectors, YOLO-World and Grounding DINO, and on challenging infrared (LLVIP, FLIR) and depth (NYUv2) datasets, achieving performance comparable to full fine-tuning while preserving the model's zero-shot capability. Code available at: https://github.com/heitorrapela/ModPrompt.
Heitor Rapela Medeiros, Atif Belal, Srikanth Muralidharan, Eric Granger, Marco Pedersoli
ICCV4
2025 Sparsity Outperforms Low-Rank Projections in Few-Shot Adaptation
abstract
Adapting Vision-Language Models (VLMs) to new domains with few labeled samples remains a significant challenge due to severe overfitting and computational constraints. State-of-the-art solutions, such as low-rank reparameterization, mitigate these issues but often struggle with generalization and require extensive hyperparameter tuning. In this paper, a novel Sparse Optimization (SO) framework is proposed. Unlike low-rank approaches that typically constrain updates to a fixed subspace, our SO method leverages high sparsity to dynamically adjust very few parameters. We introduce two key paradigms. First, we advocate for \textit{local sparsity and global density}, which updates a minimal subset of parameters per iteration while maintaining overall model expressiveness. As a second paradigm, we advocate for \textit{local randomness and global importance}, which sparsifies the gradient using random selection while pruning the first moment based on importance. This combination significantly mitigates overfitting and ensures stable adaptation in low-data regimes. Extensive experiments on 11 diverse datasets show that SO achieves state-of-the-art few-shot adaptation performance while reducing memory overhead.
Nairouz Mrabah, Nicolas Richet, Ismail Ben Ayed, Eric Granger
ICCV4
2025 TD-Paint: Faster Diffusion Inpainting Through Time-Aware Pixel Conditioning
abstract
Diffusion models have emerged as highly effective techniques for inpainting, however, they remain constrained by slow sampling rates. While recent advances have enhanced generation quality, they have also increased sampling time, thereby limiting scalability in real-world applications. We investigate the generative sampling process of diffusion-based inpainting models and observe that these models make minimal use of the input condition during the initial sampling steps. As a result, the sampling trajectory deviates from the data manifold, requiring complex synchronization mechanisms to realign the generation process. To address this, we propose Time-aware Diffusion Paint (TD-Paint), a novel approach that adapts the diffusion process by modeling variable noise levels at the pixel level. This technique allows the model to efficiently use known pixel values from the start, guiding the generation process toward the target manifold. By embedding this information early in the diffusion process, TD-Paint significantly accelerates sampling without compromising image quality. Unlike conventional diffusion-based inpainting models, which require a dedicated architecture or an expensive generation loop, TD-Paint achieves faster sampling times without architectural modifications. Experimental results across three datasets show that TD-Paint outperforms state-of-the-art diffusion models while maintaining lower complexity.
Tsiry Mayet, Pourya Shamsolmoali, Simon Bernard 0001, Eric Granger, Romain Hérault, Clément Chatelain 0001
ICLR4
2025 LT-Soups: Bridging Head and Tail Classes via Subsampled Model Soups
abstract
Real-world datasets typically exhibit long-tailed (LT) distributions, where a few head classes dominate and many tail classes are severely underrepresented. While recent work shows that parameter-efficient fine-tuning (PEFT) methods like LoRA and AdaptFormer preserve tail-class performance on foundation models such as CLIP, we find that they do so at the cost of head-class accuracy. We identify the head-tail ratio, the proportion of head to tail classes, as a crucial but overlooked factor influencing this trade-off. Through controlled experiments on CIFAR100 with varying imbalance ratio ($\rho$) and head-tail ratio ($\eta$), we show that PEFT excels in tail-heavy scenarios but degrades in more balanced and head-heavy distributions. To overcome these limitations, we propose LT-Soups, a two-stage model soups framework designed to generalize across diverse LT regimes. In the first stage, LT-Soups averages models fine-tuned on balanced subsets to reduce head-class bias; in the second, it fine-tunes only the classifier on the full dataset to restore head-class accuracy. Experiments across six benchmark datasets show that LT-Soups achieves superior trade-offs compared to both PEFT and traditional model soups across a wide range of imbalance regimes.
Masih Aminbeidokhti, Subhankar Roy, Eric Granger, Elisa Ricci 0001, Marco Pedersoli
NeurIPS3
2025 Learning Task-Agnostic Representations through Multi-Teacher Distillation
abstract
Casting complex inputs into tractable representations is a critical step across various fields. Diverse embedding models emerge from differences in architectures, loss functions, input modalities and datasets, each capturing unique aspects of the input. Multi-teacher distillation leverages this diversity to enrich representations but often remains tailored to specific tasks. We introduce a task-agnostic framework based on a ``majority vote" objective function. We demonstrate that this function is bounded by the mutual information between the student and the teachers' embeddings, leading to a task-agnostic distillation loss that eliminates dependence on task-specific labels or prior knowledge. Comprehensive evaluations across text, vision models, and molecular modeling show that our method effectively leverages teacher diversity, resulting in representations enabling better performance for a wide range of downstream tasks such as classification, clustering, or regression. Additionally, we train and release state-of-the-art embedding models, enhancing downstream performance in various modalities.
Philippe Formont, Maxime Darrin, Banafsheh Karimian, Eric Granger, Jackie Chi Kit Cheung, Ismail Ben Ayed, Mohammadhadi Shateri, Pablo Piantanida
NeurIPS4
2025 Learning from Stochastic Teacher Representations Using Student-Guided Knowledge Distillation
Muhammad Haseeb Aslam, Clara Martinez, Marco Pedersoli, Alessandro L. Koerich, Ali Etemad, Eric Granger
ECML/PKDD (6)6
2025 Bidirectional Multi-Step Domain Generalization for Visible-Infrared Person Re-Identification
abstract
A key challenge in visible-infrared person re-identification (V-I ReID) is training a backbone model capable of effectively addressing the significant discrepancies across modalities. State-of-the-art methods that generate a single intermediate bridging domain are often less effective, as this generated domain may not adequately capture sufficient common discriminant information. This paper introduces Bidirectional Multi-step Domain Generalization (BMDG), a novel approach for unifying feature representations across diverse modalities. BMDG creates multiple virtual intermediate domains by learning and aligning body part features extracted from both I and V modalities. In particular, our method aims to minimize the cross-modal gap in two steps. First, BMDG aligns modalities in the feature space by learning shared and modality-invariant body part prototypes from V and I images. Then, it generalizes the feature representation by applying bidirectional multi-step learning, which progressively refines feature representations in each step and incorporates more prototypes from both modalities. Based on these prototypes, multiple bridging steps enhance the feature representation. Experiments11Our code is available at: alehdaghi.github.io/BMDG conducted on V-I ReID datasets indicate that our BMDG approach can outperform state-of-the-art part-based and intermediate generation methods, and can be integrated into other part-based methods to enhance their V-I ReID performance.
Mahdi Alehdaghi, Pourya Shamsolmoali, Rafael M. O. Cruz, Eric Granger
WACV4
2025 Attention-Based Class-Conditioned Alignment for Multi-Source Domain Adaptation of Object Detectors
abstract
Domain adaptation methods for object detection (OD) strive to mitigate the impact of distribution shifts by promoting feature alignment across source and target domains. Multi-source domain adaptation (MSDA) allows leveraging multiple annotated source datasets and unlabeled target data to improve the accuracy and robustness of the detection model. Most state-of-the-art MSDA methods for OD perform feature alignment in a class-agnostic manner. This is challenging since the objects have unique modality information due to variations in object appearance across domains. A recent prototype-based approach proposed a class-wise alignment, yet it suffers from error accumulation caused by noisy pseudo-labels that can negatively affect adaptation with imbalanced data. To overcome these limitations, we propose an attention-based class-conditioned alignment method for MSDA, designed to align instances of each object category across domains. In particular, an attention module combined with an adversarial domain classifier allows learning domain-invariant and class-specific instance representations. Experimental results on multiple benchmarking MSDA datasets indicate that our method outperforms state-of-the-art methods and exhibits robustness to class imbalance, achieved through a conceptually simple class-conditioning strategy. Our code is available at: https://github.com/imatif17/ACIA.
Atif Belal, Akhil Meethal, Francisco Perdigón Romero, Marco Pedersoli, Eric Granger
WACV5
2025 Mixed Patch Visible-Infrared Modality Agnostic Object Detection
abstract
In real-world scenarios, using multiple modalities like visible (RGB) and infrared (IR) can greatly improve the performance of a predictive task such as object detection (OD). Multimodal learning is a common way to leverage these modalities, where multiple modality-specific encoders and a fusion module are used to improve performance. In this paper, we tackle a different way to employ RGB and IR modalities, where only one modality or the other is observed by a single shared vision encoder. This realistic setting requires a lower memory footprint and is more suitable for applications such as autonomous driving and surveillance, which commonly rely on RGB and IR data. However, when learning a single encoder on multiple modalities, one modality can dominate the other, producing un-even recognition results. This work investigates how to efficiently leverage RGB and IR modalities to train a common transformer-based OD vision encoder while countering the effects of modality imbalance. For this, we introduce a novel training technique to Mix Patches (MiPa)from the two modalities, in conjunction with a patch-wise modality agnostic module, for learning a common representation of both modalities. Our experiments show that MiPa can learn a representation to reach competitive results on traditional RGB/IR benchmarks while only requiring a single modality during inference. Our code is available at: https://github.com/heitorrapela/MiPa.
Heitor Rapela Medeiros, David Latortue, Eric Granger, Marco Pedersoli
WACV3
2025 A Realistic Protocol for Evaluation of Weakly Supervised Object Localization
abstract
Weakly Supervised Object Localization (WSOL) allows training deep learning models for classification and localization (LOC) using only global class-level labels. The absence of bounding box (bbox) supervision during training raises challenges in the literature for hyper-parameter tuning, model selection, and evaluation. WSOL methods rely on a validation set with bbox annotations for model selection, and a test set with bbox annotations for threshold estimation for producing bboxes from localization maps. This approach, however, is not aligned with the WSOL setting as these annotations are typically unavailable in real-world scenarios. Our initial empirical analysis shows a significant decline in LOC performance when model selection and threshold estimation rely solely on class labels and the image itself, respectively, compared to using manual bbox annotations. This highlights the importance of incorporating bbox labels for optimal model performance. In this paper11Our scope. This work focuses on WSOL [2],[28]–[30], as opposed to other related tasks such as weakly supervised detection, segmentation, or instance segmentation, which are often mixed in earlier works [11], [18], [40]., a new WSOL evaluation protocol is proposed that provides LOC information without the need for manual bbox annotations. In particular, we generated noisy pseudo-boxes from a pretrained off-the-shelf region proposal method such as Selective Search, CLIP, and RPN for model selection. These bboxes are also employed to estimate the threshold from LOC maps, circumventing the need for test-set bbox annotations. Our experiments22Our code and generated pseudo-bounding boxes can be accessed at github.com/shakeebmurtaza/wsol_model_selection. with several WSOL methods on challenging natural and medical image datasets show that using the proposed pseudo-bboxes for validation facilitates the model selection and threshold estimation, with LOC performance comparable to models selected using GT bboxes on the validation set and threshold estimation on the test set. It also outperforms models selected using class-level labels, and then dynamically thresholded with only LOC maps.
Shakeeb Murtaza, Soufiane Belharbi, Marco Pedersoli, Eric Granger
WACV4
2025 Fusion for Visual-Infrared Person ReID in Real-World Surveillance Using Corrupted Multimodal Data
Arthur Josi, Mahdi Alehdaghi, Rafael M. O. Cruz, Eric Granger
Int. J. Comput. Vis.4
2025 CoLo-CAM: Class activation mapping for object co-localization in weakly-labeled unconstrained videos
abstract
Leveraging spatiotemporal information in videos is critical for weakly supervised video object localization (WSVOL) tasks. However, state-of-the-art methods only rely on visual and motion cues, while discarding discriminative information, making them susceptible to inaccurate localizations. Recently, discriminative models have been explored for WSVOL tasks using a temporal class activation mapping (CAM) method. Although their results are promising, objects are assumed to have limited movement from frame to frame, leading to degradation in performance for relatively long-term dependencies. This paper proposes a novel CAM method for WSVOL that exploits spatiotemporal information in activation maps during training without constraining an object’s position. Its training relies on Co - Lo calization, hence, the name CoLo-CAM . Given a sequence of frames, localization is jointly learned based on color cues extracted across the corresponding maps, by assuming that an object has similar color in consecutive frames. CAM activations are constrained to respond similarly over pixels with similar colors, achieving co-localization. This improves localization performance because the joint learning creates direct communication among pixels across all image locations and over all frames, allowing for transfer, aggregation, and correction of localizations. Co-localization is integrated into training by minimizing the color term of a conditional random field (CRF) loss over a sequence of frames/CAMs. Extensive experiments 1 on two challenging YouTube-Objects datasets of unconstrained videos show the merits of our CoLo-CAM method, and its robustness to long-term dependencies, leading to new state-of-the-art performance for WSVOL task.
Soufiane Belharbi, Shakeeb Murtaza, Marco Pedersoli, Ismail Ben Ayed, Luke McCaffrey, Eric Granger
Pattern Recognit.6
2025 Adaptive Generation of Privileged Intermediate Information for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (V-I ReID) seeks to retrieve images of the same individual captured over a distributed network of RGB and IR sensors. Several V-I ReID approaches directly integrate the V and I modalities to represent images within a shared space. However, given the significant gap in the data distributions between V and I modalities, cross-modal V-I ReID remains challenging. A solution is to involve a privileged intermediate space to bridge between modalities, but in practice, such data is not available and requires selecting or creating effective mechanisms for informative intermediate domains. This paper introduces the Adaptive Generation of Privileged Intermediate Information (AGPI2) training approach to adapt and generate a virtual domain that bridges discriminative information between the V and I modalities. AGPI2enhances the training of a deep V-I ReID backbone by generating and then leveraging bridging privileged information without modifying the model in the inference phase. This information captures shared discriminative attributes that are not easily ascertainable for the model within individual V or I modalities. Towards this goal, a non-linear generative module is trained with adversarial objectives, transforming V attributes into intermediate spaces that also contain I features. This domain exhibits less domain shift relative to the I domain compared to the V domain. Meanwhile, the embedding module within AGPI2aims to extract discriminative modality-invariant features for both modalities by leveraging modality-free descriptors from generated images, making them a bridge between the main modalities. Experiments conducted on challenging V-I ReID datasets indicate that AGPI2consistently increases matching accuracy without additional computational resources during inference.
Mahdi Alehdaghi, Arthur Josi, Rafael M. O. Cruz, Pourya Shamsolmoali, Eric Granger
IEEE Trans. Inf. Forensics Secur.5
2025 Hybrid Gromov-Wasserstein Embedding for Capsule Learning
abstract
Capsule networks (CapsNets) aim to parse images into a hierarchy of objects, parts, and their relationships using a two-step process involving part-whole transformation and hierarchical component routing. However, this hierarchical relationship modeling is computationally expensive, which has limited the wider use of CapsNet despite its potential advantages. The current state of CapsNet models primarily focuses on comparing their performance with capsule baselines, falling short of achieving the same level of proficiency as deep convolutional neural network (CNN) variants in intricate tasks. To address this limitation, we present an efficient approach for learning capsules that surpasses canonical baseline models and even demonstrates superior performance compared with high-performing convolution models. Our contribution can be outlined in two aspects: first, we introduce a group of subcapsules onto which an input vector is projected. Subsequently, we present the hybrid Gromov-Wasserstein (HGW) framework, which initially quantifies the dissimilarity between the input and the components modeled by the subcapsules, followed by determining their alignment degree through optimal transport (OT). This innovative mechanism capitalizes on new insights into defining alignment between the input and subcapsules, based on the similarity of their respective component distributions. This approach enhances CapsNets' capacity to learn from intricate, high-dimensional data while retaining their interpretability and hierarchical structure. Our proposed model offers two distinct advantages: 1) its lightweight nature facilitates the application of capsules to more intricate vision tasks, including object detection; and 2) it outperforms baseline approaches in these demanding tasks. Our empirical findings illustrate that HGW capsules (HGWCapsules) exhibit enhanced robustness against affine transformations, scale effectively to larger datasets, and surpass CNN and CapsNet models across various vision tasks.
Pourya Shamsolmoali, Masoumeh Zareapoor, Swagatam Das, Eric Granger, Salvador García 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 SeTformer Is What You Need for Vision and Language
abstract
The dot product self-attention (DPSA) is a fundamental component of transformers. However, scaling them to long sequences, like documents or high-resolution images, becomes prohibitively expensive due to the quadratic time and memory complexities arising from the softmax operation. Kernel methods are employed to simplify computations by approximating softmax but often lead to performance drops compared to softmax attention. We propose SeTformer, a novel transformer where DPSA is purely replaced by Self-optimal Transport (SeT) for achieving better performance and computational efficiency. SeT is based on two essential softmax properties: maintaining a non-negative attention matrix and using a nonlinear reweighting mechanism to emphasize important tokens in input sequences. By introducing a kernel cost function for optimal transport, SeTformer effectively satisfies these properties. In particular, with small and base-sized models, SeTformer achieves impressive top-1 accuracies of 84.7% and 86.2% on ImageNet-1K. In object detection, SeTformer-base outperforms the FocalNet counterpart by +2.2 mAP, using 38% fewer parameters and 29% fewer FLOPs. In semantic segmentation, our base-size model surpasses NAT by +3.5 mIoU with 33% fewer parameters. SeTformer also achieves state-of-the-art results in language modeling on the GLUE benchmark. These findings highlight SeTformer applicability for vision and language tasks.
Pourya Shamsolmoali, Masoumeh Zareapoor, Eric Granger, Michael Felsberg
AAAI3
2024 Modality Translation for Object Detection Adaptation Without Forgetting Prior Knowledge
Heitor Rapela Medeiros, Masih Aminbeidokhti, Fidel A. Guerrero-Peña, David Latortue, Eric Granger, Marco Pedersoli
ECCV (89)5
2024 Distilling Privileged Multimodal Information for Expression Recognition using Optimal Transport
abstract
Deep learning models for multimodal expression recognition have reached remarkable performance in controlled laboratory environments because of their ability to learn complementary and redundant semantic information. However, these models struggle in the wild, mainly because of the unavailability and quality of modalities used for training. In practice, only a subset of the training-time modalities may be available at test time. Learning with privileged information enables models to exploit data from additional modalities that are only available during training. State-of-the-art knowledge distillation (KD) methods have been proposed to distill information from multiple teacher models (each trained on a modality) to a common student model. These privileged KD methods typically utilize point-to-point matching, yet have no explicit mechanism to capture the structural information in the teacher representation space formed by introducing the privileged modality. We argue that encoding this same structure in the student space may lead to enhanced student performance. This paper introduces a new structural KD mechanism based on optimal transport (OT), where entropy-regularized OT distills the structural dark knowledge. Our privileged KD with OT (PKDOT) method captures the local structures in the multimodal teacher representation by calculating a cosine similarity matrix and selecting the top-k anchors to allow for sparse OT solutions, resulting in a more stable distillation process. Experiments1were performed on two challenging problems - pain estimation on the Biovid dataset (ordinal classification) and arousal-valance prediction on the Affwild2 dataset (regression). Results show that our proposed method can outperform state-of-the-art privileged KD methods on these problems. The diversity among modalities and fusion architectures indicates that PKDOT is modality-and model-agnostic.
Muhammad Haseeb Aslam, Muhammad Osama Zeeshan, Soufiane Belharbi, Marco Pedersoli, Alessandro L. Koerich, Simon Bacon, Eric Granger
FG7
2024 Guided Interpretable Facial Expression Recognition via Spatial Action Unit Cues
abstract
Although state-of-the-art classifiers for facial expression recognition (FER) can achieve a high level of accuracy, they lack interpretability, an important feature for end-users. Experts typically associate spatial action units (AUs) from a codebook to facial regions for the visual interpretation of expressions. In this paper, the same expert steps are followed. A new learning strategy is proposed to explicitly incorporate AU cues into classifier training, allowing to train deep interpretable models. During training, this AU codebook is used, along with the input image expression label, and facial landmarks, to construct a AU heatmap that indicates the most discriminative image regions of interest w.r.t the facial expression. This valuable spatial cue is leveraged to train a deep interpretable classifier for FER. This is achieved by constraining the spatial layer features of a classifier to be correlated with AU heatmaps. Using a composite loss, the classifier is trained to correctly classify an image while yielding interpretable visual layer-wise attention correlated with AU maps, simulating the expert decision process. Our strategy only relies on image class expression for supervision, without additional manual annotations. Our new strategy is generic, and can be applied to any deep CNN - or transformer-based classifier without requiring any architectural change or significant additional training time. Our extensive evaluation11Our code is available at:https://github.com/sbelharbi/interpretable-fer-aus. on two public benchmarks RAF-DB, and AffectNet datasets shows that our proposed strategy can improve layer-wise interpretability without degrading classification performance. In addition, we explore a common type of interpretable classifiers that rely on class activation mapping (CAM) methods, and show that our approach can also improve CAM interpretability.
Soufiane Belharbi, Marco Pedersoli, Alessandro L. Koerich, Simon Bacon, Eric Granger
FG5
2024 Subject-Based Domain Adaptation for Facial Expression Recognition
abstract
Adapting a deep learning model to a specific target individual is a challenging facial expression recognition (FER) task that may be achieved using unsupervised domain adaptation (UDA) methods. Although several UDA methods have been proposed to adapt deep FER models across source and target data sets, multiple subject-specific source domains are needed to accurately represent the intra-and inter-person variability in subject-based adaption. This paper considers the setting where domains correspond to individuals, not entire datasets. Unlike UDA, multi-source domain adaptation (MSDA) methods can leverage multiple source datasets to improve the accuracy and robustness of the target model. However, previous methods for MSDA adapt image classification models across datasets and do not scale well to a more significant number of source domains. This paper introduces a new MSDA method for subject-based domain adaptation in FER. It efficiently leverages information from multiple source subjects (labeled source domain data) to adapt a deep FER model to a single target individual (unlabeled target domain data). During adaptation, our subject-based MSDA first computes a between-source discrepancy loss to mitigate the domain shift among data from several source subjects. Then, a new strategy is employed to generate augmented confident pseudo-labels for the target subject, allowing a reduction in the domain shift between source and target subjects. Experiments1performed on the challenging BioVid heat and pain dataset with 87 subjects and the UNBC-McMaster shoulder pain dataset with 25 subjects show that our subject-based MSDA can outperform state-of-the-art methods yet scale well to multiple subject-based source domains.
Muhammad Osama Zeeshan, Muhammad Haseeb Aslam, Soufiane Belharbi, Alessandro L. Koerich, Marco Pedersoli, Simon Bacon, Eric Granger
FG7
2024 SR-CACO-2: A Dataset for Confocal Fluorescence Microscopy Image Super-Resolution
abstract
Confocal fluorescence microscopy is one of the most accessible and widely used imaging techniques for the study of biological processes at the cellular and subcellular levels. Scanning confocal microscopy allows the capture of high-quality images from thick three-dimensional (3D) samples, yet suffers from well-known limitations such as photobleaching and phototoxicity of specimens caused by intense light exposure, which limits its use in some applications, especially for living cells. Cellular damage can be alleviated by changing imaging parameters to reduce light exposure, often at the expense of image quality.Machine/deep learning methods for single-image super-resolution (SISR) can be applied to restore image quality by upscaling lower-resolution (LR) images to produce high-resolution images (HR). These SISR methods have been successfully applied to photo-realistic images due partly to the abundance of publicly available datasets. In contrast, the lack of publicly available data partly limits their application and success in scanning confocal microscopy.In this paper, we introduce a large scanning confocal microscopy dataset named SR-CACO-2 that is comprised of low- and high-resolution image pairs marked for three different fluorescent markers. It allows to evaluate the performance of SISR methods on three different upscaling levels (X2, X4, X8). SR-CACO-2 contains the human epithelial cell line Caco-2 (ATCC HTB-37), and it is composed of 2,200 unique images, captured with four resolutions and three markers, that have been translated in the form of 9,937 patches for experiments with SISR methods. Given the new SR-CACO-2 dataset, we also provide benchmarking results for 16 state-of-the-art methods that are representative of the main SISR families. Results show that these methods have limited success in producing high-resolution textures, indicating that SR-CACO-2 represents a challenging problem. The dataset is released under a Creative Commons license (CC BY-NC-SA 4.0), and it can be accessed freely. Our dataset, code and pretrained weights for SISR methods are publicly available: https://github.com/sbelharbi/sr-caco-2.
Soufiane Belharbi, Mara KM Whitford, Phuong Hoang, Shakeeb Murtaza, Luke McCaffrey, Eric Granger
NeurIPS6
2024 Domain Generalization by Rejecting Extreme Augmentations
abstract
Data augmentation is one of the most effective techniques for regularizing deep learning models and improving recognition performance in various tasks and domains. However, this holds for standard in-domain settings, in which the training and test data follow the same distribution. For the out-of-domain case, where the test data follow a different and unknown distribution, the best recipe for data augmentation is unclear. In this paper, we show that for out-of-domain and domain generalization settings, data augmentation can provide a conspicuous and robust improvement in performance. To do that, we propose a simple training procedure: (i) use uniform sampling on standard data augmentation transformations; (ii) increase the strength transformations to account for the higher data variance expected when working out-of-domain, and (iii) devise a new reward function to reject extreme transformations that can harm the training. With this procedure, our data augmentation scheme achieves a level of accuracy comparable to or better than state-of-the-art methods on benchmark domain generalization datasets. Code: https://github.com/Masseeh/DCAug
Masih Aminbeidokhti, Fidel A. Guerrero-Peña, Heitor Rapela Medeiros, Thomas Dubail, Eric Granger, Marco Pedersoli
WACV5
2024 Multi-Source Domain Adaptation for Object Detection with Prototype-based Mean Teacher
abstract
Adapting visual object detectors to operational target domains is a challenging task, commonly achieved using unsupervised domain adaptation (UDA) methods. Recent studies have shown that when the labeled dataset comes from multiple source domains, treating them as separate domains and performing a multi-source domain adaptation (MSDA) improves the accuracy and robustness over blending these source domains and performing a UDA. For adaptation, existing MSDA methods learn domain-invariant and domain-specific parameters (for each source domain). However, unlike single-source UDA methods, learning domain-specific parameters makes them grow significantly in proportion to the number of source domains. This paper proposes a novel MSDA method called Prototype-based Mean Teacher (PMT), which uses class prototypes instead of domain-specific subnets to encode domain-specific information. These prototypes are learned using a contrastive loss, aligning the same categories across domains and separating different categories far apart. Given the use of prototypes, the number of parameters required for our PMT method does not increase significantly with the number of source domains, thus reducing memory issues and possible overfitting. Empirical studies indicate that PMT outperforms state-of-the-art MSDA methods on several challenging object detection datasets. Our code is available at https://github.com/imatif17/Prototype-Mean-Teacher
Atif Belal, Akhil Meethal, Francisco Perdigón Romero, Marco Pedersoli, Eric Granger
WACV5
2024 HalluciDet: Hallucinating RGB Modality for Person Detection Through Privileged Information
abstract
A powerful way to adapt a visual recognition model to a new domain is through image translation. However, common image translation approaches only focus on generating data from the same distribution as the target domain. Given a cross-modal application, such as pedestrian detection from aerial images, with a considerable shift in data distribution between infrared (IR) to visible (RGB) images, a translation focused on generation might lead to poor performance as the loss focuses on irrelevant details for the task. In this paper, we propose HalluciDet, an IR-RGB image translation model for object detection. Instead of focusing on reconstructing the original image on the IR modality, it seeks to reduce the detection loss of an RGB detector, and therefore avoids the need to access RGB data. This model produces a new image representation that enhances objects of interest in the scene and greatly improves detection performance. We empirically compare our approach against state-of-the-art methods for image translation and for fine-tuning on IR, and show that our HalluciDet improves detection accuracy in most cases by exploiting the privileged information encoded in a pre-trained RGB detector. Code: https://github.com/heitorrapela/HalluciDet.
Heitor Rapela Medeiros, Fidel A. Guerrero-Peña, Masih Aminbeidokhti, Thomas Dubail, Eric Granger, Marco Pedersoli
WACV5
2024 Facial expression analysis using Decomposed Multiscale Spatiotemporal Networks
Wheidima C. Melo, Eric Granger, Miguel Bordallo López
Expert Syst. Appl.2
2024 Special section: Best papers of the international conference on pattern recognition and artificial intelligence (ICPRAI) 2022
Mounim A. El-Yacoubi, Umapada Pal 0001, Eric Granger, Pong C. Yuen
Pattern Recognit. Lett.3
2023 Re-basin via implicit Sinkhorn differentiation
abstract
The recent emergence of new algorithms for permuting models into functionally equivalent regions of the solution space has shed some light on the complexity of error surfaces and some promising properties like mode connectivity. However, finding the permutation that minimizes some ob-jectives is challenging, and current optimization techniques are not differentiable, which makes it difficult to integrate into a gradient-based optimization, and often leads to sub-optimal solutions. In this paper, we propose a Sinkhorn re-basin network with the ability to obtain the transportation plan that better suits a given objective. Unlike the current state-of-art, our method is differentiable and, there-fore, easy to adapt to any task within the deep learning do-main. Furthermore, we show the advantage of our re-basin method by proposing a new cost function that allows per-forming incremental learning by exploiting the linear mode connectivity property. The benefit of our method is compared against similar approaches from the literature under several conditions for both optimal transport and linear mode connectivity. The effectiveness of our continual learning method based on re-basin is also shown for several common benchmark datasets, providing experimental results that are competitive with the state-of-art. The source code is provided at https://github.com/jagp/sinkhorn-rebasin.
Fidel A. Guerrero-Peña, Heitor Rapela Medeiros, Thomas Dubail, Masih Aminbeidokhti, Eric Granger, Marco Pedersoli
CVPR5
2023 Recursive Joint Attention for Audio-Visual Fusion in Regression Based Emotion Recognition
abstract
In video-based emotion recognition (ER), it is important to effectively leverage the complementary relationship among audio (A) and visual (V) modalities, while retaining the intramodal characteristics of individual modalities. In this paper, a recursive joint attention model is proposed along with long short-term memory (LSTM) modules for the fusion of vocal and facial expressions in regression-based ER. Specifically, we investigated the possibility of exploiting the complementary nature of A and V modalities using a joint cross-attention model in a recursive fashion with LSTMs to capture the intramodal temporal dependencies within the same modalities as well as among the A-V feature representations. By integrating LSTMs with recursive joint cross-attention, our model can efficiently leverage both intra- and inter-modal relationships for the fusion of A and V modalities. The results of extensive experiments1performed on the challenging Affwild2 and Fatigue (private) datasets indicate that the proposed A-V fusion model can significantly outperform state-of-art-methods.
Gnana Praveen Rajasekhar, Eric Granger, Patrick Cardinal
ICASSP2
2023 Image Completion Via Dual-Path Cooperative Filtering
abstract
Given the recent advances with image-generating algorithms, deep image completion methods have made significant progress. However, state-of-art methods typically provide poor cross-scene generalization, and generated masked areas often contain blurry artifacts. Predictive filtering is a method for restoring images, which predicts the most effective kernels based on the input scene. Motivated by this approach, we address image completion as a filtering problem. Deep feature-level semantic filtering is introduced to fill in missing information, while preserving local structure and generating visually realistic content. In particular, a Dual-path Cooperative Filtering (DCF) model is proposed, where one path predicts dynamic kernels, and the other path extracts multi-level features by using Fast Fourier Convolution to yield semantically coherent reconstructions. Experiments on three challenging image completion datasets show that our proposed DCF outperforms state-of-art methods.
Pourya Shamsolmoali, Masoumeh Zareapoor, Eric Granger
ICASSP3
2023 TCAM: Temporal Class Activation Maps for Object Localization in Weakly-Labeled Unconstrained Videos
abstract
Weakly supervised video object localization (WSVOL) allows locating object in videos using only global video tags such as object classes. State-of-art methods rely on multiple independent stages, where initial spatio-temporal proposals are generated using visual and motion cues, and then prominent objects are identified and refined. The localization involves solving an optimization problem over one or more videos, and video tags are typically used for video clustering. This process requires a model per video or per class making for costly inference. Moreover, localized regions are not necessary discriminant because these methods rely on unsupervised motion methods like optical flow, or discarded video tags from optimization. In this paper, we leverage the successful class activation mapping (CAM) methods, designed for WSOL based on still images. A new Temporal CAM (TCAM) method is introduced for training a discriminant deep learning (DL) model to exploit spatio-temporal information in videos, using an CAM-Temporal Max Pooling (CAM-TMP) aggregation mechanism over consecutive CAMs. In particular, activations of regions of interest (ROIs) are collected from CAMs produced by a pretrained CNN classifier, and generate pixel-wise pseudo-labels for training a decoder. In addition, a global unsupervised size constraint, and local constraint such as CRF are used to yield more accurate CAMs. Inference over single independent frames allows parallel processing of a clip of frames, and real-time localization. Extensive experiments1on two challenging YouTube-Objects datasets with unconstrained videos indicate that CAM methods (trained on independent frames) can yield decent localization accuracy. Our proposed TCAM method achieves a new state-of-art in WSVOL accuracy, and visual results suggest that it can be adapted for subsequent tasks, such as object detection and tracking.
Soufiane Belharbi, Ismail Ben Ayed, Luke McCaffrey, Eric Granger
WACV4
2023 Camera Alignment and Weighted Contrastive Learning for Domain Adaptation in Video Person ReID
abstract
Systems for person re-identification (ReID) can achieve a high accuracy when trained on large fully-labeled image datasets. However, the domain shift typically associated with diverse operational capture conditions (e.g., camera viewpoints and lighting) may translate to a significant decline in performance. This paper focuses on unsupervised domain adaptation (UDA) for video-based ReID – a relevant scenario that is less explored in the literature. In this scenario, the ReID model must adapt to a complex target domain defined by a network of diverse video cameras based on track-let information. State-of-art methods cluster unlabeled target data, yet domain shifts across target cameras (sub-domains) can lead to poor initialization of clustering methods that propagates noise across epochs, thus preventing the ReID model to accurately associate samples of same identity. In this paper, an UDA method is introduced for video person ReID that leverages knowledge on video tracklets, and on the distribution of frames captured over target cameras to improve the performance of CNN backbones trained using pseudo-labels. Our method relies on an adversarial approach, where a camera-discriminator network is introduced to extract discriminant camera-independent representations, facilitating the subsequent clustering. In addition, a weighted contrastive loss is proposed to leverage the confidence of clusters, and mitigate the risk of incorrect identity associations. Experimental results obtained on three challenging video-based person ReID datasets – PRID2011, iLIDS-VID, and MARS – indicate that our proposed method can outperform related state-of-the-art methods. Our code is available at: https://github.com/dmekhazni/CAWCL-ReID
Djebril Mekhazni, Maximilien Dufau, Christian Desrosiers, Marco Pedersoli, Eric Granger
WACV5
2023 TransVLAD: Multi-Scale Attention-Based Global Descriptors for Visual Geo-Localization
abstract
Visual geo-localization remains a challenging task due to variations in the appearance and perspective among captured images. This paper introduces an efficient TransVLAD module, which aggregates attention-based feature maps into a discriminative and compact global descriptor. Unlike existing methods that generate feature maps using only convolutional neural networks (CNNs), we propose a sparse transformer to encode global dependencies and compute attention-based feature maps, which effectively reduces visual ambiguities that occurs in large-scale geo-localization problems. A positional embedding mechanism is used to learn the corresponding geometric configurations between query and gallery images. A grouped VLAD layer is also introduced to reduce the number of parameters, and thus construct an efficient module. Finally, rather than only learning from the global descriptors on entire images, we propose a self-supervised learning method to further encode more information from multi-scale patches between the query and positive gallery images. Extensive experiments on three challenging large-scale datasets indicate that our model outperforms state-of-the-art models, and has lower computational complexity. The code is available at: https://github.com/wacv-23/TVLAD.
Yifan Xu 0034, Pourya Shamsolmoali, Eric Granger, Claire Nicodeme, Laurent Gardes, Jie Yang 0002
WACV3
2023 DiPS: Discriminative pseudo-label sampling with self-supervised transformers for weakly supervised object localization
Shakeeb Murtaza, Soufiane Belharbi, Marco Pedersoli, Aydin Sarraf, Eric Granger
Image Vis. Comput.5
2023 MDN: A Deep Maximization-Differentiation Network for Spatio-Temporal Depression Detection
abstract
Deep learning (DL) models have been successfully applied in video-based affective computing, allowing, for instance, to recognize emotions and mood, or to estimate the intensity of pain or stress of individuals based on their facial expressions. Despite the recent advances with state-of-the-art DL models for spatio-temporal recognition of facial expressions associated with depressive behaviour, some key challenges remain in the cost-effective application of 3D-CNNs: (1) 3D convolutions usually employ structures with fixed temporal depth that decreases the potential to extract discriminative representations due to the usually small difference of spatio-temporal variations along different depression levels; and (2) the computational complexity of these models with consequent susceptibility to overfitting. To address these challenges, we propose a novel DL architecture called the Maximization and Differentiation Network (MDN) in order to effectively represent facial expression variations that are relevant for depression assessment. The MDN, operating without 3D convolutions, explores multiscale temporal information using a maximization block that captures smooth facial variations and a difference block that encodes sudden facial variations. Extensive experiments using our proposed MDN with models with 100 and 152 layers result in improved performance while reducing the number of parameters by more than$3\times$when compared with 3D ResNet models. Our model also outperforms other 3D models and achieves state-of-the-art results for depression detection. Code available at:https://github.com/wheidima/MDN.
Wheidima C. Melo, Eric Granger, Miguel Bordallo López
IEEE Trans. Affect. Comput.2
2023 GEN: Generative Equivariant Networks for Diverse Image-to-Image Translation
abstract
Image-to-image (I2I) translation has become a key asset for generative adversarial networks. Convolutional neural networks (CNNs), despite having a significant performance, are not able to capture the spatial relationships among different parts of an object and, thus, do not qualify as the ideal representative model for image translation tasks. As a remedy to this problem, capsule networks have been proposed to represent patterns for a visual object in such a way that preserves hierarchical spatial relationships. The training of capsules is constrained by learning all pairwise relationships between capsules of consecutive layers. This design would be prohibitively expensive both in time and memory. In this article, we present a new framework for capsule networks to provide a full description of the input components at various levels of semantics, which can successfully be applied to the generator-discriminator architectures without incurring computational overhead compared to the CNNs. To successfully apply the proposed capsules in the generative adversarial network, we put forth a novel Gromov-Wasserstein (GW) distance as a differentiable loss function that compares the dissimilarity between two distributions and then guides the learned distribution toward target properties, using optimal transport (OT) discrepancy. The proposed method-which is called generative equivariant network (GEN)-is an alternative architecture for GANs with equivariance capsule layers. The proposed model is evaluated through a comprehensive set of experiments on I2I translation and image generation tasks and compared with several state-of-the-art models. Results indicate that there is a principled connection between generative and capsule models that allows extracting discriminant and invariant information from image data.
Pourya Shamsolmoali, Masoumeh Zareapoor, Swagatam Das, Salvador García 0001, Eric Granger, Jie Yang 0002
IEEE Trans. Cybern.5
2022 Knowledge Distillation for Multi-Target Domain Adaptation in Real-Time Person Re-Identification
abstract
Despite the recent success of deep learning architectures, person re-identification (ReID) remains a challenging problem in real-word applications. Several unsupervised single-target domain adaptation (STDA) methods have recently been proposed to limit the decline in ReID accuracy caused by the domain shift that typically occurs between source and target video data. Given the multimodal nature of person ReID data (due to variations across camera viewpoints and capture conditions), training a common CNN backbone to address domain shifts across multiple target domains, can provide an efficient solution for real-time ReID applications. Although multi-target domain adaptation (MTDA) has not been widely addressed in the ReID literature, a straightforward approach consists in blending different target datasets, and performing STDA on the mixture to train a common CNN. However, this approach may lead to poor generalization, especially when blending a growing number of distinct target domains to train a smaller CNN. To alleviate this problem, we introduce a new MTDA method based on knowledge distillation (KD-ReID) that is suitable for real-time person ReID applications. Our method adapts a common lightweight student backbone CNN over the target domains by alternatively distilling from multiple specialized teacher CNNs, each one adapted on data from a specific target domain. Extensive experiments1conducted on several challenging person ReID datasets indicate that our approach outperforms state-of-art methods for MTDA, including blending methods, particularly when training a compact CNN backbone like OSNet. Results suggest that our flexible MTDA approach can be employed to design cost-effective ReID systems for real-time video surveillance applications.
Félix Remigereau, Djebril Mekhazni, Sajjad Abdoli, Le Thanh Nguyen-Meidine, Rafael M. O. Cruz, Eric Granger
ICIP6
2022 Salient Skin Lesion Segmentation via Dilated Scale-Wise Feature Fusion Network
abstract
Skin lesion detection in dermoscopic images is essential in the accurate and early diagnosis of skin cancer by a computerized apparatus. Current skin lesion segmentation approaches show poor performance in challenging circumstances such as indistinct lesion boundaries, low contrast between the lesion and the surrounding area, or heterogeneous background that causes over/under segmentation of the skin lesion. To accurately recognize the lesion from the neighboring regions, we propose a dilated scale-wise feature fusion network based on convolution factorization. Our network is designed to simultaneously extract features at different scales which are systematically fused for better detection. The proposed model has satisfactory accuracy and efficiency. Various experiments for lesion segmentation are performed along with comparisons with the state-of-the-art models. Our proposed model consistently showcases state-of-the-art results.
Pourya Shamsolmoali, Masoumeh Zareapoor, Jie Yang 0002, Eric Granger, Huiyu Zhou 0001
ICPR4
2022 Enhanced Single-Shot Detector for Small Object Detection in Remote Sensing Images
abstract
Small-object detection is a challenging problem. In the last few years, the convolution neural networks methods have been achieved considerable progress. However, the current detectors struggle with effective features extraction for small-scale objects. To address this challenge, we propose image pyramid single-shot detector (IPSSD). In IPSSD, single-shot detector is adopted combined with an image pyramid network to extract semantically strong features for generating candidate regions. The proposed network can enhance the small-scale features from a feature pyramid network. We evaluated the performance of the proposed model on two public datasets and the results show the superior performance of our model compared to the other state-of-the-art object detectors.
Pourya Shamsolmoali, Masoumeh Zareapoor, Jie Yang 0002, Eric Granger, Jocelyn Chanussot
IGARSS4
2022 Dynamic Template Selection Through Change Detection for Adaptive Siamese Tracking
abstract
Deep Siamese trackers have recently gained much attention in recent years since they can track visual objects at high speed. Additionally, adaptive tracking methods, where target samples collected by the tracker are employed for online learning, have achieved state-of-the-art accuracy. However, single object tracking (SOT) remains a challenging task in real-world application due to changes and deformations in a target object's appearance. Learning on all the collected samples may lead to catastrophic forgetting, and thereby corrupt the tracking model. In this paper, SOT is formulated as an online incremental learning problem. A new method is proposed for dynamic sample selection and memory replay, preventing template corruption. In particular, we propose a change detection mechanism to detect gradual changes in object appearance, and select the corresponding samples for online adaption. In addition, an entropy-based sample selection strategy is introduced to maintain a diversified auxiliary buffer for memory replay. Our proposed method can be integrated into any object tracking algorithm that leverages online learning for model adaptation. Extensive experiments conducted on the OTB-100, LaSOT, UAV123, and TrackingNet datasets highlight the cost-effectiveness of our method, along with the contribution of its key components. Results indicate that integrating our proposed method into state-of-art adaptive Siamese trackers can increase the potential benefits of a template update strategy, and significantly improve performance. Code: https://github.com/madhukiranets/Adaptive-Siamese-Dimp
Madhu Kiran, Le Thanh Nguyen-Meidine, Rajat Sahay, Rafael M. O. Cruz, Louis-Antoine Blais-Morin, Eric Granger
IJCNN6
2022 Semi-Weakly Supervised Object Detection by Sampling Pseudo Ground-Truth Boxes
abstract
Semi- and weakly-supervised learning have recently attracted considerable attention in the object detection literature since they can alleviate the cost of annotation needed to successfully train deep learning models. State-of-art approaches for semi-supervised learning rely on student-teacher models trained using a multi-stage process, and considerable data augmentation. Custom networks have been developed for the weakly-supervised setting, making it difficult to adapt to different detectors. In this paper, a weakly semi-supervised training method is introduced that reduces these training challenges, yet achieves state-of-the-art performance by leveraging only a small fraction of fully-labeled images with information in weakly-labeled images. In particular, our generic sampling-based learning strategy produces pseudo ground-truth (GT) bounding box annotations in an online fashion, eliminating the need for multi-stage training, and student-teacher network configurations. These pseudo GT boxes are sampled from weakly-labeled images based on the categorical score of object proposals accumulated via a score propagation process. Empirical results11Our code is available at: https://github.com/akhilpm/SemiWSOD on the Pascal VOC dataset, indicates that the proposed approach improves performance by 5.0% when using VOC 2007 as fully-labeled, and VOC 2012 as weak-labeled data. Also, with 5-10% fully annotated images, we observed an improvement of more than 10% in$m$AP, showing that a modest investment in image-level annotation, can substantially improve detection performance.
Akhil Meethal, Marco Pedersoli, Zhongwen Zhu, Francisco Perdigón Romero, Eric Granger
IJCNN5
2022 F-CAM: Full Resolution Class Activation Maps via Guided Parametric Upscaling
abstract
Class Activation Mapping (CAM) methods have recently gained much attention for weakly-supervised object localization (WSOL) tasks. They allow for CNN visualization and interpretation without training on fully annotated image datasets. CAM methods are typically integrated within off-the-shelf CNN backbones, such as ResNet50. Due to convolution and pooling operations, these backbones yield low resolution CAMs with a down-scaling factor of up to 32, contributing to inaccurate localizations. Interpolation is required to restore full size CAMs, yet it does not consider the statistical properties of objects, such as color and texture, leading to activations with inconsistent boundaries, and inaccurate localizations. As an alternative, we introduce a generic method for parametric upscaling of CAMs that allows constructing accurate full resolution CAMs (FCAMs). In particular, we propose a trainable decoding architecture that can be connected to any CNN classifier to produce highly accurate CAM localizations. Given an original low resolution CAM, foreground and background pixels are randomly sampled to fine-tune the decoder. Additional priors such as image statistics and size constraints are also considered to expand and refine object boundaries. Extensive experiments1, over three CNN backbones and six WSOL baselines on the CUB-200-2011 and OpenImages datasets, indicate that our F-CAM method yields a significant improvement in CAM localization accuracy. F-CAM performance is competitive with state-of-art WSOL methods, yet it requires fewer computations during inference.
Soufiane Belharbi, Aydin Sarraf, Marco Pedersoli, Ismail Ben Ayed, Luke McCaffrey, Eric Granger
WACV6
2022 Cross-modal distillation for RGB-depth person re-identification
Frank Hafner, Amran Bhuyian, Julian F. P. Kooij, Eric Granger
Comput. Vis. Image Underst.4
2022 Incremental multi-target domain adaptation for object detection with efficient domain transfer
Le Thanh Nguyen-Meidine, Madhu Kiran, Marco Pedersoli, Jose Dolz, Louis-Antoine Blais-Morin, Eric Granger
Pattern Recognit.6
2022 A Deep Multiscale Spatiotemporal Network for Assessing Depression From Facial Dynamics
abstract
Recently, deep learning models have been successfully employed in many video-based affective computing applications (e.g., detecting pain, stress, and Alzheimer’s disease). One key application is automatic depression recognition – recognition of facial expressions associated with depressive behaviour. State-of-the-art deep learning algorithms to recognize depression typically explore spatial and temporal information individually, by using 2D convolutional neural networks (CNNs) to analyze appearance information, and then by either mapping facial feature variations or averaging the depression level over video frames. This approach has limitations in terms of its ability to represent dynamic information that can help to accurately discriminate between depression levels. In contrast, models based on 3D CNNs allow to directly encode the spatio-temporal relationships, although these models rely on temporal information with fixed range and single receptive field. This approach limits the ability to capture variations of facial expression with diverse ranges, and the exploitation of diverse facial areas. In this article, a novel 3D CNN architecture – the Multiscale Spatiotemporal Network (MSN) – is introduced to effectively represent facial information related to depressive behaviours from videos. The basic structure of the model is composed of parallel convolutional layers with different temporal depths and sizes of receptive field, which allows the MSN to explore a wide range of spatio-temporal variations in facial expressions. Experimental results on two benchmark datasets show that our MSN architecture is effective, outperforming state-of-the-art methods in automatic depression recognition.
Wheidima C. Melo, Eric Granger, Abdenour Hadid
IEEE Trans. Affect. Comput.2
2022 Deep Interpretable Classification and Weakly-Supervised Segmentation of Histology Images via Max-Min Uncertainty
abstract
Weakly-supervised learning (WSL) has recently triggered substantial interest as it mitigates the lack of pixel-wise annotations. Given global image labels, WSL methods yield pixel-level predictions (segmentations), which enable to interpret class predictions. Despite their recent success, mostly with natural images, such methods can face important challenges when the foreground and background regions have similar visual cues, yielding high false-positive rates in segmentations, as is the case in challenging histology images. WSL training is commonly driven by standard classification losses, which implicitly maximize model confidence, and locate the discriminative regions linked to classification decisions. Therefore, they lack mechanisms for modeling explicitly non-discriminative regions and reducing false-positive rates. We propose novel regularization terms, which enable the model to seek both non-discriminative and discriminative regions, while discouraging unbalanced segmentations. We introduce high uncertainty as a criterion to localize non-discriminative regions that do not affect classifier decision, and describe it with original Kullback-Leibler (KL) divergence losses evaluating the deviation of posterior predictions from the uniform distribution. Our KL terms encourage high uncertainty of the model when the latter inputs the latent non-discriminative regions. Our loss integrates: (i) a cross-entropy seeking a foreground, where model confidence about class prediction is high; (ii) a KL regularizer seeking a background, where model uncertainty is high; and (iii) log-barrier terms discouraging unbalanced segmentations. Comprehensive experiments and ablation studies over the public GlaS colon cancer data and a Camelyon16 patch-based benchmark for breast cancer show substantial improvements over state-of-the-art WSL methods, and confirm the effect of our new regularizers (our code is publicly available at https://github.com/sbelharbi/deep-wsl-histo-min-max-uncertainty).
Soufiane Belharbi, Jérôme Rony, Jose Dolz, Ismail Ben Ayed, Luke McCaffrey, Eric Granger
IEEE Trans. Medical Imaging6
2021 Variational Fair Clustering
abstract
We propose a general variational framework of fair clustering, which integrates an original Kullback-Leibler (KL) fairness term with a large class of clustering objectives, including prototype or graph based. Fundamentally different from the existing combinatorial and spectral solutions, our variational multi-term approach enables to control the trade-off levels between the fairness and clustering objectives. We derive a general tight upper bound based on a concave-convex decomposition of our fairness term, its Lipschitz-gradient property and the Pinsker’s inequality. Our tight upper bound can be jointly optimized with various clustering objectives, while yielding a scalable solution, with convergence guarantee. Interestingly, at each iteration, it performs an independent update for each assignment variable. Therefore, it can be easily distributed for large-scale datasets. This scalability is important as it enables to explore different trade-off levels between the fairness and clustering objectives. Unlike spectral relaxation, our formulation does not require computing its eigenvalue decomposition. We report comprehensive evaluations and comparisons with state-of-the-art methods over various fair clustering benchmarks, which show that our variational formulation can yield highly competitive solutions in terms of fairness and clustering objectives.
Imtiaz Masud Ziko, Jing Yuan 0001, Eric Granger, Ismail Ben Ayed
AAAI3
2021 Holistic Guidance for Occluded Person Re-Identification
Madhu Kiran, Gnana Praveen Rajasekhar, Le Thanh Nguyen-Meidine, Soufiane Belharbi, Louis-Antoine Blais-Morin, Eric Granger
BMVC6
2021 Cross Attentional Audio-Visual Fusion for Dimensional Emotion Recognition
abstract
Multimodal analysis has recently drawn much interest in affective computing, since it can improve the overall accuracy of emotion recognition over isolated uni-modal approaches. The most effective techniques for multimodal emotion recognition efficiently leverage diverse and complimentary sources of information, such as facial, vocal, and physiological modalities, to provide comprehensive feature representations. In this paper, we focus on dimensional emotion recognition based on the fusion of facial and vocal modalities extracted from videos, where complex spatiotemporal relationships may be captured. Most of the existing fusion techniques rely on recurrent networks or conventional attention mechanisms that do not effectively leverage the complimentary nature of audiovisual (A-V) modalities. We introduce a cross-attentional fusion approach to extract the salient features across A - V modalities, allowing for accurate prediction of continuous values of valence and arousal. Our new cross-attentional A - V fusion model efficiently leverages the inter-modal relationships. In particular, it computes cross-attention weights to focus on the more contributive features across individual modalities, and thereby combine contributive feature representations, which are then fed to fully connected layers for the prediction of valence and arousal. The effectiveness of the proposed approach is validated experimentally on videos from the RECOLA and Fatigue (private) data-sets. Results indicate that our cross-attentional A - V fusion model is a cost-effective approach that outperforms state-of-the-art fusion approaches. Code is available: https://github.com/praveena2j/Cross-Attentional-AV-Fusion.
Gnana Praveen Rajasekhar, Eric Granger, Patrick Cardinal
FG2
2021 Augmented Lagrangian Adversarial Attacks
abstract
Adversarial attack algorithms are dominated by penalty methods, which are slow in practice, or more efficient distance-customized methods, which are heavily tailored to the properties of the distance considered. We propose a white-box attack algorithm to generate minimally perturbed adversarial examples based on Augmented Lagrangian principles. We bring several algorithmic modifications, which have a crucial effect on performance. Our attack enjoys the generality of penalty methods and the computational efficiency of distance-customized algorithms, and can be readily used for a wide set of distances. We compare our attack to state-of-the-art methods on three datasets and several models, and consistently obtain competitive performances with similar or lower computational complexity.
Jérôme Rony, Eric Granger, Marco Pedersoli, Ismail Ben Ayed
ICCV2
2021 Temporal Stochastic Softmax for 3D CNNs: An Application in Facial Expression Recognition
abstract
Training deep learning models for accurate spatiotemporal recognition of facial expressions in videos requires significant computational resources. For practical reasons, 3D Convolutional Neural Networks (3D CNNs) are usually trained with relatively short clips randomly extracted from videos. However, such uniform sampling is generally sub-optimal because equal importance is assigned to each temporal clip. In this paper, we present a strategy for efficient video-based training of 3D CNNs. It relies on softmax temporal pooling and a weighted sampling mechanism to select the most relevant training clips. The proposed softmax strategy provides several advantages – a reduced computational complexity due to efficient clip sampling, and an improved accuracy since temporal weighting focuses on more relevant clips during both training and inference. Experimental results obtained with the proposed method on several facial expression recognition benchmarks show the benefits of focusing on more informative clips in training videos. In particular, our approach improves performance and computational cost by reducing the impact of inaccurate trimming and coarse annotation of videos, and heterogeneous distribution of visual information across time.
Théo Ayral, Marco Pedersoli, Simon Bacon, Eric Granger
WACV4
2021 Deep Active Learning for Joint Classification & Segmentation with Weak Annotator
abstract
CNN visualization and interpretation methods, like class-activation maps (CAMs), are typically used to highlight the image regions linked to class predictions. These models allow to simultaneously classify images and extract class-dependent saliency maps, without the need for costly pixel-level annotations. However, they typically yield segmentations with high false-positive rates and, therefore, coarse visualisations, more so when processing challenging images, as encountered in histology. To mitigate this issue, we propose an active learning (AL) framework, which progressively integrates pixel-level annotations during training. Given training data with global image-level labels, our deep weakly-supervised learning model jointly performs supervised image-level classification and active learning for segmentation, integrating pixel annotations by an oracle. Unlike standard AL methods that focus on sample selection, we also leverage large numbers of unlabeled images via pseudo-segmentations (i.e., self-learning at the pixel level), and integrate them with the oracle-annotated samples during training. We report extensive experiments over two challenging benchmarks -- high-resolution medical images (histology GlaS data for colon cancer) and natural images (CUB-200-2011 for bird species). Our results indicate that, by simply using random sample selection, the proposed approach can significantly outperform state-of the-art CAMs and AL methods, with an identical oracle-supervision budget. Our code is publicly available.
Soufiane Belharbi, Ismail Ben Ayed, Luke McCaffrey, Eric Granger
WACV4
2021 Unsupervised Multi-Target Domain Adaptation Through Knowledge Distillation
abstract
Unsupervised domain adaptation (UDA) seeks to alleviate the problem of domain shift between the distribution of unlabeled data from the target domain w.r.t. labeled data from the source domain. While the single-target UDA scenario is well studied in the literature, Multi-Target Domain Adaptation (MTDA) remains largely unexplored despite its practical importance, e.g., in multi-camera video-surveillance applications. The MTDA problem can be addressed by adapting one specialized model per target domain, although this solution is too costly in many real-world applications. Blending multiple targets for MTDA has been proposed, yet this solution may lead to a reduction in model specificity and accuracy. In this paper, we propose a novel unsupervised MTDA approach to train a CNN that can generalize well across multiple target domains. Our Multi-Teacher MTDA (MT-MTDA) method relies on multi-teacher knowledge distillation (KD) to iteratively distill target domain knowledge from multiple teachers to a common student. The KD process is performed in a progressive manner, where the student is trained by each teacher on how to perform UDA for a specific target, instead of directly learning domain adapted features. Finally, instead of combining the knowledge from each teacher, MT-MTDA alternates between teachers that distill knowledge, thereby preserving the specificity of each target (teacher) when learning to adapt to the student. MT-MTDA is compared against state- of-the-art methods on several challenging UDA benchmarks, and empirical results show that our proposed model can provide a considerably higher level of accuracy across multiple target domains. Our code is available at: https://gi.thub.com/LIVIAETS/MT-MTDA.
Le Thanh Nguyen-Meidine, Atif Belal, Madhu Kiran, Jose Dolz, Louis-Antoine Blais-Morin, Eric Granger
WACV6
2021 Flow guided mutual attention for person re-identification
Madhu Kiran, Amran Bhuiyan, Le Thanh Nguyen-Meidine, Louis-Antoine Blais-Morin, Ismail Ben Ayed, Eric Granger
Image Vis. Comput.6
2021 Knowledge distillation methods for efficient unsupervised adaptation across multiple domains
Le Thanh Nguyen-Meidine, Atif Belal, Madhu Kiran, Jose Dolz, Louis-Antoine Blais-Morin, Eric Granger
Image Vis. Comput.6
2021 Deep domain adaptation with ordinal regression for pain assessment using weakly-labeled videos
Gnana Praveen Rajasekhar, Eric Granger, Patrick Cardinal
Image Vis. Comput.2
2021 Boundary loss for highly unbalanced segmentation
Hoel Kervadec, Jihene Bouchtiba, Christian Desrosiers, Eric Granger, Jose Dolz, Ismail Ben Ayed
Medical Image Anal.4
2020 A Unifying Mutual Information View of Metric Learning: Cross-Entropy vs. Pairwise Losses
Malik Boudiaf, Jérôme Rony, Imtiaz Masud Ziko, Eric Granger, Marco Pedersoli, Pablo Piantanida, Ismail Ben Ayed
ECCV (6)4
2020 Unsupervised Domain Adaptation in the Dissimilarity Space for Person Re-identification
Djebril Mekhazni, Amran Bhuiyan, George S. Eskander Ekladious, Eric Granger
ECCV (27)4
2020 Deep Weakly Supervised Domain Adaptation for Pain Localization in Videos
abstract
Automatic pain assessment has an important potential diagnostic value for populations that are incapable of articulating their pain experiences. As one of the dominating nonverbal channels for eliciting pain expression events, facial expressions has been widely investigated for estimating the pain intensity of individual. However, using state-of-the-art deep learning (DL) models in real-world pain estimation applications poses several challenges related to the subjective variations of facial expressions, operational capture conditions, and lack of representative training videos with labels. Given the cost of annotating intensity levels for every video frame, we propose a weakly-supervised domain adaptation (WSDA) technique that allows for training 3D CNNs for spatiotemporal pain intensity estimation using weakly labeled videos, where labels are provided on a periodic basis. In particular, WSDA integrates multiple instance learning into an adversarial deep domain adaptation framework to train an Inflated 3D-CNN (I3D) model such that it can accurately estimate pain intensities in the target operational domain. The training process relies on weak target loss, along with domain loss and source loss for domain adaptation of the I3D model. Experimental results obtained using labeled source domain RECOLA videos and weakly-labeled target domain UNBC-McMaster videos indicate that the proposed deep WSDA approach can achieve significantly higher level of sequence (bag)-level and frame (instance)-level pain localization accuracy than related state-of-the-art approaches.
Gnana Praveen Rajasekhar, Eric Granger, Patrick Cardinal
FG2
2020 Encoding Temporal Information For Automatic Depression Recognition From Facial Analysis
abstract
Depression is a mental illness that may be harmful to an individual's health. Using deep learning models to recognize the facial expressions of individuals captured in videos has shown promising results for automatic depression detection. Typically, depression levels are recognized using 2D-Convolutional Neural Networks (CNNs) that are trained to extract static features from video frames, which impairs the capture of dynamic spatio-temporal relations. As an alternative, 3D-CNNs may be employed to extract spatiotemporal features from short video clips, although the risk of overfitting increases due to the limited availability of labeled depression video data. To address these issues, we propose a novel temporal pooling method to capture and encode the spatio-temporal dynamic of video clips into an image map. This approach allows fine-tuning a pre-trained 2D CNN to model facial variations, and thereby improving the training process and model accuracy. Our proposed method is based on two-stream model that performs late fusion of appearance and dynamic information. Extensive experiments on two benchmark AVEC datasets indicate that the proposed method is efficient and outperforms the state-of-the-art schemes.
Wheidima C. Melo, Eric Granger, Miguel Bordallo López
ICASSP2
2020 Laplacian Regularized Few-Shot Learning
abstract
We propose a transductive Laplacian-regularized inference for few-shot tasks. Given any feature embedding learned from the base classes, we minimize a quadratic binary-assignment function containing two terms: (1) a unary term assigning query samples to the nearest class prototype, and (2) a pairwise Laplacian term encouraging nearby query samples to have consistent label assignments. Our transductive inference does not re-train the base model, and can be viewed as a graph clustering of the query set, subject to supervision constraints from the support set. We derive a computationally efficient bound optimizer of a relaxation of our function, which computes independent (parallel) updates for each query sample, while guaranteeing convergence. Following a simple cross-entropy training on the base classes, and without complex meta-learning strategies, we conducted comprehensive experiments over five few-shot learning benchmarks. Our LaplacianShot consistently outperforms state-of-the-art methods by significant margins across different models, settings, and data sets. Furthermore, our transductive inference is very fast, with computational times that are close to inductive inference, and can be used for large-scale few-shot tasks.
Imtiaz Masud Ziko, Jose Dolz, Eric Granger, Ismail Ben Ayed
ICML3
2020 Convolutional STN for Weakly Supervised Object Localization
abstract
Weakly-supervised object localization is a challenging task in which the object of interest should be localized while learning its appearance. State-of-the-art methods recycle the architecture of a standard CNN by using the activation maps of the last layer for localizing the object. While this approach is simple and works relatively well, object localization relies on different features than classification, thus, a specialized localization mechanism is required during training to improve performance. In this paper, we propose a convolutional, multi-scale spatial localization network that provides accurate localization for the object of interest. Experimental results on CUB-200-2011 and ImageNet datasets show that our proposed approach provides competitive performance for weakly supervised localization.
Akhil Meethal, Marco Pedersoli, Soufiane Belharbi, Eric Granger
ICPR4
2020 Progressive Gradient Pruning for Classification, Detection and Domain Adaptation
abstract
Although deep neural networks (NNs) have achieved state-of-the-art accuracy in many visual recognition tasks, the growing computational complexity and energy consumption of networks remains an issue, especially for applications on platforms with limited resources and requiring real-time processing. Filter pruning techniques have recently shown promising results for the compression and acceleration of convolutional NNs (CNNs). However, these techniques involve numerous steps and complex optimisations because some only prune after training CNNs, while others prune from scratch during training, by integrating sparsity constraints or by modifying the loss function. In this paper, we introduce a new Progressive Gradient Pruning (PGP) technique for iterative filter pruning during training. In contrast to previous progressive pruning techniques, it relies on a novel filter selection criterion that measures the change in filter weights, uses a new hard and soft pruning strategy, and effectively adapts momentum tensors during the backward propagation pass. Experimental results obtained after training various CNNs on benchmark datasets for image classification, object detection and domain adaptation indicate that our PGP technique can achieve a better trade-off between classification accuracy and network (time and memory) complexity than PSFP and other state-of-the-art filter pruning techniques. Code is available on GitHub link: https://github.com/Anon6627/Pruning-PGP.
Le Thanh Nguyen-Meidine, Eric Granger, Madhu Kiran, Marco Pedersoli, Louis-Antoine Blais-Morin
ICPR2
2020 Dual-Triplet Metric Learning for Unsupervised Domain Adaptation in Video Face Recognition
abstract
The scalability and complexity of deep learning models remains a key issue in many of visual recognition applications. For instance, in video surveillance, fine tuning of a model with labeled image data from each new camera is required to reduce the domain shift between videos captured from the source domain (laboratory setting) and the target domain (operational environment). In many video surveillance applications, like face recognition and person re-identification, a pair-wise matcher is typically employed to assign a query image captured using a video camera to the corresponding reference images in a gallery. The different configuration, viewpoint, and operational conditions of each camera can introduce significant shifts in pair-wise distance distributions, resulting in a decline in recognition performance for new cameras. In this paper, a new deep domain adaptation (DA) method is proposed to adapt the CNN embedding of a Siamese network using unlabeled tracklets captured with a new video camera. To this end, a dual-triplet loss is introduced for metric learning, where two triplets are constructed using video data from a source camera, and a new target camera. In order to constitute the dual triplets, a mutual-supervised learning approach is introduced where the source camera acts as a teacher, providing the target camera with an initial embedding. Then, the student relies on the teacher to iteratively label the positive and negative pairs collected during, e.g., initial camera calibration. Both source and target embeddings continue to simultaneously learn such that their pair-wise distance distributions become aligned. For validation, the proposed metric learning technique is used to train deep Siamese networks under different training scenarios, and is compared to state-of-the-art techniques for still-to-video FR on the COXS2V and a private video-based FR dataset. Results indicate that the proposed method can provide a level of accuracy that is comparable to the upper bound performance, in training scenario where labeled target data is employed to fine-tune the Siamese network.
George S. Eskander Ekladious, Hugo Lemoine, Eric Granger, Kaveh Kamali, Salim Moudache
IJCNN3
2020 Joint Progressive Knowledge Distillation and Unsupervised Domain Adaptation
abstract
Currently, the divergence in distributions of design and operational data, and large computational complexity are limiting factors in the adoption of CNNs in real-world applications. For instance, person re-identification systems typically rely on a distributed set of cameras, where each camera has different capture conditions. This can translate to a considerable shift between source (e.g. lab setting) and target (e.g. operational camera) domains. Given the cost of annotating image data captured for fine-tuning in each target domain, unsupervised domain adaptation (UDA) has become a popular approach to adapt CNNs. Moreover, state-of-the-art deep learning models that provide a high level of accuracy often rely on architectures that are too complex for real-time applications. Although several compression and UDA approaches have recently been proposed to overcome these limitations, they do not allow optimizing a CNN to simultaneously address both. In this paper, we propose an unexplored direction - the joint optimization of CNNs to provide a compressed model that is adapted to perform well for a given target domain. In particular, the proposed approach performs unsupervised knowledge distillation (KD) from a complex teacher model to a compact student model, by leveraging both source and target data. It also improves upon existing UDA techniques by progressively teaching the student about domain-invariant features, instead of directly adapting a compact model on target domain data. Our method is compared against state-of-the-art compression and UDA techniques, using two popular classification datasets for UDA - Office31 and ImageClef-DA. In both datasets, results indicate that our method can achieve the highest level of accuracy while requiring a comparable or lower time complexity.
Le Thanh Nguyen-Meidine, Eric Granger, Madhu Kiran, Jose Dolz, Louis-Antoine Blais-Morin
IJCNN2
2020 Pose Guided Gated Fusion for Person Re-identification
abstract
Person re-identification is an important yet challenging problem in visual recognition. Despite the recent advances with deep learning (DL) models for spatio-temporal and multi-modal fusion, re-identification approaches often fail to leverage the contextual information (e.g., pose and illumination) to dynamically select the most discriminant con-volutional filters (i.e., appearance features) for feature representation and inference. State-of-the-art techniques for gated fusion employ complex dedicated part- or attention-based architectures for late fusion, and do not incorporate pose and appearance information to train the backbone network. In this paper, a new DL model is proposed for pose-guided re-identification, comprised of a deep backbone, pose estimation, and gated fusion network. Given a query image of an individual, the backbone convolutional NN produces a feature embedding required for pair-wise matching with embeddings for reference images, where feature maps from the pose network and from mid-level CNN layers are combined by the gated fusion network to generate pose-guided gating. The proposed framework allows to dynamically activate the most discriminant CNN filters based on pose information in order to perform a finer grained recognition. Extensive experiments on three challenging benchmark datasets indicate that integrating the pose-guided gated fusion into the state-of-the-art re-identification backbone architecture allows to improve their recognition accuracy. Experimental results also support our intuition on the advantages of gating backbone appearance information using the pose feature maps at mid-level CNN layers.
Amran Bhuiyan, Yang Liu 0093, Parthipan Siva, Mehrsan Javan Roshtkhari, Ismail Ben Ayed, Eric Granger
WACV6
2020 Cross-Domain Face Synthesis using a Controllable GAN
abstract
The performance of face recognition (FR) systems for video surveillance has been shown to improve when the design data is augmented through synthetic face generation. This is true, for instance, with pair-wise matchers (e.g., deep Siamese networks) that rely on a reference gallery, typically with one still image per individual. However, generating synthetic images based on stills (from the source domain) may not improve performance during operations due to the domain shift w.r.t. the target domain. Moreover, despite the emergence of Generative Adversarial Networks (GANs) for realistic synthetic generation, it is often difficult to control the conditions under which synthetic faces are generated. In this paper, a cross-domain face synthesis approach is proposed that integrates a new Controllable GAN (C-GAN). It employs an off-the-shelf 3D face model as a simulator to generate facial images under various poses. The simulated images and noise are input to the C-GAN for realism refinement. It relies on an additional adversarial game as a third player to preserve the identity and specific facial attributes of the refined images. This allows generating realistic synthetic face images that reflect capture conditions in the target domain, while controlling the GAN output such that faces may be generated under desired pose conditions. Experiments were performed using videos from the Chokepoint and COX-S2V datasets, and a deep Siamese network for FR with a single reference still per person. Results indicate that the proposed approach can provide a higher level of accuracy compared to state- of-the-art approaches for synthetic data augmentation1.
Fania Mokhayeri, Kaveh Kamali, Eric Granger
WACV3
2020 Robust face tracking using multiple appearance models and graph relational learning
Tanushri Chakravorty, Guillaume-Alexandre Bilodeau, Eric Granger
Mach. Vis. Appl.3
2020 A paired sparse representation model for robust face recognition from a single sample
Fania Mokhayeri, Eric Granger
Pattern Recognit.2
2020 F-measure curves: A tool to visualize classifier performance under imbalance
Roghayeh Soleymani, Eric Granger, Giorgio Fumera
Pattern Recognit.2
2020 Feature Learning from Spectrograms for Assessment of Personality Traits
abstract
Several methods have recently been proposed to analyze speech and automatically infer the personality of the speaker. These methods often rely on prosodic and other hand crafted speech processing features extracted with off-the-shelf toolboxes. To achieve high accuracy, numerous features are typically extracted using complex and highly parameterized algorithms. In this paper, a new method based on feature learning and spectrogram analysis is proposed to simplify the feature extraction process while maintaining a high level of accuracy. The proposed method learns a dictionary of discriminant features from patches extracted in the spectrogram representations of training speech segments. Each speech segment is then encoded using the dictionary, and the resulting feature set is used to perform classification of personality traits. Experiments indicate that the proposed method achieves state-of-the-art results with an important reduction in complexity when compared to the most recent reference methods. The number of features, and difficulties linked to the feature extraction process are greatly reduced as only one type of descriptors is used, for which the 7 parameters can be tuned automatically. In contrast, the simplest reference method uses 4 types of descriptors to which 6 functionals are applied, resulting in over 20 parameters to be tuned.
Marc-André Carbonneau, Eric Granger, Yazid Attabi, Ghyslain Gagnon
IEEE Trans. Affect. Comput.2
2019 RGB-Depth Cross-Modal Person Re-identification
abstract
Person re-identification is a key challenge for surveillance across multiple sensors. Prompted by the advent of powerful deep learning models for visual recognition, and inexpensive RGBD cameras and sensor-rich mobile robotic platforms, e.g. self-driving vehicles, we investigate the relatively unexplored problem of cross-modal re-identification of persons between RGB (color) and depth images. The considerable divergence in data distributions across different sensor modalities introduces additional challenges to the typical difficulties like distinct viewpoints, occlusions, and pose and illumination variation. While some work has investigated re-identification across RGB and infrared, we take inspiration from successes in transfer learning from RGB to depth in object detection tasks. Our main contribution is a novel cross-modal distillation network for robust person re-identification, which learns a shared feature representation space of person's appearance in both RGB and depth images. The proposed network was compared to conventional and deep learning approaches proposed for other cross-domain re-identification tasks. Results obtained on the public BIWI and RobotPKU datasets indicate that the proposed method can significantly outperform the state-of-the-art approaches by up to 10.5% mAp, demonstrating the benefit of the proposed distillation paradigm.
Frank Hafner, Amran Bhuiyan, Julian F. P. Kooij, Eric Granger
AVSS4
2019 On the Interaction Between Deep Detectors and Siamese Trackers in Video Surveillance
abstract
Visual object tracking is an important function in many real-time video surveillance applications, such as localization and spatio-temporal recognition of persons. In realworld applications, an object detector and tracker must interact on a periodic basis to discover new objects, and thereby to initiate tracks. Periodic interactions with the detector can also allow the tracker to validate and/or update its object template with new bounding boxes. However, bounding boxes provided by a state-of-the-art detector are noisy, due to changes in appearance, background and occlusion, which can cause the tracker to drift. Moreover, CNN-based detectors can provide a high level of accuracy at the expense of computational complexity, so interactions should be minimized for real-time applications. In this paper, a new approach is proposed to manage detector-tracker interactions for trackers from the Siamese-FC family. By integrating a change detection mechanism into a deep Siamese-FC tracker, its template can be adapted in response to changes in a target's appearance that lead to drifts during tracking. An abrupt change detection triggers an update of tracker template using the bounding box produced by the detector, while in the case of a gradual change, the detector is used to update an evolving set of templates for robust matching. Experiments were performed using state-of-the-art Siamese-FC trackers and the YOLOv3 detector on a subset of videos from the OTB-100 dataset that mimic video surveillance scenarios. Results highlight the importance for reliable VOT of using accurate detectors. They also indicate that our adaptive Siamese trackers are robust to noisy object detections, and can significantly improve the performance of Siamese-FC tracking.
Madhu Kiran, Vivek Tiwari, Le Thanh Nguyen-Meidine, Louis-Antoine Blais-Morin, Eric Granger
AVSS5
2019 Decoupling Direction and Norm for Efficient Gradient-Based L2 Adversarial Attacks and Defenses
abstract
Research on adversarial examples in computer vision tasks has shown that small, often imperceptible changes to an image can induce misclassification, which has security implications for a wide range of image processing systems. Considering L2 norm distortions, the Carlini and Wagner attack is presently the most effective white-box attack in the literature. However, this method is slow since it performs a line-search for one of the optimization terms, and often requires thousands of iterations. In this paper, an efficient approach is proposed to generate gradient-based attacks that induce misclassifications with low L2 norm, by decoupling the direction and the norm of the adversarial perturbation that is added to the image. Experiments conducted on the MNIST, CIFAR-10 and ImageNet datasets indicate that our attack achieves comparable results to the state-of-the-art (in terms of L2 norm) with considerably fewer iterations (as few as 100 iterations), which opens the possibility of using these attacks for adversarial training. Models trained with our attack achieve state-of-the-art robustness against white-box gradient-based L2 attacks on the MNIST and CIFAR-10 datasets, outperforming the Madry defense when the attacks are limited to a maximum norm.
Jérôme Rony, Luiz G. Hafemann, Luiz Eduardo Soares de Oliveira, Ismail Ben Ayed, Robert Sabourin, Eric Granger
CVPR6
2019 Combining Global and Local Convolutional 3D Networks for Detecting Depression from Facial Expressions
abstract
Deep learning architectures have been successfully applied in video-based health monitoring, to recognize distinctive variations in the facial appearance of subjects. To detect patterns of variation linked to depressive behavior, deep neural networks (NNs) typically exploit spatial and temporal information separately by, e.g., cascading a 2D convolutional NN (CNN) with a recurrent NN (RNN), although the intrinsic spatio-temporal relationships can deteriorate. With the recent advent of 3D CNNs like the convolutional 3D (C3D) network, these spatio-temporal relationships can be modeled to improve performance. However, the accuracy of C3D networks remain an issue when applied to depression detection. In this paper, the fusion of diverse C3D predictions are proposed to improve accuracy, where spatio-temporal features are extracted from global (full-face) and local (eyes) regions of subject. This allows to increasingly focus on a local facial region that is highly relevant for analyzing depression. Additionally, the proposed network integrates 3D Global Average Pooling in order to efficiently summarize spatio-temporal features without using fully-connected layers, and thereby reduce the number of model parameters and potential over-fitting. Experimental results on the Audio Visual Emotion Challenge (AVEC 2013 and AVEC 2014) depression datasets indicates that combining the responses of global and local C3D networks achieves a higher level of accuracy than state-of-the-art systems.
Wheidima C. Melo, Eric Granger, Abdenour Hadid
FG2
2019 Robust Video Face Recognition From a Single Still Using a Synthetic Plus Variational Model
abstract
Sparse representation-based classification (SRC) techniques have been shown to achieve a high level of performance in video-based face recognition (FR). However, matching faces captured in uncontrolled video conditions against a gallery with a single reference facial still per individual typically yields low accuracy. To improve robustness to intra-class variations, SRC techniques for FR have recently been extended to incorporate variational information from an external generic set into an auxiliary variational dictionary. Despite their success in handling linear variations, probe facial images with non-linear variations due to e.g., changes in pose and expressions, cannot be accurately reconstructed with a linear combination of images from gallery and auxiliary dictionaries because they do not share the same type of variations. In this paper, a new synthetic plus variational model is proposed to account for the non-linearities, particularly with pose variations. It reconstructs a probe image using (1) an auxiliary variational dictionary and (2) an augmented gallery dictionary enriched with a set of synthetic images generated from the reference faces with a wide diversity of pose angles. By solving a newly formulated simultaneous sparsity-based optimization problem, the augmented gallery dictionary is encouraged to share the same sparsity pattern with the variational dictionary for the same pose angles. In this way, each synthetic face in the augmented dictionary is combined with similar facial viewpoint in the variational dictionary. Experimental results obtained on Chokepoint and COX-S2V datasets, using different face representations, indicate that the proposed approach can outperform state-of-the-art SRC-based methods for still-to-video FR with a SSPP.
Fania Mokhayeri, Eric Granger
FG2
2019 Depression Detection Based on Deep Distribution Learning
abstract
Major depressive disorder is among the most common and harmful mental health problems. Several deep learning architectures have been proposed for video-based detection of depression based on the facial expressions of subjects. To predict the depression level, these architectures are often modeled for regression with Euclidean loss. Consequently, they do not leverage the data distribution, nor explore the ordinal relationship between facial images and depression levels, and have limited robustness to noisy and uncertain labeling. This paper introduces a deep learning architecture for accurately predicting depression levels through distribution learning. It relies on a new expectation loss function that allows to estimate the underlying data distribution over depression levels, where expected values of the distribution are optimized to approach the ground-truth levels. The proposed approach can produce accurate predictions of depression levels even under label uncertainty. Extensive experiments on the AVEC2013 and AVEC2014 datasets indicate that the proposed architecture represents an effective approach that can outperform state-of-the-art techniques.
Wheidima C. Melo, Eric Granger, Abdenour Hadid
ICIP2
2019 Curriculum Semi-supervised Segmentation
Hoel Kervadec, Jose Dolz, Eric Granger, Ismail Ben Ayed
MICCAI (2)3
2019 Constrained-CNN losses for weakly supervised segmentation
Hoel Kervadec, Jose Dolz, Meng Tang 0001, Eric Granger, Yuri Boykov, Ismail Ben Ayed
Medical Image Anal.4
2019 Domain-Specific Face Synthesis for Video Face Recognition From a Single Sample Per Person
abstract
In video surveillance, face recognition (FR) systems are employed to detect individuals of interest appearing over a distributed network of cameras. The performance of still-to-video FR systems can decline significantly because faces captured in unconstrained operational domain (OD) over multiple video cameras have a different underlying data distribution compared to faces captured under controlled conditions in the enrollment domain with a still camera. This is particularly true when individuals are enrolled to the system using a single reference still. To improve the robustness of these systems, it is possible to augment the reference set by generating synthetic faces based on the original still. However, without the knowledge of the OD, many synthetic images must be generated to account for all possible capture conditions. FR systems may, therefore, require complex implementations and yield lower accuracy when training on many less relevant images. This paper introduces an algorithm for domain-specific face synthesis (DSFS) that exploits the representative intra-class variation information available from the OD. Prior to operation (during camera calibration), a compact set of faces from unknown persons appearing in the OD is selected through affinity propagation clustering in the captured condition space (defined by pose and illumination estimation). The domain-specific variations of these face images are then projected onto the reference still of each individual by integrating an image-based face relighting technique inside the 3-D reconstruction framework. A compact set of synthetic faces is generated that resemble individuals of interest under the capture conditions relevant to the OD. In a particular implementation based on sparse representation classification, the synthetic faces generated with the DSFS are employed to form a cross-domain dictionary that accounts for structured sparsity, where the dictionary blocks combine the original and synthetic faces of each individual. Experimental results obtained with videos from the Chokepoint and COX-S2V data sets reveal that augmenting the reference gallery set of still-to-video FR systems using the proposed DSFS approach can provide a significantly higher level of accuracy compared with the state-of-the-art approaches, with only a moderate increase in its computational complexity.
Fania Mokhayeri, Eric Granger, Guillaume-Alexandre Bilodeau
IEEE Trans. Inf. Forensics Secur.2
2019 Bag-Level Aggregation for Multiple-Instance Active Learning in Instance Classification Problems
abstract
A growing number of applications, e.g., video surveillance and medical image analysis, require training recognition systems from large amounts of weakly annotated data, while some targeted interactions with a domain expert are allowed to improve the training process. In such cases, active learning (AL) can reduce labeling costs for training a classifier by querying the expert to provide the labels of most informative instances. This paper focuses on AL methods for instance classification problems in multiple instance learning (MIL), where data are arranged into sets, called bags, which are weakly labeled. Most AL methods focus on single-instance learning problems. These methods are not suitable for MIL problems because they cannot account for the bag structure of data. In this paper, new methods for bag-level aggregation of instance informativeness are proposed for multiple instance AL (MIAL). The aggregated informativeness method identifies the most informative instances based on classifier uncertainty and queries bags incorporating the most information. The other proposed method, called cluster-based aggregative sampling, clusters data hierarchically in the instance space. The informativeness of instances is assessed by considering bag labels, inferred instance labels, and the proportion of labels that remain to be discovered in clusters. Both proposed methods significantly outperform reference methods in extensive experiments using benchmark data from several application domains. Results indicate that using an appropriate strategy to address MIAL problems yields a significant reduction in the number of queries needed to achieve the same level of performance as single-instance AL methods.
Marc-André Carbonneau, Eric Granger, Ghyslain Gagnon
IEEE Trans. Neural Networks Learn. Syst.2
2018 Contextual Weighting of Patches for Local Matching in Still-to-Video Face Recognition
abstract
Still-to-video face recognition (FR) systems for watchlist screening seek to recognize individuals of interest given faces captured over a network of video surveillance cameras. Screening faces against a watchlist is a challenging application because only a limited number of reference stills is available per individual during enrollment, and the appearance of face captures in videos changes from camera to camera, due to variations in illumination, pose, blur, scale, expression and occlusion. In order to improve the robustness of FR systems, several local matching techniques have been proposed that rely on static or dynamic weighting of patches. However, these approaches are not suitable for watchlist screening applications where the capturing conditions vary significantly over different camera fields of view (FoV). In this paper, a new dynamic weighting technique is proposed for weighting facial patches based on video data collected a priori from the specific operational domain (camera FoV) and on image quality assessment. Results obtained on videos from the Chokepoint dataset indicate that the proposed approach can significantly outperform the reference local matching methods because patch weights tend to grow for discriminant facial regions.
Ibtihel Amara, Eric Granger, Abdenour Hadid
FG2
2018 Scalable Laplacian K-modes
abstract
We advocate Laplacian K-modes for joint clustering and density mode finding, and propose a concave-convex relaxation of the problem, which yields a parallel algorithm that scales up to large datasets and high dimensions. We optimize a tight bound (auxiliary function) of our relaxation, which, at each iteration, amounts to computing an independent update for each cluster-assignment variable, with guar- anteed convergence. Therefore, our bound optimizer can be trivially distributed for large-scale data sets. Furthermore, we show that the density modes can be obtained as byproducts of the assignment variables via simple maximum-value operations whose additional computational cost is linear in the number of data points. Our formulation does not need storing a full affinity matrix and computing its eigenvalue decomposition, neither does it perform expensive projection steps and Lagrangian-dual inner iterates for the simplex constraints of each point. Fur- thermore, unlike mean-shift, our density-mode estimation does not require inner- loop gradient-ascent iterates. It has a complexity independent of feature-space dimension, yields modes that are valid data points in the input set and is appli- cable to discrete domains as well as arbitrary kernels. We report comprehensive experiments over various data sets, which show that our algorithm yields very competitive performances in term of optimization quality (i.e., the value of the discrete-variable objective at convergence) and clustering accuracy.
Imtiaz Masud Ziko, Eric Granger, Ismail Ben Ayed
NeurIPS2
2018 Progressive boosting for class imbalance and its application to face re-identification
Roghayeh Soleymani, Eric Granger, Giorgio Fumera
Expert Syst. Appl.2
2018 Tracking using Numerous Anchor Points
Tanushri Chakravorty, Guillaume-Alexandre Bilodeau, Eric Granger
Mach. Vis. Appl.3
2018 Multiple instance learning: A survey of problem characteristics and applications
Marc-André Carbonneau, Veronika Cheplygina, Eric Granger, Ghyslain Gagnon
Pattern Recognit.3
2017 CNNs with cross-correlation matching for face recognition in video surveillance using a single training sample per person
abstract
In video surveillance, face recognition (FR) systems seek to detect individuals of interest appearing over a distributed network of cameras. Still-to-video FR systems match faces captured in videos under challenging conditions against facial models, often designed using one reference still per individual. Although CNNs can achieve among the highest levels of accuracy in many real-world FR applications, state-of-the-art CNNs that are suitable for still-to-video FR, like trunk-branch ensemble (TBE) CNNs, represent complex solutions for real-time applications. In this paper, an efficient CNN architecture is proposed for accurate still-to-video FR from a single reference still. The CCM-CNN is based on new cross-correlation matching (CCM) and triplet-loss optimization methods that provide discriminant face representations. The matching pipeline exploits a matrix Hadamard product followed by a fully connected layer inspired by adaptive weighted cross-correlation. A triplet-based training approach is proposed to optimize the CCM-CNN parameters such that the inter-class variations are increased, while enhancing robustness to intra-class variations. To further improve robustness, the network is fine-tuned using synthetically-generated faces based on still and videos of non-target individuals. Experiments on videos from the COX Face and Chokepoint datasets indicate that the CCM-CNN can achieve a high level of accuracy that is comparable to TBE-CNN and HaarNet, but with a significantly lower time and memory complexity. It may therefore represent the better trade-off between accuracy and complexity for real-time video surveillance applications.
Mostafa Parchami, Saman Bashbaghi, Eric Granger
AVSS3
2017 Using deep autoencoders to learn robust domain-invariant representations for still-to-video face recognition
abstract
Video-based face recognition (FR) is a challenging task in real-world applications. In still-to-video FR, probe facial regions of interest (ROIs) are typically captured with lower-quality video cameras under unconstrained conditions, where facial appearances vary according to pose, illumination, scale, expression, etc. These video ROIs are typically compared against facial models designed with high-quality reference still ROI of each target individual enrolled to the system. In this paper, an efficient Canonical Face Representation CNN (CFR-CNN) is proposed for accurate still-to-video FR from a single sample per person, where still and video ROIs are captured in different conditions. Given a facial ROI captured under unconstrained video conditions, the CRF-CNN reconstructs it as a high-quality canonical ROI for matching that corresponds to the conditons of reference still ROIs (e.g., well-illuminated, sharp, frontal views with neutral expression). A deep autoencoder network is trained using a novel weighted loss function that can robustly generate similar face embeddings for the same subjects. Then, during operations, those face embeddings belonging to pairs of still and video ROIs from a target individual are accurately matched using a fully-connected classification network. Experimental results obtained with the COX Face and Chokepoint datasets indicate that the proposed CFR-CNN can achieve convincing level of accuracy. The computational complexity (number of operations, network parameters and layers) is significantly lower than state-of-the-art CNNs for video FR, and suggests that the CFR-CNN represents a cost-effective solution for real-time applications.
Mostafa Parchami, Saman Bashbaghi, Eric Granger, Saif Iftekar Sayed
AVSS3
2017 Multi-orientation Scene Text Detection Leveraging Background Suppression
Xihan Wang, Xiaoyi Feng, Zhaoqiang Xia, Jinye Peng 0001, Eric Granger
ICIG (1)5
2017 Dynamic Selection of Exemplar-SVMs for Watch-list Screening through Domain Adaptation
Saman Bashbaghi, Eric Granger, Robert Sabourin, Guillaume-Alexandre Bilodeau
ICPRAM2
2017 Video-based face recognition using ensemble of haar-like deep convolutional neural networks
abstract
Growing number of surveillance and biometric applications seek to recognize the face of individuals appearing in the viewpoint of video cameras. Systems for video-based FR can be subjected to challenging operational environments, where the appearance of faces captured with video cameras varies significantly due to changes in pose, illumination, scale, blur, expression, occlusion, etc. In particular, with still-to-video FR, a limited number of high-quality facial images are typically captured for enrollment of an individual to the system, whereas an abundance facial trajectories can be captured using video cameras during operations, under different viewpoints and uncontrolled conditions. This paper presents a deep learning architecture that can learn a robust facial representation for each target individual during enrollment, and then accurately compare the facial regions of interest (ROIs) extracted from a still reference image (of the target individual) with ROIs extracted from live or archived videos. An ensemble of deep convolutional neural networks (DCNNs) named HaarNet is proposed, where a trunk network first extracts features from the global appearance of the facial ROIs (holistic representation). Then, three branch networks effectively embed asymmetrical and complex facial features (local representations) based on Haar-like features. In order to increase the discriminativness of face representations, a novel regularized triplet-loss function is proposed that reduces the intra-class variations, while increasing the inter-class variations. Given the single reference still per target individual, the robustness of the proposed DCNN is further improved by fine-tuning the HaarNet with synthetically-generated facial still ROIs that emulate capture conditions found in operational environments. The proposed system is evaluated on stills and videos from the challenging COX Face and Chokepoint datasets according to accuracy and complexity. Experimental results indicate that the proposed method can significantly improve performance with respect to state-of-the-art systems for video-based FR.
Mostafa Parchami, Saman Bashbaghi, Eric Granger
IJCNN3
2017 Robust watch-list screening using dynamic ensembles of SVMs based on multiple face representations
Saman Bashbaghi, Eric Granger, Robert Sabourin, Guillaume-Alexandre Bilodeau
Mach. Vis. Appl.2
2017 Dynamic ensembles of exemplar-SVMs for still-to-video face recognition
Saman Bashbaghi, Eric Granger, Robert Sabourin, Guillaume-Alexandre Bilodeau
Pattern Recognit.2
2016 Witness identification in multiple instance learning using random subspaces
abstract
Multiple instance learning (MIL) is a form of weakly-supervised learning where instances are organized in bags. A label is provided for bags, but not for instances. MIL literature typically focuses on the classification of bags seen as one object, or as a combination of their instances. In both cases, performance is generally measured using labels assigned to entire bags. In this paper, the MIL problem is formulated as a knowledge discovery task for which algorithms seek to discover the witnesses (i.e. identifying positive instances), using the weak supervision provided by bag labels. Some MIL methods are suitable for instance classification, but perform poorly in application where the witness rate is low, or when the positive class distribution is multimodal. A new method that clusters data projected in random subspaces is proposed to perform witness identification in these adverse settings. The proposed method is assessed on MIL data sets from three application domains, and compared to 7 reference MIL algorithms for the witness identification task. The proposed algorithm constantly ranks among the best methods in all experiments, while all other methods perform unevenly across data sets.
Marc-André Carbonneau, Eric Granger, Ghyslain Gagnon
ICPR2
2016 Loss factors for learning Boosting ensembles from imbalanced data
abstract
Class imbalance is an issue in many real world applications because classification algorithms tend to misclassify instances from the class of interest when its training samples are outnumbered by those of other classes. Several variations of AdaBoost ensemble method have been proposed in literature to learn from imbalanced data based on re-sampling. However, their loss factor is based on standard accuracy, which still biases performance towards the majority class. This problem is mitigated using cost-sensitive Boosting algorithms, although it can be avoided at the outset by modifying the loss factor calculation. In this paper, two loss factors, based on F-measure and G-mean are proposed that are more suitable to deal with imbalanced data during the Boosting learning process. The performance of standard AdaBoost and of three specialized versions for class imbalance (SMOTEBoost, RUSBoost, and RB-Boost) are empirically evaluated using the proposed loss factors, both on synthetic data and on a real-world face re-identification task. Experimental results show a significant performance improvement on AdaBoost and RUSBoost with the proposed loss factors.
Roghayeh Soleymani, Eric Granger, Giorgio Fumera
ICPR2
2016 Learning of Graph Compressed Dictionaries for Sparse Representation Classification
abstract
Despite the limited target data available to design face models in video surveillance applications, many faces of non-target individuals may be captured in operational environments, and over multiple cameras, to improve robustness to variations. This paper focuses on Sparse Representation Classification (SRC) techniques that are suitable for the design of still-to-video FR systems based on under-sampled dictionaries. The limited reference data available during enrolment is complemented by an over-complete external dictionary that is formed with an abundance of faces from non-target individuals. In this paper, the Graph-Compressed Dictionary Learning (GCDL) technique is proposed to learn compact auxiliary dictionaries for SRC. GCDL is based on matrix factorization, and allows to maintain a high level of accuracy with compressed dictionaries because it exploits structural information to represent intra-class variations. Graph factorization compression has been shown to efficiently compress data, and can therefore rapidly construct compressed dictionaries. Accuracy and efficiency of the proposed technique is assessed and compared to reference sparse coding and dictionary learning technique using videos from the CAS-PEAL database. GCDL is shown to provide fast matching and adaptation of compressed dictionaries to new reference faces from the video surveillance environments.
Farshad Nourbakhsh, Eric Granger
ICPRAM2
2016 Classifier Ensembles with Trajectory Under-Sampling for Face Re-Identification
abstract
Class imbalance is an issue in many real world applications because classification algorithms tend to misclassify instances from the class of interest when its training samples are outnumbered by those of other classes. Several variations of AdaBoost ensemble method have been proposed in literature to learn from imbalanced data based on re-sampling. However, their loss factor is based on standard accuracy, which still biases performance towards the majority class. This problem is mitigated using cost-sensitive Boosting algorithms, although it can be avoided at the outset by modifying the loss factor calculation. In this paper, two loss factors, based on F-measure and G-mean are proposed that are more suitable to deal with imbalanced data during the Boosting learning process. The performance of standard AdaBoost and of three specialized versions for class imbalance (SMOTEBoost, RUSBoost, and RB-Boost) are empirically evaluated using the proposed loss factors, both on synthetic data and on a real-world face re-identification task. Experimental results show a significant performance improvement on AdaBoost and RUSBoost with the proposed loss factors.
Roghayeh Soleymani, Eric Granger, Giorgio Fumera
ICPRAM2
2016 Robust multiple-instance learning ensembles using random subspace instance selection
Marc-André Carbonneau, Eric Granger, Alexandre J. Raymond, Ghyslain Gagnon
Pattern Recognit.2
2016 Adaptive appearance model tracking for still-to-video face recognition
abstract
Systems for still-to-video face recognition (FR) seek to detect the presence of target individuals based on reference facial still images or mug-shots. These systems encounter several challenges in video surveillance applications due to variations in capture conditions (e.g., pose, scale, illumination, blur and expression) and to camera inter-operability. Beyond these issues, few reference stills are available during enrollment to design representative facial models of target individuals. Systems for still-to-video FR must therefore rely on adaptation, multiple face representation, or synthetic generation of reference stills to enhance the intra-class variability of face models . Moreover, many FR systems only match high quality faces captured in video, which further reduces the probability of detecting target individuals. Instead of matching faces captured through segmentation to reference stills, this paper exploits Adaptive Appearance Model Tracking (AAMT) to gradually learn a track-face-model for each individual appearing in the scene. The Sequential Karhunen–Loeve technique is used for online learning of these track-face-models within a particle filter-based face tracker. Meanwhile, these models are matched over successive frames against the reference still images of each target individual enrolled to the system, and then matching scores are accumulated over several frames for robust spatiotemporal recognition. A target individual is recognized if scores accumulated for a track-face-model over a fixed time surpass some decision threshold. The main advantage of AAMT over traditional still-to-video FR systems is the greater diversity of facial representation that may be captured during operations, and this can lead to better discrimination for spatiotemporal recognition. Compared to state-of-the-art adaptive biometric systems, the proposed method selects facial captures to update an individual׳s face model more reliably because it relies on information from tracking. Simulation results obtained with the Chokepoint video dataset indicate that the proposed method provides a significantly higher level of performance compared state-of-the-art systems when a single reference still per individual is available for matching. This higher level of performance is achieved when the diverse facial appearances that are captured in video through AAMT correspond to that of reference stills.
M. Ali Akber Dewan, Eric Granger, Gian Luca Marcialis, Robert Sabourin, Fabio Roli
Pattern Recognit.2
2015 Ensembles of exemplar-SVMs for video face recognition from a single sample per person
abstract
Recognizing the face of target individuals in a watch-list is among the most challenging applications in video surveillance, especially when enrollment is based on one reference still facial image. Besides the limited representativeness of facial models used for matching, the appearance of faces captured in videos varies due to changes in illumination, pose, scales, etc., and to camera inter-operability. A multi-classifier system is proposed in this paper for robust still-to-video face recognition (FR) based on multiple diverse face representations. An individual-specific ensemble of exemplar-SVMs (e-SVMs) classifiers is assigned to each target person, where each classifier is trained using a high-quality reference face still versus many lower-quality faces of non-target individuals captured in videos. Diverse face representations are generated from different patches isolated in facial images and face descriptors that are robust to various nuisance factors (e.g., illumination and pose) commonly encountered in surveillance environments. Discriminant feature subsets, training samples, and ensemble fusion functions are selected using faces of non-target individuals captured in videos of the scene. Experiments on videos from the Chokepoint dataset reveal that the proposed ensemble of e-SVMs outperforms state-of-the-art FR systems specialized for the single sample per person problem.
Saman Bashbaghi, Eric Granger, Robert Sabourin, Guillaume-Alexandre Bilodeau
AVSS2
2015 Contextual object tracker with structure encoding
abstract
Motivated by the problem of object tracking in video sequences, this paper presents a new Contextual Object Tracker with Structural Encoding (CTSE). The novelty in our tracking approach lies in the application of contextual and structural information (that is specific to a target object) into a model-free tracker. This is first achieved by including features from a complementary region having correlated motion with the target object. Second, a local structure that represents a spatial constraint between features within the target object are included. SIFT keypoints are used as features to encode both these information. The tracking is done in three steps. Firstly, keypoints are detected and described to encode object structure. Secondly, they are matched in every frame. Finally, each matched keypoint votes for the target object location locally in a voting matrix by using the encoded object structure. The voting method gives more priority to the keypoints that have been matched more often and are closest to the target's center than the rest. The proposed tracker is competitive with state-of-the art trackers while being significantly faster. It ranks as first or second most accurate tracker in experiments with standard datasets.
Tanushri Chakravorty, Guillaume-Alexandre Bilodeau, Eric Granger
ICIP3
2015 Synthetic face generation under various operational conditions in video surveillance
abstract
In still-to-video face recognition (FR), the faces captured with surveillance cameras are matched against reference stills of target individuals enrolled to the system. FR is a challenging problem in video surveillance due to uncontrolled capture conditions (variations in pose, expression, illumination, blur, scale, etc.), and the limited number of reference stills to model target individuals. This paper introduces a new approach to generate multiple synthetic face images per reference still based on camera-specific capture conditions to deal with illumination variations. For each reference still, a diverse set of faces from non-target individuals appearing in the camera viewpoint are selected based on luminance and contrast distortion. These face images are then decomposed into detail layer and large scale layer using an edge-preserving image decomposition to obtain their illumination dependent component. Finally, the large scale layers of these images are morphed with each reference still image to generate multiple synthetic reference stills that incorporate illumination and contrast conditions. Experimental results obtained with the ChokePoint dataset reveal that these synthetic faces produce an enhanced face model. As the number of synthetic faces grows, the proposed approach provides a higher level of accuracy and robustness across a range of capture conditions.
Fania Mokhayeri, Eric Granger, Guillaume-Alexandre Bilodeau
ICIP2
2015 Adaptive Classification for Person Re-identification Driven by Change Detection
Christophe Pagano, Eric Granger, Robert Sabourin, Gian Luca Marcialis, Fabio Roli
ICPRAM (1)2
2015 Adaptive skew-sensitive fusion of ensembles and their application to face re-identification
abstract
Adaptive classifier ensembles have been shown to improve the accuracy and robustness of systems for face recognition (FR) in video surveillance. However, it is often assumed that the proportions of faces captured for target and non-target individuals are balanced, or they are known a priori, and constant over time. Some active approaches have been proposed to update the ensemble during operations according to class imbalance of the input data stream. Beyond the estimation operational class imbalance, these approaches commonly generate diverse pools of classifiers by selecting balanced training data, limiting the potential diversity provided by the abundant non-target data. In this paper, a skew-sensitive ensemble is proposed to adaptively combine classifiers trained with data selected to have varying levels of imbalance and complexity. Given a face re-identification application, faces captured for each person appearing in the scene are tracked and regrouped into trajectories. During enrollment, faces in a reference trajectory are combined with those of selected non-target trajectories to generate a pool of 2-class classifiers using data with various levels of imbalance and complexity. During operations, the level of imbalance is periodically estimated by comparing input trajectories and pre-computed histograms using Hellinger distance quantification. Ensemble fusion functions are then adapted based on the imbalance and complexity of operational data. Finally, ensemble scores are accumulated over trajectories for robust spatio-temporal FR. Results obtained in experiments with synthetic data and Face in Action videos reveal that the proposed approach can significantly improve performance across operational imbalances.
Miguel De-la-Torre, Eric Granger, Robert Sabourin
IJCNN2
2015 Real-time visual play-break detection in sport events using a context descriptor
abstract
The detection of play and break segments in team sports is an essential step towards the automation of live game capture and broadcast. This paper presents a two-stage hierarchical method for play-break detection in non-edited video feeds of sport events. Unlike most existing methods, this algorithm performs action and event recognition on content, and thus does not rely on production cues of broadcast feeds. Moreover, the method does not require player tracking, can be used in real-time, and can be easily adapted to different sports. In the first stage, bag-of-words event detectors are trained to recognize key events such as line changes, face-offs and preliminary play-breaks. In the second stage, the output of the detectors along with a novel feature based on spatio-temporal interest points are used to create a context descriptor for the final decision. Experiments demonstrate the efficiency of the proposed method on real hockey game footage, achieving 90% accuracy.
Marc-André Carbonneau, Alexandre J. Raymond, Eric Granger, Ghyslain Gagnon
ISCAS3
2015 Individual-specific management of reference data in adaptive ensembles for face re-identification
abstract
During video surveillance, face re‐identification allows recognition and targeting of individuals of interest from faces captured across a network of video cameras. In such applications, face recognition is challenging because faces are captured under limited spatial and temporal constraints. In addition, facial models for recognition are commonly designed using a limited number of representative reference samples from faces captured under specific conditions, regrouped into facial trajectories. Given new reference samples (provided by an operator or through some self‐updating process), updating facial models may allow maintenance of a high level of performance over time. Although adaptive ensembles have been successfully applied to robust modelling of an individual's facial appearance, reference data samples from a trajectory must be stored for validation. In this study, a memory management strategy based on Kullback–Leiber (KL) divergence is proposed to rank and select the most relevant validation samples over time in adaptive individual‐specific ensembles. When new reference samples become available for an individual, updates to the corresponding ensemble are validated using a mixture of new and previously‐stored samples. Only the samples with the highest KL divergence are preserved in memory for future adaptations. This strategy is compared with reference classifiers using videos from the face in action data. Simulation results show that the proposed strategy tends to select discriminative samples from wolf‐like individuals for validation. It allows maintenance of a high level of performance, while reducing the number of samples per individual by up to 80%.
Miguel De-la-Torre, Eric Granger, Robert Sabourin, Dmitry O. Gorodnichy
IET Comput. Vis.2
2015 An adaptive ensemble-based system for face recognition in person re-identification
Miguel De-la-Torre, Eric Granger, Robert Sabourin, Dmitry O. Gorodnichy
Mach. Vis. Appl.2
2015 Adaptive skew-sensitive ensembles for face recognition in video surveillance
Miguel De-la-Torre, Eric Granger, Robert Sabourin, Dmitry O. Gorodnichy
Pattern Recognit.2
2014 Improving Signature-Based Biometric Cryptosystems Using Cascaded Signature Verification-Fuzzy Vault (SV-FV) Approach
abstract
Biometric cryptosystems have been applied to secure secret keys for encryption and digital signatures by means of biometric traits, e.g., Fingerprint, face, etc., where the fuzzy vault (FV) mechanism has been extensively employed. Recently, the authors proposed a FV system based on the offline signature images, so that digitized documents can be secured with the embedded handwritten signatures. However, the FV design concerns mostly with alleviating biometric variability with less focusing on its power in discriminating forgeries. Accordingly, the decoding accuracy of implementations is below the level required in practical banking transactions. On the other hand, signature verification (SV) systems have shown higher accuracy in discriminating forgeries. In this paper, accuracy of signature-based biometric cryptosystems is enhanced by cascading SV and FV modules. Signature samples are first verified by the SV module. Then, only verified samples are processed by FV decoders for unlocking cryptographic keys. Hence, the upper limit of the false accept rate is determined by the more accurate SV module. Simulation results obtained with the Brazilian signature database indicate the viability of the proposed approach. Cascaded SV-FV system increases decoding accuracy by about 35% compared to the pure FV systems.
George S. Eskander, Robert Sabourin, Eric Granger
ICFHR3
2014 Watch-List Screening Using Ensembles Based on Multiple Face Representations
abstract
Still-to-video face recognition (FR) is an important function in watch list screening, where faces captured over a network of video surveillance cameras are matched against reference stills of target individuals. Recognizing faces in a watch list is a challenging problem in semi -- and unconstrained surveillance environments due to the lack of control over capture and operational conditions, and to the limited number of reference stills. This paper provides a performance baseline and guidelines for ensemble-based systems using a single high-quality reference still per individual, as found in many watch list screening applications. In particular, modular systems are considered, where an ensemble of template matchers based on multiple face representations is assigned to each individual of interest. During enrollment, multiple feature extraction (FE) techniques are applied to patches isolated in the reference still to generate diverse face-part representations that are robust to various nuisance factors (e.g., illumination and pose) encountered in video surveillance. The selection of relevant feature subsets, decision thresholds, and fusion functions of ensembles are achieved using faces of non-target individuals selected from reference videos (forming a universal background model). During operations, a face tracker gradually regroups faces captured from different people appearing in a scene, while each user-specific ensemble generates a decision per face capture. This leads to robust spatio-temporal FR when accumulated ensemble predictions surpass a detection threshold. Simulation results obtained with the Chokepoint video dataset show a significant improvement to accuracy, (1) when performing score-level fusion of matchers, where patches-based and FE techniques generate ensemble diversity, (2) when defining feature subsets and decision thresholds for each individual matcher of an ensemble using non-target videos, and (3) when accumulating positive detections over multiple frames.
Saman Bashbaghi, Eric Granger, Robert Sabourin, Guillaume-Alexandre Bilodeau
ICPR2
2014 Self-Updating with Facial Trajectories for Video-to-Video Face Recognition
abstract
For applications of face recognition (FR) in video surveillance, it is often costly or unfeasible to collect several high quality reference samples a priori to design representative facial models. Moreover, changes in capture conditions and human physiology create divergence between facial models and input captures. Multiple classifier systems (MCS) have been successfully applied to video-to-video FR, where the face of each individual of interest is modeled using an ensemble of 2-class classifiers (trained on target vs. non-target samples). However, the reliable self-update of these individual-specific ensembles with relevant target and non-target samples raises several challenges. In this paper, an adaptive MCS is proposed that allows for self-updating facial models given face trajectories captured during operations. Different faces appearing in a camera viewpoint are tracked, and ensemble predictions for facial captures are accumulated along each track for robust video-to-video FR. When the number of positive predictions over time surpasses an update threshold, the target face samples extracted from the trajectory are combined with non-target samples selected from the cohort and universal models for efficient self-update the corresponding face model. A learn-and-combine strategy is then employed to avoid knowledge corruption during self-update of an ensemble. At a transaction level, the adaptive MCS outperforms the reference systems that do not allow self-updating on Face in Action videos. Analysis at a trajectory level indicates that the proposed system allows for robust spatio-temporal recognition, which translates to enhanced security and situation analysis.
Miguel De-la-Torre, Eric Granger, Paulo Vinicius Wolski Radtke, Robert Sabourin, Dmitry O. Gorodnichy
ICPR2
2014 A bio-cryptographic system based on offline signature images
George S. Eskander, Robert Sabourin, Eric Granger
Inf. Sci.3
2014 Adaptive ensembles for face recognition in changing video surveillance environments
Christophe Pagano, Eric Granger, Robert Sabourin, Gian Luca Marcialis, Fabio Roli
Inf. Sci.2
2014 Rapid blockwise multi-resolution clustering of facial images for intelligent watermarking
Bassem S. Rabil, Robert Sabourin, Eric Granger
Mach. Vis. Appl.3
2013 Evolving Classifier Ensembles using Dynamic Multi-objective Swarm Intelligence
Jean-François Connolly, Eric Granger, Robert Sabourin
ICPRAM2
2013 Securing high resolution grayscale facial captures using a blockwise coevolutionary GA
Bassem S. Rabil, Safa Tliba, Eric Granger, Robert Sabourin
Expert Syst. Appl.3
2013 DS-DPSO: A dual surrogate approach for intelligent watermarking of bi-tonal document image streams
Eduardo Vellasques, Robert Sabourin, Eric Granger
Expert Syst. Appl.3
2013 Multi-feature extraction and selection in writer-independent off-line signature verification
Dominique Rivard, Eric Granger, Robert Sabourin
Int. J. Document Anal. Recognit.2
2012 A comparison of adaptive matchers for screening of faces in video surveillance
abstract
Video-based face screening is essentially a detection problem where faces captured in video sequences are matched against the facial models of individuals of interest. This problem is associated with several operational challenges, from lighting and pose changes, to natural aging of target individuals, and to the limited availability of reference samples from changing environments to design facial models. Some matchers proposed in literature may be employed to adapt facial models of individuals enrolled to the system in response to new reference samples. This paper reviews and compares the performance of these matchers, focusing on their ability for adapting to new data. An experimental methodology is proposed to assess their performance for video surveillance applications. This methodology is focused on transactional and subject-based performance, and considers the imbalance of positive and negative samples. Experiments are then performed with the Canegie Mellon University Face in Action video dataset, according to matching accuracy and resource requirements. Results indicate that ensemble-based matchers outperform traditional monolithic approaches, maintaining a higher level of accuracy over time when adapting to new reference samples.
Miguel De-la-Torre, Paulo Vinicius Wolski Radtke, Eric Granger, Robert Sabourin, Dmitry O. Gorodnichy
CISDA3
2012 Gaussian mixture modeling for dynamic particle swarm optimization of recurrent problems
abstract
In dynamic optimization problems, the optima location and fitness value change over time. Techniques in literature for dynamic optimization involve tracking one or more peaks moving in a sequential manner through the parameter space. However, many practical applications in, e.g., video and image processing involve optimizing a stream of recurrent problems, subject to noise. In such cases, rather than tracking one or more moving peaks, the focus is on managing a memory of solutions along with information allowing to associate these solutions with their respective problem instances. In this paper, Gaussian Mixture Modeling (GMM) of Dynamic Particle Swarm Optimization (DPSO) solutions is proposed for fast optimization of streams of recurrent problems. In order to avoid costly re-optimizations over time, a compact density representation of previously-found DPSO solutions is created through mixture modeling in the optimization space, and stored in memory. For proof of concept simulation, the proposed hybrid GMM-DPSO technique is employed to optimize embedding parameters of a bi-tonal watermarking system on a heterogeneous database of document images. Results indicate that the computational burden of this watermarking problem is reduced by up to 90.4% with negligible impact on accuracy.
Eduardo Vellasques, Robert Sabourin, Eric Granger
GECCO3
2012 Adaptation of Writer-Independent Systems for Offline Signature Verification
abstract
Although writer-independent offline signature verification (WI-SV) systems may provide a high level of accuracy, they are not secure due to the need to store user templates for authentication. Moreover, state-of-the-art writer-dependent (WD) and writer-independent (WI) systems provide enhanced accuracy through information fusion at either feature, score or decision levels, but they increase computational complexity. In this paper, a method for adapting WI-SV systems to different users is proposed, leading to secure and compact WD-SV systems. Feature representations embedded within WI classifiers are extracted and tuned to each enrolled user while building a user-specific classifier. Simulation results on the Brazilian signature database indicate that the proposed method yields WD classifiers that provide the same level of accuracy as that of the baseline WI classifiers (AER of about 5.38), while reducing complexity by about 99.5%.
George S. Eskander, Robert Sabourin, Eric Granger
ICFHR3
2012 On the correlation between genotype and classifier diversity
Jean-François Connolly, Eric Granger, Robert Sabourin
ICPR2
2012 Adaptive selection of ensembles for imbalanced class distributions
Paulo Vinicius Wolski Radtke, Eric Granger, Robert Sabourin, Dmitry O. Gorodnichy
ICPR2
2012 A dual-staged classification-selection approach for automated update of biometric templates
Ajita Rattani, Gian Luca Marcialis, Eric Granger, Fabio Roli
ICPR3
2012 Incremental update of biometric models in face-based video surveillance
abstract
Video-based face recognition of individuals involves matching facial regions captured in video sequences against the model of individuals enrolled to a face recognition system. Due to a limited control over operational conditions, classification systems applied to face matching are confronted with complex pattern recognition environments that change over time. Therefore, the facial model of an individual tends to diverge from the underlying data distribution. Although a limited amount of reference data is often collected during initial enrollment, new samples often become available over time to update and refine models. In this paper, an adaptive ensemble of classifiers is proposed to update facial models in response to new reference samples. To avoid knowledge corruption linked to incremental learning of monolithic classifiers, and maintain a high level of performance, this ensemble exploits a learn-and-combine approach. In response to new reference samples, a new 2-class Probabilistic Fuzzy ARTMAP classifier is trained and combined to previously-trained classifiers in the ROC space. Iterative Boolean Combination is employed for fusion of 2-class classifiers of each individual in the decision space. Performance is assessed in terms of AUC accuracy and resource requirements under different incremental learning scenarios with new data extracted from the Faces in Action data set. Simulation results indicate that the proposed system significantly outperforms reference classifiers and ensembles for incremental learning.
Miguel De-la-Torre, Eric Granger, Paulo Vinicius Wolski Radtke, Robert Sabourin, Dmitry O. Gorodnichy
IJCNN2
2012 Detector ensembles for face recognition in video surveillance
abstract
Biometric systems for recognizing faces in video streams have become relevant in a growing number of private and public sector applications, among them screening for individuals of interest in dense and moving crowds. In practice, the performance of these systems typically declines because they encounter a variety of uncontrolled conditions that change during operations, and they are designed a priori using limited data and knowledge of underlying data distributions. This paper presents multi-classifier system that can achieve a high level of performance in real-world video surveillance applications. This system assigns an ensemble of detectors (2-class classifiers) per individual, where base detectors are co-jointly trained using population-based evolutionary optimization. During enrolment of an individual, an aggregative Dynamic Niching Particle Swarm Optimization (DNPSO)-based training strategy generates a diversified homogenous pool of ARTMAP neural network classifiers using reference data samples. Classifiers associated with local optima of the aggregative DNPSO are directly selected and efficiently combined in the Receiver Operating Characteristic (ROC) space. Performance is assessed in terms of both accuracy and resource requirements on facial regions extracted from video streams of the Face in Action database. A comparison between a standard global and modular classification architectures is provided in this paper. Simulation results indicate that recognizing an individual using the aforementioned ensemble of detectors provides a scalable architecture that maintains a significantly higher level of accuracy and robustness as the number of individuals grows.
Christophe Pagano, Eric Granger, Robert Sabourin, Dmitry O. Gorodnichy
IJCNN2
2012 An adaptive classification system for video-based face recognition
Jean-François Connolly, Eric Granger, Robert Sabourin
Inf. Sci.2
2012 A survey of techniques for incremental learning of HMM parameters
Wael Khreich, Eric Granger, Ali Miri, Robert Sabourin
Inf. Sci.2
2012 Dynamic selection of generative-discriminative ensembles for off-line signature verification
Luana Batista, Eric Granger, Robert Sabourin
Pattern Recognit.2
2012 Evolution of heterogeneous ensembles through dynamic particle swarm optimization for video-based face recognition
Jean-François Connolly, Eric Granger, Robert Sabourin
Pattern Recognit.2
2012 Adaptive ROC-based ensembles of HMMs applied to anomaly detection
Wael Khreich, Eric Granger, Ali Miri, Robert Sabourin
Pattern Recognit.2
2010 An adaptive ensemble of fuzzy ARTMAP neural networks for video-based face classification
abstract
A key feature in population based optimization algorithms is the ability to explore a search space and make a decision based on multiple solutions. In this paper, an incremental learning strategy based on a dynamic particle swarm optimization (DPSO) algorithm allows to produce heterogeneous ensembles of classifiers for video-based face recognition. This strategy is applied to an adaptive classification system (ACS) comprised of a swarm of fuzzy ARTMAP (FAM) neural network classifiers, a DPSO algorithm, and a long term memory (LTM). The performance of this ACS with an ensemble of FAM networks selected among local bests of the swarm, is compared to that of the ACS with the global best network under different incremental learning scenarios. Performance is assessed in terms of classification rate and resource requirements for incremental learning of new data blocks extracted from real-world video streams, and are given along with reference kNN and FAM classifier optimized for batch learning. Simulation results indicate that the learning strategy maintains diversity within the ensemble classifiers, providing a significantly higher classification rate than that of the best FAM network alone. However, classification with an ensemble requires more resources.
Jean-François Connolly, Eric Granger, Robert Sabourin
IEEE Congress on Evolutionary Computation2
2010 Evolving ARTMAP neural networks using Multi-Objective Particle Swarm Optimization
abstract
In this paper, a supervised learning strategy based on a Multi-Objective Particle Swarm Optimization (MOPSO) is introduced for ARTMAP neural networks. It is based on the concept of neural network evolution in that particles of a MOPSO swarm (i.e., network solutions) seek to determine user-defined parameters and network (weights and architecture) such that generalisation error and network resources are minimized. The performance of this strategy has been assessed with fuzzy ARTMAP using synthetic and real-world data for video-based face classification. Simulation results indicate that when the MOPSO strategy is used to train fuzzy ARTMAP, it produces a significantly lower classification error than when trained using standard hyper-parameter settings. Furthermore, the non-dominated MOPSO solutions represent a better compromise between error and resource allocation than mono-objective PSO-based strategies that minimizes only classification error. Overall, results obtained with the MOPSO strategy reveal the importance of optimizing parameters and network for each problem, where both error and resources are minimized during fitness evaluation.
Eric Granger, Donavan Prieur, Jean-François Connolly
IEEE Congress on Evolutionary Computation1
2010 Applying Dissimilarity Representation to Off-Line Signature Verification
abstract
In this paper, a two-stage off-line signature verification system based on dissimilarity representation is proposed. In the first stage, a set of discrete left-to-right HMMs trained with different number of states and codebook sizes is used to measure similarity values that populate new feature vectors. Then, these vectors are input to the second stage, which provides the final classification. Experiments were performed using two different classification techniques - AdaBoost, and Random Subspaces with SVMs - and a real-world signature verification database. Results indicate that the performance is significantly better with the proposed system over other reference signature verification systems from literature.
Luana Batista, Eric Granger, Robert Sabourin
ICPR2
2010 Boolean Combination of Classifiers in the ROC Space
abstract
Using Boolean AND and OR functions to combine the responses of multiple one- or two-class classifiers in the ROC space may significantly improve performance of a detection system over a single best classifier. However, techniques found in literature assume that the classifiers are conditionally independent, and that their ROC curves are convex. These assumptions are not valid in most real-world applications, where classifiers are designed using limited and imbalanced training data. A new Iterative Boolean Combination (IBC) technique applies all Boolean functions to combine the ROC curves produced by multiple classifiers without prior assumptions, and its time complexity is linear according to the number of classifiers. The results of computer simulations conducted on synthetic and real-world host-based intrusion detection data indicate that combining the responses from multiple HMMs with IBC can achieve a significantly higher level of performance than with the AND and OR combinations, especially when training data is limited and imbalanced.
Wael Khreich, Eric Granger, Ali Miri, Robert Sabourin
ICPR2
2010 Improving performance of HMM-based off-line signature verification systems through a multi-hypothesis approach
Luana Batista, Eric Granger, Robert Sabourin
Int. J. Document Anal. Recognit.2
2010 Iterative Boolean combination of classifiers in the ROC space: An application to anomaly detection with HMMs
Wael Khreich, Eric Granger, Ali Miri, Robert Sabourin
Pattern Recognit.2
2010 On the memory complexity of the forward-backward algorithm
Wael Khreich, Eric Granger, Ali Miri, Robert Sabourin
Pattern Recognit. Lett.2
2009 Incremental adaptation of fuzzy ARTMAP neural networks for video-based face classification
abstract
In many practical applications, new training data is acquired at different points in time, after a classification system has originally been trained. For instance, in face recognition systems, new training data may become available to enroll or to update knowledge of an individual. In this paper, a neural network classifier applied to video-based face recognition is adapted through supervised incremental learning of real-world video data. A training strategy based on particle swarm optimization is employed to co-optimize the weights, architecture and hyperparameters of the fuzzy ARTMAP network during incremental learning of new data. The performance of fuzzy ARTMAP is compared under different class update scenarios when incremental learning is performed according to 3 cases-(A) hyperparameters set to standard values, (B) hyperparameters optimized only at the beginning of the learning process with all classes, and (C) hyperparameters re-optimized whenever new training data becomes available. Overall results indicate that when samples from each individual enrolled to the system are employed for optimization, a higher classification rate is achieved and the solutions produced are more robust to variations caused by pattern presentation order. When all classes are refined equally, this is true with incremental learning according to case (C), whereas, if one class is refined at a time, best performance is obtained with case (B). However, optimizing hyperparameters requires more resources: several training sequences are needed to find the optimal solution and fuzzy ARTMAP with hyperparameters optimized according to classification rate tends to generate a high number of category nodes over longer convergence time.
Jean-François Connolly, Eric Granger, Robert Sabourin
CISDA2
2009 A comparison of techniques for on-line incremental learning of HMM parameters in anomaly detection
abstract
Hidden Markov Models (HMMs) have been shown to provide a high level performance for detecting anomalies in intrusion detection systems. Since incomplete training data is always employed in practice, and environments being monitored are susceptible to changes, a system for anomaly detection should update its HMM parameters in response to new training data from the environment. Several techniques have been proposed in literature for on-line learning of HMM parameters. However, the theoretical convergence of these algorithms is based on an infinite stream of data for optimal performances. When learning sequences with a finite length, on-line incremental versions of these algorithms can improve discrimination by allowing for convergence over several training iterations. In this paper, the performance of these techniques is compared for learning new sequences of training data in host-based intrusion detection. The discrimination of HMMs trained with different techniques is assessed from data corresponding to sequences of system calls to the operating system kernel. In addition, the resource requirements are assessed through an analysis of time and memory complexity. Results suggest that the techniques for online incremental learning of HMM parameters can provide a higher level of discrimination than those for on-line learning, yet require significantly fewer resources than with batch training. On-line incremental learning techniques may provide a promising solution for adaptive intrusion detection systems.
Wael Khreich, Eric Granger, Ali Miri, Robert Sabourin
CISDA2
2009 Combining Hidden Markov Models for Improved Anomaly Detection
abstract
In host-based intrusion detection systems (HIDS), anomaly detection involves monitoring for significant deviations from normal system behavior. Hidden Markov Models (HMMs) have been shown to provide a high level performance for detecting anomalies in sequences of system calls to the operating system kernel. Although the number of hidden states is a critical parameter for HMM performance, it is often chosen heuristically or empirically, by selecting the single value that provides the best performance on training data. However, this single best HMM does not typically provide a high level of performance over the entire detection space. This paper presents a multiple-HMMs approach, where each HMM is trained using a different number of hidden states, and where HMM responses are combined in the receiver operating characteristics (ROC) space according to the maximum realizable ROC (MRROC) technique. The performance of this approach is compared favorably to that of a single best HMM and to a traditional sequence matching technique called STIDE, using different synthetic HIDS data sets. Results indicate that this approach provides a higher level of performance over a wide range of training set sizes with various alphabet sizes and irregularity indices, and different anomaly sizes, without a significant computational and storage overhead.
Wael Khreich, Eric Granger, Robert Sabourin, Ali Miri
ICC2
2009 A Multi-Hypothesis Approach for Off-Line Signature Verification with HMMs
abstract
In this paper, an approach based on the combination of discrete Hidden Markov Models (HMMs) in the ROC space is proposed to improve the performance of off-line signature verification (SV) systems designed from limited and unbalanced training data. This approach is inspired by the multiple-hypothesis principle, and allows the system to choose, from a set of different HMMs, the most suitable solution for a given input sample. By training an ensemble of user-specific HMMs with different number of states, and then combining these models in the ROC space, it is pos-sible to construct a composite ROC curve that provides a more accurate estimation of system’s performance during training and significantly reduces the error rates during op-erations. The experiments performed by using a real-world SV database with random, simple and skilled forgeries, in-dicated that the proposed approach can reduce the average error rates by more than 17%. 1
Luana Batista, Eric Granger, Robert Sabourin
ICDAR2
2008 A comparison of fuzzy ARTMAP and Gaussian ARTMAP neural networks for incremental learning
abstract
Automatic pattern classifiers that allow for incremental learning can adapt internal class models efficiently in response to new information, without having to retrain from the start using all the cumulative training data. In this paper, the performance of two such classifiers - the fuzzy ARTMAP and Gaussian ARTMAP neural networks - are characterize and compared for supervised incremental learning in environments where class distributions are fixed. Their potential for incremental learning of new blocks of training data, after previously been trained, is assessed in terms of generalization error and resource requirements, for several synthetic pattern recognition problems. The advantages and drawbacks of these architectures are discussed for incremental learning with different data block sizes and data set structures. Overall results indicate that Gaussian ARTMAP is the more suitable for incremental learning as it usually provides an error rate that is comparable to that of batch learning for the data sets, and for a wide range of training block sizes. The better performance is a result of the representation of categories as Gaussian distributions, and of using category-specific learning rate that decreases during the training process. With all the data sets, the error rate obtained by training through incremental learning is usually significantly higher than through batch learning for fuzzy ARTMAP. Training fuzzy ARTMAP and Gaussian ARTMAP through incremental learning often requires fewer training epochs to converge, and leads to more compact networks.
Eric Granger, Jean-François Connolly, Robert Sabourin
IJCNN1
2007 Incremental Learning of Stochastic Grammars with Graphical EM in Radar Electronic Support
abstract
Although stochastic context-free grammars (SCFGs) appear promising for recognition of radar emitters, and for estimation of their level of threat in radar electronic support (ES) systems, well-known techniques for learning their production rule probabilities are computationally demanding, and cannot efficiently reflect changes in operational environments. Some techniques have been proposed for fast learning of SCFGs probabilities, yet, of those, only the HOLA technique can perform learning incrementally. In this paper, two incremental versions of the graphical EM (gEM) technique are proposed. The incremental gEM (igEM) and on-line incremental gEM (oigEM) allow for adapting production rule probabilities from new data, without having to retrain from the start on all accumulated training data. These new techniques are compared to HOLA using radar signal data. An experimental protocol has been defined such that the impact on performance of factors like the size of new data blocks for incremental learning, and the level of ambiguity of MFR grammars, may be observed. Results indicate that, contrary to HOLA, incremental learning of training data blocks with igEM and oigEM provides the same level of accuracy as learning from all cumulative data from scratch, even for small data blocks. As expected, incremental learning significantly reduces the overall time and memory complexities. Finally, it appears that while the computational complexity and memory requirements of igEM and oigEM may be greater than that of HOLA, they both provide a higher level of accuracy.
Guillaume Latombe, Eric Granger, Fred A. Dilkes
ICASSP (2)2
2007 Face Recognition in Video Using a What-and-Where Fusion Neural Network
abstract
A what-and-where fusion neural network is applied to the recognition of human faces from video sequences. The spatio-temporal information contained in successive video frames allows to effectively accumulate a classifier's predictions for each person being tracked in an environment. In a particular realization of this network, a fuzzy ARTMAP neural network is used to classify faces detected in each frame, while a bank of Kalman filters is used to track blobs that contain the extracted faces moving in the environment. Performance of the what-and-where fusion neural network is compared to that of the fuzzy ARTMAP and k-nearest-neighbor (k-NN) classifiers in terms of classification rate, convergence time and compression. In this paper, the impact on performance of setting different region of interest (ROI), of optimizing fuzzy ARTMAP parameters, and of selecting different training subset sizes, is assessed. Simulation results on real-world video sequences indicate that this network can achieve a classification rate that is significantly higher (by approximately 50% in some cases) than that of fuzzy ARTMAP alone, and than that of the k-NN.
Mamoudou Barry, Eric Granger
IJCNN2
2006 Fast Incremental Techniques for Learning Production Rule Probabilities in Radar Electronic Support
abstract
Although Stochastic context-free grammars appear promising for recognition of radar emitters, and for estimation of their respective level of threat in radar electronic support systems, well-known techniques for learning their production rule probabilities are computationally demanding. In this paper, three fast incremental alternatives, called graphical EM (gEM), tree scanning (TS), and HOLA, are compared from several perspectives - perplexity, generalization error, time and space complexity, and convergence time. Estimation of the execution time and storage requirements allows for the assessment of complexity, while computer simulation using a radar pulse data set allows to asses the other performance measures. Results indicate that gEM and TS may provide a greater level of accuracy than HOLA, and that computational complexity may be orders of magnitude lower with HOLA. Furthermore, HOLA is an on-line technique that allows for incremental learning of probabilities to reflect changes in operational environments
Guillaume Latombe, Eric Granger, Fred A. Dilkes
ICASSP (5)2
2006 Particle Swarm Optimization of Fuzzy ARTMAP Parameters
abstract
In this paper a particle swarm optimization (PSO)-based training strategy is introduced for fuzzy ARTMAP that minimizes generalization error while optimizing parameter values. Through a comprehensive set simulations, it has been shown that this training strategy allows fuzzy ARTMAP to achieve a significantly lower generalization error than when it uses typical training strategies. Furthermore, the PSO strategy eliminates degradation of generalization error due to overtraining resulting from the training set size, number of training epochs, and data set structure. Overall results obtained with the PSO strategy reveal the importance of optimizing parameters and weights using a consistent objective function. In fact, the parameters found using this strategy vary significantly according to, e.g., training set size and data set structure, and always differ considerably from the popular choice of parameters that allows to minimize resources.
Eric Granger, Philippe Henniges, Luiz Eduardo Soares de Oliveira, Robert Sabourin
IJCNN1
2005 Factors of overtraining with fuzzy ARTMAP neural networks
abstract
In this paper, the impact of overtraining on the performance of fuzzy ARTMAP neural networks is assessed for pattern recognition problems consisting of overlapping class distributions, and consisting of complex decision boundaries with no overlap. Computer simulations are performed with fuzzy ARTMAP networks trained for one epoch, through cross-validation, and until network convergence, using several data sets representing these pattern recognition problems. By comparing the generalisation error and resources required by these networks, the extent of overtraining due to factors such as data set structure, training strategy, number of training epochs, data normalisation, and training set size, is demonstrated. A significant degradation in fuzzy ARTMAP performance due to overtraining is shown to depend on the training set size and the number of training epochs for pattern recognition problems with overlapping class distributions.
Philippe Henniges, Eric Granger, Robert Sabourin
IJCNN2
2003 A Pattern Reordering Approach Based on Ambiguity Detection for Online Category Learning
abstract
Pattern reordering is proposed as an alternative to sequential and batch processing for online category learning. Upon detecting that the categorization of a new input pattern is ambiguous, the input is postponed for a predefined time, after which it is reexamined and categorized for good. This approach is shown to improve the categorization performance over purely sequential processing, while yielding a shorter input response time, or latency, than batch processing. In order to examine the response time of processing schemes, the latency of a typical implementation is derived and compared to lower bounds. Gaussian and softmax models are derived from reject option theory and are considered for detecting ambiguity and triggering pattern postponement. The average latency and Rand Adjusted clustering score of reordered, sequential, and batch processing are compared through computer simulation using two unsupervised competitive learning neural networks and a radar pulse data set.
Eric Granger, Yvon Savaria, Pierre Lavoie
IEEE Trans. Pattern Anal. Mach. Intell.1
2001 A What-and-Where fusion neural network for recognition and tracking of multiple radar emitters
Eric Granger, Mark A. Rubin, Stephen Grossberg, Pierre Lavoie
Neural Networks1
2000 Classification of Incomplete Data Using the Fuzzy ARTMAP Neural Network
abstract
The fuzzy ARTMAP neural network is used to classify data that is incomplete in one or more ways. These include a limited number of training cases, missing components, missing class labels, and missing classes. Modifications for dealing with such incomplete data are introduced, and performance is assessed on an emitter identification task using a database of radar pulses.
Eric Granger, Mark A. Rubin, Stephen Grossberg, Pierre Lavoie
IJCNN (6)1
2000 Analysis of quantization effects in a digital hardware implementation of a fuzzy ART neural network algorithm
abstract
A reformulated Adaptive Resonance Theory (ART) neural network algorithm has recently been implemented in digital hardware. Naturally, the fixed point, fixed word length data format used causes some output differences with respect to floating point computer simulation. These differences are observed when using realistic input data. The effects of input quantization and the accumulation of round off errors in the arithmetic operations making up the algorithm are analyzed. Even a small quantization or round off error can trigger a change in the clustering produced. This does not mean that the clustering is not valid. Indeed, the validity of the clustering can be comparable to that obtained by floating point computer simulation, provided the word length is sufficient. This is verified on realistic input data consisting of radar pulses received from a number of emitters.
Marc-André Cantin, Yves Blaquière, Yvon Savaria, Pierre Lavoie, Eric Granger
ISCAS5
1998 Familiarity Discrimination of Radar Pulses
Eric Granger, Stephen Grossberg, Mark A. Rubin, William W. Streilein
NIPS1
1998 A comparison of self-organizing neural networks for fast clustering of radar pulses
Eric Granger, Yvon Savaria, Pierre Lavoie, Marc-André Cantin
Signal Process.1