VLDB 2026 Research / reviewers in the wild / expert
Ismail Ben Ayed
dblp:68/4478
· DBLP profile ↗
147ranked-venue papers
21as first author
77since 2021 · last 2026
0000-0002-9668-8027ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 98 · 15 first-author · 50 since 2021Artificial intelligence and machine learning · 70 · 9 first-author · 42 since 2021Applied, interdisciplinary, general and emerging computing · 43 · 6 first-author · 18 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Histopath-C: Towards Realistic Domain Shifts for Histopathology Vision-Language AdaptationabstractMedical Vision-language models (VLMs) have shown remarkable performances in various medical imaging domains such as histopathology by leveraging pre-trained, contrastive models that exploit visual and textual information. However, histopathology images may exhibit severe domain shifts, such as staining, contamination, blurring, and noise, which may severely degrade the VLM’s downstream performance. In this work, we introduce Histopath-C, a new benchmark with realistic synthetic corruptions designed to mimic real-world distribution shifts observed in digital histopathology. Our framework dynamically applies corruptions to any available dataset and evaluates Test-Time Adaptation (TTA) mechanisms on the fly. We then propose LATTE, a transductive, low-rank adaptation strategy that exploits multiple text templates, mitigating the sensitivity of histopathology VLMs to diverse text inputs. Our approach outperforms state-of-the-art TTA methods originally designed for natural images across a breadth of histopathology datasets, demonstrating the effectiveness of our proposed design for robust adaptation in histopathology images. Code and data are available at https://github.com/Mehrdad-Noori/Histopath-C. Mehrdad Noori, Gustavo Adolfo Vargas Hakim, David Osowiechi, Fereshteh Shakeri, Ali Bahri, Moslem Yazdanpanah, Sahar Dastani, Ismail Ben Ayed, Christian Desrosiers |
WACV | 8 |
| 2026 | Revisiting Layer Normalization for Point Cloud Test Time AdaptationabstractWe analyze Layer Normalization (LN) from a domain (batch) perspective and explain why BatchNorm-style test-time fixes often fail on Transformer backbones. As feature dimension and batch size grow, the per-feature batch marginals after LN’s pre-affine step concentrate at mean ≈ 0 and variance ≈ 1, making cross-batch re-standardization unnecessary and often harmful. This yields a simple rule: keep the pre-affine LN intact and adjust only the post-affine mean and gain. We instantiate this with LN-TTA, a backpropagation-free and source-free, test-time adaptation that performs a single forward pass and uniformly reparameterizes each LN layer. On three corrupted 3D point-cloud suites (ScanObjectNN-C, ModelNet40-C, ShapeNet-C), LN-TTA improves over Source-Only by +12.35, +15.58, and +3.03 points, surpasses backpropagation baselines (e.g., TENT), and sustains up to 93 samples/s, on average 39× faster and 5× more memory-efficient than the next-best backprop-free method. Code is available at: github.com/MosyMosy/LN_TTA. Moslem Yazdanpanah, Ali Bahri, Mehrdad Noori, Sahar Dastani, Samuel Barbeau, David Osowiechi, Gustavo Adolfo Vargas Hakim, Ismail Ben Ayed, Christian Desrosiers |
WACV | 8 |
| 2025 | AttackBench: Evaluating Gradient-based Attacks for Adversarial ExamplesabstractWhile novel gradient-based attacks are continuously proposed to improve the optimization of adversarial examples, each is shown to outperform its predecessors using different experimental setups, implementations, and computational budgets, leading to biased and unfair comparisons. In this work, we overcome this issue by proposing AttackBench, i.e., an attack evaluation framework that evaluates the effectiveness of each attack (along with its different library implementations) under the same maximum available computational budget. To this end, we (i) define a novel optimality metric that quantifies how close each attack is to the optimal solution (empirically estimated by ensembling all attacks), and (ii) limit the maximum number of forward and backward queries that each attack can execute on the target model. Our extensive experimental analysis compares more than 100 attack implementations over 800 different configurations, considering both CIFAR-10 and ImageNet models, and shows that only a few attack implementations outperform all the remaining approaches. These findings suggest that novel defenses should be evaluated against different attacks than those normally used in the literature to avoid overly-optimistic robustness evaluations. We release AttackBench as a publicly-available benchmark that will be continuously updated with new attack implementations to maintain an up-to-date ranking of the best gradient-based attacks. We release AttackBench as a publicly available benchmark, including a continuously updated leaderboard and source code to maintain an up-to-date ranking of the best gradient-based attacks. Antonio Emanuele Cinà, Jérôme Rony, Maura Pintor, Luca Demetrio, Ambra Demontis, Battista Biggio, Ismail Ben Ayed, Fabio Roli |
AAAI | 7 |
| 2025 | Spectral Informed Mamba for Robust Point Cloud ProcessingabstractState Space Models (SSMs) have shown significant promise in Natural Language Processing (NLP) and, more recently, computer vision. This paper introduces a new methodology that leverages Mamba and Masked Autoencoder (MAE) networks for point-cloud data in both supervised and self-supervised learning. We propose three key contributions to enhance Mamba’s capability in processing complex point-cloud structures. First, we exploit the spectrum of a graph Laplacian to capture patch connectivity, defining an isometry-invariant traversal order that is robust to viewpoints and captures shape manifolds better than traditional 3D grid-based traversals. Second, we adapt segmentation via a recursive patch partitioning strategy informed by Laplacian spectral components, allowing finer integration and segment analysis. Third, we address token placement in MAE for Mamba by restoring tokens to their original positions, which preserves essential order and improves learning. Extensive experiments demonstrate the improvements that our approach brings over state-of-the-art baselines in classification, segmentation, and few-shot tasks. The implementation is available at: https://github.com/AliBahri94/SI-Mamba.git. Ali Bahri, Moslem Yazdanpanah, Mehrdad Noori, Sahar Dastani, Milad Cheraghalikhani, Gustavo Adolfo Vargas Hakim, David Osowiechi, Farzad Beizaee, Ismail Ben Ayed, Christian Desrosiers |
CVPR | 9 |
| 2025 | Conformal Prediction for Zero-Shot ModelsabstractVision-Language models pre-trained at large scale have shown unprecedented adaptability and generalization to downstream tasks. Although its discriminative potential has been widely explored, its reliability and uncertainty are still overlooked. In this work, we investigate the capabilities of CLIP models under the split conformal prediction paradigm, which provides theoretical guarantees to black-box models based on a small, labeled calibration set. In contrast to the main body of literature on conformal predictors in vision classifiers, foundation models exhibit a particular characteristic: they are pre-trained on a one-time basis on an inaccessible source domain, different from the transferred task. This domain drift negatively affects the efficiency of the conformal sets and poses additional challenges. To alleviate this issue, we propose Conf-Ot, a transfer learning setting that operates transductive over the combined calibration and query sets. Solving an optimal transport problem, the proposed method bridges the domain gap between pre-training and adaptation without requiring additional data splits but still maintaining coverage guarantees. We comprehensively explore this conformal prediction strategy on a broad span of 15 datasets and three nonconformity scores. Conf-Ot provides consistent relative improvements of up to 20% on set efficiency while being ×15 faster than popular transductive approaches. We make the code available1. Julio Silva-Rodríguez, Ismail Ben Ayed, Jose Dolz |
CVPR | 2 |
| 2025 | Realistic Test-Time Adaptation of Vision-Language ModelsabstractThe zero-shot capabilities of Vision-Language Models (VLMs) have been widely leveraged to improve predictive performance. However, previous works on transductive or test-time adaptation (TTA) often make strong assumptions about the data distribution, such as the presence of all classes. Our work challenges these favorable deployment scenarios and introduces a more realistic evaluation framework, including (i) a variable number of effective classes for adaptation within a single batch, and (ii) non-i.i.d. batches of test samples in online adaptation settings. We provide comprehensive evaluations, comparisons, and ablation studies that demonstrate how current transductive or TTA methods for VLMs systematically compromise the models’ initial zero-shot robustness across various realistic scenarios, favoring performance gains under advantageous assumptions about the test sample distributions. Furthermore, we introduce StatA, a versatile method that can handle a wide range of deployment scenarios, including those with a variable number of effective classes at test time. Our approach incorporates a novel regularization term designed specifically for VLMs, which acts as a statistical anchor preserving the initial text-encoder knowledge, particularly in low-data regimes. Code available at https://github.com/MaxZanella/StatA. Maxime Zanella, Clément Fuchs, Christophe De Vleeschouwer, Ismail Ben Ayed |
CVPR | 4 |
| 2025 | UNEM: UNrolled Generalized EM for Transductive Few-Shot LearningabstractTransductive few-shot learning has recently triggered wide attention in computer vision. Yet, current methods introduce key hyper-parameters, which control the prediction statistics of the test batches, such as the level of class balance, affecting performances significantly. Such hyper-parameters are empirically grid-searched over validation data, and their configurations may vary substantially with the target dataset and pre-training model, making such empirical searches both sub-optimal and computationally intractable. In this work, we advocate and introduce the unrolling paradigm, also referred to as "learning to optimize", in the context of few-shot learning, thereby learning efficiently and effectively a set of optimized hyperparameters. Specifically, we unroll a generalization of the ubiquitous Expectation-Maximization (EM) optimizer into a neural network architecture, mapping each of its iterates to a layer and learning a set of key hyper-parameters over validation data. Our unrolling approach covers various statistical feature distributions and pre-training paradigms, including recent foundational vision-language models and standard vision-only classifiers. We report comprehensive experiments, which cover a breadth of fine-grained downstream image classification tasks, showing significant gains brought by the proposed unrolled EM algorithm over iterative variants. The achieved improvements reach up to 10% and 7.5% on vision-only and vision-language benchmarks, respectively. The source code and learned parameters are available at https://github.com/ZhouLong0/UNEM-Transductive. Fereshteh Shakeri, Aymen Sadraoui, Mounir Kaaniche, Jean-Christophe Pesquet, Ismail Ben Ayed |
CVPR | 6 |
| 2025 | Enhancing Remote Sensing Vision-Language Models for Zero-Shot Scene Classificationabstractpeer reviewed Karim El Khoury, Maxime Zanella, Benoît Gérin, Tiffanie Godelaine, Benoît Macq, Saïd Mahmoudi, Christophe De Vleeschouwer, Ismail Ben Ayed |
ICASSP | 8 |
| 2025 | ViLU: Learning Vision-Language Uncertainties for Failure PredictionabstractReliable Uncertainty Quantification (UQ) and failure prediction remain open challenges for Vision-Language Models (VLMs). We introduce ViLU, a new Vision-Language Uncertainty quantification framework that contextualizes uncertainty estimates by leveraging all task-relevant textual representations. ViLU constructs an uncertainty-aware multi-modal representation by integrating the visual embedding, the predicted textual embedding, and an image-conditioned textual representation via cross-attention. Unlike traditional UQ methods based on loss prediction, ViLU trains an uncertainty predictor as a binary classifier to distinguish correct from incorrect predictions using a weighted binary cross-entropy loss, making it loss-agnostic. In particular, our proposed approach is well-suited for post-hoc settings, where only vision and text embeddings are available without direct access to the model itself. Extensive experiments on diverse datasets show the significant gains of our method compared to state-of-the-art failure prediction methods. We apply our method to standard classification datasets, such as ImageNet-1k, as well as large-scale image-caption datasets like CC12M and LAION-400M. Ablation studies highlight the critical role of our architecture and training in achieving effective uncertainty quantification. Our code is publicly available and can be found here: https://github.com/ykrmm/ViLU. Marc Lafon, Yannis Karmim, Julio Silva-Rodríguez, Paul Couairon, Clément Rambour, Raphaël Fournier-S'niehotta, Ismail Ben Ayed, Jose Dolz, Nicolas Thome |
ICCV | 7 |
| 2025 | Sparsity Outperforms Low-Rank Projections in Few-Shot AdaptationabstractAdapting Vision-Language Models (VLMs) to new domains with few labeled samples remains a significant challenge due to severe overfitting and computational constraints. State-of-the-art solutions, such as low-rank reparameterization, mitigate these issues but often struggle with generalization and require extensive hyperparameter tuning. In this paper, a novel Sparse Optimization (SO) framework is proposed. Unlike low-rank approaches that typically constrain updates to a fixed subspace, our SO method leverages high sparsity to dynamically adjust very few parameters. We introduce two key paradigms. First, we advocate for \textit{local sparsity and global density}, which updates a minimal subset of parameters per iteration while maintaining overall model expressiveness. As a second paradigm, we advocate for \textit{local randomness and global importance}, which sparsifies the gradient using random selection while pruning the first moment based on importance. This combination significantly mitigates overfitting and ensures stable adaptation in low-data regimes. Extensive experiments on 11 diverse datasets show that SO achieves state-of-the-art few-shot adaptation performance while reducing memory overhead. Nairouz Mrabah, Nicolas Richet, Ismail Ben Ayed, Eric Granger |
ICCV | 3 |
| 2025 | Purge-Gate: Backpropagation-Free Test-Time Adaptation for Point Clouds Classification via Token PurgingabstractTest-time adaptation (TTA) is crucial for mitigating performance degradation caused by distribution shifts in 3D point cloud classification. In this work, we introduce Token Purging (PG), a novel backpropagation-free approach that removes tokens highly affected by domain shifts before they reach attention layers. Unlike existing TTA methods, PG operates at the token level, ensuring robust adaptation without iterative updates. We propose two variants: PG-SP, which leverages source statistics, and PG-SF, a fully source-free version relying on CLS-token-driven adaptation. Extensive evaluations on ModelNet40-C, ShapeNet-C, and ScanObjectNN-C demonstrate that PG-SP achieves an average of +10.3\% higher accuracy than state-of-the-art backpropagation-free methods, while PG-SF sets new benchmarks for source-free adaptation. Moreover, PG is 12.4 times faster and 5.5 times more memory efficient than our baseline, making it suitable for real-world deployment. Code is available at \hyperlink{https://github.com/MosyMosy/Purge-Gate}{https://github.com/MosyMosy/Purge-Gate} Moslem Yazdanpanah, Ali Bahri, Mehrdad Noori, Sahar Dastani, Gustavo Adolfo Vargas Hakim, David Osowiechi, Ismail Ben Ayed, Christian Desrosiers |
ICCV | 7 |
| 2025 | SMART-PC: Skeletal Model Adaptation for Robust Test-Time Training in Point CloudsabstractTest-Time Training has emerged as a promising solution to address distribution shifts in 3D point cloud classification. However, existing methods often rely on computationally expensive backpropagation during adaptation, limiting their applicability in real-world, time-sensitive scenarios. In this paper, we introduce SMART-PC, a skeleton-based framework that enhances resilience to corruptions by leveraging the geometric structure of 3D point clouds. During pre-training, our method predicts skeletal representations, enabling the model to extract robust and meaningful geometric features that are less sensitive to corruptions, thereby improving adaptability to test-time distribution shifts.
Unlike prior approaches, SMART-PC achieves real-time adaptation by eliminating backpropagation and updating only BatchNorm statistics, resulting in a lightweight and efficient framework capable of achieving high frame-per-second rates while maintaining superior classification performance. Extensive experiments on benchmark datasets, including ModelNet40-C, ShapeNet-C, and ScanObjectNN-C, demonstrate that SMART-PC achieves state-of-the-art results, outperforming existing methods such as MATE in terms of both accuracy and computational efficiency. The implementation is available at: \url{https://github.com/AliBahri94/SMART-PC}. Ali Bahri, Moslem Yazdanpanah, Sahar Dastani, Mehrdad Noori, Gustavo Adolfo Vargas Hakim, David Osowiechi, Farzad Beizaee, Ismail Ben Ayed, Christian Desrosiers |
ICML | 8 |
| 2025 | Regularized Low-Rank Adaptation for Few-Shot Organ Segmentation
Ghassen Baklouti, Julio Silva-Rodríguez, Jose Dolz, Houda Bahig, Ismail Ben Ayed |
MICCAI (7) | 5 |
| 2025 | Reflect: Rectified Flows for Efficient Brain Anomaly Correction Transport
Farzad Beizaee, Sina Hajimiri, Ismail Ben Ayed, Gregory A. Lodygensky, Christian Desrosiers, Jose Dolz |
MICCAI (4) | 3 |
| 2025 | Trustworthy Few-Shot Transfer of Medical VLMs Through Split Conformal Prediction
Julio Silva-Rodríguez, Ismail Ben Ayed, Jose Dolz |
MICCAI (7) | 2 |
| 2025 | Few-Shot, Now for Real: Medical VLMs Adaptation Without Balanced Sets or Validation
Julio Silva-Rodríguez, Fereshteh Shakeri, Houda Bahig, Jose Dolz, Ismail Ben Ayed |
MICCAI (7) | 5 |
| 2025 | TRUST: Test-Time Refinement using Uncertainty-Guided SSM TraversesabstractState Space Models (SSMs) have emerged as efficient alternatives to Vision Transformers (ViTs), with VMamba standing out as a pioneering architecture designed for vision tasks. However, their generalization performance degrades significantly under distribution shifts. To address this limitation, we propose TRUST (Test-Time Refinement using Uncertainty-Guided SSM Traverses), a novel test-time adaptation (TTA) method that leverages diverse traversal permutations to generate multiple causal perspectives of the input image. Model predictions serve as pseudo-labels to guide updates of the Mamba-specific parameters, and the adapted weights are averaged to integrate the learned information across traversal scans. Altogether, TRUST is the first approach that explicitly leverages the unique architectural properties of SSMs for adaptation. Experiments on seven benchmarks show that TRUST consistently improves robustness and outperforms existing TTA methods. Sahar Dastani, Ali Bahri, Gustavo Adolfo Vargas Hakim, Moslem Yazdanpanah, Mehrdad Noori, David Osowiechi, Samuel Barbeau, Ismail Ben Ayed, Hervé Lombaert, Christian Desrosiers |
NeurIPS | 8 |
| 2025 | Learning Task-Agnostic Representations through Multi-Teacher DistillationabstractCasting complex inputs into tractable representations is a critical step across various fields. Diverse embedding models emerge from differences in architectures, loss functions, input modalities and datasets, each capturing unique aspects of the input. Multi-teacher distillation leverages this diversity to enrich representations but often remains tailored to specific tasks. We introduce a task-agnostic framework based on a ``majority vote" objective function.
We demonstrate that this function is bounded by the mutual information between the student and the teachers' embeddings, leading to a task-agnostic distillation loss that eliminates dependence on task-specific labels or prior knowledge.
Comprehensive evaluations across text, vision models, and molecular modeling show that our method effectively leverages teacher diversity, resulting in representations enabling better performance for a wide range of downstream tasks such as classification, clustering, or regression. Additionally, we train and release state-of-the-art embedding models, enhancing downstream performance in various modalities. Philippe Formont, Maxime Darrin, Banafsheh Karimian, Eric Granger, Jackie Chi Kit Cheung, Ismail Ben Ayed, Mohammadhadi Shateri, Pablo Piantanida |
NeurIPS | 6 |
| 2025 | Test-Time Adaptation of Vision-Language Models for Open-Vocabulary Semantic SegmentationabstractRecently, test-time adaptation has attracted wide interest in the context of vision-language models for image classification. However, to the best of our knowledge, the problem is completely overlooked in dense prediction tasks such as Open-Vocabulary Semantic Segmentation (OVSS). In response, we propose a novel TTA method tailored to adapting VLMs for segmentation during test time. Unlike TTA methods for image classification, our Multi-Level and Multi-Prompt (MLMP) entropy minimization integrates features from intermediate vision-encoder layers and is performed with different text-prompt templates at both the global CLS token and local pixel-wise levels.
Our approach could be used as plug-and-play for any segmentation network, does not require additional training data or labels, and remains effective even with a single test sample. Furthermore, we introduce a comprehensive OVSS TTA benchmark suite, which integrates a rigorous evaluation protocol, nine segmentation datasets, 15 common synthetic corruptions, and additional real and rendered domain shifts, with a total of 87 distinct test scenarios, establishing a standardized and comprehensive testbed for future TTA research in open-vocabulary segmentation. Our experiments on this suite demonstrate that our segmentation-tailored method consistently delivers significant gains over direct adoption of TTA classification baselines. Code and data are available at https://github.com/dosowiechi/MLMP. Mehrdad Noori, David Osowiechi, Gustavo Adolfo Vargas Hakim, Ali Bahri, Moslem Yazdanpanah, Sahar Dastani, Farzad Beizaee, Ismail Ben Ayed, Christian Desrosiers |
NeurIPS | 8 |
| 2025 | FDS: Feedback-Guided Domain Synthesis with Multi-Source Conditional Diffusion Models for Domain GeneralizationabstractDomain Generalization techniques aim to enhance model robustness by simulating novel data distributions during training, typically through various augmentation or stylization strategies. However, these methods frequently suffer from limited control over the diversity of generated images and lack assurance that these images span distinct distributions. To address these challenges, we propose FDS, Feedback-guided Domain Synthesis, a novel strategy that employs diffusion models to synthesize novel, pseudo-domains by training a single model on all source domains and performing domain mixing based on learned features. By incorporating images that pose classification challenges to models trained on original samples, alongside the original dataset, we ensure the generation of a training set that spans a broad distribution spectrum. Our comprehensive evaluations demonstrate that this methodology sets new benchmarks in domain generalization performance across a range of challenging datasets, effectively managing diverse types of domain shifts. The code can be found at: https://github.com/Mehrdad-Noori/FDS Ali Bahri, Mehrdad Noori, Gustavo Adolfo Vargas Hakim, Ismail Ben Ayed, Milad Cheraghalikhani, David Osowiechi, Christian Desrosiers, Moslem Yazdanpanah |
WACV | 4 |
| 2025 | Test-Time Adaptation in Point Clouds: Leveraging Sampling Variation with Weight AveragingabstractTest-Time Adaptation (TTA) addresses distribution shifts during testing by adapting a pretrained model without access to source data. In this work, we propose a novel TTA approach for 3D point cloud classification, combining sampling variation with weight averaging. Our method leverages Farthest Point Sampling (FPS) and K-Nearest Neighbors (KNN) to create multiple point cloud representations, adapting the model for each variation using the TENT algorithm. The final model parameters are obtained by averaging the adapted weights, leading to improved robustness against distribution shifts. Extensive experiments on ModelNet40-C, ShapeNet-C, and ScanObjectNN-C datasets, with different backbones (Point-MAE, Point-Net, DGCNN), demonstrate that our approach consistently outperforms existing methods while maintaining minimal resource overhead. The proposed method effectively enhances model generalization and stability in challenging real-world conditions. The implementation is available at: https://github.com/AliBahri94/SVWA_TTA.git. Ali Bahri, Moslem Yazdanpanah, Mehrdad Noori, Sahar Dastani, Milad Cheraghalikhani, David Osowiechi, Farzad Beizaee, Gustavo Adolfo Vargas Hakim, Ismail Ben Ayed, Christian Desrosiers |
WACV | 9 |
| 2025 | Pay Attention to Your Neighbours: Training-Free Open-Vocabulary Semantic SegmentationabstractDespite the significant progress in deep learning for dense visual recognition problems, such as semantic segmentation, traditional methods are constrained by fixed class sets. Meanwhile, vision-language foundation models, such as CLIP, have showcased remarkable effectiveness in numerous zero-shot image-level tasks, owing to their robust generalizability. Recently, a body of work has investigated utilizing these models in open-vocabulary semantic segmentation (OVSS). However, existing approaches often rely on impractical supervised pretraining or access to additional pretrained networks. In this work, we propose a strong baseline for training-free OVSS, termed Neighbour-Aware CLIP (NACLIP), representing a straightforward adaptation of CLIP tailored for this scenario. Our method enforces localization of patches in the self-attention of CLIP's vision transformer which, despite being crucial for dense prediction tasks, has been overlooked in the OVSS literature. By incorporating design choices favouring segmentation, our approach significantly improves performance without requiring additional data, auxiliary pretrained networks, or extensive hyperparameter tuning, making it highly practicalfor real-world applications. Experiments are performed on 8 popular semantic segmentation benchmarks, yielding state-of-the-art performance on most scenarios. Our code is publicly available at https://github.com/sinahmr/NACLIP. Sina Hajimiri, Ismail Ben Ayed, Jose Dolz |
WACV | 2 |
| 2025 | CLIPArTT: Adaptation of CLIP to New Domains at Test TimeabstractPre-trained vision-language models (VLMs), exemplified by CLIP, demonstrate remarkable adaptability across zero-shot classification tasks without additional training. However, their performance diminishes in the presence of domain shifts. In this study, we introduce CLIP Adaptation duRing Test-Time (CLIPArTT), a fully test-time adaptation (TTA) approach for CLIP, which involves automatic text prompts construction during inference for their use as text supervision. Our method employs a unique, minimally invasive text prompt tuning process, wherein multiple predicted classes are aggregated into a single new text prompt, used as pseudo label to re-classify inputs in a transductive manner. Additionally, we pioneer the standardization of TTA benchmarks (e.g., TENT) in the realm of VLMs. Our findings demonstrate that, without requiring additional transformations nor new trainable modules, CLIPArTT enhances performance dynamically across non-corrupted datasets such as CIFAR-100, corrupted datasets like CIFAR-100-C and ImageNet-C, alongside synthetic datasets such as VisDA-C. This research underscores the potential for improving VLMs' adaptability through novel test-time strategies, offering insights for robust performance across varied datasets and environments. The code can be found at: https://github.com/dosowiechi/CLIPArTT.git Gustavo Adolfo Vargas Hakim, David Osowiechi, Mehrdad Noori, Milad Cheraghalikhani, Ali Bahri, Moslem Yazdanpanah, Ismail Ben Ayed, Christian Desrosiers |
WACV | 7 |
| 2025 | Neighbor-aware calibration of segmentation networks with penalty-based constraints
Balamurali Murugesan, Sukesh Adiga V, Bingyuan Liu, Hervé Lombaert, Ismail Ben Ayed, Jose Dolz |
Medical Image Anal. | 5 |
| 2025 | A Foundation Language-Image Model of the Retina (FLAIR): encoding expert knowledge in text supervision
Julio Silva-Rodríguez, Hadi Chakor, Riadh Kobbi, Jose Dolz, Ismail Ben Ayed |
Medical Image Anal. | 5 |
| 2025 | Towards Foundation Models and Few-Shot Parameter-Efficient Fine-Tuning for Volumetric Organ Segmentation
Julio Silva-Rodríguez, Jose Dolz, Ismail Ben Ayed |
Medical Image Anal. | 3 |
| 2025 | CoLo-CAM: Class activation mapping for object co-localization in weakly-labeled unconstrained videosabstractLeveraging spatiotemporal information in videos is critical for weakly supervised video object localization (WSVOL) tasks. However, state-of-the-art methods only rely on visual and motion cues, while discarding discriminative information, making them susceptible to inaccurate localizations. Recently, discriminative models have been explored for WSVOL tasks using a temporal class activation mapping (CAM) method. Although their results are promising, objects are assumed to have limited movement from frame to frame, leading to degradation in performance for relatively long-term dependencies. This paper proposes a novel CAM method for WSVOL that exploits spatiotemporal information in activation maps during training without constraining an object’s position. Its training relies on Co - Lo calization, hence, the name CoLo-CAM . Given a sequence of frames, localization is jointly learned based on color cues extracted across the corresponding maps, by assuming that an object has similar color in consecutive frames. CAM activations are constrained to respond similarly over pixels with similar colors, achieving co-localization. This improves localization performance because the joint learning creates direct communication among pixels across all image locations and over all frames, allowing for transfer, aggregation, and correction of localizations. Co-localization is integrated into training by minimizing the color term of a conditional random field (CRF) loss over a sequence of frames/CAMs. Extensive experiments 1 on two challenging YouTube-Objects datasets of unconstrained videos show the merits of our CoLo-CAM method, and its robustness to long-term dependencies, leading to new state-of-the-art performance for WSVOL task. Soufiane Belharbi, Shakeeb Murtaza, Marco Pedersoli, Ismail Ben Ayed, Luke McCaffrey, Eric Granger |
Pattern Recognit. | 4 |
| 2025 | Aggregatedf-average neural network applied to few-shot class incremental learning
Mathieu Vu, Emilie Chouzenoux, Ismail Ben Ayed, Jean-Christophe Pesquet |
Signal Process. | 3 |
| 2024 | LP++: A Surprisingly Strong Linear Probe for Few-Shot CLIPabstractIn a recent, strongly emergent literature on few-shot CLIP adaptation, Linear Probe (LP) has been often reported as a weak baseline. This has motivated intensive research building convoluted prompt learning or feature adaptation strategies. In this work, we propose and exam-ine from convex-optimization perspectives a generalization of the standard LP baseline, in which the linear classifier weights are learnable functions of the text embedding, with class-wise multipliers blending image and text knowledge. As our objective function depends on two types of variables, i.e., the class visual prototypes and the learnable blending parameters, we propose a computationally efficient block coordinate Majorize-Minimize (MM) descent algorithm. In our full-batch MM optimizer, which we coin LP++, step sizes are implicit, unlike standard gradient descent practices where learning rates are intensively searched over validation sets. By examining the mathematical properties of our loss (e.g., Lipschitz gradient continuity), we build ma-jorizing functions yielding data-driven learning rates and derive approximations of the loss's minima, which provide data-informed initialization of the variables. Our image-language objective function, along with these non-trivial optimization insights and ingredients, yields, surprisingly, highly competitive few-shot CLIP performances. Furthermore, LP++ operates in black-box, relaxes intensive validation searches for the optimization hyper-parameters, and runs orders-of-magnitudes faster than state-of-the-art few-shot CLIP adaptation methods. Our code is available at: https://github.com/FereshteShakeri/FewShot-CLIP-Strong-Baseline.git. Yunshi Huang, Fereshteh Shakeri, Jose Dolz, Malik Boudiaf, Houda Bahig, Ismail Ben Ayed |
CVPR | 6 |
| 2024 | Transductive Zero-Shot and Few-Shot CLIPabstractTransductive inference has been widely investigated in few-shot image classification, but completely overlooked in the recent, fast growing literature on adapting vision-langage models like CLIP. This paper addresses the transductive zero-shot and few-shot CLIP classification challenge, in which inference is performed jointly across a mini-batch of unlabeled query samples, rather than treating each instance independently. We initially construct informative vision-text probability features, leading to a classification problem on the unit simplex set. Inspired by Expectation-Maximization (EM), our optimization-based classification objective models the data probability distribution for each class using a Dirichlet law. The minimization problem is then tackled with a novel block Majorization-Minimization algorithm, which simultaneously estimates the distribution parameters and class assignments. Extensive numerical experiments on 11 datasets underscore the benefits and efficacy of our batch inference approach. On zero-shot tasks with test batches of 75 samples, our approach yields near 20% improvement in ImageNet accuracy over CLIP's zero-shot performance. Additionally, we outperform state-of-the-art methods in the few-shot setting. The code is available at: https://github.com/SegoleneMartin/transductive-CLIP. Ségolène Martin, Yunshi Huang, Fereshteh Shakeri, Jean-Christophe Pesquet, Ismail Ben Ayed |
CVPR | 5 |
| 2024 | NC-TTT: A Noise Constrastive Approach for Test-Time TrainingabstractDespite their exceptional performance in vision tasks, deep learning models often struggle when faced with domain shifts during testing. Test-Time Training (TTT) methods have recently gained popularity by their ability to enhance the robustness of models through the addition of an auxiliary objective that is jointly optimized with the main task. Being strictly unsupervised, this auxiliary objective is used at test time to adapt the model without any access to labels. In this work, we propose Noise-Contrastive TestTime Training (NC-TTT), a novel unsupervised TTT technique based on the discrimination of noisy feature maps. By learning to classify noisy views of projected feature maps, and then adapting the model accordingly on new domains, classification performance can be recovered by an important margin. Experiments on several popular testtime adaptation baselines demonstrate the advantages of our method compared to recent approaches for this task. The code can be found at: https://github.com/GustavoVargasHakim/NCTTT.git David Osowiechi, Gustavo Adolfo Vargas Hakim, Mehrdad Noori, Milad Cheraghalikhani, Ali Bahri, Moslem Yazdanpanah, Ismail Ben Ayed, Christian Desrosiers |
CVPR | 7 |
| 2024 | A Closer Look at the Few-Shot Adaptation of Large Vision-Language ModelsabstractEfficient transfer learning (ETL) is receiving increasing attention to adapt large pre-trained language-vision models on downstream tasks with a few labeled samples. While significant progress has been made, we reveal that state-of-the-art ETL approaches exhibit strong performance only in narrowly-defined experimental setups, and with a careful adjustment of hyperparameters based on a large corpus of labeled samples. In particular, we make two interesting, and surprising empirical observations. First, to out-perform a simple Linear Probing baseline, these methods require to optimize their hyper-parameters on each target task. And second, they typically underperform -sometimes dramatically- standard zero-shot predictions in the presence of distributional drifts. Motivated by the unrealistic assumptions made in the existing literature, i.e., access to a large validation set and case-specific grid-search for optimal hyperparameters, we propose a novel approach that meets the requirements of real-world scenarios. More concretely, we introduce a CLass-Adaptive linear Probe (CLAP) objective, whose balancing term is optimized via an adaptation of the general Augmented Lagrangian method tailored to this context. We comprehensively evaluate CLAP on a broad span of datasets and scenarios, demonstrating that it consistently outperforms SoTA approaches, while yet being a much more efficient alternative. Code available at https://github.com/jusiro/CLAP. Julio Silva-Rodríguez, Sina Hajimiri, Ismail Ben Ayed, Jose Dolz |
CVPR | 3 |
| 2024 | On the Test-Time Zero-Shot Generalization of Vision-Language Models: Do we Really need Prompt Learning?abstractThe development of large vision-language models, notably CLIP, has catalyzed research into effective adaptation techniques, with a particular focus on soft prompt tuning. Conjointly, test-time augmentation, which utilizes multiple augmented views of a single image to enhance zero-shot generalization, is emerging as a significant area of interest. This has predominantly directed research efforts toward test-time prompt tuning. In contrast, we introduce a robust MeanShift for Test-time Augmentation (MTA), which surpasses prompt-based methods without requiring this intensive training procedure. This positions MTA as an ideal solution for both standalone and API-based applications. Additionally, our method does not rely on ad hoc rules (e.g., confidence threshold) used in some previous test-time augmentation techniques to filter the augmented views. Instead, MTA incorporates a quality assessment variable for each view directly into its optimization process, termed as the inlierness score. This score is jointly optimized with a density mode seeking process, leading to an efficient trainingand hyperparameter-free approach. We extensively benchmark our method on 15 datasets and demonstrate MTA's superiority and computational efficiency. Deployed easily as plug-and-play module on top of zero-shot models and state-of-the-art few-shot methods, MTA shows systematic and consistent improvements. Maxime Zanella, Ismail Ben Ayed |
CVPR | 2 |
| 2024 | Robust Calibration of Large Vision-Language Adapters
Balamurali Murugesan, Julio Silva-Rodríguez, Ismail Ben Ayed, Jose Dolz |
ECCV (24) | 3 |
| 2024 | Class and Region-Adaptive Constraints for Network Calibration
Balamurali Murugesan, Julio Silva-Rodríguez, Ismail Ben Ayed, Jose Dolz |
MICCAI (8) | 3 |
| 2024 | Few-Shot Adaptation of Medical Vision-Language Models
Fereshteh Shakeri, Yunshi Huang, Julio Silva-Rodríguez, Houda Bahig, An Tang, Jose Dolz, Ismail Ben Ayed |
MICCAI (12) | 7 |
| 2024 | When is an Embedding Model More Promising than Another?abstractEmbedders play a central role in machine learning, projecting any object into numerical representations that can, in turn, be leveraged to perform various downstream tasks. The evaluation of embedding models typically depends on domain-specific empirical approaches utilizing downstream tasks, primarily because of the lack of a standardized framework for comparison. However, acquiring adequately large and representative datasets for conducting these assessments is not always viable and can prove to be prohibitively expensive and time-consuming. In this paper, we present a unified approach to evaluate embedders. First, we establish theoretical foundations for comparing embedding models, drawing upon the concepts of sufficiency and informativeness. We then leverage these concepts to devise a tractable comparison criterion (information sufficiency), leading to a task-agnostic and self-supervised ranking procedure. We demonstrate experimentally that our approach aligns closely with the capability of embedding models to facilitate various downstream tasks in both natural language processing and molecular biology. This effectively offers practitioners a valuable tool for prioritizing model trials. Maxime Darrin, Philippe Formont, Ismail Ben Ayed, Jackie Chi Kit Cheung, Pablo Piantanida |
NeurIPS | 3 |
| 2024 | WATT: Weight Average Test Time Adaptation of CLIPabstractVision-Language Models (VLMs) such as CLIP have yielded unprecedented performances for zero-shot image classification, yet their generalization capability may still be seriously challenged when confronted to domain shifts. In response, we present Weight Average Test-Time Adaptation (WATT) of CLIP, a new approach facilitating full test-time adaptation (TTA) of this VLM. Our method employs a diverse set of templates for text prompts, augmenting the existing framework of CLIP. Predictions are utilized as pseudo labels for model updates, followed by weight averaging to consolidate the learned information globally. Furthermore, we introduce a text ensemble strategy, enhancing the overall test performance by aggregating diverse textual cues.
Our findings underscore the effectiveness of WATT across diverse datasets, including CIFAR-10-C, CIFAR-10.1, CIFAR-100-C, VisDA-C, and several other challenging datasets, effectively covering a wide range of domain shifts. Notably, these enhancements are achieved without the need for additional model transformations or trainable modules. Moreover, compared to other TTA methods, our approach can operate effectively with just a single image. The code is available at: https://github.com/Mehrdad-Noori/WATT. David Osowiechi, Mehrdad Noori, Gustavo Adolfo Vargas Hakim, Moslem Yazdanpanah, Ali Bahri, Milad Cheraghalikhani, Sahar Dastani, Farzad Beizaee, Ismail Ben Ayed, Christian Desrosiers |
NeurIPS | 9 |
| 2024 | Boosting Vision-Language Models with TransductionabstractTransduction is a powerful paradigm that leverages the structure of unlabeled data to boost predictive accuracy. We present TransCLIP, a novel and computationally efficient transductive approach designed for Vision-Language Models (VLMs). TransCLIP is applicable as a plug-and-play module on top of popular inductive zero- and few-shot models, consistently improving their performances. Our new objective function can be viewed as a regularized maximum-likelihood estimation, constrained by a KL divergence penalty that integrates the text-encoder knowledge and guides the transductive learning process. We further derive an iterative Block Majorize-Minimize (BMM) procedure for optimizing our objective, with guaranteed convergence and decoupled sample-assignment updates, yielding computationally efficient transduction for large-scale datasets. We report comprehensive evaluations, comparisons, and ablation studies that demonstrate: (i) Transduction can greatly enhance the generalization capabilities of inductive pretrained zero- and few-shot VLMs; (ii) TransCLIP substantially outperforms standard transductive few-shot learning methods relying solely on vision features, notably due to the KL-based language constraint. Maxime Zanella, Benoît Gérin, Ismail Ben Ayed |
NeurIPS | 3 |
| 2024 | Bag of Tricks for Fully Test-Time AdaptationabstractFully Test-Time Adaptation (TTA), which aims at adapting models to data drifts, has recently attracted wide interest. Numerous tricks and techniques have been proposed to ensure robust learning on arbitrary streams of unlabeled data. However, assessing the true impact of each individual technique and obtaining a fair comparison still constitutes a significant challenge. To help consolidate the community’s knowledge, we present a categorization of selected orthogonal TTA techniques, including small batch normalization, stream rebalancing, reliable sample selection, and network confidence calibration. We meticulously dissect the effect of each approach on different scenarios of interest. Through our analysis, we shed light on trade-offs induced by those techniques between accuracy, the computational power required, and model complexity. We also uncover the synergy that arises when combining techniques and are able to establish new state-of-the-art results. Saypraseuth Mounsaveng, Florent Chiaroni, Malik Boudiaf, Marco Pedersoli, Ismail Ben Ayed |
WACV | 5 |
| 2024 | Prompting classes: Exploring the Power of Prompt Class Learning in Weakly Supervised Semantic SegmentationabstractRecently, CLIP-based approaches have exhibited remarkable performance on generalization and few-shot learning tasks, fueled by the power of contrastive language-vision pre-training. In particular, prompt tuning has emerged as an effective strategy to adapt the pre-trained language-vision models to downstream tasks by employing task-related textual tokens. Motivated by this progress, in this work we question whether other fundamental problems, such as weakly supervised semantic segmentation (WSSS), can benefit from prompt tuning. Our findings reveal two interesting observations that shed light on the impact of prompt tuning on WSSS. First, modifying only the class token of the text prompt results in a greater impact on the Class Activation Map (CAM), compared to arguably more complex strategies that optimize the context. And second, the class token associated with the image ground truth does not necessarily correspond to the category that yields the best CAM. Motivated by these observations, we introduce a novel approach based on a PrOmpt cLass lEarning (POLE) strategy. Through extensive experiments we demonstrate that our simple, yet efficient approach achieves SOTA performance in a well-known WSSS benchmark. These results highlight not only the benefits of language-vision models in WSSS but also the potential of prompt learning for this problem. The code is available at code link Balamurali Murugesan, Rukhshanda Hussain, Rajarshi Bhattacharya, Ismail Ben Ayed, Jose Dolz |
WACV | 4 |
| 2024 | Structure-aware feature stylization for domain generalizationabstractGeneralizing to out-of-distribution (OOD) data is a challenging task for existing deep learning approaches. This problem largely comes from the common but often incorrect assumption of statistical learning algorithms that the source and target data come from the same i.i.d. distribution. To tackle the limited variability of domains available during training, as well as domain shifts at test time, numerous approaches for domain generalization have focused on generating samples from new domains. Recent studies on this topic suggest that feature statistics from instances of different domains can be mixed to simulate synthesized images from a novel domain. While this simple idea achieves state-of-art results on various domain generalization benchmarks, it ignores structural information which is key to transferring knowledge across different domains. In this paper, we leverage the ability of humans to recognize objects using solely their structural information (prominent region contours) to design a Structural-Aware Feature Stylization method for domain generalization. Our method improves feature stylization based on mixing instance statistics by enforcing structural consistency across the different style-augmented samples. This is achieved via a multi-task learning model which classifies original and augmented images while also reconstructing their edges in a secondary task. The edge reconstruction task helps the network preserve image structure during feature stylization, while also acting as a regularizer for the classification task. Through quantitative comparisons, we verify the effectiveness of our method upon existing state-of-the-art methods on PACS, VLCS, OfficeHome, DomainNet and Digits-DG. The implementation is available at this repository. Milad Cheraghalikhani, Mehrdad Noori, David Osowiechi, Gustavo Adolfo Vargas Hakim, Ismail Ben Ayed, Christian Desrosiers |
Comput. Vis. Image Underst. | 5 |
| 2024 | Do we really need dice? The hidden region-size biases of segmentation losses
Bingyuan Liu, Jose Dolz, Adrian Galdran, Riadh Kobbi, Ismail Ben Ayed |
Medical Image Anal. | 5 |
| 2024 | Simplex Clustering via sBeta With Applications to Online Adjustment of Black-Box PredictionsabstractWe explore clustering the softmax predictions of deep neural networks and introduce a novel probabilistic clustering method, referred to ask-sBetas. In the general context of clustering discrete distributions, the existing methods focused on exploring distortion measures tailored to simplex data, such as the KL divergence, as alternatives to the standard euclidean distance. We provide a general maximum a posteriori (MAP) perspective of clustering distributions, emphasizing that the statistical models underlying the existing distortion-based methods may not be descriptive enough. Instead, we optimize a mixed-variable objective measuring data conformity within each cluster to the introduced$\mathtt {sBeta}$density function, whose parameters are constrained and estimated jointly with binary assignment variables. Our versatile formulation approximates various parametric densities for modeling simplex data and enables the control of the cluster-balance bias. This yields highly competitive performances for the unsupervised adjustment of black-box model predictions in various scenarios. Florent Chiaroni, Malik Boudiaf, Amar Mitiche, Ismail Ben Ayed |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | TFS-ViT: Token-level feature stylization for domain generalization
Mehrdad Noori, Milad Cheraghalikhani, Ali Bahri, Gustavo Adolfo Vargas Hakim, David Osowiechi, Ismail Ben Ayed, Christian Desrosiers |
Pattern Recognit. | 6 |
| 2023 | Exploring the Transferability of a Foundation Model for Fundus Images: Application to Hypertensive Retinopathy
Julio Silva-Rodríguez, Jihed Chelbi, Waziha Kabir, Hadi Chakor, Jose Dolz, Ismail Ben Ayed, Riadh Kobbi |
CGI (3) | 6 |
| 2023 | Open-Set Likelihood Maximization for Few-Shot LearningabstractWe tackle the Few-Shot Open-Set Recognition (FSOSR) problem, i.e. classifying instances among a set of classes for which we only have a few labeled samples, while simultaneously detecting instances that do not belong to any known class. We explore the popular transductive setting, which leverages the unlabelled query instances at inference. Motivated by the observation that existing transductive methods perform poorly in open-set scenarios, we propose a generalization of the maximum likelihood principle, in which latent scores down-weighing the influence of potential outliers are introduced alongside the usual parametric model. Our formulation embeds supervision constraints from the support set and additional penalties discouraging overconfident predictions on the query set. We proceed with a block-coordinate descent, with the latent scores and parametric model co-optimized alternately, thereby benefiting from each other. We call our resulting formulation Open-Set Likelihood Optimization (OSLO). OSLO is interpretable and fully modular; it can be applied on top of any pre-trained model seamlessly. Through extensive experiments, we show that our method surpasses existing inductive and transductive methods on both aspects of open-set recognition, namely inlier classification and outlier detection. Code is available at https://github.com/ebennequin/few-shot-open-set. Malik Boudiaf, Etienne Bennequin, Myriam Tami, Antoine Toubhans, Pablo Piantanida, Céline Hudelot, Ismail Ben Ayed |
CVPR | 7 |
| 2023 | A Strong Baseline for Generalized Few-Shot Semantic SegmentationabstractThis paper introduces a generalized few-shot segmentation framework with a straightforward training process and an easy-to-optimize inference phase. In particular, we propose a simple yet effective model based on the well-known InfoMax principle, where the Mutual Information (MI) between the learned feature representations and their corresponding predictions is maximized. In addition, the terms derived from our MI-based formulation are coupled with a knowledge distillation term to retain the knowledge on base classes. With a simple training process, our inference model can be applied on top of any segmentation network trained on base classes. The proposed inference yields substantial improvements on the popular few-shot segmentation benchmarks, PASCAL-5iand COCO-20i. Particularly, for novel classes, the improvement gains range from 7% to 26% (PASCAL-5i) and from 3% to 12% (COCO-20i) in the 1-shot and 5-shot scenarios, respectively. Furthermore, we propose a more challenging setting, where performance gaps are further exacerbated. Our code is publicly available at https://github.com/sinahmr/DIaM. Sina Hajimiri, Malik Boudiaf, Ismail Ben Ayed, Jose Dolz |
CVPR | 3 |
| 2023 | Class Adaptive Network CalibrationabstractRecent studies have revealed that, beyond conventional accuracy, calibration should also be considered for training modern deep neural networks. To address miscalibration during learning, some methods have explored different penalty functions as part of the learning objective, along-side a standard classification loss, with a hyper-parameter controlling the relative contribution of each term. Nevertheless, these methods share two major drawbacks: 1) the scalar balancing weight is the same for all classes, hindering the ability to address different intrinsic difficulties or imbalance among classes; and 2) the balancing weight is usually fixed without an adaptive strategy, which may prevent from reaching the best compromise between accuracy and calibration, and requires hyper-parameter search for each application. We propose Class Adaptive Label Smoothing (CALS) for calibrating deep networks, which allows to learn class-wise multipliers during training, yielding a powerful alternative to common label smoothing penalties. Our method builds on a general Augmented Lagrangian approach, a well-established technique in constrained optimization, but we introduce several modifications to tailor it for large-scale, class-adaptive training. Comprehensive evaluation and multiple comparisons on a variety of benchmarks, including standard and long-tailed image classification, semantic segmentation, and text classification, demonstrate the superiority of the proposed method. The code is available at https://github.com/by-liu/CALS. Bingyuan Liu, Jérôme Rony, Adrian Galdran, Jose Dolz, Ismail Ben Ayed |
CVPR | 5 |
| 2023 | Proximal Splitting Adversarial Attack for Semantic SegmentationabstractClassification has been the focal point of research on adversarial attacks, but only a few works investigate methods suited to denser prediction tasks, such as semantic segmentation. The methods proposed in these works do not accurately solve the adversarial segmentation problem and, therefore, overestimate the size of the perturbations required to fool models. Here, we propose a white-box attack for these models based on a proximal splitting to produce adversarial perturbations with much smaller$\ell_{\infty}$norms. Our attack can handle large numbers of constraints within a nonconvex minimization framework via an Augmented Lagrangian approach, coupled with adaptive constraint scaling and masking strategies. We demonstrate that our attack significantly outperforms previously proposed ones, as well as classification attacks that we adapted for segmentation, providing a first comprehensive benchmark for this dense task. Jérôme Rony, Jean-Christophe Pesquet, Ismail Ben Ayed |
CVPR | 3 |
| 2023 | Transductive Learning for Textual Few-Shot Classification in API-based Embedding ModelsabstractPierre Colombo, Victor Pellegrain, Malik Boudiaf, Myriam Tami, Victor Storchan, Ismail Ayed, Pablo Piantanida. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Pierre Colombo, Victor Pellegrain, Malik Boudiaf, Myriam Tami, Victor Storchan, Ismail Ben Ayed, Pablo Piantanida |
EMNLP | 6 |
| 2023 | Parametric Information Maximization for Generalized Category DiscoveryabstractWe introduce a Parametric Information Maximization (PIM) model for the Generalized Category Discovery (GCD) problem. Specifically, we propose a bi-level optimization formulation, which explores a parameterized family of objective functions, each evaluating a weighted mutual information between the features and the latent labels, subject to supervision constraints from the labeled samples. Our formulation mitigates the class-balance bias encoded in standard information maximization approaches, thereby handling effectively both short-tailed and long-tailed data sets. We report extensive experiments and comparisons demonstrating that our PIM model consistently sets new state-of-the-art performances in GCD across six different datasets, more so when dealing with challenging fine-grained problems. Our code: https://github.com/ThalesGroup/pim-generalized-category-discovery. Florent Chiaroni, Jose Dolz, Imtiaz Masud Ziko, Amar Mitiche, Ismail Ben Ayed |
ICCV | 5 |
| 2023 | ClusT3: Information Invariant Test-Time TrainingabstractDeep Learning models have shown remarkable performance in a broad range of vision tasks. However, they are often vulnerable to domain shifts at test-time. Test-time training (TTT) methods have been developed in an attempt to mitigate these vulnerabilities, where a secondary task is solved at training time, simultaneously with the main task, to be later used as an self-supervised proxy task at test-time. In this work, we propose a novel unsupervised TTT technique based on the maximization of Mutual Information between multi-scale feature maps and a discrete latent representation, which can be integrated to the standard training as an auxiliary clustering task. Experimental results demonstrate competitive classification performance on different popular test-time adaptation benchmarks. The code can be found at: https://github.com/dosowiechi/ClusT3.git Gustavo Adolfo Vargas Hakim, David Osowiechi, Mehrdad Noori, Milad Cheraghalikhani, Ali Bahri, Ismail Ben Ayed, Christian Desrosiers |
ICCV | 6 |
| 2023 | Trust Your Neighbours: Penalty-Based Constraints for Model Calibration
Balamurali Murugesan, Sukesh Adiga V, Bingyuan Liu, Hervé Lombaert, Ismail Ben Ayed, Jose Dolz |
MICCAI (3) | 5 |
| 2023 | TCAM: Temporal Class Activation Maps for Object Localization in Weakly-Labeled Unconstrained VideosabstractWeakly supervised video object localization (WSVOL) allows locating object in videos using only global video tags such as object classes. State-of-art methods rely on multiple independent stages, where initial spatio-temporal proposals are generated using visual and motion cues, and then prominent objects are identified and refined. The localization involves solving an optimization problem over one or more videos, and video tags are typically used for video clustering. This process requires a model per video or per class making for costly inference. Moreover, localized regions are not necessary discriminant because these methods rely on unsupervised motion methods like optical flow, or discarded video tags from optimization. In this paper, we leverage the successful class activation mapping (CAM) methods, designed for WSOL based on still images. A new Temporal CAM (TCAM) method is introduced for training a discriminant deep learning (DL) model to exploit spatio-temporal information in videos, using an CAM-Temporal Max Pooling (CAM-TMP) aggregation mechanism over consecutive CAMs. In particular, activations of regions of interest (ROIs) are collected from CAMs produced by a pretrained CNN classifier, and generate pixel-wise pseudo-labels for training a decoder. In addition, a global unsupervised size constraint, and local constraint such as CRF are used to yield more accurate CAMs. Inference over single independent frames allows parallel processing of a clip of frames, and real-time localization. Extensive experiments1on two challenging YouTube-Objects datasets with unconstrained videos indicate that CAM methods (trained on independent frames) can yield decent localization accuracy. Our proposed TCAM method achieves a new state-of-art in WSVOL accuracy, and visual results suggest that it can be adapted for subsequent tasks, such as object detection and tracking. Soufiane Belharbi, Ismail Ben Ayed, Luke McCaffrey, Eric Granger |
WACV | 2 |
| 2023 | TTTFlow: Unsupervised Test-Time Training with Normalizing FlowabstractA major problem of deep neural networks for image classification is their vulnerability to domain changes at test-time. Recent methods have proposed to address this problem with test-time training (TTT), where a two-branch model is trained to learn a main classification task and also a self-supervised task used to perform test-time adaptation. However, these techniques require defining a proxy task specific to the target application. To tackle this limitation, we propose TTTFlow: a Y-shaped architecture using an unsupervised head based on Normalizing Flows to learn the nor-mal distribution of latent features and detect domain shifts in test examples. At inference, keeping the unsupervised head fixed, we adapt the model to domain-shifted examples by maximizing the log likelihood of the Normalizing Flow. Our results show that our method can significantly improve the accuracy with respect to previous works. David Osowiechi, Gustavo Adolfo Vargas Hakim, Mehrdad Noori, Milad Cheraghalikhani, Ismail Ben Ayed, Christian Desrosiers |
WACV | 5 |
| 2023 | Segmentation with mixed supervision: Confidence maximization helps knowledge distillation
Bingyuan Liu, Christian Desrosiers, Ismail Ben Ayed, Jose Dolz |
Medical Image Anal. | 3 |
| 2023 | Calibrating segmentation networks with margin-based label smoothing
Balamurali Murugesan, Bingyuan Liu, Adrian Galdran, Ismail Ben Ayed, Jose Dolz |
Medical Image Anal. | 4 |
| 2023 | Adversarial Robustness Via Fisher-Rao RegularizationabstractAdversarial robustness has become a topic of growing interest in machine learning since it was observed that neural networks tend to be brittle. We propose an information-geometric formulation of adversarial defense and introduce Fire, a new Fisher-Rao regularization for the categorical cross-entropy loss, which is based on the geodesic distance between the softmax outputs corresponding to natural and perturbed input features. Based on the information-geometric properties of the class of softmax distributions, we derive an explicit characterization of the Fisher-Rao Distance (FRD) for the binary and multiclass cases, and draw some interesting properties as well as connections with standard regularization metrics. Furthermore, we verify on a simple linear and Gaussian model, that all Pareto-optimal points in the accuracy-robustness region can be reached by Fire while other state-of-the-art methods fail. Empirically, we evaluate the performance of various classifiers trained with the proposed loss on standard datasets, showing up to a simultaneous 1% of improvement in terms of clean and robust performances while reducing the training time by 20% over the best-performing methods. Marine Picot, Francisco Messina, Malik Boudiaf, Fabrice Labeau, Ismail Ben Ayed, Pablo Piantanida |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Parameter-free Online Test-time AdaptationabstractTraining state-of-the-art vision models has become prohibitively expensive for researchers and practitioners. For the sake of accessibility and resource reuse, it is important to focus on adapting these models to a variety of down-stream scenarios. An interesting and practical paradigm is online test-time adaptation, according to which training data is inaccessible, no labelled data from the test distribution is available, and adaptation can only happen at test time and on a handful of samples. In this paper, we investigate how test-time adaptation methods fare for a number of pre-trained models on a variety of real-world scenarios, significantly extending the way they have been originally evaluated. We show that they perform well only in narrowly-defined experimental setups and sometimes fail catastrophically when their hyperparameters are not selected for the same scenario in which they are being tested. Motivated by the inherent uncertainty around the conditions that will ultimately be encountered at test time, we propose a particularly “conservative” approach, which addresses the problem with a Laplacian Adjusted Maximum-likelihood Estimation (LAME) objective. By adapting the model's output (not its parameters), and solving our objective with an efficient concave-convex procedure, our approach exhibits a much higher average accuracy across scenarios than existing methods, while being notably faster and have a much lower memory footprint. The code is available at https://github.com/fiveai/LAME. Malik Boudiaf, Romain Müller, Ismail Ben Ayed, Luca Bertinetto |
CVPR | 3 |
| 2022 | The Devil is in the Margin: Margin-based Label Smoothing for Network CalibrationabstractIn spite of the dominant performances of deep neural networks, recent works have shown that they are poorly calibrated, resulting in over-confident predictions. Miscalibration can be exacerbated by overfitting due to the minimization of the cross-entropy during training, as it promotes the predicted softmax probabilities to match the one-hot label assignments. This yields a pre-softmax activation of the correct class that is significantly larger than the remaining activations. Recent evidence from the literature suggests that loss functions that embed implicit or explicit maximization of the entropy of predictions yield state-of-the-art calibration performances. We provide a unifying constrained-optimization perspective of current state-of-the-art calibration losses. Specifically, these losses could be viewed as approximations of a linear penalty (or a Lagrangian term) imposing equality constraints on logit distances. This points to an important limitation of such underlying equality constraints, whose ensuing gradients constantly push towards a non-informative solution, which might prevent from reaching the best compromise between the discriminative performance and calibration of the model during gradient-based optimization. Following our observations, we propose a simple and flexible generalization based on inequality constraints, which imposes a controllable margin on logit distances. Comprehensive experiments on a variety of image classification, semantic segmentation and NLP benchmarks demonstrate that our method sets novel state-of-the-art results on these tasks in terms of network calibration, without affecting the discriminative performance. The code is available at https://github.com/by-liu/MbLS. Bingyuan Liu, Ismail Ben Ayed, Adrian Galdran, Jose Dolz |
CVPR | 2 |
| 2022 | Test-Time Adaptation with Shape Moments for Image Segmentation
Mathilde Bateson, Hervé Lombaert, Ismail Ben Ayed |
MICCAI (4) | 3 |
| 2022 | Towards Practical Few-shot Query Sets: Transductive Minimum Description Length InferenceabstractStandard few-shot benchmarks are often built upon simplifying assumptions on the query sets, which may not always hold in practice. In particular, for each task at testing time, the classes effectively present in the unlabeled query set are known a priori, and correspond exactly to the set of classes represented in the labeled support set. We relax these assumptions and extend current benchmarks, so that the query-set classes of a given task are unknown, but just belong to a much larger set of possible classes. Our setting could be viewed as an instance of the challenging yet practical problem of extremely imbalanced $K$-way classification, $K$ being much larger than the values typically used in standard benchmarks, and with potentially irrelevant supervision from the support set. Expectedly, our setting incurs drops in the performances of state-of-the-art methods. Motivated by these observations, we introduce a \textbf{P}rim\textbf{A}l \textbf{D}ual Minimum \textbf{D}escription \textbf{LE}ngth (\textbf{PADDLE}) formulation, which balances data-fitting accuracy and model complexity for a given few-shot task, under supervision constraints from the support set. Our constrained MDL-like objective promotes competition among a large set of possible classes, preserving only effective classes that befit better the data of a few-shot task. It is hyper-parameter free, and could be applied on top of any base-class training. Furthermore, we derive a fast block coordinate descent algorithm for optimizing our objective, with convergence guarantee, and a linear computational complexity at each iteration. Comprehensive experiments over the standard few-shot datasets and the more realistic and challenging \textit{i-Nat} dataset show highly competitive performances of our method, more so when the numbers of possible classes in the tasks increase. Our code is publicly available at \url{https://github.com/SegoleneMartin/PADDLE}. Ségolène Martin, Malik Boudiaf, Emilie Chouzenoux, Jean-Christophe Pesquet, Ismail Ben Ayed |
NeurIPS | 5 |
| 2022 | F-CAM: Full Resolution Class Activation Maps via Guided Parametric UpscalingabstractClass Activation Mapping (CAM) methods have recently gained much attention for weakly-supervised object localization (WSOL) tasks. They allow for CNN visualization and interpretation without training on fully annotated image datasets. CAM methods are typically integrated within off-the-shelf CNN backbones, such as ResNet50. Due to convolution and pooling operations, these backbones yield low resolution CAMs with a down-scaling factor of up to 32, contributing to inaccurate localizations. Interpolation is required to restore full size CAMs, yet it does not consider the statistical properties of objects, such as color and texture, leading to activations with inconsistent boundaries, and inaccurate localizations. As an alternative, we introduce a generic method for parametric upscaling of CAMs that allows constructing accurate full resolution CAMs (FCAMs). In particular, we propose a trainable decoding architecture that can be connected to any CNN classifier to produce highly accurate CAM localizations. Given an original low resolution CAM, foreground and background pixels are randomly sampled to fine-tune the decoder. Additional priors such as image statistics and size constraints are also considered to expand and refine object boundaries. Extensive experiments1, over three CNN backbones and six WSOL baselines on the CUB-200-2011 and OpenImages datasets, indicate that our F-CAM method yields a significant improvement in CAM localization accuracy. F-CAM performance is competitive with state-of-art WSOL methods, yet it requires fewer computations during inference. Soufiane Belharbi, Aydin Sarraf, Marco Pedersoli, Ismail Ben Ayed, Luke McCaffrey, Eric Granger |
WACV | 4 |
| 2022 | Source-free domain adaptation for image segmentation
Mathilde Bateson, Hoel Kervadec, Jose Dolz, Hervé Lombaert, Ismail Ben Ayed |
Medical Image Anal. | 5 |
| 2022 | Deep Interpretable Classification and Weakly-Supervised Segmentation of Histology Images via Max-Min UncertaintyabstractWeakly-supervised learning (WSL) has recently triggered substantial interest as it mitigates the lack of pixel-wise annotations. Given global image labels, WSL methods yield pixel-level predictions (segmentations), which enable to interpret class predictions. Despite their recent success, mostly with natural images, such methods can face important challenges when the foreground and background regions have similar visual cues, yielding high false-positive rates in segmentations, as is the case in challenging histology images. WSL training is commonly driven by standard classification losses, which implicitly maximize model confidence, and locate the discriminative regions linked to classification decisions. Therefore, they lack mechanisms for modeling explicitly non-discriminative regions and reducing false-positive rates. We propose novel regularization terms, which enable the model to seek both non-discriminative and discriminative regions, while discouraging unbalanced segmentations. We introduce high uncertainty as a criterion to localize non-discriminative regions that do not affect classifier decision, and describe it with original Kullback-Leibler (KL) divergence losses evaluating the deviation of posterior predictions from the uniform distribution. Our KL terms encourage high uncertainty of the model when the latter inputs the latent non-discriminative regions. Our loss integrates: (i) a cross-entropy seeking a foreground, where model confidence about class prediction is high; (ii) a KL regularizer seeking a background, where model uncertainty is high; and (iii) log-barrier terms discouraging unbalanced segmentations. Comprehensive experiments and ablation studies over the public GlaS colon cancer data and a Camelyon16 patch-based benchmark for breast cancer show substantial improvements over state-of-the-art WSL methods, and confirm the effect of our new regularizers (our code is publicly available at https://github.com/sbelharbi/deep-wsl-histo-min-max-uncertainty). Soufiane Belharbi, Jérôme Rony, Jose Dolz, Ismail Ben Ayed, Luke McCaffrey, Eric Granger |
IEEE Trans. Medical Imaging | 4 |
| 2021 | Variational Fair ClusteringabstractWe propose a general variational framework of fair clustering, which integrates an original Kullback-Leibler (KL) fairness term with a large class of clustering objectives, including prototype or graph based. Fundamentally different from the existing combinatorial and spectral solutions, our variational multi-term approach enables to control the trade-off levels between the fairness and clustering objectives. We derive a general tight upper bound based on a concave-convex decomposition of our fairness term, its Lipschitz-gradient property and the Pinsker’s inequality. Our tight upper bound can be jointly optimized with various clustering objectives, while yielding a scalable solution, with convergence guarantee. Interestingly, at each iteration, it performs an independent update for each assignment variable. Therefore, it can be easily distributed for large-scale datasets. This scalability is important as it enables to explore different trade-off levels between the fairness and clustering objectives. Unlike spectral relaxation, our formulation does not require computing its eigenvalue decomposition. We report comprehensive evaluations and comparisons with state-of-the-art methods over various fair clustering benchmarks, which show that our variational formulation can yield highly competitive solutions in terms of fairness and clustering objectives. Imtiaz Masud Ziko, Jing Yuan 0001, Eric Granger, Ismail Ben Ayed |
AAAI | 4 |
| 2021 | Few-Shot Segmentation Without Meta-Learning: A Good Transductive Inference Is All You Need?abstractWe show that the way inference is performed in few-shot segmentation tasks has a substantial effect on performances—an aspect often overlooked in the literature in favor of the meta-learning paradigm. We introduce a transductive inference for a given query image, leveraging the statistics of its unlabeled pixels, by optimizing a new loss containing three complementary terms: i) the cross-entropy on the labeled support pixels; ii) the Shannon entropy of the posteriors on the unlabeled query-image pixels; and iii) a global KL-divergence regularizer based on the proportion of the predicted foreground. As our inference uses a simple linear classifier of the extracted features, its computational load is comparable to inductive inference and can be used on top of any base training. Foregoing episodic training and using only standard cross-entropy training on the base classes, our inference yields competitive performances on standard benchmarks in the 1-shot scenarios. As the number of available shots increases, the gap in performances widens: on PASCAL-5i, our method brings about 5% and 6% improvements over the state-of-the-art, in the 5- and 10-shot scenarios, respectively. Furthermore, we introduce a new setting that includes domain shifts, where the base and novel classes are drawn from different datasets. Our method achieves the best performances in this more realistic setting. Our code is freely available online: https://github.com/mboudiaf/RePRI-for-Few-Shot-Segmentation. Malik Boudiaf, Hoel Kervadec, Imtiaz Masud Ziko, Pablo Piantanida, Ismail Ben Ayed, Jose Dolz |
CVPR | 5 |
| 2021 | Augmented Lagrangian Adversarial AttacksabstractAdversarial attack algorithms are dominated by penalty methods, which are slow in practice, or more efficient distance-customized methods, which are heavily tailored to the properties of the distance considered. We propose a white-box attack algorithm to generate minimally perturbed adversarial examples based on Augmented Lagrangian principles. We bring several algorithmic modifications, which have a crucial effect on performance. Our attack enjoys the generality of penalty methods and the computational efficiency of distance-customized algorithms, and can be readily used for a wide set of distances. We compare our attack to state-of-the-art methods on three datasets and several models, and consistently obtain competitive performances with similar or lower computational complexity. Jérôme Rony, Eric Granger, Marco Pedersoli, Ismail Ben Ayed |
ICCV | 4 |
| 2021 | Realistic evaluation of transductive few-shot learningabstractTransductive inference is widely used in few-shot learning, as it leverages the statistics of the unlabeled query set of a few-shot task, typically yielding substantially better performances than its inductive counterpart. The current few-shot benchmarks use perfectly class-balanced tasks at inference. We argue that such an artificial regularity is unrealistic, as it assumes that the marginal label probability of the testing samples is known and fixed to the uniform distribution. In fact, in realistic scenarios, the unlabeled query sets come with arbitrary and unknown label marginals. We introduce and study the effect of arbitrary class distributions within the query sets of few-shot tasks at inference, removing the class-balance artefact. Specifically, we model the marginal probabilities of the classes as Dirichlet-distributed random variables, which yields a principled and realistic sampling within the simplex. This leverages the current few-shot benchmarks, building testing tasks with arbitrary class distributions. We evaluate experimentally state-of-the-art transductive methods over 3 widely used data sets, and observe, surprisingly, substantial performance drops, even below inductive methods in some cases. Furthermore, we propose a generalization of the mutual-information loss, based on α-divergences, which can handle effectively class-distribution variations. Empirically, we show that our transductive α-divergence optimization outperforms state-of-the-art methods across several data sets, models and few-shot settings. Olivier Veilleux, Malik Boudiaf, Pablo Piantanida, Ismail Ben Ayed |
NeurIPS | 4 |
| 2021 | On the Texture Bias for Few-Shot CNN SegmentationabstractDespite the initial belief that Convolutional Neural Networks (CNNs) are driven by shapes to perform visual recognition tasks, recent evidence suggests that texture bias in CNNs provides higher performing models when learning on large labeled training datasets. This contrasts with the perceptual bias in the human visual cortex, which has a stronger preference towards shape components. Perceptual differences may explain why CNNs achieve human-level performance when large labeled datasets are available, but their performance significantly degrades in low-labeled data scenarios, such as few-shot semantic segmentation. To remove the texture bias in the context of few-shot learning, we propose a novel architecture that integrates a set of Difference of Gaussians (DoG) to attenuate high-frequency local components in the feature space. This produces a set of modified feature maps, whose high-frequency components are diminished at different standard deviation values of the Gaussian distribution in the spatial domain. As this results in multiple feature maps for a single image, we employ a bi-directional convolutional long-short-term-memory to efficiently merge the multi scale-space representations. We perform extensive experiments on three well-known few-shot segmentation benchmarks -Pascal i5, COCO-20i and FSS-1000- and demonstrate that our method outperforms state-of-the-art approaches in two datasets under the same conditions. Reza Azad, Abdur Razzaq Fayjie, Claude Kauffmann, Ismail Ben Ayed, Marco Pedersoli, Jose Dolz |
WACV | 4 |
| 2021 | Deep Active Learning for Joint Classification & Segmentation with Weak AnnotatorabstractCNN visualization and interpretation methods, like class-activation maps (CAMs), are typically used to highlight the image regions linked to class predictions. These models allow to simultaneously classify images and extract class-dependent saliency maps, without the need for costly pixel-level annotations. However, they typically yield segmentations with high false-positive rates and, therefore, coarse visualisations, more so when processing challenging images, as encountered in histology. To mitigate this issue, we propose an active learning (AL) framework, which progressively integrates pixel-level annotations during training. Given training data with global image-level labels, our deep weakly-supervised learning model jointly performs supervised image-level classification and active learning for segmentation, integrating pixel annotations by an oracle. Unlike standard AL methods that focus on sample selection, we also leverage large numbers of unlabeled images via pseudo-segmentations (i.e., self-learning at the pixel level), and integrate them with the oracle-annotated samples during training. We report extensive experiments over two challenging benchmarks -- high-resolution medical images (histology GlaS data for colon cancer) and natural images (CUB-200-2011 for bird species). Our results indicate that, by simply using random sample selection, the proposed approach can significantly outperform state-of the-art CAMs and AL methods, with an identical oracle-supervision budget. Our code is publicly available. Soufiane Belharbi, Ismail Ben Ayed, Luke McCaffrey, Eric Granger |
WACV | 2 |
| 2021 | Learning Data Augmentation with Online Bilevel Optimization for Image ClassificationabstractData augmentation is a key practice in machine learning for improving generalization performance. However, finding the best data augmentation hyperparameters requires domain knowledge or a computationally demanding search. We address this issue by proposing an efficient approach to automatically train a network that learns an effective distribution of transformations to improve its generalization. Using bilevel optimization, we directly optimize the data augmentation parameters using a validation set. This framework can be used as a general solution to learn the optimal data augmentation jointly with an end task model like a classifier. Results show that our joint training method produces an image classification accuracy that is comparable to or better than carefully hand-crafted data augmentation. Yet, it does not need an expensive external validation loop on the data augmentation hyperparameters. Saypraseuth Mounsaveng, Issam H. Laradji, Ismail Ben Ayed, David Vázquez 0001, Marco Pedersoli |
WACV | 3 |
| 2021 | Flow guided mutual attention for person re-identification
Madhu Kiran, Amran Bhuiyan, Le Thanh Nguyen-Meidine, Louis-Antoine Blais-Morin, Ismail Ben Ayed, Eric Granger |
Image Vis. Comput. | 5 |
| 2021 | Boundary loss for highly unbalanced segmentation
Hoel Kervadec, Jihene Bouchtiba, Christian Desrosiers, Eric Granger, Jose Dolz, Ismail Ben Ayed |
Medical Image Anal. | 6 |
| 2021 | Deep Clustering: On the Link Between Discriminative Models and K-MeansabstractIn the context of recent deep clustering studies, discriminative models dominate the literature and report the most competitive performances. These models learn a deep discriminative neural network classifier in which the labels are latent. Typically, they use multinomial logistic regression posteriors and parameter regularization, as is very common in supervised learning. It is generally acknowledged that discriminative objective functions (e.g., those based on the mutual information or the KL divergence) are more flexible than generative approaches (e.g., K-means) in the sense that they make fewer assumptions about the data distributions and, typically, yield much better unsupervised deep learning results. On the surface, several recent discriminative models may seem unrelated to K-means. This study shows that these models are, in fact, equivalent to K-means under mild conditions and common posterior models and parameter regularization. We prove that, for the commonly used logistic regression posteriors, maximizing the L2L2 regularized mutual information via an approximate alternating direction method (ADM) is equivalent to minimizing a soft and regularized K-means loss. Our theoretical analysis not only connects directly several recent state-of-the-art discriminative models to K-means, but also leads to a new soft and regularized deep K-means algorithm, which yields competitive performance on several image clustering benchmarks. Mohammed Jabi, Marco Pedersoli, Amar Mitiche, Ismail Ben Ayed |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Constrained Domain Adaptation for Image SegmentationabstractDomain Adaption tasks have recently attracted substantial attention in computer vision as they improve the transferability of deep network models from a source to a target domain with different characteristics. A large body of state-of-the-art domain-adaptation methods was developed for image classification purposes, which may be inadequate for segmentation tasks. We propose to adapt segmentation networks with a constrained formulation, which embeds domain-invariant prior knowledge about the segmentation regions. Such knowledge may take the form of anatomical information, for instance, structure size or shape, which can be known a priori or learned from the source samples via an auxiliary task. Our general formulation imposes inequality constraints on the network predictions of unlabeled or weakly labeled target samples, thereby matching implicitly the prediction statistics of the target and source domains, with permitted uncertainty of prior knowledge. Furthermore, our inequality constraints easily integrate weak annotations of the target data, such as image-level tags. We address the ensuing constrained optimization problem with differentiable penalties, fully suited for conventional stochastic gradient descent approaches. Unlike common two-step adversarial training, our formulation is based on a single segmentation network, which simplifies adaptation, while improving training quality. Comparison with state-of-the-art adaptation methods reveals considerably better performance of our model on two challenging tasks. Particularly, it consistently yields a performance gain of 1-4% Dice across architectures and datasets. Our results also show robustness to imprecision in the prior knowledge. The versatility of our novel approach can be readily used in various segmentation problems, with code available publicly. Mathilde Bateson, Jose Dolz, Hoel Kervadec, Hervé Lombaert, Ismail Ben Ayed |
IEEE Trans. Medical Imaging | 5 |
| 2020 | A Unifying Mutual Information View of Metric Learning: Cross-Entropy vs. Pairwise Losses
Malik Boudiaf, Jérôme Rony, Imtiaz Masud Ziko, Eric Granger, Marco Pedersoli, Pablo Piantanida, Ismail Ben Ayed |
ECCV (6) | 7 |
| 2020 | Laplacian Regularized Few-Shot LearningabstractWe propose a transductive Laplacian-regularized inference for few-shot tasks. Given any feature embedding learned from the base classes, we minimize a quadratic binary-assignment function containing two terms: (1) a unary term assigning query samples to the nearest class prototype, and (2) a pairwise Laplacian term encouraging nearby query samples to have consistent label assignments. Our transductive inference does not re-train the base model, and can be viewed as a graph clustering of the query set, subject to supervision constraints from the support set. We derive a computationally efficient bound optimizer of a relaxation of our function, which computes independent (parallel) updates for each query sample, while guaranteeing convergence. Following a simple cross-entropy training on the base classes, and without complex meta-learning strategies, we conducted comprehensive experiments over five few-shot learning benchmarks. Our LaplacianShot consistently outperforms state-of-the-art methods by significant margins across different models, settings, and data sets. Furthermore, our transductive inference is very fast, with computational times that are close to inductive inference, and can be used for large-scale few-shot tasks. Imtiaz Masud Ziko, Jose Dolz, Eric Granger, Ismail Ben Ayed |
ICML | 4 |
| 2020 | Source-Relaxed Domain Adaptation for Image Segmentation
Mathilde Bateson, Hoel Kervadec, Jose Dolz, Hervé Lombaert, Ismail Ben Ayed |
MICCAI (1) | 5 |
| 2020 | Cost-Sensitive Regularization for Diabetic Retinopathy Grading from Eye Fundus Images
Adrian Galdran, Jose Dolz, Hadi Chakor, Hervé Lombaert, Ismail Ben Ayed |
MICCAI (5) | 5 |
| 2020 | Information Maximization for Few-Shot LearningabstractWe introduce Transductive Infomation Maximization (TIM) for few-shot learning. Our method maximizes the mutual information between the query features and their label predictions for a given few-shot task, in conjunction with a supervision loss based on the support set. Furthermore, we propose a new alternating-direction solver for our mutual-information loss, which substantially speeds up transductive inference convergence over gradient-based optimization, while yielding similar accuracy. TIM inference is modular: it can be used on top of any base-training feature extractor. Following standard transductive few-shot settings, our comprehensive experiments demonstrate that TIM outperforms state-of-the-art methods significantly across various datasets and networks, while used on top of a fixed feature extractor trained with simple cross-entropy on the base classes, without resorting to complex meta-learning schemes. It consistently brings between 2% and 5% improvement in accuracy over the best performing method, not only on all the well-established few-shot benchmarks but also on more challenging scenarios, with domain shifts and larger numbers of classes. Malik Boudiaf, Imtiaz Masud Ziko, Jérôme Rony, Jose Dolz, Pablo Piantanida, Ismail Ben Ayed |
NeurIPS | 6 |
| 2020 | Pose Guided Gated Fusion for Person Re-identificationabstractPerson re-identification is an important yet challenging problem in visual recognition. Despite the recent advances with deep learning (DL) models for spatio-temporal and multi-modal fusion, re-identification approaches often fail to leverage the contextual information (e.g., pose and illumination) to dynamically select the most discriminant con-volutional filters (i.e., appearance features) for feature representation and inference. State-of-the-art techniques for gated fusion employ complex dedicated part- or attention-based architectures for late fusion, and do not incorporate pose and appearance information to train the backbone network. In this paper, a new DL model is proposed for pose-guided re-identification, comprised of a deep backbone, pose estimation, and gated fusion network. Given a query image of an individual, the backbone convolutional NN produces a feature embedding required for pair-wise matching with embeddings for reference images, where feature maps from the pose network and from mid-level CNN layers are combined by the gated fusion network to generate pose-guided gating. The proposed framework allows to dynamically activate the most discriminant CNN filters based on pose information in order to perform a finer grained recognition. Extensive experiments on three challenging benchmark datasets indicate that integrating the pose-guided gated fusion into the state-of-the-art re-identification backbone architecture allows to improve their recognition accuracy. Experimental results also support our intuition on the advantages of gating backbone appearance information using the pose feature maps at mid-level CNN layers. Amran Bhuiyan, Yang Liu 0093, Parthipan Siva, Mehrsan Javan Roshtkhari, Ismail Ben Ayed, Eric Granger |
WACV | 5 |
| 2020 | Discretely-constrained deep network for weakly supervised segmentation
Jizong Peng, Hoel Kervadec, Jose Dolz, Ismail Ben Ayed, Marco Pedersoli, Christian Desrosiers |
Neural Networks | 4 |
| 2020 | An ILP Model for Multi-Label MRFs With Connectivity ConstraintsabstractInteger Linear Programming (ILP) formulations of multi-label Markov random fields (MRFs) models with global connectivity priors were investigated previously in computer vision. In these works, only Linear Programming (LP) relaxations [1] or simplified versions [2] of the problem were solved. This paper investigates the ILP of MRF with exact connectivity priors via a branch-and-cut method, which provably finds globally optimal solutions. It enforces connectivity priors iteratively by a cutting plane method, and provides feasible solutions with a guarantee on sub-optimality even if we terminate it earlier. The proposed ILP can be applied as a post-processing method on top of any existing multi-label segmentation approach. As it provides globally optimal solution, it can be used off-line to serve as quality check for any fast on-line algorithm. Furthermore, the scribble based model presented in this paper could be potentially used to generate ground-truth proposals for any deep learning based segmentation. We demonstrate the power and usefulness of our model by extensive experiments on the BSDS500 and PASCAL VOC dataset. The experiments show that our proposed model achieves great performance, yielding provably global optimum in most instances and that provably good optimization solutions also provide good segmentation accuracy, even with the limited computing time of few seconds. Ruobing Shen, Bo Tang 0017, Andrea Lodi 0001, Andrea Tramontani, Ismail Ben Ayed |
IEEE Trans. Image Process. | 5 |
| 2019 | Beyond Gradient Descent for Regularized Segmentation LossesabstractThe simplicity of gradient descent (GD) made it the default method for training ever-deeper and complex neural networks. Both loss functions and architectures are often explicitly tuned to be amenable to this basic local optimization. In the context of weakly-supervised CNN segmentation, we demonstrate a well-motivated loss function where an alternative optimizer (ADM) achieves the state-of-the-art while GD performs poorly. Interestingly, GD obtains its best result for a "smoother" tuning of the loss function. The results are consistent across different network architectures. Our loss is motivated by well-understood MRF/CRF regularization models in "shallow" segmentation and their known global solvers. Our work suggests that network design/training should pay more attention to optimization methods. Dmitrii Marin, Meng Tang 0001, Ismail Ben Ayed, Yuri Boykov |
CVPR | 3 |
| 2019 | Decoupling Direction and Norm for Efficient Gradient-Based L2 Adversarial Attacks and DefensesabstractResearch on adversarial examples in computer vision tasks has shown that small, often imperceptible changes to an image can induce misclassification, which has security implications for a wide range of image processing systems. Considering L2 norm distortions, the Carlini and Wagner attack is presently the most effective white-box attack in the literature. However, this method is slow since it performs a line-search for one of the optimization terms, and often requires thousands of iterations. In this paper, an efficient approach is proposed to generate gradient-based attacks that induce misclassifications with low L2 norm, by decoupling the direction and the norm of the adversarial perturbation that is added to the image. Experiments conducted on the MNIST, CIFAR-10 and ImageNet datasets indicate that our attack achieves comparable results to the state-of-the-art (in terms of L2 norm) with considerably fewer iterations (as few as 100 iterations), which opens the possibility of using these attacks for adversarial training. Models trained with our attack achieve state-of-the-art robustness against white-box gradient-based L2 attacks on the MNIST and CIFAR-10 datasets, outperforming the Madry defense when the attacks are limited to a maximum norm. Jérôme Rony, Luiz G. Hafemann, Luiz Eduardo Soares de Oliveira, Ismail Ben Ayed, Robert Sabourin, Eric Granger |
CVPR | 4 |
| 2019 | Constrained Domain Adaptation for Segmentation
Mathilde Bateson, Hoel Kervadec, Jose Dolz, Hervé Lombaert, Ismail Ben Ayed |
MICCAI (2) | 5 |
| 2019 | Curriculum Semi-supervised Segmentation
Hoel Kervadec, Jose Dolz, Eric Granger, Ismail Ben Ayed |
MICCAI (2) | 4 |
| 2019 | Kernel Cuts: Kernel and Spectral Clustering Meet Regularization
Meng Tang 0001, Dmitrii Marin, Ismail Ben Ayed, Yuri Boykov |
Int. J. Comput. Vis. | 3 |
| 2019 | Constrained-CNN losses for weakly supervised segmentation
Hoel Kervadec, Jose Dolz, Meng Tang 0001, Eric Granger, Yuri Boykov, Ismail Ben Ayed |
Medical Image Anal. | 6 |
| 2019 | Kernel Clustering: Density Biases and SolutionsabstractKernel methods are popular in clustering due to their generality and discriminating power. However, we show that many kernel clustering criteria have density biases theoretically explaining some practically significant artifacts empirically observed in the past. For example, we provide conditions and formally prove the density mode isolation bias in kernel K-means for a common class of kernels. We call it Breiman's bias due to its similarity to the histogram mode isolation previously discovered by Breiman in decision tree learning with Gini impurity. We also extend our analysis to other popular kernel clustering methods, e.g., average/normalized cut or dominant sets, where density biases can take different forms. For example, splitting isolated points by cut-based criteria is essentially the sparsest subset bias, which is the opposite of the density mode bias. Our findings suggest that a principled solution for density biases in kernel clustering should directly address data inhomogeneity. We show that density equalization can be implicitly achieved using either locally adaptive weights or locally adaptive kernels. Moreover, density equalization makes many popular kernel clustering objectives equivalent. Our synthetic and real data experiments illustrate density biases and proposed solutions. We anticipate that theoretical understanding of kernel clustering limitations and their principled solutions will be important for a broad spectrum of data analysis applications across the disciplines. Dmitrii Marin, Meng Tang 0001, Ismail Ben Ayed, Yuri Boykov |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | HyperDense-Net: A Hyper-Densely Connected CNN for Multi-Modal Image SegmentationabstractRecently, dense connections have attracted substantial attention in computer vision because they facilitate gradient flow and implicit deep supervision during training. Particularly, DenseNet that connects each layer to every other layer in a feed-forward fashion and has shown impressive performances in natural image classification tasks. We propose HyperDenseNet, a 3-D fully convolutional neural network that extends the definition of dense connectivity to multi-modal segmentation problems. Each imaging modality has a path, and dense connections occur not only between the pairs of layers within the same path but also between those across different paths. This contrasts with the existing multi-modal CNN approaches, in which modeling several modalities relies entirely on a single joint layer (or level of abstraction) for fusion, typically either at the input or at the output of the network. Therefore, the proposed network has total freedom to learn more complex combinations between the modalities, within and in-between all the levels of abstraction, which increases significantly the learning representation. We report extensive evaluations over two different and highly competitive multi-modal brain tissue segmentation challenges, iSEG 2017 and MRBrainS 2013, with the former focusing on six month infant data and the latter on adult images. HyperDenseNet yielded significant improvements over many state-of-the-art segmentation networks, ranking at the top on both benchmarks. We further provide a comprehensive experimental analysis of features re-use, which confirms the importance of hyper-dense connections in multi-modal representation learning. Our code is publicly available. Jose Dolz, Karthik Gopinath, Jing Yuan 0001, Hervé Lombaert, Christian Desrosiers, Ismail Ben Ayed |
IEEE Trans. Medical Imaging | 6 |
| 2019 | Benchmark on Automatic Six-Month-Old Infant Brain Segmentation Algorithms: The iSeg-2017 ChallengeabstractAccurate segmentation of infant brain magnetic resonance (MR) images into white matter (WM), gray matter (GM), and cerebrospinal fluid (CSF) is an indispensable foundation for early studying of brain growth patterns and morphological changes in neurodevelopmental disorders. Nevertheless, in the isointense phase (approximately 6-9 months of age), due to inherent myelination and maturation process, WM and GM exhibit similar levels of intensity in both T1-weighted (T1w) and T2-weighted (T2w) MR images, making tissue segmentation very challenging. Despite many efforts were devoted to brain segmentation, only few studies have focused on the segmentation of 6-month infant brain images. With the idea of boosting methodological development in the community, iSeg-2017 challenge (http://iseg2017.web.unc.edu) provides a set of 6-month infant subjects with manual labels for training and testing the participating methods. Among the 21 automatic segmentation methods participating in iSeg-2017, we review the 8 top-ranked teams, in terms of Dice ratio, modified Hausdorff distance and average surface distance, and introduce their pipelines, implementations, as well as source codes. We further discuss limitations and possible future directions. We hope the dataset in iSeg-2017 and this review article could provide insights into methodological development for the community. Li Wang 0026, Dong Nie, Élodie Puybareau, Jose Dolz, Qian Zhang 0066, Fan Wang 0023, Zhengwang Wu, Jiawei Chen 0001, Kim-Han Thung, Toan Duc Bui, Jitae Shin, Guodong Zeng, Guoyan Zheng, Vladimir S. Fonov, Andrew Doyle, Yongchao Xu, Pim Moeskops, Josien P. W. Pluim, Christian Desrosiers, Ismail Ben Ayed, Gerard Sanroma, Oualid M. Benkarim, Adrià Casamitjana, Verónica Vilaplana, Weili Lin, Gang Li 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 22 |
| 2018 | On Regularized Losses for Weakly-supervised CNN Segmentation
Meng Tang 0001, Federico Perazzi, Abdelaziz Djelouah, Ismail Ben Ayed, Christopher Schroers, Yuri Boykov |
ECCV (16) | 4 |
| 2018 | Scalable Laplacian K-modesabstractWe advocate Laplacian K-modes for joint clustering and density mode finding, and propose a concave-convex relaxation of the problem, which yields a parallel algorithm that scales up to large datasets and high dimensions. We optimize a tight bound (auxiliary function) of our relaxation, which, at each iteration, amounts to computing an independent update for each cluster-assignment variable, with guar- anteed convergence. Therefore, our bound optimizer can be trivially distributed for large-scale data sets. Furthermore, we show that the density modes can be obtained as byproducts of the assignment variables via simple maximum-value operations whose additional computational cost is linear in the number of data points. Our formulation does not need storing a full affinity matrix and computing its eigenvalue decomposition, neither does it perform expensive projection steps and Lagrangian-dual inner iterates for the simplex constraints of each point. Fur- thermore, unlike mean-shift, our density-mode estimation does not require inner- loop gradient-ascent iterates. It has a complexity independent of feature-space dimension, yields modes that are valid data points in the input set and is appli- cable to discrete domains as well as arbitrary kernels. We report comprehensive experiments over various data sets, which show that our algorithm yields very competitive performances in term of optimization quality (i.e., the value of the discrete-variable objective at convergence) and clustering accuracy. Imtiaz Masud Ziko, Eric Granger, Ismail Ben Ayed |
NeurIPS | 3 |
| 2017 | DOPE: Distributed Optimization for Pairwise Energies
Jose Dolz, Ismail Ben Ayed, Christian Desrosiers |
CVPR | 2 |
| 2017 | Unbiased Shape Compactness for Segmentation
Jose Dolz, Ismail Ben Ayed, Christian Desrosiers |
MICCAI (1) | 2 |
| 2017 | Regularised differentiation for image derivativesabstractThis study investigates a regularised differentiation method to estimate image derivatives. The scheme minimises an integral functional containing an anti‐differentiation data discrepancy term and a smoothness regularisation term. When discretised, the Euler–Lagrange necessary conditions for a minimum of the functional yield a large scale sparse system of linear equations, which can be solved efficiently by Jacobi/Gauss–Seidel iterations. The authors investigate the impact of the method in the context of two important problems in computer vision: optical flow and scene flow estimation. Quantitative results, using the Middlebury dataset and other real and synthetic images, show that the authors’ regularised differentiation scheme outperforms standard derivative definitions by smoothed finite differences, which are commonly used in motion analysis. The method can be readily used in various other image analysis problems. Yosra Mathlouthi, Amar Mitiche, Ismail Ben Ayed |
IET Image Process. | 3 |
| 2017 | Local Submodularization for Binary Pairwise EnergiesabstractMany computer vision problems require optimization of binary non-submodular energies. We propose a general optimization framework based on local submodular approximations (LSA). Unlike standard LP relaxation methods that linearize the whole energy globally, our approach iteratively approximates the energy locally. On the other hand, unlike standard local optimization methods (e.g., gradient descent or projection techniques) we use non-linear submodular approximations and optimize them without leaving the domain of integer solutions. We discuss two specific LSA algorithms based on trust region and auxiliary function principles, LSA-TR and LSA-AUX. The proposed methods obtain state-of-the-art results on a wide range of applications such as binary deconvolution, curvature regularization, inpainting, segmentation with repulsion and two types of shape priors. Finally, we discuss a move-making extension to the LSA-TR approach. While our paper is focused on pairwise energies, our ideas extend to higher-order problems. The code is available online. Lena Gorelick, Yuri Boykov, Olga Veksler, Ismail Ben Ayed, Andrew Delong |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2016 | Normalized Cut Meets MRF
Meng Tang 0001, Dmitrii Marin, Ismail Ben Ayed, Yuri Boykov |
ECCV (2) | 3 |
| 2015 | Volumetric Bias in Segmentation and Reconstruction: Secrets and SolutionsabstractMany standard optimization methods for segmentation and reconstruction compute ML model estimates for appearance or geometry of segments, e.g. Zhu-Yuille [23], Torr [20], Chan-Vese [6], GrabCut [18], Delong et al. [8]. We observe that the standard likelihood term in these formu-lations corresponds to a generalized probabilistic K-means energy. In learning it is well known that this energy has a strong bias to clusters of equal size [11], which we express as a penalty for KL divergence from a uniform distribution of cardinalities. However, this volumetric bias has been mostly ignored in computer vision. We demonstrate signif- icant artifacts in standard segmentation and reconstruction methods due to this bias. Moreover, we propose binary and multi-label optimization techniques that either (a) remove this bias or (b) replace it by a KL divergence term for any given target volume distribution. Our general ideas apply to continuous or discrete energy formulations in segmenta- tion, stereo, and other reconstruction problems. Yuri Boykov, Hossam Isack, Carl Olsson, Ismail Ben Ayed |
ICCV | 4 |
| 2015 | Secrets of GrabCut and Kernel K-MeansabstractThe log-likelihood energy term in popular model-fitting segmentation methods, e.g. [39, 8, 28, 10], is presented as a generalized "probabilistic K-means" energy [16] for color space clustering. This interpretation reveals some limitations, e.g. over-fitting. We propose an alternative approach to color clustering using kernel K-means energy with well-known properties such as non-linear separation and scalability to higher-dimensional feature spaces. Our bound formulation for kernel K-means allows to combine general pair-wise feature clustering methods with image grid regularization using graph cuts, similarly to standard color model fitting techniques for segmentation. Unlike histogram or GMM fitting [39, 28], our approach is closely related to average association and normalized cut. But, in contrast to previous pairwise clustering algorithms, our approach can incorporate any standard geometric regularization in the image domain. We analyze extreme cases for kernel bandwidth (e.g. Gini bias) and demonstrate effectiveness of KNN-based adaptive bandwidth strategies. Our kernel K-means approach to segmentation benefits from higher-dimensional features where standard model fitting fails. Meng Tang 0001, Ismail Ben Ayed, Dmitrii Marin, Yuri Boykov |
ICCV | 2 |
| 2015 | Right ventricle segmentation from cardiac MRI: A collation study
Caroline Petitjean, Maria A. Zuluaga, Wenjia Bai, Jean-Nicolas Dacher, Damien Grosgeorge, Jérôme Caudron, Su Ruan, Ismail Ben Ayed, Manuel Jorge Cardoso, Hsiang-Chou Chen, Daniel Jimenez-Carretero, María J. Ledesma-Carbayo, Christos Davatzikos, Jimit Doshi, Güray Erus, Oskar M. O. Maier, Cyrus M. S. Nambakhsh, Yangming Ou, Sébastien Ourselin, Chun-Wei Peng, Nicholas S. Peters, Terry M. Peters, Martin Rajchl, Daniel Rueckert, Wenzhe Shi, Ching-Wei Wang, Haiyan Wang 0018, Jing Yuan 0001 |
Medical Image Anal. | 8 |
| 2015 | Distribution Matching with the Bhattacharyya Similarity: A Bound Optimization FrameworkabstractWe present efficient graph cut algorithms for three problems: (1) finding a region in an image, so that the histogram (or distribution) of an image feature within the region most closely matches a given model; (2) co-segmentation of image pairs and (3) interactive image segmentation with a user-provided bounding box. Each algorithm seeks the optimum of a global cost function based on the Bhattacharyya measure, a convenient alternative to other matching measures such as the Kullback-Leibler divergence. Our functionals are not directly amenable to graph cut optimization as they contain non-linear functions of fractional terms, which make the ensuing optimization problems challenging. We first derive a family of parametric bounds of the Bhattacharyya measure by introducing an auxiliary labeling. Then, we show that these bounds are auxiliary functions of the Bhattacharyya measure, a result which allows us to solve each problem efficiently via graph cuts. We show that the proposed optimization procedures converge within very few graph cut iterations. Comprehensive and various experiments, including quantitative and comparative evaluations over two databases, demonstrate the advantages of the proposed algorithms over related works in regard to optimality, computational load, accuracy and flexibility. Ismail Ben Ayed, Kumaradevan Punithakumar, Shuo Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Submodularization for Binary Pairwise EnergiesabstractMany computer vision problems require optimization of binary non-submodular energies. We propose a general optimization framework based on local submodular approximations (LSA). Unlike standard LP relaxation methods that linearize the whole energy globally, our approach iteratively approximates the energies locally. On the other hand, unlike standard local optimization methods (e.g. gradient descent or projection techniques) we use non-linear submodular approximations and optimize them without leaving the domain of integer solutions. We discuss two specific LSA algorithms based on trust region and auxiliary function principles, LSA-TR and LSA-AUX. These methods obtain state-of-the-art results on a wide range of applications outperforming many standard techniques such as LBP, QPBO, and TRWS. While our paper is focused on pairwise energies, our ideas extend to higher-order problems. The code is available online. Lena Gorelick, Yuri Boykov, Olga Veksler, Ismail Ben Ayed, Andrew Delong |
CVPR | 4 |
| 2014 | Pseudo-bound Optimization for Binary Energies
Meng Tang 0001, Ismail Ben Ayed, Yuri Boykov |
ECCV (5) | 2 |
| 2014 | TRIC: Trust Region for Invariant Compactness and Its Application to Abdominal Aorta Segmentation
Ismail Ben Ayed, Brandon Miles, Gregory J. Garvin |
MICCAI (1) | 1 |
| 2014 | Convex-Relaxed Kernel Mapping for Image SegmentationabstractThis paper investigates a convex-relaxed kernel mapping formulation of image segmentation. We optimize, under some partition constraints, a functional containing two characteristic terms: 1) a data term, which maps the observation space to a higher (possibly infinite) dimensional feature space via a kernel function, thereby evaluating nonlinear distances between the observations and segments parameters and 2) a total-variation term, which favors smooth segment surfaces (or boundaries). The algorithm iterates two steps: 1) a convex-relaxation optimization with respect to the segments by solving an equivalent constrained problem via the augmented Lagrange multiplier method and 2) a convergent fixed-point optimization with respect to the segments parameters. The proposed algorithm can bear with a variety of image types without the need for complex and application-specific statistical modeling, while having the computational benefits of convex relaxation. Our solution is amenable to parallelized implementations on graphics processing units (GPUs) and extends easily to high dimensions. We evaluated the proposed algorithm with several sets of comprehensive experiments and comparisons, including: 1) computational evaluations over 3D medical-imaging examples and high-resolution large-size color photographs, which demonstrate that a parallelized implementation of the proposed method run on a GPU can bring a significant speed-up and 2) accuracy evaluations against five state-of-the-art methods over the Berkeley color-image database and a multimodel synthetic data set, which demonstrates competitive performances of the algorithm. Mohamed Ben Salah, Ismail Ben Ayed, Jing Yuan 0001, Hong Zhang 0013 |
IEEE Trans. Image Process. | 2 |
| 2014 | Regional Assessment of Cardiac Left Ventricular Myocardial Function via MRI Statistical FeaturesabstractAutomating the detection and localization of segmental (regional) left ventricle (LV) abnormalities in magnetic resonance imaging (MRI) has recently sparked an impressive research effort, with promising performances and a breadth of techniques. However, despite such an effort, the problem is still acknowledged to be challenging, with much room for improvements in regard to accuracy. Furthermore, most of the existing techniques are labor intensive, requiring delineations of the endo- and/or epi-cardial boundaries in all frames of a cardiac sequence. The purpose of this study is to investigate a real-time machine-learning approach which uses some image features that can be easily computed, but that nevertheless correlate well with the segmental cardiac function. Starting from a minimum user input in only one frame in a subject dataset, we build for all the regional segments and all subsequent frames a set of statistical MRI features based on a measure of similarity between distributions. We demonstrate that, over a cardiac cycle, the statistical features are related to the proportion of blood within each segment. Therefore, they can characterize segmental contraction without the need for delineating the LV boundaries in all the frames. We first seek the optimal direction along which the proposed image features are most descriptive via a linear discriminant analysis. Then, using the results as inputs to a linear support vector machine classifier, we obtain an abnormality assessment of each of the standard cardiac segments in real-time. We report a comprehensive experimental evaluation of the proposed algorithm over 928 cardiac segments obtained from 58 subjects. Compared against ground-truth evaluations by experienced radiologists, the proposed algorithm performed competitively, with an overall classification accuracy of 86.09% and a kappa measure of 0.73. Mariam Afshin, Ismail Ben Ayed, Kumaradevan Punithakumar, Max W. K. Law, Ali Islam, Aashish Goela, Terry M. Peters, Shuo Li 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2013 | Auxiliary Cuts for General Classes of Higher Order FunctionalsabstractSeveral recent studies demonstrated that higher order (non-linear) functionals can yield outstanding performances in the contexts of segmentation, co-segmentation and tracking. In general, higher order functionals result in difficult problems that are not amenable to standard optimizers, and most of the existing works investigated particular forms of such functionals. In this study, we derive general bounds for a broad class of higher order functionals. By introducing auxiliary variables and invoking the Jensen's inequality as well as some convexity arguments, we prove that these bounds are auxiliary functionals for various non-linear terms, which include but are not limited to several affinity measures on the distributions or moments of segment appearance and shape, as well as soft constraints on segment volume. From these general-form bounds, we state various non-linear problems as the optimization of auxiliary functionals by graph cuts. The proposed bound optimizers are derivative-free, and consistently yield very steep functional decreases, thereby converging within a few graph cuts. We report several experiments on color and medical data, along with quantitative comparisons to state of-the-art methods. The results demonstrate competitive performances of the proposed algorithms in regard to accuracy and convergence speed, and confirm their potential in various vision and medical applications. Ismail Ben Ayed, Lena Gorelick, Yuri Boykov |
CVPR | 1 |
| 2013 | Right Ventricle Segmentation with Probability Product Kernel Constraints
Cyrus M. S. Nambakhsh, Terry M. Peters, Ali Islam, Ismail Ben Ayed |
MICCAI (1) | 4 |
| 2013 | Left ventricle segmentation in MRI via convex relaxed distribution matching
Cyrus M. S. Nambakhsh, Jing Yuan 0001, Kumaradevan Punithakumar, Aashish Goela, Martin Rajchl, Terry M. Peters, Ismail Ben Ayed |
Medical Image Anal. | 7 |
| 2013 | Regional heart motion abnormality detection: An information theoretic approach
Kumaradevan Punithakumar, Ismail Ben Ayed, Ali Islam, Aashish Goela, Ian G. Ross, Jaron Chong, Shuo Li 0001 |
Medical Image Anal. | 2 |
| 2012 | Convex relaxation for image segmentation by kernel mappingabstractThis study proposes a novel multiregion image segmentation method using convex relaxation optimization and kernel mapping of the image data. The image data is transformed by a kernel function in order to support various image models while avoiding complex modeling. This is embedded implicitly in an objective function which is optimized by iterating a two-step strategy. First, a fixed point sequence is used to evaluate the regions parameters. Second, the image partition is updated by an efficient multiplier-based algorithm which uses the standard augmented Lagrangian method. A thorough experimental study is carried out over a multi-model synthetic dataset, the Berkeley database, as well as cardiac 3D data to show the effectiveness of the proposed method. Mohamed Ben Salah, Ismail Ben Ayed, Jing Yuan 0001, Zhijie Wang 0003, Hong Zhang 0013 |
ICIP | 2 |
| 2012 | Global Assessment of Cardiac Function Using Image Statistics in MRI
Mariam Afshin, Ismail Ben Ayed, Ali Islam, Aashish Goela, Terry M. Peters, Shuo Li 0001 |
MICCAI (2) | 2 |
| 2012 | Vertebral Body Segmentation in MRI via Convex Relaxation and Distribution Matching
Ismail Ben Ayed, Kumaradevan Punithakumar, Rashid Minhas, Rohit Joshi, Gregory J. Garvin |
MICCAI (1) | 1 |
| 2012 | Regional Heart Motion Abnormality Detection via Multiview Fusion
Kumaradevan Punithakumar, Ismail Ben Ayed, Ali Islam, Aashish Goela, Shuo Li 0001 |
MICCAI (2) | 2 |
| 2012 | Max-flow segmentation of the left ventricle by recovering subject-specific distributions via a bound of the Bhattacharyya measure
Ismail Ben Ayed, Huamei Chen, Kumaradevan Punithakumar, Ian G. Ross, Shuo Li 0001 |
Medical Image Anal. | 1 |
| 2012 | Active Curve Recovery of Region Boundary PatternsabstractThis study investigates the recovery of region boundary patterns in an image by a variational level set method which drives an active curve to coincide with boundaries on which a feature distribution matches a reference distribution. We formulate the scheme for both the Kullback-Leibler and the Bhattacharyya similarities, and apply it in two conditions: the simultaneous recovery of all region boundaries consistent with a given outline pattern, and segmentation in the presence of faded boundary segments. The first task uses an image-based geometric feature, and the second a photometric feature. In each case, the corresponding curve evolution equation can be viewed as a geodesic active contour (GAC) flow having a variable stopping function which depends on the feature distribution on the active curve. This affords a potent global representation of the target boundaries, which can effectively drive active curve segmentation in a variety of otherwise adverse conditions. Detailed experimentation shows that the scheme can significantly improve on current region and edge-based formulations. Mohamed Ben Salah, Ismail Ben Ayed, Amar Mitiche |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | A Convex Max-Flow Approach to Distribution-Based Figure-Ground SeparationabstractThis study investigates a convex relaxation approach to figure-ground separation with a global distribution matching prior evaluated by the Bhattacharyya measure. The problem amounts to finding a region that most closely matches a known model distribution. It has been previously addressed by curve evolution, which leads to suboptimal and computationally intensive algorithms, or by graph cuts, which result in metrication errors. Solving a sequence of convex subproblems, the proposed relaxation is based on a novel bound of the Bhattacharyya measure which yields an algorithm robust to initial conditions. Furthermore, we propose a novel flow configuration that accounts for labeling-function variations, unlike existing configurations. This leads to a new max-flow formulation which is dual to the convex relaxed subproblems we obtained. We further prove that such a formulation yields exact and global solutions to the original, nonconvex subproblems. A comprehensive experimental evaluation on the Microsoft GrabCut database demonstrates that our approach yields improvements in optimality and accuracy over related recent methods. Kumaradevan Punithakumar, Jing Yuan 0001, Ismail Ben Ayed, Shuo Li 0001, Yuri Boykov |
SIAM J. Imaging Sci. | 3 |
| 2011 | Assessment of Regional Myocardial Function via Statistical Features in MR Images
Mariam Afshin, Ismail Ben Ayed, Kumaradevan Punithakumar, Max W. K. Law, Ali Islam, Aashish Goela, Ian G. Ross, Terry M. Peters, Shuo Li 0001 |
MICCAI (3) | 2 |
| 2011 | Multiregion Image Segmentation by Parametric Kernel Graph CutsabstractThe purpose of this study is to investigate multiregion graph cut image partitioning via kernel mapping of the image data. The image data is transformed implicitly by a kernel function so that the piecewise constant model of the graph cut formulation becomes applicable. The objective function contains an original data term to evaluate the deviation of the transformed data, within each segmentation region, from the piecewise constant model, and a smoothness, boundary preserving regularization term. The method affords an effective alternative to complex modeling of the original image data while taking advantage of the computational benefits of graph cuts. Using a common kernel function, energy minimization typically consists of iterating image partitioning by graph cut iterations and evaluations of region parameters via fixed point computation. A quantitative and comparative performance assessment is carried out over a large number of experiments using synthetic grey level data as well as natural images from the Berkeley database. The effectiveness of the method is also demonstrated through a set of experiments with real images of a variety of types such as medical, synthetic aperture radar, and motion maps. Mohamed Ben Salah, Amar Mitiche, Ismail Ben Ayed |
IEEE Trans. Image Process. | 3 |
| 2010 | Graph cut segmentation with a global constraint: Recovering region distribution via a bound of the Bhattacharyya measureabstractThis study investigates an efficient algorithm for image segmentation with a global constraint based on the Bhattacharyya measure. The problem consists of finding a region consistent with an image distribution learned a priori. We derive an original upper bound of the Bhattacharyya measure by introducing an auxiliary labeling. From this upper bound, we reformulate the problem as an optimization of an auxiliary function by graph cuts. Then, we demonstrate that the proposed procedure converges and give a statistical interpretation of the upper bound. The algorithm requires very few iterations to converge, and finds nearly global optima. Quantitative evaluations and comparisons with state-of-the-art methods on the Microsoft GrabCut segmentation database demonstrated that the proposed algorithm brings improvements in regard to segmentation accuracy, computational efficiency, and optimality. We further demonstrate the flexibility of the algorithm in object tracking. Ismail Ben Ayed, Huamei Chen, Kumaradevan Punithakumar, Ian G. Ross, Shuo Li 0001 |
CVPR | 1 |
| 2010 | Finding image distributions on active curvesabstractThis study investigates an active curve functional which measures a similarity between the distribution of an image feature on the curve and a model distribution learned a priori. The curve evolution equation resulting from the minimization of this contour-based functional can be viewed as a geodesic active contour with a variable stopping function. The variable stopping function depends on the distribution of image feature on the curve and, therefore, can deal with difficult cases where the desired boundary corresponds to very weak image transitions. We ran several experiments supported by quantitative performance evaluations over several examples of segmentation and tracking of the left ventricle inner and outer boundaries in cardiac magnetic resonance image sequences. The results are significantly more accurate than with region-based and edge-based functionals. Ismail Ben Ayed, Amar Mitiche, Mohamed Ben Salah, Shuo Li 0001 |
CVPR | 1 |
| 2010 | Image partitioning with kernel mapping and graph cutsabstractA novel multiregion graph cut image partitioning method combined with kernel mapping is presented. A kernel function transforms implicitly the image data into data of a higher dimension so that the piecewise constant model of the graph cut formulation becomes applicable. The method yields an effective alternative to complex modeling of the original image data while taking advantage of the rapidity of graph cuts. A variety of noise models are, thus, considered by a single model. Using a common kernel function, we minimize the objective functional by iterating (1) regions parameters update and (2) image partitioning by graph cut iterations. A comparative performance evaluation is carried out over a large set of experiments using synthetic grey level data. Besides, a set of tests with real images such as SAR and medical images is shown to demonstrate the validity of the method. Mohamed Ben Salah, Amar Mitiche, Ismail Ben Ayed |
ICIP | 3 |
| 2010 | Regional Heart Motion Abnormality Detection via Information Measures and Unscented Kalman Filtering
Kumaradevan Punithakumar, Ismail Ben Ayed, Ali Islam, Ian G. Ross, Shuo Li 0001 |
MICCAI (1) | 2 |
| 2010 | Effective Level Set Image Segmentation With a Kernel Induced Data TermabstractThis study investigates level set multiphase image segmentation by kernel mapping and piecewise constant modeling of the image data thereof. A kernel function maps implicitly the original data into data of a higher dimension so that the piecewise constant model becomes applicable. This leads to a flexible and effective alternative to complex modeling of the image data. The method uses an active curve objective functional with two terms: an original term which evaluates the deviation of the mapped image data within each segmentation region from the piecewise constant model and a classic length regularization term for smooth region boundaries. Functional minimization is carried out by iterations of two consecutive steps: 1) minimization with respect to the segmentation by curve evolution via Euler-Lagrange descent equations and 2) minimization with respect to the regions parameters via fixed point iterations. Using a common kernel function, this step amounts to a mean shift parameter update. We verified the effectiveness of the method by a quantitative and comparative performance evaluation over a large number of experiments on synthetic images, as well as experiments with a variety of real images such as medical, satellite, and natural images, as well as motion maps. Mohamed Ben Salah, Amar Mitiche, Ismail Ben Ayed |
IEEE Trans. Image Process. | 3 |
| 2010 | Detection of left ventricular motion abnormality via information measures and Bayesian filteringabstractWe present an original information theoretic measure of heart motion based on the Shannon's differential entropy (SDE), which allows heart wall motion abnormality detection. Based on functional images, which are subject to noise and segmentation inaccuracies, heart wall motion analysis is acknowledged as a difficult problem, and as such, incorporation of prior knowledge is crucial for improving accuracy. Given incomplete, noisy data and a dynamic model, the Kalman filter, a well-known recursive Bayesian filter, is devised in this study to the estimation of the left ventricular (LV) cavity points. However, due to similarity between the statistical information of normal and abnormal heart motions, detecting and classifying abnormality is a challenging problem, which we investigate with a global measure based on the SDE. We further derive two other possible information theoretic abnormality detection criteria, one is based on Rényi entropy and the other on Fisher information. The proposed methods analyze wall motion quantitatively by constructing distributions of the normalized radial distance estimates of the LV cavity. Using 269 x 20 segmented LV cavities of short-axis MRI obtained from 30 subjects, the experimental analysis demonstrates that the proposed SDE criterion can lead to a significant improvement over other features that are prevalent in the literature related to the LV cavity, namely, mean radial displacement and mean radial velocity. Kumaradevan Punithakumar, Ismail Ben Ayed, Ian G. Ross, Ali Islam, Jaron Chong, Shuo Li 0001 |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2009 | Tracking Endocardial Boundary and Motion via Graph Cut Distribution Matching and Multiple Model Filtering
Kumaradevan Punithakumar, Ismail Ben Ayed, Ali Islam, Ian G. Ross, Shuo Li 0001 |
ACCV (3) | 2 |
| 2009 | Image segmentation in a kernel-induced spaceabstractA novel level set multiphase image segmentation method combined with kernel mapping is presented. A kernel function maps implicitly the original data into data of a higher dimension so that the piecewise constant model becomes applicable. The goal is to consider several types of noise by a single model. Gradient flow equations are iteratively derived in order to minimize the segmentation functional with respect to the partition, in a first step, and the regions parameters in a second step. Using a common kernel function, we verified the effectiveness of the method by a quantitative and comparative performance evaluation over experiments on synthetic images, as well as a variety of real images such as medical, SAR, and natural images. Mohamed Ben Salah, Amar Mitiche, Ismail Ben Ayed |
ICIP | 3 |
| 2009 | Left Ventricle Segmentation via Graph Cut Distribution Matching
Ismail Ben Ayed, Kumaradevan Punithakumar, Shuo Li 0001, Ali Islam, Jaron Chong |
MICCAI (1) | 1 |
| 2009 | Heart Motion Abnormality Detection via an Information Measure and Bayesian Filtering
Kumaradevan Punithakumar, Shuo Li 0001, Ismail Ben Ayed, Ian G. Ross, Ali Islam, Jaron Chong |
MICCAI (1) | 3 |
| 2009 | A Statistical Overlap Prior for Variational Image Segmentation
Ismail Ben Ayed, Shuo Li 0001, Ian G. Ross |
Int. J. Comput. Vis. | 1 |
| 2009 | Embedding Overlap Priors in Variational Left Ventricle TrackingabstractWe propose to embed overlap priors in variational tracking of the left ventricle (LV) in cardiac magnetic resonance (MR) sequences. The method consists of evolving two curves toward the LV endo- and epicardium boundaries. We derive the curve evolution equations by minimizing two functionals each containing an original overlap prior constraint. The latter measures the conformity of the overlap between the nonparametric (kernel-based) intensity distributions within the three target regions--LV cavity, myocardium and background-to a prior learned from a given segmentation of the first frame. The Bhattacharyya coefficient is used as an overlap measure. Different from existing intensity-driven constraints, the proposed priors do not assume implicitly that the overlap between the intensity distributions within different regions has to be minimal. This prevents both the papillary muscles from being included erroneously in the myocardium and the curves from spilling into the background. Although neither geometric training nor preprocessing were used, quantitative evaluation of the similarities between automatic and independent manual segmentations showed that the proposed method yields a competitive score in comparison with existing methods. This allows more flexibility in clinical use because our solution is based only on the current intensity data, and consequently, the results are not bounded to the characteristics, variability, and mathematical description of a finite training set. We also demonstrate experimentally that the overlap measures are approximately constant over a cardiac sequence, which allows to learn the overlap priors from a single frame. Ismail Ben Ayed, Shuo Li 0001, Ian G. Ross |
IEEE Trans. Medical Imaging | 1 |
| 2008 | Tracking distributions with an overlap priorabstractRecent studies have shown that embedding similarity/dissimilarity measures between distributions in the variational level set framework can lead to effective object segmentation/tracking algorithms. In this connection, existing methods assume implicitly that the overlap between the distributions of image data within the object and its background has to be minimal. Unfortunately, such assumption may not be valid in many important applications. This study investigates an overlap prior, which embeds knowledge about the overlap between the distributions of the object and the background in level set tracking. It consists of evolving a curve to delineate the target object in the current frame. The level set curve evolution equation is sought following the maximization of a functional containing three terms: (1) an original overlap prior which measures the conformity of overlap between the nonparametric (kernel-based) distributions within the object and the background to a learned description, (2) a term which measures the similarity between a model distribution of the object and the sample distribution inside the curve, and (3) a regularization term for smooth segmentation boundaries. The Bhattacharyya coefficient is used as an overlap measure. Apart from leading to a method which is more versatile than current ones, the overlap prior speeds up significantly the curve evolution. Comparisons and results demonstrate the advantages of the proposed prior over related methods, and its usefulness in important applications such as the left ventricle tracking in magnetic resonance (MR) images. Ismail Ben Ayed, Shuo Li 0001, Ian G. Ross |
CVPR | 1 |
| 2008 | Left Ventricle Tracking Using Overlap Priors
Ismail Ben Ayed, Yingli Lu, Shuo Li 0001, Ian G. Ross |
MICCAI (1) | 1 |
| 2008 | A Region Merging Prior for Variational Level Set Image SegmentationabstractIn current level set image segmentation methods, the number of regions is assumed to known beforehand. As a result, it remains constant during the optimization of the objective functional. How to allow it to vary is an important question which has been generally avoided. This study investigates a region merging prior related to regions area to allow the number of regions to vary automatically during curve evolution, thereby optimizing the objective functional implicitly with respect to the number of regions. We give a statistical interpretation to the coefficient of this prior to balance its effect systematically against the other functional terms. We demonstrate the validity and efficiency of the method by testing on real images of intensity, color, and motion. Ismail Ben Ayed, Amar Mitiche |
IEEE Trans. Image Process. | 1 |
| 2007 | Embedding a Region Merging Prior in Level Set Vector-Valued Image Segmentation
Ismail Ben Ayed, Amar Mitiche |
ACCV (1) | 1 |
| 2006 | A Partition Constrained Minimization Scheme for Efficient Multiphase Level Set Image SegmentationabstractThis study investigates a new multiphase minimization scheme which embeds a simple, efficient partition constraint directly in multiple level set evolution. Starting from an arbitrary initial partition, the minimization of the N-region segmentation functional is carried out following a first order expansion of the data term with an embedded partition constraint: if a point leaves a region, it goes to single other region. The method has a computational advantage over previous multiphase schemes, is stepwise optimal, i.e., permit to effect the maximum decrease of the functional at each evolution step, and is robust to initialization. The method is discussed by comparison with previous methods and experimental results are included to this effect. Ismail Ben Ayed, Amar Mitiche |
ICIP | 1 |
| 2006 | Variational Unsupervised Segmentation of Multi-Look Complex Polarimetric Images using a Wishart Observation ModelabstractWe address unsupervised variational segmentation of multi-look complex polarimetric images using a Wishart observation model via level sets. The methods consists of minimizing a functional containing an original data term derived from maximum likelihood Wishart approximation and a classical boundary length prior. The minimization is carried out efficiently by first order expansion of the data term and a new multiphase method which embeds a simple partition constraint directly in curve evolution. Results are shown on both synthetic and real images. Quantitative performance evaluation and comparisons with another method are also given. Ismail Ben Ayed, Amar Mitiche, Ziad Belhadj |
ICIP | 1 |
| 2006 | Polarimetric Image Segmentation via Maximum-Likelihood Approximation and Efficient Multiphase Level-SetsabstractThis study investigates a level set method for complex polarimetric image segmentation. It consists of minimizing a functional containing an original observation term derived from maximum-likelihood approximation and a complex Wishart/Gaussian image representation and a classical boundary length prior. The minimization is carried out efficiently by a new multiphase method which embeds a simple partition constraint directly in curve evolution to guarantee a partition of the image domain from an arbitrary initial partition. Results are shown on both synthetic and real images. Quantitative performance evaluation and comparisons are also given. Ismail Ben Ayed, Amar Mitiche, Ziad Belhadj |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2006 | Unsupervised Variational Image Segmentation/Classification Using a Weibull Observation ModelabstractStudies have shown that the Weibull distribution can model accurately a wide variety of images. Its parameters index a family of distributions which includes the exponential and approximations of the Gaussian and the Raleigh models widely used in image segmentation. This study investigates the Weibull distribution in unsupervised image segmentation and classification by a variational method. The data term of the segmentation functional measures the conformity of the image intensity in each region to a Weibull distribution whose parameters are determined jointly with the segmentation. Minimization of the functional is implemented by active curves via level sets and consists of iterations of two consecutive steps: curve evolution via Euler-Lagrange descent equations and evaluation of the Weibull distribution parameters. Experiments with synthetic and real images are described which verify the validity of method and its implementation. Ismail Ben Ayed, N. Hennane, Amar Mitiche |
IEEE Trans. Image Process. | 1 |
| 2005 | Level set curve evolution partitioning of polarimetric imagesabstractWe investigate a method of segmentation of multichannel polarimetric images, such as in laser illuminated or synthetic aperture radar, into a given but arbitrary number of regions via curve evolution and level sets. The algorithm consists of evolving closed curve, within an explicit correspondence between the interiors of curves and regions segmentation, to minimize a multivariate criterion corresponding to the complex Gaussian polarimetric model and a term of smoothness of the boundaries of regions. Results are shown on a polarimetric image. Ismail Ben Ayed, Amar Mitiche, Ziad Belhadj |
ICIP (1) | 1 |
| 2005 | Multiregion Level-Set Partitioning of Synthetic Aperture Radar ImagesabstractThe purpose of this study is to investigate Synthetic Aperture Radar (SAR) image segmentation into a given but arbitrary number of gamma homogeneous regions via active contours and level sets. The segmentation of SAR images is a difficult problem due to the presence of speckle which can be modeled as strong, multiplicative noise. The proposed algorithm consists of evolving simple closed planar curves within an explicit correspondence between the interiors of curves and regions of segmentation to minimize a criterion containing a term of conformity of data to a speckle model of noise and a term of regularization. Results are shown on both synthetic and real images. Ismail Ben Ayed, Amar Mitiche, Ziad Belhadj |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2004 | Sar image segmentation with active contours and level sets
Ismail Ben Ayed, Carlos Vázquez 0001, Amar Mitiche, Ziad Belhadj |
ICIP | 1 |
| 2004 | Image segmentation as regularized clustering: a fully global curve evolution methodabstractThe purpose of this study is to investigate image segmentation from the viewpoint of image data regularized clustering. From this viewpoint, segmentation into a fixed but arbitrary number N of regions is stated as the simultaneous minimization of N - 1 energy functional, each involving a single region and its complement. The resulting Euler-Lagrange curve evolution equations yield a partition at convergence provided the curves are initialized so as to define an arbitrary partition of the image domain. The method is implemented via level sets, and results are shown on synthetic and natural vectorial images. Carlos Vázquez 0001, Amar Mitiche, Ismail Ben Ayed |
ICIP | 3 |