EDBT 2026 Demo / reviewers in the wild / expert
Jean-Philippe Thiran
dblp:t/JeanPhilippeThiran
· DBLP profile ↗
168ranked-venue papers
4as first author
27since 2021 · last 2025
0000-0003-2938-9657ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 103 · 4 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 50 · 12 since 2021Artificial intelligence and machine learning · 46 · 12 since 2021Systems, architecture and hardware · 5Human-computer interaction and ubiquitous computing · 2Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | What to align in multimodal contrastive learning?abstractHumans perceive the world through multisensory integration, blending the information of different modalities to adapt their behavior.
Contrastive learning offers an appealing solution for multimodal self-supervised learning. Indeed, by considering each modality as a different view of the same entity, it learns to align features of different modalities in a shared representation space. However, this approach is intrinsically limited as it only learns shared or redundant information between modalities, while multimodal interactions can arise in other ways. In this work, we introduce CoMM, a Contrastive Multimodal learning strategy that enables the communication between modalities in a single multimodal space. Instead of imposing cross- or intra- modality constraints, we propose to align multimodal representations by maximizing the mutual information between augmented versions of these multimodal features. Our theoretical analysis shows that shared, synergistic and unique terms of information naturally emerge from this formulation, allowing us to estimate multimodal interactions beyond redundancy. We test CoMM both in a controlled and in a series of real-world settings: in the former, we demonstrate that CoMM effectively captures redundant, unique and synergistic information between modalities. In the latter, CoMM learns complex multimodal interactions and achieves state-of-the-art results on seven multimodal tasks. Benoit Dufumier, Javiera Castillo-Navarro, Devis Tuia, Jean-Philippe Thiran |
ICLR | 4 |
| 2025 | A Simple Framework for Open-Vocabulary Zero-Shot SegmentationabstractZero-shot classification capabilities naturally arise in models trained within a vision-language contrastive framework. Despite their classification prowess, these models struggle in dense tasks like zero-shot open-vocabulary segmentation. This deficiency is often attributed to the absence of localization cues in captions and the intertwined nature of the learning process, which encompasses both image/text representation learning and cross-modality alignment. To tackle these issues, we propose SimZSS, a $\textbf{Sim}$ple framework for open-vocabulary $\textbf{Z}$ero-$\textbf{S}$hot $\textbf{S}$egmentation. The method is founded on two key principles: i) leveraging frozen vision-only models that exhibit spatial awareness while exclusively aligning the text encoder and ii) exploiting the discrete nature of text and linguistic knowledge to pinpoint local concepts within captions. By capitalizing on the quality of the visual representations, our method requires only image-caption pair datasets and adapts to both small curated and large-scale noisy datasets. When trained on COCO Captions across 8 GPUs, SimZSS achieves state-of-the-art results on 7 out of 8 benchmark datasets in less than 15 minutes. Our code and pretrained models are publicly available at https://github.com/tileb1/simzss. Thomas Stegmüller, Tim Lebailly, Nikola Dukic, Behzad Bozorgtabar, Tinne Tuytelaars, Jean-Philippe Thiran |
ICLR | 6 |
| 2025 | Uncertainty modeling for fine-tuned implicit functionsabstractImplicit functions such as Neural Radiance Fields (NeRFs), occupancy networks, and signed distance functions (SDFs) have become pivotal in computer vision for reconstructing detailed object shapes from sparse views. Achieving optimal performance with these models can be challenging due to the extreme sparsity of inputs and distribution shifts induced by data corruptions. To this end, large, noise-free synthetic datasets can serve as shape priors to help models fill in gaps, but the resulting reconstructions must be approached with caution. Uncertainty estimation is crucial for assessing the quality of these reconstructions, particularly in identifying areas where the model is uncertain about the parts it has inferred from the prior. In this paper, we introduce Dropsembles, a novel method for uncertainty estimation in tuned implicit functions. We demonstrate the efficacy of our approach through a series of experiments, starting with toy examples and progressing to a real-world scenario. Specifically, we train a Convolutional Occupancy Network on synthetic anatomical data and test it on low-resolution MRI segmentations of the lumbar spine. Our results show that Dropsembles achieve the accuracy and calibration levels of deep ensembles but with significantly less computational cost. Anna Susmelj, Mael Macuglia, Natasa Tagasovska, Reto Sutter, Sebastiano Caprara, Jean-Philippe Thiran, Ender Konukoglu |
ICLR | 6 |
| 2025 | ReservoirTTA: Prolonged Test-time Adaptation for Evolving and Recurring DomainsabstractThis paper introduces **ReservoirTTA**, a novel plug–in framework designed for prolonged test–time adaptation (TTA) in scenarios where the test domain continuously shifts over time, including cases where domains recur or evolve gradually.
At its core, ReservoirTTA maintains a reservoir of domain-specialized models—an adaptive test-time model ensemble—that both detects new domains via online clustering over style features of incoming samples and routes each sample to the appropriate specialized model, and thereby enables domain-specific adaptation.
This multi-model strategy overcomes key limitations of single model adaptation, such as catastrophic forgetting, inter-domain interference, and error accumulation, ensuring robust and stable performance on sustained non-stationary test distributions.
Our theoretical analysis reveals key components that bound parameter variance and prevent model collapse, while our plug–in TTA module mitigates catastrophic forgetting of previously encountered domains. Extensive experiments on scene-level corruption benchmarks (ImageNet-C, CIFAR-10/100-C), object-level style shifts (DomainNet-126, PACS), and semantic segmentation (Cityscapes→ACDC) — covering recurring and continuously evolving domain shifts — show that ReservoirTTA substantially improves adaptation accuracy and maintains stable performance across prolonged, recurring shifts, outperforming state-of-the-art methods. Our code is publicly available at https://github.com/LTS5/ReservoirTTA. Guillaume Vray, Devavrat Tomar, Xufeng Gao, Jean-Philippe Thiran, Evan Shelhamer, Behzad Bozorgtabar |
NeurIPS | 4 |
| 2025 | A Path-Based Model for Aberration Correction in Ultrasound ImagingabstractPulse-Echo Ultrasound Imaging suffers from several sources of image degradation. In clinical conditions, superficial layers made of different tissues (e.g. skin, fat or muscles) create aberrations that can severely deteriorate image quality. To correct such aberrations, the majority of existing methods use either phase screens or speed of sound maps. However, a technique that is both accurate in real-world scenarios and compatible with near-real time imaging is lacking. Indeed phase screens are too simplistic to be physically accurate and speed of sound maps are computationally costly to estimate. We propose a new model of aberrations driven by the paths followed by ultrasound waves in the aberrating layer. With this new representation, we formulate an optimization problem in which a coherence factor is maximized with respect to a grid of aberrating paths. This problem is solved via a gradient ascent algorithm with variable splitting, in which all necessary gradients are expressed analytically. Using simulations of aberrating layers, we show that the proposed method can correct strong aberrations (i.e. of several periods) and outperforms a state-of-the-art technique based on speed of sound maps. Using in vivo experiments, we demonstrate that the proposed method is able to correct real aberrations in a few seconds which represents a major step forward towards a broader use of aberration correction methods. Baptiste Hériard-Dubreuil, Adrien Besson, Claude Cohen-Bacrie, Jean-Philippe Thiran |
IEEE Trans. Medical Imaging | 4 |
| 2024 | Combining Graph Transformers Based Multi-Label Active Learning and Informative Data Augmentation for Chest Xray ClassificationabstractInformative sample selection in active learning (AL) helps a machine learning system attain optimum performance with minimum labeled samples, thus improving human-in-the-loop computer-aided diagnosis systems with limited labeled data. Data augmentation is highly effective for enlarging datasets with less labeled data. Combining informative sample selection and data augmentation should leverage their respective advantages and improve performance of AL systems. We propose a novel approach to combine informative sample selection and data augmentation for multi-label active learning. Conventional informative sample selection approaches have mostly focused on the single-label case which do not perform optimally in the multi-label setting. We improve upon state-of-the-art multi-label active learning techniques by representing disease labels as graph nodes, use graph attention transformers (GAT) to learn more effective inter-label relationships and identify most informative samples. We generate transformations of these informative samples which are also informative. Experiments on public chest xray datasets show improved results over state-of-the-art multi-label AL techniques in terms of classification performance, learning rates, and robustness. We also perform qualitative analysis to determine the realism of generated images. Dwarikanath Mahapatra, Behzad Bozorgtabar, ZongYuan Ge, Mauricio Reyes 0001, Jean-Philippe Thiran |
AAAI | 5 |
| 2024 | CrIBo: Self-Supervised Learning via Cross-Image Object-Level BootstrappingabstractLeveraging nearest neighbor retrieval for self-supervised representation learning has proven beneficial with object-centric images. However, this approach faces limitations when applied to scene-centric datasets, where multiple objects within an image are only implicitly captured in the global representation. Such global bootstrapping can lead to undesirable entanglement of object representations. Furthermore, even object-centric datasets stand to benefit from a finer-grained bootstrapping approach. In response to these challenges, we introduce a novel $\textbf{Cr}$oss-$\textbf{I}$mage Object-Level $\textbf{Bo}$otstrapping method tailored to enhance dense visual representation learning. By employing object-level nearest neighbor bootstrapping throughout the training, CrIBo emerges as a notably strong and adequate candidate for in-context learning, leveraging nearest neighbor retrieval at test time. CrIBo shows state-of-the-art performance on the latter task while being highly competitive in more standard downstream segmentation tasks. Our code and pretrained models are publicly available at https://github.com/tileb1/CrIBo. Tim Lebailly, Thomas Stegmüller, Behzad Bozorgtabar, Jean-Philippe Thiran, Tinne Tuytelaars |
ICLR | 4 |
| 2024 | Un-Mixing Test-Time Normalization Statistics: Combatting Label Temporal CorrelationabstractRecent test-time adaptation methods heavily rely on nuanced adjustments of batch normalization (BN) parameters. However, one critical assumption often goes overlooked: that of independently and identically distributed (i.i.d.) test batches with respect to unknown labels. This oversight leads to skewed BN statistics and undermines the reliability of the model under non-i.i.d. scenarios. To tackle this challenge, this paper presents a novel method termed '$\textbf{Un-Mix}$ing $\textbf{T}$est-Time $\textbf{N}$ormalization $\textbf{S}$tatistics' (UnMix-TNS). Our method re-calibrates the statistics for each instance within a test batch by $\textit{mixing}$ it with multiple distinct statistics components, thus inherently simulating the i.i.d. scenario. The core of this method hinges on a distinctive online $\textit{unmixing}$ procedure that continuously updates these statistics components by incorporating the most similar instances from new test batches. Remarkably generic in its design, UnMix-TNS seamlessly integrates with a wide range of leading test-time adaptation methods and pre-trained architectures equipped with BN layers. Empirical evaluations corroborate the robustness of UnMix-TNS under varied scenarios—ranging from single to continual and mixed domain shifts, particularly excelling with temporally correlated test data and corrupted non-i.i.d. real-world streams. This adaptability is maintained even with very small batch sizes or single instances. Our results highlight UnMix-TNS's capacity to markedly enhance stability and performance across various benchmarks. Our code is publicly available at https://github.com/devavratTomar/unmixtns. Devavrat Tomar, Guillaume Vray, Jean-Philippe Thiran, Behzad Bozorgtabar |
ICLR | 3 |
| 2024 | Windowed Radon Transform for Robust Speed-of-Sound Imaging With Pulse-Echo UltrasoundabstractIn recent years, methods estimating the spatial distribution of tissue speed of sound with pulse-echo ultrasound are gaining considerable traction. They can address limitations of B-mode imaging, for instance in diagnosing fatty liver diseases. Current state-of-the-art methods relate the tissue speed of sound to local echo shifts computed between images that are beamformed using restricted transmit and receive apertures. However, the aperture limitation affects the robustness of phase-shift estimations and, consequently, the accuracy of reconstructed speed-of-sound maps. Here, we propose a method based on the Radon transform of image patches able to estimate local phase shifts from full-aperture images. We validate our technique on simulated, phantom and in-vivo data acquired on a liver and compare it with a state-of-the-art method. We show that the proposed method enhances the stability to changes of beamforming speed of sound and to a reduction of the number of insonifications. In particular, the deployment of pulse-echo speed-of-sound estimation methods onto portable ultrasound devices can be eased by the reduction of the number of insonifications allowed by the proposed method. Samuel Beuret, Baptiste Hériard-Dubreuil, Naiara Korta Martiartu, Michael Jaeger, Jean-Philippe Thiran |
IEEE Trans. Medical Imaging | 5 |
| 2024 | Windowed Radon Transform and Tensor Rank-1 Decomposition for Adaptive Beamforming in Ultrafast UltrasoundabstractUltrafast ultrasound has recently emerged as an alternative to traditional focused ultrasound. By virtue of the low number of insonifications it requires, ultrafast ultrasound enables the imaging of the human body at potentially very high frame rates. However, unaccounted for speed-of-sound variations in the insonified medium often result in phase aberrations in the reconstructed images. The diagnosis capability of ultrafast ultrasound is thus ultimately impeded. Therefore, there is a strong need for adaptive beamforming methods that are resilient to speed-of-sound aberrations. Several of such techniques have been proposed recently but they often lack parallelizability or the ability to directly correct both transmit and receive phase aberrations. In this article, we introduce an adaptive beamforming method designed to address these shortcomings. To do so, we compute the windowed Radon transform of several complex radio-frequency images reconstructed using delay-and-sum. Then, we apply to the obtained local sinograms weighted tensor rank-1 decompositions and their results are eventually used to reconstruct a corrected image. We demonstrate using simulated and in-vitro data that our method is able to successfully recover aberration-free images and that it outperforms both coherent compounding and the recently introduced SVD beamformer. Finally, we validate the proposed beamforming technique on in-vivo data, resulting in a significant improvement of image quality compared to the two reference methods. Samuel Beuret, Jean-Philippe Thiran |
IEEE Trans. Medical Imaging | 2 |
| 2024 | Distill-SODA: Distilling Self-Supervised Vision Transformer for Source-Free Open-Set Domain Adaptation in Computational PathologyabstractDeveloping computational pathology models is essential for reducing manual tissue typing from whole slide images, transferring knowledge from the source domain to an unlabeled, shifted target domain, and identifying unseen categories. We propose a practical setting by addressing the above-mentioned challenges in one fell swoop, i.e., source-free open-set domain adaptation. Our methodology focuses on adapting a pre-trained source model to an unlabeled target dataset and encompasses both closed-set and open-set classes. Beyond addressing the semantic shift of unknown classes, our framework also deals with a covariate shift, which manifests as variations in color appearance between source and target tissue samples. Our method hinges on distilling knowledge from a self-supervised vision transformer (ViT), drawing guidance from either robustly pre-trained transformer models or histopathology datasets, including those from the target domain. In pursuit of this, we introduce a novel style-based adversarial data augmentation, serving as hard positives for self-training a ViT, resulting in highly contextualized embeddings. Following this, we cluster semantically akin target images, with the source model offering weak pseudo-labels, albeit with uncertain confidence. To enhance this process, we present the closed-set affinity score (CSAS), aiming to correct the confidence levels of these pseudo-labels and to calculate weighted class prototypes within the contextualized embedding space. Our approach establishes itself as state-of-the-art across three public histopathological datasets for colorectal cancer assessment. Notably, our self-training method seamlessly integrates with open-set detection methods, resulting in enhanced performance in both closed-set and open-set recognition tasks. Guillaume Vray, Devavrat Tomar, Behzad Bozorgtabar, Jean-Philippe Thiran |
IEEE Trans. Medical Imaging | 4 |
| 2023 | CrOC: Cross-View Online Clustering for Dense Visual Representation LearningabstractLearning dense visual representations without labels is an arduous task and more so from scene-centric data. We propose to tackle this challenging problem by proposing a Cross-view consistency objective with an Online Clustering mechanism (CrOC) to discover and segment the semantics of the views. In the absence of hand-crafted priors, the resulting method is more generalizable and does not require a cumbersome pre-processing step. More importantly, the clustering algorithm conjointly operates on the features of both views, thereby elegantly bypassing the issue of content not represented in both views and the ambiguous matching of objects from one crop to the other. We demonstrate excellent performance on linear and unsupervised segmentation transfer tasks on various datasets and similarly for video object segmentation. Our code and pre-trained models are publicly available at https://github.com/stegmuel/CrOC. Thomas Stegmüller, Tim Lebailly, Behzad Bozorgtabar, Tinne Tuytelaars, Jean-Philippe Thiran |
CVPR | 5 |
| 2023 | TeSLA: Test-Time Self-Learning With Automatic Adversarial AugmentationabstractMost recent test-time adaptation methods focus on only classification tasks, use specialized network architectures, destroy model calibration or rely on lightweight information from the source domain. To tackle these issues, this paper proposes a novel Test-time Self-Learning method with automatic Adversarial augmentation dubbed TeSLA for adapting a pre-trained source model to the unlabeled streaming test data. In contrast to conventional self-learning methods based on cross-entropy, we introduce a new test-time loss function through an implicitly tight connection with the mutual information and online knowledge distillation. Furthermore, we propose a learnable efficient adversarial augmentation module that further enhances online knowledge distillation by simulating high entropy augmented images. Our method achieves state-of-the-art classification and segmentation results on several benchmarks and types of domain shifts, particularly on challenging measurement shifts of medical images. TeSLA also benefits from several desirable properties compared to competing methods in terms of calibration, uncertainty metrics, insensitivity to model architectures, and source training strategies, all supported by extensive ablations. Our code and models are available at https://github.com/devavratTomar/TeSLA. Devavrat Tomar, Guillaume Vray, Behzad Bozorgtabar, Jean-Philippe Thiran |
CVPR | 4 |
| 2023 | Adaptive Similarity Bootstrapping for Self-Distillation based Representation LearningabstractMost self-supervised methods for representation learning leverage a cross-view consistency objective i.e. they maximize the representation similarity of a given image’s augmented views. Recent work NNCLR goes beyond the cross-view paradigm and uses positive pairs from different images obtained via nearest neighbor bootstrapping in a contrastive setting. We empirically show that as opposed to the contrastive learning setting which relies on negative samples, incorporating nearest neighbor bootstrapping in a self-distillation scheme can lead to a performance drop or even collapse. We scrutinize the reason for this unexpected behavior and provide a solution. We propose to adaptively bootstrap neighbors based on the estimated quality of the latent space. We report consistent improvements compared to the naive bootstrapping approach and the original baselines. Our approach leads to performance improvements for various self-distillation method/backbone combinations and standard downstream tasks. Our code is publicly available at https://github.com/tileb1/AdaSim. Tim Lebailly, Thomas Stegmüller, Behzad Bozorgtabar, Jean-Philippe Thiran, Tinne Tuytelaars |
ICCV | 4 |
| 2023 | AMAE: Adaptation of Pre-trained Masked Autoencoder for Dual-Distribution Anomaly Detection in Chest X-Rays
Behzad Bozorgtabar, Dwarikanath Mahapatra, Jean-Philippe Thiran |
MICCAI (1) | 3 |
| 2023 | ScoreNet: Learning Non-Uniform Attention and Augmentation for Transformer-Based Histopathological Image ClassificationabstractProgress in digital pathology is hindered by high-resolution images and the prohibitive cost of exhaustive localized annotations. The commonly used paradigm to categorize pathology images is patch-based processing, which often incorporates multiple instance learning (MIL) to aggregate local patch-level representations yielding image-level prediction. Nonetheless, diagnostically relevant regions may only take a small fraction of the whole tissue, and current MIL-based approaches often process images uniformly, discarding the inter-patches interactions. To alleviate these issues, we propose ScoreNet, a new efficient transformer that exploits a differentiable recommendation stage to extract discriminative image regions and dedicate computational resources accordingly. The proposed transformer leverages the local and global attention of a few dynamically recommended high-resolution regions at an efficient computational cost. We further introduce a novel mixing data-augmentation, namely ScoreMix, by leveraging the image’s semantic distribution to guide the data mixing and produce coherent sample-label pairs. ScoreMix is embarrassingly simple and mitigates the pitfalls of previous augmentations, which assume a uniform semantic distribution and risk mislabeling the samples. Thorough experiments and ablation studies on three breast cancer histology datasets of Haematoxylin & Eosin (H&E) have validated the superiority of our approach over prior arts, including transformer-based models on tumour regions-of-interest (TRoIs) classification. ScoreNet equipped with proposed ScoreMix augmentation demonstrates better generalization capabilities and achieves new state-of-the-art (SOTA) results with only 50% of the data compared to other mixing augmentation variants. Finally, ScoreNet yields high efficacy and outperforms SOTA efficient transformers, namely TransPath [37] and SwinTransformer [20], with throughput around 3× and 4× higher than the aforementioned architectures, respectively. Our code is publicly available1. Thomas Stegmüller, Behzad Bozorgtabar, Antoine Spahr, Jean-Philippe Thiran |
WACV | 4 |
| 2022 | Anomaly Detection and Localization Using Attention-Guided Synthetic Anomaly and Test-Time Adaptation
Behzad Bozorgtabar, Dwarikanath Mahapatra, Jean-Philippe Thiran |
BMVC | 3 |
| 2022 | Self-Supervised Generative Style Transfer for One-Shot Medical Image SegmentationabstractIn medical image segmentation, supervised deep networks’ success comes at the cost of requiring abundant labeled data. While asking domain experts to annotate only one or a few of the cohort’s images is feasible, annotating all available images is impractical. This issue is further exacerbated when pre-trained deep networks are exposed to a new image dataset from an unfamiliar distribution. Using available open-source data for ad-hoc transfer learning or hand-tuned techniques for data augmentation only provides suboptimal solutions. Motivated by atlas-based segmentation, we propose a novel volumetric self-supervised learning for data augmentation capable of synthesizing volumetric image-segmentation pairs via learning transformations from a single labeled atlas to the unlabeled data. Our work’s central tenet benefits from a combined view of one-shot generative learning and the proposed self-supervised training strategy that cluster unlabeled volumetric images with similar styles together. Unlike previous methods, our method does not require input volumes at inference time to synthesize new images. Instead, it can generate diversified volumetric image-segmentation pairs from a prior distribution given a single or multi-site dataset. Augmented data generated by our method used to train the segmentation network provide significant improvements over state-of-the-art deep one-shot learning methods on the task of brain MRI segmentation. Ablation studies further exemplified that the proposed appearance model and joint training are crucial to synthesize realistic examples compared to existing medical registration methods. The code, data, and models are available at https://github.com/devavratTomar/SST/. Devavrat Tomar, Behzad Bozorgtabar, Manana Lortkipanidze, Guillaume Vray, Mohammad Saeed Rad, Jean-Philippe Thiran |
WACV | 6 |
| 2022 | Self-rule to multi-adapt: Generalized multi-source feature learning using unsupervised domain adaptation for colorectal cancer tissue detectionabstractSupervised learning is constrained by the availability of labeled data, which are especially expensive to acquire in the field of digital pathology. Making use of open-source data for pre-training or using domain adaptation can be a way to overcome this issue. However, pre-trained networks often fail to generalize to new test domains that are not distributed identically due to tissue stainings, types, and textures variations. Additionally, current domain adaptation methods mainly rely on fully-labeled source datasets. In this work, we propose Self-Rule to Multi-Adapt (SRMA), which takes advantage of self-supervised learning to perform domain adaptation, and removes the necessity of fully-labeled source datasets. SRMA can effectively transfer the discriminative knowledge obtained from a few labeled source domain's data to a new target domain without requiring additional tissue annotations. Our method harnesses both domains' structures by capturing visual similarity with intra-domain and cross-domain self-supervision. Moreover, we present a generalized formulation of our approach that allows the framework to learn from multiple source domains. We show that our proposed method outperforms baselines for domain adaptation of colorectal tissue type classification in single and multi-source settings, and further validate our approach on an in-house clinical cohort. The code and trained models are available open-source: https://github.com/christianabbet/SRA. Christian Abbet, Linda Studer, Andreas Fischer 0002, Heather Dawson, Inti Zlobec, Behzad Bozorgtabar, Jean-Philippe Thiran |
Medical Image Anal. | 7 |
| 2022 | Hierarchical graph representations in digital pathologyabstractCancer diagnosis, prognosis, and therapy response predictions from tissue specimens highly depend on the phenotype and topological distribution of constituting histological entities. Thus, adequate tissue representations for encoding histological entities is imperative for computer aided cancer patient care. To this end, several approaches have leveraged cell-graphs, capturing the cell-microenvironment, to depict the tissue. These allow for utilizing graph theory and machine learning to map the tissue representation to tissue functionality, and quantify their relationship. Though cellular information is crucial, it is incomplete alone to comprehensively characterize complex tissue structure. We herein treat the tissue as a hierarchical composition of multiple types of histological entities from fine to coarse level, capturing multivariate tissue information at multiple levels. We propose a novel multi-level hierarchical entity-graph representation of tissue specimens to model the hierarchical compositions that encode histological entities as well as their intra- and inter-entity level interactions. Subsequently, a hierarchical graph neural network is proposed to operate on the hierarchical entity-graph and map the tissue structure to tissue functionality. Specifically, for input histology images, we utilize well-defined cells and tissue regions to build HierArchical Cell-to-Tissue (HACT) graph representations, and devise HACT-Net, a message passing graph neural network, to classify the HACT representations. As part of this work, we introduce the BReAst Carcinoma Subtyping (BRACS) dataset, a large cohort of Haematoxylin & Eosin stained breast tumor regions-of-interest, to evaluate and benchmark our proposed methodology against pathologists and state-of-the-art computer-aided diagnostic approaches. Through comparative assessment and ablation studies, our proposed method is demonstrated to yield superior classification results compared to alternative methods as well as individual pathologists. The code, data, and models can be accessed at https://github.com/histocartography/hact-net. Pushpak Pati, Guillaume Jaume, Antonio Foncubierta-Rodríguez, Florinda Feroce, Anna Maria Anniciello, Giosue Scognamiglio, Nadia Brancati, Maryse Fiche, Estelle Dubruc, Daniel Riccio, Maurizio Di Bonito, Giuseppe De Pietro, Gerardo Botti, Jean-Philippe Thiran, Maria Frucci, Orcun Goksel, Maria Gabrani |
Medical Image Anal. | 14 |
| 2021 | Quantifying Explainers of Graph Neural Networks in Computational PathologyabstractExplainability of deep learning methods is imperative to facilitate their clinical adoption in digital pathology. However, popular deep learning methods and explainability techniques (explainers) based on pixel-wise processing disregard biological entities’ notion, thus complicating comprehension by pathologists. In this work, we address this by adopting biological entity-based graph processing and graph explainers enabling explanations accessible to pathologists. In this context, a major challenge becomes to discern meaningful explainers, particularly in a standardized and quantifiable fashion. To this end, we propose herein a set of novel quantitative metrics based on statistics of class separability using pathologically measurable concepts to characterize graph explainers. We employ the proposed metrics to evaluate three types of graph explainers, namely the layer-wise relevance propagation, gradient-based saliency, and graph pruning approaches, to explain Cell-Graph representations for Breast Cancer Subtyping. The proposed metrics are also applicable in other domains by using domain-specific intuitive concepts. We validate the qualitative and quantitative findings on the BRACS dataset, a large cohort of breast cancer RoIs, by expert pathologists. The code, data, and models can be accessed here1. Guillaume Jaume, Pushpak Pati, Behzad Bozorgtabar, Antonio Foncubierta-Rodríguez, Anna Maria Anniciello, Florinda Feroce, Tilman Rau, Jean-Philippe Thiran, Maria Gabrani, Orcun Goksel |
CVPR | 8 |
| 2021 | Learning Whole-Slide Segmentation from Inexact and Incomplete Labels Using Tissue Graphs
Valentin Anklin, Pushpak Pati, Guillaume Jaume, Behzad Bozorgtabar, Antonio Foncubierta-Rodríguez, Jean-Philippe Thiran, Mathilde Sibony, Maria Gabrani, Orcun Goksel |
MICCAI (2) | 6 |
| 2021 | Benefiting from Bicubically Down-Sampled Images for Learning Real-World Image Super-ResolutionabstractSuper-resolution (SR) has traditionally been based on pairs of high-resolution images (HR) and their low-resolution (LR) counterparts obtained artificially with bicubic downsampling. However, in real-world SR, there is a large variety of realistic image degradations and analytically modeling these realistic degradations can prove quite difficult. In this work, we propose to handle real-world SR by splitting this ill-posed problem into two comparatively more well-posed steps. First, we train a network to transform real LR images to the space of bicubically down-sampled images in a supervised manner, by using both real LR/HR pairs and synthetic pairs. Second, we take a generic SR network trained on bicubically downsampled images to super-resolve the transformed LR image. The first step of the pipeline addresses the problem by registering the large variety of degraded images to a common, well understood space of images. The second step then leverages the already impressive performance of SR on bicubically downsampled images, sidestepping the issues of end-to-end training on datasets with many different image degradations. We demonstrate the effectiveness of our proposed method by comparing it to recent methods in real-world SR and show that our proposed approach outperforms the state-of-the-art works in terms of both qualitative and quantitative results, as well as results of an extensive user study conducted on several real image datasets. Mohammad Saeed Rad, Thomas Yu, Claudiu Cristian Musat, Hazim Kemal Ekenel, Behzad Bozorgtabar, Jean-Philippe Thiran |
WACV | 6 |
| 2021 | Comparison of non-parametric T2 relaxometry methods for myelin water quantificationabstractMulti-component T2 relaxometry allows probing tissue microstructure by assessing compartment-specific T2 relaxation times and water fractions, including the myelin water fraction. Non-negative least squares (NNLS) with zero-order Tikhonov regularization is the conventional method for estimating smooth T2 distributions. Despite the improved estimation provided by this method compared to non-regularized NNLS, the solution is still sensitive to the underlying noise and the regularization weight. This is especially relevant for clinically achievable signal-to-noise ratios. In the literature of inverse problems, various well-established approaches to promote smooth solutions, including first-order and second-order Tikhonov regularization, and different criteria for estimating the regularization weight have been proposed, such as L-curve, Generalized Cross-Validation, and Chi-square residual fitting. However, quantitative comparisons between the available reconstruction methods for computing the T2 distribution, and between different approaches for selecting the optimal regularization weight, are lacking. In this study, we implemented and evaluated ten reconstruction algorithms, resulting from the individual combinations of three penalty terms with three criteria to estimate the regularization weight, plus non-regularized NNLS. Their performance was evaluated both in simulated data and real brain MRI data acquired from healthy volunteers through a scan-rescan repeatability analysis. Our findings demonstrate the need for regularization. As a result of this work, we provide a list of recommendations for selecting the optimal reconstruction algorithms based on the acquired data. Moreover, the implemented methods were packaged in a freely distributed toolbox to promote reproducible research, and to facilitate further research and the use of this promising quantitative technique in clinical practice. Erick Jorge Canales-Rodríguez, Marco Pizzolato, Gian Franco Piredda, Tom Hilbert, Nicolas Kunz, Caroline Pot, Thomas Yu, Raymond Salvador, Edith Pomarol-Clotet, Tobias Kober, Jean-Philippe Thiran, Alessandro Daducci |
Medical Image Anal. | 11 |
| 2021 | Model-informed machine learning for multi-component T2 relaxometryabstractRecovering the T2 distribution from multi-echo T2 magnetic resonance (MR) signals is challenging but has high potential as it provides biomarkers characterizing the tissue micro-structure, such as the myelin water fraction (MWF). In this work, we propose to combine machine learning and aspects of parametric (fitting from the MRI signal using biophysical models) and non-parametric (model-free fitting of the T2 distribution from the signal) approaches to T2 relaxometry in brain tissue by using a multi-layer perceptron (MLP) for the distribution reconstruction. For training our network, we construct an extensive synthetic dataset derived from biophysical models in order to constrain the outputs with a priori knowledge of in vivo distributions. The proposed approach, called Model-Informed Machine Learning (MIML), takes as input the MR signal and directly outputs the associated T2 distribution. We evaluate MIML in comparison to a Gaussian Mixture Fitting (parametric) and Regularized Non-Negative Least Squares algorithms (non-parametric) on synthetic data, an ex vivo scan, and high-resolution scans of healthy subjects and a subject with Multiple Sclerosis. In synthetic data, MIML provides more accurate and noise-robust distributions. In real data, MWF maps derived from MIML exhibit the greatest conformity to anatomical scans, have the highest correlation to a histological map of myelin volume, and the best unambiguous lesion visualization and localization, with superior contrast between lesions and normal appearing tissue. In whole-brain analysis, MIML is 22 to 4980 times faster than the non-parametric and parametric methods, respectively. Thomas Yu, Erick Jorge Canales-Rodríguez, Marco Pizzolato, Gian Franco Piredda, Tom Hilbert, Elda Fischi Gomez, Matthias Weigel, Muhamed Barakovic, Meritxell Bach Cuadra, Cristina Granziera, Tobias Kober, Jean-Philippe Thiran |
Medical Image Anal. | 12 |
| 2021 | CNN-Based Ultrasound Image Reconstruction for Ultrafast Displacement TrackingabstractThanks to its capability of acquiring full-view frames at multiple kilohertz, ultrafast ultrasound imaging unlocked the analysis of rapidly changing physical phenomena in the human body, with pioneering applications such as ultrasensitive flow imaging in the cardiovascular system or shear-wave elastography. The accuracy achievable with these motion estimation techniques is strongly contingent upon two contradictory requirements: a high quality of consecutive frames and a high frame rate. Indeed, the image quality can usually be improved by increasing the number of steered ultrafast acquisitions, but at the expense of a reduced frame rate and possible motion artifacts. To achieve accurate motion estimation at uncompromised frame rates and immune to motion artifacts, the proposed approach relies on single ultrafast acquisitions to reconstruct high-quality frames and on only two consecutive frames to obtain 2-D displacement estimates. To this end, we deployed a convolutional neural network-based image reconstruction method combined with a speckle tracking algorithm based on cross-correlation. Numerical and in vivo experiments, conducted in the context of plane-wave imaging, demonstrate that the proposed approach is capable of estimating displacements in regions where the presence of side lobe and grating lobe artifacts prevents any displacement estimation with a state-of-the-art technique that relies on conventional delay-and-sum beamforming. The proposed approach may therefore unlock the full potential of ultrafast ultrasound, in applications such as ultrasensitive cardiovascular motion and flow analysis or shear-wave elastography. Dimitris Perdios, Manuel Vonlanthen, Florian Martinez, Marcel Arditi, Jean-Philippe Thiran |
IEEE Trans. Medical Imaging | 5 |
| 2021 | Self-Attentive Spatial Adaptive Normalization for Cross-Modality Domain AdaptationabstractDespite the successes of deep neural networks on many challenging vision tasks, they often fail to generalize to new test domains that are not distributed identically to the training data. The domain adaptation becomes more challenging for cross-modality medical data with a notable domain shift. Given that specific annotated imaging modalities may not be accessible nor complete. Our proposed solution is based on the cross-modality synthesis of medical images to reduce the costly annotation burden by radiologists and bridge the domain gap in radiological images. We present a novel approach for image-to-image translation in medical images, capable of supervised or unsupervised (unpaired image data) setups. Built upon adversarial training, we propose a learnable self-attentive spatial normalization of the deep convolutional generator network's intermediate activations. Unlike previous attention-based image-to-image translation approaches, which are either domain-specific or require distortion of the source domain's structures, we unearth the importance of the auxiliary semantic information to handle the geometric changes and preserve anatomical structures during image translation. We achieve superior results for cross-modality segmentation between unpaired MRI and CT data for multi-modality whole heart and multi-modal brain tumor MRI (T1/T2) datasets compared to the state-of-the-art methods. We also observe encouraging results in cross-modality conversion for paired MRI and CT images on a brain dataset. Furthermore, a detailed analysis of the cross-modality image translation, thorough ablation studies confirm our proposed method's efficacy. Devavrat Tomar, Manana Lortkipanidze, Guillaume Vray, Behzad Bozorgtabar, Jean-Philippe Thiran |
IEEE Trans. Medical Imaging | 5 |
| 2020 | Divide-and-Rule: Self-Supervised Learning for Survival Analysis in Colorectal Cancer
Christian Abbet, Inti Zlobec, Behzad Bozorgtabar, Jean-Philippe Thiran |
MICCAI (5) | 4 |
| 2020 | SALAD: Self-supervised Aggregation Learning for Anomaly Detection on X-Rays
Behzad Bozorgtabar, Dwarikanath Mahapatra, Guillaume Vray, Jean-Philippe Thiran |
MICCAI (1) | 4 |
| 2020 | T2 Mapping from Super-Resolution-Reconstructed Clinical Fast Spin Echo Magnetic Resonance Acquisitions
Hélène Lajous, Tom Hilbert, Christopher W. Roy, Sébastien Tourbier, Priscille de Dumast, Thomas Yu, Jean-Philippe Thiran, Jean-Baptiste Ledoux, Davide Piccini, Patric Hagmann, Reto Meuli, Tobias Kober, Matthias Stuber, Ruud B. van Heeswijk, Meritxell Bach Cuadra |
MICCAI (2) | 7 |
| 2020 | Structure Preserving Stain Normalization of Histopathology Images Using Self Supervised Semantic Guidance
Dwarikanath Mahapatra, Behzad Bozorgtabar, Jean-Philippe Thiran, Ling Shao 0001 |
MICCAI (5) | 3 |
| 2020 | Automated Detection of Cortical Lesions in Multiple Sclerosis Patients with 7T MRI
Francesco La Rosa, Erin S. Beck, Ahmed Abdulkadir, Jean-Philippe Thiran, Daniel S. Reich, Pascal Sati, Meritxell Bach Cuadra |
MICCAI (4) | 4 |
| 2020 | An Evolutionary Framework for Microstructure-Sensitive Generalized Diffusion Gradient Waveforms
Raphaël Truffet, Jonathan Rafael-Patino, Gabriel Girard, Marco Pizzolato, Christian Barillot, Jean-Philippe Thiran, Emmanuel Caruyer |
MICCAI (2) | 6 |
| 2020 | Benefiting from multitask learning to improve single image super-resolution
Mohammad Saeed Rad, Behzad Bozorgtabar, Claudiu Cristian Musat, Urs-Viktor Marti, Max Basler, Hazim Kemal Ekenel, Jean-Philippe Thiran |
Neurocomputing | 7 |
| 2020 | ExprADA: Adversarial domain adaptation for facial expression analysis
Behzad Bozorgtabar, Dwarikanath Mahapatra, Jean-Philippe Thiran |
Pattern Recognit. | 3 |
| 2019 | Real-Time Textureless-Region Tolerant High-Resolution Depth Estimation SystemabstractThis study presents a real-time depth estimation hardware system aiming to provide high-resolution depth data and to eliminate its noise on textureless regions without causing any interference problem by utilizing artificial pattern projection. The system generates up to 2K resolution depth data and reaches up to 256 disparity range which are configurable by the end-user owing to its parameterized design. It is capable of streaming depth data with 21 frames per second (fps) with 2K resolution and 128 pixel disparity range, and its throughput performance changes depending on configuration of the output resolution and the disparity range. Bilal Demir, Jean-Philippe Thiran, Yusuf Leblebici |
DSD | 2 |
| 2019 | G2-VER: Geometry Guided Model Ensemble for Video-based Facial Expression RecognitionabstractThis paper addresses the problem of automatic facial expression recognition in videos, where the goal is to predict discrete emotion labels best describing the emotions expressed in short video clips. Building on a pre-trained convolutional neural network (CNN) model dedicated to analyzing the video frames and LSTM network designed to process the trajectories of the facial landmarks, this paper investigates several novel directions. First of all, improved face descriptors based on 2D CNNs and facial landmarks are proposed. Second, the paper investigates fusion methods of the features temporally, including a novel hierarchical recurrent neural network combining facial landmark trajectories over time. In addition, we propose a modification to state-of-the-art expression recognition architectures to adapt them to video processing in a simple way. In both ensemble approaches, the temporal information is integrated. Comparative experiments on publicly available video-based facial expression recognition datasets verified that the proposed framework outperforms state-of-the-art methods. Moreover, we introduce a near-infrared video dataset containing facial expressions from subjects driving their cars, which are recorded in real world conditions. Tanguy Albrici, Mandana Fasounaki, Saleh Bagher Salimi, Guillaume Vray, Behzad Bozorgtabar, Hazim Kemal Ekenel, Jean-Philippe Thiran |
FG | 7 |
| 2019 | Using Photorealistic Face Synthesis and Domain Adaptation to Improve Facial Expression AnalysisabstractCross-domain synthesizing realistic faces to learn deep models has attracted increasing attention for facial expression analysis as it helps to improve the performance of expression recognition accuracy despite having small number of real training images. However, learning from synthetic face images can be problematic due to the distribution discrepancy between low-quality synthetic images and real face images and may not achieve the desired performance when the learned model applies to real world scenarios. To this end, we propose a new attribute guided face image synthesis to perform a translation between multiple image domains using a single model. In addition, we adopt the proposed model to learn from synthetic faces by matching the feature distributions between different domains while preserving each domain's characteristics. We evaluate the effectiveness of the proposed approach on several face datasets on generating realistic face images. We demonstrate that the expression recognition performance can be enhanced by benefiting from our face synthesis model. Moreover, we also conduct experiments on a near-infrared dataset containing facial expression videos of drivers to assess the performance using in-the-wild data for driver emotion recognition. Behzad Bozorgtabar, Mohammad Saeed Rad, Hazim Kemal Ekenel, Jean-Philippe Thiran |
FG | 4 |
| 2019 | SynDeMo: Synergistic Deep Feature Alignment for Joint Learning of Depth and Ego-MotionabstractDespite well-established baselines, learning of scene depth and ego-motion from monocular video remains an ongoing challenge, specifically when handling scaling ambiguity issues and depth inconsistencies in image sequences. Much prior work uses either a supervised mode of learning or stereo images. The former is limited by the amount of labeled data, as it requires expensive sensors, while the latter is not always readily available as monocular sequences. In this work, we demonstrate the benefit of using geometric information from synthetic images, coupled with scene depth information, to recover the scale in depth and ego-motion estimation from monocular videos. We developed our framework using synthetic image-depth pairs and unlabeled real monocular images. We had three training objectives: first, to use deep feature alignment to reduce the domain gap between synthetic and monocular images to yield more accurate depth estimation when presented with only real monocular images at test time. Second, we learn scene specific representation by exploiting self-supervision coming from multi-view synthetic images without the need for depth labels. Third, our method uses single-view depth and pose networks, which are capable of jointly training and supervising one another mutually, yielding consistent depth and ego-motion estimates. Extensive experiments demonstrate that our depth and ego-motion models surpass the state-of-the-art, unsupervised methods and compare favorably to early supervised deep models for geometric understanding. We validate the effectiveness of our training objectives against standard benchmarks thorough an ablation study. Behzad Bozorgtabar, Mohammad Saeed Rad, Dwarikanath Mahapatra, Jean-Philippe Thiran |
ICCV | 4 |
| 2019 | SROBB: Targeted Perceptual Loss for Single Image Super-ResolutionabstractBy benefiting from perceptual losses, recent studies have improved significantly the performance of the superresolution task, where a high-resolution image is resolved from its low-resolution counterpart. Although such objective functions generate near-photorealistic results, their capability is limited, since they estimate the reconstruction error for an entire image in the same way, without considering any semantic information. In this paper, we propose a novel method to benefit from perceptual loss in a more objective way. We optimize a deep network-based decoder with a targeted objective function that penalizes images at different semantic levels using the corresponding terms. In particular, the proposed method leverages our proposed OBB (Object, Background and Boundary) labels, generated from segmentation labels, to estimate a suitable perceptual loss for boundaries, while considering texture similarity for backgrounds. We show that our proposed approach results in more realistic textures and sharper edges, and outperforms other state-of-the-art algorithms in terms of both qualitative results on standard benchmarks and results of extensive user studies. Mohammad Saeed Rad, Behzad Bozorgtabar, Urs-Viktor Marti, Max Basler, Hazim Kemal Ekenel, Jean-Philippe Thiran |
ICCV | 6 |
| 2019 | Learning Global Brain Microstructure Maps Using Trainable Sparse EncodersabstractDiffusion-Weighted Magnetic Resonance Imaging is the only non-invasive technique available to infer the underlying brain tissue microstructure. Currently, one of the promising methods for microstructure imaging is signal modelling using convex formulation, e.g. using the COMMIT framework. Despite the benefits introduced with such framework, an important limitation is the long convergence time, making the method unappealing for clinical applications. In order to address this limitation, we propose to use a neural network to learn the sparse representation of the data and perform an end-to-end reconstruction of the microstructure estimates directly from the diffusion-weighted data. Our results show that the neural network can accurately estimate the microstructure maps, 4 orders of magnitude faster than the convex formulation. Jonathan Rafael-Patino, Muhamed Barakovic, Gabriel Girard, Alessandro Daducci, Jean-Philippe Thiran |
ICIP | 5 |
| 2019 | Informative sample generation using class aware generative adversarial networks for classification of chest Xrays
Behzad Bozorgtabar, Dwarikanath Mahapatra, Hendrik von Tengg-Kobligk, Alexander Pollinger, Lukas Ebner, Jean-Philippe Thiran, Mauricio Reyes 0001 |
Comput. Vis. Image Underst. | 6 |
| 2019 | Learn to synthesize and synthesize to learn
Behzad Bozorgtabar, Mohammad Saeed Rad, Hazim Kemal Ekenel, Jean-Philippe Thiran |
Comput. Vis. Image Underst. | 4 |
| 2019 | Joint Sparsity With Partially Known Support and Application to Ultrasound ImagingabstractWe investigate the benefits of known partial support for the recovery of joint-sparse signals and demonstrate that it is advantageous in terms of recovery performance for both rank-blind and rank-aware algorithms. We suggest extensions of several joint-sparse recovery algorithms, e.g., simultaneous normalized iterative hard thresholding, subspace greedy methods and subspace-augmented multiple signal classification techniques. We describe a direct application of the proposed methods for compressive multiplexing of ultrasound (US) signals. The technique exploits the compressive multiplexer architecture for signal compression and relies on joint-sparsity of US signals in the frequency domain for signal reconstruction. We validate the proposed algorithms on numerical experiments and show their superiority against state-of-the-art approaches in rank-defective cases. We also demonstrate that the techniques lead to a significant increase of the image quality on in vivo carotid images compared to reconstruction without partially known support. The supporting code is available on https://github.com/AdriBesson/spl2018_joint_sparse. Adrien Besson, Dimitris Perdios, Yves Wiaux, Jean-Philippe Thiran |
IEEE Signal Process. Lett. | 4 |
| 2018 | Pulse-Stream Models in Time-of-Flight ImagingabstractThis paper considers the problem of reconstructing raw signals from random projections in the context of time-of-flight imaging with an array of sensors. It presents a new signal model, coined as multi-channel pulse-stream model, which exploits pulse-stream models and accounts for additional structure induced by inter-sensor dependencies. We propose a sampling theorem and a reconstruction algorithm, based on ℓ -minimization, for signals belonging to such a model. We demonstrate the benefits of the proposed approach by means of numerical simulations and on a real non-destructive-evaluation application where the peak-signal-to-noise-ratio is increased by 3 dB compared to standard compressed-sensing strategies. Adrien Besson, Dimitris Perdios, Yves Wiaux, Jean-Philippe Thiran |
ICASSP | 4 |
| 2018 | Efficient Active Learning for Image Classification and Segmentation Using a Sample Selection and Conditional Generative Adversarial Network
Dwarikanath Mahapatra, Behzad Bozorgtabar, Jean-Philippe Thiran, Mauricio Reyes 0001 |
MICCAI (2) | 3 |
| 2017 | Combining LiDAR space clustering and convolutional neural networks for pedestrian detectionabstractPedestrian detection is an important component for safety of autonomous vehicles, as well as for traffic and street surveillance. There are extensive benchmarks on this topic and it has been shown to be a challenging problem when applied on real use-case scenarios. In purely image-based pedestrian detection approaches, the state-of-the-art results have been achieved with convolutional neural networks (CNN) and surprisingly few detection frameworks have been built upon multi-cue approaches. In this work, we develop a new pedestrian detector for autonomous vehicles that exploits LiDAR data, in addition to visual information. In the proposed approach, LiDAR data is utilized to generate region proposals by processing the three dimensional point cloud that it provides. These candidate regions are then further processed by a state-of-the-art CNN classifier that we have fine-tuned for pedestrian detection. We have extensively evaluated the proposed detection process on the KITTI dataset. The experimental results show that the proposed LiDAR space clustering approach provides a very efficient way of generating region proposals leading to higher recall rates and fewer misses for pedestrian detection. This indicates that LiDAR data can provide auxiliary information for CNN-based approaches. Damien Matti, Hazim Kemal Ekenel, Jean-Philippe Thiran |
AVSS | 3 |
| 2017 | 1024-Channel 3D ultrasound digital beamformer in a single 5W FPGAabstract3D ultrasound, an emerging medical imaging technique that is presently only used in hospitals, has the potential to enable breakthrough telemedicine applications, provided that its cost and power dissipation can be minimized. In this paper, we present an FPGA architecture suitable for a portable medical 3D ultrasound device. We show an optimized design for the digital part of the imager, including the delay calculation block, which is its most critical part. Our computationally efficient approach requires a single FPGA for 3D imaging, which is unprecedented. The design is scalable; a configuration supporting a 32×32-channel probe, which enables high-quality imaging, consumes only about 5W. Federico Angiolini, Aya Ibrahim, William Andrew Simon, Ahmet Caner Yuzuguler, Marcel Arditi, Jean-Philippe Thiran, Giovanni De Micheli |
DATE | 6 |
| 2017 | Learning the weight matrix for sparsity averaging in compressive imagingabstractWe propose to map the fast iterative soft thresholding algorithm to a deep neural network (DNN), with a sparsity prior in a concatenation of wavelet bases, in the context of compressive imaging. We exploit the DNN architecture to learn the optimal weight matrix of the corresponding reweighted ℓ1-minimization problem. We later use the learned weight matrix for the image reconstruction process, which is recast as a simple ℓ1-minimization problem. The approach, denoted as learned extended FISTA, shows promising results in terms of image quality, compared to state-of-the art algorithms, and significantly reduces the reconstruction time required to solve the reweighted ℓ1-minimization problem. Dimitris Perdios, Adrien Besson, Philippe Rossinelli, Jean-Philippe Thiran |
ICIP | 4 |
| 2017 | A Computer Vision System to Localize and Classify Wastes on the Streets
Mohammad Saeed Rad, Andreas von Kaenel, Andre Droux, François Tièche, Nabil Ouerhani, Hazim Kemal Ekenel, Jean-Philippe Thiran |
ICVS | 7 |
| 2017 | Action Units and Their Cross-Correlations for Prediction of Cognitive Load during DrivingabstractDriving requires the constant coordination of many body systems and full attention of the person. Cognitive distraction (subsidiary mental load) of the driver is an important factor that decreases attention and responsiveness, which may result in human error and accidents. In this paper, we present a study of facial expressions of such mental diversion of attention. First, we introduce a multi-camera database of 46 people recorded while driving a simulator in two conditions, baseline and induced cognitive load using a secondary task. Then, we present an automatic system to differentiate between the two conditions, where we use features extracted from Facial Action Unit (AU) values and their cross-correlations in order to exploit recurring synchronization and causality patterns. Both the recording and detection system are suitable for integration in a vehicle and a real-world application, e.g., an early warning system. We show that when the system is trained individually on each subject we achieve a mean accuracy and F-score of$\sim 95$percent, and for the subject independent tests$\sim 68$percent accuracy and$\sim 66$percent F-score, with person-specific normalization to handle subject dependency. Based on the results, we discuss the universality of the facial expressions of such states and possible real-world uses of the system. Anil Yüce, Hua Gao, Gabriel Louis Cuendet, Jean-Philippe Thiran |
IEEE Trans. Affect. Comput. | 4 |
| 2017 | A Regression-Based User Calibration Framework for Real-Time Gaze EstimationabstractEye movements play a very significant role in human-computer interaction (HCI) as they are natural and fast, and contain important cues for human cognitive state and visual attention. Over the last two decades, many techniques have been proposed to accurately estimate the gaze. Among these, video-based remote eye trackers have attracted much interest, since they enable nonintrusive gaze estimation. To achieve high estimation accuracies for remote systems, user calibration is inevitable in order to compensate for the estimation bias caused by person-specific eye parameters. Although several explicit and implicit user calibration methods have been proposed to ease the calibration burden, the procedure is still cumbersome and needs further improvement. In this paper, we present a comprehensive analysis of regression-based user calibration techniques. We propose a novel weighted least squares regression-based user calibration method together with a real-time cross-ratio based gaze estimation framework. The proposed system enables to obtain high estimation accuracy with minimum user effort, which leads to user-friendly HCI applications. Experimental results conducted on both simulations and user experiments show that our framework achieves a significant performance improvement over the state-of-the-art user calibration methods when only a few points are available for the calibration. Nuri Murat Arar, Hua Gao, Jean-Philippe Thiran |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | Soil Moisture Estimation by SAR in Alpine Fields Using Gaussian Process Regressor Trained by Model SimulationsabstractIn this paper, we address the problem of retrieving soil moisture over a grassland alpine area from Synthetic Aperture Radar (SAR) data using a statistical algorithm trained by simulations of a physical model. A time series of C-band VV-polarized Wide Swath images acquired by Envisat Advanced SAR (ASAR) in the snow-free periods of 2010 and 2011 was simulated using a discrete radiative transfer model (RTM). The test area was located in the Mazia valley, South Tyrol (Italy), where the main land types are meadows and pastures. Soil moisture was collected from five meteorological stations, two of which situated in meadows and the rest in pastures. The smallest and the highest RMSEs of the RTM simulations were 0.78 dB and 1.91 dB, respectively. After backscattering simulation, the top soil moisture was estimated using Gaussian Process Regression (GPR). GPR was trained with the backscatter model simulations (including terrain features) for 2010, and then used to predict moisture from radar observations acquired in 2011. The relative importance of different input features was also assessed. The RMSE of the predicted soil moisture for the largest training data set (including aspect as a terrain feature) was 5.6% Vol. and the corresponding correlation coefficient was 0.84. Jelena Stamenkovic, Leila Guerriero, Paolo Ferrazzoli, Claudia Notarnicola, Felix Greifeneder, Jean-Philippe Thiran |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2016 | Estimating fusion weights of a multi-camera eye tracking system by leveraging user calibration dataabstractCross-ratio (CR)-based eye tracking has been attracting much interest due to its simple setup, yet its accuracy is lower than that of the model-based approaches. In order to improve the estimation accuracy, a multi-camera setup can be exploited rather than the traditional single camera systems. The overall gaze point can be computed by fusion of available gaze information from all cameras. This paper presents a real-time multi-camera eye tracking system in which the estimation of gaze relies on simple CR geometry. A novel weighted fusion method is proposed, which leverages the user calibration data to learn the fusion weights. Experimental results conducted on real data show that the proposed method achieves a significant accuracy improvement over single camera systems. The real-time system achieves 0.82° of visual angle accuracy error with very few calibration data (5 points) under natural head movements, which is competitive with more complex model-based systems. Nuri Murat Arar, Jean-Philippe Thiran |
ETRA | 2 |
| 2016 | Single-FPGA, scalable, low-power, and high-quality 3D ultrasound beamformerabstractWe present an efficient FPGA architecture suitable for a medical 3D ultrasound beamformer. We tackle the delay calculation bottleneck, which is the heart and the most critical part of the beamformer, by proposing a computationally efficient design that is able to perform volumetric real-time beamforming on a single-chip FPGA. The design has been demonstrated for a 32×32-channel receive probe, and we extrapolated the requirements of the architecture for 80×80 channels. William Andrew Simon, Ahmet Caner Yuzuguler, Aya Ibrahim, Federico Angiolini, Marcel Arditi, Jean-Philippe Thiran, Giovanni De Micheli |
FPL | 6 |
| 2016 | Single-FPGA 3D ultrasound beamformerabstractIn medical diagnosis, ultrasound (US) imaging is one of the most common, safe, and powerful techniques. Volumetric (3D) US imaging, an emerging technique, is even more attractive than standard 2D imaging, as it allows for imaging without the local presence of a trained sonographer finely positioning the probe. This would be particularly useful in rescue operations, remote areas and developing countries. Unfortunately, present-day 3D imagers are expensive, bulky and power-hungry, confining them to hospitals. There is therefore a strong motivation to develop efficient electronics to enable a portable US platform that is small, cheap, and battery-operated. Beamforming (BF) is the most computationally expensive of 3D imaging. Both commercial [1] and research [2] imagers have dealt with the challenge by reducing the number of receive channels, hence simplifying the computation through the usage of far fewer elements. This comes at the cost of image quality, and the resulting machines are nonetheless still non-portable and expensive. In turn, the bottleneck of the BF process is the calculation of acoustic delays, which requires up to trillions of square roots per second. We propose a drastically more efficient architecture [3]. With geometric considerations, each delay is calculated from a small set of square roots (mapped onto CORDICs), plus two additions. In this demo, we will show the reconstruction of a 2.5M-voxel volume, supporting a transducer with 32×32 receive channels. We have fitted the architecture into a single Kintex UltraScale KU040 [4], which is unprecedented. We also extrapolated the utilization of a 80×80 instance on a Virtex UltraScale XCVU190 [4]. Table I shows the implementation results. Fig. 1 shows our beamformer custom block connected to the other FPGA subsystems. The delay calculation architecture is shown in Fig. 2. The demo setup is presented in Fig. 3, where the 3D beamformer is implemented on the FPGA, while the pre- and post-processing stages are currently performed on Matlab. Ahmet Caner Yuzuguler, William Andrew Simon, Aya Ibrahim, Federico Angiolini, Marcel Arditi, Jean-Philippe Thiran, Giovanni De Micheli |
FPL | 6 |
| 2016 | Compressed delay-and-sum beamforming for ultrafast ultrasound imagingabstractThe theory of compressed sensing (CS) leverages upon structure of signals in order to reduce the number of samples needed to reconstruct a signal, compared to the Nyquist rate. Although CS approaches have been proposed for ultrasound (US) imaging with promising results, practical implementations are hard to achieve due to the impossibility to mimic random sampling on a US probe and to the high memory requirements of the measurement model. In this paper, we propose a CS framework for US imaging based on an easily implementable acquisition scheme and on a delay-and-sum measurement model. Adrien Besson, Rafael E. Carrillo, Olivier Bernard 0001, Yves Wiaux, Jean-Philippe Thiran |
ICIP | 5 |
| 2016 | Kernel Low-Rank and Sparse Graph for Unsupervised and Semi-Supervised Classification of Hyperspectral ImagesabstractIn this paper, we present a graph representation that is based on the assumption that data live on a union of manifolds. Such a representation is based on sample proximities in reproducing kernel Hilbert spaces and is thus linear in the feature space and nonlinear in the original space. Moreover, it also expresses sample relationships under sparse and low-rank constraints, meaning that the resulting graph will have limited connectivity (sparseness) and that samples belonging to the same group will be likely to be connected together and not with those from other groups (low rankness). We present this graph representation as a general representation that can be then applied to any graph-based method. In the experiments, we consider the clustering of hyperspectral images and semi-supervised classification (one class and multiclass). Frank de Morsier, Maurice Borgeaud, Volker Gass, Jean-Philippe Thiran, Devis Tuia |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2015 | Accelerated microstructure imaging via convex optimisation for regions with multiple fibres (AMICOx)abstractThis paper reviews and extends our previous work to enable fast axonal diameter mapping from diffusion MRI data in the presence of multiple fibre populations within a voxel. Most of the existing mi-crostructure imaging techniques use non-linear algorithms to fit their data models and consequently, they are computationally expensive and usually slow. Moreover, most of them assume a single axon orientation while numerous regions of the brain actually present more complex configurations, e.g. fiber crossing. We present a flexible framework, based on convex optimisation, that enables fast and accurate reconstructions of the microstructure organisation, not limited to areas where the white matter is coherently oriented. We show through numerical simulations the ability of our method to correctly estimate the microstructure features (mean axon diameter and intra-cellular volume fraction) in crossing regions. Anna Auría, David Romascano, E. Canales-Rodriguen, Yves Wiaux, T. B. Dirby, Daniel C. Alexander, Jean-Philippe Thiran, Alessandro Daducci |
ICIP | 7 |
| 2015 | Towards Convenient Calibration for Cross-Ratio Based Gaze EstimationabstractEye gaze movements are considered as a salient modality for human computer interaction applications. Recently, cross-ratio (CR) based eye tracking methods have attracted increasing interest because they provide remote gaze estimation using a single uncalibrated camera. However, due to the simplification assumptions in CR-based methods, their performance is lower than the model-based approaches [8]. Several efforts have been made to improve the accuracy by compensating for the assumptions with subject specific calibration. This paper presents a CR-based automatic gaze estimation system that accurately works under natural head movements. A subject-specific calibration method based on regularized least-squares regression (LSR) is introduced for achieving higher accuracy compared to other state-of-the-art calibration methods. Experimental results also show that the proposed calibration method generalizes better when fewer calibration points are used. This enables user friendly applications with minimum calibration effort without sacrificing too much accuracy. In addition, we adaptively fuse the estimation of the point of regard (PoR) from both eyes based on the visibility of eye features. The adaptive fusion scheme reduces accuracy error by around 20% and also increases the estimation coverage under natural head movements. Nuri Murat Arar, Hua Gao, Jean-Philippe Thiran |
WACV | 3 |
| 2015 | Cluster validity measure and merging system for hierarchical clustering considering outliers
Frank de Morsier, Devis Tuia, Maurice Borgeaud, Volker Gass, Jean-Philippe Thiran |
Pattern Recognit. | 5 |
| 2015 | Prediction of asynchronous dimensional emotion ratings from audiovisual and physiological data
Fabien Ringeval, Florian Eyben, Eleni Kroupi, Anil Yüce, Jean-Philippe Thiran, Touradj Ebrahimi, Denis Lalanne, Björn W. Schuller |
Pattern Recognit. Lett. | 5 |
| 2015 | COMMIT: Convex Optimization Modeling for Microstructure Informed TractographyabstractTractography is a class of algorithms aiming at in vivo mapping the major neuronal pathways in the white matter from diffusion magnetic resonance imaging (MRI) data. These techniques offer a powerful tool to noninvasively investigate at the macroscopic scale the architecture of the neuronal connections of the brain. However, unfortunately, the reconstructions recovered with existing tractography algorithms are not really quantitative even though diffusion MRI is a quantitative modality by nature. As a matter of fact, several techniques have been proposed in recent years to estimate, at the voxel level, intrinsic microstructural features of the tissue, such as axonal density and diameter, by using multicompartment models. In this paper, we present a novel framework to reestablish the link between tractography and tissue microstructure. Starting from an input set of candidate fiber-tracts, which are estimated from the data using standard fiber-tracking techniques, we model the diffusion MRI signal in each voxel of the image as a linear combination of the restricted and hindered contributions generated in every location of the brain by these candidate tracts. Then, we seek for the global weight of each of them, i.e., the effective contribution or volume, such that they globally fit the measured signal at best. We demonstrate that these weights can be easily recovered by solving a global convex optimization problem and using efficient algorithms. The effectiveness of our approach has been evaluated both on a realistic phantom with known ground-truth and in vivo brain data. Results clearly demonstrate the benefits of the proposed formulation, opening new perspectives for a more quantitative and biologically plausible assessment of the structural connectivity of the brain. Alessandro Daducci, Alessandro Dal Palù, Alia Lemkaddem, Jean-Philippe Thiran |
IEEE Trans. Medical Imaging | 4 |
| 2014 | Sparsity in tensor optimization for optical-interferometric imagingabstractImage recovery in optical interferometry is an ill-posed nonlinear inverse problem arising from incomplete power spectrum and bispectrum measurements. We review our previous work, which reformulates this nonlinear problem in the framework of tensor recovery and studies two different approaches to solve it: one is nonlinear and nonconvex while the other is linear and convex. We extend the linear convex procedure to account for signal sparsity and we also present numerical simulations that show the improvement in the quality of reconstruction of sparse images when including a sparsity prior. Anna Auría, Rafael E. Carrillo, Jean-Philippe Thiran, Yves Wiaux |
ICIP | 3 |
| 2014 | Detecting emotional stress from facial expressions for driving safetyabstractMonitoring the attentive and emotional status of the driver is critical for the safety and comfort of driving. In this work a real-time non-intrusive monitoring system is developed, which detects the emotional states of the driver by analyzing facial expressions. The system considers two negative basic emotions, anger and disgust, as stress related emotions. We detect an individual emotion in each video frame and the decision on the stress level is made on sequence level. Experimental results show that the developed system operates very well on simulated data even with generic models. An additional pose normalization step reduces the impact of pose mismatch due to camera setup and pose variation, and hence improves the detection accuracy further. Hua Gao, Anil Yüce, Jean-Philippe Thiran |
ICIP | 3 |
| 2014 | Non-linear low-rank and sparse representation for hyperspectral image analysisabstractIn this paper, we tackle the problem of unsupervised classification of hyperspectral images. We propose a clustering method based on graphs representing the data structure, which is assumed to be an union of multiple manifolds. The method constraints the pixels to be expressed as a low-rank and sparse combination of the others in a reproducing kernel Hilbert spaces (RKHS). This captures the global (low-rank) and local (sparse) structures. Spectral clustering is applied on the graph to assign the pixels to the different manifolds. A large scale approach is proposed, in which the optimization is first performed on a subset of the data and then it is applied to the whole image using a non-linear collaborative representation respecting the manifolds structure. Experiments on two hyperspectral images show very good unsupervised classification results compared to competitive approaches. Frank de Morsier, Devis Tuia, Maurice Borgeaucft, Volker Gass, Jean-Philippe Thiran |
IGARSS | 5 |
| 2014 | Crop backscatter modeling and soil moisture estimation with support vector regressionabstractIn this paper, we used an improved version of the Tor Vergata radiative transfer model to simulate the backscattering coefficient for the L-band SAR signals over areas covered with vegetation. Fields of winter wheat, maize and sugar beet observed during the AgriSAR2006 campaign were investigated. For maize field, the presence of periodic soil surface profiles played an important role in determining the total backscattering. Soil moisture was also estimated using an inverse algorithm based on a supervised, non-parametric learning technique, v-SVR. v-SVR proved good generalization properties even with a limited number of training samples available. Dependence to the origin of training samples, as well as the influence of different features, was thoroughly considered. Jelena Stamenkovic, Paolo Ferrazzoli, Leila Guerriero, Devis Tuia, Jean-Philippe Thiran, Maurice Borgeaud |
IGARSS | 5 |
| 2014 | Efficient Total Variation Algorithm for Fetal Brain MRI Reconstruction
Sébastien Tourbier, Xavier Bresson, Patric Hagmann, Jean-Philippe Thiran, Reto Meuli, Meritxell Bach Cuadra |
MICCAI (2) | 4 |
| 2014 | Sparse regularization for fiber ODF reconstruction: From the suboptimality of l2 and l1 priors to l0
Alessandro Daducci, Dimitri Van De Ville, Jean-Philippe Thiran, Yves Wiaux |
Medical Image Anal. | 3 |
| 2014 | Surface Reconstruction From Microscopic Images in Optical LithographyabstractThis paper presents a method to reconstruct 3D surfaces of silicon wafers from 2D images of printed circuits taken with a scanning electron microscope. Our reconstruction method combines the physical model of the optical acquisition system with prior knowledge about the shapes of the patterns in the circuit; the result is a shape-from-shading technique with a shape prior. The reconstruction of the surface is formulated as an optimization problem with an objective functional that combines a data-fidelity term on the microscopic image with two prior terms on the surface. The data term models the acquisition system through the irradiance equation characteristic of the microscope; the first prior is a smoothness penalty on the reconstructed surface, and the second prior constrains the shape of the surface to agree with the expected shape of the pattern in the circuit. In order to account for the variability of the manufacturing process, this second prior includes a deformation field that allows a nonlinear elastic deformation between the expected pattern and the reconstructed surface. As a result, the minimization problem has two unknowns, and the reconstruction method provides two outputs: 1) a reconstructed surface and 2) a deformation field. The reconstructed surface is derived from the shading observed in the image and the prior knowledge about the pattern in the circuit, while the deformation field produces a mapping between the expected shape and the reconstructed surface that provides a measure of deviation between the circuit design models and the real manufacturing process. Virginia Estellers, Jean-Philippe Thiran, Maria Gabrani |
IEEE Trans. Image Process. | 2 |
| 2014 | Harmonic Active ContoursabstractWe propose a segmentation method based on the geometric representation of images as 2-D manifolds embedded in a higher dimensional space. The segmentation is formulated as a minimization problem, where the contours are described by a level set function and the objective functional corresponds to the surface of the image manifold. In this geometric framework, both data-fidelity and regularity terms of the segmentation are represented by a single functional that intrinsically aligns the gradients of the level set function with the gradients of the image and results in a segmentation criterion that exploits the directional information of image gradients to overcome image inhomogeneities and fragmented contours. The proposed formulation combines this robust alignment of gradients with attractive properties of previous methods developed in the same geometric framework: 1) the natural coupling of image channels proposed for anisotropic diffusion and 2) the ability of subjective surfaces to detect weak edges and close fragmented boundaries. The potential of such a geometric approach lies in the general definition of Riemannian manifolds, which naturally generalizes existing segmentation methods (the geodesic active contours, the active contours without edges, and the robust edge integrator) to higher dimensional spaces, non-flat images, and feature spaces. Our experiments show that the proposed technique improves the segmentation of multi-channel images, images subject to inhomogeneities, and images characterized by geometric structures like ridges or valleys. Virginia Estellers, Dominique Zosso, Xavier Bresson, Jean-Philippe Thiran |
IEEE Trans. Image Process. | 4 |
| 2014 | Fast Geodesic Active Fields for Image Registration Based on Splitting and Augmented Lagrangian ApproachesabstractIn this paper, we present an efficient numerical scheme for the recently introduced geodesic active fields (GAF) framework for geometric image registration. This framework considers the registration task as a weighted minimal surface problem. Hence, the data-term and the regularization-term are combined through multiplication in a single, parametrization invariant and geometric cost functional. The multiplicative coupling provides an intrinsic, spatially varying and data-dependent tuning of the regularization strength, and the parametrization invariance allows working with images of nonflat geometry, generally defined on any smoothly parametrizable manifold. The resulting energy-minimizing flow, however, has poor numerical properties. Here, we provide an efficient numerical scheme that uses a splitting approach; data and regularity terms are optimized over two distinct deformation fields that are constrained to be equal via an augmented Lagrangian approach. Our approach is more flexible than standard Gaussian regularization, since one can interpolate freely between isotropic Gaussian and anisotropic TV-like smoothing. In this paper, we compare the geodesic active fields method with the popular Demons method and three more recent state-of-the-art algorithms: NL-optical flow, MRF image registration, and landmark-enhanced large displacement optical flow. Thus, we can show the advantages of the proposed FastGAF method. It compares favorably against Demons, both in terms of registration speed and quality. Over the range of example applications, it also consistently produces results not far from more dedicated state-of-the-art methods, illustrating the flexibility of the proposed framework. Dominique Zosso, Xavier Bresson, Jean-Philippe Thiran |
IEEE Trans. Image Process. | 3 |
| 2014 | Quantitative Comparison of Reconstruction Methods for Intra-Voxel Fiber Recovery From Diffusion MRIabstractValidation is arguably the bottleneck in the diffusion magnetic resonance imaging (MRI) community. This paper evaluates and compares 20 algorithms for recovering the local intra-voxel fiber structure from diffusion MRI data and is based on the results of the "HARDI reconstruction challenge" organized in the context of the "ISBI 2012" conference. Evaluated methods encompass a mixture of classical techniques well known in the literature such as diffusion tensor, Q-Ball and diffusion spectrum imaging, algorithms inspired by the recent theory of compressed sensing and also brand new approaches proposed for the first time at this contest. To quantitatively compare the methods under controlled conditions, two datasets with known ground-truth were synthetically generated and two main criteria were used to evaluate the quality of the reconstructions in every voxel: correct assessment of the number of fiber populations and angular accuracy in their orientation. This comparative study investigates the behavior of every algorithm with varying experimental conditions and highlights strengths and weaknesses of each approach. This information can be useful not only for enhancing current algorithms and develop the next generation of reconstruction methods, but also to assist physicians in the choice of the most adequate technique for their studies. Alessandro Daducci, Erick Jorge Canales-Rodríguez, Maxime Descoteaux, Eleftherios Garyfallidis, Yaniv Gur, Ying-Chia Lin, Merry Mani, Sylvain Merlet, Michael Paquette, Alonso Ramirez-Manzanares, Marco Reisert, Paulo Reis Rodrigues, Farshid Sepehrband, Emmanuel Caruyer, Jeiran Choupan, Rachid Deriche, Mathews Jacob, Gloria Menegaz, Vesna Prckovska, Mariano Rivera, Yves Wiaux, Jean-Philippe Thiran |
IEEE Trans. Medical Imaging | 22 |
| 2013 | Sparsity Averaging for Compressive ImagingabstractWe discuss a novel sparsity prior for compressive imaging in the context of the theory of compressed sensing with coherent redundant dictionaries, based on the observation that natural images exhibit strong average sparsity over multiple coherent frames. We test our prior and the associated algorithm, based on an analysis reweighted formulation, through extensive numerical simulations on natural images for spread spectrum and random Gaussian acquisition schemes. Our results show that average sparsity outperforms state-of-the-art priors that promote sparsity in a single orthonormal basis or redundant frame, or that promote gradient sparsity. Code and test data are available at https://github.com/basp-group/sopt. Rafael E. Carrillo, Jason D. McEwen, Dimitri Van De Ville, Jean-Philippe Thiran, Yves Wiaux |
IEEE Signal Process. Lett. | 4 |
| 2013 | Weighted Shape-Based Averaging With Neighborhood Prior Model for Multiple Atlas Fusion-Based Medical Image SegmentationabstractIn medical imaging, merging automated segmentations obtained from multiple atlases has become a standard practice for improving the accuracy. In this letter, we propose two new fusion methods: “Global Weighted Shape-Based Averaging” (GWSBA) and “Local Weighted Shape-Based Averaging” (LWSBA). These methods extend the well known Shape-Based Averaging (SBA) by additionally incorporating the similarity information between the reference (i.e., atlas) images and the target image to be segmented. We also propose a new spatially-varying similarity-weighted neighborhood prior model, and an edge-preserving smoothness term that can be used with many of the existing fusion methods. We first present our new Markov Random Field (MRF) based fusion framework that models the above mentioned information. The proposed methods are evaluated in the context of segmentation of lymph nodes in the head and neck 3D CT images, and they resulted in more accurate segmentations compared to the existing SBA. Subrahmanyam Gorthi, Meritxell Bach Cuadra, Pierre-Alain Tercier, Abdelkarim Allal, Jean-Philippe Thiran |
IEEE Signal Process. Lett. | 5 |
| 2013 | Sparse Reverberant Audio Source Separation via Reweighted AnalysisabstractWe propose a novel algorithm for source signals estimation from an underdetermined convolutive mixture assuming known mixing filters. Most of the state-of-the-art methods are dealing with anechoic or short reverberant mixture, assuming a synthesis sparse prior in the time-frequency domain and a narrowband approximation of the convolutive mixing process. In this paper, we address the source estimation of convolutive mixtures with a new algorithm based on i) an analysis sparse prior, ii) a reweighting scheme so as to increase the sparsity, iii) a wideband data-fidelity term in a constrained form. We show, through theoretical discussions and simulations, that this algorithm is particularly well suited for source separation of realistic reverberation mixtures. Particularly, the proposed algorithm outperforms state-of-the-art methods on reverberant mixtures of audio sources by more than 2 dB of signal-to-distortion ratio on the BSS Oracle dataset. Simon Arberet, Pierre Vandergheynst, Rafael E. Carrillo, Jean-Philippe Thiran, Yves Wiaux |
IEEE Trans. Speech Audio Process. | 4 |
| 2013 | Source/Filter Factorial Hidden Markov Model, With Application to Pitch and Formant TrackingabstractTracking vocal tract formant frequencies (fp) and estimating the fundamental frequency (f0) are two tracking problems that have been tackled in many speech processing works, often independently, with applications to articulatory parameters estimations, speech analysis/synthesis or linguistics. Many works assume an auto-regressive (AR) model to fit the spectral envelope, hence indirectly estimating the formant tracks from the AR parameters. However, directly estimating the formant frequencies, or equivalently the poles of the AR filter, allows to further model the smoothness of the desired tracks. In this paper, we propose a Factorial Hidden Markov Model combined with a vocal source/filter model, with parameters naturally encoding the f0and fptracks. Two algorithms are proposed, with two different strategies: first, a simplification of the underlying model, with a parameter estimation based on variational methods, and second, a sparse decomposition of the signal, based on Non-negative Matrix Factorization methodology. The results are comparable to state-of-the-art formant tracking algorithms. With the use of a complete production model, the proposed systems provide robust formant tracks which can be used in various applications. The algorithms could also be extended to deal with multiple-speaker signals. Jean-Louis Durrieu, Jean-Philippe Thiran |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2013 | Sample and Pixel Weighting Strategies for Robust Incremental Visual TrackingabstractIn this paper, we introduce the incremental temporally weighted principal component analysis (ITWPCA) algorithm, based on singular value decomposition update, and the incremental temporally weighted visual tracking with spatial penalty (ITWVTSP) algorithm for robust visual tracking. ITWVTSP uses ITWPCA for computing incrementally a robust low dimensional subspace representation (model) of the tracked object. The robustness is based on the capacity of weighting the contribution of each single sample to the subspace generation to reduce the impact of bad quality samples, reducing the risk of model drift. Furthermore, ITWVTSP can exploit the a priori knowledge about important regions of a tracked object. This is done by penalizing the tracking error on some predefined regions of the tracked object, which increases the accuracy of tracking. Several tests are performed on several challenging video sequences, showing the robustness and accuracy of the proposed algorithm, as well as its superiority with respect to state-of-the-art techniques. Javier Cruz-Mota, Michel Bierlaire, Jean-Philippe Thiran |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2013 | Semi-Supervised Novelty Detection Using SVM Entire Solution PathabstractVery often, the only reliable information available to perform change detection is the description of some “unchanged” regions. Since, sometimes, these regions do not contain all the relevant information to identify their counterpart (the changes), we consider the use of unlabeled data to perform semi-supervised novelty detection (SSND). SSND can be seen as an unbalanced classification problem solved using the cost-sensitive support vector machine (CS-SVM), but this requires a heavy parameter search. Here, we propose the use of entire solution path algorithms for the CS-SVM in order to facilitate and accelerate parameter selection for SSND. Two algorithms are considered and evaluated. The first algorithm is an extension of the CS-SVM algorithm that returns the entire solution path in a single optimization. This way, optimization of a separate model for each hyperparameter set is avoided. The second algorithm forces the solution to be coherent through the solution path, thus producing classification boundaries that are nested (included in each other). We also present a low-density (LD) criterion for selecting optimal classification boundaries, thus avoiding recourse to cross validation (CV) that usually requires information about the “change” class. Experiments are performed on two multitemporal change detection data sets (flood and fire detection). Both algorithms tracing the solution path provide similar performances than the standard CS-SVM while being significantly faster. The proposed LD criterion achieves results that are close to the ones obtained by CV but without using information about the changes. Frank de Morsier, Devis Tuia, Maurice Borgeaud, Volker Gass, Jean-Philippe Thiran |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2013 | Enhanced Compressed Sensing Recovery With Level Set NormalsabstractWe propose a compressive sensing algorithm that exploits geometric properties of images to recover images of high quality from few measurements. The image reconstruction is done by iterating the two following steps: 1) estimation of normal vectors of the image level curves, and 2) reconstruction of an image fitting the normal vectors, the compressed sensing measurements, and the sparsity constraint. The proposed technique can naturally extend to nonlocal operators and graphs to exploit the repetitive nature of textured images to recover fine detail structures. In both cases, the problem is reduced to a series of convex minimization problems that can be efficiently solved with a combination of variable splitting and augmented Lagrangian methods, leading to fast and easy-to-code algorithms. Extended experiments show a clear improvement over related state-of-the-art algorithms in the quality of the reconstructed images and the robustness of the proposed method to noise, different kind of images, and reduced measurements. Virginia Estellers, Jean-Philippe Thiran, Xavier Bresson |
IEEE Trans. Image Process. | 2 |
| 2013 | Sparse Image Reconstruction on the Sphere: Implications of a New Sampling TheoremabstractWe study the impact of sampling theorems on the fidelity of sparse image reconstruction on the sphere. We discuss how a reduction in the number of samples required to represent all information content of a band-limited signal acts to improve the fidelity of sparse image reconstruction, through both the dimensionality and sparsity of signals. To demonstrate this result, we consider a simple inpainting problem on the sphere and consider images sparse in the magnitude of their gradient. We develop a framework for total variation inpainting on the sphere, including fast methods to render the inpainting problem computationally feasible at high resolution. Recently a new sampling theorem on the sphere was developed, reducing the required number of samples by a factor of two for equiangular sampling schemes. Through numerical simulations, we verify the enhanced fidelity of sparse image reconstruction due to the more efficient sampling of the sphere provided by the new sampling theorem. Jason D. McEwen, Gilles Puy, Jean-Philippe Thiran, Pierre Vandergheynst, Dimitri Van De Ville, Yves Wiaux |
IEEE Trans. Image Process. | 3 |
| 2012 | Lower and upper bounds for approximation of the Kullback-Leibler divergence between Gaussian Mixture ModelsabstractMany speech technology systems rely on Gaussian Mixture Models (GMMs). The need for a comparison between two GMMs arises in applications such as speaker verification, model selection or parameter estimation. For this purpose, the Kullback-Leibler (KL) divergence is often used. However, since there is no closed form expression to compute it, it can only be approximated. We propose lower and upper bounds for the KL divergence, which lead to a new approximation and interesting insights into previously proposed approximations. An application to the comparison of speaker models also shows how such approximations can be used to validate assumptions on the models. Jean-Louis Durrieu, Jean-Philippe Thiran, Finnian Kelly |
ICASSP | 2 |
| 2012 | Fast globally supervised segmentation by active contours with shape and texture descriptorsabstractWe present a new globally supervised segmentation method in the characteristic function framework based on an active contours (AC) model incorporating both shape prior and texture descriptors. The shape prior descriptor is formulated as the traditional Legendre moment and the texture descriptor as a linear combination of local inside/outside texture descriptor. Using these two descriptors, the AC energy incorporates both learned textures and training shapes. This formulation has two main advantages: 1) by discriminating independently the foreground/background textures. 2) by incorporating both the learned inside/outside texture and the training shape. The trade-off between inside and outside texture descriptor is ensured by balancing descriptor. We illustrate the performance of our segmentation algorithm using some challenging textured images. Foued Derraz, Jean-Philippe Thiran, Abdelmalik Taleb-Ahmed, Laurent Peyrodie, Gérard Forzy |
ICIP | 2 |
| 2012 | Semi-supervised and unsupervised novelty detection using nested support vector machinesabstractVery often in change detection only few labels or even none are available. In order to perform change detection in these extreme scenarios, they can be considered as novelty detection problems, semi-supervised (SSND) if some labels are available otherwise unsupervised (UND). SSND can be seen as an unbalanced classification between labeled and unlabeled samples using the Cost-Sensitive Support Vector Machine (CS-SVM). UND assumes novelties in low density regions and can be approached using the One-Class SVM (OC-SVM). We propose here to use nested entire solution path algorithms for the OC-SVM and CS-SVM in order to accelerate the parameter selection and alleviate the dependency to labeled “changed” samples. Experiments are performed on two multitemporal change detection datasets (flood and fire detection) and the performance of the two methods proposed compared. Frank de Morsier, Maurice Borgeaud, Christoph Küchler, Volker Gass, Jean-Philippe Thiran |
IGARSS | 5 |
| 2012 | Scale Invariant Feature Transform on the Sphere: Theory and Applications
Javier Cruz-Mota, Iva Bogdanova, Benoît Paquier, Michel Bierlaire, Jean-Philippe Thiran |
Int. J. Comput. Vis. | 5 |
| 2012 | On Dynamic Stream Weighting for Audio-Visual Speech RecognitionabstractThe integration of audio and visual information improves speech recognition performance, specially in the presence of noise. In these circumstances it is necessary to introduce audio and visual weights to control the contribution of each modality to the recognition task. We present a method to set the value of the weights associated to each stream according to their reliability for speech recognition, allowing them to change with time and adapt to different noise and working conditions. Our dynamic weights are derived from several measures of the stream reliability, some specific to speech processing and others inherent to any classification task, and take into account the special role of silence detection in the definition of audio and visual weights. In this paper, we propose a new confidence measure, compare it to existing ones, and point out the importance of the correct detection of silence utterances in the definition of the weighting system. Experimental results support our main contribution: the inclusion of a voice activity detector in the weighting scheme improves speech recognition over different system architectures and confidence measures, leading to an increase in performance more relevant than any difference between the proposed confidence measures. Virginia Estellers, Mihai Gurban, Jean-Philippe Thiran |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | Efficient Algorithm for Level Set Method Preserving Distance FunctionabstractThe level set method is a popular technique for tracking moving interfaces in several disciplines, including computer vision and fluid dynamics. However, despite its high flexibility, the original level set method is limited by two important numerical issues. First, the level set method does not implicitly preserve the level set function as a distance function, which is necessary to estimate accurately geometric features, s.a. the curvature or the contour normal. Second, the level set algorithm is slow because the time step is limited by the standard Courant-Friedrichs-Lewy (CFL) condition, which is also essential to the numerical stability of the iterative scheme. Recent advances with graph cut methods and continuous convex relaxation methods provide powerful alternatives to the level set method for image processing problems because they are fast, accurate, and guaranteed to find the global minimizer independently to the initialization. These recent techniques use binary functions to represent the contour rather than distance functions, which are usually considered for the level set method. However, the binary function cannot provide the distance information, which can be essential for some applications, s.a. the surface reconstruction problem from scattered points and the cortex segmentation problem in medical imaging. In this paper, we propose a fast algorithm to preserve distance functions in level set methods. Our algorithm is inspired by recent efficient l(1) optimization techniques, which will provide an efficient and easy to implement algorithm. It is interesting to note that our algorithm is not limited by the CFL condition and it naturally preserves the level set function as a distance function during the evolution, which avoids the classical re-distancing problem in level set methods. We apply the proposed algorithm to carry out image segmentation, where our methods prove to be 5-6 times faster than standard distance preserving level set techniques. We also present two applications where preserving a distance function is essential. Nonetheless, our method stays generic and can be applied to any level set methods that require the distance information. Virginia Estellers, Dominique Zosso, Rongjie Lai, Stanley J. Osher, Jean-Philippe Thiran, Xavier Bresson |
IEEE Trans. Image Process. | 5 |
| 2012 | Spread Spectrum Magnetic Resonance ImagingabstractWe propose a novel compressed sensing technique to accelerate the magnetic resonance imaging (MRI) acquisition process. The method, coined spread spectrum MRI or simply s(2)MRI, consists of premodulating the signal of interest by a linear chirp before random k-space under-sampling, and then reconstructing the signal with nonlinear algorithms that promote sparsity. The effectiveness of the procedure is theoretically underpinned by the optimization of the coherence between the sparsity and sensing bases. The proposed technique is thoroughly studied by means of numerical simulations, as well as phantom and in vivo experiments on a 7T scanner. Our results suggest that s(2)MRI performs better than state-of-the-art variable density k-space under-sampling approaches. Gilles Puy, José P. Marques, Rolf Gruetter, Jean-Philippe Thiran, Dimitri Van De Ville, Pierre Vandergheynst, Yves Wiaux |
IEEE Trans. Medical Imaging | 4 |
| 2011 | Sparse non-negative decomposition of speech power spectra for formant trackingabstractMany works on speech processing have dealt with auto-regressive (AR) models for spectral envelope and formant frequency estimation, mostly focusing on the estimation of the AR parameters. How ever, it is also interesting to be able to directly estimate the formant frequencies, or equivalently the poles of the AR filter. To tackle this issue, we propose in this paper to decompose the signal onto several bases, one for each formant, taking advantage of recent works on nonnegative matrix factorization (NMF) for the estimation stage, further refined by sparsity and smoothness penalties. The results are encouraging, and the proposed system provides formant tracks which seem robust enough to be used in different applications such as phonetic analysis, emotion detection or as visual cue for computer-aided pronunciation training applications. The model can also be extended to deal with multiple-speaker signals. Jean-Louis Durrieu, Jean-Philippe Thiran |
ICASSP | 2 |
| 2011 | Harmonic active contours for multichannel image segmentationabstractWe propose a segmentation method based on the geometric representation of images as surfaces embedded in a higher dimensional space, handling naturally multichannel images. The segmentation is based on an active contour embedded in the image manifold, along with a set of image features. Hence, both data-fidelity and regularity terms of the active contour are jointly optimized minimizing a single Polaykov energy representing the hyper-surface of this manifold. Compared to previous methods, our approach is purely geometrical and does not require additional weighting of the energy functional to drive the segmentation to the image contours. The potential of such a geometric approach lies in the general definition of Riemannian manifolds, validating the proposed technique for scale-space methods, volumetric data or catadioptric images. We present here the segmentation technique called Harmonic Active Contours, give an implementation for multichannel images including gradient and region-based segmentation criteria and apply it to color images. Virginia Estellers, Dominique Zosso, Xavier Bresson, Jean-Philippe Thiran |
ICIP | 4 |
| 2011 | Towards a diffusion image processing validation and accuracy prediction frameworkabstractValidation is the main bottleneck preventing the adoption of many medical image processing algorithms in the clinical practice. In the classical approach, a-posteriori analysis is performed based on some objective metrics. In this work, a different approach based on Petri Nets (PN) is proposed. The basic idea consists in predicting the accuracy that will result from a given processing based on the characterization of the sources of inaccuracy of the system. Here we propose a proof of concept in the scenario of a diffusion imaging analysis pipeline. A PN is built after the detection of the possible sources of inaccuracy. By integrating the first qualitative insights based on the PN with quantitative measures, it is possible to optimize the PN itself, to predict the inaccuracy of the system in a different setting. Results show that the proposed model provides a good prediction performance and suggests the optimal processing approach. Francesca Pizzorni Ferrarese, Alessandro Daducci, Meritxell Bach Cuadra, Alia Lemkaddem, Cristina Granziera, Jean-Philippe Thiran, Gloria Menegaz |
ICIP | 6 |
| 2011 | Comparison of energy minimization methods for 3-D brain tissue classificationabstractThis paper presents 3-D brain tissue classification schemes using three recent promising energy minimization methods for Markov random fields: graph cuts, loopy belief propagation and tree-reweighted message passing. The classification is performed us ng the well known finite Gaussian mixture Markov Random Field model. Results from the above methods are compared with widely used iterative conditional modes algorithm. The evaluation is per formed on a dataset containing simulated Tl-weighted MR brain volumes with varying noise and intensity non-uniformities. The comparisons are performed in terms of energies as well as based on ground truth segmentations, using various quantitative metrics. Subrahmanyam Gorthi, Jean-Philippe Thiran, Meritxell Bach Cuadra |
ICIP | 2 |
| 2011 | Active deformation fields: Dense deformation field estimation for atlas-based segmentation using the active contour framework
Subrahmanyam Gorthi, Valerie Duay, Xavier Bresson, Meritxell Bach Cuadra, Francisco Javier Sánchez Castro, Claudio Pollo, Abdelkarim Allal, Jean-Philippe Thiran |
Medical Image Anal. | 8 |
| 2011 | Geodesic Active Fields - A Geometric Framework for Image RegistrationabstractIn this paper we present a novel geometric framework called geodesic active fields for general image registration. In image registration, one looks for the underlying deformation field that best maps one image onto another. This is a classic ill-posed inverse problem, which is usually solved by adding a regularization term. Here, we propose a multiplicative coupling between the registration term and the regularization term, which turns out to be equivalent to embed the deformation field in a weighted minimal surface problem. Then, the deformation field is driven by a minimization flow toward a harmonic map corresponding to the solution of the registration problem. This proposed approach for registration shares close similarities with the well-known geodesic active contours model in image segmentation, where the segmentation term (the edge detector function) is coupled with the regularization term (the length functional) via multiplication as well. As a matter of fact, our proposed geometric model is actually the exact mathematical generalization to vector fields of the weighted length problem for curves and surfaces introduced by Caselles-Kimmel-Sapiro. The energy of the deformation field is measured with the Polyakov energy weighted by a suitable image distance, borrowed from standard registration models. We investigate three different weighting functions, the squared error and the approximated absolute error for monomodal images, and the local joint entropy for multimodal images. As compared to specialized state-of-the-art methods tailored for specific applications, our geometric framework involves important contributions. Firstly, our general formulation for registration works on any parametrizable, smooth and differentiable surface, including nonflat and multiscale images. In the latter case, multiscale images are registered at all scales simultaneously, and the relations between space and scale are intrinsically being accounted for. Second, this method is, to the best of our knowledge, the first reparametrization invariant registration method introduced in the literature. Thirdly, the multiplicative coupling between the registration term, i.e. local image discrepancy, and the regularization term naturally results in a data-dependent tuning of the regularization strength. Finally, by choosing the metric on the deformation field one can freely interpolate between classic Gaussian and more interesting anisotropic, TV-like regularization. Dominique Zosso, Xavier Bresson, Jean-Philippe Thiran |
IEEE Trans. Image Process. | 3 |
| 2010 | Spread spectrum for interferometric and magnetic resonance imagingabstractWe consider images probed through incomplete and noisy Fourier coverages, both in the context of radio interferometry (RI) and of magnetic resonance imaging (MRI). We show that the quality of signal reconstruction can be significantly enhanced by the introduction of a linear chirp modulation, which induces a spread spectrum phenomenon. Gilles Puy, Yves Wiaux, Rolf Gruetter, Jean-Philippe Thiran, Dimitri Van De Ville, Pierre Vandergheynst |
ICASSP | 4 |
| 2010 | Geodesic Active Fields on the SphereabstractIn this paper, we propose a novel method to register images defined on spherical meshes. Instances of such spherical images include inflated cortical feature maps in brain medical imaging or images from omni directional cameras. We apply the Geodesic Active Fields (GAF) framework locally at each vertex of the mesh. Therefore we define a dense deformation field, which is embedded in a higher dimensional manifold, and minimize the weighted Polyakov energy. While the Polyakov energy itself measures the hyper area of the embedded deformation field, its weighting allows to account for the quality of the current image alignment. Iteratively minimizing the energy drives the deformation field towards a smooth solution of the registration problem. Although the proposed approach does not necessarily outperform state-of-the-art methods that are tightly tailored to specific applications, it is of methodological interest due to its high degree of flexibility and versatility. Dominique Zosso, Jean-Philippe Thiran |
ICPR | 2 |
| 2010 | Overcoming asynchrony in Audio-Visual Speech RecognitionabstractIn this paper we propose two alternatives to overcome the natural asynchrony of modalities in Audio-Visual Speech Recognition. We first investigate the use of asynchronous statistical models based on Dynamic Bayesian Networks with different levels of asynchrony. We show that audio-visual models should consider asynchrony within word boundaries and not at phoneme level. The second approach to the problem includes an additional processing of the features before being used for recognition. The proposed technique aligns the temporal evolution of the audio and video streams in terms of a speech-recognition system and enables the use of simpler statistical models for classification. On both cases we report experiments with the CUAVE database, showing the improvements obtained with the proposed asynchronous model and feature processing technique compared to traditional systems. Virginia Estellers, Jean-Philippe Thiran |
MMSP | 2 |
| 2010 | Modelling human perception of static facial expressions
Matteo Sorci, Gianluca Antonini, Javier Cruz, Thomas Robin 0001, Michel Bierlaire, Jean-Philippe Thiran |
Image Vis. Comput. | 6 |
| 2010 | Information theoretic combination of pattern classifiers
Julien Meynet, Jean-Philippe Thiran |
Pattern Recognit. | 2 |
| 2009 | Classification of tensors and fiber tracts using Mercer-kernels encoding soft probabilistic spatial and diffusion informationabstractIn this paper, we present a kernel-based approach to the clustering of diffusion tensors and fiber tracts. We propose to use a Mercer kernel over the tensor space where both spatial and diffusion information are taken into account. This kernel highlights implicitly the connectivity along fiber tracts. Tensor segmentation is performed using kernel-PCA compounded with a landmark-Isomap embedding and k-means clustering. Based on a soft fiber representation, we extend the tensor kernel to deal with fiber tracts using the multi-instance kernel that reflects not only interactions between points along fiber tracts, but also the interactions between diffusion tensors. This unsupervised method is further extended by way of an atlas-based registration of diffusion-free images, followed by a classification of fibers based on nonlinear kernel Support Vector Machines (SVMs). Promising experimental results of tensor and fiber classification of the human skeletal muscle over a significant set of healthy and diseased subjects demonstrate the potential of our approach. Radhouène Neji, Nikos Paragios, Gilles Fleury, Jean-Philippe Thiran, Georg Langs |
CVPR | 4 |
| 2009 | Non-Euclidean image-adaptive Radial Basis Functions for 3D interactive segmentationabstractIn the context of variational image segmentation, we propose a new finite-dimensional implicit surface representation. The key idea is to span a subset of implicit functions with linear combinations of spatially-localized kernels that follow image features. This is achieved by replacing the Euclidean distance in conventional Radial Basis Functions with non-Euclidean, image-dependent distances. For the minimization of an objective region-based criterion, this representation yields more accurate results with fewer control points than its Euclidean counterpart. If the user positions these control points, the non-Euclidean distance enables to further specify our localized kernels for a target object in the image. Moreover, an intuitive control of the result of the segmentation is obtained by casting inside/outside labels as linear inequality constraints. Finally, we discuss several algorithmic aspects needed for a responsive interactive workflow. We have applied this framework to 3D medical imaging and built a real-time prototype with which the segmentation of whole organs is only a few clicks away. Benoit Mory, Roberto Ardon, Anthony J. Yezzi, Jean-Philippe Thiran |
ICCV | 4 |
| 2009 | Selecting relevant visual features for speechreadingabstractA quantitative measure of relevance is proposed for the task of constructing visual feature sets which are at the same time relevant and compact. A feature's relevance is given by the amount of information that it contains about the problem, while compactness is achieved by preventing the replication of information between features. To achieve these goals, we use mutual information both for assessing relevance and measuring the redundancy between features. Our application is speechreading, that is, speech recognition performed on the video of the speaker. This is justified by the fact that the performance of audio speech recognition can be improved by augmenting the audio features with visual ones, especially when there is noise in the audio channel. We report significant improvements compared to the most common method of dimensionality reduction for speechreading, Linear Discriminant Analysis (LDA). Virginia Estellers, Mihai Gurban, Jean-Philippe Thiran |
ICIP | 3 |
| 2009 | Cooperative Object Segmentation and Behavior Inference in Image Sequences
Laura Gui, Jean-Philippe Thiran, Nikos Paragios |
Int. J. Comput. Vis. | 2 |
| 2009 | Sequential anisotropic multichannel Wiener filtering with Rician bias correction applied to 3D regularization of DWI data
Marcos Martín-Fernández, Emma Muñoz-Moreno, Leila Cammoun, Jean-Philippe Thiran, Carl-Fredrik Westin, Carlos Alberola-López |
Medical Image Anal. | 4 |
| 2009 | Addendum to "Sequential anisotropic multichannel Wiener filtering with Rician bias correction applied to 3D regularization of DWI data" [Medical Image Analysis 13 (2009) 19-35]
Marcos Martín-Fernández, Emma Muñoz-Moreno, Leila Cammoun, Jean-Philippe Thiran, Carl-Fredrik Westin, Carlos Alberola-López |
Medical Image Anal. | 4 |
| 2009 | A Scale-Space of Cortical Feature MapsabstractIn this paper we define a scale-space for cortical mean curvature maps on the sphere, that offers a hierarchical representation of the brain cortical structures, useful in multiscale registration and analysis algorithms. A spherical feature map was obtained through inflation of the cortical surface of one hemisphere, extracted from structural MR images. Using the Beltrami framework, we embedded this spherical mesh in a higher dimensional space and the feature assigned to a mesh vertex became an additional component of its coordinates. This enhanced mesh then evolved under Beltrami flow. Imposing an appropriate aspect ratio for the feature components, we thus minimized an interpolation between theL2and TV-norm of the map. The collection of all maps produced by this PDE formed a scale-space. Our results suggest that this scale-space provides a generalization of the brain map suitable for use e.g., within a multiscale registration framework. Dominique Zosso, Jean-Philippe Thiran |
IEEE Signal Process. Lett. | 2 |
| 2008 | Fast texture segmentation model based on the shape operator and active contourabstractWe present an approach for unsupervised segmentation of natural and textural images based on active contour, differential geometry and information theoretical concept. More precisely, we propose a new texture descriptor which intrinsically defines the geometry of textural regions using the shape operator borrowed from differential geometry. Then, we use the popular Kullback-Leibler distance to define an active contour model which distinguishes the background and textural objects of interest represented by the probability density functions of our new texture descriptor. We prove the existence of a solution to the proposed segmentation model. Finally, a fast and easy to implement texture segmentation algorithm is introduced to extract meaningful objects. We present promising synthetic and real-world results and compare our algorithm to other state-of-the-art techniques. Nawal Houhou, Jean-Philippe Thiran, Xavier Bresson |
CVPR | 2 |
| 2008 | CEC designer: Domain specific modelling for the industrial automation based on the IEC 61499 standardabstractSupport for evolutionary design approaches is not a very much investigated topic in the automation domain. Advanced facilities for semantic validation, refactoring and model transformations, which are necessary for rapid prototyping of control applications, are seldom found in design tools. Implementing such facilities is even more difficult when using the IEC 61499 standard, through some characteristics of the persistence model. This paper presents a specialized development environment for modelling and automatically generating applications for the industrial automation field, its unique facilities for supporting evolutionary design and the suggested improvements to the IEC 61499 model. This environment was developed inside a project for assessing the IEC 61499 principles and for verifying design approaches and execution models on a real plant. Marco Colla, Tiziano Leidi, Murat Kunt, Jean-Philippe Thiran |
ETFA | 4 |
| 2008 | Modelling human perception of static facial expressionsabstractData collected through a recent web-based survey show that the perception (i.e. labeling) of a human facial expression by a human observer is a subjective process, which results in a lack of a unique ground-truth, as intended in the standard classification framework. In this paper we propose the use of discrete choice models (DCM) for human perception of static facial expressions. Random utility functions are defined in order to capture the attractiveness, perceived by the human observer for an expression class, when asked to assign a label to an actual expression image. The utilities represent a natural way for the modeler to formalize her prior knowledge on the process. Starting with a model based on facial action coding systems (FACS), we subsequently defines two other models by adding two new sets of explanatory variables. The model parameters are learned through maximum likelihood estimation and a cross-validation procedure is used for validation purposes. Matteo Sorci, Jean-Philippe Thiran, Javier Cruz, Thomas Robin 0001, Michel Bierlaire |
FG | 2 |
| 2008 | Shape prior based on statistical map for active contour segmentationabstractWe propose a new method for performing active contour segmentation based on the statistical prior knowledge of the object to detect. From a binary training set of objects, a statistical map describes the possible shapes of the object by computing the probability for each point to belong to the object. This statistical map is treated as a prior distribution and an energy functional is defined such that the object reaches the most probable shape knowing the model. The optimization is done in the level-set framework. Results on both synthetic and medical images are shown. Nawal Houhou, Alia Lemkaddem, Valerie Duay, Abdelkarim Allal, Jean-Philippe Thiran |
ICIP | 5 |
| 2008 | Dynamic modality weighting for multi-stream hmms inaudio-visual speech recognitionabstractMerging decisions from different modalities is a crucial problem in Audio-Visual Speech Recognition. To solve this, state synchronous multi-stream HMMs have been proposed for their important advantage of incorporating stream reliability in their fusion scheme. This paper focuses on stream weight adaptation based on modality confidence estimators. We assume different and time-varying environment noise, as can be encountered in realistic applications, and, for this, adaptive methods are best suited. Stream reliability is assessed directly through classifier outputs since they are not specific to either noise type or level. The influence of constraining the weights to sum to one is also discussed. Mihai Gurban, Jean-Philippe Thiran, Thomas Drugman, Thierry Dutoit |
ICMI | 2 |
| 2008 | An Active Contour-Based Atlas Registration Model Applied to Automatic Subthalamic Nucleus Targeting on MRI: Method and Validation
Valerie Duay, Xavier Bresson, Francisco Javier Sánchez Castro, Claudio Pollo, Meritxell Bach Cuadra, Jean-Philippe Thiran |
MICCAI (2) | 6 |
| 2008 | A Surface-Based Approach to Quantify Local Cortical GyrificationabstractThe high complexity of cortical convolutions in humans is very challenging both for engineers to measure and compare it, and for biologists and physicians to understand it. In this paper, we propose a surface-based method for the quantification of cortical gyrification. Our method uses accurate 3-D cortical reconstruction and computes local measurements of gyrification at thousands of points over the whole cortical surface. The potential of our method to identify and localize precisely gyral abnormalities is illustrated by a clinical study on a group of children affected by 22q11 Deletion Syndrome, compared to control individuals. Marie Schaer, Meritxell Bach Cuadra, Lucas Tamarit, François Lazeyras, Stephan Eliez, Jean-Philippe Thiran |
IEEE Trans. Medical Imaging | 6 |
| 2008 | Extraction of Audio Features Specific to Speech Production for Multimodal Speaker DetectionabstractA method that exploits an information theoretic framework to extract optimized audio features using video information is presented. A simple measure of mutual information (MI) between the resulting audio and video features allows the detection of the active speaker among different candidates. This method involves the optimization of an Mi-based objective function. No approximation is needed to solve this optimization problem, neither for the estimation of the probability density functions (pdfs) of the features, nor for the cost function itself. The pdfs are estimated from the samples using a nonparametric approach. The challenging optimization problem is solved using a global method: the differential evolution algorithm. Two information theoretic optimization criteria are compared and their ability to extract audio features specific to speech production is discussed. Using these specific audio features, candidate video features are then classified as member of the "speaker" or "non-speaker" class, resulting in a speaker detection scheme. As a result, our method achieves a speaker detection rate of 100% on in-house test sequences, and of 85% on most commonly used sequences. Patricia Besson, Vlad Popovici, Jean-Marc Vesin, Jean-Philippe Thiran, Murat Kunt |
IEEE Trans. Multim. | 4 |
| 2007 | Joint Object Segmentation and Behavior Classification in Image SequencesabstractIn this paper, we propose a general framework for fusing bottom-up segmentation with top-down object behavior classification over an image sequence. This approach is beneficial for both tasks, since it enables them to cooperate so that knowledge relevant to each can aid in the resolution of the other, thus enhancing the final result. In particular, classification offers dynamic probabilistic priors to guide segmentation, while segmentation supplies its results to classification, ensuring that they are consistent both with prior knowledge and with new image information. We demonstrate the effectiveness of our framework via a particular implementation for a hand gesture recognition application. The prior models are learned from training data using principal components analysis and they adapt dynamically to the content of new images. Our experimental results illustrate the robustness of our joint approach to segmentation and behavior classification in challenging conditions involving occlusions of the target object before a complex background. Laura Gui, Jean-Philippe Thiran, Nikos Paragios |
CVPR | 2 |
| 2007 | Variational Segmentation using Fuzzy Region Competition and Local Non-Parametric Probability Density FunctionsabstractWe describe a novel variational segmentation algorithm designed to split an image in two regions based on their intensity distributions. A functional is proposed to integrate the unknown probability density functions of both regions within the optimization process. The method simultaneously performs segmentation and non-parametric density estimation. It does not make any assumption on the underlying distributions, hence it is flexible and can be applied to a wide range of applications. Although a boundary evolution scheme may be used to minimize the functional, we choose to consider an alternative formulation with a membership function. The latter has the advantage of being convex in each variable, so that the minimization is faster and less sensitive to initial conditions. Finally, to improve the accuracy and the robustness to low-frequency artifacts, we present an extension for the more general case of local space-varying probability densities. The approach readily extends to vectorial images and 3D volumes, and we show several results on synthetic and photographic images, as well as on 3D medical data. Benoit Mory, Roberto Ardon, Jean-Philippe Thiran |
ICCV | 3 |
| 2007 | Relevant Feature Selection for Audio-Visual Speech RecognitionabstractWe present a feature selection method based on information theoretic measures, targeted at multimodal signal processing, showing how we can quantitatively assess the relevance of features from different modalities. We are able to find the features with the highest amount of information relevant for the recognition task, and at the same having minimal redundancy. Our application is audio-visual speech recognition, and in particular selecting relevant visual features. Experimental results show that our method outperforms other feature selection algorithms from the literature by improving recognition accuracy even with a significantly reduced number of features. Thomas Drugman, Mihai Gurban, Jean-Philippe Thiran |
MMSP | 3 |
| 2007 | Face detection with boosted Gaussian features
Julien Meynet, Vlad Popovici, Jean-Philippe Thiran |
Pattern Recognit. | 3 |
| 2007 | A level set method for segmentation of the thalamus and its nuclei in DT-MRI
Lisa Jonasson, Patric Hagmann, Claudio Pollo, Xavier Bresson, Cecilia Richero Wilson, Reto Meuli, Jean-Philippe Thiran |
Signal Process. | 7 |
| 2007 | Scale Space Analysis and Active Contours for Omnidirectional ImagesabstractA new generation of optical devices that generate images covering a larger part of the field of view than conventional cameras, namely catadioptric cameras, is slowly emerging. These omnidirectional images will most probably deeply impact computer vision in the forthcoming years, provided that the necessary algorithmic background stands strong. In this paper, we propose a general framework that helps define various computer vision primitives. We show that geometry, which plays a central role in the formation of omnidirectional images, must be carefully taken into account while performing such simple tasks as smoothing or edge detection. Partial differential equations (PDEs) offer a very versatile tool that is well suited to cope with geometrical constraints. We derive new energy functionals and PDEs for segmenting images obtained from catadioptric cameras and show that they can be implemented robustly using classical finite difference schemes. Various experimental results illustrate the potential of these new methods on both synthetic and natural images. Iva Bogdanova, Xavier Bresson, Jean-Philippe Thiran, Pierre Vandergheynst |
IEEE Trans. Image Process. | 3 |
| 2007 | Representing Diffusion MRI in 5-D Simplifies Regularization and Segmentation of White Matter TractsabstractWe present a new five-dimensional (5-D) space representation of diffusion magnetic resonance imaging (dMRI) of high angular resolution. This 5-D space is basically a non-Euclidean space of position and orientation in which crossing fiber tracts can be clearly disentangled, that cannot be separated in three-dimensional position space. This new representation provides many possibilities for processing and analysis since classical methods for scalar images can be extended to higher dimensions even if the spaces are not Euclidean. In this paper, we show examples of how regularization and segmentation of dMRI is simplified with this new representation. The regularization is used with the purpose of denoising and but also to facilitate the segmentation task by using several scales, each scale representing a different level of resolution. We implement in five dimensions the Chan-Vese method combined with active contours without edges for the segmentation and the total variation functional for the regularization. The purpose of this paper is to explore the possibility of segmenting white matter structures directly as entirely separated bundles in this 5-D space. We will present results from a synthetic model and results on real data of a human brain acquired with diffusion spectrum magnetic resonance imaging (MRI), one of the dMRI of high angular resolution available. These results will lead us to the conclusion that this new high-dimensional representation indeed simplifies the problem of segmentation and regularization. Lisa Jonasson, Xavier Bresson, Jean-Philippe Thiran, Van J. Wedeen, Patric Hagmann |
IEEE Trans. Medical Imaging | 3 |
| 2006 | Discrete Choice Models for Static Facial Expression Recognition
Gianluca Antonini, Matteo Sorci, Michel Bierlaire, Jean-Philippe Thiran |
ACIVS | 4 |
| 2006 | Image Segmentation Model using Active Contour and Image DecompositionabstractThis paper proposes an image segmentation model based on the active contour model, the Mumford-Shah functional and the image decomposition process. Generally speaking, the active contour model detects boundaries in images from sharp intensities variations and the Mumford-Shah model finds smooth regions from homogeneous intensities. Our model merges these two complementary approaches while considering the Four Color Theorem to globally partition any given image. We also consider the textural part lying in natural images by separating it from the geometric part, which contains the meaningful objects, to help the segmentation process. Our segmentation model is experimented with a 1-D signal and 2-D images. Xavier Bresson, Jean-Philippe Thiran |
ICIP | 2 |
| 2006 | Automatic Extraction of Geometric Lip Features with Application to Multi-Modal Speaker IdentificationabstractIn this paper we consider the problem of automatic extraction of the geometric lip features for the purposes of multi-modal speaker identification. The use of visual information from the mouth region can be of great importance for improving the speaker identification system performance in noisy conditions. We propose a novel method for automated lip features extraction that utilizes color space transformation and a fuzzy-based c-means clustering technique. Using the obtained visual cues closed-set audio-visual speaker identification experiments are performed on the CUAVE database, showing promising results Ivana Arsic, Roger Vilagut, Jean-Philippe Thiran |
ICME | 3 |
| 2006 | Behavioral Priors for Detection and Tracking of Pedestrians in Video Sequences
Gianluca Antonini, Santiago Venegas-Martinez, Michel Bierlaire, Jean-Philippe Thiran |
Int. J. Comput. Vis. | 4 |
| 2006 | A Variational Model for Object Segmentation Using Boundary Information and Shape Prior Driven by the Mumford-Shah Functional
Xavier Bresson, Pierre Vandergheynst, Jean-Philippe Thiran |
Int. J. Comput. Vis. | 3 |
| 2006 | Multiscale Active Contours
Xavier Bresson, Pierre Vandergheynst, Jean-Philippe Thiran |
Int. J. Comput. Vis. | 3 |
| 2006 | Counting Pedestrians in Video Sequences Using Trajectory ClusteringabstractIn this paper, we propose the use of lustering methods for automatic counting of pedestrians in video sequences. As input, we consider the output of those detection/tracking systems that overestimate the number of targets. Clustering techniques are applied to the resulting trajectories in order to reduce the bias between the number of tracks and the real number of targets. The main hypothesis is that those trajectories belonging to the same human body are more similar than trajectories belonging to different individuals. Several data representations and different distance/similarity measures are proposed and compared, under a common hierarchical clustering framework, and both quantitative and qualitative results are presented Gianluca Antonini, Jean-Philippe Thiran |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2006 | A Cross Validation Study of Deep Brain Stimulation Targeting: From Experts to Atlas-Based, Segmentation-Based and Automatic Registration AlgorithmsabstractValidation of image registration algorithms is a difficult task and open-ended problem, usually application-dependent. In this paper, we focus on deep brain stimulation (DBS) targeting for the treatment of movement disorders like Parkinson's disease and essential tremor. DBS involves implantation of an electrode deep inside the brain to electrically stimulate specific areas shutting down the disease's symptoms. The subthalamic nucleus (STN) has turned out to be the optimal target for this kind of surgery. Unfortunately, the STN is in general not clearly distinguishable in common medical imaging modalities. Usual techniques to infer its location are the use of anatomical atlases and visible surrounding landmarks. Surgeons have to adjust the electrode intraoperatively using electrophysiological recordings and macrostimulation tests. We constructed a ground truth derived from specific patients whose STNs are clearly visible on magnetic resonance (MR) T2-weighted images. A patient is chosen as atlas both for the right and left sides. Then, by registering each patient with the atlas using different methods, several estimations of the STN location are obtained. Two studies are driven using our proposed validation scheme. First, a comparison between different atlas-based and nonrigid registration algorithms with a evaluation of their performance and usability to locate the STN automatically. Second, a study of which visible surrounding structures influence the STN location. The two studies are cross validated between them and against expert's variability. Using this scheme, we evaluated the expert's ability against the estimation error provided by the tested algorithms and we demonstrated that automatic STN targeting is possible and as accurate as the expert-driven techniques currently used. We also show which structures have to be taken into account to accurately estimate the STN location. Francisco Javier Sánchez Castro, Claudio Pollo, Reto Meuli, Philippe Maeder, Olivier Cuisenaire, Meritxell Bach Cuadra, Jean-Guy Villemure, Jean-Philippe Thiran |
IEEE Trans. Medical Imaging | 8 |
| 2005 | Atlas-based segmentation of medical images locally constrained by level setsabstractAtlas-based segmentation has become a standard paradigm for exploiting prior knowledge in medical image segmentation. In this paper, we propose a method to exploit both the robustness of global registration techniques and the accuracy of a local registration based on level set tracking. First, the atlas is globally put in correspondence with the patient image by an affine and an intensity-based non rigid registration. Based on this rough initialisation, the level set functions corresponding to particular objects of interest of the deformed atlas are used to segment the corresponding objects in the patient image. We propose a technique to derive a dense deformation field from the motion of these level set functions. This is particularly important when we want to infer the position of invisible structures like the brain sub-thalamic nuclei from the position of visible surrounding structures. This can also be advantageously exploited to register an atlas following a hierarchical approach. Results are shown on 2D synthetic images and 2D real images extracted from brain and prostate MR volumes and neck CT volumes. Valerie Duay, Nawal Houhou, Jean-Philippe Thiran |
ICIP (2) | 3 |
| 2005 | Region-based satellite image classification: method and validationabstractWe propose an algorithm for very high-resolution satellite image classification that combines non-supervised segmentation with a supervised classification. Both multi-spectral data and local spatial priors are used in the Gaussian hidden Markov random field (GHMRF) model for the segmentation. Then, two classifiers, Mahalanobis distance classifier and SVM, are studied using intensity, texture and shape features. Validation is done qualitatively and quantitatively by comparison with a manual classification used as a ground truth. Results show very good performance of our approach in comparison to existing techniques. Also, we demonstrate that spectral and spatial features calculated on segmented regions are much more discriminant than the spectral features of the pixels taken individually for the classification task. Xavier Gigandet, Meritxell Bach Cuadra, Abram Pointet, Leila Cammoun, Régis Caloz, Jean-Philippe Thiran |
ICIP (3) | 6 |
| 2005 | Cross Validation of Experts Versus Registration Methods for Target Localization in Deep Brain Stimulation
Francisco Javier Sánchez Castro, Claudio Pollo, Reto Meuli, Philippe Maeder, Meritxell Bach Cuadra, Olivier Cuisenaire, Jean-Guy Villemure, Jean-Philippe Thiran |
MICCAI | 8 |
| 2005 | Monte Carlo video text segmentationabstractThis paper presents a probabilistic algorithm for segmenting and recognizing text embedded in video sequences based on adaptive thresholding using a Bayes filtering method. The algorithm approximates the posterior distribution of segmentation thresholds of video text by a set of weighted samples. The set of samples is initialized by applying a classical segmentation algorithm on the first video frame and further refined by random sampling under a temporal Bayesian framework. This framework allows us to evaluate a text image segmentor on the basis of recognition result instead of visual segmentation result, which is directly relevant to our character recognition task. Results on a database of 6944 images demonstrate the validity of the algorithm. Datong Chen, Jean-Marc Odobez, Jean-Philippe Thiran |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2005 | White matter fiber tract segmentation in DT-MRI using geometric flows
Lisa Jonasson, Xavier Bresson, Patric Hagmann, Olivier Cuisenaire, Reto Meuli, Jean-Philippe Thiran |
Medical Image Anal. | 6 |
| 2005 | Kernel matching pursuit for large datasets
Vlad Popovici, Samy Bengio, Jean-Philippe Thiran |
Pattern Recognit. | 3 |
| 2005 | From error probability to information theoretic (multi-modal) signal processing
Torsten Butz, Jean-Philippe Thiran |
Signal Process. | 2 |
| 2005 | Comparison and validation of tissue modelization and statistical classification methods in T1-weighted MR brain imagesabstractThis paper presents a validation study on statistical nonsupervised brain tissue classification techniques in magnetic resonance (MR) images. Several image models assuming different hypotheses regarding the intensity distribution model, the spatial model and the number of classes are assessed. The methods are tested on simulated data for which the classification ground truth is known. Different noise and intensity nonuniformities are added to simulate real imaging conditions. No enhancement of the image quality is considered either before or during the classification process. This way, the accuracy of the methods and their robustness against image artifacts are tested. Classification is also performed on real data where a quantitative validation compares the methods' results with an estimated ground truth from manual segmentations by experts. Validity of the various classification methods in the labeling of the image as well as in the tissue volume is estimated with different local and global measures. Results demonstrate that methods relying on both intensity and spatial information are more robust to noise and field inhomogeneities. We also demonstrate that partial volume is not perfectly modeled, even though methods that account for mixture classes outperform methods that only consider pure Gaussian classes. Finally, we show that simulated data results can also be extended to real data. Meritxell Bach Cuadra, Leila Cammoun, Torsten Butz, Olivier Cuisenaire, Jean-Philippe Thiran |
IEEE Trans. Medical Imaging | 5 |
| 2004 | Bayesian.integration of a discrete choice pedestrian behavioral model and image correlation techniques for automatic multi object trackingabstractIn this paper we deal with the multiobject tracking problem in the particular case of pedestrians, assuming the detection step already done. We use a Bayesian framework to combine the likelihood term provided by an image correlation algorithm with a prior distribution given by a discrete choice model for pedestrian behavior, calibrated on real data. We aim to show how the combination of the image information with a model of pedestrian behavior can provides appreciable results in real and complex scenarios. Santiago Venegas-Martinez, Gianluca Antonini, Jean-Philippe Thiran, Michel Bierlaire |
ICIP | 3 |
| 2004 | Adaptive Hough transform for the detection of natural shapes under weak affine transformations
Olivier Ecabert, Jean-Philippe Thiran |
Pattern Recognit. Lett. | 2 |
| 2004 | Pattern recognition using higher-order local autocorrelation coefficients
Vlad Popovici, Jean-Philippe Thiran |
Pattern Recognit. Lett. | 2 |
| 2004 | A localization/verification scheme for finding text in images and video frames based on contrast independent features and machine learning methods
Datong Chen, Jean-Marc Odobez, Jean-Philippe Thiran |
Signal Process. Image Commun. | 3 |
| 2004 | Atlas-based segmentation of pathological MR brain images using a model of lesion growthabstractWe propose a method for brain atlas deformation in the presence of large space-occupying tumors, based on an a priori model of lesion growth that assumes radial expansion of the lesion from its starting point. Our approach involves three steps. First, an affine registration brings the atlas and the patient into global correspondence. Then, the seeding of a synthetic tumor into the brain atlas provides a template for the lesion. The last step is the deformation of the seeded atlas, combining a method derived from optical flow principles and a model of lesion growth. Results show that a good registration is performed and that the method can be applied to automatic segmentation of structures and substructures in brains with gross deformation, with important medical applications in neurosurgery, radiosurgery, and radiotherapy. Meritxell Bach Cuadra, Claudio Pollo, Anton Bardera, Olivier Cuisenaire, Jean-Guy Villemure, Jean-Philippe Thiran |
IEEE Trans. Medical Imaging | 6 |
| 2003 | A priori information in image segmentation: energy functional based on shape statistical model and image informationabstractIn this paper, we propose an energy functional to segment objects whose global shape is a priori known thanks to a statistical model. Our work aims at extending the variational approach of Chen et al. [Y. Chen, et al., 2002] by integrating the statistical shape model of Leventon et al. [M. Leventon, et al., 2000]. The proposed energy functional allows us to capture an object that exhibits high image gradients and a shape compatible with the statistical model which best fits the segmented object. The minimization of the functional provides a system of coupled equations whose steady-state solution is the solution of the segmentation problem. Results are presented on synthetic and medical images. Xavier Bresson, Pierre Vandergheynst, Jean-Philippe Thiran |
ICIP (3) | 3 |
| 2003 | Atlas-based segmentation of pathological brain MR imagesabstractA method for brain atlas deformation in presence of large space-occupying tumors, based on an a priori model of lesion growth that assumes radial expansion of the lesion from its starting point is proposed. First, an affine registration brings the atlas and the patient into global correspondence. Then, the seeding of a synthetic tumor into the brain atlas provides a template for the lesion. Finally, the seeded atlas is deformed, combining a method derived from optical flow principles and a model of lesion growth (MLG). Results show that the method can be applied to the automatic segmentation of structures and substructures in brains with gross deformation, with important medical applications in neurosurgery, radiosurgery and radiotherapy. Meritxell Bach Cuadra, Claudio Pollo, Anton Bardera, Olivier Cuisenaire, Jean-Guy Villemure, Jean-Philippe Thiran |
ICIP (1) | 6 |
| 2003 | A New Brain Segmentation Framework
Torsten Butz, Patric Hagmann, Eric Tardif, Reto Meuli, Jean-Philippe Thiran |
MICCAI (2) | 5 |
| 2003 | Brain Shift Correction Based on a Boundary Element Biomechanical Model with Different Material Properties
Olivier Ecabert, Torsten Butz, Arya Nabavi, Jean-Philippe Thiran |
MICCAI (1) | 4 |
| 2003 | In Apologiam - rules of the game and plagiarism
Philippe Salembier, Jean-Philippe Thiran, Jean-Marc Vesin, Pierre Vandergheynst, Murat Kunt |
Signal Process. | 2 |
| 2003 | 3D Encoding/2D Decoding of Medical DataabstractWe propose a fully three-dimensional (3-D) wavelet-based coding system featuring 3-D encoding/two-dimensional (2-D) decoding functionalities. A fully 3-D transform is combined with context adaptive arithmetic coding; 2-D decoding is enabled by encoding every 2-D subband image independently. The system allows a finely graded up to lossless quality scalability on any 2-D image of the dataset. Fast access to 2-D images is obtained by decoding only the corresponding information thus avoiding the reconstruction of the entire volume. The performance has been evaluated on a set of volumetric data and compared to that provided by other 3-D as well as 2-D coding systems. Results show a substantial improvement in coding efficiency (up to 33%) on volumes featuring good correlation properties along the z axis. Even though we did not address the complexity issue, we expect a decoding time of the order of one second/image after optimization. In summary, the proposed 3-D/2-D multidimensional layered zero coding system provides the improvement in compression efficiency attainable with 3-D systems without sacrificing the effectiveness in accessing the single images characteristic of 2-D ones. Gloria Menegaz, Jean-Philippe Thiran |
IEEE Trans. Medical Imaging | 2 |
| 2002 | Feature space mutual information in speech-video sequencesabstractKeywords: LTS5 Reference EPFL-CONF-86903View record in Web of Science Record created on 2006-06-14, modified on 2017-05-10 Torsten Butz, Jean-Philippe Thiran |
ICME (2) | 2 |
| 2002 | Atlas-Based Segmentation of Pathological Brains Using a Model of Tumor Growth
Meritxell Bach Cuadra, Patric Hagmann, Claudio Pollo, Jean-Guy Villemure, Benoit M. Dawant, Jean-Philippe Thiran |
MICCAI (1) | 7 |
| 2002 | Validation of Tissue Modelization and Classification Techniques in T1-Weighted MR Brain Images
Meritxell Bach Cuadra, Bram Platel, Eduardo Solanas, Torsten Butz, Jean-Philippe Thiran |
MICCAI (1) | 5 |
| 2002 | Lossy to lossless object-based coding of 3-D MRI dataabstractWe propose a fully three-dimensional (3-D) object-based coding system exploiting the diagnostic relevance of the different regions of the volumetric data for rate allocation. The data are first decorrelated via a 3-D discrete wavelet transform. The implementation via the lifting steps scheme allows to map integer-to-integer values, enabling lossless coding, and facilitates the definition of the object-based inverse transform. The coding process assigns disjoint segments of the bitstream to the different objects, which can be independently accessed and reconstructed at any up-to-lossless quality. Two fully 3-D coding strategies are considered: embedded zerotree coding (EZW-3D) and multidimensional layered zero coding (MLZC), both generalized for region of interest (ROI)-based processing. In order to avoid artifacts along region boundaries, some extra coefficients must be encoded for each object. This gives rise to an overheading of the bitstream with respect to the case where the volume is encoded as a whole. The amount of such extra information depends on both the filter length and the decomposition depth. The system is characterized on a set of head magnetic resonance images. Results show that MLZC and EZW-3D have competitive performances. In particular, the best MLZC mode outperforms the others state-of-the-art techniques on one of the datasets for which results are available in the literature. Gloria Menegaz, Jean-Philippe Thiran |
IEEE Trans. Image Process. | 2 |
| 2001 | Text Identification in Complex Background Using SVMabstractThe paper presents a fast and robust algorithm to identify text in image or video frames with complex backgrounds and compression effects. The algorithm first extracts the candidate text line on the basis of edge analysis, baseline location and heuristic constraints. Support Vector Machine (SVM) is then used to identify text line from the candidates in edge-based distance map feature space. Experiments based on a large amount of images and video frames from different sources showed the advantages of this algorithm compared to conventional methods in both identification quality and computation time. Datong Chen, Hervé Bourlard, Jean-Philippe Thiran |
CVPR (2) | 3 |
| 2001 | Shot boundary detection with mutual informationabstractWe present a novel approach for shot boundary detection that uses mutual information (MI) and affine image registration. The MI measures the statistical difference between consecutive frames, while the applied affine registration compensates for camera panning and zooming. Results for different sequences are presented to illustrate the motion and zoom compensation and the robustness of MI to illumination changes. Furthermore we show that the affine registration has no effect at the shot boundaries themselves and therefore doesn't corrupt the result. Because the presented algorithm analyses the frames sequentially, its parallelization is straightforward even for distributed memory architectures. We quantify the speed-up on a LINUX cluster and show that the communicational load of the implementation is almost negligible, resulting in an almost linear speed-up. Torsten Butz, Jean-Philippe Thiran |
ICIP (3) | 2 |
| 2001 | Automatic segmentation of internal structures of the brain in MR images using a tandem of affine and non-rigid registration of an anatomical brain atlasabstractIn the study of many neurological pathologies, the accurate quantization of the white matter (WM) and gray matter (GM) volumes of the brain is essential Moreover, regional volume calculations may bring even more useful diagnostic information. We present therefore the segmentation of internal structures of the brain for further regional WM and GM volume quantization. A priori information about the brain anatomy is included in the segmentation process by the registration of the patient MR images with a computerized brain atlas. We propose the combination of a global affine transformation used to initialize key boundary surfaces (lateral ventricles and cortical surfaces) of both images with a local free-form transformation based on an optical flow algorithm. We apply this technique to segment the cerebellum and the cerebral trunk in order to exclude them from our WM and GM volume quantization. Validation has been conducted on a large number of images, showing excellent results. Meritxell Bach Cuadra, Olivier Cuisenaire, Reto Meuli, Jean-Philippe Thiran |
ICIP (3) | 4 |
| 2001 | Higher order autocorrelations for pattern classificationabstractThe use of higher-order local autocorrelations as features for pattern recognition has been acknowledged for many years, but their applicability was restricted to relatively low orders (2 or 3) and small local neighborhoods, due to combinatorial increase in computational costs. A new method for using these features is presented, which allows the use of autocorrelations of any order and of larger neighborhoods. The method is closely related to the classifier used, a support vector machine (SVM), and exploits the special form of the inner products of autocorrelations and the properties of some kernel functions used by SVM. Using SVM, linear and nonlinear classification functions can be learned, extending the previous works on higher-order autocorrelations which were based on linear classifiers. Vlad Popovici, Jean-Philippe Thiran |
ICIP (3) | 2 |
| 2001 | Relative anatomical location for statistical non-parametric brain tissue classification in MR imagesabstractWe propose a statistical nonparametric classification of brain tissues from an MR image based on the voxel intensities and on the relative anatomical location of the different tissues. We generate an artificial image component as the distance from the edges of the segmented brain. The nonparametric k-nearest neighbors rule (k-NN) is used since it requires no a priori information on the probability distribution of this distance component. The k-NN rule is also tested using different metrics (Euclidean, weighted Euclidean, Mahalanobis) in the classification space to define what "nearest neighbors" are. The results are twofold: firstly we show that all metrics perform well in ideal conditions, but that the Mahalanobis (and to some extent the weighted Euclidean) metric is more robust in the case of under-training of the classifier. Secondly we show that using the relative anatomical location in combination with the intensity information improves the classification of the tissues. Eduardo Solanas, Valerie Duay, Olivier Cuisenaire, Jean-Philippe Thiran |
ICIP (2) | 4 |
| 2001 | Affine Registration with Feature Space Mutual Information
Torsten Butz, Jean-Philippe Thiran |
MICCAI | 2 |
| 2001 | Surface Based Atlas Matching of the Brain Using Deformable Surfaces and Volumetric Finite Elements
Matthieu Ferrant, Olivier Cuisenaire, Benoît Macq, Jean-Philippe Thiran, Martha Elizabeth Shenton, Ron Kikinis, Simon K. Warfield |
MICCAI | 4 |
| 2001 | Exploiting Voxel Correlation for Automated MRI Bias Field Correction by Conditional Entropy Minimization
Eduardo Solanas, Jean-Philippe Thiran |
MICCAI | 2 |
| 2000 | Multirate Coding of 3D Medical DataabstractThe last generation medical imaging equipment produce multidimensional (3D or 3D+time) data distributions. On a coding perspective, it is reasonable to expect that the exploitation of the full dimensional correlation among data samples would lead to a sensible improvement in compression performances, especially for isotropic datasets. We propose a fully three-dimensional wavelet-based coding system providing a finely-graded up to lossless data representation in a single bistream. The data are first decorrelated by a 3D discrete wavelet transform, performed by the non-linear lifting scheme mapping integers to integers. This enables the lossless mode and permits the in-place implementation of the transform at a reduced computational complexity. The coding scheme is inspired to the layered-zero coding proposed by Taubman and Zakhor (1994), extended to handle fully 3D subband structures. Performances are characterized with respect to both the 2D version of the same algorithm and the JPEG standard. The rate-saving is strongly influenced by the amount of the data correlation in the z dimension, ranging between 16.5% and 5.5% for the considered datasets. Gloria Menegaz, Laurent Grewe, Jean-Philippe Thiran |
ICIP | 3 |
| 1999 | Object-Based Coding of Volumetric Medical Data
Gloria Menegaz, Vincent Vaerman, Jean-Philippe Thiran |
ICIP (3) | 3 |
| 1999 | A Parametric Hybrid Model Used for Multidimensional Object RepresentationabstractIn this paper, we present a parametric hybrid model used in the framework of multidimensional object representation, for applications to both object visualization and object-based data compression. Our model is defined as a set of hybrid ellipsoids suitable for both globally and locally deforming the reconstructed shape. Its new parameterization, as compared to classical techniques, allows us to preserve its analytical representation during the fitting process. It is fitted to the object contours by means of a genetic algorithm minimizing a mean-square error criterion. Several criteria are proposed and discussed according to the stability of the optimization process, as well as the ability to efficiently initialize the model parameters. Finally, fitting results are presented for 2D and 3D data and different applications are proposed. Vincent Vaerman, Gloria Menegaz, Jean-Philippe Thiran |
ICIP (1) | 3 |
| 1998 | Interactive DICOM image transmission and telediagnosis over the European ATM networkabstractThe European High-Performance Information Infrastructure in Medicine, n(o)B3014 (HIM3) project of the Trans-European Network--Integrated Broadband Communications (TEN-IBC) program, started on March 1996 and finished on February 1997, aimed to test the medical usability of the European asynchronous transfer mode (ATM) network in medical image transmission. The Department of Radiology, University of Pisa, Pisa, Italy, and St-Luc University Hospital, Brussels, Belgium, involved in the project as healthcare partners in the radiological domain, established several connection sessions finalized to test the usability of Digital Imaging and Communication (DICOM) image transmission and interactive telediagnosis tools in the daily radiological practice. The Pisa site was connected to the Italian ATM pilot (Sirius Network) through the Tuscany metropolitan area network (MAN), while St-Luc University Hospital was connected to Belgium ATM network through the Brussels MAN. By means of international connections provided by the European JAMES project, a link between the two sites was established, connecting both national ATM networks. Due to the large variety of hardware present in the medical centers, multiplatform software tools were used and tested: central test node (CTN) release 2.8 [3], VAT [6], NV-3.3 [7], and IDI (UCL homemade multiplatform teleradiology tool for interactive visualization and processing of DICOM images). During the telediagnosis session, lead by radiologists in both hospitals, each site submitted neuroradiological clinical cases to the other for remote consultation. The connection, available for a period of two weeks, at 2-Mbit/s bandwidth, allowed the transmission of MR images (256 x 256 x 12 bit) and simultaneous multimedia interactive discussion of the cases. Both off-line transmission and review of the images, using the CTN DICOM transfer routines, and on-line interactive image discussion, using the IDI telediagnosis software, were tested successfully from the technical and medical point of view. Emanuele Neri, Jean-Philippe Thiran, Davide Caramella, Claudio Petri, Carlo Bartolozzi, Bruno Piscaglia, Benoît Macq, Thierry Duprez, Guy Cosnard, Baudouin Maldague, Johan De Pauw |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 1997 | A queue-based region growing algorithm for accurate segmentation of multi-dimensional digital images
Jean-Philippe Thiran, Vincent Warscotte, Benoît Macq |
Signal Process. | 1 |
| 1996 | Morphological registration of 3D medical imagesabstractRegistration of three-dimensional medical image modalities (PET and MRI) of the human brain is performed by a multi-scale morphological shape description combined with 3D Chamfer matching. Objects of Interest are first segmented. For MR images of the head, a Directional Watershed algorithm is used to obtain an accurate segmentation of the brain. PET images are segmented by thresholding combined with region growing. The shape description is then performed by generalized morphological skeletons, introduced in this article, providing a multi-scale representation of the object shape. Registration is finally operated by Chamfer matching, using a rigid transform for PET-MRI registration. Non-rigid registration is also evoked. Jean-Philippe Thiran, Benoît Macq, Christian Michel |
ICIP (2) | 1 |
| 1996 | IMIS: A multi-platform software package for telediagnosis and 3D medical image processingabstractWe present a project developed in our University in order to provide Medical Imaging Departments with efficient software tools for 3D medical image processing and transmission. In this context, the IMIS software package has been developed, combining, in a modular programming strategy, an easily upgradable graphical user interface, using the Tcl/Tk toolkit, with high performance image processing techniques, such as 3D lossless multiresolution image compression for telediagnosis on narrow-band ISDN networks. IMIS is based on a set of tools allowing the handling of 3D medical images such as magnetic resonance images (MRI), positron emission tomography (PET), and computed tomography (CT) of functional MRI (fMRI). Jean-Philippe Thiran, Bruno Piscaglia, Patrick Piscaglia, Benoît Macq, Jean-Francois Goudemant, Roger Demeure |
ICIP (2) | 1 |
| 1994 | Morphological Classification of Cancerous CellsabstractWe present a new method for the automatic recognition of cancerous cells from a digitized picture of a microscopic section. The method is based on the analysis of four criteria of malignancy, in relation with the shape and the size of the observed cells. It provides the physician with nonsubjective numerical values for the four concerned criteria of malignancy, in order to help him to decide whether the tissue is cancerous or not. The automatic approach described uses mathematical morphology, first to remove the background noise from the image and to operate a segmentation of the nuclei of the cells, which contain most of the features of malignancy, next to analyse the shape and the size of these nuclei and lastly to evaluate their texture. From the values of the four extracted criteria, an automatic classification of the image is using a Kohonen neural network.> Jean-Philippe Thiran, Benoît Macq, Jacques Mairesse |
ICIP (3) | 1 |