EDBT 2026 Demo / reviewers in the wild / expert
Camille Kurtz
dblp:98/9838
· DBLP profile ↗
45ranked-venue papers
10as first author
20since 2021 · last 2026
0000-0001-9254-7537ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Augmenting medical vision-language pretraining via domain-aware retrieval-augmented caption enrichmentabstractContrastive Language-Image Pre-training (CLIP) has demonstrated strong potential in learning transferable visual models / representations by aligning paired images and textual descriptions. Nevertheless, the quality of training data remains a significant bottleneck. In many real-world scenarios, image-text pairs are often noisy or accompanied by captions that are either too short or too generic to convey key visual attributes. For example, in medical imaging, most available data come from illustrative figures in public literature instead of detailed clinical reports, resulting in captions that lack the precision and context provided by expert annotations. Recent efforts to improve caption quality using large language models have largely focused on natural images and overlook the integration of domain-specific knowledge. In this study, we propose a retrieval-augmented generation framework guided by expert semantic knowledge to enrich image captions in the medical context. We further introduce a multi-text training strategy that effectively incorporates these enriched descriptions into CLIP training. Our approach, demonstrated in the medical domain as a proof of concept, achieves state-of-the-art performances on multiple downstream tasks, highlighting its broader potential for vision-language pretraining in specialized domains. Our code is available at: https://github.com/Wxy-24/RaceCLIP. Xiaoyang Wei 0004, Camille Kurtz, Florence Cloppet |
ICMR | 2 |
| 2026 | Recognition of action phases from spatio-temporal intermediate representations in amateur sports competition videosabstractThe video broadcasting of sports events leads daily to a mass of data. In this work, we are interested in the recognition of action phases (versus non-action, e.g. timeout) in videos of amateur sports competitions. The automation of this task can be related to the classification of actions from videos where the state of the art relies on the use of 3D deep neural approaches. Although providing good performances in the context of professional sports and offline analysis, deploying these approaches in an amateur/live processing setting can be problematic. In this study, we investigate the interest of an approach relying on intermediate local planar representations capturing both spatial and temporal information from the input videos. These representations are compatible with lightweight 2D neural models that are optimized for video classification. We evaluate the performances of the proposed approach combined with different visual backbones such as traditional convolutional neural networks (e.g., ResNet, MobileNet) as well as Transformer-based architectures and we compare our results with different baseline methods (2D and 3D models), on a dataset of 1000 annotated videos from amateur football competitions. Furthermore, we demonstrate the generalization capabilities of our method by evaluating it on a different sport (rugby) in a zero-shot setting. The results obtained highlight the interest of the proposed approach for sport competition videos, with acceptable performance and computation times allowing scaling in amateur/live processing settings. Lazhar Bouacha, Thierry Magnien, Laurent Wendling, Camille Kurtz |
Comput. Vis. Image Underst. | 4 |
| 2026 | SpIRL: Spatially-aware image representation learning under the supervision of relative position descriptorsabstractExtracting good visual representations from image contents is essential for solving many computer vision problems (e.g. image retrieval , object detection, classification). In this context, state-of-the-art approaches are mainly based on learning a representation using a neural network optimized for a given task. The encoders optimized in this way can then be deployed as backbones for various downstream tasks. When the latter involves reasoning about spatial information from the image content (e.g. retrieve similar structured scenes or compare spatial configurations), this may be suboptimal since models like convolutional neural networks struggle to reason about the relative position of objects in images. Previous studies on building hand-crafted spatial representations , thanks to Relative Position Descriptors (RPD), showed they were powerful to discriminate spatial relations between crisp objects, but such spatial descriptors have rarely been integrated into deep neural networks . We propose in this article different strategies embedded in a common framework called SpIRL (SPatially-aware Image Representation Learning) to guide the optimization of encoders to make them learn more spatial information, under the supervision of an RPD and with the help of a novel dataset (44k images) that does not induce learning semantic information. By using these strategies, we aim to help encoders build more spatially-aware representations. Our experimental results showcase that encoders trained under the SpIRL framework can capture accurate information about the spatial configurations of objects in images on two selected downstream tasks and public datasets. Logan Servant, Michaël Clément, Laurent Wendling, Camille Kurtz |
Pattern Recognit. | 4 |
| 2025 | Automating Geospatial Vision Tasks with a Large Language Model Agent
Camille Kurtz, Sylvain Lobry |
ECML/PKDD (8) | 3 |
| 2025 | Contrastive Learning of Image Representations Guided by Spatial Relations
Logan Servant, Michaël Clément, Laurent Wendling, Camille Kurtz |
WACV | 4 |
| 2025 | Relaxing Binary Constraints in Contrastive Vision-Language Medical Representation LearningabstractBy aligning paired image and caption embeddings as input, contrastive vision-language representation learning has witnessed significant advances as illustrated by CLIP, allowing visual encoders to learn from textual supervision and vice versa. Benefiting from millions of image-caption pairs collected from the Internet, CLIP-like models show competitive performances against fully supervised baselines. However, the learned visual representations are still undermined due to the binary constraint, as most contrastive learning frameworks follow strict one-to-one corre-spondence for the input pairs of data and optimize the models using the InfoNCE loss function. The embeddings of the paired image-text are aligned while the unpaired image-text are pushed away from each other. In fact, there are natu-rally many “false negatives” among these negative pairs since unpaired data can also have a high similarity. In this work, we aim to overcome the impact offalse negatives in vision-language representation learning by introducing soft targets for estimating the similarity between unpaired images and texts using external semantic knowledge structured in the form of graphs. The interest of such a method is demonstrated in the application context of medical imaging. Xiaoyang Wei 0004, Camille Kurtz, Florence Cloppet |
WACV | 2 |
| 2025 | Enhancing vision-language contrastive representation learning using domain knowledgeabstractVisual representation learning plays a key role in solving medical computer vision tasks. Recent advances in the literature often rely on vision-language models aiming to learn the representation of medical images from the supervision of paired captions in a label-free manner. The training of such models is however very data/time intensive and the alignment strategies involved in the contrastive loss functions may not capture the full richness of information carried by inter-data relationships. We assume here that considering expert knowledge from the medical domain can provide solutions to these problems during model optimization. To this end, we propose a novel knowledge-augmented vision-language contrastive representation learning framework consisting of the following steps: (1) Modeling the hierarchical relationships between various medical concepts using expert knowledge and medical images in a dataset through a knowledge graph, followed by translating each node into a knowledge embedding; And (2) integrating knowledge embeddings into a vision-language contrastive learning framework, either by introducing an additional alignment loss between visual and knowledge embeddings or by relaxing binary constraints of vision-language alignment using knowledge embeddings. Our results demonstrate that the proposed solution achieves competitive performances against state-of-the-art approaches for downstream tasks while requiring significantly less training data. Our code is available at https://github.com/Wxy-24/KL-CVR . Xiaoyang Wei 0004, Camille Kurtz, Florence Cloppet |
Comput. Vis. Image Underst. | 2 |
| 2025 | Semantic aware representation learning for optimizing image retrieval systems in radiologyabstractContent-based image retrieval (CBIR), which consists of ranking a set of images with respect to a query image based on visual similarity, can assist diagnostic radiologists in assessing medical images, by identifying similar digital images in large image databases. Despite the many recent advances and innovations in CBIR for general images, their adoption in radiology has been slow and limited. In the current paper we attempt to close the gap between the two domains and wisely adapt modern CBIR techniques to radiology images: by extending the latest representation learning techniques in a way that can overcome the unique challenges and at the same time take advantage of the specific opportunities that are present in radiology we were able to come up with novel and effective medical image retrieval methods. Our method achieves the highest CUI@5 scores (18.48, 15.95) on two widely used datasets (ROCO and MEDICAT respectively), showcasing the superiority of the proposed method in comparison with state-of-the-art relevant alternatives. • Metadata about image semantics is beneficial for representation learning. • External medical knowledge can be integrated for better semantics understanding. • Severity of mistakes can to be taken into account to improve image retrieval. • Differentiable average precision-based loss is suitable for representation learning. Zografoula Vagena, Xiaoyang Wei 0004, Camille Kurtz, Florence Cloppet |
Pattern Recognit. | 3 |
| 2024 | New Algorithms for Multivalued Component Trees
Nicolas Passat, Romain Perrin, Jimmy Francky Randrianasoa, Camille Kurtz, Benoît Naegel |
ICPR (23) | 4 |
| 2024 | Segmentation-Guided Attention for Visual Question Answering from Remote Sensing ImagesabstractVisual Question Answering for Remote Sensing (RSVQA) is a task that aims at answering natural language questions about the content of a remote sensing image. The visual features extraction is therefore an essential step in a VQA pipeline. By incorporating attention mechanisms into this process, models gain the ability to focus selectively on salient regions of the image, prioritizing the most relevant visual information for a given question. In this work, we propose to embed an attention mechanism guided by segmentation into a RSVQA pipeline. We argue that segmentation plays a crucial role in guiding attention by providing a contextual understanding of the visual information, underlying specific objects or areas of interest. To evaluate this methodology, we provide a new VQA dataset that exploits very high-resolution RGB orthophotos annotated with 16 segmentation classes and question/answer pairs. Our study shows promising results of our new methodology, gaining almost 10% of overall accuracy compared to a classical method on the proposed dataset. Lucrezia Tosato, Hichem Boussaid, Flora Weissgerber, Camille Kurtz, Laurent Wendling, Sylvain Lobry |
IGARSS | 4 |
| 2022 | Fine-Tune Your Classifier: Finding Correlations with TemperatureabstractTemperature is a widely used hyperparameter in various tasks involving neural networks, such as classification or metric learning, whose choice can have a direct impact on the model performance. Most of existing works select its value using hyperparameter optimization methods requiring several runs to find the optimal value. We propose to analyze the impact of temperature on classification tasks by describing a dataset as a set of statistics computed on representations on which we can build a heuristic giving us a default value of temperature. We study the correlation between these extracted statistics and the observed optimal temperatures. This preliminary study on more than a hundred combinations of different datasets and features extractors highlights promising results towards the construction of a general heuristic for temperature. Benjamin Chamand, Olivier Risser-Maroix, Camille Kurtz, Philippe Joly, Nicolas Loménie |
ICIP | 3 |
| 2022 | Embedding Spatial Relations in Visual Question Answering for Remote SensingabstractRemote sensing images carry a wealth of information that is not easily accessible to end-users as it requires strong technical skills and knowledge. Visual Question Answering (VQA), a task that aims at answering an open-ended question in natural language from an image, can provide an easier access to this information. Considering the geographical information contained in remote sensing images, questions often embed an important spatial aspect, for instance regarding the relative position of two objects. Our objective is to better model the spatial relations in the construction of a ground-truth database of image/question/answer triplets and to assess the capacity a VQA model has to answer these questions. In this article, we propose to use histograms of forces to model the directional spatial relations between geo-localized objects. This allows a finer modeling of ambiguous relationships between objects and to provide different levels of assessment of a relation (e.g. object A is slightly/strictly to the west of object B). Using this new dataset, we evaluate the performances of a classical VQA model and propose a curriculum learning strategy to better take into account the varying difficulty of questions embedding spatial relations. With this approach, we show an improvement in the performances of our model, highlighting the interest of embedding spatial relations in VQA for remote sensing applications. Maxime Faure, Sylvain Lobry, Camille Kurtz, Laurent Wendling |
ICPR | 3 |
| 2022 | Text-guided visual representation learning for medical image retrieval systemsabstractRadiologists are now confronted with the difficulty of interpreting cross-sectional studies composed of thousands of images. In hospitals, all acquired imaging data are stored in a picture archiving and communication system (PACS). To take advantage of these masses of previously interpreted images, with the ultimate goal to facilitate the diagnosis of new cases, a promising approach would be to integrate a Content-Based Image Retrieval (CBIR) system into PACS. CBIR system performances are inherently limited by the features considered to represent the images. The current state of the art for the extraction of visual features relies on deep learning which requires a sufficient amount of annotated data to learn a generalizable model, such annotated data being rare and difficult to use in medical imaging. At the same time, PACS contain additional information such as radiological reports which supplement visual information carried by the images. We study here how such semantic information, hidden in these reports, can be used to supervise the learning of neuronal models to build a better visual representation of images. In this context our contribution is threefold. We first adapted a contrastive learning approach, which is usually used to learn representation from pairs of positive images in an unsupervised manner, to deal with in-domain medical data. Second, to train such a model to be robust and generalizable with a sufficient amount of data, we propose to re-employ the "dormant" medical imaging literature. Finally, the visual features and the deep models learned in this way, can be considered in CBIR systems as coarse-grained information which can then be fine-tuned in PACS, with more specific images depending on the applications. The obtained experimental result with state of the art contrastive learning methods highlight the interest of this approach. Guillaume Sérieys, Camille Kurtz, Laure Fournier, Florence Cloppet |
ICPR | 2 |
| 2022 | Description and recognition of complex spatial configurations of object pairs with Force Banner 2D features
Robin Deléarde, Camille Kurtz, Laurent Wendling |
Pattern Recognit. | 2 |
| 2021 | Violence Detection from Video under 2D Spatio-Temporal RepresentationsabstractAction recognition in videos, especially for violence detection, is now a hot topic in computer vision. The interest of this task is related to the multiplication of videos from surveillance cameras or live television content producing complex $2D+t$ data. State-of-the-art methods rely on end-to-end learning from 3D neural network approaches that should be trained with a large amount of data to obtain discriminating features. To face these limitations, we present in this article a method to classify videos for violence recognition purpose, by using a classical 2D convolutional neural network (CNN). The strategy of the method is two-fold: (1) we start by building several 2D spatio-temporal representations from an input video, (2) the new representations are considered to feed the CNN to the train/test process. The classification decision of the video is carried out by aggregating the individual decisions from its different 2D spatio-temporal representations. An experimental study on public datasets containing violent videos highlights the interest of the presented method. Mohamed Chelali, Camille Kurtz, Nicole Vincent |
ICIP | 2 |
| 2021 | Learning an Adaptation Function to Assess Image Visual SimilaritiesabstractHuman perception is routinely assessing the similarity between images, both for decision making and creative thinking. But the underlying cognitive process is not really well understood yet, hence difficult to be mimicked by computer vision systems. State-of-the-art approaches using deep architectures are often based on the comparison of images described as feature vectors learned for image categorization task. As a consequence, such features are powerful to compare semantically related images but not really efficient to compare images visually similar but semantically unrelated. Inspired by previous works on neural features adaptation to psycho-cognitive representations, we focus here on the specific task of learning visual image similarities when analogy matters. We propose to use different layers of a categorization-based CNN (pre-trained on ImageNet) as a rough approximation of the visual cortex and learn only an adaptation function corresponding to the approximation of the the primate IT cortex through the metric learning framework. Our experiments conducted on the Totally Looks Like image dataset highlight the interest of our method, by increasing the retrieval scores @5 by 1.75×. Olivier Risser-Maroix, Camille Kurtz, Nicolas Loménie |
ICIP | 2 |
| 2021 | Combination of visual and semantic criteria for automated selection of region proposals in a bounding boxabstractDeep learning based techniques have been widely used for semantic segmentation. The underlying voluminous DNN models are trained on large datasets that have been annotated at the pixel level by humans. Such low-level annotation tasks are expensive to obtain for newly collected datasets. Alternatively, we propose ComViSe, a segmentation pipeline that requires only high-level annotations that remain relatively accessible (e.g., bounding boxes and labels of a detection, labels of a legend) to segment a given image. ComViSe embeds a segmentation framework, pre-trained on a semantically different dataset, to generate image region proposals. The pipeline relies then on several semantic, visual and geometric criteria to characterize each proposed region, and combines them to select the optimal segmentation mask, comparing diverse aggregation strategies from handcrafted formula to automatic ones, supervised or not. An experimental study conducted on the PASCAL VOC dataset shows that these effectively combined criteria are enough to select the mask proposals with the best IoU score in most cases, and that the aggregation can be done automatically. Mohamed-Hicham Leghettas, Robin Deléarde, Camille Kurtz, Laurent Wendling |
ICMV | 3 |
| 2021 | Deep-STaR: Classification of image time series based on spatio-temporal representations
Mohamed Chelali, Camille Kurtz, Anne Puissant, Nicole Vincent |
Comput. Vis. Image Underst. | 2 |
| 2021 | Influence of Data Representations and Deep Architectures in Image Time Series ClassificationabstractImage time series, such as Satellite Image Time Series (SITS) or MRI functional sequences in the medical domain, carry both spatial and temporal information. In many pattern recognition applications such as image classification, taking into account such rich information may be crucial and discriminative during the decision making stage. However, the extraction of spatio-temporal features from image time series is difficult to handle due to the complex representation of the data cube. In this paper, we present a strategy based on Random Walk to build a novel segment-based representation of the data, passing from a 2D[Formula: see text] dimension to a 2D one, more easily manipulable and without losing too much spatial information. Such new representation is then used to feed a classical Convolutional Neural Network (CNN) in order to learn spatio-temporal features with only 2D convolutions and to classify image time series data for a particular classification problem. The influence of the way the 2D[Formula: see text] data are represented, as well as the impact of the network architectures on the results, are carefully studied. The interest of this approach is highlighted on a remote sensing application for the classification of complex agricultural crops. Mohamed Chelali, Camille Kurtz, Anne Puissant, Nicole Vincent |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2021 | Supervised quality evaluation of binary partition trees for object segmentation
Jimmy Francky Randrianasoa, Pierre Cettour-Janet, Camille Kurtz, Eric Desjardin, Pierre Gançarski, Nathalie Bednarek, François Rousseau 0002, Nicolas Passat |
Pattern Recognit. | 3 |
| 2020 | Classification of spatially enriched pixel time series with convolutional neural networksabstractSatellite Image Time Series (SITS), MRI sequences, and more generally image time series, constitute 2D+t data providing spatial and temporal information about an observed scene. Given a pattern recognition task such as image classification, considering jointly such rich information is crucial during the decision process. Nevertheless, due to the complex representation of the data-cube, spatio-temporal features extraction from 2D+t data remains difficult to handle. We present in this article an approach to learn such features from this data, and then to proceed to their classification. Our strategy consists in enriching pixel time series with spatial information. It is based on Random Walk to build a novel segment-based representation of the data, passing from a 2D+t dimension to a 2D one, without loosing too much spatial information. Such new representation is then involved in an end-to-end learning process with a classical 2D Convolutional Neural Network (CNN) in order to learn spatiotemporal features for the classification of image time series. Our approach is evaluated on a remote sensing application for the mapping of agricultural crops. Thanks to a visual attention mechanism, the proposed 2D spatio-temporal representation makes also easier the interpretation of a SITS to understand spatiotemporal phenomenons related to soil management practices. Mohamed Chelali, Camille Kurtz, Anne Puissant, Nicole Vincent |
ICPR | 2 |
| 2020 | Force Banner for the recognition of spatial relationsabstractStudying the spatial organization of objects in images is fundamental to increase both the understanding of a sensed scene and the explainability of the perceived similarity between images. This leads to the fundamental problem of handling spatial relations: given two objects depicted in an image, or two parts in an object, how to extract and describe efficiently their spatial configuration? Dedicated descriptors already exist for this task, like the efficient force histogram. In this article, we introduce the Force Banner, which extends it to two dimensions by using a panel of forces (attraction and repulsion), so as to benefit from more expressiveness and to model rich spatial information. This descriptor can be used as an intermediate representation of the image dedicated to the spatial configuration, and feed a classical 2D Convolutional Neural Network (CNN) to benefit from their powerful performances. As an illustration of this, we used it to solve a classification problem aiming to discriminate simple spatial relations, but with variable configuration complexities. Experimental results obtained on datasets of images with various shapes highlight the interest of this approach, in particular for complex spatial configurations. Robin Deléarde, Camille Kurtz, Philippe Dejean, Laurent Wendling |
ICPR | 2 |
| 2020 | Fuzzy directional enlacement landscapes for the evaluation of complex spatial relations
Michaël Clément, Camille Kurtz, Laurent Wendling |
Pattern Recognit. | 2 |
| 2020 | Editorial - Virtual Special Issue: "Hierarchical Representations: New Results and Challenges for Image Analysis"
Nicolas Passat, Camille Kurtz, Antoine Vacavant |
Pattern Recognit. Lett. | 2 |
| 2018 | Relevance feedback for enhancing content based image retrieval and automatic prediction of semantic image features: Application to bone tumor radiographs
Imon Banerjee, Camille Kurtz, Alon Edward Devorah, Bao H. Do, Daniel L. Rubin, Christopher F. Beaulieu |
J. Biomed. Informatics | 2 |
| 2018 | Learning spatial relations and shapes for structural object description and scene recognition
Michaël Clément, Camille Kurtz, Laurent Wendling |
Pattern Recognit. | 2 |
| 2018 | Binary Partition Tree construction from multiple features for image segmentation
Jimmy Francky Randrianasoa, Camille Kurtz, Eric Desjardin, Nicolas Passat |
Pattern Recognit. | 2 |
| 2017 | A Document Straight Line Based Segmentation for Complex Layout ExtractionabstractDocument layout extraction is a difficult step in the image interpretation process due to the high complexity of documents. The main challenge relies on the huge gap between both the physical and the logical structures of document images. In order to loose as few as possible information, most existing methods are working at pixel level. In this paper, we present a new framework for complex layout extraction based on features of high levels obtained from a document straight line based segmentation. We propose to capture the straight line segments thanks to a new transform integrating the local spatial organization of the segments contained in the document content. Such transform can be applied either on the foreground (related to the document content) or the background pixels, in order to take advantage of the duality of information present in both document parts. Experimental results obtained on the PRImA Layout Analysis dataset illustrate the robustness of our framework for the extraction of specific components of the document including text areas, images and separators. Héloïse Alhéritière, Florence Cloppet, Camille Kurtz, Jean-Marc Ogier, Nicole Vincent |
ICDAR | 3 |
| 2017 | Local Enlacement Histograms for Historical Drop Caps Style RecognitionabstractThis article focuses on the specific issue of drop caps image recognition in the context of cultural heritage preservation. Due to their heterogeneity and their weakly structured properties, these historical images represent challenging data. An important aspect in the recognition process of drop caps is their background styles, which can be considered as discriminative features to identify both the printer and the period. Most existing methods for style recognition are based on low-level features such as color or texture properties. In this article, we present a novel framework for the recognition of drop caps style based on features of higher levels. We propose to capture the spatial structure carried by these images using relative position descriptors modeling the enlacement between local cells of pixel layers obtained from a document segmentation step. Such descriptors are then exploited in an efficient bag-of-features learning procedure. Experimental results obtained on a dataset of historical drop caps images highlight the interest of this approach, and in particular the benefit of considering spatial information. Michaël Clément, Mickaël Coustaty, Camille Kurtz, Laurent Wendling |
ICDAR | 3 |
| 2017 | Evaluating the quality of binary partition trees based on uncertain semantic ground-truth for image segmentationabstractThe binary partition tree (BPT) is a hierarchical data-structure that models the content of an image in a multiscale way. In particular, a cut of the BPT of an image provides a segmentation, as a partition of the image support. Actually, building a BPT allows for dramatically reducing the search space for segmentation purposes, based on intrinsic (image signal) and extrinsic (construction metric) information. A large literature has been devoted to the construction on such metrics, and the associated choice of criteria (spectral, spatial, geometric, etc.) for building relevant BPTs, in particular in the challenging context of remote sensing. But, surprisingly, there exists few works dedicated to evaluate the quality of BPTs, i.e. their ability to further provide a satisfactory segmentation. In this paper, we propose a framework for BPT quality evaluation, in a supervised paradigm. Indeed, we assume that ground-truth segments are provided by an expert, possibly with a semantic labelling and a given uncertainty. Then, we describe local evaluation metrics, BPT nodes / ground-truth segments fitting strategies, and global quality score computation considering semantic information, leading to a complete evaluation framework. This framework is illustrated in the context of BPT segmentation of multispectral satellite images. Jimmy Francky Randrianasoa, Camille Kurtz, Pierre Gançarski, Eric Desjardin, Nicolas Passat |
ICIP | 2 |
| 2017 | Directional Enlacement Histograms for the Description of Complex Spatial Configurations between ObjectsabstractThe analysis of spatial relations between objects in digital images plays a crucial role in various application domains related to pattern recognition and computer vision. Classical models for the evaluation of such relations are usually sufficient for the handling of simple objects, but can lead to ambiguous results in more complex situations. In this article, we investigate the modeling of spatial configurations where the objects can be imbricated in each other. We formalize this notion with the term enlacement, from which we also derive the term interlacement, denoting a mutual enlacement of two objects. Our main contribution is the proposition of new relative position descriptors designed to capture the enlacement and interlacement between two-dimensional objects. These descriptors take the form of circular histograms allowing to characterize spatial configurations with directional granularity, and they highlight useful invariance properties for typical image understanding applications. We also show how these descriptors can be used to evaluate different complex spatial relations, such as the surrounding of objects. Experimental results obtained in the different application domains of medical imaging, document image analysis and remote sensing, confirm the genericity of this approach. Michaël Clément, Adrien Poulenard, Camille Kurtz, Laurent Wendling |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2016 | Bags of spatial relations and shapes features for structural object descriptionabstractWe introduce a novel bags-of-features framework based on relative position descriptors, modeling both spatial relations and shape information between the pairwise structural subparts of objects. First, we propose a hierarchical approach for the decomposition of complex objects into structural subparts, as well as their description using the concept of Force Histogram Decomposition (FHD). Then, an original learning methodology is presented, in order to produce discriminative hierarchical spatial features for object classification tasks. The cornerstone is to build an homogeneous vocabulary of shapes and spatial configurations occurring across the objects at different scales of decomposition. An advantage of this learning procedure is its compatibility with traditional bags-of-features frameworks, allowing for hybrid representations of both structural and local features. Classification results obtained on two datasets of images highlight the interest of this approach based on hierarchical spatial relations descriptors to recognize structured objects. Michaël Clément, Camille Kurtz, Laurent Wendling |
ICPR | 2 |
| 2015 | Tree Leaves Extraction in Natural Images: Comparative Study of Preprocessing Tools and Segmentation MethodsabstractIn this paper, we propose a comparative study of various segmentation methods applied to the extraction of tree leaves from natural images. This study follows the design of a mobile application, developed by Cerutti et al. (published in ReVeS Participation--Tree Species Classification Using Random Forests and Botanical Features. CLEF 2012), to highlight the impact of the choices made for segmentation aspects. All the tests are based on a database of 232 images of tree leaves depicted on natural background from smartphones acquisitions. We also propose to study the improvements, in terms of performance, using preprocessing tools, such as the interaction between the user and the application through an input stroke, as well as the use of color distance maps. The results presented in this paper shows that the method developed by Cerutti et al. (denoted Guided Active Contour), obtains the best score for almost all observation criteria. Finally, we detail our online benchmark composed of 14 unsupervised methods and 6 supervised ones. Manuel Grand-Brochier, Antoine Vacavant, Guillaume Cerutti, Camille Kurtz, Jonathan Weber, Laure Tougne |
IEEE Trans. Image Process. | 4 |
| 2014 | A semantic framework for the retrieval of similar radiological images based on medical annotationsabstractImage retrieval approaches can assist radiologists by finding similar images in databases as a means to providing decision support. In general, images are indexed using low-level imaging features, and a distance function is used to find the best matches in the feature space. However, using low-level features to capture the appearance of diseases in images is challenging and the semantic gap between these features and the high-level visual concepts in radiology may impair the system performance. We present a semantic framework that enables retrieving similar images based on high-level semantic image annotations. This framework relies on (1) an automatic approach to predict the annotations as semantic terms from Riesz texture image features and (2) a distance function to compare images considering both texture-based and radiodensity-based similarities among image annotations. Experiments performed on CT images emphasize the relevance of this framework. Camille Kurtz, Adrien Depeursinge, Christopher F. Beaulieu, Daniel L. Rubin |
ICIP | 1 |
| 2014 | Multivalued Component-Tree FilteringabstractWe introduce the new notion of multivalued component-tree, that extends the classical component-tree initially devoted to grey-level images, in the mathematical morphology framework. We prove that multivalued component-trees can model images whose values are hierarchically organized. We also show that they can be efficiently built from standard component-tree construction algorithms, and involved in antiextensive filtering procedures. The relevance and usefulness of multivalued component-trees is illustrated by an applicative example on hierarchically classified remote sensing images. Camille Kurtz, Benoît Naegel, Nicolas Passat |
ICPR | 1 |
| 2014 | A hierarchical knowledge-based approach for retrieving similar medical images described with semantic annotations
Camille Kurtz, Christopher F. Beaulieu, Sandy Napel, Daniel L. Rubin |
J. Biomed. Informatics | 1 |
| 2014 | On combining image-based and ontological semantic dissimilarities for medical image retrieval applications
Camille Kurtz, Adrien Depeursinge, Sandy Napel, Christopher F. Beaulieu, Daniel L. Rubin |
Medical Image Anal. | 1 |
| 2014 | Connected Filtering Based on Multivalued Component-TreesabstractIn recent papers, a new notion of component-graph was introduced. It extends the classical notion of component-tree initially proposed in mathematical morphology to model the structure of gray-level images. Component-graphs can indeed model the structure of any-gray-level or multivalued-images. We now extend the antiextensive filtering scheme based on component-trees, to make it tractable in the framework of component-graphs. More precisely, we provide solutions for building a component-graph, reducing it based on selection criteria, and reconstructing a filtered image from a reduced component-graph. In this paper, we first consider the cases where component-graphs still have a tree structure; they are then called multivalued component-trees. The relevance and usefulness of such multivalued component-trees are illustrated by applicative examples on hierarchically classified remote sensing images. Camille Kurtz, Benoît Naegel, Nicolas Passat |
IEEE Trans. Image Process. | 1 |
| 2014 | Predicting Visual Semantic Descriptive Terms From Radiological Image Data: Preliminary Results With Liver Lesions in CTabstractWe describe a framework to model visual semantics of liver lesions in CT images in order to predict the visual semantic terms (VST) reported by radiologists in describing these lesions. Computational models of VST are learned from image data using linear combinations of high-order steerable Riesz wavelets and support vector machines (SVM). In a first step, these models are used to predict the presence of each semantic term that describes liver lesions. In a second step, the distances between all VST models are calculated to establish a nonhierarchical computationally-derived ontology of VST containing inter-term synonymy and complementarity. A preliminary evaluation of the proposed framework was carried out using 74 liver lesions annotated with a set of 18 VSTs from the RadLex ontology. A leave-one-patient-out cross-validation resulted in an average area under the ROC curve of 0.853 for predicting the presence of each VST. The proposed framework is expected to foster human-computer synergies for the interpretation of radiological images while using rotation-covariant computational models of VSTs to 1) quantify their local likelihood and 2) explicitly link them with pixel-based image content in the context of a given imaging domain. Adrien Depeursinge, Camille Kurtz, Christopher F. Beaulieu, Sandy Napel, Daniel L. Rubin |
IEEE Trans. Medical Imaging | 2 |
| 2013 | A hierarchical semantic-based distance for nominal histogram comparison
Camille Kurtz, Pierre Gançarski, Nicolas Passat, Anne Puissant |
Data Knowl. Eng. | 1 |
| 2012 | A histogram semantic-based distance for multiresolution image classificationabstractImage classification methods based on histogram analysis generally require to use relevant distances for histogram comparison. In this article, we propose a new distance devoted to compare histograms associated to semantic concepts linked by (dis)similarity correlations. This distance, whose computation relies on a hierarchical strategy, captures the multilevel semantic relations between these concepts. It also inherits from the low complexity properties of standard bin-to-bin distances, thus leading to fast and accurate results in the context of multiresolution image classification. Experiments performed on satellite images emphasize the relevance and usefulness of the proposed distance. Camille Kurtz, Nicolas Passat, Pierre Gançarski, Anne Puissant |
ICIP | 1 |
| 2012 | Domain adaptation for the extraction of complex urban patterns from multiresolution satellite imagesabstractThe extraction of complex urban patterns from Very High Spatial Resolution (VHSR) images presents several challenges related to the complexity of the data. Based on the availability of images of a same scene at various resolutions (Medium to Very High Spatial resolutions), a hierarchical approach has been recently proposed to segment/classify objects of interest in a top-down fashion in order to determine patterns from VHSR images. To perform, this method requires the interactive definition of segmentation examples for each considered resolution image. In the context of large dataset processing, such interactive task becomes time consuming. To deal with this issue, we propose in this article, an extension of the domain adaptation paradigm enabling the transfer of the segmentation examples defined on a source dataset to automatically process a target one. Experiments performed on urban images provide satisfactory results which may be further used for operational needs. Camille Kurtz, Anne Puissant, Nicolas Passat, Pierre Gançarski |
IGARSS | 1 |
| 2012 | Extraction of complex patterns from multiresolution remote sensing images: A hierarchical top-down methodology
Camille Kurtz, Nicolas Passat, Pierre Gançarski, Anne Puissant |
Pattern Recognit. | 1 |
| 2012 | Spatio-temporal reasoning for the classification of satellite image time series
François Petitjean, Camille Kurtz, Nicolas Passat, Pierre Gançarski |
Pattern Recognit. Lett. | 2 |
| 2011 | A context-based approach for the classification of Satellite Image Time SeriesabstractSatellite Image Time Series (SITS) analysis is an important domain with various applications in land study. In the coming years, both high temporal and high spatial resolution SITS will be available. This article aims at providing both temporal and spatial analysis of SITS. We propose first segmenting each image of the series, and then using these segmentations in order to characterize each pixel of the data with a spatial dimension (i.e. with contextual information). Providing spatially characterized pixels, pixel-based temporal analysis can be performed. Experiments carried out with this methodology show the relevance of this approach and the significance of the resulting extracted patterns in the context of the analysis of SITS. Camille Kurtz, François Petitjean, Pierre Gançarski |
IGARSS | 1 |