Mariella Dimiccoli

dblp:86/7595 · DBLP profile ↗
← Back
32ranked-venue papers
11as first author
10since 2021 · last 2025
0000-0002-2669-400XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 7 first-author · 6 since 2021Artificial intelligence and machine learning · 16 · 5 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 CLOT: Closed Loop Optimal Transport for Unsupervised Action Segmentation
abstract
© 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.
Elena Belén Bueno-Benito, Mariella Dimiccoli
ICCV2
2024 3M-Transformer: A Multi-Stage Multi-Stream Multimodal Transformer for Embodied Turn-Taking Prediction
abstract
Predicting turn-taking in multiparty conversations has many practical applications in human-computer/robot interaction. However, the complexity of human communication makes it a challenging task. Recent advances have shown that synchronous multi-perspective egocentric data can significantly improve turn-taking prediction compared to asynchronous, single-perspective transcriptions. Building on this research, we propose a new multimodal transformer-based architecture for predicting turn-taking in embodied, synchronized multi-perspective data. Our experimental results on the recently introduced EgoCom dataset show a substantial performance improvement of up to 14.01% on average compared to existing baselines and alternative transformer-based approaches. The source code, and the pre-trained models of our 3T-Transformer will be available upon acceptance1.
Mehdi Fatan, Emanuele Mincato, Dimitra Pintzou, Mariella Dimiccoli
ICASSP4
2024 2by2: Weakly-Supervised Learning for Global Action Segmentation
Elena Belén Bueno-Benito, Mariella Dimiccoli
ICPR (32)2
2023 Enhancing egocentric 3D pose estimation with third person views
abstract
We propose a novel approach to enhance the 3D body pose estimation of a person computed from videos captured from a single wearable camera. The main technical contribution consists of leveraging high-level features linking first- and third-views in a joint embedding space. To learn such embedding space we introduce First2Third-Pose, a new paired synchronized dataset of nearly 2000 videos depicting human activities captured from both first- and third-view perspectives. We explicitly consider spatial- and motion-domain features, combined using a semi-Siamese architecture trained in a self-supervised fashion. Experimental results demonstrate that the joint multi-view embedded space learned with our dataset is useful to extract discriminatory features from arbitrary single-view egocentric videos, with no need to perform any sort of domain adaptation or knowledge of camera parameters. An extensive evaluation demonstrates that we achieve significant improvement in egocentric 3D body pose estimation performance on two unconstrained datasets, over three supervised state-of-the-art approaches. The collected dataset and pre-trained model are available for research purposes.1
Ameya Dhamanaskar, Mariella Dimiccoli, Enric Corona, Albert Pumarola, Francesc Moreno-Noguer
Pattern Recognit.2
2022 Debiased-CAM to mitigate image perturbations with faithful visual explanations of machine learning
abstract
Model explanations such as saliency maps can improve user trust in AI by highlighting important features for a prediction. However, these become distorted and misleading when explaining predictions of images that are subject to systematic error (bias) by perturbations and corruptions. Furthermore, the distortions persist despite model fine-tuning on images biased by different factors (blur, color temperature, day/night). We present Debiased-CAM to recover explanation faithfulness across various bias types and levels by training a multi-input, multi-task model with auxiliary tasks for explanation and bias level predictions. In simulation studies, the approach not only enhanced prediction accuracy, but also generated highly faithful explanations about these predictions as if the images were unbiased. In user studies, debiased explanations improved user task performance, perceived truthfulness and perceived helpfulness. Debiased training can provide a versatile platform for robust performance and explanation faithfulness for a wide range of applications with data biases.
Wencan Zhang, Mariella Dimiccoli, Brian Y. Lim
CHI2
2022 Recognizing object surface material from impact sounds for robot manipulation
abstract
We investigated the use of impact sounds generated during exploratory behaviors in a robotic manipulation setup as cues for predicting object surface material and for recognizing individual objects. We collected and make available the YCB-impact sounds dataset which includes over 3,000 impact sounds for the YCB set of everyday objects lying on a table. Impact sounds were generated in three modes: (i) human holding a gripper and hitting, scratching, or dropping the object; (ii) gripper attached to a teleoperated robot hitting the object from the top; (iii) autonomously operated robot hitting the objects from the side with two different speeds. A convolutional neural network is trained from scratch to recognize the object material (steel, aluminium, hard plastic, soft plastic, other plastic, ceramic, wood, paper/cardboard, foam, glass, rubber) from a single impact sound. On the manually collected dataset with more variability in the speed of the action, nearly 60% accuracy for the test set (not presented objects) was achieved. On a robot setup and a stereotypical poking action from top, accuracy of 85% was achieved. This performance drops to 79% if multiple exploratory actions are combined. Individual objects from the set of 75 objects can be recognized with a 79% accuracy. This work demonstrates promising results regarding the possibility of using impact sound for recognition in tasks like single-stream recycling where objects have to be sorted based on their material composition.
Mariella Dimiccoli, Shubhan P. Patni, Matej Hoffmann, Francesc Moreno-Noguer
IROS1
2021 Learning grounded word meaning representations on similarity graphs
abstract
This paper introduces a novel approach to learn visually grounded meaning representations of words as low-dimensional node embeddings on an underlying graph hierarchy.The lower level of the hierarchy models modality-specific word representations through dedicated but communicating graphs, while the higher level puts these representations together on a single graph to learn a representation jointly from both modalities.The topology of each graph models similarity relations among words, and is estimated jointly with the graph embedding.The assumption underlying this model is that words sharing similar meaning correspond to communities in an underlying similarity graph in a lowdimensional space.We named this model Hierarchical Multi-Modal Similarity Graph Embedding (HM-SGE).Experimental results validate the ability of HM-SGE to simulate human similarity judgements and concept categorization, outperforming the state of the art. 1
Mariella Dimiccoli, Herwig Wendt, Pau Batlle Franch
EMNLP (1)1
2021 Graph Constrained Data Representation Learning for Human Motion Segmentation
abstract
Recently, transfer subspace learning based approaches have shown to be a valid alternative to unsupervised subspace clustering and temporal data clustering for human motion segmentation (HMS). These approaches leverage prior knowledge from a source domain to improve clustering performance on a target domain, and currently they represent the state of the art in HMS. Bucking this trend, in this paper, we propose a novel unsupervised model that learns a representation of the data and digs clustering information from the data itself. Our model is reminiscent of temporal subspace clustering, but presents two critical differences. First, we learn an auxiliary data matrix that can deviate from the initial data, hence confers more degrees of freedom to the coding matrix. Second, we introduce a regularization term for this auxiliary data matrix that preserves the local geometrical structure present in the high-dimensional space. The proposed model is efficiently optimized by using an original Alternating Direction Method of Multipliers (ADMM) formulation allowing to learn jointly the auxiliary data representation, a nonnegative dictionary and a coding matrix. Experimental results on four benchmark datasets for HMS demonstrate that our approach achieves significantly better clustering performance then state-of-the-art methods, including both unsupervised and more recent semi-supervised transfer learning approaches1.
Mariella Dimiccoli, Lluís Garrido, Guillem Rodríguez Corominas, Herwig Wendt
ICCV1
2021 Interaction-GCN: A Graph Convolutional Network Based Framework for Social Interaction Recognition in Egocentric Videos
abstract
In this paper we propose a new framework to categorize social interactions in egocentric videos, we named InteractionGCN. Our method extracts patterns of relational and non-relational cues at the frame level and uses them to build a relational graph from which the interactional context at the frame level is estimated via a Graph Convolutional Network (GCN) based approach. Then it propagates this context over time, together with first-person motion information, through a Gated Recurrent Unit architecture. Ablation studies and experimental evaluation on two publicly available datasets validate the proposed approach and establish state of the art results.11Features and trained model available at https://github.com/simonefelicioni/InteractionGCN
Simone Felicioni, Mariella Dimiccoli
ICIP2
2021 Learning Event Representations for Temporal Segmentation of Image Sequences by Dynamic Graph Embedding
abstract
Recently, self-supervised learning has proved to be effective to learn representations of events suitable for temporal segmentation in image sequences, where events are understood as sets of temporally adjacent images that are semantically perceived as a whole. However, although this approach does not require expensive manual annotations, it is data hungry and suffers from domain adaptation problems. As an alternative, in this work, we propose a novel approach for learning event representations named Dynamic Graph Embedding (DGE). The assumption underlying our model is that a sequence of images can be represented by a graph that encodes both semantic and temporal similarity. The key novelty of DGE is to learn jointly the graph and its graph embedding. At its core, DGE works by iterating over two steps: 1) updating the graph representing the semantic and temporal similarity of the data based on the current data representation, and 2) updating the data representation to take into account the current data graph structure. The main advantage of DGE over state-of-the-art self-supervised approaches is that it does not require any training set, but instead learns iteratively from the data itself a low-dimensional embedding that reflects their temporal and semantic similarity. Experimental results on two benchmark datasets of real image sequences captured at regular time intervals demonstrate that the proposed DGE leads to event representations effective for temporal segmentation. In particular, it achieves robust temporal segmentation on the EDUBSeg and EDUBSeg-Desc benchmark datasets, outperforming the state of the art. Additional experiments on two Human Motion Segmentation benchmark datasets demonstrate the generalization capabilities of the proposed DGE.
Mariella Dimiccoli, Herwig Wendt
IEEE Trans. Image Process.1
2020 Modeling Long-Term Interactions to Enhance Action Recognition
abstract
In this paper, we propose a new approach to understand actions in egocentric videos that exploit the semantics of object interactions at both frame and temporal levels. At the frame level, we use a region-based approach that takes as input a primary region roughly corresponding to the user hands and a set of secondary regions potentially corresponding to the interacting objects and calculates the action score through a CNN formulation. This information is then fed to a Hierarchical Long Short-Term Memory Network (HLSTM) that captures temporal dependencies between actions within and across shots. Ablation studies thoroughly validate the proposed approach, showing in particular that both levels of the HLSTM architecture contribute to performance improvement. Furthermore, quantitative comparisons show that the proposed approach outperforms the state-of-the-art in terms of action recognition on standard benchmarks, without relying on motion information.
Alejandro Cartas Ayala, Petia Radeva, Mariella Dimiccoli
ICPR3
2019 Social Relation Recognition in Egocentric Photostreams
abstract
This paper proposes an approach to automatically categorize the social interactions of a user wearing a photo-camera (2fpm), by relying solely on what the camera is seeing. The problem is challenging due to the overwhelming complexity of social life and the extreme intra-class variability of social interactions captured under unconstrained conditions. We adopt the formalization proposed in Bugental's social theory, that groups human relations into five social domains with related categories. Our method is a new deep learning architecture that exploits the hierarchical structure of the label space and relies on a set of social attributes estimated at frame level to provide a semantic representation of social interactions. Experimental results on the new EgoSocialRelation dataset demonstrate the effectiveness of our proposal.
Emanuel Sanchez Aimar, Petia Radeva, Mariella Dimiccoli
ICIP3
2019 Enhancing Temporal Segmentation by Nonlocal Self-Similarity
abstract
Temporal segmentation of untrimmed videos and photo-streams is currently an active area of research in computer vision and image processing. This paper proposes a new approach to improve the temporal segmentation of photo-streams. The method consists in enhancing image representations by encoding long-range temporal dependencies. Our key contribution is to take advantage of the temporal stationarity assumption of photostreams for modeling each frame by its nonlocal self-similarity function. The proposed approach is put to test on the EDUB-Seg dataset, a standard benchmark for egocentric photostream temporal segmentation. Starting from seven different (CNN based) image features, the method yields consistent improvements in event segmentation quality, leading to an average increase of F-measure of 3.71% with respect to the state of the art.
Mariella Dimiccoli, Herwig Wendt
ICIP1
2019 Smartphone picture organization: A hierarchical approach
Stefan Lonn, Petia Radeva, Mariella Dimiccoli
Comput. Vis. Image Underst.3
2018 Towards social pattern characterization in egocentric photo-streams
Maedeh Aghaei, Mariella Dimiccoli, Cristian Canton, Petia Radeva
Comput. Vis. Image Underst.2
2018 Introduction to the special issue: Egocentric Vision and Lifelogging
Mariella Dimiccoli, Cathal Gurrin, David Crandall, Xavier Giró-i-Nieto, Petia Radeva
J. Vis. Commun. Image Represent.1
2018 Learning quantification from images: A structured neural architecture
abstract
Abstract Major advances have recently been made in merging language and vision representations. Most tasks considered so far have confined themselves to the processing of objects and lexicalised relations amongst objects (content words). We know, however, that humans (even pre-school children) can abstract over raw multimodal data to perform certain types of higher level reasoning, expressed in natural language byfunction words. A case in point is given by their ability to learn quantifiers, i.e. expressions likefew,someandall. From formal semantics and cognitive linguistics, we know that quantifiers are relations over sets which, as a simplification, we can see as proportions. For instance, inmost fish are red,mostencodes the proportion of fish which are red fish. In this paper, we study how well current neural network strategies model such relations. We propose a task where, given an image and a query expressed by an object–property pair, the system must return a quantifier expressing which proportions of the queried object have the queried property. Our contributions are twofold. First, we show that the best performance on this task involves coupling state-of-the-art attention mechanisms with a network architecture mirroring the logical structure assigned to quantifiers by classic linguistic formalisation. Second, we introduce a new balanced dataset of image scenarios associated with quantification queries, which we hope will foster further research in this area.
Ionut Sorodoc, Sandro Pezzelle, Aurélie Herbelot, Mariella Dimiccoli
Nat. Lang. Eng.4
2018 Batch-based activity recognition from egocentric photo-streams revisited
Alejandro Cartas Ayala, Juan Marín, Petia Radeva, Mariella Dimiccoli
Pattern Anal. Appl.4
2017 Clothing and People - A Social Signal Processing Perspective
abstract
In our society and century, clothing is not anymore used only as a means for body protection. Our paper builds upon the evidence, studied within the social sciences, that clothing brings a clear communicative message in terms of social signals, influencing the impression and behaviour of others towards a person. In fact, clothing correlates with personality traits, both in terms of self-assessment and assessments that unacquainted people give to an individual. The consequences of these facts are important: the influence of clothing on the decision making of individuals has been investigated in the literature, showing that it represents a discriminative factor to differentiate among diverse groups of people. Unfortunately, this has been observed after cumbersome and expensive manual annotations, on very restricted populations, limiting the scope of the resulting claims. With this position paper, we want to sketch the main steps of the very first systematic analysis, driven by social signal processing techniques, of the relationship between clothing and social signals, both sent and perceived. Thanks to human parsing technologies, which exhibit high robustness owing to deep learning architectures, we are now capable to isolate visual patterns characterising a large types of garments. These algorithms will be used to capture statistical relations on a large corpus of evidence to confirm the sociological findings and to go beyond the state of the art.
Maedeh Aghaei, Federico Parezzan, Mariella Dimiccoli, Petia Radeva, Marco Cristani
FG3
2017 All the people around me: Face discovery in egocentric photo-streams
abstract
Given an unconstrained stream of images captured by a wearable photo-camera (2fpm), we propose an unsupervised bottom-up approach for automatic clustering appearing faces into the individual identities present in these data. The problem is challenging since images are acquired under real world conditions; hence the visible appearance of the people in the images undergoes intensive variations. Our proposed pipeline consists of first arranging the photo-stream into events, later, localizing the appearance of multiple people in them, and finally, grouping various appearances of the same person across different events. Experimental results performed on a dataset acquired by wearing a photo-camera during one month, demonstrate the effectiveness of the proposed approach for the considered purpose.
Maedeh Aghaei, Mariella Dimiccoli, Petia Radeva
ICIP2
2017 LTA 2017: The Second Workshop on Lifelogging Tools and Applications
abstract
The organisation of personal data is receiving increasing research attention due to the challenges we face in gathering, enriching, searching, and visualising such data. Given the increasing ease with which personal data being gathered by individuals, the concept of a lifelog digital library of rich multimedia and sensory content for every individual is fast becoming a reality. The LTA 2017 workshop aims to bring together academics and practitioners to discuss approaches to lifelog data analytics and applications; and to debate the opportunities and challenges for researchers in this new and challenging area.
Cathal Gurrin, Xavier Giró-i-Nieto, Petia Radeva, Mariella Dimiccoli, Duc-Tien Dang-Nguyen, Hideo Joho
ACM Multimedia4
2017 SR-clustering: Semantic regularized clustering for egocentric photo streams segmentation
abstract
While wearable cameras are becoming increasingly popular, locating relevant information in large unstructured collections of egocentric images is still a tedious and time consuming process. This paper addresses the problem of organizing egocentric photo streams acquired by a wearable camera into semantically meaningful segments, hence making an important step towards the goal of automatically annotating these photos for browsing and retrieval. In the proposed method, first, contextual and semantic information is extracted for each image by employing a Convolutional Neural Networks approach. Later, a vocabulary of concepts is defined in a semantic space by relying on linguistic information. Finally, by exploiting the temporal coherence of concepts in photo streams, images which share contextual and semantic attributes are grouped together. The resulting temporal segmentation is particularly suited for further analysis, ranging from event recognition to semantic indexing and summarization. Experimental results over egocentric set of nearly 31,000 images, show the prominence of the proposed approach over state-of-the-art methods.
Mariella Dimiccoli, Marc Bolaños, Estefanía Talavera, Maedeh Aghaei, Stavri G. Nikolov, Petia Radeva
Comput. Vis. Image Underst.1
2017 Toward Storytelling From Visual Lifelogging: An Overview
abstract
Visual lifelogging consists of acquiring images that capture the daily experiences of the user by wearing a camera over a long period of time. The pictures taken offer considerable potential for knowledge mining concerning how people live their lives; hence, they open up new opportunities for many potential applications in fields including healthcare, security, leisure, and the quantified self. However, automatically building a story from a huge collection of unstructured egocentric data presents major challenges. This paper provides a thorough review of advances made so far in egocentric data analysis and, in view of the current state of the art, indicates new lines of research to move us toward storytelling from visual lifelogging.
Marc Bolaños, Mariella Dimiccoli, Petia Radeva
IEEE Trans. Hum. Mach. Syst.2
2016 With whom do I interact? Detecting social interactions in egocentric photo-streams
abstract
Given a user wearing a low frame rate wearable camera during a day, this work aims to automatically detect the moments when the user gets engaged into a social interaction solely by reviewing the automatically captured photos by the worn camera. The proposed method, inspired by the sociological concept of F-formation, exploits distance and orientation of the appearing individuals -with respect to the user- in the scene from a bird-view perspective. As a result, the interaction pattern over the sequence can be understood as a two-dimensional time series that corresponds to the temporal evolution of the distance and orientation features over time. A Long-Short Term Memory-based Recurrent Neural Network is then trained to classify each time series. Experimental evaluation over a dataset of 30.000 images has shown promising results on the proposed method for social interaction detection in egocentric photo-streams.
Maedeh Aghaei, Mariella Dimiccoli, Petia Radeva
ICPR2
2016 LTA 2016: The First Workshop on Lifelogging Tools and Applications
abstract
The organisation of personal data is receiving increasing research attention due to the challenges we face in gathering, enriching, searching, and visualising such data. Given the increasing ease with which personal data being gathered by individuals, the concept of a lifelog digital library of rich multimedia and sensory content for every individual is fast becoming a reality. The LTA~2016 workshop aims to bring together academics and practitioners to discuss approaches to lifelog data analytics and applications; and to debate the opportunities and challenges for researchers in this new and challenging area.
Cathal Gurrin, Xavier Giró-i-Nieto, Petia Radeva, Mariella Dimiccoli, Håvard D. Johansen, Hideo Joho, Vivek K. Singh 0001
ACM Multimedia4
2016 Multi-face tracking by extended bag-of-tracklets in egocentric photo-streams
Maedeh Aghaei, Mariella Dimiccoli, Petia Radeva
Comput. Vis. Image Underst.2
2016 Particle detection and tracking in fluorescence time-lapse imaging: a contrario approach
Mariella Dimiccoli, Jean-Pascal Jacob, Lionel Moisan
Mach. Vis. Appl.1
2015 Towards social interaction detection in egocentric photo-streams
abstract
Detecting social interaction in videos relying solely on visual cues is a valuable task that is receiving increasing attention in recent years. In this work, we address this problem in the challenging domain of egocentric photo-streams captured by a low temporal resolution wearable camera (2fpm). The major difficulties to be handled in this context are the sparsity of observations as well as unpredictability of camera motion and attention orientation due to the fact that the camera is worn as part of clothing. Our method consists of four steps: multi-faces localization and tracking, 3D localization, pose estimation and analysis of f-formations. By estimating pair-to-pair interaction probabilities over the sequence, our method states the presence or absence of interaction with the camera wearer and specifies which people are more involved in the interaction. We tested our method over a dataset of 18.000 images and we show its reliability on our considered purpose.
Maedeh Aghaei, Mariella Dimiccoli, Petia Radeva
ICMV2
2013 Localization of Protein Aggregation in Escherichia coli Is Governed by Diffusion and Nucleoid Macromolecular Crowding Effect
abstract
Aggregates of misfolded proteins are a hallmark of many age-related diseases. Recently, they have been linked to aging of Escherichia coli (E. coli) where protein aggregates accumulate at the old pole region of the aging bacterium. Because of the potential of E. coli as a model organism, elucidating aging and protein aggregation in this bacterium may pave the way to significant advances in our global understanding of aging. A first obstacle along this path is to decipher the mechanisms by which protein aggregates are targeted to specific intercellular locations. Here, using an integrated approach based on individual-based modeling, time-lapse fluorescence microscopy and automated image analysis, we show that the movement of aging-related protein aggregates in E. coli is purely diffusive (Brownian). Using single-particle tracking of protein aggregates in live E. coli cells, we estimated the average size and diffusion constant of the aggregates. Our results provide evidence that the aggregates passively diffuse within the cell, with diffusion constants that depend on their size in agreement with the Stokes-Einstein law. However, the aggregate displacements along the cell long axis are confined to a region that roughly corresponds to the nucleoid-free space in the cell pole, thus confirming the importance of increased macromolecular crowding in the nucleoids. We thus used 3D individual-based modeling to show that these three ingredients (diffusion, aggregation and diffusion hindrance in the nucleoids) are sufficient and necessary to reproduce the available experimental data on aggregate localization in the cells. Taken together, our results strongly support the hypothesis that the localization of aging-related protein aggregates in the poles of E. coli results from the coupling of passive diffusion-aggregation with spatially non-homogeneous macromolecular crowding. They further support the importance of "soft" intracellular structuring (based on macromolecular crowding) in diffusion-based protein localization in E. coli.
Anne-Sophie Coquel, Jean-Pascal Jacob, Maël Primet, Alice Demarez, Mariella Dimiccoli, Thomas Julou, Lionel Moisan, Ariel B. Lindner, Hugues Berry
PLoS Comput. Biol.5
2009 Exploiting T-junctions for depth segregation in single images
abstract
Occlusion is one of the major consequences of the physical image generation process: it occurs when an opaque object partly obscures the view of another object further away from the viewpoint. Local signatures of occlusion in the projected image plane are T-shaped junctions. They represent, in some sense, one of the most primitive depth information. In this paper, we investigate the usefulness of T-junctions for depth segregation in single images. Our strategy consists in incorporating ordering information provided by T-junctions into a region merging algorithm and then reasoning about the depth relations between the regions of the final partition using a graph model. Experimental results demonstrate the effectiveness of the proposed approach.
Mariella Dimiccoli, Philippe Salembier
ICASSP1
2009 Hierarchical region-based representation for segmentation and filtering with depth in single images
abstract
This paper presents an algorithm for tree-based representation of single images and its applications to segmentation and filtering with depth. In a our recent work, we have addressed the problem of segmentation with depth by incorporating depth ordering information into a region merging algorithm and by reasoning about depth relations through a graph model. In this paper, we extend this previous work giving a two-fold contribution. First, we propose to model each pixel statistically by its probability distribution instead of deterministically by its color value. Second, we propose a depth-oriented filter, which allows to remove foreground regions and to replace them with a plausible background. Experimental results are satisfactory.
Mariella Dimiccoli, Philippe Salembier
ICIP1
2007 Geometrical image filtering with connected operators and image inpainting
abstract
This paper deals with the joint use of connected operators and image inpainting for image filtering. Connected operators filter the image by merging its flat zones while preserving contour information. Image inpainting restores the values of an image for a destroyed or consciously masked subregion of the image domain. In the present paper, it will be shown that image inpainting can be combined with connected operators to perform an efficient geometrical filtering technique. First, connected operators are presented and their drawbacks for certain applications are highlighted. Second, image inpainting methodology is introduced and a structural image inpainting algorithm is described. Finally, a general filtering scheme is proposed to show how the drawbacks of connected operators can be efficiently solved by structural image inpainting.
Mariella Dimiccoli, Philippe Salembier
VCIP1