Guillermo Sapiro

dblp:82/5175 · also Guillermo R. Sapiro · DBLP profile ↗
← Back
265ranked-venue papers
19as first author
14since 2021 · last 2025
0000-0001-9190-6964ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 176 · 14 first-author · 2 since 2021Artificial intelligence and machine learning · 108 · 8 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 28 · 2 since 2021Systems, architecture and hardware · 4 · 1 since 2021Databases, data management, data science and information retrieval · 3Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2025 SSOLE: Rethinking Orthogonal Low-rank Embedding for Self-Supervised Learning
abstract
Self-supervised learning (SSL) aims to learn meaningful representations from unlabeled data. Orthogonal Low-rank Embedding (OLE) shows promise for SSL by enhancing intra-class similarity in a low-rank subspace and promoting inter-class dissimilarity in a high-rank subspace, making it particularly suitable for multi-view learning tasks. However, directly applying OLE to SSL poses significant challenges: (1) the virtually infinite number of "classes" in SSL makes achieving the OLE objective impractical, leading to representational collapse; and (2) low-rank constraints may fail to distinguish between positively and negatively correlated features, further undermining learning. To address these issues, we propose SSOLE (Self-Supervised Orthogonal Low-rank Embedding), a novel framework that integrates OLE principles into SSL by (1) decoupling the low-rank and high-rank enforcement to align with SSL objectives; and (2) applying low-rank constraints to feature deviations from their mean, ensuring better alignment of positive pairs by accounting for the signs of cosine similarities. Our theoretical analysis and empirical results demonstrate that these adaptations are crucial to SSOLE’s effectiveness. Moreover, SSOLE achieves competitive performance across SSL benchmarks without relying on large batch sizes, memory banks, or dual-encoder architectures, making it an efficient and scalable solution for self-supervised tasks. Code is available at https://github.com/husthuaan/ssole.
Lun Huang, Qiang Qiu 0001, Guillermo Sapiro
ICLR3
2025 Addressing Misspecification in Simulation-based Inference through Data-driven Calibration
abstract
Driven by steady progress in deep generative modeling, simulation-based inference (SBI) has emerged as the workhorse for inferring the parameters of stochastic simulators. However, recent work has demonstrated that model misspecification can harm SBI's reliability, preventing its adoption in important applications where only misspecified simulators are available. This work introduces robust posterior estimation (RoPE), a framework that overcomes model misspecification with a small real-world calibration set of ground truth parameter measurements. We formalize the misspecification gap as the solution of an optimal transport (OT) problem between learned representations of real-world and simulated observations, allowing RoPE to learn a model of the misspecification without placing additional assumptions on its nature. RoPE shows how the calibration set and OT together offer a controllable balance between calibrated uncertainty and informative inference even under severely misspecified simulators. Results on four synthetic tasks and two real-world problems with ground-truth labels demonstrate that RoPE outperforms baselines and consistently returns informative and calibrated credible intervals.
Antoine Wehenkel, Juan L. Gamella, Ozan Sener, Jens Behrmann, Guillermo Sapiro, Jörn-Henrik Jacobsen, Marco Cuturi
ICML5
2024 Achieving Group Distributional Robustness and Minimax Group Fairness with Interpolating Classifiers
Natalia Martínez, Martín Bertrán, Guillermo Sapiro
AISTATS3
2024 From Geometry to Causality- Ricci Curvature and the Reliability of Causal Inference on Networks
abstract
Causal inference on networks faces challenges posed in part by violations of standard identification assumptions due to dependencies between treatment units. Although graph geometry fundamentally influences such dependencies, the potential of geometric tools for causal inference on networked treatment units is yet to be unlocked. Moreover, despite significant progress utilizing graph neural networks (GNNs) for causal inference on networks, methods for evaluating their achievable reliability without ground truth are lacking. In this work we establish for the first time a theoretical link between network geometry, the graph Ricci curvature in particular, and causal inference, formalizing the intrinsic challenges that negative curvature poses to estimating causal parameters. The Ricci curvature can then be used to assess the reliability of causal estimates in structured data, as we empirically demonstrate. Informed by this finding, we propose a method using the geometric Ricci flow to reduce causal effect estimation error in networked data, showcasing how this newfound connection between graph geometry and causal inference could improve GNN-based causal inference. Bridging graph geometry and causal inference, this paper opens the door to geometric techniques for improving causal estimation on networks.
Amirhossein Farzam, Allen R. Tannenbaum, Guillermo Sapiro
ICML3
2023 Enhancing early autism prediction based on electronic records using clinical narratives
Junya Chen, Matthew Engelhard, Ricardo Henao, Samuel Berchuck, Brian Eichner, Eliana M. Perrin, Guillermo Sapiro, Geraldine Dawson
J. Biomed. Informatics7
2023 Motion robust magnetic resonance imaging via efficient Fourier aggregation
Oren Solomon, Rémi Patriat, Henry Braun, Tara E. Palnitkar, Steen Moeller, Edward J. Auerbach, Kâmil Ugurbil, Guillermo Sapiro, Noam Harel
Medical Image Anal.8
2023 Exploring Complexity of Facial Dynamics in Autism Spectrum Disorder
abstract
Atypical facial expression is one of the early symptoms of autism spectrum disorder (ASD) characterized by reduced regularity and lack of coordination of facial movements. Automatic quantification of these behaviors can offer novel biomarkers for screening, diagnosis, and treatment monitoring of ASD. In this work, 40 toddlers with ASD and 396 typically developing toddlers were shown developmentally-appropriate and engaging movies presented on a smart tablet during a well-child pediatric visit. The movies consisted of social and non-social dynamic scenes designed to evoke certain behavioral and affective responses. The front-facing camera of the tablet was used to capture the toddlers' face. Facial landmarks' dynamics were then automatically computed using computer vision algorithms. Subsequently, the complexity of the landmarks' dynamics was estimated for the eyebrows and mouth regions using multiscale entropy. Compared to typically developing toddlers, toddlers with ASD showed higher complexity (i.e., less predictability) in these landmarks' dynamics. This complexity in facial dynamics contained novel information not captured by traditional facial affect analyses. These results suggest that computer vision analysis of facial landmark movements is a promising approach for detecting and quantifying early behavioral symptoms associated with ASD.
Pradeep Raj Krishnappa Babu, Matías Di Martino, Zhuoqing Chang, Sam Perochon, Kimberly L. H. Carpenter, Scott Compton, Steven Espinosa, Geraldine Dawson, Guillermo Sapiro
IEEE Trans. Affect. Comput.9
2022 Scaling-Translation-Equivariant Networks with Decomposed Convolutional Filters
abstract
Encoding the scale information explicitly into the representation learned by a convolutional neural network (CNN) is beneficial for many computer vision tasks especially when dealing with multiscale inputs. We study, in this paper, a scaling-translation-equivariant ($\mathcal{ST}$-equivariant) CNN with joint convolutions across the space and the scaling group, which is shown to be both sufficient and necessary to achieve equivariance for the regular representation of the scaling-translation group $\mathcal{ST}$. To reduce the model complexity and computational burden, we decompose the convolutional filters under two pre-fixed separable bases and truncate the expansion to low-frequency components. A further benefit of the truncated filter expansion is the improved deformation robustness of the equivariant representation, a property which is theoretically analyzed and empirically verified. Numerical experiments demonstrate that the proposed scaling-translation-equivariant network with decomposed convolutional filters (ScDCFNet) achieves significantly improved performance in multiscale image classification and better interpretability than regular CNNs at a reduced model size.
Wei Zhu 0007, Qiang Qiu 0001, A. Robert Calderbank, Guillermo Sapiro, Xiuyuan Cheng
J. Mach. Learn. Res.4
2021 Nested Learning for Multi-Level Classification
abstract
Deep neural networks models are generally designed and trained for a specific type and quality of data. In this work, we address this problem in the context of nested learning. For many applications, both the input data, at training and testing, and the prediction can be conceived at multiple nested quality/resolutions. We show that by leveraging this multi-scale information, the problem of poor generalization and prediction overconfidence, as well as the exploitation of multiple training data quality, can be efficiently addressed. We evaluate the proposed ideas in six public datasets: MNIST, Fashion-MNIST, CIFAR10, CIFAR100, Plantvillage, and DBPEDIA. We observe that coarsely annotated data can help to solve fine predictions and reduce overconfidence significantly. We also show that hierarchical learning produces models intrinsically more robust to adversarial attacks and data perturbations.
Raphaël Achddou, Matías Di Martino, Guillermo Sapiro
ICASSP3
2021 Blind Pareto Fairness and Subgroup Robustness
abstract
Much of the work in the field of group fairness addresses disparities between predefined groups based on protected features such as gender, age, and race, which need to be available at train, and often also at test, time. These approaches are static and retrospective, since algorithms designed to protect groups identified a priori cannot anticipate and protect the needs of different at-risk groups in the future. In this work we analyze the space of solutions for worst-case fairness beyond demographics, and propose Blind Pareto Fairness (BPF), a method that leverages no-regret dynamics to recover a fair minimax classifier that reduces worst-case risk of any potential subgroup of sufficient size, and guarantees that the remaining population receives the best possible level of service. BPF addresses fairness beyond demographics, that is, it does not rely on predefined notions of at-risk groups, neither at train nor at test time. Our experimental results show that the proposed framework improves worst-case risk in multiple standard datasets, while simultaneously providing better levels of service for the remaining population. The code is available at github.com/natalialmg/BlindParetoFairness
Natalia Martínez, Martín Bertrán, Afroditi Papadaki, Miguel R. D. Rodrigues, Guillermo Sapiro
ICML5
2021 Cirrus: A Long-range Bi-pattern LiDAR Dataset
abstract
In this paper, we introduce Cirrus, a new long-range bi-pattern LiDAR public dataset for autonomous driving tasks such as 3D object detection, critical to highway driving and timely decision making. Our platform is equipped with a high-resolution video camera and a pair of LiDAR sensors with a 250-meter effective range, which is significantly longer than existing public datasets. We record paired point clouds simultaneously using both Gaussian and uniform scanning patterns. Point density varies significantly across such a long range, and different scanning patterns further diversify object representation in LiDAR. In Cirrus, eight categories of objects are exhaustively annotated in the LiDAR point clouds for the entire effective range. To illustrate the kind of studies supported by this new dataset, we introduce LiDAR model adaptation across different ranges, scanning patterns, and sensor devices. Promising results show the great potential of this new dataset to the robotics and computer vision communities.
Ze Wang 0008, Sihao Ding 0002, Ying Li 0139, Jonas Fenn, Sohini Roychowdhury, Andreas Wallin, Lane Martin, Scott Ryvola, Guillermo Sapiro, Qiang Qiu 0001
ICRA9
2021 A Scalable Off-the-Shelf Framework for Measuring Patterns of Attention in Young Children and Its Application in Autism Spectrum Disorder
abstract
Autism spectrum disorder (ASD) is associated with deficits in the processing of social information and difficulties in social interaction, and individuals with ASD exhibit atypical attention and gaze. Traditionally, gaze studies have relied upon precise and constrained means of monitoring attention using expensive equipment in laboratories. In this work we develop a low-cost off-the-shelf alternative for measuring attention that can be used in natural settings. The head and iris positions of 104 16-31 months children, an age range appropriate for ASD screening and diagnosis, 22 of them diagnosed with ASD, were recorded using the front facing camera in an iPad while they watched on the device screen a movie displaying dynamic stimuli, social stimuli on the left and nonsocial stimuli on the right. The head and iris position were then automatically analyzed via computer vision algorithms to detect the direction of attention. Children in the ASD group paid less attention to the movie, showed less attention to the social as compared to the nonsocial stimuli, and often fixated their attention to one side of the screen. The proposed method provides a low-cost means of monitoring attention to properly designed stimuli, demonstrating that the integration of stimuli design and automatic response analysis results in the opportunity to use off-the-shelf cameras to assess behavioral biomarkers.
Matthieu Bovery, Geraldine Dawson, Jordan Hashemi, Guillermo Sapiro
IEEE Trans. Affect. Comput.4
2021 Computer Vision Analysis for Quantification of Autism Risk Behaviors
abstract
Observational behavior analysis plays a key role for the discovery and evaluation of risk markers for many neurodevelopmental disorders. Research on autism spectrum disorder (ASD) suggests that behavioral risk markers can be observed at 12 months of age or earlier, with diagnosis possible at 18 months. To date, these studies and evaluations involving observational analysis tend to rely heavily on clinical practitioners and specialists who have undergone intensive training to be able to reliably administer carefully designed behavioural-eliciting tasks, code the resulting behaviors, and interpret such behaviors. These methods are therefore extremely expensive, time-intensive, and are not easily scalable for large population or longitudinal observational analysis. We developed a self-contained, closed-loop, mobile application with movie stimuli designed to engage the child's attention and elicit specific behavioral and social responses, which are recorded with a mobile device camera and then analyzed via computer vision algorithms. Here, in addition to presenting this paradigm, we validate the system to measure engagement, name-call responses, and emotional responses of toddlers with and without ASD who were presented with the application. Additionally, we show examples of how the proposed framework can further risk marker research with fine-grained quantification of behaviors. The results suggest these objective and automatic methods can be considered to aid behavioral analysis, and can be suited for objective automatic analysis for future studies.
Jordan Hashemi, Geraldine Dawson, Kimberly L. H. Carpenter, Kathleen Campbell, Qiang Qiu 0001, Steven Espinosa, Samuel Marsan, Jeffrey P. Baker, Helen Link Egger, Guillermo Sapiro
IEEE Trans. Affect. Comput.10
2021 Rethinking Shape From Shading for Spoofing Detection
abstract
Spoofing attacks are critical threats to modern face recognition systems, and most common countermeasures exploit 2D texture features as they are easy to extract and deploy. 3D shape-based methods can substantially improve spoofing prevention, but extracting the 3D shape of the face often requires complex hardware such as a 3D scanner and expensive computation. Motivated by the classical shape-from-shading model, we propose to obtain 3D facial features that can be used to recognize the presence of an actual 3D face, without explicit shape reconstruction. Such shading-based 3D features are extracted highly efficiently from a pair of images captured under different illumination, e.g., two images captured with and without flash. Thus the proposed method provides a rich 3D geometrical representation at negligible computational cost and minimal to none additional hardware. A theoretical analysis is provided to support why such simple 3D features can effectively describe the presence of an actual 3D shape while avoiding complicated calibration steps or hardware setup. Experimental validation shows that the proposed method can produce state-of-the-art spoofing prevention and enhance existing texture-based solutions.
Matías Di Martino, Qiang Qiu 0001, Guillermo Sapiro
IEEE Trans. Image Process.3
2020 Detecting Adversarial Samples Using Influence Functions and Nearest Neighbors
abstract
Deep neural networks (DNNs) are notorious for their vulnerability to adversarial attacks, which are small perturbations added to their input images to mislead their prediction. Detection of adversarial examples is, therefore, a fundamental requirement for robust classification frameworks. In this work, we present a method for detecting such adversarial attacks, which is suitable for any pre-trained neural network classifier. We use influence functions to measure the impact of every training sample on the validation set data. From the influence scores, we find the most supportive training samples for any given validation example. A k-nearest neighbor (k-NN) model fitted on the DNN's activation layers is employed to search for the ranking of these supporting training samples. We observe that these samples are highly correlated with the nearest neighbors of the normal inputs, while this correlation is much weaker for adversarial inputs. We train an adversarial detector using the k-NN ranks and distances and show that it successfully distinguishes adversarial examples, getting state-of-the-art results on six attack methods with three datasets. Code is available at https://github.com/giladcohen/NNIF_adv_defense.
Gilad Cohen, Guillermo Sapiro, Raja Giryes
CVPR2
2020 Stochastic Conditional Generative Networks with Basis Decomposition
Ze Wang 0008, Xiuyuan Cheng, Guillermo Sapiro, Qiang Qiu 0001
ICLR3
2020 Minimax Pareto Fairness: A Multi Objective Perspective
abstract
In this work we formulate and formally characterize group fairness as a multi-objective optimization problem, where each sensitive group risk is a separate objective. We propose a fairness criterion where a classifier achieves minimax risk and is Pareto-efficient w.r.t. all groups, avoiding unnecessary harm, and can lead to the best zero-gap model if policy dictates so. We provide a simple optimization algorithm compatible with deep neural networks to satisfy these constraints. Since our method does not require test-time access to sensitive attributes, it can be applied to reduce worst-case classification errors between outcomes in unbalanced classification problems. We test the proposed methodology on real case-studies of predicting income, ICU patient mortality, skin lesions classification, and assessing credit risk, demonstrating how our framework compares favorably to other approaches.
Natalia Martínez, Martín Bertrán, Guillermo Sapiro
ICML3
2020 Instance-based Generalization in Reinforcement Learning
abstract
Agents trained via deep reinforcement learning (RL) routinely fail to generalize to unseen environments, even when these share the same underlying dynamics as the training levels. Understanding the generalization properties of RL is one of the challenges of modern machine learning. Towards this goal, we analyze policy learning in the context of Partially Observable Markov Decision Processes (POMDPs) and formalize the dynamics of training levels as instances. We prove that, independently of the exploration strategy, reusing instances introduces significant changes on the effective Markov dynamics the agent observes during training. Maximizing expected rewards impacts the learned belief state of the agent by inducing undesired instance-specific speed-running policies instead of generalizable ones, which are sub-optimal on the training set. We provide generalization bounds to the value gap in train and test environments based on the number of training instances, and use insights based on these to improve performance on unseen levels. We propose training a shared belief representation over an ensemble of specialized policies, from which we compute a consensus policy that is used for data collection, disallowing instance-specific exploitation. We experimentally validate our theory, observations, and the proposed computational solution over the CoinRun benchmark.
Martín Bertrán, Natalia Martínez, Mariano Phielipp, Guillermo Sapiro
NeurIPS4
2020 A Dictionary Approach to Domain-Invariant Learning in Deep Networks
abstract
In this paper, we consider domain-invariant deep learning by explicitly modeling domain shifts with only a small amount of domain-specific parameters in a Convolutional Neural Network (CNN). By exploiting the observation that a convolutional filter can be well approximated as a linear combination of a small set of dictionary atoms, we show for the first time, both empirically and theoretically, that domain shifts can be effectively handled by decomposing a convolutional layer into a domain-specific atom layer and a domain-shared coefficient layer, while both remain convolutional. An input channel will now first convolve spatially only with each respective domain-specific dictionary atom to ``absorb" domain variations, and then output channels are linearly combined using common decomposition coefficients trained to promote shared semantics across domains. We use toy examples, rigorous analysis, and real-world examples with diverse datasets and architectures, to show the proposed plug-in framework's effectiveness in cross and joint domain performance and domain adaptation. With the proposed architecture, we need only a small set of dictionary atoms to model each additional domain, which brings a negligible amount of additional parameters, typically a few hundred.
Ze Wang 0008, Xiuyuan Cheng, Guillermo Sapiro, Qiang Qiu 0001
NeurIPS3
2020 Differential 3D Facial Recognition: Adding 3D to Your State-of-the-Art 2D Method
abstract
Active illumination is a prominent complement to enhance 2D face recognition and make it more robust, e.g., to spoofing attacks and low-light conditions. In the present work we show that it is possible to adopt active illumination to enhance state-of-the-art 2D face recognition approaches with 3D features, while bypassing the complicated task of 3D reconstruction. The key idea is to project over the test face a high spatial frequency pattern, which allows us to simultaneously recover real 3D information plus a standard 2D facial image. Therefore, state-of-the-art 2D face recognition solution can be transparently applied, while from the high frequency component of the input image, complementary 3D facial features are extracted. Experimental results on ND-2006 dataset show that the proposed ideas can significantly boost face recognition performance and dramatically improve the robustness to spoofing attacks.
Matías Di Martino, Fernando Suzacq, Mauricio Delbracio, Qiang Qiu 0001, Guillermo Sapiro
IEEE Trans. Pattern Anal. Mach. Intell.5
2019 Non-Contact Photoplethysmogram and Instantaneous Heart Rate Estimation from Infrared Face Video
abstract
Extracting the instantaneous heart rate (iHR) from face videos has been well studied in recent years. It is well known that changes in skin color due to blood flow can be captured using conventional cameras. One of the main limitations of methods that rely on this principle is the need of an illumination source. Moreover, they have to be able to operate under different light conditions. One way to avoid these constraints is using infrared cameras, allowing the monitoring of iHR under low light conditions. In this work, we present a simple, principled signal extraction method that recovers the iHR from infrared face videos. We tested the procedure on 7 participants, for whom we recorded an electrocardiogram simultaneously with their infrared face video. We checked that the recovered signal matched the ground truth iHR, showing that infrared is a promising alternative to conventional video imaging for heart rate monitoring, especially in low light conditions. Code is available at https://github.com/natalialmg/IR_iHR.
Natalia Martínez, Martín Bertrán, Guillermo Sapiro, Hau-Tieng Wu
ICIP3
2019 RotDCF: Decomposition of Convolutional Filters for Rotation-Equivariant Deep Networks
Xiuyuan Cheng, Qiang Qiu 0001, A. Robert Calderbank, Guillermo Sapiro
ICLR (Poster)4
2019 Adversarially Learned Representations for Information Obfuscation and Inference
abstract
Data collection and sharing are pervasive aspects of modern society. This process can either be voluntary, as in the case of a person taking a facial image to unlock his/her phone, or incidental, such as traffic cameras collecting videos on pedestrians. An undesirable side effect of these processes is that shared data can carry information about attributes that users might consider as sensitive, even when such information is of limited use for the task. It is therefore desirable for both data collectors and users to design procedures that minimize sensitive information leakage. Balancing the competing objectives of providing meaningful individualized service levels and inference while obfuscating sensitive information is still an open problem. In this work, we take an information theoretic approach that is implemented as an unconstrained adversarial game between Deep Neural Networks in a principled, data-driven manner. This approach enables us to learn domain-preserving stochastic transformations that maintain performance on existing algorithms while minimizing sensitive information leakage.
Martín Bertrán, Natalia Martínez, Afroditi Papadaki, Qiang Qiu 0001, Miguel R. D. Rodrigues, Galen Reeves, Guillermo Sapiro
ICML7
2018 Natural Human Exploration under Approach and Avoidance Motivation in a Real-Life Spatial Environment
Deeksha Malhotra, Kimberly Chiew, Mai-Anh Vu, Nicole Heller, Guillermo Sapiro, R. Alison Adcock
CogSci5
2018 OLÉ: Orthogonal Low-Rank Embedding - A Plug and Play Geometric Loss for Deep Learning
abstract
Deep neural networks trained using a softmax layer at the top and the cross-entropy loss are ubiquitous tools for image classification. Yet, this does not naturally enforce intra-class similarity nor inter-class margin of the learned deep representations. To simultaneously achieve these two goals, different solutions have been proposed in the literature, such as the pairwise or triplet losses. However, these carry the extra task of selecting pairs or triplets, and the extra computational burden of computing and learning for many combinations of them. In this paper, we propose a plug-and-play loss term for deep networks that explicitly reduces intra-class variance and enforces inter-class margin simultaneously, in a simple and elegant geometric manner. For each class, the deep features are collapsed into a learned linear subspace, or union of them, and inter-class subspaces are pushed to be as orthogonal as possible. Our proposed Orthogonal Low-rank Embedding (OLE) does not require carefully crafting pairs or triplets of samples for training, and works standalone as a classification loss, being the first reported deep metric learning framework of its kind. Because of the improved margin between features of different classes, the resulting deep networks generalize better, are more discriminative, and more robust. We demonstrate improved classification performance in general object recognition, plugging the proposed loss term into existing off-the-shelf architectures. In particular, we show the advantage of the proposed loss in the small data/model scenario, and we significantly advance the state-of-the-art on the Stanford STL-10 benchmark.
José Lezama, Qiang Qiu 0001, Pablo Musé, Guillermo Sapiro
CVPR4
2018 LDMNet: Low Dimensional Manifold Regularized Neural Networks
abstract
Deep neural networks have proved very successful on archetypal tasks for which large training sets are available, but when the training data are scarce, their performance suffers from overfitting. Many existing methods of reducing overfitting are data-independent. Data-dependent regularizations are mostly motivated by the observation that data of interest lie close to a manifold, which is typically hard to parametrize explicitly. These methods usually only focus on the geometry of the input data, and do not necessarily encourage the networks to produce geometrically meaningful features. To resolve this, we propose the Low-Dimensional-Manifold-regularized neural Network (LDMNet), which incorporates a feature regularization method that focuses on the geometry of both the input data and the output features. In LDMNet, we regularize the network by encouraging the combination of the input data and the output features to sample a collection of low dimensional manifolds, which are searched efficiently without explicit parametrization. To achieve this, we directly use the manifold dimension as a regularization term in a variational functional. The resulting Euler-Lagrange equation is a Laplace-Beltrami equation over a point cloud, which is solved by the point integral method without increasing the computational complexity. In the experiments, we show that LDMNet significantly outperforms widely-used regularizers. Moreover, LDMNet can extract common features of an object imaged via different modalities, which is very useful in real-world applications such as cross-spectral face recognition.
Wei Zhu 0007, Qiang Qiu 0001, Jiaji Huang, A. Robert Calderbank, Guillermo Sapiro, Ingrid Daubechies
CVPR5
2018 ForestHash: Semantic Hashing with Shallow Random Forests and Tiny Convolutional Networks
Qiang Qiu 0001, José Lezama, Alexander M. Bronstein, Guillermo Sapiro
ECCV (2)4
2018 A Practical Guide to Multi-Image Alignment
abstract
Multi - image alignment, bringing a group of images into common register, is an ubiquitous problem and the first step of many applications in a wide variety of domains. As a result, a great amount of effort is being invested in developing efficient multi-image alignment algorithms. Little has been done, however, to answer fundamental practical questions such as: what is the comparative performance of existing methods? is there still room for improvement? under which conditions should one technique be preferred over another? does adding more images or prior image information improve the registration results? In this work, we present a thorough analysis and evaluation of the main multi-image alignment methods which, combined with theoretical limits in multi-image alignment performance, allows us to organize them under a common framework and provide practical answers to these essential questions.
Cecilia Aguerrebere, Mauricio Delbracio, Alberto Bartesaghi, Guillermo Sapiro
ICASSP4
2018 Classifying Pump-Probe Images of Melanocytic Lesions Using the WEYL Transform
abstract
Diagnosis of melanoma is fraught with uncertainty, and discordance rates among physicians remain high because of the lack of a definitive criterion. Motivated by this challenge, this paper first introduces the Patch Weyl transform (PWT), a 2-dimensional variant of the Weyl transform. It then presents a method for classifying pump-probe images of melanocytic lesions based on the PWT coefficients. Performance of the PWT coefficients is shown to be superior to classification based on baseline intensity, on standard descriptors such as the Histogram of Oriented Gradients (HOG) and Local Binary Patterns (LBP), and on coefficients derived from PCA and Fourier representations of the data.
Hyun Keun Ahn, Qiang Qiu 0001, Edward Bosch, Andrew Thompson 0001, Francisco E. Robles, Guillermo Sapiro, Warren S. Warren, A. Robert Calderbank
ICASSP6
2018 The Learned Inexact Project Gradient Descent Algorithm
abstract
Accelerating iterative algorithms for solving inverse problems using neural networks have become a very popular strategy in the recent years. In this work, we propose a theoretical analysis that may provide an explanation for its success. Our theory relies on the usage of inexact projections with the projected gradient descent (PGD) method. It is demonstrated in various problems including image super-resolution.
Raja Giryes, Yonina C. Eldar, Alexander M. Bronstein, Guillermo Sapiro
ICASSP4
2018 DCFNet: Deep Neural Network with Decomposed Convolutional Filters
abstract
Filters in a Convolutional Neural Network (CNN) contain model parameters learned from enormous amounts of data. In this paper, we suggest to decompose convolutional filters in CNN as a truncated expansion with pre-fixed bases, namely the Decomposed Convolutional Filters network (DCFNet), where the expansion coefficients remain learned from data. Such a structure not only reduces the number of trainable parameters and computation, but also imposes filter regularity by bases truncation. Through extensive experiments, we consistently observe that DCFNet maintains accuracy for image classification tasks with a significant reduction of model parameters, particularly with Fourier-Bessel (FB) bases, and even with random bases. Theoretically, we analyze the representation stability of DCFNet with respect to input variations, and prove representation stability under generic assumptions on the expansion coefficients. The analysis is consistent with the empirical observations.
Qiang Qiu 0001, Xiuyuan Cheng, A. Robert Calderbank, Guillermo Sapiro
ICML4
2017 Generalization Error of Invariant Classifiers
abstract
This paper studies the generalization error of invariant classifiers. In particular, we consider the common scenario where the classification task is invariant to certain transformations of the input, and that the classifier is constructed (or learned) to be invariant to these transformations. Our approach relies on factoring the input space into a product of a base space and a set of transformations. We show that whereas the generalization error of a non-invariant classifier is proportional to the complexity of the input space, the generalization error of an invariant classifier is proportional to the complexity of the base space. We also derive a set of sufficient conditions on the geometry of the base space and the set of transformations that ensure that the complexity of the base space is much smaller than the complexity of the input space. Our analysis applies to general classifiers such as convolutional neural networks. We demonstrate the implications of the developed theory for such classifiers with experiments on the MNIST and CIFAR-10 datasets.
Jure Sokolic, Raja Giryes, Guillermo Sapiro, Miguel R. D. Rodrigues
AISTATS3
2017 Not Afraid of the Dark: NIR-VIS Face Recognition via Cross-Spectral Hallucination and Low-Rank Embedding
abstract
Surveillance cameras today often capture NIR (near infrared) images in low-light environments. However, most face datasets accessible for training and verification are only collected in the VIS (visible light) spectrum. It remains a challenging problem to match NIR to VIS face images due to the different light spectrum. Recently, breakthroughs have been made for VIS face recognition by applying deep learning on a huge amount of labeled VIS face samples. The same deep learning approach cannot be simply applied to NIR face recognition for two main reasons: First, much limited NIR face images are available for training compared to the VIS spectrum. Second, face galleries to be matched are mostly available only in the VIS spectrum. In this paper, we propose an approach to extend the deep learning breakthrough for VIS face recognition to the NIR spectrum, without retraining the underlying deep models that see only VIS faces. Our approach consists of two core components, cross-spectral hallucination and low-rank embedding, to optimize respectively input and output of a VIS deep model for cross-spectral face recognition. Cross-spectral hallucination produces VIS faces from NIR images through a deep learning approach. Low-rank embedding restores a low-rank structure for faces deep features across both NIR and VIS spectrum. We observe that it is often equally effective to perform hallucination to input NIR images or low-rank embedding to output deep features for a VIS deep model for cross-spectral recognition. When hallucination and low-rank embedding are deployed together, we observe significant further improvement, we obtain state-of-the-art accuracy on the CASIA NIR-VIS v2.0 benchmark, without the need at all to re-train the recognition system.
José Lezama, Qiang Qiu 0001, Guillermo Sapiro
CVPR3
2017 Deep Video Deblurring for Hand-Held Cameras
abstract
Motion blur from camera shake is a major problem in videos captured by hand-held devices. Unlike single-image deblurring, video-based approaches can take advantage of the abundant information that exists across neighboring frames. As a result the best performing methods rely on the alignment of nearby frames. However, aligning images is a computationally expensive and fragile procedure, and methods that aggregate information must therefore be able to identify which regions have been accurately aligned and which have not, a task that requires high level scene understanding. In this work, we introduce a deep learning solution to video deblurring, where a CNN is trained end-to-end to learn how to accumulate information across frames. To train this network, we collected a dataset of real videos recorded with a high frame rate camera, which we use to generate synthetic motion blur for supervision. We show that the features learned from this dataset extend to deblurring motion blur that arises due to camera shake in a wide range of videos, and compare the quality of results to a number of other baselines.
Shuochen Su, Mauricio Delbracio, Jue Wang 0001, Guillermo Sapiro, Wolfgang Heidrich, Oliver Wang
CVPR4
2017 Nonnegative Matrix Underapproximation for Robust Multiple Model Fitting
abstract
In this work, we introduce a highly efficient algorithm to address the nonnegative matrix underapproximation (NMU) problem, i.e., nonnegative matrix factorization (NMF) with an additional underapproximation constraint. NMU results are interesting as, compared to traditional NMF, they present additional sparsity and part-based behavior, explaining unique data features. To show these features in practice, we first present an application to the analysis of climate data. We then present an NMU-based algorithm to robustly fit multiple parametric models to a dataset. The proposed approach delivers state-of-the-art results for the estimation of multiple fundamental matrices and homographies, outperforming other alternatives in the literature and exemplifying the use of efficient NMU computations.
Mariano Tepper, Guillermo Sapiro
CVPR2
2017 Self-Learning Scene-Specific Pedestrian Detectors Using a Progressive Latent Model
abstract
In this paper, a self-learning approach is proposed towards solving scene-specific pedestrian detection problem without any human annotation involved. The self-learning approach is deployed as progressive steps of object discovery, object enforcement, and label propagation. In the learning procedure, object locations in each frame are treated as latent variables that are solved with a progressive latent model (PLM). Compared with conventional latent models, the proposed PLM incorporates a spatial regularization term to reduce ambiguities in object proposals and to enforce object localization, and also a graph-based label propagation to discover harder instances in adjacent frames. With the difference of convex (DC) objective functions, PLM can be efficiently optimized with a concave-convex programming and thus guaranteeing the stability of self-learning. Extensive experiments demonstrate that even without annotation the proposed self-learning approach outperforms weakly supervised learning approaches, while achieving comparable performance with transfer learning and fully supervised approaches.
Qixiang Ye, Tianliang Zhang 0003, Wei Ke 0003, Qiang Qiu 0001, Jie Chen 0001, Guillermo Sapiro, Baochang Zhang 0001
CVPR6
2017 A Sparse Bayesian Learning Algorithm for White Matter Parameter Estimation from Compressed Multi-shell Diffusion MRI
Pramod Kumar P., Stamatios N. Sotiropoulos, Guillermo Sapiro, Christophe Lenglet
MICCAI (1)3
2017 Probabilistic fluorescence-based synapse detection
abstract
Deeper exploration of the brain's vast synaptic networks will require new tools for high-throughput structural and molecular profiling of the diverse populations of synapses that compose those networks. Fluorescence microscopy (FM) and electron microscopy (EM) offer complementary advantages and disadvantages for single-synapse analysis. FM combines exquisite molecular discrimination capacities with high speed and low cost, but rigorous discrimination between synaptic and non-synaptic fluorescence signals is challenging. In contrast, EM remains the gold standard for reliable identification of a synapse, but offers only limited molecular discrimination and is slow and costly. To develop and test single-synapse image analysis methods, we have used datasets from conjugate array tomography (cAT), which provides voxel-conjugate FM and EM (annotated) images of the same individual synapses. We report a novel unsupervised probabilistic method for detection of synapses from multiplex FM (muxFM) image data, and evaluate this method both by comparison to EM gold standard annotated data and by examining its capacity to reproduce known important features of cortical synapse distributions. The proposed probabilistic model-based synapse detector accepts molecular-morphological synapse models as user queries, and delivers a volumetric map of the probability that each voxel represents part of a synapse. Taking human annotation of cAT EM data as ground truth, we show that our algorithm detects synapses from muxFM data alone as successfully as human annotators seeing only the muxFM data, and accurately reproduces known architectural features of cortical synapse distributions. This approach opens the door to data-driven discovery of new synapse types and their density. We suggest that our probabilistic synapse detector will also be useful for analysis of standard confocal and super-resolution FM images, where EM cross-validation is not practical.
Anish K. Simhal, Cecilia Aguerrebere, Forrest Collman, Joshua T. Vogelstein, Kristina D. Micheva, Richard J. Weinberg, Stephen J. Smith, Guillermo Sapiro
PLoS Comput. Biol.8
2016 A short-graph fourier transform via personalized pagerank vectors
abstract
The short-time Fourier transform (STFT) is widely used to analyze the spectra of temporal signals that vary through time. Signals defined over graphs, due to their intrinsic complexity, exhibit large variations in their patterns. In this work we propose a new formulation for an STFT for signals defined over graphs. This formulation draws on recent ideas from spectral graph theory, using personalized PageRank vectors as its fundamental building block. Furthermore, this work establishes and explores the connection between local spectral graph theory and localized spectral analysis of graph signals. We accompany the presentation with synthetic and real-world examples, showing the suitability of the proposed approach.
Mariano Tepper, Guillermo Sapiro
ICASSP2
2016 Dissimilarity-Based Sparse Subset Selection
abstract
Finding an informative subset of a large collection of data points or models is at the center of many problems in computer vision, recommender systems, bio/health informatics as well as image and natural language processing. Given pairwise dissimilarities between the elements of a 'source set' and a 'target set,' we consider the problem of finding a subset of the source set, called representatives or exemplars, that can efficiently describe the target set. We formulate the problem as a row-sparsity regularized trace minimization problem. Since the proposed formulation is, in general, NP-hard, we consider a convex relaxation. The solution of our optimization finds representatives and the assignment of each element of the target set to each representative, hence, obtaining a clustering. We analyze the solution of our proposed optimization as a function of the regularization parameter. We show that when the two sets jointly partition into multiple groups, our algorithm finds representatives from all groups and reveals clustering of the sets. In addition, we show that the proposed framework can effectively deal with outliers. Our algorithm works with arbitrary dissimilarities, which can be asymmetric or violate the triangle inequality. To efficiently implement our algorithm, we consider an Alternating Direction Method of Multipliers (ADMM) framework, which results in quadratic complexity in the problem size. We show that the ADMM implementation allows to parallelize the algorithm, hence further reducing the computational time. Finally, by experiments on real-world datasets, we show that our proposed algorithm improves the state of the art on the two problems of scene categorization using representative images and time-series modeling and segmentation using representative models.
Ehsan Elhamifar, Guillermo Sapiro, S. Shankar Sastry
IEEE Trans. Pattern Anal. Mach. Intell.2
2016 Graph Matching: Relax at Your Own Risk
abstract
Graph matching-aligning a pair of graphs to minimize their edge disagreements-has received wide-spread attention from both theoretical and applied communities over the past several decades, including combinatorics, computer vision, and connectomics. Its attention can be partially attributed to its computational difficulty. Although many heuristics have previously been proposed in the literature to approximately solve graph matching, very few have any theoretical support for their performance. A common technique is to relax the discrete problem to a continuous problem, therefore enabling practitioners to bring gradient-descent-type algorithms to bear. We prove that an indefinite relaxation (when solved exactly) almost always discovers the optimal permutation, while a common convex relaxation almost always fails to discover the optimal permutation. These theoretical results suggest that initializing the indefinite algorithm with the convex optimum might yield improved practical performance. Indeed, experimental results illuminate and corroborate these theoretical findings, demonstrating that excellent results are achieved in both benchmark and real data problems by amalgamating the two approaches.
Vince Lyzinski, Donniell E. Fishkind, Marcelo Fiori, Joshua T. Vogelstein, Carey E. Priebe, Guillermo Sapiro
IEEE Trans. Pattern Anal. Mach. Intell.6
2015 Low-Rank Spatio-Temporal Video Segmentation
abstract
Recently, a great deal of interest has been generated by the technique known as Robust Principle Component Analysis (RPCA) of Candes et al. [1], which addresses the problem of separating a matrix into a low-rank and a sparse component. This very general formulation can be used for tasks such as background estimation in videos and face recognition. In the case of background estimation, the low-rank matrix models the background, and the sparse matrix corresponds to the foreground. A considerable drawback of this approach is its poor robustness to local lighting conditions. If lighting conditions vary locally, one of two things may happen. Either the method incorporates the lighting variation into the foreground, which is clearly undesirable, or the rank of the background model is allowed to increase. Unfortunately, this second option means that the true foreground is likely to become included in the background, especially for objects which are static for a short while. Here, we propose to model the background as a piece-wise low-rank matrix. In this manner, it will be possible to extract several localised models which correspond to coherent lighting conditions. However, for this we need to segment the input video into such coherent regions. We refer to this problem as a low-rank spatio-temporal video segmentation. We present an algorithm to address this segmentation problem, based on region merging and spectral clustering techniques. We show that by carrying out a local RPCA in each region, the results of foreground/background separation are greatly improved, in comparison with both the standard RPCA and several other well-known background estimation techniques. Let X ∈ Rm×n represent an input video, in matrix form. Each frame contains m pixels, and there are a total of n frames in our video. The goal of RPCA is to decompose X as X≈ L+S, where L is the low-rank matrix and S is the sparse matrix. Unfortunately, the rank of a matrix is a non-convex function, so a surrogate function, the nuclear norm is used. Thus, the background/foreground separation problem may be formulated as follows:
Alasdair Newson, Mariano Tepper, Guillermo Sapiro
BMVC3
2015 From Local to Global Communities in Large Networks Through Consensus
Mariano Tepper, Guillermo Sapiro
CIARP2
2015 Burst deblurring: Removing camera shake through fourier burst accumulation
abstract
Numerous recent approaches attempt to remove image blur due to camera shake, either with one or multiple input images, by explicitly solving an inverse and inherently ill-posed deconvolution problem. If the photographer takes a burst of images, a modality available in virtually all modern digital cameras, we show that it is possible to combine them to get a clean sharp version. This is done without explicitly solving any blur estimation and subsequent inverse problem. The proposed algorithm is strikingly simple: it performs a weighted average in the Fourier domain, with weights depending on the Fourier spectrum magnitude. The method's rationale is that camera shake has a random nature and therefore each image in the burst is generally blurred differently. Experiments with real camera data show that the proposed Fourier Burst Accumulation algorithm achieves state-of-the-art results an order of magnitude faster, with simplicity for on-board implementation on camera phones.
Mauricio Delbracio, Guillermo Sapiro
CVPR2
2015 Alignment with intra-class structure can improve classification
abstract
High dimensional data is modeled using low-rank subspaces, and the probability of misclassification is expressed in terms of the principal angles between subspaces. The form taken by this expression motivates the design of a new feature extraction method that enlarges inter-class separation, while preserving intra-class structure. The method can be tuned to emphasize different features shared by members within the same class. Classification performance is compared to that of state-of-the-art methods on synthetic data and on the real face database. The probability of misclassification is decreased when intra-class structure is taken into account.
Jiaji Huang, Qiang Qiu 0001, A. Robert Calderbank, Miguel R. D. Rodrigues, Guillermo Sapiro
ICASSP5
2015 Geometry-Aware Deep Transform
abstract
Many recent efforts have been devoted to designing sophisticated deep learning structures, obtaining revolutionary results on benchmark datasets. The success of these deep learning methods mostly relies on an enormous volume of labeled training samples to learn a huge number of parameters in a network; therefore, understanding the generalization ability of a learned deep network cannot be overlooked, especially when restricted to a small training set, which is the case for many applications. In this paper, we propose a novel deep learning objective formulation that unifies both the classification and metric learning criteria. We then introduce a geometry-aware deep transform to enable a non-linear discriminative and robust feature transform, which shows competitive performance on small training sets for both synthetic and real-world data. We further support the proposed framework with a formal (K, ϵ)-robustness analysis.
Jiaji Huang, Qiang Qiu 0001, A. Robert Calderbank, Guillermo Sapiro
ICCV4
2015 Intel realsense = Real low cost gaze
abstract
Intel's newly-announced low-cost RealSense 3D camera claims significantly better precision than other currently available low-cost platforms and is expected to become ubiquitous in laptops and mobile devices starting this year. In this paper, we demonstrate for the first time that the RealSense camera can be easily converted into a real low-cost gaze tracker. Gaze has become increasingly relevant as an input for human-computer interaction due to its association with attention. It is also critical in clinical mental health diagnosis. We present a novel 3D gaze and fixation tracker based on the eye surface geometry captured with the RealSense 3D camera. First, eye surface 3D point clouds are segmented to extract the pupil center and iris using registered infrared images. With non-ellipsoid eye surface and single fixation point assumptions, pupil centers and iris normal vectors are used to first estimate gaze (for each eye), and then a single fixation point for both eyes simultaneously using a RANSAC-based approach. With a simple learned bias field correction model, the fixation tracker demonstrates mean error of approximately 1 cm at 20-30 cm, which is sufficiently adequate for gaze and fixation tracking in human-computer interaction and mental health diagnosis applications.
Mark Draelos, Qiang Qiu 0001, Alexander M. Bronstein, Guillermo Sapiro
ICIP4
2015 Cross-modality pose-invariant facial expression
abstract
In this work, we present a dictionary learning based framework for robust, cross-modality, and pose-invariant facial expression recognition. The proposed framework first learns a dictionary that i) contains both 3D shape and morphological information as well as 2D texture and geometric information, ii) enforces coherence across both 2D and 3D modalities and different poses, and iii) is robust in the sense that a learned dictionary can be applied across multiple facial expression datasets. We demonstrate that enforcing domain specific block structures on the dictionary, given a test expression sample, we can transform such sample across different domains for tasks such as pose alignment. We validate our approach on the task of pose-invariant facial expression recognition on the standard BU3D-FE and MultiPie datasets, achieving state of the art performance.
Jordan Hashemi, Qiang Qiu 0001, Guillermo Sapiro
ICIP3
2015 Clinical deep brain stimulation region prediction using regression forests from high-field MRI
abstract
This paper presents a prediction framework of brain subcortical structures which are invisible on clinical low-field MRI, learning detailed information from ultrahigh-field MR training data. Volumetric segmentation of Deep Brain Stimulation (DBS) structures within the Basal ganglia is a prerequisite process for reliable DBS surgery. While ultrahigh-field MR imaging (7 Tesla) allows direct visualization of DBS targeting structures, such ultrahigh-fields are not always clinically available, and therefore the relevant structures need to be predicted from the clinical data. We address the shape prediction problem with a regression forest, non-linearly mapping predictors to target structures with high confidence, exploiting ultrahigh-field MR training data. We consider an application for the subthalamic nucleus (STN) prediction as a crucial DBS target. Experimental results on Parkinson's patients validate that the proposed approach enables reliable estimation of the STN from clinical 1.5T MRI.
Yuval Duchin, Guillermo Sapiro, Jerrold L. Vitek, Noam Harel
ICIP3
2015 Multi-temporal foreground detection in videos
abstract
A common task in video processing is the binary separation of a video's content into either background or moving foreground. However, many situations require a foreground analysis with a finer temporal granularity, in particular for objects or people which remain immobile for a certain period of time. We propose an efficient method which detects foreground at different timescales, by exploiting the desirable theoretical and practical properties of Robust Principal Component Analysis. Our algorithm can be used in a variety of scenarios such as detecting people who have fallen in a video, or analysing the fluidity of road traffic, while avoiding costly computations needed for nearest neighbours searches or optical flow analysis. Finally, our algorithm has the useful ability to perform motion analysis without explicitly requiring computationally expensive motion estimation.
Mariano Tepper, Alasdair Newson, Pablo Sprechmann, Guillermo Sapiro
ICIP4
2015 Robust Prediction of Clinical Deep Brain Stimulation Target Structures via the Estimation of Influential High-Field MR Atlases
Yuval Duchin, Jerrold L. Vitek, Noam Harel, Guillermo Sapiro
MICCAI (2)6
2015 Sparse Bayesian Inference of White Matter Fiber Orientations from Compressed Multi-resolution Diffusion MRI
Pramod Kumar P., Julio Martin Duarte-Carvajalino, Stamatios N. Sotiropoulos, Guillermo Sapiro, Christophe Lenglet
MICCAI (1)4
2015 Discriminative Robust Transformation Learning
abstract
This paper proposes a framework for learning features that are robust to data variation, which is particularly important when only a limited number of trainingsamples are available. The framework makes it possible to tradeoff the discriminative value of learned features against the generalization error of the learning algorithm. Robustness is achieved by encouraging the transform that maps data to features to be a local isometry. This geometric property is shown to improve (K, \epsilon)-robustness, thereby providing theoretical justification for reductions in generalization error observed in experiments. The proposed optimization frameworkis used to train standard learning algorithms such as deep neural networks. Experimental results obtained on benchmark datasets, such as labeled faces in the wild,demonstrate the value of being able to balance discrimination and robustness.
Jiaji Huang, Qiang Qiu 0001, Guillermo Sapiro, A. Robert Calderbank
NIPS3
2015 Learning transformations for clustering and classification
Qiang Qiu 0001, Guillermo Sapiro
J. Mach. Learn. Res.2
2015 Learning Efficient Sparse and Low Rank Models
abstract
Parsimony, including sparsity and low rank, has been shown to successfully model data in numerous machine learning and signal processing tasks. Traditionally, such modeling approaches rely on an iterative algorithm that minimizes an objective function with parsimony-promoting terms. The inherently sequential structure and data-dependent complexity and latency of iterative optimization constitute a major limitation in many applications requiring real-time performance or involving large-scale data. Another limitation encountered by these modeling techniques is the difficulty of their inclusion in discriminative learning scenarios. In this work, we propose to move the emphasis from the model to the pursuit algorithm, and develop a process-centric view of parsimonious modeling, in which a learned deterministic fixed-complexity pursuit process is used in lieu of iterative optimization. We show a principled way to construct learnable pursuit process architectures for structured sparse and robust low rank models, derived from the iteration of proximal descent algorithms. These architectures learn to approximate the exact parsimonious representation at a fraction of the complexity of the standard optimization methods. We also show that appropriate training regimes allow to naturally extend parsimonious models to discriminative settings. State-of-the-art results are demonstrated on several challenging problems in image and audio processing with several orders of magnitude speed-up compared to the exact optimization algorithms.
Pablo Sprechmann, Alexander M. Bronstein, Guillermo Sapiro
IEEE Trans. Pattern Anal. Mach. Intell.3
2015 Removing Camera Shake via Weighted Fourier Burst Accumulation
abstract
Numerous recent approaches attempt to remove image blur due to camera shake, either with one or multiple input images, by explicitly solving an inverse and inherently ill-posed deconvolution problem. If the photographer takes a burst of images, a modality available in virtually all modern digital cameras, we show that it is possible to combine them to get a clean sharp version. This is done without explicitly solving any blur estimation and subsequent inverse problem. The proposed algorithm is strikingly simple: it performs a weighted average in the Fourier domain, with weights depending on the Fourier spectrum magnitude. The method can be seen as a generalization of the align and average procedure, with a weighted average, motivated by hand-shake physiology and theoretically supported, taking place in the Fourier domain. The method's rationale is that camera shake has a random nature, and therefore, each image in the burst is generally blurred differently. Experiments with real camera data, and extensive comparisons, show that the proposed Fourier burst accumulation algorithm achieves state-of-the-art results an order of magnitude faster, with simplicity for on-board implementation on camera phones. Finally, we also present experiments in real high dynamic range (HDR) scenes, showing how the method can be straightforwardly extended to HDR photography.
Mauricio Delbracio, Guillermo Sapiro
IEEE Trans. Image Process.2
2015 Compressive Sensing by Learning a Gaussian Mixture Model From Measurements
abstract
Compressive sensing of signals drawn from a Gaussian mixture model (GMM) admits closed-form minimum mean squared error reconstruction from incomplete linear measurements. An accurate GMM signal model is usually not available a priori, because it is difficult to obtain training signals that match the statistics of the signals being sensed. We propose to solve that problem by learning the signal model in situ, based directly on the compressive measurements of the signals, without resorting to other signals to train a model. A key feature of our method is that the signals being sensed are treated as random variables and are integrated out in the likelihood. We derive a maximum marginal likelihood estimator (MMLE) that maximizes the likelihood of the GMM of the underlying signals given only their linear compressive measurements. We extend the MMLE to a GMM with dominantly low-rank covariance matrices, to gain computational speedup. We report extensive experimental results on image inpainting, compressive sensing of high-speed video, and compressive hyperspectral imaging (the latter two based on real compressive cameras). The results demonstrate that the proposed methods outperform state-of-the-art methods by significant margins.
Jianbo Yang, Xuejun Liao, Xin Yuan 0002, Patrick Llull, David J. Brady, Guillermo Sapiro, Lawrence Carin
IEEE Trans. Image Process.6
2014 Low-Cost Compressive Sensing for Color Video and Depth
abstract
A simple and inexpensive (low-power and low-bandwidth) modification is made to a conventional off-the-shelf color video camera, from which we recover multiple color frames for each of the original measured frames, and each of the recovered frames can be focused at a different depth. The recovery of multiple frames for each measured frame is made possible via high-speed coding, manifested via translation of a single coded aperture, the inexpensive translation is constituted by mounting the binary code on a piezoelectric device. To simultaneously recover depth information, a liquid lens is modulated at high speed, via a variable voltage. Consequently, during the aforementioned coding process, the liquid lens allows the camera to sweep the focus through multiple depths. In addition to designing and implementing the camera, fast recovery is achieved by an anytime algorithm exploiting the group-sparsity of wavelet/DCT coefficients.
Xin Yuan 0002, Patrick Llull, Xuejun Liao, Jianbo Yang, David J. Brady, Guillermo Sapiro, Lawrence Carin
CVPR6
2014 Questionnaire simplification for fast risk analysis of children's mental health
abstract
Early detection and treatment of psychiatric disorders on children has shown significant impact in their subsequent development and quality of life. The assessment of psychopathology in childhood is commonly carried out by performing long comprehensive interviews such as the widely used Preschool Age Psychiatric Assessment (PAPA). Unfortunately, the time required to complete a full interview is too long to apply it at the scale of the actual population at risk, and most of the population goes undiagnosed or is diagnosed significantly later than desired. In this work, we aim to learn from unique and very rich previously collected PAPA examples the inter-correlations between different questions in order to provide a reliable risk analysis in the form of a much shorter interview. This helps to put such important risk analysis at the hands of regular practitioners, including teachers and family doctors. We use for this purpose the alternating decision trees algorithm, which combines decision trees with boosting to produce small and interpretable decision rules. Rather than a binary prediction, the algorithm provides a measure of confidence in the classification outcome. This is highly desirable from a clinical perspective, where it is preferable to abstain a decision on the low-confidence cases and recommend further screening. In order to prevent over-fitting, we propose to use network inference analysis to predefine a set of candidate question with consistent high correlation with the diagnosis. We report encouraging results with high levels of prediction using two independently collected datasets. The length and accuracy of the developed method suggests that it could be a valuable tool for preliminary evaluation in everyday care.
Kimberly L. H. Carpenter, Pablo Sprechmann, Marcelo Fiori, A. Robert Calderbank, Helen Link Egger, Guillermo Sapiro
ICASSP6
2014 All for one, one for all: Consensus community detection in networks
abstract
Given an universe of distinct, low-level communities of a network, we aim at identifying the “meaningful” and consistent communities in this universe. We address this as the process of obtaining consensual community detections and formalize it as a bi-clustering problem. While most consensus algorithms only take into account pairwise relations and end up analyzing a huge matrix, our proposed characterization of the consensus problem (1) does not drop useful information, and (2) analyzes a much smaller matrix, rendering the problem tractable for large networks. We also propose a new parameterless bi-clustering algorithm, fit for the type of matrices we analyze. The approach has proven successful in a very diverse set of experiments, ranging from unifying the results of multiple community detection algorithms to finding common communities from multi-modal or noisy networks.
Mariano Tepper, Guillermo Sapiro
ICASSP2
2014 Learning Transformations
abstract
A low-rank transformation learning framework for subspace clustering and classification is here proposed. Many high-dimensional data, such as face images and motion sequences, approximately lie in a union of low-dimensional subspaces. The corresponding subspace clustering problem has been extensively studied in the literature, partitioning such high-dimensional data into clusters corresponding to their underlying low-dimensional subspaces. However, low-dimensional intrinsic structures are often violated for real-world observations, as they can be corrupted by errors or deviate from ideal models. We propose to address this by learning a linear transformation on subspaces using matrix rank, via its convex surrogate nuclear norm, as the optimization criteria. The learned linear transformation restores a low-rank structure for data from the same subspace, and, at the same time, forces a high-rank structure for data from different subspaces. In this way, we reduce variations within the subspaces, and increase separation between the subspaces for improved subspace clustering and classification.
Qiang Qiu 0001, Guillermo Sapiro
ICIP2
2014 Learning compressed image classification features
abstract
Learning a transformation-based dimension reduction, thereby compressive, technique for classification is here proposed. High-dimensional data often approximately lie in a union of low-dimensional subspaces. We propose to perform dimension reduction by learning a “fat” linear transformation matrix on subspaces using nuclear norm as the optimization criteria. The learned transformation enables dimension reduction, and, at the same time, restores a low-rank structure for data from the same class and maximizes the separation between different classes, thereby improving classification via learned low-dimensional features. Theoretical and experimental results support the proposed framework, which can be interpreted as learning compressing sensing matrices for classification.
Qiang Qiu 0001, Guillermo Sapiro
ICIP2
2014 Intersecting 2D lines: A simple method for detecting vanishing points
abstract
We present a simple and powerful technique for testing with a prescribed precision whether a set of 2D lines meet at a given point. The method is based on a probabilistic framework and has a fundamental geometric interpretation. We use this technique for detecting vanishing points in images. We developed a very simple algorithm that yields state-of-the-art results at a much lower computational cost than its competitors. The presentation of the proposed formulation is complemented with numerous examples.
Mariano Tepper, Guillermo Sapiro
ICIP2
2014 A Complete System for Candidate Polyps Detection in Virtual Colonoscopy
abstract
We present a computer-aided detection pipeline for polyp detection in Computer tomographic colonography. The first stage of the pipeline consists of a simple colon segmentation technique that enhances polyps, which is followed by an adaptive-scale candidate polyp delineation, in order to capture the appropriate polyp size. In the last step, candidates are classified based on new texture and geometric features that consider both the information in the candidate polyp location and its immediate surrounding area. The system is tested with ground truth data, including flat and small polyps which are hard to detect even with optical colonoscopy. We achieve 100% sensitivity for polyps larger than 6 mm in size with just 0.9 false positives per case, and 93% sensitivity with 2.8 false positives per case for polyps larger than 3 mm in size.
Marcelo Fiori, Pablo Musé, Guillermo Sapiro
Int. J. Pattern Recognit. Artif. Intell.3
2014 A Biclustering Framework for Consensus Problems
abstract
We consider grouping as a general characterization for problems such as clustering, community detection in networks, and multiple parametric model estimation. We are interested in merging solutions from different grouping algorithms, distilling all their good qualities into a consensus solution. In this paper, we propose a biclustering framework and perspective for reaching consensus in such grouping problems. In particular, this is the first time that the task of finding/fitting multiple parametric models to a dataset is formally posed as a consensus problem. We highlight the equivalence of these tasks and establish the connection with the computational Gestalt program, which seeks to provide a psychologically inspired detection theory for visual events. We also present a simple but powerful biclustering algorithm, specially tuned to the nature of the problem we address, though general enough to handle many different instances inscribed within our characterization. The presentation is accompanied with diverse and extensive experimental results in clustering, community detection, and multiple parametric model estimation in image processing applications.
Mariano Tepper, Guillermo Sapiro
SIAM J. Imaging Sci.2
2014 Video Compressive Sensing Using Gaussian Mixture Models
abstract
A Gaussian mixture model (GMM)-based algorithm is proposed for video reconstruction from temporally compressed video measurements. The GMM is used to model spatio-temporal video patches, and the reconstruction can be efficiently computed based on analytic expressions. The GMM-based inversion method benefits from online adaptive learning and parallel computation. We demonstrate the efficacy of the proposed inversion method with videos reconstructed from simulated compressive video measurements, and from a real compressive video camera. We also use the GMM as a tool to investigate adaptive video compressive sensing, i.e., adaptive rate of temporal compression.
Jianbo Yang, Xin Yuan 0002, Xuejun Liao, Patrick Llull, David J. Brady, Guillermo Sapiro, Lawrence Carin
IEEE Trans. Image Process.6
2014 Semiautomatic Segmentation of Brain Subcortical Structures From High-Field MRI
abstract
Volumetric segmentation of subcortical structures, such as the basal ganglia and thalamus, is necessary for noninvasive diagnosis and neurosurgery planning. This is a challenging problem due in part to limited boundary information between structures, similar intensity profiles across the different structures, and low contrast data. This paper presents a semiautomatic segmentation system exploiting the superior image quality of ultrahigh field (7 T) MRI. The proposed approach utilizes the complementary edge information in the multiple structural MRI modalities. It combines optimally selected two modalities from susceptibility-weighted, T2-weighted, and diffusion MRI, and introduces a tailored new edge indicator function. In addition to this, we employ prior shape and configuration knowledge of the subcortical structures in order to guide the evolution of geometric active surfaces. Neighboring structures are segmented iteratively, constraining oversegmentation at their borders with a nonoverlapping penalty. Several experiments with data acquired on a 7 T MRI scanner demonstrate the feasibility and power of the approach for the segmentation of basal ganglia components critical for neurosurgery applications such as deep brain stimulation surgery.
Christophe Lenglet, Yuval Duchin, Guillermo Sapiro, Noam Harel
IEEE J. Biomed. Health Informatics4
2013 Reflective Symmetry Detection by Rectifying Randomized Correspondences
Zhongwei Tang, Mariano Tepper, Guillermo Sapiro
BMVC3
2013 Polyps Flagging in Virtual Colonoscopy
Marcelo Fiori, Pablo Musé, Guillermo Sapiro
CIARP (2)3
2013 Ants Crawling to Discover the Community Structure in Networks
Mariano Tepper, Guillermo Sapiro
CIARP (2)2
2013 Learnable low rank sparse models for speech denoising
abstract
In this paper we present a framework for real time enhancement of speech signals. Our method leverages a new process-centric approach for sparse and parsimonious models, where the representation pursuit is obtained applying a deterministic function or process rather than solving an optimization problem. We first propose a rank-regularized robust version of non-negative matrix factorization (NMF) for modeling time-frequency representations of speech signals in which the spectral frames are decomposed as sparse linear combinations of atoms of a low-rank dictionary. Then, a parametric family of pursuit processes is derived from the iteration of the proximal descent method for solving this model. We present several experiments showing successful results and the potential of the proposed framework. Incorporating discriminative learning makes the proposed method significantly outperform exact NMF algorithms, with fixed latency and at a fraction of it's computational complexity.
Pablo Sprechmann, Alexander M. Bronstein, Michael M. Bronstein, Guillermo Sapiro
ICASSP4
2013 Audio restoration from multiple copies
abstract
A method for removing impulse noise from audio signals by fusing multiple copies of the same recording is introduced in this paper. The proposed algorithm exploits the fact that while in general multiple copies of a given recording are available, all sharing the same master, most degradations in audio signals are record-dependent. Our method first seeks for the optimal non-rigid alignment of the signals that is robust to the presence of sparse outliers with arbitrary magnitude. Unlike previous approaches, we simultaneously find the optimal alignment of the signals and impulsive degradation. This is obtained via continuous dynamic time warping computed solving an Eikonal equation. We propose to use our approach in the derivative domain, reconstructing the signal by solving an inverse problem that resembles the Poisson image editing technique. The proposed framework is here illustrated and tested in the restoration of old gramophone recordings showing promising results; however, it can be used in other applications where different copies of the signal of interest are available and the degradations are copy-dependent.
Pablo Sprechmann, Alexander M. Bronstein, Jean-Michel Morel, Guillermo Sapiro
ICASSP4
2013 A Convex Optimization Framework for Active Learning
abstract
In many image/video/web classification problems, we have access to a large number of unlabeled samples. However, it is typically expensive and time consuming to obtain labels for the samples. Active learning is the problem of progressively selecting and annotating the most informative unlabeled samples, in order to obtain a high classification performance. Most existing active learning algorithms select only one sample at a time prior to retraining the classifier. Hence, they are computationally expensive and cannot take advantage of parallel labeling systems such as Mechanical Turk. On the other hand, algorithms that allow the selection of multiple samples prior to retraining the classifier, may select samples that have significant information overlap or they involve solving a non-convex optimization. More importantly, the majority of active learning algorithms are developed for a certain classifier type such as SVM. In this paper, we develop an efficient active learning framework based on convex programming, which can select multiple samples at a time for annotation. Unlike the state of the art, our algorithm can be used in conjunction with any type of classifiers, including those of the family of the recently proposed Sparse Representation-based Classification (SRC). We use the two principles of classifier uncertainty and sample diversity in order to guide the optimization program towards selecting the most informative unlabeled samples, which have the least information overlap. Our method can incorporate the data distribution in the selection process by using the appropriate dissimilarity between pairs of samples. We show the effectiveness of our framework in person detection, scene categorization and face recognition on real-world datasets.
Ehsan Elhamifar, Guillermo Sapiro, Allen Y. Yang, S. Shankar Sastry
ICCV2
2013 Fast L1 smoothing splines with an application to Kinect depth data
abstract
Splines are a popular and attractive way of smoothing noisy data. Computing splines involves minimizing a functional which is a linear combination of a fitting term and a regularization term. The former is classically computed using a (sometimes weighted) L2 norm while the latter ensures smoothness. In this work we propose to replace the L2 norm in the fitting term with an L1 norm, leading to automatic robustness to outliers. To solve the resulting minimization problem we propose an extremely simple and efficient numerical scheme based on split-Bregman iteration and a DCT-based filter. The algorithm is applied to the problem of smoothing and impainting range data, where high-quality results are obtained in short processing times.
Mariano Tepper, Guillermo Sapiro
ICIP2
2013 Gaussian mixture model for video compressive sensing
abstract
A Gaussian Mixture Model (GMM)-based algorithm is proposed for video reconstruction from temporal compressed measurements. The GMM is used to model spatio-temporal video patches, and the reconstruction can be efficiently computed based on analytic expressions. The developed GMM reconstruction method benefits from online adaptive learning and parallel computation. We demonstrate the efficacy of the proposed GMM with videos reconstructed from simulated compressive video measurements and from a real compressive video camera.
Jianbo Yang, Xin Yuan 0002, Xuejun Liao, Patrick Llull, Guillermo Sapiro, David J. Brady, Lawrence Carin
ICIP5
2013 Adaptive temporal compressive sensing for video
abstract
This paper introduces the concept of adaptive temporal compressive sensing (CS) for video. We propose a CS algorithm to adapt the compression ratio based on the scene's temporal complexity, computed from the compressed data, without compromising the quality of the reconstructed video. The temporal adaptivity is manifested by manipulating the integration time of the camera, opening the possibility to realtime implementation. The proposed algorithm is a generalized temporal CS approach that can be incorporated with a diverse set of existing hardware systems.
Xin Yuan 0002, Jianbo Yang, Patrick Llull, Xuejun Liao, Guillermo Sapiro, David J. Brady, Lawrence Carin
ICIP5
2013 Locating occupants in preschool classrooms using a multiple RGB-D sensor system
abstract
Presented are results demonstrating that, in developing a system with its first objective being the sustained detection of adults and young children as they move and interact in a normal preschool setting, the direct application of the straightforward RGB-D innovations presented here significantly outperforms even far more algorithmically advanced methods relying solely on images. The use of multiple RGB-D sensors by this project for depth-aware object localization economically resolves numerous issues regularly frustrating earlier vision-only detection and human surveillance methods, issues such as occlusions, illumination changes, unexpected postures, atypical morphologies, erratic or unanticipated motions, reflections, and misleading textures and colorations. This multiple RGB-D installation forms the front-end for a multi-step pipeline, the first portion of which seeks to isolate, in situ, 3D renderings of classroom occupants sufficient for a later analysis of their behaviors and interactions. Towards this end, a voxel-based approach to foreground/background separation and an effective adaptation of supervoxel clustering for 3D were developed, and 3D and image-only methods were tested and compared. The project's setting is highly challenging, but then so are its longer term goals: the automated detection of early childhood precursors, ofttimes very subtle, to a number of increasingly common developmental disorders.
Nicholas Walczak, Joshua Fasching, William D. Toczyski, Vassilios Morellas, Guillermo Sapiro, Nikolaos Papanikolopoulos
IROS5
2013 Robust Multimodal Graph Matching: Sparse Coding Meets Graph Matching
abstract
Graph matching is a challenging problem with very important applications in a wide range of fields, from image and video analysis to biological and biomedical problems. We propose a robust graph matching algorithm inspired in sparsity-related techniques. We cast the problem, resembling group or collaborative sparsity formulations, as a non-smooth convex optimization problem that can be efficiently solved using augmented Lagrangian techniques. The method can deal with weighted or unweighted graphs, as well as multimodal data, where different graphs represent different types of data. The proposed approach is also naturally integrated with collaborative graph inference techniques, solving general network inference problems where the observed variables, possibly coming from different modalities, are not in correspondence. The algorithm is tested and compared with state-of-the-art graph matching techniques in both synthetic and real graphs. We also present results on multimodal graphs and applications to collaborative inference of brain connectivity from alignment-free functional magnetic resonance imaging (fMRI) data.
Marcelo Fiori, Pablo Sprechmann, Joshua T. Vogelstein, Pablo Musé, Guillermo Sapiro
NIPS5
2013 Supervised Sparse Analysis and Synthesis Operators
abstract
In this paper, we propose a new and computationally efficient framework for learning sparse models. We formulate a unified approach that contains as particular cases models promoting sparse synthesis and analysis type of priors, and mixtures thereof. The supervised training of the proposed model is formulated as a bilevel optimization problem, in which the operators are optimized to achieve the best possible performance on a specific task, e.g., reconstruction or classification. By restricting the operators to be shift invariant, our approach can be thought as a way of learning analysis+synthesis sparsity-promoting convolutional operators. Leveraging recent ideas on fast trainable regressors designed to approximate exact sparse codes, we propose a way of constructing feed-forward neural networks capable of approximating the learned models at a fraction of the computational cost of exact solvers. In the shift-invariant case, this leads to a principled way of constructing task-specific convolutional networks. We illustrate the proposed models on several experiments in music analysis and image processing applications.
Pablo Sprechmann, Roee Litman, Tal Ben Yakar, Alexander M. Bronstein, Guillermo Sapiro
NIPS5
2013 Sparse Modeling of Intrinsic Correspondences
abstract
Abstract We present a novel sparse modeling approach to non‐rigid shape matching using only the ability to detect repeatable regions. As the input to our algorithm, we are given only two sets of regions in two shapes; no descriptors are provided so the correspondence between the regions is not know, nor we know how many regions correspond in the two shapes. We show that even with such scarce information, it is possible to establish very accurate correspondence between the shapes by using methods from the field of sparse modeling, being this, the first non‐trivial use of sparse models in shape correspondence. We formulate the problem ofpermuted sparse coding, in which we solve simultaneously for an unknown permutation ordering the regions on two shapes and for an unknown correspondence in functional representation. We also propose a robust variant capable of handling incomplete matches. Numerically, the problem is solved efficiently by alternating the solution of a linear assignment and a sparse coding problem. The proposed methods are evaluated qualitatively and quantitatively on standard benchmarks containing both synthetic and scanned objects.
Jonathan Pokrass, Alexander M. Bronstein, Michael M. Bronstein, Pablo Sprechmann, Guillermo Sapiro
Comput. Graph. Forum5
2013 Deep Learning with Hierarchical Convolutional Factor Analysis
abstract
Unsupervised multilayered (“deep”) models are considered for imagery. The model is represented using a hierarchical convolutional factor-analysis construction, with sparse factor loadings and scores. The computation of layer-dependent model parameters is implemented within a Bayesian setting, employing a Gibbs sampler and variational Bayesian (VB) analysis that explicitly exploit the convolutional nature of the expansion. To address large-scale and streaming data, an online version of VB is also developed. The number of dictionary elements at each layer is inferred from the data, based on a beta-Bernoulli implementation of the Indian buffet process. Example results are presented for several image-processing applications, with comparisons to related models in the literature.
Bo Chen 0001, Gungor Polatkan, Guillermo Sapiro, David M. Blei, David B. Dunson, Lawrence Carin
IEEE Trans. Pattern Anal. Mach. Intell.3
2012 If you are happy and you know it... tweet
abstract
Extracting sentiment from Twitter data is one of the fundamental problems in social media analytics. Twitter's length constraint renders determining the positive/negative sentiment of a tweet difficult, even for a human judge. In this work we present a general framework for per-tweet (in contrast with batches of tweets) sentiment analysis which consists of: (1) extracting tweets about a desired target subject, (2) separating tweets with sentiment, and (3) setting apart positive from negative tweets. For each step, we study the performance of a number of classical and new machine learning algorithms. We also show that the intrinsic sparsity of tweets allows performing classification in a low dimensional space, via random projections, without losing accuracy. In addition, we present weighted variants of all employed algorithms, exploiting the available labeling uncertainty, which further improve classification accuracy. Finally, we show that spatially aggregating our per-tweet classification results produces a very satisfactory outcome, making our approach a good candidate for batch tweet sentiment analysis.
Amir Asiaee T., Mariano Tepper, Arindam Banerjee 0001, Guillermo Sapiro
CIKM4
2012 See all by looking at a few: Sparse modeling for finding representative objects
abstract
We consider the problem of finding a few representatives for a dataset, i.e., a subset of data points that efficiently describes the entire dataset. We assume that each data point can be expressed as a linear combination of the representatives and formulate the problem of finding the representatives as a sparse multiple measurement vector problem. In our formulation, both the dictionary and the measurements are given by the data matrix, and the unknown sparse codes select the representatives via convex optimization. In general, we do not assume that the data are low-rank or distributed around cluster centers. When the data do come from a collection of low-rank models, we show that our method automatically selects a few representatives from each low-rank model. We also analyze the geometry of the representatives and discuss their relationship to the vertices of the convex hull of the data. We show that our framework can be extended to detect and reject outliers in datasets, and to efficiently deal with new observations and large datasets. The proposed framework and theoretical foundations are illustrated with examples in video summarization and image classification using representatives.
Ehsan Elhamifar, Guillermo Sapiro, René Vidal
CVPR2
2012 Adapted statistical compressive sensing: Learning to sense gaussian mixture models
abstract
A framework for learning sensing kernels adapted to signals that follow a Gaussian mixture model (GMM) is introduced in this paper. This follows the paradigm of statistical compressive sensing (SCS), where a statistical model, a GMM in particular, replaces the standard sparsity model of classical compressive sensing (CS), leading to both theoretical and practical improvements. We show that the optimized sensing matrix outperforms random sampling matrices originally exploited both in CS and SCS.
Julio Martin Duarte-Carvajalino, Guoshen Yu, Lawrence Carin, Guillermo Sapiro
ICASSP4
2012 Semi-supervised multi-domain regression with distinct training sets
abstract
We address the problems of multi-domain and single-domain regression based on distinct labeled training sets for each of the domains and a large unlabeled training set from all domains. We formulate these problems as ones of Bayesian estimation with partial knowledge of statistical relations. We propose a worst-case design strategy and study the resulting estimators. Our analysis explicitly accounts for the cardinality of the labeled sets and includes the special cases in which one of the labeled sets is very large or, in the other extreme, completely missing. We demonstrate our estimators in the context of audio-visual word recognition and provide comparisons to several recently proposed multi-modal learning algorithms.
Tomer Michaeli, Yonina C. Eldar, Guillermo Sapiro
ICASSP3
2012 LOw-rank data modeling via the minimum description length principle
abstract
Robust low-rank matrix estimation is a topic of increasing interest, with promising applications in a variety of fields, from computer vision to data mining and recommender systems. Recent theoretical results establish the ability of such data models to recover the true underlying low-rank matrix when a large portion of the measured matrix is either missing or arbitrarily corrupted. However, if low rank is not a hypothesis about the true nature of the data, but a device for extracting regularity from it, no current guidelines exist for choosing the rank of the estimated matrix. In this work we address this problem by means of the Minimum Description Length (MDL) principle - a well established information-theoretic approach to statistical inference - as a guideline for selecting a model for the data at hand. We demonstrate the practical usefulness of our formal approach with results for complex background extraction in video sequences.
Ignacio Ramírez, Guillermo Sapiro
ICASSP2
2012 Gaussian mixture models for score-informed instrument separation
abstract
A new framework for representing quasi-harmonic signals, and its application to score-informed single channel musical instruments separation, is introduced in this paper. In the proposed approach, the signal's pitch and spectral envelope are modeled separately. The model combines parametric filters enforcing an harmonic structure in the representation, with Gaussian modeling for representing the spectral envelope. The estimation of the signal's model is cast as an inverse problem efficiently solved via a maximum a posteriori expectation-maximization algorithm. The relation of the proposed framework with common non-negative factorization methods is also discussed. The algorithm is evaluated with both real and synthetic instruments mixtures, and comparisons with recently proposed techniques are presented.
Pablo Sprechmann, Pablo Cancela, Guillermo Sapiro
ICASSP3
2012 Decoupled coarse-to-fine matching and nonlinear regularization for efficient motion estimation
abstract
A simple motion estimation algorithm, light-weighted both in memory and in time, is presented in this paper. This simplicity is achieved by decoupling the matching and the regularization stages in the estimation process. Experiments show that the obtained results are comparable with state-of-the-art algorithms that are much more computationally demanding.
Mariano Tepper, Guillermo Sapiro
ICIP2
2012 Learning Efficient Structured Sparse Models
Alexander M. Bronstein, Pablo Sprechmann, Guillermo Sapiro
ICML3
2012 A multi-sensor visual tracking system for behavior monitoring of at-risk children
abstract
Clinical studies confirm that mental illnesses such as autism, Obsessive Compulsive Disorder (OCD), etc. show behavioral abnormalities even at very young ages; the early diagnosis of which can help steer effective treatments. Most often, the behavior of such at-risk children deviate in very subtle ways from that of a normal child; correct diagnosis of which requires prolonged and continuous monitoring of their activities by a clinician, which is a difficult and time intensive task. As a result, the development of automation tools for assisting in such monitoring activities will be an important step towards effective utilization of the diagnostic resources. In this paper, we approach the problem from a computer vision standpoint, and propose a novel system for the automatic monitoring of the behavior of children in their natural environment through the deployment of multiple non-invasive sensors (cameras and depth sensors). We provide details of our system, together with algorithms for the robust tracking of the activities of the children. Our experiments, conducted in the Shirley G. Moore Laboratory School, demonstrate the effectiveness of our methodology.
Ravishankar Sivalingam, Anoop Cherian, Joshua Fasching, Nicholas Walczak, Nathaniel D. Bird, Vassilios Morellas, Barbara Murphy, Kathryn Cullen, Kelvin O. Lim, Guillermo Sapiro, Nikolaos Papanikolopoulos
ICRA10
2012 Sparse modeling for hyperspectral imagery with LiDAR data fusion for subpixel mapping
abstract
Several studies suggest that the use of geometric features along with spectral information improves the classification and visualization quality of hyperspectral imagery. These studies normally make use of spatial neighborhoods of hyperspectral pixels for extracting these geometric features. In this work, we merge point cloud Light Detection and Ranging (LiDAR) data and hyperspectral imagery (HSI) into a single sparse modeling pipeline for subpixel mapping and classification. The model accounts for material variability and noise by using learned dictionaries that act as spectral endmembers. Additionally, the estimated abundances are influenced by the LiDAR point cloud density, particularly helpful in spectral mixtures involving partial occlusions and illumination changes caused by elevation differences. We demonstrate the advantages of the proposed algorithm with co-registered LiDAR-HSI data.
Alexey Castrodad, Timothy Khuon, Robert Rand 0002, Guillermo Sapiro
IGARSS4
2012 Detecting risk-markers in children in a preschool classroom
abstract
Early intervention in mental disorders can dramatically increase an individual's quality of life. Additionally, when symptoms of mental illness appear in childhood or adolescence, they represent the later stages of a process that began years earlier. One goal of psychiatric research is to identify risk-markers: genetic, neural, behavioral and/or social deviations that indicate elevated risk of a particular mental disorder. Ideally, screening of risk-markers should occur in a community setting, and not a clinical setting which may be time-consuming and resource-intensive. Given this situation, a system for automatically detecting risk-markers in children would be highly valuable. In this paper, we describe such a system that has been installed at the Shirley G. Moore Lab School, a research pre-school at the University of Minnesota. This system consists of multiple RGB+D sensors and is able to detect children and adults in the classroom, tracking them as they move around the room. We use the tracking results to extract high-level information about the behavior and social interaction of children, that can then be used to screen for early signs of mental disorders.
Joshua Fasching, Nicholas Walczak, Ravishankar Sivalingam, Kathryn Cullen, Barbara Murphy, Guillermo Sapiro, Vassilios Morellas, Nikolaos Papanikolopoulos
IROS6
2012 Finding Exemplars from Pairwise Dissimilarities via Simultaneous Sparse Recovery
abstract
Given pairwise dissimilarities between data points, we consider the problem of finding a subset of data points called representatives or exemplars that can efficiently describe the data collection. We formulate the problem as a row-sparsity regularized trace minimization problem which can be solved efficiently using convex programming. The solution of the proposed optimization program finds the representatives and the probability that each data point is associated to each one of the representatives. We obtain the range of the regularization parameter for which the solution of the proposed optimization program changes from selecting one representative to selecting all data points as the representatives. When data points are distributed around multiple clusters according to the dissimilarities, we show that the data in each cluster select only representatives from that cluster. Unlike metric-based methods, our algorithm does not require that the pairwise dissimilarities be metrics and can be applied to dissimilarities that are asymmetric or violate the triangle inequality. We demonstrate the effectiveness of the proposed algorithm on synthetic data as well as real-world datasets of images and text.
Ehsan Elhamifar, Guillermo Sapiro, René Vidal
NIPS2
2012 Topology Constraints in Graphical Models
abstract
Graphical models are a very useful tool to describe and understand natural phenomena, from gene expression to climate change and social interactions. The topological structure of these graphs/networks is a fundamental part of the analysis, and in many cases the main goal of the study. However, little work has been done on incorporating prior topological knowledge onto the estimation of the underlying graphical models from sample data. In this work we propose extensions to the basic joint regression model for network estimation, which explicitly incorporate graph-topological constraints into the corresponding optimization approach. The first proposed extension includes an eigenvector centrality constraint, thereby promoting this important prior topological property. The second developed extension promotes the formation of certain motifs, triangle-shaped ones in particular, which are known to exist for example in genetic regulatory networks. The presentation of the underlying formulations, which serve as examples of the introduction of topological constraints in network estimation, is complemented with examples in diverse datasets demonstrating the importance of incorporating such critical prior knowledge.
Marcelo Fiori, Pablo Musé, Guillermo Sapiro
NIPS3
2012 Kernelized Probabilistic Matrix Factorization: Exploiting Graphs and Side Information
abstract
We propose a new matrix completion algorithm—Kernelized Probabilistic Matrix Factorization (KPMF), which effectively incorporates external side information into the matrix factorization process. Unlike Probabilistic Matrix Factorization (PMF) [14], which assumes an independent latent vector for each row (and each column) with Gaussian priors, KMPF works with latent vectors spanning all rows (and columns) with Gaussian Process (GP) priors. Hence, KPMF explicitly captures the underlying (nonlinear) covariance structures across rows and columns. This crucial difference greatly boosts the performance of KPMF when appropriate side information, e.g., users' social network in recommender systems, is incorporated. Furthermore, GP priors allow the KPMF model to fill in a row that is entirely missing in the original matrix based on the side information alone, which is not feasible for standard PMF formulation. In our paper, we mainly work on the matrix completion problem with a graph among the rows and/or columns as side information, but the proposed framework can be easily used with other types of side information as well. Finally, we demonstrate the efficacy of KPMF through two different applications: 1) recommender systems and 2) image restoration.
Tinghui Zhou, Hanhuai Shan, Arindam Banerjee 0001, Guillermo Sapiro
SDM4
2012 A nonintrusive system for behavioral analysis of children using multiple RGB+depth sensors
abstract
In developmental disorders such as autism and schizophrenia, observing behavioral precursors in very early childhood can allow for early intervention and can improve patient outcomes. While such precursors open the possibility of broad and large-scale screening, until now they have been identified only through experts' painstaking examinations and their manual annotations of limited, unprocessed video footage. Here we introduce a system to automate and assist in such procedures. Employing multiple inexpensive real-time rgb+depth (rgb+d) sensors recording from multiple viewpoints, our non-invasive system - now installed at the Shirley G. Moore Lab School, a research preschool - is being developed to monitor and reconstruct the play and interactions of preschoolers. The system's role is to help in assessing the growing volumes of its on-site recordings and to provide the data needed to uncover additional neuromotor behavioral markers via techniques such as data mining.
Nicholas Walczak, Joshua Fasching, William D. Toczyski, Ravishankar Sivalingam, Nathaniel D. Bird, Kathryn Cullen, Vassilios Morellas, Barbara Murphy, Guillermo Sapiro, Nikolaos Papanikolopoulos
WACV9
2012 Sparse Modeling of Human Actions from Motion Imagery
Alexey Castrodad, Guillermo Sapiro
Int. J. Comput. Vis.2
2012 Dictionary Learning for Noisy and Incomplete Hyperspectral Images
abstract
We consider analysis of noisy and incomplete hyperspectral imagery, with the objective of removing the noise and inferring the missing data. The noise statistics may be wavelength dependent, and the fraction of data missing (at random) may be substantial, including potentially entire bands, offering the potential to significantly reduce the quantity of data that need be measured. To achieve this objective, the imagery is divided into contiguous three-dimensional (3D) spatio-spectral blocks of spatial dimension much less than the image dimension. It is assumed that each such 3D block may be represented as a linear combination of dictionary elements of the same dimension, plus noise, and the dictionary elements are learned in situ based on the observed data (no a priori training). The number of dictionary elements needed for representation of any particular block is typically small relative to the block dimensions, and all the image blocks are processed jointly (“collaboratively") to infer the underlying dictionary. We address dictionary learning from a Bayesian perspective, considering two distinct means of imposing sparse dictionary usage. These models allow inference of the number of dictionary elements needed as well as the underlying wavelength-dependent noise statistics. It is demonstrated that drawing the dictionary elements from a Gaussian process prior, imposing structure on the wavelength dependence of the dictionary elements, yields significant advantages, relative to the more conventional approach of using an independent and identically distributed Gaussian prior for the dictionary elements; this advantage is particularly evident in the presence of noise. The framework is demonstrated by processing hyperspectral imagery with a significant number of voxels missing uniformly at random, with imagery at specific wavelengths missing entirely, and in the presence of substantial additive noise.
Zhengming Xing, Mingyuan Zhou, Alexey Castrodad, Guillermo Sapiro, Lawrence Carin
SIAM J. Imaging Sci.4
2012 A Convex Model for Nonnegative Matrix Factorization and Dimensionality Reduction on Physical Space
abstract
A collaborative convex framework for factoring a data matrix X into a nonnegative product AS , with a sparse coefficient matrix S, is proposed. We restrict the columns of the dictionary matrix A to coincide with certain columns of the data matrix X, thereby guaranteeing a physically meaningful dictionary and dimensionality reduction. We use l(1, ∞) regularization to select the dictionary from the data and show that this leads to an exact convex relaxation of l(0) in the case of distinct noise-free data. We also show how to relax the restriction-to- X constraint by initializing an alternating minimization approach with the solution of the convex model, obtaining a dictionary close to but not necessarily in X. We focus on applications of the proposed framework to hyperspectral endmember and abundance identification and also show an application to blind source separation of nuclear magnetic resonance data.
Ernie Esser, Michael Möller 0001, Stanley J. Osher, Guillermo Sapiro, Jack Xin
IEEE Trans. Image Process.4
2012 Sparse Representations for Range Data Restoration
abstract
In this paper, the problem of denoising and occlusion restoration of 3-D range data based on dictionary learning and sparse representation methods is explored. We apply these techniques after converting the noisy 3-D surface into one or more images. We present experimental results on the proposed approaches.
Mona Mahmoudi, Guillermo Sapiro
IEEE Trans. Image Process.2
2012 Universal Regularizers for Robust Sparse Coding and Modeling
abstract
Sparse data models, where data is assumed to be well represented as a linear combination of a few elements from a dictionary, have gained considerable attention in recent years, and their use has led to state-of-the-art results in many signal and image processing tasks. It is now well understood that the choice of the sparsity regularization term is critical in the success of such models. Based on a codelength minimization interpretation of sparse coding, and using tools from universal coding theory, we propose a framework for designing sparsity regularization terms which have theoretical and practical advantages when compared with the more standard l(0) or l(1) ones. The presentation of the framework and theoretical foundations is complemented with examples that show its practical advantages in image denoising, zooming and classification.
Ignacio Ramírez, Guillermo Sapiro
IEEE Trans. Image Process.2
2012 Solving Inverse Problems With Piecewise Linear Estimators: From Gaussian Mixture Models to Structured Sparsity
abstract
A general framework for solving image inverse problems with piecewise linear estimations is introduced in this paper. The approach is based on Gaussian mixture models, which are estimated via a maximum a posteriori expectation-maximization algorithm. A dual mathematical interpretation of the proposed framework with a structured sparse estimation is described, which shows that the resulting piecewise linear estimate stabilizes the estimation when compared with traditional sparse inverse problem techniques. We demonstrate that, in a number of image inverse problems, including interpolation, zooming, and deblurring of narrow kernels, the same simple and computationally efficient algorithm yields results in the same ballpark as that of the state of the art.
Guoshen Yu, Guillermo Sapiro, Stéphane Mallat
IEEE Trans. Image Process.2
2012 Nonparametric Bayesian Dictionary Learning for Analysis of Noisy and Incomplete Images
abstract
Nonparametric Bayesian methods are considered for recovery of imagery based upon compressive, incomplete, and/or noisy measurements. A truncated beta-Bernoulli process is employed to infer an appropriate dictionary for the data under test and also for image recovery. In the context of compressive sensing, significant improvements in image recovery are manifested using learned dictionaries, relative to using standard orthonormal image expansions. The compressive-measurement projections are also optimized for the learned dictionary. Additionally, we consider simpler (incomplete) measurements, defined by measuring a subset of image pixels, uniformly selected at random. Spatial interrelationships within imagery are exploited through use of the Dirichlet and probit stick-breaking processes. Several example results are presented, with comparisons to other methods in the literature.
Mingyuan Zhou, Haojun Chen, John W. Paisley, Lingbo Li 0002, Zhengming Xing, David B. Dunson, Guillermo Sapiro, Lawrence Carin
IEEE Trans. Image Process.8
2012 Toward Multiple Catheters Detection in Fluoroscopic Image Guided Interventions
abstract
Catheters are routinely inserted via vessels to cavities of the heart during fluoroscopic image guided interventions for electrophysiology (EP) procedures such as ablation. During such interventions, the catheter undergoes nonrigid deformation due to physician interaction, patient's breathing, and cardiac motions. EP clinical applications can benefit from fast and accurate automatic catheter tracking in the fluoroscopic images. The typical low quality in fluoroscopic images and the presence of other medical instruments in the scene make the automatic detection and tracking of catheters in clinical environments very challenging. Toward the development of such an application, a robust and efficient method for detecting and tracking the catheter sheath is developed. The proposed approach exploits the clinical setup knowledge to constrain the search space while boosting both tracking speed and accuracy, and is based on a computationally efficient framework to trace the sheath and simultaneously detect one or multiple catheter tips. The algorithm is based on a modification of the fast marching weighted distance computation that efficiently calculates, on the fly, important geodesic properties in relevant regions of the image. This is followed by a cascade classifier for detecting the catheter tips. The proposed technique is validated on 1107 fluoroscopic images acquired on multiple patients across four different clinics, achieving multiple catheter tracking at a rate of 10 images/s with a very low false positive rate of 1.06.
Liron Yatziv, Mathieu Chartouni, Saurabh Datta, Guillermo Sapiro
IEEE Trans. Inf. Technol. Biomed.4
2011 Efficient matrix completion with Gaussian models
abstract
A general framework based on Gaussian models and a MAP-EM algorithm is introduced in this paper for solving matrix/table completion problems. The numerical experiments with the standard and challenging movie ratings data show that the proposed approach, based on probably one of the simplest probabilistic models, leads to the results in the same ballpark as the state-of-the-art, at a lower computational cost.
Flavien Léger, Guoshen Yu, Guillermo Sapiro
ICASSP3
2011 Sparse coding and dictionary learning based on the MDL principle
abstract
The power of sparse signal coding with learned overcomplete dictionaries has been demonstrated in a variety of applications and fields, from signal processing to statistical inference and machine learning. However, the statistical properties of these models, such as underfitting or overfitting given sets of data, are still not well characterized in the literature. This work aims at filling this gap by means of the Minimum Description Length (MDL) principle a well established information-theoretic approach to statistical inference. The resulting framework derives a family of efficient sparse coding and modeling (dictionary learning) algorithms, which by virtue of the MDL principle, are completely parameter free. Furthermore, such framework allows to incorporate additional prior information in the model, such as Markovian dependencies, in a natural way. We demonstrate the performance of the proposed framework with results for image de noising and classification tasks.
Ignacio Ramírez, Guillermo Sapiro
ICASSP2
2011 Collaborative sources identification in mixed signals via hierarchical sparse modeling
abstract
A collaborative framework for detecting the different sources in mixed signals is presented in this paper. The approach is based on C-HiLasso, a convex collaborative hierarchical sparse model, and proceeds as follows. First, we build a structured dictionary for mixed signals by concatenating a set of sub-dictionaries, each one of them learned to sparsely model one of a set of possible classes. Then, the coding of the mixed signal is performed by efficiently solving a convex optimization problem that combines standard sparsity with group and collaborative sparsity. The present sources are identified by looking at the sub-dictionaries automatically selected in the coding. The collaborative filtering in C-HiLasso takes advantage of the temporal/spatial redundancy in the mixed signals, letting collections of samples collaborate in identifying the classes, while allowing in dividualsamples to have different internal sparse representations. This collaboration is critical to further stabilize the sparse representation of signals, in particular the class/sub-dictionary selection. The internal sparsity inside the sub-dictionaries, as naturally incorporated by the hierarchical aspects of C-HiLasso, is critical to make the model consistent with the essence of the sub-dictionaries that have been trained for sparse representation of each individual class. We present applications from speaker and instrument identification and texture separation. In the case of audio signals, we use sparse modeling to describe the short-term power spectrum envelopes of harmonic sounds. The proposed pitch independent method automatically detects the number of sources on a recording.
Pablo Sprechmann, Ignacio Ramírez, Pablo Cancela, Guillermo Sapiro
ICASSP4
2011 Statistical compressive sensing of Gaussian mixture models
abstract
A new framework of compressive sensing (CS), namely statistical compressive sensing (SCS), that aims at efficiently sampling a collection of signals that follow a statistical distribution and achieving accurate reconstruction on average, is introduced. For signals following a Gaussian distribution, with Gaussian or Bernoulli sensing matrices of O(k) measurements, considerably smaller than the O(k log(N/k)) required by conventional CS, where N is the signal dimension, and with an optimal decoder implemented with linear filtering, significantly faster than the pursuit decoders applied in conventional CS, the error of SCS is shown tightly upper bounded by a constant times the best k-term approximation error, with overwhelming probability. The failure probability is also significantly smaller than that of conventional CS. Stronger yet simpler results further show that for any sensing matrix, the error of Gaussian SCS is upper bounded by a constant times the best k-term approximation with probability one, and the bound constant can be efficiently calculated. For signals following Gaussian mixture models, SCS with a piecewise linear decoder is introduced and shown to produce for real images better results than conventional CS based on sparse models.
Guoshen Yu, Guillermo Sapiro
ICASSP2
2011 Covariate-dependent dictionary learning and sparse coding
abstract
A dependent hierarchical beta process (dHBP) is developed as a prior for data that may be represented in terms of a sparse set of latent features (dictionary elements), with covariate dependent feature usage. The dHBP is applicable to general covariates and data models, imposing that signals with similar covariates are likely to be manifested in terms of similar features. As an application, we consider the simultaneous sparse modeling of multiple images, with the covariate of a given image linked to its similarity to all other images (as applied in manifold learning). Efficient inference is performed using hybrid Gibbs, Metropolis-Hastings and slice sampling.
Mingyuan Zhou, Hongxia Yang, Guillermo Sapiro, David B. Dunson, Lawrence Carin
ICASSP3
2011 Hierarchical invariant sparse modeling for image analysis
abstract
Sparse representation theory has been increasingly used in signal processing and machine learning. In this paper we introduce a hierarchical sparse modeling approach which integrates information from the image patch level to derive a mid-level invariant image and pattern representation. The proposed framework is based on a hierarchical architecture of dictionary learning for sparse coding in a cortical (log-polar) space, combined with a novel pooling operator which incorporates the Rapid transform and max pooling to attain rotation and scale invariance. The invariant sparse representation of patterns here presented - can be used in different object recognition tasks. Promising results are obtained for three applications - 2D shapes classification, texture recognition and object detection.
Leah Bar, Guillermo Sapiro
ICIP2
2011 The Hierarchical Beta Process for Convolutional Factor Analysis and Deep Learning
Bo Chen 0001, Gungor Polatkan, Guillermo Sapiro, David B. Dunson, Lawrence Carin
ICML3
2011 On the Integration of Topic Modeling and Dictionary Learning
Lingbo Li 0002, Mingyuan Zhou, Guillermo Sapiro, Lawrence Carin
ICML3
2011 A Variational Framework for Exemplar-Based Image Inpainting
Pablo Arias 0001, Gabriele Facciolo, Vicent Caselles, Guillermo Sapiro
Int. J. Comput. Vis.4
2011 A Continuum Mechanical Approach to Geodesics in Shape Space
Benedikt Wirth, Leah Bar, Martin Rumpf, Guillermo Sapiro
Int. J. Comput. Vis.4
2011 A Hough transform global probabilistic approach to multiple-subject diffusion MRI tractography
Iman Aganj, Christophe Lenglet, Neda Jahanshad, Essa Yacoub, Noam Harel, Paul M. Thompson, Guillermo Sapiro
Medical Image Anal.7
2011 Learning Discriminative Sparse Representations for Modeling, Source Separation, and Mapping of Hyperspectral Imagery
abstract
A method is presented for subpixel modeling, mapping, and classification in hyperspectral imagery using learned block-structured discriminative dictionaries, where each block is adapted and optimized to represent a material in a compact and sparse manner. The spectral pixels are modeled by linear combinations of subspaces defined by the learned dictionary atoms, allowing for linear mixture analysis. This model provides flexibility in source representation and selection, thus accounting for spectral variability, small-magnitude errors, and noise. A spatial-spectral coherence regularizer in the optimization allows pixel classification to be influenced by similar neighbors. We extend the proposed approach for cases for which there is no knowledge of the materials in the scene, unsupervised classification, and provide experiments and comparisons with simulated and real data. We also present results when the data have been significantly undersampled and then reconstructed, still retaining high-performance classification, showing the potential role of compressive sensing and sparse modeling techniques in efficient acquisition/transmission missions for hyperspectral imagery.
Alexey Castrodad, Zhengming Xing, John B. Greer, Edward Bosch, Lawrence Carin, Guillermo Sapiro
IEEE Trans. Geosci. Remote. Sens.6
2011 Esophagus Silhouette Extraction and Reconstruction From Fluoroscopic Views for Cardiac Ablation Procedure Guidance
abstract
Cardiac ablation involves the risk of serious complications when thermal injury to the esophagus occurs. This paper proposes to reduce the risk of such injuries by a proactive visualization technique, improving physician awareness of the esophagus location in the absence of or in addition to a reactive monitoring device such as a thermal probe. This is achieved by combining a graphical representation of the esophagus with live fluoroscopy. Toward this goal, we present an automated method to reconstruct and visualize a 3-D esophagus model from fluoroscopy image sequences acquired using different C-arm viewing directions. In order to visualize the esophagus under fluoroscopy, it is first biomarked by swallowing a contrast agent such as barium. Images obtained in this procedure are then used to automatically extract the 2-D esophagus silhouette and reconstruct a 3-D surface of the esophagus internal wall. Once the 3-D representation has been computed, it can be visualized using fluoroscopy overlay techniques. Compared to 3-D esophagus imaging using CT or C-arm CT, our proposed fluoroscopy method requires low radiation dose and enables a simpler workflow on geometry-calibrated standard C-arm systems.
Liron Yatziv, Julian Ibarz, Norbert Strobel, Saurabh Datta, Guillermo Sapiro
IEEE Trans. Inf. Technol. Biomed.5
2010 Classification and clustering via dictionary learning with structured incoherence and shared features
abstract
A clustering framework within the sparse modeling and dictionary learning setting is introduced in this work. Instead of searching for the set of centroid that best fit the data, as in k-means type of approaches that model the data as distributions around discrete points, we optimize for a set of dictionaries, one for each cluster, for which the signals are best reconstructed in a sparse coding manner. Thereby, we are modeling the data as a union of learned low dimensional subspaces, and data points associated to subspaces spanned by just a few atoms of the same learned dictionary are clustered together. An incoherence promoting term encourages dictionaries associated to different classes to be as independent as possible, while still allowing for different classes to share features. This term directly acts on the dictionaries, thereby being applicable both in the supervised and unsupervised settings. Using learned dictionaries for classification and clustering makes this method robust and well suited to handle large datasets. The proposed framework uses a novel measurement for the quality of the sparse representation, inspired by the robustness of the ℓ1regularization term in sparse coding. In the case of unsupervised classification and/or clustering, a new initialization based on combining sparse coding with spectral clustering is proposed. This initialization clusters the dictionary atoms, and therefore is based on solving a low dimensional eigen-decomposition problem, being applicable to large datasets. We first illustrate the proposed framework with examples on standard image and speech datasets in the supervised classification setting, obtaining results comparable to the state-of-the-art with this simple approach. We then present experiments for fully unsupervised clustering on extended standard datasets and texture images, obtaining excellent performance.
Ignacio Ramírez, Pablo Sprechmann, Guillermo Sapiro
CVPR3
2010 Dynamic Color Flow: A Motion-Adaptive Color Model for Object Segmentation in Video
Jue Wang 0001, Guillermo Sapiro
ECCV (5)3
2010 Hierarchical dictionary learning for invariant classification
abstract
Sparse representation theory has been increasingly used in the fields of signal processing and machine learning. The standard sparse models are not invariant to spatial transformations such as image rotations, and the representation is very sensitive even under small such distortions. Most studies addressing this problem proposed algorithms which either use transformed data as part of the training set, or are invariant or robust only under minor transformations. In this paper we suggest a framework which extracts sparse features invariant under significant rotations and scalings. The algorithm is based on a hierarchical architecture of dictionary learning for sparse coding in a cortical (log-polar) space. The proposed model is tested in supervised classification applications and proved to be robust under transformed data.
Leah Bar, Guillermo Sapiro
ICASSP2
2010 Dictionary learning and sparse coding for unsupervised clustering
abstract
A clustering framework within the sparse modeling and dictionary learning setting is introduced in this work. Instead of searching for the set of centroid that best fit the data, as in k-means type of approaches that model the data as distributions around discrete points, we optimize for a set of dictionaries, one for each cluster, for which the signals are best reconstructed in a sparse coding manner. Thereby, we are modeling the data as the of union of learned low dimensional subspaces, and data points associated to subspaces spanned by just a few atoms of the same learned dictionary are clustered together. Using learned dictionaries makes this method robust and well suited to handle large datasets. The proposed clustering algorithm uses a novel measurement for the quality of the sparse representation, inspired by the robustness of the ℓ1regularization term in sparse coding. We first illustrate this measurement with examples on standard image and speech datasets in the supervised classification setting, showing with a simple approach its discriminative power and obtaining results comparable to the state-of-the-art. We then conclude with experiments for fully unsupervised clustering on extended standard datasets and texture images, obtaining excellent performance.
Pablo Sprechmann, Guillermo Sapiro
ICASSP2
2010 Discriminative sparse representations in hyperspectral imagery
abstract
Recent advances in sparse modeling and dictionary learning for discriminative applications show high potential for numerous classification tasks. In this paper, we show that highly accurate material classification from hyperspectral imagery (HSI) can be obtained with these models, even when the data is reconstructed from a very small percentage of the original image samples. The proposed supervised HSI classification is performed using a measure that accounts for both reconstruction errors and sparsity levels for sparse representations based on class-dependent learned dictionaries. Combining the dictionaries learned for the different materials, a linear mixing model is derived for sub-pixel classification. Results with real hyperspectral data cubes are shown both for urban and non-urban terrain.
Alexey Castrodad, Zhengming Xing, John B. Greer, Edward Bosch, Lawrence Carin, Guillermo Sapiro
ICIP6
2010 Nonparametric image interpolation and dictionary learning using spatially-dependent Dirichlet and beta process priors
abstract
We present a Bayesian model for image interpolation and dictionary learning that uses two nonparametric priors for sparse signal representations: the beta process and the Dirichlet process. Additionally, the model uses spatial information within the image to encourage sharing of information within image subregions. We derive a hybrid MAP/Gibbs sampler, which performs Gibbs sampling for the latent indicator variables and MAP estimation for all other parameters. We present experimental results, where we show an improvement over other state-of-the-art algorithms in the low-measurement regime.
John W. Paisley, Mingyuan Zhou, Guillermo Sapiro, Lawrence Carin
ICIP3
2010 Image modeling and enhancement via structured sparse model selection
abstract
An image representation framework based on structured sparse model selection is introduced in this work. The corresponding modeling dictionary is comprised of a family of learned orthogonal bases. For an image patch, a model is first selected from this dictionary through linear approximation in a best basis, and the signal estimation is then calculated with the selected model. The model selection leads to a guaranteed near optimal denoising estimator. The degree of freedom in the model selection is equal to the number of the bases, typically about 10 for natural images, and is significantly lower than with traditional overcomplete dictionary approaches, stabilizing the representation. For an image patch of size √N × √N, the computational complexity of the proposed framework is O (N2), typically 2 to 3 orders of magnitude faster than estimation in an overcomplete dictionary. The orthogonal bases are adapted to the image of interest and are computed with a simple and fast procedure. State-of-the-art results are shown in image denoising, deblurring, and inpainting.
Guoshen Yu, Guillermo Sapiro, Stéphane Mallat
ICIP2
2010 ODF Maxima Extraction in Spherical Harmonic Representation via Analytical Search Space Reduction
Iman Aganj, Christophe Lenglet, Guillermo Sapiro
MICCAI (2)3
2010 A Gromov-Hausdorff Framework with Diffusion Geometry for Topologically-Robust Non-rigid Shape Matching
Alexander M. Bronstein, Michael M. Bronstein, Ron Kimmel, Mona Mahmoudi, Guillermo Sapiro
Int. J. Comput. Vis.5
2010 Online Learning for Matrix Factorization and Sparse Coding
Julien Mairal, Francis R. Bach, Jean Ponce, Guillermo Sapiro
J. Mach. Learn. Res.4
2010 Sparse Representation for Computer Vision and Pattern Recognition
abstract
Techniques from sparse signal representation are beginning to see significant impact in computer vision, often on nontraditional applications where the goal is not just to obtain a compact high-fidelity representation of the observed signal, but also to extract semantic information. The choice of dictionary plays a key role in bridging this gap: unconventional dictionaries consisting of, or learned from, the training samples themselves provide the key to obtaining state-of-the-art results and to attaching semantic meaning to sparse signal representations. Understanding the good performance of such unconventional dictionaries in turn demands new algorithmic and analytical techniques. This review paper highlights a few representative examples of how the interaction between sparse signal representation and computer vision can enrich both fields, and raises a number of open questions for further study.
John Wright 0001, Yi Ma 0001, Julien Mairal, Guillermo Sapiro, Thomas S. Huang, Shuicheng Yan
Proc. IEEE4
2010 Introduction to the Special Section on Optimization in Imaging Sciences
abstract
Many image analysis problems can be formulated as estimation problems that require minimization of a discrete or discretized energy function. Commonly used energy functions are based on models in physics, geometry, or statistics. With some exceptions, however, many formulations correspond to NP-hard optimization problems. Recently there has been a surge in research in image sciences where either known or newly developed combinatorial optimization algorithms were successfully applied to a wide spectrum of problems in imaging. There are several aspects of this development. Some of the new combinatorial algorithms could guarantee globally optimal solutions for certain special cases, and several such special classes were identified in image. Furthermore, new optimization methods with proven performance guarantees have been developed and applied to a much wider spectrum of more difficult imaging problems. Imaging science has been enjoying both from the use of advanced optimization techniques and the development of new techniques with contributions beyond imaging problems. This special section includes a number of representative papers in this topic. The diversity of these papers showcases the large spectrum of optimization topics and challenges currently addressed by the imaging community: These papers also show a unique characteristic of this area in imaging sciences: the combination of deep theoretical contributions with very relevant practical applications. The tools being used for proving the important theoretical results and the actual applications are very broad and diverse as well. I would like to thank the guest editors—Professor Endre Boros, Professor Yuri Boykov, Professor Jerome Darbon, and Professor Philip Torr—as well as the special section advisor, Professor Andrew Blake, for their work with this section. Special thanks go also to the SIAM staff for the support for this section, as well as to the SIIMS editorial board for guidance.
Guillermo Sapiro
SIAM J. Imaging Sci.1
2010 A Comprehensive Framework for Image Inpainting
abstract
Inpainting is the art of modifying an image in a form that is not detectable by an ordinary observer. There are numerous and very different approaches to tackle the inpainting problem, though as explained in this paper, the most successful algorithms are based upon one or two of the following three basic techniques: copy-and-paste texture synthesis, geometric partial differential equations (PDEs), and coherence among neighboring pixels. We combine these three building blocks in a variational model, and provide a working algorithm for image inpainting trying to approximate the minimum of the proposed energy functional. Our experiments show that the combination of all three terms of the proposed energy works better than taking each term separately, and the results obtained are within the state-of-the-art.
Aurélie Bugeau, Marcelo Bertalmío, Vicent Caselles, Guillermo Sapiro
IEEE Trans. Image Process.4
2010 Simultaneous Object Classification and Segmentation With High-Order Multiple Shape Models
abstract
Shape models (SMs), capturing the common features of a set of training shapes, represent a new incoming object based on its projection onto the corresponding model. Given a set of learned SMs representing different objects classes, and an image with a new shape, this work introduces a joint classification-segmentation framework with a twofold goal. First, to automatically select the SM that best represents the object, and second, to accurately segment the image taking into account both the image information and the features and variations learned from the online selected model. A new energy functional is introduced that simultaneously accomplishes both goals. Model selection is performed based on a shape similarity measure, online determining which model to use at each iteration of the steepest descent minimization, allowing for model switching and adaptation to the data. High-order SMs are used in order to deal with very similar object classes and natural variability within them. Position and transformation invariance is included as part of the modeling as well. The presentation of the framework is complemented with examples for the difficult task of simultaneously classifying and segmenting closely related shapes, such as stages of human activities, in images with severe occlusions.
Federico Lecumberry, Alvaro Pardo, Guillermo Sapiro
IEEE Trans. Image Process.3
2009 Non-local sparse models for image restoration
abstract
We propose in this paper to unify two different approaches to image restoration: On the one hand, learning a basis set (dictionary) adapted to sparse signal descriptions has proven to be very effective in image reconstruction and classification tasks. On the other hand, explicitly exploiting the self-similarities of natural images has led to the successful non-local means approach to image restoration. We propose simultaneous sparse coding as a framework for combining these two approaches in a natural manner. This is achieved by jointly decomposing groups of similar signals on subsets of the learned dictionary. Experimental results in image denoising and demosaicking tasks with synthetic and real noise show that the proposed method outperforms the state of the art, making it possible to effectively restore raw images from digital cameras at a reasonable speed and memory cost.
Julien Mairal, Francis R. Bach, Jean Ponce, Guillermo Sapiro, Andrew Zisserman
ICCV4
2009 Seeing 3D objects in a single 2D image
abstract
A general framework simultaneously addressing pose estimation, 2D segmentation, object recognition, and 3D reconstruction from a single image is introduced in this paper. The proposed approach partitions 3D space into voxels and estimates the voxel states that maximize a likelihood integrating two components: the object fidelity, that is, the probability that an object occupies the given voxels, here encoded as a 3D shape prior learned from 3D samples of objects in a class; and the image fidelity, meaning the probability that the given voxels would produce the input image when properly projected to the image plane. We derive a loop-less graphical model for this likelihood and propose a computationally efficient optimization algorithm that is guaranteed to produce the global likelihood maximum. Furthermore, we derive a multi-resolution implementation of this algorithm that permits to trade reconstruction and estimation accuracy for computation. The presentation of the proposed framework is complemented with experiments on real data demonstrating the accuracy of the proposed approach.
Diego Rother, Guillermo Sapiro
ICCV2
2009 Multiple shape models for simultaneous object classification and segmentation
abstract
Shape models (SMs), capturing the common features of a set of training shapes, represent a new incoming object based on its projection onto the corresponding model. Given a set of learned SMs representing different objects, and an image with a new shape, this work introduces a joint classification-segmentation framework with a twofold goal. First, to automatically select the SM that best represents the object, and second, to accurately segment the image taking into account both the image information and the features and variations learned from the on-line selected model. A new energy functional is introduced that simultaneously accomplishes both goals. Model selection is performed based on a shape similarity measure, determining which model to use at each iteration of the steepest descent minimization, allowing for model switching and adaptation to the data. High-order SMs are used in order to deal with very similar object classes and natural variability within them. The presentation of the framework is complemented with examples for the difficult task of simultaneously classifying and segmenting closely related shapes, stages of human activities, in images with severe occlusions.
Federico Lecumberry, Alvaro Pardo, Guillermo Sapiro
ICIP3
2009 Online dictionary learning for sparse coding
abstract
Sparse coding---that is, modelling data vectors as sparse linear combinations of basis elements---is widely used in machine learning, neuroscience, signal processing, and statistics. This paper focuses on learning the basis set, also called dictionary, to adapt it to specific data, an approach that has recently proven to be very effective for signal reconstruction and classification in the audio and image processing domains. This paper proposes a new online optimization algorithm for dictionary learning, based on stochastic approximations, which scales up gracefully to large datasets with millions of training samples. A proof of convergence is presented, along with experiments with natural images demonstrating that it leads to faster performance and better dictionaries than classical batch algorithms for both small and large datasets.
Julien Mairal, Francis R. Bach, Jean Ponce, Guillermo Sapiro
ICML4
2009 Discriminative k-metrics
abstract
The k q-flats algorithm is a generalization of the popular k-means algorithm where q dimensional best fit affine sets replace centroids as the cluster prototypes. In this work, a modification of the k q-flats framework for pattern classification is introduced. The basic idea is to replace the original reconstruction only energy, which is optimized to obtain the k affine spaces, by a new energy that incorporates discriminative terms. This way, the actual classification task is introduced as part of the design and optimization. The presentation of the proposed framework is complemented with experimental results, showing that the method is computationally very efficient and gives excellent results on standard supervised learning benchmarks.
Arthur Szlam, Guillermo Sapiro
ICML2
2009 Multiple Q-Shell ODF Reconstruction in Q-Ball Imaging
Iman Aganj, Christophe Lenglet, Guillermo Sapiro, Essa Yacoub, Kâmil Ugurbil, Noam Harel
MICCAI (1)3
2009 Non-Parametric Bayesian Dictionary Learning for Sparse Image Representations
abstract
Non-parametric Bayesian techniques are considered for learning dictionaries for sparse image representations, with applications in denoising, inpainting and compressive sensing (CS). The beta process is employed as a prior for learning the dictionary, and this non-parametric method naturally infers an appropriate dictionary size. The Dirichlet process and a probit stick-breaking process are also considered to exploit structure within an image. The proposed method can learn a sparse dictionary in situ; training images may be exploited if available, but they are not required. Further, the noise variance need not be known, and can be non-stationary. Another virtue of the proposed method is that sequential inference can be readily employed, thereby allowing scaling to large images. Several example results are presented, using both Gibbs and variational Bayesian inference, with comparisons to other state-of-the-art approaches.
Mingyuan Zhou, Haojun Chen, John W. Paisley, Guillermo Sapiro, Lawrence Carin
NIPS5
2009 Three-dimensional point cloud recognition via distributions of geometric distances
Mona Mahmoudi, Guillermo Sapiro
Graph. Model.2
2009 Geodesic Matting: A Framework for Fast Interactive Image and Video Segmentation and Matting
Guillermo Sapiro
Int. J. Comput. Vis.2
2009 New Possibilities with Sobolev Active Contours
Ganesh Sundaramoorthi, Anthony J. Yezzi, Andrea Mennucci, Guillermo Sapiro
Int. J. Comput. Vis.4
2009 Generalized Newton-Type Methods for Energy Formulations in Image Processing
abstract
Many problems in image processing are addressed via the minimization of a cost functional. The most prominently used optimization technique is gradient-descent, often used due to its simplicity and applicability where other techniques, e.g., those coming from discrete optimization, cannot be applied. Yet, gradient-descent suffers from slow convergence, and often to just local minima which highly depend on the initialization and the condition number of the functional Hessian. Newton-type methods, on the other hand, are known to have a faster, quadratic convergence. In its classical form, the Newton method relies on the $L^2$-type norm to define the descent direction. In this paper, we generalize and reformulate this very important optimization method by introducing Newton-type methods based on more general norms. Such norms are introduced both in the descent computation (Newton step) and in the corresponding stabilizing trust-region. This generalization opens up new possibilities in the extraction of the Newton step, including benefits such as mathematical stability and the incorporation of smoothness constraints. We first present the derivation of the modified Newton step in the calculus of variation framework needed for image processing. Then, we demonstrate the method with two common objective functionals: variational image deblurring and geometric active contours for image segmentation. We show that in addition to the fast convergence, norms adapted to the problem at hand yield different and superior results.
Leah Bar, Guillermo Sapiro
SIAM J. Imaging Sci.2
2009 Learning to Sense Sparse Signals: Simultaneous Sensing Matrix and Sparsifying Dictionary Optimization
abstract
Sparse signal representation, analysis, and sensing have received a lot of attention in recent years from the signal processing, optimization, and learning communities. On one hand, learning overcomplete dictionaries that facilitate a sparse representation of the data as a liner combination of a few atoms from such dictionary leads to state-of-the-art results in image and video restoration and classification. On the other hand, the framework of compressed sensing (CS) has shown that sparse signals can be recovered from far less samples than those required by the classical Shannon-Nyquist Theorem. The samples used in CS correspond to linear projections obtained by a sensing projection matrix. It has been shown that, for example, a nonadaptive random sampling matrix satisfies the fundamental theoretical requirements of CS, enjoying the additional benefit of universality. On the other hand, a projection sensing matrix that is optimally designed for a certain class of signals can further improve the reconstruction accuracy or further reduce the necessary number of samples. In this paper, we introduce a framework for the joint design and optimization, from a set of training images, of the nonparametric dictionary and the sensing matrix. We show that this joint optimization outperforms both the use of random sensing matrices and those matrices that are optimized independently of the learning of the dictionary. Particular cases of the proposed framework include the optimization of the sensing matrix for a given dictionary as well as the optimization of the dictionary for a predefined sensing environment. The presentation of the framework and its efficient numerical optimization is complemented with numerous examples on classical image datasets.
Julio Martin Duarte-Carvajalino, Guillermo Sapiro
IEEE Trans. Image Process.2
2009 Video SnapCut: robust video object cutout using localized classifiers
abstract
Although tremendous success has been achieved for interactive object cutout in still images, accurately extracting dynamic objects in video remains a very challenging problem. Previous video cutout systems present two major limitations: (1) reliance on global statistics, thus lacking the ability to deal with complex and diverse scenes; and (2) treating segmentation as a global optimization, thus lacking a practical workflow that can guarantee the convergence of the systems to the desired results. We present Video SnapCut , a robust video object cutout system that significantly advances the state-of-the-art. In our system segmentation is achieved by the collaboration of a set of local classifiers, each adaptively integrating multiple local image features. We show how this segmentation paradigm naturally supports local user editing and propagates them across time. The object cutout system is completed with a novel coherent video matting technique. A comprehensive evaluation and comparison is presented, demonstrating the effectiveness of the proposed system at achieving high quality results, as well as the robustness of the system against various types of inputs.
Jue Wang 0001, David Simons, Guillermo Sapiro
ACM Trans. Graph.4
2008 Discriminative learned dictionaries for local image analysis
abstract
Sparse signal models have been the focus of much recent research, leading to (or improving upon) state-of-the-art results in signal, image, and video restoration. This article extends this line of research into a novel framework for local image discrimination tasks, proposing an energy formulation with both sparse reconstruction and class discrimination components, jointly optimized during dictionary learning. This approach improves over the state of the art in texture segmentation experiments using the Brodatz database, and it paves the way for a novel scene analysis and recognition framework based on simultaneously learning discriminative and reconstructive dictionaries. Preliminary results in this direction using examples from the Pascal VOC06 and Graz02 datasets are presented as well.
Julien Mairal, Francis R. Bach, Jean Ponce, Guillermo Sapiro, Andrew Zisserman
CVPR4
2008 Generalized Newton methods for energy formulations in image procesing
abstract
Many problems in image processing are solved via the minimization of a cost functional. The most widely used optimization technique is the gradient descent, often used due to its simplicity and applicability where other optimization techniques, e.g., those coming from discrete optimization, can not be used. Yet, gradient descent suffers from a slow convergence, and often to just local minima which highly depends on the condition number of the functional Hessian. Newton- type methods, on the other hand, are known to have a rapid (quadratic) convergence. In its classical form, the Newton method relies on the L2-type norm to define the descent direction. In this paper, we generalize and reformulate this very important optimization method by introducing a novel Newton method based on general norms. This generalization opens up new possibilities in the extraction of the Newton step, including benefits such as mathematical stability and smoothness constraints. We first present the derivation of the modified Newton step in the calculus of variation framework. Then we demonstrate the method with two common objective functionals: variational image deblurring and geodesic active contours. We show that in addition to the fast convergence, different selections norm yield different and superior results.
Leah Bar, Guillermo Sapiro
ICIP2
2008 Supervised Dictionary Learning
abstract
It is now well established that sparse signal models are well suited to restoration tasks and can effectively be learned from audio, image, and video data. Recent research has been aimed at learning discriminative sparse models instead of purely reconstructive ones. This paper proposes a new step in that direction with a novel sparse representation for signals belonging to different classes in terms of a shared dictionary and multiple decision functions. It is shown that the linear variant of the model admits a simple probabilistic interpretation, and that its most general variant also admits a simple interpretation in terms of kernels. An optimization framework for learning all the components of the proposed model is presented, along with experiments on standard handwritten digit and texture classification tasks.
Julien Mairal, Francis R. Bach, Jean Ponce, Guillermo Sapiro, Andrew Zisserman
NIPS4
2008 On geometric variational models for inpainting surface holes
Vicent Caselles, Gloria Haro, Guillermo Sapiro, Joan Verdera
Comput. Vis. Image Underst.3
2008 Translated Poisson Mixture Model for Stratification Learning
Gloria Haro, Gregory Randall, Guillermo Sapiro
Int. J. Comput. Vis.3
2008 Robust Foreground Detection In Video Using Pixel Layers
abstract
A framework for robust foreground detection that works under difficult conditions such as dynamic background and moderately moving camera is presented in this paper. The proposed method includes two main components: coarse scene representation as the union of pixel layers, and foreground detection in video by propagating these layers using a maximum-likelihood assignment. We first cluster into "layers" those pixels that share similar statistics. The entire scene is then modeled as the union of such non-parametric layer-models. An in-coming pixel is detected as foreground if it does not adhere to these adaptive models of the background. A principled way of computing thresholds is used to achieve robust detection performance with a pre-specified number of false alarms. Correlation between pixels in the spatial vicinity is exploited to deal with camera motion without precise registration or optical flow. The proposed technique adapts to changes in the scene, and allows to automatically convert persistent foreground objects to background and re-convert them to foreground when they become interesting. This simple framework addresses the important problem of robust foreground and unusual region detection, at about 10 frames per second on a standard laptop computer. The presentation of the proposed approach is complemented by results on challenging real data and comparisons with other standard techniques.
Kedar A. Patwardhan, Guillermo Sapiro, Vassilios Morellas
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 Message from the Editor-in-Chief
abstract
Dear SIIMS Reader, It is a great pleasure and privilege to welcome you to the first issue of the SIAM Journal on Imaging Sciences. This journal, with the help of its editorial board, contributors, readers, and SIAM staff, follows the outstanding tradition of high-quality SIAM journals and will become a leading journal in the broad aspects of fundamental imaging sciences. The journal began accepting submissions in April of 2007. By the end of February 2008, SIIMS had received 88 manuscripts submitted for publication. Of these, you will find the first six accepted papers in this first all-electronic issue. Additional papers have been accepted recently and will appear in subsequent issues. The papers in this issue provide a representative glance of the diversity of topics covered by the journal, something clearly observable from the outstanding editorial board itself. A number of papers currently in the pipeline and to be published in future issues further stress the broad vision of the journal. SIIMS came to light thanks to the tremendous effort of numerous people, and I want to publicly thank them. Prof. Tony Chan worked very hard with SIAM to define this new journal. The whole SIAM staff is superb, and without them my life as editor-in-chief would be impossible. In particular, Mitch Chernoff and Cherie Trebisky have been a critical source of support. The current editorial board has been very supportive and worked very hard to guarantee timely and high-quality reviews. The whole cycle of reviewing and production has been extraordinarily fast on average; some of the papers in this issue were submitted just a few short months ago, and some in this issue were accepted less than two months ago. I appreciate very much the support of all the people involved with this journal and their hard work in making this possible. Please enjoy your reading, and we look forward to your high-quality submissions in the future.
Guillermo Sapiro
SIAM J. Imaging Sci.1
2008 Statistical Characterization of Protein Ensembles
abstract
When accounting for structural fluctuations or measurement errors, a single rigid structure may not be sufficient to represent a protein. One approach to solve this problem is to represent the possible conformations as a discrete set of observed conformations, an ensemble. In this work, we follow a different richer approach, and introduce a framework for estimating probability density functions in very high dimensions, and then apply it to represent ensembles of folded proteins. This proposed approach combines techniques such as kernel density estimation, maximum likelihood, cross-validation, and bootstrapping. We present the underlying theoretical and computational framework and apply it to artificial data and protein ensembles obtained from molecular dynamics simulations. We compare the results with those obtained experimentally, illustrating the potential and advantages of this representation.
Diego Rother, Guillermo Sapiro, Vijay Pande
IEEE ACM Trans. Comput. Biol. Bioinform.2
2008 Multiscale Representation and Segmentation of Hyperspectral Imagery Using Geometric Partial Differential Equations and Algebraic Multigrid Methods
abstract
A fast algorithm for multiscale representation and segmentation of hyperspectral imagery is introduced in this paper. The multiscale/scale-space representation is obtained by solving a nonlinear diffusion partial differential equation (PDE) for vector-valued images. We use algebraic multigrid techniques to obtain a fast and scalable solution of the PDE and to segment the hyperspectral image following the intrinsic multigrid structure. We test our algorithm on four standard hyperspectral images that represent different environments commonly found in remote sensing applications: agricultural, urban, mining, and marine. The experimental results show that the segmented images lead to better classification than using the original data directly, in spite of the use of simple similarity metrics and piecewise constant approximations obtained from the segmentation maps.
Julio Martin Duarte-Carvajalino, Guillermo Sapiro, Miguel Velez-Reyes, Paul Castillo 0001
IEEE Trans. Geosci. Remote. Sens.2
2008 Sparse Representation for Color Image Restoration
abstract
Sparse representations of signals have drawn considerable interest in recent years. The assumption that natural signals, such as images, admit a sparse decomposition over a redundant dictionary leads to efficient algorithms for handling such sources of data. In particular, the design of well adapted dictionaries for images has been a major challenge. The K-SVD has been recently proposed for this task and shown to perform very well for various grayscale image processing tasks. In this paper, we address the problem of learning dictionaries for color images and extend the K-SVD-based grayscale image denoising algorithm that appears in. This work puts forward ways for handling nonhomogeneous noise and missing information, paving the way to state-of-the-art results in applications such as color image denoising, demosaicing, and inpainting, as demonstrated in this paper.
Julien Mairal, Michael Elad, Guillermo Sapiro
IEEE Trans. Image Process.3
2007 Connecting the Out-of-Sample and Pre-Image Problems in Kernel Methods
abstract
Kernel methods have been widely studied in the field of pattern recognition. These methods implicitly map, "the kernel trick," the data into a space which is more appropriate for analysis. Many manifold learning and dimensionality reduction techniques are simply kernel methods for which the mapping is explicitly computed. In such cases, two problems related with the mapping arise: The out-of-sample extension and the pre-image computation. In this paper we propose a new pre-image method based on the Nystrom formulation for the out-of-sample extension, showing the connections between both problems. We also address the importance of normalization in the feature space, which has been ignored by standard pre-image algorithms. As an example, we apply these ideas to the Gaussian kernel, and relate our approach to other popular pre-image methods. Finally, we show the application of these techniques in the study of dynamic shapes.
Pablo Arias 0001, Gregory Randall, Guillermo Sapiro
CVPR3
2007 Regularized Mixed Dimensionality and Density Learning in Computer Vision
abstract
A framework for the regularized estimation of nonuniform dimensionality and density in high dimensional data is introduced in this work. This leads to learning stratifications, that is, mixture of manifolds representing different characteristics and complexities in the data set. The basic idea relies on modeling the high dimensional sample points as a process of Poisson mixtures, with regularizing restrictions and spatial continuity constraints. Theoretical asymptotic results for the model are presented as well. The presentation of the framework is complemented with artificial and real examples showing the importance of regularized stratification learning in computer vision applications.
Gloria Haro, Gregory Randall, Guillermo Sapiro
CVPR3
2007 A Geodesic Framework for Fast Interactive Image and Video Segmentation and Matting
abstract
An interactive framework for soft segmentation and matting of natural images and videos is presented in this paper. The proposed technique is based on the optimal, linear time, computation of weighted geodesic distances to the user-provided scribbles, from which the whole data is automatically segmented. The weights are based on spatial and/or temporal gradients, without explicit optical flow or any advanced and often computationally expensive feature detectors. These could be naturally added to the proposed framework as well if desired, in the form of weights in the geodesic distances. A localized refinement step follows this fast segmentation in order to accurately compute the corresponding matte function. Additional constraints into the distance definition permit to efficiently handle occlusions such as people or objects crossing each other in a video sequence. The presentation of the framework is complemented with numerous and diverse examples, including extraction of moving foreground from dynamic background, and comparisons with the recent literature.
Guillermo Sapiro
ICCV2
2007 A Variational Framework for Simultaneous Motion Estimation and Restoration of Motion-Blurred Video
abstract
The problem of motion estimation and restoration of objects in a blurred video sequence is addressed in this paper. Fast movement of the objects, together with the aperture time of the camera, result in a motion-blurred image. The direct velocity estimation from this blurred video is inaccurate. On the other hand, an accurate estimation of the velocity of the moving objects is critical for restoration of motion-blurred video. Therefore, restoration needs accurate motion estimation and vice versa, and a joint process is called for. To address this problem we derive a novel model of the blurring process and propose a Mumford-Shah type of variational framework, acting on consecutive frames, for joint object deblurring and velocity estimation. The proposed procedure distinguishes between the moving object and the background and is accurate also close to the boundary of the moving object. Experimental results both on simulated and real data show the importance of this joint estimation and its superior performance when compared to the independent estimation of motion and restoration.
Leah Bar, Benjamin Berkels, Martin Rumpf, Guillermo Sapiro
ICCV4
2007 What Can Casual Walkers Tell Us About A 3D Scene?
abstract
An approach for incremental learning of a 3D scene from a single static video camera is presented in this paper. In particular, we exploit the presence of casual people walking in the scene to infer relative depth, learn shadows, and segment the critical ground structure. Considering that this type of video data is so ubiquitous, this work provides an important step towards 3D scene analysis from single cameras in readily available ordinary videos and movies. On-line 3D scene learning, as presented here, is very important for applications such as scene analysis, foreground refinement, tracking, biometrics, automated camera collaboration, activity analysis, identification, and real-time computer-graphics applications. The main contributions of this work are then two-fold. First, we use the people in the scene to continuously learn and update the 3D scene parameters using an incremental robust (L1) error minimization. Secondly, models of shadows in the scene are learned using a statistical framework. A symbiotic relationship between the shadow model and the estimated scene geometry is exploited towards incremental mutual improvement. We illustrate the effectiveness of the proposed framework with applications in foreground refinement, automatic segmentation as well as relative depth mapping of the floor/ground, and estimation of 3D trajectories of people in the scene.
Diego Rother, Kedar A. Patwardhan, Guillermo Sapiro
ICCV3
2007 Distancecut: Interactive Segmentation and Matting of Images and Videos
abstract
An interactive algorithm for soft segmentation and matting of natural images and videos is presented in this paper. The technique follows and extends Protiere et al. (2007), where the user first roughly scribbles/labels different regions of interest, and from them the whole data is automatically segmented. The segmentation and alpha matte are obtained from the fast, linear complexity, computation of weighted distances to the user-provided scribbles. These weighted distances assign probabilities to each labeled class for every pixel. The weights are derived from models of the image regions obtained from the user provided scribbles via kernel density estimation. The matting results follow from combining this density and the computed weighted distances. We present the underlying framework and examples showing the capability of the algorithm to segment and compute alpha mattes, in interactive real time, for difficult natural data.
Guillermo Sapiro
ICIP (2)2
2007 Multiscale Sparse Image Representationwith Learned Dictionaries
abstract
This paper introduces a new framework for learning multiscale sparse representations of natural images with overcomplete dictionaries. Our work extends the K-SVD algorithm [1], which learns sparse single-scale dictionaries for natural images. Recent work has shown that the K-SVD can lead to state-of-the-art image restoration results [2, 3]. We show that these are further improved with a multi-scale approach, based on a Quadtree decomposition. Our framework provides an alternative to multiscale pre-defined dictionaries such as wavelets, curvelets, and contourlets, with dictionaries optimized for the data and application instead of pre-modelled ones.
Julien Mairal, Guillermo Sapiro, Michael Elad
ICIP (3)2
2007 A Graph-Based Foreground Representation and its Application in Example Based People Matching in Video
abstract
In this work, we propose a framework for foreground representation in video and illustrate it with a multi-camera people matching application. We first decompose the video into foreground and background. A low-level coarse segmentation of the foreground is then used to generate a simple graph representation. A vertex in the graph represents the "appearance" of a corresponding segment in the foreground, while the relationship between two segments is encoded by an edge between the corresponding vertices. This provides a simple yet powerful and general representation of the foreground, which can be very useful in problems such as people detection and tracking. We illustrate the effectiveness of this model using an "example based query" type of application for people matching in videos. Matching results are provided in multiple-camera situations and also under occlusion.
Kedar A. Patwardhan, Guillermo Sapiro, Vassilios Morellas
ICIP (5)2
2007 Meshless geometric subdivision
Carsten Moenning, Facundo Mémoli, Guillermo Sapiro, Nira Dyn, Neil A. Dodgson
Graph. Model.3
2007 Spatially Coherent Nonlinear Dimensionality Reduction and Segmentation of Hyperspectral Images
abstract
The nonlinear dimensionality reduction and its effects on vector classification and segmentation of hyperspectral images are investigated in this letter. In particular, the way dimensionality reduction influences and helps classification and segmentation is studied. The proposed framework takes into account the nonlinear nature of high-dimensional hyperspectral images and projects onto a lower dimensional space via a novel spatially coherent locally linear embedding technique. The spatial coherence is introduced by comparing pixels based on their local surrounding structure in the image domain and not just on their individual values as classically done. This spatial coherence in the image domain across the multiple bands defines the high-dimensional local neighborhoods used for the dimensionality reduction. This spatial coherence concept is also extended to the segmentation and classification stages that follow the dimensionality reduction, introducing a modified vector angle distance. We present the underlying concepts of the proposed framework and experimental results showing the significant classification improvements
Anish Mohan, Guillermo Sapiro, Edward Bosch
IEEE Geosci. Remote. Sens. Lett.2
2007 Video Inpainting Under Constrained Camera Motion
abstract
A framework for inpainting missing parts of a video sequence recorded with a moving or stationary camera is presented in this work. The region to be inpainted is general: it may be still or moving, in the background or in the foreground, it may occlude one object and be occluded by some other object. The algorithm consists of a simple preprocessing stage and two steps of video inpainting. In the preprocessing stage, we roughly segment each frame into foreground and background. We use this segmentation to build three image mosaics that help to produce time consistent results and also improve the performance of the algorithm by reducing the search space. In the first video inpainting step, we reconstruct moving objects in the foreground that are "occluded" by the region to be inpainted. To this end, we fill the gap as much as possible by copying information from the moving foreground in other frames, using a priority-based scheme. In the second step, we inpaint the remaining hole with the background. To accomplish this, we first align the frames and directly copy when possible. The remaining pixels are filled in by extending spatial texture synthesis techniques to the spatiotemporal domain. The proposed framework has several advantages over state-of-the-art algorithms that deal with similar types of data and constraints. It permits some camera motion, is simple to implement, fast, does not require statistical models of background nor foreground, works well in the presence of rich and cluttered backgrounds, and the results show that there is no visible blurring or motion artifacts. A number of real examples taken with a consumer hand-held camera are shown supporting these findings.
Kedar A. Patwardhan, Guillermo Sapiro, Marcelo Bertalmío
IEEE Trans. Image Process.2
2007 Interactive Image Segmentation via Adaptive Weighted Distances
abstract
An interactive algorithm for soft segmentation of natural images is presented in this paper. The user first roughly scribbles different regions of interest, and from them, the whole image is automatically segmented. This soft segmentation is obtained via fast, linear complexity computation of weighted distances to the user-provided scribbles. The adaptive weights are obtained from a series of Gabor filters, and are automatically computed according to the ability of each single filter to discriminate between the selected regions of interest. We present the underlying framework and examples showing the capability of the algorithm to segment diverse images.
Alexis Protiere, Guillermo Sapiro
IEEE Trans. Image Process.2
2007 A Geometric Method for Automatic Extraction of Sulcal Fundi
abstract
Sulcal fundi are 3-D curves that lie in the depths of the cerebral cortex and, in addition to their intrinsic value in brain research, are often used as landmarks for downstream computations in brain imaging. In this paper, we present a geometric algorithm that automatically extracts the sulcal fundi from magnetic resonance images and represents them as spline curves lying on the extracted triangular mesh representing the cortical surface. The input to our algorithm is a triangular mesh representation of an extracted cortical surface as computed by one of several available software packages for performing automated and semi-automated cortical surface extraction. Given this input we first compute a geometric depth measure for each triangle on the cortical surface mesh, and based on this information we extract sulcal regions by checking for connected regions exceeding a depth threshold. We then identify endpoints of each region and delineate the fundus by thinning the connected region while keeping the endpoints fixed. The curves, thus, defined are regularized using weighted splines on the surface mesh to yield high-quality representations of the sulcal fundi. We present the geometric framework and validate it with real data from human brains. Comparisons with expert-labeled sulcal fundi are part of this validation process.
Chiu Yen Kao, Michael Hofer, Guillermo Sapiro, Josh Stern, Kelly Rehm, David A. Rottenberg
IEEE Trans. Medical Imaging3
2006 Stratification Learning: Detecting Mixed Density and Dimensionality in High Dimensional Point Clouds
abstract
The study of point cloud data sampled from a stratification, a collection of manifolds with possible different dimensions, is pursued in this paper. We present a technique for simultaneously soft clustering and estimating the mixed dimensionality and density of such structures. The framework is based on a maximum likelihood estimation of a Poisson mixture model. The presentation of the approach is completed with artificial and real examples demonstrating the importance of extending manifold learning to stratification learning.
Gloria Haro, Gregory Randall, Guillermo Sapiro
NIPS3
2006 Constrained regularization of digital terrain elevation data
abstract
A framework for geometric regularization of elevation maps is introduced in this letter. The framework takes into account errors in the data, which form part of standard elevation maps specifications, as well as possible additional user/application-dependent constraints. The algorithm is based on adapting the theory of geometric active surfaces to the problem of regularizing elevation maps. We present the underlying concepts and numerical experiments showing the effectiveness and potential of this theory.
Anish Mohan, Alberto Bartesaghi, Guillermo Sapiro
IEEE Geosci. Remote. Sens. Lett.3
2006 Statistical Analysis of RNA Backbone
abstract
Local conformation is an important determinant of RNA catalysis and binding. The analysis of RNA conformation is particularly difficult due to the large number of degrees of freedom (torsion angles) per residue. Proteins, by comparison, have many fewer degrees of freedom per residue. In this work, we use and extend classical tools from statistics and signal processing to search for clusters in RNA conformational space. Results are reported both for scalar analysis, where each torsion angle is separately studied, and for vectorial analysis, where several angles are simultaneously clustered. Adapting techniques from vector quantization and clustering to the RNA structure, we find torsion angle clusters and RNA conformational motifs. We validate the technique using well-known conformational motifs, showing that the simultaneous study of the total torsion angle space leads to results consistent with known motifs reported in the literature and also to the finding of new ones.
Eli Hershkovitz, Guillermo Sapiro, Allen R. Tannenbaum, Loren Dean Williams
IEEE ACM Trans. Comput. Biol. Bioinform.2
2006 Fair polyline networks for constrained smoothing of digital terrain elevation data
abstract
In this paper, a framework for smoothing gridlike digital terrain elevation data, which achieves a fair shape by means of minimizing an energy functional, is presented. The minimization is performed under the side condition of hard constraints, which comes from available horizontal and vertical accuracy bounds in the standard elevation specification. In this paper, the framework is introduced, and the suitability of this method for the tasks of accuracy-constrained smoothing, feature-preserving smoothing, and filling of data voids is demonstrated
Michael Hofer, Guillermo Sapiro, Johannes Wallner 0001
IEEE Trans. Geosci. Remote. Sens.2
2006 Fast image and video colorization using chrominance blending
abstract
Colorization, the task of coloring a grayscale image or video, involves assigning from the single dimension of intensity or luminance a quantity that varies in three dimensions, such as red, green, and blue channels. Mapping between intensity and color is, therefore, not unique, and colorization is ambiguous in nature and requires some amount of human interaction or external information. A computationally simple, yet effective, approach of colorization is presented in this paper. The method is fast and it can be conveniently used "on the fly," permitting the user to interactively get the desired results promptly after providing a reduced set of chrominance scribbles. Based on the concepts of luminance-weighted chrominance blending and fast intrinsic distance computations, high-quality colorization results for still images and video are obtained at a fraction of the complexity and computational cost of previously reported techniques. Possible extensions of the algorithm introduced here included the capability of changing the colors of an existing color image or video, as well as changing the underlying luminance, and many other special effects demonstrated here.
Liron Yatziv, Guillermo Sapiro
IEEE Trans. Image Process.2
2005 Tracking of moving objects under severe and total occlusions
abstract
We present an algorithm for tracking moving objects using intrinsic minimal surfaces which handles particularly well the presence of severe and total occlusions even in the presence of weak object boundaries. We adopt an edge based approach and find the segmentation as a minimal surface in 3D space-time, the metric being dictated by the image gradient. Object boundaries are represented implicitly as the level set of a higher dimensional function, and no particular object model is assumed. We also avoid explicit estimation of a dynamic model since the problem is regarded as one of static energy minimization. A set of interior points provided by the user is used to constrain the optimization, which basically corresponds to selecting the object of interest within the video sequence. The constraints are such that they restrict the resulting surface to be star-shaped in the 3D spatio-temporal space. We present some challenging examples that show the robustness of the technique.
Alberto Bartesaghi, Guillermo Sapiro
ICIP (1)2
2005 Video inpainting of occluding and occluded objects
abstract
We present a basic technique to fill-in missing parts of a video sequence taken from a static camera. Two important cases are considered. The first case is concerned with the removal of non-stationary objects that occlude stationary background. We use a priority based spatio-temporal synthesis scheme for inpainting the stationary background. The second and more difficult case involves filling-in moving objects when they are partially occluded. For this, we propose a priority scheme to first inpaint the occluded moving objects and then fill-in the remaining area with stationary background using the method proposed for the first case. We use as input an optical-flow based mask, which tells if an undamaged pixel is moving or is stationary. The moving object is inpainted by copying patches from undamaged frames, and this copying is independent of the background of the moving object in either frame. This work has applications in a variety of different areas, including video special effects and restoration and enhancement of damaged videos. The examples shown in the paper illustrate these ideas.
Kedar A. Patwardhan, Guillermo Sapiro, Marcelo Bertalmío
ICIP (2)2
2005 Inpainting the colors
abstract
A framework for automatic image colorization, the art of adding color to a monochrome image or movie, is presented in this paper. The approach is based on considering the geometry and structure of the monochrome luminance input, given by its gradient information, as representing the geometry and structure of the whole colored version. The color is then obtained by solving a partial differential equation that propagates a few color scribbles provided by the user or by side information, while considering the gradient information brought in by the monochrome data. This way, the color is inpainted, constrained both by the monochrome image geometry and the provided color samples. We present the underlying framework and examples for still images and movies.
Guillermo Sapiro
ICIP (2)1
2005 Level Set and PDE Methods for Visualization
David E. Breen, Robert M. Kirby, Aaron E. Lefohn, Ken Museth, Tobias Preußer, Guillermo Sapiro, Ross T. Whitaker
IEEE Visualization6
2005 Three-dimensional shape rendering from multiple images
Alberto Bartesaghi, Guillermo Sapiro, Thomas Malzbender, Dan Gelb
Graph. Model.2
2005 Fast image and video denoising via nonlocal means of similar neighborhoods
abstract
In this letter, improvements to the nonlocal means image denoising method introduced by Buades et al. are presented. The original nonlocal means method replaces a noisy pixel by the weighted average of pixels with related surrounding neighborhoods. While producing state-of-the-art denoising results, this method is computationally impractical. In order to accelerate the algorithm, we introduce filters that eliminate unrelated neighborhoods from the weighted average. These filters are based on local average gray values and gradients, preclassifying neighborhoods and thereby reducing the original quadratic complexity to a linear one and reducing the influence of less-related areas in the denoising of a given pixel. We present the underlying framework and experimental results for gray level and color images as well as for video.
Mona Mahmoudi, Guillermo Sapiro
IEEE Signal Process. Lett.2
2005 An Energy-Based Three-Dimensional Segmentation Approach for the Quantitative Interpretation of Electron Tomograms
abstract
Electron tomography allows for the determination of the three-dimensional structures of cells and tissues at resolutions significantly higher than that which is possible with optical microscopy. Electron tomograms contain, in principle, vast amounts of information on the locations and architectures of large numbers of subcellular assemblies and organelles. The development of reliable quantitative approaches for the analysis of features in tomograms is an important problem, and a challenging prospect due to the low signal-to-noise ratios that are inherent to biological electron microscopic images. This is, in part, a consequence of the tremendous complexity of biological specimens. We report on a new method for the automated segmentation of HIV particles and selected cellular compartments in electron tomograms recorded from fixed, plastic-embedded sections derived from HIV-infected human macrophages. Individual features in the tomogram are segmented using a novel robust algorithm that finds their boundaries as global minimal surfaces in a metric space defined by image features. The optimization is carried out in a transformed spherical domain with the center an interior point of the particle of interest, providing a proper setting for the fast and accurate minimization of the segmentation energy. This method provides tools for the semi-automated detection and statistical evaluation of HIV particles at different stages of assembly in the cells and presents opportunities for correlation with biochemical markers of HIV infection. The segmentation algorithm developed here forms the basis of the automated analysis of electron tomograms and will be especially useful given the rapid increases in the rate of data acquisition. It could also enable studies of much larger data sets, such as those which might be obtained from the tomographic analysis of HIV-infected cells from studies of large populations.
Alberto Bartesaghi, Guillermo Sapiro, Sriram Subramaniam
IEEE Trans. Image Process.2
2005 Is image steganography natural?
abstract
Steganography is the art of secret communication. Its purpose is to hide the presence of information, using, for example, images as covers. We experimentally investigate if stego-images, bearing a secret message, are statistically "natural." For this purpose, we use recent results on the statistics of natural images and investigate the effect of some popular steganography techniques. We found that these fundamental statistics of natural images are, in fact, generally altered by the hidden "nonnatural" information. Frequently, the change is consistently biased in a given direction. However, for the class of natural images considered, the change generally falls within the intrinsic variability of the statistics, and, thus, does not allow for reliable detection, unless knowledge of the data hiding process is taken into account. In the latter case, significant levels of detection are demonstrated.
Alvaro Martín, Guillermo Sapiro, Gadiel Seroussi
IEEE Trans. Image Process.2
2004 Non-photorealistic rendering from multiple images
abstract
A new paradigm for automatic nonphotorealistic rendering (NPR) is introduced in this paper. Existing NPR approaches can be categorized in two groups depending on the type of input they use: image based and object based. Using multiple images as input to the NPR scheme, we propose a novel hybrid model that simultaneously uses information from the image and object domains. The benefit not only comes from combining the features of each approach, but most important, it minimizes the need for manual or user assisted tasks in extracting scene features and geometry, as employed in virtually all state-of-the-art NPR approaches. We describe a particular implementation of such an hybrid system and present a number of automatically generated pen-and-ink style drawings. This work then shows how to use and extend well developed techniques in computer vision to address fundamental problems in image representation and rendering.
Alberto Bartesaghi, Guillermo Sapiro, Thomas Malzbender, Dan Gelb
ICIP2
2004 Automatic image decomposition
abstract
The decomposition of an image into its primitive components, such as cartoon plus texture, is a fundamental problem in image processing. In previous works, various authors have proposed a technique to achieve this decomposition into structure and texture. These two components are competing ones, and their proposed model has a critical parameter that controls this decomposition. In this paper, we show how to automatically select this parameter, and demonstrate with examples the importance of this optimal selection.
Kedar A. Patwardhan, Guillermo Sapiro
ICIP2
2004 Lightfield completion
Liron Yatziv, Guillermo Sapiro, Marc Levoy
ICIP2
2004 Comparing Point Clouds
Facundo Mémoli, Guillermo Sapiro
Symposium on Geometry Processing2
2004 Area-Based Medial Axis of Planar Curves
Marc Niethammer, Santiago Betelú, Guillermo Sapiro, Allen R. Tannenbaum, Peter J. Giblin
Int. J. Comput. Vis.3
2004 Morse description and geometric encoding of digital elevation maps
abstract
Two complementary geometric structures for the topographic representation of an image are developed in this work. The first one computes a description of the Morse-topological structure of the image, while the second one computes a simplified version of its drainage structure. The topographic significance of the Morse and drainage structures of digital elevation maps (DEMs) suggests that they can been used as the basis of an efficient encoding scheme. As an application, we combine this geometric representation with an interpolation algorithm and lossless data compression schemes to develop a compression scheme for DEMs. This algorithm achieves high compression while controlling the maximum error in the decoded elevation map, a property that is necessary for the majority of applications dealing with DEMs. We present the underlying theory and compression results for standard DEM data.
Andres Fco. Solé, Vicent Caselles, Guillermo Sapiro, Francesc Aràndiga
IEEE Trans. Image Process.3
2003 Simultaneous Structure and Texture Image Inpainting
abstract
An algorithm for the simultaneous filling-in of texture and structure in regions of missing image information is presented. The basic idea is to first decompose the image into the sum of two functions with different basic characteristics, and then reconstruct each one of these functions separately with structure and texture filling-in algorithms. The first function used in the decomposition is of bounded variation, representing the underlying image structure, while the second function captures the texture and possible noise. The region of missing information in the bounded variation image is reconstructed using image inpainting algorithms, while the same region in the texture image is filled-in with texture synthesis techniques. The original image is then reconstructed adding back these two sub-images. The novel contribution of the paper is then in the combination of these three previously developed components: image decomposition with inpainting and texture synthesis, which permits the simultaneous use of filling-in algorithms that are suited for different image characteristics. Examples on real images show the advantages of this proposed approach.
Marcelo Bertalmío, Luminita A. Vese, Guillermo Sapiro, Stanley J. Osher
CVPR (2)3
2003 Image filling-in in a decomposition space
abstract
An algorithm for the simultaneous filling-in of texture and structure in regions of missing image information is presented in this paper. The basic idea is to first decompose the image into the sum of two functions with different basic characteristics, and then reconstruct each one of these functions separately with structure and texture filling-in algorithms. The first function used in the decomposition is of bounded variation, representing the underlying image structure, while the second function captures the texture and possible noise. The region of missing information in the bounded variation image is reconstructed using image inpainting algorithms, while the same region in the texture image is filled-in with texture synthesis techniques. The original image is then reconstructed adding back these two subimages. The novel contribution of this paper is then in the combination of these three previously developed components, image decomposition with in-painting and texture synthesis, which permits the simultaneous use of filling-in algorithms that are suited for different image characteristics. The novelty in the approach is to perform filling-in in a domain different from the original given image space. Examples on real images show the advantages of this proposed approach.
Marcelo Bertalmío, Luminita A. Vese, Guillermo Sapiro, Stanley J. Osher
ICIP (1)3
2003 Projection based image and video inpainting using wavelets
abstract
In this paper, we present a technique for automatic color image inpainting, the art of modifying an image-region in a nondetectable form. The main algorithm is based on the theory of projections onto convex sets (POCS). The image and its wavelet transform are projected onto each other after applying suitable constraints in each domain. This technique exploits the frequency-spatial representation provided by wavelets and utilizes the correlation between the damaged area in the image and its neighborhood. The resulting restored area is homogeneous with its surrounding and preserves the aesthetics of the image. The same technique is used for simple video restoration problems. Video frames are stacked and treated as a 3D volume, making a natural use of interframe correlation.
Kedar A. Patwardhan, Guillermo Sapiro
ICIP (1)2
2003 Color histogram equalization through mesh deformation
abstract
In this paper we propose an extension of grayscale histogram equalization for color images. For aesthetic reasons, previously proposed color histogram equalization techniques do not generate uniform color histograms. Our method will always generate an almost uniform color histogram thus making an optimal use of the color space. This is particularly interesting for pseudo-color scientific visualization. The method is based on deforming a mesh in color space to fit the existing histogram and then map it to a uniform histogram. It is a natural extension of grayscale histogram equalization and it can be applied to spatial and color space of any dimension.
Eric Pichon, Marc Niethammer, Guillermo Sapiro
ICIP (2)3
2003 Morse description and geometric encoding of DEM data
abstract
Two complementary geometric structures for the topographic representation of an image are developed in this work. The first one computes a description of the Morse structure of the image, while the second one computes a simplified version of its drainage structure. The topographic significance of the Morse and drainage structures of digital elevation maps (DEM) suggests that they can been used as the basis of an efficient encoding scheme. We combine this geometric representation with an interpolation algorithm and lossless data compression schemes to develop a compression scheme for DEM. This algorithm permits to obtain compression results while controlling the maximum error in the decoded elevation map, a property that is necessary for the majority of applications dealing with DEM.
Andres Fco. Solé, Vicent Caselles, Guillermo Sapiro, Francesc Aràndiga
ICIP (2)3
2003 Inpainting surface holes
abstract
An algorithm for filling-in surface holes is introduced in this paper. The basic idea is to represent the surface of interest in implicit form, and fill-in the holes with a system of geometric partial differential equations derived from image inpainting algorithms. The framework and examples with synthetic and real data are presented.
Joan Verdera, Vicent Caselles, Marcelo Bertalmío, Guillermo Sapiro
ICIP (2)4
2003 Three-Dimensional Segmentation of Brain Aneurysms in CTA Using Non-parametric Region-Based Information and Implicit Deformable Models: Method and Evaluation
Monica Hernandez, Alejandro F. Frangi, Guillermo Sapiro
MICCAI (2)3
2003 Simultaneous structure and texture image inpainting
abstract
An algorithm for the simultaneous filling-in of texture and structure in regions of missing image information is presented in this paper. The basic idea is to first decompose the image into the sum of two functions with different basic characteristics, and then reconstruct each one of these functions separately with structure and texture filling-in algorithms. The first function used in the decomposition is of bounded variation, representing the underlying image structure, while the second function captures the texture and possible noise. The region of missing information in the bounded variation image is reconstructed using image inpainting algorithms, while the same region in the texture image is filled-in with texture synthesis techniques. The original image is then reconstructed adding back these two sub-images. The novel contribution of this paper is then in the combination of these three previously developed components, image decomposition with inpainting and texture synthesis, which permits the simultaneous use of filling-in algorithms that are suited for different image characteristics. Examples on real images show the advantages of this proposed approach.
Marcelo Bertalmío, Luminita A. Vese, Guillermo Sapiro, Stanley J. Osher
IEEE Trans. Image Process.3
2003 Visualization of high dynamic range images
abstract
A novel paradigm for information visualization in high dynamic range images is presented in this paper. These images, real or synthetic, have luminance with typical ranges many orders of magnitude higher than that of standard output/viewing devices, thereby requiring some processing for their visualization. In contrast with existent approaches, which compute a single image with reduced range, close in a given sense to the original data, we propose to look for a representative set of images. The goal is then to produce a minimal set of images capturing the information all over the high dynamic range data, while at the same time preserving a natural appearance for each one of the images in the set. A specific algorithm that achieves this goal is presented and tested on natural and synthetic data.
Alvaro Pardo, Guillermo Sapiro
IEEE Trans. Image Process.2
2003 Structure and texture filling-in of missing image blocks in wireless transmission and compression applications
abstract
An approach for filling-in blocks of missing data in wireless image transmission is presented. When compression algorithms such as JPEG are used as part of the wireless transmission process, images are first tiled into blocks of 8 x 8 pixels. When such images are transmitted over fading channels, the effects of noise can destroy entire blocks of the image. Instead of using common retransmission query protocols, we aim to reconstruct the lost data using correlation between the lost block and its neighbors. If the lost block contained structure, it is reconstructed using an image inpainting algorithm, while texture synthesis is used for the textured blocks. The switch between the two schemes is done in a fully automatic fashion based on the surrounding available blocks. The performance of this method is tested for various images and combinations of lost blocks. The viability of this method for image compression, in association with lossy JPEG, is also discussed.
Shantanu Rane, Guillermo Sapiro, Marcelo Bertalmío
IEEE Trans. Image Process.2
2003 Texture Synthesis for 3D Shape Representation
abstract
Considerable evidence suggests that a viewer's perception of the 3D shape of a polygonally-defined object can be significantly affected (either masked or enhanced) by the presence of a surface texture pattern. However, investigations into the specific mechanisms of texture's effect on shape perception are still ongoing and the question of how to design and apply a texture pattern to a surface in order to best facilitate shape perception remains open. Recently, we have suggested that, for anisotropic texture patterns, the accuracy of shape judgments may be significantly affected by the orientation of the surface texture pattern anisotropy with respect to the principal directions of curvature over the surface. However, it has been difficult, until this time, to conduct controlled studies specifically investigating the effect of texture orientation on shape perception because there has been no simple and reliable method for texturing an arbitrary doubly curved surface with a specified input pattern such that the dominant orientation of the pattern everywhere follows a predefined directional vector field over the surface, while seams and projective distortion of the pattern are avoided. In this paper, we present a straightforward and highly efficient method for achieving such a texture and describe how it can potentially be used to enhance shape representation. Specifically, we describe a novel, efficient, automatic algorithm for seamlessly synthesizing, from a sample 2D pattern, a high resolution fitted surface texture in which the dominant orientation of the pattern locally follows a specified vector field over the surface at a per-pixel level and in which seams, projective distortion, and repetition artifacts in the texture pattern are nearly completely avoided. We demonstrate the robustness of our method with a variety of texture swatches applied to standard graphics data sets and we explain how our method can be used to facilitate research in the perception of shape from texture.
Gabriele Gorla, Victoria Interrante, Guillermo Sapiro
IEEE Trans. Vis. Comput. Graph.3
2002 Visualization of high dynamic range images
abstract
A novel paradigm for information visualization in high dynamic range images is presented. These images, real or synthetic, have luminance with typical ranges many orders of magnitude higher than that of standard output devices, thereby requiring some processing for their visualization. In contrast to existing approaches, which compute a single image with reduced range, close in a given sense to the original data, we propose to look for a representative set of images. The goal is then to produce a minimal set of images capturing the information all over the high dynamic range data, while at the same time preserving a natural appearance for each one of the images in the set. A specific algorithm that achieves this goal is presented and tested on natural and synthetic data.
Alvaro Pardo, Guillermo Sapiro
ICIP (1)2
2002 Structure and texture filling-in of missing image blocks in wireless transmission and compression
abstract
An approach for filling-in blocks of missing data in wireless image transmission is presented in this paper. When compression algorithms such as JPEG are used as part of the wireless transmission process, images are first tiled into blocks of 8/spl times/8 pixels. When such images are transmitted over fading channels, the effects of noise can kill entire blocks of the image. Instead of using common retransmission query protocols, we aim to reconstruct the lost data using correlation between the lost block and its neighbors. If the lost block contained structure, it is reconstructed using an image inpainting algorithm, while texture synthesis is used for the textured blocks. The switch between the two schemes is done in a fully automatic fashion based on the surrounding available blocks. The performance of this method is tested for various images and combinations of lost blocks. The viability of this method for image compression, in association with lossy JPEG, is also discussed.
Shantanu Rane, Marcelo Bertalmío, Guillermo Sapiro
ICIP (1)3
2002 Wavelet-domain reconstruction of lost blocks in wireless image transmission and packet-switched networks
abstract
A fast scheme for wavelet-domain interpolation of lost image blocks in wireless image transmission is presented in this paper. In the transmission of block-coded images, fading in wireless channels and congestion in packet-switched networks can cause entire blocks to be lost. Instead of using retransmission query protocols, we reconstruct the lost block in the wavelet-domain using the correlation between the lost block and its neighbors. The algorithm first uses simple thresholding to determine the presence or absence of edges in the lost block. This is followed by an interpolation scheme, designed to minimize the blockiness effect, while preserving the edges or texture in the interior of the block. The interpolation scheme minimizes the square of the error between the border coefficients of the lost block and those of its neighbors, at each transform scale. The performance of the algorithm on standard test images, its low computational overhead at the decoder, and its performance vis-a-vis other reconstruction schemes, is discussed.
Shantanu Rane, Jeremiah Remus, Guillermo Sapiro
ICIP (1)3
2002 Special Issue on Partial Differential Equations in Image Processing, Computer Vision, and Computer Graphics
Olivier D. Faugeras, Pietro Perona, Guillermo Sapiro
J. Vis. Commun. Image Represent.3
2001 Navier-Stokes, Fluid Dynamics, and Image and Video Inpainting
abstract
Image inpainting involves filling in part of an image or video using information from the surrounding area. Applications include the restoration of damaged photographs and movies and the removal of selected objects. We introduce a class of automated methods for digital inpainting. The approach uses ideas from classical fluid dynamics to propagate isophote lines continuously from the exterior into the region to be inpainted. The main idea is to think of the image intensity as a 'stream function for a two-dimensional incompressible flow. The Laplacian of the image intensity plays the role of the vorticity of the fluid; it is transported into the region to be inpainted by a vector field defined by the stream function. The resulting algorithm is designed to continue isophotes while matching gradient vectors at the boundary of the inpainting region. The method is directly based on the Navier-Stokes equations for fluid dynamics, which has the immediate advantage of well-developed theoretical and numerical results. This is a new approach for introducing ideas from computational fluid dynamics into problems in computer vision and image analysis.
Marcelo Bertalmío, Andrea L. Bertozzi, Guillermo Sapiro
CVPR (1)3
2001 A Variational Model for Filling-In Gray Level and Color Images
Coloma Ballester, Vicent Caselles, Joan Verdera, Marcelo Bertalmío, Guillermo Sapiro
ICCV5
2001 Affine Invariant Erosion of 3D Shapes
abstract
A new definition of affine invariant erosion of 3D surfaces is introduced. Instead of being based in terms of Euclidean distances, the volumes enclosed between the surface and its chords are used. The resulting erosion is insensitive to noise, and by construction, it is affine invariant. We prove some key properties about this erosion operation, and we propose a simple method to compute the erosion of implicit surfaces. We also discuss how the affine erosion can be used to define 3D affine invariant robust skeletons.
Santiago Betelú, Guillermo Sapiro, Allen R. Tannenbaum
ICCV2
2001 Missile tracking using knowledge-based adaptive thresholding
abstract
We apply a knowledge-based segmentation method developed for still and video images to the problem of tracking missiles and high speed projectiles. Since we are only interested in segmenting a portion of the missile (namely, the nose cone), we use our segmentation procedure as a method of adapting thresholding. The key idea is to utilize a priori knowledge about the objects present in the image, e.g. missile and background, introduced via Bayes' rule. Posterior probabilities obtained in this way are anisotropically smoothed, and the image segmentation is obtained via MAP classifications of the smoothed data. When segmenting sequences of images, the smoothed posterior probabilities of past frames are used as prior distributions in succeeding frames.
Steven Haker, Guillermo Sapiro, Allen R. Tannenbaum, Donald Washburn
ICIP (1)2
2001 Crease Enhancement Diffusion
Andres Fco. Solé, Antonio M. López 0001, Guillermo Sapiro
Comput. Vis. Image Underst.3
2001 On the computation of the affine skeletons of planar curves and the detection of skew symmetry
Santiago Betelú, Guillermo Sapiro, Allen R. Tannenbaum, Peter J. Giblin
Pattern Recognit.2
2001 Vector probability diffusion
abstract
The basic motivation of this work is to introduce contextual information into image segmentation tasks by adding spatial coherence to the posterior probabilities corresponding to the classes present in the scene. A method for isotropic and anisotropic diffusion of vector probabilities in general, and posterior probabilities in particular, is introduced. The technique is based on diffusing via coupled partial differential equations restricted to the semi-hyperplane corresponding to probability functions. Both the partial differential equations and their corresponding numerical implementation guarantee that the vector remains a probability vector, having all its components positive and adding to one. Applying the method to posterior probabilities in classification problems, spatial and contextual coherence is introduced before the maximum a posteriori (MAP) decision, thereby improving the classification results.
Alvaro Pardo, Guillermo Sapiro
IEEE Signal Process. Lett.2
2001 Evaluation of JPEG-LS, the new lossless and controlled-lossy still image compression standard, for compression of high-resolution elevation data
abstract
The compression of elevation data is studied. The performance of JPEG-LS, the new international ISO/ITU standard for lossless and near-lossless (controlled-lossy) still-image compression, is investigated both for data from the USGS digital elevation model (DEM) database and the navy-provided digital terrain model (DTM) data. Using JPEG-LS has the advantage of working with a standard algorithm. Moreover, in contrast with algorithms like the popular JPEG-lossy standard, this algorithm permits the completely lossless compression of the data as well as a controlled lossy mode where a sharp upper bound on the elevation error is selected by the user. All these are achieved at a very low computational complexity. In addition to these algorithmic advantages, they show that JPEG-LS achieves significantly better compression results than those obtained with other (nonstandard) algorithms previously investigated for the compression of elevation data. The results here reported suggest that JPEG-LS can immediately be adopted for the compression of elevation data for a number of applications.
Shantanu Rane, Guillermo Sapiro
IEEE Trans. Geosci. Remote. Sens.2
2001 Filling-in by joint interpolation of vector fields and gray levels
abstract
A variational approach for filling-in regions of missing data in digital images is introduced. The approach is based on joint interpolation of the image gray levels and gradient/isophotes directions, smoothly extending in an automatic fashion the isophote lines into the holes of missing data. This interpolation is computed by solving the variational problem via its gradient descent flow, which leads to a set of coupled second order partial differential equations, one for the gray-levels and one for the gradient orientations. The process underlying this approach can be considered as an interpretation of the Gestaltist's principle of good continuation. No limitations are imposed on the topology of the holes, and all regions of missing data can be simultaneously processed, even if they are surrounded by completely different structures. Applications of this technique include the restoration of old photographs and removal of superimposed text like dates, subtitles, or publicity. Examples of these applications are given. We conclude the paper with a number of theoretical results on the proposed variational approach and its corresponding gradient descent flow.
Coloma Ballester, Marcelo Bertalmío, Vicent Caselles, Guillermo Sapiro, Joan Verdera
IEEE Trans. Image Process.4
2001 Color image enhancement via chromaticity diffusion
abstract
A novel approach for color image denoising is proposed in this paper. The algorithm is based on separating the color data into chromaticity and brightness, and then processing each one of these components with partial differential equations or diffusion flows. In the proposed algorithm, each color pixel is considered as an n-dimensional vector. The vectors' direction, a unit vector, gives the chromaticity, while the magnitude represents the pixel brightness. The chromaticity is processed with a system of coupled diffusion equations adapted from the theory of harmonic maps in liquid crystals. This theory deals with the regularization of vectorial data, while satisfying the intrinsic unit norm constraint of directional data such as chromaticity. Both isotropic and anisotropic diffusion flows are presented for this n-dimensional chromaticity diffusion flow. The brightness is processed by a scalar median filter or any of the popular and well established anisotropic diffusion flows for scalar image enhancement. We present the underlying theory, a number of examples, and briefly compare with the current literature.
Bei Tang, Guillermo Sapiro, Vicent Caselles
IEEE Trans. Image Process.2
2001 Anisotropic 2D and 3D Averaging of fMRI Signals
abstract
A novel method for denoising functional magnetic resonance imaging temporal signals is presented in this note. The method is based on progressively enhancing the temporal signal by means of adaptive anisotropic spatial averaging. This average is based on a new metric for comparing temporal signals corresponding to active fMRI regions. Examples are presented both for simulated and real two and three-dimensional data. The software implementing the proposed technique is publicly available for the research community.
Andres Fco. Solé, Shing-Chung Ngan, Guillermo Sapiro, Xiaoping Hu 0001, Antonio M. López 0001
IEEE Trans. Medical Imaging3
2000 Noise-Resistant Affine Skeletons of Planar Curves
Santiago Betelú, Guillermo Sapiro, Allen R. Tannenbaum, Peter J. Giblin
ECCV (1)2
2000 Segmentation-Free Skeletonization of Gray-Scale Images via PDE's
abstract
A simple approach to compute the skeletons of gray-scale images using partial differential equations is presented. The proposed scheme works directly on the gray-scale images, without the necessity of pre-segmentation (binarization), or the addition of shock capturing schemes. This is accomplished by deforming the given image according to a family of modified continuous-scale erosion/dilation equations. With the scheme here proposed, the skeleton of multiple objects can be simultaneously computed. Examples on synthetic and real images are provided.
Do Hyun Chung, Guillermo Sapiro
ICIP2
2000 Segmenting Skin Lesions with Partial Differential Equations Based Image Processing Algorithms
abstract
A PDE-based system for detecting the boundary of skin lesions in digital clinical skin images is presented. The image is first pre-processed via contrast-enhancement and anisotropic diffusion. If the lesion is covered by hairs, a PDE-based continuous morphological filter that removes them is used as an additional pre-processing step. Following these steps, the skin lesion is segmented either by the geodesic active contours model or the geodesic edge tracing approach. These techniques are based on computing, again via PDE's, a geodesic curve in a space defined by the image content. Examples showing the performance of the algorithm are given.
Do Hyun Chung, Guillermo Sapiro
ICIP2
2000 Using Anisotropic Diffusion of Probability Maps for Activity Detection in Block-Design Functional MRI
abstract
A new approach for improving the detection of pixels associated with neural activity in functional magnetic resonance imaging (fMRI) is presented. We propose to use anisotropic diffusion to exploit the spatial correlation between the active pixels in functional MRI. Specifically, in this paper the anisotropic diffusion flow is applied to a probability image, obtained either from t-map statistics or via Bayes rule. In general, this information diffusion technique can be incorporated into other activity detection algorithms before the active/non-active hard decision is made. Examples with simulated and real data show improvements over classical techniques.
Hong Shan Neoh, Guillermo Sapiro
ICIP2
2000 Vector Probability Diffusion
abstract
A method for isotropic and anisotropic diffusion of vector probabilities in general, and posterior probabilities in particular, is introduced. The technique is based on diffusing via coupled partial differential equations restricted to the semi-hyperplane corresponding to probability functions. Both the partial differential equations and their corresponding numerical implementation guarantee that the vector remains a probability vector, having all its components positive and adding to one. Applying the method to posterior probabilities in classification problems, spatial and contextual coherence is introduced before the MAP decision, thereby improving the classification results.
Alvaro Pardo, Guillermo Sapiro
ICIP2
2000 Chromaticity Diffusion
abstract
A novel approach for color image denoising is proposed. The algorithm is based on separating the color data into chromaticity and brightness, and then processing each one of these components with partial differential equations or diffusion flows. In the proposed algorithm, each color pixel is considered as an n-dimensional vector. The vectors' direction, a unit vector, gives the chromaticity, while the magnitude represents the pixel brightness. The chromaticity is processed with a system of coupled diffusion equations adapted from the theory of harmonic maps in liquid crystals. This theory deals with the regularization of vectorial data, while satisfying the intrinsic unit norm constraint of directional data such as chromaticity. Both isotropic and anisotropic diffusion flows are presented for this n-dimensional chromaticity diffusion flow. The brightness is processed by a scalar median filter or any of the popular and well established anisotropic diffusion flows for scalar image enhancement. We present the underlying theory, a number of examples, and comparison with the current literature.
Bei Tang, Guillermo Sapiro, Vicent Caselles
ICIP2
2000 Image inpainting
abstract
Inpainting, the technique of modifying an image in an undetectable form, is as ancient as art itself. The goals and applications of inpainting are numerous, from the restoration of damaged paintings and photographs to the removal/replacement of selected objects. In this paper, we introduce a novel algorithm for digital inpainting of still images that attempts to replicate the basic techniques used by professional restorators. After the user selects the regions to be restored, the algorithm automatically fills-in these regions with information surrounding them. The fill-in is done in such a way that isophote lines arriving at the regions' boundaries are completed inside. In contrast with previous approaches, the technique here introduced does not require the user to specify where the novel information comes from. This is automatically done (and in a fast way), thereby allowing to simultaneously fill-in numerous regions containing completely different structures and surrounding backgrounds. In addition, no limitations are imposed on the topology of the region to be inpainted. Applications of this technique include the restoration of old photographs and damaged film; removal of superimposed text like dates, subtitles, or publicity; and the removal of entire objects from the image like microphones or wires in special effects.
Marcelo Bertalmío, Guillermo Sapiro, Vicent Caselles, Coloma Ballester
SIGGRAPH2
2000 Diffusion of General Data on Non-Flat Manifolds via Harmonic Maps Theory: The Direction Diffusion Case
Bei Tang, Guillermo Sapiro, Vicent Caselles
Int. J. Comput. Vis.2
2000 Special Issue on the Second International Conference on Scale Space Theory in Computer Vision: Guest Editors' Comments
Olivier D. Faugeras, Mads Nielsen, Pietro Perona, Bart M. ter Haar Romeny, Guillermo Sapiro
J. Vis. Commun. Image Represent.5
2000 Call for Papers: Special Issue on Partial Differential Equations in Image Processing, Computer Vision, and Computer Graphics
Olivier D. Faugeras, Pietro Perona, Guillermo Sapiro
J. Vis. Commun. Image Represent.3
2000 Morphing Active Contours
abstract
A method for deforming curves in a given image to a desired position in a second image is introduced. The algorithm is based on deforming the first image toward the second one via a partial differential equation (PDE), while tracking the deformation of the curves of interest in the first image with an additional, coupled PDE; both the images and the curves on the frame/slices of interest are used for tracking. The technique can be applied to object tracking and sequential segmentation. The topology of the deforming curve can change without any special topology handling procedures added to the scheme. This permits, for example, the automatic tracking of scenes where, due to occlusions, the topology of the objects of interest changes from frame to frame. In addition, this work introduces the concept of projecting velocities to obtain systems of coupled PDEs for image analysis applications. We show examples for object tracking and segmentation of electronic microscopy.
Marcelo Bertalmío, Guillermo Sapiro, Gregory Randall
IEEE Trans. Pattern Anal. Mach. Intell.2
2000 On the level lines and geometry of vector-valued images
abstract
In this letter, we extend the concept of level lines of scalar images to vector-valued data. Consistent with the scalar case, we define the level-lines of vector-valued images as the integral curves of the directions of minimal vectorial change. This direction, and the magnitude of the change, are computed using classical Riemannian geometry. As an example of the use of this new concept, we show how to visualize the basic geometry of vector-valued images with a scalar image.
Do Hyun Chung, Guillermo Sapiro
IEEE Signal Process. Lett.2
2000 Knowledge-based segmentation of SAR data with learned priors
abstract
An approach for the segmentation of still and video synthetic aperture radar (SAR) images is described. A priori knowledge about the objects present in the image, e.g., target, shadow and background terrain, is introduced via Bayes' rule. Posterior probabilities obtained in this way are then anisotropically smoothed, and the image segmentation is obtained via MAP classifications of the smoothed data. When segmenting sequences of images, the smoothed posterior probabilities of past frames are used to learn the prior distributions in the succeeding frame. We show with examples from public data sets that this method provides an efficient and fast technique for addressing the segmentation of SAR data.
Steven Haker, Guillermo Sapiro, Allen R. Tannenbaum
IEEE Trans. Image Process.2
2000 The LOCO-I lossless image compression algorithm: principles and standardization into JPEG-LS
abstract
LOCO-I (LOw COmplexity LOssless COmpression for Images) is the algorithm at the core of the new ISO/ITU standard for lossless and near-lossless compression of continuous-tone images, JPEG-LS. It is conceived as a "low complexity projection" of the universal context modeling paradigm, matching its modeling unit to a simple coding unit. By combining simplicity with the compression potential of context models, the algorithm "enjoys the best of both worlds." It is based on a simple fixed context model, which approaches the capability of the more complex universal techniques for capturing high-order dependencies. The model is tuned for efficient performance in conjunction with an extended family of Golomb-type codes, which are adaptively chosen, and an embedded alphabet extension for coding of low-entropy image regions. LOCO-I attains compression ratios similar or superior to those obtained with state-of-the-art schemes based on arithmetic coding. Moreover, it is within a few percentage points of the best available compression ratios, at a much lower complexity level. We discuss the principles underlying the design of LOCO-I, and its standardization into JPEC-LS.
Marcelo J. Weinberger, Gadiel Seroussi, Guillermo Sapiro
IEEE Trans. Image Process.3
2000 Conformal Surface Parameterization for Texture Mapping
abstract
We give an explicit method for mapping any simply connected surface onto the sphere in a manner which preserves angles. This technique relies on certain conformal mappings from differential geometry. Our method provides a new way to automatically assign texture coordinates to complex undulating surfaces. We demonstrate a finite element method that can be used to apply our mapping technique to a triangulated geometric description of a surface.
Steven Haker, Sigurd B. Angenent, Allen R. Tannenbaum, Ron Kikinis, Guillermo Sapiro, Michael Halle
IEEE Trans. Vis. Comput. Graph.5
1999 Direction Diffusion
abstract
In a number of disciplines, directional data provides a fundamental source of information. A novel framework for isotropic and anisotropic diffusion of directions is presented in this paper. The framework can be applied both to regularize directional data and to obtain multiscale representations of it. The basic idea is to apply and extend results from the theory of harmonic maps in liquid crystals. This theory deals with the regularization of vectorial data, while satisfying the unit norm constraint of directional data. We show the corresponding variational and partial differential equations formulations for isotropic diffusion, obtained from an L/sub 2/ norm, and edge preserving diffusion, obtained from an L/sub 1/ norm. In contrast with previous approaches, the framework is valid for directions in any dimensions, supports non-smooth data, and gives both isotropic and anisotropic formulations. We present a number of theoretical results, open questions, and examples for gradient vectors, optical flow, and color images.
Bei Tang, Guillermo Sapiro, Vicent Caselles
ICCV2
1999 Vector Median Filters, Morphology, and PDE's: Theoretical Connections
abstract
We formally connect between vector median filters, morphological operators, and geometric partial differential equations. Considering a lexicographic order, which permits us to define an order between vectors in IR/sup N/, we first show that the vector median filter of a vector valued image is equivalent to a collection of infimum-supremum morphological operations. We then proceed and study the asymptotic behavior of this filter. We also provide an interpretation of the infinitesimal iteration of this vectorial median filter in terms of systems of coupled geometric partial differential equations. The main component of the vector evolves according to curvature motion, while, intuitively, the others regularly deform their level sets toward those of this main component. These results extend to the vector case classical connections between scalar median filters, mathematical morphology, and mean curvature motion.
Vicent Caselles, Guillermo Sapiro, Do Hyun Chung
ICIP (4)2
1999 From LOCO-I to the JPEG-LS Standard
abstract
LOGO-I (LOw COmplexity LOssless COmpression for Images) is the algorithm at the core of the new ISO/ITU standard for lossless and near-lossless compression of continuous-tone images, JPEG-LS. The algorithm was conceived as a "low complexity projection" of the universal context modeling paradigm, matching its modeling unit to a simple coding unit based on Golomb codes. The JPEG-LS standard evolved after successive refinements of the core algorithm, and a description of its design principles and main algorithmic components is presented in this paper. LOCO-I/JPEG-LS attains compression ratios similar or superior to those obtained with state-of-the-art schemes based on arithmetic coding. Moreover, it is within a few percentage points of the best available compression ratios, at a much lower complexity level.
Marcelo J. Weinberger, Gadiel Seroussi, Guillermo Sapiro
ICIP (4)3
1999 Color and Illuminant Voting
abstract
A geometric-vision approach to color constancy and illuminant estimation is presented in this paper. We show a general framework, based on ideas from the generalized probabilistic Hough transform, to estimate the illuminant and reflectance of natural images. Each image pixel "votes" for possible illuminants and the estimation is based on cumulative votes. The framework is natural for the introduction of physical constraints in the color constancy problem. We show the relationship of this work to previous algorithms for color constancy and present examples.
Guillermo Sapiro
IEEE Trans. Pattern Anal. Mach. Intell.1
1999 Shape preserving local histogram modification
abstract
A novel approach for shape preserving contrast enhancement is presented in this paper. Contrast enhancement is achieved by means of a local histogram equalization algorithm which preserves the level-sets of the image. This basic property is violated by common local schemes, thereby introducing spurious objects and modifying the image information. The scheme is based on equalizing the histogram in all the connected components of the image, which are defined based both on the grey-values and spatial relations between pixels in the image, and following mathematical morphology, constitute the basic objects in the scene. We give examples for both grey-value and color images.
Vicent Caselles, Jose Luis Lisani, Jean-Michel Morel, Guillermo Sapiro
IEEE Trans. Image Process.4
1999 Region Tracking on Level-Sets Methods
abstract
Since the work by Osher and Sethian on level-sets algorithms for numerical shape evolutions, this technique has been used for a large number of applications in numerous fields. In medical imaging, this numerical technique has been successfully used, for example, in segmentation and cortex unfolding algorithms. The migration from a Lagrangian implementation to a Eulerian one via implicit representations or level-sets brought some of the main advantages of the technique, i.e., topology independence and stability. This migration means also that the evolution is parametrization free. Therefore, we do not know exactly how each part of the shape is deforming and the point-wise correspondence is lost. In this note we present a technique to numerically track regions on surfaces that are being deformed using the level-sets method. The basic idea is to represent the region of interest as the intersection of two implicit surfaces and then track its deformation from the deformation of these surfaces. This technique then solves one of the main shortcomings of the very useful level-sets approach. Applications include lesion localization in medical images, region tracking in functional MRI (fMRI) visualization, and geometric surface mapping.
Marcelo Bertalmío, Guillermo Sapiro, Gregory Randall
IEEE Trans. Medical Imaging2
1998 Affine Invariant Medial Axis and Skew Symmetry
abstract
Affine invariant medial axes and symmetry sets of planar shapes are introduced and studied in this paper. Two different approaches are presented. The first one is based on affine invariant distances, and defines the symmetry set, a set containing the medial axis; as the closure of the locus of points on (at least) two affine normals an affine-equidistant from the corresponding points on the curve. The second approach is based on affine bitangent conics. In this case the symmetry set is defined as the closure of the locus of centers of conics with (at least) three-point contact with two or more distinct points on the curve. This is equivalent to conic and curve having, at those points, the same affine tangent, or the same Euclidean tangent and curvature. Although the two analogous definitions for the classical Euclidean symmetry set (medial axis) are equivalent, this is not the case for the affine group. We then show how to use the symmetry set to detect affine skew symmetry, proving that the contact based symmetry set is a straight line if and only if the given shape is the affine transformation of a symmetric object.
Peter J. Giblin, Guillermo Sapiro
ICCV2
1998 Bilinear Voting
abstract
A geometric-vision approach to solve bilinear problems in general, and the color constancy and illuminant estimation problem in particular, is presented in this paper. We show a general framework, based on ideas from the generalized (probabilistic) Hough transform, to estimate the unknown variables in the bilinear form. In the case of illuminant and reflectance estimation in natural images, each image pixel "votes" for possible illuminants (or reflectance), and the estimation is based on cumulative votes. In the general case, the voting is for the parameters of the bilinear model. The framework is natural for the introduction of physical constraints. For the case of illuminant estimation, we briefly show the relation of this work with previous algorithms for color constancy, and present examples.
Guillermo Sapiro
ICCV1
1998 Segmentating Cortical Gray Matter for Functional MRI Visualization
abstract
We describe a system that is being used to segment gray matter and create connected cortical representations from MRI. The method exploits knowledge of the anatomy of the cortex and incorporates structural constraints into the segmentation. First, the white matter and CSF regions in the MR volume are segmented using some novel techniques of posterior anisotropic diffusion. Then, the user selects the cortical white matter component of interest, and its structure is verified by checking for cavities and handles. After this, a connected representation of the gray matter is created by a constrained growing-out from the white matter boundary. Because the connectivity is computed, the segmentation can be used as input to several methods of visualizing the spatial pattern of cortical activity within gray matter. In our case, the connected representation of gray matter is used to create a representation of the flattened cortex. Then, fMRI measurements are overlaid on the flattened representation, yielding a representation of the volumetric data within a single image.
P. C. Teo, Guillermo Sapiro, Brian A. Wandell
ICCV2
1998 Morphing Active Contours: A Geometric Approach to Topology-Independent Image Segmentation and Tracking
Marcelo Bertalmío, Guillermo Sapiro, Gregory Randall
ICIP (3)2
1998 Knowledge-based Segmentation of SAR Images
abstract
A new approach for the segmentation of still and video SAR images is described. A priori knowledge about the objects present in the image, e.g., target, shadow, and background terrain, is introduced via Bayes' rule. Posterior probabilities obtained in this way are then anisotropically smoothed, and the image segmentation is obtained via MAP classifications of the smoothed data. When segmenting sequences of images, the smoothed posterior probabilities of past frames are used to learn the prior distributions in the succeeding frame. We show, via a large number of examples from public data sets, that this method provides an efficient and fast technique for addressing the segmentation of SAR data.
Steven Haker, Guillermo Sapiro, Allen R. Tannenbaum
ICIP (1)2
1998 Segmenting Neurons in Electronic Microscopy via Geometric Tracing
Luis Vázquez, Guillermo Sapiro, Gregory Randall
ICIP (3)2
1998 Robust anisotropic diffusion
abstract
Relations between anisotropic diffusion and robust statistics are described in this paper. Specifically, we show that anisotropic diffusion can be seen as a robust estimation procedure that estimates a piecewise smooth image from a noisy input image. The "edge-stopping" function in the anisotropic diffusion equation is closely related to the error norm and influence function in the robust estimation framework. This connection leads to a new "edge-stopping" function based on Tukey's biweight robust estimator that preserves sharper boundaries than previous formulations and improves the automatic stopping of the diffusion. The robust statistical interpretation also provides a means for detecting the boundaries (edges) between the piecewise smooth regions in an image that has been smoothed with anisotropic diffusion. Additionally, we derive a relationship between anisotropic diffusion and regularization with line processes. Adding constraints on the spatial organization of the line processes allows us to develop new anisotropic diffusion equations that result in a qualitative improvement in the continuity of edges.
Michael J. Black, Guillermo Sapiro, David H. Marimont, David Heeger
IEEE Trans. Image Process.2
1997 Robust Anisotropic Diffusion and Sharpening of Scalar and Vector Images
abstract
Relations between anisotropic diffusion and robust statistics are described. We show that anisotropic diffusion can be seen as a robust estimation procedure that estimates a piecewise smooth image from a noisy input image. The "edge-stopping" function in the anisotropic diffusion equation is closely related to the error norm and influence function in the robust estimation framework. This connection leads to a new "edge-stopping" function based on Tukey's biweight robust estimator, that preserves sharper boundaries than previous formulations and improves the automatic stopping of the diffusion. The robust statistical interpretation also provides a means for detecting the boundaries (edges) between the piecewise smooth regions in the image. We extend the framework to vector-valued images and show applications to robust image sharpening.
Michael J. Black, Guillermo Sapiro, David H. Marimont, David Heeger
ICIP (1)2
1997 Shape Preserving Local Contrast Enhancement
abstract
A novel approach for shape preserving contrast enhancement is presented. Contrast enhancement is achieved by means of a local histogram equalization algorithm which preserves the level-sets of the image. This basic property is violated by common local schemes, thereby introducing spurious objects and modifying the image information. The scheme is based on equalizing the histogram in all the connected components of the image, which are defined based on the image grey-values and spatial relations between its pixels. Following mathematical morphology, these constitute the basic objects in the scene. We give examples for both grey-valued and color images.
Vicent Caselles, Jose Luis Lisani, Jean-Michel Morel, Guillermo Sapiro
ICIP (1)4
1997 Anisotropic Smoothing of Posterior Probabilities
abstract
Teo et al. (see IEEE Trans. on Medical Imaging, 1997) proposed an efficient image segmentation technique that anisotropically smoothes the homogeneous posterior probabilities before independent pixel wise MAP classification is carried out. In this paper we develop the mathematical theory underlying the technique. We demonstrate that prior anisotropic smoothing of the posterior probabilities yields the MAP solution of a discrete MRF with a non-interacting, analog discontinuity field. In contrast, isotropic smoothing of the posterior probabilities is equivalent to computing the MAP solution of a single, discrete MRF using continuous relaxation labeling. Combining a discontinuity field with a discrete MRF is important as it allows the disabling of clique potentials across discontinuities. Furthermore, explicit representation of the discontinuity field suggests new algorithms that incorporate properties like hysteresis and non-maximal suppression.
P. C. Teo, Guillermo Sapiro, Brian A. Wandell
ICIP (1)2
1997 Contrast Enhancement via Image Evolution Flow
Guillermo Sapiro, Vicent Caselles
CVGIP Graph. Model. Image Process.1
1997 Color Snakes
Guillermo Sapiro
Comput. Vis. Image Underst.1
1997 Geodesic Active Contours
Vicent Caselles, Ron Kimmel, Guillermo Sapiro
Int. J. Comput. Vis.3
1997 Minimal Surfaces Based Object Segmentation
abstract
A geometric approach for 3D object segmentation and representation is presented. The segmentation is obtained by deformable surfaces moving towards the objects to be detected in the 3D image. The model is based on curvature motion and the computation of surfaces with minimal areas, better known as minimal surfaces. The space where the surfaces are computed is induced from the 3D image (volumetric data) in which the objects are to be detected. The model links between classical deformable surfaces obtained via energy minimization, and intrinsic ones derived from curvature based flows. The new approach is stable, robust, and automatically handles changes in the surface topology during the deformation.
Vicent Caselles, Ron Kimmel, Guillermo Sapiro, Catalina Sbert
IEEE Trans. Pattern Anal. Mach. Intell.3
1997 Creating Connected Representations of Cortical Gray Matter for Functional MRI Visualization
abstract
We describe a system that is being used to segment gray matter from magnetic resonance imaging (MRI) and to create connected cortical representations for functional MRI visualization (fMRI). The method exploits knowledge of the anatomy of the cortex and incorporates structural constraints into the segmentation. First, the white matter and cerebral spinal fluid (CSF) regions in the MR volume are segmented using a novel techniques of posterior anisotropic diffusion. Then, the user selects the cortical white matter component of interest, and its structure is verified by checking for cavities and handles. After this, a connected representation of the gray matter is created by a constrained growing-out from the white matter boundary. Because the connectivity is computed, the segmentation can be used as input to several methods of visualizing the spatial pattern of cortical activity within gray matter. In our case, the connected representation of gray matter is used to create a flattened representation of the cortex. Then, fMRI measurements are overlaid on the flattened representation, yielding a representation of the volumetric data within a single image. The software is freely available to the research community.
P. C. Teo, Guillermo Sapiro, Brian A. Wandell
IEEE Trans. Medical Imaging2
1996 Affine Invariant Detection: Edges, Active Contours, and Segments
abstract
In this paper we undertake a systematic investigation of affine invariant object detection. Edge detection is first presented from the point of view of the affine invariant scale-space obtained by curvature based motion of the image level-sets. In this case, affine invariant edges are obtained as a weighted difference of images at different scales. We then introduce the affine gradient as the simplest possible affine invariant differential function which has the same qualitative behavior as the Euclidean gradient magnitude. These edge detectors are the basis both to extend the affine invariant scale-space to a complete affine flow for image denoising and simplification, and to define affine invariant active contours for object detection and edge integration. The active contours are obtained as a gradient flow in a conformally Euclidean space defined by the image on which the object is to be detected. That is, we show that objects can be segmented in an affine invariant manner by computing a path of minimal weighted affine distance, the weight being given by functions of the affine edge detectors. The geodesic path is computed via an algorithm which allows to simultaneously detect any number of objects independently of the initial curve topology.
Peter J. Olver, Guillermo Sapiro, Allen R. Tannenbaum
CVPR2
1996 Vector-Valued Active Contours
abstract
A framework for object segmentation in vector-valued images is presented in this paper. The first scheme proposed is based on geometric active contours moving towards the objects to be detected in the vector-valued image. Objects boundaries are obtained as geodesics or minimal weighted distance curves in a Riemannian space. The metric in this space is given by a definition of edges in vector-valued images. The curve flow corresponding to the proposed active contours holds formal existence, uniqueness, stability, and correctness results. The technique is applicable for example to color and texture images. The scheme automatically handles changes in the deforming curve topology. We conclude the paper presenting an extension of the color active contours which leads to a possible image flow for vector-valued image segmentation. The algorithm is based on moving each one of the image level-sets according to the proposed color active contours. This extension also shows the relation of the color geodesic active contours with a number of partial-differential-equations based image processing algorithms as anisotropic diffusion and shock filters.
Guillermo Sapiro
CVPR1
1996 Loco-I: A Low Complexity, Context-Based, Lossless Image Compression Algorithm
abstract
LOCO-I (low complexity lossless compression for images) is a novel lossless compression algorithm for continuous-tone images which combines the simplicity of Huffman coding with the compression potential of context models, thus "enjoying the best of both worlds." The algorithm is based on a simple fixed context model, which approaches the capability of the more complex universal context modeling techniques for capturing high-order dependencies. The model is tuned for efficient performance in conjunction with a collection of (context-conditioned) Huffman codes, which is realized with an adaptive, symbol-wise, Golomb-Rice code. LOCO-I attains, in one pass, and without recourse to the higher complexity arithmetic coders, compression ratios similar or superior to those obtained with state-of-the-art schemes based on arithmetic coding. In fact, LOCO-I is being considered by the ISO committee as a replacement for the current lossless standard in low-complexity applications.
Marcelo J. Weinberger, Gadiel Seroussi, Guillermo Sapiro
Data Compression Conference3
1996 Three Dimensional Object Modeling via Minimal Surfaces
Vicent Caselles, Ron Kimmel, Guillermo Sapiro, Catalina Sbert
ECCV (1)3
1996 From active contours to anisotropic diffusion: connections between basic PDE's in image processing
abstract
We present mathematical and qualitative relations between a number of partial differential equations frequently used in image processing and computer vision. We show for example that classical active contours introduced for object detection by Terzopoulos (1988) and colleagues are connected to anisotropic diffusion flows as those defined by Perona and Malik (1990). We also deal with the relation of these flows with shock filters and variational approaches for image restoration.
Guillermo Sapiro
ICIP (1)1
1996 Vector (self) snakes: a geometric framework for color, texture, and multiscale image segmentation
abstract
A partial differential equations (PDEs) based geometric framework for segmentation of vector-valued images is described. The first component of this approach is based on two dimensional geometric active contours deforming from their initial position towards objects in the image. The boundaries of these objects are then obtained as geodesics or minimal weighted distance curves in a Riemannian space. The metric in this space is given by a definition of edges in vector-valued images, incorporating information from all the image components. The curve flow corresponding to these active contours holds formal existence, uniqueness, stability, and correctness results. Then, embedding the deforming curve as the level-set of the image, that is, deforming each one of the image components level-sets according to these active contours, a system of coupled PDEs is obtained. This system deforms the image towards uniform regions, obtaining a simplified (or segmented) image. The flow is related to a number of PDEs based image processing algorithms as anisotropic diffusion and shock filters. The technique is applicable to color and texture images, as well as to vector data obtained from general image decompositions.
Guillermo Sapiro
ICIP (1)1
1996 Anisotropic diffusion of multivalued images with applications to color filtering
abstract
A general framework for anisotropic diffusion of multivalued images is presented. We propose an evolution equation where, at each point in time, the directions and magnitudes of the maximal and minimal rate of change in the vector-image are first evaluated. These are given by eigenvectors and eigenvalues of the first fundamental form in the given image metric. Then, the image diffuses via a system of coupled differential equations in the direction of minimal change. The diffusion "strength" is controlled by a function that measures the degree of dissimilarity between the eigenvalues. We apply the proposed framework to the filtering of color images represented in CIE-L*a*b* space.
Guillermo Sapiro, Dario L. Ringach
IEEE Trans. Image Process.1
1995 Geodesic Active Contours
abstract
A novel scheme for the detection of object boundaries is presented. The technique is based on active contours deforming according to intrinsic geometric measures of the image. The evolving contours naturally split and merge, allowing the simultaneous detection of several objects and both interior and exterior boundaries. The proposed approach is based on the relation between active contours and the computation of geodesics or minimal distance curves. The minimal distance curve lays in a Riemannian space whose metric as defined by the image content. This geodesic approach for object segmentation allows to connect classical "snakes" based on energy minimization and geometric active contours based on the theory of curve evolution. Previous models of geometric active contours are improved as showed by a number of examples. Formal results concerning existence, uniqueness, stability, and correctness of the evolution are presented as well.>
Vicent Caselles, Ron Kimmel, Guillermo Sapiro
ICCV3
1995 Geometric partial differential equations in image analysis: past, present, and future
abstract
The author discusses the main characteristics of the use of partial differential equations and curve/surface evolution theory in computer vision and image processing. The approach and its main advantages are described, together with a number of examples.
Guillermo Sapiro
ICIP (3)1
1995 Histogram modification via partial differential equations
abstract
An algorithm for histogram modification via image evolution equations is presented. We show that the image histogram can be modified to achieve any given distribution as the steady state solution of this partial differential equation. We then prove that this equation corresponds to a gradient descent flow of a variational problem. That is, the proposed PDE is solving an energy minimization problem. This gives a new interpretation to histogram modification and contrast enhancement in general. This interpretation is completely formulated in the image domain, in contrast with classical techniques for histogram modification which are formulated in a probabilistic domain. From this, new algorithms for contrast enhancement, which include for example image modeling, can be derived. Based on the energy formulation and its corresponding PDE, we show that the proposed histogram modification algorithm can be combined with denoising schemes. This allows one to perform simultaneous contrast enhancement and denoising, avoiding common noise sharpening effects in classical algorithms. The approach is extended to focal contrast enhancement as well. Theoretical results regarding the existence of solutions to the proposed equations are presented.
Guillermo Sapiro, Vicent Caselles
ICIP (3)1
1995 Evolutions of Planar Polygons
abstract
Evolutions of closed planar polygons are studied in this work. In the first part of the paper, the general theory of linear polygon evolutions is presented, and two specific problems are analyzed. The first one is a polygonal analog of a novel affine-invariant differential curve evolution, for which the convergence of planar curves to ellipses was proved. In the polygon case, convergence to polygonal approximation of ellipses, polygo nal ellipses, is proven. The second one is related to cyclic pursuit problems, and convergence, either to polygonal ellipses or to polygonal circles, is proven. In the second part, two possible polygonal analogues of the well-known Euclidean curve shortening flow are presented. The models follow from geometric considerations. Experimental results show that an arbitrary initial polygon converges to either regular or irregular polygonal approximations of circles when evolving according to the proposed Euclidean flows.
Alfred M. Bruckstein, Guillermo Sapiro, Doron Shaked
Int. J. Pattern Recognit. Artif. Intell.2
1995 Area and Length Preserving Geometric Invariant Scale-Spaces
abstract
In this paper, area preserving multi-scale representations of planar curves are described. This allows smoothing without shrinkage at the same time preserving all the scale-space properties. The representations are obtained deforming the curve via geometric heat flows while simultaneously magnifying the plane by a homethety which keeps the enclosed area constant. When the Euclidean geometric heat flow is used, the resulting representation is Euclidean invariant, and similarly it is affine invariant when the affine one is used. The flows are geometrically intrinsic to the curve, and exactly satisfy all the basic requirements of scale-space representations. In the case of the Euclidean heat flow, it is completely local as well. The same approach is used to define length preserving geometric flows. A similarity (scale) invariant geometric heat flow is studied as well in this work.>
Guillermo Sapiro, Allen R. Tannenbaum
IEEE Trans. Pattern Anal. Mach. Intell.1
1994 Area and Lenght Preserving Geometric Invariant Scale-Spaces
Guillermo Sapiro, Allen R. Tannenbaum
ECCV (2)1
1994 Experiments on Geometric Image Enhancement
abstract
In this paper we experiments with geometric algorithms for image smoothing. Examples are given for MRI and ATR data. We emphasize experiments with the affine invariant geometric smoother or affine heat equation, originally developed for binary shape smoothing, and found to be efficient for gray-level images as well. Efficient numerical implementations of these flows give anisotropic diffusion processes which preserve edges.>
Guillermo Sapiro, Allen R. Tannenbaum, Yu-Li You, Mostafa Kaveh
ICIP (2)1
1994 Morphological Image Coding Based on a Geometric Sampling Theorem and a Modified Skeleton Representation
Guillermo Sapiro, David Malah
J. Vis. Commun. Image Represent.1
1993 Affine invariant scale-space
Guillermo Sapiro, Allen R. Tannenbaum
Int. J. Comput. Vis.1
1993 Implementing continuous-scale morphology via curve evolution
Guillermo Sapiro, Ron Kimmel, Doron Shaked, Benjamin B. Kimia, Alfred M. Bruckstein
Pattern Recognit.1