EDBT 2026 Demo / reviewers in the wild / expert
Gérard G. Medioni
dblp:02/3016 · also Gérard Guy Medioni
· DBLP profile ↗
243ranked-venue papers
12as first author
11since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 185 · 8 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 146 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 1 first-authorSystems, architecture and hardware · 12 · 2 first-authorHuman-computer interaction and ubiquitous computing · 12 · 2 first-authorDatabases, data management, data science and information retrieval · 3Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Group-Aware Reinforcement Learning for Output Diversity in Large Language ModelsabstractOron Anschel, Alon Shoshan, Adam Botach, Shunit Haviv Hakimi, Asaf Gendler, Emanuel Ben Baruch, Nadav Bhonker, Igor Kviatkovsky, Manoj Aggarwal, Gerard Medioni. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Oron Anschel, Alon Shoshan, Adam Botach, Shunit Haviv Hakimi, Asaf Gendler, Emanuel Ben Baruch, Nadav Bhonker, Igor Kviatkovsky, Manoj Aggarwal, Gérard G. Medioni |
EMNLP | 10 |
| 2025 | LV-MAE: Learning Long Video Representations Through Masked-Embedding AutoencodersabstractIn this work, we introduce long-video masked-embedding autoencoders (LV-MAE), a self-supervised learning framework for long video representation. Our approach treats short- and long-span dependencies as two separate tasks. Such decoupling allows for a more intuitive video processing where short-span spatiotemporal primitives are first encoded and are then used to capture long-range dependencies across consecutive video segments. To achieve this, we leverage advanced off-the-shelf multimodal encoders to extract representations from short segments within the long video, followed by pre-training a masked-embedding autoencoder capturing high-level interactions across segments. LV-MAE is highly efficient to train and enables the processing of much longer videos by alleviating the constraint on the number of input frames. Furthermore, unlike existing methods that typically pre-train on short-video datasets, our approach offers self-supervised pre-training using long video samples (e.g., 20+ minutes video clips) at scale. Using LV-MAE representations, we achieve state-of-the-art results on three long-video benchmarks -- LVU, COIN, and Breakfast -- employing only a simple classification head for either attentive or linear probing. Finally, to assess LV-MAE pre-training and visualize its reconstruction quality, we leverage the video-language aligned space of short video representations to monitor LV-MAE through video-text retrieval. Code is available at https://github.com/amazon-science/lv-mae. Ilan Naiman, Emanuel Ben Baruch, Oron Anschel, Alon Shoshan, Igor Kviatkovsky, Manoj Aggarwal, Gérard G. Medioni |
ICCV | 7 |
| 2025 | Distilling the Knowledge in Data PruningabstractWith the increasing size of datasets used for training neural networks, data pruning has gained traction in recent years. However, most current data pruning algorithms are limited in their ability to preserve accuracy compared to models trained on the full data, especially in high pruning regimes. In this paper we explore the application of data pruning while incorporating knowledge distillation (KD) when training on a pruned subset. That is, rather than relying solely on ground-truth labels, we also use the soft predictions from a teacher network pre-trained on the complete data. By integrating KD into training, we demonstrate significant improvement across datasets, pruning methods, and on all pruning fractions. We first establish a theoretical motivation for employing self-distillation to improve training on pruned data. Then, we empirically make a compelling and highly practical observation: using KD, simple random pruning is comparable or superior to sophisticated pruning methods across all pruning regimes. On ImageNet for example, we achieve superior accuracy despite training on a random subset of only 50% of the data. Additionally, we demonstrate a crucial connection between the pruning factor and the optimal knowledge distillation weight. This helps mitigate the impact of samples with noisy labels and low-quality images retained by typical pruning algorithms. Finally, we make an intriguing observation: when using lower pruning fractions, larger teachers lead to accuracy degradation, while surprisingly, employing teachers with a smaller capacity than the student’s may improve results. Our code will be made available. Emanuel Ben Baruch, Adam Botach, Igor Kviatkovsky, Manoj Aggarwal, Gérard G. Medioni |
ICML | 5 |
| 2024 | Data Pruning via Separability, Integrity, and Model Uncertainty-Aware Importance Sampling
Steven A. Grosz, Manoj Aggarwal, Gérard G. Medioni, Anil K. Jain 0001 |
ICPR (2) | 6 |
| 2024 | FPGAN-Control: A Controllable Fingerprint Generator for Training with Synthetic DataabstractTraining fingerprint recognition models using synthetic data has recently gained increased attention in the biometric community as it alleviates the dependency on sensitive personal data. Existing approaches for fingerprint generation are limited in their ability to generate diverse impressions of the same finger, a key property for providing effective data for training recognition models. To address this gap, we present FPGAN-Control, an identity preserving image generation framework which enables control over the fingerprint’s image appearance (e.g., fingerprint type, acquisition device, pressure level) of generated fingerprints. We introduce a novel appearance loss that encourages disentanglement between the fingerprint’s identity and appearance properties. In our experiments, we used the publicly available NIST SD302 (N2N) dataset for training the FPGAN-Control model. We demonstrate the merits of FPGAN-Control, both quantitatively and qualitatively, in terms of identity preservation level, degree of appearance control, and low synthetic-to-real domain gap. Finally, training recognition models using only synthetic datasets generated by FPGAN-Control lead to recognition accuracies that are on par or even surpass models trained using real data. To the best of our knowledge, this is the first work to demonstrate this. Alon Shoshan, Nadav Bhonker, Emanuel Ben Baruch, Ori Nizan, Igor Kviatkovsky, Joshua J. Engelsma, Manoj Aggarwal, Gérard G. Medioni |
WACV | 8 |
| 2024 | Asymmetric Image Retrieval with Cross Model Compatible EnsemblesabstractThe asymmetrical retrieval setting is a well suited solution for resource constrained applications such as face recognition and image retrieval. In this setting, a large model is used for indexing the gallery while a lightweight model is used for querying. The key principle in such systems is ensuring that both models share the same embedding space. Most methods in this domain are based on knowledge distillation. While useful, they suffer from several drawbacks: they are upper-bounded by the performance of the single best model found and cannot be extended to use an ensemble of models in a straightforward manner. In this paper we present an approach that does not rely on knowledge distillation, rather it utilizes embedding transformation models. This allows the use of N independently trained and diverse gallery models (e.g., trained on different datasets or having a different architecture) and a single query model. As a result, we improve the overall accuracy beyond that of any single model while maintaining a low computational budget for querying. Additionally, we propose a gallery image rejection method that utilizes the diversity between multiple transformed embeddings to estimate the uncertainty of gallery images. Alon Shoshan, Ori Linial, Nadav Bhonker, Elad Hirsch, Lior Zamir, Igor Kviatkovsky, Gérard G. Medioni |
WACV | 7 |
| 2023 | Synthetic data for model selectionabstractRecent breakthroughs in synthetic data generation approaches made it possible to produce highly photorealistic images which are hardly distinguishable from real ones. Furthermore, synthetic generation pipelines have the potential to generate an unlimited number of images. The combination of high photorealism and scale turn synthetic data into a promising candidate for improving various machine learning (ML) pipelines. Thus far, a large body of research in this field has focused on using synthetic images for training, by augmenting and enlarging training data. In contrast to using synthetic data for training, in this work we explore whether synthetic data can be beneficial for model selection. Considering the task of image classification, we demonstrate that when data is scarce, synthetic data can be used to replace the held out validation set, thus allowing to train on a larger dataset. We also introduce a novel method to calibrate the synthetic error estimation to fit that of the real domain. We show that such calibration significantly improves the usefulness of synthetic data for model selection. Alon Shoshan, Nadav Bhonker, Igor Kviatkovsky, Matan Fintz, Gérard G. Medioni |
ICML | 5 |
| 2023 | OutfitTransformer: Learning Outfit Representations for Fashion RecommendationabstractLearning an effective outfit-level representation is critical for predicting the compatibility of items in an outfit, and retrieving complementary items for a partial outfit. We present a framework, OutfitTransformer, that uses the pro-posed task-specific tokens and leverages the self-attention mechanism to learn effective outfit-level representations en-coding the compatibility relations between all items in the entire outfit for addressing both compatibility prediction and complementary item retrieval. For compatibility pre-diction, we design an outfit token to capture a global out-fit representation and train the framework using a classification loss. For complementary item retrieval, we design a target item token that additionally takes the target item specification (in the form of a category or text description) into consideration. We train our framework using a pro-posed set-wise outfit ranking loss to generate a target item embedding given an outfit, and a target item specification as inputs. The generated target item embedding is then used to retrieve compatible items that match the rest of the out-fit. Additionally, we adopt a pre-training approach and a curriculum learning strategy to improve retrieval performance. Experiments show that our approach outperforms state-of-the-art methods on compatibility prediction, fill-in-the-blank, and complementary item retrieval tasks. Rohan Sarkar, Navaneeth Bodla, Mariya I. Vasileva, Yen-Liang Lin, Anurag Beniwal, Alan Lu, Gérard G. Medioni |
WACV | 7 |
| 2022 | Efficient Video Instance Segmentation via Tracklet Query and ProposalabstractVideo Instance Segmentation (VIS) aims to simultaneously classify, segment, and track multiple object instances in videos. Recent clip-level VIS takes a short video clip as input each time showing stronger performance than frame-level VIS (tracking-by-segmentation), as more temporal context from multiple frames is utilized. Yet, most clip-level methods are neither end-to-end learnable nor real-time. These limitations are addressed by the recent VIS transformer (VisTR) [25] which performs VIS end-to-end within a clip. However, VisTR suffers from long training time due to its frame-wise dense attention. In addition, VisTR is not fully end-to-end learnable in multiple video clips as it requires a hand-crafted data association to link instance tracklets between successive clips. This paper proposes EfficientVIS, a fully end-to-end framework with efficient training and inference. At the core are tracklet query and tracklet proposal that associate and segment regions-of-interest (RoIs) across space and time by an iterative query-video interaction. We further propose a correspondence learning that makes tracklets linking between clips end-to-end learnable. Compared to VisTR, EfficientVIS requires$15\times$fewer training epochs while achieving state-of-the-art accuracy on the YouTube-VIS benchmark. Meanwhile, our method enables whole video instance segmentation in a single end-to-end pass without data association at all. Jialian Wu, Sudhir Yarram, Hui Liang 0003, Junsong Yuan 0001, Jayan Eledath, Gérard G. Medioni |
CVPR | 7 |
| 2021 | Energy-Based Learning for Scene Graph GenerationabstractTraditional scene graph generation methods are trained using cross-entropy losses that treat objects and relationships as independent entities. Such a formulation, however, ignores the structure in the output space, in an inherently structured prediction problem. In this work, we introduce a novel energy-based learning framework for generating scene graphs. The proposed formulation allows for efficiently incorporating the structure of scene graphs in the output space. This additional constraint in the learning framework acts as an inductive bias and allows models to learn efficiently from a small number of labels. We use the proposed energy-based framework†to train existing stateof-the-art models and obtain a significant performance improvement, of up to 21% and 27%, on the Visual Genome [9] and GQA [5] benchmark datasets, respectively. Furthermore, we showcase the learning efficiency of the proposed framework by demonstrating superior performance in the zero- and few-shot settings where data is scarce. Mohammed Suhail, Abhay Mittal, Behjat Siddiquie, Chris Broaddus, Jayan Eledath, Gérard G. Medioni, Leonid Sigal |
CVPR | 6 |
| 2021 | GAN-Control: Explicitly Controllable GANsabstractWe present a framework for training GANs with explicit control over generated facial images. We are able to control the generated image by settings exact attributes such as age, pose, expression, etc. Most approaches for manipulating GAN-generated images achieve partial control by leveraging the latent space disentanglement properties, obtained implicitly after standard GAN training. Such methods are able to change the relative intensity of certain attributes, but not explicitly set their values. Recently proposed methods, designed for explicit control over human faces, harness morphable 3D face models (3DMM) to allow fine-grained control capabilities in GANs. Unlike these methods, our control is not constrained to 3DMM parameters and is extendable beyond the domain of human faces. Using contrastive learning, we obtain GANs with an explicitly disentangled latent space. This disentanglement is utilized to train control-encoders mapping human-interpretable inputs to suitable latent vectors, thus allowing explicit control. In the domain of human faces we demonstrate control over identity, age, pose, expression, hair color and illumination. We also demonstrate control capabilities of our framework in the domains of painted portraits and dog image generation. We demonstrate that our approach achieves state-of-the-art performance both qualitatively and quantitatively. Alon Shoshan, Nadav Bhonker, Igor Kviatkovsky, Gérard G. Medioni |
ICCV | 4 |
| 2020 | AOWS: Adaptive and Optimal Network Width Search With Latency ConstraintsabstractNeural architecture search (NAS) approaches aim at automatically finding novel CNN architectures that fit computational constraints while maintaining a good performance on the target platform. We introduce a novel efficient one-shot NAS approach to optimally search for channel numbers, given latency constraints on a specific hardware. We first show that we can use a black-box approach to estimate a realistic latency model for a specific inference platform, without the need for low-level access to the inference computation. Then, we design a pairwise MRF to score any channel configuration and use dynamic programming to efficiently decode the best performing configuration, yielding an optimal solution for the network width search. Finally, we propose an adaptive channel configuration sampling scheme to gradually specialize the training phase to the target computational constraints. Experiments on ImageNet classification show that our approach can find networks fitting the resource constraints on different target platforms while improving accuracy over the state-of-the-art efficient networks. Maxim Berman, Leonid Pishchulin, Matthew B. Blaschko, Gérard G. Medioni |
CVPR | 5 |
| 2020 | Learning and recognition for assistive computer vision
Giovanni Maria Farinella, Marco Leo, Gérard G. Medioni, Mohan M. Trivedi |
Pattern Recognit. Lett. | 3 |
| 2019 | Deep, Landmark-Free FAME: Face Alignment, Modeling, and Expression Estimation
Feng-Ju Chang, Anh Tuan Tran 0001, Tal Hassner, Iacopo Masi, Ramakant Nevatia, Gérard G. Medioni |
Int. J. Comput. Vis. | 6 |
| 2019 | Face-Specific Data Augmentation for Unconstrained Face Recognition
Iacopo Masi, Anh Tuan Tran 0001, Tal Hassner, Gozde Sahin, Gérard G. Medioni |
Int. J. Comput. Vis. | 5 |
| 2019 | Learning Pose-Aware Models for Pose-Invariant Face Recognition in the WildabstractWe propose a method designed to push the frontiers of unconstrained face recognition in the wild with an emphasis on extreme out-of-plane pose variations. Existing methods either expect a single model to learn pose invariance by training on massive amounts of data or else normalize images by aligning faces to a single frontal pose. Contrary to these, our method is designed to explicitly tackle pose variations. Our proposed Pose-Aware Models (PAM) process a face image using several pose-specific, deep convolutional neural networks (CNN). 3D rendering is used to synthesize multiple face poses from input images to both train these models and to provide additional robustness to pose variations at test time. Our paper presents an extensive analysis of the IARPA Janus Benchmark A (IJB-A), evaluating the effects that landmark detection accuracy, CNN layer selection, and pose model selection all have on the performance of the recognition pipeline. It further provides comparative evaluations on IJB-A and the PIPA dataset. These tests show that our approach outperforms existing methods, even surprisingly matching the accuracy of methods that were specifically fine-tuned to the target dataset. Parts of this work previously appeared in [1] and [2]. Iacopo Masi, Feng-Ju Chang, Jongmoo Choi, Shai Harel, Jungyeon Kim, KangGeon Kim, Jatuporn Toy Leksut, Stephen Rawls, Yue Wu 0001, Tal Hassner, Wael Abd-Almageed, Gérard G. Medioni, Louis-Philippe Morency, Premkumar Natarajan, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 12 |
| 2019 | Benchmarking parts based face processing in-the-wild for gender recognition and head pose estimation
Flávio H. de Bittencourt Zavan, Olga R. P. Bellon, Luciano Silva, Gérard G. Medioni |
Pattern Recognit. Lett. | 4 |
| 2018 | Extreme 3D Face Reconstruction: Seeing Through OcclusionsabstractExisting single view, 3D face reconstruction methods can produce beautifully detailed 3D results, but typically only for near frontal, unobstructed viewpoints. We describe a system designed to provide detailed 3D reconstructions of faces viewed under extreme conditions, out of plane rotations, and occlusions. Motivated by the concept of bump mapping, we propose a layered approach which decouples estimation of a global shape from its mid-level details (e.g., wrinkles). We estimate a coarse 3D face shape which acts as a foundation and then separately layer this foundation with details represented by a bump map. We show how a deep convolutional encoder-decoder can be used to estimate such bump maps. We further show how this approach naturally extends to generate plausible details for occluded facial regions. We test our approach and its components extensively, quantitatively demonstrating the invariance of our estimated facial details. We further provide numerous qualitative examples showing that our method produces detailed 3D face shapes in viewing conditions where existing state of the art often break down. Anh Tuan Tran 0001, Tal Hassner, Iacopo Masi, Eran Paz, Yuval Nirkin, Gérard G. Medioni |
CVPR | 6 |
| 2018 | ExpNet: Landmark-Free, Deep, 3D Facial ExpressionsabstractWe describe a deep learning based method for estimating 3D facial expression coefficients. Unlike previous work, our process does not relay on facial landmark detection methods as a proxy step. Recent methods have shown that a CNN can be trained to regress accurate and discriminative 3D morphable model (3DMM) representations, directly from image intensities. By foregoing landmark detection, these methods were able to estimate shapes for occluded faces appearing in unprecedented viewing conditions. We build on those methods by showing that facial expressions can also be estimated by a robust, deep, landmark-free approach. Our ExpNet CNN is applied directly to the intensities of a face image and regresses a 29D vector of 3D expression coefficients. We propose a unique method for collecting data to train our network, leveraging on the robustness of deep networks to training label noise. We further offer a novel means of evaluating the accuracy of estimated expression coefficients: by measuring how well they capture facial emotions on the CK+ and EmotiW-17 emotion recognition benchmarks. We show that our ExpNet produces expression coefficients which better discriminate between facial emotions than those obtained using state of the art, facial landmark detectors. Moreover, this advantage grows as image scales drop, demonstrating that our ExpNet is more robust to scale changes than landmark detectors. Finally, our ExpNet is orders of magnitude faster than its alternatives. Feng-Ju Chang, Anh Tuan Tran 0001, Tal Hassner, Iacopo Masi, Ramakant Nevatia, Gérard G. Medioni |
FG | 6 |
| 2018 | On Face Segmentation, Face Swapping, and Face PerceptionabstractWe show that even when face images are unconstrained and arbitrarily paired, face swapping between them is quite simple. To this end, we make the following contributions. (a) Instead of tailoring systems for face segmentation, as others previously proposed, we show that a standard fully convolutional network (FCN) can achieve remarkably fast and accurate segmentations, provided that it is trained on a rich enough example set. For this purpose, we describe novel data collection and generation routines which provide challenging segmented face examples. (b) We use our segmentations for robust face swapping under unprecedented conditions. (c) Unlike previous work, our swapping is robust enough to allow for extensive quantitative tests. To this end, we use the Labeled Faces in the Wild (LFW) benchmark and measure the effect of intra- and inter-subject face swapping on recognition. We show that our intra-subject swapped faces remain as recognizable as their sources, testifying to the effectiveness of our method. In line with established perceptual studies, we show that better face swapping produces less recognizable inter-subject results. This is the first time this effect was quantitatively demonstrated by machine vision systems. Yuval Nirkin, Iacopo Masi, Anh Tuan Tran 0001, Tal Hassner, Gérard G. Medioni |
FG | 5 |
| 2018 | Robust Denoising of Piece-Wise Smooth ManifoldsabstractA common smoothness model used in graph based regularization approaches is to require the energy of signals to be small with respect to the graph Laplacian of the graph. In this paper, we suggest an alternative approach which can effectively incorporate the high frequency information of the graph for unsupervised piece-wise smooth manifold denoising using Spectral Graph Wavelets. Our approach is based on a novel technique to remove noise from SGW coefficients estimated from a local tangent space based graph, which allows us to effectively regularize manifolds with singularities, such as for example intersecting manifolds. Experimental results on synthetic and real datasets in computer vision applications show that our proposed approach outperforms the state of the art, and is an effective tool to remove noise from manifolds with complex structures without over-smoothing at discontinuities. Shay Deutsch, Antonio Ortega, Gérard G. Medioni |
ICASSP | 3 |
| 2018 | Face and Body Association for Video-Based Face RecognitionabstractIn recent years face recognition has made extraordinary leaps, yet unconstrained video-based face identification in the wild remains an open and interesting problem. Videos, unlike still-images, offer a myriad of data for face modeling, sampling, and recognition, but, on the other hand, contain low-quality frames and motion blur. A key component in video-based face recognition is the way in which faces are associated through the video sequence before being used for recognition. In this paper, we present a video-based face recognition method taking advantage of face and body association (FBA). To track and associate subjects that appear across frames in multiple shots, we solve a data association problem using both face and body appearance. The final recovered track is then used to build a face representation for recognition. We evaluate our FBA method for video-based face recognition on a challenging dataset. Our experiments show up to 5% improvement in the identification rate over the state-of-the-art. KangGeon Kim, Zhenheng Yang, Iacopo Masi, Ramakant Nevatia, Gérard G. Medioni |
WACV | 5 |
| 2018 | Facial Landmark Detection with Tweaked Convolutional Neural NetworksabstractThis paper concerns the problem of facial landmark detection. We provide a unique new analysis of the features produced at intermediate layers of a convolutional neural network (CNN) trained to regress facial landmark coordinates. This analysis shows that while being processed by the CNN, face images can be partitioned in an unsupervised manner into subsets containing faces in similar poses (i.e., 3D views) and facial properties (e.g., presence or absence of eye-wear). Based on this finding, we describe a novel CNN architecture, specialized to regress the facial landmark coordinates of faces in specific poses and appearances. To address the shortage of training data, particularly in extreme profile poses, we additionally present data augmentation techniques designed to provide sufficient training examples for each of these specialized sub-networks. The proposed Tweaked CNN (TCNN) architecture is shown to outperform existing landmark detection methods in an extensive battery of tests on the AFW, ALFW, and 300W benchmarks. Finally, to promote reproducibility of our results, we make code and trained models publicly available through our project webpage. Yue Wu 0001, Tal Hassner, KangGeon Kim, Gérard G. Medioni, Premkumar Natarajan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | Regressing Robust and Discriminative 3D Morphable Models with a Very Deep Neural NetworkabstractThe 3D shapes of faces are well known to be discriminative. Yet despite this, they are rarely used for face recognition and always under controlled viewing conditions. We claim that this is a symptom of a serious but often overlooked problem with existing methods for single view 3D face reconstruction: when applied in the wild, their 3D estimates are either unstable and change for different photos of the same subject or they are over-regularized and generic. In response, we describe a robust method for regressing discriminative 3D morphable face models (3DMM). We use a convolutional neural network (CNN) to regress 3DMM shape and texture parameters directly from an input photo. We overcome the shortage of training data required for this purpose by offering a method for generating huge numbers of labeled examples. The 3D estimates produced by our CNN surpass state of the art accuracy on the MICC data set. Coupled with a 3D-3D face matching pipeline, we show the first competitive face recognition results on the LFW, YTF and IJB-A benchmarks using 3D face shapes as representations, rather than the opaque deep feature vectors used by other modern systems. Anh Tuan Tran 0001, Tal Hassner, Iacopo Masi, Gérard G. Medioni |
CVPR | 4 |
| 2017 | Local-Global Landmark Confidences for Face RecognitionabstractA key to successful face recognition is accurate and reliable face alignment using automatically-detected facial landmarks. Given this strong dependency between face recognition and facial landmark detection, robust face recognition requires knowledge of when the facial landmark detection algorithm succeeds and when it fails. Facial landmark confidence represents this measure of success. In this paper, we propose two methods to measure landmark detection confidence: local confidence based on local predictors of each facial landmark, and global confidence based on a 3D rendered face model. A score fusion approach is also introduced to integrate these two confidences effectively. We evaluate both confidence metrics on two datasets for face recognition: JANUS CS2 and IJB-A datasets. Our experiments show up to 9% improvements when face recognition algorithm integrates the local-global confidence metrics. KangGeon Kim, Feng-Ju Chang, Jongmoo Choi, Louis-Philippe Morency, Ramakant Nevatia, Gérard G. Medioni |
FG | 6 |
| 2017 | Rapid Synthesis of Massive Face Sets for Improved Face RecognitionabstractRecent work demonstrated that computer graphics techniques can be used to improve face recognition performances by synthesizing multiple new views of faces available in existing face collections. By so doing, more images and more appearance variations are available for training, thereby improving the deep models trained on these images. Similar rendering techniques were also applied at test time to align faces in 3D and reduce appearance variations when comparing faces. These previous results, however, did not consider the computational cost of rendering: At training, rendering millions of face images can be prohibitive; at test time, rendering can quickly become a bottleneck, particularly when multiple images represent a subject. This paper builds on a number of observations which, under certain circumstances, allow rendering new 3D views of faces at a computational cost which is equivalent to simple 2D image warping. We demonstrate this by showing that the run-time of an optimized OpenGL rendering engine is slower than the simple Python implementation we designed for the same purpose. The proposed rendering is used in a face recognition pipeline and tested on the challenging IJB-A and Janus CS2 benchmarks. Our results show that our rendering is not only fast, but improves recognition accuracy. Iacopo Masi, Tal Hassner, Anh Tuan Tran 0001, Gérard G. Medioni |
FG | 4 |
| 2017 | Deep 3D face identificationabstractWe propose a novel 3D face recognition algorithm using a deep convolutional neural network (DCNN) and a 3D face expression augmentation technique. The performance of 2D face recognition algorithms has significantly increased by leveraging the representational power of deep neural networks and the use of large-scale labeled training data. In this paper, we show that transfer learning from a CNN trained on 2D face images can effectively work for 3D face recognition by fine-tuning the CNN with an extremely small number of 3D facial scans. We also propose a 3D face expression augmentation technique which synthesizes a number of different facial expressions from a single 3D face scan. Our proposed method shows excellent recognition results on Bosphorus, BU-3DFE, and 3D-TEC datasets without using hand-crafted features. The 3D face identification using our deep features also scales well for large databases. Donghyun Kim 0006, Matthias Hernandez, Jongmoo Choi, Gérard G. Medioni |
IJCB | 4 |
| 2017 | Exploring Local Context for Multi-target Tracking in Wide Area Aerial SurveillanceabstractTracking many vehicles in wide coverage aerial imagery is crucial for understanding events in a large field of view. Most approaches aim to associate detections from frame differencing into tracks. However, slow or stopped vehicles result in long-term missing detections and further cause tracking discontinuities. Relying merely on appearance clue to recover missing detections is difficult as targets are extremely small and in grayscale. In this paper, we address the limitations of detection association trackers by coupling it with a local context tracker (LCT), which does not rely on motion detections. On one hand, our LCT learns neighboring spatial relation and tracks each target in consecutive frames using graph optimization. It takes the advantage of context constraints to avoid drifting to nearby targets. We generate hypotheses from sparse and dense flow efficiently to keep solutions tractable. On the other hand, we use detection association strategy to extract short tracks in batch processing. We explicitly handle merged detections by generating additional hypotheses from them. Our evaluation on wide area aerial imagery sequences shows significant improvement over state-of-the-art methods. Bor-Jeng Chen, Gérard G. Medioni |
WACV | 2 |
| 2017 | Accurate 3D face reconstruction via prior constrained structure from motion
Matthias Hernandez, Tal Hassner, Jongmoo Choi, Gérard G. Medioni |
Comput. Graph. | 4 |
| 2017 | Computer vision for assistive technologies
Marco Leo, Gérard G. Medioni, Mohan M. Trivedi, Takeo Kanade, Giovanni Maria Farinella |
Comput. Vis. Image Underst. | 2 |
| 2016 | Holistically Constrained Local Model: Going Beyond Frontal Poses for Facial Landmark Detection
KangGeon Kim, Tadas Baltrusaitis, Amir Zadeh 0001, Louis-Philippe Morency, Gérard G. Medioni |
BMVC | 5 |
| 2016 | Pose-Aware Face Recognition in the WildabstractWe propose a method to push the frontiers of unconstrained face recognition in the wild, focusing on the problem of extreme pose variations. As opposed to current techniques which either expect a single model to learn pose invariance through massive amounts of training data, or which normalize images to a single frontal pose, our method explicitly tackles pose variation by using multiple pose specific models and rendered face images. We leverage deep Convolutional Neural Networks (CNNs) to learn discriminative representations we call Pose-Aware Models (PAMs) using 500K images from the CASIA WebFace dataset. We present a comparative evaluation on the new IARPA Janus Benchmark A (IJB-A) and PIPA datasets. On these datasets PAMs achieve remarkably better performance than commercial products and surprisingly also outperform methods that are specifically fine-tuned on the target dataset. Iacopo Masi, Stephen Rawls, Gérard G. Medioni, Premkumar Natarajan |
CVPR | 3 |
| 2016 | Do We Really Need to Collect Millions of Faces for Effective Face Recognition?
Iacopo Masi, Anh Tuan Tran 0001, Tal Hassner, Jatuporn Toy Leksut, Gérard G. Medioni |
ECCV (5) | 5 |
| 2016 | Capturing Dynamic Textured Surfaces of Moving Targets
Ruizhe Wang 0002, Lingyu Wei, Etienne Vouga, Qixing Huang, Duygu Ceylan, Gérard G. Medioni, Hao Li 0015 |
ECCV (7) | 6 |
| 2016 | Manifold denoising based on spectral graph waveletsabstractWe propose a new framework for manifold denoising using the Spectral Graph Wavelet transform, which enables non-iterative denoising directly in the graph frequency domain, an approach inspired by conventional wavelet-based signal denoising methods. We theoretically justify our approach, based on the fact that for smooth manifolds the coordinate information tends to create energy in the low spectral graph wavelet coefficients, while the noise affects all frequency bands in a similar way. Experimental results show that our suggested manifold frequency denoising (MFD) approach significantly outperforms the state of the art manifold denosing methods, and is robust to a wide range of parameter selections, e.g., the choice of k nearest neighbor connectivity of the graph. Shay Deutsch, Antonio Ortega, Gérard G. Medioni |
ICASSP | 3 |
| 2016 | Accurate 3D face modeling and recognition from RGB-D stream in the presence of large pose changesabstractWe propose a 3D face modeling and recognition system using an RGB-D stream in the presence of large pose changes. In the previous work, all facial data points are registered with a reference to improve the accuracy of 3D face model from a low-resolution depth sequence. This registration often fails when applied to non-frontal faces. It causes inaccurate 3D face models and poor performance of matching. We address this problem by pre-aligning each input face (`frontalization') before the registration, which avoids registration failures. For each frame, our method estimates the 3D face pose, assesses the quality of data, segments the facial region, frontalizes it, and performs an accurate registration with the previous 3D model. The 3D-3D recognition system using accurate 3D models from our method outperforms other face recognition systems and shows 100% rank 1 recognition accuracy on a dataset with 30 subjects. Donghyun Kim 0006, Jongmoo Choi, Jatuporn Toy Leksut, Gérard G. Medioni |
ICIP | 4 |
| 2016 | Expression invariant 3D face modeling from an RGB-D videoabstractWe aim to reconstruct an accurate neutral 3D face model from an RGB-D video in the presence of extreme expression changes. Since each depth frame, taken by a low-cost sensor, is noisy, point clouds from multiple frames can be registered and aggregated to build an accurate 3D model. However, direct aggregation of multiple data produces erroneous results in natural interaction (e.g., talking and showing expressions). We propose to analyze facial expression from an RGB frame and neutralize the corresponding 3D point cloud if needed. We first estimate the person's expression by fitting blendshape coefficients using 2D facial landmarks for each frame and calculate an expression deformity (expression score). With the estimated expression score, we determine whether an input face is neutral or non-neutral. If the face is non-neutral, we proceed to neutralize the expression of the 3D point cloud in that frame. To neutralize the 3D point cloud of a face, we deform our generic 3D face model by applying the estimated blendshape coefficients, find displacement vectors from the deformed generic face to a neutral generic face, and apply the displacement vectors to the input 3D point cloud. After preprocessing frames in a video, we rank frames based on the expression scores and register the ranked frames into a single 3D model. Our system produces a neutral 3D face model in the presence of extreme expression changes even when neutral faces do not exist in the video. Donghyun Kim 0006, Jongmoo Choi, Jatuporn Toy Leksut, Gérard G. Medioni |
ICPR | 4 |
| 2016 | Face recognition using deep multi-pose representationsabstractWe introduce our method and system for face recognition using multiple pose-aware deep learning models. In our representation, a face image is processed by several pose-specific deep convolutional neural network (CNN) models to generate multiple pose-specific features. 3D rendering is used to generate multiple face poses from the input image. Sensitivity of the recognition system to pose variations is reduced since we use an ensemble of pose-specific CNN features. The paper presents extensive experimental results on the effect of landmark detection, CNN layer selection and pose model selection on the performance of the recognition pipeline. Our novel representation achieves better results than the state-of-the-art on IARPA's CS2 and NIST's IJB-A in both verification and identification (i.e. search) tasks. Wael Abd-Almageed, Yue Wu 0001, Stephen Rawls, Shai Harel, Tal Hassner, Iacopo Masi, Jongmoo Choi, Jatuporn Toy Leksut, Jungyeon Kim, Premkumar Natarajan, Ramakant Nevatia, Gérard G. Medioni |
WACV | 12 |
| 2016 | Persistent 3D stabilization for aerial imageryabstractIn this paper, we present a novel 3D stabilization method for aerial imagery. While most existing aerial surveillance applications rely on 2D stabilization, which generates parallax errors in urban scenes, we aim to compensate camera motion perfectly as if the scene had been captured by a static camera. It is not only a better way to visualize events but also a more reliable inputs for higher level surveillance tasks. We tackle this problem by proposing a novel approach which generates accurate long-term dense mapping and handles occlusion robustly. Accurate dense mapping from a given frame to the reference frame is achieved in two steps: 1. Identify and compensate outlier flow from independent moving objects. 2. Accumulate the compensated mapping in two paths to avoid drifting. To handle occlusion regions, we formulate occlusion segmentation as an optimization problem considering both mapping relation and smoothness. Seamless rendering is achieved by gradient guidance of a dynamic background model. We evaluate our method on real-world aerial sequences with more than 100 frames with strong parallax. Our method outperforms state-of-the-art 2D and 3D stabilization approaches both qualitatively and quantitatively. Moreover, we demonstrate that when using our stabilization result as an input, the number of false detection significantly reduces in an independent moving object detection method, which is widely applied to wide area aerial surveillance applications. Bor-Jeng Chen, Gérard G. Medioni |
WACV | 2 |
| 2016 | Special issue on Assistive Computer Vision and Robotics - Part I
Giovanni Maria Farinella, Takeo Kanade, Marco Leo, Gérard G. Medioni, Mohan M. Trivedi |
Comput. Vis. Image Underst. | 4 |
| 2016 | Special Issue on Assistive Computer Vision and Robotics - "Assistive Solutions for Mobility, Communication and HMI"
Giovanni Maria Farinella, Takeo Kanade, Marco Leo, Gérard G. Medioni, Mohan M. Trivedi |
Comput. Vis. Image Underst. | 4 |
| 2016 | RGB-D camera based wearable navigation system for the visually impaired
Young Hoon Lee, Gérard G. Medioni |
Comput. Vis. Image Underst. | 2 |
| 2016 | Persistent people tracking and face capture using a PTZ camera
Yinghao Cai, Gérard G. Medioni |
Mach. Vis. Appl. | 2 |
| 2015 | Motion propagation detection association for multi-target tracking in wide area aerial surveillanceabstractWe propose a novel approach to track multiple targets with weak appearance in low frame rate wide area aerial videos. In real world scenarios, non-linear motion such as sharp turns after slowing down or U-shape trajectories occur. Performing accurate matching without introducing undesired trajectories is very challenging. To tackle various motion patterns, we sequentially optimizing an objective function and propagating motion information at each time step in a sliding temporal window. We show how to exploit an optimal short track (tracklet) for each detection in the first frame of each window using dynamic programming. Tracklets obtained in the window are then associated with existing tracks iteratively to form final tracks. We reduce false alarms in background subtraction motion detection with the aid of optical flow. Our system is tested on two challenging datasets. The quantitative evaluation on a long annotated aerial video sequence shows that the proposed approach outperforms state-of-the-art detection and tracking methods in all common axes of evaluation metrics. Bor-Jeng Chen, Gérard G. Medioni |
AVSS | 2 |
| 2015 | The multi-strand graph for a PTZ trackerabstractHigh-resolution images can be used to resolve matching ambiguities between trajectory fragments (tracklets), which is one of the main challenges in multiple target tracking. A PTZ camera, which can pan, tilt and zoom, is a powerful and efficient tool that offers both close-up views and wide area coverage on demand. The wide-area view makes it possible to track many targets while the close-up view allows individuals to be identified from high-resolution images of their faces. A central component of a PTZ tracking system is a scheduling algorithm that determines which target to zoom in on. In this paper we study this scheduling problem from a theoretical perspective, where the high resolution images are also used for tracklet matching. We propose a novel data structure, the Multi-Strand Tracking Graph (MSG), which represents the set of tracklets computed by a tracker and the possible associations between them. The MSG allows efficient scheduling as well as resolving - directly or by elimination - matching ambiguities between tracklets. The main feature of the MSG is the auxiliary data saved in each vertex, which allows efficient computation while avoiding time-consuming graph traversal. Synthetic data simulations are used to evaluate our scheduling algorithm and to demonstrate its superiority over a naïve one. Shachaf Melman, Yael Moses, Gérard G. Medioni, Yinghao Cai |
AVSS | 3 |
| 2015 | Geolocalization using mobile phone and street grid map in dynamic environmentabstractWe present a robust system for geolocalization in dynamic environments. Our application is a camera system designed to help the visually impaired navigate. It is also suitable for healthy eyesight users to find their way around unfamiliar areas. We combine visual odometry (VO) with the semantic information available in map to estimate the global coordinates of the walking users. In order to handle dynamic environments, our approach exploits the ground plane constraint in estimating the visual odometry. The motion estimation results are fed into a Monte Carlo Localization framework which localizes the user by matching the local motion trajectory with the shape of the street network found in the map. We validated our system with video sequences captured in different crowded urban environments. Experimental results show that our method not only corrects the cumulative drifting error but also manages to recover from serious visual odometry failure. Tung-Sing Leung, Gérard G. Medioni |
ICME | 2 |
| 2015 | Intersecting Manifolds: Detection, Segmentation, and Labeling
Shay Deutsch, Gérard G. Medioni |
IJCAI | 2 |
| 2015 | Convex Cut: A realtime pseudo-structure extraction algorithm for 3D point cloud dataabstractIn this paper, a realtime pseudo-structure extraction algorithm for 3D indoor point cloud data (PCD) is proposed. This algorithm is called Convex Cut (CC) because of its two main steps: cutting the PCD with arbitrary planes, and extracting convex parts. CC can be used as a preprocessing module for other existing algorithms to extract static parts in dynamic environments or to represent a principal 3D model of a given PCD. Its calculation time is 24 milliseconds for 50k PCD on a consumer PC, and it yields a precision value of 0.90 and a recall value of 0.99 on average in highly dynamic and cluttered environments. Some possible applications are explained such as simultaneous localization and mapping in dynamic environments, efficient dense map representation, robust 3D scan matching with plane features, and natural motion planning. ChangHyun Jun, Jihwan Youn, Jongmoo Choi, Gérard G. Medioni, Nakju Lett Doh |
IROS | 4 |
| 2015 | Surface Oriented Traverse for robust instance detection in RGB-DabstractWe address the problem of robust instance detection in RGB-D image in the presence of noisy data, cluttering, partial occlusion and large pose variation. We extract contour points from the depth image, construct a Surface Oriented Traverse (SOT) feature for each contour point and further classify it as either belonging or not belonging to the instance of interest. Starting from each contour point, its SOT feature is constructed by traversing and uniformly sampling along an oriented geodesic path on the object surface. After classification, all contour points vote for an instance-specific saliency map, from which the instance of interest is finally localized. Compared with the holistic template-based and learning-based methods, our method inherits advantages of the feature-based methods in dealing with cluttering, partial occlusion, and large pose variation. Furthermore, our method does not require accurate 3D models or high quality laser scan data as input and takes noisy data from commodity 3D sensors. Experimental results on the public RGB-D Object Dataset and our FindMe RGB-D Dataset demonstrate the effectiveness and robustness of our proposed instance detection algorithm. Ruizhe Wang 0002, Gérard G. Medioni, Wenyi Zhao |
IROS | 2 |
| 2015 | 3-D Mediated Detection and Tracking in Wide Area Aerial SurveillanceabstractWe address the problem of tracking many moving objects in wide area aerial surveillance. We propose that using 3-D information significantly improves performance over standard state of the art trackers, which rely on 2-D stabilization. We present and contrast two 3-D mediated approaches. The first approach assumes a dense 3-D model has previously been obtained. Given the camera poses, we can predict the image flow between frames and perform pixel-level classification for detection. The second approach instead computes the image flow, and relies on the epipolar geometry constraint to distinguish object motion from parallax. Our experiments on real imagery show a significant improvement in probability of detection (from 36% to 77%/67%), false alarm rate (from 75% to 3%/4%), and a speedup of an order of magnitude for the tracker itself. Using explicit 3-D improves the results, at a higher computational cost. Bor-Jeng Chen, Gérard G. Medioni |
WACV | 2 |
| 2015 | Multimodal Registration of Multiple Retinal Images Based on Line StructuresabstractWe propose a framework to perform multimodal registration of multiple images. In retinal imaging, this alignment enables the physician to correlate the features across modalities, which can help formulate a diagnosis. The images appear very different and there are few reliable modality-invariant features. We base our registration on the salient line structures extracted with a tensor-voting approach and aligned to minimize the Chamfer distance. For every pair of images, we match the line junctions and extremities to get a candidate transformation that is further refined with an Iterative Closest Point approach. We use a global chained registration framework to recover from failed registration and we account for non-planarities with a Thin-Plate Splines deformation. Our approach can handle large variations across modalities and is evaluated on real-world retinal images with 5 modalities per eye. We achieve an average error of 52 µm on our dataset. Matthias Hernandez, Gérard G. Medioni, Zhihong Hu, Srinivas R. Sadda |
WACV | 2 |
| 2015 | Progressive 3D Model Acquisition with a Commodity Hand-Held CameraabstractWe present a system for progressive 3D model acquisition with a commodity hand-held camera. The pipeline starts with a real-time scanning stage accomplished using a sparse point-based tracker for camera pose estimation and a dense patch-based tracker for dense reconstruction. While the user scans the target object, a model composed of local planar patches is continuously updated and displayed for visual feedback. This live feedback loop allows the user to choose new viewpoints based on the state of the current reconstruction, and determine if the model is completely covered in desired details. After live scanning is completed, our system refines the reconstructed patches into denser and more accurate patches through an offline model refinement procedure. We demonstrate the ability of our system on various real datasets. Zhuoliang Kang, Gérard G. Medioni |
WACV | 2 |
| 2015 | Near laser-scan quality 3-D face reconstruction from a low-quality depth stream
Matthias Hernandez, Jongmoo Choi, Gérard G. Medioni |
Image Vis. Comput. | 3 |
| 2014 | Persistent Tracking for Wide Area Aerial SurveillanceabstractPersistent surveillance of large geographic areas from unmanned aerial vehicles allows us to learn much about the daily activities in the region of interest. Nearly all of the approaches addressing tracking in this imagery are detection-based and rely on background subtraction or frame differencing to provide detections. This, however, makes it difficult to track targets once they slow down or stop, which is not acceptable for persistent tracking, our goal. We present a multiple target tracking approach that does not exclusively rely on background subtraction and is better able to track targets through stops. It accomplishes this by effectively running two trackers in parallel: one based on detections from background subtraction providing target initialization and reacquisition, and one based on a target state regressor providing frame to frame tracking. We evaluated the proposed approach on a long sequence from a wide area aerial imagery dataset, and the results show improved object detection rates and ID-switch rates with limited increases in false alarms compared to the competition. Jan Prokaj, Gérard G. Medioni |
CVPR | 2 |
| 2014 | 3D Modeling from Wide Baseline Range Scans Using Contour CoherenceabstractRegistering 2 or more range scans is a fundamental problem, with application to 3D modeling. While this problem is well addressed by existing techniques such as ICP when the views overlap significantly at a good initialization, no satisfactory solution exists for wide baseline registration. We propose here a novel approach which leverages contour coherence and allows us to align two wide baseline range scans with limited overlap from a poor initialization. Inspired by ICP, we maximize the contour coherence by building robust corresponding pairs on apparent contours and minimizing their distances in an iterative fashion. We use the contour coherence under a multi-view rigid registration framework, and this enables the reconstruction of accurate and complete 3D models from as few as 4 frames. We further extend it to handle articulations, and this allows us to model articulated objects such as human body. Experimental results on both synthetic and real data demonstrate the effectiveness and robustness of our contour coherence based registration approach to wide baseline range scans, and to 3D modeling. Ruizhe Wang 0002, Jongmoo Choi, Gérard G. Medioni |
CVPR | 3 |
| 2014 | Learning symbolic descriptions of activities from examples in WAASabstractWe present an automatic system that learns symbolic representations of activities from examples in Wide Area Aerial Surveillance (WAAS). In the previous work, we presented an ERM (Entity Relationship Models)-based activity recognition system in which finding an activity is equivalent to sending a query, defined by SQL statements, to a Relational DataBase Management System (RDBMS). The system enables us to identify spatial and geo-spatial activities in WAAS as long as activities are carefully defined by human operators. Here, we show how to infer a structured definition of an activity from examples provided by a user. Our system randomly generates a set of possible SQL statements using a logic generator in a MCMC framework, uses a memory-based RDBMS to validate generated SQL statements with the input data/database, and selects the best answer that allows the RDBMS to explain the input positive examples while excluding negative examples. We have evaluated our system on real visual tracks. Our system can find activity definitions from input examples and associated query results including motion patterns (e.g., "loop") and geospatial activities (e.g., "parking in a lot"). Jongmoo Choi, Gérard G. Medioni |
SIGSPATIAL/GIS | 2 |
| 2014 | An Examination of Multivariate Time Series Hashing with Applications to Health CareabstractAs large-scale multivariate time series data become increasingly common in application domains, such as health care and traffic analysis, researchers are challenged to build efficient tools to analyze it and provide useful insights. Similarity search, as a basic operator for many machine learning and data mining algorithms, has been extensively studied before, leading to several efficient solutions. However, similarity search for multivariate time series data is intrinsically challenging because (1) there is no conclusive agreement on what is a good similarity metric for multivariate time series data and (2) calculating similarity scores between two time series is often computationally expensive. In this paper, we address this problem by applying a generalized hashing framework, namely kernelized locality sensitive hashing, to accelerate time series similarity search with a series of representative similarity metrics. Experiment results on three large-scale clinical data sets demonstrate the effectiveness of the proposed approach. David C. Kale, Dian Gong, Zhengping Che, Yan Liu 0002, Gérard G. Medioni, Randall C. Wetzel, Patrick Ross |
ICDM | 5 |
| 2014 | Aerial Implicit 3D Video Stabilization Using Epipolar Geometry ConstraintabstractWe present an accurate video stabilization method on aerial videos using the epipolar geometry constraint. Most previous methods used 2D homography for stabilization, but failed to overcome the parallax problem. In this work, we propose to use dense correspondences for stabilization and the epipolar constraint to deal with the parallax effect. We start by estimating the dense correspondences between two frames. The dense correspondences are then used to estimate the epipolar geometry. The epipolar geometry has an implicit 3D constraint that can be used to improve the dense correspondences and handle the parallax. We evaluate our method on a real-life database containing three aerial image sequences. We also compare our method with the dominant 2D method for aerial video stabilization. The quantitative result demonstrates the effectiveness of our approach. Loc Huynh, Jongmoo Choi, Gérard G. Medioni |
ICPR | 3 |
| 2014 | Pose Independent Face Recognition by Localizing Local Binary Patterns via Deformation ComponentsabstractIn this paper we address the problem of pose independent face recognition with a gallery set containing one frontal face image per enrolled subject while the probe set is composed by just a face image undergoing pose variations. The approach uses a set of aligned 3D models to learn deformation components using a 3D Morph able Model (3DMM). This further allows fitting a 3DMM efficiently on an image using a Ridge regression solution, regularized on the face space estimated via PCA. Then the approach describes each profile face by computing Local Binary Pattern (LBP) histograms localized on each deformed vertex, projected on a rendered frontal view. In the experimental result we evaluate the proposed method on the CMU Multi-PIE to assess face recognition algorithm across pose. We show how our process leads to higher performance than regular baselines reporting high recognition rate considering a range of facial poses in the probe set, up to ±45°. Finally we remark that our approach can handle continuous pose variations and it is comparable with recent state-of-the-art approaches. Iacopo Masi, Claudio Ferrari, Alberto Del Bimbo, Gérard G. Medioni |
ICPR | 4 |
| 2014 | Automatic acquisition and animation of virtual avatarsabstractThe USC Institute for Creative Technologies will demonstrate a pipline for automatic reconstruction and animation of lifelike 3D avatars acquired by rotating the user's body in front of a single Microsoft Kinect sensor. Based on a fusion of state-of-the-art techniques in computer vision, graphics, and animation, this approach can produce a fully rigged character model suitable for real-time virtual environments in less than four minutes. Ari Shapiro, Andrew W. Feng, Ruizhe Wang 0002, Gérard G. Medioni, Mark T. Bolas, Evan A. Suma |
VR | 4 |
| 2014 | Exploring context information for inter-camera multiple target trackingabstractIn this paper, we present a new solution to inter-camera multiple target tracking with non-overlapping fields of view. The identities of people are maintained when they are moving from one camera to another. Instead of matching snapshots of people across cameras, we mainly explore what kind of context information from videos can be used for inter-camera tracking. We introduce two kinds of context information, spatio-temporal context and relative appearance context in this paper. The spatio-temporal context indicates a way of collecting samples for discriminative appearance learning where target-specific appearance models are learned to distinguish different people from each other. The relative appearance context models inter-object appearance similarities for people walking in proximity. The relative appearance model helps disambiguate individual appearance matching across cameras. We show improved performance with context information for inter-camera tracking. Our method achieves promising results in two crowded scenes compared with state-of-art methods. Yinghao Cai, Gérard G. Medioni |
WACV | 2 |
| 2014 | Real-time 3-D face tracking and modeling framework for mid-res camabstractWe present a robust, real-time 3-D face tracking and modeling system providing accurate 6 degree-of-freedom head pose in the presence of large out-of-plane motion, strong expression changes, and partial occlusions. In this paper, we have extended the previous 3-D face tracking and modeling framework [10] with automatic initialization, reacquisition, and automatic pose correction. Our system first generates a 3-D face model from a single frontal image. We then extract uniformly distributed random points and track them in 2-D. Given these correspondences, the 3-D head pose is robustly estimated using a RANSAC-PnP process. As the head moves, we dynamically add new feature points to handle a large range of poses. A measure of the accumulated error over time allows an auto-correction mechanism to recover from drift when necessary. If the tracker gets lost, due to motion blur or strong occlusions, the system re-initializes. We present live demo results, which shows excellent tracking under large motion (roll: 360°, yaw: ±90°, pitch: −60° to +90°), fast movement, occlusion and facial expression variations. The system runs at 14 fps on a laptop CPU. By experiments on different datasets, our method shows state of the art results. Jongmoo Choi, Anh Tuan Tran 0001, Yann Dumortier, Gérard G. Medioni |
WACV | 4 |
| 2014 | Fast dense 3D reconstruction using an adaptive multiscale discrete-continuous variational methodabstractWe present a system for fast dense 3D reconstruction with a hand-held camera. Walking around a target object, we shoot sequential images using continuous shooting mode. High-quality camera poses are obtained offline using structure-from-motion (SfM) algorithm with Bundle Adjustment. Multi-view stereo is solved using a new, efficient adaptive multiscale discrete-continuous variational method to generate depth maps with sub-pixel accuracy. Depth maps are then fused into a 3D model using volumetric integration with truncated signed distance function (TSDF). Our system is accurate, efficient and flexible: accurate depth maps are estimated with sub-pixel accuracy in stereo matching; dense models can be achieved within minutes as major algorithms parallelized on multi-core processor and GPU; various tasks can be handled (e.g. reconstruction of objects in both indoor and outdoor environment with different scales) without specific hand-tuning parameters. We evaluate our system quantitatively and qualitatively on Middlebury benchmark and another dataset collected with a smartphone camera. Zhuoliang Kang, Gérard G. Medioni |
WACV | 2 |
| 2014 | Co-trained generative and discriminative trackers with cascade particle filter
Thang Ba Dinh, Gérard G. Medioni |
Comput. Vis. Image Underst. | 3 |
| 2014 | Rapid avatar capture and simulation using commodity depth sensorsabstractABSTRACT We demonstrate a method of acquiring a 3D model of a human using commodity scanning hardware and then controlling that 3D figure in a simulated environment in only a few minutes. The model acquisition requires four static poses taken at 90° angles relative to each other. The 3D model is then given a skeleton and smooth binding information necessary for control and simulation. The 3D models that are captured are suitable for use in applications where recognition and distinction among characters by shape, form, or clothing is important, such as small group or crowd simulations or other socially oriented applications. Because of the speed at which a human figure can be captured and the low hardware requirements, this method can be used to capture, track, and model human figures as their appearances change over time. Copyright © 2014 John Wiley & Sons, Ltd. Ari Shapiro, Andrew W. Feng, Ruizhe Wang 0002, Hao Li 0015, Mark T. Bolas, Gérard G. Medioni, Evan A. Suma |
Comput. Animat. Virtual Worlds | 6 |
| 2014 | Structured Time Series Analysis for Human Action Segmentation and RecognitionabstractWe address the problem of structure learning of human motion in order to recognize actions from a continuous monocular motion sequence of an arbitrary person from an arbitrary viewpoint. Human motion sequences are represented by multivariate time series in the joint-trajectories space. Under this structured time series framework, we first propose Kernelized Temporal Cut (KTC), an extension of previous works on change-point detection by incorporating Hilbert space embedding of distributions, to handle the nonparametric and high dimensionality issues of human motions. Experimental results demonstrate the effectiveness of our approach, which yields realtime segmentation, and produces high action segmentation accuracy. Second, a spatio-temporal manifold framework is proposed to model the latent structure of time series data. Then an efficient spatio-temporal alignment algorithm Dynamic Manifold Warping (DMW) is proposed for multivariate time series to calculate motion similarity between action sequences (segments). Furthermore, by combining the temporal segmentation algorithm and the alignment algorithm, online human action recognition can be performed by associating a few labeled examples from motion capture data. The results on human motion capture data and 3D depth sensor data demonstrate the effectiveness of the proposed approach in automatically segmenting and recognizing motion sequences, and its ability to handle noisy and partially occluded data, in the transfer learning module. Dian Gong, Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Comparing strategies for 3D face recognition from a 3D sensorabstractWe address the problem of 3D face recognition from 3D data, using different strategies. One strategy (1F-NF), explored earlier, is to match each individual frame to a set of reference frames. A second one (1F-3D) is to replace the set of reference frames by a 3D model resulting from the integration of individual frames. A third strategy (3D-3D) is to use a 3D face model inferred from multiple frames as the input probe. We show that the recognition performance using 3D model to 3D model outperforms the others, at the cost of a delay in response, due to the model building step. Jongmoo Choi, Ayush Sharma, Gérard G. Medioni |
RO-MAN | 3 |
| 2013 | Towards a practical PTZ face detection and tracking systemabstractWe address the problem of automatic face detection and tracking in uncontrolled scenarios using a pan-tilt-zoom (PTZ) network camera, which could prove most helpful in forensic applications. The detected faces are associated with the corresponding people and trajectories. The dynamic nature of real-world scenarios and real-time restrictions complicate our task. Different from previous work which use a mixture of wide angle cameras and PTZ cameras, we explore the limits to what can be expected from a single PTZ camera. The system first detects and tracks pedestrians in zoomed-out mode, then selects, using a scheduler, a person to zoom in to. After zoom in, we come back to wide area mode, and solve the person-to-person, face-to-person and face-to-face data association problems. Extensive experiments in challenging indoor and outdoor uncontrolled conditions demonstrate the effectiveness of the proposed system. Yinghao Cai, Gérard G. Medioni, Thang Ba Dinh |
WACV | 2 |
| 2012 | Kernelized Temporal Cut for Online Temporal Segmentation and Recognition
Dian Gong, Gérard G. Medioni, Sikai Zhu |
ECCV (3) | 2 |
| 2012 | Tracking Using Motion Patterns for Very Crowded Scenes
Dian Gong, Gérard G. Medioni |
ECCV (2) | 3 |
| 2012 | Activity recognition in wide aerial video surveillance using entity relationship modelsabstractWe present the design and implementation of an activity recognition system in wide area aerial video surveillance using Entity Relationship Models (ERM). In this approach, finding an activity is equivalent to sending a query to a Relational DataBase Management System (RDBMS). By incorporating reference imagery and Geographic Information System (GIS) data, tracked objects can be associated with physical meanings, and several high levels of reasoning, such as traffic patterns or abnormal activity detection, can be performed. We demonstrate that different types of activities, with hierarchical structure, multiple actors, and context information, are effectively and efficiently defined and inferred using the ERM framework. We also show how visual tracks can be better interpreted as activities by using geo information. Experimental results on both real visual tracks and GPS traces validate our approach. Jongmoo Choi, Yann Dumortier, Jan Prokaj, Gérard G. Medioni |
SIGSPATIAL/GIS | 4 |
| 2012 | Robust Multiple Manifold Structure Learning
Dian Gong, Gérard G. Medioni |
ICML | 3 |
| 2012 | Real-time staircase detection from a wearable stereo system
Young Hoon Lee, Tung-Sing Leung, Gérard G. Medioni |
ICPR | 3 |
| 2012 | Real-time 3D face identification from a depth camera
Rui Min 0002, Jongmoo Choi, Gérard G. Medioni, Jean-Luc Dugelay |
ICPR | 3 |
| 2012 | Real-time 3-D face tracking and modeling from awebcamabstractWe first infer a 3-D face model from a single frontal image using automatically extracted 2-D landmarks and deforming a generic 3-D model. Then, for any input image, we extract feature points and track them in 2-D. Given these correspondences, sometimes noisy and incorrect, we robustly estimate the 3-D head pose using PnP and a RANSAC process. As the head moves, we dynamically add new feature points to handle a large range of poses. When the tracker gets lost, due to motion blur or occlusions, the system re-initializes by matching feature points to the reference frontal image feature points. Our system runs in real-time (>;15Hz) on a standard CPU with a GPU card. We present results on stored video and will present a live demo, showing excellent tracking under large motion, fast movement, occlusion and facial expression variations. We also show comparative results with the ground truth BU head tracking dataset. Jongmoo Choi, Yann Dumortier, Sang-Il Choi, Muhammad Bilal Ahmad, Gérard G. Medioni |
WACV | 5 |
| 2012 | Accurate efficient mosaicking for Wide Area Aerial SurveillanceabstractWide Area Aerial Surveillance (WAAS) imagery is captured by an array of smaller sensors sharing an optical center, instead of one large sensor. It is desirable to generate a single image (mosaic) from the sensor array, since it simplifies higher level vision tasks. It is important that the mosaic be of high quality, without noticeable seams, and be estimated efficiently for every frame of the video. We propose a piecewise affine model to handle image distortions not captured by a homography. This model has more degrees of freedom than a standard lens distortion model used in previous work, yet can be estimated just as efficiently by carefully selecting a small number of constraints in the optimization of model parameters. We have evaluated our algorithm on publicly available WAAS imagery and the results show our proposed model produces more accurate mosaics than the widely used lens distortion model, thus simplifying further stages of image analysis. Jan Prokaj, Gérard G. Medioni |
WACV | 2 |
| 2012 | A Closed-Form Solution to Tensor Voting: Theory and ApplicationsabstractWe prove a closed-form solution to tensor voting (CFTV): Given a point set in any dimensions, our closed-form solution provides an exact, continuous, and efficient algorithm for computing a structure-aware tensor that simultaneously achieves salient structure detection and outlier attenuation. Using CFTV, we prove the convergence of tensor voting on a Markov random field (MRF), thus termed as MRFTV, where the structure-aware tensor at each input site reaches a stationary state upon convergence in structure propagation. We then embed structure-aware tensor into expectation maximization (EM) for optimizing a single linear structure to achieve efficient and robust parameter estimation. Specifically, our EMTV algorithm optimizes both the tensor and fitting parameters and does not require random sampling consensus typically used in existing robust statistical techniques. We performed quantitative evaluation on its accuracy and robustness, showing that EMTV performs better than the original TV and other state-of-the-art techniques in fundamental matrix estimation for multiview stereo matching. The extensions of CFTV and EMTV for extracting multiple and nonlinear structures are underway. Tai-Pang Wu, Sai-Kit Yeung, Jiaya Jia, Chi-Keung Tang, Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2012 | A Framework for Robust Online Video Contrast Enhancement Using Modularity OptimizationabstractWe address the problem of video contrast enhancement. Existing techniques either do not exploit temporal information at all or do not exploit it correctly. This results in inconsistency that causes undesirable flash and flickering artifacts. Our method analyzes video streams and cluster frames that are similar to each other. Our method does not have omniscient information about the entire video sequence. It is an online process with a fixed delay. A sliding window mechanism successfully detects shot boundaries “on-the-fly” in a video. A graph-based technique called “modularity” performs automatic clustering of video frames without a priori information about clusters. For every cluster in the video, we extract key frames belonging to each cluster using eigen analysis and estimate enhancement parameters for only the key frame, then use these parameters to enhance frames belonging to that cluster, thus making our method robust. We evaluate the clustering method on video sequences from the TRECVid 2001 dataset and compare it with existing methods. We show reduction of flash artifacts in enhanced videos. We show statistically significant improvement in perceived video quality and validate that by conducting experiments on human observers. We show application of our clustering process to perform robust video segmentation. Anustup Choudhury, Gérard G. Medioni |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | Context tracker: Exploring supporters and distracters in unconstrained environmentsabstractVisual tracking in unconstrained environments is very challenging due to the existence of several sources of varieties such as changes in appearance, varying lighting conditions, cluttered background, and frame-cuts. A major factor causing tracking failure is the emergence of regions having similar appearance as the target. It is even more challenging when the target leaves the field of view (FoV) leading the tracker to follow another similar object, and not reacquire the right target when it reappears. This paper presents a method to address this problem by exploiting the context on-the-fly in two terms: Distracters and Supporters. Both of them are automatically explored using a sequential randomized forest, an online template-based appearance model, and local features. Distracters are regions which have similar appearance as the target and consistently co-occur with high confidence score. The tracker must keep tracking these distracters to avoid drifting. Supporters, on the other hand, are local key-points around the target with consistent co-occurrence and motion correlation in a short time span. They play an important role in verifying the genuine target. Extensive experiments on challenging real-world video sequences show the tracking improvement when using this context information. Comparisons with several state-of-the-art approaches are also provided. Thang Ba Dinh, Nam Sy Vo, Gérard G. Medioni |
CVPR | 3 |
| 2011 | Using 3D scene structure to improve trackingabstractIn this work we consider the problem of tracking objects from a moving airborne platform in wide area surveillance through long occlusions and/or when their motion is unpredictable. The main idea is to take advantage of the known 3D scene structure to estimate a dynamic occlusion map, and to use the occlusion map to determine traffic entry and exit into these zones, which we call sources and sinks. Then the track linking problem is formulated as an alignment of sequences of tracks entering a sink and leaving a source. The sequence alignment problem is solved optimally and efficiently using dynamic programming. We have evaluated our algorithm on a vehicle tracking task in wide area motion imagery and have shown that track fragmentation is significantly decreased and outperforms the Hungarian algorithm. Jan Prokaj, Gérard G. Medioni |
CVPR | 2 |
| 2011 | High resolution face sequences from a PTZ network cameraabstractWe propose here to acquire high resolution sequences of a person's face using a pan-tilt-zoom (PTZ) network camera. This capability should prove helpful in forensic analysis of video sequences as frames containing faces are tagged, and within a frame, windows containing faces can be retrieved. The system starts in pedestrian detector mode, where the lens angle is set widest, and detects people using a pedestrian detector module. The camera then changes to the region of interest (ROI) focusing mode where the parameters are automatically tuned to put the upper body of the detected person, where the face should appear, in the field of view (FOV). Then, in the face detection mode, the face is detected using a face detector module, and the system switches to an active tracking mode consisting a control loop to actively follow the detected face with two different modules: a tracker to track the face in the image, and a camera control module to adjust the camera parameters. During this loop, our tracker learns online the face appearance in multiple views under all condition changes. It runs robustly at 15 fps and is able to reacquire the face of interest after total occlusion or leaving FOV. We compare our tracker with various state-of-the-art tracking methods in terms of precision and running time performance. Extensive experiments in challenging indoor and outdoor conditions are also demonstrated to validate the complete system. Thang Ba Dinh, Nam Sy Vo, Gérard G. Medioni |
FG | 3 |
| 2011 | Dynamic Manifold Warping for view invariant action recognitionabstractWe address the problem of learning view-invariant 3D models of human motion from motion capture data, in order to recognize human actions from a monocular video sequence with arbitrary viewpoint. We propose a Spatio-Temporal Manifold (STM) model to analyze non-linear multivariate time series with latent spatial structure and apply it to recognize actions in the joint-trajectories space. Based on STM, a novel alignment algorithm Dynamic Manifold Warping (DMW) and a robust motion similarity metric are proposed for human action sequences, both in 2D and 3D. DMW extends previous works on spatio-temporal alignment by incorporating manifold learning. We evaluate and compare the approach to state-of-the-art methods on motion capture data and realistic videos. Experimental results demonstrate the effectiveness of our approach, which yields visually appealing alignment results, produces higher action recognition accuracy, and can recognize actions from arbitrary views with partial occlusion. Dian Gong, Gérard G. Medioni |
ICCV | 2 |
| 2011 | Aerial 3D reconstruction with line-constrained dynamic programmingabstractAerial imagery of an urban environment is often characterized by significant occlusions, sharp edges, and textureless regions, leading to poor 3D reconstruction using conventional multi-view stereo methods. In this paper, we propose a novel approach to 3D reconstruction of urban areas from a set of uncalibrated aerial images. A very general structural prior is assumed that urban scenes consist mostly of planar surfaces oriented either in a horizontal or an arbitrary vertical orientation. In addition, most structural edges associated with such surfaces are also horizontal or vertical. These two assumptions provide powerful constraints on the underlying 3D geometry. The main contribution of this paper is to translate the two constraints on 3D structure into intra-image-column and inter-image-column constraints, respectively, and to formulate the dense reconstruction as a 2-pass Dynamic Programming problem, which is solved in complete parallel on a GPU. The result is an accurate cloud of 3D dense points of the underlying urban scene. Our algorithm completes the reconstruction of 1M points with 160 available discrete height levels in under a hundred seconds. Results on multiple datasets show that we are capable of preserving a high level of structural detail and visual quality. Huei-Hung Liao, Yuping Lin, Gérard G. Medioni |
ICCV | 3 |
| 2011 | Robust unsupervised motion pattern inference from video and applicationsabstractWe propose an unsupervised learning framework to infer motion patterns in videos and in turn use them to improve tracking of moving objects in sequences from static cameras. Based on tracklets, we use a manifold learning method Tensor Voting to infer the local geometric structures in (x, y) space, and embed tracklet points into (x, y, θ) space, where θ represents motion direction. In this space, points automatically form intrinsic manifold structures, each of which corresponds to a motion pattern. To define each group, a novel robustmanifold grouping algorithm is proposed. Tensor Voting is performed to provide multiple geometric cues which formulate multiple similarity kernels between any pair of points, and a spectral clustering technique is used in this multiple kernel setting. The grouping algorithm achieves better performance than state-of-the-art methods in our applications. Extracted motion patterns can then be used as a prior to improve the performance of any object tracker. It is especially useful to reduce false alarms and ID switches. Experiments are performed on challenging real-world sequences, and a quantitative analysis of the results shows the framework effectively improves state-of-the-art tracker. Gérard G. Medioni |
ICCV | 2 |
| 2011 | 3D object recognition in range images using visibility contextabstractRecognizing and localizing queried objects in range images plays an important role for robotic manipulation and navigation. Even though it has been steadily studied, it is still a challenging task for scenes with occlusion and clutter. We present a novel approach to object recognition that boosts dissimilarity between queried objects and similar-shaped background objects in the scene by maximizing use of the visibility context. We design a new point pair feature containing discriminative description inferred from the visibility context. Also, we propose a pose estimation method that accurately localizes objects using these point pair matches. Finally, two measures of validity are suggested to discard false detections. With 10 query objects, our approach is evaluated on depth images of cluttered office scenes captured from a real-time range sensor. The experimental results demonstrate that our method remarkably outperforms two state-of-the-art methods in terms of recognition (recall & precision) and runtime performance. Gérard G. Medioni |
IROS | 2 |
| 2011 | Co-training framework of generative and discriminative trackers with partial occlusion handlingabstractPartial occlusion is a challenging problem in object tracking. In online visual tracking, it is the critical factor causing drift. To address this problem, we propose a novel approach using a co-training framework of generative and discriminative trackers. Our approach is able to detect the occluding region and continuously update both the generative and discriminative models using the information from the non-occluded part. The generative model encodes all of the appearance variations using a low dimension subspace, which helps provide a strong reacquisition ability. Meanwhile, the discriminative classifer, an online support vector machine, focuses on separating the object from the background using a Histograms of Oriented Gradients (HOG) feature set. For each search window, an occlusion likelihood map is generated by the two trackers through a co-decision process. If there is disagreement between these two trackers, the movement vote of KLT local features is used as a referee. Precise occlusion segmentation is performed using MeanShift. Finally, each tracker recovers the occluded part and updates its own model using the new non-occluded information. Experimental results on challenging sequences with different types of objects are presented. We also compare with other state-of-the-art methods to demonstrate the superiority and robustness of our tracking framework. Thang Ba Dinh, Gérard G. Medioni |
WACV | 2 |
| 2011 | Urban scene understanding from aerial and ground LIDAR data
Gérard G. Medioni |
Mach. Vis. Appl. | 2 |
| 2011 | Efficient detection and tracking of moving objects in geo-coordinates
Yuping Lin, Gérard G. Medioni |
Mach. Vis. Appl. | 3 |
| 2010 | Accurate 3D face reconstruction from weakly calibrated wide baseline images with profile contoursabstractWe propose a method to generate a highly accurate 3D face model from a set of wide-baseline images in a weakly calibrated setup. Our approach is purely data driven, and produces faithful 3D models without any pre-defined models, unlike other statistical model-based approaches. Our results do not rely upon a critical initialization step nor parameters for optimization steps. We process 5 images (including profile views), infer the accurate poses of cameras in all views, and then infer a dense 3D face model. The quality of 3D face models depends on the accuracy of estimated head-camera motion. First, we propose to use an iterative bundle adjustment approach to remove outliers in corresponding points. Contours in the profile views are matched to provide reliable correspondences that link two opposite side of views together. For dense reconstruction, we propose to use a face-specific cylindrical representation which allows us to solve a global optimization problem for N-view dense aggregation. Profile contours are used once again to provide constraints in the optimization step. Experimental results using synthetic and real images show that our method provides accurate and stable reconstruction results on wide-baseline images. We compare our method with state of the art methods, and show that it provides significantly better results in terms of both accuracy and efficiency. Yuping Lin, Gérard G. Medioni, Jongmoo Choi |
CVPR | 2 |
| 2010 | 3D Face Reconstruction Using a Single or Multiple ViewsabstractWe present a 3D face reconstruction system that takes as input either one single view or several different views. Given a facial image, we first classify the facial pose into one of five predefined poses, then detect two anchor points that are then used to detect a set of predefined facial landmarks. Based on these initial steps, for a single view we apply a warping process using a generic 3D face model to build a 3D face. For multiple views, we apply sparse bundle adjustment to reconstruct 3D landmarks which are used to deform the generic 3D face model. Experimental results on the Color FERET and CMU multi-PIE databases confirm our framework is effective in creating realistic 3D face models that can be used in many computer vision applications, such as 3D face recognition at a distance. Jongmoo Choi, Gérard G. Medioni, Yuping Lin, Luciano Silva, Olga R. P. Bellon, Maurício Pamplona Segundo, Timothy C. Faltemier |
ICPR | 2 |
| 2010 | Color Constancy Using Standard Deviation of Color ChannelsabstractWe address here the problem of color constancy and propose a new method to achieve color constancy based on the statistics of images with color cast. Images with color cast have standard deviation of one color channel significantly different from that of other color channels. This observation is also applicable to local patches of images and ratio of the maximum and minimum standard deviation of color channels of local patches is used as a prior to select a pixel color as illumination color. We provide extensive validation of our method on commonly used datasets having images under varying illumination conditions and show our method to be robust to choice of dataset and at least as good as current state-of-the-art color constancy approaches. Anustup Choudhury, Gérard G. Medioni |
ICPR | 2 |
| 2010 | Dense Structure Inference for Object Classification in Aerial LIDAR DatasetabstractWe present a framework to classify small freeform objects in 3D aerial scans of a large urban area. The system first identifies large structures such as the ground surface and roofs of buildings densely built in the scene, by fitting planar patches and grouping adjacent patches similar in pose together. Then, it segments initial object candidates which represent the visible surface of an object using the identified structures. To deal with sparse density in points representing each candidate, we also propose a novel method to infer a dense 3D structure from the given sparse and noisy points without any meshes and iterations. To label object candidates, we build a tree-structure database of object classes, which captures latent patterns in shape of 3D objects in a hierarchical manner. We demonstrate our system on the aerial LIDAR dataset acquired from a few square kilometers of Ottawa. Gérard G. Medioni |
ICPR | 2 |
| 2010 | Dimensionality Estimation, Manifold Learning and Function Approximation using Tensor Voting
Philippos Mordohai, Gérard G. Medioni |
J. Mach. Learn. Res. | 2 |
| 2009 | StaRSaC: Stable random sample consensus for parameter estimationabstractWe address the problem of parameter estimation in presence of both uncertainty and outlier noise. This is a common occurrence in computer vision: feature localization is performed with an inherent uncertainty which can be described as Gaussian, with unknown variance; feature matching in multiple images produces incorrect data points. RANSAC is the preferred method to reject outliers if the variance of the uncertainty noise is known, but fails otherwise, by producing either a tight fit to an incorrect solution, or by computing a solution which includes outliers. We thus propose a new estimator which enforces stability of the solution with respect to the uncertainty bound. We show that the variance of the estimated parameters (VoP) exhibits ranges of stability with respect to this bound. Within this range of stability, we can accurately segment the inliers, and estimate the parameters, the variance of the Gaussian noise. We show how to compute this stable range using RANSAC and a search. We validate our results by extensive tests and comparison with state of the art estimators on both synthetic and real data sets. These include line fitting, homography estimation, and fundamental matrix estimation. The proposed method outperforms all others. Jongmoo Choi, Gérard G. Medioni |
CVPR | 2 |
| 2009 | Visual loop closing using multi-resolution SIFT grids in metric-topological SLAMabstractWe present an image based simultaneous localization and mapping (SLAM) framework with online, appearance only loop closing. We adopt a layered approach with metric maps over small areas at the local level and a global, graph based abstract topological framework to build consistent maps over large distances. Rao-Blackwellised particle filtering and sparse bundle adjustment are efficiently coupled with a stereo vision based odometry module to construct conditionally independent `submaps' using SIFT features. By extracting keyframes from these submaps, a multiresolution dictionary of distinct features is built online to learn a generative model of appearance and perform loop closure. Creating such a dictionary also enables the system to distinguish between similar regions during loop closure without requiring any offline training, as has been described in other approaches. Furthermore, instead of occupancy or grid maps, we build 3D reconstructions of the world; a model we plan to use as input to a scene interpretation module for providing navigational cues to the visually impaired. We demonstrate the robustness of our SLAM system with indoor and outdoor experiments for full 6 degrees of freedom motion using only a stereo camera in hand, running at 1 Hz on a standard PC. Vivek Pradeep, Gérard G. Medioni, James D. Weiland |
CVPR | 2 |
| 2009 | Motion pattern interpretation and detection for tracking moving vehicles in airborne videoabstractDetection and tracking of moving vehicles in airborne videos is a challenging problem. Many approaches have been proposed to improve motion segmentation on frame-by-frame and pixel-by-pixel bases, however, little attention has been paid to analyze the long-term motion pattern, which is a distinctive property for moving vehicles in airborne videos. In this paper, we provide a straightforward geometric interpretation of a general motion pattern in 4D space (x, y, vx, vy). We propose to use the tensor voting computational framework to detect and segment such motion patterns in 4D space. Specifically, in airborne videos, we analyze the essential difference in motion patterns caused by parallax and independent moving objects, which leads to a practical method for segmenting motion patterns (flows) created by moving vehicles in stabilized airborne videos. The flows are used in turn to facilitate detection and tracking of each individual object in the flow. Conceptually, this approach is similar to “track-before-detect” techniques, which involves temporal information in the process as early as possible. As shown in the experiments, many difficult cases in airborne videos, such as parallax, noisy background modeling and long term occlusions, can be addressed by our approach. Gérard G. Medioni |
CVPR | 2 |
| 2009 | Color constancy using denoising methods and cepstral analysisabstractWe address here the problem of color constancy and propose two new methods for achieving color constancy-the first method uses denoising techniques such as a Gaussian filter, Median filter, Bilateral filter and Non-local means filter to smooth the image for illuminant estimation, while the second method acts in the frequency domain by doing a cepstral analysis of the image. We provide extensive validation tests for our illuminant estimation on commonly used datasets having images under different illumination conditions, and the results show that both new methods outperform current state-of-the-art color constancy approaches, at a very low computational cost. Anustup Choudhury, Gérard G. Medioni |
ICIP | 2 |
| 2009 | Real time tracking using an active pan-tilt-zoom network cameraabstractWe present here a real time active vision system on a PTZ network camera to track an object of interest. We address two critical issues in this paper. One is the control of the camera through network communication to follow a selected object. The other is to track an arbitrary type of object in real time under conditions of pose, viewpoint and illumination changes. We analyze the difficulties in the control through the network and propose a practical solution for tracking using a PTZ network camera. Moreover, we propose a robust real time tracking approach, which enhances the effectiveness by using complementary features under a two-stage particle filtering framework and a multi-scale mechanism. To improve time performance, the tracking algorithm is implemented as a multi-threaded process in OpenMP. Comparative experiments with state-of-the-art methods demonstrate the efficiency and robustness of our system in various applications such as pedestrian tracking, face tracking, and vehicle tracking. Thang Ba Dinh, Gérard G. Medioni |
IROS | 3 |
| 2009 | 3-D model based vehicle recognitionabstractWe present a method for recognizing a vehicle's make and model in a video clip taken from an arbitrary viewpoint. This is an improvement over existing methods which require a front view. In addition, we present a Bayesian approach for establishing accurate correspondences in multiple view geometry. We take a model-based, top-down approach to classify vehicles. First, the vehicle pose is estimated in every frame by calculating its 3-D motion on a plane using a structure from motion algorithm. Then, exemplars from a database of 3-D models are rotated to the same pose as the vehicle in the video, and projected to the image. Features in the model images and the vehicle image are matched, and a model matching score is computed. The model with the best score is identified as the model of the vehicle in the video. Results on real video sequences are presented. Jan Prokaj, Gérard G. Medioni |
WACV | 2 |
| 2009 | Multiple-Target Tracking by Spatiotemporal Monte Carlo Markov Chain Data AssociationabstractWe propose a framework for tracking multiple targets, where the input is a set of candidate regions in each frame, as obtained from a state-of-the-art background segmentation module, and the goal is to recover trajectories of targets over time. Due to occlusions by targets and static objects, as also by noisy segmentation and false alarms, one foreground region may not correspond to one target faithfully. Therefore, the one-to-one assumption used in most data association algorithms is not always satisfied. Our method overcomes the one-to-one assumption by formulating the visual tracking problem in terms of finding the best spatial and temporal association of observations, which maximizes the consistency of both motion and appearance of trajectories. To avoid enumerating all possible solutions, we take a Data-Driven Markov Chain Monte Carlo (DD-MCMC) approach to sample the solution space efficiently. The sampling is driven by an informed proposal scheme controlled by a joint probability model combining motion and appearance. Comparative experiments with quantitative evaluations are provided. Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2009 | Identifying Noncooperative Subjects at a Distance Using Face Images and Inferred Three-Dimensional Face ModelsabstractWe present an approach to identify noncooperative individuals at a distance from a sequence of images, using 3-D face models. Most biometric features (such as fingerprints, hand shape, iris, or retinal scans) require cooperative subjects in close proximity to the biometric system. We process images acquired with an ultrahigh-resolution video camera, infer the location of the subjects' head, use this information to crop the region of interest, build a 3-D face model, and use this 3-D model to perform biometric identification. To build the 3-D model, we use an image sequence, as natural head and body motion provides enough viewpoint variation to perform stereomotion for 3-D face reconstruction. We have conducted experiments on a 2-D and 3-D databases collected in our laboratory. First, we found that metric 3-D face models can be used for recognition by using simple scaling method even though there is no exact scale in the 3-D reconstruction. Second, experiments using a commercial 3-D matching engine suggest the feasibility of the proposed approach for recognition against 3-D galleries at a distance (3, 6, and 9 m). Moreover, we show initial 3-D face modeling results on various factors including head motion, outdoor lighting conditions, and glasses. The evaluation results suggest that video data alone, at a distance of 3 to 9 meters, can provide a 3-D face shape that supports successful face recognition. The performance of 3-D-3-D recognition with the currently generated models does not quite match that of 2-D-2-D. We attribute this to the quality of the inferred models, and this suggests a clear path for future research. Gérard G. Medioni, Jongmoo Choi, Cheng-Hao Kuo, Douglas Fidaleo |
IEEE Trans. Syst. Man Cybern. Part A | 1 |
| 2008 | 3D face tracking and expression inference from a 2D sequence using manifold learningabstractWe propose a person-dependent, manifold-based approach for modeling and tracking rigid and nonrigid 3D facial deformations from a monocular video sequence. The rigid and nonrigid motions are analyzed simultaneously in 3D, by automatically fitting and tracking a set of landmarks. We do not represent all nonrigid facial deformations as a simple complex manifold, but instead decompose them on a basis of eight 1D manifolds. Each 1D manifold is learned offline from sequences of labeled expressions, such as smile, surprise, etc. Any expression is then a linear combination of values along these 8 axes, with coefficient representing the level of activation. We experimentally verify that expressions can indeed be represented this way, and that individual manifolds are indeed 1D. The manifold dimensionality estimation, manifold learning, and manifold traversal operation are all implemented in the N-D Tensor Voting framework. Using simple local operations, this framework gives an estimate of the tangent and normal spaces at every sample, and provides excellent robustness to noise and outliers. The output of our system, besides the tracked landmarks in 3D, is a labeled annotation of the expression. We demonstrate results on a number of challenging sequences. Wei-Kai Liao, Gérard G. Medioni |
CVPR | 2 |
| 2008 | Retinal image registration from 2D to 3DabstractWe propose a 2D registration method for multi-modal image sequences of the retinal fundus, and a 3D metric reconstruction of near planar surface from multiple views. There are two major contributions in our paper. For 2D registration, our method produces high registration rates while accounting for large modality differences. Compared with the state of the art method [5], our approach has higher registration rate (97.2% vs. 82.31%) while the computation time is much less. This is achieved by extracting features from the edge maps of the contrast enhanced images, and performing pairwise registration by matching the features in an iterative manner, maximizing the number of matches and estimating homographies accurately. The pairwise registration result is further globally optimized by an indirect registration process. For 3D registration part, images are registered to the reference frame by transforming points via a reconstructed 3D surface. The challenge is the reconstruction of a near planar surface, in which the shallow depth makes it a quasi-degenerate case for estimating the geometry from images. Our contribution is the proposed 4-pass bundle adjustment method that gives optimal estimation of all camera poses. With accurate camera poses, the 3D surface can be reconstructed using the images associated with the cameras with the largest baseline. Compared with state of the art 3D retinal image registration methods, our approach produces better results in all image sets. Yuping Lin, Gérard G. Medioni |
CVPR | 2 |
| 2008 | Online Tracking and Reacquisition Using Co-trained Generative and Discriminative Trackers
Thang Ba Dinh, Gérard G. Medioni |
ECCV (2) | 3 |
| 2008 | Distributed Visual Processing for a Home Visual Sensor NetworkabstractWe address issues dealing with distributed visual processing for a personal service robot in the Intelligent Home environment. We propose an efficient and reliable framework to organize and coordinate the vision sensor nodes: fixed cameras mounted on walls, and camera(s) on the mobile robot. We also propose key visual functionalities necessary for the robot to perform its activities. They include people detection and identification, action recognition, gesture recognition, and self-localization. We propose solutions to the different vision tasks, and present our implementation within this framework, validated with experimental results. Kwangsu Kim, Gérard G. Medioni |
WACV | 2 |
| 2008 | 2-D registration and 3-D shape inference of the retinal fundus from fluorescein images
Tae Eun Choe, Gérard G. Medioni, Isaac Cohen, Alexander C. Walsh, Srinivas R. Sadda |
Medical Image Anal. | 2 |
| 2008 | Inferring Segmented Dense Motion Layers Using 5D Tensor VotingabstractWe present a novel local spatiotemporal approach to produce motion segmentation and dense temporal trajectories from an image sequence. A common representation of image sequences is a 3D spatiotemporal volume, (x,y,t), and its corresponding mathematical formalism is the fiber bundle. However, directly enforcing the spatiotemporal smoothness constraint is difficult in the fiber bundle representation. Thus, we convert the representation into a new 5D space (x,y,t,vx,vy) with an additional velocity domain, where each moving object produces a separate 3D smooth layer. The smoothness constraint is now enforced by extracting 3D layers using the tensor voting framework in a single step that solves both correspondence and segmentation simultaneously. Motion segmentation is achieved by identifying those layers, and the dense temporal trajectories are obtained by converting the layers back into the fiber bundle representation. We proceed to address three applications (tracking, mosaic, and 3D reconstruction) that are hard to solve from the video stream directly because of the segmentation and dense matching steps, but become straightforward with our framework. The approach does not make restrictive assumptions about the observed scene or camera motion and is therefore generally applicable. We present results on a number of data sets. Changki Min, Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Map-Enhanced UAV Image Sequence Registration and Synchronization of Multiple Image SequencesabstractRegistering consecutive images from an airborne sensor into a mosaic is an essential tool for image analysts. Strictly local methods tend to accumulate errors, resulting in distortion. We propose here to use a reference image (such as a high resolution map image) to overcome this limitation. In our approach, we register a frame in an image sequence to the map using both frame-to-frame registration and frame-to-map registration iteratively. In frame-to-frame registration, a frame is registered to its previous frame. With its previous frame been registered to the map in the previous iteration, we can derive an estimated transformation from the frame to the map. In frame-to-map registration, we warp the frame to the map by this transformation to compensate for scale and rotation difference and then perform an area based matching using mutual information to find correspondences between this warped frame and the map. These correspondences together with the correspondences in previous frames could be regarded as correspondences between the partial local mosaic and the map. By registering the partial local mosaic to the map, we derive a transformation from the frame to the map. With this two-step registration, the errors between each consecutive frames are not accumulated. We then extend our approach to synchronize multiple image sequences by tracking moving objects in each image sequence, and aligning the frames based on the object's coordinates in the reference image. Yuping Lin, Gérard G. Medioni |
CVPR | 2 |
| 2007 | Moving Object Detection on a Runway Prior to Landing Using an Onboard Infrared CameraabstractDetermining the status of a runway prior to landing is essential for any aircraft, whether manned or unmanned. In this paper, we present a method that can detect moving objects on the runway from an onboard infrared camera prior to the landing phase. Since the runway is a planar surface, we first locally stabilize the sequence to automatically selected reference frames using feature points in the neighborhood of the runway. Next, we normalize the stabilized sequence to compensate for the global intensity variation caused by the gain control of the infrared camera. We then create a background model to learn an appearance model of the runway. Finally, we identify moving objects by comparing the image sequence with the background model. We have tested our system with both synthetic and real world data and show that it can detect distant moving objects on the runway. We also provide a quantitative analysis of the performance with respect to variations in size, direction and speed of the target. Cheng-Hua Pai, Yuping Lin, Gérard G. Medioni, Ray Rida Hamza |
CVPR | 3 |
| 2007 | Multiple Target Tracking Using Spatio-Temporal Markov Chain Monte Carlo Data AssociationabstractWe propose a framework for general multiple target tracking, where the input is a set of candidate regions in each frame, as obtained from a state of the art background learning, and the goal is to recover trajectories of targets over time from noisy observations. Due to occlusions by targets and static objects, noisy segmentation and false alarms, one foreground region may not correspond to one target faithfully. Therefore the one-to-one assumption used in most data association algorithm is not always satisfied. Our method overcomes the one-to-one assumption by formulating the visual tracking problem in terms of finding the best spatial and temporal association of observations, which maximizes the consistency of both motion and appearance of trajectories. To avoid enumerating all possible solutions, we take a data driven Markov chain Monte Carlo (DD-MCMC) approach to sample the solution space efficiently. The sampling is driven by an informed proposal scheme controlled by a joint probability model combining motion and appearance. To make sure the Markov chain to converge to a desired distribution, we propose an automatic approach to determine the parameters in the target distribution. Comparative experiments with quantitative evaluations are provided. Gérard G. Medioni, Isaac Cohen |
CVPR | 2 |
| 2007 | Inferring 3D Volumetric Shape of Both Moving Objects and Static Background Observed by a Moving CameraabstractWe present a novel approach to inferring 3D volumetric shape of both moving objects and static background from video sequences shot by a moving camera, with the assumption that the objects move rigidly on a ground plane. The 3D scene is divided into a set of volume elements, termed as voxels, organized in an adaptive octree structure. Each voxel is assigned a label at each time instant, either as empty, or belonging to background structure, or a moving object. The task of shape inference is then formulated as assigning each voxel a dynamic label which minimizes photo and motion variance between voxels and the original sequence. We propose a three-step voxel labeling method based on a robust photo-motion variance measure. First, a sparse set of surface points are utilized to initialize a subset of voxels. Then, a deterministic voxel coloring scheme carves away the voxels with large variance. Finally, the labeling results are refined by a graph cuts based optimization method to enforce global smoothness. Experimental results on both indoor and outdoor sequences demonstrate the effectiveness and robustness of our method. Chang Yuan, Gérard G. Medioni |
CVPR | 2 |
| 2007 | 3-D Metric Reconstruction and Registration of Images of Near-planar SurfacesabstractIn this study, we address the problem of 3-D dense metric reconstruction and registration from multiple images, given that the observed surface is nearly planar. This is difficult, as classical methods work well only if the scene is truly planar (mosaicing) or the scene has certain significant depth variations (classical Structure-from- Motion (SfM)). One domain in which this problem occurs is image analysis of the retinal fundus. Our approach is to first assume planarity, and perform 2-D global registration. A first bundle adjustment is applied to find the camera positions in metric space. We then select two images and compute the epipolar geometry between them using plane+parallax approach. These images are matched to generate a dense disparity map using mutual information. A second bundle adjustment is applied to transform the disparity map into a dense metric depth map, fixing the 2 camera positions. A third bundle adjustment is performed to refine both camera positions and a 3-D structure. All images are back-projected to the 3-D structure for the final registration. The entire process is fully automatic. In addition, a clear definition of "near-planarity " is provided. 3-D reconstruction is shown visually. The method is general, and can be applied to other domains, as shown in the experiments. Tae Eun Choe, Gérard G. Medioni |
ICCV | 2 |
| 2007 | Robust Real-Time Vision for a Personal Service Robot in a Home Visual Sensor NetworkabstractWe address issues dealing with visual perception for a personal service robot in the intelligent home environment. We identify key visual functionalities necessary for the robot to perform its activities. They include people detection and identification, gesture recognition, and self-localization. We propose an efficient and reliable framework to organize and coordinate the vision sensor nodes: fixed cameras mounted on walls, and earnera(s) on the mobile robot. We propose solutions to the different vision tasks, and present our implementation within this framework, validated with experimental results. Kwangsu Kim, Gérard G. Medioni |
RO-MAN | 2 |
| 2007 | Planar Patch based 3D Environment Modeling with Stereo CameraabstractWe present two robust and novel algorithms to model a 3D environment using both intensity and range data provided by an off-the-shelf stereo camera. The main issue we need to address is that the output of the stereo system is both sparse and noisy. To overcome this limitation,, we detect planar patches in the environment by region segmentation in 2D and plane extraction in 3D. The extracted planar patches are used not only to represent the workspace, but also to fill holes in range data. We also suggest a new planar patch based scan matching algorithm to register multiple views, and to incrementally augment the description of the 3D workspace in a sequence of scenes. Experimental results on real data show that planar patch segmentation and 3D scene registration for environment modeling can be robustly achieved by the proposed approaches. Gérard G. Medioni, Sukhan Lee 0001 |
RO-MAN | 2 |
| 2007 | Efficient Articulated Model Fitting on a Single Image or a SequenceabstractModels that can efficiently, compactly, and semantically represent potential users are important tools for human-robot interaction applications. We model a person as a projection of a generic 3D articulated model and propose a method to estimate its joint positions from image data in an optimization framework. This is done by constructing a function that grades a configuration of joints according to how well it matches the underlying image and model based priors. We then search for local optimum in this space both efficiently and exhaustively by assembling partial configurations in a bottom-up manner. Working from the leaves of the tree to its root, we maintain a list of locally optimal, yet sufficiently distinct candidate configurations for the body pose. We then adapt this algorithm for use on a sequence of images to make it even more efficient by considering configurations that are near their position in the previous frame. This way, the number of partial configurations generated and evaluated significantly reduces. These algorithms are validated on real image data. Matheen Siddiqui, Gérard G. Medioni |
RO-MAN | 2 |
| 2007 | Map-Enhanced UAV Image Sequence RegistrationabstractRegistering consecutive images from an airborne sensor into a mosaic is an essential tool for image analysts. Strictly local methods tend to accumulate errors, resulting in distortion. We propose here to use a reference image (such as a high resolution map image) to overcome this limitation. In our approach, we register a frame in an image sequence to the map using both frame-to-frame registration and frame-to-map registration iteratively. In frame-to-frame registration, a frame is registered to its previous frame. With its previous frame been registered to the map in the previous iteration, we can derive an estimated transformation from the frame to the map. In frame-to-map registration, we warp the frame to the map by this transformation to compensate for scale and rotation difference and then perform an area based matching using mutual information to find correspondences between this warped frame and the map. From these correspondences, we derive a transformation that further registers the warped frame to the map. With this two-step registration, the errors between each consecutive frames are not accumulated. We present results on real image sequences from a hot air balloon Yuping Lin, Gérard G. Medioni |
WACV | 3 |
| 2007 | Robust real-time vision for a personal service robot
Gérard G. Medioni, Alexandre R. J. François, Matheen Siddiqui, Kwangsu Kim, Hosub Yoon |
Comput. Vis. Image Underst. | 1 |
| 2007 | Registration of 3D Points Using Geometric Algebra and Tensor Voting
Leo Reyes, Gérard G. Medioni, Eduardo Bayro-Corrochano |
Int. J. Comput. Vis. | 2 |
| 2007 | Text segmentation in color images using tensor voting
JaeGuyn Lim, Gérard G. Medioni |
Image Vis. Comput. | 3 |
| 2007 | Detecting Motion Regions in the Presence of a Strong Parallax from a Moving Camera by Multiview Geometric ConstraintsabstractWe present a method for detecting motion regions in video sequences observed by a moving camera, in the presence of strong parallax due to static 3D structures. The proposed method classifies each image pixel into planar background, parallax or motion regions by sequentially applying 2D planar homographies, the epipolar constraint and a novel geometric constraint, called "structure consistency constraint". The structure consistency constraint, as the main contribution of this paper, is derived from the relative camera poses among three frames and implemented within the "Plane+Parallax" framework. Unlike previous planar-parallax constraints, the proposed constraint does not require the reference plane to be constant across multiple views. It directly measures the inconsistency between the projective structures from the same point under camera motion and reference plane change. The structure consistency constraint is capable of detecting moving objects followed by a moving camera in the same direction, a so called degenerate configuration where the epipolar constraint fails. We demonstrate the effectiveness and robustness of our method with experimental results on real-world video sequences. Chang Yuan, Gérard G. Medioni, Jinman Kang, Isaac Cohen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | 3-D Shape Reconstruction of Retinal FundusabstractWe present a method for 3-D shape reconstruction of retinal fundus from fluorescein images. Our method extracts the location of vessels’ bifurcation as a reliable feature for estimating the epipolar geometry using a plane-and-parallax approach. The proposed solution robustly estimates the fundamental matrix for nearly planar surfaces, such as the retinal fundus. We propose the use of mutual information criteria for accurate estimation of the disparity maps, where the matched Y-features are used for automatically estimating the bounds of the disparities range space. Our experimental results validate the proposed method on sets of difficult fluorescein image pairs. Tae Eun Choe, Isaac Cohen, Gérard G. Medioni |
CVPR (2) | 3 |
| 2006 | 3D Reconstruction of Background and Objects Moving on Ground Plane Viewed from a Moving CameraabstractWe present a novel method to obtain a 3D Euclidean reconstruction of both the background and moving objects in a video sequence. We assume that, multiple objects are moving rigidly on a ground plane observed by a moving camera. The video sequence is first segmented into static background and motion blobs by a homography-based motion segmentation method. Then classical "Structure from Motion" (SfM) techniques are applied to obtain a Euclidean reconstruction of the static background. The motion blob corresponding to each moving object is treated as if there were a static object observed by a hypothetical moving camera, called a "virtual camera". This virtual camera shares the same intrinsic parameters with the real camera but moves differently due to object motion. The same SfM techniques are applied to estimate the 3D shape of each moving object and the pose of the virtual camera. We show that the unknown scale of moving objects can be approximately determined by the ground plane, which is a key contribution of this paper. Another key contribution is that we prove that the 3D motion of moving objects can be solved from the virtual camera motion with a linear constraint imposed on the object translation. In our approach, a planartranslation constraint is formulated: "the 3D instantaneous translation of moving objects must be parallel to the ground plane". Results on real-world video sequences demonstrate the effectiveness and robustness of our approach. Chang Yuan, Gérard G. Medioni |
CVPR (2) | 2 |
| 2006 | Evaluation of 3-D Shape Reconstruction of Retinal Fundus
Tae Eun Choe, Isaac Cohen, Gérard G. Medioni, Alexander C. Walsh, Srinivas R. Sadda |
MICCAI (1) | 3 |
| 2006 | Stereo Using Monocular Cues within the Tensor Voting FrameworkabstractWe address the fundamental problem of matching in two static images. The remaining challenges are related to occlusion and lack of texture. Our approach addresses these difficulties within a perceptual organization framework, considering both binocular and monocular cues. Initially, matching candidates for all pixels are generated by a combination of matching techniques. The matching candidates are then embedded in disparity space, where perceptual organization takes place in 3D neighborhoods and, thus, does not suffer from problems associated with scanline or image neighborhoods. The assumption is that correct matches produce salient, coherent surfaces, while wrong ones do not. Matching candidates that are consistent with the surfaces are kept and grouped into smooth layers. Thus, we achieve surface segmentation based on geometric and not photometric properties. Surface overextensions, which are due to occlusion, can be corrected by removing matches whose projections are not consistent in color with their neighbors of the same surface in both images. Finally, the projections of the refined surfaces on both images are used to obtain disparity hypotheses for unmatched pixels. The final disparities are selected after a second tensor voting stage, during which information is propagated from more reliable pixels to less reliable ones. We present results on widely used benchmark stereo pairs. Philippos Mordohai, Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2005 | Detection and Tracking of Moving Objects from a Moving Platform in Presence of Strong ParallaxabstractWe present a novel approach to detect and track independently moving regions in a 3D scene observed by a moving camera in the presence of strong parallax. Detected moving pixels are classified into independently moving regions or parallax regions by analyzing two geometric constraints: the commonly used epipolar constraint, and the structure consistency constraint. The second constraint is implemented within a "plane+parallax" framework and represented by a bilinear relationship which relates the image points to their relative depths. This newly derived relationship is related to trilinear tensor, but can be enforced into more than three frames. It does not assume a constant reference plane in the scene and therefore eliminates the need for manual selection of reference plane. Then, a robust parallax filtering scheme is proposed to accumulate the geometric constraint errors within a sliding window and estimate a likelihood map for pixel classification. The likelihood map is integrated into our tracking framework based on the spatio-temporal joint probability data association filter (JPDAF). This tracking approach infers the trajectory and bounding box of the moving objects by searching the optimal path with maximum joint probability within a fixed size of buffer. We demonstrate the performance of the proposed approach on real video sequences where parallax effects are significant. Jinman Kang, Isaac Cohen, Gérard G. Medioni, Chang Yuan |
ICCV | 3 |
| 2005 | Unsupervised Dimensionality Estimation and Manifold Learning in high-dimensional Spaces by Tensor Voting
Philippos Mordohai, Gérard G. Medioni |
IJCAI | 2 |
| 2005 | A Voting-Based Computational Framework for Visual Motion Analysis and InterpretationabstractMost approaches for motion analysis and interpretation rely on restrictive parametric models and involve iterative methods which depend heavily on initial conditions and are subject to instability. Further difficulties are encountered in image regions where motion is not smooth-typically around motion boundaries. This work addresses the problem of visual motion analysis and interpretation by formulating it as an inference of motion layers from a noisy and possibly sparse point set in a 4D space. The core of the method is based on a layered 4D representation of data and a voting scheme for affinity propagation. The inherent problem caused by the ambiguity of 2D to 3D interpretation is usually handled by adding additional constraints, such as rigidity. However, enforcing such a global constraint has been problematic in the combined presence of noise and multiple independent motions. By decoupling the processes of matching, outlier rejection, segmentation, and interpretation, we extract accurate motion layers based on the smoothness of image motion, then locally enforce rigidity for each layer in order to infer its 3D structure and motion. The proposed framework is noniterative and consistently handles both smooth moving regions and motion discontinuities without using any prior knowledge of the motion model. Mircea Nicolescu, Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2005 | Parametric reconstruction of generalized cylinders from limb edgesabstractThe three-dimensional (3-D) reconstruction of generalized cylinders (GCs) is an important research field in computer vision. One of the main difficulties is that some contour features in images cannot be reconstructed by traditional stereovision because they do not correspond to reflectance discontinuities of surface in space. In this paper, we present a novel, parametric approach for the 3-D reconstruction of circular generalized cylinders (CGCs) only from the limb edges of CGCs in two images. Instead of exploiting the invariant and quasiinvariant properties of some specific subclasses of GCs in projections, our reconstruction is achieved by some general assumptions on GCs, and can, therefore, be applied to a broader subclass of GCs. In order to improve robustness, we perform the extraction and labeling of the limb edge interactively, and estimate the epipolar geometry between two images by an optimal algorithm. Then, for different types of GCs, three kinds of symmetries (parallel symmetry, skew symmetry, and local smooth symmetry) are employed to compute the symmetry of limb edges. The surface points corresponding to limb edges in images are reconstructed by integrating the recovered epipolar geometry and the properties induced from the assumptions that we make on the GCs. Finally, a homography-based method is exploited to further refine the 3-D description of the GC with a coplanar curved axis. Chunhong Pan, Hongping Yan, Gérard G. Medioni, Songde Ma |
IEEE Trans. Image Process. | 3 |
| 2004 | Stereo Using Monocular Cues within the Tensor Voting Framework
Philippos Mordohai, Gérard G. Medioni |
ECCV (4) | 2 |
| 2004 | A robust and non-iterative estimation method of multiple 2d motionsabstractThis paper focuses on extracting 2D parametric motion regions from uncalibrated images. Our approach simultaneously infers and detects multiple image regions characterized by 2D motions, affine or homography transformations, from noisy initial matches. This approach is based on: (1) a parametric method to detect and extract 2D affine or homography motion regions; (2) the representation of the matching points in decoupled joint image spaces; (3) the characterization of the property associated with affine transformation in the defined spaces; (4) a non-iterative process to extract multiple 2D motions simultaneously based on tensor-voting; (5) local affine to global homography estimation; (6) region refinement based on a hybrid property: motion and color homogeneity. The robustness of the approach is demonstrated with several results. Eun-Young Kang 0002, Isaac Cohen, Gérard G. Medioni |
ICIP | 3 |
| 2004 | A layer extraction system based on dominant motion estimation and global registrationabstractWe describe a system that extracts layers from a video sequence based on a method estimating and stabilizing the dominant motion. Our approach performs pair-wise registration of images parameterized by 2D affine or projective transformation as the first step. This pair-wise registration is based on: (1) hierarchical parameter estimation and refinement; (2) feature-matching; (3) FFT (fast Fourier transformation)-based global matching; and (4) RANSAC-based parameter estimation. A global registration is then performed among overlapping multi-frames. The global registration is based on inferring the topology (i.e. spatio-temporal relation among frames) and the characterization of the likelihood within overlapping areas. As the last step, we extract the background layer that consists of static components of the video stream after motion estimation. Our background extraction is based on deriving a pixel-wise color distribution in time and provides a basis for a compact description of a video. The presented approach is illustrated by a set of challenging examples. Eun-Young Kang 0002, Isaac Cohen, Gérard G. Medioni |
ICME | 3 |
| 2004 | Simultaneous Two-View Epipolar Geometry Estimation and Motion Segmentation by 4D Tensor VotingabstractWe address the problem of simultaneous two-view epipolar geometry estimation and motion segmentation from nonstatic scenes. Given a set of noisy image pairs containing matches of n objects, we propose an unconventional, efficient, and robust method, 4D tensor voting, for estimating the unknown n epipolar geometries, and segmenting the static and motion matching pairs into n independent motions. By considering the 4D isotropic and orthogonal joint image space, only two tensor voting passes are needed, and a very high noise to signal ratio (up to five) can be tolerated. Epipolar geometries corresponding to multiple, rigid motions are extracted in succession. Only two uncalibrated frames are needed, and no simplifying assumption (such as affine camera model or homographic model between images) other than the pin-hole camera model is made. Our novel approach consists of propagating a local geometric smoothness constraint in the 4D joint image space, followed by global consistency enforcement for extracting the fundamental matrices corresponding to independent motions. We have performed extensive experiments to compare our method with some representative algorithms to show that better performance on nonstatic scenes are achieved. Results on challenging data sets are presented. Wai-Shun Tong, Chi-Keung Tang, Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2004 | First Order Augmentation to Tensor Voting for Boundary Inference and Multiscale Analysis in 3DabstractMost computer vision applications require the reliable detection of boundaries. In the presence of outliers, missing data, orientation discontinuities, and occlusion, this problem is particularly challenging. We propose to address it by complementing the tensor voting framework, which was limited to second order properties, with first order representation and voting. First order voting fields and a mechanism to vote for 3D surface and volume boundaries and curve endpoints in 3D are defined. Boundary inference is also useful for a second difficult problem in grouping, namely, automatic scale selection. We propose an algorithm that automatically infers the smallest scale that can preserve the finest details. Our algorithm then proceeds with progressively larger scales to ensure continuity where it has not been achieved. Therefore, the proposed approach does not oversmooth features or delay the handling of boundaries and discontinuities until model misfit occurs. The interaction of smooth features, boundaries, and outliers is accommodated by the unified representation, making possible the perceptual organization of data in curves, surfaces, volumes, and their boundaries simultaneously. We present results on a variety of data sets to show the efficacy of the improved formalism. Wai-Shun Tong, Chi-Keung Tang, Philippos Mordohai, Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2003 | Continuous Tracking Within and Across Camera StreamsabstractThis paper presents a new approach for continuous tracking of moving objects observed by multiple, heterogeneous cameras. Our approach simultaneously processes video streams from stationary and pan-tilt-zoom cameras. The detection of moving objects from moving camera streams is performed by defining an adaptive background model that takes into account the camera motion approximated by an affine transformation. We address the tracking problem by separately modeling motion and appearance of the moving objects using two probabilistic models. For the appearance model, multiple color distribution components are proposed for ensuring a more detailed description of the object being tracked. The motion model is obtained using a Kalman filter (KF) process, which predicts the position of the moving object. The tracking is performed by the maximization of a joint probability model. The novelty of our approach consists in modeling the multiple trajectories observed by the moving and stationary cameras in the same KF framework. It allows deriving a more accurate motion measurement for objects simultaneously viewed by the two cameras and an automatic handling of occlusions, errors in the detection and camera handoff. We demonstrate the performances of the system on several video surveillance sequences. Jinman Kang, Isaac Cohen, Gérard G. Medioni |
CVPR (1) | 3 |
| 2003 | Motion Segmentation with Accurate Boundaries - A Tensor Voting ApproachabstractProducing an accurate motion flow field is very difficult at motion boundaries. We present a noniterative approach for segmentation from image motion, based on two voting processes, in different dimensional spaces. By expressing the motion layers as surfaces in a 4D (four-dimensional) space, a voting process is first used to enforce the smoothness of motion and determine an estimation of pixel velocities, motion regions and boundaries. The boundary estimation is then combined with intensity information from the original images in order to locally define a boundary tensor field. The correct boundary is inferred by a 2D (two-dimensional) voting process within this field that enforces the smoothness of boundaries. Finally, correct velocities are computed for the pixels near boundaries, as they are reassigned to different regions. We demonstrate our contribution by analyzing several image sequences, containing multiple types of motion. Mircea Nicolescu, Gérard G. Medioni |
CVPR (1) | 2 |
| 2003 | Mirror symmetry => 2-view stereo geometry
Alexandre R. J. François, Gérard G. Medioni, Roman Waupotitsch |
Image Vis. Comput. | 2 |
| 2003 | Layered 4D Representation and Voting for Grouping from MotionabstractWe address the problem of perceptual grouping from motion cues by formulating it as a motion layers inference from a sparse and noisy point set in a 4D space. Our approach is based on a layered 4D representation of data, and a voting scheme for token communication, within a tensor voting computational framework. Given two sparse sets of point tokens, the image position and potential velocity of each token are encoded into a 4D tensor. By enforcing the smoothness of motion through a voting process, the correct velocity is selected for each input point as the most salient token. An additional dense voting step allows for the inference of a dense representation in terms of pixel velocities, motion regions, and boundaries. Using a 4D space for this tensor voting approach is essential since it allows for a spatial separation of the points according to both their velocities and image coordinates. Unlike most other methods that optimize certain objective functions, our approach is noniterative and, therefore, does not suffer from local optima or poor convergence problems. We demonstrate our method with synthetic and real images, by analyzing several difficult cases-opaque and transparent motion, rigid and nonrigid motion, curves and surfaces in motion. Mircea Nicolescu, Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2002 | Perceptual Grouping from Motion Cues Using Tensor Voting in 4-D
Mircea Nicolescu, Gérard G. Medioni |
ECCV (3) | 2 |
| 2002 | Practical algorithms for stratified structure-from-motion
George Q. Chen, Gérard G. Medioni |
Image Vis. Comput. | 2 |
| 2002 | Inference of Segmented Overapping Surfaces from Binocular StereoabstractPresents an integrated approach to the derivation of scene descriptions from a pair of stereo images, where the steps of feature correspondence and surface reconstruction are addressed within the same framework. Special attention is given to the development of a methodology with general applicability. In order to handle the issues of noise, lack of image features, surface discontinuities and regions that are visible in one image only, we adopt a tensor representation for the data and introduce a robust computational technique called tensor voting for information propagation. The key contributions of this paper are twofold. First, we introduce "saliency" instead of correlation scores as the criterion to determine the correctness of matches and the integration of feature matching and structure extraction. Second, our tensor representation and voting as a tool enables us to perform the complex computations associated with the formulation of the stereo problem in 3D at a reasonable computational cost. We illustrate the steps on an example, then provide results on both random dot stereograms and real stereo pairs, all processed with the same parameter set. Mi-Suen Lee, Gérard G. Medioni, Philippos Mordohai |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2002 | Curvature-Augmented Tensor Voting for Shape Inference from Noisy 3D DataabstractImproves the basic tensor voting formalism to infer the sign and direction of principal curvatures at each input site from noisy 3D data. Unlike most previous approaches, no local surface fitting, partial derivative computation, nor oriented normal vector recovery is performed in our method. These approaches are known to be noise-sensitive, since accurate partial derivative information is often required, which is usually unavailable from real data. Also, unlike approaches that detect signs of Gaussian curvature, we can handle points with zero Gaussian curvature uniformly, without first localizing them in a separate process. The tensor-voting curvature estimation is non-iterative, does not require initialization, and is robust to a considerable amount of outlier noise, as its effect is reduced by collecting a large number of tensor votes. Qualitative and quantitative results on synthetic and real complex data are presented. Chi-Keung Tang, Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2001 | First Order Tensor Voting, and Application to 3-D Scale AnalysisabstractMany computer vision systems depend on reliable detection of 3D boundaries and regions in order to proceed. In the presence of outliers, missing data and orientation discontinuities due to occlusion, it is difficult to detect boundaries and interpolate data without over-smoothing important feature curves. The authora address these problems by incorporating first order tensor information into the tensor voting formalism, which is second-order based. To propagate an adaptive smoothness constraint at a preferred orientation non-iteratively, we vote for a first order tensor (or vector) to capture polarity and orientation information. To integrate first and second order tensors, we propose an algorithm for inferring the proper scale based on the continuity constraint, and preserving the finest details. Given a noisy 3D point set, the new and improved formalism can better localize boundary curves and orientation discontinuities. Unlike many approaches that over-smooth features, or delay the handling of boundaries and discontinuities until model misfit occurs, the interaction of smooth features, boundaries, discontinuities, outliers are encoded at the representation level. We present results from a variety of datasets to show the efficacy of the improved formalism. Wai-Shun Tong, Chi-Keung Tang, Gérard G. Medioni |
CVPR (1) | 3 |
| 2001 | Epipolar Geometry Estimation for Non-Static Scenes by 4D Tensor VotingabstractIn the presence of false matches and moving objects, image registration is challenging, as outlier rejection, matching and registration become interdependent. We present an efficient and robust method, 4D tensor voting to estimate epipolar geometries for non-static scenes, and identify matching points due to salient and independent motions. Unlike other optimization techniques, data communication in 4D tensor voting does not involve any iterative search. Thus, initialization, local optimum, convergence, and dimensionality of parameter space are not problematic. Like the 8D counterpart, the only assumption we make is the pinhole camera model. Two advancements are made in this work. First, we reduce the dimensionality, and the 4D joint image space is an isotropic and orthogonal one, validating the general assumptions of tensor voting. This improvement is evidenced by the facts that only two passes are needed, and that 4D tensor voting can tolerate an even larger noise/signal ratio (up to a ratio of five). Second, instead of discarding motion pixels as outliers, we successively extract the epipolar geometries contributed by the static background and by the matching points due to salient motions. Only two frames are needed, and no simplifying assumption (such as affine camera model or homographic model between images) is made. Our 4D algorithm consists of two stages: local continuity constraint propagation to remove outliers, and global consistency checking to localize a 4D topological point cone. Results on challenging datasets are presented. Wai-Shun Tong, Chi-Keung Tang, Gérard G. Medioni |
CVPR (1) | 3 |
| 2001 | Interactive 3D model extraction from a single image
Alexandre R. J. François, Gérard G. Medioni |
Image Vis. Comput. | 2 |
| 2001 | Crater detection for autonomous landing on asteroids
B. Leroy, Gérard G. Medioni, Larry H. Matthies |
Image Vis. Comput. | 2 |
| 2001 | Event Detection and Analysis from Video StreamsabstractWe present a system which takes as input a video stream obtained from an airborne moving platform and produces an analysis of the behavior of the moving objects in the scene. To achieve this functionality, our system relies on two modular blocks. The first one detects and tracks moving regions in the sequence. It uses a set of features at multiple scales to stabilize the image sequence, that is, to compensate for the motion of the observer, then extracts regions with residual motion and uses an attribute graph representation to infer their trajectories. The second module takes as input these trajectories, together with user-provided information in the form of geospatial context and goal context to instantiate likely scenarios. We present details of the system, together with results on a number of real video sequences and also provide a quantitative analysis of the results. Gérard G. Medioni, Isaac Cohen, François Brémond, Somboon Hongeng, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2001 | N-Dimensional Tensor Voting and Application to Epipolar Geometry EstimationabstractWe address the problem of epipolar geometry estimation by formulating it as one of hyperplane inference from a sparse and noisy point set in an 8D space. Given a set of noisy point correspondences in two images of a static scene without correspondences, even in the presence of moving objects, our method extracts good matches and rejects outliers. The methodology is novel and unconventional, since, unlike most other methods optimizing certain scalar, objective functions, our approach does not involve initialization or any iterative search in the parameter space. Therefore, it is free of the problem of local optima or poor convergence. Further, since no search is involved, it is unnecessary to impose simplifying assumption to the scene being analyzed for reducing the search complexity. Subject to the general epipolar constraint only, we detect wrong matches by a computation scheme, 8D tensor voting, which is an instance of the more general N-dimensional tensor voting framework. In essence, the input set of matches is first transformed into a sparse 8D point set. Dense, 8D tensor kernels are then used to vote for the most salient hyperplane that captures all inliers inherent in the input. With this filtered set of matches, the normalized eight-point algorithm can be used to estimate the fundamental matrix accurately. By making use of efficient data structure and locality, our method is both time and space efficient despite the higher dimensionality. We demonstrate the general usefulness of our method using example image pairs for aerial image analysis, with widely different views, and from nonstatic 3D scenes. Each example contains a considerable number of wrong matches. Chi-Keung Tang, Gérard G. Medioni, Mi-Suen Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2000 | Tracking Segmented Objects Using Tensor VotingabstractThe paper presents a new approach to track objects in motion when observed by a fixed camera, with severe occlusions, merging/splitting objects and defects in the detection. We first detect regions corresponding to moving objects in each frame, then try to establish their trajectory. We propose to implement the temporal continuity constraint efficiently, and apply it to tracking problems in realistic scenarios. The method is based on a spatiotemporal (2D+t) representation of the moving regions, and uses the tensor voting methodology to enforce smoothness in space and table of the tracked objects. Although other characteristics may be considered, only the connected components of the moving regions are used, without further assumptions about the object being tracked. We demonstrate the performance of the system on several real sequences. Pierre Kornprobst, Gérard G. Medioni |
CVPR | 2 |
| 2000 | A Graph-Based Global Registration for 2D MosaicsabstractWe describe a graph-based global registration method for creating 2D mosaic images. When multi-frames overlap in space, global registration is necessary to minimize the accumulated registration errors. We use a graph to represent the temporal and spatial connectivity and show that global registration can be obtained through the search for an optimal path in the constructed graph. The definition of an adequate objective function characterizing the global registration provides a direct manipulation of the graph. The framework presented here allows the automatic construction of the graph, and the construction of a consistent mosaic from a collection of frames using projective transformations. Eun-Young Kang 0002, Isaac Cohen, Gérard G. Medioni |
ICPR | 3 |
| 2000 | A 2D + t Tensor Voting Based Approach for TrackingabstractPresents an approach to tracking objects in motion when observed by a fixed camera, with severe occlusions, merging/splitting objects and defects in the detection. We first detect regions corresponding to moving objects in each frame, then try to establish their trajectory. The method is based on a spatiotemporal (2D+t) representation of the moving regions, and uses the tensor voting methodology to enforce smoothness in space and time of the tracked regions. We demonstrate the performance of the system on several real sequences. Pierre Kornprobst, Gérard G. Medioni |
ICPR | 2 |
| 2000 | 3-D Structures for Generic Object RecognitionabstractWe discuss the issues and challenges of generic object recognition. We argue that high-level, volumetric part-based descriptions are essential in the process of recognizing objects that might never have been observed before, and for which no exact geometric model is available. We discuss the representation scheme and its relationships to the three main tasks to solve: extracting descriptions from real images, under a wide variety of viewing conditions; learning new objects by storing their description in a database; and recognizing objects by matching their description to that of similar previously observed objects. Gérard G. Medioni, Alexandre R. J. François |
ICPR | 1 |
| 2000 | GlobeAll: Panoramic Video for an Intelligent RoomabstractChoosing the appropriate type of video input is an important issue for any vision-based system and the right decision must take into account the specific requirements of the intended application. In the context of intelligent room systems, we establish several qualitative criteria to evaluate the video input component and we use them to compare three current solutions: mobile pan-tilt-zoom cameras, wide-angle lens cameras and electronic pan-tilt-zoom cameras. We show that electronic pan-tilt-zoom systems best satisfy our criteria. To support this claim, we present GlobeAll, a modular four-component prototype for a vision-based intelligent room: a video input component that uses an electronic pan-tilt-zoom camera array, a background learning and foreground extraction component, a tracking component and an interpretation component. Mircea Nicolescu, Gérard G. Medioni |
ICPR | 2 |
| 2000 | A modular middleware flow scheduling framework (poster session)abstractImmersive, interactive applications require on-line processing and mixing of multimedia data. In order to realize the Immersipresence vision, we propose a generic, extensible, modular multimedia system software architecture. We describe here the Flow Scheduling Framework (FSF), that constitutes the core of its middleware layer. The FSF is an extensible set of classes that provide basic synchronization functionality and composition mechanisms to develop data-stream processing components. In this dataflow approach, applications are implemented by specifying data streams and their path through processing nodes, where they can undergo various manipulations. We describe the details of the FSF data and processing model that supports stream synchronization in a concurrent processing framework. We illustrate the FSF concepts with a real-time video stream processing application. Alexandre R. J. François, Gérard G. Medioni |
ACM Multimedia | 2 |
| 1999 | A Volumetric Stereo Matching Method: Application to Image-Based ModelingabstractWe formulate stereo matching as an extremal surface extraction problem. This is made possible by embedding the disparity surface inside a volume where the surface is composed of voxels with locally maximal similarity values. This formulation naturally implements the coherence principle, and allows us to incorporate most known global constraints. Time efficiency is achieved by executing the algorithm in a coarse-to-fine fashion, and only populating the full volume at the coarsest level. To make the system more practical, we present a rectification algorithm based on the fundamental matrix, avoiding full camera calibration. We present results on standard stereo pairs, and on our own data set. The results are qualitatively evaluated in terms of both the generated disparity maps and the 3-D models. Qian Chen 0022, Gérard G. Medioni |
CVPR | 2 |
| 1999 | Efficient Iterative Solution to M-View Projective Reconstruction ProblemabstractWe propose an efficient solution to the general M-view projective reconstruction problem, using matrix factorization and iterative least squares. The method can accept input with missing data, meaning that not all points are necessarily visible in all views. It runs much faster than the often-used non-linear minimization method, while preserving the accuracy of the latter. The key idea is to convert the minimization problem into a series of weighted least squares sub-problems with drastically reduced matrix sizes. Additionally, we show that good initial values can always be obtained. Experimental results on both synthetic and real data are presented. Potential applications are also demonstrated. Qian Chen 0022, Gérard G. Medioni |
CVPR | 2 |
| 1999 | Detecting and Tracking Moving Objects for Video SurveillanceabstractWe address the problem of detection and tracking of moving objects in a video stream obtained from a moving airborne platform. The proposed method relies on a graph representation of moving objects which allows to derive and maintain a dynamic template of each moving object by enforcing their temporal coherence. This inferred template along with the graph representation used in our approach allows us to characterize objects trajectories as an optimal path in a graph. The proposed tracker allows to deal with partial occlusions, stop and go motion in very challenging situations. We demonstrate results on a number of different real sequences. We then define an evaluation methodology to quantify our results and show how tracking overcome detection errors. Isaac Cohen, Gérard G. Medioni |
CVPR | 2 |
| 1999 | Accurate Motion Flow Estimation with DiscontinuitiesabstractWe address the problem of motion flow estimation for a scene with multiple moving objects, observed from a possibly moving camera. We take as input a (possibly sparse) noisy velocity field, as obtained from local matching, produce a set of motion boundaries, and identify pixels with different velocities in overlapping layers. For a fixed observer, these overlapping layers capture occlusion information. For a moving observer, further processing is required to segment independent objects and infer structure. Unlike previous approaches, which generate layers by iteratively fitting data to a set of predefined parameters, we instead find boundaries first, then infer regions and address occlusion overlap relationships. All computational steps use a common framework of tensors to represent velocity information, together with saliency (confidence), and uncertainty. Communication between sites is performed by convolution-like tensor voting. The scheme is non-iterative, and the only free parameter is the scale, related to neighborhood size. We illustrate the approach with results obtained from synthetic sequences and from real images. The quantitative results compare favorably with those of other methods, especially in the presence of occlusion. Lionel Gaucher, Gérard G. Medioni |
ICCV | 2 |
| 1999 | Robust Estimation of Curvature Information from Noisy 3D Data for Shape DescriptionabstractWe describe an effective and novel approach to infer sign and direction of principal curvatures at each input site from noisy 3D data. Unlike most previous approaches, no local surface fitting, partial derivative computation of any kind, nor oriented normal vector recovery is performed in our method. These approaches are noise-sensitive since accurate, local, partial derivative information is often required, which is usually unavailable from real data because of the unavoidable outlier noise inherent in many measurement phases. Also, we can handle points with zero Gaussian curvature uniformly (i.e., without the need to localize and handle them first as a separate process). Our approach is based on Tensor Voting, a unified, salient structure inference process. Both the sign and the direction of principal curvatures are inferred directly from the input. Each input is first transformed into a synthetic tensor A novel and robust approach based on tensor voting is proposed for curvature information estimation. With faithfully inferred curvature information, each input ellipsoid is aligned with curvature-based dense tensor kernels to produce a dense tensor field. Surfaces and crease curves are extracted from this dense field, by using an extremal feature extraction process. The computation is non-iterative, does not require initialization, and robust to considerable amounts of outlier noise as its effect is reduced by collecting a large number of tensor votes. qualitative and quantitative results on synthetic as well as real and complex data are presented. Chi-Keung Tang, Gérard G. Medioni |
ICCV | 2 |
| 1999 | Epipolar Geometry Estimation by Tensor Voting in 8DabstractWe present a novel, efficient, initialization free approach to the problem of epipolar geometry estimation, by formulating it as one of hyperplane inference from a sparse and noisy point set in an 8D space. Given a set of noisy point correspondences in two images as obtained from two views of a static scene without correspondences, even in the presence of moving objects, our method pulls out inlier matches while rejecting outliers. Unlike most methods which optimize certain objective function, our approach does not involve initialization or any search in the parameter space, and therefore is free of the problem of local optima or poor convergence. Since no search is involved, it is unnecessary to impose simplifying assumption (such as affine camera or local planar homography) to the scene being analyzed for reducing the search complexity. Subject to the general epipolar constraint only, we detect wrong matches by establishing salient "extremalities" via a naval approach, 8D Tensor Voting: the input set of matches is first transformed into a sparse and discrete 8D point set. Dense tensor kernels are then applied to vote for the most salient hyperplane (normal and intercept) that captures all inliers inherent in the input. With this filtered set of matches, the normalized Eight-point Algorithm suffices for the accurate estimation of the fundamental matrix. By using efficient data structure and locality, our method is both time and space efficient despite the higher dimensionality. We demonstrate the general usefulness of our method using example image pairs (i) for aerial image analysis, (ii) with widely different views, and (iii) from non-static 3D scenes (e.g. basketball game in an indoor stadium). Each example contains a considerable amount of wrong matches. Chi-Keung Tang, Gérard G. Medioni, Mi-Suen Lee |
ICCV | 2 |
| 1999 | Grouping ., -, ->, [formula], into Regions, Curves, and Junctions
Mi-Suen Lee, Gérard G. Medioni |
Comput. Vis. Image Underst. | 2 |
| 1999 | Simultaneous Surface Approximation and Segmentation of Complex Objects
Chia-Wei Liao, Gérard G. Medioni |
Comput. Vis. Image Underst. | 2 |
| 1998 | A Unified Framework for Salient Curves, Regions, and Junctions Inference
Mi-Suen Lee, Gérard G. Medioni |
ACCV (2) | 2 |
| 1998 | Inferring Segmented Surface Description from Stereo DataabstractWe present an integrated approach to the derivation of scene description from binocular stereo images. By inferring the scene description directly from local measurements of both point and line correspondences, we address both the stereo correspondence problem and the surface reconstruction problem simultaneously. We introduce a robust computational technique called tensor voting for the inference of scene description in terms of surfaces, junctions, and region boundaries. The methodology is grounded in two elements: tensor calculus for representation, and non-linear voting for data communication. By efficiently and effectively collecting and analyzing neighborhood information, we are able to handle the tasks of interpolation, discontinuity detection, and outlier identification simultaneously. The proposed method is non-iterative, robust to initialization and thresholding in the preprocessing stage, and the only critical free parameter is the size of the neighborhood. We illustrate the approach with results on a variety of images. Mi-Suen Lee, Gérard G. Medioni |
CVPR | 2 |
| 1998 | Integrated Surface, Curve and Junction Inference from Sparse 3-D Data SetsabstractWe are interested in descriptions of 3-D data sets, as obtained from stereo or a 3-D digitizer. We therefore consider as input a sparse set of points, possibly associated with orientation information. In this paper, we address the problem of inferring integrated high-level descriptions such as surfaces, curves, and junctions from a sparse point set. While the method described previously provides excellent results for smooth structures, it only detects discontinuities, but does not localize them. For precise localization, we propose a non-iterative cooperative algorithm in which surfaces, curves, and junctions work together: Initial estimates are computed based on previous results, where each point in the given sparse and possibly noisy point set is convolved with a predefined vector mask to produce dense saliency maps. These maps serve as input to our novel maximal surface and curve marching algorithms for initial surface and curve extraction. Refinement of initial estimates is achieved by hybrid voting using excitatory and inhibitory fields for inferring reliable and natural extension so that surface/curve and curve/junction discontinuities are preserved. Results on several synthetic as well as real data sets are presented. Chi-Keung Tang, Gérard G. Medioni |
ICCV | 2 |
| 1998 | Automatic, Accurate Surface Model Inference for Dental CAD/Cam
Chi-Keung Tang, Gérard G. Medioni, François Duret |
MICCAI | 2 |
| 1998 | Building human face models from two imagesabstractWe present a practical technique for building 3-D human face models from two photographs. Rather than using expensive 3-D scanners, we show that frontal face models can be faithfully reconstructed with unsophisticated digital cameras in a totally non-invasive setup. We propose a rectification algorithm based on the fundamental matrix by computing the dual of the point transformation matrix. The image matching problem is converted into a maximal surface extraction problem which is then solved efficiently. Finally, the Euclidean approximation is achieved with the help of a novel factorization method for perspective cameras. Two examples are presented. Qian Chen 0022, Gérard G. Medioni |
MMSP | 2 |
| 1998 | Extremal feature extraction from 3-D vector and noisy scalar fieldsabstractWe are interested in feature extraction from volume data in terms of coherent surfaces and 3D space curves. The input can be an inaccurate scalar or vector field, sampled densely or sparsely on a regular 3D grid, in which poor resolution and the presence of spurious noisy samples make traditional iso-surface techniques inappropriate. In this paper, we present a general-purpose methodology to extract surfaces or curves from a digital 3D potential vector field {(s,v~)}, in which each voxel holds a scalar s designating the strength and a vector v~ indicating the direction. For scalar, sparse or low-resolution data, we "vectorize" and "densify" the volume by tensor voting to produce dense vector fields that are suitable as input to our algorithms, the extremal surface and curve algorithms. Both algorithms extract, with sub-voxel precision, coherent features representing local extrema in the given vector field. These coherent features are a hole-free triangulation mesh (in the surface case), and a set of connected, oriented and non-intersecting polyline segments (in the curve case). We demonstrate the general usefulness of both extremal algorithms on a variety of real data by properly extracting their inherent extremal properties, such as (a) shock waves induced by abrupt velocity or direction changes in a flow field, (b) interacting vortex cores and vorticity lines in a velocity field, (c) crest-lines and ridges implicit in a digital terrain map, and (d) grooves, anatomical lines and complex surfaces from noisy dental data. Chi-Keung Tang, Gérard G. Medioni |
IEEE Visualization | 2 |
| 1998 | Volumetric description of dip solder joints from range dataabstractDip solder joints exhibit a variety of complex shapes, making inspection by traditional representation methods difficult. Traditional approaches use various features of 2D images to represent joint shape; however, correct 3-D representation is essential when analyzing joint shape for visual inspection. We present a method of using generalized cylinders (GCs) to obtain a 3-D representation of a solder joint from the joint's range image. This method is based on symmetry detection. B-spline contour representation allows analytical detection of a symmetry axis that is used as a basis for the GC axis. In addition, we present a B-spline fitting method that uses an initial line fitting procedure. The contour of the cross section is expressed as an ellipse using an ellipse-specified conic fitting method. Correct representations for over 53 test sample range images were obtained. Yuji Takagi, Gérard G. Medioni |
WACV | 2 |
| 1998 | Full Volumetric Descriptions From Three Intensity ImagesabstractWe address the problem of recovering high-level, volumetric and segmented (or part-based) descriptions of objects from intensity images. As input we use three closely spaced images of an object and recover descriptions based on generalized cylinders (GCs). We start by extracting a hierarchy of groups from contour images in the three views. Grouping is based on proximity, parallelism, and symmetry. The groups in the three views are matched and their contours are labeled as "true" edges. We then infer the GC axis, its cross-section, and the scaling function. The cross-section is recovered if seen in the images, else it is inferred using the visible surfaces and GC properties. We consider groups with true edges, limb edges, or a combination of both. The coarse volumetric descriptions obtained are refined to include surface details as seen in the intensity images. We demonstrate results on real images of moderately complex objects with texture and shadows. Parag Havaldar, Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1998 | Inference of Integrated Surface, Curve, and Junction Descriptions From Sparse 3D DataabstractWe address the problem of inferring integrated high-level descriptions such as surfaces, 3D curves, and junctions from a sparse point set. For precise localization, we propose a noniterative cooperative algorithm in which surfaces, curves, and junctions work together. Initial estimates are computed based on the work by Guy and Medioni (1997), where each point in the given sparse and possibly noisy point set is convolved with a predefined vector mask to produce dense saliency maps. These maps serve as input to our novel extremal surface and curve algorithms for initial surface and curve extraction. These initial features are refined and integrated by using excitatory and inhibitory fields. Consequently, intersecting surfaces (resp. curves) are fused precisely at their intersection curves (resp. junctions). Results on several synthetic as well as real data sets are presented. Chi-Keung Tang, Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1997 | 3DSketch: Modeling by Digitizing with a Smart 3D PenabstractWe describe a novel system "3DSketch" which demonstrates a two-handed 31) sketching paradigm for 3D modeling by casually digitizing an existing object.The conceptual model of the interface is based on the everyday experience in sketching with a pen on a piece of paper, but in our system, the user holds a 3D digitizing stylus as a 3D pen to sketch in 3D space.We call the pen a "smart pen", since in the Prototyper module, the userjust sketches a few strokes on the object, and immediately sees a 3D prototype made of clay lumps; then, in the Retiner module, when the user adds more random strokes over the object, the prototype surface automatically adapts to follow the pen, and the surface features (edges and comers) dig21 with the user specified ones, as if the pen tip applied magnetic attractive force to the prototype: in the autoTracer module, the user sketches over smooth regions, and the system performs intelligent reasoning to infer smooth surfaces, and also extract discontinuity edges and comers from the user's inaccurate and fragmented strokes.The internal surface representation is triangular splines (TriBezier, TriB, TYiNURElS), whose advantages in arbitrary triangulation and local subdivision make it flexible to model general surfaces. Gérard G. Medioni |
ACM Multimedia | 2 |
| 1997 | Image synthesis from a sparse set of viewsabstractThe authors present an image synthesis methodology and a system built around it. Given a sparse set of photographs taken from unknown viewpoints, the system generates images from new, different viewpoints with correct perspective, and handles occlusion. It achieves this without requiring any knowledge about the 3D structure of the scene nor the intrinsic camera parameters. The photo-realistic rendering process is polygon based and can be potentially implemented as real time texture mapping. The system is robust to noise by taking advantage of duplicate information from multiple views. They present results on several example scenes. Qian Chen 0022, Gérard G. Medioni |
IEEE Visualization | 2 |
| 1997 | Synthesizing Novel Views from Unregistered 2-D ImagesabstractSynthesizing the image of a 3‐D scene as it would be captured by a camera from an arbitrary viewpoint is a central problem in Computer Graphics. Given a complete 3‐D model, it is possible to render the scene from any viewpoint. The construction of models is a tedious task. Here, we propose to bypass the model construction phase altogether, and to generate images of a 3‐D scene from any novel viewpoint from prestored images. Unlike methods presented so far, we propose to completely avoid inferring and reasoning in 3‐D by using projective invariants. These invariants are derived from corresponding points in the prestored images. The correspondences between features are established off‐line in a semi‐automated way. It is then possible to generate wireframe animation in real time on a standard computing platform. Well understood texture mapping methods can be applied to the wireframes to realistically render new images from the prestored ones. The method proposed here should allow the integration of computer generated and real imagery for applications such as walkthroughs in realistic virtual environments. We illustrate our approach on synthetic and real indoor and outdoor images. Parag Havaldar, Mi-Suen Lee, Gérard G. Medioni |
Comput. Graph. Forum | 3 |
| 1997 | Inference of Surfaces, 3D Curves, and Junctions From Sparse, Noisy, 3D DataabstractWe address the problem of obtaining dense surface information from a sparse set of 3D data in the presence of spurious noise samples. The input can be in the form of points, or points with an associated tangent or normal, allowing both position and direction to be corrupted by noise. Most approaches treat the problem as an interpolation problem, which is solved by fitting a surface such as a membrane or thin plate to minimize some function. We argue that these physical constraints are not sufficient, and propose to impose additional perceptual constraints such as good continuity and "cosurfacity". These constraints allow us to not only infer surfaces, but also to detect surface orientation discontinuities, as well as junctions, all at the same time. The approach imposes no restriction on genus, number of discontinuities, number of objects, and is noniterative. The result is in the form of three dense saliency maps for surfaces, intersections between surfaces (i.e., 3D curves), and 3D junctions, respectively. These saliency maps are then used to guide a "marching" process to generate a description (e.g., a triangulated mesh) making information about surfaces, space curves, and 3D junctions explicit. The traditional marching process needs to be refined as the polarity of the surface orientation is not necessarily locally consistent. These three maps are currently not integrated, and this is the topic of our ongoing research. We present results on a variety of computer-generated and real data, having varying curvature, of different genus, and multiple objects. Gideon Guy, Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1996 | Inference of segmented, volumetric shape from three intensity imagesabstractWe present a method to infer segmented and full volumetric descriptions of objects from intensity images. We use three weakly calibrated images from closely spaced viewpoints as input. Deriving full volumetric descriptions requires the development of robust inference rules. The inference rules are based on local properties of generalized cylinders (GCs). We first detect groups in each image based on proximity, parallelism and symmetry. The groups in the three images are matched and their contours are labelled as "true" and "limb" edges. We use the information about groups and the label associated with their contours to recover visible surfaces and their surface axes. To extract the complete volume in terms of a GC, we need to infer the GC axis, its cross section and the scaling function. The properties of straight and curved axis generalized cylinders are used locally on the visible surfaces to obtain the GC axis. The cross section is recovered if seen in the images, else it is inferred using the visible surfaces and GC properties. We consider groups with true edges, limb edges or a combination of both. The final descriptions are volumetric and in terms of parts. Sometimes, when not enough information is present to make volumetric inferences, the descriptions remain at the surface level. We demonstrate results on real images of moderately complex objects with texture and shadows. Parag Havaldar, Gérard G. Medioni |
CVPR | 2 |
| 1996 | View Synthesis from Unregistered 2-D Images
Parag Havaldar, Mi-Suen Lee, Gérard G. Medioni |
Graphics Interface | 3 |
| 1996 | Spherical winged B-snakesabstractWe introduce spherical triangular B-splines for closed shape representation, and discuss the shape reconstruction using our new model "winged B-snakes", which are deformable surfaces coupled with active edges and junctions. We show the results of using the spherical winged B-snakes for simultaneous surface reconstruction and feature detection from range images or scattered 3D data. Gérard G. Medioni |
ICIP (2) | 2 |
| 1996 | Reconstructing free-form surfaces from sparse dataabstractWe propose a scheme to recover general free-form surfaces from sparse data, and the data may contain unknown discontinuities. We use a global voting method to infer from sparse data three dense potential fields for surfaces, edges, and junctions. We then use a new model called "winged B-snakes", which are deformable triangular B-spline surfaces embedded with active curves, to fit the surfaces and align the edges and junctions. A smooth C/sup 1/ surface with preserved discontinuity edges and junctions is obtained after the "winged B-snakes" have evolved and converged in the three potential fields using energy minimization. The triangular B-splines are state-of-the-art free-form surface representations and have good properties of arbitrary triangulation, lowest degree, local control, convex hull, automatic continuity, and affine invariance. Gérard G. Medioni |
ICPR | 2 |
| 1996 | Triangular NURBS Surface Modeling of Scattered DataabstractWe propose to fit triangular NURBS surfaces to noisy, sparse, scattered 3D data while simultaneously localizing and preserving sharp edges. We use a vector voting method to interpolate, from sparse data, three dense potential fields for surfaces, edges, and junctions. The global voting interpolants encode several human perceptual grouping principles such as cosurfacity, proximity, and constancy of curvature. The inferred potential fields are stored in three volumetric grids, giving each voxel the probability of being a surface point, an edge point, and a junction point. Then we use a new model called "winged B snakes", which are deformable triangular NURBS surfaces embedded with active curves, to fit the surfaces and align the edges and junctions. Finally, a smooth C/sup 1/ surface which preserves discontinuity edges and junctions is constructed. Fine tuning and surface fairing is done by adjusting the weights. Gérard G. Medioni |
IEEE Visualization | 2 |
| 1996 | Inferring global pereeptual contours from local features
Gideon Guy, Gérard G. Medioni |
Int. J. Comput. Vis. | 2 |
| 1996 | Perceptual grouping for generic recognition
Parag Havaldar, Gérard G. Medioni, Fridtjof Stein |
Int. J. Comput. Vis. | 2 |
| 1996 | Computer vision research at the University of Southern California
Ramakant Nevatia, Gérard G. Medioni |
Int. J. Comput. Vis. | 2 |
| 1996 | B-rep object description from multiple range views
Bahram Parvin, Gérard G. Medioni |
Int. J. Comput. Vis. | 2 |
| 1995 | Segmented Shape Descriptions from 3-View StereoabstractWe address the recovery of segmented, 3-D descriptions of an object from intensity images. We use three views of an object from slightly different viewpoints as our input. For each image we extract a hierarchy of groups based on proximity, parallelism and symmetry in a robust manner. The groups in the three images are matched by computing the epipolar geometry. For each set of matched groups from the three images, we then label the contours of the groups as "true" or "limb" edges. Using the information about groups, the label associated with their contours and projective properties of subclasses of Generalized Cylinders, we infer the 3-D structure of these groups. The proposed method not only allows robust shape recovery but also produces segmented parts. Our approach can also deal with groups generated as a result of texture or shadows on the object. We present results on real images of moderately complex objects.> Parag Havaldar, Gérard G. Medioni |
ICCV | 2 |
| 1995 | Surface Approximation of a Cloud of 3D Points
Chia-Wei Liao, Gérard G. Medioni |
CVGIP Graph. Model. Image Process. | 2 |
| 1995 | Description of Complex Objects from Multiple Range Images Using an Inflating Balloon Model
Gérard G. Medioni |
Comput. Vis. Image Underst. | 2 |
| 1995 | Finding Waldo, or Focus of Attention Using Local Color InformationabstractWe present a method to locate an "object" in a color image, or more precisely, to select a set of likely locations for the object. The model is assumed to be of known color distribution, which permits the use color-space processing. A new method is presented, which exploits more information than the previous backprojection algorithm of Swain and Ballard (1990) at a competitive complexity. Precisely, the new algorithm is based on matching local histograms with the model, instead of directly replacing pixels with a confidence that they belong to the object. We prove that a simple version of this algorithm degenerates into backprojection in the worst case. In addition, we show how to estimate the scale of the model. Results are shown on pictures digitized from the famous "Where is Waldo" books. Issues concerning the optimal choice of a color space and its quantization are carefully considered and studied in this application. We also propose to use co-occurrence histograms to deal with cases where important color variations can be expected.> François Ennesser, Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1995 | Map-based localization using the panoramic horizonabstractPresents an approach to solve the localization problem, in which an observer is given a topographic map of an area and dropped off at an unknown location. The solution to this problem requires establishing correspondences between viewer-centered observable features and their location on the map. The feature the authors select is the panoramic horizon curve, defined as the sky-ground boundary perceived by the observer as he performs a full 360/spl deg/ in place. In the authors' approach, they first precompute, offline, these horizon curves at a set of locations on a grid, from the topological map. These curves are approximated by polygons with different line fitting tolerances to gain robustness to noise in the authors' representation. These polygons are grouped into overlapping super segments, which are then encoded and stored in a table. The online computation consists of acquiring the panoramic view and extracting (with human help) the horizon curve. This curve is approximated by a polygon and the resulting super segments, used as indices in the data base, allow one to retrieve candidate locations. The best candidate is selected during a verification step which applies geometric constraints. This process uses local features and can therefore tolerate significant occlusion likely to occur in real environments. The authors illustrate the performance of the approach on results obtained from real data. Fridtjof Stein, Gérard G. Medioni |
IEEE Trans. Robotics Autom. | 2 |
| 1994 | Surface description of complex objects from multiple range imagesabstractAddresses the problem of constructing a complete surface model of an object using a set of registered range images. Our approach is based on a dynamic balloon model represented by using a triangulated mesh. The vertices in the mesh are linked to their neighboring vertices by springs to simulate the surface tension and to keep the shell smooth. Unlike other dynamic models proposed by previous researchers, our balloon model is driven purely by an applied inflation force towards the object surface from inside the object, until the mesh elements reach the object surface. The system includes an adaptive local triangle mesh subdivision scheme that results in an evenly distributed mesh. Since our approach is not based on global minimization, it can handle complex, non-star-shaped objects without relying on a carefully selected initial state or encountering a local minimum problem. It also allows us to adapt the mesh surface to changes in local surface shapes and to handle any holes that are present in the input data by adjusting certain system parameters adaptively and locally. We present results on some complex, non-star-shaped objects from real range images.> Gérard G. Medioni |
CVPR | 2 |
| 1994 | Simultaneous segmentation and approximation of complex patternsabstractDeformable model have been widely used to approximate objects from collected data points, but most algorithms based on the deformable model can only handle geometrically and topologically simple objects. They are inadequate for objects with deep cavities or multi-part objects. Furthermore, they always assume there is only one underlying object for the collected data, which means the segmentation has been done ahead of time. Unlike most deformable algorithms which approximate one object at a time, our proposed approach can apply simultaneously more than one curve to approximate multiple objects. Using (1) the residual data points, (2) the bad parts of the fitting curve, and (3) appropriate Boolean operations, our approach is able to detect patterns with holes or cavities, and can perform segmentation by itself for more than one underlying object. We currently present experiments mainly on 2D data. These 2D algorithms can be extended to 3D without theoretical difficulties. An experiment on 3D data, composed of two genus I toruses, is also presented. Also, a new method for defining the external energy is presented, which helps capture the shape more accurately with low time and reasonable space complexities, and a method to prevent self-intersection of the curve during evolution is also introduced.> Chia-Wei Liao, Gérard G. Medioni |
CVPR | 2 |
| 1994 | Extraction Of Groups For Recognition
Parag Havaldar, Gérard G. Medioni, Fridtjof Stein |
ECCV (1) | 2 |
| 1994 | Learning, recognition and navigation from a sequence of infrared imagesabstractWe address the problem in which an autonomous system equipped with a single infrared camera "learns" a designated rigid 3D scene so that, at another time, it can recognize it and guide itself relative to the reconstructed scene, starting from an approximately known viewpoint, to reach a give destination. This scenario is relevant to several domains, in particular military missions and robotic navigation. The goal of our system is to realise a real-time implementation (on special hardware). In this paper, we describe the software version of such an automated system and show results on real infrared images. Nicolas Milhaud, Gérard G. Medioni |
ICPR (1) | 2 |
| 1994 | Part decomposition and description of 3D shapesabstractAddresses the problem of obtaining natural descriptions of 3D shapes. The authors present one of the first attempts to address the description of 3D compound objects, where the parts are connected smoothly. The input the authors consider is either complete 3D data or range data from a single view. The authors suggest a volumetric graph representation of the object, where the nodes represent individual parts and the edges represent connectivity information. The authors suggest the use of properties of the parabolic curves for performing the part decomposition. The authors consider parts with tubular structure with a straight or curved axis. They are also interested in the internal description of the parts. The authors study two classes of shapes, namely straight homogeneous GCs, and planar right constant GCs. The authors suggest the use of properties of the parabolic curves for recovering natural descriptions of these classes in terms of their cross sections and axes. Hillel Rom, Gérard G. Medioni |
ICPR (1) | 2 |
| 1994 | Model validation for change detection [machine vision]abstractAn important application of machine vision is to provide a means to monitor a scene over a period of time and report changes in the content of the scene. We have developed a validation mechanism that implements the first step towards a system for detecting changes in images of aerial scenes. By validation we mean the confirmation of the presence of model objects in the image. Our system uses a 3-D site model of the scene as a basis for model validation, and eventually for detecting changes and to update the site model. The scenario for our present validation system consists of adding a new image to a database associated with the site. The validation process is implemented in three steps: registration of the image to the model, or equivalently, determination of the position and orientation of the camera; matching of model features to image features; and validation of the objects in the model. Our system processes the new image monocularly and uses shadows as 3-D clues to help validate the model. The system has been tested using a hand-generated site model and several images of a 500:1 scale model of the site, acquired form several viewpoints.> Mathias Bejanin, Andres Huertas, Gérard G. Medioni, Ramakant Nevatia |
WACV | 3 |
| 1993 | Finding Waldo, or focus of attention using local color informationabstractA method is presented to locate an object in a color image, or more precisely, to select a set of likely locations for the object. The model is assumed to be of known color, which permits the use of color-space processing. A new method is presented, which exploits more information than the previous backprojection algorithm of Swain and Ballard at a competitive complexity. The new algorithm is based on matching local histograms with the model, instead of directly replacing pixels with a confidence that they belong to the object. It is proved that a simple version of this algorithm degenerates into backprojection in the worst case. The authors show how to estimate the scale of the model. The use of co-occurrence histograms is proposed to deal with cases where important color variations can be expected.> François Ennesser, Gérard G. Medioni |
CVPR | 2 |
| 1993 | Inferring global perceptual contours from local featuresabstractAn attempt is made to solve the problem of imperfect data produced by state-of-the-art edge detectors through the implementation of laws of perceptual grouping, derived from psychology. A saliency-enhancing operator is introduced. It is capable of highlighting features (edges, junctions, etc.) which are considered important psychologically. It also infers features which are not detected by low-level detectors. It is shown how to extract salient curves and junctions and generate a description ranking these features by the likelihood of them occurring accidentally. The problem of illusory contours apparent in end-point formations is discussed. All operations are parameter-free, noniterative and are linear with the number of edges in the input image.> Gideon Guy, Gérard G. Medioni |
CVPR | 2 |
| 1993 | Hierarchical Decomposition and Axial Shape DescriptionabstractA method for producing a segmented axial description of a given shape together with a hierarchical decomposition of the shape into its parts is presented. The novelty of this approach lies in the combination of several competing approaches and tools into a unified scheme and an efficient implementation producing natural descriptions. Smooth local symmetries are used for the axial description of parts, which are suggested by curvature sign changes. Parallel symmetries are used to provide information on global relationships within the shape. This information is used for parsing shape into a hierarchy of parts. This approach uses both region and contour information, can handle shapes with corners, and addresses the issue of local versus global information, the issue of scale, and the notion of part. The method is computationally efficient, robust, and stable. Results that show that it provides an intuitive shape description are included.> Hillel Rom, Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1993 | B-spline Contour Representation and Symmetry DetectionabstractThe detection of edges is only one of many steps in the understanding of images. Further processing necessarily involves grouping operations between contours. We present a representation of edge contours by approximating B-splines and show that such a representation facilitates the extraction of symmetries between contours. Our representation is rich, compact, stable, and does not critically depend on feature extraction. We turn our attention to the detection of three types of symmetries: skew symmetries and parallel symmetries, which have proven to be of great importance in inferring shape from contour, and smooth local symmetries, which have been used for planar shape description. We show that our representation facilitates the computation of these symmetries.> Philippe Saint-Marc, Hillel Rom, Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 1992 | Hierarchical decomposition and axial shape descriptionabstractThe problem of obtaining intuitive descriptions of planar shapes is addressed. In particular, a method for producing a segmented axial description of a given shape together with a hierarchical decomposition of the shape into its parts is suggested. Smooth local symmetries are used for the axial description of parts. Parallel symmetries are used to provide information on global relationships within the shape. It is assumed that the shape is a closed planar curve. The approach uses both region and contour information, can handle shapes with corners, and addresses the issues of local versus global information, the issue of scale and the notion of part. The method is computationally efficient, robust, and stable. Results showing that it provides an intuitive shape description are presented.> Hillel Rom, Gérard G. Medioni |
CVPR | 2 |
| 1992 | Perceptual grouping using global saliency-enhancing operatorsabstractIntroduces saliency-enhancing operators capable of highlighting features which are considered perceptually relevant. One is able to extract salient curves and junctions and generate a description ranking these features by their likelihood of coming from the original scene. The authors suggest the global extension field as means of describing the behavior of a curve segment, in terms of its continuation. It is shown that a directional convolution of an edge image with the above field can produce useful descriptions. Other fields are also used in the same manner to produce similar results for domain-specific applications. The scheme is particularly useful and robust as a gap filler and in the presence of noise.> Gideon Guy, Gérard G. Medioni |
ICPR (1) | 2 |
| 1992 | Representation of range data with B-spline surface patchesabstractPresents an implementation of deformable models for approximation of 3-D surfaces. It is an extension of work on B-snakes (Menet, Saint-Marc, and Gerard (1990), which approximate curves using B-splines. The user provides an initial simple surface, such as a cylinder or a plane, which is subject to internal forces and external forces which attracts it toward features. The problem is cast in terms of energy minimization, and is solved iteratively. The authors choice of basis functions leads to reasonable complexity and good numerical stability. The authors give results on real range images to illustrate the applicability of their approach. The advantages of this approach are that it provides a compact representation of the approximated data, it gives a C/sub 1/ continuous analytical description of the data, which allows computation of differential properties, and lends itself to applications such as non-rigid motion tacking and object recognition.> Chia-Wei Liao, Gérard G. Medioni |
ICPR (3) | 2 |
| 1992 | B-rep from unregistered multiple range imagesabstractThe authors present a method to produce an integrated description of an object given multiple range views without registration. The multiple view representation is in the form of a B-rep (boundary representation). They describe each view of the object as an attributed graph whose nodes are surface patches and links are the relations between surfaces. Any two attributed graphs, each corresponding to a given view, are matched, and the rigid motion transformation between them is computed. The basic strategy for multiple view integration is composed of two aspects: first, a composite graph which contains the bounding surfaces and their corresponding attributes is built and then these surfaces are intersected so that the edges and vertices corresponding to the B-rep description are identified. Results on objects with polyhedral as well as quadratic curved surfaces are presented.> Bahram Parvin, Gérard G. Medioni |
ICRA | 2 |
| 1992 | Map-based localization using the panoramic horizonabstractAn approach is presented to solve the localization problem, in which an observer is given a topographic map of an area and dropped off at an unknown location. The solution to this problem requires establishing correspondences between viewer-centered observable features and their location on the map. The feature selected is the panoramic horizon curve, defined as the sky-ground boundary perceived by the observer as a full 360 degrees rotation is performed. The authors propose to extract from many locations in the map the panoramic horizon curves which would be observed by the observer at each location. Such curves are encoded and stored in a table. To locate an unknown location, the panoramic horizon curve of the unknown location is first extracted, and then approximated by a family of polygons with different line fitting tolerances. By indexing into the table, candidate locations are retrieved. The correct candidate is found by applying further geometrical constraints in the verification step. The claims are validated by showing some results from a real map.> Fridtjof Stein, Gérard G. Medioni |
ICRA | 2 |
| 1992 | The RegiStar Machine: from conception to installationabstractThe authors have developed a machine to perform the task of automatic registration of color separation films, a process manually performed by skilled professionals in the graphics arts printing industry. The development of such a machine requires overcoming significant challenges: designing a sound computer vision methodology while respecting hard timing constraints, transferring software across platforms and languages, validating the software, building the actual machine around the algorithms, testing the conformity to tolerances, educating operators on the use of such a machine, and having a system robust enough to operate around the clock with no technical supervision. The authors present a brief overview of the problem, followed by the answers they provided to the challenges above.> Gérard G. Medioni, Andres Huertas, Monti R. Wilson |
WACV | 1 |
| 1992 | Object modelling by registration of multiple range images
Gérard G. Medioni |
Image Vis. Comput. | 2 |
| 1992 | 3-D Surface Description from Binocular StereoabstractA stereo vision system that attempts to achieve robustness with respect to scene characteristics, from textured outdoor scenes to environments composed of highly regular man-made objects is presented. It integrates area-based and feature-based primitives. The area-based processing provides a dense disparity map, and the feature-based processing provides an accurate location of discontinuities. An area-based cross correlation, an ordering constraint, and a weak surface smoothness assumption are used to produce an initial disparity map. This disparity map is only a blurred version of the true one because of the smoothing introduced by the cross correlation. The problem can be reduced by introducing edge information. The disparity map is smoothed and the unsupported points removed. This method gives an active role to edgels parallel to the epipolar lines, whereas they are discarded in most feature-based systems. Very good results have been obtained on complex scenes in different domains.> Steven D. Cochran, Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1992 | Structural Indexing: Efficient 3-D Object RecognitionabstractThe authors present an approach for the recognition of multiple 3-D object models from three 3-D scene data. The approach uses two different types of primitives for matching: small surface patches, where differential properties can be reliably computed, and lines corresponding to depth or orientation discontinuities. These are represented by splashes and 3-D curves, respectively. It is shown how both of these primitives can be encoded by a set of super segments, consisting of connected linear segments. These super segments are entered into a table and provide the essential mechanism for fast retrieval and matching. The issues of robustness and stability of the features are addressed in detail. The acquisition of the 3-D models is performed automatically by computing splashes in highly structured areas of the objects and by using boundary and surface edges for the generation of 3-D curves. The authors present results with the current system (3-D object recognition based on super segments) and discuss further extensions.> Fridtjof Stein, Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1992 | Structural Indexing: Efficient 2D Object RecognitionabstractThe problem of recognition of multiple flat objects in a cluttered environment from an arbitrary viewpoint is addressed. The models are acquired automatically and approximated by polygons with multiple line tolerances for robustness. Groups of consecutive segments (super segments) are then encoded and entered into a table. This provides the essential mechanism for indexing and fast retrieval. Once the database of all models is built, the recognition proceeds by segmenting the scene into a polygonal approximation; the code for each super segment retrieves model hypotheses from the table. Hypotheses are clustered if they are mutually consistent and represent the instance of a model. Finally, the estimate of the transformation is refined. This methodology makes it possible to recognize models despite noise, occlusion, scale rotation translation, and a restricted range of weak perspective. A complexity bound is obtained.> Fridtjof Stein, Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1991 | A dynamic system for object description and correspondenceabstractA unified framework for object description and correspondence in range images is proposed. This is based on modeling the ambiguities with a dynamic system at every level of hierarchy. The critical issue in object description is boundary completion. The authors show how this step can be modeled as the interaction of long and short term variables operating on partial contours, represented by parametric splines. In a sense, this network operates on an evolutionary basis where boundary completion is modeled as a synergistic process, and mutation is modeled as differential decay. Once reliable surface features are extracted, the same framework, with appropriate constraints, can find the corresponding surfaces and compute the transformation.> Bahram Parvin, Gérard G. Medioni |
CVPR | 2 |
| 1991 | Structural hashing: efficient three dimensional object recognitionabstractAn approach for the recognition of multiple three-dimensional object models from three-dimensional scene data is presented. The authors work on dense data, but neither the models nor the scene data have to be complete. The problem is addressed in a realistic environment: the viewpoint is arbitrary, the objects vary widely in complexity, and no assumptions about the structure of the surface are made. The approach is novel in that it uses two different types of primitives for matching: small surface patches, where differential properties can be reliably computed, and lines corresponding to depth or orientation discontinuities. These are represented by splashes and 3-D curves respectively. It is shown how both of these primitives can be encoded by a set of super segments, consisting of connected linear segments. These super segments are entered into a hash table, and provide the essential mechanism for fast retrieval and matching.> Fridtjof Stein, Gérard G. Medioni |
CVPR | 2 |
| 1991 | Object modeling by registration of multiple range imagesabstractThe problem of creating a complete model of a physical object is studied. Although this may be possible using intensity images, the authors use range images which directly provide access to three-dimensional information. The first problem that needs to be solved is to find the transformation between the different views. Previous approaches have either assumed this transformation to be known (which is extremely difficult for a complete model) or computed it with feature matching (which is not accurate enough for integration. The authors propose an approach that works on range data directly and registers successive views with enough overlapping area to get an accurate transformation between views. This is performed by minimizing a functional that does not require point-to-point matches. Details are given of the registration method and modeling procedure, and they are illustrated on range images of complex objects.> Gérard G. Medioni |
ICRA | 2 |
| 1991 | A layered network for the correspondence of 3D objectsabstractA computational approach for solving the correspondence problem between different views of objects in range images is presented. This is modeled as a layered constraint satisfaction network which can be implemented on a parallel analog neural network. In this approach, each view of an object is represented by an attributed graph with nodes as surfaces and their bounding vertices, and links as relations between adjacent surfaces. The matching strategy is a two-step process. Each step is formulated with a constraint satisfaction network, and implemented on a Hopfield network. At each level, a set of local, adjacency and global constraints is specified, and an appropriate energy function to be minimized is defined. At the first level of this hierarchy, surface patches are matched and clusters of rotation transformations are hypothesized. At the second level, the computed rotation transformation is applied to the corresponding vertices, and the translation vector is computed.> Bahram Parvin, Gérard G. Medioni |
ICRA | 2 |
| 1991 | Adaptive Smoothing: A General Tool for Early VisionabstractA method to smooth a signal while preserving discontinuities is presented. This is achieved by repeatedly convolving the signal with a very small averaging mask weighted by a measure of the signal continuity at each point. Edge detection can be performed after a few iterations, and features extracted from the smoothed signal are correctly localized (hence, no tracking is needed). This last property allows the derivation of a scale-space representation of a signal using the adaptive smoothing parameter k as the scale dimension. The relation of this process to anisotropic diffusion is shown. A scheme to preserve higher-order discontinuities and results on range images is proposed. Different implementations of adaptive smoothing are presented, first on a serial machine, for which a multigrid algorithm is proposed to speed up the smoothing effect, then on a single instruction multiple data (SIMD) parallel machine such as the Connection Machine. Various applications of adaptive smoothing such as edge detection, range image feature extraction, corner detection, and stereo matching are discussed.> Philippe Saint-Marc, Jer-Sen Chen, Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 1991 | A versatile PC-based range finding systemabstractThe authors present an active triangulation-based range finding system composed of an independent laser system generating a sheet of light projected on the object to be measured, which is placed on a linear or a rotary table driven by a personal computer. This computer includes a video digitizer board to which two cameras, looking at the scene from both sides of the sheet of light, are connected. Besides its low cost, this system has several advantages over similar systems. First of all, two cameras are used to limit the occlusion problem, and a method is proposed to integrate range data obtained from these cameras into a single range image. The calibration of each camera is very simple, provides subpixel accuracy, and is performed only once as the laser or the camera does not move. The data acquisition uses an interpolation technique that produces very accurate depth measurements. The system also provides intensity data in registration with the range data. The application of all these techniques is illustrated by showing numerous examples of the range and intensity data acquisition from various complex objects.> Philippe Saint-Marc, Jean-Luc Jezouin, Gérard G. Medioni |
IEEE Trans. Robotics Autom. | 3 |
| 1990 | Parallel Multiscale Stereo Matching Using Adaptive Smoothing
Jer-Sen Chen, Gérard G. Medioni |
ECCV | 2 |
| 1990 | B-Spline Contour Representation and Symmetry Detection
Philippe Saint-Marc, Gérard G. Medioni |
ECCV | 2 |
| 1990 | Object recognition using geometric hashing on the Connection MachineabstractA parallel implementation of a system to recognize 2D objects under realistic scenarios (occlusion, rotation, translation, and perspective) is presented. A preprocessing phase and a recognition phase are used. Both phases have been implemented on the Connection Machine, achieving O(n/sup -x/) with n/sup x/ processors (x> Olivier Bourdon, Gérard G. Medioni |
ICPR (2) | 2 |
| 1990 | Efficient two dimensional object recognitionabstractThe problem of recognition of multiple flat objects in a cluttered environment from an arbitrary viewpoint (weak perspective) is addressed. The models are acquired automatically and initially approximated by polygons with multiple line tolerances for robustness. Groups of consecutive segments (supersegments) are then gray-coded and entered into a hash table. This provides the essential mechanism for indexing and fast retrieval. Once the database of all models is built, the recognition proceeds by segmenting the scene into a polygonal approximation; the gray code for each supersegment retrieves model hypotheses from the hash table. Hypotheses are clustered if they are mutually consistent and represent the instance of a model. The estimate of the transformation is refined. This methodology makes it possible to recognize models in the presence of noise, occlusion, scale, rotation, translation, and weak perspective. Unlike most of the current systems, its complexity grows as O(kN), where N is the number of models and k> Fridtjof Stein, Gérard G. Medioni |
ICPR (1) | 2 |
| 1990 | Automatic registration of color separation filmsabstractThe problem of registration of four-color halftone separations for color printing is addressed. A method is presented for automatic registration of color separation films. An operator manually selects two windows from the reference negative (generally cyan) with a digitizing cursor, and each window covering approximately 6 mm/sup 2/ (0.25 in/sup 2/) is digitized into a 640*640 array. On each negative and for each window, the macro edges are extracted, and the contours are approximated by line segments. The segments from corresponding windows on different negatives are then matched with the reference ones to provide an estimate of the translation between them. The two translations (from the two windows) provide the parameters of the rigid planar transform between negatives (rotation and translation) and permit the punching of registration holes into the pictures for each negative. The system has been implemented in the RegiStar machine, built to perform the mechanical tasks associated with the algorithm. It is able to handle a set of four-color separations in about 5 min, from image acquisition to punching of registration holes on the films, maintaining an accuracy of 12 mu m (0.5 mil) for binary patterns and 25 mu m (1 mil) for true halftones. This speed is obtained by using an off-the-shelf Mercury array processor attached to an IBM personal computer.> Gérard G. Medioni, Monti R. Wilson, Andres Huertas |
ICRA | 1 |
| 1990 | Refining edges detected by a LoG operator
Fatih Ulupinar, Gérard G. Medioni |
Comput. Vis. Graph. Image Process. | 2 |
| 1990 | Automatic registration of color separation films
Gérard G. Medioni, Andres Huertas, Monti R. Wilson |
Mach. Vis. Appl. | 1 |
| 1989 | Adaptive smoothing: a general tool for early visionabstractThe authors present a method to smooth a signal-whether it is an intensity image, a range image, or a contour-which preserves discontinuities and thus facilitates their detection. This is achieved by repeatedly convolving the signal with a very small averaging filter modulated by a measure of the signal discontinuity at each point. This process is related to the anisotropic diffusion reported by P. Perona and J. Malik (1987) but it has a much simpler formulation and is not subject to instability or divergence. Real examples show how this approach can be applied to the smoothing of various types of signals. The detected features do not move, and thus no tracking is needed. The last property makes it possible to derive a novel scale-space representation of a signal using a small number of scales. Finally, this process is easily implemented on parallel architectures: the running time on a 16 K connection machine is three orders of magnitude faster than on a serial machine.> Philippe Saint-Marc, Jer-Sen Chen, Gérard G. Medioni |
CVPR | 3 |
| 1989 | Issues in geometric reasoning from range imageryabstractThe authors discuss some of the issues that have to be tackled in order to perform geometric reasoning from range imagery. They begin by pointing out that a successful system must deal with real-world data, and therefore take into account the effects of noise and quantization. They suggest that adaptive smoothing may prove to be a helpful tool for such a task. The next stage of processing involves a symbolic representation of the original data. The authors spell out criteria for shape description, discuss current representation schemes and point out their limitations, and then propose some ideas for overcoming such limitations, illustrated on real examples. Finally, they look at the issues in recognition, and more specifically the matching part, with reference to different methodologies, tree search and constraint satisfaction network.> Gérard G. Medioni, Philippe Saint-Marc |
SMC | 1 |
| 1989 | Adaptive multiscale feature extraction from range data
Bahram Parvin, Gérard G. Medioni |
Comput. Vis. Graph. Image Process. | 2 |
| 1989 | Author's Reply
Jer-Sen Chen, Andres Huertas, Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 1989 | Detection, Localization, and Estimation of EdgesabstractA method to detect, locate, and estimate edges in a one-dimensional signal is presented. It is inherently more accurate than all previous schemes as it explicitly models and corrects interaction between nearby edges. The method is iterative with initial estimation of edges provided by the zero crossings of the signal convolved with Laplacian of Gaussian (LoG) filter. The necessary computations necessitate knowledge of this convolved output only in a neighborhood around each zero crossing and in most cases, could be performed locally by independent parallel processors. Results on one-dimensional slices extracted from real images, and on images which have been proposed independently in the row and column directions are shown. An analysis of the method is provided including issues of complexity and convergence, and directions of future research are outlined.> Jer-Sen Chen, Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1989 | Recognizing 3-D Objects Using Surface DescriptionsabstractThe authors provide a complete method for describing and recognizing 3-D objects, using surface information. Their system takes as input dense range date and automatically produces a symbolic description of the objects in the scene in terms of their visible surface patches. This segmented representation may be viewed as a graph whose nodes capture information about the individual surface patches and whose links represent the relationships between them, such as occlusion and connectivity. On the basis of these relations, a graph for a given scene is decomposed into subgraphs corresponding to different objects. A model is represented by a set of such descriptions from multiple viewing angles, typically four to six. Models can therefore be acquired and represented automatically. Matching between the objects in a scene and the models is performed by three modules: the screener, in which the most likely candidate views for each object are found; the graph matcher, which compares the potential matching graphs and computes the 3-D transformation between them; and the analyzer, which takes a critical look at the results and proposes to split and merge object graphs.> Ting-Jun Fan, Gérard G. Medioni, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1989 | Stereo Error Detection, Correction, and EvaluationabstractAn algorithm is presented for error detection and correction of disparity, as a process separate from stereo matching, with the contention that matching is not necessarily the best way to utilize all the physical constraints characteristic to stereopsis. As a result of the bias in stereo research towards matching, vision tasks like surface interpolation and object modeling have to accept erroneous data from the stereo matchers without the benefits of any intervening stage of error correction. An algorithm which identifies all errors in disparity data that can be detected on the basis of figural continuity and corrects them is presented. The algorithm can be used as a postprocessor to any edged-based stereo matching algorithm, and can additionally be used to automatically provide quantitative evaluations on the performance of matching algorithms of this class.> Rakesh Mohan, Gérard G. Medioni, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1988 | Building an accurate range finder with off the shelf componentsabstractThe authors present an active triangulation range finding system composed of an independent laser system generating a plane of light projected on an object placed on a rotary table driven by a personal computer. This computer includes a video digitizer board connected to a camera looking at the scene. Besides its low cost, this system has other advantages over the comparable existing systems. First, the authors have designed a simple, fast and accurate calibration procedure which does not require any knowledge about the camera parameters or the relative position of the camera with the laser plane. Furthermore, this calibration procedure is performed only once, ensuring stable and accurate results. The result of the scanning of a given object is given in cylindrical coordinates. Choosing different viewpoints, Cartesian range images of the same object are computed in order to show, with shaded and perspective views of the scanned object, the quality of the results.> Jean-Luc Jezouin, Philippe Saint-Marc, Gérard G. Medioni |
CVPR | 3 |
| 1988 | Useful geometric properties of the generalized coneabstractThe authors present results on geometric properties of the generalized cone, in an effort to utilize it for a shape description system. They first derive the relationship between the generalized cone description and the surface description given by differential geometry. Then they derive expressions for the Gaussian and mean curvatures of a generalized cone, in general, and obtain expressions for some special cases like the torus, the solid of revolution etc. They study the planarity property of the contour generators of a generalized cone, in particular, one with a planar axis. They find that homogeneous generalized cones with planar axes and circular cross sections or constant-size cross sections have planar contour generators in an orthographic side view. An example of such a generalized cone is the torus. However, the contour generators are not planar in a general view. They also study symmetry properties of some generalized cones and find, in particular, that in orthographic projection the contour of the solid of revolution is symmetric about the projection of its axis from any point of view.> Kashipati Rao, Gérard G. Medioni |
CVPR | 2 |
| 1988 | Refining edges detected by a LoG operatorabstractThe Laplacian of Gaussian (LoG) operator is one of the most popular operators used in edge detection. This operator, however, has some problems: zero-crossings do not always correspond to edges, and edges with an asymmetric profile introduce a symmetric bias between edge and zero-crossing locations. The authors offer solutions to these two problems. First, for one-dimensional signals, such as slices from images, they propose a simple test to detect true edges, and, for the problem of bias, they propose different techniques: the first one combines the results of the convolution of two LoG operators of different deviations, whereas the others sample the convolution with a single LoG filter at two points besides the zero-crossing. In addition to localization, these methods allow them to further characterize the shape of the edge. The authors present an implementation of these techniques for edges in 2-D images.> Fatih Ulupinar, Gérard G. Medioni |
CVPR | 2 |
| 1988 | Spatio-temporal analysis for velocity estimation of contours in an image sequence with occlusionabstractA method is presented for analyzing a sequence of closely sampled images, forming a spatio-temporal volume. It is proposed to find the normal velocity component of each edge point by making assumptions about this scene, such as smooth motion, common motion and constant illumination. The resulting normal velocity field can then be used to resolve the real velocity field and to perform segmentation. The system is quite robust and is capable of handling occlusion as well as disocclusion. Results are presented on synthetic data consisting of two objects moving with occlusion and on real image sequences.> Shou-Ling Peng, Gérard G. Medioni |
ICPR | 2 |
| 1988 | Matching 3-D objects using surface descriptionsabstractA method is developed to extract important curves, corresponding to physical boundaries of objects, from a range image. It is shown how to infer, from these curves, a segmentation of the scene into surface patches, and how to use these descriptions to establish correspondences between two scenes. In a first step, labeled curves corresponding to jump boundaries, creases, and limbs of objects are grouped into boundaries of regions. Each region is therefore described by its boundaries and by a polynomial approximation, which allows each path individually and also the complete scene to be reconstructed. In a second step, objects (or partial objects) are inferred from surface patches, and then two range images are at this partial object level. Graphs of objects are matched using a best-first search under three types of constraints: unary constraints between corresponding nodes, binary constraints between corresponding linked pairs of nodes, and constraints imposed by the computed geometric transformation. Substantial partial occlusion is allowed. The generality and robustness of this approach is illustrated by several examples.> Ting-Jun Fan, Gérard G. Medioni, Ramakant Nevatia |
ICRA | 2 |
| 1988 | Robot hand-eye coordination: shape description and graspingabstractThe successful execution of grasps by a robot hand requires a translation of visual information into control signals to the hand which produce the desired spatial orientation and preshape for an arbitrary object. An approach to this problem, based on separation of the task into two modules, is presented. A vision module is used to transform an image into a volumetric shape description using generalized cones. The data structure containing this geometric information becomes an input to the grasping module, which obtains a list of feasible grasp modes and a set of control signals for the robot hand. Various features of both modules are discussed.> Kashipati Rao, Gérard G. Medioni, Huan Liu 0001, George A. Bekey |
ICRA | 2 |
| 1987 | Fast Convolution with Laplacian-of-Gaussian MasksabstractWe present a technique for computing the convolution of an image with LoG (Laplacian-of-Gaussian) masks. It is well known that a LoG of variance a can be decomposed as a Gaussian mask and a LoG of variance a1 < a. We take advantage of the specific spectral characteristics of these filters in our computation: the LoG is a bandpass filter; we can therefore fold the spectrum of the image (after low pass filtering) without loss of information, which is equivalent to reducing the resolution. We present a complete evaluation of the parameters involved, together with a complexity analysis that leads to the paradoxical result that the computation time decreases when a increases. We illustrate the method on two images. Jer-Sen Chen, Andres Huertas, Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 1987 | Segmented descriptions of 3-D surfacesabstractA method to segment and describe visible surfaces of three-dimensional (3-D) objects is presented by first segmenting the surfaces into simple surface patches and then using these patches and their boundaries to describe the 3-D surfaces. First, distinguished points are extracted which will comprise the edges of segmented surface patches, using the zero-crossings and extrema of curvature along a given direction. Two different methods are used: if the sensor provides relatively noise-free range images, the principal curvatures are computed at only one resolution, otherwise, a multiple scale approach is used and curvature is computed in four directions 45° apart to facilitate interscale tracking. These points are then grouped into curves and these curves are classified into different classes which correspond to significant physical properties such as jump boundaries, folds, and ridge lines (or smooth extrema). Then jump boundaries and folds are used to segment the surfaces into surface patches, and a simple surface is fitted to each patch to reconstruct the original objects. These descriptions not only make explicit most of the salient properties present in the original input, but are more suited to further processing, such as matching with a given model. The generality and robustness of this approach is illustrated on scene images with different available range sensors. Ting-Jun Fan, Gérard G. Medioni, Ramakant Nevatia |
IEEE J. Robotics Autom. | 2 |
| 1986 | Corner detection and curve representation using cubic B-splinesabstractIn this paper, we propose to use B-Splines to represent digital curves. We have developed an efficient algorithm to locate corners and at the same time encode curve segments between them using B-Splines. Used in conjunction with our subpixel edge detector, [1], it allows us to obtain accurate position of the corners, as needed in many registration problems such as stereo matching and motion parameter estimation. In addition to corners, we detect points of significant curvature between them. The resulting representation is a good approximation of the original, in the sense that it makes interesting points explicit, and achieves significant data compression. Gérard G. Medioni, Yoshio Yasumoto |
ICRA | 1 |
| 1986 | Detection of Intensity Changes with Subpixel Accuracy Using Laplacian-Gaussian MasksabstractWe present a system that takes a gray level image as input, locates edges with subpixel accuracy, and links them into lines. Edges are detected by finding zero-crossings in the convolution of the image with Laplacian-of-Gaussian (LoG) masks. The implementation differs markedly from M.I.T.'s as we decompose our masks exactly into a sum of two separable filters instead of the usual approximation by a difference of two Gaussians (DOG). Subpixel accuracy is obtained through the use of the facet model [1]. We also note that the zero-crossings obtained from the full resolution image using a space constant ¿ for the Gaussian, and those obtained from the 1/n resolution image with 1/n pixel accuracy and a space constant of ¿/n for the Gaussian, are very similar, but the processing times are very different. Finally, these edges are grouped into lines using the technique described in [2]. Andres Huertas, Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1986 | Robust Estimation of Three-Dimensional Motion Parameters from a Sequence of Image Frames Using RegularizationabstractIn this paper, we look at the issue of accurate estimation of the three-dimensional motion parameters of a rigid body from a sequence of synthetic images, and relate the effect of some parameters to the shape of an error function. We first consider the case where only a small set of corresponding points is identified and suggest that a technique called regularization improves the quality and stability of a solution. We then observe that, if more pairs of corresponding points are available, the error function becomes smooth and the solution stable. Finally, we try to improve the quality of estimation by considering more than two consecutive frames for a moving camera looking at a stationary scene, and summing the error functions obtained for any two consecutive frames. Surprisingly enough, this technique does not improve stability unless we use regularization again. Yoshio Yasumoto, Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1985 | Segment-based stereo matching
Gérard G. Medioni, Ramakant Nevatia |
Comput. Vis. Graph. Image Process. | 1 |
| 1984 | Matching Images Using Linear FeaturesabstractWe describe techniques for matching two images or an image and a map. This operation is basic for machine vision and is needed for the tasks of object recognition, change detection, map up-dating, passive navigation, and other tasks. Our system uses line-based descriptions, and matching is accomplished by a relaxation operation which computes most similar geometrical structures. A more efficient variation, called the ``kernel'' method, is also described. We give results on complex aerial images which contain many image differences, caused by varying sun position, different seasons, and imaging environments, and also structural changes caused by man-made alterations such as new construction. Gérard G. Medioni, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1984 | Visual inspection using linear features
Gérard G. Medioni, Ramakant Nevatia |
Pattern Recognit. | 2 |
| 1982 | Segmentation of Images Into Regions Using Edge Information
Gérard G. Medioni |
AAAI | 1 |