VLDB 2026 Research / reviewers in the wild / expert
Jean Ponce
dblp:p/JeanPonce
· DBLP profile ↗
191ranked-venue papers
29as first author
24since 2021 · last 2025
0009-0000-5449-7620ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 180 · 26 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 108 · 13 first-author · 10 since 2021Systems, architecture and hardware · 26 · 8 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Online 3D Scene Reconstruction Using Neural Object PriorsabstractThis paper addresses the problem of reconstructing a scene online at the level of objects given an RGB-D video sequence. While current object-aware neural implicit rep-resentations hold promise, they are limited in online reconstruction efficiency and shape completion. Our main contributions to alleviate the above limitations are twofold. First, we propose a feature grid interpolation mechanism to continuously update grid-based object-centric neural implicit representations as new object parts are revealed. Second, we construct an object library with previously mapped objects in advance and leverage the corresponding shape priors to initialize geometric object models in new videos, sub-sequently completing them with novel views as well as synthesized past views to avoid losing original object details. Extensive experiments on synthetic environments from the Replica dataset, real-world ScanNet sequences and videos captured in our laboratory demonstrate that our approach outperforms state-of-the-art neural implicit models for this task in terms of reconstruction accuracy and completeness. Thomas Chabal, Shizhe Chen, Jean Ponce, Cordelia Schmid |
3DV | 3 |
| 2025 | A New Statistical Model of Star Speckles for Learning to Detect and Characterize Exoplanets in Direct Imaging ObservationsabstractThe search for exoplanets is an active field in astronomy, with direct imaging as one of the most challenging methods due to faint exoplanet signals buried within stronger residual starlight. Successful detection requires advanced image processing to separate the exoplanet signal from this nuisance component. This paper presents a novel statistical model that captures nuisance fluctuations using a multi-scale approach, leveraging problem symmetries and a joint spectral channel representation grounded in physical principles. Our model integrates into an interpretable, end-to-end learnable framework for simultaneous exoplanet detection and flux estimation. The proposed algorithm is evaluated against the state of the art using datasets from the SPHERE instrument operating at the Very Large Telescope (VLT). It significantly improves the precision-recall trade-off, notably on challenging datasets that are otherwise unusable by astronomers. The proposed approach is computationally efficient, robust to varying data quality, and well suited for large-scale observational surveys.1 Théo Bodrito, Olivier Flasseur, Julien Mairal, Jean Ponce, Maud Langlois, Anne-Marie Lagrange |
CVPR | 4 |
| 2025 | Generalized portrait quality assessment
Nicolas Chahine, Sira Ferradans, Javier Vazquez-Corral, Jean Ponce |
Pattern Recognit. Lett. | 4 |
| 2024 | Dense Optical Tracking: Connecting the DotsabstractRecent approaches to point tracking are able to recover the trajectory of any scene point through a large portion of a video despite the presence of occlusions. They are, how-ever, too slow in practice to track every point observed in a single frame in a reasonable amount of time. This paper introduces DOT, a novel, simple and efficient method for solving this problem. It first extracts a small set of tracks from key regions at motion boundaries using an off-the-shelf point tracking algorithm. Given source and target frames, DOT then computes rough initial estimates of a dense flow field and visibility mask through nearest-neighbor inter-polation, before refining them using a learnable optical flow estimator that explicitly handles occlusions and can be trained on synthetic data with ground-truth correspon-dences. We show that DOT is significantly more accurate than current optical flow techniques, outperforms sophis-ticated “universal” trackers like OmniMotion, and is on par with, or better than, the best point tracking algorithms like CoTracker while being at least two orders of magnitude faster. Quantitative and qualitative experiments with syn-thetic and real videos validate the promise of the proposed approach. Code, data, and videos showcasing the capabili-ties of our approach are available in the project webpage.11https://161ernoing.github.io/dot Guillaume Le Moing, Jean Ponce, Cordelia Schmid |
CVPR | 2 |
| 2023 | An Image Quality Assessment Dataset for PortraitsabstractYear after year, the demand for ever-better smartphone photos continues to grow, in particular in the domain of portrait photography. Manufacturers thus use perceptual quality criteria throughout the development of smartphone cameras. This costly procedure can be partially replaced by automated learning-based methods for image quality assessment (IQA). Due to its subjective nature, it is necessary to estimate and guarantee the consistency of the IQA process, a characteristic lacking in the mean opinion scores (MOS) widely used for crowdsourcing IQA. In addition, existing blind IQA (BIQA) datasets pay little attention to the difficulty of cross-content assessment, which may degrade the quality of annotations. This paper introduces PIQ23, a portrait-specific IQA dataset of 5116 images of 50 predefined scenarios acquired by 100 smartphones, covering a high variety of brands, models, and use cases. The dataset includes individuals of various genders and ethnicities who have given explicit and informed consent for their photographs to be used in public research. It is annotated by pairwise comparisons (PWC) collected from over 30 image quality experts for three image attributes: face detail preservation, face target exposure, and overall image quality. An in-depth statistical analysis of these annotations allows us to evaluate their consistency over PIQ23. Finally, we show through an extensive comparison with existing baselines that semantic information (image context) can be used to improve IQA predictions. The dataset along with the proposed statistical analysis and BIQA algorithms are available: https://github.com/DXOMARK-Research/PIQ2023 Nicolas Chahine, Ana-Stefania Calarasanu, Davide Garcia-Civiero, Théo Cayla, Sira Ferradans, Jean Ponce |
CVPR | 6 |
| 2023 | WALDO: Future Video Synthesis using Object Layer Decomposition and Parametric Flow PredictionabstractThis paper presents WALDO (WArping Layer-Decomposed Objects), a novel approach to the prediction of future video frames from past ones. Individual images are decomposed into multiple layers combining object masks and a small set of control points. The layer structure is shared across all frames in each video to build dense inter-frame connections. Complex scene motions are modeled by combining parametric geometric transformations associated with individual layers, and video synthesis is broken down into discovering the layers associated with past frames, predicting the corresponding transformations for upcoming ones and warping the associated object regions accordingly, and filling in the remaining image parts. Extensive experiments on multiple benchmarks including urban videos (Cityscapes and KITTI) and videos featuring nonrigid motions (UCF-Sports and H3.6M), show that our method consistently outperforms the state of the art by a significant margin in every case. Code, pretrained models, and video samples synthesized by our approach can be found in the project webpage.1 Guillaume Le Moing, Jean Ponce, Cordelia Schmid |
ICCV | 2 |
| 2023 | Learning Reward Functions for Robotic Manipulation by Observing HumansabstractObserving a human demonstrator manipulate objects provides a rich, scalable and inexpensive source of data for learning robotic policies. However, transferring skills from human videos to a robotic manipulator poses several challenges, not least a difference in action and observation spaces. In this work, we use unlabeled videos of humans solving a wide range of manipulation tasks to learn a task-agnostic reward function for robotic manipulation policies. Thanks to the diversity of this training data, the learned reward function sufficiently generalizes to image observations from a previously unseen robot embodiment and environment to provide a meaningful prior for directed exploration in reinforcement learning. We propose two methods for scoring states relative to a goal image: through direct temporal regression, and through distances in an embedding space obtained with time-contrastive learning. By conditioning the function on a goal image, we are able to reuse one model across a variety of tasks. Unlike prior work on leveraging human videos to teach robots, our method, Human Offline Learned Distances (HOLD) requires neither a priori data from the robot environment, nor a set of task-specific human demonstrations, nor a predefined notion of correspondence across morphologies, yet it is able to accelerate training of several manipulation tasks on a simulated robot arm compared to using only a sparse reward obtained from task completion. Minttu Alakuijala, Gabriel Dulac-Arnold, Julien Mairal, Jean Ponce, Cordelia Schmid |
ICRA | 4 |
| 2023 | A minimum swept-volume metric structure for configuration spaceabstractBorrowing elementary ideas from solid mechanics and differential geometry, this presentation shows that the volume swept by a regular solid undergoing a wide class of volume-preserving deformations induces a rather natural metric structure with well-defined and computable geodesics on its configuration space. This general result applies to concrete classes of articulated objects such as robot manipulators, and we demonstrate as a proof of concept the computation of geodesic paths for a free flying rod and planar robotic arms as well as their use in path planning with many obstacles. Yann de Mont-Marin, Jean Ponce, Jean-Paul Laumond |
ICRA | 2 |
| 2023 | Revisiting Deformable Convolution for Depth CompletionabstractDepth completion, which aims to generate high-quality dense depth maps from sparse depth maps, has attracted increasing attention in recent years. Previous work usually employs RGB images as guidance, and introduces iterative spatial propagation to refine estimated coarse depth maps. However, most of the propagation refinement methods require several iterations and suffer from a fixed receptive field, which may contain irrelevant and useless information with very sparse input. In this paper, we address these two challenges simultaneously by revisiting the idea of deformable convolution. We propose an effective architecture that leverages deformable kernel convolution as a single-pass refinement module, and empirically demonstrate its superiority. To better understand the function of deformable convolution and exploit it for depth completion, we further systematically investigate a variety of representative strategies. Our study reveals that, different from prior work, deformable convolution needs to be applied on an estimated depth map with a relatively high density for better performance. We evaluate our model on the large-scale KITTI dataset and achieve state-of-the-art level performance in both accuracy and inference speed. Our code is available at https://github.com/AlexSunNiklReDC. Xinglong Sun, Jean Ponce, Yu-Xiong Wang |
IROS | 2 |
| 2022 | Active Learning Strategies for Weakly-Supervised Object Detection
Huy V. Vo, Oriane Siméoni, Spyros Gidaris, Andrei Bursuc, Patrick Pérez, Jean Ponce |
ECCV (30) | 6 |
| 2022 | VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning
Adrien Bardes, Jean Ponce, Yann LeCun |
ICLR | 2 |
| 2022 | Assembly Planning from Observations under Physical ConstraintsabstractThis paper addresses the problem of copying an unknown assembly of primitives with known shape and appearance using information extracted from a single photograph by an off-the-shelf procedure for object detection and pose estimation. The proposed algorithm uses a simple combination of physical stability constraints, convex optimization and Monte Carlo tree search to plan assemblies as sequences of pick-and-place operations represented by STRIPS operators. It is efficient and, most importantly, robust to the errors in object detection and pose estimation unavoidable in any real robotic system. The proposed approach is demonstrated with thorough experiments on a UR5 manipulator. Thomas Chabal, Robin Strudel, Etienne Arlaud, Jean Ponce, Cordelia Schmid |
IROS | 4 |
| 2022 | VICRegL: Self-Supervised Learning of Local Visual FeaturesabstractMost recent self-supervised methods for learning image representations focus on either producing a global feature with invariance properties, or producing a set of local features. The former works best for classification tasks while the latter is best for detection and segmentation tasks. This paper explores the fundamental trade-off between learning local and global features. A new method called VICRegL is proposed that learns good global and local features simultaneously, yielding excellent performance on detection and segmentation tasks while maintaining good performance on classification tasks. Concretely, two identical branches of a standard convolutional net architecture are fed two differently distorted versions of the same image. The VICReg criterion is applied to pairs of global feature vectors. Simultaneously, the VICReg criterion is applied to pairs of local feature vectors occurring before the last pooling layer. Two local feature vectors are attracted to each other if their l2-distance is below a threshold or if their relative locations are consistent with a known geometric transformation between the two input images. We demonstrate strong performance on linear classification and segmentation transfer tasks. Code and pretrained models are publicly available at: https://github.com/facebookresearch/VICRegL Adrien Bardes, Jean Ponce, Yann LeCun |
NeurIPS | 2 |
| 2022 | Homography-Based Minimal-Case Relative Pose Estimation With Known Gravity DirectionabstractIn this paper, we propose a novel approach to two-view minimal-case relative pose problems based on homography with known gravity direction. This case is relevant to smart phones, tablets, and other camera-IMU (Inertial measurement unit) systems which have accelerometers to measure the gravity vector. We explore the rank-1 constraint on the difference between the euclidean homography matrix and the corresponding rotation, and propose an efficient two-step solution for solving both the calibrated and semi-calibrated (unknown focal length) problems. Based on the hidden variable technique, we convert the problems to the polynomial eigenvalue problems, and derive new 3.5-point, 3.5-point, 4-point solvers for two cameras such that the two focal lengths are unknown but equal, one of them is unknown, and both are unknown and possibly different, respectively. We present detailed analyses and comparisons with the existing 6- and 7-point solvers, including results with smart phone images. Yaqing Ding 0001, Jian Yang 0003, Jean Ponce, Hui Kong 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Learning Semantic Correspondence Exploiting an Object-Level PriorabstractWe address the problem of semantic correspondence, that is, establishing a dense flow field between images depicting different instances of the same object or scene category. We propose to use images annotated with binary foreground masks and subjected to synthetic geometric deformations to train a convolutional neural network (CNN) for this task. Using these masks as part of the supervisory signal provides an object-level prior for the semantic correspondence task and offers a good compromise between semantic flow methods, where the amount of training data is limited by the cost of manually selecting point correspondences, and semantic alignment ones, where the regression of a single global geometric transformation between images may be sensitive to image-specific details such as background clutter. We propose a new CNN architecture, dubbed SFNet, which implements this idea. It leverages a new and differentiable version of the argmax function for end-to-end training, with a loss that combines mask and flow consistency with smoothness terms. Experimental results demonstrate the effectiveness of our approach, which significantly outperforms the state of the art on standard benchmarks. Junghyup Lee, Dohyung Kim 0006, Wonkyung Lee, Jean Ponce, Bumsub Ham |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | High dynamic range and super-resolution from raw image burstsabstractPhotographs captured by smartphones and mid-range cameras have limited spatial resolution and dynamic range, with noisy response in underexposed regions and color artefacts in saturated areas. This paper introduces the first approach (to the best of our knowledge) to the reconstruction of highresolution, high-dynamic range color images from raw photographic bursts captured by a handheld camera with exposure bracketing. This method uses a physically-accurate model of image formation to combine an iterative optimization algorithm for solving the corresponding inverse problem with a learned image representation for robust alignment and a learned natural image prior. The proposed algorithm is fast, with low memory requirements compared to state-of-the-art learning-based approaches to image restoration, and features that are learned end to end from synthetic yet realistic data. Extensive experiments demonstrate its excellent performance with super-resolution factors of up to ×4 on real photographs taken in the wild with hand-held cameras, and high robustness to low-light conditions, noise, camera shake, and moderate object motion. Bruno Lecouat, Thomas Eboli, Jean Ponce, Julien Mairal |
ACM Trans. Graph. | 3 |
| 2021 | Localizing Objects with Self-supervised Transformers and no Labels
Oriane Siméoni, Gilles Puy, Huy V. Vo, Simon Roburin, Spyros Gidaris, Andrei Bursuc, Patrick Pérez, Renaud Marlet, Jean Ponce |
BMVC | 9 |
| 2021 | Lucas-Kanade Reloaded: End-to-End Super-Resolution from Raw Image BurstsabstractThis presentation addresses the problem of reconstructing a high-resolution image from multiple lower-resolution snapshots captured from slightly different viewpoints in space and time. Key challenges for solving this super-resolution problem include (i) aligning the input pictures with sub-pixel accuracy, (ii) handling raw (noisy) images for maximal faithfulness to native camera data, and (iii) designing/learning an image prior (regularizer) well suited to the task. We address these three challenges with a hybrid algorithm building on the insight from [45] that aliasing is an ally in this setting, with parameters that can be learned end to end, while retaining the interpretability of classical approaches to inverse problems. The effectiveness of our approach is demonstrated on synthetic and real image bursts, setting a new state of the art on several benchmarks and delivering excellent qualitative results on real raw bursts captured by smartphones and prosumer cameras. Our code is available at https://github.com/bruno-31/lkburst.git. Bruno Lecouat, Jean Ponce, Julien Mairal |
ICCV | 2 |
| 2021 | Unsupervised Layered Image Decomposition into Object PrototypesabstractWe present an unsupervised learning framework for decomposing images into layers of automatically discovered object models. Contrary to recent approaches that model image layers with autoencoder networks, we represent them as explicit transformations of a small set of prototypical images. Our model has three main components: (i) a set of object prototypes in the form of learnable images with a transparency channel, which we refer to as sprites; (ii) differentiable parametric functions predicting occlusions and transformation parameters necessary to instantiate the sprites in a given image; (iii) a layered image formation model with occlusion for compositing these instances into complete images including background. By jointly learning the sprites and occlusion/transformation predictors to reconstruct images, our approach not only yields accurate layered image decompositions, but also identifies object categories and instance parameters. We first validate our approach by providing results on par with the state of the art on standard multiobject synthetic benchmarks (Tetrominoes, Multi-dSprites, CLEVR6). We then demonstrate the applicability of our model to real images in tasks that include clustering (SVHN, GTSRB), cosegmentation (Weizmann Horse) and object discovery from unfiltered social network images. To the best of our knowledge, our approach is the first layered image decomposition algorithm that learns an explicit and shared concept of object type, and is robust enough to be applied to real images. Tom Monnier, Elliot Vincent, Jean Ponce, Mathieu Aubry |
ICCV | 3 |
| 2021 | Equality Constrained Differential Dynamic ProgrammingabstractTrajectory optimization is an important tool in task-based robot motion planning, due to its generality and convergence guarantees under some mild conditions. It is often used as a post-processing operation to smooth out trajectories that are generated by probabilistic methods or to directly control the robot motion. Unconstrained trajectory optimization problems have been well studied, and are commonly solved using Differential Dynamic Programming methods that allow for fast convergence at a relatively low computational cost. In this paper, we propose an augmented Lagrangian approach that extends these ideas to equality-constrained trajectory optimization problems, while maintaining a balance between convergence speed and numerical stability. We illustrate our contributions on various standard robotic problems and highlights their benefits compared to standard approaches. Sarah El Kazdadi, Justin Carpentier, Jean Ponce |
ICRA | 3 |
| 2021 | Online Learning and Control of Complex Dynamical Systems from Sensory InputabstractIdentifying an effective model of a dynamical system from sensory data and using it for future state prediction and control is challenging. Recent data-driven algorithms based on Koopman theory are a promising approach to this problem, but they typically never update the model once it has been identified from a relatively small set of observation, thus making long-term prediction and control difficult for realistic systems, in robotics or fluid mechanics for example. This paper introduces a novel method for learning an embedding of the state space with linear dynamics from sensory data. Unlike previous approaches, the dynamics model can be updated online and thus easily applied to systems with non-linear dynamics in the original configuration space. The proposed approach is evaluated empirically on several classical dynamical systems and sensory modalities, with good performance on long-term prediction and control. Oumayma Bounou, Jean Ponce, Justin Carpentier |
NeurIPS | 2 |
| 2021 | CCVS: Context-aware Controllable Video SynthesisabstractThis presentation introduces a self-supervised learning approach to the synthesis of new videos clips from old ones, with several new key elements for improved spatial resolution and realism: It conditions the synthesis process on contextual information for temporal continuity and ancillary information for fine control. The prediction model is doubly autoregressive, in the latent space of an autoencoder for forecasting, and in image space for updating contextual information, which is also used to enforce spatio-temporal consistency through a learnable optical flow module. Adversarial training of the autoencoder in the appearance and temporal domains is used to further improve the realism of its output. A quantizer inserted between the encoder and the transformer in charge of forecasting future frames in latent space (and its inverse inserted between the transformer and the decoder) adds even more flexibility by affording simple mechanisms for handling multimodal ancillary information for controlling the synthesis process (e.g., a few sample frames, an audio track, a trajectory in image space) and taking into account the intrinsically uncertain nature of the future by allowing multiple predictions. Experiments with an implementation of the proposed approach give very good qualitative and quantitative results on multiple tasks and standard benchmarks. Guillaume Le Moing, Jean Ponce, Cordelia Schmid |
NeurIPS | 2 |
| 2021 | Large-Scale Unsupervised Object DiscoveryabstractExisting approaches to unsupervised object discovery (UOD) do not scale up to large datasets without approximations that compromise their performance. We propose a novel formulation of UOD as a ranking problem, amenable to the arsenal of distributed methods available for eigenvalue problems and link analysis. Through the use of self-supervised features, we also demonstrate the first effective fully unsupervised pipeline for UOD. Extensive experiments on COCO~\cite{Lin2014cocodataset} and OpenImages~\cite{openimages} show that, in the single-object discovery setting where a single prominent object is sought in each image, the proposed LOD (Large-scale Object Discovery) approach is on par with, or better than the state of the art for medium-scale datasets (up to 120K images), and over 37\% better than the only other algorithms capable of scaling up to 1.7M images. In the multi-object discovery setting where multiple objects are sought in each image, the proposed LOD is over 14\% better in average precision (AP) than all other methods for datasets ranging from 20K to 1.7M images. Using self-supervised features, we also show that the proposed method obtains state-of-the-art UOD performance on OpenImages. Huy V. Vo, Elena Sizikova, Cordelia Schmid, Patrick Pérez, Jean Ponce |
NeurIPS | 5 |
| 2021 | Deformable Kernel Networks for Joint Image Filtering
Jean Ponce, Bumsub Ham |
Int. J. Comput. Vis. | 2 |
| 2020 | Minimal Solutions to Relative Pose Estimation From Two Views Sharing a Common Direction With Unknown Focal LengthabstractWe propose minimal solutions to relative pose estimation problem from two views sharing a common direction with unknown focal length. This is relevant for cameras equipped with an IMU (inertial measurement unit), e.g., smart phones, tablets. Similar to the 6-point algorithm for two cameras with unknown but equal focal lengths and 7-point algorithm for two cameras with different and unknown focal lengths, we derive new 4- and 5-point algorithms for these two cases, respectively. The proposed algorithms can cope with coplanar points, which is a degenerate configuration for these 6- and 7-point counterparts. We present a detailed analysis and comparisons with the state of the art. Experimental results on both synthetic data and real images from a smart phone demonstrate the usefulness of the proposed algorithms. Yaqing Ding 0001, Jian Yang 0003, Jean Ponce, Hui Kong 0001 |
CVPR | 3 |
| 2020 | End-to-end Interpretable Learning of Non-blind Image Deblurring
Thomas Eboli, Jian Sun 0009, Jean Ponce |
ECCV (17) | 3 |
| 2020 | Fully Trainable and Interpretable Non-local Sparse Models for Image Restoration
Bruno Lecouat, Jean Ponce, Julien Mairal |
ECCV (22) | 2 |
| 2020 | Learning to Compose Hypercolumns for Visual Correspondence
Juhong Min, Jongmin Lee 0005, Jean Ponce, Minsu Cho |
ECCV (15) | 3 |
| 2020 | Toward Unsupervised, Multi-object Discovery in Large-Scale Image Collections
Huy V. Vo, Patrick Pérez, Jean Ponce |
ECCV (23) | 3 |
| 2020 | A Flexible Framework for Designing Trainable Priors with Adaptive Smoothing and Game EncodingabstractWe introduce a general framework for designing and training neural network layers whose forward passes can be interpreted as solving non-smooth convex optimization problems, and whose architectures are derived from an optimization algorithm. We focus on convex games, solved by local agents represented by the nodes of a graph and interacting through regularization functions. This approach is appealing for solving imaging problems, as it allows the use of classical image priors within deep models that are trainable end to end. The priors used in this presentation include variants of total variation, Laplacian regularization, bilateral filtering, sparse coding on learned dictionaries, and non-local self similarities. Our models are fully interpretable as well as parameter and data efficient. Our experiments demonstrate their effectiveness on a large diversity of tasks ranging from image denoising and compressed sensing for fMRI to dense stereo matching. Bruno Lecouat, Jean Ponce, Julien Mairal |
NeurIPS | 2 |
| 2019 | SFNet: Learning Object-Aware Semantic CorrespondenceabstractWe address the problem of semantic correspondence, that is, establishing a dense flow field between images depicting different instances of the same object or scene category. We propose to use images annotated with binary foreground masks and subjected to synthetic geometric deformations to train a convolutional neural network (CNN) for this task. Using these masks as part of the supervisory signal offers a good compromise between semantic flow methods, where the amount of training data is limited by the cost of manually selecting point correspondences, and semantic alignment ones, where the regression of a single global geometric transformation between images may be sensitive to image-specific details such as background clutter. We propose a new CNN architecture, dubbed SFNet, which implements this idea. It leverages a new and differentiable version of the argmax function for end-to-end training, with a loss that combines mask and flow consistency with smoothness terms. Experimental results demonstrate the effectiveness of our approach, which significantly outperforms the state of the art on standard benchmarks. Junghyup Lee, Dohyung Kim 0006, Jean Ponce, Bumsub Ham |
CVPR | 3 |
| 2019 | Coordinate-Free Carlsson-Weinshall Duality and Relative Multi-View GeometryabstractWe present a coordinate-free description of Carlsson-Weinshall duality between scene points and camera pinholes and use it to derive a new characterization of primal/dual multi-view geometry. In the case of three views, a particular set of reduced trilinearities provide a novel parameterization of camera geometry that, unlike existing ones, is subject only to very simple internal constraints. These trilinearities lead to new "quasi-linear" algorithms for primal and dual structure from motion. We include some preliminary experiments with real and synthetic data. Matthew Trager, Martial Hebert, Jean Ponce |
CVPR | 3 |
| 2019 | Unsupervised Image Matching and Object Discovery as OptimizationabstractLearning with complete or partial supervision is power- ful but relies on ever-growing human annotation efforts. As a way to mitigate this serious problem, as well as to serve specific applications, unsupervised learning has emerged as an important field of research. In computer vision, unsu- pervised learning comes in various guises. We focus here on the unsupervised discovery and matching of object cate- gories among images in a collection, following the work of Cho et al. [12]. We show that the original approach can be reformulated and solved as a proper optimization problem. Experiments on several benchmarks establish the merit of our approach. Huy V. Vo, Francis R. Bach, Minsu Cho, Kai Han 0001, Yann LeCun, Patrick Pérez, Jean Ponce |
CVPR | 7 |
| 2019 | An Efficient Solution to the Homography-Based Relative Pose Problem With a Common Reference DirectionabstractIn this paper, we propose a novel approach to two-view minimal-case relative pose problems based on homography with a common reference direction. We explore the rank-1 constraint on the difference between the Euclidean homography matrix and the corresponding rotation, and propose an efficient two-step solution for solving both the calibrated and partially calibrated (unknown focal length) problems. We derive new 3.5-point, 3.5-point, 4-point solvers for two cameras such that the two focal lengths are unknown but equal, one of them is unknown, and both are unknown and possibly different, respectively. We present detailed analyses and comparisons with existing 6 and 7-point solvers, including results with smart phone images. Yaqing Ding 0001, Jian Yang 0003, Jean Ponce, Hui Kong 0001 |
ICCV | 3 |
| 2019 | Hyperpixel Flow: Semantic Correspondence With Multi-Layer Neural FeaturesabstractEstablishing visual correspondences under large intra-class variations requires analyzing images at different levels, from features linked to semantics and context to local patterns, while being invariant to instance-specific details. To tackle these challenges, we represent images by “hyperpixels” that leverage a small number of relevant features selected among early to late layers of a convolutional neural network. Taking advantage of the condensed features of hyperpixels, we develop an effective real-time matching algorithm based on Hough geometric voting. The proposed method, hyperpixel flow, sets a new state of the art on three standard benchmarks as well as a new dataset, SPair-71k, which contains a significantly larger number of image pairs than existing datasets, with more accurate and richer annotations for in-depth analysis. Juhong Min, Jongmin Lee 0005, Jean Ponce, Minsu Cho |
ICCV | 3 |
| 2019 | Build your own hybrid thermal/EO camera for autonomous vehicleabstractIn this work, we propose a novel paradigm to design a hybrid thermal/EO (Electro-Optical or visible-light) camera, whose thermal and RGB frames are pixel-wisely aligned and temporally synchronized. Compared with the existing schemes, we innovate in three ways in order to make it more compact in dimension, and thus more practical and extendable for real-world applications. The first is a redesign of the structure layout of the thermal and EO cameras. The second is on obtaining a pixel-wise spatial registration of the thermal and RGB frames by a coarse mechanical adjustment and a fine alignment through a constant homography warping. The third innovation is on extending one single hybrid camera to a hybrid camera array, through which we can obtain wide-view spatially aligned thermal, RGB and disparity images simultaneously. The experimental results show that the average error of spatial-alignment of two image modalities can be less than one pixel. Yigong Zhang, Shuo Gu, Yubin Guo, Minghao Liu 0003, Zezhou Sun, Zhixing Hou, Ying Wang 0007, Jian Yang 0003, Jean Ponce, Hui Kong 0001 |
ICRA | 11 |
| 2018 | On the Solvability of Viewing Graphs
Matthew Trager, Brian Osserman, Jean Ponce |
ECCV (16) | 3 |
| 2018 | Dijkstra Model for Stereo-Vision Based Road Detection: A Non-Parametric MethodabstractThis paper proposes a new method for detecting a road from a stereo pair of images. First, the horizon is accurately estimated by a robust, weighted-sampling RANSAC-like method in the improved v-disparity map. The vanishing point of the road region is located using both the horizon information and road flatness constraints. Then it is used as the source node of a weighted graph formed by the pixels of the left stereo-image and their adjacency relationships. The weight of each edge measures the inconsistency of adjacent pixels, and is computed using both the gray-scale and disparity information. Detecting road borders is thus reduced to finding two shortest paths from the source node to the bottom row of the image by the Dijkstra algorithm. The proposed method has been tested on 2621 image pairs of different road scenes from the KITTI dataset. Our experiments demonstrate that this training free approach detects horizon, vanishing point, and road region accurately and robustly, and compares favorably with the state of the art on the KITTI benchmark. Yigong Zhang, Jian Yang 0003, Jean Ponce, Hui Kong 0001 |
ICRA | 3 |
| 2018 | Robust Guided Image Filtering Using Nonconvex PotentialsabstractFiltering images using a guidance signal, a process called guided or joint image filtering, has been used in various tasks in computer vision and computational photography, particularly for noise reduction and joint upsampling. This uses an additional guidance signal as a structure prior, and transfers the structure of the guidance signal to an input image, restoring noisy or altered image structure. The main drawbacks of such a data-dependent framework are that it does not consider structural differences between guidance and input images, and that it is not robust to outliers. We propose a novel SD (for static/dynamic) filter to address these problems in a unified framework, and jointly leverage structural information from guidance and input images. Guided image filtering is formulated as a nonconvex optimization problem, which is solved by the majorize-minimization algorithm. The proposed algorithm converges quickly while guaranteeing a local minimum. The SD filter effectively controls the underlying image structure at different scales, and can handle a variety of types of data from different sensors. It is robust to outliers and other artifacts such as gradient reversal and global intensity shift, and has good edge-preserving smoothing properties. We demonstrate the flexibility and effectiveness of the proposed SD filter in a variety of applications, including depth upsampling, scale-space filtering, texture removal, flash/non-flash denoising, and RGB/NIR denoising. Bumsub Ham, Minsu Cho, Jean Ponce |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Proposal Flow: Semantic Correspondences from Object ProposalsabstractFinding image correspondences remains a challenging problem in the presence of intra-class variations and large changes in scene layout. Semantic flow methods are designed to handle images depicting different instances of the same object or scene category. We introduce a novel approach to semantic flow, dubbed proposal flow, that establishes reliable correspondences using object proposals. Unlike prevailing semantic flow approaches that operate on pixels or regularly sampled local regions, proposal flow benefits from the characteristics of modern object proposals, that exhibit high repeatability at multiple scales, and can take advantage of both local and geometric consistency constraints among proposals. We also show that the corresponding sparse proposal flow can effectively be transformed into a conventional dense flow field. We introduce two new challenging datasets that can be used to evaluate both general semantic flow techniques and region-based approaches such as proposal flow. We use these benchmarks to compare different matching algorithms, object proposals, and region features within proposal flow, to the state of the art in semantic flow. This comparison, along with experiments on standard datasets, demonstrates that proposal flow significantly outperforms existing semantic flow methods in various settings. Bumsub Ham, Minsu Cho, Cordelia Schmid, Jean Ponce |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2018 | When Dijkstra Meets Vanishing Point: A Stereo Vision Approach for Road DetectionabstractIn this paper, we propose a vanishing-point constrained Dijkstra road model for road detection in a stereo-vision paradigm. First, the stereo-camera is used to generate the u- and v-disparity maps of road image, from which the horizon can be extracted. With the horizon and ground region constraints, we can robustly locate the vanishing point of road region. Second, a weighted graph is constructed using all pixels of the image, and the detected vanishing point is treated as the source node of the graph. By computing a vanishing-point constrained Dijkstra minimum-cost map, where both disparity and gradient of gray image are used to calculate cost between two neighbor pixels, the problem of detecting road borders in image is transformed into that of finding two shortest paths that originate from the vanishing point to two pixels in the last row of image. The proposed approach has been implemented and tested over 2600 grayscale images of different road scenes in the KITTI data set. The experimental results demonstrate that this training-free approach can detect horizon, vanishing point, and road regions very accurately and robustly. It can achieve promising performance. Yigong Zhang, Yingna Su, Jian Yang 0003, Jean Ponce, Hui Kong 0001 |
IEEE Trans. Image Process. | 4 |
| 2017 | Kernel Square-Loss Exemplar Machines for Image RetrievalabstractZepeda and Perez [41] have recently demonstrated the promise of the exemplar SVM (ESVM) as a feature encoder for image retrieval. This paper extends this approach in several directions: We first show that replacing the hinge loss by the square loss in the ESVM cost function significantly reduces encoding time with negligible effect on accuracy. We call this model square-loss exemplar machine, or SLEM. We then introduce a kernelized SLEM which can be implemented efficiently through low-rank matrix decomposition, and displays improved performance. Both SLEM variants exploit the fact that the negative examples are fixed, so most of the SLEM computational complexity is relegated to an offline process independent of the positive examples. Our experiments establish the performance and computational advantages of our approach using a large array of base features and standard image retrieval datasets. Rafael S. Rezende, Joaquin Zepeda, Jean Ponce, Francis R. Bach, Patrick Pérez |
CVPR | 3 |
| 2017 | General Models for Rational Cameras and the Case of Two-Slit ProjectionsabstractThe rational camera model recently introduced in [18] provides a general methodology for studying abstract nonlinear imaging systems and their multi-view geometry. This paper builds on this framework to study physical realizations of rational cameras. More precisely, we give an explicit account of the mapping between between physical visual rays and image points (missing in the original description), which allows us to give simple analytical expressions for direct and inverse projections. We also consider primitive camera models, that are orbits under the action of various projective transformations, and lead to a general notion of intrinsic parameters. The methodology is general, but it is illustrated concretely by an in-depth study of two-slit cameras, that we model using pairs of linear projections. This simple analytical form allows us to describe models for the corresponding primitive cameras, to introduce intrinsic parameters with a clear geometric meaning, and to define an epipolar tensor characterizing two-view correspondences. In turn, this leads to new algorithms for structure from motion and self-calibration. Matthew Trager, Bernd Sturmfels, John F. Canny, Martial Hebert, Jean Ponce |
CVPR | 5 |
| 2017 | SCNet: Learning Semantic CorrespondenceabstractThis paper addresses the problem of establishing semantic correspondences between images depicting different instances of the same object or scene category. Previous approaches focus on either combining a spatial regularizer with hand-crafted features, or learning a correspondence model for appearance only. We propose instead a convolutional neural network architecture, called SCNet, for learning a geometrically plausible model for semantic correspondence. SCNet uses region proposals as matching primitives, and explicitly incorporates geometric consistency in its loss function. It is trained on image pairs obtained from the PASCAL VOC 2007 keypoint dataset, and a comparative evaluation on several standard benchmarks demonstrates that the proposed approach substantially outperforms both recent deep learning architectures and previous methods based on hand-crafted features. Kai Han 0001, Rafael S. Rezende, Bumsub Ham, Kwan-Yee Kenneth Wong, Minsu Cho, Cordelia Schmid, Jean Ponce |
ICCV | 7 |
| 2016 | Proposal FlowabstractFinding image correspondences remains a challenging problem in the presence of intra-class variations and large changes in scene layout. Semantic flow methods are designed to handle images depicting different instances of the same object or scene category. We introduce a novel approach to semantic flow, dubbed proposal flow, that establishes reliable correspondences using object proposals. Unlike prevailing semantic flow approaches that operate on pixels or regularly sampled local regions, proposal flow benefits from the characteristics of modern object proposals, that exhibit high repeatability at multiple scales, and can take advantage of both local and geometric consistency constraints among proposals. We also show that proposal flow can effectively be transformed into a conventional dense flow field. We introduce a new dataset that can be used to evaluate both general semantic flow techniques and region-based approaches such as proposal flow. We use this benchmark to compare different matching algorithms, object proposals, and region features within proposal flow, to the state of the art in semantic flow. This comparison, along with experiments on standard datasets, demonstrates that proposal flow significantly outperforms existing semantic flow methods in various settings. Bumsub Ham, Minsu Cho, Cordelia Schmid, Jean Ponce |
CVPR | 4 |
| 2016 | Consistency of Silhouettes and Their DualsabstractSilhouettes provide rich information on three-dimensional shape, since the intersection of the associated visual cones generates the "visual hull", which encloses and approximates the original shape. However, not all silhouettes can actually be projections of the same object in space: this simple observation has implications in object recognition and multi-view segmentation, and has been (often implicitly) used as a basis for camera calibration. In this paper, we investigate the conditions for multiple silhouettes, or more generally arbitrary closed image sets, to be geometrically "consistent". We present this notion as a natural generalization of traditional multi-view geometry, which deals with consistency for points. After discussing some general results, we present a "dual" formulation for consistency, that gives conditions for a family of planar sets to be sections of the same object. Finally, we introduce a more general notion of silhouette "compatibility" under partial knowledge of the camera projections, and point out some possible directions for future research. Matthew Trager, Martial Hebert, Jean Ponce |
CVPR | 3 |
| 2016 | Learning Dictionary of Discriminative Part Detectors for Image Categorization and Cosegmentation
Jian Sun 0009, Jean Ponce |
Int. J. Comput. Vis. | 2 |
| 2016 | Trinocular Geometry Revisited
Matthew Trager, Jean Ponce, Martial Hebert |
Int. J. Comput. Vis. | 2 |
| 2015 | Unsupervised object discovery and localization in the wild: Part-based matching with bottom-up region proposalsabstractThis paper addresses unsupervised discovery and localization of dominant objects from a noisy image collection with multiple object classes. The setting of this problem is fully unsupervised, without even image-level annotations or any assumption of a single dominant class. This is far more general than typical colocalization, cosegmentation, or weakly-supervised localization tasks. We tackle the discovery and localization problem using a part-based region matching approach: We use off-the-shelf region proposals to form a set of candidate bounding boxes for objects and object parts. These regions are efficiently matched across images using a probabilistic Hough transform that evaluates the confidence for each candidate correspondence considering both appearance and spatial consistency. Dominant objects are discovered and localized by comparing the scores of candidate regions and selecting those that stand out over other regions containing them. Extensive experimental evaluations on standard benchmarks demonstrate that the proposed approach significantly outperforms the current state of the art in colocalization, and achieves robust object discovery in challenging mixed-class datasets. Minsu Cho, Suha Kwak, Cordelia Schmid, Jean Ponce |
CVPR | 4 |
| 2015 | Robust image filtering using joint static and dynamic guidanceabstractRegularizing images under a guidance signal has been used in various tasks in computer vision and computational photography, particularly for noise reduction and joint upsampling. The aim is to transfer fine structures of guidance signals to input images, restoring noisy or altered structures. One of main drawbacks in such a data-dependent framework is that it does not handle differences in structure between guidance and input images. We address this problem by jointly leveraging structural information of guidance and input images. Image filtering is formulated as a nonconvex optimization problem, which is solved by the majorization-minimization algorithm. The proposed algorithm converges quickly while guaranteeing a local minimum. It effectively controls image structures at different scales and can handle a variety of types of data from different sensors. We demonstrate the flexibility and effectiveness of our model in several applications including depth super-resolution, scale-space filtering, texture removal, flash/non-flash denoising, and RGB/NIR denoising. Bumsub Ham, Minsu Cho, Jean Ponce |
CVPR | 3 |
| 2015 | Learning a convolutional neural network for non-uniform motion blur removalabstractIn this paper, we address the problem of estimating and removing non-uniform motion blur from a single blurry image. We propose a deep learning approach to predicting the probabilistic distribution of motion blur at the patch level using a convolutional neural network (CNN). We further extend the candidate set of motion kernels predicted by the CNN using carefully designed image rotations. A Markov random field model is then used to infer a dense non-uniform motion blur field enforcing motion smoothness. Finally, motion blur is removed by a non-uniform deblurring model using patch-level image prior. Experimental evaluations show that our approach can effectively estimate and remove complex non-uniform motion blur that is not handled well by previous approaches. Jian Sun 0009, Wenfei Cao, Zongben Xu, Jean Ponce |
CVPR | 4 |
| 2015 | Weakly-Supervised Alignment of Video with TextabstractSuppose that we are given a set of videos, along with natural language descriptions in the form of multiple sentences (e.g., manual annotations, movie scripts, sport summaries etc.), and that these sentences appear in the same temporal order as their visual counterparts. We propose in this paper a method for aligning the two modalities, i.e., automatically providing a time (frame) stamp for every sentence. Given vectorial features for both video and text, this can be cast as a temporal assignment problem, with an implicit linear mapping between the two feature modalities. We formulate this problem as an integer quadratic program, and solve its continuous convex relaxation using an efficient conditional gradient algorithm. Several rounding procedures are proposed to construct the final integer solution. After demonstrating significant improvements over the state of the art on the related task of aligning video with symbolic labels [7], we evaluate our method on a challenging dataset of videos with associated textual descriptions [37], and explore bag-of-words and continuous representations for text. Piotr Bojanowski, Rémi Lajugie, Edouard Grave, Francis R. Bach, Ivan Laptev, Jean Ponce, Cordelia Schmid |
ICCV | 6 |
| 2015 | Unsupervised Object Discovery and Tracking in Video CollectionsabstractThis paper addresses the problem of automatically localizing dominant objects as spatio-temporal tubes in a noisy collection of videos with minimal or even no supervision. We formulate the problem as a combination of two complementary processes: discovery and tracking. The first one establishes correspondences between prominent regions across videos, and the second one associates similar object regions within the same video. Interestingly, our algorithm also discovers the implicit topology of frames associated with instances of the same object class across different videos, a role normally left to supervisory information in the form of class labels in conventional image and video understanding methods. Indeed, as demonstrated by our experiments, our method can handle video collections featuring multiple object classes, and substantially outperforms the state of the art in colocalization, even though it tackles a broader problem with much less supervision. Suha Kwak, Minsu Cho, Ivan Laptev, Jean Ponce, Cordelia Schmid |
ICCV | 4 |
| 2015 | The Joint Image HandbookabstractGiven multiple perspective photographs, point correspondences form the "joint image", effectively a replica of three dimensional space distributed across its two-dimensional projections. This set can be characterized by multilinear equations over image coordinates, such as epipolar and trifocal constraints. We revisit in this paper the geometric and algebraic properties of the joint image, and address fundamental questions such as how many and which multilinearities are necessary and/or sufficient to determine camera geometry and/or image correspondences. The new theoretical results in this paper answer these questions in a very general setting and, in turn, are intended to serve as a "handbook" reference about multilinearities for practitioners. Matthew Trager, Martial Hebert, Jean Ponce |
ICCV | 3 |
| 2014 | Finding Matches in a Haystack: A Max-Pooling Strategy for Graph Matching in the Presence of OutliersabstractA major challenge in real-world feature matching problems is to tolerate the numerous outliers arising in typical visual tasks. Variations in object appearance, shape, and structure within the same object class make it harder to distinguish inliers from outliers due to clutters. In this paper, we propose a max-pooling approach to graph matching, which is not only resilient to deformations but also remarkably tolerant to outliers. The proposed algorithm evaluates each candidate match using its most promising neighbors, and gradually propagates the corresponding scores to update the neighbors. As final output, it assigns a reliable score to each match together with its supporting neighbors, thus providing contextual information for further verification. We demonstrate the robustness and utility of our method with synthetic and real image experiments. Minsu Cho, Jian Sun 0009, Olivier Duchenne, Jean Ponce |
CVPR | 4 |
| 2014 | Trinocular Geometry RevisitedabstractWhen do the visual rays associated with triplets of point correspondences converge, that is, intersect in a common point? Classical models of trinocular geometry based on the fundamental matrices and trifocal tensor associated with the corresponding cameras only provide partial answers to this fundamental question, in large part because of underlying, but seldom explicit, general configuration assumptions. This paper uses elementary tools from projective line geometry to provide necessary and sufficient geometric and analytical conditions for convergence in terms of transversals to triplets of visual rays, without any such assumptions. In turn, this yields a novel and simple minimal parameterization of trinocular geometry for cameras with non-collinear or collinear pinholes. Jean Ponce, Martial Hebert |
CVPR | 1 |
| 2014 | Weakly Supervised Action Labeling in Videos under Ordering Constraints
Piotr Bojanowski, Rémi Lajugie, Francis R. Bach, Ivan Laptev, Jean Ponce, Cordelia Schmid, Josef Sivic |
ECCV (5) | 5 |
| 2014 | On Image Contours of Projective Shapes
Jean Ponce, Martial Hebert |
ECCV (4) | 1 |
| 2013 | Learning to Estimate and Remove Non-uniform Image BlurabstractThis paper addresses the problem of restoring images subjected to unknown and spatially varying blur caused by defocus or linear (say, horizontal) motion. The estimation of the global (non-uniform) image blur is cast as a multi-label energy minimization problem. The energy is the sum of unary terms corresponding to learned local blur estimators, and binary ones corresponding to blur smoothness. Its global minimum is found using Ishikawa's method by exploiting the natural order of discretized blur values for linear motions and defocus. Once the blur has been estimated, the image is restored using a robust (non-uniform) deblurring algorithm based on sparse regularization with global image statistics. The proposed algorithm outputs both a segmentation of the image into uniform-blur layers and an estimate of the corresponding sharp image. We present qualitative results on real images, and use synthetic data to quantitatively compare our approach to the publicly available implementation of Chakrabarti~et al. Florent Couzinie-Devy, Jian Sun 0009, Karteek Alahari, Jean Ponce |
CVPR | 4 |
| 2013 | Finding Actors and Actions in MoviesabstractWe address the problem of learning a joint model of actors and actions in movies using weak supervision provided by scripts. Specifically, we extract actor/action pairs from the script and use them as constraints in a discriminative clustering framework. The corresponding optimization problem is formulated as a quadratic program under linear constraints. People in video are represented by automatically extracted and tracked faces together with corresponding motion features. First, we apply the proposed framework to the task of learning names of characters in the movie and demonstrate significant improvements over previous methods used for this task. Second, we explore the joint actor/action constraint and show its advantage for weakly supervised action learning. We validate our method in the challenging setting of localizing and recognizing characters and their actions in feature length movies Casablanca and American Beauty. Piotr Bojanowski, Francis R. Bach, Ivan Laptev, Jean Ponce, Cordelia Schmid, Josef Sivic |
ICCV | 4 |
| 2013 | Learning Graphs to MatchabstractMany tasks in computer vision are formulated as graph matching problems. Despite the NP-hard nature of the problem, fast and accurate approximations have led to significant progress in a wide range of applications. Learning graph models from observed data, however, still remains a challenging issue. This paper presents an effective scheme to parameterize a graph model, and learn its structural attributes for visual object matching. For this, we propose a graph representation with histogram-based attributes, and optimize them to increase the matching accuracy. Experimental evaluations on synthetic and real image datasets demonstrate the effectiveness of our approach, and show significant improvement in matching accuracy over graphs with pre-defined structures. Minsu Cho, Karteek Alahari, Jean Ponce |
ICCV | 3 |
| 2013 | Learning Discriminative Part Detectors for Image Classification and CosegmentationabstractIn this paper, we address the problem of learning discriminative part detectors from image sets with category labels. We propose a novel latent SVM model regularized by group sparsity to learn these part detectors. Starting from a large set of initial parts, the group sparsity regularizer forces the model to jointly select and optimize a set of discriminative part detectors in a max-margin framework. We propose a stochastic version of a proximal algorithm to solve the corresponding optimization problem. We apply the proposed method to image classification and co segmentation, and quantitative experiments with standard benchmarks show that it matches or improves upon the state of the art. Jian Sun 0009, Jean Ponce |
ICCV | 2 |
| 2012 | Multi-class cosegmentationabstractBottom-up, fully unsupervised segmentation remains a daunting challenge for computer vision. In the cosegmentation context, on the other hand, the availability of multiple images assumed to contain instances of the same object classes provides a weak form of supervision that can be exploited by discriminative approaches. Unfortunately, most existing algorithms are limited to a very small number of images and/or object classes (typically two of each). This paper proposes a novel energy-minimization approach to cosegmentation that can handle multiple classes and a significantly larger number of images. The proposed cost function combines spectral- and discriminative-clustering terms, and it admits a probabilistic interpretation. It is optimized using an efficient EM method, initialized using a convex quadratic approximation of the energy. Comparative experiments show that the proposed approach matches or improves the state of the art on several standard datasets. Armand Joulin, Francis R. Bach, Jean Ponce |
CVPR | 3 |
| 2012 | Non-uniform Deblurring for Shaken Images
Oliver Whyte, Josef Sivic, Andrew Zisserman, Jean Ponce |
Int. J. Comput. Vis. | 4 |
| 2012 | Task-Driven Dictionary LearningabstractModeling data with linear combinations of a few elements from a learned dictionary has been the focus of much recent research in machine learning, neuroscience, and signal processing. For signals such as natural images that admit such sparse representations, it is now well established that these models are well suited to restoration tasks. In this context, learning the dictionary amounts to solving a large-scale matrix factorization problem, which can be done efficiently with classical optimization tools. The same approach has also been used for learning features from data for other purposes, e.g., image classification, but tuning the dictionary in a supervised way for these tasks has proven to be more difficult. In this paper, we present a general formulation for supervised dictionary learning adapted to a wide variety of tasks, and present an efficient algorithm for solving the corresponding optimization problem. Experiments on handwritten digit classification, digital art identification, nonlinear inverse image problems, and compressed sensing demonstrate that our approach is effective in large-scale settings, and is well suited to supervised and semi-supervised classification, as well as regression tasks for data that admit sparse representations. Julien Mairal, Francis R. Bach, Jean Ponce |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | Sparse image representation with epitomesabstractSparse coding, which is the decomposition of a vector using only a few basis elements, is widely used in machine learning and image processing. The basis set, also called dictionary, is learned to adapt to specific data. This approach has proven to be very effective in many image processing tasks. Traditionally, the dictionary is an unstructured “flat” set of atoms. In this paper, we study structured dictionaries which are obtained from an epitome, or a set of epitomes. The epitome is itself a small image, and the atoms are all the patches of a chosen size inside this image. This considerably reduces the number of parameters to learn and provides sparse image decompositions with shift-invariance properties. We propose a new formulation and an algorithm for learning the structured dictionaries associated with epitomes, and illustrate their use in image de-noising tasks. Louise Benoît, Julien Mairal, Francis R. Bach, Jean Ponce |
CVPR | 4 |
| 2011 | Ask the locals: Multi-way local pooling for image recognitionabstractInvariant representations in object recognition systems are generally obtained by pooling feature vectors over spatially local neighborhoods. But pooling is not local in the feature vector space, so that widely dissimilar features may be pooled together if they are in nearby locations. Recent approaches rely on sophisticated encoding methods and more specialized codebooks (or dictionaries), e.g., learned on subsets of descriptors which are close in feature space, to circumvent this problem. In this work, we argue that a common trait found in much recent work in image recognition or retrieval is that it leverages locality in feature space on top of purely spatial locality. We propose to apply this idea in its simplest form to an object recognition system based on the spatial pyramid framework, to increase the performance of small dictionaries with very little added engineering. State-of-the-art results on several object recognition benchmarks show the promise of this approach. Y-Lan Boureau, Nicolas Le Roux, Francis R. Bach, Jean Ponce, Yann LeCun |
ICCV | 4 |
| 2011 | A graph-matching kernel for object categorizationabstractThis paper addresses the problem of category-level image classification. The underlying image model is a graph whose nodes correspond to a dense set of regions, and edges reflect the underlying grid structure of the image and act as springs to guarantee the geometric consistency of nearby regions during matching. A fast approximate algorithm for matching the graphs associated with two images is presented. This algorithm is used to construct a kernel appropriate for SVM-based image classification, and experiments with the Caltech 101, Caltech 256, and Scenes datasets demonstrate performance that matches or exceeds the state of the art for methods using a single type of features. Olivier Duchenne, Armand Joulin, Jean Ponce |
ICCV | 3 |
| 2011 | A Tensor-Based Algorithm for High-Order Graph MatchingabstractThis paper addresses the problem of establishing correspondences between two sets of visual features using higher order constraints instead of the unary or pairwise ones used in classical methods. Concretely, the corresponding hypergraph matching problem is formulated as the maximization of a multilinear objective function over all permutations of the features. This function is defined by a tensor representing the affinity between feature tuples. It is maximized using a generalization of spectral techniques where a relaxed problem is first solved by a multidimensional power method and the solution is then projected onto the closest assignment matrix. The proposed approach has been implemented, and it is compared to state-of-the-art algorithms on both synthetic and real data. Olivier Duchenne, Francis R. Bach, In-So Kweon, Jean Ponce |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2010 | Sparse coding and dictionary learning for image understanding
Jean Ponce |
BMVC | 1 |
| 2010 | Admissible linear map models of linear camerasabstractThis paper presents a complete analytical characterization of a large class of central and non-central imaging devices dubbed linear cameras by Ponce. Pajdla has shown that a subset of these, the oblique cameras, can be modelled by a certain type of linear map. We give here a full tabulation of all admissible maps that induce cameras in the general sense of Grossberg and Nayar, and show that these cameras are exactly the linear ones. Combining these two models with a new notion of intrinsic parameters and normalized coordinates for linear cameras allows us to give simple analytical formulas for direct and inverse projections. We also show that the epipolar geometry of any two linear cameras can be characterized by a fundamental matrix whose size is at most 6 × 6 when the cameras are uncalibrated, or by an essential matrix of size at most 4 × 4 when their internal parameters are known. Similar results hold for trinocular constraints. Guillaume Batog, Xavier Goaoc, Jean Ponce |
CVPR | 3 |
| 2010 | Learning mid-level features for recognitionabstractMany successful models for scene or object recognition transform low-level descriptors (such as Gabor filter responses, or SIFT descriptors) into richer representations of intermediate complexity. This process can often be broken down into two steps: (1) a coding step, which performs a pointwise transformation of the descriptors into a representation better adapted to the task, and (2) a pooling step, which summarizes the coded features over larger neighborhoods. Several combinations of coding and pooling schemes have been proposed in the literature. The goal of this paper is threefold. We seek to establish the relative importance of each step of mid-level feature extraction through a comprehensive cross evaluation of several types of coding modules (hard and soft vector quantization, sparse coding) and pooling schemes (by taking the average, or the maximum), which obtains state-of-the-art performance or better on several recognition benchmarks. We show how to improve the best performing coding scheme by learning a supervised discriminative dictionary for sparse coding. We provide theoretical and empirical insight into the remarkable performance of max pooling. By teasing apart components shared by modern mid-level feature extractors, our approach aims to facilitate the design of better recognition architectures. Y-Lan Boureau, Francis R. Bach, Yann LeCun, Jean Ponce |
CVPR | 4 |
| 2010 | Discriminative clustering for image co-segmentationabstractPurely bottom-up, unsupervised segmentation of a single image into foreground and background regions remains a challenging task for computer vision. Co-segmentation is the problem of simultaneously dividing multiple images into regions (segments) corresponding to different object classes. In this paper, we combine existing tools for bottom-up image segmentation such as normalized cuts, with kernel methods commonly used in object recognition. These two sets of techniques are used within a discriminative clustering framework: the goal is to assign foreground/background labels jointly to all images, so that a supervised classifier trained with these labels leads to maximal separation of the two classes. In practice, we obtain a combinatorial optimization problem which is relaxed to a continuous convex optimization problem, that can itself be solved efficiently for up to dozens of images. We illustrate the proposed method on images with very similar foreground objects, as well as on more challenging problems with objects with higher intra-class variations. Armand Joulin, Francis R. Bach, Jean Ponce |
CVPR | 3 |
| 2010 | Non-uniform deblurring for shaken imagesabstractBlur from camera shake is mostly due to the 3D rotation of the camera, resulting in a blur kernel that can be significantly non-uniform across the image. However, most current deblurring methods model the observed image as a convolution of a sharp image with a uniform blur kernel. We propose a new parametrized geometric model of the blurring process in terms of the rotational velocity of the camera during exposure. We apply this model to two different algorithms for camera shake removal: the first one uses a single blurry image (blind deblurring), while the second one uses both a blurry image and a sharp but noisy image of the same scene. We show that our approach makes it possible to model and remove a wider class of blurs than previous approaches, including uniform blur as a special case, and demonstrate its effectiveness with experiments on real images. Oliver Whyte, Josef Sivic, Andrew Zisserman, Jean Ponce |
CVPR | 4 |
| 2010 | A Theoretical Analysis of Feature Pooling in Visual Recognition
Y-Lan Boureau, Jean Ponce, Yann LeCun |
ICML | 2 |
| 2010 | Efficient Optimization for Discriminative Latent Class ModelsabstractDimensionality reduction is commonly used in the setting of multi-label supervised classification to control the learning capacity and to provide a meaningful representation of the data. We introduce a simple forward probabilistic model which is a multinomial extension of reduced rank regression; we show that this model provides a probabilistic interpretation of discriminative clustering methods with added benefits in terms of number of hyperparameters and optimization. While expectation-maximization (EM) algorithm is commonly used to learn these models, its optimization usually leads to local minimum because it relies on a non-convex cost function with many such local minima. To avoid this problem, we introduce a local approximation of this cost function, which leads to a quadratic non-convex optimization problem over a product of simplices. In order to minimize such functions, we propose an efficient algorithm based on convex relaxation and low-rank representation of our data, which allows to deal with large instances. Experiments on text document classification show that the new model outperforms other supervised dimensionality reduction methods, while simulations on unsupervised clustering show that our probabilistic formulation has better properties than existing discriminative clustering methods. Armand Joulin, Francis R. Bach, Jean Ponce |
NIPS | 3 |
| 2010 | Online Learning for Matrix Factorization and Sparse Coding
Julien Mairal, Francis R. Bach, Jean Ponce, Guillermo Sapiro |
J. Mach. Learn. Res. | 3 |
| 2010 | Accurate, Dense, and Robust Multiview StereopsisabstractThis paper proposes a novel algorithm for multiview stereopsis that outputs a dense set of small rectangular patches covering the surfaces visible in the images. Stereopsis is implemented as a match, expand, and filter procedure, starting from a sparse set of matched keypoints, and repeatedly expanding these before using visibility constraints to filter away false matches. The keys to the performance of the proposed algorithm are effective techniques for enforcing local photometric consistency and global visibility constraints. Simple but effective methods are also proposed to turn the resulting patch model into a mesh which can be further refined by an algorithm that enforces both photometric consistency and regularization constraints. The proposed approach automatically detects and discards outliers and obstacles and does not require any initialization in the form of a visual hull, a bounding box, or valid depth ranges. We have tested our algorithm on various data sets including objects with fine surface details, deep concavities, and thin structures, outdoor scenes observed from a restricted set of viewpoints, and "crowded" scenes where moving obstacles appear in front of a static structure of interest. A quantitative evaluation on the Middlebury benchmark shows that the proposed method outperforms all others submitted so far for four out of the six data sets. Yasutaka Furukawa, Jean Ponce |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Detecting Abandoned Objects With a Moving CameraabstractThis paper presents a novel framework for detecting nonflat abandoned objects by matching a reference and a target video sequences. The reference video is taken by a moving camera when there is no suspicious object in the scene. The target video is taken by a camera following the same route and may contain extra objects. The objective is to find these objects. GPS information is used to roughly align the two videos and find the corresponding frame pairs. Based upon the GPS alignment, four simple but effective ideas are proposed to achieve the objective: an intersequence geometric alignment based upon homographies, which is computed by a modified RANSAC, to find all possible suspicious areas, an intrasequence geometric alignment to remove false alarms caused by high objects, a local appearance comparison between two aligned intrasequence frames to remove false alarms in flat areas, and a temporal filtering step to confirm the existence of suspicious objects. Experiments on fifteen pairs of videos show the promise of the proposed method. Hui Kong 0001, Jean-Yves Audibert, Jean Ponce |
IEEE Trans. Image Process. | 3 |
| 2010 | General Road Detection From a Single ImageabstractGiven a single image of an arbitrary road, that may not be well-paved, or have clearly delineated edges, or some a priori known color or texture distribution, is it possible for a computer to find this road? This paper addresses this question by decomposing the road detection process into two steps: the estimation of the vanishing point associated with the main (straight) part of the road, followed by the segmentation of the corresponding road area based upon the detected vanishing point. The main technical contributions of the proposed approach are a novel adaptive soft voting scheme based upon a local voting region using high-confidence voters, whose texture orientations are computed using Gabor filters, and a new vanishing-point-constrained edge detection technique for detecting road boundaries. The proposed method has been implemented, and experiments with 1003 general road images demonstrate that it is effective at detecting road regions in challenging conditions. Hui Kong 0001, Jean-Yves Audibert, Jean Ponce |
IEEE Trans. Image Process. | 3 |
| 2009 | A tensor-based algorithm for high-order graph matchingabstractThis paper addresses the problem of establishing correspondences between two sets of visual features using higher-order constraints instead of the unary or pairwise ones used in classical methods. Concretely, the corresponding hypergraph matching problem is formulated as the maximization of a multilinear objective function over all permutations of the features. This function is defined by a tensor representing the affinity between feature tuples. It is maximized using a generalization of spectral techniques where a relaxed problem is first solved by a multi-dimensional power method, and the solution is then projected onto the closest assignment matrix. The proposed approach has been implemented, and it is compared to state-of-the-art algorithms on both synthetic and real data. Olivier Duchenne, Francis R. Bach, In-So Kweon, Jean Ponce |
CVPR | 4 |
| 2009 | Dense 3D motion capture for human facesabstractThis paper proposes a novel approach to motion capture from multiple, synchronized video streams, specifically aimed at recording dense and accurate models of the structure and motion of highly deformable surfaces such as skin, that stretches, shrinks, and shears in the midst of normal facial expressions. Solving this problem is a key step toward effective performance capture for the entertainment industry, but progress so far has been hampered by the lack of appropriate local motion and smoothness models. The main technical contribution of this paper is a novel approach to regularization adapted to nonrigid tangential deformations. Concretely, we estimate the nonrigid deformation parameters at each vertex of a surface mesh, smooth them over a local neighborhood for robustness, and use them to regularize the tangential motion estimation. To demonstrate the power of the proposed approach, we have integrated it into our previous work for markerless motion capture [9], and compared the performances of the original and new algorithms on three extremely challenging face datasets that include highly nonrigid skin deformations, wrinkles, and quickly changing expressions. Additional experiments with a dataset featuring fast-moving cloth with complex and evolving fold structures demonstrate that the adaptability of the proposed regularization scheme to nonrigid tangential motion does not hamper its robustness, since it successfully recovers the shape and motion of the cloth without overfitting it despite the absence of stretch or shear in this case. Yasutaka Furukawa, Jean Ponce |
CVPR | 2 |
| 2009 | Vanishing point detection for road detectionabstractGiven a single image of an arbitrary road, that may not be well-paved, or have clearly delineated edges, or some a priori known color or texture distribution, is it possible for a computer to find this road? This paper addresses this question by decomposing the road detection process into two steps: the estimation of the vanishing point associated with the main (straight) part of the road, followed by the segmentation of the corresponding road area based on the detected vanishing point. The main technical contributions of the proposed approach are a novel adaptive soft voting scheme based on variable-sized voting region using confidence-weighted Gabor filters, which compute the dominant texture orientation at each pixel, and a new vanishing-point-constrained edge detection technique for detecting road boundaries. The proposed method has been implemented, and experiments with 1003 general road images demonstrate that it is both computationally efficient and effective at detecting road regions in challenging conditions. Hui Kong 0001, Jean-Yves Audibert, Jean Ponce |
CVPR | 3 |
| 2009 | What is a camera?abstractThis paper addresses the problem of characterizing a general class of cameras under reasonable, “linear” assumptions. Concretely, we use the formalism and terminology of classical projective geometry to model cameras by two-parameter linear families of straight lines-that is, degenerate reguli (rank-3 families) and non-degenerate linear congruences (rank-4 families). This model captures both the general linear cameras of Yu and McMillan and the linear oblique cameras of Pajdla. From a geometric perspective, it affords a simple classification of all possible camera configurations. From an analytical viewpoint, it also provides a simple and unified methodology for deriving general formulas for projection and inverse projection, triangulation, and binocular and trinocular geometry. Jean Ponce |
CVPR | 1 |
| 2009 | Automatic annotation of human actions in videoabstractThis paper addresses the problem of automatic temporal annotation of realistic human actions in video using minimal manual supervision. To this end we consider two associated problems: (a) weakly-supervised learning of action models from readily available annotations, and (b) temporal localization of human actions in test videos. To avoid the prohibitive cost of manual annotation for training, we use movie scripts as a means of weak supervision. Scripts, however, provide only implicit, noisy, and imprecise information about the type and location of actions in video. We address this problem with a kernel-based discriminative clustering algorithm that locates actions in the weakly-labeled training data. Using the obtained action samples, we train temporal action detectors and apply them to locate actions in the raw video data. Our experiments demonstrate that the proposed method for weakly-supervised learning of action models leads to significant improvement in action detection. We present detection results for three action classes in four feature length movies with challenging and realistic video data. Olivier Duchenne, Ivan Laptev, Josef Sivic, Francis R. Bach, Jean Ponce |
ICCV | 5 |
| 2009 | Non-local sparse models for image restorationabstractWe propose in this paper to unify two different approaches to image restoration: On the one hand, learning a basis set (dictionary) adapted to sparse signal descriptions has proven to be very effective in image reconstruction and classification tasks. On the other hand, explicitly exploiting the self-similarities of natural images has led to the successful non-local means approach to image restoration. We propose simultaneous sparse coding as a framework for combining these two approaches in a natural manner. This is achieved by jointly decomposing groups of similar signals on subsets of the learned dictionary. Experimental results in image denoising and demosaicking tasks with synthetic and real noise show that the proposed method outperforms the state of the art, making it possible to effectively restore raw images from digital cameras at a reasonable speed and memory cost. Julien Mairal, Francis R. Bach, Jean Ponce, Guillermo Sapiro, Andrew Zisserman |
ICCV | 3 |
| 2009 | Online dictionary learning for sparse codingabstractSparse coding---that is, modelling data vectors as sparse linear combinations of basis elements---is widely used in machine learning, neuroscience, signal processing, and statistics. This paper focuses on learning the basis set, also called dictionary, to adapt it to specific data, an approach that has recently proven to be very effective for signal reconstruction and classification in the audio and image processing domains. This paper proposes a new online optimization algorithm for dictionary learning, based on stochastic approximations, which scales up gracefully to large datasets with millions of training samples. A proof of convergence is presented, along with experiments with natural images demonstrating that it leads to faster performance and better dictionaries than classical batch algorithms for both small and large datasets. Julien Mairal, Francis R. Bach, Jean Ponce, Guillermo Sapiro |
ICML | 3 |
| 2009 | Carved Visual Hulls for Image-Based Modeling
Yasutaka Furukawa, Jean Ponce |
Int. J. Comput. Vis. | 2 |
| 2009 | Accurate Camera Calibration from Multi-View Stereo and Bundle Adjustment
Yasutaka Furukawa, Jean Ponce |
Int. J. Comput. Vis. | 2 |
| 2008 | Segmentation by transductionabstractThis paper addresses the problem of segmenting an image into regions consistent with user-supplied seeds (e.g., a sparse set of broad brush strokes). We view this task as a statistical transductive inference, in which some pixels are already associated with given zones and the remaining ones need to be classified. Our method relies on the Laplacian graph regularizer, a powerful manifold learning tool that is based on the estimation of variants of the Laplace-Beltrami operator and is tightly related to diffusion processes. Segmentation is modeled as the task of finding matting coefficients for unclassified pixels given known matting coefficients for seed pixels. The proposed algorithm essentially relies on a high margin assumption in the space of pixel characteristics. It is simple, fast, and accurate, as demonstrated by qualitative results on natural images and a quantitative comparison with state-of-the-art methods on the Microsoft GrabCut segmentation database. Olivier Duchenne, Jean-Yves Audibert, Renaud Keriven, Jean Ponce, Florent Ségonne |
CVPR | 4 |
| 2008 | Dense 3D motion capture from synchronized video streamsabstractThis paper proposes a novel approach to non-rigid, markerless motion capture from synchronized video streams acquired by calibrated cameras. The instantaneous geometry of the observed scene is represented by a polyhedral mesh with fixed topology. The initial mesh is constructed in the first frame using the publicly available PMVS software for multi-view stereo [7]. Its deformation is captured by tracking its vertices over time, using two optimization processes at each frame: a local one using a rigid motion model in the neighborhood of each vertex, and a global one using a regularized nonrigid model for the whole mesh. Qualitative and quantitative experiments using seven real datasets show that our algorithm effectively handles complex nonrigid motions and severe occlusions. Yasutaka Furukawa, Jean Ponce |
CVPR | 2 |
| 2008 | Accurate camera calibration from multi-view stereo and bundle adjustmentabstractThe advent of high-resolution digital cameras and sophisticated multi-view stereo algorithms offers the promises of unprecedented geometric fidelity in image-based modeling tasks, but it also puts unprecedented demands on camera calibration to fulfill these promises. This paper presents a novel approach to camera calibration where top-down information from rough camera parameter estimates and the output of a publicly available multiview-stereo system [6] on scaled-down input images are used to effectively guide the search for additional image correspondences and significantly improve camera calibration parameters using a standard bundle adjustment algorithm [14]. The proposed method has been tested on several real datasets—including objects without salient features for which image correspondences cannot be found in a purely bottom-up fashion, and image-based modeling tasks-including the construction of visual hulls where thin structures are lost without our calibration procedure. Yasutaka Furukawa, Jean Ponce |
CVPR | 2 |
| 2008 | Discriminative learned dictionaries for local image analysisabstractSparse signal models have been the focus of much recent research, leading to (or improving upon) state-of-the-art results in signal, image, and video restoration. This article extends this line of research into a novel framework for local image discrimination tasks, proposing an energy formulation with both sparse reconstruction and class discrimination components, jointly optimized during dictionary learning. This approach improves over the state of the art in texture segmentation experiments using the Brodatz database, and it paves the way for a novel scene analysis and recognition framework based on simultaneously learning discriminative and reconstructive dictionaries. Preliminary results in this direction using examples from the Pascal VOC06 and Graz02 datasets are presented as well. Julien Mairal, Francis R. Bach, Jean Ponce, Guillermo Sapiro, Andrew Zisserman |
CVPR | 3 |
| 2008 | Discriminative Sparse Image Models for Class-Specific Edge Detection and Image Interpretation
Julien Mairal, Marius Leordeanu, Francis R. Bach, Martial Hebert, Jean Ponce |
ECCV (3) | 5 |
| 2008 | Supervised Dictionary LearningabstractIt is now well established that sparse signal models are well suited to restoration tasks and can effectively be learned from audio, image, and video data. Recent research has been aimed at learning discriminative sparse models instead of purely reconstructive ones. This paper proposes a new step in that direction with a novel sparse representation for signals belonging to different classes in terms of a shared dictionary and multiple decision functions. It is shown that the linear variant of the model admits a simple probabilistic interpretation, and that its most general variant also admits a simple interpretation in terms of kernels. An optimization framework for learning all the components of the proposed model is presented, along with experiments on standard handwritten digit and texture classification tasks. Julien Mairal, Francis R. Bach, Jean Ponce, Guillermo Sapiro, Andrew Zisserman |
NIPS | 3 |
| 2007 | Accurate, Dense, and Robust Multi-View StereopsisabstractThis paper proposes a novel algorithm for calibrated multi-view stereopsis that outputs a (quasi) dense set of rectangular patches covering the surfaces visible in the input images. This algorithm does not require any initialization in the form of a bounding volume, and it detects and discards automatically outliers and obstacles. It does not perform any smoothing across nearby features, yet is currently the top performer in terms of both coverage and accuracy for four of the six benchmark datasets presented in [20]. The keys to its performance are effective techniques for enforcing local photometric consistency and global visibility constraints. Stereopsis is implemented as a match, expand, and filter procedure, starting from a sparse set of matched keypoints, and repeatedly expanding these to nearby pixel correspondences before using visibility constraints to filter away false matches. A simple but effective method for turning the resulting patch model into a mesh appropriate for image-based modeling is also presented. The proposed approach is demonstrated on various datasets including objects with fine surface details, deep concavities, and thin structures, outdoor scenes observed from a restricted set of viewpoints, and "crowded" scenes where moving obstacles appear in different places in multiple images of a static structure of interest. Yasutaka Furukawa, Jean Ponce |
CVPR | 2 |
| 2007 | Flexible Object Models for Category-Level 3D Object RecognitionabstractToday's category-level object recognition systems largely focus on fronto-parallel views of objects with characteristic texture patterns. To overcome these limitations, we propose a novel framework for visual object recognition where object classes are represented by assemblies of partial surface models (PSMs) obeying loose local geometric constraints. The PSMs themselves are formed of dense, locally rigid assemblies of image features. Since our model only enforces local geometric consistency, both at the level of model parts and at the level of individual features within the parts, it is robust to viewpoint changes and intra-class variability. The proposed approach has been implemented, and it outperforms the state-of-the-art algorithms for object detection and localization recently compared in [14] on the Pascal 2005 VOC Challenge Cars Test 1 data. Akash Kushal, Cordelia Schmid, Jean Ponce |
CVPR | 3 |
| 2007 | Projective Visual Hulls
Svetlana Lazebnik, Yasutaka Furukawa, Jean Ponce |
Int. J. Comput. Vis. | 3 |
| 2007 | Segmenting, Modeling, and Matching Video Clips Containing Multiple Moving ObjectsabstractThis paper presents a novel representation for dynamic scenes composed of multiple rigid objects that may undergo different motions and are observed by a moving camera. Multiview constraints associated with groups of affine-covariant scene patches and a normalized description of their appearance are used to segment a scene into its rigid components, construct three-dimensional models of these components, and match instances of models recovered from different image sequences. The proposed approach has been applied to the detection and matching of moving objects in video sequences and to shot matching, i.e., the identification of shots that depict the same scene in a video clip. Fred Rothganger, Svetlana Lazebnik, Cordelia Schmid, Jean Ponce |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2007 | Capturing a Convex Object With Three DiscsabstractThis paper addresses the problem of capturing an arbitrary convex object P in the plane with three congruent disc-shaped robots. Given two stationary robots in contact with P, we characterize the set of positions of a third robot, the so-called capture region, that prevent P from escaping to infinity via continuous rigid motion. We show that the computation of the capture region reduces to a visibility problem. We present two algorithms for solving this problem, and for computing the capture region when P is a polygon and the robots are points (zero-radius discs). The first algorithm is exact and has polynomial time complexity. The second one uses simple hidden surface removal techniques from computer graphics to output an arbitrarily accurate approximation of the capture region; it has been implemented, and examples are presented. Jeff Erickson 0001, Shripad Thite, Fred Rothganger, Jean Ponce |
IEEE Trans. Robotics | 4 |
| 2006 | Beyond Bags of Features: Spatial Pyramid Matching for Recognizing Natural Scene CategoriesabstractThis paper presents a method for recognizing scene categories based on approximate global geometric correspondence. This technique works by partitioning the image into increasingly fine sub-regions and computing histograms of local features found inside each sub-region. The resulting "spatial pyramid" is a simple and computationally efficient extension of an orderless bag-of-features image representation, and it shows significantly improved performance on challenging scene categorization tasks. Specifically, our proposed method exceeds the state of the art on the Caltech-101 database and achieves high accuracy on a large database of fifteen natural scene categories. The spatial pyramid framework also offers insights into the success of several recently proposed image descriptions, including Torralba’s "gist" and Lowe’s SIFT descriptors. Svetlana Lazebnik, Cordelia Schmid, Jean Ponce |
CVPR (2) | 3 |
| 2006 | A Geodesic Active Contour Framework for Finding GlassabstractThis paper addresses the problem of finding objects made of glass (or other transparent materials) in images. Since the appearance of glass objects depends for the most part on what lies behind them, we propose to use binary criteria ("are these two regions made of the same material?") rather than unary ones ("is this glass?") to guide the segmentation process. Concretely, we combine two complementary measures of affinity between regions made of the same material and discrepancy between regions made of different ones into a single objective function, and use the geodesic active contour framework to minimize this function over pixel labels. The proposed approach has been implemented, and qualitative and quantitative experimental results are presented. Kenton McHenry, Jean Ponce |
CVPR (1) | 2 |
| 2006 | Carved Visual Hulls for Image-Based Modeling
Yasutaka Furukawa, Jean Ponce |
ECCV (1) | 2 |
| 2006 | Modeling 3D Objects from Stereo Views and Recognizing Them in Photographs
Akash Kushal, Jean Ponce |
ECCV (2) | 2 |
| 2006 | 3D Object Modeling and Recognition Using Local Affine-Invariant Image Descriptors and Multi-View Spatial Constraints
Fred Rothganger, Svetlana Lazebnik, Cordelia Schmid, Jean Ponce |
Int. J. Comput. Vis. | 4 |
| 2006 | Robust Structure and Motion from Outlines of Smooth Curved SurfacesabstractThis paper addresses the problem of estimating the motion of a camera as it observes the outline (or apparent contour) of a solid bounded by a smooth surface in successive image frames. In this context, the surface points that project onto the outline of an object depend on the viewpoint and the only true correspondences between two outlines of the same object are the projections of frontier points where the viewing rays intersect in the tangent plane of the surface. In turn, the epipolar geometry is easily estimated once these correspondences have been identified. Given the apparent contours detected in an image sequence, a robust procedure based on RANSAC and a voting strategy is proposed to simultaneously estimate the camera configurations and a consistent set of frontier point projections by enforcing the redundancy of multiview epipolar geometry. The proposed approach is, in principle, applicable to orthographic, weak-perspective, and affine projection models. Experiments with nine real image sequences are presented for the orthographic projection case, including a quantitative comparison with the ground-truth data for the six data sets for which the latter information is available. Sample visual hulls have been computed from all image sequences for qualitative evaluation. Yasutaka Furukawa, Amit Sethi, Jean Ponce, David J. Kriegman |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2005 | Finding GlassabstractThis paper addresses the problem of finding glass objects in images. Visual cues obtained by combining the systematic distortions in background texture occurring at the boundaries of transparent objects with the strong highlights typical of glass surfaces are used to train a hierarchy of classifiers, identify glass edges, and find consistent support regions for these edges. Qualitative and quantitative experiments involving a number of different classifiers and real images are presented. Kenton McHenry, Jean Ponce, David A. Forsyth |
CVPR (2) | 2 |
| 2005 | On the Absolute Quadratic Complex and Its Application to AutocalibrationabstractThis article introduces the absolute quadratic complex formed by all lines that intersect the absolute conic. If /spl omega/ denotes the 3 /spl times/ 3 symmetric matrix representing the image of that conic under the action of a camera with projection matrix P, it is shown that /spl omega/ /spl ap/ P/sup ~//spl Omega//sub /spl I.bar//P/sup ~T/ where V is the 3 /spl times/ 6 line projection matrix associated with P and /spl Omega//sub /spl I.bar// is a 6 /spl times/ 6 symmetric matrix of rank 3 representing the absolute quadratic complex. This simple relation between a camera's intrinsic parameters, its projection matrix expressed in a projective coordinate frame, and the metric upgrade separating this frame from a metric one - as respectively captured by the matrices /spl omega/, P/sup ~/ and /spl Omega//sub /spl I.bar// - provides a new framework for autocalibration, particularly well suited to typical digital cameras with rectangular or square pixels since the skew and aspect ratio are decoupled from the other intrinsic parameters in /spl omega/. Jean Ponce, Kenton McHenry, Théodore Papadopoulo, Monique Teillaud, Bill Triggs |
CVPR (1) | 1 |
| 2005 | A Maximum Entropy Framework for Part-Based Texture and Object RecognitionabstractThis paper presents a probabilistic part-based approach for texture and object recognition. Textures are represented using a part dictionary found by quantizing the appearance of scale- or affine- invariant keypoints. Object classes are represented using a dictionary of composite semi-local parts, or groups of neighboring keypoints with stable and distinctive appearance and geometric layout. A discriminative maximum entropy framework is used to learn the posterior distribution of the class label given the occurrences of parts from the dictionary in the training set. Experiments on two texture and two object databases demonstrate the effectiveness of this framework for visual classification. Svetlana Lazebnik, Cordelia Schmid, Jean Ponce |
ICCV | 3 |
| 2005 | The Local Projective Shape of Smooth Surfaces and Their Outlines
Svetlana Lazebnik, Jean Ponce |
Int. J. Comput. Vis. | 2 |
| 2005 | A Sparse Texture Representation Using Local Affine RegionsabstractThis paper introduces a texture representation suitable for recognizing images of textured surfaces under a wide range of transformations, including viewpoint changes and nonrigid deformations. At the feature extraction stage, a sparse set of affine Harris and Laplacian regions is found in the image. Each of these regions can be thought of as a texture element having a characteristic elliptic shape and a distinctive appearance pattern. This pattern is captured in an affine-invariant fashion via a process of shape normalization followed by the computation of two novel descriptors, the spin image and the RIFT descriptor. When affine invariance is not required, the original elliptical shape servee as an additional discriminative feature for texture recognition. The proposed approach is evaluated in retrieval and classification tasks using the entire Brodatz database and a publicly available collection of 1,000 photographs of textured surfaces taken from different viewpoints. Svetlana Lazebnik, Cordelia Schmid, Jean Ponce |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2004 | Semi-Local Affine Parts for Object RecognitionabstractHAL is a multi-disciplinary open access archive for the deposit and dissemination of sci-entific research documents, whether they are pub-lished or not. The documents may come from teaching and research institutions in France or abroad, or from public or private research centers. L’archive ouverte pluridisciplinaire HAL, est destinée au dépôt et a ̀ la diffusion de documents scientifiques de niveau recherche, publiés ou non, émanant des établissements d’enseignement et de recherche français ou étrangers, des laboratoires publics ou privés. Svetlana Lazebnik, Cordelia Schmid, Jean Ponce |
BMVC | 3 |
| 2004 | Segmenting, Modeling, and Matching Video Clips Containing Multiple Moving Objects
Fred Rothganger, Svetlana Lazebnik, Cordelia Schmid, Jean Ponce |
CVPR (2) | 4 |
| 2004 | Structure and Motion from Images of Smooth Textureless Objects
Yasutaka Furukawa, Amit Sethi, Jean Ponce, David J. Kriegman |
ECCV (2) | 3 |
| 2004 | Guest Editorial: Computer Vision Research at the Beckman Institute for Advanced Science and Technology
Jean Ponce |
Int. J. Comput. Vis. | 1 |
| 2004 | Editorial
Jean Ponce |
Int. J. Comput. Vis. | 1 |
| 2004 | Curve and Surface Duals and the Recognition of Curved 3D Objects from their Silhouettes
Amit Sethi, David Renaudie, David J. Kriegman, Jean Ponce |
Int. J. Comput. Vis. | 4 |
| 2003 | A Sparse Texture Representation Using Affine-Invariant RegionsabstractThis paper introduces a texture representation suitable for recognizing images of textured surfaces under a wide range of transformations, including viewpoint changes and nonrigid deformations. At the feature extraction stage, a sparse set of affine-invariant local patches is extracted from the image. This spatial selection process permits the computation of characteristic scale and neighborhood shape for every texture element. The proposed texture representation is evaluated in retrieval and classification tasks using the entire Brodatz database and a collection of photographs of textured surfaces taken from different viewpoints. Svetlana Lazebnik, Cordelia Schmid, Jean Ponce |
CVPR (2) | 3 |
| 2003 | 3D Object Modeling and Recognition Using Affine-Invariant Patches and Multi-View Spatial ConstraintsabstractThis paper presents a representation for three-dimensional objects in terms of affine-invariant image patches and their spatial relationships. Multi-view constraints associated with groups of patches are combined with a normalized representation of their appearance to guide matching and reconstruction, allowing the acquisition of true three-dimensional affine and Euclidean models from multiple images and their recognition in a single photograph taken from an arbitrary viewpoint. The proposed approach does not require a separate segmentation stage and is applicable to cluttered scenes. Preliminary modeling and recognition results are presented. Fred Rothganger, Svetlana Lazebnik, Cordelia Schmid, Jean Ponce |
CVPR (2) | 4 |
| 2003 | The Local Projective Shape of Smooth Surfaces and their OutlinesabstractWe examine projectively invariant local properties of smooth curves and surfaces. Oriented projective differential geometry is proposed as a theoretical framework for establishing such invariants and describing the local shape of surfaces and their outlines. This framework is applied to two problems: a projective proof of Koenderink's famous characterization of convexities, concavities, and inflections of apparent contours; and the determination of the relative orientation of rim tangents at frontier points. Svetlana Lazebnik, Jean Ponce |
ICCV | 2 |
| 2003 | Affine-Invariant Local Descriptors and Neighborhood Statistics for Texture RecognitionabstractWe present a framework for texture recognition based on local affine-invariant descriptors and their spatial layout. At modelling time, a generative model of local descriptors is learned from sample images using the EM algorithm. The EM framework allows the incorporation of unsegmented multitexture images into the training set. The second modelling step consists of gathering co-occurrence statistics of neighboring descriptors. At recognition time, initial probabilities computed from the generative model are refined using a relaxation step that incorporates co-occurrence statistics. Performance is evaluated on images of an indoor scene and pictures of wild animals. Svetlana Lazebnik, Cordelia Schmid, Jean Ponce |
ICCV | 3 |
| 2003 | Binocular Helmholtz StereopsisabstractHelmholtz stereopsis has been introduced recently as a surface reconstruction technique that does not assume a model of surface reflectance. In the reported formulation, correspondence was established using a rank constraint, necessitating at least three viewpoints and three pairs of images. Here, it is revealed that the fundamental Helmholtz stereopsis constraint defines a nonlinear partial differential equation, which can be solved using only two images. It is shown that, unlike conventional stereo, binocular Helmholtz stereopsis is able to establish correspondence (and thereby recover surface depth) for objects having an arbitrary and unknown BRDF and in textureless regions (i.e., regions of constant or slowly varying BRDF). An implementation and experimental results validate the method for specular surfaces with and without texture. Todd E. Zickler, Jeffrey Ho, David J. Kriegman, Jean Ponce, Peter N. Belhumeur |
ICCV | 4 |
| 2003 | Capturing a convex object with three discsabstractThis paper addresses the problem of capturing an arbitrary convex object P in the plane with three congruent disc-shaped robots. Given two stationary robots in contact with P, we characterize the set of positions of a third robot that prevent P from escaping to infinity and show that the computation of this so-called capture region reduces to the resolution of a visibility problem. We present two algorithms for solving this problem and computing the capture region when P is a polygon and the robots are points (zero-radius discs). The first algorithm is exact and has polynomial-time complexity. The second one uses simple hidden-surface removal techniques from computer graphics to output an arbitrarily accurate approximation of the capture region; it has been implemented and examples are presented. Jeff Erickson 0001, Shripad Thite, Fred Rothganger, Jean Ponce |
ICRA | 4 |
| 2002 | On Pencils of Tangent Planes and the Recognition of Smooth 3D Shapes from Silhouettes
Svetlana Lazebnik, Amit Sethi, Cordelia Schmid, David J. Kriegman, Jean Ponce, Martial Hebert |
ECCV (3) | 5 |
| 2002 | Motion planning for disc-shaped robots pushing a polygonal object in the planeabstractThis paper addresses the problem of using three disc-shaped robots to manipulate a polygonal object in the plane in the presence of obstacles. The proposed approach is based on the computation of maximal discs (dubbed maximum independent capture discs, or MICaDs) where the robots can move independently while preventing the object from escaping their grasp. It is shown that, in the absence of obstacles, it is always possible to bring a polygonal object from any configuration to any other one with robot motions constrained to lie in a set of overlapping MICaDs. This approach is generalized to the case where obstacles are present by decomposing the corresponding motion planning task into the construction of a collision-free path for a modified form of the object, and the execution of this path by a sequence of simultaneous and independent robot motions within overlapping MICaDs. The proposed algorithm is guaranteed to generate a valid plan, provided a collision-free path exists for the modified form of the object. It has been implemented and experiments with Nomadic Scout mobile robots are presented. Attawith Sudsang, Fred Rothganger, Jean Ponce |
IEEE Trans. Robotics Autom. | 3 |
| 2001 | On Computing Exact Visual Hulls of Solids Bounded by Smooth SurfacesabstractThis paper presents a method for computing the visual hull that is based on two novel representations: the rim mesh, which describes the connectivity of contour generators on the object surface; and the visual hull mesh, which describes the exact structure of the surface of the solid formed by intersecting a finite number of visual cones. We describe the topological features of these meshes and show how they can be identified in the image using epipolar constraints. These constraints are used to derive an image-based practical reconstruction algorithm that works with weakly calibrated cameras. Experiments on synthetic and real data validate the proposed approach. Svetlana Lazebnik, Edmond Boyer, Jean Ponce |
CVPR (1) | 3 |
| 2001 | Provably-Convergent Iterative Methods for Projective Structure from MotionabstractThe estimation of the projective structure of a scene from image correspondences can be formulated as the minimization of the mean-squared distance between predicted and observed image points with respect to the projection matrices, the scene point positions, and their depths. Since these unknowns are not independent, constraints must be chosen to ensure that the optimization process. is well posed. This paper examines three plausible choices, and shows that the first one leads to the Sturm-Triggs projective factorization algorithm, while the other two lead to new provably-convergent approaches. Experiments with synthetic and real data are used to compare the proposed techniques to the Sturm-Triggs algorithm and bundle adjustment. Shyjan Mahamud, Martial Hebert, Yasuhiro Omori, Jean Ponce |
CVPR (1) | 4 |
| 2001 | An implemented planner for manipulating a polygonal object in the plane with three disc-shaped mobile robotsabstractPresents an implementation of a planner that uses three disc-shaped robots to manipulate a polygonal object in the plane in the presence of obstacles. The approach is based on the computation of the maximal discs (maximal independent capture discs or MICaDs) where the robots can move independently while preventing the object from escaping their grasp. It has been shown that, in the absence of obstacles, it is always possible to bring a polygonal object from any configuration to any other one with robot motions constrained to lie in a set of overlapping MICaDs. This approach is generalized to the case where obstacles are present by decomposing the motion planning task into (1) the construction of a collision-free path for a modified form of the object, and (2) the execution of this path by a sequence of simultaneous and independent robot motions within overlapping MICaDs. The approach is guaranteed to work provided a collision free path exists for the modified form of the object. Experiments with Nomadic Scouts and a visual localization system are presented. Attawith Sudsang, Fred Rothganger, Jean Ponce |
IROS | 3 |
| 2001 | Image-Based Rendering Using Parameterized Image Varieties
Yakup Genc, Jean Ponce |
Int. J. Comput. Vis. | 2 |
| 2001 | On Computing Structural Changes in Evolving Surfaces and their Appearance
Sung-Il Pae, Jean Ponce |
Int. J. Comput. Vis. | 2 |
| 2000 | Duals, Invariants, and the Recognition of Smooth Objects from their Occlucing Contours
David Renaudie, David J. Kriegman, Jean Ponce |
ECCV (1) | 3 |
| 2000 | A Reconfigurable Parts Feeder with an Array of PinsabstractThis paper presents a simple parts feeder consisting of a grid of retractable pins on a vertical plate to manipulate polygonal parts. This reconfigurable "Pachinko machine" is intended as a parts feeding device for flexible assembly. A part dropped on this device may come to rest on the actuated pins, or bounce out or fall through. We can control the set of equilibrium part configurations by selecting the set of actuated pins. The objective is to automatically compute sequences of pin actuation that bring the part to a goal configuration without predicting the exact object motion between equilibria. Our approach is based on the construction of the capture region of each part equilibrium. Reorienting a part reduces to building a directed graph whose nodes consist of equilibria and whose area link pairs of nodes such that the first equilibrium lies in the capture region of the second one, and then exploring this graph to find paths from initial to goal states. We have implemented an algorithm to generate the capture regions and these paths, and have conducted experiments on a prototype Pachinko machine. Sebastien J. Blind, Christopher C. McCullough, Srinivas Akella, Jean Ponce |
ICRA | 4 |
| 2000 | Constructing Geometric Object Models from ImagesabstractThis paper addresses the problem of constructing object models from various types of images. After a brief discussion of current approaches to this problem, we focus on two of its instances: the construction of three-dimensional surface models from object outlines found in a small set of registered photographs; and the synthesis of new images of a scene without any explicit three-dimensional reconstruction (image-based rendering). In both cases, we discuss the state of the art and present some of our recent work as an illustration of what can be achieved today. Jean Ponce, Yakup Genc, Steve Sullivan |
ICRA | 1 |
| 2000 | A New Approach to Motion Planning for Disc-Shaped Robots Manipulating a Polygonal Object in the PlaneabstractThis paper addresses the problem of using three disc-shaped robots to manipulate a polygonal object in the plane in the presence of obstacles. The proposed approach is based on the characterization of the maximal discs (maximum independent capture discs, or MICaDS) where the robots can move independently while preventing the object from escaping their grasp. It is shown that, in the absence of obstacles, it is always possible to bring a polygonal object from any configuration to any other one with robot motions constrained to lie in a set of overlapping MICaDS. A strategy for computing these motions is used in conjunction with an exact motion planner to devise an algorithm guaranteed to find a motion plan avoiding collisions with obstacles as long as a collision-free path exists for the object grown by the diameter of the robots plus some arbitrary positive number /spl epsiv/. Attawith Sudsang, Jean Ponce |
ICRA | 2 |
| 2000 | Probabilistic 3D Object Recognition
Ilan Shimshoni, Jean Ponce |
Int. J. Comput. Vis. | 2 |
| 1999 | Parameterized Image Varieties and Estimation with Bilinear ConstraintsabstractThis paper addresses the problem of reliably estimating the coefficients of the parameterized image variety (PIV) associated with the set of weak perspective images of a rigid scene, with applications in image-based rendering. Exploiting the fact that the constraints defining the PIV are linear in its coefficients and bilinear in the image data, the estimation procedure is cast in the errors-in-variables framework and solved using the method proposed by Y. Leedan and P. Meer (1998) for this type of problems. The proposed approach has been implemented, and experiments with real data are shown to yield much better prediction power than the original method based on singular value decomposition. Extensions to the more difficult case of paraperspective projection are briefly discussed. Yakup Genc, Jean Ponce, Yoram Leedan, Peter Meer |
CVPR | 2 |
| 1999 | Toward a Scale-Space Aspect Graph: Solids of RevolutionabstractThis paper addresses the problem of constructing the scale-space aspect graph of a solid of revolution whose surface is the zero set of a polynomial volumetric density undergoing a Gaussian diffusion process. Equations for the associated visual event surfaces are derived, and polynomial curve tracing techniques are used to delineate these surfaces. An implementation and examples are presented, and limitations as well as extensions of the proposed approach are discussed. Sung-Il Pae, Jean Ponce |
CVPR | 2 |
| 1999 | On Manipulating Polygonal Objects with Three 2-DOF Robots in the PlaneabstractAddresses the problem of grasping and manipulating a polygonal object with three disc-shaped robots in the plane. These robots may be the fingertips of a gripper or mobile platforms. The proposed approach is based on the characterization of the range of possible object motions when two of the effectors are fixed and the third one is allowed to move in the plane with two degrees of freedom. This technique does not assume that contact is maintained during the execution of the grasping/manipulation task, nor does it rely on detailed (and a priori unverifiable) models of friction or contact dynamics, but it allows the construction of manipulation plans guaranteed to succeed under the weaker assumption that jamming does not occur during the task execution. The proposed approach is validated by simulation examples and preliminary experiments with Nomadic Scout robots. Attawith Sudsang, Jean Ponce, Mark Hyman, David J. Kriegman |
ICRA | 2 |
| 1999 | Structure and Motion Estimation from Dynamic Silhouettes under Perspective Projection
Tanuja Joshi, Narendra Ahuja, Jean Ponce |
Int. J. Comput. Vis. | 3 |
| 1998 | Parameterized Image Varieties: A Novel Approach to the Analysis and Synthesis of Image SequencesabstractThis paper addresses the problem of characterizing the space formed by all images of a rigid set of n points observed by a weak perspective or paraperspective camera. By taking explicitly into account the Euclidean constraints associated with calibrated cameras, we show that this space is a six-dimensional variety embedded in R/sup 2n/, and parameterize it using the image positions of three reference points. This parameterization is constructed via linear least squares from point correspondences established across a sequence of images, and it is used to synthesize new pictures without any explicit three-dimensional model. Degenerate scene and camera configurations are analyzed, and experiments with real image sequences are presented. Yakup Genc, Jean Ponce |
ICCV | 2 |
| 1998 | Automatic Model Construction, Pose Estimation, and Object Recognition from Photographs using Triangular SplinesabstractThis paper proposes a method for automatically constructing triangular G/sup 1/ spline models of complex three-dimensional objects from a few registered photographs. These models are used for pose estimation from monocular silhouette data and they form the basis for a simple recognition strategy. The proposed approach is demonstrated by several experiments. Steve Sullivan, Jean Ponce |
ICCV | 2 |
| 1998 | On Grasping and Manipulating Polygonal Objects with Disc-Shaped Robots in the PlaneabstractThis paper addresses the problem of grasping and manipulating a polygonal object with three disc-shaped robots capable of translating in arbitrary directions in the plane. The main novelty of the proposed approach is that it does not assume that contact is maintained during the execution of the grasping/manipulation task, nor does it rely on detailed (and a priori unverifiable) models of friction or contact dynamics. Instead, the range of possible object motions for a given position of the robots is characterized in configuration space. This allows the construction of manipulation plans guaranteed to succeed under the weaker assumption that jamming does not occur during the task execution. Attawith Sudsang, Jean Ponce |
ICRA | 2 |
| 1998 | Invariant-Based Recognition of Complex Curved 3D Objects from Image ContoursabstractThis paper addresses the problem of recognizing three-dimensional objects bounded by smooth curved surfaces from image contours found in a single photograph. The proposed approach is based on a viewpoint-invariant relationship between object geometry and certain image features under weak perspective projection. The image features themselves are viewpoint-dependent. Concretely, the set of all possible silhouette bitangents, along with the contour points sharing the same tangent direction, is the projection of a one-dimensional set of surface points where each point lies on the occluding contour for a five-parameter family of viewpoints. These image features form a one-parameter family of equivalence classes, and it is shown that each class can be characterized by a set of numerical attributes that remain constant across the corresponding five-dimensional set of viewpoints. This is the basis for describing objects by “invariant” curves embedded in high-dimensional spaces. Modeling is achieved by moving an object in front of a camera and does not require knowing the object-to-camera transformation; nor does it involve implicit or explicit three-dimensional shape reconstruction. At recognition time, attributes computed from a single image are used to index the model database, and both qualitative and quantitative verification procedures eliminate potential false matches. The approach has been implemented and examples are presented. B. Vijayakumar, David J. Kriegman, Jean Ponce |
Comput. Vis. Image Underst. | 3 |
| 1998 | Epipolar Geometry and Linear Subspace Methods: A New Approach to Weak Calibration
Jean Ponce, Yakup Genc |
Int. J. Comput. Vis. | 1 |
| 1998 | Automatic Model Construction and Pose Estimation From Photographs Using Triangular SplinesabstractThis paper addresses the automatic construction of complex spline object models from a few photographs. Our approach combines silhouettes from registered images to construct a G/sup 1/-continuous triangular spline approximation of an object with unknown topology. We apply a similar optimization procedure to estimate the pose of a modeled object from a single image. Experimental examples of model construction and pose estimation are presented for several complex objects. Steve Sullivan, Jean Ponce |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1997 | In-hand manipulation: geometry and algorithmsabstractAddresses the problem of manipulating three-dimensional objects with a reconfigurable gripper. A detailed analysis of the problem geometry in configuration space is used to devise a simple and efficient algorithm for manipulation planning. The proposed approach has been implemented and preliminary simulation experiments are discussed. Attawith Sudsang, Jean Ponce |
IROS | 2 |
| 1997 | On planning immobilizing grasps for a reconfigurable gripperabstractWe propose a reconfigurable gripper that consists of two parallel plates whose distance can be adjusted by a computer-controlled actuator. The bottom plate is a bare plane, and the top plate carries a rectangular grid of actuated pins that can translate in discrete increments under computer control. We propose to use this gripper to immobilize objects through frictionless contacts with three of the pins and the bottom plate. We present an efficient grasp planning algorithm, describe the design of the gripper, which is currently under construction, and report preliminary simulation experiments. Attawith Sudsang, Narayan Srinivasa, Jean Ponce |
IROS | 3 |
| 1997 | On Computing Aspect Graphs of Smooth Shapes from Volumetric Data
J. Alison Noble, Dale L. Wilson, Jean Ponce |
Comput. Vis. Image Underst. | 3 |
| 1997 | Recovering the Shape of Polyhedra Using Line-Drawing Analysis and Complex Reflectance Models
Ilan Shimshoni, Jean Ponce |
Comput. Vis. Image Underst. | 2 |
| 1997 | Hot curves for modelling and recognition of smooth curved 3D objectsabstract: We represent arbitrary smooth curved 3D shapes by a discrete set of HOT curves where a surface admits High Order Tangents. These curves determine the structure of the image contours and its catastrophic changes, and there is a natural correspondence between some of them and monocular contour features such as inflections and bitangents. We present a method for automatically constructing the HOT curves from continuous sequences of video images and describe an approach to object recognition using viewpoint-dependent monocular image features as indices into a database of models and as a basis for pose estimation. We have implemented both the methods and present results obtained from real images. 1 Introduction While implemented recognition systems based on parametric shape representations such as algebraic surfaces or superquadrics have demonstrated their usefulness, the ultimate utility of a representation is limited by its scope. This suggests looking for a more general representation ... Tanuja Joshi, B. Vijayakumar, David J. Kriegman, Jean Ponce |
Image Vis. Comput. | 4 |
| 1997 | Finite-Resolution Aspect Graphs of Polyhedral ObjectsabstractWe address the problem of computing the aspect graph of a polyhedral object observed by an orthographic camera with limited spatial resolution, such that two image points separated by a distance smaller than a preset threshold cannot be resolved. Under this model, views that would differ under normal orthographic projection may become equivalent, while "accidental" views may occur over finite areas of the view space. We present a catalogue of visual events for polyhedral objects and give an algorithm for computing the aspect graph and enumerating all qualitatively different aspects. The algorithm has been fully implemented and results are presented. Ilan Shimshoni, Jean Ponce |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1996 | Epipolar Geometry and Linear Subspace Methods: A New Approach to Weak CalibrationabstractThis paper addresses the problem of estimating the epipolar geometry from point correspondences between two images taken by uncalibrated perspective cameras. It is shown that Jepson's and Heeger's linear subspace technique for infinitesimal motion estimation can be generalized to the finite motion case by choosing an appropriate basis for projective space. This yields a linear method for weak calibration. The proposed algorithm has been implemented and tested on both real and synthetic images, and it is compared to other linear and non-linear approaches to weak calibration. Jean Ponce, Yakup Genc |
CVPR | 1 |
| 1996 | Structure and motion of curved 3D objects from monocular silhouettesabstractThe silhouette of a smooth 3D object observed by a moving camera changes over time. Past work has shown how surface geometry can be recovered using the deformation of the silhouette when the camera motion is known. This paper addresses the problem of estimating both the full Euclidean surface structure and the camera motion from a dense set of silhouettes captured under orthographic or scaled orthographic projection. The approach relies on a viewpoint-invariant representation of curves swept by viewpoint-dependent features such as bitangents, inflections and contour points with parallel tangents. Feature points, which form stereo frontier points between non-consecutive images, are matched using this representation. The camera's angular velocity is computed from constraints derived from this correspondence along with the image velocity of these features. From the angular velocity, the epipolar geometry is ascertained, and infinitesimal motion frontier points can be detected. In turn, the motion of these frontier points constrains the translation component of camera motion. Finally, the surface is reconstructed using established techniques once the camera motion has been estimated. B. Vijayakumar, David J. Kriegman, Jean Ponce |
CVPR | 3 |
| 1996 | On planning immobilizing fixtures for three-dimensional polyhedral partsabstractWe propose a simple three-dimensional modular fixturing device and present an algorithm for enumerating all of the immobilizing fixtures of a polyhedral object that can be achieved with this device and four frictionless contacts. Our approach is based on the second-order mobility theory of Rimon and Burdick (1993, 1994). Jean Ponce |
ICRA | 1 |
| 1995 | Structure and Motion Estimation from Dynamic Silhouettes under Perspective ProjectionabstractAddresses the problem of estimating the structure and motion of a smooth curved object from its silhouettes observed over time by a trinocular stereo rig under perspective projection. We first construct a model for the local structure along the silhouette for each frame in the temporal sequence. Successive local models are then integrated into a global surface description by estimating the motion between successive time instants. The algorithm tracks certain surface features (parabolic points) and image features (silhouette inflections and frontier points) which are used to bootstrap the motion estimation process. The entire silhouette along with the reconstructed local structure are then used to refine the initial motion estimate. We have implemented the proposed approach and report results on real images.> Tanuja Joshi, Narendra Ahuja, Jean Ponce |
ICCV | 3 |
| 1995 | Probabilistic 3D Object RecognitionabstractA probabilistic 3D object recognition algorithm is presented. In order to guide the recognition process the probability that match hypotheses between image features and model features are correct is computed. A model is developed which uses the probabilistic peaking effect of measured angles and ratios of lengths by tracing iso angle and iso ratio curves on the viewing sphere. The model also accounts for various types of uncertainty in the input such as incomplete and inexact edge detection. For each match hypothesis the pose of the object and the pose uncertainty which is due to the uncertainty in vertex position are recovered. This is used to find sets of hypotheses which reinforce each other by matching features of the same object with compatible uncertainty subsets. A probabalistic expression is used to rank these hypothesis sets. The hypothesis sets with the highest rank are output. The algorithm has been fully implemented, and tested on real images.> Ilan Shimshoni, Jean Ponce |
ICCV | 2 |
| 1995 | Invariant-Based Recognition of Complex Curved 3D Objects from Image ContoursabstractTo recognize three-dimensional objects bounded by smooth curved surfaces from monocular image contours, viewpoint-dependent image features must be related to object geometry. Contour bitangents and inflections along with associated parallel tangents points are the projection of surface points that lie on the occluding contour for a five-parameter family of scaled orthographic projection viewpoints. An invariant representation can be computed from these image features and seen for modeling and recognizing objects. Modeling is achieved by moving an object in front of a camera to obtain a curve of possible invariants. The relative camera-object motion is not required, and 3D models are not utilized. At recognition time, invariants computed from a single image are used to index the model database. Using the matched features, independent qualitative and quantitative verification procedures eliminate potential false matches. Examples from an implementation are presented.> B. Vijayakumar, David J. Kriegman, Jean Ponce |
ICCV | 3 |
| 1995 | New Techniques for Computing Four-Finger-Force-Closure Grasps of Polyhedral ObjectsabstractIt was shown in Ponce et al. (1993) that four-finger force-closure grasps fall into three categories: concurrent, pencil, and regulus grasps. The authors propose new techniques for computing these three types of grasps. The authors have implemented them and present examples. Attawith Sudsang, Jean Ponce |
ICRA | 2 |
| 1995 | On computing three-finger force-closure grasps of polygonal objectsabstractThis paper addresses the problem of computing stable grasps of 2-D polygonal objects. We consider the case of a hand equipped with three hard fingers and assume point contact with friction. We prove new sufficient conditions for equilibrium and force closure that are linear in the unknown grasp parameters. This reduces computing the stable grasp regions in configuration space to constructing the three-dimensional projection of a five-dimensional polytope. We present an efficient projection algorithm based on linear programming and variable elimination among linear constraints. Maximal object segments where fingers can be positioned independently while ensuring force closure are found by linear optimization within the grasp regions. The approach has been implemented and several examples are presented. Jean Ponce, Bernard Faverjon |
IEEE Trans. Robotics Autom. | 1 |
| 1994 | HOT curves for modelling and recognition of smooth curved 3D objectsabstractArbitrary smooth curved 3D shapes are represented by a discrete set of high-order tangent (HOT) curves, where a surface admits HOTs. These curves determine the structure of the image contours and its catastrophic changes, and there is a natural correspondence between some of them and monocular contour features such as inflections and bitangents. We present a method for automatically constructing the HOT curves from continuous sequences of video images and describe an approach to object recognition using viewpoint-dependent monocular image features as indices into a database of models and as a basis for pose estimation. We have implemented both of the methods, and present results obtained from real images.> Tanuja Joshi, Jean Ponce, B. Vijayakumar, David J. Kriegman |
CVPR | 2 |
| 1994 | Object representation for object recognitionabstractThis paper discusses some representation issues and challenges involved in object recognition. It is intended as a step toward assessing current object representation schemes and proposing design and evaluation criteria for future ones.> Jean Ponce, Ruzena Bajcsy, Dimitris N. Metaxas, Thomas O. Binford, David A. Forsyth, Martial Hebert, Katsushi Ikeuchi, Avinash C. Kak, Linda G. Shapiro, Stan Sclaroff, Alex Pentland, George C. Stockman |
CVPR | 1 |
| 1994 | Recovering the shape of polyhedra using line-drawing analysis and complex reflectance modelsabstractFollowing Sugihara, we represent the geometric constraints imposed by the line-drawing of a polyhedron as a set of linear equalities and inequalities. Unlike him, we explicitly take into account the uncertainty in vertex position. This allows us to circumvent the superstrictness of the constraints without deleting any of them. For a given error bound, deciding whether a line-drawing is the correct projection of a polyhedron is reduced to linear programming, and 3D shape recovery is reduced to optimization under linear constraints. Our method can be used for recovering the shape of polyhedral objects whose reflectance can be modelled accurately. We have implemented if for the Lambertian model and the Lambertian model with interreflections. We present results obtained using real images.> Ilan Shimshoni, Jean Ponce |
CVPR | 2 |
| 1994 | Analytical Methods for Uncalibrated Stereo and Motion Reconstruction
Jean Ponce, David H. Marimont, Todd A. Cass |
ECCV (1) | 1 |
| 1994 | Geometric Methods for Relative Reconstruction from Weakly Calibrated ImagesabstractWe present several new geometric methods for relative stereo and motion reconstruction using a discrete set of point correspondences. We suppose that the epipoles are known but do not assume any knowledge of the cameras' intrinsic or extrinsic parameters. In each case, we choose a set of five points as a basis for projective space and perform reconstruction relative to these five points. We also present a new technique for reprojection without reconstruction. We have implemented the proposed methods and present several examples using real images.> Jean Ponce, David H. Marimont, Todd A. Cass |
ICRA | 1 |
| 1994 | Using Geometric Distance Fits for 3-D Object Modeling and RecognitionabstractAddresses the problems of automatically constructing algebraic surface models from sets of 2D and 3D images and using these models in pose computation, motion and deformation estimation, and object recognition. We propose using a combination of constrained optimization and nonlinear least-squares estimation techniques to minimize the mean-squared geometric distance between a set of points or rays and a parameterized surface. In modeling tasks, the unknown parameters are the surface coefficients, while in pose and deformation estimation tasks they represent the transformation which maps the observer's coordinate system onto the modeled surface's own coordinate system. We have applied this approach to a variety of real range, computerized tomography and video images.> Steve Sullivan, Lorraine Sandford, Jean Ponce |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 1994 | Parameterized Families of Polynomials for Bounded Algebraic Curve and Surface FittingabstractInterest in algebraic curves and surfaces of high degree as geometric models or shape descriptors for different model-based computer vision tasks has increased in recent years, and although their properties make them a natural choice for object recognition and positioning applications, algebraic curve and surface fitting algorithms often suffer from instability problems. One of the main reasons for these problems is that, while the data sets are always bounded, the resulting algebraic curves or surfaces are, in most cases, unbounded. In this paper, the authors propose to constrain the polynomials to a family with bounded zero sets, and use only members of this family in the fitting process. For every even number d the authors introduce a new parameterized family of polynomials of degree d whose level sets are always bounded, in particular, its zero sets. This family has the same number of degrees of freedom as a general polynomial of the same degree. Three methods for fitting members of this polynomial family to measured data points are introduced. Experimental results of fitting curves to sets of points in R/sup 2/ and surfaces to sets of points in R/sup 3/ are presented.> Gabriel Taubin, Fernando Cukierman, Steve Sullivan, Jean Ponce, David J. Kriegman |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 1993 | Reconstruction of HOT curves from image sequencesabstractAn approach is presented for reconstructing two types of 3-D higher order tangency (HOT) curves from a sequence of images. These curves are useful for object recognition. The reconstruction results for bitangents are encouraging in comparison to those of inflections. They are probably more accurately reconstructed because they are readily located in images, and their common tangent is very accurately estimated from the point locations. This makes bitangents a good feature choice for recognition.> David J. Kriegman, B. Vijayakumar, Jean Ponce |
CVPR | 3 |
| 1993 | On using geometric distance fits to estimate 3D object shape, pose, and deformation from range, CT, and video imagesabstractThe problems of automatically constructing algebraic surface models from sets of 3D and 2D images and using these models in pose computation, motion and deformation estimation, and object recognition are addressed. It is proposed that a combination of constrained optimization and nonlinear least-squares estimation techniques be used to minimize the mean-squared geometric distance between a set of points or rays and a parameterized surface. In modeling tasks, the unknown parameters are the surface coefficients, while in pose and deformation estimation tasks they represent the transformation mapping the observer's coordinate system onto the modeled surface's own coordinate system. This approach is applied to a variety of real range, computerized tomography (CT), and video images.> Steve Sullivan, Lorraine Sandford, Jean Ponce |
CVPR | 3 |
| 1992 | Parametrizing and fitting bounded algebraic curves and surfacesabstractAn approach to fitting of implicit algebraic curves and surfaces to point data is introduced. Two families of polynomials with bounded zero sets are presented. Members of these families have the same number of degrees of freedom as general polynomials of the same degree. Methods for fitting members of these families of polynomials to measured data points are described. Experimental results for sets of points in R/sup 2/ and R/sup 3/ for curves and surfaces, respectively, are presented.> Gabriel Taubin, Fernando Cukierman, Steve Sullivan, Jean Ponce, David J. Kriegman |
CVPR | 4 |
| 1992 | Constraints for Recognizing and Locating Curved 3D Objects from Monocular Image Features
David J. Kriegman, B. Vijayakumar, Jean Ponce |
ECCV | 3 |
| 1992 | Computing Exact Aspect Graphs of Curved Objects: Algebraic Surfaces
Jean Ponce, Sylvain Petitjean, David J. Kriegman |
ECCV | 1 |
| 1992 | An algebraic approach to line-drawing analysis in the presence of uncertaintyabstractFollowing the work of K. Sugihara (1984), the authors represent the geometric constraints imposed by the line-drawing of a polyhedron as a set of linear equalities and inequalities. They, however, explicitly take into account the uncertainty in the vertex position. This allows the circumvention of the superstrictness of the constraints without deleting any constraints. For a given error bound, the condition whether a line-drawing is the correct projection of a polyhedron is reduced to linear programing, and the 3D shape recovery is reduced to optimization under linear constraints. The approach has been implemented, and examples are presented.> Jean Ponce, Ilan Shimshoni |
ICRA | 1 |
| 1992 | A System For Planning And Executing Two-finger Force-Closure Grasps Of Curved 2D ObjectsabstractThzs paper presents a system for plannzng and executzng stable grasps of curved two-damenszonal objects. We conszder the case of a hand equzpped wzth two hard fingers and assume poznt contact wzth frzctzon Objects are modelled by parametrzc curves, and force-closure grasps are characterzzed by systems of polynomzal constraints an the parameters of these curves. All configuratzon space regzons satzsfyzng these constraznfs are found by a parallel numerzcal cell decomposztzon algoriihni based on curve tracing and continadzon techniques ~IULLIILU~ object segiiienls diel e fingers can be positzoned an depeii dently are found by optimzzateon wzlhzn the grasp regaons The approuch has been implemented uszng a dzstrzbuted archztecture, and experzinents using a PUMA rohot equipped ubiih a pneumatzc two-finger grippe? aid U vzszoii system are presented Darrell Stam, Jean Ponce, Bernard Faverjon |
IROS | 2 |
| 1992 | On using CAD models to compute the pose of curved 3D objects
Jean Ponce, Anthony Hoogs, David J. Kriegman |
CVGIP Image Underst. | 1 |
| 1992 | Computing exact aspect graphs of curved objects: Algebraic surfaces
Sylvain Petitjean, Jean Ponce, David J. Kriegman |
Int. J. Comput. Vis. | 2 |
| 1991 | On computing two-finger force-closure grasps of curved 2D objectsabstractAn approach to the computation of stable grasps of curved two-dimensional objects is presented. The authors consider the case of a hand equipped with two hard fingers and assume point contact with friction. Objects are modeled by parametric curves, and force-closure grasps are characterized by systems of polynomial constraints in the parameters of these curves. All configuration space regions satisfying these constraints are found by a numerical cell decomposition algorithm based on curve tracing and continuation techniques. Maximal object segments where fingers can be positioned independently are found by optimization within the grasp regions. The approach has been implemented and examples are presented.> Bernard Faverjon, Jean Ponce |
ICRA | 2 |
| 1990 | Computing Exact Aspect Graphs of Curved Objects: Parametric Surfaces
Jean Ponce, David J. Kriegman |
AAAI | 1 |
| 1990 | On characterizing ribbons and finding skewed symmetries
Jean Ponce |
Comput. Vis. Graph. Image Process. | 1 |
| 1990 | Computing exact aspect graphs of curved objects: Solids of revolution
David J. Kriegman, Jean Ponce |
Int. J. Comput. Vis. | 2 |
| 1990 | Straight homogeneous generalized cylinders: Differential geometry and uniqueness results
Jean Ponce |
Int. J. Comput. Vis. | 1 |
| 1990 | On Recognizing and Positioning Curved 3-D Objects from Image ContoursabstractAn approach for explicitly relating the shape of image contours to models of curved three-dimensional objects is presented. This relationship is used for object recognition and positioning. Object models consist of collections of parametric surface patches and their intersection curves; this includes nearly all representations used in computer-aided geometric design and computer vision. The image contours considered are the projections of surface discontinuities and occluding contours. Elimination theory provides a method for constructing the implicit equation of these contours for an object observed under orthographic or perspective projection. This equation is parameterized by the object's position and orientation with respect to the observer. Determining these parameters is reduced to a fitting problem between the theoretical contour and the observed data points. The proposed approach readily extends to parameterized models. It has been implemented for a simple world composed of various surfaces of revolution and tested on several real images.> David J. Kriegman, Jean Ponce |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1989 | On characterizing ribbons and finding skewed symmetriesabstractThe author compares Blum, Brooks, and Brady ribbons, and proves that Blum and Brady ribbons are not, in general, Brooks ribbons. Conversely, he proves that Brook ribbons are, in general, neither Blum nor Brady ribbons. For Blum and Brooks ribbons, it is in principle trivial to decide whether two contour points may form a ribbon pair; they have to form a local symmetry. This property is not true for Brooks ribbons. Attention is also given to whether it is possible to characterize locally the pairs of contour points which form a Brooks ribbon pair. Using the curvature of a Brooks ribbon, it is shown that this is possible for some classes of Brooks ribbons, including skewed symmetries. This result is used in an implemented algorithm for finding skewed symmetries in an image, and examples of segmentation of real images are given.> Jean Ponce |
ICRA | 1 |
| 1989 | Invariant Properties of Straight Homogeneous Generalized Cylinders and Their ContoursabstractA fundamental group in computer vision is the recovery of three-dimensional shape from image data. While this problem is in general underconstrained, the authors show that it can be simplified in the case where the objects being viewed are generalized cylinders. They consider the class of straight homogeneous generalized cylinders (SHGCs), without any further assumption on the viewing direction or the precise shape of these objects. They present a rigorous mathematical study of the geometry of SHGCs and characterize their Gaussian curvature and occluding contours, and use these results to prove several new invariant properties of the contours of SHGCs. These properties are, in turn, used in two implemented algorithms for recovering SHGC descriptions from image contours. Several examples of segmentation of real images are given. Other applications are also discussed.> Jean Ponce, David M. Chelberg, Wallace B. Mann |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1988 | Straight homogeneous generalized cylinders: differential geometry and uniqueness resultsabstractThe author studies the differential geometry of straight homogeneous generalized cylinders (SHGCs). He derives a necessary and sufficient condition that an SHGC must verify to parameterize a regular surface, computes the Gaussian curvature of a regular SHGC, and proves that the parabolic lines of an SHGC are either meridians or parallels. Using these results, he addresses the following problem: under which conditions can a given surface have several descriptions by SHGCs? He proves several results. In particular, he proves that two SHGCs with the same cross-section plane and axis direction are necessarily deduced from each other through inverse scalings of their cross-sections and sweeping rule curve. He extends Shafer's pivot and slant theorems. Finally, he proves that a surface with at least two parabolic lines has at most three different SHGC descriptions, and that a surface with at least four parabolic lines has at most a unique SHGC description.> Jean Ponce |
CVPR | 1 |
| 1988 | Finding the limbs and cusps of generalized cylinders
Jean Ponce, David M. Chelberg |
Int. J. Comput. Vis. | 1 |
| 1987 | Finding the limbs and cusps of generalized cylindersabstractThis paper addresses the problem of finding analytically the limbs and cusps of generalized cylinders. Orthographic projections of generalized cylinders whose axis is straight and whose axis is an arbitrary 3D curve are considered in turn. In both cases, the general equations of the limbs and cusps are given. They are solved for three classes of generalized cylinders: solids of revolution, straight homogeneous generalized cylinders whose scaling sweeping rule is a polynomial of degree less than or equal to 5 and generalized cylinders whose axis is an arbitrary 3D curve but the cross section is circular and constant. Examples of limbs and cusps found for each class are given. Extensions and applications of the results presented are discussed. Jean Ponce, David M. Chelberg |
ICRA | 1 |
| 1987 | Localized intersections computation for solid modelling with straight homogenous generalized cylindersabstractThis paper reports progress in the development of a solid modelling system combining straight homogeneous generalized cylinders through set operations. Two basic components of this system are the modules which compute the set operations between primitives and display the resulting solids using ray tracing. These two modules are also very computationally intensive as they involve a large number of surface-surface and ray-surface intersections computations. We introduce a novel hierarchical representation for straight homogeneous cylinders called Box Tree. The Box Tree is analogous to a Quadtree in parameter space. It is an exact boundary representation which describes the surface of the associated generalized cylinder by a hierarchy of enclosing boxes. We use the Box Tree to efficiently compute the set operations and ray tracing algorithms by localizing the search for intersections to the regions where they may occur. We discuss complexity issues and illustrate the performances of our modelling system on a variety of examples. Jean Ponce, David M. Chelberg |
ICRA | 1 |
| 1987 | An object centered hierarchical representation for 3D objects: The prism tree
Jean Ponce, Olivier D. Faugeras |
Comput. Vis. Graph. Image Process. | 1 |
| 1985 | Toward a surface primal sketchabstractThis paper reports progress toward the development of a representation of significant surface changes in dense depth maps. We call tile representation the Surface Primal Sketch by analogy with representations of intensity changes, image structure, and changes in curvature of planar curves. We describe an implemented program that detects, localizes, and symbolically describes: steps, where the surface height function is discontinuous, and roofs, where the surface is continuous but the surface normal is discontinuous. We illustrate the performance of the program on range maps of objects of varying complexity. Jean Ponce, J. Michael Brady |
ICRA | 1 |
| 1985 | Describing surfaces
J. Michael Brady, Jean Ponce, Alan L. Yuille, Haruo Asada |
Comput. Vis. Graph. Image Process. | 2 |
| 1983 | Prism Trees: A Hierarchical Representation for 3-D Objects
Olivier D. Faugeras, Jean Ponce |
IJCAI | 2 |