Jean Ponce

dblp:p/JeanPonce · DBLP profile ↗
← Back
191ranked-venue papers
29as first author
24since 2021 · last 2025
0009-0000-5449-7620ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 180 · 26 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 108 · 13 first-author · 10 since 2021Systems, architecture and hardware · 26 · 8 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author
YearPublicationVenuePosition
2025 Online 3D Scene Reconstruction Using Neural Object Priors
abstract
This paper addresses the problem of reconstructing a scene online at the level of objects given an RGB-D video sequence. While current object-aware neural implicit rep-resentations hold promise, they are limited in online reconstruction efficiency and shape completion. Our main contributions to alleviate the above limitations are twofold. First, we propose a feature grid interpolation mechanism to continuously update grid-based object-centric neural implicit representations as new object parts are revealed. Second, we construct an object library with previously mapped objects in advance and leverage the corresponding shape priors to initialize geometric object models in new videos, sub-sequently completing them with novel views as well as synthesized past views to avoid losing original object details. Extensive experiments on synthetic environments from the Replica dataset, real-world ScanNet sequences and videos captured in our laboratory demonstrate that our approach outperforms state-of-the-art neural implicit models for this task in terms of reconstruction accuracy and completeness.
Thomas Chabal, Shizhe Chen, Jean Ponce, Cordelia Schmid
3DV3
2025 A New Statistical Model of Star Speckles for Learning to Detect and Characterize Exoplanets in Direct Imaging Observations
abstract
The search for exoplanets is an active field in astronomy, with direct imaging as one of the most challenging methods due to faint exoplanet signals buried within stronger residual starlight. Successful detection requires advanced image processing to separate the exoplanet signal from this nuisance component. This paper presents a novel statistical model that captures nuisance fluctuations using a multi-scale approach, leveraging problem symmetries and a joint spectral channel representation grounded in physical principles. Our model integrates into an interpretable, end-to-end learnable framework for simultaneous exoplanet detection and flux estimation. The proposed algorithm is evaluated against the state of the art using datasets from the SPHERE instrument operating at the Very Large Telescope (VLT). It significantly improves the precision-recall trade-off, notably on challenging datasets that are otherwise unusable by astronomers. The proposed approach is computationally efficient, robust to varying data quality, and well suited for large-scale observational surveys.1
Théo Bodrito, Olivier Flasseur, Julien Mairal, Jean Ponce, Maud Langlois, Anne-Marie Lagrange
CVPR4
2025 Generalized portrait quality assessment
Nicolas Chahine, Sira Ferradans, Javier Vazquez-Corral, Jean Ponce
Pattern Recognit. Lett.4
2024 Dense Optical Tracking: Connecting the Dots
abstract
Recent approaches to point tracking are able to recover the trajectory of any scene point through a large portion of a video despite the presence of occlusions. They are, how-ever, too slow in practice to track every point observed in a single frame in a reasonable amount of time. This paper introduces DOT, a novel, simple and efficient method for solving this problem. It first extracts a small set of tracks from key regions at motion boundaries using an off-the-shelf point tracking algorithm. Given source and target frames, DOT then computes rough initial estimates of a dense flow field and visibility mask through nearest-neighbor inter-polation, before refining them using a learnable optical flow estimator that explicitly handles occlusions and can be trained on synthetic data with ground-truth correspon-dences. We show that DOT is significantly more accurate than current optical flow techniques, outperforms sophis-ticated “universal” trackers like OmniMotion, and is on par with, or better than, the best point tracking algorithms like CoTracker while being at least two orders of magnitude faster. Quantitative and qualitative experiments with syn-thetic and real videos validate the promise of the proposed approach. Code, data, and videos showcasing the capabili-ties of our approach are available in the project webpage.11https://161ernoing.github.io/dot
Guillaume Le Moing, Jean Ponce, Cordelia Schmid
CVPR2
2023 An Image Quality Assessment Dataset for Portraits
abstract
Year after year, the demand for ever-better smartphone photos continues to grow, in particular in the domain of portrait photography. Manufacturers thus use perceptual quality criteria throughout the development of smartphone cameras. This costly procedure can be partially replaced by automated learning-based methods for image quality assessment (IQA). Due to its subjective nature, it is necessary to estimate and guarantee the consistency of the IQA process, a characteristic lacking in the mean opinion scores (MOS) widely used for crowdsourcing IQA. In addition, existing blind IQA (BIQA) datasets pay little attention to the difficulty of cross-content assessment, which may degrade the quality of annotations. This paper introduces PIQ23, a portrait-specific IQA dataset of 5116 images of 50 predefined scenarios acquired by 100 smartphones, covering a high variety of brands, models, and use cases. The dataset includes individuals of various genders and ethnicities who have given explicit and informed consent for their photographs to be used in public research. It is annotated by pairwise comparisons (PWC) collected from over 30 image quality experts for three image attributes: face detail preservation, face target exposure, and overall image quality. An in-depth statistical analysis of these annotations allows us to evaluate their consistency over PIQ23. Finally, we show through an extensive comparison with existing baselines that semantic information (image context) can be used to improve IQA predictions. The dataset along with the proposed statistical analysis and BIQA algorithms are available: https://github.com/DXOMARK-Research/PIQ2023
Nicolas Chahine, Ana-Stefania Calarasanu, Davide Garcia-Civiero, Théo Cayla, Sira Ferradans, Jean Ponce
CVPR6
2023 WALDO: Future Video Synthesis using Object Layer Decomposition and Parametric Flow Prediction
abstract
This paper presents WALDO (WArping Layer-Decomposed Objects), a novel approach to the prediction of future video frames from past ones. Individual images are decomposed into multiple layers combining object masks and a small set of control points. The layer structure is shared across all frames in each video to build dense inter-frame connections. Complex scene motions are modeled by combining parametric geometric transformations associated with individual layers, and video synthesis is broken down into discovering the layers associated with past frames, predicting the corresponding transformations for upcoming ones and warping the associated object regions accordingly, and filling in the remaining image parts. Extensive experiments on multiple benchmarks including urban videos (Cityscapes and KITTI) and videos featuring nonrigid motions (UCF-Sports and H3.6M), show that our method consistently outperforms the state of the art by a significant margin in every case. Code, pretrained models, and video samples synthesized by our approach can be found in the project webpage.1
Guillaume Le Moing, Jean Ponce, Cordelia Schmid
ICCV2
2023 Learning Reward Functions for Robotic Manipulation by Observing Humans
abstract
Observing a human demonstrator manipulate objects provides a rich, scalable and inexpensive source of data for learning robotic policies. However, transferring skills from human videos to a robotic manipulator poses several challenges, not least a difference in action and observation spaces. In this work, we use unlabeled videos of humans solving a wide range of manipulation tasks to learn a task-agnostic reward function for robotic manipulation policies. Thanks to the diversity of this training data, the learned reward function sufficiently generalizes to image observations from a previously unseen robot embodiment and environment to provide a meaningful prior for directed exploration in reinforcement learning. We propose two methods for scoring states relative to a goal image: through direct temporal regression, and through distances in an embedding space obtained with time-contrastive learning. By conditioning the function on a goal image, we are able to reuse one model across a variety of tasks. Unlike prior work on leveraging human videos to teach robots, our method, Human Offline Learned Distances (HOLD) requires neither a priori data from the robot environment, nor a set of task-specific human demonstrations, nor a predefined notion of correspondence across morphologies, yet it is able to accelerate training of several manipulation tasks on a simulated robot arm compared to using only a sparse reward obtained from task completion.
Minttu Alakuijala, Gabriel Dulac-Arnold, Julien Mairal, Jean Ponce, Cordelia Schmid
ICRA4
2023 A minimum swept-volume metric structure for configuration space
abstract
Borrowing elementary ideas from solid mechanics and differential geometry, this presentation shows that the volume swept by a regular solid undergoing a wide class of volume-preserving deformations induces a rather natural metric structure with well-defined and computable geodesics on its configuration space. This general result applies to concrete classes of articulated objects such as robot manipulators, and we demonstrate as a proof of concept the computation of geodesic paths for a free flying rod and planar robotic arms as well as their use in path planning with many obstacles.
Yann de Mont-Marin, Jean Ponce, Jean-Paul Laumond
ICRA2
2023 Revisiting Deformable Convolution for Depth Completion
abstract
Depth completion, which aims to generate high-quality dense depth maps from sparse depth maps, has attracted increasing attention in recent years. Previous work usually employs RGB images as guidance, and introduces iterative spatial propagation to refine estimated coarse depth maps. However, most of the propagation refinement methods require several iterations and suffer from a fixed receptive field, which may contain irrelevant and useless information with very sparse input. In this paper, we address these two challenges simultaneously by revisiting the idea of deformable convolution. We propose an effective architecture that leverages deformable kernel convolution as a single-pass refinement module, and empirically demonstrate its superiority. To better understand the function of deformable convolution and exploit it for depth completion, we further systematically investigate a variety of representative strategies. Our study reveals that, different from prior work, deformable convolution needs to be applied on an estimated depth map with a relatively high density for better performance. We evaluate our model on the large-scale KITTI dataset and achieve state-of-the-art level performance in both accuracy and inference speed. Our code is available at https://github.com/AlexSunNiklReDC.
Xinglong Sun, Jean Ponce, Yu-Xiong Wang
IROS2
2022 Active Learning Strategies for Weakly-Supervised Object Detection
Huy V. Vo, Oriane Siméoni, Spyros Gidaris, Andrei Bursuc, Patrick Pérez, Jean Ponce
ECCV (30)6
2022 VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning
Adrien Bardes, Jean Ponce, Yann LeCun
ICLR2
2022 Assembly Planning from Observations under Physical Constraints
abstract
This paper addresses the problem of copying an unknown assembly of primitives with known shape and appearance using information extracted from a single photograph by an off-the-shelf procedure for object detection and pose estimation. The proposed algorithm uses a simple combination of physical stability constraints, convex optimization and Monte Carlo tree search to plan assemblies as sequences of pick-and-place operations represented by STRIPS operators. It is efficient and, most importantly, robust to the errors in object detection and pose estimation unavoidable in any real robotic system. The proposed approach is demonstrated with thorough experiments on a UR5 manipulator.
Thomas Chabal, Robin Strudel, Etienne Arlaud, Jean Ponce, Cordelia Schmid
IROS4
2022 VICRegL: Self-Supervised Learning of Local Visual Features
abstract
Most recent self-supervised methods for learning image representations focus on either producing a global feature with invariance properties, or producing a set of local features. The former works best for classification tasks while the latter is best for detection and segmentation tasks. This paper explores the fundamental trade-off between learning local and global features. A new method called VICRegL is proposed that learns good global and local features simultaneously, yielding excellent performance on detection and segmentation tasks while maintaining good performance on classification tasks. Concretely, two identical branches of a standard convolutional net architecture are fed two differently distorted versions of the same image. The VICReg criterion is applied to pairs of global feature vectors. Simultaneously, the VICReg criterion is applied to pairs of local feature vectors occurring before the last pooling layer. Two local feature vectors are attracted to each other if their l2-distance is below a threshold or if their relative locations are consistent with a known geometric transformation between the two input images. We demonstrate strong performance on linear classification and segmentation transfer tasks. Code and pretrained models are publicly available at: https://github.com/facebookresearch/VICRegL
Adrien Bardes, Jean Ponce, Yann LeCun
NeurIPS2
2022 Homography-Based Minimal-Case Relative Pose Estimation With Known Gravity Direction
abstract
In this paper, we propose a novel approach to two-view minimal-case relative pose problems based on homography with known gravity direction. This case is relevant to smart phones, tablets, and other camera-IMU (Inertial measurement unit) systems which have accelerometers to measure the gravity vector. We explore the rank-1 constraint on the difference between the euclidean homography matrix and the corresponding rotation, and propose an efficient two-step solution for solving both the calibrated and semi-calibrated (unknown focal length) problems. Based on the hidden variable technique, we convert the problems to the polynomial eigenvalue problems, and derive new 3.5-point, 3.5-point, 4-point solvers for two cameras such that the two focal lengths are unknown but equal, one of them is unknown, and both are unknown and possibly different, respectively. We present detailed analyses and comparisons with the existing 6- and 7-point solvers, including results with smart phone images.
Yaqing Ding 0001, Jian Yang 0003, Jean Ponce, Hui Kong 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Learning Semantic Correspondence Exploiting an Object-Level Prior
abstract
We address the problem of semantic correspondence, that is, establishing a dense flow field between images depicting different instances of the same object or scene category. We propose to use images annotated with binary foreground masks and subjected to synthetic geometric deformations to train a convolutional neural network (CNN) for this task. Using these masks as part of the supervisory signal provides an object-level prior for the semantic correspondence task and offers a good compromise between semantic flow methods, where the amount of training data is limited by the cost of manually selecting point correspondences, and semantic alignment ones, where the regression of a single global geometric transformation between images may be sensitive to image-specific details such as background clutter. We propose a new CNN architecture, dubbed SFNet, which implements this idea. It leverages a new and differentiable version of the argmax function for end-to-end training, with a loss that combines mask and flow consistency with smoothness terms. Experimental results demonstrate the effectiveness of our approach, which significantly outperforms the state of the art on standard benchmarks.
Junghyup Lee, Dohyung Kim 0006, Wonkyung Lee, Jean Ponce, Bumsub Ham
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 High dynamic range and super-resolution from raw image bursts
abstract
Photographs captured by smartphones and mid-range cameras have limited spatial resolution and dynamic range, with noisy response in underexposed regions and color artefacts in saturated areas. This paper introduces the first approach (to the best of our knowledge) to the reconstruction of highresolution, high-dynamic range color images from raw photographic bursts captured by a handheld camera with exposure bracketing. This method uses a physically-accurate model of image formation to combine an iterative optimization algorithm for solving the corresponding inverse problem with a learned image representation for robust alignment and a learned natural image prior. The proposed algorithm is fast, with low memory requirements compared to state-of-the-art learning-based approaches to image restoration, and features that are learned end to end from synthetic yet realistic data. Extensive experiments demonstrate its excellent performance with super-resolution factors of up to ×4 on real photographs taken in the wild with hand-held cameras, and high robustness to low-light conditions, noise, camera shake, and moderate object motion.
Bruno Lecouat, Thomas Eboli, Jean Ponce, Julien Mairal
ACM Trans. Graph.3
2021 Localizing Objects with Self-supervised Transformers and no Labels
Oriane Siméoni, Gilles Puy, Huy V. Vo, Simon Roburin, Spyros Gidaris, Andrei Bursuc, Patrick Pérez, Renaud Marlet, Jean Ponce
BMVC9
2021 Lucas-Kanade Reloaded: End-to-End Super-Resolution from Raw Image Bursts
abstract
This presentation addresses the problem of reconstructing a high-resolution image from multiple lower-resolution snapshots captured from slightly different viewpoints in space and time. Key challenges for solving this super-resolution problem include (i) aligning the input pictures with sub-pixel accuracy, (ii) handling raw (noisy) images for maximal faithfulness to native camera data, and (iii) designing/learning an image prior (regularizer) well suited to the task. We address these three challenges with a hybrid algorithm building on the insight from [45] that aliasing is an ally in this setting, with parameters that can be learned end to end, while retaining the interpretability of classical approaches to inverse problems. The effectiveness of our approach is demonstrated on synthetic and real image bursts, setting a new state of the art on several benchmarks and delivering excellent qualitative results on real raw bursts captured by smartphones and prosumer cameras. Our code is available at https://github.com/bruno-31/lkburst.git.
Bruno Lecouat, Jean Ponce, Julien Mairal
ICCV2
2021 Unsupervised Layered Image Decomposition into Object Prototypes
abstract
We present an unsupervised learning framework for decomposing images into layers of automatically discovered object models. Contrary to recent approaches that model image layers with autoencoder networks, we represent them as explicit transformations of a small set of prototypical images. Our model has three main components: (i) a set of object prototypes in the form of learnable images with a transparency channel, which we refer to as sprites; (ii) differentiable parametric functions predicting occlusions and transformation parameters necessary to instantiate the sprites in a given image; (iii) a layered image formation model with occlusion for compositing these instances into complete images including background. By jointly learning the sprites and occlusion/transformation predictors to reconstruct images, our approach not only yields accurate layered image decompositions, but also identifies object categories and instance parameters. We first validate our approach by providing results on par with the state of the art on standard multiobject synthetic benchmarks (Tetrominoes, Multi-dSprites, CLEVR6). We then demonstrate the applicability of our model to real images in tasks that include clustering (SVHN, GTSRB), cosegmentation (Weizmann Horse) and object discovery from unfiltered social network images. To the best of our knowledge, our approach is the first layered image decomposition algorithm that learns an explicit and shared concept of object type, and is robust enough to be applied to real images.
Tom Monnier, Elliot Vincent, Jean Ponce, Mathieu Aubry
ICCV3
2021 Equality Constrained Differential Dynamic Programming
abstract
Trajectory optimization is an important tool in task-based robot motion planning, due to its generality and convergence guarantees under some mild conditions. It is often used as a post-processing operation to smooth out trajectories that are generated by probabilistic methods or to directly control the robot motion. Unconstrained trajectory optimization problems have been well studied, and are commonly solved using Differential Dynamic Programming methods that allow for fast convergence at a relatively low computational cost. In this paper, we propose an augmented Lagrangian approach that extends these ideas to equality-constrained trajectory optimization problems, while maintaining a balance between convergence speed and numerical stability. We illustrate our contributions on various standard robotic problems and highlights their benefits compared to standard approaches.
Sarah El Kazdadi, Justin Carpentier, Jean Ponce
ICRA3
2021 Online Learning and Control of Complex Dynamical Systems from Sensory Input
abstract
Identifying an effective model of a dynamical system from sensory data and using it for future state prediction and control is challenging. Recent data-driven algorithms based on Koopman theory are a promising approach to this problem, but they typically never update the model once it has been identified from a relatively small set of observation, thus making long-term prediction and control difficult for realistic systems, in robotics or fluid mechanics for example. This paper introduces a novel method for learning an embedding of the state space with linear dynamics from sensory data. Unlike previous approaches, the dynamics model can be updated online and thus easily applied to systems with non-linear dynamics in the original configuration space. The proposed approach is evaluated empirically on several classical dynamical systems and sensory modalities, with good performance on long-term prediction and control.
Oumayma Bounou, Jean Ponce, Justin Carpentier
NeurIPS2
2021 CCVS: Context-aware Controllable Video Synthesis
abstract
This presentation introduces a self-supervised learning approach to the synthesis of new videos clips from old ones, with several new key elements for improved spatial resolution and realism: It conditions the synthesis process on contextual information for temporal continuity and ancillary information for fine control. The prediction model is doubly autoregressive, in the latent space of an autoencoder for forecasting, and in image space for updating contextual information, which is also used to enforce spatio-temporal consistency through a learnable optical flow module. Adversarial training of the autoencoder in the appearance and temporal domains is used to further improve the realism of its output. A quantizer inserted between the encoder and the transformer in charge of forecasting future frames in latent space (and its inverse inserted between the transformer and the decoder) adds even more flexibility by affording simple mechanisms for handling multimodal ancillary information for controlling the synthesis process (e.g., a few sample frames, an audio track, a trajectory in image space) and taking into account the intrinsically uncertain nature of the future by allowing multiple predictions. Experiments with an implementation of the proposed approach give very good qualitative and quantitative results on multiple tasks and standard benchmarks.
Guillaume Le Moing, Jean Ponce, Cordelia Schmid
NeurIPS2
2021 Large-Scale Unsupervised Object Discovery
abstract
Existing approaches to unsupervised object discovery (UOD) do not scale up to large datasets without approximations that compromise their performance. We propose a novel formulation of UOD as a ranking problem, amenable to the arsenal of distributed methods available for eigenvalue problems and link analysis. Through the use of self-supervised features, we also demonstrate the first effective fully unsupervised pipeline for UOD. Extensive experiments on COCO~\cite{Lin2014cocodataset} and OpenImages~\cite{openimages} show that, in the single-object discovery setting where a single prominent object is sought in each image, the proposed LOD (Large-scale Object Discovery) approach is on par with, or better than the state of the art for medium-scale datasets (up to 120K images), and over 37\% better than the only other algorithms capable of scaling up to 1.7M images. In the multi-object discovery setting where multiple objects are sought in each image, the proposed LOD is over 14\% better in average precision (AP) than all other methods for datasets ranging from 20K to 1.7M images. Using self-supervised features, we also show that the proposed method obtains state-of-the-art UOD performance on OpenImages.
Huy V. Vo, Elena Sizikova, Cordelia Schmid, Patrick Pérez, Jean Ponce
NeurIPS5
2021 Deformable Kernel Networks for Joint Image Filtering
Jean Ponce, Bumsub Ham
Int. J. Comput. Vis.2
2020 Minimal Solutions to Relative Pose Estimation From Two Views Sharing a Common Direction With Unknown Focal Length
abstract
We propose minimal solutions to relative pose estimation problem from two views sharing a common direction with unknown focal length. This is relevant for cameras equipped with an IMU (inertial measurement unit), e.g., smart phones, tablets. Similar to the 6-point algorithm for two cameras with unknown but equal focal lengths and 7-point algorithm for two cameras with different and unknown focal lengths, we derive new 4- and 5-point algorithms for these two cases, respectively. The proposed algorithms can cope with coplanar points, which is a degenerate configuration for these 6- and 7-point counterparts. We present a detailed analysis and comparisons with the state of the art. Experimental results on both synthetic data and real images from a smart phone demonstrate the usefulness of the proposed algorithms.
Yaqing Ding 0001, Jian Yang 0003, Jean Ponce, Hui Kong 0001
CVPR3
2020 End-to-end Interpretable Learning of Non-blind Image Deblurring
Thomas Eboli, Jian Sun 0009, Jean Ponce
ECCV (17)3
2020 Fully Trainable and Interpretable Non-local Sparse Models for Image Restoration
Bruno Lecouat, Jean Ponce, Julien Mairal
ECCV (22)2
2020 Learning to Compose Hypercolumns for Visual Correspondence
Juhong Min, Jongmin Lee 0005, Jean Ponce, Minsu Cho
ECCV (15)3
2020 Toward Unsupervised, Multi-object Discovery in Large-Scale Image Collections
Huy V. Vo, Patrick Pérez, Jean Ponce
ECCV (23)3
2020 A Flexible Framework for Designing Trainable Priors with Adaptive Smoothing and Game Encoding
abstract
We introduce a general framework for designing and training neural network layers whose forward passes can be interpreted as solving non-smooth convex optimization problems, and whose architectures are derived from an optimization algorithm. We focus on convex games, solved by local agents represented by the nodes of a graph and interacting through regularization functions. This approach is appealing for solving imaging problems, as it allows the use of classical image priors within deep models that are trainable end to end. The priors used in this presentation include variants of total variation, Laplacian regularization, bilateral filtering, sparse coding on learned dictionaries, and non-local self similarities. Our models are fully interpretable as well as parameter and data efficient. Our experiments demonstrate their effectiveness on a large diversity of tasks ranging from image denoising and compressed sensing for fMRI to dense stereo matching.
Bruno Lecouat, Jean Ponce, Julien Mairal
NeurIPS2
2019 SFNet: Learning Object-Aware Semantic Correspondence
abstract
We address the problem of semantic correspondence, that is, establishing a dense flow field between images depicting different instances of the same object or scene category. We propose to use images annotated with binary foreground masks and subjected to synthetic geometric deformations to train a convolutional neural network (CNN) for this task. Using these masks as part of the supervisory signal offers a good compromise between semantic flow methods, where the amount of training data is limited by the cost of manually selecting point correspondences, and semantic alignment ones, where the regression of a single global geometric transformation between images may be sensitive to image-specific details such as background clutter. We propose a new CNN architecture, dubbed SFNet, which implements this idea. It leverages a new and differentiable version of the argmax function for end-to-end training, with a loss that combines mask and flow consistency with smoothness terms. Experimental results demonstrate the effectiveness of our approach, which significantly outperforms the state of the art on standard benchmarks.
Junghyup Lee, Dohyung Kim 0006, Jean Ponce, Bumsub Ham
CVPR3
2019 Coordinate-Free Carlsson-Weinshall Duality and Relative Multi-View Geometry
abstract
We present a coordinate-free description of Carlsson-Weinshall duality between scene points and camera pinholes and use it to derive a new characterization of primal/dual multi-view geometry. In the case of three views, a particular set of reduced trilinearities provide a novel parameterization of camera geometry that, unlike existing ones, is subject only to very simple internal constraints. These trilinearities lead to new "quasi-linear" algorithms for primal and dual structure from motion. We include some preliminary experiments with real and synthetic data.
Matthew Trager, Martial Hebert, Jean Ponce
CVPR3
2019 Unsupervised Image Matching and Object Discovery as Optimization
abstract
Learning with complete or partial supervision is power- ful but relies on ever-growing human annotation efforts. As a way to mitigate this serious problem, as well as to serve specific applications, unsupervised learning has emerged as an important field of research. In computer vision, unsu- pervised learning comes in various guises. We focus here on the unsupervised discovery and matching of object cate- gories among images in a collection, following the work of Cho et al. [12]. We show that the original approach can be reformulated and solved as a proper optimization problem. Experiments on several benchmarks establish the merit of our approach.
Huy V. Vo, Francis R. Bach, Minsu Cho, Kai Han 0001, Yann LeCun, Patrick Pérez, Jean Ponce
CVPR7
2019 An Efficient Solution to the Homography-Based Relative Pose Problem With a Common Reference Direction
abstract
In this paper, we propose a novel approach to two-view minimal-case relative pose problems based on homography with a common reference direction. We explore the rank-1 constraint on the difference between the Euclidean homography matrix and the corresponding rotation, and propose an efficient two-step solution for solving both the calibrated and partially calibrated (unknown focal length) problems. We derive new 3.5-point, 3.5-point, 4-point solvers for two cameras such that the two focal lengths are unknown but equal, one of them is unknown, and both are unknown and possibly different, respectively. We present detailed analyses and comparisons with existing 6 and 7-point solvers, including results with smart phone images.
Yaqing Ding 0001, Jian Yang 0003, Jean Ponce, Hui Kong 0001
ICCV3
2019 Hyperpixel Flow: Semantic Correspondence With Multi-Layer Neural Features
abstract
Establishing visual correspondences under large intra-class variations requires analyzing images at different levels, from features linked to semantics and context to local patterns, while being invariant to instance-specific details. To tackle these challenges, we represent images by “hyperpixels” that leverage a small number of relevant features selected among early to late layers of a convolutional neural network. Taking advantage of the condensed features of hyperpixels, we develop an effective real-time matching algorithm based on Hough geometric voting. The proposed method, hyperpixel flow, sets a new state of the art on three standard benchmarks as well as a new dataset, SPair-71k, which contains a significantly larger number of image pairs than existing datasets, with more accurate and richer annotations for in-depth analysis.
Juhong Min, Jongmin Lee 0005, Jean Ponce, Minsu Cho
ICCV3
2019 Build your own hybrid thermal/EO camera for autonomous vehicle
abstract
In this work, we propose a novel paradigm to design a hybrid thermal/EO (Electro-Optical or visible-light) camera, whose thermal and RGB frames are pixel-wisely aligned and temporally synchronized. Compared with the existing schemes, we innovate in three ways in order to make it more compact in dimension, and thus more practical and extendable for real-world applications. The first is a redesign of the structure layout of the thermal and EO cameras. The second is on obtaining a pixel-wise spatial registration of the thermal and RGB frames by a coarse mechanical adjustment and a fine alignment through a constant homography warping. The third innovation is on extending one single hybrid camera to a hybrid camera array, through which we can obtain wide-view spatially aligned thermal, RGB and disparity images simultaneously. The experimental results show that the average error of spatial-alignment of two image modalities can be less than one pixel.
Yigong Zhang, Shuo Gu, Yubin Guo, Minghao Liu 0003, Zezhou Sun, Zhixing Hou, Ying Wang 0007, Jian Yang 0003, Jean Ponce, Hui Kong 0001
ICRA11
2018 On the Solvability of Viewing Graphs
Matthew Trager, Brian Osserman, Jean Ponce
ECCV (16)3
2018 Dijkstra Model for Stereo-Vision Based Road Detection: A Non-Parametric Method
abstract
This paper proposes a new method for detecting a road from a stereo pair of images. First, the horizon is accurately estimated by a robust, weighted-sampling RANSAC-like method in the improved v-disparity map. The vanishing point of the road region is located using both the horizon information and road flatness constraints. Then it is used as the source node of a weighted graph formed by the pixels of the left stereo-image and their adjacency relationships. The weight of each edge measures the inconsistency of adjacent pixels, and is computed using both the gray-scale and disparity information. Detecting road borders is thus reduced to finding two shortest paths from the source node to the bottom row of the image by the Dijkstra algorithm. The proposed method has been tested on 2621 image pairs of different road scenes from the KITTI dataset. Our experiments demonstrate that this training free approach detects horizon, vanishing point, and road region accurately and robustly, and compares favorably with the state of the art on the KITTI benchmark.
Yigong Zhang, Jian Yang 0003, Jean Ponce, Hui Kong 0001
ICRA3
2018 Robust Guided Image Filtering Using Nonconvex Potentials
abstract
Filtering images using a guidance signal, a process called guided or joint image filtering, has been used in various tasks in computer vision and computational photography, particularly for noise reduction and joint upsampling. This uses an additional guidance signal as a structure prior, and transfers the structure of the guidance signal to an input image, restoring noisy or altered image structure. The main drawbacks of such a data-dependent framework are that it does not consider structural differences between guidance and input images, and that it is not robust to outliers. We propose a novel SD (for static/dynamic) filter to address these problems in a unified framework, and jointly leverage structural information from guidance and input images. Guided image filtering is formulated as a nonconvex optimization problem, which is solved by the majorize-minimization algorithm. The proposed algorithm converges quickly while guaranteeing a local minimum. The SD filter effectively controls the underlying image structure at different scales, and can handle a variety of types of data from different sensors. It is robust to outliers and other artifacts such as gradient reversal and global intensity shift, and has good edge-preserving smoothing properties. We demonstrate the flexibility and effectiveness of the proposed SD filter in a variety of applications, including depth upsampling, scale-space filtering, texture removal, flash/non-flash denoising, and RGB/NIR denoising.
Bumsub Ham, Minsu Cho, Jean Ponce
IEEE Trans. Pattern Anal. Mach. Intell.3
2018 Proposal Flow: Semantic Correspondences from Object Proposals
abstract
Finding image correspondences remains a challenging problem in the presence of intra-class variations and large changes in scene layout. Semantic flow methods are designed to handle images depicting different instances of the same object or scene category. We introduce a novel approach to semantic flow, dubbed proposal flow, that establishes reliable correspondences using object proposals. Unlike prevailing semantic flow approaches that operate on pixels or regularly sampled local regions, proposal flow benefits from the characteristics of modern object proposals, that exhibit high repeatability at multiple scales, and can take advantage of both local and geometric consistency constraints among proposals. We also show that the corresponding sparse proposal flow can effectively be transformed into a conventional dense flow field. We introduce two new challenging datasets that can be used to evaluate both general semantic flow techniques and region-based approaches such as proposal flow. We use these benchmarks to compare different matching algorithms, object proposals, and region features within proposal flow, to the state of the art in semantic flow. This comparison, along with experiments on standard datasets, demonstrates that proposal flow significantly outperforms existing semantic flow methods in various settings.
Bumsub Ham, Minsu Cho, Cordelia Schmid, Jean Ponce
IEEE Trans. Pattern Anal. Mach. Intell.4
2018 When Dijkstra Meets Vanishing Point: A Stereo Vision Approach for Road Detection
abstract
In this paper, we propose a vanishing-point constrained Dijkstra road model for road detection in a stereo-vision paradigm. First, the stereo-camera is used to generate the u- and v-disparity maps of road image, from which the horizon can be extracted. With the horizon and ground region constraints, we can robustly locate the vanishing point of road region. Second, a weighted graph is constructed using all pixels of the image, and the detected vanishing point is treated as the source node of the graph. By computing a vanishing-point constrained Dijkstra minimum-cost map, where both disparity and gradient of gray image are used to calculate cost between two neighbor pixels, the problem of detecting road borders in image is transformed into that of finding two shortest paths that originate from the vanishing point to two pixels in the last row of image. The proposed approach has been implemented and tested over 2600 grayscale images of different road scenes in the KITTI data set. The experimental results demonstrate that this training-free approach can detect horizon, vanishing point, and road regions very accurately and robustly. It can achieve promising performance.
Yigong Zhang, Yingna Su, Jian Yang 0003, Jean Ponce, Hui Kong 0001
IEEE Trans. Image Process.4
2017 Kernel Square-Loss Exemplar Machines for Image Retrieval
abstract
Zepeda and Perez [41] have recently demonstrated the promise of the exemplar SVM (ESVM) as a feature encoder for image retrieval. This paper extends this approach in several directions: We first show that replacing the hinge loss by the square loss in the ESVM cost function significantly reduces encoding time with negligible effect on accuracy. We call this model square-loss exemplar machine, or SLEM. We then introduce a kernelized SLEM which can be implemented efficiently through low-rank matrix decomposition, and displays improved performance. Both SLEM variants exploit the fact that the negative examples are fixed, so most of the SLEM computational complexity is relegated to an offline process independent of the positive examples. Our experiments establish the performance and computational advantages of our approach using a large array of base features and standard image retrieval datasets.
Rafael S. Rezende, Joaquin Zepeda, Jean Ponce, Francis R. Bach, Patrick Pérez
CVPR3
2017 General Models for Rational Cameras and the Case of Two-Slit Projections
abstract
The rational camera model recently introduced in [18] provides a general methodology for studying abstract nonlinear imaging systems and their multi-view geometry. This paper builds on this framework to study physical realizations of rational cameras. More precisely, we give an explicit account of the mapping between between physical visual rays and image points (missing in the original description), which allows us to give simple analytical expressions for direct and inverse projections. We also consider primitive camera models, that are orbits under the action of various projective transformations, and lead to a general notion of intrinsic parameters. The methodology is general, but it is illustrated concretely by an in-depth study of two-slit cameras, that we model using pairs of linear projections. This simple analytical form allows us to describe models for the corresponding primitive cameras, to introduce intrinsic parameters with a clear geometric meaning, and to define an epipolar tensor characterizing two-view correspondences. In turn, this leads to new algorithms for structure from motion and self-calibration.
Matthew Trager, Bernd Sturmfels, John F. Canny, Martial Hebert, Jean Ponce
CVPR5
2017 SCNet: Learning Semantic Correspondence
abstract
This paper addresses the problem of establishing semantic correspondences between images depicting different instances of the same object or scene category. Previous approaches focus on either combining a spatial regularizer with hand-crafted features, or learning a correspondence model for appearance only. We propose instead a convolutional neural network architecture, called SCNet, for learning a geometrically plausible model for semantic correspondence. SCNet uses region proposals as matching primitives, and explicitly incorporates geometric consistency in its loss function. It is trained on image pairs obtained from the PASCAL VOC 2007 keypoint dataset, and a comparative evaluation on several standard benchmarks demonstrates that the proposed approach substantially outperforms both recent deep learning architectures and previous methods based on hand-crafted features.
Kai Han 0001, Rafael S. Rezende, Bumsub Ham, Kwan-Yee Kenneth Wong, Minsu Cho, Cordelia Schmid, Jean Ponce
ICCV7
2016 Proposal Flow
abstract
Finding image correspondences remains a challenging problem in the presence of intra-class variations and large changes in scene layout. Semantic flow methods are designed to handle images depicting different instances of the same object or scene category. We introduce a novel approach to semantic flow, dubbed proposal flow, that establishes reliable correspondences using object proposals. Unlike prevailing semantic flow approaches that operate on pixels or regularly sampled local regions, proposal flow benefits from the characteristics of modern object proposals, that exhibit high repeatability at multiple scales, and can take advantage of both local and geometric consistency constraints among proposals. We also show that proposal flow can effectively be transformed into a conventional dense flow field. We introduce a new dataset that can be used to evaluate both general semantic flow techniques and region-based approaches such as proposal flow. We use this benchmark to compare different matching algorithms, object proposals, and region features within proposal flow, to the state of the art in semantic flow. This comparison, along with experiments on standard datasets, demonstrates that proposal flow significantly outperforms existing semantic flow methods in various settings.
Bumsub Ham, Minsu Cho, Cordelia Schmid, Jean Ponce
CVPR4
2016 Consistency of Silhouettes and Their Duals
abstract
Silhouettes provide rich information on three-dimensional shape, since the intersection of the associated visual cones generates the "visual hull", which encloses and approximates the original shape. However, not all silhouettes can actually be projections of the same object in space: this simple observation has implications in object recognition and multi-view segmentation, and has been (often implicitly) used as a basis for camera calibration. In this paper, we investigate the conditions for multiple silhouettes, or more generally arbitrary closed image sets, to be geometrically "consistent". We present this notion as a natural generalization of traditional multi-view geometry, which deals with consistency for points. After discussing some general results, we present a "dual" formulation for consistency, that gives conditions for a family of planar sets to be sections of the same object. Finally, we introduce a more general notion of silhouette "compatibility" under partial knowledge of the camera projections, and point out some possible directions for future research.
Matthew Trager, Martial Hebert, Jean Ponce
CVPR3
2016 Learning Dictionary of Discriminative Part Detectors for Image Categorization and Cosegmentation
Jian Sun 0009, Jean Ponce
Int. J. Comput. Vis.2
2016 Trinocular Geometry Revisited
Matthew Trager, Jean Ponce, Martial Hebert
Int. J. Comput. Vis.2
2015 Unsupervised object discovery and localization in the wild: Part-based matching with bottom-up region proposals
abstract
This paper addresses unsupervised discovery and localization of dominant objects from a noisy image collection with multiple object classes. The setting of this problem is fully unsupervised, without even image-level annotations or any assumption of a single dominant class. This is far more general than typical colocalization, cosegmentation, or weakly-supervised localization tasks. We tackle the discovery and localization problem using a part-based region matching approach: We use off-the-shelf region proposals to form a set of candidate bounding boxes for objects and object parts. These regions are efficiently matched across images using a probabilistic Hough transform that evaluates the confidence for each candidate correspondence considering both appearance and spatial consistency. Dominant objects are discovered and localized by comparing the scores of candidate regions and selecting those that stand out over other regions containing them. Extensive experimental evaluations on standard benchmarks demonstrate that the proposed approach significantly outperforms the current state of the art in colocalization, and achieves robust object discovery in challenging mixed-class datasets.
Minsu Cho, Suha Kwak, Cordelia Schmid, Jean Ponce
CVPR4
2015 Robust image filtering using joint static and dynamic guidance
abstract
Regularizing images under a guidance signal has been used in various tasks in computer vision and computational photography, particularly for noise reduction and joint upsampling. The aim is to transfer fine structures of guidance signals to input images, restoring noisy or altered structures. One of main drawbacks in such a data-dependent framework is that it does not handle differences in structure between guidance and input images. We address this problem by jointly leveraging structural information of guidance and input images. Image filtering is formulated as a nonconvex optimization problem, which is solved by the majorization-minimization algorithm. The proposed algorithm converges quickly while guaranteeing a local minimum. It effectively controls image structures at different scales and can handle a variety of types of data from different sensors. We demonstrate the flexibility and effectiveness of our model in several applications including depth super-resolution, scale-space filtering, texture removal, flash/non-flash denoising, and RGB/NIR denoising.
Bumsub Ham, Minsu Cho, Jean Ponce
CVPR3
2015 Learning a convolutional neural network for non-uniform motion blur removal
abstract
In this paper, we address the problem of estimating and removing non-uniform motion blur from a single blurry image. We propose a deep learning approach to predicting the probabilistic distribution of motion blur at the patch level using a convolutional neural network (CNN). We further extend the candidate set of motion kernels predicted by the CNN using carefully designed image rotations. A Markov random field model is then used to infer a dense non-uniform motion blur field enforcing motion smoothness. Finally, motion blur is removed by a non-uniform deblurring model using patch-level image prior. Experimental evaluations show that our approach can effectively estimate and remove complex non-uniform motion blur that is not handled well by previous approaches.
Jian Sun 0009, Wenfei Cao, Zongben Xu, Jean Ponce
CVPR4
2015 Weakly-Supervised Alignment of Video with Text
abstract
Suppose that we are given a set of videos, along with natural language descriptions in the form of multiple sentences (e.g., manual annotations, movie scripts, sport summaries etc.), and that these sentences appear in the same temporal order as their visual counterparts. We propose in this paper a method for aligning the two modalities, i.e., automatically providing a time (frame) stamp for every sentence. Given vectorial features for both video and text, this can be cast as a temporal assignment problem, with an implicit linear mapping between the two feature modalities. We formulate this problem as an integer quadratic program, and solve its continuous convex relaxation using an efficient conditional gradient algorithm. Several rounding procedures are proposed to construct the final integer solution. After demonstrating significant improvements over the state of the art on the related task of aligning video with symbolic labels [7], we evaluate our method on a challenging dataset of videos with associated textual descriptions [37], and explore bag-of-words and continuous representations for text.
Piotr Bojanowski, Rémi Lajugie, Edouard Grave, Francis R. Bach, Ivan Laptev, Jean Ponce, Cordelia Schmid
ICCV6
2015 Unsupervised Object Discovery and Tracking in Video Collections
abstract
This paper addresses the problem of automatically localizing dominant objects as spatio-temporal tubes in a noisy collection of videos with minimal or even no supervision. We formulate the problem as a combination of two complementary processes: discovery and tracking. The first one establishes correspondences between prominent regions across videos, and the second one associates similar object regions within the same video. Interestingly, our algorithm also discovers the implicit topology of frames associated with instances of the same object class across different videos, a role normally left to supervisory information in the form of class labels in conventional image and video understanding methods. Indeed, as demonstrated by our experiments, our method can handle video collections featuring multiple object classes, and substantially outperforms the state of the art in colocalization, even though it tackles a broader problem with much less supervision.
Suha Kwak, Minsu Cho, Ivan Laptev, Jean Ponce, Cordelia Schmid
ICCV4
2015 The Joint Image Handbook
abstract
Given multiple perspective photographs, point correspondences form the "joint image", effectively a replica of three dimensional space distributed across its two-dimensional projections. This set can be characterized by multilinear equations over image coordinates, such as epipolar and trifocal constraints. We revisit in this paper the geometric and algebraic properties of the joint image, and address fundamental questions such as how many and which multilinearities are necessary and/or sufficient to determine camera geometry and/or image correspondences. The new theoretical results in this paper answer these questions in a very general setting and, in turn, are intended to serve as a "handbook" reference about multilinearities for practitioners.
Matthew Trager, Martial Hebert, Jean Ponce
ICCV3
2014 Finding Matches in a Haystack: A Max-Pooling Strategy for Graph Matching in the Presence of Outliers
abstract
A major challenge in real-world feature matching problems is to tolerate the numerous outliers arising in typical visual tasks. Variations in object appearance, shape, and structure within the same object class make it harder to distinguish inliers from outliers due to clutters. In this paper, we propose a max-pooling approach to graph matching, which is not only resilient to deformations but also remarkably tolerant to outliers. The proposed algorithm evaluates each candidate match using its most promising neighbors, and gradually propagates the corresponding scores to update the neighbors. As final output, it assigns a reliable score to each match together with its supporting neighbors, thus providing contextual information for further verification. We demonstrate the robustness and utility of our method with synthetic and real image experiments.
Minsu Cho, Jian Sun 0009, Olivier Duchenne, Jean Ponce
CVPR4
2014 Trinocular Geometry Revisited
abstract
When do the visual rays associated with triplets of point correspondences converge, that is, intersect in a common point? Classical models of trinocular geometry based on the fundamental matrices and trifocal tensor associated with the corresponding cameras only provide partial answers to this fundamental question, in large part because of underlying, but seldom explicit, general configuration assumptions. This paper uses elementary tools from projective line geometry to provide necessary and sufficient geometric and analytical conditions for convergence in terms of transversals to triplets of visual rays, without any such assumptions. In turn, this yields a novel and simple minimal parameterization of trinocular geometry for cameras with non-collinear or collinear pinholes.
Jean Ponce, Martial Hebert
CVPR1
2014 Weakly Supervised Action Labeling in Videos under Ordering Constraints
Piotr Bojanowski, Rémi Lajugie, Francis R. Bach, Ivan Laptev, Jean Ponce, Cordelia Schmid, Josef Sivic
ECCV (5)5
2014 On Image Contours of Projective Shapes
Jean Ponce, Martial Hebert
ECCV (4)1
2013 Learning to Estimate and Remove Non-uniform Image Blur
abstract
This paper addresses the problem of restoring images subjected to unknown and spatially varying blur caused by defocus or linear (say, horizontal) motion. The estimation of the global (non-uniform) image blur is cast as a multi-label energy minimization problem. The energy is the sum of unary terms corresponding to learned local blur estimators, and binary ones corresponding to blur smoothness. Its global minimum is found using Ishikawa's method by exploiting the natural order of discretized blur values for linear motions and defocus. Once the blur has been estimated, the image is restored using a robust (non-uniform) deblurring algorithm based on sparse regularization with global image statistics. The proposed algorithm outputs both a segmentation of the image into uniform-blur layers and an estimate of the corresponding sharp image. We present qualitative results on real images, and use synthetic data to quantitatively compare our approach to the publicly available implementation of Chakrabarti~et al.
Florent Couzinie-Devy, Jian Sun 0009, Karteek Alahari, Jean Ponce
CVPR4
2013 Finding Actors and Actions in Movies
abstract
We address the problem of learning a joint model of actors and actions in movies using weak supervision provided by scripts. Specifically, we extract actor/action pairs from the script and use them as constraints in a discriminative clustering framework. The corresponding optimization problem is formulated as a quadratic program under linear constraints. People in video are represented by automatically extracted and tracked faces together with corresponding motion features. First, we apply the proposed framework to the task of learning names of characters in the movie and demonstrate significant improvements over previous methods used for this task. Second, we explore the joint actor/action constraint and show its advantage for weakly supervised action learning. We validate our method in the challenging setting of localizing and recognizing characters and their actions in feature length movies Casablanca and American Beauty.
Piotr Bojanowski, Francis R. Bach, Ivan Laptev, Jean Ponce, Cordelia Schmid, Josef Sivic
ICCV4
2013 Learning Graphs to Match
abstract
Many tasks in computer vision are formulated as graph matching problems. Despite the NP-hard nature of the problem, fast and accurate approximations have led to significant progress in a wide range of applications. Learning graph models from observed data, however, still remains a challenging issue. This paper presents an effective scheme to parameterize a graph model, and learn its structural attributes for visual object matching. For this, we propose a graph representation with histogram-based attributes, and optimize them to increase the matching accuracy. Experimental evaluations on synthetic and real image datasets demonstrate the effectiveness of our approach, and show significant improvement in matching accuracy over graphs with pre-defined structures.
Minsu Cho, Karteek Alahari, Jean Ponce
ICCV3
2013 Learning Discriminative Part Detectors for Image Classification and Cosegmentation
abstract
In this paper, we address the problem of learning discriminative part detectors from image sets with category labels. We propose a novel latent SVM model regularized by group sparsity to learn these part detectors. Starting from a large set of initial parts, the group sparsity regularizer forces the model to jointly select and optimize a set of discriminative part detectors in a max-margin framework. We propose a stochastic version of a proximal algorithm to solve the corresponding optimization problem. We apply the proposed method to image classification and co segmentation, and quantitative experiments with standard benchmarks show that it matches or improves upon the state of the art.
Jian Sun 0009, Jean Ponce
ICCV2
2012 Multi-class cosegmentation
abstract
Bottom-up, fully unsupervised segmentation remains a daunting challenge for computer vision. In the cosegmentation context, on the other hand, the availability of multiple images assumed to contain instances of the same object classes provides a weak form of supervision that can be exploited by discriminative approaches. Unfortunately, most existing algorithms are limited to a very small number of images and/or object classes (typically two of each). This paper proposes a novel energy-minimization approach to cosegmentation that can handle multiple classes and a significantly larger number of images. The proposed cost function combines spectral- and discriminative-clustering terms, and it admits a probabilistic interpretation. It is optimized using an efficient EM method, initialized using a convex quadratic approximation of the energy. Comparative experiments show that the proposed approach matches or improves the state of the art on several standard datasets.
Armand Joulin, Francis R. Bach, Jean Ponce
CVPR3
2012 Non-uniform Deblurring for Shaken Images
Oliver Whyte, Josef Sivic, Andrew Zisserman, Jean Ponce
Int. J. Comput. Vis.4
2012 Task-Driven Dictionary Learning
abstract
Modeling data with linear combinations of a few elements from a learned dictionary has been the focus of much recent research in machine learning, neuroscience, and signal processing. For signals such as natural images that admit such sparse representations, it is now well established that these models are well suited to restoration tasks. In this context, learning the dictionary amounts to solving a large-scale matrix factorization problem, which can be done efficiently with classical optimization tools. The same approach has also been used for learning features from data for other purposes, e.g., image classification, but tuning the dictionary in a supervised way for these tasks has proven to be more difficult. In this paper, we present a general formulation for supervised dictionary learning adapted to a wide variety of tasks, and present an efficient algorithm for solving the corresponding optimization problem. Experiments on handwritten digit classification, digital art identification, nonlinear inverse image problems, and compressed sensing demonstrate that our approach is effective in large-scale settings, and is well suited to supervised and semi-supervised classification, as well as regression tasks for data that admit sparse representations.
Julien Mairal, Francis R. Bach, Jean Ponce
IEEE Trans. Pattern Anal. Mach. Intell.3
2011 Sparse image representation with epitomes
abstract
Sparse coding, which is the decomposition of a vector using only a few basis elements, is widely used in machine learning and image processing. The basis set, also called dictionary, is learned to adapt to specific data. This approach has proven to be very effective in many image processing tasks. Traditionally, the dictionary is an unstructured “flat” set of atoms. In this paper, we study structured dictionaries which are obtained from an epitome, or a set of epitomes. The epitome is itself a small image, and the atoms are all the patches of a chosen size inside this image. This considerably reduces the number of parameters to learn and provides sparse image decompositions with shift-invariance properties. We propose a new formulation and an algorithm for learning the structured dictionaries associated with epitomes, and illustrate their use in image de-noising tasks.
Louise Benoît, Julien Mairal, Francis R. Bach, Jean Ponce
CVPR4
2011 Ask the locals: Multi-way local pooling for image recognition
abstract
Invariant representations in object recognition systems are generally obtained by pooling feature vectors over spatially local neighborhoods. But pooling is not local in the feature vector space, so that widely dissimilar features may be pooled together if they are in nearby locations. Recent approaches rely on sophisticated encoding methods and more specialized codebooks (or dictionaries), e.g., learned on subsets of descriptors which are close in feature space, to circumvent this problem. In this work, we argue that a common trait found in much recent work in image recognition or retrieval is that it leverages locality in feature space on top of purely spatial locality. We propose to apply this idea in its simplest form to an object recognition system based on the spatial pyramid framework, to increase the performance of small dictionaries with very little added engineering. State-of-the-art results on several object recognition benchmarks show the promise of this approach.
Y-Lan Boureau, Nicolas Le Roux, Francis R. Bach, Jean Ponce, Yann LeCun
ICCV4
2011 A graph-matching kernel for object categorization
abstract
This paper addresses the problem of category-level image classification. The underlying image model is a graph whose nodes correspond to a dense set of regions, and edges reflect the underlying grid structure of the image and act as springs to guarantee the geometric consistency of nearby regions during matching. A fast approximate algorithm for matching the graphs associated with two images is presented. This algorithm is used to construct a kernel appropriate for SVM-based image classification, and experiments with the Caltech 101, Caltech 256, and Scenes datasets demonstrate performance that matches or exceeds the state of the art for methods using a single type of features.
Olivier Duchenne, Armand Joulin, Jean Ponce
ICCV3
2011 A Tensor-Based Algorithm for High-Order Graph Matching
abstract
This paper addresses the problem of establishing correspondences between two sets of visual features using higher order constraints instead of the unary or pairwise ones used in classical methods. Concretely, the corresponding hypergraph matching problem is formulated as the maximization of a multilinear objective function over all permutations of the features. This function is defined by a tensor representing the affinity between feature tuples. It is maximized using a generalization of spectral techniques where a relaxed problem is first solved by a multidimensional power method and the solution is then projected onto the closest assignment matrix. The proposed approach has been implemented, and it is compared to state-of-the-art algorithms on both synthetic and real data.
Olivier Duchenne, Francis R. Bach, In-So Kweon, Jean Ponce
IEEE Trans. Pattern Anal. Mach. Intell.4
2010 Sparse coding and dictionary learning for image understanding
Jean Ponce
BMVC1
2010 Admissible linear map models of linear cameras
abstract
This paper presents a complete analytical characterization of a large class of central and non-central imaging devices dubbed linear cameras by Ponce. Pajdla has shown that a subset of these, the oblique cameras, can be modelled by a certain type of linear map. We give here a full tabulation of all admissible maps that induce cameras in the general sense of Grossberg and Nayar, and show that these cameras are exactly the linear ones. Combining these two models with a new notion of intrinsic parameters and normalized coordinates for linear cameras allows us to give simple analytical formulas for direct and inverse projections. We also show that the epipolar geometry of any two linear cameras can be characterized by a fundamental matrix whose size is at most 6 × 6 when the cameras are uncalibrated, or by an essential matrix of size at most 4 × 4 when their internal parameters are known. Similar results hold for trinocular constraints.
Guillaume Batog, Xavier Goaoc, Jean Ponce
CVPR3
2010 Learning mid-level features for recognition
abstract
Many successful models for scene or object recognition transform low-level descriptors (such as Gabor filter responses, or SIFT descriptors) into richer representations of intermediate complexity. This process can often be broken down into two steps: (1) a coding step, which performs a pointwise transformation of the descriptors into a representation better adapted to the task, and (2) a pooling step, which summarizes the coded features over larger neighborhoods. Several combinations of coding and pooling schemes have been proposed in the literature. The goal of this paper is threefold. We seek to establish the relative importance of each step of mid-level feature extraction through a comprehensive cross evaluation of several types of coding modules (hard and soft vector quantization, sparse coding) and pooling schemes (by taking the average, or the maximum), which obtains state-of-the-art performance or better on several recognition benchmarks. We show how to improve the best performing coding scheme by learning a supervised discriminative dictionary for sparse coding. We provide theoretical and empirical insight into the remarkable performance of max pooling. By teasing apart components shared by modern mid-level feature extractors, our approach aims to facilitate the design of better recognition architectures.
Y-Lan Boureau, Francis R. Bach, Yann LeCun, Jean Ponce
CVPR4
2010 Discriminative clustering for image co-segmentation
abstract
Purely bottom-up, unsupervised segmentation of a single image into foreground and background regions remains a challenging task for computer vision. Co-segmentation is the problem of simultaneously dividing multiple images into regions (segments) corresponding to different object classes. In this paper, we combine existing tools for bottom-up image segmentation such as normalized cuts, with kernel methods commonly used in object recognition. These two sets of techniques are used within a discriminative clustering framework: the goal is to assign foreground/background labels jointly to all images, so that a supervised classifier trained with these labels leads to maximal separation of the two classes. In practice, we obtain a combinatorial optimization problem which is relaxed to a continuous convex optimization problem, that can itself be solved efficiently for up to dozens of images. We illustrate the proposed method on images with very similar foreground objects, as well as on more challenging problems with objects with higher intra-class variations.
Armand Joulin, Francis R. Bach, Jean Ponce
CVPR3
2010 Non-uniform deblurring for shaken images
abstract
Blur from camera shake is mostly due to the 3D rotation of the camera, resulting in a blur kernel that can be significantly non-uniform across the image. However, most current deblurring methods model the observed image as a convolution of a sharp image with a uniform blur kernel. We propose a new parametrized geometric model of the blurring process in terms of the rotational velocity of the camera during exposure. We apply this model to two different algorithms for camera shake removal: the first one uses a single blurry image (blind deblurring), while the second one uses both a blurry image and a sharp but noisy image of the same scene. We show that our approach makes it possible to model and remove a wider class of blurs than previous approaches, including uniform blur as a special case, and demonstrate its effectiveness with experiments on real images.
Oliver Whyte, Josef Sivic, Andrew Zisserman, Jean Ponce
CVPR4
2010 A Theoretical Analysis of Feature Pooling in Visual Recognition
Y-Lan Boureau, Jean Ponce, Yann LeCun
ICML2
2010 Efficient Optimization for Discriminative Latent Class Models
abstract
Dimensionality reduction is commonly used in the setting of multi-label supervised classification to control the learning capacity and to provide a meaningful representation of the data. We introduce a simple forward probabilistic model which is a multinomial extension of reduced rank regression; we show that this model provides a probabilistic interpretation of discriminative clustering methods with added benefits in terms of number of hyperparameters and optimization. While expectation-maximization (EM) algorithm is commonly used to learn these models, its optimization usually leads to local minimum because it relies on a non-convex cost function with many such local minima. To avoid this problem, we introduce a local approximation of this cost function, which leads to a quadratic non-convex optimization problem over a product of simplices. In order to minimize such functions, we propose an efficient algorithm based on convex relaxation and low-rank representation of our data, which allows to deal with large instances. Experiments on text document classification show that the new model outperforms other supervised dimensionality reduction methods, while simulations on unsupervised clustering show that our probabilistic formulation has better properties than existing discriminative clustering methods.
Armand Joulin, Francis R. Bach, Jean Ponce
NIPS3
2010 Online Learning for Matrix Factorization and Sparse Coding
Julien Mairal, Francis R. Bach, Jean Ponce, Guillermo Sapiro
J. Mach. Learn. Res.3
2010 Accurate, Dense, and Robust Multiview Stereopsis
abstract
This paper proposes a novel algorithm for multiview stereopsis that outputs a dense set of small rectangular patches covering the surfaces visible in the images. Stereopsis is implemented as a match, expand, and filter procedure, starting from a sparse set of matched keypoints, and repeatedly expanding these before using visibility constraints to filter away false matches. The keys to the performance of the proposed algorithm are effective techniques for enforcing local photometric consistency and global visibility constraints. Simple but effective methods are also proposed to turn the resulting patch model into a mesh which can be further refined by an algorithm that enforces both photometric consistency and regularization constraints. The proposed approach automatically detects and discards outliers and obstacles and does not require any initialization in the form of a visual hull, a bounding box, or valid depth ranges. We have tested our algorithm on various data sets including objects with fine surface details, deep concavities, and thin structures, outdoor scenes observed from a restricted set of viewpoints, and "crowded" scenes where moving obstacles appear in front of a static structure of interest. A quantitative evaluation on the Middlebury benchmark shows that the proposed method outperforms all others submitted so far for four out of the six data sets.
Yasutaka Furukawa, Jean Ponce
IEEE Trans. Pattern Anal. Mach. Intell.2
2010 Detecting Abandoned Objects With a Moving Camera
abstract
This paper presents a novel framework for detecting nonflat abandoned objects by matching a reference and a target video sequences. The reference video is taken by a moving camera when there is no suspicious object in the scene. The target video is taken by a camera following the same route and may contain extra objects. The objective is to find these objects. GPS information is used to roughly align the two videos and find the corresponding frame pairs. Based upon the GPS alignment, four simple but effective ideas are proposed to achieve the objective: an intersequence geometric alignment based upon homographies, which is computed by a modified RANSAC, to find all possible suspicious areas, an intrasequence geometric alignment to remove false alarms caused by high objects, a local appearance comparison between two aligned intrasequence frames to remove false alarms in flat areas, and a temporal filtering step to confirm the existence of suspicious objects. Experiments on fifteen pairs of videos show the promise of the proposed method.
Hui Kong 0001, Jean-Yves Audibert, Jean Ponce
IEEE Trans. Image Process.3
2010 General Road Detection From a Single Image
abstract
Given a single image of an arbitrary road, that may not be well-paved, or have clearly delineated edges, or some a priori known color or texture distribution, is it possible for a computer to find this road? This paper addresses this question by decomposing the road detection process into two steps: the estimation of the vanishing point associated with the main (straight) part of the road, followed by the segmentation of the corresponding road area based upon the detected vanishing point. The main technical contributions of the proposed approach are a novel adaptive soft voting scheme based upon a local voting region using high-confidence voters, whose texture orientations are computed using Gabor filters, and a new vanishing-point-constrained edge detection technique for detecting road boundaries. The proposed method has been implemented, and experiments with 1003 general road images demonstrate that it is effective at detecting road regions in challenging conditions.
Hui Kong 0001, Jean-Yves Audibert, Jean Ponce
IEEE Trans. Image Process.3
2009 A tensor-based algorithm for high-order graph matching
abstract
This paper addresses the problem of establishing correspondences between two sets of visual features using higher-order constraints instead of the unary or pairwise ones used in classical methods. Concretely, the corresponding hypergraph matching problem is formulated as the maximization of a multilinear objective function over all permutations of the features. This function is defined by a tensor representing the affinity between feature tuples. It is maximized using a generalization of spectral techniques where a relaxed problem is first solved by a multi-dimensional power method, and the solution is then projected onto the closest assignment matrix. The proposed approach has been implemented, and it is compared to state-of-the-art algorithms on both synthetic and real data.
Olivier Duchenne, Francis R. Bach, In-So Kweon, Jean Ponce
CVPR4
2009 Dense 3D motion capture for human faces
abstract
This paper proposes a novel approach to motion capture from multiple, synchronized video streams, specifically aimed at recording dense and accurate models of the structure and motion of highly deformable surfaces such as skin, that stretches, shrinks, and shears in the midst of normal facial expressions. Solving this problem is a key step toward effective performance capture for the entertainment industry, but progress so far has been hampered by the lack of appropriate local motion and smoothness models. The main technical contribution of this paper is a novel approach to regularization adapted to nonrigid tangential deformations. Concretely, we estimate the nonrigid deformation parameters at each vertex of a surface mesh, smooth them over a local neighborhood for robustness, and use them to regularize the tangential motion estimation. To demonstrate the power of the proposed approach, we have integrated it into our previous work for markerless motion capture [9], and compared the performances of the original and new algorithms on three extremely challenging face datasets that include highly nonrigid skin deformations, wrinkles, and quickly changing expressions. Additional experiments with a dataset featuring fast-moving cloth with complex and evolving fold structures demonstrate that the adaptability of the proposed regularization scheme to nonrigid tangential motion does not hamper its robustness, since it successfully recovers the shape and motion of the cloth without overfitting it despite the absence of stretch or shear in this case.
Yasutaka Furukawa, Jean Ponce
CVPR2
2009 Vanishing point detection for road detection
abstract
Given a single image of an arbitrary road, that may not be well-paved, or have clearly delineated edges, or some a priori known color or texture distribution, is it possible for a computer to find this road? This paper addresses this question by decomposing the road detection process into two steps: the estimation of the vanishing point associated with the main (straight) part of the road, followed by the segmentation of the corresponding road area based on the detected vanishing point. The main technical contributions of the proposed approach are a novel adaptive soft voting scheme based on variable-sized voting region using confidence-weighted Gabor filters, which compute the dominant texture orientation at each pixel, and a new vanishing-point-constrained edge detection technique for detecting road boundaries. The proposed method has been implemented, and experiments with 1003 general road images demonstrate that it is both computationally efficient and effective at detecting road regions in challenging conditions.
Hui Kong 0001, Jean-Yves Audibert, Jean Ponce
CVPR3
2009 What is a camera?
abstract
This paper addresses the problem of characterizing a general class of cameras under reasonable, “linear” assumptions. Concretely, we use the formalism and terminology of classical projective geometry to model cameras by two-parameter linear families of straight lines-that is, degenerate reguli (rank-3 families) and non-degenerate linear congruences (rank-4 families). This model captures both the general linear cameras of Yu and McMillan and the linear oblique cameras of Pajdla. From a geometric perspective, it affords a simple classification of all possible camera configurations. From an analytical viewpoint, it also provides a simple and unified methodology for deriving general formulas for projection and inverse projection, triangulation, and binocular and trinocular geometry.
Jean Ponce
CVPR1
2009 Automatic annotation of human actions in video
abstract
This paper addresses the problem of automatic temporal annotation of realistic human actions in video using minimal manual supervision. To this end we consider two associated problems: (a) weakly-supervised learning of action models from readily available annotations, and (b) temporal localization of human actions in test videos. To avoid the prohibitive cost of manual annotation for training, we use movie scripts as a means of weak supervision. Scripts, however, provide only implicit, noisy, and imprecise information about the type and location of actions in video. We address this problem with a kernel-based discriminative clustering algorithm that locates actions in the weakly-labeled training data. Using the obtained action samples, we train temporal action detectors and apply them to locate actions in the raw video data. Our experiments demonstrate that the proposed method for weakly-supervised learning of action models leads to significant improvement in action detection. We present detection results for three action classes in four feature length movies with challenging and realistic video data.
Olivier Duchenne, Ivan Laptev, Josef Sivic, Francis R. Bach, Jean Ponce
ICCV5
2009 Non-local sparse models for image restoration
abstract
We propose in this paper to unify two different approaches to image restoration: On the one hand, learning a basis set (dictionary) adapted to sparse signal descriptions has proven to be very effective in image reconstruction and classification tasks. On the other hand, explicitly exploiting the self-similarities of natural images has led to the successful non-local means approach to image restoration. We propose simultaneous sparse coding as a framework for combining these two approaches in a natural manner. This is achieved by jointly decomposing groups of similar signals on subsets of the learned dictionary. Experimental results in image denoising and demosaicking tasks with synthetic and real noise show that the proposed method outperforms the state of the art, making it possible to effectively restore raw images from digital cameras at a reasonable speed and memory cost.
Julien Mairal, Francis R. Bach, Jean Ponce, Guillermo Sapiro, Andrew Zisserman
ICCV3
2009 Online dictionary learning for sparse coding
abstract
Sparse coding---that is, modelling data vectors as sparse linear combinations of basis elements---is widely used in machine learning, neuroscience, signal processing, and statistics. This paper focuses on learning the basis set, also called dictionary, to adapt it to specific data, an approach that has recently proven to be very effective for signal reconstruction and classification in the audio and image processing domains. This paper proposes a new online optimization algorithm for dictionary learning, based on stochastic approximations, which scales up gracefully to large datasets with millions of training samples. A proof of convergence is presented, along with experiments with natural images demonstrating that it leads to faster performance and better dictionaries than classical batch algorithms for both small and large datasets.
Julien Mairal, Francis R. Bach, Jean Ponce, Guillermo Sapiro
ICML3
2009 Carved Visual Hulls for Image-Based Modeling
Yasutaka Furukawa, Jean Ponce
Int. J. Comput. Vis.2
2009 Accurate Camera Calibration from Multi-View Stereo and Bundle Adjustment
Yasutaka Furukawa, Jean Ponce
Int. J. Comput. Vis.2
2008 Segmentation by transduction
abstract
This paper addresses the problem of segmenting an image into regions consistent with user-supplied seeds (e.g., a sparse set of broad brush strokes). We view this task as a statistical transductive inference, in which some pixels are already associated with given zones and the remaining ones need to be classified. Our method relies on the Laplacian graph regularizer, a powerful manifold learning tool that is based on the estimation of variants of the Laplace-Beltrami operator and is tightly related to diffusion processes. Segmentation is modeled as the task of finding matting coefficients for unclassified pixels given known matting coefficients for seed pixels. The proposed algorithm essentially relies on a high margin assumption in the space of pixel characteristics. It is simple, fast, and accurate, as demonstrated by qualitative results on natural images and a quantitative comparison with state-of-the-art methods on the Microsoft GrabCut segmentation database.
Olivier Duchenne, Jean-Yves Audibert, Renaud Keriven, Jean Ponce, Florent Ségonne
CVPR4
2008 Dense 3D motion capture from synchronized video streams
abstract
This paper proposes a novel approach to non-rigid, markerless motion capture from synchronized video streams acquired by calibrated cameras. The instantaneous geometry of the observed scene is represented by a polyhedral mesh with fixed topology. The initial mesh is constructed in the first frame using the publicly available PMVS software for multi-view stereo [7]. Its deformation is captured by tracking its vertices over time, using two optimization processes at each frame: a local one using a rigid motion model in the neighborhood of each vertex, and a global one using a regularized nonrigid model for the whole mesh. Qualitative and quantitative experiments using seven real datasets show that our algorithm effectively handles complex nonrigid motions and severe occlusions.
Yasutaka Furukawa, Jean Ponce
CVPR2
2008 Accurate camera calibration from multi-view stereo and bundle adjustment
abstract
The advent of high-resolution digital cameras and sophisticated multi-view stereo algorithms offers the promises of unprecedented geometric fidelity in image-based modeling tasks, but it also puts unprecedented demands on camera calibration to fulfill these promises. This paper presents a novel approach to camera calibration where top-down information from rough camera parameter estimates and the output of a publicly available multiview-stereo system [6] on scaled-down input images are used to effectively guide the search for additional image correspondences and significantly improve camera calibration parameters using a standard bundle adjustment algorithm [14]. The proposed method has been tested on several real datasets—including objects without salient features for which image correspondences cannot be found in a purely bottom-up fashion, and image-based modeling tasks-including the construction of visual hulls where thin structures are lost without our calibration procedure.
Yasutaka Furukawa, Jean Ponce
CVPR2
2008 Discriminative learned dictionaries for local image analysis
abstract
Sparse signal models have been the focus of much recent research, leading to (or improving upon) state-of-the-art results in signal, image, and video restoration. This article extends this line of research into a novel framework for local image discrimination tasks, proposing an energy formulation with both sparse reconstruction and class discrimination components, jointly optimized during dictionary learning. This approach improves over the state of the art in texture segmentation experiments using the Brodatz database, and it paves the way for a novel scene analysis and recognition framework based on simultaneously learning discriminative and reconstructive dictionaries. Preliminary results in this direction using examples from the Pascal VOC06 and Graz02 datasets are presented as well.
Julien Mairal, Francis R. Bach, Jean Ponce, Guillermo Sapiro, Andrew Zisserman
CVPR3
2008 Discriminative Sparse Image Models for Class-Specific Edge Detection and Image Interpretation
Julien Mairal, Marius Leordeanu, Francis R. Bach, Martial Hebert, Jean Ponce
ECCV (3)5
2008 Supervised Dictionary Learning
abstract
It is now well established that sparse signal models are well suited to restoration tasks and can effectively be learned from audio, image, and video data. Recent research has been aimed at learning discriminative sparse models instead of purely reconstructive ones. This paper proposes a new step in that direction with a novel sparse representation for signals belonging to different classes in terms of a shared dictionary and multiple decision functions. It is shown that the linear variant of the model admits a simple probabilistic interpretation, and that its most general variant also admits a simple interpretation in terms of kernels. An optimization framework for learning all the components of the proposed model is presented, along with experiments on standard handwritten digit and texture classification tasks.
Julien Mairal, Francis R. Bach, Jean Ponce, Guillermo Sapiro, Andrew Zisserman
NIPS3
2007 Accurate, Dense, and Robust Multi-View Stereopsis
abstract
This paper proposes a novel algorithm for calibrated multi-view stereopsis that outputs a (quasi) dense set of rectangular patches covering the surfaces visible in the input images. This algorithm does not require any initialization in the form of a bounding volume, and it detects and discards automatically outliers and obstacles. It does not perform any smoothing across nearby features, yet is currently the top performer in terms of both coverage and accuracy for four of the six benchmark datasets presented in [20]. The keys to its performance are effective techniques for enforcing local photometric consistency and global visibility constraints. Stereopsis is implemented as a match, expand, and filter procedure, starting from a sparse set of matched keypoints, and repeatedly expanding these to nearby pixel correspondences before using visibility constraints to filter away false matches. A simple but effective method for turning the resulting patch model into a mesh appropriate for image-based modeling is also presented. The proposed approach is demonstrated on various datasets including objects with fine surface details, deep concavities, and thin structures, outdoor scenes observed from a restricted set of viewpoints, and "crowded" scenes where moving obstacles appear in different places in multiple images of a static structure of interest.
Yasutaka Furukawa, Jean Ponce
CVPR2
2007 Flexible Object Models for Category-Level 3D Object Recognition
abstract
Today's category-level object recognition systems largely focus on fronto-parallel views of objects with characteristic texture patterns. To overcome these limitations, we propose a novel framework for visual object recognition where object classes are represented by assemblies of partial surface models (PSMs) obeying loose local geometric constraints. The PSMs themselves are formed of dense, locally rigid assemblies of image features. Since our model only enforces local geometric consistency, both at the level of model parts and at the level of individual features within the parts, it is robust to viewpoint changes and intra-class variability. The proposed approach has been implemented, and it outperforms the state-of-the-art algorithms for object detection and localization recently compared in [14] on the Pascal 2005 VOC Challenge Cars Test 1 data.
Akash Kushal, Cordelia Schmid, Jean Ponce
CVPR3
2007 Projective Visual Hulls
Svetlana Lazebnik, Yasutaka Furukawa, Jean Ponce
Int. J. Comput. Vis.3
2007 Segmenting, Modeling, and Matching Video Clips Containing Multiple Moving Objects
abstract
This paper presents a novel representation for dynamic scenes composed of multiple rigid objects that may undergo different motions and are observed by a moving camera. Multiview constraints associated with groups of affine-covariant scene patches and a normalized description of their appearance are used to segment a scene into its rigid components, construct three-dimensional models of these components, and match instances of models recovered from different image sequences. The proposed approach has been applied to the detection and matching of moving objects in video sequences and to shot matching, i.e., the identification of shots that depict the same scene in a video clip.
Fred Rothganger, Svetlana Lazebnik, Cordelia Schmid, Jean Ponce
IEEE Trans. Pattern Anal. Mach. Intell.4
2007 Capturing a Convex Object With Three Discs
abstract
This paper addresses the problem of capturing an arbitrary convex object P in the plane with three congruent disc-shaped robots. Given two stationary robots in contact with P, we characterize the set of positions of a third robot, the so-called capture region, that prevent P from escaping to infinity via continuous rigid motion. We show that the computation of the capture region reduces to a visibility problem. We present two algorithms for solving this problem, and for computing the capture region when P is a polygon and the robots are points (zero-radius discs). The first algorithm is exact and has polynomial time complexity. The second one uses simple hidden surface removal techniques from computer graphics to output an arbitrarily accurate approximation of the capture region; it has been implemented, and examples are presented.
Jeff Erickson 0001, Shripad Thite, Fred Rothganger, Jean Ponce
IEEE Trans. Robotics4
2006 Beyond Bags of Features: Spatial Pyramid Matching for Recognizing Natural Scene Categories
abstract
This paper presents a method for recognizing scene categories based on approximate global geometric correspondence. This technique works by partitioning the image into increasingly fine sub-regions and computing histograms of local features found inside each sub-region. The resulting "spatial pyramid" is a simple and computationally efficient extension of an orderless bag-of-features image representation, and it shows significantly improved performance on challenging scene categorization tasks. Specifically, our proposed method exceeds the state of the art on the Caltech-101 database and achieves high accuracy on a large database of fifteen natural scene categories. The spatial pyramid framework also offers insights into the success of several recently proposed image descriptions, including Torralba’s "gist" and Lowe’s SIFT descriptors.
Svetlana Lazebnik, Cordelia Schmid, Jean Ponce
CVPR (2)3
2006 A Geodesic Active Contour Framework for Finding Glass
abstract
This paper addresses the problem of finding objects made of glass (or other transparent materials) in images. Since the appearance of glass objects depends for the most part on what lies behind them, we propose to use binary criteria ("are these two regions made of the same material?") rather than unary ones ("is this glass?") to guide the segmentation process. Concretely, we combine two complementary measures of affinity between regions made of the same material and discrepancy between regions made of different ones into a single objective function, and use the geodesic active contour framework to minimize this function over pixel labels. The proposed approach has been implemented, and qualitative and quantitative experimental results are presented.
Kenton McHenry, Jean Ponce
CVPR (1)2
2006 Carved Visual Hulls for Image-Based Modeling
Yasutaka Furukawa, Jean Ponce
ECCV (1)2
2006 Modeling 3D Objects from Stereo Views and Recognizing Them in Photographs
Akash Kushal, Jean Ponce
ECCV (2)2
2006 3D Object Modeling and Recognition Using Local Affine-Invariant Image Descriptors and Multi-View Spatial Constraints
Fred Rothganger, Svetlana Lazebnik, Cordelia Schmid, Jean Ponce
Int. J. Comput. Vis.4
2006 Robust Structure and Motion from Outlines of Smooth Curved Surfaces
abstract
This paper addresses the problem of estimating the motion of a camera as it observes the outline (or apparent contour) of a solid bounded by a smooth surface in successive image frames. In this context, the surface points that project onto the outline of an object depend on the viewpoint and the only true correspondences between two outlines of the same object are the projections of frontier points where the viewing rays intersect in the tangent plane of the surface. In turn, the epipolar geometry is easily estimated once these correspondences have been identified. Given the apparent contours detected in an image sequence, a robust procedure based on RANSAC and a voting strategy is proposed to simultaneously estimate the camera configurations and a consistent set of frontier point projections by enforcing the redundancy of multiview epipolar geometry. The proposed approach is, in principle, applicable to orthographic, weak-perspective, and affine projection models. Experiments with nine real image sequences are presented for the orthographic projection case, including a quantitative comparison with the ground-truth data for the six data sets for which the latter information is available. Sample visual hulls have been computed from all image sequences for qualitative evaluation.
Yasutaka Furukawa, Amit Sethi, Jean Ponce, David J. Kriegman
IEEE Trans. Pattern Anal. Mach. Intell.3
2005 Finding Glass
abstract
This paper addresses the problem of finding glass objects in images. Visual cues obtained by combining the systematic distortions in background texture occurring at the boundaries of transparent objects with the strong highlights typical of glass surfaces are used to train a hierarchy of classifiers, identify glass edges, and find consistent support regions for these edges. Qualitative and quantitative experiments involving a number of different classifiers and real images are presented.
Kenton McHenry, Jean Ponce, David A. Forsyth
CVPR (2)2
2005 On the Absolute Quadratic Complex and Its Application to Autocalibration
abstract
This article introduces the absolute quadratic complex formed by all lines that intersect the absolute conic. If /spl omega/ denotes the 3 /spl times/ 3 symmetric matrix representing the image of that conic under the action of a camera with projection matrix P, it is shown that /spl omega/ /spl ap/ P/sup ~//spl Omega//sub /spl I.bar//P/sup ~T/ where V is the 3 /spl times/ 6 line projection matrix associated with P and /spl Omega//sub /spl I.bar// is a 6 /spl times/ 6 symmetric matrix of rank 3 representing the absolute quadratic complex. This simple relation between a camera's intrinsic parameters, its projection matrix expressed in a projective coordinate frame, and the metric upgrade separating this frame from a metric one - as respectively captured by the matrices /spl omega/, P/sup ~/ and /spl Omega//sub /spl I.bar// - provides a new framework for autocalibration, particularly well suited to typical digital cameras with rectangular or square pixels since the skew and aspect ratio are decoupled from the other intrinsic parameters in /spl omega/.
Jean Ponce, Kenton McHenry, Théodore Papadopoulo, Monique Teillaud, Bill Triggs
CVPR (1)1
2005 A Maximum Entropy Framework for Part-Based Texture and Object Recognition
abstract
This paper presents a probabilistic part-based approach for texture and object recognition. Textures are represented using a part dictionary found by quantizing the appearance of scale- or affine- invariant keypoints. Object classes are represented using a dictionary of composite semi-local parts, or groups of neighboring keypoints with stable and distinctive appearance and geometric layout. A discriminative maximum entropy framework is used to learn the posterior distribution of the class label given the occurrences of parts from the dictionary in the training set. Experiments on two texture and two object databases demonstrate the effectiveness of this framework for visual classification.
Svetlana Lazebnik, Cordelia Schmid, Jean Ponce
ICCV3
2005 The Local Projective Shape of Smooth Surfaces and Their Outlines
Svetlana Lazebnik, Jean Ponce
Int. J. Comput. Vis.2
2005 A Sparse Texture Representation Using Local Affine Regions
abstract
This paper introduces a texture representation suitable for recognizing images of textured surfaces under a wide range of transformations, including viewpoint changes and nonrigid deformations. At the feature extraction stage, a sparse set of affine Harris and Laplacian regions is found in the image. Each of these regions can be thought of as a texture element having a characteristic elliptic shape and a distinctive appearance pattern. This pattern is captured in an affine-invariant fashion via a process of shape normalization followed by the computation of two novel descriptors, the spin image and the RIFT descriptor. When affine invariance is not required, the original elliptical shape servee as an additional discriminative feature for texture recognition. The proposed approach is evaluated in retrieval and classification tasks using the entire Brodatz database and a publicly available collection of 1,000 photographs of textured surfaces taken from different viewpoints.
Svetlana Lazebnik, Cordelia Schmid, Jean Ponce
IEEE Trans. Pattern Anal. Mach. Intell.3
2004 Semi-Local Affine Parts for Object Recognition
abstract
HAL is a multi-disciplinary open access archive for the deposit and dissemination of sci-entific research documents, whether they are pub-lished or not. The documents may come from teaching and research institutions in France or abroad, or from public or private research centers. L’archive ouverte pluridisciplinaire HAL, est destinée au dépôt et a ̀ la diffusion de documents scientifiques de niveau recherche, publiés ou non, émanant des établissements d’enseignement et de recherche français ou étrangers, des laboratoires publics ou privés.
Svetlana Lazebnik, Cordelia Schmid, Jean Ponce
BMVC3
2004 Segmenting, Modeling, and Matching Video Clips Containing Multiple Moving Objects
Fred Rothganger, Svetlana Lazebnik, Cordelia Schmid, Jean Ponce
CVPR (2)4
2004 Structure and Motion from Images of Smooth Textureless Objects
Yasutaka Furukawa, Amit Sethi, Jean Ponce, David J. Kriegman
ECCV (2)3
2004 Guest Editorial: Computer Vision Research at the Beckman Institute for Advanced Science and Technology
Jean Ponce
Int. J. Comput. Vis.1
2004 Editorial
Jean Ponce
Int. J. Comput. Vis.1
2004 Curve and Surface Duals and the Recognition of Curved 3D Objects from their Silhouettes
Amit Sethi, David Renaudie, David J. Kriegman, Jean Ponce
Int. J. Comput. Vis.4
2003 A Sparse Texture Representation Using Affine-Invariant Regions
abstract
This paper introduces a texture representation suitable for recognizing images of textured surfaces under a wide range of transformations, including viewpoint changes and nonrigid deformations. At the feature extraction stage, a sparse set of affine-invariant local patches is extracted from the image. This spatial selection process permits the computation of characteristic scale and neighborhood shape for every texture element. The proposed texture representation is evaluated in retrieval and classification tasks using the entire Brodatz database and a collection of photographs of textured surfaces taken from different viewpoints.
Svetlana Lazebnik, Cordelia Schmid, Jean Ponce
CVPR (2)3
2003 3D Object Modeling and Recognition Using Affine-Invariant Patches and Multi-View Spatial Constraints
abstract
This paper presents a representation for three-dimensional objects in terms of affine-invariant image patches and their spatial relationships. Multi-view constraints associated with groups of patches are combined with a normalized representation of their appearance to guide matching and reconstruction, allowing the acquisition of true three-dimensional affine and Euclidean models from multiple images and their recognition in a single photograph taken from an arbitrary viewpoint. The proposed approach does not require a separate segmentation stage and is applicable to cluttered scenes. Preliminary modeling and recognition results are presented.
Fred Rothganger, Svetlana Lazebnik, Cordelia Schmid, Jean Ponce
CVPR (2)4
2003 The Local Projective Shape of Smooth Surfaces and their Outlines
abstract
We examine projectively invariant local properties of smooth curves and surfaces. Oriented projective differential geometry is proposed as a theoretical framework for establishing such invariants and describing the local shape of surfaces and their outlines. This framework is applied to two problems: a projective proof of Koenderink's famous characterization of convexities, concavities, and inflections of apparent contours; and the determination of the relative orientation of rim tangents at frontier points.
Svetlana Lazebnik, Jean Ponce
ICCV2
2003 Affine-Invariant Local Descriptors and Neighborhood Statistics for Texture Recognition
abstract
We present a framework for texture recognition based on local affine-invariant descriptors and their spatial layout. At modelling time, a generative model of local descriptors is learned from sample images using the EM algorithm. The EM framework allows the incorporation of unsegmented multitexture images into the training set. The second modelling step consists of gathering co-occurrence statistics of neighboring descriptors. At recognition time, initial probabilities computed from the generative model are refined using a relaxation step that incorporates co-occurrence statistics. Performance is evaluated on images of an indoor scene and pictures of wild animals.
Svetlana Lazebnik, Cordelia Schmid, Jean Ponce
ICCV3
2003 Binocular Helmholtz Stereopsis
abstract
Helmholtz stereopsis has been introduced recently as a surface reconstruction technique that does not assume a model of surface reflectance. In the reported formulation, correspondence was established using a rank constraint, necessitating at least three viewpoints and three pairs of images. Here, it is revealed that the fundamental Helmholtz stereopsis constraint defines a nonlinear partial differential equation, which can be solved using only two images. It is shown that, unlike conventional stereo, binocular Helmholtz stereopsis is able to establish correspondence (and thereby recover surface depth) for objects having an arbitrary and unknown BRDF and in textureless regions (i.e., regions of constant or slowly varying BRDF). An implementation and experimental results validate the method for specular surfaces with and without texture.
Todd E. Zickler, Jeffrey Ho, David J. Kriegman, Jean Ponce, Peter N. Belhumeur
ICCV4
2003 Capturing a convex object with three discs
abstract
This paper addresses the problem of capturing an arbitrary convex object P in the plane with three congruent disc-shaped robots. Given two stationary robots in contact with P, we characterize the set of positions of a third robot that prevent P from escaping to infinity and show that the computation of this so-called capture region reduces to the resolution of a visibility problem. We present two algorithms for solving this problem and computing the capture region when P is a polygon and the robots are points (zero-radius discs). The first algorithm is exact and has polynomial-time complexity. The second one uses simple hidden-surface removal techniques from computer graphics to output an arbitrarily accurate approximation of the capture region; it has been implemented and examples are presented.
Jeff Erickson 0001, Shripad Thite, Fred Rothganger, Jean Ponce
ICRA4
2002 On Pencils of Tangent Planes and the Recognition of Smooth 3D Shapes from Silhouettes
Svetlana Lazebnik, Amit Sethi, Cordelia Schmid, David J. Kriegman, Jean Ponce, Martial Hebert
ECCV (3)5
2002 Motion planning for disc-shaped robots pushing a polygonal object in the plane
abstract
This paper addresses the problem of using three disc-shaped robots to manipulate a polygonal object in the plane in the presence of obstacles. The proposed approach is based on the computation of maximal discs (dubbed maximum independent capture discs, or MICaDs) where the robots can move independently while preventing the object from escaping their grasp. It is shown that, in the absence of obstacles, it is always possible to bring a polygonal object from any configuration to any other one with robot motions constrained to lie in a set of overlapping MICaDs. This approach is generalized to the case where obstacles are present by decomposing the corresponding motion planning task into the construction of a collision-free path for a modified form of the object, and the execution of this path by a sequence of simultaneous and independent robot motions within overlapping MICaDs. The proposed algorithm is guaranteed to generate a valid plan, provided a collision-free path exists for the modified form of the object. It has been implemented and experiments with Nomadic Scout mobile robots are presented.
Attawith Sudsang, Fred Rothganger, Jean Ponce
IEEE Trans. Robotics Autom.3
2001 On Computing Exact Visual Hulls of Solids Bounded by Smooth Surfaces
abstract
This paper presents a method for computing the visual hull that is based on two novel representations: the rim mesh, which describes the connectivity of contour generators on the object surface; and the visual hull mesh, which describes the exact structure of the surface of the solid formed by intersecting a finite number of visual cones. We describe the topological features of these meshes and show how they can be identified in the image using epipolar constraints. These constraints are used to derive an image-based practical reconstruction algorithm that works with weakly calibrated cameras. Experiments on synthetic and real data validate the proposed approach.
Svetlana Lazebnik, Edmond Boyer, Jean Ponce
CVPR (1)3
2001 Provably-Convergent Iterative Methods for Projective Structure from Motion
abstract
The estimation of the projective structure of a scene from image correspondences can be formulated as the minimization of the mean-squared distance between predicted and observed image points with respect to the projection matrices, the scene point positions, and their depths. Since these unknowns are not independent, constraints must be chosen to ensure that the optimization process. is well posed. This paper examines three plausible choices, and shows that the first one leads to the Sturm-Triggs projective factorization algorithm, while the other two lead to new provably-convergent approaches. Experiments with synthetic and real data are used to compare the proposed techniques to the Sturm-Triggs algorithm and bundle adjustment.
Shyjan Mahamud, Martial Hebert, Yasuhiro Omori, Jean Ponce
CVPR (1)4
2001 An implemented planner for manipulating a polygonal object in the plane with three disc-shaped mobile robots
abstract
Presents an implementation of a planner that uses three disc-shaped robots to manipulate a polygonal object in the plane in the presence of obstacles. The approach is based on the computation of the maximal discs (maximal independent capture discs or MICaDs) where the robots can move independently while preventing the object from escaping their grasp. It has been shown that, in the absence of obstacles, it is always possible to bring a polygonal object from any configuration to any other one with robot motions constrained to lie in a set of overlapping MICaDs. This approach is generalized to the case where obstacles are present by decomposing the motion planning task into (1) the construction of a collision-free path for a modified form of the object, and (2) the execution of this path by a sequence of simultaneous and independent robot motions within overlapping MICaDs. The approach is guaranteed to work provided a collision free path exists for the modified form of the object. Experiments with Nomadic Scouts and a visual localization system are presented.
Attawith Sudsang, Fred Rothganger, Jean Ponce
IROS3
2001 Image-Based Rendering Using Parameterized Image Varieties
Yakup Genc, Jean Ponce
Int. J. Comput. Vis.2
2001 On Computing Structural Changes in Evolving Surfaces and their Appearance
Sung-Il Pae, Jean Ponce
Int. J. Comput. Vis.2
2000 Duals, Invariants, and the Recognition of Smooth Objects from their Occlucing Contours
David Renaudie, David J. Kriegman, Jean Ponce
ECCV (1)3
2000 A Reconfigurable Parts Feeder with an Array of Pins
abstract
This paper presents a simple parts feeder consisting of a grid of retractable pins on a vertical plate to manipulate polygonal parts. This reconfigurable "Pachinko machine" is intended as a parts feeding device for flexible assembly. A part dropped on this device may come to rest on the actuated pins, or bounce out or fall through. We can control the set of equilibrium part configurations by selecting the set of actuated pins. The objective is to automatically compute sequences of pin actuation that bring the part to a goal configuration without predicting the exact object motion between equilibria. Our approach is based on the construction of the capture region of each part equilibrium. Reorienting a part reduces to building a directed graph whose nodes consist of equilibria and whose area link pairs of nodes such that the first equilibrium lies in the capture region of the second one, and then exploring this graph to find paths from initial to goal states. We have implemented an algorithm to generate the capture regions and these paths, and have conducted experiments on a prototype Pachinko machine.
Sebastien J. Blind, Christopher C. McCullough, Srinivas Akella, Jean Ponce
ICRA4
2000 Constructing Geometric Object Models from Images
abstract
This paper addresses the problem of constructing object models from various types of images. After a brief discussion of current approaches to this problem, we focus on two of its instances: the construction of three-dimensional surface models from object outlines found in a small set of registered photographs; and the synthesis of new images of a scene without any explicit three-dimensional reconstruction (image-based rendering). In both cases, we discuss the state of the art and present some of our recent work as an illustration of what can be achieved today.
Jean Ponce, Yakup Genc, Steve Sullivan
ICRA1
2000 A New Approach to Motion Planning for Disc-Shaped Robots Manipulating a Polygonal Object in the Plane
abstract
This paper addresses the problem of using three disc-shaped robots to manipulate a polygonal object in the plane in the presence of obstacles. The proposed approach is based on the characterization of the maximal discs (maximum independent capture discs, or MICaDS) where the robots can move independently while preventing the object from escaping their grasp. It is shown that, in the absence of obstacles, it is always possible to bring a polygonal object from any configuration to any other one with robot motions constrained to lie in a set of overlapping MICaDS. A strategy for computing these motions is used in conjunction with an exact motion planner to devise an algorithm guaranteed to find a motion plan avoiding collisions with obstacles as long as a collision-free path exists for the object grown by the diameter of the robots plus some arbitrary positive number /spl epsiv/.
Attawith Sudsang, Jean Ponce
ICRA2
2000 Probabilistic 3D Object Recognition
Ilan Shimshoni, Jean Ponce
Int. J. Comput. Vis.2
1999 Parameterized Image Varieties and Estimation with Bilinear Constraints
abstract
This paper addresses the problem of reliably estimating the coefficients of the parameterized image variety (PIV) associated with the set of weak perspective images of a rigid scene, with applications in image-based rendering. Exploiting the fact that the constraints defining the PIV are linear in its coefficients and bilinear in the image data, the estimation procedure is cast in the errors-in-variables framework and solved using the method proposed by Y. Leedan and P. Meer (1998) for this type of problems. The proposed approach has been implemented, and experiments with real data are shown to yield much better prediction power than the original method based on singular value decomposition. Extensions to the more difficult case of paraperspective projection are briefly discussed.
Yakup Genc, Jean Ponce, Yoram Leedan, Peter Meer
CVPR2
1999 Toward a Scale-Space Aspect Graph: Solids of Revolution
abstract
This paper addresses the problem of constructing the scale-space aspect graph of a solid of revolution whose surface is the zero set of a polynomial volumetric density undergoing a Gaussian diffusion process. Equations for the associated visual event surfaces are derived, and polynomial curve tracing techniques are used to delineate these surfaces. An implementation and examples are presented, and limitations as well as extensions of the proposed approach are discussed.
Sung-Il Pae, Jean Ponce
CVPR2
1999 On Manipulating Polygonal Objects with Three 2-DOF Robots in the Plane
abstract
Addresses the problem of grasping and manipulating a polygonal object with three disc-shaped robots in the plane. These robots may be the fingertips of a gripper or mobile platforms. The proposed approach is based on the characterization of the range of possible object motions when two of the effectors are fixed and the third one is allowed to move in the plane with two degrees of freedom. This technique does not assume that contact is maintained during the execution of the grasping/manipulation task, nor does it rely on detailed (and a priori unverifiable) models of friction or contact dynamics, but it allows the construction of manipulation plans guaranteed to succeed under the weaker assumption that jamming does not occur during the task execution. The proposed approach is validated by simulation examples and preliminary experiments with Nomadic Scout robots.
Attawith Sudsang, Jean Ponce, Mark Hyman, David J. Kriegman
ICRA2
1999 Structure and Motion Estimation from Dynamic Silhouettes under Perspective Projection
Tanuja Joshi, Narendra Ahuja, Jean Ponce
Int. J. Comput. Vis.3
1998 Parameterized Image Varieties: A Novel Approach to the Analysis and Synthesis of Image Sequences
abstract
This paper addresses the problem of characterizing the space formed by all images of a rigid set of n points observed by a weak perspective or paraperspective camera. By taking explicitly into account the Euclidean constraints associated with calibrated cameras, we show that this space is a six-dimensional variety embedded in R/sup 2n/, and parameterize it using the image positions of three reference points. This parameterization is constructed via linear least squares from point correspondences established across a sequence of images, and it is used to synthesize new pictures without any explicit three-dimensional model. Degenerate scene and camera configurations are analyzed, and experiments with real image sequences are presented.
Yakup Genc, Jean Ponce
ICCV2
1998 Automatic Model Construction, Pose Estimation, and Object Recognition from Photographs using Triangular Splines
abstract
This paper proposes a method for automatically constructing triangular G/sup 1/ spline models of complex three-dimensional objects from a few registered photographs. These models are used for pose estimation from monocular silhouette data and they form the basis for a simple recognition strategy. The proposed approach is demonstrated by several experiments.
Steve Sullivan, Jean Ponce
ICCV2
1998 On Grasping and Manipulating Polygonal Objects with Disc-Shaped Robots in the Plane
abstract
This paper addresses the problem of grasping and manipulating a polygonal object with three disc-shaped robots capable of translating in arbitrary directions in the plane. The main novelty of the proposed approach is that it does not assume that contact is maintained during the execution of the grasping/manipulation task, nor does it rely on detailed (and a priori unverifiable) models of friction or contact dynamics. Instead, the range of possible object motions for a given position of the robots is characterized in configuration space. This allows the construction of manipulation plans guaranteed to succeed under the weaker assumption that jamming does not occur during the task execution.
Attawith Sudsang, Jean Ponce
ICRA2
1998 Invariant-Based Recognition of Complex Curved 3D Objects from Image Contours
abstract
This paper addresses the problem of recognizing three-dimensional objects bounded by smooth curved surfaces from image contours found in a single photograph. The proposed approach is based on a viewpoint-invariant relationship between object geometry and certain image features under weak perspective projection. The image features themselves are viewpoint-dependent. Concretely, the set of all possible silhouette bitangents, along with the contour points sharing the same tangent direction, is the projection of a one-dimensional set of surface points where each point lies on the occluding contour for a five-parameter family of viewpoints. These image features form a one-parameter family of equivalence classes, and it is shown that each class can be characterized by a set of numerical attributes that remain constant across the corresponding five-dimensional set of viewpoints. This is the basis for describing objects by “invariant” curves embedded in high-dimensional spaces. Modeling is achieved by moving an object in front of a camera and does not require knowing the object-to-camera transformation; nor does it involve implicit or explicit three-dimensional shape reconstruction. At recognition time, attributes computed from a single image are used to index the model database, and both qualitative and quantitative verification procedures eliminate potential false matches. The approach has been implemented and examples are presented.
B. Vijayakumar, David J. Kriegman, Jean Ponce
Comput. Vis. Image Underst.3
1998 Epipolar Geometry and Linear Subspace Methods: A New Approach to Weak Calibration
Jean Ponce, Yakup Genc
Int. J. Comput. Vis.1
1998 Automatic Model Construction and Pose Estimation From Photographs Using Triangular Splines
abstract
This paper addresses the automatic construction of complex spline object models from a few photographs. Our approach combines silhouettes from registered images to construct a G/sup 1/-continuous triangular spline approximation of an object with unknown topology. We apply a similar optimization procedure to estimate the pose of a modeled object from a single image. Experimental examples of model construction and pose estimation are presented for several complex objects.
Steve Sullivan, Jean Ponce
IEEE Trans. Pattern Anal. Mach. Intell.2
1997 In-hand manipulation: geometry and algorithms
abstract
Addresses the problem of manipulating three-dimensional objects with a reconfigurable gripper. A detailed analysis of the problem geometry in configuration space is used to devise a simple and efficient algorithm for manipulation planning. The proposed approach has been implemented and preliminary simulation experiments are discussed.
Attawith Sudsang, Jean Ponce
IROS2
1997 On planning immobilizing grasps for a reconfigurable gripper
abstract
We propose a reconfigurable gripper that consists of two parallel plates whose distance can be adjusted by a computer-controlled actuator. The bottom plate is a bare plane, and the top plate carries a rectangular grid of actuated pins that can translate in discrete increments under computer control. We propose to use this gripper to immobilize objects through frictionless contacts with three of the pins and the bottom plate. We present an efficient grasp planning algorithm, describe the design of the gripper, which is currently under construction, and report preliminary simulation experiments.
Attawith Sudsang, Narayan Srinivasa, Jean Ponce
IROS3
1997 On Computing Aspect Graphs of Smooth Shapes from Volumetric Data
J. Alison Noble, Dale L. Wilson, Jean Ponce
Comput. Vis. Image Underst.3
1997 Recovering the Shape of Polyhedra Using Line-Drawing Analysis and Complex Reflectance Models
Ilan Shimshoni, Jean Ponce
Comput. Vis. Image Underst.2
1997 Hot curves for modelling and recognition of smooth curved 3D objects
abstract
: We represent arbitrary smooth curved 3D shapes by a discrete set of HOT curves where a surface admits High Order Tangents. These curves determine the structure of the image contours and its catastrophic changes, and there is a natural correspondence between some of them and monocular contour features such as inflections and bitangents. We present a method for automatically constructing the HOT curves from continuous sequences of video images and describe an approach to object recognition using viewpoint-dependent monocular image features as indices into a database of models and as a basis for pose estimation. We have implemented both the methods and present results obtained from real images. 1 Introduction While implemented recognition systems based on parametric shape representations such as algebraic surfaces or superquadrics have demonstrated their usefulness, the ultimate utility of a representation is limited by its scope. This suggests looking for a more general representation ...
Tanuja Joshi, B. Vijayakumar, David J. Kriegman, Jean Ponce
Image Vis. Comput.4
1997 Finite-Resolution Aspect Graphs of Polyhedral Objects
abstract
We address the problem of computing the aspect graph of a polyhedral object observed by an orthographic camera with limited spatial resolution, such that two image points separated by a distance smaller than a preset threshold cannot be resolved. Under this model, views that would differ under normal orthographic projection may become equivalent, while "accidental" views may occur over finite areas of the view space. We present a catalogue of visual events for polyhedral objects and give an algorithm for computing the aspect graph and enumerating all qualitatively different aspects. The algorithm has been fully implemented and results are presented.
Ilan Shimshoni, Jean Ponce
IEEE Trans. Pattern Anal. Mach. Intell.2
1996 Epipolar Geometry and Linear Subspace Methods: A New Approach to Weak Calibration
abstract
This paper addresses the problem of estimating the epipolar geometry from point correspondences between two images taken by uncalibrated perspective cameras. It is shown that Jepson's and Heeger's linear subspace technique for infinitesimal motion estimation can be generalized to the finite motion case by choosing an appropriate basis for projective space. This yields a linear method for weak calibration. The proposed algorithm has been implemented and tested on both real and synthetic images, and it is compared to other linear and non-linear approaches to weak calibration.
Jean Ponce, Yakup Genc
CVPR1
1996 Structure and motion of curved 3D objects from monocular silhouettes
abstract
The silhouette of a smooth 3D object observed by a moving camera changes over time. Past work has shown how surface geometry can be recovered using the deformation of the silhouette when the camera motion is known. This paper addresses the problem of estimating both the full Euclidean surface structure and the camera motion from a dense set of silhouettes captured under orthographic or scaled orthographic projection. The approach relies on a viewpoint-invariant representation of curves swept by viewpoint-dependent features such as bitangents, inflections and contour points with parallel tangents. Feature points, which form stereo frontier points between non-consecutive images, are matched using this representation. The camera's angular velocity is computed from constraints derived from this correspondence along with the image velocity of these features. From the angular velocity, the epipolar geometry is ascertained, and infinitesimal motion frontier points can be detected. In turn, the motion of these frontier points constrains the translation component of camera motion. Finally, the surface is reconstructed using established techniques once the camera motion has been estimated.
B. Vijayakumar, David J. Kriegman, Jean Ponce
CVPR3
1996 On planning immobilizing fixtures for three-dimensional polyhedral parts
abstract
We propose a simple three-dimensional modular fixturing device and present an algorithm for enumerating all of the immobilizing fixtures of a polyhedral object that can be achieved with this device and four frictionless contacts. Our approach is based on the second-order mobility theory of Rimon and Burdick (1993, 1994).
Jean Ponce
ICRA1
1995 Structure and Motion Estimation from Dynamic Silhouettes under Perspective Projection
abstract
Addresses the problem of estimating the structure and motion of a smooth curved object from its silhouettes observed over time by a trinocular stereo rig under perspective projection. We first construct a model for the local structure along the silhouette for each frame in the temporal sequence. Successive local models are then integrated into a global surface description by estimating the motion between successive time instants. The algorithm tracks certain surface features (parabolic points) and image features (silhouette inflections and frontier points) which are used to bootstrap the motion estimation process. The entire silhouette along with the reconstructed local structure are then used to refine the initial motion estimate. We have implemented the proposed approach and report results on real images.>
Tanuja Joshi, Narendra Ahuja, Jean Ponce
ICCV3
1995 Probabilistic 3D Object Recognition
abstract
A probabilistic 3D object recognition algorithm is presented. In order to guide the recognition process the probability that match hypotheses between image features and model features are correct is computed. A model is developed which uses the probabilistic peaking effect of measured angles and ratios of lengths by tracing iso angle and iso ratio curves on the viewing sphere. The model also accounts for various types of uncertainty in the input such as incomplete and inexact edge detection. For each match hypothesis the pose of the object and the pose uncertainty which is due to the uncertainty in vertex position are recovered. This is used to find sets of hypotheses which reinforce each other by matching features of the same object with compatible uncertainty subsets. A probabalistic expression is used to rank these hypothesis sets. The hypothesis sets with the highest rank are output. The algorithm has been fully implemented, and tested on real images.>
Ilan Shimshoni, Jean Ponce
ICCV2
1995 Invariant-Based Recognition of Complex Curved 3D Objects from Image Contours
abstract
To recognize three-dimensional objects bounded by smooth curved surfaces from monocular image contours, viewpoint-dependent image features must be related to object geometry. Contour bitangents and inflections along with associated parallel tangents points are the projection of surface points that lie on the occluding contour for a five-parameter family of scaled orthographic projection viewpoints. An invariant representation can be computed from these image features and seen for modeling and recognizing objects. Modeling is achieved by moving an object in front of a camera to obtain a curve of possible invariants. The relative camera-object motion is not required, and 3D models are not utilized. At recognition time, invariants computed from a single image are used to index the model database. Using the matched features, independent qualitative and quantitative verification procedures eliminate potential false matches. Examples from an implementation are presented.>
B. Vijayakumar, David J. Kriegman, Jean Ponce
ICCV3
1995 New Techniques for Computing Four-Finger-Force-Closure Grasps of Polyhedral Objects
abstract
It was shown in Ponce et al. (1993) that four-finger force-closure grasps fall into three categories: concurrent, pencil, and regulus grasps. The authors propose new techniques for computing these three types of grasps. The authors have implemented them and present examples.
Attawith Sudsang, Jean Ponce
ICRA2
1995 On computing three-finger force-closure grasps of polygonal objects
abstract
This paper addresses the problem of computing stable grasps of 2-D polygonal objects. We consider the case of a hand equipped with three hard fingers and assume point contact with friction. We prove new sufficient conditions for equilibrium and force closure that are linear in the unknown grasp parameters. This reduces computing the stable grasp regions in configuration space to constructing the three-dimensional projection of a five-dimensional polytope. We present an efficient projection algorithm based on linear programming and variable elimination among linear constraints. Maximal object segments where fingers can be positioned independently while ensuring force closure are found by linear optimization within the grasp regions. The approach has been implemented and several examples are presented.
Jean Ponce, Bernard Faverjon
IEEE Trans. Robotics Autom.1
1994 HOT curves for modelling and recognition of smooth curved 3D objects
abstract
Arbitrary smooth curved 3D shapes are represented by a discrete set of high-order tangent (HOT) curves, where a surface admits HOTs. These curves determine the structure of the image contours and its catastrophic changes, and there is a natural correspondence between some of them and monocular contour features such as inflections and bitangents. We present a method for automatically constructing the HOT curves from continuous sequences of video images and describe an approach to object recognition using viewpoint-dependent monocular image features as indices into a database of models and as a basis for pose estimation. We have implemented both of the methods, and present results obtained from real images.>
Tanuja Joshi, Jean Ponce, B. Vijayakumar, David J. Kriegman
CVPR2
1994 Object representation for object recognition
abstract
This paper discusses some representation issues and challenges involved in object recognition. It is intended as a step toward assessing current object representation schemes and proposing design and evaluation criteria for future ones.>
Jean Ponce, Ruzena Bajcsy, Dimitris N. Metaxas, Thomas O. Binford, David A. Forsyth, Martial Hebert, Katsushi Ikeuchi, Avinash C. Kak, Linda G. Shapiro, Stan Sclaroff, Alex Pentland, George C. Stockman
CVPR1
1994 Recovering the shape of polyhedra using line-drawing analysis and complex reflectance models
abstract
Following Sugihara, we represent the geometric constraints imposed by the line-drawing of a polyhedron as a set of linear equalities and inequalities. Unlike him, we explicitly take into account the uncertainty in vertex position. This allows us to circumvent the superstrictness of the constraints without deleting any of them. For a given error bound, deciding whether a line-drawing is the correct projection of a polyhedron is reduced to linear programming, and 3D shape recovery is reduced to optimization under linear constraints. Our method can be used for recovering the shape of polyhedral objects whose reflectance can be modelled accurately. We have implemented if for the Lambertian model and the Lambertian model with interreflections. We present results obtained using real images.>
Ilan Shimshoni, Jean Ponce
CVPR2
1994 Analytical Methods for Uncalibrated Stereo and Motion Reconstruction
Jean Ponce, David H. Marimont, Todd A. Cass
ECCV (1)1
1994 Geometric Methods for Relative Reconstruction from Weakly Calibrated Images
abstract
We present several new geometric methods for relative stereo and motion reconstruction using a discrete set of point correspondences. We suppose that the epipoles are known but do not assume any knowledge of the cameras' intrinsic or extrinsic parameters. In each case, we choose a set of five points as a basis for projective space and perform reconstruction relative to these five points. We also present a new technique for reprojection without reconstruction. We have implemented the proposed methods and present several examples using real images.>
Jean Ponce, David H. Marimont, Todd A. Cass
ICRA1
1994 Using Geometric Distance Fits for 3-D Object Modeling and Recognition
abstract
Addresses the problems of automatically constructing algebraic surface models from sets of 2D and 3D images and using these models in pose computation, motion and deformation estimation, and object recognition. We propose using a combination of constrained optimization and nonlinear least-squares estimation techniques to minimize the mean-squared geometric distance between a set of points or rays and a parameterized surface. In modeling tasks, the unknown parameters are the surface coefficients, while in pose and deformation estimation tasks they represent the transformation which maps the observer's coordinate system onto the modeled surface's own coordinate system. We have applied this approach to a variety of real range, computerized tomography and video images.>
Steve Sullivan, Lorraine Sandford, Jean Ponce
IEEE Trans. Pattern Anal. Mach. Intell.3
1994 Parameterized Families of Polynomials for Bounded Algebraic Curve and Surface Fitting
abstract
Interest in algebraic curves and surfaces of high degree as geometric models or shape descriptors for different model-based computer vision tasks has increased in recent years, and although their properties make them a natural choice for object recognition and positioning applications, algebraic curve and surface fitting algorithms often suffer from instability problems. One of the main reasons for these problems is that, while the data sets are always bounded, the resulting algebraic curves or surfaces are, in most cases, unbounded. In this paper, the authors propose to constrain the polynomials to a family with bounded zero sets, and use only members of this family in the fitting process. For every even number d the authors introduce a new parameterized family of polynomials of degree d whose level sets are always bounded, in particular, its zero sets. This family has the same number of degrees of freedom as a general polynomial of the same degree. Three methods for fitting members of this polynomial family to measured data points are introduced. Experimental results of fitting curves to sets of points in R/sup 2/ and surfaces to sets of points in R/sup 3/ are presented.>
Gabriel Taubin, Fernando Cukierman, Steve Sullivan, Jean Ponce, David J. Kriegman
IEEE Trans. Pattern Anal. Mach. Intell.4
1993 Reconstruction of HOT curves from image sequences
abstract
An approach is presented for reconstructing two types of 3-D higher order tangency (HOT) curves from a sequence of images. These curves are useful for object recognition. The reconstruction results for bitangents are encouraging in comparison to those of inflections. They are probably more accurately reconstructed because they are readily located in images, and their common tangent is very accurately estimated from the point locations. This makes bitangents a good feature choice for recognition.>
David J. Kriegman, B. Vijayakumar, Jean Ponce
CVPR3
1993 On using geometric distance fits to estimate 3D object shape, pose, and deformation from range, CT, and video images
abstract
The problems of automatically constructing algebraic surface models from sets of 3D and 2D images and using these models in pose computation, motion and deformation estimation, and object recognition are addressed. It is proposed that a combination of constrained optimization and nonlinear least-squares estimation techniques be used to minimize the mean-squared geometric distance between a set of points or rays and a parameterized surface. In modeling tasks, the unknown parameters are the surface coefficients, while in pose and deformation estimation tasks they represent the transformation mapping the observer's coordinate system onto the modeled surface's own coordinate system. This approach is applied to a variety of real range, computerized tomography (CT), and video images.>
Steve Sullivan, Lorraine Sandford, Jean Ponce
CVPR3
1992 Parametrizing and fitting bounded algebraic curves and surfaces
abstract
An approach to fitting of implicit algebraic curves and surfaces to point data is introduced. Two families of polynomials with bounded zero sets are presented. Members of these families have the same number of degrees of freedom as general polynomials of the same degree. Methods for fitting members of these families of polynomials to measured data points are described. Experimental results for sets of points in R/sup 2/ and R/sup 3/ for curves and surfaces, respectively, are presented.>
Gabriel Taubin, Fernando Cukierman, Steve Sullivan, Jean Ponce, David J. Kriegman
CVPR4
1992 Constraints for Recognizing and Locating Curved 3D Objects from Monocular Image Features
David J. Kriegman, B. Vijayakumar, Jean Ponce
ECCV3
1992 Computing Exact Aspect Graphs of Curved Objects: Algebraic Surfaces
Jean Ponce, Sylvain Petitjean, David J. Kriegman
ECCV1
1992 An algebraic approach to line-drawing analysis in the presence of uncertainty
abstract
Following the work of K. Sugihara (1984), the authors represent the geometric constraints imposed by the line-drawing of a polyhedron as a set of linear equalities and inequalities. They, however, explicitly take into account the uncertainty in the vertex position. This allows the circumvention of the superstrictness of the constraints without deleting any constraints. For a given error bound, the condition whether a line-drawing is the correct projection of a polyhedron is reduced to linear programing, and the 3D shape recovery is reduced to optimization under linear constraints. The approach has been implemented, and examples are presented.>
Jean Ponce, Ilan Shimshoni
ICRA1
1992 A System For Planning And Executing Two-finger Force-Closure Grasps Of Curved 2D Objects
abstract
Thzs paper presents a system for plannzng and executzng stable grasps of curved two-damenszonal objects. We conszder the case of a hand equzpped wzth two hard fingers and assume poznt contact wzth frzctzon Objects are modelled by parametrzc curves, and force-closure grasps are characterzzed by systems of polynomzal constraints an the parameters of these curves. All configuratzon space regzons satzsfyzng these constraznfs are found by a parallel numerzcal cell decomposztzon algoriihni based on curve tracing and continadzon techniques ~IULLIILU~ object segiiienls diel e fingers can be positzoned an depeii dently are found by optimzzateon wzlhzn the grasp regaons The approuch has been implemented uszng a dzstrzbuted archztecture, and experzinents using a PUMA rohot equipped ubiih a pneumatzc two-finger grippe? aid U vzszoii system are presented
Darrell Stam, Jean Ponce, Bernard Faverjon
IROS2
1992 On using CAD models to compute the pose of curved 3D objects
Jean Ponce, Anthony Hoogs, David J. Kriegman
CVGIP Image Underst.1
1992 Computing exact aspect graphs of curved objects: Algebraic surfaces
Sylvain Petitjean, Jean Ponce, David J. Kriegman
Int. J. Comput. Vis.2
1991 On computing two-finger force-closure grasps of curved 2D objects
abstract
An approach to the computation of stable grasps of curved two-dimensional objects is presented. The authors consider the case of a hand equipped with two hard fingers and assume point contact with friction. Objects are modeled by parametric curves, and force-closure grasps are characterized by systems of polynomial constraints in the parameters of these curves. All configuration space regions satisfying these constraints are found by a numerical cell decomposition algorithm based on curve tracing and continuation techniques. Maximal object segments where fingers can be positioned independently are found by optimization within the grasp regions. The approach has been implemented and examples are presented.>
Bernard Faverjon, Jean Ponce
ICRA2
1990 Computing Exact Aspect Graphs of Curved Objects: Parametric Surfaces
Jean Ponce, David J. Kriegman
AAAI1
1990 On characterizing ribbons and finding skewed symmetries
Jean Ponce
Comput. Vis. Graph. Image Process.1
1990 Computing exact aspect graphs of curved objects: Solids of revolution
David J. Kriegman, Jean Ponce
Int. J. Comput. Vis.2
1990 Straight homogeneous generalized cylinders: Differential geometry and uniqueness results
Jean Ponce
Int. J. Comput. Vis.1
1990 On Recognizing and Positioning Curved 3-D Objects from Image Contours
abstract
An approach for explicitly relating the shape of image contours to models of curved three-dimensional objects is presented. This relationship is used for object recognition and positioning. Object models consist of collections of parametric surface patches and their intersection curves; this includes nearly all representations used in computer-aided geometric design and computer vision. The image contours considered are the projections of surface discontinuities and occluding contours. Elimination theory provides a method for constructing the implicit equation of these contours for an object observed under orthographic or perspective projection. This equation is parameterized by the object's position and orientation with respect to the observer. Determining these parameters is reduced to a fitting problem between the theoretical contour and the observed data points. The proposed approach readily extends to parameterized models. It has been implemented for a simple world composed of various surfaces of revolution and tested on several real images.>
David J. Kriegman, Jean Ponce
IEEE Trans. Pattern Anal. Mach. Intell.2
1989 On characterizing ribbons and finding skewed symmetries
abstract
The author compares Blum, Brooks, and Brady ribbons, and proves that Blum and Brady ribbons are not, in general, Brooks ribbons. Conversely, he proves that Brook ribbons are, in general, neither Blum nor Brady ribbons. For Blum and Brooks ribbons, it is in principle trivial to decide whether two contour points may form a ribbon pair; they have to form a local symmetry. This property is not true for Brooks ribbons. Attention is also given to whether it is possible to characterize locally the pairs of contour points which form a Brooks ribbon pair. Using the curvature of a Brooks ribbon, it is shown that this is possible for some classes of Brooks ribbons, including skewed symmetries. This result is used in an implemented algorithm for finding skewed symmetries in an image, and examples of segmentation of real images are given.>
Jean Ponce
ICRA1
1989 Invariant Properties of Straight Homogeneous Generalized Cylinders and Their Contours
abstract
A fundamental group in computer vision is the recovery of three-dimensional shape from image data. While this problem is in general underconstrained, the authors show that it can be simplified in the case where the objects being viewed are generalized cylinders. They consider the class of straight homogeneous generalized cylinders (SHGCs), without any further assumption on the viewing direction or the precise shape of these objects. They present a rigorous mathematical study of the geometry of SHGCs and characterize their Gaussian curvature and occluding contours, and use these results to prove several new invariant properties of the contours of SHGCs. These properties are, in turn, used in two implemented algorithms for recovering SHGC descriptions from image contours. Several examples of segmentation of real images are given. Other applications are also discussed.>
Jean Ponce, David M. Chelberg, Wallace B. Mann
IEEE Trans. Pattern Anal. Mach. Intell.1
1988 Straight homogeneous generalized cylinders: differential geometry and uniqueness results
abstract
The author studies the differential geometry of straight homogeneous generalized cylinders (SHGCs). He derives a necessary and sufficient condition that an SHGC must verify to parameterize a regular surface, computes the Gaussian curvature of a regular SHGC, and proves that the parabolic lines of an SHGC are either meridians or parallels. Using these results, he addresses the following problem: under which conditions can a given surface have several descriptions by SHGCs? He proves several results. In particular, he proves that two SHGCs with the same cross-section plane and axis direction are necessarily deduced from each other through inverse scalings of their cross-sections and sweeping rule curve. He extends Shafer's pivot and slant theorems. Finally, he proves that a surface with at least two parabolic lines has at most three different SHGC descriptions, and that a surface with at least four parabolic lines has at most a unique SHGC description.>
Jean Ponce
CVPR1
1988 Finding the limbs and cusps of generalized cylinders
Jean Ponce, David M. Chelberg
Int. J. Comput. Vis.1
1987 Finding the limbs and cusps of generalized cylinders
abstract
This paper addresses the problem of finding analytically the limbs and cusps of generalized cylinders. Orthographic projections of generalized cylinders whose axis is straight and whose axis is an arbitrary 3D curve are considered in turn. In both cases, the general equations of the limbs and cusps are given. They are solved for three classes of generalized cylinders: solids of revolution, straight homogeneous generalized cylinders whose scaling sweeping rule is a polynomial of degree less than or equal to 5 and generalized cylinders whose axis is an arbitrary 3D curve but the cross section is circular and constant. Examples of limbs and cusps found for each class are given. Extensions and applications of the results presented are discussed.
Jean Ponce, David M. Chelberg
ICRA1
1987 Localized intersections computation for solid modelling with straight homogenous generalized cylinders
abstract
This paper reports progress in the development of a solid modelling system combining straight homogeneous generalized cylinders through set operations. Two basic components of this system are the modules which compute the set operations between primitives and display the resulting solids using ray tracing. These two modules are also very computationally intensive as they involve a large number of surface-surface and ray-surface intersections computations. We introduce a novel hierarchical representation for straight homogeneous cylinders called Box Tree. The Box Tree is analogous to a Quadtree in parameter space. It is an exact boundary representation which describes the surface of the associated generalized cylinder by a hierarchy of enclosing boxes. We use the Box Tree to efficiently compute the set operations and ray tracing algorithms by localizing the search for intersections to the regions where they may occur. We discuss complexity issues and illustrate the performances of our modelling system on a variety of examples.
Jean Ponce, David M. Chelberg
ICRA1
1987 An object centered hierarchical representation for 3D objects: The prism tree
Jean Ponce, Olivier D. Faugeras
Comput. Vis. Graph. Image Process.1
1985 Toward a surface primal sketch
abstract
This paper reports progress toward the development of a representation of significant surface changes in dense depth maps. We call tile representation the Surface Primal Sketch by analogy with representations of intensity changes, image structure, and changes in curvature of planar curves. We describe an implemented program that detects, localizes, and symbolically describes: steps, where the surface height function is discontinuous, and roofs, where the surface is continuous but the surface normal is discontinuous. We illustrate the performance of the program on range maps of objects of varying complexity.
Jean Ponce, J. Michael Brady
ICRA1
1985 Describing surfaces
J. Michael Brady, Jean Ponce, Alan L. Yuille, Haruo Asada
Comput. Vis. Graph. Image Process.2
1983 Prism Trees: A Hierarchical Representation for 3-D Objects
Olivier D. Faugeras, Jean Ponce
IJCAI2