EDBT 2026 Demo / reviewers in the wild / expert
Enrique Dunn
dblp:27/5341
· DBLP profile ↗
45ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0002-5436-0023ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 38 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 36 · 6 first-author · 9 since 2021Systems, architecture and hardware · 3 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Camera Resection from Known Line Pencils and a Radially Distorted ScanlineabstractWe present a marker-based geometric estimation framework for the absolute pose of a camera by analyzing the 1D observations in a single radially distorted pixel scanline. We leverage a pair of known co-planar pencils of lines, along with lens distortion parameters, to propose an ensemble of solvers exploring the space of estimation strategies applicable to our setup. First, we present a minimal algebraic solver requiring only six measurements and yielding eight solutions, which relies on the intersection of two conics defined by one of the pencils of lines. Then, we present a unique closed-form geometric solver from seven measurements. Finally, we present an homography-based formulation amenable to linear least-squares from eight or more measurements. Our geometric framework constitutes a theoretical analysis on the minimum geometric context necessary to solve in closed form for the absolute pose of a single camera from a single radially distorted scanline. Juan Carlos Dibene, Enrique Dunn |
CVPR | 2 |
| 2025 | DRaM-LHM: A Quaternion Framework for Iterative Camera Pose Estimation
Weizhi Du, Zhixiang Min, Baochen She, Enrique Dunn, Sonya M. Hanson |
ICCV | 5 |
| 2024 | Hybrid and Non-minimal Planar Motion Estimation from Point Correspondences
Juan Carlos Dibene, Enrique Dunn |
ACCV (9) | 2 |
| 2023 | NeurOCS: Neural NOCS Supervision for Monocular 3D Object LocalizationabstractMonocular 3D object localization in driving scenes is a crucial task, but challenging due to its ill-posed nature. Estimating 3D coordinates for each pixel on the object surface holds great potential as it provides dense 2D-3D geometric constraints for the underlying PnP problem. However, high-quality ground truth supervision is not available in driving scenes due to sparsity and various artifacts of Lidar data, as well as the practical infeasibility of collecting per-instance CAD models. In this work, we present NeurOCS, a framework that uses instance masks and 3D boxes as input to learn 3D object shapes by means of differentiable rendering, which further serves as supervision for learning dense object coordinates. Our approach rests on insights in learning a category-level shape prior directly from real driving scenes, while properly handling single-view ambiguities. Furthermore, we study and make critical design choices to learn object coordinates more effectively from an object-centric view. Altogether, our framework leads to new state-of-the-art in monocular 3D localization that ranks 1st on the KITTI-Object [16] benchmark among published monocular methods. Zhixiang Min, Bingbing Zhuang, Samuel Schulter, Buyu Liu, Enrique Dunn, Manmohan Krishna Chandraker |
CVPR | 5 |
| 2023 | General Planar Motion from a Pair of 3D CorrespondencesabstractWe present a novel 2-point method for estimating the relative pose of a camera undergoing planar motion from 3D data (e.g. from a calibrated stereo setup or an RGBD sensor). Unlike prior art, our formulation does not assume knowledge of the plane of motion, (e.g. parallelism between the optical axis and motion plane) to resolve the under-constrained nature of SE(3) motion estimation in this context. Instead, we enforce geometric constraints identifying, in closed-form, a unique planar motion solution from an orbital set of geometrically consistent SE(3) motion estimates. We explore the set of special and degenerate geometric cases arising from our formulation. Experiments on synthetic data characterize the sensitivity of our estimation framework to measurement noise and different types of observed motion. We integrate our solver within a RANSAC framework and demonstrate robust operation on standard benchmark sequences of real-world imagery. Code is available at: https://github.com/jdibenes/gpm. Juan Carlos Dibene, Zhixiang Min, Enrique Dunn |
ICCV | 3 |
| 2023 | Geometric Viewpoint Learning with Hyper-Rays and Harmonics EncodingabstractViewpoint is a fundamental modality that carries the interaction between observers and their environment. This paper proposes the first deep-learning framework for the viewpoint modality. The challenge in formulating learning frameworks for viewpoints resides in a suitable multimodal representation that links across the camera viewing space and 3D environment. Traditional approaches reduce the problem to image analysis instances, making them computationally expensive and not adequately modelling the intrinsic geometry and environmental context of 6DoF viewpoints. We improve these issues in two ways. 1) We propose a generalized viewpoint representation forgoing the analysis of photometric pixels in favor of encoded viewing ray embeddings attained from point cloud learning frameworks. 2) We propose a novel SE(3)-bijective 6D viewing ray, hyper-ray, that addresses the DoF deficiency problem of using 5DoF viewing rays representing 6DoF viewpoints. We demonstrate our approach has both efficiency and accuracy superiority over existing methods in novel real-world environments. Zhixiang Min, Juan Carlos Dibene, Enrique Dunn |
ICCV | 3 |
| 2022 | LASER: LAtent SpacE Rendering for 2D Visual LocalizationabstractWe present LASER, an image-based Monte Carlo Localization (MCL) framework for 2D floor maps. LASER introduces the concept of latent space rendering, where 2D pose hypotheses on the floor map are directly rendered into a geometrically-structured latent space by aggregating viewing ray features. Through a tightly coupled rendering codebook scheme, the viewing ray features are dynamically determined at rendering-time based on their geometries (i.e. length, incident-angle), endowing our representation with view-dependent fine-grain variability. Our codebook scheme effectively disentangles feature encoding from rendering, allowing the latent space rendering to run at speeds above 10KHz. Moreover, through metric learning, our geometrically-structured latent space is common to both pose hypotheses and query images with arbitrary field of views. As a result, LASER achieves state-of-the-art performance on large-scale indoor localization datasets (i. e. ZInD [5] and Structured3D [38]) for both panorama and perspective image queries, while significantly outperforming existing learning-based methods in speed. Zhixiang Min, Naji Khosravan, Zachary Bessinger, Manjunath Narayana, Sing Bing Kang, Enrique Dunn, Ivaylo Boyadzhiev |
CVPR | 6 |
| 2022 | Prepare for Ludicrous Speed: Marker-based Instantaneous Binocular Rolling Shutter LocalizationabstractWe propose a marker-based geometric framework for the high-frequency absolute 3D pose estimation of a binocular camera system by using the data captured during the exposure of a single rolling shutter scanline. In contrast to existing approaches enforcing temporal or motion models among scanlines (e.g. linear motion, constant velocity or small motion assumptions), we strive to determine the pose from instantaneous binocular capture (i.e. without using data from previous scanlines) and achieve drift-free pose estimation. We leverage the projective invariants of a novel rigid planar pattern, to both define a geometric reference as well as to determine 2D-3D correspondences from raw edge detection measurements from individual scanlines. Moreover, to tackle the ensuing multi-view estimation problem, achieve real-time operation, and minimize latency, we develop a pair of custom solvers leveraging our geometric setup. To mitigate sensitivity to noise, we propose a geometrically consistent measurement refinement mechanism. We verify the quality of our solvers by comparing with state of the art general solvers for absolute pose estimation of generalized cameras. Finally, we demonstrate the effectiveness of our proposed approach with an FPGA-based implementation which achieves a localization throughput of 129.6 KHz with a 1.5 μs latency. Juan Carlos Dibene, Yazmín Maldonado, Leonardo Trujillo 0001, Enrique Dunn |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2021 | GTT-Net: Learned Generalized Trajectory TriangulationabstractWe present GTT-Net, a supervised learning framework for the reconstruction of sparse dynamic 3D geometry. We build on a graph-theoretic formulation of the generalized trajectory triangulation problem, where non-concurrent multi-view imaging geometry is known but global image sequencing is not provided. GTT-Net learns pairwise affinities modeling the spatio-temporal relationships among our input observations and leverages them to determine 3D geometry estimates. Experiments reconstructing 3D motion-capture sequences show GTT-Net outperforms the state of the art in terms of accuracy and robustness. Within the context of articulated motion reconstruction, our proposed architecture is 1) able to learn and enforce semantic 3D motion priors for shared training and test domains, while being 2) able to generalize its performance across different training and test domains. Moreover, GTT-Net provides a computationally streamlined framework for trajectory triangulation with applications to multi-instance reconstruction and event segmentation. Xiangyu Xu 0004, Enrique Dunn |
ICCV | 2 |
| 2021 | VOLDOR+SLAM: For the times when feature-based or direct methods are not good enoughabstractWe present a dense-indirect SLAM system using external dense optical flows as input. We extend the recent probabilistic visual odometry model VOLDOR [1], by incorporating the use of geometric priors to 1) robustly bootstrap estimation from monocular capture, while 2) seamlessly supporting stereo and/or RGB-D input imagery. Our customized back-end tightly couples our intermediate geometric estimates with an adaptive priority scheme managing the connectivity of an incremental pose graph. We leverage recent advances in dense optical flow methods to achieve accurate and robust camera pose estimates, while constructing fine-grain globally-consistent dense environmental maps. Our open source implementation [https://github.com/htkseason/VOLDOR] operates online at around 15 FPS on a single GTX1080Ti GPU. Zhixiang Min, Enrique Dunn |
ICRA | 2 |
| 2020 | VOLDOR: Visual Odometry From Log-Logistic Dense Optical Flow ResidualsabstractWe propose a dense indirect visual odometry method taking as input externally estimated optical flow fields instead of hand-crafted feature correspondences. We define our problem as a probabilistic model and develop a generalized-EM formulation for the joint inference of camera motion, pixel depth, and motion-track confidence. Contrary to traditional methods assuming Gaussian-distributed observation errors, we supervise our inference framework under an (empirically validated) adaptive log-logistic distribution model. Moreover, the log-logistic residual model generalizes well to different state-of-the-art optical flow methods, making our approach modular and agnostic to the choice of optical flow estimators. Our method achieved top-ranking results on both TUM RGB-D and KITTI odometry benchmarks. Our open-sourced implementation is inherently GPU-friendly with only linear computational and storage growth. Zhixiang Min, Yiding Yang, Enrique Dunn |
CVPR | 3 |
| 2019 | Discrete Laplace Operator Estimation for Dynamic 3D ReconstructionabstractWe present a general paradigm for dynamic 3D reconstruction from multiple independent and uncontrolled image sources having arbitrary temporal sampling density and distribution. Our graph-theoretic formulation models the spatio-temporal relationships among our observations in terms of the joint estimation of their 3D geometry and its discrete Laplace operator. Towards this end, we define a tri-convex optimization framework that leverages the geometric properties and dependencies found among a Euclidean shape-space and the discrete Laplace operator describing its local and global topology. We present a reconstructability analysis, experiments on motion capture data and multi-view image datasets, as well as explore applications to geometry-based event segmentation and data association. Xiangyu Xu 0004, Enrique Dunn |
ICCV | 2 |
| 2018 | Dynamic Visual Sequence Prediction with Motion Flow NetworksabstractWe target the problem of synthesizing future motion sequences from a temporally ordered set of input images. Previous methods tackled this problem in two manners: predicting the future image pixel values and predicting the dense time-space trajectory of pixels. Towards this end, generative encoder-decoder networks have been widely adopted in both kinds of methods. However, pixel prediction with these networks has been shown to suffer from blurry outputs, since images are generated from scratch and there is no explicit enforcement of visual coherency. Alternately, crisp details can be achieved by transferring pixels from the input image through dense trajectory predictions, but this process requires pre-computed motion fields for training, which limit the learning ability for the neural networks. To synthesize realistic movement of objects under weak supervision (without pre-computed dense motion fields), we propose two novel network structures. Our first network encodes the input images as feature maps, and uses a decoder network to predict the future pixel correspondences for a series of subsequent time steps. The attained correspondence fields are then used to synthesize future views. Our second network focuses on human-centered capture by augmenting our framework to include sparse pose estimates [30] to guide our dense correspondence prediction. Compared with state-of-the-art pixel generating and dense trajectories predicting networks, our model performs better on synthetic as well as on real-world human body movement sequences. Dinghuang Ji, Enrique Dunn, Jan-Michael Frahm |
WACV | 3 |
| 2018 | Self-Expressive Dictionary Learning for Dynamic 3D ReconstructionabstractWe target the problem of sparse 3D reconstruction of dynamic objects observed by multiple unsynchronized video cameras with unknown temporal overlap. To this end, we develop a framework to recover the unknown structure without sequencing information across video sequences. Our proposed compressed sensing framework poses the estimation of 3D structure as the problem of dictionary learning, where the dictionary is defined as an aggregation of the temporally varying 3D structures. Given the smooth motion of dynamic objects, we observe any element in the dictionary can be well approximated by a sparse linear combination of other elements in the same dictionary (i.e., self-expression). Our formulation optimizes a biconvex cost function that leverages a compressed sensing formulation and enforces both structural dependency coherence across video streams, as well as motion smoothness across estimates from common video sources. We further analyze the reconstructability of our approach under different capture scenarios, and its comparison and relation to existing methods. Experimental results on large amounts of synthetic data as well as real imagery demonstrate the effectiveness of our approach. Enliang Zheng, Dinghuang Ji, Enrique Dunn, Jan-Michael Frahm |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | Learned Contextual Feature Reweighting for Image Geo-LocalizationabstractWe address the problem of large scale image geo-localization where the location of an image is estimated by identifying geo-tagged reference images depicting the same place. We propose a novel model for learning image representations that integrates context-aware feature reweighting in order to effectively focus on regions that positively contribute to geo-localization. In particular, we introduce a Contextual Reweighting Network (CRN) that predicts the importance of each region in the feature map based on the image context. Our model is learned end-to-end for the image geo-localization task, and requires no annotation other than image geo-tags for training. In experimental results, the proposed approach significantly outperforms the previous state-of-the-art on the standard geo-localization benchmark datasets. We also demonstrate that our CRN discovers task-relevant contexts without any additional supervision. Hyo Jin Kim 0004, Enrique Dunn, Jan-Michael Frahm |
CVPR | 2 |
| 2016 | Bringing 3D Models Together: Mining Video Liaisons in Crowdsourced Reconstructions
Ke Wang 0021, Enrique Dunn, Mikel Rodriguez, Jan-Michael Frahm |
ACCV (4) | 2 |
| 2016 | Spatio-Temporally Consistent Correspondence for Dense Dynamic Scene Modeling
Dinghuang Ji, Enrique Dunn, Jan-Michael Frahm |
ECCV (6) | 2 |
| 2016 | Efficient joint stereo estimation and land usage classification for multiview satellite dataabstractWe propose an efficient algorithm to jointly estimate geometry and semantics for a given geographical region observed by multiple satellite images. Our joint estimation leverages an efficient PatchMatch inference framework defined over lattice discretization of the environment. Our cost function relies on the local planarity assumption to model scene geometry and neural network classification to determine semantic (e.g. land use) labels for geometric structures. By utilizing the commonly available direct (i.e. space to image) rational polynomial coefficients (RPC) satellite camera models, our approach effectively circumvents the need for estimating or refining inverse RPC models. Experiments illustrate both the computational efficiency and high quality scene geometry estimates attained by our approach for satellite imagery. To further illustrate the generality of our representation and inference framework, experiments on standard benchmarks for ground-level imagery are also included. Ke Wang 0021, Craig Stutts, Enrique Dunn, Jan-Michael Frahm |
WACV | 3 |
| 2016 | Towards Kilo-Hertz 6-DoF Visual Tracking Using an Egocentric Cluster of Rolling Shutter CamerasabstractTo maintain a reliable registration of the virtual world with the real world, augmented reality (AR) applications require highly accurate, low-latency tracking of the device. In this paper, we propose a novel method for performing this fast 6-DOF head pose tracking using a cluster of rolling shutter cameras. The key idea is that a rolling shutter camera works by capturing the rows of an image in rapid succession, essentially acting as a high-frequency 1D image sensor. By integrating multiple rolling shutter cameras on the AR device, our tracker is able to perform 6-DOF markerless tracking in a static indoor environment with minimal latency. Compared to state-of-the-art tracking systems, this tracking approach performs at significantly higher frequency, and it works in generalized environments. To demonstrate the feasibility of our system, we present thorough evaluations on synthetically generated data with tracking frequencies reaching 56.7 kHz. We further validate the method's accuracy on real-world images collected from a prototype of our tracking system against ground truth data using standard commodity GoPro cameras capturing at 120 Hz frame rate. Akash Bapat, Enrique Dunn, Jan-Michael Frahm |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2015 | Reconstructing the world* in six daysabstractWe propose a novel, large-scale, structure-from-motion framework that advances the state of the art in data scalability from city-scale modeling (millions of images) to world-scale modeling (several tens of millions of images) using just a single computer. The main enabling technology is the use of a streaming-based framework for connected component discovery. Moreover, our system employs an adaptive, online, iconic image clustering approach based on an augmented bag-of-words representation, in order to balance the goals of registration, comprehensiveness, and data compactness. We demonstrate our proposal by operating on a recent publicly available 100 million image crowd-sourced photo collection containing images geographically distributed throughout the entire world. Results illustrate that our streaming-based approach does not compromise model completeness, but achieves unprecedented levels of efficiency and scalability. Jared Heinly, Johannes L. Schönberger, Enrique Dunn, Jan-Michael Frahm |
CVPR | 3 |
| 2015 | Synthesizing Illumination Mosaics from Internet Photo-CollectionsabstractWe propose a framework for the automatic creation of time-lapse mosaics of a given scene. We achieve this by leveraging the illumination variations captured in Internet photo-collections. In order to depict and characterize the illumination spectrum of a scene, our method relies on building discrete representations of the image appearance space through connectivity graphs defined over a pairwise image distance function. The smooth appearance transitions are found as the shortest path in the similarity graph among images, and robust image alignment is achieved by leveraging scene semantics, multi-view geometry, and image warping techniques. The attained results present an insightful and compact visualization of the scene illuminations captured in crowd-sourced imagery. Dinghuang Ji, Enrique Dunn, Jan-Michael Frahm |
ICCV | 2 |
| 2015 | Predicting Good Features for Image Geo-Localization Using Per-Bundle VLADabstractWe address the problem of recognizing a place depicted in a query image by using a large database of geo-tagged images at a city-scale. In particular, we discover features that are useful for recognizing a place in a data-driven manner, and use this knowledge to predict useful features in a query image prior to the geo-localization process. This allows us to achieve better performance while reducing the number of features. Also, for both learning to predict features and retrieving geo-tagged images from the database, we propose per-bundle vector of locally aggregated descriptors (PBVLAD), where each maximally stable region is described by a vector of locally aggregated descriptors (VLAD) on multiple scale-invariant features detected within the region. Experimental results show the proposed approach achieves a significant improvement over other baseline methods. Hyo Jin Kim 0004, Enrique Dunn, Jan-Michael Frahm |
ICCV | 2 |
| 2015 | Sparse Dynamic 3D Reconstruction from Unsynchronized VideosabstractWe target the sparse 3D reconstruction of dynamic objects observed by multiple unsynchronized video cameras with unknown temporal overlap. To this end, we develop a framework to recover the unknown structure without sequencing information across video sequences. Our proposed compressed sensing framework poses the estimation of 3D structure as the problem of dictionary learning. Moreover, we define our dictionary as the temporally varying 3D structure, while we define local sequencing information in terms of the sparse coefficients describing a locally linear 3D structural interpolation. Our formulation optimizes a biconvex cost function that leverages a compressed sensing formulation and enforces both structural dependency coherence across video streams, as well as motion smoothness across estimates from common video sources. Experimental results demonstrate the effectiveness of our approach in both synthetic data and captured imagery. Enliang Zheng, Dinghuang Ji, Enrique Dunn, Jan-Michael Frahm |
ICCV | 3 |
| 2015 | Minimal Solvers for 3D Geometry from Satellite ImageryabstractWe propose two novel minimal solvers which advance the state of the art in satellite imagery processing. Our methods are efficient and do not rely on the prior existence of complex inverse mapping functions to correlate 2D image coordinates and 3D terrain. Our first solver improves on the stereo correspondence problem for satellite imagery, in that we provide an exact image-to-object space mapping (where prior methods were inaccurate). Our second solver provides a novel mechanism for 3D point triangulation, which has improved robustness and accuracy over prior techniques. Given the usefulness and ubiquity of satellite imagery, our proposed methods allow for improved results in a variety of existing and future applications. Enliang Zheng, Ke Wang 0021, Enrique Dunn, Jan-Michael Frahm |
ICCV | 3 |
| 2014 | Recovering Correct Reconstructions from Indistinguishable GeometryabstractStructure-from-motion (SFM) is widely utilized to generate 3D reconstructions from unordered photo-collections. However, in the presence of non unique, symmetric, or otherwise indistinguishable structure, SFM techniques often incorrectly reconstruct the final model. We propose a method that not only determines if an error is present, but automatically corrects the error in order to produce a correct representation of the scene. We find that by exploiting the co-occurrence information present in the scene's geometry, we can successfully isolate the 3D points causing the incorrect result. This allows us to split an incorrect reconstruction into error-free sub-models that we then correctly merge back together. Our experimental results show that our technique is efficient, robust to a variety of scenes, and outperforms existing methods. Jared Heinly, Enrique Dunn, Jan-Michael Frahm |
3DV | 2 |
| 2014 | Stereo under Sequential Optimal Sampling: A Statistical Analysis Framework for Search Space ReductionabstractWe develop a sequential optimal sampling framework for stereo disparity estimation by adapting the Sequential Probability Ratio Test (SPRT) model. We operate over local image neighborhoods by iteratively estimating single pixel disparity values until sufficient evidence has been gathered to either validate or contradict the current hypothesis regarding local scene structure. The output of our sampling is a set of sampled pixel positions along with a robust and compact estimate of the set of disparities contained within a given region. We further propose an efficient plane propagation mechanism that leverages the pre-computed sampling positions and the local structure model described by the reduced local disparity set. Our sampling framework is a general pre-processing mechanism aimed at reducing computational complexity of disparity search algorithms by ascertaining a reduced set of disparity hypotheses for each pixel. Experiments demonstrate the effectiveness of the proposed approach when compared to state of the art methods. Yilin Wang 0001, Ke Wang 0021, Enrique Dunn, Jan-Michael Frahm |
CVPR | 3 |
| 2014 | PatchMatch Based Joint View Selection and Depthmap EstimationabstractWe propose a multi-view depthmap estimation approach aimed at adaptively ascertaining the pixel level data associations between a reference image and all the elements of a source image set. Namely, we address the question, what aggregation subset of the source image set should we use to estimate the depth of a particular pixel in the reference image? We pose the problem within a probabilistic framework that jointly models pixel-level view selection and depthmap estimation given the local pairwise image photoconsistency. The corresponding graphical model is solved by EM-based view selection probability inference and PatchMatch-like depth sampling and propagation. Experimental results on standard multi-view benchmarks convey the state-of-the art estimation accuracy afforded by mitigating spurious pixel level data associations. Additionally, experiments on large Internet crowd sourced data demonstrate the robustness of our approach against unstructured and heterogeneous image capture characteristics. Moreover, the linear computational and storage requirements of our formulation, as well as its inherent parallelism, enables an efficient and scalable GPU-based implementation. Enliang Zheng, Enrique Dunn, Vladimir Jojic, Jan-Michael Frahm |
CVPR | 2 |
| 2014 | Correcting for Duplicate Scene Structure in Sparse 3D Reconstruction
Jared Heinly, Enrique Dunn, Jan-Michael Frahm |
ECCV (4) | 2 |
| 2014 | 3D Reconstruction of Dynamic Textures in Crowd Sourced Data
Dinghuang Ji, Enrique Dunn, Jan-Michael Frahm |
ECCV (1) | 2 |
| 2014 | Joint Object Class Sequencing and Trajectory Triangulation (JOST)
Enliang Zheng, Ke Wang 0021, Enrique Dunn, Jan-Michael Frahm |
ECCV (7) | 3 |
| 2014 | P-HRTF: Efficient personalized HRTF computation for high-fidelity spatial soundabstractAccurate rendering of 3D spatial audio for interactive virtual auditory displays requires the use of personalized head-related transfer functions (HRTFs). We present a new approach to compute personalized HRTFs for any individual using a method that combines state-of-the-art image-based 3D modeling with an efficient numerical simulation pipeline. Our 3D modeling framework enables capture of the listener's head and torso using consumer-grade digital cameras to estimate a high-resolution non-parametric surface representation of the head, including the extended vicinity of the listener's ear. We leverage sparse structure from motion and dense surface reconstruction techniques to generate a 3D mesh. This mesh is used as input to a numeric sound propagation solver, which uses acoustic reciprocity and Kirchhoff surface integral representation to efficiently compute an individual's personalized HRTF. The overall computation takes tens of minutes on multi-core desktop machine. We have used our approach to compute the personalized HRTFs of few individuals, and we present our preliminary evaluation here. To the best of our knowledge, this is the first commodity technique that can be used to compute personalized HRTFs in a lab or home setting. Alok Meshram, Ravish Mehra, Hongsheng Yang, Enrique Dunn, Jan-Michael Frahm, Dinesh Manocha |
ISMAR | 4 |
| 2014 | Rotation estimation from cloud trackingabstractWe address the problem of online relative orientation estimation from streaming video captured by a sky-facing camera on a mobile device. Namely, we rely on the detection and tracking of visual features attained from cloud structures. Our proposed method achieves robust and efficient operation by combining realtime visual odometry modules, learning based feature classification, and Kalman filtering within a robustness-driven data management framework, while achieving framerate processing on a mobile device. The relatively large 3D distance between the camera and the observed cloud features is leveraged to simplify our processing pipeline. First, as an efficiency driven optimization, we adopt a homography based motion model and focus on estimating relative rotations across adjacent keyframes. To this end, we rely on efficient feature extraction, KLT tracking, and RANSAC based model fitting. Second, to ensure the validity of our simplified motion model, we segregate detected cloud features from scene features through SVM classification. Finally, to make tracking more robust, we employ predictive Kalman filtering to enable feature persistence through temporary occlusions and manage feature spatial distribution to foster tracking robustness. Results exemplify the accuracy and robustness of the proposed approach and highlight its potential as a passive orientation sensor. Sangwoo Cho, Enrique Dunn, Jan-Michael Frahm |
WACV | 2 |
| 2014 | Combining semantic scene priors and haze removal for single image depth estimationabstractWe consider the problem of estimating the relative depth of a scene from a monocular image. The dark channel prior, used as a statistical observation of haze free images, has been previously leveraged for haze removal and relative depth estimation tasks. However, as a local measure, it fails to account for higher order semantic relationship among scene elements. We propose a dual channel prior used for identifying pixels that are unlikely to comply with the dark channel assumption, leading to erroneous depth estimates. We further leverage semantic segmentation information and patch match label propagation to enforce semantically consistent geometric priors. Experiments illustrate the quantitative and qualitative advantages of our approach when compared to state of the art methods. Ke Wang 0021, Enrique Dunn, Joseph Tighe, Jan-Michael Frahm |
WACV | 2 |
| 2012 | Efficient and Scalable Depthmap Fusion
Enliang Zheng, Enrique Dunn, Rahul Raguram, Jan-Michael Frahm |
BMVC | 2 |
| 2012 | Comparative Evaluation of Binary Features
Jared Heinly, Enrique Dunn, Jan-Michael Frahm |
ECCV (2) | 2 |
| 2011 | Adaptive Scale Selection for Hierarchical StereoabstractHierarchical stereo provides an efficient coarse-to-fine mechanism for disparity map estimation.However, common drawbacks of such an approach include the loss of high frequency structures not observable at coarse scale levels, as well as the unrecoverable propagation of erroneous disparity estimates through the scale space.This paper presents an adaptive scale selection mechanism to determine a suitable resolution level from which to begin the hierarchical depth estimation process for each pixel.The proposed scale selection mechanism allows us to robustly implement variable cost aggregation in order to reduce the variability of the photo-consistency measure across scale space.We also incorporate a weighted shiftable window mechanism to enable error correction during coarse-to-fine depth refinement.Experiments illustrate the effectiveness of our approach in terms of disparity accuracy, while attaining a computational efficiency compromise between full resolution and hierarchical disparity map estimation. Yi-Hung Jen, Enrique Dunn, Pierre Fite Georgel, Jan-Michael Frahm |
BMVC | 2 |
| 2011 | A geometric solver for calibrated stereo egomotionabstractThis paper introduces a novel geometrical solution for the pose estimation of a stereo camera system as commonly used in robotics, where the camera system balances between coverage and overlap. The proposed approach considers a set of features observed, respectively, in four, three and two views. In contrast to most algebraic solutions our constraints are geometrically meaningful. Initially, we use a four view feature to restrict our translation vector to lie on the surface of a sphere while setting orientation as a function of translation up to a single rotational degree of freedom. Next, we use a three view feature to restrict the translation vector to lie on a circle on the sphere, while completely defining orientation as a function of translation. Finally, we use a two view feature to determine the translation vector lying on the intersection of the circle and one of the generator lines of a doubly ruled quadric. We show how for this final step, the problem can be reduced to the intersection of two coplanar circles. We also analyze the degenerate configurations of the proposed solver and perform an experimental evaluation. Enrique Dunn, Brian Clipp, Jan-Michael Frahm |
ICCV | 1 |
| 2010 | Building Rome on a Cloudless Day
Jan-Michael Frahm, Pierre Fite Georgel, David Gallup, Tim Johnson, Rahul Raguram, Changchang Wu, Yi-Hung Jen, Enrique Dunn, Brian Clipp, Svetlana Lazebnik |
ECCV (4) | 8 |
| 2009 | Next Best View Planning for Active Model ImprovementabstractWe propose a novel approach to determining the Next Best View (NBV) for the task of efficiently building highly accurate 3D models from images. Our proposed method deploys a hierarchical uncertainty driven model refinement process designed to select vantage viewpoints based on the model’s covariance structure and appearance, as well as the camera characteristics. The developed NBV planning system incrementally builds a sensing strategy by sequentially finding the single camera placement, which best reduces an existing model’s 3D uncertainty. The generic nature of our system’s design and internal data representation makes it well suited to be applied to a wide variety of 3D modeling algorithms. It can be used within active computer vision systems as well as for optimized view selection from the set of available views. Experimental results are presented to illustrate the effectiveness and versatility of our approach. Enrique Dunn, Jan-Michael Frahm |
BMVC | 1 |
| 2009 | Developing visual sensing strategies through next best view planningabstractWe propose an approach for acquiring geometric 3D models using cameras mounted on autonomous vehicles and robots. Our method uses structure from motion techniques from computer vision to obtain the geometric structure of the scene. To achieve an efficient goal-driven resource deployment, we develop an incremental approach, which alternates between an accuracy-driven next best view determination and recursive path planning. The next best view is determined by a novel cost function that quantifies the expected contribution of future viewing configurations. A sensing path for robot motion towards the next best view is then achieved by a cost-driven recursive search of intermediate viewing configurations. We discuss some of the properties of our view cost function in the context of an iterative view planning process and present experimental results on a synthetic environment. Enrique Dunn, Jur P. van den Berg, Jan-Michael Frahm |
IROS | 1 |
| 2006 | Parisian camera placement for vision metrology
Enrique Dunn, Gustavo Olague, Evelyne Lutton |
Pattern Recognit. Lett. | 1 |
| 2005 | Pareto optimal camera placement for automated visual inspectionabstractIn this work the problem of camera placement for automated visual inspection is studied under a multi-objective framework. Reconstruction accuracy and operational costs are incorporated into our methodology as separate criteria to optimize. Our approach is based on the initial assumption of conflict among the considered objectives. Hence, the expected results are in the form of Pareto optimal compromise solutions. In order to solve our optimization problem an evolutionary based technique is implemented. Experimental results confirm the conflict among the considered objectives and offer important insights into the relationships between solution quality and process efficiency for high-accurate 3D reconstruction systems. Enrique Dunn, Gustavo Olague |
IROS | 1 |
| 2004 | Pareto optimal sensing strategies for an active vision systemabstractWe present a multiobjective methodology, based on evolutionary computation, for solving the sensor planning problem for an active vision system. The application of different representation schemes, that allow to consider either fixed or variable size camera networks in a single evolutionary process, is studied. Furthermore, a novel representation of the recombination and mutation operators is brought forth. The developed methodology is incorporated into a 3D simulation environment and experimental results shown. Results validate the flexibility and effectiveness of our approach and offer new research alternatives in the field of sensor planning. Enrique Dunn, Gustavo Olague, Evelyne Lutton, Marc Schoenauer |
IEEE Congress on Evolutionary Computation | 1 |
| 2003 | Hybrid Evolutionary Ridge Regression Approach for High-Accurate Corner ExtractionabstractCorner measurement is of main concern within the following tasks: camera calibration, image matching, object tracking, recognition and reconstruction. This paper presents a hybrid evolutionary ridge regression approach for the problem of corner modeling. We search model parameters characterizing L-corner models by means of fitting the model to the image data. As the model fitting relies on an initial parameter estimation, we use a global approach to find the global minimum. Experimental results applied to an L-corner using several levels of noise show the advantages and disadvantages of our evolutionary algorithm compared to down-hill simplex and simulated annealing. Gustavo Olague, Benjamín Hernández, Enrique Dunn |
CVPR (1) | 3 |
| 2001 | Multiple robot task distribution: towards an autonomous photogrammetric systemabstractAutomation of photogrammetric tasks by means of manipulator robots is a complex problem. It involves many planning and controlling aspects that reflect on the overall system performance in terms of precision and efficiency. This paper deals with the problem of task distribution for a multiple manipulator work cell with the goal of obtaining highly accurate object measurements. Task distribution is separated into two independent combinatorial optimization problems: activity assignment and tour planning. These problems are solved simultaneously by an optimization method based on genetic algorithms. This method implements a series of restriction-based heuristics in order to utilize a simple genetic representation similar to random keys. Experiments that validate the effectiveness of our approach are presented. Gustavo Olague, Enrique Dunn |
SMC | 2 |