Kostas Daniilidis

dblp:d/KostasDaniilidis · also Konstantinos Daniilidis · DBLP profile ↗
← Back
208ranked-venue papers
12as first author
51since 2021 · last 2026
0000-0003-0498-0758ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 185 · 10 first-author · 47 since 2021Graphics, computer vision, multimedia, augmented reality and games · 107 · 8 first-author · 26 since 2021Systems, architecture and hardware · 48 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4
YearPublicationVenuePosition
2026 Next Best View Selections for Semantic and Dynamic 3D Gaussian Splatting
abstract
Understanding semantics and dynamics has been crucial for embodied agents in various tasks. Both tasks have much more data redundancy than the static scene understanding task. We formulate the view selection problem as an active learning problem, where the goal is to prioritize frames that provide the greatest information gain for model training. To this end, we propose an active learning algorithm with Fisher Information that quantifies the informativeness of candidate views with respect to both semantic Gaussian parameters and deformation networks. This formulation allows our method to jointly handle semantic reasoning and dynamic scene modeling, providing a principled alternative to heuristic or random strategies. We evaluate our method on large-scale static images and dynamic video datasets by selecting informative frames from multi-camera setups. Experimental results demonstrate that our approach consistently improves rendering quality and semantic segmentation performance, outperforming baseline methods based on random selection and uncertainty-based heuristics.
Yiqian Li, Wen Jiang 0008, Kostas Daniilidis
3DV3
2025 ETAP: Event-based Tracking of Any Point
abstract
Tracking any point (TAP) recently shifted the motion estimation paradigm from focusing on individual salient points with local templates to tracking arbitrary points with global image contexts. However, while research has mostly focused on driving the accuracy of models in nominal settings, addressing scenarios with difficult lighting conditions and high-speed motions remains out of reach due to the limitations of the sensor. This work addresses this challenge with the first event camera-based TAP method. It leverages the high temporal resolution and high dynamic range of event cameras for robust high-speed tracking, and the global contexts in TAP methods to handle asynchronous and sparse event measurements. We further extend the TAP framework to handle event feature variations induced by motion — thereby addressing an open challenge in purely event-based tracking — with a novel feature alignment-loss which ensures the learning of motion-robust features. Our method is trained with data from a new data generation pipeline and systematically ablated across all design decisions. Our method shows strong cross-dataset generalization and performs 136% better on the average Jaccard metric than the baselines. Moreover, on an established feature tracking benchmark, it achieves a 20% improvement over the previous best event-only method and even surpasses the previous best events-and-frames method by 4.1%. Our code is available at https://github.com/tub-rip/ETAP.
Friedhelm Hamann, Daniel Gehrig, Filbert Febryanto, Kostas Daniilidis, Guillermo Gallego 0002
CVPR4
2025 MoSca: Dynamic Gaussian Fusion from Casual Videos via 4D Motion Scaffolds
abstract
We introduce 4D Motion Scaffolds (MoSca), a modern 4D reconstruction system designed to reconstruct and synthesize novel views of dynamic scenes from monocular videos captured casually in the wild. To address such a challenging and ill-posed inverse problem, we leverage prior knowledge from foundational vision models and lift the video data to a novel Motion Scaffold (MoSca) representation, which compactly and smoothly encodes the underlying motions/deformations. The scene geometry and appearance are then disentangled from the deformation field and are encoded by globally fusing the Gaussians anchored onto the MoSca and optimized via Gaussian Splatting. Additionally, camera focal length and poses can be solved using bundle adjustment without the need of any other pose estimation tools. Experiments demonstrate state-of-the-art performance on dynamic rendering benchmarks and its effectiveness on real videos.
Jiahui Lei, Yijia Weng, Adam W. Harley, Leonidas J. Guibas, Kostas Daniilidis
CVPR5
2025 PromptHMR: Promptable Human Mesh Recovery
abstract
Human pose and shape (HPS) estimation presents challenges in diverse scenarios such as crowded scenes, person-person interactions, and single-view reconstruction. Existing approaches lack mechanisms to incorporate auxiliary "side information" that could enhance reconstruction accuracy in such challenging scenarios. Furthermore, the most accurate methods rely on cropped person detections and cannot exploit scene context while methods that process the whole image often fail to detect people and are less accurate than methods that use crops. While recent language-based methods explore HPS reasoning through large language or vision-language models, their metric accuracy is well below the state of the art. In contrast, we present PromptHMR, a transformer-based promptable method that reformulates HPS estimation through spatial and semantic prompts. Our method processes full images to maintain scene context and accepts multiple input modalities: spatial prompts like bounding boxes and masks, and semantic prompts like language descriptions or interaction labels. PromptHMR demonstrates robust performance across challenging scenarios: estimating people from bounding boxes as small as faces in crowded scenes, improving body shape estimation through language descriptions, modeling person-person interactions, and producing temporally coherent motions in videos. Experiments on benchmarks show that PromptHMR achieves state-of-the-art performance while offering flexible prompt-based control over the HPS estimation process.
Yufu Wang, Yu Sun 0030, Priyanka Patel, Kostas Daniilidis, Michael J. Black, Muhammed Kocabas
CVPR4
2025 Multimodal LLM Guided Exploration and Active Mapping Using Fisher Information
Wen Jiang 0008, Boshu Lei, Katrina Ashton, Kostas Daniilidis
ICCV4
2025 DIMO: Diverse 3D Motion Generation for Arbitrary Objects
Linzhan Mou, Jiahui Lei, Chen Wang 0049, Lingjie Liu, Kostas Daniilidis
ICCV5
2025 Continuous-Time Human Motion Field from Event Cameras
Ziyun Wang 0001, Ruijun Zhang, Yufu Wang, Kostas Daniilidis
ICCV5
2025 EqNIO: Subequivariant Neural Inertial Odometry
abstract
Neural network-based odometry using accelerometer and gyroscope readings from a single IMU can achieve robust, and low-drift localization capabilities, through the use of _neural displacement priors (NDPs)_. These priors learn to produce denoised displacement measurements but need to ignore data variations due to specific IMU mount orientation and motion directions, hindering generalization. This work introduces EqNIO, which addresses this challenge with _canonical displacement priors_, i.e., priors that are invariant to the orientation of the gravity-aligned frame in which the IMU data is expressed. We train such priors on IMU measurements, that are mapped into a learnable canonical frame, which is uniquely defined via three axes: the first is gravity, making the frame gravity aligned, while the second and third are predicted from IMU data. The outputs (displacement and covariance) are mapped back to the original gravity-aligned frame. To maximize generalization, we find that these learnable frames must transform equivariantly with global gravity-preserving roto-reflections from the subgroup $O_g(3)\subset O(3)$, acting on the trajectory, rendering the NDP $O(3)$-_subequivariant_. We tailor specific linear, convolutional, and non-linear layers that commute with the actions of the group. Moreover, we introduce a bijective decomposition of angular rates into vectors that transform similarly to accelerations, allowing us to leverage both measurement types. Natively, angular rates would need to be inverted upon reflection, unlike acceleration, which hinders their joint processing. We highlight EqNIO's flexibility and generalization capabilities by applying it to both filter-based (TLIO), and end-to-end (RONIN) architectures, and outperforming existing methods that use _soft equivariance from auxiliary losses or data augmentation on various datasets. We believe this work paves the way for low-drift and generalizable neural inertial odometry on edge devices. The project details and code can be found at [https://github.com/RoyinaJayanth/EqNIO](https://github.com/RoyinaJayanth/EqNIO).
Royina Karegoudra Jayanth, Yinshuang Xu, Ziyun Wang 0001, Evangelos Chatzipantazis, Kostas Daniilidis, Daniel Gehrig
ICLR5
2025 Next Best Sense: Guiding Vision and Touch with FisherRF for 3D Gaussian Splatting
abstract
We propose a framework for active next best view and touch selection for robotic manipulators using 3D Gaussian Splatting (3DGS). 3DGS is emerging as a useful explicit 3D scene representation for robotics, as it has the ability to represent scenes in a both photorealistic and geometrically accurate manner. However, in real-world, online robotic scenes where the number of views is limited given efficiency requirements, random view selection for 3DGS becomes impractical as views are often overlapping and redundant. We address this issue by proposing an end-to-end online training and active view selection pipeline, which enhances the performance of 3DGS in few-view robotics settings. We first elevate the performance of few-shot 3DGS with a novel semantic depth alignment method using Segment Anything Model 2 (SAM2) that we supplement with Pearson depth and surface normal loss to improve color and depth reconstruction of real-world scenes. We then extend FisherRF, a next-best-view selection method for 3DGS, to select views and touch poses based on depth uncertainty. We perform online view selection on a real robot system during live 3DGS training. We motivate our improvements to few-shot GS scenes, and extend depth-based FisherRF to them, where we demonstrate both qualitative and quantitative improvements on challenging robot scenes. For more information, please see our project page at arm.stanford.edu/next-best-sense.
Matthew Strong, Boshu Lei, Aiden Swann, Wen Jiang 0008, Kostas Daniilidis, Monroe Kennedy III
ICRA5
2025 A Scalable, Causal, and Energy Efficient Framework for Neural Decoding with Spiking Neural Networks
abstract
Brain-computer interfaces (BCIs) promise to enable vital functions, such as speech and prosthetic control, for individuals with neuromotor impairments. Central to their success are neural decoders, models that map neural activity to intended behavior. Current learning-based decoding approaches fall into two classes: simple, causal models that lack generalization, or complex, non-causal models that generalize and scale offline but struggle in real-time settings. Both face a common challenge, their reliance on power-hungry artificial neural network backbones, which makes integration into real-world, resource-limited systems difficult. Spiking neural networks (SNNs) offer a promising alternative. Because they operate causally (i.e. only on present and past inputs) these models are suitable for real-time use, and their low energy demands make them ideal for battery-constrained environments. To this end, we introduce **Spikachu: a scalable, causal, and energy-efficient neural decoding framework based on SNNs**. Our approach processes binned spikes directly by projecting them into a shared latent space, where spiking modules, adapted to the timing of the input, extract relevant features; these latent representations are then integrated and decoded to generate behavioral predictions. We evaluate our approach on 113 recording sessions from 6 non-human primates, totaling 43 hours of recordings. Our method outperforms causal baselines when trained on single sessions using between 2.26× and 418.81× less energy. Furthermore, we demonstrate that scaling up training to multiple sessions and subjects improves performance and enables few-shot transfer to unseen sessions, subjects, and tasks. Overall, Spikachu introduces a scalable, online-compatible neural decoding framework based on SNNs, whose performance is competitive relative to state-of-the-art models while consuming orders of magnitude less energy.
Georgios Mentzelopoulos, Ioannis Asmanis, Konrad P. Kording, Eva L. Dyer, Kostas Daniilidis, Flavia Vitale
NeurIPS5
2024 GART: Gaussian Articulated Template Models
abstract
We introduce Gaussian Articulated Template Model (GART), an explicit, efficient, and expressive representation for non-rigid articulated subject capturing and rendering from monocular videos. GART utilizes a mixture of moving 3D Gaussians to explicitly approximate a deformable subject's geometry and appearance. It takes advantage of a categorical template model prior (SMPL, SMAL, etc.) with learnable forward skinning while further generalizing to more complex non-rigid deformations with novel latent bones. GART can be reconstructed via differentiable rendering from monocular videos in seconds or minutes and rendered in novel poses faster than 150fps.
Jiahui Lei, Yufu Wang, Georgios Pavlakos, Lingjie Liu, Kostas Daniilidis
CVPR5
2024 Motion-Prior Contrast Maximization for Dense Continuous-Time Motion Estimation
Friedhelm Hamann, Ziyun Wang 0001, Ioannis Asmanis, Kenneth Chaney, Guillermo Gallego 0002, Kostas Daniilidis
ECCV (3)6
2024 FisherRF: Active View Selection and Mapping with Radiance Fields Using Fisher Information
Wen Jiang 0008, Boshu Lei, Kostas Daniilidis
ECCV (13)3
2024 DynMF: Neural Motion Factorization for Real-Time Dynamic View Synthesis with 3D Gaussian Splatting
Agelos Kratimenos, Jiahui Lei, Kostas Daniilidis
ECCV (79)3
2024 Track Everything Everywhere Fast and Robustly
Yunzhou Song, Jiahui Lei, Ziyun Wang 0001, Lingjie Liu, Kostas Daniilidis
ECCV (3)5
2024 Un-EVIMO: Unsupervised Event-Based Independent Motion Segmentation
Ziyun Wang 0001, Kostas Daniilidis
ECCV (16)3
2024 TRAM: Global Trajectory and Motion of 3D Humans from in-the-Wild Videos
Yufu Wang, Ziyun Wang 0001, Lingjie Liu, Kostas Daniilidis
ECCV (11)4
2024 Scalable Networked Feature Selection with Randomized Algorithm for Robot Navigation
abstract
We address the problem of sparse selection of visual features for localizing a team of robots navigating in an unknown environment, where robots can exchange relative position measurements with neighbors. We select a set of the most informative features by anticipating their importance in robots localization by simulating trajectories of robots over a prediction horizon. Through theoretical proofs, we establish a crucial connection between graph Laplacian and the importance of features. We leverage a scalable randomized algorithm for sparse sums of positive semidefinite matrices to efficiently select a set of the most informative features.
Vivek Pandey, Arash Amini, Guangyi Liu 0004, Ufuk Topcu, Qiyu Sun, Kostas Daniilidis, Nader Motee
IROS6
2024 Uncertainty-Aware Deployment of Pre-trained Language-Conditioned Imitation Learning Policies
abstract
Large-scale robotic policies trained on data from diverse tasks and robotic platforms hold great promise for enabling general-purpose robots; however, reliable generalization to new environment conditions remains a major challenge. Toward addressing this challenge, we propose a novel approach for uncertainty-aware deployment of pre-trained language-conditioned imitation learning agents. Specifically, we use temperature scaling to calibrate these models and exploit the calibrated model to make uncertainty-aware decisions by aggregating the local information of candidate actions. We implement our approach in simulation using three such pre-trained models, and showcase its potential to significantly enhance task completion rates. The accompanying code is accessible at the link: https://github.com/BobWu1998/uncertainty_quant_all.git
Bruce D. Lee, Kostas Daniilidis, Bernadette Bucher, Nikolai Matni
IROS3
2024 Neural decoding from stereotactic EEG: accounting for electrode variability across subjects
abstract
Deep learning based neural decoding from stereotactic electroencephalography (sEEG) would likely benefit from scaling up both dataset and model size. To achieve this, combining data across multiple subjects is crucial. However, in sEEG cohorts, each subject has a variable number of electrodes placed at distinct locations in their brain, solely based on clinical needs. Such heterogeneity in electrode number/placement poses a significant challenge for data integration, since there is no clear correspondence of the neural activity recorded at distinct sites between individuals. Here we introduce seegnificant: a training framework and architecture that can be used to decode behavior across subjects using sEEG data. We tokenize the neural activity within electrodes using convolutions and extract long-term temporal dependencies between tokens using self-attention in the time dimension. The 3D location of each electrode is then mixed with the tokens, followed by another self-attention in the electrode dimension to extract effective spatiotemporal neural representations. Subject-specific heads are then used for downstream decoding tasks. Using this approach, we construct a multi-subject model trained on the combined data from 21 subjects performing a behavioral task. We demonstrate that our model is able to decode the trial-wise response time of the subjects during the behavioral task solely from neural data. We also show that the neural representations learned by pretraining our model across individuals can be transferred in a few-shot manner to new subjects. This work introduces a scalable approach towards sEEG data integration for multi-subject model training, paving the way for cross-subject generalization for sEEG decoding.
Georgios Mentzelopoulos, Evangelos Chatzipantazis, Ashwin G. Ramayya, Michelle J. Hedlund, Vivek P. Buch, Kostas Daniilidis, Konrad P. Kording, Flavia Vitale
NeurIPS6
2024 Improving Equivariant Model Training via Constraint Relaxation
abstract
Equivariant neural networks have been widely used in a variety of applications due to their ability to generalize well in tasks where the underlying data symmetries are known. Despite their successes, such networks can be difficult to optimize and require careful hyperparameter tuning to train successfully. In this work, we propose a novel framework for improving the optimization of such models by relaxing the hard equivariance constraint during training: We relax the equivariance constraint of the network's intermediate layers by introducing an additional non-equivariant term that we progressively constrain until we arrive at an equivariant solution. By controlling the magnitude of the activation of the additional relaxation term, we allow the model to optimize over a larger hypothesis space containing approximate equivariant networks and converge back to an equivariant solution at the end of training. We provide experimental results on different state-of-the-art network architectures, demonstrating how this training framework can result in equivariant models with improved generalization performance. Our code is available at https://github.com/StefanosPert/Equivariant_Optimization_CR
Stefanos Pertigkiozoglou, Evangelos Chatzipantazis, Shubhendu Trivedi, Kostas Daniilidis
NeurIPS4
2024 $SE(3)$ Equivariant Ray Embeddings for Implicit Multi-View Depth Estimation
abstract
Incorporating inductive bias by embedding geometric entities (such as rays) as input has proven successful in multi-view learning. However, the methods adopting this technique typically lack equivariance, which is crucial for effective 3D learning. Equivariance serves as a valuable inductive prior, aiding in the generation of robust multi-view features for 3D scene understanding. In this paper, we explore the application of equivariant multi-view learning to depth estimation, not only recognizing its significance for computer vision and robotics but also addressing the limitations of previous research. Most prior studies have either overlooked equivariance in this setting or achieved only approximate equivariance through data augmentation, which often leads to inconsistencies across different reference frames. To address this issue, we propose to embed $SE(3)$ equivariance into the Perceiver IO architecture. We employ Spherical Harmonics for positional encoding to ensure 3D rotation equivariance, and develop a specialized equivariant encoder and decoder within the Perceiver IO architecture. To validate our model, we applied it to the task of stereo depth estimation, achieving state of the art results on real-world datasets without explicit geometric constraints or extensive data augmentation.
Yinshuang Xu, Dian Chen 0005, Katherine Liu, Sergey Zakharov, Rares Ambrus, Kostas Daniilidis, Vitor Campagnolo Guizilini
NeurIPS6
2024 EvDNeRF: Reconstructing Event Data with Dynamic Neural Radiance Fields
abstract
We present EvDNeRF, a pipeline for generating event data and training an event-based dynamic NeRF, for the purpose of faithfully reconstructing eventstreams on scenes with rigid and non-rigid deformations that may be too fast to capture with a standard camera. Event cameras register asynchronous per-pixel brightness changes at MHz rates with high dynamic range, making them ideal for observing fast motion with almost no motion blur. Neural radiance fields (NeRFs) offer visual-quality geometric-based learnable rendering, but prior work with events has only considered reconstruction of static scenes. Our EvDNeRF can predict eventstreams of dynamic scenes from a static or moving viewpoint between any desired timestamps, thereby allowing it to be used as an event-based simulator for a given scene. We show that by training on varied batch sizes of events, we can improve test-time predictions of events at fine time resolutions, outperforming baselines that pair standard dynamic NeRFs with event generators. We release our simulated and real datasets, as well as code for multi-view event-based data generation and the training and evaluation of EvDNeRF models1.
Anish Bhattacharya, Ratnesh Madaan, Fernando Cladera Ojeda, Sai Vemprala, Rogerio Bonatti, Kostas Daniilidis, Ashish Kapoor, Vijay Kumar 0001, Nikolai Matni, Jayesh K. Gupta
WACV6
2023 EFEM: Equivariant Neural Field Expectation Maximization for 3D Object Segmentation Without Scene Supervision
abstract
We introduce Equivariant Neural Field Expectation Maximization (EFEM), a simple, effective, and robust geometric algorithm that can segment objects in 3D scenes without annotations or training on scenes. We achieve such unsupervised segmentation by exploiting single object shape priors. We make two novel steps in that direction. First, we introduce equivariant shape representations to this problem to eliminate the complexity induced by the variation in object configuration. Second, we propose a novel EM algorithm that can iteratively refine segmentation masks using the equivariant shape prior. We collect a novel real dataset Chairs and Mugs that contains various object configurations and novel scenes in order to verify the effectiveness and robustness of our method. Experimental results demonstrate that our method achieves consistent and robust performance across different scenes where the (weakly) supervised methods may fail. Code and data available at https://www.cis.upenn.edu/~leijh/projects/EFEM
Jiahui Lei, Congyue Deng, Karl Schmeckpeper, Leonidas J. Guibas, Kostas Daniilidis
CVPR5
2023 ReFit: Recurrent Fitting Network for 3D Human Recovery
abstract
We present Recurrent Fitting (ReFit), a neural network architecture for single-image, parametric 3D human reconstruction. ReFit learns a feedback-update loop that mirrors the strategy of solving an inverse problem through optimization. At each iterative step, it reprojects keypoints from the human model to feature maps to query feedback, and uses a recurrent-based updater to adjust the model to fit the image better. Because ReFit encodes strong knowledge of the inverse problem, it is faster to train than previous regression models. At the same time, ReFit improves state-of-the-art performance on standard benchmarks. Moreover, ReFit applies to other optimization settings, such as multi-view fitting and single-view shape fitting. Project website: https://yufu-wang.github.io/refit_humans/
Yufu Wang, Kostas Daniilidis
ICCV2
2023 NeuS2: Fast Learning of Neural Implicit Surfaces for Multi-view Reconstruction
abstract
Recent methods for neural surface representation and rendering, for example NeuS [59], have demonstrated the remarkably high-quality reconstruction of static scenes. However, the training of NeuS takes an extremely long time (8 hours), which makes it almost impossible to apply them to dynamic scenes with thousands of frames. We propose a fast neural surface reconstruction approach, called NeuS2, which achieves two orders of magnitude improvement in terms of acceleration without compromising reconstruction quality. To accelerate the training process, we parameterize a neural surface representation by multi-resolution hash encodings and present a novel lightweight calculation of second-order derivatives tailored to our networks to leverage CUDA parallelism, achieving a factor two speed up. To further stabilize and expedite training, a progressive learning strategy is proposed to optimize multi-resolution hash encodings from coarse to fine. We extend our method for fast training of dynamic scenes, with a proposed incremental training strategy and a novel global transformation prediction component, which allow our method to handle challenging long sequences with large movements and deformations. Our experiments on various datasets demonstrate that NeuS2 significantly outperforms the state-of-the-arts in both surface reconstruction accuracy and training speed for both static and dynamic scenes. The code is available at our website: https://vcai.mpi-inf.mpg.de/projects/NeuS2/.
Yiming Wang 0009, Qin Han, Marc Habermann, Kostas Daniilidis, Christian Theobalt, Lingjie Liu
ICCV4
2023 SE(3)-Equivariant Attention Networks for Shape Reconstruction in Function Space
Evangelos Chatzipantazis, Stefanos Pertigkiozoglou, Edgar Dobriban, Kostas Daniilidis
ICLR4
2023 Banana: Banach Fixed-Point Network for Pointcloud Segmentation with Inter-Part Equivariance
abstract
Equivariance has gained strong interest as a desirable network property that inherently ensures robust generalization. However, when dealing with complex systems such as articulated objects or multi-object scenes, effectively capturing inter-part transformations poses a challenge, as it becomes entangled with the overall structure and local transformations. The interdependence of part assignment and per-part group action necessitates a novel equivariance formulation that allows for their co-evolution. In this paper, we present Banana, a Banach fixed-point network for equivariant segmentation with inter-part equivariance by construction. Our key insight is to iteratively solve a fixed-point problem, where point-part assignment labels and per-part SE(3)-equivariance co-evolve simultaneously. We provide theoretical derivations of both per-step equivariance and global convergence, which induces an equivariant final convergent state. Our formulation naturally provides a strict definition of inter-part equivariance that generalizes to unseen inter-part configurations. Through experiments conducted on both articulated objects and multi-object scans, we demonstrate the efficacy of our approach in achieving strong generalization under inter-part transformations, even when confronted with substantial changes in pointcloud geometry and topology.
Congyue Deng, Jiahui Lei, Bokui Shen, Kostas Daniilidis, Leonidas J. Guibas
NeurIPS4
2023 NAP: Neural 3D Articulated Object Prior
abstract
We propose Neural 3D Articulated object Prior (NAP), the first 3D deep generative model to synthesize 3D articulated object models. Despite the extensive research on generating 3D static objects, compositions, or scenes, there are hardly any approaches on capturing the distribution of articulated objects, a common object category for human and robot interaction. To generate articulated objects, we first design a novel articulation tree/graph parameterization and then apply a diffusion-denoising probabilistic model over this representation where articulated objects can be generated via denoising from random complete graphs. In order to capture both the geometry and the motion structure whose distribution will affect each other, we design a graph denoising network for learning the reverse diffusion process. We propose a novel distance that adapts widely used 3D generation metrics to our novel task to evaluate generation quality. Experiments demonstrate our high performance in articulated object generation as well as its applications on conditioned generation, including Part2Motion, PartNet-Imagination, Motion2Part, and GAPart2Object.
Jiahui Lei, Congyue Deng, Bokui Shen, Leonidas J. Guibas, Kostas Daniilidis
NeurIPS5
2023 SE(3) Equivariant Convolution and Transformer in Ray Space
abstract
3D reconstruction and novel view rendering can greatly benefit from geometric priors when the input views are not sufficient in terms of coverage and inter-view baselines. Deep learning of geometric priors from 2D images requires each image to be represented in a $2D$ canonical frame and the prior to be learned in a given or learned $3D$ canonical frame. In this paper, given only the relative poses of the cameras, we show how to learn priors from multiple views equivariant to coordinate frame transformations by proposing an $SE(3)$-equivariant convolution and transformer in the space of rays in 3D. We model the ray space as a homogeneous space of $SE(3)$ and introduce the $SE(3)$-equivariant convolution in ray space. Depending on the output domain of the convolution, we present convolution-based $SE(3)$-equivariant maps from ray space to ray space and to $\mathbb{R}^3$. Our mathematical framework allows us to go beyond convolution to $SE(3)$-equivariant attention in the ray space. We showcase how to tailor and adapt the equivariant convolution and transformer in the tasks of equivariant $3D$ reconstruction and equivariant neural rendering from multiple views. We demonstrate $SE(3)$-equivariance by obtaining robust results in roto-translated datasets without performing transformation augmentation.
Yinshuang Xu, Jiahui Lei, Kostas Daniilidis
NeurIPS3
2023 Multi-view Tracking, Re-ID, and Social Network Analysis of a Flock of Visually Similar Birds in an Outdoor Aviary
Shiting Xiao, Yufu Wang, Ammon Perkes, Bernd Pfrommer, Marc F. Schmidt, Kostas Daniilidis, Marc Badger
Int. J. Comput. Vis.6
2022 Cross-modal Map Learning for Vision and Language Navigation
abstract
We consider the problem of Vision-and-Language Navigation (VLN). The majority of current methods for VLN are trained end-to-end using either unstructured memory such as LSTM, or using cross-modal attention over the egocentric observations of the agent. In contrast to other works, our key insight is that the association between language and vision is stronger when it occurs in explicit spatial representations. In this work, we propose a cross-modal map learning model for vision-and-language navigation that first learns to predict the top-down semantics on an egocentric map for both observed and unobserved regions, and then predicts a path towards the goal as a set of way-points. In both cases, the prediction is informed by the language through cross-modal attention mechanisms. We experimentally test the basic hypothesis that language-driven navigation can be solved given a map, and then show competitive results on the full VLN-CE benchmark.
Georgios Georgakis, Karl Schmeckpeper, Karan Wanchoo, Soham Dan, Eleni Miltsakaki, Dan Roth 0001, Kostas Daniilidis
CVPR7
2022 CaDeX: Learning Canonical Deformation Coordinate Space for Dynamic Surface Representation via Neural Homeomorphism
abstract
While neural representations for static 3D shapes are widely studied, representations for deformable surfaces are limited to be template-dependent or to lack efficiency. We introduce Canonical Deformation Coordinate Space (CaDeX), a unified representation of both shape and nonrigid motion. Our key insight is the factorization of the deformation between frames by continuous bijective canonical maps (homeomorphisms) and their inverses that go through a learned canonical shape. Our novel deformation representation and its implementation are simple, efficient, and guarantee cycle consistency, topology preservation, and, if needed, volume conservation. Our modelling of the learned canonical shapes provides a flexible and stable space for shape prior learning. We demonstrate state-of-the-art performance in modelling a wide range of deformable geometries: human bodies, animal bodies, and articulated objects.11https://www.cis.upenn.edu/-leijh/projects/cadex
Jiahui Lei, Kostas Daniilidis
CVPR2
2022 EvAC3D: From Event-Based Apparent Contours to 3D Models via Continuous Visual Hulls
Ziyun Wang 0001, Kenneth Chaney, Kostas Daniilidis
ECCV (7)3
2022 Learning to Map for Active Semantic Goal Navigation
Georgios Georgakis, Bernadette Bucher, Karl Schmeckpeper, Kostas Daniilidis
ICLR5
2022 Unified Fourier-based Kernel and Nonlinearity Design for Equivariant Networks on Homogeneous Spaces
abstract
We introduce a unified framework for group equivariant networks on homogeneous spaces derived from a Fourier perspective. We consider tensor-valued feature fields, before and after a convolutional layer. We present a unified derivation of kernels via the Fourier domain by leveraging the sparsity of Fourier coefficients of the lifted feature fields. The sparsity emerges when the stabilizer subgroup of the homogeneous space is a compact Lie group. We further introduce a nonlinear activation, via an elementwise nonlinearity on the regular representation after lifting and projecting back to the field through an equivariant convolution. We show that other methods treating features as the Fourier coefficients in the stabilizer subgroup are special cases of our activation. Experiments on $SO(3)$ and $SE(3)$ show state-of-the-art performance in spherical vector field regression, point cloud classification, and molecular completion.
Yinshuang Xu, Jiahui Lei, Edgar Dobriban, Kostas Daniilidis
ICML4
2022 Level Set Mesher: Single-image to 3D reconstruction by following the level sets of the signed distance function
abstract
We present Level Set Mesher, a single-image to 3D reconstruction strategy to deform an initial spherical triangular mesh into the 3D geometry of a target shape. Level Set Mesher offsets each vertex in a discrete number of steps following a learned velocity vector field V modeled as a Graph Attention Network. At each step k, we constrain the deformed vertices to lie on the lklevel set of the target shape’s signed distance function to guide the deformation process. We show that our approach accurately estimates the surface’s normal of the predicted shapes and reduces mesh’s artifacts in the final prediction. We compare our approach with state-of-the-art single-image to 3D reconstruction methods and show improvements in accuracy predictions, resulting in better quality and manifold meshes.
Diego Alberto Patiño Cortes, Carlos Esteves, Kostas Daniilidis
ICPR3
2022 Uncertainty-driven Planner for Exploration and Navigation
abstract
We consider the problems of exploration and pointgoal navigation in previously unseen environments, where the spatial complexity of indoor scenes and partial observability constitute these tasks challenging. We argue that learning occupancy priors over indoor maps provides significant advantages towards addressing these problems. To this end, we present a novel planning framework that first learns to generate occupancy maps beyond the field-of-view of the agent, and second leverages the model uncertainty over the generated areas to formulate path selection policies for each task of interest. For pointgoal navigation the policy chooses paths with an upper confidence bound policy for efficient and traversable paths, while for exploration the policy maximizes model uncertainty over candidate paths. We perform experiments in the visually realistic environments of Matterport3D using the Habitat simulator and demonstrate: 1) Improved results on exploration and map quality metrics over competitive methods, and 2) The effectiveness of our planning module when paired with the state-of-the-art DD-PPO method for the point-goal navigation task.
Georgios Georgakis, Bernadette Bucher, Anton Arapin, Karl Schmeckpeper, Nikolai Matni, Kostas Daniilidis
ICRA6
2022 Single-camera 3D head fitting for mixed reality clinical applications
Tejas Mane, Aylar Bayramova, Kostas Daniilidis, Philippos Mordohai, Elena Bernardis
Comput. Vis. Image Underst.3
2022 Event-Based Vision: A Survey
abstract
Event cameras are bio-inspired sensors that differ from conventional frame cameras: Instead of capturing images at a fixed rate, they asynchronously measure per-pixel brightness changes, and output a stream of events that encode the time, location and sign of the brightness changes. Event cameras offer attractive properties compared to traditional cameras: high temporal resolution (in the order of μs), very high dynamic range (140 dB versus 60 dB), low power consumption, and high pixel bandwidth (on the order of kHz) resulting in reduced motion blur. Hence, event cameras have a large potential for robotics and computer vision in challenging scenarios for traditional cameras, such as low-latency, high speed, and high dynamic range. However, novel methods are required to process the unconventional output of these sensors in order to unlock their potential. This paper provides a comprehensive overview of the emerging field of event-based vision, with a focus on the applications and the algorithms developed to unlock the outstanding properties of event cameras. We present event cameras from their working principle, the actual sensors that are available and the tasks that they have been used for, from low-level vision (feature detection and tracking, optic flow, etc.) to high-level vision (reconstruction, segmentation, recognition). We also discuss the techniques developed to process events, including learning-based techniques, as well as specialized processors for these novel sensors, such as spiking neural networks. Additionally, we highlight the challenges that remain to be tackled and the opportunities that lie ahead in the search for a more efficient, bio-inspired way for machines to perceive and interact with the world.
Guillermo Gallego 0002, Tobi Delbruck, Garrick Orchard, Chiara Bartolozzi, Brian Taba, Andrea Censi, Stefan Leutenegger, Andrew J. Davison, Jörg Conradt, Kostas Daniilidis, Davide Scaramuzza 0001
IEEE Trans. Pattern Anal. Mach. Intell.10
2021 Birds of a Feather: Capturing Avian Shape Models From Images
abstract
Animals are diverse in shape, but building a deformable shape model for a new species is not always possible due to the lack of 3D data. We present a method to capture new species using an articulated template and images of that species. In this work, we focus mainly on birds. Although birds represent almost twice the number of species as mammals, no accurate shape model is available. To capture a novel species, we first fit the articulated template to each training sample. By disentangling pose and shape, we learn a shape space that captures variation both among species and within each species from image evidence. We learn models of multiple species from the CUB dataset, and contribute new species-specific and multi-species shape models that are useful for downstream reconstruction tasks. Using a low-dimensional embedding, we show that our learned 3D shape space better reflects the phylogenetic relationships among birds than learned perceptual features.
Yufu Wang, Nikos Kolotouros, Kostas Daniilidis, Marc Badger
CVPR3
2021 EventGAN: Leveraging Large Scale Image Datasets for Event Cameras
abstract
Event cameras provide a number of benefits over traditional cameras, such as the ability to track incredibly fast motions, high dynamic range, and low power consumption. However, their application into computer vision problems, many of which are primarily dominated by deep learning solutions, has been limited by the lack of labeled training data for events. In this work, we propose a method which leverages the existing labeled data for images by simulating events from a pair of temporal image frames, using a convolutional neural network. We train this network on pairs of images and events, using an adversarial discriminator loss and a pair of cycle consistency losses. The cycle consistency losses utilize a pair of pre-trained self-supervised networks which perform optical flow estimation and image reconstruction from events, and constrain our network to generate events which result in accurate outputs from both of these networks. Trained fully end to end, our network learns a generative model for events from images without the need for accurate modeling of the motion in the scene, exhibited by modeling based methods, while also implicitly modeling event noise. Using this simulator, we train a pair of downstream networks on object detection and 2D human pose estimation from events, using simulated data from large scale image datasets, and demonstrate the networks' abilities to generalize to datasets with real events. The code and dataset in this paper are available here: https://github.com/alexzzhu/EventGAN.
Alex Zihao Zhu, Ziyun Wang 0001, Kaung Khant, Kostas Daniilidis
ICCP4
2021 Probabilistic Modeling for Human Mesh Recovery
abstract
This paper focuses on the problem of 3D human reconstruction from 2D evidence. Although this is an inherently ambiguous problem, the majority of recent works avoid the uncertainty modeling and typically regress a single estimate for a given input. In contrast to that, in this work, we propose to embrace the reconstruction ambiguity and we recast the problem as learning a mapping from the input to a distribution of plausible 3D poses. Our approach is based on the normalizing flows model and offers a series of advantages. For conventional applications, where a single 3D estimate is required, our formulation allows for efficient mode computation. Using the mode leads to performance that is comparable with the state of the art among deterministic unimodal regression models. Simultaneously, since we have access to the likelihood of each sample, we demonstrate that our model is useful in a series of downstream tasks, where we leverage the probabilistic nature of the prediction as a tool for more accurate estimation. These tasks include reconstruction from multiple uncalibrated views, as well as human model fitting, where our model acts as a powerful image-based prior for mesh recovery. Our results validate the importance of probabilistic modeling, and indicate state-of-the-art performance across a variety of settings. Code and models are available at: https://www.seas.upenn.edu/~nkolot/projects/prohmr.
Nikos Kolotouros, Georgios Pavlakos, Dinesh Jayaraman, Kostas Daniilidis
ICCV4
2021 Simple and Effective VAE Training with Calibrated Decoders
abstract
Variational autoencoders (VAEs) provide an effective and simple method for modeling complex distributions. However, training VAEs often requires considerable hyperparameter tuning to determine the optimal amount of information retained by the latent variable. We study the impact of calibrated decoders, which learn the uncertainty of the decoding distribution and can determine this amount of information automatically, on the VAE performance. While many methods for learning calibrated decoders have been proposed, many of the recent papers that employ VAEs rely on heuristic hyperparameters and ad-hoc modifications instead. We perform the first comprehensive comparative analysis of calibrated decoder and provide recommendations for simple and effective VAE training. Our analysis covers a range of datasets and several single-image and sequential VAE models. We further propose a simple but novel modification to the commonly used Gaussian decoder, which computes the prediction variance analytically. We observe empirically that using heuristic modifications is not necessary with our method.
Oleh Rybkin, Kostas Daniilidis, Sergey Levine
ICML2
2021 Model-Based Reinforcement Learning via Latent-Space Collocation
abstract
The ability to plan into the future while utilizing only raw high-dimensional observations, such as images, can provide autonomous agents with broad and general capabilities. However, realistic tasks require performing temporally extended reasoning, and cannot be solved with only myopic, short-sighted planning. Recent work in model-based reinforcement learning (RL) has shown impressive results on tasks that require only short-horizon reasoning. In this work, we study how the long-horizon planning abilities can be improved with an algorithm that optimizes over sequences of states, rather than actions, which allows better credit assignment. To achieve this, we draw on the idea of collocation and adapt it to the image-based setting by leveraging probabilistic latent variable models, resulting in an algorithm that optimizes trajectories over latent variables. Our latent collocation method (LatCo) provides a general and effective visual planning approach, and significantly outperforms prior model-based approaches on challenging visual control tasks with sparse rewards and long-term goals. See the videos on the supplementary website \url{https://sites.google.com/view/latco-mbrl/.}
Oleh Rybkin, Chuning Zhu, Anusha Nagabandi, Kostas Daniilidis, Igor Mordatch, Sergey Levine
ICML4
2021 Object-centric Video Prediction without Annotation
abstract
In order to interact with the world, agents must be able to predict the results of the world’s dynamics. A natural approach to learn about these dynamics is through video prediction, as cameras are ubiquitous and powerful sensors. Direct pixel-to-pixel video prediction is difficult, does not take advantage of known priors, and does not provide an easy interface to utilize the learned dynamics. Object-centric video prediction offers a solution to these problems by taking advantage of the simple prior that the world is made of objects and by providing a more natural interface for control. However, existing object-centric video prediction pipelines require dense object annotations in training video sequences. In this work, we present Object-centric Prediction without Annotation (OPA), an object-centric video prediction method that takes advantage of priors from powerful computer vision models. We validate our method on a dataset comprised of video sequences of stacked objects falling, and demonstrate how to adapt a perception model in an environment through end-to-end video prediction training.
Karl Schmeckpeper, Georgios Georgakis, Kostas Daniilidis
ICRA3
2021 Deformable Linear Object Prediction Using Locally Linear Latent Dynamics
abstract
We propose a framework for deformable linear object prediction. Prediction of deformable objects (e.g., rope) is challenging due to their non-linear dynamics and infinite-dimensional configuration spaces. By mapping the dynamics from a non-linear space to a linear space, we can use the good properties of linear dynamics for easier learning and more efficient prediction. We learn a locally linear, action-conditioned dynamics model that can be used to predict future latent states. Then, we decode the predicted latent state into the predicted state. We also apply a sampling-based optimization algorithm to select the optimal control action. We empirically demonstrate that our approach can predict the rope state accurately up to ten steps into the future and that our algorithm can find the optimal action given an initial state and a goal state.
Karl Schmeckpeper, Pratik Chaudhari, Kostas Daniilidis
ICRA4
2021 An Adversarial Objective for Scalable Exploration
abstract
Collecting new experience is costly in many robotic tasks, so determining how to efficiently explore in a new environment to learn as much as possible in as few trials as possible is an important problem for robotics. In this paper, we propose a method for exploring for the purpose of learning a dynamics model. Our key idea is to minimize a score given by a discriminator network as an objective for a planner which chooses actions. This discriminator is optimized jointly with a prediction model and enables our active learning approach to sample sequences of observations and actions which result in predictions considered the least realistic by the discriminator. Comparable existing exploration methods cannot operate in many prediction-planning pipelines used in robotic learning without hardware modifications to standard robotics platforms in order to accommodate their large compute requirements, so the primary contribution of our adversarial exploration method is scalability. We demonstrate progressively increased performance of our adversarial exploration approach compared to leading model-based exploration strategies as compute is restricted in simulated environments. We further demonstrate the ability of our adversarial method to scale to a robotic manipulation prediction-planning pipeline where we improve sample efficiency and prediction performance for a domain transfer problem.
Bernadette Bucher, Karl Schmeckpeper, Nikolai Matni, Kostas Daniilidis
IROS4
2021 Self-Supervised Optical Flow with Spiking Neural Networks and Event Based Cameras
abstract
Optical flow can be leveraged in robotic systems for obstacle detection where low latency solutions are critical in highly dynamic settings. While event-based cameras have changed the dominant paradigm of sending by encoding stimuli into spike trails, offering low bandwidth and latency, events are still processed with traditional convolutional networks in GPUs defeating, thus, the promise of efficient low capacity low power processing that inspired the design of event sensors. In this work, we introduce a shallow spiking neural network for the computation of optical flow consisting of Leaky Integrate and Fire neurons.Optical flow is predicted as the synthesis of motion orientation selective channels. Learning is accomplished by Back-propapagation Through Time. We present promising results on events recorded in real "in the wild" scenes that has the capability to use only a small fraction of the energy consumed in CNNs deployed on GPUs.
Kenneth Chaney, Artemis Panagopoulou, Chankyu Lee, Kaushik Roy 0001, Kostas Daniilidis
IROS5
2021 Discovering and Achieving Goals via World Models
abstract
How can artificial agents learn to solve many diverse tasks in complex visual environments without any supervision? We decompose this question into two challenges: discovering new goals and learning to reliably achieve them. Our proposed agent, Latent Explorer Achiever (LEXA), addresses both challenges by learning a world model from image inputs and using it to train an explorer and an achiever policy via imagined rollouts. Unlike prior methods that explore by reaching previously visited states, the explorer plans to discover unseen surprising states through foresight, which are then used as diverse targets for the achiever to practice. After the unsupervised phase, LEXA solves tasks specified as goal images zero-shot without any additional learning. LEXA substantially outperforms previous approaches to unsupervised goal reaching, both on prior benchmarks and on a new challenging benchmark with 40 test tasks spanning across four robotic manipulation and locomotion domains. LEXA further achieves goals that require interacting with multiple objects in sequence. Project page: https://orybkin.github.io/lexa/
Russell Mendonca, Oleh Rybkin, Kostas Daniilidis, Danijar Hafner, Deepak Pathak
NeurIPS3
2021 Towards Statistically Provable Geometric 3D Human Pose Recovery
abstract
Recovering three-dimensional (3D) structures such as object poses from limited two-dimensional (2D) information is an important research problem in computer vision, graphics, and robotics. The estimation of object pose from single images or multiple casual images could be ill-conditioned math problems. There is a popular family of algorithms of geometric sparse representation for 3D pose recovery (GSR-3D) that pretrains an overcomplete dictionary of 3D basis poses $B$, and then matches the detected 2D object pose $Y$ by jointly estimating the transformation $R$, projection $\Pi,$ and combination coefficients $c$, assuming $Y \approx \Pi \sum_i c_i R B_i$. In this paper, we make the first step of analyzing to which extent could we solve this ill-conditioned problem, and of understanding how the recovery error is affected by fundamental factors, e.g., dictionary size, observation noise, and running time. As these factors are implicit in objective functions, we analyze with the help of various sparse regularizers and a multistage optimizer, and prove that the recovery error $\mathcal L(l)$ decays w.r.t. the number of stages $l$ with a high probability, $Prob\left(\mathcal L(l) < \rho^{l-1} \mathcal L(0) + \delta \right) \geq 1- \epsilon$, where the constants $0< \rho <1, 0<\delta, 0<\epsilon \ll 1$ are related to the aforementioned factors. To the best of our knowledge, this is the first theoretical analysis in this line of research. Experiments are conducted to support our improvement upon previous regularization within the same framework. This will further characterize the trade-off between speed and accuracy towards real-time geometric inference in applications.
Jianqiao Wangni, Dahua Lin, Kostas Daniilidis, Jianbo Shi
SIAM J. Imaging Sci.4
2020 Coherent Reconstruction of Multiple Humans From a Single Image
abstract
In this work, we address the problem of multi-person 3D pose estimation from a single image. A typical regression approach in the top-down setting of this problem would first detect all humans and then reconstruct each one of them independently. However, this type of prediction suffers from incoherent results, e.g., interpenetration and inconsistent depth ordering between the people in the scene. Our goal is to train a single network that learns to avoid these problems and generate a coherent 3D reconstruction of all the humans in the scene. To this end, a key design choice is the incorporation of the SMPL parametric body model in our top-down framework, which enables the use of two novel losses. First, a distance field-based collision loss penalizes interpenetration among the reconstructed people. Second, a depth ordering-aware loss reasons about occlusions and promotes a depth ordering of people that leads to a rendering which is consistent with the annotated instance segmentation. This provides depth supervision signals to the network, even if the image has no explicit 3D annotations. The experiments show that our approach outperforms previous methods on standard 3D pose benchmarks, while our proposed losses enable more coherent reconstruction in natural images. The project website with videos, results, and code can be found at: https://jiangwenpl.github.io/multiperson.
Wen Jiang 0008, Nikos Kolotouros, Georgios Pavlakos, Xiaowei Zhou 0001, Kostas Daniilidis
CVPR5
2020 3D Bird Reconstruction: A Dataset, Model, and Shape Recovery from a Single View
Marc Badger, Yufu Wang, Adarsh Modh, Ammon Perkes, Nikos Kolotouros, Bernd Pfrommer, Marc F. Schmidt, Kostas Daniilidis
ECCV (18)8
2020 Spike-FlowNet: Event-Based Optical Flow Estimation with Energy-Efficient Hybrid Neural Networks
Chankyu Lee, Adarsh Kosta, Alex Zihao Zhu, Kenneth Chaney, Kostas Daniilidis, Kaushik Roy 0001
ECCV (29)5
2020 Learning Predictive Models from Observation and Interaction
Karl Schmeckpeper, Annie Xie, Oleh Rybkin, Stephen Tian, Kostas Daniilidis, Sergey Levine, Chelsea Finn
ECCV (20)5
2020 Planning to Explore via Self-Supervised World Models
abstract
Reinforcement learning allows solving complex tasks, however, the learning tends to be task-specific and the sample efficiency remains a challenge. We present Plan2Explore, a self-supervised reinforcement learning agent that tackles both these challenges through a new approach to self-supervised exploration and fast adaptation to new tasks, which need not be known during exploration. During exploration, unlike prior methods which retrospectively compute the novelty of observations after the agent has already reached them, our agent acts efficiently by leveraging planning to seek out expected future novelty. After exploration, the agent quickly adapts to multiple downstream tasks in a zero or a few-shot manner. We evaluate on challenging control tasks from high-dimensional image inputs. Without any training supervision or task-specific interaction, Plan2Explore outperforms prior self-supervised exploration methods, and in fact, almost matches the performances oracle which has access to rewards. Videos and code: https://ramanans1.github.io/plan2explore/
Ramanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel, Danijar Hafner, Deepak Pathak
ICML3
2020 A Low-Rank Matrix Approximation Approach to Multiway Matching with Applications in Multi-Sensory Data Association
abstract
Consider the case of multiple visual sensors perceiving the same scene from different viewpoints. In order to achieve consistent visual perception, the problem of data association, in this case establishing correspondences between observed features, must be first solved. In this work, we consider multiway matching which is a specific instance of multi-sensory data association. Multiway matching refers to the problem of establishing correspondences among a set of images from noisy pairwise correspondences, typically by exploiting cycle- consistency. We propose a novel optimization-based formulation of multiway matching problem as a nonconvex low-rank matrix approximation problem. We propose two novel algorithms for numerically solving the problem at hand. The first one is an algorithm based on the Alternating Direction Method of Multipliers (ADMM). The second one is a Riemannian trust- region method on the multinomial manifold, the manifold of strictly positive stochastic matrices, equipped with the Fisher information metric. Experimental results demonstrate that the proposed methods have the state of the art performance in multiway matching while reducing the computational complexity compared to the state of the art.
Spyridon Leonardos, Xiaowei Zhou 0001, Kostas Daniilidis
ICRA3
2020 The Tiercel: A novel autonomous micro aerial vehicle that can map the environment by flying into obstacles
abstract
Autonomous flight through unknown environments in the presence of obstacles is a challenging problem for micro aerial vehicles (MAVs). A majority of the current state-of-art research assumes obstacles as opaque objects that can be easily sensed by optical sensors such as cameras or LiDARs. However in indoor environments with glass walls and windows, or scenarios with smoke and dust, robots (even birds) have a difficult time navigating through the unknown space.In this paper, we present the design of a new class of micro aerial vehicles that achieves autonomous navigation and are robust to collisions. In particular, we present the Tiercel MAV: a small, agile, light weight and collision-resilient robot powered by a cellphone grade CPU. Our design exploits contact to infer the presence of transparent or reflective obstacles like glass walls, integrating touch with visual perception for SLAM. The Tiercel is able to localize using visual-inertial odometry (VIO) running on board the robot with a single downward facing fisheye camera and an IMU. We show how our collision detector design and experimental set up enable us to characterize the impact of collisions on VIO. We further develop a planning strategy to enable the Tiercel to fly autonomously in an unknown space, sustaining collisions and creating a 2D map of the environment. Finally we demonstrate a swarm of three autonomous Tiercel robots safely navigating and colliding through an obstacle field to reach their objectives.
Yash Mulgaonkar, Wenxin Liu 0002, Dinesh Thakur, Kostas Daniilidis, Camillo J. Taylor, Vijay Kumar 0001
ICRA4
2020 Spin-Weighted Spherical CNNs
abstract
Learning equivariant representations is a promising way to reduce sample and model complexity and improve the generalization performance of deep neural networks. The spherical CNNs are successful examples, producing SO(3)-equivariant representations of spherical inputs. There are two main types of spherical CNNs. The first type lifts the inputs to functions on the rotation group SO(3) and applies convolutions on the group, which are computationally expensive since SO(3) has one extra dimension. The second type applies convolutions directly on the sphere, which are limited to zonal (isotropic) filters, and thus have limited expressivity. In this paper, we present a new type of spherical CNN that allows anisotropic filters in an efficient way, without ever leaving the spherical domain. The key idea is to consider spin-weighted spherical functions, which were introduced in physics in the study of gravitational waves. These are complex-valued functions on the sphere whose phases change upon rotation. We define a convolution between spin-weighted functions and build a CNN based on it. The spin-weighted functions can also be interpreted as spherical vector fields, allowing applications to tasks where the inputs or outputs are vector fields. Experiments show that our method outperforms previous methods on tasks like classification of spherical images, classification of 3D shapes and semantic segmentation of spherical panoramas.
Carlos Esteves, Ameesh Makadia, Kostas Daniilidis
NeurIPS3
2020 Learning SO(3) Equivariant Representations with Spherical CNNs
Carlos Esteves, Christine Allen-Blanchette, Ameesh Makadia, Kostas Daniilidis
Int. J. Comput. Vis.4
2019 Convolutional Mesh Regression for Single-Image Human Shape Reconstruction
abstract
This paper addresses the problem of 3D human pose and shape estimation from a single image. Previous approaches consider a parametric model of the human body, SMPL, and attempt to regress the model parameters that give rise to a mesh consistent with image evidence. This parameter regression has been a very challenging task, with model-based approaches underperforming compared to nonparametric solutions in terms of pose estimation. In our work, we propose to relax this heavy reliance on the model's parameter space. We still retain the topology of the SMPL template mesh, but instead of predicting model parameters, we directly regress the 3D location of the mesh vertices. This is a heavy task for a typical network, but our key insight is that the regression becomes significantly easier using a Graph-CNN. This architecture allows us to explicitly encode the template mesh structure within the network and leverage the spatial locality the mesh has to offer. Image-based features are attached to the mesh vertices and the Graph-CNN is responsible to process them on the mesh structure, while the regression target for each vertex is its 3D location. Having recovered the complete 3D geometry of the mesh, if we still require a specific model parametrization, this can be reliably regressed from the vertices locations. We demonstrate the flexibility and the effectiveness of our proposed graph-based mesh regression by attaching different types of features on the mesh vertices. In all cases, we outperform the comparable baselines relying on model parameter regression, while we also achieve state-of-the-art results among model-based pose estimation approaches.
Nikos Kolotouros, Georgios Pavlakos, Kostas Daniilidis
CVPR3
2019 Unsupervised Event-Based Learning of Optical Flow, Depth, and Egomotion
abstract
In this work, we propose a novel framework for unsupervised learning for event cameras that learns motion information from only the event stream. In particular, we propose an input representation of the events in the form of a discretized volume that maintains the temporal distribution of the events, which we pass through a neural network to predict the motion of the events. This motion is used to attempt to remove any motion blur in the event image. We then propose a loss function applied to the motion compensated event image that measures the motion blur in this image. We train two networks with this framework, one to predict optical flow, and one to predict egomotion and depths, and evaluate these networks on the Multi Vehicle Stereo Event Camera dataset, along with qualitative results from a variety of different scenes.
Alex Zihao Zhu, Liangzhe Yuan, Kenneth Chaney, Kostas Daniilidis
CVPR4
2019 Equivariant Multi-View Networks
abstract
Several popular approaches to 3D vision tasks process multiple views of the input independently with deep neural networks pre-trained on natural images, where view permutation invariance is achieved through a single round of pooling over all views. We argue that this operation discards important information and leads to subpar global descriptors. In this paper, we propose a group convolutional approach to multiple view aggregation where convolutions are performed over a discrete subgroup of the rotation group, enabling, thus, joint reasoning over all views in an equivariant (instead of invariant) fashion, up to the very last layer. We further develop this idea to operate on smaller discrete homogeneous spaces of the rotation group, where a polar view representation is used to maintain equivariance with only a fraction of the number of input views. We set the new state of the art in several large scale 3D shape retrieval tasks, and show additional applications to panoramic scene classification.
Carlos Esteves, Yinshuang Xu, Christine Allen-Blanchette, Kostas Daniilidis
ICCV4
2019 Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the Loop
abstract
Model-based human pose estimation is currently approached through two different paradigms. Optimization-based methods fit a parametric body model to 2D observations in an iterative manner, leading to accurate image-model alignments, but are often slow and sensitive to the initialization. In contrast, regression-based methods, that use a deep network to directly estimate the model parameters from pixels, tend to provide reasonable, but not pixel accurate, results while requiring huge amounts of supervision. In this work, instead of investigating which approach is better, our key insight is that the two paradigms can form a strong collaboration. A reasonable, directly regressed estimate from the network can initialize the iterative optimization making the fitting faster and more accurate. Similarly, a pixel accurate fit from iterative optimization can act as strong supervision for the network. This is the core of our proposed approach SPIN (SMPL oPtimization IN the loop). The deep network initializes an iterative optimization routine that fits the body model to 2D joints within the training loop, and the fitted estimate is subsequently used to supervise the network. Our approach is self-improving by nature, since better network estimates can lead the optimization to better solutions, while more accurate optimization fits provide better supervision for the network. We demonstrate the effectiveness of our approach in different settings, where 3D ground truth is scarce, or not available, and we consistently outperform the state-of-the-art model-based pose estimation approaches by significant margins. The project website with videos, results, and code can be found at https://seas.upenn.edu/~nkolot/projects/spin.
Nikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas Daniilidis
ICCV4
2019 TexturePose: Supervising Human Mesh Estimation With Texture Consistency
abstract
This work addresses the problem of model-based human pose estimation. Recent approaches have made significant progress towards regressing the parameters of parametric human body models directly from images. Because of the absence of images with 3D shape ground truth, relevant approaches rely on 2D annotations or sophisticated architecture designs. In this work, we advocate that there are more cues we can leverage, which are available for free in natural images, i.e., without getting more annotations, or modifying the network architecture. We propose a natural form of supervision, that capitalizes on the appearance constancy of a person among different frames (or viewpoints). This seemingly insignificant and often overlooked cue goes a long way for model-based pose estimation. The parametric model we employ allows us to compute a texture map for each frame. Assuming that the texture of the person does not change dramatically between frames, we can apply a novel texture consistency loss, which enforces that each point in the texture map has the same texture value across all frames. Since the texture is transferred in this common texture map space, no camera motion computation is necessary, or even an assumption of smoothness among frames. This makes our proposed supervision applicable in a variety of settings, ranging from monocular video, to multi-view images. We benchmark our approach against strong baselines that require the same or even more annotations that we do and we consistently outperform them. Simultaneously, we achieve state-of-the-art results among model-based pose estimation approaches in different benchmarks. The project website with videos, results, and code can be found at https://seas.upenn.edu/~pavlakos/projects/texturepose.
Georgios Pavlakos, Nikos Kolotouros, Kostas Daniilidis
ICCV3
2019 Learning what you can do before doing anything
Oleh Rybkin, Karl Pertsch, Konstantinos G. Derpanis, Kostas Daniilidis, Andrew Jaegle
ICLR (Poster)4
2019 Cross-Domain 3D Equivariant Image Embeddings
abstract
Spherical convolutional networks have been introduced recently as tools to learn powerful feature representations of 3D shapes. Spherical CNNs are equivariant to 3D rotations making them ideally suited to applications where 3D data may be observed in arbitrary orientations. In this paper we learn 2D image embeddings with a similar equivariant structure: embedding the image of a 3D object should commute with rotations of the object. We introduce a cross-domain embedding from 2D images into a spherical CNN latent space. This embedding encodes images with 3D shape properties and is equivariant to 3D rotations of the observed object. The model is supervised only by target embeddings obtained from a spherical CNN pretrained for 3D shape classification. We show that learning a rich embedding for images with appropriate geometric structure is sufficient for tackling varied applications, such as relative pose estimation and novel view synthesis, without requiring additional task-specific supervision.
Carlos Esteves, Avneesh Sud, Zhengyi Luo 0002, Kostas Daniilidis, Ameesh Makadia
ICML4
2019 Learning Event-based Height from Plane and Parallax
abstract
Event-based cameras are a novel asynchronous sensing modality that provide exciting benefits, such as the ability to track fast moving objects with no motion blur and low latency, high dynamic range, and low power consumption. Given the low latency of the cameras, as well as their ability to work in challenging lighting conditions, these cameras are a natural fit for reactive problems such as fast local structure estimation. In this work, we propose a fast method to perform structure estimation for vehicles traveling in a roughly 2D environment (e.g. in an environment with a ground plane). Our method transfers the method of plane and parallax to events, which, given the homography to a ground plane and the pose of the camera, generates a warping of the events which removes the optical flow for events on the ground plane, while inducing flow for events above the ground plane. We then estimate dense flow in this warped space using a self-supervised neural network, which provides the height of all points in the scene. We evaluate our method on the Multi Vehicle Stereo Event Camera dataset, and show its ability to rapidly estimate the scene structure both at high speeds and in low lighting conditions.
Kenneth Chaney, Alex Zihao Zhu, Kostas Daniilidis
IROS3
2019 MonoCap: Monocular Human Motion Capture using a CNN Coupled with a Geometric Prior
abstract
Recovering 3D full-body human pose is a challenging problem with many applications. It has been successfully addressed by motion capture systems with body worn markers and multiple cameras. In this paper, we address the more challenging case of not only using a single camera but also not leveraging markers: going directly from 2D appearance to 3D geometry. Deep learning approaches have shown remarkable abilities to discriminatively learn 2D appearance features. The missing piece is how to integrate 2D, 3D, and temporal information to recover 3D geometry and account for the uncertainties arising from the discriminative model. We introduce a novel approach that treats 2D joint locations as latent variables whose uncertainty distributions are given by a deep fully convolutional neural network. The unknown 3D poses are modeled by a sparse representation and the 3D parameter estimates are realized via an Expectation-Maximization algorithm, where it is shown that the 2D joint location uncertainties can be conveniently marginalized out during inference. Extensive evaluation on benchmark datasets shows that the proposed approach achieves greater accuracy over state-of-the-art baselines. Notably, the proposed approach does not require synchronized 2D-3D data for training and is applicable to "in-the-wild" images, which is demonstrated with the MPII dataset.
Xiaowei Zhou 0001, Menglong Zhu, Georgios Pavlakos, Spyridon Leonardos, Konstantinos G. Derpanis, Kostas Daniilidis
IEEE Trans. Pattern Anal. Mach. Intell.6
2018 Ordinal Depth Supervision for 3D Human Pose Estimation
abstract
Our ability to train end-to-end systems for 3D human pose estimation from single images is currently constrained by the limited availability of 3D annotations for natural images. Most datasets are captured using Motion Capture (MoCap) systems in a studio setting and it is difficult to reach the variability of 2D human pose datasets, like MPII or LSP. To alleviate the need for accurate 3D ground truth, we propose to use a weaker supervision signal provided by the ordinal depths of human joints. This information can be acquired by human annotators for a wide range of images and poses. We showcase the effectiveness and flexibility of training Convolutional Networks (ConvNets) with these ordinal relations in different settings, always achieving competitive performance with ConvNets trained with accurate 3D joint coordinates. Additionally, to demonstrate the potential of the approach, we augment the popular LSP and MPII datasets with ordinal depth annotations. This extension allows us to present quantitative and qualitative evaluation in non-studio conditions. Simultaneously, these ordinal annotations can be easily incorporated in the training procedure of typical ConvNets for 3D human pose. Through this inclusion we achieve new state-of-the-art performance for the relevant benchmarks and validate the effectiveness of ordinal depth supervision for 3D human pose.
Georgios Pavlakos, Xiaowei Zhou 0001, Kostas Daniilidis
CVPR3
2018 Learning to Estimate 3D Human Pose and Shape From a Single Color Image
abstract
This work addresses the problem of estimating the full body 3D human pose and shape from a single color image. This is a task where iterative optimization-based solutions have typically prevailed, while Convolutional Networks (ConvNets) have suffered because of the lack of training data and their low resolution 3D predictions. Our work aims to bridge this gap and proposes an efficient and effective direct prediction method based on ConvNets. Central part to our approach is the incorporation of a parametric statistical body shape model (SMPL) within our end-to-end framework. This allows us to get very detailed 3D mesh results, while requiring estimation only of a small number of parameters, making it friendly for direct network prediction. Interestingly, we demonstrate that these parameters can be predicted reliably only from 2D keypoints and masks. These are typical outputs of generic 2D human analysis ConvNets, allowing us to relax the massive requirement that images with 3D shape ground truth are available for training. Simultaneously, by maintaining differentiability, at training time we generate the 3D mesh from the estimated parameters and optimize explicitly for the surface using a 3D per-vertex loss. Finally, a differentiable renderer is employed to project the 3D mesh to the image, which enables further refinement of the network, by optimizing for the consistency of the projection with 2D annotations (i.e., 2D keypoints or masks). The proposed approach outperforms previous baselines on this task and offers an attractive solution for direct prediction of3D shape from a single color image.
Georgios Pavlakos, Luyang Zhu, Xiaowei Zhou 0001, Kostas Daniilidis
CVPR4
2018 Multi-Image Semantic Matching by Mining Consistent Features
abstract
This work proposes a multi-image matching method to estimate semantic correspondences across multiple images. In contrast to the previous methods that optimize all pairwise correspondences, the proposed method identifies and matches only a sparse set of reliable features in the image collection. In this way, the proposed method is able to prune nonrepeatable features and also highly scalable to handle thousands of images. We additionally propose a low-rank constraint to ensure the geometric consistency of feature correspondences over the whole image collection. Besides the competitive performance on multi-graph matching and semantic flow benchmarks, we also demonstrate the applicability of the proposed method for reconstructing object-class models and discovering object-class landmarks from images without using any annotation.
Qianqian Wang 0002, Xiaowei Zhou 0001, Kostas Daniilidis
CVPR3
2018 Learning SO(3) Equivariant Representations with Spherical CNNs
Carlos Esteves, Christine Allen-Blanchette, Ameesh Makadia, Kostas Daniilidis
ECCV (13)4
2018 Realtime Time Synchronized Event-Based Stereo
Alex Zihao Zhu, Kostas Daniilidis
ECCV (6)3
2018 Polar Transformer Networks
Carlos Esteves, Christine Allen-Blanchette, Xiaowei Zhou 0001, Kostas Daniilidis
ICLR (Poster)4
2018 Understanding image motion with group representations
Andrew Jaegle, Stephen Phillips, Daphne Ippolito, Kostas Daniilidis
ICLR (Poster)4
2018 Semi-Dense Visual-Inertial Odometry and Mapping for Quadrotors with SWAP Constraints
abstract
Micro Aerial Vehicles have the potential to assist humans in real life tasks involving applications such as smart homes, search and rescue, and architecture construction. To enhance autonomous navigation capabilities these vehicles need to be able to create dense 3D maps of the environment, while concurrently estimating their own motion. In this paper, we are particularly interested in small vehicles that can navigate cluttered indoor environments. We address the problem of visual inertial state estimation, control and 3D mapping on platforms with Size, Weight, And Power (SWAP) constraints. The proposed approach is validated through experimental results on a 250 g, 22 cm diameter quadrotor equipped only with a stereo camera and an IMU with a computationally-limited CPU showing the ability to autonomously navigate, while concurrently creating a 3D map of the environment.
Wenxin Liu 0002, Giuseppe Loianno, Kartik Mohta, Kostas Daniilidis, Vijay Kumar 0001
ICRA4
2018 Human Motion Capture Using a Drone
abstract
Current motion capture (MoCap) systems generally require markers and multiple calibrated cameras, which can be used only in constrained environments. In this work we introduce a drone-based system for 3D human MoCap. The system only needs an autonomously flying drone with an on-board RGB camera and is usable in various indoor and outdoor environments. A reconstruction algorithm is developed to recover full-body motion from the video recorded by the drone. We argue that, besides the capability of tracking a moving subject, a flying drone also provides fast varying viewpoints, which is beneficial for motion reconstruction. We evaluate the accuracy of the proposed system using our new DroCap dataset and also demonstrate its applicability for MoCap in the wild using a consumer drone.
Xiaowei Zhou 0001, Sikang Liu 0002, Georgios Pavlakos, Vijay Kumar 0001, Kostas Daniilidis
ICRA5
2018 A Unifying View of Geometry, Semantics, and Data Association in SLAM
abstract
Traditional approaches for simultaneous localization and mapping (SLAM) rely on geometric features such as points, lines, and planes to infer the environment structure. They make hard decisions about the (data) association between observed features and mapped landmarks to update the environment model. This paper makes two contributions to the state of the art in SLAM. First, it generalizes the purely geometric model by introducing semantically meaningful objects, represented as structured models of mid-level part features. Second, instead of making hard, potentially wrong associations between semantic features and objects, it shows that SLAM inference can be performed efficiently with probabilistic data association. The approach not only allows building meaningful maps (containing doors, chairs, cars, etc.) but also offers significant advantages in ambiguous environments.
Nikolay Atanasov 0001, Sean L. Bowman, Kostas Daniilidis, George J. Pappas
IJCAI3
2017 Harvesting Multiple Views for Marker-Less 3D Human Pose Annotations
abstract
Recent advances with Convolutional Networks (ConvNets) have shifted the bottleneck for many computer vision tasks to annotated data collection. In this paper, we present a geometry-driven approach to automatically collect annotations for human pose prediction tasks. Starting from a generic ConvNet for 2D human pose, and assuming a multi-view setup, we describe an automatic way to collect accurate 3D human pose annotations. We capitalize on constraints offered by the 3D geometry of the camera setup and the 3D structure of the human body to probabilistically combine per view 2D ConvNet predictions into a globally optimal 3D pose. This 3D pose is used as the basis for harvesting annotations. The benefit of the annotations produced automatically with our approach is demonstrated in two challenging settings: (i) fine-tuning a generic ConvNet-based 2D pose predictor to capture the discriminative aspects of a subjects appearance (i.e.,personalization), and (ii) training a ConvNet from scratch for single view 3D human pose prediction without leveraging 3D pose groundtruth. The proposed multi-view pose estimator achieves state-of-the-art results on standard benchmarks, demonstrating the effectiveness of our method in exploiting the available multi-view information.
Georgios Pavlakos, Xiaowei Zhou 0001, Konstantinos G. Derpanis, Kostas Daniilidis
CVPR4
2017 Coarse-to-Fine Volumetric Prediction for Single-Image 3D Human Pose
abstract
This paper addresses the challenge of 3D human pose estimation from a single color image. Despite the general success of the end-to-end learning paradigm, top performing approaches employ a two-step solution consisting of a Convolutional Network (ConvNet) for 2D joint localization and a subsequent optimization step to recover 3D pose. In this paper, we identify the representation of 3D pose as a critical issue with current ConvNet approaches and make two important contributions towards validating the value of end-to-end learning for this task. First, we propose a fine discretization of the 3D space around the subject and train a ConvNet to predict per voxel likelihoods for each joint. This creates a natural representation for 3D pose and greatly improves performance over the direct regression of joint coordinates. Second, to further improve upon initial estimates, we employ a coarse-to-fine prediction scheme. This step addresses the large dimensionality increase and enables iterative refinement and repeated processing of the image features. The proposed approach outperforms all state-of-the-art methods on standard benchmarks achieving a relative error reduction greater than 30% on average. Additionally, we investigate using our volumetric representation in a related architecture which is suboptimal compared to our end-to-end approach, but is of practical interest, since it enables training when no image with corresponding 3D groundtruth is available, and allows us to present compelling results for in-the-wild images.
Georgios Pavlakos, Xiaowei Zhou 0001, Konstantinos G. Derpanis, Kostas Daniilidis
CVPR4
2017 Event-Based Visual Inertial Odometry
abstract
Event-based cameras provide a new visual sensing model by detecting changes in image intensity asynchronously across all pixels on the camera. By providing these events at extremely high rates (up to 1MHz), they allow for sensing in both high speed and high dynamic range situations where traditional cameras may fail. In this paper, we present the first algorithm to fuse a purely event-based tracking algorithm with an inertial measurement unit, to provide accurate metric tracking of a cameras full 6dof pose. Our algorithm is asynchronous, and provides measurement updates at a rate proportional to the camera velocity. The algorithm selects features in the image plane, and tracks spatiotemporal windows around these features within the event stream. An Extended Kalman Filter with a structureless measurement model then fuses the feature tracks with the output of the IMU. The camera poses from the filter are then used to initialize the next step of the tracker and reject failed tracks. We show that our method successfully tracks camera motion on the Event-Camera Dataset in a number of challenging situations.
Alex Zihao Zhu, Nikolay Atanasov 0001, Kostas Daniilidis
CVPR3
2017 Fast Multi-image Matching via Density-Based Clustering
Roberto Tron, Xiaowei Zhou 0001, Carlos Esteves, Kostas Daniilidis
ICCV4
2017 Probabilistic data association for semantic SLAM
abstract
Traditional approaches to simultaneous localization and mapping (SLAM) rely on low-level geometric features such as points, lines, and planes. They are unable to assign semantic labels to landmarks observed in the environment. Furthermore, loop closure recognition based on low-level features is often viewpoint-dependent and subject to failure in ambiguous or repetitive environments. On the other hand, object recognition methods can infer landmark classes and scales, resulting in a small set of easily recognizable landmarks, ideal for view-independent unambiguous loop closure. In a map with several objects of the same class, however, a crucial data association problem exists. While data association and recognition are discrete problems usually solved using discrete inference, classical SLAM is a continuous optimization over metric information. In this paper, we formulate an optimization problem over sensor states and semantic landmark positions that integrates metric information, semantic information, and data associations, and decompose it into two interconnected problems: an estimation of discrete data association and landmark class probabilities, and a continuous optimization over the metric states. The estimated landmark and robot poses affect the association and class distributions, which in turn affect the robot-landmark pose optimization. The performance of our algorithm is demonstrated on indoor and outdoor datasets.
Sean L. Bowman, Nikolay Atanasov 0001, Kostas Daniilidis, George J. Pappas
ICRA3
2017 Distributed consistent data association via permutation synchronization
abstract
Data association is one of the fundamental problems in multi-sensor systems. Most current techniques rely on pairwise data associations which can be spurious even after the employment of outlier rejection schemes. Considering multiple pairwise associations at once significantly increases accuracy and leads to consistency. In this work, we propose a fully decentralized method for globally consistent data association from pairwise data associations based on a distributed averaging scheme on the set of doubly stochastic matrices. We demonstrate the effectiveness of the proposed method using theoretical analysis and experimental evaluation.
Spyridon Leonardos, Xiaowei Zhou 0001, Kostas Daniilidis
ICRA3
2017 6-DoF object pose from semantic keypoints
abstract
This paper presents a novel approach to estimating the continuous six degree of freedom (6-DoF) pose (3D translation and rotation) of an object from a single RGB image. The approach combines semantic keypoints predicted by a convolutional network (convnet) with a deformable shape model. Unlike prior work, we are agnostic to whether the object is textured or textureless, as the convnet learns the optimal representation from the available training image data. Furthermore, the approach can be applied to instance- and class-based pose recovery. Empirically, we show that the proposed approach can accurately recover the 6-DoF object pose for both instance- and class-based scenarios with a cluttered background. For class-based object pose estimation, state-of-the-art accuracy is shown on the large-scale PASCAL3D+ dataset.
Georgios Pavlakos, Xiaowei Zhou 0001, Aaron Chan, Konstantinos G. Derpanis, Kostas Daniilidis
ICRA5
2017 PennCOSYVIO: A challenging Visual Inertial Odometry benchmark
abstract
We present PennCOSYVIO, a new challenging Visual Inertial Odometry (VIO) benchmark with synchronized data from a VI-sensor (stereo camera and IMU), two Project Tango hand-held devices, and three GoPro Hero 4 cameras. Recorded at UPenn's Singh center, the 150m long path of the hand-held rig crosses from outdoors to indoors and includes rapid rotations, thereby testing the abilities of VIO and Simultaneous Localization and Mapping (SLAM) algorithms to handle changes in lighting, different textures, repetitive structures, and large glass surfaces. All sensors are synchronized and intrinsically and extrinsically calibrated. We demonstrate the accuracy with which ground-truth poses can be obtained via optic localization off of fiducial markers. The data set can be found at https://daniilidis-group.github.io/penncosyvio/.
Bernd Pfrommer, Nitin Sanket, Kostas Daniilidis, Jonas Cleveland
ICRA3
2017 Event-based feature tracking with probabilistic data association
abstract
Asynchronous event-based sensors present new challenges in basic robot vision problems like feature tracking. The few existing approaches rely on grouping events into models and computing optical flow after assigning future events to those models. Such a hard commitment in data association attenuates the optical flow quality and causes shorter flow tracks. In this paper, we introduce a novel soft data association modeled with probabilities. The association probabilities are computed in an intertwined EM scheme with the optical flow computation that maximizes the expectation (marginalization) over all associations. In addition, to enable longer tracks we compute the affine deformation with respect to the initial point and use the resulting residual as a measure of persistence. The computed optical flow enables a varying temporal integration different for every feature and sized inversely proportional to the length of the flow. We show results in egomotion and very fast vehicle sequences and we show the superiority over standard frame-based cameras.
Alex Zihao Zhu, Nikolay Atanasov 0001, Kostas Daniilidis
ICRA3
2017 Precise dispensing of liquids using visual feedback
abstract
Robotic pouring is an important step in improving the safety, productivity and repeatability in the biotechnology industry and generally increasing the effectiveness of robotics in human based environments. In this work we present a method to autonomously dispense a precise amount of fluid using only visual feedback without using precision pouring instruments such as pipettes, syringes or pourers. We model circular and rectangular pouring container geometries. We prove that for square containers we can control the flow by only observing the fluid height in the receiving beaker. We show a systematic approach using a hybrid control scheme that is robust to the initial amount of fluid in the pouring container and inconsistent flow. Specifically we present (a) a model for pouring (b) a model based algorithm to drive a robot arm (c) visual feedback for regulating the pouring rate. We demonstrate this using the Rethink Robotics Sawyer manipulator and mvBluefox MLC202bc camera.
Monroe Kennedy III, Kendall Queen, Dinesh Thakur, Kostas Daniilidis, Vijay Kumar 0001
IROS4
2017 Shape-based object classification and recognition through continuum manipulation
abstract
We introduce a novel approach to shape-based object classification and recognition through the use of a continuum manipulator. Noticing the fact that when a continuum manipulator wraps around an object in a whole-arm grasping, its own shape is indicative of the shape of the object, our approach enables learning and recognition of object classes based on the shapes of continuum wraps. It offers the following advantages: (1) recognition of objects that are not easily detected by vision, such as transparent objects, and (2) highly efficient recognition of such objects of varied sizes due to high-level and rich shape information in each wrap, unlike recognition based on tactile sensing via conventional grasping. Simulation and experiments demonstrate the effectiveness of our approach.
Huitan Mao, Jing Xiao 0001, Mabel M. Zhang, Kostas Daniilidis
IROS4
2017 Active end-effector pose selection for tactile object recognition through Monte Carlo tree search
abstract
This paper considers the problem of active object recognition using touch only. The focus is on adaptively selecting a sequence of wrist poses that achieves accurate recognition by enclosure grasps. It seeks to minimize the number of touches and maximize recognition confidence. The actions are formulated as wrist poses relative to each other, making the algorithm independent of absolute workspace coordinates. The optimal sequence is approximated by Monte Carlo tree search. We demonstrate results in a physics engine and on a real robot. In the physics engine, most object instances were recognized in at most 16 grasps. On a real robot, our method recognized objects in 2-9 grasps and outperformed a greedy baseline.
Mabel M. Zhang, Nikolay Atanasov 0001, Kostas Daniilidis
IROS3
2017 Sparse Representation for 3D Shape Estimation: A Convex Relaxation Approach
abstract
We investigate the problem of estimating the 3D shape of an object defined by a set of 3D landmarks, given their 2D correspondences in a single image. A successful approach to alleviating the reconstruction ambiguity is the 3D deformable shape model and a sparse representation is often used to capture complex shape variability. But the model inference is still challenging due to the nonconvexity in the joint optimization of shape and viewpoint. In contrast to prior work that relies on an alternating scheme whose solution depends on initialization, we propose a convex approach to addressing this challenge and develop an efficient algorithm to solve the proposed convex program. We further propose a robust model to handle gross errors in the 2D correspondences. We demonstrate the exact recovery property of the proposed method, the advantage compared to several nonconvex baselines and the applicability to recover 3D human poses and car models from single images.
Xiaowei Zhou 0001, Menglong Zhu, Spyridon Leonardos, Kostas Daniilidis
IEEE Trans. Pattern Anal. Mach. Intell.4
2017 The Space of Essential Matrices as a Riemannian Quotient Manifold
abstract
The essential matrix, which encodes the epipolar constraint between points in two projective views, is a cornerstone of modern computer vision. Previous works have proposed different characterizations of the space of essential matrices as a Riemannian manifold. However, they either do not consider the symmetric role played by the two views or do not fully take into account the geometric peculiarities of the epipolar constraint. We address these limitations with a characterization as a quotient manifold that can be easily interpreted in terms of camera poses. While our main focus is on theoretical aspects, we include applications to optimization problems in computer vision.
Roberto Tron, Kostas Daniilidis
SIAM J. Imaging Sci.2
2017 Automated System for Semantic Object Labeling With Soft-Object Recognition and Dynamic Programming Segmentation
abstract
This paper presents an automated robotic system for generating semantic maps of inventory in retail environments. In retail settings, semantic maps are labeled maps of stores where each discrete section of shelving is assigned a department label describing the types of products on that shelf. Starting from a metric map of the store, the robot autonomously extracts the shelf boundaries, generates a distance-optimal tour of the store to view every shelf, and follows the tour while avoiding unmapped clutter and moving people. The robot creates a point cloud of the store using the data collected from this tour. We introduce a novel soft-object assignment algorithm to create a virtual map and a dynamic programming algorithm to segment this map. These algorithms use a priori information about the products to boost data from laser and camera sensors in order to recognize and semantically label objects. The primary contribution of this paper is the integration of multiple systems for automated path planning, navigation, object recognition, and semantic mapping. This paper represents an important contribution toward deploying mobile robots in dynamic human environments.
Jonas Cleveland, Dinesh Thakur, Philip M. Dames, Cody J. Phillips 0001, Terry Kientz, Kostas Daniilidis, John Bergstrom, Vijay Kumar 0001
IEEE Trans Autom. Sci. Eng.6
2016 Absolute Pose and Structure from Motion for Surfaces of Revolution: Minimal Problems Using Apparent Contours
abstract
The class of objects that can be represented by surfaces of revolution (SoRs) is highly prevalent in human work and living spaces. Due to their prevalence and convenient geometric properties, SoRs have been employed over the past thirty years for single-view camera calibration and pose estimation, and have been studied in terms of SoR object reconstruction and recognition. Such treatment has provided techniques for the automatic identification and classification of important SoR structures, such as apparent contours, cross sections, bitangent points, creases, and inflections. The presence of these structures are crucial to most SoR-based image metrology algorithms. This paper develops single-view and two-view pose recovery and reconstruction formulations that only require apparent contours, and no other SoR features.The primary objective of this paper is to present and experimentallyvalidate the minimal problems pertaining toSoR metrology from apparent contours. For a single view with a known reference model, this includes absolute pose recovery. For many views and no reference model this is extended to structure from motion (SfM). Assuming apparent contours as input that have been identified and segmented with reasonable accuracy, the minimal problems aredemonstrated to produce accurate SoR pose and shape results when used as part of a RANSAC-based hypothesis generation and evaluation pipeline.
Cody J. Phillips 0001, Kostas Daniilidis
3DV2
2016 Sparseness Meets Deepness: 3D Human Pose Estimation from Monocular Video
abstract
This paper addresses the challenge of 3D full-body human pose estimation from a monocular image sequence. Here, two cases are considered: (i) the image locations of the human joints are provided and (ii) the image locations of joints are unknown. In the former case, a novel approach is introduced that integrates a sparsity-driven 3D geometric prior and temporal smoothness. In the latter case, the former case is extended by treating the image locations of the joints as latent variables to take into account considerable uncertainties in 2D joint locations. A deep fully convolutional network is trained to predict the uncertainty maps of the 2D joint locations. The 3D pose estimates are realized via an Expectation-Maximization algorithm over the entire sequence, where it is shown that the 2D joint location uncertainties can be conveniently marginalized out during inference. Empirical evaluation on the Human3.6M dataset shows that the proposed approaches achieve greater 3D pose estimation accuracy over state-of-the-art baselines. Further, the proposed approach outperforms a publicly available 2D pose estimation baseline on the challenging PennAction dataset.
Xiaowei Zhou 0001, Menglong Zhu, Spyridon Leonardos, Konstantinos G. Derpanis, Kostas Daniilidis
CVPR5
2016 Fast, robust, continuous monocular egomotion computation
abstract
We propose robust methods for estimating camera egomotion in noisy, real-world monocular image sequences in the general case of unknown observer rotation and translation with two views and a small baseline. This is a difficult problem because of the nonconvex cost function of the perspective camera motion equation and because of non-Gaussian noise arising from noisy optical flow estimates and scene non-rigidity. To address this problem, we introduce the expected residual likelihood method (ERL), which estimates confidence weights for noisy optical flow data using likelihood distributions of the residuals of the flow field under a range of counterfactual model parameters. We show that ERL is effective at identifying outliers and recovering appropriate confidence weights in many settings. We compare ERL to a novel formulation of the perspective camera motion equation using a lifted kernel, a recently proposed optimization framework for joint parameter and confidence weight estimation with good empirical properties. We incorporate these strategies into a motion estimation pipeline that avoids falling into local minima. We find that ERL outperforms the lifted kernel method and baseline monocular egomotion estimation strategies on the challenging KITTI dataset, while adding almost no runtime cost over baseline egomotion methods.
Andrew Jaegle, Stephen Phillips, Kostas Daniilidis
ICRA3
2016 Articulated motion estimation from a monocular image sequence using spherical tangent bundles
abstract
We propose a second order stochastic dynamical model for generic articulated objects whose state space is a Riemannian manifold naturally suggested by the articulation constraints. We derive the equations of a Riemannian Extended Kalman Filter to perform the structure estimation from an image sequence captured by a perspective camera. In order to theoretically validate our approach, we prove that the proposed model is locally weakly observable. Finally, we report quantitative results on both synthetic data and on real sequences from the CMU Mocap dataset.
Spyridon Leonardos, Xiaowei Zhou 0001, Kostas Daniilidis
ICRA3
2016 Monocular 3D tracking of deformable surfaces
abstract
The problem of reconstructing deformable 3D surfaces has been studied in the non-rigid structure from motion context, where either tracked points over long sequences or an initial 3D shape are required, and also with piecewise methods, where the deformable surface is modeled as a triangulated mesh, which is fitted to an initial estimation of the 3D surface computed from correspondences in two views. In this paper we present a new scheme to reconstruct deformable surfaces by tracking relevant features that parametrize such deformation. Assuming that an initial 3D shape related to a reference frame is available, we initially match the reference and current frames using visual information. Then, these correspondences are clustered in patches with geometric characteristics in the image domain and 3D space. In order to reduce the number of parameters to be estimated, we explain each cluster using thin-plate splines (TPS) with a minimal number of control points. Then the 3D coordinates of these control points in the deformed surface are estimated using a non-linear least squares approach, deriving on the reconstruction of the full deformed patches. We perform experiments in synthetic and real data of monocular video sequences to validate our approach.
Luis Puig, Kostas Daniilidis
ICRA2
2016 A triangle histogram for object classification by tactile sensing
abstract
We present a new descriptor for tactile 3D object classification. It is invariant to object movement and simple to construct, using only the relative geometry of points on the object surface. We demonstrate successful classification of 185 objects in 10 categories, at sparse to dense surface sampling rate in point cloud simulation, with an accuracy of 77.5% at the sparsest and 90.1% at the densest. In a physics-based simulation, we show that contact clouds resembling the object shape can be obtained by a series of gripper closures using a robotic hand equipped with sparse tactile arrays. Despite sparser sampling of the object's surface, classification still performs well, at 74.7%. On a real robot, we show the ability of the descriptor to discriminate among different object instances, using data collected by a tactile hand.
Mabel M. Zhang, Monroe Kennedy III, M. Ani Hsieh, Kostas Daniilidis
IROS4
2015 Reconstruction of 3D Pose for Surfaces of Revolution from Range Data
abstract
Axial symmetry is a common property of everyday objects. Bottles, cups, cans and bowls, all usually fall in that category and can be modeled by surfaces of revolution (SOR). In this paper, we address the problem of estimating the parameters of an SOR (axis and generatrix) from range data. Although SOR reconstruction from RGB images is well studied, previous works using depth measurements are limited. We propose a formulation similar to the 3D registration problem and our solution is based on an alternating procedure that recovers the complete surface geometry, i.e. The axis and the profile curve of the SOR. We evaluate our method both quantitatively and qualitatively using four different datasets that provide depth images from a large variety of axially symmetric objects.
Georgios Pavlakos, Kostas Daniilidis
3DV2
2015 A metric parametrization for trifocal tensors with non-colinear pinholes
abstract
The trifocal tensor, which describes the relation between projections of points and lines in three views, is a fundamental entity of geometric computer vision. In this work, we investigate a new parametrization of the trifocal tensor for calibrated cameras with non-colinear pinholes obtained from a quotient Riemannian manifold. We incorporate this formulation into state-of-the art methods for optimization on manifolds, and show, through experiments in pose averaging, that it produces a meaningful way to measure distances between trifocal tensors.
Spyridon Leonardos, Roberto Tron, Kostas Daniilidis
CVPR3
2015 3D shape estimation from 2D landmarks: A convex relaxation approach
abstract
We investigate the problem of estimating the 3D shape of an object, given a set of 2D landmarks in a single image. To alleviate the reconstruction ambiguity, a widely-used approach is to confine the unknown 3D shape within a shape space built upon existing shapes. While this approach has proven to be successful in various applications, a challenging issue remains, i.e., the joint estimation of shape parameters and camera-pose parameters requires to solve a nonconvex optimization problem. The existing methods often adopt an alternating minimization scheme to locally update the parameters, and consequently the solution is sensitive to initialization. In this paper, we propose a convex formulation to address this problem and develop an efficient algorithm to solve the proposed convex program. We demonstrate the exact recovery property of the proposed method, its merits compared to alternative methods, and the applicability in human pose and car shape estimation.
Xiaowei Zhou 0001, Spyridon Leonardos, Kostas Daniilidis
CVPR4
2015 Multi-image Matching via Fast Alternating Minimization
abstract
In this paper we propose a global optimization-based approach to jointly matching a set of images. The estimated correspondences simultaneously maximize pairwise feature affinities and cycle consistency across multiple images. Unlike previous convex methods relying on semidefinite programming, we formulate the problem as a low-rank matrix recovery problem and show that the desired semidefiniteness of a solution can be spontaneously fulfilled. The low-rank formulation enables us to derive a fast alternating minimization algorithm in order to handle practical problems with thousands of features. Both simulation and real experiments demonstrate that the proposed algorithm can achieve a competitive performance with an order of magnitude speedup compared to the state-of-the-art algorithm. In the end, we demonstrate the applicability of the proposed method to match the images of different object instances and as a result the potential to reconstruct category-specific object models from those images.
Xiaowei Zhou 0001, Menglong Zhu, Kostas Daniilidis
ICCV3
2015 Single Image Pop-Up from Discriminatively Learned Parts
abstract
We introduce a new approach for estimating a fine grained 3D shape and continuous pose of an object from a single image. Given a training set of view exemplars, we learn and select appearance-based discriminative parts which are mapped onto the 3D model through a facility location optimization. The training set of 3D models is summarized into a set of basis shapes from which we can generalize by linear combination. Given a test image, we detect hypotheses for each part. The main challenge is to select from these hypotheses and compute the 3D pose and shape coefficients at the same time. To achieve this, we optimize a function that considers simultaneously the appearance matching of the parts as well as the geometric reprojection error. We apply the alternating direction method of multipliers (ADMM) to minimize the resulting convex function. Our main and novel contribution is the simultaneous solution for part localization and detailed 3D geometry estimation by maximizing both appearance and geometric compatibility with convex relaxation.
Menglong Zhu, Xiaowei Zhou 0001, Kostas Daniilidis
ICCV3
2015 Decentralized active information acquisition: Theory and application to multi-robot SLAM
abstract
This paper addresses the problem of controlling mobile sensing systems to improve the accuracy and efficiency of gathering information autonomously. It applies to scenarios such as environmental monitoring, search and rescue, surveillance and reconnaissance, and simultaneous localization and mapping (SLAM). A multi-sensor active information acquisition problem, capturing the common characteristics of these scenarios, is formulated. The goal is to design sensor control policies which minimize the entropy of the estimation task, conditioned on the future measurements. First, we provide a non-greedy centralized solution, which is computationally fast, since it exploits linearized sensing models, and memory efficient, since it exploits sparsity in the environment model. Next, we decentralize the control task to obtain linear complexity in the number of sensors and provide suboptimality guarantees. Finally, our algorithms are applied to the multi-robot active SLAM problem to enable a decentralized nonmyopic solution that exploits sparsity in the planning process.
Nikolay Atanasov 0001, Jerome Le Ny, Kostas Daniilidis, George J. Pappas
ICRA3
2015 Initialization techniques for 3D SLAM: A survey on rotation estimation and its use in pose graph optimization
abstract
Pose graph optimization is the non-convex optimization problem underlying pose-based Simultaneous Localization and Mapping (SLAM). If robot orientations were known, pose graph optimization would be a linear least-squares problem, whose solution can be computed efficiently and reliably. Since rotations are the actual reason why SLAM is a difficult problem, in this work we survey techniques for 3D rotation estimation. Rotation estimation has a rich history in three scientific communities: robotics, computer vision, and control theory. We review relevant contributions across these communities, assess their practical use in the SLAM domain, and benchmark their performance on representative SLAM problems (Fig. 1). We show that the use of rotation estimation to bootstrap iterative pose graph solvers entails significant boost in convergence speed and robustness.
Luca Carlone, Roberto Tron, Kostas Daniilidis, Frank Dellaert
ICRA3
2015 Online self-supervised monocular visual odometry for ground vehicles
abstract
This paper presents an online self-supervised approach to monocular visual odometry and ground classification applied to ground vehicles. We solve the motion and structure problem based on a constrained kinematic model. The true scale of the monocular scene is recovered by estimating the ground surface. We consider a general parametric ground surface model and use the Random Sample Consensus (RANSAC) algorithm for robust fitting of the parameters. The estimated ground surface provides training samples to learn a probabilistic appearance-based ground classifier in an online and self-supervised manner. The appearance-based classifier is then used to bias the RANSAC sampling to generate better hypotheses for parameter estimation of the ground surface model. Thus, without relying on any prior information, we combine geometric estimates with appearance-based classification to achieve an online self-learning scheme from monocular vision. Experimental results demonstrate that online learning improves the computational efficiency and accuracy compared to standard sampling in RANSAC. Evaluations on the KITTI benchmark dataset demonstrate the stability and accuracy of our overall methods in comparison to previous approaches.
Bhoram Lee, Kostas Daniilidis, Daniel D. Lee
ICRA2
2015 Grasping surfaces of revolution: Simultaneous pose and shape recovery from two views
abstract
In many scenarios, robots encounter rotationally symmetric objects for which no known 3D model exists. To be able to grasp such objects using existing grasp point computation schemes, an estimate of their 3D-pose and shape is necessary. In this paper, we address the problem of recovering 3D-pose and shape of an unknown surface of revolution from two perspective views of known relative orientation. We propose a new algorithm for simultaneous estimation of pose and shape without making use of any cross-sections or bi-tangent points needed by other approaches. Our algorithm builds upon existing single-view SOR reconstruction approaches and couples the pose estimation and reconstruction process. Pose is optimized to minimize discrepancies between reconstructed shapes from two views. Our method works even in the presence of only one of the two apparent contours of a surface of revolution. We test our algorithm by comparing to ground-truth poses and shapes as well as performing grasping experiments. We introduce a new dataset of rotationally symmetric objects in a variety of poses and backgrounds and with a measured ground-truth pose and shape.
Cody J. Phillips 0001, Matthieu Lecce, Casey Davis, Kostas Daniilidis
ICRA4
2015 MSG-cal: Multi-sensor graph-based calibration
abstract
We present a system for determining a global solution for the relative poses between multiple sensors with different modalities and varying fields of view. The final calibration result produces a tree of transforms rooted at a single sensor that allows the fusion of the sensor streams into a shared coordinate frame. The method differs from other approaches by handling any number of sensors with only minimal constraints on their fields of view, producing a global solution that is better than any pairwise solution, and by simplifying the data collection process through automatic data association.
Jason L. Owens, Philip R. Osteen, Kostas Daniilidis
IROS3
2014 Geometric Urban Geo-localization
abstract
We propose a purely geometric correspondence-free approach to urban geo-localization using 3D point-ray features extracted from the Digital Elevation Map of an urban environment. We derive a novel formulation for estimating the camera pose locus using 3D-to-2D correspondence of a single point and a single direction alone. We show how this allows us to compute putative correspondences between building corners in the DEM and the query image by exhaustively combining pairs of point-ray features. Then, we employ the two-point method to estimate both the camera pose and compute correspondences between buildings in the DEM and the query image. Finally, we show that the computed camera poses can be efficiently ranked by a simple skyline projection step using building edges from the DEM. Our experimental evaluation illustrates the promise of a purely geometric approach to the urban geo-localization problem.
Mayank Bansal, Kostas Daniilidis
CVPR2
2014 On the Quotient Representation for the Essential Manifold
abstract
The essential matrix, which encodes the epipolar constraint between points in two projective views, is a cornerstone of modern computer vision. Previous works have proposed different characterizations of the space of essential matrices as a Riemannian manifold. However, they either do not consider the symmetric role played by the two views, or do not fully take into account the geometric peculiarities of the epipolar constraint. We address these limitations with a characterization as a quotient manifold which can be easily interpreted in terms of camera poses. While our main focus in on theoretical aspects, we include experiments in pose averaging, and show that the proposed formulation produces a meaningful distance between essential matrices.
Roberto Tron, Kostas Daniilidis
CVPR2
2014 Statistical Pose Averaging with Non-isotropic and Incomplete Relative Measurements
Roberto Tron, Kostas Daniilidis
ECCV (5)2
2014 Active Deformable Part Models Inference
Menglong Zhu, Nikolay Atanasov 0001, George J. Pappas, Kostas Daniilidis
ECCV (7)4
2014 Information acquisition with sensing robots: Algorithms and error bounds
abstract
Utilizing the capabilities of configurable sensing systems requires addressing difficult information gathering problems. Near-optimal approaches exist for sensing systems without internal states. However, when it comes to optimizing the trajectories of mobile sensors the solutions are often greedy and rarely provide performance guarantees. Notably, under linear Gaussian assumptions, the problem becomes deterministic and can be solved off-line. Approaches based on submodularity have been applied by ignoring the sensor dynamics and greedily selecting informative locations in the environment. This paper presents a non-greedy algorithm with suboptimality guarantees, which relies on concavity instead of submodularity and takes the sensor dynamics into account. Coupled with linearization and model predictive control, the algorithm can be used to generate adaptive policies for mobile sensors with non-linear sensing models. Applications in gas concentration mapping and target tracking are presented.
Nikolay Atanasov 0001, Jerome Le Ny, Kostas Daniilidis, George J. Pappas
ICRA3
2014 Vision-based control of a quadrotor for perching on lines
abstract
We formulate the position-based visual servoing problem for a quadrotor equipped with a monocular camera and an IMU relying only on features on planes and lines in order to fly above and perch on arbitrarily oriented lines. We show that we are able to compute the orientation of an arbitrarily oriented line, the speed of the robot and its position with respect to the target line using two points at a known distance on the line. The direction of the velocity is derived from optical flow induced by features on a plane in the background Finally, we demonstrate fully autonomous flight and perching using a small 230 gram quadrotor with all the computations running on the robot.
Kartik Mohta, Vijay Kumar 0001, Kostas Daniilidis
ICRA3
2014 An optimization approach to bearing-only visual homing with applications to a 2-D unicycle model
abstract
We consider the problem of bearing-based visual homing: Given a mobile robot which can measure bearing directions corresponding to known landmarks, the goal is to guide the robot toward a desired “home” location. We propose a control law based on the gradient field of a Lyapunov function, and give sufficient conditions for global convergence. We show that the well-known Average Landmark Vector method (for which no convergence proof was known) can be obtained as a particular case of our framework. We then derive a sliding mode control law for a unicycle model which follows this gradient field. Both controllers do not depend on range information. Finally, we also show how our framework can be used to characterize the sensitivity of a home location with respect to noise in the specified bearings.
Roberto Tron, Kostas Daniilidis
ICRA2
2014 Single image 3D object detection and pose estimation for grasping
abstract
We present a novel approach for detecting objects and estimating their 3D pose in single images of cluttered scenes. Objects are given in terms of 3D models without accompanying texture cues. A deformable parts-based model is trained on clusters of silhouettes of similar poses and produces hypotheses about possible object locations at test time. Objects are simultaneously segmented and verified inside each hypothesis bounding region by selecting the set of superpixels whose collective shape matches the model silhouette. A final iteration on the 6-DOF object pose minimizes the distance between the selected image contours and the actual projection of the 3D model. We demonstrate successful grasps using our detection and pose estimate with a PR2 robot. Extensive evaluation with a novel ground truth dataset shows the considerable benefit of using shape-driven cues for detecting objects in heavily cluttered scenes.
Menglong Zhu, Konstantinos G. Derpanis, Yinfei Yang, Samarth Brahmbhatt, Mabel M. Zhang, Cody J. Phillips 0001, Matthieu Lecce, Kostas Daniilidis
ICRA8
2014 MEVO: Multi-environment stereo visual odometry
abstract
The ego motion estimation from an image sequence, commonly known as visual odometry, has been thoroughly studied in recent years. Different solutions have been developed depending on the particular scenario the system interacts in. In highly textured environments point features are abundant and visual odometry approaches focus on complementary steps, such as sparse bundle adjustment or keyframe techniques, to improve the accuracy of the motion estimation. In textureless scenarios, the absence of point features motivates the use of different image features. Lines have proven to be an interesting alternative to points in man-made environments, but very few visual odometry approaches have been developed using these types of features. Moreover, the combination of point and line features has not been considered in the development of real-time visual odometry algorithms. In this paper, we explore the combination of point and line features to robustly compute the six degree of freedom motion transformation between consecutive stereo frames. Additionally, we deal with the problem of line stereo matching, since our approach is based on 3D-2D correspondences to estimate motion. We develop an efficient algorithm to compute the stereo line matching, even in situations where one of the endpoints describing the line segment in the left image is not visible in the right image. Several experiments with synthetic and real image sequences show that a simple but effective combination of point and line features improves the motion estimate compared to approaches using only one type of these features with a slight increase in computational cost.
Thomas Koletschka, Luis Puig, Kostas Daniilidis
IROS3
2014 Scale Space for Camera Invariant Features
abstract
In this paper we propose a new approach to compute the scale space of any central projection system, such as catadioptric, fisheye or conventional cameras. Since these systems can be explained using a unified model, the single parameter that defines each type of system is used to automatically compute the corresponding Riemannian metric. This metric, is combined with the partial differential equations framework on manifolds, allows us to compute the Laplace-Beltrami (LB) operator, enabling the computation of the scale space of any central projection system. Scale space is essential for the intrinsic scale selection and neighborhood description in features like SIFT. We perform experiments with synthetic and real images to validate the generalization of our approach to any central projection system. We compare our approach with the best-existing methods showing competitive results in all type of cameras: catadioptric, fisheye, and perspective.
Luis Puig, Josechu J. Guerrero, Kostas Daniilidis
IEEE Trans. Pattern Anal. Mach. Intell.3
2014 Nonmyopic View Planning for Active Object Classification and Pose Estimation
abstract
One of the central problems in computer vision is the detection of semantically important objects and the estimation of their pose. Most of the work in object detection has been based on single image processing, and its performance is limited by occlusions and ambiguity in appearance and geometry. This paper proposes an active approach to object detection in which the point of view of a mobile depth camera is controlled. When an initial static detection phase identifies an object of interest, several hypotheses are made about its class and orientation. Then, a sequence of views, which balances the amount of energy used to move the sensor with the chance of identifying the correct hypothesis, is planned. We formulate an active hypothesis testing problem, which includes sensor mobility, and solve it using a point-based approximate partially observable Markov decision process algorithm. The validity of our approach is verified through simulation and realworld experiments with the PR2 robot. The results suggest that the approach outperforms the widely used greedy viewpoint selection and provides a significant improvement over static object detection.
Nikolay Atanasov 0001, Bharath Sankaran, Jerome Le Ny, George J. Pappas, Kostas Daniilidis
IEEE Trans. Robotics5
2013 Joint Spectral Correspondence for Disparate Image Matching
abstract
We address the problem of matching images with disparate appearance arising from factors like dramatic illumination (day vs. night), age (historic vs. new) and rendering style differences. The lack of local intensity or gradient patterns in these images makes the application of pixel-level descriptors like SIFT infeasible. We propose a novel formulation for detecting and matching persistent features between such images by analyzing the eigen-spectrum of the joint image graph constructed from all the pixels in the two images. We show experimental results of our approach on a public dataset of challenging image pairs and demonstrate significant performance improvements over state-of-the-art.
Mayank Bansal, Kostas Daniilidis
CVPR2
2013 Hypothesis testing framework for active object detection
abstract
One of the central problems in computer vision is the detection of semantically important objects and the estimation of their pose. Most of the work in object detection has been based on single image processing and its performance is limited by occlusions and ambiguity in appearance and geometry. This paper proposes an active approach to object detection by controlling the point of view of a mobile depth camera. When an initial static detection phase identifies an object of interest, several hypotheses are made about its class and orientation. The sensor then plans a sequence of viewpoints, which balances the amount of energy used to move with the chance of identifying the correct hypothesis. We formulate an active M-ary hypothesis testing problem, which includes sensor mobility, and solve it using a point-based approximate POMDP algorithm. The validity of our approach is verified through simulation and experiments with real scenes captured by a kinect sensor. The results suggest a significant improvement over static object detection.
Nikolay Atanasov 0001, Bharath Sankaran, Jerome Le Ny, Thomas Koletschka, George J. Pappas, Kostas Daniilidis
ICRA6
2013 Novel Representations, Methods, and Algorithms in Computer Vision
Kostas Daniilidis, Petros Maragos, Nikos Paragios
Int. J. Comput. Vis.1
2012 Dynamic scene understanding: The role of orientation features in space and time in scene classification
abstract
Natural scene classification is a fundamental challenge in computer vision. By far, the majority of studies have limited their scope to scenes from single image stills and thereby ignore potentially informative temporal cues. The current paper is concerned with determining the degree of performance gain in considering short videos for recognizing natural scenes. Towards this end, the impact of multiscale orientation measurements on scene classification is systematically investigated, as related to: (i) spatial appearance, (ii) temporal dynamics and (iii) joint spatial appearance and dynamics. These measurements in visual space, x-y, and spacetime, x-y-t, are recovered by a bank of spatiotemporal oriented energy filters. In addition, a new data set is introduced that contains 420 image sequences spanning fourteen scene categories, with temporal scene information due to objects and surfaces decoupled from camera-induced ones. This data set is used to evaluate classification performance of the various orientation-related representations, as well as state-of-the-art alternatives. It is shown that a notable performance increase is realized by spatiotemporal approaches in comparison to purely spatial or purely temporal methods.
Konstantinos G. Derpanis, Matthieu Lecce, Kostas Daniilidis, Richard P. Wildes
CVPR3
2012 Robot localization using soft object detection
abstract
In this paper, we give a new double twist to the robot localization problem. We solve the problem for the case of prior maps which are semantically annotated perhaps even sketched by hand. Data association is achieved not through the detection of visual features but the detection of object classes used in the annotation of the prior maps. To avoid the caveats of general object recognition, we propose a new representation of the query images that consists of a vector of the detection scores for each object class. Given such soft object detections we are able to create hypotheses about pose and to refine them through particle filtering. As opposed to small confined office and kitchen spaces, our experiment takes place in a large open urban rail station with multiple semantically ambiguous places. The success of our approach shows that our new representation is a robust way to exploit the plethora of existing prior maps for GPS-denied environments avoiding the data association problems when matching point clouds or visual features.
Roy Anati, Davide Scaramuzza 0001, Konstantinos G. Derpanis, Kostas Daniilidis
ICRA4
2012 Identifying maximal rigid components in bearing-based localization
abstract
We present an approach for sensor network localization when provided with a set of angular constraints. This problem arises in camera networks when angles between nearby points can be measured but depth measurements are not readily available. We provide contributions for two different variations on this problem. First, when each node is aware of a global coordinate frame, we present a novel method for identifying the components of the problem that are rigidly constrained. Second, in the more difficult case where only relative angles are known, we propose a novel spectral solution that achieves a globally-optimal embedding under transitively-triangular constraints, which we show encompass a wide range of real-world conditions. We demonstrate the utility of our algorithm on both synthetic data and data from quadrotor robot formations.
Ryan Kennedy, Kostas Daniilidis, Oleg Naroditsky, Camillo J. Taylor
IROS2
2012 Optimal pixel aspect ratio for enhanced 3D TV visualization
Hossein Azari, Irene Cheng 0001, Kostas Daniilidis, Anup Basu
Comput. Vis. Image Underst.3
2012 Reconstructing and analyzing periodic human motion from stationary monocular views
Evan Ribnick, Ravishankar Sivalingam, Nikolaos Papanikolopoulos, Kostas Daniilidis
Comput. Vis. Image Underst.4
2012 Shape-Based Object Detection via Boundary Structure Segmentation
Alexander Toshev, Ben Taskar, Kostas Daniilidis
Int. J. Comput. Vis.3
2012 Two Efficient Solutions for Visual Odometry Using Directional Correspondence
abstract
This paper presents two new, efficient solutions to the two-view, relative pose problem from three image point correspondences and one common reference direction. This three-plus-one problem can be used either as a substitute for the classic five-point algorithm, using a vanishing point for the reference direction, or to make use of an inertial measurement unit commonly available on robots and mobile devices where the gravity vector becomes the reference direction. We provide a simple, closed-form solution and a solution based on algebraic geometry which offers numerical advantages. In addition, we introduce a new method for computing visual odometry with RANSAC and four point correspondences per hypothesis. In a set of real experiments, we demonstrate the power of our approach by comparing it to the five-point method in a hypothesize-and-test visual odometry setting.
Oleg Naroditsky, Xun S. Zhou, Jean H. Gallier, Stergios I. Roumeliotis, Kostas Daniilidis
IEEE Trans. Pattern Anal. Mach. Intell.5
2011 Optimizing polynomial solvers for minimal geometry problems
abstract
In recent years polynomial solvers based on algebraic geometry techniques, and specifically the action matrix method, have become popular for solving minimal problems in computer vision. In this paper we develop a new method for reducing the computational time and improving numerical stability of algorithms using this method. To achieve this, we propose and prove a set of algebraic conditions which allow us to reduce the size of the elimination template (polynomial coefficient matrix), which leads to faster LU or QR decomposition. Our technique is generic and has potential to improve performance of many solvers that use the action matrix method. We demonstrate the approach on specific examples, including an image stitching algorithm where computation time is halved and single precision arithmetic can be used.
Oleg Naroditsky, Kostas Daniilidis
ICCV2
2011 Automatic alignment of a camera with a line scan LIDAR system
abstract
We propose a new method for extrinsic calibration of a line-scan LIDAR with a perspective projection camera. Our method is a closed-form, minimal solution to the problem. The solution is a symbolic template found via variable elimination and the multi-polynomial Macaulay resultant. It does not require initialization, and can be used in an automatic calibration setting when paired with RANSAC and least-squares refinement. We show the efficacy of our approach through a set of simulations and a real calibration.
Oleg Naroditsky, Alexander Patterson, Kostas Daniilidis
ICRA3
2011 Exploiting motion priors in visual odometry for vehicle-mounted cameras with non-holonomic constraints
abstract
This paper presents a new method to estimate the relative motion of a vehicle from images of a single camera. The biggest problem in visual motion estimation is data association; matched points contain many outliers that must be detected and removed so that the motion can be estimated accurately. A very established method for robust motion estimation in the presence of outliers is the five-point RANSAC algorithm. Five-point RANSAC operates by generating motion hypotheses from randomly-sampled minimal sets of five-point correspondences. These hypotheses are then tested against all data points and the motion hypothesis that after a given number of iterations returns the largest number of inliers is taken as the solution to the problem. A typical drawback of RANSAC is that the number of iterations required to find a suitable solution grows exponentially with the number of outliers, often requiring thousands of iterations for typical data from urban environments. Another problem is that - due to its random nature - sometimes the found solution is not the “best” solution to the motion estimation problem. In this paper, we describe an algorithm for relative motion estimation in the presence of outliers, which does not rely on RANSAC. Contrary to RANSAC, motion hypotheses are not generated from randomly-sampled point correspondences, but from a “proposal distribution” that is built by exploiting the vehicle non-holonomic constraints. We show that not only is the proposed algorithm significantly faster than RANSAC, but that the returned solution may also be better in that it favors the underlying motion model of the vehicle, thus overcoming the typical limitations of RANSAC. Additionally, the proposed algorithm provides the likelihood of the motion estimate, which can be very useful in all those applications where a probability distribution of the position of the vehicle is required (e.g., SLAM). Finally, the performance of the proposed method is compared to that of the standard five-point RANSAC on real images collected from a vehicle moving in a cluttered, urban environment.
Davide Scaramuzza 0001, Andrea Censi, Kostas Daniilidis
IROS3
2011 Geo-localization of street views with aerial image databases
abstract
We study the feasibility of solving the challenging problem of geolocalizing ground level images in urban areas with respect to a database of images captured from the air such as satellite and oblique aerial images. We observe that comprehensive aerial image databases are widely available while complete coverage of urban areas from the ground is at best spotty. As a result, localization of ground level imagery with respect to aerial collections is a technically important and practically significant problem. We exploit two key insights: (1) satellite image to oblique aerial image correspondences are used to extract building facades, and (2) building facades are matched between oblique aerial and ground images for geo-localization. Key contributions include: (1) A novel method for extracting building facades using building outlines; (2) Correspondence of building facades between oblique aerial and ground images without direct matching; and (3) Position and orientation estimation of ground images. We show results of ground image localization in a dense urban area.
Mayank Bansal, Harpreet Sawhney, Kostas Daniilidis
ACM Multimedia4
2010 Object detection via boundary structure segmentation
abstract
We address the problem of object detection and segmentation using holistic properties of object shape. Global shape representations are highly susceptible to clutter inevitably present in realistic images, and can be robustly recognized only using a precise segmentation of the object. To this end, we propose a figure/ground segmentation method for extraction of image regions that resemble the global properties of a model boundary structure and are perceptually salient. Our shape representation, called the chordiogram, is based on geometric relationships of object boundary edges, while the perceptual saliency cues we use favor coherent regions distinct from the background. We formulate the segmentation problem as an integer quadratic program and use a semidefinite programming relaxation to solve it. Obtained solutions provide the segmentation of an object as well as a detection score used for object recognition. Our single-step approach improves over state of the art methods on several object detection and segmentation benchmarks.
Alexander Toshev, Ben Taskar, Kostas Daniilidis
CVPR3
2010 Spherical Correlation of Visual Representations for 3D Model Retrieval
Ameesh Makadia, Kostas Daniilidis
Int. J. Comput. Vis.2
2009 Shape-based object recognition in videos using 3D synthetic object models
abstract
In this paper we address the problem of recognizing moving objects in videos by utilizing synthetic 3D models. We use only the silhouette space of the synthetic models making thus our approach independent of appearance. To deal with the decrease in discriminability in the absence of appearance, we align sequences of object masks from video frames to paths in silhouette space. We extract object silhouettes from video by an integration of feature tracking, motion grouping of tracks, and co-segmentation of successive frames. Subsequently, the object masks from the video are matched to 3D model silhouettes in a robust matching and alignment phase. The result is a matching score for every 3D model to the video, along with a pose alignment of the model to the video. Promising experimental results indicate that a purely shape-based matching scheme driven by synthetic 3D models can be successfully applied for object recognition in videos.
Alexander Toshev, Ameesh Makadia, Kostas Daniilidis
CVPR3
2009 Constructing Topological Maps using Markov Random Fields and Loop-Closure Detection
abstract
We present a system which constructs a topological map of an environment given a sequence of images. This system includes a novel image similarity score which uses dynamic programming to match images using both the appearance and relative positions of local features simultaneously. Additionally an MRF is constructed to model the probability of loop-closures. A locally optimal labeling is found using Loopy-BP. Finally we outline a method to generate a topological map from loop closure data. Results are presented on four urban sequences and one indoor sequence.
Roy Anati, Kostas Daniilidis
NIPS2
2009 Vision-Based Localization for Leader-Follower Formation Control
abstract
This paper deals with vision-based localization for leader–follower formation control. Each unicycle robot is equipped with a panoramic camera that only provides the view angle to the other robots. The localization problem is studied using a new observability condition valid for general nonlinear systems and based on the extended output Jacobian. This allows us to identify those robot motions that preserve the system observability and those that render it nonobservable. The state of the leader–follower system is estimated via the extended Kalman filter, and an input-state feedback control law is designed to stabilize the formation. Simulations and real-data experiments confirm the theoretical results and show the effectiveness of the proposed formation control.
Gian Luca Mariottini, Fabio Morbidi, Domenico Prattichizzo, Nicholas Vander Valk, Nathan Michael, George J. Pappas, Kostas Daniilidis
IEEE Trans. Robotics7
2009 Vision-Based, Distributed Control Laws for Motion Coordination of Nonholonomic Robots
abstract
In this paper, we study the problem of distributed motion coordination among a group of nonholonomic ground robots. We develop vision-based control laws for parallel and balanced circular formations using a consensus approach. The proposed control laws are distributed in the sense that they require information only from neighboring robots. Furthermore, the control laws are coordinate-free and do not rely on measurement or communication of heading information among neighbors but instead require measurements of bearing, optical flow, and time to collision, all of which can be measured using visual sensors. Collision-avoidance capabilities are added to the team members, and the effectiveness of the control laws are demonstrated on a group of mobile robots.
Nima Moshtagh, Nathan Michael, Ali Jadbabaie, Kostas Daniilidis
IEEE Trans. Robotics4
2008 Object Detection from Large-Scale 3D Datasets Using Bottom-Up and Top-Down Descriptors
Alexander Patterson, Philippos Mordohai, Kostas Daniilidis
ECCV (4)3
2008 Monocular visual odometry in urban environments using an omnidirectional camera
abstract
We present a system for monocular simultaneous localization and mapping (mono-SLAM) relying solely on video input. Our algorithm makes it possible to precisely estimate the camera trajectory without relying on any motion model. The estimation is completely incremental: at a given time frame, only the current location is estimated while the previous camera positions are never modified. In particular, we do not perform any simultaneous iterative optimization of the camera positions and estimated 3D structure (local bundle adjustment). The key aspect of the system is a fast and simple pose estimation algorithm that uses information not only from the estimated 3D map, but also from the epipolar constraint. We show that the latter leads to a much more stable estimation of the camera trajectory than the conventional approach. We perform high precision camera trajectory estimation in urban scenes with a large amount of clutter. Using an omnidirectional camera placed on a vehicle, we cover one of the longest distance ever reported, up to 2.5 kilometers.
Jean-Philippe Tardif, Yanis Pavlidis, Kostas Daniilidis
IROS3
2007 A Single-perspective Novel Panoramic View from Radially Distorted Non-central Images
abstract
In this paper, we propose an image-based technique for panoramic novelview generation using three uncalibrated wide-angle images as its input. State of the art in novel view generation presumes the calibration and removal of radial distortion or any other deformation resulting from t he geometry of a non-central camera. We propose a method which replaces this calibration with the assumption that the epipole corresponding to the novel viewpoint is at the center of radial distortion and that it is known. 1 Introduction In novel view synthesis, given an image pair, the intensity of each ray in the novel view is determined by finding corresponding point matches betwee n those images, using the epipolar geometries between the two views and with respect to the novel view. This process requires that the given images obey the standard perspective model, with a single viewpoint and no radial distortion. Such assumptions make the creation of novel panoramas a process of acquiring tens of images, each with relatively narrow field-of-view, with special apparatus, and applying calibration and estimatio n of circular motion [9]. In this paper, we propose an approach where a novel view can be synthesized from multiple views which might be highly radially distorted, and even non-central, without compensating explicitly for the resulting image deformations. We shed new light on the problem (Fig. 1) of “how a scene looks from a scene point?”[8] by eliminating the need for a reference plane and facilitating synthesis of omnidre ctional novel views. We assume that, for n ≥ 3 views, the location of the epipole of the novel view in each image is known, and, in the case of radially distorted views, that it is coincident with the centre of radial distortion in the image. This assumption is either enforced (for example by choosing a visible scene point as the centre of projection of the view to be synthesized, and actually fixating on this point in the case of radially dis torted views), or else it may naturally be satisfied for certain wide-angle imaging syste ms (for example multiple views of a spherical mirror). Given such a configuration, the 2D sta r of virtual rays at the novel viewpoint are mapped to a 1D star of lines in each image, which we model as a projection from P 2 (the domain of the novel view) to P 1 (the angles of the line pencil in image plane). We make use of the well known fact that three projections of a line in space (a virtual line in our case) yield a trifocal constraint, from w hich we extract the projection matrices required for the view synthesis. The main contributions of this paper are:
Rana Molana, Kostas Daniilidis
BMVC2
2007 Image Matching via Saliency Region Correspondences
abstract
We introduce the notion of co-saliency for image matching. Our matching algorithm combines the discriminative power of feature correspondences with the descriptive power of matching segments. Co-saliency matching score favors correspondences that are consistent with 'soft' image segmentation as well as with local point feature matching. We express the matching model via a joint image graph (JIG) whose edge weights represent intra-as well as inter-image relations. The dominant spectral components of this graph lead to simultaneous pixel-wise alignment of the images and saliency-based synchronization of 'soft' image segmentation. The co-saliency score function, which characterizes these spectral components, can be directly used as a similarity metric as well as a positive feedback for updating and establishing new point correspondences. We present experiments showing the extraction of matching regions and pointwise correspondences, and the utility of the global image similarity in the context of place recognition.
Alexander Toshev, Jianbo Shi, Kostas Daniilidis
CVPR3
2007 Scale-Invariant Features on the Sphere
abstract
This paper considers an application of scale-invariant feature detection using scale-space analysis suitable for use with wide field of view cameras. Rather than obtain scale- space images via convolution with the Gaussian function on the image plane, we map the image to the sphere and obtain scale-space images as the solution to the heat (diffusion) equation on the sphere which is implemented in the frequency domain using spherical harmonics. The percentage correlation of scale-invariant features that may be matched between any two wide-angle images subject to change in camera pose is then compared using each of these methods. We also present a means by which the required sampling bandwidth may be determined and propose a suitable anti-aliasing filter which may be used when this bandwidth exceeds the maximum permissible due to computational requirements. The results show improved performance using scale-space images obtained as the solution of the diffusion equation on the sphere, with additional improvements observed using the anti-aliasing filter.
Peter Hansen 0003, Peter I. Corke, Wageeh W. Boles, Kostas Daniilidis
ICCV4
2007 Leader-Follower Formations: Uncalibrated Vision-Based Localization and Control
abstract
This paper focuses on leader-follower formations of mobile robots equipped with panoramic cameras and extend earlier works in the literature addressing both the vision-based localization and control problems. First, a new sufficient analytical condition for localizability is proved and used to shed light on the geometrical meaning of formation localization using uncalibrated vision sensors, here performed with the unscented Kalman filter. Second, we design a feedback control law based on dynamic extension in order to extend the applicability of our control scheme also to the case of distant robots.
Gian Luca Mariottini, Fabio Morbidi, Domenico Prattichizzo, George J. Pappas, Kostas Daniilidis
ICRA5
2007 Scale invariant feature matching with wide angle images
abstract
Numerous scale-invariant feature matching algorithms using scale-space analysis have been proposed for use with perspective cameras, where scale-space is defined as convolution with a Gaussian. The contribution of this work is a method suitable for use with wide angle cameras. Given an input image, we map it to the unit sphere and obtain scale-space images by convolution with the solution of the spherical diffusion equation on the sphere which we implement in the spherical Fourier domain. Using such an approach, the scale-space response of a point in space is independent of its position on the image plane for a camera subject to pure rotation. Scale-invariant features are then found as local extrema in scale-space. Given this set of scale-invariant features, we then generate feature descriptors by considering a circular support region defined on the sphere whose size is selected relative to the feature scale. We compare our method to a naive implementation of SIFT where the image is treated as perspective, where our results show an improvement in matching performance.
Peter Hansen 0003, Peter I. Corke, Wageeh W. Boles, Kostas Daniilidis
IROS4
2007 Correspondence-free Structure from Motion
Ameesh Makadia, Christopher Geyer, Kostas Daniilidis
Int. J. Comput. Vis.3
2006 Epipolar Geometry of Central Projection Systems Using Veronese Maps
abstract
We study the epipolar geometry between views acquired by mixtures of central projection systems including catadioptric sensors and cameras with lens distortion. Since the projection models are in general non-linear, a new representation for the geometry of central images is proposed. This representation is the lifting through Veronese maps of the image plane to the 5D projective space. It is shown that, for most sensor combinations, there is a bilinear form relating the lifted coordinates of corresponding image points. We analyze the properties of the embedding and explicitly construct the lifted fundamental matrices in order to understand their structure. The usefulness of the framework is illustrated by estimating the epipolar geometry between images acquired by a paracatadioptric system and a camera with radial distortion.
João Pedro Barreto 0001, Kostas Daniilidis
CVPR (1)2
2006 Structure from Motion with Known Camera Positions
abstract
The wide availability of GPS sensors is changing the landscape in the applications of structure from motion techniques for localization. In this paper, we study the problem of estimating camera orientations from multiple views, given the positions of the viewpoints in a world coordinate system and a set of point correspondences across the views. Given three or more views, the above problem has a finite number of solutions for three or more point correspondences. Given six or more views, the problem has a finite number of solutions for just two or more points. In the three-view case, we show the necessary and sufficient conditions for the three essential matrices to be consistent with a set of known baselines. We also introduce a method to recover the absolute orientations of three views in world coordinates from their essential matrices. To refine these estimates we perform a least-squares minimization on the group cross product SO(3) × SO(3) × SO(3). We report experiments on synthetic data and on data from the ICCV2005 Computer Vision Contest.
Rodrigo L. Carceroni, Ankita Kumar, Kostas Daniilidis
CVPR (1)3
2006 Fully Automatic Registration of 3D Point Clouds
abstract
We propose a novel technique for the registration of 3D point clouds which makes very few assumptions: we avoid any manual rough alignment or the use of landmarks, displacement can be arbitrarily large, and the two point sets can have very little overlap. Crude alignment is achieved by estimation of the 3D-rotation from two Extended Gaussian Images even when the data sets inducing them have partial overlap. The technique is based on the correlation of the two EGIs in the Fourier domain and makes use of the spherical and rotational harmonic transforms. For pairs with low overlap which fail a critical verification step, the rotational alignment can be obtained by the alignment of constellation images generated from the EGIs. Rotationally aligned sets are matched by correlation using the Fourier transform of volumetric functions. A fine alignment is acquired in the final step by running Iterative Closest Points with just few iterations.
Ameesh Makadia, Alexander Patterson, Kostas Daniilidis
CVPR (1)3
2006 Vision-based Control Laws for Distributed Flocking of Nonholonomic Agents
abstract
We study the problem of vision-based flocking and coordination of a group of kinematic agents in 2 and 3 dimensions. It is shown that in the absence of communication among agents, and by using only visual information, a group of mobile agents can align their velocity vectors and move in a formation. A coordinate-free control law is used to develop a vision-based input for each nonholonomic agent. The vision-based input does not rely on heading measurements, but only requires measurements of bearing, optical flow and time-to-collision, all of which can be efficiently measured
Nima Moshtagh, Ali Jadbabaie, Kostas Daniilidis
ICRA3
2006 Rotation Recovery from Spherical Images without Correspondences
abstract
This paper addresses the problem of rotation estimation directly from images defined on the sphere and without correspondence. The method is particularly useful for the alignment of large rotations and has potential impact on 3D shape alignment. The foundation of the method lies in the fact that the spherical harmonic coefficients undergo a unitary mapping when the original image is rotated. The correlation between two images is a function of rotations and we show that it has an SO(3)-Fourier transform equal to the pointwise product of spherical harmonic coefficients of the original images. The resolution of the rotation space depends on the bandwidth we choose for the harmonic expansion and the rotation estimate is found through a direct search in this 3D discretized space. A refinement of the rotation estimate can be obtained from the conservation of harmonic coefficients in the rotational shift theorem. A novel decoupling of the shift theorem with respect to the Euler angles is presented and exploited in an iterative scheme to refine the initial rotation estimates. Experiments show the suitability of the method for large rotations and the dependence of the method on bandwidth and the choice of the spherical harmonic coefficients.
Ameesh Makadia, Kostas Daniilidis
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 Radon-Based Structure from Motion without Correspondences
abstract
We present a novel approach for the estimation of 3D-motion directly from two images using the Radon transform. We assume a similarity function defined on the cross-product of two images which assigns a weight to all feature pairs. This similarity function is integrated over all feature pairs that satisfy the epipolar constraint. This integration is equivalent to filtering the similarity function with a Dirac function embedding the epipolar constraint. The result of this convolution is a function of the five unknown motion parameters with maxima at the positions of compatible rigid motions. The breakthrough is in the realization that the Radon transform is a filtering operator: If we assume that images are defined on spheres and the epipolar constraint is a group action of two rotations on two spheres, then the Radon transform is a convolution/correlation integral. We propose a new algorithm to compute this integral from the spherical harmonics of the similarity and Dirac functions. The resulting resolution in the motion space depends on the bandwidth we keep from the spherical transform. The strength of the algorithm is in avoiding a commitment to correspondences, thus being robust to erroneous feature detection, outliers, and multiple motions. The algorithm has been tested in sequences of real omnidirectional images and it outperforms correspondence-based structure from motion.
Ameesh Makadia, Christopher Geyer, S. Shankar Sastry, Kostas Daniilidis
CVPR (1)4
2005 Fundamental Matrix for Cameras with Radial Distortion
abstract
When deploying a heterogeneous camera network or when we use cheap zoom cameras like in cell-phones, it is not practical, if not impossible to off-line calibrate the radial distortion of each camera using reference objects. It is rather desirable to have an automatic procedure without strong assumptions about the scene. In this paper, we present a new algorithm for estimating the epipolar geometry of two views where the two views can be radially distorted with different distortion factors. It is the first algorithm in the literature solving the case of different distortion in the left and right view linearly and without assuming the existence of lines in the scene. Points in the projective plane are lifted to a quadric in three-dimensional projective space. A radial distortion of the projective plane results to a matrix transformation in the space of lifted coordinates. The new epipolar constraint depends linearly on a 4 /spl times/ 4 radial fundamental matrix which has 9 degrees of freedom. A complete algorithm is presented and tested on real imagery.
João Pedro Barreto 0001, Kostas Daniilidis
ICCV2
2005 Correspondenceless Ego-Motion Estimation Using an IMU
abstract
Mobile robots can be easily equipped with numerous sensors which can aid in the tasks of localization and ego-motion estimation. Two such examples are Inertial Measurement Units (IMU), which provide a gravity vector via pitch and roll angular velocities, and wide-angle or panoramic imaging devices. As the number of powerful devices on a single robot increases, an important problem arises in how to fuse the information coming from multiple sources to obtain an accurate and efficient motion estimate. The IMU provides real-time readings which can be employed in orientation estimation, while in principle an Omnidirectional camera provides enough information to estimate the full rigid motion (up to translational scale). However, in addition to being computationally overwhelming, such an estimation is traditionally based on the sensitive search for feature correspondences between image frames. In this paper we present a novel algorithm that exploits information from an IMU to reduce the five parameter motion search to a three-parameter estimation. For this task we formulate a generalized Hough transform which processes image features directly to avoid searching for correspondences. The Hough space is computed rapidly by re-treating the transform as a convolution of spherical images.
Ameesh Makadia, Kostas Daniilidis
ICRA2
2005 Using skew Gabor filter in source signal separation and local spectral orientation analysis
Weichuan Yu, Gerald Sommer, Kostas Daniilidis, James S. Duncan
Image Vis. Comput.3
2004 Using Skew Gabor Filter in Source Signal Separation and Local Spectral Multi-Orientation Analysis
Weichuan Yu, Gerald Sommer, Kostas Daniilidis
CVPR (1)3
2004 Normalized Cross-Correlation for Spherical Images
Lorenzo Sorgi, Kostas Daniilidis
ECCV (2)2
2004 Hybrid control for visibility-based pursuit-evasion games
abstract
Pursuit-evasion games in complex environments have a rich but disconnected history. Continuous or differential pursuit-evasion games focus on optimal control methods, and rely on very intense computations in order to provide locally optimal controls. Discrete pursuit-evasion games on graphs are algorithmically much more appealing, but completely ignore the physical dynamics of the players, resulting in possibly infeasible motions. In this paper, we present a provable and algorithmically feasible solution for visibility-based pursuit-evasion games in simply-connected environments, for players with dynamic constraints. This is achieved by combining two recent but distant results.
Volkan Isler, Calin Belta, Kostas Daniilidis, George J. Pappas
IROS3
2004 Sampling based sensor-network deployment
abstract
In this paper, we consider the problem of placing networked sensors in a way that guarantees coverage and connectivity. We focus on sampling based deployment and present algorithms that guarantee coverage and connectivity with a small number of sensors. We consider two different scenarios based on the flexibility of deployment. If deployment has to be accomplished in one step, like airborne deployment, then the main question becomes how many sensors are needed. If deployment can be implemented in multiple steps, then awareness of coverage and connectivity can be updated. For this case, we present incremental deployment algorithms, which consider the current placement to adjust the sampling domain. The algorithms are simple, easy to implement, and require a small number of sensors. We believe the concepts and algorithms presented in this paper provide a unifying framework for existing and future deployment algorithms, which consider many practical issues not considered in the present work.
Volkan Isler, Sampath Kannan, Kostas Daniilidis
IROS3
2004 VC-Dimension of Exterior Visibility
abstract
In this paper, we study the Vapnik-Chervonenkis (VC)-dimension of set systems arising in 2D polygonal and 3D polyhedral configurations where a subset consists of all points visible from one camera. In the past, it has been shown that the VC-dimension of planar visibility systems is bounded by 23 if the cameras are allowed to be anywhere inside a polygon without holes. Here, we consider the case of exterior visibility, where the cameras lie on a constrained area outside the polygon and have to observe the entire boundary. We present results for the cases of cameras lying on a circle containing a polygon (VC-dimension= 2) or lying outside the convex hull of a polygon (VC-dimension= 5). The main result of this paper concerns the 3D case: We prove that the VC-dimension is unbounded if the cameras lie on a sphere containing the polyhedron, hence the term exterior visibility.
Volkan Isler, Sampath Kannan, Kostas Daniilidis, Pavel Valtr 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2004 Stereo-based environment scanning for immersive telepresence
abstract
The processing power and network bandwidth required for true immersive telepresence applications are only now beginning to be available. We draw from our experience developing stereo based tele-immersion prototypes to present the main issues arising when building these systems. Tele-immersion is a new medium that enables a user to share a virtual space with remote participants. The user is immersed in a rendered three-dimensional (3-D) world that is transmitted from a remote site. To acquire this 3-D description, we apply binocular and trinocular stereo techniques which provide a view-independent scene description. Slow processing cycles or long network latencies interfere with the users' ability to communicate, so the dense stereo range data must be computed and transmitted at high frame rates. Moreover, reconstructed 3-D views of the remote scene must be as accurate as possible to achieve a sense of presence. We address both issues of speed and accuracy using a variety of techniques including the power of supercomputing clusters and a method for combining motion and stereo in order to increase speed and robustness. We present the latest prototype acquiring a room-size environment in real time using a supercomputing cluster, and we discuss its strengths and current weaknesses.
Jane Mulligan, Xenophon Zabulis, Nikhil Kelshikar, Kostas Daniilidis
IEEE Trans. Circuits Syst. Video Technol.4
2003 Direct 3D-Rotation Estimation from Spherical Images via a Generalized Shift Theorem
abstract
Omnidirectional images arising from 3D-motion of a camera contain persistent structures over a large variation of motions because of their large field of view. This persistence made appearance-based methods attractive for robot localization given reference views. Assuming that central omnidirectional images can be mapped to the sphere, the question is what are the underlying mappings of the sphere that can reflect a rotational camera motion. Given such a mapping, we propose a systematic way for finding invariance and the mapping parameters themselves based on the generalization of the Fourier transform. Using results from representation theory, we can generalize the Fourier transform to any homogeneous space with a transitively acting group. Such a case is the sphere with rotation as the acting group. The spherical harmonics of an image pair are related to each other through a shift theorem involving the irreducible representation of the rotation group. We show how to extract Euler angles using this theorem. We study the effect of the number of spherical harmonic coefficients as well as the effect of violation of appearance persistence in real imagery.
Ameesh Makadia, Kostas Daniilidis
CVPR (2)2
2003 Mirrors in motion: Epipolar geometry and motion estimation
abstract
In this paper we consider the images taken from pairs of parabolic catadioptric cameras separated by discrete motions. Despite the nonlinearity of the projection model, the epipolar geometry arising from such a system, like the perspective case, can be encoded in a bilinear form, the catadioptric fundamental matrix. We show that all such matrices have equal Lorentzian singular values, and they define a nine-dimensional manifold in the space of 4 /spl times/ 4 matrices. Furthermore, this manifold can be identified with a quotient of two Lie groups. We present a method to estimate a matrix in this space, so as to obtain an estimate of the motion. We show that the estimation procedures are robust to modest deviations from the ideal assumptions.
Christopher Geyer, Kostas Daniilidis
ICCV2
2003 Local exploration: online algorithms and a probabilistic framework
abstract
Mapping an environment with an imaging sensor becomes very challenging if the environment to be mapped is unknown and has to be explored. Exploration involves the planning of views so that the entire environment is covered. The majority of implemented mapping systems use a heuristic planning while theoretical approaches regard only the traveled distance as cost. However, practical range acquisition systems spend a considerable amount of time for acquisition. In this paper, we address the problem of minimizing the cost of looking around a corner, involving the time spent in traveling as well as the time spent for reconstruction. Such a local exploration can be used as a subroutine for global algorithms. We prove competitive ratios for two online algorithms. Then, we provide two representations of local exploration as a Markov Decision Process and apply a known policy iteration algorithm. Simulation results show that for some distributions the probabilistic approach outperforms deterministic strategies.
Volkan Isler, Sampath Kannan, Kostas Daniilidis
ICRA3
2003 Multiple motion analysis: in spatial or in spectral domain?
Weichuan Yu, Gerald Sommer, Kostas Daniilidis
Comput. Vis. Image Underst.3
2003 Three dimensional orientation signatures with conic kernel filtering for multiple motion analysis
Weichuan Yu, Gerald Sommer, Kostas Daniilidis
Image Vis. Comput.3
2003 Linear Pose Estimation from Points or Lines
abstract
Estimation of camera pose from an image of n points or lines with known correspondence is a thoroughly studied problem in computer vision. Most solutions are iterative and depend on nonlinear optimization of some geometric constraint, either on the world coordinates or on the projections to the image plane. For real-time applications, we are interested in linear or closed-form solutions free of initialization. We present a general framework which allows for a novel set of linear solutions to the pose estimation problem for both n points and n lines. We then analyze the sensitivity of our solutions to image noise and show that the sensitivity analysis can be used as a conservative predictor of error for our algorithms. We present a number of simulations which compare our results to two other recent linear algorithms, as well as to iterative approaches. We conclude with tests on real imagery in an augmented reality setup.
Adnan Ansar, Kostas Daniilidis
IEEE Trans. Pattern Anal. Mach. Intell.2
2003 Omnidirectional video
Christopher Geyer, Kostas Daniilidis
Vis. Comput.2
2002 Linear Pose Estimation from Points or Lines
Adnan Ansar, Kostas Daniilidis
ECCV (4)2
2002 Properties of the Catadioptric Fundamental Matrix
Christopher Geyer, Kostas Daniilidis
ECCV (2)2
2002 Trinocular Stereo: A Real-Time Algorithm and its Evaluation
Jane Mulligan, Volkan Isler, Kostas Daniilidis
Int. J. Comput. Vis.3
2002 Paracatadioptric Camera Calibration
abstract
Catadioptric sensors refer to the combination of lens-based devices and reflective surfaces. These systems are useful because they may have a field of view which is greater than hemispherical, providing the ability to simultaneously view in any direction. Configurations which have a unique effective viewpoint are of primary interest, among these is the case where the reflective surface is a parabolic mirror and the camera is such that it induces an orthographic projection and which we call paracatadioptric. We present an algorithm for the calibration of such a device using only the images of lines in space. In fact, we show that we may obtain all of the intrinsic parameters from the images of only three lines and that this is possible without any metric information. We propose a closed-form solution for focal length, image center, and aspect ratio for skewless cameras and a polynomial root solution in the presence of skew. We also give a method for determining the orientation of a plane containing two sets of parallel lines from one uncalibrated view. Such an orientation recovery enables a rectification which is impossible to achieve in the case of a single uncalibrated view taken by a conventional camera. We study the performance of the algorithm in simulated setups and compare results on real images with an approach based on the image of the mirror's bounding circle.
Christopher Geyer, Kostas Daniilidis
IEEE Trans. Pattern Anal. Mach. Intell.2
2002 Oriented Structure of the Occlusion Distortion: Is It Reliable?
abstract
In the energy spectrum of an occlusion sequence, the distortion term has the same orientation as the velocity of the occluding signal. Other works claimed that this oriented structure can be used to distinguish the occluding velocity from the occluded one. We argue that the orientation structure of the distortion cannot always work as a reliable feature due to the rapidly decreasing energy contribution. This already weak orientation structure is further blurred by a superposition of distinct distortion components. We also indicate that the superposition principle of Shizawa and Mase (1991) for multiple motion estimation needs to be adjusted.
Weichuan Yu, Gerald Sommer, Steven S. Beauchemin, Kostas Daniilidis
IEEE Trans. Pattern Anal. Mach. Intell.4
2001 Linear Augmented Reality Registration
Adnan Ansar, Kostas Daniilidis
CAIP2
2001 Multispectral Skin Color Modeling
abstract
The automated detection of humans in computer vision as well as the realistic rendering of people in computer graphics necessitates improved modeling of the human skin color. We describe the acquisition and modeling of skin reflectance data densely sampled over the entire visible spectrum. The data collected through a spectrograph allows us to explain skin color (and its variations) and to discriminate between human skin and dyes designed to mimic human skin. We study the approximation of these data using several sets of basis functions. Our study shows that skin reflectance data can best be approximated by a linear combination of Gaussians or their first derivatives. This result has a significant practical impact on optical acquisition devices: the entire visible spectrum of skin reflectance can now be captured with a few filters of optimally chosen central wavelengths and bandwidth.
Elli Angelopoulou, Rana Molana, Kostas Daniilidis
CVPR (2)3
2001 Structure and Motion from Uncalibrated Catadioptric Views
abstract
In this paper we present a new algorithm for structure from motion from point correspondences in images taken from uncalibrated catadioptric cameras with parabolic mirrors. We assume that the unknown intrinsic parameters are three: the combined focal length of the mirror and lens and the intersection of the optical axis with the image. We introduce a new representation for images of points and lines in catadioptric images which we call the circle space. This circle space includes imaginary circles, one of which is the image of the absolute conic. We formulate the epipolar constraint in this space and establish a new 4/spl times/4 catadioptric fundamental matrix. We show that the image of the absolute conic belongs to the kernel of this matrix. This enables us to prove that Euclidean reconstruction is feasible from two views with constant parameters and from three views with varying parameters. In both cases, it is one less than the number of views necessary with perspective cameras.
Christopher Geyer, Kostas Daniilidis
CVPR (1)2
2001 3D-Orientation Signatures with Conic Kernel Filtering for Multiple Motion Analysis
abstract
In this paper we propose a new 3D kernel for the recovery of 3D-orientation signatures. The kernel is a Gaussian function defined in local spherical coordinates and its Cartesian support has the shape of a truncated cone with its axis in the radial direction and very small angular support. A set of such kernels is obtained by uniformly sampling the 2D space of polar and azimuth angles. The projection of a local neighborhood on such a kernel set produces a local 3D-orientation signature. In the case of spatiotemporal analysis, such a kernel set can be applied either on the derivative space of a local neighborhood or on the local Fourier transform. The well known planes arising from single or multiple motion produce maxima in the orientation signature. Due to the kernel's local support spatiotemporal signatures possess higher orientation resolution than 3D steerable filters and motion maxima can be detected and localized more accurately. We describe and show in experiments the superiority of the proposed kernels compared to Hough transformation or EM-based multiple motion detection.
Weichuan Yu, Gerald Sommer, Kostas Daniilidis
CVPR (1)3
2001 Performance Evaluation of Stereo for Tele-presence
abstract
In an immersive tele-presence environment a 3D remote real scene is projected from the viewpoint of the local user. This 3D world is acquired through stereo reconstruction at the remote site. In this paper we start a performance analysis of stereo algorithms with respect to the task of immersive visualization. As opposed to usual monocular image based rendering, we are also interested in the depth error in novel views because our rendering is stereoscopic. We describe an evaluation test-bed which provides a world-wide first available set of registered dense "ground-truth" laser data and image data from multiple views. We establish metrics for novel depth views that reflect discrepancies both in the image and in 3D-space. It is well known that stereo performance is affected by both erroneous matching as well as incorrect depth triangulation. We experimentally study the effects of occlusion and low texture on the distributions of the error metrics. Then, we algebraically predict the behavior of depth and novel projection error as a function of the camera set-up and the error in the disparity. These are first steps towards building a laboratory for psychophysical judgement of depth estimates which is the ultimate performance test of tele-presence stereo.
Jane Mulligan, Volkan Isler, Kostas Daniilidis
ICCV3
2001 Real time trinocular stereo for tele-immersion
abstract
Tele-immersion is a technology that augments your space with real-time 3D projections of remote spaces thus facilitating the interaction of people from different places in virtually the same environment. Tele-immersion combines 3D scene recovery from computer vision, and rendering and interaction from computer graphics. We describe the real-time 3D scene acquisition using a new algorithm for trinocular stereo. We extend this method in time by combining motion and stereo in order to increase speed and robustness.
Jane Mulligan, Kostas Daniilidis
ICIP (3)2
2001 Visual and haptic collaborative tele-presence
Adnan Ansar, Denilson Rodrigues, Jaydev P. Desai, Kostas Daniilidis, Vijay Kumar 0001, Mario Fernando Montenegro Campos
Comput. Graph.4
2001 Catadioptric Projective Geometry
Christopher Geyer, Kostas Daniilidis
Int. J. Comput. Vis.2
2001 Approximate orientation steerability based on angular Gaussians
abstract
Junctions are significant features in images with intensity variation that exhibits multiple orientations. This makes the detection and characterization of junctions a challenging problem. The characterization of junctions would ideally be given by the response of a filter at every orientation. This can be achieved by the principle of steerability that enables the decomposition of a filter into a linear combination of basis functions. However, current steerability approaches suffer from the consequences of the uncertainty principle: in order to achieve high resolution in orientation they need a large number of basis filters increasing, thus, the computational complexity. Furthermore, these functions have usually a wide support which only accentuates the computational burden. We propose a novel alternative to current steerability approaches. It is based on utilizing a set of polar separable filters with small support to sample orientation information. The orientation signature is then obtained by interpolating orientation samples using Gaussian functions with small support. Compared with current steerability techniques our approach achieves a higher orientation resolution with a lower complexity. In addition, we build a polar pyramid to characterize junctions of arbitrary inherent orientation scales.
Weichuan Yu, Kostas Daniilidis, Gerald Sommer
IEEE Trans. Image Process.2
2000 A Unifying Theory for Central Panoramic Systems and Practical Applications
Christopher Geyer, Kostas Daniilidis
ECCV (2)2
2000 Predicting Disparity Windows for Real-Time Stereo
Jane Mulligan, Kostas Daniilidis
ECCV (1)2
2000 Omnidirectional Vision: Theory and Algorithms
abstract
Surround perception is crucial for an immersive sense of presence in communication and for efficient navigation and surveillance in robotics. To enable surround perception, new omnidirectional systems were designed which gave a new impetus for rethinking the way images are acquired and analyzed. Based on insights gained from such designs, we formulate a novel unifying theory of imaging. We prove that all single viewpoint mirror-lens devices are equivalent to projective mappings from the sphere to the plane. These mappings are paired with a duality principle which relates points to line projections. The commonly used parabolic mirror projection is shown to be equivalent to the stereographic projection, providing therefore the invariants of a conformal mapping. It turns out that conventional cameras, which are only a special case in our theory, provide the barest minimum of information about the environment. We review current approaches to omnidirectional imaging and present a framework for calibration of omnidirectional cameras from single views.
Kostas Daniilidis, Christopher Geyer
ICPR1
2000 Trinocular Stereo for Non-Parallel Configurations
abstract
The constraint of a third camera in stereo vision is a useful tool for reducing ambiguity in matching. Most of the systems using trinocular stereo to date however, have used configurations where the image planes of all three cameras are coplanar, or can be rectified to be so. We explore the computation of dense trinocular disparity maps for non-planar camera configurations which arise when cameras surround the object to be modeled. Our approach rectifies the cameras as two independent stereo pairs. We start with an exhaustive lookup scheme and then consider retaining only a list of N disparities per pixel with maximal correlation values for the right and left pairs. Experimental results and comparisons demonstrate that both methods reduce outliers over binocular stereo, and that the N-hypothesis system trades large lookup tables for somewhat lower density of valid matches.
Jane Mulligan, Kostas Daniilidis
ICPR2
1999 Complex Analysis for Reconstruction from Controlled Motion
R. Andrew Hicks, David Pettey, Kostas Daniilidis, Ruzena Bajcsy
CAIP3
1999 Constrained Self-Calibration
abstract
This paper focuses on the estimation of the intrinsic camera parameters and the trajectory of the camera from an image sequence. Intrinsic camera calibration and pose estimation are the prerequisites for many applications involving navigation tasks, scene reconstruction, and merging of virtual and real environments. Proposed and evaluated is a technical solution to decrease the sensitivity of self-calibration by placing easily identifiable targets of known shape in the environment. The relative position of the targets need not be known a priori. Assuming an appropriate ratio of size to distance these targets resolve known ambiguities. Constraints on the target placement and the cameras' motions are explored. The algorithm is extensively tested in a variety of real-world scenarios.
Jeffrey Mendelsohn, Kostas Daniilidis
CVPR2
1999 Detection and Characterization of Multiple Motion Points
abstract
The computation of optical flow is a well studied topic in biological and computational vision. However, the existence of multiple motions in dynamic imagery due to occlusion or even transparency still raises challenging questions. In this paper, we propose an approach for the detection and characterization of occlusion and transparency. We propose a theoretical framework for both types of multiple motions which explicitly shows the difference between occlusion and transparency in the frequency domain. Then, we employ an EM-algorithm for the computation of one or two image velocities and a simple test for the detection of occlusion. Our approach differs from other EM-approaches which blindly assume the superposition of two models in the spatial domain without providing with a separate formal model for occlusion. We test and compare the characterization performance on synthetic and real data.
Weichuan Yu, Gerald Sommer, Steven S. Beauchemin, Kostas Daniilidis
CVPR4
1999 Catadioptric Camera Calibration
abstract
Catadioptric systems are realizations of omnidirectional vision through mirror-lens combinations. Designs preserving the uniqueness of an effective viewpoint have recently gained attraction. We present here a novel approach for estimating the intrinsic parameters of a well-known catadioptric system consisting of a paraboloid mirror and an orthographic lens. We introduce the geometry of catadioptric line projection and we show that the vanishing points lie on a conic section which encodes the entire calibration information. Projections of two sets of parallel lines suffice for intrinsic calibration from one view as well as for metric rectification of a plane. Our approach overcomes limitations of existing manual calibration methods and was successfully tested on the task of back-warping real-images images onto virtual planes.
Christopher Geyer, Kostas Daniilidis
ICCV2
1998 Rotated Wedge Averaging Method for Junction Classification
abstract
The computational cost of conventional filter methods for junction characterization is very high. This burden can be attenuated by using steerable filters. However, in order to achieve a high orientational selectivity to characterize complex junctions a large number of basis filters is necessary. From this results a yet too high computational effort for steerable filters. In this paper we present a new method for characterizing junctions which keeps the high orientational resolution and is computationally efficient. It is based on applying rotated copies of a wedge averaging filter and estimating the derivative with respect to the polar angle. The new method is compared with the steerable wedge filter method in experiments with real images. We show the superiority of our method as well as its adaptability to scale changes and robustness against noise.
Weichuan Yu, Kostas Daniilidis, Gerald Sommer
CVPR2
1998 Low-Cost Junction Characterization using Polar Averaging Filters
Weichuan Yu, Kostas Daniilidis, Gerald Sommer
ICIP (3)2
1997 Optimization of Stereo Disparity Estimation Using the Instantaneous Frequency
Michael Hansen, Kostas Daniilidis, Gerald Sommer
CAIP2
1997 Fixation Simplifies 3D Motion Estimation
Kostas Daniilidis
Comput. Vis. Image Underst.1
1996 Active Intrinsic Calibration Using Vanishing Points
abstract
During a fixed axis camera rotation every image point is moving on a conic section. If the point is a vanishing point the conic section is invariant to possible translations of the observer. Given the rotation axis and the inter-frame correspondence of a set of parallel lines we are able to compute the intrinsic parameters without knowledge of the rotation angles. We propagate the error covariances and we remove the bias in the computation of the conic. We experimentally study the sensitivity of calibration to the amount of rotation and we compare our performance to the performance of a recent active calibration technique.
Kostas Daniilidis, Jörg Ernst
CVPR1
1996 Decoupling the 3D Motion Space by Fixation
Kostas Daniilidis, Inigo Thomas
ECCV (1)1
1996 The dual quaternion approach to hand-eye calibration
abstract
In order to relate measurements made by a sensor mounted on a mechanical link to the robot's coordinate frame we must first estimate the transformation between the sensor and the link frame. In this paper we introduce the use of dual quaternions which are the algebraic counterpart of screws. We prove algebraically that if we consider the camera and motor transformations as screws, then only the line coefficients of the screw axes are relevant regarding the hand-eye calibration. This new parametrization enables us to simultaneously solve for the hand-eye rotation and translation using the singular value decomposition.
Kostas Daniilidis, Eduardo Bayro-Corrochano
ICPR1
1996 Active intrinsic calibration using vanishing points
Kostas Daniilidis, Jörg Ernst
Pattern Recognit. Lett.1
1995 Computation of 3-D-Motion Parameters Using the Log-Polar Transform
Kostas Daniilidis
CAIP1
1995 Optical Flow Computation in the Log-Polar-Plane
Kostas Daniilidis, Volker Krüger
CAIP1
1993 The coupling of rotation and translation in motion estimation of planar surfaces
abstract
The error sensitivity in the estimation of the 3-D motion and the normal of a planar surface from an instantaneous motion field is studied. The statistical theory or the Cramer-Rao lower bound is used for error covariance in the estimated motion and structure parameters. This enables the derivation of results valid for any unbiased estimator under the assumption of Gaussian noise in the motion field. The obtained lower-bound-matrix is studied analytically with respect to measurement noise, size of the field of view, and the motion-geometry configuration. The coupling between translation and rotation is exacerbated if the field of view and the slant of the plane become smaller, and the deviation of the translation from the viewing direction becomes larger. The relationships of the uncertainty bounds for every unknown motion parameter to the angle between translation and the plane-normal, the size of the field of view, and the distance from the perceived plane and the translation magnitude are discussed.>
Kostas Daniilidis, Hans-Hellmut Nagel
CVPR1
1993 Model-based object tracking in monocular image sequences of road traffic scenes
D. Roller, Kostas Daniilidis, Hans-Hellmut Nagel
Int. J. Comput. Vis.2
1992 Model-Based Object Tracking in Traffic Scenes
Dieter Koller, Kostas Daniilidis, Torfi Thórhallsson, Hans-Hellmut Nagel
ECCV2
1990 Analytical Results on Error Sensitivity of Motion Estimation from Two Views
Kostas Daniilidis, Hans-Hellmut Nagel
ECCV1
1990 Analytical results on error sensitivity of motion estimation from two views
Kostas Daniilidis, Hans-Hellmut Nagel
Image Vis. Comput.1