EDBT 2026 Demo / reviewers in the wild / expert
Philippos Mordohai
dblp:10/618
· DBLP profile ↗
56ranked-venue papers
6as first author
15since 2021 · last 2025
0000-0002-9671-4408ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 44 · 6 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 3 first-author · 3 since 2021Systems, architecture and hardware · 10 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Consensus-Driven Uncertainty for Robotic Grasping based on RGB PerceptionabstractDeep object pose estimators are notoriously overconfident. A grasping agent that both estimates the 6-DoF pose of a target object and predicts the uncertainty of its own estimate could avoid task failure by choosing not to act under high uncertainty. Even though object pose estimation improves and uncertainty quantification research continues to make strides, few studies have connected them to the downstream task of robotic grasping. We propose a method for training lightweight, deep networks to predict whether a grasp guided by an image-based pose estimate will succeed before that grasp is attempted. We generate training data for our networks via object pose estimation on real images and simulated grasping. We also find that, despite high object variability in grasping trials, networks benefit from training on all objects jointly, suggesting that a diverse variety of objects can nevertheless contribute to the same goal. Data, code, and guides are hosted at: https://github.com/EricCJoyce/Consensus-Driven-Uncertainty/ Eric C. Joyce, Qianwen Zhao, Nathaniel Burgdorfer, Philippos Mordohai |
IROS | 5 |
| 2025 | Understanding while Exploring: Semantics-driven Active MappingabstractEffective robotic autonomy in unknown environments demands proactive exploration and precise understanding of both geometry and semantics. In this paper, we propose ActiveSGM, an active semantic mapping framework designed to predict the informativeness of potential observations before execution. Built upon a 3D Gaussian Splatting (3DGS) mapping backbone, our approach employs semantic and geometric uncertainty quantification, coupled with a sparse semantic representation, to guide exploration. By enabling robots to strategically select the most beneficial viewpoints, ActiveSGM efficiently enhances mapping completeness, accuracy, and robustness to noisy semantic data, ultimately supporting more adaptive scene exploration. Our experiments on the Replica and Matterport3D datasets highlight the effectiveness of ActiveSGM in active semantic mapping tasks. Huangying Zhan, Hairong Yin, Yi Xu 0002, Philippos Mordohai |
NeurIPS | 5 |
| 2025 | Glissando-Net: Deep Single View Category Level Pose Estimation and 3D ReconstructionabstractWe present a deep learning model, dubbed Glissando-Net, to simultaneously estimate the pose and reconstruct the 3D shape of objects at the category level from a single RGB image. Previous works predominantly focused on either estimating poses (often at the instance level), or reconstructing shapes, but not both. Glissando-Net is composed of two auto-encoders that are jointly trained, one for RGB images and the other for point clouds. We embrace two key design choices in Glissando-Net to achieve a more accurate prediction of the 3D shape and pose of the object given a single RGB image as input. First, we augment the feature maps of the point cloud encoder and decoder with transformed feature maps from the image decoder, enabling effective 2D-3D interaction in both training and prediction. Second, we predict both the 3D shape and pose of the object in the decoder stage. This way, we better utilize the information in the 3D point clouds presented only in the training stage to train the network for more accurate prediction. We jointly train the two encoder-decoders for RGB and point cloud data to learn how to pass latent features to the point cloud decoder during inference. In testing, the encoder of the 3D point cloud is discarded. The design of Glissando-Net is inspired by codeSLAM. Unlike codeSLAM, which targets 3D reconstruction of scenes, we focus on pose estimation and shape reconstruction of objects, and directly predict the object pose and a pose invariant 3D reconstruction without the need of the code optimization step. Extensive experiments, involving both ablation studies and comparison with competing methods, demonstrate the efficacy of our proposed method, and compare favorably with the state-of-the-art. Hao Kang, Philippos Mordohai, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Stereo-NEC: Enhancing Stereo Visual-Inertial SLAM Initialization with Normal Epipolar ConstraintsabstractWe propose an accurate and robust initialization approach for stereo visual-inertial SLAM systems. Unlike the current state-of-the-art method, which heavily relies on the accuracy of a pure visual SLAM system to estimate inertial variables without updating camera poses, potentially compromising accuracy and robustness, our approach offers a different solution. We realize the crucial impact of precise gyroscope bias estimation on rotation accuracy. This, in turn, affects trajectory accuracy due to the accumulation of translation errors. To address this, we first independently estimate the gyroscope bias and use it to formulate a maximum a posteriori problem for further refinement. After this refinement, we proceed to update the rotation estimation by performing IMU integration with gyroscope bias removed from gyroscope measurements. We then leverage robust and accurate rotation estimates to enhance translation estimation via 3-DoF bundle adjustment. Moreover, we introduce a novel approach for determining the success of the initialization by evaluating the residual of the normal epipolar constraint. Extensive evaluations on the EuRoC dataset illustrate that our method excels in accuracy and robustness. It outperforms ORB-SLAM3, the current leading stereo visual-inertial initialization method, in terms of absolute trajectory error and relative rotation error, while maintaining competitive computational speed. Notably, even with 5 keyframes for initialization, our method consistently surpasses the state-of-the-art approach using 10 keyframes in rotation accuracy. The open source code is available at https://github.com/ApdowJN/Stereo-NEC.git. Chieh Chou, Ganesh Sevagamoorthy, Zheng Chen 0016, Ziyue Feng, Youjie Xia, Feiyang Cai, Yi Xu 0002, Philippos Mordohai |
ICRA | 10 |
| 2023 | Learning the Distribution of Errors in Stereo Matching for Joint Disparity and Uncertainty EstimationabstractWe present a new loss function for joint disparity and uncertainty estimation in deep stereo matching. Our work is motivated by the need for precise uncertainty estimates and the observation that multi-task learning often leads to improved performance in all tasks. We show that this can be achieved by requiring the distribution of uncertainty to match the distribution of disparity errors via a KL divergence term in the network's loss function. A differentiable soft-histogramming technique is used to approximate the distributions so that they can be used in the loss. We Experimentally assess the effectiveness of our approach and observe significant improvements in both disparity and uncertainty prediction on large datasets. Our code is available at https://github.com/lly00412/SEDNet.git. Philippos Mordohai |
CVPR | 3 |
| 2023 | V-FUSE: Volumetric Depth Map Fusion with Long-Range ConstraintsabstractWe introduce a learning-based depth map fusion framework that accepts a set of depth and confidence maps generated by a Multi-View Stereo (MVS) algorithm as input and improves them. This is accomplished by integrating volumetric visibility constraints that encode long-range surface relationships across different views into an end-to-end trainable architecture. We also introduce a depth search window estimation sub-network trained jointly with the larger fusion sub-network to reduce the depth hypothesis search space along each ray. Our method learns to model depth consensus and violations of visibility constraints directly from the data; effectively removing the necessity of fine-tuning fusion parameters. Extensive experiments on MVS datasets show substantial improvements in the accuracy of the output fused depth and confidence maps. Our code is available at https://github.com/nburgdorfer/V-FUSE Nathaniel Burgdorfer, Philippos Mordohai |
ICCV | 2 |
| 2023 | 3-D Reconstruction Using Monocular Camera and Lights: Multi-View Photometric Stereo for Non-Stationary RobotsabstractThis paper proposes a novel underwater Multi-View Photometric Stereo (MVPS) framework for reconstructing scenes in 3-D with a non-stationary low-cost robot equipped with a monocular camera and fixed lights. The underwater realm is the primary focus of study here, due to the challenges in utilizing underwater camera imagery and lack of low-cost reliable localization systems. Previous underwater PS approaches provided accurate scene reconstruction results, but assumed that the robot was stationary at the bottom. This assumption is limiting, as many artifacts, reefs, and man-made structures are large and meters above the bottom. Our proposed MVPS framework relaxes the stationarity assumption by utilizing a monocular SLAM system to estimate small robot motions and extract an initial sparse feature map. To compensate for the scale inconsistency in monocular SLAM output, our MVPS optimization scheme collectively estimates a high-quality, dense 3-D reconstruction and corrects the camera pose estimates. We also present an attenuation and camera-light extrinsic parameter calibration method for non-stationary robots. Finally, validation experiments with a BlueROV2 demonstrated the low-cost capability of producing high-quality scene reconstructions. Overall, this work is the foundation of an active perception pipeline for robots (i.e., underwater, ground, and aerial) to explore and map complex structures in high accuracy and resolution with an inexpensive sensor-light configuration. Monika Roznere, Philippos Mordohai, Ioannis M. Rekleitis, Alberto Quattrini Li |
ICRA | 2 |
| 2023 | Real-Time Dense 3D Mapping of Underwater EnvironmentsabstractThis paper addresses real-time dense 3D reconstruction for a resource-constrained Autonomous Underwater Vehicle (AUV). Underwater vision-guided operations are among the most challenging as they combine 3D motion in the presence of external forces, limited visibility, and absence of global positioning. Obstacle avoidance and effective path planning require online dense reconstructions of the environment. Autonomous operation is central to environmental monitoring, marine archaeology, resource utilization, and underwater cave exploration. To address this problem, we propose to use SVIn2, a robust VIO method, together with a real-time 3D reconstruction pipeline. We provide extensive evaluation on four challenging underwater datasets. Our pipeline produces comparable reconstruction with that of COLMAP, the state-of-the-art offline 3D reconstruction method, at high frame rates on a single CPU. Bharat Joshi, Nathaniel Burgdorfer, Konstantinos Batsos, Alberto Quattrini Li, Philippos Mordohai, Ioannis M. Rekleitis |
ICRA | 6 |
| 2023 | EDI: ESKF-based Disjoint Initialization for Visual-Inertial SLAM SystemsabstractVisual-inertial initialization can be classified into joint and disjoint approaches. Joint approaches tackle both the visual and the inertial parameters together by aligning observations from feature-bearing points based on IMU integration then use a closed-form solution with visual and acceleration observations to find initial velocity and gravity. In contrast, disjoint approaches independently solve the Structure from Motion (SFM) problem and determine inertial parameters from up-to-scale camera poses obtained from pure monocular SLAM. However, previous disjoint methods have limitations, like assuming negligible acceleration bias impact or accurate rotation estimation by pure monocular SLAM. To address these issues, we propose EDI, a novel approach for fast, accurate, and robust visual-inertial initialization. Our method incorporates an Error-state Kalman Filter (ESKF) to estimate gyroscope bias and correct rotation estimates from monocular SLAM, overcoming dependence on pure monocular SLAM for rotation estimation. To estimate the scale factor without prior information, we offer a closed-form solution for initial velocity, scale, gravity, and acceleration bias estimation. To address gravity and acceleration bias coupling, we introduce weights in the linear least-squares equations, ensuring acceleration bias observability and handling outliers. Extensive evaluation on the EuRoC dataset shows that our method achieves an average scale error of 5.8% in less than 3 seconds, outperforming other state-of-the-art disjoint visual-inertial initialization approaches, even in challenging environments and with artificial noise corruption. Yuhang Ming 0001, Philippos Mordohai |
IROS | 4 |
| 2023 | PRN: Panoptic Refinement NetworkabstractPanoptic segmentation is the task of uniquely assigning every pixel in an image to either a semantic label or an individual object instance, generating a coherent and complete scene description. Many current panoptic segmentation methods, however, predict masks of semantic classes and object instances in separate branches, yielding inconsistent predictions. Moreover, because state-of-the-art panoptic segmentation models rely on box proposals, the instance masks predicted are often of low-resolution. To overcome these limitations, we propose the Panoptic Refinement Network (PRN), which takes masks from base panoptic segmentation models and refines them jointly to produce coherent results. PRN extends the offset map-based architecture of Panoptic-Deeplab with several novel ideas including a foreground mask and instance bounding box offsets, as well as coordinate convolutions for improved spatial prediction. Experimental results on COCO and Cityscapes show that PRN can significantly improve already accurate results from a variety of panoptic segmentation networks. Jason Kuen, Zhe Lin 0001, Philippos Mordohai, Simon Chen |
WACV | 4 |
| 2022 | Towards Mapping of Underwater Structures by a Team of Autonomous Underwater Vehicles
Marios Xanthidis, Bharat Joshi, Monika Roznere, Nathaniel Burgdorfer, Alberto Quattrini Li, Philippos Mordohai, Srihari Nelakuditi, Ioannis M. Rekleitis |
ISRR | 7 |
| 2022 | Single-camera 3D head fitting for mixed reality clinical applications
Tejas Mane, Aylar Bayramova, Kostas Daniilidis, Philippos Mordohai, Elena Bernardis |
Comput. Vis. Image Underst. | 4 |
| 2022 | On the Synergies Between Machine Learning and Binocular Stereo for Depth Estimation From Images: A SurveyabstractStereo matching is one of the longest-standing problems in computer vision with close to 40 years of studies and research. Throughout the years the paradigm has shifted from local, pixel-level decision to various forms of discrete and continuous optimization to data-driven, learning-based methods. Recently, the rise of machine learning and the rapid proliferation of deep learning enhanced stereo matching with new exciting trends and applications unthinkable until a few years ago. Interestingly, the relationship between these two worlds is two-way. While machine, and especially deep, learning advanced the state-of-the-art in stereo matching, stereo itself enabled new ground-breaking methodologies such as self-supervised monocular depth estimation based on deep networks. In this paper, we review recent research in the field of learning-based depth estimation from single and binocular images highlighting the synergies, the successes achieved so far and the open challenges the community is going to face in the immediate future. Matteo Poggi, Fabio Tosi, Konstantinos Batsos, Philippos Mordohai, Stefano Mattoccia |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Inlier clustering based on the residuals of random hypotheses
Mohammed Kutbi, Yizhe Chang, Philippos Mordohai |
Pattern Recognit. Lett. | 3 |
| 2021 | Usability Studies of an Egocentric Vision-Based Robotic WheelchairabstractMotivated by the need to improve the quality of life for the elderly and disabled individuals who rely on wheelchairs for mobility, and who may have limited or no hand functionality at all, we propose an egocentric computer vision based co-robot wheelchair to enhance their mobility without hand usage. The robot is built using a commercially available powered wheelchair modified to be controlled by head motion. Head motion is measured by tracking an egocentric camera mounted on the user’s head and faces outward. Compared with previous approaches to hands-free mobility, our system provides a more natural human robot interface because it enables the user to control the speed and direction of motion in a continuous fashion, as opposed to providing a small number of discrete commands. This article presents three usability studies, which were conducted on 37 subjects. The first two usability studies focus on comparing the proposed control method with existing solutions while the third study was conducted to assess the effectiveness of training subjects to operate the wheelchair over several sessions. A limitation of our studies is that they have been conducted with healthy participants. Our findings, however, pave the way for further studies with subjects with disabilities. Mohammed Kutbi, Xiaoxue Du, Yizhe Chang, Nikolaos Agadakos, Gang Hua 0001, Philippos Mordohai |
ACM Trans. Hum. Robot Interact. | 8 |
| 2020 | Do End-to-end Stereo Algorithms Under-utilize Information?abstractDeep networks for stereo matching typically leverage 2D or 3D convolutional encoder-decoder architectures to aggregate cost and regularize the cost volume for accurate disparity estimation. Due to content-insensitive convolutions and down-sampling and up-sampling operations, these cost aggregation mechanisms do not take full advantage of the information available in the images. Disparity maps suffer from over-smoothing near occlusion boundaries, and erroneous predictions in thin structures. In this paper, we show how deep adaptive filtering and differentiable semi-global aggregation can be integrated in existing 2D and 3D convolutional networks for end-to-end stereo matching, leading to improved accuracy. The improvements are due to utilizing RGB information from the images as a signal to dynamically guide the matching process, in addition to being the signal we attempt to match across the images. We show extensive experimental results on the KITTI 2015 and Virtual KITTI 2 datasets comparing four stereo networks (DispNetC, GCNet, PSMNet and GANet) after integrating four adaptive filters (segmentation-aware bilateral filtering, dynamic filtering networks, pixel adaptive convolution and semi-global aggregation) into their architectures. Our code is available at https://github.com/ccj5351/DAFStereoNets. Changjiang Cai, Philippos Mordohai |
3DV | 2 |
| 2020 | Matching-space Stereo Networks for Cross-domain GeneralizationabstractEnd-to-end deep networks represent the state of the art for stereo matching. While excelling on images framing environments similar to the training set, major drops in accuracy occur in unseen domains (e.g., when moving from synthetic to real scenes). In this paper we introduce a novel family of architectures, namely Matching-Space Networks (MS-Nets), with improved generalization properties. By replacing learning-based feature extraction from image RGB values with matching functions and confidence measures from conventional wisdom, we move the learning process from the color space to the Matching Space, avoiding over-specialization to domain specific features. Extensive experimental results on four real datasets highlight that our proposal leads to superior generalization to unseen environments over conventional deep architectures, keeping accuracy on the source domain almost unaltered. Our code is available at https://qithub.com/ccj5351/MS-Nets. Changjiang Cai, Matteo Poggi, Stefano Mattoccia, Philippos Mordohai |
3DV | 4 |
| 2020 | NBVC: A Benchmark for Depth Estimation from Narrow-Baseline Video ClipsabstractWe present a benchmark for online, video-based depth estimation, a problem that is not covered by the current set of benchmarks for evaluating 3D reconstruction, which focus on offline, batch reconstruction. Online depth estimation from video captured by a moving camera is a key enabling technology for compelling applications in robotics and augmented reality. Inspired by progress in many aspects of robotics due to benchmarks and datasets, we propose a new benchmark called NBVC for evaluating methods for online depth estimation from video. Our benchmark is composed of short video sequences with corresponding high-quality ground truth depth maps, derived from the recent Tanks and Temples dataset. We are hopeful that our work will be instrumental in the development of learning-based algorithms for online depth estimation from video clips, and will also lead to improvements in conventional approaches. In addition to the benchmark, we present a superpixel-based plane sweeping stereo algorithm and use it to investigate various aspects of the problem. The paper contains our initial findings and conclusions. Philippos Mordohai, Konstantinos Batsos, Ameesh Makadia, Noah Snavely |
IROS | 1 |
| 2019 | Oriented Point Sampling for Plane Detection in Unorganized Point CloudsabstractPlane detection in 3D point clouds is a crucial pre-processing step for applications such as point cloud segmentation, semantic mapping and SLAM. In contrast to many recent plane detection methods that are only applicable on organized point clouds, our work is targeted to unorganized point clouds that do not permit a 2D parametrization. We compare three methods for detecting planes in point clouds efficiently. One is a novel method proposed in this paper that generates plane hypotheses by sampling from a set of points with estimated normals. We named this method Oriented Point Sampling (OPS) to contrast with more conventional techniques that require the sampling of three unoriented points to generate plane hypotheses. We also implemented an efficient plane detection method based on local sampling of three unoriented points and compared it with OPS and the 3D-KHT algorithm, which is based on octrees, on the detection of planes on 10,000 point clouds from the SUN RGB-D dataset. Philippos Mordohai |
ICRA | 2 |
| 2018 | RecResNet: A Recurrent Residual CNN Architecture for Disparity Map EnhancementabstractWe present a neural network architecture applied to the problem of refining a dense disparity map generated by a stereo algorithm to which we have no access. Our approach is able to learn which disparity values should be modified and how, from a training set of images, estimated disparity maps and the corresponding ground truth. Its only input at test time is a disparity map and the reference image. Two design characteristics are critical for the success of our network: (i) it is formulated as a recurrent neural network, and (ii) it estimates the output refined disparity map as a combination of residuals computed at multiple scales, that is at different up-sampling and down-sampling rates. The first property allows the network, which we named RecResNet, to progressively improve the disparity map, while the second property allows the corrections to come from different scales of analysis, addressing different types of errors in the current disparity map. We present competitive quantitative and qualitative results on the KITTI 2012 and 2015 benchmarks that surpass the accuracy of previous disparity refinement methods. Konstantinos Batsos, Philippos Mordohai |
3DV | 2 |
| 2018 | CBMV: A Coalesced Bidirectional Matching Volume for Disparity EstimationabstractRecently, there has been a paradigm shift in stereo matching with learning-based methods achieving the best results on all popular benchmarks. The success of these methods is due to the availability of training data with ground truth; training learning-based systems on these datasets has allowed them to surpass the accuracy of conventional approaches based on heuristics and assumptions. Many of these assumptions, however, had been validated extensively and hold for the majority of possible inputs. In this paper, we generate a matching volume leveraging both data with ground truth and conventional wisdom. We accomplish this by coalescing diverse evidence from a bidirectional matching process via random forest classifiers. We show that the resulting matching volume estimation method achieves similar accuracy to purely data-driven alternatives on benchmarks and that it generalizes to unseen data much better. In fact, the results we submitted to the KITTI and ETH3D benchmarks were generated using a classifier trained on the Middlebury 2014 dataset. Konstantinos Batsos, Changjiang Cai, Philippos Mordohai |
CVPR | 3 |
| 2017 | High-Resolution Stereo Matching based on Sampled Photoconsistency Computation
Chloe LeGendre, Konstantinos Batsos, Philippos Mordohai |
BMVC | 3 |
| 2017 | Special Issue on Large-Scale 3D Modeling of Urban Indoor or Outdoor Scenes from Images and Range Scans
Ioannis Stamos, Marc Pollefeys, Long Quan, Philippos Mordohai, Yasutaka Furukawa |
Comput. Vis. Image Underst. | 4 |
| 2016 | An egocentric computer vision based co-robot wheelchairabstractMotivated by the emerging needs to improve the quality of life for the elderly and disabled individuals who rely on wheelchairs for mobility, and who might have limited or no hand functionality at all, we propose an egocentric computer vision based co-robot wheelchair to enhance their mobility without hand usage. The co-robot wheelchair is built upon a typical commercial power wheelchair. The user can access 360 degrees of motion direction as well as a continuous range of speed without the use of hands via the egocentric computer vision based control we developed. The user wears an egocentric camera and collaborates with the robotic wheelchair by conveying the motion commands with head motions. Compared with previous sip-n-puff, chin-control and tongue-operated solutions to hands-free mobility, this egocentric computer vision based control system provides a more natural human robot interface. Our experiments show that this design is of higher usability and users can quickly learn to control and operate the wheelchair. Besides its convenience in manual navigation, the egocentric camera also supports novel user-robot interaction modes by enabling autonomous navigation towards a detected person or object of interest. User studies demonstrate the usability and efficiency of the proposed egocentric computer vision co-robot wheelchair. Mohammed Kutbi, Changjiang Cai, Philippos Mordohai, Gang Hua 0001 |
IROS | 5 |
| 2016 | Correctness Prediction, Accuracy Improvement and Generalization of Stereo Matching Using Supervised Learning
Aristotle Spyropoulos, Philippos Mordohai |
Int. J. Comput. Vis. | 2 |
| 2016 | Correspondence estimation for non-rigid point clouds with automatic part discovery
Hao Guo 0005, Dehai Zhu, Philippos Mordohai |
Vis. Comput. | 3 |
| 2015 | Confidence Estimation for Superpixel-Based Stereo MatchingabstractIn this paper we propose an approach for estimating the confidence of stereo matches for super pixel-based disparity estimation. To our knowledge, this is the first such method reported in the literature. Starting from a simple super pixel stereo algorithm, we present a representative set of features that can be extracted from the disparity map and the super pixel fitting process. A random forest classifier is then trained on these features to predict whether the disparity assigned to each pixel of a test disparity map is correct or not. We perform experiments on the KITTI stereo benchmark and show that our confidence estimator is very accurate in predicting which disparities are correct and which are not. We also present a post-processing algorithm for improving the accuracy of the disparity maps that exploits the confidence estimates to reject wrong disparity values and achieves significant error reduction. Rafael Gouveia, Aristotle Spyropoulos, Philippos Mordohai |
3DV | 3 |
| 2015 | Ensemble Classifier for Combining Stereo Matching AlgorithmsabstractStereo matching, as many problems in computer vision, has been addressed by a multitude of algorithms, each with its own strengths and weaknesses. Instead of following the conventional approach and trying to tune or enhance one of the algorithms so that it dominates the competition, we resign to the idea that a truly optimal algorithm may not be discovered soon and take a different approach. We present a novel methodology for combining a large number of heterogeneous algorithms that is able to clearly surpass the accuracy of the most accurate algorithms in the set. At the core of our approach is the design of an ensemble classifier trained to decide whether a particular stereo matcher is correct on a certain pixel. In addition to features describing the pixel, our feature vector encodes the agreement and disagreement between the matcher under consideration and all other matchers. This formulation leads to high accuracy in disparity estimation on the KITTI stereo benchmark. Aristotle Spyropoulos, Philippos Mordohai |
3DV | 2 |
| 2015 | Consistent 3D Background Model Estimation from Multi-viewpoint VideosabstractWe present an approach for estimating the 3D background model of a scene from a collection of synchronized videos. Unlike previous work, our method is fully automatic, does not require empty frames depicting just the background, and makes very mild assumptions about the foreground. The constraint on the cameras is that they should have sufficiently narrow baselines to enable multi-view stereo matching. Using the images and primarily the depth maps as inputs, our algorithm detects potential background pixels to generate initial per-camera background models, which are then fused to form the final, consistent 3D background model. We show results on diverse video sequences captured using different camera configurations. Despite the challenges posed by the input videos, in which some parts of the background are always occluded in the images, we are able to extract accurate models of the background that are effective in foreground segmentation. This would have been impossible using conventional background subtraction methods that operate on the frames of each camera separately. Moreover, fusion makes the per-camera background models consistent. Iraklis Tsekourakis, Philippos Mordohai |
3DV | 2 |
| 2015 | Exact bias correction and covariance estimation for stereo visionabstractWe present an approach for correcting the bias in 3D reconstruction of points imaged by a calibrated stereo rig. Our analysis is based on the observation that, due to quantization error, a 3D point reconstructed by triangulation essentially represents an entire region in space. The true location of the world point that generated the triangulated point could be anywhere in this region. We argue that the reconstructed point, if it is to represent this region in space without bias, should be located at the centroid of this region, which is not what has been done in the literature. We derive the exact geometry of these regions in space, which we call 3D cells, and we show how they can be viewed as uniform distributions of possible pre-images of the pair of corresponding pixels. By assuming a uniform distribution of points in 3D, as opposed to a uniform distribution of the projections of these 3D points on the images, we arrive at a fast and exact computation of the triangulation bias in each cell. In addition, we derive the exact covariance matrices of the 3D cells. We validate our approach in a variety of simulations ranging from 3D reconstruction to camera localization and relative motion estimation. In all cases, we are able to demonstrate a marked improvement compared to conventional techniques for small disparity values, for which bias is significant and the required corrections are large. Charles Freundlich, Michael M. Zavlanos, Philippos Mordohai |
CVPR | 3 |
| 2014 | Classification of Vehicle Parts in Unstructured 3D Point CloudsabstractUnprecedented amounts of 3D data can be acquired in urban environments, but their use for scene understanding is challenging due to varying data resolution and variability of objects in the same class. An additional challenge is due to the nature of the point clouds themselves, since they lack detailed geometric or semantic information that would aid scene understanding. In this paper we present a general algorithm for segmenting and jointly classifying object parts and the object itself. Our pipeline consists of local feature extraction, robust RANSAC part segmentation, part-level feature extraction, a structured model for parts in objects, and classification using state-of-the-art classifiers. We have tested this pipeline in a very challenging dataset that consists of real world scans of vehicles. Our contributions include the development of a segmentation and classification pipeline for objects and their parts, and a method for segmentation that is robust to the complexity of unstructured 3D points clouds, as well as a part ordering strategy for the sequential structured model and a joint feature representation between object parts. Allan Zelener, Philippos Mordohai, Ioannis Stamos |
3DV | 2 |
| 2014 | Learning to Detect Ground Control Points for Improving the Accuracy of Stereo MatchingabstractWhile machine learning has been instrumental to the ongoing progress in most areas of computer vision, it has not been applied to the problem of stereo matching with similar frequency or success. We present a supervised learning approach for predicting the correctness of stereo matches based on a random forest and a set of features that capture various forms of information about each pixel. We show highly competitive results in predicting the correctness of matches and in confidence estimation, which allows us to rank pixels according to the reliability of their assigned disparities. Moreover, we show how these confidence values can be used to improve the accuracy of disparity maps by integrating them with an MRF-based stereo algorithm. This is an important distinction from current literature that has mainly focused on sparsification by removing potentially erroneous disparities to generate quasi-dense disparity maps. Aristotle Spyropoulos, Nikos Komodakis, Philippos Mordohai |
CVPR | 3 |
| 2014 | 3D Interest Point Detection via Discriminative Learning
Leizer Teran, Philippos Mordohai |
ECCV (1) | 2 |
| 2014 | A quantitative evaluation of surface normal estimation in point cloudsabstractWe revisit a well-studied problem in the analysis of range data: surface normal estimation for a set of unorganized points. Surface normal estimation has been well-studied initially due to its theoretical appeal and more recently due to its many practical applications. The latter cover several aspects of range data analysis from plane or surface fitting to segmentation, object detection and scene analysis. Following the vast majority of the literature, we also focus our attention on techniques that operate in small neighborhoods around the point whose normal is to be estimated. We pay close attention to aspects of the implementation, such as the use of weights and normalization, that have not been studied in detail in the past. We perform quantitative evaluation on a diverse set of point clouds derived from 3D meshes, which allows us to obtain accurate ground truth. Krzysztof Jordan, Philippos Mordohai |
IROS | 2 |
| 2013 | A hybrid control approach to the Next-Best-View problem using stereo visionabstractIn this paper, we consider the problem of precisely localizing a group of stationary targets using a single stereo camera mounted on a mobile robot. In particular, assuming that at least one pair of stereo images of the targets is available, we seek to determine where to move the stereo camera so that the localization uncertainty of the targets is minimized. We call this problem the Next-Best-View problem. The advantage of using a stereo camera is that, using triangulation, the two simultaneous images can yield range and bearing measurements of the targets, as well as their uncertainty. We use a Kalman filter to fuse location and uncertainty estimates as more measurements are acquired. Our solution to the Next-Best-View problem is to iteratively minimize the fused uncertainty of the targets' locations subject to field-of-view constraints. We capture these objectives by appropriate artificial potentials on the camera's relative frame and the global frame, respectively. In particular, with every new observation, the mobile stereo camera computes the new next best view on the relative frame and subsequently realizes this view in the global frame via gradient descent on the space of robot positions and orientations, until a new observation is made. Integration of next best view with motion planning results in a hybrid system, which we illustrate in computer simulations. Charles Freundlich, Philippos Mordohai, Michael M. Zavlanos |
ICRA | 2 |
| 2013 | Integrity Verification of K-means Clustering Outsourced to Infrastructure as a Service (IaaS) ProvidersabstractThe Cloud-based infrastructure-as-a-service (IaaS) paradigm (e.g., Amazon EC2) enables a client who lacks computational resources to outsource her dataset and data mining tasks to the Cloud. However, as the Cloud may not be fully trusted, it raises serious concerns about the integrity of the mining results returned by the Cloud. To this end, in this paper, we provide a focused study about how to perform integrity verification of the k-means clustering task outsourced to an IaaS provider. We consider the untrusted sloppy IaaS service provider that intends to return wrong clustering results by terminating the iterations early to save computational cost. We develop both probabilistic and deterministic verification methods to catch the incorrect clustering result by the service provider. The deterministic method returns 100% integrity guarantee with cost that is much cheaper than executing k-means clustering locally, while the probabilistic method returns a probabilistic integrity guarantee with computational cost even cheaper than the deterministic approach. Our experimental results show that our verification methods can effectively and efficiently capture the sloppy service provider. Philippos Mordohai, Wendy Hui Wang, Hui Xiong 0001 |
SDM | 2 |
| 2012 | A Quantitative Evaluation of Confidence Measures for Stereo VisionabstractWe present an extensive evaluation of 17 confidence measures for stereo matching that compares the most widely used measures as well as several novel techniques proposed here. We begin by categorizing these methods according to which aspects of stereo cost estimation they take into account and then assess their strengths and weaknesses. The evaluation is conducted using a winner-take-all framework on binocular and multibaseline datasets with ground truth. It measures the capability of each confidence method to rank depth estimates according to their likelihood for being correct, to detect occluded pixels, and to generate low-error depth maps by selecting among multiple hypotheses for each pixel. Our work was motivated by the observation that such an evaluation is missing from the rapidly maturing stereo literature and that our findings would be helpful to researchers in binocular and multiview stereo. Philippos Mordohai |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Automatic Facial Expression Recognition using Bags of Motion WordsabstractWe present a fully automatic approach for facial expression recognition based on a representation of facial motion using a vocabulary of local motion descriptors. Previous studies have shown that motion is sufficient for recognizing expressions. Moreover, by discarding appearance after optical flow estimation, our representation is invariant to the subjects’ ethnic background, facial hair and other confounders. Unlike most facial expression recognition approaches, ours is general and not specifically tailored to faces. Annotation efforts for training are minimal, since the user does not have to label frames according to the phase of the expression, or identify facial features. Only a single expression label per sequence is required. We show results on a database of 600 video sequences. Liefei Xu, Philippos Mordohai |
BMVC | 2 |
| 2010 | Evaluation of stereo confidence indoors and outdoorsabstractWe present an extensive evaluation of 13 confidence metrics for stereo matching that compares the most widely used metrics as well as four novel techniques proposed here. We begin by categorizing the methods according to which aspects of stereo computation they take into account and, then, assess their strengths and weaknesses. The evaluation is conducted on indoor and outdoor datasets with ground truth and measures the capability of each confidence metric to rank depth estimates according to their likelihood for being correct, to detect occluded pixels and to generate low-error depth maps by selecting among multiple hypotheses for each pixel. We believe that such an evaluation is missing from the rapidly maturing stereo literature and that our findings will be helpful to researchers in binocular and multi-view stereo. Philippos Mordohai |
CVPR | 2 |
| 2010 | Detecting and parsing architecture at city scale from range dataabstractWe present a method for detecting and parsing buildings from unorganized 3D point clouds into a compact, hierarchical representation that is useful for high-level tasks. The input is a set of range measurements that cover large-scale urban environment. The desired output is a set of parse trees, such that each tree represents a semantic decomposition of a building - the nodes are roof surfaces as well as volumetric parts inferred from the observable surfaces. We model the above problem using a simple and generic grammar and use an efficient dependency parsing algorithm to generate the desired semantic description. We show how to learn the parameters of this simple grammar in order to produce correct parses of complex structures. We are able to apply our model on large point clouds and parse an entire city. Alexander Toshev, Philippos Mordohai, Ben Taskar |
CVPR | 2 |
| 2010 | Dimensionality Estimation, Manifold Learning and Function Approximation using Tensor Voting
Philippos Mordohai, Gérard G. Medioni |
J. Mach. Learn. Res. | 1 |
| 2009 | The Self-Aware Matching Measure for stereoabstractWe revisit stereo matching functions, a topic that is considered well understood, from a different angle. Our goal is to discover a transformation that operates on the cost or similarity measures between pixels in binocular stereo. This transformation should produce a new matching curve that results in higher matching accuracy. The desired transformation must have no additional parameters over those of the original matching function and must result in a new matching function that can be used by existing local, global and semi-local stereo algorithms without having to modify the algorithms. We propose a transformation that meets these requirements, taking advantage of information derived from matching the input images against themselves. We analyze the behavior of this transformation, which we call Self-Aware Matching Measure (SAMM), on a diverse set of experiments on data with ground truth. Our results show that the SAMM improves the performance of dense and semi-dense stereo. Moreover, as opposed to the current state of the art, it does not require distinctiveness to match pixels reliably. Philippos Mordohai |
ICCV | 1 |
| 2008 | Variable baseline/resolution stereoabstractWe present a novel multi-baseline, multi-resolution stereo method, which varies the baseline and resolution proportionally to depth to obtain a reconstruction in which the depth error is constant. This is in contrast to traditional stereo, in which the error grows quadratically with depth, which means that the accuracy in the near range far exceeds that of the far range. This accuracy in the near range is unnecessarily high and comes at significant computational cost. It is, however, non-trivial to reduce this without also reducing the accuracy in the far range. Many datasets, such as video captured from a moving camera, allow the baseline to be selected with significant flexibility. By selecting an appropriate baseline and resolution (realized using an image pyramid), our algorithm computes a depthmap which has these properties: 1) the depth accuracy is constant over the reconstructed volume, 2) the computational effort is spread evenly over the volume, 3) the angle of triangulation is held constant w.r.t. depth. Our approach achieves a given target accuracy with minimal computational effort, and is orders of magnitude faster than traditional stereo. David Gallup, Jan-Michael Frahm, Philippos Mordohai, Marc Pollefeys |
CVPR | 3 |
| 2008 | Object Detection from Large-Scale 3D Datasets Using Bottom-Up and Top-Down Descriptors
Alexander Patterson, Philippos Mordohai, Kostas Daniilidis |
ECCV (4) | 2 |
| 2008 | Fluid in Video: Augmenting Real Video with Simulated FluidsabstractAbstract We present a technique for coupling simulated fluid phenomena that interact with real dynamic scenes captured as a binocular video sequence. We first process the binocular video sequence to obtain a complete 3D reconstruction of the scene, including velocity information. We use stereo for the visible parts of 3D geometry and surface completion to fill the missing regions. We then perform fluid simulation within a 3D domain that contains the object, enabling one‐way coupling from the video to the fluid. In order to maintain temporal consistency of the reconstructed scene and the animated fluid across frames, we develop a geometry tracking algorithm that combines optic flow and depth information with a novel technique for “velocity completion”. The velocity completion technique uses local rigidity constraints to hypothesize a motion field for the entire 3D shape, which is then used to propagate and filter the reconstructed shape over time. This approach not only generates smoothly varying geometry across time, but also simultaneously provides the necessary boundary conditions for one‐way coupling between the dynamic geometry and the simulated fluid. Finally, we employ a GPU based scheme for rendering the synthetic fluid in the real video, taking refraction and scene texture into account. Vivek Kwatra, Philippos Mordohai, Rahul Narain, Sashi Kumar Penta, Mark T. Carlson, Marc Pollefeys, Ming C. Lin |
Comput. Graph. Forum | 2 |
| 2008 | Detailed Real-Time Urban 3D Reconstruction from Video
Marc Pollefeys, David Nistér, Jan-Michael Frahm, Amir Akbarzadeh, Philippos Mordohai, Brian Clipp, Chris Engels, David Gallup, Seon Joo Kim, Paul Merrell, C. Salmi, Sudipta N. Sinha, B. Talton, Liang Wang 0002, Qingxiong Yang, Henrik Stewénius, Ruigang Yang, Greg Welch, Herman Towles |
Int. J. Comput. Vis. | 5 |
| 2007 | Real-Time Plane-Sweeping Stereo with Multiple Sweeping DirectionsabstractRecent research has focused on systems for obtaining automatic 3D reconstructions of urban environments from video acquired at street level. These systems record enormous amounts of video; therefore a key component is a stereo matcher which can process this data at speeds comparable to the recording frame rate. Furthermore, urban environments are unique in that they exhibit mostly planar surfaces. These surfaces, which are often imaged at oblique angles, pose a challenge for many window-based stereo matchers which suffer in the presence of slanted surfaces. We present a multi-view plane-sweep-based stereo algorithm which correctly handles slanted surfaces and runs in real-time using the graphics processing unit (GPU). Our algorithm consists of (1) identifying the scene's principle plane orientations, (2) estimating depth by performing a plane-sweep for each direction, (3) combining the results of each sweep. The latter can optionally be performed using graph cuts. Additionally, by incorporating priors on the locations of planes in the scene, we can increase the quality of the reconstruction and reduce computation time, especially for uniform textureless surfaces. We demonstrate our algorithm on a variety of scenes and show the improved accuracy obtained by accounting for slanted surfaces. David Gallup, Jan-Michael Frahm, Philippos Mordohai, Qingxiong Yang, Marc Pollefeys |
CVPR | 3 |
| 2007 | Temporally Consistent Reconstruction from Multiple Video Streams Using Enhanced Belief PropagationabstractWe present an approach for 3D reconstruction from multiple video streams taken by static, synchronized and calibrated cameras that is capable of enforcing temporal consistency on the reconstruction of successive frames. Our goal is to improve the quality of the reconstruction by finding corresponding pixels in subsequent frames of the same camera using optical flow, and also to at least maintain the quality of the single time-frame reconstruction when these correspondences are wrong or cannot be found. This allows us to process scenes with fast motion, occlusions and self- occlusions where optical flow fails for large numbers of pixels. To this end, we modify the belief propagation algorithm to operate on a 3D graph that includes both spatial and temporal neighbors and to be able to discard messages from outlying neighbors. We also propose methods for introducing a bias and for suppressing noise typically observed in uniform regions. The bias encapsulates information about the background and aids in achieving a temporally consistent reconstruction and in the mitigation of occlusion related errors. We present results on publicly available real video sequences. We also present quantitative comparisons with results obtained by other researchers. E. Scott Larsen, Philippos Mordohai, Marc Pollefeys, Henry Fuchs |
ICCV | 2 |
| 2007 | Real-Time Visibility-Based Fusion of Depth MapsabstractWe present a viewpoint-based approach for the quick fusion of multiple stereo depth maps. Our method selects depth estimates for each pixel that minimize violations of visibility constraints and thus remove errors and inconsistencies from the depth maps to produce a consistent surface. We advocate a two-stage process in which the first stage generates potentially noisy, overlapping depth maps from a set of calibrated images and the second stage fuses these depth maps to obtain an integrated surface with higher accuracy, suppressed noise, and reduced redundancy. We show that by dividing the processing into two stages we are able to achieve a very high throughput because we are able to use a computationally cheap stereo algorithm and because this architecture is amenable to hardware-accelerated (GPU) implementations. A rigorous formulation based on the notion of stability of a depth estimate is presented first. It aims to determine the validity of a depth estimate by rendering multiple depth maps into the reference view as well as rendering the reference depth map into the other views in order to detect occlusions and free- space violations. We also present an approximate alternative formulation that selects and validates only one hypothesis based on confidence. Both formulations enable us to perform video-based reconstruction at up to 25 frames per second. We show results on the multi-view stereo evaluation benchmark datasets and several outdoors video sequences. Extensive quantitative analysis is performed using an accurately surveyed model of a real building as ground truth. Paul Merrell, Amir Akbarzadeh, Liang Wang 0002, Philippos Mordohai, Jan-Michael Frahm, Ruigang Yang, David Nistér, Marc Pollefeys |
ICCV | 4 |
| 2007 | Evaluation of Large Scale Scene ReconstructionabstractWe present an evaluation methodology and data for large scale video-based 3D reconstruction. We evaluate the effects of several parameters and draw conclusions that can be useful for practical systems operating in uncontrolled environments. Unlike the benchmark datasets used for the binocular stereo and multi-view reconstruction evaluations, which were collected under well-controlled conditions, our datasets are captured outdoors using video cameras mounted on a moving vehicle. As a result, the videos are much more realistic and include phenomena such as exposure changes from viewing both bright and dim surfaces, objects at varying distances from the camera, and objects of varying size and degrees of texture. The dataset includes ground truth models and precise camera pose information. We also present an evaluation methodology applicable to reconstructions of large scale environments. We evaluate the accuracy and completeness of reconstructions obtained by two fast, visibility-based depth map fusion algorithms as parameters vary. Paul Merrell, Philippos Mordohai, Jan-Michael Frahm, Marc Pollefeys |
ICCV | 2 |
| 2007 | Multi-View Stereo via Graph Cuts on the Dual of an Adaptive Tetrahedral MeshabstractWe formulate multi-view 3D shape reconstruction as the computation of a minimum cut on the dual graph of a semi- regular, multi-resolution, tetrahedral mesh. Our method does not assume that the surface lies within a finite band around the visual hull or any other base surface. Instead, it uses photo-consistency to guide the adaptive subdivision of a coarse mesh of the bounding volume. This generates a multi-resolution volumetric mesh that is densely tesselated in the parts likely to contain the unknown surface. The graph-cut on the dual graph of this tetrahedral mesh produces a minimum cut corresponding to a triangulated surface that minimizes a global surface cost functional. Our method makes no assumptions about topology and can recover deep concavities when enough cameras observe them. Our formulation also allows silhouette constraints to be enforced during the graph-cut step to counter its inherent bias for producing minimal surfaces. Local shape refinement via surface deformation is used to recover details in the reconstructed surface. Reconstructions of the Multi- View Stereo Evaluation benchmark datasets and other real datasets show the effectiveness of our method. Sudipta N. Sinha, Philippos Mordohai, Marc Pollefeys |
ICCV | 2 |
| 2006 | Stereo Using Monocular Cues within the Tensor Voting FrameworkabstractWe address the fundamental problem of matching in two static images. The remaining challenges are related to occlusion and lack of texture. Our approach addresses these difficulties within a perceptual organization framework, considering both binocular and monocular cues. Initially, matching candidates for all pixels are generated by a combination of matching techniques. The matching candidates are then embedded in disparity space, where perceptual organization takes place in 3D neighborhoods and, thus, does not suffer from problems associated with scanline or image neighborhoods. The assumption is that correct matches produce salient, coherent surfaces, while wrong ones do not. Matching candidates that are consistent with the surfaces are kept and grouped into smooth layers. Thus, we achieve surface segmentation based on geometric and not photometric properties. Surface overextensions, which are due to occlusion, can be corrected by removing matches whose projections are not consistent in color with their neighbors of the same surface in both images. Finally, the projections of the refined surfaces on both images are used to obtain disparity hypotheses for unmatched pixels. The final disparities are selected after a second tensor voting stage, during which information is propagated from more reliable pixels to less reliable ones. We present results on widely used benchmark stereo pairs. Philippos Mordohai, Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Unsupervised Dimensionality Estimation and Manifold Learning in high-dimensional Spaces by Tensor Voting
Philippos Mordohai, Gérard G. Medioni |
IJCAI | 1 |
| 2004 | Stereo Using Monocular Cues within the Tensor Voting Framework
Philippos Mordohai, Gérard G. Medioni |
ECCV (4) | 1 |
| 2004 | First Order Augmentation to Tensor Voting for Boundary Inference and Multiscale Analysis in 3DabstractMost computer vision applications require the reliable detection of boundaries. In the presence of outliers, missing data, orientation discontinuities, and occlusion, this problem is particularly challenging. We propose to address it by complementing the tensor voting framework, which was limited to second order properties, with first order representation and voting. First order voting fields and a mechanism to vote for 3D surface and volume boundaries and curve endpoints in 3D are defined. Boundary inference is also useful for a second difficult problem in grouping, namely, automatic scale selection. We propose an algorithm that automatically infers the smallest scale that can preserve the finest details. Our algorithm then proceeds with progressively larger scales to ensure continuity where it has not been achieved. Therefore, the proposed approach does not oversmooth features or delay the handling of boundaries and discontinuities until model misfit occurs. The interaction of smooth features, boundaries, and outliers is accommodated by the unified representation, making possible the perceptual organization of data in curves, surfaces, volumes, and their boundaries simultaneously. We present results on a variety of data sets to show the efficacy of the improved formalism. Wai-Shun Tong, Chi-Keung Tang, Philippos Mordohai, Gérard G. Medioni |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2002 | Inference of Segmented Overapping Surfaces from Binocular StereoabstractPresents an integrated approach to the derivation of scene descriptions from a pair of stereo images, where the steps of feature correspondence and surface reconstruction are addressed within the same framework. Special attention is given to the development of a methodology with general applicability. In order to handle the issues of noise, lack of image features, surface discontinuities and regions that are visible in one image only, we adopt a tensor representation for the data and introduce a robust computational technique called tensor voting for information propagation. The key contributions of this paper are twofold. First, we introduce "saliency" instead of correlation scores as the criterion to determine the correctness of matches and the integration of feature matching and structure extraction. Second, our tensor representation and voting as a tool enables us to perform the complex computations associated with the formulation of the stereo problem in 3D at a reasonable computational cost. We illustrate the steps on an example, then provide results on both random dot stereograms and real stereo pairs, all processed with the same parameter set. Mi-Suen Lee, Gérard G. Medioni, Philippos Mordohai |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |