Long Quan

dblp:04/575 · DBLP profile ↗
← Back
169ranked-venue papers
27as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 136 · 23 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 127 · 15 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1Theory of computation · 1
YearPublicationVenuePosition
2026 LFG: Local Controllable 3D Generation Using Latent Flexible Grid Representation
Kaiyi Zhang 0002, Long Quan
ICPR (1)3
2025 DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos
abstract
Estimating video depth in open-world scenarios is challenging due to the diversity of videos in appearance, content motion, camera movement, and length. We present DepthCrafter, an innovative method for generating temporally consistent long depth sequences with intricate details for open-world videos, without requiring any supplementary information such as camera poses or optical flow. The generalization ability to open-world videos is achieved by training the video-to-depth model from a pretrained image-to-video diffusion model, through our meticulously designed three-stage training strategy. Our training approach enables the model to generate depth sequences with variable lengths at one time, up to 110 frames, and harvest both precise depth details and rich content diversity from realistic and synthetic datasets. We also propose an inference strategy that can process extremely long videos through segment-wise estimation and seamless stitching. Comprehensive evaluations on multiple datasets reveal that DepthCrafter achieves state-of-the-art performance in open-world video depth estimation under zero-shot settings. Furthermore, DepthCrafter facilitates various downstream applications, including depth-based visual effects and conditional video generation.
Wenbo Hu 0002, Xiangjun Gao, Xiaoyu Li 0002, Sijie Zhao, Xiaodong Cun, Yong Zhang 0034, Long Quan, Ying Shan
CVPR7
2025 Mani-GS: Gaussian Splatting Manipulation with Triangular Mesh
abstract
Neural 3D representations, such as Neural Radiation Fields (NeRF), excel at producing photorealistic rendering results but lack the flexibility for manipulation and editing which is crucial for content creation. However, manipulating NeRF is not highly controllable and requires a long training and inference time. With the emergence of 3D Gaussian Splatting (3DGS), extremely high-fidelity novel view synthesis can be achieved using an explicit point-based 3D representation with much faster training and rendering speed. However, there is still a lack of effective means to manipulate 3DGS freely while maintaining rendering quality. In this work, we aim to tackle the challenge of achieving manipulable photo-realistic rendering. We propose to utilize a triangular mesh to manipulate 3DGS directly with self-adaptation. This approach reduces the need to design various algorithms for different types of 3DGS manipulation. By utilizing a triangle shape-aware Gaussian binding and adapting method, we can achieve 3DGS manipulation and preserve high-fidelity rendering. In addition, our method is also effective with inaccurate meshes extracted from 3DGS. Experiments demonstrate our method’s effectiveness and superiority over baseline approaches.
Xiangjun Gao, Xiaoyu Li 0002, Yiyu Zhuang, Qi Zhang 0029, Wenbo Hu 0002, Chaopeng Zhang, Yao Yao 0008, Ying Shan, Long Quan
CVPR9
2025 Matrix3D: Large Photogrammetry Model All-in-One
abstract
We present Matrix3D, a unified model that performs several photogrammetry subtasks, including pose estimation, depth prediction, and novel view synthesis using just the same model. Matrix3D utilizes a multi-modal diffusion transformer (DiT) to integrate transformations across several modalities, such as images, camera parameters, and depth maps. The key to Matrix3D’s large-scale multi-modal training lies in the incorporation of a mask learning strategy. This enables full-modality model training even with partially complete data, such as bi-modality data of image-pose and image-depth pairs, thus significantly increases the pool of available training data. Matrix3D demonstrates state-of-the-art performance in pose estimation and novel view synthesis tasks. Additionally, it offers fine-grained control through multi-round interactions, making it an innovative tool for 3D content creation. Project page: https://nju-3dv.github.io/projects/matrix3d.
Yuanxun Lu, Jingyang Zhang, Tian Fang, Jean-Daniel Nahmias, Yanghai Tsin, Long Quan, Xun Cao, Yao Yao 0008, Shiwei Li 0001
CVPR6
2024 ConTex-Human: Free-View Rendering of Human from a Single Image with Texture-Consistent Synthesis
abstract
In this work, we propose a method to address the chal-lenge of rendering a 3D human from a single image in a free-view manner. Some existing approaches could achieve this by using generalizable pixel-aligned implicit fields to reconstruct a textured mesh of a human or by employing a 2D diffusion model as guidance with the Score Distillation Sampling (SDS) method, to lift the 2D image into 3D space. However, a generalizable implicit field often results in an over-smooth texture field, while the SDS method tends to lead to a texture-inconsistent novel view with the input image. In this paper, we introduce a texture-consistent back view synthesis module that could transfer the reference im-age content to the back view through depth and text-guided attention injection. Moreover, to alleviate the color distortion that occurs in the side region, we propose a visibility-aware patch consistency regularization for texture mapping and refinement combined with the synthesized back view texture. With the above techniques, we can achieve high-fidelity and texture-consistent human rendering from a single image. Experiments conducted on both real and synthetic data demonstrate the effectiveness of our method and show that our approach outperforms previous baseline methods.
Xiangjun Gao, Xiaoyu Li 0002, Chaopeng Zhang, Qi Zhang 0029, Yan-Pei Cao 0001, Ying Shan, Long Quan
CVPR7
2024 Direct2.5: Diverse Text-to-3D Generation via Multi-view 2.5D Diffusion
abstract
Recent advances in generative AI have unveiled significant potential for the creation of 3D content. However, current methods either apply a pre-trained 2D diffusion model with the time-consuming score distillation sampling (SDS), or a direct 3D diffusion model trained on limited 3D data losing generation diversity. In this work, we approach the problem by employing a multi-view 2.5D diffusion fine-tuned from a pre-trained 2D diffusion model. The multi-view 2.5D diffusion directly models the structural distribution of 3D data, while still maintaining the strong generalization ability of the original 2D diffusion model, filling the gap between 2D diffusion-based and direct 3D diffusion-based methods for 3D content generation. During inference, multi-view normal maps are generated using the 2.5D diffusion, and a novel differentiable rasterization scheme is introduced to fuse the almost consistent multi-view normal maps into a consistent 3D model. We further design a normal-conditioned multi-view image generation module for fast appearance generation given the 3D geometry. Our method is a one-pass diffusion process and does not require any SDS optimization as post-processing. We demonstrate through extensive experiments that, our direct 2.5D generation with the specially-designed fusion scheme can achieve diverse, mode-seeking-free, and high-fidelity 3D content generation in only 10 seconds. Project page: https://nju-3dv.github.io/projects/direct25.
Yuanxun Lu, Jingyang Zhang, Shiwei Li 0001, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan, Xun Cao, Yao Yao 0008
CVPR7
2024 JointNet: Extending Text-to-Image Diffusion for Dense Distribution Modeling
abstract
We introduce JointNet, a novel neural network architecture for modeling the joint distribution of images and an additional dense modality (e.g., depth maps). JointNet is extended from a pre-trained text-to-image diffusion model, where a copy of the original network is created for the new dense modality branch and is densely connected with the RGB branch. The RGB branch is locked during network fine-tuning, which enables efficient learning of the new modality distribution while maintaining the strong generalization ability of the large-scale pre-trained diffusion model. We demonstrate the effectiveness of JointNet by using the RGB-D diffusion as an example and through extensive experiments, showcasing its applicability in a variety of applications, including joint RGB-D generation, dense depth prediction, depth-conditioned image generation, and high-resolution 3D panorama generation.
Jingyang Zhang, Shiwei Li 0001, Yuanxun Lu, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan, Yao Yao 0008
ICLR7
2023 NeILF++: Inter-Reflectable Light Fields for Geometry and Material Estimation
abstract
We present a novel differentiable rendering framework for joint geometry, material, and lighting estimation from multi-view images. In contrast to previous methods which assume a simplified environment map or co-located flashlights, in this work, we formulate the lighting of a static scene as one neural incident light field (NeILF) and one outgoing neural radiance field (NeRF). The key insight of the proposed method is the union of the incident and outgoing light fields through physically-based rendering and inter-reflections between surfaces, making it possible to disentangle the scene geometry, material, and lighting from image observations in a physically-based manner. The proposed incident light and inter-reflection framework can be easily applied to other NeRF systems. We show that our method can not only decompose the outgoing radiance into incident lights and surface materials, but also serve as a surface refinement module that further improves the reconstruction detail of the neural surface. We demonstrate on several datasets that the proposed method is able to achieve state-of-the-art results in terms of geometry reconstruction quality, material estimation accuracy, and the fidelity of novel view rendering.
Jingyang Zhang, Yao Yao 0008, Shiwei Li 0001, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan
ICCV8
2022 Critical Regularizations for Neural Surface Reconstruction in the Wild
abstract
Neural implicit functions have recently shown promising results on surface reconstructions from multiple views. However, current methods still suffer from excessive time complexity and poor robustness when reconstructing unbounded or complex scenes. In this paper, we present RegSDF, which shows that proper point cloud supervisions and geometry regularizations are sufficient to produce high-quality and robust reconstruction results. Specifically, RegSDF takes an additional oriented point cloud as input, and optimizes a signed distance field and a surface light field within a differentiable rendering framework. We also introduce the two critical regularizations for this optimization. The first one is the Hessian regularization that smoothly diffuses the signed distance values to the entire distance field given noisy and incomplete input. And the second one is the minimal surface regularization that compactly interpolates and extrapolates the missing geometry. Extensive experiments are conducted on DTU, Blended-MVS, and Tanks and Temples datasets. Compared with recent neural surface reconstruction approaches, RegSDF is able to reconstruct surfaces with fine details even for open scenes with complex topologies and unstructured camera trajectories.
Jingyang Zhang, Yao Yao 0008, Shiwei Li 0001, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan
CVPR7
2022 ASpanFormer: Detector-Free Image Matching with Adaptive Span Transformer
Zixin Luo, Lei Zhou 0011, Yurun Tian, Mingmin Zhen, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan
ECCV (32)9
2022 NeILF: Neural Incident Light Field for Physically-based Material Estimation
Yao Yao 0008, Jingyang Zhang, Yihang Qu, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan
ECCV (31)8
2022 OANet: Learning Two-View Correspondences and Geometry Using Order-Aware Network
abstract
Establishing correct correspondences between two images should consider both local and global spatial context. Given putative correspondences of feature points in two views, in this paper, we propose Order-Aware Network, which infers the probabilities of correspondences being inliers and regresses the relative pose encoded by the essential or fundamental matrix. Specifically, this proposed network is built hierarchically and comprises three operations. First, to capture the local context of sparse correspondences, the network clusters unordered input correspondences by learning a soft assignment matrix. These clusters are in canonical order and invariant to input permutations. Next, the clusters are spatially correlated to encode the global context of correspondences. After that, the context-encoded clusters are interpolated back to the original size and position to build a hierarchical architecture. We intensively experiment on both outdoor and indoor datasets. The accuracy of the two-view geometry and correspondences are significantly improved over the state-of-the-arts. Besides, based on the proposed method and advanced local feature, we won the first place in CVPR 2019 image matching workshop challenge and also achieve state-of-the-art results in the Visual Localization benchmark. Code is available at https://github.com/zjhthu/OANet.
Dawei Sun 0007, Zixin Luo, Anbang Yao, Lei Zhou 0011, Tianwei Shen, Yurong Chen 0001, Long Quan, Hongen Liao
IEEE Trans. Pattern Anal. Mach. Intell.9
2021 Learning to Match Features with Seeded Graph Matching Network
abstract
Matching local features across images is a fundamental problem in computer vision. Targeting towards high accuracy and efficiency, we propose Seeded Graph Matching Network, a graph neural network with sparse structure to reduce redundant connectivity and learn compact representation. The network consists of 1) Seeding Module, which initializes the matching by generating a small set of reliable matches as seeds. 2) Seeded Graph Neural Network, which utilizes seed matches to pass messages within/across images and predicts assignment costs. Three novel operations are proposed as basic elements for message passing: 1) Attentional Pooling, which aggregates keypoint features within the image to seed matches. 2) Seed Filtering, which enhances seed features and exchanges messages across images. 3) Attentional Unpooling, which propagates seed features back to original keypoints. Experiments show that our method reduces computational and memory complexity significantly compared with typical attention-based networks while competitive or higher performance is achieved.
Zixin Luo, Lei Zhou 0011, Xuyang Bai, Zeyu Hu, Chiew-Lan Tai, Long Quan
ICCV8
2021 Learning Signed Distance Field for Multi-view Surface Reconstruction
abstract
Recent works on implicit neural representations have shown promising results for multi-view surface reconstruction. However, most approaches are limited to relatively simple geometries and usually require clean object masks for reconstructing complex and concave objects. In this work, we introduce a novel neural surface reconstruction framework that leverages the knowledge of stereo matching and feature consistency to optimize the implicit surface representation. More specifically, we apply a signed distance field (SDF) and a surface light field to represent the scene geometry and appearance respectively. The SDF is directly supervised by geometry from stereo matching, and is refined by optimizing the multi-view feature consistency and the fidelity of rendered images. Our method is able to improve the robustness of geometry estimation and support reconstruction of complex scene topologies. Extensive experiments have been conducted on DTU, EPFL and Tanks and Temples datasets. Compared to previous state-of-the-art methods, our method achieves better mesh reconstruction in wide open scenes without masks as input.
Jingyang Zhang, Yao Yao 0008, Long Quan
ICCV3
2021 A Survey of Powertrain Technologies for Energy-Efficient Heavy-Duty Machinery
abstract
This article presents a comprehensive, multidisciplinary overview of the development of powertrain technologies for energy-efficient heavy-duty earthmoving machines. The heavy-duty earthmoving equipment industry has been among the biggest contributors to emissions globally. However, due to high power demand and multidisciplinary powertrain structures, improving the energy efficiency of heavy-duty mobile machines has been a pressing and challenging task in the industry. To cope with this challenge, hydraulics and power electronics (PE) have been the key driving forces. As such, the relative developments in both fields are covered in this article. For hydraulics, developments of efficient hydraulic circuits will be overviewed in detail along with the introduction of hydraulic energy recovery technologies. In addition, developments of PE architectures in hybrid and electrified machines will be introduced. Furthermore, potential medium-voltage dc mining site power distribution and the valves of wide bandgap devices will also be discussed with the hope to open up new research opportunities in PE. Moreover, emerging hybrid electrohydraulic drive technology is introduced. Based on the overview in this article, it is anticipated that electrohydraulic hybridization will be the future trend in the earthmoving machine industry. Deeper collaboration between the two areas is desirable.
Zhongyi Quan, Zhongbao Wei, Yunwei Li 0001, Long Quan
Proc. IEEE5
2020 BlendedMVS: A Large-Scale Dataset for Generalized Multi-View Stereo Networks
abstract
While deep learning has recently achieved great success on multi-view stereo (MVS), limited training data makes the trained model hard to be generalized to unseen scenarios. Compared with other computer vision tasks, it is rather difficult to collect a large-scale MVS dataset as it requires expensive active scanners and labor-intensive process to obtain ground truth 3D structures. In this paper, we introduce BlendedMVS, a novel large-scale dataset, to provide sufficient training ground truth for learning-based MVS. To create the dataset, we apply a 3D reconstruction pipeline to recover high-quality textured meshes from images of well-selected scenes. Then, we render these mesh models to color images and depth maps. To introduce the ambient lighting information during training, the rendered color images are further blended with the input images to generate the training input. Our dataset contains over 17k high-resolution images covering a variety of scenes, including cities, architectures, sculptures and small objects. Extensive experiments demonstrate that BlendedMVS endows the trained model with significantly better generalization ability compared with other MVS datasets. The dataset and pretrained models are available at https://github.com/YoYo000/BlendedMVS.
Yao Yao 0008, Zixin Luo, Shiwei Li 0001, Jingyang Zhang, Yufan Ren, Lei Zhou 0011, Tian Fang, Long Quan
CVPR8
2020 D3Feat: Joint Learning of Dense Detection and Description of 3D Local Features
abstract
A successful point cloud registration often lies on robust establishment of sparse matches through discriminative 3D local features. Despite the fast evolution of learning-based 3D feature descriptors, little attention has been drawn to the learning of 3D feature detectors, even less for a joint learning of the two tasks. In this paper, we leverage a 3D fully convolutional network for 3D point clouds, and propose a novel and practical learning mechanism that densely predicts both a detection score and a description feature for each 3D point. In particular, we propose a keypoint selection strategy that overcomes the inherent density variations of 3D point clouds, and further propose a self-supervised detector loss guided by the on-the-fly feature matching results during training. Finally, our method achieves state-of-the-art results in both indoor and outdoor scenarios, evaluated on 3DMatch and KITTI datasets, and shows its strong generalization ability on the ETH dataset. Towards practical use, we show that by adopting a reliable feature detector, sampling a smaller number of features is sufficient to achieve accurate and fast point cloud alignment.
Xuyang Bai, Zixin Luo, Lei Zhou 0011, Hongbo Fu 0001, Long Quan, Chiew-Lan Tai
CVPR5
2020 ASLFeat: Learning Local Features of Accurate Shape and Localization
abstract
This work focuses on mitigating two limitations in the joint learning of local feature detectors and descriptors. First, the ability to estimate the local shape (scale, orientation, etc.) of feature points is often neglected during dense feature extraction, while the shape-awareness is crucial to acquire stronger geometric invariance. Second, the localization accuracy of detected keypoints is not sufficient to reliably recover camera geometry, which has become the bottleneck in tasks such as 3D reconstruction. In this paper, we present ASLFeat, with three light-weight yet effective modifications to mitigate above issues. First, we resort to deformable convolutional networks to densely estimate and apply local transformation. Second, we take advantage of the inherent feature hierarchy to restore spatial resolution and low-level details for accurate keypoint localization. Finally, we use a peakiness measurement to relate feature responses and derive more indicative detection scores. The effect of each modification is thoroughly studied, and the evaluation is extensively conducted across a variety of practical scenarios. State-of-the-art results are reported that demonstrate the superiority of our methods.
Zixin Luo, Lei Zhou 0011, Xuyang Bai, Yao Yao 0008, Shiwei Li 0001, Tian Fang, Long Quan
CVPR9
2020 Joint Semantic Segmentation and Boundary Detection Using Iterative Pyramid Contexts
abstract
In this paper, we present a joint multi-task learning framework for semantic segmentation and boundary detection. The critical component in the framework is the iterative pyramid context module (PCM), which couples two tasks and stores the shared latent semantics to interact between the two tasks. For semantic boundary detection, we propose the novel spatial gradient fusion to suppress non-semantic edges. As semantic boundary detection is the dual task of semantic segmentation, we introduce a loss function with boundary consistency constraint to improve the boundary pixel accuracy for semantic segmentation. Our extensive experiments demonstrate superior performance over state-of-the-art works, not only in semantic segmentation but also in semantic boundary detection. In particular, a mean IoU score of 81.8% on Cityscapes test set is achieved without using coarse data or any external data for semantic segmentation. For semantic boundary detection, we improve over previous state-of-the-art works by 9.9% in terms of AP and 6.8% in terms of MF(ODS).
Mingmin Zhen, Jinglu Wang, Lei Zhou 0011, Shiwei Li 0001, Tianwei Shen, Jiaxiang Shang, Tian Fang, Long Quan
CVPR8
2020 KFNet: Learning Temporal Camera Relocalization Using Kalman Filtering
abstract
Temporal camera relocalization estimates the pose with respect to each video frame in sequence, as opposed to one-shot relocalization which focuses on a still image. Even though the time dependency has been taken into account, current temporal relocalization methods still generally underperform the state-of-the-art one-shot approaches in terms of accuracy. In this work, we improve the temporal relocalization method by using a network architecture that incorporates Kalman filtering (KFNet) for online camera relocalization. In particular, KFNet extends the scene coordinate regression problem to the time domain in order to recursively establish 2D and 3D correspondences for the pose determination. The network architecture design and the loss formulation are based on Kalman filtering in the context of Bayesian learning. Extensive experiments on multiple relocalization benchmarks demonstrate the high accuracy of KFNet at the top of both one-shot and temporal relocalization approaches.
Lei Zhou 0011, Zixin Luo, Tianwei Shen, Mingmin Zhen, Yao Yao 0008, Tian Fang, Long Quan
CVPR8
2020 Self-Supervised Monocular 3D Face Reconstruction by Occlusion-Aware Multi-view Geometry Consistency
Jiaxiang Shang, Tianwei Shen, Shiwei Li 0001, Lei Zhou 0011, Mingmin Zhen, Tian Fang, Long Quan
ECCV (15)7
2020 Learning Discriminative Feature with CRF for Unsupervised Video Object Segmentation
Mingmin Zhen, Shiwei Li 0001, Lei Zhou 0011, Jiaxiang Shang, Haoan Feng, Tian Fang, Long Quan
ECCV (27)7
2020 Stochastic Bundle Adjustment for Efficient and Scalable 3D Reconstruction
Lei Zhou 0011, Zixin Luo, Mingmin Zhen, Tianwei Shen, Shiwei Li 0001, Zhuofei Huang, Tian Fang, Long Quan
ECCV (15)8
2020 Learning Stereo Matchability in Disparity Regression Networks
abstract
Learning-based stereo matching has recently achieved promising results, yet still suffers difficulties in establishing reliable matches in weakly matchable regions that are textureless, non-Lambertian, or occluded. In this paper, we address this challenge by proposing a stereo matching network that considers pixel-wise matchability. Specifically, the network jointly regresses disparity and matchability maps from 3D probability volume through expectation and entropy operations. Next, a learned attenuation is applied as the robust loss function to alleviate the influence of weakly matchable pixels in the training. Finally, a matchability-aware disparity refinement is introduced to improve the depth inference in weakly matchable regions. The proposed deep stereo matchability (DSM) framework can improve the matching result or accelerate the computation while still guaranteeing the quality. Moreover, the DSM framework is portable to many recent stereo networks. Extensive experiments are conducted on Scene Flow and KITTI stereo datasets to demonstrate the effectiveness of the proposed framework over the state-of-the-art learning-based stereo methods.
Jingyang Zhang, Yao Yao 0008, Zixin Luo, Shiwei Li 0001, Tianwei Shen, Tian Fang, Long Quan
ICPR7
2020 Distributed Very Large Scale Bundle Adjustment by Global Camera Consensus
abstract
The increasing scale of Structure-from-Motion is fundamentally limited by the conventional optimization framework for the all-in-one global bundle adjustment. In this paper, we propose a distributed approach to coping with this global bundle adjustment for very large scale Structure-from-Motion computation. First, we derive the distributed formulation from the classical optimization algorithm ADMM, Alternating Direction Method of Multipliers, based on the global camera consensus. Then, we analyze the conditions under which the convergence of this distributed optimization would be guaranteed. In particular, we adopt over-relaxation and self-adaption schemes to improve the convergence rate. After that, we propose to split the large scale camera-point visibility graph in order to reduce the communication overheads of the distributed computing. The experiments on both public large scale SfM data-sets and our very large scale aerial photo sets demonstrate that the proposed distributed method clearly outperforms the state-of-the-art method in efficiency and accuracy.
Siyu Zhu 0001, Tianwei Shen, Lei Zhou 0011, Zixin Luo, Tian Fang, Long Quan
IEEE Trans. Pattern Anal. Mach. Intell.7
2019 Learning Fully Dense Neural Networks for Image Semantic Segmentation
abstract
Semantic segmentation is pixel-wise classification which retains critical spatial information. The “feature map reuse” has been commonly adopted in CNN based approaches to take advantage of feature maps in the early layers for the later spatial reconstruction. Along this direction, we go a step further by proposing a fully dense neural network with an encoderdecoder structure that we abbreviate as FDNet. For each stage in the decoder module, feature maps of all the previous blocks are adaptively aggregated to feedforward as input. On the one hand, it reconstructs the spatial boundaries accurately. On the other hand, it learns more efficiently with the more efficient gradient backpropagation. In addition, we propose the boundary-aware loss function to focus more attention on the pixels near the boundary, which boosts the “hard examples” labeling. We have demonstrated the best performance of the FDNet on the two benchmark datasets: PASCAL VOC 2012, NYUDv2 over previous works when not considering training on other datasets.
Mingmin Zhen, Jinglu Wang, Lei Zhou 0011, Tian Fang, Long Quan
AAAI5
2019 Recurrent MVSNet for High-Resolution Multi-View Stereo Depth Inference
abstract
Deep learning has recently demonstrated its excellent performance for multi-view stereo (MVS). However, one major limitation of current learned MVS approaches is the scalability: the memory-consuming cost volume regularization makes the learned MVS hard to be applied to high-resolution scenes. In this paper, we introduce a scalable multi-view stereo framework based on the recurrent neural network. Instead of regularizing the entire 3D cost volume in one go, the proposed Recurrent Multi-view Stereo Network (R-MVSNet) sequentially regularizes the 2D cost maps along the depth direction via the gated recurrent unit (GRU). This reduces dramatically the memory consumption and makes high-resolution reconstruction feasible. We first show the state-of-the-art performance achieved by the proposed R-MVSNet on the recent MVS benchmarks. Then, we further demonstrate the scalability of the proposed method on several large-scale scenarios, where previous learned approaches often fail due to the memory constraint. Code is available at https://github.com/YoYo000/MVSNet.
Yao Yao 0008, Zixin Luo, Shiwei Li 0001, Tianwei Shen, Tian Fang, Long Quan
CVPR6
2019 Cross-Atlas Convolution for Parameterization Invariant Learning on Textured Mesh Surface
abstract
We present a convolutional network architecture for direct feature learning on mesh surfaces through their atlases of texture maps. The texture map encodes the parameterization from 3D to 2D domain, rendering not only RGB values but also rasterized geometric features if necessary. Since the parameterization of texture map is not pre-determined, and depends on the surface topologies, we therefore introduce a novel cross-atlas convolution to recover the original mesh geodesic neighborhood, so as to achieve the invariance property to arbitrary parameterization. The proposed module is integrated into classification and segmentation architectures, which takes the input texture map of a mesh, and infers the output predictions. Our method not only shows competitive performances on classification and segmentation public benchmarks, but also paves the way for the broad mesh surfaces learning.
Shiwei Li 0001, Zixin Luo, Mingmin Zhen, Yao Yao 0008, Tianwei Shen, Tian Fang, Long Quan
CVPR7
2019 ContextDesc: Local Descriptor Augmentation With Cross-Modality Context
abstract
Most existing studies on learning local features focus on the patch-based descriptions of individual keypoints, whereas neglecting the spatial relations established from their keypoint locations. In this paper, we go beyond the local detail representation by introducing context awareness to augment off-the-shelf local feature descriptors. Specifically, we propose a unified learning framework that leverages and aggregates the cross-modality contextual information, including (i) visual context from high-level image representation, and (ii) geometric context from 2D keypoint distribution. Moreover, we propose an effective N-pair loss that eschews the empirical hyper-parameter search and improves the convergence. The proposed augmentation scheme is lightweight compared with the raw local feature description, meanwhile improves remarkably on several large-scale benchmarks with diversified scenes, which demonstrates both strong practicality and generalization ability in geometric matching applications.
Zixin Luo, Tianwei Shen, Lei Zhou 0011, Yao Yao 0008, Shiwei Li 0001, Tian Fang, Long Quan
CVPR8
2019 Learning Two-View Correspondences and Geometry Using Order-Aware Network
abstract
Establishing correspondences between two images requires both local and global spatial context. Given putative correspondences of feature points in two views, in this paper, we propose Order-Aware Network, which infers the probabilities of correspondences being inliers and regresses the relative pose encoded by the essential matrix. Specifically, this proposed network is built hierarchically and comprises three novel operations. First, to capture the local context of sparse correspondences, the network clusters unordered input correspondences by learning a soft assignment matrix. These clusters are in a canonical order and invariant to input permutations. Next, the clusters are spatially correlated to form the global context of correspondences. After that, the context-encoded clusters are recovered back to the original size through a proposed upsampling operator. We intensively experiment on both outdoor and indoor datasets. The accuracy of the two-view geometry and correspondences are significantly improved over the state-of-the-arts.
Dawei Sun 0007, Zixin Luo, Anbang Yao, Lei Zhou 0011, Tianwei Shen, Yurong Chen 0001, Hongen Liao, Long Quan
ICCV9
2019 Beyond Photometric Loss for Self-Supervised Ego-Motion Estimation
abstract
Accurate relative pose is one of the key components in visual odometry (VO) and simultaneous localization and mapping (SLAM). Recently, the self-supervised learning framework that jointly optimizes the relative pose and target image depth has attracted the attention of the community. Previous works rely on the photometric error generated from depths and poses between adjacent frames, which contains large systematic error under realistic scenes due to reflective surfaces and occlusions. In this paper, we bridge the gap between geometric loss and photometric loss by introducing the matching loss constrained by epipolar geometry in a self-supervised framework. Evaluated on the KITTI dataset, our method outperforms the state-of-the-art unsupervised egomotion estimation methods by a large margin. The code and data are available at https://github.com/hlzz/DeepMatchVO.
Tianwei Shen, Zixin Luo, Lei Zhou 0011, Hanyu Deng, Tian Fang, Long Quan
ICRA7
2018 Matchable Image Retrieval by Learning from Surface Reconstruction
Tianwei Shen, Zixin Luo, Lei Zhou 0011, Siyu Zhu 0001, Tian Fang, Long Quan
ACCV (1)7
2018 Reconstructing Thin Structures of Manifold Surfaces by Integrating Spatial Curves
abstract
The manifold surface reconstruction in multi-view stereo often fails in retaining thin structures due to incomplete and noisy reconstructed point clouds. In this paper, we address this problem by leveraging spatial curves. The curve representation in nature is advantageous in modeling thin and elongated structures, implying topology and connectivity information of the underlying geometry, which exactly compensates the weakness of scattered point clouds. We present a novel surface reconstruction method using both curves and point clouds. First, we propose a 3D curve reconstruction algorithm based on the initialize-optimize-extend strategy. Then, tetrahedra are constructed from points and curves, where the volumes of thin structures are robustly preserved by the Curve-conformed Delaunay Refinement. Finally, the mesh surface is extracted from tetrahedra by a graph optimization. The method has been intensively evaluated on both synthetic and real-world datasets, showing significant improvements over state-of-the-art methods.
Shiwei Li 0001, Yao Yao 0008, Tian Fang, Long Quan
CVPR4
2018 Very Large-Scale Global SfM by Distributed Motion Averaging
abstract
Global Structure-from-Motion (SfM) techniques have demonstrated superior efficiency and accuracy than the conventional incremental approach in many recent studies. This work proposes a divide-and-conquer framework to solve very large global SfM at the scale of millions of images. Specifically, we first divide all images into multiple partitions that preserve strong data association for well-posed and parallel local motion averaging. Then, we solve a global motion averaging that determines cameras at partition boundaries and a similarity transformation per partition to register all cameras in a single coordinate frame. Finally, local and global motion averaging are iterated until convergence. Since local camera poses are fixed during the global motion average, we can avoid caching the whole reconstruction in memory at once. This distributed framework significantly enhances the efficiency and robustness of large-scale motion averaging.
Siyu Zhu 0001, Lei Zhou 0011, Tianwei Shen, Tian Fang, Ping Tan 0002, Long Quan
CVPR7
2018 GeoDesc: Learning Local Descriptors by Integrating Geometry Constraints
Zixin Luo, Tianwei Shen, Lei Zhou 0011, Siyu Zhu 0001, Yao Yao 0008, Tian Fang, Long Quan
ECCV (9)8
2018 MVSNet: Depth Inference for Unstructured Multi-view Stereo
Yao Yao 0008, Zixin Luo, Shiwei Li 0001, Tian Fang, Long Quan
ECCV (8)5
2018 Learning and Matching Multi-View Descriptors for Registration of Point Clouds
Lei Zhou 0011, Siyu Zhu 0001, Zixin Luo, Tianwei Shen, Mingmin Zhen, Tian Fang, Long Quan
ECCV (15)8
2017 Relative Camera Refinement for Accurate Dense Reconstruction
abstract
Multi-view stereo (MVS) depends on the pre-determined camera geometry, often from structure from motion (SfM) or simultaneous localization and mapping (SLAM). However, cameras may not be locally optimal for dense stereo matching, especially when it comes from the large scale SfM or the SLAM with multiple sensor fusion. In this paper, we propose a local camera refinement approach for accurate dense reconstruction. Firstly, we refines the relative geometry of independent camera pair using a tailored bundle adjustment. The refinement is also extended to a multi-view version for general MVS reconstructions. Then, the non-rigid dense alignment is formulated as an inverse-distortion problem to transfer point clouds from each local coordinate system to a global coordinate system. The proposed framework has been intensively validated in both SfM and SLAM based dense reconstructions. Results on different datasets show that our method can significantly improve the dense reconstruction quality.
Yao Yao 0008, Shiwei Li 0001, Siyu Zhu 0001, Hanyu Deng, Tian Fang, Long Quan
3DV6
2017 Distributed Very Large Scale Bundle Adjustment by Global Camera Consensus
abstract
The increasing scale of Structure-from-Motion is fundamentally limited by the conventional optimization framework for the all-in-one global bundle adjustment. In this paper, we propose a distributed approach to coping with this global bundle adjustment for very large scale Structure-from-Motion computation. First, we derive the distributed formulation from the classical optimization algorithm ADMM, Alternating Direction Method of Multipliers, based on the global camera consensus. Then, we analyze the conditions under which the convergence of this distributed optimization would be guaranteed. In particular, we adopt over-relaxation and self-adaption schemes to improve the convergence rate. After that, we propose to split the large scale camera-point visibility graph in order to reduce the communication overheads of the distributed computing. The experiments on both public large scale SfM data-sets and our very large scale aerial photo sets demonstrate that the proposed distributed method clearly outperforms the state-of-the-art method in efficiency and accuracy.
Siyu Zhu 0001, Tian Fang, Long Quan
ICCV4
2017 Progressive Large Scale-Invariant Image Matching in Scale Space
abstract
The power of modern image matching approaches is still fundamentally limited by the abrupt scale changes in images. In this paper, we propose a scale-invariant image matching approach to tackling the very large scale variation of views. Drawing inspiration from the scale space theory, we start with encoding the image's scale space into a compact multi-scale representation. Then, rather than trying to find the exact feature matches all in one step, we propose a progressive two-stage approach. First, we determine the related scale levels in scale space, enclosing the inlier feature correspondences, based on an optimal and exhaustive matching in a limited scale space. Second, we produce both the image similarity measurement and feature correspondences simultaneously after restricting matching between the related scale levels in a robust way. The matching performance has been intensively evaluated on vision tasks including image retrieval, feature matching and Structurefrom- Motion (SfM). The successful integration of the challenging fusion of high aerial and low ground-level views with significant scale differences manifests the superiority of the proposed approach.
Lei Zhou 0011, Siyu Zhu 0001, Tianwei Shen, Jinglu Wang, Tian Fang, Long Quan
ICCV6
2017 A robust three-stage approach to large-scale urban scene recognition
Jinglu Wang, Yonghua Lu, Long Quan
Sci. China Inf. Sci.4
2017 Special Issue on Large-Scale 3D Modeling of Urban Indoor or Outdoor Scenes from Images and Range Scans
Ioannis Stamos, Marc Pollefeys, Long Quan, Philippos Mordohai, Yasutaka Furukawa
Comput. Vis. Image Underst.3
2017 A novel data clustering algorithm based on modified gravitational search algorithm
XiaoHong Han, Long Quan, Xiaoyan Xiong, Matt Almeter, Yuan Lan
Eng. Appl. Artif. Intell.2
2017 Fast Descriptors and Correspondence Propagation for Robust Global Point Cloud Registration
abstract
In this paper, we present a robust global approach for point cloud registration from uniformly sampled points. Based on eigenvalues and normals computed from multiple scales, we design fast descriptors to extract local structures of these points. The eigenvalue-based descriptor is effective at finding seed matches with low precision using nearest neighbor search. Generally, recovering the transformation from matches with low precision is rather challenging. Therefore, we introduce a mechanism named correspondence propagation to aggregate each seed match into a set of numerous matches. With these sets of matches, multiple transformations between point clouds are computed. A quality function formulated from distance errors is used to identify the best transformation and fulfill a coarse alignment of the point clouds. Finally, we refine the alignment result with the trimmed iterative closest point algorithm. The proposed approach can be applied to register point clouds with significant or limited overlaps and small or large transformations. More encouragingly, it is rather efficient and very robust to noise. A comparison to traditional descriptor-based methods and other global algorithms demonstrates the fine performance of the proposed approach. We also show its promising application in large-scale reconstruction with the scans of two real scenes. In addition, the proposed approach can be used to register low-resolution point clouds captured by Kinect as well.
Huan Lei, Guang Jiang, Long Quan
IEEE Trans. Image Process.3
2016 Color Correction for Image-Based Modeling in the Large
Tianwei Shen, Jinglu Wang, Tian Fang, Siyu Zhu 0001, Long Quan
ACCV (4)5
2016 Efficient Multi-view Surface Refinement with Adaptive Resolution Control
Shiwei Li 0001, Sing Yu Siu, Tian Fang, Long Quan
ECCV (1)4
2016 Graph-Based Consistent Matching for Structure-from-Motion
Tianwei Shen, Siyu Zhu 0001, Tian Fang, Long Quan
ECCV (3)5
2016 Object localization using positive features
Huan Lei, Guang Jiang, Ruiyan Wang, Long Quan
Neurocomputing4
2016 Image-Based Building Regularization Using Structural Linear Features
abstract
Reconstructed building models using stereo-based methods inevitably suffer from noise, leading to the lack of regularity which is characterized by straightness of structural linear features and smoothness of homogeneous regions. We leverage the structural linear features embedded in the mesh to construct a novel surface scaffold structure for model regularization. The regularization comprises two iterative stages: (1) the linear features are semi-automatically proposed from images by exploiting photometric and geometric clues jointly; (2) the scaffold topology represented by spatial relations among the linear features is optimized according to data fidelity and topological rules, then the mesh is refined by adjusting itself to the consolidated scaffold. Our method has two advantages. First, the proposed scaffold representation is able to concisely describe semantic building structures. Second, the scaffold structure is embedded in the mesh, which can preserve the mesh connectivity and avoid stitching or intersecting surfaces in challenging cases. We demonstrate that our method can enhance structural characteristics and suppress irregularities in the building models robustly in some challenging datasets. Moreover, the regularization can significantly improve the results of general applications such as simplification and non-photorealistic rendering.
Jinglu Wang, Tian Fang, Qingkun Su, Siyu Zhu 0001, Shengnan Cai, Chiew-Lan Tai, Long Quan
IEEE Trans. Vis. Comput. Graph.8
2015 Dual-Feature Warping-Based Motion Model Estimation
abstract
To break down the geometry assumptions of traditional motion models (e.g., homography, affine), warping-based motion model recently becomes very popular and is adopted in many latest applications (e.g., image stitching, video stabilization). With high degrees of freedom, the accuracy of model heavily relies on data-terms (keypoint correspondences). In some low-texture environments (e.g., indoor) where keypoint feature is insufficient or unreliable, the warping model is often erroneously estimated. In this paper we propose a simple and effective approach by considering both keypoint and line segment correspondences as data-term. Line segment is a prominent feature in artificial environments and it can supply sufficient geometrical and structural information of scenes, which not only helps guild to a correct warp in low-texture condition, but also prevents the undesired distortion induced by warping. The combination aims to complement each other and benefit for a wider range of scenes. Our method is general and can be ported to many existing applications. Experiments demonstrate that using dual-feature yields more robust and accurate result especially for those low-texture images.
Shiwei Li 0001, Lu Yuan 0001, Jian Sun 0001, Long Quan
ICCV4
2015 Higher-Order CRF Structural Segmentation of 3D Reconstructed Surfaces
abstract
In this paper, we propose a structural segmentation algorithm to partition multi-view stereo reconstructed surfaces of large-scale urban environments into structural segments. Each segment corresponds to a structural component describable by a surface primitive of up to the second order. This segmentation is for use in subsequent urban object modeling, vectorization, and recognition. To overcome the high geometrical and topological noise levels in the 3D reconstructed urban surfaces, we formulate the structural segmentation as a higher-order Conditional Random Field (CRF) labeling problem. It not only incorporates classical lower-order 2D and 3D local cues, but also encodes contextual geometric regularities to disambiguate the noisy local cues. A general higher-order CRF is difficult to solve. We develop a bottom-up progressive approach through a patch-based surface representation, which iteratively evolves from the initial mesh triangles to the final segmentation. Each iteration alternates between performing a prior discovery step, which finds the contextual regularities of the patch-based representation, and an inference step that leverages the regularities as higher-order priors to construct a more stable and regular segmentation. The efficiency and robustness of the proposed method is extensively demonstrated on real reconstruction models, yielding significantly better performance than classical mesh segmentation methods.
Jinglu Wang, Tian Fang, Chiew-Lan Tai, Long Quan
ICCV5
2015 Joint Camera Clustering and Surface Segmentation for Large-Scale Multi-view Stereo
abstract
In this paper, we propose an optimal decomposition approach to large-scale multi-view stereo from an initial sparse reconstruction. The success of the approach depends on the introduction of surface-segmentation-based camera clustering rather than sparse-point-based camera clustering, which suffers from the problems of non-uniform reconstruction coverage ratio and high redundancy. In details, we introduce three criteria for camera clustering and surface segmentation for reconstruction, and then we formulate these criteria into an energy minimization problem under constraints. To solve this problem, we propose a joint optimization in a hierarchical framework to obtain the final surface segments and corresponding optimal camera clusters. On each level of the hierarchical framework, the camera clustering problem is formulated as a parameter estimation problem of a probability model solved by a General Expectation-Maximization algorithm and the surface segmentation problem is formulated as a Markov Random Field model based on the probability estimated by the previous camera clustering process. The experiments on several Internet datasets and aerial photo datasets demonstrate that the proposed approach method generates more uniform and complete dense reconstruction with less redundancy, resulting in more efficient multi-view stereo algorithm.
Shiwei Li 0001, Tian Fang, Siyu Zhu 0001, Long Quan
ICCV5
2015 A modified gravitational search algorithm based on sequential quadratic programming and chaotic map for ELD optimization
XiaoHong Han, Long Quan, Xiaoyan Xiong
Knowl. Inf. Syst.2
2014 Multi-scale Tetrahedral Fusion of a Similarity Reconstruction and Noisy Positional Measurements
Tian Fang, Siyu Zhu 0001, Long Quan
ACCV (2)4
2014 Multi-view Geometry Compression
Siyu Zhu 0001, Tian Fang, Long Quan
ACCV (2)4
2014 How Fashion Talks: Clothing-Region-Based Gender Recognition
Shengnan Cai, Jingdong Wang 0001, Long Quan
CIARP3
2014 Local Readjustment for High-Resolution 3D Reconstruction
abstract
Global bundle adjustment usually converges to a non-zero residual and produces sub-optimal camera poses for local areas, which leads to loss of details for high- resolution reconstruction. Instead of trying harder to optimize everything globally, we argue that we should live with the non-zero residual and adapt the camera poses to local areas. To this end, we propose a segment-based approach to readjust the camera poses locally and improve the reconstruction for fine geometry details. The key idea is to partition the globally optimized structure from motion points into well-conditioned segments for re-optimization, reconstruct their geometry individually, and fuse everything back into a consistent global model. This significantly reduces severe propagated errors and estimation biases caused by the initial global adjustment. The results on several datasets demonstrate that this approach can significantly improve the reconstruction accuracy, while maintaining the consistency of the 3D structure between segments.
Siyu Zhu 0001, Tian Fang, Jianxiong Xiao, Long Quan
CVPR4
2014 Low-rank SIFT: An affine invariant feature for place recognition
abstract
In this paper, we study the problem of recognizing man-made objects and present a novel affine-invariant feature, Low-rank SIFT, which exploits the regular appearance property in man-made objects. The proposed feature achieves full affine invariance without needing to simulate over affine parameter space. We rectify local patches by converting them to their low-rank forms to achieve skew invariance, and perform the way similar to conventional SIFT to resolve rotation, translation and scaling ambiguity. The main contributions lie in two-fold: our method seeks to leverage low-rank prior to estimate affine parameters for local patches directly and we propose a fast algorithm to compute such parameters by introducing the Low-rank Integral Map. Besides, we describe a pipeline of constructing a geotagged building database from the ground up. We demonstrate the effectiveness of our approach in the application to place recognition.
Harry Yang, Shengnan Cai, Jingdong Wang 0001, Long Quan
ICIP4
2014 Real-Time Object Tracking with Generalized Part-Based Appearance Model and Structure-Constrained Motion Model
abstract
In this paper, we propose a real-time object tracking approach. It utilizes generalized part-based appearance model and structure-constrained motion model as auxiliary. The appearance of the target object is modeled by the proposed generalized part-based appearance model, which combines the appearance of different parts of the target object, adaptively updated by an efficient structure learning scheme based on the online Passive-Aggressive algorithm. By integrating the confidence scores of multiple parts, mutual compensation is realized, significantly enhances the robustness of our method against the structure deformation and partial occlusion during the tracking. In addition, we enhance the performance of our tracker by using a motion model. It employs a structure-constrained rule, that is, the change on the structure of the target object between consecutive frames is small. Experiments on public video sequences verify the superior performance of our algorithm.
Honghui Zhang, Shengnan Cai, Long Quan
ICPR3
2014 Feature subset selection by gravitational search algorithm optimization
XiaoHong Han, Xiaoming Chang, Long Quan, Xiaoyan Xiong, Jingxia Li, Zhaoxia Zhang
Inf. Sci.3
2014 Regularized Tree Partitioning and Its Application to Unsupervised Image Segmentation
abstract
In this paper, we propose regularized tree partitioning approaches. We study normalized cut (NCut) and average cut (ACut) criteria over a tree, forming two approaches: 1) normalized tree partitioning (NTP) and 2) average tree partitioning (ATP). We give the properties that result in an efficient algorithm for NTP and ATP. In addition, we present the relations between the solutions of NTP and ATP over the maximum weight spanning tree of a graph and NCut and ACut over this graph. To demonstrate the effectiveness of the proposed approaches, we show its application to image segmentation over the Berkeley image segmentation data set and present qualitative and quantitative comparisons with state-of-the-art methods.
Jingdong Wang 0001, Huaizu Jiang, Yangqing Jia, Xian-Sheng Hua 0001, Changshui Zhang, Long Quan
IEEE Trans. Image Process.6
2014 Joint Segmentation of Images and Scanned Point Cloud in Large-Scale Street Scenes With Low-Annotation Cost
abstract
We propose a novel method for the parsing of images and scanned point cloud in large-scale street environment. The proposed method significantly reduces the intensive labeling cost in previous works by automatically generating training data from the input data. The automatic generation of training data begins with the initialization of training data with weak priors in the street environment, followed by a filtering scheme to remove mislabeled training samples. We formulate the filtering as a binary labeling optimization problem over a conditional random filed that we call object graph, simultaneously integrating spatial smoothness preference and label consistency between 2D and 3D. Toward the final parsing, with the automatically generated training data, a CRF-based parsing method that integrates the coordination of image appearance and 3D geometry is adopted to perform the parsing of large-scale street scenes. The proposed approach is evaluated on city-scale Google Street View data, with an encouraging parsing performance demonstrated.
Honghui Zhang, Jinglu Wang, Tian Fang, Long Quan
IEEE Trans. Image Process.4
2014 Data-Driven Synthetic Modeling of Trees
abstract
In this paper, we develop a data-driven technique to model trees from a single laser scan. A multi-layer representation of the tree structure is proposed to guide the modeling process. In this process, a marching cylinder algorithm is first developed to construct visible branches from the laser scan data. Three levels of crown feature points are then extracted from the scan data to synthesize three layers of non-visible branches. Based on the hierarchical particle flow technique, the branch synthesis method has the advantage of producing visually convincing tree models that are consistent with scan data. User intervention is extremely limited. The robustness of this technique has been validated on both conifer and broadleaf trees.
Xiaopeng Zhang 0001, Hongjun Li 0002, Mingrui Dai, Wei Ma 0008, Long Quan
IEEE Trans. Vis. Comput. Graph.5
2014 Optimizing neighborhood projection with relaxation factor for inextensible cloth simulation
Xiaowu Chen 0001, Qinping Zhao, Long Quan
Vis. Comput.4
2013 Learning CRFs for Image Parsing with Adaptive Subgradient Descent
abstract
We propose an adaptive sub gradient descent method to efficiently learn the parameters of CRF models for image parsing. To balance the learning efficiency and performance of the learned CRF models, the parameter learning is iteratively carried out by solving a convex optimization problem in each iteration, which integrates a proximal term to preserve the previously learned information and the large margin preference to distinguish bad labeling and the ground truth labeling. A solution of sub gradient descent updating form is derived for the convex optimization problem, with an adaptively determined updating step-size. Besides, to deal with partially labeled training data, we propose a new objective constraint modeling both the labeled and unlabeled parts in the partially labeled training data for the parameter learning of CRF models. The superior learning efficiency of the proposed method is verified by the experiment results on two public datasets. We also demonstrate the powerfulness of our method for handling partially labeled training data.
Honghui Zhang, Jingdong Wang 0001, Ping Tan 0002, Jinglu Wang, Long Quan
ICCV5
2013 Facing the classification of binary problems with a hybrid system based on quantum-inspired binary gravitational search algorithm and K-NN method
XiaoHong Han, Long Quan, Xiaoyan Xiong
Eng. Appl. Artif. Intell.2
2013 Image-Based Modeling of Unwrappable Façades
abstract
In this paper, we propose an unwrappable representation for image-based façade modeling from multiple registered images. An unwrappable façade is represented by the mutually orthogonal baseline and profile. We first reconstruct semidense 3D points from images, then the baseline and profile are extracted from the point cloud to construct the base shape and compose the textures of the building from the images. Through our unwrapping process, the reconstructed 3D points and composed textures are further mapped to an unwrapped space that is parameterized by the baseline and profile. In doing so, the unwrapped space becomes equivalent to the planar space in which planar façade modeling techniques can be used to reconstruct the details of the buildings. Finally, the augmented details can be wrapped back to the original 3D space to generate the final model. This newly introduced unwrappable representation extends the state-of-the-art modeling for planar façades to a more general class of façades. We demonstrate the power of the unwrappable representation with a few examples in which the façade is not planar.
Tian Fang, Zhexi Wang, Honghui Zhang, Long Quan
IEEE Trans. Vis. Comput. Graph.4
2012 Quasi-regular Facade Structure Extraction
Tian Han 0001, Chiew-Lan Tai, Long Quan
ACCV (4)4
2012 Parsing façade with rank-one approximation
abstract
The binary split grammar is powerful to parse façade in a broad range of types, whose structure is characterized by repetitive patterns with different layouts. We notice that, as far as two labels are concerned, BSG parsing is equivalent to approximating a façade by a matrix with multiple rank-one patterns. Then, we propose an efficient algorithm to decompose an arbitrary matrix into a rank-one matrix and a residual matrix, whose magnitude is small in the sense of l0-norm. Next, we develop a block-wise partition method to parse a more general façade. Our method leverages on the recent breakthroughs in convex optimization that can effectively decompose a matrix into a low-rank and sparse matrix pair. The rank-one block-wise parsing not only leads to the detection of repetitive patterns, but also gives an accurate façade segmentation. Experiments on intensive façade data sets have demonstrated that our method outperforms the state-of-the-art techniques and benchmarks both in robustness and efficiency.
Tian Han 0001, Long Quan, Chiew-Lan Tai
CVPR3
2012 Per-pixel translational symmetry detection, optimization, and segmentation
abstract
We present a novel method for translational symmetry detection, optimization, and symmetry object segmentation in façade images. Unlike most previous methods, our detection algorithm accumulates pixel-level correspondence in translation space. Thus it does not rely on feature point detection and handles patterns with low repetition counts. To improve the robustness with multiple interfering symmetries, we introduce an image-space global optimization, which resolves multiple per-pixel symmetry lattices. We then propose a learning-based method that generates refined segmentation of foreground symmetry objects of arbitrary shapes, with the aid of the per-pixel symmetry information. Our proposed method is accurate, robust and efficient as demonstrated by an extensive evaluation using a large façade image database.
Lei Yang 0006, Honghui Zhang, Long Quan
CVPR4
2012 Camera calibration using identical objects
Ruiyan Wang, Guang Jiang, Long Quan, Chengke Wu 0001
Mach. Vis. Appl.3
2011 Partial similarity based nonparametric scene parsing in certain environment
abstract
In this paper we propose a novel nonparametric image parsing method for the image parsing problem in certain environment. A novel and efficient nearest neighbor matching scheme, the ANN bilateral matching scheme, is proposed. Based on the proposed matching scheme, we first retrieve some partially similar images for each given test image from the training image database. The test image can be well explained by these retrieved images, with similar regions existing in the retrieved images for each region in the test image. Then, we match the test image to the retrieved training images with the ANN bilateral matching scheme, and parse the test image by integrating multiple cues in a markov random field. Experiment on three datasets shows our method achieved promising parsing accuracy and outperformed two state-of-the-art nonparametric image parsing methods.
Honghui Zhang, Tian Fang, Xiaowu Chen 0001, Qinping Zhao, Long Quan
CVPR5
2011 Translation symmetry detection in a fronto-parallel view
abstract
In this paper, we present a method of detecting translation symmetries from a fronto-parallel image. The proposed method automatically detects unknown multiple repetitive patterns of arbitrary shapes, which are characterized by translation symmetries on a plane. The central idea of our approach is to take advantage of the interesting properties of translation symmetries in both image space and the space of transformation group. We first detect feature points in input image as sampling points. Then for each sampling point, we search for the most probable corresponding lattice structures in the image and transform spaces using scale-space similarity maps. Finally, using a MRF formulation, we optimally partition the graph of all sampling points associated with the estimated lattices into subgraphs of sampling points and lattices belonging to the same symmetry pattern. Our method is robust because of the joint analysis in image and transform spaces, and the MRF optimization. We demonstrate the robustness and effectiveness of our method on a large variety of images.
Long Quan
CVPR2
2011 The Geometry of Reflectance Symmetries
abstract
Different materials reflect light in different ways, and this reflectance interacts with shape, lighting, and viewpoint to determine an object's image. Common materials exhibit diverse reflectance effects, and this is a significant source of difficulty for image analysis. One strategy for dealing with this diversity is to build computational tools that exploit reflectance symmetries, such as reciprocity and isotropy, that are exhibited by broad classes of materials. By building tools that exploit these symmetries, one can create vision systems that are more likely to succeed in real-world, non-Lambertian environments. In this paper, we develop a framework for representing and exploiting reflectance symmetries. We analyze the conditions for distinct surface points to have local view and lighting conditions that are equivalent under these symmetries, and we represent these conditions in terms of the geometric structure they induce on the Gaussian sphere and its abstraction, the projective plane. We also study the behavior of these structures under perturbations of surface shape and explore applications to both calibrated and uncalibrated photometric stereo.
Long Quan, Todd E. Zickler
IEEE Trans. Pattern Anal. Mach. Intell.2
2010 Rectilinear parsing of architecture in urban environment
abstract
We propose an approach that parses registered images captured at ground level into architectural units for large-scale city modeling. Each parsed unit has a regularized shape, which can be used for further modeling purposes. In our approach, we first parse the environment into buildings, the ground, and the sky using a joint 2D-3D segmentation method. Then, we partition buildings into individual façades. The partition problem is formulated as a dynamic programming optimization for a sequence of natural vertical separating lines. Each façade is regularized by a floor line and a roof line. The floor line is the intersection line of the vertical plane of buildings and the horizontal plane of the ground. The roof line links edge points of roof region. The parsed results provide a first geometric approximation to the city environment, and can be further analyzed if necessary. The approach is demonstrated and validated on several large-scale city datasets.
Tian Fang, Jianxiong Xiao, Honghui Zhang, Qinping Zhao, Long Quan
CVPR6
2010 Resampling Structure from Motion
Tian Fang, Long Quan
ECCV (2)2
2010 Supervised Label Transfer for Semantic Segmentation of Street Scenes
Honghui Zhang, Jianxiong Xiao, Long Quan
ECCV (5)3
2009 Multiple view semantic segmentation for street view images
abstract
We propose a simple but powerful multi-view semantic segmentation framework for images captured by a camera mounted on a car driving along streets. In our approach, a pair-wise Markov Random Field (MRF) is laid out across multiple views. Both 2D and 3D features are extracted at a super-pixel level to train classifiers for the unary data terms of MRF. For smoothness terms, our approach makes use of color differences in the same image to identify accurate segmentation boundaries, and dense pixel-to-pixel correspondences to enforce consistency across different views. To speed up training and to improve the recognition quality, our approach adaptively selects the most similar training data for each scene from the label pool. Furthermore, we also propose a powerful approach within the same framework to enable large-scale labeling in both the 3D space and 2D images. We demonstrate our approach on more than 10,000 images from Google Maps Street View.
Jianxiong Xiao, Long Quan
ICCV2
2009 Linear Neighborhood Propagation and Its Applications
abstract
In this paper, a novel graph-based transductive classification approach, called Linear Neighborhood Propagation, is proposed. The basic idea is to predict the label of a data point according to its neighbors in a linear way. This method can be cast into a second-order intrinsic Gaussian Markov random field framework. Its result corresponds to a solution to an approximate inhomogeneous biharmonic equation with Dirichlet boundary conditions. Different from existing approaches, our approach provides a novel graph structure construction method by introducing multiple-wise edges instead of pairwise edges, and presents an effective scheme to estimate the weights for such multiple-wise edges. To the best of our knowledge, these two contributions are novel for semi-supervised classification. The experimental results on image segmentation and transductive classification demonstrate the effectiveness and efficiency of the proposed approach.
Jingdong Wang 0001, Fei Wang 0001, Changshui Zhang, Helen C. Shen, Long Quan
IEEE Trans. Pattern Anal. Mach. Intell.5
2009 Automatic 3D face recognition from depth and intensity Gabor features
Chenghua Xu, Stan Z. Li, Tieniu Tan, Long Quan
Pattern Recognit.4
2009 Image-based street-side city modeling
abstract
We propose an automatic approach to generate street-side 3D photo-realistic models from images captured along the streets at ground level. We first develop a multi-view semantic segmentation method that recognizes and segments each image at pixel level into semantically meaningful areas, each labeled with a specific object class, such as building, sky, ground, vegetation and car. A partition scheme is then introduced to separate buildings into independent blocks using the major line structures of the scene. Finally, for each block, we propose an inverse patch-based orthographic composition and structure analysis method for façade modeling that efficiently regularizes the noisy and missing reconstructed 3D data. Our system has the distinct advantage of producing visually compelling results by imposing strong priors of building regularity. We demonstrate the fully automatic system on a typical city example to validate our methodology.
Jianxiong Xiao, Tian Fang, Maxime Lhuillier, Long Quan
ACM Trans. Graph.5
2008 Robust dual motion deblurring
abstract
This paper presents a robust algorithm to deblur two consecutively captured blurred photos from camera shaking. Previous dual motion deblurring algorithms succeeded in small and simple motion blur and are very sensitive to noise. We develop a robust feedback algorithm to perform iteratively kernel estimation and image deblurring. In kernel estimation, the stability and capability of the algorithm is greatly improved by incorporating a robust cost function and a set of kernel priors. The robust cost function serves to reject outliers and noise, while kernel priors, including sparseness and continuity, remove ambiguity and maintain kernel shape. In deblurring, we propose a novel and robust approach which takes two blurred images as input to infer the clear image. The deblurred image is then used as feedback to refine kernel estimation. Our method can successfully estimate large and complex motion blurs which cannot be handled by previous dual or single image motion deblurring algorithms. The results are shown to be significantly better than those of previous approaches.
Jia Chen 0026, Lu Yuan 0001, Chi-Keung Tang, Long Quan
CVPR4
2008 Normalized tree partitioning for image segmentation
abstract
In this paper, we propose a novel graph based clustering approach with satisfactory clustering performance and low computational cost. It consists of two main steps: tree fitting and partitioning. We first introduce a probabilistic method to fit a tree to a data graph under the sense of minimum entropy. Then, we propose a novel tree partitioning method under a normalized cut criterion, called Normalized Tree Partitioning (NTP), in which a fast combinatorial algorithm is designed for exact bipartitioning. Moreover, we extend it to k-way tree partitioning by proposing an efficient best-first recursive bipartitioning scheme. Compared with spectral clustering, NTP produces the exact global optimal bipartition, introduces fewer approximations for k-way partitioning and can intrinsically produce superior performance. Compared with bottom-up aggregation methods, NTP adopts a global criterion and hence performs better. Last, experimental results on image segmentation demonstrate that our approach is more powerful compared with existing graph-based approaches.
Jingdong Wang 0001, Yangqing Jia, Xian-Sheng Hua 0001, Changshui Zhang, Long Quan
CVPR5
2008 Learning Two-View Stereo Matching
Jianxiong Xiao, Jingni Chen, Dit-Yan Yeung, Long Quan
ECCV (3)4
2008 Structuring Visual Words in 3D for Arbitrary-View Object Localization
Jianxiong Xiao, Jingni Chen, Dit-Yan Yeung, Long Quan
ECCV (3)4
2008 Formal Use of Design Patterns and Refactoring
Long Quan, Zongyan Qiu, Zhiming Liu 0001
ISoLA1
2008 Subpixel Photometric Stereo
abstract
Conventional photometric stereo recovers one normal direction per pixel of the input image. This fundamentally limits the scale of recovered geometry to the resolution of the input image, and cannot model surfaces with subpixel geometric structures. In this paper, we propose a method to recover subpixel surface geometry by studying the relationship between the subpixel geometry and the reflectance properties of a surface. We first describe a generalized physically-based reflectance model that relates the distribution of surface normals inside each pixel area to its reflectance function. The distribution of surface normals can be computed from the reflectance functions recorded in photometric stereo images. A convexity measure of subpixel geometry structure is also recovered at each pixel, through an analysis of the shadowing attenuation. Then, we use the recovered distribution of surface normals and the surface convexity to infer subpixel geometric structures on a surface of homogeneous material by spatially arranging the normals among pixels at a higher resolution than that of the input image. Finally, we optimize the arrangement of normals using a combination of belief propagation and MCMC based on a minimum description length criterion on 3D textons over the surface. The experiments demonstrate the validity of our approach and show superior geometric resolution for the recovered surfaces.
Ping Tan 0002, Stephen Lin 0001, Long Quan
IEEE Trans. Pattern Anal. Mach. Intell.3
2008 Single image tree modeling
abstract
In this paper, we introduce a simple sketching method to generate a realistic 3D tree model from a single image. The user draws at least two strokes in the tree image: the first crown stroke around the tree crown to mark up the leaf region, the second branch stroke from the tree root to mark up the main trunk, and possibly few other branch strokes for refinement. The method automatically generates a 3D tree model including branches and leaves. Branches are synthesized by a growth engine from a small library of elementary subtrees that are pre-defined or built on the fly from the recovered visible branches. The visible branches are automatically traced from the drawn branch strokes according to image statistics on the strokes. Leaves are generated from the region bounded by the first crown stroke to complete the tree. We demonstrate our method on a variety of examples.
Tian Fang, Jianxiong Xiao, Long Quan
ACM Trans. Graph.5
2008 Image-based façade modeling
abstract
We propose in this paper a semi-automatic image-based approach to façade modeling that uses images captured along streets and relies on structure from motion to recover camera positions and point clouds automatically as the initial stage for modeling. We start by considering a building façade as a flat rectangular plane or a developable surface with an associated texture image composited from the multiple visible images. A façade is then decomposed and structured into a Directed Acyclic Graph of rectilinear elementary patches. The decomposition is carried out top-down by a recursive subdivision, and followed by a bottom-up merging with the detection of the architectural bilateral symmetry and repetitive patterns. Each subdivided patch of the flat façade is augmented with a depth optimized using the 3D points cloud. Our system also allows for an easy user feedback in the 2D image space for the proposed decomposition and augmentation. Finally, our approach is demonstrated on a large number of façades from a variety of street-side images.
Jianxiong Xiao, Tian Fang, Eyal Ofek, Long Quan
ACM Trans. Graph.6
2008 Progressive inter-scale and intra-scale non-blind image deconvolution
abstract
Ringing is the most disturbing artifact in the image deconvolution. In this paper, we present a progressive inter-scale and intra-scale non-blind image deconvolution approach that significantly reduces ringing. Our approach is built on a novel edge-preserving deconvolution algorithm called bilateral Richardson-Lucy (BRL) which uses a large spatial support to handle large blur. We progressively recover the image from a coarse scale to a fine scale (inter-scale), and progressively restore image details within every scale (intra-scale). To perform the inter-scale deconvolution, we propose a joint bilateral Richardson-Lucy (JBRL) algorithm so that the recovered image in one scale can guide the deconvolution in the next scale. In each scale, we propose an iterative residual deconvolution to progressively recover image details. The experimental results show that our progressive deconvolution can produce images with very little ringing for large blur kernels.
Lu Yuan 0001, Jian Sun 0001, Long Quan, Harry Shum
ACM Trans. Graph.3
2008 Filtering and Rendering of Resolution-Dependent Reflectance Models
abstract
The apparent reflectance of a surface depends upon the resolution at which it is imaged. Conventional reflectance models represent reflection at a single predetermined resolution; however, a low-resolution pixel that views a greater surface area often exhibits a reflectance more complicated than a high-resolution pixel with a smaller area. To address resolution dependency in reflectance, we utilize a generalized reflectance model based on a mixture of multiple conventional models, and present a framework for efficiently determining the reflectance mixture model of each pixel with respect to resolution. Mixture model parameters are precomputed at multiple resolutions and stored in mipmaps. Unlike color textures, these reflectance parameters cannot be accurately filtered by trilinear interpolation, so we present a technique for nonlinear mipmap filtering that minimizes aliasing in rendered results. This framework can be applied with various parametric reflectance models in graphics hardware for real-time processing. With this technique for filtering and rendering with mipmaps of reflectance mixture models, our system can rapidly render the resolution-dependent reflectance effects that are customarily disregarded in conventional rendering methods. At the end of this paper, we also describe how shadowing and masking effects can be incorporated into this framework to increase the realism of rendering.
Ping Tan 0002, Stephen Lin 0001, Long Quan, Baining Guo, Harry Shum
IEEE Trans. Vis. Comput. Graph.3
2007 Isotropy, Reciprocity and the Generalized Bas-Relief Ambiguity
abstract
A set of images of a Lambertian surface under varying lighting directions defines its shape up to a three-parameter generalized bas-relief (GBR) ambiguity. In this paper, we examine this ambiguity in the context of surfaces having an additive non-Lambertian reflectance component, and we show that the GBR ambiguity is resolved by any non-Lambertian reflectance function that is isotropic and spatially invariant. The key observation is that each point on a curved surface under directional illumination is a member of a family of points that are in isotropic or reciprocal configurations. We show that the GBR can be resolved in closed form by identifying members of these families in two or more images. Based on this idea, we present an algorithm for recovering full Euclidean geometry from a set of uncalibrated photometric stereo images, and we evaluate it empirically on a number of examples.
Satya P. Mallick, Long Quan, David J. Kriegman, Todd E. Zickler
CVPR3
2007 Joint Affinity Propagation for Multiple View Segmentation
abstract
A joint segmentation is a simultaneous segmentation of registered 2D images and 3D points reconstructed from the multiple view images. It is fundamental in structuring the data for subsequent modeling applications. In this paper, we treat this joint segmentation as a weighted graph labeling problem. First, we construct a 3D graph for the joint 3D and 2D points using a joint similarity measure. Then, we propose a hierarchical sparse affinity propagation algorithm to automatically and jointly segment 2D images and group 3D points. Third, a semi-supervised affinity propagation algorithm is proposed to refine the automatic results with the user assistance. Finally, intensive experiments demonstrate the effectiveness of the proposed approaches.
Jianxiong Xiao, Jingdong Wang 0001, Ping Tan 0002, Long Quan
ICCV4
2007 Blurred/Non-Blurred Image Alignment using Sparseness Prior
abstract
Aligning a pair of blurred and non-blurred images is a prerequisite for many image and video restoration and graphics applications. The traditional alignment methods such as direct and feature-based approaches cannot be used due to the presence of motion blur in one image of the pair. In this paper, we present an effective and accurate alignment approach for a blurred/non-blurred image pair. We exploit a statistical characteristic of the real blur kernel - the marginal distribution of kernel value is sparse. Using this sparseness prior, we can search the best alignment which produces the sparsest blur kernel. The search is carried out in scale space with a coarse-to-fine strategy for efficiency. Finally, we demonstrate the effectiveness of our algorithm for image deblurring, video restoration, and image matting.
Lu Yuan 0001, Jian Sun 0001, Long Quan, Harry Shum
ICCV3
2007 Image-Based Modeling by Joint Segmentation
Long Quan, Jingdong Wang 0001, Ping Tan 0002, Lu Yuan 0001
Int. J. Comput. Vis.1
2007 Accurate and Scalable Surface Representation and Reconstruction from Images
abstract
We introduce a new surface representation method, called patchwork, to extend three-dimensional surface reconstruction capabilities from multiple images. A patchwork is the combination of several patches that are built one by one. This design potentially allows for the reconstruction of an object with arbitrarily large dimensions while preserving a fine level of detail. We formally demonstrate that this strategy leads to a spatial complexity independent of the dimensions of the reconstructed object and to a time complexity that is linear with respect to the object area. The former property ensures that we never run out of storage and the latter means that reconstructing an object can be done in a reasonable amount of time. In addition, we show that the patchwork representation handles equivalently open and closed surfaces, whereas most of the existing approaches are limited to a specific scenario, an open or closed surface, but not both. The patchwork concept is orthogonal to the method chosen for surface optimization. Most of the existing optimization techniques can be cast into this framework. To illustrate the possibilities offered by this approach, we propose two applications that demonstrate how our method dramatically extends a recent accurate graph technique based on minimal cuts. We first revisit the popular carving techniques. This results in a well-posed reconstruction problem that still enjoys the tractability of voxel space. We also show how we can advantageously combine several image-driven criteria to achieve a finely detailed geometry by surface propagation. These two examples demonstrate the versatility and flexibility of patchwork reconstruction. They underscore other properties inherited from patchwork representation: Although some min-cut methods have difficulty in handling complex shapes (e.g., with complex topologies), they can naturally manipulate any geometry through the patchwork representation while preserving their intrinsic qualities. The above properties of patchwork representation and reconstruction are demonstrated with real image sequences.
Sylvain Paris, Long Quan, François X. Sillion
IEEE Trans. Pattern Anal. Mach. Intell.3
2007 Image-based tree modeling
abstract
In this paper, we propose an approach for generating 3D models of natural-looking trees from images that has the additional benefit of requiring little user intervention. While our approach is primarily image-based, we do not model each leaf directly from images due to the large leaf count, small image footprint, and widespread occlusions. Instead, we populate the tree with leaf replicas from segmented source images to reconstruct the overall tree shape. In addition, we use the shape patterns of visible branches to predict those of obscured branches. We demonstrate our approach on a variety of trees.
Ping Tan 0002, Jingdong Wang 0001, Sing Bing Kang, Long Quan
ACM Trans. Graph.5
2007 Image deblurring with blurred/noisy image pairs
abstract
Taking satisfactory photos under dim lighting conditions using a hand-held camera is challenging. If the camera is set to a long exposure time, the image is blurred due to camera shake. On the other hand, the image is dark and noisy if it is taken with a short exposure time but with a high camera gain. By combining information extracted from both blurred and noisy images, however, we show in this paper how to produce a high quality image that cannot be obtained by simply denoising the noisy image, or deblurring the blurred image alone. Our approach is image deblurring with the help of the noisy image. First, both images are used to estimate an accurate blur kernel, which otherwise is difficult to obtain from a single blurred image. Second, and again using both images, a residual deconvolution is proposed to significantly reduce ringing artifacts inherent to image deconvolution. Third, the remaining ringing artifacts in smooth image regions are further suppressed by a gain-controlled deconvolution process. We demonstrate the effectiveness of our approach using a number of indoor and outdoor images taken by off-the-shelf hand-held cameras in poor lighting environments.
Lu Yuan 0001, Jian Sun 0001, Long Quan, Harry Shum
ACM Trans. Graph.3
2006 Separation of Highlight Reflections on Textured Surfaces
abstract
We present a method for separating highlight reflections on textured surfaces. In contrast to previous techniques that use diffuse color information from outside the highlight area to constrain the solution, the proposed method further capitalizes on the spatial distributions of colors to resolve ambiguities in separation that often arise in real images. For highlight pixels in which a clear-cut separation cannot be determined from color space analysis, we evaluate possible separation solutions based on their consistency with diffuse texture characteristics outside the highlight. With consideration of color distributions in both the color space and the image space, appreciably enhanced separation performance can be attained in challenging cases.
Ping Tan 0002, Long Quan, Stephen Lin 0001
CVPR (2)2
2006 Picture Collage
abstract
In this paper, we address a novel problem of automatically creating a picture collage from a group of images. Picture collage is a kind of visual image summary - to arrange all input images on a given canvas, allowing overlay, to maximize visible visual information. We formulate the picture collage creation problem in a Bayesian framework. The salient regions of each image are firstly extracted and represented as a set of weighted rectangles. Then, the image arrangement is formulated as a Maximum a Posterior (MAP) problem such that the output picture collage shows as many visible salient regions (without being overlaid by others) from all images as possible. Moreover, a very efficientMarkov chain Monte Carlo (MCMC) method is designed for the optimization. Applications to desktop image browsing and image search result summarization demonstrate the effectiveness of our approach.
Jingdong Wang 0001, Long Quan, Jian Sun 0001, Xiaoou Tang, Harry Shum
CVPR (1)2
2006 Resolution-Enhanced Photometric Stereo
Ping Tan 0002, Stephen Lin 0001, Long Quan
ECCV (3)3
2006 A Surface Reconstruction Method Using Global Graph Cut Optimization
Sylvain Paris, François X. Sillion, Long Quan
Int. J. Comput. Vis.3
2006 Combining local features for robust nose location in 3D facial data
Chenghua Xu, Tieniu Tan, Yunhong Wang 0001, Long Quan
Pattern Recognit. Lett.4
2006 Image-based plant modeling
abstract
In this paper, we propose a semi-automatic technique for modeling plants directly from images. Our image-based approach has the distinct advantage that the resulting model inherits the realistic shape and complexity of a real plant. We designed our modeling system to be interactive, automating the process of shape recovery while relying on the user to provide simple hints on segmentation. Segmentation is performed in both image and 3D spaces, allowing the user to easily visualize its effect immediately. Using the segmented image and 3D data, the geometry of each leaf is then automatically recovered from the multiple views by fitting a deformable leaf model. Our system also allows the user to easily reconstruct branches in a similar manner. We show realistic reconstructions of a variety of plants, and demonstrate examples of plant editing.
Long Quan, Ping Tan 0002, Lu Yuan 0001, Jingdong Wang 0001, Sing Bing Kang
ACM Trans. Graph.1
2005 Asymmetrical Occlusion Handling Using Graph Cut for Multi-View Stereo
abstract
Occlusion is usually modelled in two images symmetrically in previous stereo algorithms which cannot work for multi-view stereo efficiently. In this paper, we present a novel formulation that handles occlusion using only one depth map in an asymmetrical way. Consequently, multi-view information is efficiently accumulated to achieve high accuracy. The resulting energy function is complex and approximate graph cut based solutions are proposed. Our approach complements the theory and extends the applicability of using graph cut in stereo. The experiments demonstrate that the approach is comparable with the state of the art and potentially more efficient for multi-view stereo.
Long Quan
CVPR (2)2
2005 Interactive Shape from Shading
abstract
Shape from shading (SfS) has always been difficult for real applications due to its intrinsic ill-posedness. In this paper, we propose an interactive SfS method which efficiently uses human knowledge in order to resolve ambiguity. We propose a global solution of continuous surfaces with a few constraints of surface normals that are interactively imposed to regularize the problem. A surface is divided into local patches, and each local solution is estimated with a fast marching SfS. It is shown that the boundaries of local solutions constitute a weighted Voronoi diagram, which allows for the formation of a global solution from the local ones. Finally, we optimize this global estimation by minimizing an energy functional based on shading and smoothness priors. Reconstruction results from both synthetic and real images demonstrate the usability of the new approach for various modeling applications.
Yasuyuki Matsushita, Long Quan, Harry Shum
CVPR (1)3
2005 Detection of Concentric Circles for Camera Calibration
abstract
The geometry of plane-based calibration methods is well understood, but some user interaction is often needed in practice for feature detection. This paper presents a fully automatic calibration system that uses patterns of pairs of concentric circles. The key observation is to introduce a geometric method that constructs a sequence of points strictly convergent to the image of the circle center from an arbitrary point. The method automatically detects the points of the pattern features by the construction method, and identify them by invariants. It then takes advantage of homological constraints to consistently and optimally estimate the features in the image. The experiments demonstrate the robustness and the accuracy of the new method.
Guang Jiang, Long Quan
ICCV2
2005 Progressive Surface Reconstruction from Images Using a Local Prior
abstract
This paper introduces a new method for surface reconstruction from multiple calibrated images. The primary contribution of this work is the notion of local prior to combine the flexibility of the carving approach with the accuracy of graph-cut optimization. A progressive refinement scheme is used to recover the topology and reason the visibility of the object. Within each voxel, a detailed surface patch is optimally reconstructed using a graph-cut method. The advantage of this technique is its ability to handle complex shape similarly to level sets while enjoying a higher precision. Compared to carving techniques, the addressed problem is well-posed, and the produced surface does not suffer from aliasing. In addition, our approach seamlessly handles complete and partial reconstructions: If the scene is only partially visible, the process naturally produces an open surface; otherwise, if the scene is fully visible, it creates a complete shape. These properties are demonstrated on real image sequences
Sylvain Paris, Long Quan, François X. Sillion
ICCV3
2005 Data-dependent kernels for high-dimensional data classification
abstract
For high-dimensional data classification problems such as face recognition, one of the most efficient classifiers is the nearest neighbor (NN) classifier. What mostly affects the NN classification performance is the feature extracted by some methods. And the kernel method is one of the efficient methods for extracting features. However, the selection of kernel parameters is still difficult. In this paper, we propose a so-called data dependent kernel (DDK) which is defined by generalizing the Gaussian kernel. Also an efficient and practical method is presented to calculate the DDK parameters. Moreover, one DDK based on subspaces is given to improve the recognition performance. Experiments show that the proposed DDK can achieve promising classification performance in face recognition and SPECT heart diagnosis.
Jingdong Wang 0001, James T. Kwok, Helen C. Shen, Long Quan
IJCNN4
2005 Multiresolution Reflectance Filtering
abstract
Physically-based reflectance models typically represent light scattering as a function of surface geometry at the pixel level. With changes in viewing resolution, the geometry imaged within a pixel can undergo significant variations that result in changing reflectance characteristics. To address these transformations, we present a multiresolution reflectance framework based on microfacet normal distributions within a pixel over different scales. Since these distributions must be efficiently determined with respect to resolution, they are recorded at multiple resolution levels in mipmaps. The main contribution of this work is a real-time mipmap filtering technique for these distribution-based parameters that not only provides smooth reflectance transitions in scale, but also minimizes aliasing. With this multiresolution reflectance technique, our system can rapidly and accurately incorporate fine reflectance detail that is customarily disregarded in multiresolution rendering methods.
Ping Tan 0002, Stephen Lin 0001, Long Quan, Baining Guo, Harry Shum
Rendering Techniques3
2005 Outward-Looking Circular Motion Analysis of Large Image Sequences
abstract
This paper presents a novel and simple method of analyzing the motion of a large image sequence captured by a calibrated outward-looking video camera moving on a circular trajectory for large-scale environment applications. Previous circular motion algorithms mainly focus on inward-looking turntable-like setups. They are not suitable for outward-looking motion where the conic trajectory of corresponding points degenerates to straight lines. The circular motion of a calibrated camera essentially has only one unknown rotation angle for each frame. The motion recovery for the entire sequence computes only one fundamental matrix of a pair of frames to extract the angular motion of the pair using Laguerre's formula and then propagates the computation of the unknown rotation angles to the other frames by tracking one point over at least three frames. Finally, a maximum-likelihood estimation is developed for the optimization of the whole sequence. Extensive experiments demonstrate the validity of the method and the feasibility of the application in image-based rendering.
Guang Jiang, Long Quan, Hung-Tat Tsui, Harry Shum
IEEE Trans. Pattern Anal. Mach. Intell.3
2005 A Quasi-Dense Approach to Surface Reconstruction from Uncalibrated Images
abstract
This paper proposes a quasi-dense approach to 3D surface model acquisition from uncalibrated images. First, correspondence information and geometry are computed based on new quasi-dense point features that are resampled subpixel points from a disparity map. The quasi-dense approach gives more robust and accurate geometry estimations than the standard sparse approach. The robustness is measured as the success rate of full automatic geometry estimation with all involved parameters fixed. The accuracy is measured by a fast gauge-free uncertainty estimation algorithm. The quasi-dense approach also works for more largely separated images than the sparse approach, therefore, it requires fewer images for modeling. More importantly, the quasidense approach delivers a high density of reconstructed 3D points on which a surface representation can be reconstructed. This fills the gap of insufficiency of the sparse approach for surface reconstruction, essential for modeling and visualization applications. Second, surface reconstruction methods from the given quasi-dense geometry are also developed. The algorithm optimizes new unified functionals integrating both 3D quasi-dense points and 2D image information, including silhouettes. Combining both 3D data and 2D images is more robust than the existing methods using only 2D information or only 3D data. An efficient bounded regularization method is proposed to implement the surface evolution by level-set methods. Its properties are discussed and proven for some cases. As a whole, a complete automatic and practical system of 3D modeling from raw images captured by hand-held cameras to surface representation is proposed. Extensive experiments demonstrate the superior performance of the quasi-dense approach with respect to the standard sparse approach in robustness, accuracy, and applicability.
Maxime Lhuillier, Long Quan
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 Modeling hair from multiple views
abstract
In this paper, we propose a novel image-based approach to model hair geometry from images taken at multiple viewpoints. Unlike previous hair modeling techniques that require intensive user interactions or rely on special capturing setup under controlled illumination conditions, we use a handheld camera to capture hair images under uncontrolled illumination conditions. Our multi-view approach is natural and flexible for capturing. It also provides inherent strong and accurate geometric constraints to recover hair models.In our approach, the hair fibers are synthesized from local image orientations. Each synthesized fiber segment is validated and optimally triangulated from all visible views. The hair volume and the visibility of synthesized fibers can also be reliably estimated from multiple views. Flexibility of acquisition, little user interaction, and high quality results of recovered complex hair models are the key advantages of our method.
Eyal Ofek, Long Quan, Harry Shum
ACM Trans. Graph.3
2004 Region-Based Progressive Stereo Matching
Long Quan
CVPR (1)2
2004 Surface Reconstruction by Propagating 3D Stereo Data in Multiple 2D Images
Sylvain Paris, Long Quan, Maxime Lhuillier
ECCV (1)3
2004 Adaptive Multi-Resolution Fitting and its Application to Realistic Head Modeling
abstract
The general approach for object modeling is to construct the surface from the high-quality range points obtained from laser scanners. In this paper, we face the noise point cloud obtained from image sequences by a common camera and develop a novel algorithm of adaptive multi-resolution fitting (AMRF) for object modeling. This algorithm combines the adaptive subdivision scheme with multi-resolution fitting so that the control model is subdivided locally and adaptively according to the local complexity of the point cloud and approximates the 3D data level by level. The proposed method can conquer the holes and outliers efficiently and create full compatibility between the complexity of the mesh model and the representation of the local details. We apply the proposed method to the complete head modeling with the real data, and the results seem very promising.
Chenghua Xu, Long Quan, Yunhong Wang 0001, Tieniu Tan, Maxime Lhuillier
GMP2
2004 Robust nose detection in 3d facial data using local characteristics
Chenghua Xu, Yunhong Wang 0001, Tieniu Tan, Long Quan
ICIP4
2004 Constrained planar motion analysis by decomposition
Long Quan, Le Lu 0001, Harry Shum
Image Vis. Comput.1
2004 Circular Motion Geometry Using Minimal Data
abstract
Circular motion or single axis motion is widely used in computer vision and graphics for 3D model acquisition. This paper describes a new and simple method for recovering the geometry of uncalibrated circular motion from a minimal set of only two points in four images. This problem has been previously solved using nonminimal data either by computing the fundamental matrix and trifocal tensor in three images or by fitting conics to tracked points in five or more images. It is first established that two sets of tracked points in different images under circular motion for two distinct space points are related by a homography. Then, we compute a plane homography from a minimal two points in four images. After that, we show that the unique pair of complex conjugate eigenvectors of this homography are the image of the circular points of the parallel planes of the circular motion. Subsequently, all other motion and structure parameters are computed from this homography in a straighforward manner. The experiments on real image sequences demonstrate the simplicity, accuracy, and robustness of the new method.
Guang Jiang, Long Quan, Hung-Tat Tsui
IEEE Trans. Pattern Anal. Mach. Intell.2
2003 Circular Motion Geometry by Minimal 2 Points in 4 Images
abstract
We describe a new and simple method of recovering the geometry of uncalibrated circular motion or single axis motion using a minimal data set of 2 points in 4 images. This problem has been solved using nonminimal data either by computing the fundamental matrix and trifocal tensor in 3 images, or by fitting conics to tracked points in 5 images. Our new method first computes a planar homography from a minimum of 2 points in 4 images. It is shown that two eigenvectors of this homography are the images of the circular points. Then, other fixed image entities and rotation angles can be straightforwardly computed. The crux of the method lies in relating this planar homography from two different points to a homology naturally induced by corresponding points on different conic loci from a circular motion. The experiments on real image sequences demonstrate the simplicity, accuracy and robustness of the new method.
Guang Jiang, Long Quan, Hung-Tat Tsui
ICCV2
2003 Surface Reconstruction by Integrating 3D and 2D Data of Multiple Views
abstract
Surface representation is needed for almost all modeling and visualization applications, but unfortunately, 3D data from a passive vision system are often insufficient for a traditional surface reconstruction technique that is designed for densely scanned 3D point data. In this paper, we develop a new method for surface reconstruction by combining both 3D data and 2D image information. The silhouette information extracted from 2D images can also be integrated as an option if it is available. The new method is a variational approach with a new functional integrating 3D stereo data with 2D image information. This gives a more robust approach than existing methods using only pure 2D information or 3D stereo data. We also propose a bounded regularization method to implement efficiently the surface evolution by level-set methods. The properties of the algorithms are discussed, proved for some cases, and empirically demonstrated through intensive experiments on real sequences.
Maxime Lhuillier, Long Quan
ICCV2
2003 Highlight Removal by Illumination-Constrained Inpainting
abstract
We present a single-image highlight removal method that incorporates illumination-based constraints into image inpainting. Unlike occluded image regions filled by traditional inpainting, highlight pixels contain some useful information for guiding the inpainting process. Constraints provided by observed pixel colors, highlight color analysis and illumination color uniformity are employed in our method to improve estimation of the underlying diffuse color. The inclusion of these illumination constraints allows for better recovery of shading and textures by inpainting. Experimental results are given to demonstrate the performance of our method.
Ping Tan 0002, Stephen Lin 0001, Long Quan, Harry Shum
ICCV3
2003 Lightweight Face Relighting
abstract
We present a method to relight human faces in real time, using consumer-grade graphics cards even with limited 3D capabilities. We show how to render faces using a combination of a simple, hardware-accelerated parametric model simulating skin shading and a detail texture map, and provide robust procedures to estimate all the necessary parameters for a given face. Our model strikes a balance between the difficulty of realistic face rendering (given the very specific reflectance properties of skin) and the goal of real-time rendering with limited hardware capabilities. This is accomplished by automatically generating an optimal set of parameters for a simple rendering model. We offer a discussion of the issues in face rendering to discern the pros and cons of various rendering models and to generalize our approach to most of the current hardware constraints. We provide results demonstrating the usability of our approach and the improvements we introduce both in the performance and in the visual quality of the resulting faces.
Sylvain Paris, François X. Sillion, Long Quan
PG3
2003 Geometry of Single Axis Motions Using Conic Fitting
abstract
Previous algorithms for recovering 3D geometry from an uncalibrated image sequence of a single axis motion of unknown rotation angles are mainly based on the computation of two-view fundamental matrices and three-view trifocal tensors. We propose three new methods that are based on fitting a conic locus to corresponding image points over multiple views. The main advantage is that determining only five parameters of a conic from one corresponding point over at least five views is simpler and more robust than determining a fundamental matrix from two views or a trifocal tensor from three views. It is shown that the geometry of single axis motion can be recovered either by computing one conic locus and one fundamental matrix or by computing at least two conic loci. A maximum likelihood solution based on this parametrization of the single axis motion is also described for optimal estimation using three or more loci. The experiments on real image sequences demonstrate the simplicity, accuracy, and robustness of the new methods.
Guang Jiang, Hung-Tat Tsui, Long Quan, Andrew Zisserman
IEEE Trans. Pattern Anal. Mach. Intell.3
2003 Image-based rendering by joint view triangulation
abstract
The creation of novel views using prestored images or image-based rendering has many potential applications, such as visual simulation, virtual reality, and telepresence, for which traditional computer graphics based on geometric modeling would be unsatisfactory particularly with very complex three-dimensional scenes. This paper presents a new image-based rendering system that tackles the two most difficult problems of image-based modeling: pixel matching and visibility handling. We first introduce the joint view triangulation (JVT), a novel representation for pairs of images that handles the visibility and occlusion problems created by the parallaxes between the images. The JVT is built from matched planar patches regularized by local smooth constraints encoded by plane homographies. Then, we introduce an incremental edge-constrained construction algorithm. Finally, we present a pseudo-painter's rendering algorithm for the JVT and demonstrate the performance of these methods experimentally.
Maxime Lhuillier, Long Quan
IEEE Trans. Circuits Syst. Video Technol.2
2002 Single Axis Geometry by Fitting Conics
Guang Jiang, Hung-Tat Tsui, Long Quan, Andrew Zisserman
ECCV (1)3
2002 Quasi-Dense Reconstruction from Image Sequence
Maxime Lhuillier, Long Quan
ECCV (2)2
2002 Match Propagation for Image-Based Modeling and Rendering
abstract
This paper presents a quasi-dense matching algorithm between images based on the match propagation principle. The algorithm starts from a set of sparse seed matches, then propagates to the neighboring pixels by the best-first strategy, and produces a quasi-dense disparity map. The quasi-dense matching aims at broad modeling and visualization applications which rely heavily on matching information. Our algorithm is robust to initial sparse match outliers due to the best-first strategy. It is efficient in time and space as it is only output sensitive. It handles half-occluded areas because of the simultaneous enforcement of newly introduced discrete 2D gradient disparity limit and the uniqueness constraint. The properties of the algorithm are discussed and empirically demonstrated. The quality of quasi-dense matching are validated through intensive real examples.
Maxime Lhuillier, Long Quan
IEEE Trans. Pattern Anal. Mach. Intell.2
2001 Recovering the Geometry of Single Axis Motions by Conic Fitting
abstract
In this paper, we propose a new approach for recovering 3D geometry from uncalibrated and unknown rotation angles of single axis motions. Unlike previous approaches, the computation of the trifocal tensor is not required. The new approach is based on fitting a conic to the corresponding points in different images of the same space point. It is then shown that the essential geometry of single axis motion is encoded by a conic. Since we are determining only 5 parameters from N views instead of 18 parameters from a subsequence of 3 views, our approach is much more simple and robust. The experiments on a real image sequence demonstrate the accuracy and robustness of the new algorithm.
Guang Jiang, Hung-Tat Tsui, Long Quan, Shang-Qian Liu
CVPR (1)3
2001 Relief Mosaics by Joint View Triangulation
abstract
Relief mosaics are collections of registered images that extend traditional mosaics by supporting motion parallax. A simple parallax interpolation algorithm based on computed correspondence information allows high quality blur-free and ghost-free mosaics to be created using images from moving hand-held cameras that would not be suitable for traditional mosaicing. The renderer can also display local parallax changes, giving a local but visually convincing illusion of depth. Moreover, relief mosaics can be used for approximate plenoptic modeling from hand-held cameras at lower spatial sampling rates than existing light-field methods. We present a fully automatic correspondence based construction system for relief mosaics, and show how they can be used in applications.
Maxime Lhuillier, Long Quan, Harry Shum, Hung-Tat Tsui
CVPR (1)2
2001 Concentric Mosaic(s)Planar Motion and 1D Cameras
abstract
General SFM methods give poor results for images captured by constrained motions such as planar motion of concentric mosaics (CM). In this paper we propose new SFM algorithms for both images captured by CM and composite mosaic images from CM. We first introduce ID affine camera model for completing 1D camera models. Then we show that a 2D image captured by CM can be decoupled into two 1D images: one 1D projective and one ID affine; a composite mosaic image can by rebinned into a calibrated ID panorama projective camera. Finally we describe subspace reconstruction methods and demonstrate both in theory and experiments the advantage of the decomposition method over the general SFM methods by incorporating the constrained motion into the earliest stage of motion analysis.
Long Quan, Le Lu 0001, Harry Shum, Maxime Lhuillier
ICCV1
2001 Minimal Projective Reconstruction Including Missing Data
abstract
The minimal data necessary for projective reconstruction from image points is well-known when each object point is visible in all images. We formulate and propose solutions to a family of reconstruction problems for multiple images from minimal data, where there are missing points in some of the images. The ability to handle the minimal cases with missing data is of great theoretical and practical importance. It is unavoidable to use them to bootstrap robust estimation such as RANSAC and LMS algorithms and optimal estimation such as bundle adjustment. First, we develop a framework to parameterize the multiple view geometry needed to handle the missing data cases. Then, we present a solution to the minimal case of eight points in three images, where one different point is missing in each of the three images. We prove that there are, in general, as many as 11 solutions for this minimal case. Furthermore all minimal cases with missing data for three and four images are catalogued. Finally, we demonstrate the method on both simulated and real images and show that the algorithms presented in the paper can be used for practical problems.
Fredrik Kahl, Anders Heyden, Long Quan
IEEE Trans. Pattern Anal. Mach. Intell.3
2001 Two-Way Ambiguity in 2D Projective Reconstruction from Three Uncalibrated 1D Images
abstract
We show that there is, in general, a two-way ambiguity for 2D projective reconstruction from three uncalibrated 1D views, independent of the number of point correspondences. The two distinct projective reconstructions are exactly related by a quadratic transformation with the three camera centers as fundamental points. Unique 2D reconstruction is possible only when the three camera centers are aligned. By Carlsson duality (1995), there is a dual two-way ambiguity for 2D projective reconstruction from six point correspondences, independent of the number of 1D views. The theoretical results are demonstrated on numerical examples.
Long Quan
IEEE Trans. Pattern Anal. Mach. Intell.1
2000 Edge-Constrained Joint View Triangulation for Image Interpolation
abstract
Image-based-interpolation creates smooth and photo-realistic views between two view points. The concept of joint view triangulation (JVT) has been proven to be an efficient multi-view representation to handle visibility issue. However, the existing JVT built only on a regular sampling grid, often produces undesirable artifacts for artificial objects. To tackle these problems, a new edge-constrained joint view triangulation is developed in this paper to integrate contour points and artificial rectilinear objects as triangulation constraints. Also a super-sampling technique is introduced to refine visible boundaries. The new algorithm is successfully demonstrated on many real image pairs.
Maxime Lhuillier, Long Quan
CVPR2
2000 Robust Dense Matching Using Local and Global Geometric Constraints
abstract
A new robust dense matching algorithm is introduced. The algorithm starts from matching the most textured points, then a match propagation algorithm is developed with the best first strategy to dense matching. Next, the matching map is regularised by using the local geometric constraints encoded by planar affine applications and by using the global geometric constraint encoded by the fundamental matrix. Two most distinctive features are a match propagation strategy developed by analogy to region growing and a successive regularisation by local and global geometric constraints. The algorithm is efficient, robust and can cope with wide disparity. The algorithm is demonstrated on many real image pairs, and applications on image interpolation and a creation of novel views are also presented.
Maxime Lhuillier, Long Quan
ICPR2
2000 Camera Calibration and Relative Pose Estimation from Gravity
abstract
We examine the potential use of gravity for camera calibration and pose estimation purposes. Concretely, objects being launched or dropped follow trajectories dictated by the law of gravity. We examine if video sequences of such trajectories give us exploitable constraints for estimating the imaging geometry. It is shown that it is possible to estimate the infinite homography and the epipolar geometry between pairs of views from this input, from which we can estimate (some) intrinsic parameters and relative pose. There are less singularities compared to approaches that do not use the information that the observed trajectories follow gravity. In this paper, we sketch the geometric principles of our idea and validate them by numerical simulations.
Peter F. Sturm, Long Quan
ICPR2
2000 A Unified Linear Algorithm for a Novel View Synthesis and Camera Pose Estimation in Mixed Reality
abstract
We propose a linear algorithm that is useful for realizing geometric registration between the view of a real scene and that of a virtual object in an image-based rendering framework. In a unified framework, the novel view synthesis of a virtual object based on three views' matching constraints and the recovery of the camera pose that is necessary for the base image selection can be performed. The feasibility of the algorithm is demonstrated by using ground-truth synthesized data and real scene data.
Toshihiro Kobayashi, Goki Inoue, Yuichi Ohta, Long Quan
VR4
2000 Self-Calibration of a 1D Projective Camera and Its Application to the Self-Calibration of a 2D Projective Camera
abstract
We introduce the concept of self-calibration of a 1D projective camera from point correspondences, and describe a method for uniquely determining the two internal parameters of a 1D camera, based on the trifocal tensor of three 1D images. The method requires the estimation of the trifocal tensor which can be achieved linearly with no approximation unlike the trifocal tensor of 2D images and solving for the roots of a cubic polynomial in one variable. Interestingly enough, we prove that a 2D camera undergoing planar motion reduces to a 1D camera. From this observation, we deduce a new method for self-calibrating a 2D camera using planar motions. Both the self-calibration method for a 1D camera and its applications for 2D camera calibration are demonstrated on real image sequences.
Olivier D. Faugeras, Long Quan, Peter F. Sturm
IEEE Trans. Pattern Anal. Mach. Intell.2
2000 A robust algorithm to estimate the fundamental matrix
Zezhi Chen, Chengke Wu 0001, Peiyi Shen, Long Quan
Pattern Recognit. Lett.5
1999 Image Interpolation by Joint View Triangulation
abstract
Creating novel views by interpolating prestored images or view morphing has many applications in visual simulation. We present in this paper a new method of automatically interpolating two images which tackles two most difficult problems of morphing due to the lack of depth informational pixel matching and visibility handling. We first describe a quasi-dense matching algorithm based on region growing with the best first strategy for match propagation. Then, we describe a robust construction of matched planar patches using local geometric constraints encoded by a homography. After that we introduce a novel representation, joint view triangulation, for visible and half-occluded patches in two images to handle their visibility during the creation of new view. Finally we demonstrate these techniques on real image pairs.
Maxime Lhuillier, Long Quan
CVPR2
1999 Minimal Projective Reconstruction with Missing Data
abstract
The minimal data necessary for projective reconstruction from point correspondences is well-known when the points are visible in all images. In this paper, we formulate and propose solutions to a new family of reconstruction problems from multiple images with minimal data, where there are missing points in some of the images. The ability to handle the minimal cases with missing data is of great theoretical and practical importance. It is unavoidable to use them to bootstrap robust estimation such as RANSAC and LMS algorithms and optimal estimation such as bundle adjustment. First, we develop a framework to parametrize the multiple view geometry, needed to handle the missing data cases. Then we present a solution to the minimal case of 8 points in 3 images, where one of the points is missing in one of the three images. We prove that there are in general as many as 11 solutions for this minimal case. Furthermore, all minimal cases with missing data for 3 and 4 in images are catalogued. Finally we demonstrate the method on both simulated and real images and show that the algorithms presented in this paper can be used for practical problems.
Long Quan, Anders Heyden, Fredrik Kahl
CVPR1
1999 Inherent Two-Way Ambiguity in 2D Projective Reconstruction from Three Uncalibrated 1D Images
abstract
It is shown that there always exists a two-way ambiguity for 2D projective reconstruction from three uncalibrated 1D views independent of the number of point correspondences. It is also shown that the two distinct projective reconstructions are exactly related by a quadratic transformation with the three camera centers as the fundamental points. The unique reconstruction exists only for the case where the three camera centers are aligned. The theoretical results are demonstrated on numerical examples.
Long Quan
ICCV1
1999 Linear N-Point Camera Pose Determination
abstract
The determination of camera position and orientation from known correspondences of 3D reference points and their images is known as pose estimation in computer vision and space resection in photogrammetry. It is well-known that from three corresponding points there are at most four algebraic solutions. Less appears to be known about the cases of four and five corresponding points. We propose a family of linear methods that yield a unique solution to 4- and 5-point pose determination for generic reference points. We first review the 3-point algebraic method. Then we present our two-step, 4-point and one-step, 5-point linear algorithms. The 5-point method can also be extended to handle more than five points. Finally, we demonstrate our methods on both simulated and real images. We show that they do not degenerate for coplanar configurations and even outperform the special linear algorithm for coplanar configurations in practice.
Long Quan, Zhong-Dan Lan
IEEE Trans. Pattern Anal. Mach. Intell.1
1998 Precise Matching by Robust Estimation of Deformation and Local Coherence
Zhong-Dan Lan, Roger Mohr, Long Quan
ACCV (1)3
1998 Robot Stereo-hand Coordination for Grasping Curved Parts
abstract
In this paper we present an algorithm to compute set-point (i.e. a position to be reached by the robot) automatically from conic features virtually placed by an operator onto the object. We then use a visual servoing algorithm to guide the gripper to its final position. We have tested this algorithm in trying to open a valve with a 6 degree of freedom robot arm, using only visual information and without any model of the valve. 1 Introduction Grasping is one of the most common tasks in robotics and telemanipulation. But it is also one of the most difficult to accomplish when operating on objects of unknown shape and size in hazardous or remote environments. Although this task is crucial for many applications (space exploration, nuclear plants inspection, etc.), few systems are able to operate with sufficient flexibility and reliability in complex situations. The difficulties are twofold. First, to grasp an object implies to select a grasp (a relationship between the gripper and the obj...
Yves Dufournaud, Radu Horaud, Long Quan
BMVC3
1998 A New Linear Method for Euclidean Motion/Structure from Three Calibrated Affine Views
abstract
We introduce a unified framework for developing matching constraints of multiple affine views and rederive 2-view (affine epipolar geometry) and 3-view (affine image transfer) constraints within this framework. We then describe a new linear method for Euclidean motion and structure from 3 calibrated affine images, based on insight into the particular structure of these multiple-view constraints. Compared with the existing linear method of Huang and Lee (1989), the new method uses different and more appropriate constraints. It has no failure mode of the Euclidean factorisation method of Tomasi and Kanade (1992). We demonstrate the method on real image sequences.
Long Quan, Yuichi Ohta
CVPR1
1998 Self-Calibration of a 1D Projective Camera and Its Application to the Self-Calibration of a 2D Projective Camera
Olivier D. Faugeras, Long Quan, Peter F. Sturm
ECCV (1)2
1998 Linear N>=4-Point Pose Determination
abstract
The determination of the position and the orientation of the camera from the known correspondences of the reference points and the image points is known as the problem of pose estimation in computer vision or space resection in photogrammetry. It is well known that using 3 corresponding points has at most 4 solutions. Less appears to be known about the cases of 4 and 5 points. In this paper, we describe linear solutions that always give the unique solution to 4-point and 5-point pose determination for the reference points not lying on the critical configurations. The same linear method can also be extended to any n/spl ges/5 points. The robustness and accuracy of the method are experimented both on simulated and real images.
Long Quan, Zhong-Dan Lan
ICCV1
1998 Joint Invariants of a Triplet of Coplanar Conics: Stability and Discriminating Power for Object Recognition
Long Quan, Francoise Veillon
Comput. Vis. Image Underst.1
1997 Uniqueness of 3D Affine Reconstruction of Lines with Affine Cameras
Long Quan, Roger Mohr
CAIP1
1997 Uncalibrated 1D projective camera and 3D affine reconstruction of lines
abstract
We describe a linear algorithm to recover 3D affine shape/motion from line correspondences over three views with uncalibrated affine cameras. The key idea is the introduction of a one-dimensional projective camera. This converts the 3D affine reconstruction of "lines" into 2D projective reconstruction of "points". Using the full tensorial representation of three uncalibrated 1D views, we prove that the 3D affine reconstruction of lines from minimal data is unique up to a re-ordering of the views. 3D affine line reconstruction can be performed by properly rescaling image coordinates instead of using projection matrices. The algorithm is validated on both simulated and real image sequences.
Long Quan
CVPR1
1997 How Useful is Projective Geometry?
Patrick Gros, Richard I. Hartley, Roger Mohr, Long Quan
Comput. Vis. Image Underst.4
1997 Affine structure from line correspondences with uncalibrated affine cameras
abstract
This paper presents a linear algorithm for recovering 3D affine shape and motion from line correspondences with uncalibrated affine cameras. The algorithm requires a minimum of seven line correspondences over three views. The key idea is the introduction of a one-dimensional projective camera. This converts 3D affine reconstruction of "line directions" into 2D projective reconstruction of "points". In addition, a line-based factorization method is also proposed to handle redundant views. Experimental results both on simulated and real image sequences validate the robustness and the accuracy of the algorithm.
Long Quan, Takeo Kanade
IEEE Trans. Pattern Anal. Mach. Intell.1
1996 A Factorization Method for Affine Structure from Line Correspondences
abstract
A family of structure from motion algorithms called the factorization method has been recently developed from the orthographic projection model to the affine camera model. All these algorithms are limited to handling only point features of the image stream. We propose in this paper an algorithm for the recovery of shape and motion from line correspondences by the factorization method with the affine camera. Instead of one step factorization for points, a multi-step factorization method is developed for lines based on the decomposition of the whole shape and motion into three separate substructures. Each of these substructures can then be linearly solved by factorizing the appropriate measurement matrices. It is also established that affine shape and motion with uncalibrated affine cameras can be achieved with at least seven lines over three views, which extends the previous results of Koenderink and Van Doorn (1989) for points to lines.
Long Quan, Takeo Kanade
CVPR1
1996 Self-calibration of an affine camera from multiple views
Long Quan
Int. J. Comput. Vis.1
1996 Conic Reconstruction and Correspondence From Two Views
abstract
Conics are widely accepted as one of the most fundamental image features together with points and line segments. The problem of space reconstruction and correspondence of two conics from two views is addressed in this paper. It is shown that there are two independent polynomial conditions on the corresponding pair of conics across two views, given the relative orientation of the two views. These two correspondence conditions are derived algebraically and one of them is shown to be fundamental in establishing the correspondences of conics. A unified closed-form solution is also developed for both projective reconstruction of conics in space from two uncalibrated camera views and metric reconstruction from two calibrated camera views. Experiments are conducted to demonstrate the discriminality of the correspondence conditions and the accuracy and stability of the reconstruction both for simulated and real images.
Long Quan
IEEE Trans. Pattern Anal. Mach. Intell.1
1995 Self-calibration of an Affine Camera
Long Quan, Roger Mohr
CAIP1
1995 Affine stereo calibration
Peter F. Sturm, Long Quan
CAIP2
1995 Joint Invariants of a Triplet of Coplanar Conics: Stability and Discriminating Power for Object Recognition
Francoise Veillon, Long Quan, Peter F. Sturm
CAIP2
1995 Invariant of a Pair of Non-Coplanar Conies in Space: Definition, Geometric Interpretation and Computation
abstract
The joint invariants of a pair of coplanar conics has been widely used in recent vision literature. In this paper, the algebraic invariant of a pair of non-coplanar conics in space is concerned. The algebraic invariant of a pair of non-coplanar conics is first derived from the invariant algebra of a pair of quaternary quadratic forms by using the dual representation of space conics. Then, this algebraic invariant is geometrically interpreted in terms of cross-ratios. Finally, an analytical procedure for projective reconstruction of a space conic from two uncalibrated images is developed and the correspondence conditions of the conics between two views are also explicited.>
Long Quan
ICCV1
1995 Invariants of Six Points and Projective Reconstruction From Three Uncalibrated Images
abstract
There are three projective invariants of a set of six points in general position in space. It is well known that these invariants cannot be recovered from one image, however an invariant relationship does exist between space invariants and image invariants. This invariant relationship is first derived for a single image. Then this invariant relationship is used to derive the space invariants, when multiple images are available. This paper establishes that the minimum number of images for computing these invariants is three, and the computation of invariants of six points from three images can have as many as three solutions. Algorithms are presented for computing these invariants in closed form. The accuracy and stability with respect to image noise, selection of the triplets of images and distance between viewing positions are studied both through real and simulated images. Applications of these invariants are also presented. Both the results of Faugeras (1992) and Hartley et al. (1992) for projective reconstruction and Sturm's method (1869) for epipolar geometry determination from two uncalibrated images with at least seven points are extended to the case of three uncalibrated images with only six points.>
Long Quan
IEEE Trans. Pattern Anal. Mach. Intell.1
1994 Invariants of 6 Points from 3 Uncalibrated Images
Long Quan
ECCV (2)1
1993 Affine Stereo Calibration for Relative Affine Shape Reconstruction
abstract
International audience
Long Quan
BMVC1
1993 Relative 3-D reconstruction using multiple uncalibrated images
abstract
It is shown how relative 3-D reconstruction for point correspondence of multiple images from uncalibrated cameras can be achieved through reference points. The original contributions with respect to other related works in the field are a direct global method for relative 3-D reconstruction and a geometrical method to select a correct set of reference points among all image points. Experimental results from both simulated and real image sequences are presented.>
Roger Mohr, Francoise Veillon, Long Quan
CVPR3
1992 Invariants of a pair of conics revisited
Long Quan, Patrick Gros, Roger Mohr
Image Vis. Comput.1
1991 Invariants of a Pair of Conies Revisited
Long Quan, Patrick Gros, Roger Mohr
BMVC1
1990 Geometrical learning from multiple stereo views through monocular based feature grouping
abstract
A geometrical learning system is described which constructs automatically the geometric model of indoor environments from multiple stereo views. The produced model includes not only 3-D segments initially provided by a stereo vision system, but also 3-D feature groups. Feature grouping helps to predict the position of a given view with respect to the current model and allows one to correct noisy features by the geometric constraints of feature groups. This 3-D feature grouping is based on a monocular process working on the 2-D segments associated with one source image of a stereo view. Experimental results are presented.>
Eric Thirion, Long Quan
ICCV2
1989 Determining perspective structures using hierarchical Hough transform
Long Quan, Roger Mohr
Pattern Recognit. Lett.1
1988 Matching Perspective Images Using Geometric Constraints And Perceptual Grouping
abstract
International audience
Long Quan, Roger Mohr
ICCV1