VLDB 2026 Research / reviewers in the wild / expert
Ping Tan 0002
dblp:61/6118-2
· DBLP profile ↗
133ranked-venue papers
7as first author
65since 2021 · last 2026
0000-0002-4506-6973ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 108 · 6 first-author · 51 since 2021Artificial intelligence and machine learning · 95 · 4 first-author · 47 since 2021Systems, architecture and hardware · 5 · 3 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SAIL-Recon: Large SfM by Augmenting Scene Regression with LocalizationabstractScene regression methods, such as VGGT [86], solve the Structure-from-Motion (SfM) problem by directly regressing camera poses and 3D scene structures from input images. They demonstrate impressive performance in handling images under extreme viewpoint changes. However, these methods struggle to handle a large number of input images. To address this problem, we introduce SAIL-Recon, a feed-forward Transformer for large scale SfM, by augmenting the scene regression network with visual localization capabilities. Specifically, our method first computes a neural scene representation tokens from a subset of anchor images. The regression network is then fine-tuned to reconstruct all input images conditioned on this neural scene representation. Comprehensive experiments show that our method not only scales efficiently to large-scale scenes, but also achieves state-of-the-art results on both camera pose estimation and novel view synthesis benchmarks, including TUM-RGBD, CO3Dv2, and Tanks & Temples. Code and models are publicly available here. Junyuan Deng, Heng Li 0009, Weiqiang Ren, Qian Zhang 0009, Ping Tan 0002 |
3DV | 6 |
| 2026 | SPATIALGEN: Layout-Guided 3D Indoor Scene Generation
Chuan Fang, Heng Li 0009, Yixun Liang, Jia Zheng 0002, Yongsen Mao, Yuan Liu 0025, Rui Tang 0015, Zihan Zhou 0001, Ping Tan 0002 |
3DV | 9 |
| 2026 | CTR3D: Cross-View Token Reduction for Dense Multi-View GenerationabstractRecent multi-view diffusion (MVD) methods have utilized the generative capabilities of 2D image diffusion models to produce multi-view images from a single-view input. However, existing approaches often depend on dense crossview attention layers, which hinder scalability and fidelity due to their high computational costs. In this paper, we propose CTR3D, a novel method that incorporates token reduction in multi-view attention layers to efficiently generate dense, high-resolution multi-view images without restricting the camera viewpoints of the generated views. Our approach is designed into three key steps: redundancy removal, attention interaction, and token recovery. These steps leverage lightweight, projection-based techniques for multi-view token reduction and recovery, significantly improving the computational efficiency of MVD. By reducing the number of tokens in attention layers while preserving multi-view consistency, our model achieves state-of-the-art performance in novel view synthesis and 3D reconstruction while keeping efficiency for generation of dense high-resolution images and normals. Experimental results demonstrate that our method surpasses existing approaches, providing a more efficient and effective solution for multi-view generation. https://github.com/HKUST-SAIL/CTR3D Kunming Luo, Hongyu Yan, Yuan Liu 0025, Manyuan Zhang, Wenping Wang 0001, Ping Tan 0002 |
3DV | 7 |
| 2026 | Learning Efficient Meshflow and Optical Flow From Event CamerasabstractIn this paper, we explore the problem of event-based meshflow estimation, a novel task that involves predicting a spatially smooth sparse motion field from event cameras. To start, we review the state-of-the-art in event-based flow estimation, highlighting two key areas for further research: i) the lack of meshflow-specific event datasets and methods, and ii) the underexplored challenge of event data density. First, we generate a large-scale High-Resolution Event Meshflow (HREM) dataset, which showcases its superiority by encompassing the merits of high resolution at 1280 × 720, handling dynamic objects and complex motion patterns, and offering both optical flow and meshflow labels. These aspects have not been fully explored in previous works. Besides, we propose Efficient Event-based MeshFlow (EEMFlow) network, a lightweight model featuring a specially crafted encoder-decoder architecture to facilitate swift and accurate meshflow estimation. Furthermore, we upgrade EEMFlow network to support dense event optical flow, in which a Confidence-induced Detail Completion (CDC) module is proposed to preserve sharp motion boundaries. We conduct comprehensive experiments to show the exceptional performance and runtime efficiency (30×faster) of our EEMFlow model compared to the recent state-of-the-art flow method. As an extension, we expand HREM into HREM+, a multi-density event dataset contributing to a thorough study of the robustness of existing methods across data with varying densities, and propose an Adaptive Density Module (ADM) to adjust the density of input event data to a more optimal range, enhancing the model's generalization ability. We empirically demonstrate that ADM helps to significantly improve the performance of EEMFlow and EEMFlow+ by 8% and 10%, respectively. Xinglong Luo, Ao Luo, Kunming Luo, Zhengning Wang, Ping Tan 0002, Bing Zeng 0001, Shuaicheng Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | ROVER: Robust Loop Closure Verification With Trajectory Prior in Repetitive EnvironmentsabstractLoop closure detection is important for simultaneous localization and mapping (SLAM), which associates current observations with historical keyframes, achieving drift correction and global relocalization. However, a falsely detected loop can be fatal, and this is especially difficult in repetitive environments where appearance-based features fail due to the high similarity. Therefore, verifying a loop closure is a critical step to avoid false-positive detections. Existing works in loop closure verification predominantly focus on learning invariant appearance features, neglecting the prior knowledge of the robot’s spatial-temporal motion cue, i.e., trajectory. In this article, we propose ROVER, a loop closure verification method that leverages the historical trajectory as a prior constraint to reject false loops in challenging repetitive environments. For each loop candidate, it is first used to estimate the robot trajectory with pose-graph optimization. This trajectory is then submitted to a scoring scheme that assesses its compliance with the trajectory without the loop, which we refer to as the trajectory prior constraint (TPC), to determine if the loop candidate should be accepted. Benchmark comparisons and real-world experiments demonstrate the effectiveness of the proposed method. Furthermore, we integrate ROVER into state-of-the-art SLAM systems to verify its robustness and efficiency. Jingwen Yu, Jianhao Jiao, Anjun Hu, Zhonghang Liu, Jiankun Wang 0001, Ping Tan 0002, Hong Zhang 0013 |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2026 | Iris3D: 3D Generation via Synchronized Diffusion DistillationabstractWe introduce Iris3D, a novel 3D content generation system that generates vivid textures and detailed 3D shapes while preserving the input information. Our system integrates a Multi-View Large Reconstruction Model (MVLRM [Li et al. 2023b ]) to generate a coarse 3D mesh and introduces a novel optimization scheme called Synchronized Diffusion Distillation (SDD) for refinement. Unlike previous refined methods based on Score Distillation Sampling (SDS), which suffer from unstable optimization and geometric over-smoothing due to ambiguities across different views and modalities, our method effectively distills consistent multi-view and multi-modal priors from 2D diffusion models in a training-free manner. This enables robust optimization of 3D representations. Additionally, because SDD is training-free, it preserves the diffusion’s prior knowledge and mitigates potential degradation. This characteristic makes it highly compatible with advanced 2D diffusion techniques like IP-Adapters and ControlNet, allowing for more controllable 3D generation with additional conditioning signals. Experiments demonstrate that our method produces high-quality 3D results with plausible textures and intricate geometric details. Yixun Liang, Fei-Peng Tian, Jiarui Liu 0003, Ying-Cong Chen, Ping Tan 0002, Xiaoxiao Long |
ACM Trans. Graph. | 7 |
| 2026 | RaDe-GS: Rasterizing Depth in Gaussian SplattingabstractGaussian Splatting (GS) has proven to be highly effective in novel view synthesis, achieving high-quality and real-time rendering. However, its potential for reconstructing detailed 3D shapes has not been fully explored. Existing methods often suffer from limited shape accuracy due to the discrete and unstructured nature of Gaussian primitives, which complicates the shape extraction. While recent techniques like 2D GS have attempted to improve shape reconstruction, they often reformulate the Gaussian primitives in ways that reduce both rendering quality and computational efficiency. To address these problems, our work introduces a rasterized approach to render the depth maps and surface normal maps of general 3D Gaussian primitives. Our method not only significantly enhances shape reconstruction accuracy but also maintains the computational efficiency intrinsic to Gaussian Splatting. It achieves a Chamfer distance error comparable to Neuralangelo Li et al. [ 2023 ] on the DTU dataset and maintains similar computational efficiency as the original 3D GS methods. Our method is a significant advancement in Gaussian Splatting and can be directly integrated into existing Gaussian Splatting-based methods. Baowen Zhang, Chuan Fang, Rakesh Shrestha, Yixun Liang, Xiaoxiao Long, Ping Tan 0002 |
ACM Trans. Graph. | 6 |
| 2025 | Ctrl-Room: Controllable Text-to-3D Room Meshes Generation with Layout ConstraintsabstractText-driven 3D indoor scene generation is useful for gaming, film industry, and AR/VR applications. However, existing methods cannot faithfully capture the scene layout based on text descriptions, nor do they allow flexible editing of individual objects in the room. To address these problems, we present Ctrl-Room, which can generate convincing 3D rooms with designer-style layouts and high-fidelity textures from just a text prompt. Our key insight is to separate the modeling of layouts and appearance. Our proposed method consists of two stages: a Layout Generation Stage and an Appearance Generation Stage. The Layout Generation Stage trains a text-conditional diffusion model to learn the layout distribution with our holistic scene code parameterization. Next, the Appearance Generation Stage employs a fine-tuned ControlNet to produce a vivid panoramic image of the room guided by the 3D scene layout, then further upgrades to a panoramic NeRF model. Benefiting from the scene code parameterization, we can easily edit the generated room model through our mask-guided editing module, without expensive edit-specific training. Extensive experiments on the Structured3D dataset demonstrate that our method outperforms existing methods in producing more reasonable, view-consistent, and editable 3D rooms from text prompts. Chuan Fang, Kunming Luo, Xiaotao Hu, Rakesh Shrestha, Ping Tan 0002 |
3DV | 6 |
| 2025 | Gaussianavatar-Editor: Photorealistic Animatable Gaussian Head Avatar EditorabstractWe introduce GaussianAvatar-Editor, an innovative framework for text-driven editing of animatable Gaussian head avatars that can be fully controlled in expression, pose, and viewpoint. Unlike static 3D Gaussian editing, editing animatable 4D Gaussian avatars presents challenges related to motion occlusion and spatial-temporal inconsistency. To address these issues, we propose the Weighted Alpha Blending Equation (WABE). This function enhances the blending weight of visible Gaussians while suppressing the influence on non-visible Gaussians, effectively handling motion occlusion during editing. Furthermore, to improve editing quality and ensure 4D consistency, we incorporate conditional adversarial learning into the editing process. This strategy helps to refine the edited results and maintain consistency throughout the animation. By integrating these methods, our GaussianAvatar-Editor Xiangyue Liu 0001, Kunming Luo, Heng Li 0009, Yuan Liu 0025, Li Yi 0001, Ping Tan 0002 |
3DV | 7 |
| 2025 | Universal Features Guided Zero-Shot Category-Level Object Pose EstimationabstractObject pose estimation, crucial in computer vision and robotics applications, faces challenges with the diversity of unseen categories. We propose a zero-shot method to achieve category-level 6-DOF object pose estimation, which exploits both 2D and 3D universal features of input RGB-D image to establish semantic similarity-based correspondences and can be extended to unseen categories without additional model fine-tuning. Our method begins with combining efficient 2D universal features to find sparse correspondences between intra-category objects and gets initial coarse pose. To handle the correspondence degradation of 2D universal features if the pose deviates much from the target pose, we use an iterative strategy to optimize the pose. Subsequently, to resolve pose ambiguities due to shape differences between intra-category objects, the coarse pose is refined by optimizing with dense alignment constraint of 3D universal features. Our method outperforms previous methods on the REAL275 and Wild6D benchmarks for unseen categories. Wentian Qu, Chenyu Meng, Heng Li 0009, Jian Cheng 0006, CuiXia Ma, Hongan Wang, Xiao Zhou 0023, Xiaoming Deng 0001, Ping Tan 0002 |
AAAI | 9 |
| 2025 | Dora: Sampling and Benchmarking for 3D Shape Variational Auto-EncodersabstractRecent 3D content generation pipelines commonly employ Variational Autoencoders (VAEs) to encode shapes into compact latent representations for diffusion-based generation. However, the widely adopted uniform point sampling strategy in Shape VAE training often leads to a significant loss of geometric details, limiting the quality of shape reconstruction and downstream generation tasks. We present Dora-Vae, a novel approach that enhances VAE reconstruction through our proposed sharp edge sampling strategy and a dual cross-attention mechanism. By identifying and prioritizing regions with high geometric complexity during training, our method significantly improves the preservation of fine-grained shape features. Such sampling strategy and the dual attention mechanism enable the VAE to focus on crucial geometric details that are typically missed by uniform sampling approaches. To systematically evaluate VAE reconstruction quality, we additionally propose Dora-Bench, a benchmark that quantifies shape complexity through the density of sharp edges, introducing a new metric focused on reconstruction accuracy at these salient geometric features. Extensive experiments on the Dora-Bench demonstrate that Dora-Vae achieves comparable reconstruction quality to the state-of-the-art dense XCube-Vae while requiring a latent space at least 8× smaller (1,280 vs. > 10,000 codes). Project page: https://aruichen.github.io/Dora. Yixun Liang, Guan Luo, Jiarui Liu 0003, Xiu Li 0001, Xiaoxiao Long, Jiashi Feng, Ping Tan 0002 |
CVPR | 10 |
| 2025 | CraftsMan3D: High-fidelity Mesh Generation with 3D Native Diffusion and Interactive Geometry RefinerabstractWe present a novel generative 3D modeling system, coined CraftsMan3D, which can generate high-fidelity 3D geometries with highly varied shapes, detailed surfaces, and, notably, allows for refining the geometry in an interactive manner. Despite the significant advancements in 3D generation, existing methods still struggle with lengthy optimization processes, self-occlusion, irregular mesh topologies, and difficulties in accommodating user editing, consequently impeding their widespread adoption and implementation in 3D modeling softwares. Our work is inspired by the craftsman, who usually roughs out the holistic figure of the work first and elaborates the surface details subsequently. Specifically, we first introduce a robust data preprocessing pipeline that utilizes visibility check and winding mumber to maximize the use of existing 3D data. Leveraging this data, we employ a 3D-native DiT model that directly models the distribution of 3D data in latent space, generating coarse geometries in seconds. Subsequently, a normal-based geometry refiner enhances local surface details, which can be applied automatically or interactively with user input. Extensive experiments demonstrate that our method achieves high efficacy in producing superior quality 3D meshes compared to existing methods. Jiarui Liu 0003, Hongyu Yan, Yixun Liang, Xuelin Chen, Ping Tan 0002, Xiaoxiao Long |
CVPR | 7 |
| 2025 | SkillMimic: Learning Basketball Interaction Skills from DemonstrationsabstractTraditional reinforcement learning methods for human- object interaction (HOI) rely on labor-intensive, manually designed skill rewards that do not generalize well across different interactions. We introduce SkillMimic, a unified data-driven framework that fundamentally changes how agents learn interaction skills by eliminating the need for skill-specific rewards. Our key insight is that a unified HOI imitation reward can effectively capture the essence of diverse interaction patterns from HOI datasets. This enables SkillMimic to learn a single policy that not only masters multiple interaction skills but also facilitates skill transitions, with both diversity and generalization improving as the HOI dataset grows. For evaluation, we collect and introduce two basketball datasets containing approximately 35 minutes of diverse basketball skills. Extensive experiments show that SkillMimic successfully masters a wide range of basketball skills including stylistic variations in dribbling, layup, and shooting. Moreover, these learned skills can be effectively composed by a high-level controller to accomplish complex and long-horizon tasks such as consecutive scoring, opening new possibilities for scalable and generalizable interaction skill learning. Project page: https://ingrid789.github.io/SkillMimic/ Yinhuai Wang, Qihan Zhao 0001, Runyi Yu 0003, Hok Wai Tsui, Ailing Zeng, Jiwen Yu, Xiu Li 0001, Qifeng Chen 0001, Jian Zhang 0018, Lei Zhang 0001, Ping Tan 0002 |
CVPR | 13 |
| 2025 | Boost 3D Reconstruction Using Diffusion-Based Monocular Camera CalibrationabstractIn this paper, we present DM-Calib, a diffusion-based approach for estimating pinhole camera intrinsic parameters from a single input image. Monocular camera calibration is essential for many 3D vision tasks. However, most existing methods depend on handcrafted assumptions or are constrained by limited training data, resulting in poor generalization across diverse real-world images. Recent advancements in stable diffusion models, trained on massive data, have shown the ability to generate high-quality images with varied characteristics. Emerging evidence indicates that these models implicitly capture the relationship between camera focal length and image content. Building on this insight, we explore how to leverage the powerful priors of diffusion models for monocular pinhole camera calibration. Specifically, we introduce a new image-based representation, termed Camera Image, which losslessly encodes the numerical camera intrinsics and integrates seamlessly with the diffusion framework. Using this representation, we reformulate the problem of estimating camera intrinsics as the generation of a dense Camera Image conditioned on an input image. By fine-tuning a stable diffusion model to generate a Camera Image from a single RGB input, we can extract camera intrinsics via a RANSAC operation. We further demonstrate that our monocular calibration method enhances performance across various 3D tasks, including zero-shot metric depth estimation, 3D metrology, pose estimation and sparse-view reconstruction. Extensive experiments on multiple public datasets show that our approach significantly outperforms baselines and provides broad benefits to 3D vision tasks. Junyuan Deng, Wei Yin 0006, Qian Zhang 0001, Xiaotao Hu, Weiqiang Ren, Xiao-Xiao Long, Ping Tan 0002 |
ICCV | 8 |
| 2025 | SpatialLM: Training Large Language Models for Structured Indoor ModelingabstractSpatialLM is a large language model designed to process 3D point cloud data and generate structured 3D scene understanding outputs. These outputs include architectural elements like walls, doors, windows, and oriented object boxes with their semantic categories. Unlike previous methods which exploit task-specific network designs, our model adheres to the standard multimodal LLM architecture and is fine-tuned directly from open-source LLMs.
To train SpatialLM, we collect a large-scale, high-quality synthetic dataset consisting of the point clouds of 12,328 indoor scenes (54,778 rooms) with ground-truth 3D annotations, and conduct a careful study on various modeling and training decisions. On public benchmarks, our model gives state-of-the-art performance in layout estimation and competitive results in 3D object detection. With that, we show a feasible path for enhancing the spatial understanding capabilities of modern LLMs for applications in augmented reality, embodied robotics, and more. Yongsen Mao, Junhao Zhong, Chuan Fang, Jia Zheng 0002, Rui Tang 0015, Hao Zhu 0004, Ping Tan 0002, Zihan Zhou 0001 |
NeurIPS | 7 |
| 2025 | Minimum Latency Deep Online Video Stabilization and Its ExtensionsabstractWe present a novel deep camera path optimization framework for minimum latency online video stabilization. Typically, a stabilization pipeline consists of three steps: motion estimation, path smoothing, and novel view synthesis. Most previous methods concentrate on motion estimation while path optimization receives less attention, particularly in the crucial online setting where future frames are inaccessible. In this work, we adopt off-the-shelf high-quality deep motion models for motion estimation and focus only on the path optimization. Specifically, our camera path smoothing network takes a short 2D camera path in a sliding window as input and outputs the stabilizing warp field of the last frame, which warps the coming frame to its stabilized position. We explore three motion densities: a global single camera path, local mesh-based bundled paths, and dense flow paths. A hybrid loss and an efficient motion smoothing attention (EMSA) module are proposed for spatially and temporally consistent path smoothing. Moreover, we build a motion dataset that contains stable and unstable motion pairs for training. Extensive experiments demonstrate that our method surpasses state-of-the-art online stabilization methods and rivals the performance of offline methods, offering compelling advancements in the field of video stabilization. Shuaicheng Liu, Zhuofan Zhang, Zhen Liu 0022, Ping Tan 0002, Bing Zeng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Learning Photometric Feature Transform for Free-Form Object ScanabstractWe propose a novel framework to automatically learn to aggregate and transform photometric measurements from multiple unstructured views into spatially distinctive and view-invariant low-level features, which are subsequently fed to a multi-view stereo pipeline to enhance 3D reconstruction. The illumination conditions during acquisition and the feature transform are jointly trained on a large amount of synthetic data. We further build a system to reconstruct both the geometry and anisotropic reflectance of a variety of challenging objects from hand-held scans. The effectiveness of the system is demonstrated with a lightweight prototype, consisting of a camera and an array of LEDs, as well as an off-the-shelf tablet. Our results are validated against reconstructions from a professional 3D scanner and photographs, and compare favorably with state-of-the-art techniques. Xiang Feng 0004, Kaizhang Kang, Fan Pei, Huakeng Ding, Jinjiang You, Ping Tan 0002, Kun Zhou 0001, Hongzhi Wu |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | Efficient 3D Implicit Head Avatar With Mesh-Anchored Hash Table Blendshapesabstract3D head avatars built with neural implicit volumetric representations have achieved unprecedented levels of pho-torealism. However, the computational cost of these methods remains a significant barrier to their widespread adoption, particularly in real-time applications such as virtual reality and teleconferencing. While attempts have been made to develop fast neural rendering approaches for static scenes, these methods cannot be simply employed to support realistic facial expressions, such as in the case of a dynamic facial performance. To address these challenges, we propose a novel fast 3D neural implicit head avatar model that achieves real-time rendering while maintaining fine-grained controllability and high rendering quality. Our key idea lies in the introduction of local hash table blendshapes, which are learned and attached to the vertices of an underlying face parametric model. These per-vertex hash-tables are linearly merged with weights predicted via a CNN, re-sulting in expression dependent embeddings. Our novel representation enables efficient density and color predictions using a lightweight MLP, which is further accelerated by a hierarchical nearest neighbor search method. Extensive experiments show that our approach runs in real-time while achieving comparable rendering quality to state-of-the-arts and decent results on challenging expressions. Ziqian Bai, Feitong Tan, Sean Ryan Fanello, Rohit Pandey, Mingsong Dou, Shichen Liu, Ping Tan 0002, Yinda Zhang 0001 |
CVPR | 7 |
| 2024 | PanoContext-Former: Panoramic Total Scene Understanding with a TransformerabstractPanoramic images enable deeper understanding and more holistic perception of 360° surrounding environment, which can naturally encode enriched scene context information compared to standard perspective image. Previous work has made lots of effort to solve the scene understanding task in a hybrid solution based on 2D-3D geometric reasoning, thus each sub-task is processed separately and few correlations are explored in this procedure. In this paper, we propose a fully 3D method for holistic indoor scene understanding which recovers the objects' shapes, oriented bounding boxes and the 3D room layout simultaneously from a single panorama. To maximize the exploration of the rich context information, we design a transformer-based context module to predict the representation and relationship among each component of the scene. In addition, we introduce a new dataset for scene understanding, including photo-realistic panoramas, high-fidelity depth images, accurately annotated room layouts, oriented object bounding boxes and shapes. Experiments on the synthetic and new datasets demonstrate that our method outperforms previous panoramic scene understanding methods in terms of both layout estimation and 3D object detection. Chuan Fang, Liefeng Bo, Zilong Dong, Ping Tan 0002 |
CVPR | 5 |
| 2024 | VMINer: Versatile Multi-view Inverse Rendering with Near-and Far-field Light SourcesabstractThis paper introduces a versatile multi-view inverse rendering framework with near-and far-field light sources. Tackling the fundamental challenge of inherent ambiguity in inverse rendering, our framework adopts a lightweight yet inclusive lighting model for different near-and far-field lights, thus is able to make use of input images under varied lighting conditions available during capture. It leverages observations under each lighting to disentangle the intrinsic geometry and material from the external lighting, using both neural radiance field rendering and physically-based surface rendering on the 3D implicit fields. After training, the reconstructed scene is extracted to a textured triangle mesh for seamless integration into industrial rendering soft-ware for various applications. Quantitatively and qualitatively tested on synthetic and real-world scenes, our method shows superiority to state-of-the-art multi-view inverse rendering methods in both speed and quality. Jiajun Tang 0001, Ping Tan 0002, Boxin Shi |
CVPR | 3 |
| 2024 | GenN2N: Generative NeRF2NeRF TranslationabstractWe present GenN2N, a unified NeRF-to-NeRF translation framework for various NeRF translation tasks such as text-driven NeRF editing, colorization, super-resolution, in-painting, etc. Unlike previous methods designed for individual translation tasks with task-specific schemes, GenN2N achieves all these NeRF editing tasks by employing a plug-and-play image-to-image translator to perform editing in the 2D domain and lifting 2D edits into the 3D NeRF space. Since the 3D consistency of 2D edits may not be assured, we propose to model the distribution of the underlying 3D edits through a generative model that can cover all possible edited NeRFs. To model the distribution of 3D edited NeRFs from 2D edited images, we carefully design a VAE-GAN that encodes images while decoding NeRFs. The latent space is trained to align with a Gaussian distribution and the NeRFs are supervised through an adversarial loss on its renderings. To ensure the latent code does not depend on 2D viewpoints but truly reflects the 3D edits, we also regularize the latent code through a contrastive learning scheme. Extensive experiments on various editing tasks show GenN2N, as a universal framework, performs as well or better than task-specific specialists while possessing flexible generative power. More results on our project page: https://xiangyueliu.github.io/GenN2N/. Xiangyue Liu 0001, Kunming Luo, Ping Tan 0002, Li Yi 0001 |
CVPR | 4 |
| 2024 | Bilateral Propagation Network for Depth CompletionabstractDepth completion aims to derive a dense depth map from sparse depth measurements with a synchronized color image. Current state-of-the-art (SOTA) methods are predominantly propagation-based, which work as an iterative refinement on the initial estimated dense depth. However, the initial depth estimations mostly result from direct applications of convolutional layers on the sparse depth map. In this paper, we present a Bilateral Propagation Network (BP-Net), that propagates depth at the earliest stage to avoid directly convolving on sparse data. Specifically, our approach propagates the target depth from nearby depth measurements via a non-linear model, whose coefficients are generated through a multi-layer perceptron conditioned on both radiometric difference and spatial distance. By integrating bilateral propagation with multi-modal fusion and depth refinement in a multi-scale framework, our BP-Net demonstrates outstanding performance on both indoor and outdoor scenes. It achieves SOTA on the NYUv2 dataset and ranks 1st on the KITTI depth completion benchmark at the time of submission. Experimental results not only show the effectiveness of bilateral propagation but also emphasize the significance of early-stage propagation in contrast to the refinement stage. Our code and trained models will be available on the project page. Jie Tang 0015, Fei-Peng Tian, Boshi An, Jian Li 0003, Ping Tan 0002 |
CVPR | 5 |
| 2024 | PointRegGPT: Boosting 3D Point Cloud Registration Using Generative Point-Cloud Pairs for Training
Suyi Chen, Hao Xu 0018, Haipeng Li 0001, Kunming Luo, Guanghui Liu 0001, Chi-Wing Fu, Ping Tan 0002, Shuaicheng Liu |
ECCV (51) | 7 |
| 2024 | GeoWizard: Unleashing the Diffusion Priors for 3D Geometry Estimation from a Single Image
Wei Yin 0006, Mu Hu, Yuexin Ma, Ping Tan 0002, Shaojie Shen, Dahua Lin, Xiaoxiao Long |
ECCV (22) | 6 |
| 2024 | GV-Bench: Benchmarking Local Feature Matching for Geometric Verification of Long-term Loop Closure DetectionabstractVisual loop closure detection is an important module in visual simultaneous localization and mapping (SLAM), which associates current camera observation with previously visited places. Loop closures correct drifts in trajectory estimation to build a globally consistent map. However, a false loop closure can be fatal, so verification is required as an additional step to ensure robustness by rejecting the false positive loops. Geometric verification has been a well-acknowledged solution that leverages spatial clues provided by local feature matching to find true positives. Existing feature matching methods focus on homography and pose estimation in long-term visual localization, lacking references for geometric verification. To fill the gap, this paper proposes a unified benchmark targeting geometric verification of loop closure detection under long-term conditional variations. Furthermore, we evaluate six representative local feature matching methods (handcrafted and learning-based) under the benchmark, with in-depth analysis for limitations and future directions. Jingwen Yu, Hanjing Ye, Jianhao Jiao, Ping Tan 0002, Hong Zhang 0013 |
IROS | 4 |
| 2024 | Era3D: High-Resolution Multiview Diffusion using Efficient Row-wise AttentionabstractIn this paper, we introduce **Era3D**, a novel multiview diffusion method that generates high-resolution multiview images from a single-view image. Despite significant advancements in multiview generation, existing methods still suffer from camera prior mismatch, inefficacy, and low resolution, resulting in poor-quality multiview images. Specifically, these methods assume that the input images should comply with a predefined camera type, e.g. a perspective camera with a fixed focal length, leading to distorted shapes when the assumption fails. Moreover, the full-image or dense multiview attention they employ leads to a dramatic explosion of computational complexity as image resolution increases, resulting in prohibitively expensive training costs. To bridge the gap between assumption and reality, Era3D first proposes a diffusion-based camera prediction module to estimate the focal length and elevation of the input image, which allows our method to generate images without shape distortions. Furthermore, a simple but efficient attention layer, named row-wise attention, is used to enforce epipolar priors in the multiview diffusion, facilitating efficient cross-view information fusion. Consequently, compared with state-of-the-art methods, Era3D generates high-quality multiview images with up to a 512×512 resolution while reducing computation complexity of multiview attention by 12x times. Comprehensive experiments demonstrate the superior generation power of Era3D- it can reconstruct high-quality and detailed 3D meshes from diverse single-view input images, significantly outperforming baseline multiview diffusion methods. Yuan Liu 0025, Xiaoxiao Long, Feihu Zhang, Cheng Lin 0001, Xingqun Qi, Shanghang Zhang, Wei Xue 0002, Wenhan Luo, Ping Tan 0002, Wenping Wang 0001, Yike Guo |
NeurIPS | 11 |
| 2024 | DMHomo: Learning Homography with Diffusion ModelsabstractSupervised homography estimation methods face a challenge due to the lack of adequate labeled training data. To address this issue, we propose DMHomo , a diffusion model-based framework for supervised homography learning. This framework generates image pairs with accurate labels, realistic image content, and realistic interval motion, ensuring that they satisfy adequate pairs. We utilize unlabeled image pairs with pseudo labels such as homography and dominant plane masks, computed from existing methods, to train a diffusion model that generates a supervised training dataset. To further enhance performance, we introduce a new probabilistic mask loss, which identifies outlier regions through supervised training, and an iterative mechanism to optimize the generative and homography models successively. Our experimental results demonstrate that DMHomo effectively overcomes the scarcity of qualified datasets in supervised homography learning and improves generalization to real-world scenes. The code and dataset are available at GitHub ( https://github.com/lhaippp/DMHomo ). Haipeng Li 0001, Hai Jiang 0006, Ao Luo, Ping Tan 0002, Haoqiang Fan, Bing Zeng 0001, Shuaicheng Liu |
ACM Trans. Graph. | 4 |
| 2023 | Learning Personalized High Quality Volumetric Head Avatars from Monocular RGB VideosabstractWe propose a method to learn a high-quality implicit 3D head avatar from a monocular RGB video captured in the wild. The learnt avatar is driven by a parametric face model to achieve user-controlled facial expressions and head poses. Our hybrid pipeline combines the geometry prior and dynamic tracking of a 3DMM with a neural radiance field to achieve fine-grained control and photorealism. To reduce over-smoothing and improve out-of-model expressions synthesis, we propose to predict local features anchored on the 3DMM geometry. These learnt features are driven by 3DMM deformation and interpolated in 3D space to yield the volumetric radiance at a designated query point. We further show that using a Convolutional Neural Network in the UV space is critical in incorporating spatial context and producing representative local features. Extensive experiments show that we are able to reconstruct high-quality avatars, with more accurate expression-dependent details, good generalization to out-of-training expressions, and quantitatively superior renderings compared to other state-of-the-art approaches. Ziqian Bai, Feitong Tan, Zeng Huang, Kripasindhu Sarkar, Danhang Tang, Di Qiu, Abhimitra Meka, Ruofei Du, Mingsong Dou, Sergio Orts, Rohit Pandey, Ping Tan 0002, Thabo Beeler, Sean Ryan Fanello, Yinda Zhang 0001 |
CVPR | 12 |
| 2023 | NeuMap: Neural Coordinate Mapping by Auto-Transdecoder for Camera LocalizationabstractThis paper presents an end-to-end neural mapping method for camera localization, encoding a whole scene into a grid of latent codes, with which a Transformer-based auto-decoder regresses 3D coordinates of query pixels. State-of-the-art camera localization methods require each scene to be stored as a 3D point cloud with per-point features, which takes several gigabytes of storage per scene. While compression is possible, the performance drops significantly at high compression rates. NeuMap achieves extremely high compression rates with minimal performance drop by using 1) learnable latent codes to store scene information and 2) a scene-agnostic Transformer-based auto-decoder to infer coordinates for a query pixel. The scene-agnostic network design also learns robust matching priors by training with large-scale data, and further allows us to just optimize the codes quickly for a new scene while fixing the network weights. Extensive evaluations with five benchmarks show that NeuMap outperforms all the other coordinate regression methods significantly and reaches similar performance as the feature matching methods while having a much smaller scene representation size. For example, NeuMap achieves 39.1% accuracy in Aachen night benchmark with only 6MB of data, while other compelling methods require 100MB or a few gigabytes and fail completely under high compression settings. The codes are available at https://github.com/Tangshitao/NeuMap. Shitao Tang, Sicong Tang, Andrea Tagliasacchi, Ping Tan 0002, Yasutaka Furukawa |
CVPR | 4 |
| 2023 | Learning Optical Flow from Event Camera with Rendered DatasetabstractWe study the problem of estimating optical flow from event cameras. One important issue is how to build a high-quality event-flow dataset with accurate event values and flow labels. Previous datasets are created by either capturing real scenes by event cameras or synthesizing from images with pasted foreground objects. The former case can produce real event values but with calculated flow labels, which are sparse and inaccurate. The latter case can generate dense flow labels but the interpolated events are prone to errors. In this work, we propose to render a physically correct event-flow dataset using computer graphics models. In particular, we first create indoor and outdoor 3D scenes by Blender with rich scene content variations. Second, diverse camera motions are included for the virtual capturing, producing images and accurate flow labels. Third, we render high-framerate videos between images for accurate events. The rendered dataset can adjust the density of events, based on which we further introduce an adaptive density module (ADM). Experiments show that our proposed dataset can facilitate event-flow learning, whereas previous approaches when trained on our dataset can improve their performances constantly by a relatively large margin. In addition, event-flow pipelines when equipped with our ADM can further improve performances. Our code is available at https://github.com/boomluo02/ADMFlow. Xinglong Luo, Kunming Luo, Ao Luo, Zhengning Wang, Ping Tan 0002, Shuaicheng Liu |
ICCV | 5 |
| 2023 | DPS-Net: Deep Polarimetric Stereo Depth EstimationabstractStereo depth estimation usually struggles to deal with textureless scenes for both traditional and learning-based methods due to the inherent dependence on image correspondence matching. In this paper, we propose a novel neural network, i.e., DPS-Net, to exploit both the prior geometric knowledge and polarimetric information for depth estimation with two polarimetric stereo images. Specifically, we construct both RGB and polarization correlation volumes to fully leverage the multi-domain similarity between polarimetric stereo images. Since inherent ambiguities exist in the polarization images, we introduce the iso-depth cost explicitly into the network to solve these ambiguities. Moreover, we design a cascaded dual-GRU architecture to recurrently update the disparity and effectively fuse both the multi-domain correlation features and the iso-depth cost. Besides, we present new synthetic and real polarimetric stereo datasets for evaluation. Experimental results demonstrate that our method outperforms the state-of-the-art stereo depth estimation methods. Chaoran Tian, Weihong Pan, Zimo Wang, Mao Mao, Guofeng Zhang 0001, Hujun Bao, Ping Tan 0002, Zhaopeng Cui |
ICCV | 7 |
| 2023 | Minimum Latency Deep Online Video StabilizationabstractWe present a novel camera path optimization framework for the task of online video stabilization. Typically, a stabilization pipeline consists of three steps: motion estimating, path smoothing, and novel view rendering. Most previous methods concentrate on motion estimation, proposing various global or local motion models. In contrast, path optimization receives relatively less attention, especially in the important online setting, where no future frames are available. In this work, we adopt recent off-the-shelf high-quality deep motion models for motion estimation to recover the camera trajectory and focus on the latter two steps. Our network takes a short 2D camera path in a sliding window as input and outputs the stabilizing warp field of the last frame in the window, which warps the coming frame to its stabilized position. A hybrid loss is well-defined to constrain the spatial and temporal consistency. In addition, we build a motion dataset that contains stable and unstable motion pairs for the training. Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art online methods both qualitatively and quantitatively and achieves comparable performance to offline methods. Our code and dataset are available at https://github.com/liuzhen03/NNDVS. Zhuofan Zhang, Zhen Liu 0022, Ping Tan 0002, Bing Zeng 0001, Shuaicheng Liu |
ICCV | 3 |
| 2023 | Dense RGB Slam with Neural Implicit Maps
Heng Li 0009, Xiaodong Gu 0004, Weihao Yuan 0001, Luwei Yang, Zilong Dong, Ping Tan 0002 |
ICLR | 6 |
| 2023 | Compact Real-Time Radiance Fields with Neural CodebookabstractReconstructing neural radiance fields with explicit volumetric representations, demonstrated by Plenoxels, has shown remarkable advantages on training and rendering efficiency, while grid-based representations typically induce considerable overhead for storage and transmission. In this work, we present a simple and effective framework for pursuing compact radiance fields from the perspective of compression methodology. By exploiting intrinsic properties exhibiting in grid models, a non-uniform compression stem is developed to significantly reduce model complexity and a novel parameterized module, named Neural Codebook, is introduced for better encoding high-frequency details specific to per-scene models via a fast optimization. Our approach can achieve over 40 × reduction on grid model storage with competitive rendering quality. In addition, the method can achieve real-time rendering speed with 180 fps, realizing significant advantage on storage cost compared to real-time rendering methods. Lingzhi Li 0002, Zhongshu Wang, Li Shen 0003, Ping Tan 0002 |
ICME | 5 |
| 2023 | Recurrent 3D Hand Pose Estimation Using Cascaded Pose-Guided 3D Alignmentsabstract3D hand pose estimation is a challenging problem in computer vision due to the high degrees-of-freedom of hand articulated motion space and large viewpoint variation. As a consequence, similar poses observed from multiple views can be dramatically different. In order to deal with this issue, view-independent features are required to achieve state-of-the-art performance. In this paper, we investigate the impact of view-independent features on 3D hand pose estimation from a single depth image, and propose a novel recurrent neural network for 3D hand pose estimation, in which a cascaded 3D pose-guided alignment strategy is designed for view-independent feature extraction and a recurrent hand pose module is designed for modeling the dependencies among sequential aligned features for 3D hand pose estimation. In particular, our cascaded pose-guided 3D alignments are performed in 3D space in a coarse-to-fine fashion. First, hand joints are predicted and globally transformed into a canonical reference frame; Second, the palm of the hand is detected and aligned; Third, local transformations are applied to the fingers to refine the final predictions. The proposed recurrent hand pose module for aligned 3D representation can extract recurrent pose-aware features and iteratively refines the estimated hand pose. Our recurrent module could be utilized for both single-view estimation and sequence-based estimation with 3D hand pose tracking. Experiments show that our method improves the state-of-the-art by a large margin on popular benchmarks with the simple yet efficient alignment and network architectures. Xiaoming Deng 0001, Dexin Zuo, Yinda Zhang 0001, Zhaopeng Cui, Jian Cheng 0006, Ping Tan 0002, Liang Chang 0001, Marc Pollefeys, Sean Ryan Fanello, Hongan Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Image Matting With Deep Gaussian ProcessabstractWe observe a common characteristic between the classical propagation-based image matting and the Gaussian process (GP)-based regression. The former produces closer alpha matte values for pixels associated with a higher affinity, while the outputs regressed by the latter are more correlated for more similar inputs. Based on this observation, we reformulate image matting as GP and find that this novel matting-GP formulation results in a set of attractive properties. First, it offers an alternative view on and approach to propagation-based image matting. Second, an application of kernel learning in GP brings in a novel deep matting-GP technique, which is pretty powerful for encapsulating the expressive power of deep architecture on the image relative to its matting. Third, an existing scalable GP technique can be incorporated to further reduce the computational complexity to$\mathcal {O}(n)$from$\mathcal {O}(n^{3})$of many conventional matting propagation techniques. Our deep matting-GP provides an attractive strategy toward addressing the limit of widespread adoption of deep learning techniques to image matting for which a sufficiently large labeled dataset is lacking. A set of experiments on both synthetically composited images and real-world images show the superiority of the deep matting-GP to not only the classical propagation-based matting techniques but also modern deep learning-based approaches. Yuanjie Zheng, Yunshuai Yang, Tongtong Che, Sujuan Hou, Wenhui Huang 0002, Yue Gao 0002, Ping Tan 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2023 | High-Resolution Volumetric Reconstruction for Clothed HumansabstractWe present a novel method for reconstructing clothed humans from a sparse set of, e.g., 1–6 RGB images. Despite impressive results from recent works employing deep implicit representation, we revisit the volumetric approach and demonstrate that better performance can be achieved with proper system design. The volumetric representation offers significant advantages in leveraging 3D spatial context through 3D convolutions, and the notorious quantization error is largely negligible with a reasonably large yet affordable volume resolution, e.g., 512. To handle memory and computation costs, we propose a sophisticated coarse-to-fine strategy with voxel culling and subspace sparse convolution. Our method starts with a discretized visual hull to compute a coarse shape and then focuses on a narrow band nearby the coarse shape for refinement. Once the shape is reconstructed, we adopt an image-based rendering approach, which computes the colors of surface points by blending input images with learned weights. Extensive experimental results show that our method significantly reduces the mean point-to-surface (P2S) precision of state-of-the-art methods by more than 50% to achieve approximately 2mm accuracy with a 512 volume resolution. Additionally, images rendered from our textured model achieve a higher peak signal-to-noise ratio (PSNR) compared to state-of-the-art methods. Sicong Tang, Guangyuan Wang, Qing Ran, Lingzhi Li 0002, Li Shen 0005, Ping Tan 0002 |
ACM Trans. Graph. | 6 |
| 2022 | Efficient Virtual View Selection for 3D Hand Pose Estimationabstract3D hand pose estimation from single depth is a fundamental problem in computer vision, and has wide applications. However, the existing methods still can not achieve satisfactory hand pose estimation results due to view variation and occlusion of human hand. In this paper, we propose a new virtual view selection and fusion module for 3D hand pose estimation from single depth. We propose to automatically select multiple virtual viewpoints for pose estimation and fuse the results of all and find this empirically delivers accurate and robust pose estimation. In order to select most effective virtual views for pose fusion, we evaluate the virtual views based on the confidence of virtual views using a light-weight network via network distillation. Experiments on three main benchmark datasets including NYU, ICVL and Hands2019 demonstrate that our method outperforms the state-of-the-arts on NYU and ICVL, and achieves very competitive performance on Hands2019-Task1, and our proposed virtual view selection and fusion module is both effective for 3D hand pose estimation. Jian Cheng 0006, Yanguang Wan, Dexin Zuo, CuiXia Ma, Ping Tan 0002, Hongan Wang, Xiaoming Deng 0001, Yinda Zhang 0001 |
AAAI | 6 |
| 2022 | Cluster Contrast for Unsupervised Person Re-identification
Zuozhuo Dai, Guangyuan Wang, Weihao Yuan 0001, Siyu Zhu 0001, Ping Tan 0002 |
ACCV (6) | 5 |
| 2022 | RCP: Recurrent Closest Point for Point Cloudabstract3D motion estimation including scene flow and point cloud registration has drawn increasing interest. Inspired by 2D flow estimation, recent methods employ deep neural networks to construct the cost volume for estimating accurate 3D flow. However, these methods are limited by the fact that it is difficult to define a search window on point clouds because of the irregular data structure. In this paper, we avoid this irregularity by a simple yet effective method. We decompose the problem into two interlaced stages, where the 3D flows are optimized point-wisely at the first stage and then globally regularized in a recurrent network at the second stage. Therefore, the recurrent network only receives the regular point-wise information as the input. In the experiments, we evaluate the proposed method on both the 3D scene flow estimation and the point cloud registration task. For 3D scene flow estimation, we make comparisons on the widely used FlyingThings3D [32] and KITTI [33] datasets. For point cloud registration, we follow previous works and evaluate the data pairs with large pose and partially overlapping from ModelNet40 [65]. The results show that our method outperforms the previous method and achieves a new state-of-the-art performance on both 3D scene flow estimation and point cloud registration, which demonstrates the superiority of the proposed zero-order method on irregular point cloud data. Our source code is available at https://github.com/gxd1994/RCP. Xiaodong Gu 0004, Chengzhou Tang, Weihao Yuan 0001, Zuozhuo Dai, Siyu Zhu 0001, Ping Tan 0002 |
CVPR | 6 |
| 2022 | RAGO: Recurrent Graph Optimizer For Multiple Rotation AveragingabstractThis paper proposes a deep recurrent Rotation Averaging Graph Optimizer (RAGO) for Multiple Rotation Averaging (MRA). Conventional optimization-based methods usually fail to produce accurate results due to corrupted and noisy relative measurements. Recent learning-based approaches regard MRA as a regression problem, while these methods are sensitive to initialization due to the gauge freedom problem. To handle these problems, we propose a learnable iterative graph optimizer minimizing a gauge- invariant cost function with an edge rectification strategy to mitigate the effect of inaccurate measurements. Our graph optimizer iteratively refines the global camera rotations by minimizing each node's single rotation objective function. Besides, our approach iteratively rectifies relative rotations to make them more consistent with the current camera orientations and observed relative rotations. Furthermore,$we$employ a gated recurrent unit to improve the result by tracing the temporal information of the cost graph. Our framework is a real-time learning-to-optimize rotation averaging graph optimizer with a tiny size deployed for real-world applications. RAGO outperforms previous traditional and deep methods on real-world and synthetic datasets. The code is available at github.com/sfu-gruvi-3dv/RAGO. Heng Li 0009, Zhaopeng Cui, Shuaicheng Liu, Ping Tan 0002 |
CVPR | 4 |
| 2022 | Learning to Zoom Inside Camera Imaging PipelineabstractExisting single image super-resolution methods are either designed for synthetic data, or for real data but in the RGB-to-RGB or the RAW-to-RGB domain. This paper proposes to zoom an image from RAW to RAW inside the camera imaging pipeline. The RAW-to-RAW domain closes the gap between the ideal and the real degradation models. It also excludes the image signal processing pipeline, which refocuses the model learning onto the super-resolution. To these ends, we design a method that receives a low-resolution RAW as the input and estimates the desired higher-resolution RAW jointly with the degradation model. In our method, two convolutional neural networks are learned to constrain the high-resolution image and the degradation model in lower-dimensional subspaces. This subspace constraint converts the ill-posed SISR problem to a well-posed one. To demonstrate the superiority of the proposed method and the RAW-to-RAW domain, we conduct evaluations on the RealSR and the SR-RAW datasets. The results show that our method performs superiorly over the state-of-the-arts both qualitatively and quantitatively, and it also generalizes well and enables zero-shot transfer across different sensors. Chengzhou Tang, Yuqiang Yang, Bing Zeng 0001, Ping Tan 0002, Shuaicheng Liu |
CVPR | 4 |
| 2022 | SceneSqueezer: Learning to Compress Scene for Camera RelocalizationabstractStandard visual localization methods build a priori 3D model of a scene which is used to establish correspondences against the 2D keypoints in a query image. Storing these pre-built 3D scene models can be prohibitively expensive for large-scale environments, especially on mobile devices with limited storage and communication bandwidth. We design a novel framework that compresses a scene while still maintaining localization accuracy. The scene is compressed in three stages: first, the database frames are clustered using pairwise co-visibility information. Then, a learned point selection module prunes the points in each cluster taking into account the final pose estimation accuracy. In the final stage, the features of the selected points are further compressed using learned quantization. Query image registration is done using only the compressed scene points. To the best of our knowledge, we are the first to propose learned scene compression for visual localization. We also demonstrate the effectiveness and efficiency of our method on various outdoor datasets where it can perform accurate localization with low memory consumption. Luwei Yang, Rakesh Shrestha, Shuaicheng Liu, Guofeng Zhang 0001, Zhaopeng Cui, Ping Tan 0002 |
CVPR | 7 |
| 2022 | Neural Window Fully-connected CRFs for Monocular Depth EstimationabstractEstimating the accurate depth from a single image is challenging since it is inherently ambiguous and ill-posed. While recent works design increasingly complicated and powerful networks to directly regress the depth map, we take the path of CRFs optimization. Due to the expensive computation, CRFs are usually performed between neighborhoods rather than the whole graph. To leverage the potential of fully-connected CRFs, we split the input into windows and perform the FC-CRFs optimization within each window, which reduces the computation complexity and makes FC-CRFs feasible. To better capture the relationships between nodes in the graph, we exploit the multi-head attention mechanism to compute a multi-head potential function, which is fed to the networks to output an optimized depth map. Then we build a bottom-up-top-down structure, where this neural window FC-CRFs module serves as the decoder, and a vision transformer serves as the encoder. The experiments demonstrate that our method significantly improves the performance across all metrics on both the KITTI and NYUv2 datasets, compared to previous methods. Furthermore, the proposed method can be directly applied to panorama images and outperforms all previous panorama methods on the MatterPort3D dataset.11Project page: https://weihaosky.github.io/newcrfs Weihao Yuan 0001, Xiaodong Gu 0004, Zuozhuo Dai, Siyu Zhu 0001, Ping Tan 0002 |
CVPR | 5 |
| 2022 | GAT-CADNet: Graph Attention Network for Panoptic Symbol Spotting in CAD DrawingsabstractSpotting graphical symbols from the computer-aided design (CAD) drawings is essential to many industrial applications. Different from raster images, CAD drawings are vector graphics consisting of geometric primitives such as segments, arcs, and circles. By treating each CAD drawing as a graph, we propose a novel graph attention network GAT-CADNet to solve the panoptic symbol spotting problem: vertex features derived from the GAT branch are mapped to semantic labels, while their attention scores are cascaded and mapped to instance prediction. Our key contributions are three-fold: 1) the instance symbol spotting task is formulated as a subgraph detection problem and solved by predicting the adjacency matrix; 2) a relative spatial encoding (RSE) module explicitly encodes the relative positional and geometric relation among vertices to enhance the vertex attention; 3) a cascaded edge encoding (CEE) module extracts vertex attentions from multiple stages of GAT and treats them as edge encoding to predict the adjacency matrix. The proposed GAT-CADNet is intuitive yet effective and manages to solve the panoptic symbol spotting problem in one consolidated network. Extensive experiments and ablation studies on the public benchmark show that our graph-based approach surpasses existing state-of-the-art methods by a large margin. Zhaohua Zheng, Jianfang Li 0001, Lingjie Zhu, Honghua Li, Frank Petzold, Ping Tan 0002 |
CVPR | 6 |
| 2022 | Domain Randomization-Enhanced Depth Simulation and Restoration for Perceiving and Grasping Specular and Transparent Objects
Qiyu Dai, Jiyao Zhang, Tianhao Wu 0001, Hao Dong 0003, Ziyuan Liu 0003, Ping Tan 0002, He Wang 0010 |
ECCV (39) | 7 |
| 2022 | Quadtree Attention for Vision Transformers
Shitao Tang, Siyu Zhu 0001, Ping Tan 0002 |
ICLR | 4 |
| 2022 | CycleHand: Increasing 3D Pose Estimation Ability on In-the-wild Monocular Image through Cyclic FlowabstractCurrent methods for 3D hand pose estimation fail to generalize well to in-the-wild new scenarios due to varying camera viewpoints, self-occlusions, and complex environments. To address this problem, we propose CycleHand to improve the generalization ability of the model in a self-supervised manner. Our motivation is based on an observation: if one globally rotates the whole hand and reversely rotates it back, the estimated 3D poses of fingers should keep consistent before and after the rotation because the wrist-relative hand poses stay unchanged during global 3D rotation. Hence, we propose arbitrary-rotation self-supervised consistency learning to improve the model's robustness for varying viewpoints. Another innovation of CycleHand is that we propose a high-fidelity texture map to render the photorealistic rotated hand with different lighting conditions, backgrounds, and skin tones to further enhance the effectiveness of our self-supervised task. To reduce the potential negative effects brought by the domain shift of synthetic images, we use the idea of contrastive learning to learn a synthetic-real consistent feature extractor in extracting domain-irrelevant hand representations. Experiments show that CycleHand can largely improve the hand pose estimation performance in both canonical datasets and real-world applications. Daiheng Gao, Xindi Zhang 0003, Xingyu Chen 0002, Andong Tan, Bang Zhang, Ping Tan 0002 |
ACM Multimedia | 7 |
| 2022 | Streaming Radiance Fields for 3D Video SynthesisabstractWe present an explicit-grid based method for efficiently reconstructing streaming radiance fields for novel view synthesis of real world dynamic scenes. Instead of training a single model that combines all the frames, we formulate the dynamic modeling problem with an incremental learning paradigm in which per-frame model difference is trained to complement the adaption of a base model on the current frame. By exploiting the simple yet effective tuning strategy with narrow bands, the proposed method realizes a feasible framework for handling video sequences on-the-fly with high training efficiency. The storage overhead induced by using explicit grid representations can be significantly reduced through the use of model difference based compression. We also introduce an efficient strategy to further accelerate model optimization for each frame. Experiments on challenging video sequences demonstrate that our approach is capable of achieving a training speed of 15 seconds per-frame with competitive rendering quality, which attains $1000 \times$ speedup over the state-of-the-art implicit methods. Lingzhi Li 0002, Zhongshu Wang, Li Shen 0003, Ping Tan 0002 |
NeurIPS | 5 |
| 2022 | Patch-Based Uncalibrated Photometric Stereo Under Natural IlluminationabstractThis paper presents a photometric stereo method that works with unknown natural illumination without any calibration objects or initial guess of the target shape. To solve this challenging problem, we propose the use of an equivalent directional lighting model for small surface patches consisting of slowly varying normals, and solve each patch up to an arbitrary orthogonal ambiguity. We further build the patch connections by extracting consistent surface normal pairs via spatial overlaps among patches and intensity profiles. Guided by these connections, the local ambiguities are unified to a global orthogonal one through Markov Random Field optimization and rotation averaging. After applying the integrability constraint, our solution contains only a binary ambiguity, which could be easily removed. Experiments using both synthetic and real-world datasets show our method provides even comparable results to calibrated methods. Heng Guo 0003, Zhipeng Mo, Boxin Shi, Feng Lu 0005, Sai-Kit Yeung, Ping Tan 0002, Yasuyuki Matsushita |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2021 | MeshMVS: Multi-View Stereo Guided Mesh ReconstructionabstractDeep learning based 3D shape generation methods generally utilize latent features extracted from color images to encode the semantics of objects and guide the shape generation process. These color image semantics only implicitly encode 3D information, potentially limiting the accuracy of the generated shapes. In this paper we propose a multi-view mesh generation method which incorporates geometry information explicitly by using the features from intermediate depth representations of multi-view stereo and regularizing the 3D shapes against these depth images. First, our system predicts a coarse 3D volume from the color images by probabilistically merging voxel occupancy grids from the prediction of individual views. Then the depth images from multi-view stereo along with the rendered depth images of the coarse shape are used as a contrastive input whose features guide the refinement of the coarse shape through a series of graph convolution networks. Notably, we achieve superior results than state-of-the-art multi-view shape generation methods with 34% decrease in Chamfer distance to ground truth and 14% increase in F1-score on ShapeNet dataset. Rakesh Shrestha, Zhiwen Fan, Qingkun Su, Zuozhuo Dai, Siyu Zhu 0001, Ping Tan 0002 |
3DV | 6 |
| 2021 | Riggable 3D Face Reconstruction via In-Network OptimizationabstractThis paper presents a method for riggable 3D face reconstruction from monocular images, which jointly estimates a personalized face rig and per-image parameters including expressions, poses, and illuminations. To achieve this goal, we design an end-to-end trainable network embedded with a differentiable in-network optimization. The network first parameterizes the face rig as a compact latent code with a neural decoder, and then estimates the latent code as well as per-image parameters via a learnable optimization. By estimating a personalized face rig, our method goes beyond static reconstructions and enables downstream applications such as video retargeting. In-network optimization explicitly enforces constraints derived from the first principles, thus introduces additional priors than regression-based methods. Finally, data-driven priors from deep learning are utilized to constrain the ill-posed monocular setting and ease the optimization difficulty. Experiments demonstrate that our method achieves SOTA reconstruction accuracy, reasonable robustness and generalization ability, and supports standard face rig applications. Ziqian Bai, Zhaopeng Cui, Xiaoming Liu 0002, Ping Tan 0002 |
CVPR | 4 |
| 2021 | HumanGPS: Geodesic PreServing Feature for Dense Human CorrespondencesabstractIn this paper, we address the problem of building dense correspondences between human images under arbitrary camera viewpoints and body poses. Prior art either assumes small motion between frames or relies on local descriptors, which cannot handle large motion or visually ambiguous body parts, e.g., left vs. right hand. In contrast, we propose a deep learning framework that maps each pixel to a feature space, where the feature distances reflect the geodesic distances among pixels as if they were projected onto the surface of a 3D human scan. To this end, we introduce novel loss functions to push features apart according to their geodesic distances on the surface. Without any semantic annotation, the proposed embeddings automatically learn to differentiate visually similar parts and align different subjects into an unified feature space. Extensive experiments show that the learned embeddings can produce accurate correspondences between images with remarkable generalization capabilities on both intra and inter subjects.1 Feitong Tan, Danhang Tang, Mingsong Dou, Rohit Pandey, Cem Keskin, Ruofei Du, Deqing Sun, Sofien Bouaziz, Sean Ryan Fanello, Ping Tan 0002, Yinda Zhang 0001 |
CVPR | 11 |
| 2021 | Learning Camera Localization via Dense Scene MatchingabstractCamera localization aims to estimate 6 DoF camera poses from RGB images. Traditional methods detect and match interest points between a query image and a prebuilt 3D model. Recent learning-based approaches encode scene structures into a specific convolutional neural network (CNN) and thus are able to predict dense coordinates from RGB images. However, most of them require re-training or re-adaption for a new scene and have difficulties in handling large-scale scenes due to limited network capacity. We present a new method for scene agnostic camera localization using dense scene matching (DSM), where a cost volume is constructed between a query image and a scene. The cost volume and the corresponding coordinates are processed by a CNN to predict dense coordinates. Camera poses can then be solved by PnP algorithms. In addition, our method can be extended to temporal domain, which leads to extra performance boost during testing time. Our scene-agnostic approach achieves comparable accuracy as the existing scene-specific approaches, such as KFNet, on the 7scenes and Cambridge benchmark. This approach also remarkably outperforms state-of-the-art scene-agnostic dense coordinate regression network SANet. The Code is available at https://github.com/Tangshitao/DenseScene-Matching. Shitao Tang, Chengzhou Tang, Rui Huang 0001, Siyu Zhu 0001, Ping Tan 0002 |
CVPR | 5 |
| 2021 | End-to-End Rotation Averaging With Multi-Source PropagationabstractThis paper presents an end-to-end neural network for multiple rotation averaging in SfM. Due to the manifold constraint of rotations, conventional methods usually take two separate steps involving spanning tree based initialization and iterative nonlinear optimization respectively. These methods can suffer from bad initializations due to the noisy spanning tree or outliers in input relative rotations. To handle these problems, we propose to integrate initialization and optimization together in an unified graph neural network via a novel differentiable multi-source propagation module. Specifically, our network utilizes the image context and geometric cues in feature correspondences to reduce the impact of outliers. Furthermore, unlike the methods that utilize the spanning tree to initialize orientations according to a single reference node in a top-down manner, our net-work initializes orientations according to multiple sources while utilizing information from all neighbors in a differentiable way. More importantly, our end-to-end formulation also enables iterative re-weighting of input relative orientations at test time to improve the accuracy of the final estimation by minimizing the impact of outliers. We demonstrate the effectiveness of our method on two real-world datasets, achieving state-of-the-art performance. Luwei Yang, Heng Li 0009, Jamal Ahmed Rahim, Zhaopeng Cui, Ping Tan 0002 |
CVPR | 5 |
| 2021 | FloorPlanCAD: A Large-Scale CAD Drawing Dataset for Panoptic Symbol SpottingabstractAccess to large and diverse computer-aided design (CAD) drawings is critical for developing symbol spotting algorithms. In this paper, we present FloorPlan-CAD, a large-scale real-world CAD drawing dataset containing over 10,000 floor plans, ranging from residential to commercial buildings. CAD drawings in the dataset are all represented as vector graphics, which enable us to provide line-grained annotations of 30 object categories. Equipped by such annotations, we introduce the task of panoptic symbol spotting, which requires to spot not only instances of countable things, but also the semantic of uncountable stuff. Aiming to solve this task, we propose a novel method by combining Graph Convolutional Networks (GCNs) with Convolutional Neural Networks (CNNs), which captures both non-Euclidean and Euclidean features and can be trained end-to-end. The proposed CNN-GCN method achieved state-of-the-art (SOTA) performance on the task of semantic symbol spotting, and help us build a baseline network for the panoptic symbol spotting task. Our contributions are three-fold: 1) to the best of our knowledge, the presented CAD drawing dataset is the first of its kind; 2) the panoptic symbol spotting task considers the spotting of both thing instances and stuff semantic as one recognition problem; and 3) we presented a baseline solution to the panoptic symbol spotting task based on a novel CNN-GCN method, which achieved SOTA performance on semantic symbol spotting. We believe that these contributions will boost research in related areas. The dataset and code is publicly available at https://floorplancad.github.io/. Zhiwen Fan, Lingjie Zhu, Honghua Li, Xiaohao Chen, Siyu Zhu 0001, Ping Tan 0002 |
ICCV | 6 |
| 2021 | Learning Efficient Photometric Feature Transform for Multi-view StereoabstractWe present a novel framework to learn to convert the per-pixel photometric information at each view into spatially distinctive and view-invariant low-level features, which can be plugged into existing multi-view stereo pipeline for enhanced 3D reconstruction. Both the illumination conditions during acquisition and the subsequent per-pixel feature transform can be jointly optimized in a differentiable fashion. Our framework automatically adapts to and makes efficient use of the geometric information available in different forms of input data. High-quality 3D reconstructions of a variety of challenging objects are demonstrated on the data captured with an illumination multiplexing device, as well as a point light. Our results compare favorably with state-of-the-art techniques. Kaizhang Kang, Cihui Xie, Ruisheng Zhu, Xiaohe Ma, Ping Tan 0002, Hongzhi Wu, Kun Zhou 0001 |
ICCV | 5 |
| 2021 | CondLaneNet: a Top-to-down Lane Detection Framework Based on Conditional ConvolutionabstractModern deep-learning-based lane detection methods are successful in most scenarios but struggling for lane lines with complex topologies. In this work, we propose CondLaneNet, a novel top-to-down lane detection framework that detects the lane instances first and then dynamically predicts the line shape for each instance. Aiming to resolve lane instance-level discrimination problem, we introduce a conditional lane detection strategy based on conditional convolution and row-wise formulation. Further, we design the Recurrent Instance Module(RIM) to overcome the problem of detecting lane lines with complex topologies such as dense lines and fork lines. Benefit from the end-to-end pipeline which requires little post-process, our method has real-time efficiency. We extensively evaluate our method on three benchmarks of lane detection. Results show that our method achieves state-of-the-art performance on all three benchmark datasets. Moreover, our method has the coexistence of accuracy and efficiency, e.g. a 78.14 F1 score and 220 FPS on CULane. Our code is available at https://github.com/aliyun/conditional-lane-detection. Lizhe Liu, Xiaohao Chen, Siyu Zhu 0001, Ping Tan 0002 |
ICCV | 4 |
| 2021 | Interacting Two-Hand 3D Pose and Shape Reconstruction from Single Color ImageabstractIn this paper, we propose a novel deep learning framework to reconstruct 3D hand poses and shapes of two interacting hands from a single color image. Previous methods designed for single hand cannot be easily applied for the two hand scenario because of the heavy inter-hand occlusion and larger solution space. In order to address the occlusion and similar appearance between hands that may confuse the network, we design a hand pose-aware attention module to extract features associated to each individual hand respectively. We then leverage the two hand context presented in interaction to propose a context-aware cascaded refinement that improves the hand pose and shape accuracy of each hand conditioned on the context between interacting hands. Extensive experiments on the main benchmark datasets demonstrate that our method predicts accurate 3D hand pose and shape from single color image, and achieves the state-of-the-art performance. Code is available in project webpage https://baowenz.github.io/Intershape/. Baowen Zhang, Yangang Wang 0001, Xiaoming Deng 0001, Yinda Zhang 0001, Ping Tan 0002, CuiXia Ma, Hongan Wang |
ICCV | 5 |
| 2021 | Single-Shot is Enough: Panoramic Infrastructure Based Calibration of Multiple Cameras and 3D LiDARsabstractThe integration of multiple cameras and 3D Li-DARs has become basic configuration of augmented reality devices, robotics, and autonomous vehicles. The calibration of multi-modal sensors is crucial for a system to properly function, but it remains tedious and impractical for mass production. Moreover, most devices require re-calibration after usage for certain period of time. In this paper, we propose a single-shot solution for calibrating extrinsic transformations among multiple cameras and 3D LiDARs. We establish a panoramic infrastructure, in which a camera or LiDAR can be robustly localized using data from single frame. Experiments are conducted on three devices with different camera-LiDAR configurations, showing that our approach achieved comparable calibration accuracy with the state-of-the-art approaches but with much greater efficiency. Chuan Fang, Zilong Dong, Honghua Li, Siyu Zhu 0001, Ping Tan 0002 |
IROS | 6 |
| 2021 | Stereo Matching by Self-supervision of Multiscopic VisionabstractSelf-supervised learning for depth estimation possesses several advantages over supervised learning. The benefits of no need for ground-truth depth, online fine-tuning, and better generalization with unlimited data attract researchers to seek self-supervised solutions. In this work, we propose a new self-supervised framework for stereo matching utilizing multiple images captured at aligned camera positions. A cross photometric loss, an uncertainty-aware mutual-supervision loss, and a new smoothness loss are introduced to optimize the network in learning disparity maps end-to-end without ground-truth depth information. To train this framework, we build a new multiscopic dataset consisting of synthetic images rendered by 3D engines and real images captured by real cameras. After being trained with only the synthetic images, our network can perform well in unseen outdoor scenes. Our experiment shows that our model obtains better disparity maps than previous unsupervised methods on the KITTI dataset and is comparable to supervised methods when generalized to unseen data. Our source code and dataset are available at https://sites.google.com/view/multiscopic. Weihao Yuan 0001, Yazhan Zhang, Bingkun Wu, Siyu Zhu 0001, Ping Tan 0002, Michael Yu Wang, Qifeng Chen 0001 |
IROS | 5 |
| 2021 | Hand Pose Understanding With Large-Scale Photo-Realistic Rendering DatasetabstractHand pose understanding is essential to applications such as human computer interaction and augmented reality. Recently, deep learning based methods achieve great progress in this problem. However, the lack of high-quality and large-scale dataset prevents the further improvement of hand pose related tasks such as 2D/3D hand pose from color and depth from color. In this paper, we develop a large-scale and high-quality synthetic dataset, PBRHand. The dataset contains millions of photo-realistic rendered hand images and various ground truths including pose, semantic segmentation, and depth. Based on the dataset, we firstly investigate the effect of rendering methods and used databases on the performance of three hand pose related tasks: 2D/3D hand pose from color, depth from color and 3D hand pose from depth. This study provides insights that photo-realistic rendering dataset is worthy of synthesizing and shows that our new dataset can improve the performance of the state-of-the-art on these tasks. This synthetic data also enables us to explore multi-task learning, while it is expensive to have all the ground truth available on real data. Evaluations show that our approach can achieve state-of-the-art or competitive performance on several public datasets. Xiaoming Deng 0001, Yinda Zhang 0001, Yuying Zhu 0002, Dachuan Cheng, Dexin Zuo, Zhaopeng Cui, Ping Tan 0002, Liang Chang 0001, Hongan Wang |
IEEE Trans. Image Process. | 8 |
| 2021 | Weakly Supervised Learning for Single Depth-Based Hand Shape RecoveryabstractRecent emerging technologies such AR/VR and HCI are drawing high demand on more comprehensive hand shape understanding, requiring not only 3D hand skeleton pose but also hand shape geometry. In this paper, we propose a deep learning framework to produce 3D hand shape from a single depth image. To address the challenge that capturing ground truth 3D hand shape in the training dataset is non-trivial, we leverage synthetic data to construct a statistical hand shape model and adopt weak supervision from widely accessible hand skeleton pose annotation. To bridge the gap due to the different hand skeleton definitions in the existing public datasets, we propose a joint regression network for hand pose adaptation. To reconstruct the hand shape, we use Chamfer loss between the predicted hand shape and the point cloud from the input depth to learn the shape reconstruction model in a weakly-supervised manner. Experiments demonstrate that our model adapts well to the real data and produces accurate hand shapes that outperform the state-of-the-art methods both qualitatively and quantitatively. Xiaoming Deng 0001, Yuying Zhu 0002, Yinda Zhang 0001, Zhaopeng Cui, Ping Tan 0002, Wentian Qu, CuiXia Ma, Hongan Wang |
IEEE Trans. Image Process. | 5 |
| 2021 | Learning Guided Convolutional Network for Depth CompletionabstractDense depth perception is critical for autonomous driving and other robotics applications. However, modern LiDAR sensors only provide sparse depth measurement. It is thus necessary to complete the sparse LiDAR data, where a synchronized guidance RGB image is often used to facilitate this completion. Many neural networks have been designed for this task. However, they often naïvely fuse the LiDAR data and RGB image information by performing feature concatenation or element-wise addition. Inspired by the guided image filtering, we design a novel guided network to predict kernel weights from the guidance image. These predicted kernels are then applied to extract the depth image features. In this way, our network generates content-dependent and spatially-variant kernels for multi-modal feature fusion. Dynamically generated spatially-variant kernels could lead to prohibitive GPU memory consumption and computation overhead. We further design a convolution factorization to reduce computation and memory consumption. The GPU memory reduction makes it possible for feature fusion to work in multi-stage scheme. We conduct comprehensive experiments to verify our method on real-world outdoor, indoor and synthetic datasets. Our method produces strong results. It outperforms state-of-the-art methods on the NYUv2 dataset and ranks 1st on the KITTI depth completion benchmark at the time of submission. It also presents strong generalization capability under different 3D point densities, various lighting and weather conditions as well as cross-dataset evaluations. The code will be released for reproduction. Jie Tang 0015, Fei-Peng Tian, Wei Feng 0005, Jian Li 0003, Ping Tan 0002 |
IEEE Trans. Image Process. | 5 |
| 2021 | Structure-Aware Motion Deblurring Using Multi-Adversarial Optimized CycleGANabstractRecently, Convolutional Neural Networks (CNNs) have achieved great improvements in blind image motion deblurring. However, most existing image deblurring methods require a large amount of paired training data and fail to maintain satisfactory structural information, which greatly limits their application scope. In this paper, we present an unsupervised image deblurring method based on a multi-adversarial optimized cycle-consistent generative adversarial network (CycleGAN). Although original CycleGAN can handle unpaired training data well, the generated high-resolution images are probable to lose content and structure information. To solve this problem, we utilize a multi-adversarial mechanism based on CycleGAN for blind motion deblurring to generate high-resolution images iteratively. In this multi-adversarial manner, the hidden layers of the generator are gradually supervised, and the implicit refinement is carried out to generate high-resolution images continuously. Meanwhile, we also introduce the structure-aware mechanism to enhance the structure and detail retention ability of the multi-adversarial network for deblurring by taking the edge map as guidance information and adding multi-scale edge constraint functions. Our approach not only avoids the strict need for paired training data and the errors caused by blur kernel estimation, but also maintains the structural information better with multi-adversarial learning and structure-aware mechanism. Comprehensive experiments on several benchmarks have shown that our approach prevails the state-of-the-art methods for blind image motion deblurring. Jie Chen 0097, Bin Sheng 0001, Ping Li 0016, Ping Tan 0002, Tong-Yee Lee |
IEEE Trans. Image Process. | 6 |
| 2020 | Leveraging Multi-View Image Sets for Unsupervised Intrinsic Image Decomposition and Highlight SeparationabstractWe present an unsupervised approach for factorizing object appearance into highlight, shading, and albedo layers, trained by multi-view real images. To do so, we construct a multi-view dataset by collecting numerous customer product photos online, which exhibit large illumination variations that make them suitable for training of reflectance separation and can facilitate object-level decomposition. The main contribution of our approach is a proposed image representation based on local color distributions that allows training to be insensitive to the local misalignments of multi-view images. In addition, we present a new guidance cue for unsupervised training that exploits synergy between highlight separation and intrinsic image decomposition. Over a broad range of objects, our technique is shown to yield state-of-the-art results for both of these tasks. Renjiao Yi, Ping Tan 0002, Stephen Lin 0001 |
AAAI | 2 |
| 2020 | Deep Facial Non-Rigid Multi-View StereoabstractWe present a method for 3D face reconstruction from multi-view images with different expressions. We formulate this problem from the perspective of non-rigid multi-view stereo (NRMVS). Unlike previous learning-based methods, which often regress the face shape directly, our method optimizes the 3D face shape by explicitly enforcing multi-view appearance consistency, which is known to be effective in recovering shape details according to conventional multi-view stereo methods. Furthermore, by estimating face shape through optimization based on multi-view consistency, our method can potentially have better generalization to unseen data. However, this optimization is challenging since each input image has a different expression. We facilitate it with a CNN network that learns to regularize the non-rigid 3D face according to the input image and preliminary optimization results. Extensive experiments show that our method achieves the state-of-the-art performance on various datasets and generalizes well to in-the-wild data. Ziqian Bai, Zhaopeng Cui, Jamal Ahmed Rahim, Xiaoming Liu 0002, Ping Tan 0002 |
CVPR | 5 |
| 2020 | Cascade Cost Volume for High-Resolution Multi-View Stereo and Stereo MatchingabstractThe deep multi-view stereo (MVS) and stereo matching approaches generally construct 3D cost volumes to regularize and regress the output depth or disparity. These methods are limited when high-resolution outputs are needed since the memory and time costs grow cubically as the volume resolution increases. In this paper, we propose a both memory and time efficient cost volume formulation that is complementary to existing multi-view stereo and stereo matching approaches based on 3D cost volumes. First, the proposed cost volume is built upon a standard feature pyramid encoding geometry and context at gradually finer scales. Then, we can narrow the depth (or disparity) range of each stage by the depth (or disparity) map from the previous stage. With gradually higher cost volume resolution and adaptive adjustment of depth (or disparity) intervals, the output is recovered in a coarser to fine manner. We apply the cascade cost volume to the representative MVS-Net, and obtain a 35.6% improvement on DTU benchmark (1st place), with 50.6% and 59.3% reduction in GPU memory and run-time. It is also the state-of-the-art learning-based method on Tanks and Temples benchmark. The statistics of accuracy, run-time and GPU memory on other representative stereo CNNs also validate the effectiveness of our proposed method. Our source code is available at https://github.com/alibaba/cascade-stereo. Xiaodong Gu 0004, Zhiwen Fan, Siyu Zhu 0001, Zuozhuo Dai, Feitong Tan, Ping Tan 0002 |
CVPR | 6 |
| 2020 | End-to-End Learning Local Multi-View Descriptors for 3D Point CloudsabstractIn this work, we propose an end-to-end framework to learn local multi-view descriptors for 3D point clouds. To adopt a similar multi-view representation, existing studies use hand-crafted viewpoints for rendering in a preprocessing stage, which is detached from the subsequent descriptor learning stage. In our framework, we integrate the multi-view rendering into neural networks by using a differentiable renderer, which allows the viewpoints to be optimizable parameters for capturing more informative local context of interest points. To obtain discriminative descriptors, we also design a soft-view pooling module to attentively fuse convolutional features across views. Extensive experiments on existing 3D registration benchmarks show that our method outperforms existing local descriptors both quantitatively and qualitatively. Lei Li 0038, Siyu Zhu 0001, Hongbo Fu 0001, Ping Tan 0002, Chiew-Lan Tai |
CVPR | 4 |
| 2020 | Self-Supervised Human Depth Estimation From Monocular VideosabstractPrevious methods on estimating detailed human depth often require supervised training with ‘ground truth’ depth data. This paper presents a self-supervised method that can be trained on YouTube videos without known depth, which makes training data collection simple and improves the generalization of the learned network. The self-supervised learning is achieved by minimizing a photo-consistency loss, which is evaluated between a video frame and its neighboring frames warped according to the estimated depth and the 3D non-rigid motion of the human body. To solve this non-rigid motion, we first estimate a rough SMPL model at each video frame and compute the non-rigid body motion accordingly, which enables self-supervised learning on estimating the shape details. Experiments demonstrate that our method enjoys better generalization, and performs much better on data in the wild. Feitong Tan, Hao Zhu 0004, Zhaopeng Cui, Siyu Zhu 0001, Marc Pollefeys, Ping Tan 0002 |
CVPR | 6 |
| 2020 | LSM: Learning Subspace Minimization for Low-Level VisionabstractWe study the energy minimization problem in low-level vision tasks from a novel perspective. We replace the heuristic regularization term with a data-driven learnable subspace constraint, and preserve the data term to exploit domain knowledge derived from the first principles of a task. This learning subspace minimization (LSM) framework unifies the network structures and the parameters for many different low-level vision tasks, which allows us to train a single network for multiple tasks simultaneously with shared parameters, and even generalizes the trained network to an unseen task as long as the data term can be formulated. We validate our LSM frame on four low-level tasks including edge detection, interactive segmentation, stereo matching, and optical flow, and validate the network on various datasets. The experiments demonstrate that the proposed LSM generates state-of-the-art results with smaller model size, faster training convergence, and real-time inference. Chengzhou Tang, Lu Yuan 0001, Ping Tan 0002 |
CVPR | 3 |
| 2020 | Channel Equilibrium Networks for Learning Deep RepresentationabstractConvolutional Neural Networks (CNNs) are typically constructed by stacking multiple building blocks, each of which contains a normalization layer such as batch normalization (BN) and a rectified linear function such as ReLU. However, this work shows that the combination of normalization and rectified linear function leads to inhibited channels, which have small magnitude and contribute little to the learned feature representation, impeding the generalization ability of CNNs. Unlike prior arts that simply removed the inhibited channels, we propose to “wake them up” during training by designing a novel neural building block, termed Channel Equilibrium (CE) block, which enables channels at the same layer to contribute equally to the learned representation. We show that CE is able to prevent inhibited channels both empirically and theoretically. CE has several appealing benefits. (1) It can be integrated into many advanced CNN architectures such as ResNet and MobileNet, outperforming their original networks. (2) CE has an interesting connection with the Nash Equilibrium, a well-known solution of a non-cooperative game. (3) Extensive experiments show that CE achieves state-of-the-art performance on various challenging benchmarks such as ImageNet and COCO. Wenqi Shao, Shitao Tang, Xingang Pan, Ping Tan 0002, Xiaogang Wang 0001, Ping Luo 0002 |
ICML | 4 |
| 2020 | Depth-Aware Motion Deblurring Using Loopy Belief PropagationabstractMost motion-blurred images captured in the real world have spatially-varying point-spread functions, and some are caused by different positions and depth values, which cannot be handled by most state-of-the-art deblurring methods based on deconvolution. To overcome this problem, we propose a depth-aware motion blur model that treats a blurred image as an integration of a sequence of clear images. To restore the clear latent image, we extend the Richardson-Lucy method to incorporate our blur model with a given depth image. The empty holes in the depth image, caused by occlusion or device limitations, are fixed by PatchMatch-based depth filling. We regard the depth image as a Markov random field and select candidate labels by using belief propagation to set and smooth depth values for empty areas. Deblurring and depth filling are performed iteratively to refine the results. Our method can also be applied to real-world images with the assistance of motion estimation. The deblurring process is shown to be convergent; moreover, the number of iterations and the level of noise amplification are acceptable. The experimental results show that our method can not only handle depth-variant motion blur but also refine depth images. Bin Sheng 0001, Ping Li 0016, Xiaoxin Fang, Ping Tan 0002, Enhua Wu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Multi-View Photometric Stereo: A Robust Solution and Benchmark Dataset for Spatially Varying Isotropic MaterialsabstractWe present a method to capture both 3D shape and spatially varying reflectance with a multi-view photometric stereo (MVPS) technique that works for general isotropic materials. Our algorithm is suitable for perspective cameras and nearby point light sources. Our data capture setup is simple, which consists of only a digital camera, some LED lights, and an optional automatic turntable. From a single viewpoint, we use a set of photometric stereo images to identify surface points with the same distance to the camera. We collect this information from multiple viewpoints and combine it with structure-from-motion to obtain a precise reconstruction of the complete 3D shape. The spatially varying isotropic bidirectional reflectance distribution function (BRDF) is captured by simultaneously inferring a set of basis BRDFs and their mixing weights at each surface point. In experiments, we demonstrate our algorithm with two different setups: a studio setup for highest precision and a desktop setup for best usability. According to our experiments, under the studio setting, the captured shapes are accurate to 0.5 millimeters and the captured reflectance has a relative root-mean-square error (RMSE) of 9%. We also quantitatively evaluate state-of-the-art MVPS on a newly collected benchmark dataset, which is publicly available for inspiring future research. Min Li 0049, Zhenglong Zhou, Boxin Shi, Changyu Diao, Ping Tan 0002 |
IEEE Trans. Image Process. | 6 |
| 2020 | Micrography QR CodesabstractThis paper presents a novel algorithm to generate micrography QR codes, a novel machine-readable graphic generated by embedding a QR code within a micrography image. The unique structure of micrography makes it incompatible with existing methods used to combine QR codes with natural or halftone images. We exploited the high-frequency nature of micrography in the design of a novel deformation model that enables the skillful warping of individual letters and adjustment of font weights to enable the embedding of a QR code within a micrography. The entire process is supervised by a set of visual quality metrics tailored specifically for micrography, in conjunction with a novel QR code quality measure aimed at striking a balance between visual fidelity and decoding robustness. The proposed QR code quality measure is based on probabilistic models learned from decoding experiments using popular decoders with synthetic QR codes to capture the various forms of distortion that result from image embedding. Experiment results demonstrate the efficacy of the proposed method in generating micrography QR codes of high quality from a wide variety of inputs. The ability to embed QR codes with multiple scales makes it possible to produce a wide range of diverse designs. Experiments and user studies were conducted to evaluate the proposed method from a qualitative as well as quantitative perspective. Shih-Hsuan Hung, Chih-Yuan Yao, Yu-Jen Fang, Ping Tan 0002, Ruen-Rone Lee, Alla Sheffer, Hung-Kuo Chu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2020 | Intrinsic Image Decomposition with Step and Drift Shading SeparationabstractDecomposing an image into the shading and reflectance layers remains challenging due to its severely under-constrained nature. We present an approach based on illumination decomposition that recovers the intrinsic images without additional information, e.g., depth or user interaction. Our approach is based on the rationale that the shading component contains the step and drift channels simultaneously. We decompose the illumination into two channels: the step shading, corresponding to the sharp shading changes due to cast shadow or abrupt shape changes; the drift shading, accounting for the smooth shading variations due to gradual illumination changes or slow shape changes. Due to such transformation of turning the conventional assumption that shading has smoothness as reasonable prior, our model has the advantages in handling real images, especially with the cast shadows or strong shape edges. We also apply a much stricter edge classifier along with a reinforcement process to enhance our method. We formulate the problem using a two-parameter energy function and split it into two energy functions corresponding to the reflectance and step shading. Experiments on the MIT dataset, the IIW dataset and the MPI Sintel dataset have shown the success of our approach over the state-of-the-art methods. Bin Sheng 0001, Ping Li 0016, Yuxi Jin, Ping Tan 0002, Tong-Yee Lee |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2019 | Batch DropBlock Network for Person Re-Identification and BeyondabstractSince the person re-identification task often suffers from the problem of pose changes and occlusions, some attentive local features are often suppressed when training CNNs. In this paper, we propose the Batch DropBlock (BDB) Network which is a two branch network composed of a conventional ResNet-50 as the global branch and a feature dropping branch.The global branch encodes the global salient representations.Meanwhile, the feature dropping branch consists of an attentive feature learning module called Batch DropBlock, which randomly drops the same region of all input feature maps in a batch to reinforce the attentive feature learning of local regions.The network then concatenates features from both branches and provides a more comprehensive and spatially distributed feature representation. Albeit simple, our method achieves state-of-the-art on person re-identification and it is also applicable to general metric learning tasks. For instance, we achieve 76.4% Rank-1 accuracy on the CUHK03-Detect dataset and 83.0% Recall-1 score on the Stanford Online Products dataset, outperforming the existed works by a large margin (more than 6%). Zuozhuo Dai, Mingqiang Chen, Xiaodong Gu 0004, Siyu Zhu 0001, Ping Tan 0002 |
ICCV | 5 |
| 2019 | A Neural Network for Detailed Human Depth Estimation From a Single ImageabstractThis paper presents a neural network to estimate a detailed depth map of the foreground human in a single RGB image. The result captures geometry details such as cloth wrinkles, which are important in visualization applications. To achieve this goal, we separate the depth map into a smooth base shape and a residual detail shape and design a network with two branches to regress them respectively. We design a training strategy to ensure both base and detail shapes can be faithfully learned by the corresponding network branches. Furthermore, we introduce a novel network layer to fuse a rough depth map and surface normals to further improve the final result. Quantitative comparison with fused `ground truth' captured by real depth cameras and qualitative examples on unconstrained Internet images demonstrate the strength of the proposed method. Sicong Tang, Feitong Tan, Kelvin Cheng 0003, Siyu Zhu 0001, Ping Tan 0002 |
ICCV | 6 |
| 2019 | Learned Map Prediction for Enhanced Mobile Robot ExplorationabstractWe demonstrate an autonomous ground robot capable of exploring unknown indoor environments for reconstructing their 2D maps. This problem has been traditionally tackled by geometric heuristics and information theory. More recently, deep learning and reinforcement learning based approaches have been proposed to learn exploration behavior in an end-to-end manner. We present a method that combines the strengths of these different approaches. Specifically, we employ a state-of-the-art generative neural network to predict unknown regions of a partially explored map, and use the prediction to enhance the exploration in an information-theoretic manner. We evaluate our system in simulation using floor plans of real buildings. We also present comparisons with traditional methods which demonstrate the advantage of our method in terms of exploration efficiency. We retain an advantage over end-to-end learned exploration methods in that the robot's behavior is easily explicable in terms of the predicted map. Rakesh Shrestha, Fei-Peng Tian, Wei Feng 0005, Ping Tan 0002, Richard Vaughan 0001 |
ICRA | 4 |
| 2019 | A Benchmark Dataset and Evaluation for Non-Lambertian and Uncalibrated Photometric StereoabstractClassic photometric stereo is often extended to deal with real-world materials and work with unknown lighting conditions for practicability. To quantitatively evaluate non-Lambertian and uncalibrated photometric stereo, a photometric stereo image dataset containing objects of various shapes with complex reflectance properties and high-quality ground truth normals is still missing. In this paper, we introduce the 'DiLiGenT' dataset with calibrated Directional Lightings, objects of General reflectance with different shininess, and 'ground Truth' normals from high-precision laser scanning. We use our dataset to quantitatively evaluate state-of-the-art photometric stereo methods for general materials and unknown lighting conditions, selected from a newly proposed photometric stereo taxonomy emphasizing non-Lambertian and uncalibrated methods. The dataset and evaluation results are made publicly available, and we hope it can serve as a benchmark platform that inspires future research. Boxin Shi, Zhipeng Mo, Dinglong Duan, Sai-Kit Yeung, Ping Tan 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2019 | Joint Stabilization and Direction of 360° VideosabstractThree-hundred-sixty-degree (360°) video provides an immersive experience for viewers, allowing them to freely explore the world by turning their head. However, creating high-quality 360° video content can be challenging, as viewers may miss important events by looking in the wrong direction, or they may see things that ruin the immersion, such as stitching artifacts and the film crew. We take advantage of the fact that not all directions are equally likely to be observed; most viewers are more likely to see content located at “true north,” i.e., in front of them, due to ergonomic constraints. We therefore propose 360° video direction, where the video is jointly optimized to orient important events to the front of the viewer and visual clutter behind them, while producing smooth camera motion. Unlike traditional video, viewers can still explore the space as desired, but with the knowledge that the most important content is likely to be in front of them. Constraints can be user guided, either added directly on the equirectangular projection or by recording “guidance” viewing directions while watching the video in a VR headset or automatically computed, such as via visual saliency or forward-motion direction. To accomplish this, we propose a new motion estimation technique specifically designed for 360° video that outperforms the commonly used five-point algorithm on wide-angle video. We additionally formulate the direction problem as an optimization where a novel parametrization of spherical warping allows us to correct for some degree of parallax effects. We compare our approach to recent methods that address stabilization-only and converting 360° video to narrow field-of-view video. Our pipeline can also enable the viewing of wide-angle non-360° footage in a spherical 360° space, giving an immersive “virtual cinema” experience for a wide range of existing content filmed with first-person cameras. Chengzhou Tang, Oliver Wang, Feng Liu 0015, Ping Tan 0002 |
ACM Trans. Graph. | 4 |
| 2018 | Polarimetric Dense Monocular SLAMabstractThis paper presents a novel polarimetric dense monocular SLAM (PDMS) algorithm based on a polarization camera. The algorithm exploits both photometric and polarimetric light information to produce more accurate and complete geometry. The polarimetric information allows us to recover the azimuth angle of surface normals from each video frame to facilitate dense reconstruction, especially at textureless or specular regions. There are two challenges in our approach: 1) surface azimuth angles from the polarization camera are very noisy; and 2) we need a near real-time solution for SLAM. Previous successful methods on polarimetric multi-view stereo are offline and require manually pre-segmented object masks to suppress the effects of erroneous angle information along boundaries. Our fully automatic approach efficiently iterates azimuth-based depth propagations, two-view depth consistency check, and depth optimization to produce a depthmap in real-time, where all the algorithmic steps are carefully designed to enable a GPU implementation. To our knowledge, this paper is the first to propose a photometric and polarimetric method for dense SLAM. We have qualitatively and quantitatively evaluated our algorithm against a few of competing methods, demonstrating the superior performance on various indoor and outdoor scenes. Luwei Yang, Feitong Tan, Ao Li 0009, Zhaopeng Cui, Yasutaka Furukawa, Ping Tan 0002 |
CVPR | 6 |
| 2018 | Very Large-Scale Global SfM by Distributed Motion AveragingabstractGlobal Structure-from-Motion (SfM) techniques have demonstrated superior efficiency and accuracy than the conventional incremental approach in many recent studies. This work proposes a divide-and-conquer framework to solve very large global SfM at the scale of millions of images. Specifically, we first divide all images into multiple partitions that preserve strong data association for well-posed and parallel local motion averaging. Then, we solve a global motion averaging that determines cameras at partition boundaries and a similarity transformation per partition to register all cameras in a single coordinate frame. Finally, local and global motion averaging are iterated until convergence. Since local camera poses are fixed during the global motion average, we can avoid caching the whole reconstruction in memory at once. This distributed framework significantly enhances the efficiency and robustness of large-scale motion averaging. Siyu Zhu 0001, Lei Zhou 0011, Tianwei Shen, Tian Fang, Ping Tan 0002, Long Quan |
CVPR | 6 |
| 2018 | Faces as Lighting Probes via Unsupervised Deep Highlight Extraction
Renjiao Yi, Chenyang Zhu 0002, Ping Tan 0002, Stephen Lin 0001 |
ECCV (9) | 3 |
| 2018 | Active Image-Based Modeling with a Toy DroneabstractImage-based modeling techniques [1]-[3] can now generate photo-realistic 3D models from images. But it is up to users to provide high quality images with good coverage and view overlap, which makes the data capturing process tedious and time consuming. We seek to automate data capturing for image-based modeling. The core of our system is an iterative linear method to solve the multi-view stereo (MVS) problem quickly and plan the Next-Best-View (NBV) effectively. Our fast MVS algorithm enables online model reconstruction and quality assessment to determine the NBVs on the fly. We test our system with a toy unmanned aerial vehicle (UAV) in simulated, indoor and outdoor experiments. Results show that our system improves the efficiency of data acquisition and ensures the completeness of the final model. Rui Huang 0001, Danping Zou, Richard Vaughan 0001, Ping Tan 0002 |
ICRA | 4 |
| 2018 | Active Recurrence of Lighting Condition for Fine-Grained Change DetectionabstractThis paper addresses active lighting recurrence (ALR), a new problem that actively relocalizes a light source to physically reproduce the lighting condition for a same scene from single reference image. ALR is of great importance for fine-grained visual monitoring and change detection, because some phenomena or minute changes can only be clearly observed under particular lighting conditions. Hence, effective ALR should be able to online navigate a light source toward the target pose, which is challenging due to the complexity and diversity of real-world lighting \& imaging processes. We propose to use the simple parallel lighting as an analogy model and based on Lambertian law to compose an instant navigation ball for this purpose. We theoretically prove the feasibility of this ALR strategy for realistic near point light sources and its invariance to the ambiguity of normal \& lighting decomposition. Extensive quantitative experiments and challenging real-world tasks on fine-grained change monitoring of cultural heritages verify the effectiveness of our approach. We also validate its generality to non-Lambertian scenes. Qian Zhang 0051, Wei Feng 0005, Fei-Peng Tian, Ping Tan 0002 |
IJCAI | 5 |
| 2018 | Joint Hand Detection and Rotation Estimation Using CNNabstractHand detection is essential for many hand related tasks, e.g., recovering hand pose and understanding gesture. However, hand detection in uncontrolled environments is challenging due to the flexibility of wrist joint and cluttered background. We propose a convolutional neural network (CNN), which formulates in-plane rotation explicitly to solve hand detection and rotation estimation jointly. Our network architecture adopts the backbone of faster R-CNN to generate rectangular region proposals and extract local features. The rotation network takes the feature as input and estimates an in-plane rotation which manages to align the hand, if any in the proposal, to the upward direction. A derotation layer is then designed to explicitly rotate the local spatial feature map according to the rotation network and feed aligned feature map for detection. Experiments show that our method outperforms the state-of-the-art detection models on widely-used benchmarks, such as Oxford and Egohands database. Further analysis show that rotation estimation and classification can mutually benefit each other. Xiaoming Deng 0001, Yinda Zhang 0001, Shuo Yang 0002, Ping Tan 0002, Liang Chang 0001, Hongan Wang |
IEEE Trans. Image Process. | 4 |
| 2018 | Outdoor Markerless Motion Capture with Sparse Handheld Video CamerasabstractWe present a method for outdoor markerless motion capture with sparse handheld video cameras. In the simplest setting, it only involves two mobile phone cameras following the character. This setup can maximize the flexibilities of data capture and broaden the applications of motion capture. To solve the character pose under such challenge settings, we exploit the generative motion capture methods and propose a novel model-view consistency that considers both foreground and background in the tracking stage. The background is modeled as a deformable 2D grid, which allows us to compute the background-view consistency for sparse moving cameras. The 3D character pose is tracked with a global-local optimization through minimizing our consistency cost. A novel motion regularizer is also proposed in the optimization to constrain the solution pose space. The whole process of the proposed method is simple as frame by frame video segmentation is not required. Our method outperforms several alternative methods on various examples demonstrated in the paper. Yangang Wang 0001, Yebin Liu, Xin Tong 0001, Qionghai Dai, Ping Tan 0002 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2017 | Polarimetric Multi-view StereoabstractMulti-view stereo relies on feature correspondences for 3D reconstruction, and thus is fundamentally flawed in dealing with featureless scenes. In this paper, we propose polarimetric multi-view stereo, which combines per-pixel photometric information from polarization with epipolar constraints from multiple views for 3D reconstruction. Polarization reveals surface normal information, and is thus helpful to propagate depth to featureless regions. Polarimetric multi-view stereo is completely passive and can be applied outdoors in uncontrolled illumination, since the data capture can be done simply with either a polarizer or a polarization camera. Unlike previous work on shape-from-polarization which is limited to either diffuse polarization or specular polarization only, we propose a novel polarization imaging model that can handle real-world objects with mixed polarization. We prove there are exactly two types of ambiguities on estimating surface azimuth angles from polarization, and we resolve them with graph optimization and iso-depth contour tracing. This step significantly improves the initial depth map estimate, which are later fused together for complete 3D reconstruction. Extensive experimental results demonstrate high-quality 3D reconstruction and better performance than state-of-the-art multi-view stereo methods, especially on featureless 3D objects, such as ceramic tiles, office room with white walls, and highly reflective cars in the outdoors. Zhaopeng Cui, Jinwei Gu, Boxin Shi, Ping Tan 0002, Jan Kautz |
CVPR | 4 |
| 2017 | DualGAN: Unsupervised Dual Learning for Image-to-Image TranslationabstractConditional Generative Adversarial Networks (GANs) for cross-domain image-to-image translation have made much progress recently [7, 8, 21, 12, 4, 18]. Depending on the task complexity, thousands to millions of labeled image pairs are needed to train a conditional GAN. However, human labeling is expensive, even impractical, and large quantities of data may not always be available. Inspired by dual learning from natural language translation [23], we develop a novel dual-GAN mechanism, which enables image translators to be trained from two sets of unlabeled images from two domains. In our architecture, the primal GAN learns to translate images from domain U to those in domain V, while the dual GAN learns to invert the task. The closed loop made by the primal and dual tasks allows images from either domain to be translated and then reconstructed. Hence a loss function that accounts for the reconstruction error of images can be used to train the translators. Experiments on multiple image translation tasks with unlabeled data show considerable performance gain of DualGAN over a single GAN. For some tasks, DualGAN can even achieve comparable or slightly better results than conditional GAN trained on fully labeled data. Zili Yi, Hao (Richard) Zhang, Ping Tan 0002, Minglun Gong |
ICCV | 3 |
| 2017 | Time slice video synthesis by robust video alignmentabstractTime slice photography is a popular effect that visualizes the passing of time by aligning and stitching multiple images capturing the same scene at different times together into a single image. Extending this effect to video is a difficult problem, and one where existing solutions have only had limited success. In this paper, we propose an easy-to-use and robust system for creating time slice videos from a wide variety of consumer videos. The main technical challenge we address is how to align videos taken at different times with substantially different appearances, in the presence of moving objects and moving cameras with slightly different trajectories. To achieve a temporally stable alignment, we perform a mixed 2D-3D alignment, where a rough 3D reconstruction is used to generate sparse constraints that are integrated into a pixelwise 2D registration. We apply our method to a number of challenging scenarios, and show that we can achieve a higher quality registration than prior work. We propose a 3D user interface that allows the user to easily specify how multiple videos should be composited in space and time. Finally, we show that our alignment method can be applied in more general video editing and compositing tasks, such as object removal. Zhaopeng Cui, Oliver Wang, Ping Tan 0002, Jue Wang 0001 |
ACM Trans. Graph. | 3 |
| 2017 | Learning to group discrete graphical patternsabstractWe introduce a deep learning approach for grouping discrete patterns common in graphical designs. Our approach is based on a convolutional neural network architecture that learns a grouping measure defined over a pair of pattern elements. Motivated by perceptual grouping principles, the key feature of our network is the encoding of element shape, context, symmetries, and structural arrangements. These element properties are all jointly considered and appropriately weighted in our grouping measure. To better align our measure with human perceptions for grouping, we train our network on a large, human-annotated dataset of pattern groupings consisting of patterns at varying granularity levels, with rich element relations and varieties, and tempered with noise and other data imperfections. Experimental results demonstrate that our deep-learned measure leads to robust grouping results. Zhaoliang Lun, Changqing Zou, Evangelos Kalogerakis, Ping Tan 0002, Marie-Paule Cani, Hao (Richard) Zhang |
ACM Trans. Graph. | 5 |
| 2016 | A Benchmark Dataset and Evaluation for Non-Lambertian and Uncalibrated Photometric StereoabstractRecent progress on photometric stereo extends the technique to deal with general materials and unknown illumination conditions. However, due to the lack of suitable benchmark data with ground truth shapes (normals), quantitative comparison and evaluation is difficult to achieve. In this paper, we first survey and categorize existing methods using a photometric stereo taxonomy emphasizing on non-Lambertian and uncalibrated methods. We then introduce the 'DiLiGenT' photometric stereo image dataset with calibrated Directional Lightings, objects of General reflectance, and 'ground Truth' shapes (normals). Based on our dataset, we quantitatively evaluate state-of-the-art photometric stereo methods for general non-Lambertian materials and unknown lightings to analyze their strengths and limitations. Boxin Shi, Zhipeng Mo, Dinglong Duan, Sai-Kit Yeung, Ping Tan 0002 |
CVPR | 6 |
| 2016 | Automatic Fence Segmentation in Videos of Dynamic ScenesabstractWe present a fully automatic approach to detect and segment fence-like occluders from a video clip. Unlike previous approaches that usually assume either static scenes or cameras, our method is capable of handling both dynamic scenes and moving cameras. Under a bottom-up framework, it first clusters pixels into coherent groups using color and motion features. These pixel groups are then analyzed in a fully connected graph, and labeled as either fence or non-fence using graph-cut optimization. Finally, we solve a dense Conditional Random Filed (CRF) constructed from multiple frames to enhance both spatial accuracy and temporal coherence of the segmentation. Once segmented, one can use existing hole-filling methods to generate a fencefree output. Extensive evaluation suggests that our method outperforms previous automatic and interactive approaches on complex examples captured by mobile devices. Renjiao Yi, Jue Wang 0001, Ping Tan 0002 |
CVPR | 3 |
| 2016 | MeshFlow: Minimum Latency Online Video Stabilization
Shuaicheng Liu, Ping Tan 0002, Lu Yuan 0001, Jian Sun 0001, Bing Zeng 0001 |
ECCV (6) | 2 |
| 2016 | Legible compact calligramsabstractA calligram is an arrangement of words or letters that creates a visual image, and a compact calligram fits one word into a 2D shape. We introduce a fully automatic method for the generation of legible compact calligrams which provides a balance between conveying the input shape, legibility, and aesthetics. Our method has three key elements: a path generation step which computes a global layout path suitable for embedding the input word; an alignment step to place the letters so as to achieve feature alignment between letter and shape protrusions while maintaining word legibility; and a final deformation step which deforms the letters to fit the shape while balancing fit against letter legibility. As letter legibility is critical to the quality of compact calligrams, we conduct a large-scale crowd-sourced study on the impact of different letter deformations on legibility and use the results to train a letter legibility measure which guides the letter deformation. We show automatically generated calligrams on an extensive set of word-image combinations. The legibility and overall quality of the calligrams are evaluated and compared, via user studies, to those produced by human creators, including a professional artist, and existing works. Changqing Zou, Junjie Cao 0001, Warunika Ranaweera, Ibraheem Alhashim, Ping Tan 0002, Alla Sheffer, Hao (Richard) Zhang |
ACM Trans. Graph. | 5 |
| 2015 | Linear Global Translation Estimation with Feature TracksabstractGlobal structure-from-motion (SfM) algorithms register all cameras simultaneously, which are potentially more efficient and less prone to drifting than incremental SfM methods. Global SfM methods often solve the camera orientations and positions separately. This paper focuses on the problem of global position (i.e. translation) estimation. Essential matrix based global translation estimation methods (e.g. [1]) usually degenerate at collinear camera motion because the translation scale is not determined by an essential matrix. Trifocal tensor based methods (e.g. [3]) usually rely on a strongly connected camera-triplet graph, where two triplets are connected by their common edge. The 3D reconstruction will distort or break into disconnected components when such strong association among images does not exist. The recent 1DSfM method [4] designs a smart filter to discard outlier essential matrices and solves scene points and cameras together by enforcing orientation consistency. However, this method requires abundant association between input images, e.g.∼O(n2) essential matrices for n cameras, which is more suitable for Internet images and often fails on sequentially captured data. The data association problem of [4] and [3] is exemplified in Figure 1. The Street example on the top is a sequential data where each image is only matched upto 4 neighbors. 1DSfM fails on this example due to insufficient image association. In the Seville example on the bottom, those Internet images are mostly captured from two viewpoints (see the two representative sample images) with weak affinity between images at different viewpoints. This weak data association causes seriously distorted reconstruction for the triplet-based method in [3]. This paper introduces a direct linear algorithm to address the presented challenges. It avoids degeneracy at collinear motion and deals with weakly associated data. Our method capitalizes on constraints from essential matrices and feature tracks. As shown in Figure 2 (a), the location of a scene point p can be computed as the middle point of the mutual perpendicular line segment AB of the two rays passing through p’s image projections: Zhaopeng Cui, Nianjuan Jiang, Chengzhou Tang, Ping Tan 0002 |
BMVC | 4 |
| 2015 | Global Structure-from-Motion by Similarity AveragingabstractGlobal structure-from-motion (SfM) methods solve all cameras simultaneously from all available relative motions. It has better potential in both reconstruction accuracy and computation efficiency than incremental methods. However, global SfM is challenging, mainly because of two reasons. Firstly, translation averaging is difficult, since an essential matrix only tells the direction of relative translation. Secondly, it is also hard to filter out bad essential matrices due to feature matching failures. We propose to compute a sparse depth image at each camera to solve both problems. Depth images help to upgrade an essential matrix to a similarity transformation, which can determine the scale of relative translation. Thus, camera registration is formulated as a well-posed similarity averaging problem. Depth images also make the filtering of essential matrices simple and effective. In this way, translation averaging can be solved robustly in two convex L1 optimization problems, which reach the global optimum rapidly. We demonstrate this method in various examples including sequential data, Internet data, and ambiguous data with repetitive scene structures. Zhaopeng Cui, Ping Tan 0002 |
ICCV | 2 |
| 2015 | Where2Stand: A Human Position Recommendation System for Souvenir PhotographyabstractPeople often take photographs at tourist sites and these pictures usually have two main elements: a person in the foreground and scenery in the background. This type of “souvenir photo” is one of the most common photos clicked by tourists. Although algorithms that aid a user-photographer in taking a well-composed picture of a scene exist [Ni et al. 2013], few studies have addressed the issue of properly positioning human subjects in photographs. In photography, the common guidelines of composing portrait images exist. However, these rules usually do not consider the background scene. Therefore, in this article, we investigate human-scenery positional relationships and construct a photographic assistance system to optimize the position of human subjects in a given background scene, thereby assisting the user in capturing high-quality souvenir photos. We collect thousands of well-composed portrait photographs to learn human-scenery aesthetic composition rules. In addition, we define a set of negative rules to exclude undesirable compositions. Recommendation results are achieved by combining the first learned positive rule with our proposed negative rules. We implement the proposed system on an Android platform in a smartphone. The system demonstrates its efficacy by producing well-composed souvenir photos. Yinting Wang, Mingli Song, Dacheng Tao, Yong Rui, Jiajun Bu, Ah Chung Tsoi, Shaojie Zhuo, Ping Tan 0002 |
ACM Trans. Intell. Syst. Technol. | 8 |
| 2014 | Photometric Stereo Using Internet ImagesabstractPhotometric stereo using unorganized Internet images is very challenging, because the input images are captured under unknown general illuminations, with uncontrolled cameras. We propose to solve this difficult problem by a simple yet effective approach that makes use of a coarse shape prior. The shape prior is obtained from multi-view stereo and will be useful in twofold: resolving the shape-light ambiguity in uncalibrated photometric stereo and guiding the estimated normals to produce the high quality 3D surface. By assuming the surface albedo is not highly contrasted, we also propose a novel linear approximation of the nonlinear camera responses with our normal estimation algorithm. We evaluate our method using synthetic data and demonstrate the surface improvement on real data over multi-view stereo results. Boxin Shi, Kenji Inose, Yasuyuki Matsushita, Ping Tan 0002, Sai-Kit Yeung, Katsushi Ikeuchi |
3DV | 4 |
| 2014 | SteadyFlow: Spatially Smooth Optical Flow for Video StabilizationabstractWe propose a novel motion model, SteadyFlow, to represent the motion between neighboring video frames for stabilization. A SteadyFlow is a specific optical flow by enforcing strong spatial coherence, such that smoothing feature trajectories can be replaced by smoothing pixel profiles, which are motion vectors collected at the same pixel location in the SteadyFlow over time. In this way, we can avoid brittle feature tracking in a video stabilization system. Besides, SteadyFlow is a more general 2D motion model which can deal with spatially-variant motion. We initialize the SteadyFlow by optical flow and then discard discontinuous motions by a spatial-temporal analysis and fill in missing regions by motion completion. Our experiments demonstrate the effectiveness of our stabilization on real-world challenging videos. Shuaicheng Liu, Lu Yuan 0001, Ping Tan 0002, Jian Sun 0001 |
CVPR | 3 |
| 2014 | A New Perspective on Material Classification and Ink IdentificationabstractThe surface bi-directional reflectance distribution function (BRDF) can be used to distinguish different materials. The BRDFs of many real materials are near isotropic and can be approximated well by a 2D function. When the camera principal axis is coincident with the surface normal of the material sample, the captured BRDF slice is nearly 1D, which suffers from significant information loss. Thus, improvement in classification performance can be achieved by simply setting the camera at a slanted view to capture a larger portion of the BRDF domain. We further use a handheld flashlight camera to capture a 1D BRDF slice for material classification. This 1D slice captures important reflectance properties such as specular reflection and retro-reflectance. We apply these results on ink classification, which can be used in forensics and analyzing historical manuscripts. For the first time, we show that most of the inks on the market can be well distinguished by their reflectance properties. Rakesh Shiradkar, Li Shen 0003, George V. Landon, Sim Heng Ong, Ping Tan 0002 |
CVPR | 5 |
| 2014 | PanoContext: A Whole-Room 3D Context Model for Panoramic Scene Understanding
Yinda Zhang 0001, Shuran Song, Ping Tan 0002, Jianxiong Xiao |
ECCV (6) | 3 |
| 2014 | Auto-calibrating photometric stereo using ring light constraints
Rakesh Shiradkar, Ping Tan 0002, Sim Heng Ong |
Mach. Vis. Appl. | 2 |
| 2014 | Bi-Polynomial Modeling of Low-Frequency ReflectancesabstractWe present a bi-polynomial reflectance model that can precisely represent the low-frequency component of reflectance. Most existing reflectance models aim at accurately representing the complete reflectance domain for photo-realistic rendering purposes. In contrast, our bi-polynomial model is developed for the purpose of accurately solving inverse problems by effectively discarding the high-frequency component while retaining nonlinear variations in the low-frequency part. The bi-polynomial reflectance model is useful for estimating reflectance and shape of an object. Experimental evaluation in comparison with other parametric reflectance models demonstrates that the proposed model achieves better performance in reflectometry and photometric stereo applications. Boxin Shi, Ping Tan 0002, Yasuyuki Matsushita, Katsushi Ikeuchi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | Video Tonal Stabilization via Color States SmoothingabstractWe address the problem of removing video color tone jitter that is common in amateur videos recorded with hand-held devices. To achieve this, we introduce color state to represent the exposure and white balance state of a frame. The color state of each frame can be computed by accumulating the color transformations of neighboring frame pairs. Then, the tonal changes of the video can be represented by a time-varying trajectory in color state space. To remove the tone jitter, we smooth the original color state trajectory by solving an L1 optimization problem with PCA dimensionality reduction. In addition, we propose a novel selective strategy to remove small tone jitter while retaining extreme exposure and white balance changes to avoid serious artifacts. Quantitative evaluation and visual comparison with previous work demonstrate the effectiveness of our tonal stabilization method. This system can also be used as a preprocessing tool for other video editing methods. Yinting Wang, Dacheng Tao, Xiang Li 0205, Mingli Song, Jiajun Bu, Ping Tan 0002 |
IEEE Trans. Image Process. | 6 |
| 2014 | TrackCam: 3D-aware tracking shots from consumer videoabstractPanning and tracking shots are popular photography techniques in which the camera tracks a moving object and keeps it at the same position, resulting in an image where the moving foreground is sharp but the background is blurred accordingly, creating an artistic illustration of the foreground motion. Such shots however are hard to capture even for professionals, especially when the foreground motion is complex (e.g., non-linear motion trajectories). In this work we propose a system to generate realistic, 3D-aware tracking shots from consumer videos. We show how computer vision techniques such as segmentation and structure-from-motion can be used to lower the barrier and help novice users create high quality tracking shots that are physically plausible. We also introduce a pseudo 3D approach for relative depth estimation to avoid expensive 3D reconstruction for improved robustness and a wider application range. We validate our system through extensive quantitative and qualitative evaluations. Shuaicheng Liu, Jue Wang 0001, Sunghyun Cho, Ping Tan 0002 |
ACM Trans. Graph. | 4 |
| 2013 | FrameBreak: Dramatic Image Extrapolation by Guided Shift-MapsabstractWe significantly extrapolate the field of view of a photograph by learning from a roughly aligned, wide-angle guide image of the same scene category. Our method can extrapolate typical photos into complete panoramas. The extrapolation problem is formulated in the shift-map image synthesis framework. We analyze the self-similarity of the guide image to generate a set of allowable local transformations and apply them to the input image. Our guided shift-map method reserves to the scene layout of the guide image when extrapolating a photograph. While conventional shift-map methods only support translations, this is not expressive enough to characterize the self-similarity of complex scenes. Therefore we additionally allow image transformations of rotation, scaling and reflection. To handle this increase in complexity, we introduce a hierarchical graph optimization method to choose the optimal transformation at each output pixel. We demonstrate our approach on a variety of indoor, outdoor, natural, and man-made scenes. Yinda Zhang 0001, Jianxiong Xiao, James Hays, Ping Tan 0002 |
CVPR | 4 |
| 2013 | A Global Linear Method for Camera Pose RegistrationabstractWe present a linear method for global camera pose registration from pair wise relative poses encoded in essential matrices. Our method minimizes an approximate geometric error to enforce the triangular relationship in camera triplets. This formulation does not suffer from the typical `unbalanced scale' problem in linear methods relying on pair wise translation direction constraints, i.e. an algebraic error, nor the system degeneracy from collinear motion. In the case of three cameras, our method provides a good linear approximation of the trifocal tensor. It can be directly scaled up to register multiple cameras. The results obtained are accurate for point triangulation and can serve as a good initialization for final bundle adjustment. We evaluate the algorithm performance with different types of data and demonstrate its effectiveness. Our system produces good accuracy, robustness, and outperforms some well-known systems on efficiency. Nianjuan Jiang, Zhaopeng Cui, Ping Tan 0002 |
ICCV | 3 |
| 2013 | Learning CRFs for Image Parsing with Adaptive Subgradient DescentabstractWe propose an adaptive sub gradient descent method to efficiently learn the parameters of CRF models for image parsing. To balance the learning efficiency and performance of the learned CRF models, the parameter learning is iteratively carried out by solving a convex optimization problem in each iteration, which integrates a proximal term to preserve the previously learned information and the large margin preference to distinguish bad labeling and the ground truth labeling. A solution of sub gradient descent updating form is derived for the convex optimization problem, with an adaptively determined updating step-size. Besides, to deal with partially labeled training data, we propose a new objective constraint modeling both the labeled and unlabeled parts in the partially labeled training data for the parameter learning of CRF models. The superior learning efficiency of the proposed method is verified by the experiment results on two public datasets. We also demonstrate the powerfulness of our method for handling partially labeled training data. Honghui Zhang, Jingdong Wang 0001, Ping Tan 0002, Jinglu Wang, Long Quan |
ICCV | 3 |
| 2013 | Bundled camera paths for video stabilizationabstractWe present a novel video stabilization method which models camera motion with a bundle of (multiple) camera paths. The proposed model is based on a mesh-based, spatially-variant motion representation and an adaptive, space-time path optimization. Our motion representation allows us to fundamentally handle parallax and rolling shutter effects while it does not require long feature trajectories or sparse 3D reconstruction. We introduce the 'as-similar-as-possible' idea to make motion estimation more robust. Our space-time path smoothing adaptively adjusts smoothness strength by considering discontinuities, cropping size and geometrical distortion in a unified optimization framework. The evaluation on a large variety of consumer videos demonstrates the merits of our method. Shuaicheng Liu, Lu Yuan 0001, Ping Tan 0002, Jian Sun 0001 |
ACM Trans. Graph. | 3 |
| 2012 | Video stabilization with a depth cameraabstractPrevious video stabilization methods often employ homographies to model transitions between consecutive frames, or require robust long feature tracks. However, the homography model is invalid for scenes with significant depth variations, and feature point tracking is fragile in videos with textureless objects, severe occlusion or camera rotation. To address these challenging cases, we propose to solve video stabilization with an additional depth sensor such as the Kinect camera. Though the depth image is noisy, incomplete and low resolution, it facilitates both camera motion estimation and frame warping, which make the video stabilization a much well posed problem. The experiments demonstrate the effectiveness of our algorithm. Shuaicheng Liu, Yinting Wang, Lu Yuan 0001, Jiajun Bu, Ping Tan 0002, Jian Sun 0001 |
CVPR | 5 |
| 2012 | A biquadratic reflectance model for radiometric image analysisabstractRadiometric image analysis methods heavily rely on reflectance models. Due to the complexity of real materials, methods based on simple models such as the Lambertian model often suffer from inaccuracy. On the other hand, more advanced models such as the Cook-Torrance model severely complicate the analysis problem. We tackle this dilemma by focusing on the low-frequency component of the reflectance. We propose a compact biquadratic reflectance model to represent the reflectance of a broad class of materials precisely in the low-frequency domain. We validate our model by fitting to both existing parametric models and non-parametric measured data, and show that our model outperforms existing parametric diffuse models. We show applications of reflectometry using general diffuse surfaces and photometric stereo for general isotropic materials. Experimental results show the effectiveness of our biquadratic model and its usefulness in radiometric image analysis. Boxin Shi, Ping Tan 0002, Yasuyuki Matsushita, Katsushi Ikeuchi |
CVPR | 2 |
| 2012 | 3D Reconstruction of Dynamic Scenes with Multiple Handheld Cameras
Hanqing Jiang, Haomin Liu, Ping Tan 0002, Guofeng Zhang 0001, Hujun Bao |
ECCV (2) | 3 |
| 2012 | Estimation of Intrinsic Image Sequences from Image+Depth Video
Kyong Joon Lee, Xin Tong 0001, Minmin Gong, Shahram Izadi, Sang Uk Lee, Ping Tan 0002, Stephen Lin 0001 |
ECCV (6) | 7 |
| 2012 | Elevation Angle from Reflectance Monotonicity: Photometric Stereo for General Isotropic Reflectances
Boxin Shi, Ping Tan 0002, Yasuyuki Matsushita, Katsushi Ikeuchi |
ECCV (3) | 2 |
| 2012 | Detecting discontinuities for surface reconstruction
Yinting Wang, Jiajun Bu, Mingli Song, Ping Tan 0002 |
ICPR | 5 |
| 2012 | A Closed-Form Solution to Retinex with Nonlocal Texture ConstraintsabstractWe propose a method for intrinsic image decomposition based on retinex theory and texture analysis. While most previous methods approach this problem by analyzing local gradient properties, our technique additionally identifies distant pixels with the same reflectance through texture analysis, and uses these nonlocal reflectance constraints to significantly reduce ambiguity in decomposition. We formulate the decomposition problem as the minimization of a quadratic function which incorporates both the retinex constraint and our nonlocal texture constraint. This optimization can be solved in closed form with the standard conjugate gradient algorithm. Extensive experimentation with comparisons to previous techniques validate our method in terms of both decomposition accuracy and runtime efficiency. Ping Tan 0002, Li Shen 0003, Enhua Wu, Stephen Lin 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Richardson-Lucy Deblurring for Scenes under a Projective Motion PathabstractThis paper addresses how to model and correct image blur that arises when a camera undergoes ego motion while observing a distant scene. In particular, we discuss how the blurred image can be modeled as an integration of the clear scene under a sequence of planar projective transformations (i.e., homographies) that describe the camera's path. This projective motion path blur model is more effective at modeling the spatially varying motion blur exhibited by ego motion than conventional methods based on space-invariant blur kernels. To correct the blurred image, we describe how to modify the Richardson-Lucy (RL) algorithm to incorporate this new blur model. In addition, we show that our projective motion RL algorithm can incorporate state-of-the-art regularization priors to improve the deblurred results. The projective motion path blur model, along with the modified RL algorithm, is detailed, together with experimental results demonstrating its overall effectiveness. Statistical analysis on the algorithm's convergence properties and robustness to noise is also provided. Yu-Wing Tai, Ping Tan 0002, Michael S. Brown |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Semantic colorization with internet imagesabstractColorization of a grayscale photograph often requires considerable effort from the user, either by placing numerous color scribbles over the image to initialize a color propagation algorithm, or by looking for a suitable reference image from which color information can be transferred. Even with this user supplied data, colorized images may appear unnatural as a result of limited user skill or inaccurate transfer of colors. To address these problems, we propose a colorization system that leverages the rich image content on the internet. As input, the user needs only to provide a semantic text label and segmentation cues for major foreground objects in the scene. With this information, images are downloaded from photo sharing websites and filtered to obtain suitable reference images that are reliable for color transfer to the given grayscale photo. Different image colorizations are generated from the various reference images, and a graphical user interface is provided to easily select the desired result. Our experiments and user study demonstrate the greater effectiveness of this system in comparison to previous techniques. Alex Yong Sang Chia, Shaojie Zhuo, Raj Kumar Gupta, Yu-Wing Tai, Siu-Yeung Cho, Ping Tan 0002, Stephen Lin 0001 |
ACM Trans. Graph. | 6 |
| 2010 | Self-calibrating photometric stereoabstractWe present a self-calibrating photometric stereo method. From a set of images taken from a fixed viewpoint under different and unknown lighting conditions, our method automatically determines a radiometric response function and resolves the generalized bas-relief ambiguity for estimating accurate surface normals and albedos. We show that color and intensity profiles, which are obtained from registered pixels across images, serve as effective cues for addressing these two calibration problems. As a result, we develop a complete auto-calibration method for photometric stereo. The proposed method is useful in many practical scenarios where calibrations are difficult. Experimental results validate the accuracy of the proposed method using various real-world scenes. Boxin Shi, Yasuyuki Matsushita, Chao Xu 0006, Ping Tan 0002 |
CVPR | 5 |
| 2009 | Photometric stereo and weather estimation using internet imagesabstractWe extend photometric stereo to make it work with internet images, which are typically associated with different viewpoints and significant noise. For popular tourism sites, thousands of images can be obtained from internet search engines. With these images, our method computes the global illumination for each image and the surface orientation at some scene points. The illumination information can then be used to estimate the weather conditions (such as sunny or cloudy) for each image, since there is a strong correlation between weather and scene illumination. We demonstrate our method on several challenging examples. Li Shen 0003, Ping Tan 0002 |
CVPR | 2 |
| 2008 | Intrinsic image decomposition with non-local texture cuesabstractWe present a method for decomposing an image into its intrinsic reflectance and shading components. Different from previous work, our method examines texture information to obtain constraints on reflectance among pixels that may be distant from one another in the image. We observe that distinct points with the same intensity-normalized texture configuration generally have the same reflectance value. The separation of shading and reflectance components should thus be performed in a manner that guarantees these non-local constraints. We formulate intrinsic image decomposition by adding these non-local texture constraints to the local derivative analysis employed in conventional techniques. Our results show a significant improvement in performance, with better recovery of global reflectance and shading structure than by previous methods. Li Shen 0003, Ping Tan 0002, Stephen Lin 0001 |
CVPR | 2 |
| 2008 | Subpixel Photometric StereoabstractConventional photometric stereo recovers one normal direction per pixel of the input image. This fundamentally limits the scale of recovered geometry to the resolution of the input image, and cannot model surfaces with subpixel geometric structures. In this paper, we propose a method to recover subpixel surface geometry by studying the relationship between the subpixel geometry and the reflectance properties of a surface. We first describe a generalized physically-based reflectance model that relates the distribution of surface normals inside each pixel area to its reflectance function. The distribution of surface normals can be computed from the reflectance functions recorded in photometric stereo images. A convexity measure of subpixel geometry structure is also recovered at each pixel, through an analysis of the shadowing attenuation. Then, we use the recovered distribution of surface normals and the surface convexity to infer subpixel geometric structures on a surface of homogeneous material by spatially arranging the normals among pixels at a higher resolution than that of the input image. Finally, we optimize the arrangement of normals using a combination of belief propagation and MCMC based on a minimum description length criterion on 3D textons over the surface. The experiments demonstrate the validity of our approach and show superior geometric resolution for the recovered surfaces. Ping Tan 0002, Stephen Lin 0001, Long Quan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2008 | Filtering and Rendering of Resolution-Dependent Reflectance ModelsabstractThe apparent reflectance of a surface depends upon the resolution at which it is imaged. Conventional reflectance models represent reflection at a single predetermined resolution; however, a low-resolution pixel that views a greater surface area often exhibits a reflectance more complicated than a high-resolution pixel with a smaller area. To address resolution dependency in reflectance, we utilize a generalized reflectance model based on a mixture of multiple conventional models, and present a framework for efficiently determining the reflectance mixture model of each pixel with respect to resolution. Mixture model parameters are precomputed at multiple resolutions and stored in mipmaps. Unlike color textures, these reflectance parameters cannot be accurately filtered by trilinear interpolation, so we present a technique for nonlinear mipmap filtering that minimizes aliasing in rendered results. This framework can be applied with various parametric reflectance models in graphics hardware for real-time processing. With this technique for filtering and rendering with mipmaps of reflectance mixture models, our system can rapidly render the resolution-dependent reflectance effects that are customarily disregarded in conventional rendering methods. At the end of this paper, we also describe how shadowing and masking effects can be incorporated into this framework to increase the realism of rendering. Ping Tan 0002, Stephen Lin 0001, Long Quan, Baining Guo, Harry Shum |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2007 | Joint Affinity Propagation for Multiple View SegmentationabstractA joint segmentation is a simultaneous segmentation of registered 2D images and 3D points reconstructed from the multiple view images. It is fundamental in structuring the data for subsequent modeling applications. In this paper, we treat this joint segmentation as a weighted graph labeling problem. First, we construct a 3D graph for the joint 3D and 2D points using a joint similarity measure. Then, we propose a hierarchical sparse affinity propagation algorithm to automatically and jointly segment 2D images and group 3D points. Third, a semi-supervised affinity propagation algorithm is proposed to refine the automatic results with the user assistance. Finally, intensive experiments demonstrate the effectiveness of the proposed approaches. Jianxiong Xiao, Jingdong Wang 0001, Ping Tan 0002, Long Quan |
ICCV | 3 |
| 2007 | Image-Based Modeling by Joint Segmentation
Long Quan, Jingdong Wang 0001, Ping Tan 0002, Lu Yuan 0001 |
Int. J. Comput. Vis. | 3 |
| 2007 | Image-based tree modelingabstractIn this paper, we propose an approach for generating 3D models of natural-looking trees from images that has the additional benefit of requiring little user intervention. While our approach is primarily image-based, we do not model each leaf directly from images due to the large leaf count, small image footprint, and widespread occlusions. Instead, we populate the tree with leaf replicas from segmented source images to reconstruct the overall tree shape. In addition, we use the shape patterns of visible branches to predict those of obscured branches. We demonstrate our approach on a variety of trees. Ping Tan 0002, Jingdong Wang 0001, Sing Bing Kang, Long Quan |
ACM Trans. Graph. | 1 |
| 2006 | Separation of Highlight Reflections on Textured SurfacesabstractWe present a method for separating highlight reflections on textured surfaces. In contrast to previous techniques that use diffuse color information from outside the highlight area to constrain the solution, the proposed method further capitalizes on the spatial distributions of colors to resolve ambiguities in separation that often arise in real images. For highlight pixels in which a clear-cut separation cannot be determined from color space analysis, we evaluate possible separation solutions based on their consistency with diffuse texture characteristics outside the highlight. With consideration of color distributions in both the color space and the image space, appreciably enhanced separation performance can be attained in challenging cases. Ping Tan 0002, Long Quan, Stephen Lin 0001 |
CVPR (2) | 1 |
| 2006 | Resolution-Enhanced Photometric Stereo
Ping Tan 0002, Stephen Lin 0001, Long Quan |
ECCV (3) | 1 |
| 2006 | Image-based plant modelingabstractIn this paper, we propose a semi-automatic technique for modeling plants directly from images. Our image-based approach has the distinct advantage that the resulting model inherits the realistic shape and complexity of a real plant. We designed our modeling system to be interactive, automating the process of shape recovery while relying on the user to provide simple hints on segmentation. Segmentation is performed in both image and 3D spaces, allowing the user to easily visualize its effect immediately. Using the segmented image and 3D data, the geometry of each leaf is then automatically recovered from the multiple views by fitting a deformable leaf model. Our system also allows the user to easily reconstruct branches in a similar manner. We show realistic reconstructions of a variety of plants, and demonstrate examples of plant editing. Long Quan, Ping Tan 0002, Lu Yuan 0001, Jingdong Wang 0001, Sing Bing Kang |
ACM Trans. Graph. | 2 |
| 2005 | Multiresolution Reflectance FilteringabstractPhysically-based reflectance models typically represent light scattering as a function of surface geometry at the pixel level. With changes in viewing resolution, the geometry imaged within a pixel can undergo significant variations that result in changing reflectance characteristics. To address these transformations, we present a multiresolution reflectance framework based on microfacet normal distributions within a pixel over different scales. Since these distributions must be efficiently determined with respect to resolution, they are recorded at multiple resolution levels in mipmaps. The main contribution of this work is a real-time mipmap filtering technique for these distribution-based parameters that not only provides smooth reflectance transitions in scale, but also minimizes aliasing. With this multiresolution reflectance technique, our system can rapidly and accurately incorporate fine reflectance detail that is customarily disregarded in multiresolution rendering methods. Ping Tan 0002, Stephen Lin 0001, Long Quan, Baining Guo, Harry Shum |
Rendering Techniques | 1 |
| 2003 | Highlight Removal by Illumination-Constrained InpaintingabstractWe present a single-image highlight removal method that incorporates illumination-based constraints into image inpainting. Unlike occluded image regions filled by traditional inpainting, highlight pixels contain some useful information for guiding the inpainting process. Constraints provided by observed pixel colors, highlight color analysis and illumination color uniformity are employed in our method to improve estimation of the underlying diffuse color. The inclusion of these illumination constraints allows for better recovery of shading and textures by inpainting. Experimental results are given to demonstrate the performance of our method. Ping Tan 0002, Stephen Lin 0001, Long Quan, Harry Shum |
ICCV | 1 |