EDBT 2026 Demo / reviewers in the wild / expert
Ruigang Yang
dblp:08/5690
· DBLP profile ↗
160ranked-venue papers
13as first author
32since 2021 · last 2026
0000-0001-5296-6307ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 115 · 10 first-author · 18 since 2021Artificial intelligence and machine learning · 107 · 6 first-author · 21 since 2021Systems, architecture and hardware · 10 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Real Garment Benchmark (RGBench): A Comprehensive Benchmark for Robotic Garment Manipulation Featuring a High-Fidelity Scalable SimulatorabstractWhile there has been significant progress to use simulated data to learn robotic manipulation of rigid objects, applying its success to deformable objects has been hindered by the lack of both deformable object models and realistic non-rigid body simulators. In this paper, we present Real Garment Benchmark (RGBench), a comprehensive benchmark for robotic manipulation of garments. It features a diverse set of over 6000 garment mesh models, a new high-performance simulator, and a comprehensive protocol to evaluate garment simulation quality with carefully measured real garment dynamics. Our experiments demonstrate that our simulator outperforms currently available cloth simulators by a large margin, reducing simulation error by 20% while maintaining a speed of 3 times faster. We will publicly release RGBench to accelerate future research in robotic garment manipulation. Wenkang Hu, Xincheng Tang, Yanzhi E, Zhengjie Shu, Wei Li 0111, Huamin Wang 0001, Ruigang Yang |
AAAI | 8 |
| 2025 | OLiDM: Object-aware LiDAR Diffusion Models for Autonomous DrivingabstractTo enhance autonomous driving, innovative approaches have been proposed to generate simulated LiDAR data. However, these methods often face challenges in producing high-quality and controllable foreground objects. To cater to the needs of object-aware tasks in 3D perception, we introduce OLiDM, a novel framework capable of generating controllable and high-fidelity LiDAR data at both the object and scene levels. OLiDM consists of two pivotal components: the Object-Scene Progressive Generation (OPG) module and the Object Semantic Alignment (OSA) module. OPG adapts to user-specific prompts to generate desired foreground objects, which are subsequently employed as conditions in scene generation, ensuring controllable and diverse output at both the object and scene levels. This also facilitates the association of user-defined object-level annotations with the generated LiDAR scenes. Moreover, OSA aims to rectify the misalignment between foreground objects and background scenes, enhancing the overall quality of the generated objects. The broad efficacy of OLiDM is demonstrated across both unconditional and conditional LiDAR generation tasks, as well as 3D perception tasks. Specifically, on the KITTI-360 dataset, OLiDM surpasses prior state-of-the-art methods such as UltraLiDAR by 11.8 in FPD, producing data that closely mirrors real-world distributions. Additionally, in sparse-to-dense LiDAR completion, OLiDM achieves a significant improvement over LiDARGen, with a 57.47% increase in semantic IoU. Moreover, in 3D object detection, OLiDM enhances the performance of mainstream detectors by 2.4% in mAP and 1.9% in NDS, underscoring its potential in advancing 3D perception models. Tianyi Yan, Junbo Yin, Xianpeng Lang, Ruigang Yang, Cheng-Zhong Xu 0001, Jianbing Shen |
AAAI | 4 |
| 2025 | Fuel-Optimal Operational Speed Planning for Autonomous Trucking on HighwaysabstractThe rapid advancement of autonomous driving technology, particularly in autonomous trucking on highways, shows great value for enhancing efficiency and reducing costs in the logistics industry. In this work, we define the full-trip speed planning problem for autonomous trucks under delivery time and fuel consumption constraints, referred to as the Operational Speed Planning (OSP) problem. To support and accelerate research on the OSP problem, we have developed a comprehensive dataset using a fleet of over 400 trucks. The dataset contains rich, diverse information covering more than 22 million kilometers of real-world highway driving data. In addition to this static dataset, we have developed a closed-loop simulator that allows for the interactive evaluation of OSP solutions, enabling researchers to test speed planning strategies in a realistic environment. Furthermore, we provide an OSP baseline method based on dynamic programming to optimize speed planning, balancing the delivery time requirements and fuel consumption. Our extensive experiments demonstrate both the accuracy of the simulation and the effectiveness of the OSP baseline in planning optimal speeds, proving its capability to meet time constraints while improving fuel efficiency. The dataset, simulator, and baseline will be made publicly available to foster further research and innovation in this area. Wei Li 0111, Jiahao Xiang, Jiaping Ren, Ruigang Yang |
ICRA | 6 |
| 2025 | LLMamba-Net: A Lightweight Network Integrating Linear Mamba for Facial Expression Recognition
Kaidi Hu, Guojiao Zhao, Ruigang Yang |
PRCV (7) | 4 |
| 2025 | Modality Confusion Learning: A Versatile Framework for Visible-Infrared Re-identification
Sanyuan Zhao, Mang Ye, Ruigang Yang, Jianbing Shen |
Int. J. Comput. Vis. | 4 |
| 2025 | IEMFormer: Internal and External Multi-Fusion Transformer for Indoor RGB-D Semantic SegmentationabstractEffectively fusing and complementing RGB and depth modalities while mitigating image noise is a critical challenge in the RGB-D semantic segmentation task. In this paper, we propose a novel Internal and External Multi-fusion Transformer (IEMFormer) to address this issue. IEMFormer incorporates stage-specific fusion strategies to enhance modal complementarity. For internal fusion, we integrate a fusion unit within the traditional Transformer block, combining matching tokens from both modalities on a pixel-by-pixel basis. For external fusion, the proposed External Adaptive Cross-modal Fusion (EACF) module filters dual-modal features across both spatial and channel dimensions, serving the purpose of adaptively weighting complementary channel information and robustly aggregating spatial patterns from both modalities, thereby facilitating the integration of multimodal information. Additionally, the Global Self-attention Guided Fusion (GSGF) module in the decoder refines the fused features from earlier stages, effectively suppressing noise. This is achieved by leveraging high-level semantic features to guide the refinement and incorporating an active noise suppression mechanism to prevent overfitting to dominant, noisy features. Extensive experiments on the NYUv2 and SUN RGB-D datasets demonstrate that IEMFormer achieves highly competitive performance in accurately understanding indoor scenes. Kaidi Hu, Wei Li 0111, Guangwei Gao, Ruigang Yang |
IEEE Signal Process. Lett. | 4 |
| 2025 | ABFE-Net: Attention-Based Feature Enhancement Network for Few-Shot Point Cloud ClassificationabstractFew-shot 3D point cloud classification has attracted significant attention due to the challenge of acquiring large-scale labeled data. Existing methods often employ network backbones tailored for fully-supervised learning, which can lead to suboptimal performance in few-shot settings. To tackle these limitations, we propose ABFE-Net, a novel method for point cloud classification with few-shot learning principles. We comprehensively summarize the drawbacks of existing network architectures into four aspects: contextual information loss, channel redundancy, overfitting, and insufficient hidden feature extraction. Accordingly, we design novel modules, such as the Attention-based Dilated Mix-up Module (ADMM) and Attention-based Comprehensive Feature Learning (ACFL), to enhance the network by addressing those issues effectively. Experiments on multiple public datasets demonstrate that ABFE-Net achieves state-of-the-art performance with superior generalization. Kaidi Hu, Mao Ye 0005, Wei Li 0111, Ruigang Yang |
IEEE Signal Process. Lett. | 5 |
| 2024 | DI-V2X: Learning Domain-Invariant Representation for Vehicle-Infrastructure Collaborative 3D Object DetectionabstractVehicle-to-Everything (V2X) collaborative perception has recently gained significant attention due to its capability to enhance scene understanding by integrating information from various agents, e.g., vehicles, and infrastructure. However, current works often treat the information from each agent equally, ignoring the inherent domain gap caused by the utilization of different LiDAR sensors of each agent, thus leading to suboptimal performance. In this paper, we propose DI-V2X, that aims to learn Domain-Invariant representations through a new distillation framework to mitigate the domain discrepancy in the context of V2X 3D object detection. DI-V2X comprises three essential components: a domain-mixing instance augmentation (DMA) module, a progressive domain-invariant distillation (PDD) module, and a domain-adaptive fusion (DAF) module. Specifically, DMA builds a domain-mixing 3D instance bank for the teacher and student models during training, resulting in aligned data representation. Next, PDD encourages the student models from different domains to gradually learn a domain-invariant feature representation towards the teacher, where the overlapping regions between agents are employed as guidance to facilitate the distillation process. Furthermore, DAF closes the domain gap between the students by incorporating calibration-aware domain-adaptive attention. Extensive experiments on the challenging DAIR-V2X and V2XSet benchmark datasets demonstrate DI-V2X achieves remarkable performance, outperforming all the previous V2X models. Code is available at https://github.com/Serenos/DI-V2X. Xiang Li 0001, Junbo Yin, Wei Li 0111, Cheng-Zhong Xu 0001, Ruigang Yang, Jianbing Shen |
AAAI | 5 |
| 2024 | IS-Fusion: Instance-Scene Collaborative Fusion for Multimodal 3D Object DetectionabstractBird's eye view (BEV) representation has emerged as a dominant solution for describing 3D space in autonomous driving scenarios. However, objects in the BEV representation typically exhibit small sizes, and the associated point cloud context is inherently sparse, which leads to great challenges for reliable 3D perception. In this paper, we propose IS-Fusion, an innovative multimodal fusion framework that jointly captures the Instance- and Scene-level contextual information. IS-Fusion essentially differs from existing approaches that only focus on the BEV scene-level fusion by explicitly incorporating instance-level multimodal information, thus facilitating the instance-centric tasks like 3D object detection. It comprises a Hierarchical Scene Fusion (HSF) module and an Instance-Guided Fusion (IGF) module. HSF applies Point-to-Grid and Grid-to-Region transformers to capture the multimodal scene context at different granularities. IGF mines instance candidates, explores their relationships, and aggregates the local multimodal context for each instance. These instances then serve as guidance to enhance the scene feature and yield an instance-aware BEV representation. On the challenging nuScenes benchmark, IS-Fusion outperforms all the published multimodal works to date. Code is available at: https://github.com/yinjunbo/IS-Fusion. Junbo Yin, Jianbing Shen, Runnan Chen, Wei Li 0111, Ruigang Yang, Pascal Frossard, Wenguan Wang |
CVPR | 5 |
| 2024 | NPC: Neural Predictive Control for Fuel-Efficient Autonomous TrucksabstractFuel efficiency is a crucial aspect of long-distance cargo transportation by oil-powered trucks that economize on costs and decrease carbon emissions. Current predictive control methods depend on an accurate model of vehicle dynamics and engine, including weight, drag coefficient, and the Brake-specific Fuel Consumption (BSFC) map of the engine. We propose a pure data-driven method, Neural Predictive Control (NPC), which does not use any physical model for the vehicle. After training with over 20,000 km of historical data, the novel proposed NVFormer implicitly models the relationship between vehicle dynamics, road slope, fuel consumption, and control commands using the attention mechanism. Based on the online sampled primitives from the past of the current freight trip and anchor-based future data synthesis, the NVFormer can infer optimal control command for reasonable fuel consumption. The physical model-free NPC outperforms the base PCC method with 2.41% and 3.45% more significant fuel saving in simulation and open-road highway testing, respectively. Jiaping Ren, Jiahao Xiang, Hongfei Gao, Yiming Ren 0001, Yuexin Ma, Ruigang Yang, Wei Li 0111 |
ICRA | 8 |
| 2024 | ESP: Extro-Spective Prediction for Long-term Behavior Reasoning in Emergency ScenariosabstractEmergent-scene safety is the key milestone for fully autonomous driving, and reliable on-time prediction is essential to maintain safety in emergency scenarios. However, these emergency scenarios are long-tailed and hard to collect, which restricts the system from getting reliable predictions. In this paper, we build a new dataset, which aims at the longterm prediction with the inconspicuous state variation in history for the emergency event, named the Extro-Spective Prediction (ESP) problem. Based on the proposed dataset, a flexible feature encoder for ESP is introduced to various prediction methods as a seamless plug-in, and its consistent performance improvement underscores its efficacy. Furthermore, a new metric named clamped temporal error (CTE) is proposed to give a more comprehensive evaluation of prediction performance, especially in time-sensitive emergency events of subseconds. Interestingly, as our ESP features can be described in human-readable language naturally, the application of integrating into ChatGPT also shows huge potential. The ESP-dataset and all benchmarks are released at https://dingrui-wang.github.io/ESP-Dataset/. Dingrui Wang, Zheyuan Lai, Yuda Li, Yuexin Ma, Johannes Betz, Ruigang Yang, Wei Li 0111 |
ICRA | 7 |
| 2024 | Vision-Centric BEV Perception: A SurveyabstractIn recent years, vision-centric Bird's Eye View (BEV) perception has garnered significant interest from both industry and academia due to its inherent advantages, such as providing an intuitive representation of the world and being conducive to data fusion. The rapid advancements in deep learning have led to the proposal of numerous methods for addressing vision-centric BEV perception challenges. However, there has been no recent survey encompassing this novel and burgeoning research field. To catalyze future research, this paper presents a comprehensive survey of the latest developments in vision-centric BEV perception and its extensions. It compiles and organizes up-to-date knowledge, offering a systematic review and summary of prevalent algorithms. Additionally, the paper provides in-depth analyses and comparative results on various BEV perception tasks, facilitating the evaluation of future works and sparking new research directions. Furthermore, the paper discusses and shares valuable empirical implementation details to aid in the advancement of related algorithms. Yuexin Ma, Xuyang Bai, Huitong Yang, Yuenan Hou, Yaming Wang, Yu Qiao 0001, Ruigang Yang, Xinge Zhu |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2023 | LWSIS: LiDAR-Guided Weakly Supervised Instance Segmentation for Autonomous DrivingabstractImage instance segmentation is a fundamental research topic in autonomous driving, which is crucial for scene understanding and road safety. Advanced learning-based approaches often rely on the costly 2D mask annotations for training. In this paper, we present a more artful framework, LiDAR-guided Weakly Supervised Instance Segmentation (LWSIS), which leverages the off-the-shelf 3D data, i.e., Point Cloud, together with the 3D boxes, as natural weak supervisions for training the 2D image instance segmentation models. Our LWSIS not only exploits the complementary information in multimodal data during training but also significantly reduces the annotation cost of the dense 2D masks. In detail, LWSIS consists of two crucial modules, Point Label Assignment (PLA) and Graph-based Consistency Regularization (GCR). The former module aims to automatically assign the 3D point cloud as 2D point-wise labels, while the atter further refines the predictions by enforcing geometry and appearance consistency of the multimodal data. Moreover, we conduct a secondary instance segmentation annotation on the nuScenes, named nuInsSeg, to encourage further research on multimodal perception tasks. Extensive experiments on the nuInsSeg, as well as the large-scale Waymo, show that LWSIS can substantially improve existing weakly supervised segmentation models by only involving 3D data during training. Additionally, LWSIS can also be incorporated into 3D object detectors like PointPainting to boost the 3D detection performance for free. The code and dataset are available at https://github.com/Serenos/LWSIS. Xiang Li 0001, Junbo Yin, Botian Shi, Yikang Li 0002, Ruigang Yang, Jianbing Shen |
AAAI | 5 |
| 2023 | SSDA3D: Semi-supervised Domain Adaptation for 3D Object Detection from Point CloudabstractLiDAR-based 3D object detection is an indispensable task in advanced autonomous driving systems. Though impressive detection results have been achieved by superior 3D detectors, they suffer from significant performance degeneration when facing unseen domains, such as different LiDAR configurations, different cities, and weather conditions. The mainstream approaches tend to solve these challenges by leveraging unsupervised domain adaptation (UDA) techniques. However, these UDA solutions just yield unsatisfactory 3D detection results when there is a severe domain shift, e.g., from Waymo (64-beam) to nuScenes (32-beam). To address this, we present a novel Semi-Supervised Domain Adaptation method for 3D object detection (SSDA3D), where only a few labeled target data is available, yet can significantly improve the adaptation performance. In particular, our SSDA3D includes an Inter-domain Adaptation stage and an Intra-domain Generalization stage. In the first stage, an Inter-domain Point-CutMix module is presented to efficiently align the point cloud distribution across domains. The Point-CutMix generates mixed samples of an intermediate domain, thus encouraging to learn domain-invariant knowledge. Then, in the second stage, we further enhance the model for better generalization on the unlabeled target set. This is achieved by exploring Intra-domain Point-MixUp in semi-supervised learning, which essentially regularizes the pseudo label distribution. Experiments from Waymo to nuScenes show that, with only 10% labeled target data, our SSDA3D can surpass the fully-supervised oracle model with 100% target label. Our code is available at https://github.com/yinjunbo/SSDA3D. Yan Wang 0116, Junbo Yin, Wei Li 0111, Pascal Frossard, Ruigang Yang, Jianbing Shen |
AAAI | 5 |
| 2023 | Transformation-Equivariant 3D Object Detection for Autonomous Drivingabstract3D object detection received increasing attention in autonomous driving recently. Objects in 3D scenes are distributed with diverse orientations. Ordinary detectors do not explicitly model the variations of rotation and reflection transformations. Consequently, large networks and extensive data augmentation are required for robust detection. Recent equivariant networks explicitly model the transformation variations by applying shared networks on multiple transformed point clouds, showing great potential in object geometry modeling. However, it is difficult to apply such networks to 3D object detection in autonomous driving due to its large computation cost and slow reasoning speed. In this work, we present TED, an efficient Transformation-Equivariant 3D Detector to overcome the computation cost and speed issues. TED first applies a sparse convolution backbone to extract multi-channel transformation-equivariant voxel features; and then aligns and aggregates these equivariant features into lightweight and compact representations for high-performance 3D object detection. On the highly competitive KITTI 3D car detection leaderboard, TED ranked 1st among all submissions with competitive efficiency. Code is available at https://github.com/hailanyi/TED. Chenglu Wen, Wei Li 0111, Xin Li 0003, Ruigang Yang, Cheng Wang 0003 |
AAAI | 5 |
| 2023 | Bridging Language and Geometric Primitives for Zero-shot Point Cloud SegmentationabstractWe investigate transductive zero-shot point cloud semantic segmentation, where the network is trained on seen objects and able to segment unseen objects. The 3D geometric elements are essential cues to imply a novel 3D object type. However, previous methods neglect the fine-grained relationship between the language and the 3D geometric elements. To this end, we propose a novel framework to learn the geometric primitives shared in seen and unseen categories' objects and employ a fine-grained alignment between language and the learned geometric primitives. Therefore, guided by language, the network recognizes the novel objects represented with geometric primitives. Specifically, we formulate a novel point visual representation, the similarity vector of the point's feature to the learnable prototypes, where the prototypes automatically encode geometric primitives via back-propagation. Besides, we propose a novel Unknown-aware InfoNCE Loss to fine-grained align the visual representation with language. Extensive experiments show that our method significantly outperforms other state-of-the-art methods in the harmonic mean-intersection-over-union (hIoU), with the improvement of 17.8%, 30.4%, 9.2% and 7.9% on S3DIS, ScanNet, SemanticKITTI and nuScenes datasets, respectively. Codes are available1 https://github.com/runnanchen/Zero-Shot-Point-Cloud-Segmentation. Runnan Chen, Xinge Zhu, Nenglun Chen, Wei Li 0111, Yuexin Ma, Ruigang Yang, Wenping Wang 0001 |
ACM Multimedia | 6 |
| 2023 | Graph Neural Network and Spatiotemporal Transformer Attention for 3D Video Object Detection From Point CloudsabstractPrevious works for LiDAR-based 3D object detection mainly focus on the single-frame paradigm. In this paper, we propose to detect 3D objects by exploiting temporal information in multiple frames, i.e., point cloud videos. We empirically categorize the temporal information into short-term and long-term patterns. To encode the short-term data, we present a Grid Message Passing Network (GMPNet), which considers each grid (i.e., the grouped points) as a node and constructs a k-NN graph with the neighbor grids. To update features for a grid, GMPNet iteratively collects information from its neighbors, thus mining the motion cues in grids from nearby frames. To further aggregate long-term frames, we propose an Attentive Spatiotemporal Transformer GRU (AST-GRU), which contains a Spatial Transformer Attention (STA) module and a Temporal Transformer Attention (TTA) module. STA and TTA enhance the vanilla GRU to focus on small objects and better align moving objects. Our overall framework supports both online and offline video object detection in point clouds. We implement our algorithm based on prevalent anchor-based and anchor-free detectors. Evaluation results on the challenging nuScenes benchmark show superior performance of our method, achieving first on the leaderboard (at the time of paper submission) without any "bells and whistles." Our source code is available at https://github.com/shenjianbing/GMP3D. Junbo Yin, Jianbing Shen, Xin Gao 0001, David Crandall, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | FaceScape: 3D Facial Dataset and Benchmark for Single-View 3D Face ReconstructionabstractIn this article, we present a large-scale detailed 3D face dataset, FaceScape, and the corresponding benchmark to evaluate single-view facial 3D reconstruction. By training on FaceScape data, a novel algorithm is proposed to predict elaborate riggable 3D face models from a single image input. FaceScape dataset releases 16,940 textured 3D faces, captured from 847 subjects and each with 20 specific expressions. The 3D models contain the pore-level facial geometry that is also processed to be topologically uniform. These fine 3D facial models can be represented as a 3D morphable model for coarse shapes and displacement maps for detailed geometry. Taking advantage of the large-scale and high-accuracy dataset, a novel algorithm is further proposed to learn the expression-specific dynamic details using a deep neural network. The learned relationship serves as the foundation of our 3D face prediction system from a single image input. Different from most previous methods, our predicted 3D models are riggable with highly detailed geometry under different expressions. We also use FaceScape data to generate the in-the-wild and in-the-lab benchmark to evaluate recent methods of single-view face reconstruction. The accuracy is reported and analyzed on the dimensions of camera pose and focal length, which provides a faithful and comprehensive evaluation and reveals new challenges. The unprecedented dataset, benchmark, and code have been released to the public for research purpose. Hao Zhu 0004, Longwei Guo, Mingkai Huang, Menghua Wu, Qiu Shen, Ruigang Yang, Xun Cao |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2023 | Fuel Rate Prediction for Heavy-Duty TrucksabstractFuel cost contributes significantly to the high operation cost of heavy-duty trucks. Developing fuel rate prediction models is the cornerstone of fuel consumption optimization approaches for heavy-duty trucks. However, limited by accurate features directly related to the truck’s fuel consumption, state-of-the-art models show poor performance and are rarely deployed in practice. In this paper, we use the truck’s engine management system (EMS) and Instant Fuel Meter (IFM) to collect a three-month dataset during the period of December 2019 to June 2020. Seven prediction models, including linear regression, polynomial regression, MLP, CNN, LSTM, CNN-LSTM, and AutoML, are investigated and evaluated to predict real-time fuel rate. The evaluation results show that the EMS and IFM dataset help to improve the coefficient of determination of traditional linear/polynomial models from 0.87 to 0.96, while learning-based approach AutoML improves the coefficient of determination to attain 0.99. Besides, we explore the actual deployment of fuel rate prediction with transfer learning and path planning for autonomous driving. Liangkai Liu, Wei Li 0111, Dawei Wang 0006, Ruigang Yang, Weisong Shi |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | STCrowd: A Multimodal Dataset for Pedestrian Perception in Crowded ScenesabstractAccurately detecting and tracking pedestrians in 3D space is challenging due to large variations in rotations, poses and scales. The situation becomes even worse for dense crowds with severe occlusions. However, existing benchmarks either only provide 2D annotations, or have limited 3D annotations with low-density pedestrian distribution, making it difficult to build a reliable pedestrian perception system especially in crowded scenes. To better evaluate pedestrian perception algorithms in crowded scenarios, we introduce a large-scale multimodal dataset, STCrowd. Specifically, in STCrowd, there are a total of 219 K pedestrian instances and 20 persons per frame on average, with various levels of occlusion. We provide synchronized LiDAR point clouds and camera images as well as their corresponding 3D labels and joint IDs. STCrowd can be used for various tasks, including LiDAR-only, image-only, and sensor-fusion based pedestrian detection and tracking. We provide baselines for most of the tasks. In addition, considering the property of sparse global distribution and density-varying local distribution of pedestrians, we further propose a novel method, Density-aware Hierarchical heatmap Aggregation (DHA), to enhance pedestrian perception in crowded scenes. Extensive experiments show that our new method achieves state-of-the-art performance for pedestrian detection on various datasets. https://github.com/4DVLab/STCrowd.git. Peishan Cong, Xinge Zhu, Feng Qiao 0001, Yiming Ren 0001, Xidong Peng, Yuenan Hou, Lan Xu 0003, Ruigang Yang, Dinesh Manocha, Yuexin Ma |
CVPR | 8 |
| 2022 | Part-Level Car Parsing and Reconstruction in Single Street View ImagesabstractPart information has been proven to be resistant to occlusions and viewpoint changes, which are main difficulties in car parsing and reconstruction. However, in the absence of datasets and approaches incorporating car parts, there are limited works that benefit from it. In this paper, we propose the first part-aware approach for joint part-level car parsing and reconstruction in single street view images. Without labor-intensive part annotations on real images, our approach simultaneously estimates pose, shape, and semantic parts of cars. There are two contributions in this paper. First, our network introduces dense part information to facilitate pose and shape estimation, which is further optimized with a novel 3D loss. To obtain part information in real images, a class-consistent method is introduced to implicitly transfer part knowledge from synthesized images. Second, we construct the first high-quality dataset containing 348 car models with physical dimensions and part annotations. Given these models, 60K synthesized images with randomized configurations are generated. Experimental results demonstrate that part knowledge can be effectively transferred with our class-consistent method, which significantly improves part segmentation performance on real street views. By fusing dense part information, our pose and shape estimation results achieve the state-of-the-art performance on the ApolloCar3D and outperform previous approaches by large margins in terms of both A3DP-Abs and A3DP-Rel. Qichuan Geng, Hong Zhang 0009, Feixiang Lu, Xinyu Huang 0001, Sen Wang 0003, Zhong Zhou, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2022 | Salient Object Detection in the Deep Learning Era: An In-Depth SurveyabstractAs an essential problem in computer vision, salient object detection (SOD) has attracted an increasing amount of research attention over the years. Recent advances in SOD are predominantly led by deep learning-based solutions (named deep SOD). To enable in-depth understanding of deep SOD, in this paper, we provide a comprehensive survey covering various aspects, ranging from algorithm taxonomy to unsolved issues. In particular, we first review deep SOD algorithms from different perspectives, including network architecture, level of supervision, learning paradigm, and object-/instance-level detection. Following that, we summarize and analyze existing SOD datasets and evaluation metrics. Then, we benchmark a large group of representative SOD models, and provide detailed analyses of the comparison results. Moreover, we study the performance of SOD algorithms under different attribute settings, which has not been thoroughly explored previously, by constructing a novel SOD dataset with rich attribute annotations covering various salient object types, challenging factors, and scene categories. We further analyze, for the first time in the field, the robustness of SOD models to random input perturbations and adversarial attacks. We also look into the generalization and difficulty of existing SOD datasets. Finally, we discuss several open issues of SOD and outline future research directions. All the saliency prediction maps, our constructed dataset with annotations, and codes for evaluation are publicly available at https://github.com/wenguanwang/SODsurvey. Wenguan Wang, Qiuxia Lai, Huazhu Fu, Jianbing Shen, Haibin Ling, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Cylindrical and Asymmetrical 3D Convolution Networks for LiDAR-Based PerceptionabstractState-of-the-art methods for driving-scene LiDAR-based perception (including point cloud semantic segmentation, panoptic segmentation and 3D detection, etc.) often project the point clouds to 2D space and then process them via 2D convolution. Although this cooperation shows the competitiveness in the point cloud, it inevitably alters and abandons the 3D topology and geometric relations. A natural remedy is to utilize the 3D voxelization and 3D convolution network. However, we found that in the outdoor point cloud, the improvement obtained in this way is quite limited. An important reason is the property of the outdoor point cloud, namely sparsity and varying density. Motivated by this investigation, we propose a new framework for the outdoor LiDAR segmentation, where cylindrical partition and asymmetrical 3D convolution networks are designed to explore the 3D geometric pattern while maintaining these inherent properties. The proposed model acts as a backbone and the learned features from this model can be used for downstream tasks such as point cloud semantic and panoptic segmentation or 3D detection. In this paper, we benchmark our model on these three tasks. For semantic segmentation, we evaluate the proposed model on several large-scale datasets, i.e., SemanticKITTI, nuScenes and A2D2. Our method achieves the state-of-the-art on the leaderboard of SemanticKITTI (both single-scan and multi-scan challenge), and significantly outperforms existing methods on nuScenes and A2D2 dataset. Furthermore, the proposed 3D framework also shows strong performance and good generalization on LiDAR panoptic segmentation and LiDAR 3D detection. Xinge Zhu, Hui Zhou 0005, Fangzhou Hong, Wei Li 0111, Yuexin Ma, Hongsheng Li 0001, Ruigang Yang, Dahua Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2022 | Detailed Avatar Recovery From Single ImageabstractThis paper presents a novel framework to recover detailed avatar from a single image. It is a challenging task due to factors such as variations in human shapes, body poses, texture, and viewpoints. Prior methods typically attempt to recover the human body shape using a parametric-based template that lacks the surface details. As such resulting body shape appears to be without clothing. In this paper, we propose a novel learning-based framework that combines the robustness of the parametric model with the flexibility of free-form 3D deformation. We use the deep neural networks to refine the 3D shape in a Hierarchical Mesh Deformation (HMD) framework, utilizing the constraints from body joints, silhouettes, and per-pixel shading information. Our method can restore detailed human body shapes with complete textures beyond skinned models. Experiments demonstrate that our method has outperformed previous state-of-the-art approaches, achieving better accuracy in terms of both 2D IoU number and 3D metric distance. Hao Zhu 0004, Xinxin Zuo, Sen Wang 0003, Xun Cao, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | WAFP-Net: Weighted Attention Fusion Based Progressive Residual Learning for Depth Map Super-ResolutionabstractDespite the remarkable progresses achieved in depth map super-resolution (DSR), it remains a major challenge to tackle with real-world degradation of low-resolution (LR) depth maps. Synthetic datasets are mainly used in existing DSR approaches, which is quite different from what would get from a real depth sensor. Besides, the enhancements of features in existing DSR approaches are not sufficiently enough, which also limit the performance. To alleviate these problems, we first propose two types of degradation models to describe the generation of LR depth maps, including bi-cubic down-sampling with noise and interval down-sampling, and different DSR models are learned correspondingly. Then, we propose a weighted attention fusion strategy that is embedded into a progressive residual learning framework, which guarantees that the high-resolution (HR) depth maps can be well recovered in a coarse-to-fine manner. The weighted attention fusion strategy can enhance the features with abundant high-frequency components in both global and local manners, thus better HR depth maps can be expected. Besides, to re-use the effective information in the progressive process sufficiently, a multi-stage fusion module is combined into the proposed framework, and the Total Generalized Variation (TGV) regularization and input loss are exploited to further improve the performance of our method. Extensive experiments of different benchmarks demonstrate the superiority of our approach over the state-of-the-art (SOTA) approaches. Xibin Song, Dingfu Zhou, Wei Li 0111, Yuchao Dai, Liu Liu 0009, Hongdong Li, Ruigang Yang, Liangjun Zhang |
IEEE Trans. Multim. | 7 |
| 2021 | Adaptive Surface Normal Constraint for Depth EstimationabstractWe present a novel method for single image depth estimation using surface normal constraints. Existing depth estimation methods either suffer from the lack of geometric constraints, or are limited to the difficulty of reliably capturing geometric context, which leads to a bottleneck of depth estimation quality. We therefore introduce a simple yet effective method, named Adaptive Surface Normal (ASN) constraint, to effectively correlate the depth estimation with geometric consistency. Our key idea is to adaptively determine the reliable local geometry from a set of randomly sampled candidates to derive surface normal constraint, for which we measure the consistency of the geometric contextual features. As a result, our method can faithfully reconstruct the 3D geometry and is robust to local shape variations, such as boundaries, sharp corners and noises. We conduct extensive evaluations and comparisons using public datasets. The experimental results demonstrate our method outperforms the state-of-the-art methods and has superior efficiency and robustness. Codes are available at: https://github.com/xxlong0/ASNDepth Xiaoxiao Long, Cheng Lin 0001, Lingjie Liu, Wei Li 0111, Christian Theobalt, Ruigang Yang, Wenping Wang 0001 |
ICCV | 6 |
| 2021 | Image Re-composition via Regional Content-Style DecouplingabstractTypical image composition harmonizes regions from different images to a single plausible image. We extend the idea of image composition by introducing the content-style decomposition and combination to form the concept of image re-composition. In other words, our image re-composition could arbitrarily combine those contents and styles decomposed from different images to generate more diverse images in a unified framework. In the decomposition stage, we incorporate the whitening normalization to obtain a more thorough content-style decoupling, which substantially improves the re-composition results. Moreover, to handle the variation of structure and texture of different objects in an image, we design the network to support regional feature representation and achieve region-aware content-style decomposition. Regarding the composition stage, we propose a cycle consistency loss to constrain the network preserving the content and style information during the composition. Our method can produce diverse re-composition results, including content-content, content-style and style-style. Our experimental results demonstrate a large improvement over the current state-of-the-art methods. Wei Li 0111, Hong Zhang 0009, Ruigang Yang, Weiwei Xu 0003 |
ACM Multimedia | 6 |
| 2021 | Invisible for both Camera and LiDAR: Security of Multi-Sensor Fusion based Perception in Autonomous Driving Under Physical-World AttacksabstractIn Autonomous Driving (AD) systems, perception is both security and safety critical. Despite various prior studies on its security issues, all of them only consider attacks on camera-or LiDAR-based AD perception alone. However, production AD systems today predominantly adopt a Multi-Sensor Fusion (MSF) based design, which in principle can be more robust against these attacks under the assumption that not all fusion sources are (or can be) attacked at the same time. In this paper, we present the first study of security issues of MSF-based perception in AD systems. We directly challenge the basic MSF design assumption above by exploring the possibility of attacking all fusion sources simultaneously. This allows us for the first time to understand how much security guarantee MSF can fundamentally provide as a general defense strategy for AD perception.We formulate the attack as an optimization problem to generate a physically-realizable, adversarial 3D-printed object that misleads an AD system to fail in detecting it and thus crash into it. To systematically generate such a physical-world attack, we propose a novel attack pipeline that addresses two main design challenges: (1) non-differentiable target camera and LiDAR sensing systems, and (2) non-differentiable cell-level aggregated features popularly used in LiDAR-based AD perception. We evaluate our attack on MSF algorithms included in representative open-source industry-grade AD systems in real-world driving scenarios. Our results show that the attack achieves over 90% success rate across different object types and MSF algorithms. Our attack is also found stealthy, robust to victim positions, transferable across MSF algorithms, and physical-world realizable after being 3D-printed and captured by LiDAR and camera devices. To concretely assess the end-to-end safety impact, we further perform simulation evaluation and show that it can cause a 100% vehicle collision rate for an industry-grade AD system. We also evaluate and discuss defense strategies. Ningfei Wang, Chaowei Xiao, Ruigang Yang, Qi Alfred Chen, Mingyan Liu, Bo Li 0026 |
SP | 6 |
| 2021 | Plane Segmentation Based on the Optimal-Vector-Field in LiDAR Point CloudsabstractOne key challenge in the point cloud segmentation is the detection and split of overlapping regions between different planes. The existing methods depend on the similarity and the dissimilarity in neighbor regions without a global constraint, which brings the 'over-' and 'under-' segmentation in the results. Hence, this paper presents a pipeline of the accurate plane segmentation for point clouds to address the shortcoming in the local optimization. There are two phases included in the proposed segmentation process. One is a local phase to calculate connectivity scores between different planes based on local variations of surface normals. In this phase, a new optimal-vector-field is formulated to detect the plane intersections. The optimal-vector-field is large in magnitude at plane intersections and vanishing at other regions. The other one is a global phase to smooth local segmentation cues to mimic leading eigenvector computation in the graph-cut. Evaluation of two datasets shows that the achieved precision and recall is 94.50 percent and 90.81 percent on the collected mobile LiDAR data and obtains an average accuracy of 75.4 percent on an open benchmark, which outperforms the state-of-the-art methods in terms of completeness and correctness. Sheng Xu 0003, Ruisheng Wang 0001, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Gated Path Selection Network for Semantic SegmentationabstractSemantic segmentation is a challenging task that needs to handle large scale variations, deformations, and different viewpoints. In this paper, we develop a novel network named Gated Path Selection Network (GPSNet), which aims to adaptively select receptive fields while maintaining the dense sampling capability. In GPSNet, we first design a two-dimensional SuperNet, which densely incorporates features from growing receptive fields. And then, a Comparative Feature Aggregation (CFA) module is introduced to dynamically aggregate discriminative semantic context. In contrast to previous works that focus on optimizing sparse sampling locations on regular grids, GPSNet can adaptively harvest free form dense semantic context information. The derived adaptive receptive fields and dense sampling locations are data-dependent and flexible which can model various contexts of objects. On two representative semantic segmentation datasets, i.e., Cityscapes and ADE20K, we show that the proposed approach consistently outperforms previous methods without bells and whistles. Qichuan Geng, Hong Zhang 0009, Xiaojuan Qi 0001, Gao Huang 0001, Ruigang Yang, Zhong Zhou |
IEEE Trans. Image Process. | 5 |
| 2021 | SparseFusion: Dynamic Human Avatar Modeling From Sparse RGBD ImagesabstractIn this paper, we propose a novel approach to reconstruct 3D human body shapes based on a sparse set of RGBD frames using a single RGBD camera. We specifically focus on the realistic settings where human subjects move freely during the capture. The main challenge is how to robustly fuse these sparse frames into a canonical 3D model, under pose changes and surface occlusions. This is addressed by our new framework consisting of the following steps. First, based on a generative human template, for every two frames having sufficient overlap, an initial pairwise alignment is performed; It is followed by a global non-rigid registration procedure, in which partial results from RGBD frames are collected into a unified 3D shape, under the guidance of correspondences from the pairwise alignment; Finally, the texture map of the reconstructed human model is optimized to deliver a clear and spatially consistent texture. Empirical evaluations on synthetic and real datasets demonstrate both quantitatively and qualitatively the superior performance of our framework in reconstructing complete 3D human models with high fidelity. It is worth noting that our framework is flexible, with potential applications going beyond shape reconstruction. As an example, we showcase its use in reshaping and reposing to a new avatar. Xinxin Zuo, Sen Wang 0003, Jiangbin Zheng 0001, Minglun Gong, Ruigang Yang, Li Cheng 0001 |
IEEE Trans. Multim. | 6 |
| 2021 | Heter-Sim: Heterogeneous Multi-Agent Systems Simulation by Interactive Data-Driven OptimizationabstractInteractive multi-agent simulation algorithms are used to compute the trajectories and behaviors of different entities in virtual reality scenarios. However, current methods involve considerable parameter tweaking to generate plausible behaviors. We introduce a novel approach (Heter-Sim) that combines physics-based simulation methods with data-driven techniques using an optimization-based formulation. Our approach is general and can simulate heterogeneous agents corresponding to human crowds, traffic, vehicles, or combinations of different agents with varying dynamics. We estimate motion states from real-world datasets that include information about position, velocity, and control direction. Our optimization algorithm considers several constraints, including velocity continuity, collision avoidance, attraction, direction control. Other constraints are implemented by introducing a novel energy function to control the motions of heterogeneous agents. To accelerate the computations, we reduce the search space for both collision avoidance and optimal solution computation. Heter-Sim can simulate tens or hundreds of agents at interactive rates and we compare its accuracy with real-world datasets and prior algorithms. We also perform user studies that evaluate the plausible behaviors generated by our algorithm and a user study that evaluates the plausibility of our algorithm via VR. Jiaping Ren, Yangxi Xiao, Ruigang Yang, Dinesh Manocha, Xiaogang Jin 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2020 | RotPredictor: Unsupervised Canonical Viewpoint Learning for Point Cloud ClassificationabstractRecently, significant progress has been achieved in analyzing the 3D point cloud with deep learning techniques. However, existing networks suffer from poor generalization and robustness to arbitrary rotations applied to the input point cloud. Different from traditional strategies that improve the rotation robustness with data augmentation or specifically designed spherical representation or harmonics-based kernels, we propose to rotate the point cloud into a canonical viewpoint for boosting the following downstream target task, e.g., object classification and part segmentation. Specifically, the canonical viewpoint is predicted by the network RotPredictor in an unsupervised way and the loss function is only built on the target task. Our RotPredictor satisfies the rotation equivariance property in (3) approximately and the predication output has the linear relationship with the applied rotation transformation. In addition, the RotPredictor is an independent plug and play module, which can be employed by any point-based deep learning framework without extra burden. Experimental results on the public model classification dataset ModelNet40 show the performance for all baselines can be boosted by integrating the proposed module. In addition, by adding our proposed module, we can achieve the state-of-the-art classification accuracy with 90.2% on the rotation-augmented ModelNet40 benchmark. Dingfu Zhou, Xibin Song, Shengze Jin, Ruigang Yang, Liangjun Zhang |
3DV | 5 |
| 2020 | CSPN++: Learning Context and Resource Aware Convolutional Spatial Propagation Networks for Depth CompletionabstractDepth Completion deals with the problem of converting a sparse depth map to a dense one, given the corresponding color image. Convolutional spatial propagation network (CSPN) is one of the state-of-the-art (SoTA) methods of depth completion, which recovers structural details of the scene. In this paper, we propose CSPN++, which further improves its effectiveness and efficiency by learning adaptive convolutional kernel sizes and the number of iterations for the propagation, thus the context and computational resource needed at each pixel could be dynamically assigned upon requests. Specifically, we formulate the learning of the two hyper-parameters as an architecture selection problem where various configurations of kernel sizes and numbers of iterations are first defined, and then a set of soft weighting parameters are trained to either properly assemble or select from the pre-defined configurations at each pixel. In our experiments, we find weighted assembling can lead to significant accuracy improvements, which we referred to as "context-aware CSPN", while weighted selection, "resource-aware CSPN" can reduce the computational resource significantly with similar or better accuracy. Besides, the resource needed for CSPN++ can be adjusted w.r.t. the computational budget automatically. Finally, to avoid the side effects of noise or inaccurate sparse depths, we embed a gated network inside CSPN++, which further improves the performance. We demonstrate the effectiveness of CSPN++ on the KITTI depth completion benchmark, where it significantly improves over CSPN and other SoTA methods 1. Xinjing Cheng, Peng Wang 0001, Chenye Guan, Ruigang Yang |
AAAI | 4 |
| 2020 | AutoRemover: Automatic Object Removal for Autonomous Driving VideosabstractMotivated by the need for photo-realistic simulation in autonomous driving, in this paper we present a video inpainting algorithm AutoRemover, designed specifically for generating street-view videos without any moving objects. In our setup we have two challenges: the first is the shadow, shadows are usually unlabeled but tightly coupled with the moving objects. The second is the large ego-motion in the videos. To deal with shadows, we build up an autonomous driving shadow dataset and design a deep neural network to detect shadows automatically. To deal with large ego-motion, we take advantage of the multi-source data, in particular the 3D data, in autonomous driving. More specifically, the geometric relationship between frames is incorporated into an inpainting deep neural network to produce high-quality structurally consistent video output. Experiments show that our method outperforms other state-of-the-art (SOTA) object removal algorithms, reducing the RMSE by over 19%. Wei Li 0111, Peng Wang 0001, Chenye Guan, Yuhang Song 0003, Baoquan Chen, Weiwei Xu 0003, Ruigang Yang |
AAAI | 10 |
| 2020 | Speech2Video Synthesis with 3D Skeleton Regularization and Expressive Body Poses
Miao Liao, Peng Wang 0001, Hao Zhu 0004, Xinxin Zuo, Ruigang Yang |
ACCV (5) | 6 |
| 2020 | 3D Part Guided Image Editing for Fine-Grained Object UnderstandingabstractHolistically understanding an object with its 3D movable parts is essential for visual models of a robot to interact with the world. For example, only by understanding many possible part dynamics of other vehicles (e.g., door or trunk opening, taillight blinking for changing lane), a self-driving vehicle can be success in dealing with emergency cases. However, existing visual models tackle rarely on these situations, but focus on bounding box detection. In this paper, we fill this important missing piece in autonomous driving by solving two critical issues. First, for dealing with data scarcity, we propose an effective training data generation process by fitting a 3D car model with dynamic parts to cars in real images. This allows us to directly edit the real images using the aligned 3D parts, yielding effective training data for learning robust deep neural networks (DNNs). Secondly, to benchmark the quality of 3D part understanding, we collected a large dataset in real driving scenario with cars in uncommon states (CUS), i.e. with door or trunk opened etc., which demonstrates that our trained network with edited images largely outperforms other baselines in terms of 2D detection and instance segmentation accuracy. Zongdai Liu, Feixiang Lu, Peng Wang 0001, Liangjun Zhang, Ruigang Yang |
CVPR | 6 |
| 2020 | Channel Attention Based Iterative Residual Learning for Depth Map Super-ResolutionabstractDespite the remarkable progresses made in deep learning based depth map super-resolution (DSR), how to tackle real-world degradation in low-resolution (LR) depth maps remains a major challenge. Existing DSR model is generally trained and tested on synthetic dataset, which is very different from what would get from a real depth sensor. In this paper, we argue that DSR models trained under this setting are restrictive and not effective in dealing with realworld DSR tasks. We make two contributions in tackling real-world degradation of different depth sensors. First, we propose to classify the generation of LR depth maps into two types: non-linear downsampling with noise and interval downsampling, for which DSR models are learned correspondingly. Second, we propose a new framework for real-world DSR, which consists of four modules : 1) An iterative residual learning module with deep supervision to learn effective high-frequency components of depth maps in a coarse-to-fine manner; 2) A channel attention strategy to enhance channels with abundant high-frequency components; 3) A multi-stage fusion module to effectively reexploit the results in the coarse-to-fine process; and 4) A depth refinement module to improve the depth map by TGV regularization and input loss. Extensive experiments on benchmarking datasets demonstrate the superiority of our method over current state-of-the-art DSR methods. Xibin Song, Yuchao Dai, Dingfu Zhou, Liu Liu 0009, Wei Li 0111, Hongdong Li, Ruigang Yang |
CVPR | 7 |
| 2020 | FaceScape: A Large-Scale High Quality 3D Face Dataset and Detailed Riggable 3D Face PredictionabstractIn this paper, we present a large-scale detailed 3D face dataset, FaceScape, and propose a novel algorithm that is able to predict elaborate riggable 3D face models from a single image input. FaceScape dataset provides 18,760 textured 3D faces, captured from 938 subjects and each with 20 specific expressions. The 3D models contain the pore-level facial geometry that is also processed to be topologically uniformed. These fine 3D facial models can be represented as a 3D morphable model for rough shapes and displacement maps for detailed geometry. Taking advantage of the large-scale and high-accuracy dataset, a novel algorithm is further proposed to learn the expression-specific dynamic details using a deep neural network. The learned relationship serves as the foundation of our 3D face prediction system from a single image input. Different than the previous methods, our predicted 3D models are riggable with highly detailed geometry under different expressions. The unprecedented dataset and code will be released to public for research purpose. Hao Zhu 0004, Mingkai Huang, Qiu Shen, Ruigang Yang, Xun Cao |
CVPR | 6 |
| 2020 | LiDAR-Based Online 3D Video Object Detection With Graph-Based Message Passing and Spatiotemporal Transformer AttentionabstractExisting LiDAR-based 3D object detectors usually focus on the single-frame detection, while ignoring the spatiotemporal information in consecutive point cloud frames. In this paper, we propose an end-to-end online 3D video object detector that operates on point cloud sequences. The proposed model comprises a spatial feature encoding component and a spatiotemporal feature aggregation component. In the former component, a novel Pillar Message Passing Network (PMPNet) is proposed to encode each discrete point cloud frame. It adaptively collects information for a pillar node from its neighbors by iterative message passing, which effectively enlarges the receptive field of the pillar feature. In the latter component, we propose an Attentive Spatiotemporal Transformer GRU (AST-GRU) to aggregate the spatiotemporal information, which enhances the conventional ConvGRU with an attentive memory gating mechanism. AST-GRU contains a Spatial Transformer Attention (STA) module and a Temporal Transformer Attention (TTA) module, which can emphasize the foreground objects and align the dynamic objects, respectively. Experimental results demonstrate that the proposed 3D video object detector achieves state-of-the-art performance on the large-scale nuScenes benchmark. Junbo Yin, Jianbing Shen, Chenye Guan, Dingfu Zhou, Ruigang Yang |
CVPR | 5 |
| 2020 | A Unified Object Motion and Affinity Model for Online Multi-Object TrackingabstractCurrent popular online multi-object tracking (MOT) solutions apply single object trackers (SOTs) to capture object motions, while often requiring an extra affinity network to associate objects, especially for the occluded ones. This brings extra computational overhead due to repetitive feature extraction for SOT and affinity computation. Meanwhile, the model size of the sophisticated affinity network is usually non-trivial. In this paper, we propose a novel MOT framework that unifies object motion and affinity model into a single network, named UMA, in order to learn a compact feature that is discriminative for both object motion and affinity measure. In particular, UMA integrates single object tracking and metric learning into a unified triplet network by means of multi-task learning. Such design brings advantages of improved computation efficiency, low memory requirement and simplified training procedure. In addition, we equip our model with a task-specific attention module, which is used to boost task-aware feature learning. The proposed UMA can be easily trained end-to-end, and is elegant - requiring only one training stage. Experimental results show that it achieves promising performance on several MOT Challenge benchmarks. Junbo Yin, Wenguan Wang, Qinghao Meng, Ruigang Yang, Jianbing Shen |
CVPR | 4 |
| 2020 | Joint 3D Instance Segmentation and Object Detection for Autonomous DrivingabstractCurrently, in Autonomous Driving (AD), most of the 3D object detection frameworks (either anchor- or anchor-free-based) consider the detection as a Bounding Box (BBox) regression problem. However, this compact representation is not sufficient to explore all the information of the objects. To tackle this problem, we propose a simple but practical detection framework to jointly predict the 3D BBox and instance segmentation. For instance segmentation, we propose a Spatial Embeddings (SEs) strategy to assemble all foreground points into their corresponding object centers. Base on the SE results, the object proposals can be generated based on a simple clustering strategy. For each cluster, only one proposal is generated. Therefore, the Non-Maximum Suppression (NMS) process is no longer needed here. Finally, with our proposed instance-aware ROI pooling, the BBox is refined by a second-stage network. Experimental results on the public KITTI dataset show that the proposed SEs can significantly improve the instance segmentation results compared with other feature embedding-based method. Meanwhile, it also outperforms most of the 3D object detectors on the KITTI testing benchmark. Dingfu Zhou, Xibin Song, Liu Liu 0009, Junbo Yin, Yuchao Dai, Hongdong Li, Ruigang Yang |
CVPR | 8 |
| 2020 | DVI: Depth Guided Video Inpainting for Autonomous Driving
Miao Liao, Feixiang Lu, Dingfu Zhou, Wei Li 0111, Ruigang Yang |
ECCV (21) | 6 |
| 2020 | AutoTrajectory: Label-Free Trajectory Extraction and Prediction from Videos Using Dynamic Points
Yuexin Ma, Xinge Zhu, Xinjing Cheng, Ruigang Yang, Jiming Liu 0001, Dinesh Manocha |
ECCV (13) | 4 |
| 2020 | Domain-Invariant Stereo Matching Networks
Feihu Zhang, Xiaojuan Qi 0001, Ruigang Yang, Victor Adrian Prisacariu, Benjamin W. Wah, Philip Torr 0001 |
ECCV (2) | 3 |
| 2020 | Angus Cattle Recognition Using Deep LearningabstractAngus cattle have significant economical values. Individualized management is expected to improve the efficiency and prevent financial loss in the farming industry. However Angus cattle, being all black, are a challenging case for visual recognition. We present a system for image segmentation and identification on Angus cattle using deep learning methods. Two databases of cattle were first collected and annotated, one is frontal face only, captured in a lab setting with controlled lighting and pose in the same day. The second was captured in a farm with natural light and background at three different days. The full body of cattle is captured from different angles. Using three popular neutral networks: PrimNet, VGG16 and ResNet50, we have evaluated a number of design choices for cattle identifications, including face only, face + body, and with/without background segmentation. The best result is obtained using face + body image without background, achieving 85.45% accuracy with the VGG16 net. If we use images captured under different days as training and testing datasets, the accuracy drops dramatically below 10%. It remains as a challenging open problem to be resolved. Shunnan Chen, Sen Wang 0003, Xinxin Zuo, Ruigang Yang |
ICPR | 4 |
| 2020 | Omnidirectional Depth Extension NetworksabstractOmnidirectional 360° camera proliferates rapidly for autonomous robots since it significantly enhances the perception ability by widening the field of view (FoV). However, corresponding 360° depth sensors, which are also critical for the perception system, are still difficult or expensive to have. In this paper, we propose a low-cost 3D sensing system that combines an omnidirectional camera with a calibrated projective depth camera, where the depth from the limited FoV can be automatically extended to the rest of recorded omnidirectional image. To accurately recover the missing depths, we design an omnidirectional depth extension convolutional neural network (ODE-CNN), in which a spherical feature transform layer (SFTL) is embedded at the end of feature encoding layers, and a deformable convolutional spatial propagation network (D-CSPN) is appended at the end of feature decoding layers. The former re-samples the neighborhood of each pixel in the omnidirectional coordination to the projective coordination, which reduce the difficulty of feature learning, and the later automatically finds a proper context to well align the structures in the estimated depths via CNN w.r.t. the reference image, which significantly improves the visual quality. Finally, we demonstrate the effectiveness of proposed ODE-CNN over the popular 360D dataset, and show that ODE-CNN significantly outperforms (relatively 33% reduction in depth error) other state-of-the-art (SoTA) methods. Xinjing Cheng, Peng Wang 0001, Yanqi Zhou, Chenye Guan, Ruigang Yang |
ICRA | 5 |
| 2020 | Learning Resilient Behaviors for Navigation Under UncertaintyabstractDeep reinforcement learning has great potential to acquire complex, adaptive behaviors for autonomous agents automatically. However, the underlying neural network polices have not been widely deployed in real-world applications, especially in these safety-critical tasks (e.g., autonomous driving). One of the reasons is that the learned policy cannot perform flexible and resilient behaviors as traditional methods to adapt to diverse environments. In this paper, we consider the problem that a mobile robot learns adaptive and resilient behaviors for navigating in unseen uncertain environments while avoiding collisions. We present a novel approach for uncertainty-aware navigation by introducing an uncertainty-aware predictor to model the environmental uncertainty, and we propose a novel uncertainty-aware navigation network to learn resilient behaviors in the prior unknown environments. To train the proposed uncertainty-aware network more stably and efficiently, we present the temperature decay training paradigm, which balances exploration and exploitation during the training process. Our experimental evaluation demonstrates that our approach can learn resilient behaviors in diverse environments and generate adaptive trajectories according to environmental uncertainties. Tingxiang Fan, Pinxin Long, Wenxi Liu, Jia Pan 0001, Ruigang Yang, Dinesh Manocha |
ICRA | 5 |
| 2020 | Instance Segmentation of LiDAR Point CloudsabstractWe propose a robust baseline method for instance segmentation which are specially designed for large-scale outdoor LiDAR point clouds. Our method includes a novel dense feature encoding technique, allowing the localization and segmentation of small, far-away objects, a simple but effective solution for single-shot instance prediction and effective strategies for handling severe class imbalances. Since there is no public dataset for the study of LiDAR instance segmentation, we also build a new publicly available LiDAR point cloud dataset to include both precise 3D bounding box and point-wise labels for instance segmentation, while still being about 3~20 times as large as other existing LiDAR datasets. The dataset will be published at https://github.com/feihuzhang/LiDARSeg. Feihu Zhang, Chenye Guan, Song Bai 0001, Ruigang Yang, Philip Torr 0001, Victor Adrian Prisacariu |
ICRA | 5 |
| 2020 | InstanceFusion: Real-time Instance-level 3D Reconstruction Using a Single RGBD CameraabstractAbstract We present InstanceFusion, a robust real‐time system to detect, segment, and reconstruct instance‐level 3D objects of indoor scenes with a hand‐held RGBD camera. It combines the strengths of deep learning and traditional SLAM techniques to produce visually compelling 3D semantic models. The key success comes from our novel segmentation scheme and the efficient instance‐level data fusion, which are both implemented on GPU. Specifically, for each incoming RGBD frame, we take the advantages of the RGBD features, the 3D point cloud, and the reconstructed model to perform instance‐level segmentation. The corresponding RGBD data along with the instance ID are then fused to the surfel‐based models. In order to sufficiently store and update these data, we design and implement a new data structure using the OpenGL Shading Language. Experimental results show that our method advances the state‐of‐the‐art (SOTA) methods in instance segmentation and data fusion by a big margin. In addition, our instance segmentation improves the precision of 3D reconstruction, especially in the loop closure. InstanceFusion system runs 20.5Hz on a consumer‐level GPU, which supports a number of augmented reality (AR) applications (e.g., 3D model registration, virtual interaction, AR map) and robot applications (e.g., navigation, manipulation, grasping). To facilitate future research and reproduce our system more easily, the source code, data, and the trained model are released on Github: https://github.com/Fancomi2017/InstanceFusion . Feixiang Lu, Haotian Peng, Xinhang Yang, Ruizhi Cao, Liangjun Zhang, Ruigang Yang |
Comput. Graph. Forum | 8 |
| 2020 | Learning Depth with Convolutional Spatial Propagation NetworkabstractIn this paper, we propose the convolutional spatial propagation network (CSPN) and demonstrate its effectiveness for various depth estimation tasks. CSPN is a simple and efficient linear propagation model, where the propagation is performed with a manner of recurrent convolutional operations, in which the affinity among neighboring pixels is learned through a deep convolutional neural network (CNN). Compare to the previous state-of-the-art (SOTA) linear propagation model, i.e., spatial propagation networks (SPN), CSPN is 2 to 5× faster in practice. We concatenate CSPN and its variants to SOTA depth estimation networks, which significantly improve the depth accuracy. Specifically, we apply CSPN to two depth estimation problems: depth completion and stereo matching, in which we design modules which adapts the original 2D CSPN to embed sparse depth samples during the propagation, operate with 3D convolution and be synergistic with spatial pyramid pooling. In our experiments, we show that all these modules contribute to the final performance. For the task of depth completion, our method reduce the depth error over 30 percent in the NYU v2 and KITTI datasets. For the task of stereo matching, our method currently ranks 1st on both the KITTI Stereo 2012 and 2015 benchmarks. Xinjing Cheng, Peng Wang 0001, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | The ApolloScape Open Dataset for Autonomous Driving and Its ApplicationabstractAutonomous driving has attracted tremendous attention especially in the past few years. The key techniques for a self-driving car include solving tasks like 3D map construction, self-localization, parsing the driving road and understanding objects, which enable vehicles to reason and act. However, large scale data set for training and system evaluation is still a bottleneck for developing robust perception models. In this paper, we present the ApolloScape dataset [1] and its applications for autonomous driving. Compared with existing public datasets from real scenes, e.g., KITTI [2] or Cityscapes [3] , ApolloScape contains much large and richer labelling including holistic semantic dense point cloud for each site, stereo, per-pixel semantic labelling, lanemark labelling, instance segmentation, 3D car instance, high accurate location for every frame in various driving videos from multiple sites, cities and daytimes. For each task, it contains at lease 15x larger amount of images than SOTA datasets. To label such a complete dataset, we develop various tools and algorithms specified for each task to accelerate the labelling process, such as joint 3D-2D segment labeling, active labelling in videos etc. Depend on ApolloScape, we are able to develop algorithms jointly consider the learning and inference of multiple tasks. In this paper, we provide a sensor fusion scheme integrating camera videos, consumer-grade motion sensors (GPS/IMU), and a 3D semantic map in order to achieve robust self-localization and semantic segmentation for autonomous driving. We show that practically, sensor fusion and joint learning of multiple tasks are beneficial to achieve a more robust and accurate system. We expect our dataset and proposed relevant algorithms can support and motivate researchers for further development of multi-sensor fusion and multi-task learning in the field of computer vision. Xinyu Huang 0001, Peng Wang 0001, Xinjing Cheng, Dingfu Zhou, Qichuan Geng, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2020 | Inferring Salient Objects from Human FixationsabstractPrevious research in visual saliency has been focused on two major types of models namely fixation prediction and salient object detection. The relationship between the two, however, has been less explored. In this work, we propose to employ the former model type to identify salient objects. We build a novel Attentive Saliency Network (ASNet)11.Available at: https://github.com/wenguanwang/ASNet. that learns to detect salient objects from fixations. The fixation map, derived at the upper network layers, mimics human visual attention mechanisms and captures a high-level understanding of the scene from a global view. Salient object detection is then viewed as fine-grained object-level saliency segmentation and is progressively optimized with the guidance of the fixation map in a top-down manner. ASNet is based on a hierarchy of convLSTMs that offers an efficient recurrent mechanism to sequentially refine the saliency features over multiple steps. Several loss functions, derived from existing saliency evaluation metrics, are incorporated to further boost the performance. Extensive experiments on several challenging datasets show that our ASNet outperforms existing methods and is capable of generating accurate segmentation maps with the help of the computed fixation prior. Our work offers a deeper insight into the mechanisms of attention and narrows the gap between salient object detection and fixation prediction. Wenguan Wang, Jianbing Shen, Xingping Dong, Ali Borji, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2020 | Detailed Surface Geometry and Albedo Recovery from RGB-D Video under Natural IlluminationabstractThis article presents a novel approach for depth map enhancement from an RGB-D video sequence. The basic idea is to exploit the photometric information in the color sequence to resolve the inherent ambiguity of shape from shading problem. Instead of making any assumption about surface albedo or controlled object motion and lighting, we use the lighting variations introduced by casual object movement. We are effectively calculating photometric stereo from a moving object under natural illuminations. One of the key technical challenges is to establish correspondences over the entire image set. We, therefore, develop a lighting insensitive robust pixel matching technique that out-performs optical flow method in presence of lighting variations. An adaptive reference frame selection procedure is introduced to get more robust to imperfect lambertian reflections. In addition, we present an expectation-maximization framework to recover the surface normal and albedo simultaneously, without any regularization term. We have validated our method on both synthetic and real datasets to show its superior performance on both surface details recovery and intrinsic decomposition. Xinxin Zuo, Sen Wang 0003, Jiangbin Zheng 0001, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2020 | Interactive free-viewpoint video generationabstractFree-viewpoint video (FVV) is processed video content in which viewers can freely select the viewing position and angle. FVV delivers an improved visual experience and can also help synthesize special effects and virtual reality content. In this paper, a complete FVV system is proposed to interactively control the viewpoints of video relay programs through multimedia terminals such as computers and tablets. The hardware of the FVV generation system is a set of synchronously controlled cameras, and the software generates videos in novel viewpoints from the captured video using view interpolation. The interactive interface is designed to visualize the generated video in novel viewpoints and enable the viewpoint to be changed interactively. Experiments show that our system can synthesize plausible videos in intermediate viewpoints with a view range of up to 180°. Hao Zhu 0004, Wei Li 0111, Xun Cao, Ruigang Yang |
Virtual Real. Intell. Hardw. | 6 |
| 2019 | IoU Loss for 2D/3D Object DetectionabstractIn the 2D/3D object detection task, Intersection-over-Union (IoU) has been widely employed as an evaluation metric to evaluate the performance of different detectors in the testing stage. However, during the training stage, the common distance loss (e.g, L_1 or L_2) is often adopted as the loss function to minimize the discrepancy between the predicted and ground truth Bounding Box (Bbox). To eliminate the performance gap between training and testing, the IoU loss has been introduced for 2D object detection in [1] and [2]. Unfortunately, all these approaches only work for axis-aligned 2D Boxes, which cannot be applied for more general object detection task with rotated Boxes. To resolve this issue, we investigate the IoU computation for two rotated Boxes first and then implement a unified framework, IoU loss layer for both 2D and 3D object detection tasks. By integrating the implemented IoU loss into several state-of-the-art 3D object detectors, consistent improvements have been achieved for both bird-eye-view 2D detection and point cloud 3D detection on the public KITTI [3] benchmark. Dingfu Zhou, Xibin Song, Chenye Guan, Junbo Yin, Yuchao Dai, Ruigang Yang |
3DV | 7 |
| 2019 | TrafficPredict: Trajectory Prediction for Heterogeneous Traffic-AgentsabstractTo safely and efficiently navigate in complex urban traffic, autonomous vehicles must make responsible predictions in relation to surrounding traffic-agents (vehicles, bicycles, pedestrians, etc.). A challenging and critical task is to explore the movement patterns of different traffic-agents and predict their future trajectories accurately to help the autonomous vehicle make reasonable navigation decision. To solve this problem, we propose a long short-term memory-based (LSTM-based) realtime traffic prediction algorithm, TrafficPredict. Our approach uses an instance layer to learn instances’ movements and interactions and has a category layer to learn the similarities of instances belonging to the same type to refine the prediction. In order to evaluate its performance, we collected trajectory datasets in a large city consisting of varying conditions and traffic densities. The dataset includes many challenging scenarios where vehicles, bicycles, and pedestrians move among one another. We evaluate the performance of TrafficPredict on our new dataset and highlight its higher accuracy for trajectory prediction by comparing with prior prediction methods. Yuexin Ma, Xinge Zhu, Ruigang Yang, Wenping Wang 0001, Dinesh Manocha |
AAAI | 4 |
| 2019 | Detailed Human Shape Estimation From a Single Image by Hierarchical Mesh DeformationabstractThis paper presents a novel framework to recover detailed human body shapes from a single image. It is a challenging task due to factors such as variations in human shapes, body poses, and viewpoints. Prior methods typically attempt to recover the human body shape using a parametric based template that lacks the surface details. As such the resulting body shape appears to be without clothing. In this paper, we propose a novel learning-based framework that combines the robustness of parametric model with the flexibility of free-form 3D deformation. We use the deep neural networks to refine the 3D shape in a Hierarchical Mesh Deformation (HMD) framework, utilizing the constraints from body joints, silhouettes, and per-pixel shading information. We are able to restore detailed human body shapes beyond skinned models. Experiments demonstrate that our method has outperformed previous state-of-the-art approaches, achieving better accuracy in terms of both 2D IoU number and 3D metric distance. The code is available in https://github.com/zhuhao-nju/hmd.git. Hao Zhu 0004, Xinxin Zuo, Sen Wang 0003, Xun Cao, Ruigang Yang |
CVPR | 5 |
| 2019 | ApolloCar3D: A Large 3D Car Instance Understanding Benchmark for Autonomous DrivingabstractAutonomous driving has attracted remarkable attention from both industry and academia. An important task is to estimate 3D properties (e.g. translation, rotation and shape) of a moving or parked vehicle on the road. This task, while critical, is still under-researched in the computer vision community – partially owing to the lack of large scale and fully-annotated 3D car database suitable for autonomous driving research. In this paper, we contribute the first large scale database suitable for 3D car instance understanding – ApolloCar3D. The dataset contains 5,277 driving images and over 60K car instances, where each car is fitted with an industry-grade 3D CAD model with absolute model size and semantically labelled keypoints. This dataset is above 20× larger than PASCAL3D+ and KITTI, the current state-of-the-art. To enable efficient labelling in 3D, we build a pipeline by considering 2D-3D keypoint correspondences for a single instance and 3D relationship among multiple instances. Equipped with such dataset, we build various baseline algorithms with the state-of-the-art deep convolutional neural networks. Specifically, we first segment each car with a pre-trained Mask R-CNN, and then regress towards its 3D pose and shape based on a deformable 3D car model with or without using semantic keypoints. We show that using keypoints significantly improves fitting performance. Finally, we develop a new 3D metric jointly considering 3D pose and 3D shape, allowing for comprehensive evaluation and ablation study. Xibin Song, Peng Wang 0001, Dingfu Zhou, Chenye Guan, Yuchao Dai, Hongdong Li, Ruigang Yang |
CVPR | 9 |
| 2019 | GA-Net: Guided Aggregation Net for End-To-End Stereo MatchingabstractIn the stereo matching task, matching cost aggregation is crucial in both traditional methods and deep neural network models in order to accurately estimate disparities. We propose two novel neural net layers, aimed at capturing local and the whole-image cost dependencies respectively. The first is a semi-global aggregation layer which is a differentiable approximation of the semi-global matching, the second is the local guided aggregation layer which follows a traditional cost filtering strategy to refine thin structures. These two layers can be used to replace the widely used 3D convolutional layer which is computationally costly and memory-consuming as it has cubic computational/memory complexity. In the experiments, we show that nets with a two-layer guided aggregation block easily outperform the state-of-the-art GC-Net which has nineteen 3D convolutional layers. We also train a deep guided aggregation network (GA-Net) which gets better accuracies than state-of-the-art methods on both Scene Flow dataset and KITTI benchmarks. Feihu Zhang, Victor Adrian Prisacariu, Ruigang Yang, Philip Torr 0001 |
CVPR | 3 |
| 2019 | Improved Techniques for Training Adaptive Deep NetworksabstractAdaptive inference is a promising technique to improve the computational efficiency of deep models at test time. In contrast to static models which use the same computation graph for all instances, adaptive networks can dynamically adjust their structure conditioned on each input. While existing research on adaptive inference mainly focuses on designing more advanced architectures, this paper investigates how to train such networks more effectively. Specifically, we consider a typical adaptive deep network with multiple intermediate classifiers. We present three techniques to improve its training efficacy from two aspects: 1) a Gradient Equilibrium algorithm to resolve the conflict of learning of different classifiers; 2) an Inline Subnetwork Collaboration approach and a One-for-all Knowledge Distillation algorithm to enhance the collaboration among classifiers. On multiple datasets (CIFAR-10, CIFAR-100 and ImageNet), we show that the proposed approach consistently leads to further improved efficiency on top of state-of-the-art adaptive deep networks. Hao Li 0069, Hong Zhang 0009, Xiaojuan Qi 0001, Ruigang Yang, Gao Huang 0001 |
ICCV | 4 |
| 2019 | Compact Reachability Map for Excavator Motion PlanningabstractIn this paper, we propose a novel compact reachability map representation for excavator motion planning. The constructed reachability map can concisely encode the bucket’s reachable pose and the translation capability limited by excavator’s kinematic structure. By explicitly exploiting the property that the basic excavation motion lies on the excavation plane determined by excavator links, we further reduce the construction of the map from 3D Euclidean space to 2D excavation plane. We show the pre-computed reachability map can be used to develop new excavator motion planning approach. By indexing on the pre-computed reachability map, we can efficiently compute the feasible full-bucket trajectory for single step excavation operation. We highlight the results of the reachability map construction and demonstrate the simulation results of motion planning using a commercial dynamic simulator. Yajue Yang, Liangjun Zhang, Xinjing Cheng, Jia Pan 0001, Ruigang Yang |
IROS | 5 |
| 2019 | Mask-off: Synthesizing Face Images in the Presence of Head-mounted DisplaysabstractWearable VR/AR devices provide users with fully immersive experience in a virtual environment, enabling possibilities to reshape the forms of entertainment and telepresence. While the body language is a crucial element in effective communication, wearing a head-mounted display (HMD) could severely hinder the eye contact and block facial expressions. We present a novel headset removal technique that enables high-quality occlusion-free communication in virtual environment. In particular, our solution synthesizes photoreal faces in the occluded region with faithful reconstruction of facial expressions and eye movements. Towards this goal, we develop a novel capture setup that consists of two near-infrared (NIR) cameras inside the HMD for eye capturing and one external RGB camera for recording visible face regions. To enable realistic face synthesis with consistent illuminations, we propose a data-driven approach to fuse the narrow-field-of-view NIR images with the RGB image captured from the external camera. In addition, to generate pho-torealistic eyes, a dedicated algorithm is proposed to colorize the NIR eye images and further rectify the color distortion caused by the non-linear mapping of IR light sensitivity. Experimental results demonstrate that our framework is capable to synthesize high-fidelity unoccluded facial images with accurate tracking of head motion, facial expression and eye movement. Qingguo Xu, Weikai Chen 0001, Jun Xing, Xinyu Huang 0001, Ruigang Yang |
VR | 7 |
| 2019 | Semi-Supervised Video Object Segmentation with Super-TrajectoriesabstractWe introduce a semi-supervised video segmentation approach based on an efficient video representation, called as "super-trajectory". A super-trajectory corresponds to a group of compact point trajectories that exhibit consistent motion patterns, similar appearances, and close spatiotemporal relationships. We generate the compact trajectories using a probabilistic model, which enables handling of occlusions and drifts effectively. To reliably group point trajectories, we adopt a modified version of the density peaks based clustering algorithm that allows capturing rich spatiotemporal relations among trajectories in the clustering process. We incorporate two intuitive mechanisms for segmentation, called as reverse-tracking and object re-occurrence, for robustness and boosting the performance. Building on the proposed video representation, our segmentation method is discriminative enough to accurately propagate the initial annotations in the first frame onto the remaining frames. Our extensive experimental analyses on three challenging benchmarks demonstrate that, given the annotation in the first frame, our method is capable of extracting the target objects from complex backgrounds, and even reidentifying them after prolonged occlusions, producing high-quality video object segments. Wenguan Wang, Jianbing Shen, Fatih Porikli, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | Fast Texture Mapping Adjustment via Local/Global OptimizationabstractThis paper deals with the texture mapping of a triangular mesh model given a set of calibrated images. Different from the traditional approach of applying projective texture mapping with model parameterizations, we develop an image-space texture optimization scheme that aims to reduce visible seams or misalignment at texture or depth boundaries. Our novel scheme starts with an efficient local (and parallel) texture adjustment scheme at these boundaries, followed by a global correction step to rectify potential texture distortions caused by the local movement. Our phased optimization scheme achieves 50$\sim$∼100 times speed up on GPU (or 6× on CPU) compared to previous state-of-the-art methods. Experiments on a variety of models showed that we achieve this significant speed-up without sacrificing texture quality. Our approach significantly improves resilience to modeling and calibration errors, thereby allowing fast and fully automatic creation of textured models using commodity depth sensors by untrained users. Wei Li 0111, Huajun Gong, Ruigang Yang |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2018 | Identity Preserving Face Completion for Large Ocular Region Occlusion
Weikai Chen 0001, Jun Xing, Xiaoming Li 0002, Zachary Bessinger, Fuchang Liu, Wangmeng Zuo, Ruigang Yang |
BMVC | 8 |
| 2018 | View Extrapolation of Human Body From a Single ImageabstractWe study how to synthesize novel views of human body from a single image. Though recent deep learning based methods work well for rigid objects, they often fail on objects with large articulation, like human bodies. The core step of existing methods is to fit a map from the observable views to novel views by CNNs; however, the rich articulation modes of human body make it rather challenging for CNNs to memorize and interpolate the data well. To address the problem, we propose a novel deep learning based pipeline that explicitly estimates and leverages the geometry of the underlying human body. Our new pipeline is a composition of a shape estimation network and an image generation network, and at the interface a perspective transformation is applied to generate a forward flow for pixel value transportation. Our design is able to factor out the space of data variation and makes learning at each step much easier. Empirically, we show that the performance for pose-varying objects can be improved dramatically. Our method can also be applied on real data captured by 3D sensors, and the flow generated by our methods can be used for generating high quality results in higher resolution. Hao Zhu 0004, Peng Wang 0001, Xun Cao, Ruigang Yang |
CVPR | 5 |
| 2018 | DeLS-3D: Deep Localization and Segmentation With a 3D Semantic MapabstractFor applications such as augmented reality, autonomous driving, self-localization/camera pose estimation and scene parsing are crucial technologies. In this paper, we propose a unified framework to tackle these two problems simultaneously. The uniqueness of our design is a sensor fusion scheme which integrates camera videos, motion sensors (GPS/IMU), and a 3D semantic map in order to achieve robustness and efficiency of the system. Specifically, we first have an initial coarse camera pose obtained from consumer-grade GPS/IMU, based on which a label map can be rendered from the 3D semantic map. Then, the rendered label map and the RGB image are jointly fed into a pose CNN, yielding a corrected camera pose. In addition, to incorporate temporal information, a multi-layer recurrent neural network (RNN) is further deployed improve the pose accuracy. Finally, based on the pose from RNN, we render a new label map, which is fed together with the RGB image into a segment CNN which produces perpixel semantic label. In order to validate our approach, we build a dataset with registered 3D point clouds and video camera images. Both the point clouds and the images are semantically-labeled. Each video frame has ground truth pose from highly accurate motion sensors. We show that practically, pose estimation solely relying on images like PoseNet [25] may fail due to street view confusion, and it is important to fuse multiple sensors. Finally, various ablation studies are performed, which demonstrate the effectiveness of the proposed system. In particular, we show that scene parsing and pose estimation are mutually beneficial to achieve a more robust and accurate system. Peng Wang 0001, Ruigang Yang, Binbin Cao, Wei Xu 0017, Yuanqing Lin |
CVPR | 2 |
| 2018 | Depth Estimation via Affinity Learned with Convolutional Spatial Propagation Network
Xinjing Cheng, Peng Wang 0001, Ruigang Yang |
ECCV (16) | 3 |
| 2018 | Learning Warped Guidance for Blind Face Restoration
Xiaoming Li 0002, Ming Liu 0018, Yuting Ye, Wangmeng Zuo, Liang Lin 0004, Ruigang Yang |
ECCV (13) | 6 |
| 2018 | Saliency-Aware Video Object SegmentationabstractVideo saliency, aiming for estimation of a single dominant object in a sequence, offers strong object-level cues for unsupervised video object segmentation. In this paper, we present a geodesic distance based technique that provides reliable and temporally consistent saliency measurement of superpixels as a prior for pixel-wise labeling. Using undirected intra-frame and inter-frame graphs constructed from spatiotemporal edges or appearance and motion, and a skeleton abstraction step to further enhance saliency estimates, our method formulates the pixel-wise segmentation task as an energy minimization problem on a function that consists of unary terms of global foreground and background models, dynamic location models, and pairwise terms of label smoothness potentials. We perform extensive quantitative and qualitative experiments on benchmark datasets. Our method achieves superior performance in comparison to the current state-of-the-art in terms of accuracy and speed. Wenguan Wang, Jianbing Shen, Ruigang Yang, Fatih Porikli |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | 3D Reconstruction in the Presence of Glass and Mirrors by Acoustic and Visual FusionabstractWe present a practical and inexpensive method to reconstruct 3D scenes that include transparent and mirror objects. Our work is motivated by the need for automatically generating 3D models of interior scenes, which commonly include glass. These large structures are often invisible to cameras or even to our human visual system. Existing 3D reconstruction methods for transparent objects are usually not applicable in such a room-sized reconstruction setting. Our simple hardware setup augments a regular depth camera (e.g., the Microsoft Kinect camera) with a single ultrasonic sensor, which is able to measure the distance to any object, including transparent surfaces. The key technical challenge is the sparse sampling rate from the acoustic sensor, which only takes one point measurement per frame. To address this challenge, we take advantage of the fact that the large scale glass structures in indoor environments are usually either piece-wise planar or a simple parametric surface. Based on these assumptions, we have developed a novel sensor fusion algorithm that first segments the (hybrid) depth map into different categories such as opaque/transparent/infinity (e.g., too far to measure) and then updates the depth map based on the segmentation outcome. We validated our algorithms with a number of challenging cases, including multiple panes of glass, mirrors, and even a curved glass cabinet. Yu Zhang 0280, Mao Ye 0005, Dinesh Manocha, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2018 | Inexact descent methods for elastic parameter optimizationabstractElastic parameter optimization has revealed its importance in 3D modeling, virtual reality, and additive manufacturing in recent years. Unfortunately, it is known to be computationally expensive, especially if there are many parameters and data samples. To address this challenge, we propose to introduce the inexactness into descent methods, by iteratively solving a forward simulation step and a parameter update step in an inexact manner. The development of such inexact descent methods is centered at two questions: 1) how accurate/inaccurate can the two steps be; and 2) what is the optimal way to implement an inexact descent method. The answers to these questions are in our convergence analysis, which proves the existence of relative error thresholds for the two inexact steps to ensure the convergence. This means we can simply solve each step by a fixed number of iterations, if the iterative solver is at least linearly convergent. While the use of the inexact idea speeds up many descent methods, we specifically favor a GPU-based one powered by state-of-the-art simulation techniques. Based on this method, we study a variety of implementation issues, including backtracking line search, initialization, regularization, and multiple data samples. We demonstrate the use of our inexact method in elasticity measurement and design applications. Our experiment shows the method is fast, reliable, memory-efficient, GPU-friendly, flexible with different elastic models, scalable to a large parameter space, and parallelizable for multiple data samples. Wei Li 0111, Ruigang Yang, Huamin Wang 0001 |
ACM Trans. Graph. | 3 |
| 2017 | Detailed Surface Geometry and Albedo Recovery from RGB-D Video under Natural IlluminationabstractIn this paper we present a novel approach for depth map enhancement from an RGB-D video sequence. The basic idea is to exploit the photometric information in the color sequence. Instead of making any assumption about surface albedo or controlled object motion and lighting, we use the lighting variations introduced by casual object movement. We are effectively calculating photometric stereo from a moving object under natural illuminations. The key technical challenge is to establish correspondences over the entire image set. We therefore develop a lighting insensitive robust pixel matching technique that out-performs optical flow method in presence of lighting variations. In addition we present an expectation-maximization framework to recover the surface normal and albedo simultaneously, without any regularization term. We have validated our method on both synthetic and real datasets to show its superior performance on both surface details recovery and intrinsic decomposition. Xinxin Zuo, Sen Wang 0003, Jiangbin Zheng 0001, Ruigang Yang |
ICCV | 4 |
| 2017 | A generative human-robot motion retargeting approach using a single depth sensorabstractThe goal of human-robot motion retargeting is to let a robot follow the movements performed by a human subject. This is traditionally achieved by applying the estimated poses from a human pose tracking system to a robot via explicit joint mapping strategies. In this paper, we present a novel approach that combine the human pose estimation and the motion retarget procedure in a unified generative framework. A 3D parametric human-robot model is proposed that has the specific joint and stability configurations as a robot while its shape resembles a human subject. Using a single depth camera to monitor human pose, we use its raw depth map as input and drive the human-robot model to fit the input 3D point cloud. The calculated joint angles of the fitted model can be applied onto the robots for retargeting. The robot's joint angles, instead of fitted individually, are fitted globally so that the transformed surface shape is as consistent as possible to the input point cloud. The robot configurations including its skeleton proportion, joint limitation, and DoF are enforced implicitly in the formulation. No explicit and pre-defined joints mapping strategies are needed. This framework is tested with both simulations and real robots that have different skeleton proportion and DoFs compared with human to show its effectiveness for motion retargeting. Sen Wang 0003, Xinxin Zuo, Runxiao Wang, Fuhua (Frank) Cheng, Ruigang Yang |
ICRA | 5 |
| 2016 | Single-Shot Time-of-Flight Phase Unwrapping Using Two Modulation FrequenciesabstractWe present a novel phase unwrapping framework for the Time-of-Flight sensor that can match the performance of systems using two modulation frequencies, within a single shot. Our framework is based on an interleaved pixel arrangement, where a pixel measures phase at a different modulation frequency from its neighboring pixels. We demonstrate that: (1) it is practical to capture ToF images that contain phases from two frequencies in a single shot, with no loss in signal fidelity, (2) phase unwrapping can be effectively performed on such an interleaved phase image, and (3) our method preserves the original spatial resolution. We find that the output of our framework is comparable to results using two shots under separate modulation frequencies, and is significantly better than using a single modulation frequency. Changpeng Ti, Ruigang Yang, James Davis 0001 |
3DV | 2 |
| 2016 | High-speed Depth Stream Generation from a Hybrid CameraabstractHigh-speed video has been commonly adopted in consumer-grade cameras, augmenting these videos with a corresponding depth stream will enable new multimedia applications, such as 3D slow-motion video. In this paper, we present a hybrid camera system that combines a high-speed color camera with a depth sensor, e.g. Kinect depth sensor, to generate a depth stream that can produce both high-speed and high-resolution RGB+depth stream. Simply interpolating the low-speed depth frames is not satisfactory, where interpolation artifacts and lose in surface details are often visible. We have developed a novel framework that utilizes both shading constraints within each frame and optical flow constraints between neighboring frames. More specifically we present (a) an effective method to find the intrinsics images to allow more accurate normal estimation; and (b) an optimization-based framework to estimate the high-resolution/high-speed depth stream, taking into consideration temporal smoothness and shading/depth consistency. We evaluated our holistic framework with both synthetic and real sequences, it showed superior performance than previous state-of-the-art. Xinxin Zuo, Sen Wang 0003, Jiangbin Zheng 0001, Ruigang Yang |
ACM Multimedia | 4 |
| 2016 | Real-Time Simultaneous Pose and Shape Estimation for Articulated Objects Using a Single Depth CameraabstractIn this paper we present a novel real-time algorithm for simultaneous pose and shape estimation for articulated objects, such as human beings and animals. The key of our pose estimation component is to embed the articulated deformation model with exponential-maps-based parametrization into a Gaussian Mixture Model. Benefiting from this probabilistic measurement model, our algorithm requires no explicit point correspondences as opposed to most existing methods. Consequently, our approach is less sensitive to local minimum and handles fast and complex motions well. Moreover, our novel shape adaptation algorithm based on the same probabilistic model automatically captures the shape of the subjects during the dynamic pose estimation process. The personalized shape model in turn improves the tracking accuracy. Furthermore, we propose novel approaches to use either a mesh model or a sphere-set model as the template for both pose and shape estimation under this unified framework. Extensive evaluations on publicly available data sets demonstrate that our method outperforms most state-of-the-art pose estimation algorithms with large margin, especially in the case of challenging motions. Furthermore, our shape estimation method achieves comparable accuracy with state of the arts, yet requires neither statistical shape model nor extra calibration procedure. Our algorithm is not only accurate but also fast, we have implemented the entire processing pipeline on GPU. It can achieve up to 60 frames per second on a middle-range graphics card. Mao Ye 0005, Yang Shen 0011, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2015 | Simultaneous Time-of-Flight sensing and photometric stereo with a single ToF sensorabstractWe present a novel system which incorporates photometric stereo with the Time-of-Flight depth sensor. Adding to the classic ToF, the system utilizes multiple point light sources that enable the capturing of a normal field whilst taking depth images. Two calibration methods are proposed to determine the light sources' positions given the ToF sensor's relatively low resolution. An iterative refinement algorithm is formulated to account for the extra phase delays caused by the positions of the light sources. We find in experiments that the system is comparable to the classic ToF in depth accuracy, and it is able to recover finer details that are lost due to the noise level of the ToF sensor. Changpeng Ti, Ruigang Yang, James Davis 0001 |
CVPR | 2 |
| 2015 | 3D Reconstruction in the presence of glasses by acoustic and stereo fusionabstractWe present a practical and inexpensive method to reconstruct 3D scenes that include piece-wise planar transparent objects. Our work is motivated by the need for automatically generating 3D models of interior scenes, in which glass structures are common. These large structures are often invisible to cameras or even our human visual system. Existing 3D reconstruction methods for transparent objects are usually not applicable in such a room-size reconstruction setting. Our approach augments a regular depth camera (e.g., the Microsoft Kinect camera) with a single ultrasonic sensor, which is able to measure distance to any objects, including transparent surfaces. We present a novel sensor fusion algorithm that first segments the depth map into different categories such as opaque/transparent/infinity (e.g., too far to measure) and then updates the depth map based on the segmentation outcome. Our current hardware setup can generate only one additional point measurement per frame, yet our fusion algorithm is able to generate satisfactory reconstruction results based on our probabilistic model. We highlight the performance in many challenging indoor benchmarks. Mao Ye 0005, Yu Zhang 0280, Ruigang Yang, Dinesh Manocha |
CVPR | 3 |
| 2015 | Interactive Visual Hull Refinement for Specular and Transparent Object Surface ReconstructionabstractIn this paper we present a method of using standard multi-view images for 3D surface reconstruction of non-Lambertian objects. We extend the original visual hull concept to incorporate 3D cues presented by internal occluding contours, i.e., occluding contours that are inside the object's silhouettes. We discovered that these internal contours, which are results of convex parts on an object's surface, can lead to a tighter fit than the original visual hull. We formulated a new visual hull refinement scheme -- Locally Convex Carving that can completely reconstruct concavity caused by two or more intersecting convex surfaces. In addition we develop a novel approach for contour tracking given labeled contours in sparse key frames. It is designed specifically for highly specular or transparent objects, for which assumptions made in traditional contour detection/tracking methods, such as highest gradient and stationary texture edges, are no longer valid. It is formulated as an energy minimization function where several novel terms are developed to increase robustness. Based on the two core algorithms, we have developed an interactive system for 3D modeling. We have validated our system, both quantitatively and qualitatively, with four datasets of different object materials. Results show that we are able to generate visually pleasing models for very challenging cases. Xinxin Zuo, Sen Wang 0003, Jiangbin Zheng 0001, Ruigang Yang |
ICCV | 5 |
| 2015 | Light field projection for lighting reproductionabstractWe propose a novel approach to generate 4D light field in the physical world for lighting reproduction. The light field is generated by projecting lighting images on a lens array. The lens array turns the projected images into a controlled anisotropic point light source array which can simulate the light field of a real scene. In terms of acquisition, we capture an array of light probe images from a real scene, based on which an incident light field is generated. The lens array and the projectors are geometric and photometrically calibrated, and an efficient resampling algorithm is developed to turn the incident light field into the images projected onto the lens array. The reproduced illumination, which allows per-ray lighting control, can produce realistic lighting result on real objects, avoiding the complex process of geometric and material modeling. We demonstrate the effectiveness of our approach with a prototype setup. Zhong Zhou, Xiaofeng Qiu, Ruigang Yang, Qinping Zhao |
VR | 4 |
| 2015 | An improved iterative back projection algorithm based on ringing artifacts suppression
Xin Yang 0002, Dake Zhou, Ruigang Yang |
Neurocomputing | 4 |
| 2014 | Real-Time Simultaneous Pose and Shape Estimation for Articulated Objects Using a Single Depth CameraabstractIn this paper we present a novel real-time algorithm for simultaneous pose and shape estimation for articulated objects, such as human beings and animals. The key of our pose estimation component is to embed the articulated deformation model with exponential-maps-based parametrization into a Gaussian Mixture Model. Benefiting from the probabilistic measurement model, our algorithm requires no explicit point correspondences as opposed to most existing methods. Consequently, our approach is less sensitive to local minimum and well handles fast and complex motions. Extensive evaluations on publicly available datasets demonstrate that our method outperforms most state-of-art pose estimation algorithms with large margin, especially in the case of challenging motions. Moreover, our novel shape adaptation algorithm based on the same probabilistic model automatically captures the shape of the subjects during the dynamic pose estimation process. Experiments show that our shape estimation method achieves comparable accuracy with state of the arts, yet requires neither parametric model nor extra calibration procedure. Mao Ye 0005, Ruigang Yang |
CVPR | 2 |
| 2014 | Quality Dynamic Human Body Modeling Using a Single Low-Cost Depth CameraabstractIn this paper we present a novel autonomous pipeline to build a personalized parametric model (pose-driven avatar) using a single depth sensor. Our method first captures a few high-quality scans of the user rotating herself at multiple poses from different views. We fit each incomplete scan using template fitting techniques with a generic human template, and register all scans to every pose using global consistency constraints. After registration, these watertight models with different poses are used to train a parametric model in a fashion similar to the SCAPE method. Once the parametric model is built, it can be used as an animitable avatar or more interestingly synthesizing dynamic 3D models from single-view depth videos. Experimental results demonstrate the effectiveness of our system to produce dynamic models. Qing Zhang 0017, Mao Ye 0005, Ruigang Yang |
CVPR | 4 |
| 2014 | Data-Driven Flower Petal Modeling with Botany PriorsabstractIn this paper we focus on the 3D modeling of flower, in particular the petals. The complex structure, severe occlusions, and wide variations make the reconstruction of their 3D models a challenging task. Therefore, even though the flower is the most distinctive part of a plant, there has been little modeling study devoted to it. We overcome these challenges by combining data driven modeling techniques with domain knowledge from botany. Taking a 3D point cloud of an input flower scanned from a single view, our method starts with a level-set based segmentation of each individual petal, using both appearance and 3D information. Each segmented petal is then fitted with a scale-invariant morphable petal shape model, which is constructed from individually scanned exemplar petals. Novel constraints based on botany studies, such as the number and spatial layout of petals, are incorporated into the fitting process for realistically reconstructing occluded regions and maintaining correct 3D spatial relations. Finally, the reconstructed petal shape is texture mapped using the registered color images, with occluded regions filled in by content from visible ones. Experiments show that our approach can obtain realistic modeling of flowers even with severe occlusions and large shape/size variations. Mao Ye 0005, Ruigang Yang |
CVPR | 4 |
| 2014 | Towards virtualized welding: Visualization and monitoring of remote weldingabstractWe present a new hybrid reality system that supports the monitoring and visualizing of a welding system. Our system first uses 3D scanning techniques to create a digital model of objects to be welded. Based on the model, a mock-up is constructed from a set of templates or 3D printed. The welding process is captured by cameras and visualized on the 3D mock-up using projectors. The welder can therefore monitor the welding process as if the welding is on the mock-up with proper spatial and 3D cues. An initial user performance evaluation of the system demonstrated several cognitive and performance benefits of the current implementation and suggests avenues for future research. Will Seidelman, Yukang Liu, Travis Kent, C. Melody Carswell, Ruigang Yang |
ICME | 7 |
| 2014 | Video face beautificationabstractThis paper presents a novel system framework of face beautification. Unlike prior works that deal with single images, the proposed beautification framework is designed for an input video and it is able to improve both the appearance and the shape of a face. Our system adopts a state-of-the-art algorithm to synthesize and track 3D face models using blendshapes. The personalized 3D model can be edited to satisfy personal preference. This interactive process is needed only once per subject. Based on the tracking result and the modified face model, we present an algorithm to beautify the face video efficiently and consistently. Furthermore we develop a variant of content preserving warping to reduce warping distortions along the face boundary. Finally we adopt real time bilateral filtering to remove wrinkles, freckles, and unwanted blemishes. This framework is evaluated on a set of videos. The experiments demonstrate that our framework can generate consistent and pleasant results over video frames while the original expressions and features are persevered naturally. Xinyu Huang 0001, Jizhou Gao, Alade O. Tokuta, Cha Zhang, Ruigang Yang |
ICME | 6 |
| 2014 | A Performance Comparison between Circular and Spline-Based Methods for Iris SegmentationabstractIris segmentation is an important module of iris recognition that can substantially affect recognition performance. Since iris and pupil boundaries usually are not exactly circular, spline-based methods have been used to model irregular iris and pupil boundaries recently. However, in most existing methods, many other factors or modules in the iris recognition pipeline are evaluated together and their mixed effects are assumed to be negligible. More importantly, the splines that model irregularity of the boundaries could not be enough to model the internal nonlinear deformations of an iris pattern (e.g., caused by iris dilation). As a result, it remains unclear whether spline-based methods can provide significant improvements. In this paper, we conduct a complete performance comparison between circular and spline-based methods. There are mainly two contributions. Firstly, for the purpose of comparison, we propose a spline estimator that is robust to outliers caused by eyelashes, eyelids, highlights, and shadows. Secondly, we analyze the relation between iris matching distances and segmentation results by using circular and spline-based methods. Based on our experiments, we found that, even with the proposed robust spline estimator, the improvement of recognition performance is still limited (around 6%). Therefore, in case that less robust spline estimators are used due to the real-time requirement in practical systems, the actual recognition improvement by using splines could be far below the expectation. Changpeng Ti, Xinyu Huang 0001, Alade O. Tokuta, Ruigang Yang |
ICPR | 5 |
| 2014 | Virtualized welding: a new paradigm for tele-operated weldingabstractWe present a new mixed reality system that supports tele-operation of a welding robot. We create a 3D mockup of the welding pieces and use projector-based displays to visualize the welding process directly on the 3D display. Multi-cameras are used to capture both the welding environment and the operator's motion. The welder can therefore monitor and control the welding process as if the welding is on the mock-up, which provides proper spatial and 3D cues. We evaluated our system with a number of control tasks and the results shows the effectiveness of our system as compared to traditional alternatives. Yukang Liu, Ruigang Yang |
VRST | 4 |
| 2014 | Real-time Human Pose and Shape Estimation for Virtual Try-On Using a Single Commodity Depth CameraabstractWe present a system that allows the user to virtually try on new clothes. It uses a single commodity depth camera to capture the user in 3D. Both the pose and the shape of the user are estimated with a novel real-time template-based approach that performs tracking and shape adaptation jointly. The result is then used to drive realistic cloth simulation, in which the synthesized clothes are overlayed on the input image. The main challenge is to handle missing data and pose ambiguities due to the monocular setup, which captures less than 50 percent of the full body. Our solution is to incorporate automatic shape adaptation and novel constraints in pose tracking. The effectiveness of our system is demonstrated with a number of examples. Mao Ye 0005, Huamin Wang 0001, Nianchen Deng, Xubo Yang, Ruigang Yang |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2014 | Personal Photograph Enhancement Using Internet Photo CollectionsabstractGiven the growth of Internet photo collections, we now have a visual index of all major cities and tourist sites in the world. However, it is still a difficult task to capture that perfect shot with your own camera when visiting these places, especially when your camera itself has limitations, such as a limited field of view. In this paper, we propose a framework to overcome the imperfections of personal photographs of tourist sites using the rich information provided by large-scale Internet photo collections. Our method deploys state-of-the-art techniques for constructing initial 3D models from photo collections. The same techniques are then used to register personal photographs to these models, allowing us to augment personal 2D images with 3D information. This strong available scene prior allows us to address a number of traditionally challenging image enhancement techniques and achieve high-quality results using simple and robust algorithms. Specifically, we demonstrate automatic foreground segmentation, mono-to-stereo conversion, field-of-view expansion, photometric enhancement, and additionally automatic annotation with geolocation and tags. Our method clearly demonstrates some possible benefits of employing the rich information contained in online photo databases to efficiently enhance and augment one's own personal photographs. Jizhou Gao, Oliver Wang, Pierre Fite Georgel, Ruigang Yang, James Davis 0001, Jan-Michael Frahm, Marc Pollefeys |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2013 | Online Building Segmentation from Ground-Based LiDAR Data in Urban ScenesabstractThe availability of active 3D sensing devices such as LiDAR has significantly increased the collection of 3D urban scenes with rich details. The sheer amount of data brings a lot of opportunities but also poses tremendous challenges for both academic research and industrial applications on point cloud classification and building reconstruction. In this paper, we present an online algorithm to automatically detect and segment buildings from large scale unorganized 3D point clouds of urban scenes acquired by ground-Based LiDAR devices. The core idea is that buildings can be observed in a street view separated by empty spaces such as alleys. By progressively projecting 3D points onto views along the scanning path, buildings can be detected as large regions with dense points. Experiments on several large scale datasets show that our approach can efficiently produce satisfactory results. Jizhou Gao, Ruigang Yang |
3DV | 2 |
| 2013 | High-Quality Stereo Video Matching via User Interaction and Space-Time PropagationabstractEven current state-of-the-art automatic stereo matching methods often struggle on natural images and videos, in great part due to fundamental matching ambiguities in low texture regions and a lack of higher level object knowledge. Stereo image matching can benefit greatly from user input to guide the matching process and help disambiguate matches. Applying interactive correction tools from scratch on each frame of a video would not only be throwing away valuable information provided by the user on other frames, but would also likely be too time consuming to be practical for video even if excellent disparity results could be obtained within a few minutes on each frame. In this work, we propose a stereo video matching system that allows user interaction to obtain high quality, dense disparity maps on key frames and then intelligently propagates the user input and key frame disparities to automatically produce high quality disparity maps on intermediate frames. The disparity maps on key frames are obtained using several novel, easy-to-use, and effective interactive tools. Our novel propagation algorithm estimates 3D transformations that map user corrected areas in key frames to intermediate frames. Experiments demonstrate the effectiveness and efficiency of our hybrid interactive/automatic approach. Brian L. Price, Scott Cohen, Ruigang Yang |
3DV | 4 |
| 2013 | Video Enhancement of People Wearing Polarized Glasses: Darkening Reversal and Reflection ReductionabstractWith the wide-spread of consumer 3D-TV technology, stereoscopic videoconferencing systems are emerging. However, the special glasses participants wear to see 3D can create distracting images. This paper presents a computational framework to reduce undesirable artifacts in the eye regions caused by these 3D glasses. More specifically, we add polarized filters to the stereo camera so that partial images of reflection can be captured. A novel Bayesian model is then developed to describe the imaging process of the eye regions including darkening and reflection, and infer the eye regions based on Classification Expectation-Maximization (EM). The recovered eye regions under the glasses are brighter and with little reflections, leading to a more nature videoconferencing experience. Qualitative evaluations and user studies are conducted to demonstrate the substantial improvement our approach can achieve. Mao Ye 0005, Cha Zhang, Ruigang Yang |
CVPR | 3 |
| 2013 | Predictive control for robot arm teleoperationabstractThis paper presents a robotic welding teleoperation system that can learn the skilled welder's behavior and transfer human intelligence into welding robots. In this system a 6-DOF robot arm is equipped with a compact 3D weld pool sensing system. The motion of the human welder movement is tracked accurately by the Leap sensor. A predictive control algorithm is proposed to control the speed of the robot arm movement. Tracking experiments are conducted to track pre-set movement with varying speed and human hand movement. It is found that the robot with the proposed speed control algorithm is able to track human hand movement with acceptable accuracy. A foundation is thus established to realize welding teleoperation and transfer human intelligence to the welding robot. Yukang Liu, Ruigang Yang |
IECON | 4 |
| 2013 | An experimental study of pupil constriction for liveness detectionabstractAs iris recognition systems have been deployed in many security areas, liveness detection that can distinguish between real iris patterns and fake ones becomes an important module. Most existing algorithms focus on the appearance difference between real and fake iris (for example, printed patterns, cosmetic contact lenses etc.) which is a very difficult problem. Instead of studying image properties of fake irises, we show that pupil constriction, the fundamental characteristic of real and live irises, can be very robust for liveness detection. In this experimental study, we first build an iris acquisition system that can acquire two eye images under two different illumination conditions in a less intrusive environment. Second, in order to model the process of pupil constriction, we propose a feature descriptor that consists of similarity measurement between iris patches and ratio of iris and pupil diameters. Third, the performance of liveness prediction is evaluated based on the training of a Support Vector Machine (SVM) classifier. The high success prediction rate shows that the classifier is effective without knowing any prior knowledge of fake irises. Xinyu Huang 0001, Changpeng Ti, Qi-zhen Hou, Alade O. Tokuta, Ruigang Yang |
WACV | 5 |
| 2013 | Guest Editorial: 3D Imaging, Processing and Modelling
Guy Godin, Michael Goesele, Yasuyuki Matsushita, Ryusuke Sagawa, Ruigang Yang |
Int. J. Comput. Vis. | 5 |
| 2013 | Measurement of mirror surfaces using specular reflection and analytical computation
Zhen Zhou Wang, Xinyu Huang 0001, Ruigang Yang |
Mach. Vis. Appl. | 3 |
| 2013 | Fusion of Median and Bilateral Filtering for Range Image UpsamplingabstractWe present a new upsampling method to enhance the spatial resolution of depth images. Given a low-resolution depth image from an active depth sensor and a potentially high-resolution color image from a passive RGB camera, we formulate it as an adaptive cost aggregation problem and solve it using the bilateral filter. The formulation synergistically combines the median and bilateral filters thus it better preserves the depth edges and is more robust to noise. Numerical and visual evaluations on a total of 37 Middlebury data sets demonstrate the effectiveness of our method. A real-time high-resolution depth capturing system is also developed using commercial active depth sensor based on the proposed upsampling method. Qingxiong Yang, Narendra Ahuja, Ruigang Yang, Kar-Han Tan, James Davis 0001, W. Bruce Culbertson, John G. Apostolopoulos, Gang Wang 0012 |
IEEE Trans. Image Process. | 3 |
| 2013 | Semantic decomposition and reconstruction of residential scenes from LiDAR dataabstractWe present a complete system to semantically decompose and reconstruct 3D models from point clouds. Different than previous urban modeling approaches, our system is designed for residential scenes, which consist of mainly low-rise buildings that do not exhibit the regularity and repetitiveness as high-rise buildings in downtown areas. Our system first automatically labels the input into distinctive categories using supervised learning techniques. Based on the semantic labels, objects in different categories are reconstructed with domain-specific knowledge. In particular, we present a novel building modeling scheme that aims to decompose and fit the building point cloud into basic blocks that are block-wise symmetric and convex . This building representation and its reconstruction algorithm are flexible, efficient, and robust to missing data. We demonstrate the effectiveness of our system on various datasets and compare our building modeling scheme with other state-of-the-art reconstruction algorithms to show its advantage in terms of both quality and speed. Jizhou Gao, Yu Zhou 0007, Guiliang Lu, Mao Ye 0005, Ruigang Yang |
ACM Trans. Graph. | 8 |
| 2012 | Edge-preserving photometric stereo via depth fusionabstractWe present a sensor fusion scheme that combines active stereo with photometric stereo. Aiming at capturing full-frame depth for dynamic scenes at a minimum of three lighting conditions, we formulate an iterative optimization scheme that (1) adaptively adjusts the contribution from photometric stereo so that discontinuity can be preserved; (2) detects shadow areas by checking the visibility of the estimated point with respect to the light source, instead of using image-based heuristics; and (3) behaves well for ill-conditioned pixels that are under shadow, which are inevitable in almost any scene. Furthermore, we decompose our non-linear cost function into subproblems that can be optimized efficiently using linear techniques. Experiments show significantly improved results over the previous state-of-the-art in sensor fusion. Qing Zhang 0017, Mao Ye 0005, Ruigang Yang, Yasuyuki Matsushita, Bennett Wilburn |
CVPR | 3 |
| 2012 | See-through Image Enhancement through Sensor FusionabstractMany hardware designs have been developed to allow a camera to be placed optically directly behind the screen. The purpose of such setups is to enable two-way video teleconferencing that maintains eye-contact. However, the image from the see-through camera usually exhibits a number of imaging artifacts such as low signal to noise ratio, incorrect color balance, and lost of details. We develop a novel image enhancement framework that utilizes an auxiliary color+depth camera that is mounted on the side of the screen. By fusing the information from both cameras, we are able to significantly improve the quality of the see-through image. Experimental results have demonstrated that our fusion method compares favorably against traditional image enhancement/warping methods that uses only a single image. Mao Ye 0005, Ruigang Yang, Cha Zhang |
ICME | 3 |
| 2012 | Simulation Guided Hair Dynamics Modeling from VideoabstractAbstract In this paper we present a hybrid approach to reconstruct hair dynamics from multi‐view video sequences, captured under uncontrolled lighting conditions. The key of this method is a refinement approach that combines image‐based reconstruction techniques with physically based hair simulation. Given an initially reconstructed sequence of hair fiber models, we develop a hair dynamics refinement system using particle‐based simulation and incompressible fluid simulation. The system allows us to improve reconstructed hair fiber motions and complete missing fibers caused by occlusion or tracking failure. The refined space‐time hair dynamics are consistent with video inputs and can be also used to generate novel hair animations of different hair styles. We validate this method through various real hair examples. Qing Zhang 0017, Jing Tong, Huamin Wang 0001, Ruigang Yang |
Comput. Graph. Forum | 5 |
| 2012 | Automatic Real-Time Video Matting Using Time-of-Flight Camera and Multichannel Poisson Equations
Liang Wang 0002, Minglun Gong, Ruigang Yang, Cha Zhang, Yee-Hong Yang |
Int. J. Comput. Vis. | 4 |
| 2012 | Video Stereolization: Combining Motion Analysis with User InteractionabstractWe present a semiautomatic system that converts conventional videos into stereoscopic videos by combining motion analysis with user interaction, aiming to transfer as much as possible labeling work from the user to the computer. In addition to the widely used structure from motion (SFM) techniques, we develop two new methods that analyze the optical flow to provide additional qualitative depth constraints. They remove the camera movement restriction imposed by SFM so that general motions can be used in scene depth estimation-the central problem in mono-to-stereo conversion. With these algorithms, the user's labeling task is significantly simplified. We further developed a quadratic programming approach to incorporate both quantitative depth and qualitative depth (such as these from user scribbling) to recover dense depth maps for all frames, from which stereoscopic view can be synthesized. In addition to visual results, we present user study results showing that our approach is more intuitive and less labor intensive, while producing 3D effect comparable to that from current state-of-the-art interactive algorithms. Miao Liao, Jizhou Gao, Ruigang Yang, Minglun Gong |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2011 | Interreflection removal for photometric stereo by using spectrum-dependent albedoabstractWe present a novel method that can separate m-bounced light and remove the interreflections in a photometric stereo setup. Under the assumption of a uniformly colored lambertian surface, the intensity of a point in the scene is the sum of 1-bounced light through m-bounced light rays. Ruled by the law of diffuse reflection, whenever a light ray is bounced by the surface, its intensity will be attenuated by the factor of albedo ρ. This implies that the measured intensity value can be written as a polynomial function of ρ, and the intensity contribution of the m-bounced light rays are expressed by the term of ρm. Therefore, when we change the surface albedo, the intensity of the m-bounced light is changed to the order of m. This non-linearity gives us the possibility to separate the m-bounced light. In practice, we illuminate the scene with different light colors to effectively simulate different surface albedos since albedo is spectrum dependent. Once the m-bounced light rays are separated, we can perform the photometric stereo algorithm on the 1-bounced light (direct lighting) images to produce the 3D shape without the impact of interreflections. Experiments have shown that we get significantly improved scene reconstruction with a minimum of two color images. Miao Liao, Xinyu Huang 0001, Ruigang Yang |
CVPR | 3 |
| 2011 | Global stereo matching leveraged by sparse ground control pointsabstractWe present a novel global stereo model that makes use of constraints from points with known depths, i.e., the Ground Control Points (GCPs) as referred to in stereo literature. Our formulation explicitly models the influences of GCPs in a Markov Random Field. A novel GCPs-based regularization term is naturally integrated into our global optimization framework in a principled way using the Bayes rule. The optimal solution of the inference problem can be approximated via existing energy minimization techniques such as graph cuts used in this paper. Our generic probabilistic framework allows GCPs to be obtained from various modalities and provides a natural way to integrate the information from multiple sensors. Quantitative evaluations demonstrate the effectiveness of the proposed formulation for regularizing the ill-posed stereo matching problem and improving reconstruction accuracy. Liang Wang 0002, Ruigang Yang |
CVPR | 2 |
| 2011 | Accurate 3D pose estimation from a single depth imageabstractThis paper presents a novel system to estimate body pose configuration from a single depth map. It combines both pose detection and pose refinement. The input depth map is matched with a set of pre-captured motion exemplars to generate a body configuration estimation, as well as semantic labeling of the input point cloud. The initial estimation is then refined by directly fitting the body configuration with the observation (e.g., the input depth). In addition to the new system architecture, our other contributions include modifying a point cloud smoothing technique to deal with very noisy input depth maps, a point cloud alignment and pose search algorithm that is view-independent and efficient. Experiments on a public dataset show that our approach achieves significantly higher accuracy than previous state-of-art methods. Mao Ye 0005, Xianwang Wang, Ruigang Yang, Liu Ren 0001, Marc Pollefeys |
ICCV | 3 |
| 2011 | A novel see-through screen based on weave fabricsabstractSee-through screens (STS) have found important applications in remote collaboration systems to enhance non-verbal communication and gaze awareness. Existing STS designs often sacrifice the display quality significantly, rendering low-contrast images that discount the overall user experience. In this paper, we present a novel see-through screen solution based on weave fabrics. Such fabrics are known to be acoustically transparent and used to build professional projection screens for Hollywood studios. We place a cam-era immediately behind the screen and synchronize it with a 120Hz projector to perform time-multiplexing display and video capture. By focusing the camera at the user 4–5 feet away from the screen, the image of the weave fabric will be severely blurred. We present the imaging principle of the setup, and derive image processing techniques to enhance the quality of the captured video. The overall system is low cost, has much better display quality than existing systems, and can be used to build wall-size see-through screens for various applications. Cha Zhang, Ruigang Yang, Tim Large, Zhengyou Zhang |
ICME | 2 |
| 2011 | Reliability Fusion of Time-of-Flight Depth and Stereo Geometry for High Quality Depth MapsabstractTime-of-flight range sensors have error characteristics, which are complementary to passive stereo. They provide real-time depth estimates in conditions where passive stereo does not work well, such as on white walls. In contrast, these sensors are noisy and often perform poorly on the textured scenes where stereo excels. We explore their complementary characteristics and introduce a method for combining the results from both methods that achieve better accuracy than either alone. In our fusion framework, the depth probability distribution functions from each of these sensor modalities are formulated and optimized. Robust and adaptive fusion is built on a pixel-wise reliability weighting function calculated for each method. In addition, since time-of-flight devices have primarily been used as individual sensors, they are typically poorly calibrated. We introduce a method that substantially improves upon the manufacturer's calibration. We demonstrate that our proposed techniques lead to improved accuracy and robustness on an extensive set of experimental results. Jiejie Zhu, Liang Wang 0002, Ruigang Yang, James Davis 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | A Uniform Framework for Estimating Illumination Chromaticity, Correspondence, and Specular ReflectionabstractBased upon a new correspondence matching invariant called illumination chromaticity constancy, we present a new solution for illumination chromaticity estimation, correspondence searching, and specularity removal. Using as few as two images, the core of our method is the computation of a vote distribution for a number of illumination chromaticity hypotheses via correspondence matching. The hypothesis with the highest vote is accepted as correct. The estimated illumination chromaticity is then used together with the new matching invariant to match highlights, which inherently provides solutions for correspondence searching and specularity removal. Our method differs from the previous approaches: those treat these vision problems separately and generally require that specular highlights be detected in a preprocessing step. Also, our method uses more images than previous illumination chromaticity estimation methods, which increases its robustness because more inputs/constraints are used. Experimental results on both synthetic and real images demonstrate the effectiveness of the proposed method. Qingxiong Yang, Narendra Ahuja, Ruigang Yang |
IEEE Trans. Image Process. | 4 |
| 2010 | Learning 3D shape from a single facial image via non-linear manifold embedding and alignmentabstractThe 3D reconstruction of a face from a single frontal image is an ill-posed problem. This is further accentuated when the face image is captured under different poses and/or complex illumination conditions. In this paper, we aim to solve the shape recovery problem from a single facial image under these challenging conditions. The local image models for each patch of facial images and the local surface models for each patch of 3D shape are learned using a non-linear dimensionality reduction technique, and the correspondences between these local models are then learned by a manifold alignment method. By combining the local shapes, the global shape of a face can be reconstructed directly using a single least-square system of equations. We perform experiments on synthetic and real data, and validate the algorithm against the ground truth. Experimental results show that our method can yield accurate shape recovery from out-of-training samples with a variety of pose and illumination variations. Xianwang Wang, Ruigang Yang |
CVPR | 2 |
| 2010 | Semantic Segmentation of Urban Scenes Using Dense Depth Maps
Liang Wang 0002, Ruigang Yang |
ECCV (4) | 3 |
| 2010 | Real-time video matting using multichannel poisson equations
Minglun Gong, Liang Wang 0002, Ruigang Yang, Yee-Hong Yang |
Graphics Interface | 3 |
| 2010 | Spatial-Temporal Fusion for High Accuracy Depth Maps Using Dynamic MRFsabstractTime-of-flight range sensors and passive stereo have complimentary characteristics in nature. To fuse them to get high accuracy depth maps varying over time, we extend traditional spatial MRFs to dynamic MRFs with temporal coherence. This new model allows both the spatial and the temporal relationship to be propagated in local neighbors. By efficiently finding a maximum of the posterior probability using Loopy Belief Propagation, we show that our approach leads to improved accuracy and robustness of depth estimates for dynamic scenes. Jiejie Zhu, Liang Wang 0002, Jizhou Gao, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2010 | Endoscopic Video Texture Mapping on Pre-Built 3-D Anatomical Objects Without Camera TrackingabstractTraditional minimally invasive surgeries use a view port provided by an endoscope or laparoscope. We argue that a useful addition to typical endoscopic imagery would be a global 3-D view providing a wider field of view with explicit depth information for both the exterior and interior of target anatomy. One technical challenge of implementing such a view is finding efficient and accurate means of registering texture images from the laparoscope on prebuilt 3-D surface models of target anatomy derived from magnetic resonance (MR) or computed tomography (CT) images. This paper presents a novel method for addressing this challenge that differs from previous approaches, which depend on tracking the position of the laparoscope. We take advantage of the fact that neighboring frames within a video sequence usually contain enough coherence to allow a 2-D-2-D registration, which is a much more tractable problem. The texturing process can be bootstrapped by an initial 2-D-3-D user-assisted registration of the first video frame followed by mostly-automatic texturing of subsequent frames. We perform experiments on phantom and real data, validate the algorithm against the ground truth, and compare it with the traditional tracking method by simulations. Experiments show that our method improves registration performance compared to the traditional tracking approach. Xianwang Wang, Qing Zhang 0017, Qiong Han, Ruigang Yang, C. Melody Carswell, W. Brent Seales, Erica Sutton |
IEEE Trans. Medical Imaging | 4 |
| 2009 | Manifold Estimation in View-Based Feature Space for Face Synthesis across Poses
Xinyu Huang 0001, Jizhou Gao, Sen-Ching S. Cheung, Ruigang Yang |
ACCV (1) | 4 |
| 2009 | Image deblurring for less intrusive iris captureabstractFor most iris capturing scenarios, captured iris images could easily blur when the user is out of the depth of field (DOF) of the camera, or when he or she is moving. The common solution is to let the user try the capturing process again as the quality of these blurred iris images is not good enough for recognition. In this paper, we propose a novel iris deblurring algorithm that can be used to improve the robustness and nonintrusiveness for iris capture. Unlike other iris deblurring algorithms, the key feature of our algorithm is that we use the domain knowledge inherent in iris images and iris capture settings to improve the performance, which could be in the form of iris image statistics, characteristics of pupils or highlights, or even depth information from the iris capturing system itself. Our experiments on both synthetic and real data demonstrate that our deblurring algorithm can significantly restore blurred iris patterns and therefore improve the robustness of iris capture. Xinyu Huang 0001, Liu Ren 0001, Ruigang Yang |
CVPR | 3 |
| 2009 | Joint depth and alpha matte optimization via fusion of stereo and time-of-flight sensorabstractWe present a new approach to iteratively estimate both high-quality depth map and alpha matte from a single image or a video sequence. Scene depth, which is invariant to illumination changes, color similarity and motion ambiguity, provides a natural and robust cue for foreground/ background segmentation - a prerequisite for matting. The image mattes, on the other hand, encode rich information near boundaries where either passive or active sensing method performs poorly. We develop a method to combine the complementary nature of scene depth and alpha matte to mutually enhance their qualities. We formulate depth inference as a global optimization problem where information from passive stereo, active range sensor and matte is merged. The depth map is used in turn to enhance the matting. In addition, we extend this approach to video matting by incorporating temporal coherence, which reduces flickering in the composite video. We show that these techniques lead to improved accuracy and robustness for both static and dynamic scenes. Jiejie Zhu, Miao Liao, Ruigang Yang |
CVPR | 3 |
| 2009 | Unsupervised learning of high-order structural semantics from imagesabstractStructural semantics are fundamental to understanding both natural and man-made objects from languages to buildings. They are manifested as repeated structures or patterns and are often captured in images. Finding repeated patterns in images, therefore, has important applications in scene understanding, 3D reconstruction, and image retrieval as well as image compression. Previous approaches in visual-pattern mining limited themselves by looking for frequently co-occurring features within a small neighborhood in an image. However, semantics of a visual pattern are typically defined by specific spatial relationships between features regardless of the spatial proximity. In this paper, semantics are represented as visual elements and geometric relationships between them. A novel unsupervised learning algorithm finds pair-wise associations of visual elements that have consistent geometric relationships sufficiently often. The algorithms are efficient - maximal matchings are determined without combinatorial search. High-order structural semantics are extracted by mining patterns that are composed of pairwise spatially consistent associations of visual elements. We demonstrate the effectiveness of our approach for discovering repeated visual patterns on a variety of image collections. Jizhou Gao, Jinze Liu, Ruigang Yang |
ICCV | 4 |
| 2009 | Modeling deformable objects from a single depth cameraabstractWe propose a novel approach to reconstruct complete 3D deformable models over time by a single depth camera, provided that most parts of the models are observed by the camera at least once. The core of this algorithm is based on the assumption that the deformation is continuous and predictable in a short temporal interval. While the camera can only capture part of a whole surface at any time instant, partial surfaces reconstructed from different times are assembled together to form a complete 3D surface for each time instant, even when the shape is under severe deformation. A mesh warping algorithm based on linear mesh deformation is used to align different partial surfaces. A volumetric method is then used to combine partial surfaces, fix missing holes, and smooth alignment errors. Our experiment shows that this approach is able to reconstruct visually plausible 3D surface deformation results with a single camera. Miao Liao, Qing Zhang 0017, Huamin Wang 0001, Ruigang Yang, Minglun Gong |
ICCV | 4 |
| 2009 | Multimedia processing on commodity graphics hardwareabstractDriven by the need for interactive entertainment, modern PCs are equipped with specialized graphics processors (GPUs) for creation and display of images. These GPUs have become increasingly programmable, to the point that they now are capable of efficiently executing a significant number of computational kernels from non-graphical applications. In this introductory paper we first present a high-level overview of modern graphics hardwares architecture, then introduce programming tools available for application development on GPUs. Finally we present several multimedia applications that have been efficiently accelerated by GPUs. Ruigang Yang |
ICME | 1 |
| 2009 | Stereo Matching with Color-Weighted Correlation, Hierarchical Belief Propagation, and Occlusion HandlingabstractIn this paper, we formulate a stereo matching algorithm with careful handling of disparity, discontinuity and occlusion. The algorithm works with a global matching stereo model based on an energy-minimization framework. The global energy contains two terms, the data term and the smoothness term. The data term is first approximated by a color-weighted correlation, then refined in occluded and low-texture areas in a repeated application of a hierarchical loopy belief propagation algorithm. The experimental results are evaluated on the Middlebury data sets, showing that our algorithm is the top performer among all the algorithms listed there. Qingxiong Yang, Liang Wang 0002, Ruigang Yang, Henrik Stewénius, David Nistér |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2009 | Physically guided liquid surface modeling from videosabstractWe present an image-based reconstruction framework to model real water scenes captured by stereoscopic video. In contrast to many image-based modeling techniques that rely on user interaction to obtain high-quality 3D models, we instead apply automatically calculated physically-based constraints to refine the initial model. The combination of image-based reconstruction with physically-based simulation allows us to model complex and dynamic objects such as fluid. Using a depth map sequence as initial conditions, we use a physically based approach that automatically fills in missing regions, removes outliers, and refines the geometric shape so that the final 3D model is consistent to both the input video data and the laws of physics. Physically-guided modeling also makes interpolation or extrapolation in the space-time domain possible, and even allows the fusion of depth maps that were taken at different times or viewpoints. We demonstrated the effectiveness of our framework with a number of real scenes, all captured using only a single pair of cameras. Huamin Wang 0001, Miao Liao, Qing Zhang 0017, Ruigang Yang, Greg Turk |
ACM Trans. Graph. | 4 |
| 2008 | Stereoscopic inpainting: Joint color and depth completion from stereo imagesabstractWe present a novel algorithm for simultaneous color and depth inpainting. The algorithm takes stereo images and estimated disparity maps as input and fills in missing color and depth information introduced by occlusions or object removal. We first complete the disparities for the occlusion regions using a segmentation-based approach. The completed disparities can be used to facilitate the user in labeling objects to be removed. Since part of the removed regions in one image is visible in the other, we mutually complete the two images through 3D warping. Finally, we complete the remaining unknown regions using a depth-assisted texture synthesis technique, which simultaneously fills in both color and depth. We demonstrate the effectiveness of the proposed algorithm on several challenging data sets. Liang Wang 0002, Hailin Jin, Ruigang Yang, Minglun Gong |
CVPR | 3 |
| 2008 | Fusion of time-of-flight depth and stereo for high accuracy depth mapsabstractTime-of-flight range sensors have error characteristics which are complementary to passive stereo. They provide real time depth estimates in conditions where passive stereo does not work well, such as on white walls. In contrast, these sensors are noisy and often perform poorly on the textured scenes for which stereo excels. We introduce a method for combining the results from both methods that performs better than either alone. A depth probability distribution function from each method is calculated and then merged. In addition, stereo methods have long used global methods such as belief propagation and graph cuts to improve results, and we apply these methods to this sensor. Since time-of-flight devices have primarily been used as individual sensors, they are typically poorly calibrated. We introduce a method that substantially improves upon the manufacturerpsilas calibration. We show that these techniques lead to improved accuracy and robustness. Jiejie Zhu, Liang Wang 0002, Ruigang Yang, James Davis 0001 |
CVPR | 3 |
| 2008 | Illumination and Person-Insensitive Head Pose Estimation Using Distance Metric Learning
Xianwang Wang, Xinyu Huang 0001, Jizhou Gao, Ruigang Yang |
ECCV (2) | 4 |
| 2008 | Search Space Reduction for MRF Stereo
Liang Wang 0002, Hailin Jin, Ruigang Yang |
ECCV (1) | 3 |
| 2008 | Real-time Light Fall-off StereoabstractWe present a real-time depth recovery system using Light Fall-off Stereo (LFS). Our system contains two co-axial point light sources (LEDs) synchronized with a video camera. The video camera captures the scene under these two LEDs in complementary states(e.g., one on, one off). Based on the inverse square law for light intensity, the depth can be directly solved using the pixel ratio from two consecutive frames. We demonstrate the effectiveness of our approach with a number of real world scenes. Quantitative evaluation shows that our system compares favorably to other commercial real-time 3D range sensors, particularly in textured areas. We believe our system offers a low-cost high-resolution alternative for depth sensing under controlled lighting. Miao Liao, Liang Wang 0002, Ruigang Yang, Minglun Gong |
ICIP | 3 |
| 2008 | Detailed Real-Time Urban 3D Reconstruction from Video
Marc Pollefeys, David Nistér, Jan-Michael Frahm, Amir Akbarzadeh, Philippos Mordohai, Brian Clipp, Chris Engels, David Gallup, Seon Joo Kim, Paul Merrell, C. Salmi, Sudipta N. Sinha, B. Talton, Liang Wang 0002, Qingxiong Yang, Henrik Stewénius, Ruigang Yang, Greg Welch, Herman Towles |
Int. J. Comput. Vis. | 17 |
| 2008 | Robust and Accurate Visual Echo Cancelation in a Full-duplex Projector-Camera SystemabstractIn this paper we study the problem of "visual echo" in a full-duplex projector-camera system for telecollaboration applications. Visual echo is defined as the appearance of projected contents observed by the camera. It can potentially saturate the projected contents, similar to audio echo in telephone conversation. Our approach to visual echo cancellation includes an offline calibration procedure that records the geometric and photometric transfer between the projector and the camera in a look-up table. During run-time, projected contents in the captured video are identified using the calibration information and suppressed, therefore achieving the goal of cancelling visual echo. Our approach can accurately handle full-color images under arbitrary reflectance of display surfaces and photometric response of the projector or camera. It is robust to geometric registration errors and quantization effects and is therefore particularly effective for high-frequency contents such as texts and hand drawings. We demonstrate the effectiveness of our approach with a variety of real images in a full-duplex projector-camera system. Miao Liao, Ruigang Yang, Zhengyou Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Toward the Light Field Display: Autostereoscopic Rendering via a Cluster of ProjectorsabstractUltimately, a display device should be capable of reproducing the visual effects observed in reality. In this paper we introduce an autostereoscopic display that uses a scalable array of digital light projectors and a projection screen augmented with microlenses to simulate a light field for a given three-dimensional scene. Physical objects emit or reflect light in all directions to create a light field that can be approximated by the light field display. The display can simultaneously provide many viewers from different viewpoints a stereoscopic effect without head tracking or special viewing glasses. This work focuses on two important technical problems related to the light field display; calibration and rendering. We present a solution to automatically calibrate the light field display using a camera and introduce two efficient algorithms to render the special multi-view images by exploiting their spatial coherence. The effectiveness of our approach is demonstrated with a four-projector prototype that can display dynamic imagery with full parallax. Ruigang Yang, Xinyu Huang 0001, Sifang Li, Christopher O. Jaynes |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2007 | Calibrating Pan-Tilt Cameras with Telephoto Lenses
Xinyu Huang 0001, Jizhou Gao, Ruigang Yang |
ACCV (1) | 3 |
| 2007 | Light Fall-off StereoabstractWe present light fall-off stereo-LFS-a new method for computing depth from scenes beyond lambertian reflectance and texture. LFS takes a number of images from a stationary camera as the illumination source moves away from the scene. Based on the inverse square law for light intensity, the ratio images are directly related to scene depth from the perspective of the light source. Using this as the invariant, we developed both local and global methods for depth recovery. Compared to previous reconstruction methods for non-lamebrain scenes, LFS needs as few as two images, does not require calibrated camera or light sources, or reference objects in the scene. We demonstrated the effectiveness of LFS with a variety of real-world scenes. Miao Liao, Liang Wang 0002, Ruigang Yang, Minglun Gong |
CVPR | 3 |
| 2007 | Flexible Pixel Compositor for Plug-and-Play Multi-Projector DisplaysabstractIn summary, we are developing the next generation compositor to satisfy the demanding needs from emerging applications. It can be used beyond multi-projector displays. The first is auto-stereoscopic (multi-view) displays, in particular lenticular-based displays. These 3D displays in fact display many views simultaneously and therefore require orders of magnitude more pixels to provide an observer adequate resolution. This can be achieved only by a rendering cluster. Furthermore, images from the rendering nodes typically need to be sliced and interleaved to form the proper composite image for display. We also envision that our flexible hardware can be used for distributed general-purpose computing on graphics processor units (GPGPU). It provides the random write capability missing in most current graphics hardware. By providing a scalable and flexible link among a cluster of GPUs, they can efficiently work in concert to solve problems, both graphical and non-graphical, on a much larger scale. Ruigang Yang, Daniel R. Rudolf, Vijai Raghunathan |
CVPR | 1 |
| 2007 | Spatial-Depth Super Resolution for Range ImagesabstractWe present a new post-processing step to enhance the resolution of range images. Using one or two registered and potentially high-resolution color images as reference, we iteratively refine the input low-resolution range image, in terms of both its spatial resolution and depth precision. Evaluation using the Middlebury benchmark shows across-the-board improvement for sub-pixel accuracy. We also demonstrated its effectiveness for spatial resolution enhancement up to 100 times with a single reference image. Qingxiong Yang, Ruigang Yang, James Davis 0001, David Nistér |
CVPR | 2 |
| 2007 | Real-Time Visibility-Based Fusion of Depth MapsabstractWe present a viewpoint-based approach for the quick fusion of multiple stereo depth maps. Our method selects depth estimates for each pixel that minimize violations of visibility constraints and thus remove errors and inconsistencies from the depth maps to produce a consistent surface. We advocate a two-stage process in which the first stage generates potentially noisy, overlapping depth maps from a set of calibrated images and the second stage fuses these depth maps to obtain an integrated surface with higher accuracy, suppressed noise, and reduced redundancy. We show that by dividing the processing into two stages we are able to achieve a very high throughput because we are able to use a computationally cheap stereo algorithm and because this architecture is amenable to hardware-accelerated (GPU) implementations. A rigorous formulation based on the notion of stability of a depth estimate is presented first. It aims to determine the validity of a depth estimate by rendering multiple depth maps into the reference view as well as rendering the reference depth map into the other views in order to detect occlusions and free- space violations. We also present an approximate alternative formulation that selects and validates only one hypothesis based on confidence. Both formulations enable us to perform video-based reconstruction at up to 25 frames per second. We show results on the multi-view stereo evaluation benchmark datasets and several outdoors video sequences. Extensive quantitative analysis is performed using an accurately surveyed model of a real building as ground truth. Paul Merrell, Amir Akbarzadeh, Liang Wang 0002, Philippos Mordohai, Jan-Michael Frahm, Ruigang Yang, David Nistér, Marc Pollefeys |
ICCV | 6 |
| 2007 | Automatic Natural Video Matting with DepthabstractVideo matting is the process of taking a sequence of frames, isolating the foreground, and replacing the background in each frame. We look at existing single-frame matting techniques and present a method that improves upon them by adding depth information acquired by a time-offlight range scanner. We use the depth information to automate the process so it can be practically used for video sequences. In addition, we show that we can improve the results from natural matting algorithms by adding a depth channel. The additional depth information allows us to reduce the artifacts that arise from ambiguities that occur when an object is a similar color to its background. Oliver Wang, Jonathan Finger, Qingxiong Yang, James Davis 0001, Ruigang Yang |
PG | 5 |
| 2007 | A Performance Study on Different Cost Aggregation Approaches Used in Real-Time Stereo Matching
Minglun Gong, Ruigang Yang, Liang Wang 0002, Mingwei Gong |
Int. J. Comput. Vis. | 2 |
| 2007 | Restoring 2D Content from Distorted DocumentsabstractThis paper presents a framework to restore the 2D content printed on documents in the presence of geometric distortion and non-uniform illumination. Compared with textbased document imaging approaches that correct distortion to a level necessary to obtain sufficiently readable text or to facilitate optical character recognition (OCR), our work targets nontextual documents where the original printed content is desired. To achieve this goal, our framework acquires a 3D scan of the document's surface together with a high-resolution image. Conformal mapping is used to rectify geometric distortion by mapping the 3D surface back to a plane while minimizing angular distortion. This conformal "deskewing" assumes no parametric model of the document's surface and is suitable for arbitrary distortions. Illumination correction is performed by using the 3D shape to distinguish content gradient edges from illumination gradient edges in the high-resolution image. Integration is performed using only the content edges to obtain a reflectance image with significantly less illumination artifacts. This approach makes no assumptions about light sources and their positions. The results from the geometric and photometric correction are combined to produce the final output. Michael S. Brown, Mingxuan Sun 0001, Ruigang Yang, Yun Lin 0011, W. Brent Seales |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2007 | BRDF Invariant Stereo Using Light Transport ConstancyabstractNearly all existing methods for stereo reconstruction assume that scene reflectance is Lambertian and make use of brightness constancy as a matching invariant. We introduce a new invariant for stereo reconstruction called light transport constancy (LTC), which allows completely arbitrary scene reflectance (bidirectional reflectance distribution functions (BRDFs)). This invariant can be used to formulate a rank constraint on multiview stereo matching when the scene is observed by several lighting configurations in which only the lighting intensity varies. In addition, we show that this multiview constraint can be used with as few as two cameras and two lighting configurations. Unlike previous methods for BRDF invariant stereo, LTC does not require precisely configured or calibrated light sources or calibration objects in the scene. Importantly, the new constraint can be used to provide BRDF invariance to any existing stereo method whenever appropriate lighting variation is available. Liang Wang 0002, Ruigang Yang, James Davis 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Space-Time Light Field RenderingabstractIn this paper, we propose a novel framework called space-time light field rendering, which allows continuous exploration of a dynamic scene in both space and time. Compared to existing light field capture/rendering systems, it offers the capability of using unsynchronized video inputs and the added freedom of controlling the visualization in the temporal domain, such as smooth slow motion and temporal integration. In order to synthesize novel views from any viewpoint at any time instant, we develop a two-stage rendering algorithm. We first interpolate in the temporal domain to generate globally synchronized images using a robust spatial-temporal image registration algorithm followed by edge-preserving image morphing. We then interpolate these software-synchronized images in the spatial domain to synthesize the final view. In addition, we introduce a very accurate and robust algorithm to estimate subframe temporal offsets among input video sequences. Experimental results from unsynchronized videos with or without time stamps show that our approach is capable of maintaining photorealistic quality from a variety of real scenes. Huamin Wang 0001, Mingxuan Sun 0001, Ruigang Yang |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2006 | Real-time Global Stereo Matching Using Hierarchical Belief PropagationabstractIn this paper, we present a belief propagation based global algorithm that generates high quality results while maintaining real-time performance. To our knowledge, it is the first BP based global method that runs at real-time speed. Our efficiency performance gains mainly from the parallelism of graphics hardware,which leads to a 45 times speedup compared to the CPU implementation. To qualify the accurancy of our approach, the experimental results are evaluated on the Middlebury data sets, showing that our approach is among the best (ranked first in the new evaluation system) for all real-time approaches. In addition, since the running time of general BP is linear to the number of iterations, adopting a large number of iterations is not feasible for practical applications. Hence a novel approach is proposed to adaptively update pixel cost. Unlike general BP methods, the running time of our proposed algorithm dramatically converges. Qingxiong Yang, Liang Wang 0002, Ruigang Yang, Miao Liao, David Nistér |
BMVC | 3 |
| 2006 | Stereo Matching with Color-Weighted Correlation, Hierarchical Belief Propagation and Occlusion HandlingabstractIn this paper, we formulate an algorithm for the stereo matching problem with careful handling of disparity, discontinuity and occlusion. The algorithm works with a global matching stereo model based on an energy- minimization framework. The global energy contains two terms, the data term and the smoothness term. The data term is first approximated by a color-weighted correlation, then refined in occluded and low-texture areas in a repeated application of a hierarchical loopy belief propagation algorithm. The experimental results are evaluated on the Middlebury data set, showing that our algorithm is the top performer. Qingxiong Yang, Liang Wang 0002, Ruigang Yang, Henrik Stewénius, David Nistér |
CVPR (2) | 3 |
| 2006 | View-dependent textured splatting
Ruigang Yang, David Guinnip, Liang Wang 0002 |
Vis. Comput. | 1 |
| 2005 | BRDF Invariant Stereo Using Light Transport ConstancyabstractNearly all existing methods for stereo reconstruction assume that scene reflectance is Lambertian, and make use of color constancy as a matching invariant. We introduce a new invariant for stereo reconstruction called light transport constancy, which allows completely arbitrary scene reflectance (BRDFs). This invariant can be used to formulate a rank constraint on multiview stereo matching when the scene is observed in several lighting configurations. In addition, we show that this multiview constraint can be used with as few as two cameras and two lighting configurations. Unlikely previous methods for BRDF invariant stereo, light transport constancy does not require precisely configured or calibrated light sources, nor calibration objects in the scene. Importantly, the new constraint can be used to provide BRDF invariance to any existing stereo method, whenever appropriate lighting variation is available. James Davis 0001, Ruigang Yang, Liang Wang 0002 |
ICCV | 2 |
| 2005 | Geometric and Photometric Restoration of Distorted DocumentsabstractWe present a system to restore the 2D content printed on distorted documents. Our system works by acquiring a 3D scan of the document's surface together with a high-resolution image. Using the 3D surface information and the 2D image, we can ameliorate unwanted surface distortion and effects from non-uniform illumination. Our system can process arbitrary geometric distortions, not requiring any pre-assumed parametric models for the document's geometry. The illumination correction uses the 3D shape to distinguish content edges from illumination edges to recover the 2D content's reflectance image while making no assumptions about light sources and their positions. Results are shown for real objects, demonstrating a complete framework capable of restoring geometric and photometric artifacts on distorted documents Mingxuan Sun 0001, Ruigang Yang, Yun Lin 0011, George V. Landon, W. Brent Seales, Michael S. Brown |
ICCV | 2 |
| 2005 | Towards space: time light field renderingabstractSo far extending light field rendering to dynamic scenes has been trivially treated as the rendering of static light fields stacked in time. This type of approaches requires input video sequences in strict synchronization and allows only discrete exploration in the temporal domain determined by the capture rate. In this paper we propose a novel framework, space-time light field rendering, which allows continuous exploration of a dynamic scene in both spatial and temporal domain with unsynchronized input video sequences.In order to synthesize novel views from any viewpoint at any time instant, we develop a two-stage rendering algorithm. We first interpolate in the temporal domain to generate globally synchronized images using a robust spatial-temporal image registration algorithm followed by edge-preserving image morphing. We then interpolate those software-synchronized images in the spatial domain to synthesize the final view. Our experimental results show that our approach is robust and capable of maintaining photo-realistic results. Huamin Wang 0001, Ruigang Yang |
SI3D | 2 |
| 2005 | Camera-Based Calibration Techniques for Seamless Multiprojector DisplaysabstractMultiprojector, large-scale displays are used in scientific visualization, virtual reality, and other visually intensive applications. In recent years, a number of camera-based computer vision techniques have been proposed to register the geometry and color of tiled projection-based display. These automated techniques use cameras to "calibrate" display geometry and photometry, computing per-projector corrective warps and intensity corrections that are necessary to produce seamless imagery across projector mosaics. These techniques replace the traditional labor-intensive manual alignment and maintenance steps, making such displays cost-effective, flexible, and accessible. In this paper, we present a survey of different camera-based geometric and photometric registration techniques reported in the literature to date. We discuss several techniques that have been proposed and demonstrated, each addressing particular display configurations and modes of operation. We overview each of these approaches and discuss their advantages and disadvantages. We examine techniques that address registration on both planar (video walls) and arbitrary display surfaces and photometric correction for different kinds of display surfaces. We conclude with a discussion of the remaining challenges and research opportunities for multiprojector displays. Michael S. Brown, Aditi Majumder, Ruigang Yang |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2004 | Eye Gaze Correction with Stereovision for Video-TeleconferencingabstractThe lack of eye contact in desktop video teleconferencing substantially reduces the effectiveness of video contents. While expensive and bulky hardware is available on the market to correct eye gaze, researchers have been trying to provide a practical software-based solution to bring video-teleconferencing one step closer to the mass market. This paper presents a novel approach: Based on stereo analysis combined with rich domain knowledge (a personalized face model), we synthesize, using graphics hardware, a virtual video that maintains eye contact. A 3D stereo head tracker with a personalized face model is used to compute initial correspondences across two views. More correspondences are then added through template and feature matching. Finally, all the correspondence information is fused together for view synthesis using view morphing techniques. The combined methods greatly enhance the accuracy and robustness of the synthesized views. Our current system is able to generate an eye-gaze corrected video stream at five frames per second on a commodity 1 GHz PC. Ruigang Yang, Zhengyou Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2003 | Multi-Resolution Real-Time Stereo on Commodity Graphics HardwareabstractIn this paper a stereo algorithm suitable for implementation on commodity graphics hardware is presented. This is important since it allows freeing up the main processor for other tasks including high-level interpretation of the stereo results. Our algorithm relies on the traditional sum-of-square-differences (SSD) dissimilarity measure between correlation windows. To achieve good results close to depth discontinuities as well as on low texture areas, a multi-resolution approach is used. The approach efficiently combines SSD measurements for windows of different sizes. Our implementation running on an NVIDIA GeForce4 graphics card achieves 50-70M disparity evaluations per second including all the overhead to download images and read-back the disparity map, which is equivalent to the fastest commercial CPU implementations available. An important advantage of our approach is that rectification is not necessary so that correspondences can just as easily be obtained for images that contain the epipoles. Another advantage is that this approach can easily be extended to multi-baseline stereo. Ruigang Yang, Marc Pollefeys |
CVPR (1) | 1 |
| 2003 | Dealing with Textureless Regions and Specular Highlights - A Progressive Space Carving Scheme Using a Novel Photo-consistency MeasureabstractWe present two extensions to the space carving framework. The first is a progressive scheme to better reconstruct surfaces lacking sufficient textures. The second is a novel photo-consistency measure that is valid for both specular and diffuse surfaces, under unknown lighting conditions. Ruigang Yang, Marc Pollefeys, Greg Welch |
ICCV | 1 |
| 2003 | Real-Time Consensus-Based Scene Reconstruction Using Commodity Graphics HardwareabstractAbstract We present a novel use of commodity graphics hardware that effectively combines a plane‐sweeping algorithm with view synthesis for real‐time, online 3D scene acquisition and view synthesis. Using real‐time imagery from a few calibrated cameras, our method can generate new images from nearby viewpoints, estimate a dense depth map from the current viewpoint, or create a textured triangular mesh. We can do each of these without any prior geometric information or requiring any user interaction, in real time and online. The heart of our method is to use programmable Pixel Shader technology to square intensity differences between reference image pixels, and then to choose final colors (or depths) that correspond to the minimum difference, i.e. the most consistent color. In this paper we describe the method, place it in the context of related work in computer graphics and computer vision, and present some results. ACM CSS: I.3.3 Computer Graphics—Bitmap and framebuffer operations, I.4.8 Image Processing and Computer Vision—Depth cues, Stereo Ruigang Yang, Greg Welch, Gary Bishop |
Comput. Graph. Forum | 1 |
| 2002 | Eye Gaze Correction with Stereovision for Video-Teleconferencing
Ruigang Yang, Zhengyou Zhang |
ECCV (2) | 1 |
| 2002 | Real-Time Consensus-Based Scene Reconstruction Using Commodity Graphics HardwareabstractWe present a novel use of commodity graphics hardware that effectively combines a plane-sweeping algorithm with view synthesis for real-time, on-line 3D scene acquisition and view synthesis. Using real-time imagery from a few calibrated cameras, our method can generate new images from nearby viewpoints, estimate a dense depth map from the current viewpoint, or create a textured triangular mesh. We can do this without prior geometric information or requiring any user interaction, in real time and on line. The heart of our method is using programmable pixel shader technology to square intensity differences between reference image pixels, and then to choose final colors (or depths) that correspond to the minimum difference, i.e. the most consistent color. In this paper we describe the method, place it in the context of related work in computer graphics and computer vision, and present results. Ruigang Yang, Greg Welch, Gary Bishop |
PG | 1 |
| 2001 | PixelFlex: A Reconfigurable Multi-Projector Display SystemabstractThis paper presents PixelFlex - a spatially reconfigurable multi-projector display system. The PixelFlex system is composed of ceiling-mounted projectors, each with computer-controlled pan, tilt, zoom and focus; and a camera for closed-loop calibration. Working collectively, these controllable projectors function as a single logical display capable of being easily modified into a variety of spatial formats of differing pixel density, size and shape. New layouts are automatically calibrated within minutes to generate the accurate warping and blending functions needed to produce seamless imagery across planar display surfaces, thus giving the user the flexibility to quickly create, save and restore multiple screen configurations. Overall, PixelFlex provides a new level of automatic reconfigurability and usage, departing from the static, one-size-fits-all design of traditional large-format displays. As a front-projection system, PixelFlex can be installed in most environments with space constraints and requires little or no post-installation mechanical maintenance because of the closed-loop calibration. Ruigang Yang, David Gotz, Justin Hensley, Herman Towles, Michael S. Brown |
IEEE Visualization | 1 |
| 1999 | Geometrically correct imagery for teleconferencingabstractCurrent camera-monitor teleconferencing applications produce unrealistic imagery and break any sense of presence for the participants. Other capture/display technologies can be used to provide more compelling teleconferencing. However, complex geometries in capture/display systems make producing geometrically correct imagery difficult. It is usually impractical to detect, model and compensate for all effects introduced by the capture/display system. Most applications simply ignore these issues and rely on the user acceptance of the camera-monitor paradigm. Ruigang Yang, Michael S. Brown, W. Brent Seales, Henry Fuchs |
ACM Multimedia (1) | 1 |
| 1999 | Multi-Projector Displays Using Camera-Based RegistrationabstractConventional projector-based display systems are typically designed around precise and regular configurations of projectors and display surfaces. While this results in rendering simplicity and speed, it also means painstaking construction and ongoing maintenance. In previously published work, we introduced a vision of projector-based displays constructed from a collection of casually-arranged projectors and display surfaces. In this paper, we present flexible yet practical methods for realizing this vision, enabling low-cost mega-pixel display systems with large physical dimensions, higher resolution, or both. The techniques afford new opportunities to build personal 3D visualization systems in offices, conference rooms, theaters, or even your living room. As a demonstration of the simplicity and effectiveness of the methods that we continue to perfect, we show in the included video that a 10-year old child can construct and calibrate a two-camera, two-projector, head-tracked display system, all in about 15 minutes. Ramesh Raskar, Michael S. Brown, Ruigang Yang, Wei-Chao Chen, Greg Welch, Herman Towles, W. Brent Seales, Henry Fuchs |
IEEE Visualization | 3 |
| 1998 | Registering, Integrating and Building CAD Models from Range DataabstractWe introduce two methods for the registration of range images when a prior estimate of the transformation between views is not available and the overlap between images is relatively small. The methods are an extension to the work of Gueziec and Ayache (1994) and Turk and Levoy (1994) and consists of 2 stages. First, we find the initial estimated transformation by extracting and matching 3D space curves from different scans of the same object. If no salient features are available on the object we use fiducial marks to find the initial transformation. This allows us to always find a satisfactory and even highly accurate transformation independent of the geometry of the object. Second, we apply a modified iterative closest points algorithm (ICP) to improve the accuracy of registration. We define a weighted distance function based on surface curvature which can reduce the number of iterations and requires a less accurate initial estimate of the transformation. Ruigang Yang, Peter K. Allen |
ICRA | 1 |