VLDB 2026 Research / reviewers in the wild / expert
Ulrich Neumann
dblp:52/4897
· DBLP profile ↗
130ranked-venue papers
6as first author
18since 2021 · last 2025
0000-0001-8977-7112ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 112 · 6 first-author · 13 since 2021Artificial intelligence and machine learning · 45 · 16 since 2021Human-computer interaction and ubiquitous computing · 31 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 5Databases, data management, data science and information retrieval · 4Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Comprehensive Relighting: Generalizable and Consistent Monocular Human Relighting and HarmonizationabstractThis paper introduces Comprehensive Relighting, the first all-in-one approach that can both control and harmonize the lighting from an image or video of humans with arbitrary body parts from any scene. Building such a generalizable model is extremely challenging due to the lack of dataset, restricting existing image-based relighting models to a specific scenario (e.g., face or static human). To address this challenge, we repurpose a pre-trained diffusion model as a general image prior and jointly model the human relighting and background harmonization in the coarse-to-fine framework. To further enhance the temporal coherence of the relighting, we introduce an unsupervised temporal lighting model that learns the lighting cycle consistency from many real-world videos without any ground truth. In inference time, our temporal lighting module is combined with the diffusion models through the spatio-temporal feature blending algorithms without extra training; and we apply a new guided refinement as a post-processing to pre-serve the high-frequency details from the input image. In the experiments, Comprehensive Relighting shows a strong generalizability and lighting temporal coherence, outperforming existing image-based human relighting and harmonization methods. Xin Sun 0014, Krishna Kumar Singh, Zhixin Shu, He Zhang 0004, Jimei Yang, Nanxuan Zhao, Tuanfeng Y. Wang, Simon S. Chen, Ulrich Neumann, Jae Shin Yoon |
CVPR | 11 |
| 2025 | Can3Tok: Canonical 3D Tokenization and Latent Modeling of Scene-Level 3D Gaussiansabstract3D generation has made significant progress, however, it still largely remains at the object-level. Feedforward 3D scene-level generation has been rarely explored due to the lack of models capable of scaling-up latent representation learning on 3D scene-level data. Unlike object-level generative models, which are trained on well-labeled 3D data in a bounded canonical space, scene-level generations with 3D scenes represented by 3D Gaussian Splatting (3DGS) are unbounded and exhibit scale inconsistency across different scenes, making unified latent representation learning for generative purposes extremely challenging. In this paper, we introduce Can3Tok, the first 3D scene-level variational autoencoder (VAE) capable of encoding a large number of Gaussian primitives into a low-dimensional latent embedding, which effectively captures both semantic and spatial information of the inputs. Beyond model design, we propose a general pipeline for 3D scene data processing to address scale inconsistency issue. We validate our method on the recent scene-level 3D dataset DL3DV-10K, where we found that only Can3Tok successfully generalizes to novel 3D scenes, while compared methods fail to converge on even a few hundred scene inputs during training and exhibit zero generalization ability during inference. Finally, we demonstrate image-to-3DGS and text-to-3DGS generation as our applications to demonstrate its ability to facilitate downstream generation tasks. Quankai Gao, Iliyan Georgiev, Tuanfeng Y. Wang, Krishna Kumar Singh, Ulrich Neumann, Jae Shin Yoon |
ICCV | 5 |
| 2025 | DGH: Dynamic Gaussian HairabstractThe creation of photorealistic dynamic hair remains a major challenge in digital human modeling because of the complex motions, occlusions, and light scattering. Existing methods often resort to static capture and physics-based models that do not scale as they require manual parameter fine-tuning to handle the diversity of hairstyles and motions, and heavy computation to obtain high-quality appearance. In this paper, we present Dynamic Gaussian Hair (DGH), a novel framework that efficiently learns hair dynamics and appearance. We propose: (1) a coarse-to-fine model that learns temporally coherent hair motion dynamics across diverse hairstyles; (2) a strand-guided optimization module that learns a dynamic 3D Gaussian representation for hair appearance with support for differentiable rendering, enabling gradient-based learning of view-consistent appearance under motion. Unlike prior simulation-based pipelines, our approach is fully data-driven, scales with training data, and generalizes across various hairstyles and head motion sequences. Additionally, DGH can be seamlessly integrated into a 3D Gaussian avatar framework, enabling realistic, animatable hair for high-fidelity avatar representation. DGH achieves promising geometry and appearance results, providing a scalable, data-driven alternative to physics-based simulation and rendering. Yuanlu Xu, Edith Tretschk, Anastasia Ianina, Aljaz Bozic, Ulrich Neumann, Tony Tung |
NeurIPS | 7 |
| 2024 | InSpaceType: Dataset and Benchmark for Reconsidering Cross-Space Type Performance in Indoor Monocular Depth
Cho-Ying Wu, Quankai Gao, Chin-Cheng Hsu, Te-Lin Wu, Jing-Wen Chen, Ulrich Neumann |
BMVC | 6 |
| 2024 | Boosting Generalizability towards Zero-Shot Cross-Dataset Single-Image Indoor Depth by Meta-InitializationabstractIndoor robots rely on depth to perform tasks like navigation or obstacle detection, and single-image depth estimation is widely used to assist perception. Most indoor single-image depth prediction focuses less on model generalizability to unseen datasets, concerned with in-the-wild robustness for system deployment. This work leverages gradient-based meta-learning to gain higher generalizability on zero-shot cross-dataset inference. Unlike the most-studied meta-learning of image classification associated with explicit class labels, no explicit task boundaries exist for continuous depth values tied to highly varying indoor environments regarding object arrangement and scene composition. We propose fine-grained task that treats each RGB-D mini-batch as a task in our meta-learning formulation. We first show that our method on limited data induces a much better prior (max 27.8% in RMSE). Then, finetuning on meta-learned initialization consistently outperforms baselines without the meta approach. Aiming at generalization, we propose zero-shot cross-dataset protocols and validate higher generalizability induced by our meta-initialization, as a simple and useful plugin to many existing depth estimation methods. The work at the intersection of depth and meta-learning potentially drives both research to step closer to practical robotic and machine perception usage. Cho-Ying Wu, Yiqi Zhong, Ulrich Neumann |
IROS | 4 |
| 2024 | Motion Graph Unleashed: A Novel Approach to Video PredictionabstractWe introduce motion graph, a novel approach to address the video prediction problem, i.e., predicting future video frames from limited past data. The motion graph transforms patches of video frames into interconnected graph nodes, to comprehensively describe the spatial-temporal relationships among them. This representation overcomes the limitations of existing motion representations such as image differences, optical flow, and motion matrix that either fall short in capturing complex motion patterns or suffer from excessive memory consumption. We further present a video prediction pipeline empowered by motion graph, exhibiting substantial performance improvements and cost reductions. Extensive experiments on various datasets, including UCF Sports, KITTI and Cityscapes, highlight the strong representative ability of motion graph. Especially on UCF Sports, our method matches and outperforms the SOTA methods with a significant reduction in model size by 78% and a substantial decrease in GPU memory utilization by 47%. Yiqi Zhong, Luming Liang, Bohan Tang, Ilya Zharkov, Ulrich Neumann |
NeurIPS | 5 |
| 2023 | Complete 3D Human Reconstruction from a Single Incomplete ImageabstractThis paper presents a method to reconstruct a complete human geometry and texture from an image of a person with only partial body observed, e.g., a torso. The core challenge arises from the occlusion: there exists no pixel to reconstruct where many existing single-view human reconstruction methods are not designed to handle such invisible parts, leading to missing data in 3D. To address this challenge, we introduce a novel coarse-to-fine human reconstruction framework. For coarse reconstruction, explicit volumetric features are learned to generate a complete human geometry with 3D convolutional neural networks conditioned by a 3D body model and the style features from visible parts. An implicit network combines the learned 3D features with the high-quality surface normals enhanced from multiviews to produce fine local details, e.g., high-frequency wrinkles. Finally, we perform progressive texture inpainting to reconstruct a complete appearance of the person in a view-consistent way, which is not possible without the reconstruction of a complete geometry. In experiments, we demonstrate that our method can reconstruct high-quality 3D humans, which is robust to occlusion. Jae Shin Yoon, Tuanfeng Y. Wang, Krishna Kumar Singh, Ulrich Neumann |
CVPR | 5 |
| 2023 | Strivec: Sparse Tri-Vector Radiance FieldsabstractWe propose Strivec, a novel neural representation that models a 3D scene as a radiance field with sparsely distributed and compactly factorized local tensor feature grids. Our approach leverages tensor decomposition, following the recent work TensoRF [6], to model the tensor grids. In contrast to TensoRF which uses a global tensor and focuses on their vector-matrix decomposition, we propose to utilize a cloud of local tensors and apply the classic CANDE-COMP/PARAFAC (CP) decomposition [4] to factorize each tensor into triple vectors that express local feature distributions along spatial axes and compactly encode a local neural field. We also apply multi-scale tensor grids to discover the geometry and appearance commonalities and exploit spatial coherence with the tri-vector factorization at multiple local scales. The final radiance field properties are regressed by aggregating neural features from multiple local tensors across all scales. Our tri-vector tensors are sparsely distributed around the actual scene surface, discovered by a fast coarse reconstruction, leveraging the sparsity of a 3D scene. We demonstrate that our model can achieve better rendering quality while using significantly fewer parameters than previous methods, including TensoRF and Instant-NGP [23]. Quankai Gao, Qiangeng Xu, Hao Su 0001, Ulrich Neumann, Zexiang Xu |
ICCV | 4 |
| 2023 | MMVP: Motion-Matrix-based Video PredictionabstractA central challenge of video prediction lies where the system has to reason the objects’ future motions from image frames while simultaneously maintaining the consistency of their appearances across frames. This work introduces an end-to-end trainable two-stream video prediction framework, Motion-Matrix-based Video Prediction (MMVP), to tackle this challenge. Unlike previous methods that usually handle motion prediction and appearance maintenance within the same set of modules, MMVP decouples motion and appearance information by constructing appearance-agnostic motion matrices. The motion matrices represent the temporal similarity of each and every pair of feature patches in the input frames, and are the sole input of the motion prediction module in MMVP. This design improves video prediction in both accuracy and efficiency, and reduces the model size. Results of extensive experiments demonstrate that MMVP outperforms state-of-the-art systems on public data sets by non-negligible large margins (≈ 1 db in PSNR, UCF Sports) in significantly smaller model sizes (84% the size or smaller). Please refer to this $link$ for the official code and the datasets used in this paper. Yiqi Zhong, Luming Liang, Ilya Zharkov, Ulrich Neumann |
ICCV | 4 |
| 2023 | Collaborative Uncertainty Benefits Multi-Agent Multi-Modal Trajectory ForecastingabstractIn multi-modal multi-agent trajectory forecasting, two major challenges have not been fully tackled: 1) how to measure the uncertainty brought by the interaction module that causes correlations among the predicted trajectories of multiple agents; 2) how to rank the multiple predictions and select the optimal predicted trajectory. In order to handle the aforementioned challenges, this work first proposes a novel concept, collaborative uncertainty (CU), which models the uncertainty resulting from interaction modules. Then we build a general CU-aware regression framework with an original permutation-equivariant uncertainty estimator to do both tasks of regression and uncertainty estimation. Furthermore, we apply the proposed framework to current SOTA multi-agent multi-modal forecasting systems as a plugin module, which enables the SOTA systems to: 1) estimate the uncertainty in the multi-agent multi-modal trajectory forecasting task; 2) rank the multiple predictions and select the optimal one based on the estimated uncertainty. We conduct extensive experiments on a synthetic dataset and two public large-scale multi-agent trajectory forecasting benchmarks. Experiments show that: 1) on the synthetic dataset, the CU-aware regression framework allows the model to appropriately approximate the ground-truth Laplace distribution; 2) on the multi-agent trajectory forecasting benchmarks, the CU-aware regression framework steadily helps SOTA systems improve their performances. Especially, the proposed framework helps VectorNet improve by 262 cm regarding the Final Displacement Error of the chosen optimal prediction on the nuScenes dataset; 3) in multi-agent multi-modal trajectory forecasting, prediction uncertainty is proportional to future stochasticity; 4) the estimated CU values are highly related to the interactive information among agents. The proposed framework can guide the development of more reliable and safer forecasting systems in the future. Bohan Tang, Yiqi Zhong, Chenxin Xu, Ulrich Neumann, Ya Zhang 0002, Siheng Chen, Yanfeng Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Behind the Curtain: Learning Occluded Shapes for 3D Object DetectionabstractAdvances in LiDAR sensors provide rich 3D data that supports 3D scene understanding. However, due to occlusion and signal miss, LiDAR point clouds are in practice 2.5D as they cover only partial underlying shapes, which poses a fundamental challenge to 3D perception. To tackle the challenge, we present a novel LiDAR-based 3D object detection model, dubbed Behind the Curtain Detector (BtcDet), which learns the object shape priors and estimates the complete object shapes that are partially occluded (curtained) in point clouds. BtcDet first identifies the regions that are affected by occlusion and signal miss. In these regions, our model predicts the probability of occupancy that indicates if a region contains object shapes and integrates this probability map with detection features and generates high-quality 3D proposals. Finally, the occupancy estimation is integrated into the proposal refinement module to generate accurate bounding boxes. Extensive experiments on the KITTI Dataset and the Waymo Open Dataset demonstrate the effectiveness of BtcDet. Particularly for the 3D detection of both cars and cyclists on the KITTI benchmark, BtcDet surpasses all of the published state-of-the-art methods by remarkable margins. Code is released. Qiangeng Xu, Yiqi Zhong, Ulrich Neumann |
AAAI | 3 |
| 2022 | Cross-Modal Perceptionist: Can Face Geometry be Gleaned from Voices?abstractThis work digs into a root question in human perception: can face geometry be gleaned from one's voices? Previous works that study this question only adopt developments in image synthesis and convert voices into face images to show correlations, but working on the image domain unavoidably involves predicting attributes that voices cannot hint, including facial textures, hairstyles, and backgrounds. We instead investigate the ability to reconstruct 3D faces to concentrate on only geometry, which is much more physiologically grounded. We propose our analysis framework, Cross-Modal Perceptionist, under both supervised and unsupervised learning. First, we construct a dataset, Voxceleb-3D, which extends Voxceleb and includes paired voices and face meshes, making supervised learning possible. Second, we use a knowledge distillation mechanism to study whether face geometry can still be gleaned from voices without paired voices and 3D face data under limited availability of 3D face scans. We break down the core question into four parts and perform visual and numerical analyses as responses to the core question. Our findings echo those in physiology and neuroscience about the correlation between voices and facial structures. The work provides future human-centric cross-modal learning with explainable foundations. See our project page. Cho-Ying Wu, Chin-Cheng Hsu, Ulrich Neumann |
CVPR | 3 |
| 2022 | Toward Practical Monocular Indoor Depth EstimationabstractThe majority of prior monocular depth estimation meth-ods without groundtruth depth guidance focus on driving scenarios. We show that such methods generalize poorly to unseen complex indoor scenes, where objects are cluttered and arbitrarily arranged in the near field. To obtain more robustness, we propose a structure distillation approach to learn knacks from an off-the-shelf relative depth estima-tor that produces structured but metric-agnostic depth. By combining structure distillation with a branch that learns metrics from left-right consistency, we attain structured and metric depth for generic indoor scenes and make inferences in real-time. To facilitate learning and evaluation, we col-lect SimSIN, a dataset from simulation with thousands of environments, and UniSIN, a dataset that contains about 500 real scan sequences of generic indoor environments. We experiment in both sim-to-real and real-to-real settings, and show improvements, as well as in downstream applications using our depth maps. This work provides a full study, covering methods, data, and applications aspects. Cho-Ying Wu, Jialiang Wang 0001, Michael Hall, Ulrich Neumann, Shuochen Su |
CVPR | 4 |
| 2022 | Point-NeRF: Point-based Neural Radiance FieldsabstractVolumetric neural rendering methods like NeRF [34] generate high-quality view synthesis results but are optimized per-scene leading to prohibitive reconstruction time. On the other hand, deep multi-view stereo methods can quickly reconstruct scene geometry via direct network inference. Point-NeRF combines the advantages of these two approaches by using neural 3D point clouds, with associated neural features, to model a radiance field. Point-NeRF can be rendered efficiently by aggregating neural point features near scene surfaces, in a ray marching-based rendering pipeline. Moreover, Point-NeRF can be initialized via direct inference of a pre-trained deep network to produce a neural point cloud; this point cloud can be finetuned to surpass the visual quality of NeRF with 30× faster training time. Point-NeRF can be combined with other 3D re-construction methods and handles the errors and outliers in such methods via a novel pruning and growing mechanism. Qiangeng Xu, Zexiang Xu, Julien Philip, Sai Bi, Zhixin Shu, Kalyan Sunkavalli, Ulrich Neumann |
CVPR | 7 |
| 2022 | Aware of the History: Trajectory Forecasting with the Local Behavior Data
Yiqi Zhong, Zhenyang Ni, Siheng Chen, Ulrich Neumann |
ECCV (22) | 4 |
| 2021 | Synergy between 3DMM and 3D Landmarks for Accurate 3D Facial GeometryabstractThis work studies learning from a synergy process of 3D Morphable Models (3DMM) and 3D facial landmarks to predict complete 3D facial geometry, including 3D alignment, face orientation, and 3D face modeling. Our synergy process leverages a representation cycle for 3DMM parameters and 3D landmarks. 3D landmarks can be extracted and refined from face meshes built by 3DMM parameters. We next reverse the representation direction and show that predicting 3DMM parameters from sparse 3D landmarks improves the information flow. Together we create a synergy process that utilizes the relation between 3D landmarks and 3DMM parameters, and they collaboratively contribute to better performance. We extensively validate our contribution on full tasks of facial geometry prediction and show our superior and robust performance on these tasks for various scenarios. Particularly, we adopt only simple and widely-used network operations to attain fast and accurate facial geometry prediction. Codes and data: https: //choyingw.github.io/works/SynergyNet/. Cho-Ying Wu, Qiangeng Xu, Ulrich Neumann |
3DV | 3 |
| 2021 | Scene Completeness-Aware Lidar Depth Completion for Driving ScenarioabstractThis paper introduces Scene Completeness-Aware Depth Completion (SCADC) to complete raw lidar scans into dense depth maps with fine and complete scene structures. Recent sparse depth completion for lidars only focuses on the lower scenes and produces irregular estimations on the upper because existing datasets, such as KITTI, do not provide groundtruth for upper areas. These areas are considered less important since they are usually sky or trees of less scene understanding interest. However, we argue that in several driving scenarios such as large trucks or cars with loads, objects could extend to the upper parts of scenes. Thus depth maps with structured upper scene estimation are important for RGBD algorithms. SCADC adopts stereo images that produce disparities with better scene completeness but are generally less precise than lidars, to help sparse lidar depth completion. To our knowledge, we are the first to focus on scene completeness of sparse depth completion. We validate our SCADC on both depth estimate precision and scene-completeness on KITTI. Moreover, we experiment on less-explored outdoor RGBD semantic segmentation with scene completeness-aware D-input to validate our method. Cho-Ying Wu, Ulrich Neumann |
ICASSP | 2 |
| 2021 | Collaborative Uncertainty in Multi-Agent Trajectory ForecastingabstractUncertainty modeling is critical in trajectory-forecasting systems for both interpretation and safety reasons. To better predict the future trajectories of multiple agents, recent works have introduced interaction modules to capture interactions among agents. This approach leads to correlations among the predicted trajectories. However, the uncertainty brought by such correlations is neglected. To fill this gap, we propose a novel concept, collaborative uncertainty (CU), which models the uncertainty resulting from the interaction module. We build a general CU-based framework to make a prediction model learn the future trajectory and the corresponding uncertainty. The CU-based framework is integrated as a plugin module to current state-of-the-art (SOTA) systems and deployed in two special cases based on multivariate Gaussian and Laplace distributions. In each case, we conduct extensive experiments on two synthetic datasets and two public, large-scale benchmarks of trajectory forecasting. The results are promising: 1) The results of synthetic datasets show that CU-based framework allows the model to nicely rebuild the ground-truth distribution. 2) The results of trajectory forecasting benchmarks demonstrate that the CU-based framework steadily helps SOTA systems improve their performances. Specially, the proposed CU-based framework helps VectorNet improve by 57 cm regarding Final Displacement Error on nuScenes dataset. 3) The visualization results of CU illustrate that the value of CU is highly related to the amount of the interactive information among agents. Bohan Tang, Yiqi Zhong, Ulrich Neumann, Siheng Chen, Ya Zhang 0002 |
NeurIPS | 3 |
| 2020 | Modeling Cross-Modal Interaction in a Multi-detector, Multi-modal Tracking Framework
Yiqi Zhong, Suya You, Ulrich Neumann |
ACCV (2) | 3 |
| 2020 | Grid-GCN for Fast and Scalable Point Cloud LearningabstractDue to the sparsity and irregularity of the point cloud data, methods that directly consume points have become popular. Among all point-based models, graph convolutional networks (GCN) lead to notable performance by fully preserving the data granularity and exploiting point interrelation. However, point-based networks spend a significant amount of time on data structuring (e.g., Farthest Point Sampling (FPS) and neighbor points querying), which limit the speed and scalability. In this paper, we present a method, named Grid-GCN, for fast and scalable point cloud learning. Grid-GCN uses a novel data structuring strategy, Coverage-Aware Grid Query (CAGQ). By leveraging the efficiency of grid space, CAGQ improves spatial coverage while reducing the theoretical time complexity. Compared with popular sampling methods such as Farthest Point Sampling (FPS) and Ball Query, CAGQ achieves up to 50 times speed-up. With a Grid Context Aggregation (GCA) module, Grid-GCN achieves state-of-the-art performance on major point cloud classification and segmentation benchmarks with significantly faster runtime than previous studies. Remarkably, Grid-GCN achieves the inference speed of 50FPS on ScanNet using 81920 points as input. The supplementary xharlie.github.io/papers/GGCN_supCamReady.pdf and the code github.com/xharlie/Grid-GCN are released. Qiangeng Xu, Xudong Sun 0005, Cho-Ying Wu, Panqu Wang, Ulrich Neumann |
CVPR | 5 |
| 2020 | Stochastic Dynamics for Video InfillingabstractIn this paper, we introduce a stochastic dynamics video infilling (SDVI) framework to generate frames between long intervals in a video. Our task differs from video interpolation which aims to produce transitional frames for a short interval between every two frames and increase the temporal resolution. Our task, namely video infilling, however, aims to infill long intervals with plausible frame sequences. Our framework models the infilling as a constrained stochastic generation process and sequentially samples dynamics from the inferred distribution. SDVI consists of two parts: (1) a bi-directional constraint propagation module to guarantee the spatial-temporal coherence among frames, (2) a stochastic sampling process to generate dynamics from the inferred distributions. Experimental results show that SDVI can generate clear frame sequences with varying contents. Moreover, motions in the generated sequence are realistic and able to transfer smoothly from the given start frame to the terminal frame. Qiangeng Xu, Hanwang Zhang, Weiyue Wang 0002, Peter N. Belhumeur, Ulrich Neumann |
WACV | 5 |
| 2019 | 3DN: 3D Deformation NetworkabstractApplications in virtual and augmented reality create a demand for rapid creation and easy access to large sets of 3D models. An effective way to address this demand is to edit or deform existing 3D models based on a reference, e.g., a 2D image which is very easy to acquire. Given such a source 3D model and a target which can be a 2D image, 3D model, or a point cloud acquired as a depth scan, we introduce 3DN, an end-to-end network that deforms the source model to resemble the target. Our method infers per-vertex offset displacements while keeping the mesh connectivity of the source model fixed. We present a training strategy which uses a novel differentiable operation, mesh sampling operator, to generalize our method across source and target models with varying mesh densities. Mesh sampling operator can be seamlessly integrated into the network to handle meshes with different topologies. Qualitative and quantitative results show that our method generates higher quality results compared to the state-of-the art learning-based methods for 3D shape generation. Weiyue Wang 0002, Duygu Ceylan, Radomír Mech, Ulrich Neumann |
CVPR | 4 |
| 2019 | Salient Building Outline Enhancement and Extraction Using Iterative L0 Smoothing and Line EnhancingabstractIn this paper, our goal is salient building outline enhancement and extraction from images taken from consumer cameras using L0smoothing. We address weak outlines and over-smoothing problem. Weak outlines are often undetected by edge extractors or easily smoothed out. We propose an iterative method, including the smoothing cell and sharpening cell. In the smoothing cell, we iteratively enlarge the smoothing level of the L0smoothing. In the sharpening cell, we use Hough Transform to extract lines, based on the assumption that salient outlines for buildings are usually straight, and enhance those extracted lines. Our goal is to enhance line structures and do the L0smoothing simultaneously. Also, we propose to create building masks from semantic segmentation using an encoder-decoder network. The masks filter out irrelevant edges. We also provide an evaluation dataset on this task. Cho-Ying Wu, Ulrich Neumann |
ICIP | 2 |
| 2019 | DISN: Deep Implicit Surface Network for High-quality Single-view 3D ReconstructionabstractReconstructing 3D shapes from single-view images has been a long-standing research problem. In this paper, we present DISN, a Deep Implicit Surface Net- work which can generate a high-quality detail-rich 3D mesh from a 2D image by predicting the underlying signed distance fields. In addition to utilizing global image features, DISN predicts the projected location for each 3D point on the 2D image and extracts local features from the image feature maps. Combin- ing global and local features significantly improves the accuracy of the signed distance field prediction, especially for the detail-rich areas. To the best of our knowledge, DISN is the first method that constantly captures details such as holes and thin structures present in 3D shapes from single-view images. DISN achieves the state-of-the-art single-view reconstruction performance on a variety of shape categories reconstructed from both synthetic and real images. Code is available at https://github.com/laughtervv/DISN. The supplemen- tary can be found at https://xharlie.github.io/images/neurips_ 2019_supp.pdf Qiangeng Xu, Weiyue Wang 0002, Duygu Ceylan, Radomír Mech, Ulrich Neumann |
NeurIPS | 5 |
| 2019 | Deep RGB-D Canonical Correlation Analysis For Sparse Depth CompletionabstractIn this paper, we propose our Correlation For Completion Network (CFCNet), an end-to-end deep learning model that uses the correlation between two data sources to perform sparse depth completion. CFCNet learns to capture, to the largest extent, the semantically correlated features between RGB and depth information. Through pairs of image pixels and the visible measurements in a sparse depth map, CFCNet facilitates feature-level mutual transformation of different data sources. Such a transformation enables CFCNet to predict features and reconstruct data of missing depth measurements according to their corresponding, transformed RGB features. We extend canonical correlation analysis to a 2D domain and formulate it as one of our training objectives (i.e. 2d deep canonical correlation, or “2D^2CCA loss"). Extensive experiments validate the ability and flexibility of our CFCNet compared to the state-of-the-art methods on both indoor and outdoor scenes with different real-life sparse patterns. Codes are available at: https://github.com/choyingw/CFCNet. Yiqi Zhong, Cho-Ying Wu, Suya You, Ulrich Neumann |
NeurIPS | 4 |
| 2018 | Recurrent Slice Networks for 3D Segmentation of Point CloudsabstractPoint clouds are an efficient data format for 3D data. However, existing 3D segmentation methods for point clouds either do not model local dependencies [21] or require added computations [14, 23]. This work presents a novel 3D segmentation framework, RSNet1, to efficiently model local structures in point clouds. The key component of the RSNet is a lightweight local dependency module. It is a combination of a novel slice pooling layer, Recurrent Neural Network (RNN) layers, and a slice unpooling layer. The slice pooling layer is designed to project features of unordered points onto an ordered sequence of feature vectors so that traditional end-to-end learning algorithms (RNNs) can be applied. The performance of RSNet is validated by comprehensive experiments on the S3DIS[1], ScanNet[3], and ShapeNet [34] datasets. In its simplest form, RSNets surpass all previous state-of-the-art methods on these benchmarks. And comparisons against previous state-of-the-art methods [21, 23] demonstrate the efficiency of RSNets. Qiangui Huang, Weiyue Wang 0002, Ulrich Neumann |
CVPR | 3 |
| 2018 | SGPN: Similarity Group Proposal Network for 3D Point Cloud Instance SegmentationabstractWe introduce Similarity Group Proposal Network (SGPN), a simple and intuitive deep learning framework for 3D object instance segmentation on point clouds. SGPN uses a single network to predict point grouping proposals and a corresponding semantic class for each proposal, from which we can directly extract instance segmentation results. Important to the effectiveness of SGPN is its novel representation of 3D instance segmentation results in the form of a similarity matrix that indicates the similarity between each pair of points in embedded feature space, thus producing an accurate grouping proposal for each point. Experimental results on various 3D scenes show the effectiveness of our method on 3D instance segmentation, and we also evaluate the capability of SGPN to improve 3D object detection and semantic segmentation results. We also demonstrate its flexibility by seamlessly incorporating 2D CNN features into the framework to boost performance. Weiyue Wang 0002, Ronald Yu, Qiangui Huang, Ulrich Neumann |
CVPR | 4 |
| 2018 | Depth-Aware CNN for RGB-D Segmentation
Weiyue Wang 0002, Ulrich Neumann |
ECCV (11) | 2 |
| 2018 | Learning to Prune Filters in Convolutional Neural NetworksabstractMany state-of-the-art computer vision algorithms use large scale convolutional neural networks (CNNs) as basic building blocks. These CNNs are known for their huge number of parameters, high redundancy in weights, and tremendous computing resource consumptions. This paper presents a learning algorithm to simplify and speed up these CNNs. Specifically, we introduce a “try-and-learn” algorithm to train pruning agents that remove unnecessary CNN filters in a data-driven way. With the help of a novel reward function, our agents removes a significant number of filters in CNNs while maintaining performance at a desired level. Moreover, this method provides an easy control of the tradeoff between network performance and its scale. Performance of our algorithm is validated with comprehensive pruning experiments on several popular CNNs for visual recognition and semantic segmentation tasks. Qiangui Huang, Shaohua Kevin Zhou, Suya You, Ulrich Neumann |
WACV | 4 |
| 2017 | Shape Inpainting Using 3D Generative Adversarial Network and Recurrent Convolutional NetworksabstractRecent advances in convolutional neural networks have shown promising results in 3D shape completion. But due to GPU memory limitations, these methods can only produce low-resolution outputs. To inpaint 3D models with semantic plausibility and contextual details, we introduce a hybrid framework that combines a 3D Encoder-Decoder Generative Adversarial Network (3D-ED-GAN) and a Longterm Recurrent Convolutional Network (LRCN). The 3DED- GAN is a 3D convolutional neural network trained with a generative adversarial paradigm to fill missing 3D data in low-resolution. LRCN adopts a recurrent neural network architecture to minimize GPU memory usage and incorporates an Encoder-Decoder pair into a Long Shortterm Memory Network. By handling the 3D model as a sequence of 2D slices, LRCN transforms a coarse 3D shape into a more complete and higher resolution volume. While 3D-ED-GAN captures global contextual structure of the 3D shape, LRCN localizes the fine-grained details. Experimental results on both real-world and synthetic data show reconstructions from corrupted models result in complete and high-resolution 3D objects. Weiyue Wang 0002, Qiangui Huang, Suya You, Ulrich Neumann |
ICCV | 5 |
| 2017 | Self-paced cross-modality transfer learning for efficient road segmentationabstractAccurate road segmentation is a prerequisite for autonomous driving. Current state-of-the-art methods are mostly based on convolutional neural networks (CNNs). Nevertheless, their good performance is at expense of abundant annotated data and high computational cost. In this work, we address these two issues by a self-paced cross-modality transfer learning framework with efficient projection CNN. To be specific, with the help of stereo images, we first tackle a relevant but easier task, i.e. free-space detection with well developed unsupervised methods. Then, we transfer these useful but noisy knowledge in depth modality to single RGB modality with self-paced CNN learning. Finally, we only need to fine-tune the CNN with a few annotated images to get good performance. In addition, we propose an efficient projection CNN, which can improve the fine-grained segmentation results with little additional cost. At last, we test our method on KITTI road benchmark. Our proposed method surpasses all published methods at a speed of 15fps. Weiyue Wang 0002, Naiyan Wang, Suya You, Ulrich Neumann |
ICRA | 5 |
| 2016 | Exemplar-Based 3D Shape Segmentation in Point CloudsabstractThis paper addresses the problem of automatic 3D shape segmentation in point cloud representation. Of particular interest are segmentations of noisy real scans, which is a difficult problem in previous works. To guide segmentation of target shape, a small set of pre-segmented exemplar shapes in the same category is adopted. The main idea is to register the target shape with exemplar shapes in a piece-wise rigid manner, so that pieces under the same rigid transformation are more likely to be in the same segment. To achieve this goal, an over-complete set of candidate transformations is generated in the first stage. Then, each transformation is treated as a label and an assignment is optimized over all points. The transformation labels, together with nearest-neighbor transferred segment labels, constitute final labels of target shapes. The method is not dependent on high-order features, and thus robust to noise as can be shown in the experiments on challenging datasets. Rongqi Qiu, Ulrich Neumann |
3DV | 2 |
| 2016 | 3D point cloud object detection with multi-view convolutional neural networkabstractEfficient detection of three dimensional (3D) objects in point clouds is a challenging problem. Performing 3D descriptor matching or 3D scanning-window search with detector are both time-consuming due to the 3-dimensional complexity. One solution is to project 3D point cloud into 2D images and thus transform the 3D detection problem into 2D space, but projection at multiple viewpoints and rotations produce a large amount of 2D detection tasks, which limit the performance and complexity of the 2D detection algorithm choice. We propose to use convolutional neural network (CNN) for the 2D detection task, because it can handle all viewpoints and rotations for the same class of object together, as well as predicting multiple classes of objects with the same network, without the need for individual detector for each object class. We further improve the detection efficiency by concatenating two extra levels of early rejection networks with binary outputs before the multi-class detection network. Experiments show that our method has competitive overall performance with at least one-order of magnitude speed-up comparing with latest 3D point cloud detection methods. Guan Pang, Ulrich Neumann |
ICPR | 2 |
| 2016 | IPDC: Iterative part-based dense correspondence between point cloudsabstractColor point clouds from 3D scanners are a representation of real-world geometry and color. However, such scan data are imperfect, containing noise, outliers, and occlusions. Noise-free point clouds can be computed by virtually scanning 3D CAD models from a pre-built library, but their geometry may differ from real-world objects. We describe a new algorithm to automatically compute dense correspondences between point cloud scans of same-type objects, thus making it possible to transfer real-world color from noisy scans (source scan) to noise-free virtual scans (target scan), even in cases where the scan objects differ. The method segments both point clouds into parts and then computes part correspondences between them. An iterative algorithm applies a set of rigid transformations to the corresponding parts to determine a dense mapping between them. The dense mapping allows color or other parameter transfers. The resulting point cloud has the geometry of the virtual scan and the color from the real-world scan, mapped in a semantically consistent manner. Rongqi Qiu, Ulrich Neumann |
WACV | 2 |
| 2015 | Fast and Robust Multi-view 3D Object Recognition in Point CloudsabstractRecognition of three dimensional (3D) objects in point clouds is a challenging problem. Existing methods often require prior segmentation or 3D descriptor training and matching, both time consuming and complex processes, especially for large-scale industrial or urban street data. We describe a new recognition approach that projects a 3D point cloud into several 2D depth images from multiple viewpoints, transforming the 3D recognition problem into a series of 2D detection problems. This method reduces complexity, stabilizes performance, and significantly speeds up the recognition process, without any requirement for object segmentation or detector training. Experiments validate the superiority of our method over several state-of-the-art methods on examples from industrial and street data scans. Guan Pang, Ulrich Neumann |
3DV | 2 |
| 2014 | A Unified Framework for Augmented Reality and Knowledge-Based Systems in Maintaining AircraftabstractAircraft maintenance and training play one of the most important roles in ensuring flight safety. The maintenance process usually involves massive numbers of components and substantial procedural knowledge of maintenance procedures. Maintenance tasks require technicians to follow rigorous procedures to prevent operational errors in the maintenance process. In addition, the maintenance time is a cost-sensitive issue for airlines. This paper proposes intelligent augmented reality (IAR) system to minimize operation errors and time-related costs and help aircraft technicians cope with complex tasks by using an intuitive UI/UX interface for their maintenance tasks. The IAR system is composed mainly of three major modules: 1) the AR module 2) the knowledge-based system (KBS) module 3) a unified platform with an integrated UI/UX module between the AR and KBS modules. The AR module addresses vision-based tracking, annotation, and recognition. The KBS module deals with ontology-based resources and context management. Overall testing of the IAR system is conducted at Korea Air Lines (KAL) hangars. Tasks involving the removal and installation of pitch trimmers in landing gear are selected for benchmarking purposes, and according to the results, the proposed IAR system can help technicians to be more effective and accurate in performing their maintenance tasks. Kyeong-Jin Oh, Inay Ha, Kee-Sung Lee, Myung-Duk Hong, Ulrich Neumann, Suya You |
AAAI | 6 |
| 2014 | Pipe-Run Extraction and Reconstruction from Point Clouds
Rongqi Qiu, Qian-Yi Zhou, Ulrich Neumann |
ECCV (3) | 3 |
| 2014 | Visualizing aerial LiDAR cities with hierarchical hybrid point-polygon structures
Zhenzhen Gao, Luciano Nocera, Ulrich Neumann |
Graphics Interface | 4 |
| 2013 | Training-Based Object Recognition in Cluttered 3D Point CloudsabstractRecognition of three dimensional (3D) objects is a challenging problem, especially in cluttered or occluded scenes. Many existing methods focus on a specific type of object or scene, or require prior segmentation. We describe a robust and efficient general purpose 3D object recognition method that combines machine learning procedures with 3D local features, without a requirement for a priori object segmentation. Experiments validate our method on various object types from engineering and street data scans. Guan Pang, Ulrich Neumann |
3DV | 2 |
| 2013 | The Gixel array descriptor (GAD) for multimodal image matchingabstractFeature description and matching is a fundamental problem for many computer vision applications. However, most existing descriptors only work well on images of a single modality with similar texture. This paper presents a novel basic descriptor unit called a Gixel, which uses an additive scoring method to sample surrounding edge information. Several Gixels in a circular array create a powerful descriptor called the Gixel Array Descriptor (GAD), excelling in multi-modal image matching, especially when one of the images is edge-dominant with little texture. Experiments demonstrate the superiority of GAD on multi-modal matching, while maintaining a performance comparable to several state-of-the-art descriptors on single modality matching. Guan Pang, Ulrich Neumann |
WACV | 2 |
| 2013 | Complete residential urban area reconstruction from dense aerial LiDAR point clouds
Qian-Yi Zhou, Ulrich Neumann |
Graph. Model. | 2 |
| 2012 | Modeling Residential Urban Areas from Dense Aerial LiDAR Point Clouds
Qian-Yi Zhou, Ulrich Neumann |
CVM | 2 |
| 2012 | 2.5D building modeling by discovering global regularitiesabstractWe introduce global regularities in the 2.5D building modeling problem, to reflect the orientation and placement similarities between planar elements in building structures. Given a 2.5D point cloud scan, we present an automatic approach that simultaneously detects locally fitted plane primitives and global regularities. While global regularities are extracted by analyzing the plane primitives, they adjust the planes in return and effectively correct local fitting errors. We explore a broad variety of global regularities between 2.5D planar elements including both planer roof patches and planar facade patches. By aligning planar elements to global regularities, our method significantly improves the model quality in terms of both geometry and human judgement. Qian-Yi Zhou, Ulrich Neumann |
CVPR | 2 |
| 2012 | Visually-complete aerial LiDAR point cloud renderingabstractAerial LiDAR (Light Detection and Ranging) point clouds are gathered by a downward scanning laser on a low-flying aircraft. Due to the imaging process, vertical surface features such as building walls, and ground areas under tree canopies are totally or partially occluded, resulting in gaps and sparsely sampled areas. These gaps produce unwanted holes and uneven point distributions that often produce artifacts when visualized using point-based rendering (PBR) techniques. We show how to extend PBR by inferring the physical nature of LiDAR points for visual realism and added comprehension. More specifically, the class of object a point is related to augments the point cloud in pre-processing and/or adapts the online rendering, to produce visualizations that are more complete and realistic. We provide examples of point cloud augmentation for building walls and ground areas under tree canopies. We show how different types of procedurally generated geometry can be used to recover building walls. These methods are generic and can be applied to any aerial LiDAR data set with buildings and trees. Our work also incorporates an out-of-core strategy for hierarchical data management and GPU-accelerated PBR with extended deferred shading. The combined system provides interactive visually-complete rendering of virtually unlimited-size LiDAR point clouds. Experimental results show that our rendering approach adds only a slight overhead to PBR and provides comparable visual cues to visualizations generated by off-line pre-computation of 3D polygonal urban models. Zhenzhen Gao, Luciano Nocera, Ulrich Neumann |
SIGSPATIAL/GIS | 3 |
| 2012 | Fusing oblique imagery with augmented aerial LiDARabstractWe present a scalable out-of-core technique for mapping colors from aerial oblique imagery to large scale aerial LiDAR (Light Detection and Ranging) point cloud. Our method does not require meshing or intensive processing of points, only fast and effective augmentation is applied to fill occluded points on building walls and under tree canopies. The presented system applies a modified visibility pass of GPU splatting to map colors, where occluded points are filtered out by projecting all points as oriented surface splats into images. A weighting scheme is utilized to accumulate colors from all contributing images while leveraging image resolution and surface orientation. The effectiveness of color mapping is demonstrated through visualizations of colored points by a GPU splatting algorithm. Zhenzhen Gao, Luciano Nocera, Ulrich Neumann |
SIGSPATIAL/GIS | 3 |
| 2012 | Efficient matchings in augmented reality applicationabstractWith fast growing popularity of smart phones in recent years, augmented reality (AR) becomes more demanding than ever before. However, one of main challenges is that while features like SIFT or SURF are robust in matchings, they are not computationally efficient. In this paper, we propose an efficient matching method for robust features. A distinctive descriptor is also proposed for performance improvements. Besides, we have developed an outdoor augmented reality system that is based on our proposed methods. The system demonstrates that not only it can achieve robust matchings efficiently, it is also capable to handle large occlusions such as passengers and moving vehicles. Wei Guan 0005, Suya You, Ulrich Neumann |
ICIP | 3 |
| 2012 | Image relighting and matching with illumination informationabstractIn recent years, interest point based feature such as SIFT and SURF are widely used for image matchings. While these features are robust to changes in scales, rotations, and local geometric deformations, they are usually less able to simultaneously handle viewpoint changes and large illumination changes. In this paper, we will cope with this challenging problem. The basic idea is to relight one of the two images so that the image has similar illumination conditions as the other one. After relighting process, the point based features become effective again in the matching process. Our experiments show that the proposed method has good performance for matching two images with very different illuminations. Wei Guan 0005, Suya You, Tanasai Sucontphunt, Ulrich Neumann |
ICIP | 4 |
| 2012 | Efficient matchings and mobile augmented realityabstractWith the fast-growing popularity of smart phones in recent years, augmented reality (AR) on mobile devices is gaining more attention and becomes more demanding than ever before. However, the limited processors in mobile devices are not quite promising for AR applications that require real-time processing speed. The challenge exists due to the fact that, while fast features are usually not robust enough in matchings, robust features like SIFT or SURF are not computationally efficient. There is always a tradeoff between robustness and efficiency and it seems that we have to sacrifice one for the other. While this is true for most existing features, researchers have been working on designing new features with both robustness and efficiency. In this article, we are not trying to present a completely new feature. Instead, we propose an efficient matching method for robust features. An adaptive scoring scheme and a more distinctive descriptor are also proposed for performance improvements. Besides, we have developed an outdoor augmented reality system that is based on our proposed methods. The system demonstrates that not only it can achieve robust matchings efficiently, it is also capable to handle large occlusions such as passengers and moving vehicles, which is another challenge for many AR applications. Wei Guan 0005, Suya You, Ulrich Neumann |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2011 | 2.5D building modeling with topology controlabstract2.5D building reconstruction aims at creating building models composed of complex roofs and vertical walls. In this paper, we define 2.5D building topology as a set of roof features, wall features, and point features; together with the associations between them. Based on this definition, we extend 2.5D dual contouring into a 2.5D modeling method with topology control. Comparing with the previous method, we put less restrictions on the adaptive simplification process. We show results under intense geometry simplifications. Our results preserve significant topology structures while the number of triangles is comparable to that of manually created models or primitive-based models. Qian-Yi Zhou, Ulrich Neumann |
CVPR | 2 |
| 2011 | GPS-aided recognition-based user tracking system with augmented reality in extreme large-scale areasabstractWe present a recognition-based user tracking and augmented reality system that works in extreme large scale areas. The system will provide a user who captures an image of a building facade with precise location of the building and augmented information about the building. While GPS cannot provide information about camera poses, it is needed to aid reducing the searching ranges in image database. A patch-retrieval method is used for efficient computations and real-time camera pose recovery. With the patch matching as the prior information, the whole image matching can be done through propagations in an efficient way so that a more stable camera pose can be generated. Augmented information such as building names and locations are then delivered to the user. The proposed system mainly contains two parts, offline database building and online user tracking. The database is composed of images for different locations of interests. The locations are clustered into groups according to their UTM coordinates. An overlapped clustering method is used to cluster these locations in order to restrict the retrieval range and avoid ping pong effects. For each cluster, a vocabulary tree is built for searching the most similar view. On the tracking part, the rough location of the user is obtained from the GPS and the exact location and camera pose are calculated by querying patches of the captured image. The patch property makes the tracking robust to occlusions and dynamics in the scenes. Moreover, due to the overlapped clusters, the system simulates the "soft handoff" feature and avoid frequent swaps in memory resource. Experiments show that the proposed tracking and augmented reality system is efficient and robust in many cases. Wei Guan 0005, Suya You, Ulrich Neumann |
MMSys | 3 |
| 2011 | Recognition-driven 3D navigation in large-scale virtual environmentsabstractWe present a recognition-driven navigation system for large-scale 3D virtual environments. The proposed system contains three parts, virtual environment reconstruction, feature database building and recognition-based navigation. The virtual environment is reconstructed automatically with LIDAR data and aerial images. The feature database is composed of image patches with features and registered location and orientation information. The database images are taken at different distances from the scenes with various viewing angles, and these images are then partitioned into smaller patches. When a user navigates the real world with a handheld camera, the captured image is used to estimate its location and orientation. These location and orientation information are also reflected in the virtual environment. With the proposed patch approach, the recognition is robust to large occlusions and can be done in real time. Experiments show that our proposed navigation system is efficient and well synchronized with real world navigation. Wei Guan 0005, Suya You, Ulrich Neumann |
VR | 3 |
| 2011 | Computationally efficient retrieval-based tracking system and augmented reality for large-scale areasabstractWe present a retrieval-based tracking system that requires less computational time and cost. The system tracks a user's location through a small portion of an image captured by the camera, and then refines the camera pose by propagating matchings to the whole image. Augmented information such as building names and locations will be delivered to the user. The progressive way to process image data not only can provide the user with location information at real-time speed, but more importantly, it reduces the feature matching time by limiting the searching ranges. The proposed system contains two parts, offline database building and online user tracking. The database is composed of image patches with features and location information. The images are captured at different locations of interests from different viewing angles and distances, and then these images are partitioned into smaller patches. The location of a user can be calculated by querying one or more patches of the captured image. Moreover, the system is capable to handle large occlusions in images due to the patch approach. Experiments show that the proposed tracking system is efficient and robust in many different environments. Wei Guan 0005, Suya You, Ulrich Neumann |
WACV | 3 |
| 2011 | Scan-Based Volume Animation Driven by Locally Adaptive Articulated RegistrationsabstractThis paper describes a complete system to create anatomically accurate example-based volume deformation and animation of articulated body regions, starting from multiple in vivo volume scans of a specific individual. In order to solve the correspondence problem across volume scans, a template volume is registered to each sample. The wide range of pose variations is first approximated by volume blend deformation (VBD), providing proper initialization of the articulated subject in different poses. A novel registration method is presented to efficiently reduce the computation cost while avoiding strong local minima inherent in complex articulated body volume registration. The algorithm highly constrains the degrees of freedom and search space involved in the nonlinear optimization, using hierarchical volume structures and locally constrained deformation based on the biharmonic clamped spline. Our registration step establishes a correspondence across scans, allowing a data-driven deformation approach in the volume domain. The results provide an occlusion-free person-specific 3D human body model, asymptotically accurate inner tissue deformations, and realistic volume animation of articulated movements driven by standard joint control estimated from the actual skeleton. Our approach also addresses the practical issues arising in using scans from living subjects. The robustness of our algorithms is tested by their applications on the hand, probably the most complex articulated region in the body, and the knee, a frequent subject area for medical imaging due to injuries. Taehyun Rhee, John P. Lewis, Ulrich Neumann, Krishna S. Nayak |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2010 | 2.5D Dual Contouring: A Robust Approach to Creating Building Models from Aerial LiDAR Point Clouds
Qian-Yi Zhou, Ulrich Neumann |
ECCV (3) | 2 |
| 2010 | Crafting 3D faces using free form portrait sketching and plausible texture inference
Tanasai Sucontphunt, Borom Tunwattanapong, Zhigang Deng 0001, Ulrich Neumann |
Graphics Interface | 4 |
| 2009 | A Dynamic Programming Approach to Maximizing Tracks for Structure from Motion
Jonathan Mooser, Suya You, Ulrich Neumann, Raphaël Grasset, Mark Billinghurst |
ACCV (2) | 3 |
| 2009 | A robust approach for automatic registration of aerial images with untextured aerial LiDAR dataabstractAirborne LiDAR technology draws increasing interest in large-scale 3D urban modeling in recent years. 3D LiDAR data typically has no texture information. To generate photo-realistic 3D models, oblique aerial images are needed for texture mapping, in which the key step is to obtain accurate registration between aerial images and untextured 3D LiDAR data. We present a robust automatic registration approach. A novel feature called 3CS is proposed which is composed of connected line segments. Putative line segment correspondences are obtained by matching 3CS features detected from both aerial images and 3D LiDAR data. Outliers are removed with a two-level RANSAC algorithm that integrates local and global processing to improve robustness and efficiency. The approach has been tested on 2290 aerial images that cover a variety of urban environments in Oakland and Atlanta areas. Its correct pose recovery rate is over 98%. Lu Wang 0013, Ulrich Neumann |
CVPR | 2 |
| 2009 | A streaming framework for seamless building reconstruction from large-scale aerial LiDAR dataabstractWe present a streaming framework for seamless building reconstruction from huge aerial LiDAR point sets. By storing data as stream files on hard disk and using main memory as only a temporary storage for ongoing computation, we achieve efficient out-of-core data management. This gives us the ability to handle data sets with hundreds of millions of points in a uniform manner. By adapting a building modeling pipeline into our streaming framework, we create the whole urban model of Atlanta from 17.7 GB LiDAR data with 683 M points in under 25 hours using less than 1 GB memory. To integrate this complex modeling pipeline with our streaming framework, we develop a state propagation mechanism, and extend current reconstruction algorithms to handle the large scale of data. Qian-Yi Zhou, Ulrich Neumann |
CVPR | 2 |
| 2009 | Wide-baseline image matching using Line SignaturesabstractWe present a wide-baseline image matching approach based on line segments. Line segments are clustered into local groups according to spatial proximity. Each group is treated as a feature called a Line Signature. Similar to local features, line signatures are robust to occlusion, image clutter, and viewpoint changes. The descriptor and similarity measure of line signatures are presented. Under our framework, the feature matching is not only robust against affine distortion but also a considerable range of 3D viewpoint changes for non-planar surfaces. When compared to matching approaches based on existing local features, our method shows improved results with low-texture scenes. Moreover, extensive experiments validate that our method has advantages in matching structured non-planar scenes under large viewpoint changes and illumination variations. Lu Wang 0013, Ulrich Neumann, Suya You |
ICCV | 2 |
| 2009 | PTZ camera calibration for Augmented Virtual EnvironmentsabstractAugmented Virtual Environments (AVE) are very effective in the application of surveillance, in which multiple video streams are projected onto a 3D urban model for better visualization and comprehension of the dynamic scenes. One of the key issues in creating such systems is to estimate the parameters of each camera including the intrinsic parameters and its pose relative to the 3D model. Nowadays, PTZ cameras are popular in this kind of applications. How to rapidly calibrate them at an arbitrary PTZ setting is not clear in the literature. We propose an efficient approach with two steps. In the first step, panoramic images are generated at a set of zooms. The images composing these panoramas are calibrated and stored in a database. In the second step, an image is acquired at an arbitrary PTZ setting. Its best matching image in the database is found by using an efficient local feature recognition technique. Based on this image, the camera parameters at the new PTZ setting can be estimated. Lu Wang 0013, Suya You, Ulrich Neumann |
ICME | 3 |
| 2009 | Robust pose estimation in untextured environments for augmented reality applicationsabstractWe present a robust camera pose estimation approach for stereo images captured in untextured environments. Unlike most of existing registration algorithms which are point-based and make use of intensities of pixels in the neighborhood, our approach imports line segments in registration process. With line segments as primitives, the proposed algorithm is capable to handle untextured images such as scenes captured in man-made environments, as well as the cases when there are large viewpoint changes or illumination changes. Furthermore, since the proposed algorithm is robust to large base-line stereos, there are improvements on the accuracy of 3D points reconstruction. With well-calculated camera pose and object positions in 3D space, we can embed virtual objects into existing scene with higher accuracy for realistic effects. In our experiments, 2D labels are embedded in the 3D scene space to achieve annotation effects as in AR. Wei Guan 0005, Lu Wang 0013, Jonathan Mooser, Suya You, Ulrich Neumann |
ISMAR | 5 |
| 2009 | Crafting Personalized Facial Avatars Using Editable Portrait and Photograph ExampleabstractComputer-generated facial avatars have been increasingly used in a variety of virtual reality applications. Emulating the real-world face sculpting process, we present an interactive system to intuitively craft personalized 3D facial avatars by using 3D portrait editing and image example-based painting techniques. Starting from a default 3D face portrait, users can conveniently perform intuitive "pulling" operations on its 3D surface to sculpt the 3D face shape towards any individual. To automatically maintain the faceness of the 3D face being crafted, novel facial anthropometry constraints and a reduced face description space are incorporated into the crafting algorithms dynamically. Once the 3D face geometry is crafted, this system can automatically generate a face texture for the crafted model using an image example-based painting algorithm. Our user studies showed that with this system, users are able to craft a personalized 3D facial avatar efficiently on average within one minute. Tanasai Sucontphunt, Zhigang Deng 0001, Ulrich Neumann |
VR | 3 |
| 2009 | Applying robust structure from motion to markerless augmented realityabstractWe demonstrate a complete system for markerless augmented reality using robust structure from motion. The proposed system includes two main components. The first is a means of learning the appearance of complex 3D objects and augmenting them with virtual annotations. Its output is database of recognizable landmarks along with 3D descriptions of accompanying virtual objects. The second component uses this data to recognize the previously learned landmarks, recover camera pose, and render the associated virtual content. Both components make use of the recently developed subtrack optimization algorithm for structure from motion, which we demonstrate to be a useful tool for both learning the structure of objects and tracking camera pose after recognition. The complete system is demonstrated on several complex real-world examples. Jonathan Mooser, Suya You, Ulrich Neumann, Quan Wang 0001 |
WACV | 3 |
| 2008 | Fast and extensible building modeling from airborne LiDAR dataabstractThis paper presents an automatic algorithm which reconstructs building models from airborne LiDAR (light detection and ranging) data of urban areas. While our algorithm inherits the typical building reconstruction pipeline, several major distinct features are developed to enhance efficiency and robustness: 1) we design a novel vegetation detection algorithm based on differential geometry properties and unbalanced SVM; 2) after roof patch segmentation, a fast boundary extraction method is introduced to produce topology-correct water tight boundaries; 3) instead of making assumptions on the angles between roof boundary lines, we propose a data-driven algorithm which automatically learns the principal directions of roof boundaries and uses them in footprint production. Furthermore, we show the extendability of our algorithm by supporting non-flat object patterns with the help of only a few user interactions. We demonstrate the efficiency and accuracy of our algorithm by showing experiment results on urban area data of several different data sets. Qian-Yi Zhou, Ulrich Neumann |
GIS | 2 |
| 2008 | Interactive 3D facial expression posing through 2D portrait manipulation
Tanasai Sucontphunt, Zhenyao Mo, Ulrich Neumann, Zhigang Deng 0001 |
Graphics Interface | 3 |
| 2008 | Supporting range and segment-based hysteresis thresholding in edge detectionabstractOne of the important steps in gradient-based edge detection is thresholding, e.g., hysteresis thresholding in the Canny detector. Traditional approaches only use gradient magnitude as the criterion to select edge pixels. We introduce a novel saliency measure called supporting range. It measures the range along the gradient direction of an edge pixel in which its gradient magnitude is a local maximum. As the result, our approach can detect the edge pixels on object boundaries even if they have very low gradient magnitude. In addition, unlike the Canny detector, the proposed hysteresis thresholding approach is based on segments instead of individual pixels. This also makes the approach more robust to detect edge pixels on weak boundaries. Lu Wang 0013, Suya You, Ulrich Neumann |
ICIP | 3 |
| 2008 | Large document, small screen: a camera driven scroll and zoom control for mobile devicesabstractWe present a three degree-of-freedom control designed for viewing large documents and images on a mobile device equipped with a camera. Tracking natural features detected in the camera's field of view, we can roughly estimate the motion of the device, using the results to scroll and zoom the current document. Central to our implementation is the manner by which we amplify the motion, allowing the user to scroll through large portions of the document with minimal hand movement. Then, using a Hidden Markov Model, we determine when the user is scrolling, zooming, or some combination of the two, thus providing smoother, more fluid control. We demonstrate a prototype of our 3DOF control that can easily navigate documents that are many times larger than the display area, and show how it might be incorporated into a larger document retrieval application. Jonathan Mooser, Suya You, Ulrich Neumann |
SI3D | 3 |
| 2008 | Rapid Creation of Large-scale Photorealistic Virtual EnvironmentsabstractThe rapid and efficient creation of virtual environments has become a crucial part of virtual reality applications. In particular, civil and defense applications often require and employ detailed models of operations areas for training, simulations of different scenarios, planning for natural or man-made events, monitoring, surveillance, games and films. A realistic representation of the large-scale environments is therefore imperative for the success of such applications since it increases the immersive experience of its users and helps reduce the difference between physical and virtual reality. However, the task of creating such large-scale virtual environments still remains a time-consuming and manual work. In this work we propose a novel method for the rapid reconstruction of photorealistic large-scale virtual environments. First, a novel parameterized geometric primitive is presented for the automatic building detection, identification and reconstruction of building structures. In addition, buildings with complex roofs containing non-linear surfaces are reconstructed interactively using a nonlinear primitive. Secondly, we present a rendering pipeline for the composition of photorealistic textures which unlike existing techniques it can recover missing or occluded texture information by integrating multiple information captured from different optical sensors (ground, aerial and satellite). Charalambos Poullis, Suya You, Ulrich Neumann |
VR | 3 |
| 2008 | A Vision-Based System For Automatic Detection and Extraction Of Road NetworksabstractIn this paper we present a novel vision-based system for automatic detection and extraction of complex road networks from various sensor resources such as aerial photographs, satellite images, and LiDAR. Uniquely, the proposed system is an integrated solution that merges the power of perceptual grouping theory (Gabor filtering, tensor voting) and optimized segmentation techniques (global optimization using graph-cuts) into a unified framework to address the challenging problems of geospatial feature detection and classification. Firstly, the local precision of the Gabor filters is combined with the global context of the tensor voting to produce accurate classification of the geospatial features. In addition, the tensorial representation used for the encoding of the data eliminates the need for any thresholds, therefore removing any data dependencies. Secondly, a novel orientation-based segmentation is presented which incorporates the classification of the perceptual grouping, and results in segmentations with better defined boundaries and continuous linear segments. Finally, a set of Gaussian-based filters are applied to automatically extract centerline information (magnitude, width and orientation). This information is then used for creating road segments and then transforming them to their polygonal representations. Charalambos Poullis, Suya You, Ulrich Neumann |
WACV | 3 |
| 2008 | Expressive Speech Animation Synthesis with Phoneme-Level ControlsabstractAbstract This paper presents a novel data‐driven expressive speech animation synthesis system with phoneme‐level controls. This system is based on a pre‐recorded facial motion capture database, where an actress was directed to recite a pre‐designed corpus with four facial expressions (neutral, happiness, anger and sadness). Given new phoneme‐aligned expressive speech and its emotion modifiers as inputs, a constrained dynamic programming algorithm is used to search for best‐matched captured motion clips from the processed facial motion database by minimizing a cost function. Users optionally specify ‘hard constraints’ (motion‐node constraints for expressing phoneme utterances) and ‘soft constraints’ (emotion modifiers) to guide this search process. We also introduce a phoneme–Isomap interface for visualizing and interacting phoneme clusters that are typically composed of thousands of facial motion capture frames. On top of this novel visualization interface, users can conveniently remove contaminated motion subsequences from a large facial motion dataset. Facial animation synthesis experiments and objective comparisons between synthesized facial motion and captured motion showed that this system is effective for producing realistic expressive speech animations. Zhigang Deng 0001, Ulrich Neumann |
Comput. Graph. Forum | 2 |
| 2008 | Reusable skinning templates using cage-based deformationsabstractCharacter skinning determines how the shape of the surface geometry changes as a function of the pose of the underlying skeleton. In this paper we describe skinning templates, which define common deformation behaviors for common joint types. This abstraction allows skinning solutions to be shared and reused, and they allow a user to quickly explore many possible alternatives for the skinning behavior of a character. The skinning templates are implemented using cage-based deformations, which offer a flexible design space within which to develop reusable skinning behaviors. We demonstrate the interactive use of skinning templates to quickly explore alternate skinning behaviors for 3D models. Qian-Yi Zhou, Michiel van de Panne, Daniel Cohen-Or, Ulrich Neumann |
ACM Trans. Graph. | 5 |
| 2007 | Linear feature extraction using perceptual grouping and graph-cutsabstractIn this paper we present a novel system for the detection and extraction of road map information from high-resolution satellite imagery. Charalambos Poullis, Suya You, Ulrich Neumann |
GIS | 3 |
| 2007 | Semiautomatic registration between ground-level panoramas and an orthorectified aerial image for building modelingabstractAerial imagery and ground-level imagery are two complementary data sources for architectural modeling. How to integrate them is a critical issue in creating complete, photo-realistic and large-scale urban models. We describe a semiautomatic approach of detecting feature correspondences between ground-level images and the building footprint in an orthorectified aerial image. The ground-level images are stitched into panoramas in order to obtain a wide camera field of view. Line segments are extracted from ground-level images. Their corresponding segments on the building footprints are automatically detected through a voting process. Meanwhile the camera pose of the ground-level images is also obtained. Wrong correspondences are corrected through user interaction. Later, the height values of the building roof corners are computed and a piece-wise planar 3D model with photo-realistic facade and roof texture is then created. Lu Wang 0013, Suya You, Ulrich Neumann |
ICCV | 3 |
| 2007 | An Augmented Reality Interface for Mobile Information RetrievalabstractRecent years have seen growing interest in mobile augmented reality. The ability to retrieve information and display it as virtual content overlaid on top of an image of the real world is a natural extension to a mobile device equipped with a camera and wireless connectivity. Such applications need to address a number of technical hurdles including target recognition and camera pose estimation. Beyond these fundamental challenges, however, there is the problem of data connectivity and user interface presentation. How do we connect mobile clients to multiple data sources that may be changing in real time and provide the user with a flexible interface for navigating through the relevant content? We present an AR user interface framework specifically designed to expose disparate data sources through a single application server. Our proposed system uses an multi-tier architecture to separate back-end data retrieval from front-end graphical presentation and UI event handling. We describe how it might be used to build an oil platform equipment maintenance system, as an example of a collaborative, data-driven mobile application. Jonathan Mooser, Lu Wang 0013, Suya You, Ulrich Neumann |
ICME | 4 |
| 2007 | Generating High-Resolution Textures for 3D Virtual Environments using View-Independent Texture MappingabstractImage based modeling and rendering techniques have become increasingly popular for creating and visualizing 3D models from a set of images. Typically, these techniques depend on view-dependent texture mapping to render the textured 3D models in which the texture of novel views is synthesized at runtime according to different view-points. This is computationally expensive and limits their application in domains where efficient computations are required, such as games and virtual reality. In this paper we present an offline technique for creating view-independent texture atlases for 3D models, given a set of registered images. The best texture map resolution is computed by considering the areas of the projected polygons in the images. Texture maps are generated by a weighted composition of all available image information in the scene.Assuming that all surfaces of the model are exhibiting Lambertian reflectance properties, ray-tracing is then employed, for creating the view-independent texture maps. Finally, all the generated texture maps are packed into texture atlases. The result is a 3D model with an associated view-independent texture atlas which can be used efficiently in any application without any knowledge of camera pose information. Charalambos Poullis, Suya You, Ulrich Neumann |
ICME | 3 |
| 2007 | A High-Performance Image Matching and Recognition System for Multimedia ApplicationsabstractThis paper presents a high-performance image matching and recognition system for rapid and robust detection, matching and recognition of scene imagery and objects in varied backgrounds. Advanced image processing and pattern recognition technologies provide the system with object distinctiveness, robustness to occlusions, and invariance to scale and geometric distortions. Extensive experiments and several commercial applications have demonstrated the system's value and utility for multimedia applications. Suya You, Ulrich Neumann |
ICME | 2 |
| 2007 | Real-Time Object Tracking for Augmented Reality Combining Graph Cuts and Optical FlowabstractWe present an efficient and accurate object tracking algorithm based on the concept of graph cut segmentation. The ability to track visible objects in real-time provides an invaluable tool for the implementation of markerless Augmented Reality. Once an object has been detected, it's location in future frames can be used to position virtual content, and thus annotate the environment. Unlike many object tracking algorithms, our approach does not rely on a preexisting 3D model or any other information about the object or its environment. It takes, as input, a set of pixels representing an object in an initial frame and uses a combination of optical flow and graph cut segmentation to determine the corresponding pixels in each future frame. Experiments show that our algorithm robustly tracks objects of disparate shapes and sizes over hundreds of frames, and can even handle difficult cases where an object contains many of the same colors as its background. We further show how this technology can be applied to practical AR applications. Jonathan Mooser, Suya You, Ulrich Neumann |
ISMAR | 3 |
| 2007 | Soft-Tissue Deformation for In Vivo Volume AnimationabstractArticulated body animation with smooth skin deformation is an important topic in computer graphics. This paper presents a pipeline that extends articulated body deformation to the volume graphics domain. The pipeline consists of in-vivo volume scans, kinematic joint estimation, volumetric joint weight computation, soft-tissue volume deformation, and direct volume rendering. The result is a fully articulated body volume driven by intuitive joint control that respects rigid deformation of the bone structures and produces smooth deformations of both the skin surface and the interior soft tissue regions. Taehyun Rhee, John P. Lewis, Ulrich Neumann, Krishna S. Nayak |
PG | 3 |
| 2007 | Single View Camera Calibration for Augmented Virtual EnvironmentsabstractAugmented virtual environments (AVE) are very effective in the application of surveillance, in which multiple video streams are projected onto a 3D urban model for better visualization and comprehension of the dynamic scenes. One of the key issues in creating such systems is to estimate the parameters of each camera including the intrinsic parameters and its pose relative to the 3D model. Existing camera pose estimation approaches require known intrinsic parameters and at least three 2D to 3D feature (point or line) correspondences. This cannot always be satisfied in an AVE system. Moreover, due to noise, the estimated camera location may be far from the expectation of the users when the number of correspondences is small. Our approach combines the users' prior knowledge about the camera location and the constraints from the parallel relationship between lines with those from feature correspondences. With at least two feature correspondences, it can always output an estimation of the camera parameters that gives an accurate alignment between the projection of the image (or video) and the 3D model Lu Wang 0013, Suya You, Ulrich Neumann |
VR | 3 |
| 2007 | Rigid Head Motion in Expressive Speech Animation: Analysis and SynthesisabstractRigid head motion is a gesture that conveys important nonverbal information in human communication, and hence it needs to be appropriately modeled and included in realistic facial animations to effectively mimic human behaviors. In this paper, head motion sequences in expressive facial animations are analyzed in terms of their naturalness and emotional salience in perception. Statistical measures are derived from an audiovisual database, comprising synchronized facial gestures and speech, which revealed characteristic patterns in emotional head motion sequences. Head motion patterns with neutral speech significantly differ from head motion patterns with emotional speech in motion activation, range, and velocity. The results show that head motion provides discriminating information about emotional categories. An approach to synthesize emotional head motion sequences driven by prosodic features is presented, expanding upon our previous framework on head motion synthesis. This method naturally models the specific temporal dynamics of emotional head motion sequences by building hidden Markov models for each emotional category (sadness, happiness, anger, and neutral state). Human raters were asked to assess the naturalness and the emotional content of the facial animations. On average, the synthesized head motion sequences were perceived even more natural than the original head motion sequences. The results also show that head motion modifies the emotional perception of the facial animation especially in the valence and activation domain. These results suggest that appropriate head motion not only significantly improves the naturalness of the animation but can also be used to enhance the emotional content of the animation to effectively engage the users Carlos Busso, Zhigang Deng 0001, Michael Grimm, Ulrich Neumann, Shri Narayanan |
IEEE Trans. Speech Audio Process. | 4 |
| 2006 | Real-time Hand Pose Recognition Using Low-Resolution Depth ImagesabstractGesture recognition methods based on intensity or color images often suffer from low efficiency and lack of robustness. In this paper, we employ a new laser-based camera that produces reliable low-resolution depth images at video rates. By decomposing and recognizing hand poses as finger states (finger poses and finger inter-relations), we achieve robust hand pose recognition in real-time (30 frames/second). Zhenyao Mo, Ulrich Neumann |
CVPR (2) | 2 |
| 2006 | Tricodes: A Barcode-Like Fiducial Design for Augmented Reality MediaabstractVisual markers, or fiducials, have become one of the most common methods of camera pose estimation in augmented reality (AR) media. Many present day fiducial-based AR systems use arbitrary patterns, such as simple line drawings or alpha-numeric characters, and require that an application be "trained" to recognize its pattern set. These techniques work well on a small scale, but as the number of fiducials grows, accuracy and performance degrade. We describe a new fiducial design called TriCodes that, like a barcode, provides a systematic way of printing and identifying a vast library of patterns. We compare TriCodes to the popular ARToolkit package, demonstrating its advantages in the presence of large numbers of fiducials Jonathan Mooser, Suya You, Ulrich Neumann |
ICME | 3 |
| 2006 | Geodec: Enabling Geospatial Decision MakingabstractThe rapid increase in the availability of geospatial data has motivated the effort to seamlessly integrate this information into an information-rich and realistic 3D environment. However, heterogeneous data sources with varying degrees of consistency and accuracy pose a challenge to such efforts. We describe the geospatial decision making (GeoDec) system, which accurately integrates satellite imagery, three-dimensional models, textures and video streams, road data, maps, point data and temporal data. The system also includes a glove-based user interface Cyrus Shahabi, Yao-Yi Chiang, Kelvin Chung, Kai-Chen Huang, Ali Khoshgozaran, Craig A. Knoblock, Sung Lee, Ulrich Neumann, Ramakant Nevatia, Arjun Rihan, Snehal Thakkar, Suya You |
ICME | 8 |
| 2006 | Lexical Gesture InterfaceabstractGesture interfaces have long been pursued in the context of portable computing and immersive environments. However, such interfaces have been difficult to realize, in part due to a lack of frameworks for their design and implementation. This paper presents a framework for automatically producing a gesture interface based on a simple interface description. Rather than defining hand positions in a low-level high-dimensional joint angle space, we describe and recognize gestures in a "lexical" space, in which each hand pose is decomposed into elements in a finger-pose alphabet. The alphabet and underlying rules are defined as a gesture notation system called GeLex. By implementing a generic hand pose recognition algorithm, and a mechanism to adapt it to a specific application based on a general interface description, developing a gesture interface becomes straightforward. Zhenyao Mo, Ulrich Neumann |
ICVS | 2 |
| 2006 | Perceiving Visual Emotions with Speech
Zhigang Deng 0001, Jeremy N. Bailenson, John P. Lewis, Ulrich Neumann |
IVA | 4 |
| 2006 | Animating blendshape faces by cross-mapping motion capture dataabstractAnimating 3D faces to achieve compelling realism is a challenging task in the entertainment industry. Previously proposed face transfer approaches generally require a high-quality animated source face in order to transfer its motion to new 3D faces. In this work, we present a semi-automatic technique to directly animate popularized 3D blendshape face models by mapping facial motion capture data spaces to 3D blendshape face spaces. After sparse markers on the face of a human subject are captured by motion capture systems while a video camera is simultaneously used to record his/her front face, then we carefully select a few motion capture frames and accompanying video frames as reference mocap-video pairs. Users manually tune blendshape weights to perceptually match the animated blendshape face models with reference facial images (the reference mocap-video pairs) in order to create reference mocap-weight pairs. Finally, the Radial Basis Function (RBF) regression technique is used to map any new facial motion capture frame to blendshape weights based on the reference mocap-weight pairs. Our results demonstrate that this technique is efficient to animate blendshape face models, while offering its generality and flexiblity. Zhigang Deng 0001, Pei-Ying Chiang, Pamela Fox, Ulrich Neumann |
SI3D | 4 |
| 2006 | Human hand modeling from surface anatomyabstractThe human hand is an important interface with complex shape and movement. In virtual reality and gaming applications the use of an individualized rather than generic hand representation can increase the sense of immersion and in some cases may lead to more effortless and accurate interaction with the virtual world. We present a method for constructing a person-specific model from a single canonically posed palm image of the hand without human guidance. Tensor voting is employed to extract the principal creases on the palmar surface. Joint locations are estimated using extracted features and analysis of surface anatomy. The skin geometry of a generic 3D hand model is deformed using radial basis functions guided by correspondences to the extracted surface anatomy and hand contours. The result is a 3D model of an individual's hand, with similar joint locations, contours, and skin texture. Taehyun Rhee, Ulrich Neumann, John P. Lewis |
SI3D | 2 |
| 2006 | Real-Time Weighted Pose-Space Deformation on the GPUabstractAbstract WPSD (Weighted Pose Space Deformation) is an example based skinning method for articulated body animation. The per‐vertex computation required in WPSD can be parallelized in a SIMD (Single Instruction Multiple Data) manner and implemented on a GPU. While such vertex‐parallel computation is often done on the GPU vertex processors, further parallelism can potentially be obtained by using the fragment processors. In this paper, we develop a parallel deformation method using the GPU fragment processors. Joint weights for each vertex are automatically calculated from sample poses, thereby reducing manual effort and enhancing the quality of WPSD as well as SSD (Skeletal Subspace Deformation). We show sufficient speed‐up of SSD, PSD (Pose Space Deformation) and WPSD to make them suitable for real‐time applications. Categories and Subject Descriptors (according to ACM CCS): I.3.1 [Computer Graphics]: Hardware Architecture‐Parallel processing, I.3.5 [Computer Graphics]: Computational Geometry and Object Modeling‐Curve, surface, solid and object modeling, I.3.7 [Computer Graphics]: Three‐Dimensional Graphics and Realism‐Animation. Taehyun Rhee, John P. Lewis, Ulrich Neumann |
Comput. Graph. Forum | 3 |
| 2006 | Expressive Facial Animation Synthesis by Learning Speech Coarticulation and Expression SpacesabstractSynthesizing expressive facial animation is a very challenging topic within the graphics community. In this paper, we present an expressive facial animation synthesis system enabled by automated learning from facial motion capture data. Accurate 3D motions of the markers on the face of a human subject are captured while he/she recites a predesigned corpus, with specific spoken and visual expressions. We present a novel motion capture mining technique that "learns" speech coarticulation models for diphones and triphones from the recorded data. A Phoneme-Independent Expression Eigenspace (PIEES) that encloses the dynamic expression signals is constructed by motion signal processing (phoneme-based time-warping and subtraction) and Principal Component Analysis (PCA) reduction. New expressive facial animations are synthesized as follows: First, the learned coarticulation models are concatenated to synthesize neutral visual speech according to novel speech input, then a texture-synthesis-based approach is used to generate a novel dynamic expression signal from the PIEES model, and finally the synthesized expression signal is blended with the synthesized neutral visual speech to create the final expressive facial animation. Our experiments demonstrate that the system can effectively synthesize realistic expressive facial animation. Zhigang Deng 0001, Ulrich Neumann, John P. Lewis, Tae-Yong Kim 0002, Murtaza Bulut, Shri Narayanan |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2005 | Synthesizing speech animation by learning compact speech co-articulation modelsabstractWhile speech animation fundamentally consists of a sequence of phonemes over time, sophisticated animation requires smooth interpolation and co-articulation effects, where the preceding and following phonemes influence the shape of a phoneme. Co-articulation has been approached in speech animation research in several ways, most often by simply smoothing the mouth geometry motion over time. Data-driven approaches tend to generate realistic speech animation, but they need to store a large facial motion database, which is not feasible for real time gaming and interactive applications on platforms such as PDAs and cell phones. In this paper we show that accurate speech co-articulation model with compact size can be learned from facial motion capture data. An initial phoneme sequence is generated automatically from text-to-speech (TTS) systems. Then, our learned co-articulation model is applied to the resulting phoneme sequence, producing natural and detailed motion. The contribution of this work is that speech co-articulation models "learned" from real human motion data can be used to generate natural-looking speech motion while simultaneously preserving the expressiveness of the animation via keyframing control. Simultaneously, this approach can be effectively applied to interactive applications due to its compact size. Zhigang Deng 0001, John P. Lewis, Ulrich Neumann |
Computer Graphics International | 3 |
| 2005 | SmartCanvas: a gesture-driven intelligent drawing desk systemabstractThis paper describes SmartCanvas, an intelligent desk system that allows a user to perform freehand drawing on a desk or similar surface with gestures. Our system requires one camera and no touch sensors. The key underlying technique is a vision-based method that distinguishes drawing gestures and transitional gestures in real time, avoiding the need for "artificial" gestures to mark the beginning and end of a drawing stroke. The method achieves an average classification accuracy of 92.17%. Pie-shaped menus and a "rotate-to-and-select" approach eliminate the need for a fixed menu display, resulting in an "invisible" interface. Zhenyao Mo, John P. Lewis, Ulrich Neumann |
IUI | 3 |
| 2005 | Reducing blendshape interference by selected motion attenuationabstractBlendshapes (linear shape interpolation models) are perhaps the most commonly employed technique in facial animation practice. A major problem in creating blendshape animation is that of blendshape interference: the adjustment of a single blendshape "slider" may degrade the effects obtained with previous slider movements, because the blendshapes have overlapping, non-orthogonal effects. Because models used in commercial practice may have 100 or more individual blendshapes, the interference problem is the subject of considerable manual effort. Modelers iteratively resculpt models to reduce interference where possible, and animators must compensate for those interference effects that remain. In this short paper we consider the blendshape interference problem from a linear algebra point of view. We find that while full orthogonality is not desirable, the goal of preserving previous adjustments to the model can be effectively approached by allowing the user to temporarily designate a set of points as representative of the previous (desired) adjustments. We then simply solve for blendshape slider values that mimic desired new movement while moving these "tagged" points as little as possible. The resulting algorithm is easy to implement and demonstrably reduces cases of blendshape interference found in existing models. John P. Lewis, Jonathan Mooser, Zhigang Deng 0001, Ulrich Neumann |
SI3D | 4 |
| 2005 | Rapid part-based 3D modelingabstractAn intuitive and easy-to-use 3D modeling system has become more crucial with the rapid growth of computer graphics in our daily lives. Image-based modeling (IBM) has been a popular alternative to pure 3D modelers (e.g. 3D Studio Max, Maya) since its introduction in the late 1990s. However, IBM techniques are inherently very slow and rarely user friendly. Most IBM techniques require either very extensive manual input and/or multiple images. In this paper, we present an IBM technique that gives high level of detail with 1-2 minutes of manipulation from a novice user using only single, un-calibrated image. Our system modifies a generic part-based model of the object under investigation. User inputs are entered via a simple interface and converted into modifications to the whole 3D model. We demonstrate the effectiveness of our modeler by modeling several vehicles, such as SUVs, sedan/hatchback/coupe cars, minivans, trucks and more. Ismail Oner Sebe, Suya You, Ulrich Neumann |
VRST | 3 |
| 2005 | Natural head motion synthesis driven by acoustic prosodic featuresabstractAbstract Natural head motion is important to realistic facial animation and engaging human–computer interactions. In this paper, we present a novel data‐driven approach to synthesize appropriate head motion by sampling from trained hidden markov models (HMMs). First, while an actress recited a corpus specifically designed to elicit various emotions, her 3D head motion was captured and further processed to construct a head motion database that included synchronized speech information. Then, an HMM for each discrete head motion representation (derived directly from data using vector quantization) was created by using acoustic prosodic features derived from speech. Finally, first‐order Markov models and interpolation techniques were used to smooth the synthesized sequence. Our comparison experiments and novel synthesis results show that synthesized head motions follow the temporal dynamic behavior of real human subjects. Copyright © 2005 John Wiley & Sons, Ltd. Carlos Busso, Zhigang Deng 0001, Ulrich Neumann, Shri Narayanan |
Comput. Animat. Virtual Worlds | 3 |
| 2004 | Face Inpainting with Local Linear RepresentationsabstractA number of investigators have had success using domain specific prior knowledge to produce improved superresolution images of faces ("hallucinating faces"). These efforts address the scenario where a face image is obtained from a low-resolution camera. A related but less studied problem occurs when the missing information is the result of occlusion rather than low camera resolution, as in the case when a person is wearing sunglasses. Recently Hwang and Lee [14] introduced the first algorithm for solving this reconstruction "inpainting" problem. In the current work we report results of a psychological study that provides independent evidence regarding the validity of the face reconstruction task, and we demonstrate an improved reconstruction approach using a positive, local linear representation. The positive, local mixture operates on real-world images without manual intervention in many cases, and provides demonstrably lower reconstruction error than is obtainable with a global representation. Zhenyao Mo, John P. Lewis, Ulrich Neumann |
BMVC | 3 |
| 2004 | Ripple-free local bases by designabstractIn some applications, a local or "parts based" representation is preferable to global basis functions, such as those used in Fourier and principal component analysis. In applications that require human understanding and editing of the data, it is also desirable that the basis functions be in some sense as "simple" as possible. This means, for example, that the basis functions should not have Gabor-like ripples if such ripples are not a prominent feature of the data to be represented. The paper introduces a direct local basis construction. Specifically, we show that local bases result from maximizing an appropriate redefinition of pairwise orthogonality while maintaining the ability to represent the data. The resulting basis functions are competitive with (and for some applications superior to) those obtained from existing algorithms, and the construction does not require that the basis coefficients be statistically non-Gaussian or independent, as would be the case with an independent component analysis approach. John P. Lewis, Zhenyao Mo, Ulrich Neumann |
ICASSP (3) | 3 |
| 2004 | Analysis of emotion recognition using facial expressions, speech and multimodal informationabstractThe interaction between human beings and computers will be more natural if computers are able to perceive and respond to human non-verbal communication such as emotions. Although several approaches have been proposed to recognize human emotions based on facial expressions or speech, relatively limited work has been done to fuse these two, and other, modalities to improve the accuracy and robustness of the emotion recognition system. This paper analyzes the strengths and the limitations of systems based only on facial expressions or acoustic information. It also discusses two approaches used to fuse these two modalities: decision level and feature level integration. Using a database recorded from an actress, four emotions were classified: sadness, anger, happiness, and neutral state. By the use of markers on her face, detailed facial motions were captured with motion capture, in conjunction with simultaneous speech recordings. The results reveal that the system based on facial expression gave better performance than the system based on just acoustic information for the emotions considered. Results also show the complementarily of the two modalities and that when these two modalities are fused, the performance and the robustness of the emotion recognition system improve measurably. Carlos Busso, Zhigang Deng 0001, Serdar Yildirim, Murtaza Bulut, Chul Min Lee, Abe Kazemzadeh, Sungbok Lee, Ulrich Neumann, Shri Narayanan |
ICMI | 8 |
| 2004 | A Robust Hybrid Tracking System for Outdoor Augmented Reality
Bolan Jiang, Ulrich Neumann, Suya You |
VR | 2 |
| 2004 | Colorplate: A Robust Hybrid Tracking System for Outdoor Augmented Reality
Bolan Jiang, Ulrich Neumann, Suya You |
VR | 2 |
| 2004 | Analysis of co-articulation regions for performance-driven facial animationabstractAbstract A facial gesture analysis procedure is presented for the control of animated faces. Facial images are partitioned into a set of local, independently actuated regions of appearance change termed co‐articulation regions (CRs). Each CR is parameterized by the activation level of a set of face gestures that affect the region. The activation of a CR is analyzed using independent component analysis (ICA) on a set of training images acquired from an actor. Gesture intensity classification is performed in ICA space by correlation to training samples. Correlation in ICA space proves to be an efficient and stable method for gesture intensity classification with limited training data. A discrete sample‐based synthesis method is also presented. An artist creates an actor‐independent reconstruction sample database that is indexed with CR state information analyzed in real time from video. Copyright © 2004 John Wiley & Sons, Ltd. Douglas Fidaleo, Ulrich Neumann |
Comput. Animat. Virtual Worlds | 2 |
| 2004 | VisualIDs: automatic distinctive icons for desktop interfacesabstractAlthough existing GUIs have a sense of space, they provide no sense of place. Numerous studies report that users misplace files and have trouble wayfinding in virtual worlds despite the fact that people have remarkable visual and spatial abilities. This issue is considered in the human-computer interface field and has been addressed with alternate display/navigation schemes. Our paper presents a fundamentally graphics based approach to this 'lost in hyperspace' problem. Specifically, we propose that spatial display of files is not sufficient to engage our visual skills; scenery (distinctive visual appearance) is needed as well. While scenery (in the form of custom icon assignments) is already possible in current operating systems, few if any users take the time to manually assign icons to all their files. As such, our proposal is to generate visually distinctive icons ("VisualIDs") automatically , while allowing the user to replace the icon if desired. The paper discusses psychological and conceptual issues relating to icons, visual memory, and the necessary relation of scenery to data. A particular icon generation algorithm is described; subjects using these icons in simulated file search and recall tasks show significantly improved performance with little effort. Although the incorporation of scenery in a graphical user interface will introduce many new (and interesting) design problems that cannot be addressed in this paper, we show that automatically created scenery is both beneficial and feasible. John P. Lewis, Ruth Rosenholtz, Nickson Fong, Ulrich Neumann |
ACM Trans. Graph. | 4 |
| 2003 | Urban Site Modeling from LiDAR
Suya You, Ulrich Neumann, Pamela Fox |
ICCSA (3) | 3 |
| 2003 | Practical eye movement model using texture synthesisabstractAs humans we are especially sensitive to the appearance of the face, and on the face, the eyes are particularly important. In fact, in attempts to animate photo-realistic CG humans, the eyes are very often what destroys the illusion [Williams 2003]. The state of art in eye movement synthesis is the Eyes Alive model [Lee et al. 2002] that develops a custom statistical model specifically for eye movement. While its results are the best to date, the model is complex and one wonders if it could be improved by using additional or different statistics. In fact the problem of generating novel animation that captures the “character” of given training data is the same problem as texture synthesis. In this sketch we describe a practical eye movement model using non-parametric texture synthesis techniques ([Efros and Leung 1999]), simulating the eye gaze motion and eye blink motion simultaneously. This approach uses the data directly and without an intervening humancrafted statistical model, yet it produces results that appear as good or better than the more complex statistical model. 2 Approach Zhigang Deng 0001, John P. Lewis, Ulrich Neumann |
SIGGRAPH | 3 |
| 2003 | Video-based virtual environmentsabstractNo abstract available. Thomas Pintaric, Albert A. Rizzo, Ulrich Neumann |
SIGGRAPH | 3 |
| 2003 | 360 Degree Panoramic HMD ImmersionabstractPanoramic video image acquisition is based on multiple overlapped sub-images. We will demonstrate high-resolution panoramic video by employing an array of five video cameras viewing the scene over a combined 360-degrees of horizontal arc and 50-degrees vertical. The five live video streams are digitized and processed in real time by a computer system. The camera lens distortions and colorimetric variations are corrected by the software application and a complete panoramic image is constructed in memory. Users can navigate the scene by wearing a head mounted display (HMD). A single window with a resolution of 800x600 is output to the HMD. A real-time (inertial-magnetic) orientation tracker is fixed to the HMD to sense the user’s head orientation. The orientation is reported to the viewing application through an IP socket, and the output display window is positioned (to mimic pan and tilt) within the full panoramic image in response to the user’s head orientation. Kambiz Ghahremani, Albert A. Rizzo, Ulrich Neumann |
VR | 3 |
| 2003 | Augmented Virtual Environments (AVE): Dynamic Fusion of Imagery and 3D ModelsabstractAn augmented virtual environment (AVE) fuses dynamic imagery with 3D models. The AVE provides a unique approach to visualize and comprehend multiple streams of temporal data or images. Models are used as a 3D substrate for the visualization of temporal imagery, providing improved comprehension of scene activities. The core elements of AVE systems include model construction, sensor tracking, real-time video/image acquisition, and dynamic texture projection for 3D visualization. This paper focuses on the integration of these components and the results that illustrate the utility and benefits of the resulting augmented virtual environment. Ulrich Neumann, Suya You, Bolan Jiang, Jong Weon Lee 0001 |
VR | 1 |
| 2002 | CoArt: Co-articulation Region Analysis for Control of 2D CharactersabstractA facial analysis-synthesis framework based on a concise set of local, independently actuated, coarticulation regions (CRs) is presented for the control of 2D animated characters. CRs are parameterized by muscle actuations and thereby provide a physically meaningful description of face state that is easily abstracted to higher-level descriptions of facial expression. An independent component analysis on a set of training images acquired from an actor is used to characterize the appearance space of each CR. Within this framework actor-independent face reconstruction databases can be created by an artist or extracted from video sequences. In addition, the muscle parameter values may be used to drive any similarly parameterized 3D facial model. The flexibility afforded by such a methodology is demonstrated with applications to 2D facial animation control and sample based video synthesis. The analysis runs in real-time on modest consumer hardware. Douglas Fidaleo, Ulrich Neumann |
CA | 2 |
| 2002 | Tracking with Omni-Directional Vision for Outdoor AR SystemsabstractMost pose (3D position and 3D orientation) tracking methods using vision require a priori knowledge about the environment and correspondences between 3D environment features and 2D images. This environmental information is difficult to acquire accurately for large working volumes or may not be available at all, especially for outdoor environments. As a result, most pose tracking methods using vision are designed for small indoor working spaces. We track the pose of a moving camera from 2D images of the world. The pose of a camera is tracked through two 5 degree-of-freedom (DOF) motion estimations, which requires only 2D-to-2D correspondences. Therefore, the presented method can be applied to varied working space sizes including outdoor environments. Jong Weon Lee 0001, Suya You, Ulrich Neumann |
ISMAR | 3 |
| 2001 | A layered approach to deformable modeling and animationabstractOur approach integrates three mechanisms needed to model and animate deformable objects: controlling mechanisms, geometric surface deformation, and mesh refinement. Most approaches focus either modeling the physically correct behavior or alternative representations for deformable models. This results in a set of method specific algorithms that best represents a particular class of deformable objects. By encapsulating each process, our system introduces the interface that allows one to integrate existing controllers in a modular fashion. We demonstrate this in our system by instantiating a hardware accelerated free form deformation for geometric deformation and midpoint subdivision for mesh refinement. Finally, we discuss the available options for the controlling mechanisms and show how this approach leads to a generalizable framework. Clint Chua, Ulrich Neumann |
CA | 2 |
| 2001 | Expression cloningabstractWe present a novel approach to producing facial expression animations for new models. Instead of creating new facial animations from scratch for each new model created, we take advantage of existing animation data in the form of vertex motion vectors. Our method allows animations created by any tools or methods to be easily retargeted to new models. We call this process expression cloning and it provides a new alternative for creating facial animations for character models. Expression cloning makes it meaningful to compile a high-quality facial animation library since this data can be reused for new models. Our method transfers vertex motion vectors from a source face model to a target model having different geometric proportions and mesh structure (vertex number and connectivity). With the aid of an automated heuristic correspondence search, expression cloning typically requires a user to select fewer than ten points in the model. Cloned expression animations preserve the relative motions, dynamics, and character of the original facial animations. Jun-yong Noh, Ulrich Neumann |
SIGGRAPH | 2 |
| 2001 | Fusion of Vision and Gyro Tracking for Robust Augmented Reality RegistrationabstractA novel framework enables accurate augmented reality (AR) registration with integrated inertial gyroscope and vision tracking technologies. The framework includes a two-channel complementary motion filter that combines the low-frequency stability of vision sensors with the high-frequency tracking of gyroscope sensors, hence achieving stable static and dynamic six-degree-of-freedom pose tracking. Our implementation uses an extended Kalman filter (EKF). Quantitative analysis and experimental results show that the fusion method achieves dramatic improvements in tracking stability and robustness over either sensor alone. We also demonstrate a new fiducial design and detection system in our example AR annotation systems that illustrate the behavior and benefits of the new tracking method. Suya You, Ulrich Neumann |
VR | 2 |
| 2000 | A Thin Shell Volume for Modeling Human HairabstractHair-to-hair interaction is often ignored in human hair modeling, due to its computational and algorithmic complexity. In this paper, we present our experimental approach to simulate the complex behavior of long human hair, taking into account the hair-to-hair interactions. For long hair, we propose the thin shell volume (TSV) model for enhancing hair realism by simulating complex hair-hair interaction. The TSV is a thin bounding volume that encloses a given hair surface. The TSV method enables virtual hair combing for surface-model hairstyles. Combing produces the effects of hair-hair interaction that occur in real human hair. Any hair model based on a surface representation can benefit from this approach. The TSV is presented mainly as a tool to modeling human hair, but this model also gives rise to global hair animation control. Tae-Yong Kim 0002, Ulrich Neumann |
CA | 2 |
| 2000 | Motion Estimation with Incomplete Information Using Omni-Directional VisionabstractWe present a new motion estimation framework and apply it to omni-directional imagery. Our method estimates motions incrementally using an implicit extended Kalman filter (IEKF). Each individual feature provides partial information about the camera motion. The motion estimate is incrementally improved as each feature is processed, similar to the SCAAT approach. The SCAAT method was developed for calibrated features and full 6DOF-pose tracking whereas our method estimates 5DOF translation and rotation motions concurrently from uncalibrated features based on the rigidity and the depth independent constraints. The main difference of our method from others is the combination of a recursive estimation framework in an IEKF and the constraints used in motion estimation. Jong Weon Lee 0001, Ulrich Neumann |
ICIP | 2 |
| 2000 | A video-based augmented reality golf simulator
Alok Govil, Suya You, Ulrich Neumann |
ACM Multimedia | 3 |
| 2000 | Immersive panoramic video
Ulrich Neumann, Thomas Pintaric, Albert A. Rizzo |
ACM Multimedia | 1 |
| 2000 | Virtual Environment Applications in Clinical NeuropsychologyabstractVirtual environment (VE) technology is increasingly being recognized as a useful medium for the study, assessment, and rehabilitation of cognitive processes and functional abilities. The capacity of VE technology to create dynamic three-dimensional (3D) stimulus environments, within which all behavioral responding can be recorded, offers clinical assessment and rehabilitation options that are not available using traditional neuropsychological methods. This work has the potential to advance the scientific study of normal cognitive and behavioral processes and to improve our capacity to understand, measure, and treat the impairments typically found in clinical populations with central nervous system (CNS) dysfunction. The paper provides a rationale for the application of VE technology in the areas of neuropsychological assessment and cognitive rehabilitation, presents a tabled summary of the VE literature targeting cognitive/functional processes in clinical CNS populations and briefly describes two of our VE applications targeting attention and visuospatial processing. Albert A. Rizzo, J. Galen Buckwalter, Cheryl van der Zaag, Ulrich Neumann, Marcus Thiébaux, Clint Chua, Andre van Rooyen, L. Humphrey, Peter Larson |
VR | 4 |
| 2000 | Animated deformations with radial basis functionsabstractWe present a novel approach to creating deformations of polygonal models using Radial Basis Functions (RBFs) to produce localized real-time deformations. Radial Basis Functions assume surface smoothness as a minimal constraint and animations produce smooth displacements of affected vertices in a model. Animations are produced by controlling an arbitrary sparse set of control points defined on or near the surface of the model. The ability to directly manipulate a facial surface with a small number of point motions facilitates an intuitive method for creating facial expressions for virtual environment applications such as an immersive teleconferencing system or entertainment. Smooth deformations of the human face or other models are possible and illustrated with examples of a variety of expressions and mouth shapes. Jun-yong Noh, Douglas Fidaleo, Ulrich Neumann |
VRST | 3 |
| 2000 | Web-Based Remote Rendering with IBRAC (Image-Based Rendering Acceleration and Compression)abstractRecent advances in Internet and computer graphics stimulate intensive use and development of 3D graphics on the World Wide Web. To increase efficiency of systems using 3D graphics on the web, the presented method utilizes previously rendered and transmitted images to accelerate the rendering and compression of new synthetic scene images. The algorithm employs ray casting and epipolar constraints to exploit spatial and temporal coherence between the current and previously rendered images. The reprojection of color and visibility data accelerates the computation of new images. The rendering method intrinsically computes a residual image, based on a user specified error tolerance that balances image quality against computation time and bandwidth. Encoding and decoding uses the same algorithm, so the transmitted residual image consists only of significant data without addresses or offsets. We measure rendering speed‐ups of four to seven without visible degradation. Compression ratios per frame are a factor of two to ten better than MPEG2 in our test cases. There is no transmission of 3D scene data to delay the first image. The efficiency of the server and client generally increases with scene complexity or data size since the rendering time is predominantly a function of image size. This approach is attractive for remote rendering applications such as web‐based scientific visualization where a client system may be a relatively low‐performance machine and limited network bandwidth makes transmission of large 3D data impractical. Ilmi Yoon, Ulrich Neumann |
Comput. Graph. Forum | 2 |
| 1999 | Hybrid Inertial and Vision Tracking for Augmented Reality RegistrationabstractThe biggest single obstacle to building effective augmented reality (AR) systems is the lack of accurate wide-area sensors for trackers that report the locations and orientations of objects in an environment. Active (sensor-emitter) tracking technologies require powered-device installation. Limiting their use to prepared areas that are relatively free of natural or man-made interference sources. Vision-based systems can use passive landmarks, but they are more computationally demanding and often exhibit erroneous behavior due to occlusion or numerical instability. Inertial sensors are completely passive, requiring no external devices or targets, however, the drift rates in portable strapdown configurations are too great for practical use. In this paper, we present a hybrid approach to AR tracking that integrates inertial and vision-based technologies. We exploit the complementary nature of the two technologies to compensate for the weaknesses in each component. Analysis and experimental results demonstrate this system's effectiveness. Suya You, Ulrich Neumann, Ronald T. Azuma |
VR | 2 |
| 1999 | Tracking in unprepared environments for augmented reality systems
Ronald T. Azuma, Jong Weon Lee 0001, Bolan Jiang, Jun Park, Suya You, Ulrich Neumann |
Comput. Graph. | 6 |
| 1999 | Natural Feature Tracking for Augmented RealityabstractNatural scene features stabilize and extend the tracking range of augmented reality (AR) pose-tracking systems. We develop robust computer vision methods to detect and track natural features in video images. Point and region features are automatically and adaptively selected for properties that lead to robust tracking. A multistage tracking algorithm produces accurate motion estimates, and the entire system operates in a closed-loop that stabilizes its performance and accuracy. We present demonstrations of the benefits of using tracked natural features for AR applications that illustrate direct scene annotation, pose stabilization, and extendible tracking range. Our system represents a step toward integrating vision with graphics to produce robust wide-area augmented realities. Ulrich Neumann, Suya You |
IEEE Trans. Multim. | 1 |
| 1998 | Integration of Region Tracking and Optical Flow for Image Motion Estimation
Ulrich Neumann, Suya You |
ICIP (3) | 1 |
| 1998 | Special Issue on High Fidelity Media Processing: Guest Editors' Comments
Tomlinson Holman, C.-C. Jay Kuo, Chris Kyriakakis, Ulrich Neumann, Antonio Ortega |
J. Vis. Commun. Image Represent. | 4 |
| 1997 | Fast color fiducial detection and dynamic workspace extension in video see-through self-tracking augmented realityabstractThe registration problem is one of the major issues in augmented reality (AR). Fiducial tracking is gaining interest as a solution to this problem in video see-through AR because of the availability of digitized real scenes. There are several AR systems using fiducial tracking, but most of them operate in small desktop workspaces. It is difficult to apply them directly to large scale applications. The wide range of work distance and non-uniform lighting conditions make fiducial detection very difficult. Adding new fiducials requires off-line processing for measuring positions of new fiducials. We propose a fast and robust fiducial detection procedure with carefully designed color fiducials and noise analysis of digitized images. We also present a dynamic workspace extension method with on-line position determination of unknown features. We present a framework for applying AR to large scale applications. Youngkwan Cho, Jun Park, Ulrich Neumann |
PG | 3 |
| 1996 | Improved Specular Highlights With Adaptive ShadingabstractGouraud shading and Phong shading are widely used interpolation methods to render a polygon mesh of a curved surface. When an illumination equation has a specular reflection term Phong shading produces more realistic results than Gouraud shading. The specular highlights produced by Phong shading give visual information about surface geometry and properties. But it is not implemented in most graphics workstations and software renderers due to its computational expense. This paper introduces an adaptive shading method that produces Phong-shaded quality images for a small increase in rendering time using Gouraud shading systems. Higher quality images can be obtained from existing graphics hardware or software through a simple modification of the application or library. Youngkwan Cho, Ulrich Neumann, Jongwook Woo |
Computer Graphics International | 2 |
| 1996 | A self-tracking augmented reality systemabstractWe present a color-video-based augmented reality (AR) system that is designed to be self-tracking that is, it requires no separate tracking subsystem. Rather, tracking is performed strictly from the video images acquired through the lens of the camera also used to view the real world. The methods for tracking are rooted in prior research in photogrammetry and computer vision. This approach to tracking for AR systems enables a variety of new applications in assembly guidance that are not feasible with current AR technology. Our initial application is in aircraft manufacturing. We outline our approaches to feature detection, correspondence, pose determination, and system calibration. The results obtained thus far are summarized along with the problems we encountered. Ulrich Neumann, Youngkwan Cho |
VRST | 1 |
| 1995 | Interactive Volume Visualization on a Heterogeneous Message-Passing MulticomputerabstractThis paper describes VOL2, an interactive general-purpose volume renderer based on ray casting and implemented on Pixel-Planes 5, a distributed-memory, message-passing multicomputer. VOL2 is a pipelined renderer using image-space task parallelism and object-space data partitioning. We describe the parallelization and load balancing techniques used in order to achieve interactive response and near-real-time frame rates. We also present a number of applications for our system and derive some general conclusions about operation of image-order rendering algorithms on message-passing multicomputers. Andrei State, Jonathan McAllister, Ulrich Neumann, Tim J. Cullip, David T. Chen, Henry Fuchs |
SI3D | 3 |
| 1992 | Interactive Volume Rendering on a MulticomputerabstractDirect volume rendering is a computationally intensive operation that has become a valued and often preferred visualization tool.For maximal data comprehension, interactive manipulation of the rendering parameters is desirable.To this end, a reasonable target would be a system capable of displaying 12E3 voxel data sets at multiple frames per second.Although the computing resources required to attain this performance are beyond those available in current uniprocessor workstations, multicomputers and VLSI rendering hardware offer a solution.This paper describes a volume rendering algorithm for MIMD message passing multicomputers.This algorithm addresses the issues of distributed rendering, data set distribution, load balancing, and contention for the routing network.An implementation on a multicomputer with a 1 D ring network is analyzed, and extension of the algorithm to a 2D mesh topology is described.In addition, the paper presents a method of exploiting screen coherence through the use of VLSI pixel processor arrays.Though not critical to the general algorithm, this rendering approach is demonstrated in the example implementation where it serves as a hardware accelerator of the rendering process.Commercial graphics workstations use pixel processors to accelerate polygon rendering; this paper proposes a new use of this hardware for accelerating volume rendering. Ulrich Neumann |
SI3D | 1 |
| 1992 | Real-Time Procedural TexturesabstractWe describe a software system on the Pixel-Planes 5 graphics engine that displays user-defined antialiased procedural textures at rates of about 30 frames per second for use in realtime graphics applications.Our system allows a user to create textures that can modulate both diffuse and specular color, the sharpness of specular highlights, the amount of transparency and the surface normals of an object.We describe a texture editor that allows a user to interactively create and edit procedural textures.Antialiasing is essential for real-time textures, and in this paper we present some techniques for antialiasing procedural textures.Another direction we are exploring is the use of dynamic textures, which are functions of time or orientation.Examples of textures we have generated include a translucent fire texture that waves and flickers and an animated water texture that shows the use of both environment mapping and normal perturbation (bump mapping). John Rhoades, Greg Turk, Andrei State, Ulrich Neumann, Amitabh Varshney |
SI3D | 5 |
| 1991 | Achieving Direct Volume Visualization with Interactive Semantic Region SelectionabstractThe authors have achieved rates as high as 15 frames per second for interactive direct visualization of 3D data by trading some function for speed, while volume rendering with a full complement of ramp classification capabilities is performed at 1.4 frames per second. These speeds have made the combination of region selection with volume rendering practical for the first time. Semantic-driven selection, rather than geometric clipping, has proved to be a natural means of interacting with 3D data. Internal organs in medical data or other regions of interest can be built from preprocessed region primitives. The resulting combined system has been applied to real 3D medical data with encouraging results.> Terry S. Yoo, Ulrich Neumann, Henry Fuchs, Stephen M. Pizer, Tim J. Cullip, John Rhoades, Ross T. Whitaker |
IEEE Visualization | 2 |