EDBT 2026 Demo / reviewers in the wild / expert
Xin Tong 0001
dblp:86/2176-1
· DBLP profile ↗
169ranked-venue papers
6as first author
53since 2021 · last 2026
0000-0001-8788-2453ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 148 · 4 first-author · 43 since 2021Artificial intelligence and machine learning · 48 · 26 since 2021Human-computer interaction and ubiquitous computing · 7 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SQuadGen: Generating Simple Quad Layouts via Chart Distance Fieldsabstract3D shapes from scanning, reconstruction, or AI-generated content often lack simple quad mesh layouts—critical for efficient editing and modeling. Existing quad-remeshing techniques typically produce complex layouts with irregular loops, leading to tedious manual cleanup and extensive algorithm tuning. We introduce SQUADGEN, a diffusion-based generative framework that leverages Chart Distance Fields (CDF) to synthesize simple quad layouts on 3D shapes. Our approach addresses two key challenges: (1) the discrete nature of mesh connectivity, which hinders learning, and (2) the scarcity of large-scale datasets with simple quad meshes. To overcome the first, we propose CDF, a continuous surface-based representation enabling effective learning and synthesis of quad layouts. To address the second, we define loop-aware simplicity metrics and construct a large-scale dataset of high-quality quad layouts recovered from public 3D repositories through a robust quad-recovery pipeline. Extensive evaluations across diverse 3D inputs show that SQUADGEN consistently outperforms existing methods, producing robust, artist-friendly simple quad layouts. Youkang Kong, Yang Liu 0014, Yue Dong 0001, Xin Tong 0001, Harry Shum |
ACM Trans. Graph. | 4 |
| 2026 | ComboStoc: Combinatorial Stochasticity for Diffusion Generative ModelsabstractIn this paper, we study an under-explored but important factor of diffusion generative models, i.e., the combinatorial complexity. Data samples are generally high-dimensional, and for various structured generation tasks, additional attributes are combined to associate with data samples. We show that the space spanned by the combination of dimensions and attributes can be insufficiently covered by existing training schemes of diffusion generative models, potentially limiting test time performance. We present a simple fix to this problem by constructing stochastic processes that fully exploit the combinatorial structures, hence the name ComboStoc. Using this simple strategy, we show that network training is significantly accelerated across diverse data modalities, including images and 3D structured shapes. Moreover, ComboStoc enables a new way of test time generation which uses asynchronous time steps for different dimensions and attributes, thus allowing for varying degrees of control over them. Our code is available at: https://github.com/Xrvitd/ComboStoc. Rui Xu 0016, Jiepeng Wang 0001, Hao Pan 0001, Yang Liu 0014, Xin Tong 0001, Shi-Qing Xin, Changhe Tu, Taku Komura, Wenping Wang 0001 |
ACM Trans. Graph. | 5 |
| 2026 | ESGaussianFace: Emotional and Stylized Audio-Driven Facial Animation via 3D Gaussian SplattingabstractMost current audio-driven facial animation research primarily focuses on generating videos with neutral emotions. While some studies have addressed the generation of facial videos driven by emotional audio, efficiently generating high-quality talking head videos that integrate both emotional expressions and style features remains a significant challenge. In this paper, we propose ESGaussianFace, an innovative framework for emotional and stylized audio-driven facial animation. Our approach leverages 3D Gaussian Splatting to reconstruct 3D scenes and render videos, ensuring efficient generation of 3D consistent results. We propose an emotion-audio-guided spatial attention method that effectively integrates emotion features with audio content features. Through emotion-guided attention, the model is able to reconstruct facial details across different emotional states more accurately. To achieve emotional and stylized deformations of the 3D Gaussian points through emotion and style features, we introduce two 3D Gaussian deformation predictors. Futhermore, we propose a multi-stage training strategy, enabling the step-by-step learning of the character's lip movements, emotional variations, and style features. Our generated results exhibit high efficiency, high quality, and 3D consistency. Extensive experimental results demonstrate that our method outperforms existing state-of-the-art techniques in terms of lip movement accuracy, expression variation, and style feature expressiveness. Chuhang Ma, Shuai Tan 0002, Jiaolong Yang, Xin Tong 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | SAMPro3D: Locating SAM Prompts in 3D for Zero-Shot Instance SegmentationabstractWe introduce SAMPro3D for zero-shot instance segmentation of 3D scenes. Given the 3D point cloud and multiple posed RGB-D frames of 3D scenes, our approach segments 3D instances by applying the pretrained Segment Anything Model (SAM) to 2D frames. Our key idea in-volves locating SAM prompts in 3D to align their projected pixel prompts across frames, ensuring the view consistency of SAM-predicted masks. Moreover, we suggest selecting prompts from the initial set guided by the information of SAM-predicted masks across all views, which enhances the overall performance. We further propose to consolidate different prompts if they are segmenting different surface parts of the same 3D instance, bringing a more comprehensive segmentation. Notably, our method does not require any additional training. Extensive experiments on diverse benchmarks show that our method achieves comparable or better performance compared to previous zero-shot or fully supervised approaches, and in many cases surpasses human annotations. Furthermore, since our fine-grained predictions often lack annotations in available datasets, we present ScanNet200-Fine50 test data which provides fine-grained annotations on 50 scenes from ScanNet200 dataset. Mutian Xu, Xingyilang Yin, Lingteng Qiu, Yang Liu 0014, Xin Tong 0001, Xiaoguang Han 0001 |
3DV | 5 |
| 2025 | MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training SupervisionabstractWe present MoGe, a powerful model for recovering 3D geometry from monocular open-domain images. Given a single image, our model directly predicts a 3D point map of the captured scene with an affine-invariant representation, which is agnostic to true global scale and shift. This new representation precludes ambiguous supervision in training and facilitates effective geometry learning. Furthermore, we propose a set of novel global and local geometry supervision techniques that empower the model to learn high-quality geometry. These include a robust, optimal, and efficient point cloud alignment solver for accurate global shape learning, and a multi-scale local geometry loss promoting precise local geometry supervision. We train our model on a large, mixed dataset and demonstrate its strong generalizability and high accuracy. In our comprehensive evaluation on diverse unseen datasets, our model significantly outperforms state-of-the-art methods across all tasks, including monocular estimation of 3D point map, depth map, and camera field of view. Ruicheng Wang, Sicheng Xu, Cassie Dai, Jianfeng Xiang, Yu Deng 0006, Xin Tong 0001, Jiaolong Yang |
CVPR | 6 |
| 2025 | Structured 3D Latents for Scalable and Versatile 3D GenerationabstractWe introduce a novel 3D generation method for versatile and high-quality 3D asset creation. The cornerstone is a unified Structured LATent (SLat) representation which allows decoding to different output formats, such as Radiance Fields, 3D Gaussians, and meshes. This is achieved by integrating a sparsely-populated 3D grid with dense multiview visual features extracted from a powerful vision foundation model, comprehensively capturing both structural (geometry) and textural (appearance) information while maintaining flexibility during decoding.We employ rectified flow transformers tailored for SLat as our 3D generation models and train models with up to 2 billion parameters on a large 3D asset dataset of 500K diverse objects. Our model generates high-quality results with text or image conditions, significantly surpassing existing methods, including recent ones at similar scales. We showcase flexible output format selection and local 3D editing capabilities which were not offered by previous models. Project Page: trellis3d.github.io. Jianfeng Xiang, Zelong Lv, Sicheng Xu, Yu Deng 0006, Ruicheng Wang, Bowen Zhang 0010, Dong Chen 0003, Xin Tong 0001, Jiaolong Yang |
CVPR | 8 |
| 2025 | Robust Optical Transceiver Manipulation in Cluttered Cable Environments Using 3D Scene Understanding and PlanningabstractRobotic manipulation in cluttered environments presents significant challenges, particularly when the clutter includes thin, deformable objects like cables, which complicate perception and decision-making processes. In the context of datacenters, the automation of networking tasks often involves the manipulation of optical transceivers within densely packed cable configurations. Such environments are characterized by an abundance of delicate, overlapping, and intersecting cables, leading to frequent occlusions. This paper introduces an innovative system designed for the manipulation of optical transceivers in environments cluttered by cables. Our integrated approach combines advanced 3D scene understanding with a heuristic-based pushing policy to effectively manipulate optical transceivers amidst clutter. The system's perception component utilizes image segmentation and 3D reconstruction to accurately model the transceivers and surrounding cables. Meanwhile, the planning aspect employs a search algorithm with task-specific heuristics, to navigate the gripper, displace obstructing cables, and safely achieve a precise pre-grasp position in front of the target transceiver. We have conducted extensive evaluations of our methodology in both simulated and real-world settings, demonstrating its high success rates, robustness, and proficiency in addressing the unique challenges posed by cable-occluded environments within datacenters. Iason Sarantopoulos, Bohong Weng, Sicheng Xu, Jiaolong Yang, Xin Tong 0001, Fabian Otto, David Sweeney, Andromachi Chatzieleftheriou, Antony I. T. Rowstron |
ICRA | 7 |
| 2025 | MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp DetailsabstractWe propose MoGe-2, an advanced open-domain geometry estimation model that recovers a metric-scale 3D point map of a scene from a single image. Our method builds upon the recent monocular geometry estimation approach, MoGe, which predicts affine-invariant point maps with unknown scales. We explore effective strategies to extend MoGe for metric geometry prediction without compromising the relative geometry accuracy provided by the affine-invariant point representation. Additionally, we discover that noise and errors in real data diminish fine-grained detail in the predicted geometry. We address this by developing a data refinement approach that filters and completes real data using sharp synthetic labels, significantly enhancing the granularity of the reconstructed geometry while maintaining the overall accuracy. We train our model on a large corpus of mixed datasets and conducted comprehensive evaluations, demonstrating its superior performance in achieving accurate relative geometry, precise metric scale, and fine-grained detail recovery -- capabilities that no previous methods have simultaneously achieved. Ruicheng Wang, Sicheng Xu, Yue Dong 0001, Yu Deng 0006, Jianfeng Xiang, Zelong Lv, Guangzhong Sun, Xin Tong 0001, Jiaolong Yang |
NeurIPS | 8 |
| 2025 | Swin3D: A Pretrained Transformer Backbone for 3D Indoor Scene UnderstandingabstractThe use of pretrained backbones with fine-tuning has shown success for 2D vision and natural language processing tasks, with advantages over task-specific networks. In this paper, we introduce a pretrained 3D backbone, called Swin3d, for 3D indoor scene understanding. We designed a 3D Swin Transformer as our backbone network, which enables efficient self-attention on sparse voxels with linear memory complexity, making the backbone scalable to large models and datasets. We also introduce a generalized contextual relative positional embedding scheme to capture various irregularities of point signals for improved network performance. We pretrained a large Swin3d model on a synthetic Structured3D dataset, which is an order of magnitude larger than the ScanNet dataset. Our model pretrained on the synthetic dataset not only generalizes well to downstream segmentation and detection on real 3D point datasets but also outperforms state-of-the-art methods on downstream tasks with +2.3 mIoU and +2.2 mIoU on S3DIS Area5 and 6-fold semantic segmentation, respectively, +1.8 mIoU on ScanNet segmentation (val), +1.9 [email protected] on ScanNet detection, and +8.1 [email protected] on S3DIS detection. A series of extensive ablation studies further validated the scalability, generality, and superior performance enabled by our approach. Yuxiao Guo 0001, Jian-Yu Xiong, Yang Liu 0014, Hao Pan 0001, Peng-Shuai Wang, Xin Tong 0001, Baining Guo |
Comput. Vis. Media | 7 |
| 2025 | MPS-NeRF: Generalizable 3D Human Rendering From Multiview ImagesabstractThere has been rapid progress recently on 3D human rendering, including novel view synthesis and pose animation, based on the advances of neural radiance fields (NeRF). However, most existing methods focus on person-specific training and their training typically requires multi-view videos. This article deals with a new challenging task - rendering novel views and novel poses for a person unseen in training, using only multiview still images as input without videos. For this task, we propose a simple yet surprisingly effective method to train a generalizable NeRF with multiview images as conditional input. The key ingredient is a dedicated representation combining a canonical NeRF and a volume deformation scheme. Using a canonical space enables our method to learn shared properties of human and easily generalize to different people. Volume deformation is used to connect the canonical space with input and target images and query image features for radiance and density prediction. We leverage the parametric 3D human model fitted on the input images to derive the deformation, which works quite well in practice when combined with our canonical NeRF. The experiments on both real and synthetic data with the novel view synthesis and pose animation tasks collectively demonstrate the efficacy of our method. Xiangjun Gao, Jiaolong Yang, Jongyoo Kim, Sida Peng, Zicheng Liu 0001, Xin Tong 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | StructRe: Rewriting for Structured Shape ModelingabstractMan-made 3D shapes are naturally organized in parts and hierarchies; such structures provide important constraints for shape reconstruction and generation. Modeling shape structures is difficult, because there can be multiple hierarchies for a given shape, causing ambiguity, and across different categories, the shape structures are correlated with semantics, limiting generalization. We present StructRe , a structure rewriting system, as a novel approach to structured shape modeling. Given a 3D object represented by points and components, StructRe can rewrite it upward into more concise structures, or downward into more detailed structures; by iterating the rewriting process, hierarchies are obtained. Such a localized rewriting process enables probabilistic modeling of ambiguous structures and robust generalization across object categories. We train StructRe on PartNet data and show its generalization to cross-category and multiple object hierarchies, and test its extension to ShapeNet. We also demonstrate the benefits of probabilistic and generalizable structure modeling for shape reconstruction, generation and editing tasks. Jiepeng Wang 0001, Hao Pan 0001, Yang Liu 0014, Xin Tong 0001, Taku Komura, Wenping Wang 0001 |
ACM Trans. Graph. | 4 |
| 2025 | A Real-Time Method for Inserting Virtual Objects Into Neural Radiance FieldsabstractWe present the first real-time method for inserting a rigid virtual object into a neural radiance field (NeRF), which produces realistic lighting and shadowing effects, as well as allows interactive manipulation of the object. By exploiting the rich information about lighting and geometry in a NeRF, our method overcomes several challenges of object insertion in augmented reality. For lighting estimation, we produce accurate and robust incident lighting that combines the 3D spatially-varying lighting from NeRF and an environment lighting to account for sources not covered by the NeRF. For occlusion, we blend the rendered virtual object with the background scene using an opacity map integrated from the NeRF. For shadows, with a precomputed field of spherical signed distance fields, we query the visibility term for any point around the virtual object, and cast soft, detailed shadows onto 3D surfaces. Compared with state-of-the-art techniques, our approach can insert virtual objects into scenes with superior fidelity, and has great potential to be further applied to augmented reality systems. Keyang Ye, Hongzhi Wu, Xin Tong 0001, Kun Zhou 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | Plan, Posture and Go: Towards Open-Vocabulary Text-to-Motion Generation
Wenxun Dai, Chunyu Wang 0001, Yiji Cheng, Yansong Tang, Xin Tong 0001 |
ECCV (27) | 6 |
| 2024 | Diffusion Models are Geometry Critics: Single Image 3D Editing Using Pre-trained Diffusion Priors
Ruicheng Wang, Jianfeng Xiang, Jiaolong Yang, Xin Tong 0001 |
ECCV (56) | 4 |
| 2024 | A Simple Baseline for Spoken Language to Sign Language Translation with 3D Avatars
Ronglai Zuo, Fangyun Wei, Zenggui Chen, Brian Kan-Wing Mak, Jiaolong Yang, Xin Tong 0001 |
ECCV (49) | 6 |
| 2024 | 3D Feature Prediction for Masked-AutoEncoder-Based Point Cloud PretrainingabstractMasked autoencoders (MAE) have recently been introduced to 3D self-supervised pretraining for point clouds due to their great success in NLP and computer vision. Unlike MAEs used in the image domain, where the pretext task is to restore features at the masked pixels, such as colors, the existing 3D MAE works reconstruct the missing geometry only, i.e, the location of the masked points. In contrast to previous studies, we advocate that point location recovery is inessential and restoring intrinsic point features is much superior. To this end, we propose to ignore point position reconstruction and recover high-order features at masked points including surface normals and surface variations, through a novel attention-based decoder which is independent of the encoder design. We validate the effectiveness of our pretext task and decoder design using different encoder structures for 3D training and demonstrate the advantages of our pretrained networks on various point cloud analysis tasks. Siming Yan, Yuxiao Guo 0001, Hao Pan 0001, Peng-Shuai Wang, Xin Tong 0001, Yang Liu 0014, Qixing Huang |
ICLR | 6 |
| 2024 | VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real TimeabstractWe introduce VASA, a framework for generating lifelike talking faces with appealing visual affective skills (VAS) given a single static image and a speech audio clip. Our premiere model, VASA-1, is capable of not only generating lip movements that are exquisitely synchronized with the audio, but also producing a large spectrum of facial nuances and natural head motions that contribute to the perception of authenticity and liveliness.
The core innovations include a diffusion-based holistic facial dynamics and head movement generation model that works in a face latent space, and the development of such an expressive and disentangled face latent space using videos.
Through extensive experiments including evaluation on a set of new metrics, we show that our method significantly outperforms previous methods along various dimensions comprehensively. Our method delivers high video quality with realistic facial and head dynamics and also supports the online generation of 512$\times$512 videos at up to 40 FPS with negligible starting latency.
It paves the way for real-time engagements with lifelike avatars that emulate human conversational behaviors. Sicheng Xu, Yuxiao Guo 0001, Jiaolong Yang, Zhenyu Zang, Xin Tong 0001, Baining Guo |
NeurIPS | 8 |
| 2024 | Touchscreen-based Hand Tracking for Remote Whiteboard InteractionabstractIn whiteboard-based remote communication, the seamless integration of drawn content and hand-screen interactions is essential for an immersive user experience. Previous methods either require bulky device setups for capturing hand gestures or fail to accurately track the hand poses from capacitive images. In this paper, we present a real-time method for precise tracking 3D poses of both hands from capacitive video frames. To this end, we develop a deep neural network to identify hands and infer hand joint positions from capacitive frames, and then recover 3D hand poses from the hand-joint positions via a constrained inverse kinematic solver. Additionally, we design a device setup for capturing high-quality hand-screen interaction data and obtained a more accurate synchronized capacitive video and hand pose dataset. Our method improves the accuracy and stability of 3D hand tracking for capacitive frames while maintaining a compact device setup for remote communication. We validate our scheme design and its superior performance on 3D hand pose tracking and demonstrate the effectiveness of our method in whiteboard-based remote communication. Xinshuang Liu, Xin Tong 0001 |
UIST | 3 |
| 2024 | Neural Path Sampling for Rendering Pure Specular Light TransportabstractAbstract Multi‐bounce, pure specular light paths produce complex lighting effects, such as caustics and sparkle highlights, which are challenging to render due to their sparse and diverse nature. We introduce a learning‐based method for the efficient rendering of pure specular light transport. The key idea is training a neural network to model the distribution of all specular light paths between pairs of endpoints for one specular object. To achieve this, for each object, our method models the distribution of sparse and diverse specular light paths between two endpoints using smooth 2D maps of ray directions from one endpoint and represents these maps with a 2D convolutional network. We design a training scheme to efficiently sample specular light paths from the scene and train the network. Once trained, our method predicts specular light paths for a given pair of endpoints using the network and employs root‐finding‐based algorithms for rendering the specular light transport. Experimental results demonstrate that our method generates high‐quality results, supports dynamic lighting and moving objects within the scene, and significantly enhances the rendering speed of existing techniques. Yue Dong 0001, Youkang Kong, Xin Tong 0001 |
Comput. Graph. Forum | 4 |
| 2024 | Spin-Weighted Spherical Harmonics for Polarized Light TransportabstractThe objective of polarization rendering is to simulate the interaction of light with materials exhibiting polarization-dependent behavior. However, integrating polarization into rendering is challenging and increases computational costs significantly. The primary difficulty lies in efficiently modeling and computing the complex reflection phenomena associated with polarized light. Specifically, frequency-domain analysis, essential for efficient environment lighting and storage of complex light interactions, is lacking. To efficiently simulate and reproduce polarized light interactions using frequency-domain techniques, we address the challenge of maintaining continuity in polarized light transport represented by Stokes vectors within angular domains. The conventional spherical harmonics method cannot effectively handle continuity and rotation invariance for Stokes vectors. To overcome this, we develop a new method called polarized spherical harmonics (PSH) based on the spin-weighted spherical harmonics theory. Our method provides a rotation-invariant representation of Stokes vector fields. Furthermore, we introduce frequency domain formulations of polarized rendering equations and spherical convolution based on PSH. We first define spherical convolution on Stokes vector fields in the angular domain, and it also provides efficient computation of polarized light transport, nearly on an entry-wise product in the frequency domain. Our frequency domain formulation, including spherical convolution, led to the development of the first real-time polarization rendering technique under polarized environmental illumination, named precomputed polarized radiance transfer, using our polarized spherical harmonics. Results demonstrate that our method can effectively and accurately simulate and reproduce polarized light interactions in complex reflection phenomena, including polarized environmental illumination and soft shadows. Shinyoung Yi 0001, Donggun Kim 0002, Jiwoong Na, Xin Tong 0001, Min H. Kim 0001 |
ACM Trans. Graph. | 4 |
| 2023 | NeRFInvertor: High Fidelity NeRF-GAN Inversion for Single-Shot Real Image AnimationabstractNerf-based Generative models have shown impressive capacity in generating high-quality images with consistent 3D geometry. Despite successful synthesis of fake identity images randomly sampled from latent space, adopting these models for generating face images of real subjects is still a challenging task due to its so-called inversion issue. In this paper, we propose a universal method to surgically finetune these NeRF-GAN models in order to achieve high-fidelity animation of real subjects only by a single image. Given the optimized latent code for an out-of-domain real image, we employ 2D loss functions on the rendered image to reduce the identity gap. Furthermore, our method leverages explicit and implicit 3D regularizations using the in-domain neighborhood samples around the optimized latent code to remove geometrical and visual artifacts. Our experiments confirm the effectiveness of our method in realistic, high-fidelity, and 3D consistent animation of real faces on multiple NeRF-GAN models across different datasets. Yu Yin 0001, Kamran Ghasedi, HsiangTao Wu, Jiaolong Yang, Xin Tong 0001, Yun Fu 0001 |
CVPR | 5 |
| 2023 | GRAM-HD: 3D-Consistent Image Generation at High Resolution with Generative Radiance ManifoldsabstractRecent works have shown that 3D-aware GANs trained on unstructured single image collections can generate multiview images of novel instances. The key underpinnings to achieve this are a 3D radiance field generator and a volume rendering process. However, existing methods either cannot generate high-resolution images (e.g., up to 256×256) due to the high computation cost of neural volume rendering, or rely on 2D CNNs for image-space upsampling which jeopardizes the 3D consistency across different views. This paper proposes a novel 3D-aware GAN that can generate high resolution images (up to 1024 1024) while keeping strict 3D consistency as in volume ×rendering. Our motivation is to achieve super-resolution directly in the 3D space to preserve 3D consistency. We avoid the otherwise prohibitively-expensive computation cost by applying 2D convolutions on a set of 2D radiance manifolds defined in the recent generative radiance manifold (GRAM) approach, and apply dedicated loss functions for effective GAN training at high resolution. Experiments on FFHQ and AFHQv2 datasets show that our method can produce high-quality 3D-consistent results that significantly outperform existing methods. It makes a significant step towards closing the gap between traditional 2D image generation and 3D-consistent free-view generation.1 Jianfeng Xiang, Jiaolong Yang, Yu Deng 0006, Xin Tong 0001 |
ICCV | 4 |
| 2023 | 3D-aware Image Generation using 2D Diffusion ModelsabstractIn this paper, we introduce a novel 3D-aware image generation method that leverages 2D diffusion models. We formulate the 3D-aware image generation task as multiview 2D image set generation, and further to a sequential unconditional–conditional multiview image generation process. This allows us to utilize 2D diffusion models to boost the generative modeling power of the method. Additionally, we incorporate depth information from monocular depth estimators to construct the training data for the conditional diffusion model using only still images.We train our method on a large-scale unstructured 2D image dataset, i.e., ImageNet, which is not addressed by previous methods. It produces high-quality images that significantly outperform prior methods. Furthermore, our approach showcases its capability to generate instances with large view angles, even though the training images are diverse and unaligned, gathered from "in-the-wild" realworld environments.1 Jianfeng Xiang, Jiaolong Yang, Binbin Huang 0004, Xin Tong 0001 |
ICCV | 4 |
| 2023 | AniPortraitGAN: Animatable 3D Portrait Generation from 2D Image CollectionsabstractPrevious animatable 3D-aware GANs for human generation have primarily focused on either the human head or full body. However, head-only videos are relatively uncommon in real life, and full body generation typically does not deal with facial expression control and still has challenges in generating high-quality results. Towards applicable video avatars, we present an animatable 3D-aware GAN that generates portrait images with controllable facial expression, head pose, and shoulder movements. It is a generative model trained on unstructured 2D image collections without using 3D or video data. For the new task, we base our method on the generative radiance manifold representation and equip it with learnable facial and head-shoulder deformations. A dual-camera rendering and adversarial learning scheme is proposed to improve the quality of the generated faces, which is critical for portrait images. A pose deformation processing network is developed to generate plausible deformations for challenging regions such as long hair. Experiments show that our method, trained on unstructured 2D images, can generate diverse and high-quality 3D portraits with desired control over different properties. Yue Wu 0012, Sicheng Xu, Jianfeng Xiang, Fangyun Wei, Qifeng Chen 0001, Jiaolong Yang, Xin Tong 0001 |
SIGGRAPH Asia | 7 |
| 2023 | RemoteTouch: Enhancing Immersive 3D Video Communication with Hand TouchabstractRecent research advance has significantly improved the visual real-ism of immersive 3D video communication. In this work we present a method to further enhance this immersive experience by adding the hand touch capability (“remote hand clapping”). In our system, each meeting participant sits in front of a large screen with haptic feedback. The local participant can reach his hand out to the screen and perform hand clapping with the remote participant as if the two participants were only separated by a virtual glass. A key challenge in emulating the remote hand touch is the realistic rendering of the participant's hand and arm as the hand touches the screen. When the hand is very close to the screen, the RGBD data required for realistic rendering is no longer available. To tackle this challenge, we present a dual representation of the user's hand. Our dual representation not only preserves the high-quality rendering usually found in recent image-based rendering systems but also allows the hand to reach to the screen. This is possible because the dual representation includes both an image-based model and a 3D geometry-based model, with the latter driven by a hand skeleton tracked by a side view camera. In addition, the dual representation provides a distance-based fusion of the image-based and 3D geometry-based models as the hand moves closer to the screen. The result is that the image-based and 3D geometry-based models mutually enhance each other, leading to realistic and seamless rendering. Our experiments demonstrate that our method provides consistent hand contact experience between remote users and improves the immersive experience of 3D video communication. Zhiqi Li 0004, Sicheng Xu, Jiaolong Yang, Xin Tong 0001, Baining Guo |
VR | 6 |
| 2023 | Semantic segmentation-assisted instance feature fusion for multi-level 3D part instance segmentationabstractRecognizing 3D part instances from a 3D point cloud is crucial for 3D structure and scene understanding. Several learning-based approaches use semantic segmentation and instance center prediction as training tasks and fail to further exploit the inherent relationship between shape semantics and part instances. In this paper, we present a new method for 3D part instance segmentation. Our method exploits semantic segmentation to fuse nonlocal instance features, such as center prediction, and further enhances the fusion scheme in a multi- and cross-level way. We also propose a semantic region center prediction task to train and leverage the prediction results to improve the clustering of instance points. Our method outperforms existing methods with a large-margin improvement in the PartNet benchmark. We also demonstrate that our feature fusion scheme can be applied to other existing methods to improve their performance in indoor scene instance segmentation tasks. Chun-Yu Sun, Xin Tong 0001, Yang Liu 0014 |
Comput. Vis. Media | 2 |
| 2023 | Semi-supervised 3D shape segmentation with multilevel consistency and part substitutionabstractThe lack of fine-grained 3D shape segmentation data is the main obstacle to developing learning-based 3D segmentation techniques. We propose an effective semi-supervised method for learning 3D segmentations from a few labeled 3D shapes and a large amount of unlabeled 3D data. For the unlabeled data, we present a novel multilevel consistency loss to enforce consistency of network predictions between perturbed copies of a 3D shape at multiple levels: point level, part level, and hierarchical level. For the labeled data, we develop a simple yet effective part substitution scheme to augment the labeled 3D shapes with more structural variations to enhance training. Our method has been extensively validated on the task of 3D object semantic segmentation on PartNet and ShapeNetPart, and indoor scene semantic segmentation on ScanNet. It exhibits superior performance to existing semi-supervised and unsupervised pre-training 3D approaches. Chun-Yu Sun, Hao-Xiang Guo 0001, Peng-Shuai Wang, Xin Tong 0001, Yang Liu 0014, Harry Shum |
Comput. Vis. Media | 5 |
| 2023 | Locally Attentional SDF Diffusion for Controllable 3D Shape GenerationabstractAlthough the recent rapid evolution of 3D generative neural networks greatly improves 3D shape generation, it is still not convenient for ordinary users to create 3D shapes and control the local geometry of generated shapes. To address these challenges, we propose a diffusion-based 3D generation framework --- locally attentional SDF diffusion , to model plausible 3D shapes, via 2D sketch image input. Our method is built on a two-stage diffusion model. The first stage, named occupancy-diffusion , aims to generate a low-resolution occupancy field to approximate the shape shell. The second stage, named SDF-diffusion , synthesizes a high-resolution signed distance field within the occupied voxels determined by the first stage to extract fine geometry. Our model is empowered by a novel view-aware local attention mechanism for image-conditioned shape generation, which takes advantage of 2D image patch features to guide 3D voxel feature learning, greatly improving local controllability and model generalizability. Through extensive experiments in sketch-conditioned and category-conditioned 3D shape generation tasks, we validate and demonstrate the ability of our method to provide plausible and diverse 3D shapes, as well as its superior controllability and generalizability over existing work. Xin-Yang Zheng, Hao Pan 0001, Peng-Shuai Wang, Xin Tong 0001, Yang Liu 0014, Harry Shum |
ACM Trans. Graph. | 4 |
| 2022 | GRAM: Generative Radiance Manifolds for 3D-Aware Image Generationabstract3D-aware image generative modeling aims to generate 3D-consistent images with explicitly controllable camera poses. Recent works have shown promising results by training neural radiance field (NeRF) generators on unstructured 2D images, but still cannot generate highly-realistic images with fine details. A critical reason is that the high memory and computation cost of volumetric representation learning greatly restricts the number of point samples for radiance integration during training. Deficient sampling not only limits the expressive power of the generator to handle fine details but also impedes effective GAN training due to the noise caused by unstable Monte Carlo sampling. We propose a novel approach that regulates point sampling and radiance field learning on 2D manifolds, embodied as a set of learned implicit surfaces in the 3D volume. For each viewing ray, we calculate ray-surface intersections and accumulate their radiance generated by the network. By training and rendering such radiance mani folds, our generator can produce high quality images with realistic fine details and strong visual 3D consistency.11Project page: https://yudeng.github.io/GRAM/ Yu Deng 0006, Jiaolong Yang, Jianfeng Xiang, Xin Tong 0001 |
CVPR | 4 |
| 2022 | AniFaceGAN: Animatable 3D-Aware Face Image Generation for Video AvatarsabstractAlthough 2D generative models have made great progress in face image generation and animation, they often suffer from undesirable artifacts such as 3D inconsistency when rendering images from different camera viewpoints. This prevents them from synthesizing video animations indistinguishable from real ones. Recently, 3D-aware GANs extend 2D GANs for explicit disentanglement of camera pose by leveraging 3D scene representations. These methods can well preserve the 3D consistency of the generated images across different views, yet they cannot achieve fine-grained control over other attributes, among which facial expression control is arguably the most useful and desirable for face animation. In this paper, we propose an animatable 3D-aware GAN for multiview consistent face animation generation. The key idea is to decompose the 3D representation of the 3D-aware GAN into a template field and a deformation field, where the former represents different identities with a canonical expression, and the latter characterizes expression variations of each identity. To achieve meaningful control over facial expressions via deformation, we propose a 3D-level imitative learning scheme between the generator and a parametric 3D face model during adversarial training of the 3D-aware GAN. This helps our method achieve high-quality animatable face image generation with strong visual 3D consistency, even though trained with only unstructured 2D images. Extensive experiments demonstrate our superior performance over prior works. Project page: \url{https://yuewuhkust.github.io/AniFaceGAN/ Yue Wu 0012, Yu Deng 0006, Jiaolong Yang, Fangyun Wei, Qifeng Chen 0001, Xin Tong 0001 |
NeurIPS | 6 |
| 2022 | Classifier Guided Temporal Supersampling for Real-time RenderingabstractAbstract We present a learning based temporal supersampling algorithm for real‐time rendering. Different from existing learning‐based approaches that adopt an end‐to‐end training of a ‘black‐box’ neural network, we design a ‘white‐box’ solution that first classifies the pixels into different categories and then generates the supersampling result based on classification. Our key observation is that the core problem in temporal supersampling for rendering is to distinguish the pixels that consist of occlusion, aliasing, or shading changes. Samples from these pixels exhibit similar temporal radiance change but require different composition strategies to produce the correct supersampling result. Based on this observation, our method first classifies the pixels into several classes. Based on the classification results, our method then blends the current frame with the warped last frame via a learned weight map to get the supersampling results. We design compact neural networks for each step and develop dedicated loss functions for pixels belonging to different classes. Compared to existing learning based methods, our classifier‐based supersampling scheme takes less computational and memory cost for real‐time supersampling and generates visually compelling temporal supersampling results with fewer flickering artifacts. We evaluate the performance and generality of our method on several rendered game sequences and our method can upsample the rendered frames from 1080P to 2160P in just 13.39ms on a single Nvidia 3090GPU. Yuxiao Guo 0001, Yue Dong 0001, Xin Tong 0001 |
Comput. Graph. Forum | 4 |
| 2022 | Generative Deformable Radiance Fields for Disentangled Image Synthesis of Topology-Varying ObjectsabstractAbstract 3D‐aware generative models have demonstrated their superb performance to generate 3D neural radiance fields (NeRF) from a collection of monocular 2D images even for topology‐varying object categories. However, these methods still lack the capability to separately control the shape and appearance of the objects in the generated radiance fields. In this paper, we propose a generative model for synthesizing radiance fields of topology‐varying objects with disentangled shape and appearance variations. Our method generates deformable radiance fields, which builds the dense correspondence between the density fields of the objects and encodes their appearances in a shared template field. Our disentanglement is achieved in an unsupervised manner without introducing extra labels to previous 3D‐aware GAN training. We also develop an effective image inversion scheme for reconstructing the radiance field of an object in a real monocular image and manipulating its shape and appearance. Experiments show that our method can successfully learn the generative model from unstructured monocular images and well disentangle the shape and appearance for objects (e.g., chairs) with large topological variance. The model trained on synthetic data can faithfully reconstruct the real object in a given single image and achieve high‐quality texture and shape editing results. Yu Deng 0006, Jiaolong Yang, Jingyi Yu 0001, Xin Tong 0001 |
Comput. Graph. Forum | 5 |
| 2022 | SDF-StyleGAN: Implicit SDF-Based StyleGAN for 3D Shape GenerationabstractAbstract We present a StyleGAN2‐based deep learning approach for 3D shape generation, called SDF‐StyleGAN, with the aim of reducing visual and geometric dissimilarity between generated shapes and a shape collection. We extend StyleGAN2 to 3D generation and utilize the implicit signed distance function (SDF) as the 3D shape representation, and introduce two novel global and local shape discriminators that distinguish real and fake SDF values and gradients to significantly improve shape geometry and visual quality. We further complement the evaluation metrics of 3D generative models with the shading‐image‐based Fréchet inception distance (FID) scores to better assess visual quality and shape distribution of the generated shapes. Experiments on shape generation demonstrate the superior performance of SDF‐StyleGAN over the state‐of‐the‐art. We further demonstrate the efficacy of SDF‐StyleGAN in various tasks based on GAN inversion, including shape reconstruction, shape completion from partial point clouds, single‐view image‐based shape generation, and shape style editing. Extensive ablation studies justify the efficacy of our framework design. Our code and trained models are available at https://github.com/Zhengxinyang/SDF‐StyleGAN . Xin-Yang Zheng, Yang Liu 0014, Peng-Shuai Wang, Xin Tong 0001 |
Comput. Graph. Forum | 4 |
| 2022 | Message from the Best Paper Award CommitteeabstractVisual Media were recommended by the associate editors as candidate papers for the Best Paper Award.The Editor-in-Chief then invited the three of us to serve as the committee for choosing the Best Paper.After careful discussion by the committee, the following paper is chosen as the winner of the Best Paper Award: EfficientPose: Efficient human pose estimation with neural architecture search [1] while two other papers are awarded the Honorable Mention Awards: Efficient fastest-path computations for road maps [2]Inferring object properties from human interaction and transferring them to new motions [3] The Best Paper Award Committee would like to offer congratulations to the winners, who in addition to the prestige conferred upon them by the awards, will also receive cash prizes: the Best Paper will receive US Ming C. Lin, Xin Tong 0001, Wenping Wang 0001 |
Comput. Vis. Media | 2 |
| 2022 | Three-dimensional shape space learning for visual concept construction: challenges and research progressabstract人类可以熟练的对真实世界中物体按照形状或者功能进行分类, 并在思维中建立每类物体的视觉概念和周围真实世界的视觉知识 (Pan, 2019). Pan (2021) 指出建立这些视觉概念和视觉知识的计算表达是发展下一代人工智能的一个关键步骤. 学习同一视觉概念下所有物体的三维形状空间是实现视觉概念计算表达的一个关键步骤. 本文提出三维形状空间学习中面临的关键技术挑战, 并围绕这些技术挑战回顾了这一领域的研究进展, 最后讨论了三维形状空间学习领域的研究趋势和未来发展方向. Xin Tong 0001 |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2022 | Face Restoration via Plug-and-Play 3D Facial PriorsabstractState-of-the-art face restoration methods employ deep convolutional neural networks (CNNs) to learn a mapping between degraded and sharp facial patterns by exploring local appearance knowledge. However, most of these methods do not well exploit facial structures and identity information, and only deal with task-specific face restoration (e.g., face super-resolution or deblurring). In this paper, we propose cross-tasks and cross-models plug-and-play 3D facial priors to explicitly embed the network with the sharp facial structures for general face restoration tasks. Our 3D priors are the first to explore 3D morphable knowledge based on the fusion of parametric descriptions of face attributes (e.g., identity, facial expression, texture, illumination, and face pose). Furthermore, the priors can easily be incorporated into any network and are very efficient in improving the performance and accelerating the convergence speed. Firstly, a 3D face rendering branch is set up to obtain 3D priors of salient facial structures and identity knowledge. Secondly, for better exploiting this hierarchical information (i.e., intensity similarity, 3D facial structure, and identity content), a spatial attention module is designed for the image restoration problems. Extensive face restoration experiments including face super-resolution and deblurring demonstrate that the proposed 3D priors achieve superior face restoration results over the state-of-the-art algorithms. Xiaobin Hu, Wenqi Ren, Jiaolong Yang, Xiaochun Cao, David P. Wipf, Bjoern Menze, Xin Tong 0001, Hongbin Zha |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2022 | SkeletonNet: A Topology-Preserving Solution for Learning Mesh Reconstruction of Object Surfaces From RGB ImagesabstractThis paper focuses on the challenging task of learning 3D object surface reconstructions from RGB images. Existing methods achieve varying degrees of success by using different surface representations. However, they all have their own drawbacks, and cannot properly reconstruct the surface shapes of complex topologies, arguably due to a lack of constraints on the topological structures in their learning frameworks. To this end, we propose to learn and use the topology-preserved, skeletal shape representation to assist the downstream task of object surface reconstruction from RGB images. Technically, we propose the novel SkeletonNet design that learns a volumetric representation of a skeleton via a bridged learning of a skeletal point set, where we use parallel decoders each responsible for the learning of points on 1D skeletal curves and 2D skeletal sheets, as well as an efficient module of globally guided subvolume synthesis for a refined, high-resolution skeletal volume; we present a differentiable Point2Voxel layer to make SkeletonNet end-to-end and trainable. With the learned skeletal volumes, we propose two models, the Skeleton-Based Graph Convolutional Neural Network (SkeGCNN) and the Skeleton-Regularized Deep Implicit Surface Network (SkeDISN), which respectively build upon and improve over the existing frameworks of explicit mesh deformation and implicit field learning for the downstream surface reconstruction task. We conduct thorough experiments that verify the efficacy of our proposed SkeletonNet. SkeGCNN and SkeDISN outperform existing methods as well, and they have their own merits when measured by different metrics. Additional results in generalized task settings further demonstrate the usefulness of our proposed methods. We have made our implementation code publicly available at https://github.com/tangjiapeng/SkeletonNet. Jiapeng Tang, Xiaoguang Han 0001, Mingkui Tan, Xin Tong 0001, Kui Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | ComplexGen: CAD reconstruction by B-rep chain complex generationabstractWe view the reconstruction of CAD models in the boundary representation (B-Rep) as the detection of geometric primitives of different orders, i.e. , vertices, edges and surface patches, and the correspondence of primitives, which are holistically modeled as a chain complex, and show that by modeling such comprehensive structures more complete and regularized reconstructions can be achieved. We solve the complex generation problem in two steps. First, we propose a novel neural framework that consists of a sparse CNN encoder for input point cloud processing and a tri-path transformer decoder for generating geometric primitives and their mutual relationships with estimated probabilities. Second, given the probabilistic structure predicted by the neural network, we recover a definite B-Rep chain complex by solving a global optimization maximizing the likelihood under structural validness constraints and applying geometric refinements. Extensive tests on large scale CAD datasets demonstrate that the modeling of B-Rep chain complex structure enables more accurate detection for learning and more constrained reconstruction for optimization, leading to structurally more faithful and complete CAD B-Rep models than previous results. Hao-Xiang Guo 0001, Hao Pan 0001, Yang Liu 0014, Xin Tong 0001, Baining Guo |
ACM Trans. Graph. | 5 |
| 2022 | Sparse ellipsometry: portable acquisition of polarimetric SVBRDF and shape with unstructured flash photographyabstractEllipsometry techniques allow to measure polarization information of materials, requiring precise rotations of optical components with different configurations of lights and sensors. This results in cumbersome capture devices, carefully calibrated in lab conditions, and in very long acquisition times, usually in the order of a few days per object. Recent techniques allow to capture polarimetric spatially-varying reflectance information, but limited to a single view, or to cover all view directions, but limited to spherical objects made of a single homogeneous material. We present sparse ellipsometry , a portable polarimetric acquisition method that captures both polarimetric SVBRDF and 3D shape simultaneously. Our handheld device consists of off-the-shelf, fixed optical components. Instead of days, the total acquisition time varies between twenty and thirty minutes per object. We develop a complete polarimetric SVBRDF model that includes diffuse and specular components, as well as single scattering, and devise a novel polarimetric inverse rendering algorithm with data augmentation of specular reflection samples via generative modeling. Our results show a strong agreement with a recent ground-truth dataset of captured polarimetric BRDFs of real-world objects. Inseung Hwang, Daniel S. Jeon, Adolfo Muñoz 0001, Diego Gutierrez, Xin Tong 0001, Min H. Kim 0001 |
ACM Trans. Graph. | 5 |
| 2022 | Dual octree graph networks for learning adaptive volumetric shape representationsabstractWe present an adaptive deep representation of volumetric fields of 3D shapes and an efficient approach to learn this deep representation for high-quality 3D shape reconstruction and auto-encoding. Our method encodes the volumetric field of a 3D shape with an adaptive feature volume organized by an octree and applies a compact multilayer perceptron network for mapping the features to the field value at each 3D position. An encoder-decoder network is designed to learn the adaptive feature volume based on the graph convolutions over the dual graph of octree nodes. The core of our network is a new graph convolution operator defined over a regular grid of features fused from irregular neighboring octree nodes at different levels, which not only reduces the computational and memory cost of the convolutions over irregular neighboring octree nodes, but also improves the performance of feature learning. Our method effectively encodes shape details, enables fast 3D shape reconstruction, and exhibits good generality for modeling 3D shapes out of training categories. We evaluate our method on a set of reconstruction tasks of 3D shapes and scenes and validate its superiority over other existing approaches. Our code, data, and trained models are available at https://wang-ps.github.io/dualocnn. Peng-Shuai Wang, Yang Liu 0014, Xin Tong 0001 |
ACM Trans. Graph. | 3 |
| 2022 | VirtualCube: An Immersive 3D Video Communication SystemabstractThe VirtualCube system is a 3D video conference system that attempts to overcome some limitations of conventional technologies. The key ingredient is VirtualCube, an abstract representation of a real-world cubicle instrumented with RGBD cameras for capturing the user's 3D geometry and texture. We design VirtualCube so that the task of data capturing is standardized and significantly simplified, and everything can be built using off-the-shelf hardware. We use VirtualCubes as the basic building blocks of a virtual conferencing environment, and we provide each VirtualCube user with a surrounding display showing life-size videos of remote participants. To achieve real-time rendering of remote participants, we develop the V-Cube View algorithm, which uses multi-view stereo for more accurate depth estimation and Lumi-Net rendering for better rendering quality. The VirtualCube system correctly preserves the mutual eye gaze between participants, allowing them to establish eye contact and be aware of who is visually paying attention to them. The system also allows a participant to have side discussions with remote participants as if they were in the same room. Finally, the system sheds lights on how to support the shared space of work items (e.g., documents and applications) and track participants' visual attention to work items. Jiaolong Yang, Zhen Liu 0031, Ruicheng Wang, Xin Tong 0001, Baining Guo |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2021 | Unsupervised 3D Learning for Shape Analysis via Multiresolution Instance DiscriminationabstractWe propose an unsupervised method for learning a generic and efficient shape encoding network for different shape analysis tasks. Our key idea is to jointly encode and learn shape and point features from unlabeled 3D point clouds. For this purpose, we adapt HRNet to octree-based convolutional neural networks for jointly encoding shape and point features with fused multiresolution subnetworks and design a simple-yet-efficient Multiresolution Instance Discrimination (MID) loss for jointly learning the shape and point features. Our network takes a 3D point cloud as input and output both shape and point features. After training, Our network is concatenated with simple task-specific back-ends and fine-tuned for different shape analysis tasks. We evaluate the efficacy and generality of our method with a set of shape analysis tasks, including shape classification, semantic shape segmentation, as well as shape registration tasks. With simple back-ends, our network demonstrates the best performance among all unsupervised methods and achieves competitive performance to supervised methods. For fine-grained shape segmentation on the PartNet dataset, our method even surpasses existing supervised methods by a large margin. Peng-Shuai Wang, Qianfang Zou 0001, Zhirong Wu, Yang Liu 0014, Xin Tong 0001 |
AAAI | 6 |
| 2021 | Learning Texture Generators for 3D Shape Collections from Internet Photo Sets
Yue Dong 0001, Pieter Peers, Xin Tong 0001 |
BMVC | 4 |
| 2021 | Deformed Implicit Field: Modeling 3D Shapes With Learned Dense CorrespondenceabstractWe propose a novel Deformed Implicit Field (DIF) representation for modeling 3D shapes of a category and generating dense correspondences among shapes. With DIF, a 3D shape is represented by a template implicit field shared across the category, together with a 3D deformation field and a correction field dedicated for each shape instance. Shape correspondences can be easily established using their deformation fields. Our neural network, dubbed DIFNet, jointly learns a shape latent space and these fields for 3D objects belonging to a category without using any correspondence or part label. The learned DIF-Net can also provides reliable correspondence uncertainty measurement reflecting shape structure discrepancy. Experiments show that DIF-Net not only produces high-fidelity 3D shapes but also builds high-quality dense correspondences across different shapes. We also demonstrate several applications such as texture transfer and shape editing, where our method achieves compelling results that cannot be achieved by previous methods.1 Yu Deng 0006, Jiaolong Yang, Xin Tong 0001 |
CVPR | 3 |
| 2021 | Deep Implicit Moving Least-Squares Functions for 3D ReconstructionabstractPoint set is a flexible and lightweight representation widely used for 3D deep learning. However, their discrete nature prevents them from representing continuous and fine geometry, posing a major issue for learning-based shape generation. In this work, we turn the discrete point sets into smooth surfaces by introducing the well-known implicit moving least-squares (IMLS) surface formulation, which naturally defines locally implicit functions on point sets. We incorporate IMLS surface generation into deep neural networks for inheriting both the flexibility of point sets and the high quality of implicit surfaces. Our IMLSNet predicts an octree structure as a scaffold for generating MLS points where needed and characterizes shape geometry with learned local priors. Furthermore, our implicit function evaluation is independent of the neural network once the MLS points are predicted, thus enabling fast runtime evaluation. Our experiments on 3D object reconstruction demonstrate that IMLSNets outperform state-of-the-art learning-based methods in terms of reconstruction quality and computational efficiency. Extensive ablation tests also validate our network design and loss functions. Hao-Xiang Guo 0001, Hao Pan 0001, Peng-Shuai Wang, Xin Tong 0001, Yang Liu 0014 |
CVPR | 5 |
| 2021 | Learning High-Fidelity Face Texture Completion without Complete Face TextureabstractFor face texture completion, previous methods typically use some complete textures captured by multiview imaging systems or 3D scanners for supervised learning. This paper deals with a new challenging problem - learning to complete invisible texture in a single face image without using any complete texture. We simply leverage a large corpus of face images of different subjects (e. g., FFHQ) to train a texture completion model in an unsupervised manner. To achieve this, we propose DSD-GAN, a novel deep neural network based method that applies two discriminators in UV map space and image space. These two discriminators work in a complementary manner to learn both facial structures and texture details. We show that their combination is essential to obtain high-fidelity results. Despite the network never sees any complete facial appearance, it is able to generate compelling full textures from single images. Jongyoo Kim, Jiaolong Yang, Xin Tong 0001 |
ICCV | 3 |
| 2021 | Group-Free 3D Object Detection via TransformersabstractRecently, directly detecting 3D objects from 3D point clouds has received increasing attention. To extract object representation from an irregular point cloud, existing methods usually take a point grouping step to assign the points to an object candidate so that a PointNet-like network could be used to derive object features from the grouped points. However, the inaccurate point assignments caused by the hand-crafted grouping scheme decrease the performance of 3D object detection.In this paper, we present a simple yet effective method for directly detecting 3D objects from the 3D point cloud. Instead of grouping local points to each object candidate, our method computes the feature of an object from all the points in the point cloud with the help of an attention mechanism in the Transformers [42], where the contribution of each point is automatically learned in the network training. With an improved attention stacking scheme, our method fuses object features in different stages and generates more accurate object detection results. With few bells and whistles, the proposed method achieves state-of-the-art 3D object detection performance on two widely used benchmarks, Scan-Net V2 and SUN RGB-D. The code and models are publicly available at https://github.com/zeliu98/Group-Free-3D Zheng Zhang 0022, Yue Cao 0001, Han Hu 0001, Xin Tong 0001 |
ICCV | 5 |
| 2021 | High-Resolution Optical Flow from 1D Attention and CorrelationabstractOptical flow is inherently a 2D search problem, and thus the computational complexity grows quadratically with respect to the search window, making large displacements matching infeasible for high-resolution images. In this paper, we take inspiration from Transformers and propose a new method for high-resolution optical flow estimation with significantly less computation. Specifically, a 1D attention operation is first applied in the vertical direction of the target image, and then a simple 1D correlation in the horizontal direction of the attended image is able to achieve 2D correspondence modeling effect. The directions of attention and correlation can also be exchanged, resulting in two 3D cost volumes that are concatenated for optical flow estimation. The novel 1D formulation empowers our method to scale to very high-resolution input images while maintaining competitive performance. Extensive experiments on Sintel, KITTI and real-world 4K (2160 × 3840) resolution images demonstrated the effectiveness and superiority of our proposed method. Code and models are available at https://github.com/haofeixu/flow1d. Haofei Xu, Jiaolong Yang, Jianfei Cai 0001, Juyong Zhang, Xin Tong 0001 |
ICCV | 5 |
| 2021 | Indoor Scene Generation from a Collection of Semantic-Segmented Depth ImagesabstractWe present a method for creating 3D indoor scenes with a generative model learned from a collection of semantic-segmented depth images captured from different unknown scenes. Given a room with a specified size, our method automatically generates 3D objects in a room from a randomly sampled latent code. Different from existing methods that represent an indoor scene with the type, location, and other properties of objects in the room and learn the scene layout from a collection of complete 3D indoor scenes, our method models each indoor scene as a 3D semantic scene volume and learns a volumetric generative adversarial network (GAN) from a collection of 2.5D partial observations of 3D scenes. To this end, we apply a differentiable projection layer to project the generated 3D semantic scene volumes into semantic-segmented depth images and design a new multiple-view discriminator for learning the complete 3D scene volume from 2.5D semantic-segmented depth images. Compared to existing methods, our method not only efficiently reduces the workload of modeling and acquiring 3D scenes for training, but also produces better object shapes and their detailed layouts in the scene. We evaluate our method with different indoor scene datasets and demonstrate the advantages of our method. We also extend our method for generating 3D indoor scenes from semantic-segmented depth images inferred from RGB images of real scenes.1 Mingjia Yang, Yuxiao Guo 0001, Xin Tong 0001 |
ICCV | 4 |
| 2021 | Spline Positional Encoding for Learning 3D Implicit Signed Distance FieldsabstractMultilayer perceptrons (MLPs) have been successfully used to represent 3D shapes implicitly and compactly, by mapping 3D coordinates to the corresponding signed distance values or occupancy values. In this paper, we propose a novel positional encoding scheme, called Spline Positional Encoding, to map the input coordinates to a high dimensional space before passing them to MLPs, which help recover 3D signed distance fields with fine-scale geometric details from unorganized 3D point clouds. We verified the superiority of our approach over other positional encoding schemes on tasks of 3D shape reconstruction and 3D shape space learning from input point clouds. The efficacy of our approach extended to image reconstruction is also demonstrated and evaluated. Peng-Shuai Wang, Yang Liu 0014, Xin Tong 0001 |
IJCAI | 4 |
| 2021 | Learning and Exploring Motor Skills with Spacetime BoundsabstractAbstract Equipping characters with diverse motor skills is the current bottleneck of physics‐based character animation. We propose a Deep Reinforcement Learning (DRL) framework that enables physics‐based characters to learn and explore motor skills from reference motions. The key insight is to use loose space‐time constraints, termed spacetime bounds, to limit the search space in an early termination fashion. As we only rely on the reference to specify loose spacetime bounds, our learning is more robust with respect to low quality references. Moreover, spacetime bounds are hard constraints that improve learning of challenging motion segments, which can be ignored by imitation‐only learning. We compare our method with state‐of‐the‐art tracking‐based DRL methods. We also show how to guide style exploration within the proposed framework. Li-Ke Ma, Zeshi Yang, Xin Tong 0001, Baining Guo, KangKang Yin |
Comput. Graph. Forum | 3 |
| 2021 | StyleCariGAN: caricature generation via StyleGAN feature map modulationabstractWe present a caricature generation framework based on shape and style manipulation using StyleGAN. Our framework, dubbed StyleCariGAN , automatically creates a realistic and detailed caricature from an input photo with optional controls on shape exaggeration degree and color stylization type. The key component of our method is shape exaggeration blocks that are used for modulating coarse layer feature maps of StyleGAN to produce desirable caricature shape exaggerations. We first build a layer-mixed StyleGAN for photo-to-caricature style conversion by swapping fine layers of the StyleGAN for photos to the corresponding layers of the StyleGAN trained to generate caricatures. Given an input photo, the layer-mixed model produces detailed color stylization for a caricature but without shape exaggerations. We then append shape exaggeration blocks to the coarse layers of the layer-mixed model and train the blocks to create shape exaggerations while preserving the characteristic appearances of the input. Experimental results show that our StyleCariGAN generates realistic and detailed caricatures compared to the current state-of-the-art methods. We demonstrate StyleCariGAN also supports other StyleGAN-based image manipulations, such as facial expression control. Wonjong Jang, Gwangjin Ju, Yucheol Jung, Jiaolong Yang, Xin Tong 0001, Seungyong Lee 0001 |
ACM Trans. Graph. | 5 |
| 2021 | Data-Driven 3D Neck Modeling and AnimationabstractIn this article, we present a data-driven approach for modeling and animation of 3D necks. Our method is based on a new neck animation model that decomposes the neck animation into local deformation caused by larynx motion and global deformation driven by head poses, facial expressions, and speech. A skinning model is introduced for modeling local deformation and underlying larynx motions, while the global neck deformation caused by each factor is modeled by its corrective blendshape set, respectively. Based on this neck model, we introduce a regression method to drive the larynx motion and neck deformation from speech. Both the neck model and the speech regressor are learned from a dataset of 3D neck animation sequences captured from different identities. Our neck model significantly improves the realism of facial animation and allows users to easily create plausible neck animations from speech and facial expressions. We verify our neck model and demonstrate its advantages in 3D neck tracking and animation. Chengwei Zheng, Feng Xu 0005, Xin Tong 0001, Baining Guo |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2020 | Disentangled and Controllable Face Image Generation via 3D Imitative-Contrastive LearningabstractWe propose an approach for face image generation of virtual people with disentangled, precisely-controllable latent representations for identity of non-existing people, expression, pose, and illumination. We embed 3D priors into adversarial learning and train the network to imitate the image formation of an analytic 3D face deformation and rendering process. To deal with the generation freedom induced by the domain gap between real and rendered faces, we further introduce contrastive learning to promote disentanglement by comparing pairs of generated images. Experiments show that through our imitative-contrastive learning, the factor variations are very well disentangled and the properties of a generated face can be precisely controlled. We also analyze the learned latent space and present several meaningful properties supporting factor disentanglement. Our method can also be used to embed real images into the disentangled latent space. We hope our method could provide new understandings of the relationship between physical properties and deep image synthesis. Yu Deng 0006, Jiaolong Yang, Dong Chen 0003, Fang Wen 0001, Xin Tong 0001 |
CVPR | 5 |
| 2020 | TextureFusion: High-Quality Texture Acquisition for Real-Time RGB-D ScanningabstractReal-time RGB-D scanning technique has become widely used to progressively scan objects with a hand-held sensor. Existing online methods restore color information per voxel, and thus their quality is often limited by the tradeoff between spatial resolution and time performance. Also, such methods often suffer from blurred artifacts in the captured texture. Traditional offline texture mapping methods with non-rigid warping assume that the reconstructed geometry and all input views are obtained in advance, and the optimization takes a long time to compute mesh parameterization and warp parameters, which prevents them from being used in real-time applications. In this work, we propose a progressive texture-fusion method specially designed for real-time RGB-D scanning. To this end, we first devise a novel texture-tile voxel grid, where texture tiles are embedded in the voxel grid of the signed distance function, allowing for high-resolution texture mapping on the low-resolution geometry volume. Instead of using expensive mesh parameterization, we associate vertices of implicit geometry directly with texture coordinates. Second, we introduce real-time texture warping that applies a spatially-varying perspective mapping to input images so that texture warping efficiently mitigates the mismatch between the intermediate geometry and the current input view. It allows us to enhance the quality of texture over time while updating the geometry in real-time. The results demonstrate that the quality of our real-time texture mapping is highly competitive to that of exhaustive offline texture warping methods. Our method is also capable of being integrated into existing RGB-D scanning frameworks. Joo Ho Lee 0003, Hyunho Ha, Yue Dong 0001, Xin Tong 0001, Min H. Kim 0001 |
CVPR | 4 |
| 2020 | Deep 3D Portrait From a Single ImageabstractIn this paper, we present a learning-based approach for recovering the 3D geometry of human head from a single portrait image. Our method is learned in an unsupervised manner without any ground-truth 3D data. We represent the head geometry with a parametric 3D face model together with a depth map for other head regions including hair and ear. A two-step geometry learning scheme is proposed to learn 3D head reconstruction from in-the-wild face images, where we first learn face shape on single images using self-reconstruction and then learn hair and ear geometry using pairs of images in a stereo-matching fashion. The second step is based on the output of the first to not only improve the accuracy but also ensure the consistency of overall head geometry. We evaluate the accuracy of our method both in 3D and with pose manipulation tasks on 2D images. We alter pose based on the recovered geometry and apply a refinement network trained with adversarial learning to ameliorate the reprojected images and translate them to the real image domain. Extensive evaluations and comparison with previous methods show that our new method can produce high-fidelity 3D head geometry and head pose manipulation results. Sicheng Xu, Jiaolong Yang, Dong Chen 0003, Fang Wen 0001, Yu Deng 0006, Yunde Jia, Xin Tong 0001 |
CVPR | 7 |
| 2020 | PFCNN: Convolutional Neural Networks on 3D Surfaces Using Parallel FramesabstractSurface meshes are widely used shape representations and capture finer geometry data than point clouds or volumetric grids, but are challenging to apply CNNs directly due to their non-Euclidean structure. We use parallel frames on surface to define PFCNNs that enable effective feature learning on surface meshes by mimicking standard convolutions faithfully. In particular, the convolution of PFCNN not only maps local surface patches onto flat tangent planes, but also aligns the tangent planes such that they locally form a flat Euclidean structure, thus enabling recovery of standard convolutions. The alignment is achieved by the tool of locally flat connections borrowed from discrete differential geometry, which can be efficiently encoded and computed by parallel frame fields. In addition, the lack of canonical axis on surface is handled by sampling with the frame directions. Experiments show that for tasks including classification, segmentation and registration on deformable geometric domains, as well as semantic scene segmentation on rigid domains, PFCNNs achieve robust and superior performances without using sophisticated input features than state-of-the-art surface based CNNs. Hao Pan 0001, Yang Liu 0014, Xin Tong 0001 |
CVPR | 5 |
| 2020 | A Closer Look at Local Aggregation Operators in Point Cloud Analysis
Han Hu 0001, Yue Cao 0001, Zheng Zhang 0022, Xin Tong 0001 |
ECCV (23) | 5 |
| 2020 | Object-Based Illumination Estimation with Rendering-Aware Neural Networks
Yue Dong 0001, Stephen Lin 0001, Xin Tong 0001 |
ECCV (15) | 5 |
| 2020 | H Space: Interactive Augmented Reality ArtabstractThis artwork exploits recent research into augmented reality systems, such as the HoloLens, for building creative interaction in augmented reality. The work is being conducted in the context of interactive art experiences. The first version of the audience experience of the artwork, "H Space", was informally tested in the SIGGRAPH 2018 Art Gallery context. Experiences with a later, improved, version was evaluated at Tsinghua University. The latest distributed version will be shown in Sydney. The paper describes the concept, the background in both the art and the technological domain and points to some of the key computer human interaction art research issues that the work highlights. Ernest A. Edmonds, Damian Hills, Xin Tong 0001 |
TEI | 4 |
| 2020 | Foreword to the special section on the international conference on computer-aided design and computer graphics (CAD/Graphics) 2019
Xin Tong 0001, Karol Myszkowski, Jin Huang 0001 |
Comput. Graph. | 1 |
| 2020 | RAS: A Data-Driven Rigidity-Aware Skinning Model For 3D Facial AnimationabstractAbstract We present a novel data‐driven skinning model—rigidity‐aware skinning (RAS) model, for simulating both active and passive 3D facial animation of different identities in real time. Our model builds upon a linear blend skinning (LBS) scheme, where the bone set and skinning weights are shared for diverse identities and learned from the data via a sparse and localized skinning decomposition algorithm. Our model characterizes the animated face into the active expression and the passive deformation: The former is represented by an LBS‐based multi‐linear model learned from the FaceWareHouse data set, and the latter is represented by a spatially varying as‐rigid‐as‐possible deformation applied to the LBS‐based multi‐linear model, whose rigidity parameters are learned from the data by a novel rigidity estimation algorithm. Our RAS model is not only generic and expressive for faithfully modelling medium‐scale facial deformation, but also compact and lightweight for generating vivid facial animation in real time. We validate the efficiency and effectiveness of our RAS model for real‐time 3D facial animation and expression editing. Yang Liu 0014, L-F. Dong, Xin Tong 0001 |
Comput. Graph. Forum | 4 |
| 2020 | Real-time hair simulation with heptadiagonal decomposition on mass spring system
Jianwei Jiang 0002, Bin Sheng 0001, Ping Li 0016, Lizhuang Ma, Xin Tong 0001, Enhua Wu |
Graph. Model. | 5 |
| 2020 | Image-based acquisition and modeling of polarimetric reflectanceabstractRealistic modeling of the bidirectional reflectance distribution function (BRDF) of scene objects is a vital prerequisite for any type of physically based rendering. In the last decades, the availability of databases containing real-world material measurements has fueled considerable innovation in the development of such models. However, previous work in this area was mainly focused on increasing the visual realism of images, and hence ignored the effect of scattering on the polarization state of light, which is normally imperceptible to the human eye. Existing databases thus only capture scattered flux, or polarimetric BRDF datasets are too directionally sparse (e.g., in-plane) to be usable for simulation. While subtle to human observers, polarization is easily perceived by any optical sensor (e.g., using polarizing filters), providing a wealth of additional information about shape and material properties of the object under observation. Given the increasing application of rendering in the solution of inverse problems via analysis-by-synthesis and differentiation, the ability to realistically model polarized radiative transport is thus highly desirable. Polarization depends on the wavelength of the spectrum, and thus we provide the first polarimetric BRDF (pBRDF) dataset that captures the polarimetric properties of real-world materials over the full angular domain, and at multiple wavelengths. Acquisition of such reflectance data is challenging due to the extremely large space of angular, spectral, and polarimetric configurations that must be observed, and we propose a scheme combining image-based acquisition with spectroscopic ellipsometry to perform measurements in a realistic amount of time. This process yields raw Mueller matrices, which we subsequently transform into Rusinkiewicz-parameterized pBRDFs that can be used for rendering. Our dataset provides 25 isotropic pBRDFs spanning a wide range of appearances: diffuse/specular, metallic/dielectric, rough/smooth, and different color albedos, captured in five wavelength ranges covering the visible spectrum. We demonstrate usage of our data-driven pBRDF model in a physically based renderer that accounts for polarized interreflection, and we investigate the relationship of polarization and material appearance, providing insights into the behavior of characteristic real-world pBRDFs. Seung-Hwan Baek, Tizian Zeltner, Hyunjin Ku, Inseung Hwang, Xin Tong 0001, Wenzel Jakob, Min H. Kim 0001 |
ACM Trans. Graph. | 5 |
| 2020 | Deferred neural lighting: free-viewpoint relighting from unstructured photographsabstractWe present deferred neural lighting, a novel method for free-viewpoint relighting from unstructured photographs of a scene captured with handheld devices. Our method leverages a scene-dependent neural rendering network for relighting a rough geometric proxy with learnable neural textures. Key to making the rendering network lighting aware are radiance cues: global illumination renderings of a rough proxy geometry of the scene for a small set of basis materials and lit by the target lighting. As such, the light transport through the scene is never explicitely modeled, but resolved at rendering time by a neural rendering network. We demonstrate that the neural textures and neural renderer can be trained end-to-end from unstructured photographs captured with a double hand-held camera setup that concurrently captures the scene while being lit by only one of the cameras' flash lights. In addition, we propose a novel augmentation refinement strategy that exploits the linearity of light transport to extend the relighting capabilities of the neural rendering network to support other lighting types (e.g., environment lighting) beyond the lighting used during acquisition (i.e., flash lighting). We demonstrate our deferred neural lighting solution on a variety of real-world and synthetic scenes exhibiting a wide range of material properties, light transport effects, and geometrical complexity. Duan Gao, Yue Dong 0001, Pieter Peers, Kun Xu 0003, Xin Tong 0001 |
ACM Trans. Graph. | 6 |
| 2019 | Deep Single-View 3D Object Reconstruction with Visual Hull Embeddingabstract3D object reconstruction is a fundamental task of many robotics and AI problems. With the aid of deep convolutional neural networks (CNNs), 3D object reconstruction has witnessed a significant progress in recent years. However, possibly due to the prohibitively high dimension of the 3D object space, the results from deep CNNs are often prone to missing some shape details. In this paper, we present an approach which aims to preserve more shape details and improve the reconstruction quality. The key idea of our method is to leverage object mask and pose estimation from CNNs to assist the 3D shape learning by constructing a probabilistic singleview visual hull inside of the network. Our method works by first predicting a coarse shape as well as the object pose and silhouette using CNNs, followed by a novel 3D refinement CNN which refines the coarse shapes using the constructed probabilistic visual hulls. Experiment on both synthetic data and real images show that embedding a single-view visual hull for shape refinement can significantly improve the reconstruction quality by recovering more shapes details and improving shape consistency with the input image. Hanqing Wang 0001, Jiaolong Yang, Wei Liang 0008, Xin Tong 0001 |
AAAI | 4 |
| 2019 | Capturing Piecewise SVBRDFs with Content Aware Lighting
Xiao Li 0030, Peiran Ren, Yue Dong 0001, Gang Hua 0001, Xin Tong 0001, Baining Guo |
CGI | 5 |
| 2019 | Synthesizing 3D Shapes From Silhouette Image Collections Using Multi-Projection Generative Adversarial NetworksabstractWe present a new weakly supervised learning-based method for generating novel category-specific 3D shapes from unoccluded image collections. Our method is weakly supervised and only requires silhouette annotations from unoccluded, category-specific objects. Our method does not require access to the object’s 3D shape, multiple observations per object from different views, intra-image pixel correspondences, or any view annotations. Key to our method is a novel multi-projection generative adversarial network (MP-GAN) that trains a 3D shape generator to be consistent with multiple 2D projections of the 3D shapes, and without direct access to these 3D shapes. This is achieved through multiple discriminators that encode the distribution of 2D projections of the 3D shapes seen from a different views. Additionally, to determine the view information for each silhouette image, we also train a view prediction network on visualizations of 3D shapes synthesized by the generator. We iteratively alternate between training the generator and training the view prediction network. We validate our multi-projection GAN on both synthetic and real image datasets. Furthermore, we also show that multi-projection GANs can aid in learning other high-dimensional distributions from lower dimensional training datasets, such as material-class specific spatially varying reflectance properties from images. Xiao Li 0030, Yue Dong 0001, Pieter Peers, Xin Tong 0001 |
CVPR | 4 |
| 2019 | A Skeleton-Bridged Deep Learning Approach for Generating Meshes of Complex Topologies From Single RGB ImagesabstractThis paper focuses on the challenging task of learning 3D object surface reconstructions from single RGB images. Existing methods achieve varying degrees of success by using different geometric representations. However, they all have their own drawbacks, and cannot well reconstruct those surfaces of complex topologies. To this end, we propose in this paper a skeleton-bridged, stage-wise learning approach to address the challenge. Our use of skeleton is due to its nice property of topology preservation, while being of lower complexity to learn. To learn skeleton from an input image, we design a deep architecture whose decoder is based on a novel design of parallel streams respectively for synthesis of curve- and surface-like skeleton points. We use different shape representations of point cloud, volume, and mesh in our stage-wise learning, in order to take their respective advantages. We also propose multi-stage use of the input image to correct prediction errors that are possibly accumulated in each stage. We conduct intensive experiments to investigate the efficacy of our proposed approach. Qualitative and quantitative results on representative object categories of both simple and complex topologies demonstrate the superiority of our approach over existing ones. We will make our ShapeNet-Skeleton dataset publicly available. Jiapeng Tang, Xiaoguang Han 0001, Junyi Pan, Kui Jia, Xin Tong 0001 |
CVPR | 5 |
| 2019 | Face Video Deblurring Using 3D Facial PriorsabstractExisting face deblurring methods only consider single frames and do not account for facial structure and identity information. These methods struggle to deblur face videos that exhibit significant pose variations and misalignment. In this paper we propose a novel face video deblurring network capitalizing on 3D facial priors. The model consists of two main branches: i) a face video deblurring sub-network based on an encoder-decoder architecture, and ii) a 3D face reconstruction and rendering branch for predicting 3D priors of salient facial structures and identity knowledge. These structures encourage the deblurring branch to generate sharp faces with detailed structures. Our method not only uses low-level information (i.e., image intensity), but also middle-level information (i.e., 3D facial structure) and high-level knowledge (i.e., identity content) to further explore spatial constraints of facial components from blurry face frames. Extensive experimental results demonstrate that the proposed algorithm performs favorably against the state-of-the-art methods. Wenqi Ren, Jiaolong Yang, Senyou Deng, David P. Wipf, Xiaochun Cao, Xin Tong 0001 |
ICCV | 6 |
| 2019 | Feature preserving GAN and multi-scale feature enhancement for domain adaption person Re-identification
Xiuping Liu, Hongchen Tan, Xin Tong 0001, Junjie Cao 0001, Jun Zhou 0023 |
Neurocomputing | 3 |
| 2019 | Deep inverse rendering for high-resolution SVBRDF estimation from an arbitrary number of imagesabstractIn this paper we present a unified deep inverse rendering framework for estimating the spatially-varying appearance properties of a planar exemplar from an arbitrary number of input photographs, ranging from just a single photograph to many photographs. The precision of the estimated appearance scales from plausible when the input photographs fails to capture all the reflectance information, to accurate for large input sets. A key distinguishing feature of our framework is that it directly optimizes for the appearance parameters in a latent embedded space of spatially-varying appearance, such that no handcrafted heuristics are needed to regularize the optimization. This latent embedding is learned through a fully convolutional auto-encoder that has been designed to regularize the optimization. Our framework not only supports an arbitrary number of input photographs, but also at high resolution. We demonstrate and evaluate our deep inverse rendering solution on a wide variety of publicly available datasets. Duan Gao, Xiao Li 0030, Yue Dong 0001, Pieter Peers, Kun Xu 0003, Xin Tong 0001 |
ACM Trans. Graph. | 6 |
| 2019 | A scalable galerkin multigrid method for real-time simulation of deformable objectsabstractWe propose a simple yet efficient multigrid scheme to simulate high-resolution deformable objects in their full spaces at interactive frame rates. The point of departure of our method is the Galerkin projection which is simple to construct. However, a naïve Galerkin multigrid does not scale well for large and irregular grids because it trades-off matrix sparsity for smaller sized linear systems which eventually stops improving the performance. Given that observation, we design our special projection criterion which is based on skinning space coordinates with piecewise constant weights, to make our Galerkin multigrid method scale for high-resolution meshes without suffering from dense linear solves. The usage of skinning space coordinates enables us to reduce the resolution of grids more aggressively, and our piecewise constant weights further ensure us to always deal with reasonably-sparse linear solves. Our projection matrices also help us to manage multi-level linear systems efficiently. Therefore, our method can be applied to different optimization schemes such as Newton's method and Projective Dynamics, pushing the resolution of a real-time simulation to orders of magnitudes higher. Our final GPU implementation outperforms the other state-of-the-art GPU deformable body simulators, enabling us to simulate large deformable objects with hundred thousands of degrees of freedom in real-time. Zangyueyang Xian, Xin Tong 0001, Tiantian Liu 0002 |
ACM Trans. Graph. | 2 |
| 2019 | Point sets joint registration and co-segmentation
Siyu Hu, Xuejin Chen, Xin Tong 0001 |
Vis. Comput. | 3 |
| 2018 | HairNet: Single-View Hair Reconstruction Using Convolutional Neural Networks
Yi Zhou 0023, Liwen Hu 0001, Jun Xing, Weikai Chen 0001, Han-Wei Kung, Xin Tong 0001, Hao Li 0015 |
ECCV (11) | 6 |
| 2018 | View-Volume Network for Semantic Scene Completion from a Single Depth ImageabstractWe introduce a View-Volume convolutional neural network (VVNet) for inferring the occupancy and semantic labels of a volumetric 3D scene from a single depth image. Our method extracts the detailed geometric features from the input depth image with a 2D view CNN and then projects the features into a 3D volume according to the input depth map via a projection layer. After that, we learn the 3D context information of the scene with a 3D volume CNN for computing the result volumetric occupancy and semantic labels. With combined 2D and 3D representations, the VVNet efficiently reduces the computational cost, enables feature extraction from multi-channel high resolution inputs, and thus significantly improve the result accuracy. We validate our method and demonstrate its efficiency and effectiveness on both synthetic SUNCG and real NYU dataset. Yuxiao Guo 0001, Xin Tong 0001 |
IJCAI | 2 |
| 2018 | Single Image Surface Appearance Modeling with Self-augmented CNNs and Inexact SupervisionabstractAbstract This paper presents a deep learning based method for estimating the spatially varying surface reflectance properties from a single image of a planar surface under unknown natural lighting trained using only photographs of exemplar materials without referencing any artist generated or densely measured spatially varying surface reflectance training data. Our method is based on an empirical study of Li et al.'s [ LDPT17 ] self‐augmentation training strategy that shows that the main role of the initial approximative network is to provide guidance on the inherent ambiguities in single image appearance estimation. Furthermore, our study indicates that this initial network can be inexact (i.e., trained from other data sources) as long as it resolves the inherent ambiguities. We show that the single image estimation network trained without manually labeled data outperforms prior work in terms of accuracy as well as generality. Xiao Li 0030, Yue Dong 0001, Pieter Peers, Xin Tong 0001 |
Comput. Graph. Forum | 5 |
| 2018 | Simultaneous acquisition of polarimetric SVBRDF and normalsabstractCapturing appearance often requires dense sampling in light-view space, which is often achieved in specialized, expensive hardware setups. With the aim of realizing a compact acquisition setup without multiple angular samples of light and view, we sought to leverage an alternative optical property of light, polarization. To this end, we capture a set of polarimetric images with linear polarizers in front of a single projector and camera to obtain the appearance and normals of real-world objects. We encountered two technical challenges: First, no complete polarimetric BRDF model is available for modeling mixed polarization of both specular and diffuse reflection. Second, existing polarization-based inverse rendering methods are not applicable to a single local illumination setup since they are formulated with the assumption of spherical illumination. To this end, we first present a complete polarimetric BRDF (pBRDF) model that can define mixed polarization of both specular and diffuse reflection. Second, by leveraging our pBRDF model, we propose a novel inverse-rendering method with joint optimization of pBRDF and normals to capture spatially-varying material appearance: per-material specular properties (including the refractive index, specular roughness and specular coefficient), per-pixel diffuse albedo and normals. Our method can solve the severely ill-posed inverse-rendering problem by carefully accounting for the physical relationship between polarimetric appearance and geometric properties. We demonstrate how our method overcomes limited sampling in light-view space for inverse rendering by means of polarization. Seung-Hwan Baek, Daniel S. Jeon, Xin Tong 0001, Min H. Kim 0001 |
ACM Trans. Graph. | 3 |
| 2018 | Image smoothing via unsupervised learningabstractImage smoothing represents a fundamental component of many disparate computer vision and graphics applications. In this paper, we present a unified unsupervised (label-free) learning framework that facilitates generating flexible and high-quality smoothing effects by directly learning from data using deep convolutional neural networks (CNNs). The heart of the design is the training signal as a novel energy function that includes an edge-preserving regularizer which helps maintain important yet potentially vulnerable image structures, and a spatially-adaptive L p flattening criterion which imposes different forms of regularization onto different image regions for better smoothing quality. We implement a diverse set of image smoothing solutions employing the unified framework targeting various applications such as, image abstraction, pencil sketching, detail enhancement, texture removal and content-aware image manipulation, and obtain results comparable with or better than previous methods. Moreover, our method is extremely fast with a modern GPU (e.g, 200 fps for 1280×720 images). Qingnan Fan, Jiaolong Yang, David P. Wipf, Baoquan Chen, Xin Tong 0001 |
ACM Trans. Graph. | 5 |
| 2018 | Appearance Modeling via Proxy-to-Image AlignmentabstractEndowing 3D objects with realistic surface appearance is a challenging and time-demanding task, as real-world surfaces typically exhibit a plethora of spatially variant geometric and photometric detail. Not surprisingly, computer artists commonly use images of real-world objects as an inspiration and a reference for their digital creations. However, despite two decades of research on image-based modeling, there are still no tools available for automatically extracting the detailed appearance (microgeometry and texture) of a 3D surface from a single image. In this article, we present a novel user-assisted approach for quickly and easily extracting a nonparametric appearance model from a single photograph of a reference object. The extraction process requires a user-provided proxy, whose geometry roughly approximates that of the object in the image. Since the proxy is just a rough approximation, it is necessary to align and deform it so as to match the reference object. The main contribution of this work is a novel technique to perform such an alignment, which enables accurate joint recovery of geometric detail and reflectance. The correlations between the recovered geometry at various scales and the spatially varying reflectance constitute a nonparametric appearance model. Once extracted, the appearance model may then be applied to various 3D shapes, whose large-scale geometry may differ considerably from that of the original reference object. Thus, our approach makes it possible to construct an appearance library, allowing users to easily enrich detail-less 3D shapes with realistic geometric detail and surface texture. Hui Huang 0004, Ke Xie 0001, Dani Lischinski, Minglun Gong, Xin Tong 0001, Daniel Cohen-Or |
ACM Trans. Graph. | 6 |
| 2018 | Robust flow-guided neural prediction for sketch-based freeform surface modelingabstractSketching provides an intuitive user interface for communicating free form shapes. While human observers can easily envision the shapes they intend to communicate, replicating this process algorithmically requires resolving numerous ambiguities. Existing sketch-based modeling methods resolve these ambiguities by either relying on expensive user annotations or by restricting the modeled shapes to specific narrow categories. We present an approach for modeling generic freeform 3D surfaces from sparse, expressive 2D sketches that overcomes both limitations by incorporating convolution neural networks (CNN) into the sketch processing workflow. Given a 2D sketch of a 3D surface, we use CNNs to infer the depth and normal maps representing the surface. To combat ambiguity we introduce an intermediate CNN layer that models the dense curvature direction, or flow, field of the surface, and produce an additional output confidence map along with depth and normal. The flow field guides our subsequent surface reconstruction for improved regularity; the confidence map trained unsupervised measures ambiguity and provides a robust estimator for data fitting. To reduce ambiguities in input sketches users can refine their input by providing optional depth values at sparse points and curvature hints for strokes. Our CNN is trained on a large dataset generated by rendering sketches of various 3D shapes using non-photo-realistic line rendering (NPR) method that mimics human sketching of free-form shapes. We use the CNN model to process both single- and multi-view sketches. Using our multi-view framework users progressively complete the shape by sketching in different views, generating complete closed shapes. For each new view, the modeling is assisted by partial sketches and depth cues provided by surfaces generated in earlier views. The partial surfaces are fused into a complete shape using predicted confidence levels as weights. We validate our approach, compare it with previous methods and alternative structures, and evaluate its performance with various modeling tasks. The results demonstrate our method is a new approach for efficiently modeling freeform shapes with succinct but expressive 2D sketches. Changjian Li 0001, Hao Pan 0001, Yang Liu 0014, Xin Tong 0001, Alla Sheffer, Wenping Wang 0001 |
ACM Trans. Graph. | 4 |
| 2018 | Language-driven synthesis of 3D scenes from scene databasesabstractWe introduce a novel framework for using natural language to generate and edit 3D indoor scenes, harnessing scene semantics and text-scene grounding knowledge learned from large annotated 3D scene databases. The advantage of natural language editing interfaces is strongest when performing semantic operations at the sub-scene level, acting on groups of objects. We learn how to manipulate these sub-scenes by analyzing existing 3D scenes. We perform edits by first parsing a natural language command from the user and transforming it into a semantic scene graph that is used to retrieve corresponding sub-scenes from the databases that match the command. We then augment this retrieved sub-scene by incorporating other objects that may be implied by the scene context. Finally, a new 3D scene is synthesized by aligning the augmented sub-scene with the user's current scene, where new objects are spliced into the environment, possibly triggering appropriate adjustments to the existing scene arrangement. A suggestive modeling interface with multiple interpretations of user commands is used to alleviate ambiguities in natural language. We conduct studies comparing our approach against both prior text-to-scene work and artist-made scenes and find that our method significantly outperforms prior work and is comparable to handmade scenes even when complex and varied natural sentences are used. Rui Ma 0011, Akshay Gadi Patil, Matthew Fisher, Manyi Li, Sören Pirk, Binh-Son Hua, Sai-Kit Yeung, Xin Tong 0001, Leonidas J. Guibas, Hao (Richard) Zhang |
ACM Trans. Graph. | 8 |
| 2018 | Adaptive O-CNN: a patch-based deep representation of 3D shapesabstractWe present an Adaptive Octree-based Convolutional Neural Network (Adaptive O-CNN) for efficient 3D shape encoding and decoding. Different from volumetric-based or octree-based CNN methods that represent a 3D shape with voxels in the same resolution, our method represents a 3D shape adaptively with octants at different levels and models the 3D shape within each octant with a planar patch. Based on this adaptive patch-based representation, we propose an Adaptive O-CNN encoder and decoder for encoding and decoding 3D shapes. The Adaptive O-CNN encoder takes the planar patch normal and displacement as input and performs 3D convolutions only at the octants at each level, while the Adaptive O-CNN decoder infers the shape occupancy and subdivision status of octants at each level and estimates the best plane normal and displacement for each leaf octant. As a general framework for 3D shape analysis and generation, the Adaptive O-CNN not only reduces the memory and computational cost, but also offers better shape generation capability than the existing 3D-CNN approaches. We validate Adaptive O-CNN in terms of efficiency and effectiveness on different shape analysis and generation tasks, including shape classification, 3D autoencoding, shape prediction from a single image, and shape completion for noisy and incomplete point clouds. Peng-Shuai Wang, Chun-Yu Sun, Yang Liu 0014, Xin Tong 0001 |
ACM Trans. Graph. | 4 |
| 2018 | Outdoor Markerless Motion Capture with Sparse Handheld Video CamerasabstractWe present a method for outdoor markerless motion capture with sparse handheld video cameras. In the simplest setting, it only involves two mobile phone cameras following the character. This setup can maximize the flexibilities of data capture and broaden the applications of motion capture. To solve the character pose under such challenge settings, we exploit the generative motion capture methods and propose a novel model-view consistency that considers both foreground and background in the tracking stage. The background is modeled as a deformable 2D grid, which allows us to compute the background-view consistency for sparse moving cameras. The 3D character pose is tracked with a global-local optimization through minimizing our consistency cost. A novel motion regularizer is also proposed in the optimization to constrain the solution pose space. The whole process of the proposed method is simple as frame by frame video segmentation is not required. Our method outperforms several alternative methods on various examples demonstrated in the paper. Yangang Wang 0001, Yebin Liu, Xin Tong 0001, Qionghai Dai, Ping Tan 0002 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2018 | 3D cartoon face rigging from sparse examples
Jingyong Zhou, Hsiang-Tao Wu, Zicheng Liu 0001, Xin Tong 0001, Baining Guo |
Vis. Comput. | 4 |
| 2017 | Modeling surface appearance from a single photograph using self-augmented convolutional neural networksabstractWe present a convolutional neural network (CNN) based solution for modeling physically plausible spatially varying surface reflectance functions (SVBRDF) from a single photograph of a planar material sample under unknown natural illumination. Gathering a sufficiently large set of labeled training pairs consisting of photographs of SVBRDF samples and corresponding reflectance parameters, is a difficult and arduous process. To reduce the amount of required labeled training data, we propose to leverage the appearance information embedded in unlabeled images of spatially varying materials to self-augment the training process. Starting from an initial approximative network obtained from a small set of labeled training pairs, we estimate provisional model parameters for each unlabeled training exemplar. Given this provisional reflectance estimate, we then synthesize a novel temporarylabeledtraining pair by rendering the exact corresponding image under a new lighting condition. After refining the network using these additional training samples, we re-estimate the provisional model parameters for the unlabeled data and repeat the self-augmentation process until convergence. We demonstrate the efficacy of the proposed network structure on spatially varying wood, metals, and plastics, as well as thoroughly validate the effectiveness of the self-augmentation training process. Xiao Li 0030, Yue Dong 0001, Pieter Peers, Xin Tong 0001 |
ACM Trans. Graph. | 4 |
| 2017 | BendSketch: modeling freeform surfaces through 2D sketchingabstractSketch-based modeling provides a powerful paradigm for geometric modeling. Recent research had shown, sketch based modeling methods are most effective when targeting a specific family of surfaces. A large and growing arsenal of sketching tools is available for different types of geometries and different target user populations. Our work augments this arsenal with a new and powerful tool for modeling complex freeform shapes by sketching sparse 2D strokes; our method complements existing approaches in enabling the generation of surfaces with complex curvature patterns that are challenging to produce with existing methods. To model a desired surface patch with our technique, the user sketches the patch boundary as well as a small number of strokes representing the major bending directions of the shape. Our method uses this input to generate a curvature field that conforms to the user strokes and then uses this field to derive a freeform surface with the desired curvature pattern. To infer the surface from the strokes we first disambiguate the convex versus concave bending directions indicated by the strokes and estimate the surface bending magnitude along the strokes. We subsequently construct a curvature field based on these estimates, using a non-orthogonal 4-direction field coupled with a scalar magnitude field, and finally construct a surface whose curvature pattern reflects this field through an iterative sequence of simple linear optimizations. Our framework is well suited for single-view modeling, but also supports multi-view interaction, necessary to model complex shapes portions of which can be occluded in many views. It effectively combines multi-view inputs to obtain a coherent 3D shape. It runs at interactive speed allowing for immediate user feedback. We demonstrate the effectiveness of the proposed method through a large collection of complex examples created by both artists and amateurs. Our framework provides a useful complement to the existing sketch-based modeling methods. Changjian Li 0001, Hao Pan 0001, Yang Liu 0014, Xin Tong 0001, Alla Sheffer, Wenping Wang 0001 |
ACM Trans. Graph. | 4 |
| 2017 | Computational design and fabrication of soft pneumatic objects with desired deformationsabstractWe present an end-to-end solution for design and fabrication of soft pneumatic objects with desired deformations. Given a 3D object with its rest and deformed target shapes, our method automatically optimizes the chamber structure and material distribution inside the object volume so that the fabricated object can deform to all the target deformed poses with controlled air injection. To this end, our method models the object volume with a set of chambers separated by material shells. Each chamber has individual channels connected to the object surface and thus can be separately controlled with a pneumatic system, while the shell is comprised of base material with an embedded frame structure. A two-step algorithm is developed to compute the geometric layout of the chambers and frame structure as well as the material properties of the frame structure from the input. The design results can be fabricated with 3D printing and deformed by a controlled pneumatic system. We validate and demonstrate the efficacy of our method with soft pneumatic objects that have different shapes and deformation behaviors. Li-Ke Ma, Yizhonc Zhang, Yang Liu 0014, Kun Zhou 0001, Xin Tong 0001 |
ACM Trans. Graph. | 5 |
| 2017 | DeepToF: off-the-shelf real-time correction of multipath interference in time-of-flight imagingabstractTime-of-flight (ToF) imaging has become a widespread technique for depth estimation, allowing affordable off-the-shelf cameras to provide depth maps in real time. However, multipath interference (MPI) resulting from indirect illumination significantly degrades the captured depth. Most previous works have tried to solve this problem by means of complex hardware modifications or costly computations. In this work, we avoid these approaches and propose a new technique to correct errors in depth caused by MPI, which requires no camera modifications and takes just 10 milliseconds per frame. Our observations about the nature of MPI suggest that most of its information is available in image space; this allows us to formulate the depth imaging process as a spatially-varying convolution and use a convolutional neural network to correct MPI errors. Since the input and output data present similar structure, we base our network on an autoencoder, which we train in two stages. First, we use the encoder (convolution filters) to learn a suitable basis to represent MPI-corrupted depth images; then, we train the decoder (deconvolution filters) to correct depth from synthetic scenes, generated by using a physically-based, time-resolved renderer. This approach allows us to tackle a key problem in ToF, the lack of ground-truth data, by using a large-scale captured training set with MPI-corrupted depth to train the encoder, and a smaller synthetic training set with ground truth depth to train the decoder stage of the network. We demonstrate and validate our method on both synthetic and real complex scenarios, using an off-the-shelf ToF camera, and with only the captured, incorrect depth as input. Julio Marco, Quercus Hernandez, Adolfo Muñoz 0001, Yue Dong 0001, Adrián Jarabo, Min H. Kim 0001, Xin Tong 0001, Diego Gutierrez |
ACM Trans. Graph. | 7 |
| 2017 | O-CNN: octree-based convolutional neural networks for 3D shape analysisabstractWe present O-CNN , an Octree-based Convolutional Neural Network (CNN) for 3D shape analysis. Built upon the octree representation of 3D shapes, our method takes the average normal vectors of a 3D model sampled in the finest leaf octants as input and performs 3D CNN operations on the octants occupied by the 3D shape surface. We design a novel octree data structure to efficiently store the octant information and CNN features into the graphics memory and execute the entire O-CNN training and evaluation on the GPU. O-CNN supports various CNN structures and works for 3D shapes in different representations. By restraining the computations on the octants occupied by 3D surfaces, the memory and computational costs of the O-CNN grow quadratically as the depth of the octree increases, which makes the 3D CNN feasible for high-resolution 3D models. We compare the performance of the O-CNN with other existing 3D CNN solutions and demonstrate the efficiency and efficacy of O-CNN in three shape analysis tasks, including object classification, shape retrieval, and shape segmentation. Peng-Shuai Wang, Yang Liu 0014, Yuxiao Guo 0001, Chun-Yu Sun, Xin Tong 0001 |
ACM Trans. Graph. | 5 |
| 2016 | Action-driven 3D indoor scene evolutionabstractWe introduce a framework for action-driven evolution of 3D indoor scenes, where the goal is to simulate how scenes are altered by human actions, and specifically, by object placements necessitated by the actions. To this end, we develop an action model with each type of action combining information about one or more human poses, one or more object categories, and spatial configurations of objects belonging to these categories which summarize the object-object and object-human relations for the action. Importantly, all these pieces of information are learned from annotated photos. Correlations between the learned actions are analyzed to guide the construction of an action graph. Starting with an initial 3D scene, we probabilistically sample a sequence of actions from the action graph to drive progressive scene evolution. Each action triggers appropriate object placements, based on object co-occurrences and spatial configurations learned for the action model. We show results of our scene evolution that lead to realistic and messy 3D scenes, as well as quantitative evaluations by user studies which compare our method to manual scene creation and state-of-the-art, data-driven methods, in terms of scene plausibility and naturalness. Rui Ma 0011, Honghua Li, Changqing Zou, Zicheng Liao, Xin Tong 0001, Hao (Richard) Zhang |
ACM Trans. Graph. | 5 |
| 2016 | Mesh denoising via cascaded normal regressionabstractWe present a data-driven approach for mesh denoising. Our key idea is to formulate the denoising process with cascaded non-linear regression functions and learn them from a set of noisy meshes and their ground-truth counterparts. Each regression function infers the normal of a denoised output mesh facet from geometry features extracted from its neighborhood facets on the input mesh and sends the result as the input of the next regression function. Specifically, we develop afiltered facet normal descriptor (FND)for modeling the geometry features around each facet on the noisy mesh and model a regression function with neural networks for mapping the FNDs to the facet normals of the denoised mesh. To handle meshes with different geometry features and reduce the training difficulty, we cluster the input mesh facets according to their FNDs and train neural networks for each cluster separately in an offline learning stage. At runtime, our method applies the learned cascaded regression functions to a noisy input mesh and reconstructs the denoised mesh from the output facet normals. Our method learns the non-linear denoising process from the training data and makes no specific assumptions about the noise distribution and geometry features in the input. The runtime denoising process is fully automatic for different input meshes. Our method can be easily adapted to meshes with arbitrary noise patterns by training a dedicated regression scheme with mesh data and the particular noise pattern. We evaluate our method on meshes with both synthetic and real scanned noise, and compare it to other mesh denoising algorithms. Results demonstrate that our method outperforms the state-of-the-art mesh denoising methods and successfully removes different kinds of noise for meshes with various geometry features. Peng-Shuai Wang, Yang Liu 0014, Xin Tong 0001 |
ACM Trans. Graph. | 3 |
| 2016 | Recovering shape and spatially-varying surface reflectance under unknown illuminationabstractWe present a novel integrated approach for estimating both spatially-varying surface reflectance and detailed geometry from a video of a rotating object under unknown static illumination. Key to our method is the decoupling of the recovery of normal and surface reflectance from the estimation of surface geometry. We define an apparent normal field with corresponding reflectance for each point (including those not on the object's surface) that best explain the observations. We observe that the object's surface goes through points where the apparent normal field and corresponding reflectance exhibit a high degree of consistency with the observations. However, estimating the apparent normal field requires knowledge of the unknown incident lighting. We therefore formulate the recovery of shape, surface reflectance, and incident lighting, as an iterative process that alternates between estimating shape and lighting, and simultaneously recovers surface reflectance at each step. To recover the shape, we first form an initial surface that passes through locations with consistent apparent temporal traces, followed by a refinement that maximizes the consistency of the surface normals with the underlying apparent normal field. To recover the lighting, we rely on appearance-from-motion using the recovered geometry from the previous step. We demonstrate our integrated framework on a variety of synthetic and real test cases exhibiting a wide variety of materials and shape. Yue Dong 0001, Pieter Peers, Xin Tong 0001 |
ACM Trans. Graph. | 4 |
| 2016 | Sparse-as-possible SVBRDF acquisitionabstractWe present a novel method for capturing real-world, spatially-varying surface reflectance from a small number of object views ( k ). Our key observation is that a specific target's reflectance can be represented by a small number of custom basis materials ( N ) convexly blended by an even smaller number of non-zero weights at each point ( n ). Based on this sparse basis/sparser blend model, we develop an SVBRDF reconstruction algorithm that jointly solves for n , N , the basis BRDFs, and their spatial blend weights with an alternating iterative optimization, each step of which solves a linearly-constrained quadratic programming problem. We develop a numerical tool that lets us estimate the number of views required and analyze the effect of lighting and geometry on reconstruction quality. We validate our method with images rendered from synthetic BRDFs, and demonstrate convincing results on real objects of pre-scanned shape and lit by uncontrolled natural illumination, from very few or even a single input image. Zhiming Zhou 0001, Yue Dong 0001, David P. Wipf, Yong Yu 0001, John M. Snyder, Xin Tong 0001 |
ACM Trans. Graph. | 7 |
| 2016 | 3D cartoon face generation by local deformation mapping
Jingyong Zhou, Xin Tong 0001, Zicheng Liu 0001, Baining Guo |
Vis. Comput. | 2 |
| 2015 | Efficient intrinsic image decomposition for RGBD imagesabstractIntrinsic image decomposition is a longstanding problem in computer vision. In this paper, we present a novel approach for efficiently decomposing an RGBD image into its reflectance and shading components. A robust super-pixel segmentation method is employed to select piece-wise constant reflectance regions and reduce the total number of unknowns. With the use of depth information, low frequency environment light can be represented by spherical harmonics and solved with super-pixels. After that, pixels that do not belong to any super-pixel are solved based on the super-pixels' shading. Compared to existing works, which often depend on the color Retinex assumption, our algorithm does not require any chromaticity-based constraints and enables us to solve many challenging cases such as color lighting environments and gray-scale textures. We also design an efficient solver for our system, and with our GPU implementation, it achieves 10-23 fps and boosts the decomposition process to real-time performance, enabling a wide range of applications such as dynamic object recoloring, re-texturing and virtual object composition. Yue Dong 0001, Xin Tong 0001, Yanyun Chen |
VRST | 3 |
| 2015 | Saliency-Preserving Slicing Optimization for Effective 3D PrintingabstractAbstract We present an adaptive slicing scheme for reducing the manufacturing time for 3D printing systems. Based on a new saliency‐based metric, our method optimizes the thicknesses of slicing layers to save printing time and preserve the visual quality of the printing results. We formulate the problem as a constrained ℓ0 optimization and compute the slicing result via a two‐step optimization scheme. To further reduce printing time, we develop a saliency‐based segmentation scheme to partition an object into subparts and then optimize the slicing of each subpart separately. We validate our method with a large set of 3D shapes ranging from CAD models to scanned objects. Results show that our method saves printing time by 30–40% and generates 3D objects that are visually similar to the ones printed with the finest resolution possible. Weiming Wang 0003, Haiyuan Chao, Jing Tong, Zhouwang Yang, Xin Tong 0001, Xiuping Liu, Ligang Liu 0001 |
Comput. Graph. Forum | 5 |
| 2015 | An efficient volumetric method for non-rigid registration
Xuejin Chen, Takaaki Shiratori, Xin Tong 0001, Ligang Liu 0001 |
Graph. Model. | 4 |
| 2015 | Irradiance regression for efficient final gathering in global illumination
Xuezhen Huang, Xin Sun 0014, Zhong Ren 0001, Xin Tong 0001, Baining Guo, Kun Zhou 0001 |
Frontiers Comput. Sci. | 4 |
| 2015 | Measurement-based editing of diffuse albedo with consistent interreflectionsabstractWe present a novel measurement-based method for editing the albedo of diffuse surfaces with consistent interreflections in a photograph of a scene under natural lighting. Key to our method is a novel technique for decomposing a photograph of a scene in several images that encode how much of the observed radiance has interacted a specified number of times with the target diffuse surface. Altering the albedo of the target area is then simply a weighted sum of the decomposed components. We estimate the interaction components by recursively applying the light transport operator and formulate the resulting radiance in each recursion as a linear expression in terms of the relevant interaction components. Our method only requires a camera-projector pair, and the number of required measurements per scene is linearly proportional to the decomposition degree for a single target area. Our method does not impose restrictions on the lighting or on the material properties in the unaltered part of the scene. Furthermore, we extend our method to accommodate editing of the albedo in multiple target areas with consistent interreflections and we introduce a prediction model for reducing the acquisition cost. We demonstrate our method on a variety of scenes and validate the accuracy on both synthetic and real examples. Bo Dong 0004, Yue Dong 0001, Xin Tong 0001, Pieter Peers |
ACM Trans. Graph. | 3 |
| 2015 | Video-audio driven real-time facial animationabstractWe present a real-time facial tracking and animation system based on a Kinect sensor with video and audio input. Our method requires no user-specific training and is robust to occlusions, large head rotations, and background noise. Given the color, depth and speech audio frames captured from an actor, our system first reconstructs 3D facial expressions and 3D mouth shapes from color and depth input with a multi-linear model. Concurrently a speaker-independent DNN acoustic model is applied to extract phoneme state posterior probabilities (PSPP) from the audio frames. After that, a lip motion regressor refines the 3D mouth shape based on both PSPP and expression weights of the 3D mouth shapes, as well as their confidences. Finally, the refined 3D mouth shape is combined with other parts of the 3D face to generate the final result. The whole process is fully automatic and executed in real time. The key component of our system is a data-driven regresor for modeling the correlation between speech data and mouth shapes. Based on a precaptured database of accurate 3D mouth shapes and associated speech audio from one speaker, the regressor jointly uses the input speech and visual features to refine the mouth shape of a new actor. We also present an improved DNN acoustic model. It not only preserves accuracy but also achieves real-time performance. Our method efficiently fuses visual and acoustic information for 3D facial performance capture. It generates more accurate 3D mouth motions than other approaches that are based on audio or video input only. It also supports video or audio only input for real-time facial animation. We evaluate the performance of our system with speech and facial expressions captured from different actors. Results demonstrate the efficiency and robustness of our method. Feng Xu 0005, Jinxiang Chai, Xin Tong 0001, Qiang Huo |
ACM Trans. Graph. | 4 |
| 2015 | Image based relighting using neural networksabstractWe present a neural network regression method for relighting realworld scenes from a small number of images. The relighting in this work is formulated as the product of the scene's light transport matrix and new lighting vectors, with the light transport matrix reconstructed from the input images. Based on the observation that there should exist non-linear local coherence in the light transport matrix, our method approximates matrix segments using neural networks that model light transport as a non-linear function of light source position and pixel coordinates. Central to this approach is a proposed neural network design which incorporates various elements that facilitate modeling of light transport from a small image set. In contrast to most image based relighting techniques, this regression-based approach allows input images to be captured under arbitrary illumination conditions, including light sources moved freely by hand. We validate our method with light transport data of real scenes containing complex lighting effects, and demonstrate that fewer input images are required in comparison to related techniques. Peiran Ren, Yue Dong 0001, Stephen Lin 0001, Xin Tong 0001, Baining Guo |
ACM Trans. Graph. | 4 |
| 2015 | Rolling guidance normal filter for geometric processingabstract3D geometric features constitute rich details of polygonal meshes. Their analysis and editing can lead to vivid appearance of shapes and better understanding of the underlying geometry for shape processing and analysis. Traditional mesh smoothing techniques mainly focus on noise filtering and they cannot distinguish different scales of features well, even mixing them up. We present an efficient method to process different scale geometric features based on a novel rolling-guidance normal filter. Given a 3D mesh, our method iteratively applies a joint bilateral filter to face normals at a specified scale, which empirically smooths small-scale geometric features while preserving large-scale features. Our method recovers the mesh from the filtered face normals by a modified Poisson-based gradient deformation that yields better surface quality than existing methods. We demonstrate the effectiveness and superiority of our method on a series of geometry processing tasks, including geometry texture removal and enhancement, coating transfer, mesh segmentation and level-of-detail meshing. Peng-Shuai Wang, Xiao-Ming Fu 0001, Yang Liu 0014, Xin Tong 0001, Baining Guo |
ACM Trans. Graph. | 4 |
| 2014 | Intrinsic Image Decomposition Using Structure-Texture Separation and Surface Normals
Junho Jeon, Sunghyun Cho, Xin Tong 0001, Seungyong Lee 0001 |
ECCV (7) | 3 |
| 2014 | Time-Lapse Photometric Stereo and ApplicationsabstractAbstract This paper presents a technique to recover geometry from time‐lapse sequences of outdoor scenes. We build upon photometric stereo techniques to recover approximate shadowing, shading and normal components allowing us to alter the material and normals of the scene. Previous work in analyzing such images has faced two fundamental difficulties: 1. the illumination in outdoor images consists of time‐varying sunlight and skylight, and 2. the motion of the sun is restricted to a near‐planar arc through the sky, making surface normal recovery unstable. We develop methods to estimate the reflection component due to skylight illumination. We also show that sunlight directions are usually non‐planar, thus making surface normal recovery possible. This allows us to estimate approximate surface normals for outdoor scenes using a single day of data. We demonstrate the use of these surface normals for a number of image editing applications including reflectance, lighting, and normal editing. Fangyang Shen, Kalyan Sunkavalli, Nicolas Bonneel, Szymon Rusinkiewicz, Hanspeter Pfister, Xin Tong 0001 |
Comput. Graph. Forum | 6 |
| 2014 | Acquisition of High Spatial and Spectral Resolution Video with a Hybrid Camera System
Chenguang Ma, Xun Cao, Xin Tong 0001, Qionghai Dai, Stephen Lin 0001 |
Int. J. Comput. Vis. | 3 |
| 2014 | Reflectance scanning: estimating shading frame and BRDF with generalized linear light sourcesabstractWe present a generalized linear light source solution to estimate both the local shading frame and anisotropic surface reflectance of a planar spatially varying material sample. We generalize linear light source reflectometry by modulating the intensity along the linear light source, and show that a constant and two sinusoidal lighting patterns are sufficient for estimating the local shading frame and anisotropic surface reflectance. We propose a novel reconstruction algorithm based on the key observation that after factoring out the tangent rotation, the anisotropic surface reflectance lies in a low rank subspace. We exploit the differences in tangent rotation between surface points to infer the low rank subspace and fit each surface point's reflectance function in the projected low rank subspace to the observations. We propose two prototype acquisition devices for capturing surface reflectance that differ on whether the camera is fixed with respect to the linear light source or fixed with respect to the material sample. We demonstrate convincing results obtained from reflectance scans of surfaces with different reflectance and shading frame variations. Yue Dong 0001, Pieter Peers, Jiawan Zhang, Xin Tong 0001 |
ACM Trans. Graph. | 5 |
| 2014 | Appearance-from-motion: recovering spatially varying surface reflectance under unknown lightingabstractWe present "appearance-from-motion", a novel method for recovering the spatially varying isotropic surface reflectance from a video of a rotating subject, with known geometry, under unknown natural illumination. We formulate the appearance recovery as an iterative process that alternates between estimating surface reflectance and estimating incident lighting. We characterize the surface reflectance by a data-driven microfacet model, and recover the microfacet normal distribution for each surface point separately from temporal changes in the observed radiance. To regularize the recovery of the incident lighting, we rely on the observation that natural lighting is sparse in the gradient domain. Furthermore, we exploit the sparsity of strong edges in the incident lighting to improve the robustness of the surface reflectance estimation. We demonstrate robust recovery of spatially varying isotropic reflectance from captured video as well as an internet video sequence for a wide variety of materials and natural lighting conditions. Yue Dong 0001, Pieter Peers, Jiawan Zhang, Xin Tong 0001 |
ACM Trans. Graph. | 5 |
| 2014 | Automatic acquisition of high-fidelity facial performances using monocular videosabstractThis paper presents a facial performance capture system that automatically captures high-fidelity facial performances using uncontrolled monocular videos ( e.g ., Internet videos). We start the process by detecting and tracking important facial features such as the nose tip and mouth corners across the entire sequence and then use the detected facial features along with multilinear facial models to reconstruct 3D head poses and large-scale facial deformation of the subject at each frame. We utilize per-pixel shading cues to add fine-scale surface details such as emerging or disappearing wrinkles and folds into large-scale facial deformation. At a final step, we iterate our reconstruction procedure on large-scale facial geometry and fine-scale facial details to further improve the accuracy of facial reconstruction. We have tested our system on monocular videos downloaded from the Internet, demonstrating its accuracy and robustness under a variety of uncontrolled lighting conditions and overcoming significant shape differences across individuals. We show our system advances the state of the art in facial performance capture by comparing against alternative methods. Fuhao Shi, Hsiang-Tao Wu, Xin Tong 0001, Jinxiang Chai |
ACM Trans. Graph. | 3 |
| 2014 | Hierarchical diffusion curves for accurate automatic image vectorizationabstractDiffusion curve primitives are a compact and powerful representation for vector images. While several vector image authoring tools leverage these representations, automatically and accurately vectorizing arbitrary raster images using diffusion curves remains a difficult problem. We automatically generate sparse diffusion curve vectorizations of raster images by fitting curves in the Laplacian domain. Our approach is fast, combines Laplacian and bilaplacian diffusion curve representations, and generates a hierarchical representation that accurately reconstructs both vector art and natural images. The key idea of our method is to trace curves in the Laplacian domain, which captures both sharp and smooth image features, across scales, more robustly than previous image- and gradient-domain fitting strategies. The sparse set of curves generated by our method accurately reconstructs images and often closely matches tediously hand-authored curve data. Also, our hierarchical curves are readily usable in all existing editing frameworks. We validate our method on a broad class of images, including natural images, synthesized images with turbulent multi-scale details, and traditional vector-art, as well as illustrating simple multi-scale abstraction and color editing results. Guofu Xie, Xin Sun 0014, Xin Tong 0001, Derek Nowrouzezahrai |
ACM Trans. Graph. | 3 |
| 2014 | Controllable high-fidelity facial performance transferabstractRecent technological advances in facial capture have made it possible to acquire high-fidelity 3D facial performance data with stunningly high spatial-temporal resolution. Current methods for facial expression transfer, however, are often limited to large-scale facial deformation. This paper introduces a novel facial expression transfer and editing technique for high-fidelity facial performance data. The key idea of our approach is to decompose high-fidelity facial performances into high-level facial feature lines, large-scale facial deformation and fine-scale motion details and transfer them appropriately to reconstruct the retargeted facial animation in an efficient optimization framework. The system also allows the user to quickly modify and control the retargeted facial sequences in the spatial-temporal domain. We demonstrate the power of our approach by transferring and editing high-fidelity facial animation data from high-resolution source models to a wide range of target models, including both human faces and non-human faces such as "monster" and "dog". Feng Xu 0005, Jinxiang Chai, Xin Tong 0001 |
ACM Trans. Graph. | 4 |
| 2014 | Sensitivity-optimized rigging for example-based real-time clothing synthesisabstractWe present a real-time solution for generating detailed clothing deformations from pre-computed clothing shape examples. Given an input pose, it synthesizes a clothing deformation by blending skinned clothing deformations of nearby examples controlled by the body skeleton. Observing that cloth deformation can be well modeled with sensitivity analysis driven by the underlying skeleton, we introduce a sensitivity based method to construct a pose-dependent rigging solution from sparse examples. We also develop a sensitivity based blending scheme to find nearby examples for the input pose and evaluate their contributions to the result. Finally, we propose a stochastic optimization based greedy scheme for sampling the pose space and generating example clothing shapes. Our solution is fast, compact and can generate realistic clothing animation results for various kinds of clothes in real time. Weiwei Xu 0003, Nobuyuki Umetani, Qianwen Chao, Jie Mao, Xiaogang Jin 0001, Xin Tong 0001 |
ACM Trans. Graph. | 6 |
| 2014 | Dynamic hair capture using spacetime optimizationabstractDynamic hair strands have complex structures and experience intricate collisions and occlusion, posing significant challenges for high-quality reconstruction of their motions. We present a comprehensive dynamic hair capture system for reconstructing realistic hair motions from multiple synchronized video sequences. To recover hair strands' temporal correspondence, we propose a motion-path analysis algorithm that can robustly track local hair motions in input videos. To ensure the spatial and temporal coherence of the dynamic capture, we formulate the global hair reconstruction as a spacetime optimization problem solved iteratively. Demonstrated using a range of real-world hairstyles driven by different wind conditions and head motions, our approach is able to reconstruct complex hair dynamics matching closely with video recordings both in terms of geometry and motion details. Zexiang Xu, Hsiang-Tao Wu, Lvdi Wang, Changxi Zheng, Xin Tong 0001 |
ACM Trans. Graph. | 5 |
| 2013 | Accurate and Robust 3D Facial Capture Using a Single RGBD CameraabstractThis paper presents an automatic and robust approach that accurately captures high-quality 3D facial performances using a single RGBD camera. The key of our approach is to combine the power of automatic facial feature detection and image-based 3D nonrigid registration techniques for 3D facial reconstruction. In particular, we develop a robust and accurate image-based nonrigid registration algorithm that incrementally deforms a 3D template mesh model to best match observed depth image data and important facial features detected from single RGBD images. The whole process is fully automatic and robust because it is based on single frame facial registration framework. The system is flexible because it does not require any strong 3D facial priors such as blend shape models. We demonstrate the power of our approach by capturing a wide range of 3D facial expressions using a single RGBD camera and achieve state-of-the-art accuracy by comparing against alternative methods. Yen-Lin Chen, Hsiang-Tao Wu, Fuhao Shi, Xin Tong 0001, Jinxiang Chai |
ICCV | 4 |
| 2013 | BodyAvatar: creating freeform 3D avatars using first-person body gesturesabstractBodyAvatar is a Kinect-based interactive system that allows users without professional skills to create freeform 3D avatars using body gestures. Unlike existing gesture-based 3D modeling tools, BodyAvatar centers around a first-person "you're the avatar" metaphor, where the user treats their own body as a physical proxy of the virtual avatar. Based on an intuitive body-centric mapping, the user performs gestures to their own body as if wanting to modify it, which in turn results in corresponding modifications to the avatar. BodyAvatar provides an intuitive, immersive, and playful creation experience for the user. We present a formative study that leads to the design of BodyAvatar, the system's interactions and underlying algorithms, and results from initial user trials. Teng Han, Zhimin Ren, Nobuyuki Umetani, Xin Tong 0001, Yang Liu 0014, Takaaki Shiratori |
UIST | 5 |
| 2013 | Bi-scale appearance fabricationabstractSurfaces in the real world exhibit complex appearance due to spatial variations in both their reflectance and local shading frames (i.e. the local coordinate system defined by the normal and tangent direction). For opaque surfaces, existing fabrication solutions can reproduce well only the spatial variations of isotropic reflectance. In this paper, we present a system for fabricating surfaces with desired spatially-varying reflectance, including anisotropic ones, and local shading frames. We approximate each input reflectance, rotated by its local frame, as a small patch of oriented facets coated with isotropic glossy inks. By assigning different ink combinations to facets with different orientations, this bi-scale material can reproduce a wider variety of reflectance than the printer gamut, including anisotropic materials. By orienting the facets appropriately, we control the local shading frame. We propose an algorithm to automatically determine the optimal facets orientations and ink combinations that best approximate a given input appearance, while obeying manufacturing constraints on both geometry and ink gamut. We fabricate the resulting surface with commercially available hardware, a 3D printer to fabricate the facets and a flatbed UV printer to coat them with inks. We validate our method by fabricating a variety of isotropic and anisotropic materials with rich variations in normals and tangents. Yanxiang Lan, Yue Dong 0001, Fabio Pellacini, Xin Tong 0001 |
ACM Trans. Graph. | 4 |
| 2013 | Dynamic element texturesabstractMany natural phenomena consist of geometric elements with dynamic motions characterized by small scale repetitions over large scale structures, such as particles, herds, threads, and sheets. Due to their ubiquity, controlling the appearance and behavior of such phenomena is important for a variety of graphics applications. However, such control is often challenging; the repetitive elements are often too numerous for manual edit, while their overall structures are often too versatile for fully automatic computation. We propose a method that facilitates easy and intuitive controls at both scales: high-level structures through spatial-temporal output constraints (e.g. overall shape and motion of the output domain), and low-level details through small input exemplars (e.g. element arrangements and movements). These controls are suitable for manual specification, while the corresponding geometric and dynamic repetitions are suitable for automatic computation. Our system takes such user controls as inputs, and generates as outputs the corresponding repetitions satisfying the controls. Our method, which we call dynamic element textures , aims to produce such controllable repetitions through a combination of constrained optimization (satisfying controls) and data driven computation (synthesizing details). We use spatial-temporal samples as the core representation for dynamic geometric elements. We propose analysis algorithms for decomposing small scale repetitions from large scale themes, as well as synthesis algorithms for generating outputs satisfying user controls. Our method is general, producing a range of artistic effects that previously required disparate and specialized techniques. Chongyang Ma, Li-Yi Wei, Sylvain Lefebvre 0001, Xin Tong 0001 |
ACM Trans. Graph. | 4 |
| 2013 | Global illumination with radiance regression functionsabstractWe present radiance regression functions for fast rendering of global illumination in scenes with dynamic local light sources. A radiance regression function (RRF) represents a non-linear mapping from local and contextual attributes of surface points, such as position, viewing direction, and lighting condition, to their indirect illumination values. The RRF is obtained from precomputed shading samples through regression analysis, which determines a function that best fits the shading data. For a given scene, the shading samples are precomputed by an offline renderer. The key idea behind our approach is to exploit the nonlinear coherence of the indirect illumination data to make the RRF both compact and fast to evaluate. We model the RRF as a multilayer acyclic feed-forward neural network, which provides a close functional approximation of the indirect illumination and can be efficiently evaluated at run time. To effectively model scenes with spatially variant material properties, we utilize an augmented set of attributes as input to the neural network RRF to reduce the amount of inference that the network needs to perform. To handle scenes with greater geometric complexity, we partition the input space of the RRF model and represent the subspaces with separate, smaller RRFs that can be evaluated more rapidly. As a result, the RRF model scales well to increasingly complex scene geometry and material variation. Because of its compactness and ease of evaluation, the RRF model enables real-time rendering with full global illumination effects, including changing caustics and multiple-bounce high-frequency glossy interreflections. Peiran Ren, Jiaping Wang, Minmin Gong, Stephen Lin 0001, Xin Tong 0001, Baining Guo |
ACM Trans. Graph. | 5 |
| 2013 | Cost-effective printing of 3D objects with skin-frame structuresabstract3D printers have become popular in recent years and enable fabrication of custom objects for home users. However, the cost of the material used in printing remains high. In this paper, we present an automatic solution to design a skin-frame structure for the purpose of reducing the material cost in printing a given 3D object. The frame structure is designed by an optimization scheme which significantly reduces material volume and is guaranteed to be physically stable, geometrically approximate, and printable. Furthermore, the number of struts is minimized by solving an l 0 sparsity optimization. We formulate it as a multi-objective programming problem and an iterative extension of the preemptive algorithm is developed to find a compromise solution. We demonstrate the applicability and practicability of our solution by printing various objects using both powder-type and extrusion-type 3D printers. Our method is shown to be more cost-effective than previous works. Weiming Wang 0003, Tuanfeng Y. Wang, Zhouwang Yang, Ligang Liu 0001, Xin Tong 0001, Weihua Tong, Jiansong Deng, Falai Chen, Xiuping Liu |
ACM Trans. Graph. | 5 |
| 2013 | Interactive chromaticity mapping for multispectral images
Yanxiang Lan, Jiaping Wang, Stephen Lin 0001, Minmin Gong, Xin Tong 0001, Baining Guo |
Vis. Comput. | 5 |
| 2012 | Estimation of Intrinsic Image Sequences from Image+Depth Video
Kyong Joon Lee, Xin Tong 0001, Minmin Gong, Shahram Izadi, Sang Uk Lee, Ping Tan 0002, Stephen Lin 0001 |
ECCV (6) | 3 |
| 2012 | Printing spatially-varying reflectance for reproducing HDR imagesabstractWe present a solution for viewing high dynamic range (HDR) images with spatially-varying distributions of glossy materials printed on reflective media. Our method exploits appearance variations of the glossy materials in the angular domain to display the input HDR image at different exposures. As viewers change the print orientation or lighting directions, the print gradually varies its appearance to display the image content from the darkest to the brightest levels. Our solution is based on a commercially available printing system and is fully automatic. Given the input HDR image and the BRDFs of a set of available inks, our method computes the optimal exposures of the HDR image for all viewing conditions and the optimal ink combinations for all pixels by minimizing the difference of their appearances under all viewing conditions. We demonstrate the effectiveness of our method with print samples generated from different inputs and visualized under different viewing and lighting conditions. Yue Dong 0001, Xin Tong 0001, Fabio Pellacini, Baining Guo |
ACM Trans. Graph. | 2 |
| 2012 | Diffusion curve textures for resolution independent texture mappingabstractWe introduce a vector representation called diffusion curve textures for mapping diffusion curve images (DCI) onto arbitrary surfaces. In contrast to the original implicit representation of DCIs [Orzan et al. 2008], where determining a single texture value requires iterative computation of the entire DCI via the Poisson equation, diffusion curve textures provide an explicit representation from which the texture value at any point can be solved directly, while preserving the compactness and resolution independence of diffusion curves. This is achieved through a formulation of the DCI diffusion process in terms of Green's functions. This formulation furthermore allows the texture value of any rectangular region (e.g. pixel area) to be solved in closed form, which facilitates anti-aliasing. We develop a GPU algorithm that renders anti-aliased diffusion curve textures in real time, and demonstrate the effectiveness of this method through high quality renderings with detailed control curves and color variations. Xin Sun 0014, Guofu Xie, Yue Dong 0001, Stephen Lin 0001, Weiwei Xu 0003, Xin Tong 0001, Baining Guo |
ACM Trans. Graph. | 7 |
| 2012 | Detail-Preserving Controllable Deformation from Sparse ExamplesabstractRecent advances in laser scanning technology have made it possible to faithfully scan a real object with tiny geometric details, such as pores and wrinkles. However, a faithful digital model should not only capture static details of the real counterpart but also be able to reproduce the deformed versions of such details. In this paper, we develop a data-driven model that has two components; the first accommodates smooth large-scale deformations and the second captures high-resolution details. Large-scale deformations are based on a nonlinear mapping between sparse control points and bone transformations. A global mapping, however, would fail to synthesize realistic geometries from sparse examples, for highly deformable models with a large range of motion. The key is to train a collection of mappings defined over regions locally in both the geometry and the pose space. Deformable fine-scale details are generated from a second nonlinear mapping between the control points and per-vertex displacements. We apply our modeling scheme to scanned human hand models, scanned face models, face models reconstructed from multiview video sequences, and manually constructed dinosaur models. Experiments show that our deformation models, learned from extremely sparse training data, are effective and robust in synthesizing highly deformable models with rich fine features, for keyframe animation as well as performance-driven animation. We also compare our results with those obtained by alternative techniques. Hao-Da Huang, KangKang Yin, Ling Zhao 0006, Yizhou Yu, Xin Tong 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2012 | Real-time rendering of deformable heterogeneous translucent objects using multiresolution splatting
Pieter Peers, Jiawan Zhang, Xin Tong 0001 |
Vis. Comput. | 4 |
| 2011 | High resolution multispectral video capture with a hybrid camera systemabstractWe present a new approach to capture video at high spatial and spectral resolutions using a hybrid camera system. Composed of an RGB video camera, a grayscale video camera and several optical elements, the hybrid camera system simultaneously records two video streams: an RGB video with high spatial resolution, and a multispectral video with low spatial resolution. After registration of the two video streams, our system propagates the multispectral information into the RGB video to produce a video with both high spectral and spatial resolution. This propagation between videos is guided by color similarity of pixels in the spectral domain, proximity in the spatial domain, and the consistent color of each scene point in the temporal domain. The propagation algorithm is designed for rapid computation to allow real-time video generation at the original frame rate, and can thus facilitate real-time video analysis tasks such as tracking and surveillance. Hardware implementation details and design tradeoffs are discussed. We evaluate the proposed system using both simulations with ground truth data and on real-world scenes. The utility of this high resolution multispectral video data is demonstrated in dynamic white balance adjustment and tracking. Xun Cao, Xin Tong 0001, Qionghai Dai, Stephen Lin 0001 |
CVPR | 2 |
| 2011 | A Prism-Mask System for Multispectral Video AcquisitionabstractThis paper presents a prism-mask system for capturing multispectral videos. The system is composed of a triangular prism, a monochrome camera, and an occlusion mask. Incoming light beams from the scene are sampled by the occlusion mask, dispersed into their constituent spectra by the triangular prism, and then captured by the monochrome camera. Our system is capable of capturing frames with high spectral resolution at video rates. It also allows for different trade-offs between spectral and spatial resolution by adjusting the focal length of the camera. We demonstrate multispectral video acquisition with various spectral resolutions and spatial resolutions, as well as different frame rates. The effectiveness of our system is further evaluated with several applications, including human skin detection, physical material recognition, video segmentation, RGB video generation, and illumination identification. Xun Cao, Hao Du 0004, Xin Tong 0001, Qionghai Dai, Stephen Lin 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | AppWarp: retargeting measured materials by appearance-space warpingabstractWe propose a method for retargeting measured materials, where a source measured material is edited by applying the reflectance functions of a template measured dataset. The resulting dataset is a material that maintains the spatial patterns of the source dataset, while exhibiting the reflectance behaviors of the template. Compared to editing materials by subsequent selections and modifications, retargeting shortens the time required to achieve a desired look by directly using template data, just as color transfer does for editing images. With our method, users have to just mark corresponding regions of source and template with rough strokes, with no need for further input. This paper introduces AppWarp , an algorithm that achieves retargeting as a user-constrained, appearance-space warping operation, that executes in tens of seconds. Our algorithm is independent of the measured material representation and supports retargeting of analytic and tabulated BRDFs as well as BSSRDFs. In addition, our method makes no assumption of the data distribution in appearance-space nor on the underlying correspondence between source and target. These characteristics make AppWarp the first general formulation for appearance retargeting. We validate our method on several types of materials, including leaves, metals, waxes, woods and greeting cards. Furthermore, we demonstrate how retargeting can be used to enhance diffuse texture with high quality reflectance. Xiaobo An, Xin Tong 0001, Jonathan D. Denning, Fabio Pellacini |
ACM Trans. Graph. | 2 |
| 2011 | AppGen: interactive material modeling from a single imageabstractWe present AppGen , an interactive system for modeling materials from a single image. Given a texture image of a nearly planar surface lit with directional lighting, our system models the detailed spatially-varying reflectance properties (diffuse, specular and roughness) and surface normal variations with minimal user interaction. We ask users to indicate global shading and reflectance information by roughly marking the image with a few user strokes, while our system assigns reflectance properties and normals to each pixel. We first interactively decompose the input image into the product of a diffuse albedo map and a shading map. A two-scale normal reconstruction algorithm is then introduced to recover the normal variations from the shading map and preserve the geometric features at different scales. We finally assign the specular parameters to each pixel guided by user strokes and the diffuse albedo. Our system generates convincing results within minutes of interaction and works well for a variety of material types that exhibit different reflectance and normal variations, including natural surfaces and man-made ones. Yue Dong 0001, Xin Tong 0001, Fabio Pellacini, Baining Guo |
ACM Trans. Graph. | 2 |
| 2011 | Leveraging motion capture and 3D scanning for high-fidelity facial performance acquisitionabstractThis paper introduces a new approach for acquiring high-fidelity 3D facial performances with realistic dynamic wrinkles and fine-scale facial details. Our approach leverages state-of-the-art motion capture technology and advanced 3D scanning technology for facial performance acquisition. We start the process by recording 3D facial performances of an actor using a marker-based motion capture system and perform facial analysis on the captured data, thereby determining a minimal set of face scans required for accurate facial reconstruction. We introduce a two-step registration process to efficiently build dense consistent surface correspondences across all the face scans. We reconstruct high-fidelity 3D facial performances by combining motion capture data with the minimal set of face scans in the blendshape interpolation framework. We have evaluated the performance of our system on both real and synthetic data. Our results show that the system can capture facial performances that match both the spatial resolution of static face scans and the acquisition speed of motion capture systems. Hao-Da Huang, Jinxiang Chai, Xin Tong 0001, Hsiang-Tao Wu |
ACM Trans. Graph. | 3 |
| 2011 | Discrete element texturesabstractA variety of phenomena can be characterized by repetitive small scale elements within a large scale domain. Examples include a stack of fresh produce, a plate of spaghetti, or a mosaic pattern. Although certain results can be produced via manual placement or procedural/physical simulation, these methods can be labor intensive, difficult to control, or limited to specific phenomena. We present discrete element textures, a data-driven method for synthesizing repetitive elements according to a small input exemplar and a large output domain. Our method preserves both individual element properties and their aggregate distributions. It is also general and applicable to a variety of phenomena, including different dimensionalities, different element properties and distributions, and different effects including both artistic and physically realistic ones. We represent each element by one or multiple samples whose positions encode relevant element attributes including position, size, shape, and orientation. We propose a sample-based neighborhood similarity metric and an energy optimization solver to synthesize desired outputs that observe not only input exemplars and output domains but also optional constraints such as physics, orientation fields, and boundary conditions. As a further benefit, our method can also be applied for editing existing element distributions. Chongyang Ma, Li-Yi Wei, Xin Tong 0001 |
ACM Trans. Graph. | 3 |
| 2011 | Pocket reflectometryabstractWe present a simple, fast solution for reflectance acquisition using tools that fit into a pocket. Our method captures video of a flat target surface from a fixed video camera lit by a hand-held, moving, linear light source. After processing, we obtain an SVBRDF. We introduce a BRDF chart , analogous to a color "checker" chart, which arranges a set of known-BRDF reference tiles over a small card. A sequence of light responses from the chart tiles as well as from points on the target is captured and matched to reconstruct the target's appearance. We develop a new algorithm for BRDF reconstruction which works directly on these LDR responses, without knowing the light or camera position, or acquiring HDR lighting. It compensates for spatial variation caused by the local (finite distance) camera and light position by warping responses over time to align them to a specular reference. After alignment, we find an optimal linear combination of the Lambertian and purely specular reference responses to match each target point's response. The same weights are then applied to the corresponding (known) reference BRDFs to reconstruct the target point's BRDF. We extend the basic algorithm to also recover varying surface normals by adding two spherical caps for diffuse and specular references to the BRDF chart. We demonstrate convincing results obtained after less than 30 seconds of data capture, using commercial mobile phone cameras in a casual environment. Peiran Ren, Jiaping Wang, John M. Snyder, Xin Tong 0001, Baining Guo |
ACM Trans. Graph. | 4 |
| 2011 | TextFlow: Towards Better Understanding of Evolving Topics in TextabstractUnderstanding how topics evolve in text data is an important and challenging task. Although much work has been devoted to topic analysis, the study of topic evolution has largely been limited to individual topics. In this paper, we introduce TextFlow, a seamless integration of visualization and topic mining techniques, for analyzing various evolution patterns that emerge from multiple topics. We first extend an existing analysis technique to extract three-level features: the topic evolution trend, the critical event, and the keyword correlation. Then a coherent visualization that consists of three new visual components is designed to convey complex relationships between them. Through interaction, the topic mining model and visualization can communicate with each other to help users refine the analysis result and gain insights into the data progressively. Finally, two case studies are conducted to demonstrate the effectiveness and usefulness of TextFlow in helping users understand the major topic evolution patterns in time-varying text data. Weiwei Cui 0001, Shixia Liu, Conglei Shi, Yangqiu Song, Zekai Gao, Huamin Qu, Xin Tong 0001 |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2010 | Turning Rust into Gold: An ancient artifact as an interactive artworkabstractTurning Rust into Gold is inspired by a Chinese antique Mao-Kung Ting (cauldron) treasured by the National Palace Museum in Taiwan. Having a five-hundred-character inscription cast inside, and its weathered appearance made the Mao-Kung very unique. Motivated by revealing the great nature of the artifact and interpreting it into a meaningful narrative, we have proposed an interactive multimedia system that facilitates effective communication between museum audiences and the Mao-Kung Ting. Three technologies have been implemented to emphasize the weathered appearance of the bronze. De-/weathering simulation techniques have been deployed to revive the bronze to its original shiny gold color; while breath-based biofeedback and haptic technology have been utilized as user interfaces to trigger the de-weathering process of the Mao-Kung Ting. Also, the interactive scenarios have been designed with the Chinese cultural context and philosophy Qi, enabling users more easily fall into the Chinese civilization. The paper aims to present the development of the artwork Turing Rust into Gold, in order to further contribute to the feasibility of incorporating new media art in a historical museum context, and bring a new horizon in the museum sector. Chun-Ko Hsieh, Xin Tong 0001, Yi-Ping Hung, Chia-Ping Chen, Liang-Chun Lin, I-Ling Liu, Meng-Chieh Yu, Chu-Song Chen, Jiaping Wang |
ICME | 2 |
| 2010 | Transformational Breathing between Present and Past: Virtual Exhibition System of the Mao-Kung Ting
Chun-Ko Hsieh, Xin Tong 0001, Yi-Ping Hung, Chia-Ping Chen, Ju-Chun Ko, Meng-Chieh Yu, Han-Hung Lin, Szu-Wei Wu, Yi-Yu Chung, Liang-Chun Lin, Ming-Sui Lee, Chu-Song Chen, Jiaping Wang, Quo-Ping Lin, I-Ling Liu |
MMM | 2 |
| 2010 | Condenser-Based Instant ReflectometryabstractAbstract We present a technique for rapid capture of high quality bidirectional reflection distribution functions(BRDFs) of surface points. Our method represents the BRDF at each point by a generalized microfacet model with tabulated normal distribution function (NDF) and assumes that the BRDF is symmetrical. A compact and light‐weight reflectometry apparatus is developed for capturing reflectance data from each surface point within one second. The device consists of a pair of condenser lenses, a video camera, and six LED light sources. During capture, the reflected rays from a surface point lit by a LED lighting are refracted by a condenser lenses and efficiently collected by the camera CCD. Taking advantage of BRDF symmetry, our reflectometry apparatus provides an efficient optical design to improve the measurement quality. We also propose a model fitting algorithm for reconstructing the generalized microfacet model from the sparse BRDF slices captured from a material surface point. Our new algorithm addresses the measurement errors and generates more accurate results than previous work. Our technique provides a practical and efficient solution for BRDF acquisition, especially for materials with anisotropic reflectance. We test the accuracy of our approach by comparing our results with ground truth. We demonstrate the efficiency of our reflectometry by measuring materials with different degrees of specularity, values of Fresnel factor, and angular variation. Yanxiang Lan, Yue Dong 0001, Jiaping Wang, Xin Tong 0001, Baining Guo |
Comput. Graph. Forum | 4 |
| 2010 | Image-based modeling of inhomogeneous single-scattering participating media
Xin Tong 0001 |
Sci. China Inf. Sci. | 3 |
| 2010 | Fabricating spatially-varying subsurface scatteringabstractMany real world surfaces exhibit translucent appearance due to subsurface scattering. Although various methods exists to measure, edit and render subsurface scattering effects, no solution exists for manufacturing physical objects with desired translucent appearance. In this paper, we present a complete solution for fabricating a material volume with a desired surface BSSRDF. We stack layers from a fixed set of manufacturing materials whose thickness is varied spatially to reproduce the heterogeneity of the input BSSRDF. Given an input BSSRDF and the optical properties of the manufacturing materials, our system efficiently determines the optimal order and thickness of the layers. We demonstrate our approach by printing a variety of homogenous and heterogenous BSSRDFs using two hardware setups: a milling machine and a 3D printer. Yue Dong 0001, Jiaping Wang, Fabio Pellacini, Xin Tong 0001, Baining Guo |
ACM Trans. Graph. | 4 |
| 2010 | Manifold bootstrapping for SVBRDF captureabstractManifold bootstrapping is a new method for data-driven modeling of real-world, spatially-varying reflectance, based on the idea that reflectance over a given material sample forms a low-dimensional manifold. It provides a high-resolution result in both the spatial and angular domains by decomposing reflectance measurement into two lower-dimensional phases. The first acquires representatives of high angular dimension but sampled sparsely over the surface, while the second acquires keys of low angular dimension but sampled densely over the surface. We develop a hand-held, high-speed BRDF capturing device for phase one measurements. A condenser-based optical setup collects a dense hemisphere of rays emanating from a single point on the target sample as it is manually scanned over it, yielding 10 BRDF point measurements per second. Lighting directions from 6 LEDs are applied at each measurement; these are amplified to a full 4D BRDF using the general (NDF-tabulated) microfacet model. The second phase captures N =20-200 images of the entire sample from a fixed view and lit by a varying area source. We show that the resulting N -dimensional keys capture much of the distance information in the original BRDF space, so that they effectively discriminate among representatives, though they lack sufficient angular detail to reconstruct the SVBRDF by themselves. At each surface position, a local linear combination of a small number of neighboring representatives is computed to match each key, yielding a high-resolution SVBRDF. A quick capture session (10-20 minutes) on simple devices yields results showing sharp and anisotropic specularity and rich spatial detail. Yue Dong 0001, Jiaping Wang, Xin Tong 0001, John M. Snyder, Yanxiang Lan, Moshe Ben-Ezra, Baining Guo |
ACM Trans. Graph. | 3 |
| 2009 | A prism-based system for multispectral video acquisitionabstractIn this paper, we propose a prism-based system for capturing multispectral videos. The system consists of a triangular prism, a monochrome camera, and an occlusion mask. Incoming light beams from the scene are sampled by the occlusion mask, dispersed into their constituent spectra by the triangular prism, and then captured by the monochrome camera. Our system is capable of capturing videos of high spectral resolution. It also allows for different tradeoffs between spectral and spatial resolution by adjusting the focal length of the camera. We demonstrate the effectiveness of our system with several applications, including human skin detection, physical material recognition, and RGB video generation. Hao Du 0004, Xin Tong 0001, Xun Cao, Stephen Lin 0001 |
ICCV | 2 |
| 2009 | Texture SplicingabstractAbstract We propose a new texture editing operation called texture splicing. For this operation, we regard a texture as having repetitive elements (textons) seamlessly distributed in a particular pattern. Taking two textures as input, texture splicing generates a new texture by selecting the texton appearance from one texture and distribution from the other. Texture splicing involves self‐similarity search to extract the distribution, distribution warping, context‐dependent warping, and finally, texture refinement to preserve overall appearance. We show a variety of results to illustrate this operation. Yiming Liu 0001, Jiaping Wang, Su Xue, Xin Tong 0001, Sing Bing Kang, Baining Guo |
Comput. Graph. Forum | 4 |
| 2009 | Edit Propagation on Bidirectional Texture FunctionsabstractAbstract We propose an efficient method for editing bidirectional texture functions (BTFs) based on edit propagation scheme. In our approach, users specify sparse edits on a certain slice of BTF. An edit propagation scheme is then applied to propagate edits to the whole BTF data. The consistency of the BTF data is maintained by propagating similar edits to points with similar underlying geometry/reflectance. For this purpose, we propose to use view independent features including normals and reflectance features reconstructed from each view to guide the propagation process. We also propose an adaptive sampling scheme for speeding up the propagation process. Since our method needn't any accurate geometry and reflectance information, it allows users to edit complex BTFs with interactive feedback. Kun Xu 0003, Jiaping Wang, Xin Tong 0001, Shi-Min Hu 0001, Baining Guo |
Comput. Graph. Forum | 3 |
| 2009 | SubEdit: a representation for editing measured heterogeneous subsurface scatteringabstractIn this paper we present SubEdit , a representation for editing the BSSRDF of heterogeneous subsurface scattering acquired from real-world samples. Directly editing measured raw data is difficult due to the non-local impact of heterogeneous subsurface scattering on the appearance. Our SubEdit representation decouples these non-local effects into the product of two local scattering profiles defined at respectively the incident and outgoing surface locations. This allows users to directly manipulate the appearance of single surface locations and to robustly make selections. To further facilitate editing, we reparameterize the scattering profiles into the local appearance concepts of albedo, scattering range, and profile shape. Our method preserves the visual quality of the measured material after editing by maintaining the consistency of subsurface transport for all edits. SubEdit fits measured data well while remaining efficient enough to support interactive rendering and manipulation. We illustrate the suitability of SubEdit as a representation for editing by applying various complex modifications on a wide variety of measured heterogeneous subsurface scattering materials. Xin Tong 0001, Fabio Pellacini, Pieter Peers |
ACM Trans. Graph. | 2 |
| 2009 | Kernel Nyström method for light transportabstractWe propose a kernel Nyström method for reconstructing the light transport matrix from a relatively small number of acquired images. Our work is based on the generalized Nyström method for low rank matrices. We introduce the light transport kernel and incorporate it into the Nyström method to exploit the nonlinear coherence of the light transport matrix. We also develop an adaptive scheme for efficiently capturing the sparsely sampled images from the scene. Our experiments indicate that the kernel Nyström method can achieve good reconstruction of the light transport matrix with a few hundred images and produce high quality relighting results. The kernel Nyström method is effective for modeling scenes with complex lighting effects and occlusions which have been challenging for existing techniques. Jiaping Wang, Yue Dong 0001, Xin Tong 0001, Zhouchen Lin, Baining Guo |
ACM Trans. Graph. | 3 |
| 2008 | Lazy Solid Texture SynthesisabstractAbstract Existing solid texture synthesis algorithms generate a full volume of color content from a set of 2D example images. We introduce a new algorithm with the unique ability to restrict synthesis to a subset of the voxels, while enforcing spatial determinism. This is especially useful when texturing objects, since only a thick layer around the surface needs to be synthesized. A major difficulty lies in reducing the dependency chain of neighborhood matching, so that each voxel only depends on a small number of other voxels. Our key idea is to synthesize a volume from a set of pre‐computed 3D candidates, each being a triple of interleaved 2D neighborhoods. We present an efficient algorithm to carefully select in a pre‐process only those candidates forming consistent triples. This significantly reduces the search space during subsequent synthesis. The result is a new parallel, spatially deterministic solid texture synthesis algorithm which runs efficiently on the GPU. Our approach generates high resolution solid textures on surfaces within seconds. Memory usage and synthesis time only depend on the output textured surface area. The GPU implementation of our method rapidly synthesizes new textures for the surfaces appearing when interactively breaking or cutting objects. Yue Dong 0001, Sylvain Lefebvre 0001, Xin Tong 0001, George Drettakis |
Comput. Graph. Forum | 3 |
| 2008 | Image-based Material WeatheringabstractAbstract The appearance manifold [WTL*06] is an efficient approach for modeling and editing time‐variant appearance of materials from the BRDF data captured at single time instance. However, this method is difficult to apply in images in which weathering and shading variations are combined. In this paper, we present a technique for modeling and editing the weathering effects of an object in a single image with appearance manifolds. In our approach, we formulate the input image as the product of reflectance and illuminance. An iterative method is then developed to construct the appearance manifold in color space (i.e., Lab space) for modeling the reflectance variations caused by weathering. Based on the appearance manifold, we propose a statistical method to robustly decompose reflectance and illuminance for each pixel. For editing, we introduce a “pixel‐walking” scheme to modify the pixel reflectance according to its position on the manifold, by which the detailed reflectance variations are well preserved. We illustrate our technique in various applications, including weathering transfer between two images that is first enabled by our technique. Results show that our technique can produce much better results than existing methods, especially for objects with complex geometry and shading effects. Su Xue, Jiaping Wang, Xin Tong 0001, Qionghai Dai, Baining Guo |
Comput. Graph. Forum | 3 |
| 2008 | Modeling and rendering of heterogeneous translucent materials using the diffusion equationabstractIn this article, we propose techniques for modeling and rendering of heterogeneous translucent materials that enable acquisition from measured samples, interactive editing of material attributes, and real-time rendering. The materials are assumed to be optically dense such that multiple scattering can be approximated by a diffusion process described by the diffusion equation. For modeling heterogeneous materials, we present the inverse diffusion algorithm for acquiring material properties from appearance measurements. This modeling algorithm incorporates a regularizer to handle the ill-conditioning of the inverse problem, an adjoint method to dramatically reduce the computational cost, and a hierarchical GPU implementation for further speedup. To render an object with known material properties, we present the polygrid diffusion algorithm , which solves the diffusion equation with a boundary condition defined by the given illumination environment. This rendering technique is based on representation of an object by a polygrid, a grid with regular connectivity and an irregular shape, which facilitates solution of the diffusion equation in arbitrary volumes. Because of the regular connectivity, our rendering algorithm can be implemented on the GPU for real-time performance. We demonstrate our techniques by capturing materials from physical samples and performing real-time rendering and editing with these materials. Jiaping Wang, Xin Tong 0001, Stephen Lin 0001, Zhouchen Lin, Yue Dong 0001, Baining Guo, Harry Shum |
ACM Trans. Graph. | 3 |
| 2008 | Modeling anisotropic surface reflectance with example-based microfacet synthesisabstractWe present a new technique for the visual modeling of spatiallyvarying anisotropic reflectance using data captured from a single view. Reflectance is represented using a microfacet-based BRDF which tabulates the facets' normal distribution (NDF) as a function of surface location. Data from a single view provides a 2D slice of the 4D BRDF at each surface point from which we fit a partial NDF. The fitted NDF is partial because the single view direction coupled with the set of light directions covers only a portion of the "half-angle" hemisphere. We complete the NDF at each point by applying a novel variant of texture synthesis using similar, overlapping partial NDFs from other points. Our similarity measure allows azimuthal rotation of partial NDFs, under the assumption that reflectance is spatially redundant but the local frame may be arbitrarily oriented. Our system includes a simple acquisition device that collects images over a 2D set of light directions by scanning a linear array of LEDs over a flat sample. Results demonstrate that our approach preserves spatial and directional BRDF details and generates a visually compelling match to measured materials. Jiaping Wang, Xin Tong 0001, John M. Snyder, Baining Guo |
ACM Trans. Graph. | 3 |
| 2007 | Rendering from compressed high dynamic range textures on programmable graphics hardwareabstractHigh dynamic range (HDR) images are increasingly employed in games and interactive applications for accurate rendering and illumination. One disadvantage of HDR images is their large data size; unfortunately, even though solutions have been proposed for future hardware, commodity graphics hardware today does not provide any native compression for HDR textures. Lvdi Wang, Peter-Pike J. Sloan, Li-Yi Wei, Xin Tong 0001, Baining Guo |
SI3D | 5 |
| 2007 | Incremental wavelet importance sampling for direct illuminationabstractMost of existing importance sampling methods for direct illumination exploit importance of illumination and surface BRDF. Without taking the visibility into consideration, they can not adaptively adjust the number of samples for each pixel during the sampling process. As a result, these methods tend to produce images with noise in partially occluded regions. In this paper, we introduce an incremental wavelet importance sampling approach, in which the visibility information is used to determine the number of samples at run time. For this purpose, we present a perceptual-based variance that is computed from visibility of samples. In the sampling process, the Halton sample points are incrementally warped for each pixel until the variance of warped samples converges. We demonstrate that our method is more efficient than existing importance sampling approaches. Hao-Da Huang, Yanyun Chen, Xin Tong 0001, Wencheng Wang 0001 |
VRST | 3 |
| 2007 | Accelerated Parallel Texture Optimization
Hao-Da Huang, Xin Tong 0001, Wencheng Wang 0001 |
J. Comput. Sci. Technol. | 2 |
| 2006 | Realistic, real-time rendering of ocean wavesabstractAbstract In computer games and other real‐time graphics applications, the ocean surface is typically modelled as a texture or bump‐mapped plane with simple lighting effects. This paper describes a system for realistically rendering the water surface in real time. Our system can render calm ocean waves with sophisticated lighting effects at 100 fps on a 680 MHz Pentium III with a GeForce 3 graphics card. The wave geometry is represented view‐dependently as a dynamic displacement map with surface detail described by a dynamic bump map. The illumination model includes reflection, refraction and Fresnel effects, which are critical for producing the look and feel of water. Copyright © 2006 John Wiley & Sons, Ltd. Luiz Velho 0001, Xin Tong 0001, Baining Guo, Harry Shum |
Comput. Animat. Virtual Worlds | 3 |
| 2006 | Appearance manifolds for modeling time-variant appearance of materialsabstractWe present a visual simulation technique called appearance manifolds for modeling the time-variant surface appearance of a material from data captured at a single instant in time. In modeling time-variant appearance, our method takes advantage of the key observation that concurrent variations in appearance over a surface represent different degrees of weathering. By reorganizing these various appearances in a manner that reveals their relative order with respect to weathering degree, our method infers spatial and temporal appearance properties of the material's weathering process that can be used to convincingly generate its weathered appearance at different points in time. Results with natural non-linear reflectance variations are demonstrated in applications such as visual simulation of weathering on 3D models, increasing and decreasing the weathering of real objects, and material transfer with weathering effects. Jiaping Wang, Xin Tong 0001, Stephen Lin 0001, Minghao Pan, Chao Wang 0063, Hujun Bao, Baining Guo, Harry Shum |
ACM Trans. Graph. | 2 |
| 2005 | Visual simulation of weathering by gamma-ton tracingabstractWeathering modeling introduces blemishes such as dirt, rust, cracks and scratches to virtual scenery. In this paper we present a visual stimulation technique that works well for a wide variety of weathering phenomena. Our technique, called γ-ton tracing, is based on a type of aging-inducing particles called γ-tons. Modeling a weathering effect with γ-ton tracing involves tracing a large number of γ-tons through the scene in a way similar to photon tracing and then generating the weathering effect using the recorded γ-ton transport information. With this technique, we can produce weathering effects that are customized to the scene geometry and tailored to the weathering sources. Several effects that are challenging for existing techniques can be readily captured by γ-ton tracing. These include global transport effects. or "stainbleeding". γ-ton tracing also enables visual simulations of complex multi-weathering effects. Lastly γ-ton tracing can generate weathering effects that not only involve texture changes but also large-scale geometry changes. We demonstrate our technique with a variety of examples. Yanyun Chen, Lin Xia, Tien-Tsin Wong, Xin Tong 0001, Hujun Bao, Baining Guo, Harry Shum |
ACM Trans. Graph. | 4 |
| 2005 | Modeling and rendering of quasi-homogeneous materialsabstractMany translucent materials consist of evenly-distributed heterogeneous elements which produce a complex appearance under different lighting and viewing directions. For these quasi-homogeneous materials, existing techniques do not address how to acquire their material representations from physical samples in a way that allows arbitrary geometry models to be rendered with these materials. We propose a model for such materials that can be readily acquired from physical samples. This material model can be applied to geometric models of arbitrary shapes, and the resulting objects can be efficiently rendered without expensive subsurface light transport simulation. In developing a material model with these attributes, we capitalize on a key observation about the subsurface scattering characteristics of quasi-homogeneous materials at different scales. Locally, the non-uniformity of these materials leads to inhomogeneous subsurface scattering. For subsurface scattering on a global scale, we show that a lengthy photon path through an even distribution of heterogeneous elements statistically resembles scattering in a homogeneous medium. This observation allows us to represent and measure the global light transport within quasi-homogeneous materials as well as the transfer of light into and out of a material volume through surface mesostructures. We demonstrate our technique with results for several challenging materials that exhibit sophisticated appearance features such as transmission of back illumination through surface mesostructures. Xin Tong 0001, Jiaping Wang, Stephen Lin 0001, Baining Guo, Harry Shum |
ACM Trans. Graph. | 1 |
| 2005 | Shell radiance texture functions
Yanyun Chen, Xin Tong 0001, Stephen Lin 0001, Jiaoying Shi, Baining Guo, Harry Shum |
Vis. Comput. | 3 |
| 2005 | Capturing and rendering geometry details for BTF-mapped surfaces
Jiaping Wang, Xin Tong 0001, John Snyder, Yanyun Chen, Baining Guo, Harry Shum |
Vis. Comput. | 2 |
| 2004 | Shell texture functionsabstractWe propose a texture function for realistic modeling and efficient rendering of materials that exhibit surface mesostructures, translucency and volumetric texture variations. The appearance of such complex materials for dynamic lighting and viewing directions is expensive to calculate and requires an impractical amount of storage to precompute. To handle this problem, our method models an object as a shell layer, formed by texture synthesis of a volumetric material sample, and a homogeneous inner core. To facilitate computation of surface radiance from the shell layer, we introduce the shell texture function (STF) which describes voxel irradiance fields based on precomputed fine-level light interactions such as shadowing by surface mesostructures and scattering of photons inside the object. Together with a diffusion approximation of homogeneous inner core radiance, the STF leads to fast and detailed raytraced renderings of complex materials. Yanyun Chen, Xin Tong 0001, Jiaping Wang, Stephen Lin 0001, Baining Guo, Harry Shum |
ACM Trans. Graph. | 2 |
| 2004 | Synthesis and Rendering of Bidirectional Texture Functions on Arbitrary SurfacesabstractThe bidirectional texture function (BTF) is a 6D function that describes the appearance of a real-world surface as a function of lighting and viewing directions. The BTF can model the fine-scale shadows, occlusions, and specularities caused by surface mesostructures. In this paper, we present algorithms for efficient synthesis of BTFs on arbitrary surfaces and for hardware-accelerated rendering. For both synthesis and rendering, a main challenge is handling the large amount of data in a BTF sample. To addresses this challenge, we approximate the BTF sample by a small number of 4D point appearance functions (PAFs) multiplied by 2D geometry maps. The geometry maps and PAFs lead to efficient synthesis and fast rendering of BTFs on arbitrary surfaces. For synthesis, a surface BTF can be generated by applying a texton-based sysnthesis algorithm to a small set of 2D geometry maps while leaving the companion 4D PAFs untouched. As for rendering, a surface BTF synthesized using geometry maps is well-suited for leveraging the programmable vertex and pixel shaders on the graphics hardware. We present a real-time BTF rendering algorithm that runs at the speed of about 30 frames/second on a mid-level PC with an ATI Radeon 8500 graphics card. We demonstrate the effectiveness of our synthesis and rendering algorithms using both real and synthetic BTF samples. Xinguo Liu, Jingdan Zhang, Xin Tong 0001, Baining Guo, Harry Shum |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2003 | Rendering driven depth reconstructionabstractPrevious work on image-based rendering suggests that there is a tradeoff between the number of images and the amount of geometry required for anti-aliased rendering. For instance, plenoptic sampling theory indicates that visually acceptable rendering can be achieved when the input images are undersampled, if sufficient depth information is available for all the pixels. In this paper, we propose a novel vision reconstruction approach, rendering-driven depth recovery, to recover the amount of geometry that is necessary for anti-aliased rendering. Our approach contrasts conventional stereo reconstruction in that we do not intend to accurately reconstruct the depth for each and every single pixel, leading to a very efficient reconstruction algorithm. Our algorithm uses a block-based multi-layer depth representation, and searches in the depth space based on the causality criterion, by detecting double images. Experiments show that rendering systems using our rendering driven depth recovery algorithm can synthesize satisfactory novel views efficiently by using 'just enough geometry' recovered from undersampled input images. Yin Li 0003, Xin Tong 0001, Chi-Keung Tang, Harry Shum |
ICASSP (4) | 2 |
| 2003 | View-dependent displacement mappingabstractSignificant visual effects arise from surface mesostructure, such as fine-scale shadowing, occlusion and silhouettes. To efficiently render its detailed appearance, we introduce a technique called view-dependent displacement mapping (VDM) that models surface displacements along the viewing direction. Unlike traditional displacement mapping, VDM allows for efficient rendering of self-shadows, occlusions and silhouettes without increasing the complexity of the underlying surface mesh. VDM is based on per-pixel processing, and with hardware acceleration it can render mesostructure with rich visual appearance in real time. Lifeng Wang 0001, Xin Tong 0001, Stephen Lin 0001, Shi-Min Hu 0001, Baining Guo, Harry Shum |
ACM Trans. Graph. | 3 |
| 2002 | Diffuse-Specular Separation and Depth Recovery from Image Sequences
Stephen Lin 0001, Yuanzhen Li, Sing Bing Kang, Xin Tong 0001, Harry Shum |
ECCV (3) | 4 |
| 2002 | Lighting Interpolation by Shadow Morphing Using Intrinsic LumigraphsabstractDensely-sampled image representations such as the light field or lumigraph have been effective in enabling photorealistic image synthesis. Unfortunately, lighting interpolation with such representations has not been shown to be possible without the use of accurate 3D geometry and surface reflectance properties. In this paper we propose an approach to image-based lighting interpolation that is based on estimates of geometry and shading from relatively few images. We decompose captured light fields at different lighting conditions into intrinsic images (reflectance and illumination images), and estimate view-dependent scene geometries using multi-view stereo. We call the resulting representation an intrinsic lumigraph. In the same way that the lumigraph uses geometry to permit more accurate view interpolation, the intrinsic lumigraph uses both geometry and intrinsic images to allow high-quality interpolation at different views and lighting conditions. Joint use of geometry and intrinsic images is effective in the computation of shadow masks for shadow prediction at new lighting conditions. We illustrate our approach with images of real scenes. Yasuyuki Matsushita, Sing Bing Kang, Stephen Lin 0001, Harry Shum, Xin Tong 0001 |
PG | 5 |
| 2002 | Rendering by Manifold Hopping
Harry Shum, Lifeng Wang 0001, Jinxiang Chai, Xin Tong 0001 |
Int. J. Comput. Vis. | 4 |
| 2002 | Layered lumigraph with LOD controlabstractAbstract The rendering performance of an image‐based rendering (IBR) system is determined by the number of images and the amount of geometrical information used. In this paper, we propose a layered lumigraph representation that is configured for optimized rendering performance based on the rendering platform (e.g., processor speed, memory) and output image resolution. The layered lumigraph is produced by classifying all pixels into a number of depth layers. Based on prior work on plenoptic sampling analysis, the layered lumigraph is constructed to achieve the same rendering quality along the minimum sampling curve by balancing the number of images and depth layers. For a given rendering platform, the best rendering performance can be obtained by choosing the optimal number of images and depth layers. Moreover, the layered lumigraph is capable of level‐of‐detail (LOD) control using the same image geometry trade‐off. Therefore, the layered lumigraph fully exploits the inherent constraints between the number of images, depth complexity, and output resolution. Finally, a backward warping technique is designed to efficiently render the layered lumigraph by taking advantage of texture mapping hardware. Copyright © 2002 John Wiley & Sons, Ltd. Xin Tong 0001, Jinxiang Chai, Harry Shum |
Comput. Animat. Virtual Worlds | 1 |
| 2002 | Synthesis of bidirectional texture functions on arbitrary surfacesabstractThe bidirectional texture function (BTF) is a 6D function that can describe textures arising from both spatially-variant surface reflectance and surface mesostructures. In this paper, we present an algorithm for synthesizing the BTF on an arbitrary surface from a sample BTF. A main challenge in surface BTF synthesis is the requirement of a consistent mesostructure on the surface, and to achieve that we must handle the large amount of data in a BTF sample. Our algorithm performs BTF synthesis based on surface textons, which extract essential information from the sample BTF to facilitate the synthesis. We also describe a general search strategy, called the k-coherent search, for fast BTF synthesis using surface textons. A BTF synthesized using our algorithm not only looks similar to the BTF sample in all viewing/lighthing conditions but also exhibits a consistent mesostructure when viewing and lighting directions change. Moreover, the synthesized BTF fits the target surface naturally and seamlessly. We demonstrate the effectiveness of our algorithm with sample BTFs from various sources, including those measured from real-world textures. Xin Tong 0001, Jingdan Zhang, Ligang Liu 0001, Baining Guo, Harry Shum |
ACM Trans. Graph. | 1 |
| 2000 | Plenoptic samplingabstractThis paper studies the problem of plenoptic sampling in image-based rendering (IBR). From a spectral analysis of light field signals and using the sampling theorem, we mathematically derive the analytical functions to determine the minimum sampling rate for light field rendering. The spectral support of a light field signal is bounded by the minimum and maximum depths only, no matter how complicated the spectral support might be because of depth variations in the scene. The minimum sampling rate for light field rendering is obtained by compacting the replicas of the spectral support of the sampled light field within the smallest interval. Given the minimum and maximum depths, a reconstruction filter with an optimal and constant depth can be designed to achieve anti-aliased light field rendering. Plenoptic sampling goes beyond the minimum number of images needed for anti-aliased light field rendering. More significantly, it utilizes the scene depth information to determine the minimum sampling curve in the joint image and geometry space. The minimum sampling curve quantitatively describes the relationship among three key elements in IBR systems: scene complexity (geometrical and textural information), the number of image samples, and the output resolution. Therefore, plenoptic sampling bridges the gap between image-based rendering and traditional geometry-based rendering. Experimental results demonstrate the effectiveness of our approach. Jinxiang Chai, S. C. Chan 0001, Harry Shum, Xin Tong 0001 |
SIGGRAPH | 4 |
| 1998 | Hardware assisted fast volume rendering with boundary enhancement
Xin Tong 0001, Zesheng Tang |
J. Comput. Sci. Technol. | 1 |
| 1996 | An intelligent multi-blackboard CAD system
Yunhe Pan, Weidong Geng, Xin Tong 0001 |
Artif. Intell. Eng. | 3 |