EDBT 2026 Demo / reviewers in the wild / expert
Baining Guo
dblp:44/3946
· DBLP profile ↗
201ranked-venue papers
13as first author
48since 2021 · last 2026
0000-0001-8349-8868ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 184 · 11 first-author · 40 since 2021Artificial intelligence and machine learning · 45 · 1 first-author · 32 since 2021Human-computer interaction and ubiquitous computing · 9 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MageBench: Bridging Large Multimodal Models to AgentsabstractRecent models like OpenAI’s O1 and DeepSeek’s R1, which utilize test-time scaling techniques, have demonstrated remarkable improvements in reasoning capabilities. We anticipate that in the near future, multimodal models will also experience significant breakthroughs in multimodal reasoning. This will require some highly challenging and specialized evaluations. As one of the most crucial real-world applications of multimodal models, visual agents require complex and comprehensive capabilities such as spatial planning and vision-in-the-chain type reasoning. These capabilities are currently lacking in existing multimodal benchmarks. In this paper, we introduce MageBench, a Multimodal reasoning benchmark built upon light-weight AGEnt environments that pose significant reasoning challenges and hold substantial practical value. The results show that only a few product-level models are better than random acting, and all of them are far inferior to human level. We analyze and summarize their errors and capability gaps in visual planning. Furthermore, we found that rule-based RL can significantly boost visual reasoning capabilities. This highlights that our benchmark could serve as a valuable testing ground for the emerging field of agentic RL research. Miaosen Zhang, Qi Dai 0001, Yifan Yang 0004, Jianmin Bao, Dongdong Chen 0001, Chong Luo 0001, Xin Geng 0001, Baining Guo |
WACV | 9 |
| 2025 | ART: Anonymous Region Transformer for Variable Multi-Layer Transparent Image GenerationabstractMulti-layer image generation is a fundamental task that enables users to isolate, select, and edit specific image layers, thereby revolutionizing interactions with generative models. In this paper, we introduce the Anonymous Region Transformer (ART), which facilitates the direct generation of variable multi-layer transparent images based on a global text prompt and an anonymous region layout. Inspired by Schema theory1, this anonymous region layout allows the generative model to autonomously determine which set of visual tokens should align with which text tokens, which is in contrast to the previously dominant semantic layout for the image generation task. In addition, the layer-wise region crop mechanism, which only selects the visual tokens belonging to each anonymous region, significantly reduces attention computation costs and enables the efficient generation of images with numerous distinct layers (e.g., 50+). When compared to the full attention approach, our method is over 12 times faster and exhibits fewer layer conflicts. Furthermore, we propose a high-quality multi-layer transparent image autoencoder that supports the direct encoding and decoding of the transparency of variable multi-layer images in a joint manner. By enabling precise control and scalable layer generation, ART establishes a new paradigm for interactive content creation. Yifan Pu, Zhicong Tang, Ruihong Yin, Haoxing Ye, Yuhui Yuan, Dong Chen 0003, Jianmin Bao, Sirui Zhang, Ji Li 0006, Xiu Li 0001, Zhouhui Lian, Gao Huang 0001, Baining Guo |
CVPR | 17 |
| 2025 | UniGraspTransformer: Simplified Policy Distillation for Scalable Dexterous Robotic GraspingabstractWe introduce UniGraspTransformer, a universal Transformer-based network for dexterous robotic grasping that simplifies training while enhancing scalability and performance. Unlike prior methods such as UniDexGrasp++, which require complex, multi-step training pipelines, UniGraspTransformer follows a streamlined process: first, dedicated policy networks are trained for individual objects using reinforcement learning to generate successful grasp trajectories; then, these trajectories are distilled into a single, universal network. Our approach enables UniGraspTransformer to scale effectively, incorporating up to 12 self-attention blocks for handling thousands of objects with diverse poses. Additionally, it generalizes well to both idealized and real-world inputs, evaluated in state-based and vision-based settings. Notably, UniGraspTransformer generates a broader range of grasping poses for objects in various shapes and orientations, resulting in more diverse grasp strategies. Experimental results demonstrate significant improvements over state-of-the-art, UniDexGrasp++, across various object categories, achieving success rate gains of 3.5%, 7.7%, and 10.1% on seen objects, unseen objects within seen categories, and completely unseen objects, respectively, in the vision-based setting. Project page: https://dexhand.github.io/UniGraspTransformer/. Fangyun Wei, Xiaohan Yi, Yaobo Liang, Chang Xu 0002, Yan Lu 0001, Jiaolong Yang, Baining Guo |
CVPR | 12 |
| 2025 | Incorporating Pre-Trained Diffusion Models in Solving the Schrödinger Bridge ProblemabstractThis paper aims to unify Score-based Generative Models (SGMs), also known as Diffusion models, and the Schrödinger Bridge (SB) problem through three reparameterization techniques: Iterative Proportional Mean-Matching (IPMM), Iterative Proportional Terminus-Matching (IPTM), and Iterative Proportional Flow-Matching (IPFM). These techniques significantly accelerate and stabilize the training of SB-based models. Furthermore, the paper introduces novel initialization strategies that use pre-trained SGMs to effectively train SB-based models. By using SGMs as initialization, we leverage the advantages of both SB-based models and SGMs, ensuring efficient training of SB-based models and further improving the performance of SGMs. Extensive experiments demonstrate the significant effectiveness and improvements of the proposed methods. We believe this work contributes to and paves the way for future research on generative models. Zhicong Tang, Tiankai Hang, Shuyang Gu, Dong Chen 0003, Baining Guo |
ECAI | 5 |
| 2025 | Improved Noise Schedule for Diffusion TrainingabstractDiffusion models have emerged as the de facto choice for generating high-quality visual signals across various domains. However, training a single model to predict noise across various levels poses significant challenges, necessitating numerous iterations and incurring significant computational costs. Various approaches, such as loss weighting strategy design and architectural refinements, have been introduced to expedite convergence and improve model performance. In this study, we propose a novel approach to design the noise schedule for enhancing the training of diffusion models. Our key insight is that the importance sampling of the logarithm of the Signal-to-Noise ratio ($\log \text{SNR}$), theoretically equivalent to a modified noise schedule, is particularly beneficial for training efficiency when increasing the sample frequency around $\log \text{SNR}=0$. This strategic sampling allows the model to focus on the critical transition point between signal dominance and noise dominance, potentially leading to more robust and accurate predictions.We empirically demonstrate the superiority of our noise schedule over the standard cosine schedule.Furthermore, we highlight the advantages of our noise schedule design on the ImageNet benchmark, showing that the designed schedule consistently benefits different prediction targets. Our findings contribute to the ongoing efforts to optimize diffusion models, potentially paving the way for more efficient and effective training paradigms in the field of generative AI. Tiankai Hang, Shuyang Gu, Jianmin Bao, Fangyun Wei, Dong Chen 0003, Xin Geng 0001, Baining Guo |
ICCV | 7 |
| 2025 | Gaussian Variation Field Diffusion for High-Fidelity Video-to-4D SynthesisabstractIn this paper, we present a novel framework for video-to-4D generation that creates high-quality dynamic 3D content from single video inputs. Direct 4D diffusion modeling is extremely challenging due to costly data construction and the high-dimensional nature of jointly representing 3D shape, appearance, and motion. We address these challenges by introducing a Direct 4DMesh-to-GS Variation Field VAE that directly encodes canonical Gaussian Splats (GS) and their temporal variations from 3D animation data without per-instance fitting, and compresses high-dimensional animations into a compact latent space. Building upon this efficient representation, we train a Gaussian Variation Field diffusion model with temporal-aware Diffusion Transformer conditioned on input videos and canonical GS. Trained on carefully-curated animatable 3D objects from the Objaverse dataset, our model demonstrates superior generation quality compared to existing methods. It also exhibits remarkable generalization to in-the-wild video inputs despite being trained exclusively on synthetic data, paving the way for generating high-quality animated 3D content. Project page: https://gvfdiffusion.github.io/. Bowen Zhang 0002, Sicheng Xu, Chuxin Wang, Jiaolong Yang, Feng Zhao 0004, Dong Chen 0003, Baining Guo |
ICCV | 7 |
| 2025 | Optimizing Large Language Model Training Using FP4 QuantizationabstractThe growing computational demands of training large language models (LLMs) necessitate more efficient methods. Quantized training presents a promising solution by enabling low-bit arithmetic operations to reduce these costs. While FP8 precision has demonstrated feasibility, leveraging FP4 remains a challenge due to significant quantization errors and limited representational capacity. This work introduces the first FP4 training framework for LLMs, addressing these challenges with two key innovations: a differentiable quantization estimator for precise weight updates and an outlier clamping and compensation strategy to prevent activation collapse. To ensure stability, the framework integrates a mixed-precision training scheme and vector-wise quantization. Experimental results demonstrate that our FP4 framework achieves accuracy comparable to BF16 and FP8, with minimal degradation, scaling effectively to 13B-parameter LLMs trained on up to 100B tokens. With the emergence of next-generation hardware supporting FP4, our framework sets a foundation for efficient ultra-low precision training. Yeyun Gong, Xiao Liu 0029, Guoshuai Zhao 0001, Ziyue Yang 0002, Baining Guo, Zhengjun Zha, Peng Cheng 0005 |
ICML | 6 |
| 2025 | VideoVLA: Video Generators Can Be Generalizable Robot ManipulatorsabstractGeneralization in robot manipulation is essential for deploying robots in open-world environments and advancing toward artificial general intelligence. While recent Vision-Language-Action (VLA) models leverage large pre-trained understanding models for perception and instruction following, their ability to generalize to novel tasks, objects, and settings remains limited. In this work, we present VideoVLA, a simple approach that explores the potential of transforming large video generation models into robotic VLA manipulators. Given a language instruction and an image, VideoVLA predicts an action sequence as well as the future visual outcomes. Built on a multi-modal Diffusion Transformer, VideoVLA jointly models video, language, and action modalities, using pre-trained video generative models for joint visual and action forecasting. Our experiments show that high-quality imagined futures correlate with reliable action predictions and task success, highlighting the importance of visual imagination in manipulation. VideoVLA demonstrates strong generalization, including imitating other embodiments' skills and handling novel objects. This dual-prediction strategy—forecasting both actions and their visual consequences—explores a paradigm shift in robot learning and unlocks generalization capabilities in manipulation systems. Yichao Shen 0001, Fangyun Wei, Zhiying Du, Yaobo Liang, Yan Lu 0001, Jiaolong Yang, Nanning Zheng 0001, Baining Guo |
NeurIPS | 8 |
| 2025 | VASA-3D: Lifelike Audio-Driven Gaussian Head Avatars from a Single ImageabstractWe propose VASA-3D, an audio-driven, single-shot 3D head avatar generator. This research tackles two major challenges: capturing the subtle expression details present in real human faces, and reconstructing an intricate 3D head avatar from a single portrait image. To accurately model expression details, VASA-3D leverages the motion latent of VASA-1, a method that yields exceptional realism and vividness in 2D talking heads. A critical element of our work is translating this motion latent to 3D, which is accomplished by devising a 3D head model that is conditioned on the motion latent. Customization of this model to a single image is achieved through an optimization framework that employs numerous video frames of the reference head synthesized from the input image. The optimization takes various training losses robust to artifacts and limited pose coverage in the generated training data. Our experiment shows that VASA-3D produces realistic 3D talking heads that cannot be achieved by prior art, and it supports the online generation of 512x512 free-viewpoint videos at up to 75 FPS, facilitating more immersive engagements with lifelike 3D avatars. Sicheng Xu, Jiaolong Yang, Yu Deng 0006, Stephen Lin 0001, Baining Guo |
NeurIPS | 7 |
| 2025 | VolumeDiffusion: Feed-forward text-to-3D generation with efficient volumetric encoderabstractThis work presents VolumeDiffusion, a novel feed-forward text-to-3D generation framework that directly synthesizes 3D objects from textual descriptions. It bypasses the conventional score distillation loss based or text-to-image-to-3D approaches. To scale up the training data for the diffusion model, a novel 3D volumetric encoder is developed to efficiently acquire feature volumes from multi-view images. The 3D volumes are then trained on a diffusion model for text-to-3D generation using a 3D U-Net. This research further addresses the challenges of inaccurate object captions and high-dimensional feature volumes. The proposed model, trained on the public Objaverse dataset, demonstrates promising outcomes in producing diverse and recognizable samples from text prompts. Notably, it empowers finer control over object part characteristics through textual cues, fostering model creativity by seamlessly combining multiple concepts within a single object. This research significantly contributes to the progress of 3D generation by introducing an efficient, flexible, and scalable representation methodology. Zhicong Tang, Shuyang Gu, Chunyu Wang 0001, Ting Zhang 0002, Jianmin Bao, Dong Chen 0003, Baining Guo |
Graph. Model. | 7 |
| 2025 | Swin3D: A Pretrained Transformer Backbone for 3D Indoor Scene UnderstandingabstractThe use of pretrained backbones with fine-tuning has shown success for 2D vision and natural language processing tasks, with advantages over task-specific networks. In this paper, we introduce a pretrained 3D backbone, called Swin3d, for 3D indoor scene understanding. We designed a 3D Swin Transformer as our backbone network, which enables efficient self-attention on sparse voxels with linear memory complexity, making the backbone scalable to large models and datasets. We also introduce a generalized contextual relative positional embedding scheme to capture various irregularities of point signals for improved network performance. We pretrained a large Swin3d model on a synthetic Structured3D dataset, which is an order of magnitude larger than the ScanNet dataset. Our model pretrained on the synthetic dataset not only generalizes well to downstream segmentation and detection on real 3D point datasets but also outperforms state-of-the-art methods on downstream tasks with +2.3 mIoU and +2.2 mIoU on S3DIS Area5 and 6-fold semantic segmentation, respectively, +1.8 mIoU on ScanNet segmentation (val), +1.9 [email protected] on ScanNet detection, and +8.1 [email protected] on S3DIS detection. A series of extensive ablation studies further validated the scalability, generality, and superior performance enabled by our approach. Yuxiao Guo 0001, Jian-Yu Xiong, Yang Liu 0014, Hao Pan 0001, Peng-Shuai Wang, Xin Tong 0001, Baining Guo |
Comput. Vis. Media | 8 |
| 2025 | CCA: collaborative competitive agents for image editing
Tiankai Hang, Shuyang Gu, Dong Chen 0003, Xin Geng 0001, Baining Guo |
Frontiers Comput. Sci. | 5 |
| 2024 | CCEdit: Creative and Controllable Video Editing via Diffusion ModelsabstractIn this paper, we present CCEdit, a versatile generative video editing framework based on diffusion models. Our approach employs a novel trident network structure that separates structure and appearance control, ensuring precise and creative editing capabilities. Utilizing the foundational ControlNet architecture, we maintain the structural integrity of the video during editing. The incorporation of an additional appearance branch enables users to exert fine-grained control over the edited key frame. These two side branches seamlessly integrate into the main branch, which is constructed upon existing text-to-image (T2I) generation models, through learnable temporal layers. The versatility of our framework is demonstrated through a diverse range of choices in both structure representations and personalized T2I models, as well as the option to provide the edited key frame. To facilitate comprehensive evaluation, we introduce the BalanceCC benchmark dataset, comprising 100 videos and 4 target prompts for each video. Our extensive user studies compare CCEdit with eight state-of-the-art video editing methods. The outcomes demonstrate CCEdit's substantial superiority over all other methods. Ruoyu Feng 0001, Wenming Weng, Yuhui Yuan, Jianmin Bao, Chong Luo 0001, Zhibo Chen 0001, Baining Guo |
CVPR | 8 |
| 2024 | InstructDiffusion: A Generalist Modeling Interface for Vision TasksabstractWe present InstructDiffusion, a unified and generic framework for aligning computer vision tasks with hu-man instructions. Unlike existing approaches that integrate prior knowledge and pre-define the output space (e.g., categories and coordinates) for each vision task, we cast diverse vision tasks into a human-intuitive image-manipulating pro-cess whose output space is a flexible and interactive pixel space. Concretely, the model is built upon the diffusion process and is trained to predict pixels according to user instructions, such as encircling the man's left shoulder in red or applying a blue mask to the left car. InstructDiffusion could handle a variety of vision tasks, including understanding tasks (such as segmentation and keypoint de-tection) and generative tasks (such as editing and enhance-ment) and outperforms prior methods on novel datasets. This represents a solid step towards a generalist modeling interface for vision tasks, advancing artificial general intelligence in the field of computer vision. Zigang Geng, Binxin Yang, Tiankai Hang, Shuyang Gu, Ting Zhang 0002, Jianmin Bao, Zheng Zhang 0022, Houqiang Li, Han Hu 0001, Dong Chen 0003, Baining Guo |
CVPR | 12 |
| 2024 | MicroCinema: A Divide-and-Conquer Approach for Text-to-Video GenerationabstractWe present MicroCinema, a straightforward yet effective framework for high-quality and coherent text-to-video generation. Unlike existing approaches that align text prompts with video directly, MicroCinema introduces a Divide-and-Conquer strategy which divides the text-to-video into a two-stage process: text-to-image generation and image&text-to-video generation. This strategy offers two significant advantages. a) It allows us to take full advantage of the recent advances in text-to-image models, such as Stable Diffusion, Midjourney, and DALLE, to generate photorealistic and highly detailed images. b) Leveraging the generated image, the model can allocate less focus to fine-grained appearance details, prioritizing the efficient learning of motion dynamics. To implement this strategy effectively, we introduce two core designs. First, we propose the Appearance Injection Network, enhancing the preservation of the appearance of the given image. Second, we introduce the appearance Noise Prior, a novel mechanism aimed at maintaining the capabilities of pre-trained 2D diffusion models. These design elements empower MicroCinema to generate high-quality videos with precise motion, guided by the provided text prompts. Extensive experiments demonstrate the superiority of the proposed framework. Concretely, MicroCinema achieves SOTA zero-shot FVD of 342.86 on UCF-JOJ and 377.40 on MSR-VTT. Jianmin Bao, Wenming Weng, Ruoyu Feng 0001, Dacheng Yin, Jingxu Zhang, Qi Dai 0001, Zhiyuan Zhao 0001, Chunyu Wang 0001, Yuhui Yuan, Xiaoyan Sun 0001, Chong Luo 0001, Baining Guo |
CVPR | 15 |
| 2024 | RodinHD: High-Fidelity 3D Avatar Generation with Diffusion Models
Bowen Zhang 0010, Yiji Cheng, Chunyu Wang 0001, Ting Zhang 0002, Jiaolong Yang, Yansong Tang, Feng Zhao 0004, Dong Chen 0003, Baining Guo |
ECCV (14) | 9 |
| 2024 | IRGen: Generative Modeling for Image Retrieval
Ting Zhang 0002, Dong Chen 0003, Yujing Wang 0002, Qi Chen 0009, Xing Xie 0001, Hao Sun 0015, Qi Zhang 0066, Fan Yang 0024, Mao Yang 0004, Qingmin Liao, Jingdong Wang 0001, Baining Guo |
ECCV (15) | 14 |
| 2024 | V-DETR: DETR with Vertex Relative Position Encoding for 3D Object DetectionabstractWe introduce a highly performant 3D object detector for point clouds using the DETR framework. The prior attempts all end up with suboptimal results because they fail to learn accurate inductive biases from the limited scale of training data. In particular, the queries often attend to points that are far away from the target objects, violating the locality principle in object detection. To address the limitation, we introduce a novel 3D Vertex Relative Position Encoding (3DV-RPE) method which computes position encoding for each point based on its relative position to the 3D boxes predicted by the queries in each decoder layer, thus providing clear information to guide the model to focus on points near the objects, in accordance with the principle of locality. Furthermore, we have systematically refined our pipeline, including data normalization, to better align with the task requirements. Our approach demonstrates remarkable performance on the demanding ScanNetV2 benchmark, showcasing substantial enhancements over the prior state-of-the-art CAGroup3D. Specifically, we achieve an increase in $AP_{25}$ from $75.1\%$ to $77.8\%$ and in ${AP}_{50}$ from $61.3\%$ to $66.0\%$. Yichao Shen 0001, Zigang Geng, Yuhui Yuan, Yutong Lin, Chunyu Wang 0001, Han Hu 0001, Nanning Zheng 0001, Baining Guo |
ICLR | 9 |
| 2024 | VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real TimeabstractWe introduce VASA, a framework for generating lifelike talking faces with appealing visual affective skills (VAS) given a single static image and a speech audio clip. Our premiere model, VASA-1, is capable of not only generating lip movements that are exquisitely synchronized with the audio, but also producing a large spectrum of facial nuances and natural head motions that contribute to the perception of authenticity and liveliness.
The core innovations include a diffusion-based holistic facial dynamics and head movement generation model that works in a face latent space, and the development of such an expressive and disentangled face latent space using videos.
Through extensive experiments including evaluation on a set of new metrics, we show that our method significantly outperforms previous methods along various dimensions comprehensively. Our method delivers high video quality with realistic facial and head dynamics and also supports the online generation of 512$\times$512 videos at up to 40 FPS with negligible starting latency.
It paves the way for real-time engagements with lifelike avatars that emulate human conversational behaviors. Sicheng Xu, Yuxiao Guo 0001, Jiaolong Yang, Zhenyu Zang, Xin Tong 0001, Baining Guo |
NeurIPS | 9 |
| 2024 | GaussianCube: A Structured and Explicit Radiance Representation for 3D Generative ModelingabstractWe introduce a radiance representation that is both structured and fully explicit and thus greatly facilitates 3D generative modeling. Existing radiance representations either require an implicit feature decoder, which significantly degrades the modeling power of the representation, or are spatially unstructured, making them difficult to integrate with mainstream 3D diffusion methods. We derive GaussianCube by first using a novel densification-constrained Gaussian fitting algorithm, which yields high-accuracy fitting using a fixed number of free Gaussians, and then rearranging these Gaussians into a predefined voxel grid via Optimal Transport. Since GaussianCube is a structured grid representation, it allows us to use standard 3D U-Net as our backbone in diffusion modeling without elaborate designs. More importantly, the high-accuracy fitting of the Gaussians allows us to achieve a high-quality representation with orders of magnitude fewer parameters than previous structured representations for comparable quality, ranging from one to two orders of magnitude. The compactness of GaussianCube greatly eases the difficulty of 3D generative modeling. Extensive experiments conducted on unconditional and class-conditioned object generation, digital avatar creation, and text-to-3D synthesis all show that our model achieves state-of-the-art generation results both qualitatively and quantitatively, underscoring the potential of GaussianCube as a highly accurate and versatile radiance representation for 3D generative modeling. Bowen Zhang 0010, Yiji Cheng, Jiaolong Yang, Chunyu Wang 0001, Feng Zhao 0004, Yansong Tang, Dong Chen 0003, Baining Guo |
NeurIPS | 8 |
| 2024 | Aligning Vision Models with Human Aesthetics in Retrieval: Benchmarks and AlgorithmsabstractModern vision models are trained on very large noisy datasets. While these models acquire strong capabilities, they may not follow the user's intent to output the desired results in certain aspects, e.g., visual aesthetic, preferred style, and responsibility. In this paper, we target the realm of visual aesthetics and aim to align vision models with human aesthetic standards in a retrieval system. Advanced retrieval systems usually adopt a cascade of aesthetic models as re-rankers or filters, which are limited to low-level features like saturation and perform poorly when stylistic, cultural or knowledge contexts are involved. We find that utilizing the reasoning ability of large language models (LLMs) to rephrase the search query and extend the aesthetic expectations can make up for this shortcoming. Based on the above findings, we propose a preference-based reinforcement learning method that fine-tunes the vision models to distill the knowledge from both LLMs reasoning and the aesthetic models to better align the vision models with human aesthetics. Meanwhile, with rare benchmarks designed for evaluating retrieval systems, we leverage large multi-modality model (LMM) to evaluate the aesthetic performance with their strong abilities. As aesthetic assessment is one of the most subjective tasks, to validate the robustness of LMM, we further propose a novel dataset named HPIR to benchmark the alignment with human aesthetics. Experiments demonstrate that our method significantly enhances the aesthetic behaviors of the vision models, under several metrics. We believe the proposed algorithm can be a general practice for aligning vision models with human values. Miaosen Zhang, Yixuan Wei, Zuxuan Wu, Ji Li 0006, Zheng Zhang 0022, Qi Dai 0001, Chong Luo 0001, Xin Geng 0001, Baining Guo |
NeurIPS | 11 |
| 2024 | Unsupervised Graphic Layout Grouping with TransformersabstractGraphic design conveys messages through the combination of text, images and other visual elements. Unstructured designs such as overloaded social media graphics may fail to communicate their intended messages effectively. To address this issue, layout grouping offers a solution by organizing design elements into perceptual groups. While most methods rely on heuristic Gestalt principles, they often lack the context modeling ability needed to handle complex layouts. In this work, we reformulate the layout grouping task as a set prediction problem. It uses Transformers to learn a set of group tokens at various hierarchies, enabling it to reason the membership of the elements more effectively. The self-attention mechanism in Transformers boosts its context modeling ability, which enables it to handle complex layouts more accurately. To reduce annotation costs, we also propose an unsupervised learning strategy that pre-trains on noisy pseudo-labels induced by a novel heuristic algorithm. This approach then bootstraps to self-refine the noisy labels, further improving the accuracy of our model. Our extensive experiments demonstrate the effectiveness of our method, which outperforms existing state-of-the-art approaches in terms of accuracy and efficiency. Danqing Huang, Chunyu Wang 0001, Mingxi Cheng, Ji Li 0006, Han Hu 0001, Xin Geng 0001, Baining Guo |
WACV | 8 |
| 2023 | PeCo: Perceptual Codebook for BERT Pre-training of Vision TransformersabstractThis paper explores a better prediction target for BERT pre-training of vision transformers. We observe that current prediction targets disagree with human perception judgment. This contradiction motivates us to learn a perceptual prediction target. We argue that perceptually similar images should stay close to each other in the prediction target space. We surprisingly find one simple yet effective idea: enforcing perceptual similarity during the dVAE training. Moreover, we adopt a self-supervised transformer model for deep feature extraction and show that it works well for calculating perceptual similarity. We demonstrate that such learned visual tokens indeed exhibit better semantic meanings, and help pre-training achieve superior transfer performance in various downstream tasks. For example, we achieve 84.5% Top-1 accuracy on ImageNet-1K with ViT-B backbone, outperforming the competitive method BEiT by +1.3% under the same pre-training epochs. Our approach also gets significant improvement on object detection and segmentation on COCO and semantic segmentation on ADE20K. Equipped with a larger backbone ViT-H, we achieve the state-of-the-art ImageNet accuracy (88.3%) among methods using only ImageNet-1K data. Xiaoyi Dong, Jianmin Bao, Ting Zhang 0002, Dongdong Chen 0001, Weiming Zhang 0001, Lu Yuan 0001, Dong Chen 0003, Fang Wen 0001, Nenghai Yu, Baining Guo |
AAAI | 10 |
| 2023 | MM-Diffusion: Learning Multi-Modal Diffusion Models for Joint Audio and Video GenerationabstractWe propose the first joint audio-video generation framework that brings engaging watching and listening experiences simultaneously, towards high-quality realistic videos. To generate joint audio-video pairs, we propose a novel Multi-Modal Diffusion model (i.e., MM-Diffusion), with two-coupled denoising autoencoders. In contrast to existing single-modal diffusion models, MM-Diffusion consists of a sequential multi-modal U-Net for a joint denoising process by design. Two subnets for audio and video learn to gradually generate aligned audio-video pairs from Gaussian noises. To ensure semantic consistency across modalities, we propose a novel random-shift based attention block bridging over the two subnets, which enables efficient cross-modal alignment, and thus reinforces the audio-video fidelity for each other. Extensive experiments show superior results in unconditional audio-video generation, and zeroshot conditional tasks (e.g., video-to-audio). In particular, we achieve the best FVD and FAD on Landscape and AIST++ dancing datasets. Turing tests of 10k votes further demonstrate dominant preferences for our model. The code and pre-trained models can be downloaded at https://github.com/researchmm/MM-Diffusion. Ludan Ruan, Yiyang Ma, Huan Yang 0005, Huiguo He, Bei Liu 0001, Jianlong Fu, Nicholas Jing Yuan, Qin Jin, Baining Guo |
CVPR | 9 |
| 2023 | RODIN: A Generative Model for Sculpting 3D Digital Avatars Using DiffusionabstractThis paper presents a 3D diffusion model that automatically generates 3D digital avatars represented as neural radiance fields (NeRFs). A significant challenge for 3D diffusion is that the memory and processing costs are prohibitive for producing high-quality results with rich details. To tackle this problem, we propose the roll-out diffusion network (RODIN), which takes a 3D NeRF model represented as multiple 2D feature maps and rolls out them onto a single 2D feature plane within which we perform 3D-aware diffusion. The RODIN model brings much-needed computational efficiency while preserving the integrity of 3D diffusion by using 3D-aware convolution that attends to projected features in the 2D plane according to their original relationships in 3D. We also use latent conditioning to orchestrate the feature generation with global coherence, leading to high-fidelity avatars and enabling semantic editing based on text prompts. Finally, we use hierarchical synthesis to further enhance details. The 3D avatars generated by our model compare favorably with those produced by existing techniques. We can generate highly detailed avatars with realistic hairstyles and facial hair. We also demonstrate 3D avatar generation from image or text, as well as text-guided editability. Tengfei Wang 0002, Bo Zhang 0025, Ting Zhang 0002, Shuyang Gu, Jianmin Bao, Tadas Baltrusaitis, Jingjing Shen, Dong Chen 0003, Fang Wen 0001, Qifeng Chen 0001, Baining Guo |
CVPR | 11 |
| 2023 | iCLIP: Bridging Image Classification and Contrastive Language-Image Pre-training for Visual RecognitionabstractThis paper presents a method that effectively combines two prevalent visual recognition methods, i.e., image classification and contrastive language-image pre-training, dubbed iCLIP. Instead of naïve multi-task learning that use two separate heads for each task, we fuse the two tasks in a deep fashion that adapts the image classification to share the same formula and the same model weights with the language-image pre-training. To further bridge these two tasks, we propose to enhance the category names in image classification tasks using external knowledge, such as their descriptions in dictionaries. Extensive experiments show that the proposed method combines the advantages of two tasks well: the strong discrimination ability in image classification tasks due to the clean category labels, and the good zero-shot ability in CLIP tasks ascribed to the richer semantics in the text descriptions. In particular, it reaches 82.9% top-1 accuracy on IN-1K, and mean-while surpasses CLIP by 1.8%, with similar model size, on zero-shot recognition of Kornblith 12-dataset benchmark. The code and models are publicly available at https://github.com/weiyx16/iCLIP. Yixuan Wei, Yue Cao 0001, Zheng Zhang 0022, Houwen Peng, Zhuliang Yao, Zhenda Xie, Han Hu 0001, Baining Guo |
CVPR | 8 |
| 2023 | Adaptive Frequency Filters As Efficient Global Token MixersabstractRecent vision transformers, large-kernel CNNs and MLPs have attained remarkable successes in broad vision tasks thanks to their effective information fusion in the global scope. However, their efficient deployments, especially on mobile devices, still suffer from noteworthy challenges due to the heavy computational costs of self-attention mechanisms, large kernels, or fully connected layers. In this work, we apply conventional convolution theorem to deep learning for addressing this and reveal that adaptive frequency filters can serve as efficient global token mixers. With this insight, we propose Adaptive Frequency Filtering (AFF) token mixer. This neural operator transfers a latent representation to the frequency domain via a Fourier transform and performs semantic-adaptive frequency filtering via an elementwise multiplication, which mathematically equals to a token mixing operation in the original latent space with a dynamic convolution kernel as large as the spatial resolution of this latent representation. We take AFF token mixers as primary neural operators to build a lightweight neural network, dubbed AFFNet. Extensive experiments demonstrate the effectiveness of our proposed AFF token mixer and show that AFFNet achieve superior accuracy and efficiency trade-offs compared to other lightweight network designs on broad visual tasks, including visual recognition and dense prediction tasks. Code is available at https://github.com/microsoft/TokenMixers. Zhipeng Huang 0014, Zhizheng Zhang 0004, Cuiling Lan, Zhengjun Zha, Yan Lu 0001, Baining Guo |
ICCV | 6 |
| 2023 | Efficient Diffusion Training via Min-SNR Weighting StrategyabstractDenoising diffusion models have been a mainstream approach for image generation, however, training these models often suffers from slow convergence. In this paper, we discovered that the slow convergence is partly due to conflicting optimization directions between timesteps. To address this issue, we treat the diffusion training as a multi-task learning problem, and introduce a simple yet effective approach referred to as Min-SNR-γ. This method adapts loss weights of timesteps based on clamped signal-to-noise ratios, which effectively balances the conflicts among timesteps. Our results demonstrate a significant improvement in converging speed, 3.4× faster than previous weighting strategies. It is also more effective, achieving a new record FID score of 2.06 on the ImageNet 256 × 256 benchmark using smaller architectures than that employed in previous state-of-the-art. The code is available at https://github.com/TiankaiHang/Min-SNR-Diffusion-Training. Tiankai Hang, Shuyang Gu, Jianmin Bao, Dong Chen 0003, Han Hu 0001, Xin Geng 0001, Baining Guo |
ICCV | 8 |
| 2023 | Improving CLIP Fine-tuning PerformanceabstractCLIP models have demonstrated impressively high zero-shot recognition accuracy, however, their fine-tuning performance on downstream vision tasks is sub-optimal. Contrarily, masked image modeling (MIM) performs exceptionally for fine-tuning on downstream tasks, despite the absence of semantic labels during training. We note that the two tasks have different ingredients: image-level targets versus token-level targets, a cross-entropy loss versus a regression loss, and full-image inputs versus partial-image inputs. To mitigate the differences, we introduce a classical feature map distillation framework, which can simultaneously inherit the semantic capability of CLIP models while constructing a task incorporated key ingredients of MIM. Experiments suggest that the feature map distillation approach significantly boosts the fine-tuning performance of CLIP models on several typical down-stream vision tasks. We also observe that the approach yields new CLIP representations which share some diagnostic properties with those of MIM. Furthermore, the feature map distillation approach generalizes to other pre-training models, such as DINO, DeiT and SwinV2-G, reaching a new record of 64.2 mAP on COCO object detection with +1.1 improvement. The code and models are publicly available at https://github.com/SwinTransformer/Feature-Distillation. Yixuan Wei, Han Hu 0001, Zhenda Xie, Zheng Zhang 0022, Yue Cao 0001, Jianmin Bao, Dong Chen 0003, Baining Guo |
ICCV | 9 |
| 2023 | Keynote Speaker: Reconstructing Reality: From Physical World to Virtual EnvironmentsabstractWith increasing availability of data in various forms from images, audio, video, 3D models, motion capture, simulation results, to satellite imagery, representative samples of the various phenomena constituting the world around us bring new opportunities and research challenges. Such availability of data has led to recent advances in data-driven modeling. However, most of the existing example-based synthesis methods offer empirical models and data reconstruction that may not provide an insightful understanding of the underlying process or may be limited to a subset of observations. In this talk, I present recent advances that integrate classical model-based methods and statistical learning techniques to tackle challenging problems that have not been previously addressed. These include flow reconstruction for traffic visualization, learning heterogeneous crowd behaviors from video, simultaneous estimation of deformation and elasticity parameters from images and video, and example-based multimodal display for VR systems. These approaches offer new insights for understanding complex collective behaviors, developing better models for complex dynamical systems from captured data, delivering more effective medical diagnosis and treatment, as well as cyber-manufacturing of customized apparel. I conclude by discussing some possible future directions and challenges Baining Guo |
VR | 1 |
| 2023 | Keynote Speaker: Digital Humans in Virtual EnvironmentabstractIn this talk, I will discuss recent technologies for creating digital humans to populate the virtual environment and their applications in video communication and collaboration. I will cover both photo-realistic humans (“digital doubles”) and 3D digital avatars and will share my thoughts about the coming content generation revolution brought by large GPT models (generative pretrained). Baining Guo |
VR | 1 |
| 2023 | RemoteTouch: Enhancing Immersive 3D Video Communication with Hand TouchabstractRecent research advance has significantly improved the visual real-ism of immersive 3D video communication. In this work we present a method to further enhance this immersive experience by adding the hand touch capability (“remote hand clapping”). In our system, each meeting participant sits in front of a large screen with haptic feedback. The local participant can reach his hand out to the screen and perform hand clapping with the remote participant as if the two participants were only separated by a virtual glass. A key challenge in emulating the remote hand touch is the realistic rendering of the participant's hand and arm as the hand touches the screen. When the hand is very close to the screen, the RGBD data required for realistic rendering is no longer available. To tackle this challenge, we present a dual representation of the user's hand. Our dual representation not only preserves the high-quality rendering usually found in recent image-based rendering systems but also allows the hand to reach to the screen. This is possible because the dual representation includes both an image-based model and a 3D geometry-based model, with the latter driven by a hand skeleton tracked by a side view camera. In addition, the dual representation provides a distance-based fusion of the image-based and 3D geometry-based models as the hand moves closer to the screen. The result is that the image-based and 3D geometry-based models mutually enhance each other, leading to realistic and seamless rendering. Our experiments demonstrate that our method provides consistent hand contact experience between remote users and improves the immersive experience of 3D video communication. Zhiqi Li 0004, Sicheng Xu, Jiaolong Yang, Xin Tong 0001, Baining Guo |
VR | 7 |
| 2023 | Language-Guided Face Animation by Recurrent StyleGAN-Based GeneratorabstractRecent works on language-guided image manipulation have shown great power of language in providing rich semantics, especially for face images. However, the other natural information, motions, in language is less explored. In this paper, we leverage the motion information and study a novel task, language-guided face animation, that aims to animate a static face image with the help of languages. To better utilize both semantics and motions from languages, we propose a simple yet effective framework. Specifically, we propose a recurrent motion generator to extract a series of semantic and motion information from the language and feed it along with visual information to a pre-trained StyleGAN to generate high-quality frames. To optimize the proposed framework, three carefully designed loss functions are proposed including a regularization loss to keep the face identity, a path length regularization loss to ensure motion smoothness, and a contrastive loss to enable video synthesis with various language guidance in one single model. Extensive experiments with both qualitative and quantitative evaluations on diverse domains (e.g.,human face, anime face, and dog face) demonstrate the superiority of our model in generating high-quality and realistic videos from one still image with the guidance of language. Code will be available athttps://github.com/researchmm/language-guided-animation.git Tiankai Hang, Huan Yang 0005, Bei Liu 0001, Jianlong Fu, Xin Geng 0001, Baining Guo |
IEEE Trans. Multim. | 6 |
| 2023 | Aggregated Contextual Transformations for High-Resolution Image InpaintingabstractImage inpainting that completes large free-form missing regions in images is a promising yet challenging task. State-of-the-art approaches have achieved significant progress by taking advantage of generative adversarial networks (GAN). However, these approaches can suffer from generating distorted structures and blurry textures in high-resolution images (e.g., 512×512). The challenges mainly drive from (1) image content reasoning from distant contexts, and (2) fine-grained texture synthesis for a large missing region. To overcome these two challenges, we propose an enhanced GAN-based model, named Aggregated COntextual-Transformation GAN (AOT-GAN), for high-resolution image inpainting. Specifically, to enhance context reasoning, we construct the generator of AOT-GAN by stacking multiple layers of a proposed AOT block. The AOT blocks aggregate contextual transformations from various receptive fields, allowing to capture both informative distant image contexts and rich patterns of interest for context reasoning. For improving texture synthesis, we enhance the discriminator of AOT-GAN by training it with a tailored mask-prediction task. Such a training objective forces the discriminator to distinguish the detailed appearances of real and synthesized patches, and in turn facilitates the generator to synthesize clear textures. Extensive comparisons on Places2, the most challenging benchmark with 1.8 million high-resolution images of 365 complex scenes, show that our model outperforms the state-of-the-art. A user study including more than 30 subjects further validates the superiority of AOT-GAN. We further evaluate the proposed AOT-GAN in practical applications, e.g., logo removal, face editing, and object removal. Results show that our model achieves promising completions in the real world. We release codes and models in https://github.com/researchmm/AOT-GAN-for-Inpainting. Yanhong Zeng, Jianlong Fu, Hongyang Chao, Baining Guo |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | Fine-Grained Image Style Transfer with Visual Transformers
Huan Yang 0005, Jianlong Fu, Toshihiko Yamasaki, Baining Guo |
ACCV (3) | 5 |
| 2022 | Protecting Celebrities from DeepFake with Identity Consistency TransformerabstractIn this work we propose Identity Consistency Transformer, a novel face forgery detection method that focuses on high-level semantics, specifically identity information, and detecting a suspect face by finding identity inconsistency in inner and outer face regions. The Identity Consistency Transformer incorporates a consistency loss for identity consistency determination. We show that Identity Consistency Transformer exhibits superior generalization ability not only across different datasets but also across various types of image degradation forms found in real-world applications including deepfake videos. The Identity Consistency Transformer can be easily enhanced with additional identity information when such information is available, and for this reason it is especially well-suited for detecting face forgeries involving celebrities.11Code will be released at https://github.com/LightDXY/ICT_DeepFake Xiaoyi Dong, Jianmin Bao, Dongdong Chen 0001, Ting Zhang 0002, Weiming Zhang 0001, Nenghai Yu, Dong Chen 0003, Fang Wen 0001, Baining Guo |
CVPR | 9 |
| 2022 | CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped WindowsabstractWe present CSWin Transformer, an efficient and effective Transformer-based backbone for general-purpose vision tasks. A challenging issue in Transformer design is that global self-attention is very expensive to compute whereas local self-attention often limits the field of interactions of each token. To address this issue, we develop the Cross-Shaped Window self-attention mechanism for computing self-attention in the horizontal and vertical stripes in parallel that form a cross-shaped window, with each stripe obtained by splitting the input feature into stripes of equal width. We provide a mathematical analysis of the effect of the stripe width and vary the stripe width for different layers of the Transformer network which achieves strong modeling capability while limiting the computation cost. We also introduce Locally-enhanced Positional Encoding (LePE), which handles the local positional information better than existing encoding schemes. LePE naturally supports arbitrary input resolutions, and is thus especially effective and friendly for downstream tasks. Incorporated with these designs and a hierarchical structure, CSWin Transformer demonstrates competitive performance on common vision tasks. Specifically, it achieves 85.4% Top-1 accuracy on ImageNet-1K without any extra training data or label, 53.9 box AP and 46.4 mask AP on the COCO detection task, and 52.2 mIOU on the ADE20K semantic segmentation task, surpassing previous state-of-the-art Swin Transformer backbone by +1.2, +2.0, +1.4, and +2.0 respectively under the similar FLOPs setting. By further pretraining on the larger dataset ImageNet-21K, we achieve 87.5% Top-1 accuracy on ImageNet-1K and high segmentation performance on ADE20K with 55.7 mIoU.11Code and pretrain model is available at https://github.com/microsoft/CSWin-Transformer Xiaoyi Dong, Jianmin Bao, Dongdong Chen 0001, Weiming Zhang 0001, Nenghai Yu, Lu Yuan 0001, Dong Chen 0003, Baining Guo |
CVPR | 8 |
| 2022 | Vector Quantized Diffusion Model for Text-to-Image SynthesisabstractWe present the vector quantized diffusion (VQ-Diffusion) model for text-to-image generation. This method is based on a vector quantized variational autoencoder (VQ-VAE) whose latent space is modeled by a conditional variant of the recently developed Denoising Diffusion Probabilistic Model (DDPM). We find that this latent-space method is well-suited for text-to-image generation tasks because it not only eliminates the unidirectional bias with existing methods but also allows us to incorporate a mask-and-replace diffusion strategy to avoid the accumulation of errors, which is a serious problem with existing methods. Our experiments show that the VQ-Diffusion produces significantly better text-to-image generation results when compared with conventional autoregressive (AR) models with similar numbers of parameters. Compared with previous GAN-based text-to-image methods, our VQ-Diffusion can handle more complex scenes and improve the synthesized image quality by a large margin. Finally, we show that the image generation computation in our method can be made highly efficient by reparameterization. With traditional AR methods, the text-to-image generation time increases linearly with the output image resolution and hence is quite time consuming even for normal size images. The VQ-Diffusion allows us to achieve a better trade-off between quality and speed. Our experiments indicate that the VQ-Diffusion model with the reparameterization is fifteen times faster than traditional AR methods while achieving a better image quality. The code and models are available at https://github.com/cientgu/VQ-Diffusion. Shuyang Gu, Dong Chen 0003, Jianmin Bao, Fang Wen 0001, Bo Zhang 0025, Dongdong Chen 0001, Lu Yuan 0001, Baining Guo |
CVPR | 8 |
| 2022 | Swin Transformer V2: Scaling Up Capacity and ResolutionabstractWe present techniques for scaling Swin Transformer [35] up to 3 billion parameters and making it capable of training with images of up to 1,536x1,536 resolution. By scaling up capacity and resolution, Swin Transformer sets new records on four representative vision benchmarks: 84.0% top-1 accuracy on ImageNet- V2 image classification, 63.1 / 54.4 box / mask mAP on COCO object detection, 59.9 mIoU on ADE20K semantic segmentation, and 86.8% top-1 accuracy on Kinetics-400 video action classification. We tackle issues of training instability, and study how to effectively transfer models pre-trained at low resolutions to higher resolution ones. To this aim, several novel technologies are proposed: 1) a residual post normalization technique and a scaled cosine attention approach to improve the stability of large vision models; 2) a log-spaced continuous position bias technique to effectively transfer models pre-trained at low-resolution images and windows to their higher-resolution counterparts. In addition, we share our crucial implementation details that lead to significant savings of GPU memory consumption and thus make it feasi-ble to train large vision models with regular GPUs. Using these techniques and self-supervised pre-training, we suc-cessfully train a strong 3 billion Swin Transformer model and effectively transfer it to various vision tasks involving high-resolution images or windows, achieving the state-of-the-art accuracy on a variety of benchmarks. Code is avail-able at https://github.com/microsoft/Swin-Transformer. Han Hu 0001, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Yue Cao 0001, Zheng Zhang 0022, Li Dong 0004, Furu Wei, Baining Guo |
CVPR | 12 |
| 2022 | Advancing High-Resolution Video-Language Representation with Large-Scale Video TranscriptionsabstractWe study joint video and language (VL) pretraining to enable cross-modality learning and benefit plentiful downstream VL tasks. Existing works either extract low-quality video features or learn limited text embedding, while neglecting that high-resolution videos and diversified semantics can significantly improve cross-modality learning. In this paper, we propose a novel High-resolution and Diversified VIdeo-LAnguage pre-training model (HD-VILA) for many visual tasks. In particular, we collect a large dataset with two distinct properties: 1) the first high-resolution dataset including 371.5k hours of 720p videos, and 2) the most diversified dataset covering 15 popular YouTube categories. To enable VL pre-training, we jointly optimize the HD-VILA model by a hybrid Transformer that learns rich spatiotemporal features, and a multimodal Transformer that enforces interactions of the learned video features with diversified texts. Our pre-training model achieves new state-of-the-art results in 10 VL understanding tasks and 2 more novel text-to-visual generation tasks. For example, we outperform SOTA models with relative increases of 40.4% R@1 in zero-shot MSR-VTT text-to-video retrieval task, and 55.4% in high-resolution dataset LSMDC. The learned VL embedding is also effective in generating visually pleasing and semantically relevant results in text-to-visual editing and super-resolution tasks. Hongwei Xue, Tiankai Hang, Yanhong Zeng, Yuchong Sun, Bei Liu 0001, Huan Yang 0005, Jianlong Fu, Baining Guo |
CVPR | 8 |
| 2022 | StyleSwin: Transformer-based GAN for High-resolution Image GenerationabstractDespite the tantalizing success in a broad of vision tasks, transformers have not yet demonstrated on-par ability as ConvNets in high-resolution image generative modeling. In this paper, we seek to explore using pure transformers to build a generative adversarial network for high-resolution image synthesis. To this end, we believe that local attention is crucial to strike the balance between computational efficiency and modeling capacity. Hence, the proposed generator adopts Swin transformer in a style-based architecture. To achieve a larger receptive field, we propose double attention which simultaneously leverages the context of the local and the shifted windows, leading to improved generation quality. Moreover, we show that offering the knowledge of the absolute position that has been lost in window-based transformers greatly benefits the generation quality. The proposed StyleSwin is scalable to high resolutions, with both the coarse geometry and fine structures benefit from the strong expressivity of transformers. However, blocking artifacts occur during high-resolution synthesis because performing the local attention in a block-wise manner may break the spatial coherency. To solve this, we empirically investigate various solutions, among which we find that employing a wavelet discriminator to examine the spectral discrepancy effectively suppresses the artifacts. Extensive experiments show the superiority over prior transformer-based GANs, especially on high resolutions, e.g.,$1024 \times$1024. The StyleSwin, without complex training strategies, excels over StyleGAN on CelebA-HQ 1024, and achieves on-par performance on FFHQ-1024, proving the promise of using transformers for high-resolution image generation. The code and pretrained models are available at https://github.com/microsoft/StyleSwin. Bowen Zhang 0010, Shuyang Gu, Bo Zhang 0025, Jianmin Bao, Dong Chen 0003, Fang Wen 0001, Baining Guo |
CVPR | 8 |
| 2022 | Implicit Conversion of Manifold B-Rep Solids by Neural Halfspace RepresentationabstractWe present a novel implicit representation ---neural halfspace representation(NH-Rep), to convert manifold B-Rep solids to implicit representations. NH-Rep is a Boolean tree built on a set of implicit functions represented by the neural network, and the composite Boolean function is capable of representing solid geometry while preserving sharp features. We propose an efficient algorithm to extract the Boolean tree from a manifold B-Rep solid and devise a neural network-based optimization approach to compute the implicit functions. We demonstrate the high quality offered by our conversion algorithm on ten thousand manifold B-Rep CAD models that contain various curved patches including NURBS, and the superiority of our learning approach over other representative implicit conversion algorithms in terms of surface reconstruction, sharp feature preservation, signed distance field approximation, and robustness to various surface geometry, as well as a set of applications supported by NH-Rep. Hao-Xiang Guo 0001, Yang Liu 0014, Hao Pan 0001, Baining Guo |
ACM Trans. Graph. | 4 |
| 2022 | ComplexGen: CAD reconstruction by B-rep chain complex generationabstractWe view the reconstruction of CAD models in the boundary representation (B-Rep) as the detection of geometric primitives of different orders, i.e. , vertices, edges and surface patches, and the correspondence of primitives, which are holistically modeled as a chain complex, and show that by modeling such comprehensive structures more complete and regularized reconstructions can be achieved. We solve the complex generation problem in two steps. First, we propose a novel neural framework that consists of a sparse CNN encoder for input point cloud processing and a tri-path transformer decoder for generating geometric primitives and their mutual relationships with estimated probabilities. Second, given the probabilistic structure predicted by the neural network, we recover a definite B-Rep chain complex by solving a global optimization maximizing the likelihood under structural validness constraints and applying geometric refinements. Extensive tests on large scale CAD datasets demonstrate that the modeling of B-Rep chain complex structure enables more accurate detection for learning and more constrained reconstruction for optimization, leading to structurally more faithful and complete CAD B-Rep models than previous results. Hao-Xiang Guo 0001, Hao Pan 0001, Yang Liu 0014, Xin Tong 0001, Baining Guo |
ACM Trans. Graph. | 6 |
| 2022 | VirtualCube: An Immersive 3D Video Communication SystemabstractThe VirtualCube system is a 3D video conference system that attempts to overcome some limitations of conventional technologies. The key ingredient is VirtualCube, an abstract representation of a real-world cubicle instrumented with RGBD cameras for capturing the user's 3D geometry and texture. We design VirtualCube so that the task of data capturing is standardized and significantly simplified, and everything can be built using off-the-shelf hardware. We use VirtualCubes as the basic building blocks of a virtual conferencing environment, and we provide each VirtualCube user with a surrounding display showing life-size videos of remote participants. To achieve real-time rendering of remote participants, we develop the V-Cube View algorithm, which uses multi-view stereo for more accurate depth estimation and Lumi-Net rendering for better rendering quality. The VirtualCube system correctly preserves the mutual eye gaze between participants, allowing them to establish eye contact and be aware of who is visually paying attention to them. The system also allows a participant to have side discussions with remote participants as if they were in the same room. Finally, the system sheds lights on how to support the shared space of work items (e.g., documents and applications) and track participants' visual attention to work items. Jiaolong Yang, Zhen Liu 0031, Ruicheng Wang, Xin Tong 0001, Baining Guo |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2021 | Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsabstractThis paper presents a new vision Transformer, called Swin Transformer, that capably serves as a general-purpose backbone for computer vision. Challenges in adapting Transformer from language to vision arise from differences between the two domains, such as large variations in the scale of visual entities and the high resolution of pixels in images compared to words in text. To address these differences, we propose a hierarchical Transformer whose representation is computed with Shifted windows. The shifted windowing scheme brings greater efficiency by limiting self-attention computation to non-overlapping local windows while also allowing for cross-window connection. This hierarchical architecture has the flexibility to model at various scales and has linear computational complexity with respect to image size. These qualities of Swin Transformer make it compatible with a broad range of vision tasks, including image classification (87.3 top-1 accuracy on ImageNet-1K) and dense prediction tasks such as object detection (58.7 box AP and 51.1 mask AP on COCO test-dev) and semantic segmentation (53.5 mIoU on ADE20K val). Its performance surpasses the previous state-of-the-art by a large margin of +2.7 box AP and +2.6 mask AP on COCO, and +3.2 mIoU on ADE20K, demonstrating the potential of Transformer-based models as vision backbones. The hierarchical design and the shifted window approach also prove beneficial for all-MLP architectures. The code and models are publicly available at https://github.com/microsoft/Swin-Transformer. Yutong Lin, Yue Cao 0001, Han Hu 0001, Yixuan Wei, Zheng Zhang 0022, Stephen Lin 0001, Baining Guo |
ICCV | 8 |
| 2021 | Learning and Exploring Motor Skills with Spacetime BoundsabstractAbstract Equipping characters with diverse motor skills is the current bottleneck of physics‐based character animation. We propose a Deep Reinforcement Learning (DRL) framework that enables physics‐based characters to learn and explore motor skills from reference motions. The key insight is to use loose space‐time constraints, termed spacetime bounds, to limit the search space in an early termination fashion. As we only rely on the reference to specify loose spacetime bounds, our learning is more robust with respect to low quality references. Moreover, spacetime bounds are hard constraints that improve learning of challenging motion segments, which can be ignored by imitation‐only learning. We compare our method with state‐of‐the‐art tracking‐based DRL methods. We also show how to guide style exploration within the proposed framework. Li-Ke Ma, Zeshi Yang, Xin Tong 0001, Baining Guo, KangKang Yin |
Comput. Graph. Forum | 4 |
| 2021 | Deep Reflectance Scanning: Recovering Spatially-varying Material Appearance from a Flash-lit Video SequenceabstractAbstract In this paper we present a novel method for recovering high‐resolution spatially‐varying isotropic surface reflectance of a planar exemplar from a flash‐lit close‐up video sequence captured with a regular hand‐held mobile phone. We do not require careful calibration of the camera and lighting parameters, but instead compute a per‐pixel flow map using a deep neural network to align the input video frames. For each video frame, we also extract the reflectance parameters, and warp the neural reflectance features directly using the per‐pixel flow, and subsequently pool the warped features. Our method facilitates convenient hand‐held acquisition of spatially‐varying surface reflectance with commodity hardware by non‐expert users. Furthermore, our method enables aggregation of reflectance features from surface points visible in only a subset of the captured video frames, enabling the creation of high‐resolution reflectance maps that exceed the native camera resolution. We demonstrate and validate our method on a variety of synthetic and real‐world spatially‐varying materials. Yue Dong 0001, Pieter Peers, Baining Guo |
Comput. Graph. Forum | 4 |
| 2021 | Data-Driven 3D Neck Modeling and AnimationabstractIn this article, we present a data-driven approach for modeling and animation of 3D necks. Our method is based on a new neck animation model that decomposes the neck animation into local deformation caused by larynx motion and global deformation driven by head poses, facial expressions, and speech. A skinning model is introduced for modeling local deformation and underlying larynx motions, while the global neck deformation caused by each factor is modeled by its corrective blendshape set, respectively. Based on this neck model, we introduce a regression method to drive the larynx motion and neck deformation from speech. Both the neck model and the speech regressor are learned from a dataset of 3D neck animation sequences captured from different identities. Our neck model significantly improves the realism of facial animation and allows users to easily create plausible neck animations from speech and facial expressions. We verify our neck model and demonstrate its advantages in 3D neck tracking and animation. Chengwei Zheng, Feng Xu 0005, Xin Tong 0001, Baining Guo |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2020 | Face X-Ray for More General Face Forgery DetectionabstractIn this paper we propose a novel image representation called face X-ray for detecting forgery in face images. The face X-ray of an input face image is a greyscale image that reveals whether the input image can be decomposed into the blending of two images from different sources. It does so by showing the blending boundary for a forged image and the absence of blending for a real image. We observe that most existing face manipulation methods share a common step: blending the altered face into an existing background image. For this reason, face X-ray provides an effective way for detecting forgery generated by most existing face manipulation algorithms. Face X-ray is general in the sense that it only assumes the existence of a blending step and does not rely on any knowledge of the artifacts associated with a specific face manipulation technique. Indeed, the algorithm for computing face X-ray can be trained without fake images generated by any of the state-of-the-art face manipulation methods. Extensive experiments show that face X-ray remains effective when applied to forgery generated by unseen face manipulation techniques, while most existing face forgery detection or deepfake detection algorithms experience a significant performance drop. Lingzhi Li 0002, Jianmin Bao, Ting Zhang 0002, Hao Yang 0036, Dong Chen 0003, Fang Wen 0001, Baining Guo |
CVPR | 7 |
| 2020 | Learning Texture Transformer Network for Image Super-ResolutionabstractWe study on image super-resolution (SR), which aims to recover realistic textures from a low-resolution (LR) image. Recent progress has been made by taking high-resolution images as references (Ref), so that relevant textures can be transferred to LR images. However, existing SR approaches neglect to use attention mechanisms to transfer high-resolution (HR) textures from Ref images, which limits these approaches in challenging cases. In this paper, we propose a novel Texture Transformer Network for Image Super-Resolution (TTSR), in which the LR and Ref images are formulated as queries and keys in a transformer, respectively. TTSR consists of four closely-related modules optimized for image generation tasks, including a learnable texture extractor by DNN, a relevance embedding module, a hard-attention module for texture transfer, and a soft-attention module for texture synthesis. Such a design encourages joint feature learning across LR and Ref images, in which deep feature correspondences can be discovered by attention, and thus accurate texture features can be transferred. The proposed texture transformer can be further stacked in a cross-scale way, which enables texture recovery from different levels (e.g., from 1x to 4x magnification). Extensive experiments show that TTSR achieves significant improvements over state-of-the-art approaches on both quantitative and qualitative evaluations. Fuzhi Yang, Huan Yang 0005, Jianlong Fu, Baining Guo |
CVPR | 5 |
| 2019 | Capturing Piecewise SVBRDFs with Content Aware Lighting
Xiao Li 0030, Peiran Ren, Yue Dong 0001, Gang Hua 0001, Xin Tong 0001, Baining Guo |
CGI | 6 |
| 2019 | Learning Pyramid-Context Encoder Network for High-Quality Image InpaintingabstractHigh-quality image inpainting requires filling missing regions in a damaged image with plausible content. Existing works either fill the regions by copying high-resolution patches or generating semantically-coherent patches from region context, while neglecting the fact that both visual and semantic plausibility are highly-demanded. In this paper, we propose a Pyramid-context Encoder Network (denoted as PEN-Net) for image inpainting by deep generative models. The proposed PEN-Net is built upon a U-Net structure with three tailored components, ie., a pyramid-context encoder, a multi-scale decoder, and an adversarial training loss. First, we adopt a U-Net as backbone which can encode the context of an image from high-resolution pixels into high-level semantic features, and decode the features reversely. Second, we propose a pyramid-context encoder, which progressively learns region affinity by attention from a high-level semantic feature map, and transfers the learned attention to its adjacent high-resolution feature map. As the missing content can be filled by attention transfer from deep to shallow in a pyramid fashion, both visual and semantic coherence for image inpainting can be ensured. Third, we further propose a multi-scale decoder with deeply-supervised pyramid losses and an adversarial loss. Such a design not only results in fast convergence in training, but more realistic results in testing. Extensive experiments on a broad range of datasets shows the superior performance of the proposed network. Yanhong Zeng, Jianlong Fu, Hongyang Chao, Baining Guo |
CVPR | 4 |
| 2019 | Towards Robust Direction Invariance in Character AnimationabstractAbstract In character animation, direction invariance is a desirable property. That is, a pose facing north and the same pose facing south are considered the same; a character that can walk to the north is expected to be able to walk to the south in a similar style. To achieve such direction invariance, the current practice is to remove the facing direction's rotation around the vertical axis before further processing. Such a scheme, however, is not robust for rotational behaviors in the sagittal plane. In search of a smooth scheme to achieve direction invariance, we prove that in general a singularity free scheme does not exist. We further connect the problem with the hairy ball theorem, which is better‐known to the graphics community. Due to the nonexistence of a singularity free scheme, a general solution does not exist and we propose a remedy by using a properly‐chosen motion direction that can avoid singularities for specific motions at hand. We perform comparative studies using two deep‐learning based methods, one builds kinematic motion representations and the other learns physics‐based controls. The results show that with our robust direction invariant features, both methods can achieve better results in terms of learning speed and/or final quality. We hope this paper can not only boost performance for character animation methods, but also help related communities currently not fully aware of the direction invariance problem to achieve more robust results. Li-Ke Ma, Zeshi Yang, Baining Guo, KangKang Yin |
Comput. Graph. Forum | 3 |
| 2018 | Compressing Neural Networks using the Variational Information BottleneckabstractNeural networks can be compressed to reduce memory and computational requirements, or to increase accuracy by facilitating the use of a larger base architecture. In this paper we focus on pruning individual neurons, which can simultaneously trim model size, FLOPs, and run-time memory. To improve upon the performance of existing compression algorithms we utilize the information bottleneck principle instantiated via a tractable variational bound. Minimization of this information theoretic bound reduces the redundancy between adjacent layers by aggregating useful information into a subset of neurons that can be preserved. In contrast, the activations of disposable neurons are shut off via an attractive form of sparse regularization that emerges naturally from this framework, providing tangible advantages over traditional sparsity penalties without contributing additional tuning parameters to the energy landscape. We demonstrate state-of-the-art compression rates across an array of datasets and network architectures. Bin Dai 0008, Baining Guo, David P. Wipf |
ICML | 3 |
| 2018 | 3D cartoon face rigging from sparse examples
Jingyong Zhou, Hsiang-Tao Wu, Zicheng Liu 0001, Xin Tong 0001, Baining Guo |
Vis. Comput. | 5 |
| 2017 | Delay-Rate-Distortion Optimization for Cloud Gaming With Hybrid StreamingabstractCloud gaming as the emerging game service has attracted significant attention. However, traditional video streaming approach suffers from high bandwidth consumption, and traditional graphics streaming approach requires a long initial period to download game models. In this paper, we propose a novel hybrid streaming framework, jointly applying video streaming and graphics streaming to provide a high-quality gaming experience. In the proposed framework, cloud servers not only transmit the encoded video frames but also progressively transmit the graphics data, which are used to render a game frame to provide an additional reference to the video encoder. Based on the proposed framework, we investigate the delay-rate-distortion optimization problem, where the source rate between the video stream and the graphics stream is optimized to minimize the overall distortion under the bandwidth and response delay constraints. The experimental results demonstrate that the proposed hybrid streaming can achieve the lowest distortion under the constraints of bandwidth and response delay, compared with the traditional video streaming and graphics streaming. Xiaoming Nan, Xun Guo 0002, Yan Lu 0001, Ling Guan, Shipeng Li 0001, Baining Guo |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2016 | TopicPanorama: A Full Picture of Relevant TopicsabstractThis paper presents a visual analytics approach to analyzing a full picture of relevant topics discussed in multiple sources, such as news, blogs, or micro-blogs. The full picture consists of a number of common topics covered by multiple sources, as well as distinctive topics from each source. Our approach models each textual corpus as a topic graph. These graphs are then matched using a consistent graph matching method. Next, we develop a level-of-detail (LOD) visualization that balances both readability and stability. Accordingly, the resulting visualization enhances the ability of users to understand and analyze the matched graph from multiple perspectives. By incorporating metric learning and feature selection into the graph matching algorithm, we allow users to interactively modify the graph matching result based on their information needs. We have applied our approach to various types of data, including news articles, tweets, and blog data. Quantitative evaluation and real-world case studies demonstrate the promise of our approach, especially in support of examining a topic-graph-based full picture at different levels of detail. Xiting Wang, Shixia Liu, Jianfei Chen 0001, Jun Zhu 0001, Baining Guo |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2016 | 3D cartoon face generation by local deformation mapping
Jingyong Zhou, Xin Tong 0001, Zicheng Liu 0001, Baining Guo |
Vis. Comput. | 4 |
| 2015 | Unsupervised Extraction of Video Highlights via Robust Recurrent Auto-EncodersabstractWith the growing popularity of short-form video sharing platforms such as Instagram and Vine, there has been an increasing need for techniques that automatically extract highlights from video. Whereas prior works have approached this problem with heuristic rules or supervised learning, we present an unsupervised learning approach that takes advantage of the abundance of user-edited videos on social media websites such as YouTube. Based on the idea that the most significant sub-events within a video class are commonly present among edited videos while less interesting ones appear less frequently, we identify the significant sub-events via a robust recurrent auto-encoder trained on a collection of user-edited videos queried for each particular class of interest. The auto-encoder is trained using a proposed shrinking exponential loss function that makes it robust to noise in the web-crawled training data, and is configured with bidirectional long short term memory (LSTM) [5] cells to better model the temporal structure of highlight segments. Different from supervised techniques, our method can infer highlights using only a set of downloaded edited videos, without also needing their pre-edited counterparts which are rarely available online. Extensive experiments indicate the promise of our proposed solution in this challenging unsupervised setting. Huan Yang 0005, Baoyuan Wang, Stephen Lin 0001, David P. Wipf, Minyi Guo, Baining Guo |
ICCV | 6 |
| 2015 | Improving Sampling-based Motion ControlabstractAbstract We address several limitations of the sampling‐based motion control method of Liu et at. [ LYvdP* 10 ]. The key insight is to learn from the past control reconstruction trials through sample distribution adaptation. Coupled with a sliding window scheme for better performance and an averaging method for noise reduction, the improved algorithm can efficiently construct open‐loop controls for long and challenging reference motions in good quality. Our ideas are intuitive and the implementations are simple. We compare the improved algorithm with the original algorithm both qualitatively and quantitatively, and demonstrate the effectiveness of the improved algorithm with a variety of motions ranging from stylized walking and dancing to gymnastic and Martial Arts routines. Libin Liu 0002, KangKang Yin, Baining Guo |
Comput. Graph. Forum | 3 |
| 2015 | Irradiance regression for efficient final gathering in global illumination
Xuezhen Huang, Xin Sun 0014, Zhong Ren 0001, Xin Tong 0001, Baining Guo, Kun Zhou 0001 |
Frontiers Comput. Sci. | 5 |
| 2015 | Evolutionary Bayesian Rose TreesabstractWe present an evolutionary multi-branch tree clustering method to model hierarchical topics and their evolutionary patterns over time. The method builds evolutionary trees in a Bayesian online filtering framework. The tree construction is formulated as an online posterior estimation problem, which well balances both the fitness of the current tree and the smoothness between trees. The state-of-the-art multi-branch clustering method, Bayesian rose trees, is employed to generate a topic tree with a high fitness value. A constraint model is also introduced to preserve the smoothness between trees. A set of comprehensive experiments on real world news data demonstrates that the proposed method better incorporates historical tree information and is more efficient and effective than the traditional evolutionary hierarchical clustering algorithm. In contrast to our previous method[31], we implement two additional baseline algorithms to compare them with our algorithm. We also evaluate the performance of the clustering algorithm based on multiple constraint trees. Furthermore, two case studies are conducted to demonstrate the effectiveness and usefulness of our algorithm in helping users understand the major hierarchical topic evolutionary patterns in text data. Shixia Liu, Xiting Wang, Yangqiu Song, Baining Guo |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2015 | Computing locally injective mappings by advanced MIPSabstractComputing locally injective mappings with low distortion in an efficient way is a fundamental task in computer graphics. By revisiting the well-known MIPS (Most-Isometric ParameterizationS) method, we introduce an advanced MIPS method that inherits the local injectivity of MIPS, achieves as low as possible distortions compared to the state-of-the-art locally injective mapping techniques, and performs one to two orders of magnitude faster in computing a mesh-based mapping. The success of our method relies on two key components. The first one is an enhanced MIPS energy function that penalizes the maximal distortion significantly and distributes the distortion evenly over the domain for both mesh-based and meshless mappings. The second is a use of the inexact block coordinate descent method in mesh-based mapping in a way that efficiently minimizes the distortion with the capability not to be trapped early by the local minimum. We demonstrate the capability and superiority of our method in various applications including mesh parameterization, mesh-based and meshless deformation, and mesh improvement. Xiao-Ming Fu 0001, Yang Liu 0014, Baining Guo |
ACM Trans. Graph. | 3 |
| 2015 | Unbiased photon gathering for light transport simulationabstractPhoton mapping (PM) has been widely regarded as an efficient solution for light transport simulation, including challenging caustics paths and many-bounce indirect lighting. The efficiency of PM comes from reusing traced photons. However, the handling of photon gathering in existing PM algorithms is universally biased -- the expected value of their results does not necessarily agree with the true solution of the rendering equation. We present a novel photon gathering method to efficiently achieve unbiased rendering with photon mapping. Instead of aggregating the gathered photons into an estimated density as in classical photon mapping, we process each photon individually and connect the corresponding light sub-path with the eye sub-path that generates the gather point, creating an unbiased path sample. The Monte Carlo estimate for such a path sample is calculated by evaluating all relevant terms in a strict and unbiased way, leading to a self-contained unbiased sampling technique. We further develop a set of multiple importance sampling (MIS) weights that allow our method to be optimally combined with bidirectional path tracing (BDPT), resulting in an unbiased rendering algorithm that can efficiently handle a wide variety of light paths and that compares favorably with previous algorithms. Experiments demonstrate the efficacy and robustness of our method. Xin Sun 0014, Qiming Hou, Baining Guo, Kun Zhou 0001 |
ACM Trans. Graph. | 4 |
| 2015 | Image based relighting using neural networksabstractWe present a neural network regression method for relighting realworld scenes from a small number of images. The relighting in this work is formulated as the product of the scene's light transport matrix and new lighting vectors, with the light transport matrix reconstructed from the input images. Based on the observation that there should exist non-linear local coherence in the light transport matrix, our method approximates matrix segments using neural networks that model light transport as a non-linear function of light source position and pixel coordinates. Central to this approach is a proposed neural network design which incorporates various elements that facilitate modeling of light transport from a small image set. In contrast to most image based relighting techniques, this regression-based approach allows input images to be captured under arbitrary illumination conditions, including light sources moved freely by hand. We validate our method with light transport data of real scenes containing complex lighting effects, and demonstrate that fewer input images are required in comparison to related techniques. Peiran Ren, Yue Dong 0001, Stephen Lin 0001, Xin Tong 0001, Baining Guo |
ACM Trans. Graph. | 5 |
| 2015 | Rolling guidance normal filter for geometric processingabstract3D geometric features constitute rich details of polygonal meshes. Their analysis and editing can lead to vivid appearance of shapes and better understanding of the underlying geometry for shape processing and analysis. Traditional mesh smoothing techniques mainly focus on noise filtering and they cannot distinguish different scales of features well, even mixing them up. We present an efficient method to process different scale geometric features based on a novel rolling-guidance normal filter. Given a 3D mesh, our method iteratively applies a joint bilateral filter to face normals at a specified scale, which empirically smooths small-scale geometric features while preserving large-scale features. Our method recovers the mesh from the filtered face normals by a modified Poisson-based gradient deformation that yields better surface quality than existing methods. We demonstrate the effectiveness and superiority of our method on a series of geometry processing tasks, including geometry texture removal and enhancement, coating transfer, mesh segmentation and level-of-detail meshing. Peng-Shuai Wang, Xiao-Ming Fu 0001, Yang Liu 0014, Xin Tong 0001, Baining Guo |
ACM Trans. Graph. | 6 |
| 2014 | Let It Flow: A Static Method for Exploring Dynamic GraphsabstractResearch into social network analysis has shown that graph metrics, such as degree and closeness, are often used to summarize structural changes in a dynamic graph. However there have been few visual analytics approaches that have been proposed to help analysts study graph evolutions in the context of graph metrics. In this paper, we present a novel approach, called GraphFlow, to visualize dynamic graphs. In contrast to previous approaches that provide users with an animated visualization, GraphFlow offers a static flow visualization that summarizes the graph metrics of the entire graph and its evolution over time. Our solution supports the discovery of high-level patterns that are difficult to identify in an animation or in individual static representations. In addition, GraphFlow provides users with a set of interactions to create filtered views. These views allow users to investigate why a particular pattern has occurred. We showcase the versatility of GraphFlow using two different datasets and describe how it can help users gain insights into complex dynamic graphs. Weiwei Cui 0001, Xiting Wang, Shixia Liu, Nathalie Henry Riche, Tara M. Madhyastha, Kwan-Liu Ma, Baining Guo |
PacificVis | 7 |
| 2014 | Orientational Pyramid Matching for Recognizing Indoor ScenesabstractScene recognition is a basic task towards image understanding. Spatial Pyramid Matching (SPM) has been shown to be an efficient solution for spatial context modeling. In this paper, we introduce an alternative approach, Orientational Pyramid Matching (OPM), for orientational context modeling. Our approach is motivated by the observation that the 3D orientations of objects are a crucial factor to discriminate indoor scenes. The novelty lies in that OPM uses the 3D orientations to form the pyramid and produce the pooling regions, which is unlike SPM that uses the spatial positions to form the pyramid. Experimental results on challenging scene classification tasks show that OPM achieves the performance comparable with SPM and that OPM and SPM make complementary contributions so that their combination gives the state-of-the-art performance. Lingxi Xie, Jingdong Wang 0001, Baining Guo, Bo Zhang 0010, Qi Tian 0001 |
CVPR | 3 |
| 2014 | A novel cloud gaming framework using joint video and graphics streamingabstractAs the popularity of smart phones and tablets, users have an increasing desire to enjoy ubiquitous game playing. The emerging cloud gaming turns this desire into reality, enabling users to play games at anywhere on any devices. However, due to the huge amount of data transmission, it is challenging to provide a high quality game experience under the limited bandwidth capacity. In this paper, we propose a novel cloud gaming framework, in which we introduce two synchronized graphics buffers at both the server and the client sides. The server not only streams the compressed frames captured from game scenes, but also progressively transmits graphics data. The received graphics data is used to generate reference frames. When compressing the next frame, the cloud server will choose the reference frame with a lower residual error, from the previous frame and the current frame rendered from the graphics buffer. With the accumulation of graphics data, the frame rendered from the graphics buffer is close to the captured frame, which greatly reduces the transmission bit rates. Based on the proposed framework, we study the rate allocation problem, in which we optimize the allocated bit rates between the compressed frame and the graphics data to minimize the total distortion under the bandwidth constraint. Experimental results demonstrate that the proposed framework can optimally allocate bit rates to achieve a minimal distortion for cloud gaming compared to the traditional video streaming and graphics streaming approaches. Xiaoming Nan, Xun Guo 0002, Yan Lu 0001, Ling Guan, Shipeng Li 0001, Baining Guo |
ICME | 7 |
| 2014 | A Visual Evaluation Framework for In-Home Physical RehabilitationabstractWe propose a novel method for in-home physical rehabilitation, where a user can visually evaluate his/her performance compared to that of an expert. Normalized joint coordinates extracted from the Kinect skeleton are used as features. A novel Incremental Dynamic Time Warping (IDTW) algorithm is used to align the user and expert sequences. IDTW extends the classic DTW by providing accurate comparison between incomplete (the user's) and complete (the expert's) sequences while significantly reducing the computational time. Instead of providing a single measurement, the proposed method maps the IDTW measurements to a color-coded skeleton frame. Different colors on the limbs provide the user with an easy-to-interpret evaluation of how he or she is performing. Preliminary analysis involving different users and exercises and comparisons against the classic DTW algorithm show the effectiveness of the proposed method. Naimul Mefraz Khan, Stephen Lin 0001, Ling Guan, Baining Guo |
ISM | 4 |
| 2014 | Exemplar-Based Human Action Pose CorrectionabstractThe launch of Xbox Kinect has built a very successful computer vision product and made a big impact on the gaming industry. This sheds lights onto a wide variety of potential applications related to action recognition. The accurate estimation of human poses from the depth image is universally a critical step. However, existing pose estimation systems exhibit failures when facing severe occlusion. In this paper, we propose an exemplar-based method to learn to correct the initially estimated poses. We learn an inhomogeneous systematic bias by leveraging the exemplar information within a specific human action domain. Furthermore, as an extension, we learn a conditional model by incorporation of pose tags to further increase the accuracy of pose correction. In the experiments, significant improvements on both joint-based skeleton correction and tag prediction are observed over the contemporary approaches, including what is delivered by the current Kinect system. Our experiments for the facial landmark correction also illustrate that our algorithm can improve the accuracy of other detection/estimation systems. Wei Shen 0002, Xiang Bai, Tommer Leyvand, Baining Guo, Zhuowen Tu |
IEEE Trans. Cybern. | 5 |
| 2014 | Anisotropic simplicial meshing using local convex functionsabstractWe present a novel method to generate high-quality simplicial meshes with specified anisotropy. Given a surface or volumetric domain equipped with a Riemannian metric that encodes the desired anisotropy, we transform the problem to one of functional approximation. We construct a convex function over each mesh simplex whose Hessian locally matches the Riemannian metric, and iteratively adapt vertex positions and mesh connectivity to minimize the difference between the target convex functions and their piecewise-linear interpolation over the mesh. Our method generalizes optimal Delaunay triangulation and leads to a simple and efficient algorithm. We demonstrate its quality and speed compared to state-of-the-art methods on a variety of domains and metrics. Xiao-Ming Fu 0001, Yang Liu 0014, John M. Snyder, Baining Guo |
ACM Trans. Graph. | 4 |
| 2013 | Fast Neighborhood Graph Search Using Cartesian ConcatenationabstractIn this paper, we propose a new data structure for approximate nearest neighbor search. This structure augments the neighborhood graph with a bridge graph. We propose to exploit Cartesian concatenation to produce a large set of vectors, called bridge vectors, from several small sets of subvectors. Each bridge vector is connected with a few reference vectors near to it, forming a bridge graph. Our approach finds nearest neighbors by simultaneously traversing the neighborhood graph and the bridge graph in the best-first strategy. The success of our approach stems from two factors: the exact nearest neighbor search over a large number of bridge vectors can be done quickly, and the reference vectors connected to a bridge (reference) vector near the query are also likely to be near the query. Experimental results on searching over large scale datasets (SIFT, GISTand HOG) show that our approach outperforms state-of-the-art ANN search algorithms in terms of efficiency and accuracy. The combination of our approach with the IVFADC system [18] also shows superior performance over the BIGANN dataset of 1 billion SIFT features compared with the best previously published result. Jing Wang 0068, Jingdong Wang 0001, Rui Gan, Shipeng Li 0001, Baining Guo |
ICCV | 6 |
| 2013 | Mining evolutionary multi-branch trees from text streamsabstractUnderstanding topic hierarchies in text streams and their evolution patterns over time is very important in many applications. In this paper, we propose an evolutionary multi-branch tree clustering method for streaming text data. We build evolutionary trees in a Bayesian online filtering framework. The tree construction is formulated as an online posterior estimation problem, which considers both the likelihood of the current tree and conditional prior given the previous tree. We also introduce a constraint model to compute the conditional prior of a tree in the multi-branch setting. Experiments on real world news data demonstrate that our algorithm can better incorporate historical tree information and is more efficient and effective than the traditional evolutionary hierarchical clustering algorithm. Xiting Wang, Shixia Liu, Yangqiu Song, Baining Guo |
KDD | 4 |
| 2013 | Computing self-supporting surfaces by regular triangulationabstractMasonry structures must be compressively self-supporting; designing such surfaces forms an important topic in architecture as well as a challenging problem in geometric modeling. Under certain conditions, a surjective mapping exists between a power diagram , defined by a set of 2D vertices and associated weights, and the reciprocal diagram that characterizes the force diagram of a discrete self-supporting network. This observation lets us define a new and convenient parameterization for the space of self-supporting networks. Based on it and the discrete geometry of this design space, we present novel geometry processing methods including surface smoothing and remeshing which significantly reduce the magnitude of force densities and homogenize their distribution. Yang Liu 0014, Hao Pan 0001, John M. Snyder, Wenping Wang 0001, Baining Guo |
ACM Trans. Graph. | 5 |
| 2013 | Simulation and control of skeleton-driven soft body charactersabstractIn this paper we present a physics-based framework for simulation and control of human-like skeleton-driven soft body characters. We couple the skeleton dynamics and the soft body dynamics to enable two-way interactions between the skeleton, the skin geometry, and the environment. We propose a novel pose-based plasticity model that extends the corotated linear elasticity model to achieve large skin deformation around joints. We further reconstruct controls from reference trajectories captured from human subjects by augmenting a sampling-based algorithm. We demonstrate the effectiveness of our framework by results not attainable with a simple combination of previous methods. Libin Liu 0002, KangKang Yin, Bin Wang 0021, Baining Guo |
ACM Trans. Graph. | 4 |
| 2013 | Global illumination with radiance regression functionsabstractWe present radiance regression functions for fast rendering of global illumination in scenes with dynamic local light sources. A radiance regression function (RRF) represents a non-linear mapping from local and contextual attributes of surface points, such as position, viewing direction, and lighting condition, to their indirect illumination values. The RRF is obtained from precomputed shading samples through regression analysis, which determines a function that best fits the shading data. For a given scene, the shading samples are precomputed by an offline renderer. The key idea behind our approach is to exploit the nonlinear coherence of the indirect illumination data to make the RRF both compact and fast to evaluate. We model the RRF as a multilayer acyclic feed-forward neural network, which provides a close functional approximation of the indirect illumination and can be efficiently evaluated at run time. To effectively model scenes with spatially variant material properties, we utilize an augmented set of attributes as input to the neural network RRF to reduce the amount of inference that the network needs to perform. To handle scenes with greater geometric complexity, we partition the input space of the RRF model and represent the subspaces with separate, smaller RRFs that can be evaluated more rapidly. As a result, the RRF model scales well to increasingly complex scene geometry and material variation. Because of its compactness and ease of evaluation, the RRF model enables real-time rendering with full global illumination effects, including changing caustics and multiple-bounce high-frequency glossy interreflections. Peiran Ren, Jiaping Wang, Minmin Gong, Stephen Lin 0001, Xin Tong 0001, Baining Guo |
ACM Trans. Graph. | 6 |
| 2013 | Interpreting concept sketchesabstractConcept sketches are popularly used by designers to convey pose and function of products. Understanding such sketches, however, requires special skills to form a mental 3D representation of the product geometry by linking parts across the different sketches and imagining the intermediate object configurations. Hence, the sketches can remain inaccessible to many, especially non-designers. We present a system to facilitate easy interpretation and exploration of concept sketches. Starting from crudely specified incomplete geometry, often inconsistent across the different views, we propose a globally-coupled analysis to extract part correspondence and inter-part junction information that best explain the different sketch views. The user can then interactively explore the abstracted object to gain better understanding of the product functions. Our key technical contribution is performing shape analysis without access to any coherent 3D geometric model by reasoning in the space of inter-part relations. We evaluate our system on various concept sketches obtained from popular product design books and websites. Tianjia Shao, Wilmot Li, Kun Zhou 0001, Weiwei Xu 0003, Baining Guo, Niloy J. Mitra |
ACM Trans. Graph. | 5 |
| 2013 | Line segment sampling with blue-noise propertiesabstractLine segment sampling has recently been adopted in many rendering algorithms for better handling of a wide range of effects such as motion blur, defocus blur and scattering media. A question naturally raised is how to generate line segment samples with good properties that can effectively reduce variance and aliasing artifacts observed in the rendering results. This paper studies this problem and presents a frequency analysis of line segment sampling. The analysis shows that the frequency content of a line segment sample is equivalent to the weighted frequency content of a point sample. The weight introduces anisotropy that smoothly changes among point samples, line segment samples and line samples according to the lengths of the samples. Line segment sampling thus makes it possible to achieve a balance between noise (point sampling) and aliasing (line sampling) under the same sampling rate. Based on the analysis, we propose a line segment sampling scheme to preserve blue-noise properties of samples which can significantly reduce noise and aliasing artifacts in reconstruction results. We demonstrate that our sampling scheme improves the quality of depth-of-field rendering, motion blur rendering, and temporal light field reconstruction. Xin Sun 0014, Kun Zhou 0001, Jie Guo 0001, Guofu Xie, Jingui Pan, Baining Guo |
ACM Trans. Graph. | 7 |
| 2013 | TransCut: Interactive Rendering of Translucent CutoutsabstractWe present TransCut, a technique for interactive rendering of translucent objects undergoing fracturing and cutting operations. As the object is fractured or cut open, the user can directly examine and intuitively understand the complex translucent interior, as well as edit material properties through painting on cross sections and recombining the broken pieces—all with immediate and realistic visual feedback. This new mode of interaction with translucent volumes is made possible with two technical contributions. The first is a novel solver for the diffusion equation (DE) over a tetrahedral mesh that produces high-quality results comparable to the state-of-art finite element method (FEM) of Arbree et al. but at substantially higher speeds. This accuracy and efficiency is obtained by computing the discrete divergences of the diffusion equation and constructing the DE matrix using analytic formulas derived for linear finite elements. The second contribution is a multiresolution algorithm to significantly accelerate our DE solver while adapting to the frequent changes in topological structure of dynamic objects. The entire multiresolution DE solver is highly parallel and easily implemented on the GPU. We believe TransCut provides a novel visual effect for heterogeneous translucent objects undergoing fracturing and cutting operations. Dongping Li, Xin Sun 0014, Zhong Ren 0001, Stephen Lin 0001, Yiying Tong, Baining Guo, Kun Zhou 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2013 | Boundary-Aware Multidomain Subspace DeformationabstractIn this paper, we propose a novel framework for multidomain subspace deformation using node-wise corotational elasticity. With the proper construction of subspaces based on the knowledge of the boundary deformation, we can use the Lagrange multiplier technique to impose coupling constraints at the boundary without overconstraining. In our deformation algorithm, the number of constraint equations to couple two neighboring domains is not related to the number of the nodes on the boundary but is the same as the number of the selected boundary deformation modes. The crack artifact is not present in our simulation result, and the domain decomposition with loops can be easily handled. Experimental results show that the single-core implementation of our algorithm can achieve real-time performance in simulating deformable objects with around quarter million tetrahedral elements. Yin Yang 0002, Weiwei Xu 0003, Xiaohu Guo, Kun Zhou 0001, Baining Guo |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2013 | Interactive chromaticity mapping for multispectral images
Yanxiang Lan, Jiaping Wang, Stephen Lin 0001, Minmin Gong, Xin Tong 0001, Baining Guo |
Vis. Comput. | 6 |
| 2012 | Exemplar-based human action pose correction and taggingabstractThe launch of Xbox Kinect has built a very successful computer vision product and made a big impact to the gaming industry; this sheds lights onto a wide variety of potential applications related to action recognition. The accurate estimation of human poses from the depth image is universally a critical step. However, existing pose estimation systems exhibit failures when faced severe occlusion. In this paper, we propose an exemplar-based method to learn to correct the initially estimated poses. We learn an inhomogeneous systematic bias by leveraging the exemplar information within specific human action domain. Our algorithm is illustrated on both joint-based skeleton correction and tag prediction. In the experiments, significant improvement is observed over the contemporary approaches, including what is delivered by the current Kinect system. Wei Shen 0002, Xiang Bai, Tommer Leyvand, Baining Guo, Zhuowen Tu |
CVPR | 5 |
| 2012 | Single-view hair modeling for portrait manipulationabstractHuman hair is known to be very difficult to model or reconstruct. In this paper, we focus on applications related to portrait manipulation and take an application-driven approach to hair modeling. To enable an average user to achieve interesting portrait manipulation results, we develop a single-view hair modeling technique with modest user interaction to meet the unique requirements set by portrait manipulation. Our method relies on heuristics to generate a plausible high-resolution strand-based 3D hair model. This is made possible by an effective high-precision 2D strand tracing algorithm, which explicitly models uncertainty and local layering during tracing. The depth of the traced strands is solved through an optimization, which simultaneously considers depth constraints, layering constraints as well as regularization terms. Our single-view hair modeling enables a number of interesting applications that were previously challenging, including transferring the hairstyle of one subject to another in a potentially different pose, rendering the original portrait in a novel view and image-space hair editing. Menglei Chai, Lvdi Wang, Yanlin Weng, Yizhou Yu, Baining Guo, Kun Zhou 0001 |
ACM Trans. Graph. | 5 |
| 2012 | Printing spatially-varying reflectance for reproducing HDR imagesabstractWe present a solution for viewing high dynamic range (HDR) images with spatially-varying distributions of glossy materials printed on reflective media. Our method exploits appearance variations of the glossy materials in the angular domain to display the input HDR image at different exposures. As viewers change the print orientation or lighting directions, the print gradually varies its appearance to display the image content from the darkest to the brightest levels. Our solution is based on a commercially available printing system and is fully automatic. Given the input HDR image and the BRDFs of a set of available inks, our method computes the optimal exposures of the HDR image for all viewing conditions and the optimal ink combinations for all pixels by minimizing the difference of their appearances under all viewing conditions. We demonstrate the effectiveness of our method with print samples generated from different inputs and visualized under different viewing and lighting conditions. Yue Dong 0001, Xin Tong 0001, Fabio Pellacini, Baining Guo |
ACM Trans. Graph. | 4 |
| 2012 | All-hex meshing using singularity-restricted fieldabstractDecomposing a volume into high-quality hexahedral cells is a challenging task in geometric modeling and computational geometry. Inspired by the use of cross field in quad meshing and the CubeCover approach in hex meshing, we present a complete all-hex meshing framework based on singularity-restricted field that is essential to induce a valid all-hex structure. Given a volume represented by a tetrahedral mesh, we first compute a boundary-aligned 3D frame field inside it, then convert the frame field to be singularity-restricted by our effective topological operations. In our all-hex meshing framework, we apply the CubeCover method to achieve the volume parametrization. For reducing degenerate elements appearing in the volume parametrization, we also propose novel tetrahedral split operations to preprocess singularity-restricted frame fields. Experimental results show that our algorithm generates high-quality all-hex meshes from a variety of 3D volumes robustly and efficiently. Yang Liu 0014, Weiwei Xu 0003, Wenping Wang 0001, Baining Guo |
ACM Trans. Graph. | 5 |
| 2012 | Terrain runner: control, parameterization, composition, and planning for highly dynamic motionsabstractIn this paper we learn the skills required by real-time physics-based avatars to perform parkour-style fast terrain crossing using a mix of running, jumping, speed-vaulting, and drop-rolling. We begin with a single motion capture example of each skill and then learn reduced-order linear feedback control laws that provide robust execution of the motions during forward dynamic simulation. We then parameterize each skill with respect to the environment, such as the height of obstacles, or with respect to the task parameters, such as running speed and direction. We employ a continuation process to achieve the required parameterization of the motions and their affine feedback laws. The continuation method uses a predictor-corrector method based on radial basis functions. Lastly, we build control laws specific to the sequential composition of different skills, so that the simulated character can robustly transition to obstacle clearing maneuvers from running whenever obstacles are encountered. The learned transition skills work in tandem with a simple online step-based planning algorithm, and together they robustly guide the character to achieve a state that is well-suited for the chosen obstacle-clearing motion. Libin Liu 0002, KangKang Yin, Michiel van de Panne, Baining Guo |
ACM Trans. Graph. | 4 |
| 2012 | An interactive approach to semantic modeling of indoor scenes with an RGBD cameraabstractWe present an interactive approach to semantic modeling of indoor scenes with a consumer-level RGBD camera. Using our approach, the user first takes an RGBD image of an indoor scene, which is automatically segmented into a set of regions with semantic labels. If the segmentation is not satisfactory, the user can draw some strokes to guide the algorithm to achieve better results. After the segmentation is finished, the depth data of each semantic region is used to retrieve a matching 3D model from a database. Each model is then transformed according to the image depth to yield the scene. For large scenes where a single image can only cover one part of the scene, the user can take multiple images to construct other parts of the scene. The 3D models built for all images are then transformed and unified into a complete scene. We demonstrate the efficiency and robustness of our approach by modeling several real-world scenes. Tianjia Shao, Weiwei Xu 0003, Kun Zhou 0001, Jingdong Wang 0001, Dongping Li, Baining Guo |
ACM Trans. Graph. | 6 |
| 2012 | Diffusion curve textures for resolution independent texture mappingabstractWe introduce a vector representation called diffusion curve textures for mapping diffusion curve images (DCI) onto arbitrary surfaces. In contrast to the original implicit representation of DCIs [Orzan et al. 2008], where determining a single texture value requires iterative computation of the entire DCI via the Poisson equation, diffusion curve textures provide an explicit representation from which the texture value at any point can be solved directly, while preserving the compactness and resolution independence of diffusion curves. This is achieved through a formulation of the DCI diffusion process in terms of Green's functions. This formulation furthermore allows the texture value of any rectangular region (e.g. pixel area) to be solved in closed form, which facilitates anti-aliasing. We develop a GPU algorithm that renders anti-aliased diffusion curve textures in real time, and demonstrate the effectiveness of this method through high quality renderings with detailed control curves and color variations. Xin Sun 0014, Guofu Xie, Yue Dong 0001, Stephen Lin 0001, Weiwei Xu 0003, Xin Tong 0001, Baining Guo |
ACM Trans. Graph. | 8 |
| 2012 | Motion-guided mechanical toy modelingabstractWe introduce a new method to synthesize mechanical toys solely from the motion of their features. The designer specifies the geometry and a time-varying rotation and translation of each rigid feature component. Our algorithm automatically generates a mechanism assembly located in a box below the feature base that produces the specified motion. Parts in the assembly are selected from a parameterized set including belt-pulleys, gears, crank-sliders, quick-returns, and various cams (snail, ellipse, and double-ellipse). Positions and parameters for these parts are optimized to generate the specified motion, minimize a simple measure of complexity, and yield a well-distributed layout of parts over the driving axes. Our solution uses a special initialization procedure followed by simulated annealing to efficiently search the complex configuration space for an optimal assembly. Lifeng Zhu, Weiwei Xu 0003, John M. Snyder, Yang Liu 0014, Baining Guo |
ACM Trans. Graph. | 6 |
| 2011 | Non-Linear Beam Tracing on a GPUabstractAbstract Beam tracing combines the flexibility of ray tracing and the speed of polygon rasterization. However, beam tracing so far only handles linear transformations; thus, it is only applicable to linear effects such as planar mirror reflections but not to non‐linear effects such as curved mirror reflection, refraction, caustics and shadows. In this paper, we introduce non‐linear beam tracing to render these non‐linear effects. Non‐linear beam tracing is highly challenging because commodity graphics hardware supports only linear vertex transformation and triangle rasterization. We overcome this difficulty by designing a non‐linear graphics pipeline and implementing it on top of a commodity GPU. This allows beams to be non‐linear where rays within the same beam do not have to be parallel or intersect at a single point. Using these non‐linear beams, real‐time GPU applications can render secondary rays via polygon streaming similar to how they render primary rays. A major strength of this methodology is that it naturally supports fully dynamic scenes without the need to pre‐store a scene database. Utilizing our approach, non‐linear ray tracing effects can be rendered in real‐time on a commodity GPU under a unified framework. Baoquan Liu, Li-Yi Wei, Chongyang Ma, Ying-Qing Xu, Baining Guo, Enhua Wu |
Comput. Graph. Forum | 6 |
| 2011 | Discriminative Sketch-based 3D Model Retrieval via Robust Shape MatchingabstractAbstract We propose a sketch‐based 3D shape retrieval system that is substantially more discriminative and robust than existing systems, especially for complex models. The power of our system comes from a combination of a contour‐based 2D shape representation and a robust sampling‐based shape matching scheme. They are defined over discriminative local features and applicable for partial sketches; robust to noise and distortions in hand drawings; and consistent when strokes are added progressively. Our robust shape matching, however, requires dense sampling and registration and incurs a high computational cost. We thus devise critical acceleration methods to achieve interactive performance: precomputing kNN graphs that record transformations between neighboring contour images and enable fast online shape alignment; pruning sampling and shape registration strategically and hierarchically; and parallelizing shape matching on multi‐core platforms or GPUs. We demonstrate the effectiveness of our system through various experiments, comparisons, and user studies. Tianjia Shao, Weiwei Xu 0003, KangKang Yin, Jingdong Wang 0001, Kun Zhou 0001, Baining Guo |
Comput. Graph. Forum | 6 |
| 2011 | AppGen: interactive material modeling from a single imageabstractWe present AppGen , an interactive system for modeling materials from a single image. Given a texture image of a nearly planar surface lit with directional lighting, our system models the detailed spatially-varying reflectance properties (diffuse, specular and roughness) and surface normal variations with minimal user interaction. We ask users to indicate global shading and reflectance information by roughly marking the image with a few user strokes, while our system assigns reflectance properties and normals to each pixel. We first interactively decompose the input image into the product of a diffuse albedo map and a shading map. A two-scale normal reconstruction algorithm is then introduced to recover the normal variations from the shading map and preserve the geometric features at different scales. We finally assign the specular parameters to each pixel guided by user strokes and the diffuse albedo. Our system generates convincing results within minutes of interaction and works well for a variety of material types that exhibit different reflectance and normal variations, including natural surfaces and man-made ones. Yue Dong 0001, Xin Tong 0001, Fabio Pellacini, Baining Guo |
ACM Trans. Graph. | 4 |
| 2011 | General planar quadrilateral mesh design using conjugate direction fieldabstractWe present a novel method to approximate a freeform shape with a planar quadrilateral (PQ) mesh for modeling architectural glass structures. Our method is based on the study of conjugate direction fields (CDF) which allow the presence of ±κ/4(κ ε Z) singularities. Starting with a triangle discretization of a freeform shape, we first compute an as smooth as possible conjugate direction field satisfying the user's directional and angular constraints, then apply mixed-integer quadrangulation and planarization techniques to generate a PQ mesh which approximates the input shape faithfully. We demonstrate that our method is effective and robust on various 3D models. Yang Liu 0014, Weiwei Xu 0003, Lifeng Zhu, Baining Guo, Falai Chen |
ACM Trans. Graph. | 5 |
| 2011 | Pocket reflectometryabstractWe present a simple, fast solution for reflectance acquisition using tools that fit into a pocket. Our method captures video of a flat target surface from a fixed video camera lit by a hand-held, moving, linear light source. After processing, we obtain an SVBRDF. We introduce a BRDF chart , analogous to a color "checker" chart, which arranges a set of known-BRDF reference tiles over a small card. A sequence of light responses from the chart tiles as well as from points on the target is captured and matched to reconstruct the target's appearance. We develop a new algorithm for BRDF reconstruction which works directly on these LDR responses, without knowing the light or camera position, or acquiring HDR lighting. It compensates for spatial variation caused by the local (finite distance) camera and light position by warping responses over time to align them to a specular reference. After alignment, we find an optimal linear combination of the Lambertian and purely specular reference responses to match each target point's response. The same weights are then applied to the corresponding (known) reference BRDFs to reconstruct the target point's BRDF. We extend the basic algorithm to also recover varying surface normals by adding two spherical caps for diffuse and specular references to the BRDF chart. We demonstrate convincing results obtained after less than 30 seconds of data capture, using commercial mobile phone cameras in a casual environment. Peiran Ren, Jiaping Wang, John M. Snyder, Xin Tong 0001, Baining Guo |
ACM Trans. Graph. | 5 |
| 2011 | Multiscale vector volumesabstractWe introduce multiscale vector volumes , a compact vector representation for volumetric objects with complex internal structures spanning a wide range of scales. With our representation, an object is decomposed into components and each component is modeled as an SDF tree , a novel data structure that uses multiple signed distance functions (SDFs) to further decompose the volumetric component into regions. Multiple signed distance functions collectively can represent non-manifold surfaces and deliver a powerful vector representation for complex volumetric features. We use multiscale embedding to combine object components at different scales into one complex volumetric object. As a result, regions with dramatically different scales and complexities can co-exist in an object. To facilitate volumetric object authoring and editing, we have also developed a scripting language and a GUI prototype. With the help of a recursively defined spatial indexing structure, our vector representation supports fast random access, and arbitrary cross sections of complex volumetric objects can be visualized in real time. Lvdi Wang, Yizhou Yu, Kun Zhou 0001, Baining Guo |
ACM Trans. Graph. | 4 |
| 2011 | Radiance Transfer Biclustering for Real-Time All-Frequency Biscale RenderingabstractWe present a real-time algorithm to render all-frequency radiance transfer at both macroscale and mesoscale. At a mesoscale, the shading is computed on a per-pixel basis by integrating the product of the local incident radiance and a bidirectional texture function. While at a macroscale, the precomputed transfer matrix, which transfers the global incident radiance to the local incident radiance at each vertex, is losslessly compressed by a novel biclustering technique. The biclustering is directly applied on the radiance transfer represented in a pixel basis, on which the BTF is naturally defined. It exploits the coherence in the transfer matrix and a property of matrix element values to reduce both storage and runtime computation cost. Our new algorithm renders at real-time frame rates realistic materials and shadows under all-frequency direct environment lighting. Comparisons show that our algorithm is able to generate images that compare favorably with reference ray tracing results, and has obvious advantages over alternative methods in storage and preprocessing time. Xin Sun 0014, Qiming Hou, Zhong Ren 0001, Kun Zhou 0001, Baining Guo |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2011 | Data-Parallel Octrees for Surface ReconstructionabstractWe present the first parallel surface reconstruction algorithm that runs entirely on the GPU. Like existing implicit surface reconstruction methods, our algorithm first builds an octree for the given set of oriented points, then computes an implicit function over the space of the octree, and finally extracts an isosurface as a watertight triangle mesh. A key component of our algorithm is a novel technique for octree construction on the GPU. This technique builds octrees in real time and uses level-order traversals to exploit the fine-grained parallelism of the GPU. Moreover, the technique produces octrees that provide fast access to the neighborhood information of each octree node, which is critical for fast GPU surface reconstruction. With an octree so constructed, our GPU algorithm performs Poisson surface reconstruction, which produces high-quality surfaces through a global optimization. Given a set of 500 K points, our algorithm runs at the rate of about five frames per second, which is over two orders of magnitude faster than previous CPU algorithms. To demonstrate the potential of our algorithm, we propose a user-guided surface reconstruction technique which reduces the topological ambiguity and improves reconstruction results for imperfect scan data. We also show how to use our algorithm to perform on-the-fly conversion from dynamic point clouds to surfaces as well as to reconstruct fluid surfaces for real-time fluid simulation. Kun Zhou 0001, Minmin Gong, Baining Guo |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2010 | Interactive Rendering of Non-Constant, Refractive Media Using the Ray Equations of Gradient-Index OpticsabstractAbstract Existing algorithms can efficiently render refractive objects of constant refractive index. For a medium with a continuously varying index of refraction, most algorithms use the ray equation of geometric optics to compute piecewise‐linear approximations of the non‐linear rays. By assuming a constant refractive index within each tracing step, these methods often need a large number of small steps to generate satisfactory images. In this paper, we present a new approach for tracing non‐constant, refractive media based on the ray equations of gradient‐index optics. We show that in a medium of constant index gradient, the ray equation has a closed‐form solution, and the intersection point between a ray and the medium boundaries can be efficiently computed using the bisection method. For general non‐constant media, we model the refractive index as a piecewise‐linear function and render the refraction by tracing the tetrahedron‐based representation of the media. Our algorithm can be easily combined with existing rendering algorithms such as photon mapping to generate complex refractive caustics at interactive frame rates. We also derive analytic ray formulations for tracing mirages – a special gradient‐index optical phenomenon. Zhong Ren 0001, Baining Guo, Kun Zhou 0001 |
Comput. Graph. Forum | 3 |
| 2010 | Condenser-Based Instant ReflectometryabstractAbstract We present a technique for rapid capture of high quality bidirectional reflection distribution functions(BRDFs) of surface points. Our method represents the BRDF at each point by a generalized microfacet model with tabulated normal distribution function (NDF) and assumes that the BRDF is symmetrical. A compact and light‐weight reflectometry apparatus is developed for capturing reflectance data from each surface point within one second. The device consists of a pair of condenser lenses, a video camera, and six LED light sources. During capture, the reflected rays from a surface point lit by a LED lighting are refracted by a condenser lenses and efficiently collected by the camera CCD. Taking advantage of BRDF symmetry, our reflectometry apparatus provides an efficient optical design to improve the measurement quality. We also propose a model fitting algorithm for reconstructing the generalized microfacet model from the sparse BRDF slices captured from a material surface point. Our new algorithm addresses the measurement errors and generates more accurate results than previous work. Our technique provides a practical and efficient solution for BRDF acquisition, especially for materials with anisotropic reflectance. We test the accuracy of our approach by comparing our results with ground truth. We demonstrate the efficiency of our reflectometry by measuring materials with different degrees of specularity, values of Fresnel factor, and angular variation. Yanxiang Lan, Yue Dong 0001, Jiaping Wang, Xin Tong 0001, Baining Guo |
Comput. Graph. Forum | 5 |
| 2010 | Real-time Rendering of Heterogeneous Translucent Objects with Arbitrary ShapesabstractAbstract We present a real‐time algorithm for rendering translucent objects of arbitrary shapes. We approximate the scattering of light inside the objects using the diffusion equation, which we solve on‐the‐fly using the GPU. Our algorithm is general enough to handle arbitrary geometry, heterogeneous materials, deformable objects and modifications of lighting, all in real‐time. In a pre‐processing step, we discretize the object into a regular 4‐connected structure (QuadGraph). Due to its regular connectivity, this structure is easily packed into a texture and stored on the GPU. At runtime, we use the QuadGraph stored on the GPU to solve the diffusion equation, in real‐time, taking into account the varying input conditions: Incoming light, object material and geometry. We handle deformable objects, provided the deformation does not change the topological structure of the objects. Jiaping Wang, Nicolas Holzschuch, Kartic Subr, Jun-Hai Yong, Baining Guo |
Comput. Graph. Forum | 6 |
| 2010 | Fabricating spatially-varying subsurface scatteringabstractMany real world surfaces exhibit translucent appearance due to subsurface scattering. Although various methods exists to measure, edit and render subsurface scattering effects, no solution exists for manufacturing physical objects with desired translucent appearance. In this paper, we present a complete solution for fabricating a material volume with a desired surface BSSRDF. We stack layers from a fixed set of manufacturing materials whose thickness is varied spatially to reproduce the heterogeneity of the input BSSRDF. Given an input BSSRDF and the optical properties of the manufacturing materials, our system efficiently determines the optimal order and thickness of the layers. We demonstrate our approach by printing a variety of homogenous and heterogenous BSSRDFs using two hardware setups: a milling machine and a 3D printer. Yue Dong 0001, Jiaping Wang, Fabio Pellacini, Xin Tong 0001, Baining Guo |
ACM Trans. Graph. | 5 |
| 2010 | Manifold bootstrapping for SVBRDF captureabstractManifold bootstrapping is a new method for data-driven modeling of real-world, spatially-varying reflectance, based on the idea that reflectance over a given material sample forms a low-dimensional manifold. It provides a high-resolution result in both the spatial and angular domains by decomposing reflectance measurement into two lower-dimensional phases. The first acquires representatives of high angular dimension but sampled sparsely over the surface, while the second acquires keys of low angular dimension but sampled densely over the surface. We develop a hand-held, high-speed BRDF capturing device for phase one measurements. A condenser-based optical setup collects a dense hemisphere of rays emanating from a single point on the target sample as it is manually scanned over it, yielding 10 BRDF point measurements per second. Lighting directions from 6 LEDs are applied at each measurement; these are amplified to a full 4D BRDF using the general (NDF-tabulated) microfacet model. The second phase captures N =20-200 images of the entire sample from a fixed view and lit by a varying area source. We show that the resulting N -dimensional keys capture much of the distance information in the original BRDF space, so that they effectively discriminate among representatives, though they lack sufficient angular detail to reconstruct the SVBRDF by themselves. At each surface position, a local linear combination of a small number of neighboring representatives is computed to match each key, yielding a high-resolution SVBRDF. A quick capture session (10-20 minutes) on simple devices yields results showing sharp and anisotropic specularity and rich spatial detail. Yue Dong 0001, Jiaping Wang, Xin Tong 0001, John M. Snyder, Yanxiang Lan, Moshe Ben-Ezra, Baining Guo |
ACM Trans. Graph. | 7 |
| 2010 | Micropolygon ray tracing with defocus and motion blurabstractWe present a micropolygon ray tracing algorithm that is capable of efficiently rendering high quality defocus and motion blur effects. A key component of our algorithm is a BVH (bounding volume hierarchy) based on 4D hyper-trapezoids that project into 3D OBBs (oriented bounding boxes) in spatial dimensions. This acceleration structure is able to provide tight bounding volumes for scene geometries, and is thus efficient in pruning intersection tests during ray traversal. More importantly, it can exploit the natural coherence on the time dimension in motion blurred scenes. The structure can be quickly constructed by utilizing the micropolygon grids generated during micropolygon tessellation. Ray tracing of defocused and motion blurred scenes is efficiently performed by traversing the structure. Both the BVH construction and ray traversal are easily implemented on GPUs and integrated into a GPU-based micropolygon renderer. In our experiments, our ray tracer performs up to an order of magnitude faster than the state-of-art rasterizers while consistently delivering an image quality equivalent to a maximum-quality rasterizer. We also demonstrate that the ray tracing algorithm can be extended to handle a variety of effects, such as secondary ray effects and transparency. Qiming Hou, Baining Guo, Kun Zhou 0001 |
ACM Trans. Graph. | 4 |
| 2010 | Interactive hair rendering under environment lightingabstractWe present an algorithm for interactive hair rendering with both single and multiple scattering effects under complex environment lighting. The outgoing radiance due to single scattering is determined by the integral of the product of the environment lighting, the scattering function, and the transmittance that accounts for self-shadowing among hair fibers. We approximate the environment light by a set of spherical radial basis functions (SRBFs) and thus convert the outgoing radiance integral into the sum of radiance contributions of all SRBF lights. For each SRBF light, we factor out the effective transmittance to represent the radiance integral as the product of two terms: the transmittance and the convolution of the SRBF light and the scattering function. Observing that the convolution term is independent of the hair geometry, we precompute it for commonly-used scattering models, and reduce the run-time computation to table lookups. We further propose a technique, called the convolution optical depth map , to efficiently approximate the effective transmittance by filtering the optical depth maps generated at the center of the SRBF using a depth-dependent kernel. As for the multiple scattering computation, we handle SRBF lights by using similar factorization and precomputation schemes, and adopt sparse sampling and interpolation to speed up the computation. Compared to off-line algorithms, our algorithm can generate images of comparable quality, but at interactive frame rates. Zhong Ren 0001, Kun Zhou 0001, Wei Hua 0002, Baining Guo |
ACM Trans. Graph. | 5 |
| 2010 | Line space gathering for single scattering in large scenesabstractWe present an efficient technique to render single scattering in large scenes with reflective and refractive objects and homogeneous participating media. Efficiency is obtained by evaluating the final radiance along a viewing ray directly from the lighting rays passing near to it, and by rapidly identifying such lighting rays in the scene. To facilitate the search for nearby lighting rays, we convert lighting rays and viewing rays into 6D points and planes according to their Plücker coordinates and coefficients, respectively. In this 6D line space, the problem of closest lines search becomes one of closest points to a plane query, which we significantly accelerate using a spatial hierarchy of the 6D points. This approach to lighting ray gathering supports complex light paths with multiple reflections and refractions, and avoids the use of a volume representation, which is expensive for large-scale scenes. This method also utilizes far fewer lighting rays than the number of photons needed in traditional volumetric photon mapping, and does not discretize viewing rays into numerous steps for ray marching. With this approach, results similar to volumetric photon mapping are obtained efficiently in terms of both storage and computation. Xin Sun 0014, Kun Zhou 0001, Stephen Lin 0001, Baining Guo |
ACM Trans. Graph. | 4 |
| 2010 | Vector solid texturesabstractIn this paper, we introduce a compact random-access vector representation for solid textures made of intermixed regions with relatively smooth internal color variations. It is feature-preserving and resolution-independent. In this representation, a texture volume is divided into multiple regions. Region boundaries are implicitly defined using a signed distance function. Color variations within the regions are represented using compactly supported radial basis functions (RBFs). With a spatial indexing structure, such RBFs enable efficient color evaluation during real-time solid texture mapping. Effective techniques have been developed for generating such a vector representation from bitmap solid textures. Data structures and techniques have also been developed to compactly store region labels and distance values for efficient random access during boundary and color evaluation. Lvdi Wang, Kun Zhou 0001, Yizhou Yu, Baining Guo |
ACM Trans. Graph. | 4 |
| 2009 | The Dual-microfacet Model for Capturing Thin Transparent SlabsabstractAbstract We present a new model, called the dual‐microfacet, for those materials such as paper and plastic formed by a thin, transparent slab lying between two surfaces of spatially varying roughness. Light transmission through the slab is represented by a microfacet‐based BTDF which tabulates the microfacet's normal distribution (NDF) as a function of surface location. Though the material is bounded by two surfaces of different roughness, we approximate light transmission through it by a virtual slab determined by a single spatially‐varying NDF. This enables efficient capturing of spatially variant transparent slices. We describe a device for measuring this model over a flat sample by shining light from a CRT behind it and capturing a sequence of images from a single view. Our method captures both angular and spatial variation in the BTDF and provides a good match to measured materials. Jiaping Wang, Yiming Liu 0001, John Snyder, Enhua Wu, Baining Guo |
Comput. Graph. Forum | 6 |
| 2009 | Texture SplicingabstractAbstract We propose a new texture editing operation called texture splicing. For this operation, we regard a texture as having repetitive elements (textons) seamlessly distributed in a particular pattern. Taking two textures as input, texture splicing generates a new texture by selecting the texton appearance from one texture and distribution from the other. Texture splicing involves self‐similarity search to extract the distribution, distribution warping, context‐dependent warping, and finally, texture refinement to preserve overall appearance. We show a variety of results to illustrate this operation. Yiming Liu 0001, Jiaping Wang, Su Xue, Xin Tong 0001, Sing Bing Kang, Baining Guo |
Comput. Graph. Forum | 6 |
| 2009 | Edit Propagation on Bidirectional Texture FunctionsabstractAbstract We propose an efficient method for editing bidirectional texture functions (BTFs) based on edit propagation scheme. In our approach, users specify sparse edits on a certain slice of BTF. An edit propagation scheme is then applied to propagate edits to the whole BTF data. The consistency of the BTF data is maintained by propagating similar edits to points with similar underlying geometry/reflectance. For this purpose, we propose to use view independent features including normals and reflectance features reconstructed from each view to guide the propagation process. We also propose an adaptive sampling scheme for speeding up the propagation process. Since our method needn't any accurate geometry and reflectance information, it allows users to edit complex BTFs with interactive feedback. Kun Xu 0003, Jiaping Wang, Xin Tong 0001, Shi-Min Hu 0001, Baining Guo |
Comput. Graph. Forum | 5 |
| 2009 | Debugging GPU stream programs through automatic dataflow recording and visualizationabstractWe present a novel framework for debugging GPU stream programs through automatic dataflow recording and visualization. Our debugging system can help programmers locate errors that are common in general purpose stream programs but very difficult to debug with existing tools. A stream program is first compiled into an instrumented program using a compiler. This instrumenting compiler automatically adds to the original program dataflow recording code that saves the information of all GPU memory operations into log files. The resulting stream program is then executed on the GPU. With dataflow recording, our debugger automatically detects common memory errors such as out-of-bound access, uninitialized data access, and race conditions -- these errors are extremely difficult to debug with existing tools. When the instrumented program terminates, either normally or due to an error, a dataflow visualizer is launched and it allows the user to examine the memory operation history of all threads and values in all streams. Thus the user can analyze error sources by tracing through relevant threads and streams using the recorded dataflow. A key ingredient of our debugging framework is the GPU interrupt , a novel mechanism that we introduce to support CPU function calls from inside GPU code. We enable interrupts on the GPU by designing a specialized compilation algorithm that translates these interrupts into GPU kernels and CPU management code. Dataflow recording involving disk I/O operations can thus be implemented as interrupt handlers. The GPU interrupt mechanism also allows the programmer to discover errors in more active ways by developing customized debugging functions that can be directly used in GPU code. As examples we show two such functions: assert for data verification and watch for visualizing intermediate results. Qiming Hou, Kun Zhou 0001, Baining Guo |
ACM Trans. Graph. | 3 |
| 2009 | Motion field texture synthesisabstractA variety of animation effects such as herds and fluids contain detailed motion fields characterized by repetitive structures. Such detailed motion fields are often visually important, but tedious to specify manually or expensive to simulate computationally. Due to the repetitive nature, some of these motion fields (e.g. turbulence in fluids) could be synthesized by procedural texturing, but procedural texturing is known for its limited generality. We apply example-based texture synthesis for motion fields. Our technique is general and can take on a variety of user inputs, including captured data, manual art, and physical/procedural simulation. This data-driven approach enables artistic effects that are difficult to achieve via previous methods, such as heart shaped swirls in fluid animation. Due to the use of texture synthesis, our method is able to populate a large output field from a small input exemplar, imposing minimum user workload. Our algorithm also allows the synthesis of output motion fields not only with the same dimension as the input (e.g. 2D to 2D) but also of higher dimension, such as 3D volumetric outputs from 2D planar inputs. This cross-dimension capability supports a convenient usage scenario, i.e. the user could simply supply 2D images and our method produces a 3D motion field with similar characteristics. The motion fields produced by our method are generic, and could be combined with a variety of large-scale low-resolution motions that are easy to specify either manually or computationally but lack the repetitive structures to be characterized as textures. We apply our technique to a variety of animation phenomena, including smoke, liquid, and group motion. Chongyang Ma, Li-Yi Wei, Baining Guo, Kun Zhou 0001 |
ACM Trans. Graph. | 3 |
| 2009 | Kernel Nyström method for light transportabstractWe propose a kernel Nyström method for reconstructing the light transport matrix from a relatively small number of acquired images. Our work is based on the generalized Nyström method for low rank matrices. We introduce the light transport kernel and incorporate it into the Nyström method to exploit the nonlinear coherence of the light transport matrix. We also develop an adaptive scheme for efficiently capturing the sparsely sampled images from the scene. Our experiments indicate that the kernel Nyström method can achieve good reconstruction of the light transport matrix with a few hundred images and produce high quality relighting results. The kernel Nyström method is effective for modeling scenes with complex lighting effects and occlusions which have been challenging for existing techniques. Jiaping Wang, Yue Dong 0001, Xin Tong 0001, Zhouchen Lin, Baining Guo |
ACM Trans. Graph. | 5 |
| 2009 | All-frequency rendering of dynamic, spatially-varying reflectanceabstractWe describe a technique for real-time rendering of dynamic, spatially-varying BRDFs in static scenes with all-frequency shadows from environmental and point lights. The 6D SVBRDF is represented with a general microfacet model and spherical lobes fit to its 4D spatially-varying normal distribution function (SVNDF). A sum of spherical Gaussians (SGs) provides an accurate approximation with a small number of lobes. Parametric BRDFs are fit on-the-fly using simple analytic expressions; measured BRDFs are fit as a preprocess using nonlinear optimization. Our BRDF representation is compact, allows detailed textures, is closed under products and rotations, and supports reflectance of arbitrarily high specularity. At run-time, SGs representing the NDF are warped to align the half-angle vector to the lighting direction and multiplied by the microfacet shadowing and Fresnel factors. This yields the relevant 2D view slice on-the-fly at each pixel, still represented in the SG basis. We account for macro-scale shadowing using a new, nonlinear visibility representation based on spherical signed distance functions (SSDFs). SSDFs allow per-pixel interpolation of high-frequency visibility without ghosting and can be multiplied by the BRDF and lighting efficiently on the GPU. Jiaping Wang, Peiran Ren, Minmin Gong, John M. Snyder, Baining Guo |
ACM Trans. Graph. | 5 |
| 2009 | Example-based hair geometry synthesisabstractWe present an example-based approach to hair modeling because creating hairstyles either manually or through image-based acquisition is a costly and time-consuming process. We introduce a hierarchical hair synthesis framework that views a hairstyle both as a 3D vector field and a 2D arrangement of hair strands on the scalp. Since hair forms wisps, a hierarchical hair clustering algorithm has been developed for detecting wisps in example hairstyles. The coarsest level of the output hairstyle is synthesized using traditional 2D texture synthesis techniques. Synthesizing finer levels of the hierarchy is based on cluster oriented detail transfer. Finally, we compute a discrete tangent vector field from the synthesized hair at every level of the hierarchy to remove undesired inconsistencies among hair trajectories. Improved hair trajectories can be extracted from the vector field. Based on our automatic hair synthesis method, we have also developed simple user-controlled synthesis and editing techniques including feature-preserving combing as well as detail transfer between different hairstyles. Lvdi Wang, Yizhou Yu, Kun Zhou 0001, Baining Guo |
ACM Trans. Graph. | 4 |
| 2009 | Joint-aware manipulation of deformable modelsabstractComplex mesh models of man-made objects often consist of multiple components connected by various types of joints. We propose a joint-aware deformation framework that supports the direct manipulation of an arbitrary mix of rigid and deformable components. First we apply slippable motion analysis to automatically detect multiple types of joint constraints that are implicit in model geometry. For single-component geometry or models with disconnected components, we support user-defined virtual joints. Then we integrate manipulation handle constraints, multiple components, joint constraints, joint limits, and deformation energies into a single volumetric-cell-based space deformation problem. An iterative, parallelized Gauss-Newton solver is used to solve the resulting nonlinear optimization. Interactive deformable manipulation is demonstrated on a variety of geometric models while automatically respecting their multi-component nature and the natural behavior of their joints. Weiwei Xu 0003, KangKang Yin, Kun Zhou 0001, Michiel van de Panne, Falai Chen, Baining Guo |
ACM Trans. Graph. | 7 |
| 2009 | RenderAnts: interactive Reyes rendering on GPUsabstractWe present RenderAnts, the first system that enables interactive Reyes rendering on GPUs. Taking RenderMan scenes and shaders as input, our system first compiles RenderMan shaders to GPU shaders. Then all stages of the basic Reyes pipeline, including bounding/splitting, dicing, shading, sampling, compositing and filtering, are executed on GPUs using carefully designed data-parallel algorithms. Advanced effects such as shadows, motion blur and depth-of-field can also be rendered. In order to avoid exhausting GPU memory, we introduce a novel dynamic scheduling algorithm to bound the memory consumption during rendering. The algorithm automatically adjusts the amount of data being processed in parallel at each stage so that all data can be maintained in the available GPU memory. This allows our system to maximize the parallelism in all individual stages of the pipeline and achieve superior performance. We also propose a multi-GPU scheduling technique based on work stealing so that the system can support scalable rendering on multiple GPUs. The scheduler is designed to minimize inter-GPU communication and balance workloads among GPUs. We demonstrate the potential of RenderAnts using several complex RenderMan scenes and an open source movie entitled Elephants Dream. Compared to Pixar's PRMan, our system can generate images of comparably high quality, but is over one order of magnitude faster. For moderately complex scenes, the system allows the user to change the viewpoint, lights and materials while producing photorealistic results at interactive speed. Kun Zhou 0001, Qiming Hou, Zhong Ren 0001, Minmin Gong, Xin Sun 0014, Baining Guo |
ACM Trans. Graph. | 6 |
| 2008 | Keyframe Based Video Object DeformationabstractThis paper presents a novel deformation system for video objects. The system is designed to minimize the amount of user interaction, while providing flexible and precise user control. It has a keyframe-based user interface. The user only needs to manipulate the video object at several keyframes. Our algorithm smoothly propagate the editing result from the keyframes to the remaining frames and automatically generate the new video object. The algorithm is able to preserve the temporal coherence as well as the shape features of the video object in the original video clips. We demonstrate the potential of our system with a variety of examples. Yanlin Weng, Shichao Hu, Baining Guo |
CW | 5 |
| 2008 | Special issue of CAD/Graphics' 2007
Hong Qin 0001, Baining Guo |
Comput. Graph. | 2 |
| 2008 | Gradient-based Interpolation and Sampling for Real-time Rendering of Inhomogeneous, Single-scattering MediaabstractAbstract We present a real‐time rendering algorithm for inhomogeneous, single scattering media, where all‐frequency shading effects such as glows, light shafts, and volumetric shadows can all be captured. The algorithm first computes source radiance at a small number of sample points in the medium, then interpolates these values at other points in the volume using a gradient‐based scheme that is efficiently applied by sample splatting. The sample points are dynamically determined based on a recursive sample splitting procedure that adapts the number and locations of sample points for accurate and efficient reproduction of shading variations in the medium. The entire pipeline can be easily implemented on the GPU to achieve real‐time performance for dynamic lighting and scenes. Rendering results of our method are shown to be comparable to those from ray tracing. Zhong Ren 0001, Kun Zhou 0001, Stephen Lin 0001, Baining Guo |
Comput. Graph. Forum | 4 |
| 2008 | Image-based Material WeatheringabstractAbstract The appearance manifold [WTL*06] is an efficient approach for modeling and editing time‐variant appearance of materials from the BRDF data captured at single time instance. However, this method is difficult to apply in images in which weathering and shading variations are combined. In this paper, we present a technique for modeling and editing the weathering effects of an object in a single image with appearance manifolds. In our approach, we formulate the input image as the product of reflectance and illuminance. An iterative method is then developed to construct the appearance manifold in color space (i.e., Lab space) for modeling the reflectance variations caused by weathering. Based on the appearance manifold, we propose a statistical method to robustly decompose reflectance and illuminance for each pixel. For editing, we introduce a “pixel‐walking” scheme to modify the pixel reflectance according to its position on the manifold, by which the detailed reflectance variations are well preserved. We illustrate our technique in various applications, including weathering transfer between two images that is first enabled by our technique. Results show that our technique can produce much better results than existing methods, especially for objects with complex geometry and shading effects. Su Xue, Jiaping Wang, Xin Tong 0001, Qionghai Dai, Baining Guo |
Comput. Graph. Forum | 5 |
| 2008 | BSGP: bulk-synchronous GPU programmingabstractWe present BSGP, a new programming language for general purpose computation on the GPU. A BSGP program looks much the same as a sequential C program. Programmers only need to supply a bare minimum of extra information to describe parallel processing on GPUs. As a result, BSGP programs are easy to read, write, and maintain. Moreover, the ease of programming does not come at the cost of performance. A well-designed BSGP compiler converts BSGP programs to kernels and combines them using optimally allocated temporary streams. In our benchmark, BSGP programs achieve similar or better performance than well-optimized CUDA programs, while the source code complexity and programming time are significantly reduced. To test BSGP's code efficiency and ease of programming, we implemented a variety of GPU applications, including a highly sophisticated X3D parser that would be extremely difficult to develop with existing GPU programming languages. Qiming Hou, Kun Zhou 0001, Baining Guo |
ACM Trans. Graph. | 3 |
| 2008 | Example-based dynamic skinning in real timeabstractIn this paper we present an approach to enrich skeleton-driven animations with physically-based secondary deformation in real time. To achieve this goal, we propose a novel, surface-based deformable model that can interactively emulate the dynamics of both low-and high-frequency volumetric effects. Given a surface mesh and a few sample sequences of its physical behavior, a set of motion parameters of the material are learned during an off-line preprocessing step. The deformable model is then applicable to any given skeleton-driven animation of the surface mesh. Additionally, our dynamic skinning technique can be entirely implemented on GPUs and executed with great efficiency. Thus, with minimal changes to the conventional graphics pipeline, our approach can drastically enhance the visual experience of skeleton-driven animations by adding secondary deformation in real time. Kun Zhou 0001, Yiying Tong, Mathieu Desbrun, Hujun Bao, Baining Guo |
ACM Trans. Graph. | 6 |
| 2008 | Interactive relighting of dynamic refractive objectsabstractWe present a new technique for interactive relighting of dynamic refractive objects with complex material properties. We describe our technique in terms of a rendering pipeline in which each stage runs entirely on the GPU. The rendering pipeline converts surfaces to volumetric data, traces the curved paths of photons as they refract through the volume, and renders arbitrary views of the resulting radiance distribution. Our rendering pipeline is fast enough to permit interactive updates to lighting, materials, geometry, and viewing parameters without any precomputation. Applications of our technique include the visualization of caustics, absorption, and scattering while running physical simulations or while manipulating surfaces in real time. Xin Sun 0014, Kun Zhou 0001, Eric J. Stollnitz, Jiaoying Shi, Baining Guo |
ACM Trans. Graph. | 5 |
| 2008 | Modeling and rendering of heterogeneous translucent materials using the diffusion equationabstractIn this article, we propose techniques for modeling and rendering of heterogeneous translucent materials that enable acquisition from measured samples, interactive editing of material attributes, and real-time rendering. The materials are assumed to be optically dense such that multiple scattering can be approximated by a diffusion process described by the diffusion equation. For modeling heterogeneous materials, we present the inverse diffusion algorithm for acquiring material properties from appearance measurements. This modeling algorithm incorporates a regularizer to handle the ill-conditioning of the inverse problem, an adjoint method to dramatically reduce the computational cost, and a hierarchical GPU implementation for further speedup. To render an object with known material properties, we present the polygrid diffusion algorithm , which solves the diffusion equation with a boundary condition defined by the given illumination environment. This rendering technique is based on representation of an object by a polygrid, a grid with regular connectivity and an irregular shape, which facilitates solution of the diffusion equation in arbitrary volumes. Because of the regular connectivity, our rendering algorithm can be implemented on the GPU for real-time performance. We demonstrate our techniques by capturing materials from physical samples and performing real-time rendering and editing with these materials. Jiaping Wang, Xin Tong 0001, Stephen Lin 0001, Zhouchen Lin, Yue Dong 0001, Baining Guo, Harry Shum |
ACM Trans. Graph. | 7 |
| 2008 | Modeling anisotropic surface reflectance with example-based microfacet synthesisabstractWe present a new technique for the visual modeling of spatiallyvarying anisotropic reflectance using data captured from a single view. Reflectance is represented using a microfacet-based BRDF which tabulates the facets' normal distribution (NDF) as a function of surface location. Data from a single view provides a 2D slice of the 4D BRDF at each surface point from which we fit a partial NDF. The fitted NDF is partial because the single view direction coupled with the set of light directions covers only a portion of the "half-angle" hemisphere. We complete the NDF at each point by applying a novel variant of texture synthesis using similar, overlapping partial NDFs from other points. Our similarity measure allows azimuthal rotation of partial NDFs, under the assumption that reflectance is spatially redundant but the local frame may be arbitrarily oriented. Our system includes a simple acquisition device that collects images over a 2D set of light directions by scanning a linear array of LEDs over a flat sample. Results demonstrate that our approach preserves spatial and directional BRDF details and generates a visually compelling match to measured materials. Jiaping Wang, Xin Tong 0001, John M. Snyder, Baining Guo |
ACM Trans. Graph. | 5 |
| 2008 | Inverse texture synthesisabstractThe quality and speed of most texture synthesis algorithms depend on a 2D input sample that is small and contains enough texture variations. However, little research exists on how to acquire such sample. For homogeneous patterns this can be achieved via manual cropping, but no adequate solution exists for inhomogeneous or globally varying textures, i.e. patterns that are local but not stationary, such as rusting over an iron statue with appearance conditioned on varying moisture levels. We present inverse texture synthesis to address this issue. Our inverse synthesis runs in the opposite direction with respect to traditional forward synthesis: given a large globally varying texture, our algorithm automatically produces a small texture compaction that best summarizes the original. This small compaction can be used to reconstruct the original texture or to re-synthesize new textures under user-supplied controls. More important, our technique allows real-time synthesis of globally varying textures on a GPU, where the texture memory is usually too small for large textures. We propose an optimization framework for inverse texture synthesis, ensuring that each input region is properly encoded in the output compaction. Our optimization process also automatically computes orientation fields for anisotropic textures containing both low- and high-frequency regions, a situation difficult to handle via existing techniques. Li-Yi Wei, Jianwei Han, Kun Zhou 0001, Hujun Bao, Baining Guo, Harry Shum |
ACM Trans. Graph. | 5 |
| 2008 | Real-time KD-tree construction on graphics hardwareabstractWe present an algorithm for constructing kd-trees on GPUs. This algorithm achieves real-time performance by exploiting the GPU's streaming architecture at all stages of kd-tree construction. Unlike previous parallel kd-tree algorithms, our method builds tree nodes completely in BFS (breadth-first search) order. We also develop a special strategy for large nodes at upper tree levels so as to further exploit the fine-grained parallelism of GPUs. For these nodes, we parallelize the computation over all geometric primitives instead of nodes at each level. Finally, in order to maintain kd-tree quality, we introduce novel schemes for fast evaluation of node split costs. As far as we know, ours is the first real-time kd-tree algorithm on the GPU. The kd-trees built by our algorithm are of comparable quality as those constructed by off-line CPU algorithms. In terms of speed, our algorithm is significantly faster than well-optimized single-core CPU algorithms and competitive with multi-core CPU algorithms. Our algorithm provides a general way for handling dynamic scenes on the GPU. We demonstrate the potential of our algorithm in applications involving dynamic scenes, including GPU ray tracing, interactive photon mapping, and point cloud modeling. Kun Zhou 0001, Qiming Hou, Rui Wang 0004, Baining Guo |
ACM Trans. Graph. | 4 |
| 2008 | Real-time smoke rendering using compensated ray marchingabstractWe present a real-time algorithm calledcompensated ray marchingfor rendering of smoke under dynamic low-frequency environment lighting. Our approach is based on a decomposition of the input smoke animation, represented as a sequence of volumetric density fields, into a set of radial basis functions (RBFs) and a sequence of residual fields. To expedite rendering, the source radiance distribution within the smoke is computed from only the low-frequency RBF approximation of the density fields, since the high-frequency residuals have little impact on global illumination under low-frequency environment lighting. Furthermore, in computing source radiances the contributions from single and multiple scattering are evaluated at only the RBF centers and then approximated at other points in the volume using an RBF-based interpolation. A slice-based integration of these source radiances along each view ray is then performed to render the final image. The high-frequency residual fields, which are a critical component in the local appearance of smoke, are compensated back into the radiance integral during this ray march to generate images of high detail. The runtime algorithm, which includes both light transfer simulation and ray marching, can be easily implemented on the GPU, and thus allows for real-time manipulation of viewpoint and lighting, as well as interactive editing of smoke attributes such as extinction cross section, scattering albedo, and phase function. Only moderate preprocessing time and storage is needed. This approach provides the first method for real-time smoke rendering that includes single and multiple scattering while generating results comparable in quality to offline algorithms like ray tracing. Kun Zhou 0001, Zhong Ren 0001, Stephen Lin 0001, Hujun Bao, Baining Guo, Harry Shum |
ACM Trans. Graph. | 5 |
| 2008 | Filtering and Rendering of Resolution-Dependent Reflectance ModelsabstractThe apparent reflectance of a surface depends upon the resolution at which it is imaged. Conventional reflectance models represent reflection at a single predetermined resolution; however, a low-resolution pixel that views a greater surface area often exhibits a reflectance more complicated than a high-resolution pixel with a smaller area. To address resolution dependency in reflectance, we utilize a generalized reflectance model based on a mixture of multiple conventional models, and present a framework for efficiently determining the reflectance mixture model of each pixel with respect to resolution. Mixture model parameters are precomputed at multiple resolutions and stored in mipmaps. Unlike color textures, these reflectance parameters cannot be accurately filtered by trilinear interpolation, so we present a technique for nonlinear mipmap filtering that minimizes aliasing in rendered results. This framework can be applied with various parametric reflectance models in graphics hardware for real-time processing. With this technique for filtering and rendering with mipmaps of reflectance mixture models, our system can rapidly render the resolution-dependent reflectance effects that are customarily disregarded in conventional rendering methods. At the end of this paper, we also describe how shadowing and masking effects can be incorporated into this framework to increase the realism of rendering. Ping Tan 0002, Stephen Lin 0001, Long Quan, Baining Guo, Harry Shum |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2007 | Fogshop: Real-Time Design and Rendering of Inhomogeneous, Single-Scattering MediaabstractWe describe a new, analytic approximation to the airlight integral from scattering media whose density is modeled as a sum of Gaussians. The approximation supports real-time rendering of inhomogeneous media including their shadowing and scattering effects. For each Gaussian, this approximation samples the scattering integrand at the projection of its center along the view ray but models attenuation and shadowing with respect to the other Gaussians by integrating density along the fixed path from light source to 3D center to view point. Our method handles isotropic, single-scattering media illuminated by point light sources or low-frequency lighting environments. We also generalize models for reflectance of surfaces from constant-density to inhomogeneous media, using simple optical depth averaging in the direction of the light source or all around the receiver point. Our real-time renderer is incorporated into a system for real-time design and preview of realistic animated fog, steam, or smoke. Kun Zhou 0001, Qiming Hou, Minmin Gong, John Snyder, Baining Guo, Harry Shum |
PG | 5 |
| 2007 | High Dynamic Range Image Hallucination
Lvdi Wang, Li-Yi Wei, Kun Zhou 0001, Baining Guo, Harry Shum |
Rendering Techniques | 4 |
| 2007 | Rendering from compressed high dynamic range textures on programmable graphics hardwareabstractHigh dynamic range (HDR) images are increasingly employed in games and interactive applications for accurate rendering and illumination. One disadvantage of HDR images is their large data size; unfortunately, even though solutions have been proposed for future hardware, commodity graphics hardware today does not provide any native compression for HDR textures. Lvdi Wang, Peter-Pike J. Sloan, Li-Yi Wei, Xin Tong 0001, Baining Guo |
SI3D | 6 |
| 2007 | Low Distortion Shell Map GenerationabstractA shell map (Porumbescu et al., 2005) is a bijective mapping between shell space (the space between a base surface and its offset) and texture space. It can be used to generate small-scale features on surfaces using a variety of modeling techniques. In this paper, we present an efficient algorithm, which reduces distortion by construction, for the offset surface generation of triangular meshes. The basic idea is to independently offset each triangle of the base mesh, and then stitch them up by solving a Poisson equation. We then introduce the details for computation of a stretch metric, which measures the distortion of shell maps. Our results show a substantial improvement compared to previous results Kun Zhou 0001, Yiying Tong, Baining Guo |
VR | 5 |
| 2007 | Context-aware texturesabstractInteresting textures form on the surfaces of objects as the result of external chemical, mechanical, and biological agents. Simulating these textures is necessary to generate models for realistic image synthesis. The textures formed are progressively variant, with the variations depending on the global and local geometric context. We present a method for capturing progressively varying textures and the relevant context parameters that control them. By relating textures and context parameters, we are able to transfer the textures to novel synthetic objects. We present examples of capturing chemical effects, such as rusting; mechanical effects, such as paint cracking; and biological effects, such as the growth of mold on a surface. We demonstrate a user interface that provides a method for specifying where an object is exposed to external agents. We show the results of complex, geometry-dependent textures evolving on synthetic objects. Jianye Lu, Athinodoros S. Georghiades, Andreas Glaser, Hongzhi Wu, Li-Yi Wei, Baining Guo, Julie Dorsey, Holly E. Rushmeier |
ACM Trans. Graph. | 6 |
| 2007 | Mesh puppetry: cascading optimization of mesh deformation with inverse kinematicsabstractWe present mesh puppetry , a variational framework for detail-preserving mesh manipulation through a set of high-level, intuitive, and interactive design tools. Our approach builds upon traditional rigging by optimizing skeleton position and vertex weights in an integrated manner. New poses and animations are created by specifying a few desired constraints on vertex positions, balance of the character, length and rigidity preservation, joint limits, and/or self-collision avoidance. Our algorithm then adjusts the skeleton and solves for the deformed mesh simultaneously through a novel cascading optimization procedure, allowing realtime manipulation of meshes with 50 K + vertices for fast design of pleasing and realistic poses. We demonstrate the potential of our framework through an interactive deformation platform and various applications such as deformation transfer and motion retargeting. Kun Zhou 0001, Yiying Tong, Mathieu Desbrun, Hujun Bao, Baining Guo |
ACM Trans. Graph. | 6 |
| 2007 | Interactive relighting with dynamic BRDFsabstractWe present a technique for interactive relighting in which source radiance, viewing direction, and BRDFs can all be changed on the fly. In handling dynamic BRDFs, our method efficiently accounts for the effects of BRDF modification on the reflectance and incident radiance at a surface point. For reflectance, we develop a BRDF tensor representation that can be factorized into adjustable terms for lighting, viewing, and BRDF parameters. For incident radiance, there exists a non-linear relationship between indirect lighting and BRDFs in a scene, which makes linear light transport frameworks such as PRT unsuitable. To overcome this problem, we introduceprecomputed transfer tensors(PTTs) which decompose indirect lighting into precomputable components that are each a function of BRDFs in the scene, and can be rapidly combined at run time to correctly determine incident radiance. We additionally describe a method for efficient handling of high-frequency specular reflections by separating them from the BRDF tensor representation and processing them using precomputed visibility information. With relighting based on PTTs, interactive performance with indirect lighting is demonstrated in applications to BRDF animation and material tuning. Xin Sun 0014, Kun Zhou 0001, Yanyun Chen, Stephen Lin 0001, Jiaoying Shi, Baining Guo |
ACM Trans. Graph. | 6 |
| 2007 | Gradient domain editing of deforming mesh sequencesabstractMany graphics applications, including computer games and 3D animated films, make heavy use of deforming mesh sequences. In this paper, we generalize gradient domain editing to deforming mesh sequences. Our framework is keyframe based. Given sparse and irregularly distributed constraints at unevenly spaced keyframes, our solution first adjusts the meshes at the keyframes to satisfy these constraints, and then smoothly propagate the constraints and deformations at keyframes to the whole sequence to generate new deforming mesh sequence. To achieve convenient keyframe editing, we have developed an efficient alternating least-squares method. It harnesses the power of subspace deformation and two-pass linear methods to achieve high-quality deformations. We have also developed an effective algorithm to define boundary conditions for all frames using handle trajectory editing. Our deforming mesh editing framework has been successfully applied to a number of editing scenarios with increasing complexity, including footprint editing, path editing, temporal filtering, handle-based deformation mixing, and spacetime morphing. Weiwei Xu 0003, Kun Zhou 0001, Yizhou Yu, Qifeng Tan, Qunsheng Peng 0001, Baining Guo |
ACM Trans. Graph. | 6 |
| 2007 | Direct manipulation of subdivision surfaces on GPUsabstractWe present an algorithm for interactive deformation of subdivision surfaces, including displaced subdivision surfaces and subdivision surfaces with geometric textures. Our system lets the user directly manipulate the surface using freely-selected surface points as handles. During deformation the control mesh vertices are automatically adjusted such that the deforming surface satisfies the handle position constraints while preserving the original surface shape and details. To best preserve surface details, we develop a gradient domain technique that incorporates the handle position constraints and detail preserving objectives into the deformation energy. For displaced subdivision surfaces and surfaces with geometric textures, the deformation energy is highly nonlinear and cannot be handled with existing iterative solvers. To address this issue, we introduce a shell deformation solver, which replaces each numerically unstable iteration step with two stable mesh deformation operations. Our deformation algorithm only uses local operations and is thus suitable for GPU implementation. The result is a real-time deformation system running orders of magnitude faster than the state-of-the-art multigrid mesh deformation solver. We demonstrate our technique with a variety of examples, including examples of creating visually pleasing character animations in real-time by driving a subdivision surface with motion capture data. Kun Zhou 0001, Weiwei Xu 0003, Baining Guo, Harry Shum |
ACM Trans. Graph. | 4 |
| 2006 | Real-Time Facial Expression Mapping for High Resolution 3D Meshes
Mingli Song, Zicheng Liu 0001, Baining Guo |
Computer Graphics International | 3 |
| 2006 | Real-time Multi-perspective Rendering on Graphics Hardware
Xianyou Hou, Li-Yi Wei, Harry Shum, Baining Guo |
Rendering Techniques | 4 |
| 2006 | Silhouette Texture
Hongzhi Wu, Li-Yi Wei, Baining Guo |
Rendering Techniques | 4 |
| 2006 | An efficient large deformation method using domain decomposition
Jin Huang 0001, Xinguo Liu, Hujun Bao, Baining Guo, Harry Shum |
Comput. Graph. | 4 |
| 2006 | Realistic, real-time rendering of ocean wavesabstractAbstract In computer games and other real‐time graphics applications, the ocean surface is typically modelled as a texture or bump‐mapped plane with simple lighting effects. This paper describes a system for realistically rendering the water surface in real time. Our system can render calm ocean waves with sophisticated lighting effects at 100 fps on a 680 MHz Pentium III with a GeForce 3 graphics card. The wave geometry is represented view‐dependently as a dynamic displacement map with surface detail described by a dynamic bump map. The illumination model includes reflection, refraction and Fresnel effects, which are critical for producing the look and feel of water. Copyright © 2006 John Wiley & Sons, Ltd. Luiz Velho 0001, Xin Tong 0001, Baining Guo, Harry Shum |
Comput. Animat. Virtual Worlds | 4 |
| 2006 | Subspace gradient domain mesh deformationabstractIn this paper we present a general framework for performing constrained mesh deformation tasks with gradient domain techniques. We present a gradient domain technique that works well with a wide variety of linear and nonlinear constraints. The constraints we introduce include the nonlinear volume constraint for volume preservation, the nonlinear skeleton constraint for maintaining the rigidity of limb segments of articulated figures, and the projection constraint for easy manipulation of the mesh without having to frequently switch between multiple viewpoints. To handle nonlinear constraints, we cast mesh deformation as a nonlinear energy minimization problem and solve the problem using an iterative algorithm. The main challenges in solving this nonlinear problem are the slow convergence and numerical instability of the iterative solver. To address these issues, we develop a subspace technique that builds a coarse control mesh around the original mesh and projects the deformation energy and constraints onto the control mesh vertices using the mean value interpolation. The energy minimization is then carried out in the subspace formed by the control mesh vertices. Running in this subspace, our energy minimization solver is both fast and stable and it provides interactive responses. We demonstrate our deformation constraints and subspace deformation technique with a variety of constrained deformation examples. Jin Huang 0001, Xinguo Liu, Kun Zhou 0001, Li-Yi Wei, Shang-Hua Teng, Hujun Bao, Baining Guo, Harry Shum |
ACM Trans. Graph. | 8 |
| 2006 | Real-time soft shadows in dynamic scenes using spherical harmonic exponentiationabstractPrevious methods for soft shadows numerically integrate over many light directions at each receiver point, testing blocker visibility in each direction. We introduce a method for real-time soft shadows in dynamic scenes illuminated by large, low-frequency light sources where such integration is impractical. Our method operates on vectors representing low-frequency visibility of blockers in the spherical harmonic basis. Blocking geometry is modeled as a set of spheres; relatively few spheres capture the low-frequency blocking effect of complicated geometry. At each receiver point, we compute the product of visibility vectors for these blocker spheres as seen from the point. Instead of computing an expensive SH product per blocker as in previous work, we perform inexpensive vector sums to accumulate the log of blocker visibility. SH exponentiation then yields the product visibility vector over all blockers. We show how the SH exponentiation required can be approximated accurately and efficiently for low-order SH, accelerating previous CPU-based methods by a factor of 10 or more, depending on blocker complexity, and allowing real-time GPU implementation. Zhong Ren 0001, Rui Wang 0004, John M. Snyder, Kun Zhou 0001, Xinguo Liu, Peter-Pike J. Sloan, Hujun Bao, Qunsheng Peng 0001, Baining Guo |
ACM Trans. Graph. | 10 |
| 2006 | Appearance manifolds for modeling time-variant appearance of materialsabstractWe present a visual simulation technique called appearance manifolds for modeling the time-variant surface appearance of a material from data captured at a single instant in time. In modeling time-variant appearance, our method takes advantage of the key observation that concurrent variations in appearance over a surface represent different degrees of weathering. By reorganizing these various appearances in a manner that reveals their relative order with respect to weathering degree, our method infers spatial and temporal appearance properties of the material's weathering process that can be used to convincingly generate its weathered appearance at different points in time. Results with natural non-linear reflectance variations are demonstrated in applications such as visual simulation of weathering on 3D models, increasing and decreasing the weathering of real objects, and material transfer with weathering effects. Jiaping Wang, Xin Tong 0001, Stephen Lin 0001, Minghao Pan, Chao Wang 0063, Hujun Bao, Baining Guo, Harry Shum |
ACM Trans. Graph. | 7 |
| 2006 | Mesh quilting for geometric texture synthesisabstractWe introduce mesh quilting , a geometric texture synthesis algorithm in which a 3D texture sample given in the form of a triangle mesh is seamlessly applied inside a thin shell around an arbitrary surface through local stitching and deformation. We show that such geometric textures allow interactive and versatile editing and animation, producing compelling visual effects that are difficult to achieve with traditional texturing methods. Unlike pixel-based image quilting, mesh quilting is based on stitching together 3D geometry elements. Our quilting algorithm finds corresponding geometry elements in adjacent texture patches, aligns elements through local deformation, and merges elements to seamlessly connect texture patches. For mesh quilting on curved surfaces, a critical issue is to reduce distortion of geometry elements inside the 3D space of the thin shell. To address this problem we introduce a low-distortion parameterization of the shell space so that geometry elements can be synthesized even on very curved objects without the visual distortion present in previous approaches. We demonstrate how mesh quilting can be used to generate convincing decorations for a wide range of geometric textures. Kun Zhou 0001, Yiying Tong, Mathieu Desbrun, Baining Guo, Harry Shum |
ACM Trans. Graph. | 6 |
| 2006 | Geometry-Driven Photorealistic Facial Expression SynthesisabstractExpression mapping (also called performance driven animation) has been a popular method for generating facial animations. A shortcoming of this method is that it does not generate expression details such as the wrinkles due to skin deformations. In this paper, we provide a solution to this problem. We have developed a geometry-driven facial expression synthesis system. Given feature point positions (the geometry) of a facial expression, our system automatically synthesizes a corresponding expression image that includes photorealistic and natural looking expression details. Due to the difficulty of point tracking, the number of feature points required by the synthesis system is, in general, more than what is directly available from a performance sequence. We have developed a technique to infer the missing feature point motions from the tracked subset by using an example-based approach. Another application of our system is expression editing where the user drags feature points while the system interactively generates facial expressions with skin deformation details. Qingshan Zhang, Zicheng Liu 0001, Baining Guo, Demetri Terzopoulos, Harry Shum |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2006 | Fast example-based surface texture synthesis via discrete optimization
Jianwei Han, Kun Zhou 0001, Li-Yi Wei, Minmin Gong, Hujun Bao, Xinming Zhang 0001, Baining Guo |
Vis. Comput. | 7 |
| 2006 | Geometrically based potential energy for simulating deformable objects
Jin Huang 0001, Xinguo Liu, Kun Zhou 0001, Baining Guo, Hujun Bao |
Vis. Comput. | 5 |
| 2006 | Spherical harmonics scaling
Jiaping Wang, Kun Xu 0003, Kun Zhou 0001, Stephen Lin 0001, Shi-Min Hu 0001, Baining Guo |
Vis. Comput. | 6 |
| 2006 | Variational sphere set approximation for solid objects
Rui Wang 0004, Kun Zhou 0001, John Snyder, Xinguo Liu, Hujun Bao, Qunsheng Peng 0001, Baining Guo |
Vis. Comput. | 7 |
| 2006 | 2D shape deformation using nonlinear least squares optimization
Yanlin Weng, Weiwei Xu 0003, Yanchen Wu, Kun Zhou 0001, Baining Guo |
Vis. Comput. | 5 |
| 2005 | Multiresolution Reflectance FilteringabstractPhysically-based reflectance models typically represent light scattering as a function of surface geometry at the pixel level. With changes in viewing resolution, the geometry imaged within a pixel can undergo significant variations that result in changing reflectance characteristics. To address these transformations, we present a multiresolution reflectance framework based on microfacet normal distributions within a pixel over different scales. Since these distributions must be efficiently determined with respect to resolution, they are recorded at multiple resolution levels in mipmaps. The main contribution of this work is a real-time mipmap filtering technique for these distribution-based parameters that not only provides smooth reflectance transitions in scale, but also minimizes aliasing. With this multiresolution reflectance technique, our system can rapidly and accurately incorporate fine reflectance detail that is customarily disregarded in multiresolution rendering methods. Ping Tan 0002, Stephen Lin 0001, Long Quan, Baining Guo, Harry Shum |
Rendering Techniques | 4 |
| 2005 | Clustering method for fast deformation with constraintsabstractWe present a fast deformation method for flexible objects. The deformation of the object is physically modeled using a linear elasticity model with a displacement based finite elements method, yielding a linear system at each time step of simulation. We solve this linear system using a precomputed force-displacement matrix, which describes the object response in terms of displacement accelerations to the forces acting on each vertex. We exploit the spatial coherence to effectively compress the force-displacement matrix to make this method practical and efficient by applying the clustered principal component analysis method. And we developed a method to efficiently handle the additional constraints for interactive user manipulation. At last large deformations are addressed based upon the compressed force-displacement matrix by combining a domain decomposition method and tracking the rotational motions. The experimental results demonstrate fast performances on complex large scale objects under interactive user manipulations. Jin Huang 0001, Xinguo Liu, Hujun Bao, Baining Guo, Harry Shum |
Symposium on Solid and Physical Modeling | 4 |
| 2005 | Polygonal Shape Blending with Topological Evolutions
Ligang Liu 0001, Bo Zhang 0025, Baining Guo, Harry Shum |
J. Comput. Sci. Technol. | 3 |
| 2005 | Visual simulation of weathering by gamma-ton tracingabstractWeathering modeling introduces blemishes such as dirt, rust, cracks and scratches to virtual scenery. In this paper we present a visual stimulation technique that works well for a wide variety of weathering phenomena. Our technique, called γ-ton tracing, is based on a type of aging-inducing particles called γ-tons. Modeling a weathering effect with γ-ton tracing involves tracing a large number of γ-tons through the scene in a way similar to photon tracing and then generating the weathering effect using the recorded γ-ton transport information. With this technique, we can produce weathering effects that are customized to the scene geometry and tailored to the weathering sources. Several effects that are challenging for existing techniques can be readily captured by γ-ton tracing. These include global transport effects. or "stainbleeding". γ-ton tracing also enables visual simulations of complex multi-weathering effects. Lastly γ-ton tracing can generate weathering effects that not only involve texture changes but also large-scale geometry changes. We demonstrate our technique with a variety of examples. Yanyun Chen, Lin Xia, Tien-Tsin Wong, Xin Tong 0001, Hujun Bao, Baining Guo, Harry Shum |
ACM Trans. Graph. | 6 |
| 2005 | Modeling and rendering of quasi-homogeneous materialsabstractMany translucent materials consist of evenly-distributed heterogeneous elements which produce a complex appearance under different lighting and viewing directions. For these quasi-homogeneous materials, existing techniques do not address how to acquire their material representations from physical samples in a way that allows arbitrary geometry models to be rendered with these materials. We propose a model for such materials that can be readily acquired from physical samples. This material model can be applied to geometric models of arbitrary shapes, and the resulting objects can be efficiently rendered without expensive subsurface light transport simulation. In developing a material model with these attributes, we capitalize on a key observation about the subsurface scattering characteristics of quasi-homogeneous materials at different scales. Locally, the non-uniformity of these materials leads to inhomogeneous subsurface scattering. For subsurface scattering on a global scale, we show that a lengthy photon path through an even distribution of heterogeneous elements statistically resembles scattering in a homogeneous medium. This observation allows us to represent and measure the global light transport within quasi-homogeneous materials as well as the transfer of light into and out of a material volume through surface mesostructures. We demonstrate our technique with results for several challenging materials that exhibit sophisticated appearance features such as transmission of back illumination through surface mesostructures. Xin Tong 0001, Jiaping Wang, Stephen Lin 0001, Baining Guo, Harry Shum |
ACM Trans. Graph. | 4 |
| 2005 | Real-time rendering of plant leavesabstractThis paper presents a framework for the real-time rendering of plant leaves with global illumination effects. Realistic rendering of leaves requires a sophisticated appearance model and accurate lighting computation. For leaf appearance we introduce a parametric model that describes leaves in terms of spatially-variant BRDFs and BTDFs. These BRDFs and BTDFs, incorporating analysis of subsurface scattering inside leaf tissues and rough surface scattering on leaf surfaces, can be measured from real leaves. More importantly, this description is compact and can be loaded into graphics hardware for fast run-time shading calculations, which are essential for achieving high frame rates. For lighting computation, we present an algorithm that extends the Precomputed Radiance Transfer (PRT) approach to all-frequency lighting for leaves. In particular, we handle the combined illumination effects due to low-frequency environment light and high-frequency sunlight. This is done by decomposing the local incident radiance of sunlight into direct and indirect components. The direct component, which contains most of the high frequencies, is not pre-computed with spherical harmonics as in PRT; instead it is evaluated on-the-fly using pre-computed light-visibility convolution data. We demonstrate our framework by the rendering of a variety of leaves and assemblies thereof. Lifeng Wang 0001, Wenle Wang, Julie Dorsey, Baining Guo, Harry Shum |
ACM Trans. Graph. | 5 |
| 2005 | Precomputed shadow fields for dynamic scenesabstractWe present a soft shadow technique for dynamic scenes with moving objects under the combined illumination of moving local light sources and dynamic environment maps. The main idea of our technique is to precompute for each scene entity a shadow field that describes the shadowing effects of the entity at points around it. The shadow field for a light source, called a source radiance field (SRF), records radiance from an illuminant as cube maps at sampled points in its surrounding space. For an occluder, an object occlusion field (OOF) conversely represents in a similar manner the occlusion of radiance by an object. A fundamental difference between shadow fields and previous shadow computation concepts is that shadow fields can be precomputed independent of scene configuration. This is critical for dynamic scenes because, at any given instant, the shadow information at any receiver point can be rapidly computed as a simple combination of SRFs and OOFs according to the current scene configuration. Applications that particularly benefit from this technique include large dynamic scenes in which many instances of an entity can share a single shadow field. Our technique enables low-frequency shadowing effects in dynamic scenes in real-time and all-frequency shadows at interactive rates. Kun Zhou 0001, Stephen Lin 0001, Baining Guo, Harry Shum |
ACM Trans. Graph. | 4 |
| 2005 | Large mesh deformation using the volumetric graph LaplacianabstractWe present a novel technique for large deformations on 3D meshes using the volumetric graph Laplacian. We first construct a graph representing the volume inside the input mesh. The graph need not form a solid meshing of the input mesh's interior; its edges simply connect nearby points in the volume. This graph's Laplacian encodes volumetric details as the difference between each point in the graph and the average of its neighbors. Preserving these volumetric details during deformation imposes a volumetric constraint that prevents unnatural changes in volume. We also include in the graph points a short distance outside the mesh to avoid local self-intersections. Volumetric detail preservation is represented by a quadric energy function. Minimizing it preserves details in a least-squares sense, distributing error uniformly over the whole deformed mesh. It can also be combined with conventional constraints involving surface positions, details or smoothness, and efficiently minimized by solving a sparse linear system.We apply this technique in a 2D curve-based deformation system allowing novice users to create pleasing deformations with little effort. A novel application of this system is to apply nonrigid and exaggerated deformations of 2D cartoon characters to 3D meshes. We demonstrate our system's potential with several examples. Kun Zhou 0001, Jin Huang 0001, John M. Snyder, Xinguo Liu, Hujun Bao, Baining Guo, Harry Shum |
ACM Trans. Graph. | 6 |
| 2005 | TextureMontageabstractWe propose a technique, called TextureMontage , to seamlessly map a patchwork of texture images onto an arbitrary 3D model. A texture atlas can be created through the specification of a set of correspondences between the model and any number of texture images. First, our technique automatically partitions the mesh and the images, driven solely by the choice of feature correspondences. Most charts will then be parameterized over their corresponding image planes through the minimization of a distortion metric based on both geometric distortion and texture mismatch across patch boundaries and images. Lastly, a surface texture inpainting technique is used to fill in the remaining charts of the surface with no corresponding texture patches. The resulting texture mapping satisfies the (sparse or dense) user-specified constraints while minimizing the distortion of the texture images and ensuring a smooth transition across the boundaries of different mesh patches. Seamless Texturing of Arbitrary Surfaces From Multiple Images Kun Zhou 0001, Yiying Tong, Mathieu Desbrun, Baining Guo, Harry Shum |
ACM Trans. Graph. | 5 |
| 2005 | Light Field Morphing Using 2D FeaturesabstractWe present a 2D feature-based technique for morphing 3D objects represented by light fields. Existing light field morphing methods require the user to specify corresponding 3D feature elements to guide morph computation. Since slight errors in 3D specification can lead to significant morphing artifacts, we propose a scheme based on 2D feature elements that is less sensitive to imprecise marking of features. First, 2D features are specified by the user in a number of key views in the source and target light fields. Then the two light fields are warped view by view as guided by the corresponding 2D features. Finally, the two warped light fields are blended together to yield the desired light field morph. Two key issues in light field morphing are feature specification and warping of light field rays. For feature specification, we introduce a user interface for delineating 2D features in key views of a light field, which are automatically interpolated to other views. For ray warping, we describe a 2D technique that accounts for visibility changes and present a comparison to the ideal morphing of light fields. Light field morphing based on 2D features makes it simple to incorporate previous image morphing techniques such as nonuniform blending, as well as to morph between an image and a light field. Lifeng Wang 0001, Stephen Lin 0001, Seungyong Lee 0001, Baining Guo, Harry Shum |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2005 | Decorating Surfaces with Bidirectional Texture FunctionsabstractWe present a system for decorating arbitrary surfaces with bidirectional texture functions (BTF). Our system generates BTFs in two steps. First, we automatically synthesize a BTF over the target surface from a given BTF sample. Then, we let the user interactively paint BTF patches onto the surface such that the painted patches seamlessly integrate with the background patterns. Our system is based on a patch-based texture synthesis approach known as quilting. We present a graphcut algorithm for BTF synthesis on surfaces and the algorithm works well for a wide variety of BTF samples, including those which present problems for existing algorithms. We also describe a graphcut texture painting algorithm for creating new surface imperfections (e.g., dirt, cracks, scratches) from existing imperfections found in input BTF samples. Using these algorithms, we can decorate surfaces with real-world textures that have spatially-variant reflectance, fine-scale geometry details, and surfaces imperfections. A particularly attractive feature of BTF painting is that it allows us to capture imperfections of real materials and paint them onto geometry models. We demonstrate the effectiveness of our system with examples. Kun Zhou 0001, Lifeng Wang 0001, Yasuyuki Matsushita, Jiaoying Shi, Baining Guo, Harry Shum |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2005 | Shell radiance texture functions
Yanyun Chen, Xin Tong 0001, Stephen Lin 0001, Jiaoying Shi, Baining Guo, Harry Shum |
Vis. Comput. | 6 |
| 2005 | Capturing and rendering geometry details for BTF-mapped surfaces
Jiaping Wang, Xin Tong 0001, John Snyder, Yanyun Chen, Baining Guo, Harry Shum |
Vis. Comput. | 5 |
| 2004 | Real-time environment map interpolationabstractEnvironment mapping, or reflection mapping, has been widely used in the game and movie industries to give objects a realistic illumination atmosphere. For moving objects, direct frame-by-frame calculation of environment maps and correspondence-based interpolation are both impractical for real-time applications due to the large computational costs. To deal with this problem, "fake" environment mapping with a fixed, pre-generated environment image has been commonly used, but clearly such an approximation is inadequate for a highly reflective object whose environment is constantly changing as it moves. In this paper, we present an approach that sparsely samples environment maps of a moving object and rapidly interpolates them for high performance. Two techniques are introduced for fast environment map interpolation without computation of scene shading. The first method utilizes scene geometry to facilitate interpolation, and the second involves geometry reconstruction from depth buffer values to reduce inefficiencies caused by complex scene geometry. These two techniques can easily be implemented in graphics hardware, and test results show that they achieve significant boosts in performance over frame-by-frame environment map computation with little loss in visual quality. Wenle Wang, Lifeng Wang 0001, Stephen Lin 0001, Jianmin Wang 0001, Baining Guo |
ICIG | 5 |
| 2004 | Perceptually Based Approach for Planar Shape MorphingabstractThis paper presents an approach for establishing vertex correspondences between two planar shapes. Correspondences are established between the perceptual feature points extracted from both source and target shapes. A similarity metric between two feature points is defined using the intrinsic properties of their local neighborhoods. The optimal correspondence is found by an efficient dynamic programming technique. Our approach treats shape noise by allowing discarding small feature points, which introduces skips in the traversal of the dynamic programming graph. Our method is fast, feature preserving, and invariant to geometric transformations. We demonstrate the superiority of our approach over other approaches by experimental results. Ligang Liu 0001, Guopu Wang, Bo Zhang 0025, Baining Guo, Harry Shum |
PG | 4 |
| 2004 | Iso-charts: Stretch-driven Mesh Parameterization using Spectral Analysis
Kun Zhou 0001, John M. Snyder, Baining Guo, Harry Shum |
Symposium on Geometry Processing | 3 |
| 2004 | Shell texture functionsabstractWe propose a texture function for realistic modeling and efficient rendering of materials that exhibit surface mesostructures, translucency and volumetric texture variations. The appearance of such complex materials for dynamic lighting and viewing directions is expensive to calculate and requires an impractical amount of storage to precompute. To handle this problem, our method models an object as a shell layer, formed by texture synthesis of a volumetric material sample, and a homogeneous inner core. To facilitate computation of surface radiance from the shell layer, we introduce the shell texture function (STF) which describes voxel irradiance fields based on precomputed fine-level light interactions such as shadowing by surface mesostructures and scattering of photons inside the object. Together with a diffusion approximation of homogeneous inner core radiance, the STF leads to fast and detailed raytraced renderings of complex materials. Yanyun Chen, Xin Tong 0001, Jiaping Wang, Stephen Lin 0001, Baining Guo, Harry Shum |
ACM Trans. Graph. | 5 |
| 2004 | Mesh editing with poisson-based gradient field manipulationabstractIn this paper, we introduce a novel approach to mesh editing with the Poisson equation as the theoretical foundation. The most distinctive feature of this approach is that it modifies the original mesh geometry implicitly through gradient field manipulation. Our approach can produce desirable and pleasing results for both global and local editing operations, such as deformation, object merging, and smoothing. With the help from a few novel interactive tools, these operations can be performed conveniently with a small amount of user interaction. Our technique has three key components, a basic mesh solver based on the Poisson equation, a gradient field manipulation scheme using local transforms, and a generalized boundary condition representation based on local frames. Experimental results indicate that our framework can outperform previous related mesh editing techniques. Yizhou Yu, Kun Zhou 0001, Dong Xu 0001, Hujun Bao, Baining Guo, Harry Shum |
ACM Trans. Graph. | 6 |
| 2004 | Synthesis and Rendering of Bidirectional Texture Functions on Arbitrary SurfacesabstractThe bidirectional texture function (BTF) is a 6D function that describes the appearance of a real-world surface as a function of lighting and viewing directions. The BTF can model the fine-scale shadows, occlusions, and specularities caused by surface mesostructures. In this paper, we present algorithms for efficient synthesis of BTFs on arbitrary surfaces and for hardware-accelerated rendering. For both synthesis and rendering, a main challenge is handling the large amount of data in a BTF sample. To addresses this challenge, we approximate the BTF sample by a small number of 4D point appearance functions (PAFs) multiplied by 2D geometry maps. The geometry maps and PAFs lead to efficient synthesis and fast rendering of BTFs on arbitrary surfaces. For synthesis, a surface BTF can be generated by applying a texton-based sysnthesis algorithm to a small set of 2D geometry maps while leaving the companion 4D PAFs untouched. As for rendering, a surface BTF synthesized using geometry maps is well-suited for leveraging the programmable vertex and pixel shaders on the graphics hardware. We present a real-time BTF rendering algorithm that runs at the speed of about 30 frames/second on a mid-level PC with an ATI Radeon 8500 graphics card. We demonstrate the effectiveness of our synthesis and rendering algorithms using both real and synthetic BTF samples. Xinguo Liu, Jingdan Zhang, Xin Tong 0001, Baining Guo, Harry Shum |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2003 | Realistic Rendering of Surface Appearance Using GPUabstractSummary form only given. We present techniques for realistic modeling surface details and efficient rendering of the associated visual effects using programmable GPUs. An important topic in rendering surface appearance is the treatment of mesostructures, which are responsible for fine-scale shadowing, occlusion, inter-reflectance, and silhouettes. One way to model surface mesostructure is by using the bi-directional texture function (BTF), a 6D function that describes the appearance of a real-world surface as a function of lighting and viewing directions. We describe algorithms for efficient synthesis of BTFs on arbitrary surfaces and for hardware-accelerated rendering. Because the BTF can be measured from real-world surfaces, these algorithms allow us to texture 3D models with real-world textures. One issue with the BTF-based approach is that silhouettes of surface mesostructures are absent in the rendered images. This problem can be addressed by view-dependent displacement mapping, a technique for real-time rendering of displacement mapped surfaces using GPUs. We will present examples of our techniques shipped with major game platforms and some upcoming game titles. Baining Guo |
PG | 1 |
| 2003 | Feature-Based Surface Light Field MorphingabstractA surface light field is a function that gives the colors of each object point viewed from different directions. Object representation with a surface light field provides a nice structure for 3D photography. This paper presents a feature-based morphing technique for two objects equipped with surface light fields. The technique consists of geometry morphing and in-between light field mapping. Geometry morphing is accomplished by 3D mesh morphing, where we introduce a vertex merging technique to generate a simpler metamesh. In in-between light field mapping, an in-between object is rendered by extracting necessary fragments from input surface light fields. We also propose an acceleration technique for rendering an in-between object. Experimental results with real and synthetic data show natural and plausible morphing between objects with surface light fields. The proposed morphing technique can be used as an editing tool for 3D photography. Eunhee Jeong, Mincheol Yoon, Yunjin Lee, Minsu Ahn, Seungyong Lee 0001, Baining Guo |
PG | 6 |
| 2003 | Interactive Modeling of Tree BarkabstractThere exist many computer graphics techniques which could achieve high quality tree generation. However, only few works focus on realistic modeling of tree bark. Difficulties lie in the complex appearance of the bark surfaces from a single image. We address three main issues here: feature specification; height field assignment; and texture correction. For feature specification, we use texton channel analysis to specify a variant of common bark features, inluding ironbark, vertical and horizontal fractures, tessellation, furrowed cork, and lenticels. For height field assignment, we develop an intuitive and easy-to-use user interface (UI). Here similarity-based texture editing is used for assigning height fields within a texton channel mask. For texture correction, we use the modeled height fields to eliminate the underlying lighting effects in a captured texture. Our modeling system is image-based: it takes as input a bark image and produces as output a textured height field representing a bark sample. We demonstrate that out method is an effective and easy-to-use technique to interactively model a variety of photo realistic bark surfaces. Lifeng Wang 0001, Ligang Liu 0001, Shi-Min Hu 0001, Baining Guo |
PG | 5 |
| 2003 | View-dependent displacement mappingabstractSignificant visual effects arise from surface mesostructure, such as fine-scale shadowing, occlusion and silhouettes. To efficiently render its detailed appearance, we introduce a technique called view-dependent displacement mapping (VDM) that models surface displacements along the viewing direction. Unlike traditional displacement mapping, VDM allows for efficient rendering of self-shadows, occlusions and silhouettes without increasing the complexity of the underlying surface mesh. VDM is based on per-pixel processing, and with hardware acceleration it can render mesostructure with rich visual appearance in real time. Lifeng Wang 0001, Xin Tong 0001, Stephen Lin 0001, Shi-Min Hu 0001, Baining Guo, Harry Shum |
ACM Trans. Graph. | 6 |
| 2003 | Synthesis of progressively-variant textures on arbitrary surfacesabstractWe present an approach for decorating surfaces with progressively-variant textures . Unlike a homogeneous texture, a progressively-variant texture can model local texture variations, including the scale, orientation, color, and shape variations of texture elements. We describe techniques for modeling progressively-variant textures in 2D as well as for synthesizing them over surfaces. For 2D texture modeling, our feature-based warping technique allows the user to control the shape variations of texture elements, making it possible to capture complex texture variations such as those seen in animal coat patterns. In addition, our feature-based blending technique can create a smooth transition between two given homogeneous textures, with progressive changes of both shapes and colors of texture elements. For synthesizing textures over surfaces, the biggest challenge is that the synthesized texture elements tend to break apart as they progressively vary. To address this issue, we propose an algorithm based on texton masks, which mark most prominent texture elements in the 2D texture sample. By leveraging the power of texton masks, our algorithm can maintain the integrity of the synthesized texture elements on the target surface. Jingdan Zhang, Kun Zhou 0001, Luiz Velho 0001, Baining Guo, Harry Shum |
ACM Trans. Graph. | 4 |
| 2003 | Realistic Rendering and Animation of KnitwearabstractWe present a framework for knitwear modeling and rendering that accounts for characteristics that are particular to knitted fabrics. We first describe a model for animation that considers knitwear features and their effects on knitwear shape and interaction. With the computed free-form knitwear configurations, we present an efficient procedure for realistic synthesis based on the observation that a single cross section of yarn can serve as the basic primitive for modeling entire articles of knitwear. This primitive, called the lumislice, describes radiance from a yarn cross section that accounts for fine-level interactions among yarn fibers. By representing yarn as a sequence of identical but rotated cross sections, the lumislice can effectively propagate local microstructure over arbitrary stitch patterns and knitwear shapes. The lumislice accommodates varying levels of detail, allows for soft shadow generation, and capitalizes on hardware-assisted transparency blending. These modeling and rendering techniques together form a complete approach for generating realistic knitwear. Yanyun Chen, Stephen Lin 0001, Ying-Qing Xu, Baining Guo, Harry Shum |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2002 | Texture Mapping with a Jacobian-Based Spatially-Variant FilterabstractIn this paper we describe a new method to map a texture on a surface with a spatially-variant filter. Our filter takes into consideration the effects of anisotropy using a Jacobian approximation while computing the sampling rate, and the interpolation weights are computed with a sinc function. We also discuss how to do forward and backward mapping with the filter and extend our algorithms to 3D meshes. Our experimental results verify our analysis. Jingdan Zhang, Lifeng Wang 0001, Baining Guo |
PG | 4 |
| 2002 | An Adaptive Sampling Scheme for Out-of-Core SimplificationabstractCurrent out‐of‐core simplification algorithms can efficiently simplify large models that are too complex to be loaded in to the main memory at one time. However, these algorithms do not preserve surface details well since adaptive sampling, a typical strategy for detail preservation, remains to be an open issue for out‐of‐core simplification. In this paper, we present an adaptive sampling scheme, called the balanced retriangulation (BR), for out‐of‐core simplification. A key idea behind BR is that we can use Garland's quadric error matrix to analyze the global distribution of surface details. Based on this analysis, a local retriangulation achieves adaptive sampling by restoring detailed areas with cell split operations while further simplifying smooth areas with edge collapse operations. For a given triangle budget, BR preserves surface details significantly better than uniform sampling algorithms such as uniform clustering. Like uniform clustering, our algorithm has linear running time and small memory requirement. Guangzheng Fei, Kangying Cai, Baining Guo, Enhua Wu |
Comput. Graph. Forum | 3 |
| 2002 | Interactive multiresolution hair modeling and editingabstractHuman hair modeling is a difficult task. This paper presents a constructive hair modeling system with which users can sculpt a wide variety of hairstyles. Our Multiresolution Hair Modeling (MHM) system is based on the observed tendency of adjacent hair strands to form clusters at multiple scales due to static attraction. In our system, initial hair designs are quickly created with a small set of hair clusters. Refinements at finer levels are achieved by subdividing these initial hair clusters. Users can edit an evolving model at any level of detail, down to a single hair strand. High level editing tools support curling, scaling, and copy/paste, enabling users to rapidly create widely varying hairstyles. Editing ease and model realism are enhanced by efficient hair rendering, shading, antialiasing, and shadowing algorithms. Yanyun Chen, Ying-Qing Xu, Baining Guo, Harry Shum |
ACM Trans. Graph. | 3 |
| 2002 | Modeling and rendering of realistic feathersabstractWe present techniques for realistic modeling and rendering of feathers and birds. Our approach is motivated by the observation that a feather is a branching structure that can be described by an L-system. The parametric L-system we derived allows the user to easily create feathers of different types and shapes by changing a few parameters. The randomness in feather geometry is also incorporated into this L-system. To render a feather realistically, we have derived an efficient form of the bidirectional texture function (BTF), which describes the small but visible geometry details on the feather blade. A rendering algorithm combining the L-system and the BTF displays feathers photorealistically while capitalizing on graphics hardware for efficiency. Based on this framework of feather modeling and rendering, we developed a system that can automatically generate appropriate feathers to cover different parts of a bird's body from a few "key feathers" supplied by the user, and produce realistic renderings of the bird. Yanyun Chen, Ying-Qing Xu, Baining Guo, Harry Shum |
ACM Trans. Graph. | 3 |
| 2002 | Synthesis of bidirectional texture functions on arbitrary surfacesabstractThe bidirectional texture function (BTF) is a 6D function that can describe textures arising from both spatially-variant surface reflectance and surface mesostructures. In this paper, we present an algorithm for synthesizing the BTF on an arbitrary surface from a sample BTF. A main challenge in surface BTF synthesis is the requirement of a consistent mesostructure on the surface, and to achieve that we must handle the large amount of data in a BTF sample. Our algorithm performs BTF synthesis based on surface textons, which extract essential information from the sample BTF to facilitate the synthesis. We also describe a general search strategy, called the k-coherent search, for fast BTF synthesis using surface textons. A BTF synthesized using our algorithm not only looks similar to the BTF sample in all viewing/lighthing conditions but also exhibits a consistent mesostructure when viewing and lighting directions change. Moreover, the synthesized BTF fits the target surface naturally and seamlessly. We demonstrate the effectiveness of our algorithm with sample BTFs from various sources, including those measured from real-world textures. Xin Tong 0001, Jingdan Zhang, Ligang Liu 0001, Baining Guo, Harry Shum |
ACM Trans. Graph. | 5 |
| 2002 | Feature-based light field morphingabstractWe present a feature-based technique for morphing 3D objects represented by light fields. Our technique enables morphing of image-based objects whose geometry and surface properties are too difficult to model with traditional vision and graphics techniques. Light field morphing is not based on 3D reconstruction; instead it relies on ray correspondence, i.e., the correspondence between rays of the source and target light fields. We address two main issues in light field morphing: feature specification and visibility changes. For feature specification, we develop an intuitive and easy-to-use user interface (UI). The key to this UI is feature polygons, which are intuitively specified as 3D polygons and are used as a control mechanism for ray correspondence in the abstract 4D ray space. For handling visibility changes due to object shape changes, we introduce ray-space warping. Ray-space warping can fill arbitrarily large holes caused by object shape changes; these holes are usually too large to be properly handled by traditional image warping. Our method can deal with non-Lambertian surfaces, including specular surfaces (with dense light fields). We demonstrate that light field morphing is an effective and easy-to-use technqiue that can generate convincing 3D morphing effects. Zhunping Zhang, Lifeng Wang 0001, Baining Guo, Harry Shum |
ACM Trans. Graph. | 3 |
| 2001 | Spoken dialogue management as planning and acting under uncertaintyabstractSome stochastic models like Markov decision process (MDP) are used to model the dialogue manager. MDP-based system degrades fast when uncertainty about user’s intention increases. We propose a novel dialogue model based on the partially observable Markov decision process (POMDP). We use hidden system states and user intentions as the state set, parser results and low-level information as the observation set, domain actions and dialogue repair actions as the action set. Here the low-level information is extracted from different input modals using Bayesian networks. Because of the limitation of exact algorithms, we focus on heuristic methods and their applicability in dialogue management. Bo Zhang 0025, Qingsheng Cai, Jianfeng Mao, Eric Chang, Baining Guo |
INTERSPEECH | 5 |
| 2001 | Photorealistic rendering of knitwear using the lumisliceabstractWe present a method for efficient synthesis of photorealistic free-form knitwear. Our approach is motivated by the observation that a single cross-section of yarn can serve as the basic primitive for modeling entire articles of knitwear. This primitive, called the lumislice, describes radiance from a yarn cross-section based on fine-level interactions — such as occlusion, shadowing, and multiple scattering — among yarn fibers. By representing yarn as a sequence of identical but rotated cross-sections, the lumislice can effectively propagate local microstructure over arbitrary stitch patterns and knitwear shapes. This framework accommodates varying levels of detail and capitalizes on hardware-assisted transparency blending. To further enhance realism, a technique for generating soft shadows from yarn is also introduced. Ying-Qing Xu, Yanyun Chen, Stephen Lin 0001, Enhua Wu, Baining Guo, Harry Shum |
SIGGRAPH | 6 |
| 2001 | Planning and Acting under Uncertainty: A New Model for Spoken Dialogue System
Bo Zhang 0025, Qingsheng Cai, Jianfeng Mao, Baining Guo |
UAI | 4 |
| 2001 | Realistic and efficient rendering of free-form knitwearabstractAbstract We present a method for rendering knitwear on free‐form surfaces. This method has three main advantages. First, it renders yarn microstructure realistically and efficiently. Second, the rendering efficiency of yarn microstructure does not come at the price of ignoring the interactions between the neighboring yarn loops. Such interactions are modeled in our system to further enhance realism. Finally, our approach gives the user intuitive control on a few key aspects of knitwear appearance: the fluffiness of the yarn and the irregularity in the positioning of the yarn loops. The result is a system that efficiently produces highly realistic rendering of free‐form knitwear with user control on key aspects of visual appearance. Copyright © 2001 John Wiley & Sons, Ltd. Ying-Qing Xu, Baining Guo, Harry Shum |
Comput. Animat. Virtual Worlds | 3 |
| 2001 | Real-time texture synthesis by patch-based samplingabstractWe present an algorithm for synthesizing textures from an input sample. This patch-based sampling algorithm is fast and it makes high-quality texture synthesis a real-time process. For generating textures of the same size and comparable quality, patch-based sampling is orders of magnitude faster than existing algorithms. The patch-based sampling algorithm works well for a wide variety of textures ranging from regular to stochastic. By sampling patches according to a nonparametric estimation of the local conditional MRF density function, we avoid mismatching features across patch boundaries. We also experimented with documented cases for which pixel-based nonparametric sampling algorithms cease to be effective but our algorithm continues to work well. Ce Liu 0001, Ying-Qing Xu, Baining Guo, Harry Shum |
ACM Trans. Graph. | 4 |
| 1998 | Progressive Radiance Evaluation Using Directional Coherence MapsabstractWe develop a progressive refinement algorithm that generates an approximate image quickly, then gradually refines it towards the final result. Our algorithm can reconstruct a high-quality image after evaluating only a small percentage of the pixels. For a typical scene, evaluating only 6 % of the pixels yields an approximate image that is visually hard to distinguish from an image with all the pixels evaluated. At this low sampling rate, previous techniques such as adaptive stochastic sampling suffer from artifacts including heavily jagged edges, missing object parts, and missing high-frequency details. A key ingredient of our algorithm is the directional coherence map (DCM), a new technique for handling general radiance discontinuities in a progressive ray tracing framework. Essentially an encoding of the directional coherence in image space, the DCM performs well on discontinuities that are usually considered extremely difficult, e.g. those involving non-polygonal geometry or caused by secondary light sources. Incorporating the DCM into a ray tracing system incurs only a negligible amount of additional computation. More importantly, the DCM uses little memory and thus it preserves the strengths of ray tracing systems in dealing with complex scenes. We have implemented our algorithm on top of RADIANCE. Our enhanced system can produce high-quality images significantly faster than RADIANCE – sometimes by orders of magnitude. Moreover, when the baseline system becomes less effective as its Monte Carlo components are challenged by difficult lighting configurations, our system will still produce high quality images by redistributing computation to the small percentage of pixels as dictated by the DCM. Baining Guo |
SIGGRAPH | 1 |
| 1998 | Direct Visible Surface Interpolation
Baining Guo |
Comput. Vis. Image Underst. | 1 |
| 1997 | Feature-Oriented Rate Shaping of Pre-Compressed Image/VideoabstractRate shaping plays an important role in transmitting compressed video on a network with dynamic bandwidth. We present a scheme for rate shaping JPEG/MPEG-pre-compressed image/video with feature analysis in scale space. The proposed scheme draws its strengths from a technique for identifying perceptually important features based on geometry-driven diffusion, an interpolation scheme allowing many intentionally dropped blocks to be recovered faithfully from neighboring transmitted blocks, and some careful observations about the close interactions between block dropping and feature analysis. Compared with a successful rate shaping scheme based on joint block dropping and uniform coefficient truncation, the new scheme not only produces images of significantly higher visual quality, but also better PSNR at low bit rates. Wenjun Zeng 0001, Baining Guo, Bede Liu |
ICIP (2) | 2 |
| 1997 | Surface reconstruction: from points to splines
Baining Guo |
Comput. Aided Des. | 1 |
| 1997 | Surface Reconstruction Using Alpha ShapesabstractWe describe a method for reconstructing an unknown surface from a set of data points. The basic approach is to extract the surface as a polygon mesh from an α‐shape. Even though alpha shapes are generalized polytopes having complicated internal structures, we show that manifold surfaces, with or without boundaries, can be efficiently generated, and these surfaces completely describe the α‐shapes to the extent that they are visible from outside. Unlike the original α‐shapes, the polygonal surfaces can be easily simplified to yield compact models suitable for a variety of geometric modeling applications such as surface fitting. Baining Guo, Jai Menon 0002, B. Willette |
Comput. Graph. Forum | 1 |
| 1996 | Local shape control for free-form solids in exact CSG representation
Baining Guo, Jai Menon 0002 |
Comput. Aided Des. | 1 |
| 1995 | Interval Set: A Volume Rendering Technique Generalizing Isosurface ExtractionabstractA scalar volume V={(x,f(x))|x/spl isin/R} is described by a function f(x) defined over some region R of the three dimensional space. The paper presents a simple technique for rendering interval sets of the form I/sub g/(a,b)={(x,f(x))|a/spl les/g(x)/spl les/b}, where a and b are either real numbers of infinities. We describe an algorithm for triangulating interval sets as /spl alpha/ shapes, which can be accurately and efficiently rendered as surfaces or semi transparent clouds. On the theoretical side, interval sets provide an unified approach to isosurface extraction and direct volume rendering. On the practical side, interval sets add flexibility to scalar volume visualization-we may choose to, for example, have an interactive, high quality display of the volume surrounding or "inside" an isosurface when such display for the entire volume is too expensive to produce. Baining Guo |
IEEE Visualization | 1 |
| 1995 | A Multiscale Model for Structure-Based Volume RenderingabstractA scalar volume V={(x,f(x))|x/spl isin/R} is described by a function f(x) defined over some region R of the 3D space. We present a simple technique for rendering multiscale interval sets of the form I/sub s/(a,b)={(x,f/sub s/(x))|a/spl les/g/sub s/(x)/spl les/b}, where a and b are either real numbers or infinities, and f/sub s/(x) is a smoothed version of f(x). At each scale s, the constraint a/spl les/g/sub s/(x)/spl les/b identifies a subvolume in which the most significant variations of V are found. We use a dyadic wavelet transform to construct g/sub s/(x) from f(x) and derive subvolumes with the following attractive properties: 1) the information contained in the subvolumes are sufficient for reconstructing the entire V; and 2) the shapes of the subvolumes provide a hierarchical description of the geometric structures of V. Numerically, the reconstruction in 1) is only an approximation, but it is visually accurate as errors reside at fine scales where our visual sensitivity is not so acute. We triangulate interval sets as /spl alpha/-shapes, which can be efficiently rendered as semi-transparent clouds. Because interval sets are extracted in the object space, their visual display can respond to changes of the view point or transfer function quite fast. The result is a volume rendering technique that provides faster, more effective user interaction with practically no loss of information from the original data. The hierarchical nature of multiscale interval sets also makes it easier to understand the usual complicated structures in scalar volumes. Baining Guo |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 1995 | Quadric and cubic bitetrahedral patches
Baining Guo |
Vis. Comput. | 1 |
| 1994 | Short communication A note on two-sided holes in implicit quadric spline surfaces
Baining Guo |
Vis. Comput. | 1 |
| 1993 | Representation of arbitrary shapes using implicit quadrics
Baining Guo |
Vis. Comput. | 1 |