Harry Shum

dblp:s/HarryShum · also Heung-Yeung Shum · DBLP profile ↗
← Back
257ranked-venue papers
30as first author
23since 2021 · last 2026
0000-0002-4684-911XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 199 · 20 first-author · 11 since 2021Artificial intelligence and machine learning · 118 · 17 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 9 · 3 first-author · 2 since 2021Systems, architecture and hardware · 8 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author
YearPublicationVenuePosition
2026 PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning
abstract
Jingcheng Hu, Yinmin Zhang, Shijie Shang, Xiaobo Yang, Yue Peng, Zhewei Huang, Hebin Zhou, Xin Wu, Jie Cheng, Fanqi Wan, Xiangwen Kong, Chengyuan Yao, Kaiwen Yan, Ailin Huang, Hongyu Zhou, Qi Han, Zheng Ge, Xiangyu Zhang, Heung-Yeung Shum. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jingcheng Hu, Yinmin Zhang, Shijie Shang, Zhewei Huang, Hebin Zhou, Fanqi Wan, Xiangwen Kong, Chengyuan Yao, Kaiwen Yan, Ailin Huang, Zheng Ge, Xiangyu Zhang 0005, Harry Shum
ACL (1)19
2026 Follow-Your-Emoji-Faster: Towards Efficient, Fine-Controllable, and Expressive Freestyle Portrait Animation
Yue Ma 0016, Zexuan Yan, Hongfa Wang, Yingqing He, Junkun Yuan, Ailing Zeng, Chengfei Cai, Harry Shum, Zhifeng Li 0001, Wei Liu 0005, Qifeng Chen 0001
Int. J. Comput. Vis.10
2026 SQuadGen: Generating Simple Quad Layouts via Chart Distance Fields
abstract
3D shapes from scanning, reconstruction, or AI-generated content often lack simple quad mesh layouts—critical for efficient editing and modeling. Existing quad-remeshing techniques typically produce complex layouts with irregular loops, leading to tedious manual cleanup and extensive algorithm tuning. We introduce SQUADGEN, a diffusion-based generative framework that leverages Chart Distance Fields (CDF) to synthesize simple quad layouts on 3D shapes. Our approach addresses two key challenges: (1) the discrete nature of mesh connectivity, which hinders learning, and (2) the scarcity of large-scale datasets with simple quad meshes. To overcome the first, we propose CDF, a continuous surface-based representation enabling effective learning and synthesis of quad layouts. To address the second, we define loop-aware simplicity metrics and construct a large-scale dataset of high-quality quad layouts recovered from public 3D repositories through a robust quad-recovery pipeline. Extensive evaluations across diverse 3D inputs show that SQUADGEN consistently outperforms existing methods, producing robust, artist-friendly simple quad layouts.
Youkang Kong, Yang Liu 0014, Yue Dong 0001, Xin Tong 0001, Harry Shum
ACM Trans. Graph.5
2025 Follow-Your-Click: Open-domain Regional Image Animation via Motion Prompts
abstract
Despite recent advances in image-to-video generation, better controllability and local animation are less explored. Most existing image-to-video methods are not locally aware and tend to move the entire scene. However, human artists may need to control the movement of different objects or regions. Additionally, current I2V methods require users not only to describe the target motion but also to provide redundant detailed descriptions of frame contents.These two issues hinder the practical utilization of current I2V tools. In this paper, we propose a practical framework, named Follow-Your-Click, to achieve image animation with a simple user click (for specifying what to move) and a motion prompt (for specifying how to move). Technically, we propose the first-frame masking strategy, which significantly improves the video generation quality, and a motion-augmented module equipped with a motion prompt dataset to improve the motion prompt following abilities of our model. To further control the motion speed, we propose flow-based motion magnitude control to control the speed of target movement more precisely. Extensive experiments compared with 7 baselines, including both commercial tools and research methods on 8 metrics, suggest the superiority of our approach.
Yue Ma 0016, Yingqing He, Hongfa Wang, Andong Wang, Leqi Shen, Jixuan Ying, Chengfei Cai, Zhifeng Li 0001, Harry Shum, Wei Liu 0005, Qifeng Chen 0001
AAAI10
2025 Towards Multiple Character Image Animation Through Enhancing Implicit Decoupling
abstract
Controllable character image animation has a wide range of applications. Although existing studies have consistently improved performance, challenges persist in the field of character image animation, particularly concerning stability in complex backgrounds and tasks involving multiple characters. To address these challenges, we propose a novel multi-condition guided framework for character image animation, employing several well-designed input modules to enhance the implicit decoupling capability of the model. First, the optical flow guider calculates the background optical flow map as guidance information, which enables the model to implicitly learn to decouple the background motion into background constants and background momentum during training, and generate a stable background by setting zero background momentum during inference. Second, the depth order guider calculates the order map of the characters, which transforms the depth information into the positional information of multiple characters. This facilitates the implicit learning of decoupling different characters, especially in accurately separating the occluded body parts of multiple characters. Third, the reference pose map is input to enhance the ability to decouple character texture and pose information in the reference image. Furthermore, to fill the gap of fair evaluation of multi-character image animation, we propose a new benchmark comprising about 4,000 frames. Extensive qualitative and quantitative evaluations demonstrate that our method excels in generating high-quality character animations, especially in scenarios of complex backgrounds and multiple characters.
Jingyun Xue, Hongfa Wang, Qi Tian 0003, Yue Ma 0016, Andong Wang, Zhiyuan Zhao 0002, Shaobo Min, Kaihao Zhang, Harry Shum, Wei Liu 0005, Mengyang Liu, Wenhan Luo
ICLR10
2025 Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
abstract
We introduce Open-Reasoner-Zero, the first open source implementation of large-scale reasoning-oriented RL training on the base model focusing on scalability, simplicity and accessibility. Through extensive experiments, we demonstrate that a minimalist approach, vanilla PPO with GAE ($\lambda=1$, $\gamma=1$) and straightforward rule-based rewards, without any KL regularization, is sufficient to scale up both benchmark performance and response length, replicating the scaling phenomenon observed in DeepSeek-R1-Zero. Using the same base model as DeepSeek-R1-Zero-Qwen-32B, our implementation achieves superior performance across AIME2024, MATH500, and GPQA Diamond, while demonstrating remarkable efficiency—requiring only 1/10 of the training steps compared to the DeepSeek-R1-Zero pipeline. We validate that this recipe generalizes well across diverse training domains and different model families without algorithmic modifications. Moreover, our analysis not only covers training dynamics and ablation for critical design choices, but also quantitatively show how the learned critic in Reasoner-Zero training effectively identifies and devalues repetitive response patterns, yielding more robust advantage estimations and enhancing training stability. Embracing the principles of open-source, we release our source code, parameter settings, training data, and model weights across various sizes, fostering reproducibility and encouraging further exploration of the properties of related models.
Jingcheng Hu, Yinmin Zhang, Daxin Jiang, Xiangyu Zhang 0005, Harry Shum
NeurIPS6
2025 Engineering and technology for low-altitude economy infrastructure
Harry Shum, Xianbin Cao 0001, Mark Hansen
Frontiers Inf. Technol. Electron. Eng.2
2025 Large investment model
abstract
Abstract Traditional quantitative investment research is encountering diminishing returns alongside rising labor and time costs. To overcome these challenges, we introduce the large investment model (LIM), a novel research paradigm designed to enhance both performance and efficiency at scale. LIM employs end-to-end learning and universal modeling to create an upstream foundation model, which is capable of autonomously learning comprehensive signal patterns from diverse financial data spanning multiple exchanges, instruments, and frequencies. These “global patterns” are subsequently transferred to downstream strategy modeling, optimizing performance for specific tasks. We detail the system architecture design of LIM, address the technical challenges inherent in this approach, and outline potential directions for future research.
Jian Guo 0016, Harry Shum
Frontiers Inf. Technol. Electron. Eng.2
2024 TOSS: High-quality Text-guided Novel View Synthesis from a Single Image
abstract
In this paper, we present TOSS, which introduces text to the task of novel view synthesis (NVS) from just a single RGB image. While Zero123 has demonstrated impressive zero-shot open-set NVS capabilities, it treats NVS as a pure image-to-image translation problem. This approach suffers from the challengingly under-constrained nature of single-view NVS: the process lacks means of explicit user control and often result in implausible NVS generations. To address this limitation, TOSS uses text as high-level semantic information to constrain the NVS solution space. TOSS fine-tunes text-to-image Stable Diffusion pre-trained on large-scale text-image pairs and introduces modules specifically tailored to image and camera pose conditioning, as well as dedicated training for pose correctness and preservation of fine details. Comprehensive experiments are conducted with results showing that our proposed TOSS outperforms Zero123 with higher-quality NVS results and faster convergence. We further support these results with comprehensive ablations that underscore the effectiveness and potential of the introduced semantic guidance and architecture design.
Yukai Shi, He Cao, Boshi Tang, Xianbiao Qi, Tianyu Yang 0003, Shilong Liu 0004, Lei Zhang 0001, Harry Shum
ICLR10
2024 Think-on-Graph: Deep and Responsible Reasoning of Large Language Model on Knowledge Graph
abstract
Although large language models (LLMs) have achieved significant success in various tasks, they often struggle with hallucination problems, especially in scenarios requiring deep and responsible reasoning. These issues could be partially addressed by introducing external knowledge graphs (KG) in LLM reasoning. In this paper, we propose a new LLM-KG integrating paradigm ``$\hbox{LLM}\otimes\hbox{KG}$'' which treats the LLM as an agent to interactively explore related entities and relations on KGs and perform reasoning based on the retrieved knowledge. We further implement this paradigm by introducing a new approach called Think-on-Graph (ToG), in which the LLM agent iteratively executes beam search on KG, discovers the most promising reasoning paths, and returns the most likely reasoning results. We use a number of well-designed experiments to examine and illustrate the following advantages of ToG: 1) compared with LLMs, ToG has better deep reasoning power; 2) ToG has the ability of knowledge traceability and knowledge correctability by leveraging LLMs reasoning and expert feedback; 3) ToG provides a flexible plug-and-play framework for different LLMs, KGs and prompting strategies without any additional training cost; 4) the performance of ToG with small LLM models could exceed large LLM such as GPT-4 in certain scenarios and this reduces the cost of LLM deployment and application. As a training-free method with lower computational cost and better generality, ToG achieves overall SOTA in 6 out of 9 datasets where most previous SOTAs rely on additional training.
Jiashuo Sun, Chengjin Xu, Lumingyuan Tang, Saizhuo Wang, Chen Lin 0001, Yeyun Gong, Lionel M. Ni, Harry Shum, Jian Guo 0016
ICLR8
2024 HumanTOMATO: Text-aligned Whole-body Motion Generation
abstract
This work targets a novel text-driven **whole-body** motion generation task, which takes a given textual description as input and aims at generating high-quality, diverse, and coherent facial expressions, hand gestures, and body motions simultaneously. Previous works on text-driven motion generation tasks mainly have two limitations: they ignore the key role of fine-grained hand and face controlling in vivid whole-body motion generation, and lack a good alignment between text and motion. To address such limitations, we propose a Text-aligned whOle-body Motion generATiOn framework, named HumanTOMATO, which is the first attempt to our knowledge towards applicable holistic motion generation in this research area. To tackle this challenging task, our solution includes two key designs: (1) a Holistic Hierarchical VQ-VAE (aka H${}^{2}$VQ) and a Hierarchical-GPT for fine-grained body and hand motion reconstruction and generation with two structured codebooks; and (2) a pre-trained text-motion-alignment model to help generated motion align with the input textual description explicitly. Comprehensive experiments verify that our model has significant advantages in both the quality of generated motions and their alignment with text.
Shunlin Lu, Ailing Zeng, Ruimao Zhang, Lei Zhang 0001, Harry Shum
ICML7
2024 Follow-Your-Emoji: Fine-Controllable and Expressive Freestyle Portrait Animation
abstract
We present Follow-Your-Emoji, a diffusion-based framework for portrait animation, which animates a reference portrait with target landmark sequences. The main challenge of portrait animation is to preserve the identity of the reference portrait and transfer the target expression to this portrait while maintaining temporal consistency and fidelity. To address these challenges, Follow-Your-Emoji equipped the powerful Stable Diffusion model with two well-designed technologies. Specifically, we first adopt a new explicit motion signal, namely expression-aware landmark, to guide the animation process. We discover this landmark can not only ensure the accurate motion alignment between the reference portrait and target motion during inference but also increase the ability to portray exaggerated expressions (i.e., large pupil movements) and avoid identity leakage. Then, we propose a facial fine-grained loss to improve the model’s ability of subtle expression perception and reference portrait appearance reconstruction by using both expression and facial masks. Accordingly, our method demonstrates significant performance in controlling the expression of freestyle portraits, including real humans, cartoons, sculptures, and even animals. By leveraging a simple and effective progressive generation strategy, we extend our model to stable long-term animation, thus increasing its potential application value. To address the lack of a benchmark for this field, we introduce EmojiBench, a comprehensive benchmark comprising diverse portrait images, driving videos, and landmarks. We show extensive evaluations on EmojiBench to verify the superiority of Follow-Your-Emoji. The code, training dataset and benchmark will be found in https://github.com/mayuelala/FollowYourEmoji.
Yue Ma 0016, Hongfa Wang, Yingqing He, Junkun Yuan, Ailing Zeng, Chengfei Cai, Harry Shum, Wei Liu 0005, Qifeng Chen 0001
SIGGRAPH Asia9
2024 Quant 4.0: engineering quantitative investment with automated, explainable, and knowledge-driven artificial intelligence
abstract
Quantitative investment (abbreviated as “quant” in this paper) is an interdisciplinary field combining financial engineering, computer science, mathematics, statistics, etc. Quant has become one of the mainstream investment methodologies over the past decades, and has experienced three generations: quant 1.0, trading by mathematical modeling to discover mis-priced assets in markets; quant 2.0, shifting the quant research pipeline from small “strategy workshops” to large “alpha factories”; quant 3.0, applying deep learning techniques to discover complex nonlinear pricing rules. Despite its advantage in prediction, deep learning relies on extremely large data volume and labor-intensive tuning of “black-box” neural network models. To address these limitations, in this paper, we introduce quant 4.0 and provide an engineering perspective for next-generation quant. Quant 4.0 has three key differentiating components. First, automated artificial intelligence (AI) changes the quant pipeline from traditional hand-crafted modeling to state-of-the-art automated modeling and employs the philosophy of “algorithm produces algorithm, model builds model, and eventually AI creates AI.” Second, explainable AI develops new techniques to better understand and interpret investment decisions made by machine learning black boxes, and explains complicated and hidden risk exposures. Third, knowledge-driven AI supplements data-driven AI such as deep learning and incorporates prior knowledge into modeling to improve investment decisions, in particular for quantitative value investing. Putting all these together, we discuss how to build a system that practices the quant 4.0 concept. We also discuss the application of large language models in quantitative finance. Finally, we propose 10 challenging research problems for quant technology, and discuss potential solutions, research directions, and future trends.
Jian Guo 0016, Saizhuo Wang, Lionel M. Ni, Harry Shum
Frontiers Inf. Technol. Electron. Eng.4
2023 Hand Avatar: Free-Pose Hand Animation and Rendering from Monocular Video
abstract
We present HandAvatar, a novel representation for hand animation and rendering, which can generate smoothly compositional geometry and self-occlusion-aware texture. Specifically, we first develop a MANO-HD model as a high-resolution mesh topology to fit personalized hand shapes. Sequentially, we decompose hand geometry into per-bone rigid parts, and then re-compose paired geometry encodings to derive an across-part consistent occupancy field. As for texture modeling, we propose a self-occlusion-aware shading field (SelF). In SelF, drivable anchors are paved on the MANO-HD surface to record albedo information under a wide variety of hand poses. Moreover, directed soft occupancy is designed to describe the ray-to-surface relation, which is leveraged to generate an illumination field for the disentanglement of pose-independent albedo and pose-dependent illumination. Trained from monocular video data, our HandAvatar can perform freepose hand animation and rendering while at the same time achieving superior appearance fidelity. We also demonstrate that HandAvatar provides a route for hand appearance editing. Project website: https://seanchenxy.github.iO/HandAvatarWeb.
Baoyuan Wang, Harry Shum
CVPR3
2023 Learning Detailed Radiance Manifolds for High-Fidelity and 3D-Consistent Portrait Synthesis from Monocular Image
abstract
A key challenge for novel view synthesis of monocular portrait images is 3D consistency under continuous pose variations. Most existing methods rely on 2D generative models which often leads to obvious 3D inconsistency artifacts. We present a 3D-consistent novel view synthesis approach for monocular portrait images based on a recent proposed 3D-aware GAN, namely Generative Radiance Manifolds (GRAM) [13], which has shown strong 3D consistency at multiview image generation of virtual subjects via the radiance manifolds representation. However, simply learning an encoder to map a real image into the latent space of GRAM can only reconstruct coarse radiance manifolds without faithful fine details, while improving the reconstruction fidelity via instance-specific optimization is time-consuming. We introduce a novel detail manifolds reconstructor to learn 3D-consistent fine details on the radiance manifolds from monocular images, and combine them with the coarse radiance manifolds for high-fidelity reconstruction. The 3D priors derived from the coarse radiance manifolds are used to regulate the learned details to ensure reasonable synthesized results at novel views. Trained on in-the-wild 2D images, our method achieves high-fidelity and 3D-consistent portrait synthesis largely outperforming the prior art. Project page: https://yudeng.github.io/GRAMInverter/
Baoyuan Wang, Harry Shum
CVPR3
2023 Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation
abstract
In this paper we present Mask DINO, a unified object detection and segmentation framework. Mask DINO extends DINO (DETR with Improved Denoising Anchor Boxes) by adding a mask prediction branch which supports all image segmentation tasks (instance, panoptic, and semantic). It makes use of the query embeddings from DINO to dot-product a high-resolution pixel embedding map to predict a set of binary masks. Some key components in DINO are extended for segmentation through a shared architecture and training process. Mask DINO is simple, efficient, and scalable, and it can benefit from joint large-scale detection and segmentation datasets. Our experiments show that Mask DINO significantly outperforms all existing specialized segmentation methods, both on a ResNet-50 backbone and a pre-trained model with SwinL backbone. Notably, Mask DINO establishes the best results to date on instance segmentation (54.5 AP on COCO), panoptic segmentation (59.4 PQ on COCO), and semantic segmentation (60.8 mIoU on ADE20K) among models under one billion parameters. Code is available at https://github.com/IDEA-Research/MaskDINO.
Feng Li 0040, Hao Zhang 0097, Huaizhe Xu, Shilong Liu 0004, Lei Zhang 0001, Lionel M. Ni, Harry Shum
CVPR7
2023 Progressive Disentangled Representation Learning for Fine-Grained Controllable Talking Head Synthesis
abstract
We present a novel one-shot talking head synthesis method that achieves disentangled and fine-grained control over lip motion, eye gaze&blink, head pose, and emotional expression. We represent different motions via disentangled latent representations and leverage an image generator to synthesize talking heads from them. To effectively disentangle each motion factor, we propose a progressive disentangled representation learning strategy by separating the factors in a coarse-to-fine manner, where we first extract unified motion feature from the driving signal, and then isolate each fine-grained motion from the unified feature. We leverage motion-specific contrastive learning and regressing for non-emotional motions, and introduce feature-level decorrelation and self-reconstruction for emotional expression, to fully utilize the inherent properties of each motion factor in unstructured video data to achieve disentanglement. Experiments show that our method provides high quality speech&lip-motion synchronization along with precise and disentangled control over multiple extra facial motions, which can hardly be achieved by previous methods.
Duomin Wang, Zixin Yin, Harry Shum, Baoyuan Wang
CVPR4
2023 Reinforced Disentanglement for Face Swapping without Skip Connection
abstract
The SOTA face swap models still suffer the problem of either target identity (i.e., shape) being leaked or the target non-identity attributes (i.e., background, hair) failing to be fully preserved in the final results. We show that this insufficient disentanglement is caused by two flawed designs that were commonly adopted in prior models: (1) counting on only one compressed encoder to represent both the semantic-level non-identity facial attributes(i.e., pose) and the pixel-level non-facial region details, which is contradictory to satisfy at the same time; (2) highly relying on long skip-connections [50] between the encoder and the final generator, leaking a certain amount of target face identity into the result. To fix them, we introduce a new face swap framework called "WSC-swap" that gets rid of skip connections and uses two target encoders to respectively capture the pixel-level non-facial region attributes and the semantic non-identity attributes in the face region. To further reinforce the disentanglement learning for the target encoder, we employ both identity removal loss via adversarial training (i.e., GAN [18]) and the non-identity preservation loss via prior 3DMM models like [11]. Extensive experiments on both FaceForensics++ and CelebA-HQ show that our results significantly outperform previous works on a rich set of metrics, including one novel metric for measuring identity consistency that was completely neglected before.
Xiaohang Ren, Pengfei Yao, Harry Shum, Baoyuan Wang
ICCV4
2023 DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection
Hao Zhang 0097, Feng Li 0040, Shilong Liu 0004, Lei Zhang 0001, Hang Su 0006, Jun Zhu 0001, Lionel M. Ni, Harry Shum
ICLR8
2023 Creating Virtual AI Beings
abstract
From Apple's Siri to Amazon's Alexa to Microsoft's Cortana, today we human beings are increasingly sharing the world with AI beings. Eventually, the population of AI beings will be many times more than that of humans. In this talk, I will present some challenges and opportunities in creating AI beings: realism, interaction and embodiment. In particular, I will focus on how we can massively scale up content creation for photorealistic AI beings. For instance, how to make results controllable? How to cross the uncanny valley? How to democratize the creation process? I will show many results using the latest neural rendering technology for creating realistic AI beings. Finally, I will use an example of creating virtual singers to illustrate the state-of-the-art AI beings currently available in the social media and on the market.
Harry Shum
VR1
2023 Semi-supervised 3D shape segmentation with multilevel consistency and part substitution
abstract
The lack of fine-grained 3D shape segmentation data is the main obstacle to developing learning-based 3D segmentation techniques. We propose an effective semi-supervised method for learning 3D segmentations from a few labeled 3D shapes and a large amount of unlabeled 3D data. For the unlabeled data, we present a novel multilevel consistency loss to enforce consistency of network predictions between perturbed copies of a 3D shape at multiple levels: point level, part level, and hierarchical level. For the labeled data, we develop a simple yet effective part substitution scheme to augment the labeled 3D shapes with more structural variations to enhance training. Our method has been extensively validated on the task of 3D object semantic segmentation on PartNet and ShapeNetPart, and indoor scene semantic segmentation on ScanNet. It exhibits superior performance to existing semi-supervised and unsupervised pre-training 3D approaches.
Chun-Yu Sun, Hao-Xiang Guo 0001, Peng-Shuai Wang, Xin Tong 0001, Yang Liu 0014, Harry Shum
Comput. Vis. Media7
2023 Locally Attentional SDF Diffusion for Controllable 3D Shape Generation
abstract
Although the recent rapid evolution of 3D generative neural networks greatly improves 3D shape generation, it is still not convenient for ordinary users to create 3D shapes and control the local geometry of generated shapes. To address these challenges, we propose a diffusion-based 3D generation framework --- locally attentional SDF diffusion , to model plausible 3D shapes, via 2D sketch image input. Our method is built on a two-stage diffusion model. The first stage, named occupancy-diffusion , aims to generate a low-resolution occupancy field to approximate the shape shell. The second stage, named SDF-diffusion , synthesizes a high-resolution signed distance field within the occupied voxels determined by the first stage to extract fine geometry. Our model is empowered by a novel view-aware local attention mechanism for image-conditioned shape generation, which takes advantage of 2D image patch features to guide 3D voxel feature learning, greatly improving local controllability and model generalizability. Through extensive experiments in sketch-conditioned and category-conditioned 3D shape generation tasks, we validate and demonstrate the ability of our method to provide plausible and diverse 3D shapes, as well as its superior controllability and generalizability over existing work.
Xin-Yang Zheng, Hao Pan 0001, Peng-Shuai Wang, Xin Tong 0001, Yang Liu 0014, Harry Shum
ACM Trans. Graph.6
2022 On the principles of Parsimony and Self-consistency for the emergence of intelligence
abstract
Ten years into the revival of deep networks and artificial intelligence, we propose a theoretical framework that sheds light on understanding deep networks within a bigger picture of intelligence in general. We introduce two fundamental principles, Parsimony and Self-consistency , which address two fundamental questions regarding intelligence: what to learn and how to learn, respectively. We believe the two principles serve as the cornerstone for the emergence of intelligence, artificial or natural. While they have rich classical roots, we argue that they can be stated anew in entirely measurable and computable ways. More specifically, the two principles lead to an effective and efficient computational framework, compressive closed-loop transcription, which unifies and explains the evolution of modern deep networks and most practices of artificial intelligence. While we use mainly visual data modeling as an example, we believe the two principles will unify understanding of broad families of autonomous intelligent systems and provide a framework for understanding the brain.
Yi Ma 0001, Doris Tsao, Harry Shum
Frontiers Inf. Technol. Electron. Eng.3
2020 The Design and Implementation of XiaoIce, an Empathetic Social Chatbot
abstract
This article describes the development of Microsoft XiaoIce, the most popular social chatbot in the world. XiaoIce is uniquely designed as an artifical intelligence companion with an emotional connection to satisfy the human need for communication, affection, and social belonging. We take into account both intelligent quotient and emotional quotient in system design, cast human–machine social chat as decision-making over Markov Decision Processes, and optimize XiaoIce for long-term user engagement, measured in expected Conversation-turns Per Session (CPS). We detail the system architecture and key components, including dialogue manager, core chat, skills, and an empathetic computing module. We show how XiaoIce dynamically recognizes human feelings and states, understands user intent, and responds to user needs throughout long conversations. Since the release in 2014, XiaoIce has communicated with over 660 million active users and succeeded in establishing long-term relationships with many of them. Analysis of large-scale online logs shows that XiaoIce has achieved an average CPS of 23, which is significantly higher than that of other chatbots and even human conversations.
Jianfeng Gao 0001, Harry Shum
Comput. Linguistics4
2018 From Search to Research: Direct Answers, Perspectives and Dialog
abstract
Advances in artificial intelligence have improved machine understanding of speech, images, and natural language. This in turn has allowed us to greatly enhance the intelligence of products such as Bing and Cortana. This keynote describes our continuing journey beyond keyword-driven systems, into dialog and intelligent agent functionality, helping our users "research more, search less".
Harry Shum
WSDM1
2018 From Eliza to XiaoIce: challenges and opportunities with social chatbots
abstract
Conversational systems have come a long way since their inception in the 1960s. After decades of research and development, we have seen progress from Eliza and Parry in the 1960s and 1970s, to task-completion systems as in the Defense Advanced Research Projects Agency (DARPA) communicator program in the 2000s, to intelligent personal assistants such as Siri, in the 2010s, to today’s social chatbots like XiaoIce. Social chatbots’ appeal lies not only in their ability to respond to users’ diverse requests, but also in being able to establish an emotional connection with users. The latter is done by satisfying users’ need for communication, affection, as well as social belonging. To further the advancement and adoption of social chatbots, their design must focus on user engagement and take both intellectual quotient (IQ) and emotional quotient (EQ) into account. Users should want to engage with a social chatbot; as such, we define the success metric for social chatbots as conversation-turns per session (CPS). Using XiaoIce as an illustrative example, we discuss key technologies in building social chatbots from core chat to visual awareness to skills. We also show how XiaoIce can dynamically recognize emotion and engage the user throughout long conversations with appropriate interpersonal responses. As we become the first generation of humans ever living with artificial intelligenc (AI), we have a responsibility to design social chatbots to be both useful and empathetic, so they will become ubiquitous and help society as a whole.
Harry Shum, Xiaodong He 0001
Frontiers Inf. Technol. Electron. Eng.1
2018 Superpixel-based color-depth restoration and dynamic environment modeling for Kinect-assisted image-based rendering systems
Chong Wang 0001, S. C. Chan 0001, Li Zhang 0041, Harry Shum
Vis. Comput.5
2014 Bing, the fastest growing image search engine
abstract
Since the launch of Bing (www.bing.com) in June 2009, we have seen Bing web search market share in the US more than doubled and Bing image search query share quadrupled. In this talk, I will share our experience building Bing image search as the fastest growing image search engine, and discuss the challenges and opportunities in image search. Specifically, I will talk about how we have significantly improved image search quality, and built differentiated image search user experience using NLP, entity, big data, machine learning and computer vision technologies. By leveraging big data from billions of search queries, billions of images on the web and from the social networks, and billions of user clicks, we have designed massive machine learning systems to continuously improve image search quality. With the focus on natural language and entity understanding, for instance, we have improved Bing's ability to understand the user intent beyond queries and keywords. I will demonstrate with many examples how Bing has delivered a superior image search user experience, quantitatively, qualitatively and aesthetically, by utilizing computer vision techniques.
Harry Shum
ACM Multimedia1
2014 Improving search relevance for short queries in community question answering
abstract
Relevant question retrieval and ranking is a typical task in community question answering (CQA). Existing methods mainly focus on long and syntactically structured queries. However, when an input query is short, the task becomes challenging, due to a lack information regarding user intent. In this paper, we mine different types of user intent from various sources for short queries. With these intent signals, we propose a new intent-based language model. The model takes advantage of both state-of-the-art relevance models and the extra intent information mined from multiple sources. We further employ a state-of-the-art learning-to-rank approach to estimate parameters in the model from training data. Experiments show that by leveraging user intent prediction, our model significantly outperforms the state-of-the-art relevance models in question search.
Haocheng Wu, Wei Wu 0014, Ming Zhou 0001, Enhong Chen, Lei Duan, Harry Shum
WSDM6
2012 Graph-based collective classification for tweets
abstract
In this paper, we address the problem of classifying tweets into topical categories. Because of the short, noisy and ambiguous nature of tweets, we propose to collectively conduct the classification by exploiting the context information (i.e. related tweets) other than individually as in conventional text classification methods. In particular, we augment the content-based representation of text with tweets sharing same #hashtag or URL, which results in a tweet graph. We then formulate the tweet classification task under a graph optimization framework. We investigate three popular approaches, namely, Loopy Belief Propagation (LBP), Relaxation Labeling (RL), and Iterative Classification Algorithm (ICA). Extensive experiment results show that the graph-based tweet classification approach remarkably improves the performance, while the ICA model with relationship of sharing the same #hashtag gives the best result on separate tweet graph.
Yajuan Duan, Furu Wei, Ming Zhou 0001, Harry Shum
CIKM4
2012 Twitter Topic Summarization by Ranking Tweets using Social Influence and Content Quality
Yajuan Duan, Furu Wei, Ming Zhou 0001, Harry Shum
COLING5
2012 QsRank: Query-sensitive hash code ranking for efficient ∊-neighbor search
abstract
Although binary hash code-based image indexing methods have been recently developed for large-scale applications, the problem of ranking such hash codes has been barely studied. In this paper, we propose a query sensitive ranking algorithm (QsRank) to rank PCA-based hash codes for the ∊-neighbor search problem. The QsRank algorithm takes the target neighborhood radius ∊ and the raw feature of a given query as input, and models the statistical properties of the target ∊-neighbors in the space of hash codes. Unlike the Hamming distance, the proposed algorithm does not compress query points to hash codes. Therefore, it suffers less information loss and is more effective than Hamming distance-based approaches. Based on the QsRank method, we developed an efficient indexing structure and retrieval algorithm for large-scale ∊-neighbor search. Evaluations on two datasets of 10 million web images and 10 million SIFT descriptors demonstrate that the proposed retrieval system achieves higher accuracy with less memory cost and faster speed.
Lei Zhang 0001, Harry Shum
CVPR3
2012 Object-Based Rendering and 3-D Reconstruction Using a Moveable Image-Based System
abstract
This paper proposes a movable image-based rendering (M-IBR) system for improving the viewing freedom and environmental modeling capability of conventional static IBR systems. The system supports object-based rendering and 3-D reconstruction capability and consists of three main components.An improved video stabilization method to reduce the shaky motion frequently encountered in movable IBR systems. It employs local polynomial regression (LPR) to automatically select an appropriate bandwidth for smoothing the estimated motion.
Shuai Zhang 0004, S. C. Chan 0001, Harry Shum
IEEE Trans. Circuits Syst. Video Technol.4
2012 A Multi-Camera Approach to Image-Based Rendering and 3-D/Multiview Display of Ancient Chinese Artifacts
abstract
This paper proposes an image-based approach for the capturing, rendering and display of ancient Chinese artifacts for cultural heritage preservation. A multiple-camera circular array is proposed to record images of the artifacts, which forms a simplified circular light field (SCLF). A systematic image-based approach and associate algorithms such as segmentation, depth estimation and shape morphing are developed for rendering new views of the Chinese artifacts. An object-based compression scheme is also proposed to reduce the data size for storage and transmission of the texture, depth maps and alpha maps associated with the object-based circular light field. Spatial redundancies among the various images are exploited to improve the coding performance, while avoiding excessive complexity in selective decoding of the light field to support fast rendering speed. To allow the Chinese artifacts to be viewed over the internet, scalable prioritized transmission and rendering schemes of the SCLF with low latency were also developed. The multiple views so synthesized enable the ancient artifacts to be displayed in 3-D/multi-view displays. Several collections from the University Museum and Art Gallery at The University of Hong Kong were captured and excellent rendering results are obtained.
King To Ng, Chong Wang 0001, S. C. Chan 0001, Harry Shum
IEEE Trans. Multim.5
2012 Finding Celebrities in Billions of Web Images
abstract
In this paper, we present a face annotation system to automatically collect and label celebrity faces from the web. With the proposed system, we have constructed a large-scale dataset called “Celebrities on the Web,” which contains 2.45 million distinct images of 421 436 celebrities and is orders of magnitude larger than previous datasets.
Lei Zhang 0001, Xin-Jing Wang, Harry Shum
IEEE Trans. Multim.4
2011 Realistic and interactive image-based rendering of ancient chinese artifacts using a multiple camera array
abstract
This paper proposes a system for photorealistic interactive rendering of ancient Chinese artifacts for cultural heritage preservation using multiview images captured by a circular multiple-camera array. It employs 3D reconstruction and precomputed shadow field techniques to enable real-time relighting and object interaction. Moreover, Gabor features are employed to improve the robustness of line matching along epipolar lines and robust radial basis function modeling is employed to suppress possible outliers arising from false matching. Using the 3D model reconstructed, the precomputed shadow field is employed to provide real-time rendering/relighting and object movement, after acceleration on a graphic processing unit (GPU). Excellent rendering results are obtained and the ancient Chinese artifacts can be displayed in modern multi-view displays and conventional stereo systems.
Chong Wang 0001, S. C. Chan 0001, Harry Shum
ISCAS4
2011 Bing dialog model: intent, knowledge and user interaction
abstract
The decade-old Internet search outcomes, manifested in the form of "ten blue links," are no longer sufficient for Internet users. Many studies have shown that when users are ushered off the conventional search result pages through blue links, their needs are often partially met at best in a "hit-or-miss" fashion. To tackle this challenge, we have designed Bing, Microsoft's decision engine, to not just navigate users to a landing page through a blue link but to continue engaging with users to clarify intent and facilitate task completion. Underlying this new paradigm is the Bing Dialog Model that consists of three building blocks: an indexing system that comprehensively collects information from the web and systematically harvests knowledge, an intent model that statistically infers user intent and predicts next action, and an interaction model that elicits user intent through mathematically optimized presentations of web information and domain knowledge that matches user needs. In this talk, I'll describe Bing Dialog Model in details and demonstrate it in action through some innovative features since the launch of www.Bing.com.
Harry Shum
WSDM1
2011 Learning to Detect a Salient Object
abstract
In this paper, we study the salient object detection problem for images. We formulate this problem as a binary labeling task where we separate the salient object from the background. We propose a set of novel features, including multiscale contrast, center-surround histogram, and color spatial distribution, to describe a salient object locally, regionally, and globally. A conditional random field is learned to effectively combine these features for salient object detection. Further, we extend the proposed approach to detect a salient object from sequential images by introducing the dynamic salient features. We collected a large image database containing tens of thousands of carefully labeled images by multiple users and a video segment database, and conducted a set of experiments over them to demonstrate the effectiveness of the proposed approach.
Zejian Yuan, Jian Sun 0001, Jingdong Wang 0001, Nanning Zheng 0001, Xiaoou Tang, Harry Shum
IEEE Trans. Pattern Anal. Mach. Intell.7
2011 Scalable Face Image Retrieval with Identity-Based Quantization and Multireference Reranking
abstract
State-of-the-art image retrieval systems achieve scalability by using a bag-of-words representation and textual retrieval methods, but their performance degrades quickly in the face image domain, mainly because they produce visual words with low discriminative power for face images and ignore the special properties of faces. The leading features for face recognition can achieve good retrieval performance, but these features are not suitable for inverted indexing as they are high-dimensional and global and thus not scalable in either computational or storage cost. In this paper, we aim to build a scalable face image retrieval system. For this purpose, we develop a new scalable face representation using both local and global features. In the indexing stage, we exploit special properties of faces to design new component-based local features, which are subsequently quantized into visual words using a novel identity-based quantization scheme. We also use a very small Hamming signature (40 bytes) to encode the discriminative global feature for each face. In the retrieval stage, candidate images are first retrieved from the inverted index of visual words. We then use a new multireference distance to rerank the candidate images using the Hamming signature. On a one millon face database, we show that our local features and global Hamming signatures are complementary--the inverted index based on local features provides candidate images with good recall, while the multireference reranking with global Hamming signature leads to good precision. As a result, our system is not only scalable but also outperforms the linear scan retrieval system using the state-of-the-art face recognition feature in term of the quality.
Qifa Ke, Jian Sun 0001, Harry Shum
IEEE Trans. Pattern Anal. Mach. Intell.4
2011 Gradient Profile Prior and Its Applications in Image Super-Resolution and Enhancement
abstract
In this paper, we propose a novel generic image prior-gradient profile prior, which implies the prior knowledge of natural image gradients. In this prior, the image gradients are represented by gradient profiles, which are 1-D profiles of gradient magnitudes perpendicular to image structures. We model the gradient profiles by a parametric gradient profile model. Using this model, the prior knowledge of the gradient profiles are learned from a large collection of natural images, which are called gradient profile prior. Based on this prior, we propose a gradient field transformation to constrain the gradient fields of the high resolution image and the enhanced image when performing single image super-resolution and sharpness enhancement. With this simple but very effective approach, we are able to produce state-of-the-art results. The reconstructed high resolution images or the enhanced images are sharp while have rare ringing or jaggy artifacts.
Jian Sun 0009, Jian Sun 0001, Zongben Xu, Harry Shum
IEEE Trans. Image Process.4
2010 An Empirical Study on Learning to Rank of Tweets
Yajuan Duan, Long Jiang, Tao Qin 0001, Ming Zhou 0001, Harry Shum
COLING5
2010 Scalable face image retrieval with identity-based quantization and multi-reference re-ranking
abstract
State-of-the-art image retrieval systems achieve scalability by using bag-of-words representation and textual retrieval methods, but their performance degrades quickly in the face image domain, mainly because they 1) produce visual words with low discriminative power for face images, and 2) ignore the special properties of the faces. The leading features for face recognition can achieve good retrieval performance, but these features are not suitable for inverted indexing as they are high-dimensional and global, thus not scalable in either computational or storage cost. In this paper we aim to build a scalable face image retrieval system. For this purpose, we develop a new scalable face representation using both local and global features. In the indexing stage, we exploit special properties of faces to design new component-based local features, which are subsequently quantized into visual words using a novel identity-based quantization scheme. We also use a very small hamming signature (40 bytes) to encode the discriminative global feature for each face. In the retrieval stage, candidate images are firstly retrieved from the inverted index of visual words. We then use a new multi-reference distance to re-rank the candidate images using the hamming signature. On a one-millon face database, we show that our local features and global hamming signatures are complementary - the inverted index based on local features provides candidate images with good recall, while the multi-reference re-ranking with global hamming signature leads to good precision. As a result, our system is not only scalable but also outperforms the linear scan retrieval system using the state-of-the-art face recognition feature in term of the quality.
Qifa Ke, Jian Sun 0001, Harry Shum
CVPR4
2010 Interest seam image
abstract
We propose interest seam image, an efficient visual synopsis for video. To extract an interest seam image, a spatiotemporal energy map is constructed for the target video shot. Then an optimal seam which encompasses the highest energy is identified by an efficient dynamic programming algorithm. The optimal seam is used to extract a seam of pixels from each video frame to form one column of an image, based on which an interest seam image is finally composited. The interest seam image is efficient both in terms of computation and memory cost. Therefore it is able to power a wide variety of web-scale video content analysis applications, such as near duplicate video clip search, video genre recognition and classification, as well as video clustering, etc. The representation capacity of the proposed interest seam image is demonstrated in a large scale video retrieval task. Its advantages are clearly exhibited when compared with previous works, as reported in our experiments.
Gang Hua 0001, Lei Zhang 0001, Harry Shum
CVPR4
2010 Image-based rendering of ancient Chinese artifacts for multi-view displays - a multi-camera approach
abstract
Image-based rendering (IBR) is an emerging and promising technology for photo-realistic rendering of scenes and objects from a collection of densely sampled images and videos. This paper proposes an image-based approach to the rendering and multi-view display of ancient Chinese artifacts for cultural heritage preservation. A multiple-camera circular array was constructed to record images of the artifacts. Novel techniques for segmenting and rendering new views of the artifacts from the sampled images are developed. The multiple views so synthesized enable the ancient artifacts to be displayed in modern multi-view displays and conventional stereo systems. Several collections from the University Museum and Art Gallery at the University of Hong Kong are captured and excellent rendering results are obtained.
King To Ng, S. C. Chan 0001, Harry Shum
ISCAS4
2010 Object-Based Coding for Plenoptic Videos
abstract
A new object-based coding system for a class of dynamic image-based representations called plenoptic videos (PVs) is proposed. PVs are simplified dynamic light fields, where the videos are taken at regularly spaced locations along line segments instead of a 2-D plane. In the proposed object-based approach, objects at different depth values are segmented to improve the rendering quality. By encoding PVs at the object level, desirable functionalities such as scalability of contents, error resilience, and interactivity with an individual image-based rendering (IBR) object can be achieved. Besides supporting the coding of texture and binary shape maps for IBR objects with arbitrary shapes, the proposed system also supports the coding of grayscale alpha maps as well as depth maps (geometry information) to respectively facilitate the matting and rendering of the IBR objects. Both temporal and spatial redundancies among the streams in the PV are exploited to improve the coding performance, while avoiding excessive complexity in selective decoding of PVs to support fast rendering speed. Advanced spatial/temporal prediction methods such as global disparity-compensated prediction, as well as direct prediction and its extensions are developed. The bit allocation and rate control scheme employing a new convex optimization-based approach are also introduced. Experimental results show that considerable improvements in coding performance are obtained for both synthetic and real scenes, while supporting the stated object-based functionalities.
King To Ng, S. C. Chan 0001, Harry Shum
IEEE Trans. Circuits Syst. Video Technol.4
2009 A multi-sample, multi-tree approach to bag-of-words image representation for image retrieval
abstract
The state-of-the-art content based image retrieval systems has been significantly advanced by the introduction of SIFT features and the bag-of-words image representation. Converting an image into a bag-of-words, however, involves three non-trivial steps: feature detection, feature description, and feature quantization. At each of these steps, there is a significant amount of information lost, and the resulted visual words are often not discriminative enough for large scale image retrieval applications. In this paper, we propose a novel multi-sample multi-tree approach to computing the visual word codebook. By encoding more information of the original image feature, our approach generates a much more discriminative visual word codebook that is also efficient in terms of both computation and space consumption, without losing the original repeatability of the visual features. We evaluate our approach using both a ground-truth data set and a real-world large scale image database. Our results show that a significant improvement in both precision and recall can be achieved by using the codebook derived from our approach.
Qifa Ke, Jian Sun 0001, Harry Shum
ICCV4
2009 Spectral error correcting output codes for efficient multiclass recognition
abstract
The error correcting output codes (ECOC) is a general framework to extend any binary classifier to the multiclass case. Finding the optimal ECOC is known as a NP hard problem. In this paper, we present a spectral analysis approach for the design of ECOC. We construct a similarity graph of the classes and generate ECOC with a subset of thresholded eigenvectors of the graph Laplacian. Using the spectral analysis, the coding efficiency, classifier's diversity, Hamming distance among codewords, and binary classifiers' accuracy can be simultaneously considered. The resulting ECOC is efficient, thus only a small set of binary classifiers are to be evaluated when making a decision. In experiments with large multiclass problems, our method is between 3 and 12 times faster comparing to one-against-all, with comparable classification accuracy. Our method also shows a better performance than the most of leading methods, e.g., ClassMap, random dense ECOC, random sparse ECOC, and discriminant ECOC.
Harry Shum
ICCV3
2009 Efficient indexing for large scale visual search
abstract
With the popularity of “bag of visual terms” representations of images, many text indexing techniques have been applied in large-scale image retrieval systems. However, due to a fundamental difference between an image query (e.g. 1500 visual terms) and a text query (e.g. 3-5 terms), the usages of some text indexing techniques, e.g. inverted list, are misleading. In this work, we develop a novel indexing technique for this problem. The basic idea is to decompose a document-like representation of an image into two components, one for dimension reduction and the other for residual information preservation. The computing of similarity of two images can be transferred to measuring similarities of their components. The decomposition has two major merits: (1) these components have good properties which enable them to be efficiently indexed and retrieved; (2) The decomposition has better generalization ability than other dimension reduction algorithms. The decomposition can be achieved by either a graphical model or a matrix factorization approach. Theoretic analysis and extensive experiments over a 2.3 million image database show that this framework is scalable to index large scale image database to support fast and accurate visual search.
Zhiwei Li 0006, Lei Zhang 0001, Wei-Ying Ma, Harry Shum
ICCV5
2009 Classification via Minimum Incremental Coding Length
abstract
We present a simple new criterion for classification, based on principles from lossy data compression. The criterion assigns a test sample to the class that uses the minimum number of additional bits to code the test sample, subject to an allowable distortion. We demonstrate the asymptotic optimality of this criterion for Gaussian distributions and analyze its relationships to classical classifiers. The theoretical results clarify the connections between our approach and popular classifiers such as maximum a posteriori (MAP), regularized discriminant analysis (RDA), k-nearest neighbor (k-NN), and support vector machine (SVM), as well as unsupervised methods based on lossy coding. Our formulation induces several good effects on the resulting classifier. First, minimizing the lossy coding length induces a regularization effect which stabilizes the (implicit) density estimate in a small sample setting. Second, compression provides a uniform means of handling classes of varying dimension. The new criterion and its kernel and local versions perform competitively on synthetic examples, as well as on real imagery data such as handwritten digits and face images. On these problems, the performance of our simple classifier approaches the best reported results, without using domain-specific information. All MATLAB code and classification results are publicly available for peer evaluation at http://perception.csl.uiuc.edu/coding/home.htm.
John Wright 0001, Yi Ma 0001, Yangyu Tao, Zhouchen Lin, Harry Shum
SIAM J. Imaging Sci.5
2009 An Object-Based Approach to Image/Video-Based Synthesis and Processing for 3-D and Multiview Televisions
abstract
This paper proposes an object-based approach to a class of dynamic image-based representations called ldquoplenoptic videos,rdquo where the plenoptic video sequences are segmented into image-based rendering (IBR) objects each with its image sequence, depth map, and other relevant information such as shape and alpha information. This allows desirable functionalities such as scalability of contents, error resilience, and interactivity with individual IBR objects to be supported. Moreover, the rendering quality in scenes with large depth variations can also be improved considerably. A portable capturing system consisting of two linear camera arrays was developed to verify the proposed approach. An important step in the object-based approach is to segment the objects in video streams into layers or IBR objects. To reduce the time for segmenting plenoptic videos under the semiautomatic technique, a new object tracking method based on the level-set method is proposed. Due to possible segmentation errors around object boundaries, natural matting with Bayesian approach is also incorporated into our system. Furthermore, extensions of conventional image processing algorithms to these IBR objects are studied and illustrated with examples. Experimental results are given to illustrate the efficiency of the tracking, matting, rendering, and processing algorithms under the proposed object-based framework.
S. C. Chan 0001, Zhi-Feng Gan, King To Ng, Ka-Leung Ho, Harry Shum
IEEE Trans. Circuits Syst. Video Technol.5
2009 Picture Collage
abstract
In this paper, we address a novel problem of automatically creating a picture collage from a group of images. Picture collage is a kind of visual image summary-to arrange all input images on a given canvas, allowing overlay, to maximize visible visual information. We formulate the picture collage creation problem in a conditional random field model, which integrates image salience, canvas constraint, natural preference, and user interaction. Each image is represented by a group of weighted rectangles, which indicate the salient regions. Then picture collage is resolved by minimizing the energy, guided by the constraints. A two-step optimization method is proposed. First, a quick initialization algorithm based on the proposed 1D collage method is presented. Second, a very efficient Markov chain Monte Carlo method is designed for the refined optimization. We also integrate user interaction in the formulation and optimization to obtain an interactive collage reflecting personalized preference. Visual and quantitative experimental evaluations indicate the efficiency of the proposed collage creation technique.
Jingdong Wang 0001, Jian Sun 0001, Nanning Zheng 0001, Xiaoou Tang, Harry Shum
IEEE Trans. Multim.6
2009 Face poser: Interactive modeling of 3D facial expressions using facial priors
abstract
This article presents an intuitive and easy-to-use system for interactively posing 3D facial expressions. The user can model and edit facial expressions by drawing freeform strokes, by specifying distances between facial points, by incrementally editing curves on the face, or by directly dragging facial points in 2D screen space. Designing such an interface for 3D facial modeling and editing is challenging because many unnatural facial expressions might be consistent with the user's input. We formulate the problem in a maximum a posteriori framework by combining the user's input with priors embedded in a large set of facial expression data. Maximizing the posteriori allows us to generate an optimal and natural facial expression that achieves the goal specified by the user. We evaluate the performance of our system by conducting a thorough comparison of our method with alternative facial modeling techniques. To demonstrate the usability of our system, we also perform a user study of our system and compare with state-of-the-art facial expression modeling software (Poser 7).
Manfred Lau, Jinxiang Chai, Ying-Qing Xu, Harry Shum
ACM Trans. Graph.4
2009 Paint selection
abstract
In this paper, we present Paint Selection, a progressive painting-based tool for local selection in images. Paint Selection facilitates users to progressively make a selection by roughly painting the object of interest using a brush. More importantly, Paint Selection is efficient enough that instant feedback can be provided to users as they drag the mouse. We demonstrate that high quality selections can be quickly and effectively "painted" on a variety of multi-megapixel images.
Jiangyu Liu, Jian Sun 0001, Harry Shum
ACM Trans. Graph.3
2008 Image super-resolution using gradient profile prior
abstract
In this paper, we propose an image super-resolution approach using a novel generic image prior - gradient profile prior, which is a parametric prior describing the shape and the sharpness of the image gradients. Using the gradient profile prior learned from a large number of natural images, we can provide a constraint on image gradients when we estimate a hi-resolution image from a low-resolution image. With this simple but very effective prior, we are able to produce state-of-the-art results. The reconstructed hi-resolution image is sharp while has rare ringing or jaggy artifacts.
Jian Sun 0009, Zongben Xu, Harry Shum
CVPR3
2008 L1 regularized projection pursuit for additive model learning
abstract
In this paper, we present a L1regularized projection pursuit algorithm for additive model learning. Two new algorithms are developed for regression and classification respectively: sparse projection pursuit regression and sparse Jensen-Shannon Boosting. The introduced L1regularized projection pursuit encourages sparse solutions, thus our new algorithms are robust to overfitting and present better generalization ability especially in settings with many irrelevant input features and noisy data. To make the optimization with L1regularization more efficient, we develop an ldquoinformative feature firstrdquo sequential optimization algorithm. Extensive experiments demonstrate the effectiveness of our proposed approach.
Xiaoou Tang, Harry Shum
CVPR4
2008 Query dependent ranking using K-nearest neighbor
abstract
Many ranking models have been proposed in information retrieval, and recently machine learning techniques have also been applied to ranking model construction. Most of the existing methods do not take into consideration the fact that significant differences exist between queries, and only resort to a single function in ranking of documents. In this paper, we argue that it is necessary to employ different ranking models for different queries and onduct what we call query-dependent ranking. As the first such attempt, we propose a K-Nearest Neighbor (KNN) method for query-dependent ranking. We first consider an online method which creates a ranking model for a given query by using the labeled neighbors of the query in the query feature space and then rank the documents with respect to the query using the created model. Next, we give two offline approximations of the method, which create the ranking models in advance to enhance the efficiency of ranking. And we prove a theory which indicates that the approximations are accurate in terms of difference in loss of prediction, if the learning algorithm used is stable with respect to minor changes in training examples. Our experimental results show that the proposed online and offline methods both outperform the baseline method of using a single ranking function.
Xiubo Geng, Tie-Yan Liu, Tao Qin 0001, Andrew Arnold, Hang Li 0001, Harry Shum
SIGIR6
2008 Sketching reality: Realistic interpretation of architectural designs
abstract
In this article, we introduce sketching reality , the process of converting a freehand sketch into a realistic-looking model. We apply this concept to architectural designs. As the sketch is being drawn, our system periodically interprets its 2.5D-geometry by identifying new junctions, edges, and faces, and then analyzing the extracted topology. The user can add detailed geometry and textures through sketches as well. This is possible through the use of databases that match partial sketches to models of detailed geometry and textures. The final product is a realistic texture-mapped 2.5D-model of the building. We show a variety of buildings that have been created using this system.
Xuejin Chen, Sing Bing Kang, Ying-Qing Xu, Julie Dorsey, Harry Shum
ACM Trans. Graph.5
2008 Texture amendment: reducing texture distortion in constrained parameterization
abstract
Constrained parameterization is an effective way to establish texture coordinates between a 3D surface and an existing image or photograph. A known drawback to constrained parameterization is visual distortion that arises when the 3D geometry is mismatched to highly textured image regions. This paper introduces an approach to reduce visual distortion by expanding image regions via texture synthesis to better fit the 3D geometry. The result is a new amended texture that maintains the essence of the input texture image but exhibits significantly less distortion when mapped onto the 3D model.
Yu-Wing Tai, Michael S. Brown, Chi-Keung Tang, Harry Shum
ACM Trans. Graph.4
2008 Modeling and rendering of heterogeneous translucent materials using the diffusion equation
abstract
In this article, we propose techniques for modeling and rendering of heterogeneous translucent materials that enable acquisition from measured samples, interactive editing of material attributes, and real-time rendering. The materials are assumed to be optically dense such that multiple scattering can be approximated by a diffusion process described by the diffusion equation. For modeling heterogeneous materials, we present the inverse diffusion algorithm for acquiring material properties from appearance measurements. This modeling algorithm incorporates a regularizer to handle the ill-conditioning of the inverse problem, an adjoint method to dramatically reduce the computational cost, and a hierarchical GPU implementation for further speedup. To render an object with known material properties, we present the polygrid diffusion algorithm , which solves the diffusion equation with a boundary condition defined by the given illumination environment. This rendering technique is based on representation of an object by a polygrid, a grid with regular connectivity and an irregular shape, which facilitates solution of the diffusion equation in arbitrary volumes. Because of the regular connectivity, our rendering algorithm can be implemented on the GPU for real-time performance. We demonstrate our techniques by capturing materials from physical samples and performing real-time rendering and editing with these materials.
Jiaping Wang, Xin Tong 0001, Stephen Lin 0001, Zhouchen Lin, Yue Dong 0001, Baining Guo, Harry Shum
ACM Trans. Graph.8
2008 Inverse texture synthesis
abstract
The quality and speed of most texture synthesis algorithms depend on a 2D input sample that is small and contains enough texture variations. However, little research exists on how to acquire such sample. For homogeneous patterns this can be achieved via manual cropping, but no adequate solution exists for inhomogeneous or globally varying textures, i.e. patterns that are local but not stationary, such as rusting over an iron statue with appearance conditioned on varying moisture levels. We present inverse texture synthesis to address this issue. Our inverse synthesis runs in the opposite direction with respect to traditional forward synthesis: given a large globally varying texture, our algorithm automatically produces a small texture compaction that best summarizes the original. This small compaction can be used to reconstruct the original texture or to re-synthesize new textures under user-supplied controls. More important, our technique allows real-time synthesis of globally varying textures on a GPU, where the texture memory is usually too small for large textures. We propose an optimization framework for inverse texture synthesis, ensuring that each input region is properly encoded in the output compaction. Our optimization process also automatically computes orientation fields for anisotropic textures containing both low- and high-frequency regions, a situation difficult to handle via existing techniques.
Li-Yi Wei, Jianwei Han, Kun Zhou 0001, Hujun Bao, Baining Guo, Harry Shum
ACM Trans. Graph.6
2008 Interactive normal reconstruction from a single image
abstract
We present an interactive system for reconstructing surface normals from a single image. Our approach has two complementary contributions. First, we introduce a novel shape-from-shading algorithm (SfS) that produces faithful normal reconstruction for local image region (high-frequency component), but it fails to faithfully recover the overall global structure (low-frequency component). Our second contribution consists of an approach that corrects low-frequency error using a simple markup procedure. This approach, aptly calledrotation palette, allows the user to specify large scale corrections of surface normals by drawing simple stroke correspondences between the normal map and a sphere image which represents rotation directions. Combining these two approaches, we can produce high-quality surfaces quickly from single images.
Tai-Pang Wu, Jian Sun 0001, Chi-Keung Tang, Harry Shum
ACM Trans. Graph.4
2008 Progressive inter-scale and intra-scale non-blind image deconvolution
abstract
Ringing is the most disturbing artifact in the image deconvolution. In this paper, we present a progressive inter-scale and intra-scale non-blind image deconvolution approach that significantly reduces ringing. Our approach is built on a novel edge-preserving deconvolution algorithm called bilateral Richardson-Lucy (BRL) which uses a large spatial support to handle large blur. We progressively recover the image from a coarse scale to a fine scale (inter-scale), and progressively restore image details within every scale (intra-scale). To perform the inter-scale deconvolution, we propose a joint bilateral Richardson-Lucy (JBRL) algorithm so that the recovered image in one scale can guide the deconvolution in the next scale. In each scale, we propose an iterative residual deconvolution to progressively recover image details. The experimental results show that our progressive deconvolution can produce images with very little ringing for large blur kernels.
Lu Yuan 0001, Jian Sun 0001, Long Quan, Harry Shum
ACM Trans. Graph.4
2008 Real-time smoke rendering using compensated ray marching
abstract
We present a real-time algorithm calledcompensated ray marchingfor rendering of smoke under dynamic low-frequency environment lighting. Our approach is based on a decomposition of the input smoke animation, represented as a sequence of volumetric density fields, into a set of radial basis functions (RBFs) and a sequence of residual fields. To expedite rendering, the source radiance distribution within the smoke is computed from only the low-frequency RBF approximation of the density fields, since the high-frequency residuals have little impact on global illumination under low-frequency environment lighting. Furthermore, in computing source radiances the contributions from single and multiple scattering are evaluated at only the RBF centers and then approximated at other points in the volume using an RBF-based interpolation. A slice-based integration of these source radiances along each view ray is then performed to render the final image. The high-frequency residual fields, which are a critical component in the local appearance of smoke, are compensated back into the radiance integral during this ray march to generate images of high detail. The runtime algorithm, which includes both light transfer simulation and ray marching, can be easily implemented on the GPU, and thus allows for real-time manipulation of viewpoint and lighting, as well as interactive editing of smoke attributes such as extinction cross section, scattering albedo, and phase function. Only moderate preprocessing time and storage is needed. This approach provides the first method for real-time smoke rendering that includes single and multiple scattering while generating results comparable in quality to offline algorithms like ray tracing.
Kun Zhou 0001, Zhong Ren 0001, Stephen Lin 0001, Hujun Bao, Baining Guo, Harry Shum
ACM Trans. Graph.6
2008 Filtering and Rendering of Resolution-Dependent Reflectance Models
abstract
The apparent reflectance of a surface depends upon the resolution at which it is imaged. Conventional reflectance models represent reflection at a single predetermined resolution; however, a low-resolution pixel that views a greater surface area often exhibits a reflectance more complicated than a high-resolution pixel with a smaller area. To address resolution dependency in reflectance, we utilize a generalized reflectance model based on a mixture of multiple conventional models, and present a framework for efficiently determining the reflectance mixture model of each pixel with respect to resolution. Mixture model parameters are precomputed at multiple resolutions and stored in mipmaps. Unlike color textures, these reflectance parameters cannot be accurately filtered by trilinear interpolation, so we present a technique for nonlinear mipmap filtering that minimizes aliasing in rendered results. This framework can be applied with various parametric reflectance models in graphics hardware for real-time processing. With this technique for filtering and rendering with mipmaps of reflectance mixture models, our system can rapidly render the resolution-dependent reflectance effects that are customarily disregarded in conventional rendering methods. At the end of this paper, we also describe how shadowing and masking effects can be incorporated into this framework to increase the realism of rendering.
Ping Tan 0002, Stephen Lin 0001, Long Quan, Baining Guo, Harry Shum
IEEE Trans. Vis. Comput. Graph.5
2007 Learning to Detect A Salient Object
abstract
We study visual attention by detecting a salient object in an input image. We formulate salient object detection as an image segmentation problem, where we separate the salient object from the image background. We propose a set of novel features including multi-scale contrast, center-surround histogram, and color spatial distribution to describe a salient object locally, regionally, and globally. A conditional random field is learned to effectively combine these features for salient object detection. We also constructed a large image database containing tens of thousands of carefully labeled images by multiple users. To our knowledge, it is the first large image database for quantitative evaluation of visual attention algorithms. We validate our approach on this image database, which is public available with this paper.
Jian Sun 0001, Nanning Zheng 0001, Xiaoou Tang, Harry Shum
CVPR5
2007 Flash Cut: Foreground Extraction with Flash and No-flash Image Pairs
abstract
In this paper, we propose a novel approach for foreground layer extraction using flash/no-flash image pairs, which we call flash cut. Flash cut is based on the simple observation that only the foreground is significantly brightened by the flash and the background appearance change is very small, if the background is distant. Changes due to flash, motion, and color information are fused in an MRF framework to produce high quality segmentation results. Flash cut handles some amount of camera shake, and foreground motion, which makes it practical for anyone with a flash-equipped camera to use. We validate our approach on a variety of indoor and outdoor examples.
Jian Sun 0009, Jian Sun 0001, Sing Bing Kang, Zongben Xu, Xiaoou Tang, Harry Shum
CVPR6
2007 Interactive Offline Tracking for Color Objects
abstract
In this paper, we present an interactive offline tracking system for generic color objects. The system achieves 60- 100 fps on a 320 times 240 video. The user can therefore easily refine the tracking result in an interactive way. To fully exploit user input and reduce user interaction, the tracking problem is addressed in a global optimization framework. The optimization is efficiently performed through three steps. First, from user's input we train a fast object detector that locates candidate objects in the video based on proposed features called boosted color bin. Second, we exploit the temporal coherence to generate multiple object trajectories based on a global best-first strategy. Last, an optimal object path is found by dynamic programming.
Jian Sun 0001, Xiaoou Tang, Harry Shum
ICCV4
2007 Blurred/Non-Blurred Image Alignment using Sparseness Prior
abstract
Aligning a pair of blurred and non-blurred images is a prerequisite for many image and video restoration and graphics applications. The traditional alignment methods such as direct and feature-based approaches cannot be used due to the presence of motion blur in one image of the pair. In this paper, we present an effective and accurate alignment approach for a blurred/non-blurred image pair. We exploit a statistical characteristic of the real blur kernel - the marginal distribution of kernel value is sparse. Using this sparseness prior, we can search the best alignment which produces the sparsest blur kernel. The search is carried out in scale space with a coarse-to-fine strategy for efficiency. Finally, we demonstrate the effectiveness of our algorithm for image deblurring, video restoration, and image matting.
Lu Yuan 0001, Jian Sun 0001, Long Quan, Harry Shum
ICCV4
2007 An Object-based Approach to Plenoptic Video Processing
abstract
Image-based rendering (IBR) is an emerging technology for photo-realistic rendering of scenes from a collection of densely sampled images and videos. Recently, an object-based approach for a class of dynamic image-based representations called plenoptic videos was proposed in order to improve the rendering quality in large environment. Since images and videos are special cases of the plenoptic function, many conventional image processing algorithms such as coding, segmentation, etc have similar analogy in IBR. This paper is devoted to the extension of some commonly used image processing algorithms to IBR using the object-based approach and their possible applications. Experimental results using plenoptic videos as an example are also given to illustrate the basic concept.
S. C. Chan 0001, Zhi-Feng Gan, Harry Shum
ISCAS3
2007 Classification via Minimum Incremental Coding Length (MICL)
abstract
We present a simple new criterion for classification, based on principles from lossy data compression. The criterion assigns a test sample to the class that uses the min- imum number of additional bits to code the test sample, subject to an allowable distortion. We prove asymptotic optimality of this criterion for Gaussian data and analyze its relationships to classical classifiers. Theoretical results provide new insights into relationships among popular classifiers such as MAP and RDA, as well as unsupervised clustering methods based on lossy compression [13]. Mini- mizing the lossy coding length induces a regularization effect which stabilizes the (implicit) density estimate in a small-sample setting. Compression also provides a uniform means of handling classes of varying dimension. This simple classi- fication criterion and its kernel and local versions perform competitively against existing classifiers on both synthetic examples and real imagery data such as hand- written digits and human faces, without requiring domain-specific information.
John Wright 0001, Yangyu Tao, Zhouchen Lin, Yi Ma 0001, Harry Shum
NIPS5
2007 Fogshop: Real-Time Design and Rendering of Inhomogeneous, Single-Scattering Media
abstract
We describe a new, analytic approximation to the airlight integral from scattering media whose density is modeled as a sum of Gaussians. The approximation supports real-time rendering of inhomogeneous media including their shadowing and scattering effects. For each Gaussian, this approximation samples the scattering integrand at the projection of its center along the view ray but models attenuation and shadowing with respect to the other Gaussians by integrating density along the fixed path from light source to 3D center to view point. Our method handles isotropic, single-scattering media illuminated by point light sources or low-frequency lighting environments. We also generalize models for reflectance of surfaces from constant-density to inhomogeneous media, using simple optical depth averaging in the direction of the light source or all around the receiver point. Our real-time renderer is incorporated into a system for real-time design and preview of realistic animated fog, steam, or smoke.
Kun Zhou 0001, Qiming Hou, Minmin Gong, John Snyder, Baining Guo, Harry Shum
PG6
2007 Natural Image Colorization
Qing Luan, Fang Wen 0001, Daniel Cohen-Or, Ying-Qing Xu, Harry Shum
Rendering Techniques6
2007 High Dynamic Range Image Hallucination
Lvdi Wang, Li-Yi Wei, Kun Zhou 0001, Baining Guo, Harry Shum
Rendering Techniques5
2007 Face Hallucination: Theory and Practice
Ce Liu 0001, Harry Shum, William T. Freeman
Int. J. Comput. Vis.2
2007 Image vectorization using optimized gradient meshes
abstract
Recently, gradient meshes have been introduced as a powerful vector graphics representation to draw multicolored mesh objects with smooth transitions. Using tools from Abode Illustrator and Corel CorelDraw, a user can manually create gradient meshes even for photo-realistic vector arts, which can be further edited, stylized and animated. In this paper, we present an easy-to-use interactive tool, called optimized gradient mesh , to semi-automatically and quickly create gradient meshes from a raster image. We obtain the optimized gradient mesh by formulating an energy minimization problem. The user can also interactively specify a few vector lines to guide the mesh generation. The resulting optimized gradient mesh is an editable and scalable mesh that otherwise would have taken many hours for a user to manually create.
Jian Sun 0001, Fang Wen 0001, Harry Shum
ACM Trans. Graph.4
2007 Natural shadow matting
abstract
This article addresses the problem of natural shadow matting , the removal or extraction of natural shadows from a single image. Because textures are maintained in the shadowless image after the extraction process, our approach produces some of the best results to date among shadow removal techniques. Using the image formation equation typical of computer vision, we advocate a new model for shadow formation where shadow effect is understood as light attenuation instead of a mixture of two colors governed by the conventional matting equation. This leads to a new shadow equation with fewer unknowns to solve, where a three-channel shadow matte and a shadowless image are considered in our optimization. Our problem is formulated as one of energy minimization guided by user-supplied hints in the form of a quadmap which can be specified easily by the user. This formulation allows for robust shadow matte extraction while maintaining texture in the shadowed region by considering color transfer, texture gradient, and shadow smoothness. We demonstrate the usefulness of our approach in shadow removal, image matting, and compositing.
Tai-Pang Wu, Chi-Keung Tang, Michael S. Brown, Harry Shum
ACM Trans. Graph.4
2007 ShapePalettes: interactive normal transfer via sketching
abstract
We present a simple interactive approach to specify 3D shape in a single view using "shape palettes". The interaction is as follows: draw a simple 2D primitive in the 2D view and then specify its 3D orientation by drawing a corresponding primitive on ashape palette. The shape palette is presented as an image of some familiar shape whose local 3D orientation is readily understood and can be easily marked over. The 3D orientation from the shape palette is transferred to the 2D primitive based on the markup. As we will demonstrate, only sparse markup is needed to generate expressive and detailed 3D surfaces. This markup approach can be used to model freehand 3D surfaces drawn in a single view, or combined with image-snapping tools to quickly extract surfaces from images and photographs.
Tai-Pang Wu, Chi-Keung Tang, Michael S. Brown, Harry Shum
ACM Trans. Graph.4
2007 Image deblurring with blurred/noisy image pairs
abstract
Taking satisfactory photos under dim lighting conditions using a hand-held camera is challenging. If the camera is set to a long exposure time, the image is blurred due to camera shake. On the other hand, the image is dark and noisy if it is taken with a short exposure time but with a high camera gain. By combining information extracted from both blurred and noisy images, however, we show in this paper how to produce a high quality image that cannot be obtained by simply denoising the noisy image, or deblurring the blurred image alone. Our approach is image deblurring with the help of the noisy image. First, both images are used to estimate an accurate blur kernel, which otherwise is difficult to obtain from a single blurred image. Second, and again using both images, a residual deconvolution is proposed to significantly reduce ringing artifacts inherent to image deconvolution. Third, the remaining ringing artifacts in smooth image regions are further suppressed by a gain-controlled deconvolution process. We demonstrate the effectiveness of our approach using a number of indoor and outdoor images taken by off-the-shelf hand-held cameras in poor lighting environments.
Lu Yuan 0001, Jian Sun 0001, Long Quan, Harry Shum
ACM Trans. Graph.4
2007 Direct manipulation of subdivision surfaces on GPUs
abstract
We present an algorithm for interactive deformation of subdivision surfaces, including displaced subdivision surfaces and subdivision surfaces with geometric textures. Our system lets the user directly manipulate the surface using freely-selected surface points as handles. During deformation the control mesh vertices are automatically adjusted such that the deforming surface satisfies the handle position constraints while preserving the original surface shape and details. To best preserve surface details, we develop a gradient domain technique that incorporates the handle position constraints and detail preserving objectives into the deformation energy. For displaced subdivision surfaces and surfaces with geometric textures, the deformation energy is highly nonlinear and cannot be handled with existing iterative solvers. To address this issue, we introduce a shell deformation solver, which replaces each numerically unstable iteration step with two stable mesh deformation operations. Our deformation algorithm only uses local operations and is thus suitable for GPU implementation. The result is a real-time deformation system running orders of magnitude faster than the state-of-the-art multigrid mesh deformation solver. We demonstrate our technique with a variety of examples, including examples of creating visually pleasing character animations in real-time by driving a subdivision surface with motion capture data.
Kun Zhou 0001, Weiwei Xu 0003, Baining Guo, Harry Shum
ACM Trans. Graph.5
2006 Accurate Face Alignment using Shape Constrained Markov Network
abstract
In this paper, we present a shape constrained Markov network for accurate face alignment. The global face shape is defined as a set of weighted shape samples which are integrated into the Markov network optimization. These weighted samples provide structural constraints to make the Markov network more robust to local image noise. We propose a hierarchical Condensation algorithm to draw the shape samples efficiently. Specifically, a proposal density incorporating the local face shape is designed to generate more samples close to the image features for accurate alignment, based on a local Markov network search. A constrained regularization algorithm is also developed to weigh favorably those points that are already accurately aligned. Extensive experiments demonstrate the accuracy and effectiveness of our proposed approach.
Fang Wen 0001, Ying-Qing Xu, Xiaoou Tang, Harry Shum
CVPR (1)5
2006 Picture Collage
abstract
In this paper, we address a novel problem of automatically creating a picture collage from a group of images. Picture collage is a kind of visual image summary - to arrange all input images on a given canvas, allowing overlay, to maximize visible visual information. We formulate the picture collage creation problem in a Bayesian framework. The salient regions of each image are firstly extracted and represented as a set of weighted rectangles. Then, the image arrangement is formulated as a Maximum a Posterior (MAP) problem such that the output picture collage shows as many visible salient regions (without being overlaid by others) from all images as possible. Moreover, a very efficientMarkov chain Monte Carlo (MCMC) method is designed for the optimization. Applications to desktop image browsing and image search result summarization demonstrate the effectiveness of our approach.
Jingdong Wang 0001, Long Quan, Jian Sun 0001, Xiaoou Tang, Harry Shum
CVPR (1)5
2006 Background Cut
Jian Sun 0001, Xiaoou Tang, Harry Shum
ECCV (2)4
2006 A Joint Motion-Image Inpainting Method for Error Concealment in Video Coding
abstract
In this paper, we propose a new method for spatial-temporal error concealment in video coding using joint motion-image inpainting. The proposed method combines motion inpainting and adaptive Markov random field (MRF) based diffusion as robust motion inpainting. Image inpainting is employed to refine the result. With robust motion inpainting, effective compromise between temporal and spatial method is achieved for each point in the missing marcroblock (MB), which provides satisfactory visual quality of the restored frame with complex motion in video.
Liyong Chen, S. C. Chan 0001, Harry Shum
ICIP3
2006 A Convex Optimization-Based Frame-Level Rate Control Algorithm for Motion Compensated Hybrid DCT/DPCM Video Coding
abstract
This paper presents a convex optimization-based frame-level rate control algorithm for motion compensated hybrid DCT/DPCM video coding. By modifying the existing rate-distortion models, an improved empirical rate-distortion model with more flexibility is proposed to explain the experimental observations. The convexity and monotonicity of the proposed model are exploited to formulate the frame-level bit allocation problem as a convex programming problem. Thus the bit allocation among the frames with different picture types can be solved using convex programming methods such as the interior-point methods, if the optimal solution exists. Different importance weights for various picture types (I-, P-, etc) can also be incorporated into the scheme in order to account for the relative importance of different picture types due to their inter-dependency. The relevant model parameters are determined using previously encoded frames by means of linear regression. Simulation results show that the proposed algorithm achieves a considerably better picture quality in terms of PSNR than the conventional approaches for the tested video sequences. Therefore, it is a good alternative to these conventional approaches.
S. C. Chan 0001, Harry Shum
ICIP3
2006 Human Intention Modeling and Interactive Computer Vision
abstract
For many years, computer vision and robotics researchers have worked hard chasing the illusive goals such as "can the robot find a boy in the scene" or "can your vision system automatically segment the cat from the background". These tasks require a lot of prior knowledge and contextual information, and perhaps more importantly, understanding of human intention. How to model human intention into vision and robotic systems is, however, very challenging and can only be solved through human-computer interaction. In this talk, we propose that many difficult vision tasks can be solved with interactive vision systems, by combining powerful and real-time vision techniques with intuitive and clever user interfaces. We will show two interactive vision systems we developed recently, Lazy Snapping (Siggraph 2004) and Image Completion (Siggraph 2005). Lazy Snapping cuts out an object from a picture using graph cut, while Image Completion recovers unknown region in a picture with belief propagation. A key element in designing such interactive systems is how we model the user's intention using conditional probability (context) and likelihood associated with user interactions. Given how ill-posed most image understanding problems are, it is proposed that interactive computer vision is the paradigm we should focus today's vision research on where the key is the understanding and modeling of human intention.
Harry Shum
IROS1
2006 Real-time Multi-perspective Rendering on Graphics Hardware
Xianyou Hou, Li-Yi Wei, Harry Shum, Baining Guo
Rendering Techniques3
2006 An efficient large deformation method using domain decomposition
Jin Huang 0001, Xinguo Liu, Hujun Bao, Baining Guo, Harry Shum
Comput. Graph.5
2006 Realistic, real-time rendering of ocean waves
abstract
Abstract In computer games and other real‐time graphics applications, the ocean surface is typically modelled as a texture or bump‐mapped plane with simple lighting effects. This paper describes a system for realistically rendering the water surface in real time. Our system can render calm ocean waves with sophisticated lighting effects at 100 fps on a 680 MHz Pentium III with a GeForce 3 graphics card. The wave geometry is represented view‐dependently as a dynamic displacement map with surface detail described by a dynamic bump map. The illumination model includes reflection, refraction and Fresnel effects, which are critical for producing the look and feel of water. Copyright © 2006 John Wiley & Sons, Ltd.
Luiz Velho 0001, Xin Tong 0001, Baining Guo, Harry Shum
Comput. Animat. Virtual Worlds5
2006 Table Detection in Online Ink Notes
abstract
In documents, tables are important structured objects that present statistical and relational information. In this paper, we present a robust system which is capable of detecting tables from free style online ink notes and extracting their structure so that they can be further edited in multiple ways. First, the primative structure of tables, i.e., candidates for ruling lines and table bounding boxes, are detected among drawing strokes. Second, the logical structure of tables is determined by normalizing the table skeletons, identifying the skeleton structure, and extracting the cell contents. The detection process is similar to a decision tree so that invalid candidates can be ruled out quickly. Experimental results suggest that our system is robust and accurate in dealing with tables having complex structure or drawn under complex situations.
Zhouchen Lin, Junfeng He, Harry Shum
IEEE Trans. Pattern Anal. Mach. Intell.5
2006 Response to the Comments on "Fundamental Limits of Reconstruction-Based Superresolution Algorithms under Local Translation'
abstract
Wang and Feng (IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 28, no. 5, p 846, May 2006) pointed out that the deduction in (Z. Lin and H. Y. Shum, IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 26, no. 1, pp. 83-97, Jan. 2004) overlooked the validity of the perturbation theorem used in (Z. Lin and H. Y. Shum, IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 26, no. 1, pp. 83-97, Jan. 2004). In this paper, we show that, when the perturbation theorem is invalid, the probability of successful superresolution is very low. Therefore, we only have to derive the limits under the condition that validates the perturbation theorem, as done in (Z. Lin and H. Y. Shum, IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 26, no. 1, pp. 83-97, Jan. 2004).
Zhouchen Lin, Harry Shum
IEEE Trans. Pattern Anal. Mach. Intell.2
2006 Full-Frame Video Stabilization with Motion Inpainting
abstract
Video stabilization is an important video enhancement technology which aims at removing annoying shaky motion from videos. We propose a practical and robust approach of video stabilization that produces full-frame stabilized videos with good visual quality. While most previous methods end up with producing smaller size stabilized videos, our completion method can produce full-frame videos by naturally filling in missing image parts by locally aligning image data of neighboring frames. To achieve this, motion inpainting is proposed to enforce spatial and temporal consistency of the completion in both static and dynamic image areas. In addition, image quality in the stabilized video is enhanced with a new practical deblurring algorithm. Instead of estimating point spread functions, our method transfers and interpolates sharper image pixels of neighboring frames to increase the sharpness of the frame. The proposed video completion and deblurring methods enabled us to develop a complete video stabilizer which can naturally keep the original image quality in the stabilized videos. The effectiveness of our method is confirmed by extensive experiments over a wide variety of videos.
Yasuyuki Matsushita, Eyal Ofek, Weina Ge, Xiaoou Tang, Harry Shum
IEEE Trans. Pattern Anal. Mach. Intell.5
2006 Rule-based cleanup of on-line English ink notes
Zhouchen Lin, Harry Shum
Pattern Recognit.3
2006 Real-Time Bayesian 3-D Pose Tracking
abstract
In this paper, we propose a novel approach for real-time 3-D tracking of object pose from a single camera. We formulate the 3-D pose tracking task in a Bayesian framework which fuses feature correspondence information from both previous frame and some selected key-frames into the posterior distribution of pose. We also developed an inter-frame motion inference algorithm which can get reliable inter-frame feature correspondences and relative pose. Finally, the maximum a posteriori estimation of pose is obtained via stochastic sampling to achieve stable and drift-free tracking. Experiments show significant improvement of our algorithm over existing algorithms especially in the cases of tracking agile motion, severe occlusion, drastic illumination change, and large object scale change
Qiang Wang 0023, Xiaoou Tang, Harry Shum
IEEE Trans. Circuits Syst. Video Technol.4
2006 Learning dynamic audio-visual mapping with input-output Hidden Markov models
abstract
In this paper, we formulate the problem of synthesizing facial animation from an input audio sequence as a dynamic audio-visual mapping. We propose that audio-visual mapping should be modeled with an input-output hidden Markov model, or IOHMM. An IOHMM is an HMM for which the output and transition probabilities are conditional on the input sequence. We train IOHMMs using the expectation-maximization(EM) algorithm with a novel architecture to explicitly model the relationship between transition probabilities and the input using neural networks. Given an input sequence, the output sequence is synthesized by the maximum likelihood estimation. Experimental results demonstrate that IOHMMs can generate natural and good-quality facial animation sequences from the input audio.
Harry Shum
IEEE Trans. Multim.2
2006 Subspace gradient domain mesh deformation
abstract
In this paper we present a general framework for performing constrained mesh deformation tasks with gradient domain techniques. We present a gradient domain technique that works well with a wide variety of linear and nonlinear constraints. The constraints we introduce include the nonlinear volume constraint for volume preservation, the nonlinear skeleton constraint for maintaining the rigidity of limb segments of articulated figures, and the projection constraint for easy manipulation of the mesh without having to frequently switch between multiple viewpoints. To handle nonlinear constraints, we cast mesh deformation as a nonlinear energy minimization problem and solve the problem using an iterative algorithm. The main challenges in solving this nonlinear problem are the slow convergence and numerical instability of the iterative solver. To address these issues, we develop a subspace technique that builds a coarse control mesh around the original mesh and projects the deformation energy and constraints onto the control mesh vertices using the mean value interpolation. The energy minimization is then carried out in the subspace formed by the control mesh vertices. Running in this subspace, our energy minimization solver is both fast and stable and it provides interactive responses. We demonstrate our deformation constraints and subspace deformation technique with a variety of constrained deformation examples.
Jin Huang 0001, Xinguo Liu, Kun Zhou 0001, Li-Yi Wei, Shang-Hua Teng, Hujun Bao, Baining Guo, Harry Shum
ACM Trans. Graph.9
2006 Drag-and-drop pasting
abstract
In this paper, we present a user-friendly system for seamless image composition, which we call drag-and-drop pasting. We observe that for Poisson image editing [Perez et al. 2003] to work well, the user must carefully draw a boundary on the source image to indicate the region of interest, such that salient structures in source and target images do not conflict with each other along the boundary. To make Poisson image editing more practical and easy to use, we propose a new objective function to compute an optimized boundary condition. A shortest closed-path algorithm is designed to search for the location of the boundary. Moreover, to faithfully preserve the object's fractional boundary, we construct a blended guidance field to incorporate the object's alpha matte. To use our system, the user needs only to simply outline a region of interest in the source image, and then drag and drop it onto the target image. Experimental results demonstrate the effectiveness of our "drag-and-drop pasting" system.
Jiaya Jia, Jian Sun 0001, Chi-Keung Tang, Harry Shum
ACM Trans. Graph.4
2006 Flash matting
abstract
In this paper, we propose a novel approach to extract mattes using a pair of flash/no-flash images. Our approach, which we call flash matting , was inspired by the simple observation that the most noticeable difference between the flash and no-flash images is the foreground object if the background scene is sufficiently distant. We apply a new matting algorithm called joint Bayesian flash matting to robustly recover the matte from flash/no-flash images, even for scenes in which the foreground and the background are similar or the background is complex. Experimental results involving a variety of complex indoors and outdoors scenes show that it is easy to extract high-quality mattes using an off-the-shelf, flash-equipped camera. We also describe extensions to flash matting for handling more general scenes.
Jian Sun 0001, Yin Li 0003, Sing Bing Kang, Harry Shum
ACM Trans. Graph.4
2006 Appearance manifolds for modeling time-variant appearance of materials
abstract
We present a visual simulation technique called appearance manifolds for modeling the time-variant surface appearance of a material from data captured at a single instant in time. In modeling time-variant appearance, our method takes advantage of the key observation that concurrent variations in appearance over a surface represent different degrees of weathering. By reorganizing these various appearances in a manner that reveals their relative order with respect to weathering degree, our method infers spatial and temporal appearance properties of the material's weathering process that can be used to convincingly generate its weathered appearance at different points in time. Results with natural non-linear reflectance variations are demonstrated in applications such as visual simulation of weathering on 3D models, increasing and decreasing the weathering of real objects, and material transfer with weathering effects.
Jiaping Wang, Xin Tong 0001, Stephen Lin 0001, Minghao Pan, Chao Wang 0063, Hujun Bao, Baining Guo, Harry Shum
ACM Trans. Graph.8
2006 Animating Chinese paintings through stroke-based decomposition
abstract
This article proposes a technique to animate a Chinese style painting given its image. We first extract descriptions of the brush strokes that hypothetically produced it. The key to the extraction process is the use of a brush stroke library, which is obtained by digitizing single brush strokes drawn by an experienced artist. The steps in our extraction technique are first to segment the input image, then to find the best set of brush strokes that fit the regions, and, finally, to refine these strokes to account for local appearance. We model a single brush stroke using its skeleton and contour, and we characterize texture variation within each stroke by sampling perpendicularly along its skeleton. Once these brush descriptions have been obtained, the painting can be animated at the brush stroke level. In this article, we focus on Chinese paintings with relatively sparse strokes. The animation is produced using a graphical application we developed. We present several animations of real paintings using our technique.
Songhua Xu, Ying-Qing Xu, Sing Bing Kang, David Salesin, Yunhe Pan, Harry Shum
ACM Trans. Graph.6
2006 Mesh quilting for geometric texture synthesis
abstract
We introduce mesh quilting , a geometric texture synthesis algorithm in which a 3D texture sample given in the form of a triangle mesh is seamlessly applied inside a thin shell around an arbitrary surface through local stitching and deformation. We show that such geometric textures allow interactive and versatile editing and animation, producing compelling visual effects that are difficult to achieve with traditional texturing methods. Unlike pixel-based image quilting, mesh quilting is based on stitching together 3D geometry elements. Our quilting algorithm finds corresponding geometry elements in adjacent texture patches, aligns elements through local deformation, and merges elements to seamlessly connect texture patches. For mesh quilting on curved surfaces, a critical issue is to reduce distortion of geometry elements inside the 3D space of the thin shell. To address this problem we introduce a low-distortion parameterization of the shell space so that geometry elements can be synthesized even on very curved objects without the visual distortion present in previous approaches. We demonstrate how mesh quilting can be used to generate convincing decorations for a wide range of geometric textures.
Kun Zhou 0001, Yiying Tong, Mathieu Desbrun, Baining Guo, Harry Shum
ACM Trans. Graph.7
2006 Geometry-Driven Photorealistic Facial Expression Synthesis
abstract
Expression mapping (also called performance driven animation) has been a popular method for generating facial animations. A shortcoming of this method is that it does not generate expression details such as the wrinkles due to skin deformations. In this paper, we provide a solution to this problem. We have developed a geometry-driven facial expression synthesis system. Given feature point positions (the geometry) of a facial expression, our system automatically synthesizes a corresponding expression image that includes photorealistic and natural looking expression details. Due to the difficulty of point tracking, the number of feature points required by the synthesis system is, in general, more than what is directly available from a performance sequence. We have developed a technique to infer the missing feature point motions from the tracked subset by using an example-based approach. Another application of our system is expression editing where the user drags feature points while the system interactively generates facial expressions with skin deformation details.
Qingshan Zhang, Zicheng Liu 0001, Baining Guo, Demetri Terzopoulos, Harry Shum
IEEE Trans. Vis. Comput. Graph.5
2005 Object tracking and matting for a class of dynamic image-based representations
abstract
Image-based rendering (IBR) is an emerging technology for photo-realistic rendering of scenes from a collection of densely sampled images and videos. Recently, an object-based approach for a class of dynamic image-based representations called plenoptic videos was proposed. This paper proposes an automatic object tracking approach using the level-set method. Our tracking method, which utilizes both local and global features of the image sequences instead of global features exploited in previous approach, can achieve better tracking results for objects, especially with non-uniform energy distribution. Due to possible segmentation errors around object boundaries, natural matting with Bayesian approach is also incorporated into our system. Furthermore, a MPEG-4 like object-based algorithm is developed for compressing the plenoptic videos, which consist of the alpha maps, depth maps and textures of the segmented image-based objects from different video plenoptic streams. Experimental results show that satisfactory renderings can be obtained by the proposed approaches.
Zhi-Feng Gan, S. C. Chan 0001, Harry Shum
AVSS3
2005 Detecting Doctored Images Using Camera Response Normality and Consistency
abstract
The advance in image/video editing techniques has facilitated people in synthesizing realistic images/videos that may hard to be distinguished from real ones by visual examination. This poses a problem: how to differentiate real images/videos from doctored ones? This is a serious problem because some legal issues may occur if there is no reliable way for doctored image/video detection when human inspection fails. Digital watermarking cannot solve this problem completely. We propose an approach that computes the response functions of the camera by selecting appropriate patches in different ways. An image may be doctored if the response functions are abnormal or inconsistent to each other. The normality of the response functions is classified by a trained support vector machine (SVM). Experiments show that our method is effective for high-contrast images with many textureless edges.
Zhouchen Lin, Xiaoou Tang, Harry Shum
CVPR (1)4
2005 Full-Frame Video Stabilization
abstract
Video stabilization is an important video enhancement technology which aims at removing annoying shaky motion from videos. We propose a practical and robust approach of video stabilization that produces full-frame stabilized videos with good visual quality. While most previous methods end up with producing low resolution stabilized videos, our completion method can produce full-frame videos by naturally filling in missing image parts by locally aligning image data of neighboring frames. To achieve this, motion inpainting is proposed to enforce spatial and temporal consistency of the completion in both static and dynamic image areas. In addition, image quality in the stabilized video is enhanced with a new practical deblurring algorithm. Instead of estimating point spread functions, our method transfers and interpolates sharper image pixels of neighbouring frames to increase the sharpness of the frame. The proposed video completion and deblurring methods enabled us to develop a complete video stabilizer which can naturally keep the original image quality in the stabilized videos. The effectiveness of our method is confirmed by extensive experiments over a wide variety of videos.
Yasuyuki Matsushita, Eyal Ofek, Xiaoou Tang, Harry Shum
CVPR (1)4
2005 Concurrent Subspaces Analysis
abstract
A representative subspace is significant for image analysis, while the corresponding techniques often suffer from the curse of dimensionality dilemma. In this paper, we propose a new algorithm, called concurrent subspaces analysis (CSA), to derive representative subspaces by encoding image objects as 2/sup nd/ or even higher order tensors. In CSA, an original higher dimensional tensor is transformed into a lower dimensional one using multiple concurrent subspaces that characterize the most representative information of different dimensions, respectively. Moreover, an efficient procedure is provided to learn these subspaces in an iterative manner. As analyzed in this paper, each sub-step of CSA takes the column vectors of the matrices, which are acquired from the k-mode unfolding of the tensors, as the new objects to be analyzed, thus the curse of dimensionality dilemma can be effectively avoided. The extensive experiments on the 3/sup rd/ order tensor data, simulated video sequences and Gabor filtered digital number image database show that CSA outperforms principal component analysis in terms of both reconstruction and classification capability.
Dong Xu 0001, Shuicheng Yan, Lei Zhang 0001, HongJiang Zhang, Zhengkai Liu, Harry Shum
CVPR (2)6
2005 Interactive Shape from Shading
abstract
Shape from shading (SfS) has always been difficult for real applications due to its intrinsic ill-posedness. In this paper, we propose an interactive SfS method which efficiently uses human knowledge in order to resolve ambiguity. We propose a global solution of continuous surfaces with a few constraints of surface normals that are interactively imposed to regularize the problem. A surface is divided into local patches, and each local solution is estimated with a fast marching SfS. It is shown that the boundaries of local solutions constitute a weighted Voronoi diagram, which allows for the formation of a global solution from the local ones. Finally, we optimize this global estimation by minimizing an energy functional based on shading and smoothness priors. Reconstruction results from both synthetic and real images demonstrate the usability of the new approach for various modeling applications.
Yasuyuki Matsushita, Long Quan, Harry Shum
CVPR (1)4
2005 A Bayesian Mixture Model for Multi-View Face Alignment
abstract
For multi-view face alignment, we have to deal with two major problems: 1) the problem of multi-modality caused by diverse shape variation when the view changes dramatically; 2) the varying number of feature points caused by self-occlusion. Previous works have used nonlinear models or view based methods for multi-view face alignment. However, they either assume all feature points are visible or apply a set of discrete models separately without a uniform criterion. In this paper, we propose a unified framework to solve the problem of multi-view face alignment, in which, both the multi-modality and variable feature points are modeled by a Bayesian mixture model. We first develop a mixture model to describe the shape distribution and the feature point visibility, and then use an efficient EM algorithm to estimate the model parameters and the regularized shape. We use a set of experiments on several datasets to demonstrate the improvement of our method over traditional methods.
Yi Zhou 0020, Wayne Zhang 0001, Xiaoou Tang, Harry Shum
CVPR (2)4
2005 Bi-Directional Tracking Using Trajectory Segment Analysis
abstract
In this paper, we present a novel approach to keyframe-based tracking, called bi-directional tracking. Given two object templates in the beginning and ending keyframes, the bi-directional tracker outputs the MAP (maximum a posterior) solution of the whole state sequence of the target object in the Bayesian framework. First, a number of 3D trajectory segments of the object are extracted from the input video, using a novel trajectory segment analysis. Second, these disconnected trajectory segments due to occlusion are linked by a number of inferred occlusion segments. Last, the MAP solution is obtained by trajectory optimization in a coarse-to-fine manner. Experimental results show the robustness of our approach with respect to sudden motion, ambiguity, and short and long periods of occlusion.
Jian Sun 0001, Xiaoou Tang, Harry Shum
ICCV4
2005 Patch Based Blind Image Super Resolution
abstract
In this paper, a novel method for learning based image super resolution (SR) is presented. The basic idea is to bridge the gap between a set of low resolution (LR) images and the corresponding high resolution (HR) image using both the SR reconstruction constraint and a patch based image synthesis constraint in a general probabilistic framework. We show that in this framework, the estimation of the LR image formation parameters is straightforward. The whole framework is implemented via an annealed Gibbs sampling method. Experiments on SR on both single image and image sequence input show that the proposed method provides an automatic and stable way to compute super-resolution and the achieved result is encouraging for both synthetic and real LR images.
Qiang Wang 0023, Xiaoou Tang, Harry Shum
ICCV3
2005 Automatic 3D Face Modeling from Video
abstract
In this paper, we develop an efficient technique for fully automatic recovery of accurate 3D face shape from videos captured by a low cost camera. The method is designed to work with a short video containing a face rotating from frontal view to profile view. The whole approach consists of three components. First, automatic initialization is performed in the first frame with approximately frontal face. Then, to handle the case of low quality image captured by low cost camera, the 2D feature matching, head poses and underlying 3D face shape are estimated and refined iteratively in an efficient way based on image sequence segmentation. Finally, to take advantage of the sparse structure of the proposed algorithm, sparse bundle adjustment technique is further employed to speed up the computation. We demonstrate the accuracy and robustness of the algorithm using a set of experiments
Le Xin, Qiang Wang 0023, Jianhua Tao 0001, Xiaoou Tang, Tieniu Tan, Harry Shum
ICCV6
2005 On object-based compression for a class of dynamic image-based representations
abstract
An object-based compression scheme for a class of dynamic image-based representations called "plenoptic videos" (PVs) is studied in this paper. PVs are simplified dynamic light fields in which the videos are taken at regularly spaced locations along a line segment instead of a 2-D plane. To improve the rendering quality in scenes with large depth variations and support the functionalities at the object level for rendering, an object-based compression scheme is employed for the coding of PVs. Besides texture and shape information, the compression of geometry information in the form of depth maps is also supported. The proposed compression scheme exploits both the temporal and spatial redundancy among video object streams in the PV to achieve higher compression efficiency. Experimental results show that considerable improvements in coding performance are obtained for both synthetic and real scenes. Moreover, object-based functionalities such as rendering individual image-based objects are also illustrated.
King To Ng, S. C. Chan 0001, Harry Shum
ICIP (3)4
2005 Multiresolution Reflectance Filtering
abstract
Physically-based reflectance models typically represent light scattering as a function of surface geometry at the pixel level. With changes in viewing resolution, the geometry imaged within a pixel can undergo significant variations that result in changing reflectance characteristics. To address these transformations, we present a multiresolution reflectance framework based on microfacet normal distributions within a pixel over different scales. Since these distributions must be efficiently determined with respect to resolution, they are recorded at multiple resolution levels in mipmaps. The main contribution of this work is a real-time mipmap filtering technique for these distribution-based parameters that not only provides smooth reflectance transitions in scale, but also minimizes aliasing. With this multiresolution reflectance technique, our system can rapidly and accurately incorporate fine reflectance detail that is customarily disregarded in multiresolution rendering methods.
Ping Tan 0002, Stephen Lin 0001, Long Quan, Baining Guo, Harry Shum
Rendering Techniques5
2005 Interactive deformation of light fields
abstract
We present a software pipeline that enables an animator to deform light fields. The pipeline can be used to deform complex objects, such as furry toys, while maintaining photo-realistic quality. Our pipeline consists of three stages. First, we split the light field into sub-light fields. To facilitate splitting of complex objects, we employ a novel technique based on projected light patterns. Second, we deform each sub-light field. To do this, we provide the animator with controls similar to volumetric free-form deformation. Third, we recombine and render each sub-light field. Our rendering technique properly handles visibility changes due to occlusion among sub-light fields. To ensure consistent illumination of objects after they have been deformed, our light fields are captured with the light source fixed to the camera, rather than being fixed to the object. We demonstrate our deformation pipeline using synthetic and photographically acquired light fields. Potential applications include animation, interior design, and interactive gaming.
Billy Chen, Eyal Ofek, Harry Shum, Marc Levoy
SI3D3
2005 Clustering method for fast deformation with constraints
abstract
We present a fast deformation method for flexible objects. The deformation of the object is physically modeled using a linear elasticity model with a displacement based finite elements method, yielding a linear system at each time step of simulation. We solve this linear system using a precomputed force-displacement matrix, which describes the object response in terms of displacement accelerations to the forces acting on each vertex. We exploit the spatial coherence to effectively compress the force-displacement matrix to make this method practical and efficient by applying the clustered principal component analysis method. And we developed a method to efficiently handle the additional constraints for interactive user manipulation. At last large deformations are addressed based upon the compressed force-displacement matrix by combining a domain decomposition method and tracking the rotational motions. The experimental results demonstrate fast performances on complex large scale objects under interactive user manipulations.
Jin Huang 0001, Xinguo Liu, Hujun Bao, Baining Guo, Harry Shum
Symposium on Solid and Physical Modeling5
2005 Combining shape and physical modelsfor online cursive handwriting synthesis
Jue Wang 0001, Ying-Qing Xu, Harry Shum
Int. J. Document Anal. Recognit.4
2005 Polygonal Shape Blending with Topological Evolutions
Ligang Liu 0001, Bo Zhang 0025, Baining Guo, Harry Shum
J. Comput. Sci. Technol.4
2005 Outward-Looking Circular Motion Analysis of Large Image Sequences
abstract
This paper presents a novel and simple method of analyzing the motion of a large image sequence captured by a calibrated outward-looking video camera moving on a circular trajectory for large-scale environment applications. Previous circular motion algorithms mainly focus on inward-looking turntable-like setups. They are not suitable for outward-looking motion where the conic trajectory of corresponding points degenerates to straight lines. The circular motion of a calibrated camera essentially has only one unknown rotation angle for each frame. The motion recovery for the entire sequence computes only one fundamental matrix of a pair of frames to extract the angular motion of the pair using Laguerre's formula and then propagates the computation of the unknown rotation angles to the other frames by tracking one point over at least three frames. Finally, a maximum-likelihood estimation is developed for the optimization of the whole sequence. Extensive experiments demonstrate the validity of the method and the feasibility of the application in image-based rendering.
Guang Jiang, Long Quan, Hung-Tat Tsui, Harry Shum
IEEE Trans. Pattern Anal. Mach. Intell.5
2005 The plenoptic video
abstract
This paper presents a system for capturing and rendering a dynamic image-based representation called the plenoptic video. It is a simplified light field for dynamic environments, where user viewpoints are constrained to the camera plane of a linear array of video cameras. Important issues such as multiple camera calibration, real-time compression, decompression and rendering are addressed. The system consists of a camera array of eight Sony CCX-Z11 CCD cameras and eight Pentium 4 1.8-GHz computers connected together through a 100 Base-T local area network. It is possible to perform software-assisted real-time MPEG-2 compression at a resolution of (720/spl times/480). Using selective transmission, we are able to stream continuously plenoptic video with (256/spl times/256) resolution at a rate of 15 f/s over the network. For rendering from raw data on the hard disk, real-time rendering can be achieved with a resolution of (720/spl times/480) and a rate of 15 f/s. A new compression algorithm using both temporal and spatial predictions is also proposed for the efficient compression of the plenoptic videos. Experimental results demonstrate the usefulness of the proposed parallel processing based system in capturing and rendering high-quality dynamic image-based representations using off-the-shelf equipment, and its potential applications in visualization and immersive television systems.
S. C. Chan 0001, King To Ng, Zhi-Feng Gan, Kin-Lok Chan, Harry Shum
IEEE Trans. Circuits Syst. Video Technol.5
2005 Data compression and transmission aspects of panoramic videos
King To Ng, S. C. Chan 0001, Harry Shum
IEEE Trans. Circuits Syst. Video Technol.3
2005 Accelerate Video Decoding With Generic GPU
abstract
Most modern computers or game consoles are equipped with powerful yet cost-effective graphics processing units (GPUs) to accelerate graphics operations. Though the graphics engines in these GPUs are specially designed for graphics operations, can we harness their computing power for more general nongraphics operations? The answer is positive. In this paper, we present our study on leveraging the GPUs graphics engine to accelerate the video decoding. Specifically, a video decoding framework that involves both the central processing unit (CPU) and the GPU is proposed. By moving the whole motion compensation feedback loop of the decoder to the GPU, the CPU and GPU have been made to work in parallel in a pipelining fashion. Several techniques are also proposed to overcome the GPUs constraints or to optimize the GPU computation. Initial experimental results show that significant speed-up can be achieved by utilizing the GPU power. We have achieved real-time playback of high definition video on a PC with an Intel Pentium III 667-MHz CPU and an nVidia GeForce3 GPU.
Guobin Shen, Guang-ping Gao, Shipeng Li 0001, Harry Shum, Ya-Qin Zhang
IEEE Trans. Circuits Syst. Video Technol.4
2005 A virtual reality system using the concentric mosaic: construction, rendering, and data compression
abstract
This paper proposes a new image-based rendering (IBR) technique called "concentric mosaic" for virtual reality applications. IBR using the plenoptic function is an efficient technique for rendering new views of a scene from a collection of sample images previously captured. It provides much better image quality and lower computational requirement for rendering than conventional three-dimensional (3-D) model-building approaches. The concentric mosaic is a 3-D plenoptic function with viewpoints constrained on a plane. Compared with other more sophisticated four-dimensional plenoptic functions such as the light field and the lumigraph, the file size of a concentric mosaic is much smaller. In contrast to a panorama, the concentric mosaic allows users to move freely in a circular region and observe significant parallax and lighting changes without recovering the geometric and photometric scene models. The rendering of concentric mosaics is very efficient, and involves the reordering and interpolating of previously captured slit images in the concentric mosaic. It typically consists of hundreds of high-resolution images which consume a significant amount of storage and bandwidth for transmission. An MPEG-like compression algorithm is therefore proposed in this paper taking into account the access patterns and redundancy of the mosaic images. The compression algorithms of two equivalent representations of the concentric mosaic, namely the multiperspective panoramas and the normal setup sequence, are investigated. A multiresolution representation of concentric mosaics using a nonlinear filter bank is also proposed.
Harry Shum, King To Ng, S. C. Chan 0001
IEEE Trans. Multim.1
2005 Visual simulation of weathering by gamma-ton tracing
abstract
Weathering modeling introduces blemishes such as dirt, rust, cracks and scratches to virtual scenery. In this paper we present a visual stimulation technique that works well for a wide variety of weathering phenomena. Our technique, called γ-ton tracing, is based on a type of aging-inducing particles called γ-tons. Modeling a weathering effect with γ-ton tracing involves tracing a large number of γ-tons through the scene in a way similar to photon tracing and then generating the weathering effect using the recorded γ-ton transport information. With this technique, we can produce weathering effects that are customized to the scene geometry and tailored to the weathering sources. Several effects that are challenging for existing techniques can be readily captured by γ-ton tracing. These include global transport effects. or "stainbleeding". γ-ton tracing also enables visual simulations of complex multi-weathering effects. Lastly γ-ton tracing can generate weathering effects that not only involve texture changes but also large-scale geometry changes. We demonstrate our technique with a variety of examples.
Yanyun Chen, Lin Xia, Tien-Tsin Wong, Xin Tong 0001, Hujun Bao, Baining Guo, Harry Shum
ACM Trans. Graph.7
2005 Video object cut and paste
abstract
In this paper, we present a system for cutting a moving object out from a video clip. The cutout object sequence can be pasted onto another video or a background image. To achieve this, we first apply a new 3D graph cut based segmentation approach on the spatial-temporal video volume. Our algorithm partitions watershed presegmentation regions into foreground and background while preserving temporal coherence. Then, the initial segmentation result is refined locally. Given two frames in the video sequence, we specify two respective windows of interest which are then tracked using a bi-directional feature tracking algorithm. For each frame in between these two given frames, the segmentation in each tracked window is refined using a 2D graph cut that utilizes a local color model. Moreover, we provide brush tools for the user to control the object boundary precisely wherever needed. Based on the accurate binary segmentation result, we apply coherent matting to extract the alpha mattes and foreground colors of the object.
Yin Li 0003, Jian Sun 0001, Harry Shum
ACM Trans. Graph.3
2005 Image completion with structure propagation
abstract
In this paper, we introduce a novel approach to image completion, which we call structure propagation. In our system, the user manually specifies important missing structure information by extending a few curves or line segments from the known to the unknown regions. Our approach synthesizes image patches along these user-specified curves in the unknown region using patches selected around the curves in the known region. Structure propagation is formulated as a global optimization problem by enforcing structure and consistency constraints. If only a single curve is specified, structure propagation is solved using Dynamic Programming. When multiple intersecting curves are specified, we adopt the Belief Propagation algorithm to find the optimal patches. After completing structure propagation, we fill in the remaining unknown regions using patch-based texture synthesis. We show that our approach works well on a number of examples that are challenging to state-of-the-art techniques.
Jian Sun 0001, Lu Yuan 0001, Jiaya Jia, Harry Shum
ACM Trans. Graph.4
2005 Modeling and rendering of quasi-homogeneous materials
abstract
Many translucent materials consist of evenly-distributed heterogeneous elements which produce a complex appearance under different lighting and viewing directions. For these quasi-homogeneous materials, existing techniques do not address how to acquire their material representations from physical samples in a way that allows arbitrary geometry models to be rendered with these materials. We propose a model for such materials that can be readily acquired from physical samples. This material model can be applied to geometric models of arbitrary shapes, and the resulting objects can be efficiently rendered without expensive subsurface light transport simulation. In developing a material model with these attributes, we capitalize on a key observation about the subsurface scattering characteristics of quasi-homogeneous materials at different scales. Locally, the non-uniformity of these materials leads to inhomogeneous subsurface scattering. For subsurface scattering on a global scale, we show that a lengthy photon path through an even distribution of heterogeneous elements statistically resembles scattering in a homogeneous medium. This observation allows us to represent and measure the global light transport within quasi-homogeneous materials as well as the transfer of light into and out of a material volume through surface mesostructures. We demonstrate our technique with results for several challenging materials that exhibit sophisticated appearance features such as transmission of back illumination through surface mesostructures.
Xin Tong 0001, Jiaping Wang, Stephen Lin 0001, Baining Guo, Harry Shum
ACM Trans. Graph.5
2005 Real-time rendering of plant leaves
abstract
This paper presents a framework for the real-time rendering of plant leaves with global illumination effects. Realistic rendering of leaves requires a sophisticated appearance model and accurate lighting computation. For leaf appearance we introduce a parametric model that describes leaves in terms of spatially-variant BRDFs and BTDFs. These BRDFs and BTDFs, incorporating analysis of subsurface scattering inside leaf tissues and rough surface scattering on leaf surfaces, can be measured from real leaves. More importantly, this description is compact and can be loaded into graphics hardware for fast run-time shading calculations, which are essential for achieving high frame rates. For lighting computation, we present an algorithm that extends the Precomputed Radiance Transfer (PRT) approach to all-frequency lighting for leaves. In particular, we handle the combined illumination effects due to low-frequency environment light and high-frequency sunlight. This is done by decomposing the local incident radiance of sunlight into direct and indirect components. The direct component, which contains most of the high frequencies, is not pre-computed with spherical harmonics as in PRT; instead it is evaluated on-the-fly using pre-computed light-visibility convolution data. We demonstrate our framework by the rendering of a variety of leaves and assemblies thereof.
Lifeng Wang 0001, Wenle Wang, Julie Dorsey, Baining Guo, Harry Shum
ACM Trans. Graph.6
2005 Modeling hair from multiple views
abstract
In this paper, we propose a novel image-based approach to model hair geometry from images taken at multiple viewpoints. Unlike previous hair modeling techniques that require intensive user interactions or rely on special capturing setup under controlled illumination conditions, we use a handheld camera to capture hair images under uncontrolled illumination conditions. Our multi-view approach is natural and flexible for capturing. It also provides inherent strong and accurate geometric constraints to recover hair models.In our approach, the hair fibers are synthesized from local image orientations. Each synthesized fiber segment is validated and optimally triangulated from all visible views. The hair volume and the visibility of synthesized fibers can also be reliably estimated from multiple views. Flexibility of acquisition, little user interaction, and high quality results of recovered complex hair models are the key advantages of our method.
Eyal Ofek, Long Quan, Harry Shum
ACM Trans. Graph.4
2005 Precomputed shadow fields for dynamic scenes
abstract
We present a soft shadow technique for dynamic scenes with moving objects under the combined illumination of moving local light sources and dynamic environment maps. The main idea of our technique is to precompute for each scene entity a shadow field that describes the shadowing effects of the entity at points around it. The shadow field for a light source, called a source radiance field (SRF), records radiance from an illuminant as cube maps at sampled points in its surrounding space. For an occluder, an object occlusion field (OOF) conversely represents in a similar manner the occlusion of radiance by an object. A fundamental difference between shadow fields and previous shadow computation concepts is that shadow fields can be precomputed independent of scene configuration. This is critical for dynamic scenes because, at any given instant, the shadow information at any receiver point can be rapidly computed as a simple combination of SRFs and OOFs according to the current scene configuration. Applications that particularly benefit from this technique include large dynamic scenes in which many instances of an entity can share a single shadow field. Our technique enables low-frequency shadowing effects in dynamic scenes in real-time and all-frequency shadows at interactive rates.
Kun Zhou 0001, Stephen Lin 0001, Baining Guo, Harry Shum
ACM Trans. Graph.5
2005 Large mesh deformation using the volumetric graph Laplacian
abstract
We present a novel technique for large deformations on 3D meshes using the volumetric graph Laplacian. We first construct a graph representing the volume inside the input mesh. The graph need not form a solid meshing of the input mesh's interior; its edges simply connect nearby points in the volume. This graph's Laplacian encodes volumetric details as the difference between each point in the graph and the average of its neighbors. Preserving these volumetric details during deformation imposes a volumetric constraint that prevents unnatural changes in volume. We also include in the graph points a short distance outside the mesh to avoid local self-intersections. Volumetric detail preservation is represented by a quadric energy function. Minimizing it preserves details in a least-squares sense, distributing error uniformly over the whole deformed mesh. It can also be combined with conventional constraints involving surface positions, details or smoothness, and efficiently minimized by solving a sparse linear system.We apply this technique in a 2D curve-based deformation system allowing novice users to create pleasing deformations with little effort. A novel application of this system is to apply nonrigid and exaggerated deformations of 2D cartoon characters to 3D meshes. We demonstrate our system's potential with several examples.
Kun Zhou 0001, Jin Huang 0001, John M. Snyder, Xinguo Liu, Hujun Bao, Baining Guo, Harry Shum
ACM Trans. Graph.7
2005 TextureMontage
abstract
We propose a technique, called TextureMontage , to seamlessly map a patchwork of texture images onto an arbitrary 3D model. A texture atlas can be created through the specification of a set of correspondences between the model and any number of texture images. First, our technique automatically partitions the mesh and the images, driven solely by the choice of feature correspondences. Most charts will then be parameterized over their corresponding image planes through the minimization of a distortion metric based on both geometric distortion and texture mismatch across patch boundaries and images. Lastly, a surface texture inpainting technique is used to fill in the remaining charts of the surface with no corresponding texture patches. The resulting texture mapping satisfies the (sparse or dense) user-specified constraints while minimizing the distortion of the texture images and ensuring a smooth transition across the boundaries of different mesh patches. Seamless Texturing of Arbitrary Surfaces From Multiple Images
Kun Zhou 0001, Yiying Tong, Mathieu Desbrun, Baining Guo, Harry Shum
ACM Trans. Graph.6
2005 Light Field Morphing Using 2D Features
abstract
We present a 2D feature-based technique for morphing 3D objects represented by light fields. Existing light field morphing methods require the user to specify corresponding 3D feature elements to guide morph computation. Since slight errors in 3D specification can lead to significant morphing artifacts, we propose a scheme based on 2D feature elements that is less sensitive to imprecise marking of features. First, 2D features are specified by the user in a number of key views in the source and target light fields. Then the two light fields are warped view by view as guided by the corresponding 2D features. Finally, the two warped light fields are blended together to yield the desired light field morph. Two key issues in light field morphing are feature specification and warping of light field rays. For feature specification, we introduce a user interface for delineating 2D features in key views of a light field, which are automatically interpolated to other views. For ray warping, we describe a 2D technique that accounts for visibility changes and present a comparison to the ideal morphing of light fields. Light field morphing based on 2D features makes it simple to incorporate previous image morphing techniques such as nonuniform blending, as well as to morph between an image and a light field.
Lifeng Wang 0001, Stephen Lin 0001, Seungyong Lee 0001, Baining Guo, Harry Shum
IEEE Trans. Vis. Comput. Graph.5
2005 Decorating Surfaces with Bidirectional Texture Functions
abstract
We present a system for decorating arbitrary surfaces with bidirectional texture functions (BTF). Our system generates BTFs in two steps. First, we automatically synthesize a BTF over the target surface from a given BTF sample. Then, we let the user interactively paint BTF patches onto the surface such that the painted patches seamlessly integrate with the background patterns. Our system is based on a patch-based texture synthesis approach known as quilting. We present a graphcut algorithm for BTF synthesis on surfaces and the algorithm works well for a wide variety of BTF samples, including those which present problems for existing algorithms. We also describe a graphcut texture painting algorithm for creating new surface imperfections (e.g., dirt, cracks, scratches) from existing imperfections found in input BTF samples. Using these algorithms, we can decorate surfaces with real-world textures that have spatially-variant reflectance, fine-scale geometry details, and surfaces imperfections. A particularly attractive feature of BTF painting is that it allows us to capture imperfections of real materials and paint them onto geometry models. We demonstrate the effectiveness of our system with examples.
Kun Zhou 0001, Lifeng Wang 0001, Yasuyuki Matsushita, Jiaoying Shi, Baining Guo, Harry Shum
IEEE Trans. Vis. Comput. Graph.7
2005 Shell radiance texture functions
Yanyun Chen, Xin Tong 0001, Stephen Lin 0001, Jiaoying Shi, Baining Guo, Harry Shum
Vis. Comput.7
2005 Capturing and rendering geometry details for BTF-mapped surfaces
Jiaping Wang, Xin Tong 0001, John Snyder, Yanyun Chen, Baining Guo, Harry Shum
Vis. Comput.6
2004 Radiometric Calibration from a Single Image
Stephen Lin 0001, Jinwei Gu, Shuntaro Yamazaki, Harry Shum
CVPR (2)4
2004 Bayesian Correction of Image Intensity with Spatial Consideration
Jiaya Jia, Jian Sun 0001, Chi-Keung Tang, Harry Shum
ECCV (3)4
2004 Estimating Intrinsic Images from Image Sequences with Biased Illumination
Yasuyuki Matsushita, Stephen Lin 0001, Sing Bing Kang, Harry Shum
ECCV (2)4
2004 Synthesizing Dynamic Texture with Closed-Loop Linear Dynamic System
Lu Yuan 0001, Fang Wen 0001, Ce Liu 0001, Harry Shum
ECCV (2)4
2004 In Search of Textons
Harry Shum
Graphics Interface1
2004 On the rendering and post-processing of simplified dynamic light fields with depth information
abstract
This paper studies the rendering and post-processing of a dynamic image-based representation called the simplified dynamic light fields (SDLF) (or plenoptic videos) with depth information. The user viewpoints are limited to a camera line to simplify the capturing and compression processes. By associating each image pixel with its depth value, methods for improving the rendering quality and detecting occlusions are proposed. Due to the limited sampling at depth discontinuities, adaptive lowpass filtering is applied to the detected occluded regions near object boundaries in order to suppress the aliasing artifacts. Rendering results using computer-generated images show that considerable improvement in rendering quality even for dynamic scenes with large depth variations.
Zhi-Feng Gan, S. C. Chan 0001, King To Ng, Kin-Lok Chan, Harry Shum
ICASSP (3)5
2004 Automatic extraction of semantic colors in sports video
abstract
Color has been widely used in sports video analysis. Previous techniques, however require color models from prior information or user interaction, and do not address the problem of how to automatically form color models from a video in an arbitrary sports setting. In this paper, we propose an automatic technique for extracting color models of the playing surface and the team uniforms, which can be used in higher-level processes such as tracking and recognition. Unlike most previous methods, our approach is capable of handling multi-colored patterns like striped uniforms and playing fields. Multiple forms of color processing are used to analyze video frame content, which are then used iteratively to refine the color models. The results of our color modeling technique have been applied to shot classification, and experiments on videos of different sports have verified our approach.
Boyi Zeng, Stephen Lin 0001, Guangyou Xu, Harry Shum
ICASSP (3)5
2004 Face alignment using intrinsic information
abstract
Previous 2-D face alignment algorithms are generally quite sensitive to illumination variation and poor initialization. To account for these two obstacles, two forms of relatively lighting invariant descriptors - intrinsic gray-level information and intrinsic edge information - rare adopted in our algorithm to direct shape search. The former is recovered from local intensity normalization and useful at localizing face contours accurately despite its dependency on initialization. The latter is extracted from normalized local regions by Canny edge filtering and is robust at coarse alignment in spite of poor initialization. The different merits of these two forms of intrinsic information motivate us to employ them at different stages of our face alignment process. Extensive experimentations show that this proposed approach allows our system to handle not only illumination variation, but also poor initialization.
Yuchi Huang, Stephen Lin 0001, Hanqing Lu, Harry Shum
ICIP4
2004 Generic slow-motion replay detection in sports video
Stephen Lin 0001, Guangyou Xu, Harry Shum
ICIP5
2004 Perceptually Based Approach for Planar Shape Morphing
abstract
This paper presents an approach for establishing vertex correspondences between two planar shapes. Correspondences are established between the perceptual feature points extracted from both source and target shapes. A similarity metric between two feature points is defined using the intrinsic properties of their local neighborhoods. The optimal correspondence is found by an efficient dynamic programming technique. Our approach treats shape noise by allowing discarding small feature points, which introduces skips in the traversal of the dynamic programming graph. Our method is fast, feature preserving, and invariant to geometric transformations. We demonstrate the superiority of our approach over other approaches by experimental results.
Ligang Liu 0001, Guopu Wang, Bo Zhang 0025, Baining Guo, Harry Shum
PG5
2004 Iso-charts: Stretch-driven Mesh Parameterization using Spectral Analysis
Kun Zhou 0001, John M. Snyder, Baining Guo, Harry Shum
Symposium on Geometry Processing4
2004 A Geometric Analysis of Light Field Rendering
Zhouchen Lin, Harry Shum
Int. J. Comput. Vis.2
2004 Constrained planar motion analysis by decomposition
Long Quan, Le Lu 0001, Harry Shum
Image Vis. Comput.4
2004 Stereo Reconstruction from Multiperspective Panoramas
Yin Li 0003, Harry Shum, Chi-Keung Tang, Richard Szeliski
IEEE Trans. Pattern Anal. Mach. Intell.2
2004 Fundamental Limits of Reconstruction-Based Superresolution Algorithms under Local Translation
Zhouchen Lin, Harry Shum
IEEE Trans. Pattern Anal. Mach. Intell.2
2004 Error Analysis of Pure Rotation-Based Self-Calibration
abstract
Self-calibration using pure rotation is a well-known technique and has been shown to be a reliable means for recovering intrinsic camera parameters. However, in practice, it is virtually impossible to ensure that the camera motion for this type of self-calibration is a pure rotation. In this paper, we present an error analysis of recovered intrinsic camera parameters due to the presence of translation. We derived closed-form error expressions for a single pair of images with nondegenerate motion; for multiple rotations for which there are no closed-form solutions, analysis was done through repeated experiments. Among others, we show that translation-independent solutions do exist under certain practical conditions. Our analysis can be used to help choose the least error-prone approach (if multiple approaches exist) for a given set of conditions.
Sing Bing Kang, Harry Shum, Guangyou Xu
IEEE Trans. Pattern Anal. Mach. Intell.3
2004 Automatic Eyeglasses Removal from Face Images
abstract
In this paper, we present an intelligent image editing and face synthesis system that automatically removes eyeglasses from an input frontal face image. Although conventional image editing tools can be used to remove eyeglasses by pixel-level editing, filling in the deleted eyeglasses region with the right content is a difficult problem. Our approach works at the object level where the eyeglasses are automatically located, removed as one piece, and the void region filled. Our system consists of three parts: eyeglasses detection, eyeglasses localization, and eyeglasses removal. First, an eye region detector, trained offline, is used to approximately locate the region of eyes, thus the region of eyeglasses. A Markov-chain Monte Carlo method is then used to accurately locate key points on the eyeglasses frame by searching for the global optimum of the posterior. Subsequently, a novel sample-based approach is used to synthesize the face image without the eyeglasses. Specifically, we adopt a statistical analysis and synthesis approach to learn the mapping between pairs of face images with and without eyeglasses from a database. Extensive experiments demonstrate that our system effectively removes eyeglasses.
Ce Liu 0001, Harry Shum, Ying-Qing Xu, Zhengyou Zhang
IEEE Trans. Pattern Anal. Mach. Intell.3
2004 Shell texture functions
abstract
We propose a texture function for realistic modeling and efficient rendering of materials that exhibit surface mesostructures, translucency and volumetric texture variations. The appearance of such complex materials for dynamic lighting and viewing directions is expensive to calculate and requires an impractical amount of storage to precompute. To handle this problem, our method models an object as a shell layer, formed by texture synthesis of a volumetric material sample, and a homogeneous inner core. To facilitate computation of surface radiance from the shell layer, we introduce the shell texture function (STF) which describes voxel irradiance fields based on precomputed fine-level light interactions such as shadowing by surface mesostructures and scattering of photons inside the object. Together with a diffusion approximation of homogeneous inner core radiance, the STF leads to fast and detailed raytraced renderings of complex materials.
Yanyun Chen, Xin Tong 0001, Jiaping Wang, Stephen Lin 0001, Baining Guo, Harry Shum
ACM Trans. Graph.6
2004 Lazy snapping
abstract
In this paper, we present Lazy Snapping , an interactive image cutout tool. Lazy Snapping separates coarse and fine scale processing, making object specification and detailed adjustment easy . Moreover, Lazy Snapping provides instant visual feedback, snapping the cutout contour to the true object boundary efficiently despite the presence of ambiguous or low contrast edges. Instant feedback is made possible by a novel image segmentation algorithm which combines graph cut with pre-computed over-segmentation. A set of intuitive user interface (UI) tools is designed and implemented to provide flexible control and editing for the users. Usability studies indicate that Lazy Snapping provides a better user experience and produces better segmentation results than the state-of-the-art interactive image cutout tool, Magnetic Lasso in Adobe Photoshop.
Yin Li 0003, Jian Sun 0001, Chi-Keung Tang, Harry Shum
ACM Trans. Graph.4
2004 Pop-up light field: An interactive image-based modeling and rendering system
abstract
In this article, we present an image-based modeling and rendering system, which we call pop-up light field , that models a sparse light field using a set of coherent layers . In our system, the user specifies how many coherent layers should be modeled or popped up according to the scene complexity. A coherent layer is defined as a collection of corresponding planar regions in the light field images. A coherent layer can be rendered free of aliasing all by itself, or against other background layers. To construct coherent layers, we introduce a Bayesian approach, coherence matting , to estimate alpha matting around segmented layer boundaries by incorporating a coherence prior in order to maintain coherence across images.We have developed an intuitive and easy-to-use user interface (UI) to facilitate pop-up light field construction. The key to our UI is the concept of human-in-the-loop where the user specifies where aliasing occurs in the rendered image. The user input is reflected in the input light field images where pop-up layers can be modified. The user feedback is instant through a hardware-accelerated real-time pop-up light field renderer. Experimental results demonstrate that our system is capable of rendering anti-aliased novel views from a sparse light field.
Harry Shum, Jian Sun 0001, Shuntaro Yamazaki, Yin Li 0003, Chi-Keung Tang
ACM Trans. Graph.1
2004 Poisson matting
abstract
In this paper, we formulate the problem of natural image matting as one of solving Poisson equations with the matte gradient field. Our approach, which we call Poisson matting , has the following advantages. First, the matte is directly reconstructed from a continuous matte gradient field by solving Poisson equations using boundary information from a user-supplied trimap. Second, by interactively manipulating the matte gradient field using a number of filtering tools, the user can further improve Poisson matting results locally until he or she is satisfied. The modified local result is seamlessly integrated into the final result. Experiments on many complex natural images demonstrate that Poisson matting can generate good matting results that are not possible using existing matting techniques.
Jian Sun 0001, Jiaya Jia, Chi-Keung Tang, Harry Shum
ACM Trans. Graph.4
2004 Video tooning
abstract
We describe a system for transforming an input video into a highly abstracted, spatio-temporally coherent cartoon animation with a range of styles. To achieve this, we treat video as a space-time volume of image data. We have developed an anisotropic kernel mean shift technique to segment the video data into contiguous volumes. These provide a simple cartoon style in themselves, but more importantly provide the capability to semi-automatically rotoscope semantically meaningful regions.In our system, the user simply outlines objects on keyframes. A mean shift guided interpolation algorithm is then employed to create three dimensional semantic regions by interpolation between the keyframes, while maintaining smooth trajectories along the time dimension. These regions provide the basis for creating smooth two dimensional edge sheets and stroke sheets embedded within the spatio-temporal video volume. The regions, edge sheets, and stroke sheets are rendered by slicing them at particular times. A variety of styles of rendering are shown. The temporal coherence provided by the smoothed semantic regions and sheets results in a temporally consistent non-photorealistic appearance.
Jue Wang 0001, Ying-Qing Xu, Harry Shum, Michael F. Cohen
ACM Trans. Graph.3
2004 Mesh editing with poisson-based gradient field manipulation
abstract
In this paper, we introduce a novel approach to mesh editing with the Poisson equation as the theoretical foundation. The most distinctive feature of this approach is that it modifies the original mesh geometry implicitly through gradient field manipulation. Our approach can produce desirable and pleasing results for both global and local editing operations, such as deformation, object merging, and smoothing. With the help from a few novel interactive tools, these operations can be performed conveniently with a small amount of user interaction. Our technique has three key components, a basic mesh solver based on the Poisson equation, a gradient field manipulation scheme using local transforms, and a generalized boundary condition representation based on local frames. Experimental results indicate that our framework can outperform previous related mesh editing techniques.
Yizhou Yu, Kun Zhou 0001, Dong Xu 0001, Hujun Bao, Baining Guo, Harry Shum
ACM Trans. Graph.7
2004 Synthesis and Rendering of Bidirectional Texture Functions on Arbitrary Surfaces
abstract
The bidirectional texture function (BTF) is a 6D function that describes the appearance of a real-world surface as a function of lighting and viewing directions. The BTF can model the fine-scale shadows, occlusions, and specularities caused by surface mesostructures. In this paper, we present algorithms for efficient synthesis of BTFs on arbitrary surfaces and for hardware-accelerated rendering. For both synthesis and rendering, a main challenge is handling the large amount of data in a BTF sample. To addresses this challenge, we approximate the BTF sample by a small number of 4D point appearance functions (PAFs) multiplied by 2D geometry maps. The geometry maps and PAFs lead to efficient synthesis and fast rendering of BTFs on arbitrary surfaces. For synthesis, a surface BTF can be generated by applying a texton-based sysnthesis algorithm to a small set of 2D geometry maps while leaving the companion 4D PAFs untouched. As for rendering, a surface BTF synthesized using geometry maps is well-suited for leveraging the programmable vertex and pixel shaders on the graphics hardware. We present a real-time BTF rendering algorithm that runs at the speed of about 30 frames/second on a mid-level PC with an ATI Radeon 8500 graphics card. We demonstrate the effectiveness of our synthesis and rendering algorithms using both real and synthetic BTF samples.
Xinguo Liu, Jingdan Zhang, Xin Tong 0001, Baining Guo, Harry Shum
IEEE Trans. Vis. Comput. Graph.6
2003 Face Alignment Using Statistical Models and Wavelet Features
abstract
Active shape model (ASM) is a powerful statistical tool for face alignment by shape. However, it can suffer from changes in illumination and facial expression changes, and local minima in optimization. In this paper, we present a method, W-ASM, in which Gabor wavelet features are used for modeling local image structure. The magnitude and phase of Gabor features contain rich information about the local structural features of face images to be aligned, and provide accurate guidance for search. To a large extent, this repairs defects in gray scale based search. An E-M algorithm is used to model the Gabor feature distribution, and a coarse-to-fine grained search is used to position local features in the image. Experimental results demonstrate the ability of W-ASM to accurately align and locate facial features.
Feng Jiao, Stan Z. Li, Harry Shum, Dale Schuurmans
CVPR (1)3
2003 An Efficient Approach to Learning Inhomogeneous Gibbs Model
abstract
The inhomogeneous Gibbs model (IGM) (Liu et al., 2001) is an effective maximum entropy model in characterizing complex high-dimensional distributions. However, its training process is so slow that the applicability of IGM has been greatly restricted. In this paper, we propose an approach for fast parameter learning of IGM. In IGM learning, features are incrementally constructed to constrain the learnt distribution. When a new feature is added, Markov-chain Monte Carlo (MCMC) sampling is repeated to draw samples for parameter learning. In contrast, our approach constructs a closed-form reference distribution using approximate information gain criteria. Because our reference distribution is very close to the optimal one, importance sampling can be used to accelerate the parameter optimization process. For problems with high-dimensional distributions, our approach typically achieves a speedup of two orders of magnitude compared to the original IGM. We further demonstrate the efficiency of our approach by learning a high-dimensional joint distribution of face images and their corresponding caricatures.
Harry Shum
CVPR (1)3
2003 Kullback-Leibler Boosting
abstract
In this paper, we develop a general classification framework called Kullback-Leibler Boosting, or KLBoosting. KLBoosting has following properties. First, classification is based on the sum of histogram divergences along corresponding global and discriminating linear features. Second, these linear features, called KL features, are iteratively learnt by maximizing the projected Kullback-Leibler divergence in a boosting manner. Third, the coefficients to combine the histogram divergences are learnt by minimizing the recognition error once a new feature is added to the classifier. This contrasts conventional AdaBoost where the coefficients are empirically set. Because of these properties, KLBoosting classifier generalizes very well. Moreover, to apply KLBoosting to high-dimensional image space, we propose a data-driven Kullback-Leibler Analysis (KLA) approach to find KL features for image objects (e.g., face patches). Promising experimental results on face detection demonstrate the effectiveness of KLBoosting.
Ce Liu 0001, Harry Shum
CVPR (1)2
2003 Directional Histogram Model for Three-Dimensional Shape Similarity
abstract
In this paper, we propose a novel shape representation we call directional histogram model (DHM). It captures the shape variation of an object and is invariant to scaling and rigid transforms. The DHM is computed by first extracting a directional distribution of thickness histogram signatures, which are translation invariant. We show how the extraction of the thickness histogram distribution can be accelerated using conventional graphics hardware. Orientation invariance is achieved by computing the spherical harmonic transform of this distribution. Extensive experiments show that the DHM is capable of high discrimination power and is robust to noise.
Xinguo Liu, Robin Sun, Sing Bing Kang, Harry Shum
CVPR (1)4
2003 Image Hallucination with Primal Sketch Priors
abstract
We propose a Bayesian approach to image hallucination. Given a generic low resolution image, we hallucinate a high resolution image using a set of training images. Our work is inspired by recent progress on natural image statistics that the priors of image primitives can be well represented by examples. Specifically, primal sketch priors (e.g., edges, ridges and corners) are constructed and used to enhance the quality of the hallucinated high resolution image. Moreover, a contour smoothness constraint enforces consistency of primitives in the hallucinated image by a Markov-chain based inference algorithm. A reconstruction constraint is also applied to further improve the quality of the hallucinated image. Experiments demonstrate that our approach can hallucinate high quality super-resolution images.
Jian Sun 0001, Nanning Zheng 0001, Harry Shum
CVPR (2)4
2003 The compression of simplified dynamic light fields
abstract
This paper studies the compression of a dynamic image-based rendering (IBR) representation called simplified dynamic light fields (SDLF). It is obtained by constraining the viewpoints in a dynamic environment along a line instead of a 2D plane. The SDLFs have a dimensionality of four, which considerably simplifies their capture and data compression. A new coding algorithm for SDLFs using a modified MPEG-2 algorithm is proposed. It employs both temporal and spatial predictions from the reference video streams to explore better the redundancy among the light field images. Experimental results, using a synthetic SDLF, show that the proposed compression scheme offers a 2 dB improvement in PSNR over a similar coding scheme using only temporal prediction.
S. C. Chan 0001, King To Ng, Zhi-Feng Gan, Kin-Lok Chan, Harry Shum
ICASSP (3)5
2003 Rendering driven depth reconstruction
abstract
Previous work on image-based rendering suggests that there is a tradeoff between the number of images and the amount of geometry required for anti-aliased rendering. For instance, plenoptic sampling theory indicates that visually acceptable rendering can be achieved when the input images are undersampled, if sufficient depth information is available for all the pixels. In this paper, we propose a novel vision reconstruction approach, rendering-driven depth recovery, to recover the amount of geometry that is necessary for anti-aliased rendering. Our approach contrasts conventional stereo reconstruction in that we do not intend to accurately reconstruct the depth for each and every single pixel, leading to a very efficient reconstruction algorithm. Our algorithm uses a block-based multi-layer depth representation, and searches in the depth space based on the causality criterion, by detecting double images. Experiments show that rendering systems using our rendering driven depth recovery algorithm can synthesize satisfactory novel views efficiently by using 'just enough geometry' recovered from undersampled input images.
Yin Li 0003, Xin Tong 0001, Chi-Keung Tang, Harry Shum
ICASSP (4)4
2003 Accelerating video decoding using GPU
abstract
Most modern computers or game consoles are equipped with powerful graphics processing units (GPU) to accelerate graphics operations. There is a trend that the power of GPU outgrows that of the CPU (central processing unit). However, the GPU engines are specially designed for graphics operations. Can we take advantage of the powerful GPU engines for more general operations other than pure graphics operations? The answer is positive. In this study, we present schemes that map other non-graphics operations into graphics engines with an example application of accelerating video decoding with the assistance of GPU. Our results show that significant speed-up can be achieved by leveraging the GPU power. Specifically, we have achieved real-time playback of high definition video on a PC with an Intel Pentium III 667 MHz CPU and an nVidia GeForce3 GPU.
Guobin Shen, Lihua Zhu, Shipeng Li 0001, Harry Shum, Ya-Qin Zhang
ICASSP (4)4
2003 Multiple-cue Illumination Estimation in Textured Scenes
abstract
In this paper, we present a method that integrates cues from shading, shadow and specular reflections for estimating directional illumination in a textured scene. Texture poses a problem for lighting estimation, since texture edges can be mistaken for changes in illumination condition, and unknown variations in albedo make reflectance model fitting impractical. Unlike previous works which all assume known or uniform reflectance, our method can deal with the effects of textures by capitalizing on physical consistencies that exist among the lighting cues. Since scene textures do not exhibit such coherence, we use this property to minimize the influence of texture on illumination direction estimation. For the recovered light source directions, a technique for estimating their intensities in the presence of texture is also proposed.
Yuanzhen Li, Stephen Lin 0001, Hanqing Lu, Harry Shum
ICCV4
2003 Highlight Removal by Illumination-Constrained Inpainting
abstract
We present a single-image highlight removal method that incorporates illumination-based constraints into image inpainting. Unlike occluded image regions filled by traditional inpainting, highlight pixels contain some useful information for guiding the inpainting process. Constraints provided by observed pixel colors, highlight color analysis and illumination color uniformity are employed in our method to improve estimation of the underlying diffuse color. The inclusion of these illumination constraints allows for better recovery of shading and textures by inpainting. Experimental results are given to demonstrate the performance of our method.
Ping Tan 0002, Stephen Lin 0001, Long Quan, Harry Shum
ICCV4
2003 Dynamic depth recovery using belief propagation
abstract
In this paper, we study the dynamic stereo problem, i.e. to recover the shape of dynamic scene from multiple synchronized image sequences. To incorporate both spatial and temporal information for depth recovery, we propose a statistical framework that uses pixel process model to encode temporal coherence, and Markov random fields (MRFs) for spatial coherence. In this framework, the dynamic depth recovery problem is finally formulated as an optimization problem, and is optimized by using the belief propagation algorithm. Experimental results with the real dynamic scenes illustrate our method's ability of robust shape recovery.
Yihua Xu, Harry Shum, Songde Ma
ICIP (1)3
2003 Learning to boost GMM based speaker verification
abstract
The Gaussian mixture models (GMM) has proved to be an effective probabilistic model for speaker verification, and has been widely used in most of state-of-the-art systems. In this paper, we introduce a new method for the task: that using AdaBoost learning based on the GMM. The motivation is the following: While a GMM linearly combines a number of Gaussian models according to a set of mixing weights, we believe that there exists a better means of combining individual Gaussian mixture models. The proposed AdaBoost-GMM method is non-parametric in which a selected set of weak classifiers, each constructed based on a single Gaussian model, is optimally combined to form a strong classifier, the optimality being in the sense of maximum margin. Experiments show that the boosted GMM classifier yields 10.81% relative reduction in equal error rate for the same handsets and 11.24% for different handsets, a significant improvement over the baseline adapted GMM system.
Stan Z. Li, Dong Zhang 0001, Chengyuan Ma, Harry Shum, Eric Chang
INTERSPEECH4
2003 In Search of Texton
abstract
Texture analysis and synthesis have been studied extensively by many vision and graphics researchers. A very useful concept in texture analysis is texton or the basic element of a texture image. However, it remains unclear how to define a texton despite the fact that Julesz proposed the term "texton" more than twenty years ago. The author proposes a two-level statistical model with textons and their distribution for texture analysis and synthesis, with the emphasis on how to define textons computationally. Specifically, the author presents recent work on searching for 2D textons in patch-based texture synthesis (ACM ToG, July 2001), 3D textons in bi-directional texture function (BTF) synthesis (Siggraph'2002) and 1D motion textons in motion texture synthesis (Siggraph'2002).
Harry Shum
Shape Modeling International1
2003 Special issue on Pacific Graphics 2002
Shi-Min Hu 0001, Sabine Coquillart, Harry Shum
Graph. Model.3
2003 Learning kernel-based HMMs for dynamic sequence synthesis
Nanning Zheng 0001, Ying-Qing Xu, Harry Shum
Graph. Model.5
2003 Face alignment using texture-constrained active shape models
Shuicheng Yan, Ce Liu 0001, Stan Z. Li, HongJiang Zhang, Harry Shum
Image Vis. Comput.5
2003 Stereo Matching Using Belief Propagation
abstract
In this paper, we formulate the stereo matching problem as a Markov network and solve it using Bayesian belief propagation. The stereo Markov network consists of three coupled Markov random fields that model the following: a smooth field for depth/disparity, a line process for depth discontinuity, and a binary process for occlusion. After eliminating the line process and the binary process by introducing two robust functions, we apply the belief propagation algorithm to obtain the maximum a posteriori (MAP) estimation in the Markov network. Other low-level visual cues (e.g., image segmentation) can also be easily incorporated in our stereo model to obtain better stereo results. Experiments demonstrate that our methods are comparable to the state-of-the-art stereo algorithms for many test cases.
Jian Sun 0001, Nanning Zheng 0001, Harry Shum
IEEE Trans. Pattern Anal. Mach. Intell.3
2003 Large environment rendering using plenoptic primitives
abstract
One of the most difficult tasks in computer graphics is to enable virtual walkthroughs in very large and complicated environments that are photorealistic, seamless, and in real time. Current image-based rendering techniques, while capable of photorealism and interactive speeds, have failed in practice to extend to visualizations of such environments. We demonstrate an approach that defines a virtual walkthrough experience using plenoptic primitives (PPs). A PP can be any type of local visual experience: 360/spl deg/ static panorama, panoramic video (PV), lumigraph/light field representation, or concentric mosaics (CMs). By combining them judiciously, user experience can be authored with significantly reduced effort while maintaining high-quality user experience. We illustrate our technique on synthetic and real environments using PVs and CMs and show how the problem of achieving smooth transitions among PVs and CMs can be solved by using position-dependent local geometries.
Sing Bing Kang, Minsheng Wu, Yin Li 0003, Harry Shum
IEEE Trans. Circuits Syst. Video Technol.4
2003 Survey of image-based representations and compression techniques
abstract
We survey the techniques for image-based rendering (IBR) and for compressing image-based representations. Unlike traditional three-dimensional (3-D) computer graphics, in which 3-D geometry of the scene is known, IBR techniques render novel views directly from input images. IBR techniques can be classified into three categories according to how much geometric information is used: rendering without geometry, rendering with implicit geometry (i.e., correspondence), and rendering with explicit geometry (either with approximate or accurate geometry). We discuss the characteristics of these categories and their representative techniques. IBR techniques demonstrate a surprising diverse range in their extent of use of images and geometry in representing 3-D scenes. We explore the issues in trading off the use of images and geometry by revisiting plenoptic-sampling analysis and the notions of view dependency and geometric proxies. Finally, we highlight compression techniques specifically designed for image-based representations. Such compression techniques are important in making IBR techniques practical.
Harry Shum, Sing Bing Kang, S. C. Chan 0001
IEEE Trans. Circuits Syst. Video Technol.1
2003 Introduction to the special issue on image-based modeling, rendering, and animation
Harry Shum, Eric Petajan, Jörn Ostermann
IEEE Trans. Circuits Syst. Video Technol.1
2003 Bi-scale radiance transfer
abstract
Radiance transfer represents how generic source lighting is shadowed and scattered by an object to produce view-dependent appearance. We generalize by rendering transfer at two scales. A macro-scale is coarsely sampled over an object's surface, providing global effects like shadows cast from an arm onto a body. A meso-scale is finely sampled over a small patch to provide local texture. Low-order (25D) spherical harmonics represent low-frequency lighting dependence for both scales. To render, a coefficient vector representing distant source lighting is first transformed at the macro-scale by a matrix at each vertex of a coarse mesh. The resulting vectors represent a spatially-varying hemisphere of lighting incident to the meso-scale. A 4D function, called a radiance transfer texture (RTT), then specifies the surface's meso-scale response to each lighting basis component, as a function of a spatial index and a view direction. Finally, a 25D dot product of the macro-scale result vector with the vector looked up from the RTT performs the correct shading integral. We use an id map to place RTT samples from a small patch over the entire object; only two scalars are specified at high spatial resolution. Results show that bi-scale decomposition makes preprocessing practical and efficiently renders self-shadowing and interreflection effects from dynamic, low-frequency light sources at both scales.
Peter-Pike J. Sloan, Xinguo Liu, Harry Shum, John M. Snyder
ACM Trans. Graph.3
2003 View-dependent displacement mapping
abstract
Significant visual effects arise from surface mesostructure, such as fine-scale shadowing, occlusion and silhouettes. To efficiently render its detailed appearance, we introduce a technique called view-dependent displacement mapping (VDM) that models surface displacements along the viewing direction. Unlike traditional displacement mapping, VDM allows for efficient rendering of self-shadows, occlusions and silhouettes without increasing the complexity of the underlying surface mesh. VDM is based on per-pixel processing, and with hardware acceleration it can render mesostructure with rich visual appearance in real time.
Lifeng Wang 0001, Xin Tong 0001, Stephen Lin 0001, Shi-Min Hu 0001, Baining Guo, Harry Shum
ACM Trans. Graph.7
2003 Synthesis of progressively-variant textures on arbitrary surfaces
abstract
We present an approach for decorating surfaces with progressively-variant textures . Unlike a homogeneous texture, a progressively-variant texture can model local texture variations, including the scale, orientation, color, and shape variations of texture elements. We describe techniques for modeling progressively-variant textures in 2D as well as for synthesizing them over surfaces. For 2D texture modeling, our feature-based warping technique allows the user to control the shape variations of texture elements, making it possible to capture complex texture variations such as those seen in animal coat patterns. In addition, our feature-based blending technique can create a smooth transition between two given homogeneous textures, with progressive changes of both shapes and colors of texture elements. For synthesizing textures over surfaces, the biggest challenge is that the synthesized texture elements tend to break apart as they progressively vary. To address this issue, we propose an algorithm based on texton masks, which mark most prominent texture elements in the 2D texture sample. By leveraging the power of texton masks, our algorithm can maintain the integrity of the synthesized texture elements on the target surface.
Jingdan Zhang, Kun Zhou 0001, Luiz Velho 0001, Baining Guo, Harry Shum
ACM Trans. Graph.5
2003 Realistic Rendering and Animation of Knitwear
abstract
We present a framework for knitwear modeling and rendering that accounts for characteristics that are particular to knitted fabrics. We first describe a model for animation that considers knitwear features and their effects on knitwear shape and interaction. With the computed free-form knitwear configurations, we present an efficient procedure for realistic synthesis based on the observation that a single cross section of yarn can serve as the basic primitive for modeling entire articles of knitwear. This primitive, called the lumislice, describes radiance from a yarn cross section that accounts for fine-level interactions among yarn fibers. By representing yarn as a sequence of identical but rotated cross sections, the lumislice can effectively propagate local microstructure over arbitrary stitch patterns and knitwear shapes. The lumislice accommodates varying levels of detail, allows for soft shadow generation, and capitalizes on hardware-assisted transparency blending. These modeling and rendering techniques together form a complete approach for generating realistic knitwear.
Yanyun Chen, Stephen Lin 0001, Ying-Qing Xu, Baining Guo, Harry Shum
IEEE Trans. Vis. Comput. Graph.6
2002 Statistical Learning of Multi-view Face Detection
Stan Z. Li, Long Zhu, ZhenQiu Zhang, Andrew Blake 0001, HongJiang Zhang, Harry Shum
ECCV (4)6
2002 Diffuse-Specular Separation and Depth Recovery from Image Sequences
Stephen Lin 0001, Yuanzhen Li, Sing Bing Kang, Xin Tong 0001, Harry Shum
ECCV (3)5
2002 Hierarchical Shape Modeling for Automatic Face Localization
Ce Liu 0001, Harry Shum, Changshui Zhang
ECCV (2)2
2002 Stereo Matching Using Belief Propagation
Jian Sun 0001, Harry Shum, Nanning Zheng 0001
ECCV (2)2
2002 The application of nonlinear filter banks to efficient rendering and progressive transmission of light fields
abstract
This paper studies the application of perfect reconstruction nonlinear filter banks (NFB) to the efficient rendering and progressive transmission of light fields. The reference pictures in the conventional disparity-compensated prediction encoder are decomposed using the NFB to reduce the amount of main memory needed to support fast rendering. The NFB has very low arithmetic complexity for reconstruction and small filter support which considerably simplifies the random access operations. It can also be applied to the predicted light field images to support progressive transmission. Different prediction and reconstruction strategies are also investigated to achieve different tradeoffs between memory requirement and decoding speed.
King To Ng, S. C. Chan 0001, Harry Shum
ICIP (2)3
2002 PicToon: a personalized image-based cartoon system
abstract
In this paper, we present PicToon, a cartoon system which can generate a personalized cartoon face from an input Picture. PicToon is easy to use and requires little user interaction. Our system consists of three major components: an image-based Cartoon Generator, an interactive Cartoon Editor for exaggeration, and a speech-driven Cartoon Animator. First, to capture an artistic style, the cartoon generation is decoupled into two processes: sketch generation and stroke rendering. An example-based approach is taken to automatically generate sketch lines which depict the facial structure. An inhomogeneous non-parametric sampling plus a flexible facial template is employed to extract the vector-based facial sketch. Various styles of strokes can then be applied. Second, with the pre-designed templates in Cartoon Editor, the user can easily make the cartoon exaggerated or more expressive. Third, a real-time lip-syncing algorithm is also developed that recovers a statistical audio-visual mapping between the character's voice and the corresponding lip configuration. Experimental results demonstrate the effectiveness of our system.
Nanning Zheng 0001, Ying-Qing Xu, Harry Shum
ACM Multimedia6
2002 FloatBoost Learning for Classification
abstract
AdaBoost [3] minimizes an upper error bound which is an exponential function of the margin on the training set [14]. However, the ultimate goal in applications of pattern classification is always minimum error rate. On the other hand, AdaBoost needs an effective procedure for learning weak classifiers, which by itself is difficult especially for high dimensional data. In this paper, we present a novel procedure, called FloatBoost, for learning a better boosted classifier. FloatBoost uses a backtrack mechanism after each iteration of AdaBoost to remove weak classifiers which cause higher error rates. The resulting float-boosted classifier consists of fewer weak classifiers yet achieves lower error rates than AdaBoost in both training and test. We also propose a statistical model for learning weak classifiers, based on a stagewise approximation of the posterior using an overcomplete set of scalar features. Experi- mental comparisons of FloatBoost and AdaBoost are provided through a difficult classification problem, face detection, where the goal is to learn from training examples a highly nonlinear classifier to differentiate be- tween face and nonface patterns in a high dimensional space. The results clearly demonstrate the promises made by FloatBoost over AdaBoost.
Stan Z. Li, ZhenQiu Zhang, Harry Shum, HongJiang Zhang
NIPS3
2002 Single-Image Reflectance Estimation for Relighting by Iterative Soft Grouping
abstract
Reflectance values for image-based relighting are often estimated from grouped pixels with similar reflectance, but such groupings are difficult to compute with certainty for sparse image data. To address this problem, we propose an iterative method that aggregates BRDF data in a single image with known geometry and lighting by soft grouping, where pixels contribute to one another's estimate according to their degree of reflectance similarity. Estimation of specular reflectance is further improved by albedo-independent soft grouping of pixels based on shape continuity. With recovered reflectances, we demonstrate realistic relighting for synthetic and real scenes, including surfaces with spatially-varying reflectance.
Yuanzhen Li, Stephen Lin 0001, Sing Bing Kang, Hanqing Lu, Harry Shum
PG5
2002 Example-Based Caricature Generation with Exaggeration
abstract
In this paper, we present a system that automatically generates caricatures from input face images. From example caricatures drawn by an artist, our caricature system learns how an artist draws caricatures. In our approach, we decouple the process of caricature generation into two parts, i.e., shape exaggeration and texture style transferring. The exaggeration of a caricature is accomplished by a prototype-based method that captures the artist's understanding of what are distinctive features of a face and the exaggeration style. Such prototypes are learnt by analyzing the correlation between the image caricature pairs using partial least-squares (PLS). Experimental results demonstrate the effectiveness of our system.
Ying-Qing Xu, Harry Shum
PG4
2002 Pattern-Based Texture Metamorphosis
abstract
In this paper we study texture metamorphosis, or how to generate texture samples that smoothly transform from a source texture image to a target. We propose a pattern-based approach to specify the feature correspondence between two textures, based on the observation that man), texture images have stochastically distributed patterns which are similar to each other First, the user selects a pattern in the source and target textures, and establishes the "local feature correspondence" between these two patterns by specifying landmarks. Then, repeated patterns are automatically detected and localized in the source and target textures. The "pattern correspondence" between two textures is formulated as an integer programming problem and solved using the Hungarian algorithm. Finally, we obtain a warp function between two textures by combining "local feature correspondence" and "pattern correspondence". Experiments demonstrate that our technique produces visually appealing morphing sequences, with moderate amount of user interaction.
Ce Liu 0001, Harry Shum, Yizhou Yu
PG3
2002 Lighting Interpolation by Shadow Morphing Using Intrinsic Lumigraphs
abstract
Densely-sampled image representations such as the light field or lumigraph have been effective in enabling photorealistic image synthesis. Unfortunately, lighting interpolation with such representations has not been shown to be possible without the use of accurate 3D geometry and surface reflectance properties. In this paper we propose an approach to image-based lighting interpolation that is based on estimates of geometry and shading from relatively few images. We decompose captured light fields at different lighting conditions into intrinsic images (reflectance and illumination images), and estimate view-dependent scene geometries using multi-view stereo. We call the resulting representation an intrinsic lumigraph. In the same way that the lumigraph uses geometry to permit more accurate view interpolation, the intrinsic lumigraph uses both geometry and intrinsic images to allow high-quality interpolation at different views and lighting conditions. Joint use of geometry and intrinsic images is effective in the computation of shadow masks for shadow prediction at new lighting conditions. We illustrate our approach with images of real scenes.
Yasuyuki Matsushita, Sing Bing Kang, Stephen Lin 0001, Harry Shum, Xin Tong 0001
PG4
2002 Learning Kernel-Based HMMs for Dynamic Sequence Synthesi
abstract
In this paper we present an approach that synthesizes a dynamic sequence from another related sequence, and apply it to a virtual conductor: to synthesize linked figure animation from an input music track. We propose that the mapping between two dynamic sequences can be modeled with a Kernel-based Hidden Markov model, or KHMM. A KHMM is an HMM for which the kernel-based functions are used to model the state observation density of the joint input and output distribution. Specifically, the state observation density is estimated by employing a likelihood-weighted sampling scheme. Our KHMM model is ideal for dynamic sequence synthesis because the global dynamics are learned by the HMM, and subtle details in the dynamic mapping are kept in the kernel-based state density. We demonstrate our virtual conductor by synthesizing extensive animation sequences from input music sequences with different styles and beat patterns.
Nanning Zheng 0001, Ying-Qing Xu, Harry Shum
PG5
2002 A Novel Volume Constrained Smoothing Method for Meshes
Xinguo Liu, Hujun Bao, Harry Shum, Qunsheng Peng 0001
Graph. Model.3
2002 Relighting with the Reflected Irradiance Field: Representation, Sampling and Reconstruction
Zhouchen Lin, Tien-Tsin Wong, Harry Shum
Int. J. Comput. Vis.3
2002 Omnivergent Stereo
Steven M. Seitz, Adam Tauman Kalai, Harry Shum
Int. J. Comput. Vis.3
2002 Correction to Construction of Panoramic Image Mosaics with Global and Local Alignment
Harry Shum, Richard Szeliski
Int. J. Comput. Vis.1
2002 Rendering by Manifold Hopping
Harry Shum, Lifeng Wang 0001, Jinxiang Chai, Xin Tong 0001
Int. J. Comput. Vis.1
2002 Layered lumigraph with LOD control
abstract
Abstract The rendering performance of an image‐based rendering (IBR) system is determined by the number of images and the amount of geometrical information used. In this paper, we propose a layered lumigraph representation that is configured for optimized rendering performance based on the rendering platform (e.g., processor speed, memory) and output image resolution. The layered lumigraph is produced by classifying all pixels into a number of depth layers. Based on prior work on plenoptic sampling analysis, the layered lumigraph is constructed to achieve the same rendering quality along the minimum sampling curve by balancing the number of images and depth layers. For a given rendering platform, the best rendering performance can be obtained by choosing the optimal number of images and depth layers. Moreover, the layered lumigraph is capable of level‐of‐detail (LOD) control using the same image geometry trade‐off. Therefore, the layered lumigraph fully exploits the inherent constraints between the number of images, depth complexity, and output resolution. Finally, a backward warping technique is designed to efficiently render the layered lumigraph by taking advantage of texture mapping hardware. Copyright © 2002 John Wiley & Sons, Ltd.
Xin Tong 0001, Jinxiang Chai, Harry Shum
Comput. Animat. Virtual Worlds3
2002 Interactive multiresolution hair modeling and editing
abstract
Human hair modeling is a difficult task. This paper presents a constructive hair modeling system with which users can sculpt a wide variety of hairstyles. Our Multiresolution Hair Modeling (MHM) system is based on the observed tendency of adjacent hair strands to form clusters at multiple scales due to static attraction. In our system, initial hair designs are quickly created with a small set of hair clusters. Refinements at finer levels are achieved by subdividing these initial hair clusters. Users can edit an evolving model at any level of detail, down to a single hair strand. High level editing tools support curling, scaling, and copy/paste, enabling users to rapidly create widely varying hairstyles. Editing ease and model realism are enhanced by efficient hair rendering, shading, antialiasing, and shadowing algorithms.
Yanyun Chen, Ying-Qing Xu, Baining Guo, Harry Shum
ACM Trans. Graph.4
2002 Modeling and rendering of realistic feathers
abstract
We present techniques for realistic modeling and rendering of feathers and birds. Our approach is motivated by the observation that a feather is a branching structure that can be described by an L-system. The parametric L-system we derived allows the user to easily create feathers of different types and shapes by changing a few parameters. The randomness in feather geometry is also incorporated into this L-system. To render a feather realistically, we have derived an efficient form of the bidirectional texture function (BTF), which describes the small but visible geometry details on the feather blade. A rendering algorithm combining the L-system and the BTF displays feathers photorealistically while capitalizing on graphics hardware for efficiency. Based on this framework of feather modeling and rendering, we developed a system that can automatically generate appropriate feathers to cover different parts of a bird's body from a few "key feathers" supplied by the user, and produce realistic renderings of the bird.
Yanyun Chen, Ying-Qing Xu, Baining Guo, Harry Shum
ACM Trans. Graph.4
2002 Motion texture: a two-level statistical model for character motion synthesis
abstract
In this paper, we describe a novel technique, called motion texture, for synthesizing complex human-figure motion (e.g., dancing) that is statistically similar to the original motion captured data. We define motion texture as a set of motion textons and their distribution, which characterize the stochastic and dynamic nature of the captured motion. Specifically, a motion texton is modeled by a linear dynamic system (LDS) while the texton distribution is represented by a transition matrix indicating how likely each texton is switched to another. We have designed a maximum likelihood algorithm to learn the motion textons and their relationship from the captured dance motion. The learnt motion texture can then be used to generate new animations automatically and/or edit animation sequences interactively. Most interestingly, motion texture can be manipulated at different levels, either by changing the fine details of a specific motion at the texton level or by designing a new choreography at the distribution level. Our approach is demonstrated by many synthesized sequences of visually compelling dance motion.
Harry Shum
ACM Trans. Graph.3
2002 Synthesis of bidirectional texture functions on arbitrary surfaces
abstract
The bidirectional texture function (BTF) is a 6D function that can describe textures arising from both spatially-variant surface reflectance and surface mesostructures. In this paper, we present an algorithm for synthesizing the BTF on an arbitrary surface from a sample BTF. A main challenge in surface BTF synthesis is the requirement of a consistent mesostructure on the surface, and to achieve that we must handle the large amount of data in a BTF sample. Our algorithm performs BTF synthesis based on surface textons, which extract essential information from the sample BTF to facilitate the synthesis. We also describe a general search strategy, called the k-coherent search, for fast BTF synthesis using surface textons. A BTF synthesized using our algorithm not only looks similar to the BTF sample in all viewing/lighthing conditions but also exhibits a consistent mesostructure when viewing and lighting directions change. Moreover, the synthesized BTF fits the target surface naturally and seamlessly. We demonstrate the effectiveness of our algorithm with sample BTFs from various sources, including those measured from real-world textures.
Xin Tong 0001, Jingdan Zhang, Ligang Liu 0001, Baining Guo, Harry Shum
ACM Trans. Graph.6
2002 Feature-based light field morphing
abstract
We present a feature-based technique for morphing 3D objects represented by light fields. Our technique enables morphing of image-based objects whose geometry and surface properties are too difficult to model with traditional vision and graphics techniques. Light field morphing is not based on 3D reconstruction; instead it relies on ray correspondence, i.e., the correspondence between rays of the source and target light fields. We address two main issues in light field morphing: feature specification and visibility changes. For feature specification, we develop an intuitive and easy-to-use user interface (UI). The key to this UI is feature polygons, which are intuitively specified as 3D polygons and are used as a control mechanism for ray correspondence in the abstract 4D ray space. For handling visibility changes due to object shape changes, we introduce ray-space warping. Ray-space warping can fill arbitrarily large holes caused by object shape changes; these holes are usually too large to be properly handled by traditional image warping. Our method can deal with non-Lambertian surfaces, including specular surfaces (with dense light fields). We demonstrate that light field morphing is an effective and easy-to-use technqiue that can generate convincing 3D morphing effects.
Zhunping Zhang, Lifeng Wang 0001, Baining Guo, Harry Shum
ACM Trans. Graph.4
2001 Physically-Based Real-Time Animation of Draped Cloth
abstract
We propose a new physically-based model for real-time animation of draped cloth which not only speeds up the rendering greatly, but also maintains visually appealing results. Our simplified model is represented as a grid object composed of mass points connected by semi-rigid rods, whose behavior is governed by non-rigid dynamics. The longitudinal (vertical) and latitudinal (horizontal) directions of our model are decoupled and processed separately, and later combined to generate the final cloth. Moreover, we provide a uniform treatment of internal and external forces.
Chiyi Cheng, Jiaoying Shi, Ying-Qing Xu, Harry Shum
Computer Graphics International4
2001 Relief Mosaics by Joint View Triangulation
abstract
Relief mosaics are collections of registered images that extend traditional mosaics by supporting motion parallax. A simple parallax interpolation algorithm based on computed correspondence information allows high quality blur-free and ghost-free mosaics to be created using images from moving hand-held cameras that would not be suitable for traditional mosaicing. The renderer can also display local parallax changes, giving a local but visually convincing illusion of depth. Moreover, relief mosaics can be used for approximate plenoptic modeling from hand-held cameras at lower spatial sampling rates than existing light-field methods. We present a fully automatic correspondence based construction system for relief mosaics, and show how they can be used in applications.
Maxime Lhuillier, Long Quan, Harry Shum, Hung-Tat Tsui
CVPR (1)3
2001 Separation of Diffuse and Specular Reflection in Color Images
abstract
The presence of specular reflections in images can lead many traditional computer vision algorithms to produce erroneous results. To address this problem, we propose a method based on the neutral interface reflection model for separating the diffuse and specular reflection components in color images. From two photometric images without calibrated lighting, the illuminant chromaticity is estimated, and the RGB intensities of the two reflection components are computed for each pixel using a linear model of surface reflectance. Unlike most previous methods, the presented technique does not assume any dependencies among pixels, such as regionally uniform surface reflectance.
Stephen Lin 0001, Harry Shum
CVPR (1)2
2001 On the Fundamental Limits of Reconstruction-Based Super-Resolution Algorithms
abstract
Super-resolution is a technique that produces higher resolution images from low resolution images (LRIs). In practice, the improvement in resolution is limited. The aim of this paper is to address the problem of whether fundamental limits exist for super-resolution? Specifically, this paper provides explicit limits for a major class of super-resolution algorithms, called reconstruction-based algorithms, under both real and synthetic conditions. Our analysis is based on perturbation theory of linear systems. We also show that a sufficient number of LRIs can be determined to reach the limit. Both real and synthetic experiments are carried out to verify our analysis.
Zhouchen Lin, Harry Shum
CVPR (1)2
2001 Relighting with the Reflected Irradiance Field: Representation, Sampling and Reconstruction
abstract
Image-based relighting (IBL) is a technique to change the illumination of an image-based object/scene. In this paper, we define a representation called the reflected irradiance field which records the reflection from an object surface irradiated by a point light source that moves on a plane. This representation is dual to that of the light field. It synthesizes a novel image under a different illumination by interpolating and superimposing appropriate recorded samples. Furthermore, we study the minimum sampling problem of the reflected irradiance field, i.e., how many point light sources are needed during sampling. We find that there exists a geometry-independent bound for the sampling interval whenever the second-order derivatives of the surface BRDF and the minimum depth of the scene are bounded. This bound ensures that the error in the reconstructed image is controlled by a given tolerance, regardless of the geometry. Experiments on both synthetic and real surfaces are conducted to verify our analysis.
Zhouchen Lin, Tien-Tsin Wong, Harry Shum
CVPR (1)3
2001 A Two-Step Approach to Hallucinating Faces: Global Parametric Model and Local Nonparametric Model
abstract
In this paper, we study face hallucination, or synthesizing a high-resolution face image from low-resolution input, with the help of a large collection of high-resolution face images. We develop a two-step statistical modeling approach that integrates both a global parametric model and a local nonparametric model. First, we derive a global linear model to learn the relationship between the high-resolution face images and their smoothed and down-sampled lower resolution ones. Second, the residual between an original high-resolution image and the reconstructed high-resolution image by a learned linear model is modeled by a patch-based nonparametric Markov network, to capture the high-frequency content of faces. By integrating both global and local models, we can generate photorealistic face images. Our approach is demonstrated by extensive experiments with high-quality hallucinated faces.
Ce Liu 0001, Harry Shum, Changshui Zhang
CVPR (1)2
2001 Optimal Texture Map Reconstruction from Multiple Views
abstract
The recovery of 3D models from multiple reference images involves not only the extraction of 3D shape, but also of texture. Assuming that all surfaces are Lambertian, the resulting final texture is typically computed as a linear combination of reference textures. This is, however, not the optimal means for reconstructing textures, since this does not model the anisotropy in the texture projection. Furthermore, the spatial image sampling may be quite variable within a fore-shortened surface. This also has important implications for computer vision techniques that involve analysis by synthesis and the image-based rendering (IBR) technique of view-dependent texture mapping (VDTM). Starting with sampling theory, we show how weights should be spatially distributed for optimal texture construction. The local weights take into consideration the effects of anisotropy and variable spatial image sampling. We also present experimental results to verify our analysis.
Lifeng Wang 0001, Sing Bing Kang, Richard Szeliski, Harry Shum
CVPR (1)4
2001 Example-Based Facial Sketch Generation with Non-parametric Sampling
abstract
In this paper, we present an example-based facial sketch system. Our system automatically generates a sketch from an input image, by learning from example sketches drawn with a particular style by an artist. There are two key elements in our system: a non-parametric sampling method and a flexible sketch model. Given an input image pixel and its neighborhood, the conditional distribution of a sketch point is computed by querying the examples and finding all similar neighborhoods. An "expected sketch image" is then drawn from the distribution to reflect the drawing style. Finally, facial sketches are obtained by incorporating the sketch model. Experimental results demonstrate the effectiveness of our techniques.
Ying-Qing Xu, Harry Shum, Song-Chun Zhu, Nanning Zheng 0001
ICCV3
2001 Efficient Dense Depth Estimation from Dense Multiperspective Panoramas
Yin Li 0003, Chi-Keung Tang, Harry Shum
ICCV3
2001 Learning Inhomogeneous Gibbs Model of Faces by Minimax Entropy
Ce Liu 0001, Song-Chun Zhu, Harry Shum
ICCV3
2001 Concentric Mosaic(s)Planar Motion and 1D Cameras
abstract
General SFM methods give poor results for images captured by constrained motions such as planar motion of concentric mosaics (CM). In this paper we propose new SFM algorithms for both images captured by CM and composite mosaic images from CM. We first introduce ID affine camera model for completing 1D camera models. Then we show that a 2D image captured by CM can be decoupled into two 1D images: one 1D projective and one ID affine; a composite mosaic image can by rebinned into a calibrated ID panorama projective camera. Finally we describe subspace reconstruction methods and demonstrate both in theory and experiments the advantage of the decomposition method over the general SFM methods by incorporating the constrained motion into the earliest stage of motion analysis.
Long Quan, Le Lu 0001, Harry Shum, Maxime Lhuillier
ICCV3
2001 Image Segmentation by Data Driven Markov Chain Monte Carlo
abstract
This paper presents a computational paradigm called Data Driven Markov Chain Monte Carlo (DDMCMC) for image segmentation in the Bayesian, statistical framework. The paper contributes to image segmentation in three aspects. Firstly, it designs effective and well balanced Markov Chain dynamics to explore the solution space and makes the split and merge process reversible at a middle level vision formulation. Thus it achieves globally optimal solution independent of initial segmentations. Secondly, instead of computing a single maximum a posteriori solution, it proposes a mathematical principle for computing multiple distinct solutions to incorporates intrinsic ambiguities in image segmentation. A k-adventurers algorithm is proposed for extracting distinct multiple solutions from the Markov chain sequence. Thirdly, it utilizes data-driven (bottom-up) techniques, such as clustering and edge detection, to compute importance proposal probabilities, which effectively drive the Markov chain dynamics and achieve tremendous speedup in comparison to traditional jump-diffusion method. Thus DDM-CMC paradigm provides a unifying framework where the role of existing segmentation algorithms, such as; edge detection, clustering, region growing, split-merge, SNAKEs, region competition, are revealed as either realizing Markov chain dynamics or computing importance proposal probabilities. We report some results on color and grey level image segmentation in this paper and refer to a detailed report and a web site for extensive discussion.
Zhuowen Tu, Song-Chun Zhu, Harry Shum
ICCV3
2001 Error Analysis of Pure Rotation-Based Self-Calibration
Leslie Wang, Sing Bing Kang, Harry Shum, Guangyou Xu
ICCV3
2001 On the data compression and transmission aspects of panoramic video
abstract
This paper proposes efficient data compression and transmission techniques for panoramic video. Panoramic videos have been used as a means for representing dynamic scenes or paths along a static environment. They allow the user to change viewpoints interactively at a point in time or space. High-resolution panoramic videos, while desirable, consume a significant amount of storage and bandwidth for transmission, and make real-time decoding very compute-intensive. A high performance MPEG-like compression algorithm, which takes into account the random access requirements and the redundancies of the panoramic video, is presented. The transmission aspects of panoramic video over cable network, LAN and Internet are also briefly discussed.
King To Ng, S. C. Chan 0001, Harry Shum, Sing-Bing Kong
ICIP (2)3
2001 Portrait video phone
abstract
The rapid development of wired and wireless networks tremendouslyfacilitates communications between people. However, most of thecurrent wireless networks still work in low bandwidths, and mobiledevices still suffer from weak computational power, short batterylifetime and limited display capability. We developed a very lowbit-rate bi-level video coding technique, which can be used invideo communications almost anywhere, anytime on any device. Thespirit of this method is that rather than giving highest priorityto the basic colors of an image as in conventional DCT-basedcompression methods, we give preference to the outline features ofscenes when we have limited bandwidths. These features can berepresented by bi-level image sequences that are converted fromgray-scale image sequences. By analyzing the temporal correlationbetween successive frames and flexibilities in the scenepresentation using bi-level images, we achieve very high ratioswith our bi-level video compression scheme. Experiments show thatin low bandwidths, our method provides clearer shape, smoothermotion, shorter initial latency and much cheaper computational costthan do DCT-based methods. Our method is especially suitable forsmall mobile devices such as handheld PCs, palm-size PCs and mobilephones that possess small display screens and light computationalpower, and work in low bandwidth wireless networks. We have builtPC and Pocket PC versions of bi-level video phone systems, whichtypically provide QCIF-size video with a frame rate of 5-15 fps fora 9.6 Kbps bandwidth.
Jiang Li 0008, Keman Yu, Harry Shum, Jizheng Xu, Hanning Zhou, King To Ng, Kaibo Wang
ACM Multimedia4
2001 Portrait video phone
abstract
As the Internet and wirless networks are developed rapidly, the demand of communicating anywhere, anytime on any device emerges. However, most of the current wireless networks still work in low bandwidths, and mobile devices still suffer from weak computational power, short battery lifetime and limited display capability. We developed portrait video phone systems that can run on Pcs and Pocket Pcs at very low bit rates through the Internet. The core technology that portrait video phones employ is the so-called portrait video (or bi-level video) codec. Portrait video codec first converts a full-color video into a black/white image sequence and then compresses it into a black/white portrait-like video. Portrait video processes clearer shape, smoother motion, shorter initial latency, and cheaper computational cost than MPEG2, MPEG4 and H.263 for low bandwidths. Typically the portrait video phone provides QCIF-size video with a frame rate of 5-15 fps for a 9.6 Kbps video bandwidth. The portrait video is so small that it can even be transmitted through an HTTP proxy as text. Experiments show that the portrait video phones work well on ordinary GSM wireless telecommunication networks.
Jiang Li 0008, Keman Yu, Hanning Zhou, Jizheng Xu, King To Ng, Kaibo Wang, Harry Shum
ACM Multimedia10
2001 Speech-driven cartoon animation with emotions
abstract
In this paper, we present a cartoon face animation system for multimedia HCI applications. We animate face cartoons not only from input speech, but also based on emotions derived from speech signal. Using a corpus of over 700 utterances from different speakers, we have trained SVMs (support vector machines) to recognize four categories of emotions: neutral, happiness, anger and sadness. Given each input speech phrase, we identify its emotion content as a mixture of all four emotions, rather than classifying it into a single emotion. Then, facial expressions are= generated from the recovered emotion for each phrase, by morphing different cartoon templates that correspond to various emotions. To ensure smooth transitions in the animation, we apply low-pass filtering to the recovered (and possibly jumpy) emotion sequence. Moreover, lip-syncing is applied to produce the lip movement from speech, by recovering a statistical audio-visual mapping. Experimental results demonstrate that cartoon animation sequences generated by our system are of good and convincing quality.
Feng Yu 0027, Ying-Qing Xu, Eric Chang, Harry Shum
ACM Multimedia5
2001 Synthesizing bidirectional texture functions for real-world surfaces
abstract
In this paper, we present a novel approach to synthetically generating bidirectional texture functions (BTFs) of real-world surfaces. Unlike a conventional two-dimensional texture, a BTF is a six-dimensional function that describes the appearance of texture as a function of illumination and viewing directions. The BTF captures the appearance change caused by visible small-scale geometric details on surfaces. From a sparse set of images under different viewing/lighting settings, our approach generates BTFs in three steps. First, it recovers approximate 3D geometry of surface details using a shape-from-shading method. Then, it generates a novel version of the geometric details that has the same statistical properties as the sample surface with a non-parametric sampling method. Finally, it employs an appearance preserving procedure to synthesize novel images for the recovered or generated geometric details under various viewing/lighting settings, which then define a BTF. Our experimental results demonstrate the effectiveness of our approach.
Xinguo Liu, Yizhou Yu, Harry Shum
SIGGRAPH3
2001 Photorealistic rendering of knitwear using the lumislice
abstract
We present a method for efficient synthesis of photorealistic free-form knitwear. Our approach is motivated by the observation that a single cross-section of yarn can serve as the basic primitive for modeling entire articles of knitwear. This primitive, called the lumislice, describes radiance from a yarn cross-section based on fine-level interactions — such as occlusion, shadowing, and multiple scattering — among yarn fibers. By representing yarn as a sequence of identical but rotated cross-sections, the lumislice can effectively propagate local microstructure over arbitrary stitch patterns and knitwear shapes. This framework accommodates varying levels of detail and capitalizes on hardware-assisted transparency blending. To further enhance realism, a technique for generating soft shadows from yarn is also introduced.
Ying-Qing Xu, Yanyun Chen, Stephen Lin 0001, Enhua Wu, Baining Guo, Harry Shum
SIGGRAPH7
2001 Realistic and efficient rendering of free-form knitwear
abstract
Abstract We present a method for rendering knitwear on free‐form surfaces. This method has three main advantages. First, it renders yarn microstructure realistically and efficiently. Second, the rendering efficiency of yarn microstructure does not come at the price of ignoring the interactions between the neighboring yarn loops. Such interactions are modeled in our system to further enhance realism. Finally, our approach gives the user intuitive control on a few key aspects of knitwear appearance: the fluffiness of the yarn and the irregularity in the positioning of the yarn loops. The result is a system that efficiently produces highly realistic rendering of free‐form knitwear with user control on key aspects of visual appearance. Copyright © 2001 John Wiley & Sons, Ltd.
Ying-Qing Xu, Baining Guo, Harry Shum
Comput. Animat. Virtual Worlds4
2001 Real-time texture synthesis by patch-based sampling
abstract
We present an algorithm for synthesizing textures from an input sample. This patch-based sampling algorithm is fast and it makes high-quality texture synthesis a real-time process. For generating textures of the same size and comparable quality, patch-based sampling is orders of magnitude faster than existing algorithms. The patch-based sampling algorithm works well for a wide variety of textures ranging from regular to stochastic. By sampling patches according to a nonparametric estimation of the local conditional MRF density function, we avoid mismatching features across patch boundaries. We also experimented with documented cases for which pixel-based nonparametric sampling algorithms cease to be effective but our algorithm continues to work well.
Ce Liu 0001, Ying-Qing Xu, Baining Guo, Harry Shum
ACM Trans. Graph.5
2000 Parallel Projections for Stereo Reconstruction
abstract
This paper proposes a novel technique to computing geometric information from images captured under parallel projections. Parallel images are desirable for stereo reconstruction because parallel projection significantly reduces foreshortening. As a result, correlation based matching becomes more effective. Since parallel projection cameras are not commonly available, we construct parallel images by rebinning a large sequence of perspective images. Epipolar geometry, depth recovery and projective invariant for both 1D and 2D parallel stereos are studied. From the uncertainty analysis of depth reconstruction, it is shown that parallel stereo is superior to both conventional perspective stereo and the recently developed multiperspective stereo for vision reconstruction, in that uniform reconstruction error is obtained in parallel stereo. Traditional stereo reconstruction techniques, e.g. multi-baseline stereo, can still be applicable to parallel stereo without any modifications because epipolar lines in a parallel stereo are perfectly straight. Experimental results further confirm the performance of our approach.
Jinxiang Chai, Harry Shum
CVPR2
2000 On the Number of Samples Needed in Light Field Rendering with Constant-Depth Assumption
abstract
While several image-based rendering techniques have been proposed to successfully render scenes/objects from a large collection (e.g., thousands) of images without explicitly recovering 3D structures, the minimum number of images needed to achieve a satisfactory rendering result remains an open problem. This paper is the first attempt to investigate the lower bound for the number of samples needed in the Lumigraph/light field rendering. To simplify the analysis, we consider an ideal scene with only a point that is between a minimum and a maximum range. Furthermore, constant-depth assumption and bilinear interpolation are used for rendering. The constant-depth assumption serves to choose "nearby" rays for interpolation. Our criterion to determine the lower bound is to avoid horizontal and vertical double images, which are caused by interpolation using multiple nearby rays. This criterion is based on the causality requirement in scale-space theory, i.e., no "spurious details" should be generated while smoothing. Using this criterion, closed-form solutions of lower bounds are obtained for both 3D plenoptic function (Concentric Mosaics) and 4D plenoptic function (light field). The bounds are derived completely from the aspect of geometry and are closely related to the resolution of the camera and the depth range of the scene. These tower bounds are further verified by our experimental results.
Zhouchen Lin, Harry Shum
CVPR2
2000 A Linear Algorithm for Camera Self-Calibration, Motion and Structure Recovery for Multi-Planar Scenes from Two Perspective Images
abstract
In this paper we show that given two homography matrices for two planes in space, there is a linear algorithm for the rotation and translation between the two cameras, the focal lengths of the two cameras and the plane equations in the space. Using the estimates as an initial guess, we can further optimize the solution by minimizing the difference between observations and reprojections. Experimental results are shown. We also provide a discussion about the relationship between this approach and the Kruppa equation.
Jun-ichi Terai, Harry Shum
CVPR3
2000 A Spectral Analysis for Light Field Rendering
abstract
Image based rendering using the plenoptic function is an efficient technique for re-rendering at different viewpoints. We study the sampling and reconstruction problem of plenoptic function as a multidimensional sampling problem. The spectral support of plenoptic function is found to be an important quantity in the efficient sampling and reconstruction of such a function. A spectral analysis for the light field, a 4D plenoptic function, is performed. Its spectrum, as a function of the depth function of the scene, is then derived. This result enables us to estimate the spectral support of the light field given some prior estimate of the depth function. Results using a piecewise constant depth model show significant improvement in rendering of the light field images. The design of the reconstruction filter is also discussed.
S. C. Chan 0001, Harry Shum
ICIP2
2000 On the Compression of Image Based Rendering Scene
abstract
In image based rendering (IBR), a 3D scene is recorded through a set of photos, and a novel view is rendered by assembling data from the photo set. Compression is essential to reduce the huge data amount of IBR. We examine three categories of IBR compression algorithms: the block coder, the reference coder and the high dimensional transform (wavelet) coder. It is observed that the block coder consumes the least computation resource, however, its compression ratio is low. The reference coder achieves a good compression ratio with reasonable computation complexity. The high dimensional wavelet coder achieves the best compression ratio, however, it is also the most complex.
Jin Li 0001, Harry Shum, Ya-Qin Zhang
ICIP2
2000 Virtual Reality Using the Concentric Mosaic: Construction, Rendering and Data Compression
abstract
This paper proposes a new image based rendering technique called concentric mosaic for virtual reality applications. It is constructed by capturing vertical slit images when a camera is moving around a set of concentric circles. Concentric mosaic allows the user to move freely in a circular region and observe significant parallax and lighting changes without recovering the geometric and photometric scene model. The rendering of concentric mosaic is very efficient, which amounts to reordering and interpolating of previously captured slit images in the concentric mosaic. Concentric mosaic typically consists of hundreds of high-resolution images, which consumes significant amount of storage and bandwidth for transmission. An MPEG-like compression algorithm is therefore proposed taking advantages of the access patterns and redundancies of the mosaic images. Experimental results show that real-time reconstruction of novel views with good image quality can be achieved in a Pentium II 300 MHz PC.
Harry Shum, King To Ng, S. C. Chan 0001
ICIP1
2000 Plenoptic sampling
abstract
This paper studies the problem of plenoptic sampling in image-based rendering (IBR). From a spectral analysis of light field signals and using the sampling theorem, we mathematically derive the analytical functions to determine the minimum sampling rate for light field rendering. The spectral support of a light field signal is bounded by the minimum and maximum depths only, no matter how complicated the spectral support might be because of depth variations in the scene. The minimum sampling rate for light field rendering is obtained by compacting the replicas of the spectral support of the sampled light field within the smallest interval. Given the minimum and maximum depths, a reconstruction filter with an optimal and constant depth can be designed to achieve anti-aliased light field rendering. Plenoptic sampling goes beyond the minimum number of images needed for anti-aliased light field rendering. More significantly, it utilizes the scene depth information to determine the minimum sampling curve in the joint image and geometry space. The minimum sampling curve quantitatively describes the relationship among three key elements in IBR systems: scene complexity (geometrical and textural information), the number of image samples, and the output resolution. Therefore, plenoptic sampling bridges the gap between image-based rendering and traditional geometry-based rendering. Experimental results demonstrate the effectiveness of our approach.
Jinxiang Chai, S. C. Chan 0001, Harry Shum, Xin Tong 0001
SIGGRAPH3
2000 Review of image-based rendering techniques
Harry Shum, Sing Bing Kang
VCIP1
2000 Real-time stereo rendering of concentric mosaics with linear interpolation
Minsheng Wu, Honghui Sun, Harry Shum
VCIP3
2000 Systems and Experiment Paper: Construction of Panoramic Image Mosaics with Global and Local Alignment
Harry Shum, Richard Szeliski
Int. J. Comput. Vis.1
1999 Efficient Bundle Adjustment with Virtual Key Frames: A Hierarchical Approach to Multi-Frame Structure from Motion
abstract
In this paper we present an efficient hierarchical approach to structure from motion for long image sequences. There are two key elements to our approach: accurate 3D reconstruction for each segment and efficient bundle adjustment for the whole sequence. The image sequence is first divided into a number of segments so that feature points can be reliably tracked across each segment. Each segment has a long baseline to ensure accurate 3D reconstruction. To efficiently bundle adjust 3D structures from ail segments, we reduce the number of frames in each segment by introducing "virtual keyframes". The virtual frames encode the 3D structure of each segment along with its uncertainty but they form a small subset of the original frames. Our method achieves significant speedup over conventional bundle adjustment methods.
Harry Shum, Zhengyou Zhang, Qifa Ke
CVPR1
1999 Omnivergent Stereo
abstract
The notion of a virtual sensor for optimal 3D reconstruction is introduced. Instead of planar perspective images that collect many rays at a fixed viewpoint, omnivergent cameras collect a small number of rays at many different viewpoints. The resulting 2D manifold of rays are arranged into two multiple-perspective images for stereo reconstruction. We call such images omnivergent images, and the process of reconstructing the scene from such images, omnivergent stereo. This procedure is shown to produce 3D scene models with minimal reconstruction error due to the fact that for any point in the 3D scene, two rays with maximum vergence angle can be found in the omnivergent images. Furthermore, omnivergent images are shown to have horizontal epipolar lines, enabling the application of traditional stereo matching algorithms, without modification. Three types of omnivergent virtual sensors are presented: spherical omnivergent cameras, center-strip cameras and dual-strip cameras.
Harry Shum, Adam Tauman Kalai, Steven M. Seitz
ICCV1
1999 Stereo Reconstruction from Multiperspective Panoramas
abstract
The paper presents a new approach to computing depth maps from a large collection of images where the camera motion has been constrained to planar concentric circles. We resample the resulting collection of regular perspective images into a set of multiperspective panoramas, and then compute depth maps directly from these resampled images. Only a small number of multiperspective panoramas is needed to obtain a dense and accurate 3D reconstruction, since our panoramas sample uniformly in three dimensions: rotation angle, inverse radial distance, and vertical elevation. Using multiperspective panoramas avoids the limited overlap between the original input images that causes problems in conventional multi-baseline stereo. Our approach differs from stereo matching of panoramic images taken from different locations, where the epipolar constraints are sine curves. For our multiperspective panoramas, the epipolar geometry to first order consists of horizontal lines. Therefore, any traditional stereo algorithm can be applied to multiperspective panoramas without modification. Experimental results show that our approach generates good depth maps that can be used for image based rendering tasks such as view interpolation and extrapolation.
Harry Shum, Richard Szeliski
ICCV1
1999 What Can Be Determined from a Full and a Weak Perspective Image?
abstract
This paper presents a first investigation on the structure from motion problem from the combination of full and weak perspective images. This problem arises in multiresolution object modeling where multiple zoomed-in or close-up views are combined with wider or distant reference views. The narrow field-of-view (FOV) images from the zoomed-in or closeup views can be approximated as weak perspective projection. Using a full perspective projection model for the narrow FOV images, although more accurate, actually leads to instabilities during the estimation process due to the non-linearities in the imaging model. The weak perspective approximation leads to more stable estimation algorithms, although at the cost of a small amount of modeling inaccuracy. Previous work in structure from motion focused either on two (or more) perspective images or on a set of weak perspective (more generally, affine) images. The main contribution of this paper is the study of the SFM problem for the much neglected case of one perspective and one (or more) weak perspective image. We show that in contrast to the case of a pair of weak perspective images, there is adequate information to recover Euclidean structure from a single perspective and a single weak perspective image. The epipolar geometry is simpler than with two perspective images leading to simpler and more stable estimation algorithms. Computer simulation shows that more stable results can be obtained with the technique presented in this paper than if two images are both considered to be full perspective.
Zhengyou Zhang, P. Anandan 0001, Harry Shum
ICCV3
1999 Rendering with Concentric Mosaics
abstract
This paper presents a novel 3D plenoptic function, which we call concentric mosaics.We constrain camera motion to planar concentric circles, and create concentric mosaics using a manifold mosaic for each circle (i.e., composing slit images taken at different locations).Concentric mosaics index all input image rays naturally in 3 parameters: radius, rotation angle and vertical elevation.Novel views are rendered by combining the appropriate captured rays in an efficient manner at rendering time.Although vertical distortions exist in the rendered images, they can be alleviated by depth correction.Like panoramas, concentric mosaics do not require recovering geometric and photometric scene models.Moreover, concentric mosaics provide a much richer user experience by allowing the user to move freely in a circular region and observe significant parallax and lighting changes.Compared with a Lightfield or Lumigraph, concentric mosaics have much smaller file size because only a 3D plenoptic function is constructed.Concentric mosaics have good space and computational efficiency, and are very easy to capture.This paper discusses a complete working system from capturing, construction, compression, to rendering of concentric mosaics from synthetic and real environments.
Harry Shum, Li-wei He
SIGGRAPH1
1998 Interactive Construction of 3D Models from Panoramic Mosaics
abstract
This paper presents an interactive modeling system that constructs 3D models from a collection of panoramic image mosaics. A panoramic mosaic consists of a set of images taken around the same viewpoint, and a transformation matrix associated with each input image. Our system first recovers the camera pose for each mosaic from known line directions and points, and then constructs the 3D model using all available geometrical constraints. We partition constraints into soft and hard linear constraints so that the modeling process can be formulated as a linearly-constrained least-squares problem, which can be solved efficiently using QR factorization. The results of extracting wire frame and texture-mapped 3D models from single and multiple panoramas are presented.
Harry Shum, Richard Szeliski
CVPR1
1998 Construction and Refinement of Panoramic Mosaics with Global and Local Alignment
abstract
This paper presents techniques for constructing full view panoramic mosaics form sequences of images. Our representation associates a rotation matrix (and optionally a focal length) with each input image, rather than explicitly projecting all of the images onto a common surface (e.g., a cylinder). In order to reduce accumulated registration errors we apply global alignment (block adjustment) to whole sequence of images, which results in an optimal image mosaic (in the least squares sense). To compensate for small amounts of motion parallax introduced by translations of the camera and other unmodeled distortions we develop a local alignment (deghosting) technique which warps each image based on the results of pairwise local image registrations. By combining both global and local alignment we significantly improve the quality of our image mosaics thereby enabling the creation of full view panoramic mosaics with hand-held cameras.
Harry Shum, Richard Szeliski
ICCV1
1998 Interactive 3D modeling from multiple images using scene regularities
abstract
Due to the complexity of real scenes and the fragility of fully automated vision techniques, results from many automated modeling systems are disappointing. Automated techniques often require manual clean-up and postprocessing to segment the scene into coherent objects and surfaces, or to triangulate sparse point matches. They may also be required to enforce geometric constraints such as known orientations of surfaces. For instance, building interiors and exteriors provide vertical and horizontal lines and parallel and perpendicular planes. In this paper, we attack the 3D modeling problem from the other side: we specify some geometric knowledge ahead of time (e.g., known orientations of lines, co-planarity of points, initial scene segmentations), and use these constraints to guide our matching and reconstruction algorithms. We present two interactive (semi-automated) systems for recovering 3D models of large-scale environments from multiple images.
Harry Shum, Richard Szeliski, Simon Baker, P. Anandan 0001
WACV1
1997 Creating full view panoramic image mosaics and environment maps
abstract
This paper presents a novel approach to creating full view panoramic mosaics from image sequences.Unlike current panoramic stitching methods, which usually require pure horizontal camera panning, our system does not require any controlled motions or constraints on how the images are taken (as long as there is no strong motion parallax).For example, images taken from a hand-held digital camera can be stitched seamlessly into panoramic mosaics.Because we represent our image mosaics using a set of transforms, there are no singularity problems such as those existing at the top and bottom of cylindrical or spherical maps.Our algorithm is fast and robust because it directly recovers 3D rotations instead of general 8 parameter planar perspective transforms.Methods to recover camera focal length are also presented.We also present an algorithm for efficiently extracting environment maps from our image mosaics.By mapping the mosaic onto an artibrary texture-mapped polyhedron surrounding the origin, we can explore the virtual environment using standard 3D graphics viewers and hardware without requiring special-purpose players.
Richard Szeliski, Harry Shum
SIGGRAPH2
1997 A Parallel Feature Tracker for Extended Image Sequences
Sing Bing Kang, Richard Szeliski, Harry Shum
Comput. Vis. Image Underst.3
1997 An Integral Approach to Free-Form Object Modeling
abstract
Presents an approach to free-form object modeling from multiple range images. In most conventional approaches, successive views are registered sequentially. In contrast to the sequential approaches, we propose an integral approach which reconstructs statistically optimal object models by simultaneously aggregating all data from multiple views into a weighted least-squares (WLS) formulation. The integral approach has two components. First, a global resampling algorithm constructs partial representations of the object from individual views, so that correspondence can be established among different views. Second, a weighted least-squares algorithm integrates resampled partial representations of multiple views, using the techniques of principal component analysis with missing data (PCAMD). Experiments show that our approach is robust against noise and mismatch.
Harry Shum, Martial Hebert, Katsushi Ikeuchi, Raj Reddy
IEEE Trans. Pattern Anal. Mach. Intell.1
1996 On 3D Shape Similarity
abstract
This paper addresses the problem of 3D shape similarity between closed surfaces. A curved or polyhedral 3D object of genus zero is represented by a mesh that has nearly uniform distribution with known connectivity among mesh nodes. A shape similarity metric is defined based on the L/sub 2/ distance between the local curvature distributions over the mesh representations of the two objects. For both convex and concave objects, the shape metric can be computed in time O(n/sup 2/), where n is the number of tessellations of the sphere or the number of meshes which approximate the surface. Experiments show that our method produces good shape similarity measurements.
Harry Shum, Martial Hebert, Katsushi Ikeuchi
CVPR1
1996 Motion Estimation with Quadtree Splines
abstract
This paper presents a motion estimation algorithm based on a new multiresolution representation, the quadtree spline. This representation describes the motion field as a collection of smoothly connected patches of varying size, where the patch size is automatically adapted to the complexity of the underlying motion. The topology of the patches is determined by a quadtree data structure, and both split and merge techniques are developed for estimating this spatial subdivision. The quadtree spline is implemented using another novel representation, the adaptive hierarchical basis spline, and combines the advantages of adaptively-sized correlation windows with the speedups obtained with hierarchical basis preconditioners. Results are presented on some standard motion sequences.
Richard Szeliski, Harry Shum
IEEE Trans. Pattern Anal. Mach. Intell.2
1995 An Integral Approach to Free-Formed Object Modeling
abstract
Presents a new approach to free-formed object modeling from multiple range images. In most conventional approaches, successive views are registered sequentially. In contrast to the sequential approaches, we propose an integral approach which reconstructs statistically optimal object models by simultaneously aggregating all data from multiple views into a weighted least-squares (WLS) formulation. The integral approach has two components. First, a global resampling algorithm constructs partial representations of the object from individual views so that correspondences can be established among different views. The global resampling algorithm is based on the spherical attribute image (SAI) previously introduced in the context of object representation and recognition. Second, a weighted least-squares algorithm integrates resampled partial representations of multiple views, using the technique of principal component analysis with missing data (PCAMD). Experiments using real range images show that our approach is robust against noise and mismatches, and generates accurate object models.>
Harry Shum, Martial Hebert, Katsushi Ikeuchi, Raj Reddy
ICCV1
1995 Motion Estimation with Quadtree Splines
abstract
This paper presents a motion estimation algorithm based on a new multiresolution representation, the quadtree spline. This representation describes the motion field as a collection of smoothly connected patches of varying size, where the patch size is automatically adapted to the complexity of the underlying motion. The topology of the patches is determined by a quadtree data structure, and both split and merge techniques are developed for estimating this spatial subdivision. The quadtree spline is implemented using another novel representation, the adaptive hierarchical basis spline, and combines the advantages of adaptively-sized correlation windows with the speedups obtained with hierarchical basis preconditioners. Results are presented on some standard motion sequences.>
Richard Szeliski, Harry Shum
ICCV2
1995 Principal Component Analysis with Missing Data and Its Application to Polyhedral Object Modeling
abstract
Observation-based object modeling often requires integration of shape descriptions from different views. To overcome the problems of errors and their accumulation, we have developed a weighted least-squares (WLS) approach which simultaneously recovers object shape and transformation among different views without recovering interframe motion. We show that object modeling from a range image sequence is a problem of principal component analysis with missing data (PCAMD), which can be generalized as a WLS minimization problem. An efficient algorithm is devised. After we have segmented planar surface regions in each view and tracked them over the image sequence, we construct a normal measurement matrix of surface normals, and a distance measurement matrix of normal distances to the origin for all visible regions over the whole sequence of views, respectively. These two matrices, which have many missing elements due to noise, occlusion, and mismatching, enable us to formulate multiple view merging as a combination of two WLS problems. A two-step algorithm is presented. After surface equations are extracted, spatial connectivity among the surfaces is established to enable the polyhedral object model to be constructed. Experiments using synthetic data and real range images show that our approach is robust against noise and mismatching and generates accurate polyhedral object models.>
Harry Shum, Katsushi Ikeuchi, Raj Reddy
IEEE Trans. Pattern Anal. Mach. Intell.1
1994 Principal component analysis with missing data and its application to object modeling
abstract
Observation-based modeling can reduce the cost and effort of model constructions for tasks such as virtual reality environment. Object modeling from a sequence of range images has been formulated as a problem of principal component analysis with missing data (PCAMD), which can be generalized as a weighted least square (WLS) minimization problem. After all visible regions appeared over the whole sequence are segmented and tracked, a normal measurement matrix of surface normals and a distance measurement matrix of normal distances to the origin are constructed respectively. These two measurement matrices, with possibly many missing elements due to occlusion and mismatching, enable us to formulate multiple view merging as a combination of two WLS problems. The solution to the first WLS problem, which employs the quaternion representation of the rotation matrix, yields surface normals and rotation matrices. Subsequently the normal distances and translation vectors are computed by solving the second WLS problem. Experiments using synthetic data and real range images show that our approach is robust against noise and mismatch because it produces a statistically optimal object model by making use of redundancy from multiple views. A toy house model from a sequence of real range images is presented.>
Harry Shum, Katsushi Ikeuchi, Raj Reddy
CVPR1
1994 Virtual reality modeling from a sequence of range images
abstract
Virtual reality object modeling from a sequence of range images has been formulated as a problem of principal component analysis with missing data (PCAMD), which can be generalized as a weighted least square (WLS) minimization problem. An efficient algorithm has been devised to solve the problem of PCAMD. After all visible P regions appeared over the whole sequence of F views are segmented and tracked, a 3F/spl times/P normal measurement matrix of surface normals and an F/spl times/P distance measurement matrix of normal distances to the origin are constructed respectively. These two measurement matrices, with possibly many missing elements due to occlusion and mismatching, enable us to formulate multiple view merging as a combination of two WLS problems. By combining information at both the signal level and the algebraic level, a modified Jarvis' march algorithm is proposed to recover the spatial connectivity among all the reconstructed surface patches. Experiments using synthetic data and real range images show that our approach is robust against noise and mismatch. A toy house model from a sequence of real range images is presented.>
Harry Shum, Katsushi Ikeuchi, Raj Reddy
IROS1
1993 Implementing model-based variable-structure controllers for robot manipulators with actuator modelling
abstract
A model-based control scheme for robot manipulators employing a variable structure control law has been found to perform well, provided that the design parameters are carefully chosen. A refinement of the system model of this original scheme in which the actuator dynamics is taken into consideration is studied. Practical experiments are carried out on a commercial revolute-joint robot manipulator.
P. L. Law, Yangsheng Xu, Harry Shum
IROS4
1992 Adaptive control of space robot system with an attitude controlled base
abstract
The authors discuss adaptive control of a space robot system with an attitude-controlled base on which the robot is attached. An adaptive control scheme in joint space is proposed. Since most tasks are specified in inertia space, instead of joint space, the authors discuss the issues associated to adaptive control in inertia space and identify two potential problems, unavailability of the joint trajectory (since mapping from inertia space trajectory is dynamics-dependent and subject to uncertainty), and nonlinear parameterization in inertia space. For a planar system, the linear parameterization problem is investigated, the design procedure of the controller is illustrated, and the validity and effectiveness of the proposed control scheme are demonstrated.>
Yangsheng Xu, Harry Shum, Ju-Jang Lee, Takeo Kanade
ICRA2
1991 Variable structure model reference adaptive control of robot manipulators
abstract
An adaptive control scheme combining the variable structure and model reference methods is presented. With the variable structure technique, all known parameters of the robot system are fully used while the unknown parameters are adaptively adjusted. The overall control system maintains the basic structure of the computer torque controller, but incorporates adaptive components in the system. The method removes the requirement for persistent excitation, essential to traditional adaptive schemes for satisfactory operation. The control algorithm ensures the robustness of the controlled system with respect to disturbance, since the tracking error always converges to zero theoretically, rather than to an ill-defined residual set as in other adaptive schemes. Using this method, the transient response can be prescribed in advance. Thus, all the outstanding issues in adaptive control are directly treated. Simulation analysis for a two-degree-of-freedom robot is conducted to compare the method with the classical model reference method and computed torque method.>
Yangsheng Xu, Harry Shum
ICRA3