Chi-Keung Tang

dblp:34/4366 · DBLP profile ↗
← Back
143ranked-venue papers
8as first author
42since 2021 · last 2025
0000-0001-6495-3685ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 118 · 6 first-author · 41 since 2021Graphics, computer vision, multimedia, augmented reality and games · 103 · 4 first-author · 27 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Stable Segment Anything Model
abstract
The Segment Anything Model (SAM) achieves remarkable promptable segmentation given high-quality prompts which, however, often require good skills to specify. To make SAM robust to casual prompts, this paper presents the first comprehensive analysis on SAM’s segmentation stability across a diverse spectrum of prompt qualities, notably imprecise bounding boxes and insufficient points. Our key finding reveals that given such low-quality prompts, SAM’s mask decoder tends to activate image features that are biased towards the background or confined to specific object parts. To mitigate this issue, our key idea consists of calibrating solely SAM’s mask attention by adjusting the sampling locations and amplitudes of image features, while the original SAM model architecture and weights remain unchanged. Consequently, our deformable sampling plugin (DSP) enables SAM to adaptively shift attention to the prompted target regions in a data-driven manner. During inference, dynamic routing plugin (DRP) is proposed that toggles SAM between the deformable and regular grid sampling modes, conditioned on the input prompt quality. Thus, our solution, termed Stable-SAM, offers several advantages: 1) improved SAM’s segmentation stability across a wide range of prompt qualities, while 2) retaining SAM’s powerful promptable segmentation efficiency and generality, with 3) minimal learnable parameters (0.08 M) and fast adaptation. Extensive experiments validate the effectiveness and advantages of our approach, underscoring Stable-SAM as a more robust solution for segmenting anything. Codes are at https://github.com/fanq15/Stable-SAM.
Xin Tao 0001, Lei Ke, Mingqiao Ye, Di Zhang 0026, Pengfei Wan 0001, Yu-Wing Tai, Chi-Keung Tang
ICLR8
2025 Motion-Agent: A Conversational Framework for Human Motion Generation with LLMs
abstract
While previous approaches to 3D human motion generation have achieved notable success, they often rely on extensive training and are limited to specific tasks. To address these challenges, we introduce **Motion-Agent**, an efficient conversational framework designed for general human motion generation, editing, and understanding. Motion-Agent employs an open-source pre-trained language model to develop a generative agent, **MotionLLM**, that bridges the gap between motion and text. This is accomplished by encoding and quantizing motions into discrete tokens that align with the language model's vocabulary. With only 1-3% of the model's parameters fine-tuned using adapters, MotionLLM delivers performance on par with diffusion models and other transformer-based methods trained from scratch. By integrating MotionLLM with GPT-4 without additional training, Motion-Agent is able to generate highly complex motion sequences through multi-turn conversations, a capability that previous models have struggled to achieve. Motion-Agent supports a wide range of motion-language tasks, offering versatile capabilities for generating and customizing human motion through interactive conversational exchanges.
Xinhang Liu, Yu-Wing Tai, Chi-Keung Tang
ICLR6
2025 Segment Anything Meets Point Tracking
abstract
Foundation models have marked a significant stride to-ward addressing generalization challenges in deep learning. While the Segment Anything Model (SAM) has established a strong foothold in image segmentation, existing video segmentation methods still require extensive mask labeling for fine-tuning, or face performance drops on unseen data domains otherwise. In this paper, we show how foundation models for image segmentation make a step toward enhancing domain generalizability in video segmentation. We discover that, combined with long-term point tracking, image segmentation models yield state-of-the-art results in zero-shot video segmentation across multiple benchmarks. Surprisingly, point trackers exhibit generalization to domains beyond their synthetic pre-training sequences, which we attribute to the trackers' ability to harness the rich local information in the vicinity of each tracked point. Thus, we introduce SAM-PT, an innovative method for point-centric video segmentation, leveraging the capabilities of SAM alongside long-term point tracking. SAM-PT extends SAM's capability to tracking and segmenting anything in dynamic videos. Unlike traditional video segmentation methods that focus on object-centric mask propagation, our approach uniquely exploits point propagation to utilize local structure information independent of object semantics. The effectiveness of point-based tracking is underscored by direct evaluation on the zero-shot open-world UVO benchmark. Our experiments on popular video object segmentation and multi-object segmentation tracking benchmarks, including DAVIS, YouTube-VOS, and BDD100K, suggest that a point-based segmentation tracker yields better zero-shot performance and efficient interactions. We release our code at https://github.com/SysCV/sam-pt.
Frano Rajic, Lei Ke, Yu-Wing Tai, Chi-Keung Tang, Martin Danelljan, Fisher Yu 0001
WACV4
2025 Robust Object Detection with Domain-Invariant Training and Continual Test-Time Adaptation
Mattia Segù, Bernt Schiele, Dengxin Dai, Yu-Wing Tai, Chi-Keung Tang
Int. J. Comput. Vis.6
2024 SANeRF-HQ: Segment Anything for NeRF in High Quality
abstract
Recently, the Segment Anything Model (SAM) has showcased remarkable capabilities of zero-shot segmentation, while NeRF (Neural Radiance Fields) has gained popularity as a method for various 3D problems beyond novel view synthesis. Though there exist initial attempts to incorporate these two methods into 3D segmentation, they face the challenge of accurately and consistently segmenting objects in complex scenarios. In this paper, we introduce the Segment Anything for NeRF in High Quality (SANeRF-HQ) to achieve high-quality 3D segmentation of any target object in a given scene. SANeRF-HQ utilizes SAM for open-world object segmentation guided by user-supplied prompts, while leveraging NeRF to aggregate information from different viewpoints. To overcome the aforementioned challenges, we employ density field and RGB similarity to enhance the accuracy of segmentation boundary during the aggregation. Emphasizing on segmentation accuracy, we evaluate our method on multiple NeRF datasets where high-quality ground-truths are available or manually annotated. SANeRF-HQ shows a significant quality improvement over state-of-the-art methods in NeRF object segmentation, provides higher flexibility for object localization, and enables more consistent object segmentation across multiple views. Results and code are available at the project site: https://lyclyc52.github.io/SANeRF-HQ/.
Yichen Liu 0001, Benran Hu 0001, Chi-Keung Tang, Yu-Wing Tai
CVPR3
2024 Gear-NeRF: Free-Viewpoint Rendering and Tracking with Motion-Aware Spatio-Temporal Sampling
abstract
Extensions of Neural Radiance Fields (NeRFs) to model dynamic scenes have enabled their near photo-realistic, free-viewpoint rendering. Although these methods have shown some potential in creating immersive experiences, two drawbacks limit their ubiquity: ( i) a significant reduction in reconstruction quality when the computing budget is limited, and (ii) a lack of semantic understanding of the underlying scenes. To address these issues, we introduce Gear-NeRF, which leverages semantic information from powerful image segmentation models. Our approach presents a principled way for learning a spatio-temporal (4D) semantic embedding, based on which we introduce the concept of gears to allow for stratified modeling of dynamic regions of the scene based on the extent of their motion. Such differentiation allows us to adjust the spatio- temporal sampling resolution for each region in proportion to its motion scale, achieving more photo-realistic dynamic novel view synthesis. At the same time, almost for free, our approach enables free-viewpoint tracking of objects of interest - a functionality not yet achieved by existing NeRF-based methods. Empirical studies validate the effectiveness of our method, where we achieve state-of-the-art rendering and tracking performance on multiple challenging datasets. The project page is available at: https://merl.com/research/highlights/gear-nerf
Xinhang Liu, Yu-Wing Tai, Chi-Keung Tang, Pedro Miraldo, Suhas Lohit, Moitreya Chatterjee
CVPR3
2024 C3Net: Compound Conditioned ControlNet for Multimodal Content Generation
abstract
We present Compound Conditioned ControlNet, C3Net, a novel generative neural architecture taking conditions from multiple modalities and synthesizing multimodal contents simultaneously (e.g., image, text, audio). C3Net adapts the ControlNet [46] architecture to jointly train and make inferences on a production-ready diffusion model and its trainable copies. Specifically, C3Net first aligns the conditions from multimodalities to the same semantic latent space using modality-specific encoders based on contrastive training. Then, it generates multimodal outputs based on the aligned latent space, whose semantic information is combined using a ControlNet-like architecture called Control C3-UNet. Correspondingly, with this system design, our model offers an improved solution for joint-modality generation through learning and explaining multimodal conditions, involving more than just linear interpolation within the latent space. Meanwhile, as we align conditions to a unified latent space, C3Net only requires one trainable Control C3-UNet to work on multimodal semantic information. Furthermore, our model employs uni-modal pretraining on the condition alignment stage, outperforming the non-pretrained alignment even on relatively scarce training data and thus demonstrating high-quality compound condition generation. We contribute the first high-quality tri-modal validation set to validate quantitatively that C3Net outperforms or is on par with the first and contemporary state-of-the-art multimodal generation [43]. Our codes and tri-modal dataset will be released here.
Yuehuai Liu, Yu-Wing Tai, Chi-Keung Tang
CVPR4
2024 DragVideo: Interactive Drag-Style Video Editing
Yufan Deng, Ruida Wang, Yu-Wing Tai, Chi-Keung Tang
ECCV (56)5
2024 Deceptive-NeRF/3DGS: Diffusion-Generated Pseudo-observations for High-Quality Sparse-View Reconstruction
Xinhang Liu, Jiaben Chen, Shiu-Hong Kao, Yu-Wing Tai, Chi-Keung Tang
ECCV (16)5
2024 Distill Gold from Massive Ores: Bi-level Data Pruning Towards Efficient Dataset Distillation
Yong-Lu Li 0001, Kaitong Cui, Ziyu Wang 0010, Cewu Lu, Yu-Wing Tai, Chi-Keung Tang
ECCV (20)7
2024 ChatCam: Empowering Camera Control through Conversational AI
abstract
Cinematographers adeptly capture the essence of the world, crafting compelling visual narratives through intricate camera movements. Witnessing the strides made by large language models in perceiving and interacting with the 3D world, this study explores their capability to control cameras with human language guidance. We introduce ChatCam, a system that navigates camera movements through conversations with users, mimicking a professional cinematographer's workflow. To achieve this, we propose CineGPT, a GPT-based autoregressive model for text-conditioned camera trajectory generation. We also develop an Anchor Determinator to ensure precise camera trajectory placement. ChatCam understands user requests and employs our proposed tools to generate trajectories, which can be used to render high-quality video footage on radiance field representations. Our experiments, including comparisons to state-of-the-art approaches and user studies, demonstrate our approach's ability to interpret and execute complex instructions for camera operation, showing promising applications in real-world production settings. Project page: https://xinhangliu.com/chatcam.
Xinhang Liu, Yu-Wing Tai, Chi-Keung Tang
NeurIPS3
2024 Semantic Image Matting: General and Specific Semantics
Yanan Sun 0005, Chi-Keung Tang, Yu-Wing Tai
Int. J. Comput. Vis.2
2024 FSODv2: A Deep Calibrated Few-Shot Object Detection Network
Wei Zhuo 0005, Chi-Keung Tang, Yu-Wing Tai
Int. J. Comput. Vis.3
2023 Ultrahigh Resolution Image/Video Matting with Spatio-Temporal Sparsity
abstract
Commodity ultrahigh definition (UHD) displays are becoming more affordable which demand imaging in ultrahigh resolution (UHR). This paper proposes SparseMat, a computationally efficient approach for UHR image/video matting. Note that it is infeasible to directly process UHR images at full resolution in one shot using existing matting algorithms without running out of memory on consumer-level computational platforms, e.g., Nvidia 1080Ti with 11G memory, while patch-based approaches can introduce unsightly artifacts due to patch partitioning. Instead, our method resorts to spatial and temporal sparsity for addressing general UHR matting. When processing videos, huge computation redundancy can be reduced by exploiting spatial and temporal sparsity. In this paper, we show how to effectively detect spatio-temporal sparsity, which serves as a gate to activate input pixels for the matting model. Under the guidance of such sparsity, our method with sparse high-resolution module (SHM) can avoid patch-based inference while memory efficient for full-resolution matte refinement. Extensive experiments demonstrate that SparseMat can effectively and efficiently generate high-quality alpha matte for UHR images and videos at the original high resolution in a single pass. Project page is in https://github.com/nowsyn/SparseMat.git.
Yanan Sun 0005, Chi-Keung Tang, Yu-Wing Tai
CVPR2
2023 NeRF-RPN: A general framework for object detection in NeRFs
abstract
This paper presents the first significant object detection framework, NeRF-RPN, which directly operates on NeRF. Given a pre-trained NeRF model, NeRF-RPN aims to detect all bounding boxes of objects in a scene. By exploiting a novel voxel representation that incorporates multi-scale 3D neural volumetric features, we demonstrate it is possible to regress the 3D bounding boxes of objects in NeRF directly without rendering the NeRF at any viewpoint. NeRF-RPN is a general framework and can be applied to detect objects without class labels. We experimented NeRF-RPN with various backbone architectures, RPN head designs and loss functions. All of them can be trained in an end-to-end manner to estimate high quality 3D bounding boxes. To facilitate future research in object detection for NeRF, we built a new benchmark dataset which consists of both synthetic and real-world data with careful labeling and clean up. Code and dataset are available at htt ps: //github.com/lyclyc52/NeRF_RPN.
Benran Hu 0001, Junkai Huang 0004, Yichen Liu 0001, Yu-Wing Tai, Chi-Keung Tang
CVPR5
2023 Mask-Free Video Instance Segmentation
abstract
The recent advancement in Video Instance Segmentation (VIS) has largely been driven by the use of deeper and increasingly data-hungry transformer-based models. However, video masks are tedious and expensive to annotate, limiting the scale and diversity of existing VIS datasets. In this work, we aim to remove the mask-annotation requirement. We propose MaskFreeVIS, achieving highly competitive VIS performance, while only using bounding box annotations for the object state. We leverage the rich temporal mask consistency constraints in videos by introducing the Temporal KNN-patch Loss (TK-Loss), providing strong mask supervision without any labels. Our TK-Loss finds one-to-many matches across frames, through an efficient patch-matching step followed by a K-nearest neighbor selection. A consistency loss is then enforced on the found matches. Our mask-free objective is simple to implement, has no trainable parameters, is computationally efficient, yet outperforms baselines employing, e.g., state-of-the-art optical flow to enforce temporal mask consistency. We validate MaskFreeVIS on the YouTube-VIS 2019/2021, OVIS and BDD100K MOTS benchmarks. The results clearly demonstrate the efficacy of our method by drastically narrowing the gap between fully and weakly-supervised VIS performance. Our code and trained models are available at http://vis.xyz/pub/maskfreevis.
Lei Ke, Martin Danelljan, Henghui Ding, Yu-Wing Tai, Chi-Keung Tang, Fisher Yu 0001
CVPR5
2023 Instance Neural Radiance Field
abstract
This paper presents one of the first learning-based NeRF 3D instance segmentation pipelines, dubbed as Instance Neural Radiance Field, or Instance-NeRF. Taking a NeRF pretrained from multi-view RGB images as input, Instance-NeRF can learn 3D instance segmentation of a given scene, represented as an instance field component of the NeRF model. To this end, we adopt a 3D proposal-based mask prediction network on the sampled volumetric features from NeRF, which generates discrete 3D instance masks. The coarse 3D mask prediction is then projected to image space to match 2D segmentation masks from different views generated by existing panoptic segmentation models, which are used to supervise the training of the instance field. Notably, beyond generating consistent 2D segmentation maps from novel views, Instance-NeRF can query instance information at any 3D point, which greatly enhances NeRF object segmentation and manipulation. Our method is also one of the first to achieve such results in pure inference. Experimented on synthetic and real-world NeRF datasets with complex indoor scenes, Instance-NeRF surpasses previous NeRF segmentation works and competitive 2D segmentation methods in segmentation performance on unseen views. Code and data are available at https://github.com/lyclyc52/Instance_NeRF.
Yichen Liu 0001, Benran Hu 0001, Junkai Huang 0004, Yu-Wing Tai, Chi-Keung Tang
ICCV5
2023 EgoPCA: A New Framework for Egocentric Hand-Object Interaction Understanding
abstract
With the surge in attention to Egocentric Hand-Object Interaction (Ego-HOI), large-scale datasets such as Ego4D and EPIC-KITCHENS have been proposed. However, most current research is built on resources derived from third-person video action recognition. This inherent domain gap between first- and third-person action videos, which have not been adequately addressed before, makes current Ego-HOI suboptimal. This paper rethinks and proposes a new framework as an infrastructure to advance Ego-HOI recognition by Probing, Curation and Adaption (EgoPCA). We contribute comprehensive pre-train sets, balanced test sets and a new baseline, which are complete with a training-finetuning strategy. With our new framework, we not only achieve state-of-the-art performance on Ego-HOI benchmarks but also build several new and effective mechanisms and settings to advance further research. We believe our data and the findings will pave a new way for Ego-HOI understanding. Code and data are available at https://mvig-rhos.com/ego_pca.
Yong-Lu Li 0001, Zhemin Huang 0001, Michael Xu Liu, Cewu Lu, Yu-Wing Tai, Chi-Keung Tang
ICCV7
2023 Cascade-DETR: Delving into High-Quality Universal Object Detection
abstract
Object localization in general environments is a fundamental part of vision systems. While dominating on the COCO benchmark, recent Transformer-based detection methods are not competitive in diverse domains. Moreover, these methods still struggle to very accurately estimate the object bounding boxes in complex environments.We introduce Cascade-DETR for high-quality universal object detection. We jointly tackle the generalization to diverse domains and localization accuracy by proposing the Cascade Attention layer, which explicitly integrates object-centric information into the detection decoder by limiting the attention to the previous box prediction. To further enhance accuracy, we also revisit the scoring of queries. Instead of relying on classification scores, we predict the expected IoU of the query, leading to substantially more well-calibrated confidences. Lastly, we introduce a universal object detection benchmark, UDB10, that contains 10 datasets from diverse domains. While also advancing the state-of-the-art on COCO, Cascade-DETR substantially improves DETR-based detectors on all datasets in UDB10, even by over 10 mAP in some cases. The improvements under stringent quality requirements are even more pronounced. Our code and pretrained models are at https://github.com/SysCV/cascade-detr.
Mingqiao Ye, Lei Ke, Siyuan Li 0008, Yu-Wing Tai, Chi-Keung Tang, Martin Danelljan, Fisher Yu 0001
ICCV5
2023 Towards Robust Object Detection Invariant to Real-World Domain Shifts
Mattia Segù, Yu-Wing Tai, Fisher Yu 0001, Chi-Keung Tang, Bernt Schiele, Dengxin Dai
ICLR5
2023 Segment Anything in High Quality
abstract
The recent Segment Anything Model (SAM) represents a big leap in scaling up segmentation models, allowing for powerful zero-shot capabilities and flexible prompting. Despite being trained with 1.1 billion masks, SAM's mask prediction quality falls short in many cases, particularly when dealing with objects that have intricate structures. We propose HQ-SAM, equipping SAM with the ability to accurately segment any object, while maintaining SAM's original promptable design, efficiency, and zero-shot generalizability. Our careful design reuses and preserves the pre-trained model weights of SAM, while only introducing minimal additional parameters and computation. We design a learnable High-Quality Output Token, which is injected into SAM's mask decoder and is responsible for predicting the high-quality mask. Instead of only applying it on mask-decoder features, we first fuse them with early and final ViT features for improved mask details. To train our introduced learnable parameters, we compose a dataset of 44K fine-grained masks from several sources. HQ-SAM is only trained on the introduced detaset of 44k masks, which takes only 4 hours on 8 GPUs. We show the efficacy of HQ-SAM in a suite of 10 diverse segmentation datasets across different downstream tasks, where 8 out of them are evaluated in a zero-shot transfer protocol. Our code and pretrained models are at https://github.com/SysCV/SAM-HQ.
Lei Ke, Mingqiao Ye, Martin Danelljan, Yifan Liu 0001, Yu-Wing Tai, Chi-Keung Tang, Fisher Yu 0001
NeurIPS6
2023 BiMatting: Efficient Video Matting via Binarization
abstract
Real-time video matting on edge devices faces significant computational resource constraints, limiting the widespread use of video matting in applications such as online conferences and short-form video production. Binarization is a powerful compression approach that greatly reduces computation and memory consumption by using 1-bit parameters and bitwise operations. However, binarization of the video matting model is not a straightforward process, and our empirical analysis has revealed two primary bottlenecks: severe representation degradation of the encoder and massive redundant computations of the decoder. To address these issues, we propose BiMatting, an accurate and efficient video matting model using binarization. Specifically, we construct shrinkable and dense topologies of the binarized encoder block to enhance the extracted representation. We sparsify the binarized units to reduce the low-information decoding computation. Through extensive experiments, we demonstrate that BiMatting outperforms other binarized video matting models, including state-of-the-art (SOTA) binarization methods, by a significant margin. Our approach even performs comparably to the full-precision counterpart in visual quality. Furthermore, BiMatting achieves remarkable savings of 12.4$\times$ and 21.6$\times$ in computation and storage, respectively, showcasing its potential and advantages in real-world resource-constrained scenarios. Our code and models are released at https://github.com/htqin/BiMatting .
Haotong Qin, Lei Ke, Xudong Ma, Martin Danelljan, Yu-Wing Tai, Chi-Keung Tang, Xianglong Liu 0001, Fisher Yu 0001
NeurIPS6
2023 FaceDNeRF: Semantics-Driven Face Reconstruction, Prompt Editing and Relighting with Diffusion Models
abstract
The ability to create high-quality 3D faces from a single image has become increasingly important with wide applications in video conferencing, AR/VR, and advanced video editing in movie industries. In this paper, we propose Face Diffusion NeRF (FaceDNeRF), a new generative method to reconstruct high-quality Face NeRFs from single images, complete with semantic editing and relighting capabilities. FaceDNeRF utilizes high-resolution 3D GAN inversion and expertly trained 2D latent-diffusion model, allowing users to manipulate and construct Face NeRFs in zero-shot learning without the need for explicit 3D data. With carefully designed illumination and identity preserving loss, as well as multi-modal pre-training, FaceDNeRF offers users unparalleled control over the editing process enabling them to create and edit face NeRFs using just single-view images, text prompts, and explicit target lighting. The advanced features of FaceDNeRF have been designed to produce more impressive results than existing 2D editing approaches that rely on 2D segmentation maps for editable attributes. Experiments show that our FaceDNeRF achieves exceptionally realistic results and unprecedented flexibility in editing compared with state-of-the-art 3D face reconstruction and editing methods. Our code will be available at https://github.com/BillyXYB/FaceDNeRF.
Hao Zhang 0106, Tianyuan Dai, Yanbo Xu, Yu-Wing Tai, Chi-Keung Tang
NeurIPS5
2023 Occlusion-Aware Instance Segmentation Via BiLayer Network Architectures
abstract
Segmenting highly-overlapping image objects is challenging, because there is typically no distinction between real object contours and occlusion boundaries on images. Unlike previous instance segmentation methods, we model image formation as a composition of two overlapping layers, and propose Bilayer Convolutional Network (BCNet), where the top layer detects occluding objects (occluders) and the bottom layer infers partially occluded instances (occludees). The explicit modeling of occlusion relationship with bilayer structure naturally decouples the boundaries of both the occluding and occluded instances, and considers the interaction between them during mask regression. We investigate the efficacy of bilayer structure using two popular convolutional network designs, namely, Fully Convolutional Network (FCN) and Graph Convolutional Network (GCN). Further, we formulate bilayer decoupling using the vision transformer (ViT), by representing instances in the image as separate learnable occluder and occludee queries. Large and consistent improvements using one/two-stage and query-based object detectors with various backbones and network layer choices validate the generalization ability of bilayer decoupling, as shown by extensive experiments on image instance segmentation benchmarks (COCO, KINS, COCOA) and video instance segmentation benchmarks (YTVIS, OVIS, BDD100 K MOTS), especially for heavy occlusion cases.
Lei Ke, Yu-Wing Tai, Chi-Keung Tang
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 GCoNet+: A Stronger Group Collaborative Co-Salient Object Detector
abstract
In this paper, we present a novel end-to-end group collaborative learning network, termed GCoNet+, which can effectively and efficiently (250 fps) identify co-salient objects in natural scenes. The proposed GCoNet+ achieves the new state-of-the-art performance for co-salient object detection (CoSOD) through mining consensus representations based on the following two essential criteria: 1) intra-group compactness to better formulate the consistency among co-salient objects by capturing their inherent shared attributes using our novel group affinity module (GAM); 2) inter-group separability to effectively suppress the influence of noisy objects on the output by introducing our new group collaborating module (GCM) conditioning on the inconsistent consensus. To further improve the accuracy, we design a series of simple yet effective components as follows: i) a recurrent auxiliary classification module (RACM) promoting model learning at the semantic level; ii) a confidence enhancement module (CEM) assisting the model in improving the quality of the final predictions; and iii) a group-based symmetric triplet (GST) loss guiding the model to learn more discriminative features. Extensive experiments on three challenging benchmarks, i.e., CoCA, CoSOD3k, and CoSal2015, demonstrate that our GCoNet+ outperforms the existing 12 cutting-edge models. Code has been released at https://github.com/ZhengPeng7/GCoNet_plus.
Peng Zheng 0004, Huazhu Fu, Deng-Ping Fan, Jie Qin 0004, Yu-Wing Tai, Chi-Keung Tang, Luc Van Gool
IEEE Trans. Pattern Anal. Mach. Intell.7
2022 Human Instance Matting via Mutual Guidance and Multi-Instance Refinement
abstract
This paper introduces a new matting task called human instance matting (HIM), which requires the pertinent model to automatically predict a precise alpha matte for each human instance. Straightforward combination of closely related techniques, namely, instance segmentation, soft segmentation and human/conventional matting, will easily fail in complex cases requiring disentangling mingled colors belonging to multiple instances along hairy and thin boundary structures. To tackle these technical challenges, we propose a human instance matting framework, called InstMatt, where a novel mutual guidance strategy working in tandem with a multi-instance refinement module is used, for delineating multi-instance relationship among humans with complex and overlapping boundaries if present. A new instance matting metric called instance matting quality (IMQ) is proposed, which addresses the absence of a unified and fair means of evaluation emphasizing both instance recognition and matting quality. Finally, we construct a HIM benchmark for evaluation, which comprises of both synthetic and natural benchmark images. In addition to thorough experimental results on complex cases with multiple and overlapping human instances each has intricate boundaries, preliminary results are presented on general instance matting. Code and benchmark are available in https://github.com/nowsyn/InstMatt.
Yanan Sun 0005, Chi-Keung Tang, Yu-Wing Tai
CVPR2
2022 Mask Transfiner for High-Quality Instance Segmentation
abstract
Two-stage and query-based instance segmentation methods have achieved remarkable results. However, their segmented masks are still very coarse. In this paper, we present Mask Transfiner for high-quality and efficient instance segmentation. Instead of operating on regular dense tensors, our Mask Transfiner decomposes and represents the image regions as a quadtree. Our transformer-based approach only processes detected error-prone tree nodes and self-corrects their errors in parallel. While these sparse pixels only constitute a small proportion of the total number, they are critical to the final mask quality. This allows Mask Transfiner to predict highly accurate instance masks, at a low computational cost. Extensive experiments demonstrate that Mask Transfiner outperforms current instance segmentation methods on three popular benchmarks, significantly improving both two-stage and query-based frameworks by a large margin of +3.0 mask AP on COCO and BDD100K, and +6.6 boundary AP on Cityscapes. Our code and trained models are available at https://github.com/SysCV/transfiner.
Lei Ke, Martin Danelljan, Xia Li 0005, Yu-Wing Tai, Chi-Keung Tang, Fisher Yu 0001
CVPR5
2022 Interactiveness Field in Human-Object Interactions
abstract
Human-Object Interaction (HOI) detection plays a core role in activity understanding. Though recent two/one-stage methods have achieved impressive results, as an essential step, discovering interactive human-object pairs remains challenging. Both one/two-stage methods fail to effectively extract interactive pairs instead of generating redundant negative pairs. In this work, we introduce a previously overlooked interactiveness bimodal prior: given an object in an image, after pairing it with the humans, the generated pairs are either mostly non-interactive, or mostly interactive, with the former more frequent than the latter. Based on this interactiveness bimodal prior we propose the “interactiveness field”. To make the learned field compatible with real HOI image considerations, we propose new energy constraints based on the cardinality and difference in the inherent “interactiveness field” underlying interactive versus non-interactive pairs. Consequently, our method can detect more precise pairs and thus significantly boost HOI detection performance, which is validated on widely-used benchmarks where we achieve decent improvements over state-of-the-arts. Our code is available at https://github.comIForuckllnteractiveness-Field.
Xinpeng Liu 0002, Yong-Lu Li 0001, Yu-Wing Tai, Cewu Lu, Chi-Keung Tang
CVPR6
2022 Self-support Few-Shot Semantic Segmentation
Wenjie Pei, Yu-Wing Tai, Chi-Keung Tang
ECCV (19)4
2022 Few-Shot Video Object Detection
Chi-Keung Tang, Yu-Wing Tai
ECCV (20)2
2022 Few-Shot Object Detection with Model Calibration
Chi-Keung Tang, Yu-Wing Tai
ECCV (19)2
2022 Video Mask Transfiner for High-Quality Video Instance Segmentation
Lei Ke, Henghui Ding, Martin Danelljan, Yu-Wing Tai, Chi-Keung Tang, Fisher Yu 0001
ECCV (28)5
2022 Unsupervised Multi-View Object Segmentation Using Radiance Field Propagation
abstract
We present radiance field propagation (RFP), a novel approach to segmenting objects in 3D during reconstruction given only unlabeled multi-view images of a scene. RFP is derived from emerging neural radiance field-based techniques, which jointly encodes semantics with appearance and geometry. The core of our method is a novel propagation strategy for individual objects' radiance fields with a bidirectional photometric loss, enabling an unsupervised partitioning of a scene into salient or meaningful regions corresponding to different object instances. To better handle complex scenes with multiple objects and occlusions, we further propose an iterative expectation-maximization algorithm to refine object masks. To the best of our knowledge, RFP is the first unsupervised approach for tackling 3D scene object segmentation for neural radiance field (NeRF) without any supervision, annotations, or other cues such as 3D bounding boxes and prior knowledge of object class. Experiments demonstrate that RFP achieves feasible segmentation results that are more accurate than previous unsupervised image/scene segmentation approaches, and are comparable to existing supervised NeRF-based methods. The segmented object representations enable individual 3D object editing operations. Codes and datasets will be made publicly available.
Xinhang Liu, Jiaben Chen, Huai Yu, Yu-Wing Tai, Chi-Keung Tang
NeurIPS5
2021 Modular Interactive Video Object Segmentation: Interaction-to-Mask, Propagation and Difference-Aware Fusion
abstract
We present Modular interactive VOS (MiVOS) framework which decouples interaction-to-mask and mask propagation, allowing for higher generalizability and better performance. Trained separately, the interaction module converts user interactions to an object mask, which is then temporally propagated by our propagation module using a novel top-k filtering strategy in reading the space-time memory. To effectively take the user’s intent into account, a novel difference-aware module is proposed to learn how to properly fuse the masks before and after each interaction, which are aligned with the target frames by employing the space-time memory. We evaluate our method both qualitatively and quantitatively with different forms of user interactions (e.g., scribbles, clicks) on DAVIS to show that our method outperforms current state-of-the-art algorithms while requiring fewer frame interactions, with the additional advantage in generalizing to different types of user interactions. We contribute a large-scale synthetic VOS dataset with pixel-accurate segmentation of 4.8M frames to accompany our source codes to facilitate future research.
Ho Kei Cheng, Yu-Wing Tai, Chi-Keung Tang
CVPR3
2021 Group Collaborative Learning for Co-Salient Object Detection
abstract
We present a novel group collaborative learning framework (GCoNet) capable of detecting co-salient objects in real time (16ms), by simultaneously mining consensus representations at group level based on the two necessary criteria: 1) intra-group compactness to better formulate the consistency among co-salient objects by capturing their inherent shared attributes using our novel group affinity module; 2) inter-group separability to effectively suppress the influence of noisy objects on the output by introducing our new group collaborating module conditioning the inconsistent consensus. To learn a better embedding space without extra computational overhead, we explicitly employ auxiliary classification supervision. Extensive experiments on three challenging benchmarks, i.e., CoCA, CoSOD3k, and Cosal2015, demonstrate that our simple GCoNet outperforms 10 cutting-edge models and achieves the new state-of-the-art. We demonstrate this paper’s new technical contributions on a number of important downstream computer vision applications including content aware co-segmentation, co-localization based automatic thumbnails, etc. Code has been made publicly available: https://github.com/fanq15/GCoNet.
Deng-Ping Fan, Huazhu Fu, Chi-Keung Tang, Ling Shao 0001, Yu-Wing Tai
CVPR4
2021 Deep Occlusion-Aware Instance Segmentation With Overlapping BiLayers
abstract
Segmenting highly-overlapping objects is challenging, because typically no distinction is made between real object contours and occlusion boundaries. Unlike previous two-stage instance segmentation methods, we model image formation as composition of two overlapping layers, and propose Bilayer Convolutional Network (BCNet), where the top GCN layer detects the occluding objects (occluder) and the bottom GCN layer infers partially occluded instance (occludee). The explicit modeling of occlusion relationship with bilayer structure naturally decouples the boundaries of both the occluding and occluded instances, and considers the interaction between them during mask regression. We validate the efficacy of bilayer decoupling on both one-stage and two-stage object detectors with different backbones and network layer choices. Despite its simplicity, extensive experiments on COCO and KINS show that our occlusion-aware BCNet achieves large and consistent performance gain especially for heavy occlusion cases. Code is available at https://github.com/lkeab/BCNet.
Lei Ke, Yu-Wing Tai, Chi-Keung Tang
CVPR3
2021 Semantic Image Matting
abstract
Natural image matting separates the foreground from background in fractional occupancy which can be caused by highly transparent objects, complex foreground (e.g., net or tree), and/or objects containing very fine details (e.g., hairs). Although conventional matting formulation can be applied to all of the above cases, no previous work has attempted to reason the underlying causes of matting due to various foreground semantics.We show how to obtain better alpha mattes by incorporating into our framework semantic classification of matting regions. Specifically, we consider and learn 20 classes of matting patterns, and propose to extend the conventional trimap to semantic trimap. The proposed semantic trimap can be obtained automatically through patch structure analysis within trimap regions. Meanwhile, we learn a multi-class discriminator to regularize the alpha prediction at semantic level, and content-sensitive weights to balance different regularization losses. Experiments on multiple benchmarks show that our method outperforms other methods and has achieved the most competitive state-of-the-art performance. Finally, we contribute a large-scale Semantic Image Matting Dataset with careful consideration of data balancing across different semantic classes. Code and dataset are available at https://github.com/nowsyn/SIM.
Yanan Sun 0005, Chi-Keung Tang, Yu-Wing Tai
CVPR2
2021 Deep Video Matting via Spatio-Temporal Alignment and Aggregation
abstract
Despite the significant progress made by deep learning in natural image matting, there has been so far no representative work on deep learning for video matting due to the inherent technical challenges in reasoning temporal domain and lack of large-scale video matting datasets. In this paper, we propose a deep learning-based video matting framework which employs a novel and effective spatio-temporal feature aggregation module (ST-FAM). As optical flow estimation can be very unreliable within matting regions, ST-FAM is designed to effectively align and aggregate information across different spatial scales and temporal frames within the network decoder. To eliminate frame-by-frame trimap annotations, a lightweight interactive trimap propagation network is also introduced. The other contribution consists of a large-scale video matting dataset with groundtruth alpha mattes for quantitative evaluation and real-world high-resolution videos with trimaps for qualitative evaluation. Quantitative and qualitative experimental results show that our framework significantly outperforms conventional video matting and deep image matting methods applied to video in presence of multi-frame temporal information. Our dataset is available at https://github.com/nowsyn/DVM.
Yanan Sun 0005, Guanzhi Wang, Qiao Gu, Chi-Keung Tang, Yu-Wing Tai
CVPR4
2021 HAA500: Human-Centric Atomic Action Dataset with Curated Videos
abstract
We contribute HAA5001, a manually annotated human-centric atomic action dataset for action recognition on 500 classes with over 591K labeled frames. To minimize ambiguities in action classification, HAA500 consists of highly diversified classes of fine-grained atomic actions, where only consistent actions fall under the same label, e.g., "Baseball Pitching" vs "Free Throw in Basketball". Thus HAA500 is different from existing atomic action datasets, where coarse-grained atomic actions were labeled with coarse action-verbs such as "Throw". HAA500 has been carefully curated to capture the precise movement of human figures with little class-irrelevant motions or spatiotemporal label noises.The advantages of HAA500 are fourfold: 1) human-centric actions with a high average of 69.7% detectable joints for the relevant human poses; 2) high scalability since adding a new class can be done under 20–60 minutes; 3) curated videos capturing essential elements of an atomic action without irrelevant frames; 4) fine-grained atomic action classes. Our extensive experiments including cross-data validation using datasets collected in the wild demonstrate the clear benefits of human-centric and atomic characteristics of HAA500, which enable training even a baseline deep learning model to improve prediction by attending to atomic human poses. We detail the HAA500 dataset statistics and collection methodology and compare quantitatively with existing action recognition datasets.
Jihoon Chung, Cheng-hsin Wuu, Hsuan-ru Yang, Yu-Wing Tai, Chi-Keung Tang
ICCV5
2021 Occlusion-Aware Video Object Inpainting
abstract
Conventional video inpainting is neither object-oriented nor occlusion-aware, making it liable to obvious artifacts when large occluded object regions are inpainted. This paper presents occlusion-aware video object inpainting, which recovers both the complete shape and appearance for occluded objects in videos given their visible mask segmentation.To facilitate this new research, we construct the first large-scale video object inpainting benchmark YouTube-VOI to provide realistic occlusion scenarios with both occluded and visible object masks available. Our technical contribution VOIN jointly performs video object shape completion and occluded texture generation. In particular, the shape completion module models long-range object coherence while the flow completion module recovers accurate flow with sharp motion boundary, for propagating temporally-consistent texture to the same moving object across frames. For more realistic results, VOIN is optimized using both T-PatchGAN and a new spatio-temporal attention-based multi-class discriminator.Finally, we compare VOIN and strong baselines on YouTube-VOI. Experimental results clearly demonstrate the efficacy of our method including inpainting complex and dynamic objects. VOIN degrades gracefully with inaccurate input visible mask.
Lei Ke, Yu-Wing Tai, Chi-Keung Tang
ICCV3
2021 Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object Segmentation
abstract
This paper presents a simple yet effective approach to modeling space-time correspondences in the context of video object segmentation. Unlike most existing approaches, we establish correspondences directly between frames without re-encoding the mask features for every object, leading to a highly efficient and robust framework. With the correspondences, every node in the current query frame is inferred by aggregating features from the past in an associative fashion. We cast the aggregation process as a voting problem and find that the existing inner-product affinity leads to poor use of memory with a small (fixed) subset of memory nodes dominating the votes, regardless of the query. In light of this phenomenon, we propose using the negative squared Euclidean distance instead to compute the affinities. We validated that every memory node now has a chance to contribute, and experimentally showed that such diversified voting is beneficial to both memory efficiency and inference accuracy. The synergy of correspondence networks and diversified voting works exceedingly well, achieves new state-of-the-art results on both DAVIS and YouTubeVOS datasets while running significantly faster at 20+ FPS for multiple objects without bells and whistles.
Ho Kei Cheng, Yu-Wing Tai, Chi-Keung Tang
NeurIPS3
2021 Prototypical Cross-Attention Networks for Multiple Object Tracking and Segmentation
abstract
Multiple object tracking and segmentation requires detecting, tracking, and segmenting objects belonging to a set of given classes. Most approaches only exploit the temporal dimension to address the association problem, while relying on single frame predictions for the segmentation mask itself. We propose Prototypical Cross-Attention Network (PCAN), capable of leveraging rich spatio-temporal information for online multiple object tracking and segmentation. PCAN first distills a space-time memory into a set of prototypes and then employs cross-attention to retrieve rich information from the past frames. To segment each object, PCAN adopts a prototypical appearance module to learn a set of contrastive foreground and background prototypes, which are then propagated over time. Extensive experiments demonstrate that PCAN outperforms current video instance tracking and segmentation competition winners on both Youtube-VIS and BDD100K datasets, and shows efficacy to both one-stage and two-stage segmentation frameworks. Code and video resources are available at http://vis.xyz/pub/pcan.
Lei Ke, Xia Li 0005, Martin Danelljan, Yu-Wing Tai, Chi-Keung Tang, Fisher Yu 0001
NeurIPS5
2020 CascadePSP: Toward Class-Agnostic and Very High-Resolution Segmentation via Global and Local Refinement
abstract
State-of-the-art semantic segmentation methods were almost exclusively trained on images within a fixed resolution range. These segmentations are inaccurate for very high-resolution images since using bicubic upsampling of low-resolution segmentation does not adequately capture high-resolution details along object boundaries. In this paper, we propose a novel approach to address the high-resolution segmentation problem without using any high-resolution training data. The key insight is our CascadePSP network which refines and corrects local boundaries whenever possible. Although our network is trained with low-resolution segmentation data, our method is applicable to any resolution even for very high-resolution images larger than 4K. We present quantitative and qualitative studies on different datasets to show that CascadePSP can reveal pixel-accurate segmentation boundaries using our novel refinement module without any finetuning. Thus, our method can be regarded as class-agnostic. Finally, we demonstrate the application of our model to scene parsing in multi-class segmentation.
Ho Kei Cheng, Jihoon Chung, Yu-Wing Tai, Chi-Keung Tang
CVPR4
2020 Few-Shot Object Detection With Attention-RPN and Multi-Relation Detector
abstract
Conventional methods for object detection typically require a substantial amount of training data and preparing such high-quality training data is very labor-intensive. In this paper, we propose a novel few-shot object detection network that aims at detecting objects of unseen categories with only a few annotated examples. Central to our method are our Attention-RPN, Multi-Relation Detector and Contrastive Training strategy, which exploit the similarity between the few shot support set and query set to detect novel objects while suppressing false detection in the background. To train our network, we contribute a new dataset that contains 1000 categories of various objects with high-quality annotations. To the best of our knowledge, this is one of the first datasets specifically designed for few-shot object detection. Once our few-shot network is trained, it can detect objects of unseen categories without further training or fine-tuning. Our method is general and has a wide range of potential applications. We produce a new state-of-the-art performance on different datasets in the few-shot setting. The dataset link is https://github.com/fanq15/Few-Shot-Object-Detection-Dataset.
Wei Zhuo 0005, Chi-Keung Tang, Yu-Wing Tai
CVPR3
2020 Fast Video Object Segmentation With Temporal Aggregation Network and Dynamic Template Matching
abstract
Significant progress has been made in Video Object Segmentation (VOS), the video object tracking task in its finest level. While the VOS task can be naturally decoupled into image semantic segmentation and video object tracking, significantly much more research effort has been made in segmentation than tracking. In this paper, we introduce “tracking-by-detection” into VOS which can coherently integrates segmentation into tracking, by proposing a new temporal aggregation network and a novel dynamic time-evolving template matching mechanism to achieve significantly improved performance. Notably, our method is entirely online and thus suitable for one-shot learning, and our end-to-end trainable model allows multiple object segmentation in one forward pass. We achieve new state-of-the-art performance on the DAVIS benchmark without complicated bells and whistles in both speed and accuracy, with a speed of 0.14 second per frame and J &F measure of 75.9% respectively.
Xuhua Huang, Yu-Wing Tai, Chi-Keung Tang
CVPR4
2020 Cascaded Deep Monocular 3D Human Pose Estimation With Evolutionary Training Data
abstract
End-to-end deep representation learning has achieved remarkable accuracy for monocular 3D human pose estimation, yet these models may fail for unseen poses with limited and fixed training data. This paper proposes a novel data augmentation method that: (1) is scalable for synthesizing massive amount of training data (over 8 million valid 3D human poses with corresponding 2D projections) for training 2D-to-3D networks, (2) can effectively reduce dataset bias. Our method evolves a limited dataset to synthesize unseen 3D human skeletons based on a hierarchical human representation and heuristics inspired by prior knowledge. Extensive experiments show that our approach not only achieves state-of-the-art accuracy on the largest public benchmark, but also generalizes significantly better to unseen and rare poses. Relevant files and tools are available at the project website.
Shichao Li 0002, Lei Ke, Kevin Pratama, Yu-Wing Tai, Chi-Keung Tang, Kwang-Ting Cheng
CVPR5
2020 FSS-1000: A 1000-Class Dataset for Few-Shot Segmentation
abstract
Over the past few years, we have witnessed the success of deep learning in image recognition thanks to the availability of large-scale human-annotated datasets such as PASCAL VOC, ImageNet, and COCO. Although these datasets have covered a wide range of object categories, there are still a significant number of objects that are not included. Can we perform the same task without a lot of human annotations? In this paper, we are interested in few-shot object segmentation where the number of annotated training examples are limited to 5 only. To evaluate and validate the performance of our approach, we have built a few-shot segmentation dataset, FSS-1000, which consists of 1000 object classes with pixelwise annotation of ground-truth segmentation. Unique in FSS-1000, our dataset contains significant number of objects that have never been seen or annotated in previous datasets, such as tiny daily objects, merchandise, cartoon characters, logos, etc. We build our baseline model using standard backbone networks such as VGG-16, ResNet-101, and Inception. To our surprise, we found that training our model from scratch using FSS-1000 achieves comparable and even better results than training with weights pre-trained by ImageNet which is more than 100 times larger than FSS-1000. Both our approach and dataset are simple, effective, and easily extensible to learn segmentation of new object classes given very few annotated training examples. Dataset is available at https://github.com/HKUSTCV/FSS-1000.
Tianhan Wei, Yau Pun Chen, Yu-Wing Tai, Chi-Keung Tang
CVPR5
2020 Commonality-Parsing Network Across Shape and Appearance for Partially Supervised Instance Segmentation
Lei Ke, Wenjie Pei, Chi-Keung Tang, Yu-Wing Tai
ECCV (8)4
2020 GSNet: Joint Vehicle Pose and Shape Reconstruction with Geometrical and Scene-Aware Supervision
Lei Ke, Shichao Li 0002, Yanan Sun 0005, Yu-Wing Tai, Chi-Keung Tang
ECCV (15)5
2019 Adversarial Attacks Beyond the Image Space
abstract
Generating adversarial examples is an intriguing problem and an important way of understanding the working mechanism of deep neural networks. Most existing approaches generated perturbations in the image space, i.e., each pixel can be modified independently. However, in this paper we pay special attention to the subset of adversarial examples that correspond to meaningful changes in 3D physical properties (like rotation and translation, illumination condition, etc.). These adversaries arguably pose a more serious concern, as they demonstrate the possibility of causing neural network failure by easy perturbations of real-world 3D objects and scenes. In the contexts of object classification and visual question answering, we augment state-of-the-art deep neural networks that receive 2D input images with a rendering module (either differentiable or not) in front, so that a 3D scene (in the physical space) is rendered into a 2D image (in the image space), and then mapped to a prediction (in the output space). The adversarial perturbations can now go beyond the image space, and have clear meanings in the 3D physical world. Though image-space adversaries can be interpreted as per-pixel albedo change, we verify that they cannot be well explained along these physically meaningful dimensions, which often have a non-local effect. But it is still possible to successfully attack beyond the image space on the physical space, though this is more difficult than image-space attacks, reflected in lower success rates and heavier perturbations required.
Xiaohui Zeng, Chenxi Liu 0001, Yu-Siang Wang, Weichao Qiu, Lingxi Xie, Yu-Wing Tai, Chi-Keung Tang, Alan L. Yuille
CVPR7
2019 LADN: Local Adversarial Disentangling Network for Facial Makeup and De-Makeup
abstract
We propose a local adversarial disentangling network (LADN) for facial makeup and de-makeup. Central to our method are multiple and overlapping local adversarial discriminators in a content-style disentangling network for achieving local detail transfer between facial images, with the use of asymmetric loss functions for dramatic makeup styles with high-frequency details. Existing techniques do not demonstrate or fail to transfer high-frequency details in a global adversarial setting, or train a single local discriminator only to ensure image structure consistency and thus work only for relatively simple styles. Unlike others, our proposed local adversarial discriminators can distinguish whether the generated local image details are consistent with the corresponding regions in the given reference image in cross-image style transfer in an unsupervised setting. Incorporating these technical contributions, we achieve not only state-of-the-art results on conventional styles but also novel results involving complex and dramatic styles with high-frequency details covering large areas across multiple facial features. A carefully designed dataset of unpaired before and after makeup images is released at https://georgegu1997.github.io/LADN-project-page.
Qiao Gu, Guanzhi Wang, Mang Tik Chiu, Yu-Wing Tai, Chi-Keung Tang
ICCV5
2019 Template-Instance Loss for Offline Handwritten Chinese Character Recognition
abstract
The long-standing challenges for offline handwritten Chinese character recognition (HCCR) are twofold: Chinese characters can be very diverse and complicated while similarly looking, and cursive handwriting (due to increased writing speed and infrequent pen lifting) makes strokes and even characters connected together in a flowing manner. In this paper, we propose the template and instance loss functions for the relevant machine learning tasks in offline handwritten Chinese character recognition. First, the character template is designed to deal with the intrinsic similarities among Chinese characters. Second, the instance loss can reduce category variance according to classification difficulty, giving a large penalty to the outlier instance of handwritten Chinese character. Trained with the new loss functions using our deep network architecture HCCR14Layer model consisting of simple layers, our extensive experiments show that it yields state-of-the-art performance and beyond for offline HCCR.
Cewu Lu, Chi-Keung Tang
ICDAR4
2019 Relative CNN-RNN: Learning Relative Atmospheric Visibility From Images
abstract
We propose a deep learning approach for directly estimating relative atmospheric visibility from outdoor photos without relying on weather images or data that require expensive sensing or custom capture. Our data-driven approach capitalizes on a large collection of Internet images to learn rich scene and visibility varieties. The relative CNN-RNN coarse-to-fine model, where CNN stands for convolutional neural network and RNN stands for recurrent neural network, exploits the joint power of relative support vector machine, which has a good ranking representation, and the data-driven deep learning features derived from our novel CNN-RNN model. The CNN-RNN model makes use of shortcut connections to bridge a CNN module and an RNN coarse-to-fine module. The CNN captures the global view while the RNN simulates human's attention shift, namely, from the whole image (global) to the farthest discerned region (local). The learned relative model can be adapted to predict absolute visibility in limited scenarios. Extensive experiments and comparisons are performed to verify our method. We have built an annotated dataset consisting of about 40000 images with 0.2 million human annotations. The large-scale, annotated visibility data set will be made available to accompany this paper.
Yang You 0004, Cewu Lu, Chi-Keung Tang
IEEE Trans. Image Process.4
2018 Beyond Holistic Object Recognition: Enriching Image Understanding With Part States
abstract
Important high-level vision tasks require rich semantic descriptions of objects at part level. Based upon previous work on part localization, in this paper, we address the problem of inferring rich semantics imparted by an object part in still images. Specifically, we propose to tokenize the semantic space as a discrete set of part states. Our modeling of part state is spatially localized, therefore, we formulate the part state inference problem as a pixel-wise annotation problem. An iterative part-state inference neural network that is efficient in time and accurate in performance is specifically designed for this task. Extensive experiments demonstrate that the proposed method can effectively predict the semantic states of parts and simultaneously improve part segmentation, thus benefiting a number of visual understanding applications. The other contribution of this paper is our part state dataset which contains rich part-level semantic annotations.
Cewu Lu, Hao Su 0001, Yong-Lu Li 0001, Yongyi Lu, Li Yi 0001, Chi-Keung Tang, Leonidas J. Guibas
CVPR6
2018 Deep Video Generation, Prediction and Completion of Human Action Sequences
Haoye Cai, Chunyan Bai, Yu-Wing Tai, Chi-Keung Tang
ECCV (2)4
2018 Attribute-Guided Face Generation Using Conditional CycleGAN
Yongyi Lu, Yu-Wing Tai, Chi-Keung Tang
ECCV (12)3
2018 Image Generation from Sketch Constraint Using Contextual GAN
Yongyi Lu, Shangzhe Wu, Yu-Wing Tai, Chi-Keung Tang
ECCV (16)4
2018 Deep High Dynamic Range Imaging with Large Foreground Motions
Shangzhe Wu, Yu-Wing Tai, Chi-Keung Tang
ECCV (2)4
2018 Annotation-Free and One-Shot Learning for Instance Segmentation of Homogeneous Object Clusters
abstract
We propose a novel approach for instance segmentation given an image of homogeneous object cluster (HOC). Our learning approach is one-shot because a single video of an object instance is captured and it requires no human annotation. Our intuition is that images of homogeneous objects can be effectively synthesized based on structure and illumination priors derived from real images. A novel solver is proposed that iteratively maximizes our structured likelihood to generate realistic images of HOC. Illumination transformation scheme is applied to make the real and synthetic images share the same illumination condition. Extensive experiments and comparisons are performed to verify our method. We build a dataset consisting of pixel-level annotated images of HOC. The dataset and code will be released.
Ruiheng Chang, Jiaxu Ma, Cewu Lu, Chi-Keung Tang
IJCAI5
2018 Real-Time Video Stylization Using Object Flows
abstract
We present a real-time video stylization system and demonstrate a variety of painterly styles rendered on real video inputs. The key technical contribution lies on the object flow, which is robust to inaccurate optical flow, unknown object transformation and partial occlusion as well. Since object flows relate regions of the same object across frames, shower-door effect can be effectively reduced where painterly strokes and textures are rendered on video objects. The construction of object flows is performed in real time and automatically after applying metric learning. To reduce temporal flickering, we extend the bilateral filtering into motion bilateral filtering. We propose quantitative metrics to measure the temporal coherence on structures and textures of our stylized videos, and perform extensive experiments to compare our stylized results with baseline systems and prior works specializing in watercolor and abstraction.
Cewu Lu, Chi-Keung Tang
IEEE Trans. Vis. Comput. Graph.3
2017 Online Video Object Detection Using Association LSTM
abstract
Video object detection is a fundamental tool for many applications. Since direct application of image-based object detection cannot leverage the rich temporal information inherent in video data, we advocate to the detection of long-range video object pattern. While the Long Short-Term Memory (LSTM) has been the de facto choice for such detection, currently LSTM cannot fundamentally model object association between consecutive frames. In this paper, we propose the association LSTM to address this fundamental association problem. Association LSTM not only regresses and classifiy directly on object locations and categories but also associates features to represent each output object. By minimizing the matching error between these features, we learn how to associate objects in two consecutive frames. Additionally, our method works in an online manner, which is important for most video tasks. Compared to the traditional video object detection methods, our approach outperforms them on standard video datasets.
Yongyi Lu, Cewu Lu, Chi-Keung Tang
ICCV3
2017 Two-Class Weather Classification
abstract
Given a single outdoor image, we propose a collaborative learning approach using novel weather features to label the image as either sunny or cloudy. Though limited, this two-class classification problem is by no means trivial given the great variety of outdoor images captured by different cameras where the images may have been edited after capture. Our overall weather feature combines the data-driven convolutional neural network (CNN) feature and well-chosen weather-specific features. They work collaboratively within a unified optimization framework that is aware of the presence (or absence) of a given weather cue during learning and classification. In this paper we propose a new data augmentation scheme to substantially enrich the training data, which is used to train a latent SVM framework to make our solution insensitive to global intensity transfer. Extensive experiments are performed to verify our method. Compared with our previous work and the sole use of a CNN classifier, this paper improves the accuracy up to 7-8 percent. Our weather image dataset is available together with the executable of our classifier.
Cewu Lu, Di Lin 0002, Jiaya Jia, Chi-Keung Tang
IEEE Trans. Pattern Anal. Mach. Intell.4
2016 Improving object recognition with the ℓ-channel
Cewu Lu, Efstratios Tsougenis, Chi-Keung Tang
Pattern Recognit.3
2016 3D Navigation on Impossible Figures via Dynamically Reconfigurable Maze
abstract
Previous research on impossible figures focuses extensively on single view modeling and rendering. Existing computer games that employ impossible figures as navigation maze for gaming either use a fixed third-person view with axonometric projection to retain the figure's impossibility perception, or simply break the figure's impossibility upon view changes. In this paper, we present a new approach towards 3D gaming with impossible figures, delivering for the first time navigation in 3D mazes constructed from impossible figures. Such result cannot be achieved by previous research work in modeling impossible figures. To deliver seamless gaming navigation and interaction, we propose i) a set of guiding principles for bringing out subtle perceptions and ii) a novel computational approach to construct 3D structures from impossible figure images and then to dynamically construct the impossible-figure maze subjected to user's view. In the end, we demonstrate and discuss our method with a variety of generic maze types.
Chi-Fu William Lai, Sai-Kit Yeung, Xiaoqi Yan, Chi-Wing Fu, Chi-Keung Tang
IEEE Trans. Vis. Comput. Graph.5
2015 Complexity-adaptive distance metric for object proposals generation
abstract
Distance metric plays a key role in grouping superpixels to produce object proposals for object detection. We observe that existing distance metrics work primarily for low complexity cases. In this paper, we develop a novel distance metric for grouping two superpixels in high-complexity scenarios. Combining them, a complexity-adaptive distance measure is produced that achieves improved grouping in different levels of complexity. Our extensive experimentation shows that our method can achieve good results in the PASCAL VOC 2012 dataset surpassing the latest state-of-the-art methods.
Cewu Lu, Efstratios Tsougenis, Yongyi Lu, Chi-Keung Tang
CVPR5
2015 Square Localization for Efficient and Accurate Object Detection
abstract
The key contribution of this paper is the compact square object localization, which relaxes the exhaustive sliding window from testing all windows of different combinations of aspect ratios. Square object localization is category scalable. By using a binary search strategy, the number of scales to test is further reduced empirically to only O(log(min{H, W})) rounds of sliding CNNs, where H and W are respectively the image height and width. In the training phase, square CNN models and object co-presence priors are learned. In the testing phase, sliding CNN models are applied which produces a set of response maps that can be effectively filtered by the learned co-presence prior to output the final bounding boxes for localizing an object. We performed extensive experimental evaluation on the VOC 2007 and 2012 datasets to demonstrate that while efficient, square localization can output precise bounding boxes to improve the final detection result.
Cewu Lu, Yongyi Lu, Hao Chen 0011, Chi-Keung Tang
ICCV4
2015 Contour Box: Rejecting Object Proposals without Explicit Closed Contours
abstract
Closed contour is an important objectness indicator. We propose a new measure subject to the completeness and tightness constraints, where the optimized closed contour should be tightly bounded within an object proposal. The closed contour measure is defined using closed path integral, and we solve the optimization problem efficiently in polar coordinate system with a global optimum guaranteed. Extensive experiments show that our method can reject a large number of false proposals, and achieve over 6% improvement in object recall at the challenging overlap threshold 0.8 on the VOC 2007 test dataset.
Cewu Lu, Shu Liu 0005, Jiaya Jia, Chi-Keung Tang
ICCV4
2015 Photometric Stereo in the Wild
abstract
Conventional photometric stereo requires to capture images or videos in a dark room to obstruct complex environment light as much as possible. This paper presents a new method that capitalizes on environment light to avail geometry reconstruction, thus bringing photometric stereo to the wild, such as an outdoor scene, with uncontrolled lighting. We do not make restrictive assumption, and only use simple capture equipments, which include a mirror sphere and a video camera. Qualitative and quantitative experiments indicate the potential and practicality of our system to generalize existing frameworks.
Chun Ho Hung, Tai-Pang Wu, Yasuyuki Matsushita, Li Xu 0001, Jiaya Jia, Chi-Keung Tang
WACV6
2015 Normal Estimation of a Transparent Object Using a Video
abstract
Reconstructing transparent objects is a challenging problem. While producing reasonable results for quite complex objects, existing approaches require custom calibration or somewhat expensive labor to achieve high precision. When an overall shape preserving salient and fine details is sufficient, we show in this paper a significant step toward solving the problem when the object's silhouette is available and simple user interaction is allowed, by using a video of a transparent object shot under varying illumination. Specifically, we estimate the normal map of the exterior surface of a given solid transparent object, from which the surface depth can be integrated. Our technical contribution lies in relating this normal estimation problem to one of graph-cut segmentation. Unlike conventional formulations, however, our graph is dual-layered, since we can see a transparent object's foreground as well as the background behind it. Quantitative and qualitative evaluation are performed to verify the efficacy of this practical solution.
Sai-Kit Yeung, Tai-Pang Wu, Chi-Keung Tang, Tony F. Chan, Stanley J. Osher
IEEE Trans. Pattern Anal. Mach. Intell.3
2014 Range-Sample Depth Feature for Action Recognition
abstract
We propose binary range-sample feature in depth. It is based on τ tests and achieves reasonable invariance with respect to possible change in scale, viewpoint, and background. It is robust to occlusion and data corruption as well. The descriptor works in a high speed thanks to its binary property. Working together with standard learning algorithms, the proposed descriptor achieves state-of-the-art results on benchmark datasets in our experiments. Impressively short running time is also yielded.
Cewu Lu, Jiaya Jia, Chi-Keung Tang
CVPR3
2014 Two-Class Weather Classification
abstract
Given a single outdoor image, this paper proposes a collaborative learning approach for labeling it as either sunny or cloudy. Never adequately addressed, this twoclass classification problem is by no means trivial given the great variety of outdoor images. Our weather feature combines special cues after properly encoding them into feature vectors. They then work collaboratively in synergy under a unified optimization framework that is aware of the presence (or absence) of a given weather cue during learning and classification. Extensive experiments and comparisons are performed to verify our method. We build a new weather image dataset consisting of 10K sunny and cloudy images, which is available online together with the executable.
Cewu Lu, Di Lin 0002, Jiaya Jia, Chi-Keung Tang
CVPR4
2014 Shadow Removal from Single RGB-D Images
abstract
We present the first automatic method to remove shadows from single RGB-D images. Using normal cues directly derived from depth, we can remove hard and soft shadows while preserving surface texture and shading. Our key assumption is: pixels with similar normals, spatial locations and chromaticity should have similar colors. A modified nonlocal matching is used to compute a shadow confidence map that localizes well hard shadow boundary, thus handling hard and soft shadows within the same framework. We compare our results produced using state-of-the-art shadow removal on single RGB images, and intrinsic image decomposition on standard RGB-D datasets.
Efstratios Tsougenis, Chi-Keung Tang
CVPR3
2013 Motion-Aware KNN Laplacian for Video Matting
abstract
This paper demonstrates how the nonlocal principle benefits video matting via the KNN Laplacian, which comes with a straightforward implementation using motion-aware K nearest neighbors. In hindsight, the fundamental problem to solve in video matting is to produce spatio-temporally coherent clusters of moving foreground pixels. When used as described, the motion-aware KNN Laplacian is effective in addressing this fundamental problem, as demonstrated by sparse user markups typically on only one frame in a variety of challenging examples featuring ambiguous foreground and background colors, changing topologies with disocclusion, significant illumination changes, fast motion, and motion blur. When working with existing Laplacian-based systems, we expect our Laplacian can benefit them immediately with an improved clustering of moving foreground pixels.
Dingzeyu Li, Qifeng Chen 0001, Chi-Keung Tang
ICCV3
2013 High-quality Kinect depth filtering for real-time 3D telepresence
abstract
3D telepresence is a next-generation multimedia application, offering remote users an immersive and natural video-conferencing environment with real-time 3D graphics. Kinect sensor, a consumer-grade range camera, facilitates the implementation of some recent 3D telepresence systems. However, conventional data filtering methods are insufficient to handle Kinect depth error because such error is quantized rather than just randomly-distributed. Hence, one could often observe large irregularly-shaped patches of pixels that receive the same depth values from Kinect. To enhance visual quality in 3D telepresence, we propose a novel depth data filtering method for Kinect by means of multi-scale and direction-aware support windows. In addition, we develop a GPU-based CUDA implementation that can perform real-time depth filtering. Results from the experiments show that our method can reconstruct hole-free surfaces that are smoother and less bumpy compared to existing methods like bilateral filtering.
Mengyao Zhao, Fuwen Tan, Chi-Wing Fu, Chi-Keung Tang, Jianfei Cai 0001, Tat-Jen Cham
ICME4
2013 KNN Matting
abstract
This paper proposes to apply the nonlocal principle to general alpha matting for the simultaneous extraction of multiple image layers; each layer may have disjoint as well as coherent segments typical of foreground mattes in natural image matting. The estimated alphas also satisfy the summation constraint. As in nonlocal matting, our approach does not assume the local color-line model and does not require sophisticated sampling or learning strategies. On the other hand, our matting method generalizes well to any color or feature space in any dimension, any number of alphas and layers at a pixel beyond two, and comes with an arguably simpler implementation, which we have made publicly available. Our matting technique, aptly called KNN matting, capitalizes on the nonlocal principle by using $(K)$ nearest neighbors (KNN) in matching nonlocal neighborhoods, and contributes a simple and fast algorithm that produces competitive results with sparse user markups. KNN matting has a closed-form solution that can leverage the preconditioned conjugate gradient method to produce an efficient implementation. Experimental evaluation on benchmark datasets indicates that our matting results are comparable to or of higher quality than state-of-the-art methods requiring more involved implementation. In this paper, we take the nonlocal principle beyond alpha estimation and extract overlapping image layers using the same Laplacian framework. Given the alpha value, our closed form solution can be elegantly generalized to solve the multilayer extraction problem. We perform qualitative and quantitative comparisons to demonstrate the accuracy of the extracted image layers.
Qifeng Chen 0001, Dingzeyu Li, Chi-Keung Tang
IEEE Trans. Pattern Anal. Mach. Intell.3
2012 KNN matting
abstract
We are interested in a general alpha matting approach for the simultaneous extraction of multiple image layers; each layer may have disjoint segments for material matting not limited to foreground mattes typical of natural image matting. The estimated alphas also satisfy the summation constraint. Our approach does not assume the local color-line model, does not need sophisticated sampling strategies, and generalizes well to any color or feature space in any dimensions. Our matting technique, aptly called KNN matting, capitalizes on the nonlocal principle by using K nearest neighbors (KNN) in matching nonlocal neighborhoods, and contributes a simple and fast algorithm giving competitive results with sparse user markups. KNN matting has a closed-form solution that can leverage on the preconditioned conjugate gradient method to produce an efficient implementation. Experimental evaluation on benchmark datasets indicates that our matting results are comparable to or of higher quality than state of the art methods.
Qifeng Chen 0001, Dingzeyu Li, Chi-Keung Tang
CVPR3
2012 A Closed-Form Solution to Tensor Voting: Theory and Applications
abstract
We prove a closed-form solution to tensor voting (CFTV): Given a point set in any dimensions, our closed-form solution provides an exact, continuous, and efficient algorithm for computing a structure-aware tensor that simultaneously achieves salient structure detection and outlier attenuation. Using CFTV, we prove the convergence of tensor voting on a Markov random field (MRF), thus termed as MRFTV, where the structure-aware tensor at each input site reaches a stationary state upon convergence in structure propagation. We then embed structure-aware tensor into expectation maximization (EM) for optimizing a single linear structure to achieve efficient and robust parameter estimation. Specifically, our EMTV algorithm optimizes both the tensor and fitting parameters and does not require random sampling consensus typically used in existing robust statistical techniques. We performed quantitative evaluation on its accuracy and robustness, showing that EMTV performs better than the original TV and other state-of-the-art techniques in fundamental matrix estimation for multiview stereo matching. The extensions of CFTV and EMTV for extracting multiple and nonlinear structures are underway.
Tai-Pang Wu, Sai-Kit Yeung, Jiaya Jia, Chi-Keung Tang, Gérard G. Medioni
IEEE Trans. Pattern Anal. Mach. Intell.4
2011 Adequate reconstruction of transparent objects on a shoestring budget
abstract
Reconstructing transparent objects is a challenging problem. While producing reasonable results for quite complex objects, existing approaches require custom calibration or somewhat expensive labor to achieve high precision. On the other hand, when an overall shape preserving salient and fine details is sufficient, we show in this paper a significant step toward solving the problem on a shoestring budget, by using only a video camera, a moving spotlight, and a small chrome sphere. Specifically, the problem we address is to estimate the normal map of the exterior surface of a given solid transparent object, from which the surface depth can be integrated. Our technical contribution lies in relating this normal reconstruction problem to one of graph-cut segmentation. Unlike conventional formulations, however, our graph is dual-layered, since we can see a transparent object's foreground as well as the background behind it. Quantitative and qualitative evaluation are performed to verify the efficacy of this practical solution.
Sai-Kit Yeung, Tai-Pang Wu, Chi-Keung Tang, Tony F. Chan, Stanley J. Osher
CVPR3
2011 Matting and compositing of transparent and refractive objects
abstract
This article introduces a new approach for matting and compositing transparent and refractive objects in photographs. The key to our work is an image-based matting model, termed the Attenuation-Refraction Matte (ARM), that encodes plausible refractive properties of a transparent object along with its observed specularities and transmissive properties. We show that an object's ARM can be extracted directly from a photograph using simple user markup. Once extracted, the ARM is used to paste the object onto a new background with a variety of effects, including compound compositing, Fresnel effect, scene depth, and even caustic shadows. User studies find our results favorable to those obtained with Photoshop as well as perceptually valid in most cases. Our approach allows photo editing of transparent and refractive objects in a manner that produces realistic effects previously only possible via 3D models or environment matting.
Sai-Kit Yeung, Chi-Keung Tang, Michael S. Brown, Sing Bing Kang
ACM Trans. Graph.2
2011 Make it home: automatic optimization of furniture arrangement
abstract
We present a system that automatically synthesizes indoor scenes realistically populated by a variety of furniture objects. Given examples of sensibly furnished indoor scenes, our system extracts, in advance, hierarchical and spatial relationships for various furniture objects, encoding them into priors associated with ergonomic factors, such as visibility and accessibility, which are assembled into a cost function whose optimization yields realistic furniture arrangements. To deal with the prohibitively large search space, the cost function is optimized by simulated annealing using a Metropolis-Hastings state search step. We demonstrate that our system can synthesize multiple realistic furniture arrangements and, through a perceptual study, investigate whether there is a significant difference in the perceived functionality of the automatically synthesized results relative to furniture arrangements produced by human designers.
Lap-Fai Yu, Sai-Kit Yeung, Chi-Keung Tang, Demetri Terzopoulos, Tony F. Chan, Stanley J. Osher
ACM Trans. Graph.3
2010 Quasi-dense 3D reconstruction using tensor-based multiview stereo
abstract
We propose tensor-based multiview stereo (TMVS) for quasi-dense 3D reconstruction from uncalibrated images. Our work is inspired by the patch-based multiview stereo (PMVS), a state-of-the-art technique in multiview stereo reconstruction. The effectiveness of PMVS is attributed to the use of 3D patches in the match-propagate-filter MVS pipeline. Our key observation is: PMVS has not fully utilized the valuable 3D geometric cue available in 3D patches which are oriented points. This paper combines the complementary advantages of photoconsistency, visibility and geometric consistency enforcement in MVS via the use of 3D tensors, where our closed-form solution to tensor voting provides a unified approach to implement the match-propagate-filter pipeline. Using PMVS as the implementation backbone where TMVS is built, we provide qualitative and quantitative evaluation to demonstrate how TMVS significantly improve the MVS pipeline.
Tai-Pang Wu, Sai-Kit Yeung, Jiaya Jia, Chi-Keung Tang
CVPR4
2010 Surface-from-Gradients without Discrete Integrability Enforcement: A Gaussian Kernel Approach
abstract
Representative surface reconstruction algorithms taking a gradient field as input enforce the integrability constraint in a discrete manner. While enforcing integrability allows the subsequent integration to produce surface heights, existing algorithms have one or more of the following disadvantages: They can only handle dense per-pixel gradient fields, smooth out sharp features in a partially integrable field, or produce severe surface distortion in the results. In this paper, we present a method which does not enforce discrete integrability and reconstructs a 3D continuous surface from a gradient or a height field, or a combination of both, which can be dense or sparse. The key to our approach is the use of kernel basis functions, which transfer the continuous surface reconstruction problem into high-dimensional space, where a closed-form solution exists. By using the Gaussian kernel, we can derive a straightforward implementation which is able to produce results better than traditional techniques. In general, an important advantage of our kernel-based method is that the method does not suffer discretization and finite approximation, both of which lead to surface distortion, which is typical of Fourier or wavelet bases widely adopted by previous representative approaches. We perform comparisons with classical and recent methods on benchmark as well as challenging data sets to demonstrate that our method produces accurate surface reconstruction that preserves salient and sharp features. The source code and executable of the system are available for downloading.
Heung-Sun Ng, Tai-Pang Wu, Chi-Keung Tang
IEEE Trans. Pattern Anal. Mach. Intell.3
2010 Photometric Stereo via Expectation Maximization
abstract
This paper presents a robust and automatic approach to photometric stereo, where the two main components, namely surface normals and visible surfaces, are respectively optimized by Expectation Maximization (EM). A dense set of input images is conveniently captured using a digital video camera while a handheld spotlight is being moved around the target object and a small mirror sphere. In our approach, the inherently complex optimization problem is simplified into a two-step optimization, where EM is employed in each step: 1) Using the dense input, the weight or importance of each observation is alternately optimized with the normal and albedo at each pixel and 2) using the optimized normals and employing the Markov Random Fields (MRFs), surface integrabilities and discontinuities are alternately optimized in visible surface reconstruction. Our mathematical derivation gives simple updating rules for the EM algorithms, leading to a stable, practical, and parameter-free implementation that is very robust even in the presence of complex geometry, shadows, highlight, and transparency. We present high-quality results on normal and visible surface reconstruction, where fine geometric details are automatically recovered by our method.
Tai-Pang Wu, Chi-Keung Tang
IEEE Trans. Pattern Anal. Mach. Intell.2
2010 Modeling and rendering of impossible figures
abstract
This article introduces an optimization approach for modeling and rendering impossible figures. Our solution is inspired by how modeling artists construct physical 3D models to produce a valid 2D view of an impossible figure. Given a set of 3D locally possible parts of the figure, our algorithm automatically optimizes a view-dependent 3D model, subject to the necessary 3D constraints for rendering the impossible figure at the desired novel viewpoint. A linear and constrained least-squares solution to the optimization problem is derived, thereby allowing an efficient computation and rendering new views of impossible figures at interactive rates. Once the optimized model is available, a variety of compelling rendering effects can be applied to the impossible figure.
Tai-Pang Wu, Chi-Wing Fu, Sai-Kit Yeung, Jiaya Jia, Chi-Keung Tang
ACM Trans. Graph.5
2009 Fast, automatic and fine-grained tampered JPEG image detection via DCT coefficient analysis
Zhouchen Lin, Junfeng He, Xiaoou Tang, Chi-Keung Tang
Pattern Recognit.4
2009 Noise brush: interactive high quality image-noise separation
abstract
This paper proposes aninteractiveapproach usingjoint image-noise filteringfor achieving high quality image-noise separation. The core of the system is our novel joint image-noise filter which operates in both image and noise domain, and can effectively separate noise from both high and low frequency image structures. A novel user interface is introduced, which allows the user to interact with both the image and the noise layer, and apply the filter adaptively and locally to achieve optimal results. A comprehensive and quantitative evaluation shows that our interactive system can significantly improve the initial image-noise separation results. Our system can also be deployed in various noise-consistent image editing tasks, where preserving the noise characteristics inherent in the input image is a desired feature.
Jia Chen 0026, Chi-Keung Tang, Jue Wang 0001
ACM Trans. Graph.2
2008 Robust dual motion deblurring
abstract
This paper presents a robust algorithm to deblur two consecutively captured blurred photos from camera shaking. Previous dual motion deblurring algorithms succeeded in small and simple motion blur and are very sensitive to noise. We develop a robust feedback algorithm to perform iteratively kernel estimation and image deblurring. In kernel estimation, the stability and capability of the algorithm is greatly improved by incorporating a robust cost function and a set of kernel priors. The robust cost function serves to reject outliers and noise, while kernel priors, including sparseness and continuity, remove ambiguity and maintain kernel shape. In deblurring, we propose a novel and robust approach which takes two blurred images as input to infer the clear image. The deblurred image is then used as feedback to refine kernel estimation. Our method can successfully estimate large and complex motion blurs which cannot be handled by previous dual or single image motion deblurring algorithms. The results are shown to be significantly better than those of previous approaches.
Jia Chen 0026, Lu Yuan 0001, Chi-Keung Tang, Long Quan
CVPR3
2008 Enforcing stochastic inverse consistency in non-rigid image registration and matching
abstract
This paper presents a new method to enforce inverse consistency in nonrigid image registration and matching. Conventional approaches assume diffeomorphic transformation, implicitly or explicitly. However, the inherent smoothness constraint discourages discontinuity consideration. We propose a post-processing algorithm that integrates the input forward and backward fields, which are output by existing registration/matching algorithms, to produce more robust results. Given such a pair of input fields, our algorithm alternately refines the fields by tensor belief propagation, and enforces inverse consistency in stochastic sense by generalized total least squares fitting. To show the efficacy of our stochastic inverse consistency approach, we first present results on very noisy fields. We then demonstrate improvement on existing stereo matching where occlusion is naturally handled by localizing violations of inverse consistency. Finally, we propose a novel application on image stitching, where stochastic inverse consistency is employed in structure deformation, in order to seamlessly align overlapping images with severe misalignment in structure and intensity.
Sai-Kit Yeung, Chi-Keung Tang, Josien P. W. Pluim, Max A. Viergever, Albert C. S. Chung, Helen C. Shen
CVPR2
2008 Extracting smooth and transparent layers from a single image
abstract
Layer decomposition from a single image is an under-constrained problem, because there are more unknowns than equations. This paper studies a slightly easier but very useful alternative where only the background layer has substantial image gradients and structures. We propose to solve this useful alternative by an expectation-maximization (EM) algorithm that employs the hidden markov model (HMM), which maintains spatial coherency of smooth and overlapping layers, and helps to preserve image details of the textured background layer. We demonstrate that, using a small amount of user input, various seemingly unrelated problems in computational photography can be effectively addressed by solving this alternative using our EM-HMM algorithm.
Sai-Kit Yeung, Tai-Pang Wu, Chi-Keung Tang
CVPR3
2008 Limits of Learning-Based Superresolution Algorithms
Zhouchen Lin, Junfeng He, Xiaoou Tang, Chi-Keung Tang
Int. J. Comput. Vis.4
2008 Image Stitching Using Structure Deformation
abstract
The aim of this paper is to achieve seamless image stitching without producing visual artifact caused by severe intensity discrepancy and structure misalignment, given that the input images are roughly aligned or globally registered. Our new approach is based on structure deformation and propagation for achieving overall consistency in image structure and intensity. The new stitching algorithm, which has found applications in image compositing, image blending, and intensity correction,consists of the following main processes. Depending on the compatibility and distinctiveness of the 2-D features detected in the image plane, single or double optimal partitions are computed subject to the constraints of intensity coherence and structure continuity. Afterwards, specific 1-D features are detected along the computed optimal partitions, from which a set of sparse deformation vectors is derived to encode 1-D feature matching between the partitions. These sparse deformation cues are robustly propagated into the input images by solving the associated minimization problem in gradient domain, thus providing a uniform framework for the simultaneous alignment of image structure and intensity. We present results in general image compositing and blending, in order to show the effectiveness of our method in producing seamless stitching results from complex input images.
Jiaya Jia, Chi-Keung Tang
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 Fast image/video upsampling
abstract
We propose a simple but effective upsampling method for automatically enhancing the image/video resolution, while preserving the essential structural information. The main advantage of our method lies in a feedback-control framework which faithfully recovers the high-resolution image information from the input data,withoutimposing additional local structure constraints learned from other examples. This makes our method independent of the quality and number of the selected examples, which are issues typical of learning-based algorithms, while producing high-quality results without observable unsightly artifacts. Another advantage is that our method naturally extends to video upsampling, where the temporal coherence is maintained automatically. Finally, our method runs very fast. We demonstrate the effectiveness of our algorithm by experimenting with different image/video data.
Qi Shan, Zhaorong Li, Jiaya Jia, Chi-Keung Tang
ACM Trans. Graph.4
2008 Texture amendment: reducing texture distortion in constrained parameterization
abstract
Constrained parameterization is an effective way to establish texture coordinates between a 3D surface and an existing image or photograph. A known drawback to constrained parameterization is visual distortion that arises when the 3D geometry is mismatched to highly textured image regions. This paper introduces an approach to reduce visual distortion by expanding image regions via texture synthesis to better fit the 3D geometry. The result is a new amended texture that maintains the essence of the input texture image but exhibits significantly less distortion when mapped onto the 3D model.
Yu-Wing Tai, Michael S. Brown, Chi-Keung Tang, Harry Shum
ACM Trans. Graph.3
2008 Interactive normal reconstruction from a single image
abstract
We present an interactive system for reconstructing surface normals from a single image. Our approach has two complementary contributions. First, we introduce a novel shape-from-shading algorithm (SfS) that produces faithful normal reconstruction for local image region (high-frequency component), but it fails to faithfully recover the overall global structure (low-frequency component). Our second contribution consists of an approach that corrects low-frequency error using a simple markup procedure. This approach, aptly calledrotation palette, allows the user to specify large scale corrections of surface normals by drawing simple stroke correspondences between the normal map and a sphere image which represents rotation directions. Combining these two approaches, we can produce high-quality surfaces quickly from single images.
Tai-Pang Wu, Jian Sun 0001, Chi-Keung Tang, Harry Shum
ACM Trans. Graph.3
2007 Spatio-Temporal Markov Random Field for Video Denoising
abstract
This paper presents a novel spatio-temporal Markov random field (MRF) for video denoising. Two main issues are addressed in this paper, namely, the estimation of noise model and the proper use of motion estimation in the denoising process. Unlike previous algorithms which estimate the level of noise, our method learns the full noise distribution nonparametrically which serves as the likelihood model in the MRF. Instead of using deterministic motion estimation to align pixels, we set up a temporal likelihood by combining a probabilistic motion field with the learned noise model. The prior of this MRF is modeled by piece-wise smoothness. The main advantage of the proposed spatio-temporal MRF is that it integrates spatial and temporal information adaptively into a statistical inference framework, where the posteriori is optimized using graph cuts with alpha expansion. We demonstrate the performance of the proposed approach on benchmark data sets and real videos to show the advantages of our algorithm compared with previous single frame and multi-frame algorithms.
Jia Chen 0026, Chi-Keung Tang
CVPR2
2007 Robust Estimation of Texture Flow via Dense Feature Sampling
abstract
Texture flow estimation is a valuable step in a variety of vision related tasks, including texture analysis, image segmentation, shape-from-texture and texture remapping. This paper describes a novel and effective technique to estimate texture flow in an image given a small example patch. The key idea consists of extracting a dense set of features from the example patch where discrete orientations are encapsulated into the feature vector such that rotation can be simulated as a linear shift of the vector. This dense feature space is then compressed by PCA and clustered using EM to produce a set of small set of principal features. Obtaining these principal features at varying image scales, we can compute the per-pixel scale and orientation likelihoods for the distorted texture. The final texture flow estimation is formulated as the MAP solution of a labeling Markov network which is solved using belief propagation. Experimental results on both synthetic and real images demonstrate good results even for highly distorted examples.
Yu-Wing Tai, Michael S. Brown, Chi-Keung Tang
CVPR3
2007 Limits of Learning-Based Superresolution Algorithms
abstract
Learning-based superresolution (SR) are popular SR techniques that use application dependent priors to infer the missing details in low resolution images (LRIs). However, their performance still deteriorates quickly when the magnification factor is moderately large. This leads us to an important problem: "Do limits of learning-based SR algorithms exist?" In this paper, we attempt to shed some light on this problem when the SR algorithms are designed for general natural images (GNIs). We first define an expected risk for the SR algorithms that is based on the root mean squared error between the superresolved images and the ground truth images. Then utilizing the statistics of GNIs, we derive a closed form estimate of the lower bound of the expected risk. The lower bound can be computed by sampling real images. By computing the curve of the lower bound w.r.t. the magnification factor, we can estimate the limits of learning-based SR algorithms, at which the lower bound of expected risk exceeds a relatively large threshold. We also investigate the sufficient number of samples to guarantee an accurate estimation of the lower bound.
Zhouchen Lin, Junfeng He, Xiaoou Tang, Chi-Keung Tang
ICCV4
2007 Surface-from-Gradients with Incomplete Data for Single View Modeling
abstract
Surface gradients are useful to surface reconstruction in single view modeling, shape-from-shading, and photometric stereo. Previous algorithms minimize a complex, nonlinear energy functional, or require dense surface gradients to perform integration to generate 3D locations, or require user-input heights to constrain the solution space, or produce severe distortion and smooth out surface details. Most single-view algorithms output a Monge patch (height-field), which may introduce further surface distortion along object silhouettes and surface orientation discontinuities. Our proposed algorithm operates on a single view of complete or incomplete data. The data can be gradients without 3D locations, or 3D locations without gradients. The output surface, which is not necessarily a height-field, preserves salient depth and orientation discontinuities. Experimental comparisons on both simple and complex data show that our method produces better surfaces with significantly less distortion and more details preserved. The implementation of our closed-form solution is very straightforward.
Heung-Sun Ng, Tai-Pang Wu, Chi-Keung Tang
ICCV3
2007 Example-Based Cosmetic Transfer
abstract
Cosmetic makeup is used worldwide as a means to enhance beauty and express moods. An art form in its own right, cosmetic styles continuously change and evolve to reflect cultural and societal trends. While countless magazines and books are dedicated to demonstrating cosmetic art, the actual application of makeup still remains a physical endeavor. In this paper, we describe a procedure to apply cosmetic makeup to the image of a person's face with the click of a mouse. Our approach works from before- and-after example images created by professional makeup artists. Using our "cosmetic-transfer" procedure, we can realistically transfer the cosmetic style captured in the example-pair to another person's face. This greatly reduces the time and effort needed to demonstrate a cosmetic style on a new person's face. In addition, our approach can be used to mix-and- match, and even fine-tune, example styles, all virtually, without the need for any physical makeup.
Wai-Shun Tong, Chi-Keung Tang, Michael S. Brown, Ying-Qing Xu
PG2
2007 Soft Color Segmentation and Its Applications
abstract
We propose an automatic approach to soft color segmentation, which produces soft color segments with appropriate amount of overlapping and transparency essential to synthesizing natural images for a wide range of image-based applications. While many state-of-the-art and complex techniques are excellent at partitioning an input image to facilitate deriving a semantic description of the scene, to achieve seamless image synthesis, we advocate to a segmentation approach designed to maintain spatial and color coherence among soft segments while preserving discontinuities, by assigning to each pixel a set of soft labels corresponding to their respective color distributions. We optimize a global objective function which simultaneously exploits the reliability given by global color statistics and flexibility of local image compositing, leading to an image model where the global color statistics of an image is represented by a Gaussian Mixture Model (GMM), while the color of a pixel is explained by a local color mixture model where the weights are defined by the soft labels to the elements of the converged GMM. Transparency is naturally introduced in our probabilistic framework which infers an optimal mixture of colors at an image pixel. To adequately consider global and local information in the same framework, an alternating optimization scheme is proposed to iteratively solve for the global and local model parameters. Our method is fully automatic, and is shown to converge to a good optimal solution. We perform extensive evaluation and comparison, and demonstrate that our method achieves good image synthesis results for image-based applications such as image matting, color transfer, image deblurring, and image colorization.
Yu-Wing Tai, Jiaya Jia, Chi-Keung Tang
IEEE Trans. Pattern Anal. Mach. Intell.3
2007 Natural shadow matting
abstract
This article addresses the problem of natural shadow matting , the removal or extraction of natural shadows from a single image. Because textures are maintained in the shadowless image after the extraction process, our approach produces some of the best results to date among shadow removal techniques. Using the image formation equation typical of computer vision, we advocate a new model for shadow formation where shadow effect is understood as light attenuation instead of a mixture of two colors governed by the conventional matting equation. This leads to a new shadow equation with fewer unknowns to solve, where a three-channel shadow matte and a shadowless image are considered in our optimization. Our problem is formulated as one of energy minimization guided by user-supplied hints in the form of a quadmap which can be specified easily by the user. This formulation allows for robust shadow matte extraction while maintaining texture in the shadowed region by considering color transfer, texture gradient, and shadow smoothness. We demonstrate the usefulness of our approach in shadow removal, image matting, and compositing.
Tai-Pang Wu, Chi-Keung Tang, Michael S. Brown, Harry Shum
ACM Trans. Graph.2
2007 ShapePalettes: interactive normal transfer via sketching
abstract
We present a simple interactive approach to specify 3D shape in a single view using "shape palettes". The interaction is as follows: draw a simple 2D primitive in the 2D view and then specify its 3D orientation by drawing a corresponding primitive on ashape palette. The shape palette is presented as an image of some familiar shape whose local 3D orientation is readily understood and can be easily marked over. The 3D orientation from the shape palette is transferred to the 2D primitive based on the markup. As we will demonstrate, only sparse markup is needed to generate expressive and detailed 3D surfaces. This markup approach can be used to model freehand 3D surfaces drawn in a single view, or combined with image-snapping tools to quickly extract surfaces from images and photographs.
Tai-Pang Wu, Chi-Keung Tang, Michael S. Brown, Harry Shum
ACM Trans. Graph.2
2006 Perceptually-Inspired and Edge-Directed Color Image Super-Resolution
abstract
Inspired by multi-scale tensor voting, a computational framework for perceptual grouping and segmentation, we propose an edge-directed technique for color image superresolution given a single low-resolution color image. Our multi-scale technique combines the advantages of edgedirected, reconstruction-based and learning-based methods, and is unique in two ways. First, we consider simultaneously all the three color channels in our multi-scale tensor voting framework to produce a multi-scale edge representation to guide the process of high-resolution color image reconstruction, which is subject to the back projection constraint. Fine details are inferred without noticeable blurry or ringing artifacts. Second, the inference of highresolution curves is achieved by multi-scale tensor voting, using the dense voting field as an edge-preserving smoothness prior which is derived geometrically without any timeconsuming learning procedure. Qualitative and quantitative results indicate that our method produces convincing results in complex test cases typically used by state-of-theart image super-resolution techniques.
Yu-Wing Tai, Wai-Shun Tong, Chi-Keung Tang
CVPR (2)3
2006 Visible Surface Reconstruction from Normals with Discontinuity Consideration
abstract
Given a dense set of imperfect normals obtained by photometric stereo or shape from shading, this paper presents an optimization algorithm which alternately optimizes until convergence the surface integrabilities and discontinuities inherent in the normal field, in order to derive a segmented surface description of the visible scene without noticeable distortion. In our Expectation-Maximization (EM) framework, we enforce discontinuity-preserving integrability so that fine details are preserved within each output segment while the occlusion boundaries are localized as sharp surface discontinuities. Using the resulting weighted discontinuity map, the estimation of a discontinuity-preserving height field can be formulated into a convex optimization problem. We compare our method and present convincing results on synthetic and real data.
Tai-Pang Wu, Chi-Keung Tang
CVPR (2)2
2006 Dense Photometric Stereo by Expectation Maximization
Tai-Pang Wu, Chi-Keung Tang
ECCV (4)2
2006 Video Repairing under Variable Illumination Using Cyclic Motions
abstract
This paper presents a complete system capable of synthesizing a large number of pixels that are missing due to occlusion or damage in an uncalibrated input video. These missing pixels may correspond to the static background or cyclic motions of the captured scene. Our system employs user-assisted video layer segmentation, while the main processing in video repair is fully automatic. The input video is first decomposed into the color and illumination videos. The necessary temporal consistency is maintained by tensor voting in the spatio-temporal domain. Missing colors and illumination of the background are synthesized by applying image repairing. Finally, the occluded motions are inferred by spatio-temporal alignment of collected samples at multiple scales. We experimented on our system with some difficult examples with variable illumination, where the capturing camera can be stationary or in motion.
Jiaya Jia, Yu-Wing Tai, Tai-Pang Wu, Chi-Keung Tang
IEEE Trans. Pattern Anal. Mach. Intell.4
2006 Dense Photometric Stereo: A Markov Random Field Approach
abstract
We address the problem of robust normal reconstruction by dense photometric stereo, in the presence of complex geometry, shadows, highlight, transparencies, variable attenuation in light intensities, and inaccurate estimation in light directions. The input is a dense set of noisy photometric images, conveniently captured by using a very simple set-up consisting of a digital video camera, a reflective mirror sphere, and a handheld spotlight. We formulate the dense photometric stereo problem as a Markov network and investigate two important inference algorithms for Markov Random Fields (MRFs)--graph cuts and belief propagation--to optimize for the most likely setting for each node in the network. In the graph cut algorithm, the MRF formulation is translated into one of energy minimization. A discontinuity-preserving metric is introduced as the compatibility function, which allows alpha-expansion to efficiently perform the maximum a posteriori (MAP) estimation. Using the identical dense input and the same MRF formulation, our tensor belief propagation algorithm recovers faithful normal directions, preserves underlying discontinuities, improves the normal estimation from one of discrete to continuous, and drastically reduces the storage requirement and running time. Both algorithms produce comparable and very faithful normals for complex scenes. Although the discontinuity-preserving metric in graph cuts permits efficient inference of optimal discrete labels with a theoretical guarantee, our estimation algorithm using tensor belief propagation converges to comparable results, but runs faster because very compact messages are passed and combined. We present very encouraging results on normal reconstruction. A simple algorithm is proposed to reconstruct a surface from a normal map recovered by our method. With the reconstructed surface, an inverse process, known as relighting in computer graphics, is proposed to synthesize novel images of the given scene under user-specified light source and direction. The synthesis is made to run in real time by exploiting the state-of-the-art graphics processing unit (GPU). Our method offers many unique advantages over previous relighting methods and can handle a wide range of novel light sources and directions.
Tai-Pang Wu, Kam-Lun Tang, Chi-Keung Tang, Tien-Tsin Wong
IEEE Trans. Pattern Anal. Mach. Intell.3
2006 Drag-and-drop pasting
abstract
In this paper, we present a user-friendly system for seamless image composition, which we call drag-and-drop pasting. We observe that for Poisson image editing [Perez et al. 2003] to work well, the user must carefully draw a boundary on the source image to indicate the region of interest, such that salient structures in source and target images do not conflict with each other along the boundary. To make Poisson image editing more practical and easy to use, we propose a new objective function to compute an optimized boundary condition. A shortest closed-path algorithm is designed to search for the location of the boundary. Moreover, to faithfully preserve the object's fractional boundary, we construct a blended guidance field to incorporate the object's alpha matte. To use our system, the user needs only to simply outline a region of interest in the source image, and then drag and drop it onto the target image. Experimental results demonstrate the effectiveness of our "drag-and-drop pasting" system.
Jiaya Jia, Jian Sun 0001, Chi-Keung Tang, Harry Shum
ACM Trans. Graph.3
2005 Local Color Transfer via Probabilistic Segmentation by Expectation-Maximization
abstract
We address the problem of regional color transfer between two natural images by probabilistic segmentation. We use a new expectation-maximization (EM) scheme to impose both spatial and color smoothness to infer natural connectivity among pixels. Unlike previous work, our method takes local color information into consideration, and segment image with soft region boundaries for seamless color transfer and compositing. Our modified EM method has two advantages in color manipulation: first, subject to different levels of color smoothness in image space, our algorithm produces an optimal number of regions upon convergence, where the color statistics in each region can be adequately characterized by a component of a Gaussian mixture model (GMM). Second, we allow a pixel to fall in several regions according to our estimated probability distribution in the EM step, resulting in a transparency-like ratio for compositing different regions seamlessly. Hence, natural color transition across regions can be achieved, where the necessary intra-region and inter-region smoothness are enforced without losing original details. We demonstrate results on a variety of applications including image deblurring, enhanced color transfer, and colorizing gray scale images. Comparisons with previous methods are also presented.
Yu-Wing Tai, Jiaya Jia, Chi-Keung Tang
CVPR (1)3
2005 Dense Photometric Stereo Using Tensorial Belief Propagation
abstract
We address the normal reconstruction problem by photometric stereo using a uniform and dense set of photometric images captured at fixed viewpoint. Our method is robust to spurious noises caused by highlight and shadows and non-Lambertian reflections. To simultaneously recover normal orientations and preserve discontinuities, we model the dense photometric stereo problem into two coupled Markov random fields (MRFs): a smooth field for normal orientations, and a spatial line process for normal orientation discontinuities. We propose a very fast tensorial belief propagation method to approximate the maximum a posteriori (MAP) solution of the Markov network. Our tensor-based message passing scheme not only improves the normal orientation estimation from one of discrete to continuous, but also reduces storage and running time drastically. A convenient handheld device was built to collect a scattered set of photometric samples, from which a dense and uniform set on the lighting direction sphere is obtained. We present very encouraging results on a wide range of difficult objects to show the efficacy of our approach.
Kam-Lun Tang, Chi-Keung Tang, Tien-Tsin Wong
CVPR (1)2
2005 A Markov Random Field Approach for Dense Photometric Stereo
abstract
We present a surprisingly simple system that allows for robust normal reconstruction by photometric stereo using a uniform and dense set of photometric images captured at fixed viewpoint, in the presense of spurious noises caused by highlight, shadows and non-Lambertian reflections. Our system consists of a mirror sphere, a spotlight and a DV camera only. Using this, a dense set of unbiased but noisy photometric data that roughly distributed uniformly on the light direction sphere is produced. To simultaneously recover normal orientations and preserve discontinuities, we model the dense photometric stereo problem into two coupled Markov random fields (MRFs): a smooth field for normal orientations, and a spatial line process for normal orientation discontinuities. A very fast tensorial belief propagation method is used to approximate the maximum a posteriori (MAP) solution of the Markov network. We present very encouraging results on a wide range of difficult objects to show the efficacy of our approach.
Kam-Lun Tang, Chi-Keung Tang, Tien-Tsin Wong
CVPR (2)2
2005 Dense Photometric Stereo Using a Mirror Sphere and Graph Cut
abstract
We present a surprisingly simple system that performs robust normal reconstruction by dense photometric stereo, in the presence of large shadows, highlight, transparencies, complex geometry, variable attenuation in light intensity and inaccurate light directions. Our system consists of a mirror sphere, a spotlight and a DV camera only. Using this, we infer a dense set of unbiased but noisy photometric data uniformly distributed on the light direction sphere. We use this dense set to derive a very robust matching cost for our MRF photometric stereo model, where the maximum a posteriori (MAP) solution is estimated. To aggregate support for candidate normals in the normal refinement process, we introduce a compatibility function that is translated into a discontinuity-preserving metric, thus speeding up the MAP estimation by energy minimization using graph cut. No reference object of similar material is used. We perform detailed comparison on our approach with conventional convex minimization. We show very good normals estimated from very noisy data on a wide range of difficult objects to show the robustness and usefulness of our method.
Tai-Pang Wu, Chi-Keung Tang
CVPR (1)2
2005 Eliminating Structure and Intensity Misalignment in Image Stitching
abstract
The aim of this paper is to achieve seamless image stitching for eliminating obvious visual artifact caused by severe intensity discrepancy, image distortion and structure misalignment, given that the input images are globally registered. Our approach is based on structure deformation and propagation while maintaining the overall appearance affinity of the result to the input images. This new approach is proven to be effective in solving the above problems, and has found applications in mosaic deghosting, image blending and intensity correction. Our new method consists of the following main processes. First, salient features or structures are robustly detected and aligned along the optimal partitioning boundary between the input images. From these features, we derive sparse deformation vectors to to uniformly encode the underlying structure and intensity misalignment. These sparse deformation cues will then be propagated robustly and smoothly into the interior of the target image by solving the associated Laplace equations in the image gradient domain. We present convincing results to show that our method can handle significant structure and intensity misalignment in image stitching.
Jiaya Jia, Chi-Keung Tang
ICCV2
2005 A Bayesian Approach for Shadow Extraction from a Single Image
abstract
This paper addresses the problem of shadow extraction from a single image of a complex natural scene. No simplifying assumption on the camera and the light source other than the Lambertian assumption is used. Our method is unique because it is capable of translating very rough user-supplied hints into the effective likelihood and prior functions for our Bayesian optimization. The likelihood function requires a decent estimation of the shadowless image, which is obtained by solving the associated Poisson equation. Our Bayesian framework allows for the optimal extraction of smooth shadows while preserving texture appearance under the extracted shadow. Thus our technique can be applied to shadow removal, producing some best results to date compared with the current state-of-the-art techniques using a single input image. We propose related applications in shadow compositing and image repair using our Bayesian technique.
Tai-Pang Wu, Chi-Keung Tang
ICCV2
2005 Tensor Voting for Image Correction by Global and Local Intensity Alignment
abstract
This paper presents a voting method to perform image correction by global and local intensity alignment. The key to our modeless approach is the estimation of global and local replacement functions by reducing the complex estimation problem to the robust 2D tensor voting in the corresponding voting spaces. No complicated model for replacement function (curve) is assumed. Subject to the monotonic constraint only, we vote for an optimal replacement function by propagating the curve smoothness constraint using a dense tensor field. Our method effectively infers missing curve segments and rejects image outliers. Applications using our tensor voting approach are proposed and described. The first application consists of image mosaicking of static scenes, where the voted replacement functions are used in our iterative registration algorithm for computing the best warping matrix. In the presence of occlusion, our replacement function can be employed to construct a visually acceptable mosaic by detecting occlusion which has large and piecewise constant color. Furthermore, by the simultaneous consideration of color matches and spatial constraints in the voting space, we perform image intensity compensation and high contrast image correction using our voting framework, when only two defective input images are given.
Jiaya Jia, Chi-Keung Tang
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 Robust Estimation of Adaptive Tensors of Curvature by Tensor Voting
abstract
Although curvature estimation from a given mesh or regularly sampled point set is a well-studied problem, it is still challenging when the input consists of a cloud of unstructured points corrupted by misalignment error and outlier noise. Such input is ubiquitous in computer vision. In this paper, we propose a three-pass tensor voting algorithm to robustly estimate curvature tensors, from which accurate principal curvatures and directions can be calculated. Our quantitative estimation is an improvement over the previous two-pass algorithm, where only qualitative curvature estimation (sign of Gaussian curvature) is performed. To overcome misalignment errors, our improved method automatically corrects input point locations at subvoxel precision, which also rejects outliers that are uncorrectable. To adapt to different scales locally, we define the RadiusHit of a curvature tensor to quantify estimation accuracy and applicability. Our curvature estimation algorithm has been proven with detailed quantitative experiments, performing better in a variety of standard error metrics (percentage error in curvature magnitudes, absolute angle difference in curvature direction) in the presence of a large amount of misalignment noise.
Wai-Shun Tong, Chi-Keung Tang
IEEE Trans. Pattern Anal. Mach. Intell.2
2004 Video Repairing: Inference of Foreground and Background under Severe Occlusion
Jiaya Jia, Tai-Pang Wu, Yu-Wing Tai, Chi-Keung Tang
CVPR (1)4
2004 Bayesian Correction of Image Intensity with Spatial Consideration
Jiaya Jia, Jian Sun 0001, Chi-Keung Tang, Harry Shum
ECCV (3)3
2004 Separating Specular, Diffuse, and Subsurface Scattering Reflectances from Photometric Images
Tai-Pang Wu, Chi-Keung Tang
ECCV (2)2
2004 Inference of Segmented Color and Texture Description by Tensor Voting
abstract
A robust synthesis method is proposed to automatically infer missing color and texture information from a damaged 2D image by (N)D tensor voting (N > 3). The same approach is generalized to range and 3D data in the presence of occlusion, missing data and noise. Our method translates texture information into an adaptive (N)D tensor, followed by a voting process that infers noniteratively the optimal color values in the (N)D texture space. A two-step method is proposed. First, we perform segmentation based on insufficient geometry, color, and texture information in the input, and extrapolate partitioning boundaries by either 2D or 3D tensor voting to generate a complete segmentation for the input. Missing colors are synthesized using (N)D tensor voting in each segment. Different feature scales in the input are automatically adapted by our tensor scale analysis. Results on a variety of difficult inputs demonstrate the effectiveness of our tensor voting approach.
Jiaya Jia, Chi-Keung Tang
IEEE Trans. Pattern Anal. Mach. Intell.2
2004 Stereo Reconstruction from Multiperspective Panoramas
Yin Li 0003, Harry Shum, Chi-Keung Tang, Richard Szeliski
IEEE Trans. Pattern Anal. Mach. Intell.3
2004 Simultaneous Two-View Epipolar Geometry Estimation and Motion Segmentation by 4D Tensor Voting
abstract
We address the problem of simultaneous two-view epipolar geometry estimation and motion segmentation from nonstatic scenes. Given a set of noisy image pairs containing matches of n objects, we propose an unconventional, efficient, and robust method, 4D tensor voting, for estimating the unknown n epipolar geometries, and segmenting the static and motion matching pairs into n independent motions. By considering the 4D isotropic and orthogonal joint image space, only two tensor voting passes are needed, and a very high noise to signal ratio (up to five) can be tolerated. Epipolar geometries corresponding to multiple, rigid motions are extracted in succession. Only two uncalibrated frames are needed, and no simplifying assumption (such as affine camera model or homographic model between images) other than the pin-hole camera model is made. Our novel approach consists of propagating a local geometric smoothness constraint in the 4D joint image space, followed by global consistency enforcement for extracting the fundamental matrices corresponding to independent motions. We have performed extensive experiments to compare our method with some representative algorithms to show that better performance on nonstatic scenes are achieved. Results on challenging data sets are presented.
Wai-Shun Tong, Chi-Keung Tang, Gérard G. Medioni
IEEE Trans. Pattern Anal. Mach. Intell.2
2004 First Order Augmentation to Tensor Voting for Boundary Inference and Multiscale Analysis in 3D
abstract
Most computer vision applications require the reliable detection of boundaries. In the presence of outliers, missing data, orientation discontinuities, and occlusion, this problem is particularly challenging. We propose to address it by complementing the tensor voting framework, which was limited to second order properties, with first order representation and voting. First order voting fields and a mechanism to vote for 3D surface and volume boundaries and curve endpoints in 3D are defined. Boundary inference is also useful for a second difficult problem in grouping, namely, automatic scale selection. We propose an algorithm that automatically infers the smallest scale that can preserve the finest details. Our algorithm then proceeds with progressively larger scales to ensure continuity where it has not been achieved. Therefore, the proposed approach does not oversmooth features or delay the handling of boundaries and discontinuities until model misfit occurs. The interaction of smooth features, boundaries, and outliers is accommodated by the unified representation, making possible the perceptual organization of data in curves, surfaces, volumes, and their boundaries simultaneously. We present results on a variety of data sets to show the efficacy of the improved formalism.
Wai-Shun Tong, Chi-Keung Tang, Philippos Mordohai, Gérard G. Medioni
IEEE Trans. Pattern Anal. Mach. Intell.2
2004 Lazy snapping
abstract
In this paper, we present Lazy Snapping , an interactive image cutout tool. Lazy Snapping separates coarse and fine scale processing, making object specification and detailed adjustment easy . Moreover, Lazy Snapping provides instant visual feedback, snapping the cutout contour to the true object boundary efficiently despite the presence of ambiguous or low contrast edges. Instant feedback is made possible by a novel image segmentation algorithm which combines graph cut with pre-computed over-segmentation. A set of intuitive user interface (UI) tools is designed and implemented to provide flexible control and editing for the users. Usability studies indicate that Lazy Snapping provides a better user experience and produces better segmentation results than the state-of-the-art interactive image cutout tool, Magnetic Lasso in Adobe Photoshop.
Yin Li 0003, Jian Sun 0001, Chi-Keung Tang, Harry Shum
ACM Trans. Graph.3
2004 Pop-up light field: An interactive image-based modeling and rendering system
abstract
In this article, we present an image-based modeling and rendering system, which we call pop-up light field , that models a sparse light field using a set of coherent layers . In our system, the user specifies how many coherent layers should be modeled or popped up according to the scene complexity. A coherent layer is defined as a collection of corresponding planar regions in the light field images. A coherent layer can be rendered free of aliasing all by itself, or against other background layers. To construct coherent layers, we introduce a Bayesian approach, coherence matting , to estimate alpha matting around segmented layer boundaries by incorporating a coherence prior in order to maintain coherence across images.We have developed an intuitive and easy-to-use user interface (UI) to facilitate pop-up light field construction. The key to our UI is the concept of human-in-the-loop where the user specifies where aliasing occurs in the rendered image. The user input is reflected in the input light field images where pop-up layers can be modified. The user feedback is instant through a hardware-accelerated real-time pop-up light field renderer. Experimental results demonstrate that our system is capable of rendering anti-aliased novel views from a sparse light field.
Harry Shum, Jian Sun 0001, Shuntaro Yamazaki, Yin Li 0003, Chi-Keung Tang
ACM Trans. Graph.5
2004 Poisson matting
abstract
In this paper, we formulate the problem of natural image matting as one of solving Poisson equations with the matte gradient field. Our approach, which we call Poisson matting , has the following advantages. First, the matte is directly reconstructed from a continuous matte gradient field by solving Poisson equations using boundary information from a user-supplied trimap. Second, by interactively manipulating the matte gradient field using a number of filtering tools, the user can further improve Poisson matting results locally until he or she is satisfied. The modified local result is seamlessly integrated into the final result. Experiments on many complex natural images demonstrate that Poisson matting can generate good matting results that are not possible using existing matting techniques.
Jian Sun 0001, Jiaya Jia, Chi-Keung Tang, Harry Shum
ACM Trans. Graph.3
2004 Binary-Space-Partitioned Images for Resolving Image-Based Visibility
abstract
We propose a novel 2D representation for 3D visibility sorting, the Binary-Space-Partitioned Image (BSPI), to accelerate real-time image-based rendering. BSPI is an efficient 2D realization of a 3D BSP tree, which is commonly used in computer graphics for time-critical visibility sorting. Since the overall structure of a BSP tree is encoded in a BSPI, traversing a BSPI is comparable to traversing the corresponding BSP tree. BSPI performs visibility sorting efficiently and accurately in the 2D image space by warping the reference image triangle-by-triangle instead of pixel-by-pixel. Multiple BSPIs can be combined to solve "disocclusion," when an occluded portion of the scene becomes visible at a novel viewpoint. Our method is highly automatic, including a tensor voting preprocessing step that generates candidate image partition lines for BSPIs, filters the noisy input data by rejecting outliers, and interpolates missing information. Our system has been applied to a variety of real data, including stereo, motion, and range images.
Chi-Wing Fu, Tien-Tsin Wong, Wai-Shun Tong, Chi-Keung Tang, Andrew J. Hanson
IEEE Trans. Vis. Comput. Graph.4
2003 Image Repairing: Robust Image Synthesis by Adaptive ND Tensor Voting
abstract
We present a robust image synthesis method to automatically infer missing information from a damaged 2D image by tensor voting. Our method translates image color and texture information into an adaptive ND tensor, followed by a voting process that infers non-iteratively the optimal color values in the ND texture space for each defective pixel. ND tensor voting can be applied to images consisting of roughly homogeneous and periodic textures (e.g. a brick wall), as well as difficult images of natural scenes, which contain complex color and texture information. To effectively tackle the latter type of difficult images, a two-step method is proposed. First, we perform texture-based segmentation in the input image, and extrapolate partitioning curves to generate a complete segmentation for the image. Then, missing colors are synthesized using ND tensor voting. Automatic tensor scale analysis is used to adapt to different feature scales inherent in the input. We demonstrate the effectiveness of our approach using a difficult set of real images.
Jiaya Jia, Chi-Keung Tang
CVPR (1)2
2003 ROD-TV: Reconstruction on Demand by Tensor Voting
abstract
A "graphics for vision" approach is proposed to address the problem of reconstruction from a large and imperfect data set: reconstruction on demand by tensor voting, or ROD-TV. ROD-TV simultaneously delivers good efficiency and robustness, by adapting to a continuum of primitive connectivity, view dependence, and levels of detail (LOD). Locally inferred surface elements are robust to noise and better capture local shapes. By inferring per-vertex normals at sub-voxel precision on the fly, we can achieve interpolative shading. Since these missing details can be recovered at the current level of detail, our result is not upper bounded by the scanning resolution. By relaxing the mesh connectivity requirement, we extend ROD-TV and propose a simple but effective multiscale feature extraction algorithm. ROD-TV consists of a hierarchical data structure that encodes different levels of detail. The local reconstruction algorithm is tensor voting. It is applied on demand to the visible subset of data at a desired level of detail, by traversing the data hierarchy and collecting tensorial support in a neighborhood. We compare our approach and present encouraging results.
Wai-Shun Tong, Chi-Keung Tang
CVPR (2)2
2003 Rendering driven depth reconstruction
abstract
Previous work on image-based rendering suggests that there is a tradeoff between the number of images and the amount of geometry required for anti-aliased rendering. For instance, plenoptic sampling theory indicates that visually acceptable rendering can be achieved when the input images are undersampled, if sufficient depth information is available for all the pixels. In this paper, we propose a novel vision reconstruction approach, rendering-driven depth recovery, to recover the amount of geometry that is necessary for anti-aliased rendering. Our approach contrasts conventional stereo reconstruction in that we do not intend to accurately reconstruct the depth for each and every single pixel, leading to a very efficient reconstruction algorithm. Our algorithm uses a block-based multi-layer depth representation, and searches in the depth space based on the causality criterion, by detecting double images. Experiments show that rendering systems using our rendering driven depth recovery algorithm can synthesize satisfactory novel views efficiently by using 'just enough geometry' recovered from undersampled input images.
Yin Li 0003, Xin Tong 0001, Chi-Keung Tang, Harry Shum
ICASSP (4)3
2003 Image Registration with Global and Local Luminance Alignment
abstract
Inspired by tensor voting, we present luminance voting, a novel approach for image registration with global and local luminance alignment. The key to our modeless approach is the direct estimation of replacement function, by reducing the complex estimation problem to the robust 2D tensor voting in the corresponding voting spaces. No model for replacement function is assumed. Luminance data are first encoded into 2D ball tensors. Subject to the monotonic constraint only, we vote for an optimal replacement function by propagating the smoothness constraint using a dense tensor field. Our method effectively infers missing curve segments and rejects image outliers without assuming any simplifying or complex curve model. The voted replacement functions are used in our iterative registration algorithm for computing the best warping matrix. Unlike previous approaches, our robust method corrects exposure disparity even if the two overlapping images are initially misaligned. Luminance voting is effective in correcting exposure difference, eliminating vignettes, and thus improving image registration. We present results on a variety of images.
Jiaya Jia, Chi-Keung Tang
ICCV2
2002 Curvature-Augmented Tensor Voting for Shape Inference from Noisy 3D Data
abstract
Improves the basic tensor voting formalism to infer the sign and direction of principal curvatures at each input site from noisy 3D data. Unlike most previous approaches, no local surface fitting, partial derivative computation, nor oriented normal vector recovery is performed in our method. These approaches are known to be noise-sensitive, since accurate partial derivative information is often required, which is usually unavailable from real data. Also, unlike approaches that detect signs of Gaussian curvature, we can handle points with zero Gaussian curvature uniformly, without first localizing them in a separate process. The tensor-voting curvature estimation is non-iterative, does not require initialization, and is robust to a considerable amount of outlier noise, as its effect is reduced by collecting a large number of tensor votes. Qualitative and quantitative results on synthetic and real complex data are presented.
Chi-Keung Tang, Gérard G. Medioni
IEEE Trans. Pattern Anal. Mach. Intell.1
2001 First Order Tensor Voting, and Application to 3-D Scale Analysis
abstract
Many computer vision systems depend on reliable detection of 3D boundaries and regions in order to proceed. In the presence of outliers, missing data and orientation discontinuities due to occlusion, it is difficult to detect boundaries and interpolate data without over-smoothing important feature curves. The authora address these problems by incorporating first order tensor information into the tensor voting formalism, which is second-order based. To propagate an adaptive smoothness constraint at a preferred orientation non-iteratively, we vote for a first order tensor (or vector) to capture polarity and orientation information. To integrate first and second order tensors, we propose an algorithm for inferring the proper scale based on the continuity constraint, and preserving the finest details. Given a noisy 3D point set, the new and improved formalism can better localize boundary curves and orientation discontinuities. Unlike many approaches that over-smooth features, or delay the handling of boundaries and discontinuities until model misfit occurs, the interaction of smooth features, boundaries, discontinuities, outliers are encoded at the representation level. We present results from a variety of datasets to show the efficacy of the improved formalism.
Wai-Shun Tong, Chi-Keung Tang, Gérard G. Medioni
CVPR (1)2
2001 Epipolar Geometry Estimation for Non-Static Scenes by 4D Tensor Voting
abstract
In the presence of false matches and moving objects, image registration is challenging, as outlier rejection, matching and registration become interdependent. We present an efficient and robust method, 4D tensor voting to estimate epipolar geometries for non-static scenes, and identify matching points due to salient and independent motions. Unlike other optimization techniques, data communication in 4D tensor voting does not involve any iterative search. Thus, initialization, local optimum, convergence, and dimensionality of parameter space are not problematic. Like the 8D counterpart, the only assumption we make is the pinhole camera model. Two advancements are made in this work. First, we reduce the dimensionality, and the 4D joint image space is an isotropic and orthogonal one, validating the general assumptions of tensor voting. This improvement is evidenced by the facts that only two passes are needed, and that 4D tensor voting can tolerate an even larger noise/signal ratio (up to a ratio of five). Second, instead of discarding motion pixels as outliers, we successively extract the epipolar geometries contributed by the static background and by the matching points due to salient motions. Only two frames are needed, and no simplifying assumption (such as affine camera model or homographic model between images) is made. Our 4D algorithm consists of two stages: local continuity constraint propagation to remove outliers, and global consistency checking to localize a 4D topological point cone. Results on challenging datasets are presented.
Wai-Shun Tong, Chi-Keung Tang, Gérard G. Medioni
CVPR (1)2
2001 Efficient Dense Depth Estimation from Dense Multiperspective Panoramas
Yin Li 0003, Chi-Keung Tang, Harry Shum
ICCV2
2001 N-Dimensional Tensor Voting and Application to Epipolar Geometry Estimation
abstract
We address the problem of epipolar geometry estimation by formulating it as one of hyperplane inference from a sparse and noisy point set in an 8D space. Given a set of noisy point correspondences in two images of a static scene without correspondences, even in the presence of moving objects, our method extracts good matches and rejects outliers. The methodology is novel and unconventional, since, unlike most other methods optimizing certain scalar, objective functions, our approach does not involve initialization or any iterative search in the parameter space. Therefore, it is free of the problem of local optima or poor convergence. Further, since no search is involved, it is unnecessary to impose simplifying assumption to the scene being analyzed for reducing the search complexity. Subject to the general epipolar constraint only, we detect wrong matches by a computation scheme, 8D tensor voting, which is an instance of the more general N-dimensional tensor voting framework. In essence, the input set of matches is first transformed into a sparse 8D point set. Dense, 8D tensor kernels are then used to vote for the most salient hyperplane that captures all inliers inherent in the input. With this filtered set of matches, the normalized eight-point algorithm can be used to estimate the fundamental matrix accurately. By making use of efficient data structure and locality, our method is both time and space efficient despite the higher dimensionality. We demonstrate the general usefulness of our method using example image pairs for aerial image analysis, with widely different views, and from nonstatic 3D scenes. Each example contains a considerable number of wrong matches.
Chi-Keung Tang, Gérard G. Medioni, Mi-Suen Lee
IEEE Trans. Pattern Anal. Mach. Intell.1
1999 Robust Estimation of Curvature Information from Noisy 3D Data for Shape Description
abstract
We describe an effective and novel approach to infer sign and direction of principal curvatures at each input site from noisy 3D data. Unlike most previous approaches, no local surface fitting, partial derivative computation of any kind, nor oriented normal vector recovery is performed in our method. These approaches are noise-sensitive since accurate, local, partial derivative information is often required, which is usually unavailable from real data because of the unavoidable outlier noise inherent in many measurement phases. Also, we can handle points with zero Gaussian curvature uniformly (i.e., without the need to localize and handle them first as a separate process). Our approach is based on Tensor Voting, a unified, salient structure inference process. Both the sign and the direction of principal curvatures are inferred directly from the input. Each input is first transformed into a synthetic tensor A novel and robust approach based on tensor voting is proposed for curvature information estimation. With faithfully inferred curvature information, each input ellipsoid is aligned with curvature-based dense tensor kernels to produce a dense tensor field. Surfaces and crease curves are extracted from this dense field, by using an extremal feature extraction process. The computation is non-iterative, does not require initialization, and robust to considerable amounts of outlier noise as its effect is reduced by collecting a large number of tensor votes. qualitative and quantitative results on synthetic as well as real and complex data are presented.
Chi-Keung Tang, Gérard G. Medioni
ICCV1
1999 Epipolar Geometry Estimation by Tensor Voting in 8D
abstract
We present a novel, efficient, initialization free approach to the problem of epipolar geometry estimation, by formulating it as one of hyperplane inference from a sparse and noisy point set in an 8D space. Given a set of noisy point correspondences in two images as obtained from two views of a static scene without correspondences, even in the presence of moving objects, our method pulls out inlier matches while rejecting outliers. Unlike most methods which optimize certain objective function, our approach does not involve initialization or any search in the parameter space, and therefore is free of the problem of local optima or poor convergence. Since no search is involved, it is unnecessary to impose simplifying assumption (such as affine camera or local planar homography) to the scene being analyzed for reducing the search complexity. Subject to the general epipolar constraint only, we detect wrong matches by establishing salient "extremalities" via a naval approach, 8D Tensor Voting: the input set of matches is first transformed into a sparse and discrete 8D point set. Dense tensor kernels are then applied to vote for the most salient hyperplane (normal and intercept) that captures all inliers inherent in the input. With this filtered set of matches, the normalized Eight-point Algorithm suffices for the accurate estimation of the fundamental matrix. By using efficient data structure and locality, our method is both time and space efficient despite the higher dimensionality. We demonstrate the general usefulness of our method using example image pairs (i) for aerial image analysis, (ii) with widely different views, and (iii) from non-static 3D scenes (e.g. basketball game in an indoor stadium). Each example contains a considerable amount of wrong matches.
Chi-Keung Tang, Gérard G. Medioni, Mi-Suen Lee
ICCV1
1998 Integrated Surface, Curve and Junction Inference from Sparse 3-D Data Sets
abstract
We are interested in descriptions of 3-D data sets, as obtained from stereo or a 3-D digitizer. We therefore consider as input a sparse set of points, possibly associated with orientation information. In this paper, we address the problem of inferring integrated high-level descriptions such as surfaces, curves, and junctions from a sparse point set. While the method described previously provides excellent results for smooth structures, it only detects discontinuities, but does not localize them. For precise localization, we propose a non-iterative cooperative algorithm in which surfaces, curves, and junctions work together: Initial estimates are computed based on previous results, where each point in the given sparse and possibly noisy point set is convolved with a predefined vector mask to produce dense saliency maps. These maps serve as input to our novel maximal surface and curve marching algorithms for initial surface and curve extraction. Refinement of initial estimates is achieved by hybrid voting using excitatory and inhibitory fields for inferring reliable and natural extension so that surface/curve and curve/junction discontinuities are preserved. Results on several synthetic as well as real data sets are presented.
Chi-Keung Tang, Gérard G. Medioni
ICCV1
1998 Automatic, Accurate Surface Model Inference for Dental CAD/Cam
Chi-Keung Tang, Gérard G. Medioni, François Duret
MICCAI1
1998 Extremal feature extraction from 3-D vector and noisy scalar fields
abstract
We are interested in feature extraction from volume data in terms of coherent surfaces and 3D space curves. The input can be an inaccurate scalar or vector field, sampled densely or sparsely on a regular 3D grid, in which poor resolution and the presence of spurious noisy samples make traditional iso-surface techniques inappropriate. In this paper, we present a general-purpose methodology to extract surfaces or curves from a digital 3D potential vector field {(s,v~)}, in which each voxel holds a scalar s designating the strength and a vector v~ indicating the direction. For scalar, sparse or low-resolution data, we "vectorize" and "densify" the volume by tensor voting to produce dense vector fields that are suitable as input to our algorithms, the extremal surface and curve algorithms. Both algorithms extract, with sub-voxel precision, coherent features representing local extrema in the given vector field. These coherent features are a hole-free triangulation mesh (in the surface case), and a set of connected, oriented and non-intersecting polyline segments (in the curve case). We demonstrate the general usefulness of both extremal algorithms on a variety of real data by properly extracting their inherent extremal properties, such as (a) shock waves induced by abrupt velocity or direction changes in a flow field, (b) interacting vortex cores and vorticity lines in a velocity field, (c) crest-lines and ridges implicit in a digital terrain map, and (d) grooves, anatomical lines and complex surfaces from noisy dental data.
Chi-Keung Tang, Gérard G. Medioni
IEEE Visualization1
1998 Inference of Integrated Surface, Curve, and Junction Descriptions From Sparse 3D Data
abstract
We address the problem of inferring integrated high-level descriptions such as surfaces, 3D curves, and junctions from a sparse point set. For precise localization, we propose a noniterative cooperative algorithm in which surfaces, curves, and junctions work together. Initial estimates are computed based on the work by Guy and Medioni (1997), where each point in the given sparse and possibly noisy point set is convolved with a predefined vector mask to produce dense saliency maps. These maps serve as input to our novel extremal surface and curve algorithms for initial surface and curve extraction. These initial features are refined and integrated by using excitatory and inhibitory fields. Consequently, intersecting surfaces (resp. curves) are fused precisely at their intersection curves (resp. junctions). Results on several synthetic as well as real data sets are presented.
Chi-Keung Tang, Gérard G. Medioni
IEEE Trans. Pattern Anal. Mach. Intell.1
1995 A Fast Algorithm for Computing Optimal Rectilinear Steiner Trees for Extremal Point Sets
Siu-Wing Cheng, Chi-Keung Tang
ISAAC2