EDBT 2026 Demo / reviewers in the wild / expert
Yu-Wing Tai
dblp:40/566
· DBLP profile ↗
172ranked-venue papers
13as first author
61since 2021 · last 2026
0000-0002-3148-0380ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 141 · 11 first-author · 52 since 2021Graphics, computer vision, multimedia, augmented reality and games · 131 · 9 first-author · 40 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RealRep: Generalized SDR-to-HDR Conversion via Attribute-Disentangled Representation LearningabstractHigh-Dynamic-Range Wide-Color-Gamut (HDR-WCG) technology is becoming increasingly widespread, driving a growing need for converting Standard Dynamic Range (SDR) content to HDR. Existing methods primarily rely on fixed tone mapping operators, which struggle to handle the diverse appearances and degradations commonly present in real-world SDR content. To address this limitation, we propose a generalized SDR-to-HDR framework that enhances robustness by learning attribute-disentangled representations. Central to our approach is Realistic Attribute-Disentangled Representation Learning (RealRep), which explicitly disentangles luminance and chrominance components to capture intrinsic content variations across different SDR distributions. Furthermore, we design a Luma-/Chroma-aware negative exemplar generation strategy that constructs degradation-sensitive contrastive pairs, effectively modeling tone discrepancies across SDR styles. Building on these attribute-level priors, we introduce the Degradation-Domain Aware Controlled Mapping Network (DDACMNet), a lightweight, two-stage framework that performs adaptive hierarchical mapping guided by a control-aware normalization mechanism. DDACMNet dynamically modulates the mapping process via degradation-conditioned features, enabling robust adaptation across diverse degradation domains. Extensive experiments demonstrate that RealRep consistently outperforms state-of-the-art methods in both generalization and perceptually faithful HDR color gamut reconstruction. Li Xu 0008, Kepeng Xu, Lin Zhang 0040, Gang He 0002, Yu-Wing Tai |
AAAI | 7 |
| 2025 | Stable Segment Anything ModelabstractThe Segment Anything Model (SAM) achieves remarkable promptable segmentation given high-quality prompts which, however, often require good skills to specify. To make SAM robust to casual prompts, this paper presents the first comprehensive analysis on SAM’s segmentation stability across a diverse spectrum of prompt qualities, notably imprecise bounding boxes and insufficient points. Our key finding reveals that given such low-quality prompts, SAM’s mask decoder tends to activate image features that are biased towards the background or confined to specific object parts. To mitigate this issue, our key idea consists of calibrating solely SAM’s mask attention by adjusting the sampling locations and amplitudes of image features, while the original SAM model architecture and weights remain unchanged. Consequently, our deformable sampling plugin (DSP) enables SAM to adaptively shift attention to the prompted target regions in a data-driven manner. During inference, dynamic routing plugin (DRP) is proposed that toggles SAM between the deformable and regular grid sampling modes, conditioned on the input prompt quality. Thus, our solution, termed Stable-SAM, offers several advantages: 1) improved SAM’s segmentation stability across a wide range of prompt qualities, while 2) retaining SAM’s powerful promptable segmentation efficiency and generality, with 3) minimal learnable parameters (0.08 M) and fast adaptation. Extensive experiments validate the effectiveness and advantages of our approach, underscoring Stable-SAM as a more robust solution for segmenting anything. Codes are at https://github.com/fanq15/Stable-SAM. Xin Tao 0001, Lei Ke, Mingqiao Ye, Di Zhang 0026, Pengfei Wan 0001, Yu-Wing Tai, Chi-Keung Tang |
ICLR | 7 |
| 2025 | Motion-Agent: A Conversational Framework for Human Motion Generation with LLMsabstractWhile previous approaches to 3D human motion generation have achieved notable success, they often rely on extensive training and are limited to specific tasks. To address these challenges, we introduce **Motion-Agent**, an efficient conversational framework designed for general human motion generation, editing, and understanding.
Motion-Agent employs an open-source pre-trained language model to develop a generative agent, **MotionLLM**, that bridges the gap between motion and text. This is accomplished by encoding and quantizing motions into discrete tokens that align with the language model's vocabulary. With only 1-3% of the model's parameters fine-tuned using adapters, MotionLLM delivers performance on par with diffusion models and other transformer-based methods trained from scratch. By integrating MotionLLM with GPT-4 without additional training, Motion-Agent is able to generate highly complex motion sequences through multi-turn conversations, a capability that previous models have struggled to achieve.
Motion-Agent supports a wide range of motion-language tasks, offering versatile capabilities for generating and customizing human motion through interactive conversational exchanges. Xinhang Liu, Yu-Wing Tai, Chi-Keung Tang |
ICLR | 5 |
| 2025 | LayerCraft: Enhancing Text-to-Image Generation with CoT Reasoning and Layered Object IntegrationabstractText-to-image (T2I) generation has made remarkable progress, yet existing systems still lack intuitive control over spatial composition, object consistency, and multi-step editing. We present **LayerCraft**, a modular framework that uses large language models (LLMs) as autonomous agents to orchestrate structured, layered image generation and editing. LayerCraft supports two key capabilities: (1) *structured generation* from simple prompts via chain-of-thought (CoT) reasoning, enabling it to decompose scenes, reason about object placement, and guide composition in a controllable, interpretable manner; and (2) *layered object integration*, allowing users to insert and customize objects---such as characters or props---across diverse images or scenes while preserving identity, context, and style. The system comprises a coordinator agent, the **ChainArchitect** for CoT-driven layout planning, and the **Object Integration Network (OIN)** for seamless image editing using off-the-shelf T2I models without retraining. Through applications like batch collage editing and narrative scene generation, LayerCraft empowers non-experts to iteratively design, customize, and refine visual content with minimal manual effort. Code will be released upon acceptance. Yu-Wing Tai |
NeurIPS | 3 |
| 2025 | Segment Anything Meets Point TrackingabstractFoundation models have marked a significant stride to-ward addressing generalization challenges in deep learning. While the Segment Anything Model (SAM) has established a strong foothold in image segmentation, existing video segmentation methods still require extensive mask labeling for fine-tuning, or face performance drops on unseen data domains otherwise. In this paper, we show how foundation models for image segmentation make a step toward enhancing domain generalizability in video segmentation. We discover that, combined with long-term point tracking, image segmentation models yield state-of-the-art results in zero-shot video segmentation across multiple benchmarks. Surprisingly, point trackers exhibit generalization to domains beyond their synthetic pre-training sequences, which we attribute to the trackers' ability to harness the rich local information in the vicinity of each tracked point. Thus, we introduce SAM-PT, an innovative method for point-centric video segmentation, leveraging the capabilities of SAM alongside long-term point tracking. SAM-PT extends SAM's capability to tracking and segmenting anything in dynamic videos. Unlike traditional video segmentation methods that focus on object-centric mask propagation, our approach uniquely exploits point propagation to utilize local structure information independent of object semantics. The effectiveness of point-based tracking is underscored by direct evaluation on the zero-shot open-world UVO benchmark. Our experiments on popular video object segmentation and multi-object segmentation tracking benchmarks, including DAVIS, YouTube-VOS, and BDD100K, suggest that a point-based segmentation tracker yields better zero-shot performance and efficient interactions. We release our code at https://github.com/SysCV/sam-pt. Frano Rajic, Lei Ke, Yu-Wing Tai, Chi-Keung Tang, Martin Danelljan, Fisher Yu 0001 |
WACV | 3 |
| 2025 | Robust Object Detection with Domain-Invariant Training and Continual Test-Time Adaptation
Mattia Segù, Bernt Schiele, Dengxin Dai, Yu-Wing Tai, Chi-Keung Tang |
Int. J. Comput. Vis. | 5 |
| 2024 | SANeRF-HQ: Segment Anything for NeRF in High QualityabstractRecently, the Segment Anything Model (SAM) has showcased remarkable capabilities of zero-shot segmentation, while NeRF (Neural Radiance Fields) has gained popularity as a method for various 3D problems beyond novel view synthesis. Though there exist initial attempts to incorporate these two methods into 3D segmentation, they face the challenge of accurately and consistently segmenting objects in complex scenarios. In this paper, we introduce the Segment Anything for NeRF in High Quality (SANeRF-HQ) to achieve high-quality 3D segmentation of any target object in a given scene. SANeRF-HQ utilizes SAM for open-world object segmentation guided by user-supplied prompts, while leveraging NeRF to aggregate information from different viewpoints. To overcome the aforementioned challenges, we employ density field and RGB similarity to enhance the accuracy of segmentation boundary during the aggregation. Emphasizing on segmentation accuracy, we evaluate our method on multiple NeRF datasets where high-quality ground-truths are available or manually annotated. SANeRF-HQ shows a significant quality improvement over state-of-the-art methods in NeRF object segmentation, provides higher flexibility for object localization, and enables more consistent object segmentation across multiple views. Results and code are available at the project site: https://lyclyc52.github.io/SANeRF-HQ/. Yichen Liu 0001, Benran Hu 0001, Chi-Keung Tang, Yu-Wing Tai |
CVPR | 4 |
| 2024 | Gear-NeRF: Free-Viewpoint Rendering and Tracking with Motion-Aware Spatio-Temporal SamplingabstractExtensions of Neural Radiance Fields (NeRFs) to model dynamic scenes have enabled their near photo-realistic, free-viewpoint rendering. Although these methods have shown some potential in creating immersive experiences, two drawbacks limit their ubiquity: ( i) a significant reduction in reconstruction quality when the computing budget is limited, and (ii) a lack of semantic understanding of the underlying scenes. To address these issues, we introduce Gear-NeRF, which leverages semantic information from powerful image segmentation models. Our approach presents a principled way for learning a spatio-temporal (4D) semantic embedding, based on which we introduce the concept of gears to allow for stratified modeling of dynamic regions of the scene based on the extent of their motion. Such differentiation allows us to adjust the spatio- temporal sampling resolution for each region in proportion to its motion scale, achieving more photo-realistic dynamic novel view synthesis. At the same time, almost for free, our approach enables free-viewpoint tracking of objects of interest - a functionality not yet achieved by existing NeRF-based methods. Empirical studies validate the effectiveness of our method, where we achieve state-of-the-art rendering and tracking performance on multiple challenging datasets. The project page is available at: https://merl.com/research/highlights/gear-nerf Xinhang Liu, Yu-Wing Tai, Chi-Keung Tang, Pedro Miraldo, Suhas Lohit, Moitreya Chatterjee |
CVPR | 2 |
| 2024 | C3Net: Compound Conditioned ControlNet for Multimodal Content GenerationabstractWe present Compound Conditioned ControlNet, C3Net, a novel generative neural architecture taking conditions from multiple modalities and synthesizing multimodal contents simultaneously (e.g., image, text, audio). C3Net adapts the ControlNet [46] architecture to jointly train and make inferences on a production-ready diffusion model and its trainable copies. Specifically, C3Net first aligns the conditions from multimodalities to the same semantic latent space using modality-specific encoders based on contrastive training. Then, it generates multimodal outputs based on the aligned latent space, whose semantic information is combined using a ControlNet-like architecture called Control C3-UNet. Correspondingly, with this system design, our model offers an improved solution for joint-modality generation through learning and explaining multimodal conditions, involving more than just linear interpolation within the latent space. Meanwhile, as we align conditions to a unified latent space, C3Net only requires one trainable Control C3-UNet to work on multimodal semantic information. Furthermore, our model employs uni-modal pretraining on the condition alignment stage, outperforming the non-pretrained alignment even on relatively scarce training data and thus demonstrating high-quality compound condition generation. We contribute the first high-quality tri-modal validation set to validate quantitatively that C3Net outperforms or is on par with the first and contemporary state-of-the-art multimodal generation [43]. Our codes and tri-modal dataset will be released here. Yuehuai Liu, Yu-Wing Tai, Chi-Keung Tang |
CVPR | 3 |
| 2024 | Boundary-Enhanced Instance SegmentationabstractDespite significant progress in instance segmentation, recent solutions still fall short of boundary accuracy especially for overlapping instances of the same category. In this paper, we propose a novel boundary-enhanced instance segmentation (BEIS) framework that explicitly models the feature relationships across object boundaries for high-quality instance segmentation. Specifically, BEIS generates boundary-enhanced features using both intra-mask and cross-image boundary discrimination learning. The intra-mask boundary discrimination learning (IBDL) employs pixel-level discrimination learning to disentangle pixel representations along boundaries. The cross-image boundary discrimination learning (CBDL) learns a boundary-aware feature bank from training data to further boost the performance. Thus, CBDL can take advantage of boundary relations across images to enhance the quality of segmented boundaries. To focus on hard-to-segment boundaries, we propose an adaptive sampling strategy to automatically construct discriminative pairs in regions with high possibilities of confusion. Extensive experiments show BEIS outperforms on various datasets. Tianxiang Pan, Yu-Wing Tai, Bin Wang 0021 |
ECAI | 3 |
| 2024 | DragVideo: Interactive Drag-Style Video Editing
Yufan Deng, Ruida Wang, Yu-Wing Tai, Chi-Keung Tang |
ECCV (56) | 4 |
| 2024 | Deceptive-NeRF/3DGS: Diffusion-Generated Pseudo-observations for High-Quality Sparse-View Reconstruction
Xinhang Liu, Jiaben Chen, Shiu-Hong Kao, Yu-Wing Tai, Chi-Keung Tang |
ECCV (16) | 4 |
| 2024 | Distill Gold from Massive Ores: Bi-level Data Pruning Towards Efficient Dataset Distillation
Yong-Lu Li 0001, Kaitong Cui, Ziyu Wang 0010, Cewu Lu, Yu-Wing Tai, Chi-Keung Tang |
ECCV (20) | 6 |
| 2024 | ChatCam: Empowering Camera Control through Conversational AIabstractCinematographers adeptly capture the essence of the world, crafting compelling visual narratives through intricate camera movements. Witnessing the strides made by large language models in perceiving and interacting with the 3D world, this study explores their capability to control cameras with human language guidance. We introduce ChatCam, a system that navigates camera movements through conversations with users, mimicking a professional cinematographer's workflow. To achieve this, we propose CineGPT, a GPT-based autoregressive model for text-conditioned camera trajectory generation. We also develop an Anchor Determinator to ensure precise camera trajectory placement. ChatCam understands user requests and employs our proposed tools to generate trajectories, which can be used to render high-quality video footage on radiance field representations. Our experiments, including comparisons to state-of-the-art approaches and user studies, demonstrate our approach's ability to interpret and execute complex instructions for camera operation, showing promising applications in real-world production settings. Project page: https://xinhangliu.com/chatcam. Xinhang Liu, Yu-Wing Tai, Chi-Keung Tang |
NeurIPS | 2 |
| 2024 | Semantic Image Matting: General and Specific Semantics
Yanan Sun 0005, Chi-Keung Tang, Yu-Wing Tai |
Int. J. Comput. Vis. | 3 |
| 2024 | FSODv2: A Deep Calibrated Few-Shot Object Detection Network
Wei Zhuo 0005, Chi-Keung Tang, Yu-Wing Tai |
Int. J. Comput. Vis. | 4 |
| 2023 | Ultrahigh Resolution Image/Video Matting with Spatio-Temporal SparsityabstractCommodity ultrahigh definition (UHD) displays are becoming more affordable which demand imaging in ultrahigh resolution (UHR). This paper proposes SparseMat, a computationally efficient approach for UHR image/video matting. Note that it is infeasible to directly process UHR images at full resolution in one shot using existing matting algorithms without running out of memory on consumer-level computational platforms, e.g., Nvidia 1080Ti with 11G memory, while patch-based approaches can introduce unsightly artifacts due to patch partitioning. Instead, our method resorts to spatial and temporal sparsity for addressing general UHR matting. When processing videos, huge computation redundancy can be reduced by exploiting spatial and temporal sparsity. In this paper, we show how to effectively detect spatio-temporal sparsity, which serves as a gate to activate input pixels for the matting model. Under the guidance of such sparsity, our method with sparse high-resolution module (SHM) can avoid patch-based inference while memory efficient for full-resolution matte refinement. Extensive experiments demonstrate that SparseMat can effectively and efficiently generate high-quality alpha matte for UHR images and videos at the original high resolution in a single pass. Project page is in https://github.com/nowsyn/SparseMat.git. Yanan Sun 0005, Chi-Keung Tang, Yu-Wing Tai |
CVPR | 3 |
| 2023 | NeRF-RPN: A general framework for object detection in NeRFsabstractThis paper presents the first significant object detection framework, NeRF-RPN, which directly operates on NeRF. Given a pre-trained NeRF model, NeRF-RPN aims to detect all bounding boxes of objects in a scene. By exploiting a novel voxel representation that incorporates multi-scale 3D neural volumetric features, we demonstrate it is possible to regress the 3D bounding boxes of objects in NeRF directly without rendering the NeRF at any viewpoint. NeRF-RPN is a general framework and can be applied to detect objects without class labels. We experimented NeRF-RPN with various backbone architectures, RPN head designs and loss functions. All of them can be trained in an end-to-end manner to estimate high quality 3D bounding boxes. To facilitate future research in object detection for NeRF, we built a new benchmark dataset which consists of both synthetic and real-world data with careful labeling and clean up. Code and dataset are available at htt ps: //github.com/lyclyc52/NeRF_RPN. Benran Hu 0001, Junkai Huang 0004, Yichen Liu 0001, Yu-Wing Tai, Chi-Keung Tang |
CVPR | 4 |
| 2023 | Mask-Free Video Instance SegmentationabstractThe recent advancement in Video Instance Segmentation (VIS) has largely been driven by the use of deeper and increasingly data-hungry transformer-based models. However, video masks are tedious and expensive to annotate, limiting the scale and diversity of existing VIS datasets. In this work, we aim to remove the mask-annotation requirement. We propose MaskFreeVIS, achieving highly competitive VIS performance, while only using bounding box annotations for the object state. We leverage the rich temporal mask consistency constraints in videos by introducing the Temporal KNN-patch Loss (TK-Loss), providing strong mask supervision without any labels. Our TK-Loss finds one-to-many matches across frames, through an efficient patch-matching step followed by a K-nearest neighbor selection. A consistency loss is then enforced on the found matches. Our mask-free objective is simple to implement, has no trainable parameters, is computationally efficient, yet outperforms baselines employing, e.g., state-of-the-art optical flow to enforce temporal mask consistency. We validate MaskFreeVIS on the YouTube-VIS 2019/2021, OVIS and BDD100K MOTS benchmarks. The results clearly demonstrate the efficacy of our method by drastically narrowing the gap between fully and weakly-supervised VIS performance. Our code and trained models are available at http://vis.xyz/pub/maskfreevis. Lei Ke, Martin Danelljan, Henghui Ding, Yu-Wing Tai, Chi-Keung Tang, Fisher Yu 0001 |
CVPR | 4 |
| 2023 | Compression-Aware Video Super-ResolutionabstractVideos stored on mobile devices or delivered on the Internet are usually in compressed format and are of various unknown compression parameters, but most video super-resolution (VSR) methods often assume ideal inputs resulting in large performance gap between experimental settings and real-world applications. In spite of a few pioneering works being proposed recently to super-resolve the compressed videos, they are not specially designed to deal with videos of various levels of compression. In this paper, we propose a novel and practical compression-aware video super-resolution model, which could adapt its video enhancement process to the estimated compression level. A compression encoder is designed to model compression levels of input frames, and a base VSR model is then conditioned on the implicitly computed representation by inserting compression-aware modules. In addition, we propose to further strengthen the VSR model by taking full advantage of meta data that is embedded naturally in compressed video streams in the procedure of information fusion. Extensive experiments are conducted to demonstrate the effectiveness and efficiency of the proposed method on compressed VSR benchmarks. The codes will be available at https://github.com/aprBlue/CAVSR Takashi Isobe, Xu Jia 0012, Xin Tao 0001, Huchuan Lu, Yu-Wing Tai |
CVPR | 6 |
| 2023 | Instance Neural Radiance FieldabstractThis paper presents one of the first learning-based NeRF 3D instance segmentation pipelines, dubbed as Instance Neural Radiance Field, or Instance-NeRF. Taking a NeRF pretrained from multi-view RGB images as input, Instance-NeRF can learn 3D instance segmentation of a given scene, represented as an instance field component of the NeRF model. To this end, we adopt a 3D proposal-based mask prediction network on the sampled volumetric features from NeRF, which generates discrete 3D instance masks. The coarse 3D mask prediction is then projected to image space to match 2D segmentation masks from different views generated by existing panoptic segmentation models, which are used to supervise the training of the instance field. Notably, beyond generating consistent 2D segmentation maps from novel views, Instance-NeRF can query instance information at any 3D point, which greatly enhances NeRF object segmentation and manipulation. Our method is also one of the first to achieve such results in pure inference. Experimented on synthetic and real-world NeRF datasets with complex indoor scenes, Instance-NeRF surpasses previous NeRF segmentation works and competitive 2D segmentation methods in segmentation performance on unseen views. Code and data are available at https://github.com/lyclyc52/Instance_NeRF. Yichen Liu 0001, Benran Hu 0001, Junkai Huang 0004, Yu-Wing Tai, Chi-Keung Tang |
ICCV | 4 |
| 2023 | EgoPCA: A New Framework for Egocentric Hand-Object Interaction UnderstandingabstractWith the surge in attention to Egocentric Hand-Object Interaction (Ego-HOI), large-scale datasets such as Ego4D and EPIC-KITCHENS have been proposed. However, most current research is built on resources derived from third-person video action recognition. This inherent domain gap between first- and third-person action videos, which have not been adequately addressed before, makes current Ego-HOI suboptimal. This paper rethinks and proposes a new framework as an infrastructure to advance Ego-HOI recognition by Probing, Curation and Adaption (EgoPCA). We contribute comprehensive pre-train sets, balanced test sets and a new baseline, which are complete with a training-finetuning strategy. With our new framework, we not only achieve state-of-the-art performance on Ego-HOI benchmarks but also build several new and effective mechanisms and settings to advance further research. We believe our data and the findings will pave a new way for Ego-HOI understanding. Code and data are available at https://mvig-rhos.com/ego_pca. Yong-Lu Li 0001, Zhemin Huang 0001, Michael Xu Liu, Cewu Lu, Yu-Wing Tai, Chi-Keung Tang |
ICCV | 6 |
| 2023 | Cascade-DETR: Delving into High-Quality Universal Object DetectionabstractObject localization in general environments is a fundamental part of vision systems. While dominating on the COCO benchmark, recent Transformer-based detection methods are not competitive in diverse domains. Moreover, these methods still struggle to very accurately estimate the object bounding boxes in complex environments.We introduce Cascade-DETR for high-quality universal object detection. We jointly tackle the generalization to diverse domains and localization accuracy by proposing the Cascade Attention layer, which explicitly integrates object-centric information into the detection decoder by limiting the attention to the previous box prediction. To further enhance accuracy, we also revisit the scoring of queries. Instead of relying on classification scores, we predict the expected IoU of the query, leading to substantially more well-calibrated confidences. Lastly, we introduce a universal object detection benchmark, UDB10, that contains 10 datasets from diverse domains. While also advancing the state-of-the-art on COCO, Cascade-DETR substantially improves DETR-based detectors on all datasets in UDB10, even by over 10 mAP in some cases. The improvements under stringent quality requirements are even more pronounced. Our code and pretrained models are at https://github.com/SysCV/cascade-detr. Mingqiao Ye, Lei Ke, Siyuan Li 0008, Yu-Wing Tai, Chi-Keung Tang, Martin Danelljan, Fisher Yu 0001 |
ICCV | 4 |
| 2023 | Towards Robust Object Detection Invariant to Real-World Domain Shifts
Mattia Segù, Yu-Wing Tai, Fisher Yu 0001, Chi-Keung Tang, Bernt Schiele, Dengxin Dai |
ICLR | 3 |
| 2023 | Scene-Generalizable Interactive Segmentation of Radiance FieldsabstractExisting methods for interactive segmentation in radiance fields entail scene-specific optimization and thus cannot generalize across different scenes, which greatly limits their applicability. In this work we make the first attempt at Scene-Generalizable Interactive Segmentation in Radiance Fields (SGISRF) and propose a novel SGISRF method, which can perform 3D object segmentation for novel (unseen) scenes represented by radiance fields, guided by only a few interactive user clicks in a given set of multi-view 2D images. In particular, the proposed SGISRF focuses on addressing three crucial challenges with three specially designed techniques. First, we devise the Cross-Dimension Guidance Propagation to encode the scarce 2D user clicks into informative 3D guidance representations. Second, the Uncertainty-Eliminated 3D Segmentation module is designed to achieve efficient yet effective 3D segmentation. Third, Concealment-Revealed Supervised Learning scheme is proposed to reveal and correct the concealed 3D segmentation errors resulted from the supervision in 2D space with only 2D mask annotations. Extensive experiments on two real-world challenging benchmarks covering diverse scenes demonstrate 1) effectiveness and scene-generalizability of the proposed method, 2) favorable performance compared to classical method requiring scene-specific optimization. Songlin Tang, Wenjie Pei, Xin Tao 0001, Tanghui Jia, Guangming Lu 0002, Yu-Wing Tai |
ACM Multimedia | 6 |
| 2023 | Feature Decoupling-Recycling Network for Fast Interactive SegmentationabstractRecent interactive segmentation methods iteratively take source image, user guidance and previously predicted mask as the input without considering the invariant nature of the source image. As a result, the process of extracting features from the source image is repeated in each interaction, resulting in substantial computational redundancy. In this work, we propose the Feature Decoupling-Recycling Network (FDRN), which decouples the modeling components based on their intrinsic discrepancies and then recycles components that can be reused for each user interaction. Thus, the efficiency of the whole interactive process can be significantly improved. To be specific, we apply the Decoupling-Recycling strategy from three perspectives to address three types of discrepancies, respectively. First, our model decouples the learning of source image semantics from the encoding of user guidance to process two types of input domains separately. Second, FDRN decouples high-level and low-level features from stratified semantic representations to enhance feature learning. Third, during the encoding of user guidance, current user guidance is decoupled from historical guidance to highlight the effect of current user guidance. We conduct extensive experiments on 6 datasets from different domains and modalities, which demonstrate the following merits of our model: 1) superior efficiency than other methods, particularly advantageous in the challenging scenarios requiring long-term interactions (up to 4.25x faster), while achieving favorable segmentation performance; 2) strong applicability to various methods serving as a universal enhancement technique; 3) well cross-task generalizability, e.g., to medical image segmentation, and robustness against misleading user guidance. Weinong Wang, Xin Tao 0001, Zhiwei Xiong, Yu-Wing Tai, Wenjie Pei |
ACM Multimedia | 5 |
| 2023 | Segment Anything in High QualityabstractThe recent Segment Anything Model (SAM) represents a big leap in scaling up segmentation models, allowing for powerful zero-shot capabilities and flexible prompting. Despite being trained with 1.1 billion masks, SAM's mask prediction quality falls short in many cases, particularly when dealing with objects that have intricate structures. We propose HQ-SAM, equipping SAM with the ability to accurately segment any object, while maintaining SAM's original promptable design, efficiency, and zero-shot generalizability. Our careful design reuses and preserves the pre-trained model weights of SAM, while only introducing minimal additional parameters and computation. We design a learnable High-Quality Output Token, which is injected into SAM's mask decoder and is responsible for predicting the high-quality mask. Instead of only applying it on mask-decoder features, we first fuse them with early and final ViT features for improved mask details. To train our introduced learnable parameters, we compose a dataset of 44K fine-grained masks from several sources. HQ-SAM is only trained on the introduced detaset of 44k masks, which takes only 4 hours on 8 GPUs. We show the efficacy of HQ-SAM in a suite of 10 diverse segmentation datasets across different downstream tasks, where 8 out of them are evaluated in a zero-shot transfer protocol. Our code and pretrained models are at https://github.com/SysCV/SAM-HQ. Lei Ke, Mingqiao Ye, Martin Danelljan, Yifan Liu 0001, Yu-Wing Tai, Chi-Keung Tang, Fisher Yu 0001 |
NeurIPS | 5 |
| 2023 | BiMatting: Efficient Video Matting via BinarizationabstractReal-time video matting on edge devices faces significant computational resource constraints, limiting the widespread use of video matting in applications such as online conferences and short-form video production. Binarization is a powerful compression approach that greatly reduces computation and memory consumption by using 1-bit parameters and bitwise operations. However, binarization of the video matting model is not a straightforward process, and our empirical analysis has revealed two primary bottlenecks: severe representation degradation of the encoder and massive redundant computations of the decoder. To address these issues, we propose BiMatting, an accurate and efficient video matting model using binarization. Specifically, we construct shrinkable and dense topologies of the binarized encoder block to enhance the extracted representation. We sparsify the binarized units to reduce the low-information decoding computation. Through extensive experiments, we demonstrate that BiMatting outperforms other binarized video matting models, including state-of-the-art (SOTA) binarization methods, by a significant margin. Our approach even performs comparably to the full-precision counterpart in visual quality. Furthermore, BiMatting achieves remarkable savings of 12.4$\times$ and 21.6$\times$ in computation and storage, respectively, showcasing its potential and advantages in real-world resource-constrained scenarios. Our code and models are released at https://github.com/htqin/BiMatting . Haotong Qin, Lei Ke, Xudong Ma, Martin Danelljan, Yu-Wing Tai, Chi-Keung Tang, Xianglong Liu 0001, Fisher Yu 0001 |
NeurIPS | 5 |
| 2023 | FaceDNeRF: Semantics-Driven Face Reconstruction, Prompt Editing and Relighting with Diffusion ModelsabstractThe ability to create high-quality 3D faces from a single image has become increasingly important with wide applications in video conferencing, AR/VR, and advanced video editing in movie industries. In this paper, we propose Face Diffusion NeRF (FaceDNeRF), a new generative method to reconstruct high-quality Face NeRFs from single images, complete with semantic editing and relighting capabilities. FaceDNeRF utilizes high-resolution 3D GAN inversion and expertly trained 2D latent-diffusion model, allowing users to manipulate and construct Face NeRFs in zero-shot learning without the need for explicit 3D data.
With carefully designed illumination and identity preserving loss, as well as multi-modal pre-training, FaceDNeRF offers users unparalleled control over the editing process enabling them to create and edit face NeRFs using just single-view images, text prompts, and explicit target lighting. The advanced features of FaceDNeRF have been designed to produce more impressive results than existing 2D editing approaches that rely on 2D segmentation maps for editable attributes. Experiments show that our FaceDNeRF achieves exceptionally realistic results and unprecedented flexibility in editing compared with state-of-the-art 3D face reconstruction and editing methods. Our code will be available at https://github.com/BillyXYB/FaceDNeRF. Hao Zhang 0106, Tianyuan Dai, Yanbo Xu, Yu-Wing Tai, Chi-Keung Tang |
NeurIPS | 4 |
| 2023 | Occlusion-Aware Instance Segmentation Via BiLayer Network ArchitecturesabstractSegmenting highly-overlapping image objects is challenging, because there is typically no distinction between real object contours and occlusion boundaries on images. Unlike previous instance segmentation methods, we model image formation as a composition of two overlapping layers, and propose Bilayer Convolutional Network (BCNet), where the top layer detects occluding objects (occluders) and the bottom layer infers partially occluded instances (occludees). The explicit modeling of occlusion relationship with bilayer structure naturally decouples the boundaries of both the occluding and occluded instances, and considers the interaction between them during mask regression. We investigate the efficacy of bilayer structure using two popular convolutional network designs, namely, Fully Convolutional Network (FCN) and Graph Convolutional Network (GCN). Further, we formulate bilayer decoupling using the vision transformer (ViT), by representing instances in the image as separate learnable occluder and occludee queries. Large and consistent improvements using one/two-stage and query-based object detectors with various backbones and network layer choices validate the generalization ability of bilayer decoupling, as shown by extensive experiments on image instance segmentation benchmarks (COCO, KINS, COCOA) and video instance segmentation benchmarks (YTVIS, OVIS, BDD100 K MOTS), especially for heavy occlusion cases. Lei Ke, Yu-Wing Tai, Chi-Keung Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | GCoNet+: A Stronger Group Collaborative Co-Salient Object DetectorabstractIn this paper, we present a novel end-to-end group collaborative learning network, termed GCoNet+, which can effectively and efficiently (250 fps) identify co-salient objects in natural scenes. The proposed GCoNet+ achieves the new state-of-the-art performance for co-salient object detection (CoSOD) through mining consensus representations based on the following two essential criteria: 1) intra-group compactness to better formulate the consistency among co-salient objects by capturing their inherent shared attributes using our novel group affinity module (GAM); 2) inter-group separability to effectively suppress the influence of noisy objects on the output by introducing our new group collaborating module (GCM) conditioning on the inconsistent consensus. To further improve the accuracy, we design a series of simple yet effective components as follows: i) a recurrent auxiliary classification module (RACM) promoting model learning at the semantic level; ii) a confidence enhancement module (CEM) assisting the model in improving the quality of the final predictions; and iii) a group-based symmetric triplet (GST) loss guiding the model to learn more discriminative features. Extensive experiments on three challenging benchmarks, i.e., CoCA, CoSOD3k, and CoSal2015, demonstrate that our GCoNet+ outperforms the existing 12 cutting-edge models. Code has been released at https://github.com/ZhengPeng7/GCoNet_plus. Peng Zheng 0004, Huazhu Fu, Deng-Ping Fan, Jie Qin 0004, Yu-Wing Tai, Chi-Keung Tang, Luc Van Gool |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Transcoded Video Restoration by Temporal Spatial Auxiliary NetworkabstractIn most video platforms, such as Youtube, Kwai, and TikTok, the played videos usually have undergone multiple video encodings such as hardware encoding by recording devices, software encoding by video editing apps, and single/multiple video transcoding by video application servers. Previous works in compressed video restoration typically assume the compression artifacts are caused by one-time encoding. Thus, the derived solution usually does not work very well in practice. In this paper, we propose a new method, temporal spatial auxiliary network (TSAN), for transcoded video restoration. Our method considers the unique traits between video encoding and transcoding, and we consider the initial shallow encoded videos as the intermediate labels to assist the network to conduct self-supervised attention training. In addition, we employ adjacent multi-frame information and propose the temporal deformable alignment and pyramidal spatial fusion for transcoded video restoration. The experimental results demonstrate that the performance of the proposed method is superior to that of the previous techniques. The code is available at https://github.com/icecherylXuli/TSAN. Li Xu 0008, Gang He 0002, Jinjia Zhou, Jie Lei 0001, Weiying Xie, Yunsong Li 0001, Yu-Wing Tai |
AAAI | 7 |
| 2022 | Human Instance Matting via Mutual Guidance and Multi-Instance RefinementabstractThis paper introduces a new matting task called human instance matting (HIM), which requires the pertinent model to automatically predict a precise alpha matte for each human instance. Straightforward combination of closely related techniques, namely, instance segmentation, soft segmentation and human/conventional matting, will easily fail in complex cases requiring disentangling mingled colors belonging to multiple instances along hairy and thin boundary structures. To tackle these technical challenges, we propose a human instance matting framework, called InstMatt, where a novel mutual guidance strategy working in tandem with a multi-instance refinement module is used, for delineating multi-instance relationship among humans with complex and overlapping boundaries if present. A new instance matting metric called instance matting quality (IMQ) is proposed, which addresses the absence of a unified and fair means of evaluation emphasizing both instance recognition and matting quality. Finally, we construct a HIM benchmark for evaluation, which comprises of both synthetic and natural benchmark images. In addition to thorough experimental results on complex cases with multiple and overlapping human instances each has intricate boundaries, preliminary results are presented on general instance matting. Code and benchmark are available in https://github.com/nowsyn/InstMatt. Yanan Sun 0005, Chi-Keung Tang, Yu-Wing Tai |
CVPR | 3 |
| 2022 | Look Back and Forth: Video Super-Resolution with Explicit Temporal Difference ModelingabstractTemporal modeling is crucial for video super-resolution. Most of the video super-resolution methods adopt the optical flow or deformable convolution for explicitly motion compensation. However, such temporal modeling techniques increase the model complexity and might fail in case of occlusion or complex motion, resulting in serious distortion and artifacts. In this paper, we propose to explore the role of explicit temporal difference modeling in both LR and HR space. Instead of directly feeding consecutive frames into a VSR model, we propose to compute the temporal difference between frames and divide those pixels into two subsets according to the level of difference. They are separately processed with two branches of different receptive fields in order to better extract complementary information. To further enhance the super-resolution result, not only spatial residual features are extracted, but the difference between consecutive frames in high-frequency domain is also computed. It allows the model to exploit intermediate SR results in both future and past to refine the current SR output. The difference at different time steps could be cached such that information from further distance in time could be propagated to the current frame for refinement. Experiments on several video super-resolution benchmark datasets demonstrate the effectiveness of the proposed method and its favorable performance against state-of-the-art methods. Takashi Isobe, Xu Jia 0012, Xin Tao 0001, Ruihuang Li, Yongjie Shi, Huchuan Lu, Yu-Wing Tai |
CVPR | 9 |
| 2022 | Mask Transfiner for High-Quality Instance SegmentationabstractTwo-stage and query-based instance segmentation methods have achieved remarkable results. However, their segmented masks are still very coarse. In this paper, we present Mask Transfiner for high-quality and efficient instance segmentation. Instead of operating on regular dense tensors, our Mask Transfiner decomposes and represents the image regions as a quadtree. Our transformer-based approach only processes detected error-prone tree nodes and self-corrects their errors in parallel. While these sparse pixels only constitute a small proportion of the total number, they are critical to the final mask quality. This allows Mask Transfiner to predict highly accurate instance masks, at a low computational cost. Extensive experiments demonstrate that Mask Transfiner outperforms current instance segmentation methods on three popular benchmarks, significantly improving both two-stage and query-based frameworks by a large margin of +3.0 mask AP on COCO and BDD100K, and +6.6 boundary AP on Cityscapes. Our code and trained models are available at https://github.com/SysCV/transfiner. Lei Ke, Martin Danelljan, Xia Li 0005, Yu-Wing Tai, Chi-Keung Tang, Fisher Yu 0001 |
CVPR | 4 |
| 2022 | Interactiveness Field in Human-Object InteractionsabstractHuman-Object Interaction (HOI) detection plays a core role in activity understanding. Though recent two/one-stage methods have achieved impressive results, as an essential step, discovering interactive human-object pairs remains challenging. Both one/two-stage methods fail to effectively extract interactive pairs instead of generating redundant negative pairs. In this work, we introduce a previously overlooked interactiveness bimodal prior: given an object in an image, after pairing it with the humans, the generated pairs are either mostly non-interactive, or mostly interactive, with the former more frequent than the latter. Based on this interactiveness bimodal prior we propose the “interactiveness field”. To make the learned field compatible with real HOI image considerations, we propose new energy constraints based on the cardinality and difference in the inherent “interactiveness field” underlying interactive versus non-interactive pairs. Consequently, our method can detect more precise pairs and thus significantly boost HOI detection performance, which is validated on widely-used benchmarks where we achieve decent improvements over state-of-the-arts. Our code is available at https://github.comIForuckllnteractiveness-Field. Xinpeng Liu 0002, Yong-Lu Li 0001, Yu-Wing Tai, Cewu Lu, Chi-Keung Tang |
CVPR | 4 |
| 2022 | Self-support Few-Shot Semantic Segmentation
Wenjie Pei, Yu-Wing Tai, Chi-Keung Tang |
ECCV (19) | 3 |
| 2022 | Few-Shot Video Object Detection
Chi-Keung Tang, Yu-Wing Tai |
ECCV (20) | 3 |
| 2022 | Few-Shot Object Detection with Model Calibration
Chi-Keung Tang, Yu-Wing Tai |
ECCV (19) | 3 |
| 2022 | Video Mask Transfiner for High-Quality Video Instance Segmentation
Lei Ke, Henghui Ding, Martin Danelljan, Yu-Wing Tai, Chi-Keung Tang, Fisher Yu 0001 |
ECCV (28) | 4 |
| 2022 | SDRTV-to-HDRTV via Hierarchical Dynamic Context Feature MappingabstractIn this work, we address the task of SDR videos to HDR videos(SDRTV-to-HDRTV conversion). Previous approaches use global feature modulation for SDRTV-to-HDRTV conversion. Feature modulation scales and shifts the features in the original feature space, which has limited mapping capability. In addition, the global image mapping cannot restore detail in HDR frames due to the luminance differences in different regions of SDR frames. To resolve the appeal, we propose a two-stage solution. The first stage is a hierarchical Dynamic Context feature mapping (HDCFM) model. HDCFM learns the SDR frame to HDR frame mapping function via hierarchical feature modulation (HME and HM ) module and a dynamic context feature transformation (DYCT) module. The HME estimates the feature modulation vector, HM is capable of hierarchical feature modulation, consisting of global feature modulation in series with local feature modulation, and is capable of adaptive mapping of local image features. The DYCT module constructs a feature transformation module in conjunction with the context, which is capable of adaptively generating a feature transformation matrix for feature mapping. Compared with simple feature scaling and shifting, the DYCT module can map features into a new feature space and thus has a more excellent feature mapping capability. In the second stage, we introduce a patch discriminator-based context generation model PDCG to obtain subjective quality enhancement of over-exposed regions. The proposed method can achieve state-of-the-art objective and subjective quality results. Specifically, HDCFM achieves a PSNR gain of 0.81 dB at about 100K parameters. The number of parameters is 1/14th of the previous state-of-the-art methods. The test code will be released on https://github.com/cooperlike/HDCFM. Gang He 0002, Kepeng Xu, Li Xu 0008, Chang Wu 0001, Ming Sun 0008, Yu-Wing Tai |
ACM Multimedia | 7 |
| 2022 | DeViT: Deformed Vision Transformers in Video InpaintingabstractThis paper presents a novel video inpainting architecture named Deformed Vision Transformers (DeViT). We make three significant contributions to this task: First, we extended previous Transformers with patch alignment by introducing Deformed Patch-based Homography Estimator (DePtH), which enriches the patch-level feature alignments in key and query with additional offsets learned from patch pairs without additional supervision. DePtH enables our method to handle challenging scenes or agile motion with in-plane or out-of-plane deformation, which previous methods usually fail. Second, we introduce the Mask Pruning-based Patch Attention (MPPA) to improve the standard patch-wised feature matching by pruning out less essential features and considering the saliency map. MPPA enhances the matching accuracy between warped tokens with invalid pixels. Third, we introduce the Spatial-Temporal weighting Adaptor (STA) module to assign more accurate attention to spatial-temporal tokens under the guidance of the Deformation Factor learned from DePtH, especially for videos with agile motions. Experimental results demonstrate that our method outperforms previous state-of-the-art methods in quality and quantity and achieves a new state-of-the-art for video inpainting. Jiayin Cai, Xin Tao 0001, Chun Yuan 0003, Yu-Wing Tai |
ACM Multimedia | 5 |
| 2022 | NeRF-SR: High Quality Neural Radiance Fields using SupersamplingabstractWe present NeRF-SR, a solution for high-resolution (HR) novel view synthesis with mostly low-resolution (LR) inputs. Our method is built upon Neural Radiance Fields (NeRF) that predicts per-point density and color with a multi-layer perceptron. While producing images at arbitrary scales, NeRF struggles with resolutions that go beyond observed images. Our key insight is that NeRF benefits from 3D consistency, which means an observed pixel absorbs information from nearby views. We first exploit it by a super-sampling strategy that shoots multiple rays at each image pixel, which further enforces multi-view constraint at a sub-pixel level. Then, we show that NeRF-SR can further boost the performance of super-sampling by a refinement network that leverages the estimated depth at hand to hallucinate details from related patches on only one HR reference image. Experiment results demonstrate that NeRF-SR generates high-quality results for novel view synthesis at HR on both synthetic and real-world datasets without any external information. Project page: https://cwchenwang.github.io/NeRF-SR Chen Wang 0049, Xian Wu 0004, Song-Hai Zhang, Yu-Wing Tai, Shi-Min Hu 0001 |
ACM Multimedia | 5 |
| 2022 | Unsupervised Multi-View Object Segmentation Using Radiance Field PropagationabstractWe present radiance field propagation (RFP), a novel approach to segmenting objects in 3D during reconstruction given only unlabeled multi-view images of a scene. RFP is derived from emerging neural radiance field-based techniques, which jointly encodes semantics with appearance and geometry. The core of our method is a novel propagation strategy for individual objects' radiance fields with a bidirectional photometric loss, enabling an unsupervised partitioning of a scene into salient or meaningful regions corresponding to different object instances. To better handle complex scenes with multiple objects and occlusions, we further propose an iterative expectation-maximization algorithm to refine object masks. To the best of our knowledge, RFP is the first unsupervised approach for tackling 3D scene object segmentation for neural radiance field (NeRF) without any supervision, annotations, or other cues such as 3D bounding boxes and prior knowledge of object class. Experiments demonstrate that RFP achieves feasible segmentation results that are more accurate than previous unsupervised image/scene segmentation approaches, and are comparable to existing supervised NeRF-based methods. The segmented object representations enable individual 3D object editing operations. Codes and datasets will be made publicly available. Xinhang Liu, Jiaben Chen, Huai Yu, Yu-Wing Tai, Chi-Keung Tang |
NeurIPS | 4 |
| 2022 | Dual Convolutional Neural Networks for Low-Level Vision
Jinshan Pan, Deqing Sun, Jiawei Zhang 0002, Jinhui Tang 0001, Jian Yang 0003, Yu-Wing Tai, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 6 |
| 2022 | Learning Sequence Representations by Non-local Recurrent Neural Memory
Wenjie Pei, Xin Feng 0005, Canmiao Fu, Qiong Cao, Guangming Lu 0002, Yu-Wing Tai |
Int. J. Comput. Vis. | 6 |
| 2022 | PRIN/SPRIN: On Extracting Point-Wise Rotation Invariant FeaturesabstractPoint cloud analysis without pose priors is very challenging in real applications, as the orientations of point clouds are often unknown. In this paper, we propose a brand new point-set learning framework PRIN, namely, Point-wise Rotation Invariant Network, focusing on rotation invariant feature extraction in point clouds analysis. We construct spherical signals by Density Aware Adaptive Sampling to deal with distorted point distributions in spherical space. Spherical Voxel Convolution and Point Re-sampling are proposed to extract rotation invariant features for each point. In addition, we extend PRIN to a sparse version called SPRIN, which directly operates on sparse point clouds. Both PRIN and SPRIN can be applied to tasks ranging from object classification, part segmentation, to 3D feature matching and label alignment. Results show that, on the dataset with randomly rotated point clouds, SPRIN demonstrates better performance than state-of-the-art methods without any data augmentation. We also provide thorough theoretical proof and analysis for point-wise rotation invariance achieved by our methods. The code to reproduce our results will be made publicly available. Yang You 0004, Yujing Lou, Ruoxi Shi, Yu-Wing Tai, Lizhuang Ma, Cewu Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Hybrid Face Reflectance, Illumination, and Shape From a Single ImageabstractWe propose HyFRIS-Net to jointly estimate the hybrid reflectance and illumination models, as well as the refined face shape from a single unconstrained face image in a pre-defined texture space. The proposed hybrid reflectance and illumination representation ensure photometric face appearance modeling in both parametric and non-parametric spaces for efficient learning. While forcing the reflectance consistency constraint for the same person and face identity constraint for different persons, our approach recovers an occlusion-free face albedo with disambiguated color from the illumination color. Our network is trained in a self-evolving manner to achieve general applicability on real-world data. We conduct comprehensive qualitative and quantitative evaluations with state-of-the-art methods to demonstrate the advantages of HyFRIS-Net in modeling photo-realistic face albedo, illumination, and shape. Yongjie Zhu, Chen Li 0031, Si Li 0001, Boxin Shi, Yu-Wing Tai |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | Modular Interactive Video Object Segmentation: Interaction-to-Mask, Propagation and Difference-Aware FusionabstractWe present Modular interactive VOS (MiVOS) framework which decouples interaction-to-mask and mask propagation, allowing for higher generalizability and better performance. Trained separately, the interaction module converts user interactions to an object mask, which is then temporally propagated by our propagation module using a novel top-k filtering strategy in reading the space-time memory. To effectively take the user’s intent into account, a novel difference-aware module is proposed to learn how to properly fuse the masks before and after each interaction, which are aligned with the target frames by employing the space-time memory. We evaluate our method both qualitatively and quantitatively with different forms of user interactions (e.g., scribbles, clicks) on DAVIS to show that our method outperforms current state-of-the-art algorithms while requiring fewer frame interactions, with the additional advantage in generalizing to different types of user interactions. We contribute a large-scale synthetic VOS dataset with pixel-accurate segmentation of 4.8M frames to accompany our source codes to facilitate future research. Ho Kei Cheng, Yu-Wing Tai, Chi-Keung Tang |
CVPR | 2 |
| 2021 | Group Collaborative Learning for Co-Salient Object DetectionabstractWe present a novel group collaborative learning framework (GCoNet) capable of detecting co-salient objects in real time (16ms), by simultaneously mining consensus representations at group level based on the two necessary criteria: 1) intra-group compactness to better formulate the consistency among co-salient objects by capturing their inherent shared attributes using our novel group affinity module; 2) inter-group separability to effectively suppress the influence of noisy objects on the output by introducing our new group collaborating module conditioning the inconsistent consensus. To learn a better embedding space without extra computational overhead, we explicitly employ auxiliary classification supervision. Extensive experiments on three challenging benchmarks, i.e., CoCA, CoSOD3k, and Cosal2015, demonstrate that our simple GCoNet outperforms 10 cutting-edge models and achieves the new state-of-the-art. We demonstrate this paper’s new technical contributions on a number of important downstream computer vision applications including content aware co-segmentation, co-localization based automatic thumbnails, etc. Code has been made publicly available: https://github.com/fanq15/GCoNet. Deng-Ping Fan, Huazhu Fu, Chi-Keung Tang, Ling Shao 0001, Yu-Wing Tai |
CVPR | 6 |
| 2021 | Deep Occlusion-Aware Instance Segmentation With Overlapping BiLayersabstractSegmenting highly-overlapping objects is challenging, because typically no distinction is made between real object contours and occlusion boundaries. Unlike previous two-stage instance segmentation methods, we model image formation as composition of two overlapping layers, and propose Bilayer Convolutional Network (BCNet), where the top GCN layer detects the occluding objects (occluder) and the bottom GCN layer infers partially occluded instance (occludee). The explicit modeling of occlusion relationship with bilayer structure naturally decouples the boundaries of both the occluding and occluded instances, and considers the interaction between them during mask regression. We validate the efficacy of bilayer decoupling on both one-stage and two-stage object detectors with different backbones and network layer choices. Despite its simplicity, extensive experiments on COCO and KINS show that our occlusion-aware BCNet achieves large and consistent performance gain especially for heavy occlusion cases. Code is available at https://github.com/lkeab/BCNet. Lei Ke, Yu-Wing Tai, Chi-Keung Tang |
CVPR | 2 |
| 2021 | Semantic Image MattingabstractNatural image matting separates the foreground from background in fractional occupancy which can be caused by highly transparent objects, complex foreground (e.g., net or tree), and/or objects containing very fine details (e.g., hairs). Although conventional matting formulation can be applied to all of the above cases, no previous work has attempted to reason the underlying causes of matting due to various foreground semantics.We show how to obtain better alpha mattes by incorporating into our framework semantic classification of matting regions. Specifically, we consider and learn 20 classes of matting patterns, and propose to extend the conventional trimap to semantic trimap. The proposed semantic trimap can be obtained automatically through patch structure analysis within trimap regions. Meanwhile, we learn a multi-class discriminator to regularize the alpha prediction at semantic level, and content-sensitive weights to balance different regularization losses. Experiments on multiple benchmarks show that our method outperforms other methods and has achieved the most competitive state-of-the-art performance. Finally, we contribute a large-scale Semantic Image Matting Dataset with careful consideration of data balancing across different semantic classes. Code and dataset are available at https://github.com/nowsyn/SIM. Yanan Sun 0005, Chi-Keung Tang, Yu-Wing Tai |
CVPR | 3 |
| 2021 | Deep Video Matting via Spatio-Temporal Alignment and AggregationabstractDespite the significant progress made by deep learning in natural image matting, there has been so far no representative work on deep learning for video matting due to the inherent technical challenges in reasoning temporal domain and lack of large-scale video matting datasets. In this paper, we propose a deep learning-based video matting framework which employs a novel and effective spatio-temporal feature aggregation module (ST-FAM). As optical flow estimation can be very unreliable within matting regions, ST-FAM is designed to effectively align and aggregate information across different spatial scales and temporal frames within the network decoder. To eliminate frame-by-frame trimap annotations, a lightweight interactive trimap propagation network is also introduced. The other contribution consists of a large-scale video matting dataset with groundtruth alpha mattes for quantitative evaluation and real-world high-resolution videos with trimaps for qualitative evaluation. Quantitative and qualitative experimental results show that our framework significantly outperforms conventional video matting and deep image matting methods applied to video in presence of multi-frame temporal information. Our dataset is available at https://github.com/nowsyn/DVM. Yanan Sun 0005, Guanzhi Wang, Qiao Gu, Chi-Keung Tang, Yu-Wing Tai |
CVPR | 5 |
| 2021 | HAA500: Human-Centric Atomic Action Dataset with Curated VideosabstractWe contribute HAA5001, a manually annotated human-centric atomic action dataset for action recognition on 500 classes with over 591K labeled frames. To minimize ambiguities in action classification, HAA500 consists of highly diversified classes of fine-grained atomic actions, where only consistent actions fall under the same label, e.g., "Baseball Pitching" vs "Free Throw in Basketball". Thus HAA500 is different from existing atomic action datasets, where coarse-grained atomic actions were labeled with coarse action-verbs such as "Throw". HAA500 has been carefully curated to capture the precise movement of human figures with little class-irrelevant motions or spatiotemporal label noises.The advantages of HAA500 are fourfold: 1) human-centric actions with a high average of 69.7% detectable joints for the relevant human poses; 2) high scalability since adding a new class can be done under 20–60 minutes; 3) curated videos capturing essential elements of an atomic action without irrelevant frames; 4) fine-grained atomic action classes. Our extensive experiments including cross-data validation using datasets collected in the wild demonstrate the clear benefits of human-centric and atomic characteristics of HAA500, which enable training even a baseline deep learning model to improve prediction by attending to atomic human poses. We detail the HAA500 dataset statistics and collection methodology and compare quantitatively with existing action recognition datasets. Jihoon Chung, Cheng-hsin Wuu, Hsuan-ru Yang, Yu-Wing Tai, Chi-Keung Tang |
ICCV | 4 |
| 2021 | Occlusion-Aware Video Object InpaintingabstractConventional video inpainting is neither object-oriented nor occlusion-aware, making it liable to obvious artifacts when large occluded object regions are inpainted. This paper presents occlusion-aware video object inpainting, which recovers both the complete shape and appearance for occluded objects in videos given their visible mask segmentation.To facilitate this new research, we construct the first large-scale video object inpainting benchmark YouTube-VOI to provide realistic occlusion scenarios with both occluded and visible object masks available. Our technical contribution VOIN jointly performs video object shape completion and occluded texture generation. In particular, the shape completion module models long-range object coherence while the flow completion module recovers accurate flow with sharp motion boundary, for propagating temporally-consistent texture to the same moving object across frames. For more realistic results, VOIN is optimized using both T-PatchGAN and a new spatio-temporal attention-based multi-class discriminator.Finally, we compare VOIN and strong baselines on YouTube-VOI. Experimental results clearly demonstrate the efficacy of our method including inpainting complex and dynamic objects. VOIN degrades gracefully with inaccurate input visible mask. Lei Ke, Yu-Wing Tai, Chi-Keung Tang |
ICCV | 2 |
| 2021 | Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object SegmentationabstractThis paper presents a simple yet effective approach to modeling space-time correspondences in the context of video object segmentation. Unlike most existing approaches, we establish correspondences directly between frames without re-encoding the mask features for every object, leading to a highly efficient and robust framework. With the correspondences, every node in the current query frame is inferred by aggregating features from the past in an associative fashion. We cast the aggregation process as a voting problem and find that the existing inner-product affinity leads to poor use of memory with a small (fixed) subset of memory nodes dominating the votes, regardless of the query. In light of this phenomenon, we propose using the negative squared Euclidean distance instead to compute the affinities. We validated that every memory node now has a chance to contribute, and experimentally showed that such diversified voting is beneficial to both memory efficiency and inference accuracy. The synergy of correspondence networks and diversified voting works exceedingly well, achieves new state-of-the-art results on both DAVIS and YouTubeVOS datasets while running significantly faster at 20+ FPS for multiple objects without bells and whistles. Ho Kei Cheng, Yu-Wing Tai, Chi-Keung Tang |
NeurIPS | 2 |
| 2021 | Prototypical Cross-Attention Networks for Multiple Object Tracking and SegmentationabstractMultiple object tracking and segmentation requires detecting, tracking, and segmenting objects belonging to a set of given classes. Most approaches only exploit the temporal dimension to address the association problem, while relying on single frame predictions for the segmentation mask itself. We propose Prototypical Cross-Attention Network (PCAN), capable of leveraging rich spatio-temporal information for online multiple object tracking and segmentation. PCAN first distills a space-time memory into a set of prototypes and then employs cross-attention to retrieve rich information from the past frames. To segment each object, PCAN adopts a prototypical appearance module to learn a set of contrastive foreground and background prototypes, which are then propagated over time. Extensive experiments demonstrate that PCAN outperforms current video instance tracking and segmentation competition winners on both Youtube-VIS and BDD100K datasets, and shows efficacy to both one-stage and two-stage segmentation frameworks. Code and video resources are available at http://vis.xyz/pub/pcan. Lei Ke, Xia Li 0005, Martin Danelljan, Yu-Wing Tai, Chi-Keung Tang, Fisher Yu 0001 |
NeurIPS | 4 |
| 2021 | Physics-Based Generative Adversarial Models for Image Restoration and BeyondabstractWe present an algorithm to directly solve numerous image restoration problems (e.g., image deblurring, image dehazing, and image deraining). These problems are ill-posed, and the common assumptions for existing methods are usually based on heuristic image priors. In this paper, we show that these problems can be solved by generative models with adversarial learning. However, a straightforward formulation based on a straightforward generative adversarial network (GAN) does not perform well in these tasks, and some structures of the estimated images are usually not preserved well. Motivated by an interesting observation that the estimated results should be consistent with the observed inputs under the physics models, we propose an algorithm that guides the estimation process of a specific task within the GAN framework. The proposed model is trained in an end-to-end fashion and can be applied to a variety of image restoration and low-level vision problems. Extensive experiments demonstrate that the proposed method performs favorably against state-of-the-art algorithms. Jinshan Pan, Jiangxin Dong, Yang Liu 0119, Jiawei Zhang 0002, Jimmy S. J. Ren, Jinhui Tang 0001, Yu-Wing Tai, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2021 | An Accurate and Lightweight Method for Human Body Image Super-ResolutionabstractIn this paper, we propose a new method to super-resolve low resolution human body images by learning efficient multi-scale features and exploiting useful human body prior. Specifically, we propose a lightweight multi-scale block (LMSB) as basic module of a coherent framework, which contains an image reconstruction branch and a prior estimation branch. In the image reconstruction branch, the LMSB aggregates features of multiple receptive fields so as to gather rich context information for low-to-high resolution mapping. In the prior estimation branch, we adopt the human parsing maps and nonsubsampled shearlet transform (NSST) sub-bands to represent the human body prior, which is expected to enhance the details of reconstructed human body images. When evaluated on the newly collected HumanSR dataset, our method outperforms state-of-the-art image super-resolution methods with ∼ 8× fewer parameters; moreover, our method significantly improves the performance of human image analysis tasks (e.g. human parsing and pose estimation) for low-resolution inputs. Yunan Liu 0001, Shanshan Zhang 0001, Jie Xu 0021, Jian Yang 0003, Yu-Wing Tai |
IEEE Trans. Image Process. | 5 |
| 2021 | Push for Center Learning via Orthogonalization and Subspace Masking for Person Re-IdentificationabstractPerson re-identification aims to identify whether pairs of images belong to the same person or not. This problem is challenging due to large differences in camera views, lighting and background. One of the mainstream in learning CNN features is to design loss functions which reinforce both the class separation and intra-class compactness. In this paper, we propose a novel Orthogonal Center Learning method with Subspace Masking for person re-identification. We make the following contributions: 1) we develop a center learning module to learn the class centers by simultaneously reducing the intra-class differences and inter-class correlations by orthogonalization; 2) we introduce a subspace masking mechanism to enhance the generalization of the learned class centers; and 3) we propose to integrate the average pooling and max pooling in a regularizing manner that fully exploits their powers. Extensive experiments show that our proposed method consistently outperforms the state-of-the-art methods on large-scale ReID datasets including Market-1501, DukeMTMC-ReID, CUHK03 and MSMT17. Weinong Wang, Wenjie Pei, Qiong Cao, Shu Liu 0005, Guangming Lu 0002, Yu-Wing Tai |
IEEE Trans. Image Process. | 6 |
| 2021 | Hierarchical Generation of Human Pose With Part-Based Layer RepresentationabstractHuman pose transfer has been becoming one of the emerging research topics in recent years. However, state-of-the-art results are still far from satisfactory. One main reason is that these end-to-end methods are often blindly trained without the semantic understanding of its content. In this paper, we propose a novel method for human pose transfer with consideration of the semantic part-based representation of a human. In particular, we propose to segment the human body into multiple parts, and each of them represents a semantic region of a human. With the proposed part-based layer generators, a high-quality result is guaranteed for each local semantic region. We design a three-stage hierarchical framework to fuse local representations into the final result in a coarse-to-fine manner, which provides adaptive attention for global consistency and local details, respectively. Via exploiting spatial guidance from 3D human model through the framework, our method can naturally handle the ambiguity of self-occlusions which always causes artifacts in previous methods. With semantic-aware and spatial-aware representations, our method outperforms previous approaches quantitatively and qualitatively in better handling self-occlusions, fine detail preservation/synthesis and a higher resolution result. Xian Wu 0004, Chen Li 0031, Shi-Min Hu 0001, Yu-Wing Tai |
IEEE Trans. Image Process. | 4 |
| 2020 | Pointwise Rotation-Invariant Network with Adaptive Sampling and 3D Spherical Voxel ConvolutionabstractPoint cloud analysis without pose priors is very challenging in real applications, as the orientations of point clouds are often unknown. In this paper, we propose a brand new point-set learning framework PRIN, namely, Pointwise Rotation-Invariant Network, focusing on rotation-invariant feature extraction in point clouds analysis. We construct spherical signals by Density Aware Adaptive Sampling to deal with distorted point distributions in spherical space. In addition, we propose Spherical Voxel Convolution and Point Re-sampling to extract rotation-invariant features for each point. Our network can be applied to tasks ranging from object classification, part segmentation, to 3D feature matching and label alignment. We show that, on the dataset with randomly rotated point clouds, PRIN demonstrates better performance than state-of-the-art methods without any data augmentation. We also provide theoretical analysis for the rotation-invariance achieved by our methods. Yang You 0004, Yujing Lou, Yu-Wing Tai, Lizhuang Ma, Cewu Lu |
AAAI | 4 |
| 2020 | CascadePSP: Toward Class-Agnostic and Very High-Resolution Segmentation via Global and Local RefinementabstractState-of-the-art semantic segmentation methods were almost exclusively trained on images within a fixed resolution range. These segmentations are inaccurate for very high-resolution images since using bicubic upsampling of low-resolution segmentation does not adequately capture high-resolution details along object boundaries. In this paper, we propose a novel approach to address the high-resolution segmentation problem without using any high-resolution training data. The key insight is our CascadePSP network which refines and corrects local boundaries whenever possible. Although our network is trained with low-resolution segmentation data, our method is applicable to any resolution even for very high-resolution images larger than 4K. We present quantitative and qualitative studies on different datasets to show that CascadePSP can reveal pixel-accurate segmentation boundaries using our novel refinement module without any finetuning. Thus, our method can be regarded as class-agnostic. Finally, we demonstrate the application of our model to scene parsing in multi-class segmentation. Ho Kei Cheng, Jihoon Chung, Yu-Wing Tai, Chi-Keung Tang |
CVPR | 3 |
| 2020 | Few-Shot Object Detection With Attention-RPN and Multi-Relation DetectorabstractConventional methods for object detection typically require a substantial amount of training data and preparing such high-quality training data is very labor-intensive. In this paper, we propose a novel few-shot object detection network that aims at detecting objects of unseen categories with only a few annotated examples. Central to our method are our Attention-RPN, Multi-Relation Detector and Contrastive Training strategy, which exploit the similarity between the few shot support set and query set to detect novel objects while suppressing false detection in the background. To train our network, we contribute a new dataset that contains 1000 categories of various objects with high-quality annotations. To the best of our knowledge, this is one of the first datasets specifically designed for few-shot object detection. Once our few-shot network is trained, it can detect objects of unseen categories without further training or fine-tuning. Our method is general and has a wide range of potential applications. We produce a new state-of-the-art performance on different datasets in the few-shot setting. The dataset link is https://github.com/fanq15/Few-Shot-Object-Detection-Dataset. Wei Zhuo 0005, Chi-Keung Tang, Yu-Wing Tai |
CVPR | 4 |
| 2020 | Fast Video Object Segmentation With Temporal Aggregation Network and Dynamic Template MatchingabstractSignificant progress has been made in Video Object Segmentation (VOS), the video object tracking task in its finest level. While the VOS task can be naturally decoupled into image semantic segmentation and video object tracking, significantly much more research effort has been made in segmentation than tracking. In this paper, we introduce “tracking-by-detection” into VOS which can coherently integrates segmentation into tracking, by proposing a new temporal aggregation network and a novel dynamic time-evolving template matching mechanism to achieve significantly improved performance. Notably, our method is entirely online and thus suitable for one-shot learning, and our end-to-end trainable model allows multiple object segmentation in one forward pass. We achieve new state-of-the-art performance on the DAVIS benchmark without complicated bells and whistles in both speed and accuracy, with a speed of 0.14 second per frame and J &F measure of 75.9% respectively. Xuhua Huang, Yu-Wing Tai, Chi-Keung Tang |
CVPR | 3 |
| 2020 | Cascaded Deep Monocular 3D Human Pose Estimation With Evolutionary Training DataabstractEnd-to-end deep representation learning has achieved remarkable accuracy for monocular 3D human pose estimation, yet these models may fail for unseen poses with limited and fixed training data. This paper proposes a novel data augmentation method that: (1) is scalable for synthesizing massive amount of training data (over 8 million valid 3D human poses with corresponding 2D projections) for training 2D-to-3D networks, (2) can effectively reduce dataset bias. Our method evolves a limited dataset to synthesize unseen 3D human skeletons based on a hierarchical human representation and heuristics inspired by prior knowledge. Extensive experiments show that our approach not only achieves state-of-the-art accuracy on the largest public benchmark, but also generalizes significantly better to unseen and rare poses. Relevant files and tools are available at the project website. Shichao Li 0002, Lei Ke, Kevin Pratama, Yu-Wing Tai, Chi-Keung Tang, Kwang-Ting Cheng |
CVPR | 4 |
| 2020 | FSS-1000: A 1000-Class Dataset for Few-Shot SegmentationabstractOver the past few years, we have witnessed the success of deep learning in image recognition thanks to the availability of large-scale human-annotated datasets such as PASCAL VOC, ImageNet, and COCO. Although these datasets have covered a wide range of object categories, there are still a significant number of objects that are not included. Can we perform the same task without a lot of human annotations? In this paper, we are interested in few-shot object segmentation where the number of annotated training examples are limited to 5 only. To evaluate and validate the performance of our approach, we have built a few-shot segmentation dataset, FSS-1000, which consists of 1000 object classes with pixelwise annotation of ground-truth segmentation. Unique in FSS-1000, our dataset contains significant number of objects that have never been seen or annotated in previous datasets, such as tiny daily objects, merchandise, cartoon characters, logos, etc. We build our baseline model using standard backbone networks such as VGG-16, ResNet-101, and Inception. To our surprise, we found that training our model from scratch using FSS-1000 achieves comparable and even better results than training with weights pre-trained by ImageNet which is more than 100 times larger than FSS-1000. Both our approach and dataset are simple, effective, and easily extensible to learn segmentation of new object classes given very few annotated training examples. Dataset is available at https://github.com/HKUSTCV/FSS-1000. Tianhan Wei, Yau Pun Chen, Yu-Wing Tai, Chi-Keung Tang |
CVPR | 4 |
| 2020 | Learning Video Object Segmentation From Unlabeled VideosabstractWe propose a new method for video object segmentation (VOS) that addresses object pattern learning from unlabeled videos, unlike most existing methods which rely heavily on extensive annotated data. We introduce a unified unsupervised/weakly supervised learning framework, called MuG, that comprehensively captures intrinsic properties of VOS at multiple granularities. Our approach can help advance understanding of visual patterns in VOS and significantly reduce annotation burden. With a carefully-designed architecture and strong representation learning ability, our learned model can be applied to diverse VOS settings, including object-level zero-shot VOS, instance-level zero-shot VOS, and one-shot VOS. Experiments demonstrate promising performance in these settings, as well as the potential of MuG in leveraging unlabeled data to further improve the segmentation accuracy. Xiankai Lu, Wenguan Wang, Jianbing Shen, Yu-Wing Tai, David Crandall, Steven C. H. Hoi |
CVPR | 4 |
| 2020 | Boosting the Transferability of Adversarial Samples via AttentionabstractThe widespread deployment of deep models necessitates the assessment of model vulnerability in practice, especially for safety- and security-sensitive domains such as autonomous driving and medical diagnosis. Transfer-based attacks against image classifiers thus elicit mounting interest, where attackers are required to craft adversarial images based on local proxy models without the feedback information from remote target ones. However, under such a challenging but practical setup, the synthesized adversarial samples often achieve limited success due to overfitting to the local model employed. In this work, we propose a novel mechanism to alleviate the overfitting issue. It computes model attention over extracted features to regularize the search of adversarial examples, which prioritizes the corruption of critical features that are likely to be adopted by diverse architectures. Consequently, it can promote the transferability of resultant adversarial instances. Extensive experiments on ImageNet classifiers confirm the effectiveness of our strategy and its superiority to state-of-the-art benchmarks in both white-box and black-box settings. Weibin Wu 0002, Yuxin Su 0001, Shenglin Zhao, Irwin King, Michael R. Lyu, Yu-Wing Tai |
CVPR | 7 |
| 2020 | Towards Global Explanations of Convolutional Neural Networks With Concept AttributionabstractWith the growing prevalence of convolutional neural networks (CNNs), there is an urgent demand to explain their behaviors. Global explanations contribute to understanding model predictions on a whole category of samples, and thus have attracted increasing interest recently. However, existing methods overwhelmingly conduct separate input attribution or rely on local approximations of models, making them fail to offer faithful global explanations of CNNs. To overcome such drawbacks, we propose a novel two-stage framework, Attacking for Interpretability (AfI), which explains model decisions in terms of the importance of user-defined concepts. AfI first conducts a feature occlusion analysis, which resembles a process of attacking models to derive the category-wide importance of different features. We then map the feature importance to concept importance through ad-hoc semantic tasks. Experimental results confirm the effectiveness of AfI and its superiority in providing more accurate estimations of concept importance than existing proposals. Weibin Wu 0002, Yuxin Su 0001, Shenglin Zhao, Irwin King, Michael R. Lyu, Yu-Wing Tai |
CVPR | 7 |
| 2020 | Dive Deeper into Box for Object Detection
Ran Chen 0001, Mengdan Zhang, Shu Liu 0005, Bei Yu 0001, Yu-Wing Tai |
ECCV (22) | 6 |
| 2020 | Fully Convolutional Networks for Continuous Sign Language Recognition
Ka Leong Cheng, Zhaoyang Yang, Qifeng Chen 0001, Yu-Wing Tai |
ECCV (24) | 4 |
| 2020 | Commonality-Parsing Network Across Shape and Appearance for Partially Supervised Instance Segmentation
Lei Ke, Wenjie Pei, Chi-Keung Tang, Yu-Wing Tai |
ECCV (8) | 5 |
| 2020 | GSNet: Joint Vehicle Pose and Shape Reconstruction with Geometrical and Scene-Aware Supervision
Lei Ke, Shichao Li 0002, Yanan Sun 0005, Yu-Wing Tai, Chi-Keung Tang |
ECCV (15) | 4 |
| 2020 | Dense Hybrid Recurrent Multi-view Stereo Net with Dynamic Consistency Checking
Jianfeng Yan, Zizhuang Wei, Hongwei Yi, Mingyu Ding, Yisong Chen, Yu-Wing Tai |
ECCV (4) | 8 |
| 2020 | Pyramid Multi-view Stereo Net with Self-adaptive View Aggregation
Hongwei Yi, Zizhuang Wei, Mingyu Ding, Yisong Chen, Yu-Wing Tai |
ECCV (9) | 7 |
| 2020 | Semi-Calibrated Photometric StereoabstractWhile conventional calibrated photometric stereo methods assume that light intensities and sensor exposures are known or unknown but identical across observed images, this assumption easily breaks down in practical settings due to individual light bulb's characteristics and limited control over sensors. This paper studies the effect of unknown and possibly non-uniform light intensities and sensor exposures among observed images on the shape recovery based on photometric stereo. This leads to the development of a "semi-calibrated" photometric stereo method, where the light directions are known but light intensities (and sensor exposures) are unknown. We show that the semi-calibrated photometric stereo becomes a bilinear problem, whose general form is difficult to solve, but in the photometric stereo context, there exists a unique solution for the surface normal and light intensities (or sensor exposures). We further show that there exists a linear solution method for the problem, and develop efficient and stable solution methods. The semi-calibrated photometric stereo is advantageous over conventional calibrated photometric stereo in accurate determination of surface normal, because it relaxes the assumption of known light intensity ratios/sensor exposures. The experimental results show superior accuracy of the semi-calibrated photometric stereo in comparison to conventional methods in practical settings. Donghyeon Cho, Yasuyuki Matsushita, Yu-Wing Tai, In-So Kweon |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Memory-Attended Recurrent Network for Video CaptioningabstractTypical techniques for video captioning follow the encoder-decoder framework, which can only focus on one source video being processed. A potential disadvantage of such design is that it cannot capture the multiple visual context information of a word appearing in more than one relevant videos in training data. To tackle this limitation, we propose the Memory-Attended Recurrent Network (MARN) for video captioning, in which a memory structure is designed to explore the full-spectrum correspondence between a word and its various similar visual contexts across videos in training data. Thus, our model is able to achieve a more comprehensive understanding for each word and yield higher captioning quality. Furthermore, the built memory structure enables our method to model the compatibility between adjacent words explicitly instead of asking the model to learn implicitly, as most existing models do. Extensive validation on two real-word datasets demonstrates that our MARN consistently outperforms state-of-the-art methods. Wenjie Pei, Xiangrong Wang 0002, Lei Ke, Xiaoyong Shen, Yu-Wing Tai |
CVPR | 6 |
| 2019 | MMFace: A Multi-Metric Regression Network for Unconstrained Face ReconstructionabstractWe propose to address the face reconstruction in the wild by using a multi-metric regression network, MMFace, to align a 3D face morphable model (3DMM) to an input image. The key idea is to utilize a volumetric sub-network to estimate an intermediate geometry representation, and a parametric sub-network to regress the 3DMM parameters. Our parametric sub-network consists of identity loss, expression loss, and pose loss which greatly improves the aligned geometry details by incorporating high level loss functions directly defined in the 3DMM parametric spaces. Our high-quality reconstruction is robust under large variations of expressions, poses, illumination conditions, and even with large partial occlusions. We evaluate our method by comparing the performance with state-of-the-art approaches on latest 3D face dataset LS3D-W and Florence. We achieve significant improvements both quantitatively and qualitatively. Due to our high-quality reconstruction, our method can be easily extended to generate high-quality geometry sequences for video inputs. Hongwei Yi, Chen Li 0031, Qiong Cao, Xiaoyong Shen, Sheng Li 0008, Yu-Wing Tai |
CVPR | 7 |
| 2019 | Adversarial Attacks Beyond the Image SpaceabstractGenerating adversarial examples is an intriguing problem and an important way of understanding the working mechanism of deep neural networks. Most existing approaches generated perturbations in the image space, i.e., each pixel can be modified independently. However, in this paper we pay special attention to the subset of adversarial examples that correspond to meaningful changes in 3D physical properties (like rotation and translation, illumination condition, etc.). These adversaries arguably pose a more serious concern, as they demonstrate the possibility of causing neural network failure by easy perturbations of real-world 3D objects and scenes. In the contexts of object classification and visual question answering, we augment state-of-the-art deep neural networks that receive 2D input images with a rendering module (either differentiable or not) in front, so that a 3D scene (in the physical space) is rendered into a 2D image (in the image space), and then mapped to a prediction (in the output space). The adversarial perturbations can now go beyond the image space, and have clear meanings in the 3D physical world. Though image-space adversaries can be interpreted as per-pixel albedo change, we verify that they cannot be well explained along these physically meaningful dimensions, which often have a non-local effect. But it is still possible to successfully attack beyond the image space on the physical space, though this is more difficult than image-space attacks, reflected in lower success rates and heavier perturbations required. Xiaohui Zeng, Chenxi Liu 0001, Yu-Siang Wang, Weichao Qiu, Lingxi Xie, Yu-Wing Tai, Chi-Keung Tang, Alan L. Yuille |
CVPR | 6 |
| 2019 | Cross-Domain Adaptation for Animal Pose EstimationabstractIn this paper, we are interested in pose estimation of animals. Animals usually exhibit a wide range of variations on poses and there is no available animal pose dataset for training and testing. To address this problem, we build an animal pose dataset to facilitate training and evaluation. Considering the heavy labor needed to label dataset and it is impossible to label data for all concerned animal species, we, therefore, proposed a novel cross-domain adaptation method to transform the animal pose knowledge from labeled animal classes to unlabeled animal classes. We use the modest animal pose dataset to adapt learned knowledge to multiple animals species. Moreover, humans also share skeleton similarities with some animals (especially four-footed mammals). Therefore, the easily available human pose dataset, which is of a much larger scale than our labeled animal dataset, provides important prior knowledge to boost up the performance on animal pose estimation. Experiments show that our proposed method leverages these pieces of prior knowledge well and achieves convincing results on animal pose estimation. Jinkun Cao, Hongyang Tang, Haoshu Fang, Xiaoyong Shen, Yu-Wing Tai, Cewu Lu |
ICCV | 5 |
| 2019 | Non-Local Recurrent Neural Memory for Supervised Sequence ModelingabstractTypical methods for supervised sequence modeling are built upon the recurrent neural networks to capture temporal dependencies. One potential limitation of these methods is that they only model explicitly information interactions between adjacent time steps in a sequence, hence the high-order interactions between nonadjacent time steps are not fully exploited. It greatly limits the capability of modeling the long-range temporal dependencies since one-order interactions cannot be maintained for a long term due to information dilution and gradient vanishing. To tackle this limitation, we propose the Non-local Recurrent Neural Memory (NRNM) for supervised sequence modeling, which performs non-local operations to learn full-order interactions within a sliding temporal block and models the global interactions between blocks in a gated recurrent manner. Consequently, our model is able to capture the long-range dependencies. Besides, the latent high-level features contained in high-order interactions can be distilled by our model. We demonstrate the merits of our NRNM approach on two different tasks: action recognition and sentiment analysis. Canmiao Fu, Wenjie Pei, Qiong Cao, Chaopeng Zhang, Yong Zhao 0010, Xiaoyong Shen, Yu-Wing Tai |
ICCV | 7 |
| 2019 | LADN: Local Adversarial Disentangling Network for Facial Makeup and De-MakeupabstractWe propose a local adversarial disentangling network (LADN) for facial makeup and de-makeup. Central to our method are multiple and overlapping local adversarial discriminators in a content-style disentangling network for achieving local detail transfer between facial images, with the use of asymmetric loss functions for dramatic makeup styles with high-frequency details. Existing techniques do not demonstrate or fail to transfer high-frequency details in a global adversarial setting, or train a single local discriminator only to ensure image structure consistency and thus work only for relatively simple styles. Unlike others, our proposed local adversarial discriminators can distinguish whether the generated local image details are consistent with the corresponding regions in the given reference image in cross-image style transfer in an unsupervised setting. Incorporating these technical contributions, we achieve not only state-of-the-art results on conventional styles but also novel results involving complex and dramatic styles with high-frequency details covering large areas across multiple facial features. A carefully designed dataset of unpaired before and after makeup images is released at https://georgegu1997.github.io/LADN-project-page. Qiao Gu, Guanzhi Wang, Mang Tik Chiu, Yu-Wing Tai, Chi-Keung Tang |
ICCV | 4 |
| 2019 | Reflective Decoding Network for Image CaptioningabstractState-of-the-art image captioning methods mostly focus on improving visual features, less attention has been paid to utilizing the inherent properties of language to boost captioning performance. In this paper, we show that vocabulary coherence between words and syntactic paradigm of sentences are also important to generate high-quality image caption. Following the conventional encoder-decoder framework, we propose the Reflective Decoding Network (RDN) for image captioning, which enhances both the long-sequence dependency and position perception of words in a caption decoder. Our model learns to collaboratively attend on both visual and textual features and meanwhile perceive each word's relative position in the sentence to maximize the information delivered in the generated caption. We evaluate the effectiveness of our RDN on the COCO image captioning datasets and achieve superior performance over the previous methods. Further experiments reveal that our approach is particularly advantageous for hard cases with complex scenes to describe by captions. Lei Ke, Wenjie Pei, Ruiyu Li, Xiaoyong Shen, Yu-Wing Tai |
ICCV | 5 |
| 2019 | Depth from a Light Field Image with Learning-Based Matching CostsabstractOne of the core applications of light field imaging is depth estimation. To acquire a depth map, existing approaches apply a single photo-consistency measure to an entire light field. However, this is not an optimal choice because of the non-uniform light field degradations produced by limitations in the hardware design. In this paper, we introduce a pipeline that automatically determines the best configuration for photo-consistency measure, which leads to the most reliable depth label from the light field. We analyzed the practical factors affecting degradation in lenslet light field cameras, and designed a learning based framework that can retrieve the best cost measure and optimal depth label. To enhance the reliability of our method, we augmented an existing light field benchmark to simulate realistic source dependent noise, aberrations, and vignetting artifacts. The augmented dataset was used for the training and validation of the proposed approach. Our method was competitive with several state-of-the-art methods for the benchmark and real-world light field datasets. Hae-Gon Jeon, Jaesik Park, Gyeongmin Choe, Jinsun Park, Yunsu Bok, Yu-Wing Tai, In-So Kweon |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2019 | Deep Convolutional Neural Network for Natural Image Matting Using Initial Alpha MattesabstractWe propose a deep convolutional neural network (CNN) method for natural image matting. Our method takes multiple initial alpha mattes of the previous methods and normalized RGB color images as inputs, and directly learns an end-to-end mapping between the inputs and reconstructed alpha mattes. Among the various existing methods, we focus on using two simple methods as initial alpha mattes: the closed-form matting and KNN matting. They are complementary to each other in terms of local and nonlocal principles. A major benefit of our method is that it can "recognize" different local image structures and then combine the results of local (closed-form matting) and nonlocal (KNN matting) mattings effectively to achieve higher quality alpha mattes than both of the inputs. Furthermore, we verify extendability of the proposed network to different combinations of initial alpha mattes from more advanced techniques such as KL divergence matting and information-flow matting. On the top of deep CNN matting, we build an RGB guided JPEG artifacts removal network to handle JPEG block artifacts in alpha matting. Extensive experiments demonstrate that our proposed deep CNN matting produces visually and quantitatively high-quality alpha mattes. We perform deeper experiments including studies to evaluate the importance of balancing training data and to measure the effects of initial alpha mattes and also consider results from variant versions of the proposed network to analyze our proposed DCNN matting. In addition, our method achieved high ranking in the public alpha matting evaluation dataset in terms of the sum of absolute differences, mean squared errors, and gradient errors. Also, our RGB guided JPEG artifacts removal network restores the damaged alpha mattes from compressed images in JPEG format. Donghyeon Cho, Yu-Wing Tai, In-So Kweon |
IEEE Trans. Image Process. | 2 |
| 2018 | Weakly and Semi Supervised Human Body Part Parsing via Pose-Guided Knowledge TransferabstractHuman body part parsing, or human semantic part segmentation, is fundamental to many computer vision tasks. In conventional semantic segmentation methods, the ground truth segmentations are provided, and fully convolutional networks (FCN) are trained in an end-to-end scheme. Although these methods have demonstrated impressive results, their performance highly depends on the quantity and quality of training data. In this paper, we present a novel method to generate synthetic human part segmentation data using easily-obtained human keypoint annotations. Our key idea is to exploit the anatomical similarity among human to transfer the parsing results of a person to another person with similar pose. Using these estimated results as additional training data, our semi-supervised model outperforms its strong-supervised counterpart by 6 mIOU on the PASCAL-Person-Part dataset [6], and we achieve state-of-the-art human parsing results. Our approach is general and can be readily extended to other object/animal parsing task assuming that their anatomical similarity can be annotated by keypoints. The proposed model and accompanying source code will be made publicly available. Haoshu Fang, Guansong Lu, Xiaolin Fang 0002, Jianwen Xie, Yu-Wing Tai, Cewu Lu |
CVPR | 5 |
| 2018 | Learning Dual Convolutional Neural Networks for Low-Level VisionabstractIn this paper, we propose a general dual convolutional neural network (DualCNN) for low-level vision problems, e.g., super-resolution, edge-preserving filtering, deraining and dehazing. These problems usually involve the estimation of two components of the target signals: structures and details. Motivated by this, our proposed DualCNN consists of two parallel branches, which respectively recovers the structures and details in an end-to-end manner. The recovered structures and details can generate the target signals according to the formation model for each particular application. The DualCNN is a flexible framework for low-level vision tasks and can be easily incorporated into existing CNNs. Experimental results show that the DualCNN can be effectively applied to numerous low-level vision tasks with favorable performance against the state-of-the-art methods. Jinshan Pan, Sifei Liu, Deqing Sun, Jiawei Zhang 0002, Yang Liu 0119, Jimmy S. J. Ren, Zechao Li, Jinhui Tang 0001, Huchuan Lu, Yu-Wing Tai, Ming-Hsuan Yang 0001 |
CVPR | 10 |
| 2018 | Deep Video Generation, Prediction and Completion of Human Action Sequences
Haoye Cai, Chunyan Bai, Yu-Wing Tai, Chi-Keung Tang |
ECCV (2) | 3 |
| 2018 | Pairwise Body-Part Attention for Recognizing Human-Object Interactions
Haoshu Fang, Jinkun Cao, Yu-Wing Tai, Cewu Lu |
ECCV (10) | 3 |
| 2018 | Attribute-Guided Face Generation Using Conditional CycleGAN
Yongyi Lu, Yu-Wing Tai, Chi-Keung Tang |
ECCV (12) | 2 |
| 2018 | Image Generation from Sketch Constraint Using Contextual GAN
Yongyi Lu, Shangzhe Wu, Yu-Wing Tai, Chi-Keung Tang |
ECCV (16) | 3 |
| 2018 | Deep High Dynamic Range Imaging with Large Foreground Motions
Shangzhe Wu, Yu-Wing Tai, Chi-Keung Tang |
ECCV (2) | 3 |
| 2018 | ELD-Net: An Efficient Deep Learning Architecture for Accurate Saliency DetectionabstractRecent advances in saliency detection have utilized deep learning to obtain high-level features to detect salient regions in scenes. These advances have yielded results superior to those reported in past work, which involved the use of hand-crafted low-level features for saliency detection. In this paper, we propose ELD-Net, a unified deep learning framework for accurate and efficient saliency detection. We show that hand-crafted features can provide complementary information to enhance saliency detection that uses only high-level features. Our method uses both low-level and high-level features for saliency detection. High-level features are extracted using GoogLeNet, and low-level features evaluate the relative importance of a local region using its differences from other regions in an image. The two feature maps are independently encoded by the convolutional and the ReLU layers. The encoded low-level and high-level features are then combined by concatenation and convolution. Finally, a linear fully connected layer is used to evaluate the saliency of a queried region. A full resolution saliency map is obtained by querying the saliency of each local region of an image. Since the high-level features are encoded at low resolution, and the encoded high-level features can be reused for every query region, our ELD-Net is very fast. Our experiments show that our method outperforms state-of-the-art deep learning-based saliency detection methods. Gayoung Lee, Yu-Wing Tai, Junmo Kim 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Fast Randomized Singular Value Thresholding for Low-Rank OptimizationabstractRank minimization can be converted into tractable surrogate problems, such as Nuclear Norm Minimization (NNM) and Weighted NNM (WNNM). The problems related to NNM, or WNNM, can be solved iteratively by applying a closed-form proximal operator, called Singular Value Thresholding (SVT), or Weighted SVT, but they suffer from high computational cost of Singular Value Decomposition (SVD) at each iteration. We propose a fast and accurate approximation method for SVT, that we call fast randomized SVT (FRSVT), with which we avoid direct computation of SVD. The key idea is to extract an approximate basis for the range of the matrix from its compressed matrix. Given the basis, we compute partial singular values of the original matrix from the small factored matrix. In addition, by developping a range propagation method, our method further speeds up the extraction of approximate basis at each iteration. Our theoretical analysis shows the relationship between the approximation bound of SVD and its effect to NNM via SVT. Along with the analysis, our empirical results quantitatively and qualitatively show that our approximation rarely harms the convergence of the host algorithms. We assess the efficiency and accuracy of the proposed method on various computer vision problems, e.g., subspace clustering, weather artifact removal, and simultaneous multi-image alignment and rectification. Tae-Hyun Oh, Yasuyuki Matsushita, Yu-Wing Tai, In-So Kweon |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | A Unified Approach of Multi-scale Deep and Hand-Crafted Features for Defocus EstimationabstractIn this paper, we introduce robust and synergetic hand-crafted features and a simple but efficient deep feature from a convolutional neural network (CNN) architecture for defocus estimation. This paper systematically analyzes the effectiveness of different features, and shows how each feature can compensate for the weaknesses of other features when they are concatenated. For a full defocus map estimation, we extract image patches on strong edges sparsely, after which we use them for deep and hand-crafted feature extraction. In order to reduce the degree of patch-scale dependency, we also propose a multi-scale patch extraction strategy. A sparse defocus map is generated using a neural network classifier followed by a probability-joint bilateral filter. The final defocus map is obtained from the sparse defocus map with guidance from an edge-preserving filtered input image. Experimental results show that our algorithm is superior to state-of-the-art algorithms in terms of defocus estimation. Our work can be used for applications such as segmentation, blur magnification, all-in-focus image generation, and 3-D estimation. Jinsun Park, Yu-Wing Tai, Donghyeon Cho, In-So Kweon |
CVPR | 2 |
| 2017 | Accurate Single Stage Detector Using Recurrent Rolling ConvolutionabstractMost of the recent successful methods in accurate object detection and localization used some variants of R-CNN style two stage Convolutional Neural Networks (CNN) where plausible regions were proposed in the first stage then followed by a second stage for decision refinement. Despite the simplicity of training and the efficiency in deployment, the single stage detection methods have not been as competitive when evaluated in benchmarks consider mAP for high IoU thresholds. In this paper, we proposed a novel single stage end-to-end trainable object detection network to overcome this limitation. We achieved this by introducing Recurrent Rolling Convolution (RRC) architecture over multi-scale feature maps to construct object classifiers and bounding box regressors which are deep in context. We evaluated our method in the challenging KITTI dataset which measures methods under IoU threshold of 0.7. We showed that with RRC, a single reduced VGG-16 based model already significantly outperformed all the previously published results. At the time this paper was written our models ranked the first in KITTI car detection (the hard level), the first in cyclist detection and the second in pedestrian detection. These results were not reached by the previous single stage methods. The code is publicly available. Jimmy S. J. Ren, Xiaohao Chen, Wenxiu Sun, Jiahao Pang, Qiong Yan, Yu-Wing Tai, Li Xu 0001 |
CVPR | 7 |
| 2017 | Exploring Heterogeneous Algorithms for Accelerating Deep Convolutional Neural Networks on FPGAsabstractConvolutional neural network (CNN) finds applications in a variety of computer vision applications ranging from object recognition and detection to scene understanding owing to its exceptional accuracy. There exist different algorithms for CNNs computation. In this paper, we explore conventional convolution algorithm with a faster algorithm using Winograd's minimal filtering theory for efficient FPGA implementation. Distinct from the conventional convolution algorithm, Winograd algorithm uses less computing resources but puts more pressure on the memory bandwidth. We first propose a fusion architecture that can fuse multiple layers naturally in CNNs, reusing the intermediate data. Based on this fusion architecture, we explore heterogeneous algorithms to maximize the throughput of a CNN. We design an optimal algorithm to determine the fusion and algorithm strategy for each layer. We also develop an automated toolchain to ease the mapping from Caffe model to FPGA bitstream using Vivado HLS. Experiments using widely used VGG and AlexNet demonstrate that our design achieves up to 1.99X performance speedup compared to the prior fusion-based FPGA accelerator for CNNs. Qingcheng Xiao, Yun Liang 0001, Liqiang Lu, Shengen Yan, Yu-Wing Tai |
DAC | 5 |
| 2017 | Weakly- and Self-Supervised Learning for Content-Aware Deep Image RetargetingabstractThis paper proposes a weakly- and self-supervised deep convolutional neural network (WSSDCNN) for content-aware image retargeting. Our network takes a source image and a target aspect ratio, and then directly outputs a retargeted image. Retargeting is performed through a shift reap, which is a pixel-wise mapping from the source to the target grid. Our method implicitly learns an attention map, which leads to r content-aware shift map for image retargeting. As a result, discriminative parts in an image are preserved, while background regions are adjusted seamlessly. In the training phase, pairs of an image and its image-level annotation are used to compute content and structure tosses. We demonstrate the effectiveness of our proposed method for a retargeting application with insightful analyses. Donghyeon Cho, Jinsun Park, Tae-Hyun Oh, Yu-Wing Tai, In-So Kweon |
ICCV | 4 |
| 2017 | RMPE: Regional Multi-person Pose EstimationabstractMulti-person pose estimation in the wild is challenging. Although state-of-the-art human detectors have demonstrated good performance, small errors in localization and recognition are inevitable. These errors can cause failures for a single-person pose estimator (SPPE), especially for methods that solely depend on human detection results. In this paper, we propose a novel regional multi-person pose estimation (RMPE) framework to facilitate pose estimation in the presence of inaccurate human bounding boxes. Our framework consists of three components: Symmetric Spatial Transformer Network (SSTN), Parametric Pose Non-Maximum-Suppression (NMS), and Pose-Guided Proposals Generator (PGPG). Our method is able to handle inaccurate bounding boxes and redundant detections, allowing it to achieve 76:7 mAP on the MPII (multi person) dataset[3]. Our model and source codes are made publicly available. Haoshu Fang, Shuqin Xie, Yu-Wing Tai, Cewu Lu |
ICCV | 3 |
| 2017 | Learning Discriminative Data Fitting Functions for Blind Image DeblurringabstractSolving blind image deblurring usually requires defining a data fitting function and image priors. While existing algorithms mainly focus on developing image priors for blur kernel estimation and non-blind deconvolution, only a few methods consider the effect of data fitting functions. In contrast to the state-of-the-art methods that use a single or a fixed data fitting term, we propose a data-driven approach to learn effective data fitting functions from a large set of motion blurred images with the associated ground truth blur kernels. The learned data fitting function facilitates estimating accurate blur kernels for generic scenes and domain-specific problems with corresponding image priors. In addition, we extend the learning approach for data fitting function to latent image restoration and nonuniform deblurring. Extensive experiments on challenging motion blurred images demonstrate the proposed algorithm performs favorably against the state-of-the-art methods. Jinshan Pan, Jiangxin Dong, Yu-Wing Tai, Zhixun Su, Ming-Hsuan Yang 0001 |
ICCV | 3 |
| 2017 | PCA Based Computation of Illumination-Invariant Space for Road DetectionabstractIllumination changes such as shadows significantly affect the accuracy of various road detection methods, especially for vision-based approaches with an on-board monocular camera. To efficiently consider such illumination changes, we propose a PCA based technique, PCA-II, that finds the minimum projection space from an input RGB image, and then use the space as the illumination-invariant space for road detection. Our PCA based method shows 20 times faster performance on average over the prior entropy based method, even with a higher detection accuracy. To demonstrate its wide applicability to the road detection problem, we test the invariant space with both bottomup and top-down approaches. For a bottom-up approach, we suggest a simple patch propagation method that utilizes the property of the invariant space, and show its higher accuracy over other state-of-the-art road detection methods running in a bottom-up manner. For a top-down approach, we consider the space as an additional feature to the original RGB to train convolutional neural networks. We were also able to observe robust performance improvement of using the invariant space over the original CNN based methods that do not use the space, only with a minor runtime overhead, e.g., 50 ms per image. These results demonstrate benefits of our PCA-based illuminationinvariant space computation. Yu-Wing Tai, Sung-Eui Yoon |
WACV | 2 |
| 2017 | Category-Specific Salient View Selection via Deep Convolutional Neural NetworksabstractAbstract In this paper, we present a new framework to determine up front orientations and detect salient views of 3D models. The salient viewpoint to human preferences is the most informative projection with correct upright orientation. Our method utilizes two Convolutional Neural Network (CNN) architectures to encode category‐specific information learnt from a large number of 3D shapes and 2D images on the web. Using the first CNN model with 3D voxel data, we generate a CNN shape feature to decide natural upright orientation of 3D objects. Once a 3D model is upright‐aligned, the front projection and salient views are scored by category recognition using the second CNN model. The second CNN is trained over popular photo collections from internet users. In order to model comfortable viewing angles of 3D models, a category‐dependent prior is also learnt from the users. Our approach effectively combines category‐specific scores and classical evaluations to produce a data‐driven viewpoint saliency map. The best viewpoints from the method are quantitatively and qualitatively validated with more than 100 objects from 20 categories. Our thumbnail images of 3D models are the most favoured among those from different approaches. Seong-Heum Kim, Yu-Wing Tai, Joon-Young Lee, Jaesik Park, In-So Kweon |
Comput. Graph. Forum | 2 |
| 2017 | Refining Geometry from Depth Sensors using IR Shading Images
Gyeongmin Choe, Jaesik Park, Yu-Wing Tai, In-So Kweon |
Int. J. Comput. Vis. | 3 |
| 2017 | Automatic Trimap Generation and Consistent Matting for Light-Field ImagesabstractIn this paper, we introduce an automatic approach to generate trimaps and consistent alpha mattes of foreground objects in a light-field image. Our method first performs binary segmentation to roughly segment a light-field image into foreground and background based on depth and color. Next, we estimate accurate trimaps through analyzing color distribution along the boundary of the segmentation using guided image filter and KL-divergence. In order to estimate consistent alpha mattes across sub-images, we utilize the epipolar plane image (EPI) where colors and alphas along the same epipolar line must be consistent. Since EPI of foreground and background are mixed in the matting area, we propagate the EPI from definite foreground/background regions to unknown regions by assuming depth variations within unknown regions are spatially smooth. Using the EPI constraint, we derive two solutions to estimate alpha when color samples along epipolar line are known, and unknown. To further enhance consistency, we refine the estimated alpha mattes by using the multi-image matting Laplacian with an additional EPI smoothness constraint. In experimental evaluations, we have created a dataset where the ground truth alpha mattes of light-field images were obtained by using the blue screen technique. A variety of experiments show that our proposed algorithm produces both visually and quantitatively high-quality alpha mattes for light-field images. Donghyeon Cho, Sunyeong Kim, Yu-Wing Tai, In-So Kweon |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | Robust Multiview Photometric Stereo Using Planar Mesh ParameterizationabstractWe propose a robust uncalibrated multiview photometric stereo method for high quality 3D shape reconstruction. In our method, a coarse initial 3D mesh obtained using a multiview stereo method is projected onto a 2D planar domain using a planar mesh parameterization technique. We describe methods for surface normal estimation that work in the parameterized 2D space that jointly incorporates all geometric and photometric cues from multiple viewpoints. Using an estimated surface normal map, a refined 3D mesh is then recovered by computing an optimal displacement map in the same 2D planar domain. Our method avoids the need of merging view-dependent surface normal maps that is often required in conventional methods. We conduct evaluation on various real-world objects containing surfaces with specular reflections, multiple albedos, and complex topologies in both controlled and uncontrolled settings and demonstrate that accurate 3D meshes with fine geometric details can be recovered by our method. Jaesik Park, Sudipta N. Sinha, Yasuyuki Matsushita, Yu-Wing Tai, In-So Kweon |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2016 | Look, Listen and Learn - A Multimodal LSTM for Speaker IdentificationabstractSpeaker identification refers to the task of localizing the face of a person who has the same identity as the ongoing voice in a video. This task not only requires collective perception over both visual and auditory signals, the robustness to handle severe quality degradations and unconstrained content variations are also indispensable. In this paper, we describe a novel multimodal Long Short-Term Memory (LSTM) architecture which seamlessly unifies both visual and auditory modalities from the beginning of each sequence input. The key idea is to extend the conventional LSTM by not only sharing weights across time steps, but also sharing weights across modalities. We show that modeling the temporal dependency across face and voice can significantly improve the robustness to content quality degradations and variations. We also found that our multimodal LSTM is robustness to distractors, namely the non-speaking identities. We applied our multimodal LSTM to The Big Bang Theory dataset and showed that our system outperforms the state-of-the-art systems in speaker identification with lower false alarm rate and higher recognition accuracy. Jimmy S. J. Ren, Yongtao Hu 0001, Yu-Wing Tai, Li Xu 0001, Wenxiu Sun, Qiong Yan |
AAAI | 3 |
| 2016 | Deep Saliency with Encoded Low Level Distance Map and High Level FeaturesabstractRecent advances in saliency detection have utilized deep learning to obtain high level features to detect salient regions in a scene. These advances have demonstrated superior results over previous works that utilize hand-crafted low level features for saliency detection. In this paper, we demonstrate that hand-crafted features can provide complementary information to enhance performance of saliency detection that utilizes only high level features. Our method utilizes both high level and low level features for saliency detection under a unified deep learning framework. The high level features are extracted using the VGG-net, and the low level features are compared with other parts of an image to form a low level distance map. The low level distance map is then encoded using a convolutional neural network(CNN) with multiple 1 1 convolutional and ReLU layers. We concatenate the encoded low level distance map and the high level features, and connect them to a fully connected neural network classifier to evaluate the saliency of a query region. Our experiments show that our method can further improve the performance of state-of-the-art deep learning-based saliency detection methods. Gayoung Lee, Yu-Wing Tai, Junmo Kim 0002 |
CVPR | 2 |
| 2016 | Efficient and Robust Color Consistency for Community Photo CollectionsabstractWe present an efficient technique to optimize color consistency of a collection of images depicting a common scene. Our method first recovers sparse pixel correspondences in the input images and stacks them into a matrix with many missing entries. We show that this matrix satisfies a rank two constraint under a simple color correction model. These parameters can be viewed as pseudo white balance and gamma correction parameters for each input image. We present a robust low-rank matrix factorization method to estimate the unknown parameters of this model. Using them, we improve color consistency of the input images or perform color transfer with any input image as the source. Our approach is insensitive to outliers in the pixel correspondences thereby precluding the need for complex pre-processing steps. We demonstrate high quality color consistency results on large photo collections of popular tourist landmarks and personal photo collections containing images of people. Jaesik Park, Yu-Wing Tai, Sudipta N. Sinha, In-So Kweon |
CVPR | 2 |
| 2016 | Photometric Stereo Under Non-uniform Light Intensities and Exposures
Donghyeon Cho, Yasuyuki Matsushita, Yu-Wing Tai, In-So Kweon |
ECCV (2) | 3 |
| 2016 | Natural Image Matting Using Deep Convolutional Neural Networks
Donghyeon Cho, Yu-Wing Tai, In-So Kweon |
ECCV (2) | 2 |
| 2016 | Partial Sum Minimization of Singular Values in Robust PCA: Algorithm and ApplicationsabstractRobust Principal Component Analysis (RPCA) via rank minimization is a powerful tool for recovering underlying low-rank structure of clean data corrupted with sparse noise/outliers. In many low-level vision problems, not only it is known that the underlying structure of clean data is low-rank, but the exact rank of clean data is also known. Yet, when applying conventional rank minimization for those problems, the objective function is formulated in a way that does not fully utilize a priori target rank information about the problems. This observation motivates us to investigate whether there is a better alternative solution when using rank minimization. In this paper, instead of minimizing the nuclear norm, we propose to minimize the partial sum of singular values, which implicitly encourages the target rank constraint. Our experimental analyses show that, when the number of samples is deficient, our approach leads to a higher success rate than conventional rank minimization, while the solutions obtained by the two approaches are almost identical when the number of samples is more than sufficient. We apply our approach to various low-level vision problems, e.g., high dynamic range imaging, motion edge detection, photometric stereo, image alignment and recovery, and show that our results outperform those obtained by the conventional nuclear norm rank minimization method. Tae-Hyun Oh, Yu-Wing Tai, Jean-Charles Bazin, Hyeongwoo Kim, In-So Kweon |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Salient Region Detection via High-Dimensional Color Transform and Local Spatial SupportabstractIn this paper, we introduce a novel approach to automatically detect salient regions in an image. Our approach consists of global and local features, which complement each other to compute a saliency map. The first key idea of our work is to create a saliency map of an image by using a linear combination of colors in a high-dimensional color space. This is based on an observation that salient regions often have distinctive colors compared with backgrounds in human perception, however, human perception is complicated and highly nonlinear. By mapping the low-dimensional red, green, and blue color to a feature vector in a high-dimensional color space, we show that we can composite an accurate saliency map by finding the optimal linear combination of color coefficients in the high-dimensional color space. To further improve the performance of our saliency estimation, our second key idea is to utilize relative location and color contrast between superpixels as features and to resolve the saliency estimation from a trimap via a learning-based algorithm. The additional local features and learning-based algorithm complement the global estimation from the high-dimensional color transform-based algorithm. The experimental results on three benchmark datasets show that our approach is effective in comparison with the previous state-of-the-art saliency estimation methods. Jiwhan Kim, Dongyoon Han, Yu-Wing Tai, Junmo Kim 0002 |
IEEE Trans. Image Process. | 3 |
| 2016 | Multi-View Object Extraction With Fractional BoundariesabstractThis paper presents an automatic method to extract a multi-view object in a natural environment. We assume that the target object is bounded by the convex volume of interest defined by the overlapping space of camera viewing frustums. There are two key contributions of our approach. First, we present an automatic method to identify a target object across different images for multi-view binary co-segmentation. The extracted target object shares the same geometric representation in space with a distinctive color and texture model from the background. Second, we present an algorithm to detect color ambiguous regions along the object boundary for matting refinement. Our matting region detection algorithm is based on information theory, which measures the Kullback-Leibler (KL) divergence of local color distribution of different pixel-bands. The local pixel-band with the largest entropy is selected for matte refinement, subject to the multi-view consistent constraint. Our results are highquality alpha mattes consistent across all different viewpoints. We demonstrate the effectiveness of the proposed method using various examples. Seong-Heum Kim, Yu-Wing Tai, Jaesik Park, In-So Kweon |
IEEE Trans. Image Process. | 2 |
| 2016 | Robust All-in-Focus Super-Resolution for Focal Stack PhotographyabstractWe present an unconventional image super-resolution algorithm targeting focal stack images. Contrary to previous works, which align multiple images with sub-pixel accuracy for image super-resolution, we analyze the correlation among the differently focused narrow depth-of-field images in a focal stack to infer high-resolution details for image super-resolution. In order to accurately model the defocus kernels at different depths, we use a cubic interpolation to parameterize the projection of defocus kernels, and apply the radon transform to accurately reconstruct the defocus kernels at arbitrary depth. In the image super-resolution, we utilize the multi-image deconvolution method with a l1 -norm regularization to suppress noise and ringing artifacts. We have also extended the depth-of-field of our inputs to produce an all-in-focus super-resolution image. The effectiveness of our algorithm is demonstrated with the quantitative analysis using synthetic examples and the qualitative analysis using real-world examples. Minhaeng Lee, Yu-Wing Tai |
IEEE Trans. Image Process. | 2 |
| 2015 | Accurate depth map estimation from a lenslet light field cameraabstractThis paper introduces an algorithm that accurately estimates depth maps using a lenslet light field camera. The proposed algorithm estimates the multi-view stereo correspondences with sub-pixel accuracy using the cost volume. The foundation for constructing accurate costs is threefold. First, the sub-aperture images are displaced using the phase shift theorem. Second, the gradient costs are adaptively aggregated using the angular coordinates of the light field. Third, the feature correspondences between the sub-aperture images are used as additional constraints. With the cost volume, the multi-label optimization propagates and corrects the depth map in the weak texture regions. Finally, the local depth map is iteratively refined through fitting the local quadratic function to estimate a non-discrete depth map. Because micro-lens images contain unexpected distortions, a method is also proposed that corrects this error. The effectiveness of the proposed algorithm is demonstrated through challenging real world examples and including comparisons with the performance of advanced depth estimation algorithms. Hae-Gon Jeon, Jaesik Park, Gyeongmin Choe, Jinsun Park, Yunsu Bok, Yu-Wing Tai, In-So Kweon |
CVPR | 6 |
| 2015 | Data-driven depth map refinement via multi-scale sparse representationabstractDepth maps captured by consumer-level depth cameras such as Kinect are usually degraded by noise, missing values, and quantization. In this paper, we present a data-driven approach for refining degraded RAWdepth maps that are coupled with an RGB image. The key idea of our approach is to take advantage of a training set of high-quality depth data and transfer its information to the RAW depth map through multi-scale dictionary learning. Utilizing a sparse representation, our method learns a dictionary of geometric primitives which captures the correlation between high-quality mesh data, RAW depth maps and RGB images. The dictionary is learned and applied in a manner that accounts for various practical issues that arise in dictionary-based depth refinement. Compared to previous approaches that only utilize the correlation between RAW depth maps and RGB images, our method produces improved depth maps without over-smoothing. Since our approach is data driven, the refinement can be targeted to a specific class of objects by employing a corresponding training set. In our experiments, we show that this leads to additional improvements in recovering depth maps of human faces. Hyeokhyen Kwon, Yu-Wing Tai, Stephen Lin 0001 |
CVPR | 2 |
| 2015 | Fast randomized Singular Value Thresholding for Nuclear Norm MinimizationabstractRank minimization problem can be boiled down to either Nuclear Norm Minimization (NNM) or Weighted NNM (WNNM) problem. The problems related to NNM (or WNNM) can be solved iteratively by applying a closed-form proximal operator, called Singular Value Thresholding (SVT) (or Weighted SVT), but they suffer from high computational cost to compute a Singular Value Decomposition (SVD) at each iteration. In this paper, we propose an accurate and fast approximation method for SVT, called fast randomized SVT (FRSVT), where we avoid direct computation of SVD. The key idea is to extract an approximate basis for the range of a matrix from its compressed matrix. Given the basis, we compute the partial singular values of the original matrix from a small factored matrix. While the basis approximation is the bottleneck, our method is already severalfold faster than thin SVD. By adopting a range propagation technique, we can further avoid one of the bottleneck at each iteration. Our theoretical analysis provides a stepping stone between the approximation bound of SVD and its effect to NNM via SVT. Along with the analysis, our empirical results on both quantitative and qualitative studies show our approximation rarely harms the convergence behavior of the host algorithms. We apply it and validate the efficiency of our method on various vision problems, e.g. subspace clustering, weather artifact removal, simultaneous multi-image alignment and rectification. Tae-Hyun Oh, Yasuyuki Matsushita, Yu-Wing Tai, In-So Kweon |
CVPR | 3 |
| 2015 | RGB-Guided Hyperspectral Image UpsamplingabstractHyperspectral imaging usually lack of spatial resolution due to limitations of hardware design of imaging sensors. On the contrary, latest imaging sensors capture a RGB image with resolution of multiple times larger than a hyperspectral image. In this paper, we present an algorithm to enhance and upsample the resolution of hyperspectral images. Our algorithm consists of two stages: spatial upsampling stage and spectrum substitution stage. The spatial upsampling stage is guided by a high resolution RGB image of the same scene, and the spectrum substitution stage utilizes sparse coding to locally refine the upsampled hyperspectral image through dictionary substitution. Experiments show that our algorithm is highly effective and has outperformed state-of-the-art matrix factorization based approaches. Hyeokhyen Kwon, Yu-Wing Tai |
ICCV | 2 |
| 2015 | Image denoising via coded aperture photographyabstractWe present a novel image denoising method utilizing coded aperture photography. Our approach captures an image that is slightly optically defocused by a coded aperture. This allows us to more effectively reduce noise while high frequency of image structures are protected by the coded aperture image. We analyze the effectiveness of coded aperture in decoupling noise frequency from high frequency of image structures. A novel frequency-aware regularization is proposed to denoise and to restore sharp image from a noisy slightly out-of-focus coded aperture image. The effectiveness of our approach is demonstrated on various challenging examples with quantitative and qualitative comparisons to results of state-of-the-art denoising methods. Minhaeng Lee, Yu-Wing Tai |
ICIP | 2 |
| 2015 | A simulation based method for vehicle motion prediction
Jae-Hyuck Park, Yu-Wing Tai |
Comput. Vis. Image Underst. | 2 |
| 2015 | Robust High Dynamic Range Imaging by Rank MinimizationabstractThis paper introduces a new high dynamic range (HDR) imaging algorithm which utilizes rank minimization. Assuming a camera responses linearly to scene radiance, the input low dynamic range (LDR) images captured with different exposure time exhibit a linear dependency and form a rank-1 matrix when stacking intensity of each corresponding pixel together. In practice, misalignments caused by camera motion, presences of moving objects, saturations and image noise break the rank-1 structure of the LDR images. To address these problems, we present a rank minimization algorithm which simultaneously aligns LDR images and detects outliers for robust HDR generation. We evaluate the performances of our algorithm systematically using synthetic examples and qualitatively compare our results with results from the state-of-the-art HDR algorithms using challenging real world examples. Tae-Hyun Oh, Joon-Young Lee, Yu-Wing Tai, In-So Kweon |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | Exploiting Shading Cues in Kinect IR Images for Geometry RefinementabstractIn this paper, we propose a method to refine geometry of 3D meshes from the Kinect fusion by exploiting shading cues captured from the infrared (IR) camera of Kinect. A major benefit of using the Kinect IR camera instead of a RGB camera is that the IR images captured by Kinect are narrow band images which filtered out most undesired ambient light that makes our system robust to natural indoor illumination. We define a near light IR shading model which describes the captured intensity as a function of surface normals, albedo, lighting direction, and distance between a light source and surface points. To resolve ambiguity in our model between normals and distance, we utilize an initial 3D mesh from the Kinect fusion and multi-view information to reliably estimate surface details that were not reconstructed by the Kinect fusion. Our approach directly operates on a 3D mesh model for geometry refinement. The effectiveness of our approach is demonstrated through several challenging real-world examples. Gyeongmin Choe, Jaesik Park, Yu-Wing Tai, In-So Kweon |
CVPR | 3 |
| 2014 | Salient Region Detection via High-Dimensional Color TransformabstractIn this paper, we introduce a novel technique to automatically detect salient regions of an image via high-dimensional color transform. Our main idea is to represent a saliency map of an image as a linear combination of high-dimensional color space where salient regions and backgrounds can be distinctively separated. This is based on an observation that salient regions often have distinctive colors compared to the background in human perception, but human perception is often complicated and highly nonlinear. By mapping a low dimensional RGB color to a feature vector in a high-dimensional color space, we show that we can linearly separate the salient regions from the background by finding an optimal linear combination of color coefficients in the high-dimensional color space. Our high dimensional color space incorporates multiple color representations including RGB, CIELab, HSV and with gamma corrections to enrich its representative power. Our experimental results on three benchmark datasets show that our technique is effective, and it is computationally efficient in comparison to previous state-of-the-art techniques. Jiwhan Kim, Dongyoon Han, Yu-Wing Tai, Junmo Kim 0002 |
CVPR | 3 |
| 2014 | Calibrating a Non-isotropic Near Point Light Source Using a PlaneabstractWe show that a non-isotropic near point light source rigidly attached to a camera can be calibrated using multiple images of a weakly textured planar scene. We prove that if the radiant intensity distribution (RID) of a light source is radially symmetric with respect to its dominant direction, then the shading observed on a Lambertian scene plane is bilaterally symmetric with respect to a 2D line on the plane. The symmetry axis detected in an image provides a linear constraint for estimating the dominant light axis. The light position and RID parameters can then be estimated using a linear method. Specular highlights if available can also be used for light position estimation. We also extend our method to handle non-Lambertian reflectances which we model using a biquadratic BRDF. We have evaluated our method on synthetic data quantitavely. Our experiments on real scenes show that our method works well in practice and enables light calibration without the need of a specialized hardware. Jaesik Park, Sudipta N. Sinha, Yasuyuki Matsushita, Yu-Wing Tai, In-So Kweon |
CVPR | 4 |
| 2014 | Consistent Matting for Light Field Images
Donghyeon Cho, Sunyeong Kim, Yu-Wing Tai |
ECCV (4) | 3 |
| 2014 | Hierarchical nonrigid model for 3D medical image registrationabstractIn this paper, we propose a hierarchical model for medical image registration with a new descriptor which considers scale, rotation, and location attributes. Our proposed algorithm is a feature-based registration technique and our features encode location and orientation for matching. Using the proposed feature, 3D medical images are registered through three phases in the hierarchical model that progressively estimate the geometric transformation to align and orient features. Our approach is evaluated on both synthetic and real data using the ground-truth evaluation. Our results show that our method improved alignment accuracy compared to traditional nonrigid image registration algorithms. Sunyeong Kim, Yu-Wing Tai |
ICIP | 2 |
| 2014 | Robust pan-sharpening via color samples relocation and edge aware interpolationabstractWe present a pan-sharpening method that can produce high quality high-resolution multispectral image by fusing a high-resolution panchromatic image with a low resolution multispectral image. The major benefits of our approach are the color samples relocation algorithm and the optimization based edge aware interpolation method which protect the reconstructed images from aliasing and color diffusion artifacts around edge areas that are commonly arose in previous pan-sharpening algorithms. Our approach is robust to misalignment errors between the high-resolution panchromatic image and the low resolution multispectral image. We evaluate our results quantitatively and qualitatively on both synthetic and real world satellite images. Minhaeng Lee, Yu-Wing Tai |
ICIP | 3 |
| 2014 | A randomized algorithm for natural object colorizationabstractAbstract Natural objects often contain vivid color distribution with wide variety of colors. Conventional colorization techniques, on the other hand, produce colors that are relatively flat with little color variation. In this paper, we introduce a randomized algorithm which considers not only the value of target color but also the distribution of target color. In essence, our algorithm paints a color distribution to a region which synthesizes color distribution of a natural object. Our approach models the correlation between intensity and color in HSV color space in terms of H – S, H – V and S – V joint histogram. During the colorization process, we randomly swap and reassign color of a pixel to minimize a cost function that measures color consistency to its neighborhood and intensity‐to‐color correlation captured in the joint histogram. We tested our algorithm extensively on many natural objects and our user study confirms that our results are more vivid and natural compared to results from previous techniques. SouYoung Jin, Ho-Jin Choi, Yu-Wing Tai |
Comput. Graph. Forum | 3 |
| 2014 | Introduction to the special issue on visual understanding and applications with RGB-D cameras
Zicheng Liu 0001, Michael Beetz, Daniel Cremers, Juergen Gall, Wanqing Li 0001, Dejan Pangercic, Jürgen Sturm, Yu-Wing Tai |
J. Vis. Commun. Image Represent. | 8 |
| 2014 | A Physically-Based Approach to Reflection Separation: From Physical Modeling to Constrained OptimizationabstractWe propose a physically-based approach to separate reflection using multiple polarized images with a background scene captured behind glass. The input consists of three polarized images, each captured from the same view point but with a different polarizer angle separated by 45 degrees. The output is the high-quality separation of the reflection and background layers from each of the input images. A main technical challenge for this problem is that the mixing coefficient for the reflection and background layers depends on the angle of incidence and the orientation of the plane of incidence, which are spatially varying over the pixels of an image. Exploiting physical properties of polarization for a double-surfaced glass medium, we propose a multiscale scheme which automatically finds the optimal separation of the reflection and background layers. Through experiments, we demonstrate that our approach can generate superior results to those of previous methods. Naejin Kong, Yu-Wing Tai, Joseph S. Shin |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | High-Quality Depth Map Upsampling and Completion for RGB-D CamerasabstractThis paper describes an application framework to perform high-quality upsampling and completion on noisy depth maps. Our framework targets a complementary system setup, which consists of a depth camera coupled with an RGB camera. Inspired by a recent work that uses a nonlocal structure regularization, we regularize depth maps in order to maintain fine details and structures. We extend this regularization by combining the additional high-resolution RGB input when upsampling a low-resolution depth map together with a weighting scheme that favors structure details. Our technique is also able to repair large holes in a depth map with consideration of structures and discontinuities utilizing edge information from the RGB input. Quantitative and qualitative results show that our method outperforms existing approaches for depth map upsampling and completion. We describe the complete process for this system, including device calibration, scene warping for input alignment, and even how our framework can be extended for video depth-map completion with the consideration of temporal coherence. Jaesik Park, Hyeongwoo Kim, Yu-Wing Tai, Michael S. Brown, In-So Kweon |
IEEE Trans. Image Process. | 3 |
| 2013 | Shading-Based Shape Refinement of RGB-D ImagesabstractWe present a shading-based shape refinement algorithm which uses a noisy, incomplete depth map from Kinect to help resolve ambiguities in shape-from-shading. In our framework, the partial depth information is used to overcome bas-relief ambiguity in normals estimation, as well as to assist in recovering relative albedos, which are needed to reliably estimate the lighting environment and to separate shading from albedo. This refinement of surface normals using a noisy depth map leads to high-quality 3D surfaces. The effectiveness of our algorithm is demonstrated through several challenging real-world examples. Lap-Fai Yu, Sai-Kit Yeung, Yu-Wing Tai, Stephen Lin 0001 |
CVPR | 3 |
| 2013 | Outdoor photometric stereoabstractWe introduce a framework for outdoor photometric stereo utilizing natural environmental illumination. Our framework extends beyond existing photometric stereo methods intended for laboratory environments to encompass robust outdoor operation in the real world. In this paper, we motivate our framework, describe the components of its processing pipeline, and assess its performance in synthetic experiments as well as in natural experiments including objects in outdoor environments with complex real-world illuminations. Lap-Fai Yu, Sai-Kit Yeung, Yu-Wing Tai, Demetri Terzopoulos, Tony F. Chan |
ICCP | 3 |
| 2013 | Modeling the Calibration Pipeline of the Lytro Camera for High Quality Light-Field Image ReconstructionabstractLight-field imaging systems have got much attention recently as the next generation camera model. A light-field imaging system consists of three parts: data acquisition, manipulation, and application. Given an acquisition system, it is important to understand how a light-field camera converts from its raw image to its resulting refocused image. In this paper, using the Lytro camera as an example, we describe step-by-step procedures to calibrate a raw light-field image. In particular, we are interested in knowing the spatial and angular coordinates of the micro lens array and the resampling process for image reconstruction. Since Lytro uses a hexagonal arrangement of a micro lens image, additional treatments in calibration are required. After calibration, we analyze and compare the performances of several resampling methods for image reconstruction with and without calibration. Finally, a learning based interpolation method is proposed which demonstrates a higher quality image reconstruction than previous interpolation methods including a method used in Lytro software. Donghyeon Cho, Minhaeng Lee, Sunyeong Kim, Yu-Wing Tai |
ICCV | 4 |
| 2013 | A Learning-Based Approach to Reduce JPEG Artifacts in Image MattingabstractSingle image matting techniques assume high-quality input images. The vast majority of images on the web and in personal photo collections are encoded using JPEG compression. JPEG images exhibit quantization artifacts that adversely affect the performance of matting algorithms. To address this situation, we propose a learning-based post-processing method to improve the alpha mattes extracted from JPEG images. Our approach learns a set of sparse dictionaries from training examples that are used to transfer details from high-quality alpha mattes to alpha mattes corrupted by JPEG compression. Three different dictionaries are defined to accommodate different object structure (long hair, short hair, and sharp boundaries). A back-projection criteria combined within an MRF framework is used to automatically select the best dictionary to apply on the object's local boundary. We demonstrate that our method can produces superior results over existing state-of-the-art matting algorithms on a variety of inputs and compression levels. Inchang Choi, Sunyeong Kim, Michael S. Brown, Yu-Wing Tai |
ICCV | 4 |
| 2013 | Partial Sum Minimization of Singular Values in RPCA for Low-Level VisionabstractRobust Principal Component Analysis (RPCA) via rank minimization is a powerful tool for recovering underlying low-rank structure of clean data corrupted with sparse noise/outliers. In many low-level vision problems, not only it is known that the underlying structure of clean data is low-rank, but the exact rank of clean data is also known. Yet, when applying conventional rank minimization for those problems, the objective function is formulated in a way that does not fully utilize a priori target rank information about the problems. This observation motivates us to investigate whether there is a better alternative solution when using rank minimization. In this paper, instead of minimizing the nuclear norm, we propose to minimize the partial sum of singular values. The proposed objective function implicitly encourages the target rank constraint in rank minimization. Our experimental analyses show that our approach performs better than conventional rank minimization when the number of samples is deficient, while the solutions obtained by the two approaches are almost identical when the number of samples is more than sufficient. We apply our approach to various low-level vision problems, e.g. high dynamic range imaging, photometric stereo and image alignment, and show that our results outperform those obtained by the conventional nuclear norm rank minimization method. Tae-Hyun Oh, Hyeongwoo Kim, Yu-Wing Tai, Jean-Charles Bazin, In-So Kweon |
ICCV | 3 |
| 2013 | Multiview Photometric Stereo Using Planar Mesh ParameterizationabstractWe propose a method for accurate 3D shape reconstruction using uncalibrated multiview photometric stereo. A coarse mesh reconstructed using multiview stereo is first parameterized using a planar mesh parameterization technique. Subsequently, multiview photometric stereo is performed in the 2D parameter domain of the mesh, where all geometric and photometric cues from multiple images can be treated uniformly. Unlike traditional methods, there is no need for merging view-dependent surface normal maps. Our key contribution is a new photometric stereo based mesh refinement technique that can efficiently reconstruct meshes with extremely fine geometric details by directly estimating a displacement texture map in the 2D parameter domain. We demonstrate that intricate surface geometry can be reconstructed using several challenging datasets containing surfaces with specular reflections, multiple albedos and complex topologies. Jaesik Park, Sudipta N. Sinha, Yasuyuki Matsushita, Yu-Wing Tai, In-So Kweon |
ICCV | 4 |
| 2013 | A 3D Imaging Framework Based on High-Resolution Photometric-Stereo and Low-Resolution Depth
Zheng Lu 0002, Yu-Wing Tai, Fanbo Deng, Moshe Ben-Ezra, Michael S. Brown |
Int. J. Comput. Vis. | 2 |
| 2013 | Nonlinear Camera Response Functions and Image Deblurring: Theoretical Analysis and PracticeabstractThis paper investigates the role that nonlinear camera response functions (CRFs) have on image deblurring. We present a comprehensive study to analyze the effects of CRFs on motion deblurring. In particular, we show how nonlinear CRFs can cause a spatially invariant blur to behave as a spatially varying blur. We prove that such nonlinearity can cause large errors around edges when directly applying deconvolution to a motion blurred image without CRF correction. These errors are inevitable even with a known point spread function (PSF) and with state-of-the-art regularization-based deconvolution algorithms. In addition, we show how CRFs can adversely affect PSF estimation algorithms in the case of blind deconvolution. To help counter these effects, we introduce two methods to estimate the CRF directly from one or more blurred images when the PSF is known or unknown. Our experimental results on synthetic and real images validate our analysis and demonstrate the robustness and accuracy of our approaches. Yu-Wing Tai, Sunyeong Kim, Seon Joo Kim, Feng Li 0005, Jie Yang 0002, Jingyi Yu 0001, Yasuyuki Matsushita, Michael S. Brown |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2012 | Color-Aware Regularization for Gradient Domain Image Manipulation
Fanbo Deng, Seon Joo Kim, Yu-Wing Tai, Michael S. Brown |
ACCV (4) | 3 |
| 2012 | Nonlinear camera response functions and image deblurringabstractThis paper investigates the role that nonlinear camera response functions (CRFs) have on image deblurring. In particular, we show how nonlinear CRFs can cause a spatially invariant blur to behave as a spatially varying blur. This can result in noticeable ringing artifacts when deconvolution is applied even with a known point spread function (PSF). In addition, we show how CRFs can adversely affect PSF estimation algorithms in the case of blind deconvolution. To help counter these effects, we introduce two methods to estimate the CRF directly from one or more blurred images when the PSF is known or unknown. While not as accurate as conventional CRF estimation algorithms based on multiple exposures or calibration patterns, our approach is still quite effective in improving deblurring results in situations where the CRF is unknown. Sunyeong Kim, Yu-Wing Tai, Seon Joo Kim, Michael S. Brown, Yasuyuki Matsushita |
CVPR | 2 |
| 2012 | A physically-based approach to reflection separationabstractWe propose a physically-based approach to separate reflection using multiple polarized images with a background scene captured behind glass. The input consists of three polarized images, each captured from the same view point but with a different polarizer angle separated by 45 degrees. The output is the high-quality separation of the reflection and background layers from each of the input images. A main technical challenge for this problem is that the mixing coefficient for the reflection and background layers depends on the angle of incidence and the orientation of the plane of incidence, which are spatially-varying over the pixels of an image. Exploiting physical properties of polarization for a double-surfaced glass medium, we propose an algorithm which automatically finds the optimal separation of the reflection and background layers. Thorough experiments, we demonstrate that our approach can generate superior results to those of previous methods. Naejin Kong, Yu-Wing Tai, Joseph S. Shin |
CVPR | 2 |
| 2012 | Identigram/watermark removal using cross-channel correlationabstractWe introduce a method to repair an image which has been stamped by an identigram or a watermark. Our method is based on the cross-channel correlation which assures the co-occurrence of image discontinuities and correlation of color distributions across different color channels of an image. Using blind source separation, we find the transformation of color space which separates the structures of identigram and that of the original image into two different individual color channels. To repair the image contents in the corrupted channel, we formulate the problem using Bayes' rule where the prior and the likelihood probabilities are defined based on the cross-channel correlation assumption. We compare our results with results from inpainting and texture synthesis-based hole filling techniques. Our results are pleasable for real-world examples and have the maximum PSNR for synthetic examples. Jaesik Park, Yu-Wing Tai, In-So Kweon |
CVPR | 2 |
| 2012 | Motion-aware noise filtering for deblurring of noisy and blurry imagesabstractImage noise can present a serious problem in motion deblurring. While most state-of-the-art motion deblurring algorithms can deal with small levels of noise, in many cases such as low-light imaging, the noise is large enough in the blurred image that it cannot be handled effectively by these algorithms. In this paper, we propose a technique for jointly denoising and deblurring such images that elevates the performance of existing motion deblurring algorithms. Our method takes advantage of estimated motion blur kernels to improve denoising, by constraining the denoised image to be consistent with the estimated camera motion (i.e., no high frequency noise features that do not match the motion blur). This improved denoising then leads to higher quality blur kernel estimation and deblurring performance. The two operations are iterated in this manner to obtain results superior to suppressing noise effects through regularization in deblurring or by applying denoising as a preprocess. This is demonstrated in experiments both quantitatively and qualitatively using various image examples. Yu-Wing Tai, Stephen Lin 0001 |
CVPR | 1 |
| 2012 | Video Matting Using Multi-frame Nonlocal Matting Laplacian
Inchang Choi, Minhaeng Lee, Yu-Wing Tai |
ECCV (6) | 3 |
| 2012 | A Tensor Voting Approach for Multi-view 3D Scene Flow Estimation and Refinement
Jaesik Park, Tae-Hyun Oh, Jiyoung Jung, Yu-Wing Tai, In-So Kweon |
ECCV (4) | 4 |
| 2012 | Modeling photo composition and its application to photo re-arrangementabstractWe introduce a learning based photo composition model and its application on photo re-arrangement. In contrast to previous approaches which evaluate quality of photo composition using the rule of thirds or the golden ratio, we train a normalized saliency map from visually pleasurable photos taken by professional photographers. We use Principal Component Analysis (PCA) to analyze training data and build a Gaussian mixture model (GMM) to describe the photo composition model. Our experimental results show that our approach is reliable and our trained photo composition model can be used to improve photo quality through photo re-arrangement. Jaesik Park, Joon-Young Lee, Yu-Wing Tai, In-So Kweon |
ICIP | 3 |
| 2012 | Registration Based Non-uniform Motion DeblurringabstractAbstract This paper proposes an algorithm which uses image registration to estimate a non‐uniform motion blur point spread function (PSF) caused by camera shake. Our study is based on a motion blur model which models blur effects of camera shakes using a set of planar perspective projections (i.e., homographies). This representation can fully describe motions of camera shakes in 3D which cause non‐uniform motion blurs. We transform the non‐uniform PSF estimation problem into a set of image registration problems which estimate homographies of the motion blur model one‐by‐one through the Lucas‐Kanade algorithm. We demonstrate the performance of our algorithm using both synthetic and real world examples. We also discuss the effectiveness and limitations of our algorithm for non‐uniform deblurring. Sunghyun Cho, Hojin Cho, Yu-Wing Tai, Seungyong Lee 0001 |
Comput. Graph. Forum | 3 |
| 2012 | Probabilistic cost model for nearest neighbor search in image retrieval
Kunho Kim, Mohammad Khairul Hasan, Jae-Pil Heo, Yu-Wing Tai, Sung-Eui Yoon |
Comput. Vis. Image Underst. | 4 |
| 2011 | High-resolution hyperspectral imaging via matrix factorizationabstractHyperspectral imaging is a promising tool for applications in geosensing, cultural heritage and beyond. However, compared to current RGB cameras, existing hyperspectral cameras are severely limited in spatial resolution. In this paper, we introduce a simple new technique for reconstructing a very high-resolution hyperspectral image from two readily obtained measurements: A lower-resolution hyper-spectral image and a high-resolution RGB image. Our approach is divided into two stages: We first apply an unmixing algorithm to the hyperspectral input, to estimate a basis representing reflectance spectra. We then use this representation in conjunction with the RGB input to produce the desired result. Our approach to unmixing is motivated by the spatial sparsity of the hyperspectral input, and casts the unmixing problem as the search for a factorization of the input into a basis and a set of maximally sparse coefficients. Experiments show that this simple approach performs reasonably well on both simulations and real data examples. Rei Kawakami, Yasuyuki Matsushita, John Wright 0001, Moshe Ben-Ezra, Yu-Wing Tai, Katsushi Ikeuchi |
CVPR | 5 |
| 2011 | High quality depth map upsampling for 3D-TOF camerasabstractThis paper describes an application framework to perform high quality upsampling on depth maps captured from a low-resolution and noisy 3D time-of-flight (3D-ToF) camera that has been coupled with a high-resolution RGB camera. Our framework is inspired by recent work that uses nonlocal means filtering to regularize depth maps in order to maintain fine detail and structure. Our framework extends this regularization with an additional edge weighting scheme based on several image features based on the additional high-resolution RGB input. Quantitative and qualitative results show that our method outperforms existing approaches for 3D-ToF upsampling. We describe the complete process for this system, including device calibration, scene warping for input alignment, and even how the results can be further processed using simple user markup. Jaesik Park, Hyeongwoo Kim, Yu-Wing Tai, Michael S. Brown, In-So Kweon |
ICCV | 3 |
| 2011 | Two-phase approach for multi-view object extractionabstractIn this paper, we propose an automatic method to extract a foreground object captured from multiple viewpoints. We consider the foreground object is within the visual hull of camera field of views. By exploring the multi-view geometric relationship and color measurements of the input images, we can estimate the foreground segmentations as well as their fractional boundaries. To facilitate efficient computation and high quality mattes, we adopt a two-phase approach. The first phase of our algorithm provides quick and rough binary segmentations of the foreground object using graph-cut; the second phase refines the segmentation boundaries using matting. Our result is the high quality alpha mattes of the foreground object consistently across all different viewpoints. We demonstrate the effectiveness of our method using challenging examples. Sungheum Kim, Yu-Wing Tai, Yunsu Bok, Hyeongwoo Kim, In-So Kweon |
ICIP | 2 |
| 2011 | Motion Regularization for Matting Motion Blurred ObjectsabstractThis paper addresses the problem of matting motion blurred objects from a single image. Existing single image matting methods are designed to extract static objects that have fractional pixel occupancy. This arises because the physical scene object has a finer resolution than the discrete image pixel and therefore only occupies a fraction of the pixel. For a motion blurred object, however, fractional pixel occupancy is attributed to the object’s motion over the exposure period. While conventional matting techniques can be used to matte motion blurred objects, they are not formulated in a manner that considers the object’s motion and tend to work only when the object is on a homogeneous background. We show how to obtain better alpha mattes by introducing a regularization term in the matting formulation to account for the object’s motion. In addition, we outline a method for estimating local object motion based on local gradient statistics from the original image. For the sake of completeness, we also discuss how user markup can be used to denote the local direction in lieu of motion estimation. Improvements to alpha mattes computed with our regularization are demonstrated on a variety of examples. Hai Ting Lin, Yu-Wing Tai, Michael S. Brown |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Richardson-Lucy Deblurring for Scenes under a Projective Motion PathabstractThis paper addresses how to model and correct image blur that arises when a camera undergoes ego motion while observing a distant scene. In particular, we discuss how the blurred image can be modeled as an integration of the clear scene under a sequence of planar projective transformations (i.e., homographies) that describe the camera's path. This projective motion path blur model is more effective at modeling the spatially varying motion blur exhibited by ego motion than conventional methods based on space-invariant blur kernels. To correct the blurred image, we describe how to modify the Richardson-Lucy (RL) algorithm to incorporate this new blur model. In addition, we show that our projective motion RL algorithm can incorporate state-of-the-art regularization priors to improve the deblurred results. The projective motion path blur model, along with the modified RL algorithm, is detailed, together with experimental results demonstrating its overall effectiveness. Statistical analysis on the algorithm's convergence properties and robustness to noise is also provided. Yu-Wing Tai, Ping Tan 0002, Michael S. Brown |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | High-Quality Reflection Separation Using Polarized ImagesabstractIn this paper, we deal with a problem of separating the effect of reflection from images captured behind glass. The input consists of multiple polarized images captured from the same view point but with different polarizer angles. The output is the high quality separation of the reflection layer and the background layer from the images. We formulate this problem as a constrained optimization problem and propose a framework that allows us to fully exploit the mutually exclusive image information in our input data. We test our approach on various images and demonstrate that our approach can generate good reflection separation results. Naejin Kong, Yu-Wing Tai, Joseph S. Shin |
IEEE Trans. Image Process. | 2 |
| 2011 | Semantic colorization with internet imagesabstractColorization of a grayscale photograph often requires considerable effort from the user, either by placing numerous color scribbles over the image to initialize a color propagation algorithm, or by looking for a suitable reference image from which color information can be transferred. Even with this user supplied data, colorized images may appear unnatural as a result of limited user skill or inaccurate transfer of colors. To address these problems, we propose a colorization system that leverages the rich image content on the internet. As input, the user needs only to provide a semantic text label and segmentation cues for major foreground objects in the scene. With this information, images are downloaded from photo sharing websites and filtered to obtain suitable reference images that are reliable for color transfer to the given grayscale photo. Different image colorizations are generated from the various reference images, and a graphical user interface is provided to easily select the desired result. Our experiments and user study demonstrate the greater effectiveness of this system in comparison to previous techniques. Alex Yong Sang Chia, Shaojie Zhuo, Raj Kumar Gupta, Yu-Wing Tai, Siu-Yeung Cho, Ping Tan 0002, Stephen Lin 0001 |
ACM Trans. Graph. | 4 |
| 2010 | A framework for ultra high resolution 3D imagingabstractWe present an imaging framework to acquire 3D surface scans at ultra high-resolutions (exceeding 600 samples per mm2). Our approach couples a standard structured-light setup and photometric stereo using a large-format ultra-high-resolution camera. While previous approaches have employed similar hybrid imaging systems to fuse positional data with surface normals, what is unique to our approach is the significant asymmetry in the resolution between the low-resolution geometry and the ultra-high-resolution surface normals. To deal with these resolution differences, we propose a multi-resolution surface reconstruction scheme that propagates the low-resolution geometric constraints through the different frequency bands while gradually fusing in the high-resolution photometric stereo data. In addition, to deal with the ultra-high-resolution images, our surface reconstruction is performed in a patch-wise fashion and additional boundary constraints are used to ensure patch coherence. Based on this multi-resolution reconstruction scheme, our imaging framework can produce 3D scans that show exceptionally detailed 3D surfaces far exceeding existing technologies. Zheng Lu 0002, Yu-Wing Tai, Moshe Ben-Ezra, Michael S. Brown |
CVPR | 2 |
| 2010 | Coded exposure imaging for projective motion deblurringabstractWe propose a method for deblurring of spatially variant object motion. A principal challenge of this problem is how to estimate the point spread function (PSF) of the spatially variant blur. Based on the projective motion blur model of, we present a blur estimation technique that jointly utilizes a coded exposure camera and simple user interactions to recover the PSF. With this spatially variant PSF, objects that exhibit projective motion can be effectively de-blurred. We validate this method with several challenging image examples. Yu-Wing Tai, Naejin Kong, Stephen Lin 0001, Joseph S. Shin |
CVPR | 1 |
| 2010 | Super resolution using edge prior and single image detail synthesisabstractEdge-directed image super resolution (SR) focuses on ways to remove edge artifacts in upsampled images. Under large magnification, however, textured regions become blurred and appear homogenous, resulting in a super-resolution image that looks unnatural. Alternatively, learning-based SR approaches use a large database of exemplar images for “hallucinating” detail. The quality of the upsampled image, especially about edges, is dependent on the suitability of the training images. This paper aims to combine the benefits of edge-directed SR with those of learning-based SR. In particular, we propose an approach to extend edge-directed super-resolution to include detail from an image/texture example provided by the user (e.g., from the Internet). A significant benefit of our approach is that only a single exemplar image is required to supply the missing detail - strong edges are obtained in the SR image even if they are not present in the example image due to the combination of the edge-directed approach. In addition, we can achieve quality results at very large magnification, which is often problematic for both edge-directed and learning-based approaches. Yu-Wing Tai, Shuaicheng Liu, Michael S. Brown, Stephen Lin 0001 |
CVPR | 1 |
| 2010 | Colorization for Single Image Super Resolution
Shuaicheng Liu, Michael S. Brown, Seon Joo Kim, Yu-Wing Tai |
ECCV (6) | 4 |
| 2010 | Interactive content-aware zooming
Pierre-Yves Laffont, Jong Yun Jun, Christian Wolf 0001, Yu-Wing Tai, Khalid Idrissi, George Drettakis, Sung-Eui Yoon |
Graphics Interface | 4 |
| 2010 | Correction of Spatially Varying Image and Video Motion Blur Using a Hybrid CameraabstractWe describe a novel approach to reduce spatially varying motion blur in video and images using a hybrid camera system. A hybrid camera is a standard video camera that is coupled with an auxiliary low-resolution camera sharing the same optical path but capturing at a significantly higher frame rate. The auxiliary video is temporally sharper but at a lower resolution, while the lower frame-rate video has higher spatial resolution but is susceptible to motion blur. Our deblurring approach uses the data from these two video streams to reduce spatially varying motion blur in the high-resolution camera with a technique that combines both deconvolution and super-resolution. Our algorithm also incorporates a refinement of the spatially varying blur kernels to further improve results. Our approach can reduce motion blur from the high-resolution video as well as estimate new high-resolution frames at a higher frame rate. Experimental results on a variety of inputs demonstrate notable improvement over current state-of-the-art methods in image/video deblurring. Yu-Wing Tai, Hao Du 0004, Michael S. Brown, Stephen Lin 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2009 | Single image defocus map estimation using local contrast priorabstractImage defocus estimation is useful for several applications including deblurring, blur magnification, measuring image quality, and depth of field segmentation. In this paper, we present a simple yet effective approach for estimating a defocus blur map based on the relationship of the contrast to the image gradient in a local image region. We call this relationship the local contrast prior. The advantage of our approach is that it does not require filter banks or frequency decomposition of the input image; instead we only need to compare local gradient profiles with the local contrast. We discuss the idea behind the local contrast prior and demonstrate its effectiveness on a variety of experiments. Yu-Wing Tai, Michael S. Brown |
ICIP | 1 |
| 2008 | Image/video deblurring using a hybrid cameraabstractWe propose a novel approach to reduce spatially varying motion blur using a hybrid camera system that simultaneously captures high-resolution video at a low-frame rate together with low-resolution video at a high-frame rate. Our work is inspired by Ben-Ezra and Nayar who introduced the hybrid camera idea for correcting global motion blur for a single still image. We broaden the scope of the problem to address spatially varying blur as well as video imagery. We also reformulate the correction process to use more information available in the hybrid camera system, as well as iteratively refine spatially varying motion extracted from the low-resolution high-speed camera. We demonstrate that our approach achieves superior results over existing work and can be extended to deblurring of moving objects. Yu-Wing Tai, Hao Du 0004, Michael S. Brown, Stephen Lin 0001 |
CVPR | 1 |
| 2008 | Texture amendment: reducing texture distortion in constrained parameterizationabstractConstrained parameterization is an effective way to establish texture coordinates between a 3D surface and an existing image or photograph. A known drawback to constrained parameterization is visual distortion that arises when the 3D geometry is mismatched to highly textured image regions. This paper introduces an approach to reduce visual distortion by expanding image regions via texture synthesis to better fit the 3D geometry. The result is a new amended texture that maintains the essence of the input texture image but exhibits significantly less distortion when mapped onto the 3D model. Yu-Wing Tai, Michael S. Brown, Chi-Keung Tang, Harry Shum |
ACM Trans. Graph. | 1 |
| 2007 | Robust Estimation of Texture Flow via Dense Feature SamplingabstractTexture flow estimation is a valuable step in a variety of vision related tasks, including texture analysis, image segmentation, shape-from-texture and texture remapping. This paper describes a novel and effective technique to estimate texture flow in an image given a small example patch. The key idea consists of extracting a dense set of features from the example patch where discrete orientations are encapsulated into the feature vector such that rotation can be simulated as a linear shift of the vector. This dense feature space is then compressed by PCA and clustered using EM to produce a set of small set of principal features. Obtaining these principal features at varying image scales, we can compute the per-pixel scale and orientation likelihoods for the distorted texture. The final texture flow estimation is formulated as the MAP solution of a labeling Markov network which is solved using belief propagation. Experimental results on both synthetic and real images demonstrate good results even for highly distorted examples. Yu-Wing Tai, Michael S. Brown, Chi-Keung Tang |
CVPR | 1 |
| 2007 | Soft Color Segmentation and Its ApplicationsabstractWe propose an automatic approach to soft color segmentation, which produces soft color segments with appropriate amount of overlapping and transparency essential to synthesizing natural images for a wide range of image-based applications. While many state-of-the-art and complex techniques are excellent at partitioning an input image to facilitate deriving a semantic description of the scene, to achieve seamless image synthesis, we advocate to a segmentation approach designed to maintain spatial and color coherence among soft segments while preserving discontinuities, by assigning to each pixel a set of soft labels corresponding to their respective color distributions. We optimize a global objective function which simultaneously exploits the reliability given by global color statistics and flexibility of local image compositing, leading to an image model where the global color statistics of an image is represented by a Gaussian Mixture Model (GMM), while the color of a pixel is explained by a local color mixture model where the weights are defined by the soft labels to the elements of the converged GMM. Transparency is naturally introduced in our probabilistic framework which infers an optimal mixture of colors at an image pixel. To adequately consider global and local information in the same framework, an alternating optimization scheme is proposed to iteratively solve for the global and local model parameters. Our method is fully automatic, and is shown to converge to a good optimal solution. We perform extensive evaluation and comparison, and demonstrate that our method achieves good image synthesis results for image-based applications such as image matting, color transfer, image deblurring, and image colorization. Yu-Wing Tai, Jiaya Jia, Chi-Keung Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2006 | Perceptually-Inspired and Edge-Directed Color Image Super-ResolutionabstractInspired by multi-scale tensor voting, a computational framework for perceptual grouping and segmentation, we propose an edge-directed technique for color image superresolution given a single low-resolution color image. Our multi-scale technique combines the advantages of edgedirected, reconstruction-based and learning-based methods, and is unique in two ways. First, we consider simultaneously all the three color channels in our multi-scale tensor voting framework to produce a multi-scale edge representation to guide the process of high-resolution color image reconstruction, which is subject to the back projection constraint. Fine details are inferred without noticeable blurry or ringing artifacts. Second, the inference of highresolution curves is achieved by multi-scale tensor voting, using the dense voting field as an edge-preserving smoothness prior which is derived geometrically without any timeconsuming learning procedure. Qualitative and quantitative results indicate that our method produces convincing results in complex test cases typically used by state-of-theart image super-resolution techniques. Yu-Wing Tai, Wai-Shun Tong, Chi-Keung Tang |
CVPR (2) | 1 |
| 2006 | Video Repairing under Variable Illumination Using Cyclic MotionsabstractThis paper presents a complete system capable of synthesizing a large number of pixels that are missing due to occlusion or damage in an uncalibrated input video. These missing pixels may correspond to the static background or cyclic motions of the captured scene. Our system employs user-assisted video layer segmentation, while the main processing in video repair is fully automatic. The input video is first decomposed into the color and illumination videos. The necessary temporal consistency is maintained by tensor voting in the spatio-temporal domain. Missing colors and illumination of the background are synthesized by applying image repairing. Finally, the occluded motions are inferred by spatio-temporal alignment of collected samples at multiple scales. We experimented on our system with some difficult examples with variable illumination, where the capturing camera can be stationary or in motion. Jiaya Jia, Yu-Wing Tai, Tai-Pang Wu, Chi-Keung Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2005 | Local Color Transfer via Probabilistic Segmentation by Expectation-MaximizationabstractWe address the problem of regional color transfer between two natural images by probabilistic segmentation. We use a new expectation-maximization (EM) scheme to impose both spatial and color smoothness to infer natural connectivity among pixels. Unlike previous work, our method takes local color information into consideration, and segment image with soft region boundaries for seamless color transfer and compositing. Our modified EM method has two advantages in color manipulation: first, subject to different levels of color smoothness in image space, our algorithm produces an optimal number of regions upon convergence, where the color statistics in each region can be adequately characterized by a component of a Gaussian mixture model (GMM). Second, we allow a pixel to fall in several regions according to our estimated probability distribution in the EM step, resulting in a transparency-like ratio for compositing different regions seamlessly. Hence, natural color transition across regions can be achieved, where the necessary intra-region and inter-region smoothness are enforced without losing original details. We demonstrate results on a variety of applications including image deblurring, enhanced color transfer, and colorizing gray scale images. Comparisons with previous methods are also presented. Yu-Wing Tai, Jiaya Jia, Chi-Keung Tang |
CVPR (1) | 1 |
| 2004 | Video Repairing: Inference of Foreground and Background under Severe Occlusion
Jiaya Jia, Tai-Pang Wu, Yu-Wing Tai, Chi-Keung Tang |
CVPR (1) | 3 |