EDBT 2026 Demo / reviewers in the wild / expert
Stephen Lin 0001
dblp:55/4755-1 · also Steve Lin 0001
· DBLP profile ↗
176ranked-venue papers
3as first author
31since 2021 · last 2026
0000-0002-5616-558XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 137 · 3 first-author · 20 since 2021Artificial intelligence and machine learning · 119 · 3 first-author · 29 since 2021Applied, interdisciplinary, general and emerging computing · 8Human-computer interaction and ubiquitous computing · 4Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Study of Finetuning Video Transformers for Multi-view Geometry TasksabstractThis paper presents an investigation of vision transformer learning for multi-view geometry tasks, such as optical flow estimation, by fine-tuning video foundation models. Unlike previous methods that involve custom architectural designs and task-specific pretraining, our research finds that general-purpose models pretrained on videos can be readily transferred to multi-view problems with minimal adaptation. The core insight is that general-purpose attention between patches learns temporal and spatial information for geometric reasoning. We demonstrate that appending a linear decoder to the Transformer backbone produces satisfactory results, and iterative refinement can further elevate performance to state-of-the-art levels. This conceptually simple approach achieves top cross-dataset generalization results for optical flow estimation with end-point error (EPE) of 0.69, 1.78, and 3.15 on the Sintel clean, Sintel final, and KITTI datasets, respectively. Our method additionally establishes a new record on the online test benchmark with EPE values of 0.79, 1.88, and F1 value of 3.79. Applications to 3D depth estimation and stereo matching also show strong performance, illustrating the versatility of video-pretrained models in addressing geometric vision tasks. Huimin Wu 0001, Kwang-Ting Cheng, Stephen Lin 0001, Zhirong Wu |
AAAI | 3 |
| 2025 | Associative TransformerabstractEmerging from the pairwise attention in conventional Transformers, there is a growing interest in sparse attention mechanisms that align more closely with localized, contextual learning in the biological brain. Existing studies such as the Coordination method employ iterative cross-attention mechanisms with a bottleneck to enable the sparse association of inputs. However, these methods are parameter inefficient and fail in more complex relational reasoning tasks. To this end, we propose Associative Transformer (AiT) to enhance the association among sparsely attended input tokens, improving parameter efficiency and performance in various vision tasks such as classification and relational reasoning. AiT leverages a learnable explicit memory comprising specialized priors that guide bottleneck attentions to facilitate the extraction of diverse localized tokens. Moreover, AiT employs an associative memory-based token reconstruction using a Hopfield energy function. The extensive empirical experiments demonstrate that AiT requires significantly fewer parameters and attention layers outperforming a broad range of sparse Transformer models. Additionally, AiT outperforms the SOTA sparse Transformer models including the Coordination method on the Sort-of-CLEVR dataset. Yuwei Sun, Hideya Ochiai, Zhirong Wu, Stephen Lin 0001, Ryota Kanai |
CVPR | 4 |
| 2025 | Hyperspherical Dataset Distillation via Contrastive Embedding AlignmentabstractDataset distillation (DD) has emerged as a promising research direction to alleviate the computational burden of training models on large datasets. By distilling large datasets into compact representations, DD enables models trained on the distilled data to achieve performance on par with those trained on the original datasets. Due to its simplicity and efficiency, distribution-based DD has become the dominant approach in DD research, as it synthesizes data to approximate the distribution of real data. However, current distribution-based DD methods often capture ambiguous semantics, resulting in the loss of class-specific semantics in the distilled data. To this end, we propose HypErspherical data distillation via contRastive embedding Alignment (HERA), aiming to distill uniform class-level semantics from original data to distilled one. In particular, we model this process by von Mises-Fisher (vMF) statistics, and use prototypical contrastive fashion to facilitate the semantic alignment. Furthermore, owing that the condensation can also be portrayed by vMF, we also align the diversity of the two hyperspherical distributions to ensure the future training stability on the distilled data. Empirical studies validate the effectiveness of our approach, highlighting our approach in improving distribution-based DD. Shuoxi Zhang, Hanpeng Liu, Stephen Lin 0001, Kun He 0001 |
ICME | 3 |
| 2025 | VASA-3D: Lifelike Audio-Driven Gaussian Head Avatars from a Single ImageabstractWe propose VASA-3D, an audio-driven, single-shot 3D head avatar generator. This research tackles two major challenges: capturing the subtle expression details present in real human faces, and reconstructing an intricate 3D head avatar from a single portrait image. To accurately model expression details, VASA-3D leverages the motion latent of VASA-1, a method that yields exceptional realism and vividness in 2D talking heads. A critical element of our work is translating this motion latent to 3D, which is accomplished by devising a 3D head model that is conditioned on the motion latent. Customization of this model to a single image is achieved through an optimization framework that employs numerous video frames of the reference head synthesized from the input image. The optimization takes various training losses robust to artifacts and limited pose coverage in the generated training data. Our experiment shows that VASA-3D produces realistic 3D talking heads that cannot be achieved by prior art, and it supports the online generation of 512x512 free-viewpoint videos at up to 75 FPS, facilitating more immersive engagements with lifelike 3D avatars. Sicheng Xu, Jiaolong Yang, Yu Deng 0006, Stephen Lin 0001, Baining Guo |
NeurIPS | 6 |
| 2025 | GMConv: Modulating Effective Receptive Fields for Convolutional KernelsabstractIn convolutional neural networks (CNNs), the convolutions are conventionally performed using a square kernel with a fixed $N \times N$ receptive field (RF). However, what matters most to the network is the effective receptive field (ERF), which indicates the extent to which input pixels contribute to an output pixel. Inspired by the property that ERFs typically exhibit a Gaussian distribution, we propose a Gaussian Mask convolutional kernel (GMConv). Specifically, GMConv utilizes the Gaussian function to generate a concentric symmetry mask that is placed over the kernel to refine the RF. We analyze the RFs of CNN kernels in different CNN layers and evaluate our approach through extensive experiments on image classification and object detection tasks. Over several tasks and standard base models, our approach compares favorably against the standard convolution. For instance, using GMConv for AlexNet and ResNet-50, the top-1 accuracy on ImageNet classification is boosted by 0.98% and 0.85%, respectively. Chao Li 0068, Stephen Lin 0001, Kun He 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | You Only Need Less Attention at Each Stage in Vision TransformersabstractThe advent of Vision Transformers (ViTs) marks a substantial paradigm shift in the realm of computer vision. ViTs capture the global information of images through self attention modules, which perform dot product computations among patchified image tokens. While self-attention modules empower ViTs to capture long-range dependencies, the computational complexity grows quadratically with the number of tokens, which is a major hindrance to the practical application of ViTs. Moreover, the self-attention mechanism in deep ViTs is also susceptible to the attention saturation issue. Accordingly, we argue against the necessity of computing the attention scores in every layer, and we propose the Less-Attention Vision Transformer (LaViT), which computes only a few attention operations at each stage and calculates the subsequent feature alignments in other layers via attention transformations that leverage the previously calculated attention scores. This novel approach can mitigate two primary issues plaguing traditional self attention modules: the heavy computational burden and attention saturation. Our proposed architecture offers superior efficiency and ease of implementation, merely requiring matrix multiplications that are highly optimized in contemporary deep learning frameworks. Moreover, our architecture demonstrates exceptional performance across various vision tasks including classification, detection and segmentation. Shuoxi Zhang, Hanpeng Liu, Stephen Lin 0001, Kun He 0001 |
CVPR | 3 |
| 2024 | Unifying Feature and Cost Aggregation with Transformers for Semantic and Visual CorrespondenceabstractThis paper introduces a Transformer-based integrative feature and cost aggregation network designed for dense matching tasks. In the context of dense matching, many works benefit from one of two forms of aggregation: feature aggregation, which pertains to the alignment of similar features, or cost aggregation, a procedure aimed at instilling coherence in the flow estimates across neighboring pixels. In this work, we first show that feature aggregation and cost aggregation exhibit distinct characteristics and reveal the potential for substantial benefits stemming from the judicious use of both aggregation processes. We then introduce a simple yet effective architecture that harnesses self- and cross-attention mechanisms to show that our approach unifies feature aggregation and cost aggregation and effectively harnesses the strengths of both techniques. Within the proposed attention layers, the features and cost volume both complement each other, and the attention layers are interleaved through a coarse-to-fine design to further promote accurate correspondence estimation. Finally at inference, our network produces multi-scale predictions, computes their confidence scores, and selects the most confident flow for final prediction. Our framework is evaluated on standard benchmarks for semantic matching, and also applied to geometric matching, where we show that our approach achieves significant improvements compared to existing methods. Sunghwan Hong, Seokju Cho, Seungryong Kim, Stephen Lin 0001 |
ICLR | 4 |
| 2023 | Randomized Quantization: A Generic Augmentation for Data Agnostic Self-supervised LearningabstractSelf-supervised representation learning follows a paradigm of withholding some part of the data and tasking the network to predict it from the remaining part. Among many techniques, data augmentation lies at the core for creating the information gap. Towards this end, masking has emerged as a generic and powerful tool where content is withheld along the sequential dimension, e.g., spatial in images, temporal in audio, and syntactic in language. In this paper, we explore the orthogonal channel dimension for generic data augmentation by exploiting precision redundancy. The data for each channel is quantized through a non-uniform quantizer, with the quantized value sampled randomly within randomly sampled quantization bins. From another perspective, quantization is analogous to channel-wise masking, as it removes the information within each bin, but preserves the information across bins. Our approach significantly surpasses existing generic data augmentation methods, while showing on par performance against modality-specific augmentations. We comprehensively evaluate our approach on vision, audio, 3D point clouds, as well as the DABS benchmark which is comprised of various data modalities. The code is available at https://github.com/microsoft/random_quantize. Huimin Wu 0001, Chenyang Lei, Xiao Sun 0001, Peng-Shuai Wang, Qifeng Chen 0001, Kwang-Ting Cheng, Stephen Lin 0001, Zhirong Wu |
ICCV | 7 |
| 2023 | Global Context NetworksabstractThe non-local network (NLNet) presents a pioneering approach for capturing long-range dependencies within an image, via aggregating query-specific global context to each query position. However, through a rigorous empirical analysis, we have found that the global contexts modeled by the non-local network are almost the same for different query positions. In this paper, we take advantage of this finding to create a simplified network based on a query-independent formulation, which maintains the accuracy of NLNet but with significantly less computation. We further replace the one-layer transformation function of the non-local block by a two-layer bottleneck, which further reduces the parameter number considerably. The resulting network element, called the global context (GC) block, effectively models global context in a lightweight manner, allowing it to be applied at multiple layers of a backbone network to form a global context network (GCNet). Experiments show that GCNet generally outperforms NLNet on major benchmarks for various recognition tasks. The code and network configurations are available at https://github.com/xvjiarui/GCNet. Yue Cao 0001, Stephen Lin 0001, Fangyun Wei, Han Hu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | A Simple Multi-Modality Transfer Learning Baseline for Sign Language TranslationabstractThis paper proposes a simple transfer learning baseline for sign language translation. Existing sign language datasets (e.g. PHOENIX-2014T, CSL-Daily) contain only about 10 K-20K pairs of sign videos, gloss annotations and texts, which are an order of magnitude smaller than typical parallel data for training spoken language translation models. Data is thus a bottleneck for training effective sign language translation models. To mitigate this problem, we propose to progressively pretrain the model from general- domain datasets that include a large amount of external supervision to within-domain datasets. Concretely, we pretrain the sign-to-gloss visual network on the general domain of human actions and the within-domain of a sign-to-gloss dataset, and pretrain the gloss-to-text translation network on the general domain of a multilingual corpus and the within-domain of a gloss-to-text corpus. The joint model is fine-tuned with an additional module named the visual-language mapper that connects the two networks. This simple baseline surpasses the previous state-of-the-art results on two sign language translation benchmarks, demonstrating the effectiveness of transfer learning. With its simplicity and strong performance, this approach can serve as a solid baseline for future research. Fangyun Wei, Xiao Sun 0001, Zhirong Wu, Stephen Lin 0001 |
CVPR | 5 |
| 2022 | Video Swin TransformerabstractThe vision community is witnessing a modeling shift from CNNs to Transformers, where pure Transformer architectures have attained top accuracy on the major video recognition benchmarks. These video models are all built on Transformer layers that globally connect patches across the spatial and temporal dimensions. In this paper, we instead advocate an inductive bias of locality in video Transformers, which leads to a better speed-accuracy trade-off compared to previous approaches which compute self-attention globally even with spatial-temporal factorization. The locality of the proposed video architecture is realized by adapting the Swin Transformer designed for the image domain, while continuing to leverage the power of pre-trained image models. Our approach achieves state-of-the-art accuracy on a broad range of video recognition benchmarks, including on action recognition (84.9 top-l accuracy on Kinetics-400 and 85.9 top-l accuracy on Kinetics-600 with ~20× less pre-training data and ~3× smaller model size) and temporal modeling (69.6 top-l accuracy on Something-Something v2). Yue Cao 0001, Yixuan Wei, Zheng Zhang 0022, Stephen Lin 0001, Han Hu 0001 |
CVPR | 6 |
| 2022 | Cross-Model Pseudo-Labeling for Semi-Supervised Action RecognitionabstractSemi-supervised action recognition is a challenging but important task due to the high cost of data annotation. A common approach to this problem is to assign unlabeled data with pseudo-labels, which are then used as additional supervision in training. Typically in recent work, the pseudo-labels are obtained by training a model on the labeled data, and then using confident predictions from the model to teach itself. In this work, we propose a more effective pseudo-labeling scheme, called Cross-Model Pseudo-Labeling (CMPL). Concretely, we introduce a lightweight auxiliary network in addition to the primary backbone, and ask them to predict pseudo-labels for each other. We observe that, due to their different structural biases, these two models tend to learn complementary representations from the same video clips. Each model can thus benefit from its counterpart by utilizing cross-model predictions as supervision. Experiments on different data partition protocols demonstrate the significant improvement of our framework over existing alternatives. For example, CMPL achieves 17.6% and 25.1% Top-1 accuracy on Kinetics-400 and UCF-101 using only the RGB modality and 1% labeled data, outperforming our baseline model, FixMatch [17], by 9.0% and 10.3%, respectively.11Project page is at https://justimyhxu.github.io/projects/cmpl/. Yinghao Xu 0001, Fangyun Wei, Xiao Sun 0001, Ceyuan Yang, Yujun Shen, Bo Dai 0002, Bolei Zhou, Stephen Lin 0001 |
CVPR | 8 |
| 2022 | Cost Aggregation with 4D Convolutional Swin Transformer for Few-Shot Segmentation
Sunghwan Hong, Seokju Cho, Jisu Nam, Stephen Lin 0001, Seungryong Kim |
ECCV (29) | 4 |
| 2022 | Unsupervised Learning of Efficient Geometry-Aware Neural Articulated Representations
Atsuhiro Noguchi, Xiao Sun 0001, Stephen Lin 0001, Tatsuya Harada |
ECCV (17) | 3 |
| 2022 | Bringing Rolling Shutter Images Alive with Dual Reversed Distortion
Zhihang Zhong, Mingdeng Cao, Xiao Sun 0001, Zhirong Wu, Zhongyi Zhou, Yinqiang Zheng, Stephen Lin 0001, Imari Sato |
ECCV (7) | 7 |
| 2022 | Animation from Blur: Multi-modal Blur Decomposition with Motion Guidance
Zhihang Zhong, Xiao Sun 0001, Zhirong Wu, Yinqiang Zheng, Stephen Lin 0001, Imari Sato |
ECCV (19) | 5 |
| 2022 | Could Giant Pre-trained Image Models Extract Universal Representations?abstractFrozen pretrained models have become a viable alternative to the pretraining-then-finetuning paradigm for transfer learning. However, with frozen models there are relatively few parameters available for adapting to downstream tasks, which is problematic in computer vision where tasks vary significantly in input/output format and the type of information that is of value. In this paper, we present a study of frozen pretrained models when applied to diverse and representative computer vision tasks, including object detection, semantic segmentation and video action recognition. From this empirical analysis, our work answers the questions of what pretraining task fits best with this frozen setting, how to make the frozen setting more flexible to various downstream tasks, and the effect of larger model sizes. We additionally examine the upper bound of performance using a giant frozen pretrained model with 3 billion parameters (SwinV2-G) and find that it reaches competitive performance on a varied set of major benchmarks with only one shared frozen base network: 60.0 box mAP and 52.2 mask mAP on COCO object detection test-dev, 57.6 val mIoU on ADE20K semantic segmentation, and 81.7 top-1 accuracy on Kinetics-400 action recognition. With this work, we hope to bring greater attention to this promising path of freezing pretrained image models. Yutong Lin, Zheng Zhang 0022, Han Hu 0001, Nanning Zheng 0001, Stephen Lin 0001, Yue Cao 0001 |
NeurIPS | 6 |
| 2021 | Learning Monocular Depth in Dynamic Scenes via Instance-Aware Projection ConsistencyabstractWe present an end-to-end joint training framework that explicitly models 6-DoF motion of multiple dynamic objects, ego-motion, and depth in a monocular camera setup without supervision. Our technical contributions are three-fold. First, we highlight the fundamental difference between inverse and forward projection while modeling the individual motion of each rigid object, and propose a geometrically correct projection pipeline using a neural forward projection module. Second, we design a unified instance-aware photometric and geometric consistency loss that holistically imposes self-supervisory signals for every background and object region. Lastly, we introduce a general-purpose auto-annotation scheme using any off-the-shelf instance segmentation and optical flow models to produce video instance segmentation maps that will be utilized as input to our training pipeline. These proposed elements are validated in a detailed ablation study. Through extensive experiments conducted on the KITTI and Cityscapes dataset, our framework is shown to outperform the state-of-the-art depth and motion estimation methods. Our code, dataset, and models are publicly available. Seokju Lee, Sunghoon Im 0001, Stephen Lin 0001, In-So Kweon |
AAAI | 3 |
| 2021 | Distilling Localization for Self-Supervised Representation LearningabstractRecent progress in contrastive learning has revolutionized unsupervised representation learning. Concretely, multiple views (augmentations) from the same image are encouraged to map to close embeddings, while views from different images are pulled apart.In this paper, through visualizing and diagnosing classification errors, we observe that current contrastive models are ineffective at localizing the foreground object, limiting their ability to extract discriminative high-level features. This is due to the fact that view generation process considers pixels in an image uniformly.To address this problem, we propose a data-driven approach for learning invariance to backgrounds. It first estimates foreground saliency in images and then creates augmentations by copy-and-pasting the foreground onto a variety of back-grounds. The learning still follows an instance discrimination approach, so that the representation is trained to disregard background content and focus on the foreground. We study a variety of saliency estimation methods, and find that most methods lead to improvements for contrastive learning. With this approach, significant performance is achieved for self-supervised learning on ImageNet classification, and also for object detection on PASCAL VOC and MSCOCO. Nanxuan Zhao, Zhirong Wu, Rynson W. H. Lau, Stephen Lin 0001 |
AAAI | 4 |
| 2021 | Propagate Yourself: Exploring Pixel-Level Consistency for Unsupervised Visual Representation LearningabstractContrastive learning methods for unsupervised visual representation learning have reached remarkable levels of transfer performance. We argue that the power of contrastive learning has yet to be fully unleashed, as current methods are trained only on instance-level pretext tasks, leading to representations that may be sub-optimal for downstream tasks requiring dense pixel predictions. In this paper, we introduce pixel-level pretext tasks for learning dense feature representations. The first task directly applies contrastive learning at the pixel level. We additionally propose a pixel-to-propagation consistency task that produces better results, even surpassing the state-of-the-art approaches by a large margin. Specifically, it achieves 60.2 AP, 41.4 / 40.5 mAP and 77.2 mIoU when transferred to Pascal VOC object detection (C4), COCO object detection (FPN / C4) and Cityscapes semantic segmentation using a ResNet-50 backbone network, which are 2.6 AP, 0.8 / 1.0 mAP and 1.0 mIoU better than the previous best methods built on instance-level contrastive learning. Moreover, the pixel-level pretext tasks are found to be effective for pre-training not only regular backbone networks but also head networks used for dense downstream tasks, and are complementary to instance-level contrastive methods. These results demonstrate the strong potential of defining pretext tasks at the pixel level, and suggest a new path forward in unsupervised visual representation learning. Code is available at https://github.com/zdaxie/PixPro. Zhenda Xie, Yutong Lin, Zheng Zhang 0022, Yue Cao 0001, Stephen Lin 0001, Han Hu 0001 |
CVPR | 5 |
| 2021 | Instance Localization for Self-Supervised Detection PretrainingabstractPrior research on self-supervised learning has led to considerable progress on image classification, but often with degraded transfer performance on object detection. The objective of this paper is to advance self-supervised pretrained models specifically for object detection. Based on the inherent difference between classification and detection, we propose a new self-supervised pretext task, called instance localization. Image instances are pasted at various locations and scales onto background images. The pretext task is to predict the instance category given the composited images as well as the foreground bounding boxes. We show that integration of bounding boxes into pretraining promotes better task alignment and architecture alignment for transfer learning. In addition, we propose an augmentation method on the bounding boxes to further enhance the feature alignment. As a result, our model becomes weaker at Imagenet semantic classification but stronger at image patch localization, with an overall stronger pretrained model for object detection. Experimental results demonstrate that our approach yields state-of-the-art transfer learning results for object detection on PASCAL VOC and MSCOCO1. Ceyuan Yang, Zhirong Wu, Bolei Zhou, Stephen Lin 0001 |
CVPR | 4 |
| 2021 | Cross-Iteration Batch NormalizationabstractA well-known issue of Batch Normalization is its significantly reduced effectiveness in the case of small mini-batch sizes. When a mini-batch contains few examples, the statistics upon which the normalization is defined cannot be reliably estimated from it during a training iteration. To address this problem, we present Cross-Iteration Batch Normalization (CBN), in which examples from multiple recent iterations are jointly utilized to enhance estimation quality. A challenge of computing statistics over multiple iterations is that the network activations from different iterations are not comparable to each other due to changes in network weights. We thus compensate for the network weight changes via a proposed technique based on Taylor polynomials, so that the statistics can be accurately estimated and batch normalization can be effectively applied. On object detection and image classification with small mini-batch sizes, CBN is found to outperform the original batch normalization and a direct calculation of statistics over previous iterations without the proposed compensation technique. Code is available at https://aka.ms/cbn. Zhuliang Yao, Yue Cao 0001, Shuxin Zheng, Gao Huang 0001, Stephen Lin 0001 |
CVPR | 5 |
| 2021 | Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsabstractThis paper presents a new vision Transformer, called Swin Transformer, that capably serves as a general-purpose backbone for computer vision. Challenges in adapting Transformer from language to vision arise from differences between the two domains, such as large variations in the scale of visual entities and the high resolution of pixels in images compared to words in text. To address these differences, we propose a hierarchical Transformer whose representation is computed with Shifted windows. The shifted windowing scheme brings greater efficiency by limiting self-attention computation to non-overlapping local windows while also allowing for cross-window connection. This hierarchical architecture has the flexibility to model at various scales and has linear computational complexity with respect to image size. These qualities of Swin Transformer make it compatible with a broad range of vision tasks, including image classification (87.3 top-1 accuracy on ImageNet-1K) and dense prediction tasks such as object detection (58.7 box AP and 51.1 mask AP on COCO test-dev) and semantic segmentation (53.5 mIoU on ADE20K val). Its performance surpasses the previous state-of-the-art by a large margin of +2.7 box AP and +2.6 mask AP on COCO, and +3.2 mIoU on ADE20K, demonstrating the potential of Transformer-based models as vision backbones. The hierarchical design and the shifted window approach also prove beneficial for all-MLP architectures. The code and models are publicly available at https://github.com/microsoft/Swin-Transformer. Yutong Lin, Yue Cao 0001, Han Hu 0001, Yixuan Wei, Zheng Zhang 0022, Stephen Lin 0001, Baining Guo |
ICCV | 7 |
| 2021 | Neural Articulated Radiance FieldabstractWe present Neural Articulated Radiance Field (NARF), a novel deformable 3D representation for articulated objects learned from images. While recent advances in 3D implicit representation have made it possible to learn models of complex objects, learning pose-controllable representations of articulated objects remains a challenge, as current methods require 3D shape supervision and are unable to render appearance. In formulating an implicit representation of 3D articulated objects, our method considers only the rigid transformation of the most relevant object part in solving for the radiance field at each 3D location. In this way, the proposed method represents pose-dependent changes without significantly increasing the computational complexity. NARF is fully differentiable and can be trained from images with pose annotations. Moreover, through the use of an autoencoder, it can learn appearance variations over multiple instances of an object class. Experiments show that the proposed method is efficient and can generalize well to novel poses. The code is available for research purposes at https://github.com/nogu-atsu/NARF. Atsuhiro Noguchi, Xiao Sun 0001, Stephen Lin 0001, Tatsuya Harada |
ICCV | 3 |
| 2021 | What Makes Instance Discrimination Good for Transfer Learning?
Nanxuan Zhao, Zhirong Wu, Rynson W. H. Lau, Stephen Lin 0001 |
ICLR | 4 |
| 2021 | The Emergence of Objectness: Learning Zero-shot Segmentation from VideosabstractHumans can easily detect and segment moving objects simply by observing how they move, even without knowledge of object semantics. Inspired by this, we develop a zero-shot unsupervised approach for learning object segmentations. The model comprises two visual pathways: an appearance pathway that segments individual RGB images into coherent object regions, and a motion pathway that predicts the flow vector for each region between consecutive video frames. The two pathways jointly reconstruct a new representation called segment flow. This decoupled representation of appearance and motion is trained in a self-supervised manner to reconstruct one frame from another.When pretrained on an unlabeled video corpus, the model can be useful for a variety of applications, including 1) primary object segmentation from a single image in a zero-shot fashion; 2) moving object segmentation from a video with unsupervised test-time adaptation; 3) image semantic segmentation by supervised fine-tuning on a labeled image dataset. We demonstrate encouraging experimental results on all of these tasks using pretrained models. Runtao Liu, Zhirong Wu, Stella X. Yu, Stephen Lin 0001 |
NeurIPS | 4 |
| 2021 | Aligning Pretraining for Detection via Object-Level Contrastive LearningabstractImage-level contrastive representation learning has proven to be highly effective as a generic model for transfer learning. Such generality for transfer learning, however, sacrifices specificity if we are interested in a certain downstream task. We argue that this could be sub-optimal and thus advocate a design principle which encourages alignment between the self-supervised pretext task and the downstream task. In this paper, we follow this principle with a pretraining method specifically designed for the task of object detection. We attain alignment in the following three aspects: 1) object-level representations are introduced via selective search bounding boxes as object proposals; 2) the pretraining network architecture incorporates the same dedicated modules used in the detection pipeline (e.g. FPN); 3) the pretraining is equipped with object detection properties such as object-level translation invariance and scale invariance. Our method, called Selective Object COntrastive learning (SoCo), achieves state-of-the-art results for transfer performance on COCO detection using a Mask R-CNN framework. Code is available at https://github.com/hologerry/SoCo. Fangyun Wei, Zhirong Wu, Han Hu 0001, Stephen Lin 0001 |
NeurIPS | 5 |
| 2021 | Bootstrap Your Object Detector via Mixed TrainingabstractWe introduce MixTraining, a new training paradigm for object detection that can improve the performance of existing detectors for free. MixTraining enhances data augmentation by utilizing augmentations of different strengths while excluding the strong augmentations of certain training samples that may be detrimental to training. In addition, it addresses localization noise and missing labels in human annotations by incorporating pseudo boxes that can compensate for these errors. Both of these MixTraining capabilities are made possible through bootstrapping on the detector, which can be used to predict the difficulty of training on a strong augmentation, as well as to generate reliable pseudo boxes thanks to the robustness of neural networks to labeling error. MixTraining is found to bring consistent improvements across various detectors on the COCO dataset. In particular, the performance of Faster R-CNN~\cite{ren2015faster} with a ResNet-50~\cite{he2016deep} backbone is improved from 41.7 mAP to 44.0 mAP, and the accuracy of Cascade-RCNN~\cite{cai2018cascade} with a Swin-Small~\cite{liu2021swin} backbone is raised from 50.9 mAP to 52.8 mAP. Mengde Xu, Zheng Zhang 0022, Fangyun Wei, Yutong Lin, Yue Cao 0001, Stephen Lin 0001, Han Hu 0001, Xiang Bai |
NeurIPS | 6 |
| 2021 | Deep Depth from Uncalibrated Small Motion ClipabstractWe propose a novel approach to infer a high-quality depth map from a set of images with small viewpoint variations. In general, techniques for depth estimation from small motion consist of camera pose estimation and dense reconstruction. In contrast to prior approaches that recover scene geometry and camera motions using pre-calibrated cameras, we introduce in this paper a self-calibrating bundle adjustment method tailored for small motion which enables computation of camera poses without the need for camera calibration. For dense depth reconstruction, we present a convolutional neural network called DPSNet (Deep Plane Sweep Network) whose design is inspired by best practices of traditional geometry-based approaches. Rather than directly estimating depth or optical flow correspondence from image pairs as done in many previous deep learning methods, DPSNet takes a plane sweep approach that involves building a cost volume from deep features using the plane sweep algorithm, regularizing the cost volume, and regressing the depth map from the cost volume. The cost volume is constructed using a differentiable warping process that allows for end-to-end training of the network. Through the effective incorporation of conventional multiview stereo concepts within a deep learning framework, the proposed method achieves state-of-the-art results on a variety of challenging datasets. Sunghoon Im 0001, Hyowon Ha, Hae-Gon Jeon, Stephen Lin 0001, In-So Kweon |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Dense Cross-Modal Correspondence Estimation With the Deep Self-Correlation DescriptorabstractWe present the deep self-correlation (DSC) descriptor for establishing dense correspondences between images taken under different imaging modalities, such as different spectral ranges or lighting conditions. We encode local self-similar structure in a pyramidal manner that yields both more precise localization ability and greater robustness to non-rigid image deformations. Specifically, DSC first computes multiple self-correlation surfaces with randomly sampled patches over a local support window, and then builds pyramidal self-correlation surfaces through average pooling on the surfaces. The feature responses on the self-correlation surfaces are then encoded through spatial pyramid pooling in a log-polar configuration. To better handle geometric variations such as scale and rotation, we additionally propose the geometry-invariant DSC (GI-DSC) that leverages multi-scale self-correlation computation and canonical orientation estimation. In contrast to descriptors based on deep convolutional neural networks (CNNs), DSC and GI-DSC are training-free (i.e., handcrafted descriptors), are robust to cross-modality, and generalize well to various modality variations. Extensive experiments demonstrate the state-of-the-art performance of DSC and GI-DSC on challenging cases of cross-modal image pairs having photometric and/or geometric variations. Seungryong Kim, Dongbo Min, Stephen Lin 0001, Kwanghoon Sohn |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | ACP++: Action Co-Occurrence Priors for Human-Object Interaction DetectionabstractA common problem in the task of human-object interaction (HOI) detection is that numerous HOI classes have only a small number of labeled examples, resulting in training sets with a long-tailed distribution. The lack of positive labels can lead to low classification accuracy for these classes. Towards addressing this issue, we observe that there exist natural correlations and anti-correlations among human-object interactions. In this paper, we model the correlations as action co-occurrence matrices and present techniques to learn these priors and leverage them for more effective training, especially on rare classes. The efficacy of our approach is demonstrated experimentally, where the performance of our approach consistently improves over the state-of-the-art methods on both of the two leading HOI detection benchmark datasets, HICO-Det and V-COCO. Dong-Jin Kim 0003, Xiao Sun 0001, Jinsoo Choi, Stephen Lin 0001, In-So Kweon |
IEEE Trans. Image Process. | 4 |
| 2020 | Message from the 3DV 2020 Program Chairs
Adrian Hilton 0001, Zuzana Kukelova, Stephen Lin 0001, Jun Sato |
3DV | 3 |
| 2020 | Leveraging Multi-View Image Sets for Unsupervised Intrinsic Image Decomposition and Highlight SeparationabstractWe present an unsupervised approach for factorizing object appearance into highlight, shading, and albedo layers, trained by multi-view real images. To do so, we construct a multi-view dataset by collecting numerous customer product photos online, which exhibit large illumination variations that make them suitable for training of reflectance separation and can facilitate object-level decomposition. The main contribution of our approach is a proposed image representation based on local color distributions that allows training to be insensitive to the local misalignments of multi-view images. In addition, we present a new guidance cue for unsupervised training that exploits synergy between highlight separation and intrinsic image decomposition. Over a broad range of objects, our technique is shown to yield state-of-the-art results for both of these tasks. Renjiao Yi, Ping Tan 0002, Stephen Lin 0001 |
AAAI | 3 |
| 2020 | Single Image Reflection Removal Through Cascaded RefinementabstractWe address the problem of removing undesirable reflections from a single image captured through a glass surface, which is an ill-posed, challenging but practically important problem for photo enhancement. Inspired by iterative structure reduction for hidden community detection in social networks, we propose an Iterative Boost Convolutional LSTM Network (IBCLN) that enables cascaded prediction for reflection removal. IBCLN is a cascaded network that iteratively refines the estimates of transmission and reflection layers in a manner that they can boost the prediction quality to each other, and information across steps of the cascade is transferred using an LSTM. The intuition is that the transmission is the strong, dominant structure while the reflection is the weak, hidden structure. They are complementary to each other in a single image and thus a better estimate and reduction on one side from the original image leads to a more accurate estimate on the other side. To facilitate training over multiple cascade steps, we employ LSTM to address the vanishing gradient problem, and propose residual reconstruction loss as further training guidance. Besides, we create a dataset of real-world images with reflection and ground-truth transmission layers to mitigate the problem of insufficient data. Comprehensive experiments demonstrate that the proposed method can effectively remove reflections in real and synthetic images compared with state-of-the-art reflection removal methods. Chao Li 0068, Yixiao Yang, Kun He 0001, Stephen Lin 0001, John E. Hopcroft |
CVPR | 4 |
| 2020 | A Transductive Approach for Video Object SegmentationabstractSemi-supervised video object segmentation aims to separate a target object from a video sequence, given the mask in the first frame. Most of current prevailing methods utilize information from additional modules trained in other domains like optical flow and instance segmentation, and as a result they do not compete with other methods on common ground. To address this issue, we propose a simple yet strong transductive method, in which additional modules, datasets, and dedicated architectural designs are not needed. Our method takes a label propagation approach where pixel labels are passed forward based on feature similarity in an embedding space. Different from other propagation methods, ours diffuses temporal information in a holistic manner which take accounts of long-term object appearance. In addition, our method requires few additional computational overhead, and runs at a fast ~37 fps speed. Our single model with a vanilla ResNet50 backbone achieves an overall score of 72.3% on the DAVIS 2017 validation set and 63.1% on the test set. This simple yet high performing and efficient method can serve as a solid baseline that facilitates future research. Code and models are available at https://github.com/ microsoft/transductive-vos.pytorch. Zhirong Wu, Houwen Peng, Stephen Lin 0001 |
CVPR | 4 |
| 2020 | Detecting Human-Object Interactions with Action Co-occurrence Priors
Dong-Jin Kim 0003, Xiao Sun 0001, Jinsoo Choi, Stephen Lin 0001, In-So Kweon |
ECCV (21) | 4 |
| 2020 | Object-Based Illumination Estimation with Rendering-Aware Neural Networks
Yue Dong 0001, Stephen Lin 0001, Xin Tong 0001 |
ECCV (15) | 4 |
| 2020 | Point-Set Anchors for Object Detection, Instance Segmentation and Pose Estimation
Fangyun Wei, Xiao Sun 0001, Hongyang Li 0001, Jingdong Wang 0001, Stephen Lin 0001 |
ECCV (10) | 5 |
| 2020 | Spatially Adaptive Inference with Stochastic Feature Sampling and Interpolation
Zhenda Xie, Zheng Zhang 0022, Xizhou Zhu, Gao Huang 0001, Stephen Lin 0001 |
ECCV (1) | 5 |
| 2020 | Dense RepPoints: Representing Visual Objects with Dense Point Sets
Ze Yang 0003, Yinghao Xu 0001, Zheng Zhang 0022, Raquel Urtasun, Liwei Wang 0001, Stephen Lin 0001, Han Hu 0001 |
ECCV (21) | 7 |
| 2020 | Disentangled Non-local Neural Networks
Minghao Yin, Zhuliang Yao, Yue Cao 0001, Xiu Li 0001, Zheng Zhang 0022, Stephen Lin 0001, Han Hu 0001 |
ECCV (15) | 6 |
| 2020 | SRNet: Improving Generalization in 3D Human Pose Estimation with a Split-and-Recombine Approach
Ailing Zeng, Xiao Sun 0001, Fuyang Huang, Minhao Liu, Qiang Xu 0001, Stephen Lin 0001 |
ECCV (14) | 6 |
| 2020 | Deformable Kernels: Adapting Effective Receptive Fields for Object Deformation
Xizhou Zhu, Stephen Lin 0001, Jifeng Dai |
ICLR | 3 |
| 2020 | RepPoints v2: Verification Meets Regression for Object DetectionabstractVerification and regression are two general methodologies for prediction in neural networks. Each has its own strengths: verification can be easier to infer accurately, and regression is more efficient and applicable to continuous target variables. Hence, it is often beneficial to carefully combine them to take advantage of their benefits. In this paper, we take this philosophy to improve state-of-the-art object detection, specifically by RepPoints. Though RepPoints provides high performance, we find that its heavy reliance on regression for object localization leaves room for improvement. We introduce verification tasks into the localization prediction of RepPoints, producing RepPoints v2, which proves consistent improvements of about 2.0 mAP over the original RepPoints on COCO object detection benchmark using different backbones and training methods. RepPoints v2 also achieves 52.1 mAP on the COCO \texttt{test-dev} by a single model. Moreover, we show that the proposed approach can more generally elevate other object detection frameworks as well as applications such as instance segmentation. Zheng Zhang 0022, Yue Cao 0001, Liwei Wang 0001, Stephen Lin 0001, Han Hu 0001 |
NeurIPS | 5 |
| 2020 | Discrete-Continuous Transformation Matching for Dense Semantic CorrespondenceabstractTechniques for dense semantic correspondence have provided limited ability to deal with the geometric variations that commonly exist between semantically similar images. While variations due to scale and rotation have been examined, there is a lack of practical solutions for more complex deformations such as affine transformations because of the tremendous size of the associated solution space. To address this problem, we present a discrete-continuous transformation matching (DCTM) framework where dense affine transformation fields are inferred through a discrete label optimization in which the labels are iteratively updated via continuous regularization. In this way, our approach draws solutions from the continuous space of affine transformations in a manner that can be computed efficiently through constant-time edge-aware filtering and a proposed affine-varying CNN-based descriptor. Furthermore, leveraging correspondence consistency and confidence-guided filtering in each iteration facilitates the convergence of our method. Experimental results show that this model outperforms the state-of-the-art methods for dense semantic correspondence on various benchmarks and applications. Seungryong Kim, Dongbo Min, Stephen Lin 0001, Kwanghoon Sohn |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Angle-Closure Detection in Anterior Segment OCT Based on Multilevel Deep NetworkabstractIrreversible visual impairment is often caused by primary angle-closure glaucoma, which could be detected via anterior segment optical coherence tomography (AS-OCT). In this paper, an automated system based on deep learning is presented for angle-closure detection in AS-OCT images. Our system learns a discriminative representation from training data that captures subtle visual cues not modeled by handcrafted features. A multilevel deep network is proposed to formulate this learning, which utilizes three particular AS-OCT regions based on clinical priors: 1) the global anterior segment structure; 2) local iris region; and 3) anterior chamber angle (ACA) patch. In our method, a sliding window-based detector is designed to localize the ACA region, which addresses ACA detection as a regression task. Then, three parallel subnetworks are applied to extract AS-OCT representations for the global image and at clinically relevant local regions. Finally, the extracted deep features of these subnetworks are concatenated into one fully connected layer to predict the angle-closure detection result. In the experiments, our system is shown to surpass previous detection methods and other deep learning systems on two clinical AS-OCT datasets. Huazhu Fu, Yanwu Xu 0001, Stephen Lin 0001, Damon Wing Kee Wong, Mani Baskaran, Meenakshi Mahesh, Tin Aung, Jiang Liu 0001 |
IEEE Trans. Cybern. | 3 |
| 2019 | Visuomotor Understanding for Representation Learning of Driving Scenes
Seokju Lee, Junsik Kim 0001, Tae-Hyun Oh, Yongseop Jeong, Donggeun Yoo, Stephen Lin 0001, In-So Kweon |
BMVC | 6 |
| 2019 | Deformable ConvNets V2: More Deformable, Better ResultsabstractThe superior performance of Deformable Convolutional Networks arises from its ability to adapt to the geometric variations of objects. Through an examination of its adaptive behavior, we observe that while the spatial support for its neural features conforms more closely than regular ConvNets to object structure, this support may nevertheless extend well beyond the region of interest, causing features to be influenced by irrelevant image content. To address this problem, we present a reformulation of Deformable ConvNets that improves its ability to focus on pertinent image regions, through increased modeling power and stronger training. The modeling power is enhanced through a more comprehensive integration of deformable convolution within the network, and by introducing a modulation mechanism that expands the scope of deformation modeling. To effectively harness this enriched modeling capability, we guide network training via a proposed feature mimicking scheme that helps the network to learn features that reflect the object focus and classification power of R-CNN features. With the proposed contributions, this new version of Deformable ConvNets yields significant performance gains over the original model and produces leading results on the COCO benchmark for object detection and instance segmentation. Xizhou Zhu, Han Hu 0001, Stephen Lin 0001, Jifeng Dai |
CVPR | 3 |
| 2019 | Local Relation Networks for Image RecognitionabstractThe convolution layer has been the dominant feature extractor in computer vision for years. However, the spatial aggregation in convolution is basically a pattern matching process that applies fixed filters which are inefficient at modeling visual elements with varying spatial distributions. This paper presents a new image feature extractor, called the local relation layer, that adaptively determines aggregation weights based on the compositional relationship of local pixel pairs. With this relational approach, it can composite visual elements into higher-level entities in a more efficient manner that benefits semantic inference. A network built with local relation layers, called the Local Relation Network (LR-Net), is found to provide greater modeling capacity than its counterpart built with regular convolution on large-scale recognition tasks such as ImageNet classification. Han Hu 0001, Zheng Zhang 0022, Zhenda Xie, Stephen Lin 0001 |
ICCV | 4 |
| 2019 | RepPoints: Point Set Representation for Object DetectionabstractModern object detectors rely heavily on rectangular bounding boxes, such as anchors, proposals and the final predictions, to represent objects at various recognition stages. The bounding box is convenient to use but provides only a coarse localization of objects and leads to a correspondingly coarse extraction of object features. In this paper, we present RepPoints (representative points), a new finer representation of objects as a set of sample points useful for both localization and recognition. Given ground truth localization and recognition targets for training, RepPoints learn to automatically arrange themselves in a manner that bounds the spatial extent of an object and indicates semantically significant local areas. They furthermore do not require the use of anchors to sample a space of bounding boxes. We show that an anchor-free object detector based on RepPoints can be as effective as the state-of-the-art anchor-based detection methods, with 46.5 AP and 67.4 AP50 on the COCO test-dev detection benchmark, using ResNet-101 model. Code is available at https://github.com/microsoft/RepPoints. Ze Yang 0003, Shaohui Liu, Han Hu 0001, Liwei Wang 0001, Stephen Lin 0001 |
ICCV | 5 |
| 2019 | An Empirical Study of Spatial Attention Mechanisms in Deep NetworksabstractAttention mechanisms have become a popular component in deep neural networks, yet there has been little examination of how different influencing factors and methods for computing attention from these factors affect performance. Toward a better general understanding of attention mechanisms, we present an empirical study that ablates various spatial attention elements within a generalized attention formulation, encompassing the dominant Transformer attention as well as the prevalent deformable convolution and dynamic convolution modules. Conducted on a variety of applications, the study yields significant findings about spatial attention in deep networks, some of which run counter to conventional understanding. For example, we find that the query and key content comparison in Transformer attention is negligible for self-attention, but vital for encoder-decoder attention. A proper combination of deformable convolution with key content only saliency achieves the best accuracy-efficiency tradeoff in self-attention. Our results suggest that there exists much room for improvement in the design of attention mechanisms. Xizhou Zhu, Dazhi Cheng, Zheng Zhang 0022, Stephen Lin 0001, Jifeng Dai |
ICCV | 4 |
| 2019 | DPSNet: End-to-end Deep Plane Sweep Stereo
Sunghoon Im 0001, Hae-Gon Jeon, Stephen Lin 0001, In-So Kweon |
ICLR (Poster) | 3 |
| 2019 | Learning Residual Flow as Dynamic Motion from Stereo VideosabstractWe present a method for decomposing the 3D scene flow observed from a moving stereo rig into stationary scene elements and dynamic object motion. Our unsupervised learning framework jointly reasons about the camera motion, optical flow, and 3D motion of moving objects. Three cooperating networks predict stereo matching, camera motion, and residual flow, which represents the flow component due to object motion and not from camera motion. Based on rigid projective geometry, the estimated stereo depth is used to guide the camera motion estimation, and the depth and camera motion are used to guide the residual flow estimation. We also explicitly estimate the 3D scene flow of dynamic objects based on the residual flow and scene depth. Experiments on the KITTI dataset demonstrate the effectiveness of our approach and show that our method outperforms other state-of-the-art algorithms on the optical flow and visual odometry tasks. Seokju Lee, Sunghoon Im 0001, Stephen Lin 0001, In-So Kweon |
IROS | 3 |
| 2019 | FCSS: Fully Convolutional Self-Similarity for Dense Semantic CorrespondenceabstractWe present a descriptor, called fully convolutional self-similarity (FCSS), for dense semantic correspondence. Unlike traditional dense correspondence approaches for estimating depth or optical flow, semantic correspondence estimation poses additional challenges due to intra-class appearance and shape variations among different instances within the same object or scene category. To robustly match points across semantically similar images, we formulate FCSS using local self-similarity (LSS), which is inherently insensitive to intra-class appearance variations. LSS is incorporated through a proposed convolutional self-similarity (CSS) layer, where the sampling patterns and the self-similarity measure are jointly learned in an end-to-end and multi-scale manner. Furthermore, to address shape variations among different object instances, we propose a convolutional affine transformer (CAT) layer that estimates explicit affine transformation fields at each pixel to transform the sampling patterns and corresponding receptive fields. As training data for semantic correspondence is rather limited, we propose to leverage object candidate priors provided in most existing datasets and also correspondence consistency between object pairs to enable weakly-supervised learning. Experiments demonstrate that FCSS significantly outperforms conventional handcrafted descriptors and CNN-based descriptors on various benchmarks. Seungryong Kim, Dongbo Min, Bumsub Ham, Stephen Lin 0001, Kwanghoon Sohn |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | Physically-Based Simulation of Cosmetics via Intrinsic Image Decomposition with Facial PriorsabstractWe present a physically-based approach for simulating makeup in face images. The key idea is to decompose the face image into intrinsic image layers - namely albedo, diffuse shading, and specular highlights - which are each differently affected by cosmetics, and then manipulate each layer according to corresponding models of reflectance. Accurate intrinsic image decompositions for faces are obtained with the help of human face priors, including statistics on skin reflectance and facial geometry. The intrinsic image layers are then transformed in appearance according to measured optical properties of cosmetics and proposed adaptations of physically-based reflectance models. With this approach, realistic results are generated in a manner that preserves the personal appearance features and lighting conditions of the target face while not requiring detailed geometric and reflectance measurements. We demonstrate this technique on various forms of cosmetics including foundation, blush, lipstick, and eye shadow. Results on both images and videos exhibit a close approximation to ground truth and compare favorably to existing techniques. Chen Li 0031, Kun Zhou 0001, Hsiang-Tao Wu, Stephen Lin 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | Makeup Removal via Bidirectional Tunable De-Makeup NetworkabstractWe present a deep learning-based method for removing makeup effects (de-makeup) in a face image. This problem poses a major challenge due to obscuring of the underlying facial features by cosmetics, which is very important in multimedia applications in the field of security, entertainment, and social networking. To address this task, we propose the bidirectional tunable de-makeup network (BTD-Net), which jointly learns the makeup process to aid in learning the de-makeup process. For tractable learning of the makeup process, which is a one-to-many mapping determined by the cosmetics that are applied, we introduce a latent variable that reflects the makeup style. This latent variable is extracted in the de-makeup process and used as a condition on the makeup process to constrain the one-to-many mapping to a specific solution. Through extensive experiments, our proposed BTD-Net is found to surpass the state-of-art techniques in estimating realistic non-makeup faces that correspond to the input makeup images. We additionally show that applications such as tuning the amount of makeup can be enhanced through the use of this method. Feng Lu 0005, Chen Li 0031, Stephen Lin 0001, Xukun Shen |
IEEE Trans. Multim. | 4 |
| 2018 | A High-Quality Denoising Dataset for Smartphone CamerasabstractThe last decade has seen an astronomical shift from imaging with DSLR and point-and-shoot cameras to imaging with smartphone cameras. Due to the small aperture and sensor size, smartphone images have notably more noise than their DSLR counterparts. While denoising for smartphone images is an active research area, the research community currently lacks a denoising image dataset representative of real noisy images from smartphone cameras with high-quality ground truth. We address this issue in this paper with the following contributions. We propose a systematic procedure for estimating ground truth for noisy images that can be used to benchmark denoising performance for smartphone cameras. Using this procedure, we have captured a dataset - the Smartphone Image Denoising Dataset (SIDD) - of~30,000 noisy images from 10 scenes under different lighting conditions using five representative smartphone cameras and generated their ground truth images. We used this dataset to benchmark a number of denoising algorithms. We show that CNN-based methods perform better when trained on our high-quality dataset than when trained using alternative strategies, such as low-ISO images used as a proxy for ground truth data. Abdelrahman Abdelhamed, Stephen Lin 0001, Michael S. Brown |
CVPR | 2 |
| 2018 | Faces as Lighting Probes via Unsupervised Deep Highlight Extraction
Renjiao Yi, Chenyang Zhu 0002, Ping Tan 0002, Stephen Lin 0001 |
ECCV (9) | 4 |
| 2018 | Multi-context Deep Network for Angle-Closure Glaucoma Screening in Anterior Segment OCT
Huazhu Fu, Yanwu Xu 0001, Stephen Lin 0001, Damon Wing Kee Wong, Mani Baskaran, Meenakshi Mahesh, Tin Aung, Jiang Liu 0001 |
MICCAI (2) | 3 |
| 2018 | CoVieW'18: The 1st Workshop and Challenge on Comprehensive Video Understanding in the WildabstractThe 1st Workshop and Challenge on Comprehensive Video Understanding in the Wild, dubbed CoVieW'18, is held in Seoul, Korea on October 22, 2018, in conjuction with ACM Multimedia 2018. The workshop aims to solve the joint and comprehensive understanding problem in untrimmed videos with a particular emphasis on joint action and scene recognition. The workshop encourages researchers to participate in joint action and scene recognition challenge in untrimmed videos and to report their results. The workshop program includes 1 keynote speech, 2 invited speakers, 6 regular and challenge papers. The developments made in the workshop will deliver a step change in a variety of video applications. Kwanghoon Sohn, Ming-Hsuan Yang 0001, Hyeran Byun, Jongwoo Lim, Gee-Sern Hsu, Stephen Lin 0001, Euntai Kim, Seungryong Kim |
ACM Multimedia | 6 |
| 2018 | Recurrent Transformer Networks for Semantic CorrespondenceabstractWe present recurrent transformer networks (RTNs) for obtaining dense correspondences between semantically similar images. Our networks accomplish this through an iterative process of estimating spatial transformations between the input images and using these transformations to generate aligned convolutional activations. By directly estimating the transformations between an image pair, rather than employing spatial transformer networks to independently normalize each individual image, we show that greater accuracy can be achieved. This process is conducted in a recursive manner to refine both the transformation estimates and the feature representations. In addition, a technique is presented for weakly-supervised training of RTNs that is based on a proposed classification loss. With RTNs, state-of-the-art performance is attained on several benchmarks for semantic correspondence. Seungryong Kim, Stephen Lin 0001, Sangryul Jeon, Dongbo Min, Kwanghoon Sohn |
NeurIPS | 2 |
| 2018 | Exposure: A White-Box Photo Post-Processing FrameworkabstractRetouching can significantly elevate the visual appeal of photos, but many casual photographers lack the expertise to do this well. To address this problem, previous works have proposed automatic retouching systems based on supervised learning from paired training images acquired before and after manual editing. As it is difficult for users to acquire paired images that reflect their retouching preferences, we present in this article a deep learning approach that is instead trained on unpaired data, namely, a set of photographs that exhibits a retouching style the user likes, which is much easier to collect. Our system is formulated using deep convolutional neural networks that learn to apply different retouching operations on an input image. Network training with respect to various types of edits is enabled by modeling these retouching operations in a unified manner as resolution-independent differentiable filters. To apply the filters in a proper sequence and with suitable parameters, we employ a deep reinforcement learning approach that learns to make decisions on what action to take next, given the current state of the image. In contrast to many deep learning systems, ours provides users with an understandable solution in the form of conventional retouching edits rather than just a “black-box” result. Through quantitative comparisons and user studies, we show that this technique generates retouching results consistent with the provided photo set. Yuanming Hu, Hao He 0011, Baoyuan Wang, Stephen Lin 0001 |
ACM Trans. Graph. | 5 |
| 2017 | FC^4: Fully Convolutional Color Constancy with Confidence-Weighted PoolingabstractImprovements in color constancy have arisen from the use of convolutional neural networks (CNNs). However, the patch-based CNNs that exist for this problem are faced with the issue of estimation ambiguity, where a patch may contain insufficient information to establish a unique or even a limited possible range of illumination colors. Image patches with estimation ambiguity not only appear with great frequency in photographs, but also significantly degrade the quality of network training and inference. To overcome this problem, we present a fully convolutional network architecture in which patches throughout an image can carry different confidence weights according to the value they provide for color constancy estimation. These confidence weights are learned and applied within a novel pooling layer where the local estimates are merged into a global solution. With this formulation, the network is able to determine what to learn and how to pool automatically from color constancy datasets without additional supervision. The proposed network also allows for end-to-end training, and achieves higher efficiency and accuracy. On standard benchmarks, our network outperforms the previous state-of-the-art while achieving 120× greater efficiency. Yuanming Hu, Baoyuan Wang, Stephen Lin 0001 |
CVPR | 3 |
| 2017 | FCSS: Fully Convolutional Self-Similarity for Dense Semantic CorrespondenceabstractWe present a descriptor, called fully convolutional self-similarity (FCSS), for dense semantic correspondence. To robustly match points among different instances within the same object class, we formulate FCSS using local self-similarity (LSS) within a fully convolutional network. In contrast to existing CNN-based descriptors, FCSS is inherently insensitive to intra-class appearance variations because of its LSS-based structure, while maintaining the precise localization ability of deep neural networks. The sampling patterns of local structure and the self-similarity measure are jointly learned within the proposed network in an end-to-end and multi-scale manner. As training data for semantic correspondence is rather limited, we propose to leverage object candidate priors provided in existing image datasets and also correspondence consistency between object pairs to enable weakly-supervised learning. Experiments demonstrate that FCSS outperforms conventional handcrafted descriptors and CNN-based descriptors on various benchmarks. Seungryong Kim, Dongbo Min, Bumsub Ham, Sangryul Jeon, Stephen Lin 0001, Kwanghoon Sohn |
CVPR | 5 |
| 2017 | Radiometric Calibration from Faces in ImagesabstractWe present a method for radiometric calibration of cameras from a single image that contains a human face. This technique takes advantage of a low-rank property that exists among certain skin albedo gradients because of the pigments within the skin. This property becomes distorted in images that are captured with a non-linear camera response function, and we perform radiometric calibration by solving for the inverse response function that best restores this low-rank property in an image. Although this work makes use of the color properties of skin pigments, we show that this calibration is unaffected by the color of scene illumination or the sensitivities of the cameras color filters. Our experiments validate this approach on a variety of images containing human faces, and show that faces can provide an important source of calibration data in images where existing radiometric calibration techniques perform poorly. Chen Li 0031, Stephen Lin 0001, Kun Zhou 0001, Katsushi Ikeuchi |
CVPR | 2 |
| 2017 | Specular Highlight Removal in Facial ImagesabstractWe present a method for removing specular highlight reflections in facial images that may contain varying illumination colors. This is accurately achieved through the use of physical and statistical properties of human skin and faces. We employ a melanin and hemoglobin based model to represent the diffuse color variations in facial skin, and utilize this model to constrain the highlight removal solution in a manner that is effective even for partially saturated pixels. The removal of highlights is further facilitated through estimation of directionally variant illumination colors over the face, which is done while taking advantage of a statistically-based approximation of facial geometry. An important practical feature of the proposed method is that the skin color model is utilized in a way that does not require color calibration of the camera. Moreover, this approach does not require assumptions commonly needed in previous highlight removal techniques, such as uniform illumination color or piecewise-constant surface colors. We validate this technique through comparisons to existing methods for removing specular highlights. Chen Li 0031, Stephen Lin 0001, Kun Zhou 0001, Katsushi Ikeuchi |
CVPR | 2 |
| 2017 | DCTM: Discrete-Continuous Transformation Matching for Semantic FlowabstractTechniques for dense semantic correspondence have provided limited ability to deal with the geometric variations that commonly exist between semantically similar images. While variations due to scale and rotation have been examined, there is a lack of practical solutions for more complex deformations such as affine transformations because of the tremendous size of the associated solution space. To address this problem, we present a discrete-continuous transformation matching (DCTM) framework where dense affine transformation fields are inferred through a discrete label optimization in which the labels are iteratively updated via continuous regularization. In this way, our approach draws solutions from the continuous space of affine transformations in a manner that can be computed efficiently through constant-time edge-aware filtering and a proposed affine-varying CNN-based descriptor. Experimental results show that this model outperforms the state-of-the-art methods for dense semantic correspondence on various benchmarks. Seungryong Kim, Dongbo Min, Stephen Lin 0001, Kwanghoon Sohn |
ICCV | 3 |
| 2017 | Face inpainting based on high-level facial attributes
Mahdi Jampour, Chen Li 0031, Lap-Fai Yu, Kun Zhou 0001, Stephen Lin 0001, Horst Bischof |
Comput. Vis. Image Underst. | 5 |
| 2017 | Object-Based Multiple Foreground Segmentation in RGBD VideoabstractWe present an RGB and Depth (RGBD) video segmentation method that takes advantage of depth data and can extract multiple foregrounds in the scene. This video segmentation is addressed as an object proposal selection problem formulated in a fully-connected graph, where a flexible number of foregrounds may be chosen. In our graph, each node represents a proposal, and the edges model intra-frame and inter-frame constraints on the solution. The proposals are selected based on an RGBD video saliency map in which depth-based features are utilized to enhance the identification of foregrounds. Experiments show that the proposed multiple foreground segmentation method outperforms related techniques, and the depth cue serves as a helpful complement to RGB features. Moreover, our method provides performance comparable to the state-of-the-art RGB video segmentation techniques on regular RGB videos with estimated depth maps. Huazhu Fu, Dong Xu 0001, Stephen Lin 0001 |
IEEE Trans. Image Process. | 3 |
| 2017 | Segmentation and Quantification for Angle-Closure Glaucoma Assessment in Anterior Segment OCTabstractAngle-closure glaucoma is a major cause of irreversible visual impairment and can be identified by measuring the anterior chamber angle (ACA) of the eye. The ACA can be viewed clearly through anterior segment optical coherence tomography (AS-OCT), but the imaging characteristics and the shapes and locations of major ocular structures can vary significantly among different AS-OCT modalities, thus complicating image analysis. To address this problem, we propose a data-driven approach for automatic AS-OCT structure segmentation, measurement, and screening. Our technique first estimates initial markers in the eye through label transfer from a hand-labeled exemplar data set, whose images are collected over different patients and AS-OCT modalities. These initial markers are then refined by using a graph-based smoothing method that is guided by AS-OCT structural information. These markers facilitate segmentation of major clinical structures, which are used to recover standard clinical parameters. These parameters can be used not only to support clinicians in making anatomical assessments, but also to serve as features for detecting anterior angle closure in automatic glaucoma screening algorithms. Experiments on Visante AS-OCT and Cirrus high-definition-OCT data sets demonstrate the effectiveness of our approach. Huazhu Fu, Yanwu Xu 0001, Stephen Lin 0001, Xiaoqin Zhang 0002, Damon Wing Kee Wong, Jiang Liu 0001, Alejandro F. Frangi, Mani Baskaran, Tin Aung |
IEEE Trans. Medical Imaging | 3 |
| 2016 | Image Deblurring Using Smartphone Inertial SensorsabstractRemoving image blur caused by camera shake is an ill-posed problem, as both the latent image and the point spread function (PSF) are unknown. A recent approach to address this problem is to record camera motion through inertial sensors, i.e., gyroscopes and accelerometers, and then reconstruct spatially-variant PSFs from these readings. While this approach has been effective for highquality inertial sensors, it has been infeasible for the inertial sensors in smartphones, which are of relatively low quality and present a number of challenging issues, including varying sensor parameters, high sensor noise, and calibration error. In this paper, we identify the issues that plague smartphone inertial sensors and propose a solution that successfully utilizes the sensor readings for image deblurring. With both the sensor data and the image itself, the proposed method is able to accurately estimate the sensor parameters online and also the spatially-variant PSFs for enhanced deblurring performance. The effectiveness of this technique is demonstrated in experiments on a popular mobile phone. With this approach, the quality of image deblurring can be appreciably raised on the most common of imaging devices. Lu Yuan 0001, Stephen Lin 0001, Ming-Hsuan Yang 0001 |
CVPR | 3 |
| 2016 | Deep Self-correlation Descriptor for Dense Cross-Modal Correspondence
Seungryong Kim, Dongbo Min, Stephen Lin 0001, Kwanghoon Sohn |
ECCV (8) | 3 |
| 2016 | Unified Depth Prediction and Intrinsic Image Decomposition from a Single Image via Joint Convolutional Neural Fields
Seungryong Kim, Kihong Park, Kwanghoon Sohn, Stephen Lin 0001 |
ECCV (8) | 4 |
| 2016 | DeepVessel: Retinal Vessel Segmentation via Deep Learning and Conditional Random Field
Huazhu Fu, Yanwu Xu 0001, Stephen Lin 0001, Damon Wing Kee Wong, Jiang Liu 0001 |
MICCAI (2) | 3 |
| 2016 | Bayesian Depth-From-Defocus With Shading ConstraintsabstractWe present a method that enhances the performance of depth-from-defocus (DFD) through the use of shading information. DFD suffers from important limitations--namely coarse shape reconstruction and poor accuracy on textureless surfaces--that can be overcome with the help of shading. We integrate both forms of data within a Bayesian framework that capitalizes on their relative strengths. Shading data, however, is challenging to accurately recover from surfaces that contain texture. To address this issue, we propose an iterative technique that utilizes depth information to improve shading estimation, which in turn is used to elevate depth estimation in the presence of textures. The shading estimation can be performed in general scenes with unknown illumination using an approximate estimate of scene lighting. With this approach, we demonstrate improvements over existing DFD techniques, as well as effective shape reconstruction of textureless surfaces. Chen Li 0031, Shuochen Su, Yasuyuki Matsushita, Kun Zhou 0001, Stephen Lin 0001 |
IEEE Trans. Image Process. | 5 |
| 2015 | Continuous Symmetric Stereo with Adaptive Outlier HandlingabstractWe present a method for symmetric stereo matching in which outliers from occlusions, texture-less regions, and repeated patterns are handled in a soft and adaptive manner. Rather than making binary outlier decisions, our model incorporates continuous-valued confidence weights that account for outlier likelihood, to promote robustness in disparity estimation. In contrast to previous outlier labeling techniques that fix the labels at the start of optimization, our method iteratively updates our outlier confidence weights as the matching results are gradually refined. By doing this, errors in an initial labeling can be rectified in the matching process. Our model is optimized in an Expectation-Maximization framework that efficiently produces continuous disparity estimates. This approach provides a good combination of accuracy and speed. Experiments show that our method compares favorably to prior outlier labeling techniques on the Middlebury benchmark, and that it can generate high-quality reconstruction for outdoor images with much more complex occlusions. Chen Li 0031, Lap-Fai Yu, Zhichao Lu, Yasuyuki Matsushita, Kun Zhou 0001, Stephen Lin 0001 |
3DV | 6 |
| 2015 | Adaptive pooling over multiple trajectory attributes for action recognitionabstractWe present a new approach for feature pooling in human action recognition. Instead of partitioning videos at predefined uniform intervals in a spatial-temporal volume as done with spatial pyramid matching, our method adaptively partitions in a pooling attribute space, defined by multiple trajectory-based cues. The pooling attributes include individual spatial and temporal coordinates of a trajectory, as well as its motion saliency, curvature, and scale. To determine partitions of the attribute space in an adaptive manner, we utilize KD-trees that separate trajectories based on their distributions within the attribute space. The generated pooling volumes are jointly utilized for action recognition via SVM weights learned by Multiple Kernel Learning. Through extensive experimentation on major benchmarks, it is shown that this adaptive pooling over multiple trajectory attributes leads to significant improvements in recognition performance. Wangjiang Zhu, Baoyuan Wang, Stephen Lin 0001 |
AVSS | 3 |
| 2015 | Object-based RGBD image co-segmentation with mutex constraintabstractWe present an object-based co-segmentation method that takes advantage of depth data and is able to correctly handle noisy images in which the common foreground object is missing. With RGBD images, our method utilizes the depth channel to enhance identification of similar foreground objects via a proposed RGBD co-saliency map, as well as to improve detection of object-like regions and provide depth-based local features for region comparison. To accurately deal with noisy images where the common object appears more than or less than once, we formulate co-segmentation in a fully-connected graph structure together with mutual exclusion (mutex) constraints that prevent improper solutions. Experiments show that this object-based RGBD co-segmentation with mutex constraints outperforms related techniques on an RGBD co-segmentation dataset, while effectively processing noisy images. Moreover, we show that this method also provides performance comparable to state-of-the-art RGB co-segmentation techniques on regular RGB images with depth maps estimated from them. Huazhu Fu, Dong Xu 0001, Stephen Lin 0001, Jiang Liu 0001 |
CVPR | 3 |
| 2015 | Data-driven depth map refinement via multi-scale sparse representationabstractDepth maps captured by consumer-level depth cameras such as Kinect are usually degraded by noise, missing values, and quantization. In this paper, we present a data-driven approach for refining degraded RAWdepth maps that are coupled with an RGB image. The key idea of our approach is to take advantage of a training set of high-quality depth data and transfer its information to the RAW depth map through multi-scale dictionary learning. Utilizing a sparse representation, our method learns a dictionary of geometric primitives which captures the correlation between high-quality mesh data, RAW depth maps and RGB images. The dictionary is learned and applied in a manner that accounts for various practical issues that arise in dictionary-based depth refinement. Compared to previous approaches that only utilize the correlation between RAW depth maps and RGB images, our method produces improved depth maps without over-smoothing. Since our approach is data driven, the refinement can be targeted to a specific class of objects by employing a corresponding training set. In our experiments, we show that this leads to additional improvements in recovering depth maps of human faces. Hyeokhyen Kwon, Yu-Wing Tai, Stephen Lin 0001 |
CVPR | 3 |
| 2015 | Simulating makeup through physics-based manipulation of intrinsic image layersabstractWe present a method for simulating makeup in a face image. To generate realistic results without detailed geometric and reflectance measurements of the user, we propose to separate the image into intrinsic image layers and alter them according to proposed adaptations of physically-based reflectance models. Through this layer manipulation, the measured properties of cosmetic products are applied while preserving the appearance characteristics and lighting conditions of the target face. This approach is demonstrated on various forms of cosmetics including foundation, blush, lipstick, and eye shadow. Experimental results exhibit a close approximation to ground truth images, without artifacts such as transferred personal features and lighting effects that degrade the results of image-based makeup transfer methods. Chen Li 0031, Kun Zhou 0001, Stephen Lin 0001 |
CVPR | 3 |
| 2015 | Unsupervised Extraction of Video Highlights via Robust Recurrent Auto-EncodersabstractWith the growing popularity of short-form video sharing platforms such as Instagram and Vine, there has been an increasing need for techniques that automatically extract highlights from video. Whereas prior works have approached this problem with heuristic rules or supervised learning, we present an unsupervised learning approach that takes advantage of the abundance of user-edited videos on social media websites such as YouTube. Based on the idea that the most significant sub-events within a video class are commonly present among edited videos while less interesting ones appear less frequently, we identify the significant sub-events via a robust recurrent auto-encoder trained on a collection of user-edited videos queried for each particular class of interest. The auto-encoder is trained using a proposed shrinking exponential loss function that makes it robust to noise in the web-crawled training data, and is configured with bidirectional long short term memory (LSTM) [5] cells to better model the temporal structure of highlight segments. Different from supervised techniques, our method can infer highlights using only a set of downloaded edited videos, without also needing their pre-edited counterparts which are rarely available online. Extensive experiments indicate the promise of our proposed solution in this challenging unsupervised setting. Huan Yang 0005, Baoyuan Wang, Stephen Lin 0001, David P. Wipf, Minyi Guo, Baining Guo |
ICCV | 3 |
| 2015 | Change-Based Image Cropping with Exclusion and Compositional Features
Jianzhou Yan, Stephen Lin 0001, Sing Bing Kang, Xiaoou Tang |
Int. J. Comput. Vis. | 2 |
| 2015 | Image Classification With Densely Sampled Image Windows and Generalized Adaptive Multiple Kernel LearningabstractWe present a framework for image classification that extends beyond the window sampling of fixed spatial pyramids and is supported by a new learning algorithm. Based on the observation that fixed spatial pyramids sample a rather limited subset of the possible image windows, we propose a method that accounts for a comprehensive set of windows densely sampled over location, size, and aspect ratio. A concise high-level image feature is derived to effectively deal with this large set of windows, and this higher level of abstraction offers both efficient handling of the dense samples and reduced sensitivity to misalignment. In addition to dense window sampling, we introduce generalized adaptive l(p)-norm multiple kernel learning (GA-MKL) to learn a robust classifier based on multiple base kernels constructed from the new image features and multiple sets of prelearned classifiers from other classes. With GA-MKL, multiple levels of image features are effectively fused, and information is shared among different classifiers. Extensive evaluation on benchmark datasets for object recognition (Caltech256 and Caltech101) and scene recognition (15Scenes) demonstrate that the proposed method outperforms the state-of-the-art under a broad range of settings. Shengye Yan, Xinxing Xu, Dong Xu 0001, Stephen Lin 0001, Xuelong Li 0001 |
IEEE Trans. Cybern. | 4 |
| 2015 | Object-Based Multiple Foreground Video Co-Segmentation via Multi-State Selection GraphabstractWe present a technique for multiple foreground video co-segmentation in a set of videos. This technique is based on category-independent object proposals. To identify the foreground objects in each frame, we examine the properties of the various regions that reflect the characteristics of foregrounds, considering the intra-video coherence of the foreground as well as the foreground consistency among the different videos in the set. Multiple foregrounds are handled via a multi-state selection graph in which a node representing a video frame can take multiple labels that correspond to different objects. In addition, our method incorporates an indicator matrix that for the first time allows accurate handling of cases with common foreground objects missing in some videos, thus preventing irrelevant regions from being misclassified as foreground objects. An iterative procedure is proposed to optimize our new objective function. As demonstrated through comprehensive experiments, this object-based multiple foreground video co-segmentation method compares well with related techniques that co-segment multiple foregrounds. Huazhu Fu, Dong Xu 0001, Stephen Lin 0001, Rabab K. Ward |
IEEE Trans. Image Process. | 4 |
| 2015 | Image based relighting using neural networksabstractWe present a neural network regression method for relighting realworld scenes from a small number of images. The relighting in this work is formulated as the product of the scene's light transport matrix and new lighting vectors, with the light transport matrix reconstructed from the input images. Based on the observation that there should exist non-linear local coherence in the light transport matrix, our method approximates matrix segments using neural networks that model light transport as a non-linear function of light source position and pixel coordinates. Central to this approach is a proposed neural network design which incorporates various elements that facilitate modeling of light transport from a small image set. In contrast to most image based relighting techniques, this regression-based approach allows input images to be captured under arbitrary illumination conditions, including light sources moved freely by hand. We validate our method with light transport data of real scenes containing complex lighting effects, and demonstrate that fewer input images are required in comparison to related techniques. Peiran Ren, Yue Dong 0001, Stephen Lin 0001, Xin Tong 0001, Baining Guo |
ACM Trans. Graph. | 3 |
| 2014 | Automatic Feature Learning to Grade Nuclear Cataracts Based on Deep Learning
Xinting Gao, Stephen Lin 0001, Tien Yin Wong |
ACCV (2) | 2 |
| 2014 | Object-Based Multiple Foreground Video Co-segmentationabstractWe present a video co-segmentation method that uses category-independent object proposals as its basic element and can extract multiple foreground objects in a video set. The use of object elements overcomes limitations of low-level feature representations in separating complex foregrounds and backgrounds. We formulate object-based co-segmentation as a co-selection graph in which regions with foreground-like characteristics are favored while also accounting for intra-video and inter-video foreground coherence. To handle multiple foreground objects, we expand the co-selection graph model into a proposed multi-state selection graph model (MSG) that optimizes the segmentations of different objects jointly. This extension into the MSG can be applied not only to our co-selection graph, but also can be used to turn any standard graph model into a multi-state selection solution that can be optimized directly by the existing energy minimization techniques. Our experiments show that our object-based multiple foreground video co-segmentation method (ObMiC) compares well to related techniques on both single and multiple foreground cases. Huazhu Fu, Dong Xu 0001, Stephen Lin 0001 |
CVPR | 4 |
| 2014 | A Learning-to-Rank Approach for Image Color EnhancementabstractWe present a machine-learned ranking approach for automatically enhancing the color of a photograph. Unlike previous techniques that train on pairs of images before and after adjustment by a human user, our method takes into account the intermediate steps taken in the enhancement process, which provide detailed information on the person's color preferences. To make use of this data, we formulate the color enhancement task as a learning-to-rank problem in which ordered pairs of images are used for training, and then various color enhancements of a novel input image can be evaluated from their corresponding rank values. From the parallels between the decision tree structures we use for ranking and the decisions made by a human during the editing process, we posit that breaking a full enhancement sequence into individual steps can facilitate training. Our experiments show that this approach compares well to existing methods for automatic color enhancement. Jianzhou Yan, Stephen Lin 0001, Sing Bing Kang, Xiaoou Tang |
CVPR | 2 |
| 2014 | Intrinsic Face Image Decomposition with Human Face Priors
Chen Li 0031, Kun Zhou 0001, Stephen Lin 0001 |
ECCV (5) | 3 |
| 2014 | Alpha Matting of Motion-Blurred Objects in Bracket Sequence Images
Heesoo Myeong, Stephen Lin 0001, Kyoung Mu Lee |
ECCV (3) | 2 |
| 2014 | A Visual Evaluation Framework for In-Home Physical RehabilitationabstractWe propose a novel method for in-home physical rehabilitation, where a user can visually evaluate his/her performance compared to that of an expert. Normalized joint coordinates extracted from the Kinect skeleton are used as features. A novel Incremental Dynamic Time Warping (IDTW) algorithm is used to align the user and expert sequences. IDTW extends the classic DTW by providing accurate comparison between incomplete (the user's) and complete (the expert's) sequences while significantly reducing the computational time. Instead of providing a single measurement, the proposed method maps the IDTW measurements to a color-coded skeleton frame. Different colors on the limbs provide the user with an easy-to-interpret evaluation of how he or she is performing. Preliminary analysis involving different users and exercises and comparisons against the classic DTW algorithm show the effectiveness of the proposed method. Naimul Mefraz Khan, Stephen Lin 0001, Ling Guan, Baining Guo |
ISM | 2 |
| 2014 | Optic Cup Segmentation for Glaucoma Detection Using Low-Rank Superpixel Representation
Yanwu Xu 0001, Lixin Duan, Stephen Lin 0001, Damon Wing Kee Wong, Tien Yin Wong, Jiang Liu 0001 |
MICCAI (1) | 3 |
| 2014 | Acquisition of High Spatial and Spectral Resolution Video with a Hybrid Camera System
Chenguang Ma, Xun Cao, Xin Tong 0001, Qionghai Dai, Stephen Lin 0001 |
Int. J. Comput. Vis. | 5 |
| 2013 | Bayesian Depth-from-Defocus with Shading ConstraintsabstractWe present a method that enhances the performance of depth-from-defocus (DFD) through the use of shading information. DFD suffers from important limitations - namely coarse shape reconstruction and poor accuracy on texture less surfaces - that can be overcome with the help of shading. We integrate both forms of data within a Bayesian framework that capitalizes on their relative strengths. Shading data, however, is challenging to recover accurately from surfaces that contain texture. To address this issue, we propose an iterative technique that utilizes depth information to improve shading estimation, which in turn is used to elevate depth estimation in the presence of textures. With this approach, we demonstrate improvements over existing DFD techniques, as well as effective shape reconstruction of texture less surfaces. Chen Li 0031, Shuochen Su, Yasuyuki Matsushita, Kun Zhou 0001, Stephen Lin 0001 |
CVPR | 5 |
| 2013 | Learning the Change for Automatic Image CroppingabstractImage cropping is a common operation used to improve the visual quality of photographs. In this paper, we present an automatic cropping technique that accounts for the two primary considerations of people when they crop: removal of distracting content, and enhancement of overall composition. Our approach utilizes a large training set consisting of photos before and after cropping by expert photographers to learn how to evaluate these two factors in a crop. In contrast to the many methods that exist for general assessment of image quality, ours specifically examines differences between the original and cropped photo in solving for the crop parameters. To this end, several novel image features are proposed to model the changes in image content and composition when a crop is applied. Our experiments demonstrate improvements of our method over recent cropping algorithms on a broad range of images. Jianzhou Yan, Stephen Lin 0001, Sing Bing Kang, Xiaoou Tang |
CVPR | 2 |
| 2013 | Shading-Based Shape Refinement of RGB-D ImagesabstractWe present a shading-based shape refinement algorithm which uses a noisy, incomplete depth map from Kinect to help resolve ambiguities in shape-from-shading. In our framework, the partial depth information is used to overcome bas-relief ambiguity in normals estimation, as well as to assist in recovering relative albedos, which are needed to reliably estimate the lighting environment and to separate shading from albedo. This refinement of surface normals using a noisy depth map leads to high-quality 3D surfaces. The effectiveness of our algorithm is demonstrated through several challenging real-world examples. Lap-Fai Yu, Sai-Kit Yeung, Yu-Wing Tai, Stephen Lin 0001 |
CVPR | 4 |
| 2013 | Semantically-Based Human Scanpath Estimation with HMMsabstractWe present a method for estimating human scan paths, which are sequences of gaze shifts that follow visual attention over an image. In this work, scan paths are modeled based on three principal factors that influence human attention, namely low-level feature saliency, spatial position, and semantic content. Low-level feature saliency is formulated as transition probabilities between different image regions based on feature differences. The effect of spatial position on gaze shifts is modeled as a Levy flight with the shifts following a 2D Cauchy distribution. To account for semantic content, we propose to use a Hidden Markov Model (HMM) with a Bag-of-Visual-Words descriptor of image regions. An HMM is well-suited for this purpose in that 1) the hidden states, obtained by unsupervised learning, can represent latent semantic concepts, 2) the prior distribution of the hidden states describes visual attraction to the semantic concepts, and 3) the transition probabilities represent human gaze shift patterns. The proposed method is applied to task-driven viewing processes. Experiments and analysis performed on human eye gaze data verify the effectiveness of this method. Dong Xu 0001, Qingming Huang, Wen Li 0001, Min Xu 0001, Stephen Lin 0001 |
ICCV | 6 |
| 2013 | Automatic Grading of Nuclear Cataracts from Slit-Lamp Lens Images Using Group Sparsity Regression
Yanwu Xu 0001, Xinting Gao, Stephen Lin 0001, Damon Wing Kee Wong, Jiang Liu 0001, Dong Xu 0001, Ching Yu Cheng, Carol Yim-lui Cheung, Tien Yin Wong |
MICCAI (2) | 3 |
| 2013 | Efficient Reconstruction-Based Optic Cup Localization for Glaucoma Screening
Yanwu Xu 0001, Stephen Lin 0001, Damon Wing Kee Wong, Jiang Liu 0001, Dong Xu 0001 |
MICCAI (3) | 2 |
| 2013 | Single-Image Vignetting Correction from Gradient Distribution SymmetriesabstractWe present novel techniques for single-image vignetting correction based on symmetries of two forms of image gradients: semicircular tangential gradients (SCTG) and radial gradients (RG). For a given image pixel, an SCTG is an image gradient along the tangential direction of a circle centered at the presumed optical center and passing through the pixel. An RG is an image gradient along the radial direction with respect to the optical center. We observe that the symmetry properties of SCTG and RG distributions are closely related to the vignetting in the image. Based on these symmetry properties, we develop an automatic optical center estimation algorithm by minimizing the asymmetry of SCTG distributions, and also present two methods for vignetting estimation based on minimizing the asymmetry of RG distributions. In comparison to prior approaches to single-image vignetting correction, our methods do not rely on image segmentation and they produce more accurate results. Experiments show our techniques to work well for a wide range of images while achieving a speed-up of 3-5 times compared to a state-of-the-art method. Yuanjie Zheng, Stephen Lin 0001, Sing Bing Kang, Rui Xiao 0001, James C. Gee, Chandra Kambhamettu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | 3D shape regression for real-time facial animationabstractWe present a real-time performance-driven facial animation system based on 3D shape regression. In this system, the 3D positions of facial landmark points are inferred by a regressor from 2D video frames of an ordinary web camera. From these 3D points, the pose and expressions of the face are recovered by fitting a user-specific blendshape model to them. The main technical contribution of this work is the 3D regression algorithm that learns an accurate, user-specific face alignment model from an easily acquired set of training data, generated from images of the user performing a sequence of predefined facial poses and expressions. Experiments show that our system can accurately recover 3D face shapes even for fast motions, non-frontal faces, and exaggerated expressions. In addition, some capacity to handle partial occlusions and changing lighting conditions is demonstrated. Yanlin Weng, Stephen Lin 0001, Kun Zhou 0001 |
ACM Trans. Graph. | 3 |
| 2013 | Global illumination with radiance regression functionsabstractWe present radiance regression functions for fast rendering of global illumination in scenes with dynamic local light sources. A radiance regression function (RRF) represents a non-linear mapping from local and contextual attributes of surface points, such as position, viewing direction, and lighting condition, to their indirect illumination values. The RRF is obtained from precomputed shading samples through regression analysis, which determines a function that best fits the shading data. For a given scene, the shading samples are precomputed by an offline renderer. The key idea behind our approach is to exploit the nonlinear coherence of the indirect illumination data to make the RRF both compact and fast to evaluate. We model the RRF as a multilayer acyclic feed-forward neural network, which provides a close functional approximation of the indirect illumination and can be efficiently evaluated at run time. To effectively model scenes with spatially variant material properties, we utilize an augmented set of attributes as input to the neural network RRF to reduce the amount of inference that the network needs to perform. To handle scenes with greater geometric complexity, we partition the input space of the RRF model and represent the subspaces with separate, smaller RRFs that can be evaluated more rapidly. As a result, the RRF model scales well to increasingly complex scene geometry and material variation. Because of its compactness and ease of evaluation, the RRF model enables real-time rendering with full global illumination effects, including changing caustics and multiple-bounce high-frequency glossy interreflections. Peiran Ren, Jiaping Wang, Minmin Gong, Stephen Lin 0001, Xin Tong 0001, Baining Guo |
ACM Trans. Graph. | 4 |
| 2013 | TransCut: Interactive Rendering of Translucent CutoutsabstractWe present TransCut, a technique for interactive rendering of translucent objects undergoing fracturing and cutting operations. As the object is fractured or cut open, the user can directly examine and intuitively understand the complex translucent interior, as well as edit material properties through painting on cross sections and recombining the broken pieces—all with immediate and realistic visual feedback. This new mode of interaction with translucent volumes is made possible with two technical contributions. The first is a novel solver for the diffusion equation (DE) over a tetrahedral mesh that produces high-quality results comparable to the state-of-art finite element method (FEM) of Arbree et al. but at substantially higher speeds. This accuracy and efficiency is obtained by computing the discrete divergences of the diffusion equation and constructing the DE matrix using analytic formulas derived for linear finite elements. The second contribution is a multiresolution algorithm to significantly accelerate our DE solver while adapting to the frequent changes in topological structure of dynamic objects. The entire multiresolution DE solver is highly parallel and easily implemented on the GPU. We believe TransCut provides a novel visual effect for heterogeneous translucent objects undergoing fracturing and cutting operations. Dongping Li, Xin Sun 0014, Zhong Ren 0001, Stephen Lin 0001, Yiying Tong, Baining Guo, Kun Zhou 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2013 | Interactive chromaticity mapping for multispectral images
Yanxiang Lan, Jiaping Wang, Stephen Lin 0001, Minmin Gong, Xin Tong 0001, Baining Guo |
Vis. Comput. | 3 |
| 2012 | Motion-aware noise filtering for deblurring of noisy and blurry imagesabstractImage noise can present a serious problem in motion deblurring. While most state-of-the-art motion deblurring algorithms can deal with small levels of noise, in many cases such as low-light imaging, the noise is large enough in the blurred image that it cannot be handled effectively by these algorithms. In this paper, we propose a technique for jointly denoising and deblurring such images that elevates the performance of existing motion deblurring algorithms. Our method takes advantage of estimated motion blur kernels to improve denoising, by constraining the denoised image to be consistent with the estimated camera motion (i.e., no high frequency noise features that do not match the motion blur). This improved denoising then leads to higher quality blur kernel estimation and deblurring performance. The two operations are iterated in this manner to obtain results superior to suppressing noise effects through regularization in deblurring or by applying denoising as a preprocess. This is demonstrated in experiments both quantitatively and qualitatively using various image examples. Yu-Wing Tai, Stephen Lin 0001 |
CVPR | 2 |
| 2012 | Estimation of Intrinsic Image Sequences from Image+Depth Video
Kyong Joon Lee, Xin Tong 0001, Minmin Gong, Shahram Izadi, Sang Uk Lee, Ping Tan 0002, Stephen Lin 0001 |
ECCV (6) | 8 |
| 2012 | Beyond Spatial Pyramids: A New Feature Extraction Framework with Dense Spatial Sampling for Image Classification
Shengye Yan, Xinxing Xu, Dong Xu 0001, Stephen Lin 0001, Xuelong Li 0001 |
ECCV (4) | 4 |
| 2012 | Removal of dust artifacts in focal stack image sequences
Chen Li 0031, Kun Zhou 0001, Stephen Lin 0001 |
ICPR | 3 |
| 2012 | Efficient Optic Cup Detection from Intra-image Learning with Retinal Structure Priors
Yanwu Xu 0001, Jiang Liu 0001, Stephen Lin 0001, Dong Xu 0001, Carol Yim-lui Cheung, Tin Aung, Tien Yin Wong |
MICCAI (1) | 3 |
| 2012 | A New In-Camera Imaging Model for Color Computer Vision and Its ApplicationabstractWe present a study of in-camera image processing through an extensive analysis of more than 10,000 images from over 30 cameras. The goal of this work is to investigate if image values can be transformed to physically meaningful values, and if so, when and how this can be done. From our analysis, we found a major limitation of the imaging model employed in conventional radiometric calibration methods and propose a new in-camera imaging model that fits well with today's cameras. With the new model, we present associated calibration procedures that allow us to convert sRGB images back to their original CCD RAW responses in a manner that is significantly more accurate than any existing methods. Additionally, we show how this new imaging model can be used to build an image correction application that converts an sRGB input image captured with the wrong camera settings to an sRGB output image that would have been recorded under the correct settings of a specific camera. Seon Joo Kim, Hai Ting Lin, Zheng Lu 0002, Sabine Süsstrunk, Stephen Lin 0001, Michael S. Brown |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2012 | A Closed-Form Solution to Retinex with Nonlocal Texture ConstraintsabstractWe propose a method for intrinsic image decomposition based on retinex theory and texture analysis. While most previous methods approach this problem by analyzing local gradient properties, our technique additionally identifies distant pixels with the same reflectance through texture analysis, and uses these nonlocal reflectance constraints to significantly reduce ambiguity in decomposition. We formulate the decomposition problem as the minimization of a quadratic function which incorporates both the retinex constraint and our nonlocal texture constraint. This optimization can be solved in closed form with the standard conjugate gradient algorithm. Extensive experimentation with comparisons to previous techniques validate our method in terms of both decomposition accuracy and runtime efficiency. Ping Tan 0002, Li Shen 0003, Enhua Wu, Stephen Lin 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2012 | Diffusion curve textures for resolution independent texture mappingabstractWe introduce a vector representation called diffusion curve textures for mapping diffusion curve images (DCI) onto arbitrary surfaces. In contrast to the original implicit representation of DCIs [Orzan et al. 2008], where determining a single texture value requires iterative computation of the entire DCI via the Poisson equation, diffusion curve textures provide an explicit representation from which the texture value at any point can be solved directly, while preserving the compactness and resolution independence of diffusion curves. This is achieved through a formulation of the DCI diffusion process in terms of Green's functions. This formulation furthermore allows the texture value of any rectangular region (e.g. pixel area) to be solved in closed form, which facilitates anti-aliasing. We develop a GPU algorithm that renders anti-aliased diffusion curve textures in real time, and demonstrate the effectiveness of this method through high quality renderings with detailed control curves and color variations. Xin Sun 0014, Guofu Xie, Yue Dong 0001, Stephen Lin 0001, Weiwei Xu 0003, Xin Tong 0001, Baining Guo |
ACM Trans. Graph. | 4 |
| 2012 | Detection of Sudden Pedestrian Crossings for Driving Assistance SystemsabstractIn this paper, we study the problem of detecting sudden pedestrian crossings to assist drivers in avoiding accidents. This application has two major requirements: to detect crossing pedestrians as early as possible just as they enter the view of the car-mounted camera and to maintain a false alarm rate as low as possible for practical purposes. Although many current sliding-window-based approaches using various features and classification algorithms have been proposed for image-/video-based pedestrian detection, their performance in terms of accuracy and processing speed falls far short of practical application requirements. To address this problem, we propose a three-level coarse-to-fine video-based framework that detects partially visible pedestrians just as they enter the camera view, with low false alarm rate and high speed. The framework is tested on a new collection of high-resolution videos captured from a moving vehicle and yields a performance better than that of state-of-the-art pedestrian detection while running at a frame rate of 55 fps. Yanwu Xu 0001, Dong Xu 0001, Stephen Lin 0001, Tony X. Han, Xianbin Cao 0001, Xuelong Li 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2011 | High resolution multispectral video capture with a hybrid camera systemabstractWe present a new approach to capture video at high spatial and spectral resolutions using a hybrid camera system. Composed of an RGB video camera, a grayscale video camera and several optical elements, the hybrid camera system simultaneously records two video streams: an RGB video with high spatial resolution, and a multispectral video with low spatial resolution. After registration of the two video streams, our system propagates the multispectral information into the RGB video to produce a video with both high spectral and spatial resolution. This propagation between videos is guided by color similarity of pixels in the spectral domain, proximity in the spatial domain, and the consistent color of each scene point in the temporal domain. The propagation algorithm is designed for rapid computation to allow real-time video generation at the original frame rate, and can thus facilitate real-time video analysis tasks such as tracking and surveillance. Hardware implementation details and design tradeoffs are discussed. We evaluate the proposed system using both simulations with ground truth data and on real-world scenes. The utility of this high resolution multispectral video data is demonstrated in dynamic white balance adjustment and tracking. Xun Cao, Xin Tong 0001, Qionghai Dai, Stephen Lin 0001 |
CVPR | 4 |
| 2011 | Sliding Window and Regression Based Cup Detection in Digital Fundus Images for Glaucoma Diagnosis
Yanwu Xu 0001, Dong Xu 0001, Stephen Lin 0001, Jiang Liu 0001, Jun Cheng 0003, Carol Yim-lui Cheung, Tin Aung, Tien Yin Wong |
MICCAI (3) | 3 |
| 2011 | Coded Aperture Pairs for Depth from Defocus and Defocus Deblurring
Changyin Zhou, Stephen Lin 0001, Shree K. Nayar |
Int. J. Comput. Vis. | 2 |
| 2011 | A Prism-Mask System for Multispectral Video AcquisitionabstractThis paper presents a prism-mask system for capturing multispectral videos. The system is composed of a triangular prism, a monochrome camera, and an occlusion mask. Incoming light beams from the scene are sampled by the occlusion mask, dispersed into their constituent spectra by the triangular prism, and then captured by the monochrome camera. Our system is capable of capturing frames with high spectral resolution at video rates. It also allows for different trade-offs between spectral and spatial resolution by adjusting the focal length of the camera. We demonstrate multispectral video acquisition with various spectral resolutions and spatial resolutions, as well as different frame rates. The effectiveness of our system is further evaluated with several applications, including human skin detection, physical material recognition, video segmentation, RGB video generation, and illumination identification. Xun Cao, Hao Du 0004, Xin Tong 0001, Qionghai Dai, Stephen Lin 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2011 | Semantic colorization with internet imagesabstractColorization of a grayscale photograph often requires considerable effort from the user, either by placing numerous color scribbles over the image to initialize a color propagation algorithm, or by looking for a suitable reference image from which color information can be transferred. Even with this user supplied data, colorized images may appear unnatural as a result of limited user skill or inaccurate transfer of colors. To address these problems, we propose a colorization system that leverages the rich image content on the internet. As input, the user needs only to provide a semantic text label and segmentation cues for major foreground objects in the scene. With this information, images are downloaded from photo sharing websites and filtered to obtain suitable reference images that are reliable for color transfer to the given grayscale photo. Different image colorizations are generated from the various reference images, and a graphical user interface is provided to easily select the desired result. Our experiments and user study demonstrate the greater effectiveness of this system in comparison to previous techniques. Alex Yong Sang Chia, Shaojie Zhuo, Raj Kumar Gupta, Yu-Wing Tai, Siu-Yeung Cho, Ping Tan 0002, Stephen Lin 0001 |
ACM Trans. Graph. | 7 |
| 2010 | Coded exposure imaging for projective motion deblurringabstractWe propose a method for deblurring of spatially variant object motion. A principal challenge of this problem is how to estimate the point spread function (PSF) of the spatially variant blur. Based on the projective motion blur model of, we present a blur estimation technique that jointly utilizes a coded exposure camera and simple user interactions to recover the PSF. With this spatially variant PSF, objects that exhibit projective motion can be effectively de-blurred. We validate this method with several challenging image examples. Yu-Wing Tai, Naejin Kong, Stephen Lin 0001, Joseph S. Shin |
CVPR | 3 |
| 2010 | Super resolution using edge prior and single image detail synthesisabstractEdge-directed image super resolution (SR) focuses on ways to remove edge artifacts in upsampled images. Under large magnification, however, textured regions become blurred and appear homogenous, resulting in a super-resolution image that looks unnatural. Alternatively, learning-based SR approaches use a large database of exemplar images for “hallucinating” detail. The quality of the upsampled image, especially about edges, is dependent on the suitability of the training images. This paper aims to combine the benefits of edge-directed SR with those of learning-based SR. In particular, we propose an approach to extend edge-directed super-resolution to include detail from an image/texture example provided by the user (e.g., from the Internet). A significant benefit of our approach is that only a single exemplar image is required to supply the missing detail - strong edges are obtained in the SR image even if they are not present in the example image due to the combination of the edge-directed approach. In addition, we can achieve quality results at very large magnification, which is often problematic for both edge-directed and learning-based approaches. Yu-Wing Tai, Shuaicheng Liu, Michael S. Brown, Stephen Lin 0001 |
CVPR | 4 |
| 2010 | Correction of Spatially Varying Image and Video Motion Blur Using a Hybrid CameraabstractWe describe a novel approach to reduce spatially varying motion blur in video and images using a hybrid camera system. A hybrid camera is a standard video camera that is coupled with an auxiliary low-resolution camera sharing the same optical path but capturing at a significantly higher frame rate. The auxiliary video is temporally sharper but at a lower resolution, while the lower frame-rate video has higher spatial resolution but is susceptible to motion blur. Our deblurring approach uses the data from these two video streams to reduce spatially varying motion blur in the high-resolution camera with a technique that combines both deconvolution and super-resolution. Our algorithm also incorporates a refinement of the spatially varying blur kernels to further improve results. Our approach can reduce motion blur from the high-resolution video as well as estimate new high-resolution frames at a higher frame rate. Experimental results on a variety of inputs demonstrate notable improvement over current state-of-the-art methods in image/video deblurring. Yu-Wing Tai, Hao Du 0004, Michael S. Brown, Stephen Lin 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2010 | Line space gathering for single scattering in large scenesabstractWe present an efficient technique to render single scattering in large scenes with reflective and refractive objects and homogeneous participating media. Efficiency is obtained by evaluating the final radiance along a viewing ray directly from the lighting rays passing near to it, and by rapidly identifying such lighting rays in the scene. To facilitate the search for nearby lighting rays, we convert lighting rays and viewing rays into 6D points and planes according to their Plücker coordinates and coefficients, respectively. In this 6D line space, the problem of closest lines search becomes one of closest points to a plane query, which we significantly accelerate using a spatial hierarchy of the 6D points. This approach to lighting ray gathering supports complex light paths with multiple reflections and refractions, and avoids the use of a volume representation, which is expensive for large-scale scenes. This method also utilizes far fewer lighting rays than the number of photons needed in traditional volumetric photon mapping, and does not discretize viewing rays into numerous steps for ray marching. With this approach, results similar to volumetric photon mapping are obtained efficiently in terms of both storage and computation. Xin Sun 0014, Kun Zhou 0001, Stephen Lin 0001, Baining Guo |
ACM Trans. Graph. | 3 |
| 2009 | Single-image optical center estimation from vignetting and tangential gradient symmetryabstractIn this paper, we propose a method for estimating the optical center of a camera given only a single image with vignetting. This is accomplished by identifying the center of the vignetting effect in the image through an analysis of semicircular tangential gradients (SCTGs). For a given image pixel, the SCTG is the image gradient along the tangential direction of the circle centered at the currently estimated optical center and passing through the pixel. We show that for natural images with vignetting, the distribution of SCTGs is generally symmetric if the optical center is estimated accurately, but is skewed otherwise. By minimizing the asymmetry of the SCTG distribution with nonlinear optimization, our method is able to obtain reliable estimates of the optical center. Experiments on simulated and real vignetting images demonstrate the effectiveness of this technique. Yuanjie Zheng, Chandra Kambhamettu, Stephen Lin 0001 |
CVPR | 3 |
| 2009 | A prism-based system for multispectral video acquisitionabstractIn this paper, we propose a prism-based system for capturing multispectral videos. The system consists of a triangular prism, a monochrome camera, and an occlusion mask. Incoming light beams from the scene are sampled by the occlusion mask, dispersed into their constituent spectra by the triangular prism, and then captured by the monochrome camera. Our system is capable of capturing videos of high spectral resolution. It also allows for different tradeoffs between spectral and spatial resolution by adjusting the focal length of the camera. We demonstrate the effectiveness of our system with several applications, including human skin detection, physical material recognition, and RGB video generation. Hao Du 0004, Xin Tong 0001, Xun Cao, Stephen Lin 0001 |
ICCV | 4 |
| 2009 | Coded aperture pairs for depth from defocusabstractThe classical approach to depth from defocus uses two images taken with circular apertures of different sizes. We show in this paper that the use of a circular aperture severely restricts the accuracy of depth from defocus. We derive a criterion for evaluating a pair of apertures with respect to the precision of depth recovery. This criterion is optimized using a genetic algorithm and gradient descent search to arrive at a pair of high resolution apertures. The two coded apertures are found to complement each other in the scene frequencies they preserve. This property enables them to not only recover depth with greater fidelity but also obtain a high quality all-focused image from the two captured images. Extensive simulations as well as experiments on a variety of scenes demonstrate the benefits of using the coded apertures over conventional circular apertures. Changyin Zhou, Stephen Lin 0001, Shree K. Nayar |
ICCV | 2 |
| 2009 | Enhancing Bilinear Subspace Learning by Element RearrangementabstractThe success of bilinear subspace learning heavily depends on reducing correlations among features along rows and columns of the data matrices. In this work, we study the problem of rearranging elements within a matrix in order to maximize these correlations so that information redundancy in matrix data can be more extensively removed by existing bilinear subspace learning algorithms. An efficient iterative algorithm is proposed to tackle this essentially integer programming problem. In each step, the matrix structure is refined with a constrained Earth Mover's Distance procedure that incrementally rearranges matrices to become more similar to their low-rank approximations, which have high correlation among features along rows and columns. In addition, we present two extensions of the algorithm for conducting supervised bilinear subspace learning. Experiments in both unsupervised and supervised bilinear subspace learning demonstrate the effectiveness of our proposed algorithms in improving data compression performance and classification accuracy. Dong Xu 0001, Shuicheng Yan, Stephen Lin 0001, Thomas S. Huang, Shih-Fu Chang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2009 | Single-Image Vignetting CorrectionabstractIn this paper, we propose a method for robustly determining the vignetting function given only a single image. Our method is designed to handle both textured and untextured regions in order to maximize the use of available information. To extract vignetting information from an image, we present adaptations of segmentation techniques that locate image regions with reliable data for vignetting estimation. Within each image region, our method capitalizes on the frequency characteristics and physical properties of vignetting to distinguish it from other sources of intensity variation. Rejection of outlier pixels is applied to improve the robustness of vignetting estimation. Comprehensive experiments demonstrate the effectiveness of this technique on a broad range of images with both simulated and natural vignetting effects. Causes of failures using the proposed algorithm are also analyzed. Yuanjie Zheng, Stephen Lin 0001, Chandra Kambhamettu, Jingyi Yu 0001, Sing Bing Kang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Intrinsic image decomposition with non-local texture cuesabstractWe present a method for decomposing an image into its intrinsic reflectance and shading components. Different from previous work, our method examines texture information to obtain constraints on reflectance among pixels that may be distant from one another in the image. We observe that distinct points with the same intensity-normalized texture configuration generally have the same reflectance value. The separation of shading and reflectance components should thus be performed in a manner that guarantees these non-local constraints. We formulate intrinsic image decomposition by adding these non-local texture constraints to the local derivative analysis employed in conventional techniques. Our results show a significant improvement in performance, with better recovery of global reflectance and shading structure than by previous methods. Li Shen 0003, Ping Tan 0002, Stephen Lin 0001 |
CVPR | 3 |
| 2008 | Image/video deblurring using a hybrid cameraabstractWe propose a novel approach to reduce spatially varying motion blur using a hybrid camera system that simultaneously captures high-resolution video at a low-frame rate together with low-resolution video at a high-frame rate. Our work is inspired by Ben-Ezra and Nayar who introduced the hybrid camera idea for correcting global motion blur for a single still image. We broaden the scope of the problem to address spatially varying blur as well as video imagery. We also reformulate the correction process to use more information available in the hybrid camera system, as well as iteratively refine spatially varying motion extracted from the low-resolution high-speed camera. We demonstrate that our approach achieves superior results over existing work and can be extended to deblurring of moving objects. Yu-Wing Tai, Hao Du 0004, Michael S. Brown, Stephen Lin 0001 |
CVPR | 4 |
| 2008 | Single-image vignetting correction using radial gradient symmetryabstractIn this paper, we present a novel single-image vignetting method based on the symmetric distribution of the radial gradient (RG). The radial gradient is the image gradient along the radial direction with respect to the image center. We show that the RG distribution for natural images without vignetting is generally symmetric. However, this distribution is skewed by vignetting. We develop two variants of this technique, both of which remove vignetting by minimizing asymmetry of the RG distribution. Compared with prior approaches to single-image vignetting correction, our method does not require segmentation and the results are generally better. Experiments show our technique works for a wide range of images and it achieves a speed-up of 4’5 times compared with a state-of-the-art method. Yuanjie Zheng, Jingyi Yu 0001, Sing Bing Kang, Stephen Lin 0001, Chandra Kambhamettu |
CVPR | 4 |
| 2008 | Gradient-based Interpolation and Sampling for Real-time Rendering of Inhomogeneous, Single-scattering MediaabstractAbstract We present a real‐time rendering algorithm for inhomogeneous, single scattering media, where all‐frequency shading effects such as glows, light shafts, and volumetric shadows can all be captured. The algorithm first computes source radiance at a small number of sample points in the medium, then interpolates these values at other points in the volume using a gradient‐based scheme that is efficiently applied by sample splatting. The sample points are dynamically determined based on a recursive sample splitting procedure that adapts the number and locations of sample points for accurate and efficient reproduction of shading variations in the medium. The entire pipeline can be easily implemented on the GPU to achieve real‐time performance for dynamic lighting and scenes. Rendering results of our method are shown to be comparable to those from ray tracing. Zhong Ren 0001, Kun Zhou 0001, Stephen Lin 0001, Baining Guo |
Comput. Graph. Forum | 3 |
| 2008 | Separating corneal reflections for illumination estimation
Huiqiong Wang, Stephen Lin 0001, Xiuqing Ye, Weikang Gu |
Neurocomputing | 2 |
| 2008 | Subpixel Photometric StereoabstractConventional photometric stereo recovers one normal direction per pixel of the input image. This fundamentally limits the scale of recovered geometry to the resolution of the input image, and cannot model surfaces with subpixel geometric structures. In this paper, we propose a method to recover subpixel surface geometry by studying the relationship between the subpixel geometry and the reflectance properties of a surface. We first describe a generalized physically-based reflectance model that relates the distribution of surface normals inside each pixel area to its reflectance function. The distribution of surface normals can be computed from the reflectance functions recorded in photometric stereo images. A convexity measure of subpixel geometry structure is also recovered at each pixel, through an analysis of the shadowing attenuation. Then, we use the recovered distribution of surface normals and the surface convexity to infer subpixel geometric structures on a surface of homogeneous material by spatially arranging the normals among pixels at a higher resolution than that of the input image. Finally, we optimize the arrangement of normals using a combination of belief propagation and MCMC based on a minimum description length criterion on 3D textons over the surface. The experiments demonstrate the validity of our approach and show superior geometric resolution for the recovered surfaces. Ping Tan 0002, Stephen Lin 0001, Long Quan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Convergent 2-D Subspace Learning With Null Space AnalysisabstractRecent research has demonstrated the success of supervised dimensionality reduction algorithms 2DLDA and 2DMFA, which are based on the image-as-matrix representation, in small sample size cases. To solve the convergence problem in 2DLDA and 2DMFA, we propose in this work two new schemes, called Null Space based 2DLDA (NS2DLDA) and Null Space based 2DMFA (NS2DMFA), and apply them to the challenging multi-view face recognition task. First, we convert each 2-D face image (matrix) into a vector and compute the first projection matrixP1from the null space of the intra-class scatter matrix, such that the samples from the same class are projected to the same point. Then the data are projected and reconstructed withP1. Finally, we re-organize the reconstructed datum into a matrix and then compute the second projection directionP2, in the form of a Kronecker product of two matrices, by maximizing the inter-class scatter. A proof of algorithmic convergence is provided. The experiments on two benchmark multi-view face databases, the CMU PIE and FERET databases, demonstrate that NS2DLDA outperforms Fisherface, Null Space LDA (NSLDA) and 2DLDA. Additionally, NS2DMFA is also demonstrated to be more accurate than MFA and 2DMFA for face recognition. Dong Xu 0001, Shuicheng Yan, Stephen Lin 0001, Thomas S. Huang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2008 | Reconstruction and Recognition of Tensor-Based Objects With Concurrent Subspaces AnalysisabstractPrincipal Components Analysis (PCA) has traditionally been utilized with data expressed in the form of 1-D vectors, but there exists much data such as gray-level images, video sequences, Gabor-filtered images and so on, that are intrinsically in the form of second or higher order tensors. For representations of image objects in their intrinsic form and order rather than concatenating all the object data into a single vector, we propose in this paper a new optimal object reconstruction criterion with which the information of a high-dimensional tensor is represented as a much lower dimensional tensor computed from projections to multiple concurrent subspaces. In each of these subspaces, correlations with respect to one of the tensor dimensions are reduced, enabling better object reconstruction performance. Concurrent subspaces analysis (CSA) is presented to efficiently learn these subspaces in an iterative manner. In contrast to techniques such as PCA which vectorize tensor data, CSA's direct use of data in tensor form brings an enhanced ability to learn a representative subspace and an increased number of available projection directions. These properties enable CSA to outperform traditional algorithms in the common case of small sample sizes, where CSA can be effective even with only a single sample per class. Extensive experiments on images of faces and digital numbers encoded as second or third order tensors demonstrate that the proposed CSA outperforms PCA-based algorithms in object reconstruction and object recognition. Dong Xu 0001, Shuicheng Yan, Lei Zhang 0001, Stephen Lin 0001, HongJiang Zhang, Thomas S. Huang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2008 | Intrinsic colorizationabstractIn this paper, we present an example-based colorization technique robust to illumination differences between grayscale target and color reference images. To achieve this goal, our method performs color transfer in an illumination-independent domain that is relatively free of shadows and highlights. It first recovers an illumination-independent intrinsic reflectance image of the target scene from multiple color references obtained by web search. The reference images from the web search may be taken from different vantage points, under different illumination conditions, and with different cameras. Grayscale versions of these reference images are then used in decomposing the grayscale target image into its intrinsic reflectance and illumination components. We transfer color from the color reflectance image to the grayscale reflectance image, and obtain the final result by relighting with the illumination component of the target image. We demonstrate via several examples that our method generates results with excellent color consistency. Xiaopei Liu, Yingge Qu, Tien-Tsin Wong, Stephen Lin 0001, Andrew Chi-Sing Leung, Pheng-Ann Heng |
ACM Trans. Graph. | 5 |
| 2008 | Modeling and rendering of heterogeneous translucent materials using the diffusion equationabstractIn this article, we propose techniques for modeling and rendering of heterogeneous translucent materials that enable acquisition from measured samples, interactive editing of material attributes, and real-time rendering. The materials are assumed to be optically dense such that multiple scattering can be approximated by a diffusion process described by the diffusion equation. For modeling heterogeneous materials, we present the inverse diffusion algorithm for acquiring material properties from appearance measurements. This modeling algorithm incorporates a regularizer to handle the ill-conditioning of the inverse problem, an adjoint method to dramatically reduce the computational cost, and a hierarchical GPU implementation for further speedup. To render an object with known material properties, we present the polygrid diffusion algorithm , which solves the diffusion equation with a boundary condition defined by the given illumination environment. This rendering technique is based on representation of an object by a polygrid, a grid with regular connectivity and an irregular shape, which facilitates solution of the diffusion equation in arbitrary volumes. Because of the regular connectivity, our rendering algorithm can be implemented on the GPU for real-time performance. We demonstrate our techniques by capturing materials from physical samples and performing real-time rendering and editing with these materials. Jiaping Wang, Xin Tong 0001, Stephen Lin 0001, Zhouchen Lin, Yue Dong 0001, Baining Guo, Harry Shum |
ACM Trans. Graph. | 4 |
| 2008 | Real-time smoke rendering using compensated ray marchingabstractWe present a real-time algorithm calledcompensated ray marchingfor rendering of smoke under dynamic low-frequency environment lighting. Our approach is based on a decomposition of the input smoke animation, represented as a sequence of volumetric density fields, into a set of radial basis functions (RBFs) and a sequence of residual fields. To expedite rendering, the source radiance distribution within the smoke is computed from only the low-frequency RBF approximation of the density fields, since the high-frequency residuals have little impact on global illumination under low-frequency environment lighting. Furthermore, in computing source radiances the contributions from single and multiple scattering are evaluated at only the RBF centers and then approximated at other points in the volume using an RBF-based interpolation. A slice-based integration of these source radiances along each view ray is then performed to render the final image. The high-frequency residual fields, which are a critical component in the local appearance of smoke, are compensated back into the radiance integral during this ray march to generate images of high detail. The runtime algorithm, which includes both light transfer simulation and ray marching, can be easily implemented on the GPU, and thus allows for real-time manipulation of viewpoint and lighting, as well as interactive editing of smoke attributes such as extinction cross section, scattering albedo, and phase function. Only moderate preprocessing time and storage is needed. This approach provides the first method for real-time smoke rendering that includes single and multiple scattering while generating results comparable in quality to offline algorithms like ray tracing. Kun Zhou 0001, Zhong Ren 0001, Stephen Lin 0001, Hujun Bao, Baining Guo, Harry Shum |
ACM Trans. Graph. | 3 |
| 2008 | Discriminant Locally Linear Embedding With High-Order Tensor DataabstractGraph-embedding along with its linearization and kernelization provides a general framework that unifies most traditional dimensionality reduction algorithms. From this framework, we propose a new manifold learning technique called discriminant locally linear embedding (DLLE), in which the local geometric properties within each class are preserved according to the locally linear embedding (LLE) criterion, and the separability between different classes is enforced by maximizing margins between point pairs on different classes. To deal with the out-of-sample problem in visual recognition with vector input, the linear version of DLLE, i.e., linearization of DLLE (DLLE/L), is directly proposed through the graph-embedding framework. Moreover, we propose its multilinear version, i.e., tensorization of DLLE, for the out-of-sample problem with high-order tensor input. Based on DLLE, a procedure for gait recognition is described. We conduct comprehensive experiments on both gait and face recognition, and observe that: 1) DLLE along its linearization and tensorization outperforms the related versions of linear discriminant analysis, and DLLE/L demonstrates greater effectiveness than the linearization of LLE; 2) algorithms based on tensor representations are generally superior to linear algorithms when dealing with intrinsically high-order data; and 3) for human gait recognition, DLLE/L generally obtains higher accuracy than state-of-the-art gait recognition algorithms on the standard University of South Florida gait database. Xuelong Li 0001, Stephen Lin 0001, Shuicheng Yan, Dong Xu 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2008 | Filtering and Rendering of Resolution-Dependent Reflectance ModelsabstractThe apparent reflectance of a surface depends upon the resolution at which it is imaged. Conventional reflectance models represent reflection at a single predetermined resolution; however, a low-resolution pixel that views a greater surface area often exhibits a reflectance more complicated than a high-resolution pixel with a smaller area. To address resolution dependency in reflectance, we utilize a generalized reflectance model based on a mixture of multiple conventional models, and present a framework for efficiently determining the reflectance mixture model of each pixel with respect to resolution. Mixture model parameters are precomputed at multiple resolutions and stored in mipmaps. Unlike color textures, these reflectance parameters cannot be accurately filtered by trilinear interpolation, so we present a technique for nonlinear mipmap filtering that minimizes aliasing in rendered results. This framework can be applied with various parametric reflectance models in graphics hardware for real-time processing. With this technique for filtering and rendering with mipmaps of reflectance mixture models, our system can rapidly render the resolution-dependent reflectance effects that are customarily disregarded in conventional rendering methods. At the end of this paper, we also describe how shadowing and masking effects can be incorporated into this framework to increase the realism of rendering. Ping Tan 0002, Stephen Lin 0001, Long Quan, Baining Guo, Harry Shum |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2007 | A Probabilistic Intensity Similarity Measure based on Noise DistributionsabstractWe derive a probabilistic similarity measure between two observed image intensities that is based on the noise properties of the camera. In many vision algorithms, the effect of camera noise is either neglected or reduced in a preprocessing stage. However, noise reduction cannot be performed with high accuracy due to lack of knowledge about the true intensity signal. Our similarity metric specifically represents the likelihood that two intensity observations correspond to the same unknown noise-free scene radiance. By directly accounting for noise in the evaluation of similarity, the proposed measure makes noise reduction unnecessary and enhances many vision algorithms that involve matching of image intensities. Real-world experiments demonstrate the effectiveness of the proposed similarity measure in comparison to the standard L2norm. Yasuyuki Matsushita, Stephen Lin 0001 |
CVPR | 2 |
| 2007 | Radiometric Calibration from Noise DistributionsabstractA method is proposed for estimating radiometric response functions from noise observations. From the statistical properties of noise sources, the noise distribution for each scene radiance value is shown to be symmetric for a radiometrically calibrated camera. However, due to the non-linearity of camera response functions, the observed noise distributions become skewed in an uncalibrated camera. In this paper, we capitalize on these asymmetric profiles of measured noise distributions to estimate radiometric response functions. Unlike prior approaches, the proposed method is not sensitive to noise level, and is therefore particularly useful when the noise level is high. Also, the proposed method does not require registered input images taken with different exposures; only statistical noise distributions at multiple intensity levels are used. Real-world experiments demonstrate the effectiveness of the proposed approach in comparison to standard calibration techniques. Yasuyuki Matsushita, Stephen Lin 0001 |
CVPR | 2 |
| 2007 | Element Rearrangement for Tensor-Based Subspace LearningabstractThe success of tensor-based subspace learning depends heavily on reducing correlations along the column vectors of the mode-k flattened matrix. In this work, we study the problem of rearranging elements within a tensor in order to maximize these correlations, so that information redundancy in tensor data can be more extensively removed by existing tensor-based dimensionality reduction algorithms. An efficient iterative algorithm is proposed to tackle this essentially integer optimization problem. In each step, the tensor structure is refined with a spatially-constrained Earth Mover's Distance procedure that incrementally rearranges tensors to become more similar to their low rank approximations, which have high correlation among features along certain tensor dimensions. Monotonic convergence of the algorithm is proven using an auxiliary function analogous to that used for proving convergence of the Expectation-Maximization algorithm. In addition, we present an extension of the algorithm for conducting supervised subspace learning with tensor data. Experiments in both unsupervised and supervised subspace learning demonstrate the effectiveness of our proposed algorithms in improving data compression performance and classification accuracy. Shuicheng Yan, Dong Xu 0001, Stephen Lin 0001, Thomas S. Huang, Shih-Fu Chang |
CVPR | 3 |
| 2007 | Removal of Image Artifacts Due to Sensor DustabstractImage artifacts that result from sensor dust are a common but annoying problem for many photographers. To reduce the appearance of dust in an image, we first formulate a model of artifact formation due to sensor dust. With this artifact formation model, we make use of contextual information in the image and a color consistency constraint on dust to remove these artifacts. When multiple images are available from the same camera, even under different camera settings, this approach can also be used to reliably detect dust regions on the sensor. In contrast to image inpainting or other hole-filling methods, the proposed technique utilizes image information within a dust region to guide the use of contextual data. Joint use of these multiple cues leads to image recovery results that are not only visually pleasing, but also faithful to the actual scene. The effectiveness of this method is demonstrated in experiments with various cameras. Changyin Zhou, Stephen Lin 0001 |
CVPR | 2 |
| 2007 | Graph Embedding and Extensions: A General Framework for Dimensionality ReductionabstractA large family of algorithms - supervised or unsupervised; stemming from statistics or geometry theory - has been designed to provide different solutions to the problem of dimensionality reduction. Despite the different motivations of these algorithms, we present in this paper a general formulation known as graph embedding to unify them within a common framework. In graph embedding, each algorithm can be considered as the direct graph embedding or its linear/kernel/tensor extension of a specific intrinsic graph that describes certain desired statistical or geometric properties of a data set, with constraints from scale normalization or a penalty graph that characterizes a statistical or geometric property that should be avoided. Furthermore, the graph embedding framework can be used as a general platform for developing new dimensionality reduction algorithms. By utilizing this framework as a tool, we propose a new supervised dimensionality reduction algorithm called marginal Fisher analysis in which the intrinsic graph characterizes the intraclass compactness and connects each data point with its neighboring points of the same class, while the penalty graph connects the marginal points and characterizes the interclass separability. We show that MFA effectively overcomes the limitations of the traditional linear discriminant analysis algorithm due to data distribution assumptions and available projection directions. Real face recognition experiments show the superiority of our proposed MFA in comparison to LDA, also for corresponding kernel and tensor extensions Shuicheng Yan, Dong Xu 0001, Benyu Zhang, HongJiang Zhang, Qiang Yang 0001, Stephen Lin 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2007 | Marginal Fisher Analysis and Its Variants for Human Gait Recognition and Content- Based Image RetrievalabstractDimensionality reduction algorithms, which aim to select a small set of efficient and discriminant features, have attracted great attention for human gait recognition and content-based image retrieval (CBIR). In this paper, we present extensions of our recently proposed marginal Fisher analysis (MFA) to address these problems. For human gait recognition, we first present a direct application of MFA, then inspired by recent advances in matrix and tensor-based dimensionality reduction algorithms, we present matrix-based MFA for directly handling 2-D input in the form of gray-level averaged images. For CBIR, we deal with the relevance feedback problem by extending MFA to marginal biased analysis, in which within-class compactness is characterized only by the distances between each positive sample and its neighboring positive samples. In addition, we present a new technique to acquire a direct optimal solution for MFA without resorting to objective function modification as done in many previous algorithms. We conduct comprehensive experiments on the USF HumanID gait database and the Corel image retrieval database. Experimental results demonstrate that MFA and its extensions outperform related algorithms in both applications. Dong Xu 0001, Shuicheng Yan, Dacheng Tao, Stephen Lin 0001, HongJiang Zhang |
IEEE Trans. Image Process. | 4 |
| 2007 | Interactive relighting with dynamic BRDFsabstractWe present a technique for interactive relighting in which source radiance, viewing direction, and BRDFs can all be changed on the fly. In handling dynamic BRDFs, our method efficiently accounts for the effects of BRDF modification on the reflectance and incident radiance at a surface point. For reflectance, we develop a BRDF tensor representation that can be factorized into adjustable terms for lighting, viewing, and BRDF parameters. For incident radiance, there exists a non-linear relationship between indirect lighting and BRDFs in a scene, which makes linear light transport frameworks such as PRT unsuitable. To overcome this problem, we introduceprecomputed transfer tensors(PTTs) which decompose indirect lighting into precomputable components that are each a function of BRDFs in the scene, and can be rapidly combined at run time to correctly determine incident radiance. We additionally describe a method for efficient handling of high-frequency specular reflections by separating them from the BRDF tensor representation and processing them using precomputed visibility information. With relighting based on PTTs, interactive performance with indirect lighting is demonstrated in applications to BRDF animation and material tuning. Xin Sun 0014, Kun Zhou 0001, Yanyun Chen, Stephen Lin 0001, Jiaoying Shi, Baining Guo |
ACM Trans. Graph. | 4 |
| 2007 | Rank-One Projections With Adaptive Margins for Face RecognitionabstractIn supervised dimensionality reduction, tensor representations of images have recently been employed to enhance classification of high dimensional data with small training sets. Previous approaches for handling tensor data have been formulated with tight restrictions on projection directions that, along with convergence issues and the assumption of Gaussian-distributed class data, limit its face-recognition performance. To overcome these problems, we propose a method of rank-one projections with adaptive margins (RPAM) that gives a provably convergent solution for tensor data over a more general class of projections, while accounting for margins between samples of different classes. In contrast to previous margin-based works which determine margin sample pairs within the original high dimensional feature space, RPAM aims instead to maximize the margins defined in the expected lower dimensional feature sub-space by progressive margin refinement after each rank-one projection. In addition to handling tensor data, vector-based variants of RPAM are presented for linear mappings and for nonlinear mappings using kernel tricks. Comprehensive experimental results demonstrate that RPAM brings significant improvement in face recognition over previous subspace learning techniques. Dong Xu 0001, Stephen Lin 0001, Shuicheng Yan, Xiaoou Tang |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2006 | Separation of Highlight Reflections on Textured SurfacesabstractWe present a method for separating highlight reflections on textured surfaces. In contrast to previous techniques that use diffuse color information from outside the highlight area to constrain the solution, the proposed method further capitalizes on the spatial distributions of colors to resolve ambiguities in separation that often arise in real images. For highlight pixels in which a clear-cut separation cannot be determined from color space analysis, we evaluate possible separation solutions based on their consistency with diffuse texture characteristics outside the highlight. With consideration of color distributions in both the color space and the image space, appreciably enhanced separation performance can be attained in challenging cases. Ping Tan 0002, Long Quan, Stephen Lin 0001 |
CVPR (2) | 3 |
| 2006 | Rank-one Projections with Adaptive Margins for Face RecognitionabstractIn supervised dimensionality reduction, tensor representations of images have recently been employed to enhance classification of high-dimensional data with small training sets. To handle tensor data, this approach has been formulated with tight restrictions on projection directions that, along with convergence issues and the assumption of Gaussian distributed class data, limits its face recognition performance. To overcome these problems, we propose a method of rank-one projections with adaptive margins (RPAM) that gives a provably convergent solution for tensor data over a more general class of projections, while accounting for margins between samples of different classes. In contrast to previous margin based works which determine margin sample pairs within the original high dimensional space, RPAM instead aims to maximize the margins defined in the expected lower dimensional feature subspace by progressive margin refinement after each rank-one projection. In addition to handling tensor data, vector-based variants of RPAM are presented for linear mappings and for nonlinear mappings using kernel tricks. Comprehensive experimental results demonstrate that RPAM brings significant improvement in face recognition over previous subspace learning techniques. Dong Xu 0001, Stephen Lin 0001, Shuicheng Yan, Xiaoou Tang |
CVPR (1) | 2 |
| 2006 | Single-Image Vignetting CorrectionabstractIn this paper, we propose a method for determining the vignetting function given only a single image. Our method is designed to handle both textured and untextured regions in order to maximize the use of available information. To extract vignetting information from an image, we present adaptations of segmentation techniques that locate image regions with reliable data for vignetting estimation. Within each image region, our method capitalizes on frequency characteristics and physical properties of vignetting to distinguish it from other sources of intensity variation. The vignetting data acquired from regions are weighted according to a presented reliability measure to promote robustness in estimation. Comprehensive experiments demonstrate the effectiveness of this technique on a broad range of images. Yuanjie Zheng, Stephen Lin 0001, Sing Bing Kang |
CVPR (1) | 2 |
| 2006 | Resolution-Enhanced Photometric Stereo
Ping Tan 0002, Stephen Lin 0001, Long Quan |
ECCV (3) | 2 |
| 2006 | Appearance manifolds for modeling time-variant appearance of materialsabstractWe present a visual simulation technique called appearance manifolds for modeling the time-variant surface appearance of a material from data captured at a single instant in time. In modeling time-variant appearance, our method takes advantage of the key observation that concurrent variations in appearance over a surface represent different degrees of weathering. By reorganizing these various appearances in a manner that reveals their relative order with respect to weathering degree, our method infers spatial and temporal appearance properties of the material's weathering process that can be used to convincingly generate its weathered appearance at different points in time. Results with natural non-linear reflectance variations are demonstrated in applications such as visual simulation of weathering on 3D models, increasing and decreasing the weathering of real objects, and material transfer with weathering effects. Jiaping Wang, Xin Tong 0001, Stephen Lin 0001, Minghao Pan, Chao Wang 0063, Hujun Bao, Baining Guo, Harry Shum |
ACM Trans. Graph. | 3 |
| 2006 | Spherical harmonics scaling
Jiaping Wang, Kun Xu 0003, Kun Zhou 0001, Stephen Lin 0001, Shi-Min Hu 0001, Baining Guo |
Vis. Comput. | 4 |
| 2005 | Separating Reflections in Human Iris Images for Illumination EstimationabstractA method is presented for separating corneal reflections in an image of human irises to estimate illumination from the surrounding scene. Previous techniques for reflection separation have demonstrated success in only limited cases, such as for uniform colored lighting and simple object textures, so they are not applicable to irises which exhibit intricate textures and complicated reflections of the environment. To make this problem feasible, we present a method that capitalizes on physical characteristics of human irises to obtain an illumination estimate that encompasses the prominent light contributors in the scene. Results of this algorithm are presented for eyes of different colors, including light colored eyes for which reflection separation is necessary to determine a valid illumination estimate. Huiqiong Wang, Stephen Lin 0001, Xiaopei Liu, Sing Bing Kang |
ICCV | 2 |
| 2005 | Multiresolution Reflectance FilteringabstractPhysically-based reflectance models typically represent light scattering as a function of surface geometry at the pixel level. With changes in viewing resolution, the geometry imaged within a pixel can undergo significant variations that result in changing reflectance characteristics. To address these transformations, we present a multiresolution reflectance framework based on microfacet normal distributions within a pixel over different scales. Since these distributions must be efficiently determined with respect to resolution, they are recorded at multiple resolution levels in mipmaps. The main contribution of this work is a real-time mipmap filtering technique for these distribution-based parameters that not only provides smooth reflectance transitions in scale, but also minimizes aliasing. With this multiresolution reflectance technique, our system can rapidly and accurately incorporate fine reflectance detail that is customarily disregarded in multiresolution rendering methods. Ping Tan 0002, Stephen Lin 0001, Long Quan, Baining Guo, Harry Shum |
Rendering Techniques | 2 |
| 2005 | Modeling and rendering of quasi-homogeneous materialsabstractMany translucent materials consist of evenly-distributed heterogeneous elements which produce a complex appearance under different lighting and viewing directions. For these quasi-homogeneous materials, existing techniques do not address how to acquire their material representations from physical samples in a way that allows arbitrary geometry models to be rendered with these materials. We propose a model for such materials that can be readily acquired from physical samples. This material model can be applied to geometric models of arbitrary shapes, and the resulting objects can be efficiently rendered without expensive subsurface light transport simulation. In developing a material model with these attributes, we capitalize on a key observation about the subsurface scattering characteristics of quasi-homogeneous materials at different scales. Locally, the non-uniformity of these materials leads to inhomogeneous subsurface scattering. For subsurface scattering on a global scale, we show that a lengthy photon path through an even distribution of heterogeneous elements statistically resembles scattering in a homogeneous medium. This observation allows us to represent and measure the global light transport within quasi-homogeneous materials as well as the transfer of light into and out of a material volume through surface mesostructures. We demonstrate our technique with results for several challenging materials that exhibit sophisticated appearance features such as transmission of back illumination through surface mesostructures. Xin Tong 0001, Jiaping Wang, Stephen Lin 0001, Baining Guo, Harry Shum |
ACM Trans. Graph. | 3 |
| 2005 | Precomputed shadow fields for dynamic scenesabstractWe present a soft shadow technique for dynamic scenes with moving objects under the combined illumination of moving local light sources and dynamic environment maps. The main idea of our technique is to precompute for each scene entity a shadow field that describes the shadowing effects of the entity at points around it. The shadow field for a light source, called a source radiance field (SRF), records radiance from an illuminant as cube maps at sampled points in its surrounding space. For an occluder, an object occlusion field (OOF) conversely represents in a similar manner the occlusion of radiance by an object. A fundamental difference between shadow fields and previous shadow computation concepts is that shadow fields can be precomputed independent of scene configuration. This is critical for dynamic scenes because, at any given instant, the shadow information at any receiver point can be rapidly computed as a simple combination of SRFs and OOFs according to the current scene configuration. Applications that particularly benefit from this technique include large dynamic scenes in which many instances of an entity can share a single shadow field. Our technique enables low-frequency shadowing effects in dynamic scenes in real-time and all-frequency shadows at interactive rates. Kun Zhou 0001, Stephen Lin 0001, Baining Guo, Harry Shum |
ACM Trans. Graph. | 3 |
| 2005 | Light Field Morphing Using 2D FeaturesabstractWe present a 2D feature-based technique for morphing 3D objects represented by light fields. Existing light field morphing methods require the user to specify corresponding 3D feature elements to guide morph computation. Since slight errors in 3D specification can lead to significant morphing artifacts, we propose a scheme based on 2D feature elements that is less sensitive to imprecise marking of features. First, 2D features are specified by the user in a number of key views in the source and target light fields. Then the two light fields are warped view by view as guided by the corresponding 2D features. Finally, the two warped light fields are blended together to yield the desired light field morph. Two key issues in light field morphing are feature specification and warping of light field rays. For feature specification, we introduce a user interface for delineating 2D features in key views of a light field, which are automatically interpolated to other views. For ray warping, we describe a 2D technique that accounts for visibility changes and present a comparison to the ideal morphing of light fields. Light field morphing based on 2D features makes it simple to incorporate previous image morphing techniques such as nonuniform blending, as well as to morph between an image and a light field. Lifeng Wang 0001, Stephen Lin 0001, Seungyong Lee 0001, Baining Guo, Harry Shum |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2005 | Shell radiance texture functions
Yanyun Chen, Xin Tong 0001, Stephen Lin 0001, Jiaoying Shi, Baining Guo, Harry Shum |
Vis. Comput. | 4 |
| 2004 | Radiometric Calibration from a Single Image
Stephen Lin 0001, Jinwei Gu, Shuntaro Yamazaki, Harry Shum |
CVPR (2) | 1 |
| 2004 | Estimating Intrinsic Images from Image Sequences with Biased Illumination
Yasuyuki Matsushita, Stephen Lin 0001, Sing Bing Kang, Harry Shum |
ECCV (2) | 2 |
| 2004 | Automatic extraction of semantic colors in sports videoabstractColor has been widely used in sports video analysis. Previous techniques, however require color models from prior information or user interaction, and do not address the problem of how to automatically form color models from a video in an arbitrary sports setting. In this paper, we propose an automatic technique for extracting color models of the playing surface and the team uniforms, which can be used in higher-level processes such as tracking and recognition. Unlike most previous methods, our approach is capable of handling multi-colored patterns like striped uniforms and playing fields. Multiple forms of color processing are used to analyze video frame content, which are then used iteratively to refine the color models. The results of our color modeling technique have been applied to shot classification, and experiments on videos of different sports have verified our approach. Boyi Zeng, Stephen Lin 0001, Guangyou Xu, Harry Shum |
ICASSP (3) | 3 |
| 2004 | Real-time environment map interpolationabstractEnvironment mapping, or reflection mapping, has been widely used in the game and movie industries to give objects a realistic illumination atmosphere. For moving objects, direct frame-by-frame calculation of environment maps and correspondence-based interpolation are both impractical for real-time applications due to the large computational costs. To deal with this problem, "fake" environment mapping with a fixed, pre-generated environment image has been commonly used, but clearly such an approximation is inadequate for a highly reflective object whose environment is constantly changing as it moves. In this paper, we present an approach that sparsely samples environment maps of a moving object and rapidly interpolates them for high performance. Two techniques are introduced for fast environment map interpolation without computation of scene shading. The first method utilizes scene geometry to facilitate interpolation, and the second involves geometry reconstruction from depth buffer values to reduce inefficiencies caused by complex scene geometry. These two techniques can easily be implemented in graphics hardware, and test results show that they achieve significant boosts in performance over frame-by-frame environment map computation with little loss in visual quality. Wenle Wang, Lifeng Wang 0001, Stephen Lin 0001, Jianmin Wang 0001, Baining Guo |
ICIG | 3 |
| 2004 | Face alignment using intrinsic informationabstractPrevious 2-D face alignment algorithms are generally quite sensitive to illumination variation and poor initialization. To account for these two obstacles, two forms of relatively lighting invariant descriptors - intrinsic gray-level information and intrinsic edge information - rare adopted in our algorithm to direct shape search. The former is recovered from local intensity normalization and useful at localizing face contours accurately despite its dependency on initialization. The latter is extracted from normalized local regions by Canny edge filtering and is robust at coarse alignment in spite of poor initialization. The different merits of these two forms of intrinsic information motivate us to employ them at different stages of our face alignment process. Extensive experimentations show that this proposed approach allows our system to handle not only illumination variation, but also poor initialization. Yuchi Huang, Stephen Lin 0001, Hanqing Lu, Harry Shum |
ICIP | 2 |
| 2004 | Generic slow-motion replay detection in sports video
Stephen Lin 0001, Guangyou Xu, Harry Shum |
ICIP | 3 |
| 2004 | Shell texture functionsabstractWe propose a texture function for realistic modeling and efficient rendering of materials that exhibit surface mesostructures, translucency and volumetric texture variations. The appearance of such complex materials for dynamic lighting and viewing directions is expensive to calculate and requires an impractical amount of storage to precompute. To handle this problem, our method models an object as a shell layer, formed by texture synthesis of a volumetric material sample, and a homogeneous inner core. To facilitate computation of surface radiance from the shell layer, we introduce the shell texture function (STF) which describes voxel irradiance fields based on precomputed fine-level light interactions such as shadowing by surface mesostructures and scattering of photons inside the object. Together with a diffusion approximation of homogeneous inner core radiance, the STF leads to fast and detailed raytraced renderings of complex materials. Yanyun Chen, Xin Tong 0001, Jiaping Wang, Stephen Lin 0001, Baining Guo, Harry Shum |
ACM Trans. Graph. | 4 |
| 2003 | Multiple-cue Illumination Estimation in Textured ScenesabstractIn this paper, we present a method that integrates cues from shading, shadow and specular reflections for estimating directional illumination in a textured scene. Texture poses a problem for lighting estimation, since texture edges can be mistaken for changes in illumination condition, and unknown variations in albedo make reflectance model fitting impractical. Unlike previous works which all assume known or uniform reflectance, our method can deal with the effects of textures by capitalizing on physical consistencies that exist among the lighting cues. Since scene textures do not exhibit such coherence, we use this property to minimize the influence of texture on illumination direction estimation. For the recovered light source directions, a technique for estimating their intensities in the presence of texture is also proposed. Yuanzhen Li, Stephen Lin 0001, Hanqing Lu, Harry Shum |
ICCV | 2 |
| 2003 | Highlight Removal by Illumination-Constrained InpaintingabstractWe present a single-image highlight removal method that incorporates illumination-based constraints into image inpainting. Unlike occluded image regions filled by traditional inpainting, highlight pixels contain some useful information for guiding the inpainting process. Constraints provided by observed pixel colors, highlight color analysis and illumination color uniformity are employed in our method to improve estimation of the underlying diffuse color. The inclusion of these illumination constraints allows for better recovery of shading and textures by inpainting. Experimental results are given to demonstrate the performance of our method. Ping Tan 0002, Stephen Lin 0001, Long Quan, Harry Shum |
ICCV | 2 |
| 2003 | View-dependent displacement mappingabstractSignificant visual effects arise from surface mesostructure, such as fine-scale shadowing, occlusion and silhouettes. To efficiently render its detailed appearance, we introduce a technique called view-dependent displacement mapping (VDM) that models surface displacements along the viewing direction. Unlike traditional displacement mapping, VDM allows for efficient rendering of self-shadows, occlusions and silhouettes without increasing the complexity of the underlying surface mesh. VDM is based on per-pixel processing, and with hardware acceleration it can render mesostructure with rich visual appearance in real time. Lifeng Wang 0001, Xin Tong 0001, Stephen Lin 0001, Shi-Min Hu 0001, Baining Guo, Harry Shum |
ACM Trans. Graph. | 4 |
| 2003 | Realistic Rendering and Animation of KnitwearabstractWe present a framework for knitwear modeling and rendering that accounts for characteristics that are particular to knitted fabrics. We first describe a model for animation that considers knitwear features and their effects on knitwear shape and interaction. With the computed free-form knitwear configurations, we present an efficient procedure for realistic synthesis based on the observation that a single cross section of yarn can serve as the basic primitive for modeling entire articles of knitwear. This primitive, called the lumislice, describes radiance from a yarn cross section that accounts for fine-level interactions among yarn fibers. By representing yarn as a sequence of identical but rotated cross sections, the lumislice can effectively propagate local microstructure over arbitrary stitch patterns and knitwear shapes. The lumislice accommodates varying levels of detail, allows for soft shadow generation, and capitalizes on hardware-assisted transparency blending. These modeling and rendering techniques together form a complete approach for generating realistic knitwear. Yanyun Chen, Stephen Lin 0001, Ying-Qing Xu, Baining Guo, Harry Shum |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2002 | Diffuse-Specular Separation and Depth Recovery from Image Sequences
Stephen Lin 0001, Yuanzhen Li, Sing Bing Kang, Xin Tong 0001, Harry Shum |
ECCV (3) | 1 |
| 2002 | Single-Image Reflectance Estimation for Relighting by Iterative Soft GroupingabstractReflectance values for image-based relighting are often estimated from grouped pixels with similar reflectance, but such groupings are difficult to compute with certainty for sparse image data. To address this problem, we propose an iterative method that aggregates BRDF data in a single image with known geometry and lighting by soft grouping, where pixels contribute to one another's estimate according to their degree of reflectance similarity. Estimation of specular reflectance is further improved by albedo-independent soft grouping of pixels based on shape continuity. With recovered reflectances, we demonstrate realistic relighting for synthetic and real scenes, including surfaces with spatially-varying reflectance. Yuanzhen Li, Stephen Lin 0001, Sing Bing Kang, Hanqing Lu, Harry Shum |
PG | 2 |
| 2002 | Lighting Interpolation by Shadow Morphing Using Intrinsic LumigraphsabstractDensely-sampled image representations such as the light field or lumigraph have been effective in enabling photorealistic image synthesis. Unfortunately, lighting interpolation with such representations has not been shown to be possible without the use of accurate 3D geometry and surface reflectance properties. In this paper we propose an approach to image-based lighting interpolation that is based on estimates of geometry and shading from relatively few images. We decompose captured light fields at different lighting conditions into intrinsic images (reflectance and illumination images), and estimate view-dependent scene geometries using multi-view stereo. We call the resulting representation an intrinsic lumigraph. In the same way that the lumigraph uses geometry to permit more accurate view interpolation, the intrinsic lumigraph uses both geometry and intrinsic images to allow high-quality interpolation at different views and lighting conditions. Joint use of geometry and intrinsic images is effective in the computation of shadow masks for shadow prediction at new lighting conditions. We illustrate our approach with images of real scenes. Yasuyuki Matsushita, Sing Bing Kang, Stephen Lin 0001, Harry Shum, Xin Tong 0001 |
PG | 3 |
| 2001 | Separation of Diffuse and Specular Reflection in Color ImagesabstractThe presence of specular reflections in images can lead many traditional computer vision algorithms to produce erroneous results. To address this problem, we propose a method based on the neutral interface reflection model for separating the diffuse and specular reflection components in color images. From two photometric images without calibrated lighting, the illuminant chromaticity is estimated, and the RGB intensities of the two reflection components are computed for each pixel using a linear model of surface reflectance. Unlike most previous methods, the presented technique does not assume any dependencies among pixels, such as regionally uniform surface reflectance. Stephen Lin 0001, Harry Shum |
CVPR (1) | 1 |
| 2001 | Photorealistic rendering of knitwear using the lumisliceabstractWe present a method for efficient synthesis of photorealistic free-form knitwear. Our approach is motivated by the observation that a single cross-section of yarn can serve as the basic primitive for modeling entire articles of knitwear. This primitive, called the lumislice, describes radiance from a yarn cross-section based on fine-level interactions — such as occlusion, shadowing, and multiple scattering — among yarn fibers. By representing yarn as a sequence of identical but rotated cross-sections, the lumislice can effectively propagate local microstructure over arbitrary stitch patterns and knitwear shapes. This framework accommodates varying levels of detail and capitalizes on hardware-assisted transparency blending. To further enhance realism, a technique for generating soft shadows from yarn is also introduced. Ying-Qing Xu, Yanyun Chen, Stephen Lin 0001, Enhua Wu, Baining Guo, Harry Shum |
SIGGRAPH | 3 |