EDBT 2026 Demo / reviewers in the wild / expert
Yung-Yu Chuang
dblp:39/2152
· DBLP profile ↗
91ranked-venue papers
5as first author
18since 2021 · last 2025
0000-0002-1383-0017ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 87 · 5 first-author · 17 since 2021Artificial intelligence and machine learning · 47 · 1 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorComputer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Training-Free Industrial Defect Generation with Diffusion Models
Ruyi Xu, Yen-Tzu Chiu, Tai-I Chen, Oscar Chew, Yung-Yu Chuang, Wen-Huang Cheng |
ICCV | 5 |
| 2025 | Bevanet: Bilateral Efficient Visual Attention Network for Real-Time Semantic SegmentationabstractReal-time semantic segmentation presents the dual challenge of designing efficient architectures that capture large receptive fields for semantic understanding while also refining detailed contours. Vision transformers model long-range dependencies effectively but incur high computational cost. To address these challenges, we introduce the Large Kernel Attention (LKA) mechanism. Our proposed Bilateral Efficient Visual Attention Network (BEVANet) expands the receptive field to capture contextual information and extracts visual and structural features using Sparse Decomposed Large Separable Kernel Attentions (SDLSKA). The Comprehensive Kernel Selection (CKS) mechanism dynamically adapts the receptive field to further enhance performance. Furthermore, the Deep Large Kernel Pyramid Pooling Module (DLKPPM) enriches contextual features by synergistically combining dilated convolutions and large kernel attention. The bilateral architecture facilitates frequent branch communication, and the Boundary Guided Attention Fusion (BGAF) module enhances boundary delineation by integrating spatial and semantic features under boundary guidance. BEVANet achieves real-time segmentation at 33 FPS, yielding 79.3% mIoU without pretraining and 81.0% mIoU on Cityscapes after ImageNet pretraining, demonstrating state-of-the-art performance. Ping-Mao Huang, I-Tien Chao, Ping-Chia Huang, Jia-Wei Liao, Yung-Yu Chuang |
ICIP | 5 |
| 2024 | Image-Text Co-Decomposition for Text-Supervised Semantic SegmentationabstractThis paper addresses text-supervised semantic segmentation, aiming to learn a model capable of segmenting arbitrary visual concepts within images by using only image-text pairs without dense annotations. Existing methods have demonstrated that contrastive learning on image-text pairs effectively aligns visual segments with the meanings of texts. We notice that there is a discrepancy between text alignment and semantic segmentation: A text often consists of multiple semantic concepts, whereas semantic segmentation strives to create semantically homogeneous segments. To address this issue, we propose a novel framework, Image-Text Co-Decomposition (CoDe), where the paired image and text are jointly decomposed into a set of image regions and a set of word segments, respectively, and contrastive learning is developed to enforce region-word alignment. To work with a vision-language model, we present a prompt learning mechanism that derives an extra representation to highlight an image segment or a word segment of interest, with which more effective features can be extracted from that segment. Comprehensive experimental results demonstrate that our method performs favorably against existing text-supervised semantic segmentation methods on six benchmark datasets. The code is available at https://github.com/072jiajia/image-text-co-decomposition. Ji-Jia Wu, Andy Chia-Hao Chang, Chieh-Yu Chuang, Chun-Pei Chen, Yu-Lun Liu 0001, Min-Hung Chen, Hou-Ning Hu, Yung-Yu Chuang, Yen-Yu Lin |
CVPR | 8 |
| 2024 | Semantic-Region Specific Lookup Tables for Image Enhancement Via Unpaired LearningabstractThis paper proposes a novel unpaired learning approach to enhance images by employing specific strategies for different semantic regions. Leveraging the generative adversarial network (GAN) framework for unpaired learning, our method incorporates a cascaded 1D and 3D lookup table (LUT) structure as the generator. Initially, context-aware 1D LUTs redistribute the input image to approach the target globally. Subsequently, category-specific 3D LUTs are merged based on the semantic category probability assigned to each pixel. The fused 3D LUTs are then applied to transform individual pixels, producing visually pleasing results. Furthermore, we introduce a semantic-attended multi-discriminator, offering more precise supervision during training. To train and evaluate our method, we have curated a semantically categorized dataset. User studies and qualitative comparisons demonstrate that our model outperforms existing methods, exhibiting better alignment with human aesthetics. Zheng-Hui Huang, Tse-Yan Lee, Li-Jen Chang, Yong-Wei Chen, Ping-Jui Chiang, Jo-Fan Wu, Yung-Yu Chuang |
ICIP | 7 |
| 2023 | Robust Dynamic Radiance FieldsabstractDynamic radiance field reconstruction methods aim to model the time-varying structure and appearance of a dynamic scene. Existing methods, however, assume that accurate camera poses can be reliably estimated by Structure from Motion (SfM) algorithms. These methods, thus, are unreliable as SfM algorithms often fail or produce erroneous poses on challenging videos with highly dynamic objects, poorly textured surfaces, and rotating camera motion. We address this robustness issue by jointly estimating the static and dynamic radiance fields along with the camera parameters (poses and focal length). We demonstrate the robustness of our approach via extensive quantitative and qualitative experiments. Our results show favorable performance over the state-of-the-art dynamic view synthesis methods. Yu-Lun Liu 0001, Chen Gao 0003, Andreas Meuleman, Hung-Yu Tseng, Ayush Saraf, Changil Kim 0001, Yung-Yu Chuang, Johannes Kopf 0001, Jia-Bin Huang 0001 |
CVPR | 7 |
| 2023 | 2D-3D Interlaced Transformer for Point Cloud Segmentation with Scene-Level SupervisionabstractWe present a Multimodal Interlaced Transformer (MIT) that jointly considers 2D and 3D data for weakly supervised point cloud segmentation. Research studies have shown that 2D and 3D features are complementary for point cloud segmentation. However, existing methods require extra 2D annotations to achieve 2D-3D information fusion. Considering the high annotation cost of point clouds, effective 2D and 3D feature fusion based on weakly supervised learning is in great demand. To this end, we propose a transformer model with two encoders and one decoder for weakly supervised point cloud segmentation using only scene-level class tags. Specifically, the two encoders compute the self-attended features for 3D point clouds and 2D multi-view images, respectively. The decoder implements interlaced 2D-3D cross-attention and carries out implicit 2D and 3D feature fusion. We alternately switch the roles of queries and key-value pairs in the decoder layers. It turns out that the 2D and 3D features are iteratively enriched by each other. Experiments show that it performs favorably against existing weakly supervised point cloud segmentation methods by a large margin on the S3DIS and ScanNet benchmarks. The project page will be available at https://jimmy15923.github.io/mit_web/. Cheng-Kun Yang, Min-Hung Chen, Yung-Yu Chuang, Yen-Yu Lin |
ICCV | 3 |
| 2023 | 360MVSNet: Deep Multi-view Stereo Network with 360° Images for Indoor Scene ReconstructionabstractRecent multi-view stereo methods have achieved promising results with the advancement of deep learning techniques. Despite of the progress, due to the limited fields of view of regular images, reconstructing large indoor environments still requires collecting many images with sufficient visual overlap, which is quite labor-intensive. 360° images cover a much larger field of view than regular images and would facilitate the capture process. In this paper, we present 360MVSNet, the first deep learning network for multi-view stereo with 360° images. Our method combines uncertainty estimation with a spherical sweeping module for 360° images captured from multiple viewpoints in order to construct multi-scale cost volumes. By regressing volumes in a coarse-to-fine manner, high-resolution depth maps can be obtained. Furthermore, we have constructed EQMVS, a large-scale synthetic dataset that consists of over 50K pairs of RGB and depth maps in equirectangular projection. Experimental results demonstrate that our method can reconstruct large synthetic and real-world indoor scenes with significantly better completeness than previous traditional and learning-based methods while saving both time and effort in the data acquisition process. Ching-Ya Chiu, Yu-Ting Wu 0001, I-Chao Shen, Yung-Yu Chuang |
WACV | 4 |
| 2023 | Physically Plausible Animation of Human Upper Body from a Single ImageabstractWe present a new method for generating controllable, dynamically responsive, and photorealistic human animations. Given an image of a person, our system allows the user to generate Physically plausible Upper Body Animation (PUBA) using interaction in the image space, such as dragging their hand to various locations. We formulate a reinforcement learning problem to train a dynamic model that predicts the person’s next 2D state (i.e., keypoints on the image) conditioned on a 3D action (i.e., joint torque), and a policy that outputs optimal actions to control the person to achieve desired goals. The dynamic model leverages the expressiveness of 3D simulation and the visual realism of 2D videos. PUBA generates 2D keypoint sequences that achieve task goals while being responsive to forceful perturbation. The sequences of keypoints are then translated by a pose-to-image generator to produce the final photorealistic video. Zhengping Zhou, Yung-Yu Chuang, Jiajun Wu 0001, C. Karen Liu |
WACV | 3 |
| 2022 | StyleFaceUV: a 3D Face UV Map Generator for View-Consistent Face Image Synthesis
Wei-Chieh Chung, Jiankai Zhu, I-Chao Shen, Yu-Ting Wu 0001, Yung-Yu Chuang |
BMVC | 5 |
| 2022 | ScannerNet: A Deep Network for Scanner-Quality Document Images under Complex Illumination
Chih-Jou Hsu, Yu-Ting Wu 0001, Ming-Sui Lee, Yung-Yu Chuang |
BMVC | 4 |
| 2022 | SearchTrack: Multiple Object Tracking with Object-Customized Search and Motion-Aware Features
Zhong-Min Tsai, Yu-Ju Tsai, Chien-Yao Wang, Hong-Yuan Mark Liao, Youn-Long Lin, Yung-Yu Chuang |
BMVC | 6 |
| 2022 | An MIL-Derived Transformer for Weakly Supervised Point Cloud SegmentationabstractWe address weakly supervised point cloud segmentation by proposing a new model, MIL-derived transformer, to mine additional supervisory signals. First, the transformer model is derived based on multiple instance learning (MIL) to explore pair-wise cloud-level supervision, where two clouds of the same category yield a positive bag while two of different classes produce a negative bag. It leverages not only individual cloud annotations but also pair-wise cloud semantics for model optimization. Second, Adaptive global weighted pooling (AdaGWP) is integrated into our transformer model to replace max pooling and average pooling. It introduces learnable weights to re-scale logits in the class activation maps. It is more robust to noise while discovering more complete foreground points under weak supervision. Third, we perform point subsampling and enforce feature equivariance between the original and subsampled point clouds for regularization. The proposed method is end-to-end trainable and is general because it can work with different backbones with diverse types of weak supervision signals, including sparsely annotated points and cloud-level labels. The experiments show that it achieves state-of-the-art performance on the S3DIS and ScanNet benchmarks. The source code will be available at https://github.com/jimmy15923/wspss_mil_transformer. Cheng-Kun Yang, Ji-Jia Wu, Kai-Syun Chen, Yung-Yu Chuang, Yen-Yu Lin |
CVPR | 4 |
| 2022 | Point MixSwap: Attentional Point Cloud Mixing via Swapping Matched Structural Divisions
Ardian Umam, Cheng-Kun Yang, Yung-Yu Chuang, Jen-Hui Chuang, Yen-Yu Lin |
ECCV (29) | 3 |
| 2022 | Learning to See Through Obstructions With Layered DecompositionabstractWe present a learning-based approach for removing unwanted obstructions, such as window reflections, fence occlusions, or adherent raindrops, from a short sequence of images captured by a moving camera. Our method leverages motion differences between the background and obstructing elements to recover both layers. Specifically, we alternate between estimating dense optical flow fields of the two layers and reconstructing each layer from the flow-warped images via a deep convolutional neural network. This learning-based layer reconstruction module facilitates accommodating potential errors in the flow estimation and brittle assumptions, such as brightness consistency. We show that the proposed approach learned from synthetically generated data performs well to real images. Experimental results on numerous challenging scenarios of reflection and fence removal demonstrate the effectiveness of the proposed method. Yu-Lun Liu 0001, Wei-Sheng Lai, Ming-Hsuan Yang 0001, Yung-Yu Chuang, Jia-Bin Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Hybrid Neural Fusion for Full-frame Video StabilizationabstractExisting video stabilization methods often generate visible distortion or require aggressive cropping of frame boundaries, resulting in smaller field of views. In this work, we present a frame synthesis algorithm to achieve full-frame video stabilization. We first estimate dense warp fields from neighboring frames and then synthesize the stabilized frame by fusing the warped contents. Our core technical novelty lies in the learning-based hybrid-space fusion that alleviates artifacts caused by optical flow inaccuracy and fast-moving objects. We validate the effectiveness of our method on the NUS, selfie, and DeepStab video datasets. Extensive experiment results demonstrate the merits of our approach over prior video stabilization methods. Yu-Lun Liu 0001, Wei-Sheng Lai, Ming-Hsuan Yang 0001, Yung-Yu Chuang, Jia-Bin Huang 0001 |
ICCV | 4 |
| 2021 | Unsupervised Point Cloud Object Co-segmentation by Co-contrastive Learning and Mutual Attention SamplingabstractThis paper presents a new task, point cloud object co-segmentation, aiming to segment the common 3D objects in a set of point clouds. We formulate this task as an object point sampling problem, and develop two techniques, the mutual attention module and co-contrastive learning, to enable it. The proposed method employs two point samplers based on deep neural networks, the object sampler and the background sampler. The former targets at sampling points of common objects while the latter focuses on the rest. The mutual attention module explores point-wise correlation across point clouds. It is embedded in both samplers and can identify points with strong cross-cloud correlation from the rest. After extracting features for points selected by the two samplers, we optimize the networks by developing the co-contrastive loss, which minimizes feature discrepancy of the estimated object points while maximizing feature separation between the estimated object and back-ground points. Our method works on point clouds of an arbitrary object class. It is end-to-end trainable and does not need point-level annotations. It is evaluated on the ScanObjectNN and S3DIS datasets and achieves promising results. The source code will be available at https://github.com/jimmy15923/unsup_point_coseg. Cheng-Kun Yang, Yung-Yu Chuang, Yen-Yu Lin |
ICCV | 2 |
| 2021 | Soft Ranking Threshold Losses For Image RetrievalabstractThis paper proposes a novel loss, soft ranking threshold loss, for driving deep networks to learn better representations for image retrieval. Instead of working in the metric space, our loss works in the rank space which has a more uniform distribution and explicit scale and bounds. Our loss reduces the ranks of the distances between anchor-positive pairs below the threshold while increasing the ones between anchor-negative pairs above the threshold. In addition to the basic form, two extensions are proposed for improving the effectiveness: hard thresholds and ranking margin. Experiments show that the proposed loss outperforms the state-of-the-art losses on image retrieval applications. Chiao-An Yang, Zhixiang Wang 0001, Yen-Yu Lin, Yung-Yu Chuang |
ICIP | 4 |
| 2021 | Learning to cluster for rendering with many lightsabstractWe present an unbiased online Monte Carlo method for rendering with many lights. Our method adapts both the hierarchical light clustering and the sampling distribution to our collected samples. Designing such a method requires us to make clustering decisions under noisy observation, and making sure that the sampling distribution adapts to our target. Our method is based on two key ideas: a coarse-to-fine clustering scheme that can find good clustering configurations even with noisy samples, and a discrete stochastic successive approximation method that starts from a prior distribution and provably converges to a target distribution. We compare to other state-of-the-art light sampling methods, and show better results both numerically and visually. Yu-Ting Wu 0001, Tzu-Mao Li, Yung-Yu Chuang |
ACM Trans. Graph. | 4 |
| 2020 | Attention-Based View Selection Networks for Light-Field Disparity EstimationabstractThis paper introduces a novel deep network for estimating depth maps from a light field image. For utilizing the views more effectively and reducing redundancy within views, we propose a view selection module that generates an attention map indicating the importance of each view and its potential for contributing to accurate depth estimation. By exploring the symmetric property of light field views, we enforce symmetry in the attention map and further improve accuracy. With the attention map, our architecture utilizes all views more effectively and efficiently. Experiments show that the proposed method achieves state-of-the-art performance in terms of accuracy and ranks the first on a popular benchmark for disparity estimation for light field images. Yu-Ju Tsai, Yu-Lun Liu 0001, Ouhyoung Ming, Yung-Yu Chuang |
AAAI | 4 |
| 2020 | BEDSR-Net: A Deep Shadow Removal Network From a Single Document ImageabstractRemoving shadows in document images enhances both the visual quality and readability of digital copies of documents. Most existing shadow removal algorithms for document images use hand-crafted heuristics and are often not robust to documents with different characteristics. This paper proposes the Background Estimation Document Shadow Removal Network (BEDSR-Net), the first deep network specifically designed for document image shadow removal. For taking advantage of specific properties of document images, a background estimation module is designed for extracting the global background color of the document. During the process of estimating the background color, the module also learns information about the spatial distribution of background and non-background pixels. We encode such information into an attention map. With the estimated global background color and attention map, the shadow removal network can better recover the shadow-free image. We also show that the model trained on synthetic images remains effective for real photos, and provide a large set of synthetic shadow images of documents along with their corresponding shadow-free images and shadow masks. Extensive quantitative and qualitative experiments on several benchmarks show that the BEDSR-Net outperforms existing methods in enhancing both the visual quality and readability of document images. Yun-Hsuan Lin, Wen-Chin Chen, Yung-Yu Chuang |
CVPR | 3 |
| 2020 | Learning to See Through ObstructionsabstractWe present a learning-based approach for removing unwanted obstructions, such as window reflections, fence occlusions or raindrops, from a short sequence of images captured by a moving camera. Our method leverages the motion differences between the background and the obstructing elements to recover both layers. Specifically, we alternate between estimating dense optical flow fields of the two layers and reconstructing each layer from the flow-warped images via a deep convolutional neural network. The learning-based layer reconstruction allows us to accommodate potential errors in the flow estimation and brittle assumptions such as brightness consistency. We show that training on synthetically generated data transfers well to real images. Our results on numerous challenging scenarios of reflection and fence removal demonstrate the effectiveness of the proposed method. Yu-Lun Liu 0001, Wei-Sheng Lai, Ming-Hsuan Yang 0001, Yung-Yu Chuang, Jia-Bin Huang 0001 |
CVPR | 4 |
| 2020 | Single-Image HDR Reconstruction by Learning to Reverse the Camera PipelineabstractRecovering a high dynamic range (HDR) image from a single low dynamic range (LDR) input image is challenging due to missing details in under-/over-exposed regions caused by quantization and saturation of camera sensors. In contrast to existing learning-based methods, our core idea is to incorporate the domain knowledge of the LDR image formation pipeline into our model. We model the HDR-to-LDR image formation pipeline as the (1) dynamic range clipping, (2) non-linear mapping from a camera response function, and (3) quantization. We then propose to learn three specialized CNNs to reverse these steps. By decomposing the problem into specific sub-tasks, we impose effective physical constraints to facilitate the training of individual sub-networks. Finally, we jointly fine-tune the entire model end-to-end to reduce error accumulation. With extensive quantitative and qualitative experiments on diverse image datasets, we demonstrate that the proposed method performs favorably against state-of-the-art single-image HDR reconstruction algorithms. Yu-Lun Liu 0001, Wei-Sheng Lai, Yu-Sheng Chen, Yi-Lung Kao, Ming-Hsuan Yang 0001, Yung-Yu Chuang, Jia-Bin Huang 0001 |
CVPR | 6 |
| 2020 | Domain-Specific Mappings for Generative Adversarial Style Transfer
Hsin-Yu Chang, Zhixiang Wang 0001, Yung-Yu Chuang |
ECCV (8) | 3 |
| 2020 | Deep Exposure Fusion with Deghosting via Homography Estimation and Attention LearningabstractModern cameras have limited dynamic ranges and often produce images with saturated or dark regions using a single exposure. Although the problem could be addressed by taking multiple images with different exposures, exposure fusion methods need to deal with ghosting artifacts and detail loss caused by camera motion or moving objects. This paper proposes a deep network for exposure fusion. For reducing the potential ghosting problem, our network only takes two images, an underexposed image and an overexposed one. Our network integrates together thew homography estimation for compensating camera motion, the attention mechanism for correcting remaining misalignment and moving pixels, and adversarial learning for alleviating other remaining artifacts. Experiments on real-world photos taken using handheld mobile phones show that the proposed method can generate high-quality images with faithful detail and vivid color rendition in both dark and bright areas. Sheng-Yeh Chen, Yung-Yu Chuang |
ICASSP | 2 |
| 2020 | Shadow Removal of Text Document Images by Estimating Local and Global Background ColorsabstractThis paper proposes a simple yet effective method for removing shadows from text document images. Assuming that the document mainly contains texts, our method estimates the global and local background colors using statistical analysis of the whole image and local neighborhoods. By estimating the global and local background colors, we obtain the shadow map indicating the shadow ratio for each pixel. With the shadow map, a shadow-free image can be recovered by intrinsic decomposition. Experiments confirm that our method effectively removes the shadow of text document images regardless of the intensity, scope, and the number of shadows. Jian-Ren Wang, Yung-Yu Chuang |
ICASSP | 2 |
| 2020 | Deep Co-Saliency Detection via Stacked Autoencoder-Enabled Fusion and Self-Trained CNNsabstractImage co-saliency detection via fusion-based or learning-based methods faces cross-cutting issues. Fusion-based methods often combine saliency proposals using a majority voting rule. Their performance hence highly depends on the quality and coherence of individual proposals. Learning-based methods typically require ground-truth annotations for training, which are not available for co-saliency detection. In this work, we present a two-stage approach to address these issues jointly. At the first stage, an unsupervised deep learning model with stacked autoencoder (SAE) is proposed to evaluate the quality of saliency proposals. It employs latent representations for image foregrounds, and auto-encodes foreground consistency and foreground-background distinctiveness in a discriminative way. The resultant model, SAE-enabled fusion (SAEF), can combine multiple saliency proposals to yield a more reliable saliency map. At the second stage, motivated by the fact that fusion often leads to over-smoothed saliency maps, we develop self-trained convolutional neural networks (STCNN) to alleviate this negative effect.STCNNtakes the saliency maps produced bySAEFas inputs. It propagates information from regions of high confidence to those of low confidence. During propagation, feature representations are distilled, resulting in sharper and better co-saliency maps. Our approach is comprehensively evaluated on three benchmarks, including MSRC, iCoseg, and Cosal2015, and performs favorably against the state-of-the-arts. In addition, we demonstrate that our method can be applied to object co-segmentation and object co-localization, achieving the state-of-the-art performance in both applications. Chung-Chi Tsai, Kuang-Jui Hsu, Yen-Yu Lin, Xiaoning Qian, Yung-Yu Chuang |
IEEE Trans. Multim. | 5 |
| 2020 | Illumination-Adaptive Person Re-IdentificationabstractMost person re-identification (ReID) approaches assume that person images are captured under relatively similar illumination conditions. In reality, long-term person retrieval is common, and person images are often captured under different illumination conditions at different times across a day. In this situation, the performances of existing ReID models often degrade dramatically. This paper addresses the ReID problem with illumination variations and names it as Illumination-Adaptive Person Re-identification (IA-ReID). We propose an Illumination-Identity Disentanglement (IID) network to dispel different scales of illuminations away while preserving individuals' identity information. To demonstrate the illumination issue and to evaluate our model, we construct two large-scale simulated datasets with a wide range of illumination variations. Experimental results on the simulated datasets and real-world images demonstrate the effectiveness of the proposed framework. Zelong Zeng, Zhixiang Wang 0001, Zheng Wang 0007, Yinqiang Zheng, Yung-Yu Chuang, Shin'ichi Satoh 0001 |
IEEE Trans. Multim. | 5 |
| 2019 | Deep Video Frame Interpolation Using Cyclic Frame GenerationabstractVideo frame interpolation algorithms predict intermediate frames to produce videos with higher frame rates and smooth view transitions given two consecutive frames as inputs. We propose that: synthesized frames are more reliable if they can be used to reconstruct the input frames with high quality. Based on this idea, we introduce a new loss term, the cycle consistency loss. The cycle consistency loss can better utilize the training data to not only enhance the interpolation results, but also maintain the performance better with less training data. It can be integrated into any frame interpolation network and trained in an end-to-end manner. In addition to the cycle consistency loss, we propose two extensions: motion linearity loss and edge-guided training. The motion linearity loss approximates the motion between two input frames to be linear and regularizes the training. By applying edge-guided training, we further improve results by integrating edge information into training. Both qualitative and quantitative experiments demonstrate that our model outperforms the state-of-the-art methods. The source codes of the proposed method and more experimental results will be available at https://github.com/alex04072000/CyclicGen. Yu-Lun Liu 0001, Yi-Tung Liao, Yen-Yu Lin, Yung-Yu Chuang |
AAAI | 4 |
| 2019 | DeepCO3: Deep Instance Co-Segmentation by Co-Peak Search and Co-Saliency DetectionabstractIn this paper, we address a new task called instance co-segmentation. Given a set of images jointly covering object instances of a specific category, instance co-segmentation aims to identify all of these instances and segment each of them, i.e. generating one mask for each instance. This task is important since instance-level segmentation is preferable for humans and many vision applications. It is also challenging because no pixel-wise annotated training data are available and the number of instances in each image is unknown. We solve this task by dividing it into two sub-tasks, co-peak search and instance mask segmentation. In the former sub-task, we develop a CNN-based network to detect the co-peaks as well as co-saliency maps for a pair of images. A co-peak has two endpoints, one in each image, that are local maxima in the response maps and similar to each other. Thereby, the two endpoints are potentially covered by a pair of instances of the same category. In the latter subtask, we design a ranking function that takes the detected co-peaks and co-saliency maps as inputs and can select the object proposals to produce the final results. Our method for instance co-segmentation and its variant for object colocalization are evaluated on four datasets, and achieve favorable performance against the state-of-the-art methods. The source codes and the collected datasets are available at https://github.com/KuangJuiHsu/DeepCO3/ Kuang-Jui Hsu, Yen-Yu Lin, Yung-Yu Chuang |
CVPR | 3 |
| 2019 | Learning to Reduce Dual-Level Discrepancy for Infrared-Visible Person Re-IdentificationabstractInfrared-Visible person RE-IDentification (IV-REID) is a rising task. Compared to conventional person re-identification (re-ID), IV-REID concerns the additional modality discrepancy originated from the different imaging processes of spectrum cameras, in addition to the person's appearance discrepancy caused by viewpoint changes, pose variations and deformations presented in the conventional re-ID task. The co-existed discrepancies make IV-REID more difficult to solve. Previous methods attempt to reduce the appearance and modality discrepancies simultaneously using feature-level constraints. It is however difficult to eliminate the mixed discrepancies using only feature-level constraints. To address the problem, this paper introduces a novel Dual-level Discrepancy Reduction Learning (D$^2$RL) scheme which handles the two discrepancies separately. For reducing the modality discrepancy, an image-level sub-network is trained to translate an infrared image into its visible counterpart and a visible image to its infrared version. With the image-level sub-network, we can unify the representations for images with different modalities. With the help of the unified multi-spectral images, a feature-level sub-network is trained to reduce the remaining appearance discrepancy through feature embedding. By cascading the two sub-networks and training them jointly, the dual-level reductions take their responsibilities cooperatively and attentively. Extensive experiments demonstrate the proposed approach outperforms the state-of-the-art methods. Zhixiang Wang 0001, Zheng Wang 0007, Yinqiang Zheng, Yung-Yu Chuang, Shin'ichi Satoh 0001 |
CVPR | 4 |
| 2019 | Polarimetric Camera Calibration Using an LCD MonitorabstractIt is crucial for polarimetric imaging to accurately calibrate the polarizer angles and the camera response function (CRF) of a polarizing camera. When this polarizing camera is used in a setting of multiview geometric imaging, it is often required to calibrate its intrinsic and extrinsic parameters as well, for which Zhang's calibration method is the most widely used with either a physical checker board, or more conveniently a virtual checker pattern displayed on a monitor. In this paper, we propose to jointly calibrate the polarizer angles and the inverse CRF (ICRF) using a slightly adapted checker pattern displayed on a liquid crystal display (LCD) monitor. Thanks to the lighting principles and the industry standards of the LCD monitors, the polarimetric and radiometric calibration can be significantly simplified, when assisted by the extrinsic parameters estimated from the checker pattern. We present a simple linear method for polarizer angle calibration and a convex method for radiometric calibration, both of which can be jointly refined in a process similar to bundle adjustment. Experiments have verified the feasibility and accuracy of the proposed calibration method. Zhixiang Wang 0001, Yinqiang Zheng, Yung-Yu Chuang |
CVPR | 3 |
| 2019 | FSA-Net: Learning Fine-Grained Structure Aggregation for Head Pose Estimation From a Single ImageabstractThis paper proposes a method for head pose estimation from a single image. Previous methods often predict head poses through landmark or depth estimation and would require more computation than necessary. Our method is based on regression and feature aggregation. For having a compact model, we employ the soft stagewise regression scheme. Existing feature aggregation methods treat inputs as a bag of features and thus ignore their spatial relationship in a feature map. We propose to learn a fine-grained structure mapping for spatially grouping features before aggregation. The fine-grained structure provides part-based information and pooled values. By utilizing learnable and non-learnable importance over the spatial location, different model variants can be generated and form a complementary ensemble. Experiments show that our method outperforms the state-of-the-art methods including both the landmark-free ones and the ones based on landmark or depth estimation. With only a single RGB frame as input, our method even outperforms methods utilizing multi-modality information (RGB-D, RGB-Time) on estimating the yaw angle. Furthermore, the memory overhead of our model is 100 times smaller than those of previous methods. Tsun-Yi Yang, Yen-Yu Lin, Yung-Yu Chuang |
CVPR | 4 |
| 2019 | Weakly Supervised Instance Segmentation using the Bounding Box Tightness PriorabstractThis paper presents a weakly supervised instance segmentation method that consumes training data with tight bounding box annotations. The major difficulty lies in the uncertain figure-ground separation within each bounding box since there is no supervisory signal about it. We address the difficulty by formulating the problem as a multiple instance learning (MIL) task, and generate positive and negative bags based on the sweeping lines of each bounding box. The proposed deep model integrates MIL into a fully supervised instance segmentation network, and can be derived by the objective consisting of two terms, i.e., the unary term and the pairwise term. The former estimates the foreground and background areas of each bounding box while the latter maintains the unity of the estimated object masks. The experimental results show that our method performs favorably against existing weakly supervised methods and even surpasses some fully supervised methods for instance segmentation on the PASCAL VOC dataset. Cheng-Chun Hsu, Kuang-Jui Hsu, Chung-Chi Tsai, Yen-Yu Lin, Yung-Yu Chuang |
NeurIPS | 5 |
| 2019 | A Low-Cost Portable Polycamera for Stereoscopic 360° ImagingabstractThis paper proposes a low-cost and portable polycamera system and accompanying methods for capturing and synthesizing stereoscopic 360° panoramas. The polycamera consists of only four cameras with fisheye lenses. Synthesizing panoramas from only four views is challenging because the cameras view very differently and the captured images have significant distortions and color degradation including vignetting, contrast loss, and blurriness. For coping with these challenges, this paper proposes methods for rectifying the polyview images, estimating depth of the scene, and synthesizing stereoscopic panoramas. The proposed camera is compact in size, light in weight, and inexpensive. The proposed methods allow the synthesis of visually pleasing stereoscopic 360° panoramas using the images captured with the proposed polycamera. We have built a prototype of the polycamera and tested it on a set of scenes with different characteristics of depth ranges and depth variations. The experiments show that the proposed camera and methods are effective in generating stereoscopic 360° panoramas that can be viewed on popular virtual reality displays. Hong Shiang Lin, Chao-Chin Chang, Hsu-Yu Chang, Yung-Yu Chuang, Tzong-Li Lin, Ouhyoung Ming |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Weakly Supervised Salient Object Detection by Learning A Classifier-Driven Map GeneratorabstractTop-down saliency detection aims to highlight the regions of a specific object category, and typically relies on pixel-wise annotated training data. In this paper, we address the high cost of collecting such training data by a weakly supervised approach to object saliency detection, where only image-level labels, indicating the presence or absence of a target object in an image, are available. The proposed framework is composed of two collaborative CNN modules, an image-level classifier and a pixel-level map generator. While the former distinguishes images with objects of interest from the rest, the latter is learned to generate saliency maps by which the images masked by the maps can be better predicted by the former. In addition to the top-down guidance from class labels, the map generator is derived by also exploring other cues, including the background prior, superpixel- and object proposal-based evidence. The background prior is introduced to reduce false positives. Evidence from superpixels helps preserve sharp object boundaries. The clue from object proposals improves the integrity of highlighted objects. These different types of cues greatly regularize the training process and reduces the risk of overfitting, which happens frequently when learning CNN models with few training data. Experiments show that our method achieves superior results, even outperforming fully supervised methods. Kuang-Jui Hsu, Yen-Yu Lin, Yung-Yu Chuang |
IEEE Trans. Image Process. | 3 |
| 2018 | Deep Photo Enhancer: Unpaired Learning for Image Enhancement From Photographs With GANsabstractThis paper proposes an unpaired learning method for image enhancement. Given a set of photographs with the desired characteristics, the proposed method learns a photo enhancer which transforms an input image into an enhanced image with those characteristics. The method is based on the framework of two-way generative adversarial networks (GANs) with several improvements. First, we augment the U-Net with global features and show that it is more effective. The global U-Net acts as the generator in our GAN model. Second, we improve Wasserstein GAN (WGAN) with an adaptive weighting scheme. With this scheme, training converges faster and better, and is less sensitive to parameters than WGAN-GP. Finally, we propose to use individual batch normalization layers for generators in two-way GANs. It helps generators better adapt to their own input distributions. All together, they significantly improve the stability of GAN training for our application. Both quantitative and visual results show that the proposed method is effective for enhancing images. Yu-Sheng Chen, Yu-Ching Wang, Man-Hsin Kao, Yung-Yu Chuang |
CVPR | 4 |
| 2018 | Unsupervised CNN-Based Co-saliency Detection with Graphical Optimization
Kuang-Jui Hsu, Chung-Chi Tsai, Yen-Yu Lin, Xiaoning Qian, Yung-Yu Chuang |
ECCV (5) | 5 |
| 2018 | Generating a Perspective Image from a Panoramic Image by the Swung-to-Cylinder ProjectionabstractThis paper proposes a swung-to-cylinder projection model for mapping a sphere to a plane. It can be used to create a semi-perspective image from a panoramic image. The model has two steps. In the first step, the sphere is projected onto a swung surface constructed by a circular profile and a rounded rectangular trajectory. In the second step, the projected image on the swung surface is mapped onto a cylinder through the perspective projection. We also propose methods for automatically determining proper parameters for the projection model based on image content. The proposed model is simple, efficient and easy to control. Experiments and analysis demonstrate its effectiveness. Che-Han Chang, Wei-Sheng Lai, Yung-Yu Chuang |
ICIP | 3 |
| 2018 | A 2.5D Approach to 360 Panorama Video StabilizationabstractThis paper presents a method for stabilizing both cylindrical and spherical panorama videos with a 360-degree field of view. We observe that rotation needs to be extremely smooth for 360 videos to maintain global motion coherency and avoid wobbling. Our method decouples the rotation from other motions and applies different strategies for smoothing them. The proposed approach is 2.5D as it estimates 3D rotations without involving 3D structure-from-motion methods. Therefore, it is more robust and can be performed in an incremental way. Experiments show that our method is effective in making steady 360 cylindrical/spherical videos. Lin-Chen Shen, Tzu-Kuei Huang, Chu-Song Chen, Yung-Yu Chuang |
ICIP | 4 |
| 2018 | Co-attention CNNs for Unsupervised Object Co-segmentationabstractObject co-segmentation aims to segment the common objects in images. This paper presents a CNN-based method that is unsupervised and end-to-end trainable to better solve this task. Our method is unsupervised in the sense that it does not require any training data in the form of object masks but merely a set of images jointly covering objects of a specific class. Our method comprises two collaborative CNN modules, a feature extractor and a co-attention map generator. The former module extracts the features of the estimated objects and backgrounds, and is derived based on the proposed co-attention loss which minimizes inter-image object discrepancy while maximizing intra-image figure-ground separation. The latter module is learned to generated co-attention maps by which the estimated figure-ground segmentation can better fit the former module. Besides, the co-attention loss, the mask loss is developed to retain the whole objects and remove noises. Experiments show that our method achieves superior results, even outperforming the state-of-the-art, supervised methods. Kuang-Jui Hsu, Yen-Yu Lin, Yung-Yu Chuang |
IJCAI | 3 |
| 2018 | SSR-Net: A Compact Soft Stagewise Regression Network for Age EstimationabstractThis paper presents a novel CNN model called Soft Stagewise Regression Network (SSR-Net) for age estimation from a single image with a compact model size. Inspired by DEX, we address age estimation by performing multi-class classification and then turning classification results into regression by calculating the expected values. SSR-Net takes a coarse-to-fine strategy and performs multi-class classification with multiple stages. Each stage is only responsible for refining the decision of its previous stage for more accurate age estimation. Thus, each stage performs a task with few classes and requires few neurons, greatly reducing the model size. For addressing the quantization issue introduced by grouping ages into classes, SSR-Net assigns a dynamic range to each age class by allowing it to be shifted and scaled according to the input face image. Both the multi-stage strategy and the dynamic range are incorporated into the formulation of soft stagewise regression. A novel network architecture is proposed for carrying out soft stagewise regression. The resultant SSR-Net model is very compact and takes only 0.32 MB. Despite its compact size, SSR-Net’s performance approaches those of the state-of-the-art methods whose model sizes are often more than 1500× larger. Tsun-Yi Yang, Yi-Hsuan Huang, Yen-Yu Lin, Pi-Cheng Hsiu, Yung-Yu Chuang |
IJCAI | 5 |
| 2017 | Weakly Supervised Saliency Detection with A Category-Driven Map Generator
Kuang-Jui Hsu, Yen-Yu Lin, Yung-Yu Chuang |
BMVC | 3 |
| 2017 | Deep Co-occurrence Feature Learning for Visual Object RecognitionabstractThis paper addresses three issues in integrating part-based representations into convolutional neural networks (CNNs) for object recognition. First, most part-based models rely on a few pre-specified object parts. However, the optimal object parts for recognition often vary from category to category. Second, acquiring training data with part-level annotation is labor-intensive. Third, modeling spatial relationships between parts in CNNs often involves an exhaustive search of part templates over multiple network streams. We tackle the three issues by introducing a new network layer, called co-occurrence layer. It can extend a convolutional layer to encode the co-occurrence between the visual parts detected by the numerous neurons, instead of a few pre-specified parts. To this end, the feature maps serve as both filters and images, and mutual correlation filtering is conducted between them. The co-occurrence layer is end-to-end trainable. The resultant co-occurrence features are rotation-and translation-invariant, and are robust to object deformation. By applying this new layer to the VGG-16 and ResNet-152, we achieve the recognition rates of 83.6% and 85.8% on the Caltech-UCSD bird benchmark, respectively. The source code is available at https://github.com/yafangshih/Deep-COOC. Ya-Fang Shih, Yang-Ming Yeh, Yen-Yu Lin, Ming-Fang Weng, Yi-Chang Lu, Yung-Yu Chuang |
CVPR | 6 |
| 2017 | DeepCD: Learning Deep Complementary Descriptors for Patch RepresentationsabstractThis paper presents the DeepCD framework which learns a pair of complementary descriptors jointly for image patch representation by employing deep learning techniques. It can be achieved by taking any descriptor learning architecture for learning a leading descriptor and augmenting the architecture with an additional network stream for learning a complementary descriptor. To enforce the complementary property, a new network layer, called data-dependent modulation (DDM) layer, is introduced for adaptively learning the augmented network stream with the emphasis on the training data that are not well handled by the leading stream. By optimizing the proposed joint loss function with late fusion, the obtained descriptors are complementary to each other and their fusion improves performance. Experiments on several problems and datasets show that the proposed method1 is simple yet effective, outperforming state-of-the-art methods. Tsun-Yi Yang, Jo-Han Hsu, Yen-Yu Lin, Yung-Yu Chuang |
ICCV | 4 |
| 2016 | Accumulated Stability Voting: A Robust Descriptor from Descriptors of Multiple ScalesabstractThis paper proposes a novel local descriptor through accumulated stability voting (ASV). The stability of feature dimensions is measured by their differences across scales. To be more robust to noise, the stability is further quantized by thresholding. The principle of maximum entropy is utilized for determining the best thresholds for maximizing discriminant power of the resultant descriptor. Accumulating stability renders a real-valued descriptor and it can be converted into a binary descriptor by an additional thresholding process. The real-valued descriptor attains high matching accuracy while the binary descriptor makes a good compromise between storage and accuracy. Our descriptors are simple yet effective, and easy to implement. In addition, our descriptors require no training. Experiments on popular benchmarks demonstrate the effectiveness of our descriptors and their superiority to the state-of-the-art descriptors. Tsun-Yi Yang, Yen-Yu Lin, Yung-Yu Chuang |
CVPR | 3 |
| 2016 | Natural Image Stitching with the Global Similarity Prior
Yu-Sheng Chen, Yung-Yu Chuang |
ECCV (5) | 2 |
| 2016 | A robust automatic object segmentation method for 3D printingabstract3D printing has become an important and prevalent tool. Image-based modeling is a popular way to acquire 3D models for further editing and printing. However, exiting tools are often not robust enough for users to obtain the 3D models they want. The constructed models are often incomplete, disjoint and noisy. This paper proposes a robust automatic method for segmenting an object out of the background using a set of multi-view images. With the segmentation, 3D reconstruction methods can be applied more robustly. The segmentation is performed by minimizing an energy function which incorporates color statistics, spatial coherency, appearance proximity, epipolar constraints and back projection consistency of 3D feature points. It can be efficiently optimized using the mincut algorithm. Experiments show that the proposed method can generate better models than some popular systems. Tzu-Kuei Huang, Ying-Hsuang Wang, Ta-Kai Lin, Yung-Yu Chuang |
ICME | 4 |
| 2016 | A Tool for Stereoscopic Parameter Setting Based on Geometric Perceived Depth PercentageabstractIt is a necessary but challenging task for creative producers to have an idea of the depth perception of the target audience when watching a stereoscopic film in a cinema during production. This paper proposes a novel metric, geometric perceived depth percentage (GPDP), to numerate and depict the depth perception of a scene before rendering. In addition to the geometric relationship between the object depth and focal distance, GPDP takes the screen width and viewing distance into account. As a result, it provides a more intuitive means for predicting stereoscopy and is universal across different viewing conditions. Based on GPDP, we design a practical tool to visualize the stereoscopic perception without the need for any 3-D device or special environment. The tool utilizes the stereoscopic comfort volume, GPDP-based shading schemes, depth perception markers, and GPDP histograms as visual cues so that animators can set stereoscopic parameters more easily. The tool is easily implemented in any modern rendering pipeline, including interactive Autodesk Maya and offline Pixar's RenderMan renderer. It has been used in several production projects including commercial ones. Finally, two user studies show that GPDP is a proper depth perception indicator and the proposed tool can make the stereoscopic parameter setting process easier and more efficient. Shuen-Huei Guan, Yu-Chi Lai, Kuo-Wei Chen, Hsuang-Ting Chou, Yung-Yu Chuang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2015 | Robust image alignment with multiple feature descriptors and matching-guided neighborhoodsabstractThis paper addresses two issues hindering the advances in accurate image alignment. First, he performance of descriptor-based approaches to image alignment relies on the chosen descriptor, but the optimal descriptor typically varies from image to image, or even pixel to pixel. Second, the neighborhood structure for smoothness enforcement is usually predefined before alignment. However, object boundaries are often better discovered during alignment. The proposed approach tackles the two issues by adaptive descriptor selection and dynamic neighborhood construction. Specifically we associate each pixel to be aligned with an affine transformation, and integrate the learning of the pixel-specific transformations into image alignment. The transformations serve as the common domain for descriptor fusion, since the local consensus of each descriptor can be estimated by accessing the corresponding affine transformation t allows us to pick the most plausible descriptor for aligning each pixel. On the other hand more object-aware neighborhoods can be produced by referencing the consistency between the learned affine transformations of neighboring pixels. The promising results on popular image alignment benchmarks manifests the effectiveness of our approach. Kuang-Jui Hsu, Yen-Yu Lin, Yung-Yu Chuang |
CVPR | 3 |
| 2015 | Blur kernel estimation using normalized color-line priorsabstractThis paper proposes a single-image blur kernel estimation algorithm that utilizes the normalized color-line prior to restore sharp edges without altering edge structures or enhancing noise. The proposed prior is derived from the color-line model, which has been successfully applied to non-blind deconvolution and many computer vision problems. In this paper, we show that the original color-line prior is not effective for blur kernel estimation and propose a normalized color-line prior which can better enhance edge contrasts. By optimizing the proposed prior, our method gradually enhances the sharpness of the intermediate patches without using heuristic filters or external patch priors. The intermediate patches can then guide the estimation of the blur kernel. A comprehensive evaluation on a large image deblurring dataset shows that our algorithm achieves the state-of-the-art results. Wei-Sheng Lai, Jian-Jiun Ding, Yen-Yu Lin, Yung-Yu Chuang |
CVPR | 4 |
| 2015 | Light field image editing by 4D patch synthesisabstractThis paper presents a patch-based synthesis framework for lightfield image editing. The core of the proposed method builds upon a patch-based optimization approach. The main contribution of the paper is to extend the versatile patch-based image editing framework to 4D lightfield images and enable many editing applications for them. Specifically, the paper introduces a novel 4D lightfield patch consistency measure for avoiding synthesis of inconsistent patches into the edited lightfield images. Combining with a joint 4D patch search, our method is able to maintain the correlation among views and render a consistent interpretation of the scene. The proposed method offers patch-based solutions to a wide variety of lightfield image editing problems, including inpainting, retargeting and reshuffling. Ke-Wei Chen, Ming-Hsu Chang, Yung-Yu Chuang |
ICME | 3 |
| 2015 | Geometrically Consistent Stereoscopic Image Editing Using Patch-Based SynthesisabstractThis paper presents a patch-based synthesis framework for stereoscopic image editing. The core of the proposed method builds upon a patch-based optimization framework with two key contributions: First, we introduce a depth-dependent patch-pair similarity measure for distinguishing and better utilizing image contents with different depth structures. Second, a joint patch-pair search is proposed for properly handling the correlation between two views. The proposed method successfully overcomes two main challenges of editing stereoscopic 3D media: (1) maintaining the depth interpretation, and (2) providing controllability of the scene depth. The method offers patch-based solutions to a wide variety of stereoscopic image editing problems, including depth-guided texture synthesis, stereoscopic NPR, paint by depth, content adaptation, and 2D to 3D conversion. Several challenging cases are demonstrated to show the effectiveness of the proposed method. The results of user studies also show that the proposed method produces stereoscopic images with good stereoscopics and visual quality. Sheng-Jie Luo, Ying-Tse Sun, I-Chao Shen, Bing-Yu Chen 0004, Yung-Yu Chuang |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2015 | Disambiguating Stereoscopic Transparency Using a Thaumatrope ApproachabstractVolume rendering is a popular visualization technique for scientific computing and medical imaging. By assigning proper transparency, it allows us to see more information inside the volume. However, because volume rendering projects complex 3D structures into the 2D domain, the resultant visualization often suffers from ambiguity and its spatial relationship could be difficult to recognize correctly, especially when the scene or setting is highly transparent. Stereoscopic displays are not the rescue to the problem even though they add an additional dimension which seems helpful for resolving the ambiguity. This paper proposes a thaumatrope method to enhance 3D understanding with stereoscopic transparency for volume rendering. Our method first generates an additional cue with less spatial ambiguity by using a high opacity setting. To avoid cluttering the actual content, we only select its prominent feature for displaying. By alternating the actual content and the selected feature quickly, the viewer only perceives a whole volume while its spatial understanding has been enhanced. A user study was performed to compare the proposed method with the original stereoscopic volume rendering and the static combination of the actual content and the selected feature using a 3D display. Results show that the proposed thaumatrope approach provides better spatial understanding than compared approaches. Yan-Jen Su, Yung-Yu Chuang |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2015 | Dual-Matrix Sampling for Scalable Translucent Material RenderingabstractThis paper introduces a scalable algorithm for rendering translucent materials with complex lighting. We represent the light transport with a diffusion approximation by a dual-matrix representation with the Light-to-Surface and Surface-to-Camera matrices. By exploiting the structures within the matrices, the proposed method can locate surface samples with little contribution by using only subsampled matrices and avoid wasting computation on these samples. The decoupled estimation of irradiance and diffuse BSSRDFs also allows us to have a tight error bound, making the adaptive diffusion approximation more efficient and accurate. Experiments show that our method outperforms previous methods for translucent material rendering, especially in large scenes with massive translucent surfaces shaded by complex illumination. Yu-Ting Wu 0001, Tzu-Mao Li, Yu-Hsun Lin, Yung-Yu Chuang |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2014 | Shape-Preserving Half-Projective Warps for Image StitchingabstractThis paper proposes a novel parametric warp which is a spatial combination of a projective transformation and a similarity transformation. Given the projective transformation relating two input images, based on an analysis of the projective transformation, our method smoothly extrapolates the projective transformation of the overlapping regions into the non-overlapping regions and the resultant warp gradually changes from projective to similarity across the image. The proposed warp has the strengths of both projective and similarity warps. It provides good alignment accuracy as projective warps while preserving the perspective of individual image as similarity warps. It can also be combined with more advanced local-warp-based alignment methods such as the as-projective-as-possible warp for better alignment accuracy. With the proposed warp, the field of view can be extended by stitching images with less projective distortion (stretched shapes and enlarged sizes). Che-Han Chang, Yoichi Sato 0001, Yung-Yu Chuang |
CVPR | 3 |
| 2014 | Spatially-Varying Image Warps for Scene AlignmentabstractThis paper proposes a method to align a set of images captured from multiple view points. Traditional methods using image warps parameterized by global transformations suffer from the problem of misalignment due to parallax effects induced by camera motions between images and depth variations of the scene. Our method parameterizes warps using mesh deformation and achieves spatially-varying transformations to alleviate the misalignment problem. The proposed method has two stages: a hybrid image alignment stage which combines direct-based methods and feature-based methods, followed by a shape-preserving aggregation stage which further refines the result. Experiments show that our method achieves better alignment and provides visually pleasing image summaries for scenes. Che-Han Chang, Chiu-Ju Chen, Yung-Yu Chuang |
ICPR | 3 |
| 2014 | Augmented Multiple Instance Regression for Inferring Object Contours in Bounding BoxesabstractIn this paper, we address the problem of the high annotation cost of acquiring training data for semantic segmentation. Most modern approaches to semantic segmentation are based upon graphical models, such as the conditional random fields, and rely on sufficient training data in form of object contours. To reduce the manual effort on pixel-wise annotating contours, we consider the setting in which the training data set for semantic segmentation is a mixture of a few object contours and an abundant set of bounding boxes of objects. Our idea is to borrow the knowledge derived from the object contours to infer the unknown object contours enclosed by the bounding boxes. The inferred contours can then serve as training data for semantic segmentation. To this end, we generate multiple contour hypotheses for each bounding box with the assumption that at least one hypothesis is close to the ground truth. This paper proposes an approach, called augmented multiple instance regression (AMIR), that formulates the task of hypothesis selection as the problem of multiple instance regression (MIR), and augments information derived from the object contours to guide and regularize the training process of MIR. In this way, a bounding box is treated as a bag with its contour hypotheses as instances, and the positive instances refer to the hypotheses close to the ground truth. The proposed approach has been evaluated on the Pascal VOC segmentation task. The promising results demonstrate that AMIR can precisely infer the object contours in the bounding boxes, and hence provide effective alternatives to manually labeled contours for semantic segmentation. Kuang-Jui Hsu, Yen-Yu Lin, Yung-Yu Chuang |
IEEE Trans. Image Process. | 3 |
| 2013 | Rectangling Stereographic Projection for Wide-Angle Image VisualizationabstractThis paper proposes a new projection model for mapping a hemisphere to a plane. Such a model can be useful for viewing wide-angle images. Our model consists of two steps. In the first step, the hemisphere is projected onto a swung surface constructed by a circular profile and a rounded rectangular trajectory. The second step maps the projected image on the swung surface onto the image plane through the perspective projection. We also propose a method for automatically determining proper parameters for the projection model based on image content. The proposed model has several advantages. It is simple, efficient and easy to control. Most importantly, it makes a better compromise between distortion minimization and line preserving than popular projection models, such as stereographic and Pannini projections. Experiments and analysis demonstrate the effectiveness of our model. Che-Han Chang, Min-Chun Hu 0001, Wen-Huang Cheng, Yung-Yu Chuang |
ICCV | 4 |
| 2013 | Target-Driven Moire Pattern Synthesis by Phase ModulationabstractThis paper investigates an approach for generating two grating images so that the moire pattern of their superposition resembles the target image. Our method is grounded on the fundamental moire theorem. By focusing on the visually most dominant (1,-1)-moire component, we obtain the phase modulation constraint on the phase shifts between the two grating images. For improving visual appearance of the grating images and hiding capability the embedded image, a smoothness term is added to spread information between the two grating images and an appearance phase function is used to add irregular structures into grating images. The grating images can be printed on transparencies and the hidden image decoding can be performed optically by overlaying them together. The proposed method enables the creation of moire art and allows visual decoding without computers. Pei-Hen Tsai, Yung-Yu Chuang |
ICCV | 2 |
| 2013 | VisibilityCluster: Average Directional Visibility for Many-Light RenderingabstractThis paper proposes the VisibilityCluster algorithm for efficient visibility approximation and representation in many-light rendering. By carefully clustering lights and shading points, we can construct a visibility matrix that exhibits good local structures due to visibility coherence of nearby lights and shading points. Average visibility can be efficiently estimated by exploiting the sparse structure of the matrix and shooting only few shadow rays between clusters. Moreover, we can use the estimated average visibility as a quality measure for visibility estimation, enabling us to locally refine VisibilityClusters with large visibility variance for improving accuracy. We demonstrate that, with the proposed method, visibility can be incorporated into importance sampling at a reasonable cost for the many-light problem, significantly reducing variance in Monte Carlo rendering. In addition, the proposed method can be used to increase realism of local shading by adding directional occlusion effects. Experiments show that the proposed technique outperforms state-of-the-art importance sampling algorithms, and successfully enhances the preview quality for lighting design. Yu-Ting Wu 0001, Yung-Yu Chuang |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2012 | A line-structure-preserving approach to image resizingabstractThis paper proposes a content-aware image resizing method which simultaneously preserves both salient image features and important line structure properties: parallelism, collinearity and orientation. When there are prominent line structures in the image, image resizing methods without explicitly taking these properties into account could produce line structure distortions in their results. Since the human visual system is very sensitive to line structures, such distortions often become noticeable and disturbing. Our method couples mesh deformations for image resizing with similarity transforms for line features. Mesh deformations are used to control content preservation while similarity transforms are analyzed in the Hough space to maintain line structure properties. Our method strikes a good balance between preserving content and maintaining line structure properties. Experiments show the proposed method often outperforms methods without taking line structures into account, especially for scenes with prominent line structures. Che-Han Chang, Yung-Yu Chuang |
CVPR | 2 |
| 2012 | Affinity aggregation for spectral clusteringabstractSpectral clustering makes use of spectral-graph structure of an affinity matrix to partition data into disjoint meaningful groups. Because of its elegance, efficiency and good performance, spectral clustering has become one of the most popular clustering methods. Traditional spectral clustering assumes a single affinity matrix. However, in many applications, there could be multiple potentially useful features and thereby multiple affinity matrices. To apply spectral clustering for these cases, a possible way is to aggregate the affinity matrices into a single one. Unfortunately, affinity measures constructed from different features could have different characteristics. Careless aggregation might make even worse clustering performance. This paper proposes an affinity aggregation spectral clustering (AASC) algorithm which extends spectral clustering to a setting with multiple affinities available. AASC seeks for an optimal combination of affinity matrices so that it is more immune to ineffective affinities and irrelevant features. This enables the construction of similarity or distance-metric measures for clustering less crucial. Experiments show that AASC is effective in simultaneous clustering and feature fusion, thus enhancing the performance of spectral clustering by employing multiple affinities. Hsin-Chien Huang, Yung-Yu Chuang, Chu-Song Chen |
CVPR | 2 |
| 2012 | Scene warping: Layer-based stereoscopic image resizingabstractThis paper proposes scene warping, a layer-based stereoscopic image resizing method using image warping. The proposed method decomposes the input stereoscopic image pair into layers according to the depth and color information. A quad mesh is placed onto each layer to guide the image warping for resizing. The warped layers are composited by their depth orders to synthesize the resized stereoscopic image. We formulate an energy function to guide the warping for each layer so that the composited image avoids distortions and holes, maintains good stereoscopic properties and contains as many important pixels as possible in the reduced image space. The proposed method offers the advantages of less discontinuous artifacts, less-distorted objects, correct depth ordering and enhanced stereoscopic quality. Experiments show that our method compares favorably with existing methods. Ken-Yi Lee, Cheng-Da Chung, Yung-Yu Chuang |
CVPR | 3 |
| 2012 | Multi-affinity spectral clusteringabstractSpectral clustering (SC) has become one of the most popular clustering methods. Given an affinity matrix, SC explores its spectral-graph structure to partition data into disjoint meaningful groups. However, in many applications, there are multiple potentially useful features and thereby multiple affinity matrices. For applying spectral clustering to such cases, these affinity matrices must be aggregated into a single one. Unfortunately, affinity measures based on different features could have different characteristics. Some are more effective than others. We propose a multi-affinity spectral clustering (MASC) algorithm which extends the SC algorithm with multiple affinities available. By automatically adjusting the weights of affinity matrices, MASC is more immune to ineffective affinities and irrelevant features. This makes the choice of similarity or distance-metric measures for clustering less crucial. Experiments show that MASC is effective in simultaneous clustering and feature fusion, thus maintaining robustness of SC for multi-affinity clustering problems. Hsin-Chien Huang, Yung-Yu Chuang, Chu-Song Chen |
ICASSP | 2 |
| 2012 | Warping-Based Novel View Synthesis from a Binocular Image for Autostereoscopic DisplaysabstractThis paper presents a warping-based method for synthesizing multiple views from a binocular stereoscopic image. Auto stereoscopic displays require multiple views while most stereoscopic cameras can only capture two. Popular novel view synthesis methods, such as depth image based rendering (DIBR), often heavily rely on accurate depth maps, which are still difficult to obtain. The proposed method requires neither depth maps nor user intervention. It extracts dense and reliable features. Feature correspondences guide image warping to synthesize novel views while simultaneously maintaining stereoscopic properties and preserving image structures. Compared to DIBR, the proposed method produces higher-quality multi-view images more efficiently without tedious parameter tuning. The method can be used to convert stereoscopic images taken by binocular cameras into multi-view images ready to be displayed on auto stereoscopic displays. Yu-Hsiang Huang, Tzu-Kuei Huang, Yan-Hsiang Huang, Wei-Chao Chen, Yung-Yu Chuang |
ICME | 5 |
| 2012 | Cross-Domain Multicue Fusion for Concept-Based Video IndexingabstractThe success of query-by-concept, proposed recently to cater to video retrieval needs, depends greatly on the accuracy of concept-based video indexing. Unfortunately, it remains a challenge to recognize the presence of concepts in a video segment or to extract an objective linguistic description from it because of the semantic gap, that is, the lack of correspondence between machine-extracted low-level features and human high-level conceptual interpretation. This paper studies three issues with the aim to reduce such a gap: 1) how to explore cues beyond low-level features, 2) how to combine diverse cues to improve performance, and 3) how to utilize the learned knowledge when applying it to a new domain. To solve these problems, we propose a framework that jointly exploits multiple cues across multiple video domains. First, recursive algorithms are proposed to learn both interconcept and intershot relationships from annotations. Second, all concept labels for all shots are simultaneously refined in a single fusion model. Additionally, unseen shots are assigned pseudolabels according to their initial prediction scores so that contextual and temporal relationships can be learned, thus requiring no additional human effort. Integration of cues embedded within training and testing video sets accommodates domain change. Experiments on popular benchmarks show that our framework is effective, achieving significant improvements over popular baselines. Ming-Fang Weng, Yung-Yu Chuang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Multiple Kernel Fuzzy ClusteringabstractWhile fuzzy c-means is a popular soft-clustering method, its effectiveness is largely limited to spherical clusters. By applying kernel tricks, the kernel fuzzy c-means algorithm attempts to address this problem by mapping data with nonlinear relationships to appropriate feature spaces. Kernel combination, or selection, is crucial for effective kernel clustering. Unfortunately, for most applications, it is uneasy to find the right combination. We propose a multiple kernel fuzzy c-means (MKFC) algorithm that extends the fuzzy c-means algorithm with a multiple kernel-learning setting. By incorporating multiple kernels and automatically adjusting the kernel weights, MKFC is more immune to ineffective kernels and irrelevant features. This makes the choice of kernels less crucial. In addition, we show multiple kernel k-means to be a special case of MKFC. Experiments on both synthetic and real-world data demonstrate the effectiveness of the proposed MKFC algorithm. Hsin-Chien Huang, Yung-Yu Chuang, Chu-Song Chen |
IEEE Trans. Fuzzy Syst. | 2 |
| 2012 | SURE-based optimization for adaptive sampling and reconstructionabstractWe apply Stein's Unbiased Risk Estimator (SURE) to adaptive sampling and reconstruction to reduce noise in Monte Carlo rendering. SURE is a general unbiased estimator for mean squared error (MSE) in statistics. With SURE, we are able to estimate error for an arbitrary reconstruction kernel, enabling us to use more effective kernels rather than being restricted to the symmetric ones used in previous work. It also allows us to allocate more samples to areas with higher estimated MSE. Adaptive sampling and reconstruction can therefore be processed within an optimization framework. We also propose an efficient and memory-friendly approach to reduce the impact of noisy geometry features where there is depth of field or motion blur. Experiments show that our method produces images with less noise and crisper details than previous methods. Tzu-Mao Li, Yu-Ting Wu 0001, Yung-Yu Chuang |
ACM Trans. Graph. | 3 |
| 2012 | Perspective-aware warping for seamless stereoscopic image cloningabstractThis paper presents a novel technique for seamless stereoscopic image cloning, which performs both shape adjustment and color blending such that the stereoscopic composite is seamless in both the perceived depth and color appearance. The core of the proposed method is an iterative disparity adaptation process which alternates between two steps: disparity estimation, which re-estimates the disparities in the gradient domain so that the disparities are continuous across the boundary of the cloned region; and perspective-aware warping, which locally re-adjusts the shape and size of the cloned region according to the estimated disparities. This process guarantees not only depth continuity across the boundary but also models local perspective projection in accordance with the disparities, leading to more natural stereoscopic composites. The proposed method allows for easy cloning of objects with intricate silhouettes and vague boundaries because it does not require precise segmentation of the objects. Several challenging cases are demonstrated to show that our method generates more compelling results compared to methods with only global shape adjustment. Sheng-Jie Luo, I-Chao Shen, Bing-Yu Chen 0004, Wen-Huang Cheng, Yung-Yu Chuang |
ACM Trans. Graph. | 5 |
| 2012 | Collaborative video reindexing via matrix factorizationabstractConcept-based video indexing generates a matrix of scores predicting the possibilities of concepts occurring in video shots. Based on the idea of collaborative filtering, this article presents unsupervised methods to refine the initial scores generated by concept classifiers by taking into account the concept-to-concept correlation and shot-to-shot similarity embedded within the score matrix. Given a noisy matrix, we refine the inaccurate scores via matrix factorization. This method is further improved by learning multiple local models and incorporating contextual-temporal structures. Experiments on the TRECVID 2006--2008 datasets demonstrate relative performance gains ranging from 13% to 52% without using any user annotations or external knowledge resources. Ming-Fang Weng, Yung-Yu Chuang |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2011 | 3D cinematography principles and their applications to stereoscopic media processingabstractThis paper introduces 3D cinematography principles to the field of multimedia and illustrates their usage in stereoscopic media processing applications. These principles include (1) maintaining coordination among views, (2) having a continuous depth chart, (3) placing rest areas between strong 3D shots, (4) using a shallow depth of field for shots with excessive depth brackets, and (5) being careful about the stereoscopic window. Taking these principles into account, we propose designs for stereoscopic extensions of two popular 2D media applications---video stabilization and photo slideshow---to provide a better 3D viewing experience. User studies show that by incorporating 3D cinematography principles, the proposed methods yield more comfortable and enjoyable 3D viewing experiences than those delivered using naive extensions of conventional 2D methods. Chun-Wei Liu, Tz-Huan Huang, Ming-Hsu Chang, Ken-Yi Lee, Chia-Kai Liang, Yung-Yu Chuang |
ACM Multimedia | 6 |
| 2011 | Content-Aware Display Adaptation and Interactive Editing for Stereoscopic ImagesabstractWe propose a content-aware stereoscopic image display adaptation method which simultaneously resizes a binocular image to the target resolution and adapts its depth to the comfort zone of the display while preserving the perceived shapes of prominent objects. This method does not require depth information or dense correspondences. Given the specification of the target display and a sparse set of correspondences, our method efficiently deforms the input stereoscopic images for display adaptation by solving a least-squares energy minimization problem. This can be used to adjust stereoscopic images to fit displays with different real estates, aspect ratios and comfort zones. In addition, with slight modifications to the energy function, our method allows users to interactively adjust the sizes, locations and depths of the selected objects, giving users aesthetic control for depth perception. User studies show that the method is effective at editing depth and reducing occurrences of diplopia and distortions. Che-Han Chang, Chia-Kai Liang, Yung-Yu Chuang |
IEEE Trans. Multim. | 3 |
| 2010 | Learning Landmarks by Exploiting Social Media
Chia-Kai Liang, Yu-Ting Hsieh, Tien-Jung Chuang, Ming-Fang Weng, Yung-Yu Chuang |
MMM | 6 |
| 2009 | A collaborative benchmark for region of interest detection algorithmsabstractThis paper presents a collaborative benchmark for region of interest (ROI) detection in images. ROI detection has many useful applications and many algorithms have been proposed to automatically detect ROIs. Unfortunately, due to the lack of benchmarks, these methods were often tested on small data sets that are not available to others, making fair comparisons of these methods difficult. Examples from many fields have shown that repeatable experiments using published benchmarks are crucial to the fast advancement of the fields. To fill the gap, this paper presents our design for a collaborative game, called Photoshoot, to collect human ROI annotations for constructing an ROI benchmark. Using this game, we have gathered a large number of annotations and fused them into aggregated ROI models. With these models, we are able to evaluate six ROI detection algorithms quantitatively. Tz-Huan Huang, Kai-Yin Cheng, Yung-Yu Chuang |
CVPR | 3 |
| 2009 | High dynamic range image reconstruction from hand-held camerasabstractThis paper presents a technique for reconstructing a high-quality high dynamic range (HDR) image from a set of differently exposed and possibly blurred images taken with a hand-held camera. Recovering an HDR image from differently exposed photographs has become very popular. However, it often requires a tripod to keep the camera still when taking photographs of different exposures. To ease the process, it is often preferred to use a hand-held camera. This, however, leads to two problems, misaligned photographs and blurred long-exposed photographs. To overcome these problems, this paper adapts an alignment method and proposes a method for HDR reconstruction from possibly blurred images. We use Bayesian framework to formulate the problem and apply a maximum-likelihood approach to iteratively perform blur kernel estimation, HDR image reconstruction and camera curve recovery. When convergence, we simultaneously obtain an HDR image with rich and clear structures, the camera response curve and blur kernels. To show the effectiveness of our method, we test our method on both synthetic and real photographs. The proposed method compares favorably to two other related methods in the experiments. Pei-Ying Lu, Tz-Huan Huang, Meng-Sung Wu, Yi-Ting Cheng, Yung-Yu Chuang |
CVPR | 5 |
| 2009 | Video stabilization using robust feature trajectoriesabstractThis paper proposes a new approach for video stabilization. Most existing video stabilization methods adopt a framework of three steps, motion estimation, motion compensation and image composition. Camera motion is often estimated based on pairwise registration between frames. Thus, these methods often assume static scenes or distant backgrounds. Furthermore, for scenes with moving objects, robust methods are required for finding the dominant motion. Such assumptions and judgements could lead to errors in motion parameters. Errors are compounded by motion compensation which smoothes motion parameters. This paper proposes a method to directly stabilize a video without explicitly estimating camera motion, thus assuming neither motion models nor dominant motion. The method first extracts robust feature trajectories from the input video. Optimization is then performed to find a set of transformations to smooth out these trajectories and stabilize the video. In addition, the optimization also considers quality of the stabilized video and selects a video with not only smooth camera motion but also less unfilled area after stabilization. Experiments show that our method can deal with complicated videos containing near, large and multiple moving objects. Ken-Yi Lee, Yung-Yu Chuang, Bing-Yu Chen 0004, Ouhyoung Ming |
ICCV | 2 |
| 2008 | Photo navigatorabstractNowadays, travel has become a popular activity for people to relax their body and mind. Taking photos is then often an inevitable and frequent event during one's trip for recording the enjoyable experience. To help people to relive the wonderful travel experience they had recorded in photos, this paper presents a system, Photo Navigator, for enhancing the photo browsing experience by creating a new browsing style with a realistic feel to users as being into the scenes and taking a trip back in time to revisit the place. The proposed system is characterized by two main features. First, it better reveals the spatial relations among photos and offers a strong sense of space by taking users to fly into the scenes. Second, it is fully automatic and makes plausible for novice users to utilize the 3D technologies that are traditionally complex to manipulate. The proposed system is compared with two other photo browsing tools, ACDSee's photo slideshow and Microsoft's PhotoStory. User studies show that people would comparatively favor the browsing style we offer and appreciate the ease to create such a style. Chi-Chang Hsieh, Wen-Huang Cheng, Chia-Hu Chang, Yung-Yu Chuang, Ja-Ling Wu |
ACM Multimedia | 4 |
| 2008 | Multi-cue fusion for semantic video indexingabstractThe huge amount of videos currently available poses a difficult problem in semantic video retrieval. The success of query-by-concept, recently proposed to handle this problem, depends greatly on the accuracy of concept-based video indexing. This paper describes a multi-cue fusion approach toward improving the accuracy of semantic video indexing. This approach is based on a unified framework that explores and integrates both contextual correlation among concepts and temporal dependency among shots. The framework is novel in two ways. First, a recursive algorithm is proposed to learn both inter-concept and inter-shot relationships from ground truth annotations of tens of thousands of shots for hundreds of concepts. Second, labels for all concepts and all shots are solved simultaneously through optimizing a graphical model. Experiments on the widely used TRECVID 2006 data set show that our framework is effective for semantic concept detection in video, achieving around a 30 % performance boost on two popular benchmarks, VIREO-374 and Columbia374, in inferred average precision. Ming-Fang Weng, Yung-Yu Chuang |
ACM Multimedia | 2 |
| 2008 | Interactive content presentation based on expressed emotion and physiological feedbackabstractIn this technical demonstration, we showcase an interactive content presentation (ICP) system that integrates media-expressed-emotion-based composition, user-perceived preference feedback, and interactive digital art creation. ICP harmonizes the browsing of multimedia contents by presenting them in the form of music videos (photos, blog articles with accompanied music) based on their expressed emotion similarity. ICP facilitates content browsing by automatically and dynamically selecting the media to be played next in real time, responding to user's preference feedback measured from physiological signals. In addition, ICP enhances the enjoyments of content browsing by incorporating interactive digital art creation. ICP achieves these goals by properly integrating recent researches on media-expressed emotion classification,cross-media composition, and physiological signal processing. Tien-Lin Wu, Hsuan-Kai Wang, Murphy Chien-Chang Ho, Yuan-Pin Lin, Ting-Ting Hu, Ming-Fang Weng, Li-Wei Chan 0001, Changhua Yang, Yi-Hsuan Yang, Yi-Ping Hung, Yung-Yu Chuang, Hsin-Hsi Chen, Homer H. Chen, Jyh-Horng Chen, Shyh-Kang Jeng |
ACM Multimedia | 11 |
| 2008 | Emotion-Based Music Visualization Using Photos
Chin-Han Chen, Ming-Fang Weng, Shyh-Kang Jeng, Yung-Yu Chuang |
MMM | 4 |
| 2008 | Semantic Analysis for Automatic Event Recognition and Segmentation of Wedding Ceremony VideosabstractWedding is one of the most important ceremonies in our lives. It symbolizes the birth and creation of a new family. In this paper, we present a system for automatically segmenting a wedding ceremony video into a sequence of recognizable wedding events, e.g., the couple's wedding kiss. Our goal is to develop an automatic tool that helps users to efficiently organize, search, and retrieve his/her treasured wedding memories. Furthermore, the obtained event descriptions could benefit and complement the current research in semantic video understanding. Based on the knowledge of wedding customs, a set of audiovisual features, relating to the wedding contexts of speech/music types, applause activities, picture-taking activities, and leading roles, are exploited to build statistical models for each wedding event. Thirteen wedding events are then recognized by a hidden Markov model, which takes into account both the fitness of observed features and the temporal rationality of event ordering to improve the segmentation accuracy. We conducted experiments on a collection of wedding videos and the promising results demonstrate the effectiveness of our approach. Comparisons with conditional random fields show that the proposed approach is more effective in this application domain. Wen-Huang Cheng, Yung-Yu Chuang, Yin-Tzu Lin, Chi-Chang Hsieh, Shao-Yen Fang, Bing-Yu Chen 0004, Ja-Ling Wu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Association and Temporal Rule Mining for Post-Filtering of Semantic Concept Detection in VideoabstractAutomatic semantic concept detection in video is important for effective content-based video retrieval and mining and has gained great attention recently. In this paper, we propose a general post-filtering framework to enhance robustness and accuracy of semantic concept detection using association and temporal analysis for concept knowledge discovery. Co-occurrence of several semantic concepts could imply the presence of other concepts. We use association mining techniques to discover such inter-concept association relationships from annotations. With discovered concept association rules, we propose a strategy to combine associated concept classifiers to improve detection accuracy. In addition, because video is often visually smooth and semantically coherent, detection results from temporally adjacent shots could be used for the detection of the current shot. We propose temporal filter designs for inter-shot temporal dependency mining to further improve detection accuracy. Experiments on the TRECVID 2005 dataset show our post-filtering framework is both efficient and effective in improving the accuracy of semantic concept detection in video. Furthermore, it is easy to integrate our framework with existing classifiers to boost their performance. Ken-Hao Liu, Ming-Fang Weng, Chi-Yao Tseng, Yung-Yu Chuang, Ming-Syan Chen |
IEEE Trans. Multim. | 4 |
| 2006 | Real-time triple product relighting using spherical local-frame parameterization
Wan-Chun Ma, Chun-Tse Hsiao, Ken-Yi Lee, Yung-Yu Chuang, Bing-Yu Chen 0004 |
Vis. Comput. | 4 |
| 2005 | Level-of-detail representation of bidirectional texture functions for real-time renderingabstractThis paper presents a new technique for rendering bidirectional texture functions (BTFs) at different levels of detail (LODs). Our method first decomposes each BTF image into multiple subbands with a Laplacian pyramid. Each vector of Laplacian coefficients of a texel at the same level is regarded as a Laplacian bidirectional reflectance distribution function (BRDF). These vectors are then further compressed by applying principal components analysis (PCA). At the rendering stage, the LOD parameter for each pixel is calculated according to the distance from the viewpoint to the surface. Our rendering algorithm uses this parameter to determine how many levels of BTF Laplacian pyramid are required for rendering. Under the same sampling resolution, a BTF gradually transits to a BRDF as the camera moves away from the surface. Our method precomputes this transition and uses it for multiresolution BTF rendering. Our Laplacian pyramid representation allows real-time anti-aliased rendering of BTFs using graphics hardware. In addition to provide visually satisfactory multiresolution rendering for BTFs, our method has a comparable compression rate to the available single-resolution BTF compression techniques. Wan-Chun Ma, Sung-Hsiang Chao, Yu-Ting Tseng, Yung-Yu Chuang, Chun-Fa Chang, Bing-Yu Chen 0004, Ouhyoung Ming |
SI3D | 4 |
| 2005 | Cubical Marching Squares: Adaptive Feature Preserving Surface Extraction from Volume DataabstractIn this paper, we present a new method for surface extraction from volume data which preserves sharp features, maintains consistent topology and generates surface adaptively without crack patching. Our approach is based on the marching cubes algorithm, a popular method to convert volumetric data to polygonal meshes. The original marching cubes algorithm suffers from problems of topological inconsistency, cracks in adaptive resolution and inability to preserve sharp features. Most of marching cubes variants only focus on one or some of these problems. Although these techniques could be combined to solve these problems altogether, such a combination might not be straightforward. Moreover, some feature-preserving variants introduce an additional problem, inter-cell dependency. Our method provides a relatively simple and easy-to-implement solution to all these problems by converting 3D marching cubes into 2D cubical marching squares, resolving topology ambiguity with sharp features and eliminating inter-cell dependency by sampling face sharp features. We compare our algorithm with other marching cubes variants and demonstrate its effectiveness on various applications. Murphy Chien-Chang Ho, Fu-Che Wu, Bing-Yu Chen 0004, Yung-Yu Chuang, Ouhyoung Ming |
Comput. Graph. Forum | 4 |
| 2005 | Animating pictures with stochastic motion texturesabstractIn this paper, we explore the problem of enhancing still pictures with subtly animated motions. We limit our domain to scenes containing passive elements that respond to natural forces in some fashion. We use a semi-automatic approach, in which a human user segments the scene into a series of layers to be individually animated. Then, a "stochastic motion texture" is automatically synthesized using a spectral method, i.e., the inverse Fourier transform of a filtered noise spectrum. The motion texture is a time-varying 2D displacement map, which is applied to each layer. The resulting warped layers are then recomposited to form the animated frames. The result is a looping video texture created from a single still image, which has the advantages of being more controllable and of generally higher image quality and resolution than a video texture created from a video source. We demonstrate the technique on a variety of photographs and paintings. Yung-Yu Chuang, Dan B. Goldman, Ke Colin Zheng, Brian Curless, David Salesin, Richard Szeliski |
ACM Trans. Graph. | 1 |
| 2003 | Shadow matting and compositingabstractIn this paper, we describe a method for extracting shadows from one natural scene and inserting them into another. We develop physically-based shadow matting and compositing equations and use these to pull a shadow matte from a source scene in which the shadow is cast onto an arbitrary planar background. We then acquire the photometric and geometric properties of the target scene by sweeping oriented linear shadows (cast by a straight object) across it. From these shadow scans, we can construct a shadow displacement map without requiring camera or light source calibration. This map can then be used to deform the original shadow matte. We demonstrate our approach for both indoor scenes with controlled lighting and for outdoor scenes using natural lighting. Yung-Yu Chuang, Dan B. Goldman, Brian Curless, David Salesin, Richard Szeliski |
ACM Trans. Graph. | 1 |
| 2002 | Video matting of complex scenesabstractThis paper describes a new framework for video matting, the process of pulling a high-quality alpha matte and foreground from a video sequence. The framework builds upon techniques in natural image matting, optical flow computation, and background estimation. User interaction is comprised of garbage matte specification if background estimation is needed, and hand-drawn keyframe segmentations into "foreground," "background" and "unknown". The segmentations, called trimaps, are interpolated across the video volume using forward and backward optical flow. Competing flow estimates are combined based on information about where flow is likely to be accurate. A Bayesian matting technique uses the flowed trimaps to yield high-quality mattes of moving foreground elements with complex boundaries filmed by a moving camera. A novel technique for smoke matte extraction is also demonstrated. Yung-Yu Chuang, Aseem Agarwala, Brian Curless, David Salesin, Richard Szeliski |
ACM Trans. Graph. | 1 |
| 2001 | A Bayesian Approach to Digital MattingabstractThis paper proposes a new Bayesian framework for solving the matting problem, i.e. extracting a foreground element from a background image by estimating an opacity for each pixel of the foreground element. Our approach models both the foreground and background color distributions with spatially-varying sets of Gaussians, and assumes a fractional blending of the foreground and background colors to produce the final output. It then uses a maximum-likelihood criterion to estimate the optimal opacity, foreground and background simultaneously. In addition to providing a principled approach to the matting problem, our algorithm effectively handles objects with intricate boundaries, such as hair strands and fur, and provides an improvement over existing techniques for these difficult cases. Yung-Yu Chuang, Brian Curless, David Salesin, Richard Szeliski |
CVPR (2) | 1 |
| 2000 | Environment matting extensions: towards higher accuracy and real-time captureabstractEnvironment matting is a generalization of traditional bluescreen matting. By photographing an object in front of a sequence of structured light backdrops, a set of approximate light-transport paths through the object can be computed. The original environment matting research chose a middle ground—using a moderate number of photographs to produce results that were reasonably accurate for many objects. In this work, we extend the technique in two opposite directions: recovering a more accurate model at the expense of using additional structured light backdrops, and obtaining a simplified matte using just a single backdrop. The first extension allows for the capture of complex and subtle interactions of light with objects, while the second allows for video capture of colorless objects in motion. Yung-Yu Chuang, Douglas E. Zongker, Joel Hindorff, Brian Curless, David Salesin, Richard Szeliski |
SIGGRAPH | 1 |
| 1996 | Reusable Radiosity ObjectsabstractAbstract Because of the view independence and photo realistic image generation in a diffuse environment, radiosity is suitable for an interactive walk through system. The drawback of radiosity is that it is time‐consuming in form factor estimation, and furthermore, inserting, deleting or moving an object makes the whole costly rendering process repeat itself. To solve this problem, we encapsulate necessary information for form factor calculation and visibility estimation in each object, which is called a reusable radiosity object. An object is defined as a cluster or clusters of triangles. Whenever a scene updates, the radiosity algorithm looks up the prestored information in each object, thus speeding itself up by two orders of magnitude. Besides, solution time based on cluster representatives is linear to the number of objects since each object is reusable, encapsulated with preprocessed data in every level of hierarchy. We also analyze the unregarded error on visibility estimation and propose a statistically optimal adaptive algorithm to maintain the same error for each link. Ouhyoung Ming, Yung-Yu Chuang, Rung-Huei Liang |
Comput. Graph. Forum | 2 |