Hwann-Tzong Chen

dblp:74/2341 · DBLP profile ↗
← Back
64ranked-venue papers
9as first author
18since 2021 · last 2025
0000-0003-2806-7090ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 56 · 8 first-author · 16 since 2021Artificial intelligence and machine learning · 45 · 7 first-author · 16 since 2021
YearPublicationVenuePosition
2025 EigenGS Representation: From Eigenspace to Gaussian Image Space
abstract
Principal Component Analysis (PCA), a classical dimensionality reduction technique, and 2D Gaussian representation, an adaptation of 3D Gaussian Splatting for image representation, offer distinct approaches to modeling visual data. We present EigenGS, a novel method that bridges these paradigms through an efficient transformation pipeline connecting eigenspace and image-space Gaussian representations. Our approach enables instant initialization of Gaussian parameters for new images without requiring per-image optimization from scratch, dramatically accelerating convergence. EigenGS introduces a frequency-aware learning mechanism that encourages Gaussians to adapt to different scales, effectively modeling varied spatial frequencies and preventing artifacts in high-resolution reconstruction. Extensive experiments demonstrate that EigenGS not only achieves superior reconstruction quality compared to direct 2D Gaussian fitting but also reduces the necessary parameter count and training time. The results highlight EigenGS’s effectiveness and generalization ability across images with varying resolutions and diverse categories, making Gaussian-based image representation both high-quality and viable for real-time applications.
Lo-Wei Tai, Ching-En Li, Cheng-Lin Chen, Chih-Jung Tsai, Hwann-Tzong Chen, Tyng-Luh Liu
CVPR5
2025 Segment Anything, Even Occluded
abstract
Amodal instance segmentation, which aims to detect and segment both visible and invisible parts of objects in images, plays a crucial role in various applications, including autonomous driving, robotic manipulation, and scene understanding. While existing methods require training both front-end detectors and mask decoders jointly, this approach lacks flexibility and fails to leverage the strengths of pre-existing modal detectors. To address this limitation, we propose SAMEO, a novel framework that adapts the Segment Anything Model (SAM) as a versatile mask decoder capable of interfacing with various front-end detectors to enable mask prediction even for partially occluded objects. Acknowledging the constraints of limited amodal segmentation datasets, we introduce Amodal-LVIS, a large-scale synthetic dataset comprising 300K images derived from the modal LVIS and LVVIS datasets. This dataset significantly expands the training data available for amodal segmentation research. Our experimental results demonstrate that our approach, when trained on the newly extended dataset, including Amodal-LVIS, achieves remarkable zero-shot performance on both COCOA-cls and D2SA benchmarks, highlighting its potential for generalization to unseen scenarios.
Wei-En Tai, Yu-Lin Shih, Cheng Sun 0004, Yu-Chiang Frank Wang, Hwann-Tzong Chen
CVPR5
2025 ErpGS: Equirectangular Image Rendering Enhanced with 3D Gaussian Regularization
abstract
The use of multi-view images acquired by a 360-degree camera can reconstruct a 3D space with a wide area. There are 3D reconstruction methods from equirectangular images based on NeRF and 3DGS, as well as Novel View Synthesis (NVS) methods. On the other hand, it is necessary to overcome the large distortion caused by the projection model of a 360-degree camera when equirectangular images are used. In 3DGS-based methods, the large distortion of the 360-degree camera model generates extremely large 3D Gaussians, resulting in poor rendering accuracy. We propose ErpGS, which is Omnidirectional GS based on 3DGS to realize NVS addressing the problems. ErpGS introduce some rendering accuracy improvement techniques: geometric regularization, scale regularization, and distortion-aware weights and a mask to suppress the effects of obstacles in equirectangular images. Through experiments on public datasets, we demonstrate that ErpGS can render novel view images more accurately than conventional methods.
Shintaro Ito, Natsuki Takama, Koichi Ito 0001, Hwann-Tzong Chen, Takafumi Aoki
ICIP4
2025 Sparse2DGS: Sparse-View Surface Reconstruction Using 2D Gaussian Splatting with Dense Point Cloud
abstract
Gaussian Splatting (GS) has gained attention as a fast and effective method for novel view synthesis. It has also been applied to 3D reconstruction using multi-view images and can achieve fast and accurate 3D reconstruction. However, GS assumes that the input contains a large number of multi-view images, and therefore, the reconstruction accuracy significantly decreases when only a limited number of input images are available. One of the main reasons is the insufficient number of 3D points in the sparse point cloud obtained through Structure from Motion (SfM), which results in a poor initialization for optimizing the Gaussian primitives. We propose a new 3D reconstruction method, called Sparse2DGS, to enhance 2DGS in reconstructing objects using only three images. Sparse2DGS employs DUSt3R, a fundamental model for stereo images, along with COLMAP MVS to generate highly accurate and dense 3D point clouds, which are then used to initialize 2D Gaussians. Through experiments on the DTU dataset, we show that Sparse2DGS can accurately reconstruct the 3D shapes of objects using just three images.
Natsuki Takama, Shintaro Ito, Koichi Ito 0001, Hwann-Tzong Chen, Takafumi Aoki
ICIP4
2024 Seg2Reg: Differentiable 2D Segmentation to 1D Regression Rendering for 360 Room Layout Reconstruction
abstract
State-of-the-art single-view 360° room layout reconstruction methods formulate the problem as a high-level 1D (per-column) regression task. On the other hand, traditional low-level 2D layout segmentation is simpler to learn and can represent occluded regions, but it requires complex post-processing for the targeting layout polygon and sacrifices accuracy. We present Seg2Reg to render 1D layout depth regression from the 2D segmentation map in a differentiable and occlusion-aware way, marrying the merits of both sides. Specifically, our model predicts floor-plan density for the input equirectangular 360° image. Formulating the 2D layout representation as a density field enables us to employ ‘flattened’ volume rendering to form 1D layout depth regression. In addition, we propose a novel 3D warping augmentation on layout to improve generalization. Finally, we re-implement recent room layout reconstruction methods into our codebase for benchmarking and explore modern backbones and training techniques to serve as the strong baseline. The code is at https://PanoLayoutStudio.github.io.
Cheng Sun 0004, Wei-En Tai, Yu-Lin Shih, Kuan-Wei Chen, Yong-Jing Syu, Kent Selwyn The, Yu-Chiang Frank Wang, Hwann-Tzong Chen
CVPR8
2024 Learning Diffusion Models for Multi-view Anomaly Detection
Chieh Liu, Yu-Min Chu, Ting-I Hsieh, Hwann-Tzong Chen, Tyng-Luh Liu
ECCV (33)4
2024 Pseudo-embedding for Generalized Few-Shot 3D Segmentation
Chih-Jung Tsai, Hwann-Tzong Chen, Tyng-Luh Liu
ECCV (36)2
2023 Hashing Neural Video Decomposition with Multiplicative Residuals in Space-Time
abstract
We present a video decomposition method that facilitates layer-based editing of videos with spatiotemporally varying lighting and motion effects. Our neural model decomposes an input video into multiple layered representations, each comprising a 2D texture map, a mask for the original video, and a multiplicative residual characterizing the spatiotemporal variations in lighting conditions. A single edit on the texture maps can be propagated to the corresponding locations in the entire video frames while preserving other contents’ consistencies. Our method efficiently learns the layer-based neural representations of a 1080p video in 25s per frame via coordinate hashing and allows real-time rendering of the edited result at 71 fps on a single GPU. Qualitatively, we run our method on various videos to show its effectiveness in generating high-quality editing effects. Quantitatively, we propose to adopt feature-tracking evaluation metrics for objectively assessing the consistency of video editing. Project page: https://lightbulb12294.github.io/hashing-nvd/
Cheng-Hung Chan, Cheng-Yang Yuan, Cheng Sun 0004, Hwann-Tzong Chen
ICCV4
2023 Shape-Guided Dual-Memory Learning for 3D Anomaly Detection
abstract
We present a shape-guided expert-learning framework to tackle the problem of unsupervised 3D anomaly detection. Our method is established on the effectiveness of two specialized expert models and their synergy to localize anomalous regions from color and shape modalities. The first expert utilizes geometric information to probe 3D structural anomalies by modeling the implicit distance fields around local shapes. The second expert considers the 2D RGB features associated with the first expert to identify color appearance irregularities on the local shapes. We use the two experts to build the dual memory banks from the anomaly-free training samples and perform shape-guided inference to pinpoint the defects in the testing samples. Owing to the per-point 3D representation and the effective fusion scheme of complementary modalities, our method efficiently achieves state-of-the-art performance on the MVTec 3D-AD dataset with better recall and lower false positive rates, as preferred in real applications.
Yu-Min Chu, Chieh Liu, Ting-I Hsieh, Hwann-Tzong Chen, Tyng-Luh Liu
ICML4
2022 Pose Adaptive Dual Mixup for Few-Shot Single-View 3D Reconstruction
abstract
We present a pose adaptive few-shot learning procedure and a two-stage data interpolation regularization, termed Pose Adaptive Dual Mixup (PADMix), for single-image 3D reconstruction. While augmentations via interpolating feature-label pairs are effective in classification tasks, they fall short in shape predictions potentially due to inconsistencies between interpolated products of two images and volumes when rendering viewpoints are unknown. PADMix targets this issue with two sets of mixup procedures performed sequentially. We first perform an input mixup which, combined with a pose adaptive learning procedure, is helpful in learning 2D feature extraction and pose adaptive latent encoding. The stagewise training allows us to build upon the pose invariant representations to perform a follow-up latent mixup under one-to-one correspondences between features and ground-truth volumes. PADMix significantly outperforms previous literature on few-shot settings over the ShapeNet dataset and sets new benchmarks on the more challenging real-world Pix3D dataset.
Ta Ying Cheng, Hsuan-Ru Yang, Agathoniki Trigoni, Hwann-Tzong Chen, Tyng-Luh Liu
AAAI4
2022 Direct Voxel Grid Optimization: Super-fast Convergence for Radiance Fields Reconstruction
abstract
We present a super-fast convergence approach to reconstructing the per-scene radiance field from a set of images that capture the scene with known poses. This task, which is often applied to novel view synthesis, is recently revolution-ized by Neural Radiance Field (NeRF) for its state-of-the-art quality and fiexibility. However, NeRF and its variants require a lengthy training time ranging from hours to days for a single scene. In contrast, our approach achieves NeRF-comparable quality and converges rapidly from scratch in less than 15 minutes with a single GPU. We adopt a representation consisting of a density voxel grid for scene geometry and a feature voxel grid with a shallow network for complex view-dependent appearance. Modeling with explicit and discretized volume representations is not new, but we propose two simple yet non-trivial techniques that contribute to fast convergence speed and high-quality output. First, we introduce the post-activation interpolation on voxel density, which is capable of producing sharp surfaces in lower grid resolution. Second, direct voxel density optimization is prone to suboptimal geometry solutions, so we robustify the optimization process by imposing several priors. Finally, evaluation on five inward-facing benchmarks shows that our method matches, if not surpasses, NeRF's quality, yet it only takes about 15 minutes to train from scratch for a new scene. Code: https://github.com/sunset1995/DirectVoxGO.
Cheng Sun 0004, Min Sun 0001, Hwann-Tzong Chen
CVPR3
2022 Multiview Regenerative Morphing with Dual Flows
Chih-Jung Tsai, Cheng Sun 0004, Hwann-Tzong Chen
ECCV (16)3
2021 DropLoss for Long-Tail Instance Segmentation
abstract
Long-tailed class distributions are prevalent among the practical applications of object detection and instance segmentation. Prior work in long-tail instance segmentation addresses the imbalance of losses between rare and frequent categories by reducing the penalty for a model incorrectly predicting a rare class label. We demonstrate that the rare categories are heavily suppressed by correct background predictions, which reduces the probability for all foreground categories with equal weight. Due to the relative infrequency of rare categories, this leads to an imbalance that biases towards predicting more frequent categories. Based on this insight, we develop DropLoss -- a novel adaptive loss to compensate for this imbalance without a trade-off between rare and frequent categories. With this loss, we show state-of-the-art mAP across rare, common, and frequent categories on the LVIS dataset. Codes are available at https://github.com/timy90022/DropLoss.
Ting-I Hsieh, Esther Robb, Hwann-Tzong Chen, Jia-Bin Huang 0001
AAAI3
2021 Text-Guided Graph Neural Networks for Referring 3D Instance Segmentation
abstract
This paper addresses a new task called referring 3D instance segmentation, which aims to segment out the target instance in a 3D scene given a query sentence. Previous work on scene understanding has explored visual grounding with natural language guidance, yet the emphasis is mostly constrained on images and videos. We propose a Text-guided Graph Neural Network (TGNN) for referring 3D instance segmentation on point clouds. Given a query sentence and the point cloud of a 3D scene, our method learns to extract per-point features and predicts an offset to shift each point toward its object center. Based on the point features and the offsets, we cluster the points to produce fused features and coordinates for the candidate objects. The resulting clusters are modeled as nodes in a Graph Neural Network to learn the representations that encompass the relation structure for each candidate object. The GNN layers leverage each object's features and its relations with neighbors to generate an attention heatmap for the input sentence expression. Finally, the attention heatmap is used to "guide" the aggregation of information from neighborhood nodes. Our method achieves state-of-the-art performance on referring 3D instance segmentation and 3D localization on ScanRefer, Nr3D, and Sr3D benchmarks, respectively.
Pin-Hao Huang, Han-Hung Lee, Hwann-Tzong Chen, Tyng-Luh Liu
AAAI3
2021 Indoor Panorama Planar 3D Reconstruction via Divide and Conquer
abstract
Indoor panorama typically consists of human-made structures parallel or perpendicular to gravity. We leverage this phenomenon to approximate the scene in a 360-degree image with (H)orizontal-planes and (V)ertical-planes. To this end, we propose an effective divide-and-conquer strategy that divides pixels based on their plane orientation estimation; then, the succeeding instance segmentation module conquers the task of planes clustering more easily in each plane orientation group. Besides, parameters of V-planes depend on camera yaw rotation, but translation-invariant CNNs are less aware of the yaw change. We thus propose a yaw-invariant V-planar reparameterization for CNNs to learn. We create a benchmark for indoor panorama planar reconstruction by extending existing 360 depth datasets with ground truth H&V-planes (referred to as "PanoH&V" dataset) and adopt state-of-the-art planar reconstruction methods to predict H&V-planes as our baselines. Our method outperforms the baselines by a large margin on the proposed dataset.
Cheng Sun 0004, Chi-Wei Hsiao, Ning-Hsu Wang, Min Sun 0001, Hwann-Tzong Chen
CVPR5
2021 HoHoNet: 360 Indoor Holistic Understanding With Latent Horizontal Features
abstract
We present HoHoNet, a versatile and efficient framework for holistic understanding of an indoor 360-degree panorama using a Latent Horizontal Feature (LHFeat). The compact LHFeat flattens the features along the vertical direction and has shown success in modeling per-column modality for room layout reconstruction. HoHoNet advances in two important aspects. First, the deep architecture is redesigned to run faster with improved accuracy. Second, we propose a novel horizon-to-dense module, which relaxes the per-column output shape constraint, allowing per-pixel dense prediction from LHFeat. HoHoNet is fast: It runs at 52 FPS and 110 FPS with ResNet-50 and ResNet-34 backbones respectively, for modeling dense modalities from a high-resolution 512 × 1024 panorama. HoHoNet is also accurate. On the tasks of layout estimation and semantic segmentation, HoHoNet achieves results on par with current state-of-the-art. On dense depth estimation, HoHoNet outperforms all the prior arts by a large margin. Code is available at https://github.com/sunset1995/HoHoNet.
Cheng Sun 0004, Min Sun 0001, Hwann-Tzong Chen
CVPR3
2021 Specialize and Fuse: Pyramidal Output Representation for Semantic Segmentation
abstract
We present a novel pyramidal output representation to ensure parsimony with our "specialize and fuse" process for semantic segmentation. A pyramidal "output" representation consists of coarse-to-fine levels, where each level is "specialize" in a different class distribution (e.g., more stuff than things classes at coarser levels). Two types of pyramidal outputs (i.e., unity and semantic pyramid) are "fused" into the final semantic output, where the unity pyramid indicates unity-cells (i.e., all pixels in such cell share the same semantic label). The process ensures parsimony by predicting a relatively small number of labels for unity-cells (e.g., a large cell of grass) to build the final semantic output. In addition to the "output" representation, we design a coarse-to-fine contextual module to aggregate the "features" representation from different levels. We validate the effectiveness of each key module in our method through comprehensive ablation studies. Finally, our approach achieves state-of-the-art performance on three widely-used semantic segmentation datasets—ADE20K, COCO-Stuff, and Pascal-Context.
Chi-Wei Hsiao, Cheng Sun 0004, Hwann-Tzong Chen, Min Sun 0001
ICCV3
2021 Bridging the Gap Between Computational Photography and Visual Recognition
abstract
What is the current state-of-the-art for image restoration and enhancement applied to degraded images acquired under less than ideal circumstances? Can the application of such algorithms as a pre-processing step improve image interpretability for manual analysis or automatic visual recognition to classify scene content? While there have been important advances in the area of computational photography to restore or enhance the visual quality of an image, the capabilities of such techniques have not always translated in a useful way to visual recognition tasks. Consequently, there is a pressing need for the development of algorithms that are designed for the joint problem of improving visual appearance and recognition, which will be an enabling factor for the deployment of visual recognition tools in many real-world scenarios. To address this, we introduce the UG$^2$dataset as a large-scale benchmark composed of video imagery captured under challenging conditions, and two enhancement tasks designed to test algorithmic impact on visual quality and automatic object recognition. Furthermore, we propose a set of metrics to evaluate the joint improvement of such tasks as well as individual algorithmic advances, including a novel psychophysics-based evaluation regime for human assessment and a realistic set of quantitative measures for object recognition performance. We introduce six new algorithms for image restoration or enhancement, which were created as part of the IARPA sponsored UG$^2$Challenge workshop held at CVPR 2018. Under the proposed evaluation regime, we present an in-depth analysis of these algorithms and a host of deep learning-based and classic baseline approaches. From the observed results, it is evident that we are in the early days of building a bridge between computational photography and visual recognition, leaving many opportunities for innovation in this area.
Rosaura G. VidalMata, Sreya Banerjee, Brandon RichardWebster, Michael Albright, Pedro Davalos, Scott McCloskey, Ben Miller, Asong Tambo, Sushobhan Ghosh, Sudarshan Nagesh, Ye Yuan 0012, Yueyu Hu, Wenhan Yang, Xiaoshuai Zhang, Jiaying Liu 0001, Zhangyang Wang, Hwann-Tzong Chen, Tzu-Wei Huang, Wen-Chi Chin, Yi-Chun Li, Mahmoud Lababidi, Charles Otto, Walter J. Scheirer
IEEE Trans. Pattern Anal. Mach. Intell.18
2020 Learning Camera-Aware Noise Models
Ke-Chi Chang, Ren Wang 0014, Hung-Jin Lin, Yu-Lun Liu 0001, Chia-Ping Chen, Yu-Lin Chang, Hwann-Tzong Chen
ECCV (24)7
2020 SwipeCut: Interactive Segmentation via Seed Grouping
abstract
Interactive image segmentation algorithms rely on the user to provide annotations as the guidance. When the task of interactive segmentation is performed on a small touchscreen device, the requirement of providing precise annotations could be cumbersome to the user. We design a new interaction mechanism that actively queries seeds for guiding the user to label. Our method enforces sparsity and diversity criteria on the selection of query seeds, and at each round of interaction, the user is only presented with a small number of informative query seeds that are far apart from each other. Therefore, the user merely has to swipe through on the ROI-relevant query seeds for checking which ones of the query seeds are inside the region of interest (ROI). This kind of interaction should be easy since those gestures are commonly used on a touchscreen. As a result, we are able to derive a user-friendly interaction mechanism for annotation on small touchscreen devices. The performance of our algorithm is evaluated on six publicly available datasets. The evaluation results show that our algorithm achieves high segmentation accuracy, with short computational time and less user feedback.
Ding-Jie Chen, Hwann-Tzong Chen, Long-Wen Chang
IEEE Trans. Circuits Syst. Video Technol.2
2019 Unsupervised Meta-Learning of Figure-Ground Segmentation via Imitating Visual Effects
Ding-Jie Chen, Jui-Ting Chien, Hwann-Tzong Chen, Tyng-Luh Liu
AAAI3
2019 Instance-Level Meta Normalization
abstract
This paper presents a normalization mechanism called Instance-Level Meta Normalization (ILM Norm) to address a learning-to-normalize problem. ILM~Norm learns to predict the normalization parameters via both the feature feed-forward and the gradient back-propagation paths. ILM Norm provides a meta normalization mechanism and has several good properties. It can be easily plugged into existing instance-level normalization schemes such as Instance Normalization, Layer Normalization, or Group Normalization. ILM~Norm normalizes each instance individually and therefore maintains high performance even when small mini-batch is used. The experimental results show that ILM~Norm well adapts to different network architectures and tasks, and it consistently improves the performance of the original models.
Songhao Jia, Ding-Jie Chen, Hwann-Tzong Chen
CVPR3
2019 HorizonNet: Learning Room Layout With 1D Representation and Pano Stretch Data Augmentation
abstract
We present a new approach to the problem of estimating the 3D room layout from a single panoramic image. We represent room layout as three 1D vectors that encode, at each image column, the boundary positions of floor-wall and ceiling-wall, and the existence of wall-wall boundary. The proposed network, HorizonNet, trained for predicting 1D layout, outperforms previous state-of-the-art approaches. The designed post-processing procedure for recovering 3D room layouts from 1D predictions can automatically infer the room shape with low computation cost - it takes less than 20ms for a panorama image while prior works might need dozens of seconds. We also propose Pano Stretch Data Augmentation, which can diversify panorama data and be applied to other panorama-related learning tasks. Due to the limited data available for non-cuboid layout, we relabel 65 general layout from the current dataset for finetuning. Our approach shows good performance on general layouts by qualitative results and cross-validation.
Cheng Sun 0004, Chi-Wei Hsiao, Min Sun 0001, Hwann-Tzong Chen
CVPR4
2019 See-Through-Text Grouping for Referring Image Segmentation
abstract
Motivated by the conventional grouping techniques to image segmentation, we develop their DNN counterpart to tackle the referring variant. The proposed method is driven by a convolutional-recurrent neural network (ConvRNN) that iteratively carries out top-down processing of bottom-up segmentation cues. Given a natural language referring expression, our method learns to predict its relevance to each pixel and derives a See-through-Text Embedding Pixelwise (STEP) heatmap, which reveals segmentation cues of pixel level via the learned visual-textual co-embedding. The ConvRNN performs a top-down approximation by converting the STEP heatmap into a refined one, whereas the improvement is expected from training the network with a classification loss from the ground truth. With the refined heatmap, we update the textual representation of the referring expression by re-evaluating its attention distribution and then compute a new STEP heatmap as the next input to the ConvRNN. Boosting by such collaborative learning, the framework can progressively and simultaneously yield the desired referring segmentation and reasonable attention distribution over the referring sentence. Our method is general and does not rely on, say, the outcomes of object detection from other DNN models, while achieving state-of-the-art performance in all of the four datasets in the experiments.
Ding-Jie Chen, Songhao Jia, Yi-Chen Lo, Hwann-Tzong Chen, Tyng-Luh Liu
ICCV4
2019 COCO-GAN: Generation by Parts via Conditional Coordinating
abstract
Humans can only interact with part of the surrounding environment due to biological restrictions. Therefore, we learn to reason the spatial relationships across a series of observations to piece together the surrounding environment. Inspired by such behavior and the fact that machines also have computational constraints, we propose COnditional COordinate GAN (COCO-GAN) of which the generator generates images by parts based on their spatial coordinates as the condition. On the other hand, the discriminator learns to justify realism across multiple assembled patches by global coherence, local appearance, and edge-crossing continuity. Despite the full images are never manipulated during training, we show that COCO-GAN can produce state-of-the-art-quality full images during inference. We further demonstrate a variety of novel applications enabled by our coordinate-aware framework. First, we perform extrapolation to the learned coordinate manifold and generate off-the-boundary patches. Combining with the originally generated full image, COCO-GAN can produce images that are larger than training samples, which we called "beyond-boundary generation". We then showcase panorama generation within a cylindrical coordinate system that inherently preserves horizontally cyclic topology. On the computation side, COCO-GAN has a built-in divide-and-conquer paradigm that reduces memory requisition during training and inference, provides high-parallelism, and can generate parts of images on-demand.
Chieh Hubert Lin, Chia-Che Chang, Yu-Sheng Chen, Da-Cheng Juan, Wei Wei 0019, Hwann-Tzong Chen
ICCV6
2019 Point-to-Point Video Generation
abstract
While image synthesis achieves tremendous breakthroughs (e.g., generating realistic faces), video generation is less explored and harder to control, which limits its applications in the real world. For instance, video editing requires temporal coherence across multiple clips and thus poses both start and end constraints within a video sequence. We introduce point-to-point video generation that controls the generation process with two control points: the targeted start- and end-frames. The task is challenging since the model not only generates a smooth transition of frames but also plans ahead to ensure that the generated end-frame conforms to the targeted end-frame for videos of various lengths. We propose to maximize the modified variational lower bound of conditional data likelihood under a skip-frame training strategy. Our model can generate end-frame-consistent sequences without loss of quality and diversity. We evaluate our method through extensive experiments on Stochastic Moving MNIST, Weizmann Action, Human3.6M, and BAIR Robot Pushing under a series of scenarios. The qualitative results showcase the effectiveness and merits of point-to-point generation.
Tsun-Hsuan Wang, Yen-Chi Cheng, Chieh Hubert Lin, Hwann-Tzong Chen, Min Sun 0001
ICCV4
2019 Learning Dense Correspondences for Video Objects
abstract
We introduce a learning based method for extracting distinctive features on video objects. Based on the extracted features, we are able to derive dense correspondences between the object in the current video frame and the reference template, and then use the correspondences to identify the grasping points on the object. We train a deep-learning model to predict dense feature maps using the training data collected via solving simultaneous localization and mapping (SLAM). Further, a new feature-aggregation technique based on the optical flow of consecutive frames is applied to the integration of multiple feature maps for alleviating uncertainties. We also use the optical flow information to assess the reliability of feature matching. The experimental results show that our approach effectively reduces unreliable correspondences and thus improves the matching accuracy.
Wen-Chi Chin, Zih-Jian Jhang, Hwann-Tzong Chen, Koichi Ito 0001
ICIP3
2019 One-Shot Object Detection with Co-Attention and Co-Excitation
abstract
This paper aims to tackle the challenging problem of one-shot object detection. Given a query image patch whose class label is not included in the training data, the goal of the task is to detect all instances of the same class in a target image. To this end, we develop a novel {\em co-attention and co-excitation} (CoAE) framework that makes contributions in three key technical aspects. First, we propose to use the non-local operation to explore the co-attention embodied in each query-target pair and yield region proposals accounting for the one-shot situation. Second, we formulate a squeeze-and-co-excitation scheme that can adaptively emphasize correlated feature channels to help uncover relevant proposals and eventually the target objects. Third, we design a margin-based ranking loss for implicitly learning a metric to predict the similarity of a region proposal to the underlying query, no matter its class label is seen or unseen in training. The resulting model is therefore a two-stage detector that yields a strong baseline on both VOC and MS-COCO under one-shot setting of detecting objects from both seen and never-seen classes.
Ting-I Hsieh, Yi-Chen Lo, Hwann-Tzong Chen, Tyng-Luh Liu
NeurIPS3
2019 Attentive and Adversarial Learning for Video Summarization
abstract
This paper aims to address the video summarization problem via attention-aware and adversarial training. We formulate the problem as a sequence-to-sequence task, where the input sequence is an original video and the output sequence is its summarization. We propose a GAN-based training framework, which combines the merits of unsupervised and supervised video summarization approaches. The generator is an attention-aware Ptr-Net that generates the cutting points of summarization fragments. The discriminator is a 3D CNN classifier to judge whether a fragment is from a ground-truth or a generated summarization. The experiments show that our method achieves state-of-the-art results on SumMe, TVSum, YouTube, and LoL datasets with 1.5% to 5.6% improvements. Our Ptr-Net generator can overcome the unbalanced training-test length in the seq2seq problem, and our discriminator is effective in leveraging unpaired summarizations to achieve better performance.
Tsu-Jui Fu, Shao-Heng Tai, Hwann-Tzong Chen
WACV3
2018 Tap and Shoot Segmentation
abstract
We present a new segmentation method that leverages latent photographic information available at the moment of taking pictures. Photography on a portable device is often done by tapping to focus before shooting the picture. This tap-and-shoot interaction for photography not only specifies the region of interest but also yields useful focus/defocus cues for image segmentation. However, most of the previous interactive segmentation methods address the problem of image segmentation in a post-processing scenario without considering the action of taking pictures. We propose a learning-based approach to this new tap-and-shoot scenario of interactive segmentation. The experimental results on various datasets show that, by training a deep convolutional network to integrate the selection and focus/defocus cues, our method can achieve higher segmentation accuracy in comparison with existing interactive segmentation methods.
Ding-Jie Chen, Jui-Ting Chien, Hwann-Tzong Chen, Long-Wen Chang
AAAI3
2018 A2A: Attention to Attention Reasoning for Movie Question Answering
Chao-Ning Liu, Ding-Jie Chen, Hwann-Tzong Chen, Tyng-Luh Liu
ACCV (6)3
2018 Escaping from Collapsing Modes in a Constrained Space
Chia-Che Chang, Chieh Hubert Lin, Che-Rung Lee, Da-Cheng Juan, Wei Wei 0019, Hwann-Tzong Chen
ECCV (7)6
2018 Toward a unified scheme for fast interactive segmentation
Ding-Jie Chen, Hwann-Tzong Chen, Long-Wen Chang
J. Vis. Commun. Image Represent.2
2018 Interactive 1-bit feedback segmentation using transductive inference
Ding-Jie Chen, Hwann-Tzong Chen, Long-Wen Chang
Mach. Vis. Appl.2
2017 Video segmentation via boundary-aware flow
abstract
We present a new algorithm for unsupervised video segmentation based on boundary-aware optical flow. Existing video segmentation methods usually tweak their segmentation model to tolerate the inaccuracy in the estimation of optical flow around object boundaries. In contrast, we directly manipulate the optical flow for better quality. We smooth the optical flow via transductive inference to make the flow consistent within the object and fit to the object boundaries. We then use the boundary-aware optical flow to estimate the initial foreground object region from each frame for learning the appearance model. The learned appearance model is consequently used to refine the segmentation result. Experiments on the DAVIS dataset show that our method performs favorably against the existing ones.
Ding-Jie Chen, Hwann-Tzong Chen, Long-Wen Chang
ICIP2
2017 Chess recognition from a single depth image
abstract
This paper presents a learning-based method for recognizing chess pieces from depth information. The proposed method is integrated in a recreational robotic system that is designed to play games of chess against humans. The robot has two arms and an Ensenso N35 Stereo 3D camera. Our goal is to provide the robot visual intelligence so that it can identify the chess pieces on the chessboard using the depth information captured by the 3D camera. We build a convolutional neural network to solve this 3D object recognition problem. While training neural networks for 3D object recognition becomes popular these days, collecting enough training data is still a time-consuming task. We demonstrate that it is much more convenient and effective to generate the required training data from 3D CAD models. The neural network trained using the rendered data performs well on real inputs during testing. More specifically, the experimental results show that using the training data rendered from the CAD models under various conditions enhances the recognition accuracy significantly. When further evaluations are done on real data captured by the 3D camera, our method achieves 90.3% accuracy.
Yu-An Wei, Tzu-Wei Huang, Hwann-Tzong Chen, JenChi Liu
ICME3
2017 Quantitative Analysis of Automatic Image Cropping Algorithms: A Dataset and Comparative Study
abstract
Automatic photo cropping is an important tool for improving visual quality of digital photos without resorting to tedious manual selection. Traditionally, photo cropping is accomplished by determining the best proposal window through visual quality assessment or saliency detection. In essence, the performance of an image cropper highly depends on the ability to correctly rank a number of visually similar proposal windows. Despite the ranking nature of automatic photo cropping, little attention has been paid to learning-to-rank algorithms in tackling such a problem. In this work, we conduct an extensive study on traditional approaches as well as ranking-based croppers trained on various image features. In addition, a new dataset consisting of high quality cropping and pairwise ranking annotations is presented to evaluate the performance of various baselines. The experimental results on the new dataset provide useful insights into the design of better photo cropping algorithms.
Yi-Ling Chen 0004, Tzu-Wei Huang, Kai-Han Chang, Yu-Chen Tsai, Hwann-Tzong Chen, Bing-Yu Chen 0004
WACV5
2016 Interactive Segmentation from 1-Bit Feedback
Ding-Jie Chen, Hwann-Tzong Chen, Long-Wen Chang
ACCV (1)2
2016 Fast defocus map estimation
abstract
This paper presents a fast algorithm for deriving the defocus map from a single image. Existing methods of defocus map estimation often include a pixel-level propagation step to spread the measured sparse defocus cues over the whole image. Since the pixel-level propagation step is time-consuming, we develop an effective method to obtain the whole-image defocus blur using oversegmentation and transductive inference. Oversegmentation produces the superpixels and hence greatly reduces the computation costs for subsequent procedures. Transductive inference provides a way to calculate the similarity between superpixels, and thus helps to infer the defocus blur of each superpixel from all other superpixels. The experimental results show that our method is efficient and able to estimate a plausible superpixel-level defocus map from a given single image.
Ding-Jie Chen, Hwann-Tzong Chen, Long-Wen Chang
ICIP2
2015 Transparent texture transfer
abstract
This paper presents a toolchain that can transfer transparent, multi-layer textures onto some unknown surface in an image. We propose alpha quilting, which is able to generate a large texture map and its corresponding alpha matte from a small template of natural texture. The synthesized texture map can then be transferred to a surface according to the estimated surface normal. The procedure of texture transfer is called texture matting, which layers multiple textures onto a surface with respect to the transparency and multi-layer characteristics of natural textures. The experimental results show that our method is able to produce visually plausible results of pasting opaque and transparent textures on surfaces.
Chan-Tai Yeh, Ting-Hui Tsai, Hwann-Tzong Chen
ICIP3
2014 Example-based motion manipulation
abstract
This paper introduces an idea of imitating camera movements and shooting styles from an example video and reproducing the effects on another carelessly shot video. Our goal is to compute a series of transformations so that we can warp the input video and make the flow of the output video resemble the example flow. We formulate and solve an optimization problem to find the required transformation for each video frame. By enforcing the resemblance between flow fields, our method can recreate different shooting styles and camera movements, from simple effect of video stabilization, to more complicated ones like anti-blur panning, smooth zoom, tracking shot, and dolly zoom.
Pin-Ching Su, Hwann-Tzong Chen, Chia-Ming Cheng
ICIP2
2013 Unsupervised scene segmentation using sparse coding context
Yen-Cheng Liu, Hwann-Tzong Chen
Mach. Vis. Appl.2
2012 Learning Feature Subspaces for Appearance-Based Bundle Adjustment
Chia-Ming Cheng, Hwann-Tzong Chen
ACCV (4)2
2012 Video object cosegmentation
abstract
We introduce and address the problem of video object cosegmentation, which concerns the task of segmenting the common object in a pair of video sequences. We present a new algorithm that works on super-voxels in videos to solve this task. The algorithm computes i the intra-video relative motion derived from dense optical flow and ii) the inter-video co-features based on Gaussian mixture models. The experimental results show that, by integrating the intra-video and inter-video information, our algorithm is able to obtain better results of segmenting video objects.
Ding-Jie Chen, Hwann-Tzong Chen, Long-Wen Chang
ACM Multimedia2
2011 Fusing generic objectness and visual saliency for salient object detection
abstract
We present a novel computational model to explore the relatedness of objectness and saliency, each of which plays an important role in the study of visual attention. The proposed framework conceptually integrates these two concepts via constructing a graphical model to account for their relationships, and concurrently improves their estimation by iteratively optimizing a novel energy function realizing the model. Specifically, the energy function comprises the objectness, the saliency, and the interaction energy, respectively corresponding to explain their individual regularities and the mutual effects. Minimizing the energy by fixing one or the other would elegantly transform the model into solving the problem of objectness or saliency estimation, while the useful information from the other concept can be utilized through the interaction term. Experimental results on two benchmark datasets demonstrate that the proposed model can simultaneously yield a saliency map of better quality and a more meaningful objectness output for salient object detection.
Kai-Yueh Chang, Tyng-Luh Liu, Hwann-Tzong Chen, Shang-Hong Lai
ICCV3
2010 A square-root sampling approach to fast histogram-based search
abstract
We present an efficient pixel-sampling technique for histogram-based search. Given a template image as a query, a typical histogram-based algorithm aims to find the location of the target in another large test image, by evaluating a similarity measure for comparing the feature histogram of the template with that of each possible subwindow in the test image. The computational cost would be high if each subwindow needs to compute its histogram and evaluate the similarity measure. In this paper, we adopt the probability-product kernels as the similarity measures, and show that the computation of histograms and the evaluation of the kernel-based similarities can be integrated through a sampling approach. Specifically, we present a square-root sampling method to avoid the computation of histograms, and meanwhile, reduce the number of pixels required for evaluating the similarity measure. The proposed approximation algorithm is time- and memory-efficient. The time complexity of computing the similarity for each subwindow is O(1). The memory requirement of our algorithm is not in proportion to the size of the template or the number of histogram bins, which allows the use of more distinctive image representations for better search accuracy.
Huang-Wei Chang, Hwann-Tzong Chen
CVPR2
2010 Preattentive co-saliency detection
abstract
This paper presents a new algorithm to solve the problem of co-saliency detection, i.e., to find the common salient objects that are present in both of a pair of input images. Unlike most previous approaches, which require correspondence matching, we seek to solve the problem of co-saliency detection under a preattentive scheme. Our algorithm does not need to perform the correspondence matching between the two input images, and is able to achieve co-saliency detection before the focused attention occurs. The joint information provided by the image pair enables our algorithm to inhibit the responses of other salient objects that appear in just one of the images. Through experiments we show that our algorithm is effective in localizing the co-salient objects inside input image pairs.
Hwann-Tzong Chen
ICIP1
2010 Probing the local-feature space of interest points
abstract
We describe an empirical study on the feature space of interest points for natural images. Although local features have been widely used in image analysis as building blocks of various algorithms, there is still a lack of study on the space of local features, in particular the distributions of local features extracted from the locations of interest points. Based on an approximate neighborhood probing scheme, we devise an efficient representation that is not only helpful in visualizing the distributions of local features, but also effective in modeling the neighborhood structures of local features for fast matching. We present various experimental results that provide useful insights into feature description and fast image matching.
Wei-Ting Lee, Hwann-Tzong Chen
ICIP2
2010 Dimensionality Reduction via Tangential Learning
Jianzhuang Liu, Hwann-Tzong Chen
ICIP4
2009 Histogram-based interest point detectors
abstract
We present a new method for detecting interest points using histogram information. Unlike existing interest point detectors, which measure pixel-wise differences in image intensity, our detectors incorporate histogram-based representations, and thus can find image regions that present a distinct distribution in the neighborhood. The proposed detectors are able to capture large-scale structures and distinctive textured patterns, and exhibit strong invariance to rotation, illumination variation, and blur. The experimental results show that the proposed histogram-based interest point detectors perform particularly well for the tasks of matching textured scenes under blur and illumination changes, in terms of repeatability and distinctiveness. An extension of our method to space-time interest point detection for action classification is also presented.
Wei-Ting Lee, Hwann-Tzong Chen
CVPR2
2009 Finding good composition in panoramic scenes
abstract
We introduce a new problem of automatic photo composition, and present an effective technique for finding good views within a panoramic scene. Instead of applying heuristic rules of photo composition, we propose to imitate good composition presented in the artworks of professional photographers. Our approach tries to model the composition styles of professional photographs by analyzing the structural features and the layout of visual saliency. The task of finding good photo composition through a viewfinder is formulated as a search problem, and we present a stochastic search algorithm to look for good viewing configurations and to choose suitable reference images from the collection of masterpiece photographs. Given any initial location in the panoramic scene, our algorithm is able to suggest a better view that would often yield professional-like photo composition.
Yuan-Yang Chang, Hwann-Tzong Chen
ICCV2
2009 Landmark-based sparse color representations for color transfer
abstract
We present a novel image representation that characterizes a color image by an intensity image and a small number of color pixels. Our idea is based on solving an inverse problem of colorization: Given a color image, we seek to obtain an intensity image and a small subset of color pixels, which are called landmark pixels, so that the input color image can be recovered faithfully using the intensity image and the color cues provided by the selected landmark pixels. We develop an algorithm to derive the landmark-based sparse color representations from color images, and use the representations in the applications of color transfer and color correction. The computational cost for these applications is low owing to the sparsity of the proposed representation. The landmark-based representation is also preferable to statistics-based representations, e.g. color histograms and Gaussian mixture models, when we need to reconstruct the color image from a given representation.
Tzu-Wei Huang, Hwann-Tzong Chen
ICCV2
2007 Finding Familiar Objects and their Depth from a Single Image
abstract
We present a classification-based method to identify objects of interest, and judge their depth in a single image. Our approach is motivated by a postulate of human depth perception that people can give a credible depth estimation for an object whose familiar size is known, even without using stereo vision. To emulate the mechanism, we categorize objects into the same class if they have similar sizes and shapes, and model the sense of discovering a familiar object by applying multiple kernel logistic regression to the conditional probability of feature types. The depth of a detected target can then be obtained by referencing its corresponding object category. Overall, the proposed algorithm is efficient in both the training and testing phases, and does not require a large amount of training images for good performances.
Hwann-Tzong Chen, Tyng-Luh Liu
ICIP (6)1
2006 Segmenting Highly Articulated Video Objects with Weak-Prior Random Forests
Hwann-Tzong Chen, Tyng-Luh Liu, Chiou-Shann Fuh
ECCV (4)1
2005 Local Discriminant Embedding and Its Variants
abstract
We present a new approach, called local discriminant embedding (LDE), to manifold learning and pattern classification. In our framework, the neighbor and class relations of data are used to construct the embedding for classification problems. The proposed algorithm learns the embedding for the submanifold of each class by solving an optimization problem. After being embedded into a low-dimensional subspace, data points of the same class maintain their intrinsic neighbor relations, whereas neighboring points of different classes no longer stick to one another. Via embedding, new test data are thus more reliably classified by the nearest neighbor rule, owing to the locally discriminating nature. We also describe two useful variants: two-dimensional LDE and kernel LDE. Comprehensive comparisons and extensive experiments on face recognition are included to demonstrate the effectiveness of our method.
Hwann-Tzong Chen, Huang-Wei Chang, Tyng-Luh Liu
CVPR (2)1
2005 Tone Reproduction: A Perspective from Luminance-Driven Perceptual Grouping
abstract
We address the tone reproduction problem by integrating local adaptation with global-contrast consistency. Many previous works have tried to compress high-dynamic-range (HDR) luminances into a displayable range in imitation of the local adaptation mechanism of human eyes. Nevertheless, while the realization of local adaptation is not theoretically defined, exaggerating such effects often causes unnatural global contrasts. We propose a luminance-driven perceptual grouping process to derive a sparse representation of HDR luminances, and use the grouped regions to approximate local properties of luminances. The advantage of incorporating a sparse representation is twofold: We can simulate local adaptation based on region information, and subsequently apply piecewise tone mappings to monotonize the relative brightness over only a few perceptually significant regions. Our experimental results show that the proposed framework gives a good balance in preserving local details and maintaining global contrasts of HDR scenes.
Hwann-Tzong Chen, Tyng-Luh Liu, Tien-Lung Chang
CVPR (2)1
2005 Learning Effective Image Metrics from Few Pairwise Examples
abstract
We present a new approach to learning image metrics. The main advantage of our method lies in a formulation that requires only a few pairwise examples. Apparently, based on the little amount of side-information, it would take a very effective learning scheme to yield a useful image metric. Our algorithm achieves this goal by addressing two key issues. First, we establish a global-local (glocal) image representation that induces two structure-meaningful vector spaces to respectively describe the global and the local image properties. Second, we develop a metric optimization framework that finds an optimal bilinear transform to best explain the given side-information. We emphasize it is the glocal image representation that makes the use of bilinear transform more powerful. Experimental results on classifications of face images and visual tracking are included to demonstrate the contributions of the proposed method.
Hwann-Tzong Chen, Tyng-Luh Liu, Chiou-Shann Fuh
ICCV1
2005 Semantic manifold learning for image retrieval
abstract
Learning the user's semantics for CBIR involves two different sources of information: the similarity relations entailed by the content-based features, and the relevance relations specified in the feedback. Given that, we propose an augmented relation embedding (ARE) to map the image space into a semantic manifold that faithfully grasps the user's preferences. Besides ARE, we also look into the issues of selecting a good feature set for improving the retrieval performance. With these two aspects of efforts we have established a system that yields far better results than those previously reported. Overall, our approach can be characterized by three key properties: 1) The framework uses one relational graph to describe the similarity relations, and the other two to encode the relevant/irrelevant relations indicated in the feedback. 2) With the relational graphs so defined, learning a semantic manifold can be transformed into solving a constrained optimization problem, and is reduced to the ARE algorithm accounting for both the representation and the classification points of views. 3) An image representation based on augmented features is introduced to couple with the ARE learning. The use of these features is significant in capturing the semantics concerning different scales of image regions. We conclude with experimental results and comparisons to demonstrate the effectiveness of our method.
Yen-Yu Lin, Tyng-Luh Liu, Hwann-Tzong Chen
ACM Multimedia3
2005 Tone Reproduction: A Perspective from Luminance-Driven Perceptual Grouping
Hwann-Tzong Chen, Tyng-Luh Liu, Chiou-Shann Fuh
Int. J. Comput. Vis.1
2004 Real-Time Tracking Using Trust-Region Methods
abstract
Optimization methods based on iterative schemes can be divided into two classes: line-search methods and trust-region methods. While line-search techniques are commonly found in various vision applications, not much attention is paid to trust-region ones. Motivated by the fact that line-search methods can be considered as special cases of trust-region methods, we propose to establish a trust-region framework for real-time tracking. Our approach is characterized by three key contributions. First, since a trust-region tracking system is more effective, it often yields better performances than the outcomes of other trackers that rely on iterative optimization to perform tracking, e.g., a line-search-based mean-shift tracker. Second, we have formulated a representation model that uses two coupled weighting schemes derived from the covariance ellipse to integrate an object's color probability distribution and edge density information. As a result, the system can address rotation and nonuniform scaling in a continuous space, rather than working on some presumably possible discrete values of rotation angle and scale. Third, the framework is very flexible in that a variety of distance functions can be adapted easily. Experimental results and comparative studies are provided to demonstrate the efficiency of the proposed method.
Tyng-Luh Liu, Hwann-Tzong Chen
IEEE Trans. Pattern Anal. Mach. Intell.2
2001 Multi-Object Tracking Using Dynamical Graph Matching
abstract
We describe a tracking algorithm to address the interactions among objects, and to track them individually and confidently via a static camera. It is achieved by constructing an invariant bipartite graph to model the dynamics of the tracking process, of which the nodes are classified into objects and profiles. The best match of the graph corresponds to an optimal assignment for resolving the identities of the detected objects. Since objects may enter/exit the scene indefinitely, or when interactions occur/conclude they could form/leave a group, the number of nodes in the graph changes dynamically. Therefore it is critical to maintain an invariant property to assure that the numbers of nodes of both types are kept the same so that the matching problem is manageable. In addition, several important issues are also discussed, including reducing the effect of shadows, extracting objects' shapes, and adapting large abrupt changes in the scene background. Finally, experimental results are provided to illustrate the efficiency of our approach.
Hwann-Tzong Chen, Horng-Horng Lin, Tyng-Luh Liu
CVPR (2)1
2001 Trust-Region Methods for Real-Time Tracking
abstract
Optimization methods based on iterative schemes can be divided into two classes: linesearch methods and trust-region methods. While linesearch techniques are commonly found in various vision applications, not much attention is paid to trust-region methods. Motivated by the fact that linesearch methods can be considered as special cases of trust-region methods, we propose to apply trust-region methods to visual tracking problems. Our approach integrates trust-region methods with the Kullback Leibler distance to track a rigid or non-rigid object in real-time. If not limited by the speed of a camera, the algorithm can achieve frame rate above 60 fps. To justify our method, a variety of experiments/comparisons are carried out for the trust-region tracker and a linesearch-based mean-shift tracker with same initial conditions. The experimental results support our conjecture that a trust-region tracker should perform superiorly to a linesearch one.
Hwann-Tzong Chen, Tyng-Luh Liu
ICCV1
2000 A Variational Approach for Digital Watermarking
abstract
We describe a digital watermarking framework by minimizing an appropriate variational problem with an energy/cost function accounting for properties such as robustness, fidelity, and so forth. The process of embedding a watermark into a multimedia data can then be treated as solving an optimization problem, where we show that it can be transformed into finding the minimum cost (integral) flow of a watermarking network. The framework is very general and flexible since the cost function can be easily modified to accommodate other criteria. Unlike other watermarking methods, our approach is implemented via solving a network flow problem, and consequently this protects the watermarking scheme from deliberate attacks designed specifically using the knowledge of the algorithm.
Tyng-Luh Liu, Hwann-Tzong Chen
ICIP2
1998 Agents that have Desires and Adaptive Behaviors
Cheng-Yuan Liou, Chung-Hao Tan, Hwann-Tzong Chen, Jiun-Hung Chen
ICONIP3