EDBT 2026 Demo / reviewers in the wild / expert
Tyng-Luh Liu
dblp:68/2368
· DBLP profile ↗
87ranked-venue papers
8as first author
30since 2021 · last 2025
0000-0002-8366-5213ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 74 · 7 first-author · 24 since 2021Artificial intelligence and machine learning · 60 · 7 first-author · 21 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multimodal Promptable Token Merging for Diffusion ModelsabstractToken compression techniques, such as token merging and pruning, are essential for alleviating the substantial computational burden caused by the proliferation of tokens within attention mechanisms. However, current methods often rely on token-to-token distances or similarity metrics to evaluate token importance, which is inadequate in the context of modern promptable designs and frameworks that are gaining prominence. To address this limitation, we introduce a novel and effective merging strategy called “Multimodal Promptable Token Merging” (MPTM). The proposed method leverages a multimodal, prompt-centric methodology, assessing the proximity between tokens of each input modality and the multimodal prompt to efficiently eliminate redundant tokens while preserving those rich in information. Extensive experiments demonstrate that MPTM significantly reduces computational costs without compromising essential information in generative image tasks. When integrated into diffusion-based detection architectures, MPTM outperforms existing state-of-the-art methods by 2.3% in object detection tasks. Additionally, when applied to multimodal diffusion models, MPTM maintains high-quality output while achieving a 2.9-fold increase in throughput, highlighting its versatility. Cheng-Yao Hong, Tyng-Luh Liu |
AAAI | 2 |
| 2025 | SLIP: Spoof-Aware One-Class Face Anti-Spoofing with Language Image PretrainingabstractFace anti-spoofing (FAS) plays a pivotal role in ensuring the security and reliability of face recognition systems. With advancements in vision-language pretrained (VLP) models, recent two-class FAS techniques have leveraged the advantages of using VLP guidance, while this potential remains unexplored in one-class FAS methods. The one-class FAS focuses on learning intrinsic liveness features solely from live training images to differentiate between live and spoof faces. However, the lack of spoof training data can lead one-class FAS models to inadvertently incorporate domain information irrelevant to the live/spoof distinction (\eg, facial content), causing performance degradation when tested with a new application domain. To address this issue, we propose a novel framework called Spoof-aware one-class face anti-spoofing with Language Image Pretraining (SLIP). Given that live faces should ideally not be obscured by any spoof-attack-related objects (\eg, paper, or masks) and are assumed to yield zero spoof cue maps, we first propose an effective language-guided spoof cue map estimation to enhance one-class FAS models by simulating whether the underlying faces are covered by attack-related objects and generating corresponding nonzero spoof cue maps. Next, we introduce a novel prompt-driven liveness feature disentanglement to alleviate live/spoof-irrelative domain variations by disentangling live/spoof-relevant and domain-dependent information. Finally, we design an effective augmentation strategy by fusing latent features from live images and spoof prompts to generate spoof-like image features and thus diversify latent spoof features to facilitate the learning of one-class FAS. Our extensive experiments and ablation studies support that SLIP consistently outperforms previous one-class FAS methods. Pei-Kai Huang, Jun-Xiong Chong, Cheng-Hsuan Chiang, Tzu-Hsien Chen, Tyng-Luh Liu, Chiou-Ting Hsu |
AAAI | 5 |
| 2025 | Tracking Everything Everywhere across Multiple CamerasabstractPixel tracking in single-view video sequences has recently emerged as a significant area of research. While previous work has primarily concentrated on tracking within a given video, we propose to expand pixel correspondence estimation into multi-view scenarios. The central concept involves utilizing a canonical space that preserves a universal 3D representation across different views and timesteps. This model allows for precise tracking of points even through prolonged occlusions and significant deformations in appearance between views. Moreover, we show that our model, through the use of an efficient training strategy incorporating distillation loss, is capable of performing incremental pixel tracking, a process often seen as complex in test-time optimization techniques. Comprehensive experiments validate the method's ability to accurately establish point correspondences across cameras. Furthermore, our method achieves promising results of multi-view pixel tracking without requiring the entire video sequences to be provided at once. Li-Heng Wang, YuJu Cheng, Tyng-Luh Liu |
AAAI | 3 |
| 2025 | EigenGS Representation: From Eigenspace to Gaussian Image SpaceabstractPrincipal Component Analysis (PCA), a classical dimensionality reduction technique, and 2D Gaussian representation, an adaptation of 3D Gaussian Splatting for image representation, offer distinct approaches to modeling visual data. We present EigenGS, a novel method that bridges these paradigms through an efficient transformation pipeline connecting eigenspace and image-space Gaussian representations. Our approach enables instant initialization of Gaussian parameters for new images without requiring per-image optimization from scratch, dramatically accelerating convergence. EigenGS introduces a frequency-aware learning mechanism that encourages Gaussians to adapt to different scales, effectively modeling varied spatial frequencies and preventing artifacts in high-resolution reconstruction. Extensive experiments demonstrate that EigenGS not only achieves superior reconstruction quality compared to direct 2D Gaussian fitting but also reduces the necessary parameter count and training time. The results highlight EigenGS’s effectiveness and generalization ability across images with varying resolutions and diverse categories, making Gaussian-based image representation both high-quality and viable for real-time applications. Lo-Wei Tai, Ching-En Li, Cheng-Lin Chen, Chih-Jung Tsai, Hwann-Tzong Chen, Tyng-Luh Liu |
CVPR | 6 |
| 2025 | Promptable 3-D Object Localization with Latent Diffusion ModelsabstractAccurate identification and localization of objects in 3-D scenes are essential for advancing comprehensive 3-D scene understanding. Although diffusion models have demonstrated impressive capabilities across a broad spectrum of computer vision tasks, their potential in both 2-D and 3-D object detection remains underexplored. Existing approaches typically formulate detection as a ''noise-to-box'' process, but they rely heavily on direct coordinate regression, which limits adaptability for more advanced tasks such as grounding-based object detection. To overcome these challenges, we propose a promptable 3-D object recognition framework, which introduces a diffusion-based paradigm for flexible and conditionally guided 3-D object detection. Our approach encodes bounding boxes into latent representations and employs latent diffusion models to realize a ''promptable noise-to-box'' transformation. This formulation enables the refinement of standard 3-D object detection using textual prompts, such as class labels. Moreover, it naturally extends to grounding object detection through conditioning on natural language descriptions, and generalizes effectively to few-shot learning by incorporating annotated exemplars as visual prompts. We conduct thorough evaluations on three key 3-D object recognition tasks: general 3-D object detection, few-shot detection, and grounding-based detection. Experimental results demonstrate that our framework achieves competitive performance relative to state-of-the-art methods, validating its effectiveness, versatility, and broad applicability in 3-D computer vision. Cheng-Yao Hong, Li-Heng Wang, Tyng-Luh Liu |
NeurIPS | 3 |
| 2024 | Contrastive Learning for DeepFake Classification and Localization via Multi-Label RankingabstractWe propose a unified approach to simultaneously addressing the conventional setting of binary deepfake classification and a more challenging scenario of uncovering what facial components have been forged as well as the exact order of the manipulations. To solve the former task, we consider mul-tiple instance learning (MIL) that takes each image as a bag and its patches as instances. A positive bag corresponds to a forged image that includes at least one manipulated patch (i.e., a pixel in the feature map). The formulation allows us to estimate the probability of an input image being a fake one and establish the corresponding contrastive MIL loss. On the other hand, tackling the component-wise deepfake problem can be reduced to solving multi-label prediction, but the requirement to recover the manipulation order further complicates the learning task into a multi-label ranking prob-lem. We resolve this difficulty by designing a tailor-made loss term to enforce that the rank order of the predicted multi-label probabilities respects the ground-truth order of the sequential modifications of a deepfake image. Through extensive experiments and comparisons with other relevant techniques, we provide extensive results and ablation studies to demonstrate that the proposed method is an overall more comprehensive solution to deepfake detection. Cheng-Yao Hong, Yen-Chi Hsu, Tyng-Luh Liu |
CVPR | 3 |
| 2024 | One-Class Face Anti-Spoofing via Spoof Cue Map-Guided Feature LearningabstractMany face anti-spoofing (FAS) methods have focused on learning discriminative features from both live and spoof training data to strengthen the security of face recognition systems. However, since not every possible attack type is available in the training stage, these FAS methods usually fail to detect unseen attacks in the inference stage. In comparison, one-class FAS, where training data comprise only live faces, aims to detect whether a test face image belongs to the live class or not. In this paper, we propose a novel One-Class Spoof Cue Map estimation Network (OC-SCMNet) to address the one-class FAS detection problem. Our first goal is to learn to extract latent spoof features from live images so that their estimated Spoof Cue Maps (SCMs) should have zero responses. To avoid trapping to a trivial solution, we devise a novel SCM-guided feature learning by combining many SCMs as pseudo ground-truths to guide a conditional generator to create latent spoof features for spoof data. Our second goal is to simulate the potential out-of-distribution spoof attacks approximately. To this end, we propose using a memory bank to dynamically preserve a set of sufficiently “independent” latent spoof features to encourage the generator to probe the latent spoof feature space. Extensive experiments conducted on eight FAS benchmark datasets demonstrate that the proposed OC-SCMNet not only outperforms previous one-class FAS approaches but also achieves performance comparable to the state-of-the-art two-class FAS methods. The code is available at https://github.com/Pei-KaiHuang/CVPR24_OC_SCMNet. Pei-Kai Huang, Cheng-Hsuan Chiang, Tzu-Hsien Chen, Jun-Xiong Chong, Tyng-Luh Liu, Chiou-Ting Hsu |
CVPR | 5 |
| 2024 | Learning Diffusion Models for Multi-view Anomaly Detection
Chieh Liu, Yu-Min Chu, Ting-I Hsieh, Hwann-Tzong Chen, Tyng-Luh Liu |
ECCV (33) | 5 |
| 2024 | Pseudo-embedding for Generalized Few-Shot 3D Segmentation
Chih-Jung Tsai, Hwann-Tzong Chen, Tyng-Luh Liu |
ECCV (36) | 3 |
| 2023 | IoU-Aware Multi-Expert Cascade Network Via Dynamic Ensemble for Long-Tailed Object DetectionabstractObject detection over a long-tailed large-scale dataset is practical, challenging, and comprehensively under-explored. Recently proposed methods mainly focus on eliminating the imbalanced classification problem. However, only a few attempts have been made to consider the quality of the predicted bounding boxes. Inspired by the observation of existing Cascade architecture, "detectors with specific IoU thresholds excel at different label frequencies of bounding boxes," this paper first pinpoints the issue in long-tailed distribution. A detector may predict inaccurate bounding boxes on the categories of fewer training data such that the corresponding extracted visual features could further degrade the classification accuracy. Thus, the predicted accuracy of bounding boxes becomes substantially different among categories in the long-tailed distribution. We introduce a Multi-Expert Cascade (MEC) framework that readjusts the weight of each category in the training process via a multi-expert loss. Furthermore, we leverage dynamic ensemble mechanisms at inference time to fully utilize expert detectors and achieve better performance. Extensive experiments on the recent long-tailed large vocabulary object detection dataset show that the proposed MEC framework significantly improves the performance of most widely-used detectors over various backbones on object detection and instance segmentation tasks. Wan-Cyuan Fan, Cheng-Yao Hong, Yen-Chi Hsu, Tyng-Luh Liu |
ICASSP | 4 |
| 2023 | One-Shot Action Detection via Attention Zooming InabstractHinted by a modest support set, few-shot action detection (FSAD) aims at localizing the action instances of unseen classes within an untrimmed query video. Existing FSAD techniques mostly rely on generating a set of class-agnostic action proposals from the query video and then finding the most plausible ones by assessing their correlation to the support set. Such two-stage approaches are feasible but not efficient, largely due to neglecting the support information in generating the proposals. This work focuses on the one-shot image scenario and introduces the attention zooming in strategy to effectively and progressively carry out support-query cross-attention while generating proposals. The resulting one-stage model yields high-quality action proposals for boosting one-shot action detection (OSAD) performance. Our extensive experiments on the ActivityNet-1.3 and THUMOS-14 datasets demonstrate that the proposed framework can achieve state-of-the-art performance in tack-ling challenging image-based OSAD tasks. He-Yen Hsieh, Ding-Jie Chen, Cheng-Wei Chang, Tyng-Luh Liu |
ICASSP | 4 |
| 2023 | Attention Discriminant Sampling for Point CloudsabstractThis paper describes an attention-driven approach to 3-D point cloud sampling. We establish our method based on a structure-aware attention discriminant analysis that explores geometric and semantic relations embodied among points and their clusters. The proposed attention discriminant sampling (ADS) starts by efficiently decomposing a given point cloud into clusters to implicitly encode its structural and geometric relatedness among points. By treating each cluster as a structural component, ADS then draws on evaluating two levels of self-attention: within-cluster and between-cluster. The former reflects the semantic complexity entailed by the learned features of points within each cluster, while the latter reveals the semantic similarity between clusters. Driven by structurally preserving the point distribution, these two aspects of self-attention help avoid sampling redundancy and decide the number of sampled points in each cluster. Extensive experiments demonstrate that ADS significantly improves classification performance to 95.1% on ModelNet40 and 87.5% on ScanObjectNN and achieves 86.9% mIoU on ShapeNet Part Segmentation. For scene segmentation, ADS yields 91.1% accuracy on S3DIS with higher mIoU to the state-of-the-art and 75.6% mIoU on ScanNetV2. Furthermore, ADS surpasses the state-of-the-art with 55.0% mAP50on ScanNetV2 object detection. Cheng-Yao Hong, Yu-Ying Chou, Tyng-Luh Liu |
ICCV | 3 |
| 2023 | Shape-Guided Dual-Memory Learning for 3D Anomaly DetectionabstractWe present a shape-guided expert-learning framework to tackle the problem of unsupervised 3D anomaly detection. Our method is established on the effectiveness of two specialized expert models and their synergy to localize anomalous regions from color and shape modalities. The first expert utilizes geometric information to probe 3D structural anomalies by modeling the implicit distance fields around local shapes. The second expert considers the 2D RGB features associated with the first expert to identify color appearance irregularities on the local shapes. We use the two experts to build the dual memory banks from the anomaly-free training samples and perform shape-guided inference to pinpoint the defects in the testing samples. Owing to the per-point 3D representation and the effective fusion scheme of complementary modalities, our method efficiently achieves state-of-the-art performance on the MVTec 3D-AD dataset with better recall and lower false positive rates, as preferred in real applications. Yu-Min Chu, Chieh Liu, Ting-I Hsieh, Hwann-Tzong Chen, Tyng-Luh Liu |
ICML | 5 |
| 2023 | Aggregating Bilateral Attention for Few-Shot Instance LocalizationabstractAttention filtering under various learning scenarios has proven advantageous in enhancing the performance of many neural network architectures. The mainstream attention mechanism is established upon the non-local block, also known as an essential component of the prominent Transformer networks, to catch long-range correlations. However, such unilateral attention is often hampered by sparse and obscure responses, revealing insufficient dependencies across images/patches, and high computational cost, especially for those employing the multi-head design. To overcome these issues, we introduce a novel mechanism of aggregating bilateral attention (ABA) and validate its usefulness in tackling the task of few-shot instance localization, reflecting the underlying query-support dependency. Specifically, our method facilitates uncovering informative features via assessing: i) an embedding norm for exploring the semantically-related cues; ii) context awareness for correlating the query data and support regions. ABA is then carried out by integrating the affinity relations derived from the two measurements to serve as a lightweight but effective query-support attention mechanism with high localization recall. We evaluate ABA on two localization tasks, namely, few-shot action localization and one-shot object detection. Extensive experiments demonstrate that the proposed ABA achieves superior performances over existing methods. He-Yen Hsieh, Ding-Jie Chen, Cheng-Wei Chang, Tyng-Luh Liu |
WACV | 4 |
| 2023 | ABC-Norm Regularization for Fine-Grained and Long-Tailed Image ClassificationabstractImage classification for real-world applications often involves complicated data distributions such as fine-grained and long-tailed. To address the two challenging issues simultaneously, we propose a new regularization technique that yields an adversarial loss to strengthen the model learning. Specifically, for each training batch, we construct an adaptive batch prediction (ABP) matrix and establish its corresponding adaptive batch confusion norm (ABC-Norm). The ABP matrix is a composition of two parts, including an adaptive component to class-wise encode the imbalanced data distribution, and the other component to batch-wise assess the softmax predictions. The ABC-Norm leads to a norm-based regularization loss, which can be theoretically shown to be an upper bound for an objective function closely related to rank minimization. By coupling with the conventional cross-entropy loss, the ABC-Norm regularization could introduce adaptive classification confusion and thus trigger adversarial learning to improve the effectiveness of model learning. Different from most of state-of-the-art techniques in solving either fine-grained or long-tailed problems, our method is characterized with its simple and efficient design, and most distinctively, provides a unified solution. In the experiments, we compare ABC-Norm with relevant techniques and demonstrate its efficacy on several benchmark datasets, including (CUB-LT, iNaturalist2018); (CUB, CAR, AIR); and (ImageNet-LT), which respectively correspond to the real-world, fine-grained, and long-tailed scenarios. Yen-Chi Hsu, Cheng-Yao Hong, Ming-Sui Lee, Davi Geiger, Tyng-Luh Liu |
IEEE Trans. Image Process. | 5 |
| 2022 | Pose Adaptive Dual Mixup for Few-Shot Single-View 3D ReconstructionabstractWe present a pose adaptive few-shot learning procedure and a two-stage data interpolation regularization, termed Pose Adaptive Dual Mixup (PADMix), for single-image 3D reconstruction. While augmentations via interpolating feature-label pairs are effective in classification tasks, they fall short in shape predictions potentially due to inconsistencies between interpolated products of two images and volumes when rendering viewpoints are unknown. PADMix targets this issue with two sets of mixup procedures performed sequentially. We first perform an input mixup which, combined with a pose adaptive learning procedure, is helpful in learning 2D feature extraction and pose adaptive latent encoding. The stagewise training allows us to build upon the pose invariant representations to perform a follow-up latent mixup under one-to-one correspondences between features and ground-truth volumes. PADMix significantly outperforms previous literature on few-shot settings over the ShapeNet dataset and sets new benchmarks on the more challenging real-world Pix3D dataset. Ta Ying Cheng, Hsuan-Ru Yang, Agathoniki Trigoni, Hwann-Tzong Chen, Tyng-Luh Liu |
AAAI | 5 |
| 2022 | Capturing Humans in Motion: Temporal-Attentive 3D Human Pose and Shape Estimation from Monocular VideoabstractLearning to capture human motion is essential to 3D human pose and shape estimation from monocular video. However, the existing methods mainly rely on recurrent or convolutional operation to model such temporal information, which limits the ability to capture non-local context relations of human motion. To address this problem, we propose a motion pose and shape network (MPS-Net) to effectively capture humans in motion to estimate accurate and temporally coherent 3D human pose and shape from a video. Specifically, we first propose a motion continuity attention (MoCA) module that leverages visual cues observed from human motion to adaptively recalibrate the range that needs attention in the sequence to better capture the motion continuity dependencies. Then, we develop a hierarchical attentive feature integration (HAFI) module to effectively combine adjacent past and future feature represen-tations to strengthen temporal correlation and refine the feature representation of the current frame. By coupling the MoCA and HAFI modules, the proposed MPS-Net excels in estimating 3D human pose and shape in the video. Though conceptually simple, our MPS-Net not only outperforms the state-of-the-art methods on the 3DPW, MPI-INF-3DHP, and Human3.6M benchmark datasets, but also uses fewer network parameters. The video demos can be found at https://mps-net.github.io/MPS-Net/. Wen-Li Wei, Jen-Chun Lin, Tyng-Luh Liu, Hong-Yuan Mark Liao |
CVPR | 3 |
| 2022 | Self-supervised Sparse Representation for Video Anomaly Detection
Jhih-Ciang Wu, He-Yen Hsieh, Ding-Jie Chen, Chiou-Shann Fuh, Tyng-Luh Liu |
ECCV (13) | 5 |
| 2022 | Decoupled Contrastive Learning
Chun-Hsiao Yeh, Cheng-Yao Hong, Yen-Chi Hsu, Tyng-Luh Liu, Yubei Chen, Yann LeCun |
ECCV (26) | 4 |
| 2022 | SAGA: Self-Augmentation with Guided Attention for Representation LearningabstractSelf-supervised training that elegantly couples contrastive learning with a wide spectrum of data augmentation techniques has been shown to be a successful paradigm for representation learning. However, current methods implicitly maximize the agreement between differently augmented views of the same sample, which may perform poorly in certain situations. For example, considering an image comprising a boat on the sea, one augmented view is cropped solely from the boat and the other from the sea, whereas linking these two to form a positive pair could be misleading. To resolve this issue, we introduce a Self-Augmentation with Guided Attention (SAGA) strategy, which augments input data based on predictive attention to learn representations rather than simply applying off-the-shelf augmentation schemes. As a result, the proposed self-augmentation framework enables feature learning to enhance the robustness of representation. Chun-Hsiao Yeh, Cheng-Yao Hong, Yen-Chi Hsu, Tyng-Luh Liu |
ICASSP | 4 |
| 2022 | Contextual Proposal Network for Action LocalizationabstractThis paper investigates the problem of Temporal Action Proposal (TAP) generation, which aims to provide a set of high-quality video segments that potentially contain actions events locating in long untrimmed videos. Based on the goal to distill available contextual information, we introduce a Contextual Proposal Network (CPN) composing of two context-aware mechanisms. The first mechanism, i.e., feature enhancing, integrates the inception-like module with long-range attention to capture the multi-scale temporal contexts for yielding a robust video segment representation. The second mechanism, i.e., boundary scoring, employs the bi-directional recurrent neural networks (RNN) to capture bi-directional temporal contexts that explicitly model actionness, background, and confidence of proposals. While generating and scoring proposals, such bi-directional temporal contexts are helpful to retrieve high-quality proposals of low false positives for covering the video action instances. We conduct experiments on two challenging datasets of ActivityNet-1.3 and THUMOS-14 to demonstrate the effectiveness of the proposed Contextual Proposal Network (CPN). In particular, our method respectively surpasses state-of-the-art TAP methods by 1.54% AUC on ActivityNet-1.3 test split and by 0.61% AR@200 on THUMOS-14 dataset. He-Yen Hsieh, Ding-Jie Chen, Tyng-Luh Liu |
WACV | 3 |
| 2021 | Text-Guided Graph Neural Networks for Referring 3D Instance SegmentationabstractThis paper addresses a new task called referring 3D instance segmentation, which aims to segment out the target instance in a 3D scene given a query sentence. Previous work on scene understanding has explored visual grounding with natural language guidance, yet the emphasis is mostly constrained on images and videos. We propose a Text-guided Graph Neural Network (TGNN) for referring 3D instance segmentation on point clouds. Given a query sentence and the point cloud of a 3D scene, our method learns to extract per-point features and predicts an offset to shift each point toward its object center. Based on the point features and the offsets, we cluster the points to produce fused features and coordinates for the candidate objects. The resulting clusters are modeled as nodes in a Graph Neural Network to learn the representations that encompass the relation structure for each candidate object. The GNN layers leverage each object's features and its relations with neighbors to generate an attention heatmap for the input sentence expression. Finally, the attention heatmap is used to "guide" the aggregation of information from neighborhood nodes. Our method achieves state-of-the-art performance on referring 3D instance segmentation and 3D localization on ScanRefer, Nr3D, and Sr3D benchmarks, respectively. Pin-Hao Huang, Han-Hung Lee, Hwann-Tzong Chen, Tyng-Luh Liu |
AAAI | 4 |
| 2021 | Adaptive Image Transformer for One-Shot Object DetectionabstractOne-shot object detection tackles a challenging task that aims at identifying within a target image all object instances of the same class, implied by a query image patch. The main difficulty lies in the situation that the class label of the query patch and its respective examples are not available in the training data. Our main idea leverages the concept of language translation to boost metric-learning-based detection methods. Specifically, we emulate the language translation process to adaptively translate the feature of each object proposal to better correlate the given query feature for discriminating the class-similarity among the proposal-query pairs. To this end, we propose the Adaptive Image Transformer (AIT) module that deploys an attention-based encoder-decoder architecture to simultaneously explore intra-coder and inter-coder (i.e., each proposal-query pair) attention. The adaptive nature of our design turns out to be flexible and effective in addressing the one-shot learning scenario. With the informative attention cues, the proposed model excels in predicting the class-similarity between the target image proposals and the query image patch. Though conceptually simple, our model significantly outperforms a state-of-the-art technique, improving the unseen-class object classification from 63.8 mAP and 22.0 AP50 to 72.2 mAP and 24.3 AP50 on the PASCAL-VOC and MS-COCO benchmark datasets, respectively. Ding-Jie Chen, He-Yen Hsieh, Tyng-Luh Liu |
CVPR | 3 |
| 2021 | Learning Unsupervised Metaformer for Anomaly DetectionabstractAnomaly detection (AD) aims to address the task of classification or localization of image anomalies. This paper addresses two pivotal issues of reconstruction-based approaches to AD in images, namely, model adaptation and reconstruction gap. The former generalizes an AD model to tackling a broad range of object categories, while the latter provides useful clues for localizing abnormal regions. At the core of our method is an unsupervised universal model, termed as Metaformer, which leverages both meta-learned model parameters to achieve high model adaptation capability and instance-aware attention to emphasize the focal regions for localizing abnormal regions, i.e., to explore the reconstruction gap at those regions of interest. We justify the effectiveness of our method with SOTA results on the MVTec AD dataset of industrial images and highlight the adaptation flexibility of the universal Metaformer with multi-class and few-shot scenarios. Jhih-Ciang Wu, Ding-Jie Chen, Chiou-Shann Fuh, Tyng-Luh Liu |
ICCV | 4 |
| 2021 | The Maximum a Posterior Estimation of DartsabstractThe DARTS approach manifests the advantages of relaxing the discrete problem of network architecture search (NAS) to the continuous domain such that network weights and architecture parameters can be optimized properly. However, it falls short in providing a justifiable and reliable solution for deciding the target architecture. In particular, the design choice of a certain operation at each layer/edge is determined without considering the distribution of operations over the overall architecture or even the neighboring layers. Our method explores such dependencies from the viewpoint of maximum a posterior (MAP) estimation. The consideration takes account of both local and global information by learning transition probabilities of network operations while enabling a greedy scheme to uncover a MAP estimate of optimal target architecture. The experiments show that our method achieves state-of-the-art results on popular benchmark datasets and also can be conveniently plugged into DARTS-related techniques to boost their performance. Our code is available at https://github.com/MAP-DARTS/MAP-DARTS. Jun-Liang Lin, Yi-Lin Sung, Cheng-Yao Hong, Han-Hung Lee, Tyng-Luh Liu |
ICIP | 5 |
| 2021 | Adaptive and Generative Zero-Shot Learning
Yu-Ying Chou, Hsuan-Tien Lin, Tyng-Luh Liu |
ICLR | 3 |
| 2021 | Referring Image Segmentation via Language-Driven AttentionabstractThis paper aims to tackle the problem of referring image segmentation, which is targeted at reasoning the region of interest referred by a query natural language sentence. One key issue to address the referring image segmentation is how to establish the cross-modal representation for encoding the two modalities, namely, the query sentence and the input image. Most existing methods are designed to concatenate the features from each modality or to gradually encode the cross-modal representation concerning each word’s effect. In contrast, our approach leverages the correlation between the two modalities for constructing the cross-modal representation. To make the resulting cross-modal representation more discriminative for the segmentation task, we propose a novel mechanism of language-driven attention to encode the cross-modal representation for reflecting the attention between every single visual element and the entire query sentence. The proposed mechanism, named as Language-Driven Attention (LDA), first decouples the cross-modal correlation to channel-attention and spatial-attention and then integrates the two attentions for obtaining the cross-modal representation. The channel attention and the spatial attention respectively reveal how sensitive each channel or each pixel of a particular feature map is with respect to the query sentence. With a proper fusion of the two kinds of feature attention, the proposed LDA model can effectively guide the generation of the final cross-modal representation. The resulting representation is further strengthened for capturing the multi-receptive-field and multi-level-semantic for the intended segmentation. We assess our referring image segmentation model on four public benchmark datasets, and the experimental results show that our model achieves state-of-the-art performance Ding-Jie Chen, He-Yen Hsieh, Tyng-Luh Liu |
ICRA | 3 |
| 2021 | ezGeno: an automatic model selection package for genomic data analysisabstractMOTIVATION: To facilitate the process of tailor-making a deep neural network for exploring the dynamics of genomic DNA, we have developed a hands-on package called ezGeno. ezGeno automates the search process of various parameters and network structures and can be applied to any kind of 1D genomic data. Combinations of multiple abovementioned 1D features are also applicable. RESULTS: For the task of predicting TF binding using genomic sequences as the input, ezGeno can consistently return the best performing set of parameters and network structure, as well as highlight the important segments within the original sequences. For the task of predicting tissue-specific enhancer activity using both sequence and DNase feature data as the input, ezGeno also regularly outperforms the hand-designed models. Furthermore, we demonstrate that ezGeno is superior in efficiency and accuracy compared to the one-layer DeepBind model and AutoKeras, an open-source AutoML package. AVAILABILITY AND IMPLEMENTATION: The ezGeno package can be freely accessed at https://github.com/ailabstw/ezGeno. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jun-Liang Lin, Tsung-Ting Hsieh, Yi-An Tung, Xuanjun Chen, Yu-Chun Hsiao, Chia-Lin Yang, Tyng-Luh Liu, Chien-Yu Chen 0001 |
Bioinform. | 7 |
| 2021 | One-class anomaly detection via novelty normalization
Jhih-Ciang Wu, Sherman Lu, Chiou-Shann Fuh, Tyng-Luh Liu |
Comput. Vis. Image Underst. | 4 |
| 2021 | Learning to Visualize Music Through Shot Sequence for Automatic Concert Video MashupabstractAn experienced director usually switches among different types of shots to make visual storytelling more touching. When filming a musical performance, appropriate switching shots can produce some special effects, such as enhancing the expression of emotion or heating up the atmosphere. However, while the visual storytelling technique is often used in making professional recordings of a live concert, amateur recordings of audiences often lack such storytelling concepts and skills when filming the same event. Thus a versatile system that can perform video mashup to create a refined high-quality video from such amateur clips is desirable. To this end, we aim at translating the music into an attractive shot (type) sequence by learning the relation between music and visual storytelling of shots. The resulting shot sequence can then be used to better portray the visual storytelling of a song and guide the concert video mashup process. To achieve the task, we first introduces a novel probabilistic-based fusion approach, named as multi-resolution fused recurrent neural networks (MF-RNNs) with film-language, which integrates multi-resolution fused RNNs and a film-language model for boosting the translation performance. We then distill the knowledge in MF-RNNs with film-language into a lightweight RNN, which is more efficient and easier to deploy. The results from objective and subjective experiments demonstrate that both MF-RNNs with film-language and lightweight RNN can generate attractive shot sequences for music, thereby enhancing the viewing and listening experience. Wen-Li Wei, Jen-Chun Lin, Tyng-Luh Liu, Hsiao-Rong Tyan, Hsin-Min Wang, Hong-Yuan Mark Liao |
IEEE Trans. Multim. | 3 |
| 2020 | Query-Driven Multi-Instance LearningabstractWe introduce a query-driven approach (qMIL) to multi-instance learning where the queries aim to uncover the class labels embodied in a given bag of instances. Specifically, it solves a multi-instance multi-label learning (MIML) problem with a more challenging setting than the conventional one. Each MIML bag in our formulation is annotated only with a binary label indicating whether the bag contains the instance of a certain class and the query is specified by the word2vec of a class label/name. To learn a deep-net model for qMIL, we construct a network component that achieves a generalized compatibility measure for query-visual co-embedding and yields proper instance attentions to the given query. The bag representation is then formed as the attention-weighted sum of the instances' weights, and passed to the classification layer at the end of the network. In addition, the qMIL formulation is flexible for extending the network to classify unseen class labels, leading to a new technique to solve the zero-shot MIML task through an iterative querying process. Experimental results on action classification over video clips and three MIML datasets from MNIST, CIFAR10 and Scene are provided to demonstrate the effectiveness of our method. Yen-Chi Hsu, Cheng-Yao Hong, Ming-Sui Lee, Tyng-Luh Liu |
AAAI | 4 |
| 2020 | Self-similarity Student for Partial Label Histopathology Image Segmentation
Hsien-Tzu Cheng, Chun-Fu Yeh, Po-Chen Kuo, Andy Wei, Keng-Chi Liu, Mong-Chi Ko, Kuan-Hua Chao, Yu-Ching Peng, Tyng-Luh Liu |
ECCV (25) | 9 |
| 2020 | Temporal Action Proposal Generation Via Deep Feature EnhancementabstractTemporal action proposal generation (TAPG) is a challenging problem for analyzing video content. It aims to localize the video segments which are likely to contain actions or events. Intuitively, making a satisfying prediction of these video segments is directly relies on their representation quality. A typical representation of a video segment is applying a two-stream feature, which comprises appearance and motion information. Rather than directly concatenating the two-stream features as the previous methods, we illustrate a feature-aggregation network (FA-Net) concerning the feature-relation among neighboring video segments for obtaining the high-quality representation that better characterizing the actions or events. Further, we design a feature-expansion network (FE-Net) to extract multi-granularity features for retrieving the proposals of high action-instance covering confidence. We evaluate our approach on two challenging datasets: ActivityNet-1.3 and THUMOS-14. The experiments showed that the proposed approach consistently outperforms the existing state-of-the-art TAPG methods. He-Yen Hsieh, Ding-Jie Chen, Tyng-Luh Liu |
ICIP | 3 |
| 2020 | Video Summarization with Anchors and Multi-Head AttentionabstractVideo summarization is a challenging task that will automatically generate a representative and attractive highlight movie from the source video. Previous works explicitly exploit the hierarchical structure of video to train a summarizer. However, their method sometimes uses fixed-length segmentation, which breaks the video structure or requires additional training data to train the segmentation model. In this paper, we propose an Anchor-Based Attention RNN (ABA-RNN) for solving the video summarization problem. ABA-RNN provides two contributions. One is that we attain the frame-level and clip-level features by the anchor-based approach, and the model only needs one layer of RNN by introducing subtraction manner used in minus-LSTM. We also use multi-head attention to let the model select suitable lengths of segments. Another contribution is that we do not need any extra video preprocessing to determine shot boundaries and our architecture is end-to-end training. In experiments, we follow the standard datasets SumMe and TVSum and achieve competitive performance against the state-of-the-art results. Yi-Lin Sung, Cheng-Yao Hong, Yen-Chi Hsu, Tyng-Luh Liu |
ICIP | 4 |
| 2020 | Learning From Music to Visual Storytelling of Shots: A Deep Interactive Learning MechanismabstractLearning from music to visual storytelling of shots is an interesting and emerging task. It produces a coherent visual story in the form of a shot type sequence, which not only expands the storytelling potential for a song but also facilitates automatic concert video mashup process and storyboard generation. In this study, we present a deep interactive learning (DIL) mechanism for building a compact yet accurate sequence-to-sequence model to accomplish the task. Different from the one-way transfer between a pre-trained teacher network (or ensemble network) and a student network in knowledge distillation (KD), the proposed method enables collaborative learning between an ensemble teacher network and a student network. Namely, the student network also teaches. Specifically, our method first learns a teacher network that is composed of several assistant networks to generate a shot type sequence and produce the soft target (shot types) distribution accordingly through KD. It then constructs the student network that learns from both the ground truth label (hard target) and the soft target distribution to alleviate the difficulty of optimization and improve generalization capability. As the student network gradually advances, it turns to feed back knowledge to the assistant networks, thereby improving the teacher network in each iteration. Owing to such interactive designs, the DIL mechanism bridges the gap between the teacher and student networks and produces more superior capability for both networks. Objective and subjective experimental results demonstrate that both the teacher and student networks can generate more attractive shot sequences from music, thereby enhancing the viewing and listening experience. Jen-Chun Lin, Wen-Li Wei, Yen-Yu Lin, Tyng-Luh Liu, Hong-Yuan Mark Liao |
ACM Multimedia | 4 |
| 2019 | Unsupervised Meta-Learning of Figure-Ground Segmentation via Imitating Visual Effects
Ding-Jie Chen, Jui-Ting Chien, Hwann-Tzong Chen, Tyng-Luh Liu |
AAAI | 4 |
| 2019 | See-Through-Text Grouping for Referring Image SegmentationabstractMotivated by the conventional grouping techniques to image segmentation, we develop their DNN counterpart to tackle the referring variant. The proposed method is driven by a convolutional-recurrent neural network (ConvRNN) that iteratively carries out top-down processing of bottom-up segmentation cues. Given a natural language referring expression, our method learns to predict its relevance to each pixel and derives a See-through-Text Embedding Pixelwise (STEP) heatmap, which reveals segmentation cues of pixel level via the learned visual-textual co-embedding. The ConvRNN performs a top-down approximation by converting the STEP heatmap into a refined one, whereas the improvement is expected from training the network with a classification loss from the ground truth. With the refined heatmap, we update the textual representation of the referring expression by re-evaluating its attention distribution and then compute a new STEP heatmap as the next input to the ConvRNN. Boosting by such collaborative learning, the framework can progressively and simultaneously yield the desired referring segmentation and reasonable attention distribution over the referring sentence. Our method is general and does not rely on, say, the outcomes of object detection from other DNN models, while achieving state-of-the-art performance in all of the four datasets in the experiments. Ding-Jie Chen, Songhao Jia, Yi-Chen Lo, Hwann-Tzong Chen, Tyng-Luh Liu |
ICCV | 5 |
| 2019 | Tell Me Where It is Still Blurry: Adversarial Blurred Region Mining and RefiningabstractMobile devices such as smart phones are ubiquitously being used to take photos and videos, thus increasing the importance of image deblurring. This study introduces a novel deep learning approach that can automatically and progressively achieve the task via adversarial blurred region mining and refining (adversarial BRMR). Starting with a collaborative mechanism of two coupled conditional generative adversarial networks (CGANs), our method first learns the image-scale CGAN, denoted as iGAN, to globally generate a deblurred image and locally uncover its still blurred regions through an adversarial mining process. Then, we construct the patch-scale CGAN, denoted as pGAN, to further improve sharpness of the most blurred region in each iteration. Owing to such complementary designs, the adversarial BRMR indeed functions as a bridge between iGAN and pGAN, and yields the performance synergy in better solving blind image deblurring. The overall formulation is self-explanatory and effective to globally and locally restore an underlying sharp image. Experimental results on benchmark datasets demonstrate that the proposed method outperforms the current state-of-the-art technique for blind image deblurring both quantitatively and qualitatively. Jen-Chun Lin, Wen-Li Wei, Tyng-Luh Liu, C.-C. Jay Kuo, Hong-Yuan Mark Liao |
ACM Multimedia | 3 |
| 2019 | One-Shot Object Detection with Co-Attention and Co-ExcitationabstractThis paper aims to tackle the challenging problem of one-shot object detection. Given a query image patch whose class label is not included in the training data, the goal of the task is to detect all instances of the same class in a target image. To this end, we develop a novel {\em co-attention and co-excitation} (CoAE) framework that makes contributions in three key technical aspects. First, we propose to use the non-local operation to explore the co-attention embodied in each query-target pair and yield region proposals accounting for the one-shot situation. Second, we formulate a squeeze-and-co-excitation scheme that can adaptively emphasize correlated feature channels to help uncover relevant proposals and eventually the target objects. Third, we design a margin-based ranking loss for implicitly learning a metric to predict the similarity of a region proposal to the underlying query, no matter its class label is seen or unseen in training. The resulting model is therefore a two-stage detector that yields a strong baseline on both VOC and MS-COCO under one-shot setting of detecting objects from both seen and never-seen classes. Ting-I Hsieh, Yi-Chen Lo, Hwann-Tzong Chen, Tyng-Luh Liu |
NeurIPS | 4 |
| 2018 | A2A: Attention to Attention Reasoning for Movie Question Answering
Chao-Ning Liu, Ding-Jie Chen, Hwann-Tzong Chen, Tyng-Luh Liu |
ACCV (6) | 4 |
| 2018 | Cube Padding for Weakly-Supervised Saliency Prediction in 360° VideosabstractAutomatic saliency prediction in 360° videos is critical for viewpoint guidance applications (e.g., Facebook 360 Guide). We propose a spatial-temporal network which is (1) weakly-supervised trained and (2) tailor-made for 360° viewing sphere. Note that most existing methods are less scalable since they rely on annotated saliency map for training. Most importantly, they convert 360° sphere to 2D images (e.g., a single equirectangular image or multiple separate Normal Field-of-View (NFoV) images) which introduces distortion and image boundaries. In contrast, we propose a simple and effective Cube Padding (CP) technique as follows. Firstly, we render the 360° view on six faces of a cube using perspective projection. Thus, it introduces very little distortion. Then, we concatenate all six faces while utilizing the connectivity between faces on the cube for image padding (i.e., Cube Padding) in convolution, pooling, convolutional LSTM layers. In this way, CP introduces no image boundary while being applicable to almost all Convolutional Neural Network (CNN) structures. To evaluate our method, we propose Wild-360, a new 360° video saliency dataset, containing challenging videos with saliency heatmap annotations. In experiments, our method outperforms baseline methods in both speed and quality. Hsien-Tzu Cheng, Chun-Hung Chao, Jin-Dong Dong, Hao-Kai Wen, Tyng-Luh Liu, Min Sun 0001 |
CVPR | 5 |
| 2018 | Seethevoice: Learning from Music to Visual Storytelling of ShotsabstractTypes of shots in the language of film are considered the key elements used by a director for visual storytelling. In filming a musical performance, manipulating shots could stimulate desired effects such as manifesting the emotion or deepening the atmosphere. However, while the visual storytelling technique is often employed in creating professional recordings of a live concert, audience recordings of the same event often lack such sophisticated manipulations. Thus it would be useful to have a versatile system that can perform video mashup to create a refined video from such amateur clips. To this end, we propose to translate the music into a near-professional shot (type) sequence by learning the relation between music and visual storytelling of shots. The resulting shot sequence can then be used to better portray the visual storytelling of a song and guide the concert video mashup process. Our method introduces a novel probabilistic-based fusion approach, named as multi-resolution fused recurrent neural networks (MF-RNNs) with film-language, which integrates multi-resolution fused RNNs and a film-language model for boosting the translation performance. The results from objective and subjective experiments demonstrate that MF-RNNs with film-language can generate an appealing shot sequence with better viewing experience. Wen-Li Wei, Jen-Chun Lin, Tyng-Luh Liu, Yi-Hsuan Yang, Hsin-Min Wang, Hsiao-Rong Tyan, Hong-Yuan Mark Liao |
ICME | 3 |
| 2018 | Coherent Deep-Net Fusion To Classify Shots In Concert VideosabstractVarying types of shots is a fundamental element in the language of film, commonly used by a visual storytelling director. The technique is often used in creating professional recordings of a live concert, but meanwhile may not be appropriately applied in audience recordings of the same event. Such variations could cause the task of classifying shots in concert videos, professional or amateur, very challenging. To achieve more reliable shot classification, we propose a novel probabilistic-based approach, named as coherent classification net (CC-Net), by addressing three crucial issues. First, we focus on learning more effective features by fusing the layer-wise outputs extracted from a deep convolutional neural network (CNN), pretrained on a large-scale data set for object recognition. Second, we introduce a frame-wise classification scheme, the error weighted deep cross-correlation model (EW-Deep-CCM), to boost the classification accuracy. Specifically, the deep neural network-based cross-correlation model (deep-CCM) is constructed to not only model the extracted feature hierarchies of CNN independently, but also relate the statistical dependencies of paired features from different layers. Then, a Bayesian error weighting scheme for a classifier combination is adopted to explore the contributions from individual Deep-CCM classifiers to enhance the accuracy of shot classification in each image frame. Third, we feed the frame-wise classification results to a linear-chain conditional random field module to refine the shot predictions by taking into account the global and temporal regularities. We provide extensive experimental results on a data set of live concert videos to demonstrate the advantage of the proposed CC-Net over existing popular fusion approaches for shot classification. Jen-Chun Lin, Wen-Li Wei, Tyng-Luh Liu, Yi-Hsuan Yang, Hsin-Min Wang, Hsiao-Rong Tyan, Hong-Yuan Mark Liao |
IEEE Trans. Multim. | 3 |
| 2017 | Deep-net fusion to classify shots in concert videosabstractVarying types of shots is a fundamental element in the language of film, commonly used by a visual storytelling director to convey the emotion, ideas, and art. To classify such types of shots from images, we present a new framework that facilitates the intriguing task by addressing two key issues. We first focus on learning more effective features by fusing the layer-wise outputs extracted from a deep convolutional neural network (CNN), pre-trained on a large-scale dataset for object recognition. We then introduce a probabilistic fusion model, termed as error weighted deep cross-correlation model (EW-Deep-CCM), to boost the classification accuracy. Specifically, the deep neural network-based cross-correlation model (Deep-CCM) is constructed to not only model the extracted feature hierarchies of CNN independently but also relate the statistical dependencies of paired features from different layers. Then, a Bayesian error weighting scheme for classifier combination is adopted to explore the contributions from individual Deep-CCM classifiers to enhance the accuracy of shot classification. We provide extensive experimental results on a dataset of live concert videos to demonstrate the advantage of the proposed EW-Deep-CCM over existing popular fusion approaches. The video demos can be found at https://sites.google.com/site/ewdeepccm2/demo. Wen-Li Wei, Jen-Chun Lin, Tyng-Luh Liu, Yi-Hsuan Yang, Hsin-Min Wang, Hsiao-Rong Tyan, Hong-Yuan Mark Liao |
ICASSP | 3 |
| 2016 | Variational Convolutional Networks for Human-Centric Annotations
Tsung-Wei Ke, Che-Wei Lin, Tyng-Luh Liu, Davi Geiger |
ACCV (4) | 3 |
| 2016 | Recursive reduction net for large-scale high-dimensional dataabstractPerforming dimensionality reduction on features is essential in tackling a majority of large-scale computer vision and pattern recognition problems. The popularity of adopting high-dimensional descriptors has caused conventional techniques such as PCA inefficient or even unfeasible. We introduce an unsupervised deep-net approach, termed as recursive reduction net (RRN), to carrying out dimensionality reduction for large-scale high-dimensional data. The proposed iterative algorithm is designed to learn how to merge piecewise reduction results effectively. To this end, we use PCA as the teacher model to establish a reduction net and a fusion net, respectively. To demonstrate the usefulness of RRN, we evaluate the property of variance explaining and carry out extensive experiments on similarity search via binary coding, which would benefit from a proper dimensionality-reduction scheme. Tsung-Wei Ke, Tyng-Luh Liu |
ICIP | 2 |
| 2016 | Guided co-training for multi-view spectral clusteringabstractWe address the problem of how to design a more effective co-training scheme to tackle the multi-view spectral clustering. The conventional co-training procedure treats information from all views equally and often converges to a compromised consensus view that does not fully utilize the multiview information. We instead propose to learn an augmented view and construct its corresponding affinity matrix from a spectral decomposition of an information-rich matrix formed by the eigenvectors of the Laplacian matrices from all views. As the augmented view is expected to be more favorable for carrying out spectral clustering, we design a new pairwise co-training procedure to guide the improvements of the given multiple views separately and iteratively. Our experimental results on three popular benchmark datasets support that the convergent augmented view by the guided co-training process is useful to multi-view spectral clustering, and can yield state-of-the-art performance. Chung-Kuei Lee, Tyng-Luh Liu |
ICIP | 2 |
| 2015 | Spatio-Temporal Learning of Basketball Offensive StrategiesabstractVideo-based group behavior analysis is drawing attention to its rich applications in sports, military, surveillance and biological observations. The recent advances in tracking techniques, based on either computer vision methodology or hardware sensors, further provide the opportunity of better solving this challenging task. Focusing specifically on the analysis of basketball offensive strategies, we introduce a systematic approach to establishing unsupervised modeling of group behaviors. In view that a possible group behavior (offensive strategy) could be of different duration and represented by dynamic player trajectories, the crux of our method is to automatically divide training data into meaningful clusters and learn their respective spatio-temporal model, which is established upon Gaussian mixture regression to account for intra-class spatio-temporal variations. The resulting strategy representation turns out to be flexible that can be used to not only establish the discriminant functions but also improve learning the models. We demonstrate the usefulness of our approach by exploring its effectiveness in analyzing a set of given basketball video clips. Ching-Hang Chen, Tyng-Luh Liu, Yu-Shuen Wang, Hung-Kuo Chu, Nick C. Tang, Hong-Yuan Mark Liao |
ACM Multimedia | 2 |
| 2014 | Efficient binary codes for extremely high-dimensional dataabstractRecent advances in tackling large-scale computer vision problems have supported the use of an extremely high-dimensional descriptor to encode the image data. Under such a setting, we focus on how to efficiently carry out similarity search via employing binary codes. Observe that most of the popular high-dimensional descriptors induce feature vectors that have an implicit 2-D structure. We exploit this property to reduce the computation cost and high complexity. Specifically, our method generalizes the Iterative Quantization (ITQ) framework to handle extremely high-dimensional data in two steps. First, we restrict the dimensionality-reduction projection to a block-diagonal form and decide it by independently solving several moderate-size PCA sub-problems. Second, we replace the full rotation in ITQ with a bilinear rotation to improve the efficiency both in training and testing. Our experimental results on a large-scale dataset and comparisons with a state-of-the-art technique are promising. Tsung-Yu Lin, Tyng-Luh Liu |
ICIP | 2 |
| 2014 | Exploring Depth Information for Object Segmentation and DetectionabstractWe propose a new framework for performing object segmentation and detection simultaneously. Our method leverages with an MRF graphical model that comprises two kinds of nodes and two types of labels for inference. Specifically, we decompose an image into super pixels and generate segment proposals from each super pixel. The super pixels are then duplicated to form the two types of nodes. For each segmentation node, the model is to predict the object class label, while it is to decide the label corresponding to the best segment proposal selection at each detection node. The former is clearly a segmentation problem and the latter a detection problem. We link the two tasks by establishing a unified energy function that has a joint energy term accounting for the compatibility of the segmentation and detection labelings. Marginalizing by fixing either type of variables, the energy function can be switched into the one specifically for detection or segmentation. This property enables an alternating procedure to conveniently obtain the optimal labelings. To better explain the geometry about the objects and the scene, we use the depth information so that 3-D distances between super pixels are available in computing each energy term. Experimental results on a dataset with depth information are provided to support the effectiveness of our method. Tyng-Luh Liu, Kai-Yueh Chang, Shang-Hong Lai |
ICPR | 1 |
| 2013 | A sparse linear model for saliency-guided decolorizationabstractDifferent from most existing decolorization techniques that emphasize preserving image features revealed in the input color space, our proposed method focuses on exploring those in a higher-dimensional feature space. The shift of paradigm is motivated by that decolorization is often sensitive to adopting the various color systems. The results of converting the same color image expressed in different color spaces could vary significantly. We instead consider constructing an image-dependent feature space by learning a representative dictionary, and carry out decolorizing an image by retaining the structures there. To this end, for a given image, the atoms of the dictionary are systematically collected to reflect the visually important/salient contents, and also to concisely reduce chromatic redundancy. A sparse linear model with respect to the learned dictionary is then assumed. Finally, a linear projection to grayscale respecting the inner products in the feature space can be optimized to accomplish the conversion. Chun-Wei Liu, Tyng-Luh Liu |
ICIP | 2 |
| 2011 | From co-saliency to co-segmentation: An efficient and fully unsupervised energy minimization modelabstractWe address two key issues of co-segmentation over multiple images. The first is whether a pure unsupervised algorithm can satisfactorily solve this problem. Without the user's guidance, segmenting the foregrounds implied by the common object is quite a challenging task, especially when substantial variations in the object's appearance, shape, and scale are allowed. The second issue concerns the efficiency if the technique can lead to practical uses. With these in mind, we establish an MRF optimization model that has an energy function with nice properties and can be shown to effectively resolve the two difficulties. Specifically, instead of relying on the user inputs, our approach introduces a co-saliency prior as the hint about possible foreground locations, and uses it to construct the MRF data terms. To complete the optimization framework, we include a novel global term that is more appropriate to co-segmentation, and results in a submodular energy function. The proposed model can thus be optimally solved by graph cuts. We demonstrate these advantages by testing our method on several benchmark datasets. Kai-Yueh Chang, Tyng-Luh Liu, Shang-Hong Lai |
CVPR | 2 |
| 2011 | Fusing generic objectness and visual saliency for salient object detectionabstractWe present a novel computational model to explore the relatedness of objectness and saliency, each of which plays an important role in the study of visual attention. The proposed framework conceptually integrates these two concepts via constructing a graphical model to account for their relationships, and concurrently improves their estimation by iteratively optimizing a novel energy function realizing the model. Specifically, the energy function comprises the objectness, the saliency, and the interaction energy, respectively corresponding to explain their individual regularities and the mutual effects. Minimizing the energy by fixing one or the other would elegantly transform the model into solving the problem of objectness or saliency estimation, while the useful information from the other concept can be utilized through the interaction term. Experimental results on two benchmark datasets demonstrate that the proposed model can simultaneously yield a saliency map of better quality and a more meaningful objectness output for salient object detection. Kai-Yueh Chang, Tyng-Luh Liu, Hwann-Tzong Chen, Shang-Hong Lai |
ICCV | 2 |
| 2011 | Multiple Kernel Learning for Dimensionality ReductionabstractIn solving complex visual learning tasks, adopting multiple descriptors to more precisely characterize the data has been a feasible way for improving performance. The resulting data representations are typically high-dimensional and assume diverse forms. Hence, finding a way of transforming them into a unified space of lower dimension generally facilitates the underlying tasks such as object recognition or clustering. To this end, the proposed approach (termed MKL-DR) generalizes the framework of multiple kernel learning for dimensionality reduction, and distinguishes itself with the following three main contributions: first, our method provides the convenience of using diverse image descriptors to describe useful characteristics of various aspects about the underlying data. Second, it extends a broad set of existing dimensionality reduction techniques to consider multiple kernel learning, and consequently improves their effectiveness. Third, by focusing on the techniques pertaining to dimensionality reduction, the formulation introduces a new class of applications with the multiple kernel learning framework to address not only the supervised learning problems but also the unsupervised and semi-supervised ones. Yen-Yu Lin, Tyng-Luh Liu, Chiou-Shann Fuh |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Regularized Background Adaptation: A Novel Learning Rate Control Scheme for Gaussian Mixture ModelingabstractTo model a scene for background subtraction, Gaussian mixture modeling (GMM) is a popular choice for its capability of adaptation to background variations. However, GMM often suffers from a tradeoff between robustness to background changes and sensitivity to foreground abnormalities and is inefficient in managing the tradeoff for various surveillance scenarios. By reviewing the formulations of GMM, we identify that such a tradeoff can be easily controlled by adaptive adjustments of the GMM's learning rates for image pixels at different locations and of distinct properties. A new rate control scheme based on high-level feedback is then developed to provide better regularization of background adaptation for GMM and to help resolving the tradeoff. Additionally, to handle lighting variations that change too fast to be caught by GMM, a heuristic rooting in frame difference is proposed to assist the proposed rate control scheme for reducing false foreground alarms. Experiments show the proposed learning rate control scheme, together with the heuristic for adaptation of over-quick lighting change, gives better performance than conventional GMM approaches. Horng-Horng Lin, Jen-Hui Chuang, Tyng-Luh Liu |
IEEE Trans. Image Process. | 3 |
| 2010 | Clustering Complex Data with Group-Dependent Feature Selection
Yen-Yu Lin, Tyng-Luh Liu, Chiou-Shann Fuh |
ECCV (6) | 2 |
| 2009 | Co-occurrence Random Forests for Object Localization and Classification
Yu-Wu Chu, Tyng-Luh Liu |
ACCV (3) | 2 |
| 2009 | Learning partially-observed hidden conditional random fields for facial expression recognitionabstractThis paper describes a novel graphical model approach to seamlessly coupling and simultaneously analyzing facial emotions and the action units. Our method is based on the hidden conditional random fields (HCRFs) where we link the output class label to the underlying emotion of a facial expression sequence, and connect the hidden variables to the image frame-wise action units. As HCRFs are formulated with only the clique constraints, their labeling for hidden variables often lacks a coherent and meaningful configuration. We resolve this matter by introducing a partially-observed HCRF model, and establish an efficient scheme via Bethe energy approximation to overcome the resulting difficulties in training. For real-time applications, we also propose an online implementation to perform incremental inference with satisfactory accuracy. Kai-Yueh Chang, Tyng-Luh Liu, Shang-Hong Lai |
CVPR | 2 |
| 2009 | Efficient discriminative local learning for object recognitionabstractAlthough object recognition methods based on local learning can reasonably resolve the difficulties caused by the large variations in images from the same category, the high risk of overfitting and the heavy computational cost in training numerous local models (classifiers or distance functions) often limit their applicability. To address these two unpleasant issues, we cast the multiple, independent training processes of local models as a correlative multi-task learning problem, and design a new boosting algorithm to accomplish it. Specifically, we establish a parametric space where these local models lie and spread as a manifold-like structure, and use boosting to perform local model training by completing the manifold embedding. Via sharing the common embedding space, the learning of each local model can be properly regularized by the extra knowledge from other models, while the training time is also significantly reduced. Experimental results on two benchmark datasets, Caltech-101 and VOC 2007, support that our approach not only achieves promising recognition rates but also gives a two order speed-up in realizing local learning. Yen-Yu Lin, Jyun-Fan Tsai, Tyng-Luh Liu |
ICCV | 3 |
| 2008 | Improving local learning for object categorization by exploring the effects of rankingabstractLocal learning for classification is useful in dealing with various vision problems. One key factor for such approaches to be effective is to find good neighbors for the learning procedure. In this work, we describe a novel method to rank neighbors by learning a local distance function, and meanwhile to derive the local distance function by focusing on the high-ranked neighbors. The two aspects of considerations can be elegantly coupled through a well-defined objective function, motivated by a supervised ranking method called P-Norm Push. While the local distance functions are learned independently, they can be reshaped altogether so that their values can be directly compared. We apply the proposed method to the Caltech-101 dataset, and demonstrate the use of proper neighbors can improve the performance of classification techniques based on nearest-neighbor selection. Tien-Lung Chang, Tyng-Luh Liu, Jen-Hui Chuang |
CVPR | 2 |
| 2008 | Dimensionality Reduction for Data in Multiple Feature RepresentationsabstractIn solving complex visual learning tasks, adopting multiple descriptors to more precisely characterize the data has been a feasible way for improving performance. These representations are typically high dimensional and assume diverse forms. Thus finding a way to transform them into a unified space of lower dimension generally facilitates the underlying tasks, such as object recognition or clustering. We describe an approach that incorporates multiple kernel learning with dimensionality reduction (MKL-DR). While the proposed framework is flexible in simultaneously tackling data in various feature representations, the formulation itself is general in that it is established upon graph embedding. It follows that any dimensionality reduction techniques explainable by graph embedding can be generalized by our method to consider data in multiple feature representations. Yen-Yu Lin, Tyng-Luh Liu, Chiou-Shann Fuh |
NIPS | 2 |
| 2007 | Local Ensemble Kernel Learning for Object Category RecognitionabstractThis paper describes a local ensemble kernel learning technique to recognize/classify objects from a large number of diverse categories. Due to the possibly large intraclass feature variations, using only a single unified kernel-based classifier may not satisfactorily solve the problem. Our approach is to carry out the recognition task with adaptive ensemble kernel machines, each of which is derived from proper localization and regularization. Specifically, for each training sample, we learn a distinct ensemble kernel constructed in a way to give good classification performance for data falling within the corresponding neighborhood. We achieve this effect by aligning each ensemble kernel with a locally adapted target kernel, followed by smoothing out the discrepancies among kernels of nearby data. Our experimental results on various image databases manifest that the technique to optimize local ensemble kernels is effective and consistent for object recognition. Yen-Yu Lin, Tyng-Luh Liu, Chiou-Shann Fuh |
CVPR | 2 |
| 2007 | Finding Familiar Objects and their Depth from a Single ImageabstractWe present a classification-based method to identify objects of interest, and judge their depth in a single image. Our approach is motivated by a postulate of human depth perception that people can give a credible depth estimation for an object whose familiar size is known, even without using stereo vision. To emulate the mechanism, we categorize objects into the same class if they have similar sizes and shapes, and model the sense of discovering a familiar object by applying multiple kernel logistic regression to the conditional probability of feature types. The depth of a detected target can then be obtained by referencing its corresponding object category. Overall, the proposed algorithm is efficient in both the training and testing phases, and does not require a large amount of training images for good performances. Hwann-Tzong Chen, Tyng-Luh Liu |
ICIP (6) | 2 |
| 2006 | Direct Energy Minimization for Super-Resolution on Nonlinear Manifolds
Tien-Lung Chang, Tyng-Luh Liu, Jen-Hui Chuang |
ECCV (4) | 2 |
| 2006 | Segmenting Highly Articulated Video Objects with Weak-Prior Random Forests
Hwann-Tzong Chen, Tyng-Luh Liu, Chiou-Shann Fuh |
ECCV (4) | 2 |
| 2005 | Local Discriminant Embedding and Its VariantsabstractWe present a new approach, called local discriminant embedding (LDE), to manifold learning and pattern classification. In our framework, the neighbor and class relations of data are used to construct the embedding for classification problems. The proposed algorithm learns the embedding for the submanifold of each class by solving an optimization problem. After being embedded into a low-dimensional subspace, data points of the same class maintain their intrinsic neighbor relations, whereas neighboring points of different classes no longer stick to one another. Via embedding, new test data are thus more reliably classified by the nearest neighbor rule, owing to the locally discriminating nature. We also describe two useful variants: two-dimensional LDE and kernel LDE. Comprehensive comparisons and extensive experiments on face recognition are included to demonstrate the effectiveness of our method. Hwann-Tzong Chen, Huang-Wei Chang, Tyng-Luh Liu |
CVPR (2) | 3 |
| 2005 | Tone Reproduction: A Perspective from Luminance-Driven Perceptual GroupingabstractWe address the tone reproduction problem by integrating local adaptation with global-contrast consistency. Many previous works have tried to compress high-dynamic-range (HDR) luminances into a displayable range in imitation of the local adaptation mechanism of human eyes. Nevertheless, while the realization of local adaptation is not theoretically defined, exaggerating such effects often causes unnatural global contrasts. We propose a luminance-driven perceptual grouping process to derive a sparse representation of HDR luminances, and use the grouped regions to approximate local properties of luminances. The advantage of incorporating a sparse representation is twofold: We can simulate local adaptation based on region information, and subsequently apply piecewise tone mappings to monotonize the relative brightness over only a few perceptually significant regions. Our experimental results show that the proposed framework gives a good balance in preserving local details and maintaining global contrasts of HDR scenes. Hwann-Tzong Chen, Tyng-Luh Liu, Tien-Lung Chang |
CVPR (2) | 2 |
| 2005 | Robust Face Detection with Multi-Class BoostingabstractWith the aim to design a general learning framework for detecting faces of various poses or under different lighting conditions, we are motivated to formulate the task as a classification problem over data of multiple classes. Specifically, our approach focuses on a new multi-class boosting algorithm, called MBHboost, and its integration with a cascade structure for effectively performing face detection. There are three main advantages of using MBHboost: 1) each MBH weak learner is derived by sharing a good projection direction such that each class of data has its own decision boundary; 2) the proposed boosting algorithm is established based on an optimal criterion for multi-class classification; and 3) since MBHboost is flexible with respect to the number of classes, it turns out that it is possible to use only one single boosted cascade for the multi-class detection. All these properties give rise to a robust system to detect faces efficiently and accurately. Yen-Yu Lin, Tyng-Luh Liu |
CVPR (1) | 2 |
| 2005 | Learning Effective Image Metrics from Few Pairwise ExamplesabstractWe present a new approach to learning image metrics. The main advantage of our method lies in a formulation that requires only a few pairwise examples. Apparently, based on the little amount of side-information, it would take a very effective learning scheme to yield a useful image metric. Our algorithm achieves this goal by addressing two key issues. First, we establish a global-local (glocal) image representation that induces two structure-meaningful vector spaces to respectively describe the global and the local image properties. Second, we develop a metric optimization framework that finds an optimal bilinear transform to best explain the given side-information. We emphasize it is the glocal image representation that makes the use of bilinear transform more powerful. Experimental results on classifications of face images and visual tracking are included to demonstrate the contributions of the proposed method. Hwann-Tzong Chen, Tyng-Luh Liu, Chiou-Shann Fuh |
ICCV | 2 |
| 2005 | Shape recognition using fast boosted filteringabstractWe address the problem of recognizing 2-D shapes in images via multi-class classifications. Our approach has three key elements. First, a signed distance transform is introduced to represent a shape more informatively. Second, a filter bank is generated such that its filters can capture multiple-scale local and global features between two shapes of different classes. We then apply boosting to combine useful filters to construct discriminant classifiers. Third, in implementing our system, a new classification architecture is developed to accomplish multi-class recognition. To examine the claimed efficiencies, we consider an example of document recognition by pinpointing the strengths of our method through experimental results and comparisons. Yen-Yu Lin, Tyng-Luh Liu |
ICIP (2) | 2 |
| 2005 | Semantic manifold learning for image retrievalabstractLearning the user's semantics for CBIR involves two different sources of information: the similarity relations entailed by the content-based features, and the relevance relations specified in the feedback. Given that, we propose an augmented relation embedding (ARE) to map the image space into a semantic manifold that faithfully grasps the user's preferences. Besides ARE, we also look into the issues of selecting a good feature set for improving the retrieval performance. With these two aspects of efforts we have established a system that yields far better results than those previously reported. Overall, our approach can be characterized by three key properties: 1) The framework uses one relational graph to describe the similarity relations, and the other two to encode the relevant/irrelevant relations indicated in the feedback. 2) With the relational graphs so defined, learning a semantic manifold can be transformed into solving a constrained optimization problem, and is reduced to the ARE algorithm accounting for both the representation and the classification points of views. 3) An image representation based on augmented features is introduced to couple with the ARE learning. The use of these features is significant in capturing the semantics concerning different scales of image regions. We conclude with experimental results and comparisons to demonstrate the effectiveness of our method. Yen-Yu Lin, Tyng-Luh Liu, Hwann-Tzong Chen |
ACM Multimedia | 2 |
| 2005 | Tone Reproduction: A Perspective from Luminance-Driven Perceptual Grouping
Hwann-Tzong Chen, Tyng-Luh Liu, Chiou-Shann Fuh |
Int. J. Comput. Vis. | 2 |
| 2004 | Fast Object Detection with Occlusions
Yen-Yu Lin, Tyng-Luh Liu, Chiou-Shann Fuh |
ECCV (1) | 2 |
| 2004 | Real-Time Tracking Using Trust-Region MethodsabstractOptimization methods based on iterative schemes can be divided into two classes: line-search methods and trust-region methods. While line-search techniques are commonly found in various vision applications, not much attention is paid to trust-region ones. Motivated by the fact that line-search methods can be considered as special cases of trust-region methods, we propose to establish a trust-region framework for real-time tracking. Our approach is characterized by three key contributions. First, since a trust-region tracking system is more effective, it often yields better performances than the outcomes of other trackers that rely on iterative optimization to perform tracking, e.g., a line-search-based mean-shift tracker. Second, we have formulated a representation model that uses two coupled weighting schemes derived from the covariance ellipse to integrate an object's color probability distribution and edge density information. As a result, the system can address rotation and nonuniform scaling in a continuous space, rather than working on some presumably possible discrete values of rotation angle and scale. Third, the framework is very flexible in that a variety of distance functions can be adapted easily. Experimental results and comparative studies are provided to demonstrate the efficiency of the proposed method. Tyng-Luh Liu, Hwann-Tzong Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2003 | Representation and Self-Similarity of ShapesabstractRepresenting shapes in a compact and informative form is a significant problem for vision systems that must recognize or classify objects. We describe a compact representation model for two-dimensional (2D) shapes by investigating their self-similarities and constructing their shape axis trees (SA-trees). Our approach can be formulated as a variational one (or, equivalently, as MAP estimation of a Markov random field). We start with a 2D shape, its boundary contour, and two different parameterizations for the contour (one parameterization is oriented counterclockwise and the other clockwise). To measure its self-similarity, the two parameterizations are matched to derive the best set of one-to-one point-to-point correspondences along the contour. The cost functional used in the matching may vary and is determined by the adopted self-similarity criteria, e.g., cocircularity, distance variation, parallelism, and region homogeneity. The loci of middle points of the pairing contour points yield the shape axis and they can be grouped into a unique free tree structure, the SA-tree. By implicitly encoding the (local and global) shape information into an SA-tree, a variety of vision tasks, e.g., shape recognition, comparison, and retrieval, can be performed in a more robust and efficient way via various tree-based algorithms. A dynamic programming algorithm gives the optimal solution in O(N/sup 1/), where N is the size of the contour. Davi Geiger, Tyng-Luh Liu, Robert Kohn |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2002 | A probabilistic SVM approach for background scene initializationabstractVisual tracking systems using background subtraction have been very popular largely due to their efficiency in extracting moving objects. However, such systems often compute the reference background by assuming no moving objects are present during the initialization stage, though the assumption may not be realistic. We propose an automatic way to perform background initialization using a probabilistic SVM (support vector machine). By formulating the problem as an on-line classification one, our approach has the potential to be real-time. SVM classification is carried out for all elements of each image frame by computing the output probabilities. Newly found background elements are evaluated and determined if they should be added to the solution. The process of background initialization continues until there are no more new background elements to be considered. As the features used in an SVM dictate the outcome of classification, we find that optical flow value and inter-frame difference are the two most important ones. Experimental results are included to demonstrate the efficiency of our method. Horng-Horng Lin, Tyng-Luh Liu, Jen-Hui Chuang |
ICIP (3) | 2 |
| 2001 | Multi-Object Tracking Using Dynamical Graph MatchingabstractWe describe a tracking algorithm to address the interactions among objects, and to track them individually and confidently via a static camera. It is achieved by constructing an invariant bipartite graph to model the dynamics of the tracking process, of which the nodes are classified into objects and profiles. The best match of the graph corresponds to an optimal assignment for resolving the identities of the detected objects. Since objects may enter/exit the scene indefinitely, or when interactions occur/conclude they could form/leave a group, the number of nodes in the graph changes dynamically. Therefore it is critical to maintain an invariant property to assure that the numbers of nodes of both types are kept the same so that the matching problem is manageable. In addition, several important issues are also discussed, including reducing the effect of shadows, extracting objects' shapes, and adapting large abrupt changes in the scene background. Finally, experimental results are provided to illustrate the efficiency of our approach. Hwann-Tzong Chen, Horng-Horng Lin, Tyng-Luh Liu |
CVPR (2) | 3 |
| 2001 | Trust-Region Methods for Real-Time TrackingabstractOptimization methods based on iterative schemes can be divided into two classes: linesearch methods and trust-region methods. While linesearch techniques are commonly found in various vision applications, not much attention is paid to trust-region methods. Motivated by the fact that linesearch methods can be considered as special cases of trust-region methods, we propose to apply trust-region methods to visual tracking problems. Our approach integrates trust-region methods with the Kullback Leibler distance to track a rigid or non-rigid object in real-time. If not limited by the speed of a camera, the algorithm can achieve frame rate above 60 fps. To justify our method, a variety of experiments/comparisons are carried out for the trust-region tracker and a linesearch-based mean-shift tracker with same initial conditions. The experimental results support our conjecture that a trust-region tracker should perform superiorly to a linesearch one. Hwann-Tzong Chen, Tyng-Luh Liu |
ICCV | 2 |
| 2000 | A Variational Approach for Digital WatermarkingabstractWe describe a digital watermarking framework by minimizing an appropriate variational problem with an energy/cost function accounting for properties such as robustness, fidelity, and so forth. The process of embedding a watermark into a multimedia data can then be treated as solving an optimization problem, where we show that it can be transformed into finding the minimum cost (integral) flow of a watermarking network. The framework is very general and flexible since the cost function can be easily modified to accommodate other criteria. Unlike other watermarking methods, our approach is implemented via solving a network flow problem, and consequently this protects the watermarking scheme from deliberate attacks designed specifically using the knowledge of the algorithm. Tyng-Luh Liu, Hwann-Tzong Chen |
ICIP | 1 |
| 2000 | A Generalized Shape-Axis Model for Planar ShapesabstractWe describe a generalized shape-axis (SA) model for representing both open and closed planar curves. The SA model is an effective way to represent shapes by comparing their self-similarities. Given a 2D shape, whether it is closed or open, we use two different parametrizations for the curve. To study the self-similarity, the two parametrizations are matched to each other via a variational framework, where the self-similarity criterion is to be defined depending on the class of shapes and human perception factors. Useful self-similarity criteria include symmetry, parallelism and convexity, etc. A match is allowed to have discontinuities, and the optimal match can be computed by a dynamic programming algorithm in O(N/sup 4/) time, where N is the size of the shape. We use a grouping process for the shape axis to construct a unique SA-tree, however, when a planar shape is open, it is possible to derive an SA-forest. The generalized SA model provides a compact and informative way for 2D shape representation. Tyng-Luh Liu |
ICPR | 1 |
| 1999 | Approximate Tree Matching and Shape SimilarityabstractWe present a framework for 2D shape contour (silhouette) comparison that can account for stretchings, occlusions and region information. Topological changes due to the original 3D scenarios and articulations are also addressed. To compare the degree of similarity between any two shapes, our approach is to represent each shape contour with a free tree structure derived from a shape axis (SA) model, which we have recently proposed. We then use a tree matching scheme to find the best approximate match and the matching cost. To deal with articulations, stretchings and occlusions, three local tree matching operations, merge, cut, and merge-and-cut, are introduced to yield optimally approximate matches, which can accommodate not only one-to-one but many-to-many mappings. The optimization process gives guaranteed globally optimal match efficiently. Experimental results on a variety of shape contours are provided. Tyng-Luh Liu, Davi Geiger |
ICCV | 1 |
| 1999 | Sparse Representations for Image Decompositions
Davi Geiger, Tyng-Luh Liu, Michael J. Donahue |
Int. J. Comput. Vis. | 2 |
| 1998 | Representation and Self-Similarity of ShapesabstractRepresenting shapes is a significant problem for vision systems that must recognize or classify objects. We derive a representation for a given shape by investigating its self-similarities, and constructing its shape axis (SA) and shape axis tree (SA-tree). We start with a shape, its boundary contour, and two different parameterizations for the contour. To measure its self-similarity we consider matching pairs of points (and their tangents) along the boundary contour, i.e., matching the two parameterizations. The matching, of self-similarity criteria may vary, e.g., co-circularity, parallelism, distance, region homogeneity. The loci of middle points of the pairing contour points are the shape axis and they can be grouped into a unique tree graph, the SA-tree. The shape axis for the co-circularity criteria is compared to the symmetry axis. An interpretation in terms of object parts is also presented. Tyng-Luh Liu, Davi Geiger, Robert Kohn |
ICCV | 1 |
| 1998 | Segmenting by seeking the symmetry axisabstractWe introduce a method for segmenting a shape from an image and simultaneously determining its symmetry axis. The symmetry is used to help the segmentation and in turn the segmentation determines the symmetry. The problem is formulated as one of minimizing a goodness of fitness function and Dijkstra's algorithm is used to find the global minimum of the cost function. The results are illustrated on real images. Tyng-Luh Liu, Davi Geiger, Alan L. Yuille |
ICPR | 1 |
| 1996 | Sparse Representations for Image Decomposition with OcclusionsabstractWe study the problem of how to detect "interesting objects" appeared in a given image, I. Our approach is to treat it as a function approximation problem based on an over-redundant basis, and also account for occlusions, where the basis superposition principle is no longer valid. Since the basis (a library of image templates) is over-redundant, there are infinitely many ways to decompose I. We are motivated to select a sparse/compact representation of I, and to account for occlusions and noise. We then study a greedy and iterative "weighted L/sup p/ Matching Pursuit" strategy, with O<p<1. We use an L/sup p/ result to compute a solution, select the best template, at each stage of the pursuit. Michael J. Donahue, Davi Geiger, Tyng-Luh Liu, Robert A. Hummel |
CVPR | 3 |
| 1996 | Image Recognition with Occlusions
Tyng-Luh Liu, Michael J. Donahue, Davi Geiger, Robert A. Hummel |
ECCV (1) | 1 |
| 1996 | Recognizing Articulated Objects with Information Theoretic MethodsabstractThis paper addresses the problem of recognizing articulated and deformable objects. In particular we are interested in human arm and leg articulations. Our approach is a Bayesian-Information integration of shape similarity and snakes, and naturally combines top-down and bottom-up algorithms. The bottom-up method extracts edges, then constructs snakes (or contours) by grouping edge elements and feeds the shape analysis. The top-down one uses shape analysis, by comparing the object model with the extracted snakes, to guide/prune the search for other snakes. The optimizations are based on Dijkstra algorithm and further pruning of this algorithm is obtained by "integration by parts". Our approach is general enough to handle three dimensional objects, but our focus here is on two dimensional contours. Davi Geiger, Tyng-Luh Liu |
FG | 2 |