EDBT 2026 Demo / reviewers in the wild / expert
Chu-Song Chen
dblp:67/1007
· DBLP profile ↗
141ranked-venue papers
10as first author
26since 2021 · last 2026
0000-0002-2959-2471ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 104 · 5 first-author · 20 since 2021Artificial intelligence and machine learning · 67 · 8 first-author · 14 since 2021Databases, data management, data science and information retrieval · 7 · 1 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Systems, architecture and hardware · 4 · 1 first-author · 1 since 2021Computer networks · 3Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented GenerationabstractRetrieval-augmented generation (RAG) enables large language models (LLMs) to dynamically access external information, which is powerful for answering questions over previously unseen documents.Nonetheless, they struggle with high-level conceptual understanding and holistic comprehension due to limited context windows, which constrain their ability to perform deep reasoning over long-form, domainspecific content such as full-length books.To solve this problem, knowledge graphs (KGs) have been leveraged to provide entity-centric structure and hierarchical summaries, offering more structured support for reasoning.However, existing KG-based RAG solutions remain restricted to text-only inputs and fail to leverage the complementary insights provided by other modalities such as vision.On the other hand, reasoning from visual documents requires textual, visual, and spatial cues into structured, hierarchical concepts.To address this issue, we introduce a multimodal knowledge graphbased RAG that enables cross-modal reasoning for better content understanding.Our method incorporates visual cues into the construction of knowledge graphs, the retrieval phase, and the answer generation process.Experimental results across both global and fine-grained question answering tasks show that our approach consistently outperforms existing approaches on both textual and multimodal benchmarks. Chi-Hsiang Hsiao, Yi-Cheng Wang, Tzung-Sheng Lin, Yi-Ren Yeh, Chu-Song Chen |
ACL (1) | 5 |
| 2025 | Continual Learning for Weakly-Supervised Histopathology Tissue SegmentationabstractWeakly supervised histopathology segmentation is a widely studied field that aims to achieve pixel-level semantic segmentation using image-level annotations, reducing the need for labor-intensive labeling. Despite significant advances in this task, existing methods assume the availability of all training data at once during training. Since medical image collections typically expand over time in practice, such methods become impractical. Meanwhile, research on continual semantic segmentation has also made significant progress. However, most existing works still rely on pixel-level annotations to train models. As a result, integrating continual learning into weakly supervised segmentation models has emerged as a promising direction. To address this challenge, we propose CL4WSeg, a novel end-to-end transformer-based framework that employs temporal distillation to leverage features from previous models for continual weakly supervised segmentation. Furthermore, we utilize a controllable diffusion model to enable generative replay and integrate an image quality filter to collect high-quality images, alleviating catastrophic forgetting. Experiments on the LUAD-HistoSeg, BCSS-WSSS, and WSSS4LUAD datasets demonstrate that our approach outperforms state-of-the-art methods. Huei-Fang Yang, Chu-Song Chen |
CIBCB | 3 |
| 2025 | Relation-Rich Visual Document Generator for Visual Information ExtractionabstractDespite advances in Large Language Models (LLMs) and Multimodal LLMs (MLLMs) for visual document understanding (VDU), visual information extraction (VIE) from relation-rich documents remains challenging due to the layout diversity and limited training data. While existing synthetic document generators attempt to address data scarcity, they either rely on manually designed layouts and templates, or adopt rule-based approaches that limit layout diversity. Besides, current layout generation methods focus solely on topological patterns without considering textual content, making them impractical for generating documents with complex associations between the contents and layouts. In this paper, we propose a Relation-rIch visual Document GEnerator (RIDGE) that addresses these limitations through a two-stage approach: (1) Content Generation, which leverages LLMs to generate document content using a carefully designed Hierarchical Structure Text format which captures entity categories and relationships, and (2) Content-driven Layout Generation, which learns to create diverse, plausible document layouts solely from easily available Optical Character Recognition (OCR) results, requiring no human labeling or annotations efforts. Experimental results have demonstrated that our method significantly enhances the performance of document understanding models on various VIE benchmarks. Zi-Han Jiang, Chien-Wei Lin, Hsuan-Tung Liu, Yi-Ren Yeh, Chu-Song Chen |
CVPR | 6 |
| 2025 | PDSeg: Patch-Wise Distillation and Controllable Image Generation for Weakly-Supervised Histopathology Tissue SegmentationabstractWeakly-supervised semantic segmentation, which achieves pixel-wise segmentation using image-level labels, has emerged as an alternative to fully supervised methods by reducing the need for detailed annotations. Inspired by the recent success of the teacher-student strategy in various vision tasks, we present a transformer-based weakly supervised framework that distills knowledge from a CNN teacher. Specifically, we incorporate a sequence of patch-wise distillation tokens into the transformer student, with each token focused on learning a specific patch under the teacher’s guidance. This design enables the teacher to provide more reliable supervision to the student. On the other hand, in pathology images, it is often observed that certain tissue types are less represented than others. This class imbalance poses a significant challenge for many WSSS algorithms. To address this issue, we further introduce a data synthesis pipeline using a diffusion model conditioned on semantic label maps to mitigate the effects of class imbalance in histopathology images. Unlike previous methods that rely on full annotations to construct semantic label maps, our approach leverages the intrinsic characteristics of histopathology images. This leads to an approach that does not require full annotations and is well-suited for weakly-supervised scenarios. Through extensive experiments on the LUAD-HistoSeg and BCSS-WSSS datasets, we demonstrate that our approach outperforms state-of-the-art methods. Yu-Hsing Hsieh, Huei-Fang Yang, Chu-Song Chen |
ICASSP | 4 |
| 2025 | Mask-aware Text-to-Image Retrieval: Referring Expression Segmentation Meets Cross-modal RetrievalabstractText-to-image retrieval (TIR) aims to find relevant images based on a textual query, but existing approaches are primarily based on whole-image captions and lack interpretability. Meanwhile, referring expression segmentation (RES) enables precise object localization based on natural language descriptions but is computationally expensive when applied across large image collections. To bridge this gap, we introduce Mask-aware TIR (MaTIR), a new task that unifies TIR and RES, requiring both efficient image search and accurate object segmentation. To address this task, we propose a two-stage framework, comprising a first stage for segmentation-aware image retrieval and a second stage for reranking and object grounding with a multimodal large language model (MLLM). We leverage SAM 2 to generate object masks and Alpha-CLIP to extract region-level embeddings offline at first, enabling effective and scalable online retrieval. Secondly, MLLM is used to refine retrieval rankings and generate bounding boxes, which are matched to segmentation masks. We evaluate our approach on COCO and D3 datasets, demonstrating significant improvements in both retrieval accuracy and segmentation quality over previous methods. Our code is available at https://github.com/AI-Application-and-Integration-Lab/MaTIR. Li-Cheng Shen, Jih-Kang Hsieh, Chu-Song Chen |
ICMR | 4 |
| 2025 | Safety Depth in Large Language Models: A Markov Chain PerspectiveabstractLarge Language Models (LLMs) are increasingly adopted in high-stakes scenarios, yet their safety mechanisms often remain fragile. Simple jailbreak prompts or even benign fine-tuning can bypass internal safeguards, underscoring the need to understand the failure modes of current safety strategies. Recent findings suggest that vulnerabilities emerge when alignment is confined to only the initial output tokens. To address this, we introduce the notion of safety depth, a designated output position where the model refuses to generate harmful content. While deeper alignment appears promising, identifying the optimal safety depth remains an open and underexplored challenge.
We leverage the equivalence between autoregressive language models and Markov chains to derive the first theoretical result on identifying the optimal safety depth. To reach this safety depth effectively, we propose a cyclic group augmentation strategy that improves safety scores across six LLMs. In addition, we uncover a critical interaction between safety depth and ensemble width, demonstrating that larger ensembles can offset shallower alignments. These results suggest that test-time computation, often overlooked in safety alignment, can play a key role. Our approach provides actionable insights for building safer LLMs. Ching-Chia Kao, Chia-Mu Yu, Chun-Shien Lu, Chu-Song Chen |
NeurIPS | 4 |
| 2025 | Defending Against Repetitive Backdoor Attacks on Semi-Supervised Learning Through Lens of Rate-Distortion-Perception Trade-OffabstractSemi-supervised learning (SSL) has achieved remarkable performance with a small fraction of labeled data by leveraging vast amounts of unlabeled data from the Internet. However, this large pool of untrusted data is extremely vulnerable to data poisoning, leading to potential backdoor attacks. Current backdoor defenses are not yet effective against such a vulnerability in SSL. In this study, we propose a novel method, Unlabeled Data Purification (UPure), to disrupt the association between trigger patterns and target classes by introducing perturbations in the frequency domain. By leveraging the Rate-Distortion-Perception (RDP) trade-off, we further identify the frequency band, where the perturbations are added, and Justify this selection. Notably, UPure purifies poisoned unlabeled data without the need of extra clean labeled data. Extensive experiments on four benchmark datasets and five SSL algorithms demonstrate that UPure effectively reduces the attack success rate from 99.78% to 0% while maintaining model accuracy. Code is available here: https://github.com/chengyi-chris/UPure. Cheng-Yi Lee 0001, Ching-Chia Kao, Cheng-Han Yeh, Chun-Shien Lu, Chia-Mu Yu, Chu-Song Chen |
WACV | 6 |
| 2024 | SAM4MLLM: Enhance Multi-Modal Large Language Model for Referring Expression Segmentation
Yi-Chia Chen, Cheng Sun 0004, Yu-Chiang Frank Wang, Chu-Song Chen |
ECCV (81) | 5 |
| 2024 | Open-Vocabulary Panoptic Segmentation Using Bert Pre-Training of Vision-Language Multiway Transformer ModelabstractOpen-vocabulary panoptic segmentation remains a challenging problem. One of the biggest difficulties lies in training models to generalize to an unlimited number of classes using limited categorized training data. Recent popular methods involve large-scale vision-language pre-trained foundation models, such as CLIP. In this paper, we propose OMTSeg for open-vocabulary segmentation using another large-scale vision-language pre-trained model called BEiT-3 and leveraging the cross-modal attention between visual and linguistic features in BEiT-3 to achieve better performance. Experiments result demonstrates that OMTSeg performs favorably against state-of-the-art models. Code is available at https://github.com/AI-Application-and-IntegrationLab/OMTSeg Yi-Chia Chen, Chu-Song Chen |
ICIP | 3 |
| 2024 | On the Higher Moment Disparity of Backdoor AttacksabstractBackdoor attacks are a significant concern in deep learning, especially in applications where models are trained on data from untrusted sources. Plenty of approaches use latent representations of a backdoor model to separate trigger samples from clean ones. However, these defenses rely on some clean data to train a classifier. Recently, researchers have designed adaptive attacks that are latently inseparable, making it even harder for the defender to prevent backdoor attacks. For these reasons, we propose a novel defense, Higher Moment Disparity (HMD), based on the higher moment inspired by latent statistics. HMD uses no clean data and all intermediate representations to avoid previous concerns. Extensive experiments show that our defense against various attacks is promising. Ching-Chia Kao, Cheng-Yi Lee 0001, Chun-Shien Lu, Chia-Mu Yu, Chu-Song Chen |
ICME | 5 |
| 2024 | Adversarially Robust Deepfake Detection via Adversarial Feature Similarity Learning
Sarwar Khan, Jun-Cheng Chen, Wen-Hung Liao, Chu-Song Chen |
MMM (3) | 4 |
| 2023 | Scalable Spatial Memory for Scene Rendering and NavigationabstractNeural scene representation and rendering methods have shown promise in learning the implicit form of scene structure without supervision. However, the implicit representation learned in most existing methods is non-expandable and cannot be inferred online for novel scenes, which makes the learned representation difficult to be applied across different reinforcement learning (RL) tasks. In this work, we introduce Scene Memory Network (SMN) to achieve online spatial memory construction and expansion for view rendering in novel scenes. SMN models the camera projection and back-projection as spatially aware memory control processes, where the memory values store the information of the partial 3D area, and the memory keys indicate the position of that area. The memory controller can learn the geometry property from observations without the camera's intrinsic parameters and depth supervision. We further apply the memory constructed by SMN to exploration and navigation tasks. The experimental results reveal the generalization ability of our proposed SMN in large-scale scene synthesis and its potential to improve the performance of spatial RL tasks. Wen-Cheng Chen, Chu-Song Chen, Walon Wei-Chen Chiu, Min-Chun Hu 0001 |
AAAI | 2 |
| 2023 | LC4SV: A Denoising Framework Learning to Compensate for Unseen Speaker Verification ModelsabstractThe performance of speaker verification (SV) models may drop dramatically in noisy environments. A speech enhancement (SE) module can be used as a front-end strategy. However, existing SE methods may fail to bring performance improvements to downstream SV systems due to artifacts in the predicted signals of SE models. To compensate for artifacts, we propose a generic denoising framework named LC4SV, which can serve as a pre-processor for various unknown downstream SV models. In LC4SV, we employ a learning-based interpolation agent to automatically generate the appropriate coefficients between the enhanced signal and its noisy input to improve SV performance in noisy environments. Our experimental results demonstrate that LC4SV consistently improves the performance of various unseen SV systems. To the best of our knowledge, this work is the first attempt to develop a learning-based interpolation scheme aiming at improving SV performance in noisy environments. Chi-Chang Lee, Chu-Song Chen, Hsin-Min Wang, Tsung-Te Liu, Yu Tsao 0001 |
ASRU | 3 |
| 2023 | Continual Cell Instance Segmentation of Microscopy ImagesabstractA continual cell instance segmenter aims to continually learn to segment new objects while preserving the ability to localize and distinguish old objects without access to previous data. Besides catastrophic forgetting, background shift, where the background class could contain objects in the old and unseen future classes, could occur. In addition, as acquiring annotations is label-intensive, cell images can be partially labeled. In this paper, we present iMRCNN, which extends Mask R-CNN with knowledge distillation and pseudo labeling, to address these challenges. To preserve the learned skills, the current student distills knowledge from the former teacher at output and feature levels. Furthermore, we employ a pseudo labeling scheme, where the teacher is utilized to identify objects with no labels provided, to deal with background shift and partially labeled data. Experiments on two microscopy image sets demonstrate the effectiveness of iMRCNN over other alternatives in various incremental learning scenarios. Tzu-Ting Chuang, Ting-Yun Wei, Yu-Hsing Hsieh, Chu-Song Chen, Huei-Fang Yang |
ICASSP | 4 |
| 2023 | Hearing and Seeing Abnormality: Self-Supervised Audio-Visual Mutual Learning for Deepfake DetectionabstractThe recent development of deepfakes has resulted in serious threats to society, such as spreading misinformation, defamation, etc. Although recent deepfake detection methods are capable of achieving satisfactory results for seen forgeries, the performance drops significantly for unseen ones. With proper supervised pretraining on auxiliary tasks as prior, the situation can be improved, but the requirement to collect a large number of additional annotations for these tasks may restrict the further development of a generalized deep-fake detector. To address this issue, we propose an Audio-Visual Temporal Synchronization for Deepfake Detection framework for detecting deepfakes that maintains reasonable detection capabilities for unseen ones. The primary objective of our framework is to determine whether there has been a forgery by evaluating the consistency between the sound and the faces in a video clip, together with the relationship between the two features. First, the spatiotemporal feature extraction network is pretrained in a self-supervised manner by exploiting the audio-visual temporal synchronization task to build up a rich representation based on the temporal synchronization relationship between the audio and its corresponding video. For pretraining, we use only real data and carefully selected negative samples with contrastive loss to train the model. A temporal classifier network is used to determine whether or not the video has been manipulated using the representations obtained from the pretrained feature extraction networks. To prevent the model from overfitting to certain manipulation-specific artifacts, we froze the feature extraction networks and only trained the final classifier network on forged data. Extensive experiments on unseen forgery categories and unseen datasets have shown the effectiveness of our method to achieve state-of-the-art results. Chang-Sung Sung, Jun-Cheng Chen, Chu-Song Chen |
ICASSP | 3 |
| 2023 | Class-incremental Continual Learning for Instance Segmentation with Image-level Weak SupervisionabstractInstance segmentation requires labor-intensive manual labeling of the contours of complex objects in images for training. The labels can also be provided incrementally in practice to balance the human labor in different time steps. However, research on incremental learning for instance segmentation with only weak labels is still lacking. In this paper, we propose a continual-learning method to segment object instances from image-level labels. Unlike most weakly-supervised instance segmentation (WSIS) which relies on traditional object proposals, we transfer the semantic knowledge from weakly-supervised semantic segmentation (WSSS) to WSIS to generate instance cues. To address the background shift problem in continual learning, we employ the old class segmentation results generated by the previous model to provide more reliable semantic and peak hypotheses. To our knowledge, this is the first work on weakly-supervised continual learning for instance segmentation of images. Experimental results show that our method can achieve better performance on Pascal VOC and COCO datasets under various incremental settings1. Yu-Hsing Hsieh, Guan-Sheng Chen, Shun-Xian Cai, Ting-Yun Wei, Huei-Fang Yang, Chu-Song Chen |
ICCV | 6 |
| 2023 | Domain-Generalized Face Anti-Spoofing with Unknown AttacksabstractAlthough face anti-spoofing (FAS) methods have achieved remarkable performance on specific domains or attack types, few studies have focused on the simultaneous presence of domain changes and unknown attacks, which is closer to real application scenarios. To handle domain-generalized unknown attacks, we introduce a new method, DGUA-FAS, which consists of a Transformer-based feature extractor and a synthetic unknown attack sample generator (SUASG). The SUASG network simulates unknown attack samples to assist the training of the feature extractor. Experimental results show that our method achieves superior performance on domain generalization FAS with known or unknown attacks. Zong-Wei Hong, Hsuan-Tung Liu, Yi-Ren Yeh, Chu-Song Chen |
ICIP | 5 |
| 2023 | D4AM: A General Denoising Framework for Downstream Acoustic Models
Chi-Chang Lee, Yu Tsao 0001, Hsin-Min Wang, Chu-Song Chen |
ICLR | 4 |
| 2023 | Domain Invariant Vision Transformer Learning for Face Anti-spoofingabstractExisting face anti-spoofing (FAS) models have achieved high performance on specific datasets. However, for the application of real-world systems, the FAS model should generalize to the data from unknown domains rather than only achieve good results on a single baseline. As vision transformer models have demonstrated astonishing performance and strong capability in learning discriminative information, we investigate applying transformers to distinguish the face presentation attacks over unknown domains. In this work, we propose the Domain-invariant Vision Transformer (DiVT) for FAS, which adopts two losses to improve the generalizability of the vision transformer. First, a concentration loss is employed to learn a domain-invariant representation that aggregates the features of real face data. Second, a separation loss is utilized to union each type of attack from different domains. The experimental results show that our proposed method achieves state-of-the-art performance on the protocols of domain-generalized FAS tasks. Compared to previous domain generalization FAS models, our proposed method is simpler but more effective. Chen-Hao Liao, Wen-Cheng Chen, Hsuan-Tung Liu, Yi-Ren Yeh, Min-Chun Hu 0001, Chu-Song Chen |
WACV | 6 |
| 2022 | Continual Learning for Visual Search with Backward Consistent Feature EmbeddingabstractIn visual search, the gallery set could be incrementally growing and added to the database in practice. However, existing methods rely on the model trained on the entire dataset, ignoring the continual updating of the model. Besides, as the model updates, the new model must re-extract features for the entire gallery set to maintain compatible feature space, imposing a high computational cost for a large gallery set. To address the issues of long-term visual search, we introduce a continual learning (CL) approach that can handle the incrementally growing gallery set with backward embedding consistency. We enforce the losses of inter-session data coherence, neighbor-session model coherence, and intra-session discrimination to conduct a continual learner. In addition to the disjoint setup, our CL solution also tackles the situation of increasingly adding new classes for the blurry boundary without assuming all categories known in the beginning and during model update. To our knowledge, this is the first CL method both tackling the issue of backward-consistent feature embedding and allowing novel classes to occur in the new sessions. Extensive experiments on various benchmarks show the efficacy of our approach under a wide range of setups11Code: https://github.com/ivclab/CVS. Timmy S. T. Wan, Jun-Cheng Chen, Tzer-Yi Wu, Chu-Song Chen |
CVPR | 4 |
| 2022 | NASTAR: Noise Adaptive Speech Enhancement with Target-Conditional ResamplingabstractFor deep learning-based speech enhancement (SE) systems, the training-test acoustic mismatch can cause notable performance degradation.To address the mismatch issue, numerous noise adaptation strategies have been derived.In this paper, we propose a novel method, called noise adaptive speech enhancement with target-conditional resampling (NASTAR), which reduces mismatches with only one sample (one-shot) of noisy speech in the target environment.NASTAR uses a feedback mechanism to simulate adaptive training data via a noise extractor and a retrieval model.The noise extractor estimates the target noise from the noisy speech, called pseudo-noise.The noise retrieval model retrieves relevant noise samples from a pool of noise signals according to the noisy speech, called relevant-cohort.The pseudo-noise and the relevant-cohort set are jointly sampled and mixed with the source speech corpus to prepare simulated training data for noise adaptation.Experimental results show that NASTAR can effectively use one noisy speech sample to adapt an SE model to a target condition.Moreover, both the noise extractor and the noise retrieval model contribute to model adaptation.To our best knowledge, NASTAR is the first work to perform one-shot noise adaptation through noise extraction and retrieval. Chi-Chang Lee, Cheng-Hung Hu, Yuchen Lin 0003, Chu-Song Chen, Hsin-Min Wang, Yu Tsao 0001 |
INTERSPEECH | 4 |
| 2022 | Learning Binary Hash Codes Based on Adaptable Label RepresentationsabstractThe goal of supervised hashing is to construct hash mappings from collections of images and semantic annotations such that semantically relevant images are embedded nearby in the learned binary hash representations. Existing deep supervised hashing approaches that employ classification frameworks with a classification training objective for learning hash codes often encode class labels as one-hot or multi-hot vectors. We argue that such label encodings do not well reflect semantic relations among classes and instead, effective class label representations ought to be learned from data, which could provide more discriminative signals for hashing. In this article, we introduce Adaptive Labeling Deep Hashing (AdaLabelHash) that learns binary hash codes based on learnable class label representations. We treat the class labels as the vertices of a K -dimensional hypercube, which are trainable variables and adapted together with network weights during the backward network training procedure. The label representations, referred to as codewords, are the target outputs of hash mapping learning. In the label space, semantically relevant images are then expressed by the codewords that are nearby regarding Hamming distances, yielding compact and discriminative binary hash representations. Furthermore, we find that the learned label representations well reflect semantic relations. Our approach is easy to realize and can simultaneously construct both the label representations and the compact binary embeddings. Quantitative and qualitative evaluations on several popular benchmarks validate the superiority of AdaLabelHash in learning effective binary codes for image search. Huei-Fang Yang, Cheng-Hao Tu 0001, Chu-Song Chen |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | 360-Degree Gaze Estimation in the Wild Using Multiple Zoom Scales
Ashesh Mishra, Chu-Song Chen, Hsuan-Tien Lin |
BMVC | 2 |
| 2021 | 3D Video Stabilization With Depth Estimation by CNN-Based OptimizationabstractVideo stabilization is an essential component of visual quality enhancement. Early methods rely on feature tracking to recover either 2D or 3D frame motion, which suffer from the robustness of local feature extraction and tracking in shaky videos. Recently, learning-based methods seek to find frame transformations with high-level information via deep neural networks to overcome the robustness issue of feature tracking. Nevertheless, to our best knowledge, no learning-based methods leverage 3D cues for the transformation inference yet; hence they would lead to artifacts on complex scene-depth scenarios. In this paper, we propose Deep3D Stabilizer, a novel 3D depth-based learning method for video stabilization. We take advantage of the recent self-supervised framework on jointly learning depth and camera ego-motion estimation on raw videos. Our approach requires no data for pre-training but stabilizes the input video via 3D reconstruction directly. The rectification stage incorporates the 3D scene depth and camera motion to smooth the camera trajectory and synthesize the stabilized video. Unlike most one-size-fits-all learning-based methods, our smoothing algorithm allows users to manipulate the stability of a video efficiently. Experimental results on challenging benchmarks show that the proposed solution consistently outperforms the state-of-the-art methods on almost all motion categories. Yao-Chih Lee, Kuan-Wei Tseng, Yu-Ta Chen, Chien-Cheng Chen, Chu-Song Chen, Yi-Ping Hung |
CVPR | 5 |
| 2021 | STR-GQN: Scene Representation and Rendering for Unknown Cameras Based on Spatial Transformation RoutingabstractGeometry-aware modules are widely applied in recent deep learning architectures for scene representation and rendering. However, these modules require intrinsic camera information that might not be obtained accurately. In this paper, we propose a Spatial Transformation Routing (STR) mechanism to model the spatial properties without applying any geometric prior. The STR mechanism treats the spatial transformation as the message passing process, and the relation between the view poses and the routing weights is modeled by an end-to-end trainable neural network. Besides, an Occupancy Concept Mapping (OCM) framework is proposed to provide explainable rationals for scene-fusion processes. We conducted experiments on several datasets and show that the proposed STR mechanism improves the performance of the Generative Query Network (GQN). The visualization results reveal that the routing process can pass the observed information from one location of some view to the associated location in the other view, which demonstrates the advantage of the proposed model in terms of spatial cognition. Wen-Cheng Chen, Min-Chun Hu 0001, Chu-Song Chen |
ICCV | 3 |
| 2021 | Defect Detection Using Deep Lifelong LearningabstractWith the rapid development of deep learning, automatic defect detection has been introduced into various manufacturing pipelines. Many studies on defect inspection focus on training an accurate model that can perform well on a certain defect type. However, as the manufacturing process evolves, new defect types may appear in practice. The model trained on old defect types will struggle to detect the new ones. To address this issue, we propose to use continual lifelong learning for defect detection. The deep model can increasingly learn to detect new defects yet keeping the learned ones non-forgetting without retraining on the previous data. Our approach can build a compact model, which increasingly learns to detect new defect types. Experimental results show that our approach can learn to detect new defect types incrementally while maintaining its original capability to detect the old defect types. Cheng-Hao Tu 0001, Jia-Da Li, Chu-Song Chen |
INDIN | 4 |
| 2020 | Label Reuse for Efficient Semi-Supervised LearningabstractIn this paper, we propose a new learning strategy for semi-supervised deep learning algorithms, called label reuse, aiming to significantly reduce the expensive computational cost of pseudo label generation and the like for each unlabeled training instance since pseudo labels require to be repeatedly evaluated through the whole training process. For label reuse, we first divide the unlabeled training data into several partitions, replicate each partition in several copies, and place them consecutively in the training queue so as to reuse the pseudo labels computed at first time before invalidation. To evaluate the effectiveness of the proposed approach, we conduct extensive experiments on CIFAR-10 [1] and SVHN [2] by applying it upon the recent state-of-the-art semi-supervised deep learning approach, MixMatch [3]. The results demonstrate the proposed approach can not only significantly reduce the cost of pseudo label computation of MixMatch by a large amount but also keep comparable classification performance. Tsung-Hung Hsieh, Jun-Cheng Chen, Chu-Song Chen |
ICASSP | 3 |
| 2020 | Activity Recognition Using First-Person-View Cameras Based on Sparse Optical FlowsabstractFirst-person-view (FPV) cameras are finding wide use in daily life to record activities and sports. In this paper, we propose a succinct and robust 3D convolutional neural network (CNN) architecture accompanied with an ensemble-learning network for activity recognition with FPV videos. The proposed 3D CNN is trained on low-resolution (32 × 32) sparse optical flows using FPV video datasets consisting of daily activities. According to the experimental results, our network achieves an average accuracy of 90%. Peng Yua Kao, Yan-Jing Lei, Chu-Song Chen, Ming-Sui Lee, Yi-Ping Hung |
ICPR | 4 |
| 2020 | Pruning Depthwise Separable Convolutions for MobileNet CompressionabstractDeep convolutional neural networks are good at accuracy while bad at efficiency. To improve the inference speed, two directions have been explored in the past, lightweight model designing and network weight pruning. Lightweight models have been proposed to improve the speed with good enough accuracy. It is, however, not trivial if we can further speed up these "compact" models by weight pruning. In this paper, we present a technique to gradually prune the depthwise separable convolution networks, such as MobileNet, for improving the speed of this kind of "dense" network. When pruning depthwise separable convolutions, we need to consider more structural constraints to ensure the speedup of inference. Instead of pruning the model with the desired ratio in one stage, the proposed multi-stage gradual pruning approach can stably prune the filters with a finer pruning ratio. Our method achieves satisfiable speedup with little accuracy drop for MobileNets. Code is available at https://github.com/ivclab/Multistage_Pruning. Cheng-Hao Tu 0001, Jia-Hong Lee, Yi-Ming Chan, Chu-Song Chen |
IJCNN | 4 |
| 2020 | Cross-Batch Reference Learning for Deep RetrievalabstractLearning effective representations that exhibit semantic content is crucial to image retrieval applications. Recent advances in deep learning have made significant improvements in performance on a number of visual recognition tasks. Studies have also revealed that visual features extracted from a deep network learned on a large-scale image data set (e.g., ImageNet) for classification are generic and perform well on new recognition tasks in different domains. Nevertheless, when applied to image retrieval, such deep representations do not attain performance as impressive as used for classification. This is mainly because the deep features are optimized for classification rather than for the desired retrieval task. We introduce the cross-batch reference (CBR), a novel training mechanism that enables the optimization of deep networks with a retrieval criterion. With the CBR, the networks leverage both the samples in a single minibatch and the samples in the others for weight updates, enhancing the stochastic gradient descent (SGD) training by enabling interbatch information passing. This interbatch communication is implemented as a cross-batch retrieval process in which the networks are trained to maximize the mean average precision (mAP) that is a popular performance measure in retrieval. Maximizing the cross-batch mAP is equivalent to centralizing the samples relevant to each other in the feature space and separating the samples irrelevant to each other. The learned features can discriminate between relevant and irrelevant samples and thus are suitable for retrieval. To circumvent the discrete, nondifferentiable mAP maximization, we derive an approximate, differentiable lower bound that can be easily optimized in deep networks. Furthermore, the mAP loss can be used alone or with a classification loss. Experiments on several data sets demonstrate that our CBR learning provides favorable performance, validating its effectiveness. Huei-Fang Yang, Ting-Yen Chen, Chu-Song Chen |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2019 | Adaptive Labeling For Hash Code Learning Via Neural NetworksabstractLearning-based hash has been widely used for large-scale similarity retrieval due to the efficient computation and condensed storage of binary representations. In this paper, we propose AdaLabelHash, a hash function learning approach via neural networks. In AdaLabelHash, class label representations are adaptable during the network training. We express the labels as hypercube vertices in a K-dimensional space, and both the network weights and class label representations are updated in the learning process. As the label representations are explored from data, semantically similar categories will be assigned with the label representations that are close to each other in terms of Hamming distance in the label space. The label representations then serve as the desired output of the hash function learning so as to yield compact and discriminating binary hash codes via the network. AdaLabelHash is simple but effective, which can jointly learn label representations and infer compact binary codes from data. It is applicable to both supervised and semi-supervised learning of hash codes. Experimental results on standard benchmarks show the effectiveness of AdaLabelHash. Huei-Fang Yang, Cheng-Hao Tu 0001, Chu-Song Chen |
ICIP | 3 |
| 2019 | Increasingly Packing Multiple Facial-Informatics Modules in A Unified Deep-Learning Model via Lifelong LearningabstractSimultaneously running multiple modules is a key requirement for a smart multimedia system for facial applications including face recognition, facial expression understanding, and gender identification. To effectively integrate them, a continual learning approach to learn new tasks without forgetting is introduced. Unlike previous methods growing monotonically in size, our approach maintains the compactness in continual learning. The proposed packing-and-expanding method is effective and easy to implement, which can iteratively shrink and enlarge the model to integrate new functions. Our integrated multitask model can achieve similar accuracy with only 39.9% of the original size. Steven C. Y. Hung, Jia-Hong Lee, Timmy S. T. Wan, Yi-Ming Chan, Chu-Song Chen |
ICMR | 6 |
| 2019 | IMMVP: An Efficient Daytime and Nighttime On-Road Object DetectorabstractIt is hard to detect on-road objects under various lighting conditions. To improve the quality of the classifier, three techniques are used. We define subclasses to separate daytime and nighttime samples. Then we skip similar samples in the training set to prevent overfitting. With the help of the outside training samples, the detection accuracy is also improved. To detect objects in an edge device, Nvidia Jetson TX2 platform, we exert the lightweight model ResNet-18 FPN as the backbone feature extractor. The FPN (Feature Pyramid Network) generates good features for detecting objects over various scales. With Cascade R-CNN technique, the bounding boxes are iteratively refined for better results. Cheng-En Wu, Yi-Ming Chan, Wen-Cheng Chen, Chu-Song Chen |
MMSP | 5 |
| 2019 | Compacting, Picking and Growing for Unforgetting Continual LearningabstractContinual lifelong learning is essential to many applications. In this paper, we propose a simple but effective approach to continual deep learning. Our approach leverages the principles of deep model compression, critical weights selection, and progressive networks expansion. By enforcing their integration in an iterative manner, we introduce an incremental learning method that is scalable to the number of sequential tasks in a continual learning process. Our approach is easy to implement and owns several favorable characteristics. First, it can avoid forgetting (i.e., learn new tasks while remembering all previous tasks). Second, it allows model expansion but can maintain the model compactness when handling sequential tasks. Besides, through our compaction and selection/expansion mechanism, we show that the knowledge accumulated through learning previous tasks is helpful to build a better model for the new tasks compared to training the models independently with tasks. Experimental results show that our approach can incrementally learn a deep model tackling multiple tasks without forgetting, while the model compactness is maintained with the performance more satisfiable than individual task training. Steven C. Y. Hung, Cheng-Hao Tu 0001, Cheng-En Wu, Yi-Ming Chan, Chu-Song Chen |
NeurIPS | 6 |
| 2019 | Unsupervised Deep Learning of Compact Binary DescriptorsabstractBinary descriptors have been widely used for efficient image matching and retrieval. However, most existing binary descriptors are designed with hand-craft sampling patterns or learned with label annotation provided by datasets. In this paper, we propose a new unsupervised deep learning approach, called DeepBit, to learn compact binary descriptor for efficient visual object matching. We enforce three criteria on binary descriptors which are learned at the top layer of the deep neural network: 1) minimal quantization loss, 2) evenly distributed codes and 3) transformation invariant bit. Then, we estimate the parameters of the network through the optimization of the proposed objectives with a back-propagation technique. Extensive experimental results on various visual recognition tasks demonstrate the effectiveness of the proposed approach. We further demonstrate our proposed approach can be realized on the simplified deep neural network, and enables efficient image matching and retrieval speed with very competitive accuracies. Jiwen Lu, Chu-Song Chen, Jie Zhou 0001, Ming-Ting Sun |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Changing Background to Foreground: An Augmentation Method Based on Conditional Generative Network for Stingray DetectionabstractImage processing has been a popular tool for biological researches. Detecting specific animals in aerial images captured by an UAV is a crucial research topic. As the rapid progress of deep learning (DL), it has been a popular approach to many image classification and object detection tasks. However, DL usually requires a large set of training samples to learn the network weights, while the biological image materials are often insufficient to fulfill the demand. To improve the detection of stingrays in aerial images, this paper presents a new training sample augmentation method called Mixed Bg-Fg Synthesis. We extend a generative network, Generative Latent Optimization (GLO) to its conditional version, namely, Conditional GLO (C-GLO), which can increase stingray samples on the background and thus improve the training efficacy of a CNN detector. Unlike traditional data augmentation methods that generate new data only for image classification, our proposed method that mixes foreground and background together can generate new data for an object detection task. Experimental results show that the C-GLO augmented stingray samples is helpful to enhance the detection capability. Yi-Min Chou, Keng-Hao Liu, Chu-Song Chen |
ICIP | 4 |
| 2018 | A 2.5D Approach to 360 Panorama Video StabilizationabstractThis paper presents a method for stabilizing both cylindrical and spherical panorama videos with a 360-degree field of view. We observe that rotation needs to be extremely smooth for 360 videos to maintain global motion coherency and avoid wobbling. Our method decouples the rotation from other motions and applies different strategies for smoothing them. The proposed approach is 2.5D as it estimates 3D rotations without involving 3D structure-from-motion methods. Therefore, it is more robust and can be performed in an incremental way. Experiments show that our method is effective in making steady 360 cylindrical/spherical videos. Lin-Chen Shen, Tzu-Kuei Huang, Chu-Song Chen, Yung-Yu Chuang |
ICIP | 3 |
| 2018 | Unifying and Merging Well-trained Deep Neural Networks for Inference StageabstractWe propose a novel method to merge convolutional neural-nets for the inference stage. Given two well-trained networks that may have different architectures that handle different tasks, our method aligns the layers of the original networks and merges them into a unified model by sharing the representative codes of weights. The shared weights are further re-trained to fine-tune the performance of the merged model. The proposed method effectively produces a compact model that may run original tasks simultaneously on resource-limited devices. As it preserves the general architectures and leverages the co-used weights of well-trained networks, a substantial training overhead can be reduced to shorten the system development time. Experimental results demonstrate a satisfactory performance and validate the effectiveness of the method. Yi-Min Chou, Yi-Ming Chan, Jia-Hong Lee, Chih-Yi Chiu, Chu-Song Chen |
IJCAI | 5 |
| 2018 | Supervised Learning of Semantics-Preserving Hash via Deep Convolutional Neural NetworksabstractThis paper presents a simple yet effective supervised deep hash approach that constructs binary hash codes from labeled data for large-scale image search. We assume that the semantic labels are governed by several latent attributes with each attribute on or off, and classification relies on these attributes. Based on this assumption, our approach, dubbed supervised semantics-preserving deep hashing (SSDH), constructs hash functions as a latent layer in a deep network and the binary codes are learned by minimizing an objective function defined over classification error and other desirable hash codes properties. With this design, SSDH has a nice characteristic that classification and retrieval are unified in a single learning model. Moreover, SSDH performs joint learning of image representations, hash codes, and classification in a point-wised manner, and thus is scalable to large-scale datasets. SSDH is simple and can be realized by a slight enhancement of an existing deep architecture for classification; yet it is effective and outperforms other hashing approaches on several benchmarks and large datasets. Compared with state-of-the-art approaches, SSDH achieves higher retrieval accuracy, while the classification performance is not sacrificed. Huei-Fang Yang, Chu-Song Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Equivalent Scanning Network of Unpadded CNNsabstractThis letter presents a theory of scanning a signal with a sliding window, where the window's mapping function is built upon a convolutional neural network (CNN). When using a CNN as the sliding window, we show that the resultant feature maps are equivalent to the maps obtained by applying another CNN (called EQ-ScanNet) to the whole signal. The EQ-ScanNet can be established by reconfiguring the original CNN with dilated (i.e., sparse kernel) convolutions. We clarify that, this property is originated from the noble identity (i.e., the swapping equivalence of downsample and FIR filter), and extend the property to the generalized convolution that subsumes CNN's window-sliding operations. We further show that an unpadded CNN is a necessary condition for formulating the EQ-ScanNet. Huei-Fang Yang, Ting-Yen Chen, Cheng-Hao Tu 0001, Chu-Song Chen |
IEEE Signal Process. Lett. | 4 |
| 2018 | Joint Estimation of Age and Expression by Combining Scattering and Convolutional NetworksabstractThis article tackles the problem of joint estimation of human age and facial expression. This is an important yet challenging problem because expressions can alter face appearances in a similar manner to human aging. Different from previous approaches that deal with the two tasks independently, our approach trains a convolutional neural network (CNN) model that unifies ordinal regression and multi-class classification in a single framework. We demonstrate experimentally that our method performs more favorably against state-of-the-art approaches. Huei-Fang Yang, Bo-Yao Lin, Kuang-Yu Chang, Chu-Song Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2017 | Learning and inferring human actions with temporal pyramid features based on conditional random fieldsabstractFinding an effective way to represent human actions is yet an open problem because it usually requires taking evidences extracted from various temporal resolutions into account. A conventional way of representing an action employs temporally ordered fine-grained movements, e.g., key poses or subtle motions. Many existing approaches model actions by directly learning the transitional relationships between those fine-grained features. Yet, an action data may have many similar observations with occasional and irregular changes, which make commonly used fine-grained features less reliable. This paper presents a set of temporal pyramid features that enriches action representation with various levels of semantic granularities. For learning and inferring the proposed pyramid features, we adopt a discriminative model with latent variables to capture the hidden dynamics in each layer of the pyramid. Our method is evaluated on a Tai-Chi Chun dataset and a daily activities dataset. Both of them are collected by us. Experimental results demonstrate that our approach achieves more favorable performance than existing methods. Shih-Yao Lin 0001, Yen-Yu Lin, Chu-Song Chen, Yi-Ping Hung |
ICASSP | 3 |
| 2017 | Aesthetic Critiques Generation for PhotosabstractIt is said that a picture is worth a thousand words. Thus, there are various ways to describe an image, especially in aesthetic quality analysis. Although aesthetic quality assessment has generated a great deal of interest in the last decade, most studies focus on providing a quality rating of good or bad for an image. In this work, we extend the task to produce captions related to photo aesthetics and/or photography skills. To the best of our knowledge, this is the first study that deals with aesthetics captioning instead of AQ scoring. In contrast to common image captioning tasks that depict the objects or their relations in a picture, our approach can select a particular aesthetics aspect and generate captions with respect to the aspect chosen. Meanwhile, the proposed aspect-fusion method further uses an attention mechanism to generate more abundant aesthetics captions. We also introduce a new dataset for aesthetics captioning called the Photo Critique Captioning Dataset (PCCD), which contains pair-wise image-comment data from professional photographers. The results of experiments on PCCD demonstrate that our approaches outperform existing methods for generating aesthetic-oriented captions for images. Kuang-Yu Chang, Kung-Hung Lu, Chu-Song Chen |
ICCV | 3 |
| 2017 | Vision-Based Positioning for Internet-of-VehiclesabstractThis paper presents an algorithm for ego-positioning by using a low-cost monocular camera for systems based on the Internet-of-Vehicles. To reduce the computational and memory requirements, as well as the communication load, we tackle the model compression task as a weighted k-cover problem for better preserving the critical structures. For real-world vision-based positioning applications, we consider the issue of large scene changes and introduce a model update algorithm to address this problem. A large positioning data set containing data collected for more than a month, 106 sessions, and 14275 images is constructed. Extensive experimental results show that submeter accuracy can be achieved by the proposed ego-positioning algorithm, which outperforms existing vision-based approaches. Kuan-Wen Chen, Chun-Hsin Wang, Qiao Liang 0002, Chu-Song Chen, Ming-Hsuan Yang 0001, Yi-Ping Hung |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2017 | Recognizing Human Actions with Outlier Frames by Observation Filtering and CompletionabstractThis article addresses the problem of recognizing partially observed human actions. Videos of actions acquired in the real world often contain corrupt frames caused by various factors. These frames may appear irregularly, and make the actions only partially observed. They change the appearance of actions and degrade the performance of pretrained recognition systems. In this article, we propose an approach to address the corrupt-frame problem without knowing their locations and durations in advance. The proposed approach includes two key components: outlier filtering and observation completion . The former identifies and filters out unobserved frames, and the latter fills up the filtered parts by retrieving coherent alternatives from training data. Hidden Conditional Random Fields (HCRFs) are then used to recognize the filtered and completed actions. Our approach has been evaluated on three datasets, which contain both fully observed actions and partially observed actions with either real or synthetic corrupt frames. The experimental results show that our approach performs favorably against the other state-of-the-art methods, especially when corrupt frames are present. Shih-Yao Lin 0001, Yen-Yu Lin, Chu-Song Chen, Yi-Ping Hung |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2016 | Learning Compact Binary Descriptors with Unsupervised Deep Neural NetworksabstractIn this paper, we propose a new unsupervised deep learning approach called DeepBit to learn compact binary descriptor for efficient visual object matching. Unlike most existing binary descriptors which were designed with random projections or linear hash functions, we develop a deep neural network to learn binary descriptors in an unsupervised manner. We enforce three criterions on binary codes which are learned at the top layer of our network: 1) minimal loss quantization, 2) evenly distributed codes and 3) uncorrelated bits. Then, we learn the parameters of the networks with a back-propagation technique. Experimental results on three different visual analysis tasks including image matching, image retrieval, and object recognition clearly demonstrate the effectiveness of the proposed approach. Jiwen Lu, Chu-Song Chen, Jie Zhou 0001 |
CVPR | 3 |
| 2016 | Style retrieval from natural imagesabstractIt has been a challenging task to identify and distinguish between images of different styles. The challenges mainly come from the extraction of high-level image semantic information, and the presence of the associated ambiguity. In this work, we propose a ranking model for style identification. Given training images of different styles, we learn a pointwise ranking model for each style based on random forests. To handle the high dimensionality of visual features and to prevent against possible ambiguity, we further introduce dimension reduction and pruning techniques for our random forests. In our experiments, we provide quantitative evaluation for style categorization in terms of mean square error (MSE) and relative ranking accuracy. Moreover, our visualization and qualitative results support the use of the proposed method for style retrieval of natural images. Ting-En Tseng, Wei-Yi Chang, Chu-Song Chen, Yu-Chiang Frank Wang |
ICASSP | 3 |
| 2016 | MVC: A Dataset for View-Invariant Clothing Retrieval and Attribute PredictionabstractClothing retrieval and clothing style recognition are important and practical problems. They have drawn a lot of attention in recent years. However, the clothing photos collected in existing datasets are mostly of front- or near-front view. There are no datasets designed to study the influences of different viewing angles on clothing retrieval performance. To address view-invariant clothing retrieval problem properly, we construct a challenge clothing dataset, called Multi-View Clothing dataset. This dataset not only has four different views for each clothing item, but also provides 264 attributes for describing clothing appearance. We adopt a state-of-the-art deep learning method to present baseline results for the attribute prediction and clothing retrieval performance. We also evaluate the method on a more difficult setting, cross-view exact clothing item retrieval. Our dataset will be made publicly available for further studies towards view-invariant clothing retrieval. Kuan-Hsien Liu, Ting-Yen Chen, Chu-Song Chen |
ICMR | 3 |
| 2016 | Cross-batch Reference Learning for Deep Classification and RetrievalabstractLearning feature representations for image retrieval is essential to multimedia search and mining applications. Recently, deep convolutional networks (CNNs) have gained much attention due to their impressive performance on object detection and image classification, and the feature representations learned from a large-scale generic dataset (e.g., ImageNet) can be transferred to or fine-tuned on the datasets of other domains. However, when the feature representations learned with a deep CNN are applied to image retrieval, the performance is still not as good as they are used for classification, which restricts their applicability to relevant image search. To ensure the retrieval capability of the learned feature space, we introduce a new idea called cross-batch reference (CBR) to enhance the stochastic-gradient-descent (SGD) training of CNNs. In each iteration of our training process, the network adjustment relies not only on the training samples in a single batch, but also on the information passed by the samples in the other batches. This inter-batches communication mechanism is formulated as a cross-batch retrieval process based on the mean average precision (MAP) criterion, where the relevant and irrelevant samples are encouraged to be placed on top and rear of the retrieval list, respectively. The learned feature space is not only discriminative to different classes, but the samples that are relevant to each other or of the same class are also enforced to be centralized. To maximize the cross-batch MAP, we design a loss function that is an approximated lower bound of the MAP on the feature layer of the network, which is differentiable and easier for optimization. By combining the intra-batch classification and inter-batch cross-reference losses, the learned features are effective for both classification and retrieval tasks. Experimental results on various benchmarks demonstrate the effectiveness of our approach. Huei-Fang Yang, Chu-Song Chen |
ACM Multimedia | 3 |
| 2015 | Automatic Age Estimation from Face Images via Deep RankingabstractThis paper focuses on automatic age estimation (AAE) from face images, which amounts to determining the exact age or age group of a face image according to features from faces. Although great effort has been devoted to AAE [1, 4, 6], it remains a challenging problem. The difficulties are due to large facial appearance variations resulting from a number of factors, e.g., aging and facial expressions. AAE algorithms need to overcome heterogeneity in facial appearance changes to provide accurate age estimates. To this end, we propose a generic, deep network model for AAE (see Figure 1). Given a face image, our network first extracts features from the face through a 3-layer scattering network (ScatNet) [2], then reduces the feature dimension by principal component analysis (PCA), and finally predicts the age via category-wise rankers constructed as a 3-layer fullyconnected network. The contributions are: (1) Our ranking method is point-wised and thus is easily scaled up to large-scale datasets; (2) our deep ranking model is general and can be applied to age estimation from faces with large facial appearance variations as a result of aging or facial expression changes; and (3) we show that the high-level concepts learned from large-scale neutral faces can be transferred to estimating ages from faces under expression changes, leading to improved performance. Our model is with the following characteristics: (1) The scattering features are invariant to translation and small deformations. ScatNet is a deep convolutional network of specific characteristics. It uses predefined wavelets and computes scattering representations via a cascade of wavelet transforms and modulus pooling operators from shallow to deep layers. With the nonlinear modulus and averaging operators, ScatNet can produce representations that are discriminative as well as invariant to translation and small deformations. As ScatNet provides fundamentally invariant representations for discriminating feature extraction, only the weights of the fully-connected layers are learned in our network model, which considerably reduces the training time. (2) The rank labels encoded in the network exploit the ordering relation among labels. Each category-wise ranker is an ordinal regression ranker. We encode the age rank based on the reduction framework [5]. Given a set of training samples X = {(xi,yi), i = 1 · · ·N}, let xi ∈ RD be the input image and yi be a rank label (yi ∈ {1, . . . ,K}), respectively, where K is the number of age ranks. For rank k, we separate X into two subsets, X k and X − k , as follows: X k = {(xi,+1)|yi > k} X− k = {(xi,−1)|yi ≤ k}. (1) Huei-Fang Yang, Bo-Yao Lin, Kuang-Yu Chang, Chu-Song Chen |
BMVC | 4 |
| 2015 | Location-aware object detection via coherent region groupingabstractWe present a scene adaptation algorithm for object detection. Our method discovers scene-dependent features discriminative to classifying foreground objects into different categories. Unlike previous works suffering from insufficient training data collected online, our approach incorporated with a similarity grouping procedure can automatically gather more consistent training examples from a neighbour area. Experimental results show that the proposed method outperforms several related works with higher detection accuracies. Shen-Chi Chen, Chu-Song Chen, Yi-Ping Hung |
ICASSP | 3 |
| 2015 | Rapid Clothing Retrieval via Deep Learning of Binary Codes and Hierarchical SearchabstractThis paper deals with the problem of clothing retrieval in a recommendation system. We develop a hierarchical deep search framework to tackle this problem. We use a pre-trained network model that has learned rich mid-level visual representations in module 1. Then, in module 2, we add a latent layer to the network and have neurons in this layer to learn hashes-like representations while fine-tuning it on the clothing dataset. Finally, module 3 achieves fast clothing retrieval using the learned hash codes and representations via a coarse-to-fine strategy. We use a large clothing dataset where 161,234 clothes images are collected and labeled. Experiments demonstrate the potential of our proposed framework for clothing retrieval in a large corpus. Huei-Fang Yang, Kuan-Hsien Liu, Jen-Hao Hsiao, Chu-Song Chen |
ICMR | 5 |
| 2015 | Linear Spectral Mixture Analysis via Multiple-Kernel Learning for Hyperspectral Image ClassificationabstractLinear spectral mixture analysis (LSMA) has received wide interests for spectral unmixing in the remote sensing community. This paper introduces a framework called multiplekernel learning-based spectral mixture analysis (MKL-SMA) that integrates a newly proposed MKL method into the training process of LSMA. MKL-SMA allows us to adopt a set of nonlinear basis kernels to better characterize the data so that it can enrich the discriminant capability in classification. Because a single kernel is often insufficient to well present all the data characteristics, MKL-SMA has the advantage of providing a broader range of representation flexibilities; it also eases the kernel selection process because the kernel combination parameters can be learned automatically. Unlike most MKL approaches where complex nonlinear optimization problems are involved in their training process, we derived a closed-form solution of the kernel combination parameters in MKL-SMA. Our method is thus efficient for training and easy to implement. The usefulness of MKL-SMA is demonstrated by conducting real hyperspectral image experiments for performance evaluation. Promising results manifest the effectiveness of the proposed MKL-SMA. Keng-Hao Liu, Yen-Yu Lin, Chu-Song Chen |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2015 | Abandoned Object Detection via Temporal Consistency Modeling and Back-Tracing Verification for Visual SurveillanceabstractThis paper presents an effective approach for detecting abandoned luggage in surveillance videos. We combine short- and long-term background models to extract foreground objects, where each pixel in an input image is classified as a 2-bit code. Subsequently, we introduce a framework to identify static foreground regions based on the temporal transition of code patterns, and to determine whether the candidate regions contain abandoned objects by analyzing the back-traced trajectories of luggage owners. The experimental results obtained based on video images from 2006 Performance Evaluation of Tracking and Surveillance and 2007 Advanced Video and Signal-based Surveillance databases show that the proposed approach is effective for detecting abandoned luggage, and that it outperforms previous methods. Shen-Chi Chen, Chu-Song Chen, Daw-Tung Lin, Yi-Ping Hung |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2015 | A Learning Framework for Age Rank Estimation Based on Face Images With Scattering TransformabstractThis paper presents a cost-sensitive ordinal hyperplanes ranking algorithm for human age estimation based on face images. The proposed approach exploits relative-order information among the age labels for rank prediction. In our approach, the age rank is obtained by aggregating a series of binary classification results, where cost sensitivities among the labels are introduced to improve the aggregating performance. In addition, we give a theoretical analysis on designing the cost of individual binary classifier so that the misranking cost can be bounded by the total misclassification costs. An efficient descriptor, scattering transform, which scatters the Gabor coefficients and pooled with Gaussian smoothing in multiple layers, is evaluated for facial feature extraction. We show that this descriptor is a generalization of conventional bioinspired features and is more effective for face-based age inference. Experimental results demonstrate that our method outperforms the state-of-the-art age estimation approaches. Kuang-Yu Chang, Chu-Song Chen |
IEEE Trans. Image Process. | 2 |
| 2015 | Face Recognition and Retrieval Using Cross-Age Reference Coding With Cross-Age Celebrity DatasetabstractThis paper introduces a method for face recognition across age and also a dataset containing variations of age in the wild. We use a data-driven method to address the cross-age face recognition problem, called cross-age reference coding (CARC). By leveraging a large-scale image dataset freely available on the Internet as a reference set, CARC can encode the low-level feature of a face image with an age-invariant reference space. In the retrieval phase, our method only requires a linear projection to encode the feature and thus it is highly scalable. To evaluate our method, we introduce a large-scale dataset called cross-age celebrity dataset (CACD). The dataset contains more than 160 000 images of 2,000 celebrities with age ranging from 16 to 62. Experimental results show that our method can achieve state-of-the-art performance on both CACD and the other widely used dataset for face recognition across age. To understand the difficulties of face recognition across age, we further construct a verification subset from the CACD called CACD-VS and conduct human evaluation using Amazon Mechanical Turk. CACD-VS contains 2,000 positive pairs and 2,000 negative pairs and is carefully annotated by checking both the associated image and web contents. Our experiments show that although state-of-the-art methods can achieve competitive performance compared to average human performance, majority votes of several humans can achieve much higher performance on this task. The gap between machine and human would imply possible directions for further improvement of cross-age face recognition in the future. Bor-Chun Chen, Chu-Song Chen, Winston H. Hsu |
IEEE Trans. Multim. | 2 |
| 2014 | A spatiotemporal background extractor using a single-layer codebook modelabstractBackground subtraction is a crucial component in visual surveillance, which has been studied over years. However, an efficient algorithm that can tolerate the environment changes such as dynamic backgrounds and sudden changes of illumination is still demanding. In this paper, we design an innovative framework called the spatiotemporal background extractor (SBE) from a single-layer codebook model. Two main extractors, the background extractor (BE) and the background gradient extractor (BGE), are constructed to extract the foreground objects. The background extractor is built for each single frame with spatial information propagated from the neighbor locations, which is useful for handling dynamic background and sudden lighting changes. The background gradient extractor is also constructed and updated, and we design a propagation forbidden policy for background updating, so as to keep the completeness of foreground shape via the background gradient information. The proposed method can efficiently capture the foreground and eliminates the noise of background. The performance of the proposed method is compared with MoG [3], Codebook [4] and ViBe [8] on the Wallflower [1] and Perception [2] datasets. Chih-Wei Lin 0004, Wei-Jie Liao, Chu-Song Chen, Yi-Ping Hung |
AVSS | 3 |
| 2014 | Cross-Age Reference Coding for Age-Invariant Face Recognition and Retrieval
Bor-Chun Chen, Chu-Song Chen, Winston H. Hsu |
ECCV (6) | 2 |
| 2014 | Exploiting low-rank structures from cross-camera images for robust person re-identificationabstractMatching individuals across non-overlapping camera views is known as the problem of person re-identification. In addition to significant visual appearance variations due to lighting, view angle, etc. changes, one might encounter corrupted data due to background clutter and occlusion, or even missing data at some camera views in practical scenarios. To address the above challenges, we present a novel approach to robust person re-identification, particularly aiming at handling missing and corrupted image data across camera views. Based on the technique of low-rank matrix decomposition, our proposed algorithm observes the low-rank structure of cross-view data, which is able to disregard extreme/sparse errors while the missing instances can be recovered automatically. Our experiments will confirm the effectiveness and robustness of our method, which is shown to outperform several baseline and state-of-the-art person re-identification approaches. Ming-Hang Fu, Yu-Chiang Frank Wang, Chu-Song Chen |
ICIP | 3 |
| 2014 | Left-Luggage Detection from Finite-State-Machine Analysis in Static-Camera VideosabstractWe present an abandoned object detection system in this paper. A finite-state-machine model is introduced to extract stationary foregrounds in a scene for visual surveillance, where the state value of each pixel is inferred via the cooperation of short-term and long-term background models constructed in the proposed approach. To identify the left-luggage event, we then verify whether the static foregrounds are abandoned objects through the analysis of owner's moving trajectory back-tracked to the static foreground locations. Experimental results reveal that the proposed approach tackles the problem well on publicly available datasets. Shen-Chi Chen, Chu-Song Chen, Daw-Tung Lin, Yi-Ping Hung |
ICPR | 3 |
| 2014 | Fisher's Discriminant with Natural Image PriorsabstractLinear discriminant analysis that takes spatial smoothness into account has been developed and widely used in image processing society. However, two questions remain unanswered. First, which is the best way to incorporate the smoothness property of images with linear discriminant analysis? Second, which is the best representation for the smoothness property of images? To answer the first question, we propose a Bayesian framework of Gaussian process in order to extend Fisher's discriminant for image data. The probability structure for our extended Fisher's discriminant is explicitly formulated, and the smoothness properties of images are utilized as prior probabilities. For the second question, we suggest a family of prior probabilities derived from natural image statistics. The unknown parameters in our model are estimated via the maximum a posteriori probability (MAP) estimation. We will show that existing methods imposing smoothness assumption of images are rough approximations to the proposed MAP estimates in this framework. Experimental results on the Yale face database and the ETH-80 object categorization dataset show that the proposed method significantly outperforms the other Fisher's discriminant methods for various image data. Yao-Hsiang Yang, Lu-Hung Chen, Chu-Song Chen, Chieh-Chih Wang |
ICPR | 3 |
| 2014 | Salient object detection via local saliency estimation and global homogeneity refinement
Hsin-Ho Yeh, Keng-Hao Liu, Chu-Song Chen |
Pattern Recognit. | 3 |
| 2013 | Target-driven video summarization in a camera networkabstractNowadays, ever expanding camera network makes it difficult to find the suspect from lengthy video records. This paper proposes a target-driven video summarization framework which provides two-step Filtered Summarized Video (FSV) for tracing suspects. Before the target is identified, users can find the target efficiently using the firststep FSV of any arbitrary camera. The first-step FSV filters all the attributes of the target including the time information and the target's categories. After identifying the target, the second-step FSV with additional spatio-temporal & appearance cues are triggered in the neighbor cameras. To enhance the accuracy of the object classification for FSV, we propose a Perspective Dependent Model (PDM) which consists of many grid-based models. Finally, the experimental results show that grid-based model is more robust than general detectors and the user study demonstrates better performance for target finding and tracking in camera network for surveillance. Shen-Chi Chen, Shih-Yao Lin 0001, Kuan-Wen Chen, Chih-Wei Lin 0004, Chu-Song Chen, Yi-Ping Hung |
ICIP | 6 |
| 2013 | Intensity Rank Estimation of Facial Expressions Based on a Single ImageabstractIn this paper, we propose a framework that estimates the discrete intensity rank of a facial expression based on a single image. For most people, judging whether an expression is more intense than others is easier than determining its real-valued intensity degree, and hence the relative order of two expressions is more distinguishable than the exact difference between them. We utilize the relative order to construct an image-based ranking approach for inferring the discrete ranks. The challenge of image-based approaches is to conduct a representation for subtle expression changes. We employ an efficient descriptor, scattering transform, which is translation invariant and can linearize deformations. This scattering representation recovers the lost high frequencies and retains discrimination under invariant property. Our experimental results demonstrate that the proposed framework with scattering transform outperforms other compared feature descriptors and algorithms. Kuang-Yu Chang, Chu-Song Chen, Yi-Ping Hung |
SMC | 2 |
| 2013 | Gender classification from unaligned facial images using support subspaces
Wen-Sheng Chu, Chun-Rong Huang, Chu-Song Chen |
Inf. Sci. | 3 |
| 2013 | Video Aesthetic Quality Assessment by Temporal Integration of Photo- and Motion-Based FeaturesabstractThis paper presents a new method for accessing the aesthetic quality of videos. It consists of two processes: aesthetic features construction and temporal integration. First, our method combines both photo-based and motion-based visual clues to extract the aesthetic features for each frame in a video. We introduce new motion-based features built from optical flow and salient region extraction, and show their effectiveness to enhance the estimation of aesthetic values. Then, a temporal-order-aware framework that integrates the frame-based features is presented to further improve the evaluation accuracy by taking the time-varying properties into consideration. The experimental results demonstrate that our approach can accomplish remarkable improvement for aesthetic quality assessment of videos. Hsin-Ho Yeh, Ming-Sui Lee, Chu-Song Chen |
IEEE Trans. Multim. | 4 |
| 2012 | Affinity aggregation for spectral clusteringabstractSpectral clustering makes use of spectral-graph structure of an affinity matrix to partition data into disjoint meaningful groups. Because of its elegance, efficiency and good performance, spectral clustering has become one of the most popular clustering methods. Traditional spectral clustering assumes a single affinity matrix. However, in many applications, there could be multiple potentially useful features and thereby multiple affinity matrices. To apply spectral clustering for these cases, a possible way is to aggregate the affinity matrices into a single one. Unfortunately, affinity measures constructed from different features could have different characteristics. Careless aggregation might make even worse clustering performance. This paper proposes an affinity aggregation spectral clustering (AASC) algorithm which extends spectral clustering to a setting with multiple affinities available. AASC seeks for an optimal combination of affinity matrices so that it is more immune to ineffective affinities and irrelevant features. This enables the construction of similarity or distance-metric measures for clustering less crucial. Experiments show that AASC is effective in simultaneous clustering and feature fusion, thus enhancing the performance of spectral clustering by employing multiple affinities. Hsin-Chien Huang, Yung-Yu Chuang, Chu-Song Chen |
CVPR | 3 |
| 2012 | Multi-affinity spectral clusteringabstractSpectral clustering (SC) has become one of the most popular clustering methods. Given an affinity matrix, SC explores its spectral-graph structure to partition data into disjoint meaningful groups. However, in many applications, there are multiple potentially useful features and thereby multiple affinity matrices. For applying spectral clustering to such cases, these affinity matrices must be aggregated into a single one. Unfortunately, affinity measures based on different features could have different characteristics. Some are more effective than others. We propose a multi-affinity spectral clustering (MASC) algorithm which extends the SC algorithm with multiple affinities available. By automatically adjusting the weights of affinity matrices, MASC is more immune to ineffective affinities and irrelevant features. This makes the choice of similarity or distance-metric measures for clustering less crucial. Experiments show that MASC is effective in simultaneous clustering and feature fusion, thus maintaining robustness of SC for multi-affinity clustering problems. Hsin-Chien Huang, Yung-Yu Chuang, Chu-Song Chen |
ICASSP | 3 |
| 2012 | From rareness to compactness: Contrast-aware image saliency detectionabstractIn this paper, we present a simple but effective method called Contrast-Aware Saliency (CAS) to detect visual saliency by utilizing two general characteristics: rareness and compactness. In our approach, multiple-salient-spots are used to find initial salient clues, which appear to be rare and unique parts in an image. Then, the salient regions are detected by aggregating the surrounding regions of the spots, which fulfil the compactness nature of salient objects. Experimental results show the proposed CAS performs well in the benchmark dataset. Hsin-Ho Yeh, Chu-Song Chen |
ICIP | 2 |
| 2012 | Applying scattering operators for face recognition: A comparative study
Kuang-Yu Chang, Cheng-Fu Lin, Chu-Song Chen, Yi-Ping Hung |
ICPR | 3 |
| 2012 | Assessment of photo aesthetics with efficiency
Kuo-Yen Lo, Keng-Hao Liu, Chu-Song Chen |
ICPR | 3 |
| 2012 | Multiple Kernel Fuzzy ClusteringabstractWhile fuzzy c-means is a popular soft-clustering method, its effectiveness is largely limited to spherical clusters. By applying kernel tricks, the kernel fuzzy c-means algorithm attempts to address this problem by mapping data with nonlinear relationships to appropriate feature spaces. Kernel combination, or selection, is crucial for effective kernel clustering. Unfortunately, for most applications, it is uneasy to find the right combination. We propose a multiple kernel fuzzy c-means (MKFC) algorithm that extends the fuzzy c-means algorithm with a multiple kernel-learning setting. By incorporating multiple kernels and automatically adjusting the kernel weights, MKFC is more immune to ineffective kernels and irrelevant features. This makes the choice of kernels less crucial. In addition, we show multiple kernel k-means to be a special case of MKFC. Experiments on both synthetic and real-world data demonstrate the effectiveness of the proposed MKFC algorithm. Hsin-Chien Huang, Yung-Yu Chuang, Chu-Song Chen |
IEEE Trans. Fuzzy Syst. | 3 |
| 2012 | Intrinsic Illumination Subspace for Lighting Insensitive Face RecognitionabstractWe introduce the intrinsic illumination subspace and its application for lighting insensitive face recognition in this paper. The intrinsic illumination subspace is constructed from illumination images of intrinsic images, which is a midlevel description of appearance images and can be useful for many visual inferences. This subspace forms a convex polyhedral cone and can be efficiently represented by a low-dimensional linear subspace, which enables an analytic generation of illumination images under varying lighting conditions. When only objects of the same class, such as faces, are concerned, a class-based generic intrinsic illumination subspace can be constructed in advance and used for novel objects of the same class. Based on this class-based generic subspace, we propose a lighting normalization method for lighting insensitive face recognition, where only a single input image is required. The generic subspace is used as a bootstrap subspace for illumination images of novel objects. Face recognition experiments are performed to demonstrate the effectiveness of the proposed lighting normalization method and verify empirically that the class-based generic subspace is applicable to novel objects. Our method is simple and fast, which makes it useful for real-time applications, embedded systems, or mobile devices with limited resources. Chia-Ping Chen, Chu-Song Chen |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2011 | Ordinal hyperplanes ranker with cost sensitivities for age estimationabstractIn this paper, we propose an ordinal hyperplane ranking algorithm called OHRank, which estimates human ages via facial images. The design of the algorithm is based on the relative order information among the age labels in a database. Each ordinal hyperplane separates all the facial images into two groups according to the relative order, and a cost-sensitive property is exploited to find better hyperplanes based on the classification costs. Human ages are inferred by aggregating a set of preferences from the ordinal hyperplanes with their cost sensitivities. Our experimental results demonstrate that the proposed approach outperforms conventional multiclass-based and regression-based approaches as well as recently developed ranking-based age estimation approaches. Kuang-Yu Chang, Chu-Song Chen, Yi-Ping Hung |
CVPR | 2 |
| 2011 | Illumination invariant feature extraction based on natural images statistics - Taking face images as an exampleabstractNatural images are known to carry several distinct properties which are not shared with randomly generated images. In this article we utilize the scale invariant property of natural images to construct a filter which extracts features invariant to illumination conditions. In contrast to most of the existing methods which assume that such features lie in high frequency part of spectrum, by analyzing the power spectra of natural images we show that some of these features could lie in low frequency part as well. From this fact, we derive a Wiener filter approach to best separate the illumination-invariant features from an image. We also provide a linear time algorithm for our proposed Wiener filter, which only involves solving linear equations with narrowly banded matrix. Our experiments on variable lighting face recognition show that our proposed method does achieve the best recognition rate and is generally faster compared to the state-of-the-art methods. Lu-Hung Chen, Yao-Hsiang Yang, Chu-Song Chen, Ming-Yen Cheng |
CVPR | 3 |
| 2011 | Video aesthetic quality assessment by combining semantically independent and dependent featuresabstractThis paper aims to accomplish the work of assessing the aesthetic quality of a video. Unlike previous assessing works focusing mainly on the extraction of aesthetic features in a film, we further study the features, discover their semantic property on videos and then come up with more useful video based features such as motion space and motion direction entropy. In the experiment, we compare the assessing accuracy between two different semantic types of features and find that the semantic-independent feature is more reliable from the results. By combining all features, our method learned a more robust and accurate assessment model. Hsin-Ho Yeh, Chu-Song Chen |
ICASSP | 3 |
| 2011 | A monotonic constrained regression framework for histogram equalization and specificationabstractThis paper introduces a general framework for image contrast enhancement based on histogram equalization (HE) and specification (HS). Traditional HE and HS are simple and effective, but they often amplify the noise level of the image while enhancing it. Furthermore, they may not utilize the entire dynamic range due to the discrete nature of the image. In our framework, image contrast enhancement is posed as a nonparametric monotonic constrained regression problem, in which both the two boundary values and the slopes of the brightness transform function are controlled. We show that such a framework provides an effective way to avoid enlarging the noise level and to utilize the entire dynamic range while performing HS (and also its special case HE). Our method can thus reduce the production of visual artifacts while enhancing the image. Lu-Hung Chen, Yao-Hsiang Yang, Chu-Song Chen |
ICIP | 3 |
| 2011 | Adaptive Learning for Target Tracking and True Linking Discovering Across Multiple Non-Overlapping CamerasabstractTo track targets across networked cameras with disjoint views, one of the major problems is to learn the spatio-temporal relationship and the appearance relationship, where the appearance relationship is usually modeled as a brightness transfer function. Traditional methods learning the relationships by using either hand-labeled correspondence or batch-learning procedure are applicable when the environment remains unchanged. However, in many situations such as lighting changes, the environment varies seriously and hence traditional methods fail to work. In this paper, we propose an unsupervised method which learns adaptively and can be applied to long-term monitoring. Furthermore, we propose a method that can avoid weak links and discover the true valid links among the entry/exit zones of cameras from the correspondence. Experimental results demonstrate that our method outperforms existing methods in learning both the spatio-temporal and the appearance relationship, and can achieve high tracking accuracy in both indoor and outdoor environment. Kuan-Wen Chen, Chih-Chuan Lai, Pei-Jyun Lee, Chu-Song Chen, Yi-Ping Hung |
IEEE Trans. Multim. | 4 |
| 2011 | Temporal Color Consistency-Based Video Reproduction for DichromatsabstractIn this paper, a video re-coloring algorithm for dichromats is presented. Different from image re-coloring schemes, reproducing a video for dichromats requires maintaining temporal color consistency between frames, i.e., the same color in different frames should be re-colored to the identical new color. To achieve this goal, we extract video key colors from shots after motion estimation at first. Based on the importance of video key colors, a process order is defined to perform efficient color remapping and solve the contrast maintaining problem. Then, the remapped frame pixel values are interpolated by the remapped video key colors with spatial-temporal constraints. Experimental results show that our method can increase the visibility for dichromats and guarantee temporal color consistency. Chun-Rong Huang, Kuo-Chuan Chiu, Chu-Song Chen |
IEEE Trans. Multim. | 3 |
| 2010 | MOMI-Cosegmentation: Simultaneous Segmentation of Multiple Objects among Multiple Images
Wen-Sheng Chu, Chia-Ping Chen, Chu-Song Chen |
ACCV (1) | 3 |
| 2010 | Learning Dense Optical-Flow Trajectory Patterns for Video Object ExtractionabstractWe proposes an unsupervised method to address video object extraction (VOE) in uncontrolled videos, i.e. videos captured by low-resolution and freely moving cameras. We advocate the use of dense optical-flow trajectories (DOTs), which are obtained by propagating the optical flow information at the pixel level. Therefore, no interest point extraction is required in our framework. To integrate color and and shape information of moving objects, we group the DOTs at the super-pixel level to extract co-motion regions, and use the associated pyramid histogram of oriented gradients (PHOG) descriptors to extract objects of interest across video frames. Our approach for VOE is easy to implement, and the use of DOTs for both motion segmentation and object tracking is more robust than existing trajectory-based methods. Experiments on several video sequences exhibit the feasibility of our proposed VOE framework. Wang-Chou Lu, Yu-Chiang Frank Wang, Chu-Song Chen |
AVSS | 3 |
| 2010 | An adaptive approach for overlapping people tracking based on foreground silhouettesabstractWe propose Binary/Appearance Tracker which consists of background subtraction, silhouette similarity and particle filter to infer pedestrians' locations under different occlusion situations with a single camera. During the period of occlusions, binary and color silhouettes are adaptively used to effectively measure the similarity between the observation and the possible combinations of silhouettes. Thus, the occluded pedestrians' locations can be simply located by the most possible combination of silhouettes. The experimental results show that the proposed BATracker can track people successfully even though she/he is fully occluded. Hsin-Ho Yeh, Jiun-Yu Chen, Chun-Rong Huang, Chu-Song Chen |
ICIP | 4 |
| 2010 | Turning Rust into Gold: An ancient artifact as an interactive artworkabstractTurning Rust into Gold is inspired by a Chinese antique Mao-Kung Ting (cauldron) treasured by the National Palace Museum in Taiwan. Having a five-hundred-character inscription cast inside, and its weathered appearance made the Mao-Kung very unique. Motivated by revealing the great nature of the artifact and interpreting it into a meaningful narrative, we have proposed an interactive multimedia system that facilitates effective communication between museum audiences and the Mao-Kung Ting. Three technologies have been implemented to emphasize the weathered appearance of the bronze. De-/weathering simulation techniques have been deployed to revive the bronze to its original shiny gold color; while breath-based biofeedback and haptic technology have been utilized as user interfaces to trigger the de-weathering process of the Mao-Kung Ting. Also, the interactive scenarios have been designed with the Chinese cultural context and philosophy Qi, enabling users more easily fall into the Chinese civilization. The paper aims to present the development of the artwork Turing Rust into Gold, in order to further contribute to the feasibility of incorporating new media art in a historical museum context, and bring a new horizon in the museum sector. Chun-Ko Hsieh, Xin Tong 0001, Yi-Ping Hung, Chia-Ping Chen, Liang-Chun Lin, I-Ling Liu, Meng-Chieh Yu, Chu-Song Chen, Jiaping Wang |
ICME | 8 |
| 2010 | A Ranking Approach for Human Ages Estimation Based on Face ImagesabstractIn our daily life, it is much easier to distinguish which person is elder between two persons than how old a person is. When inferring a person's age, we may compare his or her face with many people whose ages are known, resulting in a series of comparative results, and then we conjecture the age based on the comparisons. This process involves numerous pairwise preferences information obtained by a series of queries, where each query compares the target person's face to those faces in a database. In this paper, we propose a ranking-based framework consisting of a set of binary queries. Each query collects a binary-classification-based comparison result. All the query results are then fused to predict the age. Experimental results show that our approach performs better than traditional multi-class-based and regression-based approaches for age estimation. Kuang-Yu Chang, Chu-Song Chen, Yi-Ping Hung |
ICPR | 2 |
| 2010 | Identifying Gender from Unaligned Facial Images by Set ClassificationabstractRough face alignments lead to suboptimal performance of face identification systems. In this study, we present a novel approach for identifying genders from facial images without proper face alignments. Instead of using only one input for test, we generate an image set by randomly cropping out a set of image patches from a neighborhood of the face detection region. Each image set is represented as a subspace and compared with other image sets by measuring the canonical correlation between two associated subspaces. By finding an optimal discriminative transformation for all training subspaces, the proposed approach with unaligned facial images is shown to outperform the state-of-the-art methods with face alignment. Wen-Sheng Chu, Chun-Rong Huang, Chu-Song Chen |
ICPR | 3 |
| 2010 | Transformational Breathing between Present and Past: Virtual Exhibition System of the Mao-Kung Ting
Chun-Ko Hsieh, Xin Tong 0001, Yi-Ping Hung, Chia-Ping Chen, Ju-Chun Ko, Meng-Chieh Yu, Han-Hung Lin, Szu-Wei Wu, Yi-Yu Chung, Liang-Chun Lin, Ming-Sui Lee, Chu-Song Chen, Jiaping Wang, Quo-Ping Lin, I-Ling Liu |
MMM | 12 |
| 2010 | Two-View Motion Segmentation with Model Selection and Outlier Removal by RANSAC-Enhanced Dirichlet Process Mixture Models
Yong-Dian Jian, Chu-Song Chen |
Int. J. Comput. Vis. | 2 |
| 2010 | Fast min-hashing indexing and robust spatio-temporal matching for detecting video copiesabstractThe increase in the number of video copies, both legal and illegal, has become a major problem in the multimedia and Internet era. In this article, we propose a novel method for detecting various video copies in a video sequence. To achieve fast and robust detection, the method fully integrates several components, namely the min-hashing signature to compactly represent a video sequence, a spatio-temporal matching scheme to accurately evaluate video similarity compiled from the spatial and temporal aspects, and some speedup techniques to expedite both min-hashing indexing and spatio-temporal matching. The results of experiments demonstrate that, compared to several baseline methods with different feature descriptors and matching schemes, the proposed method which combines both global and local feature descriptors yields the best performance when encountering a variety of video transformations. The method is very fast, requiring approximately 0.06 seconds to search for copies of a thirty-second video clip in a six-hour video sequence. Chih-Yi Chiu, Hsin-Min Wang, Chu-Song Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2009 | Moving cast shadow detection using physics-based featuresabstractCast shadows induced by moving objects often cause serious problems to many vision applications. We present in this paper an online statistical learning approach to model the background appearance variations under cast shadows. Based on the bi-illuminant (i.e. direct light sources and ambient illumination) dichromatic reflection model, we derive physics-based color features under the assumptions of constant ambient illumination and light sources with common spectral power distributions. We first use one Gaussian mixture model (GMM) to learn the color features, which are constant regardless of the background surfaces or illuminant colors in a scene. Then, we build up one pixel based GMM for each pixel to learn the local shadow features. To overcome the slow convergence rate in the conventional GMM learning, we update the pixel-based GMMs through confidence-rated learning. The proposed method can rapidly learn model parameters in an unsupervised way and adapt to illumination conditions or environment changes. Furthermore, we demonstrate that our method is robust to scenes with few foreground activities and videos captured at low or unsteady frame rates. Jia-Bin Huang 0001, Chu-Song Chen |
CVPR | 2 |
| 2009 | A physical approach to Moving Cast Shadow DetectionabstractThis paper presents a physics-based approach capable of detecting cast shadows in video sequence effectively. We develop a new physical model of cast shadows without making prior assumption of the spectral power distribution (SPD) of the light sources and ambient illumination in the scene. The background appearance variation caused by cast shadows is characterized as the interaction of the blocked light sources and the background surface reflectance. We then take advantage of the statistical prevalence of cast shadows to learn and update the shadow model parameters using the Gaussian mixture model (GMM) over time. The proposed algorithm is completely unsupervised and can adapt to specific environment with complex illumination condition as well as changing shadow conditions. Experimental results on three challenging sequences demonstrate the effectiveness of the proposed method. Jia-Bin Huang 0001, Chu-Song Chen |
ICASSP | 2 |
| 2009 | Image recolorization for the colorblindabstractIn this paper, we propose a new re-coloring algorithm to enhance the accessibility for the color vision deficient (or colorblind). Compared to people with normal color vision, people with color vision deficiency (CVD) have difficulty in distinguishing between certain combinations of colors. This may hinder visual communication owing to the increasing use of colors in recent years. To address this problem, we re-color the image to preserve visual detail when perceived by people with CVD. We first extract the representing colors in an image. Then we find the optimal mapping to maintain the contrast between each pair of these representing colors. The proposed algorithm is image content dependent and completely automatic. Experimental results on natural images are illustrated to demonstrate the effectiveness of the proposed re-coloring algorithm. Jia-Bin Huang 0001, Chu-Song Chen, Tzu-Cheng Jen, Sheng-Jyh Wang |
ICASSP | 2 |
| 2009 | Fast gender recognition by using a shared-integral-image approachabstractWe develop a new approach for gender recognition. In this paper, our approach uses the rectangle feature vector (RFV) as a representation to identify humans' gender from their faces. The RFV is computationally fast and effective to encode intensity variations of local regions of human face. By only using few rectangle features learned by AdaBoost, we present a gender identifier. We then use nonlinear support vector machines for classification, and obtain more accurate identification results. Bau-Cheng Shen, Chu-Song Chen, Hui-Huang Hsu |
ICASSP | 2 |
| 2009 | Video Scene Detection by Link-constrained Affinity-propagationabstractVideo scenes provide semantic meanings for video content description and summarization. This paper explores the pair-wise visual cues of near-duplicate objects for link-constraint affinity-propagation without using keyframes. Experiments demonstrate that our method is more capable to identify scenes comparing with non-constrained clustering algorithms. Chun-Rong Huang, Chu-Song Chen |
ISCAS | 2 |
| 2009 | Tracking by Parts: A Bayesian Approach With Component CollaborationabstractInstead of using global-appearance information for visual tracking, as adopted by many methods, we propose a tracking-by-parts (TBP) approach that uses partial appearance information for the task. The proposed method considers the collaborations between parts and derives a probability propagation framework by encoding the spatial coherence in a Bayesian formulation. To resolve this formulation, a TBP particle-filtering method is introduced. Unlike existing methods that only use the spatial-coherence relationship for particle-weight estimation, our method further applies this relationship for state prediction based on system dynamics. Thus, the part-based information can be utilized efficiently, and the tracking performance can be improved. Experimental results show that our approach outperforms the factored-likelihood and particle reweight methods, which only use spatial coherence for weight estimation. Wen-Yan Chang, Chu-Song Chen, Yi-Ping Hung |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2008 | An adaptive learning method for target tracking across multiple camerasabstractThis paper proposes an adaptive learning method for tracking targets across multiple cameras with disjoint views. Two visual cues are usually employed for tracking targets across cameras: spatio-temporal cue and appearance cue. To learn the relationships among cameras, traditional methods used batch-learning procedures or hand-labeled correspondence, which can work well only within a short period of time. In this paper, we propose an unsupervised method which learns both spatio-temporal relationships and appearance relationships adaptively and can be applied to long-term monitoring. Our method performs target tracking across multiple cameras while also considering the environment changes, such as sudden lighting changes. Also, we improve the estimation of spatio-temporal relationships by using the prior knowledge of camera network topology. Kuan-Wen Chen, Chih-Chuan Lai, Yi-Ping Hung, Chu-Song Chen |
CVPR | 4 |
| 2008 | A Novel Language-Model-Based Approach for Image Object Mining and Re-rankingabstractOne leading framework for image object mining is the bag-of-words (BOW) approach. The idea is to encode an image as a collection of visual words of the quantized local patches. Objects in the image can then be retrieved through inferring the semantic topics associated with the set of visual words. However, the visual BOW mining framework is apt to suffer from the so-called term-mismatch problem (a.k.a. vocabulary problem). This is caused by the poverty of query information, and consequently becomes an obstacle to deal with synonymy (i.e., different visual words for describing the same object). In this paper, we propose a novel language-model-based approach with pseudo-relevance feedback for addressing the vocabulary problem in visual BOW mining. We employ the pseudo positive images produced in response to the original query as a set of "cues" to gradually refine the query language model. Unlike traditional approaches that only ruggedly append feedback information into the original query, the proposed approach reconstructs the query language model with finer granularities so that the query concepts can be captured more accurately. The proposed approach is experimentally evaluated using two different types of image object databases. Our algorithms are shown to bring significant improvement in the retrieval accuracy over a non-feedback baseline, and achieve better performance than conventional feedback approaches. Jen-Hao Hsiao, Chu-Song Chen, Ming-Syan Chen |
ICDM | 2 |
| 2008 | Visual-word-based duplicate image search with pseudo-relevance feedbackabstractWe aim to improve the bag-of-visual-words (BOW) model for near-duplicate image retrieval, by introducing a more fine-grained pseudo-relevance feedback process. The BOW method is based on vector quantization of affine invariant descriptors of image patches. Despite its popularity and simplicity, the retrieval performance of BOW is often unsatisfactory due to the large and diverse variations of near-duplicate images. We thus propose an information-theoretic feedback framework that employs available cues in the search result to find more relevant duplicate images which are hard to retrieve by using conventional BOW approaches. Our algorithm is experimentally evaluated under a severely attacked image database, and shown to significantly improve the retrieval accuracy over a non-feedback baseline. Jen-Hao Hsiao, Chu-Song Chen, Ming-Syan Chen |
ICME | 2 |
| 2008 | Human action recognition using temporal-state shape contextsabstractIn this paper, we present a temporal-state shape context (TSSC) method that exploits space-time shape variations for human action recognition. In our method, the silhouettes of objects in a video clip are organized into three temporal states. These states are defined by fuzzy time intervals, which can lessen the degradation of recognition performance caused by time warping effects. The TSSC features capture local characteristics of the space-time shape induced by consecutive changes of silhouettes. Experimental results show that our method is effective for human action recognition, and is reliable when there are various kinds of deformations. Moreover, our method can identify spatially inconsistent parts between two shapes of the actions, which could be useful in action analysis applications. Pei-Chi Hsiao, Chu-Song Chen, Long-Wen Chang |
ICPR | 2 |
| 2008 | Face image retrieval by using Haar featuresabstractWe propose a new method to retrieve similar face images from large face databases. The proposed method extracts a set of Haar-like features, and integrates these features with supervised manifold learning. Haar-like features are intensity-based features. The values of various Haar-like features comprise our rectangle feature vector (RFV) to describe faces. Compared with several popular unsupervised dimension reduction methods, RFV is more effective in retrieving similar faces. To further improve the performance, we combine RFV and a supervised manifold learning method and obtain satisfactory retrieval results. Bau-Cheng Shen, Chu-Song Chen, Hui-Huang Hsu |
ICPR | 2 |
| 2008 | Contrast context histogram - An efficient discriminating local descriptor for object recognition and image matching
Chun-Rong Huang, Chu-Song Chen, Pau-Choo Chung |
Pattern Recognit. | 2 |
| 2008 | A Framework for Handling Spatiotemporal Variations in Video Copy DetectionabstractAn effective video copy detection framework should be robust against spatial and temporal variations, e.g., changes in brightness and speed. To this end, a content-based approach for video copy detection is proposed. We define the problem as a partial matching problem in a probabilistic model and transform it into a shortest-path problem in a matching graph. To reduce the computation costs of the proposed framework, we introduce some methods that rapidly select key frames and candidate segments from a large amount of video data. The experiment results show that the proposed approach not only handles spatial and temporal variations well, but it also reduces the computation costs substantially. Chih-Yi Chiu, Chu-Song Chen, Lee-Feng Chien |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Visual Tracking in High-Dimensional State Space by Appearance-Guided Particle FilteringabstractIn this paper, we propose a new approach, appearance-guided particle filtering (AGPF), for high degree-of-freedom visual tracking from an image sequence. This method adopts some known attractors in the state space and integrates both appearance and motion-transition information for visual tracking. A probability propagation model based on these two types of information is derived from a Bayesian formulation, and a particle filtering framework is developed to realize it. Experimental results demonstrate that the proposed method is effective for high degree-of-freedom visual tracking problems, such as articulated hand tracking and lip-contour tracking. Wen-Yan Chang, Chu-Song Chen, Yong-Dian Jian |
IEEE Trans. Image Process. | 2 |
| 2008 | Fast Human Detection Using a Novel Boosted Cascading Structure With Meta StagesabstractWe propose a method that can detect humans in a single image based on a novel cascaded structure. In our approach, both intensity-based rectangle features and gradient-based 1-D features are employed in the feature pool for weak-learner selection. The Real AdaBoost algorithm is used to select critical features from a combined feature set and learn the classifiers from the training images for each stage of the cascaded structure. Instead of using the standard boosted cascade, the proposed method employs a novel cascaded structure that exploits both the stage-wise classification information and the interstage cross-reference information. We introduce meta-stages to enhance the detection performance of a boosted cascade. Experiment results show that the proposed approach achieves high detection accuracy and efficiency. Chu-Song Chen |
IEEE Trans. Image Process. | 2 |
| 2008 | Shot Change Detection via Local Keypoint MatchingabstractShot change detection is an essential step in video content analysis. However, automatic shot change detection often suffers from high false detection rates due to camera or object movements. To solve this problem, we propose an approach based on local keypoint matching of video frames. This approach aims to detect both abrupt and gradual transitions between shots without modeling different kinds of transitions. Our experiment results show that the proposed algorithm is effective for most kinds of shot changes. Chun-Rong Huang, Huai-Ping Lee, Chu-Song Chen |
IEEE Trans. Multim. | 3 |
| 2007 | Analyzing Facial Expression by Fusing Manifolds
Wen-Yan Chang, Chu-Song Chen, Yi-Ping Hung |
ACCV (2) | 2 |
| 2007 | A Cascade of Feed-Forward Classifiers for Fast Pedestrian Detection
Chu-Song Chen |
ACCV (1) | 2 |
| 2007 | Two-View Motion Segmentation by Mixtures of Dirichlet Process with Model Selection and Outlier RemovalabstractThis paper presents a novel motion segmentation algorithm on the basis of mixture of Dirichlet process (MDP) models, a kind of nonparametric Bayesian framework. In contrast to previous approaches, our method consider motion segmentation and its model selection regarding to the number of motion models as an indivisible problem. The proposed algorithm can simultaneously infer the number of motion models, estimate the cluster memberships of correspondence points, and identify the outliers of input data. The key idea is to use MDP models to fully exploit the epipolar constraints before making premature decisions about the number motion models. To handle outliers efficiently, we then incorporate RANSAC within the inference process of MDP models and make them take the advantages of each other. In the experiments, we compare the proposed algorithm with naive RANSAC, GPCA and Schindler's method on both synthetic data and real image data. The experimental results show that we can handle more motions and still have satisfactory performance in the presence of various levels of noise and outlier. Yong-Dian Jian, Chu-Song Chen |
ICCV | 2 |
| 2007 | Cascading Multimodal Verification using Face, Voice and Iris InformationabstractIn this paper we propose a novel fusion strategy which fuses information from multiple physical traits via a cascading verification process. In the proposed system users are verified by each individual modules sequentially in turns of face, voice and iris, and would be accepted once he/she is verified by one of the modules without performing the rest of the verifications. Through adjusting thresholds for each module, the proposed approach exhibits different behavior with respect to security and user convenience. We provide a criterion to select thresholds for different requirements and we also design an user interface which helps users find the personalized thresholds intuitively. The proposed approach is verified with experiments on our in-house face-voice-iris database. The experimental results indicate that besides the flexibility between security and convenience, the proposed system also achieves better accuracy than its most accurate module. Ping-Han Lee, Lu-Jong Chu, Yi-Ping Hung, Sheng-Wen Shih, Chu-Song Chen, Hsin-Min Wang |
ICME | 5 |
| 2007 | Efficient and Effective Video Copy Detection Based on Spatiotemporal AnalysisabstractIn this paper, a novel method is presented to detect video copies for a given video query. These copies and the query have identical or near-duplicate content, which might differ in their spatiotemporal structures slightly. To address both the efficient and effective issues, we conduct the bag-of words model for video feature representation, and apply a coarse-to-fine matching scheme to analyze the video spatiotemporal structure. The proposed method can deal with various kinds of video transformations, such as cropping, zooming, speed change, and subsequence insertion/deletion, which are not well addressed in existing methods. Besides, two indexing methods are employed to speed up the matching process. Experimental results show that the proposed method can behave in an efficient and effective manner. Chih-Yi Chiu, Cheng-Chih Yang, Chu-Song Chen |
ISM | 3 |
| 2007 | Efficient hierarchical method for background subtraction
Chu-Song Chen, Chun-Rong Huang, Yi-Ping Hung |
Pattern Recognit. | 2 |
| 2007 | A New Approach to Image Copy Detection Based on Extended Feature SetsabstractConventional image copy detection research concentrates on finding features that are robust enough to resist various kinds of image attacks. However, finding a globally effective fealure is difficult and, in many cases, domain dependent. Instead of imply extracting features from copyrighted images directly, we propose a new framework called the extended feature set for detecting copies of images. In our approach, virtual prior attacks are applied to copyrighted images to generate novel features, which serve as training data. The copy-detection problem can be solved by learning classifiers from the training data, thus, generated. Our approach can be integrated into existing copy detectors to further improve their performance. Experiment results demonstrate that the proposed approach can substantially enhance the accuracy of copy detection. Jen-Hao Hsiao, Chu-Song Chen, Lee-Feng Chien, Ming-Syan Chen |
IEEE Trans. Image Process. | 2 |
| 2006 | A Method for Calibrating a Motorized Object Rig
Pang-Hung Huang, Yu-Pao Tsai, Wan-Yen Lo, Sheng-Wen Shih, Chu-Song Chen, Yi-Ping Hung |
ACCV (1) | 5 |
| 2006 | Attractor-Guided Particle Filtering for Lip Contour Tracking
Yong-Dian Jian, Wen-Yan Chang, Chu-Song Chen |
ACCV (1) | 3 |
| 2006 | The 4-Source Photometric Stereo Under General Unknown Lighting
Chia-Ping Chen, Chu-Song Chen |
ECCV (3) | 2 |
| 2006 | Image Copy Detection via Grouping in Feature Space Based on Virtual Prior AttacksabstractIn the past, many researches on image copy detection focused on finding a feature that is robust enough for various kinds of image attacks. But it is difficult to find a globally effective feature that is appropriate for many situations. In this paper, we introduce a classification framework to this problem, instead of solving the feature-selection problem. In our approach, novel features are generated by applying virtual prior attacks to copyrighted images, and the copy-detection problem is converted to a classification one that is more robust to solve. Our approach can combine existing image copy detectors and further raise their performances. Jen-Hao Hsiao, Chu-Song Chen, Lee-Feng Chien, Ming-Syan Chen |
ICIP | 2 |
| 2006 | Integration of Background Modeling and Object TrackingabstractBackground model and tracking became critical components for many vision-based applications. Typically, background modeling and object tracking are mutually independent in many approaches. In this paper, we adopt a probabilistic framework that uses particle filtering to integrate these two approaches, and the observation model is measured by Bhattacharyya distance. Experimental results and quantitative evaluations show that the proposed integration framework is effective for moving object detection Chu-Song Chen, Yi-Ping Hung |
ICME | 2 |
| 2006 | Image Content Clustering and Summarization for Photo CollectionsabstractRapid growth of digital photography in recent years spurred the need of photo management tools. In this study, we propose an automatic organization framework for photo collections based on image content, so that a novel browsing experience is provided for users. For each photograph, human faces, together with corresponding clothes and nearby regions are located. We extract color histograms of these regions as the image content feature. Then a similarity matrix of a photo collection is generated according to temporal and content features of those photographs. We perform hierarchical clustering based on this matrix, and extract duplicate subjects of a cluster by introducing the contrast context histogram (CCH) technique. The experimental results show that the developed framework provides a promising result for photo management Cheng-Hung Li, Chih-Yi Chiu, Chun-Rong Huang, Chu-Song Chen, Lee-Feng Chien |
ICME | 4 |
| 2006 | Second-Order Belief Propagation and Its Application to Object LocalizationabstractBelief propagation (BP) has been successfully used to solve many computer vision problems, such as stereo matching, object detection and low-level vision. Although the BP algorithm provides an efficient computation framework for general graph inference, the capability of BP is limited by the fact that it only considers the first-order constraints which can simply model the distance relation between two nodes. But in many computer vision problems, this limitation will bring on serious results when the differential or the angular constraints are inherent in the problems. To resolve this limitation, we generalize the BP algorithm to consider the second-order constraints, and integrate it into the particle filtering framework to speed up the computation. In addition, we apply the proposed method to develop a face localization algorithm to demonstrate its effectiveness. Yong-Dian Jian, Chu-Song Chen |
SMC | 2 |
| 2005 | Appearance-Guided Particle Filtering for Articulated Hand TrackingabstractWe propose a model-based tracking method, called appearance-guided particle filtering (AGPF), which integrates both sequential motion transition information and appearance information. A probability propagation model is derived from a Bayesian formulation for this framework, and a sequential Monte Carlo method is introduced for its realization. We apply the proposed method to articulated hand tracking, and show that it performs better than methods that only use either sequential motion transition information or only use appearance information. Wen-Yan Chang, Chu-Song Chen, Yi-Ping Hung |
CVPR (1) | 2 |
| 2005 | Lighting Normalization with Generic Intrinsic Illumination Subspace for Face RecognitionabstractIn this paper, we introduce the concept of intrinsic illumination subspace which is based on the intrinsic images. This intrinsic illumination subspace enables an analytic generation of the illumination images under varying lighting conditions. When objects of the same class are concerned, our method allows a class-based generic intrinsic illumination subspace to be constructed in advance. We propose a lighting normalization method based on the generic intrinsic illumination subspace, which is used as a bootstrap subspace for novel images. Face recognition experiments are performed to demonstrate the effectiveness of our method. Chia-Ping Chen, Chu-Song Chen |
ICCV | 2 |
| 2005 | An Efficient Approach to Multimodal Person Identity Verification by Fusing Face and Voice InformationabstractThis paper presents an effective method to combine speech recognition, speaker verification and face verification for biometric authentication. Our method provides a light-weight enrollment process and an easy-to-use verification interface. A multi-face/single-sentence strategy is used to combine voice and face verification modules, and support vector machine is employed for information fusion. Experimental results show that our method can achieve high verification accuracies Hsien-Ting Cheng, Yi-Hsiang Chao, Shih-Liang Yeh, Chu-Song Chen, Hsin-Min Wang, Yi-Ping Hung |
ICME | 4 |
| 2004 | Using Inter-feature-Line Consistencies for Sequence-Based Object Recognition
Jiun-Hung Chen, Chu-Song Chen |
ECCV (1) | 2 |
| 2004 | Stitching and Reconstruction of Linear-Pushbroom Panoramic Images for Planar Scenes
Chu-Song Chen, Fay Huang |
ECCV (2) | 1 |
| 2004 | Image set compression through minimal-cost prediction structuresabstractWe propose a new scheme for compressing on image set by building its minimal-cost prediction structure. Existing prediction-based video coding methods can be easily extended and incorporated into this scheme to achieve higher compression efficiency. According to this prediction structure, we also develop a progressive transmission approach for interactive object movie (OM) browsing. Chia-Ping Chen, Chu-Song Chen, Kuo-Liang Chung, Hsueh-I Lu, Gregory Y. Tang |
ICIP | 2 |
| 2004 | On Pose Recovery for Generalized Visual SensorsabstractWith the advances in imaging technologies for robot or machine vision, new imaging devices are being developed for robot navigation or image-based rendering. However, to satisfy some design criterion, such as image resolution or viewing ranges, these devices are not necessarily being designed to follow the perspective rule and, thus, the imaging rays may not pass through a common point. Such generalized imaging devices may not be perspective and, therefore, their poses cannot be estimated with traditional techniques. In this paper, we propose a systematic method for pose estimation of such a generalized imaging device. We formulate it as a nonperspective n point (NPnP) problem. The case with exact solutions, n=3, is investigated comprehensively. Approximate solutions can be found for n>3 in a least-squared-error manner by combining an initial-pose-estimation procedure and an orthogonally iterative procedure. This proposed method can be applied not only to nonperspective imaging devices but also perspective ones. Results from experiments show that our approach can solve the NPnP problem accurately. Chu-Song Chen, Wen-Yan Chang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2004 | Object recognition based on image sequences by using inter-feature-line consistencies
Jiun-Hung Chen, Chu-Song Chen |
Pattern Recognit. | 2 |
| 2004 | An improved algorithm for two-image camera self-calibration and Euclidean structure recovery using absolute quadric
Chun-Rong Huang, Chu-Song Chen, Pau-Choo Chung |
Pattern Recognit. | 2 |
| 2004 | Reducing SVM classification time using multiple mirror classifiersabstractWe propose an approach that uses mirror point pairs and a multiple classifier system to reduce the classification time of a support vector machine (SVM). Decisions made with multiple simple classifiers formed from mirror pairs are integrated to approximate the classification rule of a single SVM. A coarse-to-fine approach is developed for selecting a given number of member classifiers. A clustering method, derived from the similarities between classifiers, is used for a coarse selection. A greedy strategy is then used for fine selection of member classifiers. Selected member classifiers are further refined by finding a weighted combination with a perceptron. Experiment results show that our approach can successfully speed up SVM decisions while maintaining comparable classification accuracy. Jiun-Hung Chen, Chu-Song Chen |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2004 | Panoramic appearance-based recognition of video contents using matching graphsabstractThis paper proposes a general scheme for recognizing the contents of a video using a set of panoramas recorded in a database. In essence, a panorama inherently records the appearances of an omni-directional scene from its central point to arbitrary viewing directions and, thus, can serve as a compact representation of an environment. In particular, this paper emphasizes the use of a sequence of successive frames in a video taken with a video camera, instead of a single frame, for visual recognition. The associated recognition task is formulated as a shortest-path searching problem, and a dynamic-programming technique is used to solve it. Experimental results show that our method can effectively recognize a video. Chu-Song Chen, Wen-Teng Hsieh, Jiun-Hung Chen |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2003 | A real-time robust eye tracking system for autostereoscopic displays using stereo camerasabstractAutostereoscopic display systems can provide users a natural 3D visualization environment by projecting stereo video onto the user's eyes. Eye position localization is a central module in this kind of display system when users are allowed to move freely. This paper presents robust 3D eye tracking techniques that can provide accurate eye positions in real time. Technical batteries comprise: (1) robust face detection based on eigenspace method; (2) real-time face tracking; and (3) eye detection in the obtained face region. According to our implementation on a PC with a Pentium IV 1.2 GHz CPU, the frame rate of the eye tracking process can achieve 25 Hz. Chan-Hung Su, Yong-Sheng Chen, Yi-Ping Hung, Chu-Song Chen, Jiun-Hung Chen |
ICRA | 4 |
| 2002 | Compression of 3D Objects with Multistage Color-Depth Panoramic MapsabstractSummary form only given. A new representation method, the multistage color-depth panoramic map (or panomap), is proposed for compressing 3D graphic objects. The idea of the proposed method is to transform a 3D graphic object, including both the shape and color information, into a single image. Existing image compression techniques can then be applied for compressing the panomap structure, which can achieve a highly efficient representation due to its regularity. In our experiments, compressing the color part of a CMP with a lossy method (JPEG) and the depth part with a lossless one (PNG) achieves good reconstruction quality with low bit rates. Chang-Ming Tsai, Wen-Yan Chang, Chu-Song Chen, Gregory Y. Tang |
DCC | 3 |
| 2002 | Pose Estimation for Generalized Imaging Device via Solving Non-Perspective N Point ProblemabstractIn this paper we present a systematic method for pose estimation of such a generalized imaging device. We reformulate it as a non-perspective n point (NPnP) problem. The case with exact solutions, for n=3, is investigated comprehensively. Approximate solutions can also be found for n>3 with our approach in a least-squared-error manner. The proposed method can be used not only to perspective imaging devices, but also non-perspective ones. Chu-Song Chen, Wen-Yan Chang |
ICRA | 1 |
| 2002 | Augmenting panoramas with object movies by generating novel views with disparity-based view morphingabstractAbstract Our goal is to augment a panorama with object movies in a visually 3D‐consistent way. Notice that a panorama is recorded as one single 2D image and an object movie (OM) is composed of a set of 2D images taken around a 3D object. The challenge is how to integrate the above two sources of 2D images in a 3D‐consistent way so that the user can easily manipulate object movies in a panorama. To solve this problem, we adopt a purely image‐based approach that does not have to reconstruct the geometric models of the 3D objects to be inserted in the panorama. A critical issue of this method is how to generate the novel views required for showing an OM in different places of a panorama, and we have proposed a view morphing technique, called t‐DBVM, to solve this problem. Our experiments have shown that this purely image‐based approach can effectively generate visually convincing OM‐augmented panoramas. This method has great potential for many applications that require integration of panoramas and object movies, such as virtual malls, virtual museum, and interior design. Copyright © 2002 John Wiley & Sons, Ltd. Yi-Ping Hung, Chu-Song Chen, Yu-Pao Tsai, Szu-Wei Lin |
Comput. Animat. Virtual Worlds | 2 |
| 2002 | Fuzzy kernel perceptronabstractA new learning method, the fuzzy kernel perceptron (FKP), in which the fuzzy perceptron (FP) and the Mercer kernels are incorporated, is proposed in this paper. The proposed method first maps the input data into a high-dimensional feature space using some implicit mapping functions. Then, the FP is adopted to find a linear separating hyperplane in the high-dimensional feature space. Compared with the FP, the FKP is more suitable for solving the linearly nonseparable problems. In addition, it is also more efficient than the kernel perceptron (KP). Experimental results show that the FKP has better classification performance than FP, KP, and the support vector machine. Jiun-Hung Chen, Chu-Song Chen |
IEEE Trans. Neural Networks | 2 |
| 1999 | New Calibration-free Approach for Augmented Reality based on Parameterized Cuboid StructureabstractA new method called PCS (parameterized cuboid structure) is presented for augmented reality. In particular, our method can insert animated virtual objects into a static scene, with geometric consistency, and also allow the user to interactively position and rotate the virtual objects with respect to a world coordinate system in a physically (or intuitively) meaningful way. Such capability cannot be achieved by using the existing calibration-free methods. To achieve this goal, we develop a new method for estimating camera parameters, which uses a cuboid structure (or more generally, a parallelepiped structure) as the reference object. The reference cuboid structure can be either explicit or implicit-implicit in the sense that the cuboid structure can be inferred by human perception even though it does not appear explicitly in the image. This method can determine the sizes of the cuboid (or parallelepiped) as well as the intrinsic and extrinsic parameters of the camera. To insert a virtual object into a single uncalibrated image, some human interaction is unavoidable. We have implemented an AR authoring system based on the proposed PCS method which provides an auxiliary line and a refinement criterion to assist human interaction. Experimental results have demonstrated that our method can successfully insert virtual objects into both static and dynamic scenes with highly convincing geometric consistency. Chu-Song Chen, Chi-Kuo Yu, Yi-Ping Hung |
ICCV | 1 |
| 1999 | RANSAC-Based DARCES: A New Approach to Fast Automatic Registration of Partially Overlapping Range ImagesabstractIn this paper, we propose a new method, the RANSAC-based DARCES method (data-aligned rigidity-constrained exhaustive search based on random sample consensus), which can solve the partially overlapping 3D registration problem without any initial estimation. For the noiseless case, the basic algorithm of our method can guarantee that the solution it finds is the true one, and its time complexity can be shown to be relatively low. An extra characteristic is that our method can be used even for the case that there are no local features in the 3D data sets. Chu-Song Chen, Yi-Ping Hung, Jen-Bo Cheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1998 | A Fast and Robust Approach for Registration of Partially Overlapping Range ImagesabstractA popular approach for 3D registration of partially-overlapping range images is the ICP (iterative closest point) method and many of its variations. The major drawback of this type of iterative approaches is that they require a good initial estimate to guarantee that the correct solution can always be found. In this paper, we propose a new method, the RANSAC-based DARCES (data-aligned rigidity-constrained exhaustive search) method, which can solve the partially-overlapping 3D registration problem efficiently and reliably without any initial estimation. Another important characteristic of our method is that it requires no local features in the 3D data set. An extra characteristic is that, for the noiseless case, the basic algorithm of our DARCES method can guarantee that the solution it finds is the true one, due to its exhaustive-search nature. Even with the nature of exhaustive search, its time complexity can be shown to be relatively low. Experiments have demonstrated that our method is efficient and reliable for registering partially-overlapping range images. Chu-Song Chen, Yi-Ping Hung, Jen-Bo Cheng |
ICCV | 1 |
| 1998 | Integrating virtual objects into real images for augmented realityabstractA systematic approach for integrating virtual objects into real images is developed in this paper.We propose the P3P-IC.method to solve the camera pose estimation problem.A robust tracking method is developed via the combination of the LMedS technique and the P3P-ICP method.With the 3D models and the robust tracking methods, we can determine the camera poses associated with each frame in the image sequence.Knowing the camera poses for each image frame, we can then integrate virtual objects into a video segment. 1.1 Chu-Song Chen, Yi-Ping Hung, Sheng-Wen Shih, Chen-Chiung Hsieh, Chen-Yuan Tang, Chih-Guo Yu, You-Chung Cheng |
VRST | 1 |
| 1998 | Multipass hierarchical stereo matching for generation of digital terrain models from aerial images
Yi-Ping Hung, Chu-Song Chen, Kuan-Chung Hung, Yong-Sheng Chen, Chiou-Shann Fuh |
Mach. Vis. Appl. | 2 |
| 1997 | Range data acquisition using color structured lighting and stereo vision
Chu-Song Chen, Yi-Ping Hung, Chiann-Chu Chiang, Ja-Ling Wu |
Image Vis. Comput. | 1 |
| 1996 | Model-based object recognition using range images by combining morphological feature extraction and geometric hashingabstractThis paper proposes a new approach for model-based object recognition with range images by combining morphological feature extraction and geometric hashing. In low-level processing, range images are segmented into 3D-connected surface patches. In middle-level processing: each connected component is processed by using morphological operations to extract the skeletons of high-variation regions. These skeleton points can be viewed as invariant salient feature primitives. In high-level processing, geometric hashing is used to recognize objects. To reduce the number of spurious hypotheses, we propose a basis-similarity constraint. Experimental results have shown that the proposed method is effective and has great potential for model-based object recognition using range images. Chu-Song Chen, Yi-Ping Hung, Ja-Ling Wu |
ICPR | 1 |