Chun-Rong Huang

dblp:32/828 · DBLP profile ↗
← Back
41ranked-venue papers
14as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 6 first-author · 4 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 5 since 2021Systems, architecture and hardware · 4 · 3 first-authorSoftware engineering, systems software and programming languages · 2Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Helicobacter Pylori Infection Diagnosis from Endoscopic Images via Multi-Class Token based Multiple Instance Learning
abstract
To diagnose Helicobacter pylori (H. pylori) infection from endoscopic images, conventional methods require a time-consuming labeling process to annotate individual endoscopic images based on pathological findings. In this paper, we aim to diagnose H. pylori infection for each patient from a group of endoscopic images captured during endoscopy, where only patient-level labels are available for each patient. To achieve the goal, we propose a multi-class token-based multiple instance learning method which consists of the feature extractor, the multi-class token selector module, and the aggregator. By using learnable positive class tokens and negative class tokens with transformer encoders, the multi-class token selector module selects the proper class tokens to improve the H. pylori infection prediction performance of the aggregator. Compared with supervised methods and state-of-the-art multiple instance learning methods, the proposed method achieves the best results.
Wei-Te Ting, Yu-Ching Tsai, Chun-Rong Huang, Hsiu-Chi Cheng, Bor-Shyang Sheu
CIBCB3
2025 Restore Anything Anywhere: Targeted Image Restoration with Object Segmentation and Text Guidance
abstract
Image restoration techniques are widely used in various fields, such as autonomous driving, medical imaging, and satellite imagery. These techniques typically aim to restore an entire degraded image to its original state. However, in many cases, users may wish to focus on restoring specific areas to achieve desired effects in the image, tailored to their preferences. In this work, we propose Restore Anything anyWhere (RAW), a framework that enables users to restore specific types of degradation on any selected object in an image by specifying point or text prompts. Our framework first employs the Multimodal Segmentation module to generate mask priors for the target objects. Then, the text-guided restoration model performs targeted restoration on the areas defined by the object mask priors. To improve segmentation performance, we propose a Text-To-Segmentation Refinement method that combines existing techniques with the capabilities of Segment Anything, enhancing text-to-segmentation accuracy. Additionally, we introduce a streamlined text-guided restoration method to provide better control over degradation restoration, delivering highly effective results. RAW provides excellent object-level control in restoration and significantly improves the Laion Aesthetics score.
Yen-Ku Yeh, Chun-Hao Yang, Kun-Tai Wu, Yan-Tsung Peng, Chun-Rong Huang, Jun-Cheng Chen
MMSP5
2025 Cross-Scale Guidance Network for Few-Shot Moving Foreground Object Segmentation
abstract
Foreground object segmentation is one of the most important pre-processing steps in intelligent transportation and video surveillance systems. Although background modeling methods are efficient to segment foreground objects, their results are easily affected by dynamic backgrounds and updating strategies. Recently, deep learning-based methods have achieved more effective foreground object segmentation results compared with background modeling methods. However, a large number of labeled training frames are usually required. To reduce the number of training frames, we propose a novel cross-scale guidance network (CSGNet) for few-shot moving foreground object segmentation in surveillance videos. The proposed CSGNet contains the cross-scale feature expansion encoder and cross-scale feature guidance decoder. The encoder aims to represent the scenes by extracting cross-scale expansion features based on cross-scale and multiple field-of-view information learned from a limited number of training frames. The decoder aims to obtain accurate foreground object segmentation results under the guidance of the encoder features and the foreground loss. The proposed method outperforms the state-of-the-art background modeling methods and the deep learning-based methods around 2.6% and 3.1%, and the average computation time is 0.073 and 0.046 seconds for each frame in the CDNet2014 dataset and the UCSD dataset under a single GTX 1080 GPU computer. The source code will be available at https://github.com/nchucvml/CSGNet.
Yi-Sheng Liao, Yen-Wei Lin, Ya-Han Chang, Chun-Rong Huang
IEEE Trans. Intell. Transp. Syst.4
2024 Mask Focal Modulation Network for Gastric Intestinal Metaplasia Segmentation
abstract
Gastric intestinal metaplasia (IM) is an important precancerous lesion with a high correlation to gastric cancer. Early gastric IM diagnosis is very important in reducing the risk of gastric cancer. In this paper, we propose the mask focal modulation network to achieve gastric IM segmentation from endoscopic images. The proposed network consists of the encoder module, pixel decoder module, and mask focal modulation decoder module. With the cooperation of the encoder module and pixel decoder module, the proposed mask focal modulation decoder module in the network aims to aggregate features of different field-of-views to make the network focus on the important views with learnable masks and then provide better gastric IM segmentation results. The experimental results show the superiority of the proposed method over state-of-the-art semantic segmentation methods in the gastric IM segmentation dataset. The source code is available at https://github.com/nchucvml/MFMNetwork.
Er-Hsiang Yang, Wei-Lun Chang, Hsiu-Chi Cheng, Chun-Rong Huang
IJCNN6
2024 Cross-Scale Fusion Transformer for Histopathological Image Classification
abstract
Histopathological images provide the medical evidences to help the disease diagnosis. However, pathologists are not always available or are overloaded by work. Moreover, the variations of pathological images with respect to different organs, cell sizes and magnification factors lead to the difficulty of developing a general method to solve the histopathological image classification problems. To address these issues, we propose a novel cross-scale fusion (CSF) transformer which consists of the multiple field-of-view patch embedding module, the transformer encoders and the cross-fusion modules. Based on the proposed modules, the CSF transformer can effectively integrate patch embeddings of different field-of-views to learn cross-scale contextual correlations, which represent tissues and cells of different sizes and magnification factors, with less memory usage and computation compared with the state-of-the-art transformers. To verify the generalization ability of the CSF transformer, experiments are performed on four public datasets of different organs and magnification factors. The CSF transformer outperforms the state-of-the-art task specific methods, convolutional neural network-based methods and transformer-based methods. The source code will be available in our GitHub https://github.com/nchucvml/CSFT.
Sheng-Kai Huang, Yu-Ting Yu, Chun-Rong Huang, Hsiu-Chi Cheng
IEEE J. Biomed. Health Informatics3
2024 A Model-Tuned Predictive Backstepping Control Approach for Angle Following of Steer-by-Wire
abstract
An important function of the intelligent steer-by-wire requires a desired steering angle to be followed accurately. In this article, an algorithm-hybrid control is designed to realize the angle following in an electric motor steer-by-wire system, which is named as a model-tuned predictive backstepping control that consists of model tuning control, backstepping control and model predictive control. The model tuning control is used to adjust the variable modelling-term that includes the self-aligning torque, which means that some posteriori knowledge of the control system is utilized to make the model more accurate. The model predictive control is adopted to compute the variable stepping-parameters of the backstepping control, which means that some priori knowledge of the control system is utilized to optimize the control performance. Then we discuss a series of studies on the steer-by-wire system and the control algorithms that, collectively, develop an approach of how the hybrid control algorithm steers the front wheels based on a desired angle. The designed approach has been deployed into a steering control unit, and tested in a steering test vehicle to realize the angle following of electric motor steer-by-wire system. According to experimental results and statistics analyses, it can be concluded that the model-tuned predictive backstepping control is good candidate for the angle following control of steer-by-wire system.
Lin He 0011, Chun-Rong Huang, Chaolu Guo, Qin Shi 0003
IEEE Trans. Intell. Transp. Syst.2
2023 ADMM-SRNet: Alternating Direction Method of Multipliers Based Sparse Representation Network for One-Class Classification
abstract
One-class classification aims to learn one-class models from only in-class training samples. Because of lacking out-of-class samples during training, most conventional deep learning based methods suffer from the feature collapse problem. In contrast, contrastive learning based methods can learn features from only in-class samples but are hard to be end-to-end trained with one-class models. To address the aforementioned problems, we propose alternating direction method of multipliers based sparse representation network (ADMM-SRNet). ADMM-SRNet contains the heterogeneous contrastive feature (HCF) network and the sparse dictionary (SD) network. The HCF network learns in-class heterogeneous contrastive features by using contrastive learning with heterogeneous augmentations. Then, the SD network models the distributions of the in-class training samples by using dictionaries computed based on ADMM. By coupling the HCF network, SD network and the proposed loss functions, our method can effectively learn discriminative features and one-class models of the in-class training samples in an end-to-end trainable manner. Experimental results show that the proposed method outperforms state-of-the-art methods on CIFAR-10, CIFAR-100 and ImageNet-30 datasets under one-class classification settings. Code is available at https://github.com/nchucvml/ADMM-SRNet.
Chien-Yu Chiou, Kuang-Ting Lee, Chun-Rong Huang, Pau-Choo Chung
IEEE Trans. Image Process.3
2023 Video Summarization With Spatiotemporal Vision Transformer
abstract
Video summarization aims to generate a compact summary of the original video for efficient video browsing. To provide video summaries which are consistent with the human perception and contain important content, supervised learning-based video summarization methods are proposed. These methods aim to learn important content based on continuous frame information of human-created summaries. However, simultaneously considering both of inter-frame correlations among non-adjacent frames and intra-frame attention which attracts the humans for frame importance representations are rarely discussed in recent methods. To address these issues, we propose a novel transformer-based method named spatiotemporal vision transformer (STVT) for video summarization. The STVT is composed of three dominant components including the embedded sequence module, temporal inter-frame attention (TIA) encoder, and spatial intra-frame attention (SIA) encoder. The embedded sequence module generates the embedded sequence by fusing the frame embedding, index embedding and segment class embedding to represent the frames. The temporal inter-frame correlations among non-adjacent frames are learned by the TIA encoder with the multi-head self-attention scheme. Then, the spatial intra-frame attention of each frame is learned by the SIA encoder. Finally, a multi-frame loss is computed to drive the learning of the network in an end-to-end trainable manner. By simultaneously using both inter-frame and intra-frame information, our method outperforms state-of-the-art methods in both of the SumMe and TVSum datasets. The source code of the spatiotemporal vision transformer will be available at https://github.com/nchucvml/STVT.
Tzu-Chun Hsu, Yi-Sheng Liao, Chun-Rong Huang
IEEE Trans. Image Process.3
2022 Semantic Context-Aware Image Style Transfer
abstract
To provide semantic image style transfer results which are consistent with human perception, transferring styles of semantic regions of the style image to their corresponding semantic regions of the content image is necessary. However, when the object categories between the content and style images are not the same, it is difficult to match semantic regions between two images for semantic image style transfer. To solve the semantic matching problem and guide the semantic image style transfer based on matched regions, we propose a novel semantic context-aware image style transfer method by performing semantic context matching followed by a hierarchical local-to-global network architecture. The semantic context matching aims to obtain the corresponding regions between the content and style images by using context correlations of different object categories. Based on the matching results, we retrieve semantic context pairs where each pair is composed of two semantically matched regions from the content and style images. To achieve semantic context-aware style transfer, a hierarchical local-to-global network architecture, which contains two sub-networks including the local context network and the global context network, is proposed. The former focuses on style transfer for each semantic context pair from the style image to the content image, and generates a local style transfer image storing the detailed style feature representations for corresponding semantic regions. The latter aims to derive the stylized image by considering the content, the style, and the intermediate local style transfer images, so that inconsistency between different corresponding semantic regions can be addressed and solved. The experimental results show that the stylized results using our method are more consistent with human perception compared with the state-of-the-art methods.
Yi-Sheng Liao, Chun-Rong Huang
IEEE Trans. Image Process.2
2022 A Content-Adaptive Resizing Framework for Boosting Computation Speed of Background Modeling Methods
abstract
Recently, most background modeling (BM) methods claim to achieve real-time efficiency for low-resolution and standard-definition surveillance videos. With the increasing resolutions of surveillance cameras, full high-definition (full HD) surveillance videos have become the main trend and thus processing high-resolution videos becomes a novel issue in intelligent video surveillance. In this article, we propose a novel content-adaptive resizing framework (CARF) to boost the computation speed of BM methods in high-resolution surveillance videos. For each frame, we apply superpixels to separate the content of the frame to homogeneous and boundary sets. Two novel downsampling and upsampling layers based on the homogeneous and boundary sets are proposed. The front one downsamples high-resolution frames to low-resolution frames for obtaining efficient foreground segmentation results based on BM methods. The later one upsamples the low-resolution foreground segmentation results to the original resolution frames based on the superpixels. By simultaneously coupling both layers, experimental results show that the proposed method can achieve better quantitative and qualitative results compared with state-of-the-art methods. Moreover, the computation speed of the proposed method without GPU accelerations is also significantly faster than that of the state-of-the-art methods. The source code is available athttps://github.com/nchucvml/CARF.
Chun-Rong Huang, Wei-Yun Huang, Yi-Sheng Liao, Chien-Cheng Lee, Yu-Wei Yeh
IEEE Trans. Syst. Man Cybern. Syst.1
2021 Deep Ensemble Feature Network for Gastric Section Classification
abstract
In this paper, we propose a novel deep ensemble feature (DEF) network to classify gastric sections from endoscopic images. Different from recent deep ensemble learning methods, which need to train deep features and classifiers individually to obtain fused classification results, the proposed method can simultaneously learn the deep ensemble feature from arbitrary number of convolutional neural networks (CNNs) and the decision classifier in an end-to-end trainable manner. It comprises two sub networks, the ensemble feature network and the decision network. The former sub network learns the deep ensemble feature from multiple CNNs to represent endoscopic images. The latter sub network learns to obtain the classification labels by using the deep ensemble feature. Both sub networks are optimized based on the proposed ensemble feature loss and the decision loss which guide the learning of deep features and decisions. As shown in the experimental results, the proposed method outperforms the state-of-the-art deep learning, ensemble learning, and deep ensemble learning methods.
Ting-Hsuan Lin, Jyun-Yao Jhang, Chun-Rong Huang, Yu-Ching Tsai, Hsiu-Chi Cheng, Bor-Shyang Sheu
IEEE J. Biomed. Health Informatics3
2020 Driver Monitoring Using Sparse Representation With Part-Based Temporal Face Descriptors
abstract
Many driver monitoring systems (DMSs) have been proposed to reduce the risk of human-caused accidents. Traditional DMSs focus on detecting specific predefined abnormal driving behaviors, such as drowsiness or distracted driving, using generic models trained with the data collected during abnormal driving. However, it is difficult to collect sufficient representative training data to construct generic detection models, which are applicable to all drivers. Consequently, this paper proposes a new personal-based hierarchical DMS (HDMS). During driving, the first layer of the proposed HDMS detects normal and abnormal driving behavior based on normal personal driving models represented by sparse representations. When abnormal driving behavior is detected, the second layer of the HDMS further determines whether the behavior is drowsy driving behavior or distracted driving behavior. The experimental results obtained for three datasets show that the proposed HDMS outperforms existing state-of-the-art DMS methods in detecting normal driving behavior, drowsy driving behavior, and distracted driving behavior.
Chien-Yu Chiou, Wei-Cheng Wang, Shueh-Chou Lu, Chun-Rong Huang, Pau-Choo Chung, Yun-Yang Lai
IEEE Trans. Intell. Transp. Syst.4
2019 Spatial Face Context with Age and Background Cues for Group Photo Retrieval
abstract
People usually take group photos by using cameras and smart phones to record social events. With the increasing number of group photos, searching and collecting similar group photos become difficult. In this paper, we propose using spatial face context with age and background information to achieve group photo retrieval. Visual features between group photos are assessed by applying spatial arrangement, genders, ages and backgrounds information between two group photos. Then, an effective graph matching is applied to assess the similarity between two group photos based on the visual features. Experimental results show that our method can retrieve similar group photos by combing foreground and background information.
Ting-Hsuan Lin, Tien-Ying Li, Chun-Rong Huang
SNPD3
2018 Clustering Trajectories in Heterogeneous Representations for Video Event Detection
abstract
Trajectories have been shown to be robust and widely used in surveillance video event analysis. They encode spatial and temporal evidence simultaneously. Hence, clustering trajectories in a video can detect representative events. How to effectively represent trajectories is thus essential to video event detection. However, no a single representation of trajectories suffices in increasingly complex video analysis tasks. To address this issue, this paper presents a hierarchical clustering algorithm for grouping trajectories in multiple heterogeneous representations. It turns out that our method can not only group trajectories of highly similar events but also identify rare events from the dominant events. Experimental results show that our method can retrieve both dominant events and rare events compared with the state-of-the-art methods, leading to a better performance.
Wei-Cheng Wang, Yen-Yu Lin, Hsin-Wei Cheng, Chun-Rong Huang
ICIP4
2018 Spatiotemporal Coherence-Based Annotation Placement for Surveillance Videos
abstract
In this paper, we propose a novel annotation placement approach for revealing information about foreground objects in surveillance videos. To arrange positions of annotations, spatiotemporal coherence between annotations and foreground objects is applied. The annotation placement problem is formulated as an optimization problem with respect to spatiotemporal coherence of annotations and foreground objects. The optimization problem is effectively solved using Markov random fields. To the best of our knowledge, this paper is the first work that discusses and solves the annotation placement problem for surveillance videos by considering the relationships between annotations and foreground objects with trajectories. As shown in the experiments, the proposed approach can arrange annotations based on the moving trajectories of foreground objects and prevent the occlusions between different annotations and foreground objects. It also achieves better quantitative and qualitative results compared with state-of-the-art approaches.
Wei-Cheng Wang, Chien-Yu Chiou, Chun-Rong Huang, Pau-Choo Chung, Wei-Yun Huang
IEEE Trans. Circuits Syst. Video Technol.3
2018 USEAQ: Ultra-Fast Superpixel Extraction via Adaptive Sampling From Quantized Regions
abstract
We present a novel and highly efficient superpixel extraction method called USEAQ to generate regular and compact superpixels in an image. To reduce the computational cost of iterative optimization procedures adopted in most recent approaches, the proposed USEAQ for superpixel generation works in a one-pass fashion. It firstly performs joint spatial and color quantizations and groups pixels into regions. It then takes into account the variations between regions, and adaptively samples one or a few superpixel candidates for each region. It finally employs maximum a posteriori (MAP) estimation to assign pixels to the most spatially consistent and perceptually similar superpixels. It turns out that the proposed USEAQ is quite efficient, and the extracted superpixels can precisely adhere to boundaries of objects. Experimental results show that USEAQ achieves better or equivalent performance compared to the stateof- the-art superpixel extraction approaches in terms of boundary recall, undersegmentation error, achievable segmentation accuracy, the average miss rate, average undersegmentation error, and average unexplained variation, and it is significantly faster than these approaches.
Chun-Rong Huang, Wei-Cheng Wang, Wei-An Wang, Szu-Yu Lin, Yen-Yu Lin
IEEE Trans. Image Process.1
2016 USEQ: Ultra-fast superpixel extraction via quantization
abstract
We propose a novel superpixel extraction method named USEQ to generate regular and compact superpixels. To reduce the computational burden of iterative optimization procedures used in most recent approaches, the spatial and color quantizations are performed in advance to represent pixels and superpixels. Maximum a posteriori estimation in both pixel and region levels is then adopted to aggregate pixels into spatially and visually coherent superpixels. The resultant superpixels are extremely efficient to generate and can more precisely adhere to object boundaries. Compared to the state-of-the-art approaches to superpixel extraction, USEQ can achieve better or competitive performance in terms of boundary recall, undersegmentation error and achievable segmentation accuracy, and is significantly faster than these approaches.
Chun-Rong Huang, Wei-An Wang, Szu-Yu Lin, Yen-Yu Lin
ICPR1
2015 Trajectory kinematics descriptor for trajectory clustering in surveillance videos
abstract
Trajectories provide spatial-temporal information of foreground objects for event clustering and analysis. Because of the kinematic properties of foreground objects, the lengths of trajectories will be different which lead to the length problem of assessing similarity between two or more trajectories. To solve the problem, we propose a novel descriptor named trajectory kinematics descriptor to represent trajectories based on the kinematic properties from the point-of-view of Frenet-Serret frames. As shown in the experiments, applying the proposed trajectory kinematics descriptor for event clustering can achieve better F-measure scores compared to the state-of-the-art methods.
Wei-Cheng Wang, Pau-Choo Chung, Hsin-Wei Cheng, Chun-Rong Huang
ISCAS4
2015 Video gender recognition using temporal coherent face descriptor
abstract
In this paper, we propose a new temporal coherent face descriptor for video gender recognition. The proposed face descriptor is constructed from detected faces of continuous video frames. Because it describes detected faces under variant changes in continuous video frames and provides a unified feature description, face normalization and alignment processes can be avoided during gender recognition. Based on the face descriptor, a support vector machine classifier is applied to identify the gender of the subject in the videos. As shown in the experiments, our method not only achieves better results compared to the state-of-the-art methods but also the real-time performance for video processing.
Wei-Cheng Wang, Ru-Yun Hsu, Chun-Rong Huang, Li-You Syu
SNPD3
2015 Binary Descriptor Based Nonparametric Background Modeling for Foreground Extraction by Using Detection Theory
abstract
Recently, most background modeling approaches represent distributions of background changes by using parametric models such as Gaussian mixture models. Because of significant illumination changes and dynamic moving backgrounds with time, variations of background changes are hard to be modeled by parametric background models. Moreover, how to efficiently and effectively update parameters of parametric models to reflect background changes remains a problem. In this paper, we propose a novel coarse-to-fine detection theory algorithm to extract foreground objects on the basis of nonparametric background and foreground models represented by binary descriptors. We update background and foreground models by a first-in-first-out strategy to maintain the most recent observed background and foreground instances. As shown in the experiments, our method can achieve better foreground extraction results and fewer false alarms of surveillance videos with lighting changes and dynamic backgrounds in both collected and CDnet 2012 benchmark data sets.
Min-Hsiang Yang, Chun-Rong Huang, Wan-Chen Liu, Shu-Zhe Lin, Kun-Ta Chuang
IEEE Trans. Circuits Syst. Video Technol.2
2014 Spatial Face Context with Gender Information for Group Photo Similarity Assessment
abstract
Spatial arrangements of people in a photo reflect how people regard themselves in a group. Comparing spatial arrangements of subjects in two group photos can help computer understand similar social semantics and events. In this paper, we incorporate gender information with spatial arrangements of subjects to assess the visual similarity between two group photos. For each group photo, detected faces are represented by a spatial face context. Then, corresponding vertices of spatial face contexts of two group photos are retrieved by graph matching. Gender information and relative face positions are applied to the graph matching results to compute the similarity score. The experimental results show that our method can identify group photos with similar genders of subjects and spatial arrangements compared to the state-of-the-art methods.
Yi-I Chiu, Ru-Yun Hsu, Chun-Rong Huang
ICPR3
2014 Occluded object tracking based on trajectory links in surveillance videos
abstract
Trajectories of foreground objects provide rich and continuous motion information for event analysis in surveillance videos. These trajectories can be obtained by tracking foreground objects frame by frame. When appearances of foreground objects significantly change during occlusions, tracking results may become incorrect. To solve the occlusion problem, we propose a real-time post processing method based on trajectory links. By analyzing trajectory links, our method can retrieve individual trajectories of occluded foreground objects. As shown in the experiments, our method can achieve better performance compared to previous tracking methods.
Chun-Rong Huang, Yi-I Chiu, Pau-Choo Chung, Yu-Chiao Hung
ISCAS1
2014 Maximum a Posteriori Probability Estimation for Online Surveillance Video Synopsis
abstract
To reduce human efforts in browsing long surveillance videos, synopsis videos are proposed. Traditional synopsis video generation applying optimization on video tubes is very time consuming and infeasible for real-time online generation. This dilemma significantly reduces the feasibility of synopsis video generation in practical situations. To solve this problem, the synopsis video generation problem is formulated as a maximum a posteriori probability (MAP) estimation problem in this paper, where the positions and appearing frames of video objects are chronologically rearranged in real time without the need to know their complete trajectories. Moreover, a synopsis table is employed with MAP estimation to decide the temporal locations of the incoming foreground objects in the synopsis video without needing an optimization procedure. As a result, the computational complexity of the proposed video synopsis generation method can be significantly reduced. Furthermore, as it does not require prescreening the entire video, this approach can be applied on online streaming videos.
Chun-Rong Huang, Pau-Choo Chung, Di-Kai Yang, Hsing-Cheng Chen, Guan-Jie Huang
IEEE Trans. Circuits Syst. Video Technol.1
2014 Video Saliency Map Detection by Dominant Camera Motion Removal
abstract
We present a trajectory-based approach to detect salient regions in videos by dominant camera motion removal. Our approach is designed in a general way so that it can be applied to videos taken by either stationary or moving cameras without any prior information. Moreover, multiple salient regions of different temporal lengths can also be detected. To this end, we extract a set of spatially and temporally coherent trajectories of keypoints in a video. Then, velocity and acceleration entropies are proposed to represent the trajectories. In this way, long-term object motions are exploited to filter out short-term noises, and object motions of various temporal lengths can be represented in the same way. On the other hand, we are inspired by the observation that the trajectories in backgrounds, i.e., the nonsalient trajectories, are usually consistent with the dominant camera motion no matter whether the camera is stationary or not. We make use of this property to develop a unified approach to saliency generation for both stationary and moving cameras. Specifically, one-class SVM is employed to remove the consistent trajectories in motion. It follows that the salient regions could be highlighted by applying a diffusion process to the remaining trajectories. In addition, we create a set of manually annotated ground truth on the collected videos. The annotated videos are then used for performance evaluation and comparison. The promising results on various types of videos demonstrate the effectiveness and great applicability of our approach.
Chun-Rong Huang, Yun-Jung Chang, Zhi-Xiang Yang, Yen-Yu Lin
IEEE Trans. Circuits Syst. Video Technol.1
2013 Efficient graph based spatial face context representation and matching
abstract
In this paper, we propose a novel orientation-aware Urquhart graph based spatial face context representation method to efficiently describe the spatial relationship among faces in group photos. We combine graph matching with orientations of graph edges to assess the similarity of spatial face contexts from different group photos. The experimental results show that our method can find more structurally similar group photos compared to the state-of-the-art spatial face context representation methods.
Yi-I Chiu, Chun-Rong Huang, Pau-Choo Chung, Tsuhan Chen
ICASSP3
2013 Gender classification from unaligned facial images using support subspaces
Wen-Sheng Chu, Chun-Rong Huang, Chu-Song Chen
Inf. Sci.2
2012 Hypersphere distribution discriminant analysis
abstract
Current graph embedding frameworks of supervised dimensionality reduction often preserve the intraclass local structures and maximize the interclass variance. However, this strategy fails to provide adequate results when strict within-class multimodalities contradict between-class separations. In this paper, we propose Hypersphere Distribution Discriminant Analysis (HDDA), which determines the affinity by considering not only within-class local structure but also the heteropoint distribution in the neighborhood space. If the heteropoint distribution is relatively high in the feature space, this pair should be mapped apart to avoid mixing problems. By taking both the distribution of heteropoints and the distance into account, HDDA shows more effective results compared to the state-of-the-art methods.
Yi-I Chiu, Chun-Rong Huang, Pau-Choo Chung, Ching-Hsing Luo
ICASSP2
2012 Binary invariant cross color descriptor using galaxy sampling
Guo-Hao Huang, Chun-Rong Huang
ICPR2
2012 Online surveillance video synopsis
abstract
In order to continuously keep sight of environments, taking long-time surveillance videos is required. Thus, enormous storage of videos and video browsing become important problems. To solve these problems, we propose an online surveillance video synopsis method to rearrange positions and appearing frames of foreground objects chronologically in real-time. Comparing with traditional synopsis approaches, our method can directly be applied to streaming surveillance videos and generate synopsis videos without pre-screening entire videos during the synopsis processing.
Chun-Rong Huang, Hsing-Cheng Chen, Pau-Choo Chung
ISCAS1
2012 Sparsity cue in image copy detection
abstract
Image copy detection is an art of searching duplicates from a target database. Computationally efficient and robust detection is still a challenging issue. Inspired by the recent study of sparsity in the context of compressed sensing, we propose a sparse representation-based image copy detection method exploiting sparsity as the cue for searching duplicates. We find that although sparse representation can describe an image in a compact manner, the inherent discriminable features, as far as we know, are not entirely explored. In this paper, we study the discrimination ability inherent in sparsity via online dictionary learning and compact feature descriptor representation. Experimental results show that our method, compared with state-of-the-art, is computationally efficient and attains better or comparable detection performance measured in terms of precision and recall rates.
Huan-Cheng Hsu, Chun-Rong Huang, Chun-Shien Lu
ACM Multimedia2
2011 Temporal Color Consistency-Based Video Reproduction for Dichromats
abstract
In this paper, a video re-coloring algorithm for dichromats is presented. Different from image re-coloring schemes, reproducing a video for dichromats requires maintaining temporal color consistency between frames, i.e., the same color in different frames should be re-colored to the identical new color. To achieve this goal, we extract video key colors from shots after motion estimation at first. Based on the importance of video key colors, a process order is defined to perform efficient color remapping and solve the contrast maintaining problem. Then, the remapped frame pixel values are interpolated by the remapped video key colors with spatial-temporal constraints. Experimental results show that our method can increase the visibility for dichromats and guarantee temporal color consistency.
Chun-Rong Huang, Kuo-Chuan Chiu, Chu-Song Chen
IEEE Trans. Multim.1
2010 An adaptive approach for overlapping people tracking based on foreground silhouettes
abstract
We propose Binary/Appearance Tracker which consists of background subtraction, silhouette similarity and particle filter to infer pedestrians' locations under different occlusion situations with a single camera. During the period of occlusions, binary and color silhouettes are adaptively used to effectively measure the similarity between the observation and the possible combinations of silhouettes. Thus, the occluded pedestrians' locations can be simply located by the most possible combination of silhouettes. The experimental results show that the proposed BATracker can track people successfully even though she/he is fully occluded.
Hsin-Ho Yeh, Jiun-Yu Chen, Chun-Rong Huang, Chu-Song Chen
ICIP3
2010 Identifying Gender from Unaligned Facial Images by Set Classification
abstract
Rough face alignments lead to suboptimal performance of face identification systems. In this study, we present a novel approach for identifying genders from facial images without proper face alignments. Instead of using only one input for test, we generate an image set by randomly cropping out a set of image patches from a neighborhood of the face detection region. Each image set is represented as a subspace and compared with other image sets by measuring the canonical correlation between two associated subspaces. By finding an optimal discriminative transformation for all training subspaces, the proposed approach with unaligned facial images is shown to outperform the state-of-the-art methods with face alignment.
Wen-Sheng Chu, Chun-Rong Huang, Chu-Song Chen
ICPR2
2010 Wheelchair detection using cascaded decision tree
abstract
One of the major goals of healthcare systems is to automatically monitor patients of special needs and alarm the caregivers for providing assistant. In this paper, an efficient single-camera multidirectional wheelchair detector based on a cascaded decision tree (CDT) is proposed to detect a wheelchair and its moving direction simultaneously from video frames for a healthcare system. Our approach combines a decision tree structure and boosted-cascade classifiers to construct a new CDT that can perform early confidence decisions in a hierarchical manner to rapidly reject nonwheelchairs and decide the moving directions. We also impose the tracking history to guide detection routes in the CDT to further reduce detection time and increase detection accuracy. The experiments show over 92% detection rate under cluttered scenes.
Chun-Rong Huang, Pau-Choo Chung, Kuo-Wei Lin, Sheng-Chieh Tseng
IEEE Trans. Inf. Technol. Biomed.1
2009 Video Scene Detection by Link-constrained Affinity-propagation
abstract
Video scenes provide semantic meanings for video content description and summarization. This paper explores the pair-wise visual cues of near-duplicate objects for link-constraint affinity-propagation without using keyframes. Experiments demonstrate that our method is more capable to identify scenes comparing with non-constrained clustering algorithms.
Chun-Rong Huang, Chu-Song Chen
ISCAS1
2008 Contrast context histogram - An efficient discriminating local descriptor for object recognition and image matching
Chun-Rong Huang, Chu-Song Chen, Pau-Choo Chung
Pattern Recognit.1
2008 Helicobacter Pylori-Related Gastric Histology Classification Using Support-Vector-Machine-Based Feature Selection
abstract
This study presents a computer-aided diagnosis system using sequential forward floating selection (SFFS) with support vector machine (SVM) to diagnose gastric histology of Helicobacter pylori (H. pylori) from endoscopic images. To achieve this goal, candidate image features associated with clinical symptoms are extracted from endoscopic images. With these candidate features, the SFFS method is applied to select feature subsets, which perform the best classification results under SVM with respect to different histological features. By using the classifiers obtained from the feature subsets, a new diagnosis system is implemented to provide physicians with H. pylori -related histological results from endoscopic images.
Chun-Rong Huang, Pau-Choo Chung, Bor-Shyang Sheu, Hsiu-Jui Kuo, P. Mikulas
IEEE Trans. Inf. Technol. Biomed.1
2008 Shot Change Detection via Local Keypoint Matching
abstract
Shot change detection is an essential step in video content analysis. However, automatic shot change detection often suffers from high false detection rates due to camera or object movements. To solve this problem, we propose an approach based on local keypoint matching of video frames. This approach aims to detect both abrupt and gradual transitions between shots without modeling different kinds of transitions. Our experiment results show that the proposed algorithm is effective for most kinds of shot changes.
Chun-Rong Huang, Huai-Ping Lee, Chu-Song Chen
IEEE Trans. Multim.1
2007 Efficient hierarchical method for background subtraction
Chu-Song Chen, Chun-Rong Huang, Yi-Ping Hung
Pattern Recognit.3
2006 Image Content Clustering and Summarization for Photo Collections
abstract
Rapid growth of digital photography in recent years spurred the need of photo management tools. In this study, we propose an automatic organization framework for photo collections based on image content, so that a novel browsing experience is provided for users. For each photograph, human faces, together with corresponding clothes and nearby regions are located. We extract color histograms of these regions as the image content feature. Then a similarity matrix of a photo collection is generated according to temporal and content features of those photographs. We perform hierarchical clustering based on this matrix, and extract duplicate subjects of a cluster by introducing the contrast context histogram (CCH) technique. The experimental results show that the developed framework provides a promising result for photo management
Cheng-Hung Li, Chih-Yi Chiu, Chun-Rong Huang, Chu-Song Chen, Lee-Feng Chien
ICME3
2004 An improved algorithm for two-image camera self-calibration and Euclidean structure recovery using absolute quadric
Chun-Rong Huang, Chu-Song Chen, Pau-Choo Chung
Pattern Recognit.1