Chen-Kuo Chiang

dblp:51/4692 · DBLP profile ↗
← Back
31ranked-venue papers
8as first author
11since 2021 · last 2025
0000-0001-5276-1109ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 8 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Virtual Negative Sample Generation Based on Unreliable Pseudo-Labels for Semi-Supervised Semantic Segmentation
abstract
Semi-Supervised semantic segmentation primarily involves pseudo-labeling and consistency regularization. Pseudo-Labeling creates labels for unlabeled data to provide additional supervision, while consistency regularization ensures model predictions remain consistent across different views to enhance robustness. Both methods depend on the quality of pseudo-labels, which can significantly impact performance. To avoid discarding data due to low-confidence pseudo-labels, this work identifies that such unreliable pseudo-labels often exhibit competition among top classes but still retain discriminative information for lower-probability classes. Based on this observation, contrastive learning is introduced by treating unreliable pixels as negative samples for specific categories, thereby enabling more effective utilization of unlabeled data. Furthermore, standard contrastive learning typically relies on randomly sampled negative examples, which may include unrepresentative samples or outliers. To address this issue, Virtual Negative Sample Generation (VNSG) is proposed, which leverages global information to generate virtual negative samples positioned near decision boundaries. These samples serve to push anchors and negatives further apart in the feature space, thereby enhancing the effectiveness of contrastive learning. Experimental results demonstrate that the proposed approach achieves competitive performance on the PASCAL VOC 2012 and Cityscapes datasets.
Hao-Ting Li, Chia-Hsiang Tsai, Chen-Kuo Chiang
AVSS4
2024 Prototype-Guided Masking for Unsupervised Domain Adaptation
abstract
In the domain of Unsupervised Domain Adaptation (UDA), the training process leverages both labeled data from the source domain and unlabeled data from the target domain to facilitate the transfer of knowledge to the target domain. Previous UDA methods have sought to improve robust visual recognition by capitalizing on the consistency among masked images to learn spatial context relations. However, it is imperative to acknowledge that different regions within an image may hold varying degrees of significance across various recognition tasks. The conventional approach typically involves random masking to images, neglecting the evaluation of sample difficulty or ease during the model learning process. To tackle this challenge, we introduce a novel technique, called Prototype-Guided Masking, for semantic segmentation in this paper. It enforces deep models to concentrate on learning from more challenging regions, as opposed to dedicating excessive attention to samples that are easily distinguishable in the initial stages of training. A key component of this approach is the introduction of prototype confidence score, which is utilized to assess and selectively mask images. Experimental results demonstrate that the proposed Prototype-Guided Masking yields highly competitive outcomes when compared to state-of-the-art performance on benchmarks across multiple semantic segmentation datasets.
Kai-Wen Chen, Chen-Kuo Chiang
ICASSP2
2024 Generative Extension Positive Pairs and Improving Sample Selection Based on Contrastive Learning for Unsupervised Person Re-Identification
abstract
In this paper, we present Generative Extension Positive Pairs (GEPP), a novel approach to enhance unsupervised person re-identification (re-id) through contrastive learning. Data generation and pair selection methods significantly impact model performance in contrastive learning. To improve positive pair generation, we incorporate a Generative Adversarial Network (GAN) to create novel views as augmentation samples. We also introduce a sample selection scheme in the contrastive learning process to effectively choose GAN-augmented positive samples. Leveraging our sample selection results, we construct the GEPP framework and propose a unique loss function for contrastive learning. Experimental results showcase that our generative extension of positive pairs and sample selection method offer a versatile, automated, and diverse approach, achieving higher mean average precision (mAP) in re-id tasks than conventional data augmentation techniques. Additionally, our framework outperforms existing state-of-the-art methods on the Market-1501 and MSMT17 datasets. The source code is available at https://github.com/andy412510/Contrastive-sample.
Zheng-An Zhu, Chen-Kuo Chiang
ICASSP2
2024 Imperceptible adversarial attack via spectral sensitivity of human visual system
Chen-Kuo Chiang, Ying-Dar Lin, Ren-Hung Hwang, Po-Ching Lin, Shih-Ya Chang, Hao-Ting Li
Multim. Tools Appl.1
2024 ST-LP: self-training and label propagation for semi-supervised classification
Chih-Wen Lin, Chen-Kuo Chiang, Yu-An Wang, Yue-Lin Yang, Hao-Ting Li, Tzu-Chieh Lin
Multim. Tools Appl.2
2023 Vehicle View Synthesis by Generative Adversarial Network
abstract
In recent years, novel view synthesis methods has been proposed and combined with many models in different computer vision tasks. Previous works solve the problem by using additional 3D information. In this paper, a novel view synthesis method is proposed based on Generative Adversarial Networks (GANs), named PTGAN. PTGAN generates new views by specifying keypoints of vehicles from other views. This makes the pose transformation of vehicle practical when 3D information is unavailable. The proposed PTGAN first extracts identity-related and pose-unrelated feature representations from input images and then concatenates the representation with the pose information to generate the fake image with the assigned pose to deal with the pose variation problem. Experimental results demonstrate that the proposed method achieves very competitive results to the existing view synthesis methods.
Chan-Shuo Hu, Sung-Wei Tseng, Xin-Yun Fan, Chen-Kuo Chiang
ICASSP4
2023 Robust Multi-Object Tracking With Spatial Uncertainty
abstract
Most methods address the multi-object tracking (MOT) problem by tracking-by-detection paradigm, which tracks objects from the detected windows by associating detection boxes whose scores are higher than a given threshold. As such, the confidence score becomes the only indicator of bounding boxes when handling complicated cases, such as occlusions. However, a high confidence score cannot guarantee that the bounding box does not overlap with nearby objects, especially in crowded scenarios. In this paper, spatial uncertainty is proposed for MOT. Firstly, the statistical analysis indicates that spatial uncertainty is highly correlated to the occlusion ratio, which can better represent the level of occlusion of the detection boxes. It is measured by the proposed Sparse Tracker with Spatial Uncertainty (SSUTracker). Then, it is adopted to learn robust tracklet representation. The experimental results demonstrate that it improves overall performance. As a result, our approach achieves very competitive results on popular MOT17 and MOT20 benchmarks compared to state-of-the-art methods.
Pin-Jie Liao, Yu-Cheng Huang, Chen-Kuo Chiang, Shang-Hong Lai
ICASSP3
2023 ReST: A Reconfigurable Spatial-Temporal Graph Model for Multi-Camera Multi-Object Tracking
abstract
Multi-Camera Multi-Object Tracking (MC-MOT) utilizes information from multiple views to better handle problems with occlusion and crowded scenes. Recently, the use of graph-based approaches to solve tracking problems has become very popular. However, many current graph-based methods do not effectively utilize information regarding spatial and temporal consistency. Instead, they rely on single-camera trackers as input, which are prone to fragmentation and ID switch errors. In this paper, we propose a novel reconfigurable graph model that first associates all detected objects across cameras spatially before reconfiguring it into a temporal graph for Temporal Association. This two-stage association approach enables us to extract robust spatial and temporal-aware features and address the problem with fragmented tracklets. Furthermore, our model is designed for online tracking, making it suitable for real-world applications. Experimental results show that the proposed graph model is able to extract more discriminating features for object tracking, and our model achieves state-of-the-art performance on several public datasets. Code is available at https://github.com/chengche6230/ReST.
Cheng-Che Cheng, Min-Xuan Qiu, Chen-Kuo Chiang, Shang-Hong Lai
ICCV3
2022 ELAT: Ensemble Learning with Adversarial Training in defending against evaded intrusions
Ying-Dar Lin, Jehoshua-Hanky Pratama, Didik Sudyana, Yuan-Cheng Lai, Ren-Hung Hwang, Po-Ching Lin, Hsuan-Yu Lin, Wei-Bin Lee, Chen-Kuo Chiang
J. Inf. Secur. Appl.9
2022 Action recognition and tracking via deep representation extraction and motion bases learning
Hao-Ting Li, Yung-Pin Liu, Yun-Kai Chang, Chen-Kuo Chiang
Multim. Tools Appl.4
2021 Conditional Data Augmentation For Sky Segmentation
abstract
Outdoor scene parsing is a very popular topic which algorithms seek to labels or identify objects in images. Sky segmentation is one of the popular outdoor scene parsing task. Sky segmentation models are usually trained on ideal datasets and produce high quality results. However, the performance of sky segmentation model decreases because of varying weather conditions, different time and scene changes due to seasonal weather or other issues in reality. This paper focuses on applying data augmentation methods to generate diversified images. A conditional data augmentation method based on BicycleGAN is proposed in this paper. The model considers mask loss and content loss for improving the quality and details of the generated images. The experimental results demonstrate that the quality of the generated image is better than the existing methods.
Zheng-An Zhu, Chien-Hao Chen, Chen-Kuo Chiang
SNPD3
2019 Generating Adversarial Examples By Makeup Attacks on Face Recognition
abstract
Deep Learning models have been developed rapidly and achieved great success in computer vision and natural language processing. In this paper, we propose to generate adversarial examples to attack well-trained face recognition models by applying makeup effect to face images. It consists of two generative adversarial networks (GANs) based subnetworks, Makeup Transfer Sub-network and Adversarial Attack Sub-network. Makeup Transfer Sub-network transfers the non-makeup face images to makeup faces. Adversarial Attack Sub-networks hides attack information within makeup effect. The generated face images make the well-trained face recognition models misclassified as dodge attack or target attack. The experimental results demonstrate that our method can generate high-quality face makeup images and achieve higher error rates on various face recognition models compared to the existing attack methods.
Zheng-An Zhu, Yun-Zhong Lu, Chen-Kuo Chiang
ICIP3
2019 Sparse representation for image classification via paired dictionary learning
Hui-Hung Wang, Chia-Wei Tu, Chen-Kuo Chiang
Multim. Tools Appl.3
2018 Indoor Scene Layout Estimation from a Single Image
abstract
With the popularity of the hand devices and intelligent agents, many aimed to explore machine's potential in interacting with reality. Scene understanding, among the many facets of reality interaction, has gained much attention for its relevance in applications such as augmented reality (AR). Scene understanding can be partitioned into several sub tasks (i.e., layout estimation, scene classification, saliency prediction, etc). In this paper, we propose a deep learning-based approach for estimating the layout of a given indoor image in real-time. Our method consists of a deep fully convolutional network, a novel layout-degeneration augmentation method, and a new training pipeline which integrate an adaptive edge penalty and smoothness terms into the training process. Unlike previous deep learning-based methods that depend on post-processing refinement (e.g., proposal ranking and optimization), our method motivates the generalization ability of the network and the smoothness of estimated layout edges without deploying postprocessing techniques. Moreover, the proposed approach is time-efficient since it only takes the model one forward pass to render accurate layouts. We evaluate our method on LSUN Room Layout and Hedau dataset and obtain estimation results comparable with the state-of-the-art methods.
Hung-Jin Lin, Sheng-Wei Huang, Shang-Hong Lai, Chen-Kuo Chiang
ICPR4
2017 Hierarchical Structured Dictionary Learning for image categorization
abstract
A novel Hierarchical Structured Dictionary Learning (HSDL) algorithm is proposed in this paper. It aims to learn class-specific dictionaries for all classes simultaneously in a hierarchical structure. A discriminative term based on Fisher discrimination criterion is jointly considered for both the class-specific dictionaries in the lower level and the shared dictionaries in the upper level to enhance the discrimination of dictionaries. The experimental results evaluated on the ImageNet database have shown the superior performance of HSDL over the state-of-the-art dictionary learning methods.
Tzu-Chan Chuang, Chen-Kuo Chiang, Shang-Hong Lai
ICASSP2
2017 Video synthesis from stereo videos with iterative depth refinement
Chen-Hao Wei, Shang-Hong Lai, Chen-Kuo Chiang
J. Vis. Commun. Image Represent.3
2016 Discriminative Paired Dictionary Learning for Visual Recognition
abstract
A Paired Discriminative K-SVD (PD-KSVD) dictionary learning method is presented in this paper for visual recognition. To achieve high discrimination and low reconstruction errors simultaneously for sparse coding, we propose to learn class-specific sub-dictionaries from pairs of positive and negative classes to jointly reduce the reconstruction errors of positive classes while keeping the reconstruction errors of negative classes high. Then, multiple sub-dictionaries are concatenated with respect to the same negative class so that the non-zero sparse coefficients can be discriminatively distributed to improve classification accuracy. Compared to the current dictionary learning methods, the proposed PD-KSVD method achieves very competitive performance in a variety of visual recognition tasks on several publicly available datasets.
Hui-Hung Wang, Chen-Kuo Chiang
ACM Multimedia3
2016 A Multiattribute Sparse Coding Approach for Action Recognition From a Single Unknown Viewpoint
abstract
We propose a novel approach for view-independent action recognition using multiattribute sparse representation enforced with group constraints. First, an oversegmentation-based background modeling and foreground detection approach is employed to extract silhouettes from action videos. Then multiple time intervals of motion history image are computed to capture motion and pose information in human activities. To obtain a more accurate and discriminative representation, we propose multiattribute sparse representation for multiview action video classification. Actions with multiple attributes can be represented by individual attribute matrices to describe group property for each action instance. These attribute matrices are incorporated into the formulation of l1-minimization. The sparsity property as well as the group constraints make the basis selection in sparse coding more efficient in terms of accuracy. Especially, our approach is able to operate under the condition of partially labeled attributes in the training data. Finally, we demonstrate the proposed algorithm through experiments on three multiview human action datasets to show the effectiveness and robustness of the proposed method.
Te-Feng Su, Chen-Kuo Chiang, Shang-Hong Lai
IEEE Trans. Circuits Syst. Video Technol.2
2014 Spatio-Temporally Consistent View Synthesis From Video-Plus-Depth Data With Global Optimization
abstract
We propose a novel algorithm to generate a virtual-view video from a video-plus-depth sequence. The proposed method enforces the spatial and temporal consistency in the disocclusion regions by formulating the problem as an energy minimization problem in a Markov random field (MRF) framework. At the system level, we first recover the depth images and the motion vector maps after the image warping with the preprocessed depth map. Then, we formulate the energy function for the MRF with additional shift variables for each node. To reduce the high computational complexity of applying belief propagation (BP) to this problem, we present a multilevel BPs by using BP with smaller numbers of label candidates for each level. Finally, the Poisson image reconstruction is applied to improve the color consistency along the boundary of the disocclusion region in the synthesized image. Experimental results demonstrate the performance of the proposed method on several publicly available datasets.
Hsiao-An Hsu, Chen-Kuo Chiang, Shang-Hong Lai
IEEE Trans. Circuits Syst. Video Technol.2
2013 Multi-attributed Dictionary Learning for Sparse Coding
abstract
We present a multi-attributed dictionary learning algorithm for sparse coding. Considering training samples with multiple attributes, a new distance matrix is proposed by jointly incorporating data and attribute similarities. Then, an objective function is presented to learn category-dependent dictionaries that are compact (closeness of dictionary atoms based on data distance and attribute similarity), reconstructive (low reconstruction error with correct dictionary) and label-consistent (encouraging the labels of dictionary atoms to be similar). We have demonstrated our algorithm on action classification and face recognition tasks on several publicly available datasets. Experimental results with improved performance over previous dictionary learning methods are shown to validate the effectiveness of the proposed algorithm.
Chen-Kuo Chiang, Te-Feng Su, Chih Yen, Shang-Hong Lai
ICCV1
2013 Face Verification With Local Sparse Representation
abstract
In this letter, a local sparse representation is proposed for face components to describe the local structure and characteristics of the face image for face verification. We first learn a dictionary from collected local patches of face images. Then, a novel local descriptor is presented by using sparse coefficients obtained by the learned dictionary and local face patches from face components to represent the entire human face. We demonstrate the performance of the proposed local sparse representation method on several publicly available datasets. Extensive experiments on both CMU PIE dataset and the challenging LFW database have shown the effectiveness of the proposed method.
Chih-Hsueh Duan, Chen-Kuo Chiang, Shang-Hong Lai
IEEE Signal Process. Lett.2
2013 Learning Component-Level Sparse Representation for Image and Video Categorization
abstract
A novel component-level dictionary learning framework that exploits image/video group characteristics based on sparse representation is introduced in this paper. Unlike the previous methods that select the dictionaries to best reconstruct the data, we present an energy minimization formulation that jointly optimizes the learning of both sparse dictionary and component-level importance within one unified framework to provide a discriminative and sparse representation for image/video groups. The importance measures how well each feature component represents the group property with the dictionary. Then, the dictionary is updated iteratively to reduce the influence of unimportant components, thus refining the sparse representation for each group. In the end, by keeping the top K important components, a compact representation is obtained for the sparse coding dictionary. Experimental results on several public image and video data sets are shown to demonstrate the superior performance of the proposed algorithm compared with the-state-of-the-art methods.
Chen-Kuo Chiang, Chao-Hsien Liu, Chih-Hsueh Duan, Shang-Hong Lai
IEEE Trans. Image Process.1
2012 Parallelized Random Walk algorithm for background substitution on a multi-core embedded platform
abstract
Random Walk (RW) is a popular algorithm and can be applied to many applications in computer vision. In this paper, a fast algorithm is proposed to solve the large linear system in RW based on adapting the Gauss-Seidel method on a multi-core embedded system. Two tables, TYPE and INDEX, are introduced to fast locate the required data for the close-form solution. The computational overhead, along with the memory requirement, to solve the linear system can be reduced greatly, thus making the RW algorithm feasible to many applications on an embedded system. In addition, the proposed fast method is parallelized for a heterogeneous multi-core embedded platform to make the most use of the benefits of the system architecture. Experimental results show that the computational overhead can be significantly reduced by the proposed algorithm.
Yutzu Lee, Chen-Kuo Chiang, Yu-Wei Sun, Te-Feng Su, Shang-Hong Lai
ICASSP2
2011 Learning component-level sparse representation using histogram information for image classification
abstract
A novel component-level dictionary learning framework which exploits image group characteristics within sparse coding is introduced in this work. Unlike previous methods, which select the dictionaries that best reconstruct the data, we present an energy minimization formulation that jointly optimizes the learning of both sparse dictionary and component level importance within one unified framework to give a discriminative representation for image groups. The importance measures how well each feature component represents the image group property with the dictionary by using histogram information. Then, dictionaries are updated iteratively to reduce the influence of unimportant components, thus refining the sparse representation for each image group. In the end, by keeping the top K important components, a compact representation is derived for the sparse coding dictionary. Experimental results on several public datasets are shown to demonstrate the superior performance of the proposed algorithm compared to the-state-of-the-art methods.
Chen-Kuo Chiang, Chih-Hsueh Duan, Shang-Hong Lai, Shih-Fu Chang
ICCV1
2011 Fast H.264 Encoding Based on Statistical Learning
abstract
H.264/AVC, the latest video coding standard of the Joint Video Team, greatly outperforms previous standards in terms of coding bitrate and video quality, because it adopts several new techniques. However, the computational complexity is also considerably increased due to these new components. In this paper, we propose fast algorithms based on statistical learning to reduce the computational cost involved in three main components in H.264 encoder, i.e., intermode decision, multi-reference motion estimation (ME), and intra-mode prediction. First, representative features are extracted to build the learning models. Then, an offline pre-classification approach is used to determine the best results from the extracted features, thus a significant amount of computation is reduced based on the classification strategy. The proposed statistical learning-based approach is applied to the aforementioned three main components in H.264 encoder to speed up the computation. Experimental results show that the ME time of the proposed system is significantly sped up with 12 times faster than the conventional fast ME algorithm of H.264, and the total encoding time of the proposed encoder is greatly reduced with about four times faster than the fast encoder EPZS in the H.264 reference code with negligible video quality degradation.
Chen-Kuo Chiang, Wei-Hau Pan, Chiuan Hwang, ShinShan Zhuang, Shang-Hong Lai
IEEE Trans. Circuits Syst. Video Technol.1
2009 Fast multi-reference motion estimation via statistical learning for H.264/AVC
abstract
In the H.264/AVC coding standard, motion estimation (ME) is allowed to use multiple reference frames to make full use of reducing temporal redundancy in a video sequence. Although it can further reduce the motion compensation errors, it introduces tremendous computational complexity as well. In this paper, we propose a statistical learning approach to reduce the computation involved in the multireference motion estimation. Some representative features are extracted in advance to build a learning model. Then, an off-line pre-classification approach is used to determine the best reference frame number according to the run-time features. It turns out that motion estimation will be performed only on the necessary reference frames based on the learning model. Experimental results show that the computation complexity is about three times faster than the conventional fast ME algorithm while the video quality degradation is negligible.
Chen-Kuo Chiang, Shang-Hong Lai
ICME1
2009 Surface simplification by image retargeting
abstract
Surface simplification aims to reduce the complexity of a 3D model while maintaining a good approximation to the original model. In this work, we propose a novel combination of geometry images and content-aware image resizing to achieve efficient surface simplification. There are two main advantages to simplify surface based on geometry images. First, it is relatively simple to simplify the surface in the parameterized 2D space because the features of a 3D surface can be easily represented by the gradient energy. Second, the regularity and features on 3D surface can also be preserved without additional effort. The proposed retargeting algorithm performs well both on real images and 3D surface simplification.
Shu-Fan Wang, Yi-Ling Chen 0004, Chen-Kuo Chiang, Shang-Hong Lai
SIGGRAPH ASIA Sketches3
2009 Fast JND-Based Video Carving With GPU Acceleration for Real-Time Video Retargeting
abstract
A recently developed image resizing technique, seam carving, has been proved to be a useful tool for content-adaptive spatially nonuniform image resizing with the purpose of optimal display on a screen of reduced resolution or different aspect ratio. In this paper, we present a fast algorithm for real-time content-aware video retargeting based on the improved seam carving method proposed in this paper. The proposed algorithm is designed to be highly parallelizable and suitable for running on a multicore architecture. First, two novel operators, i.e., seam update and seam split, are introduced to analyze an image for detecting the local and global seams with minimum costs very efficiently. With these operators, parallel processing can be achieved to determine multiple seams simultaneously. In addition, the saliency measure is extended with a just-noticeable-distortion model which makes the resized video more consistent with human perception. We demonstrate the efficiency of the above new components with a graphics processing unit (GPU) implementation. In addition, the proposed fast seam carving algorithm is extended for video retargeting. To the best of our knowledge, this is the first paper based on the seam carving method to achieve real-time video retargeting on a GPU. Experimental results on video sequences of various characteristics are demonstrated to show the superior performance of the proposed algorithm in comparison with the existing content-adaptive image/video resizing methods.
Chen-Kuo Chiang, Shu-Fan Wang, Yi-Ling Chen 0004, Shang-Hong Lai
IEEE Trans. Circuits Syst. Video Technol.1
2008 Fast Intermode Decision Via Statistical Learning for H.264 Video Coding
Wei-Hau Pan, Chen-Kuo Chiang, Shang-Hong Lai
MMM2
2007 Adaptive Multi-Reference Downhill Simplex Search Based on Spatial-Temporal Motion Smoothness Criterion
abstract
Multi-reference frame motion estimation improves the accuracy of motion compensation in video coding. However, it also increases computational complexity dramatically. In this paper, we propose a different approach for multi-reference motion estimation via downhill simplex search. Additionally, an adaptive reference frame selection algorithm is developed based on spatial and temporal smoothness of motion vectors. We first apply single-reference downhill simplex search to the previous frame. Then, temporal smoothness of motion vectors in collocated blocks is calculated to decide the number of reference frames to be included for motion estimation. Spatial smoothness of motion vectors in the neighboring blocks is used as a criterion for termination. Experimental results show that the proposed algorithm provides better PSNR than that of original multi-reference downhill simplex search in all testing sequences with similar computational speed. In addition, it outperforms several representative single-reference frame block matching methods in terms of estimation speed and coding quality.
Wei-Hau Pan, Chen-Kuo Chiang, Shang-Hong Lai
ICASSP (1)2
2006 Fast Multi-Reference Frame Motion Estimation via Downhill Simplex Search
abstract
Multi-reference frame motion estimation improves the accuracy of motion compensation in video compression, but it also dramatically increases computational complexity. Based on tracing motion vector trajectories, fast approximated motion estimation results can be obtained for multi-reference frames. In this paper, we extend the downhill simplex search to multiple reference frames and propose several enhanced schemes to improve its efficiency and accuracy. Experimental results show that the proposed algorithm outperforms several representative single-reference frame block matching methods
Chen-Kuo Chiang, Shang-Hong Lai
ICME1