EDBT 2026 Demo / reviewers in the wild / expert
Chih-Chung Hsu
dblp:35/547
· DBLP profile ↗
47ranked-venue papers
30as first author
23since 2021 · last 2026
0000-0002-2083-4438ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 26 first-author · 14 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WWE-UIE: A Wavelet & White Balance Efficient Network for Underwater Image EnhancementabstractUnderwater Image Enhancement (UIE) aims to restore visibility and correct color distortions caused by wavelength-dependent absorption and scattering. Recent hybrid approaches, which couple domain priors with modern deep neural architectures, have achieved strong performance but incur high computational cost, limiting their practicality in real-time scenarios. In this work, we propose WWE-UIE, a compact and efficient enhancement network that integrates three interpretable priors. First, adaptive white balance alleviates the strong wavelength-dependent color attenuation, particularly the dominance of blue-green tones. Second, a wavelet-based enhancement block (WEB) performs multi-band decomposition, enabling the network to capture both global structures and fine textures, which are critical for underwater restoration. Third, a gradient-aware module (SGFB) leverages Sobel operators with learnable gating to explicitly preserve edge structures degraded by scattering. Extensive experiments on benchmark datasets demonstrate that WWE-UIE achieves competitive restoration quality with substantially fewer parameters and FLOPs, enabling real-time inference on resource-limited platforms. Ablation studies and visualizations further validate the contribution of each component. The source code is available at https://github.com/chingheng0808/WWE-UIE. Ching-Heng Cheng, Jen-Wei Lee, Chia-Ming Lee, Chih-Chung Hsu |
WACV | 4 |
| 2026 | UMCL: Unimodal-generated Multimodal Contrastive Learning for Cross-compression-rate Deepfake DetectionabstractIn deepfake detection, the varying degrees of compression employed by social media platforms pose significant challenges for model generalization and reliability. Although existing methods have progressed from single-modal to multimodal approaches, they face critical limitations: single-modal methods struggle with feature degradation under data compression in social media streaming, while multimodal approaches require expensive data collection and labeling and suffer from inconsistent modal quality or accessibility in real-world scenarios. To address these challenges, we propose a novel Unimodal-generated Multimodal Contrastive Learning (UMCL) framework for robust cross-compression-rate (CCR) deepfake detection. In the training stage, our approach transforms a single visual modality into three complementary features: compression-robust rPPG signals, temporal landmark dynamics, and semantic embeddings from pre-trained vision-language models. These features are explicitly aligned through an affinity-driven semantic alignment (ASA) strategy, which models inter-modal relationships through affinity matrices and optimizes their consistency through contrastive learning. Subsequently, our cross-quality similarity learning (CQSL) strategy enhances feature robustness across compression rates. Extensive experiments demonstrate that our method achieves superior performance across various compression rates and manipulation types, establishing a new benchmark for robust deepfake detection. Notably, our approach maintains high detection accuracy even when individual features degrade, while providing interpretable insights into feature relationships through explicit alignment. Ching-Yi Lai, Chih-Yu Jian, Pei-Cheng Chuang, Chia-Ming Lee, Chih-Chung Hsu, Chiou-Ting Hsu, Chia-Wen Lin |
Int. J. Comput. Vis. | 5 |
| 2026 | TBCNet: Twin-branch collaborative network for hyperspectral anomaly detection
Dong Zhao 0005, Mingtao You, Pei Xiang, Jianling Hu, Yuta Asano, Xin Yu 0002, Chih-Chung Hsu, Huixin Zhou, Jinchang Ren |
Pattern Recognit. | 7 |
| 2025 | Anchoring Trends: Mitigating Social Media Popularity Prediction Drift via Feature Clustering and ExpansionabstractPredicting online video popularity faces a critical challenge: prediction drift, where models trained on historical data rapidly degrade due to evolving viral trends and user behaviors. To address this temporal distribution shift, we propose an Anchored Multi-modal Clustering and Feature Generation (AMCFG) framework that discovers temporally-invariant patterns across data distributions. Our approach employs multi-modal clustering to reveal content structure, then leverages Large Language Models (LLMs) to generate semantic Anchor Features-high-level concepts such as audience demographics, content themes, and engagement patterns-that transcend superficial trend variations. These semantic anchors, combined with cluster-derived statistical features, enable prediction based on stable principles rather than ephemeral signals. Experiments demonstrate that AMCFG significantly enhances both predictive accuracy and temporal robustness, achieving superior performance on out-of-distribution data and providing a viable solution for real-world video popularity prediction. Chia-Ming Lee, Bo-Cheng Qiu, Cheng-Jun Kang, Yi-Hsuan Wu, Jun-Lin Chen, Yu-Fan Lin, Yi-Shiuan Chou, Chih-Chung Hsu |
ACM Multimedia | 8 |
| 2025 | DenseSR: Image Shadow Removal as Dense PredictionabstractShadows are a common factor degrading image quality. Single-image shadow removal (SR), particularly under challenging indirect illumination, is hampered by non-uniform content degradation and inherent ambiguity. Consequently, traditional methods often fail to simultaneously recover intra-shadow details and maintain sharp boundaries, resulting in inconsistent restoration and blurring that negatively affect both downstream applications and the overall viewing experience. To overcome these limitations, we propose the DenseSR, approaching the problem from a dense prediction perspective to emphasize restoration quality. This framework uniquely synergizes two key strategies: (1) deep scene understanding guided by geometric-semantic priors to resolve ambiguity and implicitly localize shadows, and (2) high-fidelity restoration via a novel Dense Fusion Block (DFB) in the decoder. The DFB employs adaptive component processing-using an Adaptive Content Smoothing Module (ACSM) for consistent appearance and a Texture-Boundary Recuperation Module (TBRM) for fine textures and sharp boundaries-thereby directly tackling the inconsistent restoration and blurring issues. These purposefully processed components are effectively fused, yielding an optimized feature representation preserving both consistency and fidelity. Extensive experimental results demonstrate the merits of our approach over existing methods. Our code can be available on https://github.com/VanLinLin/DenseSR Yu-Fan Lin, Chia-Ming Lee, Chih-Chung Hsu |
ACM Multimedia | 3 |
| 2024 | Prompt-guided Multi-modal contrastive learning for Cross-compression-rate Deepfake Detection
Ching-Yi Lai, Chiou-Ting Hsu, Chih-Chung Hsu, Chia-Wen Lin |
BMVC | 3 |
| 2024 | VCDSet: A New Vehicle Collision Dataset In Asia Countries For Anticipating AccidentsabstractThe safety of autonomous vehicles is a crucial concern in the field of transportation. In recent years, a number of research approaches have been proposed to address this issue, including car accident analysis, obstacle detection, lane recognition, and sign recognition. However, there is often the possibility of detecting clues that precede a collision. To better understand driving behavior and enhance the driving experience of autonomous vehicles, a number of large-scale datasets have been created by various research groups. However, none of these datasets specifically focus on risky driving behaviors, which can directly lead to accidents. By detecting risky driving behaviors in advance, it is possible to provide additional response time for autonomous vehicles. While a few car collision datasets do exist, the unique environment in Asian countries, which often involves a high number of motorcycles or bikes, can lead to a wide range of vehicle accidents. In this paper, we introduce a new dataset, named VCDSet, which consists of 603 dashcam videos of car accidents that were collected in Asian countries and include extensive annotations, including weather, road conditions, accident types, and the time at which the accident occurred. We also propose a preliminary approach for anticipating car accidents using our VCDSet and demonstrate that our method can effectively increase response time before a collision occurs. Chih-Chung Hsu, Yun-Zhong Jiang, Wei-Hao Huang |
ICIP | 1 |
| 2024 | LFGN: Low-Level Feature-Guided Network For Adversarial DefenseabstractAdversarial attacks cause deep learning models to fail, which presents a significant challenge in the field. Consequently, the development of adversarial defense techniques has become crucial. Current defense strategies struggle to effectively address adversarial attacks, making a robust defense strategy highly desirable. State-of-the-art adversarial defense schemes mainly rely on adversarial training, which requires massive computational resources. Another strategy, the transformbased approach, is a faster and more efficient way for robust model design. The current state-of-the-art method, Deepimage-prior-based (DIP), requires online training, making fast inference impossible. This paper proposes a novel learning pipeline incorporating conventional low-level features as the transform for fast inference and achieving state-of-the-art performance for adversarial defense. First, we discover the feature transformation for reducing the impact of adversarial attacks since it is hard to approximate using gradients. Conventional low-level feature extraction, such as local binary and ternary patterns, perfectly fits this requirement, allowing us to combine moderate deep neural networks with traditional low-level features for adversarial defense, which could easily be extended to existing defense methods. We conduct comprehensive experiments and analyses to demonstrate the superiority of the proposed adversarial defense scheme and achieve the best trade-off between performance and efficiency in real-world defense scenarios. Chih-Chung Hsu, Ming-Hsuan Wu, En-Chao Liu |
ICIP | 1 |
| 2024 | Revisiting Vision-Language Features Adaptation and Inconsistency for Social Media Popularity PredictionabstractSocial media popularity (SMP) prediction is a complex task involving multi-modal data integration. While pre-trained vision-language models (VLMs) like CLIP have been widely adopted for this task, their effectiveness in capturing the unique characteristics of social media content remains unexplored. This paper critically examines the applicability of CLIP-based features in SMP prediction, focusing on the overlooked phenomenon of semantic inconsistency between images and text in social media posts. Through extensive analysis, we demonstrate that this inconsistency increases with post popularity, challenging the conventional use of VLM features. We provide a comprehensive investigation of semantic inconsistency across different popularity intervals and analyze the impact of VLM feature adaptation on SMP tasks. Our experiments reveal that incorporating inconsistency measures and adapted text features significantly improves model performance, achieving an SRC of 0.729 and an MAE of 1.227. These findings not only enhance SMP prediction accuracy but also provide crucial insights for developing more targeted approaches in social media analysis. Chih-Chung Hsu, Chia-Ming Lee, Yu-Fan Lin, Yi-Shiuan Chou, Chih-Yu Jian, Chi-Han Tsai |
ACM Multimedia | 1 |
| 2024 | Real-Time Compressed Sensing for Joint Hyperspectral Image Transmission and Restoration for CubeSatabstractThis paper addresses the challenges associated with hyperspectral image (HSI) reconstruction from miniaturized satellites, which often suffer from stripe effects and are computationally resource-limited. We propose a Real-Time Compressed Sensing (RTCS) network designed to be lightweight and require only relatively few training samples for efficient and robust HSI reconstruction in the presence of the stripe effect and under noisy transmission conditions. The RTCS network features a simplified architecture that reduces the required training samples and allows for easy implementation on integer-8-based encoders, facilitating rapid compressed sensing for stripe-like HSI, which exactly matches the moderate design of miniaturized satellites on push broom scanning mechanism. This contrasts optimization-based models that demand high-precision floating-point operations, making them difficult to deploy on edge devices. Our encoder employs an integer-8-compatible linear projection for stripe-like HSI data transmission, ensuring real-time compressed sensing. Furthermore, based on the novel two-streamed architecture, an efficient HSI restoration decoder is proposed for the receiver side, allowing for edge-device reconstruction without needing a sophisticated central server. This is particularly crucial as an increasing number of miniaturized satellites necessitates significant computing resources on the ground station. Extensive experiments validate the superior performance of our approach, offering new and vital capabilities for existing miniaturized satellite systems. Chih-Chung Hsu, Chih-Yu Jian, Eng-Shen Tu, Chia-Ming Lee, Guan-Lin Chen |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | QRCODE: Quasi-Residual Convex Deep Network for Fusing Misaligned Hyperspectral and Multispectral ImagesabstractConsidering that hyperspectral image (HSI) is often of lower spatial resolution when compared to multispectral image (MSI), an economical approach for obtaining a high-spatial-resolution (HSR) HSI is to fuse the acquired HSI and MSI, thereby greatly facilitating the subsequent material identification and classification in satellite remote sensing. As satellite-acquired HSI and MSI are often misaligned, the proposed deep neural network does not require the input HSI/MSI to be spatially co-registered, making the challenging fusion network design even more difficult. In this study, we propose a streamlined and efficient convex model integrated into the sub-network, which obviates the need for complex network structures in learning spatial-spectral relationships, effectively guiding the quasi-residual learning task in our alignment-free fusion network. The convex sub-network is a low-rank model that leverages the convex geometric structure implicitly embedded in the hyperspectral signature space. To address the misalignment between HSI and MSI effectively, we introduce a novel Shifted Window Attention Module (SWAM) that exploits the neighboring correlation in the feature domain, significantly enhancing the performance and stability of the fusion task. Capitalizing on the redundancy among spectrums, we employ grouped convolution to decrease the computational complexity without causing additional performance degradation. The proposed Quasi-residual Convex Deep Network (QRCODE) demonstrates state-of-the-art performance in alignment-free HSI/MSI fusion tasks. Chia-Hsiang Lin, Chih-Chung Hsu, Si-Sheng Young, Cheng-Ying Hsieh, Shen-Chieh Tai |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Hyperspectral Tensor Completion Using Low-Rank Modeling and Convex Functional AnalysisabstractHyperspectral tensor completion (HTC) for remote sensing, critical for advancing space exploration and other satellite imaging technologies, has drawn considerable attention from recent machine learning community. Hyperspectral image (HSI) contains a wide range of narrowly spaced spectral bands hence forming unique electrical magnetic signatures for distinct materials, and thus plays an irreplaceable role in remote material identification. Nevertheless, remotely acquired HSIs are of low data purity and quite often incompletely observed or corrupted during transmission. Therefore, completing the 3-D hyperspectral tensor, involving two spatial dimensions and one spectral dimension, is a crucial signal processing task for facilitating the subsequent applications. Benchmark HTC methods rely on either supervised learning or nonconvex optimization. As reported in recent machine learning literature, John ellipsoid (JE) in functional analysis is a fundamental topology for effective hyperspectral analysis. We therefore attempt to adopt this key topology in this work, but this induces a dilemma that the computation of JE requires the complete information of the entire HSI tensor that is, however, unavailable under the HTC problem setting. We resolve the dilemma, decouple HTC into convex subproblems ensuring computational efficiency, and show state-of-the-art HTC performances of our algorithm. We also demonstrate that our method has improved the subsequent land cover classification accuracy on the recovered hyperspectral tensor. Chia-Hsiang Lin, Yangrui Liu, Chong-Yung Chi, Chih-Chung Hsu, Hsuan Ren, Tony Q. S. Quek |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Seeing is Not Believing: Toward Forgery Detection for Hyperspectral ImageabstractAmid the rapid growth of low-earth orbits (LEOs) and hyperspectral imaging (HSI) techniques, this study addresses the overlooked area of security challenges, specifically hyperspectral content manipulation or inpainting. We present HyperForensics, a novel dataset based on NASA’s AVIRIS data, covering diverse scenarios across a wavelength range of 0.4 to 2.5 micrometers. To confront the challenge of detecting small manipulated regions within HSIs, we develop a detection method using a high-resolution network (HRNet) that maintains high resolution and spatial accuracy. Experiments demonstrate this approach outperforms existing methods, highlighting the critical role of spatial-spectral features in HSI forgery detection. This pioneering work in HSI forgery benchmarking and detection invites further research to enhance HSI security. Chih-Chung Hsu, Yu-An Jhang, Min-Tso Ko |
IGARSS | 1 |
| 2023 | Gradient Boost Tree Network based on Extensive Feature Analysis for Popularity Prediction of Social PostsabstractSocial media popularity (SMP) prediction is a complex task, affected by various features such as text, images, and spatial-temporal information. One major challenge in SMP is integrating features from multiple modalities without overemphasizing user-specific details while efficiently capturing relevant user information. This study introduces a robust multi-modality feature mining framework for predicting SMP scores by incorporating additional identity-related features sourced from the official SMP dataset when a user's path alias is accessible. Our preliminary analyses suggest these supplemental features significantly enrich the user-related context, contributing to a substantial improvement in performance and proving that non-identity features are relatively unimportant. This implies that we should focus more on discovering the identity-related features than other meta-data. To further validate our findings, we perform comprehensive experiments investigating the relationship between those identity-related features and scores. Finally, the LightGBM and TabNet are employed within our framework to effectively capture intricate semantic relationships among different modality features and user-specific data. Our experimental results confirm that these identity-related features, especially external ones, significantly improve the prediction performance of SMP tasks. Chih-Chung Hsu, Chia-Ming Lee, Xiu-Yu Hou, Chi-Han Tsai |
ACM Multimedia | 1 |
| 2023 | Adapting Object Detection to Fisheye Cameras: A Knowledge Distillation with Semi-Pseudo-Label ApproachabstractIn this paper, we introduce a lightweight object detection system, custom-designed for fisheye cameras and optimized for quick deployment on embedded systems. Given the constraints of training solely on standard images, our methodology centers on the effective knowledge transfer to accentuate object detection in fisheye scenarios. The integration of the Parallel Residual Bi-Fusion (PRB) Feature Pyramid Network (FPN) into the state-of-the-art YOLOv7 backbone specifically addresses the challenges of detecting tiny objects often present in fisheye images. Chih-Chung Hsu, Wen-Hai Tseng, Ming-Hsuan Wu, Chia-Ming Lee, Wei-Hao Huang |
MMAsia | 1 |
| 2023 | Deep learning-based vehicle trajectory prediction based on generative adversarial network for autonomous driving applications
Chih-Chung Hsu, Li-Wei Kang, Shih-Yu Chen, I-Shan Wang, Ching-Hao Hong, Chuan-Yu Chang |
Multim. Tools Appl. | 1 |
| 2023 | Jointly Defending DeepFake Manipulation and Adversarial Attack Using Decoy MechanismabstractHighly realistic imaging and video synthesis have become possible and relatively simple tasks with the rapid growth of generative adversarial networks (GANs). GAN-related applications, such as DeepFake image and video manipulation and adversarial attacks, have been used to disrupt and confound the truth in images and videos over social media. DeepFake technology aims to synthesize high visual quality image content that can mislead the human vision system, while the adversarial perturbation attempts to mislead the deep neural networks to a wrong prediction. Defense strategy becomes difficult when adversarial perturbation and DeepFake are combined. This study examined a novel deceptive mechanism based on statistical hypothesis testing against DeepFake manipulation and adversarial attacks. First, a deceptive model based on two isolated sub-networks was designed to generate two-dimensional random variables with a specific distribution for detecting the DeepFake image and video. This research proposes a maximum likelihood loss for training the deceptive model with two isolated sub-networks. Afterward, a novel hypothesis was proposed for a testing scheme to detect the DeepFake video and images with a well-trained deceptive model. The comprehensive experiments demonstrated that the proposed decoy mechanism could be generalized to compressed and unseen manipulation methods for both DeepFake and attack detection. Guan-Lin Chen, Chih-Chung Hsu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | DCSN: Deformable Convolutional Semantic Segmentation Neural Network for Non-Rigid ScenesabstractThis paper presents a novel semantic segmentation network for outdoor and unstructured scenarios for autonomous driving based on deformable convolution and geometric distortion pipelines. The semantic segmentation tasks for autonomous driving are generally designed for the urban scene, city-view, and highly structured scenarios, such as the CityScapes dataset, KITTI, and BDD, while rare study focuses on outskirts scenarios. Therefore, the performance of existing semantic segmentation networks on such datasets might be unreliable. To conquer this issue, a novel densely connected residual block (DCRB) with the deformable convolution is proposed to form our backbone for capturing the non-rigid feature representation. In this way, the gradient flow of our DCRB could be better back-propagated from the segmentation head, resulting in a stable training process. Second, geometric distortion augmentation is introduced in the data augmentation pipeline, simulating the possible deformation situations in real-world outdoor scenarios. The experiments are conducted that the proposed semantic segmentation network significantly outperforms the state-of-the-art methods for both Cityscapes and Outdoor scenarios. Bor-Sheng Huang, Chih-Chung Hsu, Wo-Ting Liao, Han-Yi Kao, Xian-Yun Wang |
ICASSP | 2 |
| 2022 | A Comprehensive Study of Spatiotemporal Feature Learning for Social Medial Popularity PredictionabstractFor accurately predicting the popularity of social media, the multi-modal approach was usually adopted to have promising performance. However, the popularity is highly correlated to its identity (i.e., user ID). Inappropriate data splitting could result in lower generalizability in real-world applications. Specifically, we observed that the training and testing datasets are partitioned on a specific timestamp, whereas some users were registered after the timestamp, implying that the partial identities could be missing in the testing phase. It turns the social media prediction (SMP) tasks temporally irrelevant, making the temporal-related feature useless due to missing identities. Therefore, we form the SMP task as an identity-preserving time-series task to observe the popularity scores for the specific identity. In addition, more valuable and essential features could be explored. In this paper, by reformulating the SMP tasks and integrating the multi-modal feature aggregation to the base learner for better performance, how identity-preserving is an essential property for SMP tasks is discussed. We firstly explore the impact of the temporal features with/without identity information in the conventional SMP tasks. Moreover, we reformulate the SMP task in the time-series data-splitting and evaluate the temporal features' importance. Comprehensive experiments are conducted to deliver the suggestions for the SMP tasks and offer the corresponding solutions for effectively predicting the popularity scores. Chih-Chung Hsu, Pi-Ju Tsai, Ting-Chun Yeh, Xiu-Yu Hou |
ACM Multimedia | 1 |
| 2022 | A comprehensive study of age-related macular degeneration detection
Chih-Chung Hsu, Chia-Yen Lee, Cheng-Jhong Lin, Hung Yeh |
Multim. Tools Appl. | 1 |
| 2022 | ADAM Challenge: Detecting Age-Related Macular Degeneration From Fundus ImagesabstractAge-related macular degeneration (AMD) is the leading cause of visual impairment among elderly in the world. Early detection of AMD is of great importance, as the vision loss caused by this disease is irreversible and permanent. Color fundus photography is the most cost-effective imaging modality to screen for retinal disorders. Cutting edge deep learning based algorithms have been recently developed for automatically detecting AMD from fundus images. However, there are still lack of a comprehensive annotated dataset and standard evaluation benchmarks. To deal with this issue, we set up the Automatic Detection challenge on Age-related Macular degeneration (ADAM), which was held as a satellite event of the ISBI 2020 conference. The ADAM challenge consisted of four tasks which cover the main aspects of detecting and characterizing AMD from fundus images, including detection of AMD, detection and segmentation of optic disc, localization of fovea, and detection and segmentation of lesions. As part of the ADAM challenge, we have released a comprehensive dataset of 1200 fundus images with AMD diagnostic labels, pixel-wise segmentation masks for both optic disc and AMD-related lesions (drusen, exudates, hemorrhages and scars, among others), as well as the coordinates corresponding to the location of the macular fovea. A uniform evaluation framework has been built to make a fair comparison of different models using this dataset. During the ADAM challenge, 610 results were submitted for online evaluation, with 11 teams finally participating in the onsite challenge. This paper introduces the challenge, the dataset and the evaluation methods, as well as summarizes the participating methods and analyzes their results for each task. In particular, we observed that the ensembling strategy and the incorporation of clinical domain knowledge were the key to improve the performance of the deep learning models. Huihui Fang, Fei Li 0021, Huazhu Fu, Xu Sun 0006, Xingxing Cao, Fengbin Lin, Jaemin Son, Gwenolé Quellec, Sarah Matta, Sharath M. Shankaranarayana, Chuen-heng Wang, Nisarg A. Shah, Chia-Yen Lee, Chih-Chung Hsu, Hai Xie, Bai Ying Lei, Ujjwal Baid, Shubham Innani, Kang Dang, Wenxiu Shi, Ravi Kamble, Nitin Singhal, Ching-Wei Wang, Shih-Chang Lo, José Ignacio Orlando, Hrvoje Bogunovic, Xiulan Zhang, Yanwu Xu 0001 |
IEEE Trans. Medical Imaging | 16 |
| 2021 | Efficient-ROD: Efficient Radar Object Detection based on Densely Connected Residual NetworkabstractRadar signal-based object detection has become a primary and critical issue for autonomous driving recently. Recently advanced radar object detectors indicated that the cross-model supervision-based approach presented a promising performance based on 3D hourglass convolutional networks. However, the trade-off between computational efficiency and performance of radar object detection tasks is rarely investigated. When higher performance is required in the detection tasks, the 3D convolutional backbone network rarely meets the real-time applications. This paper proposes a lightweight, computationally efficient, and effective network architecture to conquer this issue. First, Atrous convolution, as well-known as dilated convolution, is adopted in our backbone network to make a smaller convolutional kernel having a larger receptive field so as a larger convolutional kernel can be eliminated to reduce the number of parameters. Furthermore, a densely connected residual block (DCSB) is proposed to better deliver the gradient flow from the loss function to improve the feature representation ability. Finally, the hourglass network structure is made by stacking several DCSBs with Mish activation function to form our detection network, termed as DCSN. In this manner, we can keep a larger receptive field and reduce the number of parameters significantly, resulting in an efficient radar object detector. Experiments are demonstrated that the proposed DCSN achieves a significant improvement of inference time and computational complexity, with comparable performance for radar object detection. The source code can be found in https://github.com/jesse1029/RADER-DCSN. Chih-Chung Hsu, Chieh Lee, Min-Kai Hung, Andy Yu-Lun Lin, Xian-Yu Wang |
ICMR | 1 |
| 2021 | DCSN: Deep Compressed Sensing Network for Efficient Hyperspectral Data Transmission of Miniaturized SatelliteabstractRequirements of compressed sensing techniques targeted at miniaturized hyperspectral satellite applications include lightweight onboard hardware, high-speed sensing, low sampling rate for compressing the massive volume of typical hyperspectral data, and noise robustness for reliable data transmission to the ground station. We achieve all these aims via deep learning, and neural networks resulted from which can be implemented on-chip, thereby allowing light hardware implementation. Our neural networks were trained from small-scaled data, but, even so, the resulting encoder achieves a very low sampling rate and very high speed. Unlike typical network training, the input-output pairs are not square but stripe-like images, partly because compressed acquisition does not allow performing compression after obtaining complete data cube and partly because stripe-like acquisition well matches the popular pushbroom hyperspectral sensing schemes. Even with such hard restriction caused by nontraditional training, the resulting decoder still reconstructs the image with high accuracy. To match the requirement of pushbroom sensing, a lightweight encoder is proposed to compress the stripe-like images immediately. Meanwhile, multiscale feature fusion block (MFB) and aggregation (MFA) modules are proposed to form our decoder for enhancing the feature representation of the compressed acquisitions. Furthermore, we achieve joint spatial/spectral super-resolution (SR) progressively, ensuring accurate hyperspectral reconstruction via a low-rank-driven decoder. The encoder and decoder are trained in an end-to-end manner, where noise robustness is forced during the training stage. Comprehensive experiments demonstrate the superiority of the proposed hyperspectral compressed sensing method, as well as its one-shot transfer learning (OTL)-based extension, both quantitatively and qualitatively. Chih-Chung Hsu, Chia-Hsiang Lin, Chi-Hung Kao, Yen-Cheng Lin |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Rethinking Relation between Model Stacking and Recurrent Neural Networks for Social Media PredictionabstractPopularity prediction of social posts is one of the most critical issues for social media analysis and understanding. In this paper, we discover a more dominant feature representation of text information, as well as propose a singe ensemble learning model to obtain the popularity scores, for social media prediction challenge. However, most social media prediction techniques focus on predicting the popularity score of social posts based on a single model, such as deep learning-based or ensemble learning-based approaches. However, it is well-known that the model stacking strategy is a more effective way to boost the performance on various regression tasks. In this paper, we also show that the model stacking can be modeled as a simple recurrent neural network problem with comparable performance on predicting popularity scores. Firstly, a single strong baseline is proposed based on the deep neural network with a prediction branch. Then, the partial feature maps of the last layer of our strong baseline are used to establish a new branch with an isolated predictor. It is easy to obtain multi-prediction by repeating the above two steps. These preliminary predicted scores are then formed as the input of the recurrent unit to learn the final predicted scores, called Recurrent Stacking Model (RSM). Our experiments show that the proposed ensemble learning approach outperforms other state-of-the-art methods. Furthermore, the proposed RSM also shows the superiority over our ensemble learning approach, having verified that the model stacking problem can be transformed into the training problem of a recurrent neural network. Chih-Chung Hsu, Wen-Hai Tseng, Hao-Ting Yang, Chia-Hsiang Lin, Chi-Hung Kao |
ACM Multimedia | 1 |
| 2019 | SSSNet: Small-Scale-Aware Siamese Network for Gastric Cancer DetectionabstractIn recent years, deep neural networks have become the most powerful supervised learning method. Several advanced neural networks, such as AlexNet, ZFNet, Inception, ResNet, and DenseNet, have achieved excellent performance on image recognition tasks. However, deep neural networks rely heavily on huge training sets to obtain good performance. Many applications, such as medical image analysis, do not allow for such large training sets, and it is difficult to train such networks on small-scale training sets. Magnifying narrow band imaging (M-NBI) is widely used to assist doctors in diagnosing gastric cancer, but relatively few of these images are available, compared with the number of general images. In this paper, we propose to use a Siamese network architecture to learn discriminative feature representations based on pairs of images. Then, we use a micro neural network to recognize these features and classify the input images. Our experimental results show that the proposed network can effectively learn discriminative features from a limited number of training images, and also that it can successfully recognize gastric cancer in M-NBI images. Chih-Chung Hsu, Hsin-Ti Ma, Jun-Yi Lee |
AVSS | 1 |
| 2019 | Detecting Generated Image Based on a Coupled Network with Two-Step Pairwise LearningabstractWith the rapid growth of generative adversarial networks (GANs), a photo-realistic image can be easily generated from a low-dimensional random vector nowadays. However, the generated image can be used to synthesize several persons who may have a potential effect on society with radical contents. Considering that many techniques to produce a photo-realistic facial image based on different GANs are already available, collecting training images of all possible generative models is difficult; hence, the learning-based approach would not effectively detect a fake image generated using an excluded generative model. To overcome this shortcoming, we propose a two-step pairwise learning approach to learn common fake features over the training images generated by using different generative models. First, the triplet loss will be used to simulate the relation between fake and real images and utilized to learn the discriminative features to determine whether an image is real or fake. Then, we propose a novel coupled network to accurately capture local and global image features of the fake or real images. The experimental results demonstrate that the proposed method outperforms the baseline supervised learning methods for fake facial image detection. Yi-Xiu Zhuang, Chih-Chung Hsu |
ICIP | 2 |
| 2019 | Popularity Prediction of Social Media based on Multi-Modal Feature MiningabstractPopularity prediction of social media becomes a more attractive issue in recent years. It consists of multi-type data sources such as image, meta-data, and text information. In order to effectively predict the popularity of a specified post in the social network, fusing multi-feature from heterogeneous data is required. In this paper, a popularity prediction framework for social media based on multi-modal feature mining is presented. First, we discover image semantic features by extracting their image descriptions generated by image captioning. Second, an effective text-based feature engineering is used to construct an effective word-to-vector model. The trained word-to-vector model is used to encode the text information and the semantic image features. Finally, an ensemble regression approach is proposed to aggregate these encoded features and learn the final regressor. Extensive experiments show that the proposed method significantly outperforms other state-of-the-art regression models. We also show that the multi-modal approach could effectively improve the performance in the social media prediction challenge. Chih-Chung Hsu, Li-Wei Kang, Chia-Yen Lee, Jun-Yi Lee, Zhong-Xuan Zhan, Shao-Min Wu |
ACM Multimedia | 1 |
| 2019 | Stronger Baseline for Vehicle Re-Identification in the WildabstractRecently, re-identification tasks in computer vision field draw attention. Vehicle re-identification can be used to find the suspect car (target) from a vast surveillance video dataset. One of the most critical issues in the vehicle re-identification task is how to learn the effective feature representation. In general, pairwise learning such as the contrastive and triplet loss functions is adopted to learn the discriminative feature based on the convolution neural network. A good backbone network will lead to a significant improvement in the car re-identification task. In this paper, a stronger baseline method is proposed to achieve a better feature representation ability. First, we integrate the shift-invariant convolutional neural network with ResNet backbone to enhance the consistency feature learning. Afterward, a multi-layer feature fusion module is proposed to incorporate the middle- and high-level features to further improve the performance of car re-identification. Experimental results demonstrated that the proposed stronger baseline method achieves state-of-the-art performance in terms of mean averaging precision. Chih-Chung Hsu, Cing-Hao Hung, Chih-Yu Jian, Yi-Xiu Zhuang |
VCIP | 1 |
| 2019 | SiGAN: Siamese Generative Adversarial Network for Identity-Preserving Face HallucinationabstractThough generative adversarial networks (GANs) can hallucinate high-quality high-resolution (HR) faces from low-resolution (LR) faces, they cannot ensure identity preservation during face hallucination, making the HR faces difficult to recognize. To address this problem, we propose a Siamese GAN (SiGAN) to reconstruct HR faces that visually resemble their corresponding identities. On top of a Siamese network, the proposed SiGAN consists of a pair of two identical generators and one discriminator. We incorporate reconstruction error and identity label information in the loss function of SiGAN in a pairwise manner. By iteratively optimizing the loss functions of the generator pair and the discriminator of SiGAN, we not only achieve visually-pleasing face reconstruction but also ensure that the reconstructed information is useful for identity recognition. Experimental results demonstrate that SiGAN significantly outperforms existing face hallucination GANs in objective face verification performance while achieving promising visual-quality reconstruction. Moreover, for input LR faces with unseen identities that are not part of the training dataset, SiGAN can still achieve reasonable performance. Chih-Chung Hsu, Chia-Wen Lin, Weng-Tai Su, Gene Cheung |
IEEE Trans. Image Process. | 1 |
| 2018 | Joint Pairwise Learning and Image Clustering Based on a Siamese CNNabstractHow to use a deep convolutional neural network (CNN) to efficiently and effectively learn representations of a large unlabeled set of images and group them into clusters remains a challenging problem. To address this problem, we propose a Siamese clustering CNN (SC-CNN) to iteratively learn discriminative representations for image clustering. Based on the proposed SC-CNN, we employ a mini-batch-based joint pairwise representation learning and clustering scheme to make the computation and storage cost efficient for large-scale image clustering on a personal computer with a commercial GPU graphic card. On top of SC-CNN, the proposed pairwise learning scheme effectively learns discriminative representations by appropriately selecting same-cluster and different-cluster image pairs from the results of each clustering iteration. Experimental results demonstrate that the proposed method outperforms start-of-the-art clustering schemes in clustering accuracy on public image sets. Weng-Tai Su, Chih-Chung Hsu, Ziling Huang, Chia-Wen Lin, Gene Cheung |
ICIP | 2 |
| 2018 | An Iterative Refinement Approach for Social Media Headline PredictionabstractIn this study, we propose a novel iterative refinement approach to predict the popularity score of the social media meta-data effectively. With the rapid growth of the social media on the Internet, how to adequately forecast the view count or popularity becomes more important. Conventionally, the ensemble approach such as random forest regression achieves high and stable performance on various prediction tasks. However, most of the regression methods may not precisely predict the extreme high or low values. To address this issue, we first predict the initial popularity score and retrieve their residues. In order to correctly compensate those extreme values, we adopt an ensemble regressor to compensate the residues to further improve the prediction performance. Comprehensive experiments are conducted to demonstrate the proposed iterative refinement approach outperforms the state-of-the-art regression approach. Chih-Chung Hsu, Chia-Yen Lee, Ting-Xuan Liao, Jun-Yi Lee, Tsai-Yne Hou, Ying-Chu Kuo, Jing-Wen Lin, Ching-Yi Hsueh, Zhong-Xuan Zhan, Hsiang-Chin Chien |
ACM Multimedia | 1 |
| 2018 | CNN-Based Joint Clustering and Representation Learning with Feature Drift Compensation for Large-Scale Image DataabstractGiven a large unlabeled set of images how to efficiently and effectively group them into clusters based on extracted visual representations remains a challenging problem. To address this problem we propose a convolutional neural network (CNN) to jointly solve clustering and representation learning in an iterative manner. In the proposed method given an input image set we first randomly pick k samples and extract their features as initial cluster centroids using the proposed CNN with an initial model pretrained from the ImageNet dataset. Mini-batch k-means is then performed to assign cluster labels to individual input samples for a mini-batch of images randomly sampled from the input image set until all images are processed. Subsequently the proposed CNN simultaneously updates the parameters of the proposed CNN and the centroids of image clusters iteratively based on stochastic gradient descent. We also propose a feature drift compensation scheme to mitigate the drift error caused by feature mismatch in representation learning. Experimental results demonstrate the proposed method outperforms start-of-the-art clustering schemes in terms of accuracy and storage complexity on large-scale image sets containing millions of images. Chih-Chung Hsu, Chia-Wen Lin |
IEEE Trans. Multim. | 1 |
| 2017 | Unsupervised convolutional neural networks for large-scale image clusteringabstractThe paper proposes an unsupervised convolutional neural network (UCNN) to solve clustering and representation learning jointly in an iterative manner. The key idea behind the proposed method is that learning better feature representations of images leads to more accurate image clustering results, whereas better image clustering can benefit the feature learning with the proposed UCNN. In the proposed method, given an input image set, we first randomly pick k samples and extract their features as the initial centroids of image clusters using the proposed UCNN with an initial representation model pre-trained from the ImageNet dataset. Mini-batch k-means is then performed to assign cluster labels to individual input samples for a mini-batch of images randomly sampled from the input image set until all images are processed. Subsequently, UCNN simultaneously updates the parameters of UCNN and the centroids of image clusters iteratively based on stochastic gradient descent. Experimental results demonstrate the proposed method outperforms start-of-the-art clustering schemes in terms of accuracy and memory complexity on large-scale image sets containing millions of images. Chih-Chung Hsu, Chia-Wen Lin |
ICIP | 1 |
| 2017 | Social Media Prediction Based on Residual Learning and Random ForestabstractIn this paper, we propose a unified framework for the residual learning and random forest regression for social media prediction task. Given a post including photo and its social information, the primary goal is to predict the view count of the post. In this regression problem, we first predict the view count based on random forest regressor for the social information. Since regressor tends to learn a relative soothingness model to avoid overfitting, the extreme high/low view counts of the poses are hard to predict. We solve this problem by using residual learning to refine the prediction. Based on this initial prediction, the residual value of the prediction and its ground truth is calculated. Then, the image and its social information will feed to 13-layers ResNet to predict the residual value to compensate the initial prediction for extreme high/low view counts. Experiments show that the performance of the proposed method significantly outperforms other methods. Chih-Chung Hsu, Ying-Chin Lee, Ping-En Lu, Shian-Shin Lu, Hsiao-Ting Lai, Ching-Chu Huang, Yang-Jiun Lin, Weng-Tai Su |
ACM Multimedia | 1 |
| 2017 | Objective quality assessment for video retargeting based on spatio-temporal distortion analysisabstractThis paper proposes a novel objective quality metric for video retargeting based on spatio-temporal distortion analysis. The proposed metric combines three indices: one spatial distortion index taking into account perceptual geometric distortion and spatial information loss, and two temporal distortion indices including temporal inconsistency distortion and temporal saliency similarity. Subjective tests are conducted to evaluate the performance of the proposed metric. Our experimental results show the good consistency between the proposed objective metric and the subjective rankings. Chih-Chung Hsu, Chia-Wen Lin |
VCIP | 1 |
| 2016 | Supervised-learning based face hallucination for enhancing face recognitionabstractThis paper presents a two-step supervised face hallucination framework based on class-specific dictionary learning. Since the performance of learning-based face hallucination relies on its training set, an inappropriate training set (e.g., an input face image is very different from the training set) can reduce the visual quality of reconstructed high-resolution (HR) face significantly. To address this problem, we propose to utilize supervised learning to learn a set of class-specific dictionaries so that one of the learned dictionaries can well fit the global and local characteristics of an input low-resolution (LR) face image. Besides, the representative coefficients of the input LR face image may be unreliable due to insufficient information contained in the LR input image. To resolve this issue, we propose a maximum a posteriori estimator to infer the global HR face. Experimental results demonstrate that our method cannot only effectively enhance the visual quality of a reconstructed HR face, but also significantly improves the accuracy of face recognition compared to existing hallucination methods. Weng-Tai Su, Chih-Chung Hsu, Chia-Wen Lin, Weiyao Lin |
ICASSP | 2 |
| 2015 | Temporally Coherent Superresolution of Textured Video via Dynamic Texture SynthesisabstractThis paper addresses the problem of hallucinating the missing high-resolution (HR) details of a low-resolution (LR) video while maintaining the temporal coherence of the reconstructed HR details using dynamic texture synthesis (DTS). Most existing multiframe-based video superresolution (SR) methods suffer from the problem of limited reconstructed visual quality due to inaccurate subpixel motion estimation between frames in an LR video. To achieve high-quality reconstruction of HR details for an LR video, we propose a texture-synthesis (TS)-based video SR method, in which a novel DTS scheme is proposed to render the reconstructed HR details in a temporally coherent way, which effectively addresses the temporal incoherence problem caused by traditional TS-based image SR methods. To further reduce the complexity of the proposed method, our method only performs the TS-based SR on a set of key frames, while the HR details of the remaining nonkey frames are simply predicted using the bidirectional overlapped block motion compensation. After all frames are upscaled, the proposed DTS-SR is applied to maintain the temporal coherence in the HR video. Experimental results demonstrate that the proposed method achieves significant subjective and objective visual quality improvement over state-of-the-art video SR methods. Chih-Chung Hsu, Li-Wei Kang, Chia-Wen Lin |
IEEE Trans. Image Process. | 1 |
| 2015 | Learning-Based Joint Super-Resolution and Deblocking for a Highly Compressed ImageabstractA highly compressed image is usually not only of low resolution, but also suffers from compression artifacts (blocking artifact is treated as an example in this paper). Directly performing image super-resolution (SR) to a highly compressed image would also simultaneously magnify the blocking artifacts, resulting in an unpleasing visual experience. In this paper, we propose a novel learning-based framework to achieve joint single-image SR and deblocking for a highly-compressed image. We argue that individually performing deblocking and SR (i.e., deblocking followed by SR, or SR followed by deblocking) on a highly compressed image usually cannot achieve a satisfactory visual quality. In our method, we propose to learn image sparse representations for modeling the relationship between low- and high-resolution image patches in terms of the learned dictionaries for image patches with and without blocking artifacts, respectively . As a result, image SR and deblocking can be simultaneously achieved via sparse representation and morphological component analysis (MCA)-based image decomposition. Experimental results demonstrate the efficacy of the proposed algorithm. Li-Wei Kang, Chih-Chung Hsu, Boqi Zhuang, Chia-Wen Lin, Chia-Hung Yeh |
IEEE Trans. Multim. | 2 |
| 2014 | Video super-resolution via dynamic texture synthesisabstractThis paper addresses the problem of hallucinating the missing high-resolution (HR) details of a low-resolution (LR) video while maintaining the temporal coherence of the hallucinated HR details by using dynamic texture synthesis (DTS). Most existing multi-frame-based video super-resolution (SR) methods suffer from the problem of limited reconstructed visual quality due to inaccurate sub-pixel motion estimation between frames in a LR video. To achieve high-quality reconstruction of HR details for a LR video, we propose a texture-synthesis-based video super-resolution method, in which a novel DTS scheme is proposed to render the reconstructed HR details in a time coherent way, so as to effectively address the temporal incoherence problem caused by traditional texture synthesis based image SR methods. To further reduce the complexity of the proposed method, our method only performs the DTS-based SR on a selected set of key-frames, while the HR details of the remaining non-key-frames are simply predicted using the bi-directional overlapped block motion compensation. Experimental results demonstrate that the proposed method achieves significant subjective and objective quality improvement over state-of-the-art video SR methods. Chih-Chung Hsu, Li-Wei Kang, Chia-Wen Lin |
MMSP | 1 |
| 2013 | Self-learning-based single image super-resolution of a highly compressed imageabstractLow-quality images are usually not only with low-resolution, but also suffer from compression artifacts (blocking artifact is treated as an example in this paper). Directly performing image super-resolution (SR) to a highly compressed (low-quality) image would also simultaneously magnify the blocking artifacts, resulting in unpleasing visual quality. In this paper, we propose a self-learning-based SR framework to simultaneously achieve single-image SR and compression artifact removal for a highly-compressed image. We argue that individually performing deblocking first, followed by SR to an image, would usually inevitably lose some image details induced by deblocking, which may be useful for SR, resulting in worse SR result. In our method, we propose to self-learn image sparse representation for modeling the relationship between low and high-resolution image patches in terms of the learned dictionaries, respectively, for image patches with and without blocking artifacts. As a result, image SR and deblocking can be simultaneously achieved via sparse representation and MCA (morphological component analysis)-based image decomposition. Experimental results demonstrate the efficacy of the proposed algorithm. Li-Wei Kang, Bo-Chi Chuang, Chih-Chung Hsu, Chia-Wen Lin, Chia-Hung Yeh |
MMSP | 3 |
| 2013 | Objective quality assessment for image retargeting based on perceptual distortion and information lossabstractImage retargeting techniques aim to obtain retargeted images with different sizes or aspect ratios for various display screens. Various content-aware image retargeting algorithms have been proposed recently. However, there is still no accurate objective metric for visual quality assessment of retargeted images. In this paper, we propose a novel objective metric for assessing visual quality of retargeted images based on perceptual geometric distortion and information loss. The proposed metric measures the geometric distortion of retargeted images by SIFT flow variation. Furthermore, a visual saliency map is derived to characterize human perception of the geometric distortion. On the other hand, the information loss in a retargeted image, which is calculated based on the saliency map, is integrated into the proposed metric. A user study is conducted to evaluate the performance of the proposed metric. Experimental results show the consistency between the objective assessments from the proposed metric and subjective assessments. Chih-Chung Hsu, Chia-Wen Lin, Yuming Fang 0001, Weisi Lin |
VCIP | 1 |
| 2011 | Image super-resolution via feature-based affine transformabstractState-of-the-art image super-resolution methods usually rely on search in a comprehensive dataset for appropriate high-resolution patch candidates to achieve good visual quality of reconstructed image. Exploiting different scales and orientations in images can effectively enrich a dataset. A large dataset, however, usually leads to high computational complexity and memory requirement, which makes the implementation impractical. This paper proposes a universal framework for enriching the dataset for search-based super-resolution schemes with reasonable computation and memory cost. Toward this end, the proposed method first extracts important features with multiple scales and orientations of patches based on the SIFT (Scale-invariant feature transform) descriptors and then use the extracted features to search in the dataset for the best-match HR patch(es). Once the matched features of patches are found, the found HR patch will be aligned with LR patch using homography estimation. Experimental results demonstrate that the proposed method achieves significant subjective and objective improvement when integrated with several state-of-the-art image super-resolution methods without significantly increasing the cost. Chih-Chung Hsu, Chia-Wen Lin |
MMSP | 1 |
| 2011 | Fast deconvolution-based image super-resolution using gradient priorabstractSingle-image super-resolution (SR) is to reconstruct a high-resolution image from a low-resolution input image. Nevertheless, most SR algorithms are performed in an iterative manner and are therefore time-consuming. In this paper, we propose an iteration-free single-image SR algorithm based on fast deconvolution with gradient prior. Based on the prior calculated from the initially upsampled image via current approach (e.g., bicubic interpolation or example/learning-based approaches), we make the deconvolution process well-posed, which can be efficiently solved in FFT domain. Moreover, the proposed algorithm can be directly applied to video SR, where the temporal coherence can be automatically maintained. Experimental results demonstrate that the proposed method can simultaneously obtain significant acceleration and quality improvement over several existing SR methods. Chih-Chung Hsu, Chia-Wen Lin, Li-Wei Kang |
VCIP | 2 |
| 2010 | Face hallucination using Bayesian global estimation and local basis selectionabstractThis paper proposes a two-step prototype-face-based scheme of hallucinating the high-resolution detail of a low-resolution input face image. The proposed scheme is mainly composed of two steps: the global estimation step and the local facial-parts refinement step. In the global estimation step, the initial high-resolution face image is hallucinated via a linear combination of the global prototype faces with a coefficient vector. Instead of estimating coefficient vector in the high-dimensional raw image domain, we propose a maximum a posteriori (MAP) estimator to estimate the optimum set of coefficients in the low-dimensional coefficient domain. In the local refinement step, the facial parts (i.e., eyes, nose and mouth) are further refined using a basis selection method based on overcomplete nonnegative matrix factorization (ONMF). Experimental results demonstrate that the proposed method can achieve significant subjective and objective improvement over state-of-the-art face hallucination methods, especially when an input face does not belong to a person in the training data set. Chih-Chung Hsu, Chia-Wen Lin, Chiou-Ting Hsu, Hong-Yuan Mark Liao, Jen-Yu Yu |
MMSP | 1 |
| 2009 | Cooperative face hallucination using multiple referencesabstractThis paper proposes a cooperative example-based face hallucination method using multiple references. The proposed method first uses clustering and residual prototype faces construction to improve the performance of hallucinating a single low-resolution (LR) face to obtain a high-resolution (HR) counterpart. In the case that multiple LR face images for a person are available, a unique feature of the proposed method is to cooperatively enhance the qualities of hallucinated HR images by taking into account the multiple input face images jointly as prior models. Experimental results demonstrate that the proposed cooperative method achieve significant subjective and objective improvement over single-prior schemes. Chih-Chung Hsu, Chia-Wen Lin, Chiou-Ting Hsu, Hong-Yuan Mark Liao |
ICME | 1 |
| 2008 | Accelerating vector quantization of images using modified run length coding for adaptive block representation and difference measurementabstractIn vector quantization (VQ) techniques, the block difference measurement in the nearest neighboring search and the iteration process in codebook generation are time consuming. Since there could be many smoothing or strongly related regions in images, the pixels values in the partitioned image blocks could be identical or quite similar. In this paper, two accelerating methods for codebook training and VQ coding based on the modified run-length coding method that can transform an image block into a coded vector (CV) are proposed. Acceleration can be achieved by measuring the distance between two CVs rather than measuring the difference between two image blocks. Computer simulation results show the proposed scheme outperforms the conventional VQ and the principal component analysis based methods in speeding up both the image coding and codebook training processes. Chih-Chung Hsu, Hsuan-Ting Chang |
ISCAS | 1 |
| 2008 | Video forgery detection using correlation of noise residueabstractWe propose a new approach for locating forged regions in a video using correlation of noise residue. In our method, block-level correlation values of noise residual are extracted as a feature for classification. We model the distribution of correlation of temporal noise residue in a forged video as a Gaussian mixture model (GMM). We propose a two-step scheme to estimate the model parameters. Consequently, a Bayesian classifier is used to find the optimal threshold value based on the estimated parameters. Two video inpainting schemes are used to simulate two different types of forgery processes for performance evaluation. Simulation results show that our method achieves promising accuracy in video forgery detection. Chih-Chung Hsu, Tzu-Yi Hung, Chia-Wen Lin, Chiou-Ting Hsu |
MMSP | 1 |