Tsung-Jung Liu

dblp:98/8544 · DBLP profile ↗
← Back
45ranked-venue papers
9as first author
23since 2021 · last 2025
0000-0003-4296-0942ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 6 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 12 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 8 since 2021Systems, architecture and hardware · 6 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
YearPublicationVenuePosition
2025 Object Detection and Fruit Tree Growth Stage Identification Via YOLO with Inverted and Swin Transformer Blocks
abstract
Object detection remains a critical challenge in computer vision, demanding algorithms that are both fast and accurate. This research introduces YOLO-IST, a novel object detector that integrates YOLO, Inverted Blocks, and Swin Transformer. Two novel Inverted blocks are integrated into this architecture to enhance performance. To validate our approach, we construct a new fruit tree dataset for fruit growth stage detection and recognition. Furthermore, we evaluate our model on this newly constructed dataset and two standard benchmark datasets like Microsoft COCO 2017 and Pascal VOC, demonstrating its effectiveness. The code for our proposed YOLO-IST is available at GitHub: https://github.com/nutcliu2507/YOLO-IST.
Kuan-Hsien Liu, Mingru Wang, Tsung-Jung Liu
ICIP3
2025 Efficient Feature-Guided Approach for Image Restoration
abstract
In this paper, we propose a lightweight image restoration method that achieves computational efficiency through guided restoration and detail enhancement. Our approach introduces two key components. First, the Feature Pick (FP) module directs the restoration process by filtering out redundant features, reducing computational overhead. Second, the Detail Auxiliary Block (DAB) module enhances image details by dynamically adjusting weights, allowing finer details to bypass the main restoration network and alleviating its processing burden. Together, these modules significantly improve efficiency while maintaining high restoration quality. We evaluate our method on multiple image restoration tasks, including denoising and deblurring, demonstrating state-of-the-art performance with reduced computational cost. The source code and pretrained model are available at https://github.com/leisoul/FPNET.
Chan-Yu Wang, Tsung-Jung Liu, Kuan-Hsien Liu
ICIP2
2025 Stacked MBN: Stacked Multi-Branch Network for Human Pose Estimation
abstract
Human pose estimation plays a critical role in computer vision, demonstrating broad potential in applications such as human-computer interaction, motion analysis, and animation. In the past few years, a lot of researchers utilized Convolutional Neural Networks (CNNs) as the base for their networks. Nevertheless, the traditional CNN models face challenges in effectively capturing information at multi-scales. To solve this problem, researchers increase the depth of the networks. However, as these networks become deeper, there is a simultaneous increase in both the number of parameters and computational complexity. The increasing number of parameters and growing computational complexity gradually become another tricky problem. Therefore, we propose a multi-branch module to deal with the issue of capturing multi-scale information while maintaining lower parameters and computational complexity. Moreover, we integrate channel attention and spatial attention mechanisms into the module without excessive parameters or computational burdens.
Pin-Chuan Liu, Kuan-Hsien Liu, Tsung-Jung Liu
ISNCC3
2025 Single Image Reflection Separation by Using Reflection and Refraction Estimations
abstract
Considering the influence of reflected and transmitted light on physical imaging in real-world scenarios, we propose a method that leverages reflection and refraction coefficients. By calculating these coefficients between the captured image and the transmission image, our approach effectively guides the separation of reflection layers within the captured image. The proposed three-branch framework integrates the Feature Enhancement Module (FEM), which is specifically designed to recover additional details and produce high-quality transmission images. These enhancements significantly improve the performance of the reflection removal process. Comprehensive experiments conducted on various datasets and in comparison with state-of-the-art reflection removal methods demonstrate the effectiveness of our approach. The results show that our method excels in eliminating reflection artifacts and correcting intensity distortions, yielding superior image quality. Especially, the PSNR and SSIM scores of our model outperform existing methods. Furthermore, the simplicity of the input and architecture underscores the practicality of the proposed method. The source code of the proposed method is available at https://reurl.cc/LnAmqx.
Tsung-Jung Liu, U.-In Chan, Kuan-Hsien Liu
SMC1
2025 Lightweight Dual Attention Multi-Scale Inverted Residual Neural Network for Image Inpainting
abstract
We propose DA-MSIRNet, a lightweight yet innovative architecture for high-quality image inpainting that significantly enhances the standard U-Net through four key innovations: (1) Context Anchor Attention (CAA) for efficient global context modeling via adaptive region selection, (2) Sparse Self-Attention (SpA), inspired by Spa-former, for dynamic and precise local detail refinement by focusing on salient relationships, (3) Multi-Scale Inverted Residual (MSIR) modules for enhanced multi-scale feature fusion through optimized skip connections, and (4) Structural Similarity (SSIM) Loss for improved perceptual quality and fidelity. DA-MSIRNet effectively addresses critical limitations of existing methods, including GAN instability, U-Net’s restricted receptive field, and Transformer computational complexity. Comprehensive evaluations on Places2 and CelebA-HQ datasets demonstrate that DA-MSIRNet achieves state-of-the-art performance in both quantitative metrics (PSNR/SSIM/FID) and visual quality, while maintaining superior computational efficiency. The code for our DA-MSIRNet is publicly available on GitHub: https://github.com/nutcliu2507/DA-MSIRNet.
Kuan-Hsien Liu, Chun-Chieh Chang, Tsung-Jung Liu
SMC3
2025 SCCOME: Scene Change Capture and Optical Motion Estimation for Video Quality Assessment
abstract
The rapid growth of user-generated content (UGC) on social platforms has created a pressing need for effective outdoor video quality assessment. Evaluating video quality in uncontrolled environments is challenging due to the absence of reference videos and distortions caused by compression and transmission artifacts, limiting the applicability of traditional metrics. In this paper, we propose a novel no-reference video quality assessment model that reduces computational complexity by identifying frames with significant visual changes. Optical flow detection is then applied to these frames to capture perceptually important regions, enabling focused processing. Experiments on three public outdoor video quality databases—KoNViD-1k, LIVE-Qualcomm, and CVD2014—demonstrate the effectiveness of our method. Furthermore, ablation studies highlight the critical roles of frame selection and optical flow-based region analysis in improving model performance. The source code is available at https://github.com/Hsiang417/SCCOME.
Tsung-Jung Liu, Hao-Shiang Liao, Kuan-Hsien Liu
SMC1
2025 LLM-DETR: An Enhanced DETR with a Large Language Model-Inspired Attention Mechanism for Object Detection
abstract
We propose LLM-DETR, an enhanced version of the DETR (DEtection TRansformer) framework that integrates advanced attention mechanisms inspired by large language models (LLMs) for object detection. Applied to the MS COCO 2017 dataset, LLM-DETR demonstrates a 5% increase in average precision (AP) over the original DETR, while exhibiting efficient GPU training. This improvement underscores the potential of LLM-inspired attention mechanisms for advancing object detection accuracy in various domains. The code for our LLM-DETR is publicly available on GitHub: https://github.com/mingruWang/LLM-DETR.
Kuan-Hsien Liu, Mingru Wang, Tsung-Jung Liu
SMC3
2024 Micro-Expression Recognition Based On 3DCNN Combined With GRU and New Attention Mechanism
abstract
Micro-expressions, as a form of non-verbal emotional expressions, play a key role in interpersonal interaction. However, they are also quite challenging and not easy to analyze. In this paper, we propose a dual-branch shallow 3DCNN architecture that combines Gated Recurrent Unit (GRU) and enhances the Channel Attention Module (CAM) in Convolutional Block Attention Module (CBAM) to make it more suitable for recognizing micro facial expressions. Experiments show that the proposed method can achieve good results with a relatively simple architecture. The source code and pre-trained models are available at https://github.com/dannyFan-0201/ICIP_2024.
Chun-Ting Fang, Tsung-Jung Liu, Kuan-Hsien Liu
ICIP2
2024 LIRSRN: A Lightweight Infrared Image Super-Resolution Network
abstract
Infrared imaging is crucial for many applications, such as night vision and environmental monitoring. However, it often suffers from lower spatial resolution compared to visible light imaging, which hinders its effectiveness in scenes demanding fine detail discernment. This paper introduces a Lightweight Infrared Image Super-Resolution Network (LIRSRN), which is a novel architecture designed to enhance the resolution of IR images with minimal computation overhead. The core of LIRSRN is the Attention Enhancement Module, which synergizes various attention mechanisms to emphasize important features across both channel and spatial domains. It is further enhanced by spatial frequency processing, which helps to extract key image attributes. The training is conducted on the DIV2K dataset, and the test results from "results-A" and "results-C" demonstrate the superior performance of LIRSRN in terms of PSNR and SSIM. Ablation studies compare the trade-offs between performance and computational cost using different attention mechanisms. Overall, the proposed model strikes a balance between performance and complexity. The source code and trained model are available at https://reurl.cc/yYDnA6.
Chun-An Lin, Tsung-Jung Liu, Kuan-Hsien Liu
ISCAS2
2024 RDLNET: Residual Dense Block based Lightweight Network for Video Super-Resolution
abstract
Deep learning has been widely used in video super-resolution (VSR). Most VSR methods focus on achieving better quality, and the design of deep neural networks is also becoming more complex. Since the input data of video super-resolution is already huge, if the neural network is too complex, it will cause a higher memory load, and a higher-end graphic card device is required. In order to reduce the cost of VSR system, in the paper, we propose a new method called RDLNET, which is a residual dense block based neural network with lightweights to deal with VSR. In the experiments, our proposed lightweight VSR method reduces 25% parameters and maintains almost the same PSNR compared with other state-of-the-art VSR methods. The source code of RDLNET is made available at GitHub https://github.com/nutcliu2507/RDLNET.
Kuan-Hsien Liu, Chih-Jung Wang, Tsung-Jung Liu, Wen-Ren Liu
ISCAS3
2024 AgeSynthGAN: Advanced Facial Age Synthesis with StyleGAN2
abstract
Facial age synthesis is an important research area that aims to synthesize facial images from the past (age progression) or the future (age regression) by reflecting the age factors of a given face. Ideally, this task should be able to synthesize natural faces of different ages while maintaining identity consistency. However, existing methods suffer from background blur and inconsistency issues when processing the generated images. In this study, we propose a new method that utilizes semantic segmentation technology and attention mechanisms to solve these problems. Through this method, we can control the background of the generated image, making it clearer and more consistent, and we added an attention mechanism to better extract the features of faces. Additionally, we introduce a shape loss to simulate the shape changes of faces across different ages. We evaluate our approach through both qualitative and quantitative assessments. The results show that, compared with state-of-the-art methods, our approach is visually superior, achieves higher age accuracy, and provides reasonable identity confidence. Overall, our research presents a new method for facial age synthesis with promising applications and theoretical significance. The source code and model are available at https://reurl.cc/34QX5O.
Tung-Ke Hsieh, Tsung-Jung Liu, Kuan-Hsien Liu
VCIP2
2023 Image Inpainting by Mscswin Transformer Adversarial Autoencoder
abstract
Image inpainting has been researched for years. From deeper and larger models to models that focus on global information, all of them aim to obtain results closer to reality. In this paper, we combine the stripe window and line-by-line feature shift to modify the Vision Transformer (ViT) to reduce the computation cost and obtain global information from the oblique attention. In addition, we design a new loss function to enhance the texture and colors for inpainting. At last, to validate the efficacy of our proposed model, we conduct extensive experiments on commonly seen datasets (Places2 and CelebA) compared with other state-of-the-art methods. The source code and pretrained models are available at https: //github.com/bobo0303/MSCS-Net.
Bo-Wei Chen, Tsung-Jung Liu, Kuan-Hsien Liu
ICIP2
2023 Compound Multi-Branch Feature Fusion for Image Deraindrop
abstract
Image restoration is a challenging and ill-posed problem which also has been a long-standing issue. In this paper, we proposed a multi-branch restoration model inspired from the Human Visual System (i.e., Retinal Ganglion Cells) for image deraindrop. The experiments show that the proposed multi-branch architecture, called CMFNet, has state-of-the-art performance results. The source code and pretrained models are available at https://github.com/FanChiMao/CMFNet. And the interactive demonstration of the proposed deraindrop model can be accessed at https://reurl.cc/dXaeNg.
Chi-Mao Fan, Tsung-Jung Liu, Kuan-Hsien Liu
ICIP2
2023 Facial Expression Recognition in the Wild Using FAM, RRDB and Vision Transformer Based Convolutional Neural Networks
abstract
Facial expression recognition in the wild is a quite challenge task because facial images captured in natural settings are often affected by various factors such as irregular occlusions, inconsistent face angles and varying light levels. To overcome these factors, we presented a new triple-branch deep neural network model containing facial attention module, residual in residual dense block and vision transformer to deal with facial expression recognition problem. The facial attention module can help backbone network extract more useful features. Residual in residual dense block and vision transformer can further get more detailed features. In addition to the image problems caused by above factors, the ratio of number of images for different expressions in the dataset is also very different. We proposed a new loss function to tackle this uneven ration problem. Experimental results on two benchmark in-the-wild datasets show that our model is indeed helpful for images in the wild. We also created a new expressions dataset called Fairfaceplus, which is built by adding expression categories to the original label on FairFace dataset. The code of our proposed method will be made available on GitHub https://github.com/dreampledge/fairfaceplus.
Kuan-Hsien Liu, Xiang-Kun Shih, Tsung-Jung Liu, Wen-Ren Liu
SMC3
2022 Super-Resolution of Satellite Images by two-Dimensional RRDB and Edge-Enhancement Generative Adversarial Network
abstract
With the increasing demand for high-resolution images, image super-resolution (SR) technology has become one of the focuses in related research fields. Generally speaking, high resolution is usually achieved by increasing the density and accuracy of the sensor. However, such an approach is quite expensive for equipment and design. In particular, increasing the density of satellite sensors must be undertaken great risks. Inspired by EEGAN and based on it, the Ultra-Dense Subnet (UDSN) and Edge Enhanced Network (EEN) were modified. Among them, the UDSN is used for feature extraction and obtains high-resolution results that look clear in the intermediate but are deteriorated by artifacts, and the Edge-Enhanced Subnet (EESN) is used to purify, extract and enhance the image contour and use mask processing to eliminate images contaminated by noise. Finally, the restored intermediate image and the enhanced edge are combined to produce a high-resolution image with high credibility and clear content. We use Kaggle and AID open experimental datasets to test and compare the results among different methods. It proves the performance of the proposed model is better than other SR methods.
Yu-Zhang Chen, Tsung-Jung Liu, Kuan-Hsien Liu
ICASSP2
2022 Half Wavelet Attention on M-Net+ for Low-Light Image Enhancement
abstract
Low-Light Image Enhancement is a computer vision task which intensifies the dark images to appropriate brightness. It can also be seen as an illposed problem in image restoration domain. With the success of deep neural networks, the convolutional neural networks surpass the traditional algorithm-based methods and become the mainstream in the computer vision area. To advance the performance of enhancement algorithms, we propose an image enhancement network (HWMNet) based on an improved hierarchical model: M-Net+. Specifically, we use a half wavelet attention block on M-Net+ to enrich the features from wavelet domain. Furthermore, our HWMNet has competitive performance results on two image enhancement datasets in terms of quantitative metrics and visual quality. The source code and pretrained model are available at https://github.com/FanChiMao/HWMNet.
Chi-Mao Fan, Tsung-Jung Liu, Kuan-Hsien Liu
ICIP2
2022 SUNet: Swin Transformer UNet for Image Denoising
abstract
Image restoration is a challenging ill-posed problem which also has been a long-standing issue. In the past few years, the convolution neural networks (CNNs) almost dominated the computer vision and had achieved considerable success in different levels of vision tasks including image restoration. However, recently the Swin Transformer-based model also shows impressive performance, even surpasses the CNN-based methods to become the state-of-the-art on high-level vision tasks. In this paper, we proposed a restoration model called SUNet which uses the Swin Transformer layer as our basic block and then is applied to UNet architecture for image denoising. The source code and pre-trained models are available at https://github.com/FanChiMao/SUNet.
Chi-Mao Fan, Tsung-Jung Liu, Kuan-Hsien Liu
ISCAS2
2022 Fusion of Triple Attention to Residual in Residual Dense Block to Attention Based CNN for Facial Expression Recognition
abstract
In recent years, facial expression recognition has always been a popular research topic. Since the variation among human facial expressions is huge, facial expression recognition is still one of the challenging topics in computer vision. In the facial expression recognition, the key challenge is to capture the dynamic variation of the physical structure of the face from the videos. However, traditional machine learning based systems may have large errors due to the feature extraction on environmental factors such as postures, angles, light, occlusion or backgrounds, which can easily lead to unrecognized or error identification results. In this work, we propose a deep learning based framework, and add residual in residual dense block and triple attention mechanism to our model. On several benchmark facial expression datasets, we demonstrated the effectiveness of our method and compare our model with other state-of-the-art modalities.
Kuan-Hsien Liu, Ching-Hsiang Chiu, Tsung-Jung Liu
SMC3
2022 Clothing Retrieval from Vision Transformer and Class Aware Attention to Deep Embedding Distance Learning
abstract
Clothing is an indispensable item in human life. Since the types and names of clothing are various, sometimes texts are not enough for searching the desired clothing item. The method of searching pictures through pictures has been widely used in many online shopping platforms. In order to help customers find similar clothes accurately, we crawl images from some online shopping websites to build a new dataset to conduct clothing retrieval task. We also propose a new clothing retrieval neural network model, which integrates the advantages of Triplet Network, Perceptual Loss and Capsule Network, and extends the triplet sample learning mode to add multiple negative samples for training. In our proposed model, Vision Transformer and two newly proposed attention mechanisms are applied to the Capsule Network to strengthen the features in each capsule. The Perceptual concept helps the model to extract color and texture features, and connect them. Our newly modified loss function can help the model retrieve better results. The overall training does not need landmark annotation which makes model increase complexity. In the experiments, our model is compared with other state-of-the-art models and shows better performance.
Kuan-Hsien Liu, Yu-Hsiang Wu, Tsung-Jung Liu
SMC3
2022 Image Inpainting with Frequency Domain Wavelet Convolution
abstract
This paper used Time-Frequency Analysis (TFA) techniques for signal processing on tasks of computer vision. Our main idea is as follows: To build a simple network architecture without two or more convolutional neural networks (CNNs), ana-lyze hidden features by Discrete Wavelet Transform (DWT), and send them into filters as weights by convolutions, transformers or other methods. And we do not need to build the network with 2 or more stages to accomplish this idea. Actually, we try to directly use TFA skills on CNN to build one-stage network. Networks which build by this way not only keep their outstanding performance, but also cost lower computing resources. In this paper, we mainly use DWT on CNN to solve image inpainting problems. And the results show that our model can work stably in frequency domain to realize free-form image inpainting.
Jain-Kai Huang, Tsung-Jung Liu, Kuan-Hsien Liu
VCIP2
2022 Clothing Retrieval from Channel Attention to KN Loss Learning
abstract
Due to the high diversity of clothing types and names, using text to search may not find the desired clothes. Clothing retrieval via images can help users find their favorite clothes more conveniently. So, it is widely used in many online stores and platforms, and gradually replaces text search. Clothing retrieval performance is often affected due to factors such as shooting angle, occlusion, and lighting. Therefore, we propose a low-parameter clothing retrieval model without using any additional annotations or labels to aid learning. In the model, we use the classification capability of the Capsule Network with multiple residual convolution modules to learn features, and the VGG16 network is used as a clothing feature extractor to help the model learn features such as clothing color and texture. We add different attention mechanisms to the model for feature enhancement and design a new embedding method. We also adjust the loss function to sample multiple negative classes to improve the similarity clustering effect, and add weights to strengthen it. From the experimental results, one can find that our proposed method achieves better retrieval accuracy than several state-of-the-art models.
Kuan-Hsien Liu, Yu-Hsiang Wu, Tsung-Jung Liu
VCIP3
2021 Age Regression with Specific Facial Landmarks by Dual Discriminator Adversarial Autoencoder
abstract
Facial age conversion is to generate faces of different age groups from the input face and retain the characteristics of the original face. Most of the existing methods are exploring the aging of the face, while ignoring the rejuvenation. In addition to improving aging, we will also explore the regression of human faces. Due to the lack of images of the same person in a longer age range, it becomes a challenging task. Since the generated faces are relatively unreal, we developed a novel model based on Conditional Adversarial Autoencoder (CAAE). This model uses two discriminators to generate a more realistic image. Furthermore, by considering specific facial landmarks where the face shape has changed greatly in different age groups, the face shape belonging to the corresponding age group can be obtained. Moreover, the collected database is divided into different races for training to improve the age development of different races.
Li-Chi Lan, Tsung-Jung Liu, Kuan-Hsien Liu
ICIP2
2021 Mix Attention Based Convolutional Neural Network for Clothing Brand Logo Recognition and Classification
abstract
Over the past years, fashion clothing related research problems such as clothing style classification, clothing attribute recognition, clothing recommendation, and clothing retrieval, have received great attention in the society of computer vision, pattern recognition and image processing. Recently, approaches based on deep learning have been proposed to tackle with aforementioned problems. However, no research work focusing on clothing brand identification/recognition is investigated. To deal with clothing brand identification problem, we construct a new very large-scale clothing dataset containing brand information and the brand logo bounding box information if it has one. Totally, we collected 416 fashion clothing brands in our newly built dataset. In this work, we propose a fashion Clothing Brand Recognition Network (CBR-Net), which contains two parts. First part is the identification network for the clothing brand logo detection and recognition, and the second part is an attention module based network for classifying clothing brands without logos. In the experiments, we demonstrated that our newly proposed CBR-Net can attain much better performance than other state-of-the-art methods.
Kuan-Hsien Liu, Guan-Hong Chen, Tsung-Jung Liu
SMC3
2020 Locating Waterfowl Farms from Satellite Images with Parallel Residual U-Net Architecture
abstract
For the epidemic prevention of avian influenza, there exist lots of differences between ideality and reality. This is why the epidemic is usually out of control. One of the reasons is that many illegal waterfowl farms are built without government registration. In this work, we proposed a new method trying to directly locate waterfowl farms, including both registered and unregistered ones without the need of human labeling. This will not only save human labors, but also update the location and size information of waterfowl farms regularly due to the computing speed of computers. In this work, we proposed a new method for satellite image augmentation. The layers of the model we proposed are not deeper than the other deep neural network models. However, we show that using the existing simple U-Net combined with residual blocks has better performance than the other deep models in this task.
Keng-Chih Chang, Tsung-Jung Liu, Kuan-Hsien Liu, Day-Yu Chao
SMC2
2020 No-Reference Video Quality Assessment by A Cascade Combination of Neural Networks and Regression Model
abstract
In this paper, we propose a general-purpose no-reference (NR) video quality assessment (VQA) metric based on the cascade combination of 2D convolutional neural network (CNN), multi-layer perceptron (MLP), and support vector regression (SVR) model. The features are extracted from both spatial and spatiotemporal domains by using a 2D CNN. These features can capture different aspects of video frames for predicting quality scores, and we take these features as inputs of MLP to obtain a few estimated quality scores on different perspectives. Finally, these estimated scores are combined as a final quality score by an SVR model. The proposed method is evaluated on the well-known LIVE Video database with other state-of-the-art and well-performing VQA metrics. And the experimental result demonstrates that our method is competitive with other full-reference and NR VQA metrics.
Zheng-Lung Chu, Tsung-Jung Liu, Kuan-Hsien Liu
SMC2
2020 Clothing Brand Logo Prediction: From Residual Block to Dense Block
abstract
In this paper, we proposed a new clothing brand prediction method which is rooted on a dense-block based deep convolutional neural network for brand logo detection and recognition. To learn convolutional neural networks deeper and more accurately, we adopted dense blocks into deep convolutional neural networks to make connections between layers shorter. In this work, we propose several dense-block based designs to improve clothing brand logo detection and recognition accuracies. We also constructed a new large-scale clothing brand and price (CBP) dataset and its subset, called clothing brand logo (CBL) dataset with the brand attribute and logo information to carry out this task. To lower proposed framework complexity, two pixel search steps for the bounding box movement are implemented in the training procedure. In the experiment, we show our search reduced model can outperform several state-of-the-art methods and attain good performance.
Kuan-Hsien Liu, Tsung-Jung Liu
SMC2
2020 Identifying Poultry Farms from Satellite Images with Residual Dense U-Net
abstract
In this paper, we proposed a convolutional neural network called residual dense U-Net. This network is devised based on the original U-Net network. The encoder-decoder architecture in U-Net can restore the feature map to the resolution of the original image and obtain high-level semantic features. The skip-connection in U-Net can fuse the features after up-sampling and down-sampling to prevent both high-level semantic features and low-level semantic features from being lost after down-sampling. In the encoder and decoder parts, we utilize the residual dense block (RDB) from Residual Dense Network. Before each max-pooling, we replace the last convolutional layer in the original U-Net architecture with RDB. After each up-sampling, the last convolutional layer in the original U-Net architecture will also be replaced with RDB. The proposed method will be used to find poultry farms in Taiwan from satellite images. The prediction results will be evaluated using several indicators such as IOU, precision, recall, and F1-score.
Kai-Yu Wen, Tsung-Jung Liu, Kuan-Hsien Liu, Day-Yu Chao
SMC2
2019 Image Super-Resolution Using Complex Dense Block on Generative Adversarial Networks
abstract
The recent super-resolution (SR) techniques are divided into two directions. One is to improve PSNR and the other is to improve visual quality. We believe improving visual quality is more important and practical than blindly improving PSNR. In this paper we employ a generative adversarial network (GAN) and a new perceptual loss function for photo-realistic single image super-resolution (SISR). Our main contributions are as follows: we propose a new dense block which uses complex connections between each layer to build a more powerful generator. Next, to improve the perceptual quality, we found a new set of feature maps to compute the perceptual loss, which would make the output image look more real and natural. Finally, we compare our results with other methods by subjective evaluation. The subjects rank the image generated by various methods from good to bad. The final results show that our method can generate a more natural and realistic SR image than other state-of-the-art methods.
Bo-Xun Chen, Tsung-Jung Liu, Kuan-Hsien Liu, Hsin-Hua Liu, Soo-Chang Pei
ICIP2
2019 Image Inpainting For Random Areas Using Dense Context Features
abstract
Deep convolutional neural networks (DCNN) have demonstrated their potential to generate reasonable results in image inpainting. Some existing method uses convolution to generate surrounding features, then passes features by fully connected layers, and finally predicts missing regions. Although the final result is semantically reasonable, some blurred situations generated because the standard convolution is used, which conditioned on the effective pixels and the substitute values in the masked holes. In this paper, we introduce dense blocks for the U-Net architecture, which can alleviate the problem of gradient disappearance, while also reducing the number of parameters. The most important is that it can enhance the transfer of features and make more efficient use of them. Partial convolution is used to solve the problem of artifacts such as color differences and blurring. Experiments on the place365 dataset demonstrate our approach can generate more detailed and semantically reasonable results in random area image inpainting.
Yu-Zhe Su, Tsung-Jung Liu, Kuan-Hsien Liu, Hsin-Hua Liu, Soo-Chang Pei
ICIP2
2019 Modern Architecture Style Transfer for Ruin or Old Buildings
abstract
In this work, we focus on building style transfer, which transforms ruin or old buildings to modern architecture. Inspired by Gaty's and Goodfellow's style transfer and generative adversarial network (GAN), we use CycleGAN to conquer this type of problem. As we know, image style transfer usually generated unexpected artifacts. To avoid the artifacts and generate better images, we add so called “perception loss” into the network, which is the feature loss extracted by VGG pre-trained model. In the part of “cycle” structure, we adjust cycle loss by changing the ratio of weighting parameters. Finally, we collect images of both ruin (or old) and modern architecture from websites and use unsupervised learning to train the model. The experimental results show our proposed method indeed realize the modern architecture style transfer for ruin or old buildings.
Kuan-Hsien Liu, Tsung-Jung Liu, Chia-Ching Wang, Hsin-Hua Liu, Soo-Chang Pei
ISCAS2
2019 Spatial-Temporal Visual Attention Model for Video Quality Assessment
abstract
Objective assessment for videos has developed with mature technique. However, there are still some challenges, such as mimicking the behavior that human-beings do when they watch a video. In this paper, we introduce a model for full-reference (FR) video quality assessment (VQA) which is based on visual attention, optical flow, spatio-temporal slice (STS) images and center bias map. The experimental results show that our proposed model has better performance in wireless transmission distortion than other models in the LIVE video quality database.
Wei-Juen Suen, Hsin-Hua Liu, Soo-Chang Pei, Kuan-Hsien Liu, Tsung-Jung Liu
ISCAS5
2019 Face Aging on Realistic Photos by Generative Adversarial Networks
abstract
Inspired by Gatys and Goodfellow's style transfer and generative adversarial network (GAN), we use CycleGAN to achieve age progression. CycleGAN is good at generating fake images and also competitive with other GANs. It not only generates fake images but also increases the number of images in our database. We know the better database, the better performance of the model. We also try a deeper generator to transform youth photos to elder photos. To avoid the artifacts, we not only adopt the idea of “cycle” but also add a new loss which can tell the discriminator not too strict to generated images. Finally, we collect images of young and old people from the Internet and use unsupervised learning to train our model. The experimental results show our proposed method is indeed improved and better than before.
Chia-Ching Wang, Hsin-Hua Liu, Soo-Chang Pei, Kuan-Hsien Liu, Tsung-Jung Liu
ISCAS5
2019 Learning based no-reference metric for assessing quality of experience of stereoscopic images
Tsung-Jung Liu, Kuan-Hsien Liu, Kuan-Hung Shen
J. Vis. Commun. Image Represent.1
2019 A Structure-Based Human Facial Age Estimation Framework Under a Constrained Condition
abstract
Developing an automatic age estimation method towards human faces continues to possess an important role in computer vision and pattern recognition. Many studies regarding facial age estimation mainly focus on two aspects: facial aging feature extraction and classification/regression model learning. To set our work apart from existing age estimation approaches, we consider a different aspect -system structuring, which is, under a constrained condition: given a fixed feature type and a fixed learning method, how to design a framework to improve the age estimation performance based on the constraint? We propose a four-stage fusion framework for facial age estimation. This framework starts from gender recognition, and then go to the second phase, gender-specific age grouping, and followed by the third stage, age estimation within age groups, and finally ends at the fusion stage. In the experiment, three well-known benchmark datasets, MORPH-II, FG-NET, and CLAP2016, are adopted to validate the procedure. The experimental results show that the performance can be significantly improved by using our proposed framework and this framework also outperforms several state-of-the-art age estimation methods.
Kuan-Hsien Liu, Tsung-Jung Liu
IEEE Trans. Image Process.2
2018 No-Reference Image Quality Assessment by Wide-Perceptual-Domain Scorer Ensemble Method
abstract
A no-reference (NR) learning-based approach to assess image quality is presented in this paper. The devised features are extracted from wide perceptual domains, including brightness, contrast, color, distortion, and texture. These features are used to train a model (scorer) which can predict scores. The scorer selection algorithms are utilized to help simplify the proposed system. In the final stage, the ensemble method is used to combine the prediction results from selected scorers. Two multiple-scale versions of the proposed approach are also presented along with the single-scale one. They turn out to have better performances than the original single-scale method. Because of having features from five different domains at multiple image scales and using the outputs (scores) from selected score prediction models as features for multi-scale or cross-scale fusion (i.e., ensemble), the proposed NR image quality assessment models are robust with respect to more than 24 image distortion types. They also can be used on the evaluation of images with authentic distortions. The extensive experiments on three well-known and representative databases confirm the performance robustness of our proposed model.
Tsung-Jung Liu, Kuan-Hsien Liu
IEEE Trans. Image Process.1
2017 Visual quality prediction on distorted stereoscopic images
abstract
We propose a no-reference 3D image quality assessment (IQA) model that can automatically evaluate stereoscopic images. First, the model extracts statistical features from 2D single-view images (i.e., the stereopair) and their pseudo-disparity (i.e., absolute difference) map. Then the features are used to train an IQA model to predict the image quality score by the regression module of support vector machine (SVM). The model we proposed is tested on LIVE 3D Image Quality Database Phase I, which contains only symmetric-distorted stereoscopic images, and LIVE 3D Image Quality Database Phase II, which contains both symmetric-distorted and asymmetric-distorted stereoscopic images. The experimental results on both LIVE 3D Database Phase I and II show that our proposed model leads to significantly improved performance on quality prediction of stereoscopic images.
Ching-Ti Lin, Tsung-Jung Liu, Kuan-Hsien Liu
ICIP2
2017 A ParaBoost Method to Image Quality Assessment
abstract
An ensemble method for full-reference image quality assessment (IQA) based on the parallel boosting (ParaBoost) idea is proposed in this paper. We first extract features from existing image quality metrics and train them to form basic image quality scorers (BIQSs). Then, we select additional features to address specific distortion types and train them to construct auxiliary image quality scorers (AIQSs). Both BIQSs and AIQSs are trained on small image subsets of certain distortion types and, as a result, they are weak performers with respect to a wide variety of distortions. Finally, we adopt the ParaBoost framework, which is a statistical scorer selection scheme for support vector regression (SVR), to fuse the scores of BIQSs and AIQSs to evaluate the images containing a wide range of distortion types. This ParaBoost methodology can be easily extended to images of new distortion types. Extensive experiments are conducted to demonstrate the superior performance of the ParaBoost method, which outperforms existing IQA methods by a significant margin. Specifically, the Spearman rank order correlation coefficients (SROCCs) of the ParaBoost method with respect to the LIVE, CSIQ, TID2008, and TID2013 image quality databases are 0.98, 0.97, 0.98, and 0.96, respectively.
Tsung-Jung Liu, Kuan-Hsien Liu, Joe Yuchieh Lin, Weisi Lin, C.-C. Jay Kuo
IEEE Trans. Neural Networks Learn. Syst.1
2016 Age estimation via fusion of multiple binary age grouping systems
abstract
In this paper, we propose a new divide-and-conquer based method, called fusion of multiple binary age-grouping-estimation systems, for human facial age estimation. Under a specific constraint, such as a given facial feature or classification/regression method, what is the better framework for age estimation? First we employ multiple binary-grouping systems for age group classification. Each face image will be classified into one of the two groups. Within the two groups, two models are trained to estimate ages for the faces classified into their groups, respectively. We also investigate the effect of different age grouping systems on the performance of age grouping accuracy and age estimation error. In the last stage, we propose a sequentially selection algorithm to fuse some of the binary-grouping systems to get a final age estimation result. Experiments on the MORPH2 database demonstrate our framework for age estimation can achieve satisfying results and outperform other state-of-the-art age estimation approaches.
Tsung-Jung Liu, Kuan-Hsien Liu, Hsin-Hua Liu, Soo-Chang Pei
ICIP1
2015 Comparison of subjective viewing test methods for image quality assessment
abstract
This paper presents a comparison study on subjective quality scores obtained by both single stimulus (without reference) and triple stimulus (with reference) methods. The TID2013 database is reevaluated by single stimulus approach, which is realized by absolute category rating (ACR). And the mean opinion score (MOS) provided along with TID2013 represents the results from triple-stimulus pair comparison (3-stimulus PC) method. In the end, the correlation coefficient and hypothesis testing are used to determine if there is a significant difference between both sets of scores. The experimental results show that the differences exist and are significant for some specific distortion types or image contents. We believe the findings in this work can benefit and facilitate the future subjective visual quality viewing tests.
Tsung-Jung Liu, Kuan-Hsien Liu, Hsin-Hua Liu, Soo-Chang Pei
ICIP1
2015 Facial makeup detection via selected gradient orientation of entropy information
abstract
This work presents a novel facial makeup detection method, which includes four steps: entropy information computation, feature extraction, feature selection and classification. To carry out this objective, first all face images are subject to the entropy information computation. Once the entropy images of faces are obtained, a feature extraction step is applied to the entropy images instead of original face images. The extracted features are further processed to reduce the redundant information on the feature vector, which is done by a feature selection procedure. A statistical analysis approach is chosen to realize this feature selection purpose, which aims to lower the feature dimension and maintain higher discrimination. In the last step, the makeup is detected by classifying faces into two groups: makeup and no-makeup. The experimental results on two databases indeed demonstrate the superiority of the proposed method.
Kuan-Hsien Liu, Tsung-Jung Liu, Hsin-Hua Liu, Soo-Chang Pei
ICIP2
2015 MCL-V: A streaming video quality assessment database
Joe Yuchieh Lin, Rui Song 0003, Chihao Wu 0001, Tsung-Jung Liu, Haiqiang Wang, C.-C. Jay Kuo
J. Vis. Commun. Image Represent.4
2013 Image Quality Assessment Using Multi-Method Fusion
abstract
A new methodology for objective image quality assessment (IQA) with multi-method fusion (MMF) is presented in this paper. The research is motivated by the observation that there is no single method that can give the best performance in all situations. To achieve MMF, we adopt a regression approach. The new MMF score is set to be the nonlinear combination of scores from multiple methods with suitable weights obtained by a training process. In order to improve the regression results further, we divide distorted images into three to five groups based on the distortion types and perform regression within each group, which is called "context-dependent MMF" (CD-MMF). One task in CD-MMF is to determine the context automatically, which is achieved by a machine learning approach. To further reduce the complexity of MMF, we perform algorithms to select a small subset from the candidate method set. The result is very good even if only three quality assessment methods are included in the fusion process. The proposed MMF method using support vector regression is shown to outperform a large number of existing IQA methods by a significant margin when being tested in six representative databases.
Tsung-Jung Liu, Weisi Lin, C.-C. Jay Kuo
IEEE Trans. Image Process.1
2010 Temporal information assisted video quality metric for multimedia
abstract
This paper proposed a new objective video quality metric for multimedia videos based on a different perspective. We extended one existing image quality assessment metric to a video quality metric by considering temporal information and converted it into some compensation factor to correct the video quality score obtained in the spatial domain. After some experiments, we find out this proposed quality assessment metric does work well in the Laboratory for Image and Video Engineering (LIVE) Video Quality Database and is also competitive with the other existing state-of-the-art video quality assessment methods.
Tsung-Jung Liu, Kuan-Hsien Liu, Hsin-Hua Liu
ICME1
2010 A SIFT descriptor based method for global disparity vector estimation in multiview video coding
abstract
Disparity estimation is crucial to multiview video coding (MVC), which has attracted much attention recently. In the MVC reference software, named JMVM, the global disparity vector (GDV) was estimated by frame matching on spatial neighbor views. In this paper, a scale invariant feature transform (SIFT) based disparity estimation method is proposed to estimate the GDV. The experimental results show that benefits on peak signal-to-noise ratio (PSNR) and saved bits can be obtained by adopting our proposed method compared to the JMVM method that was often used.
Kuan-Hsien Liu, Tsung-Jung Liu, Hsin-Hua Liu
ICME2
2010 Color image watermarking using SVD
abstract
Digital watermark is an important technology for the image verification. Many authors researching on color image watermarking opine RGB color space in the spatial domain is not suitable for embedding marks. In this paper, we propose a digital color image watermarking method using Singular Value Decomposition (SVD). The whole process including embedding and extracting could be finished in RGB components in spatial domain. Simulation results show that our method is successful in resolving the rightful ownership of the watermarked image with good robustness, good imperceptibility and higher security.
Soo-Chang Pei, Hsin-Hua Liu, Tsung-Jung Liu, Kuan-Hsien Liu
ICME3