EDBT 2026 Demo / reviewers in the wild / expert
Kai-Lung Hua
dblp:08/268
· DBLP profile ↗
98ranked-venue papers
10as first author
37since 2021 · last 2026
0000-0002-7735-243XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 73 · 7 first-author · 25 since 2021Artificial intelligence and machine learning · 18 · 1 first-author · 13 since 2021Computer networks · 7 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dragonite: Single-Step Drag-based Image Editing with Geometric-Semantic GuidanceabstractRecent interactive image editing methods have made notable progress, yet achieving both precise control and real-time performance remains a challenge. Drag-based methods offer detailed geometric manipulations but suffer from low image fidelity and slow runtime performance, while text-based approaches enhance realism but limit precise and pixel-level control. To overcome these limitations, we introduce Dragonite, an intuitive and efficient framework that seamlessly unifies geometric and semantic manipulation for image editing. Dragonite leverages a Dual Guidance Module that fuses geometric deformation vectors with semantic guidance cues into a joint representation space, ensuring precise manipulation of both content and semantics. By combining a single-step latent optimization mechanism with a enhanced interpolation method, Dragonite achieves efficient interactive image editing while maintaining high precision through integrated geometric and semantic guidance. Extensive evaluations on the DragBench benchmark demonstrate that Dragonite effectively resolves the trade-off between speed and accuracy, enabling real-time, high-fidelity image editing. Meng-Ting Jhong, Tai-Ming Huang, Shung-Fu Chen, Wen-Huang Cheng, Kai-Lung Hua |
WACV | 5 |
| 2026 | DirectDrag: High-Fidelity, Mask-Free, Prompt-Free Drag-based Image Editing via Readout-Guided Feature AlignmentabstractDrag-based image editing using generative models provides intuitive control over image structures. However, existing methods rely heavily on manually provided masks and textual prompts to preserve semantic fidelity and motion precision. Removing these constraints creates a fundamental trade-off: visual artifacts without masks and poor spatial control without prompts. To address these limitations, we propose DirectDrag, a novel mask-and prompt-free editing framework. DirectDrag enables precise and efficient manipulation with minimal user input while maintaining high image fidelity and accurate point alignment. DirectDrag introduces two key innovations. First, we design an Auto Soft Mask Generation module that intelligently infers editable regions from point displacement, automatically localizing deformation along movement paths while preserving contextual integrity through the generative model’s inherent capacity. Second, we develop a Readout-Guided Feature Alignment mechanism that leverages intermediate diffusion activations to maintain structural consistency during point-based edits, substantially improving visual fidelity. Despite operating without manual mask or prompt, DirectDrag achieves superior image quality compared to existing methods while maintaining competitive drag accuracy. Extensive experiments on DragBench and real-world scenarios demonstrate the effectiveness and practicality of DirectDrag for high-quality, interactive image manipulation. Code is available at: https://github.com/frakw/DirectDrag. Sheng-Hao Liao, Shang-Fu Chen, Tai-Ming Huang, Wen-Huang Cheng, Kai-Lung Hua |
WACV | 5 |
| 2025 | Towards More General Video-based Deepfake Detection through Facial Component Guided Adaptation for Foundation ModelabstractThe current deep generative models have enabled the creation of synthetic facial images with remarkable photorealism, raising significant societal concerns over their potential misuse. Despite rapid advancements in the field of deepfake detection, developing an efficient and effective approach for the generalized deepfake detection of unseen forgery samples remains challenging. To address this challenge, we leverage the rich semantic priors of foundation models and propose a novel side-network-based decoder that extracts spatial and temporal cues using the CLIP image encoder for generalized video-based Deepfake detection. Additionally, we introduce Facial Component Guidance (FCG) to enhance spatial learning generalizability by encouraging the model to focus on key facial regions. By leveraging the generic features of a vision-language foundation model, our approach demonstrates promising generalizability on challenging Deepfake datasets while also exhibiting superiority in training data efficiency, parameter efficiency, and model robustness. The source code is available at: https://github.com/aiiu-lab/DFD-FCG Yue-Hua Han, Tai-Ming Huang, Kai-Lung Hua, Jun-Cheng Chen |
CVPR | 3 |
| 2025 | EXDF: Explainable Deepfake Detection with Vision-Language ModelabstractAlthough many deepfake detection methods have been proposed to fight against severe misuse of generative AI, none provide detailed human-interpretable explanations beyond simple real/fake responses. This limitation makes it challenging for humans to assess the accuracy of detection results, especially when the models encounter unseen deepfakes. To address this issue, we propose a novel deepfake detector based on a large Vision-Language Model (VLM), capable of explaining manipulated facial regions. We frame the deepfake detection task as Visual Question Answering (VQA) and perform visual instruction tuning to train the model on our collected Explainable Deepfake Face (ExDF) dataset. The dataset consists of fake images from diverse generative adversarial networks (GANs) and diffusion models (DMs), with explanations produced by GPT-4o guided by the corresponding ground-truth masks of the manipulated regions. Moreover, a facial mask encoder is introduced to guide the model to focus on key facial features, thereby improving the detection and explanation performances. Extensive experiments demonstrate that training the proposed model on the full ExDF dataset not only enhances detection accuracy compared to baseline methods but also provides detailed, human-interpretable explanations. To our knowledge, ExDF is the first explainable deepfake face dataset covering both GANs and DMs with comprehensive descriptions of altered facial regions. Our code and dataset are available at https://github.com/aiiu-lab/ExDF. Shu-Tzu Lo, Tai-Ming Huang, Yue-Hua Han, Kai-Lung Hua, Jun-Cheng Chen |
ICIP | 4 |
| 2025 | Hierarchical Spatiotemporal Fusion for Event-Visible Object DetectionabstractTraditional visible light cameras are prone to performance degradation under varying weather and lighting conditions. To address this challenge, we introduce an eventbased camera and propose a novel hierarchical spatiotemporal fusion approach for event-visible object detection. Our method enhances detection performance by integrating data from both event-based and visible light cameras. We have designed three key modules: The Gated Event Accumulation Representation module (GEAR), the Temporal Feature Selection module (TFS), and the Adaptive Fusion module (AF). GEAR and TFS enhance temporal feature fusion at both image and feature levels, while AF effectively integrates multi-modal features with low computational complexity. Our approach has been trained and validated on the publicly available DSEC-Detection dataset, achieving mAP50 and mAP50-95 scores of 67.2% and 45.6%, respectively, demonstrating superior detection performance and validating the effectiveness of the proposed method. Sin-Ye Jhong, Hsin-Chun Lin, Tzu-Chi Liu, Kai-Lung Hua, Yung-Yao Chen |
ICRA | 4 |
| 2025 | Deterministic Optimization-Based Path Planning Techniques for Obstacle Avoidance in Human-Robot Collaborative ScenariosabstractIn the context of Smart Manufacturing, robots are increasingly designed to operate alongside humans, with collaborative robots playing a central role. Ensuring safety in such Human-Robot Collaboration (HRC) scenarios require advanced path planning systems capable of detecting and responding to obstacles, including humans, in real time, thereby enabling safe and efficient cooperation. This paper presents a comparative study of optimization-based path planning techniques applied in collaborative robotics for obstacle avoidance. A structured comparison is conducted among different deterministic optimization strategies, such as Newton’s Method, Conjugate Gradient, and Gradient Descent, each employing various methods for obstacle pose identification, estimation, and danger factor modeling. Through an extensive review of these methods and their outcomes, the study evaluates their performance in terms of path safety, computational efficiency, and adaptability to dynamic environments. The analysis highlights the strengths and limitations of each optimization-based model and provides guidance for selecting suitable path planning approaches for different robotic applications. Brijesh Patel 0002, Yung-Chieh Chang, Po Ting Lin, Chao-Lung Yang, Yung-Yao Chen, Kai-Lung Hua, Meng-Kun Liu |
SMC | 6 |
| 2025 | Hybrid CNN-ViT architecture to exploit spatio-temporal feature for fire recognition trained through transfer learning
Hong-Cyuan Wang, Yung-Yao Chen, Kai-Lung Hua |
Multim. Tools Appl. | 4 |
| 2024 | Enhancing Anchor-Based Weakly Supervised Referring Expression Comprehension with Cross-Modality Attention
Ting-Yu Chu, Yong-Xiang Lin, Kai-Lung Hua |
ACCV (3) | 4 |
| 2024 | Representation and Boundary Enhancement for Action Segmentation Using TransformerabstractIn the task of action segmentation, the goal is to partition a lengthy, untrimmed video into a series of action segments. Recently, Transformer-based methods have outperformed the previous temporal convolutional networks (TCNs) in terms of overall performance. However, both TCNs and Transformers encounter the challenge of over-segmentation. Prior approaches often relied on post-processing techniques to address this issue, but these methods are not universally applicable to every model and may sometimes result in performance degradation. Therefore, in this paper, we propose a set of loss functions to enhance representation learning and employ a multi-task learning approach to strengthen the model’s ability to identify action boundaries. Through extensive experiments, we validate that our method demonstrates significant improvements, particularly in addressing the challenge of over-segmentation. Shang-Fu Chen, Cheng-Xun Wen, Wen-Huang Cheng, Kai-Lung Hua |
ICASSP | 4 |
| 2024 | Generalized Image-Based Deepfake Detection Through Foundation Model Adaptation
Tai-Ming Huang, Yue-Hua Han, Ernie Chu, Shu-Tzu Lo, Kai-Lung Hua, Jun-Cheng Chen |
ICPR (21) | 5 |
| 2024 | Indirect: invertible and discrete noisy image rescaling with enhancement from case-dependent texturesabstractAbstract Rescaling digital images for display on various devices, while simultaneously removing noise, has increasingly become a focus of attention. However, limited research has been done on a unified framework that can efficiently perform both tasks. In response, we propose INDIRECT (INvertible and Discrete noisy Image Rescaling with Enhancement from Case-dependent Textures), a novel method designed to address image denoising and rescaling jointly. INDIRECT leverages a jointly optimized framework to produce clean and visually appealing images using a lightweight model. It employs a discrete invertible network, DDR-Net, to perform rescaling and denoising through its reversible operations, efficiently mitigating the quantization errors typically encountered during downscaling. Subsequently, the Case-dependent Texture Module (CTM) is introduced to estimate missing high-frequency information, thereby recovering a clean and high-resolution image. Experimental results demonstrate that our method achieves competitive performance across three tasks: noisy image rescaling, image rescaling, and denoising, all while maintaining a relatively small model size. Huu-Phu Do, Yan-An Chen, Nhat-Tuong Do-Tran, Kai-Lung Hua, Wen-Hsiao Peng |
Multim. Syst. | 4 |
| 2024 | GRA: Graph Representation Alignment for Semi-Supervised Action RecognitionabstractGraph convolutional networks (GCNs) have emerged as a powerful tool for action recognition, leveraging skeletal graphs to encapsulate human motion. Despite their efficacy, a significant challenge remains the dependency on huge labeled datasets. Acquiring such datasets is often prohibitive, and the frequent occurrence of incomplete skeleton data, typified by absent joints and frames, complicates the testing phase. To tackle these issues, we present graph representation alignment (GRA), a novel approach with two main contributions: 1) a self-training (ST) paradigm that substantially reduces the need for labeled data by generating high-quality pseudo-labels, ensuring model stability even with minimal labeled inputs and 2) a representation alignment (RA) technique that utilizes consistency regularization to effectively reduce the impact of missing data components. Our extensive evaluations on the NTU RGB+D and Northwestern-UCLA (N-UCLA) benchmarks demonstrate that GRA not only improves GCN performance in data-constrained environments but also retains impressive performance in the face of data incompleteness. Kuan-Hung Huang, Yao-Bang Huang, Yong-Xiang Lin, Kai-Lung Hua, Muhammad Tanveer 0001, Xuequan Lu, Muhammad Imran Razzak |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Multi-Task Self-Blended Images for Face Forgery DetectionabstractDeepfake detection has attracted extensive attention due to widespread forged images on social media. Recently, self-supervised learning (SSL) based Deepfake detection approaches have outperformed supervised methods in terms of model generalization. However, we notice that most SSL-based methods do not take the manipulation strength levels of synthesized forgery samples into consideration according to different synthesis parameters and result in suboptimal detection performances. To address this issue, we introduce several auxiliary losses to the state-of-the-art SSL-based method based on different synthesis sub-tasks during data generation by inferring their synthesis parameters where the ground-truth labels are obtained from the synthesis pipeline for free. With comprehensive evaluations on various benchmarks, our approach has achieved noticeable performance improvement. Specifically, for the cross-dataset evaluation, the proposed approach outperforms the state-of-the-art method in terms of AUC on various datasets with improvements of 3.4%, 1.47%, 1.56%, and 1.3% on the CDF, DFDC, DFDCP, and FFIW datasets and achieves competitive performance on the DFD dataset. This further demonstrates the effectiveness of the proposed approach in its generalization ability. Yue-Hua Han, Ernie Chu, Jun-Cheng Chen, Kai-Lung Hua |
MMAsia | 5 |
| 2023 | SSRFace: a face recognition framework against shallow data
Yun-Ting Zhou, Jilyan Bianca Dy, Shang-Che Hsu, Yu-Ling Hsu, Chao-Lung Yang, Kai-Lung Hua |
Multim. Tools Appl. | 6 |
| 2023 | MCGAN: mask controlled generative adversarial network for image retargeting
Jilyan Bianca Dy, John Jethro Virtusio, Daniel Stanley Tan, Yong-Xiang Lin, Joel P. Ilao, Yung-Yao Chen, Kai-Lung Hua |
Neural Comput. Appl. | 7 |
| 2023 | DEFAEK: Domain Effective Fast Adaptive Network for Face Anti-Spoofing
Jiun-Da Lin, Yue-Hua Han, Julianne Tan, Jun-Cheng Chen, Muhammad Tanveer 0001, Kai-Lung Hua |
Neural Networks | 7 |
| 2023 | MACnet: Mask augmented counting network for class-agnostic counting
Tadhg McCarthy, John Jethro Virtusio, Jose Jaena Mari Ople, Daniel Stanley Tan, Divina Amalin, Kai-Lung Hua |
Pattern Recognit. Lett. | 6 |
| 2023 | ConCoNet: Class-agnostic counting with positive and negative exemplarsabstractClass-agnostic counting is usually phrased as a matching problem between a user-defined exemplar patch and a query image. The count is derived based on the number of objects similar to the exemplar patch. However, defining a target class using only positive exemplar patches inevitably miscounts unintended objects that are visually alike to the exemplar. In this paper, we propose to include negative exemplars that define what not to count. This allows the model to calibrate its notion of what is similar based on both positive and negative exemplars. It effectively disentangles visually similar negatives, leading to a more discriminative definition of the target object. We designed our method such that it can be incorporated with other class-agnostic counting models. Moreover, application-wise, our model can be used into a semi-automatic labeling tool to simplify the job of the annotator Adrienne Francesca O. Soliven, John Jethro Virtusio, Jose Jaena Mari Ople, Daniel Stanley Tan, Divina Amalin, Kai-Lung Hua |
Pattern Recognit. Lett. | 6 |
| 2023 | Controllable Model Compression for Roadside Camera Depth EstimationabstractIn the Cooperative Intelligent Transportation System (C-ITS) paradigm, vehicles could communicate with roadside units to augment their traffic knowledge. Smart roadside units could provide second-order information (e.g., vehicle count) from raw first-order data (e.g., visual feed, point clouds), and this “smart” feature is usually provided using deep neural network models. However, implementing these useful models implies a cost for computational complexity that could hinder the future deployment of smart roadside units needed for sustainability in transportation systems. In this paper, we propose to use model compression on deep image processing models to promote its feasibility for usage in smart sensors. We formulated a controllable convolutional model compression (CCMC) algorithm that can perform filter-wise evolutionary pruning on image processing networks, along with a predefined compression ratio. CCMC is applicable for image processing networks, which have multiple possible traffic data sources (e.g., road camera surveillance). Furthermore, CCMC has a definable target compression ratio that is useful for controlling the trade-off between resource consumption and output performance. We tested our proposed method on depth estimation, which is useful for scene understanding and mapping the locations of objects in the 3D space. Our experiments show that the pruned model has minimal performance discrepancy from the original one, supporting the sustainability features needed for intelligent transportation systems. Jose Jaena Mari Ople, Shang-Fu Chen, Yung-Yao Chen, Kai-Lung Hua, Mohammad Hijji, Po Yang 0001, Khan Muhammad 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Adjustable Model Compression Using Multiple Genetic AlgorithmabstractGenerative Adversarial Networks (GAN) is a popular machine learning method that possesses powerful image generation ability, which is useful for different multimedia applications (e.g., photographic filters, image editing). However, typical GAN models have a large memory footprint that limits their practical applications for resource-constrained devices (e.g., smartphones). To deploy GAN models on devices with various hardware constraints, we propose our method, AdjustableGAN, which can compress a pretrained GAN model to different compression ratios. Our method compresses GAN by performing filter-wise pruning that follows these objectives: (1) deactivate convolutional filters for minimal performance decrease, (2) reactivate convolutional filters for maximal performance increase. We implement multiple Genetic Algorithms (GA) to perform each of these objectives— Downsize GA for best filter deactivations, while Upsize GA searches for best filter reactivations. By selective utilization of Upsize/Downsize GA, we could explicitly control the compression ratio of the model. For finalization, we fine-tune the compressed output model using the training dataset of the original input model. Our experimental results show that our method can reliably compress generative networks with minimal accuracy drop compared to other state-of-the-art compression algorithms. Jose Jaena Mari Ople, Tai-Ming Huang, Ming-Chih Chiu, Yi-Ling Chen 0002, Kai-Lung Hua |
IEEE Trans. Multim. | 5 |
| 2022 | PixMamba: Leveraging State Space Models in a Dual-Level Architecture for Underwater Image Enhancement
Wei-Tung Lin, Yong-Xiang Lin, Jyun-Wei Chen, Kai-Lung Hua |
ACCV (4) | 4 |
| 2022 | CLIPCAM: A Simple Baseline For Zero-Shot Text-Guided Object And Action LocalizationabstractThe key for the contemporary deep learning-based object and action localization algorithms to work is the large-scale annotated data. However, in real-world scenarios, since there are infinite amounts of unlabeled data beyond the categories of publicly available datasets, it is not only time- and manpower-consuming to annotate all the data but also requires a lot of computational resources to train the detectors. To address these issues, we show a simple and reliable baseline that can be easily obtained and work directly for the zero-shot text-guided object and action localization tasks without introducing additional training costs by using Grad-CAM, the widely used class visual saliency map generator, with the help of the recently released Contrastive Language-Image Pre-Training (CLIP) model by OpenAI, which is trained contrastively using the dataset of 400 million image-sentence pairs with rich cross-modal information between text semantics and image appearances. With extensive experiments on the Open Images and HICO-DET datasets, the results demonstrate the effectiveness of the proposed approach for the text-guided unseen object and action localization tasks for images. Hsuan-An Hsia, Che-Hsien Lin, Bo-Han Kung, Jhao-Ting Chen, Daniel Stanley Tan, Jun-Cheng Chen, Kai-Lung Hua |
ICASSP | 7 |
| 2022 | Code generation from a graphical user interface via attention-based encoder-decoder model
Wen-Yin Chen, Pavol Podstreleny, Wen-Huang Cheng, Yung-Yao Chen, Kai-Lung Hua |
Multim. Syst. | 5 |
| 2022 | Correction to: HoloTube: a low-cost portable 360-degree interactive autostereoscopic display
Che-Hao Hsu, Yi-Leh Wu, Wen-Huang Cheng, Kai-Lung Hua |
Multim. Tools Appl. | 5 |
| 2022 | VDNet: video deinterlacing network based on coarse adaptive module and deformable recurrent residual network
Yin-Chen Yeh, Jilyan Bianca Dy, Tai-Ming Huang, Yung-Yao Chen, Kai-Lung Hua |
Neural Comput. Appl. | 5 |
| 2022 | Controllable and Identity-Aware Facial Attribute TransformationabstractModifying facial attributes without the paired dataset proves to be a challenging task. Previously, approaches either required supervision from a ground-truth transformed image or required training a separate model for mapping every pair of attributes. These limit the scalability of the models to accommodate a larger set of attributes since the number of models that we need to train grows exponentially large. Another major drawback of the previous approaches is the unintentional gain of the identity of the person as they transform the facial attributes. We propose a method that allows for controllable and identity-aware transformations across multiple facial attributes using only a single model. Our approach is to train a generative adversarial network (GAN) with a multitask conditional discriminator that recognizes the identity of the face, distinguishes real images from fake, as well as identifies facial attributes present in an image. This guides the generator into producing an output that is realistic while preserving the person's identity and facial attributes. Through this framework, our model also learns meaningful image representations in a lower dimensional latent space and semantically associate separate parts of the encoded vector with both the person's identity and facial attributes. This opens up the possibility of generating new faces and other transformations such as making the face thinner or chubbier. Furthermore, our model only encodes the image once and allows for multiple transformations using the encoded vector. This allows for faster transformations since it does not need to reprocess the entire image for every transformation. We show the effectiveness of our proposed method through both qualitative and quantitative evaluations, such as ablative studies, visual inspection, and face verification. Competitive results are achieved compared to the main competition (CycleGAN), however, at great space and extensibility gain by using a single model. Daniel Stanley Tan, Jonathan Hans Soeseno, Kai-Lung Hua |
IEEE Trans. Cybern. | 3 |
| 2022 | Lightweight Face Anti-Spoofing Network for Telehealth ApplicationsabstractOnline healthcare applications have grown more popular over the years. For instance, telehealth is an online healthcare application that allows patients and doctors to schedule consultations, prescribe medication, share medical documents, and monitor health conditions conveniently. Apart from this, telehealth can also be used to store a patient's personal and medical information. With its rise in usage due to COVID-19, given the amount of sensitive data it stores, security measures are necessary. A simple way of making these applications more secure is through user authentication. One of the most common and often used authentications is face recognition. It is convenient and easy to use. However, face recognition systems are not foolproof. They are prone to malicious attacks like printed photos, paper cutouts, replayed videos, and 3D masks. The goal of face anti-spoofing is to differentiate real users (live) from attackers (spoof). Although effective in terms of performance, existing methods use a significant amount of parameters, making them resource-heavy and unsuitable for handheld devices. Apart from this, they fail to generalize well to new environments like changes in lighting or background. This paper proposes a lightweight face anti-spoofing framework that does not compromise on performance. Our proposed method achieves good performance with the help of an ArcFace Classifier (AC). The AC encourages differentiation between spoof and live samples by making clear boundaries between them. With clear boundaries, classification becomes more accurate. We further demonstrate our model's capabilities by comparing the number of parameters, FLOPS, and performance with other state-of-the-art methods. Jiun-Da Lin, Hung-Hsiang Lin, Jilyan Bianca Dy, Jun-Cheng Chen, Muhammad Tanveer 0001, Muhammad Imran Razzak, Kai-Lung Hua |
IEEE J. Biomed. Health Informatics | 7 |
| 2022 | Classification of Alzheimer's Disease Using Ensemble of Deep Neural Networks Trained Through Transfer LearningabstractAlzheimer's disease (AD) is one of the deadliest neurodegenerative diseases ailing the elderly population all over the world. An ensemble of Deep learning (DL) models can learn highly complicated patterns from MRI scans for the detection of AD by utilizing diverse solutions. In this work, we propose a computationally efficient, DL-architecture agnostic, ensemble of deep neural networks, named 'Deep Transfer Ensemble (DTE)' trained using transfer learning for the classification of AD. DTE leverages the complementary feature views and diversity introduced by many different locally optimum solutions reached by individual networks through the randomization of hyper-parameters. DTE achieves an accuracy of 99.05% and 85.27% on two independent splits of the large dataset for cognitively normal (NC) vs AD classification task. For the task of mild cognitive impairment (MCI) vs AD classification, DTE achieves 98.71% and 83.11% respectively on the two independent splits. It also performs reasonable on a small dataset consisting of only 50 samples per class. It achieved a maximum accuracy of 85% for NC vs AD on the small dataset. It also outperformed snapshot ensembles along with several other existing deep models from similar kind of previous works by other researchers. Muhammad Tanveer 0001, Ashraf Haroon Rashid, M. A. Ganaie 0001, Motahar Reza, Muhammad Imran Razzak, Kai-Lung Hua |
IEEE J. Biomed. Health Informatics | 6 |
| 2021 | Naturalistic Physical Adversarial Patch for Object DetectorsabstractMost prior works on physical adversarial attacks mainly focus on the attack performance but seldom enforce any restrictions over the appearance of the generated adversarial patches. This leads to conspicuous and attention-grabbing patterns for the generated patches which can be easily identified by humans. To address this issue, we pro-pose a method to craft physical adversarial patches for object detectors by leveraging the learned image manifold of a pretrained generative adversarial network (GAN) (e.g., BigGAN and StyleGAN) upon real-world images. Through sampling the optimal image from the GAN, our method can generate natural looking adversarial patches while maintaining high attack performance. With extensive experiments on both digital and physical domains and several independent subjective surveys, the results show that our proposed method produces significantly more realistic and natural looking patches than several state-of-the-art base-lines while achieving competitive attack performance.1 Yu-Chih-Tuan Hu, Jun-Cheng Chen, Bo-Han Kung, Kai-Lung Hua, Daniel Stanley Tan |
ICCV | 4 |
| 2021 | Fire Detection using Transformer NetworkabstractTechnological breakthroughs in computing have empowered vision-based surveillance systems to detect fire using transformers framework. Over the last few decades, convolutional neural networks (CNNs) have been broadly applied for many computer vision-related problems and provided satisfactory results. However, due to the inductive prejudices embedded in convolutional operations, it cannot comprehend long-range dependencies. Vision transformers (ViT) has recently become an alternative to CNN for a vision problem by factoring an image as a patches sequence and leverage intra-attention between pixels. This paper shows that ViT is a viable tool for automated fire detection by aggregating features from the whole spatial context. The proposed method is tested on benchmark fire datasets to reveal the framework's strength and effectiveness. Kai-Lung Hua |
ICMR | 2 |
| 2021 | Cascaded atrous dual attention U-Net for tumor segmentation
Wannaporn Sarapugdi, Yong-Xiang Lin, Jyh-Cheng Chen, Kai-Lung Hua |
Multim. Tools Appl. | 6 |
| 2021 | Deep spatial-temporal networks for flame detection
I-Feng Chien, Wannaporn Sarapugdi, Lili Miao, Kai-Lung Hua |
Multim. Tools Appl. | 5 |
| 2021 | Incremental Learning of Multi-Domain Image-to-Image TranslationsabstractCurrent multi-domain image-to-image translation models assume a fixed set of domains and that all the data are always available during training. However, over time, we may want to include additional domains to our model. Existing methods either require re-training the whole model with data from all domains or require training several additional modules to accommodate new domains. To address these limitations, we present IncrementalGAN, a multi-domain image-to-image translation model that can incrementally learn new domains using only a single generator. Our approach first decouples the domain label representation from the generator to allow it to be re-used for new domains without any architectural modification. Next, we introduce a distillation loss that prevents the model from forgetting previously learned domains. Our model compares favorably against several state-of-the-art baselines. Daniel Stanley Tan, Yong-Xiang Lin, Kai-Lung Hua |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Dress With Style: Learning Style From Joint Deep Embedding of Clothing Styles and Body ShapesabstractBody shape is about proportion, and fashion style is all about dressing those proportions to look their very best. Figuring out the styles to suit a body shape can be a daunting task for many people. It is, therefore, essential to develop a framework for learning the compatibility of body shapes and clothing styles. Though fashion designers and fashion stylists have analyzed the correlation between human body shapes and fashion styles for a long time, this issue did not receive much attention in multimedia science. In this paper, we present a novel style recommender, on the basis of the user's body attributes. The rich amount of fashion styling knowledge from social big data is exploited for this purpose. We first construct a joint embedding of clothing styles and human body measurements with deep multimodal representation learning on a reference dataset that has been sorted to meet the fashion rules. We then discover the relevant semantic features by propagation and selection in clothing style and body shape graphs. Experiments demonstrate the effectiveness of the proposed framework when compared with several baseline methods. Shintami Chusnul Hidayati, Ting Wei Goh, Ji-Sheng Gary Chan, Cheng-Chun Hsu, John See, Lai-Kuan Wong, Kai-Lung Hua, Yu Tsao 0001, Wen-Huang Cheng |
IEEE Trans. Multim. | 7 |
| 2021 | Neural Style Palette: A Multimodal and Interactive Style Transfer From a Single Style ImageabstractDespite the myriad of attributes found in a single style image, existing neural style transfer methods produce outputs with limited variety–typically only a single realization of the style image. They also do not provide an easy way to control the stylization process, limiting the creative freedom of users. In this paper, we propose Neural Style Palette (NSP), a method for interactively generating a variety of stylized images from only a single style input. Our approach allows human influence in the stylization process, a design inspired by Hybrid Human-Artificial Intelligence. Like a color palette,NSPenables a meaningful interaction by presenting a collection of sub-textures, which we also refer to as anchor styles, that act as a visual guide for the users. These anchor styles capture different attributes in the single style image that the users can creatively blend to create their desired realizations. To offer a diversified selection in theNSP, we constrain the anchor styles to be distant from one another while maintaining faithfulness to the original style image. This is possible through our two proposed novel losses: a style-separation loss that encourages the sub-textures to be distinct and a unification loss to ensure that the sub-textures center around the original style while encouraging additional diversity. We perform several experiments to prove the effectiveness of our method and generalize to improve existing methods. John Jethro Virtusio, Jose Jaena Mari Ople, Daniel Stanley Tan, Muhammad Tanveer 0001, Neeraj Kumar 0001, Kai-Lung Hua |
IEEE Trans. Multim. | 6 |
| 2021 | Enabling Artistic Control Over Pattern Density and Stroke StrengthabstractDespite the remarkable results and numerous advancements in neural style transfer, achieving artistic control is still a challenging feat, primarily since existing methodologies treat the style representation as a black-box model. This oversight significantly limits the range of possible artistic manipulations. In this paper, we propose a method to enable artistic control on any correlation-based style transfer models along with guiding intuitions. Our focus is on controlling two perceptual factors: Pattern Density and Stroke Strength. To achieve this, we introduce the centered Gram style representation and manipulate it with our variance-aware adaptive weighting and correlation-based selective masking. Through several experiments and comparisons with the state-of-the-art, we show that we can achieve artistic control with competitive stylization quality. Additionally, since our method involves manipulating style representation, it can easily be adapted to popular style transfer models. We analyze different style representation properties to propose rules that govern the style transfer process, which is critical towards achieving artistic control over pattern density and stroke strength. John Jethro Virtusio, Daniel Stanley Tan, Wen-Huang Cheng, Muhammad Tanveer 0001, Kai-Lung Hua |
IEEE Trans. Multim. | 5 |
| 2021 | Explainable AI: A Multispectral Palm-Vein Identification System with New Augmentation FeaturesabstractRecently, as one of the most promising biometric traits, the vein has attracted the attention of both academia and industry because of its living body identification and the convenience of the acquisition process. State-of-the-art techniques can provide relatively good performance, yet they are limited to specific light sources. Besides, it still has poor adaptability to multispectral images. Despite the great success achieved by convolutional neural networks (CNNs) in various image understanding tasks, they often require large training samples and high computation that are infeasible for palm-vein identification. To address this limitation, this work proposes a palm-vein identification system based on lightweight CNN and adaptive multi-spectral method with explainable AI. The principal component analysis on symmetric discrete wavelet transform (SMDWT-PCA) technique for vein images augmentation method is adopted to solve the problem of insufficient data and multispectral adaptability. The depth separable convolution (DSC) has been applied to reduce the number of model parameters in this work. To ensure that the experimental result demonstrates accurately and robustly, a multispectral palm image of the public dataset (CASIA) is also used to assess the performance of the proposed method. As result, the palm-vein identification system can provide superior performance to that of the former related approaches for different spectrums. Yung-Yao Chen, Sin-Ye Jhong, Chih-Hsien Hsia, Kai-Lung Hua |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2020 | Super-Resolution by Image Enhancement Using Texture TransferabstractRecent deep learning approaches in single image super-resolution (SISR) can generate high-definition textures for super-resolved (SR) images. However, they tend to hallucinate fake textures and even produce artifacts. An alternative to SISR, reference-based SR (RefSR) approaches use high-resolution (HR) reference (Ref) images to provide HR details that are missing in the low-resolution (LR) input image. We propose a novel framework that leverages existing SISR approaches and enhances them with RefSR. Specifically, we refine the output of SISR methods using neural texture transfer, where HR features are queried from the Ref images. The query is conducted by computing the similarity of textural and semantic features between the input image and the Ref images. The most similar HR features, patch-wise, to the LR image is used to augment the SR image through an augmentation network. In the case of dissimilar Ref images from the LR input image, we prevent performance degradation by including the similarity scores in the input features of the network. Furthermore, we use random texture patches during the training to condition our augmentation network to not always trust the queried texture features. Different from past RefSR approaches, our method can use arbitrary Ref images and its lower-bound performance is based on the SR image. We showcase that our method drastically improves the performance of the base SISR approach. Jose Jaena Mari Ople, Daniel Stanley Tan, Arnulfo P. Azcarraga, Chao-Lung Yang, Kai-Lung Hua |
ICIP | 5 |
| 2020 | Pairwise Adjacency Matrix on Spatial Temporal Graph Convolution Network for Skeleton-Based Two-Person Interaction RecognitionabstractSpatial-temporal graph convolutional networks (ST-GCN) have achieved outstanding performances on human action recognition, however, it might be less superior on a two-person interaction recognition (TPIR) task due to the relationship of each skeleton is not considered. In this study, we present an improvement of the STGCN model that focused on TPIR by employing the pairwise adjacency matrix to capture the relationship of person-person skeletons (ST-GCN-PAM). To validate the effectiveness of the proposed ST-GCN-PAM model on TPIR, experiments were conducted on NTU RGB+D120. Additionally, the model was also examined on the Kinetics dataset and NTU RGB+D60. The results show that the proposed ST-GCN-PAM outperforms the-state-of-the-art methods on mutual action of NTU RGB+D120 by achieving 83.28% (cross-subject) and 88.31% (cross-view) accuracy. The model is also superior to the original ST-GCN on the multi-human action of the Kinetics dataset by achieving 41.68% in Top-l and 88.91% in Top-5. Chao-Lung Yang, Aji Setyoko, Hendrik Tampubolon, Kai-Lung Hua |
ICIP | 4 |
| 2020 | Artist-based painting classification using Markov random fields with convolution neural network
Kai-Lung Hua, Trang-Thi Ho, Kevin Alfianto Jangtjik, Mei-Chen Yeh |
Multim. Tools Appl. | 1 |
| 2020 | Sketch-guided Deep Portrait GenerationabstractGenerating a realistic human class image from a sketch is a unique and challenging problem considering that the human body has a complex structure that must be preserved. Additionally, input sketches often lack important details that are crucial in the generation process, hence making the problem more complicated. In this article, we present an effective method for synthesizing realistic images from human sketches. Our framework incorporates human poses corresponding to locations of key semantic components (e.g., arm, eyes, nose), seeing that its a strong prior for generating human class images. Our sketch-image synthesis framework consists of three stages: semantic keypoint extraction, coarse image generation, and image refinement. First, we extract the semantic keypoints using Part Affinity Fields (PAFs) and a convolutional autoencoder. Then, we integrate the sketch with semantic keypoints to generate a coarse image of a human. Finally, in the image refinement stage, the coarse image is enhanced by a Generative Adversarial Network (GAN) that adopts an architecture carefully designed to avoid checkerboard artifacts and to generate photo-realistic results. We evaluate our method on 6,300 sketch-image pairs and show that our proposed method generates realistic images and compares favorably against state-of-the-art image synthesis methods. Trang-Thi Ho, John Jethro Virtusio, Yung-Yao Chen, Chih-Ming Hsu, Kai-Lung Hua |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2019 | Fuzzy Personalized Scoring Model for Recommendation SystemabstractIn this research, we aim to propose a data preprocessing framework particularly for financial sector to generate the rating data as input to the collaborative system. First, clustering technique is applied to cluster all users based on their demographic information which might be able to differentiate the customers' background. Then, for each customer group, the importance of demographic characteristics which are highly associated with financial products purchasing are analyzed by the proposed fuzzy integral technique. The importance scores across items and customers are generated either on customer groups and individuals. The analysis shows the proposed method is able to differentiate customers based on their demographic and purchasing behaviors. Also, the generated rating matrix can be directly used for collaborative filtering model. Chao-Lung Yang, Shang-Che Hsu, Kai-Lung Hua, Wen-Huang Cheng |
ICASSP | 3 |
| 2019 | Single-Fusion Detector: Towards Faster Multi-Scale Object DetectionabstractDespite recent improvements, the arbitrary sizes of objects still impede the predictive ability of object detectors. Recent solutions combine feature maps of different receptive fields to detect multi-scale objects. However, these methods have large computational costs resulting to slower inference time, which is not practical for real-time applications. Contrarily, fusion methods depending on large networks with many skip connections demand larger memory requirement, prohibiting usage in devices with limited memory. In this paper, we propose a more computationally efficient fusion method which integrates higher-order information to low-level feature maps using a single operation. Our method can flexibly adapt to any base network, allowing tailored performance for different computational requirements. Our approach achieves 81.7% mAP at 41 FPS on the PASCAL VOC dataset using ResNet-50 as the base network, which is superior in terms of both speed and mAP as compared to several state-of-the-art baselines, even those which use larger base networks. Arren Matthew C. Antioquia, Daniel Stanley Tan, Arnulfo P. Azcarraga, Kai-Lung Hua |
ICIP | 4 |
| 2019 | Smart: A Sensor-Triggerred Interactive Mr DisplayabstractTo increase the user experience in mixed reality (MR) applications, researchers proposed to display virtual objects in mirrors and transmissive mirror devices. For example, [1] proposed to detect the moving object hold by a user in the real world and display the smoke virtual visual effect in a mirror with an MR experience. On the other hand, [2] developed the MRsionCase to display the text information near the static sculpture. In recently years, researchers paid attention to display the virtual contents in head-mount-display (HMD) devices. To name a few, [3] proposed to use a VR keyboard model with MR real hands to enhance an immersive experience. Furthermore, [4] proposed to use an Ultrahaptics to generate a sensing experience on a finger of a user in an MR application. However, to represent consistent displaying effect from a virtual space to a real space (the physical world) is still a challenging research issue. In this paper, a prototype with a hologram display is developed as an interactive MR device, triggered by the sensor values on mobile devices to change the attitude of a virtual object in the virtual world and the real object in a physical world with a synchronizing manner. Y. S. Lan, C. D. Chen, S. W. Sun, W. C. Yen, Y. T. Wang, Y. H. Yang, J. M. Day, Kai-Lung Hua |
ICIP | 8 |
| 2019 | Spatially-Aware Domain Adaptation for Semantic Segmentation of Urban ScenesabstractIt is very expensive and time consuming to collect a large enough dataset with pixel-level annotations to train a semantic segmentation model. Synthetic datasets are common alternatives for training segmentation models, however models trained on synthetic data do not necessarily perform well on real world images due to the domain shift problem. Domain adaptation techniques address this problem by leveraging on adversarial training to align features. Prior works have mostly performed global feature alignment. They do not consider the positions of objects. However, objects in urban scenes are highly correlated with their spatial locations. For example, the sky will always appear on top while cars will usually appear in the middle of the image. Based on this insight, we propose a spatial-aware discriminator that accounts for the spatial prior on the objects in order to improve the feature alignment. We demonstrate in our experiments that our model outperforms several state-of-the-art baselines in terms of mean intersection over union (mIoU). Yong-Xiang Lin, Daniel Stanley Tan, Wen-Huang Cheng, Yung-Yao Chen, Kai-Lung Hua |
ICIP | 5 |
| 2019 | Segmenting Hepatic Lesions Using Residual Attention U-Net with an Adaptive Weighted Dice LossabstractWe propose a novel network architecture called Residual Attention U-Net (ResAttU-Net) for segmenting hepatic lesions. Our model incorporates residual blocks that can extract more complex features as compared with traditional convolutional layers combined with a skip-connection attention module that learns to focus on the relevant features for the task of hepatic lesions segmentation. Moreover, we train our model using an adaptive weighted dice loss that prioritizes the pixels of the tumor class over the pixels of the background class. We evaluate our model on the MICCAI Liver Tumor Segmentation (LiTS) benchmark dataset. Our experimental results show that our method significantly improves upon several state-of-the-art baselines for hepatic lesion or liver tumor segmentation. Daniel Stanley Tan, Jyh-Cheng Chen, Wen-Huang Cheng, Kai-Lung Hua |
ICIP | 5 |
| 2019 | Layout and Context Understanding for Image Synthesis with Scene GraphsabstractAdvancements on text-to-image synthesis generate remarkable images from textual descriptions. However, these methods are designed to generate only one object with varying attributes. They face difficulties with complex descriptions having multiple arbitrary objects since it would require information on the placement and sizes of each object in the image. Recently, a method that infers object layouts from scene graphs has been proposed as a solution to this problem. However, their method uses only object labels in describing the layout, which fail to capture the appearance of some objects. Moreover, their model is biased towards generating rectangular shaped objects in the absence of ground-truth masks. In this paper, we propose an object encoding module to capture object features and use it as additional information to the image generation network. We also introduce a graph-cuts based segmentation method that can infer the masks of objects from bounding boxes to better model object shapes. Our method produces more discernible images with more realistic shapes as compared to the images generated by the current state-of-the-art method. Arces Talavera, Daniel Stanley Tan, Arnulfo P. Azcarraga, Kai-Lung Hua |
ICIP | 4 |
| 2019 | Adapting Semantic Segmentation of Urban Scenes via Mask-Aware Gated DiscriminatorabstractTraining a deep neural network for semantic segmentation relies on pixel-level ground truth labels for supervision. However, collecting large datasets with pixel-level annotations is very expensive and time consuming. One workaround is to utilize synthetic data where we can generate potentially unlimited data with their corresponding ground truth labels. Unfortunately, networks trained on synthetic data perform poorly on real images due to the domain shift problem. Domain adaptation techniques have shown potential in transferring the knowledge learned from synthetic data to real world data. Prior works have mostly leveraged on adversarial training to perform a global aligning of features. However, we observed that background objects have lesser variations across different domains as opposed to foreground objects. Using this insight, we propose a method for domain adaptation that models and adapts foreground objects and background objects separately. Our approach starts with a fast style transfer to match the appearance of the inputs. This is followed by a foreground adaptation module that learns a foreground mask that is used by our gated discriminator in order to adapt the foreground and background objects separately. We demonstrate in our experiments that our model outperforms several state-of-the-art baselines in terms of mean intersection over union (mIoU). Yong-Xiang Lin, Daniel Stanley Tan, Wen-Huang Cheng, Kai-Lung Hua |
ICME | 4 |
| 2019 | 3D Object Completion via Class-Conditional Generative Adversarial Network
Yu-Chieh Chen, Daniel Stanley Tan, Wen-Huang Cheng, Kai-Lung Hua |
MMM (2) | 4 |
| 2018 | Pedestrian Detection from Lidar Data via Cooperative Deep and Hand-Crafted FeaturesabstractAutopilot systems need to be able to detect pedestrians with high precision and recall regardless of whether it is during the day or night. This means that we cannot rely on normal cameras to sense the surroundings due to its sensitivity to lighting conditions. An alternative for images is to use light detection and ranging sensors (LiDAR) that produces three-dimensional point clouds where each point represents the distance to an object. However, most pedestrian detection systems are designed for image inputs and not on distance point clouds. In this paper, we propose a method for detecting pedestrians using only the three-dimensional point clouds generated by the LiDAR. Our approach first projects the three-dimensional point cloud into a two-dimensional plane. We then extract both hand-crafted features and learned features from a convolutional neural network in order to train a support vector machine (SVM) to detect pedestrians. Our proposed method achieved significant improvements in terms of F1-measurement over prior state-of-the-art methods. Tzu-Chieh Lin, Daniel Stanley Tan, Hsueh-Ling Tang, Shih-Che Chien, Feng-Chia Chang, Yung-Yao Chen, Wen-Huang Cheng, Kai-Lung Hua |
ICIP | 8 |
| 2018 | What Dress Fits Me Best?: Fashion Recommendation on the Clothing Style for Personal Body ShapeabstractClothing is an integral part of life. Also, it is always an uneasy task for people to make decisions on what to wear. An essential style tip is to dress for the body shape, i.e., knowing one's own body shape (e.g., hourglass, rectangle, round and inverted triangle) and selecting the types of clothes that will accentuate the body's good features. In the literature, although various fashion recommendation systems for clothing items have been developed, none of them had explicitly taken the user's basic body shape into consideration. In this paper, therefore, we proposed a first framework for learning the compatibility of clothing styles and body shapes from social big data, with the goal to recommend a user about what to wear better in relation to his/her essential body attributes. The experimental results demonstrate the superiority of our proposed approach, leading to a new aspect for research into fashion recommendation. Shintami Chusnul Hidayati, Cheng-Chun Hsu, Yu-Ting Chang, Kai-Lung Hua, Jianlong Fu, Wen-Huang Cheng |
ACM Multimedia | 4 |
| 2018 | ZipNet: ZFNet-level Accuracy with 48× Fewer ParametersabstractWith the introduction of Convolutional Neural Networks, models for image classification achieve higher classification accuracy. Based on the pattern of the design of CNN architectures, increasing the number of layers equates to a higher classification accuracy, but also increases the number of parameters and model size. This negatively affects the model training time, processing time, and memory requirement. We develop ZipNet, a CNN architecture with a higher classification accuracy than ZFNet, the winner of ILSVRC 2013, but with 48.5× smaller model size and 48.7× fewer parameters. The classification accuracy of ZipNet is higher than the performance of ZFNet and SqueezeNet on all configurations of the Caltech-256 dataset with varying number of training examples. Arren Matthew C. Antioquia, Daniel Stanley Tan, Arnulfo P. Azcarraga, Wen-Huang Cheng, Kai-Lung Hua |
VCIP | 5 |
| 2018 | Vehicle Detection in Thermal Images Using Deep Neural NetworkabstractIn today's world, it becomes critical for a self-driving car to detect the vehicles irrespective of it being a day or night. We propose a real-time vehicle detection using a sequence of night-time thermal images. Moreover, the thermal images have the capability of retaining even the minuscule vehicle details in a dim environment. For an efficient vehicle detection, the thermal image dataset collected during the dusk and night is used for training purposes. Subsequently, the contrast enhancement and sharpening of these images are performed using the Thermal Feature Enhancement (TFE). Then the concatenated images are supplied as the input to allow the model to learn more effectively. Besides, we also propose an improved convolution network model entitled as the Thermal Image Only Looked Once (TOLO) model for vehicle detection. Additionally, we propose a method called as Low Probability Candidate Filter (LPCF) to compensate the probability of not-easy-to-detect vehicles. Our proposed method produces better results for the F1-measure in comparison with existing methods. Chin-Wei Chang, Kathiravan Srinivasan, Yung-Yao Chen, Wen-Huang Cheng, Kai-Lung Hua |
VCIP | 5 |
| 2018 | A Cloud-based Intelligent Skin and Scalp Analysis SystemabstractThe love of beauty is an essential part of all healthy human nature. Not only do women pay great attention to facial care, but in recent years, men's consumption in this area has also grown year by year. In facial care, in addition to sunscreen, skin care, tattoos and other chemical-based skin care products, mechanical skin washing machine has also become one of the most popular beauty appliances in recent years. In order to learn the effectiveness of various face care products and tools, in this paper, we have proposed an intelligent system that integrates face washing, wireless cameras, smart phones, and cloud image analysis functionality to allow users to obtain product recommendations. The proposed system is easy to carry and provides seven analysis functionalities, such as skin color, pigmentation, skin texture, wrinkles, pores, texture analysis, and would produce an in-depth analysis report. The proposed system is not only helpful for personal beauty care, but also beneficial for professional medical clinics and makeup companies. The collected big data will also be utilized as a reference for future cosmetics, and potential medical and biotechnology company product development and marketing. Wen-Shiung Huang, Bing-Kai Hong, Wen-Huang Cheng, Shih-Wei Sun, Kai-Lung Hua |
VCIP | 5 |
| 2018 | Interactive Style Transfer: Towards Styling User-Specified ObjectabstractResearches dealing with the task of style transfer have focused and produced results that only transfers the style of a reference image to the entirety of another image. As this domain would be beneficial in digital art, it would be preferable for such algorithms to support a more flexible style transfer, such that the style would be applied only on a specific portion of an image. We propose a framework that can selectively apply a given style onto an object using only 4 user-defined points. Our approach combines a style transfer module and an object segmentation module to synthesize the stylized image. As the ultimate goal is to develop an artistic application tool, we also introduce a method that makes use of specific filters to preserve certain characteristics of an image, such as its high frequency components. Experiments show that our proposed method is able to produce pleasing results with minimal effort while also being able to handle more complicated tasks such as the application of multiple reference style images onto different objects in an image. John Jethro Virtusio, Arces Talavera, Daniel Stanley Tan, Kai-Lung Hua, Arnulfo P. Azcarraga |
VCIP | 4 |
| 2018 | Robust RGB-D Hand Tracking Using Deep Learning PriorsabstractWith the irruption of inexpensive depth sensor devices, hand gesture tracking has become a topic of great interest. Two main problems to face respect other tracking algorithms are the high complexity of the hand structure, which translate in a very large amount of possible gestures, and the rapidness of the movements we are able to make when moving the hand or just the fingers. Recent approaches try to fit a 3D hand model to the observed RGB-D data by an optimization function that minimizes the error between the model and the data. However, these algorithms are very dependent on the initialization point, which are impractical to run in a natural environment. To solve these kinds of problems, it is common to use an offline data set with prelearned gestures that will serve as a first rough estimate. In concrete, we present an algorithm that uses an articulated ICP minimization function that is initialized by the parameters obtained from a data set of hand gestures trained through a deep learning framework. This setup has two strong points. First, deep learning provides a very fast and accurate estimate of performed hand gestures. Second, the articulated ICP algorithm allows capturing the possible variability of a gesture performed by different persons or slightly different gestures. Our proposed algorithm is evaluated and validated in several ways. Independent evaluations for the deep learning framework and articulated ICP are performed. Moreover, different real sequences are recorded to validate our approach and, finally, quantitative and qualitative comparisons are conducted with state-of-the-art algorithms. Jordi Sanchez-Riera, Kathiravan Srinivasan, Kai-Lung Hua, Wen-Huang Cheng, M. Anwar Hossain 0001, Mohammed F. Alhamid |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Background Extraction Based on Joint Gaussian Conditional Random FieldsabstractBackground extraction is generally the first step in many computer vision and augmented reality applications. Most existing methods, which assume the existence of a clean background during the reconstruction period, are not suitable for video sequences such as highway traffic surveillance videos, whose complex foreground movements may not meet the assumption of a clean background. Therefore, we propose a novel joint Gaussian conditional random field (JGCRF) background extraction algorithm for estimating the optimal weights of frame composition for a fixed-view video sequence. A maximum a posteriori problem is formulated to describe the intra- and inter-frame relationships among all pixels of all frames based on their contrast distinctness and spatial and temporal coherence. Because all background objects and elements are assumed to be static, patches that are motionless are good candidates for the background. Therefore, in the algorithm method, a motionless extractor is designed by computing the pixel-wise differences between two consecutive frames and thresholding the accumulation of variation across the frames to remove possible moving patches. The proposed JGCRF framework can flexibly link extracted motionless patches with desired fusion weights as extra observable random variables to constrain the optimization process for more consistent and robust background extraction. The results of quantitative and qualitative experiments demonstrated the effectiveness and robustness of the proposed algorithm compared with several state-of-the-art algorithms; the proposed algorithm also produced fewer artifacts and had a lower computational cost. Hong-Cyuan Wang, Yu-Chi Lai, Wen-Huang Cheng, Chin-Yun Cheng, Kai-Lung Hua |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2018 | Learning and Recognition of Clothing Genres From Full-Body ImagesabstractAccording to the theory of clothing design, the genres of clothes can be recognized based on a set of visually differentiable style elements, which exhibit salient features of visual appearance and reflect high-level fashion styles for better describing clothing genres. Instead of using less-discriminative low-level features or ambiguous keywords to identify clothing genres, we proposed a novel approach for automatically classifying clothing genres based on the visually differentiable style elements. A set of style elements, that are crucial for recognizing specific visual styles of clothing genres, were identified based on the clothing design theory. In addition, the corresponding salient visual features of each style element were identified and formulated with variables that can be computationally derived with various computer vision algorithms. To evaluate the performance of our algorithm, a dataset containing 3250 full-body shots crawled from popular online stores was built. Recognition results show that our proposed algorithms achieved promising overall precision, recall, and -score of 88.76%, 88.53%, and 88.64% for recognizing upperwear genres, and 88.21%, 88.17%, and 88.19% for recognizing lowerwear genres, respectively. The effectiveness of each style element and its visual features on recognizing clothing genres was demonstrated through a set of experiments involving different sets of style elements or features. In summary, our experimental results demonstrate the effectiveness of the proposed method in clothing genre recognition. Shintami Chusnul Hidayati, Chuang-Wen You, Wen-Huang Cheng, Kai-Lung Hua |
IEEE Trans. Cybern. | 4 |
| 2018 | Background Extraction Using Random Walk Image FusionabstractIt is important to extract a clear background for computer vision and augmented reality. Generally, background extraction assumes the existence of a clean background shot through the input sequence, but realistically, situations may violate this assumption such as highway traffic videos. Therefore, our probabilistic model-based method formulates fusion of candidate background patches of the input sequence as a random walk problem and seeks a globally optimal solution based on their temporal and spatial relationship. Furthermore, we also design two quality measures to consider spatial and temporal coherence and contrast distinctness among pixels as background selection basis. A static background should have high temporal coherence among frames, and thus, we improve our fusion precision with a temporal contrast filter and an optical-flow-based motionless patch extractor. Experiments demonstrate that our algorithm can successfully extract artifact-free background images with low computational cost while comparing to state-of-the-art algorithms. Kai-Lung Hua, Hong-Cyuan Wang, Chih-Hsiang Yeh, Wen-Huang Cheng, Yu-Chi Lai |
IEEE Trans. Cybern. | 1 |
| 2018 | Edge-Preserving Depth Map Upsampling by Joint Trilateral FilterabstractCompared to the color images, their associated depth images captured by the RGB-D sensors are typically with lower resolution. The task of depth map super-resolution (SR) aims at increasing the resolution of the range data by utilizing the high-resolution (HR) color image, while the details of the depth information are to be properly preserved. In this paper, we present a joint trilateral filtering (JTF) algorithm for depth image SR. The proposed JTF first observes context information from the HR color image. In addition to the extracted spatial and range information of local pixels, our JTF further integrates local gradient information of the depth image, which allows the prediction and refinement of HR depth image outputs without artifacts like textural copies or edge discontinuities. Quantitative and qualitative experimental results demonstrate the effectiveness and robustness of our approach over prior depth map upsampling works. Kai-Han Lo, Yu-Chiang Frank Wang, Kai-Lung Hua |
IEEE Trans. Cybern. | 3 |
| 2018 | DeepDemosaicking: Adaptive Image Demosaicking via Multiple Deep Fully Convolutional NetworksabstractConvolutional neural networks are currently the state-of-the-art solution for a wide range of image processing tasks. Their deep architecture extracts low and high-level features from images, thus, improving the model's performance. In this paper, we propose a method for image demosaicking based on deep convolutional neural networks. Demosaicking is the task of reproducing full color images from incomplete images formed from overlaid color filter arrays on image sensors found in digital cameras. Instead of producing the output image directly, the proposed method divides the demosaicking task into an initial demosaicking step and a refinement step. The initial step produces a rough demosaicked image containing unwanted color artifacts. The refinement step then reduces these color artifacts using deep residual estimation and multi-model fusion producing a higher quality image. Experimental results show that the proposed method outperforms several existing and state-of-the-art methods in terms of both subjective and objective evaluations. Daniel Stanley Tan, Wei-Yang Chen, Kai-Lung Hua |
IEEE Trans. Image Process. | 3 |
| 2017 | A CNN-LSTM framework for authorship classification of paintingsabstractThe authenticity of digital painting image is an urgent demand in the field of art. Yet, determining the authorship of a certain painting is a challenging task due to two reasons: (1) various artists might share similar painting styles; and (2) an artist could create different styles. In this paper, we present a novel method for authorship classification of paintings based on a CNN-LSTM framework. First, a multiscale pyramid is constructed from a painting image. Second, a CNN-LSTM model is learned and it returns possibly multiple labels for one image. To aggregate the final classification result, an adaptive fusion method is employed. Experimental results show that the proposed method has superior classification performance compared with the state-of-the-art techniques. Kevin Alfianto Jangtjik, Trang-Thi Ho, Mei-Chen Yeh, Kai-Lung Hua |
ICIP | 4 |
| 2017 | Multi-cue pedestrian detection from 3D point cloud dataabstractPedestrian detection is one of the key technologies of driver assistance system. In order to prevent potential collisions, pedestrians should be always accurately identified whether during the day or at night. Since the visual images of the night are not clear, this paper proposes a method for recognizing pedestrians by using a high-definition LIDAR without visual images. In order to handle the long-distance sparse point problem, a novel solution is introduced to improve the performance. The proposed method maps the three-dimensional point cloud to the two-dimensional plane by a distance-aware expansion approach and the corresponding 2D contour and its associated 2D features are then extracted. Based on both 2D and 3D cues, the proposed method obtains significant performance boosts over state-of-the-art approaches by 13% in terms of F1-measure. Hsueh-Ling Tang, Shih-Che Chien, Wen-Huang Cheng, Yung-Yao Chen, Kai-Lung Hua |
ICME | 5 |
| 2017 | Fashion World Map: Understanding Cities Through Streetwear FashionabstractFashion is an integral part of life. Streets as a social center for people's interaction become the most important public stage to showcase the fashion culture of a metropolitan area. In this paper, therefore, we propose a novel framework based on deep neural networks (DNN) for depicting the street fashion of a city by automatically discovering fashion items (e.g., jackets) in a particular look that are most iconic for the city, directly from a large collection of geo-tagged street fashion photos. To obtain a reasonable collection of iconic items, our task is formulated as the prize-collecting Steiner tree (PCST) problem, whereby a visually intuitive summary of the world's iconic street fashion can be created. To the best of our knowledge, this is the first work devoted to investigate the world's fashion landscape in modern times through the visual analytics of big social data. It shows how the visual impression of local fashion cultures across the world can be depicted, modeled, analyzed, compared, and exploited. In the experiments, our approach achieves the best performance (43.19%) on our large collected GSFashion dataset (170K photos), with an average of two times higher than all the other algorithms (FII: 20.13%, AP: 18.76%, DC: 17.90%), in terms of the users' agreement ratio on the discovered iconic fashion items of a city. The potential of our proposed framework for advanced sociological understanding is also demonstrated via practical applications. Yu-Ting Chang, Wen-Huang Cheng, Bo Wu 0018, Kai-Lung Hua |
ACM Multimedia | 4 |
| 2017 | Popularity Meter: An Influence- and Aesthetics-aware Social Media Popularity PredictorabstractSocial media websites have become an important channel for content sharing and communication between users on social networks. The shared images on the websites, even the ones from the same user, tend to receive a quite diverse distribution of views. This raises the problem of image popularity prediction on social media. To address this important research topic, we explore three essential components that have considerable impact of the image popularity, which are user profile, post metadata, and photo aesthetics. Moreover, we make use of state-of-the-art predictive modeling approaches to demonstrate the effectiveness of our proposed features in predicting image popularity. We then evaluate the proposed method through a large number of real image posts from Flickr. The experimental results show significant statistical evidence that incorporating the proposed features with ensemble learning method that combines predictions from support vector regression (SVR) and classification and regression tree (CART) models offers a satisfactory popularity prediction. By understanding the social behavior and the underlying structure of content popularity, our research results can also contribute to designing better algorithms for important applications like content recommendation and advertisement placement. Shintami Chusnul Hidayati, Yi-Ling Chen 0002, Chao-Lung Yang, Kai-Lung Hua |
ACM Multimedia | 4 |
| 2017 | Human Pose Tracking Using Online Latent Structured Support Vector Machine
Kai-Lung Hua, Irawati Nurmala Sari, Mei-Chen Yeh |
MMM (1) | 1 |
| 2017 | i-Stylist: Finding the Right Dress Through Your Social Networks
Jordi Sanchez-Riera, Jun-Ming Lin, Kai-Lung Hua, Wen-Huang Cheng, Arvin Wen Tsui |
MMM (1) | 3 |
| 2017 | Toward an easy deployable outdoor parking system - Lessons from long-term deploymentabstractData pertaining to the availability of parking slots is crucial to the efficient operation of systems designed to monitor the state of parking spaces. Outdoor parking systems have been developed using wireless sensors, Internet of Things (IoT) technology, and cameras. Unfortunately, interference from electromagnetic fields complicates the tuning of parameters for detection algorithms and limits accuracy to only 90 percent. In this study, we investigated these problems by collecting data from magnetic sensors, light sensors, and LoRa wireless modules used in the detection transient events (car arrivals and departures) over a period of 13 months. This led to the design an adaptive occupancy detection system using a variety of sensors, which can be deployed with only minimal calibration. Yi-Chao Chen 0001, Chuang-Wen You, Dian-Xuan Wu, Yi-Ling Chen 0006, Kai-Lung Hua, Yung-Jen Hsu 0001 |
PerCom | 6 |
| 2017 | Intelligent deployment of UAVs in 5G heterogeneous communication environment for improved coverage
Vishal Sharma 0001, Kathiravan Srinivasan, Han-Chieh Chao, Kai-Lung Hua, Wen-Huang Cheng |
J. Netw. Comput. Appl. | 4 |
| 2017 | HoloTabletop: an anamorphic illusion interactive holographic-like tabletop system
Che-Hao Hsu, Wen-Huang Cheng, Kai-Lung Hua |
Multim. Tools Appl. | 3 |
| 2017 | CrossbowCam: a handheld adjustable multi-camera system
Che-Hao Hsu, Wen-Huang Cheng, Yi-Leh Wu, Wen-Hsiung Huang, Tao Mei 0001, Kai-Lung Hua |
Multim. Tools Appl. | 6 |
| 2017 | HoloTube: a low-cost portable 360-degree interactive autostereoscopic display
Che-Hao Hsu, Yi-Leh Wu, Wen-Huang Cheng, Kai-Lung Hua |
Multim. Tools Appl. | 5 |
| 2017 | Multicast scheduling for stereoscopic video in wireless networks
Kai-Lung Hua, Yeni Anistyasari, Che-Hao Hsu, Tai-Lin Chin, Chao-Lung Yang, Chun-Yen Wang |
Multim. Tools Appl. | 1 |
| 2017 | Efficient cooperative relaying in flying ad hoc networks using fuzzy-bee colony optimization
Vishal Sharma 0001, Kathiravan Srinivasan, Rajesh Kumar 0013, Han-Chieh Chao, Kai-Lung Hua |
J. Supercomput. | 5 |
| 2016 | Artist-based Classification via Deep Learning with Multi-scale Weighted PoolingabstractFor analyzing digital images of paintings we propose a new approach to categorize them based on artist. Determining the authorship of a painting is challenging because common subjects are illustrated in paintings, and paintings of an artist may not have a unique style. The proposed approach is built upon convolutional neural networks (CNN)---a class of biologically inspired vision model that recently demonstrates near-human performance on several visual recognition tasks. However, training a CNN model requires large scale training data of a fixed input image size (e.g. 224 * 224). In this paper, we propose to construct a multi-layer pyramid from an image, providing 21X more features than using a single layer (i.e., the original image) alone. We train a CNN model for each layer, and propose a new weighted fusion scheme to adaptively combine the decision results. To evaluate the proposed methods, we collect a new painting image dataset, categorized into 13 artists. As demonstrated in the experimental results, the proposed method achieves a promising result---88.08% recall rate in top-2 retrieval on the challenging classification task. Kevin Alfianto Jangtjik, Mei-Chen Yeh, Kai-Lung Hua |
ACM Multimedia | 3 |
| 2016 | Locality Constrained Sparse Representation for Cat Recognition
Shintami Chusnul Hidayati, Wen-Huang Cheng, Min-Chun Hu 0001, Kai-Lung Hua |
MMM (2) | 5 |
| 2016 | Photo sundial: Estimating the time of capture in consumer photos
Tsung-Hung Tsai, Wei-Cih Jhou, Wen-Huang Cheng, Min-Chun Hu 0001, I-Chao Shen, Tekoing Lim, Kai-Lung Hua, Ahmed Ghoneim, M. Anwar Hossain 0001, Shintami Chusnul Hidayati |
Neurocomputing | 7 |
| 2016 | Context-aware joint dictionary learning for color image demosaicking
Kai-Lung Hua, Shintami Chusnul Hidayati, Fang-Lin He, Chia-Po Wei, Yu-Chiang Frank Wang |
J. Vis. Commun. Image Represent. | 1 |
| 2016 | A comparative study of data fusion for RGB-D based visual recognition
Jordi Sanchez-Riera, Kai-Lung Hua, Yuan-Sheng Hsiao, Tekoing Lim, Shintami Chusnul Hidayati, Wen-Huang Cheng |
Pattern Recognit. Lett. | 2 |
| 2016 | Erratum to: Geometry-shader-based real-time voxelization and applications
Shu-Huai Chang, Yu-Chi Lai, Chih-Yuan Yao, Kai-Lung Hua, Yuzhen Niu, Feng Liu 0015 |
Vis. Comput. | 4 |
| 2015 | VRank: Voting system on Ranking model for human age estimationabstractRanking algorithms have proven the potential for human age estimation. Currently, a common paradigm is to compare the input face with reference faces of known age to generate a ranking relation whereby the first-rank reference is exploited for labeling the input face. In this paper, we proposed a framework to improve upon the typical ranking model, called Voting system on Ranking model (VRank), by leveraging relational information (comparative relations, i.e. if the input face is younger or older than each of the references) to make a more robust estimation. Our approach has several advantages: firstly, comparative relations can be explicitly involved to benefit the estimation task; secondly, few incorrect comparisons will not influence much the accuracy of the result, making this approach more robust than the conventional approach; finally, we propose to incorporate the deep learning architecture for training, which extracts robust facial features for increasing the effectiveness of classification. In comparison to the best results from the state-of-the-art methods, the VRank showed a significant outperformance on all the benchmarks, with a relative improvement of 5.74% ~ 69.45% (FG-NET), 19.09% ~ 68.71% (MORPH), and 0.55% ~ 17.73% (IoG). Tekoing Lim, Kai-Lung Hua, Hong-Cyuan Wang, Kai-Wen Zhao, Min-Chun Hu 0001, Wen-Huang Cheng |
MMSP | 2 |
| 2015 | Poster: Exploring the Need for Sensor Learning and Collaboration in IoT-based Parking SystemsabstractThe need to find parking contributes to road congestion and leads to unnecessary fuel consumption. Of all emerging parking systems, Internet-of-Things (IoT)-based systems have demonstrated the feasibility of real-time delivery of parking availability using magnetic sensors. However, existing magnetic-based methods are prone to false positives caused by electromagnetic fields emitted from surrounding electric facilities. In this study, we conducted a 3-month data collection in a parking area. We identified the need to introduce learning and collaboration into the design of our detection algorithm which recognizes learned patterns associated with car arrivals or departures, and to filter out unreliable events based on spatial and temporal features. Dian-Xuan Wu, Chuang-Wen You, Chi-Ling Yang, Seng-Yong Lau, Kai-Lung Hua, Wen-Huang Cheng, Yi-Ling Chen 0006, Yung-Jen Hsu 0001 |
SenSys | 6 |
| 2015 | An efficient pitch-by-pitch extraction algorithm through multimodal information
Kai-Lung Hua, Chao-Ting Lai, Chuang-Wen You, Wen-Huang Cheng |
Inf. Sci. | 1 |
| 2014 | What are the Fashion Trends in New York?abstractFashion is a reflection of the society of a period. Given that New York City is one of the world's fashion capitals, understanding its change in fashion becomes a way to know the society and the times. To keep up with fashion trends, it is important to know what's " in" and what's "out" for a season. Though the fashion trends have been analyzed by fashion designers and fashion analysts for a long time, this issue has been ignored in multimedia science. In this paper, we present a novel algorithm that automatically discovers visual style elements representing fashion trends for a certain season. The visual style elements are discovered based on the stylistic coherent and unique characteristics. The experimental results demonstrate the effectiveness of our proposed method through a large number of catwalk show videos. Shintami Chusnul Hidayati, Kai-Lung Hua, Wen-Huang Cheng, Shih-Wei Sun |
ACM Multimedia | 2 |
| 2014 | Who's the Best Charades Player? Mining Iconic Movement of Semantic Concepts
Yung-Huan Hsieh, Shintami Chusnul Hidayati, Wen-Huang Cheng, Min-Chun Hu 0001, Kai-Lung Hua |
MMM (1) | 5 |
| 2014 | LaRED: a large RGB-D extensible hand gesture datasetabstractWe present the LaRED, a Large RGB-D Extensible hand gesture Dataset, recorded with an Intel's newly-developed short range depth camera. This dataset is unique and differs from the existing ones in several aspects. Firstly, the large volume of data recorded: 243, 000 tuples where each tuple is composed of a color image, a depth image, and a mask of the hand region. Secondly, the number of different classes provided: a total of 81 classes (27 gestures in 3 different rotations). Thirdly, the extensibility of dataset: the software used to record and inspect the dataset is also available, giving the possibility for future users to increase the number of data as well as the number of gestures. Finally, in this paper, some experiments are presented to characterize the dataset and establish a baseline as the start point to develop more complex recognition algorithms. The LaRED dataset is publicly available at: http://mclab.citi.sinica.edu.tw/dataset/lared/lared.html. Yuan-Sheng Hsiao, Jordi Sanchez-Riera, Tekoing Lim, Kai-Lung Hua, Wen-Huang Cheng |
MMSys | 4 |
| 2014 | A novel multi-focus image fusion algorithm based on random walks
Kai-Lung Hua, Hong-Cyuan Wang, Aulia Hakim Rusdi, Shin-Yi Jiang |
J. Vis. Commun. Image Represent. | 1 |
| 2014 | Geometry-shader-based real-time voxelization and applications
Hsu-Huai Chang, Yu-Chi Lai, Chin-Yuan Yao, Kai-Lung Hua, Yuzhen Niu, Feng Liu 0015 |
Vis. Comput. | 4 |
| 2013 | Depth map super-resolution via Markov Random Fields without texture-copying artifactsabstractThe use of time-of-flight sensors enables the record of full-frame depth maps at video frame rate, which benefits a variety of 3D image or video processing applications. However, such depth maps are typically corrupted by noise and with limited resolution. In this paper, we present a learning-based depth map super-resolution framework by solving a MRF labeling optimization problem. With the captured depth map and the associated high-resolution color image, our proposed method exhibits the capability of preserving the edges of range data while suppressing the artifacts of texture copying due to color discontinuities. Quantitative and qualitative experimental results demonstrate the effectiveness and robustness of our approach over prior depth map upsampling works. Kai-Han Lo, Kai-Lung Hua, Yu-Chiang Frank Wang |
ICASSP | 2 |
| 2013 | Physiognomy master: a novel personality analysis system based on facial featuresabstractIn this demo, we present the proposed "Physiognomy Master." It is a novel practical personality analysis system based on facial features. We first design five facial features that are essential for face reading. We then construct a database to record the facial features' values from a number of volunteers. In the meantime, the volunteers are also invited to fill out a professional personality test. The relations between the facial features and the personality traits are then learned. Given a test subject or an input frontal face image, the proposed system will produce the associated personality report by fusing the personality scores from the people who have similar facial features in the constructed database. The fusing mechanism is based on the idea that people with similar facial features possess similar personality characteristics. The proposed system is a powerful tool in numerous kinds of social interactions, such as personnel selection, team composition, and marriage matching. Che-Hao Hsu, Kai-Lung Hua, Wen-Huang Cheng |
ACM Multimedia | 2 |
| 2013 | Joint trilateral filtering for depth map super-resolutionabstractDepth map super-resolution is an emerging topic due to the increasing needs and applications using RGB-D sensors. Together with the color image, the corresponding range data provides additional information and makes visual analysis tasks more tractable. However, since the depth maps captured by such sensors are typically with limited resolution, it is preferable to enhance its resolution for improved recognition. In this paper, we present a novel joint trilateral filtering (JTF) algorithm for solving depth map super-resolution (SR) problems. Inspired by bilateral filtering, our JTF utilizes and preserves edge information from the associated high-resolution (HR) image by taking spatial and range information of local pixels. Our proposed further integrates local gradient information of the depth map when synthesizing its HR output, which alleviates textural artifacts like edge discontinuities. Quantitative and qualitative experimental results demonstrate the effectiveness and robustness of our approach over prior depth map upsampling works. Kai-Han Lo, Yu-Chiang Frank Wang, Kai-Lung Hua |
VCIP | 3 |
| 2013 | Collaborative sequential detection in surveillance sensor networksabstractTarget detection is an important problem in wireless sensor networks where a number of sensors form a network to detect the presence or absence of a certain target or event. Data fusion is a potential method broadly used to improve detection performance when the sampling data are noisy. However, low detection probability cannot be avoided if detection decisions are made based on a collection of sampling data taken at just one particular moment. This paper adopts fusion-based sequential detection to guarantee the quality of detection results. A fusion center is used to collect local data from individual sensors periodically. A final detection decision is made only after the pre-defined constraints of false alarm and missing probability are satisfied. Rules for each sensor to make local decisions and for the fusion center to make global decisions are derived. Simulations are conducted to show the latency of making the final decisions based on the proposed fusion scheme. Tai-Lin Chin, Kai-Lung Hua, Tien-Ruey Hsiang, Ge-Ming Chiu, Shiow-Yang Wu |
WCNC | 2 |
| 2013 | Dependency-aware quality-differentiated wireless video multicastabstractVideo multicast exploits the wireless broadcast nature to transmit a video stream to multiple clients with a minimum bandwidth requirement. Assigning a suitable transmission bit-rate to each scalable coded block in a video stream is however a challenging problem because clients in a wireless network have heterogeneous channel quality and experience different packet loss probability. Prior work attempts to transmit the base-layer stream at a low transmission bit-rate to ensure a high reception probability, and hence the basic visual quality. Such methods however are over-simplified for a video stream that supports multiple quality levels in a video frame and needs explicit rate assignment for each block. We propose in this paper a dependency-aware rate scheduling scheme that assigns each block a rate according to dependency between blocks. With consideration of block dependency, we can better utilize limited wireless bandwidth to deliver important blocks, and minimize the number of undecodable blocks due to the loss of their reference blocks at the receivers. The simulation results show that since our scheme reduces the number of undecodable blocks, it achieves a higher overall video quality for a multicast group than the existing schemes under different client distributions. Han-Chiang Li, Kate Ching-Ju Lin, Kai-Lung Hua, Ge-Ming Chiu, Yu-Chin Tsai, Shan Chin |
WCNC | 3 |
| 2013 | An efficient scheduling algorithm for scalable video streaming over P2P networks
Kai-Lung Hua, Ge-Ming Chiu, Hsing-Kuo Kenneth Pao, Yi-Chi Cheng |
Comput. Networks | 1 |
| 2012 | Self-learning approach to color demosaicking via support vector regressionabstractMost digital cameras capture one primary color at each pixel by a single sensor overlaid with a color filter array. To recover a full color image from incomplete color samples, one needs to restore the two missing color values for each pixel. This restoration process is known as color demosaicking. In this paper, we present a novel self-learning approach to this problem via support vector regression. Unlike prior learning-based demosaicking methods, our approach aims at extracting image-dependent information in constructing the learning model, and we do not require any additional training data. Experimental results show that our proposed method outperforms many state-of-the-art techniques in both subjective and objective image quality measures. Fang-Lin He, Yu-Chiang Frank Wang, Kai-Lung Hua |
ICIP | 3 |
| 2012 | Clothing genre classification by exploiting the style elementsabstractThis paper presents a novel approach to automatically classify the upperwear genre from a full-body input image with no restrictions of model poses, image backgrounds, and image resolutions. Five style elements, that are crucial for clothing recognition, are identified based on the clothing design theory. The corresponding features of each of these style elements are also designed. We illustrate the effectiveness of our approach by showing that the proposed algorithm achieved overall precision of 92.04%, recall of 92.45%, and F score of 92.25% with 1,077 clothing images crawled from popular online stores. Shintami Chusnul Hidayati, Wen-Huang Cheng, Kai-Lung Hua |
ACM Multimedia | 3 |
| 2012 | Inter Frame Video Compression With Large Dictionaries of Tilings: Algorithms for Tiling Selection and Entropy CodingabstractWe propose the use of large tree-structured dictionaries of tilings for video compression. Our first contribution is the construction of a rate-distortion cost function that admits fast search algorithms to select the optimal tiling for the motion compensation stage of a video coder. The computation of the cost is enabled through novel algorithms to approximate the bit rate and the distortion. Our second contribution is an efficient arithmetic coding algorithm to encode the selected tree-structured tiling. We illustrate the effectiveness of our approach by showing that a H.264/AVC-like video coder utilizing one of the proposed tiling selection methods results in up to 16% savings in bit rate for several standard video sequences as compared to H.264/AVC. This is accomplished with only a modest increase in the computation time at the encoder. Kai-Lung Hua, Rong Zhang 0014, Mary L. Comer, Ilya Pollak |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2006 | Optimal Image Tilings with Application to Video CompressionabstractWe build on our previous work on best basis selection algorithms in large tree-structured dictionaries, previously used to construct effective image coders based on block and lapped transforms. In the present paper, we use such algorithms to select the optimal tiling for the motion compensation step in video compression. We illustrate the effectiveness of this approach by showing that our tiling selection method results in up to 18% savings in bit rate as compared to the H.264 tiling selection, for several standard video sequences. Kai-Lung Hua, Ilya Pollak, Mary L. Comer |
ICIP | 1 |