EDBT 2026 Demo / reviewers in the wild / expert
Hong-Yuan Mark Liao
dblp:05/2757 · also Mark Liao 0001
· DBLP profile ↗
160ranked-venue papers
4as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 118 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 44 · 2 first-author · 6 since 2021Systems, architecture and hardware · 6 · 1 first-authorDatabases, data management, data science and information retrieval · 4Software engineering, systems software and programming languages · 2Human-computer interaction and ubiquitous computing · 2Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | YOLO-RD: Introducing Relevant and Compact Explicit Knowledge to YOLO by Retriever-DictionaryabstractIdentifying and localizing objects within images is a fundamental challenge, and numerous efforts have been made to enhance model accuracy by experimenting with diverse architectures and refining training strategies. Nevertheless, a prevalent limitation in existing models is overemphasizing the current input while ignoring the information from the entire dataset. We introduce an innovative $\textbf{R}etriever-\textbf{D}ictionary$ (RD) module to address this issue. This architecture enables YOLO-based models to efficiently retrieve features from a Dictionary that contains the insight of the dataset, which is built by the knowledge from Visual Models (VM), Large Language Models (LLM), or Visual Language Models (VLM). The flexible RD enables the model to incorporate such explicit knowledge that enhances the ability to benefit multiple tasks, specifically, segmentation, detection, and classification, from pixel to image level. The experiments show that using the RD significantly improves model performance, achieving more than a 3\% increase in mean Average Precision for object detection with less than a 1\% increase in model parameters. Beyond 1-stage object detection models, the RD module improves the effectiveness of 2-stage models and DETR-based architectures, such as Faster R-CNN and Deformable DETR. Code is released at https://github.com/henrytsui000/YOLO. Hao-Tang Tsui, Chien-Yao Wang, Hong-Yuan Mark Liao |
ICLR | 3 |
| 2025 | Generalist YOLO: Towards Real-Time End-to-End Multi-Task Visual Language ModelsabstractGeneralist models, capable of handling multiple modalities and tasks simultaneously, are currently one of the hottest research topics. However, due to interference between different tasks during the training process, existing generalist models require a very large decoder to achieve good results in various tasks, which makes real-time prediction difficult for current generalist models. This paper introduces Generalist YOLO, which takes a significant step towards real-time prediction systems for visual language generalist models. The proposed Generalist YOLO uses a unified encoder to reduce conflicts between different tasks, thereby decreasing the complexity required by the decoder. It also introduces a primary-secondary co-attention mechanism that allows different tasks to learn together more effectively, achieving high efficiency and high accuracy. We propose a semantically consistent asymmetric training strategy, allowing various tasks to benefit from performance improvements brought by the latest research results in various fields. The proposed Generalist YOLO achieves excellent results on various vision and language tasks based on MS COCO. While maintaining high accuracy across all tasks, it is 135 times faster than existing generalist models. The source code is released on GitHub at https://github.com/WongKinYiu/GeneralistYOLO. Hung-Shuo Chang, Chien-Yao Wang, Richard Robert Wang, Gene Chou, Hong-Yuan Mark Liao |
WACV | 5 |
| 2024 | YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information
Chien-Yao Wang, I-Hau Yeh, Hong-Yuan Mark Liao |
ECCV (31) | 3 |
| 2023 | YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object DetectorsabstractReal-time object detection is one of the most important research topics in computer vision. As new approaches regarding architecture optimization and training optimization are continually being developed, we have found two research topics that have spawned when dealing with these latest state-of-the-art methods. To address the topics, we propose a trainable bag-of-freebies oriented solution. We combine the flexible and efficient training tools with the proposed architecture and the compound scaling method. YOLOv7 surpasses all known object detectors in both speed and accuracy in the range from 5 FPS to 120 FPS and has the highest accuracy 56.8% AP among all known real-time object detectors with 30 FPS or higher on GPU V100. Source code is released in https://github.com/WongKinYiu/yolov7. Chien-Yao Wang, Aleksei Bochkovskii, Hong-Yuan Mark Liao |
CVPR | 3 |
| 2022 | SearchTrack: Multiple Object Tracking with Object-Customized Search and Motion-Aware Features
Zhong-Min Tsai, Yu-Ju Tsai, Chien-Yao Wang, Hong-Yuan Mark Liao, Youn-Long Lin, Yung-Yu Chuang |
BMVC | 4 |
| 2022 | Capturing Humans in Motion: Temporal-Attentive 3D Human Pose and Shape Estimation from Monocular VideoabstractLearning to capture human motion is essential to 3D human pose and shape estimation from monocular video. However, the existing methods mainly rely on recurrent or convolutional operation to model such temporal information, which limits the ability to capture non-local context relations of human motion. To address this problem, we propose a motion pose and shape network (MPS-Net) to effectively capture humans in motion to estimate accurate and temporally coherent 3D human pose and shape from a video. Specifically, we first propose a motion continuity attention (MoCA) module that leverages visual cues observed from human motion to adaptively recalibrate the range that needs attention in the sequence to better capture the motion continuity dependencies. Then, we develop a hierarchical attentive feature integration (HAFI) module to effectively combine adjacent past and future feature represen-tations to strengthen temporal correlation and refine the feature representation of the current frame. By coupling the MoCA and HAFI modules, the proposed MPS-Net excels in estimating 3D human pose and shape in the video. Though conceptually simple, our MPS-Net not only outperforms the state-of-the-art methods on the 3DPW, MPI-INF-3DHP, and Human3.6M benchmark datasets, but also uses fewer network parameters. The video demos can be found at https://mps-net.github.io/MPS-Net/. Wen-Li Wei, Jen-Chun Lin, Tyng-Luh Liu, Hong-Yuan Mark Liao |
CVPR | 4 |
| 2021 | Scaled-YOLOv4: Scaling Cross Stage Partial NetworkabstractWe show that the YOLOv4 object detection neural network based on the CSP approach, scales both up and down and is applicable to small and large networks while maintaining optimal speed and accuracy. We propose a network scaling approach that modifies not only the depth, width, resolution, but also structure of the network. YOLOv4-large model achieves state-of-the-art results: 55.5% AP (73.4% AP50) for the MS COCO dataset at a speed of ~ 16 FPS on Tesla V100, while with the test time augmentation, YOLOv4-large achieves 56.0% AP (73.3 AP50). To the best of our knowledge, this is currently the highest accuracy on the COCO dataset among any published work. The YOLOv4-tiny model achieves 22.0% AP (42.0% AP50) at a speed of ~443 FPS on RTX 2080Ti, while by using TensorRT, batch size = 4 and FP16-precision the YOLOv4-tiny achieves 1774 FPS. Chien-Yao Wang, Aleksei Bochkovskii, Hong-Yuan Mark Liao |
CVPR | 3 |
| 2021 | Learning to Visualize Music Through Shot Sequence for Automatic Concert Video MashupabstractAn experienced director usually switches among different types of shots to make visual storytelling more touching. When filming a musical performance, appropriate switching shots can produce some special effects, such as enhancing the expression of emotion or heating up the atmosphere. However, while the visual storytelling technique is often used in making professional recordings of a live concert, amateur recordings of audiences often lack such storytelling concepts and skills when filming the same event. Thus a versatile system that can perform video mashup to create a refined high-quality video from such amateur clips is desirable. To this end, we aim at translating the music into an attractive shot (type) sequence by learning the relation between music and visual storytelling of shots. The resulting shot sequence can then be used to better portray the visual storytelling of a song and guide the concert video mashup process. To achieve the task, we first introduces a novel probabilistic-based fusion approach, named as multi-resolution fused recurrent neural networks (MF-RNNs) with film-language, which integrates multi-resolution fused RNNs and a film-language model for boosting the translation performance. We then distill the knowledge in MF-RNNs with film-language into a lightweight RNN, which is more efficient and easier to deploy. The results from objective and subjective experiments demonstrate that both MF-RNNs with film-language and lightweight RNN can generate attractive shot sequences for music, thereby enhancing the viewing and listening experience. Wen-Li Wei, Jen-Chun Lin, Tyng-Luh Liu, Hsiao-Rong Tyan, Hsin-Min Wang, Hong-Yuan Mark Liao |
IEEE Trans. Multim. | 6 |
| 2020 | Drone-Based Vehicle Flow Estimation and its Application to Traffic Conflict Hotspot Detection at IntersectionsabstractDrones can provide a wider field of view, high mobility and flexibility for monitoring and analyzing traffic flows and safety conditions. In case of a perpendicular viewing angle to the ground, there will be a very less occlusion that can occur and make vehicle tracking be easier. Thus, a drone-based solution will be better for traffic conflict hotspot detection at an interaction. However, due to its observation far from the ground, limited battery time, and bandwidth, this solution should be edge-based and have a good recognition rate in small object detection. However, current edge-based SoTA (state-of-the-art) methods are weak in a small object detection. We propose CoBiF net (Concatenated Bi-Fusion feature pyramid network), a one-stage object detection model for a real-time small object detection, which consists of SPP (spatial pyramid pooling), FE (Feature Extractor), CF (Concatenated Feature) block, and BFM (Bottom-up Fusion Module). CoBiF net is memory-and-bandwidth saving for the most edge devices. Extensive experiments on UA VDT benchmark show the proposed method achieved the SoTA results for the small object detection task in terms of accuracy and efficiency. Ping-Yang Chen, Jun-Wei Hsieh, Munkhjargal Gochoo, Ming-Ching Chang, Chien-Yao Wang, Yong-Sheng Chen, Hong-Yuan Mark Liao |
ICIP | 7 |
| 2020 | How Incompletely Segmented Information Affects Multi-Object Tracking and Segmentation (MOTS)abstractIn recent years, deep learning has made dramatic advances in computer vision field, especially in improving the performance of object detection as well as instance semantic segmentation. Still, multi-object tracking (MOT) remains a very challenging issue. Even in state-of-the-art deep learning-based object detectors, a preferred paradigm for MOT: tracking-by-detection, can only slightly improve the tracking performance. Pixel-level information is considered more precise and useful for tracking performance improvement than using conventional information, such as foreground or background content in a bounding box. However, the performance of current state-of-the-art models for automatically annotating pixel-level information is still far from the expectation of human beings. Therefore, we shall explore how multi-object tracking and segmentation (MOTS) is affected when the information obtained after applying instance semantic segmentation is incomplete. We propose a mask-guided two-streamed augmentation learning (MGTSAL) algorithm, which can be applied to TrackR-CNN to alleviate significant drop of MOTS performance when encountering incompletely segmented information. We evaluate the proposed approach on MOTS KITTI dataset, and our approach outperforms the baseline model TrackR-CNN in all our experimental settings. The promising experimental results and ablation study validate the effectiveness of the proposed approach. Yu-Sheng Chou, Chien-Yao Wang, Shou-De Lin, Hong-Yuan Mark Liao |
ICIP | 4 |
| 2020 | Learning From Music to Visual Storytelling of Shots: A Deep Interactive Learning MechanismabstractLearning from music to visual storytelling of shots is an interesting and emerging task. It produces a coherent visual story in the form of a shot type sequence, which not only expands the storytelling potential for a song but also facilitates automatic concert video mashup process and storyboard generation. In this study, we present a deep interactive learning (DIL) mechanism for building a compact yet accurate sequence-to-sequence model to accomplish the task. Different from the one-way transfer between a pre-trained teacher network (or ensemble network) and a student network in knowledge distillation (KD), the proposed method enables collaborative learning between an ensemble teacher network and a student network. Namely, the student network also teaches. Specifically, our method first learns a teacher network that is composed of several assistant networks to generate a shot type sequence and produce the soft target (shot types) distribution accordingly through KD. It then constructs the student network that learns from both the ground truth label (hard target) and the soft target distribution to alleviate the difficulty of optimization and improve generalization capability. As the student network gradually advances, it turns to feed back knowledge to the assistant networks, thereby improving the teacher network in each iteration. Owing to such interactive designs, the DIL mechanism bridges the gap between the teacher and student networks and produces more superior capability for both networks. Objective and subjective experimental results demonstrate that both the teacher and student networks can generate more attractive shot sequences from music, thereby enhancing the viewing and listening experience. Jen-Chun Lin, Wen-Li Wei, Yen-Yu Lin, Tyng-Luh Liu, Hong-Yuan Mark Liao |
ACM Multimedia | 5 |
| 2019 | Dynamic Gallery for Real-Time Multi-Target Multi-Camera TrackingabstractFor multi-target multi-camera recognition tasks, tracking of objects of interest is one of the essential yet challenging issues due to the fact that the task requires re-identifying identical targets across distinct views. Multi-target multi-camera tracking (MTMCT) applications span a wide range of variety (e.g. crowd behavior analysis, anomaly individual tracking and sport player tracking), so how to make the system perform real-time tracking becomes a crucial research issue. In this paper, we propose an online hierarchical algorithm for extreme clustering based MTMCT framework. The system can automatically create a dynamic gallery with real-time fashion by collecting appearance information of multi-object tracking in single-camera view. We evaluate the effectiveness and efficiency of our framework, and compare the state-of-the-art methods on MOT16 as well as DukeMTMC for single and multiple camera tracking. The high-frame-rate performance and promising tracking results confirm our system can be used in realworld applications. Yu-Sheng Chou, Chien-Yao Wang, Ming-Chiao Chen, Shou-De Lin, Hong-Yuan Mark Liao |
AVSS | 5 |
| 2019 | Real-Time Video-Based Person Re-Identification Surveillance with Light-Weight Deep Convolutional NetworksabstractToday's person re-ID system mostly focuses on accuracy and ignores efficiency. But in most real-world surveillance systems, efficiency is often considered the most important focus of research and development. Therefore, for a person re-ID system, the ability to perform real-time identification is the most important consideration. In this study, we implemented a real-time multiple camera video-based person re-ID system using the NVIDIA Jetson TX2 platform. This system can be used in a field that requires high privacy and immediate monitoring. This system uses YOLOv3-tiny based light-weight strategies and person re-ID technology, thus reducing 46% of computation, cutting down 39.9% of model size, and accelerating 21% of computing speed. The system also effectively upgrades the pedestrian detection accuracy. In addition, the proposed person re-ID example mining and training method improves the model's performance and enhances the robustness of cross-domain data. Our system also supports the pipeline formed by connecting multiple edge computing devices in series. The system can operate at a speed up to 18 fps at 1920×1080 surveillance video stream. The demo of our developed systems can be found at https://sites.google.com/g.ncu.edu.tw/video-based-person-re-id/. Chien-Yao Wang, Ping-Yang Chen, Ming-Chiao Chen, Jun-Wei Hsieh, Hong-Yuan Mark Liao |
AVSS | 5 |
| 2019 | What Makes You Look Like You: Learning an Inherent Feature Representation for Person Re-IdentificationabstractIn this work, we address person re-identification (ReID) by learning an inherent feature representation (inherent code) that is unique to each individual. This task is difficult because the appearance of a person may vary dramatically due to diverse factors, such as illuminations, viewpoints, and human pose changes. To tackle this issue, we propose new learning objectives to learn the inherent code for each person based on deep learning. Specifically, the proposed deep-net model is trained by jointly optimizing the multiple objectives that pulls the instances of the same person closer while pushing the instances belonging to different persons far from each other. Owing to such complementary designs, the deep-net model yields a robust code for each individual and hence better solve person ReID. Promising experimental results demonstrate the robustness and effectiveness of our proposed method. Wen-Li Wei, Jen-Chun Lin, Yen-Yu Lin, Hong-Yuan Mark Liao |
AVSS | 4 |
| 2019 | Smaller Object Detection for Real-Time Embedded Traffic Flow Estimation Using Fish-Eye CamerasabstractReal-time embedded traffic flow estimation (RETFE) systems need accurate and efficient vehicle detection models to meet limited resources in budget, dimension, memory, and computing power. In recent years, object detection became a less challenging task with latest deep CNN-based state-of-the-art models, i.e., RCNN, SSD, and YOLO; however, these models cannot provide desired performance for RETFE systems due to their complex time-consuming architecture. In addition, small object (<; 30×30 pixels) detection is still a challenging task for existing methods. Thus, we propose a shallow model named Concatenated Feature Pyramid Network (CFPN) that inspired from YOLOv3 to provide above mentioned performance for the smaller object detection. Main contribution is a proposed concatenated block (CB) which has reduced number of convolutional layers and concatenations instead of time-consuming algebraic operations. The superiority of CFPN is confirmed on the COCO and an in-house CarFlow datasets on Nvidia TX2. Thus we conclude that CFPN is useful for real-time embedded smaller object detection task. Ping-Yang Chen, Jun-Wei Hsieh, Munkhjargal Gochoo, Chien-Yao Wang, Hong-Yuan Mark Liao |
ICIP | 5 |
| 2019 | Tell Me Where It is Still Blurry: Adversarial Blurred Region Mining and RefiningabstractMobile devices such as smart phones are ubiquitously being used to take photos and videos, thus increasing the importance of image deblurring. This study introduces a novel deep learning approach that can automatically and progressively achieve the task via adversarial blurred region mining and refining (adversarial BRMR). Starting with a collaborative mechanism of two coupled conditional generative adversarial networks (CGANs), our method first learns the image-scale CGAN, denoted as iGAN, to globally generate a deblurred image and locally uncover its still blurred regions through an adversarial mining process. Then, we construct the patch-scale CGAN, denoted as pGAN, to further improve sharpness of the most blurred region in each iteration. Owing to such complementary designs, the adversarial BRMR indeed functions as a bridge between iGAN and pGAN, and yields the performance synergy in better solving blind image deblurring. The overall formulation is self-explanatory and effective to globally and locally restore an underlying sharp image. Experimental results on benchmark datasets demonstrate that the proposed method outperforms the current state-of-the-art technique for blind image deblurring both quantitatively and qualitatively. Jen-Chun Lin, Wen-Li Wei, Tyng-Luh Liu, C.-C. Jay Kuo, Hong-Yuan Mark Liao |
ACM Multimedia | 5 |
| 2018 | How Sampling Rate Affects Cross-Domain Transfer Learning for Video DescriptionabstractTranslating video to language is very challenging due to diversified video contents originated from multiple activities and complicated integration of spatio-temporal information. There are two urgent issues associated with the video-to-language translation problem. First, how to transfer knowledge learned from a more general dataset to a specific application domain dataset? Second, how to generate stable video captioning (or description) results under different sampling rates? In this paper, we propose a novel temporal embedding method to better retain temporal representation under different video sampling rates. We present a transfer learning method that combines a stacked LSTM encoder-decoder structure and a temporal embedding learning with soft-attention (TELSA) mechanism. We evaluate the proposed approach on two public datasets, including MSR-VTT and MSVD. The promising experimental results confirm the effectiveness of the proposed approach. Yu-Sheng Chou, Pai-Heng Hsiao, Shou-De Lin, Hong-Yuan Mark Liao |
ICASSP | 4 |
| 2018 | Seethevoice: Learning from Music to Visual Storytelling of ShotsabstractTypes of shots in the language of film are considered the key elements used by a director for visual storytelling. In filming a musical performance, manipulating shots could stimulate desired effects such as manifesting the emotion or deepening the atmosphere. However, while the visual storytelling technique is often employed in creating professional recordings of a live concert, audience recordings of the same event often lack such sophisticated manipulations. Thus it would be useful to have a versatile system that can perform video mashup to create a refined video from such amateur clips. To this end, we propose to translate the music into a near-professional shot (type) sequence by learning the relation between music and visual storytelling of shots. The resulting shot sequence can then be used to better portray the visual storytelling of a song and guide the concert video mashup process. Our method introduces a novel probabilistic-based fusion approach, named as multi-resolution fused recurrent neural networks (MF-RNNs) with film-language, which integrates multi-resolution fused RNNs and a film-language model for boosting the translation performance. The results from objective and subjective experiments demonstrate that MF-RNNs with film-language can generate an appealing shot sequence with better viewing experience. Wen-Li Wei, Jen-Chun Lin, Tyng-Luh Liu, Yi-Hsuan Yang, Hsin-Min Wang, Hsiao-Rong Tyan, Hong-Yuan Mark Liao |
ICME | 7 |
| 2018 | Session details: Multimedia-1 (Multimedia Recommendation & Discovery)
Hong-Yuan Mark Liao |
ACM Multimedia | 1 |
| 2018 | Automatic Image Cropping for Visual Aesthetic Enhancement Using Deep Neural Networks and Cascaded RegressionabstractDespite recent progress, computational visual aesthetic is still challenging. Image cropping, which refers to the removal of unwanted scene areas, is an important step to improve the aesthetic quality of an image. However, it is challenging to evaluate whether cropping leads to aesthetically pleasing results because the assessment is typically subjective. In this paper, we propose a novel cascaded cropping regression (CCR) method to perform image cropping by learning the knowledge from professional photographers. The proposed CCR method improves the convergence speed of the cascaded method, which directly uses random-ferns regressors. In addition, a two-step learning strategy is proposed and used in the CCR method to address the problem of lacking labelled cropping data. Specifically, a deep convolutional neural network (CNN) classifier is first trained on large-scale visual aesthetic datasets. The deep CNN model is then designed to extract features from several image cropping datasets, upon which the cropping bounding boxes are predicted by the proposed CCR method. Experimental results on public image cropping datasets demonstrate that the proposed method significantly outperforms several state-of-the-art image cropping methods. Guanjun Guo, Hanzi Wang, Chunhua Shen, Yan Yan 0001, Hong-Yuan Mark Liao |
IEEE Trans. Multim. | 5 |
| 2018 | Coherent Deep-Net Fusion To Classify Shots In Concert VideosabstractVarying types of shots is a fundamental element in the language of film, commonly used by a visual storytelling director. The technique is often used in creating professional recordings of a live concert, but meanwhile may not be appropriately applied in audience recordings of the same event. Such variations could cause the task of classifying shots in concert videos, professional or amateur, very challenging. To achieve more reliable shot classification, we propose a novel probabilistic-based approach, named as coherent classification net (CC-Net), by addressing three crucial issues. First, we focus on learning more effective features by fusing the layer-wise outputs extracted from a deep convolutional neural network (CNN), pretrained on a large-scale data set for object recognition. Second, we introduce a frame-wise classification scheme, the error weighted deep cross-correlation model (EW-Deep-CCM), to boost the classification accuracy. Specifically, the deep neural network-based cross-correlation model (deep-CCM) is constructed to not only model the extracted feature hierarchies of CNN independently, but also relate the statistical dependencies of paired features from different layers. Then, a Bayesian error weighting scheme for a classifier combination is adopted to explore the contributions from individual Deep-CCM classifiers to enhance the accuracy of shot classification in each image frame. Third, we feed the frame-wise classification results to a linear-chain conditional random field module to refine the shot predictions by taking into account the global and temporal regularities. We provide extensive experimental results on a data set of live concert videos to demonstrate the advantage of the proposed CC-Net over existing popular fusion approaches for shot classification. Jen-Chun Lin, Wen-Li Wei, Tyng-Luh Liu, Yi-Hsuan Yang, Hsin-Min Wang, Hsiao-Rong Tyan, Hong-Yuan Mark Liao |
IEEE Trans. Multim. | 7 |
| 2017 | Deep-net fusion to classify shots in concert videosabstractVarying types of shots is a fundamental element in the language of film, commonly used by a visual storytelling director to convey the emotion, ideas, and art. To classify such types of shots from images, we present a new framework that facilitates the intriguing task by addressing two key issues. We first focus on learning more effective features by fusing the layer-wise outputs extracted from a deep convolutional neural network (CNN), pre-trained on a large-scale dataset for object recognition. We then introduce a probabilistic fusion model, termed as error weighted deep cross-correlation model (EW-Deep-CCM), to boost the classification accuracy. Specifically, the deep neural network-based cross-correlation model (Deep-CCM) is constructed to not only model the extracted feature hierarchies of CNN independently but also relate the statistical dependencies of paired features from different layers. Then, a Bayesian error weighting scheme for classifier combination is adopted to explore the contributions from individual Deep-CCM classifiers to enhance the accuracy of shot classification. We provide extensive experimental results on a dataset of live concert videos to demonstrate the advantage of the proposed EW-Deep-CCM over existing popular fusion approaches. The video demos can be found at https://sites.google.com/site/ewdeepccm2/demo. Wen-Li Wei, Jen-Chun Lin, Tyng-Luh Liu, Yi-Hsuan Yang, Hsin-Min Wang, Hsiao-Rong Tyan, Hong-Yuan Mark Liao |
ICASSP | 7 |
| 2017 | Deep dictionary learning for fine-grained image classificationabstractFine-grained image classification is quite challenging due to high inter-class similarity and large intra-class variations. Another issue is the small amount of training images with a large number of classes to be identified. To address the challenges, we propose a model for fine-grained image classification with its application to bird species recognition. Based on the features extracted by bilinear convolutional neural network (BCNN), we propose an on-line dictionary learning algorithm where the principle of sparsity is integrated into classification. The features extracted by BCNN encode pairwise neuron interaction in a translation-invariant manner. This property is valuable to fine-grained classification. The proposed algorithm for dictionary learning further carries out sparsity based classification, where training data can be represented with a less number of dictionary atoms. It alleviates the problems caused by insufficient training data, and makes classification much more efficient. Our approach is evaluated and compared with the state-of-the-art approaches on the CUB-200-2011 dataset. The promising experimental results demonstrate its efficacy and superiority. Yen-Yu Lin, Hong-Yuan Mark Liao |
ICIP | 3 |
| 2017 | Recognizing offensive tactics in broadcast basketball videos via key player detectionabstractWe address offensive tactic recognition in broadcast basketball videos. As a crucial component towards basketball video content understanding, tactic recognition is quite challenging because it involves multiple independent players, each of which has respective spatial and temporal variations. Motivated by the observation that most intra-class variations are caused by non-key players, we present an approach that integrates key player detection into tactic recognition. To save the annotation cost, our approach can work on training data with only video-level tactic annotation, instead of key players labeling. Specifically, this task is formulated as an MIL (multiple instance learning) problem where a video is treated as a bag with its instances corresponding to subsets of the five players. We also propose a representation to encode the spatio-temporal interaction among multiple players. It turns out that our approach not only effectively recognizes the tactics but also precisely detects the key players. Tsung-Yu Tsai, Yen-Yu Lin, Hong-Yuan Mark Liao, Shyh-Kang Jeng |
ICIP | 3 |
| 2017 | Learning deep and sparse feature representation for fine-grained object recognitionabstractIn this paper, we address fine-grained classification which is quite challenging due to high intra-class variations and subtle inter-class variations. Most modern approaches to fine-grained recognition are established based on convolutional neural networks (CNN). Despite the effectiveness, these approaches still suffer from two major problems. First, they highly rely on large sets of training data, but manually annotating numerous training data is expensive. Second, the learned feature presentations by these approaches are often of high dimensions, leading to less efficiency. To tackle the two problems, we present an approach where on-line dictionary learning is integrated into CNN. The dictionaries can be incrementally learned by leveraging a vast amount of weakly labeled data on the Internet. With these dictionaries, all the training and testing data can be sparsely represented. Our approach is evaluated and compared with the state-of-the-art approaches on the benchmark dataset, CUB-200-2011. The promising results demonstrate its superiority in both efficiency and accuracy. Yen-Yu Lin, Hong-Yuan Mark Liao |
ICME | 3 |
| 2017 | Automatic Music Video Generation Based on Simultaneous Soundtrack Recommendation and Video EditingabstractAn automated process that can suggest a soundtrack to a user-generated video (UGV) and make the UGV a music-compliant professional-like video is challenging but desirable. To this end, this paper presents an automatic music video (MV) generation system that conducts soundtrack recommendation and video editing simultaneously. Given a long UGV, it is first divided into a sequence of fixed-length short (e.g., 2 seconds) segments, and then a multi-task deep neural network (MDNN) is applied to predict the pseudo acoustic (music) features (or called the pseudo song) from the visual (video) features of each video segment. In this way, the distance between any pair of video and music segments of same length can be computed in the music feature space. Second, the sequence of pseudo acoustic (music) features of the UGV and the sequence of the acoustic (music) features of each music track in the music collection are temporarily aligned by the dynamic time warping (DTW) algorithm with a pseudo-song-based deep similarity matching (PDSM) metric. Third, for each music track, the video editing module selects and concatenates the segments of the UGV based on the target and concatenation costs given by a pseudo-song-based deep concatenation cost (PDCC) metric according to the DTW-aligned result to generate a music-compliant professional-like video. Finally, all the generated MVs are ranked, and the best MV is recommended to the user. The MDNN for pseudo song prediction and the PDSM and PDCC metrics are trained by an annotated official music video (OMV) corpus. The results of objective and subjective experiments demonstrate that the proposed system performs well and can generate appealing MVs with better viewing and listening experiences. Jen-Chun Lin, Wen-Li Wei, Hsin-Min Wang, Hong-Yuan Mark Liao |
ACM Multimedia | 5 |
| 2016 | Precise player segmentation in team sports videos using contrast-aware co-segmentationabstractPlayer segmentation in team sports videos is challenging but crucial to video semantic understanding, such as player interaction identification and tactic analysis. We leverage the appearance similarity among players of the same team, and cast this task as a co-segmentation problem. In this way, the extra knowledge shared across players significantly reduces unfavorable uncertainty in segmenting individual players. We are also aware that the performance of co-segmentation highly depends on the used features, and further propose a contrast-based approach to estimate the discriminant power of each feature in an unsupervised manner. It turns out that our approach can properly fuse features by assigning higher weights to discriminant ones, and result in remarkable performance gains. The promising results on segmenting basketball players manifest the effectiveness of our approach. Tsung-Yu Tsai, Yen-Yu Lin, Hong-Yuan Mark Liao, Shyh-Kang Jeng |
ICASSP | 3 |
| 2016 | Visual saliency detection based on homology similarity and an experimental evaluation
Hanzi Wang, Liming Zhang 0002, Yan Yan 0001, Hong-Yuan Mark Liao |
J. Vis. Commun. Image Represent. | 5 |
| 2016 | Court Reconstruction for Camera Calibration in Broadcast Basketball VideosabstractWe introduce a technique of calibrating camera motions in basketball videos. Our method particularly transforms player positions to standard basketball court coordinates and enables applications such as tactical analysis and semantic basketball video retrieval. To achieve a robust calibration, we reconstruct the panoramic basketball court from a video, followed by warping the panoramic court to a standard one. As opposed to previous approaches, which individually detect the court lines and corners of each video frame, our technique considers all video frames simultaneously to achieve calibration; hence, it is robust to illumination changes and player occlusions. To demonstrate the feasibility of our technique, we present a stroke-based system that allows users to retrieve basketball videos. Our system tracks player trajectories from broadcast basketball videos. It then rectifies the trajectories to a standard basketball court by using our camera calibration method. Consequently, users can apply stroke queries to indicate how the players move in gameplay during retrieval. The main advantage of this interface is an explicit query of basketball videos so that unwanted outcomes can be prevented. We show the results in Figs. 1, 7, 9, 10 and our accompanying video to exhibit the feasibility of our technique. Pei-Chih Wen, Wei-Chih Cheng, Yu-Shuen Wang, Hung-Kuo Chu, Nick C. Tang, Hong-Yuan Mark Liao |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2015 | Spatio-Temporal Learning of Basketball Offensive StrategiesabstractVideo-based group behavior analysis is drawing attention to its rich applications in sports, military, surveillance and biological observations. The recent advances in tracking techniques, based on either computer vision methodology or hardware sensors, further provide the opportunity of better solving this challenging task. Focusing specifically on the analysis of basketball offensive strategies, we introduce a systematic approach to establishing unsupervised modeling of group behaviors. In view that a possible group behavior (offensive strategy) could be of different duration and represented by dynamic player trajectories, the crux of our method is to automatically divide training data into meaningful clusters and learn their respective spatio-temporal model, which is established upon Gaussian mixture regression to account for intra-class spatio-temporal variations. The resulting strategy representation turns out to be flexible that can be used to not only establish the discriminant functions but also improve learning the models. We demonstrate the usefulness of our approach by exploring its effectiveness in analyzing a set of given basketball video clips. Ching-Hang Chen, Tyng-Luh Liu, Yu-Shuen Wang, Hung-Kuo Chu, Nick C. Tang, Hong-Yuan Mark Liao |
ACM Multimedia | 6 |
| 2015 | Assistive Image Comment Robot - A Novel Mid-Level Concept-Based RepresentationabstractWe present a general framework and working system for predicting likely affective responses of the viewers in the social media environment after an image is posted online. Our approach emphasizes a mid-level concept representation, in which intended affects of the image publisher is characterized by a large pool of visual concepts (termed PACs) detected from image content directly instead of textual metadata, evoked viewer affects are represented by concepts (termed VACs) mined from online comments, and statistical methods are used to model the correlations among these two types of concepts. We demonstrate the utilities of such approaches by developing an end-to-end Assistive Comment Robot application, which further includes components for multi-sentence comment generation, interactive interfaces, and relevance feedback functions. Through user studies, we showed machine suggested comments were accepted by users for online posting in 90 percent of completed user sessions, while very favorable results were also observed in various dimensions (plausibility, preference, and realism) when assessing the quality of the generated image comments. Yan-Ying Chen, Tao Chen 0015, Taikun Liu, Hong-Yuan Mark Liao, Shih-Fu Chang |
IEEE Trans. Affect. Comput. | 4 |
| 2015 | Robust Action Recognition via Borrowing Information Across Video ModalitiesabstractThe recent advances in imaging devices have opened the opportunity of better solving the tasks of video content analysis and understanding. Next-generation cameras, such as the depth or binocular cameras, capture diverse information, and complement the conventional 2D RGB cameras. Thus, investigating the yielded multimodal videos generally facilitates the accomplishment of related applications. However, the limitations of the emerging cameras, such as short effective distances, expensive costs, or long response time, degrade their applicability, and currently make these devices not online accessible in practical use. In this paper, we provide an alternative scenario to address this problem, and illustrate it with the task of recognizing human actions. In particular, we aim at improving the accuracy of action recognition in RGB videos with the aid of one additional RGB-D camera. Since RGB-D cameras, such as Kinect, are typically not applicable in a surveillance system due to its short effective distance, we instead offline collect a database, in which not only the RGB videos but also the depth maps and the skeleton data of actions are available jointly. The proposed approach can adapt the interdatabase variations, and activate the borrowing of visual knowledge across different video modalities. Each action to be recognized in RGB representation is then augmented with the borrowed depth and skeleton features. Our approach is comprehensively evaluated on five benchmark data sets of action recognition. The promising results manifest that the borrowed information leads to remarkable boost in recognition accuracy. Nick C. Tang, Yen-Yu Lin, Ju-Hsuan Hua, Shih-En Wei, Ming-Fang Weng, Hong-Yuan Mark Liao |
IEEE Trans. Image Process. | 6 |
| 2015 | Cross-Camera Knowledge Transfer for Multiview People CountingabstractWe present a novel two-pass framework for counting the number of people in an environment, where multiple cameras provide different views of the subjects. By exploiting the complementary information captured by the cameras, we can transfer knowledge between the cameras to address the difficulties of people counting and improve the performance. The contribution of this paper is threefold. First, normalizing the perspective of visual features and estimating the size of a crowd are highly correlated tasks. Hence, we treat them as a joint learning problem. The derived counting model is scalable and it provides more accurate results than existing approaches. Second, we introduce an algorithm that matches groups of pedestrians in images captured by different cameras. The results provide a common domain for knowledge transfer, so we can work with multiple cameras without worrying about their differences. Third, the proposed counting system is comprised of a pair of collaborative regressors. The first one determines the people count based on features extracted from intracamera visual information, whereas the second calculates the residual by considering the conflicts between intercamera predictions. The two regressors are elegantly coupled and provide an accurate people counting system. The results of experiments in various settings show that, overall, our approach outperforms comparable baseline methods. The significant performance improvement demonstrates the effectiveness of our two-pass regression framework. Nick C. Tang, Yen-Yu Lin, Ming-Fang Weng, Hong-Yuan Mark Liao |
IEEE Trans. Image Process. | 4 |
| 2014 | Depth and Skeleton Associated Action Recognition without Online Accessible RGB-D CamerasabstractThe recent advances in RGB-D cameras have allowed us to better solve increasingly complex computer vision tasks. However, modern RGB-D cameras are still restricted by the short effective distances. The limitation may make RGB-D cameras not online accessible in practice, and degrade their applicability. We propose an alternative scenario to address this problem, and illustrate it with the application to action recognition. We use Kinect to offline collect an auxiliary, multi-modal database, in which not only the RGB videos but also the depth maps and skeleton structures of actions of interest are available. Our approach aims to enhance action recognition in RGB videos by leveraging the extra database. Specifically, it optimizes a feature transformation, by which the actions to be recognized can be concisely reconstructed by entries in the auxiliary database. In this way, the inter-database variations are adapted. More importantly, each action can be augmented with additional depth and skeleton images retrieved from the auxiliary database. The proposed approach has been evaluated on three benchmarks of action recognition. The promising results manifest that the augmented depth and skeleton features can lead to remarkable boost in recognition accuracy. Yen-Yu Lin, Ju-Hsuan Hua, Nick C. Tang, Min-Hung Chen, Hong-Yuan Mark Liao |
CVPR | 5 |
| 2014 | Human action recognition using associated depth and skeleton informationabstractThe recent advances in imaging devices have opened the opportunity of better solving computer vision tasks. The next-generation cameras, such as the depth or binocular cameras, capture diverse information, and complement the conventional 2D RGB cameras. Thus, investigating the yielded multi-modal images generally facilitates the accomplishment of related applications. However, the limitations of these devices, such as short effective distances, expensive costs, or long response time, degrade their applicability in practical use. Addressing this problem in this work, we aim at action recognition in RGB videos with the aid of Kinect. We improve recognition accuracy by leveraging information derived from an offline collected database, in which not only the RGB but also the depth and skeleton images of actions are available. Our approach adapts the inter-database variations, and enables the sharing of visual knowledge across different image modalities. Each action instance for recognition in RGB representation is then augmented with the borrowed depth and skeleton features. Nick C. Tang, Yen-Yu Lin, Ju-Hsuan Hua, Ming-Fang Weng, Hong-Yuan Mark Liao |
ICASSP | 5 |
| 2014 | Music Driven Human Motion Manipulation for Characters in a VideoabstractMultimedia content creation and manipulation have garnered attention in recent days due to the desires of personalization. As a content producing application, we propose a novel idea that requires the fusion of video and audio intelligence. The system is composed of at least three core techniques: 1) the capability to process the video sequence to have access to the geometric and appearance information pertaining to meaningful and representative targets, 2) a systematic way to reliably classify and identify important emotions from the music, 3) effective approaches to manipulate the video targets according to the extracted music emotions. In this paper, we report preliminary results of the proposed system. Specifically, we introduce the employed framework to manipulate the magnitude and speed of music conducting gestures of a video sequence of human skeleton according to the emotion intensity and tempo of an arbitrary music excerpt, using state-of-the-art inverse kinematics and music information retrieval techniques. We present the details of the prototype system and validate its effectiveness with a video demonstrating how we can manipulate the music conducting gestures according to the proposed manipulation rules. Che-Hua Yeh, Yi-Hsuan Yang, Ming-Hsu Chang, Hong-Yuan Mark Liao |
ISM | 4 |
| 2014 | Predicting Viewer Affective Comments Based on Image Content in Social MediaabstractVisual sentiment analysis is getting increasing attention because of the rapidly growing amount of images in online social interactions and several emerging applications such as online propaganda and advertisement. Recent studies have shown promising progress in analyzing visual affect concepts intended by the media content publisher. In contrast, this paper focuses on predicting what viewer affect concepts will be triggered when the image is perceived by the viewers. For example, given an image tagged with "yummy food," the viewers are likely to comment "delicious" and "hungry," which we refer to as viewer affect concepts (VAC) in this paper. To the best of our knowledge, this is the first work explicitly distinguishing intended publisher affect concepts and induced viewer affect concepts associated with social visual content, and aiming at understanding their correlations. We present around 400 VACs automatically mined from million-scale real user comments associated with images in social media. Furthermore, we propose an automatic visual based approach to predict VACs by first detecting publisher affect concepts in image content and then applying statistical correlations between such publisher affect concepts and the VACs. We demonstrate major benefits of the proposed methods in several real-world tasks - recommending images to invoke certain target VACs among viewers, increasing the accuracy of predicting VACs by 20.1% and finally developing a social assistant tool that may suggest plausible, content-specific and desirable comments when users view new images. Yan-Ying Chen, Tao Chen 0015, Winston H. Hsu, Hong-Yuan Mark Liao, Shih-Fu Chang |
ICMR | 4 |
| 2014 | Simultaneous Tensor Decomposition and Completion Using Factor PriorsabstractThe success of research on matrix completion is evident in a variety of real-world applications. Tensor completion, which is a high-order extension of matrix completion, has also generated a great deal of research interest in recent years. Given a tensor with incomplete entries, existing methods use either factorization or completion schemes to recover the missing parts. However, as the number of missing entries increases, factorization schemes may overfit the model because of incorrectly predefined ranks, while completion schemes may fail to interpret the model factors. In this paper, we introduce a novel concept: complete the missing entries and simultaneously capture the underlying model structure. To this end, we propose a method called simultaneous tensor decomposition and completion (STDC) that combines a rank minimization technique with Tucker model decomposition. Moreover, as the model structure is implicitly included in the Tucker model, we use factor priors, which are usually known a priori in real-world tensor objects, to characterize the underlying joint-manifold drawn from the model factors. By exploiting this auxiliary information, our method leverages two classic schemes and accurately estimates the model factors and missing entries. We conducted experiments to empirically verify the convergence of our algorithm on synthetic data and evaluate its effectiveness on various kinds of real-world data. The results demonstrate the efficacy of the proposed method and its potential usage in tensor-based applications. It also outperforms state-of-the-art methods on multilinear model analysis and visual data completion tasks. Yi-Lei Chen, Chiou-Ting Hsu, Hong-Yuan Mark Liao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | Example-Based Human Motion Extrapolation and Motion Repairing Using Contour ManifoldabstractWe propose a human motion extrapolation algorithm that synthesizes new motions of a human object in a still image from a given reference motion sequence. The algorithm is implemented in two major steps: contour manifold construction and object motion synthesis. Contour manifold construction searches for low-dimensional manifolds that represent the temporal-domain deformation of the reference motion sequence. Since the derived manifolds capture the motion information of the reference sequence, the representation is more robust to variations in shape and size. With this compact representation, we can easily modify and manipulate human motions through interpolation or extrapolation in the contour manifold space. In the object motion synthesis step, the proposed algorithm generates a sequence of new shapes of the input human object in the contour manifold space and then renders the textures of those shapes to synthesize a new motion sequence. We demonstrate the efficacy of the algorithm on different types of practical applications, namely, motion extrapolation and motion repair. Nick C. Tang, Chiou-Ting Hsu, Ming-Fang Weng, Tsung-Yi Lin, Hong-Yuan Mark Liao |
IEEE Trans. Multim. | 5 |
| 2014 | Per-Cluster Ensemble Kernel Learning for Multi-Modal Image Clustering With Group-Dependent Feature SelectionabstractIn this paper, we present a clustering approach, MK-SOM, that carries out cluster-dependent feature selection, and partitions images with multiple feature representations into clusters. This work is motivated by the observations that human visual systems (HVS) can receive various kinds of visual cues for interpreting the world. Images identified by HVS as the same category are typically coherent to each other in certain crucial visual cues, but the crucial cues vary from category to category. To account for this observation and bridge the semantic gap, the proposed MK-SOM integrates multiple kernel learning (MKL) into the training process of self-organizing map (SOM), and associates each cluster with a learnable, ensemble kernel. Hence, it can leverage information captured by various image descriptors, and discoveries the cluster-specific characteristics via learning the per-cluster ensemble kernels. Through the optimization iterations, cluster structures are gradually revealed via the features specified by the learned ensemble kernels, while the quality of these ensemble kernels is progressively improved owing to the coherent clusters by enforcing SOM. Besides, MK-SOM allows the introduction of side information to improve performance, and it hence provides a new perspective of applying MKL to address both unsupervised and semi-supervised clustering tasks. Our approach is comprehensively evaluated in the two applications. The superior and promising results manifest its effectiveness. Jeng-Tsung Tsai, Yen-Yu Lin, Hong-Yuan Mark Liao |
IEEE Trans. Multim. | 3 |
| 2013 | Multi-view face detection in videos with online adaptationabstractMost learning-based approaches to face detection suffer from the problem of performance degradation on faces that are not covered by training data. However, including all variations of faces in training is practically infeasible due to the scalability restriction of machine learning algorithms and expensive manual labeling. In this work, we focus on face detection in videos, and alleviate this problem by exploiting strong correlation among video frames. We augment a pre-trained multiview face detection with an incrementally derived Gaussian process regressor. The regressor can extract and propagate visual knowledge across frames, and adapts the detector to handle unseen faces. Testing on two datasets, the promising results manifest the effectiveness of the proposed approach. Yao-Chuan Chang, Yen-Yu Lin, Hong-Yuan Mark Liao |
ICIP | 3 |
| 2013 | Cross-camera vehicle tracking via affine invariant object matching for video forensics applicationsabstractThe recent deployment of very large-scale camera networks consisting of fixed/moving surveillance cameras and vehicle video recorders, has led to a novel field in object tracking problem. The major goal is to detect and track each vehicle within a large area, which can be applied to video forensics. For example, a suspected vehicle can be automatically identified for mining digital criminal evidences from a large amount of video data. In this paper, we propose an efficient cross-camera vehicle tracking technique via affine invariant object matching. More specifically, we formulate the problem as invariant image feature matching among different viewpoints of cameras. To achieve vehicle matching, we first extract invariant image feature based on ASIFT (affine and scale-invariant feature transform) for each detected vehicle in a camera network. Then, to improve the accuracy of ASIFT feature matching between images from different viewpoints, we propose to efficiently match feature points based on our observed spatially invariant property of ASIFT, as well as the min-hash technique. As a result, cross-camera vehicle tracking can be efficiently and accurately achieved. Experimental results demonstrate the efficacy of the proposed algorithm and the feasibility to video forensics applications. Chao-Yung Hsu, Li-Wei Kang, Hong-Yuan Mark Liao |
ICME | 3 |
| 2013 | An efficient expanding block algorithm for image copy-move forgery detection
Gavin Lynch, Frank Y. Shih, Hong-Yuan Mark Liao |
Inf. Sci. | 3 |
| 2013 | An improved DCT-based perturbation scheme for high capacity data hiding in H.264/AVC intra frames
Tseng-Jung Lin, Kuo-Liang Chung, Po-Chun Chang, Yong-Huai Huang, Hong-Yuan Mark Liao, Chiung-Yao Fang |
J. Syst. Softw. | 5 |
| 2013 | Efficient reversible data hiding algorithm based on gradient-based edge direction prediction
Wei-Jen Yang, Kuo-Liang Chung, Hong-Yuan Mark Liao, Wen-Kuang Yu |
J. Syst. Softw. | 3 |
| 2013 | Moving foreground object detection via robust SIFT trajectories
Shih-Wei Sun, Yu-Chiang Frank Wang, Fay Huang, Hong-Yuan Mark Liao |
J. Vis. Commun. Image Represent. | 4 |
| 2013 | Automatic Training Image Acquisition and Effective Feature Selection From Community-Contributed Photos for Facial Attribute DetectionabstractFacial attributes are shown effective for mining specific persons and profiling human activities in large-scale media such as surveillance videos or photo-sharing services. For comprehensive analysis, a rich number of facial attributes is required. Generally, each attribute detector is obtained by supervised learning via the use of large training data. It is promising to leverage the exponentially growing community contributed photos and the associated informative contexts to ease the burden of manual annotation; however, such huge noisy data from the Internet still pose great challenges. We propose to measure the quality of training images by discriminable visual features, which are verified with the relative discrimination between the unlabeled images and the pseudo-positives (pseudo-negatives) retrieved by textual relevance. The proposed feature selection requires no heuristic threshold, therefore, can be generalized to multiple feature modalities. We further exploit the rich context cues (e.g., tags, geo-locations, etc.) associated with the publicly available photos for mining more semantically consistent but visually diverse training images around the world. Experimenting in the benchmarks, we demonstrate that our work can successfully acquire effective training images for learning generic facial attributes, where the classification error is relatively reduced up to 23.35% compared with that of the text-based approach and shown comparable with that of costly manual annotations. (All of the face images presented in this paper except the training images in Fig. 8 attribute to various Flickr users under Creative Commons Licenses). Yan-Ying Chen, Winston H. Hsu, Hong-Yuan Mark Liao |
IEEE Trans. Multim. | 3 |
| 2012 | Action recognition using instance-specific and class-consistent cuesabstractWe aim to resolve the difficulties of action recognition arising from the large intra-class variations. These unfavorable variations make it infeasible to represent one action instance by other ones of the same action. We hence propose to extract both instance-specific and class-consistent features to facilitate action recognition. Specifically, the instance-specific features explore the self-similarities among frames of each video instance, while class-consistent features summarize within-class similarities. We introduce a generative formulation to combine the two diverse types of features. The experimental results demonstrate the effectiveness of our approach. Chin-An Lin, Yen-Yu Lin, Hong-Yuan Mark Liao, Shyh-Kang Jeng |
ICIP | 3 |
| 2012 | Who's Who in a Sports Video? An Individual Level Sports Video Indexing SystemabstractSports video analysis has attracted great attention in recent years. In the past decade, numerous sports video indexing approaches have been proposed at different semantic levels. In this paper, an individual level sports video indexing (ILSVI) scheme is proposed. The individual level refers to the indexing of a sports video on a player basis, i.e. to recognize each player in a multi-player game. Since the jersey number is always "worn'' by a player as the player's identity in a game, it is feasible to recognize jersey numbers for individual level indexing in sports videos. To solve the jersey number recognition problem, a principal-axis based contour descriptor is proposed. Compared to the state-of-the-art approaches, the proposed descriptor can achieve higher recognition rate and only consume much less computation power. In addition, we developed an interactive system to realize the individual level sports video indexing (ILSVI). This interactive system includes a player detection and a jersey number detection sub-systems. The interactive system can help complete the individual level sports video indexing task. We shall use basketball game videos as the basis to develop real-world systems. Shih-Wei Sun, Wen-Huang Cheng, Yao-Ling Hung, Ivy Fan, Chris Liu, Jacqueline Hung, Chia-Kai Lin, Hong-Yuan Mark Liao |
ICME | 8 |
| 2012 | Discovering informative social subgraphs and predicting pairwise relationships from group photosabstractAn increasing number of users are contributing the sheer amount of group photos (e.g., for family, classmates, colleagues, etc.) on social media for the purpose of photo sharing and social communication. There arise strong needs for automatically understanding the group types (e.g., family vs. classmates) for recommendation services (e.g., recommending a family-friendly restaurant) and even predicting the pairwise relationships (e.g., mother-child) between the people in the photo for mining implicit social connections. Interestingly, we observe that the group photos are composed of atomic subgroups corresponding to certain social relationships. For this work, we propose a novel framework to (1) connect faces of different attributes and positions as a face graph and (2) discover informative subgraphs to represent social subgroups in group photos. A group photo can be further represented by a bag-of-face-subgraphs (BoFG) -- the occurring frequency of social subgroups, which is informative to categorize specific group types or events. We demonstrate the effectiveness of BoFG in recognizing family photos and achieve 30.5% relative improvement over the state-of-the-art low-level features. Moreover, we propose to predict the pairwise relationships (e.g., husband-wife) in a face graph by the co-occurrence information (e.g., co-occurring with a child) in the mined subgraphs. The experiments demonstrate that the informative social subgroups significantly outperform prior work (36% relatively) which considers merely facial attributes for determining pairwise relationships. Yan-Ying Chen, Winston H. Hsu, Hong-Yuan Mark Liao |
ACM Multimedia | 3 |
| 2012 | Visual knowledge transfer among multiple cameras for people counting with occlusion handlingabstractWe present a framework to count the number of people in an environment where multiple cameras with different angles of view are available. We consider the visual cues captured by each camera as a knowledge source, and carry out cross-camera knowledge transfer to alleviate the difficulties of people counting, such as partial occlusions, low-quality images, clutter backgrounds, and so on. Specifically, this work distinguishes itself with the following contributions. First, we overcome the variations of multiple heterogeneous cameras with different perspective settings by matching the same groups of pedestrians taken by these cameras, and present an algorithm for accomplishing cross-camera correspondence. Second, the proposed counting model is composed of a pair of collaborative regressors. While one regressor measures people counts by the features extracted from intra-camera visual evidences, the other recovers the yielded residual by taking the conflicts among inter-camera predictions into account. The two regressors are elegantly coupled, and jointly lead to an accurate counting system. Additionally, we provide a set of manually annotated pedestrian labels on the PETS 2010 videos for performance evaluation. Our approach is comprehensively tested in various settings and compared with competitive baselines. The significant improvement in performance manifests the effectiveness of the proposed approach. Ming-Fang Weng, Yen-Yu Lin, Nick C. Tang, Hong-Yuan Mark Liao |
ACM Multimedia | 4 |
| 2012 | Examplar-based object posture super-resolution using manifold learningabstractThis paper proposes a learning-based approach to increase the temporal resolutions of human motion sequences. Given a set of high resolution motion sequences, our idea is first to learn the motion tendency from this learning dataset and then synthesize new postures for the low-resolution sequence according to the learned motion tendency. We summarize the proposed framework in the following steps: (1) Each motion sequence is first projected into a low-dimension manifold space, where the local distance between postures could be better preserved. We then represent each of the projected motion sequences as a motion trajectory. (2) Next, motion priors learned from the HR training sequences are used to reconstruct the motion trajectory for the input sequence. (3) Finally, we use the reconstructed motion trajectory combined with object inpainting technique to generate the final result. Our experimental results demonstrate the effectiveness of the proposed method, and also show its outperformance over existing approaches. Chih-Hung Ling, Chia-Wen Lin, Chiou-Ting Hsu, Hong-Yuan Mark Liao |
MMSP | 4 |
| 2012 | Efficient reversible data hiding for color filter array images
Wei-Jen Yang, Kuo-Liang Chung, Hong-Yuan Mark Liao |
Inf. Sci. | 3 |
| 2012 | Quality-efficient demosaicing for digital time delay and integration images using edge-sensing scheme in color difference domain
Wei-Jen Yang, Kuo-Liang Chung, Hong-Yuan Mark Liao |
J. Vis. Commun. Image Represent. | 3 |
| 2011 | Compression and protection of JPEG imagesabstractThe objective of this research is to design a new JPEG-based compression scheme which simultaneously considers the security issue. Our method starts from dividing image into non-overlapping blocks with size 8×8. Among these blocks, some are used as reference blocks and the rest are used as query blocks. A query block is the combination of the residual and the resultant of a filtered reference block. We put our emphasis on how to estimate an appropriate filter and then use it as part of a secret key. With both reference blocks and the residuals of query blocks, one is able to encode secured images using a correct secret key. The experiment results will demonstrate that how different secret keys can control the quality of restored image based on the priority of authority. Yi-Chong Zeng, Fay Huang, Hong-Yuan Mark Liao |
ICIP | 3 |
| 2011 | Automatic annotation of Web videosabstractMost Web videos are captured in uncontrolled environments (e.g. videos captured by freely-moving cameras with low resolution); this makes automatic video annotation very difficult. To address this problem, we present a robust moving foreground object detection method followed by the integration of features collected from heterogeneous domains. We advance SIFT feature matching and present a probabilistic framework to construct consensus foreground object templates (CFOT). The CFOT can detect moving foreground objects of interest across video frames, and this allows us to extract visual features from foreground regions of interest. Together with the use of audio features, we are able to improve resulting annotation accuracy. We conduct experiments and achieve promising results on a Web video dataset collected from YouTube. Shih-Wei Sun, Yu-Chiang Frank Wang, Yao-Ling Hung, Chia-Ling Chang, Kuan-Chieh Chen, Shih-Sian Cheng, Hsin-Min Wang, Hong-Yuan Mark Liao |
ICME | 8 |
| 2011 | Exemplar-based Age Progression Prediction in Children FacesabstractThis work aims to develop a system for predicting age progression in children faces. Age progression prediction in children faces is critical to assist missing children searching. An integral module including feature extraction, distance measurement, and face synthesis is devised in this paper to predict faces at different ages. In the proposed method, a curvature-weighted plus bending-energy distance is employed for selecting similar facial components from an aging database. The growth curve of each facial component is used to predict the shape, size, and location of each component at a different age. Thin plate spline method is employed to synthesize a 3-D face model from the predicted components by minimizing the bending energy. Experiments are conducted to test the proposed method with various subjects and the results show that the proposed method is very promising. Cheng-Ta Shen, Wan-Hua Lu, Sheng-Wen Shih, Hong-Yuan Mark Liao |
ISM | 4 |
| 2011 | Personalized travel recommendation by mining people attributes from community-contributed photosabstractLeveraging community-contributed data (e.g., blogs, GPS logs, and geo-tagged photos) for travel recommendation is one of the active researches since there are rich contexts and trip activities in such explosively growing data. In this work, we focus on personalized travel recommendation by leveraging the freely available community-contributed photos. We propose to conduct personalized travel recommendation by further considering specific user profiles or attributes (e.g., gender, age, race). In stead of mining photo logs only, we argue to leverage the automatically detected people attributes in the photo contents. By information-theoretic measures, we will demonstrate that such people attributes are informative and effective for travel recommendation -- especially providing a promising aspect for personalization. We effectively mine the demographics for different locations (or landmarks) and travel paths. A probabilistic Bayesian learning framework which further entails mobile recommendation on the spot is introduced. We experiment on four million photos collected for eight major worldwide cities. The experiments confirm that people attributes are promising and orthogonal to prior works using travel logs only and can further improve prior travel recommendation methods especially in difficult predictions by further leveraging user contexts in mobile devices. An-Jung Cheng, Yan-Ying Chen, Yen-Ta Huang, Winston H. Hsu, Hong-Yuan Mark Liao |
ACM Multimedia | 5 |
| 2011 | Example-based human motion extrapolation based on manifold learningabstractIn this paper, we propose a new framework to synthesize human motions based on only one single posture given in the input image. To generate visually pleasing motion sequences, the proposed framework consists of two key techniques. One is motion retrieval, which retrieves reference motions from a human motion database on a low-dimensional motion manifold. Another one is human motion extrapolation, which first generates new postures by deforming the shape of the input posture according to the retrieved motions and then synthesizes the corresponding motion sequence. To demonstrate the efficacy of the proposed method, we generate several human motion sequences using input images with different postures and show that the results are indeed visually pleasing. Nick C. Tang, Chiou-Ting Hsu, Tsung-Yi Lin, Hong-Yuan Mark Liao |
ACM Multimedia | 4 |
| 2011 | Narrative Generation by Repurposing Digital Videos
Nick C. Tang, Hsiao-Rong Tyan, Chiou-Ting Hsu, Hong-Yuan Mark Liao |
MMM (1) | 4 |
| 2011 | Human Object Inpainting Using Manifold Learning-Based Posture Sequence EstimationabstractWe propose a human object inpainting scheme that divides the process into three steps: 1) human posture synthesis; 2) graphical model construction; and 3) posture sequence estimation. Human posture synthesis is used to enrich the number of postures in the database, after which all the postures are used to build a graphical model that can estimate the motion tendency of an object. We also introduce two constraints to confine the motion continuity property. The first constraint limits the maximum search distance if a trajectory in the graphical model is discontinuous, and the second confines the search direction in order to maintain the tendency of an object's motion. We perform both forward and backward predictions to derive local optimal solutions. Then, to compute an overall best solution, we apply the Markov random field model and take the potential trajectory with the maximum total probability as the final result. The proposed posture sequence estimation model can help identify a set of suitable postures from the posture database to restore damaged/missing postures. It can also make a reconstructed motion sequence look continuous. Chih-Hung Ling, Yu-Ming Liang, Chia-Wen Lin, Yong-Sheng Chen, Hong-Yuan Mark Liao |
IEEE Trans. Image Process. | 5 |
| 2011 | Virtual Contour Guided Video Object Inpainting Using Posture Mapping and RetrievalabstractThis paper presents a novel framework for object completion in a video. To complete an occluded object, our method first samples a 3-D volume of the video into directional spatio-temporal slices, and performs patch-based image inpainting to complete the partially damaged object trajectories in the 2-D slices. The completed slices are then combined to obtain a sequence of virtual contours of the damaged object. Next, a posture sequence retrieval technique is applied to the virtual contours to retrieve the most similar sequence of object postures in the available non-occluded postures. Key-posture selection and indexing are used to reduce the complexity of posture sequence retrieval. We also propose a synthetic posture generation scheme that enriches the collection of postures so as to reduce the effect of insufficient postures. Our experiment results demonstrate that the proposed method can maintain the spatial consistency and temporal motion continuity of an object simultaneously. Chih-Hung Ling, Chia-Wen Lin, Chih-Wen Su, Yong-Sheng Chen, Hong-Yuan Mark Liao |
IEEE Trans. Multim. | 5 |
| 2011 | Video Inpainting on Digitized Vintage Films via Maintaining Spatiotemporal ContinuityabstractVideo inpainting is an important video enhancement technique used to facilitate the repair or editing of digital videos. It has been employed worldwide to transform cultural artifacts such as vintage videos/films into digital formats. However, the quality of such videos is usually very poor and often contain unstable luminance and damaged content. In this paper, we propose a video inpainting algorithm for repairing damaged content in digitized vintage films, focusing on maintaining good spatiotemporal continuity. The proposed algorithm utilizes two key techniques. Motion completion recovers missing motion information in damaged areas to maintain good temporal continuity. Frame completion repairs damaged frames to produce a visually pleasing video with good spatial continuity and stabilized luminance. We demonstrate the efficacy of the algorithm on different types of video clips. Nick C. Tang, Chiou-Ting Hsu, Chih-Wen Su, Timothy K. Shih, Hong-Yuan Mark Liao |
IEEE Trans. Multim. | 5 |
| 2010 | Linear production game solution to a PTZ camera networkabstractReconfiguring the PTZ parameters of a camera network is an combinatorial optimization problem and computing the optimal solution is very time consuming. Therefore, existing methods can only provide sub-optimal solutions. In this paper, a nonlinear objective function for better utilizing the cameras to track multiple targets is proposed. Furthermore, it is shown that by expanding the unknown parameters and imposing new constraints, the nonlinear objective function can be converted into a linear production game (LPG) problem. Since an LPG possesses an optimal solution which can be evaluated with polynomial time, the proposed method is efficient and accurate. Computer simulations have been conducted and the results show that the proposed method is very promising. Yu-Chun Lai, Yu-Ming Liang, Sheng-Wen Shih, Hong-Yuan Mark Liao, Cheng-Chung Lin |
ICIP | 4 |
| 2010 | Video object inpainting using manifold-based action predictionabstractThis paper presents a novel scheme for object completion in a video. The framework includes three steps: posture synthesis, graphical model construction, and action prediction. In the very beginning, a posture synthesis method is adopted to enrich the number of postures. Then, all postures are used to build a graphical model of object action which can provide possible motion tendency. We define two constraints to confine the motion continuity property. With the two constraints, possible candidates between every two consecutive postures are significantly reduced. Finally, we apply the Markov Random Field model to perform global matching. The proposed approach can effectively maintain the temporal continuity of the reconstructed motion. The advantage of this action prediction strategy is that it can handle the cases such as non-periodic motion or complete occlusion. Chih-Hung Ling, Yu-Ming Liang, Chia-Wen Lin, Yong-Sheng Chen, Hong-Yuan Mark Liao |
ICIP | 5 |
| 2010 | An RST-Tolerant Shape Descriptor for Object DetectionabstractIn this paper, we propose a new object detection method that does not need a learning mechanism. Given a hand-drawn model as a query, we can detect and locate objects that are similar to the query model in cluttered images. To ensure the invariance with respect to rotation, scaling, and translation (RST), high curvature points (HCPs) on edges are detected first. Each pair of HCPs is then used to determine a circular region and all edge pixels covered by the circular region are transformed into a polar histogram. Finally, we use these local descriptors to detect and locate similar objects within any images. The experiment results show that the proposed method outperforms the existing state-of-the-art work. Chih-Wen Su, Hong-Yuan Mark Liao, Yu-Ming Liang, Hsiao-Rong Tyan |
ICPR | 2 |
| 2010 | Data-Driven Foreground Object Detection from a Non-stationary CameraabstractIn this paper, we propose a data-driven foreground object detection technique which can detect foreground objects from a moving camera. We propose to build a data-driven consensus foreground object template (CFOT) and then detect the foreground object region in each frame. The proposed foreground object detection technique is equipped with the following functions: (1) the ability to detect the foreground object captured by a fast moving camera ; (2) the ability to detect a low contrast (spatially/temporally) foreground object; and (3) the ability to detect a foreground object from a dynamic background. There are three contributions of our method: (1) a newly proposed data-driven foreground region decision process for generating the CFOT has been shown robust and efficient; (2) a foreground object probability is proposed for properly dealing with the imperfect initial foreground region estimations; and (3) a CFOT is generated for precise foreground object detection. Shih-Wei Sun, Fay Huang, Hong-Yuan Mark Liao |
ICPR | 3 |
| 2010 | Face hallucination using Bayesian global estimation and local basis selectionabstractThis paper proposes a two-step prototype-face-based scheme of hallucinating the high-resolution detail of a low-resolution input face image. The proposed scheme is mainly composed of two steps: the global estimation step and the local facial-parts refinement step. In the global estimation step, the initial high-resolution face image is hallucinated via a linear combination of the global prototype faces with a coefficient vector. Instead of estimating coefficient vector in the high-dimensional raw image domain, we propose a maximum a posteriori (MAP) estimator to estimate the optimum set of coefficients in the low-dimensional coefficient domain. In the local refinement step, the facial parts (i.e., eyes, nose and mouth) are further refined using a basis selection method based on overcomplete nonnegative matrix factorization (ONMF). Experimental results demonstrate that the proposed method can achieve significant subjective and objective improvement over state-of-the-art face hallucination methods, especially when an input face does not belong to a person in the training data set. Chih-Chung Hsu, Chia-Wen Lin, Chiou-Ting Hsu, Hong-Yuan Mark Liao, Jen-Yu Yu |
MMSP | 4 |
| 2010 | Fast randomized algorithm for center-detection
Kuo-Liang Chung, Yong-Huai Huang, Jyun-Pin Wang, Ting-Chin Chang, Hong-Yuan Mark Liao |
Pattern Recognit. | 5 |
| 2010 | New orientation-based elimination approach for accurate line-detection
Kuo-Liang Chung, Zeng-Wei Lin, Shih-Ting Huang, Yong-Huai Huang, Hong-Yuan Mark Liao |
Pattern Recognit. Lett. | 5 |
| 2010 | Reversible Data Hiding-Based Approach for Intra-Frame Error Concealment in H.264/AVCabstractError concealment plays an important role in robust video transmission. Recently, Chen and Leung presented an efficient data hiding-based (DH-based) approach to recover corrupted macroblocks from the intra-frame of an H.264/AVC sequence, but it suffers from the quality degradation problem. Since the quantized discrete cosine transform coefficients of an H.264/AVC sequence tend to form a Laplace distribution, we therefore propose a reversible DH-based approach for intra-frame error concealment based on this characteristic. Our design is able to achieve no quality degradation. Experimental results demonstrate that the quality of recovered video sequences obtained by our approach is indeed superior to that of the DH-based method. In addition, the quality advantage of our approach is illustrated when compared with the previous five related methods. Kuo-Liang Chung, Yong-Huai Huang, Po-Chun Chang, Hong-Yuan Mark Liao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2009 | Video object inpainting using posture mappingabstractThis paper presents a novel framework for object-based video inpainting. To complete an occluded object, our method first samples a 3-D volume of the video into directional spatio-temporal slices, and then performs patch-based image inpainting to repair the partially damaged object trajectories in the 2-D slices. The completed slices are subsequently combined to obtain a sequence of virtual contours of the damaged object. The virtual contours and a posture sequence retrieval technique are then used to retrieve the most similar sequence of object postures in the available non-occluded postures. Key-posture selection and indexing are performed to reduce the complexity of posture sequence retrieval. We also propose a synthetic posture generation scheme that enriches the collection of key-postures so as to reduce the effect of insufficient key-postures. Our experimental results demonstrate that the proposed method can maintain the spatial consistency and temporal motion continuity of an object simultaneously. Chih-Hung Ling, Chia-Wen Lin, Chih-Wen Su, Hong-Yuan Mark Liao, Yong-Sheng Chen |
ICIP | 4 |
| 2009 | An online people counting system for electronic advertising machinesabstractThis paper presents a novel people counting system for an environment in which a stationary camera can count the number of people watching a TV-wall advertisement or an electronic billboard without counting the repetitions in video streams in real time. The people actually watching an advertisement are identified via frontal face detection techniques. To count the number of people precisely, a complementary set of features is extracted from the torso of a human subject, as that part of the body contains relatively richer information than the face. In addition, for conducting robust people recognition, an online classifier trained by Fisher's Linear Discriminant (FLD) strategy is developed. Our experiment results demonstrate the efficacy of the proposed system for the people counting task. Duan-Yu Chen, Chih-Wen Su, Yi-Chong Zeng, Shih-Wei Sun, Wei-Ru Lai, Hong-Yuan Mark Liao |
ICME | 6 |
| 2009 | Cooperative face hallucination using multiple referencesabstractThis paper proposes a cooperative example-based face hallucination method using multiple references. The proposed method first uses clustering and residual prototype faces construction to improve the performance of hallucinating a single low-resolution (LR) face to obtain a high-resolution (HR) counterpart. In the case that multiple LR face images for a person are available, a unique feature of the proposed method is to cooperatively enhance the qualities of hallucinated HR images by taking into account the multiple input face images jointly as prior models. Experimental results demonstrate that the proposed cooperative method achieve significant subjective and objective improvement over single-prior schemes. Chih-Chung Hsu, Chia-Wen Lin, Chiou-Ting Hsu, Hong-Yuan Mark Liao |
ICME | 4 |
| 2009 | Bayesian age estimation on face imagesabstractThis paper proposes to formulate the age estimation on face images as a Bayesian estimation problem. The proposed framework incorporates a probabilistic model on face aging distribution into the formulation. Individual-specific prior information as well as individual-specific aging models, if available, could then be easily included into the unified age estimation framework. We conduct experiments on the publicly available FG-NET database and compare the estimation results with existing methods. The experimental results demonstrate that, by combining the probabilistic modeling of facial feature distribution, the proposed method indeed outperforms the existing methods. Chung-Chun Wang, Yi-Chueh Su, Chiou-Ting Hsu, Chia-Wen Lin, Hong-Yuan Mark Liao |
ICME | 5 |
| 2009 | A Local Feature-based Human Motion Recognition FrameworkabstractIn this paper, we propose a local feature-based human motion analysis framework. Instead of using traditional analysis methods to characterize the global structure of human motion, we extract features directly from local regions that contain motion. To implement the above concept, we adopt the rules of visual attention theory, which assert that a human motion can be described simply by a set of local features comprised of spatial relationships rather than human postures. We select two kinds of features to represent the local variation of a human motion. First, we extract the long-term movement trend of the motion. The second feature is actually a set of rough features derived by sampling multi-scale moving edges. The two types of features are considered together during the recognition process. Our experiments demonstrate that the proposed approach can achieve very good recognition results. Yu-Chun Lai, Hong-Yuan Mark Liao, Cheng-Chung Lin, Jian-Ren Chen, Yi-Fei Peter Luo |
ISCAS | 2 |
| 2009 | Unsupervised Analysis of Human Behavior based on Manifold LearningabstractIn this paper, we propose a framework for unsupervised analysis of human behavior based on manifold learning. First, a pairwise human posture distance matrix is calculated from a training action sequence. Then, the isometric feature mapping (Isomap) algorithm is applied to construct a low-dimensional structure from the distance matrix. The data points in the Isomap space are consequently represented as a time-series of low-dimensional points. A temporal segmentation technique is then applied to segment the time series into subseries corresponding to atomic actions. Next, a dynamic time warping (DTW) approach is applied for clustering atomic action sequences. Finally, we use the clustering results to learn and classify atomic actions using the nearest neighbor rule. Experiments conducted on real data demonstrate the efficacy of the proposed method. Yu-Ming Liang, Sheng-Wen Shih, Arthur Chun-Chieh Shih, Hong-Yuan Mark Liao, Cheng-Chung Lin |
ISCAS | 4 |
| 2009 | Video Inpainting on Digitized Old Films
Nick C. Tang, Hong-Yuan Mark Liao, Chih-Wen Su, Fay Huang, Timothy K. Shih |
KES (2) | 2 |
| 2009 | Motion inpainting and extrapolation for special effect productionabstractSpecial effect production is an important technology for movie industry. Although sophisticated equipments can be used to produce scenery such as fire and smoke, usually, the approach is expensive and perhaps dangerous. Can sceneries be re-produced based on existing videos using digital technology? The answer is yes but it is difficult. In this video demonstration, we propose new mechanisms using motion inpainting technologies. Dynamic texture can be altered and reused under control. In addition, actors with repeated motion can be interpolated or extrapolated and incorporated into the falsified scenery. We demonstrate using our tool to produce realistic video narratives based on existing videos. Five sets of results as well as the source videos used are available at http://member.mine.tku.edu.tw/www/ACMMM09VideoDemo. Nick C. Tang, Hsing-Ying Zhong, Joseph C. Tsai, Timothy K. Shih, Hong-Yuan Mark Liao |
ACM Multimedia | 5 |
| 2009 | Integration of Color and Motion Features for Video RetrievalabstractThe usefulness of a video database depends on whether the video of interest can be easily located. In this paper, we propose a video retrieval algorithm based on the integration of several visual cues. In contrast to key-frame based representation of shot, our approach analyzes all frames within a shot to construct a compact representation of video shot. In the video matching step, by integrating the color and motion features, a similarity measure is defined to locate the occurrence of similar video clips in the database. Therefore, our approach is able to fully exploit the spatio-temporal information contained in video. Experimental results indicate that the proposed approach is effective and outperforms some existing technique. Liang-Hua Chen, Kuo-Hao Chin, Hong-Yuan Mark Liao |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2009 | Learning Atomic Human Actions Using Variable-Length Markov ModelsabstractVisual analysis of human behavior has generated considerable interest in the field of computer vision because of its wide spectrum of potential applications. Human behavior can be segmented into atomic actions, each of which indicates a basic and complete movement. Learning and recognizing atomic human actions are essential to human behavior analysis. In this paper, we propose a framework for handling this task using variable-length Markov models (VLMMs). The framework is comprised of the following two modules: a posture labeling module and a VLMM atomic action learning and recognition module. First, a posture template selection algorithm, based on a modified shape context matching technique, is developed. The selected posture templates form a codebook that is used to convert input posture sequences into discrete symbol sequences for subsequent processing. Then, the VLMM technique is applied to learn the training symbol sequences of atomic actions. Finally, the constructed VLMMs are transformed into hidden Markov models (HMMs) for recognizing input atomic actions. This approach combines the advantages of the excellent learning function of a VLMM and the fault-tolerant recognition ability of an HMM. Experiments on realistic data demonstrate the efficacy of the proposed system. Yu-Ming Liang, Sheng-Wen Shih, Arthur Chun-Chieh Shih, Hong-Yuan Mark Liao, Cheng-Chung Lin |
IEEE Trans. Syst. Man Cybern. Part B | 4 |
| 2008 | Smart tone reproductionabstractIn this paper, we propose an effective scheme to enhance the visual details at the minimal cost of user adjustments. The uprising importance of automatic tone reproduction comes from the increasing population of digital archive programs, which contains a large number of images/videos either old, irreproducible, or poorly captured. We attempt to solve above issues by a new local normalization step and an adaptive contrast assessment process. With those two processes, our method can effectively enhance poor quality regions and simultaneously preserving good quality ones with default parameter settings. The experimental results demonstrate that our method is superior to many existing algorithms when applied to aid digital archiving issues. Dun-Yu Hsiao, Hong-Yuan Mark Liao |
ICASSP | 2 |
| 2008 | Dynamic visual saliency modeling based on spatiotemporal analysisabstractProducing an appropriate extent of visually salient regions in video sequences is a challenging task. In this work, we propose a novel approach for modeling dynamic visual attention based on spatiotemporal analysis. Our model first detects salient points in three-dimensional video volumes, and then uses them as seeds to search the extent of salient regions in a motion attention map. To determine the extent of attended regions, the maximum entropy in the spatial domain is used to analyze the dynamics obtained from spatiotemporal analysis. The experiment results show that the proposed dynamic visual attention model can effectively detect visual saliency through successive video volumes. Duan-Yu Chen, Hsiao-Rong Tyan, Dun-Yu Hsiao, Sheng-Wen Shih, Hong-Yuan Mark Liao |
ICME | 5 |
| 2008 | A framework of spatio-temporal analysis for video surveillanceabstractThis paper presents a video surveillance system that is capable of detecting and classifying moving targets in real-time. The system extracts moving targets from a video stream and classifies them into predefined categories according to their spatiotemporal properties. Classification of the moving targets is completed via a combination of a temporal boosted classifier and spatiotemporal “motion energy” analysis. We illustrate that a temporal boosted classifier can be designed that successfully recognizes five object categories: person(s), bicycle, motorcycle, vehicle, and person with umbrella. The proposed temporal boosted classifier has the unique ability to improve weak classifiers by allowing them to make use of previous information when evaluating the current frame. In addition, we demonstrate a method to further process targets in the “person(s)” category to determine if they are single moving individuals or crowds. It is shown that this challenging task of moving crowd recognition can be effectively performed using spatiotemporal motion energies. Duan-Yu Chen, Kevin J. Cannons, Hsiao-Rong Tyan, Sheng-Wen Shih, Hong-Yuan Mark Liao |
ISCAS | 5 |
| 2008 | Target region-aware tone reproductionabstractIn this paper, we propose an automatic tone reproduction scheme. This scheme enhances the visual details of an image’s low contrast regions, while preserving the details of high contrast regions. Our method first segments an image into a number of clusters based on luminance information. Subsequently, our scheme calculates the local variation of each cluster. Using the set of computed local variations, we obtain an adaptive gain factor of each pixel to make details of the image visible. With this adaptive gain factor, our method can directly enhance low contrast regions without affecting high contrast ones. The experimental results demonstrate that our method is superior to many existing algorithms, especially when testing on images that contain both high and low contrast regions. The reason behind these results is clear: most tone reproduction algorithms recover the details of low contrast regions, but simultaneously suppress the high contrast regions. Our algorithm, however, recovers the visual details of low contrast regions without sacrificing the quality of high contrast regions. Most importantly, our proposed automatic tone reproduction scheme eliminates the effort of manual adjustment of parameters. Dun-Yu Hsiao, Hong-Yuan Mark Liao |
ISCAS | 2 |
| 2008 | Video enhancement based on saturation adjustment and contrast enhancementabstractWe propose a new algorithm to improve the visual appearance of compressed videos, which are derived and reproduced from old films. The new technique is especially useful in handling the digitized old films of digital archives. Two main techniques, saturation adjustment and contrast enhancement, are implemented to improve the chrominance and luminance of video frames on HSL color space. For contrast enhancement, we use weighted histogram separation (WHS) to improve the luminance. Subsequently, the saturation ratio transfer function is proposed to adjust the saturation level of every frame. Furthermore, the new calculation method for the error of inter-frame is adopted to avoid the incorrectly enhanced result and the intensive computation consumed in motion estimation. The experimental results demonstrate that the proposed method outperforms the existing approaches in terms of video quality. Yi-Chong Zeng, Hong-Yuan Mark Liao |
ISCAS | 2 |
| 2008 | Motion extrapolation for video story planningabstractWe create video scenes using existing videos. A panorama is generated from background video, with foreground objects removed by video inpainting technique. A video planning script is provided by the user on the panorama with accurate timing of actors. Actions of these actors are extensions of existing cyclic motions, such as walking, tracked and extrapolated using our motion analysis techniques. Although the types of video story generated are limited, however, it is possible to use the mechanism proposed to generate forgery videos. Interested readers should visit our tool demonstration and forgery videos at http://member.mine.tku.edu.tw/www/ACMMM08-VideoPlanning. Nick C. Tang, Timothy K. Shih, Hong-Yuan Mark Liao, Joseph C. Tsai, Hsing-Ying Zhong |
ACM Multimedia | 3 |
| 2008 | 3-D mesh representation and retrieval using Isomap manifoldabstractWe propose a compact 3-D object representation scheme that can greatly assist the search/retrieval process in a network environment. A 3-D mesh-based object is transformed into a new coordinate frame by using the Isomap (isometric feature mapping) method. During the transformation process, not only the structure of the salient parts of an object will be kept, but also the geometrical relationships will be preserved. From the viewpoint of cognitive psychology, the data distributed on the Isomap manifold can be regarded as a set of significant features of a 3-D mesh-based object. To perform efficient matching, we project the Isomap domain 3-D object onto two different 2-D maps, and the two 2-D feature descriptors are used as the basis to measure the degree of similarity between two 3-D mesh-based objects. Experiments demonstrate that the proposed method in retrieving similar 3-D models is very effective. Most importantly, the proposed 3-D mesh retrieval scheme is still valid even if a 3-D mesh undergoes a mesh simplification process. Jung-Shiong Chang, Arthur Chun-Chieh Shih, Hsueh-Yi Sean Lin, Hai-Feng Kao, Hong-Yuan Mark Liao, Wen-Hsien Fang |
MMSP | 5 |
| 2008 | Using dynamic programming to segment star shape based on human perception and optimization formulationabstractObjects existed in the world are usually composed of a finite number of primitive components. It’s always beneficial to decompose these objects into basic components when facing the indexing, recognition, or retrieval issues. In this work we intended to devise an algorithm to effectively and efficiently decompose 2D projections of objects into primitive components. This decomposition process is important in different application domains, such as 3D object recognition, human behavior analysis, and content-based image retrieval. Hai-Feng Kao, Hong-Yuan Mark Liao |
MMSP | 2 |
| 2008 | Movie scene segmentation using background information
Liang-Hua Chen, Yu-Chun Lai, Hong-Yuan Mark Liao |
Pattern Recognit. | 3 |
| 2008 | Special Issue on Video SurveillanceabstractThe 14 regular papers and two brief papers in this special issue capture some of the state-of-the-art research on video surveillance issues, provide comprehensive overview of existing techniques, and propose novel solutions for important research problems. Ishfaq Ahmad 0001, Zhihai He, Hong-Yuan Mark Liao, Fernando Pereira 0001, Ming-Ting Sun |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2008 | Spatiotemporal Motion Analysis for the Detection and Classification of Moving TargetsabstractThis paper presents a video surveillance system in the environment of a stationary camera that can extract moving targets from a video stream in real time and classify them into predefined categories according to their spatiotemporal properties. Targets are detected by computing the pixel-wise difference between consecutive frames, and then classified with a temporally boosted classifier and ldquospatiotemporal-oriented energyrdquo analysis. We demonstrate that the proposed classifier can successfully recognize five types of objects: a person, a bicycle, a motorcycle, a vehicle, and a person with an umbrella. In addition, we process targets that do not match any of the AdaBoost-based classifier's categories by using a secondary classification module that categorizes such targets as crowds of individuals or non-crowds. We show that the above classification task can be performed effectively by analyzing a target's spatiotemporal-oriented energies, which provide a rich description of the target's spatial and dynamic features. Our experiment results demonstrate that the proposed system is extremely effective in recognizing all predefined object classes. Duan-Yu Chen, Kevin J. Cannons, Hsiao-Rong Tyan, Sheng-Wen Shih, Hong-Yuan Mark Liao |
IEEE Trans. Multim. | 5 |
| 2008 | Video-Based Human Movement Analysis and Its Application to Surveillance SystemsabstractThis paper presents a novel posture classification system that analyzes human movements directly from video sequences. In the system, each sequence of movements is converted into a posture sequence. To better characterize a posture in a sequence, we triangulate it into triangular meshes, from which we extract two features: the skeleton feature and the centroid context feature. The first feature is used as a coarse representation of the subject, while the second is used to derive a finer description. We adopt a depth-first search (dfs) scheme to extract the skeletal features of a posture from the triangulation result. The proposed skeleton feature extraction scheme is more robust and efficient than conventional silhouette-based approaches. The skeletal features extracted in the first stage are used to extract the centroid context feature, which is a finer representation that can characterize the shape of a whole body or body parts. The two descriptors working together make human movement analysis a very efficient and accurate process because they generate a set of key postures from a movement sequence. The ordered key posture sequence is represented by a symbol string. Matching two arbitrary action sequences then becomes a symbol string matching problem. Our experiment results demonstrate that the proposed method is a robust, accurate, and powerful tool for human movement analysis. Jun-Wei Hsieh, Yung-Tai Hsu, Hong-Yuan Mark Liao, Chih-Chiang Chen |
IEEE Trans. Multim. | 3 |
| 2007 | Human Action Recognition Using 2-D Spatio-Temporal TemplatesabstractA framework for human action modeling and recognition in continuous action sequences is proposed. A star figure enclosed by a bounding convex polygon is used to effectively represent the extremities of the silhouette of a human body. Thus, human actions are recorded as a sequence of the star figure's parameters, which is then used for action modeling. To model human actions in a compact manner while characterizing their spatio-temporal patterns, star figure parameters are represented by a 2-D feature map, which is used and regarded as a spatio-temporal template. Experiments to evaluate the performance of the proposed framework show that it can recognize human actions in an efficient and effective manner. Duan-Yu Chen, Sheng-Wen Shih, Hong-Yuan Mark Liao |
ICME | 3 |
| 2007 | Principal Component Analysis-based Mesh DecompositionabstractIn this paper, we propose an automatic mesh decomposition technique based on principal component analysis (PCA) and Boolean operations. First, we calculate the normalized protrusion degree of each dual vertex on the smoothed 3-D mesh. The protrusion degree of a vertex and the vertex's 3-D coordinates form a 4-D feature vector, which we use to represent the polygon mesh. Since a 3-D object is composed of a large number of polygon meshes, we apply PCA to the set of 4-D feature vectors. We take the axes corresponding to the top three principal components as the three axes of a new coordinate system and project the set of 4-D vectors onto the system. Surprisingly, the projected data along the first axis reveals the salient structures of the 3-D object. Therefore, using the first component axis as the search basis, we can identify all the salient parts of an arbitrary 3-D object. Jung-Shiong Chang, Arthur Chun-Chieh Shih, Hong-Yuan Mark Liao, Wen-Hsien Fang |
MMSP | 3 |
| 2007 | A Language Modeling Approach to Atomic Human Action RecognitionabstractVisual analysis of human behavior has generated considerable interest in the field of computer vision because it has a wide spectrum of potential applications. Atomic human action recognition is an important part of a human behavior analysis system. In this paper, we propose a language modeling framework for this task. The framework is comprised of two modules: a posture labeling module, and an atomic action learning and recognition module. A posture template selection algorithm is developed based on a modified shape context matching technique. The posture templates form a codebook that is used to convert input posture sequences into training symbol sequences or recognition symbol sequences. Finally, a variable-length Markov model technique is applied to learn and recognize the input symbol sequences of atomic actions. Experiments on real data demonstrate the efficacy of the proposed system. Yu-Ming Liang, Sheng-Wen Shih, Arthur Chun-Chieh Shih, Hong-Yuan Mark Liao, Cheng-Chung Lin |
MMSP | 4 |
| 2007 | Visual Salience-Guided Mesh DecompositionabstractIn this paper, we propose a novel mesh-decomposition scheme called "visual salience-guided mesh decomposition". The concept of "part salience", which originated in cognitive psychology, asserts that the salience of a part can be determined by (at least) three factors: the protrusion, the boundary strength, and the relative size of the part. We try to convert these conceptual rules into real computational processes, and use them to guide a three-dimensional (3D) mesh decomposition process in such a way that the significant components can be precisely identified and efficiently extracted from a given 3D mesh. The proposed decomposition scheme not only identifies the parts' boundaries defined by the minima rule, but also labels each part with a quantitative degree of visual salience during the mesh decomposition process. The experimental results show that the proposed scheme is indeed effective and powerful in decomposing a 3D mesh into its significant components Hsueh-Yi Sean Lin, Hong-Yuan Mark Liao, Ja-Chen Lin |
IEEE Trans. Multim. | 2 |
| 2007 | Motion Flow-Based Video RetrievalabstractIn this paper, we propose the use of motion vectors embedded in MPEG bitstreams to generate so-called ldquomotion flowsrdquo, which are applied to perform video retrieval. By using the motion vectors directly, we do not need to consider the shape of a moving object and its corresponding trajectory. Instead, we simply ldquolinkrdquo the local motion vectors across consecutive video frames to form motion flows, which are then recorded and stored in a video database. In the video retrieval phase, we propose a new matching strategy to execute the video retrieval task. Motions that do not belong to the mainstream motion flows are filtered out by our proposed algorithm. The retrieval process can be triggered by query-by-sketch or query-by-example. The experiment results show that our method is indeed superb in the video retrieval process. Chih-Wen Su, Hong-Yuan Mark Liao, Hsiao-Rong Tyan, Chia-Wen Lin, Duan-Yu Chen, Kuo-Chin Fan |
IEEE Trans. Multim. | 2 |
| 2006 | Real-time event detection and its application to surveillance systemsabstractIn recent years, real-time direct detection of events by surveillance systems has attracted a great deal of attention. In this paper, we propose a new video-based surveillance system that can perform real-time event detection. In the background modeling phase, we adopt a mixture of Gaussian approach to determine the background. Meanwhile, we use color blob-based tracking to track foreground objects. Due to the self-occlusion problem, the tracking module is designed as a multi-blob tracking process to obtain similar multiple trajectories. We devise an algorithm to merge these trajectories into a representative one. After applying the Douglas-Peucker algorithm to approximate a trajectory, we can compare two arbitrary trajectories. The above mechanism enables us to conduct real-time event detection if a number of wanted trajectories are pre-stored in a video surveillance system. Hong-Yuan Mark Liao, Duan-Yu Chen, Chih-Wen Su, Hsiao-Rong Tyan |
ISCAS | 1 |
| 2006 | Continuous Human Action Segmentation and Recognition Using a Spatio-Temporal Probabilistic FrameworkabstractIn this paper, a framework of automatic human action segmentation and recognition in continuous action sequences is proposed. A star-like figure is proposed to effectively represent the extremities in the silhouette of human body. The human action, thus, is recorded as a sequence of the star-like figure parameters, which is used for action modeling. To model human actions in a compact manner while characterizing their spatio-temporal distributions, star-like figure parameters are represented by Gaussian mixture models (GMM). In addition, to address the intrinsic nature of temporal variations in a continuous action sequence, we transform the time sequence of star-like figure parameters into frequency domain by discrete cosine transform (DCT) and use only the first few coefficients to represent different temporal patterns with significant discriminating power. The performance shows that the proposed framework can recognize continuous human actions in an efficient way Duan-Yu Chen, Hong-Yuan Mark Liao, Sheng-Wen Shih |
ISM | 2 |
| 2006 | Fast coarse-to-fine video retrieval using shot-level spatio-temporal statisticsabstractIn this paper, we propose a fast coarse-to-fine video retrieval scheme using shot-level spatio-temporal statistics. The scheme consists of a two-step coarse search followed by a fine search. In the coarse search stage, the shot-level motion and color distribution is computed as spatio-temporal features for shot matching. The first-step coarse search uses the shot-level global statistics to reduce the size of the search space drastically. By adding an adjacent shot of the first query shot, the second-step coarse search introduces a "causality" relation between two consecutive shots to improve the search accuracy. Finally, the fine-search step refines the search result by using the local color features extracted from the key frames of the query shots. Our experimental results show that the proposed method achieves good retrieval performance with a much reduced complexity compared to single-pass methods. Yu-Hsuan Ho, Chia-Wen Lin, Jing-Fung Chen, Hong-Yuan Mark Liao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2005 | Fast Video Retrieval via the Statistics of MotionabstractDue to the popularity of the Internet and the powerful computing capability of computers, efficient processing/retrieval of multimedia data has become an important issue. In this paper, we propose a fast video retrieval algorithm that bases its search core on the statistics of object motion. The algorithm starts with extracting object motions from a shot and then transforms/quantizes them into the form of probability distributions. By choosing the shot that has the largest entropy value among the constituent shots of an unknown query video clip, we execute the first stage video search. By comparing two shots with different lengths, their corresponding motion probability distributions are compared by a discrete Bhattacharyya distance which is designed to measure the similarity between any two distribution functions. In the second stage, we add an adjacent shot (either preceding or subsequent) to perform a finer comparison. Experimental results demonstrate that our fast video retrieval algorithm is powerful in terms of accuracy and efficiency. Jing-Fung Chen, Hong-Yuan Mark Liao, Chia-Wen Lin |
ICASSP (2) | 2 |
| 2005 | A cognitive psychology-based approach for 3-D shape retrievalabstractIn this paper, we incorporate a set of principles originated from cognitive psychology into the design of 3-D shape analysis and retrieval algorithms. Based on the "visual salience-guided mesh decomposition" we previously proposed, a 3-D mesh-based shape is first broken up into parts such that human visual perception on parts can be appropriately mimicked. Next, the decomposed parts are individually analyzed and quantified according to the properties of visual salience. To establish the indices of 3-D meshes for shape matching, spherical parameterization is adopted to map the decomposed parts onto the surface of a unit sphere. In this way, the dissimilarity degree between the query acquired from users and a model in database can be calculated. The experimental results have shown that the matching performance of the proposed scheme is indeed efficient and powerful. Hsueh-Yi Sean Lin, Hong-Yuan Mark Liao, Ja-Chen Lin |
ICME | 2 |
| 2005 | 3-D Shape Retrieval Using Cognitive Psychology-based PrinciplesabstractIn this paper, we incorporate a set of principles that originated in cognitive psychology into the design of 3D shape analysis and retrieval systems. Based on the "visual salience guided mesh decomposition" scheme we proposed previously, a 3D shape represented in mesh form is first broken into parts such that human visual perception of the parts can be appropriately mimicked. Next, the decomposed parts are individually analyzed and quantified according to the properties of visual salience. To establish the indices of 3D meshes for the subsequent retrieval process, spherical parameterization is adopted to map the decomposed parts onto the surface of a unit sphere. In this way, the degree of similarity between a query provided by a user and models in the database can be calculated. The experiment results show that the retrieval performance of the proposed scheme is indeed efficient and powerful. Hsueh-Yi Sean Lin, Ja-Chen Lin, Hong-Yuan Mark Liao |
ISM | 3 |
| 2005 | Using Normal Vectors for Stereo Correspondence Construction
Jung-Shiong Chang, Arthur Chun-Chieh Shih, Hong-Yuan Mark Liao, Wen-Hsien Fang |
KES (1) | 3 |
| 2005 | Robust Video Retrieval Using Temporal MVMB Moments
Duan-Yu Chen, Hong-Yuan Mark Liao, Suh-Yin Lee |
KES (3) | 2 |
| 2005 | Fast Video Retrieval via the Statistics of Motion Within the Regions-of-Interest
Jing-Fung Chen, Hong-Yuan Mark Liao, Chia-Wen Lin |
KES (3) | 2 |
| 2005 | Automatic Key Posture Selection for Human Behavior AnalysisabstractA novel human posture analysis framework that can perform automatic key posture selection and template matching for human behavior analysis is proposed. The entropy measurement, which is commonly adopted as an important feature to describe the degree of disorder in thermodynamics, is used as an underlying feature for identifying key postures. First, we use cumulative entropy change as an indicator to select an appropriate set of key postures from a human behavior video sequence and then conduct a cross entropy check to remove redundant key postures. With the key postures detected and stored as human posture templates, the degree of similarity between a query posture and a database template is evaluated using a modified Hausdorff distance measure. The experiment results show that the proposed system is highly efficient and powerful Duan-Yu Chen, Hong-Yuan Mark Liao, Hsiao-Rong Tyan, Chia-Wen Lin |
MMSP | 2 |
| 2005 | Human Behavior Analysis Using Deformable TriangulationsabstractThis paper presents a new posture classification system to analyze different human behaviors directly from video sequences using the technique of triangulation. For well analyzing each posture in the video sequences, we propose a triangulation-based method to triangulate it to different triangle meshes from which two important posture features are then extracted, i.e., the ones of skeleton and centroid context. The first one is used for a coarse search and the second one is for a finer classification to classify postures in more details. For the first descriptor, we take advantages of a dfs (depth-first search) scheme to extract the skeleton features of a posture from its triangulation result. Then, with the help of skeleton information, we can define a new shape descriptor, i.e., centroid context, to describe a posture up to a semantic level. That is, the centroid context is a finer descriptor to describe a posture not only from its whole shape but also from its body parts. Since the two descriptors are complement to each other, all desired human postures can be compared and classified very accurately. The nice ability of posture classification can help us generate a set of key postures for transferring a behavior sequence to a set of symbols. Then, a novel string matching scheme is proposed to analyze different human behaviors. Experimental results have proved that the proposed method is robust, accurate, and powerful in human behavior analysis Yung-Tai Hsu, Jun-Wei Hsieh, Hai-Feng Kao, Hong-Yuan Mark Liao |
MMSP | 4 |
| 2005 | Robust video sequence retrieval using a novel object-based T2D-histogram descriptor
Duan-Yu Chen, Suh-Yin Lee, Hong-Yuan Mark Liao |
J. Vis. Commun. Image Represent. | 3 |
| 2005 | Fragile watermarking for authenticating 3-D polygonal meshesabstractDesigning a powerful fragile watermarking technique for authenticating three-dimensional (3-D) polygonal meshes is a very difficult task. Yeo and Yeung were first to propose a fragile watermarking method to perform authentication of 3-D polygonal meshes. Although their method can authenticate the integrity of 3-D polygonal meshes, it cannot be used for localization of changes. In addition, it is unable to distinguish malicious attacks from incidental data processings. In this paper, we trade off the causality problem in Yeo and Yeung's method for a new fragile watermarking scheme. The proposed scheme can not only achieve localization of malicious modifications in visual inspection, but also is immune to certain incidental data processings (such as quantization of vertex coordinates and vertex reordering). During the process of watermark embedding, a local mesh parameterization approach is employed to perturb the coordinates of invalid vertices while cautiously maintaining the visual appearance of the original model. Since the proposed embedding method is independent of the order of vertices, the hidden watermark is immune to some attacks, such as vertex reordering. In addition, the proposed method can be used to perform region-based tampering detection. The experimental results have shown that the proposed fragile watermarking scheme is indeed powerful. Hsueh-Yi Sean Lin, Hong-Yuan Mark Liao, Chun-Shien Lu, Ja-Chen Lin |
IEEE Trans. Multim. | 2 |
| 2005 | A motion-tolerant dissolve detection algorithmabstractGradual shot change detection is one of the most important research issues in the field of video indexing/retrieval. Among the numerous types of gradual transitions, the dissolve-type gradual transition is considered the most common one, but it is also the most difficult one to detect. In most of the existing dissolve detection algorithms, the false/miss detection problem caused by motion is very serious. In this paper, we present a novel dissolve-type transition detection algorithm that can correctly distinguish dissolves from disturbance caused by motion. We carefully model a dissolve based on its nature and then use the model to filter out possible confusion caused by the effect of motion. Experimental results show that the proposed algorithm is indeed powerful. Chih-Wen Su, Hong-Yuan Mark Liao, Hsiao-Rong Tyan, Kuo-Chin Fan, Liang-Hua Chen |
IEEE Trans. Multim. | 2 |
| 2004 | Visual salience-guided mesh decompositionabstractIn this paper, we propose a novel mesh decomposition scheme called "visual salience-guided mesh decomposition". The concept of part salience is originated from cognitive psychology and it asserts the salience of a part can be determined by (at least) three factors: the protrusion, the boundary strength, and the relative size of a part. We try to convert these "conceptual" rules into "real" computational processes and then use them to "guide" a 3D mesh decomposition process. The experimental results have shown that the proposed scheme is indeed effective and powerful in decomposing a 3D mesh into significant components. Hsueh-Yi Sean Lin, Hong-Yuan Mark Liao, Ja-Chen Lin |
MMSP | 2 |
| 2004 | Pedestrian detection and tracking at crossroads
Chia-Jung Pai, Hsiao-Rong Tyan, Yu-Ming Liang, Hong-Yuan Mark Liao, Sei-Wang Chen |
Pattern Recognit. | 4 |
| 2004 | Extraction of video object with complex motion
Liang-Hua Chen, Yu-Chun Lai, Chih-Wen Su, Hong-Yuan Mark Liao |
Pattern Recognit. Lett. | 4 |
| 2003 | Pedestrian detection and tracking at crossroadsabstractThis paper presents a system for pedestrian detection and tracking by using image processing techniques. A very important issue in the field of intelligent transportation system is to prevent pedestrians from being hit by vehicles. Recently, a great number of vision-based techniques have been proposed for this purpose. In this paper, we propose a vision-based method which combines the use of a pedestrian model as well as the walking rhythm of pedestrians. Through integrating these spatial and temporal information grabbed by a vision system, we are able to develop a reliable pedestrian detection and tracking system that can always produce accurate results. Experimental results obtained using the real world cases have demonstrated that the proposed model is indeed superb. Chia-Jung Pai, Hsiao-Rong Tyan, Yu-Ming Liang, Hong-Yuan Mark Liao, Sei-Wang Chen |
ICIP (2) | 4 |
| 2003 | Authentication of 3-D Polygonal Meshes
Hsueh-Yi Sean Lin, Hong-Yuan Mark Liao, Chun-Shien Lu, Ja-Chen Lin |
IWDW | 2 |
| 2003 | On the preview of digital movies
Liang-Hua Chen, Chih-Wen Su, Hong-Yuan Mark Liao, Arthur Chun-Chieh Shih |
J. Vis. Commun. Image Represent. | 3 |
| 2003 | A message-based cocktail watermarking system
Gwo-Jong Yu, Chun-Shien Lu, Hong-Yuan Mark Liao |
Pattern Recognit. | 3 |
| 2003 | A new iterated two-band diffusion equation: theory and its applicationabstractIn this paper, we propose an iterated two-band filtering method to solve the selective image smoothing problem. We prove that a discrete computation step in an iterated nonlinear diffusion-based filtering algorithm is equivalent to a sequence of operations, including decomposition, regularization, and then reconstruction, in the proposed two-band filtering scheme. To correctly separate the high frequency components from the low frequency ones in the decomposition process, we adopt a dyadic wavelet-based approximation scheme. In the regularization process, we use a diffusivity function as a guide to retain useful data and suppress noises. Finally, the signal of the next stage, which is a "smoother" version of the signal at the previous stage, can be computed by reconstructing the decomposed low frequency component and the regularized high frequency component. Based on the proposed scheme, the smoothing operation can be applied to the correct targets. Experimental results show that our new approach is really efficient in noise removing. Arthur Chun-Chieh Shih, Hong-Yuan Mark Liao, Chun-Shien Lu |
IEEE Trans. Image Process. | 2 |
| 2003 | Structural digital signature for image authentication: an incidental distortion resistant schemeabstractThe existing digital data verification methods are able to detect regions that have been tampered with, but are too fragile to resist incidental manipulations. This paper proposes a new digital signature scheme which makes use of an image's contents (in the wavelet transform domain) to construct a structural digital signature (SDS) for image authentication. The characteristic of the SDS is that it can tolerate content-preserving modifications while detecting content-changing modifications. Many incidental manipulations, which were detected as malicious modifications in the previous digital signature verification or fragile watermarking schemes, can be bypassed in the proposed scheme. Performance analysis is conducted and experimental results show that the new scheme is indeed superb for image authentication. Chun-Shien Lu, Hong-Yuan Mark Liao |
IEEE Trans. Multim. | 2 |
| 2002 | A motion-tolerant dissolve detection algorithmabstractGradual shot change detection is one of the most important research issues in the field of video indexing/retrieval. Among the numerous types of gradual transitions, dissolve is considered the most common one, but also the most difficult to be detected one. It's well known that an efficient dissolve detection algorithm which can be executed on a real video is still deficient. In this paper, we present a novel dissolve detection algorithm that can efficiently detect dissolves with different durations. In addition, global motions caused by camera movement and local motions caused by object movement can be discriminated from a real dissolve by our algorithm. The experimental results show that the new method is indeed powerful. Chih-Wen Su, Hsiao-Rong Tyan, Hong-Yuan Mark Liao, L. H. Chen |
ICME (2) | 3 |
| 2002 | Wavelet-based optical flow estimationabstractA new algorithm for accurate optical flow (OF) estimation using discrete wavelet approximation is proposed. The computation of OF depends on minimizing the image and smoothness constraints. The proposed method takes advantages of the nature of wavelet theory, which can efficiently and accurately approximate any function. OF vectors and image functions are represented by means of linear combinations of scaling basis functions. Based on such wavelet-based approximation, the leading coefficients of these basis functions carry global information about the approximated functions. The proposed method can successfully convert the problem of minimizing a constraint function into that of solving a linear system of a quadratic and convex function of scaling coefficients. Once all the corresponding coefficients are determined, the flow vectors can be obtained accordingly. Experiments have been conducted on both synthetic and real image sequences. In terms of accuracy, the results show that our approach outperforms the existing methods which adopted the same objective function as ours. Li-Fen Chen, Hong-Yuan Mark Liao, Ja-Chen Lin |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2002 | Denoising and copy attacks resilient watermarking by exploiting prior knowledge at detectorabstractWatermarking with both oblivious detection and high robustness capabilities is still a challenging problem. In this paper, we tackle the aforementioned problem. One easy way to achieve blind detection is to use denoising for filtering out the hidden watermark, which can be utilized to create either a false positive (copy attack) or false negative (denoising and remodulation attack). Our basic design methodology is to exploit prior knowledge available at the detector side and then use it to design a "nonblind" embedder. We prove that the proposed scheme can resist two famous watermark estimation-based attacks, which have successfully cracked many existing watermarking schemes. False negative and false positive analyses are conducted to verify the performance of our scheme. The experimental results show that the new method is indeed powerful. Chun-Shien Lu, Hong-Yuan Mark Liao, Martin Kutter |
IEEE Trans. Image Process. | 2 |
| 2001 | Person identification using facial motionabstractA new method for person identification based on facial motion is proposed. Facial motion is represented by a high-dimensional feature vector which is constructed by concatenating a sequence of motion flow fields. Under the proposed representation, each individual can be characterized as a feature vector which collects spatial and temporal information of a face simultaneously We then make use of this high-dimensional spatiotemporal feature vector to perform person identification. The experimental results are encouraging and exciting in terms of robustness under varying illumination conditions. Li-Fen Chen, Hong-Yuan Mark Liao, Ja-Chen Lin |
ICIP (2) | 2 |
| 2001 | Video object-based watermarking: a rotation and flipping resilient schemeabstractVideo object (VO) is a very important concept in the MPEG-4 standard. Video objects may be purposely cut and pasted for illegal use. A robust watermarking scheme for video object protection is proposed. For each segmented video object, a watermark is embedded by a new technology whose design is based on the concept of communications with side information. To solve the asynchronous problem caused by object placement, we propose to use eigenvectors of a video object for synchronization of rotation and flipping. Preliminary results have demonstrated the robustness of the proposed method. Chun-Shien Lu, Hong-Yuan Mark Liao |
ICIP (2) | 2 |
| 2001 | A message-based cocktail watermarking systemabstractA noise-type Gaussian sequence is most commonly used as a watermark to claim ownership of media data. However, only a 1 bit information payload is carried in this type of watermark. For a logo-type watermark, the situation is better because it is visually recognizable and more information can be carried. However, since the sizes and shapes of logos for different organizations are different, the flexibility of use of a logo-type watermark will certainly be degraded. We design a more flexible type of watermark, i.e., a message. Since a message is composed of a finite number of ASCII-type characters, it is by nature vulnerable to attacks. Therefore, we propose to choose a set of nonlinear Hadamard codes that has the maximum Hamming distance between any two constituent codes to replace the original ASCII-type inputs. This design will make our system much more fault-tolerant in comparison with ASCII-code based systems under direct attack. To recover an attacked Hadamard code, we use a trained backpropagation neural network to perform inexact matching. Experimental results demonstrate that our message-based cocktail watermarking system is superb in terms of robustness and flexibility. Gwo-Jong Yu, Chun-Shien Lu, Hong-Yuan Mark Liao |
ICIP (3) | 3 |
| 2001 | A new watermarking scheme resistant to denoising and copy attacksabstractWatermarking with both oblivious detection and high robustness capabilities is still a challenging problem for copyright protection up to now. In order to tackle the above mentioned problem we propose to exploit prior knowledge available at the watermark detector side to design a "non-blind" embedder. We prove that the proposed scheme can resist two famous denoising-based attacks, which have successfully cracked many existing watermarking schemes. Chun-Shien Lu, Hong-Yuan Mark Liao, Martin Kutter |
MMSP | 2 |
| 2001 | MCE-Based Face RecognitionabstractThis paper proposes a complete procedure for the extraction and recognition of human faces in complex scenes. The morphology-based face detection algorithm can locate multiple faces oriented in any direction. The recognition algorithm is based on the minimum classification error (MCE) criterion. In our work, the minimum classification error formulation is incorporated into a multilayer perceptron neural network. Experimental results show that our system is robust to noisy images and complex background. Liang-Hua Chen, Peter Shaohua Deng, Hong-Yuan Mark Liao |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2001 | Why recognition in a statistics-based face recognition system should be based on the pure face portion: a probabilistic decision-based proof
Li-Fen Chen, Hong-Yuan Mark Liao, Ja-Chen Lin, Chin-Chuan Han |
Pattern Recognit. | 2 |
| 2001 | Multipurpose watermarking for image authentication and protectionabstractWe propose a novel multipurpose watermarking scheme, in which robust and fragile watermarks are simultaneously embedded, for copyright protection and content authentication. By quantizing a host image's wavelet coefficients as masking threshold units (MTUs), two complementary watermarks are embedded using cocktail watermarking and they can be blindly extracted without access to the host image. For the purpose of image protection, the new scheme guarantees that, no matter what kind of attack is encountered, at least one watermark can survive well. On the other hand, for the purpose of image authentication, our approach can locate the part of the image that has been tampered with and tolerate some incidental processes that have been executed. Experimental results show that the performance of our multipurpose watermarking scheme is indeed superb in terms of robustness and fragility. Chun-Shien Lu, Hong-Yuan Mark Liao |
IEEE Trans. Image Process. | 2 |
| 2001 | A new image flux conduction model and its application to selective image smoothingabstractA discrete image flux conduction equation which is completely new in this field is proposed. The new approach starts with formulating a discrete image flux conduction equation based on the concept of heat conduction theory. Based on this discrete equation, the status change at a time point can be directly computed from its spatial neighborhood. To more accurately estimate an image flux, we have used an orthogonal wavelet basis to approximate the gradient of the intensity at each point. Since the proposed approach is discrete by nature, it is not necessary to formulate a continuous PDE to fit the discrete image data set. Furthermore, introduction of different numerical methods to solve the PDE can also be avoided. Since the proposed approach does not require that a PDE be solved, it is therefore more efficient and accurate than the conventional methods. Experimental results obtained using both synthetic signals and real images have demonstrated that the proposed model could effectively handle the selective image smoothing problem. Chwen-Jye Sze, Hong-Yuan Mark Liao, Kuo-Chin Fan |
IEEE Trans. Image Process. | 2 |
| 2000 | Oblivious Cocktail Watermarking by Sparse Code Shrinkage: A Regional- and Global-Based SchemeabstractWatermarking with oblivious detection and high robustness capabilities together is still a challenging problem up to now. The existing methods are either robust or oblivious but it is difficult to achieve both goals simultaneously. In this paper, oblivious detection is formulated as a blind source separation problem by regarding the hidden watermarks as noise. To keep high robustness, our non-oblivious cocktail watermarking scheme, which is very robust, is adopted to combine with the oblivious detection mechanism. The proposed oblivious cocktail watermarking can be applied to global watermarking and regional watermarking. Experimental results have demonstrated the powerfulness of our method. Chun-Shien Lu, Hong-Yuan Mark Liao |
ICIP | 2 |
| 2000 | Dyadic Wavelet-Based Nonlinear Conduction Equation: Theory and ApplicationsabstractWe proposed a new dyadic wavelet-based conduction approach to take the place of the nonlinear diffusion equation for selective image smoothing. We also proved that the proposed iterated system always satisfies the so-called maximum-minimum principle no matter what kind of wavelet basis is used. Since the proposed approach does not require one to solve a partial differential equation (PDE), it is therefore more efficient and accurate than the conventional nonlinear diffusion/conduction-based methods. Experimental results using 1-D synthetic data and a real image demonstrated that the proposed method can efficiently remove noise and preserve real data. Chwen-Jye Sze, Hong-Yuan Mark Liao, Shih-Kun Huang, Chun-Shien Lu |
ICIP | 2 |
| 2000 | Mean Quantization Blind Watermarking for Image AuthenticationabstractThe objective of this paper is to propose an image authentication scheme, which is able to detect malicious tampering of images even they have also been incidentally distorted. By modeling incidental and malicious distortions as Gaussian distributions with small and large variances, respectively, we propose to embed a watermark in the wavelet domain by a mean quantization technique. Due to the various probabilities of tamper response at each scale, these responses are integrated to make a decision on the tampered areas. Statistical analysis is conducted and experimental results are given to demonstrate that our watermarking scheme is able to detect malicious attacks while tolerating incidental distortions. Gwo-Jong Yu, Chun-Shien Lu, Hong-Yuan Mark Liao, Jang-Ping Sheu |
ICIP | 3 |
| 2000 | Object Segmentation for Video CodingabstractObject segmentation and tracking are problems within the scope of MPEG-4 standardization activities. This paper proposes a novel algorithm to extract and track the moving objects in the video sequences. A model-based tracking method is developed to localize the object in each frame of a video sequence. One distinctive property of this method is that the model of an object is acquired dynamically from the video sequence, rather than being provided a priori. Our method also allows the object to have large movement between two frames and to be partially occluded in several successive frames. To extract contiguous object boundaries, an active contour model is then applied to the binary image representation of the object model. Compared with the related work of Neri et al. (1998), the proposed algorithm can achieve a more accurate location of object boundaries. Liang-Hua Chen, Jan-Ru Chen, Hong-Yuan Mark Liao |
ICPR | 3 |
| 2000 | Wavelet-Based Optical Flow EstimationabstractIn this paper, a new algorithm for accurate optical flow estimation using discrete wavelet approximation is proposed. The proposed method takes advantages of the nature of wavelet theory, which can efficiently and accurately represent "things", to model optical flow vectors and image related functions. Each flow vector and image function are represented by linear combinations of wavelet basis functions. From such wavelet-based approximation, the leading coefficients of these basis functions carry the global information of the approximated "things". The proposed method can successfully convert the problem of minimizing a constraint function into that of solving a linear system of a quadratic and convex function of wavelet coefficients. Once all the corresponding coefficients are decided, the flow vectors can be determined accordingly. Experiments conducted on both synthetic and real image sequences show that our approach outperformed the existing methods in terms of accuracy. Li-Fen Chen, Ja-Chen Lin, Hong-Yuan Mark Liao |
ICPR | 3 |
| 2000 | Multipurpose Audio WatermarkingabstractAn audio protection and authentication scheme is proposed. By quantizing a host audio's FFT-coefficients as masking threshold units (MTUs), two complementary watermarks are designed and embedded using our cocktail watermarking method. For audio protection, high robustness can be achieved; whereas for audio authentication, tampered regions can be detected. Both of the above mentioned goals are accomplished in an oblivious manner. Experimental results indicate that our multipurpose audio watermarking scheme is remarkably effective. Chun-Shien Lu, Hong-Yuan Mark Liao, Liang-Hua Chen |
ICPR | 2 |
| 2000 | A new LDA-based face recognition system which can solve the small sample size problem
Li-Fen Chen, Hong-Yuan Mark Liao, Ming-Tat Ko, Ja-Chen Lin, Gwo-Jong Yu |
Pattern Recognit. | 2 |
| 2000 | Fast face detection via morphology-based pre-processing
Chin-Chuan Han, Hong-Yuan Mark Liao, Gwo-Jong Yu, Liang-Hua Chen |
Pattern Recognit. | 2 |
| 2000 | Cocktail Watermarking for Digital Image ProtectionabstractA novel image protection scheme called "cocktail watermarking" is proposed in this paper. We analyze and point out the inadequacy of the modulation techniques commonly used in ordinary spread spectrum watermarking methods and the visual model-based ones. To resolve the inadequacy, two watermarks which play complementary roles are simultaneously embedded into a host image. We also conduct a statistical analysis to derive the lower bound of the worst likelihood that the better watermark (out of the two) can be extracted. With this "high" lower bound, it is ensured that a "better" extracted watermark is always obtained. From extensive experiments, results indicate that our cocktail watermarking scheme is remarkably effective in resisting various attacks, including combined ones. Chun-Shien Lu, Shih-Kun Huang, Chwen-Jye Sze, Hong-Yuan Mark Liao |
IEEE Trans. Multim. | 4 |
| 1999 | Wavelet-Based Off-Line Handwritten Signature Verification
Peter Shaohua Deng, Hong-Yuan Mark Liao, Chin Wen Ho, Hsiao-Rong Tyan |
Comput. Vis. Image Underst. | 2 |
| 1999 | Automatic data capture for geographic information systemsabstractWe present a map interpretation system for automatic extraction of high level information from the scanned images of Chinese land register maps. Our map interpretation system consists of three main components: text/graphics separation, parcel extraction, and rotated character recognition. Our approach to text/graphics separation is based on a simple yet effective rule: the feature points of characters are more compact than those of graphics. In the parcel extraction process, the proposed algorithm traces the branches between feature points to extract polygon structure from line drawings. Our character recognition method is based on the matching of extracted strokes using a neural network. The techniques of text/graphics separation and character recognition are robust to the rotation and writing style of characters. Another advantage of our separation algorithm is that it can successfully extract a character connected to a graphical line. Experimental results have shown that the proposed system is effective for the data capture of geographic information systems. Liang-Hua Chen, Hong-Yuan Mark Liao, Jiing-Yuh Wang, Kuo-Chin Fan |
IEEE Trans. Syst. Man Cybern. Part C | 2 |
| 1998 | Face Recognition Using a Face-Only Database: A New Approach
Hong-Yuan Mark Liao, Chin-Chuan Han, Gwo-Jong Yu, Hsiao-Rong Tyan, Meng Chang Chen, Liang-Hua Chen |
ACCV (2) | 1 |
| 1998 | Facial feature detection using geometrical face model: An efficient approach
Shi-Hong Jeng, Hong-Yuan Mark Liao, Chin-Chuan Han, Ming-Yang Chern, Yao Tsorng Liu |
Pattern Recognit. | 2 |
| 1997 | Image Registration Using a New Edge-Based Approach
Jun-Wei Hsieh, Hong-Yuan Mark Liao, Kuo-Chin Fan, Ming-Tat Ko, Yi-Ping Hung |
Comput. Vis. Image Underst. | 2 |
| 1997 | New automatic multi-level thresholding technique for segmentation of thermal images
Jung-Shiong Chang, Hong-Yuan Mark Liao, Maw-Kae Hor, Jun-Wei Hsieh, Ming-Yang Chern |
Image Vis. Comput. | 2 |
| 1997 | A new wavelet-based edge detector via constrained optimization
Jun-Wei Hsieh, Ming-Tat Ko, Hong-Yuan Mark Liao, Kuo-Chin Fan |
Image Vis. Comput. | 3 |
| 1996 | Shape from texture based on the ridge of continuous wavelet transformabstractWe propose a new shape from texture method based on the ridge of continuous wavelet transform. This method determines the orientations of a planar surface in a direct way under the perspective projection model. The variations of the image projected from a planar surface can be accurately characterized by the ridge of the continuous wavelet transform. The ridge of the 1-D signal and 2-D image are represented as a ridge curve and ridge plane, respectively. Ridges represent the energy concentration in the time-frequency plane where the energy is a local maxima. We show that the ridge of the projected image is a parabolic plane with a rotation angle equal to the tilt angle of the planar surface. The ridge is then rotated with the angle such that the slant effect appears in the X-axis and plays no role along the Y-axis. As a result, the rotated ridge plane can be regarded as the plane composed of many 1-D ridge curves. The slant angle of the 2-D image is thus obtained from the derived slant angle of the 1-D signal. A voting method and a curve fitting method are developed to obtain the slant angle of the 1-D signal. Several synthetic and real-world images have demonstrated the robustness and accuracy of our method. Chun-Shien Lu, Wen-Liang Hwang, Hong-Yuan Mark Liao, Pau-Choo Chung |
ICIP (1) | 3 |
| 1996 | An interpretation system for cadastral mapsabstractThis paper presents a map interpretation system for automatic extraction of high level information (such as parcels and their attributes) from the scanned images of Chinese cadastral maps. Our map interpretation system consists of three main components: text/graphics separation, parcel extraction, and rotated character recognition. The techniques of text/graphics separation and character recognition are robust to the rotation and writing style of character. Another advantage of our separation algorithm is that it can successfully extract character connected to graphical line. Liang-Hua Chen, Hong-Yuan Mark Liao, Jiing-Yuh Wang, Kuo-Chin Fan, Chen-Chiung Hsieh |
ICPR | 2 |
| 1996 | A fast algorithm for image registration without predetermining correspondencesabstractA novel approach for efficient image registration is proposed. The proposed method applies wavelet transforms to extract a number of feature points as the basis for registration. From the selected feature points, a subset of possible matching pairs is selected to obtain the desired registration parameters. This subset is chosen by using the orientation difference between two target images as a criterion to eliminate spurious matching pairs. In order to predetermine the orientation difference between two target images, a so-called "angle histogram" is calculated. From the angle histogram, the orientation difference can be decided. Once the orientation difference is obtained, the desired subset can be easily determined. By randomly selecting two matching pairs from this subset, a set of registration parameters can be obtained. By checking how many matching pairs are compatible with the selected parameters, the best estimation can be determined. Compared with conventional algorithms, the proposed scheme is a great improvement in terms of efficiency as well as reliability for the image registration problem. Jun-Wei Hsieh, Hong-Yuan Mark Liao, Kuo-Chin Fan, Ming-Tak Ko |
ICPR | 2 |
| 1996 | An efficient approach for facial feature detection using geometrical face modelabstractMost of the conventional approaches for facial feature detection use the template matching and correlation techniques. These kinds of approaches are very time-consuming and therefore impractical in a real-time systems. In this paper, we propose a useful geometrical face model and an efficient facial feature detection scheme. Based on the fact that human faces are constructed in the same geometrical configuration, the proposed scheme can accurately detect facial features, especially the eyes, even when the images have complex backgrounds. Experimental results demonstrate that the proposed scheme can efficiently detect human facial features and is deal for dealing with the problems caused by bad lighting condition, skew face orientation, and even facial expression. Shi-Hong Jeng, Hong-Yuan Mark Liao, Yao Tsorng Liu, Ming-Yang Chern |
ICPR | 2 |
| 1996 | A robust algorithm for separation of Chinese characters from line drawings
Liang-Hua Chen, Jiing-Yuh Wang, Hong-Yuan Mark Liao, Kuo-Chin Fan |
Image Vis. Comput. | 3 |
| 1996 | Fractal image coding system based on an adaptive side-coupling quadtree structure
Chwen-Jye Sze, Hong-Yuan Mark Liao, Kuo-Chin Fan, Ming-Yang Chern, Eric Chen-Kuo Tsao |
Image Vis. Comput. | 2 |
| 1995 | Separation of Chinese characters from graphicsabstractIn this paper, we propose a robust algorithm to separate Chinese characters from line drawings. This approach is based on the clustering of all feature points in the images. Using our algorithm, all Chinese characters can be completely separated from graphics without regard to the size, orientation and location of Chinese characters even if the characters touching or overlapping line problem occurs. Experiments show that our algorithm can be applied to both geographical information systems and forms processing. Jiing-Yuh Wang, Liang-Hua Chen, Kuo-Chin Fan, Hong-Yuan Mark Liao |
ICDAR | 4 |
| 1995 | Wavelet-Based Shape from Shading
Jun-Wei Hsieh, Hong-Yuan Mark Liao, Ming-Tat Ko, Kuo-Chin Fan |
CVGIP Graph. Model. Image Process. | 2 |
| 1995 | Extraction of characters from form documents by feature point clustering
Kuo-Chin Fan, Jeng-Ming Lu, Liang-Sheng Wang, Hong-Yuan Mark Liao |
Pattern Recognit. Lett. | 4 |
| 1994 | Wavelet-Based Shape from ShadingabstractThis paper proposes a wavelet-based approach to solving the shape from shading (SFS) problem. The proposed method takes advantage of the nature of wavelet theory, which can be applied to efficiently and accurately represent "things", to develop a faster algorithm for reconstructing better surfaces. In order to improve the robustness of the algorithm, two new constraints are introduced into the objective function to strengthen the relation between an estimated surface and its counterpart in the original image. Thus, solving the SFS problem becomes a constrained optimization process. In the first stage of the process, the set of function variables to be solved is represented by a wavelet format. Due to this format, the set of differential operators of different orders which is involved in the whole process can be approximated with the connection coefficients of Daubechies bases. In each iteration of the optimization process an appropriate step size which will result in maximum decrease of the objective function is determined. After finding correct iterative schemes, the solution of the SFS problem will finally be decided. Compared with conventional algorithms, the proposed scheme makes great improvements on the accuracy as well as the convergence speed of the SFS problem.> Jun-Wei Hsieh, Hong-Yuan Mark Liao, Ming-Tat Ko, Kuo-Chin Fan |
ICIP (2) | 2 |
| 1994 | Recovery of superquadric primitive from stereo images
Liang-Hua Chen, Wei-Chung Lin, Hong-Yuan Mark Liao |
Image Vis. Comput. | 3 |
| 1993 | Stroke-based handwritten Chinese character recognition using neural networks
Hong-Yuan Mark Liao, Jun-Shon Huang, Shih-Ta Huang |
Pattern Recognit. Lett. | 1 |