Wen-Hung Liao

dblp:66/5753 · DBLP profile ↗
← Back
39ranked-venue papers
27as first author
10since 2021 · last 2026
0000-0002-1181-3459ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 31 · 23 first-author · 10 since 2021Artificial intelligence and machine learning · 13 · 9 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 6 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 3 first-author
YearPublicationVenuePosition
2026 Decoding-Time Fusion of OCR and Large Language Models for Traditional Chinese Historical Document Recognition
Zih-Ci Lin, Wen-Hung Liao
ICPR (8)2
2025 Enhancing Fisheye Lens Object Detection with Generative Data Augmentation
abstract
Overhead fisheye cameras offer broad spatial coverage, making them suitable for surveillance in public spaces such as libraries. However, their severe image distortion and the scarcity of publicly available, privacy-compliant datasets hinder effective object detection. This study addresses these challenges through a dual strategy: augmenting data using text-to-image generative models and correcting fisheye distortion via calibrated intrinsic camera parameters. This approach enables robust training on enriched datasets while mitigating geometric artifacts. Experimental results show substantial performance gains over the YOLOv8 baseline, with [email protected] improving from 0.246 to 0.688 and [email protected]:0.95 from 0.122 to 0.518. Detection of small objects—such as beverages—improved markedly, with [email protected] rising from 0.507 to 0.795. Furthermore, combining synthetic and real data in training not only enhances generalization but also improves model robustness under challenging visual conditions.
Wen-Hung Liao, Pin-Chieh Cheng
AVSS1
2025 Exploring the Performance Recovery of Remedial Learning Within the Federated Learning Framework
abstract
This research examines the role of remedial learning in federated image classification, focusing on performance recovery when client data is subject to poisoning attack. Remedial learning targets model weaknesses through corrective strategies, aiming to enhance accuracy and stability across heterogeneous data sources. Experimental results show that applying remedial learning collaboratively across all clients in the federated framework yields significantly better performance recovery than isolating the contaminated client. Further evaluation reveals that excluding the contaminated client and retraining the model still surpasses remedial learning applied solely on that client. As training rounds increase, performance converges closely to that of standard federated training. These findings highlight the effectiveness of distributed remedial learning in mitigating the impact of data contamination and improving federated model robustness.
Wen-Hung Liao, Che-Wei Huang
AVSS1
2024 Unveiling the Potential of SSL-Generated Audio Embeddings for Cross-Lingual Speaker Recognition
abstract
This research explores the effectiveness of SSL-based audio embeddings in cross-lingual speaker recognition. We collected speech data from 120 participants, named MET-120 in which each participant recorded in three languages (Mandarin, English, and Taiwanese). We then employ self-supervised learning (SSL) pre-trained models, including Wav2vec 2.0 and BEATs, to extract audio features that can characterize the speaker. A simple residual neural network (ResNet) is trained to perform cross-lingual speaker recognition tasks. Experimental results show that the fine-tuned Wav2vec 2.0 model achieves over 90% average performance on MET-120, obtaining the best overall results. Without fine-tuning, BEATs achieves 80% average performance on MET-120, suggesting that it might serve as a soft biometric in cross-lingual scenarios. The influence of native or proficient languages on recognition results is observed. Furthermore, we evaluate the efficacy of acoustic data augmentation schemes such as SpecAugment and ShuffleAugment. Experimental results demonstrate that ShuffleAugment, when used alongside dimensionality-reduction techniques like PCA, significantly improves performance in both same-language and cross-lingual tests.
Wen-Hung Liao, Yi-Chieh Wu
ISM1
2024 Investigation of Feature Distribution and Network Weight Updates in the Machine Unlearning Process
abstract
Machine unlearning refers to the process of expunging previously learned information and data from a machine learning model to achieve the objective of privacy protection. In this research, we explore two prevalent methods for unlearning, namely, label reassignment and model manipulation. Using the CIFAR-100 classification problem with ResNet-50 architecture as an example, we examine the efficacy of these two mechanisms and their variants in the corresponding unlearning task. We further investigate the changes in feature distribution and the extent of weight updates across various network layers throughout the unlearning process. Experimental results indicate that the degree of network variation is proportional to the number of removed classes. When employing the label reassignment method, the variation is concentrated in the final stage and fully connected layers. On the other hand, using the weight resetting strategy affects more network layers, with the impact gradually decreasing from the later layers to the middle and earlier layers. Overall, when the categories to be forgotten are less than 10%, no significant impact on feature extraction is observed.
Wen-Hung Liao, Yang-Jing Lin
ISM1
2024 Adversarially Robust Deepfake Detection via Adversarial Feature Similarity Learning
Sarwar Khan, Jun-Cheng Chen, Wen-Hung Liao, Chu-Song Chen
MMM (3)3
2023 The Impact of Parroting Mode on Cross-Lingual Speaker Recognition
abstract
People use multiple languages in their daily lives across regions worldwide, which motivated us to investigate cross-lingual speaker recognition. In this work, we propose to collect recordings of Mandarin and Spanish, namely the Mandarin-Spanish-Speech Dataset (MSSD-40), to analyze the performance of various audio embeddings for cross-lingual speaker recognition tasks. All participants are fluent in Mandarin, but none of the participants have prior knowledge of the Spanish language. As such, they have been advised to adopt a parroting mode of Spanish speech production, wherein they simply repeat the sounds emanating from the loudspeaker. Using this approach, variations resulting from individual differences in language fluency can be reduced, enabling us to focus on the anatomical aspects of the speech production mechanism.Embeddings extracted from models pre-trained with a large number of audio segments have become effective solutions for coping with audio analysis tasks using small datasets. Preliminary experimental results using two collected multi-lingual datasets indicate that both embedding methods and the language employed will affect the robustness of the speaker recognition task. Precisely, stable performance is observed when familiar languages are used. BEATs embedding generates the best outcome in all languages when no fine-tuning is exercised.
Wen-Hung Liao, Yen-Chun Ou, Yi-Chieh Wu
ISM1
2022 On the Robustness of Cross-lingual Speaker Recognition using Transformer-based Approaches
abstract
Most speaker recognition systems presume that the language for enrollment and testing is the same. Cross-lingual speaker recognition is rarely investigated. This study collected trilingual (including Mandarin, English, and Taiwanese) cross-language recordings named MET-40. A total of 40 participants (20 male, 20 female) contribute to the dataset which contains 740 minutes of audio. Spoken texts are mainly taken from elementary school textbooks, and some English texts use TIMIT.We employ ResNet, vision transformer (ViT), and convo-lutional vision transformer (CvT) in combination with three acoustic features, namely, spectrogram, Mel spectrogram, and Mel frequency cepstral coefficient for single, mixed and cross-language speaker recognition tasks. In the mixed-language setting, the language to be tested is included in the training set, while in the cross-language scenario the language to be tested is not used for training. Experimental results show that the highest accuracy is 97.16% for single language models. Mixture of two languages improves the performance to 99.17%. In cross-language situations, the accuracy drops significantly to 79.64%, as the spoken language is not present in the training data. When two languages are employed for training, the accuracy rose to 90.92%. In general, CvT-based models demonstrate the best stability in all cases.The robustness of the model is critical to security in practical applications. Therefore, we analyze how adversarial attacks impact different speaker identification models. The results show that although CvT-based model exhibits excellent performance, it is easily affected by the perturbation caused by the adversarial attack. The effect is less pronounced when more languages are used for training, with an average increase of 5.11% in accuracy. Finally, extra caution needs to be taken when MFCC is chosen to be the acoustic feature, as attacks can still take place without training data, and the recognition rate is reduced by 31.57% using FGSM cross-language attack.
Wen-Hung Liao, Wei-Yu Chen, Yi-Chieh Wu
ICPR1
2021 Far-Sighted BiSeNet V2 for Real-time Semantic Segmentation
abstract
Real-time semantic segmentation is one of the most investigated areas in the field of computer vision. In this paper, we focus on improving the performance of BiSeNet V2 by modifying its architecture. BiSeNet V2 is a two-branch segmentation model designed to extract semantic information from high-level feature maps and detailed information from low-level feature maps. The proposed enhancement remains lightweight and real-time with two main modifications: enlarging the contextual information and breaking the constraint caused by the fixed size of convolutional kernels. Specifically, additional modules known as dilated strip pooling (DSP) and dilated mixed pooling (DMP) are appended to the original BiSeNet V2 model to form the far-sighted BiSeNet V2. The proposed dilated strip pooling block and dilated mixed pooling module are adapted from modules proposed in SPNet, with extra branches composed of dilated convolutions to provide larger receptive fields. The proposed far-sighted BiSeNet V2 improves the accuracy to 76.0% from 73.4% with an FPS of 94 on Nvidia 1080Ti. Moreover, the proposed dilated mixed pooling block achieves the same performance as that of the model with two mixed pooling modules using only 2/3 of the number of parameters.
Te-Wei Chen, Yen-Ting Huang, Wen-Hung Liao
AVSS3
2021 Intelligent Voice Assistant to Facilitate Elementary School English Learning: A Case Study Using Amazon Echo Dot
abstract
This research focuses on exploring the English learning process for elementary school students. We employ Amazon Echo Dot, one of the most popular intelligent voice assistants nowadays, as a tool to facilitate language learning. We developed an Amazon Skill that incorporates the content from English textbooks for the participants to interact with using voice input. The operation logs from Echo Dot faithfully reveal student's usage patterns and preferences. A semester-long experiment has been conducted with the assistance of the Affiliated Experimental Elementary School (AEES) of National Chengchi University. After data collection has been completed, we utilize acoustic and transcript evaluation metrics to examine the voice recordings and user logs. Our initial analysis focuses on the active users, i.e., participants who have continued to engage in conversations with the voice assistant. Several questions regarding user behavior are prompted and responded based on the collected and processed data. Analyzing the content of the conversation will help disclose more detailed information regarding the learning process. The progress of individual students can also be monitored to determine if further assistance is needed.
Yi-Chieh Wu, Wen-Hung Liao
ISM2
2020 Defense Mechanism Against Adversarial Attacks Using Density-based Representation of Images
abstract
Adversarial examples are slightly modified inputs devised to cause erroneous inference of deep learning models. Protection against the intervention of adversarial examples is a fundamental issue that needs to be addressed before the wide adoption of deep-learning based intelligent systems. In this research, we utilize the method known as input recharacteri-zation to effectively eliminate the perturbations found in the adversarial examples. By converting images from the intensity domain into density-based representation using half toning operation, performance of the classifier can be properly maintained. With adversarial attacks generated using FGSM, I-FGSM, and PGD, the top-5 accuracy of the hybrid model can still achieve 80.97%, 78.77%, 81.56%, respectively. Although the accuracy has been slightly affected, the influence of adversarial examples is significantly discounted. The average improvement over existing input transform defense mechanisms is approximately 10%.
Yen-Ting Huang, Wen-Hung Liao, Chen-Wei Huang
ICPR2
2020 Investigation of DNN Model Robustness Using Heterogeneous Datasets
abstract
Deep learning frameworks have been successfully applied to tackle many challenging tasks in pattern recognition and computer vision thanks to its ability to automatically extract representative features from the training data. Such type of data-driven approach, however, is subject to the criticism of too much dependency on the training set. In this research, we attempt to investigate the validity of this statement: `deep learning is only as good as its data' by evaluating the performance of deep learning models using heterogeneous data sets, in which distinct representations of the same source data are employed for training/testing. We have examined three cases: low-resolution image, severely compressed input and halftone image in this work. Our preliminary results indicate that such dependency indeed exists. Classifier performance drops considerably when the model is tested with modified or transformed input. The best outcomes are obtained when the model is trained with hybrid input.
Wen-Hung Liao, Yen-Ting Huang
ICPR1
2020 Toward Text-independent Cross-lingual Speaker Recognition Using English-Mandarin-Taiwanese Dataset
abstract
Over 40% of the world's population is bilingual. Existing speaker identification/verification systems, however, assume the same language type for both enrollment and recognition stages. In this work, we investigate the feasibility of employing multilingual speech for biometric applications. We establish a dataset containing audio recorded in English, Mandarin and Taiwanese. Three acoustic features, namely, i-vector, d-vector and x-vector have been evaluated for both speaker verification (SV) and identification (SI) tasks. Preliminary experimental results indicate that x-vector achieves the best overall performance. Additionally, the model trained with hybrid data demonstrates the highest accuracy, at the cost of extra data collection efforts. In SI tasks, we obtained over 91 % cross-lingual accuracy in all models using 3-second audio. In SV tasks, the EER among cross-lingual test is at most 6.52 %, which is observed on the model trained by English corpus. The outcome suggests the feasibility of adopting cross-lingual speech in building text-independent speaker recognition systems.
Yi-Chieh Wu, Wen-Hung Liao
ICPR2
2020 GODoc: high-throughput protein function prediction using novel k-nearest-neighbor and voting algorithms
abstract
BACKGROUND: Biological data has grown explosively with the advance of next-generation sequencing. However, annotating protein function with wet lab experiments is time-consuming. Fortunately, computational function prediction can help wet labs formulate biological hypotheses and prioritize experiments. Gene Ontology (GO) is a framework for unifying the representation of protein function in a hierarchical tree composed of GO terms. RESULTS: We propose GODoc, a general protein GO prediction framework based on sequence information which combines feature engineering, feature reduction, and a novel ​k​-nearest-neighbor algorithm to resolve the multiple GO prediction problem. Comprehensive evaluation on CAFA2 shows that GODoc performs better than two baseline models. In the CAFA3 competition (68 teams), GODoc ranks 10th in Cellular Component Ontology. Regarding the species-specific task, the proposed method ranks 10th and 8th in the eukaryotic Cellular Component Ontology and the prokaryotic Molecular Function Ontology, respectively. In the term-centric task, GODoc performs third and is tied for first for the biofilm formation of Pseudomonas aeruginosa and the long-term memory of Drosophila melanogaster, respectively. CONCLUSIONS: We have developed a novel and effective strategy to incorporate a training procedure into the k-nearest neighbor algorithm (instance-based learning) which is capable of solving the Gene Ontology multiple-label prediction problem, which is especially notable given the thousands of Gene Ontology terms.
Yi-Wei Liu, Tz-Wei Hsu, Che-Yu Chang, Wen-Hung Liao, Jia-Ming Chang
BMC Bioinform.4
2019 Enhancing object detection in the dark using U-Net based restoration module
abstract
In recent years, we have witnessed the widespread application of deep-learning techniques to various surveillance tasks, including human tracking and counting, abnormal behavior detection, and video segmentation. In most cases, the input images/videos are assumed to possess adequate visual quality to guarantee satisfactory performance. However, accuracy may be adversely affected when the input data are degraded by factors such as excessive noise or poor lighting conditions. In the paper, we develop a deep neural network based on the U-Net architecture that acts as a pre-processing module to restore images/videos with nonuniform light sources to ensure the accuracy of the subsequent object detection process. Experimental results on the VisDrone20 19 dataset [1] demonstrate the effectiveness of the proposed method, achieving a remarkable 5% increase in average recall. We expect the framework to be universally applicable to situations that call for the enhancement of raw input data.
Yen-Ting Huang, Yan-Tsung Peng, Wen-Hung Liao
AVSS3
2019 Analyzing Social Network Data Using Deep Neural Networks: A Case Study Using Twitter Posts
abstract
The limitation on the total number of characters compels Twitter users to compose their messages more succinctly, suggesting a stronger association between text and image. In this paper, we employ computer vision and word embedding techniques to analyze the relationship between image content and text messages and explore the rich information entangled. Specifically, we collected all tweets which include keywords related to Taiwan during 2017. After data cleaning, we apply machine learning techniques to classify tweets into travel and non-travel types. This is achieved by employing deep neural networks to process and integrate text and image information. Within each class, we use hierarchical clustering to further partition the data into different clusters and investigate their characteristics. Through this research, we expect to identify the relationship between text and images in a tweet and gain more understanding of the properties of tweets on social networking platforms. The proposed framework and corresponding analytical results should also prove useful for qualitative research.
Wen-Hung Liao, Yen-Ting Huang, Tsu-Hsuan Yang, Yi-Chieh Wu
ISM1
2018 Evaluation of Student's 3D Modeling Capability Based on Model Completeness and Usage Pattern in K-12 Classrooms
abstract
As more schools incorporate 3D printing into their curriculum to stimulate the creativity of K-12 students with a learning-by-doing approach, it becomes crucial to understand how users work with 3D modeling tools. In this paper, we aim to develop model and usage-pattern-related features to quantize students' performance on 3D modeling operation. The dataset is gathered from the Affiliated Experimental Elementary School (AEES) of National Chengchi University. Participants' operation log and finished work for specific 3D modeling software are recorded and analyzed. In all our lesson plans, students are required to create structurally stable and printable 3D models. Three modeling software with different levels of difficulty have been introduced and tested. The collected data include screen recording, software operation log, experts evaluation, and interviews with students, which are employed for subsequent qualitative evaluation as well as quantitative analysis. With our proposed approach, we are able to identify the key factors affecting students' learning experience and performance in terms of model completeness and usage pattern. Through these indicators, instructors can understand student's learning status of 3D modeling software more comprehensively.
Yi-Chieh Wu, Wen-Hung Liao, Chen-Yu Liu, Tsai-Yen Li, Ming-Te Chi
ICALT2
2017 Classification of Reading Patterns Based on Gaze Information
abstract
Reading is one of the main paths to acquire knowledge, either done traditionally on paper media or practiced on electronic devices. Efficiency varies when different reading patterns are involved. It is the objective of this research to classify reading patterns from fixation data using machine learning techniques in an attempt to understand and evaluate the reading and learning process. In our experiment, a low-cost eye tracker is employed to record the eye movements during the reading process. A dispersion-based algorithm is implemented to identify fixation from the recorded data. Features pertaining to fixation including duration, path length, landing position and fixation direction are extracted for classification purposes. Five categories of reading pattern have been defined and investigated in this study, namely, speed reading, slow reading, in-depth reading, skim-and-skip, and keyword spotting. We have recruited thirty subjects to participate in our experiment. The participants are instructed to read different articles using specific styles designated by the experimenter in order to assign label to the collected data. Feature selection is achieved by analyzing the predictive results of cross-validation from the training data obtained from all subjects. The average classification accuracies in five random tests are 78.24%, 74.19%, 93.75%, 87.96%, and 96.20% respectively. Further improvements are accomplished by introducing an additional undecided class to address ambiguous reading patterns.
Wen-Hung Liao, Chin-Wen Chang, Yi-Chieh Wu
ISM1
2016 Feature descriptor based on local intensity order relations of pixel group
abstract
Robust image features are essential in building effective image recognition engines. These features can be constructed according to various principles, such the distribution of local gradients (Histogram of Oriented Gradients, HOG), the relationship between two pixels (Local Binary Descriptors, LBD), or local intensity order statistics (Local Intensity Order Patterns, LIOP). Because the feature dimension grows quickly as one considers the ordering relations of a group of N (N>2) pixels, few researchers have exploited local order statistics among a pixel set to define an image feature. In this paper, we propose a novel approach to construct a feature descriptor using local intensity order relations (LIOR) in a pixel group. In contrast to LIOP where the feature dimension increases drastically with the number of elements in a set, the size of LIOR is manageable. Moreover, LIOR ensures the stability of ordering by encoding the intensity differences as weights. Two different strategies for assigning the weights have been devised and tested. Experimental results indicate that the proposed methods yield better or comparable performance for different types of image degradation when compared to the original LIOP. Additionally, the storage requirement is significantly lower when the number of pixels in a group increases.
Wen-Hung Liao, Chia-Chen Wu, Ming-Ching Lin
ICPR1
2016 Evaluation of Interactive Data Visualization Tools Based on Gaze and Mouse Tracking
abstract
As more and more interactive data visualization tools emerge, designers need a coincident evaluation method to give timely users' feedback. In this research, we propose a systematic approach to gauge the usability of such tools. Firstly, quantitative data including gaze and mouse movement are collected for statistical analysis. Secondly, user operation is encoded as a sequence for comparison. We describe the mechanism for data analysis on gaze and mouse tracking. Finally, we present understandability, discoverability, usage frequency, and efficiency as indicators to evaluate interactive data visualization tools.
Chiu-Fang Peng, Wen-Hung Liao
ISM2
2014 Feature Description Using Center-Symmetric Extended Local Ternary Patterns
abstract
Effective recognition of objects calls for the appropriate selection of feature descriptor. In this paper, we generalize the "extended local ternary patterns" (ELTP) to form a novel and compact set of features named center-symmetric extended local ternary patterns (CS-ELTP). The newly defined CS-ELTP follows a simplified encoding procedure and has a lower dimension for a fixed neighborhood region. It achieves good balances among feature dimension, recognition rate and noise resistance according to our comparative experimental analysis. In addition, we combine binary and ternary patterns to create a class of hybrid descriptor that possesses the characteristics of both types of descriptor. Experimental results indicate that the hybrid descriptor can improve the performance in noisy conditions while maintaining a reasonable feature dimension.
Wen-Hung Liao, Chia-Yu Liu, Ming-Ching Lin
ISM1
2012 Commensurate dimensionality reduction for extended local ternary patterns
Wen-Hung Liao
ICPR1
2012 Incorporating Fuzziness in Extended Local Ternary Patterns
abstract
Local binary/ternary patterns are widely employed to describe the structure of an image region. However, local patterns are very sensitive to noise due to the thresholding process. In this paper, we propose two different approaches to incorporate fuzziness in extended local ternary patterns (ELTP) to enhance the robustness of this class of operator to interferences. The first approach replaces the ternary mapping mechanism with fuzzy member functions to arrive at a fuzzy ELTP representation. The second approach modifies the clustering operation in formulating ELTP to a fuzzy C-means procedure to construct soft histograms in the final feature representation, denoted as FCM-ELTP. Both fuzzy descriptors have proven to exhibit better resistance to noise in the experiments designed to compare the performance of ELTP and the newly proposed fuzzy ELTP and FCM-ELTP.
Wen-Hung Liao
ISM1
2011 Analysis and Interpretation of e-Reader User Logs: A Case Study of High School Students' User Behaviors
abstract
This study explores the daily life user experiences of an experimental e-book reading device among high-school students, aiming to understand how well the digital natives accept the use of e-book reading devices and the potential utilities of such devices for them, either for leisure purposes or as an assistive educational tool. Toward this goal, we have custom-designed the e-reader user interface as well as the e-book content to suit the needs of this particular user group. The unique opportunity of having access to the hardware device, software design and potential users creates an ideal experimental platform for us to unbiasedly investigate the role of this new technology through a long-term user behavior collection and analysis process. We anticipate that the new reading behaviors of the digital natives will provide clues for further improvements in the design and development of digital reader devices.
Wen-Hung Liao, Chien-Pao Chueh
ICALT1
2010 Region Description Using Extended Local Ternary Patterns
abstract
The local binary pattern (LBP) operator is a computationally efficient local texture descriptor and has found many useful applications. However, its sensitivity to noise and the high dimensionality of histogram associated with a mediocre size neighborhood have raised some concerns. In this paper, we attempt to improve the original LBP by proposing a novel extension named extended local ternary pattern (ELTP). We will investigate the characteristics of ELTP in terms of noise sensitivity, discriminability and computational efficiency. Preliminary experimental results have shown better efficacy of ELTP over the original LBP.
Wen-Hung Liao
ICPR1
2010 Texture Classification Using Uniform Extended Local Ternary Patterns
abstract
We present an extension to the well-known local binary pattern (LBP) feature descriptor. The newly defined descriptor known as extended local ternary pattern (ELTP) exhibits better noise resistivity than the original LBP, while maintaining computational simplicity. We further investigate the presence of uniform patterns in ELTP. With a slight modification in the definition of uniformity, it is found experimentally that uniform ELTP account for 80% of all patterns in texture images. Comparative performance analysis indicates that the proposed uniform ELTP is more effective than uniform LBP for texture classification tasks.
Wen-Hung Liao, Ting-Jung Young
ISM1
2009 Automatic Generation of Caricatures with Multiple Expressions Using Transformative Approach
Wen-Hung Liao, Chien-An Lai
ArtsIT1
2009 A Framework for Attention-Based Personal Photo Manager
abstract
In this paper, we propose a novel framework for the design and implementation of an attention-based personal digital photo browsing platform. The key concept that separates the proposed system from existing ones is the incorporation of user interaction patterns to infer the level of interest in a particular photo. Specifically, we use Web cameras to record and analyze the viewing behavior of the user and attempt to correlate the interest of the viewer to the effective viewing time. We also devise an updating scheme to efficiently renew the timing parameter. To build a comprehensive photo browser, external EXIF data and face detection results are utilized to coarsely classify the digital images. Moreover, measures of image quality, including sharpness and contrast, are calculated to rank the search results. Finally, a ranking-based algorithm is utilized to integrate the clues acquired from different modules.
Wen-Hung Liao
SMC1
2009 Classification of Non-Speech Human Sounds: Feature Selection and Snoring Sound Analysis
abstract
Human sounds can be roughly divided into two categories: speech and non-speech. Traditional audio scene analysis research puts more emphasis on the classification of audio signals into human speech, music, and environmental sounds. We take a different perspective in this paper. We are mainly interested in the analysis of non-speech human sounds, including laugh, screaming, sneeze, and snore. Toward this goal, we investigate many commonly used acoustic features and select useful ones for classification using multivariate adaptive regression splines (MARS) and support vector machine (SVM). To evaluate the robustness of the selected features, we also perform extensive simulations to observe the effect of noise on the accuracy of the classification. Finally, for the class of snoring sounds, we propose a robust approach to further categorize them into simple snores and snores of subjects with obstructive sleep apnea (OSA).
Wen-Hung Liao, Yu-Kai Lin
SMC1
2008 Video-based activity and movement pattern analysis in overnight sleep studies
abstract
We present a non-contact monitoring system to measure the quality of sleep using near-infrared video in this paper. We envision a smart home environment in which a processing module can be installed in the bedroom to record and monitor sleep in a noninvasive manner. We describe the procedure adopted to infer motion information and discuss the method for estimating wake/sleep status from the acquired video. Performance of the proposed system was evaluated through comparison with simultaneous recordings of actigraph and polysomnography (PSG) data.
Wen-Hung Liao, Chien-Ming Yang
ICPR1
2007 Accelerating Non-photorealistic Video Effects Using Compressed Domain Processing
Wen-Hung Liao
MMM (2)1
2006 Generation of 3D Caricature by Fusing Caricature Images
abstract
Caricatures are exaggerated, cartoon-like portraits that try to capture the essence of the subject with a bit of humor or sarcasm. In this paper, we extend the automatic caricature generation algorithm developed previously in two important aspects: 1) we improve the facial feature detection procedure by incorporating active appearance models, and 2) we extend the original 2D model to stereo by fusing caricature images obtained from different views. The result is a simple yet effective procedure for creating stylized 3D models with a high degree of flexibility in selecting artistic flavors.
Yu-Lun Chen, Wen-Hung Liao, Pei-Ying Chiang
SMC2
2004 Embedding information within dynamic visual patterns
abstract
There exist many scenarios in which the automatic distinction between human and machines is desirable. This paper describes a texture-image based approach to encode text information in such a way that machine vision algorithms will experience difficulties while humans can extract the embedded text effortlessly. We exploit both static and dynamic visual patterns to diversify the image content and discourage automatic processing. Examples of the resulting patterns are presented and their properties are discussed.
Wen-Hung Liao, Chi-Chih Chang
ICME1
2003 Homomorphic processing techniques for near-infrared images
abstract
The images of objects in total darkness can be captured using a relatively low cost camcorder with the NightShot/spl reg/ function. However, the resulting images exhibit non-uniformity due to irregular illumination. We investigate the characteristics of and propose an image formation model for near-infrared images. A homomorphic processing technique built upon the image model is then developed to reduce the artifact of the captured images. We discuss how the parameters of the homomorphic filter should be selected and demonstrate the effectiveness of the proposed image processing technique with experimental results.
Wen-Hung Liao, Dai-Yun Li
ICASSP (3)1
1998 Nonrigid Motion Analysis: Articulated and Elastic Motion
Jake K. Aggarwal, Quin Cai, Wen-Hung Liao, Bikash Sabata
Comput. Vis. Image Underst.3
1997 The reconstruction of dynamic 3D structure of biological objects using stereo microscope images
Wen-Hung Liao, Shanti J. Aggarwal, Jake K. Aggarwal
Mach. Vis. Appl.1
1996 Curve and surface interpolation using rational radial basis functions
abstract
This paper addresses the problem of reconstructing two-dimensional curves and three-dimensional surfaces from scattered, sparse measurements. We extend the rational Gaussian (RaG) functions introduced by Goshtasby (1993) to general rational radial basis functions and develop a method to compute the smoothness parameters for the shape model by considering the adjacency relation of the control points. Experimental results demonstrate substantial improvements over the original RaG-based method when the input data is sparse and the distribution of the control points is highly nonuniform.
Wen-Hung Liao, Jake K. Aggarwal
ICPR1
1995 Analysis of Left Ventricular Motion
Wen-Hung Liao, Shanti J. Aggarwal, Jake K. Aggarwal
ACCV1
1994 Reconstruction of Cynamic 3-D Structures of Biological Objects Using Stereo Microscopy
abstract
The authors address the analysis of three dimensional shape and shape change in nonrigid biological objects imaged via a stereo light microscope (SLM). Most existing stereo or motion analysis techniques cannot be applied to microscopic biological images because they usually lack salient features. The authors propose an integrated approach for the reconstruction of 3D structures and motion analysis for scenes where only a few informative features are available. The key components of this framework are: (1) image registration, (2) region-of-interest extraction, and (3) stereo and motion analysis using a cooperative spatial and temporal matching process. The authors describe these three stages of processing and illustrate the efficacy of the proposed approach using real images of a live frog's ventricle. The reconstructed dynamic 3-D structures of the ventricle are demonstrated in the authors' experimental results.>
Wen-Hung Liao, Shanti J. Aggarwal, Jake K. Aggarwal
ICIP (3)1