Yi-Chieh Wu

dblp:122/7301 · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 From Temporal to Spatial: A Transformer-GAN Approach for Fall Prediction
abstract
The study aims to construct a fall detection system. We propose a Transformer-based Generative Adversarial Network (GAN) trained on human keypoint skeleton features. We arrange temporal information spatially, allowing the attention layers to learn movement details. Moreover, the objective of the GAN is to generate corresponding future motion sequences based on the current data. Consequently, the model can predict the next movement and determine whether an individual in the scene has fallen. Finally, the fall detection process employs distance metrics, including Chebyshev distance and Dynamic Time Warping (DTW), with a threshold determined through Bayesian decision theory. Experimental results demonstrate the model’s robust performance, achieving an F1-score of approximately 0.93 when applied to the evaluation set using the Chebyshev distance threshold. The framework effectively distinguishes fall and non-fall instances, showing particular strength in handling challenging scenarios like transitional movements.
Yu-Jung Hsu, Yi-Chieh Wu
AVSS2
2024 Unveiling the Potential of SSL-Generated Audio Embeddings for Cross-Lingual Speaker Recognition
abstract
This research explores the effectiveness of SSL-based audio embeddings in cross-lingual speaker recognition. We collected speech data from 120 participants, named MET-120 in which each participant recorded in three languages (Mandarin, English, and Taiwanese). We then employ self-supervised learning (SSL) pre-trained models, including Wav2vec 2.0 and BEATs, to extract audio features that can characterize the speaker. A simple residual neural network (ResNet) is trained to perform cross-lingual speaker recognition tasks. Experimental results show that the fine-tuned Wav2vec 2.0 model achieves over 90% average performance on MET-120, obtaining the best overall results. Without fine-tuning, BEATs achieves 80% average performance on MET-120, suggesting that it might serve as a soft biometric in cross-lingual scenarios. The influence of native or proficient languages on recognition results is observed. Furthermore, we evaluate the efficacy of acoustic data augmentation schemes such as SpecAugment and ShuffleAugment. Experimental results demonstrate that ShuffleAugment, when used alongside dimensionality-reduction techniques like PCA, significantly improves performance in both same-language and cross-lingual tests.
Wen-Hung Liao, Yi-Chieh Wu
ISM3
2024 Generating and Evaluating Cursive Chinese Calligraphy by Semi-Classifying Style: A Case Study Using a Diffusion Model
abstract
Generative AI offers a promising approach to overcoming the challenges of optical character recognition (OCR) for traditional Chinese cursive calligraphy. By generating text-image data, we can address data insufficiency and imbalance, enhancing the effectiveness of deep learning models in capturing the variability of cursive styles. In this study, we developed a text-to-text-content-image generation model utilizing the Cursive Chinese Calligraphy Dataset and the WordStylist framework. By categorizing images according to style characteristics, our method reduces the learning curve and improves the quality of generated artifacts while tackling common issues in handwritten datasets, such as the absence of author annotations and mixed styles without proper labeling.We trained three models with varying style categories: four-class, eight-class, and sixteen-class, and introduced a semi-annotating method for style categorization. Experimental results indicate that fewer style categories lead to greater variation in cursive characters within each class. The sixteen-class model exhibited more consistent styles, resulting in faster convergence and more stable generation, thereby offering users a broader range of style choices.To evaluate the quality of the generated content, we employed OCR-Embedding with Maximum Mean Discrepancy(MMD), derived from CMMD, as well as expert evaluations. The results from OCR-Embedding MMD were generally consistent with those of expert assessments. According to the experts, the four-class model performed best in zero-shot sample generation. We infer that zero-shot characters benefit from features of other characters within the same style class, as these characters have more training data available, contributing to improved generation of zero-shot characters.
Yi-Chieh Wu, Yu-Jung Hsu
ISM1
2023 The Impact of Parroting Mode on Cross-Lingual Speaker Recognition
abstract
People use multiple languages in their daily lives across regions worldwide, which motivated us to investigate cross-lingual speaker recognition. In this work, we propose to collect recordings of Mandarin and Spanish, namely the Mandarin-Spanish-Speech Dataset (MSSD-40), to analyze the performance of various audio embeddings for cross-lingual speaker recognition tasks. All participants are fluent in Mandarin, but none of the participants have prior knowledge of the Spanish language. As such, they have been advised to adopt a parroting mode of Spanish speech production, wherein they simply repeat the sounds emanating from the loudspeaker. Using this approach, variations resulting from individual differences in language fluency can be reduced, enabling us to focus on the anatomical aspects of the speech production mechanism.Embeddings extracted from models pre-trained with a large number of audio segments have become effective solutions for coping with audio analysis tasks using small datasets. Preliminary experimental results using two collected multi-lingual datasets indicate that both embedding methods and the language employed will affect the robustness of the speaker recognition task. Precisely, stable performance is observed when familiar languages are used. BEATs embedding generates the best outcome in all languages when no fine-tuning is exercised.
Wen-Hung Liao, Yen-Chun Ou, Yi-Chieh Wu
ISM4
2022 On the Robustness of Cross-lingual Speaker Recognition using Transformer-based Approaches
abstract
Most speaker recognition systems presume that the language for enrollment and testing is the same. Cross-lingual speaker recognition is rarely investigated. This study collected trilingual (including Mandarin, English, and Taiwanese) cross-language recordings named MET-40. A total of 40 participants (20 male, 20 female) contribute to the dataset which contains 740 minutes of audio. Spoken texts are mainly taken from elementary school textbooks, and some English texts use TIMIT.We employ ResNet, vision transformer (ViT), and convo-lutional vision transformer (CvT) in combination with three acoustic features, namely, spectrogram, Mel spectrogram, and Mel frequency cepstral coefficient for single, mixed and cross-language speaker recognition tasks. In the mixed-language setting, the language to be tested is included in the training set, while in the cross-language scenario the language to be tested is not used for training. Experimental results show that the highest accuracy is 97.16% for single language models. Mixture of two languages improves the performance to 99.17%. In cross-language situations, the accuracy drops significantly to 79.64%, as the spoken language is not present in the training data. When two languages are employed for training, the accuracy rose to 90.92%. In general, CvT-based models demonstrate the best stability in all cases.The robustness of the model is critical to security in practical applications. Therefore, we analyze how adversarial attacks impact different speaker identification models. The results show that although CvT-based model exhibits excellent performance, it is easily affected by the perturbation caused by the adversarial attack. The effect is less pronounced when more languages are used for training, with an average increase of 5.11% in accuracy. Finally, extra caution needs to be taken when MFCC is chosen to be the acoustic feature, as attacks can still take place without training data, and the recognition rate is reduced by 31.57% using FGSM cross-language attack.
Wen-Hung Liao, Wei-Yu Chen, Yi-Chieh Wu
ICPR3
2022 A Polynomial-Time Algorithm for Minimizing the Deep Coalescence Cost for Level-1 Species Networks
abstract
Phylogenetic analyses commonly assume that the species history can be represented as a tree. However, in the presence of hybridization, the species history is more accurately captured as a network. Despite several advances in modeling phylogenetic networks, there is no known polynomial-time algorithm for parsimoniously reconciling gene trees with species networks while accounting for incomplete lineage sorting. To address this issue, we present a polynomial-time algorithm for the case of level-1 networks, in which no hybrid species is the direct ancestor of another hybrid species. This work enables more efficient reconciliation of gene trees with species networks, which in turn, enables more efficient reconstruction of species networks.
Matthew LeMay, Ran Libeskind-Hadas, Yi-Chieh Wu
IEEE ACM Trans. Comput. Biol. Bioinform.3
2021 Intelligent Voice Assistant to Facilitate Elementary School English Learning: A Case Study Using Amazon Echo Dot
abstract
This research focuses on exploring the English learning process for elementary school students. We employ Amazon Echo Dot, one of the most popular intelligent voice assistants nowadays, as a tool to facilitate language learning. We developed an Amazon Skill that incorporates the content from English textbooks for the participants to interact with using voice input. The operation logs from Echo Dot faithfully reveal student's usage patterns and preferences. A semester-long experiment has been conducted with the assistance of the Affiliated Experimental Elementary School (AEES) of National Chengchi University. After data collection has been completed, we utilize acoustic and transcript evaluation metrics to examine the voice recordings and user logs. Our initial analysis focuses on the active users, i.e., participants who have continued to engage in conversations with the voice assistant. Several questions regarding user behavior are prompted and responded based on the collected and processed data. Analyzing the content of the conversation will help disclose more detailed information regarding the learning process. The progress of individual students can also be monitored to determine if further assistance is needed.
Yi-Chieh Wu, Wen-Hung Liao
ISM1
2021 The Most Parsimonious Reconciliation Problem in the Presence of Incomplete Lineage Sorting and Hybridization Is NP-Hard
abstract
The maximum parsimony phylogenetic reconciliation problem seeks to explain incongruity between a gene phylogeny and a species phylogeny with respect to a set of evolutionary events. While the reconciliation problem is well-studied for species and gene trees subject to events such as duplication, transfer, loss, and deep coalescence, recent work has examined species phylogenies that incorporate hybridization and are thus represented by networks rather than trees. In this paper, we show that the problem of computing a maximum parsimony reconciliation for a gene tree and species network is NP-hard even when only considering deep coalescence. This result suggests that future work on maximum parsimony reconciliation for species networks should explore approximation algorithms and heuristics.
Matthew LeMay, Yi-Chieh Wu, Ran Libeskind-Hadas
WABI2
2021 eMPRess: a systematic cophylogeny reconciliation tool
abstract
SUMMARY: We describe eMPRess, a software program for phylogenetic tree reconciliation under the duplication-transfer-loss model that systematically addresses the problems of choosing event costs and selecting representative solutions, enabling users to make more robust inferences. AVAILABILITY AND IMPLEMENTATION: eMPRess is freely available at http://www.cs.hmc.edu/empress. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Santi Santichaivekin, Ross Mawhorter, Justin Jiang, Trenton Wesley, Yi-Chieh Wu, Ran Libeskind-Hadas
Bioinform.7
2021 Multiple Optimal Reconciliations Under the Duplication-Loss-Coalescence Model
abstract
Gene trees can differ from species trees due to a variety of biological phenomena, the most prevalent being gene duplication, horizontal gene transfer, gene loss, and coalescence. To explain topological incongruence between the two trees, researchers apply reconciliation methods, often relying on a maximum parsimony framework. However, while several studies have investigated the space of maximum parsimony reconciliations (MPRs) under the duplication-loss and duplication-transfer-loss models, the space of MPRs under the duplication-loss-coalescence (DLC) model remains poorly understood. To address this problem, we present new algorithms for computing the size of MPR space under the DLC model and sampling from this space uniformly at random. Our algorithms are efficient in practice, with runtime polynomial in the size of the species and gene tree when the number of genes that map to any given species is fixed, thus proving that the MPR problem is fixed-parameter tractable. We have applied our methods to a biological data set of 16 fungal species to provide the first key insights in the space of MPRs under the DLC model. Our results show that a plurality reconciliation, and underlying events, are likely to be representative of MPR space.
Haoxing Du, Yi Sheng Ong, Marina Knittel, Ross Mawhorter, Nuo Liu, Gianluca Gross, Reiko Tojo, Ran Libeskind-Hadas, Yi-Chieh Wu
IEEE ACM Trans. Comput. Biol. Bioinform.9
2020 Toward Text-independent Cross-lingual Speaker Recognition Using English-Mandarin-Taiwanese Dataset
abstract
Over 40% of the world's population is bilingual. Existing speaker identification/verification systems, however, assume the same language type for both enrollment and recognition stages. In this work, we investigate the feasibility of employing multilingual speech for biometric applications. We establish a dataset containing audio recorded in English, Mandarin and Taiwanese. Three acoustic features, namely, i-vector, d-vector and x-vector have been evaluated for both speaker verification (SV) and identification (SI) tasks. Preliminary experimental results indicate that x-vector achieves the best overall performance. Additionally, the model trained with hybrid data demonstrates the highest accuracy, at the cost of extra data collection efforts. In SI tasks, we obtained over 91 % cross-lingual accuracy in all models using 3-second audio. In SV tasks, the EER among cross-lingual test is at most 6.52 %, which is observed on the model trained by English corpus. The outcome suggests the feasibility of adopting cross-lingual speech in building text-independent speaker recognition systems.
Yi-Chieh Wu, Wen-Hung Liao
ICPR1
2019 Analyzing Social Network Data Using Deep Neural Networks: A Case Study Using Twitter Posts
abstract
The limitation on the total number of characters compels Twitter users to compose their messages more succinctly, suggesting a stronger association between text and image. In this paper, we employ computer vision and word embedding techniques to analyze the relationship between image content and text messages and explore the rich information entangled. Specifically, we collected all tweets which include keywords related to Taiwan during 2017. After data cleaning, we apply machine learning techniques to classify tweets into travel and non-travel types. This is achieved by employing deep neural networks to process and integrate text and image information. Within each class, we use hierarchical clustering to further partition the data into different clusters and investigate their characteristics. Through this research, we expect to identify the relationship between text and images in a tweet and gain more understanding of the properties of tweets on social networking platforms. The proposed framework and corresponding analytical results should also prove useful for qualitative research.
Wen-Hung Liao, Yen-Ting Huang, Tsu-Hsuan Yang, Yi-Chieh Wu
ISM4
2019 Inferring Pareto-optimal reconciliations across multiple event costs under the duplication-loss-coalescence model
abstract
BACKGROUND: Reconciliation methods are widely used to explain incongruence between a gene tree and species tree. However, the common approach of inferring maximum parsimony reconciliations (MPRs) relies on user-defined costs for each type of event, which can be difficult to estimate. Prior work has explored the relationship between event costs and maximum parsimony reconciliations in the duplication-loss and duplication-transfer-loss models, but no studies have addressed this relationship in the more complicated duplication-loss-coalescence model. RESULTS: We provide a fixed-parameter tractable algorithm for computing Pareto-optimal reconciliations and recording all events that arise in those reconciliations, along with their frequencies. We apply this method to a case study of 16 fungi to systematically characterize the complexity of MPR space across event costs and identify events supported across this space. CONCLUSION: This work provides a new framework for studying the relationship between event costs and reconciliations that incorporates both macro-evolutionary events and population effects and is thus broadly applicable across eukaryotic species.
Ross Mawhorter, Nuo Liu, Ran Libeskind-Hadas, Yi-Chieh Wu
BMC Bioinform.4
2019 Computing the Diameter of the Space of Maximum Parsimony Reconciliations in the Duplication-Transfer-Loss Model
abstract
Phylogenetic tree reconciliation is widely used in the fields of molecular evolution, cophylogenetics, parasitology, and biogeography to study the evolutionary histories of pairs of entities. In these contexts, reconciliation is often performed using maximum parsimony under the Duplication-Transfer-Loss (DTL) event model. In general, the number of maximum parsimony reconciliations (MPRs) can grow exponentially with the size of the trees. While a number of previous efforts have been made to count the number of MPRs, find representative MPRs, and compute the frequencies of events across the space of MPRs, little is known about the structure of MPR space. In particular, how different are MPRs in terms of the events that they comprise? One way to address this question is to compute the diameter of MPR space, defined to be the maximum number of DTL events that distinguish any two MPRs in the solution space. We show how to compute the diameter of MPR space in polynomial time and then apply this algorithm to a large biological dataset to study the variability of events.
Jordan Haack, Eli Zupke, Andrew Ramirez, Yi-Chieh Wu, Ran Libeskind-Hadas
IEEE ACM Trans. Comput. Biol. Bioinform.4
2018 Evaluation of Student's 3D Modeling Capability Based on Model Completeness and Usage Pattern in K-12 Classrooms
abstract
As more schools incorporate 3D printing into their curriculum to stimulate the creativity of K-12 students with a learning-by-doing approach, it becomes crucial to understand how users work with 3D modeling tools. In this paper, we aim to develop model and usage-pattern-related features to quantize students' performance on 3D modeling operation. The dataset is gathered from the Affiliated Experimental Elementary School (AEES) of National Chengchi University. Participants' operation log and finished work for specific 3D modeling software are recorded and analyzed. In all our lesson plans, students are required to create structurally stable and printable 3D models. Three modeling software with different levels of difficulty have been introduced and tested. The collected data include screen recording, software operation log, experts evaluation, and interviews with students, which are employed for subsequent qualitative evaluation as well as quantitative analysis. With our proposed approach, we are able to identify the key factors affecting students' learning experience and performance in terms of model completeness and usage pattern. Through these indicators, instructors can understand student's learning status of 3D modeling software more comprehensively.
Yi-Chieh Wu, Wen-Hung Liao, Chen-Yu Liu, Tsai-Yen Li, Ming-Te Chi
ICALT1
2017 Coestimation of Gene Trees and Reconciliations Under a Duplication-Loss-Coalescence Model
Yi-Chieh Wu
ISBRA2
2017 Classification of Reading Patterns Based on Gaze Information
abstract
Reading is one of the main paths to acquire knowledge, either done traditionally on paper media or practiced on electronic devices. Efficiency varies when different reading patterns are involved. It is the objective of this research to classify reading patterns from fixation data using machine learning techniques in an attempt to understand and evaluate the reading and learning process. In our experiment, a low-cost eye tracker is employed to record the eye movements during the reading process. A dispersion-based algorithm is implemented to identify fixation from the recorded data. Features pertaining to fixation including duration, path length, landing position and fixation direction are extracted for classification purposes. Five categories of reading pattern have been defined and investigated in this study, namely, speed reading, slow reading, in-depth reading, skim-and-skip, and keyword spotting. We have recruited thirty subjects to participate in our experiment. The participants are instructed to read different articles using specific styles designated by the experimenter in order to assign label to the collected data. Feature selection is achieved by analyzing the predictive results of cross-validation from the training data obtained from all subjects. The average classification accuracies in five random tests are 78.24%, 74.19%, 93.75%, 87.96%, and 96.20% respectively. Further improvements are accomplished by introducing an additional undecided class to address ambiguous reading patterns.
Wen-Hung Liao, Chin-Wen Chang, Yi-Chieh Wu
ISM3
2017 Reconciliation feasibility in the presence of gene duplication, loss, and coalescence with multiple individuals per species
abstract
BACKGROUND: In phylogenetics, we often seek to reconcile gene trees with species trees within the framework of an evolutionary model. While the most popular models for eukaryotic species allow for only gene duplication and gene loss or only multispecies coalescence, recent work has combined these phenomena through a reconciliation structure, the labeled coalescent tree (LCT), that simultaneously describes the duplication-loss and coalescent history of a gene family. However, the LCT makes the simplifying assumption that only one individual is sampled per species whereas, with advances in gene sequencing, we now have access to multiple samples per species. RESULTS: We demonstrate that with these additional samples, there exist gene tree topologies that are impossible to reconcile with any species tree. In particular, the multiple samples enforce new constraints on the placement of duplications within a valid reconciliation. To model these constraints, we extend the LCT to a new structure, the partially labeled coalescent tree (PLCT) and demonstrate how to use the PLCT to evaluate the feasibility of a gene tree topology. We apply our algorithm to two clades of apes and flies to characterize possible sources of infeasibility. CONCLUSION: Going forward, we believe that this model represents a first step towards understanding reconciliations in duplication-loss-coalescence models with multiple samples per species.
Jennifer Rogers, Andrew Fishberg, Nora Youngs, Yi-Chieh Wu
BMC Bioinform.4
2015 Improved gene tree error correction in the presence of horizontal gene transfer
abstract
MOTIVATION: The accurate inference of gene trees is a necessary step in many evolutionary studies. Although the problem of accurate gene tree inference has received considerable attention, most existing methods are only applicable to gene families unaffected by horizontal gene transfer. As a result, the accurate inference of gene trees affected by horizontal gene transfer remains a largely unaddressed problem. RESULTS: In this study, we introduce a new and highly effective method for gene tree error correction in the presence of horizontal gene transfer. Our method efficiently models horizontal gene transfers, gene duplications and losses, and uses a statistical hypothesis testing framework [Shimodaira-Hasegawa (SH) test] to balance sequence likelihood with topological information from a known species tree. Using a thorough simulation study, we show that existing phylogenetic methods yield inaccurate gene trees when applied to horizontally transferred gene families and that our method dramatically improves gene tree accuracy. We apply our method to a dataset of 11 cyanobacterial species and demonstrate the large impact of gene tree accuracy on downstream evolutionary analyses. AVAILABILITY AND IMPLEMENTATION: An implementation of our method is available at http://compbio.mit.edu/treefix-dtl/ CONTACT: : [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mukul S. Bansal, Yi-Chieh Wu, Eric J. Alm, Manolis Kellis
Bioinform.2
2014 Pareto-optimal phylogenetic tree reconciliation
abstract
MOTIVATION: Phylogenetic tree reconciliation is a widely used method for reconstructing the evolutionary histories of gene families and species, hosts and parasites and other dependent pairs of entities. Reconciliation is typically performed using maximum parsimony, in which each evolutionary event type is assigned a cost and the objective is to find a reconciliation of minimum total cost. It is generally understood that reconciliations are sensitive to event costs, but little is understood about the relationship between event costs and solutions. Moreover, choosing appropriate event costs is a notoriously difficult problem. RESULTS: We address this problem by giving an efficient algorithm for computing Pareto-optimal sets of reconciliations, thus providing the first systematic method for understanding the relationship between event costs and reconciliations. This, in turn, results in new techniques for computing event support values and, for cophylogenetic analyses, performing robust statistical tests. We provide new software tools and demonstrate their use on a number of datasets from evolutionary genomic and cophylogenetic studies. AVAILABILITY AND IMPLEMENTATION: Our Python tools are freely available at www.cs.hmc.edu/∼hadas/xscape. .
Ran Libeskind-Hadas, Yi-Chieh Wu, Mukul S. Bansal, Manolis Kellis
Bioinform.2