Chih-Hao Hsu

dblp:77/3973 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
3since 2021 · last 2026
0000-0001-7145-5780ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Computer networks · 2 · 1 first-authorSystems, architecture and hardware · 1Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Information extraction and text analysis · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
fact-checking
0.812024
CFEVER: A Chinese Fact Extraction and VERification Dataset · AAAI 2024
Natural language and speech › Information extraction and text analysis › fact-checking
fact extraction and verification
0.812024
CFEVER: A Chinese Fact Extraction and VERification Dataset · AAAI 2024
Information retrieval › evaluation
benchmark dataset
0.812024
CFEVER: A Chinese Fact Extraction and VERification Dataset · AAAI 2024
Bioinformatics and computational biology
comparative genomics
0.112007
HomologMiner: looking for homologous genomic groups in whole genomes · Bioinform. 2007
Bioinformatics and computational biology › genomics
genome analysis
0.112007
HomologMiner: looking for homologous genomic groups in whole genomes · Bioinform. 2007
Bioinformatics and computational biology › sequence analysis
repeat detection
0.112007
HomologMiner: looking for homologous genomic groups in whole genomes · Bioinform. 2007
Bioinformatics and computational biology › sequence analysis
sequence similarity search
0.112007
HomologMiner: looking for homologous genomic groups in whole genomes · Bioinform. 2007

Methods — techniques the papers use, named apart from their topics

inter-annotator agreement · 1.5dataset construction · 1.5tandem repeat detection · 0.1sequence clustering · 0.1
YearPublicationVenuePosition
2026 An ORID-Structured GenAI Reading Companion Integrating Structured Book Chat and Virtual Labs for Science Reading
Hui-Chun Hung, Chih-Hao Hsu, Shu-I Fang, Chen-Chung Liu, Chia-Hui Chang, Ying-Tien Wu
AIED (3)3
2025 From Persona to Person: Enhancing the Naturalness with Multiple Discourse Relations Graph Learning in Personalized Dialogue Generation
Chih-Hao Hsu, Ying-Jia Lin, Hung-Yu Kao
PAKDD (2)1
2024 CFEVER: A Chinese Fact Extraction and VERification Dataset
abstract
We present CFEVER, a Chinese dataset designed for Fact Extraction and VERification. CFEVER comprises 30,012 manually created claims based on content in Chinese Wikipedia. Each claim in CFEVER is labeled as “Supports”, “Refutes”, or “Not Enough Info” to depict its degree of factualness. Similar to the FEVER dataset, claims in the “Supports” and “Refutes” categories are also annotated with corresponding evidence sentences sourced from single or multiple pages in Chinese Wikipedia. Our labeled dataset holds a Fleiss’ kappa value of 0.7934 for five-way inter-annotator agreement. In addition, through the experiments with the state-of-the-art approaches developed on the FEVER dataset and a simple baseline for CFEVER, we demonstrate that our dataset is a new rigorous benchmark for factual extraction and verification, which can be further used for developing automated systems to alleviate human fact-checking efforts. CFEVER is available at https://ikmlab.github.io/CFEVER.
Ying-Jia Lin, Chia-Jen Yeh, Yi-Ting Li, Yun-Yu Hu, Chih-Hao Hsu, Mei-Feng Lee, Hung-Yu Kao
AAAI6
2015 SubHunter: a high-performance and scalable sub-circuit recognition method with Prüfer-encoding
Hong-Yan Su, Chih-Hao Hsu, Yih-Lang Li
DATE2
2014 Robust signal synthesis of the 12-lead ECG using 3-Lead wireless ECG systems
abstract
A new signal synthesis and tracking method is developed for a 3-Lead wireless electrocardiography (ECG) system. Exploiting the temporal correlations in ECG signals, the standard 12-lead ECG signals can be synthesized with high precision by using only three differential pairs of electrodes. The correlations between the original and synthesized ECG leads can be as high as 0.98. And the separation between the two electrodes of a pair can be drastically reduced to around 10 cm, which is about the size of a medium adhesive tape. Experiment results also show that the proposed ECG synthesis method is much more robust to variations in acquisition positions when taking into account the spatial correlations among the ECG leads in syntheses. This greatly improves the feasibility of using fully wireless ECG systems on the long-term monitoring and diagnosis of heart diseases.
Chih-Hao Hsu, Sau-Hsuan Wu
ICC1
2011 Multicast Lifetime Maximization Using Network Coding in Lossy Wireless Ad Hoc Networks
abstract
In traditional stop-and-wait strategy for reliable communications, such as ARQ, retransmission for the packet loss problem would incur a great number of packet transmissions in lossy wireless ad-hoc networks. We study the reliable multicast lifetime maximization problem by alternatively exploring the random linear network coding in this paper. We formulate such problem as a min-max problem and propose a heuristic algorithm, called maximum lifetime tree (MLT), to build a multicast tree that maximizes the network lifetime. Simulation results show that the proposed algorithms can significantly increase the network lifetime when compared with the traditional algorithms under various distributions of error probability on lossy wireless links.
Chih-Hao Hsu, Peng Li 0017, Song Guo 0001, Shui Yu 0001, Zhuzhong Qian
EUC1
2011 Energy-Aware Transmission Scheduling in Mobile Sensor Networks
abstract
The great diversity of mobile sensor networks (MSNs) has emerged in different networks, including vehicular ad-hoc networks (VANETs), underwater sensor networks (UWSNs), and wireless body area networks (WBANs), which provide ubiquitous solutions for real-time monitoring. To prolong the network lifetime of MSNs, energy conservation for mobile sensors needs to be taken into consideration while we design the scheduling for MSNs. In this paper, we define the energy minimization problem for energy- aware transmission scheduling for MSNs, and prove that the problem is NP-hard. To have an optimal solution for the problem, we first formulate the problem as an Integer Linear Programming (ILP) problem, and then propose a greedy algorithm to approximate the solution of the ILP problem. We show that the computational complexity of the proposed greedy algorithm is low. We also propose a reporting mechanism to accommodate our greedy algorithm in MSNs. Simulation experiments are conducted to investigate the performance of the proposed reporting mechanism. Our performance evaluation shows that our mechanism has the mobile sensors transmit real-time sensed data with less energy consumption.
Hou-Chun Chen, Huai-Lei Fu, Phone Lin, Chih-Hao Hsu
GLOBECOM4
2011 Evaluation of methods for detecting conversion events in gene clusters
abstract
BACKGROUND: Gene clusters are genetically important, but their analysis poses significant computational challenges. One of the major reasons for these difficulties is gene conversion among the duplicated regions of the cluster, which can obscure their true relationships. Many computational methods for detecting gene conversion events have been released, but their performance has not been assessed for wide deployment in evolutionary history studies due to a lack of accurate evaluation methods. RESULTS: We designed a new method that simulates gene cluster evolution, including large-scale events of duplication, deletion, and conversion as well as small mutations. We used this simulation data to evaluate several different programs for detecting gene conversion events. CONCLUSIONS: Our evaluation identifies strengths and weaknesses of several methods for detecting gene conversion, which can contribute to more accurate analysis of gene cluster evolution.
Giltae Song, Chih-Hao Hsu, Cathy Riemer, Webb Miller
BMC Bioinform.2
2007 HomologMiner: looking for homologous genomic groups in whole genomes
abstract
MOTIVATION: Complex genomes contain numerous repeated sequences, and genomic duplication is believed to be a main evolutionary mechanism to obtain new functions. Several tools are available for de novo repeat sequence identification, and many approaches exist for clustering homologous protein sequences. We present an efficient new approach to identify and cluster homologous DNA sequences with high accuracy at the level of whole genomes, excluding low-complexity repeats, tandem repeats and annotated interspersed repeats. We also determine the boundaries of each group member so that it closely represents a biological unit, e.g. a complete gene, or a partial gene coding a protein domain. RESULTS: We developed a program called HomologMiner to identify homologous groups applicable to genome sequences that have been properly marked for low-complexity repeats and annotated interspersed repeats. We applied it to the whole genomes of human (hg17), macaque (rheMac2) and mouse (mm8). Groups obtained include gene families (e.g. olfactory receptor gene family, zinc finger families), unannotated interspersed repeats and additional homologous groups that resulted from recent segmental duplications. Our program incorporates several new methods: a new abstract definition of consistent duplicate units, a new criterion to remove moderately frequent tandem repeats, and new algorithmic techniques. We also provide preliminary analysis of the output on the three genomes mentioned above, and show several applications including identifying boundaries of tandem gene clusters and novel interspersed repeat families. AVAILABILITY: All programs and datasets are downloadable from www.bx.psu.edu/miller_lab.
Minmei Hou, Piotr Berman, Chih-Hao Hsu, Robert S. Harris
Bioinform.3
2005 An Efficient Algorithm for Near Optimal Data Allocation on Multiple Broadcast Channels
Chih-Hao Hsu, Guanling Lee, Arbee L. P. Chen
Distributed Parallel Databases1
2002 Index and Data Allocation on Multiple Broadcast Channels Considering Data Access Frequencies
abstract
In a wireless environment, the bandwidth of the channels and the energy of the portable devices are limited. Data broadcast has become an excellent method for efficient data dissemination. In this paper the problem for generating a broadcast program of a set of data items with the associated access frequencies on multiple channels is explored. In our approach, we consider allocating index information and data items on multiple broadcast channels by extending the distributed indexing approach. Moreover global data replication and local data allocation are performed to improve the average access time of all data items. Simulation is performed to compare the performance of our approach with an existing approach. The result of the experiments shows that our approach outperforms the existing approach.
Chih-Hao Hsu, Guanling Lee, Arbee L. P. Chen
Mobile Data Management1
2001 A Near Optimal Algorithm for Generating Broadcast Programs on Multiple Channels
abstract
In a wireless environment, the bandwidth of the channels and the energy of the portable devices are limited. Data broadcast has become an excellent method for efficient data dissemination. In this paper, the problem for generating a broadcast program of a set of data items with the associated access frequencies on multiple channels is explored. In our approach, an expected average access time of the broadcast data items is first derived. The broadcast program is then generated, which minimizes the expected average access time. Simulation is performed to compare the performance of our approach with two existing approaches. The result of the experiments shows that our approach outperforms others and is in fact close to the optimal.
Chih-Hao Hsu, Guanling Lee, Arbee L. P. Chen
CIKM1