VLDB 2026 Research / reviewers in the wild / expert
Wen-Lian Hsu
dblp:10/4757
· DBLP profile ↗
119ranked-venue papers
18as first author
7since 2021 · last 2026
0000-0001-7061-3513ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 49 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 39 · 1 since 2021Theory of computation · 21 · 14 first-authorGraphics, computer vision, multimedia, augmented reality and games · 11Databases, data management, data science and information retrieval · 10 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-authorSecurity and privacy · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Large language model ensemble for automated TNM staging from radiology reportsabstractMOTIVATION: Accurate TNM staging from lung cancer radiology reports is crucial for treatment planning and prognosis assessment. Manual staging processes are time-consuming and subject to inter-observer variability. Large language models (LLMs) offer opportunities to automate TNM staging with enhanced interpretability and clinical reasoning. RESULTS: We developed two complementary systems for automated TNM staging from English radiology reports. System I employs GPT-4o with reasoning-based few-shot learning and multi-step voting. System II integrates multiple LLMs (GPT-4o and Gemini-2) using DSPy framework with MIPROv2 optimization. In NTCIR-18 RadNLP 2024 English main task, our approaches achieved first (joint accuracy: 0.6543) and second place (joint accuracy: 0.6296), demonstrating superior performance in T, N, and M classification with accuracies of 0.7037/0.9136/0.8889 and 0.7284/0.9383/0.8395, respectively. AVAILABILITY AND IMPLEMENTATION: Source code freely available at https://github.com/nlptmu/multi-expert-tnm-staging under MIT license. An archival snapshot of the version used in this study is deposited on Zenodo at https://doi.org/10.5281/zenodo.20338561. Implemented in Python 3.12+ with PyTorch 2.6 and DSPY 3.0, supporting Linux. Wen-Chao Yeh, Yi-Shin Chen, Wen-Lian Hsu, Shuntaro Yada, Yung-Chun Chang |
Bioinform. | 3 |
| 2024 | Surveying biomedical relation extraction: a critical examination of current datasets and the proposal of a new resourceabstractNatural language processing (NLP) has become an essential technique in various fields, offering a wide range of possibilities for analyzing data and developing diverse NLP tasks. In the biomedical domain, understanding the complex relationships between compounds and proteins is critical, especially in the context of signal transduction and biochemical pathways. Among these relationships, protein-protein interactions (PPIs) are of particular interest, given their potential to trigger a variety of biological reactions. To improve the ability to predict PPI events, we propose the protein event detection dataset (PEDD), which comprises 6823 abstracts, 39 488 sentences and 182 937 gene pairs. Our PEDD dataset has been utilized in the AI CUP Biomedical Paper Analysis competition, where systems are challenged to predict 12 different relation types. In this paper, we review the state-of-the-art relation extraction research and provide an overview of the PEDD's compilation process. Furthermore, we present the results of the PPI extraction competition and evaluate several language models' performances on the PEDD. This paper's outcomes will provide a valuable roadmap for future studies on protein event detection in NLP. By addressing this critical challenge, we hope to enable breakthroughs in drug discovery and enhance our understanding of the molecular mechanisms underlying various diseases. Ming-Siang Huang, Jen-Chieh Han, Pei-Yen Lin, Yu-Ting You, Richard Tzong-Han Tsai, Wen-Lian Hsu |
Briefings Bioinform. | 6 |
| 2023 | Semantic Template-based Convolutional Neural Network for Text ClassificationabstractWe propose a semantic template-based distributed representation for the convolutional neural network called Semantic Template-based Convolutional Neural Network (STCNN) for text categorization that imitates the perceptual behavior of human comprehension. STCNN is a highly automatic approach that learns semantic templates that characterize a domain from raw text and recognizes categories of documents using a semantic-infused convolutional neural network that allows a template to be partially matched through a statistical scoring system. Our experiment results show that STCNN effectively classifies documents in about 140,000 Chinese news articles into predefined categories by capturing the most prominent and expressive patterns and achieves the best performance among all compared methods for Chinese topic classification. Finally, the same knowledge can be directly used to perform a semantic analysis task. Yung-Chun Chang, Siu Hin Ng, Jung-Peng Chen, Yu-Chi Liang, Wen-Lian Hsu |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2021 | LBERT: Lexically-aware Transformers based Bidirectional Encoder Representation model for learning Universal Bio-Entity RelationsabstractMOTIVATION Natural Language Processing techniques are constantly being advanced to accommodate the influx of data as well as to provide exhaustive and structured knowledge dissemination. Within the biomedical domain, relation detection between bio-entities known as the Bio-Entity Relation Extraction (BRE) task has a critical function in knowledge structuring. Although recent advances in deep learning-based biomedical domain embedding have improved BRE predictive analytics, these works are often task selective or use external knowledge-based pre-/post-processing. In addition, deep learning-based models do not account for local syntactic contexts, which have improved data representation in many kernel classifier-based models. In this study, we propose a universal BRE model, i.e. LBERT, which is a Lexically aware Transformer-based Bidirectional Encoder Representation model, and which explores both local and global contexts representations for sentence-level classification tasks. RESULTS This article presents one of the most exhaustive BRE studies ever conducted over five different bio-entity relation types. Our model outperforms state-of-the-art deep learning models in protein-protein interaction (PPI), drug-drug interaction and protein-bio-entity relation classification tasks by 0.02%, 11.2% and 41.4%, respectively. LBERT representations show a statistically significant improvement over BioBERT in detecting true bio-entity relation for large corpora like PPI. Our ablation studies clearly indicate the contribution of the lexical features and distance-adjusted attention in improving prediction performance by learning additional local semantic context along with bi-directionally learned global context. AVAILABILITY AND IMPLEMENTATION Github. https://github.com/warikoone/LBERT. SUPPLEMENTARY INFORMATION Supplementary data are available at Bioinformatics online. Neha Warikoo, Yung-Chun Chang, Wen-Lian Hsu |
Bioinform. | 3 |
| 2021 | LBERT: Lexically aware Transformer-based Bidirectional Encoder Representation model for learning universal bio-entity relationsabstractMOTIVATION: Natural Language Processing techniques are constantly being advanced to accommodate the influx of data as well as to provide exhaustive and structured knowledge dissemination. Within the biomedical domain, relation detection between bio-entities known as the Bio-Entity Relation Extraction (BRE) task has a critical function in knowledge structuring. Although recent advances in deep learning-based biomedical domain embedding have improved BRE predictive analytics, these works are often task selective or use external knowledge-based pre-/post-processing. In addition, deep learning-based models do not account for local syntactic contexts, which have improved data representation in many kernel classifier-based models. In this study, we propose a universal BRE model, i.e. LBERT, which is a Lexically aware Transformer-based Bidirectional Encoder Representation model, and which explores both local and global contexts representations for sentence-level classification tasks. RESULTS: This article presents one of the most exhaustive BRE studies ever conducted over five different bio-entity relation types. Our model outperforms state-of-the-art deep learning models in protein-protein interaction (PPI), drug-drug interaction and protein-bio-entity relation classification tasks by 0.02%, 11.2% and 41.4%, respectively. LBERT representations show a statistically significant improvement over BioBERT in detecting true bio-entity relation for large corpora like PPI. Our ablation studies clearly indicate the contribution of the lexical features and distance-adjusted attention in improving prediction performance by learning additional local semantic context along with bi-directionally learned global context. AVAILABILITY AND IMPLEMENTATION: Github. https://github.com/warikoone/LBERT. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Neha Warikoo, Yung-Chun Chang, Wen-Lian Hsu |
Bioinform. | 3 |
| 2021 | Co-AMPpred for in silico-aided predictions of antimicrobial peptides by integrating composition-based featuresabstractBACKGROUND: Antimicrobial peptides (AMPs) are oligopeptides that act as crucial components of innate immunity, naturally occur in all multicellular organisms, and are involved in the first line of defense function. Recent studies showed that AMPs perpetuate great potential that is not limited to antimicrobial activity. They are also crucial regulators of host immune responses that can modulate a wide range of activities, such as immune regulation, wound healing, and apoptosis. However, a microorganism's ability to adapt and to resist existing antibiotics triggered the scientific community to develop alternatives to conventional antibiotics. Therefore, to address this issue, we proposed Co-AMPpred, an in silico-aided AMP prediction method based on compositional features of amino acid residues to classify AMPs and non-AMPs. RESULTS: In our study, we developed a prediction method that incorporates composition-based sequence and physicochemical features into various machine-learning algorithms. Then, the boruta feature-selection algorithm was used to identify discriminative biological features. Furthermore, we only used discriminative biological features to develop our model. Additionally, we performed a stratified tenfold cross-validation technique to validate the predictive performance of our AMP prediction model and evaluated on the independent holdout test dataset. A benchmark dataset was collected from previous studies to evaluate the predictive performance of our model. CONCLUSIONS: Experimental results show that combining composition-based and physicochemical features outperformed existing methods on both the benchmark training dataset and a reduced training dataset. Finally, our proposed method achieved 80.8% accuracies and 0.871 area under the receiver operating characteristic curve by evaluating on independent test set. Our code and datasets are available at https://github.com/onkarS23/CoAMPpred . Onkar Singh, Wen-Lian Hsu, Emily Chia-Yu Su |
BMC Bioinform. | 2 |
| 2021 | A flexible template generation and matching method with applications for publication reference metadata extractionabstractAbstract Conventional rule‐based approaches use exact template matching to capture linguistic information and necessarily need to enumerate all variations. We propose a novel flexible template generation and matching scheme called the principle‐based approach (PBA) based on sequence alignment, and employ it for reference metadata extraction (RME) to demonstrate its effectiveness. The main contributions of this research are threefold. First, we propose an automatic template generation that can capture prominent patterns using the dominating set algorithm. Second, we devise an alignment‐based template‐matching technique that uses a logistic regression model, which makes it more general and flexible than pure rule‐based approaches. Last, we apply PBA to RME on extensive cross‐domain corpora and demonstrate its robustness and generality. Experiments reveal that the same set of templates produced by the PBA framework not only deliver consistent performance on various unseen domains, but also surpass hand‐crafted knowledge (templates). We use four independent journal style test sets and one conference style test set in the experiments. When compared to renowned machine learning methods, such as conditional random fields (CRF), as well as recent deep learning methods (i.e., bi‐directional long short‐term memory with a CRF layer, Bi‐LSTM‐CRF), PBA has the best performance for all datasets. Ting-Hao Yang, Yu-Lun Hsieh, Shih-Hung Liu, Yung-Chun Chang, Wen-Lian Hsu |
J. Assoc. Inf. Sci. Technol. | 5 |
| 2020 | KIDER: Knowledge-Infused Document Embedding Representation for Text Categorization
Zheng-Wen Lin, Yung-Chun Chang, Wen-Lian Hsu |
IEA/AIE | 4 |
| 2020 | Biomedical named entity recognition and linking datasets: survey and our recent developmentabstractNatural language processing (NLP) is widely applied in biological domains to retrieve information from publications. Systems to address numerous applications exist, such as biomedical named entity recognition (BNER), named entity normalization (NEN) and protein-protein interaction extraction (PPIE). High-quality datasets can assist the development of robust and reliable systems; however, due to the endless applications and evolving techniques, the annotations of benchmark datasets may become outdated and inappropriate. In this study, we first review commonlyused BNER datasets and their potential annotation problems such as inconsistency and low portability. Then, we introduce a revised version of the JNLPBA dataset that solves potential problems in the original and use state-of-the-art named entity recognition systems to evaluate its portability to different kinds of biomedical literature, including protein-protein interaction and biology events. Lastly, we introduce an ensembled biomedical entity dataset (EBED) by extending the revised JNLPBA dataset with PubMed Central full-text paragraphs, figure captions and patent abstracts. This EBED is a multi-task dataset that covers annotations including gene, disease and chemical entities. In total, it contains 85000 entity mentions, 25000 entity mentions with database identifiers and 5000 attribute tags. To demonstrate the usage of the EBED, we review the BNER track from the AI CUP Biomedical Paper Analysis challenge. Availability: The revised JNLPBA dataset is available at https://iasl-btm.iis.sinica.edu.tw/BNER/Content/Re vised_JNLPBA.zip. The EBED dataset is available at https://iasl-btm.iis.sinica.edu.tw/BNER/Content/AICUP _EBED_dataset.rar. Contact: Email: [email protected], Tel. 886-3-4227151 ext. 35203, Fax: 886-3-422-2681 Email: [email protected], Tel. 886-2-2788-3799 ext. 2211, Fax: 886-2-2782-4814 Supplementary information: Supplementary data are available at Briefings in Bioinformatics online. Ming-Siang Huang, Po-Ting Lai, Pei-Yen Lin, Yu-Ting You, Richard Tzong-Han Tsai, Wen-Lian Hsu |
Briefings Bioinform. | 6 |
| 2019 | On the Robustness of Self-Attentive ModelsabstractThis work examines the robustness of selfattentive neural networks against adversarial input perturbations.Specifically, we investigate the attention and feature extraction mechanisms of state-of-the-art recurrent neural networks and self-attentive architectures for sentiment analysis, entailment and machine translation under adversarial attacks.We also propose a novel attack algorithm for generating more natural adversarial examples that could mislead neural models but not humans.Experimental results show that, compared to recurrent neural models, self-attentive models are more robust against adversarial perturbation.In addition, we provide theoretical explanations for their superior robustness to support our claims. Yu-Lun Hsieh, Minhao Cheng, Da-Cheng Juan, Wei Wei 0019, Wen-Lian Hsu, Cho-Jui Hsieh |
ACL (1) | 5 |
| 2019 | Medical knowledge infused convolutional neural networks for cohort selection in clinical trialsabstractOBJECTIVE: In this era of digitized health records, there has been a marked interest in using de-identified patient records for conducting various health related surveys. To assist in this research effort, we developed a novel clinical data representation model entitled medical knowledge-infused convolutional neural network (MKCNN), which is used for learning the clinical trial criteria eligibility status of patients to participate in cohort studies. MATERIALS AND METHODS: In this study, we propose a clinical text representation infused with medical knowledge (MK). First, we isolate the noise from the relevant data using a medically relevant description extractor; then we utilize log-likelihood ratio based weights from selected sentences to highlight "met" and "not-met" knowledge-infused representations in bichannel setting for each instance. The combined medical knowledge-infused representation (MK) from these modules helps identify significant clinical criteria semantics, which in turn renders effective learning when used with a convolutional neural network architecture. RESULTS: MKCNN outperforms other Medical Knowledge (MK) relevant learning architectures by approximately 3%; notably SVM and XGBoost implementations developed in this study. MKCNN scored 86.1% on F1metric, a gain of 6% above the average performance assessed from the submissions for n2c2 task. Although pattern/rule-based methods show a higher average performance for the n2c2 clinical data set, MKCNN significantly improves performance of machine learning implementations for clinical datasets. CONCLUSION: MKCNN scored 86.1% on the F1 score metric. In contrast to many of the rule-based systems introduced during the n2c2 challenge workshop, our system presents a model that heavily draws on machine-based learning. In addition, the MK representations add more value to clinical comprehension and interpretation of natural texts. Chi-Jen Chen, Neha Warikoo, Yung-Chun Chang, Jin-Hua Chen, Wen-Lian Hsu |
J. Am. Medical Informatics Assoc. | 5 |
| 2018 | DART: a fast and accurate RNA-seq mapper with a partitioning strategyabstractMOTIVATION: In recent years, the massively parallel cDNA sequencing (RNA-Seq) technologies have become a powerful tool to provide high resolution measurement of expression and high sensitivity in detecting low abundance transcripts. However, RNA-seq data requires a huge amount of computational efforts. The very fundamental and critical step is to align each sequence fragment against the reference genome. Various de novo spliced RNA aligners have been developed in recent years. Though these aligners can handle spliced alignment and detect splice junctions, some challenges still remain to be solved. With the advances in sequencing technologies and the ongoing collection of sequencing data in the ENCODE project, more efficient alignment algorithms are highly demanded. Most read mappers follow the conventional seed-and-extend strategy to deal with inexact matches for sequence alignment. However, the extension is much more time consuming than the seeding step. RESULTS: We proposed a novel RNA-seq de novo mapping algorithm, call DART, which adopts a partitioning strategy to avoid the extension step. The experiment results on synthetic datasets and real NGS datasets showed that DART is a highly efficient aligner that yields the highest or comparable sensitivity and accuracy compared to most state-of-the-art aligners, and more importantly, it spends the least amount of time among the selected aligners. AVAILABILITY AND IMPLEMENTATION: https://github.com/hsinnan75/DART. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hsin-Nan Lin, Wen-Lian Hsu |
Bioinform. | 2 |
| 2017 | Leveraging manifold learning for extractive broadcast news summarizationabstractExtractive speech summarization is intended to produce a condensed version of the original spoken document by selecting a few salient sentences from the document and concatenate them together to form a summary. In this paper, we study a novel use of manifold learning techniques for extractive speech summarization. Manifold learning has experienced a surge of research interest in various domains concerned with dimensionality reduction and data representation recently, but has so far been largely under-explored in extractive text or speech summarization. Our contributions in this paper are at least twofold. First, we explore the use of several manifold learning algorithms to capture the latent semantic information of sentences for enhanced extractive speech summarization, including isometric feature mapping (ISOMAP), locally linear embedding (LLE) and Laplacian eigenmap. Second, the merits of our proposed summarization methods and several widely-used methods are extensively analyzed and compared. The empirical results demonstrate the effectiveness of our unsupervised summarization methods, in relation to several state-of-the-art methods. In particular, a synergy of the manifold learning based methods and state-of-the-art methods, such as the integer linear programming (ILP) method, contributes to further gains in summarization performance. Shih-Hung Liu, Kuan-Yu Chen 0002, Berlin Chen, Hsin-Min Wang, Wen-Lian Hsu |
ICASSP | 5 |
| 2017 | SPIRIT: A Tree Kernel-Based Method for Topic Person Interaction Detection (Extended Abstract)abstractIn this paper, we investigate the interactions between topic persons to help readers construct the background knowledge of a topic. We proposed a rich interactive tree structure to represent syntactic, context, and semantic information of text, and this structure is incorporated into a tree-based convolution kernel to identify segments that convey person interactions and further construct person interaction networks. Empirical evaluations demonstrate that the proposed method is effective in detecting and extracting the interactions between topic persons in the text, and outperforms other extraction approaches used for comparison. Furthermore, readers will be able to easily navigate through the topic persons of interest within the interaction networks, and further construct the background knowledge of the topic to facilitate comprehension. Yung-Chun Chang, Chien Chin Chen, Wen-Lian Hsu |
ICDE | 3 |
| 2017 | Kart: a divide-and-conquer algorithm for NGS read alignmentabstractMOTIVATION: Next-generation sequencing (NGS) provides a great opportunity to investigate genome-wide variation at nucleotide resolution. Due to the huge amount of data, NGS applications require very fast and accurate alignment algorithms. Most existing algorithms for read mapping basically adopt seed-and-extend strategy, which is sequential in nature and takes much longer time on longer reads. RESULTS: We develop a divide-and-conquer algorithm, called Kart, which can process long reads as fast as short reads by dividing a read into small fragments that can be aligned independently. Our experiment result indicates that the average size of fragments requiring the more time-consuming gapped alignment is around 20 bp regardless of the original read length. Furthermore, it can tolerate much higher error rates. The experiments show that Kart spends much less time on longer reads than other aligners and still produce reliable alignments even when the error rate is as high as 15%. AVAILABILITY AND IMPLEMENTATION: Kart is available at https://github.com/hsinnan75/Kart/ . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hsin-Nan Lin, Wen-Lian Hsu |
Bioinform. | 2 |
| 2017 | FISER: A Feature-Based Detection System for Person InteractionsabstractDiscovering the interactions between the persons mentioned in a set of topic documents can help readers construct the background of the topic and facilitate document comprehension. To discover person interactions, we need a detection method that can identify text segments containing information about the interactions. Information extraction algorithms then analyze the segments to extract interaction tuples and construct a network of person interaction. In this article, we define interaction detection as a classification problem. The proposed interaction detection method, called feature‐based interactive segment recognizer (FISER), exploits 19 features covering syntactic, context‐dependent, and semantic information in text to detect intra‐clausal and inter‐clausal interactive segments in topic documents. Empirical evaluations demonstrate that FISER outperformed many well‐known relation extraction and protein–protein interaction detection methods on identifying interactive segments in topic documents. In addition, the precision, recall, and F1‐score of the best feature combination are 72.9%, 55.8%, and 63.2%, respectively. Yung-Chun Chang, Pi-Hua Chuang, Chien Chin Chen, Wen-Lian Hsu |
Comput. Intell. | 4 |
| 2017 | A semantic frame-based intelligent agent for topic detection
Yung-Chun Chang, Yu-Lun Hsieh, Cen-Chieh Chen, Wen-Lian Hsu |
Soft Comput. | 4 |
| 2017 | A Position-Aware Language Modeling Framework for Extractive Broadcast News Speech SummarizationabstractExtractive summarization, a process that automatically picks exemplary sentences from a text (or spoken) document with the goal of concisely conveying key information therein, has seen a surge of attention from scholars and practitioners recently. Using a language modeling (LM) approach for sentence selection has been proven effective for performing unsupervised extractive summarization. However, one of the major difficulties facing the LM approach is to model sentences and estimate their parameters more accurately for each text (or spoken) document. We extend this line of research and make the following contributions in this work. First, we propose a position-aware language modeling framework using various granularities of position-specific information to better estimate the sentence models involved in the summarization process. Second, we explore disparate ways to integrate the positional cues into relevance models through a pseudo-relevance feedback procedure. Third, we extensively evaluate various models originated from our proposed framework and several well-established unsupervised methods. Empirical evaluation conducted on a broadcast news summarization task further demonstrates performance merits of the proposed summarization methods. Shih-Hung Liu, Kuan-Yu Chen 0002, Yu-Lun Hsieh, Berlin Chen, Hsin-Min Wang, Hsu-Chun Yen, Wen-Lian Hsu |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 7 |
| 2016 | Statistical Principle-Based Approach for Detecting miRNA-Target Gene Interaction ArticlesabstractMicroRNAs (miRNAs) are small non-coding RNAs of approximately 23 nucleotides, which negatively regulate the gene expression at the post-transcriptional level. miRNAs have been considered as good candidates for early detection or prognosis biomarkers for various diseases. Validated miRNA targets are usually reported in literature, necessitating researchers to manually screen through the related literature to keep up-to-date with novel findings. However, the amount of miRNA-related literature is increasing rapidly which makes it difficult for researchers to keep up to date. This study develops a text mining pipeline based on the statistical principle-based approach (SPBA) to detect MiRNA-Target Interactions (MTIs) mentioned in literatures. SPBA uses a collection of principles to represent linguistic concepts or rules used by human for describing MTIs. Each principle is composed of a collection of slots, which can be automatically learned from training data by merging the labeled slot sequences into more representative principles through a dominating set algorithm. Followed by a partial matching algorithm, the proposed approach can successfully recognize miRNA mentions and extract their MTIs in articles with a promising F-score of 98.8% and an accuracy of 71.43%. Nai-Wen Chang 0001, Hong-Jie Dai, Yu-Lun Hsieh, Wen-Lian Hsu |
BIBE | 4 |
| 2016 | Exploring Word Mover's Distance and Semantic-Aware Embedding Techniques for Extractive Broadcast News Summarization
Shih-Hung Liu, Kuan-Yu Chen 0002, Yu-Lun Hsieh, Berlin Chen, Hsin-Min Wang, Hsu-Chun Yen, Wen-Lian Hsu |
INTERSPEECH | 7 |
| 2016 | SPIRIT: A Tree Kernel-Based Method for Topic Person Interaction DetectionabstractThe development of a topic in a set of topic documents is constituted by a series of person interactions at a specific time and place. Knowing the interactions of the persons mentioned in these documents is helpful for readers to better comprehend the documents. In this paper, we propose a topic person interaction detection method called SPIRIT, which classifies the text segments in a set of topic documents that convey person interactions. We design the rich interactive tree structure to represent syntactic, context, and semantic information of text, and this structure is incorporated into a tree-based convolution kernel to identify interactive segments. Experiment results based on real world topics demonstrate that the proposed rich interactive tree structure effectively detects the topic person interactions and that our method outperforms many well-known relation extraction and protein-protein interaction methods. Yung-Chun Chang, Chien Chin Chen, Wen-Lian Hsu |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2015 | Positional language modeling for extractive broadcast news speech summarizationabstractExtractive summarization, with the intention of automatically selecting a set of representative sentences from a text (or spoken) document so as to concisely express the most important theme of the document, has been an active area of experimentation and development.A recent trend of research is to employ the language modeling (LM) approach for important sentence selection, which has proven to be effective for performing extractive summarization in an unsupervised fashion.However, one of the major challenges facing the LM approach is how to formulate the sentence models and estimate their parameters more accurately for each text (or spoken) document to be summarized.This paper extends this line of research and its contributions are three-fold.First, we propose a positional language modeling framework using different granularities of position-specific information to better estimate the sentence models involved in summarization.Second, we also explore to integrate the positional cues into relevance modeling through a pseudo-relevance feedback procedure.Third, the utilities of the various methods originated from our proposed framework and several well-established unsupervised methods are analyzed and compared extensively.Empirical evaluations conducted on a broadcast news summarization task seem to demonstrate the performance merits of our summarization methods. Shih-Hung Liu, Kuan-Yu Chen 0002, Berlin Chen, Hsin-Min Wang, Hsu-Chun Yen, Wen-Lian Hsu |
INTERSPEECH | 6 |
| 2015 | A context-aware approach for progression tracking of medical concepts in electronic medical recordsabstractElectronic medical records (EMRs) for diabetic patients contain information about heart disease risk factors such as high blood pressure, cholesterol levels, and smoking status. Discovering the described risk factors and tracking their progression over time may support medical personnel in making clinical decisions, as well as facilitate data modeling and biomedical research. Such highly patient-specific knowledge is essential to driving the advancement of evidence-based practice, and can also help improve personalized medicine and care. One general approach for tracking the progression of diseases and their risk factors described in EMRs is to first recognize all temporal expressions, and then assign each of them to the nearest target medical concept. However, this method may not always provide the correct associations. In light of this, this work introduces a context-aware approach to assign the time attributes of the recognized risk factors by reconstructing contexts that contain more reliable temporal expressions. The evaluation results on the i2b2 test set demonstrate the efficacy of the proposed approach, which achieved an F-score of 0.897. To boost the approach's ability to process unstructured clinical text and to allow for the reproduction of the demonstrated results, a set of developed .NET libraries used to develop the system is available at https://sites.google.com/site/hongjiedai/projects/nttmuclinicalnet. Nai-Wen Chang 0001, Hong-Jie Dai, Jitendra Jonnagaddala, Chih-Wei Chen, Richard Tzong-Han Tsai, Wen-Lian Hsu |
J. Biomed. Informatics | 6 |
| 2015 | Extractive Broadcast News Summarization Leveraging Recurrent Neural Network Language Modeling TechniquesabstractExtractive text or speech summarization manages to select a set of salient sentences from an original document and concatenate them to form a summary, enabling users to better browse through and understand the content of the document. A recent stream of research on extractive summarization is to employ the language modeling (LM) approach for important sentence selection, which has proven to be effective for performing speech summarization in an unsupervised fashion. However, one of the major challenges facing the LM approach is how to formulate the sentence models and accurately estimate their parameters for each sentence in the document to be summarized. In view of this, our work in this paper explores a novel use of recurrent neural network language modeling (RNNLM) framework for extractive broadcast news summarization. On top of such a framework, the deduced sentence models are able to render not only word usage cues but also long-span structural information of word co-occurrence relationships within broadcast news documents, getting around the need for the strict bag-of-words assumption. Furthermore, different model complexities and combinations are extensively analyzed and compared. Experimental results demonstrate the performance merits of our summarization methods when compared to several well-studied state-of-the-art unsupervised methods. Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang, Ea-Ee Jan, Wen-Lian Hsu, Hsin-Hsi Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2015 | Combining Relevance Language Modeling and Clarity Measure for Extractive Speech SummarizationabstractExtractive speech summarization, which purports to select an indicative set of sentences from a spoken document so as to succinctly represent the most important aspects of the document, has garnered much research over the years. In this paper, we cast extractive speech summarization as an ad-hoc information retrieval (IR) problem and investigate various language modeling (LM) methods for important sentence selection. The main contributions of this paper are four-fold. First, we explore a novel sentence modeling paradigm built on top of the notion of relevance, where the relationship between a candidate summary sentence and a spoken document to be summarized is discovered through different granularities of context for relevance modeling. Second, not only lexical but also topical cues inherent in the spoken document are exploited for sentence modeling. Third, we propose a novel clarity measure for use in important sentence selection, which can help quantify the thematic specificity of each individual sentence that is deemed to be a crucial indicator orthogonal to the relevance measure provided by the LM-based methods. Fourth, in an attempt to lessen summarization performance degradation caused by imperfect speech recognition, we investigate making use of different levels of index features for LM-based sentence modeling, including words, subword-level units, and their combination. Experiments on broadcast news summarization seem to demonstrate the performance merits of our methods when compared to several existing well-developed and/or state-of-the-art methods. Shih-Hung Liu, Kuan-Yu Chen 0002, Berlin Chen, Hsin-Min Wang, Hsu-Chun Yen, Wen-Lian Hsu |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2015 | Rank correlation analysis of RITE datasets and evaluation metrics - an observation on NTCIR-10 RITE Chinese subtasksabstractTextual Entailment (TE) is the task of recognizing entailment, paraphrase, and contradiction relations between a given text pair. The goal of textual entailment research is to develop a core inference component that can be applied to various domains such as QA or IR. We observed several rank correl ations on the test data and system results in the NTCIR-10 RITE-2 task, trying to find out correlations between datasets and evaluation metrics. We also constructed RITE4QA datasets in the RITE-2 task under the scenario of QA in order to see the applicability of RITE techniques in QA systems. Although we find that datasets created from different sources and different ways can hardly predict each other, we also find that ranking by RITE metrics has moderate correlation with the ranking by QA metrics if testing on artificial pairs. Both RITE metrics and QA metrics are stable in terms of their own subtasks. Chuan-Jie Lin, Cheng-Wei Lee 0001, Cheng-Wei Shih, Wen-Lian Hsu |
Web Intell. | 4 |
| 2014 | Leveraging Effective Query Modeling Techniques for Speech Recognition and SummarizationabstractStatistical language modeling (LM) that purports to quantify the acceptability of a given piece of text has long been an interesting yet challenging research area.In particular, language modeling for information retrieval (IR) has enjoyed remarkable empirical success; one emerging stream of the LM approach for IR is to employ the pseudo-relevance feedback process to enhance the representation of an input query so as to improve retrieval effectiveness.This paper presents a continuation of such a general line of research and the main contribution is threefold.First, we propose a principled framework which can unify the relationships among several widely-used query modeling formulations.Second, on top of the successfully developed framework, we propose an extended query modeling formulation by incorporating critical query-specific information cues to guide the model estimation.Third, we further adopt and formalize such a framework to the speech recognition and summarization tasks.A series of empirical experiments reveal the feasibility of such an LM framework and the performance merits of the deduced models on these two tasks. Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Ea-Ee Jan, Hsin-Min Wang, Wen-Lian Hsu, Hsin-Hsi Chen |
EMNLP | 6 |
| 2014 | Effective pseudo-relevance feedback for language modeling in extractive speech summarizationabstractExtractive speech summarization, aiming to automatically select an indicative set of sentences from a spoken document so as to concisely represent the most important aspects of the document, has become an active area for research and experimentation. An emerging stream of work is to employ the language modeling (LM) framework along with the Kullback-Leibler divergence measure for extractive speech summarization, which can perform important sentence selection in an unsupervised manner and has shown preliminary success. This paper presents a continuation of such a general line of research and its main contribution is two-fold. First, by virtue of pseudo-relevance feedback, we explore several effective sentence modeling formulations to enhance the sentence models involved in the LM-based summarization framework. Second, the utilities of our summarization methods and several widely-used methods are analyzed and compared extensively, which demonstrates the effectiveness of our methods. Shih-Hung Liu, Kuan-Yu Chen 0002, Yu-Lun Hsieh, Berlin Chen, Hsin-Min Wang, Hsu-Chun Yen, Wen-Lian Hsu |
ICASSP | 7 |
| 2014 | A recurrent neural network language modeling framework for extractive speech summarizationabstractExtractive speech summarization, with the purpose of automatically selecting a set of representative sentences from a spoken document so as to concisely express the most important theme of the document, has been an active area of research and development. A recent school of thought is to employ the language modeling (LM) approach for important sentence selection, which has proven to be effective for performing speech summarization in an unsupervised fashion. However, one of the major challenges facing the LM approach is how to formulate the sentence models and accurately estimate their parameters for each spoken document to be summarized. This paper presents a continuation of this general line of research and its contribution is two-fold. First, we propose a novel and effective recurrent neural network language modeling (RNNLM) framework for speech summarization, on top of which the deduced sentence models are able to render not only word usage cues but also long-span structural information of word co-occurrence relationships within spoken documents, getting around the need for the strict bag-of-words assumption. Second, the utilities of the method originated from our proposed framework and several widely-used unsupervised methods are analyzed and compared extensively. A series of experiments conducted on a broadcast news summarization task seem to demonstrate the performance merits of our summarization method when compared to several state-of-the-art existing unsupervised methods. Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang, Wen-Lian Hsu, Hsin-Hsi Chen |
ICME | 5 |
| 2014 | Semantic Frame-Based Natural Language Understanding for Intelligent Topic Detection Agent
Yung-Chun Chang, Yu-Lun Hsieh, Cen-Chieh Chen, Wen-Lian Hsu |
IEA/AIE (1) | 4 |
| 2014 | Enhanced language modeling for extractive speech summarization with sentence relatedness informationabstractExtractive summarization is intended to automatically select a set of representative sentences from a text or spoken document that can concisely express the most important topics of the document. Language modeling (LM) has been proven to be a promising framework for performing extractive summarization in an unsupervised manner. However, there remain two fundamental challenges facing existing LM-based methods. One is how to construct sentence models involved in the LM framework more accurately without resorting to external information sources. The other is how to additionally take into account the sentence-level structural relationships embedded in a document for important sentence selection. To address these two challenges, in this paper we explore a novel approach that generates overlapped clusters to extract sentence relatedness information from the document to be summarized, which can be used not only to enhance the estimation of various sentence models but also to allow for the sentence-level structural relationships for better summarization performance. Further, the utilities of our proposed methods and several state-of-the-art unsupervised methods are analyzed and compared extensively. A series of experiments conducted on a Mandarin broadcast news summarization task demonstrate the effectiveness and viability of our method. Index Terms: speech summarization, language modeling, clustering, relevance, sentence relatedness Shih-Hung Liu, Kuan-Yu Chen 0002, Yu-Lun Hsieh, Berlin Chen, Hsin-Min Wang, Hsu-Chun Yen, Wen-Lian Hsu |
INTERSPEECH | 7 |
| 2014 | Semantic Frame-based Statistical Approach for Topic Detection
Yung-Chun Chang, Yu-Lun Hsieh, Cen-Chieh Chen, Chad Liu, Chun-Hung Lu, Wen-Lian Hsu |
PACLIC | 6 |
| 2013 | Lipid exposure prediction enhances the inference of rotational angles of transmembrane helicesabstractBACKGROUND: Since membrane protein structures are challenging to crystallize, computational approaches are essential for elucidating the sequence-to-structure relationships. Structural modeling of membrane proteins requires a multidimensional approach, and one critical geometric parameter is the rotational angle of transmembrane helices. Rotational angles of transmembrane helices are characterized by their folded structures and could be inferred by the hydrophobic moment; however, the folding mechanism of membrane proteins is not yet fully understood. The rotational angle of a transmembrane helix is related to the exposed surface of a transmembrane helix, since lipid exposure gives the degree of accessibility of each residue in lipid environment. To the best of our knowledge, there have been few advances in investigating whether an environment descriptor of lipid exposure could infer a geometric parameter of rotational angle. RESULTS: Here, we present an analysis of the relationship between rotational angles and lipid exposure and a support-vector-machine method, called TMexpo, for predicting both structural features from sequences. First, we observed from the development set of 89 protein chains that the lipid exposure, i.e., the relative accessible surface area (rASA) of residues in the lipid environment, generated from high-resolution protein structures could infer the rotational angles with a mean absolute angular error (MAAE) of 46.32˚. More importantly, the predicted rASA from TMexpo achieved an MAAE of 51.05˚, which is better than 71.47˚ obtained by the best of the compared hydrophobicity scales. Lastly, TMexpo outperformed the compared methods in rASA prediction on the independent test set of 21 protein chains and achieved an overall Matthew's correlation coefficient, accuracy, sensitivity, specificity, and precision of 0.51, 75.26%, 81.30%, 69.15%, and 72.73%, respectively. TMexpo is publicly available at http://bio-cluster.iis.sinica.edu.tw/TMexpo. CONCLUSIONS: TMexpo can better predict rASA and rotational angles than the compared methods. When rotational angles can be accurately predicted, free modeling of transmembrane protein structures in turn may benefit from a reduced complexity in ensembles with a significantly less number of packing arrangements. Furthermore, sequence-based prediction of both rotational angle and lipid exposure can provide essential information when high-resolution structures are unavailable and contribute to experimental design to elucidate transmembrane protein functions. Jhih-Siang Lai, Cheng-Wei Cheng, Allan Lo, Ting-Yi Sung, Wen-Lian Hsu |
BMC Bioinform. | 5 |
| 2013 | TEMPTING system: A hybrid method of rule and machine learning for temporal relation extraction in patient discharge summaries
Yung-Chun Chang, Hong-Jie Dai, Johnny Chi-Yang Wu, Jian-Ming Chen, Richard Tzong-Han Tsai, Wen-Lian Hsu |
J. Biomed. Informatics | 6 |
| 2013 | The Left and Right Context of a Word: Overlapping Chinese Syllable Word Segmentation with Minimal ContextabstractSince a Chinese syllable can correspond to many characters (homophones), the syllable-to-character conversion task is quite challenging for Chinese phonetic input methods (CPIM). There are usually two stages in a CPIM: 1. segment the syllable sequence into syllable words, and 2. select the most likely character words for each syllable word. A CPIM usually assumes that the input is a complete sentence, and evaluates the performance based on a well-formed corpus. However, in practice, most Pinyin users prefer progressive text entry in several short chunks, mainly in one or two words each (most Chinese words consist of two or more characters). Short chunks do not provide enough contexts to perform the best possible syllable-to-character conversion, especially when a chunk consists of overlapping syllable words. In such cases, a conversion system often selects the boundary of a word with the highest frequency. Short chunk input is even more popular on platforms with limited computing power, such as mobile phones. Based on the observation that the relative strength of a word can be quite different when calculated leftwards or rightwards, we propose a simple division of the word context into the left context and the right context. Furthermore, we design a double ranking strategy for each word to reduce the number of errors in Step 1. Our strategy is modeled as the minimum feedback arc set problem on bipartite tournament with approximate solutions derived from genetic algorithm. Experiments show that, compared to the frequency-based method (FBM) (low memory and fast) and the conditional random fields (CRF) model (larger memory and slower), our double ranking strategy has the benefits of less memory and low power requirement with competitive performance. We believe a similar strategy could also be adopted to disambiguate conflicting linguistic patterns effectively. Mike Tian-Jian Jiang, Tsung-Hsien Lee, Wen-Lian Hsu |
ACM Trans. Asian Lang. Inf. Process. | 3 |
| 2012 | Prediction of nuclear proteins using nuclear translocation signals proposed by probabilistic latent semantic indexingabstractBACKGROUND: Identification of subcellular localization in proteins is crucial to elucidate cellular processes and molecular functions in a cell. However, given a tremendous amount of sequence data generated in the post-genomic era, determining protein localization based on biological experiments can be expensive and time-consuming. Therefore, developing prediction systems to analyze uncharacterised proteins efficiently has played an important role in high-throughput protein analyses. In a eukaryotic cell, many essential biological processes take place in the nucleus. Nuclear proteins shuttle between nucleus and cytoplasm based on recognition of nuclear translocation signals, including nuclear localization signals (NLSs) and nuclear export signals (NESs). Currently, only a few approaches have been developed specifically to predict nuclear localization using sequence features, such as putative NLSs. However, it has been shown that prediction coverage based on the NLSs is very low. In addition, most existing approaches only attained prediction accuracy and Matthew's correlation coefficient (MCC) around 54%~70% and 0.250~0.380 on independent test set, respectively. Moreover, no predictor can generate sequence motifs to characterize features of potential NESs, in which biological properties are not well understood from existing experimental studies. RESULTS: In this study, first we propose PSLNuc (Protein Subcellular Localization prediction for Nucleus) for predicting nuclear localization in proteins. First, for feature representation, a protein is represented by gapped-dipeptides and the feature values are weighted by homology information from a smoothed position-specific scoring matrix. After that, we incorporate probabilistic latent semantic indexing (PLSI) for feature reduction. Finally, the reduced features are used as input for a support vector machine (SVM) classifier. In addition to PSLNuc, we further identify gapped-dipeptide signatures for putative NLSs and NESs to develop a prediction method, PSLNTS (Protein Subcellular Localization prediction using Nuclear Translocation Signals). We apply PLSI to generate gapped-dipeptide signatures from both nuclear and non-nuclear proteins, and propose candidate sequence motifs for putative NLSs and NESs. Then, we incorporate only the proposed gapped-dipeptide signatures in an SVM classifier to mimic biological properties of NLSs and NESs for predicting nuclear localization in PSLNTS. CONCLUSIONS: Experiment results demonstrate that the proposed method shows a significant improvement for nuclear localization prediction. To compare our predictive performance with other approaches, we incorporate two non-redundant benchmark data sets, a training set and an independent test set. Evaluated by five-fold cross-validation on the training set, PSLNuc attains an overall accuracy of 79.7%, which is 4.8% improvement over the state-of-the-art system. In addition, our method also enhances the MCC from 0.497 to 0.595. Compared on the independent test set, PSLNuc outperforms other predictors by 3.9%~19.9% on accuracy and 0.077~0.207 on MCC. This suggests that, in addition to NLSs, which have been shown important for nuclear proteins, NESs can also be an effective indicator to detect non-nuclear proteins. Most notably, using only a few proposed gapped-dipeptide signatures as input features for the SVM classifier, PSLNTS further enhances the accuracy and MCC to 80.9% and 0.618, respectively. Our results demonstrate that gapped-dipeptide signatures can better discriminate nuclear and non-nuclear proteins. Moreover, the proposed gapped-dipeptide signatures can be biologically interpreted and used in further experiment analyses of nuclear translocation signals, including NLSs and NESs. Emily Chia-Yu Su, Jia-Ming Chang, Cheng-Wei Cheng, Ting-Yi Sung, Wen-Lian Hsu |
BMC Bioinform. | 5 |
| 2012 | Coreference resolution of medical concepts in discharge summaries by exploiting contextual informationabstractOBJECTIVE: Patient discharge summaries provide detailed medical information about hospitalized patients and are a rich resource of data for clinical record text mining. The textual expressions of this information are highly variable. In order to acquire a precise understanding of the patient, it is important to uncover the relationship between all instances in the text. In natural language processing (NLP), this task falls under the category of coreference resolution. DESIGN: A key contribution of this paper is the application of contextual-dependent rules that describe relationships between coreference pairs. To resolve phrases that refer to the same entity, the authors use these rules in three representative NLP systems: one rule-based, another based on the maximum entropy model, and the last a system built on the Markov logic network (MLN) model. RESULTS: The experimental results show that the proposed MLN-based system outperforms the baseline system (exact match) by average F-scores of 4.3% and 5.7% on the Beth and Partners datasets, respectively. Finally, the three systems were integrated into an ensemble system, further improving performance to 87.21%, which is 4.5% more than the official i2b2 Track 1C average (82.7%). CONCLUSION: In this paper, the main challenges in the resolution of coreference relations in patient discharge summaries are described. Several rules are proposed to exploit contextual information, and three approaches presented. While single systems provided promising results, an ensemble approach combining the three systems produced a better performance than even the best single system. Hong-Jie Dai, Johnny Chi-Yang Wu, Po-Ting Lai, Richard Tzong-Han Tsai, Wen-Lian Hsu |
J. Am. Medical Informatics Assoc. | 6 |
| 2012 | Validating Contradiction in Texts Using Online Co-Mention Pattern CheckingabstractDetecting contradictive statements is a foundational and challenging task for text understanding applications such as textual entailment. In this article, we aim to address the problem of the shortage of specific background knowledge in contradiction detection. A novel contradiction detecting approach based on the distribution of the query composed of critical mismatch combinations on the Internet is proposed to tackle the problem. By measuring the availability of mismatch conjunction phrases (MCPs), the background knowledge about two target statements can be implicitly obtained for identifying contradictions. Experiments on three different configurations show that the MCP-based approach achieves remarkable improvement on contradiction detection and can significantly improve the performance of textual entailment recognition. Cheng-Wei Shih, Chengwei Lee, Richard Tzong-Han Tsai, Wen-Lian Hsu |
ACM Trans. Asian Lang. Inf. Process. | 4 |
| 2011 | Enhancing Search Results with Semantic Annotation Using Augmented BrowsingabstractIn this paper, we describe how we integrated an artificial intelligence (AI) system into the PubMed search website using augmented browsing technology. Our system dynamically enriches the PubMed search results displayed in a user's browser with semantic annotation provided by several natural language processing (NLP) subsystems, including a sentence splitter, a part-of-speech tagger, a named entity recognizer, a section categorizer and a gene normalizer (GN). After our system is installed, the PubMed search results page is modified on the fly to categorize sections and provide additional information on gene and gene products indentified by our NLP subsystems. In addition, GN involves three main steps: candidate ID matching, false positive filtering and disambiguation, which are highly dependent on each other. We propose a joint model using a Markov logic network (MLN) to model the dependencies found in GN. The experimental results show that our joint model outperforms a baseline system that executes the three steps separately. The developed system is available at https://sites.google.com/site/pubmedannotationtool 4ijcai/home. Hong-Jie Dai, Wei-Chi Tsai, Richard Tzong-Han Tsai, Wen-Lian Hsu |
IJCAI | 4 |
| 2011 | Text Patterns and Compression Models for Semantic Class Learning
Chung-Yao Chuang, Yi-Hsun Lee, Wen-Lian Hsu |
IJCNLP | 3 |
| 2011 | Entity Disambiguation Using a Markov-Logic Network
Hong-Jie Dai, Richard Tzong-Han Tsai, Wen-Lian Hsu |
IJCNLP | 3 |
| 2011 | Evaluation via Negativa of Chinese Word Segmentation for Information Retrieval
Mike Tian-Jian Jiang, Cheng-Wei Shih, Richard Tzong-Han Tsai, Wen-Lian Hsu |
PACLIC | 4 |
| 2011 | Iteratively Estimating Pattern Reliability and Seed Quality With Extraction Consistency
Yi-Hsun Lee, Chung-Yao Chuang, Wen-Lian Hsu |
PACLIC | 3 |
| 2011 | Integration of gene normalization stages and co-reference resolution using a Markov logic networkabstractMOTIVATION: Gene normalization (GN) is the task of normalizing a textual gene mention to a unique gene database ID. Traditional top performing GN systems usually need to consider several constraints to make decisions in the normalization process, including filtering out false positives, or disambiguating an ambiguous gene mention, to improve system performance. However, these constraints are usually executed in several separate stages and cannot use each other's input/output interactively. In this article, we propose a novel approach that employs a Markov logic network (MLN) to model the constraints used in the GN task. Firstly, we show how various constraints can be formulated and combined in an MLN. Secondly, we are the first to apply the two main concepts of co-reference resolution-discourse salience in centering theory and transitivity-to GN models. Furthermore, to make our results more relevant to developers of information extraction applications, we adopt the instance-based precision/recall/F-measure (PRF) in addition to the article-wide PRF to assess system performance. RESULTS: Experimental results show that our system outperforms baseline and state-of-the-art systems under two evaluation schemes. Through further analysis, we have found several unexplored challenges in the GN task. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hong-Jie Dai, Yen-Ching Chang, Richard Tzong-Han Tsai, Wen-Lian Hsu |
Bioinform. | 4 |
| 2011 | MeInfoText 2.0: gene methylation and cancer relation extraction from biomedical literatureabstractBACKGROUND: DNA methylation is regarded as a potential biomarker in the diagnosis and treatment of cancer. The relations between aberrant gene methylation and cancer development have been identified by a number of recent scientific studies. In a previous work, we used co-occurrences to mine those associations and compiled the MeInfoText 1.0 database. To reduce the amount of manual curation and improve the accuracy of relation extraction, we have now developed MeInfoText 2.0, which uses a machine learning-based approach to extract gene methylation-cancer relations. DESCRIPTION: Two maximum entropy models are trained to predict if aberrant gene methylation is related to any type of cancer mentioned in the literature. After evaluation based on 10-fold cross-validation, the average precision/recall rates of the two models are 94.7/90.1 and 91.8/90% respectively. MeInfoText 2.0 provides the gene methylation profiles of different types of human cancer. The extracted relations with maximum probability, evidence sentences, and specific gene information are also retrievable. The database is available at http://bws.iis.sinica.edu.tw:8081/MeInfoText2/. CONCLUSION: The previous version, MeInfoText, was developed by using association rules, whereas MeInfoText 2.0 is based on a new framework that combines machine learning, dictionary lookup and pattern matching for epigenetics information extraction. The results of experiments show that MeInfoText 2.0 outperforms existing tools in many respects. To the best of our knowledge, this is the first study that uses a hybrid approach to extract gene methylation-cancer relations. It is also the first attempt to develop a gene methylation and cancer relation corpus. Yu-Ching Fang, Po-Ting Lai, Hong-Jie Dai, Wen-Lian Hsu |
BMC Bioinform. | 4 |
| 2010 | Multivariate multi-model approach for globally multimodal problemsabstractThis paper proposes an estimation of distribution algorithm (EDA) aiming at addressing globally multimodal problems, i.e., problems that present several global optima. It can be recognized that many real-world problems are of this nature, and this property generally degrades the efficiency and effectiveness of evolutionary algorithms. To overcome this source of difficulty, we designed an EDA that builds and samples multiple probabilistic models at each generation. Different from previous studies of globally multimodal problems that also use multiple models, we adopt multivariate probabilistic models. Furthermore, we have also devised a mechanism to automatically estimate the number of models that should be employed. The empirical results demonstrate that our approach obtains more global optima per run compared to the well-known EDA that employs the same class of probabilistic models but builds a single model at each generation. Moreover, the experiments also suggest that using multiple models reduces the generations spent to reach convergence. Chung-Yao Chuang, Wen-Lian Hsu |
GECCO | 2 |
| 2009 | Predicting helix-helix interactions from residue contacts in membrane proteinsabstractMOTIVATION: Helix-helix interactions play a critical role in the structure assembly, stability and function of membrane proteins. On the molecular level, the interactions are mediated by one or more residue contacts. Although previous studies focused on helix-packing patterns and sequence motifs, few of them developed methods specifically for contact prediction. RESULTS: We present a new hierarchical framework for contact prediction, with an application in membrane proteins. The hierarchical scheme consists of two levels: in the first level, contact residues are predicted from the sequence and their pairing relationships are further predicted in the second level. Statistical analyses on contact propensities are combined with other sequence and structural information for training the support vector machine classifiers. Evaluated on 52 protein chains using leave-one-out cross validation (LOOCV) and an independent test set of 14 protein chains, the two-level approach consistently improves the conventional direct approach in prediction accuracy, with 80% reduction of input for prediction. Furthermore, the predicted contacts are then used to infer interactions between pairs of helices. When at least three predicted contacts are required for an inferred interaction, the accuracy, sensitivity and specificity are 56%, 40% and 89%, respectively. Our results demonstrate that a hierarchical framework can be applied to eliminate false positives (FP) while reducing computational complexity in predicting contacts. Together with the estimated contact propensities, this method can be used to gain insights into helix-packing in membrane proteins. Allan Lo, Yi-Yuan Chiu, Einar Andreas Rødland, Ping-Chiang Lyu, Ting-Yi Sung, Wen-Lian Hsu |
Bioinform. | 6 |
| 2009 | A neural network model for constructing endophenotypes of common complex diseases: an application to male young-onset hypertension microarray dataabstractMOTIVATION: Identification of disease-related genes using high-throughput microarray data is more difficult for complex diseases as compared with monogenic ones. We hypothesized that an endophenotype derived from transcriptional data is associated with a set of genes corresponding to a pathway cluster. We assumed that a complex disease is associated with multiple endophenotypes and can be induced by their up/downregulated gene expression patterns. Thus, a neural network model was adopted to simulate the gene-endophenotype-disease relationship in which endophenotypes were represented by hidden nodes. RESULTS: We successfully constructed a three-endophenotype model for Taiwanese hypertensive males with high identification accuracy. Of the three endophenotypes, one is strongly protective, another is weakly protective and the third is highly correlated with developing young-onset male hypertension. Sixteen of the involved 101 genes were highly and consistently influential to the endophenotypes. Identification of SLC4A5, SLC5A10 and LDOC1 indicated that sodium/bicarbonate transport, sodium/glucose transport and cell-proliferation regulation may play important upstream roles and identification of BNIP1, APOBEC3F and LDOC1 suggested that apoptosis, innate immune response and cell-proliferation regulation may play important downstream roles in hypertension. The involved genes not only provide insights into the mechanism of hypertension but should also be considered in future gene mapping endeavors. Ke-Shiuan Lynn, Li-Lan Li, Yen-Ju Lin, Chiuen-Huei Wang, Shu-Hui Sheng, Ju-Hwa Lin, Wayne Liao, Wen-Lian Hsu, Wen-Harn Pan |
Bioinform. | 8 |
| 2009 | Protein subcellular localization prediction of eukaryotes using a knowledge-based approachabstractBACKGROUND: The study of protein subcellular localization (PSL) is important for elucidating protein functions involved in various cellular processes. However, determining the localization sites of a protein through wet-lab experiments can be time-consuming and labor-intensive. Thus, computational approaches become highly desirable. Most of the PSL prediction systems are established for single-localized proteins. However, a significant number of eukaryotic proteins are known to be localized into multiple subcellular organelles. Many studies have shown that proteins may simultaneously locate or move between different cellular compartments and be involved in different biological processes with different roles. RESULTS: In this study, we propose a knowledge based method, called KnowPredsite, to predict the localization site(s) of both single-localized and multi-localized proteins. Based on the local similarity, we can identify the "related sequences" for prediction. We construct a knowledge base to record the possible sequence variations for protein sequences. When predicting the localization annotation of a query protein, we search against the knowledge base and used a scoring mechanism to determine the predicted sites. We downloaded the dataset from ngLOC, which consisted of ten distinct subcellular organelles from 1923 species, and performed ten-fold cross validation experiments to evaluate KnowPred site's performance. The experiment results show that KnowPred site achieves higher prediction accuracy than ngLOC and Blast-hit method. For single-localized proteins, the overall accuracy of KnowPred site is 91.7%. For multi-localized proteins, the overall accuracy of KnowPred site is 72.1%, which is significantly higher than that of ngLOC by 12.4%. Notably, half of the proteins in the dataset that cannot find any Blast hit sequence above a specified threshold can still be correctly predicted by KnowPred site. CONCLUSION: KnowPred site demonstrates the power of identifying related sequences in the knowledge base. The experiment results show that even though the sequence similarity is low, the local similarity is effective for prediction. Experiment results show that KnowPred site is a highly accurate prediction method for both single- and multi-localized proteins. It is worth-mentioning the prediction process of KnowPred site is transparent and biologically interpretable and it shows a set of template sequences to generate the prediction result. The KnowPred site prediction server is available at http://bio-cluster.iis.sinica.edu.tw/kbloc/. Hsin-Nan Lin, Ching-Tai Chen, Ting-Yi Sung, Shinn-Ying Ho, Wen-Lian Hsu |
BMC Bioinform. | 5 |
| 2009 | HypertenGene: extracting key hypertension genes from biomedical literature with position and automatically-generated template featuresabstractBACKGROUND: The genetic factors leading to hypertension have been extensively studied, and large numbers of research papers have been published on the subject. One of hypertension researchers' primary research tasks is to locate key hypertension-related genes in abstracts. However, gathering such information with existing tools is not easy: (1) Searching for articles often returns far too many hits to browse through. (2) The search results do not highlight the hypertension-related genes discovered in the abstract. (3) Even though some text mining services mark up gene names in the abstract, the key genes investigated in a paper are still not distinguished from other genes. To facilitate the information gathering process for hypertension researchers, one solution would be to extract the key hypertension-related genes in each abstract. Three major tasks are involved in the construction of this system: (1) gene and hypertension named entity recognition, (2) section categorization, and (3) gene-hypertension relation extraction. RESULTS: We first compare the retrieval performance achieved by individually adding template features and position features to the baseline system. Then, the combination of both is examined. We found that using position features can almost double the original AUC score (0.8140 vs.0.4936) of the baseline system. However, adding template features only results in marginal improvement (0.0197). Including both improves AUC to 0.8184, indicating that these two sets of features are complementary, and do not have overlapping effects. We then examine the performance in a different domain--diabetes, and the result shows a satisfactory AUC of 0.83. CONCLUSION: Our approach successfully exploits template features to recognize true hypertension-related gene mentions and position features to distinguish key genes from other related genes. Templates are automatically generated and checked by biologists to minimize labor costs. Our approach integrates the advantages of machine learning models and pattern matching. To the best of our knowledge, this the first systematic study of extracting hypertension-related genes and the first attempt to create a hypertension-gene relation corpus based on the GAD database. Furthermore, our paper proposes and tests novel features for extracting key hypertension genes, such as relative position, section, and template features, which could also be applied to key-gene extraction for other diseases. Richard Tzong-Han Tsai, Po-Ting Lai, Hong-Jie Dai, Chi-Hsin Huang, Yue-Yang Bow, Yen-Ching Chang, Wen-Harn Pan, Wen-Lian Hsu |
BMC Bioinform. | 8 |
| 2009 | Web-based pattern learning for named entity translation in Korean-Chinese cross-language information retrieval
Yu-Chun Wang, Richard Tzong-Han Tsai, Wen-Lian Hsu |
Expert Syst. Appl. | 3 |
| 2009 | The measurement of user satisfaction with question answering systems
Chorng-Shyong Ong, Min-Yuh Day, Wen-Lian Hsu |
Inf. Manag. | 3 |
| 2009 | New Challenges for Biological Text-Mining in the Next Decade
Hong-Jie Dai, Yen-Ching Chang, Richard Tzong-Han Tsai, Wen-Lian Hsu |
J. Comput. Sci. Technol. | 4 |
| 2008 | Learning Patterns from the Web to Translate Named Entities for Cross Language Information Retrieval
Yu-Chun Wang, Richard Tzong-Han Tsai, Wen-Lian Hsu |
IJCNLP | 3 |
| 2008 | User-centered evaluation of question answering systemsabstractWith the rapid growth of the Internet and database technologies in recent years, question answering systems (QAS) have emerged as important applications. As most evaluation models focus on system-centered evaluation, user-centered evaluation has attracted little attention. Although many QAS have been implemented, little work has been done on the development of a user-centered evaluation for QAS. User-centered evaluation is used to understand a userpsilas needs and identify important dimensions and factors in the development of an information system in order to improve its acceptance. The purpose of this study is to develop a user-centered evaluation model for QAS from the userpsilas perspective. The proposed user-centered evaluation model provides a framework for the design of question answering systems from the userpsilas perspective to enhance user satisfaction and acceptance of QAS. Chorng-Shyong Ong, Min-Yuh Day, Kuo-Tay Chen, Wen-Lian Hsu |
ISI | 4 |
| 2008 | A template alignment algorithm for question classificationabstractQuestion classification (QC) plays a key role in automated question answering (QA) systems. In Chinese QC, for example, a question is analyzed and then labeled with the question type it belongs to and the expected answer type. In this paper, we propose a novel method of Chinese QC that integrates syntactic tags and semantic tags into an alignment-based approach. We adopt a template alignment (TA) algorithm to process large collections of Chinese questions and compare the classification results with those of INFOMAP, a human annotated knowledge inference engine for Chinese questions. We experimented with two approaches for the proposed system: a majority algorithm and a machine learning method that uses Support Vector Machine (SVM). The TA algorithm performs well with both approaches. The experimental results show that the accuracy achieved by TA (85.5%) is comparable to that of INFOMAP (88%). In contrast, QC based on the SVM approach, which incorporates syntactic features and TA yields an accuracy rate of 91.5%. Cheng-Lung Sung, Min-Yuh Day, Hsu-Chun Yen, Wen-Lian Hsu |
ISI | 4 |
| 2008 | Protease substrate site predictors derived from machine learning on multilevel substrate phage display dataabstractMOTIVATION: Regulatory proteases modulate proteomic dynamics with a spectrum of specificities against substrate proteins. Predictions of the substrate sites in a proteome for the proteases would facilitate understanding the biological functions of the proteases. High-throughput experiments could generate suitable datasets for machine learning to grasp complex relationships between the substrate sequences and the enzymatic specificities. But the capability in predicting protease substrate sites by integrating the machine learning algorithms with the experimental methodology has yet to be demonstrated. RESULTS: Factor Xa, a key regulatory protease in the blood coagulation system, was used as model system, for which effective substrate site predictors were developed and benchmarked. The predictors were derived from bootstrap aggregation (machine learning) algorithms trained with data obtained from multilevel substrate phage display experiments. The experimental sampling and computational learning on substrate specificities can be generalized to proteases for which the active forms are available for the in vitro experiments. AVAILABILITY: http://asqa.iis.sinica.edu.tw/fXaWeb/ Ching-Tai Chen, Ei-Wen Yang, Hung-Ju Hsu, Yi-Kun Sun, Wen-Lian Hsu, An-Suei Yang |
Bioinform. | 5 |
| 2008 | Predicting RNA-binding sites of proteins using support vector machines and evolutionary informationabstractBACKGROUND: RNA-protein interaction plays an essential role in several biological processes, such as protein synthesis, gene expression, posttranscriptional regulation and viral infectivity. Identification of RNA-binding sites in proteins provides valuable insights for biologists. However, experimental determination of RNA-protein interaction remains time-consuming and labor-intensive. Thus, computational approaches for prediction of RNA-binding sites in proteins have become highly desirable. Extensive studies of RNA-binding site prediction have led to the development of several methods. However, they could yield low sensitivities in trade-off for high specificities. RESULTS: We propose a method, RNAProB, which incorporates a new smoothed position-specific scoring matrix (PSSM) encoding scheme with a support vector machine model to predict RNA-binding sites in proteins. Besides the incorporation of evolutionary information from standard PSSM profiles, the proposed smoothed PSSM encoding scheme also considers the correlation and dependency from the neighboring residues for each amino acid in a protein. Experimental results show that smoothed PSSM encoding significantly enhances the prediction performance, especially for sensitivity. Using five-fold cross-validation, our method performs better than the state-of-the-art systems by 4.90%-6.83%, 0.88%-5.33%, and 0.10-0.23 in terms of overall accuracy, specificity, and Matthew's correlation coefficient, respectively. Most notably, compared to other approaches, RNAProB significantly improves sensitivity by 7.0%-26.9% over the benchmark data sets. To prevent data over fitting, a three-way data split procedure is incorporated to estimate the prediction performance. Moreover, physicochemical properties and amino acid preferences of RNA-binding proteins are examined and analyzed. CONCLUSION: Our results demonstrate that smoothed PSSM encoding scheme significantly enhances the performance of RNA-binding site prediction in proteins. This also supports our assumption that smoothed PSSM encoding can better resolve the ambiguity of discriminating between interacting and non-interacting residues by modelling the dependency from surrounding residues. The proposed method can be used in other research areas, such as DNA-binding site prediction, protein-protein interaction, and prediction of posttranslational modification sites. Cheng-Wei Cheng, Emily Chia-Yu Su, Jenn-Kang Hwang, Ting-Yi Sung, Wen-Lian Hsu |
BMC Bioinform. | 5 |
| 2008 | Emerging strengths in Asia Pacific bioinformaticsabstractThe 2008 annual conference of the Asia Pacific Bioinformatics Network (APBioNet), Asia's oldest bioinformatics organisation set up in 1998, was organized as the 7th International Conference on Bioinformatics (InCoB), jointly with the Bioinformatics and Systems Biology in Taiwan (BIT 2008) Conference, Oct. 20-23, 2008 at Taipei, Taiwan. Besides bringing together scientists from the field of bioinformatics in this region, InCoB is actively involving researchers from the area of systems biology, to facilitate greater synergy between these two groups. Marking the 10th Anniversary of APBioNet, this InCoB 2008 meeting followed on from a series of successful annual events in Bangkok (Thailand), Penang (Malaysia), Auckland (New Zealand), Busan (South Korea), New Delhi (India) and Hong Kong. Additionally, tutorials and the Workshop on Education in Bioinformatics and Computational Biology (WEBCB) immediately prior to the 20th Federation of Asian and Oceanian Biochemists and Molecular Biologists (FAOBMB) Taipei Conference provided ample opportunity for inducting mainstream biochemists and molecular biologists from the region into a greater level of awareness of the importance of bioinformatics in their craft. In this editorial, we provide a brief overview of the peer-reviewed manuscripts accepted for publication herein, grouped into thematic areas. As the regional research expertise in bioinformatics matures, the papers fall into thematic areas, illustrating the specific contributions made by APBioNet to global bioinformatics efforts. Shoba Ranganathan, Wen-Lian Hsu, Ueng-Cheng Yang, Tin Wee Tan |
BMC Bioinform. | 2 |
| 2008 | Semi-automatic conversion of BioProp semantic annotation to PASBio annotationabstractBACKGROUND: Semantic role labeling (SRL) is an important text analysis technique. In SRL, sentences are represented by one or more predicate-argument structures (PAS). Each PAS is composed of a predicate (verb) and several arguments (noun phrases, adverbial phrases, etc.) with different semantic roles, including main arguments (agent or patient) as well as adjunct arguments (time, manner, or location). PropBank is the most widely used PAS corpus and annotation format in the newswire domain. In the biomedical field, however, more detailed and restrictive PAS annotation formats such as PASBio are popular. Unfortunately, due to the lack of an annotated PASBio corpus, no publicly available machine-learning (ML) based SRL systems based on PASBio have been developed. In previous work, we constructed a biomedical corpus based on the PropBank standard called BioProp, on which we developed an ML-based SRL system, BIOSMILE. In this paper, we aim to build a system to convert BIOSMILE's BioProp annotation output to PASBio annotation. Our system consists of BIOSMILE in combination with a BioProp-PASBio rule-based converter, and an additional semi-automatic rule generator. RESULTS: Our first experiment evaluated our rule-based converter's performance independently from BIOSMILE performance. The converter achieved an F-score of 85.29%. The second experiment evaluated combined system (BIOSMILE + rule-based converter). The system achieved an F-score of 69.08% for PASBio's 29 verbs. CONCLUSION: Our approach allows PAS conversion between BioProp and PASBio annotation using BIOSMILE alongside our newly developed semi-automatic rule generator and rule-based converter. Our system can match the performance of other state-of-the-art domain-specific ML-based SRL systems and can be easily customized for PASBio application development. Richard Tzong-Han Tsai, Hong-Jie Dai, Chi-Hsin Huang, Wen-Lian Hsu |
BMC Bioinform. | 4 |
| 2008 | Exploiting likely-positive and unlabeled data to improve the identification of protein-protein interaction articlesabstractBACKGROUND: Experimentally verified protein-protein interactions (PPI) cannot be easily retrieved by researchers unless they are stored in PPI databases. The curation of such databases can be made faster by ranking newly-published articles' relevance to PPI, a task which we approach here by designing a machine-learning-based PPI classifier. All classifiers require labeled data, and the more labeled data available, the more reliable they become. Although many PPI databases with large numbers of labeled articles are available, incorporating these databases into the base training data may actually reduce classification performance since the supplementary databases may not annotate exactly the same PPI types as the base training data. Our first goal in this paper is to find a method of selecting likely positive data from such supplementary databases. Only extracting likely positive data, however, will bias the classification model unless sufficient negative data is also added. Unfortunately, negative data is very hard to obtain because there are no resources that compile such information. Therefore, our second aim is to select such negative data from unlabeled PubMed data. Thirdly, we explore how to exploit these likely positive and negative data. And lastly, we look at the somewhat unrelated question of which term-weighting scheme is most effective for identifying PPI-related articles. RESULTS: To evaluate the performance of our PPI text classifier, we conducted experiments based on the BioCreAtIvE-II IAS dataset. Our results show that adding likely-labeled data generally increases AUC by 3~6%, indicating better ranking ability. Our experiments also show that our newly-proposed term-weighting scheme has the highest AUC among all common weighting schemes. Our final model achieves an F-measure and AUC 2.9% and 5.0% higher than those of the top-ranking system in the IAS challenge. CONCLUSION: Our experiments demonstrate the effectiveness of integrating unlabeled and likely labeled data to augment a PPI text classification system. Our mixed model is suitable for ranking purposes whereas our hierarchical model is better for filtering. In addition, our results indicate that supervised weighting schemes outperform unsupervised ones. Our newly-proposed weighting scheme, TFBRF, which considers documents that do not contain the target word, avoids some of the biases found in traditional weighting schemes. Our experiment results show TFBRF to be the most effective among several other top weighting schemes. Richard Tzong-Han Tsai, Hsi-Chuan Hung, Hong-Jie Dai, Jaimie Yi-Wen Lin, Wen-Lian Hsu |
BMC Bioinform. | 5 |
| 2008 | Web taxonomy integration with hierarchical shrinkage algorithm and fine-grained relations
Chia-Wei Wu, Richard Tzong-Han Tsai, Cheng-Wei Lee 0001, Wen-Lian Hsu |
Expert Syst. Appl. | 4 |
| 2008 | Boosting Chinese Question Answering with Two Lightweight Methods: ABSPs and SCO-QATabstractQuestion Answering (QA) research has been conducted in many languages. Nearly all the top performing systems use heavy methods that require sophisticated techniques, such as parsers or logic provers. However, such techniques are usually unavailable or unaffordable for under-resourced languages or in resource-limited situations. In this article, we describe how a top-performing Chinese QA system can be designed by using lightweight methods effectively. We propose two lightweight methods, namely the Sum of Co-occurrences of Question and Answer Terms (SCO-QAT) and Alignment-based Surface Patterns (ABSPs). SCO-QAT is a co-occurrence-based answer-ranking method that does not need extra knowledge, word-ignoring heuristic rules, or tools. It calculates co-occurrence scores based on the passage retrieval results. ABSPs are syntactic patterns trained from question-answer pairs with a multiple alignment algorithm. They are used to capture the relations between terms and then use the relations to filter answers. We attribute the success of the ABSPs and SCO-QAT methods to the effective use of local syntactic information and global co-occurrence information. By using SCO-QAT and ABSPs, we improved the RU-Accuracy of our testbed QA system, ASQA, from 0.445 to 0.535 on the NTCIR-5 dataset. It also achieved the top 0.5 RU-Accuracy on the NTCIR-6 dataset. The result shows that lightweight methods are not only cheaper to implement, but also have the potential to achieve state-of-the-art performances. Cheng-Wei Lee 0001, Min-Yuh Day, Cheng-Lung Sung, Yi-Hsun Lee, Mike Tian-Jian Jiang, Chia-Wei Wu, Cheng-Wei Shih, Yu-Ren Chen, Wen-Lian Hsu |
ACM Trans. Asian Lang. Inf. Process. | 9 |
| 2007 | Identifying Protein Interaction Abstracts with Contextual Bag of Words
Hsieh-Chuan Hung, Richard Tzong-Han Tsai, Wen-Lian Hsu |
AAAI | 3 |
| 2007 | Exploiting unlabeled internal data in conditional random fields to reduce word segmentation errors for Chinese textsabstractThe application of text-to-speech (TTS) conversion has become widely used in recent years. Chinese TTS faces several unique difficulties. The most critical is caused by the lack of word delimiters in written Chinese. This means that Chinese word segmentation (CWS) must be the first step in Chinese TTS. Unfortunately, due to the ambiguous nature of word boundaries in Chinese, even the best CWS systems make serious segmentation errors. Incorrect sentence interpretation causes TTS errors, preventing TTS’s wider use in applications such as automatic customer services or computer reader systems for the visually impaired. In this paper, we propose a novel method that exploits unlabeled internal data to reduce word segmentation errors without using external dictionaries. To demonstrate the generality of our method, we verify our system on the most widely recognized CWS evaluation tool--the SIGHAN bakeoff, which includes datasets in both traditional and simplified Chinese. These datasets are provided by four representative academies or industrial research institutes in HK, Taiwan, Mainland China, and the U.S. Our experimental results show that with only internal data and unlabeled test data, our approach reduces segmentation errors by an average of 15 % compared to the traditional approach. Moreover, our approach achieves comparable performance to the best CWS systems that use external resources. Further analysis shows that our method has the potential to become more accurate as the amount of test data increases. Index Terms: text-to-speech, Chinese word segmentation, segmentation errors, internal unlabeled data Richard Tzong-Han Tsai, Hsi-Chuan Hung, Hong-Jie Dai, Wen-Lian Hsu |
INTERSPEECH | 4 |
| 2007 | Korean-Chinese Person Name Translation for Cross Language Information Retrieval
Yu-Chun Wang, Yi-Hsun Lee, Chu-Cheng Lin, Richard Tzong-Han Tsai, Wen-Lian Hsu |
PACLIC | 5 |
| 2007 | Detection of the inferred interaction network in hepatocellular carcinoma from EHCO (Encyclopedia of Hepatocellular Carcinoma genes Online)abstractBACKGROUND: The significant advances in microarray and proteomics analyses have resulted in an exponential increase in potential new targets and have promised to shed light on the identification of disease markers and cellular pathways. We aim to collect and decipher the HCC-related genes at the systems level. RESULTS: Here, we build an integrative platform, the Encyclopedia of Hepatocellular Carcinoma genes Online, dubbed EHCO http://ehco.iis.sinica.edu.tw, to systematically collect, organize and compare the pileup of unsorted HCC-related studies by using natural language processing and softbots. Among the eight gene set collections, ranging across PubMed, SAGE, microarray, and proteomics data, there are 2,906 genes in total; however, more than 77% genes are only included once, suggesting that tremendous efforts need to be exerted to characterize the relationship between HCC and these genes. Of these HCC inventories, protein binding represents the largest proportion (~25%) from Gene Ontology analysis. In fact, many differentially expressed gene sets in EHCO could form interaction networks (e.g. HBV-associated HCC network) by using available human protein-protein interaction datasets. To further highlight the potential new targets in the inferred network from EHCO, we combine comparative genomics and interactomics approaches to analyze 120 evolutionary conserved and overexpressed genes in HCC. 47 out of 120 queries can form a highly interactive network with 18 queries serving as hubs. CONCLUSION: This architectural map may represent the first step toward the attempt to decipher the hepatocarcinogenesis at the systems level. Targeting hubs and/or disruption of the network formation might reveal novel strategy for HCC treatment. Chun-Nan Hsu, Chia-Hung Liu, Huei-Hun Tseng, Chih-Yun Lin, Kuan-Ting Lin, Hsu-Hua Yeh, Ting-Yi Sung, Wen-Lian Hsu, Li-Jen Su, Sheng-An Lee, Chang-Han Chen, Gen-Cher Lee, D. T. Lee, Yow-Ling Shiue, Chang-Wei Yeh, Chao-Hui Chang, Cheng-Yan Kao, Chi-Ying F. Huang |
BMC Bioinform. | 9 |
| 2007 | Protein subcellular localization prediction based on compartment-specific features and structure conservationabstractBACKGROUND: Protein subcellular localization is crucial for genome annotation, protein function prediction, and drug discovery. Determination of subcellular localization using experimental approaches is time-consuming; thus, computational approaches become highly desirable. Extensive studies of localization prediction have led to the development of several methods including composition-based and homology-based methods. However, their performance might be significantly degraded if homologous sequences are not detected. Moreover, methods that integrate various features could suffer from the problem of low coverage in high-throughput proteomic analyses due to the lack of information to characterize unknown proteins. RESULTS: We propose a hybrid prediction method for Gram-negative bacteria that combines a one-versus-one support vector machines (SVM) model and a structural homology approach. The SVM model comprises a number of binary classifiers, in which biological features derived from Gram-negative bacteria translocation pathways are incorporated. In the structural homology approach, we employ secondary structure alignment for structural similarity comparison and assign the known localization of the top-ranked protein as the predicted localization of a query protein. The hybrid method achieves overall accuracy of 93.7% and 93.2% using ten-fold cross-validation on the benchmark data sets. In the assessment of the evaluation data sets, our method also attains accurate prediction accuracy of 84.0%, especially when testing on sequences with a low level of homology to the training data. A three-way data split procedure is also incorporated to prevent overestimation of the predictive performance. In addition, we show that the prediction accuracy should be approximately 85% for non-redundant data sets of sequence identity less than 30%. CONCLUSION: Our results demonstrate that biological features derived from Gram-negative bacteria translocation pathways yield a significant improvement. The biological features are interpretable and can be applied in advanced analyses and experimental designs. Moreover, the overall accuracy of combining the structural homology approach is further improved, which suggests that structural conservation could be a useful indicator for inferring localization in addition to sequence homology. The proposed method can be used in large-scale analyses of proteomes. Emily Chia-Yu Su, Hua-Sheng Chiu, Allan Lo, Jenn-Kang Hwang, Ting-Yi Sung, Wen-Lian Hsu |
BMC Bioinform. | 6 |
| 2007 | BIOSMILE: A semantic role labeling system for biomedical verbs using a maximum-entropy model with automatically generated template featuresabstractBACKGROUND: Bioinformatics tools for automatic processing of biomedical literature are invaluable for both the design and interpretation of large-scale experiments. Many information extraction (IE) systems that incorporate natural language processing (NLP) techniques have thus been developed for use in the biomedical field. A key IE task in this field is the extraction of biomedical relations, such as protein-protein and gene-disease interactions. However, most biomedical relation extraction systems usually ignore adverbial and prepositional phrases and words identifying location, manner, timing, and condition, which are essential for describing biomedical relations. Semantic role labeling (SRL) is a natural language processing technique that identifies the semantic roles of these words or phrases in sentences and expresses them as predicate-argument structures. We construct a biomedical SRL system called BIOSMILE that uses a maximum entropy (ME) machine-learning model to extract biomedical relations. BIOSMILE is trained on BioProp, our semi-automatic, annotated biomedical proposition bank. Currently, we are focusing on 30 biomedical verbs that are frequently used or considered important for describing molecular events. RESULTS: To evaluate the performance of BIOSMILE, we conducted two experiments to (1) compare the performance of SRL systems trained on newswire and biomedical corpora; and (2) examine the effects of using biomedical-specific features. The experimental results show that using BioProp improves the F-score of the SRL system by 21.45% over an SRL system that uses a newswire corpus. It is noteworthy that adding automatically generated template features improves the overall F-score by a further 0.52%. Specifically, ArgM-LOC, ArgM-MNR, and Arg2 achieve statistically significant performance improvements of 3.33%, 2.27%, and 1.44%, respectively. CONCLUSION: We demonstrate the necessity of using a biomedical proposition bank for training SRL systems in the biomedical domain. Besides the different characteristics of biomedical and newswire sentences, factors such as cross-domain framesets and verb usage variations also influence the performance of SRL systems. For argument classification, we find that NE (named entity) features indicating if the target node matches with NEs are not effective, since NEs may match with a node of the parsing tree that does not have semantic role labels in the training set. We therefore incorporate templates composed of specific words, NE types, and POS tags into the SRL system. As a result, the classification accuracy for adjunct arguments, which is especially important for biomedical SRL, is improved significantly. Richard Tzong-Han Tsai, Wen-Chi Chou, Ying-Shan Su, Yu-Chun Lin, Cheng-Lung Sung, Hong-Jie Dai, Irene Tzu-Hsuan Yeh, Wei Ku, Ting-Yi Sung, Wen-Lian Hsu |
BMC Bioinform. | 10 |
| 2007 | Reference metadata extraction using a hierarchical knowledge representation framework
Min-Yuh Day, Richard Tzong-Han Tsai, Cheng-Lung Sung, Chiu-Chen Hsieh, Cheng-Wei Lee 0001, Shih-Hung Wu, Kuen-Pin Wu, Chorng-Shyong Ong, Wen-Lian Hsu |
Decis. Support Syst. | 9 |
| 2006 | A Knowledge-Based Approach to Protein Local Structure Prediction
Ching-Tai Chen, Hsin-Nan Lin, Kuen-Pin Wu, Ting-Yi Sung, Wen-Lian Hsu |
APBC | 5 |
| 2006 | Designing a Tutoring Agent for Facilitating Collaborative Learning with Instant Messaging
Sheng-Cheng Hsu, Min-Yuh Day, Shih-Hung Wu, Wing-Kwong Wong, Wen-Lian Hsu |
Intelligent Tutoring Systems | 5 |
| 2006 | Using Instant Messaging to Provide an Intelligent Learning Environment
Chun-Hung Lu, Guey-Fa Chiou, Min-Yuh Day, Chorng-Shyong Ong, Wen-Lian Hsu |
Intelligent Tutoring Systems | 5 |
| 2006 | Web Directory Integration Using Conditional Random FieldsabstractThe purpose of integrating web directories is to transfer instances from a source to a target directory. Unlike con-ventional text categorization, in directory integration, there is extra information about the source directory that can be used to improve the classification accuracy. Many approaches exploit the measured similarity between two corresponding classes to enhance traditional text classifi-ers. These methods perform well if the topics of two classes are very similar, but they could lead to misclassifi-cation if the topics are dissimilar. We propose a directory integration approach based on the conditional random fields (CRFs) model, and model the integration process using a finite-state model. The advantage of using CRFs is that the transition features naturally include information about the relations between classes. Our results show that CRFs outperform conven-tional text classifiers. In addition, CRFs allow us to apply complex features to integrate the information about the contents of class and their labels. The performance of our approach can be improved by applying these features, especially for instances whose source and target classes are moderately similar. Chia-Wei Wu, Wen-Lian Hsu |
Web Intelligence | 2 |
| 2006 | NERBio: using selected word conjunctions, term normalization, and global patterns to improve biomedical named entity recognitionabstractBACKGROUND: Biomedical named entity recognition (Bio-NER) is a challenging problem because, in general, biomedical named entities of the same category (e.g., proteins and genes) do not follow one standard nomenclature. They have many irregularities and sometimes appear in ambiguous contexts. In recent years, machine-learning (ML) approaches have become increasingly common and now represent the cutting edge of Bio-NER technology. This paper addresses three problems faced by ML-based Bio-NER systems. First, most ML approaches usually employ singleton features that comprise one linguistic property (e.g., the current word is capitalized) and at least one class tag (e.g., B-protein, the beginning of a protein name). However, such features may be insufficient in cases where multiple properties must be considered. Adding conjunction features that contain multiple properties can be beneficial, but it would be infeasible to include all conjunction features in an NER model since memory resources are limited and some features are ineffective. To resolve the problem, we use a sequential forward search algorithm to select an effective set of features. Second, variations in the numerical parts of biomedical terms (e.g., "2" in the biomedical term IL2) cause data sparseness and generate many redundant features. In this case, we apply numerical normalization, which solves the problem by replacing all numerals in a term with one representative numeral to help classify named entities. Third, the assignment of NE tags does not depend solely on the target word's closest neighbors, but may depend on words outside the context window (e.g., a context window of five consists of the current word plus two preceding and two subsequent words). We use global patterns generated by the Smith-Waterman local alignment algorithm to identify such structures and modify the results of our ML-based tagger. This is called pattern-based post-processing. RESULTS: To develop our ML-based Bio-NER system, we employ conditional random fields, which have performed effectively in several well-known tasks, as our underlying ML model. Adding selected conjunction features, applying numerical normalization, and employing pattern-based post-processing improve the F-scores by 1.67%, 1.04%, and 0.57%, respectively. The combined increase of 3.28% yields a total score of 72.98%, which is better than the baseline system that only uses singleton features. CONCLUSION: We demonstrate the benefits of using the sequential forward search algorithm to select effective conjunction feature groups. In addition, we show that numerical normalization can effectively reduce the number of redundant and unseen features. Furthermore, the Smith-Waterman local alignment algorithm can help ML-based Bio-NER deal with difficult cases that need longer context windows. Richard Tzong-Han Tsai, Cheng-Lung Sung, Hong-Jie Dai, Hsieh-Chuan Hung, Ting-Yi Sung, Wen-Lian Hsu |
BMC Bioinform. | 6 |
| 2006 | Various criteria in the evaluation of biomedical named entity recognitionabstractBACKGROUND: Text mining in the biomedical domain is receiving increasing attention. A key component of this process is named entity recognition (NER). Generally speaking, two annotated corpora, GENIA and GENETAG, are most frequently used for training and testing biomedical named entity recognition (Bio-NER) systems. JNLPBA and BioCreAtIvE are two major Bio-NER tasks using these corpora. Both tasks take different approaches to corpus annotation and use different matching criteria to evaluate system performance. This paper details these differences and describes alternative criteria. We then examine the impact of different criteria and annotation schemes on system performance by retesting systems participated in the above two tasks. RESULTS: To analyze the difference between JNLPBA's and BioCreAtIvE's evaluation, we conduct Experiment 1 to evaluate the top four JNLPBA systems using BioCreAtIvE's classification scheme. We then compare them with the top four BioCreAtIvE systems. Among them, three systems participated in both tasks, and each has an F-score lower on JNLPBA than on BioCreAtIvE. In Experiment 2, we apply hypothesis testing and correlation coefficient to find alternatives to BioCreAtIvE's evaluation scheme. It shows that right-match and left-match criteria have no significant difference with BioCreAtIvE. In Experiment 3, we propose a customized relaxed-match criterion that uses right match and merges JNLPBA's five NE classes into two, which achieves an F-score of 81.5%. In Experiment 4, we evaluate a range of five matching criteria from loose to strict on the top JNLPBA system and examine the percentage of false negatives. Our experiment gives the relative change in precision, recall and F-score as matching criteria are relaxed. CONCLUSION: In many applications, biomedical NEs could have several acceptable tags, which might just differ in their left or right boundaries. However, most corpora annotate only one of them. In our experiment, we found that right match and left match can be appropriate alternatives to JNLPBA and BioCreAtIvE's matching criteria. In addition, our relaxed-match criterion demonstrates that users can define their own relaxed criteria that correspond more realistically to their application requirements. Richard Tzong-Han Tsai, Shih-Hung Wu, Wen-Chi Chou, Yu-Chun Lin, Jieh Hsiang, Ting-Yi Sung, Wen-Lian Hsu |
BMC Bioinform. | 8 |
| 2006 | Integrating linguistic knowledge into a conditional random fieldframework to identify biomedical named entities
Richard Tzong-Han Tsai, Wen-Chi Chou, Shih-Hung Wu, Ting-Yi Sung, Jieh Hsiang, Wen-Lian Hsu |
Expert Syst. Appl. | 6 |
| 2005 | A Linear Time Algorithm for Finding a Maximal Planar Subgraph Based on PC-Trees
Wen-Lian Hsu |
COCOON | 1 |
| 2005 | Exploiting Full Parsing Information to Label Semantic Roles Using an Ensemble of ME and SVM via Integer Linear Programming
Richard Tzong-Han Tsai, Chia-Wei Wu, Yu-Chun Lin, Wen-Lian Hsu |
CoNLL | 4 |
| 2005 | Designing an Ontology-Based Intelligent Tutoring Agent with Instant MessagingabstractThe rapid growth of the Internet and instant messaging (IM) offers new opportunities as well as challenges to both educators and students. In this paper, we propose an intelligent tutoring agent (ITA) that uses the ontology, INFOMAP, and question answering techniques through the instant messaging platform for the "operating system " course. The ITA embeds the above techniques in the teaching process and plays the role of a tutoring agent to help a teacher track, record, and understand a student's status. The ITA interface can interpret natural language to facilitate communication between the student and the tutor. Student can query and learn the concept of "operating system" through MSN Messenger, which ITA adopts as the communication protocol. The proposed ITA is accessible by adding the contact ID: [email protected] to the MSN Messenger contacts list. Min-Yuh Day, Chun-Hung Lu, Jin-Tan Yang, Guey-Fa Chiou, Chorng-Shyong Ong, Wen-Lian Hsu |
ICALT | 6 |
| 2005 | The Design of a Diagnosis System for Problem PosingabstractIn this paper, a diagnosis system for problem posing is developed. The domain of this system is restricted to area word problems. The student poses the problems through the system with phrase combination approach. Then the system diagnoses the problem and provides the information for students to revise the posed problem, if needed. The system uses an ontology-based knowledge engineering tool, InfoMap, to represent the concept knowledge for diagnosing area word problems posed by students. Sheng-Cheng Hsu, Shih-Hung Wu, Wing-Kwong Wong, Hsi-Hsun Yang, Wen-Lian Hsu |
ICALT | 5 |
| 2005 | An Empirical Exploration of Using Wiki in an English as a Second Language CourseabstractIn this paper, we present an empirical study of using a new and cost-effective Web-based collaboration software, Wiki, in a freshman-level English as a second language (ESL) course. This paper explores and observes the scenario: what if the Wiki tool were to be used in an English as a second language course in Taiwan? Students who attended this study practiced English writing on a Wiki Web site. The data about their usage and learning achievements was collected and analyzed. Our finding of a significant, but inverse, relation between students' editing usage and academic performance challenges some idealistic hypotheses that Wiki technology is "naturally beneficial" to learning. We believe that building an instructive or constructive instructional model with Wiki in a rigorous manner requires more empirical evidence. This study provides fresh evidence that will hopefully serve as an impetus to fill that gap. Hao-Chuan Wang, Chun-Hung Lu, Jun-Yi Yang, Hsin-Wen Hu, Guey-Fa Chiou, Yueh-Tzu Chiang, Wen-Lian Hsu |
ICALT | 7 |
| 2005 | RIBRA-An Error-Tolerant Algorithm for the NMR Backbone Assignment Problem
Kuen-Pin Wu, Jia-Ming Chang, Jun-Bo Chen, Chi-Fon Chang, Wen-Jin Wu, Tai-Huang Huang, Ting-Yi Sung, Wen-Lian Hsu |
RECOMB | 8 |
| 2005 | HYPROSP II-A knowledge-based hybrid method for protein secondary structure prediction based on local prediction confidenceabstractMOTIVATION: In our previous approach, we proposed a hybrid method for protein secondary structure prediction called HYPROSP, which combined our proposed knowledge-based prediction algorithm PROSP and PSIPRED. The knowledge base constructed for PROSP contains small peptides together with their secondary structural information. The hybrid strategy of HYPROSP uses a global quantitative measure, match rate, to determine whether PROSP or PSIPRED is to be used for the prediction of a target protein. HYPROSP made slight improvement of Q(3) over PSIPRED because PROSP predicted well for proteins with match rate >80%. As the portion of proteins with match rate >80% is quite small and as the performance of PSIPRED also improves, the advantage of HYPROSP is diluted. To overcome this limitation and further improve the hybrid prediction method, we present in this paper a new hybrid strategy HYPROSP II that is based on a new quantitative measure called local match rate. RESULTS: Local match rate indicates the amount of structural information that each amino acid can extract from the knowledge base. With the local match rate, we are able to define a confidence level of the PROSP prediction results for each amino acid. Our new hybrid approach, HYPROSP II, is proposed as follows: for each amino acid in a target protein, we combine the prediction results of PROSP and PSIPRED using a hybrid function defined on their respective confidence levels. Two datasets in nrDSSP and EVA are used to perform a 10-fold cross validation. The average Q(3) of HYPROSP II is 81.8% and 80.7% on nrDSSP and EVA datasets, respectively, which is 2.0% and 1.1% better than that of PSIPRED. For local structures with match rate >80%, the average Q(3) improvement is 4.4% on the nrDSSP dataset. The use of local match rate improves the accuracy better than global match rate. There has been a long history of attempts to improve secondary structure prediction. We believe that HYPROSP II has greatly utilized the power of peptide knowledge base and raised the prediction accuracy to a new high. The method we developed in this paper could have a profound effect on the general use of knowledge base techniques for various predictionalgorithms. AVAILABILITY: The Linux executable file of HYPROSP II, as well as both nrDSSP and EVA datasets can be downloaded from http://bioinformatics.iis.sinica.edu.tw/HYPROSPII/. Hsin-Nan Lin, Jia-Ming Chang, Kuen-Pin Wu, Ting-Yi Sung, Wen-Lian Hsu |
Bioinform. | 5 |
| 2004 | An Iterative Relaxation Technique for the NMR Backbone Assignment ProblemabstractNMR spectroscopy is one of the popular experiments to determine protein structures. An important stage of protein structure determination by using NMR is protein backbone resonance assignment (or backbone assignment for short). Due to the messiness and disorder of NMR spectral data, backbone assignment is usually a tedious and time-consuming manual work. This raises a great interest in developing an efficient and automatic method to perform backbone assignment. An iterative algorithm is proposed that is equipped with two operations: grouping and linking. Grouping is responsible for peak picking and part of connectivity determination. Ideally, those peaks with the same H/sup N/ and N chemical shifts can be grouped together; peaks belonging to the same group can be used to determine the order of two consecutive spin systems. However, in real situation, grouping is a difficult task due to false positives and false negatives. We sometimes add hypothetic peaks to tackle false negatives and use linking operation to remove false positives. Linking is responsible for part of connectivity determination and backbone assignment. Given a protein sequence and partial connectivity information, we try to link connected spin systems as much as possible. False positive spin systems may create conflicts in the linking stage. To find a good assignment on a noisy dataset, the backbone assignment problem is modeled as a maximum independent set problem. Although the problem is NP-complete, there are heuristic methods that obtain pretty good results. Wen-Lian Hsu, Jia-Ming Chang, Wen-Chi Chou, Jun-Bo Chen, Kuen-Pin Wu, Ting-Yi Sung, Chi-Fon Chang, Wen-Jin Wu, Tai-Huang Huang |
BIBE | 1 |
| 2004 | The Design of An Intelligent Tutoring System Based on the Ontology of Procedural KnowledgeabstractThis paper presents a new model to simulate procedural knowledge. This method divides procedural knowledge into two parts: process control and action performer. By adopting this method, we intend to help teacher construct curriculum and teaching strategies by capturing students' problem-solving process. Using the concept of procedural knowledge in intelligent tutoring systems, we can accumulate and duplicate the knowledge of teacher/curriculum manager and student model. Teacher/curriculum manager can help the teacher create good learning maps for students. The student model can help the teacher collect students' error types and design more appropriate teaching strategies. The implementation of our system is near completion. We will design a user-friendly interface for the system and have students play with this software to collect feedback. Chun-Hung Lu, Shih-Hung Wu, LiongYu Tu, Wen-Lian Hsu |
ICALT | 4 |
| 2003 | PC trees and circular-ones arrangements
Wen-Lian Hsu, Ross M. McConnell |
Theor. Comput. Sci. | 1 |
| 2002 | Applying an NVEF Word-Pair Identifier to the Chinese Syllable-to-Word Conversion Problem
Jia-Lin Tsai, Wen-Lian Hsu |
COLING | 2 |
| 2002 | SOAT: A Semi-Automatic Domain Ontology Acquisition Tool from Chinese Corpus
Shih-Hung Wu, Wen-Lian Hsu |
COLING | 2 |
| 2002 | Exploiting Knowledge Representation in an Intelligent Tutoring System for English Lexical ErrorsabstractIntelligent tutoring systems (ITSs) construction requires lots of domain knowledge created by hand. In this paper we attempt to illustrate a central role that knowledge representation can play in automating ITS design and implementation. We propose a diagnosis, interaction and treatment (DIT) model for ITS. The entire system relies upon a knowledge representation system (InfoMap) whose structured encoding of English lexical information makes it possible to (1) initiate the relevant dialogue with learners when the system is not sure about learners' intentions, (2) trigger the appropriate lexical knowledge based on their responses, and (3) automatically generate practice and test exercises based on this knowledge. Chiu-Chen Hsieh, Richard Tzong-Han Tsai, David Wible, Wen-Lian Hsu |
ICCE | 4 |
| 2002 | NTUs: An Intelligent Tutorial System Fosters Number Concepts through Computational ScaffoldingabstractThe present article describes how scaffolding is implemented in an intelligent tutorial system call NTUs (number transcoding tutorial system), and what are the results of the system tested empirically on a group of grade students in fostering their number concepts. To use NTUs the system first analyzes a user's errors on a number transcoding task, and the results of analysis are used to infer the user's zone of proximal development (ZPD). A scaffolding process which vas designed with inspiration from how people learn Chinese calligraphy is provided next in the user's ZPD. An empirical test indicates that NTUs can not only foster students' number concepts, but can also attract them to use it. Chih-Wei Hue, Chien-Huei Kao, Ming Lo, Chien-Chih Chiang, LiongYu Tu, Wen-Lian Hsu |
ICCE | 6 |
| 2002 | A Cognitive Student Model - An Ontological ApproachabstractWe present an ontological approach to the design of the student model for a tutorial agent system (TAS). Our model emphasizes the classification and detection of error types. If the student has any systematic and predictable misconceptions, the system attempts to determine the underlying reasons for such errors. We adopt the "identification, simulation, interaction, and mapping" (ISIM) strategy to achieve this goal. The tutorial agent system first identifies which problem solving method a student is using. It then simulates the procedure in a step-by-step fashion. If there is any ambiguity in the diagnosis of error types during the simulation, the system will interact with the student to resolve it. Finally, the interaction will lead to appropriate error types. The related knowledge is constructed in an ontological framework, InfoMap. In this paper, we focus on how to construct the knowledge and how the simulation works. LiongYu Tu, Wen-Lian Hsu, Shih-Hung Wu |
ICCE | 2 |
| 2001 | PC-Trees vs. PQ-Trees
Wen-Lian Hsu |
COCOON | 1 |
| 2001 | Event identification based on the information map-INFOMAPabstractWe present a knowledge representation scheme, INFOMAP, together with a mechanism that matches the event of a natural language sentence with part of the domain ontology in the INFOMAP The design of this scheme is to facilitate both human browsing and computer processing of the domain ontology. INFOMAP is also a knowledge framework designed to facilitate knowledge sharing by different application systems. We constructed a question answering, system to demonstrate the power of INFOMAP. When the QA-system receives a user's query, it will extract the corresponding events or scripts based on the ontology in INFOMAP The understanding of a question involves extracting such information as the question type, the question subject, the question condition and the question context A dialogue on the question is triggered at the same time to guide the user to retrieve more relevant information. Wen-Lian Hsu, Shih-Hung Wu, Yi-Shiou Chen |
SMC | 1 |
| 2001 | Selected papers from COCOON 1998 - Foreword
Wen-Lian Hsu, Ming-Yang Kao |
Theor. Comput. Sci. | 1 |
| 2000 | Semantic Search on Internet Tabular Information Extraction for Answering QueriesabstractAlthough extracting information from tables is essential for Internet information agents, most tables are designed for human eyes and their layout and semantic meanings are not well defined. In practice, encoding the layout of each information source is impossible. This work presents a novel semantic search approach capable of extracting information from general tables. Semantic ontology allows our agents to read tables in the same knowledge domain with different layouts. In addition, a system of layout syntax and a set of transformation rules are defined to transform tables into databases without losing their semantic meanings. Huei-Long Wang, Shih-Hung Wu, K. K. Wang, Cheng-Lung Sung, Wen-Lian Hsu, Wei-Kuan Shih |
CIKM | 5 |
| 1999 | Fast and Simple Algorithms for Recognizing Chordal Comparability Graphs and Interval GraphsabstractIn this paper, we present a linear-time algorithm for substitution decomposition on chordal graphs. Based on this result, we develop a linear-time algorithm for transitive orientation on chordal comparability graphs, which reduces the complexity of chordal comparability recognition from O(n 2 ) to O(n+m). We also devise a simple linear-time algorithm for interval graph recognition where no complicated data structure is involved. Wen-Lian Hsu, Tze-Heng Ma |
SIAM J. Comput. | 1 |
| 1999 | A New Planarity Test
Wei-Kuan Shih, Wen-Lian Hsu |
Theor. Comput. Sci. | 2 |
| 1998 | Empirical study of Mandarin Chinese discourse analysis: an event-based approachabstractDiscourse analysis plays an important role in natural language understanding. Mandarin Chinese discourse, which has many different properties compared with English discourse, is still far behind in the construction of a basic computational model. We propose an event model to elucidate anaphora and ellipsis in Mandarin Chinese. An event based approach (EBA) based on the model is designed to resolve anaphora and ellipsis in Mandarin Chinese discourse. In this approach, we provide an event based partial parser and an event based reasoning mechanism. This approach is applied to the mathematics word problems of elementary school (MWES). Our results provide empirical evidence that the EBA can resolve many difficult problems in Mandarin Chinese discourse. Yi-Shiou Chen, Wen-Lian Hsu |
ICTAI | 3 |
| 1998 | Personal BrowserabstractNo abstract available. Yi-Shiou Chen, Schy Chiou, Wen-Lian Hsu |
SIGIR | 4 |
| 1997 | On Physical Mapping Algorithms - An Error-Tolerant Test for the Consecutive Ones Property
Wen-Lian Hsu |
COCOON | 1 |
| 1995 | A Linear Time Algorithm For Finding Maximal Planar Subgraphs
Wen-Lian Hsu |
ISAAC | 1 |
| 1995 | O(M*N) Algorithms for the Recognition and Isomorphism Problems on Circular-Arc GraphsabstractCircular-arc graphs have a rich combinatorial structure. The circular endpoint sequence of arcs in a model for a circular-arc graph is usually far from unique. We present a natural restriction on these models to make it meaningful to define the unique representations for circular-arc graphs. We characterize those circular-arc graphs which have unique restricted models and give an $O(m \cdot n)$ algorithm for recognizing circular-arc graphs. We think a more careful implementation could reduce the complexity to $O(n^{2})$. Our approach is to reduce the recognition problem of circular-arc graphs to that of circle graphs. This approach has the following advantages: it is conceptually simpler than Tucker’s $O(n^{3})$ recognition algorithm: it exploits the similarity between circle graphs and circular-arc graphs in a natural fashion: it yields an isomorphism algorithm. A main contribution of this result is an illustration of the transformed decomposition technique. The decomposition tree developed for circular-arc graphs generalizes the concept of the PQ-tree, which is a data structure that keeps track of all possible interval representations of a given interval graph. As a consequence, our approach also yields an $O(m \cdot n)$ isomorphism algorithm for circle graphs. Wen-Lian Hsu |
SIAM J. Comput. | 1 |
| 1993 | Stroke segmentation as a basis for structural matching of Chinese charactersabstractA new stroke segmentation technique is reported. It is based on two essential operations applied to a given character: grouping adjacent segments into blade-like objects according to change of width as measured from global directions, and transforming those objects according to width variation as measured from intrinsic orientations. This technique proves to be effective in dividing Chinese characters into overlapping and non-overlapping components. Strokes are then restored, with the help of templates, by linking certain non-overlapping components terminating at the same overlapping areas. The resulting framework also lays down a basis for the structural matching of Chinese characters.> Fu Chang, Ying-Chu Chen, Hon-Son Don, Wen-Lian Hsu, Ching-I Kao |
ICDAR | 4 |
| 1993 | Fast Algorithms for the Dominating Set Problem on Permutation Graphs
Kuo-Hui Tsai, Wen-Lian Hsu |
Algorithmica | 2 |
| 1992 | A Simple Test for the Consecutive Ones Property
Wen-Lian Hsu |
ISAAC | 1 |
| 1992 | A Simple Test for Interval Graphs
Wen-Lian Hsu |
WG | 1 |
| 1992 | An O(n² log n) Algorithm for the Hamiltonian Cycle Problem on Circular-Arc GraphsabstractA circular arc family F is a collection of arcs on a circle. A circular-arc graph is the intersection graph of an arc family. A Hamiltonian cycle (HC) in a graph is a cycle that passes through every vertex exactly once. This paper presents an $O(n^2 \log n)$ algorithm to determine whether a given circular-arc graph contains an HC. This algorithm is based on two subroutines for interval graphs: (i) a linear time greedy algorithm for the node disjoint path cover problem and (ii) a linear time HC algorithm. If the given graph does not contain an HC, this paper can produce a proof either through the deletion of an appropriate cutset or through the failure to obtain a specific type of HC. Wei-Kuan Shih, T. C. Chern, Wen-Lian Hsu |
SIAM J. Comput. | 3 |
| 1991 | Linear Time Algorithms on Circular-Arc Graphs
Wen-Lian Hsu, Kuo-Hui Tsai |
Inf. Process. Lett. | 1 |
| 1990 | O(m\cdotn) Isomorphism Algorithms for Circular-Arc Graphs and Circle Graphs
Wen-Lian Hsu |
IPCO | 1 |
| 1989 | An O(n1.5) algorithm to color proper circular arcs
Wei-Kuan Shih, Wen-Lian Hsu |
Discret. Appl. Math. | 2 |
| 1989 | An O(n log n+m log log n) Maximum Weight Clique Algorithm for Circular-Arc Graphs
Wei-Kuan Shih, Wen-Lian Hsu |
Inf. Process. Lett. | 2 |
| 1989 | Recognizing circle graphs in polynomial timeabstractThe main result of this paper is an 0 ([ V ] x [ E ]) time algorithm for deciding whether a given graph is a circle graph, that is, the intersection graph of a set of chords on a circle. The algorithm utilizes two new graph-theoretic results, regarding necessary induced subgraphs of graphs having neither articulation points nor similar pairs of vertices. Furthermore, as a substep of the algorithm, it is shown how to find in 0 ([ V ] x [ E ]) time a decomposition of a graph into prime graphs, thereby improving on a result of Cunningham. Csaba P. Gabor, Kenneth J. Supowit, Wen-Lian Hsu |
J. ACM | 3 |
| 1988 | The coloring and maximum independent set problems on planar perfect graphsabstractEfficient decomposition algorithms for the weighted maximum independent set, minimum coloring, and minimum clique cover problems on planar perfect graphs are presented. These planar graphs can also be characterized by the absence of induced odd cycles of length greater than 3 (odd holes). The algorithm in this paper is based on decomposing these graphs into essentially two special classes of inseparable component graphs whose optimization problems are easy to solve, finding the solutions for these components and combining them to form a solution for the original graph. These two classes are (i) planar comparability graphs and (ii) planar line graphs of those planar bipartite graphs whose maximum degrees are no greater than three. The same techniques can be applied to other classes of perfect graphs, provided that efficient algorithms are available for their inseparable component graphs. Wen-Lian Hsu |
J. ACM | 1 |
| 1987 | Recognizing planar perfect graphsabstractAn O ( n 3 ) algorithm for recognizing planar graphs that do not contain induced odd cycles of length greater than 3 (odd holes) is presented. A planar graph with this property satisfies the requirement that its maximum clique size equal the minimum number of colors required for the graph (graphs all of whose induced subgraphs satisfy the latter property are perfect as defined by Berge). The algorithm presented is based on decomposing these graphs into essentially two special classes of inseparable component graphs that are easy to recognize. They are (i) planar comparability graphs and (ii) planar line graphs of those planar bipartite graphs whose maximum degrees are no greater than 3. Composition schemes for generating planar perfect graphs from those basic components are also provided. This decomposition algorithm can also be adapted to solve the corresponding maximum independent set and minimum coloring problems. Finally, the path-parity problem on planar perfect graphs is considered. Wen-Lian Hsu |
J. ACM | 1 |
| 1985 | Recognizing Circle Graphs in Polynomial TimeabstractOur main result is a polynomialtime algorithm for deciding whether a given graph is a circle graph, that is, the intersection graph of a set of chords on a circle. Our algorithm utilizes two new graph-theoretic results, regarding necessary induced subgraphs of graphs having neither articulation points nor similar pairs of vertices. Csaba P. Gabor, Wen-Lian Hsu, Kenneth J. Supowit |
FOCS | 2 |
| 1985 | Maximum Weight Clique Algorithms for Circular-Arc Graphs and Circle GraphsabstractCircle graphs and circular-arc graphs are the intersection graphs of chords and arcs in a circle. In this paper we present algorithms for finding maximum weight cliques in these graphs. The running times of the algorithms are $O(n^2 + m\log \log n)$ for circle graphs and $O(mn)$ for circular-arc graphs. Our algorithms are based on the scanning of appropriate endpoint sequences and efficient bookkeeping of results for subproblems. Wen-Lian Hsu |
SIAM J. Comput. | 1 |
| 1984 | On the maximum empty rectangle problem
Amnon Naamad, D. T. Lee, Wen-Lian Hsu |
Discret. Appl. Math. | 3 |
| 1979 | Easy and hard bottleneck location problems
Wen-Lian Hsu, George L. Nemhauser |
Discret. Appl. Math. | 1 |