VLDB 2026 Research / reviewers in the wild / expert
Richard Tzong-Han Tsai
dblp:t/TzongHanTsai · also Tzong-Han Tsai
· DBLP profile ↗
65ranked-venue papers
21as first author
15since 2021 · last 2026
0000-0003-0513-107XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 38 · 7 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 11 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SiniticMTError: A Machine Translation Dataset with Error Annotations for Sinitic LanguagesabstractDespite major advances in machine translation (MT) in recent years, progress remains limited for many low-resource languages that lack large-scale training data and linguistic resources. In this paper, we introduce \dsname, a novel fine-grained dataset that builds on existing parallel corpora to provide error span, error type, and error severity annotations in machine-translated examples from English to Mandarin, Cantonese, and Wu Chinese, along with a Mandarin-Hokkien component derived from a non-parallel source. Our dataset serves as a resource for the MT community to fine-tune models with error detection capabilities, supporting research on translation quality estimation, error-aware generation, and low-resource language evaluation. We also establish baseline results using language models to benchmark translation error detection performance. Specifically, we evaluate multiple open source and closed source LLMs using span-level and correlation-based MQM metrics, revealing their limited precision, underscoring the need for our dataset. Finally, we report our rigorous annotation process by native speakers, with analyses on pilot studies, iterative feedback, insights, and patterns in error type and severity. Hannah Liu, Junghyun Min, Annie En-Shiun Lee, Ethan Yue Heng Cheung, Shou-Yi Hung, Elsie Chan, Shiyao Qian, Runtong Liang, Kimlan Huynh, Wing Yu Yip, York Hay Ng, Tsz Fung Yau, Ka Ieng Charlotte Lo, You-Wei Wu, Richard Tzong-Han Tsai |
LREC | 15 |
| 2025 | Unsupervised domain adaptation for cross-style, cross-year land use understanding from historical mapsabstractDigitizing historical topographic maps is essential for spatial analysis in GIS; however, conventional methods for digitizing these maps are labor-intensive and challenging due to non-explicit boundaries and inconsistent map styles. We address these challenges by proposing a new Map Style Segmentation (MapStyleSeg) method that employs unsupervised domain adaptation (UDA) from deep learning (DL) to enhance cross-style, cross-year automatic map segmentation and conversion. Our method, MapStyleSeg, is exemplified by training on a fully annotated topographic map of Taiwan in 2017 and applying it to a 2001 topographic map without annotations. We also evaluated different encoder-decoder architectures and loss functions. Our results show that using the ResNet-101 backbone with the SegFormer decoder and a mix of focal and Dice loss yields the best performance: 94.94% overall accuracy (Acc), 81.8% mean Intersection over Union (mIoU), outperforming standard U-Net models without UDA (88.23% Acc, 49.3% mIoU). Our approach addresses the challenges of digitizing historical maps with varying styles, further advancing GIS digitization of historical maps, and offering useful information for urban planning, environmental monitoring, and decision-making processes. This work highlights the novel use of DL algorithms to automate complex GIS data processing that transforms historical maps into spatial datasets. Jun-Hua Wang, Andy Da-Yu Wang, Hsiung-Ming Liao, Ming-Ching Chang, Richard Tzong-Han Tsai |
Int. J. Geogr. Inf. Sci. | 5 |
| 2025 | SMUTF: Schema Matching Using Generative Tags and Hybrid Features
Yu Zhang 0202, Mei Di, Haozheng Luo, Chenwei Xu, Richard Tzong-Han Tsai |
Inf. Syst. | 5 |
| 2024 | Automated Assessment of Fidelity and Interpretability: An Evaluation Framework for Large Language Models' Explanations (Student Abstract)abstractAs Large Language Models (LLMs) become more prevalent in various fields, it is crucial to rigorously assess the quality of their explanations. Our research introduces a task-agnostic framework for evaluating free-text rationales, drawing on insights from both linguistics and machine learning. We evaluate two dimensions of explainability: fidelity and interpretability. For fidelity, we propose methods suitable for proprietary LLMs where direct introspection of internal features is unattainable. For interpretability, we use language models instead of human evaluators, addressing concerns about subjectivity and scalability in evaluations. We apply our framework to evaluate GPT-3.5 and the impact of prompts on the quality of its explanations. In conclusion, our framework streamlines the evaluation of explanations from LLMs, promoting the development of safer models. Mu-Tien Kuo, Chih-Chung Hsueh, Richard Tzong-Han Tsai |
AAAI | 3 |
| 2024 | Chat Vector: A Simple Approach to Equip LLMs with Instruction Following and Model Alignment in New LanguagesabstractShih-Cheng Huang, Pin-Zu Li, Yu-chi Hsu, Kuang-Ming Chen, Yu Tung Lin, Shih-Kai Hsiao, Richard Tsai, Hung-yi Lee. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Shih-Cheng Huang, Pin-Zu Li, Yu-Chi Hsu, Kuang-Ming Chen, Yu-Tung Lin, Shih-Kai Hsiao, Richard Tzong-Han Tsai, Hung-yi Lee |
ACL (1) | 7 |
| 2024 | Enhancing Taiwanese Hokkien Dual Translation by Exploring and Standardizing of Four Writing SystemsabstractMachine translation focuses mainly on high-resource languages (HRLs), while low-resource languages (LRLs) like Taiwanese Hokkien are relatively under-explored. The study aims to address this gap by developing a dual translation model between Taiwanese Hokkien and both Traditional Mandarin Chinese and English. We employ a pre-trained LLaMA 2-7B model specialized in Traditional Mandarin Chinese to leverage the orthographic similarities between Taiwanese Hokkien Han and Traditional Mandarin Chinese. Our comprehensive experiments involve translation tasks across various writing systems of Taiwanese Hokkien as well as between Taiwanese Hokkien and other HRLs. We find that the use of a limited monolingual corpus still further improves the model’s Taiwanese Hokkien capabilities. We then utilize our translation model to standardize all Taiwanese Hokkien writing systems into Hokkien Han, resulting in further performance improvements. Additionally, we introduce an evaluation method incorporating back-translation and GPT-4 to ensure reliable translation quality assessment even for LRLs. The study contributes to narrowing the resource gap for Taiwanese Hokkien and empirically investigates the advantages and limitations of pre-training and fine-tuning based on LLaMA 2. Bo-Han Lu, Annie En-Shiun Lee, Richard Tzong-Han Tsai |
LREC/COLING | 4 |
| 2024 | Tri-directional Hypergraph Contrastive Learning for Session-based RecommendationabstractSession-based recommendation (SBR) aims to forecast users’ future actions by analyzing their unnamed behavioral sequences within a limited time frame. Recent research in SBR has focused on leveraging various techniques, including the incorporation of contrastive learning. Despite these developments, existing studies exhibit several limitations. First, these studies solely employ either item-level or session-level contrast, overlooking the vital correlation information between items and sessions. Second, to model the various relationships present in session data, numerous studies have designed complex processes to construct multiple augmented views, which diminish the accessibility of graph contrastive learning in SBR. To overcome these challenges, we propose Tri-Rec (Tri-directional Hypergraph Contrastive Learning for Session-based Recommendation), a model that innovatively incorporates tri-directional contrast into SBR. Tri- directional contrast consists of three distinct contrastive forms, with the aim of maximizing the similarity: (1) between the same item, (2) between the same session, and (3) between each session and its containing items in augmented views. Contrary to many prevailing methods that solely employ either item-level or session-level contrast, we not only utilize both but also introduce membership-level contrast, allowing the model to harness more comprehensive information. Furthermore, we integrate the hypergraph neural network and a self-attention based readout module to capture both high-order relationships and representative user intent among sessions. Detailed empirical evaluations conducted on three real-world datasets reveal that Tri-Rec markedly surpasses state-of-the-art approaches in performance. Da-Ren Dai, Richard Tzong-Han Tsai |
IJCNN | 2 |
| 2024 | Surveying biomedical relation extraction: a critical examination of current datasets and the proposal of a new resourceabstractNatural language processing (NLP) has become an essential technique in various fields, offering a wide range of possibilities for analyzing data and developing diverse NLP tasks. In the biomedical domain, understanding the complex relationships between compounds and proteins is critical, especially in the context of signal transduction and biochemical pathways. Among these relationships, protein-protein interactions (PPIs) are of particular interest, given their potential to trigger a variety of biological reactions. To improve the ability to predict PPI events, we propose the protein event detection dataset (PEDD), which comprises 6823 abstracts, 39 488 sentences and 182 937 gene pairs. Our PEDD dataset has been utilized in the AI CUP Biomedical Paper Analysis competition, where systems are challenged to predict 12 different relation types. In this paper, we review the state-of-the-art relation extraction research and provide an overview of the PEDD's compilation process. Furthermore, we present the results of the PPI extraction competition and evaluate several language models' performances on the PEDD. This paper's outcomes will provide a valuable roadmap for future studies on protein event detection in NLP. By addressing this critical challenge, we hope to enable breakthroughs in drug discovery and enhance our understanding of the molecular mechanisms underlying various diseases. Ming-Siang Huang, Jen-Chieh Han, Pei-Yen Lin, Yu-Ting You, Richard Tzong-Han Tsai, Wen-Lian Hsu |
Briefings Bioinform. | 5 |
| 2023 | MingOfficial: A Ming Official Career Dataset and a Historical Context-Aware Representation Learning FrameworkabstractIn Chinese studies, understanding the nuanced traits of historical figures, often not explicitly evident in biographical data, has been a key interest.However, identifying these traits can be challenging due to the need for domain expertise, specialist knowledge, and context-specific insights, making the eprocess time-consuming and difficult to scale.Our focus on studying officials from China's Ming Dynasty is no exception.To tackle this challenge, we propose MingOfficial, a large-scale multi-modal dataset consisting of both structured (career records, annotated personnel types) and text (historical texts) data for 13, 031 officials.We further couple the dataset with a graph neural network (GNN) to combine both modalities in order to allow investigation of social structures and provide features to boost down-stream tasks.Experiments show that our proposed MingOfficial could enable exploratory analysis of official identities, and also significantly boost performance in tasks such as identifying nuance identities (e.g.civil officials holding military power) from 24.6% to 98.2% F 1 score in holdout test set.By making MingOfficial publicly available at https://data.depositar.io/ en/dataset/ming_official as both a dataset and an interactive tool, we aim to stimulate further research into the role of social context and representation learning in identifying individual characteristics, and hope to provide inspiration for computational approaches in other fields beyond Chinese studies. You-Jun Chen, Hsin-Yi Hsieh, Yingtao Tian, Bert Chan, Yu-Sin Liu, Richard Tzong-Han Tsai |
EMNLP | 8 |
| 2022 | Mixed Embedding of XLM for Unsupervised Cantonese-Chinese Neural Machine Translation (Student Abstract)abstractUnsupervised Neural Machines Translation is the most ideal method to apply to Cantonese and Chinese translation because parallel data is scarce in this language pair. In this paper, we proposed a method that combined a modified cross-lingual language model and performed layer to layer attention on unsupervised neural machine translation. In our experiments, we observed that our proposed method does improve the Cantonese to Chinese and Chinese to Cantonese translation by 1.088 and 0.394 BLEU scores. We finally developed a web service based on our ideal approach to provide Cantonese to Chinese Translation and vice versa. Ka Ming Wong, Richard Tzong-Han Tsai |
AAAI | 2 |
| 2022 | BRCC and SentiBahasaRojak: The First Bahasa Rojak Corpus for Pretraining and Sentiment Analysis DatasetabstractCode-mixing refers to the mixed use of multiple languages. It is prevalent in multilingual societies and is also one of the most challenging natural language processing tasks. In this paper, we study Bahasa Rojak, a dialect popular in Malaysia that consists of English, Malay, and Chinese. Aiming to establish a model to deal with the code-mixing phenomena of Bahasa Rojak, we use data augmentation to automatically construct the first Bahasa Rojak corpus for pre-training language models, which we name the Bahasa Rojak Crawled Corpus (BRCC). We also develop a new pre-trained model called “Mixed XLM”. The model can tag the language of the input token automatically to process code-mixing input. Finally, to test the effectiveness of the Mixed XLM model pre-trained on BRCC for social media scenarios where code-mixing is found frequently, we compile a new Bahasa Rojak sentiment analysis dataset, SentiBahasaRojak, with a Kappa value of 0.77. Nanda Putri Romadhona, Sin-En Lu, Bo-Han Lu, Richard Tzong-Han Tsai |
COLING | 4 |
| 2022 | Cross-language article linking with deep neural network based paragraph encoding
Yu-Chun Wang, Chia-Min Chuang, Chun-Kai Wu, Chao-Lin Pan, Richard Tzong-Han Tsai |
Comput. Speech Lang. | 5 |
| 2022 | Improving low-resource machine transliteration by using 3-way transfer learningabstractTransfer learning brings improvement to machine translation by using a resource-rich language pair to pretrain the model and then adapting it to the desired language pair. However, to date, there have been few attempts to tackle machine transliteration with transfer learning. In this article, we propose a method of using source–pivot and pivot–target datasets to improve source–target machine transliteration. Our approach first bridges the source–pivot and pivot–target datasets by reducing the distance between source and pivot embeddings. Then, our model learns to translate from the pivot language to the target language. Finally, the source–target dataset is used to fine tune the model. Our experiments show that our method is superior to the transfer learning method. When implemented with a state-of-the-art source–target translation model from NEWS’18, our transfer learning method can improve the accuracy by 1.1%. Chun-Kai Wu, Chao-Chuang Shih, Yu-Chun Wang, Richard Tzong-Han Tsai |
Comput. Speech Lang. | 4 |
| 2022 | EPG2S: Speech Generation and Speech Enhancement Based on Electropalatography and Audio Signals Using Multimodal LearningabstractSpeech generation and enhancement based on articulatory movements facilitate communication when the scope of verbal communication is absent, e.g., in patients who have lost the ability to speak. Although various techniques have been proposed to this end, electropalatography (EPG), which is a monitoring technique that records contact between the tongue and hard palate during speech, has not been adequately explored. Herein, we propose a novel multimodal EPG-to-speech (EPG2S) system that utilizes EPG and speech signals for speech generation and enhancement. Different fusion strategies based on multiple combinations of EPG and noisy speech signals are examined, and the viability of the proposed method is investigated. Experimental results indicate that EPG2S achieves desirable speech generation outcomes based solely on EPG signals. Further, the addition of noisy speech signals is observed to improve quality and intelligibility. Additionally, EPG2S is observed to achieve high-quality speech enhancement based solely on audio signals, with the addition of EPG signals further improving the performance. The late fusion strategy is deemed to be the most effective approach for simultaneous speech generation and enhancement. Lichin Chen, Po-Hsun Chen, Richard Tzong-Han Tsai, Yu Tsao 0001 |
IEEE Signal Process. Lett. | 3 |
| 2021 | Multi-modal User Intent Classification Under the Scenario of Smart Factory (Student Abstract)abstractQuestion-answering systems are becoming increasingly popular in Natural Language Processing, especially when applied in smart factory settings. A common practice in designing those systems is through intent classification. However, in a multiple-stage task commonly seen in those settings, relying solely on intent classification may lead to erroneous answers, as questions rising from different work stages may share the same intent but have different contexts and therefore require different answers. To address this problem, we designed an interactive dialogue system that utilizes contextual information to assist intent classification in a multiple-stage task. Specifically, our system incorporates user’s utterances with real-time video feed to better situate users’ questions and analyze their intent. Yu-Ching Chiu, Bo-Hao Chang, Tzu-Yu Chen, Cheng-Fu Yang, Nanyi Bi, Richard Tzong-Han Tsai, Hung-yi Lee, Yung-Jen Hsu 0001 |
AAAI | 6 |
| 2020 | Biomedical named entity recognition and linking datasets: survey and our recent developmentabstractNatural language processing (NLP) is widely applied in biological domains to retrieve information from publications. Systems to address numerous applications exist, such as biomedical named entity recognition (BNER), named entity normalization (NEN) and protein-protein interaction extraction (PPIE). High-quality datasets can assist the development of robust and reliable systems; however, due to the endless applications and evolving techniques, the annotations of benchmark datasets may become outdated and inappropriate. In this study, we first review commonlyused BNER datasets and their potential annotation problems such as inconsistency and low portability. Then, we introduce a revised version of the JNLPBA dataset that solves potential problems in the original and use state-of-the-art named entity recognition systems to evaluate its portability to different kinds of biomedical literature, including protein-protein interaction and biology events. Lastly, we introduce an ensembled biomedical entity dataset (EBED) by extending the revised JNLPBA dataset with PubMed Central full-text paragraphs, figure captions and patent abstracts. This EBED is a multi-task dataset that covers annotations including gene, disease and chemical entities. In total, it contains 85000 entity mentions, 25000 entity mentions with database identifiers and 5000 attribute tags. To demonstrate the usage of the EBED, we review the BNER track from the AI CUP Biomedical Paper Analysis challenge. Availability: The revised JNLPBA dataset is available at https://iasl-btm.iis.sinica.edu.tw/BNER/Content/Re vised_JNLPBA.zip. The EBED dataset is available at https://iasl-btm.iis.sinica.edu.tw/BNER/Content/AICUP _EBED_dataset.rar. Contact: Email: [email protected], Tel. 886-3-4227151 ext. 35203, Fax: 886-3-422-2681 Email: [email protected], Tel. 886-2-2788-3799 ext. 2211, Fax: 886-2-2782-4814 Supplementary information: Supplementary data are available at Briefings in Bioinformatics online. Ming-Siang Huang, Po-Ting Lai, Pei-Yen Lin, Yu-Ting You, Richard Tzong-Han Tsai, Wen-Lian Hsu |
Briefings Bioinform. | 5 |
| 2019 | Using Deep-Q Network to Select Candidates from N-best Speech Recognition Hypotheses for Enhancing Dialogue State TrackingabstractMost state-of-the-art dialogue state tracking (DST) methods infer the dialogue state based on ground-truth transcriptions of utterances. In real-world situations, utterances are transcribed by automatic speech recognition (ASR) systems, which output the n-best candidate transcriptions (hypotheses). In certain noisy environments, the best transcription is often imperfect, severely influencing DST accuracy and possibly causing the dialogue system to stall or loop. The missed or misrecognized words can often be found in the runner-up candidate transcriptions from 2 to n, which could be used to improve accuracy of DST. However, looking beyond the top-ranked ASR results poses a dilemma: going too far may introduce noise, while not going far enough may not uncover any useful information. In this paper, we propose a novel approach to automatically determine the optimal time to stop reexamining runner-up ASR transcriptions based on deep reinforcement learning. Our method outperforms the baseline system, which uses only the top-1 ASR result, by 3.1%. Then, we select the dialogue rounds with the top-10 largest word error rate (WER), our method can improve DST accuracy by 15.4%, which is five times the overall improvement rate (3.1%). This improvement was expected because our proposed method is able to select informative ASR results at any rank. Richard Tzong-Han Tsai, Chia-Hao Chen, Chun-Kai Wu, Yu-Cheng Hsiao, Hung-yi Lee |
ICASSP | 1 |
| 2017 | Context-aware sentiment propagation using LDA topic modeling on Chinese ConceptNet
Po-Hao Chou, Richard Tzong-Han Tsai, Yung-Jen Hsu 0001 |
Soft Comput. | 2 |
| 2016 | Cross-language article linking with different knowledge bases using bilingual topic model and translation features
Yu-Chun Wang, Chun-Kai Wu, Richard Tzong-Han Tsai |
Knowl. Based Syst. | 3 |
| 2016 | Collective Web-Based Parenthetical Translation Extraction Using Markov Logic NetworksabstractParenthetical translations are translations of terms in otherwise monolingual text that appear inside parentheses. Parenthetical translations extraction (PTE) is the task of extracting parenthetical translations from natural language documents. One of the main difficulties in PTE is to detect the left boundary of the translated term in preparenthetical text. In this article, we propose a collective approach that employs Markov logic to model multiple constraints used in the PTE task. We show how various constraints can be formulated and combined in a Markov logic network (MLN). Our experimental results show that the proposed collective PTE approach significantly outperforms a current state-of-the-art method, improving the average F-measure up to 27.11% compared to the previous word alignment approach. It also outperforms an individual MLN-based system by 8.2% and a system based on conditional random fields by 5.9%. Richard Tzong-Han Tsai |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2015 | A context-aware approach for progression tracking of medical concepts in electronic medical recordsabstractElectronic medical records (EMRs) for diabetic patients contain information about heart disease risk factors such as high blood pressure, cholesterol levels, and smoking status. Discovering the described risk factors and tracking their progression over time may support medical personnel in making clinical decisions, as well as facilitate data modeling and biomedical research. Such highly patient-specific knowledge is essential to driving the advancement of evidence-based practice, and can also help improve personalized medicine and care. One general approach for tracking the progression of diseases and their risk factors described in EMRs is to first recognize all temporal expressions, and then assign each of them to the nearest target medical concept. However, this method may not always provide the correct associations. In light of this, this work introduces a context-aware approach to assign the time attributes of the recognized risk factors by reconstructing contexts that contain more reliable temporal expressions. The evaluation results on the i2b2 test set demonstrate the efficacy of the proposed approach, which achieved an F-score of 0.897. To boost the approach's ability to process unstructured clinical text and to allow for the reproduction of the demonstrated results, a set of developed .NET libraries used to develop the system is available at https://sites.google.com/site/hongjiedai/projects/nttmuclinicalnet. Nai-Wen Chang 0001, Hong-Jie Dai, Jitendra Jonnagaddala, Chih-Wei Chen, Richard Tzong-Han Tsai, Wen-Lian Hsu |
J. Biomed. Informatics | 5 |
| 2014 | A resource-saving collective approach to biomedical semantic role labelingabstractBACKGROUND: Biomedical semantic role labeling (BioSRL) is a natural language processing technique that identifies the semantic roles of the words or phrases in sentences describing biological processes and expresses them as predicate-argument structures (PAS's). Currently, a major problem of BioSRL is that most systems label every node in a full parse tree independently; however, some nodes always exhibit dependency. In general SRL, collective approaches based on the Markov logic network (MLN) model have been successful in dealing with this problem. However, in BioSRL such an approach has not been attempted because it would require more training data to recognize the more specialized and diverse terms found in biomedical literature, increasing training time and computational complexity. RESULTS: We first constructed a collective BioSRL system based on MLN. This system, called collective BIOSMILE (CBIOSMILE), is trained on the BioProp corpus. To reduce the resources used in BioSRL training, we employ a tree-pruning filter to remove unlikely nodes from the parse tree and four argument candidate identifiers to retain candidate nodes in the tree. Nodes not recognized by any candidate identifier are discarded. The pruned annotated parse trees are used to train a resource-saving MLN-based system, which is referred to as resource-saving collective BIOSMILE (RCBIOSMILE). Our experimental results show that our proposed CBIOSMILE system outperforms BIOSMILE, which is the top BioSRL system. Furthermore, our proposed RCBIOSMILE maintains the same level of accuracy as CBIOSMILE using 92% less memory and 57% less training time. CONCLUSIONS: This greatly improved efficiency makes RCBIOSMILE potentially suitable for training on much larger BioSRL corpora over more biomedical domains. Compared to real-world biomedical corpora, BioProp is relatively small, containing only 445 MEDLINE abstracts and 30 event triggers. It is not large enough for practical applications, such as pathway construction. We consider it of primary importance to pursue SRL training on large corpora in the future. Richard Tzong-Han Tsai, Po-Ting Lai |
BMC Bioinform. | 1 |
| 2014 | Using relation selection to improve value propagation in a ConceptNet-based sentiment dictionary
Chi-En Wu, Richard Tzong-Han Tsai |
Knowl. Based Syst. | 2 |
| 2013 | Transliteration Extraction from Classical Chinese Buddhist Literature Using Conditional Random Fields
Yu-Chun Wang, Richard Tzong-Han Tsai |
PACLIC | 2 |
| 2013 | TEMPTING system: A hybrid method of rule and machine learning for temporal relation extraction in patient discharge summaries
Yung-Chun Chang, Hong-Jie Dai, Johnny Chi-Yang Wu, Jian-Ming Chen, Richard Tzong-Han Tsai, Wen-Lian Hsu |
J. Biomed. Informatics | 5 |
| 2012 | Unsupervised Japanese-Chinese Opinion Word Translation using Dependency Distance and Feature-Opinion Association Weight
Guo-Hau Lai, Ying-Mei Guo, Richard Tzong-Han Tsai |
COLING | 3 |
| 2012 | Coreference resolution of medical concepts in discharge summaries by exploiting contextual informationabstractOBJECTIVE: Patient discharge summaries provide detailed medical information about hospitalized patients and are a rich resource of data for clinical record text mining. The textual expressions of this information are highly variable. In order to acquire a precise understanding of the patient, it is important to uncover the relationship between all instances in the text. In natural language processing (NLP), this task falls under the category of coreference resolution. DESIGN: A key contribution of this paper is the application of contextual-dependent rules that describe relationships between coreference pairs. To resolve phrases that refer to the same entity, the authors use these rules in three representative NLP systems: one rule-based, another based on the maximum entropy model, and the last a system built on the Markov logic network (MLN) model. RESULTS: The experimental results show that the proposed MLN-based system outperforms the baseline system (exact match) by average F-scores of 4.3% and 5.7% on the Beth and Partners datasets, respectively. Finally, the three systems were integrated into an ensemble system, further improving performance to 87.21%, which is 4.5% more than the official i2b2 Track 1C average (82.7%). CONCLUSION: In this paper, the main challenges in the resolution of coreference relations in patient discharge summaries are described. Several rules are proposed to exploit contextual information, and three approaches presented. While single systems provided promising results, an ensemble approach combining the three systems produced a better performance than even the best single system. Hong-Jie Dai, Johnny Chi-Yang Wu, Po-Ting Lai, Richard Tzong-Han Tsai, Wen-Lian Hsu |
J. Am. Medical Informatics Assoc. | 5 |
| 2012 | Validating Contradiction in Texts Using Online Co-Mention Pattern CheckingabstractDetecting contradictive statements is a foundational and challenging task for text understanding applications such as textual entailment. In this article, we aim to address the problem of the shortage of specific background knowledge in contradiction detection. A novel contradiction detecting approach based on the distribution of the query composed of critical mismatch combinations on the Internet is proposed to tackle the problem. By measuring the availability of mismatch conjunction phrases (MCPs), the background knowledge about two target statements can be implicitly obtained for identifying contradictions. Experiments on three different configurations show that the MCP-based approach achieves remarkable improvement on contradiction detection and can significantly improve the performance of textual entailment recognition. Cheng-Wei Shih, Chengwei Lee, Richard Tzong-Han Tsai, Wen-Lian Hsu |
ACM Trans. Asian Lang. Inf. Process. | 3 |
| 2012 | A Generative Data Augmentation Model for Enhancing Chinese Dialect Pronunciation PredictionabstractMost spoken Chinese dialects lack comprehensive digital pronunciation databases, which are crucial for speech processing tasks. Given complete pronunciation databases for related dialects, one can use supervised learning techniques to predict a Chinese character's pronunciation in a target dialect based on the character's features and its pronunciation in other related dialects. Unfortunately, Chinese dialect pronunciation databases are far from complete. We propose a novel generative model that makes use of both existing dialect pronunciation data plus medieval rime books to discover patterns that exist in multiple dialects. The proposed model can augment missing dialectal pronunciations based on existing dialect pronunciation tables (even if incomplete) and the pronunciation data in rime books. The augmented pronunciation database can then be used in supervised learning settings. We evaluate the prediction accuracy in terms of phonological features, such as tone, initial phoneme, final phoneme, etc. For each character, features are evaluated on the whole, overall pronunciation feature accuracy (OPFA). Our first experimental results show that adding features from dialectal pronunciation data to our baseline rime-book model dramatically improves OPFA using the support vector machine (SVM) model. In the second experiment, we compare the performance of the SVM model using phonological features from closely related dialects with that of the model using phonological features from non-closely related dialects. The experimental results show that using features from closely related dialects results in higher accuracy. In the third experiment, we show that using our proposed data augmentation model to fill in missing data can increase the SVM model's OPFA by up to 7.6%. Chu-Cheng Lin, Richard Tzong-Han Tsai |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Improving Protein-Protein Interaction Pair Ranking with an Integrated Global Association ScoreabstractProtein-protein interaction (PPI) database curation requires text-mining systems that can recognize and normalize interactor genes and return a ranked list of PPI pairs for each article. The order of PPI pairs in this list is essential for ease of curation. Most of the current PPI pair ranking approaches rely on association analysis between the two genes in the pair. However, we propose that ranking an extracted PPI pair by considering both the association between the paired genes and each of those genes’ global associations with all other genes mentioned in the paper can provide a more reliable ranked list. In this work, we present a composite interaction score that considers not only the association score between two interactors (pair association score) but also their global association scores. We test three representative data fusion algorithms to estimate this global association score—two Borda-Fuse models and one linear combination model (LCM). The three estimation methods are evaluated using the data set of the BioCreative II.5 Interaction Pair Task (IPT) in terms of area under the interpolated precision/recall curve (AUC iP/R). Our experimental results indicate that using LCM to estimate the global association score can boost the AUC iP/R score from 0.0175 to 0.2396, outperforming the best BioCreative II.5 IPT system. Richard Tzong-Han Tsai |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2011 | Enhancing Search Results with Semantic Annotation Using Augmented BrowsingabstractIn this paper, we describe how we integrated an artificial intelligence (AI) system into the PubMed search website using augmented browsing technology. Our system dynamically enriches the PubMed search results displayed in a user's browser with semantic annotation provided by several natural language processing (NLP) subsystems, including a sentence splitter, a part-of-speech tagger, a named entity recognizer, a section categorizer and a gene normalizer (GN). After our system is installed, the PubMed search results page is modified on the fly to categorize sections and provide additional information on gene and gene products indentified by our NLP subsystems. In addition, GN involves three main steps: candidate ID matching, false positive filtering and disambiguation, which are highly dependent on each other. We propose a joint model using a Markov logic network (MLN) to model the dependencies found in GN. The experimental results show that our joint model outperforms a baseline system that executes the three steps separately. The developed system is available at https://sites.google.com/site/pubmedannotationtool 4ijcai/home. Hong-Jie Dai, Wei-Chi Tsai, Richard Tzong-Han Tsai, Wen-Lian Hsu |
IJCAI | 3 |
| 2011 | Entity Disambiguation Using a Markov-Logic Network
Hong-Jie Dai, Richard Tzong-Han Tsai, Wen-Lian Hsu |
IJCNLP | 2 |
| 2011 | Evaluation via Negativa of Chinese Word Segmentation for Information Retrieval
Mike Tian-Jian Jiang, Cheng-Wei Shih, Richard Tzong-Han Tsai, Wen-Lian Hsu |
PACLIC | 3 |
| 2011 | Integration of gene normalization stages and co-reference resolution using a Markov logic networkabstractMOTIVATION: Gene normalization (GN) is the task of normalizing a textual gene mention to a unique gene database ID. Traditional top performing GN systems usually need to consider several constraints to make decisions in the normalization process, including filtering out false positives, or disambiguating an ambiguous gene mention, to improve system performance. However, these constraints are usually executed in several separate stages and cannot use each other's input/output interactively. In this article, we propose a novel approach that employs a Markov logic network (MLN) to model the constraints used in the GN task. Firstly, we show how various constraints can be formulated and combined in an MLN. Secondly, we are the first to apply the two main concepts of co-reference resolution-discourse salience in centering theory and transitivity-to GN models. Furthermore, to make our results more relevant to developers of information extraction applications, we adopt the instance-based precision/recall/F-measure (PRF) in addition to the article-wide PRF to assess system performance. RESULTS: Experimental results show that our system outperforms baseline and state-of-the-art systems under two evaluation schemes. Through further analysis, we have found several unexplored challenges in the GN task. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hong-Jie Dai, Yen-Ching Chang, Richard Tzong-Han Tsai, Wen-Lian Hsu |
Bioinform. | 3 |
| 2011 | The gene normalization task in BioCreative IIIabstractBACKGROUND: We report the Gene Normalization (GN) challenge in BioCreative III where participating teams were asked to return a ranked list of identifiers of the genes detected in full-text articles. For training, 32 fully and 500 partially annotated articles were prepared. A total of 507 articles were selected as the test set. Due to the high annotation cost, it was not feasible to obtain gold-standard human annotations for all test articles. Instead, we developed an Expectation Maximization (EM) algorithm approach for choosing a small number of test articles for manual annotation that were most capable of differentiating team performance. Moreover, the same algorithm was subsequently used for inferring ground truth based solely on team submissions. We report team performance on both gold standard and inferred ground truth using a newly proposed metric called Threshold Average Precision (TAP-k). RESULTS: We received a total of 37 runs from 14 different teams for the task. When evaluated using the gold-standard annotations of the 50 articles, the highest TAP-k scores were 0.3297 (k=5), 0.3538 (k=10), and 0.3535 (k=20), respectively. Higher TAP-k scores of 0.4916 (k=5, 10, 20) were observed when evaluated using the inferred ground truth over the full test set. When combining team results using machine learning, the best composite system achieved TAP-k scores of 0.3707 (k=5), 0.4311 (k=10), and 0.4477 (k=20) on the gold standard, representing improvements of 12.4%, 21.8%, and 26.6% over the best team results, respectively. CONCLUSIONS: By using full text and being species non-specific, the GN task in BioCreative III has moved closer to a real literature curation task than similar tasks in the past and presents additional challenges for the text mining community, as revealed in the overall team results. By evaluating teams using the gold standard, we show that the EM algorithm allows team submissions to be differentiated while keeping the manual annotation effort feasible. Using the inferred ground truth we show measures of comparative performance between teams. Finally, by comparing team rankings on gold standard vs. inferred ground truth, we further demonstrate that the inferred ground truth is as effective as the gold standard for detecting good team performance. Zhiyong Lu, Hung-Yu Kao, Chih-Hsuan Wei, Minlie Huang, Jingchen Liu, Cheng-Ju Kuo, Chun-Nan Hsu, Richard Tzong-Han Tsai, Hong-Jie Dai, Naoaki Okazaki, Hancheol Cho, Martin Gerner, Illés Solt, Shashank Agarwal, Dina Vishnyakova, Patrick Ruch, Martin Romacker, Fabio Rinaldi 0001, Sanmitra Bhattacharya, Padmini Srinivasan, Manabu Torii, Sérgio Matos, David Campos 0001, Karin Verspoor, Kevin M. Livingston, W. John Wilbur |
BMC Bioinform. | 8 |
| 2011 | Dynamic programming re-ranking for PPI interactor and pair extraction in full-text articlesabstractBACKGROUND: Experimentally verified protein-protein interactions (PPIs) cannot be easily retrieved by researchers unless they are stored in PPI databases. The curation of such databases can be facilitated by employing text-mining systems to identify genes which play the interactor role in PPIs and to map these genes to unique database identifiers (interactor normalization task or INT) and then to return a list of interaction pairs for each article (interaction pair task or IPT). These two tasks are evaluated in terms of the area under curve of the interpolated precision/recall (AUC iP/R) score because the order of identifiers in the output list is important for ease of curation. RESULTS: Our INT system developed for the BioCreAtIvE II.5 INT challenge achieved a promising AUC iP/R of 43.5% by using a support vector machine (SVM)-based ranking procedure. Using our new re-ranking algorithm, we have been able to improve system performance (AUC iP/R) by 1.84%. Our experimental results also show that with the re-ranked INT results, our unsupervised IPT system can achieve a competitive AUC iP/R of 23.86%, which outperforms the best BC II.5 INT system by 1.64%. Compared to using only SVM ranked INT results, using re-ranked INT results boosts AUC iP/R by 7.84%. Statistical significance t-test results show that our INT/IPT system with re-ranking outperforms that without re-ranking by a statistically significant difference. CONCLUSIONS: In this paper, we present a new re-ranking algorithm that considers co-occurrence among identifiers in an article to improve INT and IPT ranking results. Combining the re-ranked INT results with an unsupervised approach to find associations among interactors, the proposed method can boost the IPT performance. We also implement score computation using dynamic programming, which is faster and more efficient than traditional approaches. Richard Tzong-Han Tsai, Po-Ting Lai |
BMC Bioinform. | 1 |
| 2011 | Multi-stage gene normalization for full-text articles with context-based species filtering for dynamic dictionary entry selectionabstractBACKGROUND: Gene normalization (GN) is the task of identifying the unique database IDs of genes and proteins in literature. The best-known public competition of GN systems is the GN task of the BioCreative challenge, which has been held four times since 2003. The last two BioCreatives, II.5 & III, had two significant differences from earlier tasks: firstly, they provided full-length articles in addition to abstracts; and secondly, they included multiple species without providing species ID information. Full papers introduce more complex targets for GN processing, while the inclusion of multiple species vastly increases the potential size of dictionaries needed for GN. BioCreative III GN uses Threshold Average Precision at a median of k errors per query (TAP-k), a new measure closely related to the well-known average precision, but also reflecting the reliability of the score provided by each GN system. RESULTS: To use full-paper text, we employed a multi-stage GN algorithm and a ranking method which exploit information in different sections and parts of a paper. To handle the inclusion of multiple unknown species, we developed two context-based dynamic strategies to select dictionary entries related to the species that appear in the paper-section-wide and article-wide context. Our originally submitted BioCreative III system uses a static dictionary containing only the most common species entries. It already exceeds the BioCreative III average team performance by at least 24% in every evaluation. However, using our proposed dynamic dictionary strategies, we were able to further improve TAP-5, TAP-10, and TAP-20 by 16.47%, 13.57% and 6.01%, respectively in the Gold 50 test set. Our best dynamic strategy outperforms the best BioCreative III systems in TAP-10 on the Silver 50 test set and in TAP-5 on the Silver 507 set. CONCLUSIONS: Our experimental results demonstrate the superiority of our proposed dynamic dictionary selection strategies over our original static strategy and most BioCreative III participant systems. Section-wide dynamic strategy is preferred because it achieves very similar TAP-k scores to article-wide dynamic strategy but it is more efficient. Richard Tzong-Han Tsai, Po-Ting Lai |
BMC Bioinform. | 1 |
| 2011 | Visual webpage block importance prediction using conditional random fieldsabstractAbstract We have developed a system that segments web pages into blocks and predicts those blocks' importance (block importance prediction or BIP). First, we use VIPS to partition a page into a tree composed of blocks and then extracts features from each block and labels all leaf nodes. This paper makes two main contributions. Firstly, we are pioneering the formulation of BIP as a sequence tagging task. We employ DFS, which outputs a single sequence for the whole tree in which related sub‐blocks are adjacent. Our second contribution is using the conditional random fields (CRF) model for labeling these sequences. CRF's transition features model correlations between neighboring labels well, and CRF can simultaneously label all blocks in a sequence to find the global optimal solution for the whole sequence, not only the best solution for each block. In our experiments, our CRF‐based system achieves an F1‐measure of 97.41%, which significantly outperforms our ME‐based baseline (95.64%). Lastly, we tested the CRF‐based system using sites which were not covered in the training data. On completely novel sites CRF performed slightly worse than ME. However, when given only two training pages from a given site, CRF improved almost three times as much as ME. Richard Tzong-Han Tsai, Borong Chiu, Chi-En Wu |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2010 | Chinese text segmentation: A hybrid approach using transductive learning and statistical association measures
Richard Tzong-Han Tsai |
Expert Syst. Appl. | 1 |
| 2010 | Using Contextual Information to Clarify Cross-Species Gene Normalization AmbiguityabstractThe goal of Gene Normalization (GN) is to identify the unique database IDs of genes and proteins mentioned in biomedical literature. A major difficulty in GN comes from the ambiguity of gene names. That is, the same gene name can refer to different database IDs depending on the species in question. In this paper, we introduce a method to exploit contextual information in an abstract, like tissue type, chromosome location, etc., to tackle this problem. Using this technique, we have been able to improve system performance (F-score) by 14.3% on the BioCreAtIvE-II GN task test set. We also examined our method on a full-text dataset with cross-species genes. The experimental results show a promising performance (AUC) of 42.94%. Our experimental results also show that with full text, versus abstract only, the system performance was 12.24% higher. Richard Tzong-Han Tsai, Po-Ting Lai |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2010 | Multistage Gene Normalization and SVM-Based Ranking for Protein Interactor Extraction in Full-Text ArticlesabstractThe interactor normalization task (INT) is to identify genes that play the interactor role in protein-protein interactions (PPIs), to map these genes to unique IDs, and to rank them according to their normalized confidence. INT has two subtasks: gene normalization (GN) and interactor ranking. The main difficulties of INT GN are identifying genes across species and using full papers instead of abstracts. To tackle these problems, we developed a multistage GN algorithm and a ranking method, which exploit information in different parts of a paper. Our system achieved a promising AUC of 0.43471. Using the multistage GN algorithm, we have been able to improve system performance (AUC) by 1.719 percent compared to a one-stage GN algorithm. Our experimental results also show that with full text, versus abstract only, INT AUC performance was 22.6 percent higher. Hong-Jie Dai, Po-Ting Lai, Richard Tzong-Han Tsai |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2009 | WikiSense: Supersense Tagging of Wikipedia Named Entities Based WordNet
Joseph Z. Chang, Richard Tzong-Han Tsai, Jason S. Chang |
PACLIC | 2 |
| 2009 | Modeling the Relationship among Linguistic Typological Features with Hierarchical Dirichlet Process
Chu-Cheng Lin, Yu-Chun Wang, Richard Tzong-Han Tsai |
PACLIC | 3 |
| 2009 | Rule-based Korean Grapheme to Phoneme Conversion Using Sound Patterns
Yu-Chun Wang, Richard Tzong-Han Tsai |
PACLIC | 2 |
| 2009 | PubMed-EX: a web browser extension to enhance PubMed search with text mining featuresabstractAbstract Summary: PubMed-EX is a browser extension that marks up PubMed search results with additional text-mining information. PubMed-EX's page mark-up, which includes section categorization and gene/disease and relation mark-up, can help researchers to quickly focus on key terms and provide additional information on them. All text processing is performed server-side, freeing up user resources. Availability: PubMed-EX is freely available at http://bws.iis.sinica.edu.tw/PubMed-EX and http://iisr.cse.yzu.edu.tw:8000/PubMed-EX/. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Richard Tzong-Han Tsai, Hong-Jie Dai, Po-Ting Lai, Chi-Hsin Huang |
Bioinform. | 1 |
| 2009 | HypertenGene: extracting key hypertension genes from biomedical literature with position and automatically-generated template featuresabstractBACKGROUND: The genetic factors leading to hypertension have been extensively studied, and large numbers of research papers have been published on the subject. One of hypertension researchers' primary research tasks is to locate key hypertension-related genes in abstracts. However, gathering such information with existing tools is not easy: (1) Searching for articles often returns far too many hits to browse through. (2) The search results do not highlight the hypertension-related genes discovered in the abstract. (3) Even though some text mining services mark up gene names in the abstract, the key genes investigated in a paper are still not distinguished from other genes. To facilitate the information gathering process for hypertension researchers, one solution would be to extract the key hypertension-related genes in each abstract. Three major tasks are involved in the construction of this system: (1) gene and hypertension named entity recognition, (2) section categorization, and (3) gene-hypertension relation extraction. RESULTS: We first compare the retrieval performance achieved by individually adding template features and position features to the baseline system. Then, the combination of both is examined. We found that using position features can almost double the original AUC score (0.8140 vs.0.4936) of the baseline system. However, adding template features only results in marginal improvement (0.0197). Including both improves AUC to 0.8184, indicating that these two sets of features are complementary, and do not have overlapping effects. We then examine the performance in a different domain--diabetes, and the result shows a satisfactory AUC of 0.83. CONCLUSION: Our approach successfully exploits template features to recognize true hypertension-related gene mentions and position features to distinguish key genes from other related genes. Templates are automatically generated and checked by biologists to minimize labor costs. Our approach integrates the advantages of machine learning models and pattern matching. To the best of our knowledge, this the first systematic study of extracting hypertension-related genes and the first attempt to create a hypertension-gene relation corpus based on the GAD database. Furthermore, our paper proposes and tests novel features for extracting key hypertension genes, such as relative position, section, and template features, which could also be applied to key-gene extraction for other diseases. Richard Tzong-Han Tsai, Po-Ting Lai, Hong-Jie Dai, Chi-Hsin Huang, Yue-Yang Bow, Yen-Ching Chang, Wen-Harn Pan, Wen-Lian Hsu |
BMC Bioinform. | 1 |
| 2009 | Learning weights for translation candidates in Japanese-Chinese information retrieval
Chu-Cheng Lin, Yu-Chun Wang, Chih-Hao Yeh, Wei-Chi Tsai, Richard Tzong-Han Tsai |
Expert Syst. Appl. | 5 |
| 2009 | Web-based pattern learning for named entity translation in Korean-Chinese cross-language information retrieval
Yu-Chun Wang, Richard Tzong-Han Tsai, Wen-Lian Hsu |
Expert Syst. Appl. | 2 |
| 2009 | New Challenges for Biological Text-Mining in the Next Decade
Hong-Jie Dai, Yen-Ching Chang, Richard Tzong-Han Tsai, Wen-Lian Hsu |
J. Comput. Sci. Technol. | 3 |
| 2008 | Exploiting Unlabeled Text to Extract New Words of Different Semantic Transparency for Chinese Word Segmentation
Richard Tzong-Han Tsai, Hsi-Chuan Hung |
IJCNLP | 1 |
| 2008 | Learning Patterns from the Web to Translate Named Entities for Cross Language Information Retrieval
Yu-Chun Wang, Richard Tzong-Han Tsai, Wen-Lian Hsu |
IJCNLP | 2 |
| 2008 | Semi-automatic conversion of BioProp semantic annotation to PASBio annotationabstractBACKGROUND: Semantic role labeling (SRL) is an important text analysis technique. In SRL, sentences are represented by one or more predicate-argument structures (PAS). Each PAS is composed of a predicate (verb) and several arguments (noun phrases, adverbial phrases, etc.) with different semantic roles, including main arguments (agent or patient) as well as adjunct arguments (time, manner, or location). PropBank is the most widely used PAS corpus and annotation format in the newswire domain. In the biomedical field, however, more detailed and restrictive PAS annotation formats such as PASBio are popular. Unfortunately, due to the lack of an annotated PASBio corpus, no publicly available machine-learning (ML) based SRL systems based on PASBio have been developed. In previous work, we constructed a biomedical corpus based on the PropBank standard called BioProp, on which we developed an ML-based SRL system, BIOSMILE. In this paper, we aim to build a system to convert BIOSMILE's BioProp annotation output to PASBio annotation. Our system consists of BIOSMILE in combination with a BioProp-PASBio rule-based converter, and an additional semi-automatic rule generator. RESULTS: Our first experiment evaluated our rule-based converter's performance independently from BIOSMILE performance. The converter achieved an F-score of 85.29%. The second experiment evaluated combined system (BIOSMILE + rule-based converter). The system achieved an F-score of 69.08% for PASBio's 29 verbs. CONCLUSION: Our approach allows PAS conversion between BioProp and PASBio annotation using BIOSMILE alongside our newly developed semi-automatic rule generator and rule-based converter. Our system can match the performance of other state-of-the-art domain-specific ML-based SRL systems and can be easily customized for PASBio application development. Richard Tzong-Han Tsai, Hong-Jie Dai, Chi-Hsin Huang, Wen-Lian Hsu |
BMC Bioinform. | 1 |
| 2008 | Exploiting likely-positive and unlabeled data to improve the identification of protein-protein interaction articlesabstractBACKGROUND: Experimentally verified protein-protein interactions (PPI) cannot be easily retrieved by researchers unless they are stored in PPI databases. The curation of such databases can be made faster by ranking newly-published articles' relevance to PPI, a task which we approach here by designing a machine-learning-based PPI classifier. All classifiers require labeled data, and the more labeled data available, the more reliable they become. Although many PPI databases with large numbers of labeled articles are available, incorporating these databases into the base training data may actually reduce classification performance since the supplementary databases may not annotate exactly the same PPI types as the base training data. Our first goal in this paper is to find a method of selecting likely positive data from such supplementary databases. Only extracting likely positive data, however, will bias the classification model unless sufficient negative data is also added. Unfortunately, negative data is very hard to obtain because there are no resources that compile such information. Therefore, our second aim is to select such negative data from unlabeled PubMed data. Thirdly, we explore how to exploit these likely positive and negative data. And lastly, we look at the somewhat unrelated question of which term-weighting scheme is most effective for identifying PPI-related articles. RESULTS: To evaluate the performance of our PPI text classifier, we conducted experiments based on the BioCreAtIvE-II IAS dataset. Our results show that adding likely-labeled data generally increases AUC by 3~6%, indicating better ranking ability. Our experiments also show that our newly-proposed term-weighting scheme has the highest AUC among all common weighting schemes. Our final model achieves an F-measure and AUC 2.9% and 5.0% higher than those of the top-ranking system in the IAS challenge. CONCLUSION: Our experiments demonstrate the effectiveness of integrating unlabeled and likely labeled data to augment a PPI text classification system. Our mixed model is suitable for ranking purposes whereas our hierarchical model is better for filtering. In addition, our results indicate that supervised weighting schemes outperform unsupervised ones. Our newly-proposed weighting scheme, TFBRF, which considers documents that do not contain the target word, avoids some of the biases found in traditional weighting schemes. Our experiment results show TFBRF to be the most effective among several other top weighting schemes. Richard Tzong-Han Tsai, Hsi-Chuan Hung, Hong-Jie Dai, Jaimie Yi-Wen Lin, Wen-Lian Hsu |
BMC Bioinform. | 1 |
| 2008 | Web taxonomy integration with hierarchical shrinkage algorithm and fine-grained relations
Chia-Wei Wu, Richard Tzong-Han Tsai, Cheng-Wei Lee 0001, Wen-Lian Hsu |
Expert Syst. Appl. | 2 |
| 2007 | Identifying Protein Interaction Abstracts with Contextual Bag of Words
Hsieh-Chuan Hung, Richard Tzong-Han Tsai, Wen-Lian Hsu |
AAAI | 2 |
| 2007 | Exploiting unlabeled internal data in conditional random fields to reduce word segmentation errors for Chinese textsabstractThe application of text-to-speech (TTS) conversion has become widely used in recent years. Chinese TTS faces several unique difficulties. The most critical is caused by the lack of word delimiters in written Chinese. This means that Chinese word segmentation (CWS) must be the first step in Chinese TTS. Unfortunately, due to the ambiguous nature of word boundaries in Chinese, even the best CWS systems make serious segmentation errors. Incorrect sentence interpretation causes TTS errors, preventing TTS’s wider use in applications such as automatic customer services or computer reader systems for the visually impaired. In this paper, we propose a novel method that exploits unlabeled internal data to reduce word segmentation errors without using external dictionaries. To demonstrate the generality of our method, we verify our system on the most widely recognized CWS evaluation tool--the SIGHAN bakeoff, which includes datasets in both traditional and simplified Chinese. These datasets are provided by four representative academies or industrial research institutes in HK, Taiwan, Mainland China, and the U.S. Our experimental results show that with only internal data and unlabeled test data, our approach reduces segmentation errors by an average of 15 % compared to the traditional approach. Moreover, our approach achieves comparable performance to the best CWS systems that use external resources. Further analysis shows that our method has the potential to become more accurate as the amount of test data increases. Index Terms: text-to-speech, Chinese word segmentation, segmentation errors, internal unlabeled data Richard Tzong-Han Tsai, Hsi-Chuan Hung, Hong-Jie Dai, Wen-Lian Hsu |
INTERSPEECH | 1 |
| 2007 | Korean-Chinese Person Name Translation for Cross Language Information Retrieval
Yu-Chun Wang, Yi-Hsun Lee, Chu-Cheng Lin, Richard Tzong-Han Tsai, Wen-Lian Hsu |
PACLIC | 4 |
| 2007 | BIOSMILE: A semantic role labeling system for biomedical verbs using a maximum-entropy model with automatically generated template featuresabstractBACKGROUND: Bioinformatics tools for automatic processing of biomedical literature are invaluable for both the design and interpretation of large-scale experiments. Many information extraction (IE) systems that incorporate natural language processing (NLP) techniques have thus been developed for use in the biomedical field. A key IE task in this field is the extraction of biomedical relations, such as protein-protein and gene-disease interactions. However, most biomedical relation extraction systems usually ignore adverbial and prepositional phrases and words identifying location, manner, timing, and condition, which are essential for describing biomedical relations. Semantic role labeling (SRL) is a natural language processing technique that identifies the semantic roles of these words or phrases in sentences and expresses them as predicate-argument structures. We construct a biomedical SRL system called BIOSMILE that uses a maximum entropy (ME) machine-learning model to extract biomedical relations. BIOSMILE is trained on BioProp, our semi-automatic, annotated biomedical proposition bank. Currently, we are focusing on 30 biomedical verbs that are frequently used or considered important for describing molecular events. RESULTS: To evaluate the performance of BIOSMILE, we conducted two experiments to (1) compare the performance of SRL systems trained on newswire and biomedical corpora; and (2) examine the effects of using biomedical-specific features. The experimental results show that using BioProp improves the F-score of the SRL system by 21.45% over an SRL system that uses a newswire corpus. It is noteworthy that adding automatically generated template features improves the overall F-score by a further 0.52%. Specifically, ArgM-LOC, ArgM-MNR, and Arg2 achieve statistically significant performance improvements of 3.33%, 2.27%, and 1.44%, respectively. CONCLUSION: We demonstrate the necessity of using a biomedical proposition bank for training SRL systems in the biomedical domain. Besides the different characteristics of biomedical and newswire sentences, factors such as cross-domain framesets and verb usage variations also influence the performance of SRL systems. For argument classification, we find that NE (named entity) features indicating if the target node matches with NEs are not effective, since NEs may match with a node of the parsing tree that does not have semantic role labels in the training set. We therefore incorporate templates composed of specific words, NE types, and POS tags into the SRL system. As a result, the classification accuracy for adjunct arguments, which is especially important for biomedical SRL, is improved significantly. Richard Tzong-Han Tsai, Wen-Chi Chou, Ying-Shan Su, Yu-Chun Lin, Cheng-Lung Sung, Hong-Jie Dai, Irene Tzu-Hsuan Yeh, Wei Ku, Ting-Yi Sung, Wen-Lian Hsu |
BMC Bioinform. | 1 |
| 2007 | Reference metadata extraction using a hierarchical knowledge representation framework
Min-Yuh Day, Richard Tzong-Han Tsai, Cheng-Lung Sung, Chiu-Chen Hsieh, Cheng-Wei Lee 0001, Shih-Hung Wu, Kuen-Pin Wu, Chorng-Shyong Ong, Wen-Lian Hsu |
Decis. Support Syst. | 2 |
| 2006 | A Hybrid Approach to Biomedical Named Entity Recognition and Semantic Role Labeling
Richard Tzong-Han Tsai |
HLT-NAACL | 1 |
| 2006 | NERBio: using selected word conjunctions, term normalization, and global patterns to improve biomedical named entity recognitionabstractBACKGROUND: Biomedical named entity recognition (Bio-NER) is a challenging problem because, in general, biomedical named entities of the same category (e.g., proteins and genes) do not follow one standard nomenclature. They have many irregularities and sometimes appear in ambiguous contexts. In recent years, machine-learning (ML) approaches have become increasingly common and now represent the cutting edge of Bio-NER technology. This paper addresses three problems faced by ML-based Bio-NER systems. First, most ML approaches usually employ singleton features that comprise one linguistic property (e.g., the current word is capitalized) and at least one class tag (e.g., B-protein, the beginning of a protein name). However, such features may be insufficient in cases where multiple properties must be considered. Adding conjunction features that contain multiple properties can be beneficial, but it would be infeasible to include all conjunction features in an NER model since memory resources are limited and some features are ineffective. To resolve the problem, we use a sequential forward search algorithm to select an effective set of features. Second, variations in the numerical parts of biomedical terms (e.g., "2" in the biomedical term IL2) cause data sparseness and generate many redundant features. In this case, we apply numerical normalization, which solves the problem by replacing all numerals in a term with one representative numeral to help classify named entities. Third, the assignment of NE tags does not depend solely on the target word's closest neighbors, but may depend on words outside the context window (e.g., a context window of five consists of the current word plus two preceding and two subsequent words). We use global patterns generated by the Smith-Waterman local alignment algorithm to identify such structures and modify the results of our ML-based tagger. This is called pattern-based post-processing. RESULTS: To develop our ML-based Bio-NER system, we employ conditional random fields, which have performed effectively in several well-known tasks, as our underlying ML model. Adding selected conjunction features, applying numerical normalization, and employing pattern-based post-processing improve the F-scores by 1.67%, 1.04%, and 0.57%, respectively. The combined increase of 3.28% yields a total score of 72.98%, which is better than the baseline system that only uses singleton features. CONCLUSION: We demonstrate the benefits of using the sequential forward search algorithm to select effective conjunction feature groups. In addition, we show that numerical normalization can effectively reduce the number of redundant and unseen features. Furthermore, the Smith-Waterman local alignment algorithm can help ML-based Bio-NER deal with difficult cases that need longer context windows. Richard Tzong-Han Tsai, Cheng-Lung Sung, Hong-Jie Dai, Hsieh-Chuan Hung, Ting-Yi Sung, Wen-Lian Hsu |
BMC Bioinform. | 1 |
| 2006 | Various criteria in the evaluation of biomedical named entity recognitionabstractBACKGROUND: Text mining in the biomedical domain is receiving increasing attention. A key component of this process is named entity recognition (NER). Generally speaking, two annotated corpora, GENIA and GENETAG, are most frequently used for training and testing biomedical named entity recognition (Bio-NER) systems. JNLPBA and BioCreAtIvE are two major Bio-NER tasks using these corpora. Both tasks take different approaches to corpus annotation and use different matching criteria to evaluate system performance. This paper details these differences and describes alternative criteria. We then examine the impact of different criteria and annotation schemes on system performance by retesting systems participated in the above two tasks. RESULTS: To analyze the difference between JNLPBA's and BioCreAtIvE's evaluation, we conduct Experiment 1 to evaluate the top four JNLPBA systems using BioCreAtIvE's classification scheme. We then compare them with the top four BioCreAtIvE systems. Among them, three systems participated in both tasks, and each has an F-score lower on JNLPBA than on BioCreAtIvE. In Experiment 2, we apply hypothesis testing and correlation coefficient to find alternatives to BioCreAtIvE's evaluation scheme. It shows that right-match and left-match criteria have no significant difference with BioCreAtIvE. In Experiment 3, we propose a customized relaxed-match criterion that uses right match and merges JNLPBA's five NE classes into two, which achieves an F-score of 81.5%. In Experiment 4, we evaluate a range of five matching criteria from loose to strict on the top JNLPBA system and examine the percentage of false negatives. Our experiment gives the relative change in precision, recall and F-score as matching criteria are relaxed. CONCLUSION: In many applications, biomedical NEs could have several acceptable tags, which might just differ in their left or right boundaries. However, most corpora annotate only one of them. In our experiment, we found that right match and left match can be appropriate alternatives to JNLPBA and BioCreAtIvE's matching criteria. In addition, our relaxed-match criterion demonstrates that users can define their own relaxed criteria that correspond more realistically to their application requirements. Richard Tzong-Han Tsai, Shih-Hung Wu, Wen-Chi Chou, Yu-Chun Lin, Jieh Hsiang, Ting-Yi Sung, Wen-Lian Hsu |
BMC Bioinform. | 1 |
| 2006 | Integrating linguistic knowledge into a conditional random fieldframework to identify biomedical named entities
Richard Tzong-Han Tsai, Wen-Chi Chou, Shih-Hung Wu, Ting-Yi Sung, Jieh Hsiang, Wen-Lian Hsu |
Expert Syst. Appl. | 1 |
| 2005 | Exploiting Full Parsing Information to Label Semantic Roles Using an Ensemble of ME and SVM via Integer Linear Programming
Richard Tzong-Han Tsai, Chia-Wei Wu, Yu-Chun Lin, Wen-Lian Hsu |
CoNLL | 1 |
| 2002 | Exploiting Knowledge Representation in an Intelligent Tutoring System for English Lexical ErrorsabstractIntelligent tutoring systems (ITSs) construction requires lots of domain knowledge created by hand. In this paper we attempt to illustrate a central role that knowledge representation can play in automating ITS design and implementation. We propose a diagnosis, interaction and treatment (DIT) model for ITS. The entire system relies upon a knowledge representation system (InfoMap) whose structured encoding of English lexical information makes it possible to (1) initiate the relevant dialogue with learners when the system is not sure about learners' intentions, (2) trigger the appropriate lexical knowledge based on their responses, and (3) automatically generate practice and test exercises based on this knowledge. Chiu-Chen Hsieh, Richard Tzong-Han Tsai, David Wible, Wen-Lian Hsu |
ICCE | 2 |