Rey-Long Liu

dblp:55/2083 · DBLP profile ↗
← Back
35ranked-venue papers
34as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 20 first-authorDatabases, data management, data science and information retrieval · 20 · 19 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Data mining · 100%
Artificial intelligence
2 papers
Knowledge representation and reasoning · 33% Trustworthy machine learning · 29% Language models and text generation · 29%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › text mining
text classification
0.012002
Incremental context mining for adaptive document classification · KDD 2002
Machine learning › Trustworthy machine learning › interpretability
explanation-based learning
0.011992
Augmenting and Efficiently Utilizing Domain Theory in Explanation-Based Natural Language Acquisition · ML 1992
Natural language and speech › Language models and text generation
language acquisition
0.011992
Augmenting and Efficiently Utilizing Domain Theory in Explanation-Based Natural Language Acquisition · ML 1992
Natural language and speech › Information extraction and text analysis
semantic role labeling
0.011993
An Empirical Study on Thematic Knowledge Acquisition Based on Syntactic Clues and Heuristics · ACL 1993

Methods — techniques the papers use, named apart from their topics

heuristic ambiguity resolution · 0.0corpus-based validation · 0.0explanation-based learning · 0.0domain theory · 0.0
YearPublicationVenuePosition
2019 Identification of Conclusive Association Entities by Biomedical Association Mining
Rey-Long Liu
ACIIDS (1)1
2017 Identification of Biomedical Articles with Highly Related Core Contents
Rey-Long Liu
ACIIDS (1)1
2016 Citation-Based Extraction of Core Contents from Biomedical Articles
Rey-Long Liu
IEA/AIE1
2015 Retrieval of Highly Related Biomedical References by Key Passages of Citations
Rey-Long Liu
IEA/AIE1
2014 Improving Health Question Classification by Word Location Weights
Rey-Long Liu
ACIIDS (1)1
2014 Identification of highly related references about gene-disease association
abstract
BACKGROUND: Curation of gene-disease associations published in literature should be based on careful and frequent survey of the references that are highly related to specific gene-disease associations. Retrieval of the references is thus essential for timely and complete curation. RESULTS: We present a technique CRFref (Conclusive, Rich, and Focused References) that, given a gene-disease pair < g, d>, ranks high those biomedical references that are likely to provide conclusive, rich, and focused results about g and d. Such references are expected to be highly related to the association between g and d. CRFref ranks candidate references based on their scores. To estimate the score of a reference r, CRFref estimates and integrates three measures: degree of conclusiveness, degree of richness, and degree of focus of r with respect to < g, d>. To evaluate CRFref, experiments are conducted on over one hundred thousand references for over one thousand gene-disease pairs. Experimental results show that CRFref performs significantly better than several typical types of baselines in ranking high those references that expert curators select to develop the summaries for specific gene-disease associations. CONCLUSION: CRFref is a good technique to rank high those references that are highly related to specific gene-disease associations. It can be incorporated into existing search engines to prioritize biomedical references for curators and researchers, as well as those text mining systems that aim at the study of gene-disease associations.
Rey-Long Liu, Chia Chun Shih
BMC Bioinform.1
2013 Reduction of Training Noises for Text Classifiers
Rey-Long Liu
ACIIDS (2)1
2013 A passage extractor for classification of disease aspect information
abstract
Retrieval of disease information is often based on several key aspects such as etiology, diagnosis, treatment, prevention, and symptoms of diseases. Automatic identification of disease aspect information is thus essential. In this article, I model the aspect identification problem as a text classification (TC) problem in which a disease aspect corresponds to a category. The disease aspect classification problem poses two challenges to classifiers: (a) a medical text often contains information about multiple aspects of a disease and hence produces noise for the classifiers and (b) text classifiers often cannot extract the textual parts (i.e., passages) about the categories of interest. I thus develop a technique, PETC (Passage Extractor for Text Classification), that extracts passages (from medical texts) for the underlying text classifiers to classify. Case studies on thousands of Chinese and English medical texts show that PETC enhances a support vector machine (SVM) classifier in classifying disease aspect information. PETC also performs better than three state‐of‐the‐art classifier enhancement techniques, including two passage extraction techniques for text classifiers and a technique that employs term proximity information to enhance text classifiers. The contribution is of significance to evidence‐based medicine, health education, and healthcare decision support. PETC can be used in those application domains in which a text to be classified may have several parts about different categories.
Rey-Long Liu
J. Assoc. Inf. Sci. Technol.1
2011 Identifying Disease Diagnosis Factors by Proximity-Based Mining of Medical Texts
Rey-Long Liu, Shu-Yu Tung, Yun-Ling Lu
ACIIDS (2)1
2011 Medical query generation by term-category correlation
Rey-Long Liu, Yi-Chih Huang
Inf. Process. Manag.1
2011 Ranker enhancement for proximity-based ranking of biomedical texts
abstract
Biomedical decision making often requires relevant evidence from the biomedical literature. Retrieval of the evidence calls for a system that receives a natural language query for a biomedical information need and, among the huge amount of texts retrieved for the query, ranks relevant texts higher for further processing. However, state-of-the-art text rankers have weaknesses in dealing with biomedical queries, which often consist of several correlating concepts and prefer those texts that completely talk about the concepts. In this article, we present a technique, Proximity-Based Ranker Enhancer (PRE), to enhance text rankers by term-proximity information. PRE assesses the term frequency (TF) of each term in the text by integrating three types of term proximity to measure the contextual completeness of query terms appearing in nearby areas in the text being ranked. Therefore, PRE may serve as a preprocessor for (or supplement to) those rankers that consider TF in ranking, without the need to change the algorithms and development processes of the rankers. Empirical evaluation shows that PRE significantly improves various kinds of text rankers, and when compared with several state-of-the-art techniques that enhance rankers by term-proximity information, PRE may more stably and significantly enhance the rankers.
Rey-Long Liu, Yi-Chih Huang
J. Assoc. Inf. Sci. Technol.1
2010 Context-based online medical terminology navigation
Rey-Long Liu, Yun-Ling Lu
Expert Syst. Appl.1
2010 Context-based term frequency assessment for text classification
abstract
Abstract Automatic text classification (TC) is essential for the management of information. To properly classify a document d, it is essential to identify the semantics of each term t in d, while the semantics heavily depend on context (neighboring terms) of t in d. Therefore, we present a technique CTFA (Context‐based Term Frequency Assessment) that improves text classifiers by considering term contexts in test documents. The results of the term context recognition are used to assess term frequencies of terms, and hence CTFA may easily work with various kinds of text classifiers that base their TC decisions on term frequencies, without needing to modify the classifiers. Moreover, CTFA is efficient, and neither huge memory nor domain‐specific knowledge is required. Empirical results show that CTFA successfully enhances performance of several kinds of text classifiers on different experimental data.
Rey-Long Liu
J. Assoc. Inf. Sci. Technol.1
2009 Online assessment of content skill levels for medical texts
Rey-Long Liu, Yun-Ling Lu
Expert Syst. Appl.1
2009 Context recognition for hierarchical text classification
abstract
Abstract Information is often organized as a text hierarchy. A hierarchical text‐classification system is thus essential for the management, sharing, and dissemination of information. It aims to automatically classify each incoming document into zero, one, or several categories in the text hierarchy. In this paper, we present a technique called CRHTC (context recognition for hierarchical text classification) that performs hierarchical text classification by recognizing the context of discussion (COD) of each category. A category's COD is governed by its ancestor categories, whose contents indicate contextual backgrounds of the category. A document may be classified into a category only if its content matches the category's COD. CRHTC does not require any trials to manually set parameters, and hence is more portable and easier to implement than other methods. It is empirically evaluated under various conditions. The results show that CRHTC achieves both better and more stable performance than several hierarchical and nonhierarchical text‐classification methodologies.
Rey-Long Liu
J. Assoc. Inf. Sci. Technol.1
2008 Context-Based Term Frequency Assessment for Text Classification
Rey-Long Liu
PRICAI1
2008 Interactive high-quality text classification
Rey-Long Liu
Inf. Process. Manag.1
2007 Text Classification for Healthcare Information Support
Rey-Long Liu
IEA/AIE1
2007 Dynamic category profiling for text filtering and classification
Rey-Long Liu
Inf. Process. Manag.1
2006 Dynamic Category Profiling for Text Filtering and Classification
Rey-Long Liu
PAKDD1
2005 Adaptive sampling for thresholding in document filtering and classification
Rey-Long Liu, Wan-Jung Lin
Inf. Process. Manag.1
2005 Incremental mining of information interest for personalized web scanning
Rey-Long Liu, Wan-Jung Lin
Inf. Syst.1
2004 Collaborative Multiagent Adaptation for Business Environmental Scanning Through the Internet
Rey-Long Liu
Appl. Intell.1
2004 Erratum to: "Mining for interactive identification of users? information needs": [Information Systems 28 (2003) 815-833]
Rey-Long Liu, Wan-Jung Lin
Inf. Syst.1
2003 Distributed agents for cost-effective monitoring of critical success factors
Rey-Long Liu, Yun-Ling Lu
Decis. Support Syst.1
2003 Adaptive Agents for Effective Information Monitoring
abstract
Information motoring is an essential basis for management and decision making in various domains. Once an information update is detected, suitable procedures may be triggered to identify potential problems and opportunities. In this paper, we define effective information monitoring (EIM) and explore how adaptive agents may cooperate with each other in order to achieve their collective goal of EIM. Since each information item may be updated by various entities (e.g. information servers) at any time, EIM calls for a multiagent system that may detect more information updates in a timely manner using a controlled amount of system resources (e.g. loading of related information servers and the Intranet). To achieve EIM, the agents should adapt themselves to the dynamically changing behaviors of the information items being monitored. They learn to issue requests of using system resources and concede to those agents that are more likely to detect updates at the time of negotiation. The framework is theoretically and empirically evaluated. Its potential applications to management by exceptions are identified and discussed.
Rey-Long Liu
Int. J. Cooperative Inf. Syst.1
2003 Mining for interactive identification of users' information needs
Rey-Long Liu, Wan-Jung Lin
Inf. Syst.1
2002 Incremental context mining for adaptive document classification
abstract
Automatic document classification (DC) is essential for the management of information and knowledge. This paper explores two practical issues in DC: (1) each document has its context of discussion, and (2) both the content and vocabulary of the document database is intrinsically evolving. The issues call for adaptive document classification (ADC) that adapts a DC system to the evolving contextual requirement of each document category, so that input documents may be classified based on their contexts of discussion. We present an incremental context mining technique to tackle the challenges of ADC. Theoretical analyses and empirical results show that, given a text hierarchy, the mining technique is efficient in incrementally maintaining the evolving contextual requirement of each category. Based on the contextual requirements mined by the system, higher-precision DC may be achieved with better efficiency.
Rey-Long Liu, Yun-Ling Lu
KDD1
2001 An Adaptive Agent Society for Environmental Scanning through the Internet
Rey-Long Liu
PRIMA1
1995 Interactive acquisition of thematic information of Chinese verbs for judicial verdict document understanding using templates, syntactic clues, and heuristics
abstract
The thematic knowledge can bridge the gap between semantic entities and syntactic constituents. In document understanding, the correctness and the efficiency could be improved if the thematic knowledge is available. In this paper, we propose a semi-automatic method to acquire thematic knowledge of Chinese verbs by exploiting syntactic clues. The syntactic clues, which may be collected by most existing syntactic processors, reduce the hypothesis space of the theta roles. The ambiguities may be further resolved by the evidences from a trainer. A set of heuristics based on linguistic constraints are employed to guide the ambiguity resolution process. To acquire thematic information for verbs, the argument structures of the verbs must be extracted first. A template matching method is used to extract the argument structure of verbs.
Koong H. C. Lin, Rey-Long Liu, Von-Wun Soo
ICDAR2
1994 A Corpus-Based Learning Technique for Building A Self-Extensible Parser
Rey-Long Liu, Von-Wun Soo
COLING1
1993 An Empirical Study on Thematic Knowledge Acquisition Based on Syntactic Clues and Heuristics
abstract
Thematic knowledge is a basis of semantic interpretation. In this paper, we propose an acquisition method to acquire thematic knowledge by exploiting syntactic clues from training sentences. The syntactic clues, which may be easily collected by most existing syntactic processors, reduce the hypothesis space of the thematic roles. The ambiguities may be further resolved by the evidences either from a trainer or from a large corpus. A set of heuristics based on linguistic constraints is employed to guide the ambiguity resolution process. When a trainer is available, the system generates new sentences whose thematic validities can be justified by the trainer. When a large corpus is available, the thematic validity may be justified by observing the sentences in the corpus. Using this way, a syntactic processor may become a thematic recognizer by simply deriving its thematic knowledge from its own syntactic knowledge.
Rey-Long Liu, Von-Wun Soo
ACL1
1993 Parsing-Driven Generalization for Natural Language Acquisition
abstract
Parsing is an important step in natural language processing. It involves tasks of searching for applicable grammatical rules which can transform natural language sentences into their corresponding parse trees. Therefore parsing can be viewed as problem solving, and language acquisition can be achieved by generalizing problem solving heuristics. In this paper we investigate how machine learning methodologies can be integrated with a Wait-And-See Parser (the problem solver) to acquire parsing-related knowledge that is needed for the parser. We call this approach parsing-driven generalization since learning (acquisition of parsing rules and classification of lexicons) is basically derived from the parsing process. Three types of generalization are reported in this paper: simple generalization, generalization by asking questions, and generalization back-propagation. Simple generalization generalizes any two parsing rules whose action parts (right-hand sides) are the same but whose condition parts (left-hand sides) have a single difference. Generalization by asking questions is triggered when a “climbing-up” move on a concept hierarchy is attempted. It is necessary for avoiding over-generalization. Generalization back-propagation propagates a confirmed generalization of some later parsing rule back to its precedent rules in a parsing sequence and thus causes them to be generalized as well. It can reduce the number of questions asked by the system. With these three types of generalization and a mechanism for maintaining lexicon classification (the domain concept hierarchy), parsing and learning can interact to utilize and acquire parsing-related knowledge. To promote the practical performance of parsing after learning, a relaxation parsing mechanism is also designed to process unseen sentences.
Rey-Long Liu, Von-Wun Soo
Int. J. Pattern Recognit. Artif. Intell.1
1992 Augmenting and Efficiently Utilizing Domain Theory in Explanation-Based Natural Language Acquisition
Rey-Long Liu, Von-Wun Soo
ML1
1990 Dealing with Ambiguities in English Conjunctions and Comparatives by a Deterministic Parser
abstract
The major problems in parsing English conjunctions and comparatives are ambiguities of scoping and ellipsis. Scoping ambiguities occur when a parser cannot deterministically detect boundaries of constituents, while ellipsis ambiguities occur when a parser cannot deterministically detect missing components. Since simple lookahead mechanisms cannot collect adequate information to resolve these ambiguities, a parsing strategy that only employs such mechanisms will need to backtrack each time it makes incorrect assumptions. In this paper, we extend the Wait-And-See strategy to parse conjunctions and comparatives deterministically and simultaneously. Several mechanisms, such as bottom-up preparsing, suspension, and pattern matching, are implemented. The bottom-up preparsing accesses the dictionary and recognizes isolated sentence fragments which can be determined without ambiguities. The suspension, which is different from Marcus’s attention shifting, allows the parser to suspend temporally at ambiguous points and continue to parse the rest of the sentence until it obtains the necessary information to resolve the ambiguities. Pattern matching uses the concept of symmetry to detect missing components (the ellipses) in the two conjoined or compared sentence fragments.
Rey-Long Liu, Von-Wun Soo
Int. J. Pattern Recognit. Artif. Intell.1