Clement T. Yu

dblp:y/ClementTYu · DBLP profile ↗
← Back
178ranked-venue papers
44as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 141 · 31 first-authorArtificial intelligence and machine learning · 33 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 13 · 6 first-authorSystems, architecture and hardware · 6 · 3 first-authorSoftware engineering, systems software and programming languages · 5 · 3 first-authorTheory of computation · 5 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4Computer networks · 2Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
103 papers
Information retrieval · 52% Data integration and cleaning · 33% Query processing and optimization · 8%
Artificial intelligence
8 papers
Information extraction and text analysis · 94% Face, body and person analysis · 6%
Theoretical computer science
13 papers
Computational complexity · 76% Approximation and online algorithms · 6% Computational geometry · 6%

Topics — the 30 heaviest of 186, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data integration and cleaning
schema matching
0.462010
Deep Web Integration with VisQI · Proc. VLDB Endow. 2010
Stop Word and Related Problems in Web Interface Integration · Proc. VLDB Endow. 2009
WebIQ: Learning from the Web to Match Deep-Web Query Interfaces · ICDE 2006
Natural language and speech › Information extraction and text analysis
sentiment analysis
0.422015
Polarity Consistency Checking for Domain Independent Sentiment Dictionaries · IEEE Trans. Knowl. Data Eng. 2015
Polarity Consistency Checking for Sentiment Dictionaries · ACL (1) 2012
Information retrieval › distributed information retrieval
metasearch
0.372007
MySearchView: a customized metasearch engine generator · SIGMOD Conference 2007
AllInOneNews: development and evaluation of a large-scale news metasearch engine · SIGMOD Conference 2007
SE-LEGO: creating metasearch engines on demand · SIGIR 2003
Data integration and cleaning › data extraction
web data extraction
0.342013
Annotating Search Results from Web Databases · IEEE Trans. Knowl. Data Eng. 2013
Automatic Extraction of Dynamic Record Sections From Search Engine Result Pages · VLDB 2006
Fully automatic wrapper generation for search engines · WWW 2005
Data integration and cleaning › web data integration
deep web integration
0.332013
YumiInt - A deep Web integration system for local search engines for Geo-referenced objects · ICDE 2013
Deep Web Integration with VisQI · Proc. VLDB Endow. 2010
WebIQ: Learning from the Web to Match Deep-Web Query Interfaces · ICDE 2006
Information retrieval › ranking › context-aware ranking
temporal ranking
0.222013
The Impacts of Structural Difference and Temporality of Tweets on Retrieval Effectiveness · ACM Trans. Inf. Syst. 2013
AllInOneNews: development and evaluation of a large-scale news metasearch engine · SIGMOD Conference 2007
Data integration and cleaning › web data integration
web query interface integration
0.232009
Stop Word and Related Problems in Web Interface Integration · Proc. VLDB Endow. 2009
Automatic integration of Web search interfaces with WISE-Integrator · VLDB J. 2004
WISE-Integrator: An Automatic Integrator of Web Search Interfaces for E-Commerce · VLDB 2003
Information retrieval › retrieval models
ad-hoc retrieval
0.212013
The Impacts of Structural Difference and Temporality of Tweets on Retrieval Effectiveness · ACM Trans. Inf. Syst. 2013
Data integration and cleaning
entity resolution
0.212013
YumiInt - A deep Web integration system for local search engines for Geo-referenced objects · ICDE 2013
Information retrieval › web search › web information retrieval › social media retrieval
microblog retrieval
0.212013
The Impacts of Structural Difference and Temporality of Tweets on Retrieval Effectiveness · ACM Trans. Inf. Syst. 2013
Information retrieval › query understanding
query classification
0.212013
The Impacts of Structural Difference and Temporality of Tweets on Retrieval Effectiveness · ACM Trans. Inf. Syst. 2013
Information retrieval › document retrieval
temporal information retrieval
0.212013
The Impacts of Structural Difference and Temporality of Tweets on Retrieval Effectiveness · ACM Trans. Inf. Syst. 2013
Information retrieval
retrieval models
0.2182007
Knowledge-intensive conceptual retrieval and passage extraction of biomedical literature · SIGIR 2007
Efficient and Effective Metasearch for Text Databases Incorporating Linkages among Documents · SIGMOD Conference 2001
Effective keyword search in relational databases · SIGMOD Conference 2006
Information retrieval › search engines
search engine selection
0.242007
AllInOneNews: development and evaluation of a large-scale news metasearch engine · SIGMOD Conference 2007
A Statistical Method for Estimating the Usefulness of Text Databases · IEEE Trans. Knowl. Data Eng. 2002
Towards a highly-scalable and effective metasearch engine · WWW 2001
Information retrieval › distributed information retrieval
resource selection
0.252002
A Methodology to Retrieve Text Documents from Multiple Databases · IEEE Trans. Knowl. Data Eng. 2002
A highly scalable and effective method for metasearch · ACM Trans. Inf. Syst. 2001
Towards a highly-scalable and effective metasearch engine · WWW 2001
Information retrieval › search engines › web crawling
deep web crawling
0.122009
A Hierarchical Approach to Model Web Query Interfaces for Web Source Integration · Proc. VLDB Endow. 2009
Fully automatic wrapper generation for search engines · WWW 2005
Information retrieval
fact-checking
0.112011
T-verifier: Verifying truthfulness of fact statements · ICDE 2011
Data integration and cleaning › schema matching
query interface matching
0.122006
WebIQ: Learning from the Web to Match Deep-Web Query Interfaces · ICDE 2006
An Interactive Clustering-based Approach to Integrating Source Query interfaces on the Deep Web · SIGMOD Conference 2004
Information retrieval
search engines
0.132007
MySearchView: a customized metasearch engine generator · SIGMOD Conference 2007
Automatic integration of Web search interfaces with WISE-Integrator · VLDB J. 2004
SE-LEGO: creating metasearch engines on demand · SIGIR 2003
Data integration and cleaning › data extraction › web data extraction
query interface extraction
0.112009
A Hierarchical Approach to Model Web Query Interfaces for Web Source Integration · Proc. VLDB Endow. 2009
Data integration and cleaning
web data integration
0.112009
A Hierarchical Approach to Model Web Query Interfaces for Web Source Integration · Proc. VLDB Endow. 2009
Information retrieval › search interfaces
web query interface
0.112009
A Hierarchical Approach to Model Web Query Interfaces for Web Source Integration · Proc. VLDB Endow. 2009
Information retrieval
ranking
0.132013
The Impacts of Structural Difference and Temporality of Tweets on Retrieval Effectiveness · ACM Trans. Inf. Syst. 2013
AllInOneNews: development and evaluation of a large-scale news metasearch engine · SIGMOD Conference 2007
Effective keyword search in relational databases · SIGMOD Conference 2006
Information retrieval
distributed information retrieval
0.132002
A Methodology to Retrieve Text Documents from Multiple Databases · IEEE Trans. Knowl. Data Eng. 2002
Efficient and Effective Metasearch for Text Databases Incorporating Linkages among Documents · SIGMOD Conference 2001
Determining Text Databases to Search in the Internet · VLDB 1998
Information retrieval
document retrieval
0.122004
An effective approach to document retrieval via utilizing WordNet and recognizing phrases · SIGIR 2004
A Methodology to Retrieve Text Documents from Multiple Databases · IEEE Trans. Knowl. Data Eng. 2002
Information retrieval
evaluation
0.142007
AllInOneNews: development and evaluation of a large-scale news metasearch engine · SIGMOD Conference 2007
Models of IR (Panel) · SIGIR 1987
Probabilistic Models for Document Retrieval: A Comparison of Performance on Experimental and Synthetic Databases · SIGIR 1986
Natural language and speech › Information extraction and text analysis
web information extraction
0.112007
Mining templates from search result records of search engines · KDD 2007
Information retrieval › search engines › semantic search
conceptual retrieval
0.112007
Knowledge-intensive conceptual retrieval and passage extraction of biomedical literature · SIGIR 2007
Data integration and cleaning › data extraction › web data extraction
deep web data extraction
0.112007
Annotating Structured Data of the Deep Web · ICDE 2007
Data integration and cleaning › data extraction
structured data extraction
0.112007
Mining templates from search result records of search engines · KDD 2007

Methods — techniques the papers use, named apart from their topics

satisfiability reduction · 0.4SAT solver · 0.4approximation algorithm · 0.2wrapper construction · 0.2temporal relevance scoring · 0.2query translation · 0.2divide-and-conquer ranking · 0.2data unit alignment · 0.2annotation aggregation · 0.2consistency checking · 0.1feature extraction from search results · 0.1tree merging · 0.1geometric layout analysis · 0.1simulation · 0.1statistical methods · 0.1passage extraction · 0.1graph model · 0.1concept-based retrieval · 0.1
YearPublicationVenuePosition
2017 Result Merging for Structured Queries on the Deep Web with Active Relevance Weight Estimation
Jing Yuan 0006, Lihong He 0001, Eduard C. Dragut, Weiyi Meng, Clement T. Yu
Inf. Syst.5
2016 Verification of Fact Statements with Multiple Truthful Alternatives
Weiyi Meng, Clement T. Yu
WEBIST (2)3
2015 A context-aware approach to detection of short irrelevant texts
abstract
This paper presents a simple and effective framework that can detect irrelevant short text contents following blogs and news articles, etc. in a context-aware and timely fashion. Nowadays, websites such as Linkedin.com and CNN.com allow their visitors to leave comments after articles, and spammers are exploiting this feature to post irrelevant contents. Visited by millions of readers per day, these websites have extremely high visibility, and irrelevant comments have a detrimental effect on the visiting traffic and revenue of these websites. Therefore, it is critical to eliminate these irrelevant comments as accurately and early as possible. Different from traditional text mining tasks, comments following news and blog articles are characterized by briefness and context-dependent semantics, making it difficult to measure semantic relevance. What's worse, there could be only a handful of comments soon after an article is posted, leading to a severe lack of information for semantics and relevance measurement. We propose to infer “context-aware semantics” to address the above challenges in a unified framework. Specifically, we construct contexts for comments using either blocks of surrounding comments, or comments collected via a principled transfer learning approach. The constructed contexts mitigate the sparseness and sharply define context-dependent semantics of comments, even at the early stage of commenting activities, allowing traditional dimension reduction methods to better capture the semantics of short texts in a context-aware way. We confirm the effectiveness of the proposed method on two real world datasets consisting of news and blog articles and comments, with a maximal improvement of 20% in Area Under Precision-Recall Curve.
Sihong Xie, Jing Wang 0102, Mohammad Shafkat Amin, Baoshi Yan, Anmol Bhasin, Clement T. Yu, Philip S. Yu
DSAA6
2015 Automated confidence ranked classification of randomized controlled trial articles: an aid to evidence-based medicine
abstract
OBJECTIVE: For many literature review tasks, including systematic review (SR) and other aspects of evidence-based medicine, it is important to know whether an article describes a randomized controlled trial (RCT). Current manual annotation is not complete or flexible enough for the SR process. In this work, highly accurate machine learning predictive models were built that include confidence predictions of whether an article is an RCT. MATERIALS AND METHODS: The LibSVM classifier was used with forward selection of potential feature sets on a large human-related subset of MEDLINE to create a classification model requiring only the citation, abstract, and MeSH terms for each article. RESULTS: The model achieved an area under the receiver operating characteristic curve of 0.973 and mean squared error of 0.013 on the held out year 2011 data. Accurate confidence estimates were confirmed on a manually reviewed set of test articles. A second model not requiring MeSH terms was also created, and performs almost as well. DISCUSSION: Both models accurately rank and predict article RCT confidence. Using the model and the manually reviewed samples, it is estimated that about 8000 (3%) additional RCTs can be identified in MEDLINE, and that 5% of articles tagged as RCTs in Medline may not be identified. CONCLUSION: Retagging human-related studies with a continuously valued RCT confidence is potentially more useful for article ranking and review than a simple yes/no prediction. The automated RCT tagging tool should offer significant savings of time and effort during the process of writing SRs, and is a key component of a multistep text mining pipeline that we are building to streamline SR workflow. In addition, the model may be useful for identifying errors in MEDLINE publication types. The RCT confidence predictions described here have been made available to users as a web service with a user query form front end at: http://arrowsmith.psych.uic.edu/cgi-bin/arrowsmith_uic/RCT_Tagger.cgi.
Aaron M. Cohen, Neil R. Smalheiser, Marian McDonagh, Clement T. Yu, Clive E. Adams, John M. Davis, Philip S. Yu
J. Am. Medical Informatics Assoc.4
2015 A Sense-Topic Model for Word Sense Induction with Unsupervised Data Enrichment
abstract
Word sense induction (WSI) seeks to automatically discover the senses of a word in a corpus via unsupervised methods. We propose a sense-topic model for WSI, which treats sense and topic as two separate latent variables to be inferred jointly. Topics are informed by the entire document, while senses are informed by the local context surrounding the ambiguous word. We also discuss unsupervised ways of enriching the original corpus in order to improve model performance, including using neural word embeddings and external corpora to expand the context of each data instance. We demonstrate significant improvements over the previous state-of-the-art, achieving the best results reported to date on the SemEval-2013 WSI task.
Jing Wang 0102, Mohit Bansal, Kevin Gimpel, Brian D. Ziebart, Clement T. Yu
Trans. Assoc. Comput. Linguistics5
2015 Polarity Consistency Checking for Domain Independent Sentiment Dictionaries
abstract
Polarity classification of words is important for applications such as Opinion Mining and Sentiment Analysis. A number of sentiment word/sense dictionaries have been manually or (semi)automatically constructed. We notice that these sentiment dictionaries have numerous inaccuracies. Besides obvious instances, where the same word appears with different polarities in different dictionaries, the dictionaries exhibit complex cases of polarity inconsistency, which cannot be detected by mere manual inspection. We introduce the concept of polarity consistency of words/senses in sentiment dictionaries in this paper. We show that the consistency problem is NP-complete. We reduce the polarity consistency problem to the satisfiability problem and utilize two fast SAT solvers to detect inconsistencies in a sentiment dictionary. We perform experiments on five sentiment dictionaries and WordNet to show interand intra-dictionaries inconsistencies.
Eduard C. Dragut, A. Prasad Sistla, Clement T. Yu, Weiyi Meng
IEEE Trans. Knowl. Data Eng.4
2015 Diversionary Comments under Blog Posts
abstract
There has been a recent swell of interest in the analysis of blog comments. However, much of the work focuses on detecting comment spam in the blogsphere. An important issue that has been neglected so far is the identification of diversionary comments. Diversionary comments are defined as comments that divert the topic from the original post. A possible purpose is to distract readers from the original topic and draw attention to a new topic. We categorize diversionary comments into five types based on our observations and propose an effective framework to identify and flag them. To the best of our knowledge, the problem of detecting diversionary comments has not been studied so far. We solve the problem in two different ways: (i) rank all comments in descending order of being diversionary and (ii) consider it as a classification problem. Our evaluation on 4,179 comments under 40 different blog posts from Digg and Reddit shows that the proposed method achieves the high mean average precision of 91.9% when the problem is considered as a ranking problem and 84.9% of F-measure as a classification problem. Sensitivity analysis indicates that the effectiveness of the method is stable under different parameter settings.
Jing Wang 0102, Clement T. Yu, Philip S. Yu, Bing Liu 0001, Weiyi Meng
ACM Trans. Web2
2014 Merging Query Results From Local Search Engines for Georeferenced Objects
abstract
The emergence of numerous online sources about local services presents a need for more automatic yet accurate data integration techniques. Local services are georeferenced objects and can be queried by their locations on a map, for instance, neighborhoods. Typical local service queries (e.g., “French Restaurant in The Loop”) include not only information about “what” (“French Restaurant”) a user is searching for (such as cuisine) but also “where” information, such as neighborhood (“The Loop”). In this article, we address three key problems: query translation, result merging and ranking. Most local search engines provide a (hierarchical) organization of (large) cities into neighborhoods. A neighborhood in one local search engine may correspond to sets of neighborhoods in other local search engines. These make the query translation challenging. To provide an integrated access to the query results returned by the local search engines, we need to combine the results into a single list of results. Our contributions include: (1) An integration algorithm for neighborhoods. (2) A very effective business listing resolution algorithm. (3) A ranking algorithm that takes into consideration the user criteria, user ratings and rankings. We have created a prototype system, Yumi, over local search engines in the restaurant domain. The restaurant domain is a representative case study for the local services. We conducted a comprehensive experimental study to evaluate Yumi. A prototype version of Yumi is available online.
Eduard C. Dragut, Bhaskar DasGupta, Brian P. Beirne, Ali Neyestani, Badr Atassi, Clement T. Yu, Weiyi Meng
ACM Trans. Web6
2013 Faceted models of blog feeds
abstract
Faceted blog distillation aims at retrieving the blogs that are not only relevant to a query but also exhibit an interested facet. In this paper we consider personal and official facets. Personal blogs depict various topics related to the personal experiences of bloggers while official blogs deliver contents with bloggers' commercial influences. We observe that some terms, such as nouns, usually describe the topics of posts in blogs while other terms, such as pronouns and adverbs, normally reflect the facets of posts. Thus we present a model that estimates the probabilistic distributions of topics and those of facets in posts. It leverages a classifier to separate facet terms from topical terms in the posterior inference. We also observe that the posts from a blog are likely to exhibit the same facet. So we propose another model that constrains the posts from a blog to have the same facet distributions in its generative process. Experimental results using the TREC 2009-2010 queries over the TREC Blogs08 collection show the effectiveness of both models. Our results outperform the best known results for personal and official distillation.
Lifeng Jia, Clement T. Yu, Weiyi Meng
CIKM2
2013 YumiInt - A deep Web integration system for local search engines for Geo-referenced objects
abstract
We present YumiInt a deep Web integration system for local search engines for Geo-referenced objects. YumiInt consists of two systems: YumiDev and YumiMeta. YumiDev is an off-line integration system that builds the key components (e.g., query translation and entity resolution) of YumiMeta. YumiMeta is the Web application to which users post queries. It translates queries to multiple sources and gets back aggregated lists of results. We present the two systems in this paper.
Eduard C. Dragut, Brian P. Beirne, Ali Neyestani, Badr Atassi, Clement T. Yu, Bhaskar DasGupta, Weiyi Meng
ICDE5
2013 Annotating Search Results from Web Databases
abstract
An increasing number of databases have become web accessible through HTML form-based search interfaces. The data units returned from the underlying database are usually encoded into the result pages dynamically for human browsing. For the encoded data units to be machine processable, which is essential for many applications such as deep web data collection and Internet comparison shopping, they need to be extracted out and assigned meaningful labels. In this paper, we present an automatic annotation approach that first aligns the data units on a result page into different groups such that the data in the same group have the same semantic. Then, for each group we annotate it from different aspects and aggregate the different annotations to predict a final annotation label for it. An annotation wrapper for the search site is automatically constructed and can be used to annotate new result pages from the same web database. Our experiments indicate that the proposed approach is highly effective.
Yiyao Lu, Hai He, Hongkun Zhao, Weiyi Meng, Clement T. Yu
IEEE Trans. Knowl. Data Eng.5
2013 The Impacts of Structural Difference and Temporality of Tweets on Retrieval Effectiveness
abstract
To explore the information seeking behaviors in microblogosphere, the microblog track at TREC 2011 introduced a real-time ad-hoc retrieval task that aims at ranking relevant tweets in reverse-chronological order. We study this problem via a two-phase approach: 1) retrieving tweets in an ad-hoc way; 2) utilizing the temporal information of tweets to enhance the retrieval effectiveness of tweets. Tweets can be categorized into two types. One type consists of short messages not containing any URL of a Web page. The other type has at least one URL of a Web page in addition to a short message. These two types of tweets have different structures. In the first phase, to address the structural difference of tweets, we propose a method to rank tweets using the divide-and-conquer strategy. Specifically, we first rank the two types of tweets separately. This produces two rankings, one for each type. Then we merge these two rankings of tweets into one ranking. In the second phase, we first categorize queries into several types by exploring the temporal distributions of their top-retrieved tweets from the first phase; then we calculate the time-related relevance scores of tweets according to the classified types of queries; finally we combine the time scores with the IR scores from the first phase to produce a ranking of tweets. Experimental results achieved by using the TREC 2011 and TREC 2012 queries over the TREC Tweets2011 collection show that: (i) our way of ranking the two types of tweets separately and then merging them together yields better retrieval effectiveness than ranking them simultaneously; (ii) our way of incorporating temporal information into the retrieval process yields further improvements, and (iii) our method compares favorably with state-of-the-art methods in retrieval effectiveness.
Lifeng Jia, Clement T. Yu, Weiyi Meng
ACM Trans. Inf. Syst.2
2012 Polarity Consistency Checking for Sentiment Dictionaries
Eduard C. Dragut, Clement T. Yu, A. Prasad Sistla, Weiyi Meng
ACL (1)3
2012 Diversionary comments under political blog posts
abstract
An important issue that has been neglected so far is the identification of diversionary comments. Diversionary comments under political blog posts are defined as comments that deliberately twist the bloggers' intention and divert the topic to another one. The purpose is to distract readers from the original topic and draw attention to a new topic. Given that political blogs have significant impact on the society, we believe it is imperative to identify such comments. We then categorize diversionary comments into 5 types, and propose an effective technique to rank comments in descending order of being diversionary. To the best of our knowledge, the problem of detecting diversionary comments has not been studied so far. Our evaluation on 2,109 comments under 20 different blog posts from Digg.com shows that the proposed method achieves the high mean average precision (MAP) of 92.6%. Sensitivity analysis indicates that the effectiveness of the method is stable under different parameter settings.
Jing Wang 0102, Clement T. Yu, Philip S. Yu, Bing Liu 0001, Weiyi Meng
CIKM2
2012 Categorizing Search Results Using WordNet and Wikipedia
Reza Taghizadeh Hemayati, Weiyi Meng, Clement T. Yu
WAIM3
2011 T-verifier: Verifying truthfulness of fact statements
abstract
The Web has become the most popular place for people to acquire information. Unfortunately, it is widely recognized that the Web contains a significant amount of untruthful information. As a result, good tools are needed to help Web users determine the truthfulness of certain information. In this paper, we propose a two-step method that aims to determine whether a given statement is truthful, and if it is not, find out the truthful statement most related to the given statement. In the first step, we try to find a small number of alternative statements of the same topic as the given statement and make sure that one of these statements is truthful. In the second step, we identify the truthful statement from the given statement and the alternative statements. Both steps heavily rely on analysing various features extracted from the search results returned by a popular search engine for appropriate queries. Our experimental results show the best variation of the proposed method can achieve a precision of about 90%.
Weiyi Meng, Clement T. Yu
ICDE3
2010 Construction of a sentimental word dictionary
abstract
The Web has plenty of reviews, comments and reports about products, services, government policies, institutions, etc. The opinions expressed in these reviews influence how people regard these entities. For example, a product with consistently good reviews is likely to sell well, while a product with numerous bad reviews is likely to sell poorly. Our aim is to build a sentimental word dictionary, which is larger than existing sentimental word dictionaries and has high accuracy. We introduce rules for deduction, which take words with known polarities as input and produce synsets (a set of synonyms with a definition) with polarities. The synsets with deduced polarities can then be used to further deduce the polarities of other words. Experimental results show that for a given sentimental word dictionary with D words, approximately an additional 50% of D words with polarities can be deduced. An experiment is conducted to find the accuracy of a random sample of the deduced words. It is found that the accuracy is about the same as that of comparing the judgment of one human with that of another.
Eduard C. Dragut, Clement T. Yu, A. Prasad Sistla, Weiyi Meng
CIKM2
2010 Identifying and Ranking Possible Semantic and Common Usage Categories of Search Engine Queries
Reza Taghizadeh Hemayati, Weiyi Meng, Clement T. Yu
WISE3
2010 Deep Web Integration with VisQI
abstract
In this paper, we present VisQI (VISual Query interface Integration system), a Deep Web integration system. VisQI is capable of (1) transforming Web query interfaces into hierarchically structured representations, (2) of classifying them into application domains and (3) of matching the elements of different interfaces. Thus VisQI contains solutions for the major challenges in building Deep Web integration systems. The system comes along with a full-fledged evaluation system that automatically compares generated data structures against a gold standard. VisQI has a framework-like architecture such that other developers can reuse its components easily.
Thomas Kabisch, Eduard C. Dragut, Clement T. Yu, Ulf Leser
Proc. VLDB Endow.3
2009 The effect of negation on sentiment analysis and retrieval effectiveness
abstract
We investigate the problem of determining the polarity of sentiments when one or more occurrences of a negation term such as "not" appear in a sentence. The concept of the scope of a negation term is introduced. By using a parse tree and typed dependencies generated by a parser and special rules proposed by us, we provide a procedure to identify the scope of each negation term. Experimental results show that the identification of the scope of negation improves both the accuracy of sentiment analysis and the retrieval effectiveness of opinion retrieval.
Lifeng Jia, Clement T. Yu, Weiyi Meng
CIKM2
2009 Advanced metasearch engines
abstract
A metasearch engine is a system, which is connected to different search engines. In response to a user query, it invokes suitable search engines for the query, merges the information returned by these search engines and output the merged result. There are two types of metasearch engines: one type for unstructured data (mostly text) and the other for structured data. In comparison to a text search engine, a metasearch engine can have a higher coverage of the Web and can have more timely information. A metasearch engine for structured data facilitates comparison shopping and services and is convenient to use. In this talk, we discuss the problems and their potential solutions. In addition, challenges and unsolved problems are sketched.
Clement T. Yu
CIKM1
2009 Deriving Customized Integrated Web Query Interfaces
abstract
Given a set of query interfaces from providers in the same domain (e.g., car rental), the goal is to build automatically an integrated interface that makes the access to individual sources transparent to users. Our goal is to allow users to choose their preferred providers. Consequently, the integrated interface should reflect only the query interfaces of these sources. The problem scrutinized in this work is deriving customized integrated interfaces. On the hypothesis that query interfaces on the Web are easily understood by ordinary users (well-designed assumption), mainly because of the way their attributes are organized (structural property) and named (lexical property), we develop algorithms to construct customized integrated interfaces. Experiments are performed to validate our analytical studies, including a user survey.
Eduard C. Dragut, Clement T. Yu, Weiyi Meng
Web Intelligence3
2009 Stop Word and Related Problems in Web Interface Integration
abstract
The goal of recent research projects on integrating Web databases has been to enable uniform access to the large amount of data behind query interfaces. Among the tasks addressed are: source discovery, query interface extraction, schema matching, etc. There are also a number of tasks that are commonly ignored or assumed to be apriori solved either manually or by some oracle. These tasks include (1) finding the set of stop words and (2) handling occurrences of "semantic enrichment words" within labels. These two subproblems have a direct impact on determining the synonymy and hyponymy relationships between labels. In (1), a word like "from" is a stop word in general but it is a content word in domains such as Airline and Real Estate. We formulate the stop word problem , prove its complexity and provide an approximation algorithm. In (2), we study the impact of words like AND and OR on establishing semantic relationships between labels (e.g. "departure date and time" is a hypernym of "departure date"). In addition, we develop a theoretical framework to differentiate synonymy relationship from hyponymy relationship among labels involving multiple words. We scrutinize its strength and limitations both analytically and experimentally. We use real data from the Web in our experiments. We analyze over 2300 labels of 220 user interfaces in 9 distinct domains.
Eduard C. Dragut, A. Prasad Sistla, Clement T. Yu, Weiyi Meng
Proc. VLDB Endow.4
2009 A Hierarchical Approach to Model Web Query Interfaces for Web Source Integration
abstract
Much data in the Web is hidden behind Web query interfaces. In most cases the only means to "surface" the content of a Web database is by formulating complex queries on such interfaces. Applications such as Deep Web crawling and Web database integration require an automatic usage of these interfaces. Therefore, an important problem to be addressed is the automatic extraction of query interfaces into an appropriate model. We hypothesize the existence of a set of domain-independent "commonsense design rules" that guides the creation of Web query interfaces. These rules transform query interfaces into schema trees. In this paper we describe a Web query interface extraction algorithm, which combines HTML tokens and the geometric layout of these tokens within a Web page. Tokens are classified into several classes out of which the most significant ones are text tokens and field tokens. A tree structure is derived for text tokens using their geometric layout. Another tree structure is derived for the field tokens. The hierarchical representation of a query interface is obtained by iteratively merging these two trees. Thus, we convert the extraction problem into an integration problem. Our experiments show the promise of our algorithm: it outperforms the previous approaches on extracting query interfaces on about 6.5% in accuracy as evaluated over three corpora with more than 500 Deep Web interfaces from 15 different domains.
Thomas Kabisch, Eduard C. Dragut, Clement T. Yu, Ulf Leser
Proc. VLDB Endow.3
2008 Improve the effectiveness of the opinion retrieval and opinion polarity classification
abstract
Opinion retrieval is a document retrieving and ranking process. A relevant document must be relevant to the query and contain opinions toward the query. Opinion polarity classification is an extension of opinion retrieval. It classifies the retrieved document as positive, negative or mixed, according to the overall polarity of the query relevant opinions in the document. This paper (1) proposes several new techniques that help improve the effectiveness of an existing opinion retrieval system; (2) presents a novel two-stage model to solve the opinion polarity classification problem. In this model, every query relevant opinionated sentence in a document retrieved by our opinion retrieval system is classified as positive or negative respectively by a SVM classifier. Then a second classifier determines the overall opinion polarity of the document. Experimental results show that both the opinion retrieval system with the proposed opinion retrieval techniques and the polarity classification model outperformed the best reported systems respectively.
Wei Zhang 0008, Lifeng Jia, Clement T. Yu, Weiyi Meng
CIKM3
2008 A system for finding biological entities that satisfy certain conditions from texts
abstract
Finding biological entities (such as genes or proteins) that satisfy certain conditions from texts is an important and challenging task in biomedical information retrieval and text mining. It is essential for many biomedical applications, such as drug discovery which normally requires collecting existing scientific facts from documents. This paper presents an effective IR system for this task, in which 1) domain knowledge is incorporated to improve retrieval effectiveness; 2) query expansion with related concepts on multiple semantic levels is employed; 3) a gene symbol disambiguation technique is implemented. We evaluated these techniques and examined two different concept-based IR models. Experiments based upon the proposed framework yield significant improvement (22% for automatic and 16.7% for non-automatic) over the best reported results of passage retrieval in the Genomics track of TREC 2007.
Clement T. Yu, Weiyi Meng
CIKM2
2007 Recognition and classification of noun phrases in queries for effective retrieval
abstract
It has been shown that using phrases properly in the document retrieval leads to higher retrieval effectiveness. In this paper, we define four types of noun phrases and present an algorithm for recognizing these phrases in queries. The strengths of several existing tools are combined for phrase recognition. Our algorithm is tested using a set of 500 web queries from a query log, and a set of 238 TREC queries. Experimental results show that our algorithm yields high phrase recognition accuracy. We also use a baseline noun phrase recognition algorithm to recognize phrases from the TREC queries. A document retrieval experiment is conducted using the TREC queries (1) without any phrases, (2) with the phrases recognized from a baseline noun phrase recognition algorithm, and (3) with the phrases recognized from our algorithm respectively. The retrieval effectiveness of (3) is better than that of (2), which is better than that of (1). This demonstrates that utilizing phrases in queries does improve the retrieval effectiveness, and better noun phrase recognition yields higher retrieval performance.
Wei Zhang 0008, Clement T. Yu, Chaojing Sun, Fang Liu 0019, Weiyi Meng
CIKM3
2007 Opinion retrieval from blogs
abstract
Opinion retrieval is a document retrieval process, which requires documents to be retrieved and ranked according to their opinions about a query topic. A relevant document must satisfy two criteria: relevant to the query topic, and contains opinions about the query, no matter if they are positive or negative. In this paper, we describe an opinion retrieval algorithm. It has a traditional information retrieval (IR) component to find topic relevant documents from a document set, an opinion classification component to find documents having opinions from the results of the IR step, and a component to rank the documents based on their relevance to the query, and their degrees of having opinions about the query. We implemented the algorithm as a working system and tested it using TREC 2006 Blog Track data in automatic title-only runs. Our result showed 28% to 32% improvements in MAP score over the best automatic runs in this 2006 track. Our result is also 13% higher than a state-of-art opinion retrieval system, which is tested on the same data set.
Wei Zhang 0008, Clement T. Yu, Weiyi Meng
CIKM2
2007 Annotating Structured Data of the Deep Web
abstract
An increasing number of databases have become Web accessible through HTML form-based search interfaces. The data units returned from the underlying database are usually encoded into the result pages dynamically for human browsing. For the encoded data units to be machine processable, which is essential for many applications such as deep Web data collection and comparison shopping, they need to be extracted out and assigned meaningful labels. In this paper, we present a multi-annotator approach that first aligns the data units into different groups such that the data in the same group have the same semantics. Then for each group, we annotate it from different aspects and aggregate the different annotations to predict a final annotation label. An annotation wrapper for the search site is automatically constructed and can be used to annotate new result pages from the same site. Our experiments indicate that the proposed approach is highly effective.
Yiyao Lu, Hai He, Hongkun Zhao, Weiyi Meng, Clement T. Yu
ICDE5
2007 Mining templates from search result records of search engines
abstract
Metasearch engine, Comparison-shopping and Deep Web crawling applications need to extract search result records enwrapped in result pages returned from search engines in response to user queries. The search result records from a given search engine are usually formatted based on a template. Precisely identifying this template can greatly help extract and annotate the data units within each record correctly. In this paper, we propose a graph model to represent record template and develop a domain independent statistical method to automatically mine the record template for any search engine using sample search result records. Our approach can identify both template tags (HTML tags) and template texts (non-tag texts), and it also explicitly addresses the mismatches between the tag structures and the data structures of search result records. Our experimental results indicate that this approach is very effective.
Hongkun Zhao, Weiyi Meng, Clement T. Yu
KDD3
2007 Knowledge-intensive conceptual retrieval and passage extraction of biomedical literature
abstract
This paper presents a study of incorporating domain-specific knowledge (i.e., information about concepts and relationships between concepts in a certain domain) in an information retrieval (IR) system to improve its effectiveness in retrieving biomedical literature. The effects of different types of domain-specific knowledge in performance contribution are examined. Based on the TREC platform, we show that appropriate use of domain-specific knowledge in a proposed conceptual retrieval model yields about 23% improvement over the best reported result in passage retrieval in the Genomics Track of TREC 2006.
Clement T. Yu, Neil R. Smalheiser, Vetle I. Torvik
SIGIR2
2007 AllInOneNews: development and evaluation of a large-scale news metasearch engine
abstract
AllInOneNews is the largest news metasearch engine in the world, connecting to over 1,000 news sites over 150 countries. Implementing a large-scale metasearch engine like AllInOneNews needs to overcome unique challenges not faced by building small metasearch engines such as developing highly scalable search engine selection techniques. In this paper, we discuss these unique challenges and our solutions to these challenges. We also discuss some novel features of AllInOneNews such as highly automated solution and semantic query match. This paper also reports the results of a comparative evaluation of three commercial news search systems, one search engine - Google News and two metasearch engines - Mamma News and AllInOneNews. Several measures such as effectiveness, diversity and time-sensitivity are used to perform the comparison. Another contribution of this paper is that we introduce a novel scheme to compare multiple news search systems in a combined measure that takes both relevance and time-sensitivity of retrieved information into consideration.
King-Lup Liu, Weiyi Meng, Clement T. Yu, Vijay Raghavan 0001, Zonghuan Wu, Yiyao Lu, Hai He, Hongkun Zhao
SIGMOD Conference4
2007 MySearchView: a customized metasearch engine generator
abstract
In this paper, we describe MySearchView, a system for assembling search engines into metasearch engines. With this system, any user can create a metasearch engine by simply letting the system know the URLs of the search engines the user wants to be included and the metasearch engine will be built fully automatically. In this paper, the main steps of building metasearch engines will be sketched. We will also outline our plan to demonstrate all the features of MySearchView.
Yiyao Lu, Zonghuan Wu, Hongkun Zhao, Weiyi Meng, King-Lup Liu, Vijay Raghavan 0001, Clement T. Yu
SIGMOD Conference7
2007 Querying Capability Modeling and Construction of Deep Web Sources
Liangcai Shu, Weiyi Meng, Hai He, Clement T. Yu
WISE4
2006 Merging Source Query Interfaces onWeb Databases
abstract
Recently, there are many e-commerce search engines that return information from Web databases. Unlike text search engines, these e-commerce search engines have more complicated user interfaces. Our aim is to construct automatically a natural query user interface that integrates a set of interfaces over a given domain of interest. For example, each airline company has a query interface for ticket reservation and our system can construct an integrated interface for all these companies. This will permit users to access information uniformly from multiple sources. Each query interface from an e-commerce search engine is designed so as to facilitate users to provide necessary information. Specifically, (1) related pieces of information such as first name and last name are grouped together and (2) certain hierarchical relationships are maintained. In this paper, we provide an algorithm to compute an integrated interface from query interfaces of the same domain. The integrated query interface can be proved to preserve the above two types of relationships. Experiments on five domains verify our theoretical study.
Eduard C. Dragut, Wensheng Wu, A. Prasad Sistla, Clement T. Yu, Weiyi Meng
ICDE4
2006 WebIQ: Learning from the Web to Match Deep-Web Query Interfaces
abstract
Integrating Deep Web sources requires highly accurate semantic matches between the attributes of the source query interfaces. These matches are usually established by comparing the similarities of the attributes’ labels and instances. However, attributes on query interfaces often have no or very few data instances. The pervasive lack of instances seriously reduces the accuracy of current matching techniques. To address this problem, we describe WebIQ, a solution that learns from both the Surface Web and the Deep Web to automatically discover instances for interface attributes. WebIQ extends question answering techniques commonly used in the AI community for this purpose. We describe how to incorporate WebIQ into current interface matching systems. Extensive experiments over five realworld domains show the utility ofWebIQ. In particular, the results show that acquired instances help improve matching accuracy from 89.5% F-1 to 97.5%, at only a modest runtime overhead.
Wensheng Wu, AnHai Doan, Clement T. Yu
ICDE3
2006 Segmentation of Publication Records of Authors from the Web
abstract
Publication records are often found in the authors’ personal home pages. If such a record is partitioned into a list of semantic fields of authors, title, date, etc., the unstructured texts can be converted into structured data, which can be used in other applications. In this paper, we present PEPURS, a publication record segmentation system. It adopts a novel "Split and Merge" strategy. A publication record is split into segments; multiple statistical classifiers compute their likelihoods of belonging to different fields; finally adjacent segments are merged if they belong to the same field. PEPURS introduces the punctuation marks and their neighboring texts as a new feature to distinguish different roles of the marks. PEPURS yields high accuracy scores in experiments.
Wei Zhang 0008, Clement T. Yu, Neil R. Smalheiser, Vetle I. Torvik
ICDE2
2006 Effective keyword search in relational databases
abstract
With the amount of available text data in relational databases growing rapidly, the need for ordinary users to search such information is dramatically increasing. Even though the major RDBMSs have provided full-text search capabilities, they still require users to have knowledge of the database schemas and use a structured query language to search information. This search model is complicated for most ordinary users. Inspired by the big success of information retrieval (IR) style keyword search on the web, keyword search in relational databases has recently emerged as a new research topic. The differences between text databases and relational databases result in three new challenges: (1) Answers needed by users are not limited to individual tuples, but results assembled from joining tuples from multiple tables are used to form answers in the form of tuple trees. (2) A single score for each answer (i.e. a tuple tree) is needed to estimate its relevance to a given query. These scores are used to rank the most relevant answers as high as possible. (3) Relational databases have much richer structures than text databases. Existing IR strategies to rank relational outputs are not adequate. In this paper, we propose a novel IR ranking strategy for effective keyword search. We are the first that conducts comprehensive experiments on search effectiveness using a real world database and a set of keyword queries collected by a major search company. Experimental results show that our strategy is significantly better than existing strategies. Our approach can be used both at the application level and be incorporated into a RDBMS to support keyword-based search in relational databases.
Fang Liu 0019, Clement T. Yu, Weiyi Meng, Abdur Chowdhury
SIGMOD Conference2
2006 Meaningful Labeling of Integrated Query Interfaces
Eduard C. Dragut, Clement T. Yu, Weiyi Meng
VLDB2
2006 Automatic Extraction of Dynamic Record Sections From Search Engine Result Pages
Hongkun Zhao, Weiyi Meng, Clement T. Yu
VLDB3
2006 Clustering e-commerce search engines based on their search interface pages using WISE-Cluster
Yiyao Lu, Hai He, Weiyi Meng, Clement T. Yu
Data Knowl. Eng.5
2005 Mining officially unrecognized side effects of drugs by combining web search and machine learning
abstract
We consider the problem of finding officially unrecognized side effects of drugs. By submitting queries to the Web involving a given drug name, it is possible to retrieve pages concerning the drug. However, many retrieved pages are irrelevant and some relevant pages are not retrieved. More relevant pages can be obtained by adding the active ingredient of the drug to the query. In order to eliminate irrelevant pages, we propose a machine learning process to filter out the undesirable pages. The process is shown experimentally to be very effective. Since obtaining training data for the machine learning process can be time consuming and expensive, we provide an automatic method to generate the training data. The method is also shown to be very accurate. The side effects of three drugs which are not recognized by FDA are validated by an expert. We believe that the same approach can be applied to many real life problems and will yield high precision. Thus, this could lead a new way to perform retrieval with high accuracy.
Carlo Curino, Bruce Lambert, Patricia M. West, Clement T. Yu
CIKM5
2005 Database selection in intranet mediators for natural language queries
abstract
No abstract available.
Fang Liu 0019, Clement T. Yu, Weiyi Meng, Ophir Frieder, David A. Grossman
CIKM3
2005 Word sense disambiguation in queries
abstract
This paper presents a new approach to determine the senses of words in queries by using WordNet. In our approach, noun phrases in a query are determined first. For each word in the query, information associated with it, including its synonyms, hyponyms, hypernyms, definitions of its synonyms and hyponyms, and its domains, can be used for word sense disambiguation. By comparing these pieces of information associated with the words which form a phrase, it may be possible to assign senses to these words. If the above disambiguation fails, then other query words, if exist, are used, by going through exactly the same process. If the sense of a query word cannot be determined in this manner, then a guess of the sense of the word is made, if the guess has at least 50% chance of being correct. If no sense of the word has 50% or higher chance of being used, then we apply a Web search to assist in the word sense disambiguation process. Experimental results show that our approach has 100% applicability and 90% accuracy on the most recent robust track of TREC collection of 250 queries. We combine this disambiguation algorithm to our retrieval system to examine the effect of word sense disambiguation in text retrieval. Experimental results show that the disambiguation algorithm together with other components of our retrieval system yield a result which is 13.7% above that produced by the same system but without the disambiguation, and 9.2% above that produced by using Lesk's algorithm. Our retrieval effectiveness is 7% better than the best reported result in the literature.
Clement T. Yu, Weiyi Meng
CIKM2
2005 Merging Interface Schemas on the Deep Web via Clustering Aggregation
abstract
We consider the problem of integrating a large number of interface schemas over the deep Web, The scale of the problem and the diversity of the sources present serious challenges to the conventional manual or rule-based approaches to schema integration. To address these challenges, we propose a novel formulation of schema integration as an optimization problem, with the objective of maximally satisfying the constraints given by individual schemas. Since the optimization problem can be shown to be NP-complete, we develop a novel approximation algorithm LMax, which builds the unified schema via recursive applications of clustering aggregation. We further extend LMax to handle the irregularities frequently occurring among the interface schemas. Extensive evaluation on real-world data sets shows the effectiveness of our approach.
Wensheng Wu, AnHai Doan, Clement T. Yu
ICDM3
2005 WISE-Integrator: A System for Extracting and Integrating Complex Web Search Interfaces of the Deep Web
Hai He, Weiyi Meng, Clement T. Yu, Zonghuan Wu
VLDB3
2005 Constructing Interface Schemas for Search Interfaces of Web Databases
Hai He, Weiyi Meng, Clement T. Yu, Zonghuan Wu
WISE3
2005 Evaluation of Result Merging Strategies for Metasearch Engines
Yiyao Lu, Weiyi Meng, Liangcai Shu, Clement T. Yu, King-Lup Liu
WISE4
2005 Fully automatic wrapper generation for search engines
abstract
When a query is submitted to a search engine, the search engine returns a dynamically generated result page containing the result records, each of which usually consists of a link to and/or snippet of a retrieved Web page. In addition, such a result page often also contains information irrelevant to the query, such as information related to the hosting site of the search engine and advertisements. In this paper, we present a technique for automatically producing wrappers that can be used to extract search result records from dynamically generated result pages returned by search engines. Automatic search result record extraction is very important for many applications that need to interact with search engines such as automatic construction and maintenance of metasearch engines and deep Web crawling. The novel aspect of the proposed technique is that it utilizes both the visual content features on the result page as displayed on a browser and the HTML tag structures of the HTML source file of the result page. Experimental results indicate that this technique can achieve very high extraction accuracy.
Hongkun Zhao, Weiyi Meng, Zonghuan Wu, Vijay Raghavan 0001, Clement T. Yu
WWW5
2004 An effective approach to document retrieval via utilizing WordNet and recognizing phrases
abstract
Noun phrases in queries are identified and classified into four types: proper names, dictionary phrases, simple phrases and complex phrases. A document has a phrase if all content words in the phrase are within a window of a certain size. The window sizes for different types of phrases are different and are determined using a decision tree. Phrases are more important than individual terms. Consequently, documents in response to a query are ranked with matching phrases given a higher priority. We utilize WordNet to disambiguate word senses of query terms. Whenever the sense of a query term is determined, its synonyms, hyponyms, words from its definition and its compound words are considered for possible additions to the query. Experimental results show that our approach yields between 23% and 31% improvements over the best-known results on the TREC 9, 10 and 12 collections for short (title only) queries, without using Web data.
Fang Liu 0019, Clement T. Yu, Weiyi Meng
SIGIR3
2004 An Interactive Clustering-based Approach to Integrating Source Query interfaces on the Deep Web
abstract
An increasing number of data sources now become available on the Web, but often their contents are only accessible through query interfaces. For a domain of interest, there often exist many such sources with varied coverage or querying capabilities. As an important step to the integration of these sources, we consider the integration of their query interfaces. More specifically, we focus on the crucial step of the integration: accurately matching the interfaces. While the integration of query interfaces has received more attentions recently, current approaches are not sufficiently general: (a) they all model interfaces with flat schemas; (b) most of them only consider 1:1 mappings of fields over the interfaces; (c) they all perform the integration in a blackbox-like fashion and the whole process has to be restarted from scratch if anything goes wrong; and (d) they often require laborious parameter tuning. In this paper, we propose an interactive, clustering-based approach to matching query interfaces. The hierarchical nature of interfaces is captured with ordered trees. Varied types of complex mappings of fields are examined and several approaches are proposed to effectively identify these mappings. We put the human integrator back in the loop and propose several novel approaches to the interactive learning of parameters and the resolution of uncertain mappings. Extensive experiments are conducted and results show that our approach is highly effective.
Wensheng Wu, Clement T. Yu, AnHai Doan, Weiyi Meng
SIGMOD Conference2
2004 Robust content-based image indexing using contextual clues and automatic pseudofeedback
Y. Alp Aslandogan, Clement T. Yu, Ravishankar Mysore
Multim. Syst.2
2004 Personalized Web Search For Improving Retrieval Effectiveness
abstract
Current Web search engines are built to serve all users, independent of the special needs of any individual user. Personalization of Web search is to carry out retrieval for each user incorporating his/her interests. We propose a novel technique to learn user profiles from users' search histories. The user profiles are then used to improve retrieval effectiveness in Web search. A user profile and a general profile are learned from the user's search history and a category hierarchy, respectively. These two profiles are combined to map a user query into a set of categories which represent the user's search intention and serve as a context to disambiguate the words in the user's query. Web search is conducted based on both the user query and the set of categories. Several profile learning and category mapping algorithms and a fusion algorithm are provided and evaluated. Experimental results indicate that our technique to personalize Web search is both effective and efficient.
Fang Liu 0019, Clement T. Yu, Weiyi Meng
IEEE Trans. Knowl. Data Eng.2
2004 Automatic integration of Web search interfaces with WISE-Integrator
Hai He, Weiyi Meng, Clement T. Yu, Zonghuan Wu
VLDB J.3
2003 SE-LEGO: creating metasearch engines on demand
abstract
No abstract available.
Zonghuan Wu, Vijay Raghavan 0001, Chun Du, Komanduru Sai C, Weiyi Meng, Hai He, Clement T. Yu
SIGIR7
2003 WISE-Integrator: An Automatic Integrator of Web Search Interfaces for E-Commerce
Hai He, Weiyi Meng, Clement T. Yu, Zonghuan Wu
VLDB3
2003 Distributed Top-N Query Processing with Possibly Uncooperative Local Systems
Clement T. Yu, George Philip, Weiyi Meng
VLDB1
2003 Creating Customized Metasearch Engines on Demand Using SE-LEGO
Zonghuan Wu, Vijay Raghavan 0001, Weiyi Meng, Hai He, Clement T. Yu, Chun Du
WAIM5
2003 Towards Automatic Incorporation of Search Engines into a Large-Scale Metasearch Engine
abstract
A metasearch engine supports unified access to multiple component search engines. To build a very large-scale metasearch engine that can access up to hundreds of thousands of component search engines, one major challenge is to incorporate large numbers of autonomous search engines in a highly effective manner. To solve this problem, we propose automatic search engine discovery, automatic search engine connection, and automatic search engine result extraction techniques. Experiments indicate that these techniques are highly effective and efficient.
Zonghuan Wu, Vijay Raghavan 0001, Hua Qian, Rama Vuyyuru, Weiyi Meng, Hai He, Clement T. Yu
Web Intelligence7
2003 Haar Wavelets for Efficient Similarity Search of Time-Series: With and Without Time Warping
abstract
We address the handling of time series search based on two important distance definitions: Euclidean distance and time warping distance. The conventional method reduces the dimensionality by means of a discrete Fourier transform. We apply the Haar wavelet transform technique and propose the use of a proper normalization so that the method can guarantee no false dismissal for Euclidean distance. We found that this method has competitive performance from our experiments. Euclidean distance measurement cannot handle the time shifts of patterns. It fails to match the same rise and fall patterns of sequences with different scales. A distance measure that handles this problem is the time warping distance. However, the complexity of computing the time warping distance function is high. Also, as time warping distance is not a metric, most indexing techniques would not guarantee any false dismissal. We propose efficient strategies to mitigate the problems of time warping. We suggest a Haar wavelet-based approximation function for time warping distance, called Low Resolution Time Warping, which results in less computation by trading off a small amount of accuracy. We apply our approximation function to similarity search in time series databases, and show by experiment that it is highly effective in suppressing the number of false alarms in similarity search.
Kin-pong Chan, Ada Wai-Chee Fu, Clement T. Yu
IEEE Trans. Knowl. Data Eng.3
2002 Personalized web search by mapping user queries to categories
abstract
Current web search engines are built to serve all users, independent of the needs of any individual user. Personalization of web search is to carry out retrieval for each user incorporating his/her interests. We propose a novel technique to map a user query to a set of categories, which represent the user's search intention. This set of categories can serve as a context to disambiguate the words in the user's query. A user profile and a general profile are learned from the user's search history and a category hierarchy respectively. These two profiles are combined to map a user query into a set of categories. Several learning and combining algorithms are evaluated and found to be effective. Among the algorithms to learn a user profile, we choose the Rocchio-based method for its simplicity, efficiency and its ability to be adaptive. Experimental results indicate that our technique to personalize web search is both effective and efficient.
Fang Liu 0019, Clement T. Yu, Weiyi Meng
CIKM2
2002 Discovering the representative of a search engine
abstract
Given a large number of search engines on the Internet, it is difficult for a person to determine which search engines could serve his/her information needs. A common solution is to construct a metasearch engine on top of the search engines. Upon receiving a user query, the metasearch engine sends it to those underlying search engines which are likely to return the desired documents for the query. The selection algorithm used by a metasearch engine to determine whether a search engine should be sent the query typically makes the decision based on the search-engine representative, which contains characteristic information about the database of a search engine. However, an underlying search engine may not be willing to provide the needed information to the metasearch engine. This paper shows that the needed information can be estimated from an uncooperative search engine with good accuracy. Two pieces of information which permit accurate search engine selection are the number of documents indexed by the search engine and the maximum weight of each term. In this paper, we present techniques for the estimation of these two pieces of information.
King-Lup Liu, Clement T. Yu, Weiyi Meng
CIKM2
2002 Automatic feedback for content based image retrieval on the Web
abstract
We address the problem of identifying images of persons in large collections, such as the Web, without an existing face image database. We describe a method and a system that automatically constructs an initial face image database for a person using textual evidence obtained from the Web, and then uses this database for identifying images of that person. The initial retrieval results are obtained via text/HTML analysis and face detection. An internal clustering process groups visually similar faces among these initial results and builds a facial database. This database is then used by a face recognizer. The outputs of the textual and visual evidence modules are combined using Dempster-Shafer (1976) evidence combination formula. We present the results of an experimental evaluation where the system was able to improve upon the detection-only method when text/HTML analysis performed poorly.
Y. Alp Aslandogan, Clement T. Yu
ICME (1)2
2002 Concept Hierarchy-Based Text Database Categorization
Weiyi Meng, Wenxian Wang, Hongyu Sun 0003, Clement T. Yu
Knowl. Inf. Syst.4
2002 A Statistical Method for Estimating the Usefulness of Text Databases
abstract
Searching desired data on the Internet is one of the most common ways the Internet is used. No single search engine is capable of searching all data on the Internet. The approach that provides an interface for invoking multiple search engines for each user query has the potential to satisfy more users. When the number of search engines under the interface is large, invoking all search engines for each query is often not cost effective because it creates unnecessary network traffic by sending the query to a large number of useless search engines and searching these useless search engines wastes local resources. The problem can be overcome if the usefulness of every search engine with respect to each query can be predicted. We present a statistical method to estimate the usefulness of a search engine for any given query. For a given query, the usefulness of a search engine in this paper is defined to be a combination of the number of documents in the search engine that are sufficiently similar to the query and the average similarity of these documents. Experimental results indicate that our estimation method is much more accurate than existing methods.
King-Lup Liu, Clement T. Yu, Weiyi Meng, Wensheng Wu, Naphtali Rishe
IEEE Trans. Knowl. Data Eng.2
2002 A Methodology to Retrieve Text Documents from Multiple Databases
abstract
This paper presents a methodology for finding the n most similar documents across multiple text databases for any given query and for any positive integer n. This methodology consists of two steps. First, the contents of databases are indicated approximately by database representatives. Databases are ranked using their representatives with respect to the given query. We provide a necessary and sufficient condition to rank the databases optimally. In order to satisfy this condition, we provide three estimation methods. One estimation method is intended for short queries; the other two are for all queries. Second, we provide an algorithm, OptDocRetrv, to retrieve documents from the databases according to their rank and in a particular way. We show that if the databases containing the n most similar documents for a given query are ranked ahead of other databases, our methodology will guarantee the retrieval of the n most similar documents for the query. When the number of databases is large, we propose to organize database representatives into a hierarchy and employ a best-search algorithm to search the hierarchy. It is shown that the effectiveness of the best-search algorithm is the same as that of evaluating the user query against all database representatives.
Clement T. Yu, King-Lup Liu, Weiyi Meng, Zonghuan Wu, Naphtali Rishe
IEEE Trans. Knowl. Data Eng.1
2001 Discovering the Representative of a Search Engine
abstract
Given a large number of search engines on the Internet, it is difficult for a person to determine which search engines could serve his/her information needs. A common solution is to construct a metasearch engine on top of the search engines. Upon receiving a user query, the metasearch engine sends it to those underlying search engines which are likely to return the desired documents for the query. The selection algorithm used by a metasearch engine to determine whether a search engine should be sent the query typically makes the decision based on the search-engine representative, which contains characteristic information about the database of a search engine. However, an underlying search engine may not be willing to provide the needed information to the metasearch engine. This paper shows that the needed information can be estimated from an uncooperative search engine with good accuracy. Two pieces of information which permit accurate search engine selection are the number of documents indexed by the search engine and the maximum weight of each term. In this paper, we present techniques for the estimation of these two pieces of information.
King-Lup Liu, Clement T. Yu, Weiyi Meng, Adrian Santoso
CIKM2
2001 Efficient and Effective Metasearch for Text Databases Incorporating Linkages among Documents
abstract
Linkages among documents have a significant impact on the importance of documents, as it can be argued that important documents are pointed to by many documents or by other important documents. Metasearch engines can be used to facilitate ordinary users for retrieving information from multiple local sources (text databases). There is a search engine associated with each database. In a large-scale metasearch engine, the contents of each local database is represented by a representative. Each user query is evaluated against he set of representatives of all databases in order to determine the appropriate databases (search engines) to search (invoke) In previous word, the linkage information between documents has not been utilized in determining the appropriate databases to search. In this paper, such information is employed to determine the degree of relevance of a document with respect to a given query. Specifically, the importance (rank) of each document as determined by the linkages is integrated in each database representative to facilitate the selection of databases for each given query. We establish a necessary and sufficient condition to rank databases optimally, while incorporating the linkage information. A method is provided to estimate the desired quantities stated in the necessary and sufficient condition. The estimation method runs in time linearly proportional to the number of query terms. Experimental results are provided to demonstrate the high retrieval effectiveness of the method.
Clement T. Yu, Weiyi Meng, Wensheng Wu, King-Lup Liu
SIGMOD Conference1
2001 Towards a highly-scalable and effective metasearch engine
abstract
A metasearch engine is a system that supports unified access to multiple local search engines. Database selection is one of the main challenges in building a large-scale metasearch engine. The problem is to efficiently and accurately determine a small number of potentially useful local search engines to invoke for each user query. In order to enable accurate selection, metadata that reflect the contents of each search engine need to be collected and used. In this paper, we propose a highly scalable and accurate database selection method. This method has several novel features. First, the metadata for representing the contents of all search engines are organized into a single integrated representative, instead of one for each search engine by all existing approaches. Such a representative yields both computation efficiency and storage efficiency. Second, our selection method is based on a theory for ranking search engines optimally. Experimental results indicate that this new ...
Zonghuan Wu, Weiyi Meng, Clement T. Yu, Zhuogang Li
WWW3
2001 Efficient Processing of Nested Fuzzy SQL Queries in a Fuzzy Database
abstract
In a fuzzy relational database where a relation is a fuzzy set of tuples and ill-known data are represented by possibility distributions, nested fuzzy queries can be expressed in the Fuzzy SQL language. Although it provides a very convenient way for users to express complex queries, a nested fuzzy query may be very inefficient to process with the naive evaluation method based on its semantics. In conventional databases, nested queries are unnested to improve the efficiency of their evaluation. In this paper, we extend the unnesting techniques to process several types of nested fuzzy queries. An extended merge-join is used to evaluate the unnested fuzzy queries. As shown by both theoretical analysis and experimental results, the unnesting techniques with the extended merge-join significantly improve the performance of evaluating nested fuzzy queries.
Qi Yang 0011, Weining Zhang, Chengwen Liu, Clement T. Yu, Hiroshi Nakajima, Naphtali Rishe
IEEE Trans. Knowl. Data Eng.5
2001 A highly scalable and effective method for metasearch
abstract
A metasearch engine is a system that supports unified access to multiple local search engines. Database selection is one of the main challenges in building a large-scale metasearch engine. The problem is to efficiently and accurately determine a small number of potentially useful local search engines to invoke for each user query. In order to enable accurate selection, metadata that reflect the contents of each search engine need to be collected and used. This article proposes a highly scalable and accurate database selection method. This method has several novel features. First, the metadata for representing the contents of all search engines are organized into a single integrated representative. Such a representative yields both computational efficiency and storage efficiency. Second, the new selection method is based on a theory for ranking search engines optimally. Experimental results indicate that this new method is very effective. An operational prototype system has been built based on the proposed approach.
Weiyi Meng, Zonghuan Wu, Clement T. Yu, Zhuogang Li
ACM Trans. Inf. Syst.3
2000 Discovery of Similarity Computations of Search Engines
abstract
Two typical situations in which it is of practical interest to determine the similarities of text documents to a query due to a search engine are: (1) a global search engine, constructed on top of a group of local search engines, wishes to retrieve the set of local documents globally most similar to a given query; and (2) an organization wants to compare the retrieval performance of search engines. The dot-product function is a widely used similarity function. For a search engine using such a function, we can determine its similarity computations if how the search engine sets the weights of terms is known, which is usually not the case. In this paper, techniques are presented to discover certain mathematical expressions of these formulas and the values of embedded constants when the dot-product similarity function is used. Preliminary results from experiments on the WebCrawler search engine are given to illustrate our techniques. 1 1 Introduction Documents of interest are often found...
King-Lup Liu, Weiyi Meng, Clement T. Yu, Naphtali Rishe
CIKM3
2000 Evaluating strategies and systems for content based indexing of person images on the Web
abstract
Content based indexing of multimedia has always been a challenging task. The enormity and the diversity of the multimedia content on the web adds another dimension to this challenge. In this paper, we examine ways of combining visual and textual information for content based indexing of multimedia on the web. In particular, we examine different methods of combining evidences due to face detection, Text/HTML analysis and face recognition for identifying person images. We provide experimental evaluation of the following strategies: i) Face detection on the image followed by Text/HTML analysis of the containing page; ii) face detection followed by face recognition; iii) face detection followed by a linear combination of evidences due to text/HTML analysis and face recognition; and iv) face detection followed by a Dempster-Shafer combination of evidences due to text/HTML analysis and face recognition. These strategies were implemented in an automatic web search agent named Diogenes1 and compared against some well known web image search engines. The latter includes commercial systems such as Alta Vista, Lycos and Ditto, and a research prototype, WebSEEk. We report the results of our experimental retrievals where Diogenes outperformed these search engines for celebrity image queries in terms of average precision.
Y. Alp Aslandogan, Clement T. Yu
ACM Multimedia2
2000 Diogenes: a web search agent for person images
Y. Alp Aslandogan, Clement T. Yu
ACM Multimedia2
2000 Multiple evidence combination in image retrieval: diogenes searches for people on the Web
abstract
In this work, we examine evidence combination mechanisms for classifying multimedia information. In particular, we examine linear and Dempster-Shafer methods of evidence combination in the context of identifying personal images on the World Wide Web. An automatic web search engine named Diogenes1 searches the web for personal images and combines different pieces of evidence for identification. The sources of evidence consist of input from face detection/recognition and text/HTML analysis modules. A degree of uncertainty is involved with both of these sources. Diogenes automatically determines the uncertainty locally for each retrieval and uses this information to set a relative significance for each evidence. To our knowledge, Diogenes is the first image search engine using Dempster-Shafer evidence combination based on automatic object recognition and dynamic local uncertainty assessment. In our experiments Diogenes comfortably outperformed some well known commercial and research prototype image search engines for celebrity image queries.
Y. Alp Aslandogan, Clement T. Yu
SIGIR2
2000 Concept Hierarchical based Text Database Categorization in a Metasearch Engine Environment
abstract
Document categorization, as a technique to improve the retrieval of useful documents, has been extensively investigated. One important issue in a large-scale meta-search engine is to select text databases that are likely to contain useful documents for a given query. We believe that database categorization can be a potentially effective technique for good database selection, especially in the Internet environment, where short queries are usually submitted. In this paper, we propose and evaluate several database categorization algorithms. This study indicates that, while some document categorization algorithms could be adopted for database categorization, algorithms that take into consideration the special characteristics of databases may be more effective. Preliminary experimental results are provided to compare the proposed database categorization algorithms.
Wenxian Wang, Weiyi Meng, Clement T. Yu
WISE3
2000 Reasoning about Qualitative Spatial Relationships
A. Prasad Sistla, Clement T. Yu
J. Autom. Reason.2
1999 Efficient and Effective Metasearch for a Large Number of Text Databases
abstract
Metasearch engines can be used to facilitate ordinary users for retrieving information from multiple local sources (text databases). In a metasearch engine, the contents of each local database is represented by a representative. Each user query is evaluated against the set of representatives of all databases in order to determine the appropriate databases to search. When the number of databases is very large, say in the order of tens of thousands or more, then a traditional metasearch engine may become inefficient as each query needs to be evaluated against too many database representatives. Furthermore, the storage requirement on the site containing the metasearch engine can be very large. In this paper, we propose to use a hierarchy of database representatives to improve the efficiency. We provide an algorithm to search the hierarchy. We show that the retrieval effectiveness of our algorithm is the same as that of evaluating the user query against all database representatives. We also show that our algorithm is efficient. In addition, we propose an alternative way of allocating representatives to sites so that the storage burden on the site containing the metasearch engine is much reduced.
Clement T. Yu, Weiyi Meng, King-Lup Liu, Wensheng Wu, Naphtali Rishe
CIKM1
1999 Detection of Heterogeneities in a Multiple Text Database Environment
abstract
As the number of text retrieval systems (search engines) grows rapidly on the World Wide Web, there is an increasing need to build search brokers (metasearch engines) on top of them. Often, the task of building an effective and efficient metasearch engine is hindered by the heterogeneities among the underlying local search engines. We first analyze the impact of various heterogeneities on building a metasearch engine. We then present some techniques that can be used to detect the most prominent heterogeneities among multiple search engines. Applications of utilizing the detected heterogeneities in building better metasearch engines are provided.
Weiyi Meng, Clement T. Yu, King-Lup Liu
CoopIS2
1999 Estimating the Usefulness of Search Engines
abstract
In this paper, we present a statistical method to estimate the usefulness of a search engine for any given query. The estimates can be used by a metasearch engine to choose local search engines to invoke. For a given query, the usefulness of a search engine in this paper is defined to be a combination of the number of documents in the search engine that are sufficiently similar to the query and the average similarity of these documents. Experimental results indicate that the proposed estimation method is quite accurate.
Weiyi Meng, King-Lup Liu, Clement T. Yu, Wensheng Wu, Naphtali Rishe
ICDE3
1999 Segmentation of skin cancer images
Marcel P. Jackowski, A. Ardeshir Goshtasby, D. Roseman, S. Bines, Clement T. Yu, Akshaya Dhawan, Arthur C. Huntley
Image Vis. Comput.6
1999 Techniques and Systems for Image and Video Retrieval
abstract
Storage and retrieval of multimedia has become a requirement for many contemporary information systems. These systems need to provide browsing, querying, navigation, and, sometimes, composition capabilities involving various forms of media. In this survey, we review techniques and systems for image and video retrieval. We first look at visual features for image retrieval such as color, texture, shape, and spatial relationships. The indexing techniques are discussed for these features. Nonvisual features include captions, annotations, relational attributes, and structural descriptions. Temporal aspects of video retrieval and video segmentation are discussed next. We review several systems for image and video retrieval including research, commercial, and World Wide Web-based systems. We conclude with an overview of current challenges and future trends for image and video retrieval.
Y. Alp Aslandogan, Clement T. Yu
IEEE Trans. Knowl. Data Eng.2
1998 Query Processing in a Video Retrieval System
abstract
A.P. Sistla et al. (1997) designed a similarity-based video retrieval system. Queries were specified in a language called the Hierarchical Temporal Language (HTL). In this paper, we present several extensions of HTL. These extensions include queries that can have the negation operator and any other logical and temporal operators such as disjunction. Efficient algorithms for processing queries in the extended language are also presented.
King-Lup Liu, A. Prasad Sistla, Clement T. Yu, Naphtali Rishe
ICDE3
1998 Determining Text Databases to Search in the Internet
Weiyi Meng, King-Lup Liu, Clement T. Yu, Yuhsi Chang, Naphtali Rishe
VLDB3
1998 Performance Analysis of Three Text-Join Algorithms
abstract
When a multidatabase system contains textual database systems (i.e., information retrieval systems), queries against the global schema of the multidatabase system may contain a new type of joins-joins between attributes of textual type. Three algorithms for processing such a type of joins are presented and their I/O costs are analyzed in this paper. Since such a type of joins often involves document collections of very large size, it is very important to find efficient algorithms to process them. The three algorithms differ on whether the documents themselves or the inverted files on the documents are used to process the join. Our analysis and the simulation results indicate that the relative performance of these algorithms depends on the input document collections, system characteristics, and the input query. For each algorithm, the type of input document collections with which the algorithm is likely to perform well is identified. An integrated algorithm that automatically selects the best algorithm to use is also proposed.
Weiyi Meng, Clement T. Yu, Wei Wang 0010, Naphtali Rishe
IEEE Trans. Knowl. Data Eng.2
1997 Similarity Based Retrieval of Videos
abstract
The authors propose a language, called Hierarchical Temporal Logic (HTL), for specifying queries on video databases. The language is based on the hierarchical as well as temporal nature of video data. They give similarity based semantics for the logic, and give efficient methods for computing similarity values for subclasses of HTL formulas. Experimental results, indicating the effectiveness of the methods, are presented.
A. Prasad Sistla, Clement T. Yu, R. Venkatasubrahmanian
ICDE2
1997 Using Semantic Contents and WordNet in Image Retrieval
abstract
Image retrievaf based on semantic contents involves extraction, modelling and indexing of content information.While extraction of abstract contents is a hard problem, it is only part of the bigger picture.In this paper we use knowledge about the semantic contents of images to improve retrieval effectiveness.In particular we use Word Net, an electronic Iexicaf system for query and database expansion.Our content model facilitates novel uses of WordNet.We also propose a new normalization formula, an object significance scheme and evaluate their effectiveness with real user experiments.We describe the experiment setup and provide quantitative evaluation of each technique.
Y. Alp Aslandogan, Chuck Thier, Clement T. Yu, Jon Zou, Naphtali Rishe
SIGIR3
1997 A Parallel Scheme Using the Divide-and-Conquer Method
Qi Yang 0011, Son Dao, Clement T. Yu, Naphtali Rishe
Distributed Parallel Databases3
1997 Correcting the Geometry and Color of Digital Images
abstract
A unified method for correcting the geometry and color of digital images is presented. The method uses a color chart with a prearranged array of known color patches. Using the true geometry and color of the chart and geometry and color of the image of the chart, transformation functions that describe the geometry and color characteristics of the camera are determined. The transformation functions are then used to correct the geometry and color of obtained images. Experimental results of the proposed method on images acquired under different scene illuminations and camera settings are presented and evaluated.
Marcel P. Jackowski, A. Ardeshir Goshtasby, S. Bines, D. Roseman, Clement T. Yu
IEEE Trans. Pattern Anal. Mach. Intell.5
1996 Performance Evaluation of G-tree and Its Application in Fuzzy Databases
abstract
Article Free Access Share on Performance evaluation of G-tree and its application in fuzzy databases Authors: Chengwen Liu DePaul University, Chicago, Illinois DePaul University, Chicago, IllinoisView Profile , Aris Ouksel University of Illinois at Chicago University of Illinois at ChicagoView Profile , Prasad Sistla University of Illinois at Chicago University of Illinois at ChicagoView Profile , Jing Wu University of Illinois at Chicago University of Illinois at ChicagoView Profile , Clement Yu University of Illinois at Chicago University of Illinois at ChicagoView Profile , Naphtali Rishe Florida International University Florida International UniversityView Profile Authors Info & Claims CIKM '96: Proceedings of the fifth international conference on Information and knowledge managementNovember 1996 Pages 235–242https://doi.org/10.1145/238355.238503Published:12 November 1996Publication History 11citation347DownloadsMetricsTotal Citations11Total Downloads347Last 12 Months6Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Chengwen Liu, Aris M. Ouksel, A. Prasad Sistla, Clement T. Yu, Naphtali Rishe
CIKM5
1996 Performance Analysis of Several Algorithms for Processing Joins between Textual Attributes
abstract
Three algorithms for processing joins on attributes of a textual type are presented and analyzed in this paper. Since such joins often involve document collections of very large size, it is very important to find efficient algorithms to process them. The three algorithms differ according to whether the documents themselves or the inverted files on the documents are used to process the join. Our analysis and simulation results indicate that the relative performance of these algorithms depends on the input document collections, the system characteristics and the input query. For each algorithm, the type of input document collection with which the algorithm is likely to perform well is identified.
Weiyi Meng, Clement T. Yu, Wei Wang 0010, Naphtali Rishe
ICDE2
1996 Efficient Processing of One and Two Dimensional Proximity Queries in Associative Memory
abstract
Proximity queries that involve multiple object types are very common. In this paper, we present a parallel algorithm for answering proximity queries of one kind over object instances that lie in a one-dimensional metric space. The algorithm exploits a specialized hardware, the Dynamic Associative Access Memory chip. In most proximity queries of this kind, the number of object types is less than or equal to three and the distance d, within which object instances are required to locate to satisfy a given proximity condition, is small (dd=80e = 1). The execution time for such queries is linearly proportional to the number of object types and is independent of the size of the database. This allows numerous concurrent users to be serviced. The algorithm is extended to process 2-dimensional proximity queries efficiently. This research is supported in part by NASA under NAGW-4080 and ARO under BMDO grant DAAH04-0024. 1 1
King-Lup Liu, G. Jack Lipovski, Clement T. Yu, Naphtali Rishe
SIGIR3
1996 Computation of Best Bounds of Probabilities from Uncertain Data
abstract
An uncertainty reasoning method is presented in this article. The method can be used to compute from a given set of conditional probabilities the best lower bounds and the best upper bounds of those conditional probabilities that are not explicitly provided. The computation of the best upper(lower) bound of such a conditional probability relies on solution of a linear programming problem. Some reduction techniques are proposed in this article to improve the efficiency of our uncertainty reasoning method. As illustrated in Section 4.3, for many uncertainty reasoning problems in medical diagnosis, by using our reduction techniques, the best range of a conditional probability, which is specified by a lower bound and an upper bound, can be computed in polynomial time in terms of the number of basic events involved in the reasoning.
Chengjie Luo, Clement T. Yu, Jorge Lobo 0001, Gaoming Wang, Tracy Pham
Comput. Intell.2
1995 Design, Implementation and Evaluation of SCORE (a System for COntent based REtrieval of Pictures)
abstract
We make use of a refined E-R model to represent the contents of pictures. We propose remedies to handle mismatches which may arise due to differences in perception of picture contents. An iconic user interface for visual query construction is presented. A naive user can specify his/her intention without learning a query language. A function which computes the similarity between a picture and a user's description is provided. Pictures which are sufficiently close to the user description, as measured by the similarity function, are retrieved. We present the results of a user-friendliness experiment to evaluate the user interface as well as retrieval effectiveness. Encouraging retrieval results and valuable lessons are obtained.>
Y. Alp Aslandogan, Chuck Thier, Clement T. Yu, Chengwen Liu, Krishnakumar R. Nair
ICDE3
1995 Efficient Processing of Nested Fuzzy SQL Queries
abstract
Fuzzy databases have been introduced to deal with uncertain or incomplete information in many applications. The efficiency of processing fuzzy queries in fuzzy databases is a major concern. We provide techniques to unnest nested fuzzy queries of two blocks in fuzzy databases. We show both theoretically and experimentally that unnesting improves the performance of nested queries significantly. The results obtained in the paper form the basis for unnesting fuzzy queries of arbitrary blocks in fuzzy databases.>
Qi Yang 0011, Chengwen Liu, Clement T. Yu, Son Dao, Hiroshi Nakajima
ICDE4
1995 Translation of Object-Oriented Queries to Relational Queries
abstract
Proposes a formal approach for translating OODB queries to equivalent relational queries. The translation is accomplished through the use of relational predicate graphs and OODB predicate graphs. One advantage of using such a graph-based approach is that we can achieve bidirectional translation between relational queries and OODB queries.>
Clement T. Yu, Weiyi Meng, Won Kim 0001, Gaoming Wang, Tracy Pham, Son Dao
ICDE1
1995 Context-Dependent Interpretations of Linguistic Terms in Fuzzy Relational Databases
abstract
Approaches are proposed to allow fuzzy terms to be interpreted according to the context within which they are used. Such an interpretation is natural and useful. A query-dependent interpretation is proposed to allow a fuzzy term to be interpreted relative to a partial answer of a query. A scaling process is used to transform a pre-defined meaning of a fuzzy term into on appropriate meaning in the given context. Sufficient conditions are given for a nested fuzzy query with RELATIVE quantifiers to be unnested for an efficient evaluation. An attribute-dependent interpretation is proposed to model the applications in which the meaning of a fuzzy term in an attribute must be interpreted with respect to values in other related attributes. Two necessary and sufficient conditions for a tuple to have a unique attribute-dependent interpretation are provided. We describe an interpretation system that allows queries to be processed based on the attribute-dependent interpretation of the data. Two techniques, grouping and shifting, are proposed to improve the implementation.>
Weining Zhang, Clement T. Yu, Bryan Reagan, Hiroshi Nakajima
ICDE2
1995 Similarity based Retrieval of Pictures Using Indices on Spatial Relationships
A. Prasad Sistla, Clement T. Yu, Chengwen Liu, King-Lup Liu
VLDB2
1995 Supporting Inheriance Using Subclass Assertions
Wei Sun 0002, Yibei Ling, Clement T. Yu
Inf. Syst.3
1995 A Theory of Translation From Relational Queries to Hierarchical Queries
abstract
In a heterogeneous database system, a query for one type of database system (i.e., a source query) may have to be translated to an equivalent query (or queries) for execution in a different type of database system (i.e., a target query). Usually, for a given source query, there is more than one possible target query translation. Some of them can be executed more efficiently than others by the receiving database system. Developing a translation procedure for each type of database system is time-consuming and expensive. We abstract a generic hierarchical database system (GHDBS) which has properties common to database systems whose schema contains hierarchical structures (e.g., System 2000, IMS, and some object-oriented database systems). We develop principles of query translation with GHDBS as the receiving database system. Translation into any specific system can be accomplished by a translation into the general system with refinements to reflect the characteristics of the specific system. We develop rules that guarantee correctness of the target queries, where correctness means that the target query is equivalent to the source query. We also provide rules that can guarantee a minimum number of target queries in cases when one source query needs to be translated to multiple target queries. Since the minimum number of target queries implies the minimum number of times the underlying system is invoked, efficiency is taken into consideration.>
Weiyi Meng, Clement T. Yu, Won Kim 0001
IEEE Trans. Knowl. Data Eng.2
1994 A Hybrid Transitive Closure Algorithm for Sequential and Parallel Processing
abstract
A new hybrid algorithm is proposed for well-formed path problems including the transitive closure problem. The CPU time for computation is O(ne), and blocking technique is incorporated to reduce the disk I/O cost in disk-resident environment. The new features of the new algorithm are that only parents sets instead of descendant sets are loaded in from disk, and the computation can be parallelized efficiently. Simulation results show that our algorithm is superior to other existing algorithms in sequential computation, and that linear speedup is achieved in parallel computation.>
Qi Yang 0011, Clement T. Yu, Chengwen Liu, Son Dao, Gaoming Wang, Tracy Pham
ICDE2
1994 Reasoning About Spatial Relationships in Picture Retrieval Systems
A. Prasad Sistla, Clement T. Yu, R. Haddad
VLDB2
1994 An Effiecient Way to Reestablish B+ Trees in a Distributed Environment
Wei Sun 0002, Weiyi Meng, Clement T. Yu, Won Kim 0001
Inf. Sci.3
1994 Efficient Query Processing for a Subset of Linear Recursive Binary Rules
abstract
We study the complexity of processing a class of rules called simple binary rule sets. The data referenced by the rules are stored in secondary memory. A necessary and sufficient condition that a simple binary rule set can be processed in a single pass of a file containing the base relations is given. Because not all simple binary rule sets can be processed in a single pass, a necessary and sufficient condition that a simple binary rule set can be processed by a constant number of passes is also given.>
Keh-Chang Guh, Clement T. Yu
IEEE Trans. Knowl. Data Eng.2
1994 Semantic Query Optimization for Tree and Chain Queries
abstract
Semantic query optimization, or knowledge-based query optimization, has received increasing interest in recent years. The authors provide an effective and systematic approach to optimizing queries by appropriately choosing semantically equivalent transformations. Basically, there are two different types of transformations: transformations by eliminating unnecessary joins, and transformations by adding/eliminating redundant beneficial/nonbeneficial selection operations (restrictions). A necessary and sufficient condition to eliminate a single unnecessary join is provided. We prove that it is /spl Nscr//spl Pscr/-/spl Cscr/omplete to eliminate as many unnecessary joins as possible for various types of acyclic queries with the exception of the closure chain queries whose query graphs are chains and all equi-join attributes are distinct. An algorithm is provided to minimize the number of joins in tree queries. This algorithm has an important property that, when applied to a closure chain query, it will yield an optimal solution with the time complexity O(n*m), where n is the number of relations referenced in the chain query, and m is the time complexity of a restriction closure computation.>
Wei Sun 0002, Clement T. Yu
IEEE Trans. Knowl. Data Eng.2
1993 Predict Query Processing Cost in a Distributed Datbase System
Weiyi Meng, Chengwen Liu, Wei Sun 0002, Clement T. Yu
DEXA4
1993 Construction of a Relational Front-end for Object-Oriented Database Systems
abstract
Proposes a solution for the construction of a relational front-end for object-oriented database systems (OODBs). Rules are provided to transform the structural part of an OODB scheme to an equivalent relational scheme to provide relational users with a relational view of the OODB scheme. A mechanism based on a relational predicate graph and an OODB predicate graph is provided to translate relational queries to OODB queries to allow relational users access to data stored in an OODB database system.>
Weiyi Meng, Clement T. Yu, Won Kim 0001, Gaoming Wang, Tracy Pham, Son Dao
ICDE2
1993 Performance Issues in Distributed Query Processing
abstract
The authors discuss various performance issues in distributed query processing. They validate and evaluate the performance of the local reduction (LR) the fragment and replicate strategy (FRS) and the partition and replicate strategy (PRS) optimization algorithms. The experimental results reveal that the choices made by these algorithms concerning which local operations should be performed, which relation should remain fragmented or which relation should be partitioned are valid. It is shown using experimental results that various parameters, such as the number of processing sites, partitioning speed relative to join speed, and sizes of the join relations, affect the performance of PRS significantly. It is also shown that the response times of query execution are affected significantly by the degree of site autonomy, interferences among processes, interface with the local database management systems (DBMSs) and communications facilities. Pipeline strategies for processing queries in an environment where relations are fragmented are studied.>
Chengwen Liu, Clement T. Yu
IEEE Trans. Parallel Distributed Syst.2
1992 Validation and Performance Evaluation of the Partition and Replicate Algorithm
abstract
The partition-and-replicate-strategy (PRS) algorithm for distributed query processing is evaluated and its performance is validated. Although in principle PRS is better than single-site processing, early experimental results indicate the contrary. Based on experimental results, the factor which causes performance deterioration is identified and a remedy is provided. As a result, it is shown that the PRS strategy outperforms single-site processing in a realistic environment and that various parameters, such as the number of processing sites, partitioning speed relative to join speed, and sizes of the join relations, affect the performance of the PRS strategy significantly. Among these parameters, the algorithm is most sensitive to the partition speed.>
Chengwen Liu, Clement T. Yu
ICDCS2
1992 Processing Hierarchical Queries in Heterogeneous Environment
abstract
The authors investigate principles of translating relational queries to hierarchical queries. Instead of using a specific hierarchical database system they abstract a generic hierarchical database system (GHDBS) which has properties common to database systems whose schema contain hierarchical structures. Principles of query translation with GHDBS as the receiving database system are developed. Rules that guarantee the correctness of the translated queries are described. Rules are provided that can guarantee a minimum number of target queries in a case when a user-submitted source query needs to be translated to multiple target queries.>
Weiyi Meng, Clement T. Yu, Won Kim 0001
ICDE2
1992 Query Optimisation in Distributed Object-Oriented Database Systems
abstract
In this paper, query processing and optimisation in distributed object-oriented database systems are discussed. The processing and optimisation of typical queries, called chain queries, in distributed object-oriented database systems are investigated in detail. An algorithm with complexity of O(n3 * hI) to minimise the total cost is provided using dynamic programming, where n is the number of classes referenced in the query, hI ≤ min (n + 2, h) and h is the number of sites in the network. A wide range of diversified issues are addressed and uniformly integrated into our basic solution to the problem. These issues include sorted states of classes; local processing of selections and projections, allowing multiple intermediate results; arbitrary target class at an arbitrary answer site; replicated data; class hierarchies (which captures the IS-A relationship among objects); different sites with different processing speeds, and communication lines between different sites with different transfer speeds. The uniformity of this algorithm under so many diversified situations strongly demonstrates the usefulness and the flexibility of the algorithm.
Wei Sun 0002, Weiyi Meng, Clement T. Yu
Comput. J.3
1992 Efficient Management of Materialized Generalized Transitive Closure in Centralized and Parallel Environments
abstract
A data structure is used to store materialized generalized transitive closure so that the evaluation of generalized transitive closure queries, deletions, and insertions of tuples can be performed efficiently in centralized and parallel environments. Some techniques to manage materialized transitive closure are presented and generalized to more general recursions. The proposed algorithms and the associated data structures are simple conceptually and in implementation. In a multiprocessor environment, the time complexities for insertion and deletion of the authors schemes are reduced. Only two rounds of communication are needed.>
Keh-Chang Guh, Clement T. Yu
IEEE Trans. Knowl. Data Eng.2
1991 Real Time Retrieval and Update of Materialized Transitive Closure
abstract
A data structure is used to store materialized transitive closure such that the evaluation of transitive closure, deletions and insertions of tuples can be performed efficiently. Experiments have been carried out on a Sun/3/180 system. It is verified experimentally and theoretically that it takes on the average O(m") to retrieve the ancestors/descendants of the given node, where m" is the number of ancestors/descendants of the given node, and it takes on the average O(m*m') to perform an insertion or a deletion of a tuple (a,b), where m is the number of ancestors of a+1 and m' is the number of descendants of b+1. It is shown that, when the data types is integer, retrieval of the ancestors/descendants of a given node takes no more than 0.0001 s; insertion/deletion of a tuple and the corresponding update involving m*m"=elements in the data structure takes approximately 0.07 s. When data type is a string of length 20, the corresponding retrieval time and insertion/deletion times are 0.0008 s and 1.5 s respectively.>
Keh-Chang Guh, Clement T. Yu
ICDE3
1991 Image Decompression: A Hybrid Image Decompressing Algorithm
abstract
Article A hybrid bilevel image decode algorithm for group 4 FAX Share on Authors: Chengjie Luo Department of Electrical Engineering and Computer Science, University of Illinois at Chicago, Chicago, IL Department of Electrical Engineering and Computer Science, University of Illinois at Chicago, Chicago, ILView Profile , Clement Yu Department of Electrical Engineering and Computer Science, University of Illinois at Chicago, Chicago, IL Department of Electrical Engineering and Computer Science, University of Illinois at Chicago, Chicago, ILView Profile Authors Info & Claims SIGIR '91: Proceedings of the 14th annual international ACM SIGIR conference on Research and development in information retrievalSeptember 1991 Pages 82–91https://doi.org/10.1145/122860.122869Online:01 September 1991Publication History 0citation245DownloadsMetricsTotal Citations0Total Downloads245Last 12 Months4Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Chengjie Luo, Clement T. Yu
SIGIR2
1991 An adaptive mixed relation decomposition algorithm for conjunctive retrieval queries
Wojtek Kozaczynski, Leszek Lilien, Clement T. Yu
Inf. Sci.3
1991 Data compression using word encoding with Huffman code
abstract
A technique for compressing large databases is presented. The method replaces frequent variable-length byte strings (words or word fragments) in the database by minimum-redundancy codes—Huffman codes. An essential part of the technique is the construction of the dictionary to yield high compression ratios. A heuristic is used to count frequencies of word fragments. A detailed analysis is provided of our implementaton in support of high compression ratios and efficient encoding and decoding under the constraint of a fixed amount of main memory. In each phase of our implementation, we explain why certain data structures or techniques are employed. Experimental results show that our compression scheme is very effective for compressing large databases of library records. © 1991 John Wiley & Sons, Inc.
Chengwen Liu, Clement T. Yu
J. Am. Soc. Inf. Sci.2
1990 Benchmarking two types of restricted transitive closure algorithms
abstract
The authors present and evaluate two algorithms-one linear and one logarithmic-for the computation of the restricted transitive closure of a binary database relation. The algorithms are implemented in a relational database management system (Ingres), and on equipment which is fairly common in today's database application environments. The performance evaluation reveals three important points. First, unlike the case of the complete transitive closure computations where the linear (seminaive) method is outperformed by the logarithmic methods, in the computation of the restricted transitive closure the opposite is true. Second, contrary to the popular belief that the algorithms run faster if the size of the intermediate result relations is decreased by deleting excess data, the fastest algorithms are those which attempt to delete no data. Unless deletions can be handled efficiently, their potential benefits are overshadowed by the cost incurred to perform them. Third, the operations union and difference are established as being significantly more expensive than the join operation in these algorithms.>
Anestis A. Toptsis, Clement T. Yu, Peter C. Nelson
COMPSAC2
1990 Query Optimization in Object-Oriented Database Systems
Wei Sun 0002, Weiyi Meng, Clement T. Yu
DEXA3
1990 Experiences with Distributed Query Processing
abstract
Different implementations of an experimental distributed query processing system which is constructed on top of existing database management systems (DBMSs) are presented. A performance evaluation is carried out. It is shown that the response times of query execution are affected significantly by the degree of site autonomy, interferences among processes, interface with the local DBMSs, and communications facilities. Too much autonomy for data replication increases response times. However lack of processing autonomy also increases response times. A goal is to identify the type of autonomy that facilitates query processing. It is shown that pipelining of the actions for data replication, which include local reduction, data transfer, and union fragments, improves performance whereas pipelining between join and data replication causes deterioration in performance because of competition of resources. However, the simulation results indicate that, in a multiprocessor environment, pipelining always improves performance provided that processes which compete for the same resource are assigned to different processors.>
Clement T. Yu, Chengwen Liu
ICDE1
1990 Improvements to an Algorithm for Equipartitioning
abstract
Modifications to a clustering algorithm in which objects are adaptively partitioned into clusters of equal size are described. The object migration automaton has the advantage of being conceptually simple and easy to implement. Unfortunately, the algorithm may exhibit slow convergence speed and in some cases may not converge at all. The algorithm is modified to provide remedies to these conditions. Through experimental results, the modifications are shown to yield a substantial speedup in convergence (while maintaining 100% accuracy), especially as the number of objects to be partitioned increases.>
William Gale, Sumit Das 0002, Clement T. Yu
IEEE Trans. Computers3
1990 Necessary and Sufficient Conditions to Linearize Double Recursive Programs in Logic Databases
abstract
Linearization of nonlinear recursive programs is an important issue in logic databases for both practical and theoretical reasons. If a nonlinear recursive program can be transformed into an equivalent linear recursive program, then it may be computed more efficiently than when the tranformation is not possible. We provide a set of necessary and sufficient conditions for a simple doubly recursive program to be equivalent to a simple linear recursive program. The necessary and sufficient conditions can be verified effectively.
Weining Zhang, Clement T. Yu, Daniel Troy
ACM Trans. Database Syst.2
1989 Distributed query processing a multiple database system
abstract
Mermaid is a testbed system which provides integrated access to multiple databases. Two query optimization algorithms have been developed for Mermaid. The semijoin algorithm tends to reduce the data transmission cost, while the replicate algorithm reduces the processing cost. An algorithm that integrates the features of these two algorithms to optimize the processing cost as well as the transmission cost is presented. A dynamic network environment is considered where processing speeds at each site and transmission speeds at each link can be variable. Moreover, distributed processing of aggregates is considered based on the functional dependency among the fragment attribute, the aggregate attribute, and the group-by attribute. Semantic information is utilized to obtain efficient query processing.>
Arbee L. P. Chen, David Brill, Marjorie Templeton, Clement T. Yu
IEEE J. Sel. Areas Commun.4
1989 Evaluation of transitive closure in distributed database systems
abstract
A special schema is constructed to represent the data of a binary relation so that the transitive closure can be evaluated efficiently in distributed database systems. The method is economical in communication cost and at the same time preserves the fast access paths for efficient parallel processing of transitive closure at local sites. Updating of data is also discussed.>
Keh-Chang Guh, Clement T. Yu
IEEE J. Sel. Areas Commun.2
1989 Automatic Knowledge Acquisition and Maintenance for Semantic Query Optimization
abstract
The authors present an approach to acquiring knowledge from previously processed queries. By using newly acquired knowledge together with given semantic knowledge, it is possible to make the query processor and/or optimizer more intelligent so that future queries can b processed more efficiently. The acquired knowledge is in the form of constraints. While some constraints are to be enforced for all database states, others are known to be valid for the current state of the database. The former constraints are statistic integrity constraints, while the latter are called dynamic integrity constraints. Some situations in which certain dynamic semantic constraints can be automatically extracted are identified. This automatic tool for knowledge acquisition can also be used as an interactive tool for identifying potential static integrity constraints. The concept of minimal knowledge base is introduced, and a method to maintain the knowledge base is presented. An algorithm to compute the restriction (selection) closure, i.e. all deductible restrictions, from a given set of restrictions, join predicates (as given in a query), and constraints is given.>
Clement T. Yu, Wei Sun 0002
IEEE Trans. Knowl. Data Eng.1
1989 A Framework for Effective Retrieval
abstract
The aim of an effective retrieval system is to yield high recall and precision (retrieval effectiveness). The nonbinary independence model, which takes into consideration the number of occurrences of terms in documents, is introduced. It is shown to be optimal under the assumption that terms are independent. It is verified by experiments to yield significant improvement over the binary independence model. The nonbinary model is extended to normalized vectors and is applicable to more general queries. Various ways to alleviate the consequences of the term independence assumption are discussed. Estimation of parameters required for the nonbinary independence model is provided, taking into consideration that a term may have different meanings.
Clement T. Yu, Weiyi Meng
ACM Trans. Database Syst.1
1989 Linearization of Nonlinear Recursive Rules
abstract
The problem of converting a simple nonlinear recursive logic query into an equivalent linear one is considered. A general method is given to transform a nonlinear rule into a sequence of linear ones. For efficient processing it is necessary to convert a nonlinear rule into a single linear rule. For such a conversion, a necessary and sufficient condition is provided for a type of doubly recursive rule to be equivalent to the resulting linear rule. It is also shown that a restricted type of higher order recursive rule is equivalent to the linear rule obtained by its conversion.>
Daniel Troy, Clement T. Yu, Weining Zhang
IEEE Trans. Software Eng.2
1989 Partition Strategy for Distributed Query Processing in Fast Local Networks
abstract
A partition-and-replicate strategy for processing distributed queries referencing no fragmented relation is sketched. An algorithm is given to determine which relation and which copy of the relation is to be partitioned into fragments, how the relation is to be partitioned, and where the fragments are to be sent for processing. Simulation results show that the partition strategy is useful for processing queries in fast local network environments. The results also show that the number of partitions does not need to be large. The use of semijoins in the partition strategy is discussed. A necessary and sufficient condition for a semijoin to yield an improvement is provided.>
Clement T. Yu, Keh-Chang Guh, David Brill, Arbee L. P. Chen
IEEE Trans. Software Eng.1
1988 Adaptive Algorithms for Balanced Multidimensional Clustering
abstract
The G-K-D tree (generalized K-D tree) method aims at reducing the average number of data page accesses per query, but it ignores the cost of index search. The authors propose two adaptive algorithms that take into consideration both data page access cost and index page access cost. It attempts to find a minimum total cost. Experimental results indicate that the proposed algorithms are superior to the G-K-D tree method.>
Clement T. Yu, Tsang Ming Jiang
ICDE1
1988 Two Learning Schemes in Information Retrieval
abstract
Two methods are given to improve weighting schemes by using relevance information of a set of queries. The first method is to estimate parameter values of two independence models in information retrieval — the binary independence model and the non-binary independence model. The parameters estimated here are used to calculate optimal weights for terms in a different set of queries. Performance of this estimation is compared to the inverse document frequency method, the cosine measure, and the statistical similarity measure. The second method is to learn optimal weights of the non-binary independence model adaptively by a learning formula. Experiments are performed on three different document collections CISI, MEDLARS, and CRN4NUL for both methods, and results are reported. Both methods show improvements compared to the existing weighting schemes. Experimental results show that the second method gives slightly better performance than the first one, and has simpler implementation.
Clement T. Yu, Hirotaka Mizuno
SIGIR1
1987 Efficient Recursive Query Processing using Wavefront Methods
abstract
In this paper, we study the optimization of linear recursive queries using wavefront methods. The following results are obtained.(i) In spite of seemingly reasonable approach of the wavefront methods, certain linear recursive queries are not processed efficiently or correctly.(ii) A characterization of the expressions generated by linear recursive rules is given. Properties of the expressions will be useful for efficient processing of linear recursive queries.(iii) Conditions for efficient processing of linear recursive rules using wavefront methods with the properties given in (ii) are provided.
Clement T. Yu, Weining Zhang
ICDE1
1987 Models of IR (Panel)
Vijay Raghavan 0001, M. Gordon, Robert R. Korfhage, Clement T. Yu
SIGIR4
1987 A Necessary Condition for a Doubly Recursive Rule to be Equivalent to a Linear Recursive Rule
abstract
Nonlinear recursive queries are usually less efficient in processing than linear recursive queries. It is therefore of interest to transform non-linear recursive queries into linear ones. We obtain a necessary and sufficient condition for a doubly recursive rule of a certain type to be logically equivalent to a single linear recursive rule obtained in a specific way.
Weining Zhang, Clement T. Yu
SIGMOD Conference2
1987 Algorithms to Process Distributed Queries in Fast Local Networks
abstract
We propose a scheme to make use of semantic information to process distributed queries locally without data transfer with respect to the join clauses of the query. Since not all queries can be processed without data transfer, we give an algorithm to recognize the "locally processable queries." For nonlocally processable queries, a simple "fragment and replicate" algorithm is used. The algorithm chooses a relation to remain fragmented at the sites where they are situated while replicating the other relations at those sites. Our algorithm determines the chosen relation and the chosen copy of every fragment of the chosen relation such that the minimum response time is obtained. The algorithm runs in linear time. If the fragments of the relation are allowed to be processed in other sites, then the problem is NP hard. Two heuristics are given for that situation. They are compared to the optimal situation. Experimental results show that the strategies produced by the heuristics have small errors relative to the optimal strategy.
Clement T. Yu, Keh-Chang Guh, Weining Zhang, Marjorie Templeton, David Brill, Arbee L. P. Chen
IEEE Trans. Computers1
1986 Adaptive Techniques for Distributed Query Optimization
abstract
We propose new adaptive techniques for distributed query optimization. These techniques are divided into two groups: the ones that improve efficiency of query execution (directly) and the ones that improve cost estimations for query execution strategies. Some of the proposed techniques utilize semantic information and knowledge acquisition to adapt to the environment. The latter, in contrast to the former, is not a well-established idea. This is a disturbing fact since knowledge acquisition can give significant improvements in performance of a query optimization algorithm. Performing analysis manually is extrernely time consuming and tedious. Therefore, some learning capacity should be added to the system. Some knowledge acquisition techniques that result in adaptive (dynamic) adjustment to run-time changes are proposed.
Clement T. Yu, Leszek Lilien, Keh-Chang Guh, Marjorie Templeton, David Brill, Arbee L. P. Chen
ICDE1
1986 Partitioning Relation for Parallel Processing in Fast Local Networks
Clement T. Yu, Keh-Chang Guh, David Brill, Arbee L. P. Chen
ICPP1
1986 Probabilistic Models for Document Retrieval: A Comparison of Performance on Experimental and Synthetic Databases
abstract
Probabilistic document retrieval systems consistent with the two Poisson independence model outperforms the binary independence model if the terms are distributed as described by the model's assumptions. The Two Poisson Effectiveness Hypothesis suggests that retrieval models based upon the two Poisson model will outperform binary independent models when used on a “real-world” database, where independence and two Poisson term occurrence distributions fail to hold, because the added information obtained from incorporating term frequency information will more than compensate for the non-Poisson distributions of terms. Searches of the MED1033 database suggest that if terms are not independent and frequencies of term occurrence are not distributed in a two Poisson manner, the binary independence sequential retrieval model outperforms the two Poisson independence retrieval model.
Robert M. Losee, Abraham Bookstein, Clement T. Yu
SIGIR3
1986 Non-Binary Independence Model
abstract
Article Free Access Share on Non-binary independence model Authors: C. T. Yu Department of Electrical Engineering & Computer Science, University of Illinois at Chicago, Chicago, Illinois Department of Electrical Engineering & Computer Science, University of Illinois at Chicago, Chicago, IllinoisView Profile , T. C. Lee Department of Electrical Engineering & Computer Science, University of Illinois at Chicago, Chicago, Illinois Department of Electrical Engineering & Computer Science, University of Illinois at Chicago, Chicago, IllinoisView Profile Authors Info & Claims SIGIR '86: Proceedings of the 9th annual international ACM SIGIR conference on Research and development in information retrievalSeptember 1986 Pages 265–268https://doi.org/10.1145/253168.253224Published:01 September 1986Publication History 3citation584DownloadsMetricsTotal Citations3Total Downloads584Last 12 Months33Last 6 weeks5 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Clement T. Yu, T. C. Lee
SIGIR1
1985 Optimization of a Hierarchical File Organization for Spelling Correction
abstract
A spelling program using a hierarchically organized file seems to be promising, since it can correct more than common typing mistakes. However, its speed of detecting spelling errors in the inputs is rather slow. Here some techniques of modifying the program to improve the speed are presented.
Tetsuro Ito, Clement T. Yu
SIGIR2
1985 Adaptive Document Clustering
abstract
Article Free Access Share on Adaptive document clustering Authors: C. T. Yu Department of Electrical Engineering & Computer Science, University of Illinois at Chicago, Chicago, IL Department of Electrical Engineering & Computer Science, University of Illinois at Chicago, Chicago, ILView Profile , Y. T. Wang Department of Electrical Engineering & Computer Science, University of Illinois at Chicago, Chicago, IL Department of Electrical Engineering & Computer Science, University of Illinois at Chicago, Chicago, ILView Profile , C. H. Chen Department of Electrical Engineering & Computer Science, University of Illinois at Chicago, Chicago, IL Department of Electrical Engineering & Computer Science, University of Illinois at Chicago, Chicago, ILView Profile Authors Info & Claims SIGIR '85: Proceedings of the 8th annual international ACM SIGIR conference on Research and development in information retrievalJune 1985Pages 197–203https://doi.org/10.1145/253495.253525Published:05 June 1985Publication History 21citation390DownloadsMetricsTotal Citations21Total Downloads390Last 12 Months11Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Clement T. Yu
SIGIR1
1985 Adaptive Information System Design: One Query at a Time
abstract
Article Free Access Share on Adaptive information system design: one query at a time Authors: C. T. Yu Dept. of EECS, University of Illinois at Chicago Dept. of EECS, University of Illinois at ChicagoView Profile , C. H. Chen Bell Communication Research Bell Communication ResearchView Profile Authors Info & Claims SIGMOD '85: Proceedings of the 1985 ACM SIGMOD international conference on Management of dataMay 1985 Pages 280–290https://doi.org/10.1145/318898.318924Published:01 May 1985Publication History 4citation321DownloadsMetricsTotal Citations4Total Downloads321Last 12 Months11Last 6 weeks4 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Clement T. Yu
SIGMOD Conference1
1985 Adaptive Record Clustering
abstract
An algorithm for record clustering is presented. It is capable of detecting sudden changes in users' access patterns and then suggesting an appropriate assignment of records to blocks. It is conceptually simple, highly intuitive, does not need to classify queries into types, and avoids collecting individual query statistics. Experimental results indicate that it converges rapidly; its performance is about 50 percent better than that of the total sort method, and about 100 percent better than that of randomly assigning records to blocks.
Clement T. Yu, Cheing-Mei Suen, Man-Keung Siu
ACM Trans. Database Syst.1
1985 Query Processing in a Fragmented Relational Distributed System: Mermaid
abstract
This paper describes the query optimizer of the Mermaid system which provides a user with a unified view of multiple preexisting databases which may be stored under different DBMS's. The algorithm is designed for databases which may contain replicated or fragmented relations and for users who are primarily making interactive, ad hoc queries. Although the implementation of the algorithm is a front-end system, not an integrated distributed DBMS, it should be applicable to a distributed DBMS also.
Clement T. Yu, Chin-Chen Chang 0005, Marjorie Templeton, David Brill, Eric Lund
IEEE Trans. Software Eng.1
1985 Adaptive File Allocation in Star Computer Network
abstract
In this paper, we study the allocation of files in a star network. Unlike previous algorithms which assume that files are independently accessed and independently assigned, the interaction of files during the processing of queries is directly incorporated into our cost model. We present an adaptive algorithm, which is much faster than existing algorithms on file allocation, obtains solutions which are on the average only 0.1 percent away from the optimal solutions, and possesses many desirable properties such as the satisfaction of some necessary and sufficient conditions for file allocation.
Clement T. Yu, Man-Keung Siu
IEEE Trans. Software Eng.1
1984 Distributed Query Processing Strategies in Mermaid, A Frontend to Data Management Systems
abstract
The Mermaid testbed system has been developed as part of an ongoing research program at SDC to explore issues in distributed data management. The Mermaid testbed provides a uniform front end that makes the complexity of manipulating data in distributed heterogeneous databases under various data management systems (DBMSs) transparent to the user. It is being used to test query optimization algorithms as well as user interface methodology.
David Brill, Marjorie Templeton, Clement T. Yu
ICDE3
1984 Optimization of Distributed Tree Queries
Clement T. Yu, Z. Meral Özsoyoglu
J. Comput. Syst. Sci.1
1983 Evaluation of The 2-Poisson Model as a Basis for Using Term Frequency Data in Searching
abstract
The early work on the probabilistic models of retrieval assumed that the document representation is binary, indicating only the presence or absence of index terms. The 2-Poisson (TP) model which was proposed as a model of how the occurrence frequency of specialty words in a collection is distributed, has since been used to develop retrieval strategies that incorporate term frequency information. This work investigates the use of the TP model, in this context, further. It is shown that the search effectiveness, when no relevance information is assumed, can be further enhanced by using this model. Furthermore, when the term weights proposed in this work are used in conjunction with weights known as term significance weights, the results are very encouraging.
Vijay Raghavan 0001, Hong-Pao Shi, Clement T. Yu
SIGIR3
1983 On the Design of a Query Processing Strategy in a Distributed Database Environment
abstract
An algorithm is given to process a given query in a fragmented distributed data base environment. Unlike previous algorithms, it has the following desired features.(1) It makes use of redundant relations to reduce communication cost;(2) a copy of each relation referenced by the query is selected so that the set of relations are contained in the minimum number of sites;(3) an efficient algorithm to process fragments is provided;(4) all relations that need not be sent to the assembly site to produce the answer are identified. Thus, unnecessary sending of these relations across sites and processing on these relations, which are common in earlier algorithms, are avoided;(5) useless semi-joins are discarded and worse semi-joins are replaced by better ones;(6) a process to estimate the cost and the benefit of a semi-join, based on dynamic execution of semi-joins is introduced. It is expected that the new process is more accurate than earlier estimation process.The algorithm is easy to implement and is operational.
Clement T. Yu
SIGMOD Conference1
1983 File Allocation in Distributed Databases with Interaction between Files
Clement T. Yu, Man-Keung Siu
VLDB1
1982 Some Esitmation Problems in Distributed Query Processing
Clement T. Yu, Y. C. Lin
ICDCS1
1982 An Evaluation of Term Dependence Models in Information Retrieval
Gerard Salton, Chris Buckley, Clement T. Yu
SIGIR3
1982 On the Construction of Feedback Queries
abstract
Optimal feedback queries in information retrieval systems are constructed An optimal retrieval rule is derived using the Neyman-Pearson decision rule Three probabdistic models and the optimal queries to be used in the models are presented Parameters which are required to construct these queries are esUmated on the basis of relevance information from the user about the retrieved documents Finally, the effects on retrieval performance of deleting a term from the optimal query in one of the three models are analyzed Categories and Subject Descriptors.G 3 [Mathematics of Computing]" Probability and Statistics, H 3 3 ]
D. Chow, Clement T. Yu
J. ACM2
1982 Term Weighting in Information Retrieval Using the Term Precision Model
abstract
At3STRACT It iS known that the use of weighted, as opposed to binary, content identifiers attached to the records of an information file improves the effectiveness of the retrieval operations Under well-defined conditions the term precision offers the best possible term weighting system A mathematscal model is used in the present study to relate the term precision weights to the frequency of occurrence of the terms in a given document collecuon and to the number of relevant documents a user wishes to retrieve in response to a query This provides for the assignment of user-dependent weights to the content identifiers and relates the term precision weights to other well-known term weighting systems Categories and Subject Descriptors.H 3 1 [
Clement T. Yu, Gerard Salton
J. ACM1
1982 A Clustered Search Algorithm Incorporating Arbitrary Term Dependencies
abstract
The documents in a database are organized into clusters, where each cluster contains similar documents and a representative of these documents. A user query is compared with all the representatives of the clusters, and on the basis of such comparisons, those clusters having many close neighbors with respect to the query are selected for searching. This paper presents an estimation of the number of close neighbors in a cluster in relation to the given query. The estimation takes into consideration the dependencies between terms. It is demonstrated by experiments that the estimate is accurate and the time to generate the estimate is small.
Clement T. Yu
ACM Trans. Database Syst.2
1981 An Approach to Probabilistic Retrieval
abstract
The objective is to relate the effectiveness of retrieval, the fuzzy set concept and the processing of Boolean query. The use of a probabilistic retrieval scheme is motivated. It is found that there is a correspondence between probabilistic retrieval schmes and fuzzy sets. A fuzzy set corresponding to a potentially optimal probabilistic retrieval scheme is obtained. Then the retrieval scheme for the fuzzy set is constructed.
Clement T. Yu
SIGIR1
1981 The measurement of term importance in automatic indexing
abstract
Abstract The frequency characteristics of terms in the documents of a collection have been used as indicators of term importance for content analysis and indexing purposes. In particular, very rare or very frequent terms are normally believed to be less effective than medium‐frequency terms. Recently automatic indexing theories have been devised that use not only the term frequency characteristics but also the relevance properties of the terms. The major term‐weighting theories are first briefly reviewed. The term precision and term utility weights that are based on the occurrence characteristics of the terms in the relevant, as opposed to the nonrelevant, documents of a collection are then introduced. Methods are suggested for estimating the relevance properties of the terms based on their overall occurrence characteristics in the collection. Finally, experimental evaluation results are shown comparing the weighting systems using the term relevance properties with the more conventional frequency‐based methodologies.
Gerard Salton, Harry Wu, Clement T. Yu
J. Am. Soc. Inf. Sci.3
1981 A Comparison of the Stability Characteristics of Some Graph Theoretic Clustering Methods
abstract
Assessing the stability of a clustering method involves the measurement of the extent to which the generated clusters are affected by perturbations in the input data. A measure which specifies the disturbance in a set of clusters as the minimum number of operations required to restore the set of modified clusters to the original ones is adopted. A number of well-known graph theoretic clustering methods are compared in terms of their stability as determined by this measure. Specifically, it is shown that among the clustering methods in any of several families of graph theoretic methods, clusters defined as the connected components are the most stable and the clusters specified as the maximal complete subgraphs are the least stable. Furthermore, as one proceeds from the method producing the most narrow clusters (maximal complete subgraphs) to those producing relatively broader clusters, the clustering process is shown to remain at least as stable as any method in the previous stages. Finally, the lower and the upper bounds for the measure of stability, when clusters are defined as the connected components, are derived.
Vijay Raghavan 0001, Clement T. Yu
IEEE Trans. Pattern Anal. Mach. Intell.2
1981 A Generalized Counter Scheme
Man-Keung Siu, Clement T. Yu
Theor. Comput. Sci.3
1980 An Approximation Algorithm for a File-Allocation Problem in a Hierarchical Distributed System
abstract
A file allocation problem in a hierarchical distributed computer network is examined. It is shown that the problem is NP-hard. An approximation algorithm is suggested. It is estimated that the approximation algorithm has a high chance of obtaining the optimal solution. Experimental results show that the difference between the optimal solution and the solution generated by the approximation algorithm is no more than 4% away from the optimal solution.
Clement T. Yu
SIGMOD Conference2
1979 Performance Analysis of three Related Assignment Problems
abstract
The relative placement of records in a database has a significant impact on the overall system performance. In this paper we study three related assignment problems. They are (1) the assignment of records to devices so as to minimize the expected completion time and to maximize the expected utilization of the devices, (2) the assignment of records to pages or blocks in secondary memory so as to minimize the number of blocks to be accessed, and (3) the analysis of the performance of a hashing scheme under the assumption that the record-to-bucket probability varies from one bucket to another. The solutions to these problems are obtained by studying an occupancy problem. In each case, an optimal solution is obtained.
Clement T. Yu, Man-Keung Siu, Z. Meral Özsoyoglu
SIGMOD Conference1
1979 On models of information retrieval processes
Clement T. Yu, Man-Keung Siu
Inf. Syst.1
1979 Experiments on the Determination of the Relationships Between Terms
abstract
The retrieval effectiveness of an automatic method that uses relevance judgments for the determination of positive as well as negative relationships between terms is evaluated. The term relationships are incorporated into the retrieval process by using a generalized similarity function that has a term match component, a positive term relationship component, and a negative term relationship component. Two strategies, query partitioning and query clustering, for the evaluation of the effectiveness of the term relationships are investigated. The latter appears to be more attractive from linguistic as well as economic points of view. The positive and the negative relationships are verified to be effective both when used individually, and in combination. The importance attached to the term relationship components relative to that of term match component is found to have a substantial effect on the retrieval performance. The usefulness of discriminant analysis as a technique for determining the relative importance of these components is investigated.
Vijay Raghavan 0001, Clement T. Yu
ACM Trans. Database Syst.2
1978 Experiments on the Determination of the Reltionships Between Terms
abstract
The retrieval effectiveness of an automatic method that uses relevance judgements for the determination of positive as well as negative relationships between terms is evaluated. The term relationships are incorporated into the retrieval process by using a generalized similarity function that has a term match component, a positive term relationship component, and a negative term relationship component. Two strategies, query partitioning and query clustering, for the evaluation of the effectiveness of the term relationships are investigated. The latter appears to be more attractive from linguistic as well as economic points of view. The positive and the negative relationships are verified to be effective both when used individually, and in combination. The importance attached to the term relationship components relative to that of term match component is found to have a substantial effect on the retrieval performance. The usefulness of discriminant analysis as a technique for determining the relative importance of these components is investigated.
Clement T. Yu, Vijay Raghavan 0001
SIGIR1
1978 Effective Automatic Indexing Using Term Addition and Deletion
abstract
In mformaUon retrieval indexing is the task consisting of the assignment to stored records and mcommg mformatton requests of content ~dent~fiers capable of representing record or query content If the mdexmg is performed automatically and the records are wntten documents, an mmal set of index terms might be chosen by taking words extracted from document roles or abstracts, this mmal term assignment might then be improved by addmg related terms chosen from a thesaurus, by deletmg extraneous or marginal terms, and by replacing smgle terms by term combmaaons and phrases.In the present study formal proofs are given of the retrieval effectiveness under well-defined condmons of mdexmg policies based on the use of single terms, term additions and deletions, and term combmaaons or phrases.
Clement T. Yu, Gerard Salton, Man-Keung Siu
J. ACM1
1978 On the Estimation of the Number of Desired Records with Respect to a Given Query
abstract
The importance of the estimation of the number of desired records for a given query is outlined. Two algorithms for the estimation in the “closest neighbors problem” are presented. The numbers of operations of the algorithms are Ο ( ml 2 ) and Ο ( ml ), where m is the number of clusters and l is the “length” of the query.
Clement T. Yu, Man-Keung Siu
ACM Trans. Database Syst.1
1978 On a Partitioning Problem
abstract
This paper investigates the problem of locating a set of “boundary points” of a large number of records. Conceptually, the boundary points partition the records into subsets of roughly the same number of elements, such that the key values of the records in one subset are all smaller or all larger than those of the records in another subset. We guess the locations of the boundary points by linear interpolation and check their accuracy by reading the key values of the records on one pass. This process is repeated until all boundary points are determined. Clearly, this problem can also be solved by performing an external tape sort. Both analytical and empirical results indicate that the number of passes required is small in comparison with that in an external tape sort. This kind of record partitioning may be of interest in setting up a statistical database system.
Clement T. Yu, Man-Keung Siu
ACM Trans. Database Syst.1
1977 A Study on the Protection of Statistical Data Bases
abstract
We study a number of protection schemes with respect to their effectiveness in providing security for statistical data bases, their feasibility and their ease of implementation. A new method is proposed, and two implementations presented. One implementation guarantees perfect protection against leakage of information about individuals; the other requires very little implementation effort, but has a small probability of leakage.
Clement T. Yu, Francis Y. L. Chin
SIGMOD Conference1
1977 A Note on a Multidimensional Searching Problem
Vijay Raghavan 0001, Clement T. Yu
Inf. Process. Lett.2
1977 Analysis of Effectiveness of Retrieval in Clustered Files
abstract
article Free Access Share on Analysis of Effectiveness of Retrieval in Clustered Files Authors: C. T. Yu Department of Computer Science, University of Alberta, Edmonton, Canada T6G 2E1 Department of Computer Science, University of Alberta, Edmonton, Canada T6G 2E1View Profile , W. S. Luk Department of Computer Science, University of Alberta, Edmonton, Canada T6G 2E1 Department of Computer Science, University of Alberta, Edmonton, Canada T6G 2E1View Profile Authors Info & Claims Journal of the ACMVolume 24Issue 4Oct. 1977 pp 607–622https://doi.org/10.1145/322033.322039Published:01 October 1977Publication History 12citation294DownloadsMetricsTotal Citations12Total Downloads294Last 12 Months1Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Clement T. Yu
J. ACM1
1977 Single-pass method for determining the semantic relationships between terms
abstract
Abstract A fast single‐pass method for the automatic determination of the semantic relationships between terms is presented. The computing time required for the method is small enough for it to be feasible in a practical environment. The experimental results obtained indicate that the method is effective in the assessment of the relationships between terms. The improvement achieved in retrieval performance over simple keyword matching is quite significant.
Clement T. Yu, Vijay Raghavan 0001
J. Am. Soc. Inf. Sci.1
1976 On the Complexity of Finding the Set of Candidate Keys for a Given Set of Functional Dependencies
Clement T. Yu, D. T. Johnson
Inf. Process. Lett.1
1976 Automatic indexing using term discrimination and term precision measurements
Gerard Salton, Anita Wong, Clement T. Yu
Inf. Process. Manag.3
1976 A Statistical Model for Relevance Feedback in Information Retrieval
abstract
A statistical model is presented for the investigation of a practical method used in relevance feedback. A necessary and sufficient condition for the two parameters used in this method to define a better query than the original query is given. A region in the plane of the parameters is shown to satisfy the sufficient condition. While the points for producing optimal queries are not exactly located, they are shown to be lying on a finite portion of a hyperbola. Experimental results support some of the theoretical findings.
Clement T. Yu, T. Y. Cheung
J. ACM1
1976 Precision Weighting - An Effective Automatic Indexing Method
abstract
A great many automatic indexing methods have been implemented and evaluated over the last few years, and automatic procedures comparable in effectiveness to conventional manual ones are now easy to generate. Two drawbacks of the available automatic indexing methods are the absence of reliable linguistic inputs during the indexing process and the lack of formal, analytical proofs concerning the effectiveness of the proposed methods. The precision weighting procedure described in the present study uses relevance criteria to weight the terms occurring in user queries as a function of the balance between relevant and nonrelevant documents in which these terms occur; this approximates a semantic know-how of term importance. Formal mathematical proofs are given under well-defined conditions of the effectiveness of the method.
Clement T. Yu, Gerard Salton
J. ACM1
1976 The stability of two common matching functions in classification with respect to a proposed measure
abstract
Abstract A measure for the quantification of the changes in classification under small changes in data is proposed. A comparison of two common matching functions, namely the single matching function and the cosine function, is made with respect to the measure. The sensitivities of clusters defined as connected components and as maximal complete subgraphs are also compared.
Clement T. Yu
J. Am. Soc. Inf. Sci.1
1975 A Formal Construction of Term Classes
abstract
The computational complexity of a formal process for the construction of term classes is examined.While the process is proved to be difficult computatlonally, heumstic methods are applied Experimental results are obtained to illustrate the maximum possible improvement in system performance of retrieval using the formal construction over simple term retrieval
Clement T. Yu
J. ACM1
1975 A theory of term importance in automatic text analysis
abstract
Abstract A good deal of work has been done over the years in an attempt to use statistical or probabilistic techniques as a basis for automatic indexing and content analysis. (1–10) Unfortunately, many of these methods are lacking in effectiveness, and the more refined procedures are computationally unattractive. A new technique, known as discrimination value analysis, ranks the text words in accordance with how well they are able to discriminate the documents of a collection from each other; that is, the value of a term depends on how much the average separation between individual documents changes when the given term is assigned for content identification. The best words are those which achieve the greatest separation. The discrimination value analysis is computationally simple, and it assigns a specific role in content analysis to single words, juxtaposed words and phrases, and word groups or thesaurus categories. Experimental results are given showing the effectiveness of the technique.
Gerard Salton, Chung-Shu Yang, Clement T. Yu
J. Am. Soc. Inf. Sci.3
1974 A methodology for the construction of term classes
Clement T. Yu
Inf. Storage Retr.1
1974 A clustering algorithm based on user queries
abstract
Abstract A clustering algorithm which is tree‐like in structure, and is based on user queries, is presented. It is compared to Bonner's Method, Rocchio's Method, Dattola's Method and the Single Link Method in three different aspects, namely system effectiveness, system efficiency and the time required for clustering. Experimental results using the Cranfield 424 collection indicate that the proposed method is superior to the other methods.
Clement T. Yu
J. Am. Soc. Inf. Sci.1