EDBT 2026 Demo / reviewers in the wild / expert
Yiu-Kai Ng
dblp:08/4453 · also Yiu-Kai Dennis Ng
· DBLP profile ↗
81ranked-venue papers
9as first author
8since 2021 · last 2025
0000-0002-5680-2796ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 58 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 25 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 1 since 2021Theory of computation · 2Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Emotion Vector-Based Fine-Tuning of Large Language Models for Age-Aware Teenage Book Recommendations
Kate Hill, Yiu-Kai Ng, Joey Sherrill |
RecSys | 2 |
| 2024 | A Robust Random Search Approach for Matching Formulas in Math Information Retrieval SystemsabstractMath is a major contributor to many areas of study, and gives someone skills that (s)he can use across other subjects and different job roles. Unfortunately, a recent study from the National Assessment of Educational Progress shows that no more than 26% of 12thgraders in the USA have been rated proficient in math since 2005, and COVID-19 only made the situation worse. In principle, appropriate online searching could promote learning of individual math concepts to help surmount the learning gap. In practice, however, current online searching works poorly for math. While traditional information retrieval systems identify semantically related documents outside of math, such systems were not designed for handling math formulas. Although some work has been done on Mathematical Information Retrieval (MIR) recently, little has focused specifically on developing indexing schemas to quickly search for and retrieve math formulas contained within math questions and answers. The objective of indexing symbols and notations used in math equations is to organize and categorize math information in a way that makes it easier to retrieve and access relevant answers to math questions. To achieve this objective, we propose a robust random search approach for retrieving math information, offering an optimal solution to speed up the process of searching huge volume of math archive. Our design goals of indexing math equations include fast matching math answers to questions, reducing disk input/output, and optimizing the process of solving math questions with suitable answers by enhancing its processing speed. Megan Shellman, Kate Hill, Yiu-Kai Ng |
ICTAI | 3 |
| 2024 | Boolean interpretation, matching, and ranking of natural language queries in product selection systems
Matthew Moulton, Yiu-Kai Ng |
Discov. Comput. | 2 |
| 2023 | Easy to Find: A Natural Language Query Processing System on Advertisements Using an Automatically Populated Database
Yiu-Kai Ng |
WEBIST | 1 |
| 2023 | Read to grow: exploring metadata of books to make intriguing book recommendations for teenage readers
Yiu-Kai Ng |
Knowl. Inf. Syst. | 1 |
| 2022 | A Hybrid Approach for Summarizing User Reviews Based on KL-Divergence and Deep LearningabstractIn the modern day the information age makes available instant access to knowledge that would have been difficult or impossible to find previously. This leads to the problem of exploring too many user reviews on popular products for potential buyers to spend adequate time to read and extract the most salient product details and opinions of previous buyers. Multi-Document Summarization, the process of abbreviating the content of a collection of documents into a short, singular summary, is an important tool for expressing the content of these source while maintaining a level of linguistic quality. We can take advantage of the summarization approach to reduce the time necessary to read a collection of topically-related user reviews to locate their desired information needs by viewing a short summary of these reviews. In this paper, we propose a new summarizer that utilizes both extractive and abstractive techniques to yield such a high-quality summary. Our hybrid summarization approach is novel, since it includes the development of an innovative extractive summarizer that simply employes the KL-divergence measure to quantify sentences that capture the major contents of a set of user reviews on a particular product and retains these sentences in a summary that contains the most salient information of the set of reviews. Hereafter, the extractive summary is passed on to an adapted BERT deep learning model to improve the linguistic quality of the original summary and generate a new shorter text that conveys the most critical information from the original one. The proposed hybrid summarization approach is simple and easy to understand. Our summarizer has been compared against three existing summarizers and shown to outperform each one of them in terms of both expressing relevant content and linguistic features of a set of user reviews. Nathaniel Benham, Yiu-Kai Ng |
ICTAI | 3 |
| 2021 | A Simple, Concise, Query-based Approach to News Article Summarization Using Sentence ScoringabstractWith the increasing amount of information being digitized and the growing connectedness of the world, access to news and their intricated information is becoming more vital. Because of this growing need, creating news article summaries is becoming an increasingly important task to allow people to access essential information quickly. However, current summarization approaches require complex, taxing algorithms that cannot be seamlessly adopted for others to implement at the speed that we need. To remedy this, we have designed an elegant approach that allows the utilizing technology to quickly employ a multinomial classifier and sentence scoring of news articles to help with querying and filtering news to allow users to obtain a brief, efficient summary of what the articles entail. The multinomial classifier achieves very effective classification of news articles for summarization. Using various complementary sentence scores, we are able to accurately determine sentences that provide the most informative contents with respect to a user query Q. Through the use of this classification and summarization, we allow information of Q to be readily available. Experimental results verify that our news article summarization approach is effective and efficient in creating high-quality summaries. In addition, the conducted empirical study demonstrates that our summarization approach outperform a significant number of DUC summarizers. Megan Thornton, Sophie Gao, Yiu-Kai Ng |
ICTAI | 3 |
| 2021 | Looking for Jobs? Matching Adults with Autism with Potential Employers for Job OpportunitiesabstractAdults with autism face many difficulties when finding employment, such as struggling with interviews and needing accommodating environments for sensory issues. Autistic adults, however, also have unique skills to contribute to the workplace that companies have recently started to seek after, such as loyalty, close attention to detail, and trustworthiness. To work around these difficulties and help companies find the talent they are looking for we have developed a job-matching system. Our system is based around the stable matching of the Gale-Shapley algorithm to match autistic adults with employers after estimating how both adults with autism and employers would rank the other group. The system also uses filtering to approximate a stable matching even with a changing pool of users and employers, meaning the results are resistant to change as the result of competition. Such a system would be of benefit to both adults with autism and employers and would advance knowledge in recommender systems that match two parties. Joseph Bills, Yiu-Kai Ng |
IDEAS | 2 |
| 2020 | Research Paper Recommendation Based on Content Similarity, Peer Reviews, Authority, and PopularityabstractAccording to the Canadian Science Publishing, there are approximately 2.5 million scientific papers published each year. The huge volume of publications can be contributed to a substantial increase in the total number of academic journals, including the increasing number of predatory or fake scientific journals, which yield high volumes of poor-quality research work. The effect of this scenario is that there is an obsolete jungle of journals to flip through in searching for high-quality and relevant references for researchers, ranging from the ones who simply look for citations to cite or latest development and knowledge in a specific scientific area of study. In solving this problem, we propose a unique, elegant research paper recommender. Besides considering the topics and contents of related publications, our recommender also examines the peer reviews, authority, and popularity of each publication to ensure its quality. Conducted empirical study shows that our recommender outperforms existing research paper recommenders and contributes to the design of searching relevant publications. Yiu-Kai Ng |
ICTAI | 1 |
| 2020 | "Don't Judge a Book by its Cover": Exploring Book Traits Children FavorabstractWe present the preliminary exploration we conducted to identify traits that can influence children’s preferences in books. Findings offer insights for the design of recommender algorithms that would look beyond patterns inferred from traditional user-system interactions (e.g., ratings) for recommendation purposes, since when it comes to children such data is rarely, if at all, available. Ashlee Milton, Levesson Batista, Garrett Allen, Yiu-Kai Ng, Maria Soledad Pera |
RecSys | 5 |
| 2020 | Recommending Video Games to Adults with Autism Spectrum Disorder for Social-Skill EnhancementabstractAutism spectrum disorder (ASD) is a long-standing mental condition characterized by hindered mental growth and development and is a lifelong disability for the majority of affected individuals. In 2018, 2-3% of children in the USA have been diagnosed with autism. As these children move to adulthood, they have difficulty in developing a well-functioning motor skill. Some of these abnormalities, however, can be gradually improved if they are treated appropriately during their adulthood. Studies have shown that people with ASD enjoy playing video games. Educational games, however, have been primarily developed for children with ASD, which are too primitive for adults with ASD. We have developed a gaming and personalized recommender system that suggests therapeutic games to adults with ASD which can improve their social-interactive skills. The gaming system maintains the entertainment value of the games to ensure that adults are interested in playing them, whereas the recommender system suggests appropriate games for adults with ASD to play. The effectiveness and merit of our gaming and recommender system is backed up by an empirical study. Alisha Banskota, Yiu-Kai Ng |
UMAP | 2 |
| 2019 | Age-Suitability Prediction for Literature Using a Recurrent Neural Network ModelabstractDigital media holds a strong presence in society today. Providers of digital media may choose to have their content undergo an age-suitability analysis before being published. These analyses provide ratings that denote to which age group(s) a particular media item is appropriate. Content rating systems exist in many countries for television, music, video games, and mobile applications. These systems allow consumers to quickly determine whether or not a given media item is suitable to their age or preference. Literature, on the other hand, remains devoid of a comparable rating system. If a new, human-driven rating system for literature were to be implemented, it would be impeded by the fact that literary content is produced far more rapidly than are other forms of digital media; human working within such a system simply would not be able to read and analyze literature at its current rate of production. Thus, to provide fast, automated age-suitability ratings to works of literature (i.e., books), we propose a computer-driven rating system which predicts a book's content rating within each of eight categories: (a) crude humor/language; (b) drug, alcohol, and tobacco use; (c) kissing; (d) profanity; (e) nudity; (f) sex and intimacy; (g) violence and horror; and (h) gay/lesbian characters given the text of that book. Our computer-driven system circumvents the major hindrance to any theoretical human-driven rating system previously mentioned, i.e., infeasibility in time spent. Eric Brewer, Yiu-Kai Ng |
ICTAI | 2 |
| 2019 | Personalized Book Recommendation Based on a Deep Learning Model and Metadata
Yiu-Kai Ng, Urim Jung |
WISE | 1 |
| 2018 | Recommending social-interactive games for adults with autism spectrum disorders (ASD)abstractGames play a significant role in modern society, since they affect people of all ages and all walks of life, whether it be socially or mentally, and have direct impacts on adults with autism. Autism spectrum disorders (ASD) are a collection of neurodevelopmental disorders characterized by qualitative impairments in social relatedness and interaction, as well as difficulties in acquiring and using communication and language abilities. Adults with ASD often find it difficult to express and recognize emotions which makes it hard for them to interact with others socially. We have designed new interactive and collaborative games for autistic adults and developed a novel strategy to recommend games to them. Using modern computer vision and graphics techniques, we (i) track the player's speech rate, facial features, eye contact, audio communication, and emotional states, and (ii) foster their collaboration. These games are personalized and recommended to a user based on games interested to the user, besides the complexity of games at different levels according to the deficient level of the emotional understanding and social skills to which the user belongs. The objective of developing and recommending short-head (i.e., familiar) and long-tail (i.e., unfamiliar) games for adults with ASD is to enhance their social interacting skills with peers so that they can live a better life. Yiu-Kai Ng, Maria Soledad Pera |
RecSys | 1 |
| 2018 | Recommending books to be exchanged online in the absence of wish listsabstractAn online exchange system is a web service that allows communities to trade items without the burden of manually selecting them, which saves users' time and effort. Even though online book‐exchange systems have been developed, their services can further be improved by reducing the workload imposed on their users. To accomplish this task, we propose a recommendation‐based book exchange system, called EasyEx, which identifies potential exchanges for a user solely based on a list of items the user is willing to part with. EasyEx is a novel and unique book‐exchange system because unlike existing online exchange systems, it does not require a user to create and maintain a wish list, which is a list of items the user would like to receive as part of the exchange. Instead, EasyEx directly suggests items to users to increase serendipity and as a result expose them to items which may be unfamiliar, but appealing, to them. In identifying books to be exchanged, EasyEx employs known recommendation strategies, that is, personalized mean and matrix factorization, to predict book ratings, which are treated as the degrees of appeal to a user on recommended books. Furthermore, EasyEx incorporates OptaPlanner, which solves constraint satisfaction problems efficiently, as part of the recommendation‐based exchange process to create exchange cycles. Experimental results have verified that EasyEx offers users recommended books that satisfy the users' interests and contributes to the item‐exchange mechanism with a new design methodology. Maria Soledad Pera, Yiu-Kai Ng |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2017 | Personalized Table-Top Game RecommendationsabstractTable-top games have proven to enhance the lives of people of all ages. From children fostering the ability to focus to adults reducing their risk of developing Alzheimers disease, games continue to play a significant role in people's success in life. Recent prevalence in table-top games have increased production and sales for the game industry, thus providing a large variety of table-top game options to users. The volume of table-top options available to users, however, is problematic as buying new, unfamiliar games is a risk, since purchase cannot guarantee play satisfaction, and 100% refunds are not warrantied if game components have been tampered with. Existing websites such as Amazon and Barnes & Noble recommend table-top games to users, but their methodology is mainly based on consumer purchase patterns which neglect game characteristics in their non-personalized recommendations. These characteristics, which include topic relevance, complexity, and game category, can significantly affect the satisfaction level of game play when players experiment with new table-top games. In order to assist users in finding games of interest to play and enrich the player's gameplaying experience, we have developed PeGRec (Personalized Game Recommender), a novel software system that recommends the latest and most personally intriguing table-top games for users. We show in a series of evaluation tests that PeGRec's recommendations are more personalized and accurate in rankings according to user interests than the ones provided by Amazon's and Barnes & Nobles' game recommenders, respectively. Yiu-Kai Ng, Iris Seaman |
ICTAI | 1 |
| 2017 | Personalized Recipe Recommendations for Toddlers Based on Nutrient Intake and Food PreferencesabstractMany guidelines on healthy diet for toddlers have been established; however, they do not suggest recipes for daily meals. Even though recipe recommendation methodologies have been well-studied and implemented, they are not designed specifically with toddlers in mind. Instead, existing recipe recommendation systems focus on users' food preference and the nutrition of an individual dish without considering the total daily nutrient needs. In order to provide toddlers with a healthy diet, it is important for parents to know how much food toddlers should be eating and what foods satisfy their daily nutrient requirement, since feeding style and attitude of parents can influence toddlers' present and future dietary behavior and food preference. Given that toddler recipes recommended to parents with the required nutrition can reduce bias introduced by parents' own food preference, there is a need for a recipe recommendation system specifically designed for toddlers. In this paper, we propose a personalized toddler recipe recommendation system which incorporates the standard nutrition guideline published by the US government at ChooseMyPlate.gov with users' food preferences. Our recommender, which makes suggestions using initial data gathered from user profiles and MyPlate information, suggests relevant recipes for toddlers. Recommended recipes not only satisfy the daily nutrient intake from different food groups for toddlers, but also offer them a healthy eating habit. Yiu-Kai Ng, Meilan Jin |
MEDES | 1 |
| 2017 | Enhancing long tail item recommendations using tripartite graphs and Markov processabstractGiven that the Internet and sophisticated transportation networks have made an increasingly huge number of products and services available to the public, consumers are unable to identify, much less evaluate the usefulness of, such goods accessible to them. Modern recommendation systems filter out products of lesser utility to the customer, showcasing those items of higher preference to the user. While current state-of-the-art recommendation systems perform fairly well, they generally do better at recommending the popular subset of all products available rather than matching consumers with the vast amount of niche products in what has been termed the "Long Tail". In their seminal work, "Challenging the Long Tail Recommendation", Yin et al. make an eloquent argument that the long tail is where organizations can create the most value for their consumers. They also argue that existing recommender systems operate fundamentally different for long tail products than for mainstream goods. While matrix factorization, nearest-neighbors, and clustering work well for the "head" market, the long tail is better represented by a graph, specifically a bipartite graph that connects a set of users to a set of goods. In this paper, we discuss the algorithms presented by Yin et al., as well as a set of similar algorithms proposed by Shang et al., which traverse the bipartite graphs through a random walker in order to identify similar users and products. We build on elements from each work, as well as elements from a Markov process, to facilitate the random walker's traversal of tripartitle graphs into the long tail regions. This method specifically constructs paths into regions of the long tail that are favorable to users. Yiu-Kai Ng |
WI | 2 |
| 2017 | Using online data sources to make query suggestions for childrenabstractExisting popular web search engines have been widely used for retrieving information of interests by their users and offer query suggestions (QS) to assist them in exploring the wealth of information online. These search tools, however, are designed without any specific group of users in mind and thus are not tailored towards the specific needs of children, which can diminish their usability and design objectives when they are employed by children. Given the increasing use of the Web for educational and entertainment purposes by children, there is an urgent need to help them search the Web effectively. In this paper, we present a QS module, denoted CQS, which assists children in finding appropriate query keywords to capture their information needs by (i) analyzing content written for/by children, (ii) examining phrases and other metadata extracted from reputable (children’s) websites, and (iii) using a supervised learning approach to rank suggestions that are appealing to children. CQS offers suggestions with vocabulary that can be comprehended by children and with topics of interest to them. We conducted a number of empirical studies using keyword queries initiated by children, besides gathering feedback on the usefulness of CQS-generated suggestions through crowdsourcing. The performance evaluation of CQS revealed the effectiveness of the methodology of CQS. In addition, it demonstrated that CQS-generated suggestions were preferred over suggestions provided by Bing and Yahoo! and at least as comparable to queries suggested by Google. Maria Soledad Pera, Yiu-Kai Ng |
Web Intell. | 2 |
| 2016 | Recommending Books for Children Based on the Collaborative and Content-Based Filtering Approaches
Yiu-Kai Ng |
ICCSA (4) | 1 |
| 2016 | Making personalized movie recommendations for childrenabstractMultimedia have significant impact on the social and psychological development of children who are often explored to inappropriate materials, including movies that are either accessible online or through other multimedia channels. Even though not all movies are bad, there are negative effects of offensive languages, violence, and sexuality as exhibited in movies. Parents and guidance of children need all the help they can get to promote the healthy use of movies these days. To offer them appropriate movies of interest to their youths, we have developed MovReC, a personalized movie recommender for children, which is designed to provide educational and suitable entertaining opportunities for children. Unlike Amazon and other online movie recommendation systems, such as Common Sense Media, IMDb, and TasteKid, MovReC is unique, since to the best of our knowledge MovReC is the first personalized children movie recommender. Moreover, MovReC determines the appealingness of a movie for a particular user based on its children-appropriate score computed by using the Backpropagation model, pre-defined category using LDA, its predicted rating using Matrix Factorization, and sentiments based on its users' reviews, which along with its like/dislike count and genres, yield the features considered by MovReC. MovReC combines these features by using the CombMNZ model to rank and recommend movies. The performance evaluation of MovReC clearly demonstrates its effectiveness and its recommended movies are highly regarded by its users. Eunice Tan, Iris Seaman, Humphrey Leung, Yiu-Kai Ng |
iiWAS | 4 |
| 2016 | Orthogonal query recommendations for childrenabstractChildren have become an ever-popular group of users who explore and search the Web. However, with this rise in popularity, there is still a gap in the amount of research that has been done to help children receive query recommendations applicable to their age group and understanding. We propose to develop an orthogonal query recommendation system and apply it specifically to children's queries. By using this recommender, we provide queries that are semantically different but conceptually very similar to a child's initial query. We have also validated our query recommendation system through Mean Reciprocal Ranking and Mean Average Precision to determine the accuracy of the recommended queries. Through this process, we have shown that the searching and browsing experience for children have improved and help them more easily find results that match their information needs. Alicia Wood, Yiu-Kai Ng |
iiWAS | 2 |
| 2016 | A readability level prediction tool for K-12 booksabstractThe readability levels of books identify suitable reading materials. Unfortunately, the majority of published books are assigned a readability level range, which is not useful to readers who look for books at a particular grade level. Existing readability formulas/analysis tools require at least an excerpt of a book to estimate its readability level, which is a severe constraint, since copyright laws prohibit book contents from being made publicly accessible. To alleviate the constraint, we have developed TRoLL which relies on publicly accessible online book metadata, in addition to using a book's snippet, if it is available, to predict its readability level. Based on a multi‐dimensional regression analysis, TRoLL determines the grade level of any book instantly, even without a sample of its text, and considers its topical suitability, which is unique. Furthermore, TRoLL is a significant contribution to the educational community, since its computed book readability levels can enrich K‐12 readers' book selections and aid parents, teachers, and librarians in locating reading materials suitable for their K‐12 readers, which can be a time‐consuming and frustrating task that does not always yield a quality outcome. Conducted empirical studies have verified the prediction accuracy of TRoLL and demonstrated its superiority over well‐known readability formulas/analysis tools. Joel Denning, Maria Soledad Pera, Yiu-Kai Ng |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2016 | Enhancing web search by using query-based clusters and multi-document summaries
Rani Qumsiyeh, Yiu-Kai Ng |
Knowl. Inf. Syst. | 2 |
| 2015 | Clustering Retrieved Web Documents to Speed Up Web Searches
Rani Qumsiyeh, Yiu-Kai Ng |
ICCSA (1) | 2 |
| 2014 | Automating readers' advisory to make book recommendations for K-12 readersabstractThe academic performance of students is affected by their reading ability, which explains why reading is one of the most important aspects of school curriculums. Promoting good reading habits among K-12 students is essential, given the enormous influence of reading on students' development as learners and members of society. In doing so, it is indispensable to provide readers with engaging and motivating reading selections. Unfortunately, existing book recommenders have failed to offer adequate choices for K-12 readers, since they either ignore the reading abilities of their users or cannot acquire the much-needed information to make recommendations due to privacy issues. To address these problems, we have developed Rabbit, a book recommender that emulates the readers' advisory service offered at school/public libraries. Rabbit considers the readability levels of its readers and determines the facets, i.e.,appeal factors, of books that evoke subconscious, emotional reactions on a reader. The design of Rabbit is unique, since it adopts a multi-dimensional approach to capture the reading abilities, preferences, and interests of its readers, which goes beyond the traditional book content/topical analysis. Conducted empirical studies have shown that Rabbit outperforms a number of (readability-based) book recommenders. Maria Soledad Pera, Yiu-Kai Ng |
RecSys | 2 |
| 2014 | Exploiting the wisdom of social connections to make personalized recommendations on scholarly articles
Maria Soledad Pera, Yiu-Kai Ng |
J. Intell. Inf. Syst. | 2 |
| 2014 | Assisting web search using query suggestion based on word similarity measure and query modification patterns
Rani Qumsiyeh, Yiu-Kai Ng |
World Wide Web | 2 |
| 2013 | A Probabilistic Query Suggestion Approach without Using Query LogsabstractCommercial web search engines include a query suggestion module so that given a user's keyword query, alternative suggestions are offered and served as a guide to assist the user in formulating queries which capture his/her intended information need in a quick and simple manner. Majorityof these modules, however, perform an in-depth analysis oflarge query logs and thus (i) their suggestions are mostlybased on queries frequently posted by users and (ii) theirdesign methodologies cannot be applied to make suggestions oncustomized search applications for enterprises for which theirrespective query logs are not large enough or non-existent. To address these design issues, we have developed PQS, aprobabilistic query suggestion module. Unlike its counterparts, PQS is not constrained by the existence of query logs, sinceit solely relies on the availability of user-generated contentfreely accessible online, such as the Wikipedia.org documentcollection, and applies simple, yet effective, probabilistic-andinformation retrieval-based models, i.e., the Multinomial, BigramLanguage, and Vector Space Models, to provide usefuland diverse query suggestions. Empirical studies conductedusing a set of test queries and the feedbacks provided byMechanical Turk appraisers have verified that PQS makesmore useful suggestions than Yahoo! and is almost as goodas Google and Bing based on the relatively small difference inperformance measures achieved by Google and Bing over PQS. Meher T. Shaikh, Maria Soledad Pera, Yiu-Kai Ng |
ICTAI | 3 |
| 2013 | What to read next?: making personalized book recommendations for K-12 usersabstractFinding books that children/teenagers are interested in these days is a non-trivial task due to the diversity of topics covered in huge volumes of books with varied readability levels. Even though K-12 readers can turn to book recommenders to look for books, the recommended books may not satisfy their personal needs, since they could be beyond/below their readability levels or fail to match their topics of interest. To address these problems, we introduce BReK12, a book recommender that makes personalized suggestions tailored to each K-12 user U based on books available on a social book-marking site that (i) are similar in content to the ones that are known to be of interest to U, (ii) have been bookmarked by users with reading patterns similar to U's, and (iii) can be comprehended by U. BReK12 is an asset to its users, since it suggests books that are appealing to its users and at grade levels that they can cope with, which can increase their reading selection choices and motivate them to read. We have also developed ReLAT, the readability analysis tool employed by BReK12 to determine the grade level of books. ReLAT is novel, compared with existing readability formulas, since it can predict the grade level of a book even if an excerpt of the book is not available. We have conducted empirical studies which have verified the accuracy of ReLAT in predicting the grade level of a book and the effectiveness of BReK12 over existing baseline recommendation systems. Maria Soledad Pera, Yiu-Kai Ng |
RecSys | 2 |
| 2013 | Enhancing Web Search Using Query-Based Clusters and LabelsabstractCurrent web search engines, such as Google, Bing, and Yahoo!, rank the set of documents S retrieved in response to a user query and display the URL of each document D in S with a title and a snippet, which serves as an abstract of D. Snippets, however, are not as useful as they are designed for, which is supposed to assist its users to quickly identify results of interest, if they exist. These snippets fail to (i) provide distinct information and (ii) capture the main contents of the corresponding documents. Moreover, when the intended information need specified in a search query is ambiguous, it is very difficult, if not impossible, for a search engine to identify precisely the set of documents that satisfy the user's intended request without requiring additional inputs. Furthermore, a document title is not always a good indicator of the content of the corresponding document. All of these design problems can be solved by our proposed query-based cluster and labeler, called QCL. QCL generates concise clusters of documents covering various subject areas retrieved in response to a user query, which saves the user's time and effort in searching for specific information of interest without having to browse through the documents one by one. Experimental results show that QCL is effective and efficient in generating high-quality clusters of documents on specific topics with informative labels. Rani Qumsiyeh, Yiu-Kai Ng |
Web Intelligence | 2 |
| 2013 | A group recommender for movies based on content similarity and popularity
Maria Soledad Pera, Yiu-Kai Ng |
Inf. Process. Manag. | 2 |
| 2013 | Web-based closed-domain data extraction on online advertisements
Maria Soledad Pera, Rani Qumsiyeh, Yiu-Kai Ng |
Inf. Syst. | 3 |
| 2012 | BReK12: a book recommender for K-12 usersabstractIdeally, students in K-12 grade levels can turn to book recommenders to locate books that match their interests. Existing book recommenders, however, fail to take into account the readability levels of their users, and hence their recommendations may be unsuitable for the users. To address this issue, we introduce BReK12, a recommender that targets K-12 users and prioritizes the reading level of its users in suggesting books of interest. Empirical studies conducted using the Bookcrossing dataset show that BReK12 outperforms a number of existing recommenders (developed for general users) in identifying books appealing to K-12 users. Maria Soledad Pera, Yiu-Kai Ng |
SIGIR | 2 |
| 2012 | Predicting the ratings of multimedia items for making personalized recommendationsabstractExisting multimedia recommenders suggest a specific type of multimedia items rather than items of different types personalized for a user based on his/her preference. Assume that a user is interested in a particular family movie, it is appealing if a multimedia recommendation system can suggest other movies, music, books, and paintings closely related to the movie. We propose a comprehensive, personalized multimedia recommendation system, denoted MudRecS, which makes recommendations on movies, music, books, and paintings similar in content to other movies, music, books, and/or paintings that a MudRecS user is interested in. MudRecS does not rely on users' access patterns/histories, connection information extracted from social networking sites, collaborated filtering methods, or user personal attributes (such as gender and age) to perform the recommendation task. It simply considers the users' ratings, genres, role players (authors or artists), and reviews of different multimedia items, which are abundant and easy to find on the Web. MudRecS predicts the ratings of multimedia items that match the interests of a user to make recommendations. The performance ofMudRecS has been compared with current state-of-the-art multimedia recommenders using various multimedia datasets, and the experimental results show that MudRecS significantly outperforms other systems in accurately predicting the ratings of multimedia items to be recommended. Rani Qumsiyeh, Yiu-Kai Ng |
SIGIR | 2 |
| 2012 | Using maximal spanning trees and word similarity to generate hierarchical clusters of non-redundant RSS news articles
Maria Soledad Pera, Yiu-Kai Ng |
J. Intell. Inf. Syst. | 2 |
| 2011 | A personalized recommendation system on scholarly publicationsabstractResearchers, as well as ordinary users who seek information in diverse academic fields, turn to the web to search for publications of interest. Even though scholarly publication recommenders have been developed to facilitate the task of discovering literature pertinent to their users, they (i) are not personalized enough to meet users' expectations, since they provide the same suggestions to users sharing similar profiles/preferences, (ii) generate recommendations pertaining to each user's general interests as opposed to the specific need of the user, and (iii) fail to take full advantages of valuable user-generated data at social websites that can enhance their performance. To address these problems, we propose PubRec, a recommender that suggests closely-related references to a particular publication P tailored to a specific user U, which minimizes the time and efforts imposed on U in browsing through general recommended publications. Empirical studies conducted using data extracted from CiteULike (i) verify the efficiency of the recommendation and ranking strategies adopted by PubRec and (ii) show that PubRec significantly outperforms other baseline recommenders. Maria Soledad Pera, Yiu-Kai Ng |
CIKM | 2 |
| 2011 | A query-based multi-document sentiment summarizerabstractReview websites, such as Epinions.com, which offer users a platform to share their opinions on diverse products and services, provide a valuable source of opinion-rich information. Browsing through archived reviews to locate different opinions on a product or service, however, is a time-consuming and tedious task, and in most cases, the large amount of available information is difficult for users to absorb. To facilitate the process of synthesizing opinions expressed in reviews on a product or service P specified in a user query/question Q, we introduce QMSS, a query-based multi-document sentiment summarizer. QMSS creates a summary for Q, which either reflects the general opinions on P or is tailored to specific facets (i.e., features) and/or sentiment of P as specified in Q. QMSS (i) identifies the facets addressed in reviews retrieved for Q, (ii) employs a sentence-based, sentiment classifier to determine the polarity of each sentence in each review, and (iii) clusters sentences in reviews according to the facets captured in the sentences, which are identified using a keyword-label extraction algorithm. This process dictates which sentences in the reviews should be included in the summary for Q. Empirical studies have verified that QMSS is highly effective in generating summaries that satisfy users' information needs and ranks on top among the state-of-the-art query-based multi-document sentiment summarizers Maria Soledad Pera, Rani Qumsiyeh, Yiu-Kai Ng |
CIKM | 3 |
| 2011 | ReadAid: A Robust and Fully-Automated Readability Assessment ToolabstractReading is an integral part of educational development, however, it is frustrating for people who struggle to understand (are not motivated to read, respectively) text documents that are beyond (below, respectively) their readability levels. Finding appropriate reading materials, with or without first scanning through their contents, is a challenge, since there are tremendous amount of documents these days and a clear majority of them are not tagged with their readability levels. Even though existing readability assessment tools determine readability levels of text documents, they analyze solely the lexical, syntactic, and/or semantic properties of a document, which are neither fully-automated, generalized, nor well-defined and are mostly based on observations. To advance the current readability analysis technique, we propose a robust, fully-automated readability analyzer, denoted ReadAid, which employs support vector machines to combine features from the US Curriculum and College Board, traditional readability measures, and the author(s) and subject area(s) of a text document d to assess the readability level of d. ReadAid can be applied for (i) filtering documents (retrieved in response to a web query) of a particular readability level, (ii) determining the readability levels of digitalized text documents, such as book chapters, magazine articles, and news stories, or (iii) dynamically analyzing, in real time, the grade level of a text document being created. The novelty of ReadAid lies on using authorship, subject areas, and academic concepts and grammatical constructions extracted from the US Curriculum to determine the readability level of a text document. Experimental results show that ReadAid is highly effective and outperforms existing state-of-the-art readability assessment tools. Rani Qumsiyeh, Yiu-Kai Ng |
ICTAI | 2 |
| 2011 | With a Little Help from My Friends: Generating Personalized Book Recommendations Using Data Extracted from a Social WebsiteabstractWith the large amount of books available nowadays, users are overwhelmed with choices when they attempt to find books of interest. While existing book recommendation systems, which are based on either collaborative filtering, content-based, or hybrid methods, suggest books (among the millions available) that might be appealing to the users, their recommendations are not personalized enough to meet users' expectations due to their collective assumption on group preference and/or exact content matching, which is a failure. To address this problem, we have developed PReF, a Personalized Recommender that relies on Friendships established by user son a social website, such as Library Thing, to make book recommendations tailored to individual users. In selecting books to be recommended to a user U, who is interested in a book B, PReF (i) considers books belonged to U's friends, (ii) applies word-correlation factors to disclose books similar in contents to B, (iii) depends on the ratings given to books by U's friends to identify highly-regarded books, and (iv) determine show reliable individual friends of U are in providing books from their own catalogs (that are similar in content to B)to be recommended. We have conducted an empirical study and verified that (i) relying on data extracted from social websites improves the effectiveness of book recommenders and (ii) PReF outperforms the recommenders employed by Amazon and Library Thing. Maria Soledad Pera, Yiu-Kai Ng |
Web Intelligence | 2 |
| 2011 | Generating Exact- and Ranked Partially-Matched Answers to Questions in AdvertisementsabstractTaking advantage of the Web, many advertisements (ads for short) websites, which aspire to increase client's transactions and thus profits, offer searching tools which allow users to (i) post keyword queries to capture their information needs or (ii) invoke form-based interfaces to create queries by selecting search options, such as a price range, filled-in entries, check boxes, or drop-down menus. These search mechanisms, however, are inadequate, since they cannot be used to specify a natural-language query with rich syntactic and semantic content, which can only be handled by a question answering (QA) system. Furthermore, existing ads websites are incapable of evaluating arbitrary Boolean queries or retrieving partially-matched answers that might be of interest to the user whenever a user's search yields only a few or no results at all. In solving these problems, we present a QA system for ads, called CQAds, which (i) allows users to post a natural-language questionQfor retrieving relevant ads, if they exist, (ii) identifies ads as answers that partially-match the requested information expressed inQ, if insufficient or no answers toQcan be retrieved, which are ordered using asimilarity-rankingapproach, and (iii) analyzes incomplete or ambiguous questions to perform the "best guess" in retrieving answers that "best match" the selection criteria specified inQ. CQAds is also equipped with a Boolean model to evaluate Boolean operators that are eitherexplicitlyorimplicitlyspecified inQ, i.e., with or without Boolean operators specified by the users, respectively. CQAds is easy to use, scalable to all ads domains, and more powerful than search tools provided by existing ads websites, since its query-processing strategy retrieves relevant ads of higher quality and quantity. We have verified the accuracy of CQAds in retrieving ads on eight ads domains and compared its ranking strategy with other well-known ranking approaches. Rani Qumsiyeh, Maria Soledad Pera, Yiu-Kai Ng |
Proc. VLDB Endow. | 3 |
| 2011 | SimPaD: A word-similarity sentence-based plagiarism detection tool on Web documentsabstractPlagiarism is a serious problem that infringes copyrighted documents/materials, which is an unethical practice and decreases the economic incentive received by their legal owners. Unfortunately, plagiarism is getting worse due to the increasing numbe Maria Soledad Pera, Yiu-Kai Ng |
Web Intell. Agent Syst. | 2 |
| 2010 | An Unsupervised Sentiment Classifier on Summarized or Full Reviews
Maria Soledad Pera, Rani Qumsiyeh, Yiu-Kai Ng |
WISE | 3 |
| 2009 | Classifying Sentence-Based Summaries of Web DocumentsabstractText classification categories Web documents in large collections into predefined classes based on their contents. Unfortunately, the classification process can be time-consuming and users are still required to spend considerable amount of time scanning through the classified Web documents to identify the ones that satisfy their information needs. In solving this problem, we first introduce CorSum, an extractive single-document summarization approach, which is simple and effective in performing the summarization task, since it only relies on word similarity to generate high-quality summaries. Hereafter, we train a Naive Bayes classifier on CorSum-generated summaries and verify the classification accuracy using the summaries and the speed-up during the process. Experimental results on the DUC-2002 and 20 Newsgroups datasets show that CorSum outperforms other extractive summarization methods, and classification time is significantly reduced using CorSum-generated summaries with compatible accuracy. More importantly, browsing summaries, instead of entire documents, classified to topic-oriented categories facilitates the information searching process on the Web. Maria Soledad Pera, Yiu-Kai Ng |
ICTAI | 2 |
| 2009 | A sophisticated library search strategy using folksonomies and similarity matchingabstractAbstract Libraries, private and public, offer valuable resources to library patrons. As of today, the only way to locate information archived exclusively in libraries is through their catalogs. Library patrons, however, often find it difficult to formulate a proper query, which requires using specific keywords assigned to different fields of desired library catalog records, to obtain relevant results. These improperly formulated queries often yield irrelevant results or no results at all. This negative experience in dealing with existing library systems turns library patrons away from directly querying library catalogs; instead, they rely on Web search engines to perform their searches first, and upon obtaining the initial information (e.g., titles, subject headings, or authors) on the desired library materials, they query library catalogs. This searching strategy is an evidence of failure of today's library systems. In solving this problem, we propose an enhanced library system, which allows partial, similarity matching of (a) tags defined by ordinary users at a folksonomy site that describe the content of books and (b) unrestricted keywords specified by an ordinary library patron in a query to search for relevant library catalog records. The proposed library system allows patrons posting a query Q using commonly used words and ranks the retrieved results according to their degrees of resemblance with Q while maintaining the query processing time comparable with that achieved by current library search engines. Maria Soledad Pera, William B. Lund, Yiu-Kai Ng |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2009 | SpamED: A spam E-mail detection approach based on phrase similarityabstractAbstract E‐mail messages are unquestionably one of the most popular communication media these days. Not only are they fast and reliable but also free in general. Unfortunately, a significant number of e‐mail messages received by e‐mail users on a daily basis are spam. This fact is annoying since spam messages translate into a waste of the user's time in reviewing and deleting them. In addition, spam messages consume resources such as storage, bandwidth, and computer‐processing time. Many attempts have been made in the past to eradicate spam; however, none has proven highly effective. In this article, we propose a spam e‐mail detection approach, called SpamED, which uses the similarity of phrases in messages to detect spam. Conducted experiments not only verify that SpamED using trigrams in e‐mail messages is capable of minimizing false positives and false negatives in spam detection but it also outperforms a number of existing e‐mail filtering approaches with a 96% accuracy rate. Maria Soledad Pera, Yiu-Kai Ng |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2009 | A dynamic attribute-based data filtering and recovery scheme for web information processing
Amit Ahuja, Yiu-Kai Ng |
Knowl. Inf. Syst. | 2 |
| 2008 | Generating Fuzzy Equivalence Classes on RSS News Articles for Retrieving Correlated Information
Nathaniel Gustafson, Maria Soledad Pera, Yiu-Kai Ng |
ICCSA (2) | 3 |
| 2008 | Identifying Spam Web Pages Based on Content Similarity
Maria Soledad Pera, Yiu-Kai Ng |
ICCSA (2) | 2 |
| 2008 | Augmenting Data Retrieval with Information Retrieval Techniques by Using Word Similarity
Nathaniel Gustafson, Yiu-Kai Ng |
NLDB | 2 |
| 2008 | Nowhere to Hide: Finding Plagiarized Documents Based on Sentence SimilarityabstractPlagiarism is a serious problem that infringes copyrighted documents/materials, which is an unethical practice and decreases the economic incentive received by authors (owners) of the original copies. Unfortunately, plagiarism is getting worse due to the increasing number of on-line publications on the Web, which facilitates locating and paraphrasing information. In solving this problem, we propose a novel plagiarism-detection method, called SimPaD, which (i) establishes the degree of resemblance between any two documents D1and D2based on their sentence-to-sentence similarity computed by using pre-defined word-correlation factors, and (ii) generates agraphical view of sentences that are similar (or the same) in D1and D2. Experimental results verify that SimPaD is highly accurate in detecting (non-) plagiarized documents and outperforms existing plagiarism-detection approaches. Nathaniel Gustafson, Maria Soledad Pera, Yiu-Kai Ng |
Web Intelligence | 3 |
| 2008 | Answering form-based web queries using the data-mining approach
Xiaochun Yang 0001, Yiu-Kai Ng |
J. Intell. Inf. Syst. | 2 |
| 2007 | Using word similarity to eradicate junk emailsabstractEmails are one of the most commonly used modern communication media these days; however, unsolicited emails obstruct this otherwise fast and convenient technology for information exchange and jeopardize the continuity of this popular communication tool. Waste of valuable resources and time and exposure to offensive content are only a few of the problems that arise as a result of junk emails. In addition, the monetary cost of processing junk emails reaches billions of dollars per year and is absorbed by public users and Internet service providers. Even though there has been extensive work in the past dedicated to eradicate junk emails, none of the existing junk email detection approaches has been highly successful in solving these problems, since spammers have been able to infiltrate existing detection techniques. In this paper, we present a new tool, JunEX, which relies on the content similarity of emails to eradicate junk emails. JunEX compares each incoming email to a core of emails marked as junk by each individual user to identify unwanted emails while reducing the number of legitimate emails treated as junk, which is critical. Conducted experiments on JunEX verify its high accuracy. Maria Soledad Pera, Yiu-Kai Ng |
CIKM | 2 |
| 2007 | Finding Similar RSS News Articles Using Correlation-Based Phrase Matching
Maria Soledad Pera, Yiu-Kai Ng |
KSEM | 2 |
| 2006 | Associated Load Shedding Strategies for Computing Multi-joins in Sensor Networks
Xiaochun Yang 0001, Yiu-Kai Ng, Bin Wang 0015, Ge Yu 0001 |
DASFAA | 3 |
| 2006 | Eliminating Redundant and Less-Informative RSS News Articles Based on Word Similarity and a Fuzzy Equivalence RelationabstractThe Internet has marked this era as the information age. There is no precedent in the amazing amount of information, especially network news, that can be accessed by Internet users these days. As a result, the problem of seeking information in online news articles is not the lack of them but being overwhelmed by them. This brings huge challenges in processing online news feeds, e.g., how to determine which news article is important, how to determine the quality of each news article, and how to filter irrelevant and redundant information. In this paper, we propose a method for filtering redundant and less-informative RSS news articles that solves the problem of excessive number of news feeds observed in RSS news aggregators. Our filtering approach measures similarity among RSS news entries by using the fuzzy-set information retrieval model and a fuzzy equivalent relation for computing word/sentence similarity to detect redundant and less-informative news articles Ian Garcia, Yiu-Kai Ng |
ICTAI | 2 |
| 2006 | Using Word Clusters to Detect Similar Web Documents
Jonathan Koberstein, Yiu-Kai Ng |
KSEM | 2 |
| 2005 | Categorizing and Extracting Information from Multilingual HTML DocumentsabstractThe amount of online information written in different natural languages and the number of non-English speaking Internet users have been increasing tremendously during the past decade. In order to provide high-performance access of multilingual information on the Internet, we have developed a data analysis and querying system (DatAQs) that: (i) analyzes, identifies, and categorizes languages used in HTML documents; (ii) extracts information from HTML documents of interest written in different languages; (iii) allows the user to submit queries for retrieving extracted information in the same natural language provided by the query engine of DatAQs using a menu-driven user interface; and (iv) processes the user's queries (as Boolean expressions) to generate the results. DatAQs extracts information from HTML documents that belong to various data-rich, narrow-in-breadth application domains, such as car ads, house rentals, job ads, stocks, university catalogs, etc. The average F-measure on identifying HTML documents written in a particular natural language correctly is 89%, whereas the F-measure on categorizing HTML documents belonged to the car-ads application domain is 94%. SeungJin Lim, Yiu-Kai Ng |
IDEAS | 2 |
| 2004 | Selective-Splitting and Cache-Maintenance Algorithms for Associative-Client Caches
Jiaxin J. Gao, Dallan Quass, Yiu-Kai Ng |
Distributed Parallel Databases | 3 |
| 2003 | A Protein Secondary Structure Prediction Framework Based on the Support Vector Machine
Xiaochun Yang 0001, Bin Wang 0015, Yiu-Kai Ng, Ge Yu 0001, Guoren Wang |
WAIM | 3 |
| 2003 | An Ontology-Based Binary-Categorization Approach for Recognizing Multiple-Record Web Documents Using a Probabilistic Retrieval Model
Yiu-Kai Ng |
Inf. Retr. | 2 |
| 2003 | Performing Binary-Categorization on Multiple-Record Web Documents Using Information Retrieval Models and Application Ontologies
Linus W. Kwong, Yiu-Kai Ng |
World Wide Web | 2 |
| 2002 | Integrating HTML Tables Using Semantic Hierarchies And Meta-Data SetsabstractAs the Internet is a global network, there is a demand on accessing closely related data without browsing through different Web documents. A significant amount of these data are presented in HTML documents. Since data contents of HTML documents are intervened by markups, it is not trivial to integrate and provide a unified view of closely related data in different HTML documents. In this paper we present an approach for integrating semantically related data in any HTML tables that belong to a particular domain of interest (ID), such as house/apartment rental, by using the semantic hierarchies generated from the tables and the predefined meta-data sets that indicate related column names in ID. In our approach, we capture each data source as semi-structured data, called semantic hierarchy, and the end result of integrating different HTML tables of ID is a unified view of data in the tables, which is presented in an XML document. Besides HTML tables, our approach can be adopted by any system that integrates semi-structured data across different platforms. SeungJin Lim, Yiu-Kai Ng, Xiaochun Yang 0001 |
IDEAS | 2 |
| 2002 | A Client-Based Web-Cache Management System
Kelvin Lau, Yiu-Kai Ng |
WAIM | 2 |
| 2001 | A Binary-Categorization Approach for Classifying Multiple-Record Web Documents Using Application Ontologies and a Probabilistic ModelabstractThe amount of information available on the World Wide Web has been increasing dramatically in recent years. To enhance speedy searching and retrieving Web documents of interest, researchers and practitioners have partially relied on various information retrieval techniques. We propose a probabilistic model to classify Web documents into relevant documents and irrelevant documents with respect to a particular application ontology, which is a conceptual-model snippet of standard ontologies. Our probabilistic model is based on multivariate statistical analysis and is different from the conventional probabilistic information retrieval models. The experiments we have conducted on a set of representative Web documents indicate that the proposed probabilistic model is promising in binary-categorization of multiple-record Web documents. Yiu-Kai Ng, June Tang, Michael A. Goodrich |
DASFAA | 1 |
| 2001 | Recognizing Ontology-Applicable Multiple-Record Web Documents
David W. Embley, Yiu-Kai Ng |
ER | 2 |
| 2001 | An Automated Change Detection Algorithm for HTML Documents Based on Semantic HierarchiesabstractThe data at many Web sites is changing rapidly, and a significant amount of this data is presented in HTML documents that consist of markups and data contents. Although XML is becoming more popular for data exchange, the presentation of data contained in XML documents is given, by and large, in the HTML format using XSL(T). Since HTML was designed to "display" data from the human perspective, it is not trivial for a machine to detect (hierarchical) changes of data in an HTML document. In this paper, we propose a heuristic algorithm, called SCD (Semantic Change Detection), to detect semantic changes to the hierarchical data contents in any two HTML documents automatically. Semantic changes differ from syntactic changes since the latter refer to changes of data contents with respect to markup structures according to the HTML grammar. SCD does not require pre-processing, nor any knowledge of the internal structure of the source documents beforehand. The time complexity of SCD is O[(|X|/spl times/|Y|)log(|X|/spl times/|Y|)], where |X| and |Y| are the number of unique branches in the syntactic hierarchies of any two given HTML documents, respectively. SeungJin Lim, Yiu-Kai Ng |
ICDE | 2 |
| 2001 | A Hybrid Fragmentation Approach for Distributed Deductive Database Systems
SeungJin Lim, Yiu-Kai Ng |
Knowl. Inf. Syst. | 2 |
| 1999 | An Automated Approach for Retrieving Hierarchical Data from HTML TablesabstractAmong the HTML elements, HTML tables [RHJ98] encapsulate hierarchically structured data (hierarchical data in short) in a tabular structure. HTML tables do not come with a rigid schema and almost any forms of two-dimensional tables are acceptable according to the HTML grammar. This relaxation complicates the process of retrieving hierarchical data from HTML tables. In this paper, we propose an automated approach for retrieving hierarchical data from HTML tables. The proposed approach constructs the content tree of an HTML table, which captures the intended hierarchy of the data content of the table, without requiring the internal structure of the table to be known beforehand. Also, the user of the content tree does not deal with HTML tags while retrieving the desired data from the content tree. Our approach can be employed by (i) a query language written for retrieving hierarchically structured data, extracted from either the contents of HTML tables or other sources, (ii) a processor for converting HTML tables to XML documents, and (iii) a data warehousing repository for collecting hierarchical data from HTML tables and storing materialized views of the tables. The time complexity of the proposed retrieval approach is proportional to the number of HTML elements in an HTML table. SeungJin Lim, Yiu-Kai Ng |
CIKM | 2 |
| 1999 | WebView: A Tool for Retrieving Internal Structures and Extracting Information from HTML DocumentsabstractHTML is a well-accepted and widely used language for creating platform-independent documents to be posted on the Web, and HTML documents are semistructured in nature according to the HTML specification. We propose a tool, called WebView, which constructs the semistructured data graph (SDG) of an HTML document H to capture the internal structure of data embedded in H and in its (in)directly linked documents. On top of the SDG, WebView provides query processing capability for evaluating SQL-like queries that are posted against the SDG, i.e., the source document(s), for extracting information from the SDG. Existing methods for extracting structured information from certain HTML documents with static internal structure, such as wrappers and integrators for data warehousing, can benefit from WebView. SeungJin Lim, Yiu-Kai Ng |
DASFAA | 2 |
| 1999 | Record-Boundary Discovery in Web DocumentsabstractExtraction of information from unstructured or semistructured Web documents often requires a recognition and delimitation of records. (By “record” we mean a group of information relevant to some entity.) Without first chunking documents that contain multiple records according to record boundaries, extraction of record information will not likely succeed. In this paper we describe a heuristic approach to discovering record boundaries in Web documents. In our approach, we capture the structure of a document as a tree of nested HTML tags, locate the subtree containing the records of interest, identify candidate separator tags within the subtree using five independent heuristics, and select a consensus separator tag based on a combined heuristic. Our approach is fast (runs linearly for practical cases within the context of the larger data-extraction problem) and accurate (100% in the experiments we conducted). David W. Embley, Y. S. Jiang, Yiu-Kai Ng |
SIGMOD Conference | 3 |
| 1999 | Conceptual-Model-Based Data Extraction from Multiple-Record Web Pages
David W. Embley, Douglas M. Campbell, Y. S. Jiang, Stephen W. Liddle, Yiu-Kai Ng, Dallan Quass, Randy D. Smith |
Data Knowl. Eng. | 5 |
| 1998 | A Conceptual-Modeling Approach to Extracting Data from the Web
David W. Embley, Douglas M. Campbell, Y. S. Jiang, Stephen W. Liddle, Yiu-Kai Ng, Dallan Quass, Randy D. Smith |
ER | 5 |
| 1998 | A Model-Forest Based Horizontal Fragmentation Approach for Disjunctive Deductive DatabasesabstractDisjunctive deductive databases (DDDBs) can capture indefinite information, i.e., imprecise or partial knowledge, of the real world. In this paper we present a method for horizontally fragmenting a DDDB based on the minimal-model forest approach. A minimal-model forest of a DDDB D is a collection of minimal-model trees of D such that each tree represents a set of facts that is disjoint from the set of facts represented in any other tree of D (All these facts are given in D.). Eventually, each tree T in the forest is assigned to a fragment along with the rules that utilize the facts represented in T to infer new facts. This approach minimizes the amount of data in D that needs to be processed for any query of D by taking the advantage of the natural partition of data that may appear in D. Aparna Seetharaman, Yiu-Kai Ng |
IDEAS | 2 |
| 1997 | Design and Analysis of Parallel Set-Term Unification
Seung-Jin Ling, Yiu-Kai Ng |
COCOON | 2 |
| 1997 | A Minimal-Model Based Horizontal Fragmentation Algorithm for Disjunctive Deductive DatabasesabstractDisjunctive deductive databases (DDDBs) capture indefinite information, i.e., imprecise or partial knowledge of the real world, and are more general than definite deductive databases that can only represent unconditionally true facts. Formal approaches for fragmenting a DDDB, with a view to distribute the DDDB and then design a query optimization strategy for the distributed DDDB, are lacking. We present a formal approach for fragmenting a DDDB based on the minimal-model semantics. Fragments generated by the proposed algorithm facilitate query evaluation against the DDDB in a distributed system. Aparna Seetharaman, Yiu-Kai Ng |
IDEAS | 2 |
| 1997 | Vertical Fragmentation and Allocation in Distributed Deductive Database Systems
SeungJin Lim, Yiu-Kai Ng |
Inf. Syst. | 2 |
| 1996 | A Formal Approach for Horizontal Fragmentation in Distributed Deductive Database Design
SeungJin Lim, Yiu-Kai Ng |
DEXA | 2 |
| 1996 | A Normal Form for Precisely Characterizing Redundancy in Nested RelationsabstractWe give a straightforward definition for redundancy in individual nested relations and define a new normal form that precisely characterizes redundancy for nested relations. We base our definition of redundancy on an arbitrary set of functional and multivalued dependencies, and show that our definition of nested normal form generalizes standard relational normalization theory. In addition, we give a condition that can prevent an unwanted structural anomaly in nested relations, namely, embedded nested relations with at most one tuple. Like other normal forms, our nested normal form can serve as a guide for database design. Wai Yin Mok, Yiu-Kai Ng, David W. Embley |
ACM Trans. Database Syst. | 2 |
| 1995 | Set-Term Unification in a Logic Database Language
SeungJin Lim, Yiu-Kai Ng |
COCOON | 2 |
| 1995 | Set-Term Matching in a Logic Database Language
SeungJin Lim, Yiu-Kai Ng |
DASFAA | 2 |