VLDB 2026 Research / reviewers in the wild / expert
Hung-Yu Kao
dblp:64/5833
· DBLP profile ↗
66ranked-venue papers
9as first author
16since 2021 · last 2025
0000-0002-8890-8544ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 4 first-author · 11 since 2021Databases, data management, data science and information retrieval · 24 · 7 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 2 first-authorHuman-computer interaction and ubiquitous computing · 4Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Theory of computation · 2Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | How Do Position Encodings Affect Length Generalization? Case Studies On In-Context Function LearningabstractThe capability of In-Context Learning (ICL) is crucial for large language models to generalize across a wide range of tasks. By utilizing prompts, these models can accurately predict outcomes for previously unseen tasks without necessitating retraining. However, this generalization ability does not extend to the length of the inputs; the effectiveness of ICL likely diminishes with excessively long inputs, resulting in errors in the generated text. To investigate this issue, we propose a study using a dataset of In-Context functions to understand the operational mechanisms of Transformer models in ICL and length generalization. We generated data using regression and Boolean functions and employed meta-learning techniques to endow the model with ICL capabilities. Our experimental results indicate that position encodings can significantly mitigate length generalization issues, with the most effective encoding extending the maximum input length to over eight times that of the original training length. However, further analysis revealed that while position encoding enhances length generalization, it compromises the model's inherent capabilities, such as its ability to generalize across different data types. Overall, our research illustrates that position encodings have a pronounced positive effect on length generalization, though it necessitates a careful trade-off with data generalization performance. Di-Nan Lin, Jui-Feng Yao, Kun-da Wu, Chen-Hsi Huang, Hung-Yu Kao |
AAAI | 6 |
| 2025 | MAPLE: Enhancing Review Generation with Multi-Aspect Prompt LEarning in Explainable RecommendationabstractChing-Wen Yang, Zhi-Quan Feng, Ying-Jia Lin, Che Wei Chen, Kun-da Wu, Hao Xu, Yao Jui-Feng, Hung-Yu Kao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ching-Wen Yang, Zhi-Quan Feng, Ying-Jia Lin, Che Wei Chen, Kun-da Wu, Jui-Feng Yao, Hung-Yu Kao |
ACL (1) | 8 |
| 2025 | From Persona to Person: Enhancing the Naturalness with Multiple Discourse Relations Graph Learning in Personalized Dialogue Generation
Chih-Hao Hsu, Ying-Jia Lin, Hung-Yu Kao |
PAKDD (2) | 3 |
| 2025 | Exploring the Effectiveness of Pre-training Language Models with Incorporation of Diglossia for Hong Kong ContentabstractIn this article, we present our works to create the first Hong Kong content-based public pre-training dataset and the experiments which resulted in the creation of ELECTRA-based models for commonly used languages in Hong Kong. The creation of pre-training dataset is required for us to study the effect of diglossia on Hong Kong language model, and this is the first ever study on the effect starting all the way from dataset creation phase. Our experiment shows that removing diglossia from pre-training data hurts model performance. We will release our data and models to encourage future studies in Hong Kong languages. 1 Yiu Cheong Yung, Ying-Jia Lin, Hung-Yu Kao |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2024 | CFEVER: A Chinese Fact Extraction and VERification DatasetabstractWe present CFEVER, a Chinese dataset designed for Fact Extraction and VERification. CFEVER comprises 30,012 manually created claims based on content in Chinese Wikipedia. Each claim in CFEVER is labeled as “Supports”, “Refutes”, or “Not Enough Info” to depict its degree of factualness. Similar to the FEVER dataset, claims in the “Supports” and “Refutes” categories are also annotated with corresponding evidence sentences sourced from single or multiple pages in Chinese Wikipedia. Our labeled dataset holds a Fleiss’ kappa value of 0.7934 for five-way inter-annotator agreement. In addition, through the experiments with the state-of-the-art approaches developed on the FEVER dataset and a simple baseline for CFEVER, we demonstrate that our dataset is a new rigorous benchmark for factual extraction and verification, which can be further used for developing automated systems to alleviate human fact-checking efforts. CFEVER is available at https://ikmlab.github.io/CFEVER. Ying-Jia Lin, Chia-Jen Yeh, Yi-Ting Li, Yun-Yu Hu, Chih-Hao Hsu, Mei-Feng Lee, Hung-Yu Kao |
AAAI | 8 |
| 2024 | Contrastive Learning for Unsupervised Sentence Embedding with False Negative Calibration
Chi-Min Chiu, Ying-Jia Lin, Hung-Yu Kao |
PAKDD (3) | 3 |
| 2024 | GViG: Generative Visual Grounding Using Prompt-Based Language Modeling for Visual Question Answering
Yi-Ting Li, Ying-Jia Lin, Chia-Jen Yeh, Hung-Yu Kao |
PAKDD (6) | 5 |
| 2024 | EPRD: Exploiting prior knowledge for evidence-providing automatic rumor detection
Jiawen Li 0003, Ronghui Li, Shiwen Ni, Hung-Yu Kao |
Neurocomputing | 4 |
| 2024 | SUSTEM: An Improved Rule-based Sundanese StemmerabstractCurrent Sundanese stemmers either ignore reduplication words or define rules to handle only affixes. There is a significant amount of reduplication words in the Sundanese language. Because of that, it is impossible to achieve superior stemming precision in the Sundanese language without addressing reduplication words. This article presents an improved stemmer for the Sundanese language, which handles affixed and reduplicated words. With a Sundanese root word list, we use a rules-based stemming technique. In our approach, all stems produced by the affixes removal or normalization processes are added to the stem list. Using a stem list can help increase stemmer accuracy by reducing stemming errors caused by affix removal sequence errors or morphological issues. The current Sundanese language stemmer, RBSS, was used as a comparison. Two datasets with 8,218 unique affixed words and reduplication words were evaluated. The results show that our stemmer's strength and accuracy have improved noticeably. The use of stem list and word reduplication rules improved our stemmer's affixed type recognition and allowed us to achieve up to 99.30% accuracy. Irwan Setiawan, Hung-Yu Kao |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2024 | DropAttack: A Random Dropped Weight Attack Adversarial Training for Natural Language UnderstandingabstractAdversarial training has been proven to be a powerful regularization technique to improve language models. In this work, we propose a novel randomdropped weightattackadversarial training method (DropAttack) for natural language understanding. Our DropAttack improves the generalization of models by minimizing the internal adversarial risk caused by a multitude of attack combinations. Specifically, DropAttack enhances the adversarial attack space by intentionally adding worst-case adversarial perturbations to the weight parameters and randomly dropping the specific proportion of attack perturbations. To extensively validate the effectiveness of DropAttack,12public English natural language understanding datasets were used. Experiments on the GLUE benchmark show that when DropAttack is applied only to the finetuning stage, it is able to improve the overall test scores of the BERT-base pre-trained model from 78.3 to 79.7 and RoBERTa-large pre-trained model from 88.1 to 88.8. Further, DropAttack also significantly improves models trained from scratch. Theoretical analysis reveals that DropAttack performs potential gradient regularization on the input and weight parameters of the model. Moreover, visualization experiments show that DropAttack can push the minimum risk of the neural network to a lower and flatter loss landscape. Shiwen Ni, Jiawen Li 0003, Min Yang 0007, Hung-Yu Kao |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2024 | NEAR: Non-Supervised Explainability Architecture for Accurate Review-Based Collaborative FilteringabstractThere is a critical issue in explainable recommender systems that compounds the challenges of explainability yet is rarely tackled: the lack of ground-truth explanation texts for training. It is unrealistic to expect every user-item pair in a dataset to have a corresponding target explanation. Hence, we pioneer the first non-supervised explainability architecture for review-based collaborative filtering (called NEAR) as our novel contribution to the theory of explanation construction in recommender systems. While maintaining excellent recommendation performance, our approach reformulates explainability as a non-supervised (i.e., unsupervised and self-supervised) explanation generation task. We formally define two explanation types, both of which NEAR can produce. An invariant explanation, fixed for all users, is based on the unsupervised extractive summary of an item's reviews via embedding clustering. Meanwhile, a variant explanation, personalized for a specific user, is a sentence-level text generated by our customized Transformer conditioned on every user-item-rating tuple and artificial ground-truth (self-supervised label) from one of the invariant explanation's sentences. Our empirical evaluation illustrates that NEAR's rating prediction accuracy is better than the other state-of-the-art baselines. Moreover, experiments and assessments show that NEAR-generated variant explanations are more personalized and distinct than those from other Transformer-based models, and our invariant explanations are preferred over those from other contemporary models in real life. Reinald Adrian Pugoy, Hung-Yu Kao |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Advancing Multi-Criteria Chinese Word Segmentation Through Criterion Classification and DenoisingabstractRecent research on multi-criteria Chinese word segmentation (MCCWS) mainly focuses on building complex private structures, adding more handcrafted features, or introducing complex optimization processes.In this work, we show that through a simple yet elegant inputhint-based MCCWS model, we can achieve state-of-the-art (SoTA) performances on several datasets simultaneously.We further propose a novel criterion-denoising objective that hurts slightly on F1 score but achieves SoTA recall on out-of-vocabulary words.Our result establishes a simple yet strong baseline for future MCCWS research. Tzu-Hsuan Chou, Hung-Yu Kao |
ACL (1) | 3 |
| 2023 | Improved Unsupervised Chinese Word Segmentation Using Pre-trained Knowledge and Pseudo-labeling TransferabstractUnsupervised Chinese word segmentation (UCWS) has made progress by incorporating linguistic knowledge from pre-trained language models using parameter-free probing techniques.However, such approaches suffer from increased training time due to the need for multiple inferences using a pre-trained language model to perform word segmentation.This work introduces a novel way to enhance UCWS performance while maintaining training efficiency.Our proposed method integrates the segmentation signal from the unsupervised segmental language model to the pre-trained BERT classifier under a pseudo-labeling framework.Experimental results demonstrate that our approach achieves state-of-the-art performance on the seven out of eight UCWS tasks while considerably reducing the training time compared to previous approaches. Hsiu-Wen Li, Ying-Jia Lin, Yi-Ting Li, Chun Lin, Hung-Yu Kao |
EMNLP | 5 |
| 2023 | CopyCAT: Masking Strategy Conscious Augmented Text for Machine Generated Text Detection
Chien-Liang Liu, Hung-Yu Kao |
PAKDD (1) | 2 |
| 2023 | KPT++: Refined knowledgeable prompt tuning for few-shot text classification
Shiwen Ni, Hung-Yu Kao |
Knowl. Based Syst. | 2 |
| 2021 | Unsupervised Extractive Summarization-Based Representations for Accurate and Explainable Collaborative FilteringabstractReinald Adrian Pugoy, Hung-Yu Kao. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Reinald Adrian Pugoy, Hung-Yu Kao |
ACL/IJCNLP (1) | 2 |
| 2020 | Generalize Sentence Representation with Self-Inference
Kai-Chou Yang, Hung-Yu Kao |
AAAI | 2 |
| 2020 | PSForest: Improving Deep Forest via Feature Pooling and Error ScreeningabstractIn recent years, most of the research on deep learning is based on deep neural networks, which uses the backpropagation algorithm to train parameters of nonlinear layers. Recently, a non-NN style deep model called Deep Forest or gcForest was proposed by Zhou and Feng, which is a deep learning model based on random forests and the training process does not rely on backpropagation. In this paper, we propose PSForest, which can be regarded as a modification of the standard Deep Forest. The main idea for improving the efficiency and performance of the Deep Forest is to do multi-grained pooling of raw features and screening the class vector of each layer based on out-of-bag error. The experiment on different datasets shows that our proposed model achieves predictive accuracy comparable to or better than gcForest, with lower memory requirement and smaller time cost. The study significantly improves the competitiveness of deep forests, further demonstrating that deep learning is more than just deep neural networks. Shiwen Ni, Hung-Yu Kao |
ACML | 2 |
| 2020 | Exploiting Microblog Conversation Structures to Detect RumorsabstractAs one of the most popular social media platforms, Twitter has become a primary source of information for many people.Unfortunately, both valid information and rumors are propagated on Twitter due to the lack of an automatic information verification system.Twitter users communicate by replying to other users' messages, forming a conversation structure.Using this structure, users can decide whether the information in the source tweet is a rumor by reading the tweet's replies, which voice other users' stances on the tweet.The majority of rumor detection researchers process such tweets based on time, ignoring the conversation structure.To reap the benefits of the Twitter conversation structure, we developed a model to detect rumors by modeling conversation structure as a graph.Thus, our model's improved representation of the conversation structure enhances its rumor detection accuracy.The experimental results on two rumor datasets show that our model outperforms several baseline models, including a state-of-theart model. Jiawen Li 0003, Yudianto Sujana, Hung-Yu Kao |
COLING | 3 |
| 2019 | Probing Neural Network Comprehension of Natural Language ArgumentsabstractWe are surprised to find that BERT's peak performance of 77% on the Argument Reasoning Comprehension Task reaches just three points below the average untrained human baseline.However, we show that this result is entirely accounted for by exploitation of spurious statistical cues in the dataset.We analyze the nature of these cues and demonstrate that a range of models all exploit them.This analysis informs the construction of an adversarial dataset on which all models achieve random accuracy.Our adversarial dataset provides a more robust assessment of argument comprehension and should be adopted as the standard in future work. Tim Niven, Hung-Yu Kao |
ACL (1) | 2 |
| 2019 | C-3PO: Click-sequence-aware deeP neural network (DNN)-based Pop-uPs recOmmendation - I know you'll click
Hsien-De Huang, Hung-Yu Kao |
Soft Comput. | 2 |
| 2018 | R2-D2: ColoR-inspired Convolutional NeuRal Network (CNN)-based AndroiD Malware DetectionsabstractThe influence of Deep Learning on image identification and natural language processing has attracted enormous attention globally. The convolution neural network that can learn without prior extraction of features fits well in response to the rapid iteration of Android malware. The traditional solution for detecting Android malware requires continuous learning through pre-extracted features to maintain high performance of identifying the malware. In order to reduce the manpower of feature engineering prior to the condition of not to extract pre-selected features, we have developed a coloR-inspired convolutional neuRal networks (CNN)-based AndroiD malware Detection (R2-D2) system. The system can convert the bytecode of classes.dex from Android archive file to rgb color code and store it as a color image with fixed size. The color image is input to the convolutional neural network for automatic feature extraction and training. The data was collected from Jan. 2017 to Aug 2017. During the period of time, we have collected approximately 2 million of benign and malicious Android apps for our experiments with the help from our research partner Leopard Mobile Inc. Our experiment results demonstrate that the proposed system has accurate security analysis on contracts. Furthermore, we keep our research results and experiment materials on http://R2D2.TWMAN.ORG. Hsien-De Huang, Hung-Yu Kao |
IEEE BigData | 2 |
| 2017 | CDRnN: A high performance chemical-disease recognizer in biomedical literatureabstractDiseases/Chemical play central roles in many areas of biomedical research and healthcare. Consequently, aggregating the disease knowledge and treatment research reports becomes an extremely critical issue, especially in rapid-growth knowledge bases (e.g., PubMed). Thus, a framework of disease/chemical named entity recognition and normalization has become increasingly important for biomedical text mining. In this work, we not only define five diversities of disease names but also develop a system for disease/chemical mention recognition and normalization in biomedical texts. Our system utilizes an order 2 conditional random fields (CRFs) model to develop a recognition system and optimize the results by customizing several post-processing, including abbreviation resolution, consistency improvement, stopwords filtering, and adjectives reorganization. After evaluation, we obtained the best performance (86.9% of F-score) on disease normalization and (89.95% of Precision) on chemical normalization. These results suggest that our system is a high-performance and state of the art recognition system for disease/chemical recognition and normalization from biomedical literature. Hsin-Chun Lee, Hung-Yu Kao |
BIBM | 2 |
| 2017 | Word co-occurrence augmented topic model in short textabstractThe large amount of text on the Internet cause people hard to understand the meaning in a short limit time. Topic models (e.g. LDA and PLSA) have then been proposed to summarize the long text into several topic terms. In the recent years, the short text media such as Twitter is very popular. Howeve r, directly applying the transitional topic model on the short text corpus usually obtains non-coherent topics. It's because that there is no enough words to discover the word co-occurrence patterns in a short document. In this paper, we solve the problem of lack of the local word co-occurrence in LDA. Thus, we proposed an improvement of word co-occurrence method to enhance the topic models. We generate new virtual documents by re-organizing the words in documents and use it to enhance the traditional LDA. The experimental results show that our re-organized LDA (RO-LDA) method gets better results in the noisy Tweet dataset and the regular news dataset. Moreover, in our proposed augmented model, we do not need any external data. Our proposed methods are only based on the original topic model, thus our methods can easily apply to other existing LDA based models. Guan-Bin Chen, Hung-Yu Kao |
Intell. Data Anal. | 2 |
| 2017 | Special issue on soft computing for knowledge management and web applications
Chang-Shing Lee, Hung-Yu Kao |
Soft Comput. | 2 |
| 2017 | k--anonymization of multiple shortest paths
Shyue-Liang Wang, Yu-Chuan Tsai, Tzung-Pei Hong, Hung-Yu Kao |
Soft Comput. | 4 |
| 2015 | LDA based semi-supervised learning from streaming short textabstractWith the rapidly growing of real-time social media, like Twitter, many users share and discuss their interest topics through such platforms. Hashtag is a type of metadata tag which allows users to annotate their topics of tweets. For research usage, for example, hashtags can help the performance of event detection by observing the trend of hashtags. Although Twitter grows rapidly, hashtag growth is not as expected. Our dataset shows that there are less than 20% of all tweets containing hashtags. We think that it is caused by that most users may have no idea what hashtags are suitable for tweets they post. If we can recommend suitable hashtags to users, it can be one of the solutions to solve the problem of low usage rate of hashtag. Hashtag recommendation belongs to supervised learning problem. More labeled data for training the learning model can get higher performance in prediction. However, labeled data in hashtag recommendation is not so much due to low usage rate of hashtag. Thus, we want to exploit unlabeled data, i.e. non-hashtag tweets, to solve this problem. Now we have large amount of unlabeled data, but directly adding all non-hashtag tweets may not be helpful to train the model. To overcome this issue, we apply the weight-updating mechanisms to filter out the useless parts of non-hashtag tweets. These mechanisms also have to consider the temporal characteristics of hashtag due to the real-time nature of Twitter. The experimental results in this research show that adding non-hashtag tweets to extend original training data outperforms baseline methods which only exploit labeled data to train the model. Ji-De Chen, Hung-Yu Kao |
DSAA | 2 |
| 2015 | Curatable Named-Entity Recognition Using Semantic RelationsabstractNamed-entity recognition (NER) plays an important role in the development of biomedical databases. However, the existing NER tools produce multifarious named-entities which may result in both curatable and non-curatable markers. To facilitate biocuration with a straightforward approach, classifying curatable named-entities is helpful with regard to accelerating the biocuration workflow. Co-occurrence Interaction Nexus with Named-entity Recognition (CoINNER) is a web-based tool that allows users to identify genes, chemicals, diseases, and action term mentions in the Comparative Toxicogenomic Database (CTD). To further discover interactions, CoINNER uses multiple advanced algorithms to recognize the mentions in the BioCreative IV CTD Track. CoINNER is developed based on a prototype system that annotated gene, chemical, and disease mentions in PubMed abstracts at BioCreative 2012 Track I (literature triage). We extended our previous system in developing CoINNER. The pre-tagging results of CoINNER were developed based on the state-of-the-art named entity recognition tools in BioCreative III. Next, a method based on conditional random fields (CRFs) is proposed to predict chemical and disease mentions in the articles. Finally, action term mentions were collected by latent Dirichlet allocation (LDA). At the BioCreative IV CTD Track, the best F-measures reached for gene/protein, chemical/drug and disease NER were 54 percent while CoINNER achieved a 61.5 percent F-measure. System URL: http://ikmbio.csie.ncku.edu.tw/coinner/ introduction.htm. Yi-Yu Hsu, Hung-Yu Kao |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2014 | Question difficulty evaluation by knowledge gap analysis in Question Answer communitiesabstractThe Community Question Answer (CQA) service is a typical forum of Web 2.0 that shares knowledge among people. There are thousands of questions that are posted and solved every day. Because of the various users of the CQA service, question search and ranking are the most important topics of research in the CQA portal. In this study, we addressed the problem of identifying questions as being hard or easy by means of a probability model. In addition, we observed the phenomenon called knowledge gap that is related to the habit of users and used a knowledge gap diagram to illustrate how much of a knowledge gap exists in different categories. To this end, we proposed an approach called the knowledge-gap-based difficulty rank (KG-DRank) algorithm, which combines the user-user network and the architecture of the CQA service to find hard questions. We used f-measure, AUC, MAP, NDCG, precision@Top5 and concordance analysis to evaluate the experimental results. Our results show that our approach leads to better performance than other baseline approaches across all evaluation metrics. Chih-Lu Lin, Ying-Liang Chen, Hung-Yu Kao |
ASONAM | 3 |
| 2014 | User preference space partition and product filters for reverse top-k queriesabstractTop-k queries have been studied mainly from the perspective of the user. Many researchers have focused on improving the efficiency of top-k problems. However, few studies have focused on the essential factors required for manufacturers to assess the potential market. A novel query type, namely, the reverse top-k, is used to assess the potential market and help manufacturers calculate the impact of their products. Given a potential product, reverse top-k will find the user preferences for which this product is in the top-k query result set. Although several algorithms can solve the reverse top-k problem, none that are available can solve the reverse top-k problem when the number of products or users is large. In this paper, we formally define our algorithm as FSP (filtering and space partition) and explain how FSP solves the reverse top-k problem. The main idea of FSP is to use the partition of the candidate space to reduce the searching of space for products. In our experimental results, FSP can find the same results as other algorithms, but FSP reduces the time cost from 231 msec to 32 msec. Zong-Hua Yang, Hung-Yu Kao |
DSAA | 2 |
| 2014 | [K1, K2]-anonymization of Shortest PathsabstractPrivacy preserving network publishing has been studied extensively in recent years. Although more works have adopted un-weighted graphs to model network relationships, weighted graph modeling can provide deeper analysis of the degree of relationships. Previous works on weighted graph privacy have concentrated on preserving the shortest path characteristic between pairs of vertices. Two common types of privacy have been proposed. One type of privacy tried to add random noise edge weights to the graph but still maintain the same shortest path. The other privacy, k-shortest path privacy, minimally perturbed edge weights so that there exists k shortest paths. However, the k-shortest path privacy only considers anonymizing same fixed number of shortest paths for all pairs of source and destination vertices. In this work, we present a new concept called [k1, k2]-shortest path privacy to allow different number of shortest paths for different pairs of vertices. A published network graph with [k1, k2]-shortest path privacy has at least k' indistinguishable shortest paths between the source and destination vertices, where k1 ≦ k' ≦ k2. A heuristic algorithm based on modifying only Non-Visited (NV) edges is proposed and experimental results showing the feasibility and characteristics of the proposed approach are presented. Yu-Chuan Tsai, Shyue-Liang Wang, Tzung-Pei Hong, Hung-Yu Kao |
MoMM | 4 |
| 2014 | Latent Features Based Prediction on New Users' Tastes
Ming-Chu Chen, Hung-Yu Kao |
PAKDD (1) | 2 |
| 2014 | On anonymizing transactions with sensitive items
Shyue-Liang Wang, Yu-Chuan Tsai, Hung-Yu Kao, Tzung-Pei Hong |
Appl. Intell. | 3 |
| 2014 | Using hidden Markov models to predict DNA-binding proteins with sequence and structure information
Yi-Yu Hsu, Wei-Jhih Chen, Shu Hui Chen, Hung-Yu Kao |
Soft Comput. | 4 |
| 2014 | IT2FS-based ontology with soft-computing mechanism for malware behavior analysis
Hsien-De Huang, Chang-Shing Lee, Mei-Hui Wang, Hung-Yu Kao |
Soft Comput. | 4 |
| 2014 | Gene Name Disambiguation UsingMulti-Scope Species DetectionabstractSpecies detection is an important topic in the text mining field. According to the importance of the research topics (e.g., species assignment to genes and document focus species detection), some studies are dedicated to an individual topic. However, no researcher to date has discussed species detection as a general problem. Therefore, we developed a multi-scope species detection model to identify the focus species for different scopes (i.e., gene mention, sentence, paragraph, and global scope of the entire article). Species assignment is one of the bottlenecks of gene name disambiguation. In our evaluation, recognizing the focus species of a gene mention in four different scopes improved the gene name disambiguation. We used the species cue words extracted from articles to estimate the relevance between an article and a species. The relevance score was calculated by our proposed entities frequency-augmented invert species frequency (EF-AISF) formula, which represents the importance of an entity to a species. We also defined a relation guide factor (RGF) to normalize the relevance score. Our method not only achieved better performance than previous methods but also can handle the articles that do not specifically mention a species. In the DECA corpus, we outperformed previous studies and obtained an accuracy of 88.22 percent. Jui-Chen Hsiao, Chih-Hsuan Wei, Hung-Yu Kao |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2013 | Genealogical-based method for ontology self-extension in MeSHabstractDuring the last decade, the advent of Ontologies used for biomedical annotation has had a deep impact on life science. MeSH is a well-known Ontology for the purpose of indexing journal articles in PubMed, improving literature searching on multi-domain topics. Since the explosion of data growth in recent years, there are new terms, concepts that weed through the old and bring forth the new. Automatically extending sets of existing terms will enable bio-curators to systematically improve text-based ontologies level by level. However, most of the related techniques which apply symbolic patterns based on a literature corpus tend to focus on more general but not specific parts of the ontology. Therefore, in this work, we present a novel method for utilizing genealogical information from Ontology itself to find suitable siblings for ontology extension. Based on the breadth and depth dimensions, the sibling generation stage and pruning strategy are proposed in our approach. As a result, on the average, the precision of the genealogical-based method achieved 0.5, with the best 0.83 performance of category “Organisms”. We also achieve average precision 0.69 of 229 new terms in MeSH 2013 version. Yu-Wen Guo, Hung-Yu Kao |
BIBM | 2 |
| 2013 | Up or Down? Click-Through Rate Prediction from Social Intention for Search AdvertisingabstractIn search advertising, advertisers should carefully compose keywords in order to enhance the opportunity for ads to be clicked. Thus, timely presenting proper advertisements to users will encourage them to click on search ads. Until now, how to efficiently improve the ad performance to earn more clicks remains a main task. In this paper, we focus on the scope of smart phone and produce a social intentional model with advertising based features to forecast future trend on ads' click-through rate (CTR). In terms of social intentional model, we analyze Chinese text content of technology forum to derive social intentional factors which are Hotness, Sentiment, Promotion, and Event. Our results indicate that with knowing public opinions or occurring events beforehand can efficiently enhance click prediction. This will be very helpful for advertisers on adjusting bidding keywords to improve ad performance via social intention. Hung-Yu Kao |
iiWAS | 2 |
| 2013 | Automatic Domain-Specific Sentiment Lexicon Generation with Label PropagationabstractNowadays, the advance of social media has led to the explosive growth of opinion data. Therefore, sentiment analysis has attracted a lot of attentions. Currently, sentiment analysis applications are divided into two main approaches, the lexicon-based approach and the machine-learning approach. However, both of them face the challenge of obtaining a large amount of human-labeled training data and corpus. For the lexicon-based approach, it requires a sentiment lexicon to determine the opinion polarity. There are many existing benchmark sentiment lexicons, but they cannot cover all the domain-specific words meanings. Thus, automatic generation of a domain-specific sentiment lexicon becomes an important task. We propose a framework to automatically generate sentiment lexicon. First, we determine the semantic similarity between two words in the entire unlabeled corpus. We treat the words as nodes and similarities as weighted edges to construct word graphs. A graph-based semi-supervised label propagation method finally assigns the polarity to unlabeled words through the proposed propagation process. Experiments conducted on the microblog data, Twitter, show that our approach leads to a better performance than baseline approaches and general-purpose sentiment dictionaries. Yen-Jen Tai, Hung-Yu Kao |
iiWAS | 2 |
| 2013 | An IT2FLS-Based Malware Analysis Mechanism: Malware Analysis Network in Taiwan (MiT)abstractMalware is one of the problems really existing in the modern post-industrial society. Hackers continuously develop novel techniques to intrude into computer systems for various reasons, so many security researchers should analyze and track new malicious program to protect sensitive information for the computer system. In this paper, we integrate the Interval Type-2 Fuzzy Logic System (IT2FLS) with malware behavioral analysis: Malware Analysis Network in Taiwan (MAN in Taiwan, MiT, and http://MiT.TWMAN.ORG). The core techniques of MiT are as follows: (1) automatically collect the logs the difference operation system to extract unknown behavior information. Also, MiT is able to automatically provide and share samples and reports via the cloud storage mechanism, (2) integrate with IT2FLS to construct the malware analysis domain knowledge for the malware behavior. Simulation results show that the proposed approach can effectively execute the malware behavior analysis, and the constructed system has also been released under GNU General Public License version 3. Hsien-De Huang, Chang-Shing Lee, Mei-Hui Wang, Hung-Yu Kao |
SMC | 4 |
| 2013 | tmVar: a text mining approach for extracting sequence variants in biomedical literatureabstractMOTIVATION: Text-mining mutation information from the literature becomes a critical part of the bioinformatics approach for the analysis and interpretation of sequence variations in complex diseases in the post-genomic era. It has also been used for assisting the creation of disease-related mutation databases. Most of existing approaches are rule-based and focus on limited types of sequence variations, such as protein point mutations. Thus, extending their extraction scope requires significant manual efforts in examining new instances and developing corresponding rules. As such, new automatic approaches are greatly needed for extracting different kinds of mutations with high accuracy. RESULTS: Here, we report tmVar, a text-mining approach based on conditional random field (CRF) for extracting a wide range of sequence variants described at protein, DNA and RNA levels according to a standard nomenclature developed by the Human Genome Variation Society. By doing so, we cover several important types of mutations that were not considered in past studies. Using a novel CRF label model and feature set, our method achieves higher performance than a state-of-the-art method on both our corpus (91.4 versus 78.1% in F-measure) and their own gold standard (93.9 versus 89.4% in F-measure). These results suggest that tmVar is a high-performance method for mutation extraction from biomedical literature. AVAILABILITY: tmVar software and its corpus of 500 manually curated abstracts are available for download at http://www.ncbi.nlm.nih.gov/CBBresearch/Lu/pub/tmVar Chih-Hsuan Wei, Bethany R. Harris, Hung-Yu Kao, Zhiyong Lu |
Bioinform. | 3 |
| 2013 | Shortest Paths Anonymization on Weighted GraphsabstractDue to the proliferation of online social networking, a large number of personal data are publicly available. As such, personal attacks, reputational, financial, or family losses might occur once this personal and sensitive information falls into the hands of malicious hackers. Research on Privacy-Preserving Network Publishing has attracted much attention in recent years. But most work focus on node de-identification and link protection. In academic social networks, business transaction networks, and transportation networks, etc, node identities and link structures are public knowledge but weights and shortest paths are sensitive. In this work, we study the problem of k-anonymous path privacy. A published network graph with k-anonymous path privacy has at least k indistinguishable shortest paths between the source and destination vertices [21]. In order to achieve such privacy, three different strategies of modification on edge weights of directed graphs are proposed. Numerical comparisons show that weight-proportional-based strategy is more efficient than PageRank-based and degree-based strategies. In addition, it is also more efficient and causes less information loss than running on un-directed graphs. Shyue-Liang Wang, Yu-Chuan Tsai, Hung-Yu Kao, I-Hsien Ting, Tzung-Pei Hong |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2012 | TWMAN+: A Type-2 fuzzy ontology model for malware behavior analysisabstractClassical ontology is not sufficient to deal with vague or imprecise knowledge for real world applications such as malware behavioral analysis. In addition, malware has grown into a pressing problem for governments and commercial organizations. Anti-malware applications represent one of the most important research topics in the area of information security threat. As a countermeasure, enhanced systems for analyzing the behavior of malware are needed in order to predict malicious actions and minimize computer damages. Many researchers use Virtual Machine (VM) systems to monitor malware behavior, but there are many Anti-VM techniques which are used to counteract the collection, analysis, and reverse engineering features of the VM based malware analysis platform. Therefore, malware researchers are likely to obtain inaccurate analysis from the VM based approach. For this reason, we have developed the Taiwan Malware Analysis Net (TWMAN) which uses a real operating system environment to improve the accuracy of malware behavior analysis and has integrated Type-1 Fuzzy Set (T1FS), Ontology, and Fuzzy Markup Language (FML) on 2010. In this paper, we use Interval Type-2 Fuzzy Set (IT2FS), eggdrop, and glftpd as a cloud service (software as a service) on the Google App Engine along with Python and Android. We believe this system can help improve the correctness of malware analysis results and reduce the rate of malware misdiagnosis. Hsien-De Huang, Chang-Shing Lee, Hani Hagras, Hung-Yu Kao |
SMC | 4 |
| 2012 | The Retrieval of Important News Stories by Influence Propagation among Communities and CategoriesabstractNowadays, people receive information of the news stories not only from newspapers but also from online news websites. They search important news stories in order to know what happen today. However, it is hard to browse all the news stories published on a day. It is necessary to identify which news stories are more newsworthy on the specific day. In this paper, we investigate how to automatically identify the importance of news stories for different news categories on a specific day by utilizing the influence propagation among communities and news categories. In particular, we build an influence propagation model which consists of three features: category relevance, bloggers' attention and bursty influence. Based on this influence propagation model, we propose a Cross-Category Social Influence Propagation (C-SIP) approach for scoring the importance of news stories on a specific day. We evaluate our approach by using the judgment of Story Ranking Task in TREC 2010 Blog Track. The experiment shows our approach attains a prominent performance in the retrieval of important news stories and gets 9.94% improvement over the best performance of participating systems in TREC 2010 Blog Track. Yu-Fan Lin, Hung-Yu Kao |
Web Intelligence | 2 |
| 2011 | CAIS: Community Based Annotation Insight Search in a Folksonomy NetworkabstractFolksonomy systems provide a way for users to share and organize bookmarks. The social relationship among users has become stronger with the rapid development of new technologies. Finding the leading objects has become an important topic. These research topics are always centered around finding the most popular pages or experts. In this paper, we propose a new notion of expertise, which we call user insight. User insight denotes the user's expertise in finding Web pages that are useful or have the potential to be popular pages before other users find them. To address the issue, we refer to three major types of Web pages, namely, isolated, well-known, and burgeoning. Burgeoning pages are exceptionally useful and attractive for users in a folksonomy system. In our paper, we build a time-based algorithm to estimate user insight. In addition, we discuss the social relationship within fan networks, and we propose a link-based algorithm called CAIS (Community-based Annotation Insight Search) to realize the reinforcement between users, communities and pages. Finally, we design several experiments to evaluate the performance of CAIS and compare it to other approaches. We prove that CAIS has a better performance for the user ranking of simulated data and real data from Delicious. Han-Chang Huang, Hung-Yu Kao |
ASONAM | 2 |
| 2011 | Evolutional Dependency Parse Trees for Biological Relation ExtractionabstractDue to the rapid growth in biological technology, the development of high-quality information extraction systems is needed and still remains a challenge. Several recently proposed approaches to biological relation extraction are based on machine learning techniques on lexical and syntactic information. Most use the dependency path between two genes/proteins instead of the whole dependency tree of a sentence for identifying relationships. However, the dependency path may not have any node between two entities. If a limited set of annotated training corpora is used for the construction of tree information of biological relationships, the training corpus will lack some sentence structures and cannot predict whether the sentence has a biological relationship. In this paper, we developed a biological relation extraction system called Evolutional Tree Extraction System - ETree. We extended the dependency path to the dependency subtree and developed a method that can automatically expand and prune these existing dependency subtrees into various dependency subtrees. These dependency subtrees are called "Evolutional Trees" and are used to predict the biological relationship sentences. Hung-Yu Kao, Yi-Tsung Tang, Jian-Fu Wang |
BIBE | 1 |
| 2011 | Augmented Transitive Relationships in Direct Protein-Protein Interaction PredictionabstractThe prediction of new protein-protein interactions is important to the discovery of the currently unknown function of various biological pathways. In addition, many databases of protein-protein interactions contain different types of interactions, including protein associations, physical protein associations and direct protein interactions. There are only a few studies that consider the issues inherent to the prediction of direct protein-protein interactions, that is, interactions between proteins that are actually in direct physical contact and are listed in known protein interaction databases. Predicting these interactions is a crucial and challenging task. Therefore, it is increasingly important to discover not only protein associations but also direct interactions. Many studies have predicted protein-protein interactions directly, by using biological features such as Gene Ontology (GO) functions and protein structural domains of two proteins with unknown interactions. In this article, we proposed an augmented transitive relationships predictor (ATRP), a new method of predicting potential direct protein-protein interactions by using transitive relationships and annotations of protein interactions. Our results demonstrate that ATRP can effectively predict unknown direct protein-protein interactions from existing protein interaction relationships. The average accuracy of this method outperformed GO-based prediction methods by a factor ranging from 28% to 62%. Yi-Tsung Tang, Hung-Yu Kao |
CISIS | 2 |
| 2011 | Applying FML and Fuzzy Ontologies to malware behavioural analysisabstractAntimalware applications represent one of the most important research topic in the area of information security threat. Indeed, most computer network issues have malwares as their underlying cause. As a consequence, enhanced systems for analyzing the behavior of malwares are needed in order to try to predict their malicious actions and minimize eventual computer damages. However, because the environments where malwares operate are characterized by high levels of imprecision and vagueness, the conventional data analysis tools lack to deal with these computer safety applications. This work tries to bridge this gap by integrating semantic technologies and computational intelligence methods, such as the Fuzzy Ontologies and Fuzzy Markup Language (FML), in order to propose an advanced semantic decision making system that, as shown by experimental results, achieves good performances in terms of malicious programs identification. Hsien-De Huang, Giovanni Acampora, Vincenzo Loia, Chang-Shing Lee, Hung-Yu Kao |
FUZZ-IEEE | 5 |
| 2011 | TransDomain: A Transitive Domain-Based Method in Protein-Protein Interaction Prediction
Yi-Tsung Tang, Hung-Yu Kao |
ISBRA | 2 |
| 2011 | Inference of transcriptional regulatory network by bootstrapping patternsabstractMOTIVATION: Transcriptional regulatory networks, which consist of linkages between transcription factors (TF) and target genes (TGene), control the expression of a genome and play important roles in all aspects of an organism's life cycle. Accurate prediction of transcriptional regulatory networks is critical in providing useful information for biologists to determine what to do next. Currently, there is a substantial amount of fragmented gene regulation information described in the medical literature. However, current related text analysis methods designed to identify protein-protein interactions are not entirely suitable for finding transcriptional regulatory networks. RESULT: In this article, we propose an automatic regulatory network inference method that uses bootstrapping of description patterns to predict the relationship between a TF and its TGenes. The proposed method differs from other regulatory network generators in that it makes use of both positive and negative patterns for different vector combinations in a sentence. Moreover, the positive pattern learning process can be fully automatic. Furthermore, patterns for active and passive voice sentences are learned separately. The experiments use 609 HIF-1 expert-tagged articles from PubMed as the gold standard. The results show that the proposed method can automatically generate a predicted regulatory network for a transcription factor. Our system achieves an F-measure of 72.60%. AVAILABILITY: The software, training/test datasets and learned patterns are available at http://140.116.99.138/∼hcw0901/PubMedSearch.php. Hei-Chia Wang, Yi-Hsiu Chen, Hung-Yu Kao, Shaw-Jenq Tsai |
Bioinform. | 3 |
| 2011 | The gene normalization task in BioCreative IIIabstractBACKGROUND: We report the Gene Normalization (GN) challenge in BioCreative III where participating teams were asked to return a ranked list of identifiers of the genes detected in full-text articles. For training, 32 fully and 500 partially annotated articles were prepared. A total of 507 articles were selected as the test set. Due to the high annotation cost, it was not feasible to obtain gold-standard human annotations for all test articles. Instead, we developed an Expectation Maximization (EM) algorithm approach for choosing a small number of test articles for manual annotation that were most capable of differentiating team performance. Moreover, the same algorithm was subsequently used for inferring ground truth based solely on team submissions. We report team performance on both gold standard and inferred ground truth using a newly proposed metric called Threshold Average Precision (TAP-k). RESULTS: We received a total of 37 runs from 14 different teams for the task. When evaluated using the gold-standard annotations of the 50 articles, the highest TAP-k scores were 0.3297 (k=5), 0.3538 (k=10), and 0.3535 (k=20), respectively. Higher TAP-k scores of 0.4916 (k=5, 10, 20) were observed when evaluated using the inferred ground truth over the full test set. When combining team results using machine learning, the best composite system achieved TAP-k scores of 0.3707 (k=5), 0.4311 (k=10), and 0.4477 (k=20) on the gold standard, representing improvements of 12.4%, 21.8%, and 26.6% over the best team results, respectively. CONCLUSIONS: By using full text and being species non-specific, the GN task in BioCreative III has moved closer to a real literature curation task than similar tasks in the past and presents additional challenges for the text mining community, as revealed in the overall team results. By evaluating teams using the gold standard, we show that the EM algorithm allows team submissions to be differentiated while keeping the manual annotation effort feasible. Using the inferred ground truth we show measures of comparative performance between teams. Finally, by comparing team rankings on gold standard vs. inferred ground truth, we further demonstrate that the inferred ground truth is as effective as the gold standard for detecting good team performance. Zhiyong Lu, Hung-Yu Kao, Chih-Hsuan Wei, Minlie Huang, Jingchen Liu, Cheng-Ju Kuo, Chun-Nan Hsu, Richard Tzong-Han Tsai, Hong-Jie Dai, Naoaki Okazaki, Hancheol Cho, Martin Gerner, Illés Solt, Shashank Agarwal, Dina Vishnyakova, Patrick Ruch, Martin Romacker, Fabio Rinaldi 0001, Sanmitra Bhattacharya, Padmini Srinivasan, Manabu Torii, Sérgio Matos, David Campos 0001, Karin Verspoor, Kevin M. Livingston, W. John Wilbur |
BMC Bioinform. | 2 |
| 2011 | Cross-species gene normalization by species inferenceabstractBACKGROUND: To access and utilize the rich information contained in the biomedical literature, the ability to recognize and normalize gene mentions referenced in the literature is crucial. In this paper, we focus on improvements to the accuracy of gene normalization in cases where species information is not provided. Gene names are often ambiguous, in that they can refer to the genes of many species. Therefore, gene normalization is a difficult challenge. METHODS: We define "gene normalization" as a series of tasks involving several issues, including gene name recognition, species assignation and species-specific gene normalization. We propose an integrated method, GenNorm, consisting of three modules to handle the issues of this task. Every issue can affect overall performance, though the most important is species assignation. Clearly, correct identification of the species can decrease the ambiguity of orthologous genes. RESULTS: In experiments, the proposed model attained the top-1 threshold average precision (TAP-k) scores of 0.3297 (k=5), 0.3538 (k=10), and 0.3535 (k=20) when tested against 50 articles that had been selected for their difficulty and the most divergent results from pooled team submissions. In the silver-standard-507 evaluation, our TAP-k scores are 0.4591 for k=5, 10, and 20 and were ranked 2nd, 2nd, and 3rd respectively. AVAILABILITY: A web service and input, output formats of GenNorm are available at http://ikmbio.csie.ncku.edu.tw/GN/. Chih-Hsuan Wei, Hung-Yu Kao |
BMC Bioinform. | 2 |
| 2010 | Represented indicator measurement and corpus distillation on focus species detectionabstractIn extraction of information from the biomedical literature, name disambiguation of domain-specific entities, such as proteins, is one of the most important issues. The entity ambiguity with the highest dimension is the species to which an entity is associated with. Furthermore, one of the bottlenecks in inter-species gene name normalization is species disambiguation. To enhance the performance of species disambiguation, the detection of focus species detection remains a substantial challenge. This study presents a method addressing this issue. The results present evaluations of all articles from the BioCreaTive I&II GN task. Our method is robust for all types of articles, particularly those without explicit species entity information. Since our method requires a training corpus to be the indicator vector, we developed an iterative corpus distillation method to extend the corpus. In the conducted experiments, the proposed method achieved a high accuracy of 85.64% and 84.32% without species entity information. Chih-Hsuan Wei, Hung-Yu Kao |
BIBM | 2 |
| 2010 | Multi-table association rules hidingabstractMany approaches for preserving association rule privacy, such as association rule mining outsourcing, association rule hiding, and anonymity, have been proposed. In particular, association rule hiding on single transaction table has been well studied. However, hiding multi-relational association rule in data warehouses is not yet investigated. This work presents a novel algorithm to hide predictive association rules on multiple tables. Given a target predictive item, a technique is proposed to hide multi-relational association rules containing the target item without joining the multiple tables. Examples and analyses are given to demonstrate the efficiency of the approach. Shyue-Liang Wang, Tzung-Pei Hong, Yu-Chuan Tsai, Hung-Yu Kao |
ISDA | 4 |
| 2010 | A Categorized Sentiment Analysis of Chinese Reviews by Mining Dependency in Product Features and Opinions from BlogsabstractIn the past, there have been many documents focusing on English reviews for sentiment analysis. These contain abundant research results which extract features and opinions, identify semantic orientation, and associate features with opinions. Although this approach has performed well for English reviews, it is not as successful with Chinese reviews. In this paper, we aim to develop a sentiment analysis system that is suitable for Chinese reviews. This system would extract features that users are interested in and detect those opinions with semantic orientations that accord with the dependency of certain features and opinions in one specific category. We then present users with the integrated results. Our experiments show that the derived system can effectively measure the dependency between features and opinions. The prominent performance of review sentiment analysis also validates the applicability of the proposed method. Hung-Yu Kao, Zi-Yu Lin |
Web Intelligence | 1 |
| 2009 | Normalizing Biomedical Name Entities by Similarity-Based Inference Network and De-ambiguity MiningabstractTo construct an intelligent biomedical knowledge management system, researchers had proposed many relation extraction methods in past. Before applying these methods, the system has to recognize the name entities in the literature and map the entities to the relative EntrezIDs. The purpose of this study is to automatically and exactly identify the relative EntrezIDs which are mentioned in literatures. We employ the similarity-based inference network to calculate the similarity score with the entities, and this EntrezID is a solution to the term variation problem. The proposed de-ambiguity strategy increases the confidence of EntrezID in literature. The strategy provides researchers a good utilization of information for mapping the entity to the EntrezID. As a result, the precision of system increase about 75.1%, and it makes the identified entity even more meaningful. The system using the proposed strategies outperforms the previous methods in biomedical entity normalization. Chih-Hsuan Wei, I-Chin Huang, Yi-Yu Hsu, Hung-Yu Kao |
BIBE | 4 |
| 2009 | Entropy-Based Visual Tree Evaluation on Block ExtractionabstractMore and More people use Cascading Style Sheets (CSS) to manage their Web pages, because CSS is easy and convenient to typesetting. However, CSS makes a Web page displayed in an ambiguous structure. The data extraction systems that based on mining the Web page structure would generate false judgments for these CSS-rich pages. For solving this issue, we propose a system that applies properties of CSS Web pages to extract data blocks. In this system, Web pages are converted into a visual tree and the entropy attributes of each node in a visual tree is calculated. In the experiment, the result shows the node attributes and the visual tree are useful to extract blocks on CSS Web pages. Our system also outperforms with other systems on container block extraction. Wei-Ting Cho, Yu-Min Lin, Hung-Yu Kao |
Web Intelligence | 3 |
| 2008 | DRANK+: A Directory Based Pagerank Prediction Method for Fast Pagerank Convergence
Hung-Yu Kao, Chia-Sheng Liu, Yu-Chuan Tsai, Chia Chun Shih, Tse-Ming Tsai |
WEBIST (2) | 1 |
| 2007 | A Fast PageRank Convergence Method based on the Cluster PredictionabstractIn recent years, search engines have already played the key roles among Web applications, and link analysis algorithms are the major methods to measure the important values of Web pages. These algorithms employ the conventional flat Web graph built by Web pages and link relations of Web pages to obtain the relative importance of Web objects. Previous researches have observed that PageRank-like link analysis algorithms have a bias against newly created Web pages. A new ranking algorithm called Page Quality was then proposed to solve this issue. Page Quality predicates future ranking values by the difference rate between the current ranking value and the previous ranking value. In this paper, we propose a new algorithm called DRank to diminish the bias of PageRank-like link analysis algorithms, and attain the better performance than Page Quality. In this algorithm, we model Web graph as a three-layer graph which includes Host Graph, Directory Graph and Page Graph by using the hierarchical structure of URLs and the structure of link relation of Web pages. We calculate the importance of Hosts, Directories and Pages by weighted graph we built and then the clustering distribution of PageRank values of pages within directories is observed. We can then predicate the more accurate values of page importance to diminish the bias of newly created pages by the clustering characteristic of PageRank. Experiment results show that DRank algorithm works well on predicating future ranking values of pages and outperform Page Quality. Hung-Yu Kao, Seng-Feng Lin |
Web Intelligence | 1 |
| 2006 | The Mining and Extraction of Primary Informative Blocks and Data Objects from Systematic Web PagesabstractWith the fast development of Internet, the Web has already been an enormous database so far, which contains extremely abundant information. Most of Web pages are represented their content by using a list of objects, such as search engine results, product information of shopping Web sites and so on, and these objects form the primary information of each page. In this paper, we focus on the issues of mining primary information and the constituted object groups. The system is divided into three major phases: (1) By transforming each Web page into corresponding tree structures, our system can visit all regions of the Web page in an efficient way, and detects the informative parts. (2) We design and quantize several novel features according to the characters of regions of a Web page. (3) A weighting model is proposed that calculates the important degree of each region, we then extract the primary information of the Web pages. The experimental result proves our system can be applied to a large number of Web pages with different themes and styles to find the correct primary information and the list of corresponding objects Yi-Feng Tseng, Hung-Yu Kao |
Web Intelligence | 2 |
| 2005 | WISDOM: Web Intrapage Informative Structure Mining Based on Document Object ModelabstractTo increase the commercial value and accessibility of pages, most content sites tend to publish their pages with intrasite redundant information, such as navigation panels, advertisements, and copyright announcements. Such redundant information increases the index size of general search engines and causes page topics to drift. In this paper, we study the problem of mining intrapage informative structure in news Web sites in order to find and eliminate redundant information. Note that intrapage informative structure is a subset of the original Web page and is composed of a set of fine-grained and informative blocks. The intrapage informative structures of pages in a news Web site contain only anchors linking to news pages or bodies of news articles. We propose an intrapage informative structure mining system called WISDOM (Web intrapage informative structure mining based on the document object model) which applies Information Theory to DOM tree knowledge in order to build the structure. WISDOM splits a DOM tree into many small subtrees and applies a top-down informative block searching algorithm to select a set of candidate informative blocks. The structure is built by expanding the set using proposed merging methods. Experiments on several real news Web sites show high precision and recall rates which validates WISDOM'S practical applicability. Hung-Yu Kao, Jan-Ming Ho, Ming-Syan Chen |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2004 | Extracting Citation Metadata from Online Publication Lists Using BLAST
I-Ane Huang, Jan-Ming Ho, Hung-Yu Kao, Shian-Hua Lin |
PAKDD | 3 |
| 2004 | DOMISA: DOM-Based Information Space Adsorption of Web Information Hierarchy MiningabstractDue to the growth of dynamic page generation techniques, the amount and the complexity of Web pages has been increasing explosively, as has the information contained within Web pages. Redundant and irrelevant information is distributed and mixed throughout a page, making it difficult to automatically identify the useful information in that page. Consequently, we propose an information hierarchy in this paper, and, from that hierarchy, we can extract the significance and the relationship value of information contained within a Web page. We can then use this hierarchical structure to create a new browsing process. Our DOM-based Information Space Adsorption (DOMISA) system applies information theory to map information in a page into an information space, and our gradient tree adsorption (GTA) process uses the document object model (DOM) trees of pages to build information hierarchies. Experiments on several commercial news Web sites show high precision and recall rates achieved by DOMISA in determining information clusters of pages which validates its practical applicability to Web sites. Hung-Yu Kao, Jan-Ming Ho, Ming-Syan Chen |
SDM | 1 |
| 2004 | Mining Web Informative Structures and Contents Based on Entropy AnalysisabstractWe study the problem of mining the informative structure of a news Web site that consists of thousands of hyperlinked documents. We define the informative structure of a news Web site as a set of index pages (or referred to as TOC, i.e., table of contents, pages) and a set of article pages linked by these TOC pages. Based on the Hyperlink Induced Topics Search (HITS) algorithm, we propose an entropy-based analysis (LAMIS) mechanism for analyzing the entropy of anchor texts and links to eliminate the redundancy of the hyperlinked structure so that the complex structure of a Web site can be distilled. However, to increase the value and the accessibility of pages, most of the content sites tend to publish their pages with intrasite redundant information, such as navigation panels, advertisements, copy announcements, etc. To further eliminate such redundancy, we propose another mechanism, called InfoDiscoverer, which applies the distilled structure to identify sets of article pages. InfoDiscoverer also employs the entropy information to analyze the information measures of article sets and to extract informative content blocks from these sets. Our result is useful for search engines, information agents, and crawlers to index, extract, and navigate significant information from a Web site. Experiments on several real news Web sites show that the precision and the recall of our approaches are much superior to those obtained by conventional methods in mining the informative structures of news Web sites. On the average, the augmented LAMIS leads to prominent performance improvement and increases the precision by a factor ranging from 122 to 257 percent when the desired recall falls between 0.5 and 1. In comparison with manual heuristics, the precision and the recall of InfoDiscoverer are greater than 0.956. Hung-Yu Kao, Shian-Hua Lin, Jan-Ming Ho, Ming-Syan Chen |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2003 | Clustering for Web Information Hierarchy MiningabstractBenefiting from the growth of techniques of dynamic page generation, the amount and the complexity of Web pages increase explosively. The structures of Web pages which are dynamically generated by the same templates are thus similar to one another and are usually assembled by a set of fundamental information clusters These neighboring information clusters usually represent the similar semantics and form a larger cluster with the more generalized information. The hierarchical structure generated by information clusters in a bottom-up manner is called the information hierarchy of a page. We study the problem of mining the information hierarchies of pages in Web sites to recognize the information distribution of pages within the multilevel, multigranularity configurations. Explicitly, we propose an information clustering system that applies a top-down information centroid searching algorithm and a multigranularity centroid converging process on the document object model (DOM) trees of pages to build the information hierarchies of pages. Experiments on several real news Web sites show the high precision and recall rates of the proposed method on determining information clusters of pages and also validate its practical applicability to real Web sites. Hung-Yu Kao, Jan-Ming Ho, Ming-Syan Chen |
Web Intelligence | 1 |
| 2002 | Entropy-based link analysis for mining web informative structuresabstractIn this paper, we study the problem of mining the informative structure of a news Web site which consists of thousands of hyperlinked documents. We define the informative structure of a news Web site as a set of index pages (or referred to as TOC, i.e., table of contents, pages) and a set of article pages linked by TOC pages through informative links. It is noted that the Hyperlink Induced Topics Search (HITS) algorithm has been employed to provide a solution to analyzing authorities and hubs of pages. However, most of the content sites tend to contain some extra hyperlinks, such as navigation panels, advertisements and banners, so as to increase the add-on values of their Web pages. Therefore, due to the structure induced by these extra hyperlinks, HITS is found to be insufficient to provide a good precision in solving the problem. To remedy this, we develop an algorithm to utilize entropy-based Link Analysis on Mining Web Informative Structures. This algorithm is referred to as LAMIS. The key idea of LAMIS is to utilize information entropy for representing the knowledge that corresponds to the amount of information in a link or a page in the link analysis. Experiments on several real news Web sites show that the precision and the recall of LAMIS are much superior to those obtained by heuristic methods and conventional ink analysis methods. Hung-Yu Kao, Ming-Syan Chen, Shian-Hua Lin, Jan-Ming Ho |
CIKM | 1 |